This release has been yanked by its maintainers, and will be ignored by installers, except when explicitly specified.
Consider using release 0.1.2 instead.
Dharma Transcription Pipeline
A headless multilingual transcription pipeline designed for Buddhist dharma teachings. Handles mixed-language audio (Tibetan, Sanskrit, English, Japanese) with per-language forced alignment, speaker diarization, a Tibetan second-pass using dharma-trained models, and optional LLM post-correction.
Privacy-first: all processing runs locally. Sacred content never touches a cloud unless you explicitly configure LLM correction with a cloud API.
Architecture — 7 Stages
Audio/Video → Ingest → WhisperX → Alignment → Diarization → Tibetan 2nd-Pass → LLM Correction → Output
| Stage | Purpose | Model | GPU Memory |
|---|---|---|---|
| 1. Ingest | ffmpeg extract to 16kHz mono WAV, checksum idempotency | ffmpeg | — |
| 2. Transcription | Primary ASR with auto language detection | WhisperX large-v3 (int8) | ~4 GB |
| 3. Alignment | Word-level timestamps per language | wav2vec2 (per-language) | ~1–2 GB |
| 4. Diarization | Speaker identification | pyannote diarization-community-1 | ~1 GB |
| 5. Tibetan 2nd-Pass | Re-transcribe Tibetan segments with dharma-trained model | OpenPecha op-whisper_small-ft-v2 | ~1 GB |
| 6. LLM Correction | Post-correction with dharma domain knowledge | Any OpenAI-compatible API | (cloud or local) |
| 7. Output | JSON, SRT, VTT, TXT, review queue | — | — |
Stages run serially with GPU memory flushing between each — designed for 6 GB consumer GPUs.
Setup
Prerequisites
- Python 3.10+
- ffmpeg + ffprobe (system install)
- NVIDIA GPU with CUDA (optional —
--device cpuworks for all stages) - HuggingFace token (for diarization only — transcription works without it)
Installation
# Clone
git clone https://github.com/guan-tends/dharma-transcribe.git
cd dharma-transcribe
# Create virtual environment
python3 -m venv venv
source venv/bin/activate
# Install with pip
pip install -e ".[dev]"
# Or install from PyPI (when published)
pip install dharma-transcribe[dev]
GPU Setup (optional but recommended)
# Install PyTorch with CUDA support (adjust cuXXX for your CUDA version)
pip install torch torchaudio --index-url https://download.pytorch.org/whl/cu128
HuggingFace Token (for diarization)
The pipeline needs a HuggingFace token to download the pyannote diarization model. Without it, diarization is skipped — transcription still completes normally.
# Get a token: https://huggingface.co/settings/tokens
# Also accept the model license: https://huggingface.co/pyannote/speaker-diarization-community-1
export HF_TOKEN=hf_your_token_here
LLM Correction Configuration (optional)
Stage 6 uses any OpenAI-compatible API. Configure via environment variables:
# Cloud API example (Synthetic.new)
export DHARMA_LLM_API_URL=https://api.example.com/v1
export DHARMA_LLM_API_KEY=your-key
export DHARMA_LLM_MODEL=hf:openai/gpt-oss-120b
# Local model example (Ollama)
ollama pull qwen3.5:9b
export DHARMA_LLM_API_URL=http://localhost:11434/v1
export DHARMA_LLM_API_KEY=ollama
export DHARMA_LLM_MODEL=qwen3.5:9b
# vLLM example (local GPU)
# Start vLLM server: vllm serve Qwen/Qwen3.5-9B
export DHARMA_LLM_API_URL=http://localhost:8000/v1
export DHARMA_LLM_API_KEY=none
export DHARMA_LLM_MODEL=Qwen/Qwen3.5-9B
Or skip LLM correction entirely: --skip-llm
Usage
# Single file
dharma-transcribe /path/to/teaching.mp4
# Directory (batch — finds all audio/video recursively)
dharma-transcribe /path/to/recordings/
# Skip LLM correction (ASR only, faster)
dharma-transcribe /path/to/teaching.mp4 --skip-llm
# CPU-only mode (no GPU required)
dharma-transcribe /path/to/teaching.mp4 --device cpu --skip-llm
# With explicit HF token
dharma-transcribe /path/to/teaching.mp4 --hf-token $HF_TOKEN
CLI Flags
| Flag | Default | Description |
|---|---|---|
input (positional) |
— | File or directory to process |
--source-dir |
— | Default source directory |
--skip-llm |
off | Skip LLM correction stage |
--hf-token |
$HF_TOKEN |
HuggingFace token for diarization |
--device |
cuda |
Compute device: cuda or cpu |
Environment Variables
See .env.example for the full list. Key variables:
| Variable | Default | Description |
|---|---|---|
DHARMA_LLM_API_URL |
(empty) | OpenAI-compatible API endpoint |
DHARMA_LLM_API_KEY |
(empty) | API key for LLM correction |
DHARMA_LLM_MODEL |
(empty) | Model name for LLM correction |
HF_TOKEN |
(empty) | HuggingFace token for diarization |
DHARMA_DEVICE |
cuda |
Compute device |
DHARMA_WHISPER_MODEL |
large-v3 |
WhisperX model size |
DHARMA_OUTPUT_DIR |
./output |
Output directory |
DHARMA_TIBETAN_MODEL |
openpecha/op-whisper_small-ft-v2 |
Tibetan second-pass model |
Output Formats
Each processed file generates:
- JSON — full structured transcript with word-level timestamps, speaker labels, confidence scores
- SRT — subtitle file with speaker labels
- VTT — WebVTT subtitles with speaker tags
- TXT — plain text reading copy
- Review queue — JSON of low-confidence segments for manual review
Outputs are written to output/{json,srt,vtt,txt,review}/.
Corrections Dictionary
output/corrections/corrections.json — case-insensitive string replacement applied before LLM correction. Grows from manual review.
{
"corrections": [
{"pattern": "bodichita", "replacement": "bodhicitta"},
{"pattern": "shun yata", "replacement": "shunyata"}
]
}
Idempotency
The manifest (output/manifest.json) tracks processed files. Re-running the pipeline skips completed files. Delete the manifest to reprocess everything.
VRAM Management
Designed for consumer GPUs (6 GB+). Stages run serially with gc.collect() + torch.cuda.empty_cache() between each. No CPU fallback needed — the GPU is flushed fully before loading the next model.
Why This Exists
Most transcription tools handle single languages. Dharma teachings commonly mix Tibetan, Sanskrit, English, and sometimes Japanese in a single recording. This pipeline:
- Detects language per segment — not per file
- Aligns each language separately — wav2vec2 models for bo, sa, en, ja
- Re-transcribes Tibetan — OpenPecha's model (trained on Garchen Rinpoche's teachings) often outperforms WhisperX on Tibetan
- Corrects with dharma knowledge — LLM post-correction knows bodhicitta from bodichita
License
MIT — see LICENSE.
Support
If this pipeline helps preserve dharma teachings, consider supporting its continued development:
- Solana:
Eu8wQcW68TKMs1a6eqzZu8znzU52QLqQugAMG8uCD6y6 - EVM (Ethereum / Base / Arbitrum / Optimism / Polygon):
0x2733ff7c865C56d565a99BE1DC11B81cc76850A5 - XRP Ledger:
r4X6e7McAQj7e8vBCeued1RYu4mCJrREDG
Release files for dharma-transcribe 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| dharma_transcribe-0.1.0.tar.gz | 37.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| dharma_transcribe-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 61.0 kB
Release files / dharma_transcribe-0.1.0.tar.gz
| Download URL | dharma_transcribe-0.1.0.tar.gz |
|---|---|
| Size | 37.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
e50f683ccb5448401e509909947d9f9594e217d8ba4d0cedb3f478c3d7a53036
|
|
BLAKE2b-256 checksum How to use checksums |
ce220119816896c21f3f29b8f039cce9392a415a0d72256b49c9ac889691375f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 24, 2026.
Transparency logRelease files / dharma_transcribe-0.1.0-py3-none-any.whl
| Download URL | dharma_transcribe-0.1.0-py3-none-any.whl |
|---|---|
| Size | 23.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
208c5d7725543ae5c614af7544d8cd97e9d925b9e39f692a5b09ced93a78265f
|
|
BLAKE2b-256 checksum How to use checksums |
0cae4219618c617b21706821ec236e4f625c71754dcd988a12f77a7041b46177
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 24, 2026.
Transparency log