transcribe
Drop-in CLI transcription. Point it at an audio/video file (or a folder), get a transcript next to it. Groq by default, local Whisper opt-in (planned).
transcribe podcast.mp3 # -> podcast.txt
transcribe ./voice-notes/ # -> .txt next to every clip
transcribe meeting.mp4 -t srt # -> meeting.srt
transcribe lecture.m4a -t md --timestamps # -> lecture.md with [mm:ss]
transcribe ./media/ -r -t json -o ./transcripts/ # recursive, JSON, one output dir
transcribe interview.mp3 --diarize # label speakers: "Speaker 1: ..."
Install
pipx install aria-transcribe
# or
uv tool install aria-transcribe
From source (dev): pipx install -e . · uv tool install -e .
Optional extras:
[local]— addsfaster-whisperfor offline transcription (Phase 3, not built yet)[keyring]— opt-in system keyring for API key storage[diarize]— addspyannote.audio+torchfor speaker labels (--diarize); ~2 GB, see Speaker diarization[all]— everything
Prerequisites
- Python 3.11+
- ffmpeg + ffprobe on
PATH- Windows:
winget install Gyan.FFmpeg·scoop install ffmpeg·choco install ffmpeg - macOS:
brew install ffmpeg - Linux:
pacman -S ffmpeg·apt install ffmpeg·dnf install ffmpeg
- Windows:
- GROQ_API_KEY for cloud transcription — grab one from console.groq.com/keys
Setup
Set the API key one of three ways (precedence: CLI > env > config):
# 1. env var (recommended — survives across sessions)
export GROQ_API_KEY="gsk_..." # bash/zsh
$env:GROQ_API_KEY = "gsk_..." # PowerShell
# 2. one-off
transcribe clip.mp3 --api-key gsk_...
# 3. config file (see below) — handy on machines you fully control
Config file
First run writes a commented template to:
$XDG_CONFIG_HOME/transcribe/config.tomlif set- otherwise
~/.config/transcribe/config.toml(works on Windows too)
Edit the values you want; everything is optional. CLI flags and env vars always win.
[defaults]
backend = "groq"
file_type = "txt"
timestamps = false
# output_dir = "~/transcripts"
# prompt_file = "~/path/to/custom-glossary.txt"
[groq]
# api_key = "gsk_..." # prefer GROQ_API_KEY env var
model = "whisper-large-v3-turbo"
upload_limit_mb = 24
[diarize]
enabled = false # default for --diarize
engine = "pyannote"
model = "pyannote/speaker-diarization-3.1"
device = "auto" # auto | cpu | cuda
# hf_token = "hf_..." # prefer HF_TOKEN env var; license must be accepted
[ui]
accent_color = "#ff2e88"
Glossary trick (proper-noun accuracy)
Drop a ~/.config/transcribe/glossary.txt with your proper nouns:
Nvikelo Nyathi, AriaCode Tech, Ableton, OpenClaw, Zellij,
vanity arias, Bulawayo, Ndebele, Shaka Zulu, Whisper, Groq
It auto-loads as the default --prompt so Whisper spells your names right
across every job. Trimmed to the trailing ~224 tokens (Whisper's prompt
budget) if longer. Override per-job with --prompt "..." or
--initial-prompt-file path.
Speaker diarization
Label who said what. Whisper only transcribes — diarization is a separate
local pass (pyannote.audio) joined
onto the transcript by timestamp, so every output format gains Speaker N labels.
uv tool install -e '.[diarize]' # one-time: adds pyannote + torch (~2 GB)
export HF_TOKEN="hf_..." # accept the model license first (link below)
transcribe interview.mp3 --diarize # -> "Speaker 1: ...", "Speaker 2: ..."
transcribe call.m4a --diarize --speakers 2 # exact speaker-count hint
transcribe panel.mp4 --diarize --min-speakers 3 --max-speakers 5
transcribe chat.wav --diarize --speaker-names "Alice,Bob" # rename Speaker 1,2
One-time setup:
- Install the extra:
uv tool install -e '.[diarize]'. - Create a Hugging Face token and accept the license for BOTH gated models
the pipeline pulls:
pyannote/speaker-diarization-3.1andpyannote/segmentation-3.0. (Accepting only the first gives a 403 on the second at model-load time.) - Provide the token via
HF_TOKEN,--hf-token, or[diarize] hf_tokenin config.
Notes:
- Word-level assignment. A Whisper segment spanning a speaker change is split, so turn boundaries land in the right place instead of mislabeling the whole line.
- Fail-hard by default. If diarization errors at runtime, the run fails (exit 6)
rather than silently writing an unlabeled transcript. Pass
--diarize-best-effortto warn-and-write the transcript without labels instead. - GPU strongly preferred for long files (
device = "cuda"; ROCm works on the RX 6800). - Missing extra → exit 3; missing/unauthorized token → exit 4 (caught as a preflight before any file is processed).
All flags
| Flag | Short | Default | Notes |
|---|---|---|---|
PATH |
— | (required) | File or directory. Directory → batch. |
--file-type |
-t |
txt |
txt · md · srt · json |
--output |
-o |
next to source | Existing dir → auto-named inside; else verbatim path. |
--backend |
-b |
groq |
groq · local (local TBD) |
--model |
-m |
per-backend default | e.g. whisper-large-v3-turbo |
--language |
-l |
auto-detect | ISO code; skips detection → slightly faster. |
--translate |
— | off | Whisper translate-to-English mode. |
--timestamps / --no-timestamps |
— | from config | mm:ss markers in txt/md only. srt/json always timestamped. |
--prompt |
-p |
none | Initial prompt to bias vocabulary. |
--initial-prompt-file |
— | none | Load prompt text from file. |
--vad / --no-vad |
— | on | Silero VAD silence stripping (local backend only). |
--diarize / --no-diarize |
— | from config | Label speakers (needs [diarize] extra + HF_TOKEN). |
--diarize-best-effort |
— | off | On diarizer runtime failure, write transcript unlabeled + warn instead of failing. |
--speakers |
— | auto | Exact speaker-count hint. Exclusive with --min/--max-speakers. |
--min-speakers |
— | auto | Lower bound on speaker count. |
--max-speakers |
— | auto | Upper bound on speaker count. |
--speaker-names |
— | none | Comma-separated names mapped to Speaker 1,2,.... |
--hf-token |
— | env / config | Hugging Face token for pyannote (highest precedence). |
--recursive |
-r |
off | Recurse into subdirectories in batch mode. |
--api-key |
— | env / config | Groq key override (highest precedence). |
--quiet |
-q |
off | Errors + final path only. Pipe-friendly. |
--verbose |
-v |
off | Sizes, timings, language detection, segment count. |
--version |
— | — | Print version, exit. |
Exit codes
| Code | Meaning |
|---|---|
| 0 | Success (batch: every file succeeded) |
| 1 | Generic / batch had failures |
| 2 | Bad usage (mutually exclusive flags, etc.) |
| 3 | Missing dependency (ffmpeg/ffprobe, or pyannote.audio when --diarize) |
| 4 | Missing or invalid Groq API key (or Hugging Face token when --diarize) |
| 5 | Input not found / unsupported file type |
| 6 | Backend, transcription, or diarization error after retries |
Output formats
- txt — plain text, paragraph-wrapped at 88 cols. With
--timestamps,[mm:ss]per segment. - md — header (title, language, duration, model, backend) + body. With
--timestamps,`[mm:ss]`markers per segment. - srt — standard SubRip subtitles; opens in any video editor.
- json — lossless dump of every segment + per-word timestamps + probabilities (+
speakerwhen--diarize). Round-trips back viaTranscriptionResult.from_dict.
With --diarize, every format gains speaker labels: txt Speaker 1: ... blocks, md **Speaker 1:** headings, srt [Speaker 1] cue prefixes, json per-segment/word speaker.
Status
- ✅ Phase 0 — skeleton (CLI parses, version, help)
- ✅ Phase 1 — Groq happy path (single short file)
- ✅ Phase 2 — all four formats + collision-safe writes + output-path resolution
- ✅ Phase 5 — batch (folder + recursive + continue-on-error)
- ✅ Phase 6 — glossary auto-load + prompt trim + README polish
- ✅ Diarization — optional
--diarizespeaker labels (pyannote, word-level, all formats) - ⬜ Phase 3 — local backend (
faster-whisper) - ⬜ Phase 4 — long-file silence chunking + timestamp stitching (only needed for files >24 MB transcoded, ~7+ hrs)
See transcribe-cli-plan.md for the full spec.
License
MIT. See LICENSE.
Metadata
Release files for aria-transcribe 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| aria_transcribe-0.1.0.tar.gz | 266.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| aria_transcribe-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 299.2 kB
Release files / aria_transcribe-0.1.0.tar.gz
| Download URL | aria_transcribe-0.1.0.tar.gz |
|---|---|
| Size | 266.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
f02389b03b7eb9edbb47702bb4e4035e36e2859c8574d5c42b97260b7c2896c0
|
|
BLAKE2b-256 checksum How to use checksums |
4f3efa9538c71c3fc58d945d5de3c1412fe93c43ac7edc7f6f4a4b7bee25a087
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.11.21 {"installer":{"name":"uv","version":"0.11.21","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"CachyOS Linux","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|
Release files / aria_transcribe-0.1.0-py3-none-any.whl
| Download URL | aria_transcribe-0.1.0-py3-none-any.whl |
|---|---|
| Size | 32.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
4f962adc35755c0e51873007bb8e2d5c8319d5ca0ca8c7ad2fa1526503066147
|
|
BLAKE2b-256 checksum How to use checksums |
23606f0ac7329f731921036c050f7b4f4e58bb0d35a55ec32f7b798e0dd22adb
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.11.21 {"installer":{"name":"uv","version":"0.11.21","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"CachyOS Linux","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|