Skip to main content

localcaption

Offline Whisper transcription for YouTube and local files. Writes SRT, VTT, and JSON. No API key.

Local, offline Whisper transcription for YouTube, Vimeo, Twitch, Twitter/X, and 1000+ other sites via yt-dlp, plus any video or audio file on disk. Paste a URL or a path; get .txt, .srt, .vtt, and .json without an API key and without uploading audio to the cloud. Default engine is whisper.cpp; faster-whisper is optional.

PyPI version Python versions License: MIT

CI Release Hatch project

GitHub stars Last commit Open issues PRs welcome

pipx install localcaption && localcaption doctor --fix

localcaption is a tiny orchestrator over three battle-tested tools:

Stage Tool
Download best audio yt-dlp (YouTube, Vimeo, Twitch, 1000+ sites)
Re-encode to 16 kHz mono WAV ffmpeg
Transcribe locally whisper.cpp (default) or faster-whisper

Nothing is uploaded to a third-party service. No OpenAI / Google / DeepL keys required. Unlike pasting a clip into ChatGPT or calling the Whisper API, transcription stays on your laptop after the download.

Pipeline overview

Why localcaption

  • Offline Whisper. Audio is transcribed on your machine. No API key, no account.
  • URLs and local files. Any URL yt-dlp supports, plus local .mp4 / .wav / .mp3 / similar.
  • Captions you can use. One run writes .txt, .srt, .vtt, and .json.
  • Batch. --batch urls.txt walks a list of URLs or files.
  • Chapters. YouTube chapter markers become .chapters.json and .chaptered.md.
  • Search. localcaption search <term> greps past transcripts with timestamps.
  • Local summaries. --summary talks to a local Ollama, not a hosted LLM.
  • Two backends. whisper.cpp by default; pip install 'localcaption[faster]' for faster-whisper.
  • doctor --fix. Installs missing ffmpeg/cmake, builds whisper.cpp, and downloads the default small.en model.

Who this is for

  • Podcasters and video folks who need SRT/VTT without uploading episodes.
  • Researchers transcribing interviews or lectures they cannot send to a cloud API.
  • Anyone who wants a YouTube transcript without logging into Google or pasting audio into ChatGPT.

Install

Prerequisites

  • Python 3.10+
  • git, ffmpeg, cmake on your $PATH (macOS: brew install ffmpeg cmake)

Recommended: pipx (one line)

The most Pythonic install. pipx creates an isolated virtualenv for localcaption and drops the console script on your $PATH, so you can run localcaption <url-or-file> from anywhere without polluting your system Python.

pipx install localcaption

The first time you run localcaption <url-or-file> it will tell you it can't find whisper.cpp. The fastest way to set it up is to let localcaption do it itself: clone, build, and download the default model in one shot:

localcaption doctor --fix          # ~2 min on an M-series Mac

doctor --fix is idempotent and end-to-end: it installs missing system tools (ffmpeg/cmake via brew/apt), clones + builds whisper.cpp at the canonical XDG location, downloads the default model, and re-runs the diagnostics to confirm everything works. Pick a faster model with --model tiny.en.

Prefer to do it yourself? Two equivalent options:

# Option A: bootstrap script (also installs pipx + the localcaption package):
curl -fsSL https://raw.githubusercontent.com/jatinkrmalik/localcaption/main/scripts/install.sh | bash

# Option B: DIY, anywhere you like:
git clone https://github.com/ggerganov/whisper.cpp /path/to/whisper.cpp
cd /path/to/whisper.cpp && cmake -B build && cmake --build build -j --config Release
bash models/download-ggml-model.sh small.en
export LOCALCAPTION_WHISPER_DIR=/path/to/whisper.cpp   # add to your shell rc

💡 The install.sh bootstrap is just pipx install localcaption followed by localcaption doctor --fix, same logic, single source of truth. Override the default model with WHISPER_MODEL=tiny.en bash install.sh.

After install, verify everything is wired up:

localcaption doctor                # read-only diagnostic
localcaption doctor --fix          # diagnostic + auto-repair anything missing

Uninstall

To completely remove localcaption and everything it installed (the binary, whisper.cpp build, and ggml models (about 500 MB total):

# pipx + whisper.cpp + models, with confirmation prompts:
curl -fsSL https://raw.githubusercontent.com/jatinkrmalik/localcaption/main/scripts/uninstall.sh | bash

# Or, if you cloned the repo:
bash scripts/uninstall.sh

Useful flags: --dry-run (preview), --yes (skip prompts), --keep-models (uninstall the binary but keep the ~500 MB whisper.cpp + models cache for next time).

Sample output:

localcaption 0.4.0

System tools:
  ✅ python  (3.12.3)
  ✅ ffmpeg  (/opt/homebrew/bin/ffmpeg)
  ✅ cmake   (/opt/homebrew/bin/cmake)
  ✅ git     (/opt/homebrew/bin/git)

Python dependencies:
  ✅ yt-dlp  (2025.10.14)

whisper.cpp:
  searching: /Users/you/.local/share/localcaption/whisper.cpp
  ✅ directory exists
  ✅ binary built  (.../build/bin/whisper-cli)
  ✅ models present  (ggml-small.en.bin)

All checks passed. You're good to go: localcaption <url-or-file>

If anything is missing, re-run with --fix and localcaption will install the missing system deps (via brew/apt), clone+build whisper.cpp, and download the default model, then re-verify:

localcaption doctor --fix                      # repair everything
localcaption doctor --fix --model tiny.en      # …with a faster/smaller model

Dev install (contributors)

If you're hacking on localcaption itself, install editable from a clone:

git clone https://github.com/jatinkrmalik/localcaption
cd localcaption
./scripts/setup.sh           # creates .venv, pip install -e .[dev], clones+builds whisper.cpp HERE
source .venv/bin/activate
pytest                        # the suite should pass

The dev setup keeps whisper.cpp/ inside the repo (so you can poke at it), and editable-installs the package so source edits take effect immediately.

Usage

CLI

# YouTube
localcaption "https://www.youtube.com/watch?v=dQw4w9WgXcQ"

# Vimeo, Twitch, Twitter/X, and 1000+ other sites work too
localcaption "https://vimeo.com/148751763"

# Local video/audio files
localcaption /path/to/video.mp4
localcaption ./recording.wav

# Batch: one URL or local path per line (# comments and blank lines ignored)
localcaption --batch urls.txt -o transcripts/ -m small.en

# Transcript + local summary (requires a running Ollama)
localcaption "https://www.youtube.com/watch?v=dQw4w9WgXcQ" --summary
localcaption ./talk.mp4 --summary --summary-model llama3.1:8b
flag default what it does
-m, --model small.en whisper model name (tiny.en, base.en, small.en, medium.en, large-v3, …)
-o, --out ./transcripts output directory
-l, --language auto ISO language code, or auto to let whisper detect it
--backend whisper-cpp transcription backend: whisper-cpp or faster-whisper. $LOCALCAPTION_BACKEND if the flag is omitted
--whisper-dir auto-detect¹ path to a built whisper.cpp checkout (whisper-cpp backend)
--keep-audio off keep the downloaded audio + intermediate WAV in <out>/.work/
--no-print off don't echo the transcript to stdout
--batch FILE off transcribe each non-empty, non-# line in FILE sequentially
--summary off after transcription, write <id>.summary.md via local Ollama
--summary-model llama3.1:8b Ollama model used by --summary
--summary-prompt built-in path to a prompt template ({transcript} is replaced if present)

¹ --whisper-dir resolution order:

  1. The explicit flag value, if given.
  2. $LOCALCAPTION_WHISPER_DIR env var.
  3. ./whisper.cpp (dev checkout).
  4. ~/.local/share/localcaption/whisper.cpp (where install.sh puts it).

Outputs <videoId>.txt, .srt, .vtt, and .json in the chosen directory. For local files, the output filename is derived from the input file's name. With --summary, also writes <videoId>.summary.md. When the source has chapter markers (typical on YouTube), also writes <videoId>.chapters.json and <videoId>.chaptered.md. The raw whisper .txt is left unchanged.

--batch FILE writes each item into <out>/<videoId>/ and skips any video whose .txt is already there, so you can re-run a list after a failure. Local paths are relative to the list file (and ~ is expanded). The process is sequential (whisper.cpp already saturates the machine). Exit 0 if everything succeeded or was skipped, 1 otherwise.

You can also invoke it as a module: python -m localcaption <url-or-file>.

faster-whisper (optional)

whisper.cpp is the default backend and needs no extra Python packages. To use faster-whisper instead (CTranslate2, typically faster on CPU/CUDA, including Windows):

pip install 'localcaption[faster]'
# pipx:
pipx inject localcaption faster-whisper

localcaption --backend faster-whisper "https://www.youtube.com/watch?v=..."
# or:
export LOCALCAPTION_BACKEND=faster-whisper

faster-whisper downloads its own CTranslate2 weights on first use; it does not read ggml files from --whisper-dir. --model names (base.en, small.en, large-v3, ...) match the usual Whisper sizes.

Summaries (optional)

If Ollama is running locally, --summary sends the .txt transcript to http://localhost:11434/api/generate and writes <id>.summary.md next to it. The built-in prompt asks for a TL;DR, key points, notable quotes, and action items.

localcaption <url-or-file> --summary
localcaption <url-or-file> --summary --summary-model mistral
localcaption <url-or-file> --summary --summary-prompt ./my_prompt.txt

If Ollama isn't reachable, localcaption logs a warning and still exits 0. The transcript files are unchanged.

Subcommands

Subcommand What it does
(default) localcaption <url-or-file> Transcribe a URL or local video/audio file.
localcaption doctor Read-only diagnostic: prereqs, whisper.cpp, available models. Useful before filing a bug.
localcaption doctor --fix Self-heal: install missing system deps, clone+build whisper.cpp, download the default model, then re-verify. Idempotent.
localcaption model list List every supported whisper model with size + install status.
localcaption model info <name> Show metadata about a single model.
localcaption model download <name> Download a model with progress bar + atomic writes.
localcaption model rm <name> Remove an installed model to free disk space.
localcaption search <term> Grep previously transcribed videos. Ranked matches with timestamps.

Search past transcripts

Each successful transcription upserts one JSON line in ~/.local/share/localcaption/index.jsonl (id, url, title, duration, chapters, transcript). Re-running the same id replaces that row. Override the path with LOCALCAPTION_INDEX_PATH.

localcaption search "install"
# vid123  02:30  First, let's install pip
#         Lecture on Python tooling

Search is a case-insensitive substring. Hits are ranked by how often the term appears (title matches get a small boost). Timestamps come from the sibling whisper .json or .srt when those files are still next to the .txt.

Managing models

localcaption defaults to small.en (~466 MB), downloaded by doctor --fix or on first use. For a faster run use --model tiny.en; for non-English audio, pick a multilingual model. If the model isn't already installed, you'll be prompted to download it:

$ localcaption --model small.en "https://www.youtube.com/watch?v=..."

Model 'small.en' is not installed (~466 MB).
  Download it now? [Y/n] y
  small.en       [████████████████████░░░░░░░░░░░░░░░░] 290.0/466.0 MB · 18.4 MB/s · ETA 9s

Or download/manage models explicitly:

localcaption model list                  # see what's available
localcaption model info small.en         # check size before committing
localcaption model download small.en     # ~466 MB, ~25 sec on a fast connection
localcaption model rm large-v3           # free 3 GB after experimenting

For scripted/CI use, pass --auto-download to skip the prompt:

localcaption --model small.en --auto-download "https://www.youtube.com/..."

Quick model picker:

Model Size Best for
tiny.en 75 MB Fast fallback, English only, low-resource environments
base.en 142 MB Faster than small.en, lower accuracy
small.en 466 MB Install default, English, accuracy/speed balance
medium.en 1.5 GB High accuracy English, ~3× slower than small.en
large-v3 3.0 GB Best accuracy, multilingual, slow
large-v3-turbo 1.6 GB Near-large quality at ~half the size, great compromise

Models without the .en suffix are multilingual (required for non-English audio).

Python API

from pathlib import Path
from localcaption.pipeline import transcribe_url

result = transcribe_url(
    "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
    out_dir=Path("transcripts"),
    whisper_dir=Path("whisper.cpp"),
    model="small.en",
    summary=True,  # optional; writes .summary.md via local Ollama
)
print(result.transcripts.txt.read_text())

# faster-whisper (pip install 'localcaption[faster]') does not need whisper_dir:
# transcribe_url(url, out_dir=Path("transcripts"), backend="faster-whisper")

Batch from Python:

from pathlib import Path
from localcaption.batch import read_url_list, transcribe_urls

result = transcribe_urls(
    read_url_list(Path("urls.txt")),
    out_dir=Path("transcripts"),
    whisper_dir=Path("whisper.cpp"),
    model="small.en",
)
print(result.summary())

Architecture

localcaption is intentionally tiny: an orchestrator (pipeline.py) drives three single-responsibility stages, each wrapping one external tool. The transcribe stage is a small Backend protocol; whisper.cpp is the default implementation and faster-whisper is an optional extra. Swapping backends does not touch download.py or audio.py.

Module map

Module architecture

Layer Files Responsibility
Entry points cli.py, __main__.py argparse, exit codes, stdout formatting
Orchestration pipeline.py, batch.py public Python API: transcribe_url(...), transcribe_urls(...)
Pipeline stages download.py, audio.py, whisper.py, backends/, summary.py download, re-encode, transcribe (pluggable), optional Ollama summary
Chapters & search chapters.py, index.py YouTube chapter sidecars + JSONL search index
Support errors.py, _logging.py exception hierarchy, tiny logger

Runtime sequence

End-to-end call flow for a single localcaption <url> invocation, including the subprocess hops to yt-dlp, ffmpeg, and whisper.cpp. The intermediate .work/ directory is cleaned up at the end unless --keep-audio is passed.

Sequence diagram

Diagrams live in docs/diagrams/ as Mermaid .mmd source files alongside the rendered PNGs. Regenerate with:

mmdc -i docs/diagrams/<name>.mmd -o docs/diagrams/<name>.png \
  -t default -b white --width 1600 --scale 2

Benchmarks

Wall-clock times for the complete pipeline (yt-dlp download → ffmpeg re-encode → whisper.cpp transcription), measured with base.en (the previous default). These have not been re-run on small.en; expect transcription to take longer. Numbers will vary with your network speed and CPU/GPU; treat them as order-of-magnitude reference, not a competitive benchmark.

Video Length Wall-clock Speed vs. realtime Hardware
TED-Ed: How does your immune system work? 5:23 7.5 s ~43× MacBook Pro M4 Pro, 48 GB
3Blue1Brown: But what is a Neural Network? 18:40 19.3 s ~58× MacBook Pro M4 Pro, 48 GB
Hasan Minhaj × Neil deGrasse Tyson: Why AI is Overrated 54:17 49.8 s ~65× MacBook Pro M4 Pro, 48 GB
Reproduce
# Apple Silicon, macOS, whisper.cpp built with Metal,
# model: ggml-base.en (matches the table above; not the current default),
# language: auto, no other heavy processes.

time localcaption --model base.en --no-print -o /tmp/lc-bench-1 \
  "https://www.youtube.com/watch?v=PSRJfaAYkW4"

time localcaption --model base.en --no-print -o /tmp/lc-bench-2 \
  "https://www.youtube.com/watch?v=aircAruvnKk"

time localcaption --model base.en --no-print -o /tmp/lc-bench-3 \
  "https://www.youtube.com/watch?v=BYizgB2FcAQ"

If you'd like to contribute numbers from a different machine (Linux + CUDA, Windows + WSL, x86 macOS, etc.), open a PR adding a row above with your hardware in the Hardware column.

Notes

  • Bigger models = better quality but slower. small.en is the default; use --model tiny.en when you want speed over accuracy.
  • Apple Silicon: whisper.cpp's CMake build uses Metal automatically, you'll see ggml_metal_init in the logs.
  • The pipeline accepts any URL yt-dlp supports (Vimeo, Twitch VODs, Twitter/X, podcast pages, and 1000+ more), not just YouTube.
  • If you hit HTTP 403 Forbidden, your yt-dlp is probably stale. pip install -U yt-dlp usually fixes it.

Roadmap

The roadmap lives on GitHub Issues so it's easy to track, comment on, and contribute to:

👉 Open roadmap items

A snapshot of what's planned (click through for full descriptions, acceptance criteria, and discussion):

# Item Labels
#7 localcaption model {list,download,rm,info} subcommand shipped in v0.2.0
#2 Batch mode (--batch urls.txt) shipped in v0.4.0
#3 Local auto-summary via Ollama (--summary) shipped in v0.4.0
#4 Speaker diarization with pyannote.audio (--diarize) stretch, help wanted
#5 YouTube chapters & grep-able search index shipped in v0.4.0
#6 Pluggable transcription backends (faster-whisper / MLX) faster-whisper shipped in v0.4.0
#1 Switch default model from base.en to small.en shipped in v0.4.0

Have an idea? Open a feature request, or jump into Discussions if you want to chat about it first.

FAQ

Does it need an OpenAI API key? No. Whisper runs locally via whisper.cpp or faster-whisper.

Does audio leave my machine? No, except the download of a URL you asked for. Transcription and optional Ollama summaries stay on localhost.

YouTube only? No. Any site yt-dlp supports (Vimeo, Twitch, Twitter/X, podcasts, and many more), plus local video and audio files.

Does it write SRT and VTT? Yes. Each run writes .txt, .srt, .vtt, and .json.

Windows? macOS and Linux are the supported platforms (doctor --fix uses Homebrew or apt). Native Windows is not supported. WSL is the realistic path if you are on Windows.

Related projects

localcaption deliberately stays tiny. If you want more, check out:

  • whishper: full web UI for local transcription with translation and editing.
  • transcribe-anything: multi-backend, Mac-arm optimised, supports URLs.
  • WhisperX: word-level timestamps and diarisation on top of openai-whisper.

Contributing

Pull requests welcome! See CONTRIBUTING.md. By participating you agree to abide by our Code of Conduct.

License

MIT.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

localcaption-0.4.1.tar.gz (676.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

localcaption-0.4.1-py3-none-any.whl (48.6 kB view details)

Uploaded Python 3

File details

Details for the file localcaption-0.4.1.tar.gz.

File metadata

  • Download URL: localcaption-0.4.1.tar.gz
  • Upload date:
  • Size: 676.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for localcaption-0.4.1.tar.gz
Algorithm Hash digest
SHA256 1877b0632fcd9788d8eb2e561788b956277808157dc04a9699ac4c1f33f4e1a1
MD5 2077206728b3239fee7c788fa9409b9b
BLAKE2b-256 ef4914bd186e4b065352e19d8114a7b253aa217b0bd7f0f99ef9f19025a53a7d

See more details on using hashes here.

Provenance

The following attestation bundles were made for localcaption-0.4.1.tar.gz:

Publisher: release.yml on jatinkrmalik/localcaption

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file localcaption-0.4.1-py3-none-any.whl.

File metadata

  • Download URL: localcaption-0.4.1-py3-none-any.whl
  • Upload date:
  • Size: 48.6 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for localcaption-0.4.1-py3-none-any.whl
Algorithm Hash digest
SHA256 dbd18c1aec7371c69aecd979b1c05cd93f5b2d81e074d57a4a5995a06e457833
MD5 4973d91eee80f85090a3f959262364c5
BLAKE2b-256 3eeb42e6574154c3abad394bfc40f747e77639e46335029c0f71b82e0ddfb15a

See more details on using hashes here.

Provenance

The following attestation bundles were made for localcaption-0.4.1-py3-none-any.whl:

Publisher: release.yml on jatinkrmalik/localcaption

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

This release

0.4.1 This release

2 files

0.4.0

2 files

0.3.0

2 files

0.2.0

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page