Skip to main content

Sub-align

CI Status License: MIT Python Versions PyPI Version

Align .srt / .lrc / .txt subtitles (or generate them) to audio/video with WhisperX forced alignment — so each cue can move independently instead of only applying one global timeline shift.

Why not only a global offset?

Tools like ffsubsync typically find a constant offset (or stretch) between speech activity and subtitle “on” times. That works well for whole-track drift, but leading/trailing silence or local timing errors can still leave lines early or late.

sub-align picks a strategy from the input type, then runs WhisperX phoneme / word-level forced alignment so each line is refined against the audio:

Input Strategy
Media only Whisper ASR → word-align → split into timed cues
.txt script ASR only for search windows → forced-align original lines
.srt / .lrc Optional global offset → expand windows by --margin → forced-align

Limitations: subtitle text must roughly match spoken content (no translation). Refine can still nudge already-good cues; very short lines may be merged with close neighbors (large pauses stay alone); VAD only blocks refine starts pulled earlier into silence. There is no speaker diarization. See docs/pipeline.md § Limitations.

More detail: docs/pipeline.md · scenarios & flags: docs/usage.md

Install

Requires Python 3.10+ and ffmpeg on PATH. First run downloads WhisperX alignment models (disk/RAM).

pip install 'sub-align[align]'
# or
uv pip install 'sub-align[align]'

Extras [align], [cpu], and [gpu] all install WhisperX. Install a matching PyTorch build first when you need a specific CPU/CUDA wheel:

# CPU
uv pip install torch --index-url https://download.pytorch.org/whl/cpu
uv pip install 'sub-align[cpu]'

# CUDA (example: cu124)
uv pip install torch --index-url https://download.pytorch.org/whl/cu124
uv pip install 'sub-align[gpu]'

Development

uv venv
uv sync --group dev                 # unit tests / lint (no WhisperX)
uv sync --group dev --extra align   # full local alignment
uv run pytest
uv run ruff check src tests

CI runs the same unit tests without the align extra. Path A/B/C coverage also uses checked-in fixtures under tests/fixtures/ (timed WAVs + recorded WhisperX ASR/align JSON), replayed via mocks in tests/test_fixture_replay.py.

To regenerate those fixtures locally (needs ffmpeg; recording needs WhisperX):

# 1) Timed TTS clips (clip 1 = main coverage, clip 2 = dash dialogue)
uv run --with edge-tts --with numpy --with soundfile \
  scripts/generate_clip1_tts.py --clip 1
uv run --with edge-tts --with numpy --with soundfile \
  scripts/generate_clip1_tts.py --clip 2

# 2) Record ASR + word-align JSON for replay tests
uv run --extra align scripts/record_whisperx_fixtures.py \
  tests/fixtures/clip1_timed.wav tests/fixtures/clip2_timed.wav

Companion subtitle inputs (clip1_script.txt, clip1_drift.srt, clip2_dialogue.srt) live beside the WAVs; edit those if the cue sheet changes.

Usage

# Timed subtitles: auto global offset + per-cue refine
sub-align media.mp4 subs.srt --language zh -o out.srt

# Lyrics (LRC): same strategy as SRT
sub-align audio.wav lyrics.lrc --language en --margin 1.0

# Untimed script: ASR windows, then force-align original lines
sub-align media.mkv script.txt --language en --model small

# Skip podcast intro/outro before aligning a script
sub-align media.mp3 script.txt --language en --trim-start 13 --trim-end 5

# Known whole-track shift (skips auto-offset ASR)
sub-align media.mkv subs.srt --language en --offset 12.5

# Audio only: transcribe + word-align into an SRT
sub-align lecture.mp4 --language en -o lecture.asr.srt

Always pass --language (e.g. en, zh) or --detect-language.

See docs/usage.md for when to use --model, --margin, --offset, --fill-gaps, --trim-*, audio-only line limits, and a Whisper model size / VRAM cheat sheet.

Python API

from sub_align import align_file

align_file(
    media="a.mp4",
    subtitle="a.srt",   # omit for audio-only transcription
    output="a.aligned.srt",
    language="zh",
    device="auto",
)

How it works (short)

  1. Load media as 16 kHz mono audio (via WhisperX / ffmpeg); optional --trim-start / --trim-end.
  2. Resolve language (--language or tiny-model detection).
  3. Build search windows by input type (ASR token match for .txt; offset + margin refine for .srt/.lrc; full ASR for media-only).
  4. Run WhisperX forced alignment; remap word times onto original cues; trim overlaps; optional --fill-gaps; write .srt or .lrc.

Full diagram and tech notes: docs/pipeline.md.

Build and publish

uv build
uv publish   # requires PyPI credentials / UV_PUBLISH_TOKEN

GitHub Actions publishes on tags matching v* (see .github/workflows/publish.yml). Configure Trusted Publishing on PyPI or repository secret UV_PUBLISH_TOKEN.

License

MIT

Metadata

Release files for sub-align 0.2.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for sub-align 0.2.2
File Size Uploaded
sub_align-0.2.2.tar.gz 586.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for sub-align 0.2.2
File Interpreter ABI Platform
sub_align-0.2.2-py3-none-any.whl Python 3 none any Details

Total release size: 618.2 kB

Release files / sub_align-0.2.2.tar.gz

Download URL sub_align-0.2.2.tar.gz
Size 586.5 kB
Tags Source
SHA-256 checksum
How to use checksums
c45cb49947584b45652e775598c771dbce914814cdf38d279fe6b03e80063371
BLAKE2b-256 checksum
How to use checksums
99ed208eadc3347177c299bb6f7aed949d6d5d90a7f89be2f867746a75044963
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.10 {"installer":{"name":"uv","version":"0.12.10","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release files / sub_align-0.2.2-py3-none-any.whl

Download URL sub_align-0.2.2-py3-none-any.whl
Size 31.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
62ee5a4c7fdb1010d5925c2fdd45514c2b3b25be3d6a5c7674c9450d47fdfed2
BLAKE2b-256 checksum
How to use checksums
e5790645db46ca198828dfe25449aba29163470dc47dee643cb87ca7fca74c7a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.10 {"installer":{"name":"uv","version":"0.12.10","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release history Release notifications | RSS feed

This release

0.2.2 This release

2 release files

0.2.1

2 release files

0.2.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page