Sub-align
Align .srt / .lrc / .txt subtitles (or generate them) to audio/video with
WhisperX forced alignment — so each cue can
move independently instead of only applying one global timeline shift.
Why not only a global offset?
Tools like ffsubsync typically find a constant offset (or stretch) between speech activity and subtitle “on” times. That works well for whole-track drift, but leading/trailing silence or local timing errors can still leave lines early or late.
sub-align picks a strategy from the input type, then runs WhisperX
phoneme / word-level forced alignment so each line is refined against the
audio:
| Input | Strategy |
|---|---|
| Media only | Whisper ASR → word-align → split into timed cues |
.txt script |
ASR only for search windows → forced-align original lines |
.srt / .lrc |
Optional global offset → expand windows by --margin → forced-align |
Limitation: subtitle text must roughly match spoken content. Alignment does not translate or correct wrong words.
More detail: docs/pipeline.md · scenarios & flags: docs/usage.md
Install
Requires Python 3.10+ and ffmpeg on PATH. First run
downloads WhisperX alignment models (disk/RAM).
pip install 'sub-align[align]'
# or
uv pip install 'sub-align[align]'
Extras [align], [cpu], and [gpu] all install WhisperX. Install a matching
PyTorch build first when you need a specific CPU/CUDA wheel:
# CPU
uv pip install torch --index-url https://download.pytorch.org/whl/cpu
uv pip install 'sub-align[cpu]'
# CUDA (example: cu124)
uv pip install torch --index-url https://download.pytorch.org/whl/cu124
uv pip install 'sub-align[gpu]'
Development
uv venv
uv sync --group dev # unit tests / lint (no WhisperX)
uv sync --group dev --extra align # full local alignment
uv run pytest
uv run ruff check src tests
Usage
# Timed subtitles: auto global offset + per-cue refine
sub-align media.mp4 subs.srt --language zh -o out.srt
# Lyrics (LRC): same strategy as SRT
sub-align audio.wav lyrics.lrc --language en --margin 1.0
# Untimed script: ASR windows, then force-align original lines
sub-align media.mkv script.txt --language en --model small
# Skip podcast intro/outro before aligning a script
sub-align media.mp3 script.txt --language en --trim-start 13 --trim-end 5
# Known whole-track shift (skips auto-offset ASR)
sub-align media.mkv subs.srt --language en --offset 12.5
# Audio only: transcribe + word-align into an SRT
sub-align lecture.mp4 --language en -o lecture.asr.srt
Always pass --language (e.g. en, zh) or --detect-language.
See docs/usage.md for when to use --model, --margin,
--offset, --fill-gaps, --trim-*, audio-only line limits, and a Whisper
model size / VRAM cheat sheet.
Python API
from sub_align import align_file
align_file(
media="a.mp4",
subtitle="a.srt", # omit for audio-only transcription
output="a.aligned.srt",
language="zh",
device="auto",
)
How it works (short)
- Load media as 16 kHz mono audio (via WhisperX / ffmpeg); optional
--trim-start/--trim-end. - Resolve language (
--languageor tiny-model detection). - Build search windows by input type (ASR token match for
.txt; offset + margin refine for.srt/.lrc; full ASR for media-only). - Run WhisperX forced alignment; remap word times onto original cues; trim
overlaps; optional
--fill-gaps; write.srtor.lrc.
Full diagram and tech notes: docs/pipeline.md.
Build and publish
uv build
uv publish # requires PyPI credentials / UV_PUBLISH_TOKEN
GitHub Actions publishes on tags matching v* (see .github/workflows/publish.yml).
Configure Trusted Publishing on PyPI or repository secret UV_PUBLISH_TOKEN.
License
MIT
Metadata
Release files for sub-align 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| sub_align-0.2.0.tar.gz | 31.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| sub_align-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 59.6 kB
Release files / sub_align-0.2.0.tar.gz
| Download URL | sub_align-0.2.0.tar.gz |
|---|---|
| Size | 31.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
c4777298556c04c41cac9a35b9b2a8ca46402f1ac5731d13d7ddf003ce7c156e
|
|
BLAKE2b-256 checksum How to use checksums |
73bf13d3aa11f81f2f0a36d6bfc3daf9cd65ac7ee14f697655cc1d6496fa496a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.12.3 {"installer":{"name":"uv","version":"0.12.3","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|
Release files / sub_align-0.2.0-py3-none-any.whl
| Download URL | sub_align-0.2.0-py3-none-any.whl |
|---|---|
| Size | 27.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
48e70f1394ee76b9edfa310a106349da1437e783d5af909cdd760a42872e6850
|
|
BLAKE2b-256 checksum How to use checksums |
6781984d9c6487259bc7ba225e8c2c339ca1fb9c29108a6489b9e854a3579d0f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.12.3 {"installer":{"name":"uv","version":"0.12.3","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|