speaker-align
Attach speaker labels to transcript segments by temporal overlap.
Diarization tells you when each person spoke. ASR tells you what was said and when. Neither tells you who said which line — that join is this package.
from speaker_align import get_diarizer, align_speakers_to_transcript
turns = get_diarizer("pyannote", auth_token=TOKEN).diarize("meeting.wav")
segments = align_speakers_to_transcript(turns.segments, whisper_segments)
# [{"start": 0.0, "end": 4.2, "text": "...", "speaker": "SPEAKER_00"}, ...]
Two halves, independently usable
| needs | use it when | |
|---|---|---|
align_speakers_to_transcript |
nothing but stdlib | you already have speaker turns |
get_diarizer(...) |
pip install 'speaker-align[pyannote]' |
you need to produce them from audio |
Importing the alignment half never imports a backend, so a caller that gets speaker turns elsewhere — a vendor API, a different diarizer, a human pass — installs no torch and no pyannote. That split is enforced by a test.
Install
pip install speaker-align # alignment only
pip install 'speaker-align[pyannote]' # + local diarization
pyannote needs a Hugging Face token for
speaker-diarization-3.1.
Pass it as auth_token=, or set one of SPEAKER_ALIGN_AUTH_TOKEN,
PYANNOTE_AUTH_TOKEN, HUGGINGFACE_TOKEN. Prefer passing it explicitly when
embedding this in an application — it keeps that application's config in one
place instead of splitting it between a config object and ambient environment.
How a label is decided
Each transcript segment goes to whichever speaker turn overlaps it most, but only when the winning overlap is convincing:
- at least
toleranceseconds (default0.5), or - any overlap at all, when the segment is shorter than
tolerance * 2
The second rule matters: a 0.3s interjection can never accumulate 0.5s of overlap, so judging it by the same absolute bar would leave every short segment unlabelled.
When neither holds, the segment gets UNKNOWN_SPEAKER ("") instead of a guess.
A wrong attribution is worse than an absent one — downstream, a multicam
editor may cut to the wrong camera because of it, and a missing label is visible
where a confident wrong one is not.
API
align_speakers_to_transcript(speaker_segments, transcript_segments, tolerance=0.5) -> list[dict]
align_segments_json(speaker_segments, transcript_segments_json, tolerance=0.5) -> str
speaker_turn_count(aligned_segments) -> int
get_diarizer(backend="pyannote", auth_token=None) -> BaseDiarizer
Transcript segments are plain dicts with start / end in seconds; every other
key is carried through untouched. Inputs are never mutated — you get copies.
align_segments_json is for callers that persist segments as a JSON column.
speaker_turn_count is a cheap downstream signal: a high count means dialogue
(cut between angles), a low one means monologue (hold, or cut to b-roll).
Adding a backend
Subclass BaseDiarizer, return a DiarizationResult, and register it in
get_diarizer. The alignment half neither knows nor cares which backend ran.
Origin
Extracted from reel-scout's diarization module so that other tools can share it. Behaviour is identical to the original: the extraction was verified against it across 2,018 generated cases plus boundary inputs (zero-length segments, segments straddling a turn boundary, missing timing keys, empty inputs) with no divergence.
Sibling package: whisper-guard, which filters hallucinations out of the transcript this one labels.
License
MIT
Metadata
Release files for speaker-align 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| speaker_align-0.1.0.tar.gz | 8.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| speaker_align-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 16.8 kB
Release files / speaker_align-0.1.0.tar.gz
| Download URL | speaker_align-0.1.0.tar.gz |
|---|---|
| Size | 8.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
fcd935c56e302ab0094dd4f39e176af9b99881288297816ebeb8d2aed1f3c0de
|
|
BLAKE2b-256 checksum How to use checksums |
a37e41cbaf0e6a57ed04c1d50ef39290c87c16f0911c03b6c077484a6a486db2
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.13
|
Release files / speaker_align-0.1.0-py3-none-any.whl
| Download URL | speaker_align-0.1.0-py3-none-any.whl |
|---|---|
| Size | 8.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
72776a53817f212bc7fe76cd11164b6d10ca3bcb9ecce69801eeca3486a6698f
|
|
BLAKE2b-256 checksum How to use checksums |
105034e4831d0809ed581bae505e687257d154f3ed8fd7eade5b63997caec194
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.13
|