Skip to main content

speaker-align

Attach speaker labels to transcript segments by temporal overlap.

Diarization tells you when each person spoke. ASR tells you what was said and when. Neither tells you who said which line — that join is this package.

from speaker_align import get_diarizer, align_speakers_to_transcript

turns = get_diarizer("pyannote", auth_token=TOKEN).diarize("meeting.wav")
segments = align_speakers_to_transcript(turns.segments, whisper_segments)
# [{"start": 0.0, "end": 4.2, "text": "...", "speaker": "SPEAKER_00"}, ...]

Two halves, independently usable

needs use it when
align_speakers_to_transcript nothing but stdlib you already have speaker turns
get_diarizer(...) pip install 'speaker-align[pyannote]' you need to produce them from audio

Importing the alignment half never imports a backend, so a caller that gets speaker turns elsewhere — a vendor API, a different diarizer, a human pass — installs no torch and no pyannote. That split is enforced by a test.

Install

pip install speaker-align                  # alignment only
pip install 'speaker-align[pyannote]'      # + local diarization

pyannote needs a Hugging Face token for speaker-diarization-3.1. Pass it as auth_token=, or set one of SPEAKER_ALIGN_AUTH_TOKEN, PYANNOTE_AUTH_TOKEN, HUGGINGFACE_TOKEN. Prefer passing it explicitly when embedding this in an application — it keeps that application's config in one place instead of splitting it between a config object and ambient environment.

How a label is decided

Each transcript segment goes to whichever speaker turn overlaps it most, but only when the winning overlap is convincing:

  • at least tolerance seconds (default 0.5), or
  • any overlap at all, when the segment is shorter than tolerance * 2

The second rule matters: a 0.3s interjection can never accumulate 0.5s of overlap, so judging it by the same absolute bar would leave every short segment unlabelled.

When neither holds, the segment gets UNKNOWN_SPEAKER ("") instead of a guess. A wrong attribution is worse than an absent one — downstream, a multicam editor may cut to the wrong camera because of it, and a missing label is visible where a confident wrong one is not.

API

align_speakers_to_transcript(speaker_segments, transcript_segments, tolerance=0.5) -> list[dict]
align_segments_json(speaker_segments, transcript_segments_json, tolerance=0.5) -> str
speaker_turn_count(aligned_segments) -> int
get_diarizer(backend="pyannote", auth_token=None) -> BaseDiarizer

Transcript segments are plain dicts with start / end in seconds; every other key is carried through untouched. Inputs are never mutated — you get copies.

align_segments_json is for callers that persist segments as a JSON column. speaker_turn_count is a cheap downstream signal: a high count means dialogue (cut between angles), a low one means monologue (hold, or cut to b-roll).

Adding a backend

Subclass BaseDiarizer, return a DiarizationResult, and register it in get_diarizer. The alignment half neither knows nor cares which backend ran.

Origin

Extracted from reel-scout's diarization module so that other tools can share it. Behaviour is identical to the original: the extraction was verified against it across 2,018 generated cases plus boundary inputs (zero-length segments, segments straddling a turn boundary, missing timing keys, empty inputs) with no divergence.

Sibling package: whisper-guard, which filters hallucinations out of the transcript this one labels.

License

MIT

Metadata

Release files for speaker-align 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for speaker-align 0.1.0
File Size Uploaded
speaker_align-0.1.0.tar.gz 8.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for speaker-align 0.1.0
File Interpreter ABI Platform
speaker_align-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 16.8 kB

Release files / speaker_align-0.1.0.tar.gz

Download URL speaker_align-0.1.0.tar.gz
Size 8.3 kB
Tags Source
SHA-256 checksum
How to use checksums
fcd935c56e302ab0094dd4f39e176af9b99881288297816ebeb8d2aed1f3c0de
BLAKE2b-256 checksum
How to use checksums
a37e41cbaf0e6a57ed04c1d50ef39290c87c16f0911c03b6c077484a6a486db2
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.13

Release files / speaker_align-0.1.0-py3-none-any.whl

Download URL speaker_align-0.1.0-py3-none-any.whl
Size 8.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
72776a53817f212bc7fe76cd11164b6d10ca3bcb9ecce69801eeca3486a6698f
BLAKE2b-256 checksum
How to use checksums
105034e4831d0809ed581bae505e687257d154f3ed8fd7eade5b63997caec194
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.13

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page