Skip to main content

voxsolo 🎚️

The solo button for any voice in a recording. Diarize → keep one speaker bit-exact, mute everyone else (overlaps included) → NLE/DAW timeline exports → interactive HTML review player.

On a mixing console, solo plays one channel and silences the rest. voxsolo does that to people: pick a speaker, get a full-length stem where only they exist — every kept sample copied verbatim from your source.

🎧 Hear it now — no install: dhayanithi-svaromance.github.io/voxsolo — a scene from His Girl Friday (1940, public domain, famous for its overlapping rapid-fire dialogue), soloed by this tool. Switch tracks to hear each actor alone; 12.9s of cross-talk silenced. The page itself is voxsolo's own HTML output.

CI License: MIT

No established open-source tool does the middle step. Source-separation models (Demucs, SepFormer, ClearVoice) resynthesize audio — fine for karaoke, wrong for film dialogue where you must keep the original noise floor, room tone, and every bit of the recording exactly as shot. voxsolo instead copies the target speaker's samples verbatim and silences everyone else, including every region where two people talk at once. What you keep is bit-identical to the source — and the tool proves it.

┌─────────────┐   ┌──────────────────┐   ┌───────────────────────────────┐
│ any media   │ → │ pyannote         │ → │ per-speaker stems (verbatim)  │
│ .mov .wav … │   │ diarization      │   │ + overlaps silenced            │
└─────────────┘   └──────────────────┘   │ + SRT/VTT/EDL/CSV/labels      │
                                         │ + HTML review player          │
                                         │ + video with isolated audio   │
                                         └───────────────────────────────┘

Why this exists

Diarization tools (pyannote, NeMo, WhisperX) output labels — RTTM files and timestamps. Editors need audio and timelines. The gap between them is real: pyannote's own issue tracker tells people asking for per-speaker audio to script it themselves, and the only repos that do exactly this are single-file utilities. voxsolo closes that gap properly:

  • Verbatim isolation. Kept samples are copied untouched — no denoising, no normalization, no resampling, no gain change. Bit depth is preserved end to end (PCM_16/PCM_24/FLOAT in → same out).
  • Overlap handling that is actually correct. Overlap is computed with an exact sweep-line algorithm over N speakers — nested overlaps, triple-talk, self-overlapping turns, touching boundaries. Simultaneous speech is silenced by default (--keep-overlap to keep it).
  • Provable output. --verify asserts, sample by sample, that kept audio is bit-identical to the source and other speakers are digitally silent — and fails the exit code if not.
  • Click-free cuts. Half-cosine micro-fades (default 10 ms) at region edges only; --fade 0 gives bit-exact edges. Pre/post-roll handles (--pad) recover word onsets without ever bleeding into another speaker's turn.
  • Filmmaker handoff. CMX 3600 EDL, SRT/WebVTT subtitles, marker CSV, Audacity labels, RTTM — plus the isolated audio muxed back into your video with the video stream copied, not re-encoded.
  • Visual review. A self-contained HTML dashboard: color-coded speaker timeline, click-to-seek, live active-speaker highlight, and a track switcher between the full mix and each isolated stem.

Install

pip install voxsolo              # or: pip install "voxsolo[transcribe]" for subtitles with text
# ffmpeg must be on PATH:  sudo apt install ffmpeg  /  brew install ffmpeg
Install from source instead
git clone https://github.com/dhayanithi-svaromance/voxsolo
cd voxsolo && pip install .

Models. The default engine is pyannote speaker-diarization-community-1 (the current standard, best accuracy), falling back to speaker-diarization-3.1. Both are gated on Hugging Face — a free, one-time terms acceptance:

  1. Accept conditions at hf.co/pyannote/speaker-diarization-community-1 (and/or 3.1 + segmentation-3.0)
  2. huggingface-cli login (or export HF_TOKEN=...)

No token? --allow-mirrors rebuilds the 3.1 recipe from community re-uploads of the weights, pinned by SHA-256 so a tampered mirror is rejected, not trusted. It works fully offline once cached. This is opt-in by design — you're choosing to trust mirror repos instead of the official gated ones.

Use

# 1. See who speaks when (+ all timeline exports)
voxsolo diarize film_scene.mov

# 2. Isolate one speaker — everyone else and all cross-talk silenced
voxsolo isolate film_scene.mov --speaker SPEAKER_01 --verify

# 3. Everything: stems for all speakers, subtitles, EDL, HTML player
voxsolo full film_scene.mov -o scene_output/

# Diarize once, render many times (no second model run)
voxsolo diarize film_scene.mov --save-timeline tl.json
voxsolo isolate film_scene.mov --timeline tl.json --all --clips

The flags that matter most:

Flag Why you'd use it
--num-speakers N Pass it whenever you know the count. On noisy material, auto-detection can collapse two similar voices into one; pinning N prevents it.
--verify Proof, not vibes: bit-identical check + leakage check, exit 3 on failure.
--keep-overlap Keep the target's overlapped speech instead of silencing it.
--keep-background Keep room tone between turns instead of digital silence.
--fps 25 Timecode rate for EDL/marker exports.
--transcribe / --no-transcribe Faster-Whisper text in SRT/VTT/labels (pip install ".[transcribe]").
--fade 0 Bit-exact region edges (accepting possible clicks).

Outputs land next to your input (or in -o DIR): <name>_<SPEAKER>_timeline.wav stems aligned to the original timeline, <name>_<SPEAKER>_only.mov video with isolated audio, <name>.srt/.vtt/.edl, <name>_markers.csv, <name>_labels.txt (Audacity), <name>.rttm, <name>_timeline.json, <name>_player.html.

Measured, not promised

Real clips, real numbers (RTX 3050 Laptop 4 GB unless noted):

Clip Length Overlap silenced Stems Diarization time --verify
His Girl Friday (1940) office scene — hear it 75 s 12.9 s in 15 regions 30.0 s + 29.9 s 11.0 s (~7× realtime) ✅ bit-identical
1980s Tamil film scene, 44.1 kHz stereo, heavy tape hiss 68 s 6.4 s in 4 regions 26.2 s + 18.3 s ✅ bit-identical

What --verify prints — for every stem, every render:

SPEAKER_00: kept 30.02s of 75.00s (40.0%) in 15 region(s)
  verify: bit-identical=True others-silent=True length=True [OK]

bit-identical compares every kept sample against the source; others-silent asserts digital zero wherever any other speaker talks. If either fails, the exit code fails. That's the whole fidelity claim, and it's machine-checked on your file, not ours.

How it compares

voxsolo pyannote.audio WhisperX Auto-Editor ClearVoice / Demucs AudioShake (commercial)
Diarization ✅ (pyannote engine) ✅ (via pyannote) ❌ (loudness only)
Per-speaker audio out ✅ verbatim ❌ labels only ❌ labels only ✅ resynthesized ✅ resynthesized
Original noise/fidelity kept ✅ bit-identical, proven ❌ neural output ❌ neural output
Overlap silencing (N-speaker exact) detects, doesn't render
NLE/DAW exports EDL, SRT, VTT, CSV, Audacity, RTTM RTTM SRT, VTT, JSON, labels FCP7 XML, FCPXML, MLT
Video mux (stream-copy)
Visual review UI ✅ HTML player

voxsolo does not separate mixed voices: silencing an overlap removes the target's words there too. That is the intended trade — verbatim fidelity over reconstruction. When you need the words back from a mix, use a source-separation tool (ClearVoice, TIGER, Bandit-v2) and accept resynthesized audio.

Honest limitations

Everything downstream depends on diarization quality, and diarization is content-dependent. Expect degradation with heavy noise/music/reverb, band-limited or archival sources, similar-sounding voices, many speakers, or sub-second turns. Concretely, on a noisy 1980s film clip during development: auto-detect merged two speakers into one; --num-speakers 2 produced the correct split. The interval algebra and rendering are exact and tested — the model in front of them is not magic. Run diarize first, sanity-check the turn structure, listen to the result, and score against a reference (--rttm + pyannote.metrics) when you need a number.

Boundaries are also approximate: --pad (default 50 ms) recovers clipped word onsets, but only into space no other speaker occupies.

Library API

from voxsolo import load_pipeline, diarize, render, verify

loaded = load_pipeline(allow_mirrors=True)          # or token="hf_..."
timeline = diarize("analysis_16k.wav", duration=68.0, pipeline=loaded.pipeline,
                   num_speakers=2)
result = render("master.wav", timeline, "SPEAKER_01", "out.wav")
report = verify("master.wav", "out.wav", timeline, "SPEAKER_01")
assert report["bit_identical_in_kept"] and report["other_speakers_silent"]

Timeline serializes to JSON and RTTM; the interval algebra (voxsolo.intervals) is dependency-free and fully unit-tested (sweep-line overlap for any N, exact subtraction, contiguity-safe padding).

Project layout

voxsolo/
  intervals.py    exact interval algebra (merge/subtract/overlap-N) — fully tested
  pipeline.py     pyannote loading: community-1 → 3.1 → SHA-256-pinned mirrors
  diarize.py      Timeline model, JSON round-trip, RTTM export
  isolate.py      keep-region computation, verbatim rendering, verification
  audio.py        ffmpeg extraction/muxing + EDL/SRT/VTT/CSV/Audacity exports
  transcribe.py   optional Faster-Whisper per-turn transcription
  html_player.py  self-contained review dashboard
  cli.py          diarize / isolate / full / player commands
tests/            30 tests, model-free, run in CI

License

MIT (see LICENSE). The pyannote models carry their own terms — accept them on their Hugging Face pages before commercial use; pyannote.audio itself is MIT, speaker-diarization-community-1 weights are CC-BY-4.0.


If voxsolo saved you a re-record, a paid tool, or an afternoon of manual muting — a star helps the next editor find it.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

voxsolo-0.1.0.tar.gz (41.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

voxsolo-0.1.0-py3-none-any.whl (36.8 kB view details)

Uploaded Python 3

File details

Details for the file voxsolo-0.1.0.tar.gz.

File metadata

  • Download URL: voxsolo-0.1.0.tar.gz
  • Upload date:
  • Size: 41.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for voxsolo-0.1.0.tar.gz
Algorithm Hash digest
SHA256 56c5e2d3df1c25e4c18ab036b04447f1286c112d66bd2bb430f15df4ebadcac5
MD5 5c7effa00054b1e6c1ec0de1e00779d8
BLAKE2b-256 f1b99f6fc0fe1ffc110042269b101d3d1e4deb101d7a4b854a0ae6850c930d2e

See more details on using hashes here.

Provenance

The following attestation bundles were made for voxsolo-0.1.0.tar.gz:

Publisher: publish.yml on dhayanithi-svaromance/voxsolo

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file voxsolo-0.1.0-py3-none-any.whl.

File metadata

  • Download URL: voxsolo-0.1.0-py3-none-any.whl
  • Upload date:
  • Size: 36.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for voxsolo-0.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 e0fbdc4f0945725a3035ce361ccbb99d3cbebec0ca64e7bf86bce9765b16c68a
MD5 0029fa38d5ce7b9566710780c900161a
BLAKE2b-256 a0c65aaaa885e36fdbf005d94921fb254ee901c416d0508a04adb99b280ea002

See more details on using hashes here.

Provenance

The following attestation bundles were made for voxsolo-0.1.0-py3-none-any.whl:

Publisher: publish.yml on dhayanithi-svaromance/voxsolo

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page