Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

speakeronnx

speakeronnx is a speaker embedding library built on onnxruntime. It does not need torch at runtime.

The library extracts speaker embeddings from audio, computes cosine similarity between them, and verifies speaker identity. It downloads ONNX models from HuggingFace automatically.

Model collection: OpenVoiceOS/speaker-embeddings-onnx

Install

pip install speakeronnx

Optional high-quality resampling:

pip install speakeronnx soxr

Quick start

from speakeronnx import SpeakerEmbedder, cosine, verify

embedder = SpeakerEmbedder(model="wespeaker-resnet34")

alice1 = embedder.embed("alice_clip1.wav")
alice2 = embedder.embed("alice_clip2.wav")
bob    = embedder.embed("bob_clip1.wav")

print(cosine(alice1, alice2))   # e.g. 0.82  - same speaker
print(cosine(alice1, bob))      # e.g. 0.21  - different speaker

ok, score = verify(alice1, alice2, threshold=0.45)
print(ok, score)  # True 0.82

More examples in examples/.

CLI

speakeronnx list                              # list available models
speakeronnx embed clip.wav                    # extract embedding
speakeronnx verify a.wav b.wav               # same-speaker check (exit 0/1)
speakeronnx verify a.wav b.wav --threshold 0.5
speakeronnx embed clip.wav --model wespeaker-ecapa512

Full CLI reference in docs/cli.md.

Models

The library registers 9 models in MODEL_REGISTRY. Each one downloads on first use:

Alias Embed dim Frontend License
wespeaker-resnet34 256 fbank80 cc-by-4.0
wespeaker-ecapa512 192 fbank80 cc-by-4.0
wespeaker-resnet293 256 fbank80 cc-by-4.0
campplus 512 fbank80 cc-by-4.0
campplus-zh-en 192 fbank80 apache-2.0
eres2net 192 fbank80 apache-2.0
titanet-small 192 fbank80 cc-by-4.0
titanet-large 192 fbank80 cc-by-4.0
redimnet-b2 192 raw apache-2.0

Full model comparison and selection guide in docs/models.md.

Documentation

Document Description
docs/index.md Full getting-started guide
docs/models.md Model comparison, selection, frontend/layout details
docs/api.md Complete API reference
docs/cli.md CLI usage reference
docs/frontend.md Feature frontend (fbank80 vs raw) technical details
docs/advanced.md Custom models, GPU, threshold tuning

Examples

Script Description
examples/basic_embedding.py Extract embedding from a single WAV
examples/verify_speakers.py Verify two clips, try multiple thresholds
examples/compare_models.py Compare all models on same utterances
examples/batch_enrollment.py Enroll speakers from directories, match unknown
examples/custom_model.py Load a custom ONNX model from disk
examples/gpu_inference.py CUDA / CoreML inference

Tests

# Unit tests (mocked, no downloads, no network)
pytest tests/test_unit.py tests/test_audio.py tests/test_frontend.py \
      tests/test_embedder.py tests/test_cli.py tests/test_model_registry.py -v

# End-to-end tests (downloads models + generates TTS audio)
pytest tests/test_e2e.py -v -s

Each registered model has an HF source, license, and embedding dimension:

Alias HF repo License Embed dim Description
wespeaker-resnet34 Wespeaker/wespeaker-voxceleb-resnet34-LM cc-by-4.0 256 ResNet34 r-vector, VoxCeleb2 Dev - recommended default
wespeaker-ecapa512 Wespeaker/wespeaker-ecapa-tdnn512-LM cc-by-4.0 192 ECAPA-TDNN-512 x-vector, VoxCeleb2 Dev
wespeaker-resnet293 Wespeaker/wespeaker-voxceleb-resnet293-LM cc-by-4.0 256 ResNet293 r-vector - highest accuracy, 28M params
campplus csukuangfj/speaker-embedding-models cc-by-4.0 512 CAM++ (D-TDNN backbone), VoxCeleb2 Dev
campplus-zh-en csukuangfj/speaker-embedding-models apache-2.0 192 3D-Speaker CAM++ multilingual (zh+en)
eres2net csukuangfj/speaker-embedding-models apache-2.0 192 ERes2Net, VoxCeleb
titanet-small csukuangfj/speaker-embedding-models cc-by-4.0 192 NVIDIA NeMo TitaNet-small (~40 MB)
titanet-large csukuangfj/speaker-embedding-models cc-by-4.0 192 NVIDIA NeMo TitaNet-large (~101 MB)
redimnet-b2 OpenVoiceOS/redimnet-b2-vox2-onnx apache-2.0 192 ReDimNet b2 (1.8M params), raw audio input

Run speakeronnx list to print descriptions and metadata for all registered models.

Feature frontend

  • fbank80 models use an 80-dim log-Mel filterbank with per-utterance CMN, implemented in pure numpy. See docs/frontend.md.
  • raw models (redimnet-b2) pass the raw 16 kHz waveform directly to ONNX. The model has an internal MelSpectrogram.

Audio requirements

  • Mono PCM WAV, any bit depth (8/16/24/32-bit int, 32-bit float)
  • Any sample rate (the library resamples internally to 16 kHz)
  • The library downmixes stereo files to mono
  • Minimum length ~1 second
  • Recommended enrollment length 5-30 seconds per speaker

Dependencies

  • onnxruntime
  • numpy
  • huggingface_hub
  • soxr (optional, for high-quality resampling)

Related projects

Metadata

Release files for speakeronnx 0.0.2a2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for speakeronnx 0.0.2a2
File Size Uploaded
speakeronnx-0.0.2a2.tar.gz 29.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for speakeronnx 0.0.2a2
File Interpreter ABI Platform
speakeronnx-0.0.2a2-py3-none-any.whl Python 3 none any Details

Total release size: 63.8 kB

Release files / speakeronnx-0.0.2a2.tar.gz

Download URL speakeronnx-0.0.2a2.tar.gz
Size 29.3 kB
Tags Source
SHA-256 checksum
How to use checksums
40258421a45efff2620f2d64c6beda6193d284db1f559433a74432a414c39fe7
BLAKE2b-256 checksum
How to use checksums
59e241344bb18be3edb504e6e58c87f8c638402648b7798b6d4320c1f1acc1c6
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / speakeronnx-0.0.2a2-py3-none-any.whl

Download URL speakeronnx-0.0.2a2-py3-none-any.whl
Size 34.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
b6ca00696cf89f887ecde6d88fcb68ceb56af1f1d9a140f4cf76557f09d4bd6c
BLAKE2b-256 checksum
How to use checksums
df6a46fc2dddaa9bae1f9de6f8e4a9b6e9a51af5ca5dca8ae8ea0b21354b2497
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page