This release is a pre-release and may not be stable for production use.
vadonnx
Load arbitrary Voice Activity Detection models behind a single, unified API. Every model runs through ONNX Runtime. You get one streaming and batch interface, one audio-format story, and pluggable models.
from vadonnx import load_vad
vad = load_vad("silero") # bundled, works fully offline
prob = vad.process_chunk(pcm_bytes) # streaming → float in [0, 1]
segments = vad.get_speech_segments(audio, sample_rate=16000)
# -> [SpeechSegment(start=0.32, end=2.27), SpeechSegment(start=3.27, end=4.45), ...]
Why
Every VAD ships its own loader, audio format, feature pipeline, and state handling.
vadonnx hides that behind one VADModel interface. Feed it audio (raw int16
bytes, numpy arrays, any sample rate) and get back per-frame speech probabilities or
ready-made speech segments. An IOSignature describes each
model declaratively, so a single generic engine drives most of them, and you can point
the same API at any custom .onnx file.
- Lightweight runtime: only
numpy,onnxruntime, andhuggingface_hub. - Offline by default: a small Silero model is bundled in the wheel.
- Streaming and batch:
process_chunk()for live audio,get_speech_segments()andprobabilities()for whole buffers. - Bring your own model: load any ONNX VAD by path or URL with a signature.
- Extensible: third parties register backends and models through entry points.
Install
uv pip install vadonnx # runtime (numpy + onnxruntime + huggingface_hub)
uv pip install "vadonnx[mic]" # + microphone examples
Models
| name | rate | parity vs upstream | notes |
|---|---|---|---|
silero / silero-8k / silero-op15 |
16k / 8k / 16k | MAE 0 | bundled default, raw PCM |
marblenet / marblenet-int8 |
16k | MAE 4e-4 | NVIDIA NeMo Frame-VAD, multilingual (license) |
pyannote / pyannote-int8 |
16k | MAE 0 | pyannote segmentation-3.0, windowed |
fsmn / fsmn-quant |
16k | tracks upstream | FunASR FSMN-VAD, needs vadonnx[fsmn] |
speechbrain |
16k | MAE 0 | SpeechBrain CRDNN, LibriParty-trained |
ten |
16k | n/a | feature extractor provided by TEN's native library |
See docs/backends.md for per-model detail and the benchmark for measured comparisons across datasets, including WebRTC and energy baselines.
Models other than the bundled Silero are downloaded on first use from the
TigreGotico HuggingFace org and cached under
$XDG_DATA_HOME/vadonnx. See docs/backends.md for per-model detail
and parity notes.
CLI
vadonnx list # list available models
vadonnx probe silero # print a model's ONNX input/output signature
vadonnx segment speech.wav # print detected speech segments of a WAV
Documentation
- Quickstart
- Streaming
- Custom models &
IOSignature - Backends & parity notes
- Plugins
- Model conversion
- Licensing
- API reference
Related projects
- TigreGotico/phoonnx: multilingual phonemization and ONNX text-to-speech, the TTS counterpart to
vadonnx. - TigreGotico/onnx-asr: offline speech recognition on ONNX models.
License
Apache-2.0. Bundled and downloaded model weights retain their upstream licenses. See docs/licensing.md.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file vadonnx-0.1.1a1.tar.gz.
File metadata
- Download URL: vadonnx-0.1.1a1.tar.gz
- Upload date:
- Size: 2.3 MB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
2208fb5e92b4eb74ac81d01058820c2e5674f48043c1faf5576b68f130f08183
|
|
| MD5 |
1805f66f654e1be7a94eb3ed8209c192
|
|
| BLAKE2b-256 |
4e47ba039dccb4f98c0efcfd52b9f2327c2443c6691eb9bf77e04b988f4348e3
|
File details
Details for the file vadonnx-0.1.1a1-py3-none-any.whl.
File metadata
- Download URL: vadonnx-0.1.1a1-py3-none-any.whl
- Upload date:
- Size: 2.0 MB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
71e7afd600d1e14328ee1849cb05b426d8cf6f7eef0f7cacf4892bec827bb7b1
|
|
| MD5 |
8f64e44e69384898c6c6dfb586d7a586
|
|
| BLAKE2b-256 |
9360f5607a8eb92350b23fb9984f373d2bc400670729b949f1f993cef581fefc
|