Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

vadonnx

Load arbitrary Voice Activity Detection models behind a single, unified API. Every model runs through ONNX Runtime. You get one streaming and batch interface, one audio-format story, and pluggable models.

from vadonnx import load_vad

vad = load_vad("silero")                              # bundled, works fully offline
prob = vad.process_chunk(pcm_bytes)                   # streaming → float in [0, 1]
segments = vad.get_speech_segments(audio, sample_rate=16000)
# -> [SpeechSegment(start=0.32, end=2.27), SpeechSegment(start=3.27, end=4.45), ...]

Why

Every VAD ships its own loader, audio format, feature pipeline, and state handling. vadonnx hides that behind one VADModel interface. Feed it audio (raw int16 bytes, numpy arrays, any sample rate) and get back per-frame speech probabilities or ready-made speech segments. An IOSignature describes each model declaratively, so a single generic engine drives most of them, and you can point the same API at any custom .onnx file.

  • Lightweight runtime: only numpy, onnxruntime, and huggingface_hub.
  • Offline by default: a small Silero model is bundled in the wheel.
  • Streaming and batch: process_chunk() for live audio, get_speech_segments() and probabilities() for whole buffers.
  • Bring your own model: load any ONNX VAD by path or URL with a signature.
  • Extensible: third parties register backends and models through entry points.

Install

uv pip install vadonnx          # runtime (numpy + onnxruntime + huggingface_hub)
uv pip install "vadonnx[mic]"   # + microphone examples

Models

name rate parity vs upstream notes
silero / silero-8k / silero-op15 16k / 8k / 16k MAE 0 bundled default, raw PCM
marblenet / marblenet-int8 16k MAE 4e-4 NVIDIA NeMo Frame-VAD, multilingual (license)
pyannote / pyannote-int8 16k MAE 0 pyannote segmentation-3.0, windowed
fsmn / fsmn-quant 16k tracks upstream FunASR FSMN-VAD, needs vadonnx[fsmn]
speechbrain 16k MAE 0 SpeechBrain CRDNN, LibriParty-trained
ten 16k n/a feature extractor provided by TEN's native library
pulsevad / pulsevad-fp32 / pulsevad-81k 16k MAE 2e-4 / 1.5e-7 / 2.4e-7 PulseVAD tiny causal CNN, 200 ms windows

See docs/backends.md for per-model detail and the benchmark for measured comparisons across datasets, including WebRTC and energy baselines.

Models other than the bundled Silero are downloaded on first use from the TigreGotico HuggingFace org and cached under $XDG_DATA_HOME/vadonnx. The PulseVAD models download from a pinned commit of the upstream GitHub repository, and a sha256 digest checks each file. See docs/backends.md for per-model detail and parity notes.

CLI

vadonnx list                       # list available models
vadonnx probe silero               # print a model's ONNX input/output signature
vadonnx segment speech.wav         # print detected speech segments of a WAV

Documentation

License

Apache-2.0. Bundled and downloaded model weights retain their upstream licenses. See docs/licensing.md.

Release files for vadonnx 0.2.0a1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for vadonnx 0.2.0a1
File Size Uploaded
vadonnx-0.2.0a1.tar.gz 2.3 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for vadonnx 0.2.0a1
File Interpreter ABI Platform
vadonnx-0.2.0a1-py3-none-any.whl Python 3 none any Details

Total release size: 4.3 MB

Release files / vadonnx-0.2.0a1.tar.gz

Download URL vadonnx-0.2.0a1.tar.gz
Size 2.3 MB
Tags Source
SHA-256 checksum
How to use checksums
8d42fd6c8bbb55a8f269820cff9764c53859b1ed8a8c757eb01fcfa9f2419dfa
BLAKE2b-256 checksum
How to use checksums
63d28ded1e66862ce88ed518f72b2e3c0361aebf047dd4b8a5e3b522acc38256
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / vadonnx-0.2.0a1-py3-none-any.whl

Download URL vadonnx-0.2.0a1-py3-none-any.whl
Size 2.0 MB
Tags Python 3
SHA-256 checksum
How to use checksums
b37905403e32f12629a55ebe081cfcb742223030720caec7ceb66d1baab6b8f2
BLAKE2b-256 checksum
How to use checksums
c12585653980223ffe419d176d0daca0930cefd0be7451a20daacad14b27755b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page