Skip to main content

Mailing list : test Mailing list : test License: CC BY-NC 4.0 downloads

Open In Colab Test Package Pypi version Python version

header


Silero VAD


Silero VAD - pre-trained enterprise-grade Voice Activity Detector (also see our STT models).


Real Time Example

https://user-images.githubusercontent.com/36505480/144874384-95f80f6d-a4f1-42cc-9be7-004c891dd481.mp4

Please note, that video loads only if you are logged in your GitHub account.


Fast start


Dependencies

System requirements to run python examples on x86-64 systems:

  • python 3.8+;
  • 1G+ RAM;
  • A modern CPU with AVX, AVX2, AVX-512 or AMX instruction sets.

Required:

  • torch>=1.12.0.

Everything else is optional and installed via extras:

Extra Installs Needed for
silero-vad[audio] torchaudio>=0.12.0,<2.10, plus torchcodec on Python 3.9+ read_audio / save_audio
silero-vad[codec] torchcodec read_audio / save_audio with no torchaudio at all
silero-vad[onnx-cpu] onnxruntime>=1.16.1, numpy ONNX and sequence models
silero-vad[onnx-gpu] onnxruntime-gpu>=1.16.1, numpy ONNX on GPU
silero-vad[all] all of the above everything

[audio] pulls in torchcodec on Python 3.9+ because torchaudio>=2.9 hands decoding over to it and does not work without it; on Python 3.8 pip resolves an older torchaudio that decodes on its own. torchcodec itself needs FFmpeg (4-9) on the system.

The model itself needs only torch — if you already load audio yourself, pass a 1-D float32 torch.Tensor straight to get_speech_timestamps and install nothing extra.

The bundled I/O helpers (read_audio, save_audio) work with either backend: torchaudio if present, otherwise torchcodec. Note that torchcodec resamples through FFmpeg rather than torchaudio.transforms.Resample, so results can differ by ~1e-3 on files that need resampling.

With torchaudio, a proper audio backend is required:

  • Option №1 - FFmpeg backend. conda install -c conda-forge 'ffmpeg<7';
  • Option №2 - sox_io backend. apt-get install sox, TorchAudio is tested on libsox 14.4.2;
  • Option №3 - soundfile backend. pip install soundfile.

If you are planning to run the VAD using solely the onnx-runtime, it will run on any other system architectures where onnx-runtume is supported. In this case please note that:

  • You will have to implement the I/O;
  • You will have to adapt the existing wrappers / examples / post-processing for your use-case.

Using pip: pip install silero-vad[audio] (or plain pip install silero-vad if you load audio yourself)

from silero_vad import load_silero_vad, read_audio, get_speech_timestamps
model = load_silero_vad()
wav = read_audio('path_to_audio_file')
speech_timestamps = get_speech_timestamps(
  wav,
  model,
  return_seconds=True,  # Return speech timestamps in seconds (default is samples)
)

Using torch.hub:

import torch
torch.set_num_threads(1)

model, utils = torch.hub.load(repo_or_dir='snakers4/silero-vad', model='silero_vad')
(get_speech_timestamps, _, read_audio, _, _) = utils

wav = read_audio('path_to_audio_file')
speech_timestamps = get_speech_timestamps(
  wav,
  model,
  return_seconds=True,  # Return speech timestamps in seconds (default is samples)
)

Key Features


  • Stellar accuracy

    Silero VAD has excellent results on speech detection tasks.

  • Fast

    One audio chunk (30+ ms) takes less than 1ms to be processed on a single CPU thread. Using batching or GPU can also improve performance considerably. Under certain conditions ONNX may even run up to 4-5x faster.

  • Lightweight

    JIT model is around two megabytes in size.

  • General

    Silero VAD was trained on huge corpora that include over 6000 languages and it performs well on audios from different domains with various background noise and quality levels.

  • Flexible sampling rate

    Silero VAD supports 8000 Hz and 16000 Hz sampling rates.

  • Highly Portable

    Silero VAD reaps benefits from the rich ecosystems built around PyTorch and ONNX running everywhere where these runtimes are available.

  • No Strings Attached

    Published under permissive license (MIT) Silero VAD has zero strings attached - no telemetry, no keys, no registration, no built-in expiration, no keys or vendor lock.


Typical Use Cases


  • Voice activity detection for IOT / edge / mobile use cases
  • Data cleaning and preparation, voice detection in general
  • Telephony and call-center automation, voice bots
  • Voice interfaces

Links



Get In Touch


Try our models, create an issue, start a discussion, join our telegram chat, email us, read our news.

Please see our wiki for relevant information and email us directly.

Citations

@misc{Silero VAD,
  author = {Silero Team},
  title = {Silero VAD: pre-trained enterprise-grade Voice Activity Detector (VAD), Number Detector and Language Classifier},
  year = {2024},
  publisher = {GitHub},
  journal = {GitHub repository},
  howpublished = {\url{https://github.com/snakers4/silero-vad}},
  commit = {insert_some_commit_here},
  email = {hello@silero.ai}
}

Examples and VAD-based Community Apps


Metadata

Release files for silero-vad 6.2.3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for silero-vad 6.2.3
File Size Uploaded
silero_vad-6.2.3.tar.gz 29.3 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for silero-vad 6.2.3
File Interpreter ABI Platform
silero_vad-6.2.3-py3-none-any.whl Python 3 none any Details

Total release size: 40.6 MB

Release files / silero_vad-6.2.3.tar.gz

Download URL silero_vad-6.2.3.tar.gz
Size 29.3 MB
Tags Source
SHA-256 checksum
How to use checksums
d7a23b09202be4953915cb9173089bb307eed0418a9bd4794d9ec5af07fa1996
BLAKE2b-256 checksum
How to use checksums
3217df6bfc8f659fc5c1e082c18275be67c5a581bd2d07734a65b141fc983ccf
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / silero_vad-6.2.3-py3-none-any.whl

Download URL silero_vad-6.2.3-py3-none-any.whl
Size 11.3 MB
Tags Python 3
SHA-256 checksum
How to use checksums
7b7f5436cfcb02fae583a05b512ea96467fd449fe54cb49a5e4f06c51a1e43b8
BLAKE2b-256 checksum
How to use checksums
84ef9099037ed6f180ea33220178df4107112c0ce2bf5fb4d6f6ab19db2844ed
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release history Release notifications | RSS feed

This release

6.2.3 This release

2 release files

6.2.2

2 release files

6.2.1

2 release files

6.2.0

2 release files

6.1.0

2 release files

6.0.0

2 release files

5.1.2

2 release files

5.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page