Skip to main content

Neural building blocks for speaker diarization

Project description

Neural speaker diarization with pyannote-audio

pyannote.audio is an open-source toolkit written in Python for speaker diarization. Based on PyTorch machine learning framework, it provides a set of trainable end-to-end neural building blocks that can be combined and jointly optimized to build speaker diarization pipelines:

pyannote.audio also comes with pretrained models covering a wide range of domains for voice activity detection, speaker change detection, overlapped speech detection, and speaker embedding:

segmentation

Open In Colab

Installation

pyannote.audio only supports Python 3.7 (or later) on Linux and macOS. It might work on Windows but there is no garantee that it does, nor any plan to add official support for Windows.

The instructions below assume that pytorch has been installed using the instructions from https://pytorch.org.

$ pip install pyannote.audio==1.1

Documentation and tutorials

Until a proper documentation is released, note that part of the API is described in this tutorial.

Citation

If you use pyannote.audio please use the following citation

@inproceedings{Bredin2020,
  Title = {{pyannote.audio: neural building blocks for speaker diarization}},
  Author = {{Bredin}, Herv{\'e} and {Yin}, Ruiqing and {Coria}, Juan Manuel and {Gelly}, Gregory and {Korshunov}, Pavel and {Lavechin}, Marvin and {Fustes}, Diego and {Titeux}, Hadrien and {Bouaziz}, Wassim and {Gill}, Marie-Philippe},
  Booktitle = {ICASSP 2020, IEEE International Conference on Acoustics, Speech, and Signal Processing},
  Address = {Barcelona, Spain},
  Month = {May},
  Year = {2020},
}

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

pyannote.audio-1.1.1.tar.gz (137.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

pyannote.audio-1.1.1-py3-none-any.whl (230.9 kB view details)

Uploaded Python 3

File details

Details for the file pyannote.audio-1.1.1.tar.gz.

File metadata

  • Download URL: pyannote.audio-1.1.1.tar.gz
  • Upload date:
  • Size: 137.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/3.2.0 pkginfo/1.6.1 requests/2.25.0 setuptools/49.2.1 requests-toolbelt/0.9.1 tqdm/4.53.0 CPython/3.9.0

File hashes

Hashes for pyannote.audio-1.1.1.tar.gz
Algorithm Hash digest
SHA256 a17aa16e9967fc2ae7f935009981524feffbd16f5e24aa82bd0148577f35430e
MD5 e41a98d483767edff726d3388da80b8d
BLAKE2b-256 87378658728839156e77ca2fc062fa140337a9b4a6f64c6f63da01c4a714fe89

See more details on using hashes here.

File details

Details for the file pyannote.audio-1.1.1-py3-none-any.whl.

File metadata

  • Download URL: pyannote.audio-1.1.1-py3-none-any.whl
  • Upload date:
  • Size: 230.9 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/3.2.0 pkginfo/1.6.1 requests/2.25.0 setuptools/49.2.1 requests-toolbelt/0.9.1 tqdm/4.53.0 CPython/3.9.0

File hashes

Hashes for pyannote.audio-1.1.1-py3-none-any.whl
Algorithm Hash digest
SHA256 d2d0132f6722f13bb7b96624cfbe04952362b90ec2eaf6a11fa3bc8c5eb5e690
MD5 612857831440cb46741e2f4528d47901
BLAKE2b-256 b89e3539c9d74a477eba4051aadb2a686b99249358f9d8f513780acab2d2a54e

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page