Skip to main content

Neural building blocks for speaker diarization

Project description

Neural speaker diarization with pyannote-audio

pyannote.audio is an open-source toolkit written in Python for speaker diarization. Based on PyTorch machine learning framework, it provides a set of trainable end-to-end neural building blocks that can be combined and jointly optimized to build speaker diarization pipelines:

pyannote.audio also comes with pretrained models covering a wide range of domains for voice activity detection, speaker change detection, overlapped speech detection, and speaker embedding:

segmentation

Open In Colab

Installation

pyannote.audio only supports Python 3.7 (or later) on Linux and macOS. It might work on Windows but there is no garantee that it does, nor any plan to add official support for Windows.

The instructions below assume that pytorch has been installed using the instructions from https://pytorch.org.

Until a proper release of pyannote.audio is available on PyPI, it must be installed from source using the develop branch of the official repository:

$ git clone https://github.com/pyannote/pyannote-audio.git
$ cd pyannote-audio
$ git checkout develop
$ pip install .

Documentation and tutorials

Until a proper documentation is released, note that part of the API is described in this tutorial.

Citation

If you use pyannote.audio please use the following citation

@inproceedings{Bredin2020,
  Title = {{pyannote.audio: neural building blocks for speaker diarization}},
  Author = {{Bredin}, Herv{\'e} and {Yin}, Ruiqing and {Coria}, Juan Manuel and {Gelly}, Gregory and {Korshunov}, Pavel and {Lavechin}, Marvin and {Fustes}, Diego and {Titeux}, Hadrien and {Bouaziz}, Wassim and {Gill}, Marie-Philippe},
  Booktitle = {ICASSP 2020, IEEE International Conference on Acoustics, Speech, and Signal Processing},
  Address = {Barcelona, Spain},
  Month = {May},
  Year = {2020},
}

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

pyannote.audio-1.1.tar.gz (137.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

pyannote.audio-1.1-py3-none-any.whl (230.8 kB view details)

Uploaded Python 3

File details

Details for the file pyannote.audio-1.1.tar.gz.

File metadata

  • Download URL: pyannote.audio-1.1.tar.gz
  • Upload date:
  • Size: 137.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/3.2.0 pkginfo/1.6.1 requests/2.24.0 setuptools/49.2.1 requests-toolbelt/0.9.1 tqdm/4.51.0 CPython/3.9.0

File hashes

Hashes for pyannote.audio-1.1.tar.gz
Algorithm Hash digest
SHA256 c8743b51e4d2688b102600afd32628efb0defecbcdcb3612f023eafdfa397a5b
MD5 cc09d2a4854b63f82b8849efc0c8de3a
BLAKE2b-256 da8acffbe3ad80d63fbe2aef0b2427f27da10ffa774fa82c73fe863a388e217b

See more details on using hashes here.

File details

Details for the file pyannote.audio-1.1-py3-none-any.whl.

File metadata

  • Download URL: pyannote.audio-1.1-py3-none-any.whl
  • Upload date:
  • Size: 230.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/3.2.0 pkginfo/1.6.1 requests/2.24.0 setuptools/49.2.1 requests-toolbelt/0.9.1 tqdm/4.51.0 CPython/3.9.0

File hashes

Hashes for pyannote.audio-1.1-py3-none-any.whl
Algorithm Hash digest
SHA256 80317f80abc3bde44d9c14116bfa484f1e3026b8cb89eb39b3c324f2029cca8f
MD5 80e91c0648169eb33f5087638723f9d3
BLAKE2b-256 5a6ffba20e05a52e7beb9177af8050c84b0c5e279645bf29d16672931b5a1ef3

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page