Skip to main content

Neural building blocks for speaker diarization

Project description

:warning: Checkout develop branch to see what is coming in pyannote.audio 2.0:

Neural speaker diarization with pyannote-audio

pyannote.audio is an open-source toolkit written in Python for speaker diarization. Based on PyTorch machine learning framework, it provides a set of trainable end-to-end neural building blocks that can be combined and jointly optimized to build speaker diarization pipelines:

pyannote.audio also comes with pretrained models covering a wide range of domains for voice activity detection, speaker change detection, overlapped speech detection, and speaker embedding:

segmentation

Open In Colab

Installation

pyannote.audio only supports Python 3.7 (or later) on Linux and macOS. It might work on Windows but there is no garantee that it does, nor any plan to add official support for Windows.

The instructions below assume that pytorch has been installed using the instructions from https://pytorch.org.

$ pip install pyannote.audio==1.1.1

Documentation and tutorials

Until a proper documentation is released, note that part of the API is described in this tutorial.

Citation

If you use pyannote.audio please use the following citation

@inproceedings{Bredin2020,
  Title = {{pyannote.audio: neural building blocks for speaker diarization}},
  Author = {{Bredin}, Herv{\'e} and {Yin}, Ruiqing and {Coria}, Juan Manuel and {Gelly}, Gregory and {Korshunov}, Pavel and {Lavechin}, Marvin and {Fustes}, Diego and {Titeux}, Hadrien and {Bouaziz}, Wassim and {Gill}, Marie-Philippe},
  Booktitle = {ICASSP 2020, IEEE International Conference on Acoustics, Speech, and Signal Processing},
  Address = {Barcelona, Spain},
  Month = {May},
  Year = {2020},
}

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

pyannote.audio-1.1.2.tar.gz (138.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

pyannote.audio-1.1.2-py3-none-any.whl (231.2 kB view details)

Uploaded Python 3

File details

Details for the file pyannote.audio-1.1.2.tar.gz.

File metadata

  • Download URL: pyannote.audio-1.1.2.tar.gz
  • Upload date:
  • Size: 138.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/3.3.0 pkginfo/1.7.0 requests/2.25.1 setuptools/49.2.1 requests-toolbelt/0.9.1 tqdm/4.56.0 CPython/3.9.1

File hashes

Hashes for pyannote.audio-1.1.2.tar.gz
Algorithm Hash digest
SHA256 b60f6bb899711278f6884fccd4a10cbf3249408e308b82a662dc6d22613625e7
MD5 55da6b71d26a35d99ffd25a6fca7900c
BLAKE2b-256 d1ff2b9ea1d947460f28d4832ca980c803ac35e2ea883fa54c1d0a8a7e558a6c

See more details on using hashes here.

File details

Details for the file pyannote.audio-1.1.2-py3-none-any.whl.

File metadata

  • Download URL: pyannote.audio-1.1.2-py3-none-any.whl
  • Upload date:
  • Size: 231.2 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/3.3.0 pkginfo/1.7.0 requests/2.25.1 setuptools/49.2.1 requests-toolbelt/0.9.1 tqdm/4.56.0 CPython/3.9.1

File hashes

Hashes for pyannote.audio-1.1.2-py3-none-any.whl
Algorithm Hash digest
SHA256 a85ac37471ab4d5e5aa98a24f13a3dec966b6aac222330fd883d6cd2f0a00ff3
MD5 70b913e906dd84ef840a1acbb3a89420
BLAKE2b-256 1b141009910780ab9a7eda0f213de08b117b48687df8f8e7f598c3941a3383f5

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page