simple_diarizer
Simplified diarization pipeline using some pretrained models.
Made to be a simple as possible to go from an input audio file to diarized segments.
import soundfile as sf
import matplotlib.pyplot as plt
from simple_diarizer.diarizer import Diarizer
from simple_diarizer.utils import combined_waveplot
diar = Diarizer(
embed_model='xvec', # 'xvec' and 'ecapa' supported
cluster_method='sc' # 'ahc' and 'sc' supported
)
segments = diar.diarize(WAV_FILE, num_speakers=NUM_SPEAKERS)
signal, fs = sf.read(WAV_FILE)
combined_waveplot(signal, fs, segments)
plt.show()
Install
Simplified diarization is available on PyPI:
pip install simple-diarizer
Source Video
"Some Quick Advice from Barack Obama!"
Pre-trained Models
The following pretrained models are used:
- Voice Activity Detection (VAD)
- Deep speaker embedding extraction
- (Optional/Experimental) Speech-to-text
- ESPnet Model Zoo
- English ASR model
- ESPnet Model Zoo
Demo
It can be checked out in the above link, where it will try and diarize any input YouTube URL.
Other References
- Spectral clustering methods lifted from https://github.com/wq2012/SpectralCluster
Planned Features
Release files for simple-diarizer 0.0.13
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| simple_diarizer-0.0.13-py3-none-any.whl | Python 3 | none | any | Details |
Release files / simple_diarizer-0.0.13-py3-none-any.whl
| Download URL | simple_diarizer-0.0.13-py3-none-any.whl |
|---|---|
| Size | 23.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
e331532a63ca9cd9ce98157c2a5c48e9ed2fbee683757298f678dd954e13278e
|
|
BLAKE2b-256 checksum How to use checksums |
ee333f214ea395176ccd2be3de075d6e242c3eb70c729b9ae832a5f0b6958533
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/4.0.2 CPython/3.9.16
|