x-vector extractor for Japanese speech
This repository provides a pre-trained model for extracting the x-vector (speaker representation vector). The model is trained using JTubeSpeech corpus, a Japanese speech corpus collected from YouTube.
このリポジトリは,x-vector (話者表現ベクトル) を抽出するための学習済みモデルを提供します.このモデルは,JTubeSpeechコーパスと呼ばれる,YouTubeから収集した日本語音声から学習されています.
Training configures / 学習時の設定
- The number of speakers: 1,233
- Sampling frequency: 16,000Hz
- Speaker recognition accuracy: 91% (test data)
- Feature: 24-dimensional MFCC
- Dimensionality of x-vector: 512
- Other configurations: followed the ASV recipe for VoxCeleb in Kaldi.
- In the opensourced model, model parameters of recognition layers following to the x-vector layer were randomized to protect data privacy.
Installation
pip install xvector-jtubespeech
Usage / 使い方
import numpy as np
from scipy.io import wavfile
import torch
from torchaudio.compliance import kaldi
from xvector_jtubespeech import XVector
def extract_xvector(
model, # xvector model
wav # 16kHz mono
):
# extract mfcc
wav = torch.from_numpy(wav.astype(np.float32)).unsqueeze(0)
mfcc = kaldi.mfcc(wav, num_ceps=24, num_mel_bins=24) # [1, T, 24]
mfcc = mfcc.unsqueeze(0)
# extract xvector
xvector = model.vectorize(mfcc) # (1, 512)
xvector = xvector.to("cpu").detach().numpy().copy()[0]
return xvector
_, wav = wavfile.read("sample.wav") # 16kHz mono
model = XVector("xvector.pth")
xvector = extract_xvector(model, wav) # (512, )
Contributors / 貢献者
- Takaki Hamada / 濱田 誉輝 (The University of Tokyo / 東京大学)
- Shinnosuke Takamichi / 高道 慎之介 (The University of Tokyo / 東京大学)
License / ライセンス
MIT
Others / その他
- The audio sample
sample.wavwas copied from PJS corpus.
Release files for xvector-jtubespeech 0.0.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| xvector_jtubespeech-0.0.2.tar.gz | 5.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| xvector_jtubespeech-0.0.2-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 10.1 kB
Release files / xvector_jtubespeech-0.0.2.tar.gz
| Download URL | xvector_jtubespeech-0.0.2.tar.gz |
|---|---|
| Size | 5.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
8b76fd5b701056b21658741ba71e92abc47fe4416431a7e0e7ffddfcfa32f364
|
|
BLAKE2b-256 checksum How to use checksums |
a704b904a8430fe75c39946ab47a604c0bb7f5f96be86d24d58beff1d7814a68
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/4.0.2 CPython/3.10.11
|
Release files / xvector_jtubespeech-0.0.2-py3-none-any.whl
| Download URL | xvector_jtubespeech-0.0.2-py3-none-any.whl |
|---|---|
| Size | 5.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
a3cf90ffe4e434995e1a8000f6a2c10ad6f67b748887a2736d9bf018d62ff853
|
|
BLAKE2b-256 checksum How to use checksums |
a0d1a49388abf8f1f587f49fef463e6b1bab444d7c0b76752ad25d2c3aded3d8
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/4.0.2 CPython/3.10.11
|