Skip to main content

x-vector extractor for Japanese speech

This repository provides a pre-trained model for extracting the x-vector (speaker representation vector). The model is trained using JTubeSpeech corpus, a Japanese speech corpus collected from YouTube.

このリポジトリは,x-vector (話者表現ベクトル) を抽出するための学習済みモデルを提供します.このモデルは,JTubeSpeechコーパスと呼ばれる,YouTubeから収集した日本語音声から学習されています.

Training configures / 学習時の設定

  • The number of speakers: 1,233
  • Sampling frequency: 16,000Hz
  • Speaker recognition accuracy: 91% (test data)
  • Feature: 24-dimensional MFCC
  • Dimensionality of x-vector: 512
  • Other configurations: followed the ASV recipe for VoxCeleb in Kaldi.
    • In the opensourced model, model parameters of recognition layers following to the x-vector layer were randomized to protect data privacy.

Installation

pip install xvector-jtubespeech

Usage / 使い方

import numpy as np
from scipy.io import wavfile
import torch
from torchaudio.compliance import kaldi

from xvector_jtubespeech import XVector

def extract_xvector(
  model, # xvector model
  wav   # 16kHz mono
):
  # extract mfcc
  wav = torch.from_numpy(wav.astype(np.float32)).unsqueeze(0)
  mfcc = kaldi.mfcc(wav, num_ceps=24, num_mel_bins=24) # [1, T, 24]
  mfcc = mfcc.unsqueeze(0)

  # extract xvector
  xvector = model.vectorize(mfcc) # (1, 512)
  xvector = xvector.to("cpu").detach().numpy().copy()[0]  

  return xvector

_, wav = wavfile.read("sample.wav") # 16kHz mono
model = XVector("xvector.pth")
xvector = extract_xvector(model, wav) # (512, )

Contributors / 貢献者

  • Takaki Hamada / 濱田 誉輝 (The University of Tokyo / 東京大学)
  • Shinnosuke Takamichi / 高道 慎之介 (The University of Tokyo / 東京大学)

License / ライセンス

MIT

Others / その他

  • The audio sample sample.wav was copied from PJS corpus.

Release files for xvector-jtubespeech 0.0.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for xvector-jtubespeech 0.0.2
File Size Uploaded
xvector_jtubespeech-0.0.2.tar.gz 5.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for xvector-jtubespeech 0.0.2
File Interpreter ABI Platform
xvector_jtubespeech-0.0.2-py3-none-any.whl Python 3 none any Details

Total release size: 10.1 kB

Release files / xvector_jtubespeech-0.0.2.tar.gz

Download URL xvector_jtubespeech-0.0.2.tar.gz
Size 5.1 kB
Tags Source
SHA-256 checksum
How to use checksums
8b76fd5b701056b21658741ba71e92abc47fe4416431a7e0e7ffddfcfa32f364
BLAKE2b-256 checksum
How to use checksums
a704b904a8430fe75c39946ab47a604c0bb7f5f96be86d24d58beff1d7814a68
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.2 CPython/3.10.11

Release files / xvector_jtubespeech-0.0.2-py3-none-any.whl

Download URL xvector_jtubespeech-0.0.2-py3-none-any.whl
Size 5.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
a3cf90ffe4e434995e1a8000f6a2c10ad6f67b748887a2736d9bf018d62ff853
BLAKE2b-256 checksum
How to use checksums
a0d1a49388abf8f1f587f49fef463e6b1bab444d7c0b76752ad25d2c3aded3d8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.2 CPython/3.10.11

Release history Release notifications | RSS feed

This release

0.0.2 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page