Skip to main content

wav2x

PyPI Build Status License

A lightweight Python package for audio representation learning, spoken language identification, speaker recognition (d-vector), and voice activity detection (VAD) with TensorFlow Lite.

wav2x replaces the heavy C++ Lingvo dependency with a pure, standalone TensorFlow/NumPy frontend that runs on standard Python (including Python 3.10 and 3.11+) across macOS, Linux, and Windows.


Installation

Install the package from PyPI:

pip install wav2x

Or install from the repository in development mode:

git clone https://github.com/wq2012/wav2x.git
cd wav2x
pip install -r requirements.txt
pip install -e .

Features

  • Lingvo-compatible Log-Mel Spectrogram Frontend: High-fidelity pure TensorFlow reimplementation of Lingvo's MelAsrFrontend (framing, pre-emphasis, windowing, overdrive FFT, Mel filterbank, log compression, frame stacking, and subsampling).
  • Spoken Language Identification (Lang-ID): Conformer-based streaming model classifying speech into 100+ supported languages.
  • Speaker Recognition (Speaker-ID / d-vector): Conformer-based speaker encoder producing robust d-vector representations for cross-lingual speaker verification and identification.
  • Integrated Voice Activity Detection (VAD): Multi-stage neural VAD for speech frame filtering prior to encoder inference.
  • Convenient HuggingFace Hub Integration: Auto-download pretrained TFLite models on the fly via from_pretrained().

Quickstart & Usage

1. Spoken Language Identification

Identify the spoken language of an audio file using the Conformer Lang-ID model:

from wav2x import WavToLangRunner

# Automatically downloads models from HuggingFace Hub if not present locally
runner = WavToLangRunner.from_pretrained(model_dir="models")

# Predict language
language_code, probs = runner.wav_to_lang("speech.wav")
print(f"Predicted language: {language_code}")

You can also pass explicit model paths:

runner = WavToLangRunner(
    vad_model_file="models/vad_short_model.tflite",
    vad_mean_stddev_file="models/vad_short_mean_stddev.csv",
    langid_model_file="models/conformer_langid_medium.tflite",
    vad_threshold=0.1,
)
language_code, probs = runner.wav_to_lang("speech.wav")

2. Speaker Recognition (d-vector & Speaker Verification)

Compute d-vector embeddings and speaker verification similarity scores:

from wav2x import WavToDvectorRunner

runner = WavToDvectorRunner.from_pretrained(model_dir="models")

# Compute d-vector sequence for an audio file
dvectors = runner.wav_to_dvector("utterance1.wav")
last_dvector = dvectors[-1, :]

# Compute cosine similarity between enrolled audio and test audio
similarity = runner.compute_score(
    enroll_audio_list=["enroll1.wav", "enroll2.wav"],
    test_audio="test.wav"
)
print(f"Speaker similarity score: {similarity:.4f}")

3. Speaker Enrollment & Identification

Enroll multiple known speakers and identify speakers in test audio:

from wav2x import WavToDvectorRunner

runner = WavToDvectorRunner.from_pretrained(model_dir="models")

# Enroll speakers
runner.enroll_speaker("Alice", ["alice_audio_1.wav", "alice_audio_2.wav"])
runner.enroll_speaker("Bob", ["bob_audio_1.wav"])

# Identify speaker in unknown audio
speaker_name, score = runner.identify_speaker("unknown.wav", threshold=0.50)
if speaker_name:
    print(f"Identified as {speaker_name} with score {score:.4f}")
else:
    print(f"Unknown speaker (best score: {score:.4f})")

4. Audio Feature Extraction (Log-Mel Spectrogram)

Extract the exact 512-dimensional stacked log-mel features used by Conformer models:

from wav2x import LogMelFeatureExtractor
import soundfile as sf

extractor = LogMelFeatureExtractor(
    frame_size_ms=32.0,
    frame_step_ms=10.0,
    num_bins=128,
    sample_rate=16000.0,
    stack_left_context=3,
    frame_stride=3,
)

# audio samples shaped [1, time] with int16 range [-32768, 32767]
data, sr = sf.read("audio.wav")
samples = (data * 32768.0).astype("int16").reshape(1, -1)

# features shape: [1, num_frames, 512]
features = extractor.extract(samples)
print(f"Extracted feature shape: {features.shape}")

Interactive Web Demos

Interactive Gradio web interfaces are available under demos/:

pip install gradio
python demos/lang-id-demo.py
python demos/speaker-id-demo.py

Running Tests

Run the full unit and end-to-end test suite:

bash run_tests.sh

Linting:

flake8 --indent-size 2 --max-line-length 80 .

Models

Pretrained models are hosted on Hugging Face Hub:


Citations

The underlying conformer architectures and training methodologies are described in:

@inproceedings{pelecanos2022parameter,
  title={Parameter-Free Attentive Scoring for Speaker Verification},
  author={Jason Pelecanos and Quan Wang and Yiling Huang and Ignacio Lopez Moreno},
  booktitle={Odyssey: The Speaker and Language Recognition Workshop},
  year={2022}
}

@inproceedings{wang2022attentive,
  title={Attentive Temporal Pooling for Conformer-based Streaming Language Identification in Long-form Speech},
  author={Quan Wang and Yang Yu and Jason Pelecanos and Yiling Huang and Ignacio Lopez Moreno},
  booktitle={Odyssey: The Speaker and Language Recognition Workshop},
  year={2022}
}

@inproceedings{chojnacka2021speakerstew,
  title={{SpeakerStew: Scaling to many languages with a triaged multilingual text-dependent and text-independent speaker verification system}},
  author={Chojnacka, Roza and Pelecanos, Jason and Wang, Quan and Moreno, Ignacio Lopez},
  booktitle={Prod. Interspeech},
  pages={1064--1068},
  year={2021},
  doi={10.21437/Interspeech.2021-646},
  issn={2958-1796},
}

License

Apache License 2.0

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

wav2x-0.1.0.tar.gz (12.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

wav2x-0.1.0-py3-none-any.whl (11.3 kB view details)

Uploaded Python 3

File details

Details for the file wav2x-0.1.0.tar.gz.

File metadata

  • Download URL: wav2x-0.1.0.tar.gz
  • Upload date:
  • Size: 12.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.11.16

File hashes

Hashes for wav2x-0.1.0.tar.gz
Algorithm Hash digest
SHA256 bcd3f9b15b80380fd598997496f138c0d659469b97b4a38990bd6395bca5105c
MD5 05ce5b7ac02ebc5534a8bb0f9c270be0
BLAKE2b-256 814d74eb3be3197da60db9623dcc7bebac30724b4e41ddde4a8954090b8dbf7c

See more details on using hashes here.

File details

Details for the file wav2x-0.1.0-py3-none-any.whl.

File metadata

  • Download URL: wav2x-0.1.0-py3-none-any.whl
  • Upload date:
  • Size: 11.3 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.11.16

File hashes

Hashes for wav2x-0.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 e77c21f22782c647f4145f82d5bc6d116d209a222116545f8d71c9941bc553fa
MD5 3a6fdc784a1b9d8a48f07ebf6c6c9be2
BLAKE2b-256 2a6567dece2135dba46df8dab18209ca2ef9a3351d922d44bee098310dc94f27

See more details on using hashes here.

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page