Skip to main content

A simple speech enhancement library using ONNX models

Project description

SeLib

SeLib is a small speech-enhancement library for ONNX models. It provides a simple enhance() function, lightweight model wrappers, and utility functions for audio/noise experiments.

Install

Install the package directly from GitHub for the latest version.

pip install git+https://github.com/nipponjo/selib.git

or for the latest release:

pip install speech-selib

Basic Use

The shortest path is selib.enhance(wave, model_id). Pass a NumPy waveform at the sample rate expected by the model.

import librosa
import selib

from selib.utils import NoiseGenerator, snr_mixer

# Load clean speech at the model sample rate.
wave_clean, sr = librosa.load("./data/clean_freesound_33711.wav", sr=48000)

# Create artificial noise and mix it with the clean speech at 5 dB SNR.
noise_gen = NoiseGenerator(sample_rate=sr)
noise = noise_gen.sample(n_samples=len(wave_clean))
wave_noisy = snr_mixer(wave_clean, noise, snr=5)

# Enhance with a registered model id.
# The ONNX model is downloaded and cached automatically on first use.
wave_enhanced = selib.enhance(wave_noisy, "deepfilternet3")

You can also save the enhanced waveform directly. The loaded model is cached by default, so repeated calls with the same model_id do not reload the ONNX session.

# Save as 24-bit PCM WAV while returning the enhanced NumPy array.
wave_enhanced = selib.enhance(
    wave_noisy,
    "deepfilternet3",    
    save_to="enhanced.wav",
    bits_per_sample=24,
)

Model Objects

For more control, load the model wrapper once and call its enhance() method directly.

from selib import load_model
from selib.models import DeepFilterNetOnnx

# load_model() chooses the correct wrapper from the registry metadata.
model: DeepFilterNetOnnx = load_model("deepfilternet3")

# DeepFilterNetOnnx supports extra options such as attenuation limiting.
wave_enhanced = model.enhance(wave_noisy, atten_lim_db=12)

Magnitude-mask models can also be used directly. They predict a mask, multiply it with the noisy magnitude spectrogram, and reuse the noisy phase.

from selib.models import MagnitudeMaskModel

# ul_unas_16k is a 16 kHz magnitude-mask model.
mask_model = MagnitudeMaskModel("ul_unas_16k")
wave_enhanced = mask_model.enhance(wave_noisy_16k)

Metrics

SeLib includes a few quick NumPy-only metrics that do not require PESQ/STOI or other external scoring libraries.

from selib.metrics import snr, segsnr, si_sdr

# Compare clean reference speech against an enhanced waveform.
print("SNR:", snr(wave_clean, wave_enhanced))
print("segSNR:", segsnr(wave_clean, wave_enhanced, sample_rate=sr))
print("SI-SDR:", si_sdr(wave_clean, wave_enhanced))

Available Models

Model ID Type Sample rate #params Paper Repository
deepfilternet3 DeepFilterNet 48 kHz 2.13M arXiv:2305.08227 GitHub
deepfilternet2 DeepFilterNet 48 kHz 2.31M arXiv:2205.05474 GitHub
deepfilternet1 DeepFilterNet 48 kHz 1.78M arXiv:2110.05588 GitHub
ul_unas_16k Magnitude mask 16 kHz 0.171M arXiv:2503.00340 GitHub

DeepFilterNet models use ERB masks plus deep-filter coefficients. Magnitude-mask models predict a spectrogram mask and reconstruct audio with the noisy phase.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

speech_selib-0.2.0.tar.gz (28.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

speech_selib-0.2.0-py3-none-any.whl (31.5 kB view details)

Uploaded Python 3

File details

Details for the file speech_selib-0.2.0.tar.gz.

File metadata

  • Download URL: speech_selib-0.2.0.tar.gz
  • Upload date:
  • Size: 28.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.13

File hashes

Hashes for speech_selib-0.2.0.tar.gz
Algorithm Hash digest
SHA256 ec85787c1afc3838e2feff7771514746e47d432715f1a72a357b0d03207654c8
MD5 2875bd31f55ecadce99e96c06fb63520
BLAKE2b-256 0298487b6ca9e237440bf0b00931c06e3febec672c7ff9ef6b9b482baca95117

See more details on using hashes here.

File details

Details for the file speech_selib-0.2.0-py3-none-any.whl.

File metadata

  • Download URL: speech_selib-0.2.0-py3-none-any.whl
  • Upload date:
  • Size: 31.5 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.13

File hashes

Hashes for speech_selib-0.2.0-py3-none-any.whl
Algorithm Hash digest
SHA256 ed93142f45170f7f3818bc144b8e8851bc9a99a98591955cc92b0694f2806d24
MD5 35a4d438675f29bb9b7862bd656223fe
BLAKE2b-256 6062aa2d9f8dfe80267ce4d43db6a20172576b7154c882e4fc9b91b690e8c251

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page