Skip to main content

A simple speech enhancement library using ONNX models

Project description

SeLib

SeLib is a small speech-enhancement library for ONNX models. It provides a simple enhance() function, lightweight model wrappers, and utility functions for audio/noise experiments.

Install

Install the package directly from GitHub while the project is still small and moving quickly.

# Install SeLib and its Python dependencies.
pip install git+https://github.com/nipponjo/selib.git

Basic Use

The shortest path is selib.enhance(wave, model_id). Pass a mono NumPy waveform at the sample rate expected by the model.

import librosa
import selib

from selib.utils import NoiseGenerator, snr_mixer

# Load clean speech at the model sample rate.
wave_clean, sr = librosa.load("./data/clean_freesound_33711.wav", sr=48000)

# Create artificial noise and mix it with the clean speech at 5 dB SNR.
noise_gen = NoiseGenerator(sample_rate=sr)
noise = noise_gen.sample(n_samples=len(wave_clean))
wave_noisy = snr_mixer(wave_clean, noise, snr=5)

# Enhance with a registered model id.
# The ONNX model is downloaded and cached automatically on first use.
wave_enhanced = selib.enhance(wave_noisy, "deepfilternet3")

You can also save the enhanced waveform directly. The loaded model is cached by default, so repeated calls with the same model_id do not reload the ONNX session.

# Save as 24-bit PCM WAV while returning the enhanced NumPy array.
wave_enhanced = selib.enhance(
    wave_noisy,
    "deepfilternet3",    
    save_to="enhanced.wav",
    bits_per_sample=24,
)

Model Objects

For more control, load the model wrapper once and call its enhance() method directly.

from selib import load_model
from selib.models import DeepFilterNetOnnx

# load_model() chooses the correct wrapper from the registry metadata.
model: DeepFilterNetOnnx = load_model("deepfilternet3")

# DeepFilterNetOnnx supports extra options such as attenuation limiting.
wave_enhanced = model.enhance(wave_noisy, atten_lim_db=12)

Magnitude-mask models can also be used directly. They predict a mask, multiply it with the noisy magnitude spectrogram, and reuse the noisy phase.

from selib import MagnitudeMaskModel

# ul_unas_16k is a 16 kHz magnitude-mask model.
mask_model = MagnitudeMaskModel("ul_unas_16k")
wave_enhanced = mask_model.enhance(wave_noisy_16k)

Metrics

SeLib includes a few quick NumPy-only metrics that do not require PESQ/STOI or other external scoring libraries.

from selib.metrics import snr, segsnr, si_sdr

# Compare clean reference speech against an enhanced waveform.
print("SNR:", snr(wave_clean, wave_enhanced))
print("segSNR:", segsnr(wave_clean, wave_enhanced, sample_rate=sr))
print("SI-SDR:", si_sdr(wave_clean, wave_enhanced))

Available Models

Model ID Type Sample rate #params Paper Repository
deepfilternet3 DeepFilterNet 48 kHz 2.13M arXiv:2305.08227 GitHub
deepfilternet2 DeepFilterNet 48 kHz 2.31M arXiv:2205.05474 GitHub
deepfilternet1 DeepFilterNet 48 kHz 1.78M arXiv:2110.05588 GitHub
ul_unas_16k Magnitude mask 16 kHz 0.171M arXiv:2503.00340 GitHub

DeepFilterNet models use ERB masks plus deep-filter coefficients. Magnitude-mask models predict a spectrogram mask and reconstruct audio with the noisy phase.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

speech_selib-0.1.1.tar.gz (27.7 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

speech_selib-0.1.1-py3-none-any.whl (30.6 kB view details)

Uploaded Python 3

File details

Details for the file speech_selib-0.1.1.tar.gz.

File metadata

  • Download URL: speech_selib-0.1.1.tar.gz
  • Upload date:
  • Size: 27.7 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.13

File hashes

Hashes for speech_selib-0.1.1.tar.gz
Algorithm Hash digest
SHA256 50c5fa6880e863972e019b0071a0ab3c224c62305eeb1498a9e7226e7491c0d5
MD5 7afecc295c1e81b98ad2ce11d1dafee0
BLAKE2b-256 55a9ee7609aeff5a63d9fff900d4f3cc84e25bbcc70bc4d7797208a8f877a617

See more details on using hashes here.

File details

Details for the file speech_selib-0.1.1-py3-none-any.whl.

File metadata

  • Download URL: speech_selib-0.1.1-py3-none-any.whl
  • Upload date:
  • Size: 30.6 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.13

File hashes

Hashes for speech_selib-0.1.1-py3-none-any.whl
Algorithm Hash digest
SHA256 d34cc101cf18097b544ed391d6c5472ed75e9ad1ac9db62539f3c122ab8eac91
MD5 886921c9f25923039aea8917a0d8ab9c
BLAKE2b-256 8ba93ac07fda3efa505df9a0fa86719f11c686a592fec70ee7ea50d26f029106

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page