Skip to main content

A simple speech enhancement library using ONNX models

Project description

SeLib

SeLib is a small speech-enhancement library for ONNX models. It provides a simple enhance() function, lightweight model wrappers, and utility functions for audio/noise experiments.

Install

Install the package directly from GitHub for the latest version.

pip install git+https://github.com/nipponjo/selib.git

or for the latest release:

pip install speech-selib

Basic Use

The shortest path is selib.enhance(wave, model_id). Pass a NumPy waveform at the sample rate expected by the model.

import librosa
import selib

from selib.utils import NoiseGenerator, snr_mixer

# Load clean speech at the model sample rate.
wave_clean, sr = librosa.load("./data/clean_freesound_33711.wav", sr=48000)

# Create artificial noise and mix it with the clean speech at 5 dB SNR.
noise_gen = NoiseGenerator(sample_rate=sr)
noise = noise_gen.sample(n_samples=len(wave_clean))
wave_noisy = snr_mixer(wave_clean, noise, snr=5)

# Enhance with a registered model id.
# The ONNX model is downloaded and cached automatically on first use.
wave_enhanced = selib.enhance(wave_noisy, "deepfilternet3")

You can also save the enhanced waveform directly. The loaded model is cached by default, so repeated calls with the same model_id do not reload the ONNX session.

# Save as 24-bit PCM WAV while returning the enhanced NumPy array.
wave_enhanced = selib.enhance(
    wave_noisy,
    "deepfilternet3",    
    save_to="enhanced.wav",
    bits_per_sample=24,
)

Model Objects

For more control, load the model wrapper once and call its enhance() method directly.

from selib import load_model
from selib.models import DeepFilterNetOnnx

# load_model() chooses the correct wrapper from the registry metadata.
model: DeepFilterNetOnnx = load_model("deepfilternet3")

# DeepFilterNetOnnx supports extra options such as attenuation limiting.
wave_enhanced = model.enhance(wave_noisy, atten_lim_db=12)

Magnitude-mask models can also be used directly. They predict a mask, multiply it with the noisy magnitude spectrogram, and reuse the noisy phase.

from selib.models import MagnitudeMaskModel

# ul_unas_16k is a 16 kHz magnitude-mask model.
mask_model = MagnitudeMaskModel("ul_unas_16k")
wave_enhanced = mask_model.enhance(wave_noisy_16k)

Metrics

SeLib includes a few quick NumPy-only metrics that do not require PESQ/STOI or other external scoring libraries.

from selib.metrics import snr, segsnr, si_sdr

# Compare clean reference speech against an enhanced waveform.
print("SNR:", snr(wave_clean, wave_enhanced))
print("segSNR:", segsnr(wave_clean, wave_enhanced, sample_rate=sr))
print("SI-SDR:", si_sdr(wave_clean, wave_enhanced))

Available Models

Model ID Type Sample rate #params Paper Repository
deepfilternet3 DeepFilterNet 48 kHz 2.13M arXiv:2305.08227 GitHub
deepfilternet2 DeepFilterNet 48 kHz 2.31M arXiv:2205.05474 GitHub
deepfilternet1 DeepFilterNet 48 kHz 1.78M arXiv:2110.05588 GitHub
ul_unas_16k Magnitude mask 16 kHz 0.171M arXiv:2503.00340 GitHub

DeepFilterNet models use ERB masks plus deep-filter coefficients. Magnitude-mask models predict a spectrogram mask and reconstruct audio with the noisy phase.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

speech_selib-0.2.1.tar.gz (33.1 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

speech_selib-0.2.1-py3-none-any.whl (37.0 kB view details)

Uploaded Python 3

File details

Details for the file speech_selib-0.2.1.tar.gz.

File metadata

  • Download URL: speech_selib-0.2.1.tar.gz
  • Upload date:
  • Size: 33.1 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.13

File hashes

Hashes for speech_selib-0.2.1.tar.gz
Algorithm Hash digest
SHA256 23c922f23dcadfef4eb4bc9f8ec0a572ebb738ff6872048b28a35bb96d9f9234
MD5 ea4cdea9a3f8c7f8051d226fbcb2617c
BLAKE2b-256 2833d067abb3c04e95e109332444c795a45618196d2b0b549555c320fa5287a2

See more details on using hashes here.

File details

Details for the file speech_selib-0.2.1-py3-none-any.whl.

File metadata

  • Download URL: speech_selib-0.2.1-py3-none-any.whl
  • Upload date:
  • Size: 37.0 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.13

File hashes

Hashes for speech_selib-0.2.1-py3-none-any.whl
Algorithm Hash digest
SHA256 f1dcf95aa7f24d3c3f17141a0c2d8de0c1e4727e760a35eec26138972f616581
MD5 2c43a39060294f68bbe05eb85fcecf53
BLAKE2b-256 8b6e9f6371f806a0d276aa4aad1a1ae624dd92bb07972369b7a1471e25674617

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page