Skip to main content

Pure-ONNX multi-engine audio super-resolution / bandwidth extension (upscale speech to 48 kHz)

Project description

audiosronnx

Speech restoration, denoising and bandwidth extension in pure ONNX. Clean up noisy recordings, and upscale narrowband speech — 8 kHz telephony, muffled captures, low-fidelity TTS — to a full 48 kHz signal.

Inference runs entirely on onnxruntime. There is no torch at runtime: every STFT, filterbank and resampler lives in numpy/scipy, so the ONNX graphs stay portable and the front-ends stay inspectable.

pip install audiosronnx
from audiosronnx import load_denoise, load_sr

clean, rate = load_denoise("dpdfnet").denoise("noisy_call.wav")   # remove noise
wide,  _    = load_sr("lavasr").upscale(clean, rate)              # extend to 48 kHz
audiosronnx denoise noisy.wav clean.wav
audiosronnx upscale clean.wav wide_48k.wav

Runtime dependencies are numpy, onnxruntime, scipy, soundfile and huggingface_hub. Weights download on first use and cache under ~/.local/share/audiosronnx, pinned by revision.

Two jobs

Denoising removes background noise and keeps the sample rate. Bandwidth extension invents high-frequency content that was never in the input and always outputs 48 kHz. They have separate entry points — load_denoise() and load_sr() — which refuse each other's engines.

They compose, in that order: denoise first, then extend. Run them the other way round and the extender reconstructs a high band out of the noise.

Denoisers

Engine Rate Size License Reach for it when
dpdfnet (default) 8 / 16 / 48 kHz 8.7–14.9 MB Apache-2.0 you have no other reason to choose
mossformer2 48 kHz 229 MB Apache-2.0 the source is fullband and quality wins
frcrn 16 kHz 57.5 MB Apache-2.0 reproducing published benchmarks
mpsenet 16 kHz 9.7 MB MIT you want a spectral model that stays small
gtcrn 16 kHz 0.54 MB MIT footprint is the binding constraint
cmgan 16 kHz 7.8 MB MIT you want a conformer metric-GAN specifically
metadenoiser 16 kHz 19–34 MB CC-BY-NC-4.0 you want a time-domain model and NC is acceptable
mossformergan 16 kHz 17.7 MB Apache-2.0 highest published PESQ (3.47)
deepfilternet 48 kHz ~2 MB MIT you already depend on libdf

SNR recovered on one clip at 19 / 11 / 5 dB input SNR: dpdfnet +4.9 / +10.3 / +13.7 dB, mossformer2 +5.9 / +10.4 / +13.4, mpsenet +4.8 / +8.9 / +11.7, frcrn +4.6 / +8.8 / +11.7, gtcrn +3.5 / +7.3 / +7.5. That is broadband Gaussian noise — a hostile synthetic case that ranks engines consistently but predicts little about babble or codec artefacts.

Recommendations: start with dpdfnet — no extra dependencies, 8/16/48 kHz, top-tier measured gain. Take mossformer2 for the best fullband quality, mossformergan for the best benchmark scores, gtcrn when footprint is the constraint, metadenoiser for a time-domain model. Engines are kept even when something else beats them, so published results stay reproducible — cmgan is dominated and stays for that reason.

Full detail in docs/denoising.md.

Bandwidth extension

Engine Input Size CPU speed License
lavasr (default) 8–48 kHz ~52 MB ~50× realtime Apache-2.0
novasr 16 kHz ~0.2 MB ~1000× realtime Apache-2.0
flowhigh any ~200 MB slow (4 ODE steps) MIT
hifiganbwe any ~4 MB fast MIT
apbwe any (12 kHz band) ~120 MB moderate MIT
sidon 16 kHz ~410 MB ~0.6× realtime MIT
callenhancer 8–16 kHz ~3 GB / ~1.3 GB int8 ~0.3–0.4× realtime CC-BY-NC-4.0

lavasr is the general-purpose choice. sidon and callenhancer do something different: they resynthesise speech rather than extend its band, repairing codec damage a bandwidth extender cannot, at a much higher cost. callenhancer is trained on telephony specifically, and its weights are non-commercial.

Full detail in docs/engines.md.

Documentation

Page Contents
docs/engines.md Every bandwidth-extension engine, options, precision
docs/denoising.md Every denoiser, and how to choose
docs/api.md load_sr, load_denoise, the model contracts, the registry
docs/cli.md Command-line reference
docs/custom-engines.md Registering your own engine
docs/conversion.md How the ONNX graphs are produced and validated
docs/not-shipped.md Models evaluated and rejected, with reasons

Runnable scripts are in examples/.

API

sr = load_sr("lavasr")
out, rate = sr.upscale("telephone_8k.wav")   # path, int16 bytes, or numpy array
sr.upscale_file("in.wav", "out_48k.wav")
sr.upscale_dir("clips_in/", "clips_out/")

dn = load_denoise("dpdfnet")
clean, rate = dn.denoise(audio, in_rate)     # returns the ENGINE's rate
dn.denoise_file("in.wav", "clean.wav")
dn.denoise_dir("clips_in/", "clips_out/")

available_models(); available_denoisers(); get_engine(alias); register_engine(entry)

Output is always mono float32. Empty, silent and non-finite input produce finite output rather than raising. See docs/api.md.

Which engines ship

An engine ships when it exports to a single static ONNX graph, runs on CPU through onnxruntime, carries a permissive license, and is validated end-to-end against the upstream pipeline — not merely against the model, since reproducing a network but not its surrounding normalisation yields audio that sounds plausible and is quietly wrong.

That rules out diffusion and flow-matching models, location-variable convolutions, and unlicensed checkpoints. docs/not-shipped.md names every candidate considered and which condition it failed.

License

Apache-2.0.

Model weights carry their upstream licenses, listed per engine in the docs and reported by audiosronnx list. Most are MIT or Apache-2.0; callenhancer is CC-BY-NC-4.0, which covers the weights rather than audio processed with them.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

audiosronnx-0.8.0a1.tar.gz (76.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

audiosronnx-0.8.0a1-py3-none-any.whl (77.2 kB view details)

Uploaded Python 3

File details

Details for the file audiosronnx-0.8.0a1.tar.gz.

File metadata

  • Download URL: audiosronnx-0.8.0a1.tar.gz
  • Upload date:
  • Size: 76.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for audiosronnx-0.8.0a1.tar.gz
Algorithm Hash digest
SHA256 324d1f4998efb2fd73342bc745d6adf27fd45372f9d84f7e5606b329b6107088
MD5 b4282ceacd33dcc23ee94ef9cb807a08
BLAKE2b-256 0e24542f6ba63089b294aa2cf8edc5f5683575e40ef12efba2a67113e944e10d

See more details on using hashes here.

File details

Details for the file audiosronnx-0.8.0a1-py3-none-any.whl.

File metadata

  • Download URL: audiosronnx-0.8.0a1-py3-none-any.whl
  • Upload date:
  • Size: 77.2 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for audiosronnx-0.8.0a1-py3-none-any.whl
Algorithm Hash digest
SHA256 c4ca357dae27528f3506b16ed370420ff32e8693c071dda442bf41415dc27940
MD5 a4ffc6439b030824a7ac66b9bda0d654
BLAKE2b-256 8efed6a485923f5e83f315a14180e10188e68a3d5889fe4b035c6660e51c37b9

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page