Skip to main content

Kudio

PyPI Python License: MIT

Lightweight, composable audio toolkit for acoustic inspection.

Kudio gives you audio I/O, streaming/recording, feature extraction, noisy-data synthesis, effects & augmentation, enhancement, and quality metrics — as small composable functions with a clean, typed API.

The core stays light (numpy + soundfile + librosa); heavy or niche features (plotting, dataframes, PESQ/STOI, live audio) are opt-in extras.

Install

pip install kudio            # core: I/O, features, effects, synthesis, metrics
pip install kudio[audio]     # + live capture / playback (PyAudio, sounddevice)
pip install kudio[eval]      # + PESQ / STOI / SDR back-ends
pip install kudio[viz]       # + plotting helpers (matplotlib)
pip install kudio[data]      # + Excel / dataframe export (pandas, openpyxl)
pip install kudio[all]       # everything

Requires Python 3.9+ (tested on CPython 3.9–3.13).

The Python badge above reflects the version currently published on PyPI; it updates automatically once v3.0 is released.

Quick start

import kudio

# --- I/O (soundfile-backed; float32) --------------------------------------
y, sr = kudio.file_load("clip.wav")             # native rate, no resample
kudio.save_wave("out.wav", y, sr, subtype="PCM_24")   # 16/24-bit or FLOAT

# --- STFT geometry as a value ---------------------------------------------
stft = kudio.STFT(sr=16000, n_fft=512, hop_length=256)
spec = stft.forward(y)                                # (frames, bins)
back = stft.inverse(y, spec)                          # phase taken from y
stft.to_dict()                                        # store beside a checkpoint

# --- features -------------------------------------------------------------
spec = kudio.waveform_to_spectrogram(y, norm=True)    # log-power spectrogram
feat = kudio.mfcc(y, sr, n_mfcc=40)                    # (1, frames, 40)
mel  = kudio.melspectrogram("clip.wav", n_mels=64)     # path or waveform…
mel  = kudio.melspectrogram(y, sr=sr, n_mels=64)       # …arrays need sr

# --- model-facing helpers -------------------------------------------------
x   = kudio.stack_context(spec, context=2)            # ±2 frames per row
win = kudio.frame_windows(spec, n_frames=64)          # (n, 64, bins)
std = kudio.Standardizer().fit(spec)                  # save it with the model
z   = std.transform(spec); spec_again = std.inverse(z)

# --- loudness (BS.1770) ---------------------------------------------------
lufs = kudio.loudness(y, sr)                          # what it sounds like,
y    = kudio.normalize_lufs(y, sr, lufs=-23.0)        # not what its peak is
fair = kudio.match_loudness(enhanced, y, sr)          # before you A/B them

# --- is this recording usable at all? no reference needed -----------------
for problem in kudio.audio_report(y, sr).problems():
    print(problem)     # "content stops at 3812 Hz although the file claims
                       #  16000 Hz — very likely upsampled from 8000 Hz"

# --- effects & augmentation ----------------------------------------------
y    = kudio.normalize(y, peak=0.99)                  # or normalize_db(y, -1)
y16  = kudio.resample(y, orig_sr=8000, target_sr=16000)
clip, (a, b) = kudio.trim_silence(y, top_db=30)       # isolate the event
segments      = kudio.split_on_silence(y)
aug = kudio.pitch_shift(y, sr, n_steps=2)
aug = kudio.add_noise_snr(y, noise, snr_db=5, seed=0) # reproducible
spec = kudio.spec_augment(spec, freq_mask_width=8, time_mask_width=16, seed=0)

# --- noisy-data synthesis -------------------------------------------------
syx = kudio.Synthesizer("data/clean", "data/noise",
                        out_path="data/mixed", snr_ratio=(-5, 0, 5))
syx.syn(mode="inc", seed=17)                          # reproducible
syx.syn(mode="inc", target_sr=16000)                  # or resample the output

# --- metrics (pure numpy, no extra deps) ----------------------------------
print(kudio.si_sdr(ref, est), kudio.snr(ref, est), kudio.segmental_snr(ref, est))

# --- recording (needs kudio[audio]) --------------------------------------
mics = kudio.list_devices("input")                    # sounddevice indices
wave = kudio.record(seconds=3, sr=16000, device=mics[0]["index"])

Device indices are backend-specific. record() runs on sounddevice, so its device= comes from kudio.list_devices(). The streaming classes (Recorder, LocalStreamReader) run on PyAudio and take indices from CheckDevice.system_devices(). The two are numbered independently — an index from one must never be handed to the other. kudio devices prints both, labelled.

Command line

kudio devices                                  # list audio devices
kudio info clip.wav                            # sr / channels / duration / peak
kudio convert in.wav out.wav --rate 16000 --subtype PCM_16
kudio synth --clean C --noise N --out O --snr -5 0 5 --seed 17
kudio trim in.wav out.wav --top-db 30

Modules

Module What's inside
kudio.core.io file_load, save_wave, resample, check_input, load_waves, copy_waves
kudio.core.stft STFT — geometry + forward/inverse, storable next to a model
kudio.core.feature waveform_to_spectrogram, spectrogram_to_waveform, mfcc, melspectrogram, stack_context, frame_windows, Standardizer, ...
kudio.effects trim_silence, split_on_silence, time_stretch, pitch_shift, normalize, add_noise_snr, reverb, spec_augment
kudio.core.loudness loudness, normalize_lufs, match_loudness — ITU-R BS.1770-4, the perceptual answer normalize's peak scaling cannot give
kudio.core.report audio_report → AudioReport — clipping, DC, silence, noise floor, real bandwidth, with no clean reference needed
kudio.core.synth Synthesizer (SNR mixing, seedable)
kudio.core.evaluator si_sdr, snr, segmental_snr (dep-free); AudioEvaluate (PESQ/STOI/SDR); check_metrics_install
kudio.core.stream record, play_audio, Recorder, LocalStreamReader, RemoteStreamReader
kudio.util list_devices, CheckDevice, map_waves, colored console helpers, timers

Everything commonly used is importable straight from the top level (kudio.…).

Errors

All library errors derive from kudio.KudioError (AudioIOError, FeatureError, DeviceError, SynthesisError, DependencyError), so you can catch them in one place. Missing an optional extra raises a DependencyError that tells you exactly what to install.

Logging

kudio uses the standard logging module (loggers named kudio.*) and prints nothing by default:

import logging; logging.basicConfig(level=logging.INFO)

Migrating from v2

v3 is a cleanup release. The short cryptic names still work but now emit a DeprecationWarning — switch to the canonical names:

v2 (deprecated) v3 canonical
w2s, wavform2spec waveform_to_spectrogram
spec2wavform spectrogram_to_waveform
f2s, wav2spec file_to_spectrogram
s2w, spec2wav save_spectrogram_as_wave
w2mfcc, wav2mfcc mfcc
concat_mfcc_ mfcc_from_files
concat_logspec_, contextual_LogSpectrogram logspec_from_files
wav2mel melspectrogram

Also: matplotlib / pandas / openpyxl are no longer installed by default (use the [viz] / [data] extras), and the deprecated _config / _enh / version compatibility shims were removed.

License

MIT — see LICENSE.txt.

Metadata

Release files for Kudio 3.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for Kudio 3.3.0
File Size Uploaded
kudio-3.3.0.tar.gz 70.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for Kudio 3.3.0
File Interpreter ABI Platform
kudio-3.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 135.2 kB

Release files / kudio-3.3.0.tar.gz

Download URL kudio-3.3.0.tar.gz
Size 70.2 kB
Tags Source
SHA-256 checksum
How to use checksums
00c2e0e8c47106aaf105497407bfa337e07a7f7fe1e3b8305148e6264ca89502
BLAKE2b-256 checksum
How to use checksums
8fd69ba787e3afe1fc1c22d8321de61ce1212f2c9426c718293ee67994de08ef
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 15, 2026.

Transparency log

Release files / kudio-3.3.0-py3-none-any.whl

Download URL kudio-3.3.0-py3-none-any.whl
Size 64.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
c87f37763560e4e0352781be5c7901ac21a79a590d0eed85c8eeb2e944210b5d
BLAKE2b-256 checksum
How to use checksums
fd3b0cc86f767265e6fa9f48f1d523ca1681e3f370baf70610e9233720ab80d5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 15, 2026.

Transparency log

Release history Release notifications | RSS feed

3.7.2

2 release files

3.7.1

2 release files

3.7.0

2 release files

3.6.0

2 release files

3.5.0

2 release files

This release

3.3.0 This release

2 release files

3.2.0

2 release files

3.0.1

2 release files

3.0.0

2 release files

2.0

2 release files

1.1.4.4

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page