Skip to main content

Audio Helper

🇫🇷 · 🇬🇧

CI License: BSD-3-Clause Python Local-first

Audio Helper belongs to a collection of libraries called AI Helpers developed for building Artificial Intelligence.

🌍 AI Helpers

logo

Audio Helper is a Python library that provides utility functions for processing audio files. It includes features like loading audio, converting formats, separating audio sources, and splitting and concatenating audio files.

The Promise

Audio Helper is local-first by design. Three honest cases:

  1. Guaranteed local. Every operation — including the browser GUI at GET /gui — runs on your machine via ffmpeg and local Demucs. Your audio is never uploaded to any third party. There is no telemetry, no account, no SaaS dependency.
  2. The one caveat: model weights. Source separation downloads the Demucs model weights once, on first run (a normal Hugging Face / PyTorch cache fetch). After that it is fully offline. Nothing else needs the network.
  3. Your decision. Nothing here forces the cloud. If you ever want to run behind a proxy or in a container, the FastAPI surface makes that easy — but that is a choice you make, not a default we impose.

Documentation

💻 Documentation

🗺️ Landscape

📋 Examples

Features

  • Audio Loading: load files with optional resampling and mono downmix.
  • Sound Conversion: ffmpeg-backed format/sample-rate/channels conversion.
  • Source Separation: vocals / drums / bass / other via Demucs (optional [demucs] extra).
  • Audio Splitting: fixed-duration chunks and arbitrary [start, end] slices.
  • Concatenation: head-to-tail join into any ffmpeg-supported container.
  • Silent Audio Generation: write silence of a specified duration.
  • Room-Tone Mixing: pink/white/brown ambient noise to mask edits between cuts.
  • Similarity: MFCC-based sound_resemblance score for A/B comparison.
  • Feature Extraction: scipy-based Mel / MFCC primitives.

Four surfaces, one toolkit — every operation above is reachable as:

  • Library: import audio_helper as ah.
  • CLI ×2: audio-helper (argparse, always installed) and audio-helper-click (click twin, [cli] extra) with identical flags.
  • HTTP API: FastAPI app ([api] extra), OpenAPI docs at /docs.
  • GUI: a build-step-free browser Recipe Canvas served at GET /gui — chain the eight verbs into a sequential pipeline, hear every intermediate step (WaveSurfer waveforms), bypass any step for instant A/B, use the ear-first before/after comparator (Space bar toggles), and export the pipeline as a committable recipe.yaml. See GUI.md.

For the exhaustive trigger catalogue, see TRIGGERS.md.

Installation

PrerequisitesPython 3.10–3.13 and git, ffmpeg, cross-platform:

  • 🍎 macOS (Homebrew): brew install python git ffmpeg
  • 🐧 Ubuntu/Debian: sudo apt update && sudo apt install -y python3 python3-pip git ffmpeg
  • 🪟 Windows (PowerShell): winget install Python.Python.3.12 Git.Git Gyan.FFmpeg

We recommend using Python environments. Check this link if you're unfamiliar with setting one up: 🥸 Tech tips.

From PyPI (recommended)

# Core audio utilities only (load, convert, split, concatenate, silent audio, chunks)
pip install audio-helper

# Add source separation (pulls in torch + torchaudio, ~2 GB)
pip install "audio-helper[demucs]"

# Optional surfaces
pip install "audio-helper[cli]"       # click-based CLI twin
pip install "audio-helper[api]"       # FastAPI HTTP surface

From source (no PyPI)

# Core audio utilities only
pip install audio-helper

# Add source separation (pulls in torch + torchaudio, ~2 GB)
pip install "audio-helper[demucs]"

# Optional surfaces
pip install "audio-helper[cli]"
pip install "audio-helper[api]"

If you call separate_sources without the [demucs] extra, the function raises an ImportError pointing you back here.

Usage

For the full catalog of recipes, see 📋 EXAMPLES.md.

Here's an example of how to use Audio Helper to load, convert, and split an audio file:

(download example.mp3 )

It is part of a JFK speech that is badly recorded

import audio_helper as ah

# Load an audio file
audio_file = "example.mp3"
audio, sample_rate = ah.load_audio(audio_file)

# Convert the audio file to a different format
output_audio = "audio_tests/example.wav"
ah.sound_converter(audio_file, output_audio)

# Split the audio file into chunks of 30 seconds
chunks = ah.split_audio_regularly(audio_file, "audio_tests/chunks_folder", split_time=30.0, overwrite = True)
# Concatenate the chunks back together
new_concatenated_audio = "audio_tests/concatenated.wav"
concatenated_audio = ah.audio_concatenation(chunks, output_audio_filename = new_concatenated_audio)

Another cool example is about source separation (DEMUCS from META) with AI separating one audio track into 4 tracks:

  • vocals
  • drums
  • bass
  • other

It works with speech and songs

import audio_helper as ah

audio_path = "input_audio.m4a"

sources = ah.separate_sources(
    audio_path,
    output_folder="audio_tests",
    device = "cpu", # or "cuda" if GPU or nothing to let it decide
    nb_workers = 4, # ignored if not cpu
    output_format = "mp3",
)

print(sources)
# {'vocals': 'audio_tests/vocals.mp3', 'drums': 'audio_tests/drums.mp3', 'bass': 'audio_tests/bass.mp3', 'other': 'audio_tests/other.mp3'}

Multi-surface exposure

audio-helper is not just a library — the same functions are exposed as a CLI and a FastAPI HTTP surface:

# Python library (default)
import audio_helper as ah

# argparse-based CLI (installed automatically)
audio-helper convert --input in.mp3 --output out.wav --freq 44100
audio-helper split --input in.mp3 --output-dir chunks/ --seconds 30
audio-helper separate --input mix.mp3 --output-dir stems/
audio-helper resemblance --a a.mp3 --b b.mp3

# click-based CLI twin (needs the [cli] extra)
pip install "audio-helper[cli]"
# or from source:
pip install "audio-helper[cli]"
audio-helper-click convert --input in.mp3 --output out.wav --freq 44100

# FastAPI HTTP surface (needs the [api] extra)
pip install "audio-helper[api]"
# or from source:
pip install "audio-helper[api]"
uvicorn audio_helper.api:app --port 8000
# → OpenAPI docs at http://localhost:8000/docs

Docker image (light, without Demucs by default):

docker build -t audio-helper .
docker run --rm -p 8000:8000 audio-helper
# with Demucs:
docker build --build-arg WITH_DEMUCS=1 -t audio-helper:demucs .

A minimal browser GUI ("audition bench") ships now — it is served by the FastAPI app at GET /gui (open http://localhost:8000/gui after starting the server). The ambitious future GUI (canvas-based recipe editor, ear-first comparator, MFCC-cluster batch view) is documented as a roadmap in GUI.md.

Author

Acknowledgements

Special thanks to Mohamed Chelali and Bachir Zerroug for fruitful discussions.

License

This project is licensed under the BSD-3-Clause License — see the LICENSE file for details.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

audio_helper-2.0.1.tar.gz (58.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

audio_helper-2.0.1-py3-none-any.whl (50.2 kB view details)

Uploaded Python 3

File details

Details for the file audio_helper-2.0.1.tar.gz.

File metadata

  • Download URL: audio_helper-2.0.1.tar.gz
  • Upload date:
  • Size: 58.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.13.13

File hashes

Hashes for audio_helper-2.0.1.tar.gz
Algorithm Hash digest
SHA256 c5647d706dd93fe25baeef056c01a4f9402a2e29c72f1a20c44dbd7a50dac0aa
MD5 bafb3899c8d6dbcd06bcb7146cada995
BLAKE2b-256 e3dd69ba7965471afe7a3e4cafe641218badb7d1f0e9137d0ae37aac46d7c678

See more details on using hashes here.

File details

Details for the file audio_helper-2.0.1-py3-none-any.whl.

File metadata

  • Download URL: audio_helper-2.0.1-py3-none-any.whl
  • Upload date:
  • Size: 50.2 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.13.13

File hashes

Hashes for audio_helper-2.0.1-py3-none-any.whl
Algorithm Hash digest
SHA256 7fe0b8790348184b2c5fb369be351ceb24fd7335756a8490c747520e8ffe3d6d
MD5 b6e2945d6c60d56c2dc19794f1bff32f
BLAKE2b-256 07ec1509042a7aabaa82710b14b74595edc67f77182d9a3cc56c33a674fee80f

See more details on using hashes here.

Release history Release notifications | RSS feed

2.1.4

2 files

2.1.3

2 files

2.1.2

2 files

2.1.1

2 files

2.1.0

2 files

This release

2.0.1 This release

2 files

2.0.0

2 files

1.6.1

2 files

1.6.0

2 files

1.5.9

2 files

1.5.8

2 files

1.5.7

2 files

1.5.6

2 files

1.5.5

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page