Skip to main content

audio-prep

CI PyPI Docs Python 3.10+ License: MIT Code style: Ruff Type checked: mypy Requires FFmpeg

Convert source audio into pretraining-ready WAV/FLAC, with validation and a dataset manifest as the handoff artifact to the pretraining pipeline. Discovery supports common FFmpeg-readable audio/video containers by default and can scan all files with --extensions all.

Defaults: 16 kHz, mono — the standard input spec for Wav2Vec2 / XLS-R style self-supervised speech pretraining. Override via CLI flags if a different downstream model needs something else.

Repo layout

audio-prep-pipeline/
├── src/audio_prep/
│   ├── config.py        # ConversionConfig: target spec + behavior knobs
│   ├── converter.py      # find_audio_files, convert_file, convert_batch
│   ├── validator.py      # probe_duration, validate_output
│   ├── chunker.py         # ChunkConfig, chunk_file, chunk_batch -- VAD speech chunking
│   ├── profiler.py        # Profiler, ProfileRecord -- runtime resource reports
│   ├── manifest.py       # build_manifest, write_manifest (JSONL output)
│   ├── exceptions.py     # ConversionError, ProbeError, ChunkingError
│   └── cli.py             # `audio-prep convert ...` / `audio-prep chunk ...` entry point
├── tests/
│   ├── conftest.py        # synthetic-audio fixtures (no binary files checked in)
│   ├── test_converter.py
│   ├── test_validator.py
│   ├── test_chunker.py    # fake-detector tests, no real VAD model needed
│   ├── test_manifest.py
│   └── test_cli.py
├── .github/workflows/ci.yml   # lint + typecheck + test matrix (3.10-3.12)
├── .pre-commit-config.yaml     # ruff + mypy + basic hygiene hooks, runs on every commit
├── pyproject.toml               # deps, ruff config, mypy config, pytest config
└── Makefile                      # `make check` runs everything CI runs, locally

Commands

audio-prep has two independent subcommands. Neither depends on the other running first -- both scan --input-dir for supported source files directly.

  • audio-prep convert - conversion only: ffmpeg resample/remix/re-encode into the target WAV/FLAC spec, then validation, then an optional manifest. Does not chunk.
  • audio-prep chunk - chunking only: VAD speech chunking straight from source audio, with Silero by default and Pyannote available for comparison. Does not convert or validate.

convert pipeline stages

  1. discovery (find_audio_files) — recursively find supported source files under an input directory.
  2. conversion (convert_file / convert_batch) — shell out to ffmpeg to resample/remix/re-encode into the target WAV/FLAC spec. Mirrors the input directory's subfolder structure on output. Runs in a process pool since each conversion is an independent subprocess call.
  3. validation (validate_output) — re-opens each converted file with soundfile and checks it actually matches the requested sample rate, channel count, and minimum duration. This catches the case where ffmpeg exits 0 but silently produced something degenerate.
  4. manifest (build_manifest / write_manifest) — JSONL file, one row per source file, recording status (ok / conversion_failed / validation_failed), output path, duration, sample rate, and any error.

Conversion failures don't abort the batch — a bad file in a 50,000-file corpus shows up as one conversion_failed row in the manifest, not a crashed job three hours in.

chunk pipeline stages

  1. discovery (find_audio_files) — recursively find supported source files under an input directory.
  2. chunking (chunk_file / chunk_batch) — runs the selected VAD backend over each file (decoding/resampling via ffmpeg) and splits it into speech-focused chunks near a [min, max] duration window, so silence-heavy source recordings don't waste pretraining compute.
  3. manifest (build_chunk_manifest / write_manifest), optional — JSONL file, one row per source file, recording status (ok / chunking_failed), chunk count, and chunk paths.

Chunking failures (e.g. no speech detected) work the same way: chunk_batch returns a ChunkResult per file instead of raising.

Setup

# ffmpeg is a system dependency, not a pip package
sudo apt-get install ffmpeg   # or: brew install ffmpeg

pip install audio-prep-pipeline

Install directly from GitHub:

pip install "audio-prep-pipeline @ git+https://github.com/nattkorat/audio-prep-pipeline.git"

For local development:

make install   # pip install -e ".[dev]" + pre-commit install

The same install provides audio-prep convert, audio-prep chunk, Silero VAD, and Pyannote VAD support. FFmpeg/FFprobe are still system dependencies and must be available on PATH. Some Pyannote models require a Hugging Face token and accepted model terms. If the selected VAD backend cannot load in an offline environment, pass --allow-energy-fallback to use a lower-quality offline detector instead.

Usage

audio-prep convert

CLI:

audio-prep convert \
    --input-dir data/raw_mp3 \
    --output-dir data/wav16k \
    --format wav \
    --sample-rate 16000 \
    --workers 8 \
    --manifest data/manifest.jsonl \
    --profile profiles/convert.json

Python:

from pathlib import Path

from audio_prep import (
    ConversionConfig,
    build_manifest,
    convert_batch,
    validate_output,
    write_manifest,
)

config = ConversionConfig(
    output_format="wav",
    sample_rate=16_000,
    channels=1,
    num_workers=8,
)

results = convert_batch(Path("data/raw_mp3"), Path("data/wav16k"), config)
validations = {
    result.output: validate_output(result.output, config)
    for result in results
    if result.success and result.output is not None
}
records = build_manifest(results, validations)
write_manifest(records, Path("data/manifest.jsonl"))
Flag Default Meaning
--input-dir (required) directory of source audio files
--output-dir (required) where converted output is written
--extensions common FFmpeg audio/video extensions comma-separated source extensions, or all to pass every regular file to FFmpeg
--format wav output format (wav or flac)
--sample-rate 16000 target sample rate
--channels 1 target channel count
--workers 4 parallel conversion workers
--min-duration-sec 0.5 validation fails files shorter than this
--overwrite off re-convert even if output already exists and passes validation
--normalize-loudness off apply EBU R128 loudness normalization (-23 LUFS)
--manifest none path to write a JSONL manifest
--profile none path to write a JSON profile with wall time, CPU time, and peak RSS

audio-prep chunk

Independent of convert -- scans --input-dir for supported source files and runs VAD chunking directly against them, decoding (and resampling, if --sample-rate doesn't match the source) via ffmpeg:

CLI:

audio-prep chunk \
    --input-dir data/raw_mp3 \
    --output-dir data/chunks \
    --sample-rate 16000 \
    --format flac \
    --min-duration-sec 5 \
    --max-duration-sec 20 \
    --merge-gap-sec 1.0 \
    --workers 4 \
    --manifest data/chunk_manifest.jsonl \
    --profile profiles/chunk-silero.json

Python:

from pathlib import Path

from audio_prep import ChunkConfig, build_chunk_manifest, chunk_batch, write_manifest

config = ChunkConfig(
    min_duration_sec=5,
    max_duration_sec=20,
    merge_gap_sec=1.0,
    output_format="flac",
    sample_rate=16_000,
    num_workers=4,
    vad_backend="silero",
)

results = chunk_batch(Path("data/raw_mp3"), Path("data/chunks"), config)
records = build_chunk_manifest(results)
write_manifest(records, Path("data/chunk_manifest.jsonl"))
Flag Default Meaning
--input-dir (required) directory of source audio files to scan and chunk
--output-dir <input-dir>/chunks where chunks are written
--extensions common FFmpeg audio/video extensions comma-separated source extensions, or all to pass every regular file to FFmpeg
--format wav output chunk format (wav or flac)
--sample-rate 16000 resample (via ffmpeg) to this rate before chunking if the source doesn't already match it
--min-duration-sec 5.0 minimum target chunk duration after nearby speech spans are packed
--max-duration-sec 20.0 maximum target chunk duration; long spans are split evenly to avoid short tails
--merge-gap-sec 1.0 silence budget for packing nearby VAD speech spans before applying duration rules
--workers 4 parallel chunking workers
--overwrite off re-chunk even if valid output already exists
--vad-backend silero speech detector backend: silero, pyannote, or energy
--pyannote-model pyannote/speaker-diarization-community-1 Hugging Face model id used with --vad-backend pyannote
--pyannote-revision none optional Hugging Face model revision; repo/model@revision is also accepted
--hf-token none Hugging Face token for gated Pyannote models; falls back to HF_TOKEN or HUGGINGFACE_TOKEN
--allow-energy-fallback off fall back to a low-quality energy detector if the selected VAD can't load, instead of raising
--manifest none path to write a JSONL chunk manifest (source file, status, chunk count/paths)
--profile none path to write a JSON profile with wall time, CPU time, and peak RSS

Compare VAD backends with the same input/settings:

audio-prep chunk \
    --input-dir data/raw_mp3 \
    --output-dir data/chunks-pyannote \
    --vad-backend pyannote \
    --pyannote-model pyannote/speaker-diarization-community-1 \
    --hf-token "$HF_TOKEN" \
    --profile profiles/chunk-pyannote.json

Python profiling:

from pathlib import Path

from audio_prep import ChunkConfig, Profiler, chunk_batch

profiler = Profiler()
config = ChunkConfig(vad_backend="pyannote", sample_rate=16_000, num_workers=1)

with profiler.measure("chunk", {"vad_backend": config.vad_backend}):
    results = chunk_batch(Path("data/raw_mp3"), Path("data/chunks-pyannote"), config)

profiler.write_json(
    Path("profiles/chunk-pyannote.json"),
    operation="chunk",
    metadata={"files": len(results), "vad_backend": config.vad_backend},
)

chunk_file itself doesn't care about source extension -- it decodes whatever path it's given -- so the Python API can also chunk an existing WAV/FLAC corpus (e.g. convert_batch output) by passing source_files explicitly instead of relying on chunk_batch discovery. This is an advanced Python API case:

from pathlib import Path

from audio_prep import ChunkConfig, ConversionConfig, build_manifest, chunk_batch
from audio_prep import convert_batch, validate_output

config = ConversionConfig(output_format="wav", sample_rate=16_000, channels=1, num_workers=8)
results = convert_batch(Path("data/raw_mp3"), Path("data/wav16k"), config)
validations = {
    result.output: validate_output(result.output, config)
    for result in results
    if result.success and result.output is not None
}
records = build_manifest(results, validations)

# source_files bypasses chunk_batch discovery, so this works
# directly against the already-converted WAV output above.
valid_outputs = [path for path, v in validations.items() if v.valid]
chunk_config = ChunkConfig(
    min_duration_sec=5,
    max_duration_sec=20,
    merge_gap_sec=1.0,
    num_workers=4,
)
chunk_results = chunk_batch(
    Path("data/wav16k"),
    Path("data/wav16k/chunks"),
    chunk_config,
    source_files=valid_outputs,
)

Development workflow

make format     # ruff format + autofix
make lint       # ruff check
make typecheck  # mypy --strict
make test       # pytest with coverage
make check      # all of the above -- run this before opening a PR

pre-commit (installed via make install) runs ruff + mypy + basic hygiene checks automatically on every commit. CI (.github/workflows/ci.yml) re-runs the same checks plus the full test matrix across Python 3.10–3.12 on every push and PR.

Extending this

Natural next additions, in roughly the order they'd come up:

  • Streaming manifest writes for very large corpora, instead of holding all ConversionResults in memory before writing.

Note: This is the template, that you have to extend from. Main branch is protected so you have to create another branch to work on.

Metadata

Release files for audio-prep-pipeline 0.1.7

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for audio-prep-pipeline 0.1.7
File Size Uploaded
audio_prep_pipeline-0.1.7.tar.gz 586.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for audio-prep-pipeline 0.1.7
File Interpreter ABI Platform
audio_prep_pipeline-0.1.7-py3-none-any.whl Python 3 none any Details

Total release size: 613.7 kB

Release files / audio_prep_pipeline-0.1.7.tar.gz

Download URL audio_prep_pipeline-0.1.7.tar.gz
Size 586.0 kB
Tags Source
SHA-256 checksum
How to use checksums
67a058c7ad6becc985a25e5b152311f35ec5b8c1573f1f3e94dd9af8e6cca5a5
BLAKE2b-256 checksum
How to use checksums
d21970a4458da58eb0eb88e6d77d9490070f03ad04d5b9b0a874279a6719845a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.8

Release files / audio_prep_pipeline-0.1.7-py3-none-any.whl

Download URL audio_prep_pipeline-0.1.7-py3-none-any.whl
Size 27.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
fe09173d9c2c07ca5141c0e74c53b4d83e3c0b0a5cbb87ab1e9f7c84f3793b77
BLAKE2b-256 checksum
How to use checksums
2f84e19af18a4b6b2bf45602dc6f9f8a3d960c8f357d17c3f0528f83fd6a207e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.8

Release history Release notifications | RSS feed

This release

0.1.7 This release

2 release files

0.1.6

2 release files

0.1.5

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page