Skip to main content

audio-prep

CI PyPI Docs Python 3.10+ License: MIT Code style: Ruff Type checked: mypy Requires FFmpeg

Convert source audio into pretraining-ready WAV/FLAC, with validation and a dataset manifest as the handoff artifact to the pretraining pipeline. Discovery supports common FFmpeg-readable audio/video containers by default and can scan all files with --extensions all.

Defaults: 16 kHz, mono — the standard input spec for Wav2Vec2 / XLS-R style self-supervised speech pretraining. Override via CLI flags if a different downstream model needs something else.

Repo layout

audio-prep-pipeline/
├── src/audio_prep/
│   ├── config.py        # ConversionConfig: target spec + behavior knobs
│   ├── converter.py      # find_audio_files, convert_file, convert_batch
│   ├── validator.py      # probe_duration, validate_output
│   ├── chunker.py         # ChunkConfig, chunk_file, chunk_batch -- VAD speech chunking
│   ├── profiler.py        # Profiler, ProfileRecord -- runtime resource reports
│   ├── manifest.py       # build_manifest, write_manifest (JSONL output)
│   ├── exceptions.py     # ConversionError, ProbeError, ChunkingError
│   └── cli.py             # `audio-prep convert ...` / `audio-prep chunk ...` entry point
├── tests/
│   ├── conftest.py        # synthetic-audio fixtures (no binary files checked in)
│   ├── test_converter.py
│   ├── test_validator.py
│   ├── test_chunker.py    # fake-detector tests, no real VAD model needed
│   ├── test_manifest.py
│   └── test_cli.py
├── .github/workflows/ci.yml   # lint + typecheck + test matrix (3.10-3.12)
├── .pre-commit-config.yaml     # ruff + mypy + basic hygiene hooks, runs on every commit
├── pyproject.toml               # deps, ruff config, mypy config, pytest config
└── Makefile                      # `make check` runs everything CI runs, locally

Commands

audio-prep has two independent subcommands. Neither depends on the other running first -- both scan --input-dir for supported source files directly.

  • audio-prep convert - conversion only: ffmpeg resample/remix/re-encode into the target WAV/FLAC spec, then validation, then an optional manifest. Does not chunk.
  • audio-prep chunk - chunking only: VAD speech chunking straight from source audio, with Silero by default and Pyannote available for comparison. Does not convert or validate.

convert pipeline stages

  1. discovery (find_audio_files) — recursively find supported source files under an input directory.
  2. conversion (convert_file / convert_batch) — shell out to ffmpeg to resample/remix/re-encode into the target WAV/FLAC spec. Mirrors the input directory's subfolder structure on output. Runs in a process pool since each conversion is an independent subprocess call.
  3. validation (validate_output) — re-opens each converted file with soundfile and checks it actually matches the requested sample rate, channel count, and minimum duration. This catches the case where ffmpeg exits 0 but silently produced something degenerate.
  4. manifest (build_manifest / write_manifest) — JSONL file, one row per source file, recording status (ok / conversion_failed / validation_failed), output path, duration, sample rate, and any error.

Conversion failures don't abort the batch — a bad file in a 50,000-file corpus shows up as one conversion_failed row in the manifest, not a crashed job three hours in.

chunk pipeline stages

  1. discovery (find_audio_files) — recursively find supported source files under an input directory.
  2. chunking (chunk_file / chunk_batch) — runs the selected VAD backend over each file (decoding/resampling via ffmpeg) and splits it into speech-only chunks bounded by a [min, max] duration window, so silence-heavy source recordings don't waste pretraining compute.
  3. manifest (build_chunk_manifest / write_manifest), optional — JSONL file, one row per source file, recording status (ok / chunking_failed), chunk count, and chunk paths.

Chunking failures (e.g. no speech detected) work the same way: chunk_batch returns a ChunkResult per file instead of raising.

Setup

# ffmpeg is a system dependency, not a pip package
sudo apt-get install ffmpeg   # or: brew install ffmpeg

pip install audio-prep-pipeline

Install directly from GitHub:

pip install "audio-prep-pipeline @ git+https://github.com/nattkorat/audio-prep-pipeline.git"

For local development:

make install   # pip install -e ".[dev]" + pre-commit install

The same install provides audio-prep convert, audio-prep chunk, Silero VAD, and Pyannote VAD support. FFmpeg/FFprobe are still system dependencies and must be available on PATH. Some Pyannote models require a Hugging Face token and accepted model terms. If the selected VAD backend cannot load in an offline environment, pass --allow-energy-fallback to use a lower-quality offline detector instead.

Usage

audio-prep convert

CLI:

audio-prep convert \
    --input-dir data/raw_mp3 \
    --output-dir data/wav16k \
    --format wav \
    --sample-rate 16000 \
    --workers 8 \
    --manifest data/manifest.jsonl \
    --profile profiles/convert.json

Python:

from pathlib import Path

from audio_prep import (
    ConversionConfig,
    build_manifest,
    convert_batch,
    validate_output,
    write_manifest,
)

config = ConversionConfig(
    output_format="wav",
    sample_rate=16_000,
    channels=1,
    num_workers=8,
)

results = convert_batch(Path("data/raw_mp3"), Path("data/wav16k"), config)
validations = {
    result.output: validate_output(result.output, config)
    for result in results
    if result.success and result.output is not None
}
records = build_manifest(results, validations)
write_manifest(records, Path("data/manifest.jsonl"))
Flag Default Meaning
--input-dir (required) directory of source audio files
--output-dir (required) where converted output is written
--extensions common FFmpeg audio/video extensions comma-separated source extensions, or all to pass every regular file to FFmpeg
--format wav output format (wav or flac)
--sample-rate 16000 target sample rate
--channels 1 target channel count
--workers 4 parallel conversion workers
--min-duration-sec 0.5 validation fails files shorter than this
--overwrite off re-convert even if output already exists and passes validation
--normalize-loudness off apply EBU R128 loudness normalization (-23 LUFS)
--manifest none path to write a JSONL manifest
--profile none path to write a JSON profile with wall time, CPU time, and peak RSS

audio-prep chunk

Independent of convert -- scans --input-dir for supported source files and runs VAD chunking directly against them, decoding (and resampling, if --sample-rate doesn't match the source) via ffmpeg:

CLI:

audio-prep chunk \
    --input-dir data/raw_mp3 \
    --output-dir data/chunks \
    --sample-rate 16000 \
    --format flac \
    --min-duration-sec 5 \
    --max-duration-sec 20 \
    --workers 4 \
    --manifest data/chunk_manifest.jsonl \
    --profile profiles/chunk-silero.json

Python:

from pathlib import Path

from audio_prep import ChunkConfig, build_chunk_manifest, chunk_batch, write_manifest

config = ChunkConfig(
    min_duration_sec=5,
    max_duration_sec=20,
    output_format="flac",
    sample_rate=16_000,
    num_workers=4,
    vad_backend="silero",
)

results = chunk_batch(Path("data/raw_mp3"), Path("data/chunks"), config)
records = build_chunk_manifest(results)
write_manifest(records, Path("data/chunk_manifest.jsonl"))
Flag Default Meaning
--input-dir (required) directory of source audio files to scan and chunk
--output-dir <input-dir>/chunks where chunks are written
--extensions common FFmpeg audio/video extensions comma-separated source extensions, or all to pass every regular file to FFmpeg
--format wav output chunk format (wav or flac)
--sample-rate 16000 resample (via ffmpeg) to this rate before chunking if the source doesn't already match it
--min-duration-sec 5.0 drop chunks shorter than this
--max-duration-sec 20.0 split longer speech into windows this size
--workers 4 parallel chunking workers
--overwrite off re-chunk even if valid output already exists
--vad-backend silero speech detector backend: silero, pyannote, or energy
--pyannote-model pyannote/voice-activity-detection Hugging Face model id used with --vad-backend pyannote
--hf-token none Hugging Face token for gated Pyannote models; falls back to HF_TOKEN or HUGGINGFACE_TOKEN
--allow-energy-fallback off fall back to a low-quality energy detector if the selected VAD can't load, instead of raising
--manifest none path to write a JSONL chunk manifest (source file, status, chunk count/paths)
--profile none path to write a JSON profile with wall time, CPU time, and peak RSS

Compare VAD backends with the same input/settings:

audio-prep chunk \
    --input-dir data/raw_mp3 \
    --output-dir data/chunks-pyannote \
    --vad-backend pyannote \
    --pyannote-model pyannote/voice-activity-detection \
    --hf-token "$HF_TOKEN" \
    --profile profiles/chunk-pyannote.json

Python profiling:

from pathlib import Path

from audio_prep import ChunkConfig, Profiler, chunk_batch

profiler = Profiler()
config = ChunkConfig(vad_backend="pyannote", sample_rate=16_000, num_workers=1)

with profiler.measure("chunk", {"vad_backend": config.vad_backend}):
    results = chunk_batch(Path("data/raw_mp3"), Path("data/chunks-pyannote"), config)

profiler.write_json(
    Path("profiles/chunk-pyannote.json"),
    operation="chunk",
    metadata={"files": len(results), "vad_backend": config.vad_backend},
)

chunk_file itself doesn't care about source extension -- it decodes whatever path it's given -- so the Python API can also chunk an existing WAV/FLAC corpus (e.g. convert_batch output) by passing source_files explicitly instead of relying on chunk_batch discovery. This is an advanced Python API case:

from pathlib import Path

from audio_prep import ChunkConfig, ConversionConfig, build_manifest, chunk_batch
from audio_prep import convert_batch, validate_output

config = ConversionConfig(output_format="wav", sample_rate=16_000, channels=1, num_workers=8)
results = convert_batch(Path("data/raw_mp3"), Path("data/wav16k"), config)
validations = {
    result.output: validate_output(result.output, config)
    for result in results
    if result.success and result.output is not None
}
records = build_manifest(results, validations)

# source_files bypasses chunk_batch discovery, so this works
# directly against the already-converted WAV output above.
valid_outputs = [path for path, v in validations.items() if v.valid]
chunk_config = ChunkConfig(min_duration_sec=5, max_duration_sec=20, num_workers=4)
chunk_results = chunk_batch(
    Path("data/wav16k"),
    Path("data/wav16k/chunks"),
    chunk_config,
    source_files=valid_outputs,
)

Development workflow

make format     # ruff format + autofix
make lint       # ruff check
make typecheck  # mypy --strict
make test       # pytest with coverage
make check      # all of the above -- run this before opening a PR

pre-commit (installed via make install) runs ruff + mypy + basic hygiene checks automatically on every commit. CI (.github/workflows/ci.yml) re-runs the same checks plus the full test matrix across Python 3.10–3.12 on every push and PR.

Extending this

Natural next additions, in roughly the order they'd come up:

  • Streaming manifest writes for very large corpora, instead of holding all ConversionResults in memory before writing.

Note: This is the template, that you have to extend from. Main branch is protected so you have to create another branch to work on.

Metadata

Release files for audio-prep-pipeline 0.1.3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for audio-prep-pipeline 0.1.3
File Size Uploaded
audio_prep_pipeline-0.1.3.tar.gz 43.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for audio-prep-pipeline 0.1.3
File Interpreter ABI Platform
audio_prep_pipeline-0.1.3-py3-none-any.whl Python 3 none any Details

Total release size: 69.7 kB

Release files / audio_prep_pipeline-0.1.3.tar.gz

Download URL audio_prep_pipeline-0.1.3.tar.gz
Size 43.3 kB
Tags Source
SHA-256 checksum
How to use checksums
2afb5a1097cf84ac1a796751ee5d69330ce0ac32a65eea5aa52704c04abc4507
BLAKE2b-256 checksum
How to use checksums
c5a8e329b7260d6a346051d64fd4d1c39aeeea160abf9c3671381224510f6c47
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.8

Release files / audio_prep_pipeline-0.1.3-py3-none-any.whl

Download URL audio_prep_pipeline-0.1.3-py3-none-any.whl
Size 26.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
8ff06bdffcd758e9939630e26b62a7a18b13873bb7c37db6a52d6125353f4925
BLAKE2b-256 checksum
How to use checksums
6808d75c79798c7b182092045bcc35230e42c3f60d53c93d725ec7a69534cd6d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.8

Release history Release notifications | RSS feed

0.1.7

2 release files

0.1.6

2 release files

0.1.5

2 release files

0.1.4

2 release files

This release

0.1.3 This release

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page