audio-prep
Convert source audio into pretraining-ready WAV/FLAC, with validation and a
dataset manifest as the handoff artifact to the pretraining pipeline. Discovery
supports common FFmpeg-readable audio/video containers by default and can scan
all files with --extensions all.
Defaults: 16 kHz, mono — the standard input spec for Wav2Vec2 / XLS-R style self-supervised speech pretraining. Override via CLI flags if a different downstream model needs something else.
Repo layout
audio-prep-pipeline/
├── src/audio_prep/
│ ├── config.py # ConversionConfig: target spec + behavior knobs
│ ├── converter.py # find_audio_files, convert_file, convert_batch
│ ├── validator.py # probe_duration, validate_output
│ ├── chunker.py # ChunkConfig, chunk_file, chunk_batch -- VAD speech chunking
│ ├── profiler.py # Profiler, ProfileRecord -- runtime resource reports
│ ├── manifest.py # build_manifest, write_manifest (JSONL output)
│ ├── exceptions.py # ConversionError, ProbeError, ChunkingError
│ └── cli.py # `audio-prep convert ...` / `audio-prep chunk ...` entry point
├── tests/
│ ├── conftest.py # synthetic-audio fixtures (no binary files checked in)
│ ├── test_converter.py
│ ├── test_validator.py
│ ├── test_chunker.py # fake-detector tests, no real VAD model needed
│ ├── test_manifest.py
│ └── test_cli.py
├── .github/workflows/ci.yml # lint + typecheck + test matrix (3.10-3.12)
├── .pre-commit-config.yaml # ruff + mypy + basic hygiene hooks, runs on every commit
├── pyproject.toml # deps, ruff config, mypy config, pytest config
└── Makefile # `make check` runs everything CI runs, locally
Commands
audio-prep has two independent subcommands. Neither depends on the other
running first -- both scan --input-dir for supported source files directly.
audio-prep convert- conversion only:ffmpegresample/remix/re-encode into the target WAV/FLAC spec, then validation, then an optional manifest. Does not chunk.audio-prep chunk- chunking only: VAD speech chunking straight from source audio, with Silero by default and Pyannote available for comparison. Does not convert or validate.
convert pipeline stages
- discovery (
find_audio_files) — recursively find supported source files under an input directory. - conversion (
convert_file/convert_batch) — shell out toffmpegto resample/remix/re-encode into the target WAV/FLAC spec. Mirrors the input directory's subfolder structure on output. Runs in a process pool since each conversion is an independent subprocess call. - validation (
validate_output) — re-opens each converted file withsoundfileand checks it actually matches the requested sample rate, channel count, and minimum duration. This catches the case where ffmpeg exits 0 but silently produced something degenerate. - manifest (
build_manifest/write_manifest) — JSONL file, one row per source file, recordingstatus(ok/conversion_failed/validation_failed), output path, duration, sample rate, and any error.
Conversion failures don't abort the batch — a bad file in a 50,000-file
corpus shows up as one conversion_failed row in the manifest, not a crashed
job three hours in.
chunk pipeline stages
- discovery (
find_audio_files) — recursively find supported source files under an input directory. - chunking (
chunk_file/chunk_batch) — runs the selected VAD backend over each file (decoding/resampling via ffmpeg) and splits it into speech-only chunks bounded by a[min, max]duration window, so silence-heavy source recordings don't waste pretraining compute. - manifest (
build_chunk_manifest/write_manifest), optional — JSONL file, one row per source file, recordingstatus(ok/chunking_failed), chunk count, and chunk paths.
Chunking failures (e.g. no speech detected) work the same way: chunk_batch
returns a ChunkResult per file instead of raising.
Setup
# ffmpeg is a system dependency, not a pip package
sudo apt-get install ffmpeg # or: brew install ffmpeg
pip install audio-prep-pipeline
Install directly from GitHub:
pip install "audio-prep-pipeline @ git+https://github.com/nattkorat/audio-prep-pipeline.git"
For local development:
make install # pip install -e ".[dev]" + pre-commit install
The same install provides audio-prep convert, audio-prep chunk, Silero VAD,
and Pyannote VAD support.
FFmpeg/FFprobe are still system dependencies and must be available on PATH.
Some Pyannote models require a Hugging Face token and accepted model terms.
If the selected VAD backend cannot load in an offline environment, pass
--allow-energy-fallback to use a lower-quality offline detector instead.
Usage
audio-prep convert
CLI:
audio-prep convert \
--input-dir data/raw_mp3 \
--output-dir data/wav16k \
--format wav \
--sample-rate 16000 \
--workers 8 \
--manifest data/manifest.jsonl \
--profile profiles/convert.json
Python:
from pathlib import Path
from audio_prep import (
ConversionConfig,
build_manifest,
convert_batch,
validate_output,
write_manifest,
)
config = ConversionConfig(
output_format="wav",
sample_rate=16_000,
channels=1,
num_workers=8,
)
results = convert_batch(Path("data/raw_mp3"), Path("data/wav16k"), config)
validations = {
result.output: validate_output(result.output, config)
for result in results
if result.success and result.output is not None
}
records = build_manifest(results, validations)
write_manifest(records, Path("data/manifest.jsonl"))
| Flag | Default | Meaning |
|---|---|---|
--input-dir |
(required) | directory of source audio files |
--output-dir |
(required) | where converted output is written |
--extensions |
common FFmpeg audio/video extensions | comma-separated source extensions, or all to pass every regular file to FFmpeg |
--format |
wav |
output format (wav or flac) |
--sample-rate |
16000 | target sample rate |
--channels |
1 | target channel count |
--workers |
4 | parallel conversion workers |
--min-duration-sec |
0.5 | validation fails files shorter than this |
--overwrite |
off | re-convert even if output already exists and passes validation |
--normalize-loudness |
off | apply EBU R128 loudness normalization (-23 LUFS) |
--manifest |
none | path to write a JSONL manifest |
--profile |
none | path to write a JSON profile with wall time, CPU time, and peak RSS |
audio-prep chunk
Independent of convert -- scans --input-dir for supported source files and runs VAD
chunking directly against them, decoding (and resampling, if --sample-rate
doesn't match the source) via ffmpeg:
CLI:
audio-prep chunk \
--input-dir data/raw_mp3 \
--output-dir data/chunks \
--sample-rate 16000 \
--format flac \
--min-duration-sec 5 \
--max-duration-sec 20 \
--workers 4 \
--manifest data/chunk_manifest.jsonl \
--profile profiles/chunk-silero.json
Python:
from pathlib import Path
from audio_prep import ChunkConfig, build_chunk_manifest, chunk_batch, write_manifest
config = ChunkConfig(
min_duration_sec=5,
max_duration_sec=20,
output_format="flac",
sample_rate=16_000,
num_workers=4,
vad_backend="silero",
)
results = chunk_batch(Path("data/raw_mp3"), Path("data/chunks"), config)
records = build_chunk_manifest(results)
write_manifest(records, Path("data/chunk_manifest.jsonl"))
| Flag | Default | Meaning |
|---|---|---|
--input-dir |
(required) | directory of source audio files to scan and chunk |
--output-dir |
<input-dir>/chunks |
where chunks are written |
--extensions |
common FFmpeg audio/video extensions | comma-separated source extensions, or all to pass every regular file to FFmpeg |
--format |
wav |
output chunk format (wav or flac) |
--sample-rate |
16000 | resample (via ffmpeg) to this rate before chunking if the source doesn't already match it |
--min-duration-sec |
5.0 | drop chunks shorter than this |
--max-duration-sec |
20.0 | split longer speech into windows this size |
--workers |
4 | parallel chunking workers |
--overwrite |
off | re-chunk even if valid output already exists |
--vad-backend |
silero |
speech detector backend: silero, pyannote, or energy |
--pyannote-model |
pyannote/speaker-diarization-community-1 |
Hugging Face model id used with --vad-backend pyannote |
--pyannote-revision |
none | optional Hugging Face model revision; repo/model@revision is also accepted |
--hf-token |
none | Hugging Face token for gated Pyannote models; falls back to HF_TOKEN or HUGGINGFACE_TOKEN |
--allow-energy-fallback |
off | fall back to a low-quality energy detector if the selected VAD can't load, instead of raising |
--manifest |
none | path to write a JSONL chunk manifest (source file, status, chunk count/paths) |
--profile |
none | path to write a JSON profile with wall time, CPU time, and peak RSS |
Compare VAD backends with the same input/settings:
audio-prep chunk \
--input-dir data/raw_mp3 \
--output-dir data/chunks-pyannote \
--vad-backend pyannote \
--pyannote-model pyannote/speaker-diarization-community-1 \
--hf-token "$HF_TOKEN" \
--profile profiles/chunk-pyannote.json
Python profiling:
from pathlib import Path
from audio_prep import ChunkConfig, Profiler, chunk_batch
profiler = Profiler()
config = ChunkConfig(vad_backend="pyannote", sample_rate=16_000, num_workers=1)
with profiler.measure("chunk", {"vad_backend": config.vad_backend}):
results = chunk_batch(Path("data/raw_mp3"), Path("data/chunks-pyannote"), config)
profiler.write_json(
Path("profiles/chunk-pyannote.json"),
operation="chunk",
metadata={"files": len(results), "vad_backend": config.vad_backend},
)
chunk_file itself doesn't care about source extension -- it decodes
whatever path it's given -- so the Python API can also chunk an existing
WAV/FLAC corpus (e.g. convert_batch output) by passing source_files
explicitly instead of relying on chunk_batch discovery. This is an advanced
Python API case:
from pathlib import Path
from audio_prep import ChunkConfig, ConversionConfig, build_manifest, chunk_batch
from audio_prep import convert_batch, validate_output
config = ConversionConfig(output_format="wav", sample_rate=16_000, channels=1, num_workers=8)
results = convert_batch(Path("data/raw_mp3"), Path("data/wav16k"), config)
validations = {
result.output: validate_output(result.output, config)
for result in results
if result.success and result.output is not None
}
records = build_manifest(results, validations)
# source_files bypasses chunk_batch discovery, so this works
# directly against the already-converted WAV output above.
valid_outputs = [path for path, v in validations.items() if v.valid]
chunk_config = ChunkConfig(min_duration_sec=5, max_duration_sec=20, num_workers=4)
chunk_results = chunk_batch(
Path("data/wav16k"),
Path("data/wav16k/chunks"),
chunk_config,
source_files=valid_outputs,
)
Development workflow
make format # ruff format + autofix
make lint # ruff check
make typecheck # mypy --strict
make test # pytest with coverage
make check # all of the above -- run this before opening a PR
pre-commit (installed via make install) runs ruff + mypy + basic hygiene
checks automatically on every commit. CI (.github/workflows/ci.yml) re-runs
the same checks plus the full test matrix across Python 3.10–3.12 on every
push and PR.
Extending this
Natural next additions, in roughly the order they'd come up:
- Streaming manifest writes for very large corpora, instead of holding
all
ConversionResults in memory before writing.
Note: This is the template, that you have to extend from. Main branch is protected so you have to create another branch to work on.
Metadata
Release files for audio-prep-pipeline 0.1.5
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| audio_prep_pipeline-0.1.5.tar.gz | 44.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| audio_prep_pipeline-0.1.5-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 70.8 kB
Release files / audio_prep_pipeline-0.1.5.tar.gz
| Download URL | audio_prep_pipeline-0.1.5.tar.gz |
|---|---|
| Size | 44.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
67cf66cad20d14a6b12a15204ca3d9d6b8705d3737d3d010c97b7998aeae43c6
|
|
BLAKE2b-256 checksum How to use checksums |
768fb69a6614e08f588bba586f3750e2220211b3d3f1deb19a8fc2a91d4de322
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.8
|
Release files / audio_prep_pipeline-0.1.5-py3-none-any.whl
| Download URL | audio_prep_pipeline-0.1.5-py3-none-any.whl |
|---|---|
| Size | 26.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
94d4125396a433d13b6a944f44f34d862c1304d051f21ae4f6367ab7e39bcae8
|
|
BLAKE2b-256 checksum How to use checksums |
5871a4ffe9acea9d7eac4531d6237bb610b0950664a763b760c79e01b5e42702
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.8
|