Skip to main content

generic-audio-transcriber

Drop-in speech-to-text for any Python project. Audio bytes in, JSON-ready text out.

from audio_transcriber import transcribe

result = transcribe(audio_bytes)   # mp3, wav, webm, ogg, m4a... detected automatically
print(result.text)                 # "Hello, this is a test."
print(result.to_json())            # {"text": "...", "language": "en", ...}

No API keys, no servers, no cloud. It runs locally on the CPU, and the model is downloaded automatically the first time you use it.

Why

Many apps need the same small feature: the user sends or records audio, and the app needs the text. This package is that feature as one function call, so you can add it to a project without learning anything about speech recognition.

Under the hood it uses faster-whisper (OpenAI's Whisper model), which is accurate, supports about 99 languages, and runs well on a plain CPU. You never have to train or tune a model.

Install

pip install generic-audio-transcriber

Requires Python 3.9+. You do not need to install ffmpeg: audio decoding is bundled.

Usage

From bytes (the main use case)

from audio_transcriber import transcribe

result = transcribe(audio_bytes)

From a file path or file object

transcribe("recording.mp3")

with open("recording.wav", "rb") as f:
    transcribe(f)

Inside a web endpoint

from fastapi import FastAPI, UploadFile
from audio_transcriber import transcribe

app = FastAPI()

@app.post("/transcribe")
async def endpoint(file: UploadFile):
    return transcribe(await file.read()).to_dict()

The result

result.text                  # full transcript
result.language              # detected language code, e.g. "pt"
result.language_probability  # confidence of the language detection
result.duration              # audio length in seconds
result.segments              # timed pieces: .start, .end, .text
result.to_dict()             # plain dict
result.to_json()             # JSON string (non-ASCII kept readable)

Silent audio gives an empty text, not an error.

Options

transcribe(
    audio,
    model="small",         # tiny | base | small | medium | large-v3
    language=None,         # "pt", "en", ... None = auto-detect
    device="cpu",          # "cpu" | "cuda" | "auto"
    compute_type="int8",   # "int8" is light; "float16" suits GPUs
    beam_size=5,
    vad_filter=True,       # skip silence, reduces made-up text
)

Choosing a model is a trade-off between weight and accuracy:

Model Size on disk Speed Accuracy
tiny smallest fastest lowest
base small fast fair
small (default) medium good good
medium / large-v3 large slow on CPU best

The model is loaded once per process and reused, so only the first call is slow. Setting language explicitly is faster and more accurate than auto-detection.

Errors

from audio_transcriber import transcribe, TranscriptionError

try:
    transcribe(data)
except TranscriptionError as e:
    ...  # empty input or audio that cannot be decoded

Command line

audio-transcriber recording.mp3            # prints JSON
audio-transcriber recording.mp3 --text     # prints plain text
audio-transcriber recording.mp3 --model base --language pt
cat recording.mp3 | audio-transcriber -    # read bytes from stdin

Development

python -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"
pytest                                # fast unit tests, no model needed
RUN_SLOW=1 pytest -m slow             # end-to-end tests, downloads the tiny model

Notes

  • The default device is cpu because it works on every machine. Pass device="cuda" if you have a working CUDA setup and want GPU speed.
  • av is pinned below version 19 because faster-whisper still uses an argument that newer releases removed. The pin can be lifted once faster-whisper fixes it.

License

MIT

Metadata

Release files for generic-audio-transcriber 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for generic-audio-transcriber 0.1.1
File Size Uploaded
generic_audio_transcriber-0.1.1.tar.gz 10.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for generic-audio-transcriber 0.1.1
File Interpreter ABI Platform
generic_audio_transcriber-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 18.4 kB

Release files / generic_audio_transcriber-0.1.1.tar.gz

Download URL generic_audio_transcriber-0.1.1.tar.gz
Size 10.0 kB
Tags Source
SHA-256 checksum
How to use checksums
7d0c14b6aa9d0b66b694e792dadbebddef2d7a81edc337808976270906428bfc
BLAKE2b-256 checksum
How to use checksums
3fab502edc60bd9a7851d27541ffa5d14b9366255421718db0e1c6313f639df0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.4

Release files / generic_audio_transcriber-0.1.1-py3-none-any.whl

Download URL generic_audio_transcriber-0.1.1-py3-none-any.whl
Size 8.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
19b1b1918e3728c996f0331b828ab02d15687829a6ea43a1d4c2094e36f67256
BLAKE2b-256 checksum
How to use checksums
53283c4eff0da1f986ca505c204a7544d39d672a879dde8000702127609a9cb4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.4

Release history Release notifications | RSS feed

This release

0.1.1 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page