Skip to main content

generic-audio-transcriber

Drop-in speech-to-text for any Python project. Audio bytes in, JSON-ready text out.

from audio_transcriber import transcribe

result = transcribe(audio_bytes)   # mp3, wav, webm, ogg, m4a... detected automatically
print(result.text)                 # "Hello, this is a test."
print(result.to_json())            # {"text": "...", "language": "en", ...}

No API keys, no servers, no cloud. It runs locally on the CPU, and the model is downloaded automatically the first time you use it.

Why

Many apps need the same small feature: the user sends or records audio, and the app needs the text. This package is that feature as one function call, so you can add it to a project without learning anything about speech recognition.

Under the hood it uses faster-whisper (OpenAI's Whisper model), which is accurate, supports about 99 languages, and runs well on a plain CPU. You never have to train or tune a model.

Install

pip install git+https://github.com/EduardoMilani8/generic-audio-transcriber.git

Requires Python 3.9+. You do not need to install ffmpeg: audio decoding is bundled.

Usage

From bytes (the main use case)

from audio_transcriber import transcribe

result = transcribe(audio_bytes)

From a file path or file object

transcribe("recording.mp3")

with open("recording.wav", "rb") as f:
    transcribe(f)

Inside a web endpoint

from fastapi import FastAPI, UploadFile
from audio_transcriber import transcribe

app = FastAPI()

@app.post("/transcribe")
async def endpoint(file: UploadFile):
    return transcribe(await file.read()).to_dict()

The result

result.text                  # full transcript
result.language              # detected language code, e.g. "pt"
result.language_probability  # confidence of the language detection
result.duration              # audio length in seconds
result.segments              # timed pieces: .start, .end, .text
result.to_dict()             # plain dict
result.to_json()             # JSON string (non-ASCII kept readable)

Silent audio gives an empty text, not an error.

Options

transcribe(
    audio,
    model="small",         # tiny | base | small | medium | large-v3
    language=None,         # "pt", "en", ... None = auto-detect
    device="cpu",          # "cpu" | "cuda" | "auto"
    compute_type="int8",   # "int8" is light; "float16" suits GPUs
    beam_size=5,
    vad_filter=True,       # skip silence, reduces made-up text
)

Choosing a model is a trade-off between weight and accuracy:

Model Size on disk Speed Accuracy
tiny smallest fastest lowest
base small fast fair
small (default) medium good good
medium / large-v3 large slow on CPU best

The model is loaded once per process and reused, so only the first call is slow. Setting language explicitly is faster and more accurate than auto-detection.

Errors

from audio_transcriber import transcribe, TranscriptionError

try:
    transcribe(data)
except TranscriptionError as e:
    ...  # empty input or audio that cannot be decoded

Command line

audio-transcriber recording.mp3            # prints JSON
audio-transcriber recording.mp3 --text     # prints plain text
audio-transcriber recording.mp3 --model base --language pt
cat recording.mp3 | audio-transcriber -    # read bytes from stdin

Development

python -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"
pytest                                # fast unit tests, no model needed
RUN_SLOW=1 pytest -m slow             # end-to-end tests, downloads the tiny model

Notes

  • The default device is cpu because it works on every machine. Pass device="cuda" if you have a working CUDA setup and want GPU speed.
  • av is pinned below version 19 because faster-whisper still uses an argument that newer releases removed. The pin can be lifted once faster-whisper fixes it.

License

MIT

Metadata

Release files for generic-audio-transcriber 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for generic-audio-transcriber 0.1.0
File Size Uploaded
generic_audio_transcriber-0.1.0.tar.gz 11.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for generic-audio-transcriber 0.1.0
File Interpreter ABI Platform
generic_audio_transcriber-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 19.6 kB

Release files / generic_audio_transcriber-0.1.0.tar.gz

Download URL generic_audio_transcriber-0.1.0.tar.gz
Size 11.3 kB
Tags Source
SHA-256 checksum
How to use checksums
8360c6dd3993c27290e4a8225168b7a286a282a1be27b71eeb40845c07ab93bf
BLAKE2b-256 checksum
How to use checksums
f6a4dc972e88464c21ebc28bb45478ac590d51c4108331821d3d5060aff0b571
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.4

Release files / generic_audio_transcriber-0.1.0-py3-none-any.whl

Download URL generic_audio_transcriber-0.1.0-py3-none-any.whl
Size 8.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
75aaeb6f16b025d9de0422efc439c5d73657186204269b93772e0ca6b7fa21de
BLAKE2b-256 checksum
How to use checksums
3d43d96fa2340f0eabc54bf9ecb5c8d6ca55c517dba82f974155fca849e5744a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.4

Release history Release notifications | RSS feed

0.1.1

2 release files

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page