generic-audio-transcriber
Drop-in speech-to-text for any Python project. Audio bytes in, JSON-ready text out.
from audio_transcriber import transcribe
result = transcribe(audio_bytes) # mp3, wav, webm, ogg, m4a... detected automatically
print(result.text) # "Hello, this is a test."
print(result.to_json()) # {"text": "...", "language": "en", ...}
No API keys, no servers, no cloud. It runs locally on the CPU, and the model is downloaded automatically the first time you use it.
Why
Many apps need the same small feature: the user sends or records audio, and the app needs the text. This package is that feature as one function call, so you can add it to a project without learning anything about speech recognition.
Under the hood it uses faster-whisper (OpenAI's Whisper model), which is accurate, supports about 99 languages, and runs well on a plain CPU. You never have to train or tune a model.
Install
pip install generic-audio-transcriber
Requires Python 3.9+. You do not need to install ffmpeg: audio decoding is bundled.
Usage
From bytes (the main use case)
from audio_transcriber import transcribe
result = transcribe(audio_bytes)
From a file path or file object
transcribe("recording.mp3")
with open("recording.wav", "rb") as f:
transcribe(f)
Inside a web endpoint
from fastapi import FastAPI, UploadFile
from audio_transcriber import transcribe
app = FastAPI()
@app.post("/transcribe")
async def endpoint(file: UploadFile):
return transcribe(await file.read()).to_dict()
The result
result.text # full transcript
result.language # detected language code, e.g. "pt"
result.language_probability # confidence of the language detection
result.duration # audio length in seconds
result.segments # timed pieces: .start, .end, .text
result.to_dict() # plain dict
result.to_json() # JSON string (non-ASCII kept readable)
Silent audio gives an empty text, not an error.
Options
transcribe(
audio,
model="small", # tiny | base | small | medium | large-v3
language=None, # "pt", "en", ... None = auto-detect
device="cpu", # "cpu" | "cuda" | "auto"
compute_type="int8", # "int8" is light; "float16" suits GPUs
beam_size=5,
vad_filter=True, # skip silence, reduces made-up text
)
Choosing a model is a trade-off between weight and accuracy:
| Model | Size on disk | Speed | Accuracy |
|---|---|---|---|
tiny |
smallest | fastest | lowest |
base |
small | fast | fair |
small (default) |
medium | good | good |
medium / large-v3 |
large | slow on CPU | best |
The model is loaded once per process and reused, so only the first call is slow.
Setting language explicitly is faster and more accurate than auto-detection.
Errors
from audio_transcriber import transcribe, TranscriptionError
try:
transcribe(data)
except TranscriptionError as e:
... # empty input or audio that cannot be decoded
Command line
audio-transcriber recording.mp3 # prints JSON
audio-transcriber recording.mp3 --text # prints plain text
audio-transcriber recording.mp3 --model base --language pt
cat recording.mp3 | audio-transcriber - # read bytes from stdin
Development
python -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"
pytest # fast unit tests, no model needed
RUN_SLOW=1 pytest -m slow # end-to-end tests, downloads the tiny model
Notes
- The default device is
cpubecause it works on every machine. Passdevice="cuda"if you have a working CUDA setup and want GPU speed. avis pinned below version 19 because faster-whisper still uses an argument that newer releases removed. The pin can be lifted once faster-whisper fixes it.
License
MIT
Metadata
Release files for generic-audio-transcriber 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| generic_audio_transcriber-0.1.1.tar.gz | 10.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| generic_audio_transcriber-0.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 18.4 kB
Release files / generic_audio_transcriber-0.1.1.tar.gz
| Download URL | generic_audio_transcriber-0.1.1.tar.gz |
|---|---|
| Size | 10.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
7d0c14b6aa9d0b66b694e792dadbebddef2d7a81edc337808976270906428bfc
|
|
BLAKE2b-256 checksum How to use checksums |
3fab502edc60bd9a7851d27541ffa5d14b9366255421718db0e1c6313f639df0
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.4
|
Release files / generic_audio_transcriber-0.1.1-py3-none-any.whl
| Download URL | generic_audio_transcriber-0.1.1-py3-none-any.whl |
|---|---|
| Size | 8.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
19b1b1918e3728c996f0331b828ab02d15687829a6ea43a1d4c2094e36f67256
|
|
BLAKE2b-256 checksum How to use checksums |
53283c4eff0da1f986ca505c204a7544d39d672a879dde8000702127609a9cb4
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.4
|