transcribe-cpp
Python bindings for transcribe.cpp, a C/C++ speech-to-text library built on ggml.
Status: in development. Until wheels are published, use a locally built
libtranscribethrough repo auto-discovery orTRANSCRIBE_LIBRARY.
Upgrading from 0.1? See the
0.2 migration guide,
including the replacement of gpu_device= with exact device objects.
import transcribe_cpp
with transcribe_cpp.Model("model.gguf") as model:
with model.session() as session:
result = session.run(pcm_float32_16k_mono)
print(result.text)
run() takes mono 16 kHz float32 PCM (buffer-protocol object or sequence). It
does not decode containers or resample; convert audio before calling it.
import numpy as np
pcm = np.asarray(audio, dtype=np.float32) # 1-D, 16 kHz mono
# Downmix stereo first; 2-D input is rejected:
# pcm = audio.mean(axis=1).astype(np.float32)
result = session.run(pcm)
Punctuation, capitalization, and text normalization
Generic run controls use "default" to preserve each model family's shipped
behavior. Models advertising model.supports("pnc") accept pnc="off" or
pnc="on"; models advertising model.supports("itn") accept the equivalent
itn values. The options are available on run(), run_batch(), stream(),
and the one-shot transcribe() helper.
result = session.run(pcm, pnc="off", itn="on")
Streaming models expose incremental transcription with committed/tentative
text views — see examples/stream_wav.py:
with model.session() as session, session.stream() as stream:
for chunk in pcm_chunks:
stream.feed(chunk)
text = stream.text() # .committed (stable) + .tentative
stream.finalize()
result = stream.snapshot() # language, segments, words, tokens, timings
Long transcriptions can be cancelled from another thread with
session.cancel() — the run raises Aborted with the partial transcript on
exc.partial_result (same for OutputTruncated).
Backends
Model(backend=...) applies a backend policy ("auto" uses the best
available). transcribe_cpp.backends() returns process-local device objects;
pass one as Model(device=device) for exact selection with no fallback. Persist
a device's device_id, not its runtime handle or index. backend_available(kind)
checks whether a backend policy can currently be satisfied.
device = next(d for d in transcribe_cpp.backends() if d.device_type == "cpu")
with transcribe_cpp.Model("model.gguf", device=device) as model:
print(model.device)
| Variable | Effect |
|---|---|
TRANSCRIBE_BACKEND |
overrides the "auto" default; explicit backend= still wins |
TRANSCRIBE_NATIVE_PROVIDER |
forces an installed native provider package, for example cu12 |
TRANSCRIBE_LIBRARY |
loads exactly this shared library |
Planned wheels will bundle CPU plus platform accelerators;
transcribe-cpp[cu12] will add the CUDA 12 provider.
Running from a working tree
The binding loads the native library at import and verifies its ABI layout and
version before use. Build a shared library, then run from the repo or point
TRANSCRIBE_LIBRARY at it:
cmake -B build-shared -DTRANSCRIBE_BUILD_SHARED=ON
cmake --build build-shared --target transcribe
cd bindings/python
PYTHONPATH=src uv run --no-project python examples/transcribe_wav.py \
../../models/whisper-tiny.en/whisper-tiny.en-Q5_K_M.gguf ../../samples/jfk.wav
No-model tests always run; model tests skip unless smoke assets are present.
Override paths with TRANSCRIBE_SMOKE_MODEL, TRANSCRIBE_SMOKE_AUDIO, and
TRANSCRIBE_SMOKE_STREAMING_MODEL.
cd bindings/python
TRANSCRIBE_LIBRARY=../../build-shared/src/libtranscribe.dylib \
uv run --extra test pytest
Notes
- One run/stream at a time per
Modelin 0.x: sessions share the model's compute backend, so serialize runs across sessions (or load one model per worker). See theModeldocstring. - Import package:
transcribe_cpp - Distribution:
transcribe-cpp - License: MIT
Release files for transcribe-cpp 0.2.4
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| transcribe_cpp-0.2.4.tar.gz | 94.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| transcribe_cpp-0.2.4-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 129.3 kB
Release files / transcribe_cpp-0.2.4.tar.gz
| Download URL | transcribe_cpp-0.2.4.tar.gz |
|---|---|
| Size | 94.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
5ca7433520282f58e4435ca865cfa4cd8a46a2ee630de0e3301683efbdfbf370
|
|
BLAKE2b-256 checksum How to use checksums |
d35380b7be4dd250f86d0757f792c3e52642114231b5df31cddd7f19fe5c38c1
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 25, 2026.
Transparency logRelease files / transcribe_cpp-0.2.4-py3-none-any.whl
| Download URL | transcribe_cpp-0.2.4-py3-none-any.whl |
|---|---|
| Size | 34.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
e4bde0002fea09dc2b573f9b18c9d5d1163630b3096392b3dc82eb5dce25be9d
|
|
BLAKE2b-256 checksum How to use checksums |
2c0ad55134fb1acb43d3cc2d2dd46e23aace14e7396d364744373263fdc05db9
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 25, 2026.
Transparency log