tomfun-speech-to-text-client
Minimal async Python client for the tomfun speech-to-text web API.
- Token authentication (
Authorization: Bearer <token>) - Language hint passed to Whisper
- Upload local files (chunked) or submit a remote URL
- Streams status via Server-Sent Events
Single runtime dependency: aiohttp.
Install
pip install tomfun-speech-to-text-client
Usage
import asyncio
from tomfun_speech_to_text_client import Client
async def main():
async with Client("https://speech-to-text.tomfun.co", token="YOUR_TOKEN") as c:
r = await c.transcribe_file("audio.mp3", language="en")
print(r.text)
r = await c.transcribe_url("https://example.com/clip.wav", language="uk")
print(r.srt)
asyncio.run(main())
With progress callback:
async def on_status(ev):
print(ev["status"], ev.get("progressLast"))
await c.transcribe_file("audio.mp3", language="en", on_status=on_status)
Cancel a job:
await c.cancel(job_id)
Speaker diarization (server-side pyannote):
r = await c.transcribe_file("meeting.mp3", language="ru", diarize=True)
print(r.diarized) # SRT with a [SPEAKER_xx] prefix per cue, or None
print(r.speaker_embeddings) # base64 .npy, shape (numSpeakers, embeddingDim), or None
Both diarization fields are None when diarize wasn't requested — and also when the
server's diarization failed on an otherwise successful job, so always check them.
allow_experimental
Off unless you pass it, and that default is deliberate: without it you always get genuine whisper.cpp output.
Turning it on admits models that serve a reduced result contract — for
parakeet that means no language selection, no VAD, and a WhisperFullJSON that is
schema-valid but synthesized (systeminfo / model / params are placeholders).
Diarization is not part of the reduction, but a worker built with
ENABLE_DIARIZATION=0 has none regardless.
What it buys: an experimental model can be the most accurate thing in the cluster
and still be the only free worker. With the flag off, such a manager is not merely
ranked lower — it is not a candidate at all, so a job can sit in the queue for
hours next to an idle machine that could have run it. Pass
allow_experimental=True when you care more about the job finishing than about
the exact shape of the JSON.
TranscriptionResult
| field | notes |
|---|---|
text |
plain transcript. Empty on silence; when the server sends it empty while the whisper JSON still holds segments, the client rebuilds it from json.transcription |
srt / vtt / csv |
subtitle/table renderings |
json |
parsed WhisperFullJSON (whisper-cli --output-json-full): transcription[].text / .offsets / .tokens |
diarized |
SRT annotated with [SPEAKER_xx], or None |
speaker_embeddings |
base64-encoded .npy, one row per detected speaker, or None |
accuracy |
accuracy of the model that actually ran (lower than preferred_accuracy when the job was routed to a degraded manager) |
raw |
the final job payload as received |
Development (uv)
cd python_client
uv sync --group dev
uv run pytest
Publishing to PyPI
- Create a PyPI API token at https://pypi.org/manage/account/token/.
- Bump
versioninpyproject.toml. - Build:
uv build - Upload.
uv publishdoes not read~/.pypirc(only--token/UV_PUBLISH_TOKEN,--username/--password, keyring, trusted publishing), so with the token stored there the shortest path istwine, which picks it up on its own — name the new version explicitly, otherwise twine trips over the already-published files indist/:uv tool run --with twine twine upload dist/tomfun_speech_to_text_client-0.0.5*
Or pass the token touvdirectly:uv publish --token "$PYPI_TOKEN"
Metadata
Release files for tomfun-speech-to-text-client 0.0.6
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| tomfun_speech_to_text_client-0.0.6.tar.gz | 100.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| tomfun_speech_to_text_client-0.0.6-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 110.2 kB
Release files / tomfun_speech_to_text_client-0.0.6.tar.gz
| Download URL | tomfun_speech_to_text_client-0.0.6.tar.gz |
|---|---|
| Size | 100.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
0d85b1790bf48bbb63dcdec05ae61e7bfe475763bfeb64bb57cda16bba2ecd63
|
|
BLAKE2b-256 checksum How to use checksums |
a379cdec5ab09e2e983f19eeae72d44723ee9a9dbb2b69e0142387c510f35a38
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.12
|
Release files / tomfun_speech_to_text_client-0.0.6-py3-none-any.whl
| Download URL | tomfun_speech_to_text_client-0.0.6-py3-none-any.whl |
|---|---|
| Size | 10.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
7a664b9b91585f5980fc98a414dc63eca76bc3e3d2d1533f5d575aafe78180ba
|
|
BLAKE2b-256 checksum How to use checksums |
9b03db62e11c42982d22275009f2a73624be9bf022c36e8d8941b1aa58bea5b8
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.12
|