tomfun-speech-to-text-client
Minimal async Python client for the tomfun speech-to-text web API.
- Token authentication (
Authorization: Bearer <token>) - Language hint passed to Whisper
- Upload local files (chunked) or submit a remote URL
- Streams status via Server-Sent Events
Single runtime dependency: aiohttp.
Install
pip install tomfun-speech-to-text-client
Usage
import asyncio
from tomfun_speech_to_text_client import Client
async def main():
async with Client("https://speech-to-text.tomfun.co", token="YOUR_TOKEN") as c:
r = await c.transcribe_file("audio.mp3", language="en")
print(r.text)
r = await c.transcribe_url("https://example.com/clip.wav", language="uk")
print(r.srt)
asyncio.run(main())
With progress callback:
async def on_status(ev):
print(ev["status"], ev.get("progressLast"))
await c.transcribe_file("audio.mp3", language="en", on_status=on_status)
Cancel a job:
await c.cancel(job_id)
Speaker diarization (server-side pyannote):
r = await c.transcribe_file("meeting.mp3", language="ru", diarize=True)
print(r.diarized) # SRT with a [SPEAKER_xx] prefix per cue, or None
print(r.speaker_embeddings) # base64 .npy, shape (numSpeakers, embeddingDim), or None
Both diarization fields are None when diarize wasn't requested — and also when the
server's diarization failed on an otherwise successful job, so always check them.
TranscriptionResult
| field | notes |
|---|---|
text |
plain transcript. Empty on silence; when the server sends it empty while the whisper JSON still holds segments, the client rebuilds it from json.transcription |
srt / vtt / csv |
subtitle/table renderings |
json |
parsed WhisperFullJSON (whisper-cli --output-json-full): transcription[].text / .offsets / .tokens |
diarized |
SRT annotated with [SPEAKER_xx], or None |
speaker_embeddings |
base64-encoded .npy, one row per detected speaker, or None |
accuracy |
accuracy of the model that actually ran (lower than preferred_accuracy when the job was routed to a degraded manager) |
raw |
the final job payload as received |
Development (uv)
cd python_client
uv sync --group dev
uv run pytest
Publishing to PyPI
- Create a PyPI API token at https://pypi.org/manage/account/token/.
- Bump
versioninpyproject.toml. - Build:
uv build - Upload:
uv publish --token "$PYPI_TOKEN"
Or viatwine:uv tool run --with twine twine upload dist/*
Metadata
Release files for tomfun-speech-to-text-client 0.0.4
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| tomfun_speech_to_text_client-0.0.4.tar.gz | 98.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| tomfun_speech_to_text_client-0.0.4-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 107.4 kB
Release files / tomfun_speech_to_text_client-0.0.4.tar.gz
| Download URL | tomfun_speech_to_text_client-0.0.4.tar.gz |
|---|---|
| Size | 98.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
37a5a7dc2249f738a234bbe4d0e17e6da35dae5a16e095a7ff6ad1f65308981d
|
|
BLAKE2b-256 checksum How to use checksums |
b748c5be63011703b95033736ef8b9b9796a395a37fd488d647828bc092dd3e7
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.12
|
Release files / tomfun_speech_to_text_client-0.0.4-py3-none-any.whl
| Download URL | tomfun_speech_to_text_client-0.0.4-py3-none-any.whl |
|---|---|
| Size | 8.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
454dabcd2ca03e96ee849668d625b650357080cb8bbb1593de8c2188ef27ae30
|
|
BLAKE2b-256 checksum How to use checksums |
549ffefc3251162d59ac2cd2c7923e0a8456f491f4375718e61e59d5c753c446
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.12
|