Skip to main content

tomfun-speech-to-text-client

Minimal async Python client for the tomfun speech-to-text web API.

  • Token authentication (Authorization: Bearer <token>)
  • Language hint passed to Whisper
  • Upload local files (chunked) or submit a remote URL
  • Streams status via Server-Sent Events

Single runtime dependency: aiohttp.

Install

pip install tomfun-speech-to-text-client

Usage

import asyncio
from tomfun_speech_to_text_client import Client

async def main():
    async with Client("https://speech-to-text.tomfun.co", token="YOUR_TOKEN") as c:
        r = await c.transcribe_file("audio.mp3", language="en")
        print(r.text)

        r = await c.transcribe_url("https://example.com/clip.wav", language="uk")
        print(r.srt)

asyncio.run(main())

With progress callback:

async def on_status(ev):
    print(ev["status"], ev.get("progressLast"))

await c.transcribe_file("audio.mp3", language="en", on_status=on_status)

Cancel a job:

await c.cancel(job_id)

Speaker diarization (server-side pyannote):

r = await c.transcribe_file("meeting.mp3", language="ru", diarize=True)
print(r.diarized)            # SRT with a [SPEAKER_xx] prefix per cue, or None
print(r.speaker_embeddings)  # base64 .npy, shape (numSpeakers, embeddingDim), or None

Both diarization fields are None when diarize wasn't requested — and also when the server's diarization failed on an otherwise successful job, so always check them.

allow_experimental

Off unless you pass it, and that default is deliberate: without it you always get genuine whisper.cpp output.

Turning it on admits models that serve a reduced result contract — for parakeet that means no language selection, no VAD, and a WhisperFullJSON that is schema-valid but synthesized (systeminfo / model / params are placeholders). Diarization is not part of the reduction, but a worker built with ENABLE_DIARIZATION=0 has none regardless.

What it buys: an experimental model can be the most accurate thing in the cluster and still be the only free worker. With the flag off, such a manager is not merely ranked lower — it is not a candidate at all, so a job can sit in the queue for hours next to an idle machine that could have run it. Pass allow_experimental=True when you care more about the job finishing than about the exact shape of the JSON.

TranscriptionResult

field notes
text plain transcript. Empty on silence; when the server sends it empty while the whisper JSON still holds segments, the client rebuilds it from json.transcription
srt / vtt / csv subtitle/table renderings
json parsed WhisperFullJSON (whisper-cli --output-json-full): transcription[].text / .offsets / .tokens
diarized SRT annotated with [SPEAKER_xx], or None
speaker_embeddings base64-encoded .npy, one row per detected speaker, or None
accuracy accuracy of the model that actually ran (lower than preferred_accuracy when the job was routed to a degraded manager)
raw the final job payload as received

Development (uv)

cd python_client
uv sync --group dev
uv run pytest

Publishing to PyPI

  1. Create a PyPI API token at https://pypi.org/manage/account/token/.
  2. Bump version in pyproject.toml.
  3. Build:
    uv build
    
  4. Upload. uv publish does not read ~/.pypirc (only --token / UV_PUBLISH_TOKEN, --username/--password, keyring, trusted publishing), so with the token stored there the shortest path is twine, which picks it up on its own — name the new version explicitly, otherwise twine trips over the already-published files in dist/:
    uv tool run --with twine twine upload dist/tomfun_speech_to_text_client-0.0.5*
    
    Or pass the token to uv directly:
    uv publish --token "$PYPI_TOKEN"
    

Metadata

Release files for tomfun-speech-to-text-client 0.0.6

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for tomfun-speech-to-text-client 0.0.6
File Size Uploaded
tomfun_speech_to_text_client-0.0.6.tar.gz 100.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for tomfun-speech-to-text-client 0.0.6
File Interpreter ABI Platform
tomfun_speech_to_text_client-0.0.6-py3-none-any.whl Python 3 none any Details

Total release size: 110.2 kB

Release files / tomfun_speech_to_text_client-0.0.6.tar.gz

Download URL tomfun_speech_to_text_client-0.0.6.tar.gz
Size 100.0 kB
Tags Source
SHA-256 checksum
How to use checksums
0d85b1790bf48bbb63dcdec05ae61e7bfe475763bfeb64bb57cda16bba2ecd63
BLAKE2b-256 checksum
How to use checksums
a379cdec5ab09e2e983f19eeae72d44723ee9a9dbb2b69e0142387c510f35a38
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.12

Release files / tomfun_speech_to_text_client-0.0.6-py3-none-any.whl

Download URL tomfun_speech_to_text_client-0.0.6-py3-none-any.whl
Size 10.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
7a664b9b91585f5980fc98a414dc63eca76bc3e3d2d1533f5d575aafe78180ba
BLAKE2b-256 checksum
How to use checksums
9b03db62e11c42982d22275009f2a73624be9bf022c36e8d8941b1aa58bea5b8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.12

Release history Release notifications | RSS feed

This release

0.0.6 This release

2 release files

0.0.5

2 release files

0.0.4

2 release files

0.0.3

2 release files

0.0.2

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page