Skip to main content

autourgos-whisperstt

Framework: Autourgos Python License: Apache 2.0 Author

Local, offline speech-to-text for the Autourgos framework using faster-whisper (a CTranslate2 reimplementation of OpenAI's Whisper). No API key, no per-call network cost — runs entirely on-device (CPU or GPU) after a one-time model download. Cross-platform.

For zero-download, zero-setup transcription on Windows specifically (at the cost of noticeably lower accuracy), see the sibling package autourgos-windowstt.

from autourgos_micinput import MicrophoneStream
from autourgos_whisperstt import WhisperSTT

stt = WhisperSTT(model_size="base")  # tiny/base/small/medium/large-v3

async def listen(seconds: float = 2.0):
    frames_needed = int(seconds * 1000 / 100)  # default chunk_ms=100
    chunks = []
    async with MicrophoneStream(sample_rate=16000) as mic:
        async for chunk in mic:
            chunks.append(chunk)
            if len(chunks) >= frames_needed:
                break
    return stt.transcribe(b"".join(chunks), sample_rate=16000)

# text = asyncio.run(listen())

Install

pip install "autourgos-whisperstt[whisper]"

faster-whisper is required to actually call .transcribe()/.atranscribe() and is gated behind the whisper extra — import autourgos_whisperstt alone never requires it. The model itself downloads (once, cached by faster-whisper/huggingface_hub) and loads lazily on first use, not at WhisperSTT() construction. Requires Python 3.10+.


Usage

WhisperSTT takes raw 16-bit PCM bytes (exactly what autourgos-micinput's MicrophoneStream yields) and returns transcribed text:

stt = WhisperSTT(model_size="base")
text = stt.transcribe(pcm_bytes, sample_rate=16000)          # sync, blocking
text = await stt.atranscribe(pcm_bytes, sample_rate=16000)   # async, offloads to a worker thread

transcribe() is blocking — it runs Whisper inference on the calling thread. Use atranscribe() from inside an event loop (e.g. right after capturing from MicrophoneStream) instead of calling transcribe() directly, same reasoning as SpeakerPlayer.awrite() in autourgos-live.

Audio is written to a temporary WAV file and handed to faster-whisper as a file path rather than a raw array — its own decoding (via av/ffmpeg) resamples correctly to the 16kHz Whisper expects regardless of your sample_rate, so you don't need to resample yourself.

With autourgos-openaichat / autourgos-responses / autourgos-agent

None of those packages depend on this one (same reasoning as autourgos-micinput — no forced dependency on callers who don't need local speech). Wire them together yourself:

from autourgos_micinput import MicrophoneStream
from autourgos_whisperstt import WhisperSTT
from autourgos_openaichat import OpenAIChatModel

stt = WhisperSTT(model_size="base")
llm = OpenAIChatModel(model="gpt-4o")

async def voice_turn():
    chunks = []
    async with MicrophoneStream(sample_rate=16000) as mic:
        async for chunk in mic:
            chunks.append(chunk)
            if len(chunks) >= 20:  # ~2s
                break
    text = await stt.atranscribe(b"".join(chunks), sample_rate=16000)
    return llm.invoke(text)

API Reference

WhisperSTT(*, model_size="base", device="cpu", compute_type="int8", language=None)

Parameter Type Default Description
model_size str "base" Any faster-whisper model name — "tiny", "base", "small", "medium", "large-v3", etc. Bigger = more accurate, slower, larger download
device str "cpu" "cpu" or "cuda" (GPU, if available and ctranslate2 has CUDA support)
compute_type str "int8" Quantization — "int8" (fastest/smallest, good for CPU), "float16" (typical for GPU), "float32"
language str None ISO 639-1 code (e.g. "en") to skip auto-detection; None auto-detects per call. Overridable per-call via transcribe(..., language=...)
Method Description
transcribe(pcm_bytes, *, sample_rate=16000, channels=1, language=None) -> str Blocking. Lazily creates the model on first call, then reuses it. Returns "" if nothing was recognized (e.g. silence).
atranscribe(pcm_bytes, *, sample_rate=16000, channels=1, language=None) -> str Async-safe equivalent — runs transcribe() in a worker thread.

Errors (autourgos_whisperstt)

Name Raised when
WhisperSTTError Base class
WhisperSTTUnavailableError faster-whisper isn't installed

Accuracy note

Live-tested round-trip (Windows SAPI text-to-speech generating a WAV, fed back through this package's recognition): spoken "Testing one two three, this is a speech recognition check." came back as "Testing 1, 2, 3, this is a speech recognition check." — near-perfect, including punctuation. autourgos-windowstt's SAPI-based recognition of the identical audio returned "Testing 1 to 3 this is a speech recognition check" — noticeably rougher. Use this package when accuracy matters more than the one-time model download; use autourgos-windowstt when you want zero setup / zero download / instant offline results on Windows specifically and can tolerate lower accuracy.


License

Apache License 2.0, Copyright (c) 2026 Jitin Kumar Sengar

Metadata

Release files for autourgos-whisperstt 0.1.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for autourgos-whisperstt 0.1.2
File Size Uploaded
autourgos_whisperstt-0.1.2.tar.gz 18.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for autourgos-whisperstt 0.1.2
File Interpreter ABI Platform
autourgos_whisperstt-0.1.2-py3-none-any.whl Python 3 none any Details

Total release size: 34.1 kB

Release files / autourgos_whisperstt-0.1.2.tar.gz

Download URL autourgos_whisperstt-0.1.2.tar.gz
Size 18.5 kB
Tags Source
SHA-256 checksum
How to use checksums
f6f964d2f2f988a4a9b43434d5c5bd8fd633de1937617a67bf7e0366c5f01329
BLAKE2b-256 checksum
How to use checksums
107d420adbcacd55a99cb4d7a099963806aec38adfd37cacf04ecca8d817b09f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.9

Release files / autourgos_whisperstt-0.1.2-py3-none-any.whl

Download URL autourgos_whisperstt-0.1.2-py3-none-any.whl
Size 15.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
e86279e4bd3646d0f90c3fbb1dd2ac9ef917b0d0819cd62c4c74658b73967d13
BLAKE2b-256 checksum
How to use checksums
d8462dca3788faaee3f97b9d19ada73750525ef2cc88e7435b6fe5c445980529
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.9

Release history Release notifications | RSS feed

This release

0.1.2 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page