autourgos-windowstt
Offline speech-to-text for the Autourgos framework using Windows' own built-in SAPI recognizer, via pywin32 COM automation. No network call, no API key, no model download — it uses whatever recognizer is already installed on the machine (Control Panel > Speech Recognition). Windows only.
For meaningfully higher accuracy at the cost of a one-time model download (and true cross-platform support), see the sibling package autourgos-whisperstt.
from autourgos_micinput import MicrophoneStream
from autourgos_windowstt import WindowsSTT
stt = WindowsSTT()
async def listen(seconds: float = 2.0):
frames_needed = int(seconds * 1000 / 100) # default chunk_ms=100
chunks = []
async with MicrophoneStream(sample_rate=16000) as mic:
async for chunk in mic:
chunks.append(chunk)
if len(chunks) >= frames_needed:
break
return stt.transcribe(b"".join(chunks), sample_rate=16000)
# text = asyncio.run(listen())
Install
pip install "autourgos-windowstt[win]"
pywin32 is required to actually call .transcribe()/.atranscribe() and is gated behind the win extra — import autourgos_windowstt alone never requires it. Windows only; requires Python 3.10+.
Usage
WindowsSTT takes raw 16-bit PCM bytes (exactly what autourgos-micinput's MicrophoneStream yields) and returns transcribed text:
stt = WindowsSTT()
text = stt.transcribe(pcm_bytes, sample_rate=16000) # sync, blocking
text = await stt.atranscribe(pcm_bytes, sample_rate=16000) # async, offloads to a worker thread
transcribe() is blocking — it pumps Windows COM messages on the calling thread until SAPI reports the end of the audio stream or timeout elapses (default 15s). Use atranscribe() from inside an event loop (e.g. right after capturing from MicrophoneStream) instead of calling transcribe() directly, same reasoning as SpeakerPlayer.awrite() in autourgos-live.
With autourgos-openaichat / autourgos-responses / autourgos-agent
None of those packages depend on this one (same reasoning as autourgos-micinput — no forced dependency on callers who don't need local speech). Wire them together yourself:
from autourgos_micinput import MicrophoneStream
from autourgos_windowstt import WindowsSTT
from autourgos_openaichat import OpenAIChatModel
stt = WindowsSTT()
llm = OpenAIChatModel(model="gpt-4o")
async def voice_turn():
chunks = []
async with MicrophoneStream(sample_rate=16000) as mic:
async for chunk in mic:
chunks.append(chunk)
if len(chunks) >= 20: # ~2s
break
text = await stt.atranscribe(b"".join(chunks), sample_rate=16000)
return llm.invoke(text)
API Reference
WindowsSTT
| Method | Description |
|---|---|
transcribe(pcm_bytes, *, sample_rate=16000, channels=1, timeout=15.0) -> str |
Blocking. Writes pcm_bytes to a temp WAV, runs it through SAPI's dictation recognizer, returns the recognized text ("" if nothing was recognized, e.g. silence). |
atranscribe(pcm_bytes, *, sample_rate=16000, channels=1, timeout=15.0) -> str |
Async-safe equivalent — runs transcribe() in a worker thread. |
sample_rate accepts any of 8000/11025/12000/16000/22050/24000/32000/44100/48000 (SAPI's supported file-stream rates); other values fall back to 16000, which is MicrophoneStream's own default — the common case needs no thought.
Errors (autourgos_windowstt)
| Name | Raised when |
|---|---|
WindowSTTError |
Base class |
WindowSTTUnavailableError |
Not running on Windows, or pywin32 isn't installed |
Accuracy note
SAPI is a legacy, non-neural dictation engine — noticeably less accurate than a modern model. Live-tested round-trip (same SAPI engine's own text-to-speech, fed back through this package's recognition): spoken "Testing one two three, this is a speech recognition check." came back as "Testing 1 to 3 this is a speech recognition check" — usably correct, but with real transcription errors, and no punctuation. autourgos-whisperstt transcribed the identical audio as "Testing 1, 2, 3, this is a speech recognition check." — near-perfect. Use this package when you want zero setup / zero download / instant offline results and can tolerate lower accuracy; use autourgos-whisperstt when accuracy matters more than the one-time model download.
License
Apache License 2.0, Copyright (c) 2026 Jitin Kumar Sengar
Metadata
Release files for autourgos-windowstt 0.1.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| autourgos_windowstt-0.1.2.tar.gz | 20.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| autourgos_windowstt-0.1.2-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 36.5 kB
Release files / autourgos_windowstt-0.1.2.tar.gz
| Download URL | autourgos_windowstt-0.1.2.tar.gz |
|---|---|
| Size | 20.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
de3a3b05e654ac926421292bfc8d59bb8e571fa720f562288678dbe674df9cc8
|
|
BLAKE2b-256 checksum How to use checksums |
c5736cfaa959e1b225d9723185c8ee7e7c605353b5ebe8c21134a49ef360dc20
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.9
|
Release files / autourgos_windowstt-0.1.2-py3-none-any.whl
| Download URL | autourgos_windowstt-0.1.2-py3-none-any.whl |
|---|---|
| Size | 16.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
910fca264ce66dfc7b796765d65895f1bc2cba19efa1c59d25ee8fe7abc5e682
|
|
BLAKE2b-256 checksum How to use checksums |
d3f198c826e5770b4b41e3129153e0d1daae139ee4bdd5ce36e79a81bb4f4e32
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.9
|