voice-agent
Reusable Python voice-agent components with Kokoro and Piper TTS, local STT, and an asynchronous OpenRouter LLM adapter. A turn-based conversation session connects these adapters.
Setup
Use uv and Python 3.12 or newer. Add only the providers you need to a consuming project, for example:
uv add "voice-agent-56[openrouter,stt]"
From this repository checkout, install the TTS providers with:
uv sync --locked --extra kokoro --extra piper
Select one extra to install only that provider. Add dependencies with uv add
and commit pyproject.toml with uv.lock. Run scripts with uv run.
TTS
from voice_agent.tts import create
engine = create("piper:en_US-lessac-medium", models_dir="/path/to/models")
for audio, sample_rate in engine.stream("Hello. How can I help?"):
print(len(audio), sample_rate)
Chunks contain mono float32 NumPy audio at the returned sample rate. Streaming is sentence-based. The application owns playback and resampling.
Supply model files through models_dir; they are not stored in Git:
- Kokoro needs
kokoro-v1.0.onnxandvoices-v1.0.bin. Its default voice isaf_heart. - Piper needs
<voice>.onnxand<voice>.onnx.json. Its default voice isen_US-lessac-medium.
STT
Install local recognition with uv sync --locked --extra stt.
The adapters accept mono float32 samples at 16 kHz and return a Transcript.
The application owns decoding, resampling, and recording limits.
from voice_agent.stt import create
engine = create("whisper", models_dir="models/stt", num_threads=2)
result = engine.transcribe(samples, language="en")
print(result.text)
Choose whisper for multilingual Whisper base or zipformer for English.
Whisper also accepts language=None for detection. Zipformer accepts only
English. Calls return completed-utterance results. They do not provide a live
microphone API. Whisper includes segment timestamps; Zipformer currently returns
text without segment timestamps. Callers must serialize access to each engine.
Model directories under models_dir are whisper-base and zipformer-en-full.
The checked-out model assets are in models/stt/.
See STT research and local measurements for the exact
models, versions, limitations, and benchmark command. Model files stay outside Git.
LLM through OpenRouter
In this repository checkout, install the provider and configure your key:
uv sync --locked --extra openrouter
cp .env.example .env
# Set OPENROUTER_API_KEY in .env, then run:
uv run --env-file .env --extra openrouter python scripts/chat_llm.py
The default model is stealth/space-bunny-alpha. This is the model described in
the supplied AI/ML API documentation,
and OpenRouter lists the same model ID.
Requests go directly to https://openrouter.ai/api/v1/chat/completions and require
an OpenRouter key, not an AI/ML API key. Set OPENROUTER_MODEL to change models,
or pass model= explicitly. The library reads environment variables; the
uv run --env-file option loads the example configuration for the script.
import asyncio
from contextlib import aclosing
from voice_agent.llm import create
async def main():
messages = [
{"role": "system", "content": "Reply briefly in natural spoken language. Avoid markdown."},
{"role": "user", "content": "Hello. Can you help me?"},
]
async with create("openrouter") as engine:
async with aclosing(engine.stream(messages)) as response:
async for text in response:
print(text, end="", flush=True)
# Or await engine.complete(messages) for one complete string.
asyncio.run(main())
stream() yields text deltas, which may be fragments of words. Buffer them into
sentences before passing them to TTS. The application owns the system prompt,
conversation history, sentence buffering, and playback. This adapter supports
text messages with system, user, and assistant roles; tool execution and full
audio conversation orchestration are not implemented.
Provider options are api_key, model, max_tokens with a default of 2048,
reasoning_effort with a default of low, and timeout with a default of 60
seconds of network inactivity. The token budget also covers reasoning. Use a
larger budget if the provider hits its limit. Set reasoning_effort=None to omit
the reasoning option for another model. Only answer content is yielded.
Use aclosing around a stream if you might stop early. Cancelling the task or
closing the iterator closes its HTTP response. Use the engine as an async
context manager, or call await engine.aclose() when finished. Requests are not
retried automatically. LLMError reports HTTP failures, timeouts, stream errors,
truncation, and empty answers. Partial text may already have been yielded when
an error occurs; the application decides what to play or retain in history.
Run the provider tests without credentials or model downloads:
uv run --locked --extra openrouter python -m unittest discover -s tests -v
Conversation sessions
voice_agent.session.ConversationSession provides turn-based orchestration and
isolated in-memory history. Supply a frozen SessionConfig, an asynchronous LLM
adapter, and async transcription and synthesis callbacks. The caller owns file
decoding, native-model concurrency, transport, deadlines, and session expiry.
Call begin() before consuming turn(audio, generation). It emits stage,
transcript, text delta, WAV audio, and completion events. Call acknowledge with
the generation only after playback completes. Cancel invalidates old generations
and closes active LLM work. It does not stop synchronous native inference that
is already running. Interrupted replies are marked in history.
This first session implementation uses completed English utterances and waits
for a full generated reply before synthesis. It does not implement streaming
microphone input, automatic barge-in, tool calls, or a remote worker protocol.
The live integration is in the sibling voice-lab application.
Design
See architecture, provider research, and NVIDIA model notes. Applications supply transport, context, and domain tools. Keep application setup in the consuming repository.
The package is licensed under Apache-2.0.
Metadata
Release files for voice-agent-56 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| voice_agent_56-0.1.0.tar.gz | 20.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| voice_agent_56-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 40.1 kB
Release files / voice_agent_56-0.1.0.tar.gz
| Download URL | voice_agent_56-0.1.0.tar.gz |
|---|---|
| Size | 20.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
ead583b00aed84ec6127b7c8e1fa4817907b2eb0031847877c9a27a9cd1dc696
|
|
BLAKE2b-256 checksum How to use checksums |
4e20135bff3a6b068c7a7a77ae614549864286e8319412573f022ebb9c1f9a68
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.12.19 {"installer":{"name":"uv","version":"0.12.19","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|
Release files / voice_agent_56-0.1.0-py3-none-any.whl
| Download URL | voice_agent_56-0.1.0-py3-none-any.whl |
|---|---|
| Size | 19.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
984ebebb590d2aa77b2697c0e34b17d9eb0f66b18a7b0189d068f3336147ca94
|
|
BLAKE2b-256 checksum How to use checksums |
bcc98f7b8f1a78a6188516fce3e3d04f75ae1d605a01a3c8dc55d7517ab1e48b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.12.19 {"installer":{"name":"uv","version":"0.12.19","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|