voice-agent
Reusable Python voice-agent components with Kokoro and Piper TTS, local STT, and an asynchronous OpenRouter LLM adapter. A turn-based conversation session connects these adapters.
Setup
Use uv and Python 3.12 or newer. Add only the providers you need to a consuming project, for example:
uv add "voice-agent-56[openrouter,stt]"
From this repository checkout, install the TTS providers with:
uv sync --locked --extra kokoro --extra piper
Select one extra to install only that provider. Add dependencies with uv add
and commit pyproject.toml with uv.lock. Run scripts with uv run.
TTS
from voice_agent.tts import create
engine = create("piper:en_US-lessac-medium", models_dir="/path/to/models")
for audio, sample_rate in engine.stream("Hello. How can I help?"):
print(len(audio), sample_rate)
Chunks contain mono float32 NumPy audio at the returned sample rate. Streaming is sentence-based. The application owns playback and resampling.
Supply model files through models_dir; they are not stored in Git:
- Kokoro needs
kokoro-v1.0.onnxandvoices-v1.0.bin. Its default voice isaf_heart. - Piper needs
<voice>.onnxand<voice>.onnx.json. Its default voice isen_US-lessac-medium.
STT
Install local recognition with uv sync --locked --extra stt.
The adapters accept mono float32 samples at 16 kHz and return a Transcript.
The application owns decoding, resampling, and recording limits.
from voice_agent.stt import create
engine = create("whisper", models_dir="models/stt", num_threads=2)
result = engine.transcribe(samples, language="en")
print(result.text)
Choose whisper for multilingual Whisper base or zipformer for English.
Whisper also accepts language=None for detection. Zipformer accepts only
English. Calls return completed-utterance results. They do not provide a live
microphone API. Whisper includes segment timestamps; Zipformer currently returns
text without segment timestamps. Callers must serialize access to each engine.
Model directories under models_dir are whisper-base and zipformer-en-full.
The checked-out model assets are in models/stt/.
See STT research and local measurements for the exact
models, versions, limitations, and benchmark command. Model files stay outside Git.
LLM through OpenRouter
In this repository checkout, install the provider and configure your key:
uv sync --locked --extra openrouter
cp .env.example .env
# Set OPENROUTER_API_KEY in .env, then run:
uv run --env-file .env --extra openrouter python scripts/chat_llm.py
The default model is stealth/space-bunny-alpha. This is the model described in
the supplied AI/ML API documentation,
and OpenRouter lists the same model ID.
Requests go directly to https://openrouter.ai/api/v1/chat/completions and require
an OpenRouter key, not an AI/ML API key. Set OPENROUTER_MODEL to change models,
or pass model= explicitly. The library reads environment variables; the
uv run --env-file option loads the example configuration for the script.
import asyncio
from contextlib import aclosing
from voice_agent.llm import create
async def main():
messages = [
{"role": "system", "content": "Reply briefly in natural spoken language. Avoid markdown."},
{"role": "user", "content": "Hello. Can you help me?"},
]
async with create("openrouter") as engine:
async with aclosing(engine.stream(messages)) as response:
async for text in response:
print(text, end="", flush=True)
# Or await engine.complete(messages) for one complete string.
asyncio.run(main())
stream() yields text deltas, which may be fragments of words. Buffer them into
sentences before passing them to TTS. The application owns the system prompt,
conversation history, sentence buffering, and playback. This adapter supports
text messages with system, user, and assistant roles. Tools work only on the
non-streaming path: when a session supplies tool schemas, the adapter returns one
completion carrying tool calls instead of a text stream. Audio turn orchestration
belongs to voice_agent.session, described below.
Provider options are api_key, model, max_tokens with a default of 2048,
reasoning_effort with a default of low, and timeout with a default of 60
seconds of network inactivity. The token budget also covers reasoning. Use a
larger budget if the provider hits its limit. Set reasoning_effort=None to omit
the reasoning option for another model. Only answer content is yielded.
Use aclosing around a stream if you might stop early. Cancelling the task or
closing the iterator closes its HTTP response. Use the engine as an async
context manager, or call await engine.aclose() when finished. Requests are not
retried automatically. LLMError reports HTTP failures, timeouts, stream errors,
truncation, and empty answers. Partial text may already have been yielded when
an error occurs; the application decides what to play or retain in history.
Run the provider tests without credentials or model downloads:
uv run --locked --extra openrouter python -m unittest discover -s tests -v
Conversation sessions
voice_agent.session.ConversationSession provides turn-based orchestration and
isolated in-memory history. Supply a frozen SessionConfig, an asynchronous LLM
adapter, and async transcription and synthesis callbacks. The caller owns file
decoding, native-model concurrency, transport, deadlines, and session expiry.
For an application-supplied greeting, call begin() and consume
speak(text, generation) before the first caller turn. Like turn(), it emits
audio and completion events and requires a playback acknowledgement. It does not
add a fabricated caller message to history.
Call begin() before consuming turn(audio, generation). It emits stage,
transcript, text delta, WAV audio, and completion events. Call acknowledge with
the generation only after playback completes. Cancel invalidates old generations
and closes active LLM work. It does not stop synchronous native inference that
is already running. Interrupted replies are marked in history.
voice_agent.endpoint.EndpointDetector accepts 20 ms mono float32 frames at
16 kHz and emits speech-start and completed-utterance events. It uses an RMS
threshold with configurable start, silence, pre-roll, and maximum-duration
limits. It does not distinguish caller speech from echo or background voices;
test and tune it against real call recordings before enabling barge-in.
This session implementation uses completed English utterances and waits for a full generated reply before synthesis. It does not implement streaming microphone input, automatic barge-in, or a remote worker protocol.
An application may pass typed tool schemas and an execute_tool callback to
ConversationSession. The OpenRouter adapter uses a non-streaming completion
when tools are available. The session emits tool_request, awaits the callback,
and emits tool_result. It ends that turn without synthesizing model text when
a tool is requested. The application must authorize the request and provide any
speech the caller should hear before executing it. Only one tool request is
supported per turn. The package does not perform campaign or call-control actions.
The telephony consumer of these tools is the sibling vicidial-ai-agent
repository. voice-lab also builds on this session and installs the same release.
Design
See architecture, provider research, and NVIDIA model notes. Applications supply transport, context, and domain tools. Keep application setup in the consuming repository.
The package is licensed under Apache-2.0.
Metadata
Release files for voice-agent-56 0.3.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| voice_agent_56-0.3.1.tar.gz | 27.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| voice_agent_56-0.3.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 50.9 kB
Release files / voice_agent_56-0.3.1.tar.gz
| Download URL | voice_agent_56-0.3.1.tar.gz |
|---|---|
| Size | 27.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
790acf875e0b79007921d516e787b9e0ead8aa4e9a859a463d8e0c2b58625e9a
|
|
BLAKE2b-256 checksum How to use checksums |
5b64c81623ff1322d58583c98beb113c4ebb356cb8af175ad90e5f58c6a83adb
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.12.19 {"installer":{"name":"uv","version":"0.12.19","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|
Release files / voice_agent_56-0.3.1-py3-none-any.whl
| Download URL | voice_agent_56-0.3.1-py3-none-any.whl |
|---|---|
| Size | 23.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
ec86feaf5357612e6623ec20eebc848e5fdeeb81fac61b550d94c0b99f5a4068
|
|
BLAKE2b-256 checksum How to use checksums |
7372a63c9942d139a8a3fc55b413dfe082448fc31545b61dfac4260007044547
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.12.19 {"installer":{"name":"uv","version":"0.12.19","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|