Skip to main content

voice-agent

Reusable Python voice-agent components with Kokoro and Piper TTS, local STT, and an asynchronous OpenRouter LLM adapter. A turn-based conversation session connects these adapters.

Setup

Use uv and Python 3.12 or newer. Add only the providers you need to a consuming project, for example:

uv add "voice-agent-56[openrouter,stt]"

From this repository checkout, install the TTS providers with:

uv sync --locked --extra kokoro --extra piper

Select one extra to install only that provider. Add dependencies with uv add and commit pyproject.toml with uv.lock. Run scripts with uv run.

TTS

from voice_agent.tts import create

engine = create("piper:en_US-lessac-medium", models_dir="/path/to/models")
for audio, sample_rate in engine.stream("Hello. How can I help?"):
    print(len(audio), sample_rate)

Chunks contain mono float32 NumPy audio at the returned sample rate. Streaming is sentence-based. The application owns playback and resampling.

Supply model files through models_dir; they are not stored in Git:

  • Kokoro needs kokoro-v1.0.onnx and voices-v1.0.bin. Its default voice is af_heart.
  • Piper needs <voice>.onnx and <voice>.onnx.json. Its default voice is en_US-lessac-medium.

STT

Install local recognition with uv sync --locked --extra stt. The adapters accept mono float32 samples at 16 kHz and return a Transcript. The application owns decoding, resampling, and recording limits.

from voice_agent.stt import create

engine = create("whisper", models_dir="models/stt", num_threads=2)
result = engine.transcribe(samples, language="en")
print(result.text)

Choose whisper for multilingual Whisper base or zipformer for English. Whisper also accepts language=None for detection. Zipformer accepts only English. Calls return completed-utterance results. They do not provide a live microphone API. Whisper includes segment timestamps; Zipformer currently returns text without segment timestamps. Callers must serialize access to each engine.

Model directories under models_dir are whisper-base and zipformer-en-full. The checked-out model assets are in models/stt/. See STT research and local measurements for the exact models, versions, limitations, and benchmark command. Model files stay outside Git.

LLM through OpenRouter

In this repository checkout, install the provider and configure your key:

uv sync --locked --extra openrouter
cp .env.example .env
# Set OPENROUTER_API_KEY in .env, then run:
uv run --env-file .env --extra openrouter python scripts/chat_llm.py

The default model is stealth/space-bunny-alpha. This is the model described in the supplied AI/ML API documentation, and OpenRouter lists the same model ID. Requests go directly to https://openrouter.ai/api/v1/chat/completions and require an OpenRouter key, not an AI/ML API key. Set OPENROUTER_MODEL to change models, or pass model= explicitly. The library reads environment variables; the uv run --env-file option loads the example configuration for the script.

import asyncio
from contextlib import aclosing
from voice_agent.llm import create

async def main():
    messages = [
        {"role": "system", "content": "Reply briefly in natural spoken language. Avoid markdown."},
        {"role": "user", "content": "Hello. Can you help me?"},
    ]
    async with create("openrouter") as engine:
        async with aclosing(engine.stream(messages)) as response:
            async for text in response:
                print(text, end="", flush=True)
        # Or await engine.complete(messages) for one complete string.

asyncio.run(main())

stream() yields text deltas, which may be fragments of words. Buffer them into sentences before passing them to TTS. The application owns the system prompt, conversation history, sentence buffering, and playback. This adapter supports text messages with system, user, and assistant roles. Tools work only on the non-streaming path: when a session supplies tool schemas, the adapter returns one completion carrying tool calls instead of a text stream. Audio turn orchestration belongs to voice_agent.session, described below.

Provider options are api_key, model, max_tokens with a default of 2048, reasoning_effort with a default of low, and timeout with a default of 60 seconds of network inactivity. The token budget also covers reasoning. Use a larger budget if the provider hits its limit. Set reasoning_effort=None to omit the reasoning option for another model. Only answer content is yielded.

Use aclosing around a stream if you might stop early. Cancelling the task or closing the iterator closes its HTTP response. Use the engine as an async context manager, or call await engine.aclose() when finished. Requests are not retried automatically. LLMError reports HTTP failures, timeouts, stream errors, truncation, and empty answers. Partial text may already have been yielded when an error occurs; the application decides what to play or retain in history.

Run the provider tests without credentials or model downloads:

uv run --locked --extra openrouter python -m unittest discover -s tests -v

Conversation sessions

voice_agent.session.ConversationSession provides turn-based orchestration and isolated in-memory history. Supply a frozen SessionConfig, an asynchronous LLM adapter, and async transcription and synthesis callbacks. The caller owns file decoding, native-model concurrency, transport, deadlines, and session expiry.

For an application-supplied greeting, call begin() and consume speak(text, generation) before the first caller turn. Like turn(), it emits audio and completion events and requires a playback acknowledgement. It does not add a fabricated caller message to history.

Call begin() before consuming turn(audio, generation). It emits stage, transcript, text delta, WAV audio, and completion events. Call acknowledge with the generation only after playback completes. Cancel invalidates old generations and closes active LLM work. It does not stop synchronous native inference that is already running. Interrupted replies are marked in history.

voice_agent.endpoint.EndpointDetector accepts 20 ms mono float32 frames at 16 kHz and emits speech-start and completed-utterance events. It uses an RMS threshold with configurable start, silence, pre-roll, and maximum-duration limits. It does not distinguish caller speech from echo or background voices; test and tune it against real call recordings before enabling barge-in.

This session implementation uses completed English utterances and waits for a full generated reply before synthesis. It does not implement streaming microphone input, automatic barge-in, or a remote worker protocol.

An application may pass typed tool schemas and an execute_tool callback to ConversationSession. The OpenRouter adapter uses a non-streaming completion when tools are available. The session emits tool_request, awaits the callback, and emits tool_result. It ends that turn without synthesizing model text when a tool is requested. The application must authorize the request and provide any speech the caller should hear before executing it. Only one tool request is supported per turn. The package does not perform campaign or call-control actions. The telephony consumer of these tools is the sibling vicidial-ai-agent repository. voice-lab also builds on this session and installs the same release.

Design

See architecture, provider research, and NVIDIA model notes. Applications supply transport, context, and domain tools. Keep application setup in the consuming repository.

The package is licensed under Apache-2.0.

Metadata

Release files for voice-agent-56 0.3.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for voice-agent-56 0.3.1
File Size Uploaded
voice_agent_56-0.3.1.tar.gz 27.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for voice-agent-56 0.3.1
File Interpreter ABI Platform
voice_agent_56-0.3.1-py3-none-any.whl Python 3 none any Details

Total release size: 50.9 kB

Release files / voice_agent_56-0.3.1.tar.gz

Download URL voice_agent_56-0.3.1.tar.gz
Size 27.6 kB
Tags Source
SHA-256 checksum
How to use checksums
790acf875e0b79007921d516e787b9e0ead8aa4e9a859a463d8e0c2b58625e9a
BLAKE2b-256 checksum
How to use checksums
5b64c81623ff1322d58583c98beb113c4ebb356cb8af175ad90e5f58c6a83adb
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.19 {"installer":{"name":"uv","version":"0.12.19","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release files / voice_agent_56-0.3.1-py3-none-any.whl

Download URL voice_agent_56-0.3.1-py3-none-any.whl
Size 23.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
ec86feaf5357612e6623ec20eebc848e5fdeeb81fac61b550d94c0b99f5a4068
BLAKE2b-256 checksum
How to use checksums
7372a63c9942d139a8a3fc55b413dfe082448fc31545b61dfac4260007044547
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.19 {"installer":{"name":"uv","version":"0.12.19","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release history Release notifications | RSS feed

This release

0.3.1 This release

2 release files

0.3.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page