voice-agent-kit
An open-source, provider-agnostic, low-latency realtime voice agent framework in Python.
Build conversational voice AI for customer support, telephony, browser voice apps, and self-hosted local enterprise clusters without vendor lock-in or cloud intermediary proxies.
Key Highlights
- Provider-agnostic: swap STT, LLM and TTS providers via config or
register_providerwithout changing orchestration code. - OpenAI-compatible LLMs: one SSE-streaming adapter for Ollama, vLLM, LM Studio, OpenRouter, OpenAI and gateways.
- Runs fully local (no API key): faster-whisper STT (optional extra) → Ollama → macOS
sayTTS (macOS only). No portable local TTS yet. - OS-agnostic core: pure-Python
py3-none-anywheel with no mandatory OS, audio-hardware or GPU dependencies; platform-specific code lives only in optional providers (docs/PLATFORMS.md). - No project server: audio and credentials go directly to the endpoints you configure.
- Barge-in: user speech cancels the in-flight LLM/TTS response and drains queued output (regression-tested).
- Tool calling:
@tooldecorator with JSON-schema extraction, argument validation and deadlines. - Evaluation framework:
voice-agent evalruns YAML scenarios in strictly separatedsynthetic/local/remotemodes and reports WER/CER, latencies, reliability, barge-in and reply script; never ranks providers. - Deterministic offline tests: fake STT/LLM/TTS/VAD for GPU-free, key-free CI.
- Languages: BCP-47 language metadata and per-provider language declarations; English/Hindi/Bengali scenarios included. No language is verified for a real model yet (see docs/LANGUAGES.md).
Not implemented: telephony transports (only interface stubs in voice_agent.telephony), microphone I/O, ML-based VAD.
Architecture Overview
flowchart TD
subgraph Transports ["Transport Layer"]
WS["WebSocketServerTransport (one shared conversation)"]
WSS["WebSocketSessionServer (one agent per connection)"]
MEM[In-memory / custom Transport]
end
subgraph CoreEngine ["VoiceAgent Orchestrator"]
VAD[Energy VAD]
STT[STT Adapter]
BUS[Async Event Bus]
LLM[Streaming LLM / OpenAI-Compatible]
TOOLS[Tool Engine]
TTS[TTS Adapter]
DRAIN[Barge-In Cancellation]
end
Transports <-->|Audio Chunks| VAD
VAD -->|Voice Activity| STT
STT -->|Transcripts| BUS
BUS <--> LLM
LLM <--> TOOLS
LLM -->|Tokens| TTS
TTS -->|PCM Audio Chunks| Transports
VAD -.->|Interruption Signal| DRAIN
DRAIN -.->|Cancel Task & Flush Queue| Transports
30-Second Quickstart
1. Installation
# Core package (ultra-lightweight, < 25MB)
pip install voice-agent-kit
# With OpenAI support
pip install "voice-agent-kit[openai]"
# With the WebSocket server transports
pip install "voice-agent-kit[websocket]"
# With all adapters
pip install "voice-agent-kit[all]"
2. Basic Agent (Pure Python)
import asyncio
import os
from voice_agent import VoiceAgent, tool
from voice_agent.providers import OpenAICompatibleLLM, DeepgramSTT, ElevenLabsTTS
from voice_agent.transports import WebSocketSecurityConfig, WebSocketServerTransport
@tool(name="check_order", description="Look up shipping status of an order")
def check_order(order_id: str) -> dict:
return {"order_id": order_id, "status": "Out for delivery"}
agent = VoiceAgent(
stt=DeepgramSTT(api_key="..."),
llm=OpenAICompatibleLLM(
base_url="http://localhost:8000/v1", # Local vLLM or Ollama instance
model="Qwen/Qwen2.5-7B-Instruct",
api_key="EMPTY",
),
tts=ElevenLabsTTS(api_key="..."),
system_prompt="You are an AI customer concierge. Keep answers concise.",
# Set VOICE_AGENT_WS_TOKEN before starting. Bind to loopback; use a TLS proxy for remote clients.
transport=WebSocketServerTransport(
host="127.0.0.1",
port=8765,
security=WebSocketSecurityConfig(auth_token=os.environ["VOICE_AGENT_WS_TOKEN"]),
),
)
agent.register_tool(check_order)
if __name__ == "__main__":
async def main() -> None:
await agent.start()
try:
await asyncio.Event().wait()
finally:
await agent.stop()
asyncio.run(main())
3. One conversation vs. many callers
WebSocketServerTransport belongs to one VoiceAgent and therefore one shared conversation.
With the default max_connections=1 it serves a single caller. max_connections on this transport
controls how many sockets may join that same conversation; it does not create independent
callers. With max_connections > 1, every socket shares conversation history, agent state, tools,
provider state, and the broadcast audio output (a UserWarning is emitted for this configuration).
Use it only when several sockets really are the same session.
For independent concurrent callers, use WebSocketSessionServer, which builds a new VoiceAgent
per connection. Its max_connections bounds the number of concurrent isolated sessions:
from voice_agent.transports import WebSocketConnectionTransport, WebSocketSessionServer
def make_agent(transport: WebSocketConnectionTransport) -> VoiceAgent:
return VoiceAgent(stt=..., llm=..., tts=..., transport=transport) # fresh providers/state per caller
server = WebSocketSessionServer(make_agent, security=WebSocketSecurityConfig(max_connections=8))
await server.start()
WebSocketServerTransport |
WebSocketSessionServer |
|
|---|---|---|
| Agents | one, shared by all sockets | one per connection (built by your factory) |
| Transport | one, output broadcast to every socket | one WebSocketConnectionTransport per connection |
max_connections > 1 means |
extra sockets in the same shared conversation | independent, isolated callers |
| History / tools / audio | shared | isolated per caller |
| Session reset | when the first socket joins an idle transport and when the last one leaves | each session starts fresh and is stopped when its connection ends |
WebSocketSessionServer also rejects (close code 1011) a session whose agent uses an EventBus, VAD,
memory, pipeline, STT/LLM/TTS provider, or tool still owned by another active session, bounds factory +
agent.start() by security.session_setup_timeout_s, cancels setup if the caller disconnects or the
server stops, and cancels any agent.stop() exceeding security.session_shutdown_timeout_s. Build
providers and tools inside the factory; share one only if it declares session_shareable = True
(@tool(session_shareable=True) for tools), i.e. keeps no per-caller state. Ownership also covers mutable
objects reachable through wrappers (adapters, partials, closures, bound methods), so a fresh wrapper per
session around one shared stateful backend is rejected too; see docs/API_DESIGN.md.
voice-agent run --ws uses WebSocketServerTransport. See
examples/websocket_server.py for both server types (--sessions for the per-caller server).
Command Line Interface (CLI)
The package includes a comprehensive diagnostics and testing CLI:
# 1. Inspect Python runtime, installed codecs, and credential status
voice-agent doctor
# 2. Run deterministic turn-taking and barge-in test scenarios
voice-agent test --scenario barge-in
# 3. Benchmark latency: SYNTHETIC (fake providers, framework overhead only) or REAL (live providers)
voice-agent benchmark --trials 5
voice-agent benchmark --mode real --config prod.yaml --audio speech_16k.wav --network "describe it"
# 4. Run interactive local conversation demo
voice-agent demo
# 5. Launch agent declaratively from a YAML file (optionally over authenticated WebSocket)
voice-agent run config.yaml
VOICE_AGENT_WS_TOKEN=... voice-agent run config.yaml --ws
Declarative YAML Configuration
version: "1.0"
agent:
name: "support-agent"
system_prompt: "You are a customer service voice assistant."
stt:
provider: "deepgram"
model: "nova-2"
llm:
provider: "openai-compatible"
base_url: "http://localhost:8000/v1"
model: "Qwen/Qwen2.5-7B-Instruct"
tts:
provider: "elevenlabs"
voice_id: "21m00Tcm4TlvDq8ikWAM"
realtime:
allow_interruptions: true
sample_rate: 16000
Development & Testing
# Clone the repository
git clone https://github.com/sandidas/voice-agent-kit.git
cd voice-agent-kit
# Create virtual environment
python3 -m venv .venv
source .venv/bin/activate
# Install in editable mode with development dependencies
pip install -e ".[dev,all]"
# Run offline unit and e2e test suite
pytest tests/unit tests/e2e
Documentation Index
- Ecosystem Research & Analysis
- Architecture & ADRs
- API Design Specification
- Provider System Specification
- Realtime Audio Pipeline
- Testing Strategy
- Security & Telemetry Policy
- Deployment Models
- Benchmarking Methodology
- Project Roadmap
- Developer & Agent Guidelines
- Contributing
- Evaluation · Languages · Local models
- Adding a provider · Adding a language
- Platform support
License
This project is licensed under the Apache License 2.0. See the LICENSE file for details.
Project Status (v0.1, pre-release)
This is not production-ready. See docs/FINAL_AUDIT.md for which
items are verified and which are not. Live cloud adapters (Deepgram, ElevenLabs, OpenAI) have
not been exercised against real APIs by the maintainers. A fully local pipeline
(faster-whisper base + Ollama llama3 + macOS say) has been run end to end; results, including
unusable Hindi/Bengali recognition with Whisper base, are in docs/EVALUATION.md.
Metadata
Release files for voice-agent-kit 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| voice_agent_kit-0.1.0.tar.gz | 410.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| voice_agent_kit-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 503.3 kB
Release files / voice_agent_kit-0.1.0.tar.gz
| Download URL | voice_agent_kit-0.1.0.tar.gz |
|---|---|
| Size | 410.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
e23ab9541478fb2d821fd35340da72983548770f21339a58d8a9be89c8ef29c3
|
|
BLAKE2b-256 checksum How to use checksums |
74bbb8a31db016fb28a644452dff465a9d2da9d4ff87b51cd67c59b4a06fc7db
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 26, 2026.
Transparency logRelease files / voice_agent_kit-0.1.0-py3-none-any.whl
| Download URL | voice_agent_kit-0.1.0-py3-none-any.whl |
|---|---|
| Size | 93.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
35afe578755e3985d90b4db3f556aff7c3e14fa5cc9d355807ef7e37b476d026
|
|
BLAKE2b-256 checksum How to use checksums |
790026c306dc0e4bacc0f98aa5548d794ef1bd26a71cb1283cb2c6d5c9f807b4
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 26, 2026.
Transparency log