Skip to main content

voice-agent-kit

CI License: Apache-2.0 Python 3.10+

An open-source, provider-agnostic, low-latency realtime voice agent framework in Python.

Build conversational voice AI for customer support, telephony, browser voice apps, and self-hosted local enterprise clusters without vendor lock-in or cloud intermediary proxies.


Key Highlights

  • Provider-agnostic: swap STT, LLM and TTS providers via config or register_provider without changing orchestration code.
  • OpenAI-compatible LLMs: one SSE-streaming adapter for Ollama, vLLM, LM Studio, OpenRouter, OpenAI and gateways.
  • Runs fully local (no API key): faster-whisper STT (optional extra) → Ollama → macOS say TTS (macOS only). No portable local TTS yet.
  • OS-agnostic core: pure-Python py3-none-any wheel with no mandatory OS, audio-hardware or GPU dependencies; platform-specific code lives only in optional providers (docs/PLATFORMS.md).
  • No project server: audio and credentials go directly to the endpoints you configure.
  • Barge-in: user speech cancels the in-flight LLM/TTS response and drains queued output (regression-tested).
  • Tool calling: @tool decorator with JSON-schema extraction, argument validation and deadlines.
  • Evaluation framework: voice-agent eval runs YAML scenarios in strictly separated synthetic / local / remote modes and reports WER/CER, latencies, reliability, barge-in and reply script; never ranks providers.
  • Deterministic offline tests: fake STT/LLM/TTS/VAD for GPU-free, key-free CI.
  • Languages: BCP-47 language metadata and per-provider language declarations; English/Hindi/Bengali scenarios included. No language is verified for a real model yet (see docs/LANGUAGES.md).

Not implemented: telephony transports (only interface stubs in voice_agent.telephony), microphone I/O, ML-based VAD.


Architecture Overview

flowchart TD
    subgraph Transports ["Transport Layer"]
        WS["WebSocketServerTransport (one shared conversation)"]
        WSS["WebSocketSessionServer (one agent per connection)"]
        MEM[In-memory / custom Transport]
    end

    subgraph CoreEngine ["VoiceAgent Orchestrator"]
        VAD[Energy VAD]
        STT[STT Adapter]
        BUS[Async Event Bus]
        LLM[Streaming LLM / OpenAI-Compatible]
        TOOLS[Tool Engine]
        TTS[TTS Adapter]
        DRAIN[Barge-In Cancellation]
    end

    Transports <-->|Audio Chunks| VAD
    VAD -->|Voice Activity| STT
    STT -->|Transcripts| BUS
    BUS <--> LLM
    LLM <--> TOOLS
    LLM -->|Tokens| TTS
    TTS -->|PCM Audio Chunks| Transports
    VAD -.->|Interruption Signal| DRAIN
    DRAIN -.->|Cancel Task & Flush Queue| Transports

30-Second Quickstart

1. Installation

# Core package (ultra-lightweight, < 25MB)
pip install voice-agent-kit

# With OpenAI support
pip install "voice-agent-kit[openai]"

# With the WebSocket server transports
pip install "voice-agent-kit[websocket]"

# With all adapters
pip install "voice-agent-kit[all]"

2. Basic Agent (Pure Python)

import asyncio
import os
from voice_agent import VoiceAgent, tool
from voice_agent.providers import OpenAICompatibleLLM, DeepgramSTT, ElevenLabsTTS
from voice_agent.transports import WebSocketSecurityConfig, WebSocketServerTransport


@tool(name="check_order", description="Look up shipping status of an order")
def check_order(order_id: str) -> dict:
    return {"order_id": order_id, "status": "Out for delivery"}


agent = VoiceAgent(
    stt=DeepgramSTT(api_key="..."),
    llm=OpenAICompatibleLLM(
        base_url="http://localhost:8000/v1",  # Local vLLM or Ollama instance
        model="Qwen/Qwen2.5-7B-Instruct",
        api_key="EMPTY",
    ),
    tts=ElevenLabsTTS(api_key="..."),
    system_prompt="You are an AI customer concierge. Keep answers concise.",
    # Set VOICE_AGENT_WS_TOKEN before starting. Bind to loopback; use a TLS proxy for remote clients.
    transport=WebSocketServerTransport(
        host="127.0.0.1",
        port=8765,
        security=WebSocketSecurityConfig(auth_token=os.environ["VOICE_AGENT_WS_TOKEN"]),
    ),
)

agent.register_tool(check_order)

if __name__ == "__main__":

    async def main() -> None:
        await agent.start()
        try:
            await asyncio.Event().wait()
        finally:
            await agent.stop()

    asyncio.run(main())

3. One conversation vs. many callers

WebSocketServerTransport belongs to one VoiceAgent and therefore one shared conversation. With the default max_connections=1 it serves a single caller. max_connections on this transport controls how many sockets may join that same conversation; it does not create independent callers. With max_connections > 1, every socket shares conversation history, agent state, tools, provider state, and the broadcast audio output (a UserWarning is emitted for this configuration). Use it only when several sockets really are the same session.

For independent concurrent callers, use WebSocketSessionServer, which builds a new VoiceAgent per connection. Its max_connections bounds the number of concurrent isolated sessions:

from voice_agent.transports import WebSocketConnectionTransport, WebSocketSessionServer


def make_agent(transport: WebSocketConnectionTransport) -> VoiceAgent:
    return VoiceAgent(stt=..., llm=..., tts=..., transport=transport)  # fresh providers/state per caller


server = WebSocketSessionServer(make_agent, security=WebSocketSecurityConfig(max_connections=8))
await server.start()
WebSocketServerTransport WebSocketSessionServer
Agents one, shared by all sockets one per connection (built by your factory)
Transport one, output broadcast to every socket one WebSocketConnectionTransport per connection
max_connections > 1 means extra sockets in the same shared conversation independent, isolated callers
History / tools / audio shared isolated per caller
Session reset when the first socket joins an idle transport and when the last one leaves each session starts fresh and is stopped when its connection ends

WebSocketSessionServer also rejects (close code 1011) a session whose agent uses an EventBus, VAD, memory, pipeline, STT/LLM/TTS provider, or tool still owned by another active session, bounds factory + agent.start() by security.session_setup_timeout_s, cancels setup if the caller disconnects or the server stops, and cancels any agent.stop() exceeding security.session_shutdown_timeout_s. Build providers and tools inside the factory; share one only if it declares session_shareable = True (@tool(session_shareable=True) for tools), i.e. keeps no per-caller state. Ownership also covers mutable objects reachable through wrappers (adapters, partials, closures, bound methods), so a fresh wrapper per session around one shared stateful backend is rejected too; see docs/API_DESIGN.md.

voice-agent run --ws uses WebSocketServerTransport. See examples/websocket_server.py for both server types (--sessions for the per-caller server).


Command Line Interface (CLI)

The package includes a comprehensive diagnostics and testing CLI:

# 1. Inspect Python runtime, installed codecs, and credential status
voice-agent doctor

# 2. Run deterministic turn-taking and barge-in test scenarios
voice-agent test --scenario barge-in

# 3. Benchmark latency: SYNTHETIC (fake providers, framework overhead only) or REAL (live providers)
voice-agent benchmark --trials 5
voice-agent benchmark --mode real --config prod.yaml --audio speech_16k.wav --network "describe it"

# 4. Run interactive local conversation demo
voice-agent demo

# 5. Launch agent declaratively from a YAML file (optionally over authenticated WebSocket)
voice-agent run config.yaml
VOICE_AGENT_WS_TOKEN=... voice-agent run config.yaml --ws

Declarative YAML Configuration

version: "1.0"
agent:
  name: "support-agent"
  system_prompt: "You are a customer service voice assistant."

stt:
  provider: "deepgram"
  model: "nova-2"

llm:
  provider: "openai-compatible"
  base_url: "http://localhost:8000/v1"
  model: "Qwen/Qwen2.5-7B-Instruct"

tts:
  provider: "elevenlabs"
  voice_id: "21m00Tcm4TlvDq8ikWAM"

realtime:
  allow_interruptions: true
  sample_rate: 16000

Development & Testing

# Clone the repository
git clone https://github.com/sandidas/voice-agent-kit.git
cd voice-agent-kit

# Create virtual environment
python3 -m venv .venv
source .venv/bin/activate

# Install in editable mode with development dependencies
pip install -e ".[dev,all]"

# Run offline unit and e2e test suite
pytest tests/unit tests/e2e

Documentation Index


License

This project is licensed under the Apache License 2.0. See the LICENSE file for details.


Project Status (v0.1, pre-release)

This is not production-ready. See docs/FINAL_AUDIT.md for which items are verified and which are not. Live cloud adapters (Deepgram, ElevenLabs, OpenAI) have not been exercised against real APIs by the maintainers. A fully local pipeline (faster-whisper base + Ollama llama3 + macOS say) has been run end to end; results, including unusable Hindi/Bengali recognition with Whisper base, are in docs/EVALUATION.md.

Metadata

Release files for voice-agent-kit 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for voice-agent-kit 0.1.0
File Size Uploaded
voice_agent_kit-0.1.0.tar.gz 410.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for voice-agent-kit 0.1.0
File Interpreter ABI Platform
voice_agent_kit-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 503.3 kB

Release files / voice_agent_kit-0.1.0.tar.gz

Download URL voice_agent_kit-0.1.0.tar.gz
Size 410.0 kB
Tags Source
SHA-256 checksum
How to use checksums
e23ab9541478fb2d821fd35340da72983548770f21339a58d8a9be89c8ef29c3
BLAKE2b-256 checksum
How to use checksums
74bbb8a31db016fb28a644452dff465a9d2da9d4ff87b51cd67c59b4a06fc7db
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 26, 2026.

Transparency log

Release files / voice_agent_kit-0.1.0-py3-none-any.whl

Download URL voice_agent_kit-0.1.0-py3-none-any.whl
Size 93.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
35afe578755e3985d90b4db3f556aff7c3e14fa5cc9d355807ef7e37b476d026
BLAKE2b-256 checksum
How to use checksums
790026c306dc0e4bacc0f98aa5548d794ef1bd26a71cb1283cb2c6d5c9f807b4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 26, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page