Skip to main content

Voice Connector middleware for LangGraph agents

Project description

AutoVox: Voice Connector for LangGraph Agents

AutoVox is a middleware package that enables real-time voice conversations with LangGraph agents by connecting voice engines (OpenAI, Gemini) to your AI workflows, providing a streamlined interface for bidirectional audio communication.

Features

  • 🎙️ Real-Time Conversations: True bidirectional voice conversations with streaming audio
  • 🔊 Multiple Voice Engines: Support for OpenAI and Google Gemini real-time voice APIs
  • 🧠 LangGraph Integration: Connect any LangGraph agent or multi-agent supervisor to voice capabilities
  • 🖥️ Web Interface: Browser-based UI for voice interactions with no coding required
  • 📱 Cross-Platform: Run on desktop or integrate into web applications
  • 🛠️ Easy Customization: Configure voices, models, and system instructions

Examples

Example Description File
Basic Voice Conversation Simple real-time conversation with a voice engine examples/realtime_conversation.py
Simple LangGraph Agent Connect a basic LangGraph agent to voice examples/simple_langgraph.py
LangGraph Conversation Advanced conversation with a LangGraph agent examples/langgraph_conversation.py
LangGraph Supervisor Connect a multi-agent supervisor to voice examples/langgraph_supervisor.py
Web Interface Browser-based voice interface with LangGraph integration examples/web/

Installation

pip install autovox

Quick Start: Real-Time Voice Conversation

Here's how to create a real-time voice conversation with a basic voice engine:

import asyncio
import os
from autovox.engines.openai_realtime import OpenAIRealTime
from autovox.core.protocol import VoiceSession, StreamSettings

async def main():
    # Create and initialize the voice engine
    engine = OpenAIRealTime()
    await engine.initialize(os.environ["OPENAI_API_KEY"])

    # Create a voice session with callbacks
    session = VoiceSession(
        engine=engine,
        on_transcription=lambda text: print(f"User: {text}"),
        on_response_chunk=lambda content: print_response(content),
        on_error=lambda error: print(f"Error: {error}")
    )

    # Configure and start the session
    settings = StreamSettings(
        voice="alloy",    # Voice for the AI assistant
        model="gpt-4o"    # Model for processing
    )
    await session.start(settings)

    print("Voice session started. Speak into your microphone...")

    # Keep the session running
    try:
        while True:
            await asyncio.sleep(1)
    except KeyboardInterrupt:
        print("Ending session...")
    finally:
        await session.stop()

def print_response(content):
    if isinstance(content, str):
        print(f"AI: {content}", end="", flush=True)

if __name__ == "__main__":
    asyncio.run(main())

Voice Engines

AutoVox supports the following real-time voice engines:

OpenAI RealTime

Enables bidirectional real-time conversations using OpenAI's WebSocket API:

from autovox.engines.openai_realtime import OpenAIRealTime

engine = OpenAIRealTime()
await engine.initialize(api_key)

Features:

  • Bidirectional real-time conversations
  • Full-duplex communication
  • Voice interruptions
  • Streaming transcriptions
  • Multiple voice options

Gemini RealTime

Leverages Google's Gemini model for real-time voice interactions:

from autovox.engines.gemini_realtime import GeminiRealTime

engine = GeminiRealTime()
await engine.initialize(api_key)

Features:

  • Real-time bidirectional audio streaming
  • Multiple voice options
  • Integrated with Gemini's multimodal capabilities

LangGraph Integration

AutoVox's core functionality is its seamless integration with LangGraph, making it easy to connect any LangGraph agent to real-time voice capabilities:

import asyncio
import os
from langchain_openai import ChatOpenAI
from langgraph.graph import StateGraph
from autovox.core.protocol import ConnectionType, EngineConfig
from autovox.agents.langgraph import LangGraphConnector, VoiceSessionConfig

async def main():
    # 1. Create a LangGraph agent with tools
    # ... your LangGraph agent code here ...
    agent = create_your_langgraph_agent()

    # 2. Configure the voice engine
    engine_config = EngineConfig(
        engine_type="openai",
        api_key=os.environ["OPENAI_API_KEY"],
        connection_type=ConnectionType.WEBSOCKET,
        settings={"model": "gpt-4o"}
    )

    # 3. Connect the agent to the voice engine
    connector = await LangGraphConnector.create(engine_config, agent)

    # 4. Configure the voice session
    config = VoiceSessionConfig(
        voice="alloy",
        model="gpt-4o",
        system_prompt="You are a helpful voice assistant who responds concisely."
    )

    # 5. Set up callbacks
    callbacks = {
        "on_transcription": lambda text: print(f"User: {text}"),
        "on_thinking": lambda thought: print(f"Thinking: {thought}"),
        "on_response_start": lambda: print("AI: ", end=""),
        "on_response_chunk": lambda content: print(content if isinstance(content, str) else "", end="", flush=True),
        "on_response_end": lambda: print("\n"),
        "on_error": lambda error: print(f"Error: {error}")
    }

    # 6. Start a real-time voice session
    session = await connector.start_voice_session(config, callbacks)

    # 7. Keep the session running
    try:
        print("Voice session started. Speak into your microphone...")
        while True:
            await asyncio.sleep(1)
    except KeyboardInterrupt:
        print("Ending session...")
    finally:
        await session.stop()

if __name__ == "__main__":
    asyncio.run(main())

Key features of LangGraph integration:

  • Use the full power of LangGraph with voice interactions
  • Access agent "thinking" steps for visibility into reasoning
  • Support for LangGraph Supervisors with multi-agent orchestration
  • Customizable voice settings per session
  • Works with all supported voice engines

Multi-Agent Supervisor Integration

AutoVox provides special integration with LangGraph's Multi-Agent Supervisor for orchestrating complex agent workflows:

from autovox.agents.langgraph import SupervisorConnector

# After creating your multi-agent supervisor...
supervisor = create_multi_agent_supervisor()

# Connect it to voice
connector = await SupervisorConnector.create(engine_config, supervisor)

# Start a voice session
session = await connector.start_voice_session(config, callbacks)

This allows users to create powerful voice interfaces to multi-agent systems that can:

  • Decompose complex tasks across specialized agents
  • Coordinate multiple experts to solve problems
  • Track reasoning across different agent roles
  • Provide unified responses through voice

Web Interface

The package includes a web-based interface for voice conversations:

# Set your API keys in .env file first (create a .env file in examples/web)
OPENAI_API_KEY=your_openai_key_here
GEMINI_API_KEY=your_gemini_key_here

Running the Web Interface

Unix/Linux/macOS

# Navigate to the web example directory
cd examples/web

# Make the script executable (if needed)
chmod +x run.sh

# Run the server
./run.sh

Windows

# Navigate to the web example directory
cd examples\web

# Run the server
run.bat

Manual Setup

# Install required dependencies
pip install fastapi uvicorn websockets python-dotenv autovox

# Run the web server
python examples/web/server.py

Then open your browser to http://localhost:8000 to interact with the voice assistant.

Features:

  • Browser-based UI for voice conversations
  • Support for both OpenAI and Gemini engines
  • No coding required to use
  • Real-time audio streaming and responses

License

MIT

Acknowledgements

The AutoVox project was inspired by:

  • OpenAI's real-time voice API
  • Google's Gemini API
  • LangGraph for agent orchestration

Voice Interaction Basics

AutoVox provides a simple, unified interface for voice interactions:

from autovox.core.protocol import VoiceSession

# Create a session with callbacks
session = VoiceSession(
    engine=engine,
    on_transcription=lambda text: print(f"User said: {text}"),
    on_response_chunk=lambda chunk: print(f"AI: {chunk}", end=""),
    on_error=lambda error: print(f"Error: {error}")
)

# Start the session
await session.start()

# Send audio data
await session.send_audio(audio_bytes)

# Stop the session when done
await session.stop()

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distributions

No source distribution files available for this release.See tutorial on generating distribution archives.

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

autovox-0.0.1-py3-none-any.whl (21.6 kB view details)

Uploaded Python 3

File details

Details for the file autovox-0.0.1-py3-none-any.whl.

File metadata

  • Download URL: autovox-0.0.1-py3-none-any.whl
  • Upload date:
  • Size: 21.6 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.12.2

File hashes

Hashes for autovox-0.0.1-py3-none-any.whl
Algorithm Hash digest
SHA256 5854d6ad9942a6a5499f538ffca528ef1d886b680e3b7a2c93a0d9531241dc96
MD5 c0c2a8c9d307212be38eede39735e89d
BLAKE2b-256 8a5038fabf189db0b5232c6818176f8abb2a06253e0bc794032cec9eb303ddc7

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page