Production infrastructure for voice agents — memory, tool orchestration, and observability
Project description
VoxInfra
Production memory, tool safety, and call replay for LiveKit voice agents.
The Problem
- No memory — Your agent forgets users the moment the call ends, making every interaction feel cold and repetitive.
- Latency kills — A single CRM lookup timing out can freeze the audio stream and drop the call.
- No replay — When a call goes wrong, you have no structured trace to debug what the agent heard, recalled, or decided.
Install
pip install voxinfra
Quick Start
from dotenv import load_dotenv
from livekit.agents import Agent, AgentSession, JobContext, WorkerOptions, cli
from livekit.plugins import deepgram, elevenlabs, openai, silero
from voxinfra import VoxSession, vox_tool
load_dotenv()
vox = VoxSession() # reads MEM0_API_KEY, SUPABASE_URL from env
@vox_tool(sla_ms=200, fallback={"name": "there"})
async def get_user_info(user_id: str) -> dict:
return await crm.lookup(user_id)
class MyAgent(Agent):
async def on_enter(self):
call_id = str(__import__("uuid").uuid4())
self.session.userdata["call_id"] = call_id
vox.obs.start_call(call_id, self.session.userdata["user_id"])
async def on_exit(self):
await vox.end_call(self.session.userdata["call_id"])
async def llm_node(self, chat_ctx, tools, model_settings):
user_id = self.session.userdata["user_id"]
call_id = self.session.userdata["call_id"]
async with vox.turn(user_id=user_id, chat_ctx=chat_ctx, call_id=call_id) as turn:
await turn.call_tool(get_user_info, user_id)
return await turn.run_llm(tools, model_settings)
VoxMemory
Persistent cross-call memory backed by Mem0 (hosted) or Qdrant (self-hosted). At the start of every turn, VoxInfra semantically searches past conversations and injects the top results into the LLM context — automatically, in 3 lines or fewer to keep voice prompts lean.
# Automatic via vox.turn() — no extra code needed
async with vox.turn(user_id="u-123", chat_ctx=chat_ctx, call_id=call_id) as turn:
# Memory was recalled and injected before this line
return await turn.run_llm(tools, model_settings)
# Manual DPDP right-to-erasure
await vox.memory.delete_user("u-123")
VoxTools
A registry of async tools with per-tool SLA enforcement. If a tool exceeds its SLA, VoxInfra serves a cached result or fallback — the agent keeps talking and the call never drops.
@vox_tool(sla_ms=200, fallback={"status": "unavailable"})
async def get_appointments(user_id: str) -> dict:
return await calendar_api.fetch(user_id)
# Inside llm_node
async with vox.turn(...) as turn:
# If get_appointments takes >200ms, fallback is returned silently
result = await turn.call_tool(get_appointments, user_id)
The @vox_tool decorator is transparent — it only attaches metadata. The function remains directly callable without going through VoxTools.
VoxObs
Structured call traces persisted to a Supabase table (vox_call_traces). Every turn records user transcript, memory recall latency, LLM latency, TTS latency, and all tool call results with SLA breach flags. Use vox.end_call() to finalise and flush the trace.
# In Agent.on_exit
await vox.end_call(self.session.userdata["call_id"])
The Supabase write is always fire-and-forget — it never blocks the call path.
Config
All configuration via environment variables or constructor kwargs.
| Variable | Description | Required |
|---|---|---|
MEM0_API_KEY |
Mem0 hosted API key | Yes (mem0 backend) |
QDRANT_URL |
Qdrant server URL | Yes (qdrant backend) |
QDRANT_API_KEY |
Qdrant API key | No |
SUPABASE_URL |
Supabase project URL | No (obs disabled without it) |
SUPABASE_SERVICE_KEY |
Supabase service role key | No |
VOXINFRA_API_KEY |
Reserved for future cloud features | No |
# From env (default)
vox = VoxSession()
# Explicit
vox = VoxSession(
mem0_api_key="m0-...",
supabase_url="https://xyz.supabase.co",
supabase_key="service-key",
default_tool_sla_ms=250,
)
# Qdrant backend (self-hosted)
vox = VoxSession(
memory_backend="qdrant",
qdrant_url="http://localhost:6333",
qdrant_collection="my_agent_memories",
)
DPDP Compliance
For deployments subject to India's Digital Personal Data Protection Act, VoxInfra supports a fully self-hosted data path. Use the Qdrant backend pointed at a Mumbai-region instance to ensure all memory vectors stay within Indian jurisdiction. Call traces can be stored in a Supabase project provisioned in the ap-south-1 region. Users have the right to erasure under DPDP Section 13 — invoke it with a single call:
await vox.memory.delete_user(user_id)
This calls Mem0's delete_all or Qdrant's collection filter delete under the hood, removing all stored memories for that user permanently.
License
MIT © Ashu Singhania
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file voxinfra-0.1.0.tar.gz.
File metadata
- Download URL: voxinfra-0.1.0.tar.gz
- Upload date:
- Size: 13.5 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.5
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
b9bc8d7b389f100998a8d6649813c804d2ef25acc65a289515179fa742de7a76
|
|
| MD5 |
215654ca51f89533812f93be79bef56f
|
|
| BLAKE2b-256 |
321b5a39a9bf427e009d985beb0a1be2062b9fde10ff70cd6aff7e99818f80f9
|
File details
Details for the file voxinfra-0.1.0-py3-none-any.whl.
File metadata
- Download URL: voxinfra-0.1.0-py3-none-any.whl
- Upload date:
- Size: 11.3 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.5
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
fac8b1a5eb59f4c8827422b4acb164fdd911a46d02e89f699f82aafbaaff7250
|
|
| MD5 |
24791da4039a8d23aef80d538c29d0f6
|
|
| BLAKE2b-256 |
af70e8b98174095de29eb44cc3873bf30d873aad2675a989805803a3f9efce1e
|