Skip to main content

VoiceEval SDK (Python)

Python License OpenTelemetry

VoiceEval is an enterprise-grade observability and evaluation SDK for Voice Agents and LLM-powered applications. Built on OpenTelemetry, it provides zero-config auto-instrumentation with detailed tracing, latency breakdown, and cost analysis.

Key Features

  • Zero-Config Auto-Instrumentation: Automatically traces calls from major LLM providers (OpenAI, Anthropic, Google Gemini) and LiveKit Agents — no code changes needed.
  • LiveKit Native: Automatically integrates with LiveKit's tracing infrastructure. Just initialize the Client and all agent spans are captured.
  • Selective Monitoring: Control which calls are traced with auto_monitor, sample_rate, monitor_call(), and skip_call().
  • High Performance: Built on OpenTelemetry with async batch exports (OTLP/HTTP), ensuring negligible runtime overhead.

Installation

pip install voiceeval-sdk
# or
uv add voiceeval-sdk

Quickstart

1. Initialize the Client

Add a single Client(...) call at the top of your agent file. This sets up OTel tracing and auto-instruments all installed LLM libraries and LiveKit.

from voiceeval import Client

client = Client(
    api_key="your_voiceeval_api_key",   # or set VOICE_EVAL_API_KEY env var
    agent_name="my-booking-agent",      # identifies this agent in the dashboard
)

2. LiveKit Agent Example

from livekit.agents import Agent, AgentSession, JobContext, cli
from voiceeval import Client

# Initialize VoiceEval — auto-instruments all LLM calls and LiveKit spans
client = Client(
    api_key="your_voiceeval_api_key",
    agent_name="my-booking-agent",
)

class MyAgent(Agent):
    def __init__(self):
        super().__init__(instructions="You are a helpful voice assistant.")

@server.rtc_session(agent_name="my-agent")
async def entrypoint(ctx: JobContext):
    session = AgentSession(
        stt=...,
        llm=...,
        tts=...,
    )
    await session.start(agent=MyAgent(), room=ctx.room)
    await ctx.connect()

3. Standalone LLM Example

Works without LiveKit too — any OpenAI/Anthropic/Gemini calls are automatically traced:

from voiceeval import Client
from openai import OpenAI

client = Client(api_key="your_voiceeval_api_key")

openai_client = OpenAI()
response = openai_client.chat.completions.create(
    model="gpt-4o",
    messages=[{"role": "user", "content": "Hello world"}]
)
# Trace is automatically captured and exported

Client Options

Parameter Type Default Description
api_key str VOICE_EVAL_API_KEY env var Your VoiceEval API key
base_url str https://api.voiceeval.com/v1/traces VoiceEval ingestion endpoint
agent_name str None Agent identifier shown in the dashboard
auto_monitor bool True Monitor all calls automatically
sample_rate float 1.0 Fraction of calls to monitor (0.0 to 1.0)
span_post_processors list None Custom span post-processing functions

Selective Monitoring

By default, every call is monitored (auto_monitor=True). You can control this at the client level or per-call.

Sample a fraction of calls

client = Client(
    api_key="your_voiceeval_api_key",
    agent_name="my-booking-agent",
    sample_rate=0.1,  # Randomly monitor 10% of calls
)

Skip specific calls

With the default auto_monitor=True, all calls are monitored. Use skip_call() inside your session handler to opt out a specific call:

from voiceeval import Client, skip_call

client = Client(
    api_key="your_voiceeval_api_key",
    agent_name="my-booking-agent",
)

@server.rtc_session(agent_name="my-agent")
async def entrypoint(ctx: JobContext):
    # Decide based on room metadata, participant info, etc.
    if ctx.room.name.startswith("internal-"):
        skip_call()  # This call won't be monitored or evaluated

    session = AgentSession(stt=..., llm=..., tts=...)
    await session.start(agent=MyAgent(), room=ctx.room)
    await ctx.connect()

Monitor only specific calls

Set auto_monitor=False so no calls are monitored by default, then use monitor_call() to opt in:

from voiceeval import Client, monitor_call

client = Client(
    api_key="your_voiceeval_api_key",
    agent_name="my-booking-agent",
    auto_monitor=False,
)

@server.rtc_session(agent_name="my-agent")
async def entrypoint(ctx: JobContext):
    # Only monitor production calls, not test rooms
    if not ctx.room.name.startswith("test-"):
        monitor_call()  # This call will be traced and evaluated

    session = AgentSession(stt=..., llm=..., tts=...)
    await session.start(agent=MyAgent(), room=ctx.room)
    await ctx.connect()

When a call is skipped (or not opted in), spans still flow to Langfuse for the dashboard but won't create backend records or trigger evaluations.

Manual Tracing (Optional)

For non-LLM functions like business logic or RAG pipelines, use the @observe decorator:

from voiceeval import observe

@observe(name_override="rag_retrieval")
def retrieve_documents(query: str):
    # Your logic here
    return docs

License

MIT

Metadata

Release files for voiceeval-sdk 0.1.9

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for voiceeval-sdk 0.1.9
File Size Uploaded
voiceeval_sdk-0.1.9.tar.gz 67.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for voiceeval-sdk 0.1.9
File Interpreter ABI Platform
voiceeval_sdk-0.1.9-py3-none-any.whl Python 3 none any Details

Total release size: 85.3 kB

Release files / voiceeval_sdk-0.1.9.tar.gz

Download URL voiceeval_sdk-0.1.9.tar.gz
Size 67.5 kB
Tags Source
SHA-256 checksum
How to use checksums
4f5ada785fbc87cfa117cbfaf5896dffd0be38b49287bbc72304418db56a3736
BLAKE2b-256 checksum
How to use checksums
fae431cb93b77eba29ea93042c11deeae0816c2cb7fb7944d1a833911302bea9
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.3

Release files / voiceeval_sdk-0.1.9-py3-none-any.whl

Download URL voiceeval_sdk-0.1.9-py3-none-any.whl
Size 17.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
4ebd146169c2b1e46fc13879502aeba271d1f8b046575962f67568621e606e0d
BLAKE2b-256 checksum
How to use checksums
c911951e9351250d9798125890de125aedb2b2061640be0a88ca5a03da890a3f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.3

Release history Release notifications | RSS feed

This release

0.1.9 This release

2 release files

0.1.8

2 release files

0.1.7

2 release files

0.1.6

2 release files

0.1.5

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page