Skip to main content

voice-evals: evaluation tools for the whole voice conversation. Replay scoring, live probes and CI gates.

voice-evals

Reproducible evaluation for voice agents. Score recorded calls, run scripted conversations over WebSocket, and turn transcription, latency, interruptions and task outcomes into CI gates.

PyPI Python Tests License

Watch the demo · Quick start · Kaggle notebooks · Live probe · Documentation

A voice agent can produce the right words and still respond too slowly, talk over a caller, or miss a correction. voice-evals makes those failures inspectable: a scripted caller talks to your agent, interrupts it, records the session, and exports a dataset that the offline evaluator can score again.

Start without API keys. Replay evaluation and the mock probe run offline. Real calls need your agent's endpoint and a caller TTS provider.

Watch the demo

Animated 22-second voice-evals demo showing replay scoring, CI gates and recording artifacts.

The demo loops automatically without sound. The scores in it come from synthetic calls. Transcript · Download MP4 with sound · Media credits.

Play the full-quality video with sound

https://github.com/user-attachments/assets/3b67eb82-26ce-4247-b652-b0f5b7513b20

Press play and enable sound for the music.

Choose your starting point

You have… Start here What you get
Recorded transcripts and timings voice-eval run calls.jsonl Metrics, per-call details and configurable gates
A live WebSocket voice agent voice-eval probe scenario.json A scripted call, interruption observations and replayable artifacts
No agent or credentials yet Offline quick start or Kaggle A complete offline run you can reproduce yourself

Quick start

Python 3.10+. Clone the repository to get the example datasets, then install the package:

git clone https://github.com/Gjusev/voice-evals.git
cd voice-evals
python -m pip install voice-evals
voice-eval run evals/data/demo_calls.jsonl

If you already have a dataset, installing from PyPI is enough: voice-eval run calls.jsonl.

See the three-call demo output
samples=3 failures=0
wer mean=0.0303 max=0.0909
task_completion=0.6667
fact_coverage=0.8333
hallucination_rate=0.3333
e2e_ms p50=950.0 p95=1103.0 p99=1116.6
stage means: e2e_ms=963.3 llm_ttft_ms=330.0 stt_ms=200.0 tts_ttfa_ms=236.7
interruptions=1 median_barge_in_stop_ms=210.0
gate: PASSED

These are synthetic example calls, not provider benchmarks. With no thresholds configured, a passing gate does not imply production readiness. failures=0 is not a count of successful tasks; task completion is reported separately.

Make quality a CI gate

voice-eval run calls.jsonl --max-wer 0.05 --min-task-completion 0.90 --max-e2e-p95-ms 1000 --json --output result.json

Pick thresholds that fit your use case. Replay exits with 0 when the gates pass, 1 when one fails, and 2 on a dataset error. See the repository's CI workflow for executable examples.

Reproduce in Kaggle

Notebook / kernel What it exercises Links
Offline benchmark Pinned wheel bundle, three-call demo, 123-call regression corpus, mock and fixture probes Open on Kaggle · Source
Live probe Scripted WSS session with your credentials; a clearly labeled mock run when secrets are absent Open on Kaggle · Source

Both notebooks are published and verified. The offline kernel passed its full check suite with internet disabled on 2026-10-04, and the live kernel completed its no-secrets mock path. If a notebook is unavailable to you, its local source and the Kaggle reproduction guide describe the same steps. The offline bundle supplies pinned wheels and checksums.

The offline kernel runs with internet disabled. The live kernel needs an externally reachable WSS endpoint for a live call; localhost on your machine is not reachable from Kaggle. Its latency numbers include the Kaggle datacenter's network path, and its mock timings are simulated.

What it measures

Layer Measurement Interpretation
Transcription Per-call WER, mean and maximum Compares ASR text with a supplied reference; mean WER weights calls equally
Latency E2E p50/p95/p99; available STT, LLM TTFT and TTS TTFA means Stage timings require explicit observations; missing stages are not inferred
Interruptions Barge-in observations and measured stop time A stop that is never confirmed is reported as censored
Task outcome Completion, required-fact coverage, forbidden-phrase rate Deterministic checks against your scenario, not an LLM judge

The output field hallucination_rate measures calls containing a configured forbidden phrase. It does not detect every possible hallucination. Read the metric definitions.

How it works

Voice evaluation workflow: a scripted caller connects to your voice agent through a live probe. Recording artifacts or existing recordings supply calls.jsonl to the replay evaluator, which produces metrics and CI gates.

View full-size diagram · Mermaid source

One evaluator, two entry points. A scored probe exports the same replay format used for offline calls. Replay reproduces the legacy scores from that recording; a new live call can vary with the agent, provider and network.

Live probe

From the cloned repository, try the four-turn appointment scenario with a conditional branch and deliberate interruption:

python -m pip install "voice-evals[probe]"
voice-eval probe src/voice_evals/resources/scenarios/appointment-v2.json --mock --output-dir out/probe-demo --max-wer 0.05 --max-barge-in-stop-ms 500
voice-eval run out/probe-demo/calls.jsonl

The mock demo takes about 25 seconds and needs no credentials or network. Its timing is simulated.

For a real call, configure ELEVENLABS_API_KEY, ELEVENLABS_VOICE_ID and PROBE_TRANSPORT_URL, plus PROBE_AGENT_API_KEY if your endpoint requires authentication:

voice-eval probe src/voice_evals/resources/scenarios/appointment-v2.json --caller elevenlabs --output-dir out/probe-live --json

Your endpoint must match a protocol map. The current transport supports mono s16le PCM at 16/24 kHz; arbitrary telephony or binary protocols need an adapter. Start with the local reference agent to exercise real sockets.

Caller and integration options

Python   Kaggle   PyPI   ElevenLabs   GitHub

Caller Use it for Setup
Mock Offline harness checks --mock
Audio fixtures Prepared caller recordings --caller fixture --fixture-manifest …
ElevenLabs Hosted caller TTS --caller elevenlabs and provider credentials
HTTP TTS An externally hosted speech model --caller http, CALLER_TTS_URL and a compatible service

Open-model examples include Qwen3-TTS, Kokoro and Chatterbox caller profiles, plus a Voxtral STT gateway example. Model services run separately. Logos identify technologies and integrations, not endorsements.

Inspect the evidence

out/probe-demo/
  manifest.json       sanitized config, provenance and file hashes
  events.jsonl        normalized event journal
  calls.jsonl         replay dataset for eligible sessions
  result.json         scores, probe observations and gate result
  audio/caller/       sent audio, per utterance
  audio/agent/        received audio, per response

Ineligible sessions report NOT SCORED with reasons. Manifests record credential environment-variable names, not secret values. Missing stage observations stay unavailable; a timed-out interruption is never assigned an invented stop duration.

Recorded live validation: one ElevenLabs TTS → real WSS → reference-agent Scribe STT session completed with WER 0.1379 and barge-in stop 232.6 ms. It failed its WER gate of 0.10. The reference dialogue policy was scripted, and this single run is transport evidence rather than a model ranking. Method, latency aggregation and reproduction tests.

Documentation

Resource Contents
Live probe guide Real calls, scenario branches, transport maps, artifacts and scoring eligibility
Dataset format and Python API JSONL schema, metric definitions and programmatic examples
Kaggle reproduction Secrets, offline bundles, runtime behavior and publishing commands
Regression corpus 123 frozen calls, hand-checked WER cases and dataset limits
Reference agent Local WebSocket endpoint with scripted policy and STT modes
Open-model examples HTTP caller services, model profiles and Voxtral gateway

Development

uv sync --extra probe
uv run pytest -q -m "not live"
uv run ruff check src tests scripts kaggle-kernel examples
uv build

The Makefile also provides make demo-probe, make test-integration and make kaggle-bundle. Offline tests never call real services. Live integration tests require VOICE_EVALS_LIVE_TESTS=1 and relevant credentials, and are blocked under CI.

Found an integration gap? Open an issue with the protocol, a minimal sanitized example, and the expected behavior.

Roadmap

  • Twilio/Telnyx media-stream adapters.
  • A dedicated OpenAI Realtime adapter.
  • An optional LLM judge for graded outcomes.
  • A GitHub Action for comparison against a committed baseline.

These are planned capabilities, not current integrations.

License

Apache 2.0. Built by Gjusev.

Metadata

Release files for voice-evals 0.2.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for voice-evals 0.2.1
File Size Uploaded
voice_evals-0.2.1.tar.gz 9.0 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for voice-evals 0.2.1
File Interpreter ABI Platform
voice_evals-0.2.1-py3-none-any.whl Python 3 none any Details

Total release size: 9.1 MB

Release files / voice_evals-0.2.1.tar.gz

Download URL voice_evals-0.2.1.tar.gz
Size 9.0 MB
Tags Source
SHA-256 checksum
How to use checksums
075898ad0087281fa61daa90d362883d13cef330cb6110b5698c42b1d612f7ed
BLAKE2b-256 checksum
How to use checksums
d2ad602c5126d208d63e7556d4acfa242a667a20352ca4f7021d71b3cb9ce49f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.23 {"installer":{"name":"uv","version":"0.12.23","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release files / voice_evals-0.2.1-py3-none-any.whl

Download URL voice_evals-0.2.1-py3-none-any.whl
Size 82.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
1ac6cf1b4c94878c99d2b28c945f5b4298a917ca4741d02b789f45c547105e51
BLAKE2b-256 checksum
How to use checksums
149b95f62eb699a872fa3ba8cc154978bcf497eb7087987e643dd44ee890d1f4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.23 {"installer":{"name":"uv","version":"0.12.23","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release history Release notifications | RSS feed

This release

0.2.1 This release

2 release files

0.2.0

2 release files

0.1.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page