voice-evals
Evaluation harness for voice agents: replay scoring plus a scripted live probe — WER, latency budgets, barge-in behavior and task outcomes, with CI gates.
Text-agent evals are everywhere. Voice adds four layers that nobody has open-sourced well: transcription quality under accents and noise, per-stage latency (voice has a hard "feels instant" budget around 800ms), interruption handling, and whether the call actually achieved its goal. voice-evals measures them from recorded calls, offline, with zero credentials — and since v0.2 it can also make the calls: a deterministic scripted caller speaks to your live agent over its real WebSocket transport, barges in, records everything, and scores it with the same replay evaluator.
Built by someone who runs a production voice agent (HeizPro KI, real-time STT/LLM/TTS), not from a spec sheet.
What it measures
| Layer | Metrics |
|---|---|
| Transcription | per-call WER (live STT output vs ground truth), mean and max |
| Latency | end-to-end p50/p95/p99 plus per-stage means: STT, LLM time-to-first-token, TTS time-to-first-audio |
| Behavior | interruption count, median time until agent audio stops after a barge-in |
| Outcome | task completion vs expected outcome, required-fact coverage, hallucination rate (forbidden claims) |
Quick start (replay, no credentials)
pip install voice-evals
voice-eval run evals/data/demo_calls.jsonl
samples=3 failures=0
wer mean=0.0303 max=0.0909
task_completion=0.6667
fact_coverage=0.8333
hallucination_rate=0.3333
e2e_ms p50=950.0 p95=1103.0 p99=1118.6
stage means: e2e_ms=963.3 llm_ttft_ms=330.0 stt_ms=200.0 tts_ttfa_ms=236.7
interruptions=1 median_barge_in_stop_ms=210.0
Gate in CI:
voice-eval run calls.jsonl --max-wer 0.05 --min-task-completion 0.90 --max-e2e-p95-ms 1000
Exit codes: 0 passed, 1 gate failed, 2 dataset error.
v0.2: live probe
pip install "voice-evals[probe]" # adds httpx + websockets
The harness becomes a synthetic caller: it TTS-generates caller lines from a
scenario script (v2 JSON, bundled example:
src/voice_evals/resources/scenarios/appointment-v2.json),
calls your agent over its real transport, records the session, and scores it
with the same evaluator. Try the fully offline mock demo (~25s, no
credentials, no network):
make demo-probe
# or: voice-eval probe src/voice_evals/resources/scenarios/appointment-v2.json \
# --mock --output-dir out/probe-demo --max-wer 0.05 --max-barge-in-stop-ms 500
A real call against an agent speaking the bundled reference protocol:
export ELEVENLABS_API_KEY=... # caller TTS
export ELEVENLABS_VOICE_ID=... # an API key alone does not identify a voice
export PROBE_TRANSPORT_URL=wss://your-agent.example.com/voice
export PROBE_AGENT_API_KEY=... # agent auth, never the ElevenLabs key
voice-eval probe src/voice_evals/resources/scenarios/appointment-v2.json \
--caller elevenlabs --voice-id "$ELEVENLABS_VOICE_ID" \
--output-dir out/probe-live --json
Open-model caller instead of ElevenLabs (the model service runs elsewhere; see
examples/open-models/):
export CALLER_TTS_URL=http://your-tts-service:8080/tts
voice-eval probe scenario.json --caller http \
--caller-model Qwen/Qwen3-TTS-12Hz-0.6B-CustomVoice \
--transport "$PROBE_TRANSPORT_URL" --output-dir out/probe-open
Prepared real speech instead of TTS:
--caller fixture --fixture-manifest evals/fixtures/audio/synthetic/manifest.json.
Local reference agent (real sockets, no third-party agent needed)
examples/reference-agent/ is a
localhost WebSocket agent speaking exactly default-v1 — scripted policy,
tone TTS, and three STT modes (none = honest NOT SCORED, watermark = pairs
with --mock, scribe = real ElevenLabs Scribe STT). It validates the
probe's real wire path end to end without any external agent service.
Live validation performed (2026-10-04, this repository)
All three live layers were exercised and are committed as opt-in tests
(VOICE_EVALS_LIVE_TESTS=1 + secrets, never in CI):
| Layer | What ran | Result |
|---|---|---|
| Caller TTS | ElevenLabsCallerVoice against the real API: native pcm_16000/pcm_24000, MP3-body rejection, 401-not-retried |
4/4 passed (tests/integration/test_live_elevenlabs.py) |
| Real WebSocket E2E | Full SessionRunner over a real socket vs the reference agent (watermark mode) |
scored, replay-equal, barge-in observed (tests/integration/test_live_probe_local_agent.py) |
| Full real probe | CLI voice-eval probe with the real ElevenLabs caller (eleven_multilingual_v2, stock voice Sarah) → real WSS transport → reference agent with real Scribe STT |
completed + scored: WER 0.1379 (real TTS→STT round trip: Scribe heard "hour" for "instead", dropped punctuation), task 1.0, E2E turn p50 1086ms/p95 2034ms, barge-in stop 232.6ms; exported calls.jsonl replays to identical scores; no secret in any artifact |
That WER is a genuine measurement of the ElevenLabs-TTS → Scribe-STT round
trip through the probe's real wire path; it says nothing about any agent's
intelligence (the reference agent's dialogue policy is scripted). The
originally supplied private voice id was not present in the account
(voice_not_found); the documented stock premade voice Sarah
(EXAVITQu4vr4xnSDxMaL) was used instead.
What the probe records
Every run writes a complete, replayable recording directory:
out/probe/manifest.json sanitized manifest: config, chosen alternatives, timings,
interruption observations, provenance, file hashes
out/probe/events.jsonl append-only normalized event journal
out/probe/calls.jsonl exported v0.1 replay record (empty+marked when not replayable)
out/probe/result.json v0.1 result fields + additive "probe" object + gate_passed
out/probe/audio/caller/*.wav exactly the bytes sent, per utterance
out/probe/audio/agent/*.wav exactly the bytes received, per response
A live session and its exported calls.jsonl replay to identical legacy
scores. Exit codes: 0 observed+scored and gates passed, 1 measured or
behavioral failure (including NOT SCORED runs with explicit reasons), 2
config/transport/provider/recording error.
Honesty rules baked into the output
- Client-observed events cannot reveal hidden STT/LLM/TTS stages. Legacy
llm_ttft_ms/tts_ttfa_msare populated only when explicit stage events exist; otherwise proxies are reported under separate names and stage fields staynullwith a recorded reason. - STT latency is labeled
client_final_asr(includes endpointing + network). - A barge-in that never confirms a stop is right-censored (
not_stopped), with a lower bound — never a fabricated duration. - Ineligible sessions print NOT SCORED with exclusion reasons; placeholder zero-fields are conventions, not measurements.
- Secrets are resolved at execution time and never serialized: manifests record env variable names, never values or signed URLs.
- Mock latency is simulated and is never presented as a provider benchmark.
Scenario v2 in one minute
Four caller turns with a conditional branch and a deliberate interruption
(full schema: src/voice_evals/resources/schemas/scenario-v2.schema.json):
{"id": "correct_time", "intent": "correct_appointment_time",
"cue": {"mode": "interrupt", "response_to": "provide_name",
"after_ms": 600, "timeout_ms": 15000, "if_missed": "fail"},
"utterance": {"text": "Sorry to interrupt. I need {{corrected_time}}, not ten.",
"alternatives": ["Sorry, could we make that {{corrected_time}} instead?"],
"reveals": ["corrected_time"]}}
- Alternatives are chosen by a stable hash of
seed+step id — deterministic, independent of branch execution order. {{placeholders}}must be declared facts, and every substituted fact must be listed inreveals(mismatch is a validation error).- Branches (
when/on_unmatched) are bounded and deterministic: literal substring predicates over one completed response. No LLM branching in v0.2. outcome_rule.source: final_agent_textis an explicit proxy for what the agent reported — not evidence a calendar write succeeded.
Transport: one protocol map, no magic
A WebSocket URL does not specify an audio protocol. The probe speaks any
endpoint covered by a declarative JSON map (bundled reference:
default-v1.json):
PCM formats per direction, handshake, outbound envelopes with a fixed
substitution allowlist, an inbound discriminator with JSON-pointer selectors,
declared capabilities, and env-referenced auth. Write your own map against
protocol-v1.schema.json
— no Python, Jinja, or scripts are ever evaluated from a map. Native PCM only
(mono s16le 16/24 kHz): no hidden codecs or resampling; unsupported formats
fail preflight. "Works with any agent" means any agent whose protocol fits the
map; a telephony webhook or arbitrary binary protocol needs a future adapter.
Dataset format (replay)
JSONL, one recorded call per line (JSON arrays also work):
{
"id": "call-001",
"scenario": {
"name": "book-appointment",
"expected_outcome": "booked",
"required_facts": ["Tuesday", "appointment"],
"forbidden_facts": ["discount"]
},
"asr_transcript": "what the agent's STT heard",
"ground_truth_transcript": "what the caller actually said",
"agent_transcript": "everything the agent said",
"outcome": "booked",
"stage_timings_ms": {"stt_ms": 190, "llm_ttft_ms": 310, "tts_ttfa_ms": 220, "e2e_ms": 820},
"interruptions": [{"at_ms": 4200, "agent_stopped_ms": 210}]
}
Only id, scenario.name, scenario.expected_outcome and ground_truth_transcript are required. Everything else scores n/a when absent, so you can start with transcripts only and add timings later. v0.2 is fully API- and schema-compatible with v0.1.1; all changes are additive.
A larger frozen regression corpus (123 calls incl. hand-checked WER cases) ships in evals/data/.
Reproduce in Kaggle
Publication pending — the links activate when the kernels and bundle dataset are published with the v0.2.x release.
Kernel A (kaggle-kernel/offline/) proves
harness reproducibility with enable_internet=false: it installs the exact
published wheel and pinned dependency wheels from a checksummed Kaggle dataset
(--no-index --find-links --require-hashes; never PyPI at runtime), then
asserts the unchanged v0.1 demo metrics, the frozen regression corpus with
hand-checked WERs, a full four-turn mock probe (conditional branch, overlap
barge-in, artifacts, live-to-replay equality), the fixture-caller path, and
the offline test suite. Outputs land in /kaggle/working with provenance and
artifact checksums.
Kernel B (kaggle-kernel/probe/) is a live
diagnostic: with Kaggle Secrets ELEVENLABS_API_KEY + PROBE_TRANSPORT_URL
it runs one scripted probe over outbound WSS; with secrets missing it runs the
full mock session through the same runner/recorder/evaluator and prints only
the missing secret names — no network attempt. A failed live attempt keeps its
artifacts and additionally writes a separately named mock diagnostic; the mock
score is never substituted for the live result. The open-model variant uses
CALLER_TTS_URL with caller=http instead of ElevenLabs credentials.
Latency location warning: live Kernel B latency is measured from the Kaggle datacenter and includes network transit/RTT and client scheduling. It is not directly comparable with local runs; mock latency is simulated and is not a provider benchmark.
Secrets setup: notebook sidebar → Add-ons → Secrets → attach
ELEVENLABS_API_KEY, ELEVENLABS_VOICE_ID (a stock voice id), and your
agent's PROBE_TRANSPORT_URL (+ optional PROBE_AGENT_API_KEY). A service on
your localhost is not reachable from Kaggle — expose an authenticated public
WSS endpoint. Your endpoint must match a protocol map; the default map covers
agents implementing the reference protocol.
Publishing (from the repository, host-side):
KAGGLE_API_TOKEN=... python -m kaggle kernels push -p kaggle-kernel/offline
KAGGLE_API_TOKEN=... python -m kaggle kernels status gjusev/voice-evals-offline-benchmark
KAGGLE_API_TOKEN=... python -m kaggle kernels output gjusev/voice-evals-offline-benchmark -p out/kaggle-offline/
KAGGLE_API_TOKEN=... python -m kaggle kernels push -p kaggle-kernel/probe
KAGGLE_API_TOKEN=... python -m kaggle kernels status gjusev/voice-evals-live-probe
KAGGLE_API_TOKEN=... python -m kaggle kernels output gjusev/voice-evals-live-probe -p out/kaggle-probe/
PowerShell: set the token once with $env:KAGGLE_API_TOKEN="...", then run
the same commands without the Bash assignment prefix. Forks must change the
metadata id and the commands' owner. KAGGLE_API_TOKEN is a publisher
credential, not a kernel runtime secret. Bundle build details:
kaggle-kernel/bundle/README.md.
Python API
from voice_evals import evaluate, load_dataset
result = evaluate(load_dataset("calls.jsonl"))
print(result.summary())
print(result.e2e_p95_ms, result.hallucination_rate)
Programmatic probe (same machinery as the CLI):
import asyncio
from pathlib import Path
from voice_evals.probe import (
MockCallerVoice, MockTransport, ProbeConfig, ScenarioScript, SessionRunner,
)
async def main() -> None:
script = ScenarioScript.load("scenario.json")
config = ProbeConfig(environment="mock", output_dir=Path("out/probe"))
runner = SessionRunner(caller=MockCallerVoice(), transport=MockTransport(), config=config)
session = await runner.run(script, output_dir=config.output_dir)
runner.last_report.print_summary()
asyncio.run(main())
The base install (jiwer only) imports the whole probe package, runner, mocks
and fixtures; httpx/websockets are imported only inside their adapters via
the [probe] extra.
Roadmap
- Twilio/Telnyx media-stream adapters behind the same transport interface.
- Dedicated OpenAI Realtime adapter (append/commit/truncate semantics).
- LLM judge for graded outcomes instead of exact outcome matching.
- Regression-gate GitHub Action comparing against a committed baseline.
Development
make install # editable install incl. probe extra + dev tools
make test # offline pytest (default: -m "not live")
make test-integration # opt-in live tests (VOICE_EVALS_LIVE_TESTS=1 + secrets)
make lint # ruff
make build # wheel + sdist
make demo-probe # offline four-turn mock probe demo
make kaggle-bundle # stage the internet-disabled bundle (validation mode)
Unit tests and CI never call real services. Live integration tests require
both VOICE_EVALS_LIVE_TESTS=1 and the relevant secrets, and are additionally
guarded against running under CI.
License
Apache 2.0. See LICENSE.
Metadata
Release files for voice-evals 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| voice_evals-0.2.0.tar.gz | 238.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| voice_evals-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 323.0 kB
Release files / voice_evals-0.2.0.tar.gz
| Download URL | voice_evals-0.2.0.tar.gz |
|---|---|
| Size | 238.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
602f55eac14e4cd8732551fd655e89eea4768a5fcfa01b25053f24d39bd5bd56
|
|
BLAKE2b-256 checksum How to use checksums |
d228028efe9ec3dcf0b64693bb5fac6767f820e13d2d4869172cd62cfb27cda1
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.12.23 {"installer":{"name":"uv","version":"0.12.23","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|
Release files / voice_evals-0.2.0-py3-none-any.whl
| Download URL | voice_evals-0.2.0-py3-none-any.whl |
|---|---|
| Size | 84.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
463c3c844da63f1c387a5387cd5ee8a5956861c54904009bb4910aa9d0ef0723
|
|
BLAKE2b-256 checksum How to use checksums |
b08a2cc978bbe28c429890b1fb990822b71bc831bf581ae9341284079d91a836
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.12.23 {"installer":{"name":"uv","version":"0.12.23","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|