voice-evals
Evaluation harness for voice agents: WER, latency budgets, barge-in behavior and task outcomes, with CI gates.
Text-agent evals are everywhere. Voice adds four layers that nobody has open-sourced well: transcription quality under accents and noise, per-stage latency (voice has a hard "feels instant" budget around 800ms), interruption handling, and whether the call actually achieved its goal. voice-evals measures them from recorded calls, offline, with zero credentials.
Built by someone who runs a production voice agent (HeizPro KI, real-time STT/LLM/TTS), not from a spec sheet.
What it measures
| Layer | Metrics |
|---|---|
| Transcription | per-call WER (live STT output vs ground truth), mean and max |
| Latency | end-to-end p50/p95/p99 plus per-stage means: STT, LLM time-to-first-token, TTS time-to-first-audio |
| Behavior | interruption count, median time until agent audio stops after a barge-in |
| Outcome | task completion vs expected outcome, required-fact coverage, hallucination rate (forbidden claims) |
Quick start
pip install voice-evals
Score the bundled demo dataset, no credentials needed:
voice-eval run evals/data/demo_calls.jsonl
samples=3 failures=0
wer mean=0.0303 max=0.0909
task_completion=0.6667
fact_coverage=0.8333
hallucination_rate=0.3333
e2e_ms p50=950.0 p95=1103.0 p99=1118.6
stage means: e2e_ms=963.3 llm_ttft_ms=330.0 stt_ms=200.0 tts_ttfa_ms=236.7
interruptions=1 median_barge_in_stop_ms=210.0
Gate in CI:
voice-eval run calls.jsonl --max-wer 0.05 --min-task-completion 0.90 --max-e2e-p95-ms 1000
Exit codes: 0 passed, 1 gate failed, 2 dataset error.
Dataset format
JSONL, one recorded call per line (JSON arrays also work):
{
"id": "call-001",
"scenario": {
"name": "book-appointment",
"expected_outcome": "booked",
"required_facts": ["Tuesday", "appointment"],
"forbidden_facts": ["discount"]
},
"asr_transcript": "what the agent's STT heard",
"ground_truth_transcript": "what the caller actually said",
"agent_transcript": "everything the agent said",
"outcome": "booked",
"stage_timings_ms": {"stt_ms": 190, "llm_ttft_ms": 310, "tts_ttfa_ms": 220, "e2e_ms": 820},
"interruptions": [{"at_ms": 4200, "agent_stopped_ms": 210}]
}
Only id, scenario.name, scenario.expected_outcome and ground_truth_transcript are required. Everything else scores n/a when absent, so you can start with transcripts only and add timings later.
Python API
from voice_evals import evaluate, load_dataset
result = evaluate(load_dataset("calls.jsonl"))
print(result.summary())
print(result.e2e_p95_ms, result.hallucination_rate)
Roadmap
- v0.2 live probe: the harness becomes a synthetic caller. It TTS-generates caller lines from a scenario script, calls your agent over its real transport (WebSocket first, Twilio/Telnyx next), records the session and scores it live, including deliberate barge-ins.
- LLM judge for graded outcomes instead of exact outcome matching.
- Regression-gate GitHub Action comparing against a committed baseline.
Development
make install # editable install with dev extras
make test # pytest, no network
make lint # ruff
make build # wheel + sdist
License
Apache 2.0. See LICENSE.
Metadata
Release files for voice-evals 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| voice_evals-0.1.1.tar.gz | 12.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| voice_evals-0.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 25.0 kB
Release files / voice_evals-0.1.1.tar.gz
| Download URL | voice_evals-0.1.1.tar.gz |
|---|---|
| Size | 12.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
2d50dcfc62516274d83b58ea1cc8b22e359ae6c3de546f5aa54aee51e19de495
|
|
BLAKE2b-256 checksum How to use checksums |
5734d6881441096c56ae0c52d723a2a42576742ccecce0e830399bc039dc5809
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.12.23 {"installer":{"name":"uv","version":"0.12.23","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|
Release files / voice_evals-0.1.1-py3-none-any.whl
| Download URL | voice_evals-0.1.1-py3-none-any.whl |
|---|---|
| Size | 12.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
75e88b6080877dabcb896990b26e411ee7191cd4694ebf3d850f3480b7ca6b63
|
|
BLAKE2b-256 checksum How to use checksums |
f506c974ebf02f44b85f50f9212cc1ddc2ab50be55eb091d6c086ea3cf3a2d7a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.12.23 {"installer":{"name":"uv","version":"0.12.23","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|