Open-source, self-hosted conversation QA for voice agents: simulate, evaluate, review, and track calls across five dimensions with the evidence behind every result. MIT.
Project description
hotato
Open-source, self-hosted conversation QA for voice agents.
Simulate, evaluate, review, and track every call across five dimensions — outcome, policy, conversation, speech, reliability — with the evidence behind every result. Offline and self-hosted; deterministic checks stay separate from the model-judged rubric; never one blended score.
Offline by default · MIT · zero runtime dependencies.
Quickstart -- no account, no keys, no network
pip install hotato && hotato start --demo
Already have uv? Zero-install, same command:
uvx hotato start --demo
It sweeps two bundled recorded calls a provider's default agent failed, writes the candidate dashboard, and turns one missed-interruption candidate into a demo failure contract it verifies on the spot (output abridged):
[start] demo: swept 2 bundled calls, 5 candidate moments; wrote hotato-sweep.json, hotato-sweep.html, hotato-no-single-threshold.svg, contracts/demo-missed-interruption.hotato/contract.json
hotato start: swept the 2 bundled demo calls offline.
sweep dashboard: hotato-sweep.html
demo contract: contracts/demo-missed-interruption.hotato
verified contract: FAIL as expected -- the demo call missed the interruption
[ ... then the exact next commands: promote a candidate, gate it in CI, re-check it ... ]
Open hotato-sweep.html -- the ranked candidate moments, each with a hear-the-bug playhead:
hotato-sweep.html · candidate moments ranked by salience, each with a playhead that sweeps the timeline in sync with the embedded audio. Candidates you review and label, never a decided verdict.
start --demo promoted one of these candidates into a portable .hotato contract and ran contract verify on it (FAIL, as expected). It then prints the exact next commands to save the candidate as a permanent fixture, gate it in CI, and re-check it.
Point it at your own recording to walk the same loop end to end:
hotato start --stereo my-call.wav # trust -> scan -> review -> label -> contract
Every command takes a two-channel recording (caller on one channel, agent on the other). A mono file or a bad export is marked NOT SCORABLE, never turned into a confident but meaningless verdict.
How the loop works
Catch a moment, confirm what it should have done, gate it in CI, then re-check today's agent.
1. Catch -- surface the candidate moments
sweep ranks the talk-over and false-stop moments across your recent calls by how far the timing missed. Two candidate types, straight from the open scorer, no accuracy score:
| Level 1 -- candidate. Hotato reports what it measured (0.32s of overlap; 0.46s of trailing silence), never a guess at intent. Each is a timing moment worth review, not a bug yet. | |
2. Confirm -- you label yield or hold
You decide the expected behavior for a candidate: yield (stop for the caller) or hold (keep talking through a backchannel). No single sensitivity dial decides it for you -- a threshold that stops missing interruptions starts false-stopping on backchannels, so when both fail in one run diagnose refuses to name one threshold:
Level 2 -- human-labeled failure. A reviewer confirms the recording broke an explicit yield-or-hold expectation. When a missed interruption and a false stop collide, the fix is the engagement-control class, not one dial.
3. Gate -- pin it to a CI contract
fixture promote saves your labeled call as a permanent regression test. On every push, CI re-scores that recording under the pinned thresholds and scorer, and exits non-zero if the stored evidence stops matching your policy:
Level 3 -- stored-evidence check. One recording, the pinned scorer, a FAIL against the labeled yield expectation. The same audio and config reproduce every number in this report.
4. Prove -- re-check the current agent
The frozen recording catches the evidence, thresholds, or scorer drifting. To check that the CURRENT agent still behaves, recapture the same scenario and score a NEW contract under the same policy:
# place the same call against today's agent, capture dual-channel, then:
hotato contract create --stereo fresh-call.wav --onset 41.90 --expect yield \
--id refund-cutoff-001-recapture --out contracts
hotato contract verify contracts/refund-cutoff-001-recapture.hotato
Level 4 -- fresh-recapture comparison. A newly captured call meets the same labelled policy and no submitted paired guard regressed. This is the claim the frozen-recording gate cannot make. Walkthrough: docs/RECAPTURE.md.
Five levels of evidence
Hotato never calls the weakest level a verdict. Every card, report, and CLI result names its level; the public tier is the weakest one its inputs support, never a blended score.
| Level | Name | What it means |
|---|---|---|
| 1 | Candidate | A timing moment worth human review. Not a bug yet. |
| 2 | Human-labeled failure | A reviewer confirmed this recording broke an explicit yield-or-hold expectation. |
| 3 | Stored-evidence check | The historical audio still produces the expected result under the pinned policy and scorer. |
| 4 | Fresh-recapture comparison | A newly captured call passed the same contract, and no submitted paired guard regressed. |
| 5 | External proof | An independent team confirmed a caught regression or a fresh recapture. Not yet published. |
A before/after experiment (hotato fix trial, and the fleet loop) re-derives every verdict from the on-disk audio under one pinned trial manifest. It refuses a proof built from an edited verdict, a re-encoded old call, a dropped fixture, or unrelated audio. The number comes from re-scoring the recordings every time, never from a stored field.
Conversation QA -- the five-dimension scorecard
A voice agent can pass every text assertion and still lose the call: the refund never fires, the recording disclosure gets skipped, the caller gets talked over. hotato test run grades one call against a conversation-test file across five dimensions, kept apart and never summed into one number:
- Outcome -- did the job get done, graded on tool-call and state evidence, not the transcript's say-so.
- Policy -- required disclosures, PII handling, and the compliance phrases your team owns.
- Conversation -- the deterministic turn-taking core: did the agent yield when the caller took the floor, and how fast.
- Speech -- response latency and the timing around each turn.
- Reliability -- pass@1 / pass@k / pass^k across repeated runs with a Wilson interval, so a flaky check reads as flaky.
Two lanes stay structurally separate. Deterministic checks (phrase, PII, policy, tool-call, sequence, latency, outcome -- pure regex, checksum, and span-lookup) live behind a wall from the model-judged rubric lane; a rubric verdict is deterministic: false, advisory, and never merged into a deterministic count. No output carries an overall_score, including --format json.
hotato scenario init refund-check --out conversation-test.yaml # a starter you edit, not a claim about your call
hotato test run conversation-test.yaml --agent support-bot
success: FAIL (required: all_deterministic_assertions_pass, no_rubric_failure)
per-dimension (grouped view; never blended):
outcome 0 pass / 0 fail / 1 inconclusive
policy 0 pass / 0 fail / 1 inconclusive
conversation 0 pass / 0 fail / 1 inconclusive
speech 0 pass / 0 fail / 1 inconclusive
reliability 0 pass / 0 fail / 0 inconclusive
Supply the call as --transcript, --trace, --state, and/or --audio and each check turns to pass or fail; a check whose evidence is absent stays INCONCLUSIVE, never guessed. Full walkthrough: docs/CONVERSATION-TEST.md.
Then scale the one call into a release gate:
hotato suite run suite.yaml --agent support-bot-- run a wholesuite.v1; scenario-driven tests execute offline through the deterministic scripted-caller simulator, and every run records into the local registry.hotato simulate --matrix scenario.yaml --out ./conv-- render ascenario.v1with a deterministic scripted caller intoorigin=simulatedconversation artifacts; the variation matrix expands into hundreds of runs, seeded and byte-identical on replay.hotato rubric run --rubrics rubrics.yaml --transcript call.json-- the model-judged lane on a pinned local model; zero egress, advisory unless--gate.docs/RUBRIC.md.hotato release compare BASELINE CANDIDATE-- diff two recorded releases per dimension and per scenario, digest-exact, surfacing new failures and fixed-since.hotato serve-- a self-hosted, read-only web app over the local registry (release readiness, scenario matrix, conversation inspector, failure clusters, production health) on127.0.0.1, bearer-token authenticated.docs/WORKSPACE.md.
Connect a production stack
The demo needs nothing. To point Hotato at real calls, connect once, then sweep on a schedule:
hotato connect vapi # credentials stored 0600, local only
hotato sweep --stack vapi --since 7d --out hotato-sweep.html # cron, CI, wherever
Run sweep on a timer and it becomes a scheduled batch scanner. Your audio stays on your machine unless you explicitly pull it from your stack. Full guide: docs/SET-AND-FORGET.md · runnable examples/set-and-forget/.
Get notified
sweep and hotato fleet run can POST a one-line JSON summary to a webhook when they finish -- off by default, opt in with --notify (repeatable for more than one URL):
hotato sweep --stack vapi --since 7d --notify https://hooks.slack.com/services/...
The payload carries counts, the top candidate moments (id, kind, timing numbers only), and local artifact paths -- never audio, a credential, or transcript text -- plus a text field a Slack incoming webhook renders directly, no template work needed. A down or slow webhook never breaks the run: a delivery failure is one warning line on stderr. Egress details: docs/EGRESS.md.
Fleet -- private, self-hosted
hotato fleet runs the loop across every agent from one local workspace: ingest calls, surface candidates, label them, and run a before/after experiment that recomputes both sides from audio under a pinned manifest. It recommends a change; it never deploys one. Local mode is stdlib-only (SQLite plus a content-addressed store) with no account and no hosted dependency, and no product limit on how many agents you register.
hotato fleet init -w acme
hotato fleet agent add -w acme --name support-bot --stack vapi --assistant-id asst_123
hotato fleet ingest -w acme --agent support-bot call.wav
hotato fleet discover -w acme --agent support-bot call.wav
hotato fleet review -w acme
A before/after experiment refuses a proof built from an edited verdict, a re-encoded old call, a dropped fixture, or unrelated audio: the number comes from re-scoring the recordings, under one pinned policy, every time. Full guide: docs/GUARDIAN-FLEET.md.
Once a workspace has some history, hotato fleet trend -w acme reads the same local SQLite registry and writes one self-contained HTML page: per-agent talk-over and time-to-yield trend lines (p50/p95 per day), candidate moments discovered over time, and experiment outcomes (improved/inconclusive/refused) -- offline, no external assets, hand-rendered inline SVG in the same house style as the sweep dashboard. A day with no measurements gets no point, and a series with fewer than two days of history is reported plainly as "not enough history to trend" instead of a faked or interpolated line.
Run it in your own cloud / VPC
The whole team workspace ships as a container. One command stands up the read-only,
token-authenticated workspace (hotato serve) on host loopback over a private
volume; the default stack makes no external call at run time.
docker compose up -d # workspace on 127.0.0.1:8321
docker compose run --rm hotato-init # optional: seed example data
docker compose --profile judge up -d # optional: a local Ollama model judge
Full walk-through -- build, connect your own calls, the local judge, air-gap,
backup, and the zero-migration promise (self-host and cloud share one set of
schemas): docs/SELF-HOST.md. Prove the no-external-calls
posture on your own machine with deploy/verify-zero-egress.sh.
Choose your path
| You want to | Run this |
|---|---|
| Try the full loop, no credentials | hotato start --demo |
| Sweep the bundled demo calls | hotato sweep --demo |
| Sweep recent calls from a real stack | hotato connect vapi then hotato sweep --stack vapi --since 7d |
| Add Hotato to an existing repo, CI gate included | hotato init starter --stack vapi --out . (docs/STARTER.md) |
| Turn a confirmed failure into a portable contract | hotato contract create --from-candidate hotato-sweep.json#1 --expect yield --id refund-cutoff-001 --out contracts (docs/CONTRACTS.md) |
| Verify contracts in CI | hotato contract verify contracts/ --junit contracts-junit.xml |
| Attach observability traces to a contract | hotato trace attach contracts/refund-cutoff-001.hotato --trace voice_trace.jsonl (docs/TRACE.md) |
| Test a candidate fix, before/after, fail-closed | hotato fix trial patch.json --name staging-x --before before/ --after after/ (docs/FIX-TRIAL.md) |
| Share a finding in a PR or slide | hotato card hotato-sweep.json#1 --out finding.svg |
| Drive it from a coding agent | uvx --from "hotato[mcp]" hotato-mcp (the voice_eval_run scorer + eleven fleet tools; configs in docs/MCP.md) |
contract verify and a promoted fixture in CI are two different guarantees, depending on which recording goes in:
| On the frozen recording (every push) | On a fresh recapture (by hand, see docs/RECAPTURE.md) |
|
|---|---|---|
| Proves | The evidence, policy, and scorer are intact | The CURRENT agent's behavior still matches the label |
| Does not prove | That the deployed agent hasn't changed | -- |
A contract bundle contains call audio. Do not commit a raw customer contract to a public repository; use sanitized fixtures for anything public. See docs/CONTRACTS.md.
Install
Add hotato to a project with pip; or run any command zero-install with uvx:
pip install hotato # core: stdlib-only, zero dependencies
pip install 'hotato[neural]' # optional Silero VAD cross-check
pip install 'hotato[transcribe]' # optional ASR transcript, context only -- never scored
pip install 'hotato[livekit]' # LiveKit live capture
pip install 'hotato[pipecat]' # Pipecat live capture
What Hotato is not
- A conversation-QA system that shows its work. It evaluates calls across five dimensions -- Outcome (grounded in tool/state evidence, never the transcript's claim), Policy, Conversation (the deterministic turn-taking core), Speech, and Reliability (pass^k across repetitions) -- as a per-dimension scorecard, never one blended score. Deterministic checks stay structurally separate from the model-judged rubric lane. See
docs/COMPARE.mdfor how it fits alongside runtime voice layers. - Not transcript scoring. It measures audio timing, not what was said. The opt-in
--transcribeflag attaches an ASR transcript purely as context to read next to the verdict; it never changes the timing score. Seedocs/TRANSCRIBE.md. - Not speaker ID. Channels are anonymous; nothing identifies who a person is.
- Not semantic intent detection. It produces candidate timing evidence. Humans label intent. CI enforces confirmed contracts.
- Not a hand on production config. It never sits in the live audio path and never changes a running agent.
Docs
- Set-and-forget monitoring (connect once, sweep on a schedule, promote confirmed bugs into fixtures):
docs/SET-AND-FORGET.md· runnableexamples/set-and-forget/ - Bad call to CI regression test, step by step:
docs/BAD-CALL-TO-CI.md· runnableexamples/bad-call-to-ci/ - What it measures (the three timing signals, re-derivable by hand):
METHODOLOGY.md· Python APIdocs/API.md - The fix ladder (each failure names a likely fix class; when the evidence maps cleanly to stack config, Hotato names the setting family and direction):
docs/FIX-PLANS.md - Rule out the non-turn-taking bugs first (STT, buffering, verbosity, refusals, wrong-language):
docs/WHY.md - Pull a call from your stack (Vapi, Twilio, Retell, LiveKit, Pipecat):
adapters/README.md· statusdocs/ADAPTER-STATUS.md - CI gates: GitHub Action
docs/CI.md· pytest plugindocs/PYTEST.md - Recorded-call battery: 12 scripted calls against a live voice agent on its provider's default settings, where a missed interruption and a false stop on a backchannel fail in the same run, so
diagnoserefuses to name one threshold:corpus/vapi-defaults/README.md - Failure contracts and traces: turn a labelled candidate into a portable, CI-verified bundle and attach observability evidence:
docs/CONTRACTS.md·docs/TRACE.md·docs/OTEL.md - Deterministic assertions on the transcript/trace (phrase, PII, policy, tool-call, outcome; never a blended score):
docs/ASSERTIONS.md - Checking the CURRENT agent, not just the frozen recording: the recapture walkthrough:
docs/RECAPTURE.md - Egress: a per-command network table derived from the code -- what's local, what reaches your vendor, what optional extras add a hosted call:
docs/EGRESS.md - Layered failure evidence and a before/after fix trial:
hotato explainbreaks a failing result down by layer, andhotato fix trialtests a candidate change before/after, fail-closed:docs/EXPLAIN.md·docs/FIX-TRIAL.md·docs/APPLY.md·docs/FIX-LOOP.md - Evidence: what Hotato validates, the input-condition trust matrix, every card and CLI block reproducible, and where Hotato fits alongside broader voice-agent testing tools:
docs/VALIDATION.md·docs/TRUST-MATRIX.md·docs/GALLERY.md·docs/EVIDENCE-PACK.md·docs/COMPARE.md - For coding agents:
AGENTS.md·llms.txt·llms-full.txt· MCP serverdocs/MCP.md· SecuritySECURITY.md - Contributing: the highest-value PR is a labelled call fixture:
docs/SUBMITTING.md
Why "hotato": good turn-taking is a game of hot potato. Speak, then pass the turn the moment the caller wants it. MIT licensed (LICENSE); the open core stays open.
mcp-name: io.github.attenlabs/hotato
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file hotato-1.3.1.tar.gz.
File metadata
- Download URL: hotato-1.3.1.tar.gz
- Upload date:
- Size: 8.3 MB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.3
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
4c969c38d3cd7c5d7e66e2ec3e08f86fe1eec54eff92144b0742c7f56be3ff12
|
|
| MD5 |
437e45d11eda30d27b9c1f2ac212ae5c
|
|
| BLAKE2b-256 |
55b267437aa9fb5feff3a5ac346dd6fb497270fcabae49a79ec34bb9db444a85
|
File details
Details for the file hotato-1.3.1-py3-none-any.whl.
File metadata
- Download URL: hotato-1.3.1-py3-none-any.whl
- Upload date:
- Size: 3.4 MB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.3
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
4a85cea067414179c5cd8ef3655eddf52cb9facf807b8585243037f9c4fc10ca
|
|
| MD5 |
8883d113e797dba0e132708e6e6aadab
|
|
| BLAKE2b-256 |
73af129665af465214e113a14398b851e5aa20f1bf4f8c9b0f193c64c243ec5b
|