Skip to main content

IndicOrderBench

Your voice agent sounds convincing. Did it actually place the correct order?
An open-source benchmark and CI gate that catches voice ordering agents committing the wrong order,
in English and Hinglish, by checking the order they submitted, not the words they said.

CI License: Apache-2.0 Python 3.11+ No API keys needed for the demo Languages: English, Hinglish

Quick start · Connect your agent · What it measures · How it compares · Methodology · Adapters · Write scenarios


The problem in one call

caller:  Do paneer wrap... actually ek hi karo. Bina pyaaz. Aur ek mango lassi.
agent:   2 Paneer Wrap add kar diya. Paneer Wrap 1 kar diya. Paneer Wrap No onion kar diya.
         1 Mango Lassi add kar diya. Aur kuch?
caller:  Bas itna hi, thanks.
agent:   Aapka order: 1 Paneer Wrap, No onion; 1 Mango Lassi. Order place ho gaya, shukriya!

The agent read back one wrap without onion. Here is what it committed to the order system:

Field Expected Actual Result
Submitted orders 1 1 Pass
Paneer Wrap quantity 1 2 Fail
Paneer Wrap · Onion No onion With onion Fail
Mango Lassi quantity 1 1 Pass

Transcript-based testing and LLM judges would have passed this call. IndicOrderBench fails it, because it compares the committed backend state against the scenario's acceptable end states. The table is real output from iob demo.

IndicOrderBench HTML report: order check table, conversation and tool-call trace for a failed Hinglish correction scenario

Why this exists

Voice ordering agents for restaurants, quick commerce and pharmacies are shipping on LiveKit, Pipecat, Vapi, Retell and custom stacks. Their hardest failures are not accent or latency. They are mid-sentence corrections ("two... actually one"), modifiers ("no onion"), cancellations, and duplicate submissions, especially in code-mixed speech like Hinglish, where the number words and the negations do not look like English.

Existing tools judge the conversation. Nobody was checking the cart. IndicOrderBench is a small, deterministic, reproducible benchmark for exactly that question, with a CI exit code so an agent change that starts placing wrong orders fails the build before it reaches a customer.

Quick start

pip install indicorderbench          # PyPI release coming; until then install from git:
pip install "git+https://github.com/Lingikaushikreddy/indicorderbench.git"

iob demo --out demo-out              # buggy vs fixed reference agent, no network, no keys
iob run starter --agent builtin:correct --tag smoke --trials 3

iob run writes results.json (every transcript, tool call and backend snapshot), junit.xml (for any CI system) and a single-file report.html that works from file://. Open docs/demo/fixed/report.html and docs/demo/buggy/report.html to see both sides of the demo.

Connect your agent

Three ways. Payloads, timeouts and the sandbox backend API are in docs/adapters.md.

In-process Python. Any agent whose logic you can import. The sandbox backend gives you JSON-schema tool definitions to hand to your LLM:

# my_agent.py
from indicorderbench.adapters.protocol import AgentReply, CallerUtterance, SessionInfo
from indicorderbench.backend.state import OrderBackend

class MyAgent:
    def __init__(self, backend: OrderBackend, session: SessionInfo) -> None:
        self.backend = backend                  # add_item, update_line, submit_order, ...
        self.tools = OrderBackend.tool_specs()  # function-calling schemas for your LLM

    async def handle(self, u: CallerUtterance) -> AgentReply:
        text = u.text or transcribe(u.audio_bytes)
        ...  # your LLM loop; run tool calls with self.backend.call(name, args)
        return AgentReply(text="Added two paneer wraps. Anything else?")

def make_agent(backend, session):
    return MyAgent(backend, session)
iob run starter --agent python:my_agent:make_agent --trials 3

HTTP. Your agent runs anywhere it can reach the sandbox backend: it receives each caller turn by HTTP and calls the backend's tools by HTTP. examples/http_agent_shim.py is a complete, tested example:

python examples/http_agent_shim.py --port 8900
iob run starter --agent http:http://127.0.0.1:8900/iob --tag smoke
# remote agent: iob run ... --backend-host 0.0.0.0 --backend-port 8765 --backend-url http://<your host>:8765

CI gate. Fail the build when ordering accuracy drops below a threshold or regresses against the baseline you committed from the last good run:

iob run starter --agent http:http://127.0.0.1:8900/iob --tag smoke \
    --baseline baselines/smoke.json --fail-under 0.9 --max-regression 0.0 --allow-infra 0

Exit codes: 0 ok, 1 infra or usage error, 2 accuracy below threshold or regression. A ready-made GitHub Actions workflow is in examples/ci/github-actions.yml.

What it measures

Outcome Meaning In the pass rate?
pass committed state matches an acceptable end state yes
fail caller was valid, infrastructure worked, order is wrong or never placed yes
simulator_invalid the agent asked something the script could not answer; the agent is not blamed no, reported separately
infra_error adapter raised, timed out, or a clip was missing no, reported separately
  • Pass rate and pass^k (τ-bench's repeated-run metric) by language and by failure category, with Wilson 95% intervals. --trials 3 tells you whether the agent is right three times in a row, not just once.
  • Five failure categories: quantity, modifier, mid-order correction, cancellation, duplicate submission.
  • Per-turn latency (p50 / p95) and a timestamped tool-call trace, including refused calls.
  • Reproducibility fields in every results.json: benchmark version, pack content hash, agent spec, trials, modality, timeouts, host and Python version. iob compare refuses to compare runs from different packs.

Validity rules, metric definitions and how to publish numbers responsibly: docs/methodology.md.

How it compares

IndicOrderBench VoiceTest EVA-Bench VoiceAgentBench τ-bench
Scores the committed order state yes no (LLM judge on transcript) partly (composite metrics) no (tool-call correctness on static queries) yes (text only)
Live, multi-turn agent under test yes yes yes no yes
Indian languages English + Hinglish, more planned no no 7 Indic languages no
Simulator-invalid as a separate outcome yes no yes n/a no
Deterministic pass/fail, no judge model yes no no yes yes
Runs offline with no API keys yes no no no no
CI exit code and baseline regression gate yes partly no no no

Links: VoiceTest · EVA-Bench · VoiceAgentBench · τ-bench. IndicOrderBench borrows the end-state comparison and pass^k from τ-bench and the simulator-validation idea from EVA-Bench, and narrows the task to ordering so the verdict can be deterministic.

The starter pack

42 scenarios: 21 English (en-IN) and 21 romanised Hinglish (hi-en), four per failure category per language plus two demo-tagged correction scenarios. Ten are tagged smoke for a 30-second gate. Every scenario lists its acceptable end states and carries clarification rules, so an agent that asks "how many?" before acting gets a fair run. A rule-based reference agent with five switchable bugs validates the pack: the correct agent passes 42/42, and each bug fails exactly its own category.

Audio: iob synth starter --provider sarvam generates caller clips with Sarvam bulbul (set SARVAM_API_KEY); iob perturb derives noisy, quiet or telephone-band variants; iob run --modality audio sends only audio. --provider silence makes placeholder clips so the whole audio pipeline runs offline (pipeline tests, not evaluations).

Write your own scenarios or a new language pack: docs/scenarios.md.

Honest limits

  • The runner is turn-based: no barge-in or overlapping speech is measured.
  • The Hinglish scenarios are marked unreviewed until native speakers review them (see packs/starter/REVIEW.md). Treat Hinglish numbers as provisional until then.
  • No LLM-based reference agent, generative caller or failure explanations in v0.1. The interfaces exist: OrderBackend.tool_specs(), AgentAdapter, Transcriber.
  • No LiveKit or Pipecat native adapter yet; the HTTP turn protocol is the integration point.

Roadmap

  • Native-speaker review of the Hinglish pack, then Telugu and Tamil packs.
  • A LiveKit Agents adapter and a Pipecat adapter.
  • A generative caller for unscripted clarifications.
  • An LLM reference agent, to publish measured ordering accuracy for real STT + LLM + TTS stacks.
  • PyPI release.

If you build or run voice ordering agents and want a category or a language covered, open an issue with a real failure you have seen. That is the most useful contribution right now.

Contributing and development

Language review, scenarios, adapters and report improvements are all welcome and small. See CONTRIBUTING.md.

uv sync --all-extras --dev
uv run pytest -q          # ~300 tests, no network
uv run iob validate starter

License

Apache-2.0. See LICENSE.

Keywords: voice agent testing, voice AI evaluation, conversational AI QA, speech agent benchmark, LLM agent benchmark, tool calling evaluation, function calling accuracy, order accuracy, Hinglish, Hindi, Indian languages, Indic NLP, LiveKit agents, Pipecat, Vapi, Retell, Sarvam, τ-bench, pass^k, regression testing for voice agents, CI for AI agents.

Metadata

Release files for indicorderbench 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for indicorderbench 0.1.0
File Size Uploaded
indicorderbench-0.1.0.tar.gz 215.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for indicorderbench 0.1.0
File Interpreter ABI Platform
indicorderbench-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 337.7 kB

Release files / indicorderbench-0.1.0.tar.gz

Download URL indicorderbench-0.1.0.tar.gz
Size 215.3 kB
Tags Source
SHA-256 checksum
How to use checksums
f49af4331a93a0a7966cf7bbe1bf08ae6ad10ebe8b3a97da736c38a4b5b2b0ef
BLAKE2b-256 checksum
How to use checksums
4f643d2a37471f6b7a96d143318c6e579af6d88ba184a7817e1de3d26ff3cea8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 1, 2026.

Transparency log

Release files / indicorderbench-0.1.0-py3-none-any.whl

Download URL indicorderbench-0.1.0-py3-none-any.whl
Size 122.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
a2ea8d0fbc8bc85a6e2c3dbdf4519979d9aa43524b27659048cef6380e41b472
BLAKE2b-256 checksum
How to use checksums
9896fd1164861bf2c08a8f8ab2d1c94742b2ca5e2963eeefa8ca57d8299abcc0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 1, 2026.

Transparency log

Release history Release notifications | RSS feed

0.1.1

2 release files

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page