Skip to main content

multivon-eval

PyPI Python License Tests

Did your AI application get better—or did the measurement change?

multivon-eval is a Python library for testing LLM applications, RAG systems, and agents. Define task-specific cases, grade outputs, compare changes, and inspect quality failures separately from errors and missing evidence. Runs and reports stay local; LLM graders call the judge provider you configure. No hosted account is required.

Documentation · Examples · Benchmarks · Changelog

Current release: 0.19.0 — September 17, 2026. Python 3.10+, Apache 2.0. Migration notes.

Case manifests and trial evidence support safer comparisons and regrading. Use Hugging Face for dataset operations, Inspect for execution, and Multivon's acceptance policies for required checks and task slices. These features ship in 0.19.0. The implementation program tracks unfinished work; reuse decisions keep the integration boundaries explicit.

Start in 30 seconds

pip install multivon-eval

Run this complete example. It needs no API key:

from multivon_eval import EvalCase, EvalSuite, ExactMatch

suite = EvalSuite("capital lookup")
suite.add_cases([
    EvalCase(input="Capital of France?", expected_output="Paris"),
    EvalCase(input="Capital of Japan?", expected_output="Tokyo"),
])
suite.add_evaluators(ExactMatch())

# Replace this fixture with your application: a function from str to str.
answers = {"Capital of France?": "Paris", "Capital of Japan?": "Tokyo"}
report = suite.run(
    answers.__getitem__,
    fail_threshold=1.0,
    save_json="results.json",
    save_html="results.html",
    verbose=False,
)
print(f"{report.passed}/{report.evaluated} passed; {report.errors} errors")
# 2/2 passed; 0 errors

Open results.html to inspect each verdict. The example verifies a small lookup fixture; it does not establish the quality of a real AI application. Start your own suite by defining task success.

The Label Studio review bridge exchanges saved text/trace trials and preserves review disagreements and missing coverage. Saved-score calibration reuses scikit-learn and SciPy for development fitting and held-out source analysis. OpenTelemetry interoperability grades retained OTLP traces and emits standard evaluation events through your existing SDK. Environment outcome checks reuse Gymnasium and independently observed state to catch missing writes, duplicate writes and forbidden changes. The failure investigation workflow connects saved trial comparison to Label Studio review and development regression cases. The vision grader audit corrects empty-claim perfect scores, invalid-judgment handling and Anthropic SDK 1.x compatibility. Content-bound media connects verified image, PDF, audio and video inputs to native Inspect logs and W3C verdict references. Controlled robustness validates candidate oracles and preserves rejected or unknown transformations. Hardness filtering keeps missing measurements separate from model failures. Experimental world-model evaluation binds numeric state forecasts to native environment transitions and tests action effects, uncertainty and actual planning outcomes. Its tested scope is fully observed state-space models, not video generation or real robots. These experimental interfaces ship in 0.19.0 with their documented scope limits.

A real workflow example

The document-to-ledger study reuses CORD receipts, pdfhell generators and Inspect execution. Models write to SQLite; independent checks verify the saved state. Both models fail the frozen policy on 39 held-out source documents. The report separates wrong amounts, missing posts, transcription-only mismatches and integration errors, with raw-log hashes, cost accounting and a reproducible protocol. It is a sandbox study, not proof of production readiness or a new dataset.

The execution-control study shows why a successful text reply can coexist with an unfinished task: native Inspect limits stop six local writes, and independent SQLite checks catch the missing commits. A completed control verifies the positive path.

Why use it

Need Available today
Catch regressions Deterministic checks, LLM graders, and saved baseline/proposal comparisons.
Detect missing evidence Separate quality failures, model/judge errors, and skipped checks. Active quality gates block errors and skipped coverage by default.
Check the graders Validate reference answers, measure agreement with reviewed labels, and inspect grader reasons.
Measure variability Repeated trials, flakiness, pass@k/pass^k, and confidence intervals with documented assumptions.
Keep results portable Local JSON, HTML, CSV, JUnit, and optional audit artifacts.

Pick your path

Task Starting point
First offline suite multivon-eval init -t quickstart -d my-eval
RAG / question answering multivon-eval init -t rag
Agent tool use multivon-eval init -t agent
LangGraph agent multivon-eval init -t agent-langgraph
OpenAI Agents SDK agent multivon-eval init -t agent-openai-sdk
Multi-turn conversations multivon-eval init -t conversation
Existing logs Score recorded outputs
Unsure what to measure Bootstrap a starter suite

Bootstrap suggests evaluators and synthetic cases. Its p25 threshold suggestions are provisional score summaries, not calibration against human acceptance labels. Review them before using them as release criteria.

Add an LLM judge

Use deterministic checks for exact requirements and judges for qualities that need interpretation. Install the relevant provider SDK and set its API key; configure a local judge if you want to avoid hosted calls.

from multivon_eval import EvalCase, EvalSuite, Faithfulness, JudgeConfig, configure

configure(JudgeConfig(provider="anthropic", model="claude-haiku-4-5"))
suite = EvalSuite("policy answers")
suite.add_cases([EvalCase(
    input="What is the refund window?",
    context="Refunds are available within 30 days of purchase.",
)])
suite.add_evaluators(Faithfulness())

# your_app(prompt) must call your application, including its retrieval step.
report = suite.run(your_app, fail_threshold=0.90, save_json="policy-results.json")

Judge providers include Anthropic, OpenAI, Google, Ollama, and LiteLLM; OpenAI-compatible endpoints support local servers. See judge configuration for setup, historical threshold packs, and the current temperature-forwarding limitation. A judge score is an estimate to validate on your task, not ground truth.

Evaluators — 44 across 7 tiers

Family Examples Judge calls?
Deterministic ExactMatch, Contains, JSONSchemaEval, BLEU, ROUGE, Latency No
LLM judge Faithfulness, Hallucination, Relevance, AnswerAccuracy, GEval Yes
Agent trace ToolCallAccuracy, ToolArgumentAccuracy, TaskCompletion Some
Conversation KnowledgeRetention, ConversationCompleteness, TurnConsistency Yes
Compliance checks PIIEvaluator, SchemaEvaluator No
Multimodal VQAFaithfulness, DocumentGrounding (experimental) Yes
Consistency SelfConsistency Yes

Browse the evaluator reference. For tool expectations, None means unspecified, [] means no calls expected, and require_order=True checks an ordered subsequence. A matching trace alone does not prove the task changed external state correctly.

Use it in CI

fail_threshold checks absolute quality. In 0.17.0, an active gate returns:

Exit Meaning
0 The configured gate passed.
1 Completed measurements failed the quality threshold.
2 Evidence is indeterminate: errors, empty runs, or skipped coverage.

Set max_error_rate= explicitly to allow an error budget. Pass save_json=, save_html=, or save_junit_xml= into suite.run() so reports are written before a failing gate raises.

For a saved baseline/proposal comparison:

multivon-eval compare baseline.json proposal.json --fail-on-regression

This also blocks incomplete or unmatched comparisons. It flags case regressions without waiting for statistical significance; no detected regression is not proof of equivalence. See CI/CD and statistical assumptions.

Evidence and limitations

The repository publishes benchmark scripts and historical results. The first full RAGChecker judge study measures current AnswerAccuracy at Pearson 0.499 against overall human preference on 280 cases (95% case-bootstrap interval 0.413–0.572). It does not establish an advantage over a same-model direct judge: the paired interval includes zero, and the ranking reverses under a post-hoc formatting check. QAG used four calls per response versus one for the direct judge.

One cross-task measurement reported F1 0.830 on 60 HaluEval summarization outputs from 30 source examples, using a QA-selected threshold. It is a small, maintainer-run result with generated hallucination labels; it does not establish state-of-the-art accuracy. Historical live benchmarks have not been rerun after the 0.17.0 grader changes.

Confidence intervals do not correct biased labels, correlated cases, or grader mistakes. Validate your success criteria and inspect errors, skips, and important task slices. Optional compliance reports organize evidence; they do not certify legal compliance.

A case with one measured check and another skipped check still counts as evaluated. The built-in gate blocks wholly skipped cases; enforce per-check coverage separately when every evaluator is required.

Useful commands

multivon-eval validate eval.py                 # check reference outputs against graders
multivon-eval view --dir runs/                 # browse saved reports
multivon-eval bootstrap --product PRODUCT.md --traces TRACES.jsonl
multivon-eval generate --from docs/faq.md --n 20
multivon-eval assess traces.jsonl              # inspect input quality locally
multivon-eval staleness .                      # inspect prompt/case drift
multivon-eval doctor --no-ping --json          # check configuration offline

doctor exits 0 when clean, 2 when it finds warnings, and 1 when it finds an error. Use multivon-eval --help for all commands.

Current release — 0.19.0

  • Retain provider attempts and native usage in a durable journal; reconcile costs only when coverage is complete.
  • Bind target, grader and Inspect retry compatibility to recorded settings and declared dependencies.
  • Reuse Label Studio, OpenTelemetry, Gymnasium, Hugging Face and Inspect through explicit adapters.
  • Retain content-bound media, environment outcomes and controlled-robustness evidence.
  • Block incomplete claim extraction, capped prefixes and missing claim verdicts from passing Faithfulness.
  • Validate report envelopes with a packaged JSON Schema and preserve historical report loading.

See migration notes, the worked document study, and the changelog. This release does not establish complete provider capture, opaque callback compatibility, customer validation or SoTA accuracy.

Strict agent judgment evidence and declared grader dependencies and native provider evidence ship in 0.19.0. Instrumented SDK calls retain request attempts and complete native usage fields; an optional SQLite journal preserves dispatched attempts after process death. Unobserved transports, streaming usage and missing responses stay explicit gaps. The accounting workflow reuses LiteLLM pricing and rejects incomplete provider budget evidence. Recorded judge subtotals do not establish a complete run cost. Target/runner snapshots and declare_target expose intended interventions and observed changes during execution; regrading preserves original target evidence. Execution controls reuse native Inspect limits and retain stop reasons. Saved-output regrading cannot erase an undeclared stop or invalidation; bounded-task acceptance requires an explicit policy and outcome checks. Native async cancellation drains owned tasks, with synchronous worker-thread limits documented. Declared Inspect tasks bind retry compatibility to task inputs, grader settings and dependency revisions. The native retry study includes a changed-rule control where preserved scores look perfect while fresh grading rejects the task; mixed evidence cannot pass acceptance. Agent judge failures remain missing measurements; opaque grader callback contracts and observed dependency drift block verified comparisons. These capabilities do not establish judge accuracy or complete provider accounting.

Use pdfhell for adversarial document fixtures, multivon-mcp for MCP access, and eval-action for GitHub workflows. The library can sit alongside your existing tracing system.

Issues and pull requests are welcome. See CONTRIBUTING.md for setup and testing, and extension contracts for custom graders, native environment integrations, report migration and tested dependency combinations. Apache 2.0 — Multivon.

Research direction: regulated enterprise workflows and moat, with public benchmarks for verifier quality. The first RAGChecker baseline reproduction matches published correlations on 280 cases using released predictions and the unchanged upstream scorer. Our subsequent same-model study above establishes an initial measurement, not a SoTA claim on those evaluator benchmarks.

Metadata

Release files for multivon-eval 0.19.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for multivon-eval 0.19.0
File Size Uploaded
multivon_eval-0.19.0.tar.gz 699.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for multivon-eval 0.19.0
File Interpreter ABI Platform
multivon_eval-0.19.0-py3-none-any.whl Python 3 none any Details

Total release size: 1.2 MB

Release files / multivon_eval-0.19.0.tar.gz

Download URL multivon_eval-0.19.0.tar.gz
Size 699.5 kB
Tags Source
SHA-256 checksum
How to use checksums
4dd19641723178372caa7d0dc2d1651a66c3a601ed431349cef9fdb3d831fdd9
BLAKE2b-256 checksum
How to use checksums
1d80d67d887bc26d3fac4e6991314ff71e225487125e6afb0dc7d636b4260034
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.13

Release files / multivon_eval-0.19.0-py3-none-any.whl

Download URL multivon_eval-0.19.0-py3-none-any.whl
Size 505.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
de8675d7c9da352495634397fce615ea7e19e8d0c09133a0d15092a669df79b2
BLAKE2b-256 checksum
How to use checksums
17807460751988b2a34e5a4472ccdd85363d1ab3d4f3faa5a83ef1216eea25a8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.13

Release history Release notifications | RSS feed

0.20.0

2 release files

This release

0.19.0 This release

2 release files

0.18.0

2 release files

0.17.0

2 release files

0.16.1

2 release files

0.16.0

2 release files

0.15.2

2 release files

0.15.1

2 release files

0.15.0

2 release files

0.14.0

2 release files

0.13.0

2 release files

0.12.3

2 release files

0.12.2

2 release files

0.12.1

2 release files

0.12.0

2 release files

0.11.1

2 release files

0.11.0

2 release files

0.10.1

2 release files

0.10.0

2 release files

0.9.8

2 release files

0.9.7

2 release files

0.9.6

2 release files

0.9.5

2 release files

0.9.4

2 release files

0.9.3

2 release files

0.9.2

2 release files

0.9.1

2 release files

0.9.0

2 release files

0.8.2

2 release files

0.8.1

2 release files

0.8.0

2 release files

0.7.8

2 release files

0.7.7

2 release files

0.7.6

2 release files

0.7.5

2 release files

0.7.4

2 release files

0.7.3

2 release files

0.7.2

2 release files

0.7.1

2 release files

0.7.0

2 release files

0.6.1

2 release files

0.6.0

2 release files

0.5.0

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.0

1 release file

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page