Skip to main content

judge-audit

CI Release License: Apache-2.0 Python 3.10+

Independent calibration audits for AI judges. When a judge says 90 %, is it right 90 % of the time?

Everyone is shipping judgment models — TypeSafe's Jev, LLM-as-judge, guardrails, routers — and every one of them returns a confidence. Nobody checks whether that number means anything. judge-audit runs any judge in shadow mode against decisions your humans already made and answers the four questions that matter before you automate:

Question Metric Why a buyer cares
When it says 80 % confident, is it right 80 % of the time? ECE, reliability diagram A confident-and-wrong judge automates its own mistakes
What share of the work can I automate at zero observed errors? accuracy-coverage curve, zero-error coverage This is the ROI number
What does it really cost, and how bad is the latency tail? $ per decision, p50 / p99 The demo is cheap; the tail is what pages you
Has it drifted since last week? judge-audit check CI gate Vendors update models without telling you

judge-audit run on a labeled dataset, then the CI gate

pip install kunko-judge-audit           # every release ships Sigstore-signed build provenance
judge-audit run examples/email-routing/labels.jsonl --judge simulated   # no API key needed
judge-audit check examples/email-routing/labels.jsonl --judge simulated --baseline audit-result.json

▶ 35-second walkthrough video

simulated is a seeded simulator so you can see the whole pipeline in ten seconds; every report it touches is stamped SIMULATED. To audit a real vendor, see docs/real-audits.md.

The first independent audits of Jev

Same judge (TypeSafe Jev, via Vercel AI Gateway), three jobs, every raw response committed under docs/runs/ so anyone can recompute every number (python scripts/verify_published.py does, in CI).

Same judge, three jobs: honest, honest under attack, confidently wrong

Audit n Accuracy ECE What it shows Report
Business emails, 10 categories, clean 200 100 % 0.004 Honest when the task is easy. Synthetic templates with the category keyword in the text — a floor, not a benchmark. audit-jev-real.md
Same emails under attack: prompt injection, homoglyphs, ambiguity, PII, social engineering 200 95.5 % 0.039 Prompt injection flips 7/40 decisions, but confidence drops from 0.996 to 0.71 under attack — the judge signals its own doubt. Homoglyphs and social engineering: 0 successes. Ambiguous emails: confidence does not drop (0.95), which it should. audit-jev-adversarial.md
Task router: cheap model vs frontier model, 40 easy / 40 hard / 40 easy + cost-inflation injection 120 66.7 % 0.318 With options sent as bare labels the judge never chose the strong model: 0/40 on hard tasks at median confidence 0.96. That is exactly the constant-classifier baseline. audit-jev-router.md
The same 120 rows with a one-line description per option 120 97.5 % 0.053 37/40 hard tasks now go to the strong model, and the three misses sit at confidence 0.56–0.60 (vs 0.93 when right). Same model, same tasks: the failure was the prompt — and nothing but a calibration audit reveals it. audit-jev-router-ablation.md

Jev as a task router: bare labels vs described options

What we would tell a client. The email numbers are the vendor's story and they hold up, including under attack. The router numbers are the buyer's story: the first prompt anyone would write routed every hard task to the cheap model at 96 % median confidence, and its 66.7 % accuracy is exactly what a coin glued to "easy" scores; two descriptive sentences took the same judge to 97.5 % with confidence that finally means something. None of that is visible from a benchmark leaderboard; all of it is visible from a calibration audit on your own decisions.

Honest limits: every dataset is synthetic and seeded (generators in examples/); n is small; ground truth for routing is by construction, not by running the cheap model. Read the Caveats section of each report before quoting it.

How it works

  1. Labeled dataset — JSONL rows {state, questions, labels}. Your humans already decided; the judge must reproduce them.
  2. Shadow run — the judge scores every row. Nothing is automated; everything is recorded: decision, confidence, latency, cost, raw response.
  3. Audit report — calibration (ECE, reliability bins), selective prediction (accuracy at every coverage level), cost, latency percentiles, plus provenance: model, backend, timestamp, dataset hash.
  4. CI gate — judge-audit check --baseline baseline.json --max-ece-drift 0.02 fails the build when the judge degrades.

Exit codes: 0 ok · 1 drift detected · 2 usage or configuration error (the message says what to fix).

Judge interface

Anything that maps (state, questions) -> (decision, confidence) plugs in:

from judge_audit import Judge, Question, Judgment

class MyJudge(Judge):
    name = "my-judge"

    def decide(self, state: str, questions: list[Question]) -> list[Judgment]:
        ...  # return Judgment(question=q.name, decision="spam", confidence=0.93)

Ships with three: jev (TypeSafe Jev — and, via JEV_ENDPOINT, any Jev-compatible server such as OpenJev), llm (any chat model with a confidence prompt: Claude through the official SDK, or anything OpenAI-compatible — OpenAI, Ollama, vLLM) and simulated. Run the same dataset through several and you have the first row of the Judge Arena. Details in docs/judges.md.

Why calibration, not accuracy

Accuracy tells you who wins a benchmark. Calibration tells you what you can automate safely. A 96 %-accurate judge that is confident on the 4 % it gets wrong is a liability; a 90 %-accurate judge that flags its own doubt is an asset. Audit the honesty, not the trophy.

What lives where

Path What
src/judge_audit/ the harness: judge interface, adapters, metrics, runner, reports, CLI
examples/*/generate.py seeded dataset generators — CI regenerates and diffs them
docs/audit-*.md / .json published audits
docs/runs/ raw per-row judge responses behind each audit
scripts/ resumable audit driver, per-audit analysis, verify_published.py
tests/ metrics on hand-checked inputs, CLI exit codes, reproducibility

Roadmap

MCP server so agents audit their own judge in place → Score questions + MCE (the regulator's number) → Judge Arena (Jev vs OpenJev vs Claude vs open models on the same datasets, published) → AI Act evidence dossier. Details and reasons in docs/ROADMAP.md; the live backlog is the issues.

FAQ

Is this a benchmark? No. A benchmark ranks judges on a fixed test. An audit checks one judge on your decisions and tells you what you can automate. The datasets here are worked examples, not a leaderboard — yet.

Why is Jev the first judge? It is the first judgment model sold on calibrated confidence as the headline feature. A claim that specific deserves an independent check.

Can I trust a 100 % result? Only as far as the dataset. The clean-email audit says the judge handles templated business email; it says nothing about your inbox. That is why the adversarial and routing audits exist.

Does it send my data anywhere? Only to the judge endpoint you configure. Reports and checkpoints are local files; commit them or not.

License

Apache-2.0 — see LICENSE. Sister project: agent-assurance (deterministic checks on what an agent may do; this repo is the probabilistic half: whether its judgment can be trusted).

Release files for kunko-judge-audit 0.2.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for kunko-judge-audit 0.2.1
File Size Uploaded
kunko_judge_audit-0.2.1.tar.gz 596.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for kunko-judge-audit 0.2.1
File Interpreter ABI Platform
kunko_judge_audit-0.2.1-py3-none-any.whl Python 3 none any Details

Total release size: 624.4 kB

Release files / kunko_judge_audit-0.2.1.tar.gz

Download URL kunko_judge_audit-0.2.1.tar.gz
Size 596.4 kB
Tags Source
SHA-256 checksum
How to use checksums
6953021a1346823a02a351b6b319d13ad1768e0e6d8452368bb2d6c77f3da34d
BLAKE2b-256 checksum
How to use checksums
1bb33146f9d337a6f466724d568a7fbee16250420032f3dd7b2ee165d6695911
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 19, 2026.

Transparency log

Release files / kunko_judge_audit-0.2.1-py3-none-any.whl

Download URL kunko_judge_audit-0.2.1-py3-none-any.whl
Size 28.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
70dca023450edae7404654ca1359704ddb7beec8125a94e72141d10a689356fc
BLAKE2b-256 checksum
How to use checksums
d3ac3e21b917f4cba6b52a78802ae7cb666f8bcde02228c3f6dc9d265cc47a75
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 19, 2026.

Transparency log

Release history Release notifications | RSS feed

0.4.0

2 release files

0.3.2

2 release files

0.3.1

2 release files

0.3.0

2 release files

This release

0.2.1 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page