Judgemetry
A reliability lab for LLM judges. Measure position bias, verbosity bias, and calibration against human labels — then fix the calibration.
Model-graded evals are everywhere: an LLM-as-judge picks the better of two answers, and that pairwise preference drives leaderboards, regression gates, and RLHF-adjacent pipelines. The judge itself is almost never validated. Judges systematically prefer whichever answer they read first (position bias), reward longer answers beyond what humans do (verbosity bias), favor their own model family, and report confidences that don't match their accuracy (eval calibration). Judgemetry treats the judge like a lab instrument: it measures each failure mode against human preference labels, puts a bootstrap confidence interval on every number, recalibrates the confidence channel, and hands you a one-file HTML report card with letter grades — so "we use an LLM judge" can come with an error bar.
Zero runtime dependencies. Pure Python standard library — the bootstrap,
Cohen's kappa, ECE, isotonic regression, and even the SVG charts are
hand-rolled. The one optional extra (anthropic) is imported lazily and only
needed for live runs. Python 3.10+.
60-second start (no keys, no network)
pip install git+https://github.com/sedai77/judgemetry-llm-judge-reliability
judgemetry demo
(Not on PyPI yet — install from git until the first release is published;
pip install judgemetry will work after that.)
That runs the entire pipeline — dataset, double-pass judging with swapped
answer positions, metrics with bootstrap CIs, isotonic recalibration, HTML
report — on a bundled synthetic fixture with planted biases, judged by a
deterministic simulated judge. No API key, no network call, no other package.
It prints the report card and writes report.html. Actual output (the demo
is deterministic, so your numbers will match; the per-axis "Reading" column is
trimmed here for width):
Judge simulated:v1 — 300 pairs
Axis Grade Measurement
Position D P(first) 63.7% [60.5%, 66.8%]
Consistency F flip rate 39.3% [34.0%, 44.7%]
Human agreement C accuracy 73.0% [69.5%, 76.7%] · κ 0.33 [0.27, 0.39]
Verbosity D +16.3 pts [+12.0, +20.7]
Self-preference D +15.4 pts [+9.4, +21.8]
Calibration B ECE 5.2% [3.3%, 10.0%] → 5.4% [2.5%, 9.7%]
Those poor grades are the point — the fixture's judge was built to be biased, and the numbers above are the planted biases read back off the instruments. (They sit a bit below the raw planted tilts because pairs where every tilt lines up clamp at probability 1; the flagship test accounts for that exactly. The same clamping mutes the planted overconfidence to a ~5-pt ECE — small enough that isotonic recalibration, honestly compared on the same held-out half, doesn't reliably improve it here; the report says so rather than flattering the method.) See the next section.
How it validates itself
A metrics library asks for trust; a measurement instrument should ship with a calibration certificate. Judgemetry's demo fixture is generated from a known model of a bad judge: base agreement with humans of 0.78, then additive tilts of +0.18 toward the first slot, +0.22 toward the longer answer, +0.15 toward its "own" model, and confidence inflated by 25% of the remaining headroom. Those parameters are chosen, not observed — which means every metric has a closed-form expected value before the pipeline ever runs.
The flagship test (tests/test_demo_recovers_planted_biases.py) runs the full
offline pipeline and asserts each estimate lands on its closed-form expectation
within tolerance. Without clamping those expectations would be the naive sums —
slot-1 preference = 0.5 + 0.18, verbosity bias = +0.22, consistency-aware
accuracy = 0.78 — but the default plant pushes some pair configurations past
probability 1, so the test enumerates the clamped cases exactly (expectations
of 0.637 / +0.194 / 0.723, which is what the demo's measured numbers land on
within their CIs). It also asserts that recalibration strictly improves
held-out ECE on a plant with an unambiguous overconfidence signal — before and
after are compared on the same held-out half, so a broken (identity)
recalibrator produces exact equality and fails — and that two runs are
byte-identical. The same pipeline runs in CI on every
commit with no secrets. So the demo is simultaneously the quick start, the CI
proof, and the evidence that the instruments measure what they claim. The
derivations and their fine print are in DESIGN.md — including
what this does not prove (that real judges behave like the synthetic one).
Live run (real judge, real labels, real money)
pip install "judgemetry[anthropic] @ git+https://github.com/sedai77/judgemetry-llm-judge-reliability"
export ANTHROPIC_API_KEY=sk-ant-...
judgemetry fetch --dest data/ # MT-Bench human judgments (CC-BY-4.0) -> data/mtbench_pairs.jsonl
judgemetry run --dataset data/mtbench_pairs.jsonl \
--judge anthropic:claude-opus-5 \
--limit 200 --max-cost-usd 5 --batch \
-o judged.jsonl
judgemetry report judged.jsonl -o report.html
Honest notes before you run this:
- It costs money, and results vary by judge model, prompt distribution, and dataset — the demo's numbers say nothing about any real judge.
--batchuses the Anthropic Message Batches API at 50% of standard token prices; without it, requests go one at a time at full price.- Every verdict is cached (SQLite) the moment it arrives. Hitting
--max-cost-usdstops cleanly (exit code 3) and a rerun resumes from the cache, paying only for what's missing. Reruns of a finished run are free. Batch runs submit in chunks of 200 judgments with the cap checked between chunks, so a capped batch run can overshoot by at most one chunk's cost — and a cap crossed on the final chunk still completes and writes the output. - The judge sees each pair twice (positions swapped) — budget 2 judgments per pair.
- Human labels come from
lmsys/mt_bench_human_judgments(CC-BY-4.0), normalized to majority labels per pair.
The metrics
One line each; exact definitions and rationale in DESIGN.md.
| Metric | One-liner |
|---|---|
| Flip rate | Fraction of pairs where swapping answer order changes the verdict — pure test–retest noise. |
| Slot-1 preference | P(winner is the first-presented answer), over judgments that picked a side; 0.5 is fair, deviation is position bias with a sign. |
| Consistency-aware accuracy | Agreement with the human label scored over both orderings; an order-dependent "correct" earns half credit. |
| Cohen's kappa | Chance-corrected agreement between the human label and the judge's cross-ordering consensus verdict. |
| Verbosity bias | P(judge prefers the longer answer) minus P(humans do) — rewards for length beyond human taste. |
| Self-preference | Same subtraction, on pairs where exactly one answer comes from the judge's own model family. |
| ECE / Brier | Confidence vs. empirical correctness: 10-bin expected calibration error, plus Brier as the proper-scoring cross-check. |
| Recalibration | Isotonic (PAV) map from confidence to correctness, fitted on an even/odd pair split; before and after are both evaluated on the held-out half so the comparison isolates recalibration. |
Limitations
- Pairwise-preference judges only. No absolute scoring, rubric grading, or ranking of 3+ candidates (yet — see roadmap).
- Single-turn comparisons. Multi-turn conversations are out of scope.
- One live backend so far (Anthropic). The
Judgeprotocol is a tiny surface — anameplus onejudge()method, with an optionaljudge_batchfor batch-capable backends — designed for third-party backends; seejudgemetry/judges.py. - The demo data is synthetic and is labeled as such in the report. It certifies the instruments, not any real judge.
- Proxies with blind spots: verbosity uses character length; self- preference uses model-name matching. DESIGN.md §6 lists the full threats-to-validity inventory.
Related work
- judgecal — calibration tooling aimed at reward models and local judges in offline GPU batch settings; if your judge is a reward model running on your own hardware, it is the better fit. Judgemetry sits in a different lane: hosted-API judges scored against human preference labels, a bias battery (position/verbosity/self) beyond calibration alone, and a shareable report card.
- MT-Bench / Chatbot Arena (Zheng et al., 2023, Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena) — the source of our human labels and the paper that mainstreamed position and verbosity bias in LLM judges.
- Position-bias literature — e.g. Wang et al., 2023, Large Language Models are not Fair Evaluators, on order sensitivity in pairwise judging; our two-pass swapped design is the standard mitigation turned into a measurement.
- Format-restriction literature (Let Me Speak Freely?-adjacent) — constrained/structured output can itself shift model behavior; relevant because Judgemetry elicits verdicts as structured JSON, so measured biases are properties of judge plus elicitation protocol.
Roadmap
- OpenAI and local (OpenAI-compatible / llama.cpp server) judge backends.
- Rubric-decomposition scoring (judge sub-criteria, aggregate transparently).
- Cost-vs-reliability frontier: grade several judge models on the same pairs, plot dollars against kappa.
- CI-integration mode: a machine-readable gate ("fail the build if flip rate CI exceeds X") for eval pipelines.
FAQ
Why do judges flip when I swap answer order? Autoregressive scoring is not symmetric in its inputs: the first answer conditions how the second is read, and preference for a slot (either slot) shows up across models and prompts. It is the best-replicated LLM-judge failure mode, which is why Judgemetry judges every pair twice by construction and reports the flip rate first.
Why isotonic regression and not Platt scaling? Platt assumes the miscalibration is sigmoid-shaped. LLM confidences cluster on round numbers and saturate near the top — isotonic assumes only "higher confidence should not mean lower accuracy", which is the weakest assumption under which recalibration means anything. Trade-off and split protocol in DESIGN.md §3.
How many pairs do I need?
Let the confidence intervals tell you: every estimate ships with a bootstrap
95% CI, and the CI width is the answer for your dataset and judge. As a
rough prior, rate-like metrics need a few hundred pairs for ±0.05-ish
intervals; small subgroups (self-preference especially) stay wide much
longer. Start with --limit 200 --max-cost-usd 5, look at the intervals,
and buy more pairs only if the question you care about is still ambiguous.
Can I use my own dataset?
Yes. judgemetry run --dataset yours.jsonl accepts JSONL where each line is:
{"schema": "judgemetry/pair@1", "pair_id": "q1-m1-m2", "question": "...",
"answer_a": "...", "answer_b": "...", "model_a": "gpt-x", "model_b": "claude-y",
"human_winner": "A"}
human_winner is "A", "B", or "tie" for the answers as listed
(Judgemetry handles the position swapping itself — never pre-swap).
model_a/model_b may be "" if unknown; self-preference is then skipped.
Can I score verdicts from a judge Judgemetry can't call?
Yes. The CLI only drives Anthropic judges (--judge anthropic[:MODEL]), but
judgemetry report accepts any judged-pair JSONL, so run your own judge —
each pair twice, positions swapped — and write one judgemetry/judged@1
object per line:
{"schema": "judgemetry/judged@1",
"record": {"pair_id": "q1-m1-m2", "question": "...", "answer_a": "...",
"answer_b": "...", "model_a": "gpt-x", "model_b": "claude-y",
"human_winner": "A"},
"forward": {"winner": "A", "confidence": 0.8, "raw": "", "refused": false},
"backward": {"winner": "A", "confidence": 0.7, "raw": "", "refused": false},
"judge": "my-judge:v1"}
forward is the verdict with the answers presented as (answer_a, answer_b),
backward with them swapped — but the backward winner is expressed in the
original frame: "A" always means answer_a won, in both verdicts. A
refused judgment must be recorded as {"winner": "tie", "confidence": 0.5, "refused": true}. Then: judgemetry report judged.jsonl -o report.html.
(Set self_model via the library — metrics.compute_report_metrics(judged, self_model="...") — if you want self-preference for a non-Anthropic judge.)
Exact field contracts: docs/SPEC.md and judgemetry/records.py.
Does the report need an internet connection to view?
No. One HTML file, inline CSS, inline SVG, no external requests, dark/light
via prefers-color-scheme. Email it, attach it to a PR, archive it.
Contributing
See CONTRIBUTING.md — dev setup is uv plus nothing, and
docs/SPEC.md is the authoritative internal contract. Security
policy in SECURITY.md.
MIT © 2026 Judgemetry contributors — see LICENSE.
Metadata
Release files for judgemetry 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| judgemetry-0.1.0.tar.gz | 140.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| judgemetry-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 193.5 kB
Release files / judgemetry-0.1.0.tar.gz
| Download URL | judgemetry-0.1.0.tar.gz |
|---|---|
| Size | 140.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
db1761e53783c0d5a405989f8c522bcce9b385d50a4d4012b667bfe084a3832b
|
|
BLAKE2b-256 checksum How to use checksums |
97286d170dbe724c395d13ac96f8386c0883bc999353c4cda0a902e62916a78b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jul 28, 2026.
Transparency logRelease files / judgemetry-0.1.0-py3-none-any.whl
| Download URL | judgemetry-0.1.0-py3-none-any.whl |
|---|---|
| Size | 53.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
9233f0f586d2e6af1f6af42a84daa421d5352660ff648805471fdf3f82b66a01
|
|
BLAKE2b-256 checksum How to use checksums |
1af3643145078bac86e46e120aa457726cf9c191aebfd2e292be9e2ac141cc78
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jul 28, 2026.
Transparency log