Skip to main content

Judgemetry

A reliability lab for LLM judges. Measure position bias, verbosity bias, and calibration against human labels — then fix the calibration.

CI Python 3.10+ License: MIT

Model-graded evals are everywhere: an LLM-as-judge picks the better of two answers, and that pairwise preference drives leaderboards, regression gates, and RLHF-adjacent pipelines. The judge itself is almost never validated. Judges systematically prefer whichever answer they read first (position bias), reward longer answers beyond what humans do (verbosity bias), favor their own model family, and report confidences that don't match their accuracy (eval calibration). Judgemetry treats the judge like a lab instrument: it measures each failure mode against human preference labels, puts a bootstrap confidence interval on every number, recalibrates the confidence channel, and hands you a one-file HTML report card with letter grades — so "we use an LLM judge" can come with an error bar.

Zero runtime dependencies. Pure Python standard library — the bootstrap, Cohen's kappa, ECE, isotonic regression, and even the SVG charts are hand-rolled. The one optional extra (anthropic) is imported lazily and only needed for live runs. Python 3.10+.

60-second start (no keys, no network)

pip install git+https://github.com/sedai77/judgemetry-llm-judge-reliability
judgemetry demo

(Not on PyPI yet — install from git until the first release is published; pip install judgemetry will work after that.)

That runs the entire pipeline — dataset, double-pass judging with swapped answer positions, metrics with bootstrap CIs, isotonic recalibration, HTML report — on a bundled synthetic fixture with planted biases, judged by a deterministic simulated judge. No API key, no network call, no other package. It prints the report card and writes report.html. Actual output (the demo is deterministic, so your numbers will match; the per-axis "Reading" column is trimmed here for width):

Judge simulated:v1 — 300 pairs
Axis             Grade  Measurement
Position         D      P(first) 63.7% [60.5%, 66.8%]
Consistency      F      flip rate 39.3% [34.0%, 44.7%]
Human agreement  C      accuracy 73.0% [69.5%, 76.7%] · κ 0.33 [0.27, 0.39]
Verbosity        D      +16.3 pts [+12.0, +20.7]
Self-preference  D      +15.4 pts [+9.4, +21.8]
Calibration      B      ECE 5.2% [3.3%, 10.0%] → 5.4% [2.5%, 9.7%]

Those poor grades are the point — the fixture's judge was built to be biased, and the numbers above are the planted biases read back off the instruments. (They sit a bit below the raw planted tilts because pairs where every tilt lines up clamp at probability 1; the flagship test accounts for that exactly. The same clamping mutes the planted overconfidence to a ~5-pt ECE — small enough that isotonic recalibration, honestly compared on the same held-out half, doesn't reliably improve it here; the report says so rather than flattering the method.) See the next section.

How it validates itself

A metrics library asks for trust; a measurement instrument should ship with a calibration certificate. Judgemetry's demo fixture is generated from a known model of a bad judge: base agreement with humans of 0.78, then additive tilts of +0.18 toward the first slot, +0.22 toward the longer answer, +0.15 toward its "own" model, and confidence inflated by 25% of the remaining headroom. Those parameters are chosen, not observed — which means every metric has a closed-form expected value before the pipeline ever runs.

The flagship test (tests/test_demo_recovers_planted_biases.py) runs the full offline pipeline and asserts each estimate lands on its closed-form expectation within tolerance. Without clamping those expectations would be the naive sums — slot-1 preference = 0.5 + 0.18, verbosity bias = +0.22, consistency-aware accuracy = 0.78 — but the default plant pushes some pair configurations past probability 1, so the test enumerates the clamped cases exactly (expectations of 0.637 / +0.194 / 0.723, which is what the demo's measured numbers land on within their CIs). It also asserts that recalibration strictly improves held-out ECE on a plant with an unambiguous overconfidence signal — before and after are compared on the same held-out half, so a broken (identity) recalibrator produces exact equality and fails — and that two runs are byte-identical. The same pipeline runs in CI on every commit with no secrets. So the demo is simultaneously the quick start, the CI proof, and the evidence that the instruments measure what they claim. The derivations and their fine print are in DESIGN.md — including what this does not prove (that real judges behave like the synthetic one).

Live run (real judge, real labels, real money)

pip install "judgemetry[anthropic] @ git+https://github.com/sedai77/judgemetry-llm-judge-reliability"
export ANTHROPIC_API_KEY=sk-ant-...

judgemetry fetch --dest data/          # MT-Bench human judgments (CC-BY-4.0) -> data/mtbench_pairs.jsonl
judgemetry run --dataset data/mtbench_pairs.jsonl \
    --judge anthropic:claude-opus-5 \
    --limit 200 --max-cost-usd 5 --batch \
    -o judged.jsonl
judgemetry report judged.jsonl -o report.html

Honest notes before you run this:

  • It costs money, and results vary by judge model, prompt distribution, and dataset — the demo's numbers say nothing about any real judge.
  • --batch uses the Anthropic Message Batches API at 50% of standard token prices; without it, requests go one at a time at full price.
  • Every verdict is cached (SQLite) the moment it arrives. Hitting --max-cost-usd stops cleanly (exit code 3) and a rerun resumes from the cache, paying only for what's missing. Reruns of a finished run are free. Batch runs submit in chunks of 200 judgments with the cap checked between chunks, so a capped batch run can overshoot by at most one chunk's cost — and a cap crossed on the final chunk still completes and writes the output.
  • The judge sees each pair twice (positions swapped) — budget 2 judgments per pair.
  • Human labels come from lmsys/mt_bench_human_judgments (CC-BY-4.0), normalized to majority labels per pair.

The metrics

One line each; exact definitions and rationale in DESIGN.md.

Metric One-liner
Flip rate Fraction of pairs where swapping answer order changes the verdict — pure test–retest noise.
Slot-1 preference P(winner is the first-presented answer), over judgments that picked a side; 0.5 is fair, deviation is position bias with a sign.
Consistency-aware accuracy Agreement with the human label scored over both orderings; an order-dependent "correct" earns half credit.
Cohen's kappa Chance-corrected agreement between the human label and the judge's cross-ordering consensus verdict.
Verbosity bias P(judge prefers the longer answer) minus P(humans do) — rewards for length beyond human taste.
Self-preference Same subtraction, on pairs where exactly one answer comes from the judge's own model family.
ECE / Brier Confidence vs. empirical correctness: 10-bin expected calibration error, plus Brier as the proper-scoring cross-check.
Recalibration Isotonic (PAV) map from confidence to correctness, fitted on an even/odd pair split; before and after are both evaluated on the held-out half so the comparison isolates recalibration.

Limitations

  • Pairwise-preference judges only. No absolute scoring, rubric grading, or ranking of 3+ candidates (yet — see roadmap).
  • Single-turn comparisons. Multi-turn conversations are out of scope.
  • One live backend so far (Anthropic). The Judge protocol is a tiny surface — a name plus one judge() method, with an optional judge_batch for batch-capable backends — designed for third-party backends; see judgemetry/judges.py.
  • The demo data is synthetic and is labeled as such in the report. It certifies the instruments, not any real judge.
  • Proxies with blind spots: verbosity uses character length; self- preference uses model-name matching. DESIGN.md §6 lists the full threats-to-validity inventory.

Related work

  • judgecal — calibration tooling aimed at reward models and local judges in offline GPU batch settings; if your judge is a reward model running on your own hardware, it is the better fit. Judgemetry sits in a different lane: hosted-API judges scored against human preference labels, a bias battery (position/verbosity/self) beyond calibration alone, and a shareable report card.
  • MT-Bench / Chatbot Arena (Zheng et al., 2023, Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena) — the source of our human labels and the paper that mainstreamed position and verbosity bias in LLM judges.
  • Position-bias literature — e.g. Wang et al., 2023, Large Language Models are not Fair Evaluators, on order sensitivity in pairwise judging; our two-pass swapped design is the standard mitigation turned into a measurement.
  • Format-restriction literature (Let Me Speak Freely?-adjacent) — constrained/structured output can itself shift model behavior; relevant because Judgemetry elicits verdicts as structured JSON, so measured biases are properties of judge plus elicitation protocol.

Roadmap

  • OpenAI and local (OpenAI-compatible / llama.cpp server) judge backends.
  • Rubric-decomposition scoring (judge sub-criteria, aggregate transparently).
  • Cost-vs-reliability frontier: grade several judge models on the same pairs, plot dollars against kappa.
  • CI-integration mode: a machine-readable gate ("fail the build if flip rate CI exceeds X") for eval pipelines.

FAQ

Why do judges flip when I swap answer order? Autoregressive scoring is not symmetric in its inputs: the first answer conditions how the second is read, and preference for a slot (either slot) shows up across models and prompts. It is the best-replicated LLM-judge failure mode, which is why Judgemetry judges every pair twice by construction and reports the flip rate first.

Why isotonic regression and not Platt scaling? Platt assumes the miscalibration is sigmoid-shaped. LLM confidences cluster on round numbers and saturate near the top — isotonic assumes only "higher confidence should not mean lower accuracy", which is the weakest assumption under which recalibration means anything. Trade-off and split protocol in DESIGN.md §3.

How many pairs do I need? Let the confidence intervals tell you: every estimate ships with a bootstrap 95% CI, and the CI width is the answer for your dataset and judge. As a rough prior, rate-like metrics need a few hundred pairs for ±0.05-ish intervals; small subgroups (self-preference especially) stay wide much longer. Start with --limit 200 --max-cost-usd 5, look at the intervals, and buy more pairs only if the question you care about is still ambiguous.

Can I use my own dataset? Yes. judgemetry run --dataset yours.jsonl accepts JSONL where each line is:

{"schema": "judgemetry/pair@1", "pair_id": "q1-m1-m2", "question": "...",
 "answer_a": "...", "answer_b": "...", "model_a": "gpt-x", "model_b": "claude-y",
 "human_winner": "A"}

human_winner is "A", "B", or "tie" for the answers as listed (Judgemetry handles the position swapping itself — never pre-swap). model_a/model_b may be "" if unknown; self-preference is then skipped.

Can I score verdicts from a judge Judgemetry can't call? Yes. The CLI only drives Anthropic judges (--judge anthropic[:MODEL]), but judgemetry report accepts any judged-pair JSONL, so run your own judge — each pair twice, positions swapped — and write one judgemetry/judged@1 object per line:

{"schema": "judgemetry/judged@1",
 "record": {"pair_id": "q1-m1-m2", "question": "...", "answer_a": "...",
            "answer_b": "...", "model_a": "gpt-x", "model_b": "claude-y",
            "human_winner": "A"},
 "forward":  {"winner": "A", "confidence": 0.8, "raw": "", "refused": false},
 "backward": {"winner": "A", "confidence": 0.7, "raw": "", "refused": false},
 "judge": "my-judge:v1"}

forward is the verdict with the answers presented as (answer_a, answer_b), backward with them swapped — but the backward winner is expressed in the original frame: "A" always means answer_a won, in both verdicts. A refused judgment must be recorded as {"winner": "tie", "confidence": 0.5, "refused": true}. Then: judgemetry report judged.jsonl -o report.html. (Set self_model via the library — metrics.compute_report_metrics(judged, self_model="...") — if you want self-preference for a non-Anthropic judge.) Exact field contracts: docs/SPEC.md and judgemetry/records.py.

Does the report need an internet connection to view? No. One HTML file, inline CSS, inline SVG, no external requests, dark/light via prefers-color-scheme. Email it, attach it to a PR, archive it.

Contributing

See CONTRIBUTING.md — dev setup is uv plus nothing, and docs/SPEC.md is the authoritative internal contract. Security policy in SECURITY.md.

MIT © 2026 Judgemetry contributors — see LICENSE.

Metadata

Release files for judgemetry 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for judgemetry 0.1.0
File Size Uploaded
judgemetry-0.1.0.tar.gz 140.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for judgemetry 0.1.0
File Interpreter ABI Platform
judgemetry-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 193.5 kB

Release files / judgemetry-0.1.0.tar.gz

Download URL judgemetry-0.1.0.tar.gz
Size 140.1 kB
Tags Source
SHA-256 checksum
How to use checksums
db1761e53783c0d5a405989f8c522bcce9b385d50a4d4012b667bfe084a3832b
BLAKE2b-256 checksum
How to use checksums
97286d170dbe724c395d13ac96f8386c0883bc999353c4cda0a902e62916a78b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 28, 2026.

Transparency log

Release files / judgemetry-0.1.0-py3-none-any.whl

Download URL judgemetry-0.1.0-py3-none-any.whl
Size 53.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
9233f0f586d2e6af1f6af42a84daa421d5352660ff648805471fdf3f82b66a01
BLAKE2b-256 checksum
How to use checksums
1af3643145078bac86e46e120aa457726cf9c191aebfd2e292be9e2ac141cc78
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 28, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page