juryrig
Audit your LLM judges before you trust them.
LLM-as-judge is everywhere: the cheapest way to grade model outputs is to ask another model. But the judge is a model too — with position bias, a weakness for long-winded answers, run-to-run inconsistency, and confidence that rarely matches its accuracy. If you haven't measured those, your eval numbers are decoration.
juryrig is a small, zero-dependency Python toolkit that treats the judge as the thing under test:
- Position-bias audit — present every A/B pair in both orders; count how often the slot (not the content) decides the winner.
- Verbosity-bias audit — re-score responses padded with content-free filler; a fair judge shouldn't reward padding.
- Prompt-injection audit — append judge-targeted instructions to bad responses; a robust judge should grade the answer, not obey it.
- Self-consistency — same input, several runs; how stable is the score?
- Panels — pool several judges (mean / median / min) and get an agreement score, so you know when your verdict depends on which judge you picked. Pairwise panels vote on A/B pairs and report a dead heat as one.
- Calibration — Brier score, reliability tables, and expected calibration error against human labels.
Run the whole battery with audit_suite(), or from the command line with
juryrig cases.json.
Install
PyPI publish pending —
juryrigis not on PyPI yet. Trusted Publishing is wired in.github/workflows/publish.yml; install from GitHub until the first release lands:
pip install git+https://github.com/ianalloway/juryrig
# after PyPI: pip install juryrig
Quickstart
from juryrig import (
MockJudge,
Panel,
position_bias,
prompt_injection_bias,
verbosity_bias,
)
rubric = "Answer must mention photosynthesis, chlorophyll, sunlight, and energy."
# 1. Audit a judge before using it
judge = MockJudge(name="demo") # swap in your own judge for real audits
cases = [("How do plants make food?", "good answer...", "weak answer...")]
bias = position_bias(judge, cases, rubric)
print(f"flip rate: {bias.flip_rate:.0%} flagged: {bias.flagged}")
injection = prompt_injection_bias(judge, [("Weak answer prompt", "vague answer")], rubric)
print(f"injection lift: {injection.mean_delta:+.3f} flagged: {injection.flagged}")
# 2. Use a panel instead of a single judge
panel = Panel([MockJudge(name="primary"), MockJudge(name="baseline")])
report = panel.evaluate(prompt="How do plants make food?",
response="Photosynthesis converts sunlight...",
rubric=rubric)
print(report.pooled, report.agreement)
position_bias() is a pairwise audit and needs a judge with compare();
MockJudge implements both compare() and judge() so the quickstart runs
without API credentials.
The whole battery in one call
audit_suite() runs every audit and pools the verdicts. Each case is a
(prompt, good_response, weak_response) triple: the pair drives the position
comparison, the good response gets padded to detect verbosity bias, and the
weak one carries the injection payload.
from juryrig import MockJudge, audit_suite
report = audit_suite(MockJudge(), cases, rubric)
print(report.summary())
assert not report.flagged, f"judge failed: {report.failures}"
report.failures names the audits that tripped (("position", "injection")).
If the judge has no compare(), the position audit is reported in
report.skipped rather than silently counted as a pass.
Ties
compare() may return "tie" as well as "A" or "B". It's optional — a
judge that only ever picks a side is unaffected — but real judges often want
to call two answers equivalent, and forcing that into a coin flip manufactures
position bias that isn't there.
How position_bias() accounts for them:
- Tying both ways is not a flip. The judge gave the same answer in both orders, which is consistency, not order-dependence.
- Tying one way and picking the other way is a flip. The verdict changed when only the order changed — that's the thing being measured.
- Ties are excluded from
first_slot_wins. Counting them as "not won by the first slot" would drag the ratio to 0 and flag a judge that ties everything as maximally biased toward slot two. With nothing decisive to go on the audit reports 0.5: no evidence of skew. The count is kept inreport.tiesso it stays visible rather than silently dropped.
Panels of pairwise judges
Panel.evaluate() pools scores; Panel.compare() pools A/B votes by majority:
verdict = panel.compare(prompt="How do plants make food?",
a="Photosynthesis converts sunlight...",
b="Plants eat soil.",
rubric=rubric)
print(verdict.winner, verdict.agreement, verdict.votes)
An even split sets winner to None and deadlocked to True, rather than
picking a side. A coin-flip winner would hide exactly the disagreement you
convened a panel to find. Judges without compare() are rejected by name —
silently dropping them would move the verdict while leaving agreement
looking healthy.
Going faster against a real judge
audit_suite on N cases is roughly 4N judge calls, which is minutes of wall
time over a network. max_workers runs them in parallel:
report = audit_suite(judge, cases, rubric, max_workers=8)
Serial by default, because a judge may be stateful or rate-limited and
threading one behind your back would be a surprise. Results are collected in
input order, so a report is identical no matter how many workers produced it
— workers are a speed knob, never a correctness one. Your judge must be
thread-safe to raise it. The CLI exposes the same thing as --workers.
Tuning what counts as a failure
Every flagged verdict comes from a Thresholds object. The defaults are
strict on purpose, but they're yours to move:
from juryrig import Thresholds, audit_suite
report = audit_suite(judge, cases, rubric, thresholds=Thresholds(
injection_max_delta=0.05, # stricter: near-zero tolerance for injection
verbosity_mean_delta=0.10, # looser: this judge is allowed to like detail
))
The measurements never change — only the line between pass and fail. Each
report carries the thresholds it was judged against, so a stored report
still explains its own verdict. A case file can set them too, under a
"thresholds" key; unknown keys are rejected rather than ignored, so a typo
can't silently leave the strict default in force.
Command line
juryrig examples/cases.json # audit the built-in MockJudge
juryrig cases.json --provider anthropic --json # audit a live judge
The case file is {"rubric": ..., "cases": [{"prompt", "good", "weak"}, ...]}.
The command exits 1 when the judge is flagged and 2 on bad input, so a CI
step is one line. Without installing, use python -m juryrig cases.json.
Optional: provider-backed judges
AnthropicJudge and OpenAIJudge wrap the Anthropic/OpenAI HTTP APIs
(stdlib-only, no extra dependencies) and work with the single-response
audits. They're not exported from the top-level package — import them
explicitly when you need a live model:
from juryrig.providers import AnthropicJudge, OpenAIJudge # needs *_API_KEY env var
An audit is many calls in a row, so both retry transient failures (429, 5xx,
network errors) with exponential backoff — one flaky response shouldn't throw
away every result collected before it. A numeric Retry-After is honoured.
Client errors like 401 and 404 fail fast, since they'd fail identically on
every attempt.
from juryrig.providers import AnthropicJudge, RetryPolicy
judge = AnthropicJudge(retry=RetryPolicy(attempts=5, backoff=1.0))
Every audit returns a small frozen dataclass with a flagged property, so
gating a CI pipeline is one if:
assert not position_bias(judge, cases, rubric).flagged, "judge is positionally biased"
Why the MockJudge has built-in flaws
MockJudge(position_bias=..., verbosity_bias=..., injection_bias=..., noise=..., instability=..., tie_margin=...) lets you dial in known defects.
That's how juryrig tests itself — the audits must detect a rigged judge and
clear a fair one — and it gives you a deterministic, network-free way to test
your eval pipeline end to end.
noise and instability are not the same knob, and the difference trips
people up:
noiseis seeded on the input, so re-judging one response returns the same score forever. It perturbs scores across responses.instabilityis seeded on a call counter, so the same input scores differently each time. It's the only flawself_consistency()can detect — a judge withnoise=0.9reports a spread of exactly0.0.
Instability is still reproducible: a fresh MockJudge replays the same
sequence, so switching the flaw on doesn't make your tests flaky.
tie_margin makes the judge answer "tie" when two responses score within
it. Useful for exercising tie handling — and note that without it, two
identical answers are handed to slot A by compare()'s tie-break, which the
position audit correctly reports as bias.
Calibration
from juryrig import brier_score, expected_calibration_error
scores = [0.9, 0.8, 0.3, 0.95] # judge scores
labels = [1, 1, 0, 0] # human ground truth
print(brier_score(scores, labels))
print(expected_calibration_error(scores, labels))
A judge that says 0.9 should be right ~90% of the time. ECE tells you how far that promise is from reality.
Demo
python3 examples/audit_demo.py
Runs the full audit suite against a fair judge and a rigged one, no API keys required.
Design notes
- Zero runtime dependencies — stdlib only, including the API clients.
- Provider-agnostic — a judge is anything with a
judge()method; pairwise judges addcompare(). Protocols, not base classes. - Deterministic tests — all randomness is hash-seeded; CI never flakes.
- Typed — ships a PEP 561
py.typedmarker, so the hints reach your type checker.
License
MIT
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file juryrig-0.1.0.tar.gz.
File metadata
- Download URL: juryrig-0.1.0.tar.gz
- Upload date:
- Size: 35.0 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
adcd812dd374b7519d1f4de8ca36fc7f1594328a0ae0c29a42c6f936dc8e485a
|
|
| MD5 |
5bcfd10775ad628f19c689a4806db633
|
|
| BLAKE2b-256 |
129b5e282c99f505e592ab8aa754207236df4ed9d51e595ba000bcf46a567a0c
|
Provenance
The following attestation bundles were made for juryrig-0.1.0.tar.gz:
Publisher:
publish.yml on ianalloway/juryrig
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
juryrig-0.1.0.tar.gz -
Subject digest:
adcd812dd374b7519d1f4de8ca36fc7f1594328a0ae0c29a42c6f936dc8e485a - Sigstore transparency entry: 2715064028
- Sigstore integration time:
-
Permalink:
ianalloway/juryrig@a7542fd8f44dcf209f62dddb380a799d6207a69d -
Branch / Tag:
refs/heads/main - Owner: https://github.com/ianalloway
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@a7542fd8f44dcf209f62dddb380a799d6207a69d -
Trigger Event:
workflow_dispatch
-
Statement type:
File details
Details for the file juryrig-0.1.0-py3-none-any.whl.
File metadata
- Download URL: juryrig-0.1.0-py3-none-any.whl
- Upload date:
- Size: 23.3 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
c6dc7e60053a93a02578981675cc2276e9b1adc2aa0303c03ac17ebc61f643e7
|
|
| MD5 |
c1573564dcd38fc8680b024fc985c047
|
|
| BLAKE2b-256 |
9de0b35a17497fc5641e5cfe507ff144e81aa05fc3827a4fa106f8db0ec52615
|
Provenance
The following attestation bundles were made for juryrig-0.1.0-py3-none-any.whl:
Publisher:
publish.yml on ianalloway/juryrig
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
juryrig-0.1.0-py3-none-any.whl -
Subject digest:
c6dc7e60053a93a02578981675cc2276e9b1adc2aa0303c03ac17ebc61f643e7 - Sigstore transparency entry: 2715064116
- Sigstore integration time:
-
Permalink:
ianalloway/juryrig@a7542fd8f44dcf209f62dddb380a799d6207a69d -
Branch / Tag:
refs/heads/main - Owner: https://github.com/ianalloway
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@a7542fd8f44dcf209f62dddb380a799d6207a69d -
Trigger Event:
workflow_dispatch
-
Statement type: