judgekeeper
Checks whether your LLM-as-judge agrees with human labels, and keeps checking: TPR, TNR and kappa against a frozen human-labeled set, run-to-run noise, position bias, drift and a CI gate.
pip install judgekeeper
# before the first PyPI release: pip install "judgekeeper @ git+https://github.com/judgekeeper/judgekeeper"
judgekeeper demo
demo validates a recorded judge on 100 bundled synthetic items: no API key, no network, a few seconds. It prints TPR, TNR and kappa and writes judgekeeper-demo/report.html. Without installing: uvx judgekeeper demo (before the PyPI release: uvx --from git+https://github.com/judgekeeper/judgekeeper judgekeeper demo).
See a real report without installing anything: LLMBar judged by Claude Haiku 4.5, also published at https://judgekeeper.github.io/judgekeeper/examples/llmbar-haiku/ (the pattern is https://<owner>.github.io/judgekeeper/examples/llmbar-haiku/).
New to judge evaluation? The website in website/ explains it step by step, with a setup guide for API keys and a hands-on tutorial. Preview it locally: python3 -m http.server --directory website 8000, then open http://localhost:8000.
Already have judge verdicts and human labels in a CSV? One command:
judgekeeper check results.csv --judge verdict --human label --out reports/my-judge/
One row per judgment. id, run, input, output and reason columns are used when present. Verdicts can be pass/fail, true/false, yes/no, correct/incorrect, 1/0 or text starting with PASS/FAIL; scores need a rule (--pass-if "score>=0.5"); other spellings need --label-map "good=pass,bad=fail". judgekeeper never guesses.
- No labels yet?
judgekeeper label items.jsonlopens a local labeling page (keys 1/2, defer, undo, notes) that writeslabels.csvas you go. Details. - Gate from pytest:
pytest --judgekeeper-report reports/my-judge/report.json, ajudgekeeper_gatefixture or a@pytest.mark.judgekeepermarker. Details. - Let a coding agent wire it in:
skills/judgekeeper/SKILL.mdtells Claude Code and other agents how, step by step.
Status: alpha (0.1.0). The validation report, CI gate and GitHub Action, pytest plugin, labeling page, judge migration, drift attribution and the promptfoo, DeepEval, Inspect AI, MLflow and Langfuse readers work.
docs/reference.mdhas every command, flag, exit code and config key;CHANGELOG.mdwhat is in 0.1.0.
Why
Teams grade AI outputs with another model, the "judge". Judges disagree with humans more than raw agreement suggests, flip verdicts between identical runs, and change silently when the provider updates the model. judgekeeper measures a judge against a frozen set of human labels, reports TPR and TNR (with 95% intervals) and kappa rather than raw agreement, and tells you, later, whether the judge is still right.
The report leads with TPR (how often the judge passes what humans passed) and TNR (how often it fails what humans failed). Either below 0.80 is "not trustworthy as a gate"; 0.80 to 0.90 is "usable with care". It also shows kappa, a confusion matrix, every disagreement with the judge's rationale, the noise floor across repeated runs ("unknown" with one run, never zero), AB/BA position bias for pairwise items, per-slice numbers, and warnings when there are fewer than 60 labels or the classes split worse than 80/20.
I have a spreadsheet
Judge verdicts and human labels already in one table: judgekeeper check above. Give it a run column (or several rows per id with different runs) to measure run-to-run noise; with one run the noise floor is reported as unknown and gate returns FLAKY. Without an id column, ids are a hash of input and output and the report says so. In a notebook:
import judgekeeper
report = judgekeeper.check_table(df, judge="verdict", human="label") # a DataFrame, a list of dicts or a path
No labels yet? Label in the local page or in Excel or Google Sheets, then freeze the labels as an anchor set:
judgekeeper label items.jsonl --out labels.csv # or: judgekeeper template items.jsonl -o labels.csv
judgekeeper import-labels labels.csv -o anchors.jsonl # checks every label, skips and lists unlabeled rows, freezes
I have a judge function
Run it 3 times over the anchor set (pairwise items in both AB and BA order) and get a report:
import judgekeeper
def my_judge(item): # item: id, input, output (never the human label)
return call_my_model(item) # bool, "PASS: ...", 0.8, (verdict, reason) or {"verdict": ..., "reason": ...}
report = judgekeeper.check_judge(my_judge, "anchors.jsonl", runs=3, fingerprint={"model": "my-model"})
Async functions work too. A judge that raises or returns nothing is recorded as an error, excluded from the metrics and counted, never scored as a fail. From the command line, a Python function or a program in any language (one item as JSON on stdin, a verdict on stdout):
judgekeeper judge anchors.jsonl --callable mypkg.judges:my_judge --runs 3 --out runs/mine/
judgekeeper judge anchors.jsonl --exec "node judge.js" --runs 3 --out runs/mine/
judgekeeper validate anchors.jsonl runs/mine/ --out reports/mine/
judgekeeper prints the number of judge calls first and asks for --yes above 1,000. The --exec contract, with 10-line Node and Python judges, is in docs/reference.md. Built-in Anthropic and OpenAI-compatible runners (--runner anthropic|openai, any OpenAI-compatible endpoint via --base-url) are documented there too. Keys stay in environment variables; nothing judgekeeper writes or prints contains one.
I use a framework
Point judgekeeper at the files your eval tool already writes and add human labels. Your eval code does not change.
promptfoo (human labels from web-UI ratings, or --labels; repeats from --repeat 3):
promptfoo eval -o results.json --repeat 3
judgekeeper import promptfoo results.json --metric helpfulness --out reports/helpfulness/
DeepEval (one test_run_*.json per run in a results folder; labels from --labels):
export DEEPEVAL_RESULTS_FOLDER=deepeval-results # then run your DeepEval tests 3 times
judgekeeper import deepeval deepeval-results/ --metric "Correctness [GEval]" --labels labels.csv --out reports/correctness/
Inspect AI (epochs are runs; labels from score edits or --labels; .eval logs need the judgekeeper[inspect] extra):
inspect eval task.py --epochs 3 --log-format json
judgekeeper import inspect logs/ --metric model_graded_qa --labels labels.csv --out reports/qa/
MLflow (judge and human assessments on your traces; each evaluation run is a run; needs the judgekeeper[mlflow] extra):
judgekeeper import mlflow --experiment my-app-eval --metric correctness --out reports/correctness/
Langfuse (judge and human scores from the public API; keys from LANGFUSE_PUBLIC_KEY and LANGFUSE_SECRET_KEY):
judgekeeper import langfuse --judge-score helpfulness --human-score helpfulness_human --from 2026-09-01 --pass-if "score>=0.5" --out reports/helpfulness/
Each writes the same report.json and report.html as check. With --anchors-out anchors.jsonl, the MLflow and Langfuse imports also freeze the labeled items as an anchor set, so you can re-judge them with judgekeeper judge for a real noise floor. How to get each file, how labels get in and each tool's traps: docs/integrations/. Any other tool can export ScoreRecords (target_id, name, annotator_kind, label, score, explanation, run, input, output, evaluator, created_at) and use judgekeeper import records, with --map for renamed columns; judgekeeper export records writes judgekeeper's runs in that format.
Keep checking
judgekeeper baseline set reports/my-judge/report.json # commit .judgekeeper/baseline.json
judgekeeper gate reports/my-judge/report.json # PASS, FAIL, FLAKY, JUDGE_CHANGED, ANCHORS_CHANGED
judgekeeper migrate anchors.jsonl runs/old/ runs/new/ --out reports/migration/
judgekeeper attribute reports/now/report.json --app-score-before X --app-score-after Y
gate never fails on a change inside the noise band. In CI, the GitHub Action (action.yml) runs judge, validate and gate, and the pytest plugin gates a report your pipeline already wrote. A judge field that is unknown on either side (imported data rarely records temperature or snapshot) is a warning, not a block, unless you pass --require-fingerprint. migrate compares an old and a new judge item by item; attribute says whether a score moved because your system changed or the judge did. Details: docs/reference.md.
First results
The judge is strong on ordinary items and weak on adversarial ones: claude-haiku-4-5-20251001 reaches kappa 0.94 against human labels on LLMBar's Natural slice but 0.56 on Adversarial/GPTOut (report, data in docs/examples/llmbar-haiku/).
The run: LLMBar (Natural and Adversarial subsets, 419 human-labeled pairwise items) judged by claude-haiku-4-5-20251001 at temperature 0, 3 runs, AB and BA. scripts/llmbar_haiku_demo.sh runs it end to end and writes the report to docs/examples/llmbar-haiku/.
Run on 2026-10-01 against api.anthropic.com. Every number below is quoted from docs/examples/llmbar-haiku/report.json (human-readable version: report.html in the same folder). TPR and TNR treat human label A as the positive class.
Verdict: Usable as a gate: kappa 0.83, TPR 0.95, TNR 0.88 against human labels.
| Run | Kappa | TPR | TNR |
|---|---|---|---|
| 1 | 0.82 | 0.95 | 0.87 |
| 2 | 0.85 | 0.96 | 0.89 |
| 3 | 0.82 | 0.95 | 0.87 |
| Mean | 0.83 | 0.95 | 0.88 |
By slice (mean over runs):
| Slice | Items | Kappa | TPR | TNR |
|---|---|---|---|---|
| Natural | 100 | 0.94 | 0.98 | 0.97 |
| Adversarial/GPTInst | 92 | 0.92 | 1.00 | 0.92 |
| Adversarial/Neighbor | 134 | 0.84 | 0.96 | 0.88 |
| Adversarial/Manual | 46 | 0.64 | 0.91 | 0.74 |
| Adversarial/GPTOut | 47 | 0.56 | 0.86 | 0.71 |
- Noise floor: mean pairwise kappa between runs 0.95. 3.6% of items changed verdict in at least one run.
- Position bias: 9.9% of judgments changed when the two outputs were swapped (AB vs BA). Kappa on the BA order alone was 0.80.
- The judge's TNR is lower than its TPR. Its errors cluster in the hardest adversarial slices (GPTOut, Manual), where kappa drops to 0.56 and 0.64.
Why this and not X
- Vendor "align evals" features (LangSmith, Arize, MLflow, Ragas, Confident AI) do one-off calibration. judgekeeper covers what happens after: time, change, and gating.
- RAND's Judge Reliability Harness generates bias probes but has no human-kappa, no time series, no CI, and only supports OpenAI.
- Several zero-star scripts from Aug to Sep 2026 sketch anchor-set drift checks. judgekeeper aims to be the maintained, multi-provider, framework-integrated version with published numbers.
Full landscape and sources: docs/research/00-judge-reliability-synthesis.md.
Roadmap
- Core validation report on a public human-labeled dataset (done; LLMBar above).
- GitHub Action and noise-floor gate (done).
- Judge migration and drift attribution (built; the live migration demo is pending a key).
- Use without a framework:
check,check_table,check_judge,--callable,--exec,demo, spreadsheet labels (done). - Readers for promptfoo, DeepEval and Inspect AI files, and the ScoreRecord import/export format (done). MLflow and Langfuse readers, with re-judging of imported labels (done).
- Launch readiness: labeling page, pytest plugin, agent skill, release and Pages workflows (done; publishing is the maintainer's manual step,
docs/RELEASING.md). - Demo agent (Claude Agent SDK), clarification (ask-vs-act) judge, MCP server.
License
MIT.
Metadata
Release files for judgekeeper 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| judgekeeper-0.1.0.tar.gz | 557.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| judgekeeper-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 682.0 kB
Release files / judgekeeper-0.1.0.tar.gz
| Download URL | judgekeeper-0.1.0.tar.gz |
|---|---|
| Size | 557.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
31e30041e3fc92af7e5f559fa2e7b36b8952adb1eafc734751a0a04449fa65ae
|
|
BLAKE2b-256 checksum How to use checksums |
8ad0fbf8a6a9165d3b1c7b7f8fb48dd350a7a359f4ecf7915d3f1fc4d6eae563
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.
Transparency logRelease files / judgekeeper-0.1.0-py3-none-any.whl
| Download URL | judgekeeper-0.1.0-py3-none-any.whl |
|---|---|
| Size | 124.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
f0c7081a88b7303e8268e1e7b7a533785b7734c0c6d467bd5184866a58638d0e
|
|
BLAKE2b-256 checksum How to use checksums |
28f8f2461782ede9c7364038f04aefa3fcb41915f7afae5f3e97557782c766e9
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.
Transparency log