Dependency-free statistical checks for AI eval claims (multiple-comparisons, judge-bias, resolution, fragility) — as a CLI and an MCP server an agent calls before trusting a benchmark number.
Project description
evalgate
Cheap statistical checks for AI eval claims — run them before you publish.
Most benchmark headlines overstate themselves in one of a few nameable ways. evalgate is four tiny, dependency-free checks, one per failure mode — the same checks behind a set of independent eval-integrity audits that caught these mistakes in published work.
Pure Python, zero dependencies, runs anywhere.
pip install git+https://github.com/ipezygj/evalgate
Use it as an MCP tool (for agents)
If you're an AI agent — or you run one (Claude, Cursor, Claude Code, Windsurf…) — evalgate ships an MCP server so the model can call these checks itself before it trusts, reports, or acts on any eval number: a benchmark score, a leaderboard #1, an LLM-as-judge verdict, or a claimed trend.
pip install "eval-gate[mcp] @ git+https://github.com/ipezygj/evalgate" # once on PyPI: pip install "eval-gate[mcp]"
Add it to your MCP client (e.g. claude_desktop_config.json / Cursor / Claude Code):
{
"mcpServers": {
"evalgate": { "command": "evalgate-mcp" }
}
}
Tools the agent gets, each with a "call this when…" description it can reason about:
| tool | the agent calls it before… |
|---|---|
check_top_rank |
claiming a model is #1 / SOTA on a benchmark (is the top rank real or a tie?) |
check_subset_win |
trusting a "best on subset/metric X" claim (look-elsewhere correction) |
check_judge_bias |
trusting an LLM-as-judge / A-B preference result (length / self-preference / position bias) |
check_resolution |
calling one of two close models better (can the benchmark even tell them apart?) |
check_trend_fragility |
reporting a fitted trend / scaling exponent (does one point carry it?) |
The five checks above work on summary numbers (scores, p-values, win counts). If the agent has the raw per-item results — which items each model solved, or the head-to-head battles — three deeper tools do the real audit instead of an approximation:
| tool | the agent calls it when… |
|---|---|
audit_leaderboard |
it has per-item results ({model: [solved item-ids]}) — the real version of check_top_rank: bootstrapped rank confidence intervals, the paired-McNemar tie group, resolvable tiers, split-half stability |
audit_preferences |
a ranking comes from pairwise votes ([winner, loser] battles) — Bradley-Terry rank CIs and a Condorcet check that preferences are transitive, not rock-paper-scissors cycles |
check_dimensions |
deciding whether one number fairly summarizes a multi-skill benchmark — counts the latent skills in the result matrix (eigenspectrum vs a shuffled null) |
The point: an agent that produces an eval number should sanity-check it, and now it can — in one
call, with a plain verdict and a recommendation. Reproducible, zero-dependency checks (the mcp
extra is only for the server transport).
The four checks
1. "We lead on subset X" — corrected for look-elsewhere
Report the subset/metric/checkpoint where a model looks best and you are reporting the maximum of many noisy tests. Correct for how many you could have picked.
evalgate correct --p 0.009 --n 23
# raw p=0.009 over 23 tests -> sidak p=0.19 (does NOT survive correction at alpha=0.05)
from evalgate import correct_best_of
correct_best_of(0.009, n_tested=23).significant # False
(A real RewardBench "best subset" win: raw p=0.009 → p=0.19 after correcting for the 23 subsets. Not a finding.)
2. Is the judge winning, or just longer / first / same-family?
An LLM-as-judge that "prefers" your model may be preferring the longer answer, the first-listed one, or its own family. Feed it the count and test against chance.
evalgate bias --wins 68 --n 100 --label "longer answer wins"
# longer answer wins: 68/100 = 68.0% (p=0.0004) -> BIAS
from evalgate import bias_rate
bias_rate(68, 100).biased # True
(A widely-used GPT-4 judge preferred the longer answer 68% of the time and its own model family 71.5% — both at p≈0.)
3. Does one data point flip your slope?
A scaling exponent or trend that hangs on a single high-leverage point isn't one. Leave each point out and refit.
evalgate loo examples/points.txt --power-law --threshold 1.0
# slope=1.08, leave-one-out range [0.87, 1.26] -> CROSSES 1 (most influential point: index 5)
from evalgate import leave_one_out, power_law_exponent
leave_one_out(xs, ys, fit=power_law_exponent, threshold=1.0).crosses_threshold # True
(A reported "super-linear" grokking exponent, α=1.13, fell to 0.97 — with a better fit — when one point was dropped.)
4. Is the gap bigger than the sample can resolve?
A leaderboard orders two models by a two-point accuracy gap on a finite test set. Ask whether that gap is even detectable at this sample size — or smaller than the minimum detectable effect, i.e. a coin flip.
evalgate power --n 200 --p1 0.85 --p2 0.83
# gap=+0.02 on n=200 (NOT significant, p=0.585); MDE at 80% power=0.103 -> UNDERPOWERED (gap < MDE)
from evalgate import power_check
power_check(200, 0.85, 0.83).resolvable # False — 2pp over 200 items can't be resolved
power_check(2000, 0.85, 0.80).resolvable # True — 5pp over 2000 items can
(Frontier models on a fixed benchmark routinely sit a task or two apart — inside the MDE — so the #1 rank is noise. More votes, not a better model, resolves the tie.)
Library API
from evalgate import (
correct_best_of, sidak, bonferroni, # look-elsewhere
bias_rate, binomial_test, # judge / metric bias
leave_one_out, ols_slope, power_law_exponent, # fragility + fits
power_check, min_detectable_effect, # power / minimum detectable effect
)
Every function returns a small dataclass that prints a one-line verdict and exposes the numbers (.corrected_p, .p_value, .loo_min …) so you can gate CI on them.
Reproduce the case studies:
python -m evalgate.checks # -> evalgate selftest: OK (reproduced all 3 case studies + power check)
Why this exists
These are textbook checks — the value isn't the math, it's running all of them, adversarially, on a number you're too close to. evalgate is the open, do-it-yourself version. When a launch, a paper, or a fundraise rides on a figure and you want it audited independently first, that's the paid practice.
Want the full checklist and the client-grade report template that wrap these checks? The Eval Integrity Kit — the 9-check audit checklist, the report template I ship to clients, an evalgate quickstart, and three worked case studies.
The fuller story — why AI benchmark scores and trading backtests overpromise, and how to catch them — is in the book Measured, Not Believed (pay what you want).
License
MIT — see LICENSE.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file eval_integrity-0.4.0.tar.gz.
File metadata
- Download URL: eval_integrity-0.4.0.tar.gz
- Upload date:
- Size: 23.7 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.14.6
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
0625639c65b62959dbe5785d3f02fec3381770386166ac8896c1a866e4770b85
|
|
| MD5 |
4dc2eab2ed15d6c6f92ea46c4b9c5b96
|
|
| BLAKE2b-256 |
64bc3c534eef3f5e256ae025cf25f6b56010ce56d29c7a1c9b316a93e38d2105
|
File details
Details for the file eval_integrity-0.4.0-py3-none-any.whl.
File metadata
- Download URL: eval_integrity-0.4.0-py3-none-any.whl
- Upload date:
- Size: 24.4 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.14.6
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
1c06e79cebb055db3926d74928273017e89b26f7984e5a1d84927d2cabb5a063
|
|
| MD5 |
470a9d3d067587c369d7eed1f699aeea
|
|
| BLAKE2b-256 |
8354bb18795bd937c4ab76f54eadd243a1bf13ac6c43324bb65f23dcde55f8b0
|