9 projects
rankkit
Ranking evaluation with error bars: NDCG, MRR and MAP with confidence intervals, plus position-bias correction for click logs
arenakit
Audit a pairwise model leaderboard before you trust its order
raterkit
Audit a labeled dataset before you trust it
abkit
Audit an A/B-test readout before you ship the decision
judgepanel
Estimate LLM-judge accuracy without gold labels and aggregate judge panels: Dawid-Skene EM, agreement statistics, bootstrap uncertainty
calikit
Calibration auditing for probabilistic predictions: reliability diagrams, ECE, Brier decomposition, and temperature scaling
abeval
A/B-test statistics for LLM evals: error bars, paired comparisons, and sample-size planning
judgekit
Audit an LLM judge before you trust it: agreement, bias probes, calibration and consistency, with bootstrap confidence intervals.
trajectory-judge
Measuring what LLM judges miss when an agent reaches the right answer the wrong way