Last released Jul 27, 2026
Statistically valid evaluation for LLM-judged benchmarks: bias-corrected scores, honest intervals, defensible comparisons
Reference semantics, conformance testing, and pathology diagnostics for LLM policy-gradient post-training