Last released Aug 15, 2026
Rank LLMs on a task suite, and measure whether the automated judge can be trusted.
Supported by