Last released Apr 26, 2026
Turn LLM failure hypotheses into minimum discriminating regression benchmarks.
Supported by