Last released Jan 19, 2025
A benchmark to evaluate implicit reasoning in LLMs using guess-the-rule games
Supported by