Last released Oct 3, 2026
Find out what your RLVR reward function actually accepts.
Diagnose why your GRPO/RLVR run is burning GPU-hours without learning.
Did your fine-tune overfit? Which checkpoint should you keep? What can you delete? Reads any HuggingFace Trainer run, local or on the Hub.