Last released Feb 7, 2026
A benchmarking toolkit for evaluating paper revision quality using LLM-as-a-judge
Supported by