Last released Jul 16, 2026
Build a retrieval eval for your own corpus: pool candidates, LLM-judge them under a rubric you own, and measure judge-vs-human alignment before trusting any metric.
Supported by