Last released Jul 23, 2026
Zero-cost LLM evaluation toolkit: score model outputs against rubrics using a local LLM-as-judge, and track consistency and drift over time.