Last released Oct 5, 2026
Evaluator toolkit for LLM outputs: pluggable LLM-as-a-judge with structured verdicts