Last released Jun 13, 2026
A lightweight toolkit for evaluating LLM outputs: BLEU, ROUGE, semantic similarity, and LLM-as-a-Judge
Supported by