4 projects
limen-eval
The same-configuration noise floor of an evaluation: per-item verdict flakiness, model-pair sign-stability rulings, and a CI gate, from the repeated runs you already have.
punchmark
Point it at the response archive a benchmark number was computed on and get a producer-identity verdict -- SAME-PRODUCER / SUBSTITUTED / UNDETERMINED at a declared false-alarm rate -- with a one-line certificate attachable to the published score.
codecaliper
A cross-language code readability + complexity measurement instrument with a versioned metric-to-syntax specification
nonius
Compose a benchmark's own committed items into depth-graded composites whose gold is a deterministic function of the component golds -- and audit, for free, whether that is possible at all.