3 projects
evalpower
How many runs before your eval means anything? Repeat-run reliability statistics for stochastic evals: audit miss rates, exact Clopper-Pearson stochasticity, runs-needed, and pass^k / safe^k curves from per-cell run counts.
agentrelbench
k-run wrapper for the EnterpriseOps-Gym benchmark: drives evaluate.py k times per task, archiving per-run DB state exports plus a batch manifest.
sqlpup
A small text-to-SQL language model, trained from scratch.