4 projects
benchspec
benchspec is a framework for evaluating AI agents with repeatable, isolated benchmarks. Write evals as Markdown, run each eval across named benchmark arms, and compare how agent behavior changes by harness, model, effort, and environment.
houserules
Natural-language linting: write your rules in English, let an LLM enforce them.
errorhandlr
UNKNOWN
BreakfastSerial
Python Framework for interacting with Arduino