5 projects
cigate
Eval-gated CI/CD for AI products: gate merges on the confidence-interval lower bound of a bias-corrected LLM-judge score, per failure-mode axis.
agenteval-py
Lightweight evaluation and observability toolkit for LLM agents
smartmemo
Semantic memory for LLM agent calls with an equivalence-first cache architecture.
guardloop
A production runtime guardrail for AI agents: budget caps, timeouts, tool limits, circuit breakers, verifier retries, and OpenTelemetry traces.
orchflow
A lightweight Python framework for readable multi-agent pipelines.