13 projects
cauterule
Automated standing-rule extraction from agent failures — extract, test, promote.
agent-self-edit
An agent that rewrites its own system prompt from execution feedback
planner-critic
Hierarchical task planning with an independent LLM critic. A planner decomposes a goal into a typed plan; a critic audits every subtask; the plan is revised until approval — or escalated to a human.
adversarial-debate
Two isolated LLM reviewers, bounded adversarial debate, preserved dissent.
mcp-fabric-toolmesh
Composable tool mesh for MCP ecosystems
agent-tooltrust
A pre-execution, deterministic policy decision point for tool-using AI agents. Evaluate every proposed tool call against declarative policy and return allow / audit / escalate / deny.
agent-eval-forge
A framework-agnostic evaluation harness for tool-using AI agents. Define scenario packs, run agents consistently, score outcomes and trajectories, and gate releases with evidence.
agent-exec-trace-analytics
Behavior analytics service for agent-exec-trace - transforms raw traces into operator-ready signals
agent-exec-trace-api
Read API service for agent-exec-trace - serves product-facing views from the normalized Postgres read model
agent-exec-trace
Execution traces for agent behavior - OpenTelemetry-style observability for AI agent workflows
mcplex-backplane
YAML-configured MCP proxy that maps simple JSON REST endpoints into MCP tools
ai-flake-sleuth
LangGraph agent that diagnoses flaky CI tests
ai-tierforge
Multi-model LLM tier router with cost-per-completed-task accounting