Skip to main content

AgentKit — A Governed MCP Tool Server

CI License: AGPL v3

An MCP server where tools are declarative, effects are typed, and every action is policy-gated and audited — usable by any MCP client (Claude Desktop, Cursor, LangGraph, Claude Agent SDK, CrewAI).

Three things distinguish it from a typical MCP server:

  • Declarative tools — define tools in YAML over your own Postgres or HTTP API. No Python, no fork. (docs/REUSE.md)
  • Real actions, not just reads — tools declare an effect (read / write / destructive) and mutating tools genuinely mutate.
  • Guardrails that hold regardless of the prompt — writes are off by default, destructive actions need a human-held approval token the model never sees, everything supports dry-run, and every call (allowed and denied) is audited. (SECURITY.md)

The bundled business-intelligence tools below are the reference pack that demonstrates all of this — not the limit of what the server does.

🔗 Self-hosting: see SELF_HOSTING.md to run your own instance.

What It Does

Reference BI pack (built in):

  • 6 MCP Tools: query_kpis, get_company_health, detect_kpi_anomalies, forecast_metric, list_available_metrics, get_executive_summary
  • 6 MCP Resources: kpi://Finance/latest and similar for Growth, Operations, People, ESG, IT_Ops
  • 1 Reusable Prompt: monthly_executive_briefing

These come from the reference pack — the core server ships with no hardcoded resources or prompts. You can add your own @mcp.resource / @mcp.prompt decorators, or load them from a tool pack. See docs/REUSE.md.

Platform capabilities:

  • Declarative tool packs — add tools over your own Postgres/HTTP in YAML (packs/)
  • Typed effects + policy engineGET /api/policy publishes the capability envelope
  • Audit trailGET /api/audit, allowed and denied, with deny reasons
  • Multi-provider LLM routing incl. self-hostedGET /api/llm-routing
  • LangGraph 3-agent workflow in workflow.py (Planner → Analyst → Reporter)
  • Claude Agent SDK demo in demos/claude_agent_sdk_demo.py
  • CrewAI demo in demos/crewai_demo.py
  • DSPy research scaffold in research/dspy_experiment.py
  • 34 tests across smoke, API, integration, and LangGraph workflow

PyPI Package

pip install agentkit-mcp   # v0.1.9
agentkit-mcp               # CLI entrypoint

Quick Start

python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
cp .env.example .env  # fill in keys + POSTGRES_URL
python mcp_server.py

Claude Desktop Setup

Add to ~/.config/Claude/claude_desktop_config.json:

{
  "mcpServers": {
    "agentkit": {
      "command": "python",
      "args": ["/abs/path/to/agentkit/mcp_server.py"],
      "env": {
        "MCP_TRANSPORT": "stdio",
        "POSTGRES_URL": "postgresql://...",
        "LOG_LEVEL": "DEBUG",
        "TELEMETRY_OPT_OUT": "true"
      }
    }
  }
}

MCP_TRANSPORT=stdio is required here — without it mcp_server.py defaults to serving over SSE (a network port) instead of talking JSON-RPC over the pipes Claude Desktop spawns it with, and no tools will appear. Local stdio mode doesn't need MCP_AUTH_TOKEN (the OS process boundary is the auth boundary); that variable only matters for the SSE/network path.

Multi-Provider LLM Routing

The 3-agent LangGraph workflow (workflow.py) and the demos/research scripts route each role to its own model via LiteLLM, configured with plain provider/model strings — no code changes to switch providers:

  • LLM_REASONING — planner + reporter agents (defaults to anthropic/claude-sonnet-4-6)
  • LLM_DEFAULT — the tool-calling analyst agent (defaults to groq/openai/gpt-oss-120b)
  • LLM_JUDGE — used by the eval suite (defaults to anthropic/claude-haiku-4-5)
  • LLM_LOCAL + INFERENCE_MODE=local — route to a local/self-hosted model (e.g. Ollama) instead of a hosted provider

Set the matching provider API key(s) (GROQ_API_KEY, ANTHROPIC_API_KEY, OPENAI_API_KEY) for whichever models you reference above. See .env.example.

  • Diagnostics: adjust LOG_LEVEL to DEBUG for verbose logs.
  • Telemetry: an anonymous startup ping is sent by default; disable with TELEMETRY_OPT_OUT=true.

Restart Claude Desktop, then ask:

  • "What's our company health right now?"
  • "Forecast revenue for the next 6 months."
  • "Are there anomalies in the Finance KPIs?"

LangGraph Workflow

from agentkit_mcp.workflow import analyze
result = analyze("What drove gross margin in Q1?")
print(result["report"])

Architecture

        Claude Desktop / Cursor / LangGraph
                      │
                      ▼ MCP
              ┌──────────────────┐
              │  mcp_server.py   │
              │   6 tools        │
              │   6 resources    │
              │   1 prompt       │
              └────────┬─────────┘
                       │
        ┌──────────────┼──────────────┐
        ▼              ▼              ▼
   pg_store      insights      forecasting
   (KPIs)        (health,      (LinearReg
                 anomalies)    + Monte Carlo)

Research Novelty & Scientific Contributions

AgentKit is an industry-proof intelligence engine:

  • Standardized Model Context Protocol (MCP) Middleware: Unified stdio and SSE transport for hot-swappable agent tools.
  • Capability Policy Engine: Formal effect separation (read/write/destructive) and prompt-independent guardrails.
  • Multi-Agent Interoperability: Tested and verified across Claude Desktop, Cursor IDE, and Devin AI.

For full theoretical formulation, math bounds, and citation details, see RESEARCH.md.

Benchmark Replication Suite

Run the reproducible benchmark evaluation suites:

# Test MCP framework overhead
python3 eval/run_benchmarks.py --seed 42

# Test Agent Tool Selection & Quality
python3 eval/run_agent_eval.py

# Test Comprehensive MCP Tool Execution Metrics
python3 eval/run_mcp_tools_benchmark.py

Integration Guides (Claude Desktop, Cursor, Devin)

Automated client verification:

python3 tests/test_mcp_client.py

License & Enterprise Use (Dual-License)

This project is open-source under the AGPL-3.0 License. It is completely free for researchers, students, and open-source hobbyists. Commercial license: see COMMERCIAL.md.

Release files for agentkit-mcp 0.1.11

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for agentkit-mcp 0.1.11
File Size Uploaded
agentkit_mcp-0.1.11.tar.gz 73.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for agentkit-mcp 0.1.11
File Interpreter ABI Platform
agentkit_mcp-0.1.11-py3-none-any.whl Python 3 none any Details

Total release size: 140.7 kB

Release files / agentkit_mcp-0.1.11.tar.gz

Download URL agentkit_mcp-0.1.11.tar.gz
Size 73.3 kB
Tags Source
SHA-256 checksum
How to use checksums
ca28b857210d89995739a2d70e531b373a841bfaf9bf64d9a076f6722a91c7e4
BLAKE2b-256 checksum
How to use checksums
983671281f971801e35838c7752e7f545c915f141a68726230a991b40f8306df
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.3

Release files / agentkit_mcp-0.1.11-py3-none-any.whl

Download URL agentkit_mcp-0.1.11-py3-none-any.whl
Size 67.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
396b9c7d45608698c1e8b28b7094d88d437bc730af2ae248cb55c1aef3e8b03d
BLAKE2b-256 checksum
How to use checksums
7f4cd07e4bcc6dd4d31b22d219d9fd1e6f861155b75d0b33a614a490fe7bb38e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.3

Release history Release notifications | RSS feed

0.1.14

2 release files

0.1.13

2 release files

0.1.12

2 release files

This release

0.1.11 This release

2 release files

0.1.10

2 release files

0.1.9

2 release files

0.1.8

2 release files

0.1.7

2 release files

0.1.6

2 release files

0.1.5

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

1 release file

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page