Skip to main content

LLM Council MCP

CI License: MIT MCP Badge

A hybrid of Andrej Karpathy's LLM Council and the Model Context Protocol: real multi-model deliberation over OpenRouter, exposed as an MCP server you call from inside Claude Code (or any MCP client).

Unlike the single-model "5 sub-agents" Claude Code skill, this runs genuine cross-model deliberation — GPT, Claude, Gemini, Grok, etc. all answer, peer-review each other's anonymized responses, and a chairman model synthesizes the verdict.

The pipeline

  1. Stage 1 — First opinions. Every council model answers your question independently (parallel).
  2. Stage 2 — Anonymized peer review. Each model sees the others' responses as "Response A/B/C…" (identities hidden so no model favors its own family) and ranks them.
  3. Stage 3 — Chairman synthesis. A designated chairman model reads all responses + rankings and produces the final answer.

You also get a peer leaderboard (average rank per model) computed from the parsed rankings.

Tools exposed

Tool What it does
council_deliberate Full 3-stage council on a hard question. Returns a markdown report (or format="json"). Optional models / chairman_model overrides, and html_path to also write a standalone HTML report ("auto" to auto-name it).
council_deliberate_streaming Same as above but emits live MCP progress + log events as each stage completes (dispatch → peer review → chairman), so the client shows a status bar. Also supports html_path.
council_jury Fast go/no-go: each model gives VERDICT: YES/NO, returns a vote tally + chairman synthesis.
council_config Shows the active roster, chairman, and whether the API key is set.

HTML reports

Pass html_path to council_deliberate / council_deliberate_streaming to also write a self-contained, shareable HTML report (dark-themed, with the final verdict, peer leaderboard, and collapsible per-model reviews). Use "auto" to drop a timestamped council-report-<ts>.html in the working directory.

Prerequisites

  • Python ≥ 3.10
  • An OpenRouter API key with credits (each council run hits N models + 1 chairman, so ≈ N+1× the tokens of a single query).

Prompts

Reusable templates your client can invoke directly (/mcp__llm-council__<name> in Claude Code).

Prompt Arguments What it does
deliberate question, context (optional) Frames a hard decision for full 3-stage deliberation and asks for the disagreements, not just the verdict.
jury question, stakes (optional) Frames a go/no-go so the tally leads and dissenters are named.
compare_options options, criteria (optional) Compares named options and forces a recommendation plus its strongest counterargument.

Resources

Read-only context you can attach to a conversation.

Resource Contents
council://roster Active roster, chairman, timeout/retry settings, and whether the API key is set (JSON).
council://methodology How the 3-stage protocol works and how to read the peer leaderboard.

Install

Recommended — no clone, no venv

Requires uv. uvx fetches and runs the server in a throwaway environment:

uvx mcp-llm-council

From PyPI

pip install mcp-llm-council

From source

git clone https://github.com/JeremyGracey-AI/llm-council-mcp
cd llm-council-mcp
python3 -m venv .venv
.venv/bin/pip install -e .

All three give you the mcp-llm-council console script (entry point llm_council_mcp.server:main). The old llm-council-mcp script name still works, so existing config keeps running.

Note on the package name. The PyPI name llm-council-mcp belongs to an unrelated project, so this one publishes as mcp-llm-council. The import path is still llm_council_mcp and the GitHub repo is unchanged.

Register with Claude Code

Claude Code reads MCP servers from ~/.claude.json (or a project .mcp.json). Easiest way:

claude mcp add llm-council \
  --env OPENROUTER_API_KEY=sk-or-v1-... \
  -- uvx mcp-llm-council

Or add it manually to ~/.claude.json:

{
  "mcpServers": {
    "llm-council": {
      "command": "uvx",
      "args": ["mcp-llm-council"],
      "env": {
        "OPENROUTER_API_KEY": "sk-or-v1-...",
        "COUNCIL_MODELS": "openai/gpt-5.1,google/gemini-3-pro-preview,anthropic/claude-sonnet-4.5,x-ai/grok-4",
        "CHAIRMAN_MODEL": "google/gemini-3-pro-preview"
      }
    }
  }
}

Running from a source checkout instead? Point command at the script in your venv:

{
  "mcpServers": {
    "llm-council": {
      "command": "/ABSOLUTE/PATH/llm-council-mcp/.venv/bin/mcp-llm-council",
      "env": { "OPENROUTER_API_KEY": "sk-or-v1-..." }
    }
  }
}

Restart Claude Code. Then just ask it to use the tools, e.g.:

Use the llm-council council_deliberate tool: should Remy v1 stay a single-agent reasoning loop or move to a multi-agent orchestrator before my Anthropic Fellowship demo?

Usage example

A typical council_deliberate call returns a markdown report shaped like this:

# LLM Council Verdict

**Question:** Should I cache embeddings in SQLite or Redis for a single-box demo?

## Final Answer (Chairman: google/gemini-3-pro-preview)
For a single-box demo, SQLite is the better default: zero extra services to run,
persistence for free, and ample throughput at demo scale. Reach for Redis only if
you later need sub-millisecond reads under concurrency or cross-process sharing.

## Peer Leaderboard (lower avg rank = better)
- **anthropic/claude-sonnet-4.5** — avg rank 1.33 (ranked by 4)
- **openai/gpt-5.1** — avg rank 1.67 (ranked by 4)
- **google/gemini-3-pro-preview** — avg rank 3.0 (ranked by 4)
- **x-ai/grok-4** — avg rank 4.0 (ranked by 4)

## Stage 1 — Individual Responses
### openai/gpt-5.1
...
## Stage 2 — Peer Reviews & Rankings
### anthropic/claude-sonnet-4.5
...
FINAL RANKING:
1. Response B
2. Response A
...

Fast go/no-go decisions use council_jury instead — each model returns VERDICT: YES/NO and you get a tally plus the chairman's synthesis:

# Council Jury Verdict

**Question:** Should we ship the v1 demo this Friday?

**Tally:** 3 YES / 1 NO
- openai/gpt-5.1: YES
- google/gemini-3-pro-preview: YES
- anthropic/claude-sonnet-4.5: YES
- x-ai/grok-4: NO

## Chairman Synthesis (google/gemini-3-pro-preview)
Ship it — three of four advisors agree the core path is solid. The lone NO flags
thin error handling on the upload route; gate the Friday release on that one fix.

Inspect the active roster any time with council_config (no API call, no cost).

Configuration (env vars)

Var Default Notes
OPENROUTER_API_KEY — Required.
COUNCIL_MODELS openai/gpt-5.1,google/gemini-3-pro-preview,anthropic/claude-sonnet-4.5,x-ai/grok-4 Comma-separated OpenRouter model IDs.
CHAIRMAN_MODEL google/gemini-3-pro-preview Synthesizer.
LLM_COUNCIL_TIMEOUT 120 Per-request seconds.
LLM_COUNCIL_MAX_RETRIES 2 Retries on 408/429/5xx.
OPENROUTER_API_URL https://openrouter.ai/api/v1/chat/completions Override the OpenRouter chat-completions endpoint (e.g. a proxy).
OPENROUTER_REFERER https://github.com/ HTTP-Referer attribution header sent to OpenRouter.
OPENROUTER_TITLE llm-council-mcp X-Title attribution header sent to OpenRouter.

Run the server directly (debug)

OPENROUTER_API_KEY=sk-or-v1-... .venv/bin/mcp-llm-council
# speaks MCP over stdio — Ctrl-C to exit

Test offline (no API cost)

PYTHONPATH=. .venv/bin/python tests/test_pipeline_mock.py   # pipeline, streaming & HTML, OpenRouter mocked
PYTHONPATH=. .venv/bin/python tests/test_mcp_boot.py        # boots server, lists all 4 tools
PYTHONPATH=. .venv/bin/python tests/test_mcp_surface.py     # asserts every tool, prompt & resource

Releasing / publishing

Pushing a v* tag triggers .github/workflows/publish.yml, which builds the sdist + wheel and publishes to PyPI via Trusted Publishing (OIDC — no stored tokens). One-time PyPI setup: on the mcp-llm-council PyPI project, add a trusted publisher for repo JeremyGracey-AI/llm-council-mcp, workflow publish.yml, environment pypi.

git tag v0.2.0
git push origin v0.2.0

Contributing

Contributions are welcome. The project is small and the test suite runs offline (no API key or credits needed), so the loop is fast:

git clone https://github.com/JeremyGracey-AI/llm-council-mcp.git
cd llm-council-mcp
python3 -m venv .venv && .venv/bin/pip install -e .
PYTHONPATH=. .venv/bin/python tests/test_pipeline_mock.py
PYTHONPATH=. .venv/bin/python tests/test_mcp_boot.py

Guidelines:

  • Open an issue first for anything beyond a small fix, so we can align on approach.
  • Branch off main, keep PRs focused, and make sure both tests pass — CI runs them on Python 3.10–3.12.
  • Add or update a test for any behavior change. Keep tests offline by mocking OpenRouter (see tests/test_pipeline_mock.py); never commit real API keys or hit the live API in tests.
  • Match the existing style: type hints, docstrings on public functions, and structured {ok, error} results rather than raised exceptions in the request path.
  • New tools belong in server.py; core pipeline logic in council.py; provider/transport details in openrouter.py.

By contributing you agree your contributions are licensed under the MIT License.

Credits

3-stage methodology and prompts adapted from karpathy/llm-council (MIT). MCP wrapper, retries, jury mode, and leaderboard added here.

Metadata

Release files for mcp-llm-council 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for mcp-llm-council 0.2.0
File Size Uploaded
mcp_llm_council-0.2.0.tar.gz 18.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for mcp-llm-council 0.2.0
File Interpreter ABI Platform
mcp_llm_council-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 36.2 kB

Release files / mcp_llm_council-0.2.0.tar.gz

Download URL mcp_llm_council-0.2.0.tar.gz
Size 18.0 kB
Tags Source
SHA-256 checksum
How to use checksums
72d261e91a84cd5e2d383f8a7457ee168a7ae8800681f5c2de87df5b914100d2
BLAKE2b-256 checksum
How to use checksums
df880a88e499a4a452c8a99ad7dc0c8de65051647b9965937e9ddaa456cd3e32
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 14, 2026.

Transparency log

Release files / mcp_llm_council-0.2.0-py3-none-any.whl

Download URL mcp_llm_council-0.2.0-py3-none-any.whl
Size 18.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
f6f0adeb4ed6ffa23e1c9028c1752c97538268169e63be0a90c81813d5fcb6a6
BLAKE2b-256 checksum
How to use checksums
4e595bbe6dca630bc586518420d05378b84089d52323d3e6efa62f705a896084
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 14, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.2.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page