An agentic evaluation harness for MCP servers — can an AI agent actually accomplish real tasks with this server's tools?
Project description
mcp-gauntlet
An agentic evaluation harness for MCP servers. Point it at any Model Context Protocol server and it answers the question the static analyzers don't: can an AI agent actually accomplish real tasks using this server's tools?
Why
The existing MCP quality tools (mcp-lighthouse, mcp-scorecard, mcp-checkup)
are all static — they inspect schemas, count tokens, and lint descriptions.
None of them run an LLM agent against the server to see whether it can actually
complete tasks. That dynamic, agent-in-the-loop evaluation — with a real
task-success rate, not just a conformance check — is what mcp-gauntlet does.
Google Lighthouse tells you your web page is well-formed. mcp-gauntlet tells you your MCP server is usable by an agent — with a task-success rate to prove it.
What it scores
Each run produces a graded report card (JSON + Markdown) across:
- Schema Health — valid JSON schemas, typed and described parameters.
- Description Quality — can an agent tell when and how to use each tool?
- Security Signals — tool-poisoning / prompt-injection markers and hidden characters; a critical finding caps the overall grade.
- Agent Task Success — a live LLM agent attempts generated tasks using only the server's tools; LLM-judged and repeated for a success rate.
- Tool-Selection Accuracy — did the agent call the tools it was expected to?
- Tool Reliability — did the server's tools execute without error?
- Robustness — does the server reject malformed input gracefully?
Leaderboard
A live leaderboard ranks popular public MCP servers by their gauntlet score: ghalebdweikat.github.io/mcp-gauntlet
Generate one yourself across any set of servers listed in a JSON file:
uv run mcp-gauntlet leaderboard --servers leaderboard.servers.json --out docs
Quickstart
uv sync --extra dev
# Static + robustness checks only — no API key required
uv run mcp-gauntlet run "python -m mcp_gauntlet.fixtures.good_server" --no-agentic
# Full gauntlet, including the live agent (Groq's free tier works)
echo "GROQ_API_KEY=gsk_..." > .env
uv run mcp-gauntlet run "npx -y @modelcontextprotocol/server-everything"
The LLM backend is provider-agnostic — any OpenAI-compatible endpoint (Groq by
default; also OpenRouter, Together, or a local Ollama / vLLM). Runs are safe by
default: only read-only tools are exercised unless you pass --allow-writes, and
generated task sets are cached so scores are reproducible across runs.
Bundled good / bad fixture servers make it easy to see the difference:
uv run mcp-gauntlet run "python -m mcp_gauntlet.fixtures.bad_server" # capped C — tool poisoning
uv run mcp-gauntlet run "python -m mcp_gauntlet.fixtures.good_server" # A
Configuration
Configure via a .env file (copy .env.example and fill it in)
or real environment variables:
| Variable | Purpose |
|---|---|
GROQ_API_KEY / GEMINI_API_KEY / OPENAI_API_KEY / OPENROUTER_API_KEY |
API key for the provider the agent should use (only one needed). A free Groq key: console.groq.com/keys. |
MCP_GAUNTLET_PROVIDER |
Which provider: groq (default), gemini, openai, openrouter, or ollama (local). |
MCP_GAUNTLET_MODEL |
Model override for that provider (e.g. gemini-flash-latest). Defaults to a sensible per-provider model. |
The --provider / --model CLI flags override these. The backend is any
OpenAI-compatible endpoint, so the same setup covers cloud providers and a local
Ollama / vLLM.
No API key? Static mode
Everything except the live agent runs without an LLM. mcp-gauntlet run <server>
with no key configured reports a static grade from the LLM-free checks —
schema health, description quality, security signals, and robustness probes:
uv run mcp-gauntlet run "npx -y @modelcontextprotocol/server-everything" --no-agentic
Add --no-probe for a pure inspection that never executes any of the server's
tools. The leaderboard behaves the same way — with no key it ranks servers on the
static + robustness checks alone.
License
MIT © Ghaleb Dweikat
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file mcp_gauntlet-0.1.0.tar.gz.
File metadata
- Download URL: mcp_gauntlet-0.1.0.tar.gz
- Upload date:
- Size: 115.9 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
f6ebb19bd20deeef79672f9a32d400c1f08632547f23c4bba5d9aff6c9f0a459
|
|
| MD5 |
3da4b699731959da9d095e8a86adbe30
|
|
| BLAKE2b-256 |
02ae0f38a6d6c681c1771ecefdfd04985ef62db5cd8c5b0c71e6210348957515
|
Provenance
The following attestation bundles were made for mcp_gauntlet-0.1.0.tar.gz:
Publisher:
publish.yml on GhalebDweikat/mcp-gauntlet
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
mcp_gauntlet-0.1.0.tar.gz -
Subject digest:
f6ebb19bd20deeef79672f9a32d400c1f08632547f23c4bba5d9aff6c9f0a459 - Sigstore transparency entry: 2229561796
- Sigstore integration time:
-
Permalink:
GhalebDweikat/mcp-gauntlet@497b42a0b2d8a1b7e8d99d3a281cdbe13e099391 -
Branch / Tag:
refs/tags/v0.1.0 - Owner: https://github.com/GhalebDweikat
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@497b42a0b2d8a1b7e8d99d3a281cdbe13e099391 -
Trigger Event:
release
-
Statement type:
File details
Details for the file mcp_gauntlet-0.1.0-py3-none-any.whl.
File metadata
- Download URL: mcp_gauntlet-0.1.0-py3-none-any.whl
- Upload date:
- Size: 41.4 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
c9ff5db630da45da2b3cc950b1f4e7cc3067f521cdecd062b19e36a6254b5dc6
|
|
| MD5 |
3fd19e5ae9ea471904d2c7265230a1ed
|
|
| BLAKE2b-256 |
edd84aa240233d0388336ed66b6983bcfcc963d25249afecaa614dda449009b7
|
Provenance
The following attestation bundles were made for mcp_gauntlet-0.1.0-py3-none-any.whl:
Publisher:
publish.yml on GhalebDweikat/mcp-gauntlet
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
mcp_gauntlet-0.1.0-py3-none-any.whl -
Subject digest:
c9ff5db630da45da2b3cc950b1f4e7cc3067f521cdecd062b19e36a6254b5dc6 - Sigstore transparency entry: 2229562219
- Sigstore integration time:
-
Permalink:
GhalebDweikat/mcp-gauntlet@497b42a0b2d8a1b7e8d99d3a281cdbe13e099391 -
Branch / Tag:
refs/tags/v0.1.0 - Owner: https://github.com/GhalebDweikat
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@497b42a0b2d8a1b7e8d99d3a281cdbe13e099391 -
Trigger Event:
release
-
Statement type: