evalmcp
Evaluation suite for MCP AI agents -- golden datasets, LLM-as-judge, security benchmarks
Part of the MCP AI Suite.
Features
- Built-in benchmark suites -- 6 golden datasets: memory, security, reasoning, tool-use, HumanEval-style code, and MMLU-style knowledge
- Pluggable judges -- exact match, contains, or LLM-as-judge for semantic evaluation
- Library, CLI & MCP server -- use it from Python, the
evalmcpCLI, theevalmcp-serverMCP server (3 tools:list_suites,run_suite,evaluate), or theevalmcp-apiFastAPI service - Regression detection -- compare successive runs to catch quality drops above a threshold
- Standard metrics -- accuracy, precision, recall, F1, per-tag breakdowns
- HTML dashboard export -- visual report of evaluation results
- JSON and CSV export -- machine-readable result output
- Model comparison -- side-by-side evaluation of different models or configurations
- Persistent store -- track evaluation runs over time for trend analysis
Installation
pip install mcpaisuite-evalmcp # runtime deps: click + mcp
pip install "mcpaisuite-evalmcp[api]" # adds the FastAPI server (evalmcp-api)
Quick Start
from evalmcp import EvalPipeline, EvalSuite, EvalCase
suite = EvalSuite(name="my_tests", cases=[
EvalCase(input="What is 2+2?", expected_output="4", tool="run_task", tags=["math"]),
EvalCase(input="Capital of France?", expected_output="Paris", tool="run_task", tags=["geography"]),
])
pipeline = EvalPipeline(judge="contains")
results = await pipeline.run_suite(suite)
summary = pipeline.summary(results)
print(f"Pass rate: {summary['pass_rate']:.0%}")
CLI
# List available benchmark suites:
evalmcp list
# Run a benchmark suite:
evalmcp run memory_basic --judge contains
# CI mode with regression detection:
evalmcp run security --judge exact --ci --threshold 0.1
# Export HTML dashboard:
evalmcp run reasoning --html report.html
Configuration
EvalPipeline is configured programmatically via constructor parameters.
| Parameter | Description |
|---|---|
kernel_pipeline |
Optional MCP kernel pipeline for generating outputs |
judge |
Judge strategy: "exact", "contains", "llm", or a BaseJudge instance |
llm_fn |
Async callable (prompt) -> str, required when judge="llm" |
store |
Optional EvalStore for persistence and regression detection |
API Reference
EvalPipeline
Orchestrates evaluation of MCP agent outputs against expected results.
pipeline = EvalPipeline(judge="llm", llm_fn=my_llm, store=EvalStore())
results = await pipeline.run_suite(suite) -> list[EvalResult]
result = await pipeline.run_case(case) -> EvalResult
summary = pipeline.summary(results) -> dict
metrics = pipeline.metrics(results) -> dict
regression = pipeline.detect_regression(suite_name, threshold=0.1) -> dict
Utility Functions
from evalmcp import export_json, export_csv, compute_metrics, generate_dashboard
export_json(results, summary, "results.json")
export_csv(results, "results.csv")
metrics = compute_metrics(results)
generate_dashboard(results, summary, metrics, output_path="report.html")
Architecture
EvalPipeline iterates over EvalCases in a suite, optionally generating outputs via a kernel_pipeline, then scoring each result through a pluggable judge (ExactMatchJudge, ContainsJudge, or LLMJudge). Results are aggregated into summary statistics with per-tag breakdowns. EvalStore persists run history to enable regression detection by comparing the two most recent runs.
Testing
pip install -e ".[dev]"
pytest tests/ -v
License
Apache-2.0 — see LICENSE.
Open source for individuals and open-source projects. For commercial use in closed-source products, a commercial license is available — contact contact@mcpaisuite.com.
Release files for mcpaisuite-evalmcp 1.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| mcpaisuite_evalmcp-1.1.0.tar.gz | 32.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| mcpaisuite_evalmcp-1.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 69.2 kB
Release files / mcpaisuite_evalmcp-1.1.0.tar.gz
| Download URL | mcpaisuite_evalmcp-1.1.0.tar.gz |
|---|---|
| Size | 32.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
cc7885b674014bd81e2ec5a91c5b30564fde59268b5291b1cea95849d830182b
|
|
BLAKE2b-256 checksum How to use checksums |
8702bb74bef6e390dfda0c589b6bedd8b57d78f804f38a2c3254f5b41abcdc05
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.12
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jun 22, 2026.
Transparency logRelease files / mcpaisuite_evalmcp-1.1.0-py3-none-any.whl
| Download URL | mcpaisuite_evalmcp-1.1.0-py3-none-any.whl |
|---|---|
| Size | 37.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
9fd40d2bbaa3c4abf6e0a992c57c70da9b67e4ef4fed80dce63a93fe76847fd0
|
|
BLAKE2b-256 checksum How to use checksums |
07a4eb6eef2aa47f8281af62368b1abdbef39fb0116d0aec69f694b11086e9a4
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.12
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jun 22, 2026.
Transparency log