Skip to main content

evalmcp

Evaluation suite for MCP AI agents -- golden datasets, LLM-as-judge, security benchmarks

Part of the MCP AI Suite.

Features

  • Built-in benchmark suites -- 6 golden datasets: memory, security, reasoning, tool-use, HumanEval-style code, and MMLU-style knowledge
  • Pluggable judges -- exact match, contains, or LLM-as-judge for semantic evaluation
  • Library, CLI & MCP server -- use it from Python, the evalmcp CLI, the evalmcp-server MCP server (3 tools: list_suites, run_suite, evaluate), or the evalmcp-api FastAPI service
  • Regression detection -- compare successive runs to catch quality drops above a threshold
  • Standard metrics -- accuracy, precision, recall, F1, per-tag breakdowns
  • HTML dashboard export -- visual report of evaluation results
  • JSON and CSV export -- machine-readable result output
  • Model comparison -- side-by-side evaluation of different models or configurations
  • Persistent store -- track evaluation runs over time for trend analysis

Installation

pip install mcpaisuite-evalmcp          # runtime deps: click + mcp
pip install "mcpaisuite-evalmcp[api]"   # adds the FastAPI server (evalmcp-api)

Quick Start

from evalmcp import EvalPipeline, EvalSuite, EvalCase

suite = EvalSuite(name="my_tests", cases=[
    EvalCase(input="What is 2+2?", expected_output="4", tool="run_task", tags=["math"]),
    EvalCase(input="Capital of France?", expected_output="Paris", tool="run_task", tags=["geography"]),
])

pipeline = EvalPipeline(judge="contains")
results = await pipeline.run_suite(suite)
summary = pipeline.summary(results)
print(f"Pass rate: {summary['pass_rate']:.0%}")

CLI

# List available benchmark suites:
evalmcp list

# Run a benchmark suite:
evalmcp run memory_basic --judge contains

# CI mode with regression detection:
evalmcp run security --judge exact --ci --threshold 0.1

# Export HTML dashboard:
evalmcp run reasoning --html report.html

Configuration

EvalPipeline is configured programmatically via constructor parameters.

Parameter Description
kernel_pipeline Optional MCP kernel pipeline for generating outputs
judge Judge strategy: "exact", "contains", "llm", or a BaseJudge instance
llm_fn Async callable (prompt) -> str, required when judge="llm"
store Optional EvalStore for persistence and regression detection

API Reference

EvalPipeline

Orchestrates evaluation of MCP agent outputs against expected results.

pipeline = EvalPipeline(judge="llm", llm_fn=my_llm, store=EvalStore())
results = await pipeline.run_suite(suite) -> list[EvalResult]
result = await pipeline.run_case(case) -> EvalResult
summary = pipeline.summary(results) -> dict
metrics = pipeline.metrics(results) -> dict
regression = pipeline.detect_regression(suite_name, threshold=0.1) -> dict

Utility Functions

from evalmcp import export_json, export_csv, compute_metrics, generate_dashboard

export_json(results, summary, "results.json")
export_csv(results, "results.csv")
metrics = compute_metrics(results)
generate_dashboard(results, summary, metrics, output_path="report.html")

Architecture

EvalPipeline iterates over EvalCases in a suite, optionally generating outputs via a kernel_pipeline, then scoring each result through a pluggable judge (ExactMatchJudge, ContainsJudge, or LLMJudge). Results are aggregated into summary statistics with per-tag breakdowns. EvalStore persists run history to enable regression detection by comparing the two most recent runs.

Testing

pip install -e ".[dev]"
pytest tests/ -v

License

Apache-2.0 — see LICENSE.

Open source for individuals and open-source projects. For commercial use in closed-source products, a commercial license is available — contact contact@mcpaisuite.com.

Release files for mcpaisuite-evalmcp 1.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for mcpaisuite-evalmcp 1.1.0
File Size Uploaded
mcpaisuite_evalmcp-1.1.0.tar.gz 32.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for mcpaisuite-evalmcp 1.1.0
File Interpreter ABI Platform
mcpaisuite_evalmcp-1.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 69.2 kB

Release files / mcpaisuite_evalmcp-1.1.0.tar.gz

Download URL mcpaisuite_evalmcp-1.1.0.tar.gz
Size 32.2 kB
Tags Source
SHA-256 checksum
How to use checksums
cc7885b674014bd81e2ec5a91c5b30564fde59268b5291b1cea95849d830182b
BLAKE2b-256 checksum
How to use checksums
8702bb74bef6e390dfda0c589b6bedd8b57d78f804f38a2c3254f5b41abcdc05
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jun 22, 2026.

Transparency log

Release files / mcpaisuite_evalmcp-1.1.0-py3-none-any.whl

Download URL mcpaisuite_evalmcp-1.1.0-py3-none-any.whl
Size 37.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
9fd40d2bbaa3c4abf6e0a992c57c70da9b67e4ef4fed80dce63a93fe76847fd0
BLAKE2b-256 checksum
How to use checksums
07a4eb6eef2aa47f8281af62368b1abdbef39fb0116d0aec69f694b11086e9a4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jun 22, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

1.1.0 This release

2 release files

1.0.5

2 release files

1.0.4

2 release files

1.0.3

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page