Skip to main content

LLM Benchmark MCP Server

MCP server that gives AI agents access to LLM benchmark data, pricing comparisons, and model recommendations.

Features

  • compare_models — Side-by-side benchmark comparison of LLMs (MMLU, HumanEval, MATH, GPQA, ARC, HellaSwag)
  • get_model_details — Detailed info about a specific model including strengths/weaknesses
  • recommend_model — Get the best model recommendation for your task and budget
  • list_top_models — Top models ranked by category (coding, math, reasoning, chat)
  • get_pricing — Pricing comparison via OpenRouter API

Supported Models

GPT-4o, GPT-4o-mini, GPT-4 Turbo, o1, o3-mini, Claude 3.5 Sonnet, Claude 3.5 Haiku, Claude 3 Opus, Gemini 2.0 Flash, Gemini 2.0 Pro, Gemini 1.5 Pro, Llama 3.1 (8B/70B/405B), Llama 3.3 70B, Mistral Large, Mistral Small, Mixtral 8x22B, DeepSeek V3, DeepSeek R1, Qwen 2.5 72B

Installation

pip install llm-benchmark-mcp-server

Usage with Claude Desktop

Add to your claude_desktop_config.json:

{
  "mcpServers": {
    "llm-benchmark": {
      "command": "benchmark-server"
    }
  }
}

Or via uvx (no install needed):

{
  "mcpServers": {
    "llm-benchmark": {
      "command": "uvx",
      "args": ["llm-benchmark-mcp-server"]
    }
  }
}

Example Queries

  • "Compare GPT-4o vs Claude 3.5 Sonnet vs Gemini 2.0 Pro"
  • "Which model is best for coding on a low budget?"
  • "Show me the top 10 models for math"
  • "What does GPT-4o cost compared to Claude?"
  • "Give me details about DeepSeek R1"

Data Sources

  • Benchmarks: Hardcoded from official papers and public leaderboards (MMLU, HumanEval, MATH, GPQA, ARC-Challenge, HellaSwag)
  • Pricing: Live data from OpenRouter API
  • Arena Rankings: Chatbot Arena Leaderboard (when available)

License

MIT

Metadata

Release files for llm-benchmark-mcp-server 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for llm-benchmark-mcp-server 0.1.0
File Size Uploaded
llm_benchmark_mcp_server-0.1.0.tar.gz 9.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for llm-benchmark-mcp-server 0.1.0
File Interpreter ABI Platform
llm_benchmark_mcp_server-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 21.1 kB

Release files / llm_benchmark_mcp_server-0.1.0.tar.gz

Download URL llm_benchmark_mcp_server-0.1.0.tar.gz
Size 9.4 kB
Tags Source
SHA-256 checksum
How to use checksums
ec988809043ad7eb093471ef5c11396fd05c07a08baa320c550f6f5d6a3a7155
BLAKE2b-256 checksum
How to use checksums
01be2cd9a39b0a0bf5d672e2c06e62017b602eb0b618b6a0fe2ba7518a0a50f5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.13.3

Release files / llm_benchmark_mcp_server-0.1.0-py3-none-any.whl

Download URL llm_benchmark_mcp_server-0.1.0-py3-none-any.whl
Size 11.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
3da7aa55b3c5249cd487b3ce05b705c5718a281cbc061fd48df4ac1a96f0be98
BLAKE2b-256 checksum
How to use checksums
bad3638a41393384487cf6dd23511b75bda33c24c3076e9ba3e0eccb2cb4fe51
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.13.3

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page