toolfit
toolfit finds the specific places your MCP server confuses models, rewrites the tool descriptions to fix them, and proves the fix with a before-and-after eval.
uvx toolfit scan "npx -y @modelcontextprotocol/server-github" # free, static, seconds
uvx toolfit eval examples/toy_server.py --seeds 10 --fix --badge # model-graded, minutes, your key
Every existing MCP grader stops at a number and a list of complaints. toolfit measures which tool pairs a model actually confuses, proposes a description rewrite for each failing tool, re-runs the same tasks against the rewrite, and reports the delta with a p-value — accepted and rejected. The rejected ones are what make the accepted ones believable.
What a run looks like
toolfit eval examples/toy_server.py --seeds 10 --fix --badge, model under test claude-sonnet-5
(full output: docs/examples/toy-server/). The toy server has two
deliberately confusable pairs: create_task/update_task share the description "Add a new task.",
and list_tasks/count_tasks share "Get tasks by status."
## Confusion Matrix
| Intended \ Called | count_tasks | create_reminder | create_task | list_tasks | update_task |
|---|---|---|---|---|---|
| count_tasks | 0 | 0 | 0 | 10 | 0 |
| create_reminder | 0 | 10 | 0 | 0 | 0 |
| create_task | 0 | 0 | 10 | 0 | 0 |
| list_tasks | 0 | 0 | 0 | 10 | 0 |
| update_task | 0 | 0 | 0 | 0 | 10 |
## Pass Rates
- count_tasks: 0/10 (0%), 95% CI [0%, 28%]
- create_reminder: 8/10 (80%), 95% CI [49%, 94%]
- create_task: 9/10 (90%), 95% CI [60%, 98%]
- list_tasks: 10/10 (100%), 95% CI [72%, 100%]
- update_task: 10/10 (100%), 95% CI [72%, 100%]
## Proposed Fixes
### create_task — REJECTED
- Before: 'Add a new task.'
- After: 'Creates a brand-new task by specifying its title and priority, distinct from
update_task (which modifies existing tasks) or create_reminder (...)'
- Pass rate: 9/10 → 7/10, p-value 1.0000
- Reason: rejected: made things worse
### count_tasks — REJECTED
- Before: 'Get tasks by status.'
- After: 'Return the number of tasks matching a given status, rather than the tasks
themselves, by accepting a required status argument.'
- Pass rate: 0/10 → 0/10, p-value 1.0000
- Reason: rejected: no change
Three things this run shows, none of them flattering, all of them the point:
- The matrix names the real problem.
count_tasksis called zero times out of ten; every request meant for it goes tolist_tasks. The solvability check flagged the identicalcreate_task/update_taskdescriptions on 14 of 20 tasks — yet Sonnet 5 still picked the right one every time, because it readstask_idin the schema. Descriptions aren't the whole story, and toolfit measures what the model does, not what the text says. - The fix loop refused to claim a win. Three rewrites were proposed and measured; one made things worse, two changed nothing. They are printed anyway. Every failure here is at the argument level or in the generator's "count vs. list" phrasing — things a description rewrite can't fix — and the numbers say so instead of the tool saying so.
- Every number carries its uncertainty.
9/10and7/10overlap almost entirely at n=10; the exact test gives p=1.0 for a change in the wrong direction. Nothing gets reported as a finding because it looked good once.
On real servers
Same command against three public servers, model under test claude-sonnet-5, 10 seeds per tool
(full reports in docs/examples/):
| Server | Tools | Pass | What the matrix showed |
|---|---|---|---|
@modelcontextprotocol/server-memory |
9 | 90/90 | Perfect diagonal. Nine crisp descriptions, zero warnings, nothing to fix. |
mcp-server-git |
12 | 112/120 | git_commit 4/10: five requests went to git_add first. The rewrite that stressed "commits staged changes" measured worse, 4→2, and was rejected. |
@modelcontextprotocol/server-filesystem |
14 | 77/140 | read_file 0/10 — its description says DEPRECATED, use read_text_file and the model obeys (scan now flags this). Thirty-plus requests across seven tools went to list_allowed_directories first. |
The git and filesystem results share a cause that a description can't fix: the model takes a correct precondition step — stage before commit, check allowed directories before touching a path — and a single-step eval scores it as the wrong tool. That column in the confusion matrix is the finding; it tells you which tools need their precondition stated ("paths are validated for you") or a multi-step harness. toolfit does not paper over it by grading the first call leniently.
Twelve rewrites were proposed for the filesystem server. Five improved the number (8→10,
6→8); none were accepted, because one Bonferroni correction across twelve proposals at n=10
sets α=0.004 and the report says so. Run --fix-tool list_directory --fix-tool move_file --seeds 20 on the tools the matrix names, not --fix on the whole catalog at once.
Two commands, two budgets
scan |
eval |
|
|---|---|---|
| What | Static lint over tools/list: missing, too-short, and duplicated descriptions |
Live model behaviour: which tool it calls, with which arguments, for a request that should lead to each tool |
| Cost | Free — no model calls, no key | Your API key: roughly tools × seeds × 3 calls, plus seeds per proposed fix |
| Time | Under a second after the server starts | Minutes |
| Output | Findings list (never a letter grade) | Confusion matrix, per-tool pass rates with 95% CIs, mutation/fix verdicts, optional badge and toolfit-fixes.json |
They are deliberately separate. scan is the zero-config front door; on mature servers it finds
little (1 finding across 166 tools on 15 public servers) because the bug it
catches is the copy-paste class. The confusion is what eval is for.
Install
uvx toolfit --help # zero-install, one-off
pipx install toolfit # persistent / CI
Python 3.10+. Talks to servers over the MCP protocol via the official mcp SDK, so the
server can be in any language.
Pointing it at a server
toolfit scan path/to/server.py # run via `uv run`
toolfit scan "npx -y @modelcontextprotocol/server-filesystem ." # any command line
toolfit scan https://your-host/mcp # Streamable HTTP
The subprocess gets your environment (minus toolfit's own *_API_KEYs, which a third-party server
binary has no business seeing), so servers that read a token from GITHUB_TOKEN or
STRIPE_SECRET_KEY work unchanged. toolfit never calls a tool — it only ever asks for the
catalog (tools/list), then asks a model what it would call. Nothing touches your backend.
How the measurement works
The credibility problem with LLM-generated evals is circularity: if one model writes the task and also decides the expected answer, you are measuring agreement with that model. toolfit inverts it:
- Sample a concrete, schema-valid argument set for one tool (
gen/schema_sampler.py), honouringrequired, enums, formats, and nullables. - Ask a generator model to write the sentence a user would type that leads to exactly those arguments, without naming the tool. Ground truth is the sampled tuple, not the model's opinion.
- Send that sentence plus the whole catalog to the model under test.
- Grade structurally (
grade/grader.py): right tool, and arguments equal after canonicalising dates, case, whitespace, and array order. No LLM judge, ever.
Guardrails: a second pass checks each task for tool-name leakage and for solvability against the catalog; both are reported as warnings, never silently dropped. Tools whose schema the sampler can't handle are excluded and listed — a partial number is never printed as a complete one.
Mutation testing is the same grader run twice. --mutate 'tool:new description' re-runs a
tool's own base tasks against a catalog where only that description is patched (protocol-level;
your source is never touched) and compares pass rates on paired trials. --fix does the same with
a proposed rewrite for every failing tool.
Significance is an exact one-sided McNemar test on the discordant pairs, Bonferroni-corrected
across everything re-measured in one run. It is deterministic and honest at small n: with 5 seeds
the smallest attainable p-value is 1/32, so --seeds 10 is the practical floor for a verdict —
the CLI warns if you go lower. Every rate carries n and a Wilson 95% interval.
CI
- uses: sreshtalluri/toolfit@main
with:
server: "npx -y @modelcontextprotocol/server-github"
eval: true # omit for the free scan only
anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
scan --strict exits 1 on any finding; eval --strict exits 1 if any tool's pass rate is below
--strict-threshold (default 0.9). Tools excluded by a schema warning are named on stderr and
not counted — they don't fail the gate, but you'll see them.
--badge writes toolfit-badge.svg with the pass rate (or the before→after delta of a single
mutation or accepted fix), coloured by rate, with the model, generator, seed count, and a hash of
the task suite embedded so the number is never separable from what produced it.
Models
--model picks the model under test and the provider is inferred from the name: claude* →
Anthropic, gpt*/o* → OpenAI, vendor/model → OpenRouter. Keys come from ANTHROPIC_API_KEY,
OPENAI_API_KEY, OPENROUTER_API_KEY. Task generation and fix proposals always use Anthropic, so
that key is required for eval regardless of --model.
Data handling. The tool catalog and the generated task text are sent to whichever provider you configure, with your key. toolfit itself stores nothing and phones nowhere. If your server is internal, that is the one place its schema leaves your machine.
Landscape
| Tool | What it does | Where it stops |
|---|---|---|
| mcpgrade | Static lint + single-step eval, A–F grade | No re-measured fix |
| mcpx | ESLint-style schema lint, CI gate | Never runs a model |
| MCProbe | Usability rules + input fuzzing | Fuzzing, not task performance |
| lastmile-ai/mcp-eval | Assertion framework with LLM judges | You write the tests |
| MCPJam Inspector | Hosted evals with cross-model comparison | GUI product, not a CLI |
| MCP-Atlas / MCP-Bench | Benchmarks ranking models | Not pointable at your server |
toolfit is not a model leaderboard and does not test for prompt injection. One server per run.
Development
uv sync --extra dev
uv run pytest -q
uv run toolfit scan examples/toy_server.py
examples/toy_server.py has two deliberately confusable tool pairs — create_task/update_task
share a description, list_tasks/count_tasks share a vague one — so the fix loop has something
real to find. Design history and the methodology decisions behind every number are in
docs/designs/toolfit-v0-scope.md.
MIT.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file toolfit-0.1.0.tar.gz.
File metadata
- Download URL: toolfit-0.1.0.tar.gz
- Upload date:
- Size: 232.7 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
uv/0.12.10 {"installer":{"name":"uv","version":"0.12.10","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
38a3968b59a8f3ad687a369c69a4b64cf0cc1b42a66eaf0c0244d82580269d0c
|
|
| MD5 |
7e8c8e1b35d3e8005e1788e5bd41f3e2
|
|
| BLAKE2b-256 |
03fca953cb3303536354054f31ae50ac6a8ca3573cf44fd4fbc6b7881c5af316
|
File details
Details for the file toolfit-0.1.0-py3-none-any.whl.
File metadata
- Download URL: toolfit-0.1.0-py3-none-any.whl
- Upload date:
- Size: 40.8 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
uv/0.12.10 {"installer":{"name":"uv","version":"0.12.10","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
2f3fe9d778ade07336eb828e8e34dbd5b43c0d2c56f82dca24d2c8df3a8423a6
|
|
| MD5 |
16beef090e175320f1b0948fc6fd08f1
|
|
| BLAKE2b-256 |
58bfa1ffc02f3d96c6daac6dba1442197dbcbdfe89ef84a715bb6c092528d21b
|