toolrank
Tool retrieval for LLM agents with hundreds of tools. Instead of putting every tool definition into the prompt, toolrank picks the few a request needs, with an embedding model (Qwen3-Embedding-8B, trained further on tool retrieval), and serves them to your agent over MCP or REST. Every retriever it ships is measured on the same public benchmarks.
Status: alpha (0.2); interfaces may still change. Documentation: https://yaman.dev/toolrank/
Quick start
Index MCP servers or OpenAPI specs, then serve them to any MCP client as two tools, search_tools
and call_tool:
pip install "toolrank[mcp]"
toolrank ingest mcp --server time="uvx mcp-server-time" --out tools/
toolrank serve --data tools/ # MCP at http://127.0.0.1:8765/mcp, REST at /v1
toolrank ranks with its backbone, Qwen3-Embedding-8B trained further on tool retrieval
(yasinyaman/toolrank-emb-8b), served by vLLM (--emb-url, by default
http://127.0.0.1:8091/v1). On a GPU host, Docker runs both:
cd deploy/docker && cp .env.example .env # set TOOLRANK_API_KEY
docker compose run --rm toolrank ingest mcp --config /config/toolrank.json --out /data
docker compose up -d
The quick start connects Claude Code, Claude Desktop and REST clients.
Results
Retrieval quality on three benchmarks, every number from toolrank eval under ToolRet's protocol
(the reports are in docs/results/;
scripts/readme_table.py checks each one before it prints it).
| Retriever | ToolRet NDCG@10 | ToolRet NDCG@10 cat-macro | LiveMCPBench Recall@5 | MCP-Zero top-1 |
|---|---|---|---|---|
| BM25, without instruction | 29.01 | 22.24 | 31.68 | 80.44 |
| BM25, with instruction | 39.27 | 36.41 | 22.92 | 45.63 |
| Qwen3-Embedding-8B | 51.11 | 46.54 | 50.82 | 78.19 |
| Qwen3-Embedding-8B + toolrank heads v0.1 | 54.03 | 47.13 | 53.03 | 79.87 |
| Qwen3-Embedding-8B in FP8 + toolrank heads v0.1 | 53.94 | 47.27 | 53.48 | 79.51 |
| toolrank backbone v0.2 (Qwen3-Embedding-8B + LoRA) | 58.90 | 54.36 | 52.06 | 88.57 |
| toolrank backbone v0.2 in FP8 (the default) | 59.02 | 54.53 | 52.06 | 87.71 |
| NV-Embed-v1 (ToolRet paper) | — | 42.71 | — | — |
| gte-Qwen2-1.5B-instruct (ToolRet paper) | — | 45.96 | — | — |
| StackOne v2, a fine-tuned 109M BGE-base (StackOne) | — | 54.40 | — | — |
- ToolRet: 7,961 queries over 44,453 tools, top 100 over the whole corpus. NDCG@10 is the micro-average of the paper's released code; cat-macro is the paper's own aggregation (the mean of the web, code and customized categories) and the only column with published numbers. Our BM25 reproduces the paper's BM25s within 0.1 (22.24 / 36.41 against 22.32 / 36.46).
- LiveMCPBench (94 queries, 525 tools) and MCP-Zero (2,792 tools): the tool text includes the MCP server's name (
toolrank data server-names). One LiveMCPBench query is about one point. MCP-Zero ships no queries: ours were written by Qwen3-8B, one per tool (toolrank data pull mcp-zero), so its column does not compare with the MCP-Zero paper. Top-1 is Precision@1. - With instruction, each query carries its task's instruction (ToolRet) or a generic one (the MCP sets), as the embedding model is served; the generic instruction costs BM25 on the MCP sets. BM25 is bm25s without stemming, the paper's setting.
- The heads (29.9M parameters,
docs/heads/MODEL_CARD.md) were trained on ToolRet's training pairs, so ToolRet is in-domain for them and the MCP sets are not. On MCP-Zero, BM25 without instruction still wins at top-1: each generated query opens with aserver:line that usually names the server, and exact matching rewards that. - The toolrank backbone v0.2 (
docs/backbone/MODEL_CARD.md) is Qwen3-Embedding-8B with a LoRA trained on 20,000 of ToolRet's training pairs, served without heads (the v0.1 heads cost it 1–2 points). Its checkpoint was picked on MCP-Zero, so that column is its selection set; ToolRet is in-domain, LiveMCPBench is held out. Beyond these sets, on generated requests over a GitHub + Stripe catalogue it gains 12–20 NDCG@10 points on tasks that need two or three tools and ties with the base model on requests for one tool (the model card has the numbers). - FP8: the bf16 weights quantized as vLLM loads them (
--quantization fp8). Every column is within a point of bf16, at about half the weight memory and batch-1 latency. - Reproduce:
bash scripts/readme_results.shwhere the backbone is served (reports indocs/results/;EMB_URL,EMB_MODEL,TAGandROWSselect another endpoint and rows), thenuv run python scripts/readme_table.py --write.
What's inside
- Ingestion of MCP servers (stdio and streamable HTTP) and OpenAPI 3.x specs; a re-run syncs only what changed. Guide
- Search and serve: adaptive K, a persistent vector index (numpy, FAISS HNSW or pgvector), an MCP proxy with two tools, a REST API, API keys and a usage log. Guide
- Agent platforms: toolrank as Claude's (
tool_reference) and OpenAI's (client-sidetool_search) tool search. Guide - Frameworks: LangGraph (langgraph-bigtool), LlamaIndex agents and the LiteLLM proxy. Guide
- Fine-tuning: heads trained on your own request-to-tool pairs, the epoch picked on a dev set. Guide
- Benchmarks: ToolRet, LiveMCPBench and MCP-Zero with BM25, dense and head scorers (and CLM, for comparison). Benchmarks
- Docker: the
toolrankimage for amd64 and arm64, compose files with vLLM, and a Dockerfile that puts vLLM and toolrank in one container. Guide
Why
- Tool definitions are expensive context. Anthropic measured 58 tools at about 55K tokens per request and reports tool-selection accuracy falling past 30 to 50 tools. RAG-MCP lifted selection accuracy from 13.6% to 43.1% on a large MCP set by retrieving tools first.
- Hosted tool searches are tied to one model provider or cloud, and most are lexical. toolrank is model-agnostic and runs on your own hardware.
- Small heads on a frozen embedding model (29.9M parameters for the request and the tool side together). You embed your tools once and rank with one dot product. The heads start as the identity, so heads fine-tuned on your data start from the base model's quality, not below it.
Feedback
Tried it? Tell us how it went: what you set up, what worked and what did not. Questions and ideas go to Discussions.
Contributing
Set-up, tests and conventions are in
CONTRIBUTING.md. The weekly
reports behind every number, in Turkish, are in
docs/reports/.
License
Apache-2.0 (LICENSE). NOTICE credits what toolrank takes from CLM (the head
architecture), ToolRet (its task metadata) and MCP-Zero (its query prompts);
THIRD_PARTY_NOTICES.md lists the dependencies, models and container
images. Contributions are welcome: see CONTRIBUTING.md and the
code of conduct; report vulnerabilities as SECURITY.md says.
Metadata
Release files for toolrank 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| toolrank-0.2.0.tar.gz | 287.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| toolrank-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 496.0 kB
Release files / toolrank-0.2.0.tar.gz
| Download URL | toolrank-0.2.0.tar.gz |
|---|---|
| Size | 287.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
85308041f80d3108db9c8c096440a0d47e44b0152add810e74caa427d80c8251
|
|
BLAKE2b-256 checksum How to use checksums |
f89296d14134c08d035c4559566b4c74904bb3f395a2c95bb345ec1d98ea3c45
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.
Transparency logRelease files / toolrank-0.2.0-py3-none-any.whl
| Download URL | toolrank-0.2.0-py3-none-any.whl |
|---|---|
| Size | 208.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
99e758cef260532b95b99c5a7755af0c6613948ec2cbd18e5479de92e9d5082a
|
|
BLAKE2b-256 checksum How to use checksums |
f7a16a57665125b484123bb84c103851af552b74d8d380b59a140160a85efeed
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.
Transparency log