Skip to main content

toolrank

PyPI CI License Heads

Tool retrieval for LLM agents with hundreds of tools. Instead of putting every tool definition into the prompt, toolrank picks the few a request needs, with an embedding model (Qwen3-Embedding-8B, trained further on tool retrieval), and serves them to your agent over MCP or REST. Every retriever it ships is measured on the same public benchmarks.

Status: alpha (0.2); interfaces may still change. Documentation: https://yaman.dev/toolrank/

Quick start

Index MCP servers or OpenAPI specs, then serve them to any MCP client as two tools, search_tools and call_tool:

pip install "toolrank[mcp]"
toolrank ingest mcp --server time="uvx mcp-server-time" --out tools/
toolrank serve --data tools/                          # MCP at http://127.0.0.1:8765/mcp, REST at /v1

toolrank ranks with its backbone, Qwen3-Embedding-8B trained further on tool retrieval (yasinyaman/toolrank-emb-8b), served by vLLM (--emb-url, by default http://127.0.0.1:8091/v1). On a GPU host, Docker runs both:

cd deploy/docker && cp .env.example .env              # set TOOLRANK_API_KEY
docker compose run --rm toolrank ingest mcp --config /config/toolrank.json --out /data
docker compose up -d

The quick start connects Claude Code, Claude Desktop and REST clients.

Results

Retrieval quality on three benchmarks, every number from toolrank eval under ToolRet's protocol (the reports are in docs/results/; scripts/readme_table.py checks each one before it prints it).

Retriever ToolRet NDCG@10 ToolRet NDCG@10 cat-macro LiveMCPBench Recall@5 MCP-Zero top-1
BM25, without instruction 29.01 22.24 31.68 80.44
BM25, with instruction 39.27 36.41 22.92 45.63
Qwen3-Embedding-8B 51.11 46.54 50.82 78.19
Qwen3-Embedding-8B + toolrank heads v0.1 54.03 47.13 53.03 79.87
Qwen3-Embedding-8B in FP8 + toolrank heads v0.1 53.94 47.27 53.48 79.51
toolrank backbone v0.2 (Qwen3-Embedding-8B + LoRA) 58.90 54.36 52.06 88.57
toolrank backbone v0.2 in FP8 (the default) 59.02 54.53 52.06 87.71
NV-Embed-v1 (ToolRet paper) — 42.71 — —
gte-Qwen2-1.5B-instruct (ToolRet paper) — 45.96 — —
StackOne v2, a fine-tuned 109M BGE-base (StackOne) — 54.40 — —
  • ToolRet: 7,961 queries over 44,453 tools, top 100 over the whole corpus. NDCG@10 is the micro-average of the paper's released code; cat-macro is the paper's own aggregation (the mean of the web, code and customized categories) and the only column with published numbers. Our BM25 reproduces the paper's BM25s within 0.1 (22.24 / 36.41 against 22.32 / 36.46).
  • LiveMCPBench (94 queries, 525 tools) and MCP-Zero (2,792 tools): the tool text includes the MCP server's name (toolrank data server-names). One LiveMCPBench query is about one point. MCP-Zero ships no queries: ours were written by Qwen3-8B, one per tool (toolrank data pull mcp-zero), so its column does not compare with the MCP-Zero paper. Top-1 is Precision@1.
  • With instruction, each query carries its task's instruction (ToolRet) or a generic one (the MCP sets), as the embedding model is served; the generic instruction costs BM25 on the MCP sets. BM25 is bm25s without stemming, the paper's setting.
  • The heads (29.9M parameters, docs/heads/MODEL_CARD.md) were trained on ToolRet's training pairs, so ToolRet is in-domain for them and the MCP sets are not. On MCP-Zero, BM25 without instruction still wins at top-1: each generated query opens with a server: line that usually names the server, and exact matching rewards that.
  • The toolrank backbone v0.2 (docs/backbone/MODEL_CARD.md) is Qwen3-Embedding-8B with a LoRA trained on 20,000 of ToolRet's training pairs, served without heads (the v0.1 heads cost it 1–2 points). Its checkpoint was picked on MCP-Zero, so that column is its selection set; ToolRet is in-domain, LiveMCPBench is held out. Beyond these sets, on generated requests over a GitHub + Stripe catalogue it gains 12–20 NDCG@10 points on tasks that need two or three tools and ties with the base model on requests for one tool (the model card has the numbers).
  • FP8: the bf16 weights quantized as vLLM loads them (--quantization fp8). Every column is within a point of bf16, at about half the weight memory and batch-1 latency.
  • Reproduce: bash scripts/readme_results.sh where the backbone is served (reports in docs/results/; EMB_URL, EMB_MODEL, TAG and ROWS select another endpoint and rows), then uv run python scripts/readme_table.py --write.

What's inside

  • Ingestion of MCP servers (stdio and streamable HTTP) and OpenAPI 3.x specs; a re-run syncs only what changed. Guide
  • Search and serve: adaptive K, a persistent vector index (numpy, FAISS HNSW or pgvector), an MCP proxy with two tools, a REST API, API keys and a usage log. Guide
  • Agent platforms: toolrank as Claude's (tool_reference) and OpenAI's (client-side tool_search) tool search. Guide
  • Frameworks: LangGraph (langgraph-bigtool), LlamaIndex agents and the LiteLLM proxy. Guide
  • Fine-tuning: heads trained on your own request-to-tool pairs, the epoch picked on a dev set. Guide
  • Benchmarks: ToolRet, LiveMCPBench and MCP-Zero with BM25, dense and head scorers (and CLM, for comparison). Benchmarks
  • Docker: the toolrank image for amd64 and arm64, compose files with vLLM, and a Dockerfile that puts vLLM and toolrank in one container. Guide

Why

  • Tool definitions are expensive context. Anthropic measured 58 tools at about 55K tokens per request and reports tool-selection accuracy falling past 30 to 50 tools. RAG-MCP lifted selection accuracy from 13.6% to 43.1% on a large MCP set by retrieving tools first.
  • Hosted tool searches are tied to one model provider or cloud, and most are lexical. toolrank is model-agnostic and runs on your own hardware.
  • Small heads on a frozen embedding model (29.9M parameters for the request and the tool side together). You embed your tools once and rank with one dot product. The heads start as the identity, so heads fine-tuned on your data start from the base model's quality, not below it.

Feedback

Tried it? Tell us how it went: what you set up, what worked and what did not. Questions and ideas go to Discussions.

Contributing

Set-up, tests and conventions are in CONTRIBUTING.md. The weekly reports behind every number, in Turkish, are in docs/reports/.

License

Apache-2.0 (LICENSE). NOTICE credits what toolrank takes from CLM (the head architecture), ToolRet (its task metadata) and MCP-Zero (its query prompts); THIRD_PARTY_NOTICES.md lists the dependencies, models and container images. Contributions are welcome: see CONTRIBUTING.md and the code of conduct; report vulnerabilities as SECURITY.md says.

Metadata

Release files for toolrank 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for toolrank 0.2.0
File Size Uploaded
toolrank-0.2.0.tar.gz 287.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for toolrank 0.2.0
File Interpreter ABI Platform
toolrank-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 496.0 kB

Release files / toolrank-0.2.0.tar.gz

Download URL toolrank-0.2.0.tar.gz
Size 287.7 kB
Tags Source
SHA-256 checksum
How to use checksums
85308041f80d3108db9c8c096440a0d47e44b0152add810e74caa427d80c8251
BLAKE2b-256 checksum
How to use checksums
f89296d14134c08d035c4559566b4c74904bb3f395a2c95bb345ec1d98ea3c45
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.

Transparency log

Release files / toolrank-0.2.0-py3-none-any.whl

Download URL toolrank-0.2.0-py3-none-any.whl
Size 208.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
99e758cef260532b95b99c5a7755af0c6613948ec2cbd18e5479de92e9d5082a
BLAKE2b-256 checksum
How to use checksums
f7a16a57665125b484123bb84c103851af552b74d8d380b59a140160a85efeed
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.2.0 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page