Skip to main content

toolhunch

Your agent has too many tools and calls the wrong one.
toolhunch searches them, lets a fast decision model pick, and can say “none”.

PyPI pre-alpha Python 3.12+ MIT license Pydantic AI integration Jev-compatible: Laya Benchmarked on ToolRet, 44,453 tools

Which tool should the agent call? GPT-6 Luna agent 80%, Jev 74%, GPT-4.1 mini agent 61%, search only 54%

On catalogs of 40–101 tools, a GPT-6 Luna agent given every tool as a function picked a relevant tool for 80% of 290 requests, Jev 74% and a GPT-4.1 mini agent 61%. Jev answered in 0.32 s against Luna's 1.45 s, and Luna never answered “none”. Results and limits below.

Why search tools at all?

An agent usually receives every tool definition on every call. That is fine for a handful of tools, but:

  • Context and cost. Every definition takes context-window space and input tokens on every call.
  • Hard limits. OpenAI rejected our request with 200 functions (array_above_max_length).
  • Caching helps, within limits. Prompt caching makes a fixed tool list cheaper, but the list still fills the context, and changing it breaks the cache.

What we measured: at 40–101 tools, both GPT agents did about as well with every tool as with 20 searched ones, and caching made the whole catalog cheap (GPT-6 Luna: $0.07 against $0.18 per 1,000 requests). Search earns its place as catalogs grow: hundreds of MCP tools, or ToolRet's 44,453.

How it works

request ──► search (BM25 + embeddings) ──► 20 candidates ──► decider ──► one tool, or “none”
  • Search ranks tools by keywords (BM25) and by meaning (embeddings), merged with reciprocal rank fusion.
  • A decider reads the request and the candidates, then picks one or answers “none”. It can be a System-1 decision model such as TypeSafe's Jev, a Jev-compatible server (Laya, rizzo-flow), or an LLM that ranks the options with its token probabilities (logprobs) or names one as structured output.
  • Two ways to use a decider: after search, on the 20 candidates, at any catalog size; or instead of search, reading the whole catalog, when it is small.

The pipeline ranks tools; an integration decides how they reach the agent. The core is framework-free.

Results

The data is ToolRet (Shi et al., Findings of ACL 2025): requests labelled with the tools that solve them, over a 44,453-tool corpus. “Relevant” means one of the labelled tools. These tests measure tool selection on single-turn requests, not task completion. Intervals are 95% bootstrap intervals over tasks; every figure is generated from a published summary.

1. Small catalogs: read everything, or search first?

Four ToolRet catalogs of 40–101 tools, with the same 290 requests for every strategy (chart above):

Strategy Relevant tool picked (95% CI) Cost / 1,000 requests Median latency
GPT-6 Luna agent, whole catalog 79.7% (74.8–84.1) $0.07 (99.6% cached) 1.45 s
GPT-6 Luna agent, 20 searched tools 77.6% (72.4–82.1) $0.18 1.24 s
Jev reads the whole catalog 74.1% (69.0–79.0) $0.28 0.32 s
Search, then Jev picks among 20 71.4% (65.9–76.6) $0.08 0.45 s
GPT-4.1 mini agent, 20 searched tools 61.7% (55.9–66.9) $0.68 0.94 s
GPT-4.1 mini agent, whole catalog 60.7% (55.2–66.2) $0.79 (92% cached) 0.86 s
Search only, top result 54.5% (48.6–60.0) ≈ $0 0.16 s

The agent makes one function-calling request; its first call is scored, never executed. GPT-6 Luna runs with reasoning off, in a later run on the same requests. Costs and latency are for warm requests, after the first of each catalog, and include search where a strategy searches first.

2. 44,453 tools: search, then decide

Search only 22%, Jev 32%, GPT-4.1 mini 33%, GPT-6 Luna 34%; ceiling 59%

On 200 held-out requests, search alone puts a relevant tool first 22% of the time. A decider over the 20 candidates raises that to 32% (Jev), 33% (GPT-4.1 mini logprobs) or 34% (GPT-6 Luna): the same precision within the intervals, with Jev the cheapest and fastest ($0.05 per 1,000 searches and 0.25 s, against $0.12 and 1.06 s for Luna, $0.52 and 1.31 s for GPT-4.1 mini). Search is the limit: a relevant tool is among the 20 candidates only 59% of the time. With 50 candidates the ceiling is 71.5%, and Jev reaches 37.0%.

Why GPT-4.1 mini for logprobs? The decider reads the probability of every option letter, and GPT-4.1 mini is the newest OpenAI model we found that returns 20 of them: GPT-5-mini refuses logprobs, GPT-5.4-mini returns at most 5, and GPT-6 Luna at most 5, only with reasoning off. So Luna answers the same prompt with one letter as structured output: no probabilities, hence no confidence threshold and no averaging below.

3. When no tool fits

To test “none”, we remove the labelled tools from some requests and check who notices. The same answer can be read in three ways:

Reading The decider…
Answer always must pick a tool
Allow “none” may answer “none of these”
Confidence threshold answers only when its probability is at least 0.95 (chosen on dev data), otherwise “none”

Outcomes with and without a right tool in the catalog

On small catalogs, with “none” allowed: when no right tool existed, Jev said “none” 45% of the time and the GPT-4.1 mini agent 41%, so both usually picked something anyway. The GPT-6 Luna agent never said it: it called a tool for all 145 requests without a right one. When a right tool existed, the GPT-4.1 mini agent said “none” to 27% of requests, Jev to 12%.

A threshold is no cure either. At 44,453 tools, on 200 requests with a relevant tool and 200 without:

Jev, 20 candidates Right Wrong “None”
Answer always 64 336 0
Allow “none” 64 265 71
Threshold 0.95 20 29 351

Without a relevant tool any pick is wrong, so “answer always” is wrong at least 200 times by design. The threshold removed about 9 in 10 wrong picks, and 2 in 3 right ones. GPT-4.1 mini gives 66/334/0, 61/272/67 and 52/215/133 for the same readings; GPT-6 Luna, allowed “none”, gives 60/257/83.

4. Does the order of the candidates matter?

A common criticism of decision models is that shuffling the options changes the answer. We asked the same 200 questions with the 20 candidates in five orders: search order and four shuffles.

Pairwise agreement: Jev 96% same order, 76.5% shuffled; GPT-4.1 mini 98% and 60.5%; GPT-6 Luna 92.1% and 62.9%

  • All three change their pick. Two orders agree on the top tool 76.5% of the time for Jev, 60.5% for GPT-4.1 mini logprobs and 62.9% for GPT-6 Luna, against 92–98% when the same order is repeated.
  • GPT-4.1 mini's precision moved most. A relevant tool came first 30.5% → 31.2% (mean of the shuffles) for Jev, 34.3% → 33.0% for Luna, and 33.0% → 29.0% for GPT-4.1 mini, lower in all four shuffles.
  • Position bias. GPT-4.1 mini and Luna picked one of the first three slots 23% of the time, Jev 18.2%; a uniform pick gives 15%.
  • Averaging five orders did not help (31.5% Jev, 31.0% GPT-4.1 mini) and costs five decisions per search.

The same-order repeats come from a separate run, so this is an observational comparison. Luna's figures use the 198 tasks it answered in every order (it returned an empty answer on 3 of 1,000 asks).

5. Cost and latency

Precision against median latency and cost: Luna most precise on small catalogs but slowest; at 44,453 tools all deciders tie and Jev is fastest and cheapest

On small catalogs the GPT-6 Luna agent was the most precise and, with its prompt cache, cheap, but took 1.45 s against Jev's 0.32 s. At 44,453 tools the deciders tie on precision, and Jev is both the fastest and the cheapest.

For the agents given the whole catalog, OpenAI served most input tokens from its prompt cache on warm requests: 92% for GPT-4.1 mini ($0.79 per 1,000 instead of $2.43 at list price) and 99.6% for GPT-6 Luna ($0.07 instead of $0.62). With 20 searched tools the list changes on every request, and nothing was cached ($0.68 and $0.18). Jev's provider-side caching was not measured. Multi-turn agent loops, where caching and tool reveal interact, come next.

The technical page has the protocol, every table, per-catalog results, latency, thresholds and limitations.

Quickstart

pip install toolhunch                  # or: uv add toolhunch
pip install "toolhunch[pydantic-ai]"   # with the Pydantic AI integration

Set TYPESAFE_API_KEY for Jev (and OPENAI_API_KEY for the agent below) in your environment:

import anyio

from toolhunch import Abstention, BM25Retriever, ChoiceDecider, ToolCard, ToolCatalog, ToolSearchPipeline, jev

catalog = ToolCatalog(
    [
        ToolCard(name="get_weather", description="Get the current weather in a city."),
        ToolCard(name="send_email", description="Send an e-mail message to a recipient."),
        # ... the rest of your tools
    ]
)
pipeline = ToolSearchPipeline(
    BM25Retriever(),  # 1. search: keep 20 candidates
    decider=ChoiceDecider(jev("jev-1.13.0"), abstention=Abstention()),  # 2. pick one, or "none"
    k=20,
    top_n=1,
)
result = anyio.run(pipeline.search, ["What's the weather in Milan?"], catalog)
print("none" if result.abstained else result.names[0])

BM25 needs no key; the benchmark uses HybridRetriever with BM25 and OpenAI embeddings. Abstention() adds “none” without a probability threshold: a threshold belongs to one model, prompt and payload shape, so calibrate your own.

With Pydantic AI

The same pipeline plugs into Pydantic AI's tool search: deferred tools stay hidden until the decider picks one.

from pydantic_ai import Agent
from pydantic_ai.capabilities import ToolSearch

from toolhunch.integrations.pydantic_ai import reveal_strategy

agent = Agent("openai:gpt-4.1-mini", capabilities=[ToolSearch(strategy=reveal_strategy(pipeline))])


@agent.tool_plain(defer_loading=True)  # hidden until the search reveals it
def get_weather(city: str) -> str:
    """Get the current weather in a city."""
    return f"{city}: sunny, 18 °C (demo)"


print(agent.run_sync("What's the weather in Milan?").output)

A complete script, run offline by the tests, is in examples/pydantic_ai_tool_search.py.

With a Jev-compatible server

Swap the decision model; everything else stays. For a local Laya server (laya-serve, English checkpoint):

from datetime import date

from toolhunch import JevWireModel, ModelLimits

laya = JevWireModel(
    "english",
    base_url="http://127.0.0.1:8000/v1",
    api_key_env=None,  # local server, no key
    limits=ModelLimits(  # declared, with their source and date
        max_options_per_choice=100,
        max_state_plus_question_tokens=512,
        max_questions_per_request=64,
        source="laya-serve 0.3.22",
        checked=date(2026, 9, 30),
    ),
)
pipeline = ToolSearchPipeline(BM25Retriever(), decider=ChoiceDecider(laya, abstention=Abstention()), k=20, top_n=1)

For rizzo-flow, use model="rizzo-latest", its URL, 26 options including “none” and max_state_plus_question_tokens=8192 at --ctx 8192. Both servers were smoke-tested on 2026-09-30 with a two-tool example and were not benchmarked (protocol, summary); a small token window can make the planner lower card detail or split a choice. OpenAILogprobModel covers OpenAI-compatible endpoints that return top_logprobs.

Status and roadmap

Pre-alpha: search, deciders and the Pydantic AI integration work; the API may change.

  • Multi-turn agents: measure whole agent loops, where prompt caching and tool reveal interact.
  • OpenAI Decisions API: announced at DevDay 2026 in limited preview, it answers questions with predefined answers in about 150 ms and is built on GPT-6 Luna. It is the same kind of model as Jev, and the next decider to benchmark against the Luna numbers above once it is available.

Reproduce

The benchmark is not on PyPI. Clone the repository, run the offline checks and regenerate the figures from the published summaries, without API calls:

git clone https://github.com/dfm88/toolhunch && cd toolhunch && uv sync --all-packages
uv run pytest -q
uv run toolhunch-bench readme-charts

The technical page gives report regeneration and paid rerun commands: about $2 for the main 44,453-tool decision calls, $1 for the small-catalog run and under $1 for the GPT-6 Luna runs. Every paid command prints an estimate first. See AGENTS.md for development conventions.

License

MIT.

Metadata

Release files for toolhunch 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for toolhunch 0.1.0
File Size Uploaded
toolhunch-0.1.0.tar.gz 47.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for toolhunch 0.1.0
File Interpreter ABI Platform
toolhunch-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 108.2 kB

Release files / toolhunch-0.1.0.tar.gz

Download URL toolhunch-0.1.0.tar.gz
Size 47.8 kB
Tags Source
SHA-256 checksum
How to use checksums
628853e87eeadea884d9dd082c45eb7bead7704201506a09f500b878056c63ee
BLAKE2b-256 checksum
How to use checksums
449edf1dd51ca0a8aee16e16aa8f1d185a7479d5ba796c4f7446e4e70ea077b4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.4 {"installer":{"name":"uv","version":"0.12.4","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release files / toolhunch-0.1.0-py3-none-any.whl

Download URL toolhunch-0.1.0-py3-none-any.whl
Size 60.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
8607591540a3788665ae2c52b36de1f2e04005ea227cb281235d46f586ba475c
BLAKE2b-256 checksum
How to use checksums
b3fceb81f9fbe18f6d36c1616992c3e460de1dba4f4dabd55924ec1699c3d2df
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.4 {"installer":{"name":"uv","version":"0.12.4","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page