ToolRecall — Deterministic Execution Layer for Agent Tools
🌐 toolrecall.dev — documentation, benchmarks, downloads
You run agents. Every session spawns its own MCP servers, every test run hits live APIs, every tool call is unrepeatable, and your agent can read ~/.ssh if it feels like it.
ToolRecall is one shared daemon that pools your MCP servers, records and replays tool results, caches repeated API calls, and enforces filesystem/terminal policy for any agent framework.
−36% on a 450-turn session — verified with separate billed API keys, not estimates. Additive to any model's prefix caching. Under 200 KB install. Python 3.11+ stdlib only.
⚠️ Who this is for: ToolRecall's file cache shines for stateless agents (Hermes, OpenCode, Cline, Google ADK) — agents with limited or no built-in context management. If your agent already manages its own context (Claude Code, Cursor, Codex CLI), the forward proxy and MCP multiplexer still save real money, but file caching through MCP may increase costs. See Agent Compatibility.
pipx install toolrecall
toolrecall setup # One-shot: config -> systemd -> daemon start
Zero config mode: Every
toolrecallcommand auto-starts the daemon if it isn't running. You never need to think about it.
Why ToolRecall — stop paying for tokens you already saw
Two capabilities are the reason to run it. Both are measured, both target the same waste: your agent re-encountering content that's already in its context.
| Capability | What it does | Measured |
|---|---|---|
| Context Tracker | Tells the agent which files it only read (vs edited) are safe to drop from its context window, so a long session stops growing larger every turn. By keeping your context small it helps prefix and non-prefix models alike. | 9.5× fewer request tokens per turn and 7.4× longer sessions before the context wall (measured on real runs, prefix-caching model) — Context Tracker |
| Input Dedup Hook | Agents re-paste the same file contents into the message over and over. The hook spots a repeat and sends a short "same content as before" placeholder instead of the full copy again — you pay for each block once, and the model still sees it. | −32.3% input tokens / −30% billed cost on real coding tasks (SWE-bench Lite, billing-verified) — Dedup Hook |
Quickstart — Forward Proxy (10 seconds, any agent)
Set one environment variable and all your agent's API calls route through the proxy. Cache hit = zero tokens billed.
export OPENAI_BASE_URL=http://localhost:8569/v1
# or for Anthropic-compatible agents:
export ANTHROPIC_BASE_URL=http://localhost:8569
That's it. Every identical API call across sessions costs $0. Works with any OpenAI/Anthropic-compatible agent — no per-agent config.
What that saves (billed API keys, not estimates):
| Workload | Without TR | With TR | Savings |
|---|---|---|---|
| Bugfix (450 turns, both completed) | $8.83 | $5.62 | −36% |
| Review (200 turns, naive died at 112) | $2.53 | $1.26 | −50% |
| Analysis (400 turns, naive died at 145) | — | Completed | full run |
Separate OpenRouter keys per arm. Full methodology: Benchmark. No token estimates — only billed dollars from the provider dashboard.
See Forward Proxy for configuration and provider routing.
Option B — MCP Bridge (for agents that support MCP)
If your agent uses MCP, register one server:
{
"mcpServers": {
"toolrecall": {
"command": "toolrecall",
"args": ["mcp"]
}
}
}
# ~/.config/toolrecall/toolrecall.toml
[mcp_multiplex]
servers = ["time", "github", "fetch"]
Now every MCP-capable agent shares one warm pool of servers. No more N×M cold Node processes.
Features: lazy loading, idle timeout, failure isolation, auto-resolution. See MCP Multiplexer.
What ToolRecall Does
| Feature | What it solves |
|---|---|
| Forward API Proxy | Cache API responses by body hash — hit = zero tokens billed. −36% on 450-turn sessions (billed keys). Additive to any provider's prefix caching. |
| Context Tracker | Track dirty/clean files, auto-hint agents what to drop from context. 9.5× fewer request tokens per turn, 7.4× longer sessions. |
| Input Dedup Hook | LiteLLM hook removes repeated file content before billing — −32.3% prompt tokens / −30% cost on SWE-bench Lite (billing-verified). |
| Replay Mode | Record agent sessions, replay deterministically in CI |
| Security Gate | Path allowlist, terminal policy, sensitive-file blocklist — any agent |
| MCP Multiplexer | One shared pool of MCP servers instead of N processes per agent session |
| File / Terminal Cache | Reduce redundant reads within a turn; bounded context growth for stateless agents without built-in context management |
| Framework Adapters | Drop-in wrappers for ADK, LangChain, herdr, Odysseus, LiteLLM |
Full detail in Architecture.
Context Tracker
TL;DR: ToolRecall caches file reads so re-reading is instant (~0.1ms). The Context Tracker adds dirty-file awareness: the agent drops old file content from its context window and re-reads on demand from cache — keeping context bounded and breaking the O(n²) attention-cost snowball.
Every turn, an agent appends all prior tool output to its history, and the LLM computes attention over the whole sequence — O(n²) in tokens. ToolRecall caches the I/O but not the context window; without help, file content the agent read ten turns ago still sits in context as redundant overhead.
The Context Tracker records which files were written (made dirty) since a user-defined checkpoint. Clean files (read but not modified) are safe to drop: a cache hit returns the same content in ~0.1ms, so dropping costs nothing.
| Category | Meaning | Agent action |
|---|---|---|
| Dirty | Modified by the agent since checkpoint | Keep — uncommitted work |
| Clean | Read but not modified | Drop from context, re-read from cache if needed |
| Untracked | Never read | Not in context — no action |
Available in the MCP Bridge as five tools — context_set_checkpoint, context_get_dirty, context_get_stats, context_reset, context_get_hint. The bridge auto-appends a hint to every tool response telling the agent which clean files to drop, so no agent-side config is required beyond the pattern.
Measured, not modeled. On real runs (Hermes agent, DeepSeek V4 Flash — a model with prefix caching already on), the tracker sent 9.5× fewer request tokens per turn (8,077 vs 76,430 at turn 10) and the session ran 7.4× longer (140 vs 19 turns) before hitting the context wall. Because it shrinks the context window itself — not just what the provider caches — the benefit holds for prefix and non-prefix models alike.
How far it goes depends on the workload: an agent that rewrites whole files every turn saves less than one that re-reads the same files. The table below is the modeled ceiling — it assumes an idealized re-read-heavy agent that drops every clean file each turn (~7 files/turn):
| Agents × Turns | Baseline (attention pairs) | With Tracker (every-turn drops) | Reduction |
|---|---|---|---|
| 1 × 30 | 1.27T | 127B | 90% |
| 5 × 30 | 6.35T | 635B | 90% |
| 10 × 30 | 12.7T | 1.27T | 90% |
| 20 × 30 | 25.4T | 2.54T | 90% |
| 10 × 100 | 171T | 4.23T | 97.5% |
Read this number carefully: the ~90% (up to 97.5%) is a modeled upper bound for an idealized re-read-heavy session — not a measured benchmark. The measured headline is the 9.5× fewer tokens / 7.4× longer endurance above. Two caveats hold either way: the daemon can't force the agent to drop — it provides the data, the agent must act on it — and append-only harnesses (Claude Code, Cursor) can't use the tracker at all. See Agent Compatibility.
Full detail: Context Tracker · Agent integration · Stale-file detection
Recall Tier (opt-in, experimental)
TL;DR: Keep only a pointer to output you probably won't need; restore it on the rare turn you're wrong.
Most blocks an agent sees are reproducible — a file at a path, a command, an API call keyed by request hash. Re-reading is byte-identical, so dropping the content is free. Non-reproducible content — a one-shot API response, a live web snapshot, ephemeral tool output that would never come back identical — can't be re-fetched, so it has historically been forced to sit in the context window at full token cost, every turn, just in case it's needed again.
The Recall Tier lets an agent do something else entirely: evict it by default, restore on demand. It works like a scratchpad that's always there but only billed when opened:
recall_storepersists the raw content out-of-band and returns a tiny deterministicnode_idpointer.- The agent keeps only the pointer in context.
recall_get(node_id)restores the raw bytes on demand — the exact lossless-recoverable eviction contract the Context Tracker already gives reproducible files, extended to the non-reproducible tail.
How this is different from a cache hit. A cache hit avoids an LLM round-trip; a recall_get is one, because it re-inserts the bytes. The win is never in the turn you call get — it's in all the turns you don't have to. A 10k-token block stored on turn 5 and never re-read costs 0 for the remaining 195 turns of a 200-turn task instead of 1.95M token-turns of sitting in context.
When to use it (and when not to)
| Use it when… | Don't bother when… |
|---|---|
| The block is a non-reproducible one-shot (web/API response, ephemeral output) | Content that is reproducible — the normal cache already handles it, losslessly |
| You need to bound context size but can't depend on a re-fetch | You'll definitely need the block again soon (eviction is only worth it if eviction usually stands) |
| The pointer is meaningfully smaller than the content | The content is tiny to begin with |
The feature is off by default and adds zero runtime dependencies. Enable with [recall].enabled = true (or TOOLRECALL_RECALL_ENABLED=true). The default TTL is 0 = never expire — set [recall].ttl (or TOOLRECALL_RECALL_TTL) in seconds to bound how long entries live. Expired entries are treated as cache misses, purged lazily on read, and swept from disk by the regular GC cycle; they are never returned and never count as cached.
Status: experimental. The tier works and is tested (roundtrip, dedup, TTL expiry, lazy purge, GC sweep), but it is not yet driven by any first-party shim or adapter — no agent calls
recall_storeautomatically today. You opt into it explicitly (via CLI or MCP) or not at all. See docs/RECALL_TIER.md for the contract.
Accounting and honesty
Every recall_get hit records the entry's token count in a dedicated recall sink in cache_status, tracked separately from file-cache hits. This is a "bytes served" counter, not a savings claim. It means "this much content was restored via the recall tier" — useful for understanding what the pool is doing, not a number that belongs in a cost-savings banner. Real savings from eviction-only use (never restoring) are invisible to accounting by definition: the win is that the context window stayed small, and there's nothing to count.
Full detail: Recall Tier
Input Dedup Hook
TL;DR: AI agents re-read the same files over and over, and every read pastes that file into the message they send to the model. This hook removes the repeated copies before they're billed — cutting input tokens with cost measured, not estimated.
Why this matters, in plain English. When an agent works on a task it re-reads the same files many times, and each read sends that file's contents to the model again. On a long session the same file can be sent five, ten, twenty times — and normal billing charges you for every copy. The hook keeps the first copy (so the model still has the information, and the provider's own caching stays intact) and turns every later repeat into a short note like "same content as before — see message 4." You pay for each block once, not once per read. How much you save depends on your agent: one that re-reads the same files a lot saves the most.
| Metric (80 req/arm, SWE-bench Lite × 8 turns) | WITH dedup | WITHOUT dedup | Saved |
|---|---|---|---|
| Total prompt tokens | 282,688 | 417,256 | 134,568 (−32.3%) |
| Billed cost (OpenRouter) | $0.0134 | $0.0191 | $0.0057 (−30.0%) |
Prefix caching preserved. Effective per-token rate is near-identical between arms (Δ $0.0015/M) — the keep-first design stubs only later duplicates, so each block's first occurrence is byte-identical to the non-dedup arm. The honest shape of the method: it saves on re-reads (savings appear from turn 4, growing to −49.9% by turn 8), not first reads.
Honesty (stated explicitly): token savings are billing-verified; task-quality is not. A SWE-bench pass@1 A/B was attempted but inconclusive (the baseline model scored 0 on the chosen tasks even in isolation), so the defensible claim is: "the hook removes wasted input tokens; its effect on task success is unverified." Savings are also workload-dependent — an agent that rewrites whole files each turn saves less.
Zero-trust customer triage: bench/litellm_dedup/measure_duplicates.py measures your own duplicate ratio from a JSONL export of your request bodies, entirely inside your perimeter, no network, no API key — so you know what you'd save before any pilot. It reports volume stubbable, deliberately not billed-$, because real savings depend on prefix-cache economics.
Full benchmark & methodology: LiteLLM Dedup Benchmark · ready-to-use config: litellm-proxy-config.yaml
How It Works
flowchart LR
subgraph Agents
A1["Claude Code"]
A2["Cursor"]
A3["Aider"]
A4["Hermes"]
end
subgraph Daemon["ToolRecall Daemon"]
MP["MCP Multiplexer"]
CA["Cache (LRU + SQLite)"]
SG["Security Gate"]
FP["Forward Proxy"]
end
subgraph OS["OS Layer"]
FS["Filesystem / Network"]
end
A1 --> MP
A2 --> MP
A3 --> MP
A4 --> MP
MP --> CA
MP --> SG
MP <--> FS
A1 --> FP
FP --> CA
FP <--> FS
One daemon, five access paths: Python client, MCP bridge, HTTP bridge, forward proxy, OS-level shim. All share one cache, one security gate, one multiplexer. See Architecture.
When To Use It
| You want this... | Use this... | Works for |
|---|---|---|
| $0 dev loops — repeated API calls cost nothing | Forward Proxy | Any agent |
| Lower per-turn cost on long agent sessions | Context Tracker | Stateless agents (Hermes, Cline, ADK) |
| Cached file reads, lower context bloat | File / Terminal Cache | Stateless agents — bounded context growth. Not for agents with built-in context management (Claude Code, Cursor, Codex CLI) |
| Deterministic CI tests for agent behavior | Replay Mode | Any agent |
| Guardrails between agents and your machine | Security Gate | Any agent |
| Warm MCP servers across sessions | MCP Multiplexer | Any agent |
| All of the above | toolrecall setup then add the MCP bridge |
See per-agent notes |
Installation
One-time setup
pipx install toolrecall # or: uv tool install toolrecall
# or: pip install toolrecall (inside a venv)
toolrecall setup # config -> systemd service -> daemon start
PATH check: After installation, make sure
toolrecallis on your$PATH.
pipxputs binaries in~/.local/bin/,uv tool installin~/.local/share/uv/tools/.
Iftoolrecallisn't found, add the right directory to your PATH or reinstall inside the venv your agent uses.
Shim in the right venv:
toolrecall shim --installinstalls the.pthshim into the current Python environment. If you installed viapipxoruv tool install, the shim goes into that isolated environment — not your agent's venv. The agent won't see it.toolrecall setupauto-detects common agent venvs and installs the shim there too.toolrecall shim --install --allscans for agent venvs (Hermes, OpenCode) and installs into all of them at once. If you need to target a specific venv manually:toolrecall shim --install --venv ~/.hermes/hermes-agent/venv toolrecall shim --install --venv ~/.local/share/uv/tools/hermes-agentThe
toolrecallpackage must also be installed in that venv (import toolrecallmust work).
Agent type → mechanism: pick based on what your agent is:
| Agent type | Mechanism | Needs toolrecall in the venv? |
Setup action |
|---|---|---|---|
| Python agent, own venv (Hermes, Codex, OpenCode, Cline) | .pth shim in the agent venv (transparent open()/subprocess cache) |
yes | toolrecall shim --install --venv <path> (opt-in) |
| Non-Python agent (Claude Code, Cursor, Cline, Windsurf) | MCP bridge (toolrecall mcp) |
n/a | register an MCP server |
| System python / global interpreter | shim in user site-packages | yes | toolrecall shim --install |
The
.pthshim is opt-in, default off —toolrecall shim --installor--venv/--allprompts before enabling. Use--yesto skip the prompt. Verify withtoolrecall shim --status [--all](printsprobe: passonly when the shim actually imports in that venv from a neutral cwd).
toolrecall setup creates ~/.config/toolrecall/toolrecall.toml with default-deny security, generates a systemd user unit, and starts the daemon. After this, every toolrecall command "just works".
Daemon auto-start fallback: systemd -> os.fork() -> DETACHED_PROCESS (Linux -> Docker/macOS -> Windows).
Per-agent integration
| Method | How | When to use |
|---|---|---|
| MCP Bridge | toolrecall mcp in agent's MCP config |
Any MCP-capable agent (recommended) |
| Go Client (tr) | tr read file.py, tr term "hostname" |
Shell scripts, CI, any language |
| Python Shim | toolrecall shim --install |
Every Python process auto-caches open/subprocess |
| Python Client | from toolrecall.client import cached_read |
Direct embedding in Python code |
| HTTP Bridge | toolrecall serve on :8569 |
Any HTTP client (curl, Go, Rust...) |
| Forward Proxy | Set OPENAI_BASE_URL=http://localhost:8569/v1 |
Cache API responses, zero tokens on hit |
Extra storage backends
pip install toolrecall[libsql] # libSQL local backend
pip install toolrecall[libsql-sync] # libSQL + Turso Cloud sync
CLI Quick Reference
toolrecall setup One-shot: config + systemd + daemon start [required once]
toolrecall status Cache status and stats [auto-starts]
toolrecall stats Detailed cache statistics (JSON) [auto-starts]
toolrecall invalidate Clear all caches [auto-starts]
toolrecall mcp Start MCP Bridge [auto-starts]
toolrecall serve Forward proxy (cache API responses) [auto-starts]
toolrecall serve --9000 Custom port forward proxy
toolrecall replay Record/replay agent sessions
toolrecall shim --install [--venv <path>|--all] Install OS-level cache shim (.pth) — opt-in
toolrecall shim --status [--venv <path>|--all] Check shim presence + import probe
toolrecall shim --uninstall [--venv <path>|--all] Remove .pth shim
toolrecall turso Turso Cloud sync: init, enable, disable, status
toolrecall init Create default config.toml and .env
toolrecall config-set Set a config value
toolrecall context Inspect Context Tracker / Recall Tier [auto-starts]
toolrecall index Index knowledge DB (FTS5 search) [not file cache pre-warm]
toolrecall index-memory Index agent memory stores
toolrecall index-dir Index a directory for FTS5 search [not file cache pre-warm]
Knowledge indexing ≠ cache warming:
toolrecall index*commands build an FTS5 search index for knowledge retrieval (docs_search()). They do NOT pre-warm the file/terminal/API response cache. The daemon's file cache warms naturally as the agent reads files — no separate command needed.
Full reference: CLI.md
Configuration
# ~/.config/toolrecall/toolrecall.toml
[mcp]
allowed_paths = ["/home/user/projects"] # Default-deny!
allow_terminal = false
[cache]
terminal_default_ttl = 60
[mcp_multiplex]
enabled = true
servers = ["time", "sequential-thinking"]
[recall]
enabled = false # opt-in Recall Tier (lossless-recoverable eviction)
[forward_proxy]
# Starts on :8569 automatically with the daemon
TOOLRECALL_* env vars override TOML. Full reference: Configuration Reference
Platform Support
| Platform | Transport | Status |
|---|---|---|
| Linux | Unix Domain Sockets | Tested in CI |
| macOS | Unix Domain Sockets | Should work (POSIX) |
| Windows | TCP localhost:8568 | Experimental |
Documentation
- toolrecall.dev — documentation portal, benchmarks, downloads
- Architecture — system design, components, data flow, token costs
- MCP Multiplexer — daemon-managed MCP server pool
- Forward Proxy — API response caching, provider list, auth routing
- Replay Mode — record/replay tool calls for deterministic CI
- Security Architecture — policy gate, trust boundary
- Agent Compatibility — per-agent value, config, caveats
- Benchmark — three-arm controlled measurement (naive vs prefix vs toolrecall), context efficiency, billed cost
- LiteLLM Dedup Benchmark — gateway dedup hook: −32.3% prompt tokens / −30% cost, billing-verified
- LiteLLM Proxy Example Config — ready-to-use
litellm_settings.callbackshook wiring - Go Dedup Reference — pure-Go request-level
dedup_messagesreference implementation - Bench Infrastructure — reproduce the three-arm benchmark
- Test Suite — test runner documentation
- CLI Reference — all subcommands
- Configuration Reference — config.toml, env vars
- Context Stale — provably stale files in agent conversations
- Context Tracker — checkpoint-based dirty-file tracking
- AGENTS.md — agent instructions for MCP context tracker integration
- Recall Tier — opt-in lossless-recoverable eviction for non-reproducible content
- Testing Guide — test philosophy, per-file coverage
- How It Works — quick technical overview
- libSQL Backend — multi-writer, vector search, cloud sync
- Docker Deployment — containerized stack
- Troubleshooting — common fixes
- Changelog — version history
- Go Client — standalone
trbinary for any language/shell - Agent Configs — ready-to-use MCP configs for popular agents
- Framework Adapters:
- Google ADK —
@cached_tooldecorator + forward proxy - LangChain / LangGraph —
ToolRecallCacheBaseCache + callback - herdr —
trbinary + MCP bridge for any pane - Odysseus —
cached_tooldecorator + MCP server caching
- Google ADK —
- Hermes Transparent Cache — auto-patching for Hermes
- Normalizer — cache key normalization, deterministic JSON
- Knowledge DB — FTS5 indexing guide
- Real-Agent Benchmark — edit-heavy session results
- Appendix — comparison tables, OSI model, ROI, audit
Contributing
git clone https://github.com/whiskybeer/toolrecall.git
cd toolrecall
make setup # one-time dev deps
make test # run tests
make check # lint + format
See Testing Guide and Makefile.
Uninstall
systemctl --user stop toolrecall-daemon
systemctl --user disable toolrecall-daemon
pipx uninstall toolrecall
rm -rf ~/.toolrecall ~/.config/toolrecall
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file toolrecall-0.8.20.tar.gz.
File metadata
- Download URL: toolrecall-0.8.20.tar.gz
- Upload date:
- Size: 392.1 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.11.2
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
c2ff15d3622f972858fa1155ecdb8ced6abfe807fced87a0dd8994eac4250919
|
|
| MD5 |
6a4f30a3728937c96dc3d4ba4fde3d06
|
|
| BLAKE2b-256 |
15d47d98ec78b385529d3aba4fb35d62556828ea681ebffdaf56a2040181b995
|
File details
Details for the file toolrecall-0.8.20-py3-none-any.whl.
File metadata
- Download URL: toolrecall-0.8.20-py3-none-any.whl
- Upload date:
- Size: 254.2 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.11.2
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
0eeeb91e74b7d87f2272761e9df2514d8de6e9da92d73a6edc266bc2481ceed1
|
|
| MD5 |
fa0d8843340605eb8f176234bf9933af
|
|
| BLAKE2b-256 |
e1ae734dc23963520a5b8472d4877acac65142fb7bb118c8bf2e8e3ad4107ff0
|