Skip to main content

Graph-structured tool retrieval for LLM agents — zero-dependency, ontology-aware hybrid search

Project description

graph-tool-call

LLM agents can't fit thousands of tool definitions into context.
Vector search finds similar tools, but misses the workflow they belong to.
graph-tool-call builds a tool graph and retrieves the right chain — not just one match.


7-case commerce regression Target only graph-tool-call
Required-producer recall 14.3% 100%
Candidate plan coverage 47.6% 100%
Target Recall@5 - 100%

Deterministic, model-free engine benchmark. Case-level evidence and full methodology.


graph-tool-call demo

PyPI Docs License: MIT Python 3.10+ CI Zero Dependencies

English · 한국어 · 中文 · 日本語


Table of Contents

Why

LLM agents need tools. But as tool count grows, two things break:

  1. Context overflow — Large catalogs spend model context on tools that cannot help the current request.
  2. Target-only retrieval misses prerequisites — Searching "refund my order" finds refundOrder, but that tool requires an order_id. The usable flow starts with the tool that produces that ID.

graph-tool-call solves both. It models tool relationships as a graph, retrieves multi-step workflows via hybrid search (BM25 + graph traversal + embedding + MCP annotations), and admits only the schemas that fit an explicit planner token budget.

Scenario Vector-only graph-tool-call
"refund my order" Returns refundOrder findOrdersByEmail + refundOrder from typed contract evidence
"read and save file" Returns read_file read_file + write_file (COMPLEMENTARY relation)
"delete old records" Returns any tool matching "delete" Destructive tools ranked first via MCP annotations
"now cancel it" (after listing orders) No context from history Demotes used tools, boosts next-step tools
Multiple Swagger specs with overlapping tools Duplicate tools in results Cross-source auto-deduplication
1,200 API endpoints Slow, noisy results Categorized + graph traversal for precise retrieval

How it works

OpenAPI / MCP / Python functions → Ingest → Build tool graph → Hybrid retrieve → Agent

Example — User says "cancel my order and process a refund"

Vector search finds cancelOrder. But the actual workflow is:

                    ┌──────────┐
          PRECEDES  │listOrders│  PRECEDES
         ┌─────────┤          ├──────────┐
         ▼         └──────────┘          ▼
   ┌──────────┐                    ┌───────────┐
   │ getOrder │                    │cancelOrder│
   └──────────┘                    └─────┬─────┘
                                        │ COMPLEMENTARY
                                        ▼
                                 ┌──────────────┐
                                 │processRefund │
                                 └──────────────┘

graph-tool-call returns the entire chain, not just one tool. Retrieval combines four signals via weighted Reciprocal Rank Fusion (wRRF):

  • BM25 — keyword matching
  • Graph traversal — relation-based expansion (PRECEDES, REQUIRES, COMPLEMENTARY)
  • Embedding similarity — semantic search (optional, any provider)
  • MCP annotations — read-only / destructive / idempotent hints

Installation

The core package has zero dependencies — just Python standard library. Install only what you need:

pip install graph-tool-call                # core (BM25 + graph) — no dependencies
pip install graph-tool-call[embedding]     # + embedding, cross-encoder reranker
pip install graph-tool-call[openapi]       # + YAML support for OpenAPI specs
pip install graph-tool-call[mcp]           # + MCP server / proxy mode
pip install graph-tool-call[all]           # everything
All extras
Extra Installs When to use
openapi pyyaml YAML OpenAPI specs
embedding numpy Semantic search (connect to Ollama/OpenAI/vLLM)
embedding-local numpy, sentence-transformers Local sentence-transformers models
similarity rapidfuzz Duplicate detection
langchain langchain-core LangChain integration
visualization pyvis, networkx HTML graph export, GraphML
dashboard dash, dash-cytoscape Interactive dashboard
lint ai-api-lint Auto-fix bad API specs
mcp mcp MCP server / proxy mode

Quick Start

Try it in 30 seconds (no install)

uvx graph-tool-call demo dependency-chain
Query: "Refund the order for alice@example.com"

Selected target:
  refundOrder(order_id)

Required producer:
  findOrdersByEmail(email) -> order_id
  evidence: api_contract, openapi_link

Execution order:
  1. findOrdersByEmail
  2. refundOrder

Python API

from graph_tool_call import ToolGraph

# Build a tool graph from the official Petstore API
tg = ToolGraph.from_url(
    "https://petstore3.swagger.io/api/v3/openapi.json",
    cache="petstore.json",
)
print(tg)
# → ToolGraph(tools=19, nodes=22, edges=100)

# Search for tools
tools = tg.retrieve("create a new pet", top_k=5)
for t in tools:
    print(f"{t.name}: {t.description}")

# Search with workflow guidance
results = tg.retrieve_with_scores("process an order", top_k=5)
for r in results:
    print(f"{r.tool.name} [{r.confidence}]")
    for rel in r.relations:
        print(f"  → {rel.hint}")

# Execute an OpenAPI tool directly
result = tg.execute(
    "addPet", {"name": "Buddy", "status": "available"},
    base_url="https://petstore3.swagger.io/api/v3",
)

OpenAPI ingest keeps execution metadata such as parameter locations, content types, candidate request-body fields, examples, security schemes, response catalogs, and error responses under tool.metadata["openapi"]. The HTTP executor uses those facts for parameter serialization and JSON/form/multipart request bodies, and returns matched response metadata for success/error diagnostics. HttpExecutor.validate_request() provides missing-required, missing-security, invalid-argument, and unused-argument preflight diagnostics without network I/O; see docs/api-reference.md. Request contracts exclude OpenAPI readOnly fields and response contracts exclude writeOnly fields, keeping generated tool inputs and graph edges aligned with the direction in which each field can actually travel. OpenAPI oneOf / anyOf request and response schemas are read as a union of branch fields with branch evidence preserved, so graph construction and request validation do not silently drop every branch after the first one. Discriminator mappings and JSON Schema const values are preserved as branch selection evidence; if a request chooses a discriminator value, preflight diagnostics can report the missing fields for that selected branch only. When a Swagger/OpenAPI document declares only a weak object schema but provides concrete request or response examples, ingest derives additive schema_inferred_from="example" contract rows from those examples. Common response envelopes such as code/message/data also record wrapper, collection, and value-path aliases, so XGEN-style adapters can recover produced values from either raw OpenAPI bodies or normalized body wrappers. Retrieval indexes tool descriptions, tags, parameter names/descriptions, AI metadata, and curated/indexable IO fields. Promoted raw OpenAPI contract rows remain planning-first by default, so large Swagger specs do not flood BM25 with common identifier fields unless the caller explicitly opts in.

Workflow planning

plan_workflow() returns ordered execution chains with prerequisites — reducing agent round-trips from 3-4 to 1.

plan = tg.plan_workflow("process a refund")
for step in plan.steps:
    print(f"{step.order}. {step.tool.name}{step.reason}")
# 1. getOrder      — prerequisite for requestRefund
# 2. requestRefund — primary action

plan.save("refund_workflow.json")

Edit, parameterize, and visualize workflows — see Direct API guide.

Other tool sources

# From an MCP server (HTTP JSON-RPC tools/list)
tg.ingest_mcp_server("https://mcp.example.com/mcp")

# From an MCP tool list (annotations preserved)
tg.ingest_mcp_tools(mcp_tools, server_name="filesystem")

# From Python callables (type hints + docstrings)
tg.ingest_functions([read_file, write_file])

MCP annotations (readOnlyHint, destructiveHint, idempotentHint, openWorldHint) are used as retrieval signals — query intent is automatically classified, and read queries prioritize read-only tools while delete queries prioritize destructive tools.


Choose your integration

graph-tool-call ships several integration patterns. Pick the one that matches your stack:

You're using... Pattern Token win Guide
Claude Code / Cursor / Windsurf MCP Proxy (aggregate N MCP servers → 3 meta-tools) ~1,200 tok/turn docs/integrations/mcp-proxy.md
Any MCP-compatible client MCP Server (single source as MCP) varies docs/integrations/mcp-server.md
LangChain / LangGraph (50+ tools) Gateway tools (N tools → 2 meta-tools) 92% docs/integrations/langchain.md
OpenAI / Anthropic SDK (existing code) Middleware (1-line monkey-patch) 76–91% docs/integrations/middleware.md
Direct control over retrieval Python API (retrieve() + format adapter) varies docs/integrations/direct-api.md

MCP Proxy (most common)

When you have many MCP servers, their tool names pile up in every LLM turn. Bundle them behind one server: 172 tools → 3 meta-tools.

# 1. Create ~/backends.json listing your MCP servers
# 2. Register the proxy with Claude Code
claude mcp add -s user tool-proxy -- \
  uvx "graph-tool-call[mcp]" proxy --config ~/backends.json

Full setup, passthrough mode, remote transport → MCP Proxy guide.

LangChain Gateway

from graph_tool_call.langchain import create_gateway_tools

# 62 tools from Slack, GitHub, Jira, MS365...
gateway = create_gateway_tools(all_tools, top_k=10)
# → [search_tools, call_tool] — only 2 tools in context

agent = create_react_agent(model=llm, tools=gateway)

92% token reduction vs binding all 62 tools. See LangChain guide for auto-filter and manual patterns.

SDK middleware

from graph_tool_call.middleware import patch_openai

patch_openai(client, graph=tg, top_k=5)  # ← add this one line

# Existing code unchanged — 248 tools go in, only 5 relevant ones are sent
response = client.chat.completions.create(
    model="gpt-4o",
    tools=all_248_tools,
    messages=messages,
)

Also works with Anthropic via patch_anthropic. See Middleware guide.


Benchmark

The release headline uses a deterministic seven-case commerce regression. It asks whether graph expansion adds every required producer after the target has been selected. No LLM or external API is involved.

Metric Target only Graph with producers
Required-producer recall 0.1429 1.0000
Candidate plan coverage 0.4762 1.0000
Candidate binding support 0.1429 1.0000
Unneeded expansion cases 0 0

The checked-in case-level artifact contains input hashes, expected targets and producers, candidate lists, metrics, limitations, and replay commands. Historical model-in-the-loop results remain in the full benchmark document but are not used as the current release headline.

→ Full results (pipeline / retrieval-only / competitive / 1068-scale / 200-tool LangChain agent across GPT and Claude): docs/benchmarks.md

# Reproduce the release claim
make launch-evidence
make launch-evidence-check

Optional BFCL-derived retrieval check

For a public-data sanity check, graph-tool-call can also run a deterministic retrieval benchmark over the official BFCL v4 function-calling JSONL files. This is not the BFCL leaderboard model AST score; it only asks whether the ground-truth function names land in the retrieved top-K.

make bfcl-benchmark

An experimental native tool-call loop is also available when you want to attach a real model to BFCL data through graph-tool-call retrieval:

make bfcl-llm-benchmark

It can optionally use the official bfcl-eval AST checker when that package is installed in an isolated benchmark environment, and a sweep runner is available for row-vs-retrieved / top-K comparisons. Full model-in-the-loop runs support case caching, repeat-safe cache namespaces, concurrency, progress output, and BFCL-compatible result JSONL export for local official-checker reruns. Current qwen3.6 full numbers are local BFCL-compatible evidence, not a BFCL leaderboard claim. Detailed methodology, commands, limitations, and current numbers live in docs/benchmarks.md.

XGEN-style quality checks

For API Collection / Planflow work, there are three focused checks: a deterministic engine benchmark, a live large-OpenAPI scale acceptance check, and a BFCL-style model-in-the-loop benchmark.

Benchmark Model used What it evaluates
make paper-corpus-check none public OpenAPI/GraphQL/MCP corpus hashes, licenses, family splits, annotations, and ingest conformance
make paper-corpus-claim-check none stricter paper gate, including independent annotation-review coverage
make paper-adapter-conformance none request/response/auth/execution/IO-contract preservation, deterministic replay, and structured unsupported diagnostics
make paper-baseline-run pinned E5 encoder B-1 through B7 paired retrieval, token-budget, and confidence-interval artifact
make paper-graph-ablation pinned E5 encoder B4→B5→B6→B7 topology, typed-contract, selector, and producer-expansion deltas
make paper-producer-coverage pinned E5 encoder ground-truth-only producer contract, edge, path, seed, and failure-reason diagnostics
make paper-output-promotion pinned E5 encoder B6→B6a required-consumer-aligned output promotion and producer-edge coverage delta
make xgen-benchmark none graph-tool-call engine search, target selector exactness, producer expansion, plan synthesis across commerce/admin/workflow fixtures
make xgen-scale-acceptance none X2BEE-scale Swagger UI discovery, dedupe, ingest, graph build, Korean product-case search
make xgen-scale-sweep none one X2BEE-scale graph build, then top-K compression diagnostics for k=3,5,10
make xgen-scale-contract-ablation none one X2BEE-scale spec load, then baseline vs promoted OpenAPI contract signal comparison
make xgen-scale-028-gate-check REPORT=... none saved XGEN scale sweep artifact check for the stricter snapshot-provenance xgen-scale-0.28 profile
make bfcl-028-gate-check REPORT=... none saved BFCL sweep artifact check for the stricter paper-ready xgen-0.28 profile
make xgen-llm-benchmark CLI --model value whether that model actually calls search_tools and selects the right plan
make xgen-benchmark
make paper-corpus-check
make paper-adapter-conformance
make paper-graph-ablation
make paper-producer-coverage
make paper-output-promotion
# Expected to fail until an independent reviewer signs the corpus annotations.
make paper-corpus-claim-check
make xgen-scale-acceptance
make xgen-scale-sweep
MANIFEST=/tmp/gtc-x2bee-openapi-snapshot/manifest.json \
  GATE_PROFILE=xgen-scale-0.28 \
  make xgen-scale-sweep
make xgen-scale-028-gate-check REPORT=/tmp/gtc-x2bee-scale-snapshot-sweep.json
make xgen-scale-contract-ablation
make xgen-llm-benchmark
poetry run python -m benchmarks.xgen_tool_graph.llm_loop \
  --model qwen3.6-27b \
  --llm-url http://127.0.0.1:8000/v1 \
  --disable-thinking

Current scores, caveats, and model-specific notes are documented in docs/benchmarks.md. For the XGEN tool graph research direction, use docs/research/xgen-tool-graph-goals.md as the roadmap and docs/research/validation-loop.md as the day-to-day validation loop instead of running full model benchmarks after every change. Claims, public datasets, baselines, ablations, and submission gates for a research paper are defined separately in the canonical paper-readiness protocol.


Advanced Features

Embedding-based hybrid search

Add semantic search on top of BM25 + graph. No heavy dependencies needed — connect to any external embedding server.

tg.enable_embedding("ollama/qwen3-embedding:0.6b")        # Ollama (recommended)
tg.enable_embedding("openai/text-embedding-3-large")      # OpenAI
tg.enable_embedding("vllm/Qwen/Qwen3-Embedding-0.6B")     # vLLM
tg.enable_embedding("sentence-transformers/all-MiniLM-L6-v2")  # local
tg.enable_embedding(lambda texts: my_embed_fn(texts))     # custom callable

Weights are auto-rebalanced. See API reference for all provider forms.

Retrieval tuning

tg.enable_reranker()                                      # cross-encoder rerank
tg.enable_diversity(lambda_=0.7)                          # MMR diversity
tg.set_weights(keyword=0.2, graph=0.5, embedding=0.3, annotation=0.2)

History-aware retrieval

Pass previously called tools to demote them and boost next-step candidates.

tools = tg.retrieve("now cancel it", history=["listOrders", "getOrder"])
# → [cancelOrder, processRefund, ...]

Save / load (preserves embeddings + weights)

tg.save("my_graph.json")
tg = ToolGraph.load("my_graph.json")
# Or use cache= in from_url() for automatic save/load
tg = ToolGraph.from_url(url, cache="my_graph.json")

LLM-enhanced ontology

tg.auto_organize(llm="ollama/qwen2.5:7b")
tg.auto_organize(llm="litellm/claude-sonnet-4-20250514")
tg.auto_organize(llm=openai.OpenAI())

Builds richer categories, relations, and search keywords. Supports Ollama, OpenAI clients, litellm, and any callable. See API reference.

Other features

Feature API Docs
Duplicate detection across specs find_duplicates / merge_duplicates API ref
Conflict detection apply_conflicts API ref
Operational analysis analyze API ref
Interactive dashboard dashboard() API ref
HTML / GraphML / Cypher export export_html / export_graphml / export_cypher API ref
Auto-fix bad OpenAPI specs from_url(url, lint=True) ai-api-lint

Documentation

Doc Description
CLI reference All graph-tool-call CLI commands
Python API reference ToolGraph methods, helpers, middleware, LangChain
Integrations MCP server / proxy, LangChain, middleware, direct API
Benchmark results Full pipeline / retrieval / competitive / scale tables
Architecture System overview, pipeline layers, data model
Design notes Algorithm design — normalization, dependency detection, ontology
Research Competitive analysis, API scale data
Release checklist Release process, changelog flow

Contributing

Contributions are welcome.

git clone https://github.com/SonAIengine/graph-tool-call.git
cd graph-tool-call
pip install poetry pre-commit
poetry install --with dev --all-extras
pre-commit install   # auto-runs ruff on every commit

# Test, lint, benchmark
poetry run pytest -v
poetry run ruff check . && poetry run ruff format --check .
python -m benchmarks.run_benchmark -v

License

MIT

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

graph_tool_call-0.36.0.tar.gz (365.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

graph_tool_call-0.36.0-py3-none-any.whl (413.3 kB view details)

Uploaded Python 3

File details

Details for the file graph_tool_call-0.36.0.tar.gz.

File metadata

  • Download URL: graph_tool_call-0.36.0.tar.gz
  • Upload date:
  • Size: 365.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for graph_tool_call-0.36.0.tar.gz
Algorithm Hash digest
SHA256 a6d6f410a24579763dab5ebb1e06704e9ff536576bc4d55669b16e784d3a14e3
MD5 e9a465458efcc4f8348801991306f5d0
BLAKE2b-256 02210cdc34f4d6b9f2232f7cf1e62fdcf567d008e1680ba43f05f5df71c847fa

See more details on using hashes here.

Provenance

The following attestation bundles were made for graph_tool_call-0.36.0.tar.gz:

Publisher: publish.yml on SonAIengine/graph-tool-call

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file graph_tool_call-0.36.0-py3-none-any.whl.

File metadata

File hashes

Hashes for graph_tool_call-0.36.0-py3-none-any.whl
Algorithm Hash digest
SHA256 01a5e6347244d2df6501ebc05db59d7a2930cdd018e62379d498b3789aafe8dc
MD5 226ea59a9a6d87d59e5ce4986e172da0
BLAKE2b-256 28338a4458f42fb4f67f733180c930522fdde836342fe9554b4f490a523aef2c

See more details on using hashes here.

Provenance

The following attestation bundles were made for graph_tool_call-0.36.0-py3-none-any.whl:

Publisher: publish.yml on SonAIengine/graph-tool-call

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page