Skip to main content

Contract testing for LLM agent tool-calls — catch tool-calling regressions when a provider ships a new model version.

Project description

toolcontract

Contract testing for LLM agent tool-calls.

When a provider ships a new model version, an agent's tool-calling behavior can silently change — invented or dropped arguments, a different tool chosen for the same input, values that no longer match a schema your downstream code depends on. Nothing in a typical CI pipeline catches this before it reaches production.

toolcontract is built around one trigger event: a model version changed — did your agent's tool calls still do what you expect? Pin a golden set of expected tool-call trajectories, run them against a live model, get a pass/fail/inconclusive verdict and a diff. Think Pact for microservice contracts, or Percy for visual regressions, but for tool calls.

Status

v0.1, in active development. Contract model, comparators/matching engine, verification ledger, OpenAI/Anthropic/LiteLLM adapters, the CLI (run/accept/check-version), a real pytest plugin, and the LLM-judge semantic tier are all built and tested. Not yet published to PyPI.

Quick start

from toolcontract import ArgKind, ArgSpec, Contract, ExpectedCall
from toolcontract.adapters.base import ToolSchema
from toolcontract.adapters.openai_adapter import OpenAIAdapter
from toolcontract.runner.engine import run_contract

log_water = ToolSchema(
    name="log_water",
    description="Log that the user drank water",
    parameters={
        "type": "object",
        "properties": {"amount_ml": {"type": "number"}},
        "required": ["amount_ml"],
        "additionalProperties": False,
    },
    strict=True,
)

contract = Contract(
    id="log-water-basic",
    description="User mentions drinking water, agent logs it",
    input_messages=({"role": "user", "content": "I just drank a glass of water"},),
    expected_calls=(
        ExpectedCall(tool_name="log_water", args={"amount_ml": ArgSpec.exact(250, kind=ArgKind.NUMBER, tolerance=20)}),
    ),
)

result = run_contract(contract, OpenAIAdapter(), [log_water], "gpt-4o")
print(result.verdict)  # Verdict.PASS / FAIL / INCONCLUSIVE

Or from the CLI, against a directory of compiled contracts:

toolcontract run contracts/ --provider openai --model gpt-4o --tools tools.json
toolcontract accept contracts/one.json --provider openai --model gpt-4o   # promote a new baseline
toolcontract check-version contracts/ --provider openai --model gpt-4o-2027 --ledger .toolcontract_ledger.json

See examples/tap_health_agent_get_meal_suggestions.py for a real contract authored against a production tool schema, including a custom comparator and an optional argument.

Providers

  • OpenAI, Anthropic — native adapters, pip install "toolcontract[openai]" / [anthropic].
  • Any OpenAI-compatible endpoint — a self-hosted LiteLLM proxy, vLLM, Ollama, OpenRouter, Azure OpenAI — already works today with zero new code:
    import openai
    from toolcontract.adapters.openai_adapter import OpenAIAdapter
    
    adapter = OpenAIAdapter(client=openai.OpenAI(base_url="http://localhost:4000", api_key="sk-..."))
    
  • LiteLLM (direct SDK routing)pip install "toolcontract[litellm]", then LiteLLMAdapter() with a provider-prefixed model string ("cerebras/zai-glm-4.7", "gemini/gemini-2.0-flash", and 100+ others LiteLLM supports). Capability (whether the underlying provider actually honors strict/grammar-constrained tool calling) is queried per-model at call time, not assumed statically — see adapters/litellm_adapter.py's module docstring for why, and for the honest caveat about how precise that signal currently is.
  • Missing your provider? See CONTRIBUTING.md — adding a native adapter is the single most valuable contribution this project can take right now.

Non-goals

This is deliberately not a general agent evaluation framework. If a request is really about one of the following, a different tool is the right fit — agentevals, DeepEval, Ragas, PydanticAI Evals, LangSmith, or Phoenix all already cover this territory well:

  • RAG retrieval or answer-quality evaluation
  • Hallucination detection or general output-quality scoring
  • Multi-turn simulated-user agent evaluation
  • General-purpose "score my agent" metrics unrelated to tool-call structure

toolcontract owns exactly one job: did this model version's tool-calling behavior change against a pinned contract. Everything in scope should trace back to that job — anything that doesn't belongs in a different package.

Design principles (from the architecture review)

  • Core engine is framework-agnostic. The pytest plugin is a thin consumer of a plain Python/CLI-callable core — not the other way around.
  • Three-valued verdicts. PASS / FAIL / INCONCLUSIVE. An argument that can't be resolved structurally (e.g. free-form natural language) reports INCONCLUSIVE, never a false FAIL.
  • Type-aware comparison, not naive equality. Dates, numbers, and strings are normalized before comparison so formatting variance doesn't produce false regressions.
  • Contracts compile to portable, versioned JSON. Python is the authoring ergonomics; the JSON file is the diffable artifact checked into a repo and hashed for the verification ledger.
  • Custom comparators are registered, never dynamically imported. A contract file's comparator_ref may not share the trust level of the code running it — a PR from an external contributor, or a contract pulled from a shared source. @toolcontract.register_comparator requires the trusted test suite to explicitly opt a function in; anything else in a contract file fails closed instead of being imported and called.
  • Capability is per-model, not per-adapter. A single adapter instance (especially a gateway-style one like LiteLLMAdapter) can front providers with genuinely different tool-calling guarantees — capability_for(model) is queried at call time, never assumed as a static fact.

Contributing

See CONTRIBUTING.md — provider adapters especially welcome.

Development

python -m venv .venv
.venv/bin/pip install -e ".[dev]"
.venv/bin/python -m pytest

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

toolcontract-0.1.0.tar.gz (41.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

toolcontract-0.1.0-py3-none-any.whl (37.0 kB view details)

Uploaded Python 3

File details

Details for the file toolcontract-0.1.0.tar.gz.

File metadata

  • Download URL: toolcontract-0.1.0.tar.gz
  • Upload date:
  • Size: 41.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.12.3

File hashes

Hashes for toolcontract-0.1.0.tar.gz
Algorithm Hash digest
SHA256 49d81b1c4fa2977964170fa559e3913153b6595895fa7d0261e6df9623db613d
MD5 c6764f82d20f74af9e85404a3a57834f
BLAKE2b-256 7726b201ac9cfb9ef7d152bd8d591086950248b617d17edf3f37e4f90a9b25c5

See more details on using hashes here.

File details

Details for the file toolcontract-0.1.0-py3-none-any.whl.

File metadata

  • Download URL: toolcontract-0.1.0-py3-none-any.whl
  • Upload date:
  • Size: 37.0 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.12.3

File hashes

Hashes for toolcontract-0.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 df97e270009e9776d628182628a9a55019197009e6ed67321fb19c6b102e44df
MD5 024fc690862d559ccd94a3b72a4fad95
BLAKE2b-256 8e5f35818dc562c87d1fca446637f025ea7af8c827237af1ca9e439af8738153

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page