Skip to main content

agenttoolkit

agenttoolkit provides one provider-neutral definition for tools exposed to LLM agents. Define a tool once — schema, availability, metadata, and execution logic — and expose it to OpenAI, Anthropic, or any other provider without duplicating definitions.

It intentionally contains no application-specific tools and no agent loop: it is a building block, not a framework.

Table of contents

Features

  • Registration through a @tools.action decorator — no hand-written JSON Schema, for either plain function signatures or Pydantic models.
  • Runtime metadata (effects, status, tags, custom fields) and an requires_approval flag, kept out of the model-facing schema but readable by the host loop that dispatches calls.
  • Context-based dependency injection (Inject[T]) so tools can receive application services without the model ever seeing them.
  • Conditional tool availability and dynamic, context-aware descriptions.
  • Sync and async tool execution behind a single async API.
  • A composable middleware pipeline (error boundary, resolution, validation, logging) that applications can extend or replace.
  • Thin, dependency-free schema adapters for OpenAI and Anthropic tool-call formats.
  • Generic ActionResult type so applications can add project-specific result fields without falling back to Any.
  • Async filesystem and shell ports with local, Docker, and Bubblewrap implementations for common agent capabilities.
  • Local Agent Skills discovery and progressive loading, compatible with the SKILL.md convention.

Installation

uv add agenttoolkit

Requires Python 3.13–3.14. Modules that use forward references should add from __future__ import annotations, since lazy annotation evaluation (PEP 649) is native only to 3.14+.

Quickstart

This is the shape of code you actually write and run — define tools with the decorator, hand their schema to the model, execute whichever call it makes, and feed the result back:

from pydantic import BaseModel, Field

from agenttoolkit import ActionResult, Inject, ToolContext, Tools, ToolSchemaFormat


class SearchParams(BaseModel):
    query: str = Field(description="What to search for")
    limit: int = Field(default=5, ge=1, le=20)


class SearchClient:
    async def search(self, query: str, limit: int) -> list[str]:
        return [query] * limit


tools = Tools(context=ToolContext(SearchClient()))


@tools.action(
    "Search the connected knowledge base.",
    params=SearchParams,
    status="Searching for {query}...",
)
async def search(params: SearchParams, client: Inject[SearchClient]) -> list[str]:
    return await client.search(params.query, params.limit)


# 1. Send the schema to the model.
schema = tools.get_schema(ToolSchemaFormat.ANTHROPIC)

# 2. The model asks to call "search" with {"query": "tool middleware"}.
result: ActionResult[object] = await tools.execute(
    "search", {"query": "tool middleware"}
)

# 3. Feed the outcome back to the model.
if result.ok:
    matches = result.result
else:
    error_message = result.error

tools.execute(...) never raises for expected failures — an unknown tool name, invalid arguments, or an exception inside the tool all come back as a failed ActionResult, ready to hand to the model as-is.

Defining tools

The @tools.action(...) decorator is the entire surface most code touches. Parameters come from a plain function signature or, for validation and richer schemas, a Pydantic model passed as params=:

@tools.action("Add two integers.")
def add(a: int, b: int) -> int:
    return a + b


class RefundParams(BaseModel):
    order_id: str
    amount: float = Field(gt=0, description="Amount to refund, in USD")


@tools.action(
    "Issue a refund for an order.",
    params=RefundParams,
    effects=(ToolEffect.NETWORK,),
    status=lambda params: (
        f"Refunding {params.amount} for order {params.order_id}..."
    ),
    tags=["billing", "write"],
    requires_approval=True,
    metadata={"owner": "billing-team"},
)
def refund(params: RefundParams, client: Inject[BillingClient]) -> str:
    client.refund(params.order_id, params.amount)
    return "refunded"

None of effects, status, tags, requires_approval, or metadata are visible to the model — they never appear in the generated JSON Schema. They exist for the host loop that dispatches the call:

  • effects — a frozenset[ToolEffect] declaring what a call does to the world (READS_WORKSPACE, WRITES_WORKSPACE, NETWORK, SPAWNS_PROCESS), readable as tool.effects or via tool.has_effect(...). Effects let middleware and host policies react to behaviour rather than to tool names, which are part of the model-facing API and change over time.
  • status — a human-readable status message, either a str.format template referencing parameter names or a callable. For callables, the parameter type is inferred from params, providing type checking and IDE navigation. Render it with tool.format_status(args) (e.g. to show "Refunding 20.0 for order o-123..." while the call runs).
  • tags — a frozenset[str] for grouping or filtering tools, readable as tool.tags.
  • metadata — an arbitrary read-only mapping for anything else the host application needs, readable as tool.extra.
  • requires_approval — readable as tool.requires_approval; check it before calling tools.execute(...) if the action needs user confirmation first. agenttoolkit does not enforce approval itself.
tool = tools.get("refund")
tool.effects             # frozenset({ToolEffect.NETWORK})
tool.tags                # frozenset({"billing", "write"})
tool.extra["owner"]      # "billing-team"
tool.requires_approval   # True
tool.format_status({"order_id": "o-123", "amount": 20.0})
# "Refunding 20.0 for order o-123..."

String status templates are validated against params at registration time. Callable field access is checked statically by the IDE or type checker.

Prefer tools.action(...) for registering tools. Direct registration is an internal implementation detail.

Dependency injection with ToolContext

ToolContext carries application services that tools need but that should never appear in the model-facing schema. Wrap a parameter in Inject[T] and it is resolved from context at call time instead of being part of the argument schema:

context = ToolContext(SearchClient(), some_other_service)
tools.set_context(context)

ToolContext.resolve(T) returns the most recently provided instance of type T (or a subclass), searching in reverse insertion order. Useful mutators:

context.provide(extra_service)  # append more dependencies
context.without(SearchClient)   # drop instances of a type
context.clear()                 # remove everything

If an Inject[T] parameter has no default and no matching dependency is found in context, execution raises ValueError rather than silently passing None.

Conditional availability and descriptions

Use provided(...) and requires(...) to expose a tool only when its dependency is present (and, optionally, satisfies a predicate). Predicates compose with &, |, and ~:

from agenttoolkit import provided, requires


@tools.action(
    "Issue a refund (admin only).",
    available_when=provided(BillingClient)
    & requires(UserInfo, predicate=lambda user: user.is_admin),
)
def refund(order_id: str, amount: float) -> str: ...

Use description_from_context(...) when a tool's description itself should depend on context (e.g. embedding a resolved account name), with a fallback for when the dependency isn't provided:

from agenttoolkit import description_from_context

description = description_from_context(
    BankingClient,
    render=lambda client: f"Look up the balance for {client.account_name}.",
    fallback="Look up account balance.",
)


@tools.action(description)
def balance() -> float: ...

Driving an agent loop

Tools.get_schema(...) returns the schema for every tool available in the active (or a given) context; Tools.execute(...) dispatches a model-produced call:

openai_schemas = tools.get_schema(ToolSchemaFormat.OPENAI)
anthropic_schemas = tools.get_schema(ToolSchemaFormat.ANTHROPIC)

result = await tools.execute("search", {"query": "tool middleware"}, context=context)

A typical loop confirms approval-gated tools before executing, and reports status while a call is in flight:

tool = tools.get(name)
if tool is not None and tool.requires_approval and not confirm(name, arguments):
    result = ActionResult[object].fail("Declined by user")
else:
    print(tool.format_status(arguments) if tool else name)
    result = await tools.execute(name, arguments, context=context)

Tools.get_available() returns the underlying Tool objects for the active registry context instead of schemas — handy for printing a catalog of what's currently exposed (tool.name, tool.resolve_description(context), tool.effects, ...).

Results (ActionResult)

class ActionResult[ResultT = str](BaseModel):
    ok: bool
    result: ResultT | None = None
    error: str | None = None

ResultT defaults to str, the common case for a tool that hands text back to the model. A bare ActionResult therefore is ActionResult[str] and validates as one — ActionResult.success(3) raises. Payloads of any other type must parametrize explicitly.

Raw tool return values are wrapped as successful results automatically. A tool may instead return an ActionResult directly — e.g. to fail without raising, or to populate a typed result:

def run_command(command: str) -> ActionResult:  # ActionResult[str]
    result = shell(command)
    if not result.ok:
        return ActionResult.fail(result.output)
    return ActionResult.success(result.output)
WeatherActionResult = ActionResult[WeatherResult]


def get_weather(city: str) -> WeatherActionResult:
    temp_c = KNOWN_CITIES.get(city.lower())
    if temp_c is None:
        return WeatherActionResult.fail(f"Unknown city: {city!r}")
    return WeatherActionResult.success(WeatherResult(city=city, temp_c=temp_c))

Because tool dispatch by name is dynamic and a registry can contain heterogeneous return types, Tools.execute() always returns ActionResult[object]; narrow result at the call site, or return a specialized ActionResult directly from the tool as above.

ActionResult rejects unknown fields. When an application needs additional result fields (trace IDs, citations, usage info), define them in a typed subclass and wire it up via Tools(result_type=...) — every result the middleware chain produces (validation failures, unknown-tool errors, the internal-error fallback) is then built through that subclass too:

class ProjectActionResult[ResultT = str](ActionResult[ResultT]):
    trace_id: str | None = None
    citations: tuple[str, ...] = ()


tools = Tools(result_type=ProjectActionResult[object])

Middleware

Every call passes through a fixed core — error boundary, tool resolution, argument validation — followed by call logging. Pass middleware= to run additional steps between the core and logging, e.g. a timeout:

from agenttoolkit import ToolCall, ToolMiddleware


class TimeoutMiddleware(ToolMiddleware):
    def __init__(self, seconds: float) -> None:
        self._seconds = seconds

    async def __call__(self, call: ToolCall, next):
        return await asyncio.wait_for(next(call), timeout=self._seconds)


tools = Tools(middleware=[TimeoutMiddleware(5.0)])

Custom middleware runs after resolution and validation, so call.tool and call.params are already populated — which is what makes filtering on tool metadata (see effects above) possible.

Merging registries

Combine tools from multiple Tools instances — e.g. when composing a registry from several feature modules:

tools.merge(other_tools)               # raises on name collisions
tools.merge(other_tools, replace=True)  # other_tools wins on collisions

Filesystem and shell primitives

agenttoolkit.builtins contains raw async implementations rather than a predefined set of model-facing tools. Applications can use them directly, inject them through ToolContext, or expose only the operations appropriate for a particular agent.

from pathlib import Path

from agenttoolkit.builtins import (
    BindMount,
    DockerSandbox,
    LocalWorkspace,
    SandboxPolicy,
)

workspace = LocalWorkspace("./project")
await workspace.write_file("src/example.py", "print('hello')\n")

entries = await workspace.list_dir("src")
source = await workspace.read_file(entries[0].path)

output = workspace.root / "output"
output.mkdir(exist_ok=True)
cli_config = Path.home() / ".config" / "my-cli"

policy = SandboxPolicy.for_workspace(
    workspace.root,
    writable=True,
    enable_network_access=True,
)
sandbox = DockerSandbox(
    "my-cli:latest",
    policy,
    inherit_environment=("MY_CLI_TOKEN",),
    mounts=(
        BindMount.read_only(cli_config, "/home/agent/.config/my-cli"),
        BindMount.read_write(output, "/output"),
    ),
    user="host",
)
result = await sandbox.execute("my-cli build --output /output")

The Workspace port provides read_file, write_file, edit_file, glob, list_dir, and stat. Exploration returns Entry values with a root-relative POSIX path, directory and symlink flags, size, and modification time. Local reads and writes are confined to the workspace root and bounded by a configurable file-size limit.

The Sandbox port returns a common SandboxResult from all backends. SandboxPolicy controls readable and writable paths, network access, environment values, timeout, captured output, memory, process, and CPU limits. DockerSandbox enforces all of these resource limits; BubblewrapSandbox supports filesystem/network isolation and host-side timeout/output limits. UnsafeLocalSandbox is useful for trusted commands but deliberately does not claim to enforce path or network isolation.

DockerSandbox also supports named bind mounts and an explicit allowlist of host environment variables. BindMount.read_write(...) writes directly back to the host. inherit_environment fails fast when a requested variable is missing and forwards its name without embedding the secret value in the generated Docker arguments. On POSIX hosts, user="host" maps the container process to the host UID and GID so generated files remain owned by the developer. Use environment={...} on SandboxPolicy or env={...} on execute(...) for explicit values and per-call overrides.

Skills

Local Agent Skills are discovered from directories containing one subdirectory per skill, each with a SKILL.md file using YAML frontmatter (name, description, and optional license, compatibility, metadata, allowed-tools) followed by Markdown instructions:

skills/
  internet-research/
    SKILL.md
    references/
      guide.md
    scripts/
      search.py

name must be 1–64 lowercase letters, numbers, or hyphens, and must match its parent directory name.

from agenttoolkit import Skills

skills = Skills.from_local_dir("./skills")

# Render the compact skill listing for the agent's system prompt.
system_prompt = f"You are helpful.\n\n{skills.render_prompt()}"

# Progressive loading returns full instructions and relative resource paths.
loaded = skills.load("internet-research")
system_prompt += f"\n\n{loaded.instructions}"

# Re-scan the configured directories after skills are added or removed.
changes = skills.refresh()
print(changes.added, changes.updated, changes.removed)

# The application decides which general filesystem and process tools to expose.
guide = read_file(loaded.directory / "references/guide.md")
output = await run_process(
    ["python", "scripts/search.py", "python packaging"],
    cwd=loaded.directory,
)

Skills.from_local_dir accepts multiple directories; a skill discovered later overrides one with the same name from an earlier directory (logged as a warning). SKILL.md is re-parsed from disk on each load, so instructions can be edited without restarting the process. refresh() rebuilds the registry from the configured directories, picking up added, changed, and removed skills. If discovery fails, the previous registry remains available.

refresh() returns an immutable SkillChanges value containing the registry revision and the added, updated, and removed skill names. refresh_if_changed() first compares a lightweight fingerprint of the SKILL.md paths, modification times, and sizes, avoiding parsing when no skill document changed.

Agents that can write their own skills can attach SkillRefreshMiddleware. By default it refreshes the registry after every tool that declares ToolEffect.WRITES_WORKSPACE, silently and without touching the tool's own result — so a newly added write tool is covered as soon as it declares its effect, with no list of tool names to keep in sync:

from agenttoolkit import SkillRefreshMiddleware

tools = Tools(
    context=ToolContext(skills),
    middleware=[SkillRefreshMiddleware()],
)

The registry is resolved from the call's ToolContext, not captured at construction — swapping the context via set_context(...) or a per-call context= argument refreshes the registry actually in use, and the middleware is a no-op when the context holds no Skills.

Pass when= to select tools by any other predicate over the Tool, e.g. SkillRefreshMiddleware(when=lambda tool: "skills" in tool.tags). Invalid skill edits are not activated and the previous registry remains available. Applications that embed skills.render_prompt() in model context should render that dynamic portion again before each model invocation.

load() returns an immutable LoadedSkill containing name, instructions, the absolute skill directory, and sorted relative resources. Resource reading, process execution, timeouts, sandboxing, and permissions deliberately belong to the application's general filesystem and process tools instead of the Skills API. Skill directories and their scripts must still be treated as trusted code.

Development

Install the locked development environment and run all quality checks:

uv sync --locked
uv run --locked ruff check .
uv run --locked pytest

The test command measures branch coverage for agenttoolkit and fails below 90%. Dependabot groups Python dependency updates into one weekly pull request; the same CI matrix validates every update on Python 3.13 and 3.14.

See CONTRIBUTING.md for the full contribution workflow and conventions.

License

MIT — see LICENSE.md.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

mm_agenttoolkit-0.1.0.tar.gz (99.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

mm_agenttoolkit-0.1.0-py3-none-any.whl (42.2 kB view details)

Uploaded Python 3

File details

Details for the file mm_agenttoolkit-0.1.0.tar.gz.

File metadata

  • Download URL: mm_agenttoolkit-0.1.0.tar.gz
  • Upload date:
  • Size: 99.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.9.2

File hashes

Hashes for mm_agenttoolkit-0.1.0.tar.gz
Algorithm Hash digest
SHA256 33ff0404f98d0e38069959de9bc0548a3c31d2912d673c8afde03dc4cf55583b
MD5 5dd6c09f0b4770578e4d0a0cdf81e443
BLAKE2b-256 2b9b2e133424a6a83efe7de70cb516d82fde73792bb5d6809a318981812a541a

See more details on using hashes here.

File details

Details for the file mm_agenttoolkit-0.1.0-py3-none-any.whl.

File metadata

File hashes

Hashes for mm_agenttoolkit-0.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 34f8f6da5b0a123dc5ba519832078b42907d80f051dbb68ef65c8da3322dc002
MD5 5a01766cd73f974301090d34dce98f21
BLAKE2b-256 77120463ea4d7a502e80b97dd49aee14cde082853bc590f61d459dd9ccd1822b

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page