agenttoolkit
agenttoolkit provides one provider-neutral definition for tools exposed to
LLM agents. Define a tool once — schema, availability, metadata, and
execution logic — and expose it to OpenAI, Anthropic, or any other provider
without duplicating definitions.
It intentionally contains no application-specific tools and no agent loop: it is a building block, not a framework.
Table of contents
- Features
- Installation
- Quickstart
- Defining tools
- Dependency injection with
ToolContext - Conditional availability and descriptions
- Driving an agent loop
- Results (
ActionResult) - Middleware
- Merging registries
- Filesystem and shell primitives
- Skills
- Development
- License
Features
- Registration through a
@tools.actiondecorator — no hand-written JSON Schema, for either plain function signatures or Pydantic models. - Runtime metadata (
effects,status,tags, custom fields) and anrequires_approvalflag, kept out of the model-facing schema but readable by the host loop that dispatches calls. - Context-based dependency injection (
Inject[T]) so tools can receive application services without the model ever seeing them. - Conditional tool availability and dynamic, context-aware descriptions.
- Sync and async tool execution behind a single async API.
- A composable middleware pipeline (error boundary, resolution, validation, logging) that applications can extend or replace.
- Thin, dependency-free schema adapters for OpenAI and Anthropic tool-call formats.
- Generic
ActionResulttype so applications can add project-specific result fields without falling back toAny. - Async filesystem and shell ports with local, Docker, and Bubblewrap implementations for common agent capabilities.
- Local Agent Skills discovery and progressive loading, compatible with the
SKILL.mdconvention.
Installation
uv add agenttoolkit
Requires Python 3.13–3.14. Modules that use forward references should add
from __future__ import annotations, since lazy annotation evaluation
(PEP 649) is native only to 3.14+.
Quickstart
This is the shape of code you actually write and run — define tools with the decorator, hand their schema to the model, execute whichever call it makes, and feed the result back:
from pydantic import BaseModel, Field
from agenttoolkit import ActionResult, Inject, ToolContext, Tools, ToolSchemaFormat
class SearchParams(BaseModel):
query: str = Field(description="What to search for")
limit: int = Field(default=5, ge=1, le=20)
class SearchClient:
async def search(self, query: str, limit: int) -> list[str]:
return [query] * limit
tools = Tools(context=ToolContext(SearchClient()))
@tools.action(
"Search the connected knowledge base.",
params=SearchParams,
status="Searching for {query}...",
)
async def search(params: SearchParams, client: Inject[SearchClient]) -> list[str]:
return await client.search(params.query, params.limit)
# 1. Send the schema to the model.
schema = tools.get_schema(ToolSchemaFormat.ANTHROPIC)
# 2. The model asks to call "search" with {"query": "tool middleware"}.
result: ActionResult[object] = await tools.execute(
"search", {"query": "tool middleware"}
)
# 3. Feed the outcome back to the model.
if result.ok:
matches = result.result
else:
error_message = result.error
tools.execute(...) never raises for expected failures — an unknown tool
name, invalid arguments, or an exception inside the tool all come back as a
failed ActionResult, ready to hand to the model as-is.
Defining tools
The @tools.action(...) decorator is the entire surface most code touches.
Parameters come from a plain function signature or, for validation and
richer schemas, a Pydantic model passed as params=:
@tools.action("Add two integers.")
def add(a: int, b: int) -> int:
return a + b
class RefundParams(BaseModel):
order_id: str
amount: float = Field(gt=0, description="Amount to refund, in USD")
@tools.action(
"Issue a refund for an order.",
params=RefundParams,
effects=(ToolEffect.NETWORK,),
status=lambda params: (
f"Refunding {params.amount} for order {params.order_id}..."
),
tags=["billing", "write"],
requires_approval=True,
metadata={"owner": "billing-team"},
)
def refund(params: RefundParams, client: Inject[BillingClient]) -> str:
client.refund(params.order_id, params.amount)
return "refunded"
None of effects, status, tags, requires_approval, or metadata are
visible to the model — they never appear in the generated JSON Schema. They
exist for the host loop that dispatches the call:
effects— afrozenset[ToolEffect]declaring what a call does to the world (READS_WORKSPACE,WRITES_WORKSPACE,NETWORK,SPAWNS_PROCESS), readable astool.effectsor viatool.has_effect(...). Effects let middleware and host policies react to behaviour rather than to tool names, which are part of the model-facing API and change over time.status— a human-readable status message, either astr.formattemplate referencing parameter names or a callable. For callables, the parameter type is inferred fromparams, providing type checking and IDE navigation. Render it withtool.format_status(args)(e.g. to show "Refunding 20.0 for order o-123..." while the call runs).tags— afrozenset[str]for grouping or filtering tools, readable astool.tags.metadata— an arbitrary read-only mapping for anything else the host application needs, readable astool.extra.requires_approval— readable astool.requires_approval; check it before callingtools.execute(...)if the action needs user confirmation first.agenttoolkitdoes not enforce approval itself.
tool = tools.get("refund")
tool.effects # frozenset({ToolEffect.NETWORK})
tool.tags # frozenset({"billing", "write"})
tool.extra["owner"] # "billing-team"
tool.requires_approval # True
tool.format_status({"order_id": "o-123", "amount": 20.0})
# "Refunding 20.0 for order o-123..."
String status templates are validated against params at registration
time. Callable field access is checked statically by the IDE or type checker.
Prefer tools.action(...) for registering tools. Direct registration is an
internal implementation detail.
Dependency injection with ToolContext
ToolContext carries application services that tools need but that should
never appear in the model-facing schema. Wrap a parameter in Inject[T] and
it is resolved from context at call time instead of being part of the
argument schema:
context = ToolContext(SearchClient(), some_other_service)
tools.set_context(context)
ToolContext.resolve(T) returns the most recently provided instance of type
T (or a subclass), searching in reverse insertion order. Useful mutators:
context.provide(extra_service) # append more dependencies
context.without(SearchClient) # drop instances of a type
context.clear() # remove everything
If an Inject[T] parameter has no default and no matching dependency is
found in context, execution raises ValueError rather than silently
passing None.
Conditional availability and descriptions
Use provided(...) and requires(...) to expose a tool only when its
dependency is present (and, optionally, satisfies a predicate). Predicates
compose with &, |, and ~:
from agenttoolkit import provided, requires
@tools.action(
"Issue a refund (admin only).",
available_when=provided(BillingClient)
& requires(UserInfo, predicate=lambda user: user.is_admin),
)
def refund(order_id: str, amount: float) -> str: ...
Use description_from_context(...) when a tool's description itself should
depend on context (e.g. embedding a resolved account name), with a fallback
for when the dependency isn't provided:
from agenttoolkit import description_from_context
description = description_from_context(
BankingClient,
render=lambda client: f"Look up the balance for {client.account_name}.",
fallback="Look up account balance.",
)
@tools.action(description)
def balance() -> float: ...
Driving an agent loop
Tools.get_schema(...) returns the schema for every tool available in the
active (or a given) context; Tools.execute(...) dispatches a model-produced
call:
openai_schemas = tools.get_schema(ToolSchemaFormat.OPENAI)
anthropic_schemas = tools.get_schema(ToolSchemaFormat.ANTHROPIC)
result = await tools.execute("search", {"query": "tool middleware"}, context=context)
A typical loop confirms approval-gated tools before executing, and reports status while a call is in flight:
tool = tools.get(name)
if tool is not None and tool.requires_approval and not confirm(name, arguments):
result = ActionResult[object].fail("Declined by user")
else:
print(tool.format_status(arguments) if tool else name)
result = await tools.execute(name, arguments, context=context)
Tools.get_available() returns the underlying Tool objects for the active
registry context instead of schemas — handy for printing a catalog of what's
currently exposed (tool.name, tool.resolve_description(context),
tool.effects, ...).
Results (ActionResult)
class ActionResult[ResultT = str](BaseModel):
ok: bool
result: ResultT | None = None
error: str | None = None
ResultT defaults to str, the common case for a tool that hands text
back to the model. A bare ActionResult therefore is ActionResult[str]
and validates as one — ActionResult.success(3) raises. Payloads of any
other type must parametrize explicitly.
Raw tool return values are wrapped as successful results automatically. A
tool may instead return an ActionResult directly — e.g. to fail without
raising, or to populate a typed result:
def run_command(command: str) -> ActionResult: # ActionResult[str]
result = shell(command)
if not result.ok:
return ActionResult.fail(result.output)
return ActionResult.success(result.output)
WeatherActionResult = ActionResult[WeatherResult]
def get_weather(city: str) -> WeatherActionResult:
temp_c = KNOWN_CITIES.get(city.lower())
if temp_c is None:
return WeatherActionResult.fail(f"Unknown city: {city!r}")
return WeatherActionResult.success(WeatherResult(city=city, temp_c=temp_c))
Because tool dispatch by name is dynamic and a registry can contain
heterogeneous return types, Tools.execute() always returns
ActionResult[object]; narrow result at the call site, or return a
specialized ActionResult directly from the tool as above.
ActionResult rejects unknown fields. When an application needs additional
result fields (trace IDs, citations, usage info), define them in a typed
subclass and wire it up via Tools(result_type=...) — every result the
middleware chain produces (validation failures, unknown-tool errors, the
internal-error fallback) is then built through that subclass too:
class ProjectActionResult[ResultT = str](ActionResult[ResultT]):
trace_id: str | None = None
citations: tuple[str, ...] = ()
tools = Tools(result_type=ProjectActionResult[object])
Middleware
Every call passes through a fixed core — error boundary, tool resolution,
argument validation — followed by call logging. Pass middleware= to run
additional steps between the core and logging, e.g. a timeout:
from agenttoolkit import ToolCall, ToolMiddleware
class TimeoutMiddleware(ToolMiddleware):
def __init__(self, seconds: float) -> None:
self._seconds = seconds
async def __call__(self, call: ToolCall, next):
return await asyncio.wait_for(next(call), timeout=self._seconds)
tools = Tools(middleware=[TimeoutMiddleware(5.0)])
Custom middleware runs after resolution and validation, so call.tool and
call.params are already populated — which is what makes filtering on tool
metadata (see effects above) possible.
Merging registries
Combine tools from multiple Tools instances — e.g. when composing a
registry from several feature modules:
tools.merge(other_tools) # raises on name collisions
tools.merge(other_tools, replace=True) # other_tools wins on collisions
Filesystem and shell primitives
agenttoolkit.builtins contains raw async implementations rather than a
predefined set of model-facing tools. Applications can use them directly,
inject them through ToolContext, or expose only the operations appropriate
for a particular agent.
from pathlib import Path
from agenttoolkit.builtins import (
BindMount,
DockerSandbox,
LocalWorkspace,
SandboxPolicy,
)
workspace = LocalWorkspace("./project")
await workspace.write_file("src/example.py", "print('hello')\n")
entries = await workspace.list_dir("src")
source = await workspace.read_file(entries[0].path)
output = workspace.root / "output"
output.mkdir(exist_ok=True)
cli_config = Path.home() / ".config" / "my-cli"
policy = SandboxPolicy.for_workspace(
workspace.root,
writable=True,
enable_network_access=True,
)
sandbox = DockerSandbox(
"my-cli:latest",
policy,
inherit_environment=("MY_CLI_TOKEN",),
mounts=(
BindMount.read_only(cli_config, "/home/agent/.config/my-cli"),
BindMount.read_write(output, "/output"),
),
user="host",
)
result = await sandbox.execute("my-cli build --output /output")
The Workspace port provides read_file, write_file, edit_file, glob,
list_dir, and stat. Exploration returns Entry values with a root-relative
POSIX path, directory and symlink flags, size, and modification time. Local
reads and writes are confined to the workspace root and bounded by a
configurable file-size limit.
The Sandbox port returns a common SandboxResult from all backends.
SandboxPolicy controls readable and writable paths, network access,
environment values, timeout, captured output, memory, process, and CPU limits.
DockerSandbox enforces all of these resource limits; BubblewrapSandbox
supports filesystem/network isolation and host-side timeout/output limits.
UnsafeLocalSandbox is useful for trusted commands but deliberately does not
claim to enforce path or network isolation.
DockerSandbox also supports named bind mounts and an explicit allowlist of
host environment variables. BindMount.read_write(...) writes directly back
to the host. inherit_environment fails fast when a requested variable is
missing and forwards its name without embedding the secret value in the
generated Docker arguments. On POSIX hosts, user="host" maps the container
process to the host UID and GID so generated files remain owned by the
developer. Use environment={...} on SandboxPolicy or env={...} on
execute(...) for explicit values and per-call overrides.
Skills
Local Agent Skills are discovered from directories containing one
subdirectory per skill, each with a SKILL.md file using YAML frontmatter
(name, description, and optional license, compatibility, metadata,
allowed-tools) followed by Markdown instructions:
skills/
internet-research/
SKILL.md
references/
guide.md
scripts/
search.py
name must be 1–64 lowercase letters, numbers, or hyphens, and must match
its parent directory name.
from agenttoolkit import Skills
skills = Skills.from_local_dir("./skills")
# Render the compact skill listing for the agent's system prompt.
system_prompt = f"You are helpful.\n\n{skills.render_prompt()}"
# Progressive loading returns full instructions and relative resource paths.
loaded = skills.load("internet-research")
system_prompt += f"\n\n{loaded.instructions}"
# Re-scan the configured directories after skills are added or removed.
changes = skills.refresh()
print(changes.added, changes.updated, changes.removed)
# The application decides which general filesystem and process tools to expose.
guide = read_file(loaded.directory / "references/guide.md")
output = await run_process(
["python", "scripts/search.py", "python packaging"],
cwd=loaded.directory,
)
Skills.from_local_dir accepts multiple directories; a skill discovered
later overrides one with the same name from an earlier directory (logged as
a warning). SKILL.md is re-parsed from disk on each load, so instructions
can be edited without restarting the process. refresh() rebuilds the registry
from the configured directories, picking up added, changed, and removed skills.
If discovery fails, the previous registry remains available.
refresh() returns an immutable SkillChanges value containing the registry
revision and the added, updated, and removed skill names. refresh_if_changed()
first compares a lightweight fingerprint of the SKILL.md paths, modification
times, and sizes, avoiding parsing when no skill document changed.
Agents that can write their own skills can attach SkillRefreshMiddleware.
By default it refreshes the registry after every tool that declares
ToolEffect.WRITES_WORKSPACE, silently and without touching the tool's own
result — so a newly added write tool is covered as soon as it declares its
effect, with no list of tool names to keep in sync:
from agenttoolkit import SkillRefreshMiddleware
tools = Tools(
context=ToolContext(skills),
middleware=[SkillRefreshMiddleware()],
)
The registry is resolved from the call's ToolContext, not captured at
construction — swapping the context via set_context(...) or a per-call
context= argument refreshes the registry actually in use, and the
middleware is a no-op when the context holds no Skills.
Pass when= to select tools by any other predicate over the Tool, e.g.
SkillRefreshMiddleware(when=lambda tool: "skills" in tool.tags).
Invalid skill edits are not activated and
the previous registry remains available. Applications that embed
skills.render_prompt() in model context should render that dynamic portion
again before each model invocation.
load() returns an immutable LoadedSkill containing name, instructions,
the absolute skill directory, and sorted relative resources. Resource
reading, process execution, timeouts, sandboxing, and permissions deliberately
belong to the application's general filesystem and process tools instead of
the Skills API. Skill directories and their scripts must still be treated as
trusted code.
Development
Install the locked development environment and run all quality checks:
uv sync --locked
uv run --locked ruff check .
uv run --locked pytest
The test command measures branch coverage for agenttoolkit and fails
below 90%. Dependabot groups Python dependency updates into one weekly pull
request; the same CI matrix validates every update on Python 3.13 and 3.14.
See CONTRIBUTING.md for the full contribution workflow and conventions.
License
MIT — see LICENSE.md.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file mm_agenttoolkit-0.1.0.tar.gz.
File metadata
- Download URL: mm_agenttoolkit-0.1.0.tar.gz
- Upload date:
- Size: 99.6 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: uv/0.9.2
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
33ff0404f98d0e38069959de9bc0548a3c31d2912d673c8afde03dc4cf55583b
|
|
| MD5 |
5dd6c09f0b4770578e4d0a0cdf81e443
|
|
| BLAKE2b-256 |
2b9b2e133424a6a83efe7de70cb516d82fde73792bb5d6809a318981812a541a
|
File details
Details for the file mm_agenttoolkit-0.1.0-py3-none-any.whl.
File metadata
- Download URL: mm_agenttoolkit-0.1.0-py3-none-any.whl
- Upload date:
- Size: 42.2 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: uv/0.9.2
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
34f8f6da5b0a123dc5ba519832078b42907d80f051dbb68ef65c8da3322dc002
|
|
| MD5 |
5a01766cd73f974301090d34dce98f21
|
|
| BLAKE2b-256 |
77120463ea4d7a502e80b97dd49aee14cde082853bc590f61d459dd9ccd1822b
|