agent-harness-adk
A fast, lightweight harness for building production AI agents in Python.
Agents, sub-agents, skills, prompts, tools, MCP servers, memory — and the runtime rails underneath them: permissions, budgets, hooks, guardrails, tracing, checkpoints and isolated workspaces. Twenty LLM providers, one loop, no framework lock-in.
pip install agent-harness-adk # or: uv add agent-harness-adk
import agent_harness # installed as agent-harness-adk, imported as agent_harness
Python 3.10 – 3.14. Three dependencies (pydantic, httpx, pyyaml), ~100 ms
to import, and no vendor SDKs — the provider adapters speak HTTP directly so
every backend travels the same retry, circuit-breaker, cost and tracing path.
60 seconds
from agent_harness import Agent, tool
@tool
def order_status(order_id: str) -> str:
"""Look up the status of a customer order.
Args:
order_id: the order number, digits only.
"""
return db.lookup(order_id)
agent = Agent(
"support",
"Answer customer questions about orders. Look the order up before answering.",
tools=[order_status],
)
result = agent.run_sync("Where is order 4182?")
print(result.output, result.cost_usd, result.steps)
The decorator reads your signature and docstring and builds the JSON Schema the
model needs. Arguments coming back from the model are validated before your
function is called. await agent.run(...) is the real implementation;
run_sync is the wrapper for scripts and notebooks.
The shape of the system
┌──────────────────────────────────────────────────────────┐
│ Orchestrator — plan · staff · run · consolidate · review │
└───────────────┬──────────────────────────────────────────┘
│ staffing decision: reuse or create?
┌──────────────┴───────────────┐
┌──────▼──────┐ ┌───────▼────────┐
│ The bench │ │ The factory │
│ pre-defined │ │ a new spec │
│ sub-agents │ │ written at run │
└──────┬──────┘ └───────┬────────┘
└──────────────┬───────────────┘
┌──────▼───────┐
│ Agent loop │ think → act → observe → repeat
└──────┬───────┘
┌─────────────────────┼─────────────────────────┐
│ context assembler │ tools · skills · MCP │ memory: user · session
│ context compactor │ workspace · providers │ orchestrator · sub-agent
└─────────────────────┴─────────────────────────┘
rails: permissions · budget · hooks · guardrails · tracing · journal ·
cache · checkpoints · sessions · scheduler · router
governance: policy · identity · residency · oversight · evidence
Tools
from agent_harness import tool, ToolContext
@tool(permission="ask", cacheable=True, tags=["billing"])
async def issue_refund(order_id: str, amount: float, ctx: ToolContext) -> str:
"""Refund a customer. Costs real money.
Args:
order_id: the order to refund.
amount: how much, in EUR.
"""
ctx.log("refunding", order=order_id)
return await billing.refund(order_id, amount)
- Sync or async, it makes no difference.
- A parameter named
ctx(or annotatedToolContext) is injected and hidden from the model. - A pydantic model as a parameter type is validated and passed through as a model, not a dict.
permissioncan tighten the policy for one tool. It can never loosen it.- Tools run in parallel when the model asks for several at once.
Built-ins in agent_harness.toolkits:
| Tool | Notes |
|---|---|
now, calculate |
exact arithmetic, no eval |
make_corpus_search |
keyword search over documents you hand it |
make_fetch_tool, make_http_tool |
domain allowlist, private-address refusal, HTML stripping |
parse_document |
text, Markdown, CSV, TSV, JSON, JSONL, HTML, XML with no dependencies; PDF, DOCX and OCR with an optional install each |
bar_chart, line_chart, render_report |
inline SVG that works in light and dark, plus markdown reports |
make_python_tool |
run code in the workspace — asks for approval every time |
A workspace brings fs_read, fs_write, fs_list, fs_delete and — only when
you ask for it — shell.
Skills
A skill is packaged know-how: a folder with SKILL.md and, optionally, its own
tools and reference files.
skills/refunds/SKILL.md
---
name: refunds
description: How we process a refund, including the approval thresholds.
---
1. Check the order is inside the 30-day window...
agent = Agent("support", "Answer support questions.", skills="./skills")
Only each skill's name and description go into the system prompt. The body
is loaded on demand through the load_skill tool, so twenty skills cost twenty
lines of context instead of twenty documents. A tools.py in the skill folder is
imported and its tools come along with it.
Prompts
from agent_harness import Prompt, PromptLibrary
triage = Prompt("triage", "Sort {ticket} into {buckets}.", version="2")
triage.render(ticket="T-1", buckets="p1/p2/p3")
library = PromptLibrary.from_dir("./prompts") # .md files with YAML frontmatter
library.render("triage", ticket="T-1")
Versioned, reviewable, .partial()-able, composable with +. Jinja is used
only when a template contains a {% %} statement and jinja2 is installed.
Memory — four scopes
| Scope | Stored | Loaded | Lifetime |
|---|---|---|---|
user (user.md) |
preferences, standards, settled decisions | in full, every message | permanent, rewritten at session close |
| session | the whole conversation plus its artefacts | in full | this session |
| orchestrator | plans, staffing decisions, spend, findings | a digest only | this job, then distilled into user memory |
| sub-agent | only the resources its task produced | nothing carried in | the task |
agent = Agent("assistant", memory=True) # the default
await agent.run("I bill my customers in EUR")
await agent.run("What currency do I use?", messages=[]) # clean run, still knows
print(await agent.close_session()) # session close → user.md rewritten
Whose memory is it? — trace
A trace says who a memory belongs to. Pass one and every record is stamped with it on the way in and filtered by it on the way out, so one store serves any number of users without them ever seeing each other.
agent = Agent("support", trace="alice") # a bare user id
agent = Agent("support", trace=Trace(user_id="alice", session_id="s-42"))
agent = Agent("support", trace=Trace(tenant_id="acme", user_id="alice"))
alice = MemoryManager(store, trace="alice")
bob = MemoryManager(store, trace="bob")
await alice.user.remember("bills in EUR")
await bob.user.load() # "" — bob never sees it
One manager, many users, one backend:
shared = MemoryManager(store)
await shared.for_trace(request.user_id).user.remember(fact)
for_trace reuses the store and its vector index, so serving a request per user
costs a small object rather than a rebuilt index.
scope decides what documents like user.md are namespaced by — "user" (the
default: preferences follow the person across sessions), "session", "tenant"
or "global". A record written without a trace stays visible to everyone: it is
shared, not orphaned.
Where is it stored? — memory/providers
Thirteen backends, one contract. The agent loop never learns which is behind it.
from agent_harness import memory_provider
store = memory_provider("postgresql://user:pass@host/agents")
store = memory_provider("mongodb://localhost:27017", database="agents")
store = memory_provider("s3://my-bucket/agent-memory")
store = memory_provider("sqlite:///./memory.db")
agent = Agent("support", memory=MemoryManager(store, trace="alice"))
| Backend | Class | Needs |
|---|---|---|
| in-process | InMemoryStore |
— |
| files | FileStore |
— |
| SQLite | SQLiteMemory |
— (standard library) |
| PostgreSQL | PostgresMemory |
asyncpg or psycopg |
| MySQL / MariaDB | MySQLMemory |
aiomysql |
| MongoDB | MongoMemory |
motor or pymongo |
| Redis | RedisMemory |
redis |
| DynamoDB | DynamoDBMemory |
boto3 |
| Elasticsearch / OpenSearch | ElasticsearchMemory |
elasticsearch |
| Amazon S3 | S3Memory |
boto3 |
| Azure Blob Storage | AzureBlobMemory |
azure-storage-blob |
| Google Cloud Storage | GCSMemory |
google-cloud-storage |
| your own API | HTTPMemory |
— (httpx already ships) |
Nothing is imported until you ask for it — a driver you do not use costs nothing
at import, and one you have not installed names its own pip install rather than
raising ImportError somewhere deep in a run:
from agent_harness import available_backends
available_backends()
# {'sqlite': True, 'postgres': False, 's3': False, ...}
Choosing between them:
- SQLite is the right default for a single service — durable, indexed, no server to run.
- Postgres, MySQL and Mongo index on the trace, so reading one user's memory is one query however many users you have.
- Redis suits session-scoped memory: pass
ttl=and it expires itself. - Elasticsearch is the only backend where
search()is ranked by the engine, so recall is good without an embedder. - S3, Azure Blob and GCS lay keys out so a trace is a prefix. That makes one user's memory a single listing, but anything narrower is filtered after the fetch — treat them as durable archival rather than a hot query path.
- HTTPMemory is for when memory must live behind a service you already run.
register_backend("cassandra", "myapp.memory", "CassandraMemory") adds your own.
The agent gets remember and recall tools. Recall is semantic: embeddings
come from whatever you configure, and the default is a deterministic offline
hashing embedder so semantic recall works with no extra dependency and no
network. Swap it for the real thing when you want to:
from agent_harness import MemoryManager, ProviderEmbedder, OpenAIProvider, FileStore
memory = MemoryManager(FileStore(".harness/memory"),
embedder=ProviderEmbedder(OpenAIProvider()))
Context: when it compresses
Every step, the conversation is measured and compacted if it is over the threshold. You choose where that is:
Agent("a", compact_at=10_000) # an absolute token count
Agent("b", compact_at=50_000)
Agent("c", compact_at=0.5) # or a fraction of the model's context window
Agent("d") # default: two thirds of the window
Compaction happens in two stages, so the cheap thing is tried first:
- Evict — oversized tool results are hollowed out, keeping their first 400
characters. Cheapest tokens to lose, and it cannot break a
tool_use/tool_resultpair because nothing is removed. - Summarise — if it is still over, the head of the conversation is summarised by a cheap model call and the tail kept verbatim. The cut moves forward until no tool result is left without its call.
Agent("a",
compact_at=10_000, # start compacting here
compact_target=0.6, # compress down to 60% of that
compact_keep_last=8) # never touch the last 8 messages
Pass compactor=ContextCompactor(...) to replace the strategy wholesale, and
memory.session.pin("the deadline is Friday") for facts that must survive it.
Sub-agents: the bench and the factory
Before staffing a task, the orchestrator asks one question: is there already a sub-agent that covers this?
from agent_harness import Agent, SubAgentSpec
manager = Agent(
"manager",
"Delegate the lookups, then consolidate what comes back.",
tools=[lookup],
subagents=[
SubAgentSpec(name="revenue_reader", description="Finds revenue figures.",
instructions="Look up the figure and report it with its source.",
tools=["lookup"], tier="fast"),
SubAgentSpec(name="cost_reader", description="Finds cost figures.",
tools=["lookup"], tier="fast"),
],
)
result = await manager.run("How did Q3 go?")
for child in result.children:
print(child.agent, child.steps, child.cost_usd)
A delegate tool appears automatically. Sub-agents start clean — no parent
transcript, no parent memory — and hand back a result, not a conversation.
Delegation does not cascade by default, and a spec's tools list is a hard
allowlist. Ask for several delegations in one turn and they run in parallel
under the concurrency cap.
Agents it writes for itself, at run time
An ordinary agent can build its own specialists mid-run, within a budget you set:
agent = Agent(
"core",
"Break the work up and give each part its own specialist.",
tools=[lookup, publish],
runtime_agents="enable", # or "disable", or a plain bool
max_runtime_agents=5, # 0-100; the ceiling for one run
)
That adds a spawn_agent tool. When the agent decides it needs five workers, it
calls it five times in one turn and they run in parallel — each one written for
its task by the factory (name, instructions, tool allowlist, model tier, step
ceiling), then run, with only its result handed back.
result = await agent.run("Reconcile these five ledgers.")
print(agent.total_spawned, [c.agent for c in result.children])
# 5 ['ledger_2024', 'ledger_2025', ...]
The rules around it:
- The budget is per run and refreshes on the next one.
agent.runtime_agents_remainingis what is left; past the ceiling the tool says so and the agent finishes with what it has rather than failing. - It does not cascade. A spawned specialist cannot spawn its own.
- Least privilege. A specialist gets the tools its spec asked for, narrowed to
what the parent holds;
runtime_agent_tools=[...]caps that further. - Every spin-up is audited, counted against
Budget(max_subagents=...), and refused once the run is stopped. - Spawned specialists stay addressable by name through
delegatefor the rest of the run, so the second task for the same worker costs nothing extra to set up.
Nothing on the bench fits? The factory writes a new specialist during the run — name, instructions, tool allowlist, model tier, step ceiling and workspace isolation — and that specialist exists only for this job.
from agent_harness import Bench
Bench.standard().names
# ['compliance_checker', 'data_analyst', 'document_extractor', 'drafting',
# 'planner', 'report_writer', 'research', 'validator']
The orchestrator
from agent_harness import Orchestrator, Budget
boss = Orchestrator("boss", max_concurrency=4, review=True, max_rework=1,
budget=Budget(max_usd=2.00))
result = await boss.run("Summarise how Q3 went, with the numbers cited.")
- Plan — acceptance tests are written before any work starts, then the task graph, then a cost estimate.
- Staff — reuse from the bench, else build with the factory.
- Run — dependency-ordered waves, parallel inside each wave, per-task retries, dependent tasks receive only what they depend on.
- Consolidate — merge, de-duplicate, rank, attribute.
- Review — an independent critic checks the deliverable against the definition of done; a rejection becomes new tasks and one rework round.
result.data["plan"] and result.data["review"] carry the full record.
MCP
from agent_harness import Agent, MCPManager, MCPServer
servers = [
MCPServer(name="files", command="npx",
args=["-y", "@modelcontextprotocol/server-filesystem", "/data"]),
MCPServer(name="api", url="https://mcp.internal/rpc",
headers={"authorization": "Bearer ..."}),
]
async with MCPManager(servers) as mcp:
agent = Agent("analyst", "Answer from the files.", tools=mcp.tools())
print((await agent.run("What is in /data/report.md?")).output)
Both transports (stdio and streamable HTTP), tools, resources and prompts. A
server that will not connect is reported in mcp.errors, not raised into your
run. allowed_tools trims what a server may expose.
Budgets that stop instead of failing
SubAgentSpec(name="researcher", description="Finds things out.",
budget=Budget(max_input_tokens=10_000, max_output_tokens=2_000))
Reaching a ceiling is not an error. The run ends cleanly, whatever the agent produced is kept, and a line is appended saying why it stopped:
I got through three of the five documents...
[The budget for this agent is exceeded — output tokens 2,048 of 2,000.
The answer above is what it completed before stopping.]
The parent gets that as the sub-agent's result and carries on. result.stop_reason
is "budget" and result.budget_exceeded names the axis. Pass
Budget(..., on_exceed="raise") if you would rather it were an error.
When a model cannot be reached
harness.router = ModelRouter(fallbacks=["claude-sonnet-5", "gpt-4.1"])
The loop walks the chain, resolving each model's provider as it goes. A 4xx is not retried elsewhere — the request is wrong and the next model will reject it the same way. Every switch lands in the journal and the audit trail.
Versions
One agent, several configurations:
agent = Agent(
"support",
tools=[order_status, issue_refund, lookup],
version="v2",
versions={
"v1": {"instructions": "Answer order questions.",
"tools": ["order_status"], "model": "claude-sonnet-5"},
"v2": {"instructions": "Answer order questions. Cite the order.",
"tools": ["order_status", "lookup"],
"guardrails": {"require_tools": ["order_status"]},
"model": "claude-opus-5"},
},
)
await agent.run(task) # v2
await agent.run(task, version="v1") # the old one, unchanged
A version says what is different; everything it leaves out falls through. The harness, provider and memory are shared, so switching is cheap and the two are comparable — run the same golden tasks against each:
v1 = await suite.run(agent.use("v1"), label="v1")
v2 = await suite.run(agent.use("v2"), label="v2")
print(v2.compare(v1).render())
Declaring it all in a file
# agents.yaml
defaults: {model: claude-opus-5}
prompts:
house_style: Answer in plain sentences and cite the order.
guardrails:
strict: {require_tools: [order_status], no_pii: true, no_placeholders: true}
subagents:
researcher:
description: Finds things out, read-only.
instructions: "{house_style} Cite every claim."
tools: [lookup]
tier: fast
budget: {max_input_tokens: 10000, max_output_tokens: 2000}
guardrails: {require_citation: true}
agents:
support:
instructions: "{house_style}"
tools: [order_status, lookup]
subagents: [researcher]
guardrails: strict
versions:
v1: {instructions: Answer order questions., tools: [order_status]}
v2: {instructions: "{house_style}"}
blueprint = Blueprint.from_file("agents.yaml")
agent = blueprint.build("support", tools=[order_status, lookup])
everything = blueprint.build_all(tools=[order_status, lookup]) # one harness
JSON works the same way. Tools stay in code — they are code — so you either
hand them in or let the file name them as import paths
(myapp.tools:order_status). Everything else is declaration, and belongs
somewhere it can be reviewed and diffed.
Guardrails: what an agent must do to be done
The content engine (Guardrails) polices text — secrets, injection, size, on
every path in and out. AgentGuardrails polices behaviour: which tools an
agent may touch, and what has to be true of its answer before that answer is
accepted.
from agent_harness import Agent, AgentGuardrails
support = Agent(
"support",
"Answer order questions.",
tools=[order_status, issue_refund],
guardrails=AgentGuardrails(
require_tools=["order_status"], # look it up, never guess
forbid_tools=["issue_refund"], # not this agent's job
must_include=["order"],
require_citation=True,
no_placeholders=True, # no "TODO", no "[insert name]"
max_cost_usd=0.25,
on_violation="retry", # tell it what is missing, let it fix it
),
)
A forbidden tool is refused before it runs. Everything else is checked when the agent tries to finish: if something is unmet the agent is told, in words, and gets another turn —
That answer does not meet this task's requirements yet:
- you answered without calling order_status — call order_status and answer from what it returns
- your answer cites nothing — give the source for each claim, or say you could not find one
Put it right and answer again.
which is usually all it needs. on_violation decides what happens when it does
not: "retry" (the default, up to max_retries), "fail" (stop the run), or
"warn" (deliver it, record the problem in result.violations).
Deterministic detectors
Exact where they can be, so you can leave them switched on:
AgentGuardrails(no_pii=True, no_secrets=True, no_injection=True,
grounded=0.6, not_toxic=True, no_repetition=True)
| Detector | What makes it usable |
|---|---|
PIIDetector |
cards are Luhn-checked, IBANs mod-97-checked, and matches are precedence-ordered — a card is never also reported as a phone number |
SecretDetector |
known key formats, plus Shannon entropy for keys nobody has published a pattern for |
InjectionDetector |
weighted signals scored 0-1, because one suspicious phrase is weak evidence and three together are not |
GroundednessDetector |
which content words in the answer appear nowhere in the sources |
ToxicityDetector |
a screen, including character substitution — not a classifier |
RepetitionDetector |
n-gram repetition, for a model looping on itself |
They report findings with a severity, a confidence and the spans they matched,
so PIIDetector().redact(text) removes exactly the value and leaves the sentence.
LLM judges
For what an algorithm cannot decide:
from agent_harness import LLMGuard, POLICIES
rails = AgentGuardrails(
LLMGuard(cheap_agent, POLICIES["safety"]),
LLMGuard(cheap_agent, "never name a competitor", block_at="high"),
no_pii=True, # the deterministic checks run first
)
Three things this gets right:
- A structured verdict. The judge returns JSON with a severity, not a mood, and a judge that will not answer in JSON has failed rather than passed.
- It fails the way you choose.
on_error="block"(the default),"allow"or"raise". A guard that silently passes when it breaks is not a guard. - Cheap checks first.
check_asyncruns the deterministic checks and only pays for a judge if they are all happy — no reason to spend a model call confirming what a regex just proved.
Ready policies: safety, pii, relevance, groundedness, jailbreak,
tone, compliance. Or pass your own sentence.
The checks ship as objects, so you can compose them directly or write your own:
| Check | Fails when |
|---|---|
RequireTools(*names) |
it answered without calling them |
ForbidTools(*names) |
it called one anyway (post-hoc audit) |
MustInclude / MustNotInclude |
the answer misses, or contains, a phrase |
MustMatch(pattern) |
the answer is not in the shape asked for |
MinLength(chars) |
a one-word answer to a question that needed working through |
RequireCitation() |
nothing in the answer points at a source |
RequireJSON() / RequireStructured() |
the output contract was not met |
NoPlaceholders() |
it handed back TODO, [insert x], lorem ipsum |
MaxSteps(n) / MaxCost(usd) |
it got there, but not within budget |
Custom(fn) |
your own rule — return False or (False, "why") |
Sub-agents carry their own, declared in the spec so it stays serialisable:
SubAgentSpec(
name="researcher",
description="Finds things out.",
guardrails={"require_citation": True, "forbid_tools": ["publish"],
"max_retries": 1},
)
And AgentGuardrails(content=Guardrails(...)) gives one agent stricter text
rules than the rest of the harness.
The rails
from agent_harness import (Harness, Budget, PolicyGate, HookEngine, Guardrails,
console_exporter)
harness = Harness.local(".harness") # sessions, memory, traces, checkpoints
harness.policy = PolicyGate("allow", ask=["issue_refund"], deny=["shell"],
approver=my_approver)
harness.guardrails = Guardrails(strict=True)
harness.tracer.add_exporter(console_exporter())
harness.reset_budget(Budget(max_usd=0.50, max_steps=8, max_subagents=4))
hooks = HookEngine()
@hooks.on("pre_tool")
def cap_refunds(ctx):
if ctx.data["tool"] == "issue_refund" and ctx.data["args"]["amount"] > 100:
ctx.block("refunds over 100 EUR need a manager")
agent = Agent("refunds", harness=harness, hooks=hooks, tools=[issue_refund])
print(harness.report()) # spend by agent and task, cache hit rate, concurrency
| Rail | What it does |
|---|---|
PolicyGate |
allow / ask / deny per action, glob rules, conditional on arguments, approver callback |
BudgetGuard |
spend, token, step, tool-call and sub-agent ceilings; child guards roll up to the parent |
RateGuard |
requests- and tokens-per-minute pacing, so you are not rate-limited by the provider |
HookEngine |
13 events; pre_tool can block or rewrite arguments, post_tool can rewrite the result, model_egress sees the real provider and region of every model call |
Guardrails |
secret redaction, private-key blocking, injection warnings, size caps — on tool output and final answers |
StopController |
abort a run and drain the sub-agents; a human is always in charge |
Tracer |
one span per run, step, model call, tool and sub-agent; console and JSONL exporters |
AuditTrail |
immutable, hash-chained who-did-what; verify() names the first tampered entry |
ServiceHealth |
latency, failure rate and saturation per model, tool and MCP server |
RunJournal |
what each agent was asked and what it returned, append-only |
ResultCache |
identical task + identical input served from cache, memory and disk tiers |
Checkpointer + Replayer |
step snapshots, a timeline, and resume-from-any-step |
RecordingProvider / ReplayProvider |
record a run once, reproduce it exactly with no network and no spend |
DeliverableStore |
the documents and reports a run produced, versioned and digested |
SessionStore |
resume, fork or branch a run; a long job survives a restart |
WorkspaceBroker |
a jailed directory per sub-agent (or a shared one for handovers), local or Docker |
ConcurrencyScheduler |
semaphore, queue, backpressure, peak tracking |
ModelRouter |
per-task model and effort tier instead of one model for everything |
SpecCompiler |
a sub-agent blueprint → the exact provider payload, inspectable before you spend |
Path safety is enforced, not clamped: a workspace tool given ../../etc/passwd
refuses rather than resolving it. shell is absent unless the workspace was
created with allow_shell=True, and even then it asks for approval.
Governance: the laws your agents run under
import os
from agent_harness import Agent, Harness
from agent_harness.governance import Governance, AgentIdentity
gov = Governance.from_packs(
["eu-ai-act", "gdpr", "ksa-pdpl", "owasp-agentic"],
policy="governance.yaml", # your rules on top of the packs
home="eu",
regions={"azure": "eu"}, # where endpoints that do not say, are
signing_key=os.environ["AUDIT_KEY"],
)
harness = Harness.local(".harness", governance=gov)
support = Agent("support", "Resolve order problems.", harness=harness,
tools=[order_status, issue_refund],
identity=AgentIdentity(owner="cx-lead@acme.com", purpose="customer_support",
tools=["order_status", "issue_refund"]))
print(gov.report("eu-ai-act").markdown()) # control → status → evidence
Packs switch on what a law or framework asks for. Your policy adds rules on top. Combining them can only make things stricter: the strictest effect wins, so adding a pack never loosens anything. Twenty-six packs ship, each dated and sourced:
| Region | Packs |
|---|---|
| EU & UK | eu-ai-act · gdpr · dora · nis2 · uk-gdpr |
| Gulf | uae-pdpl · difc-reg10 · adgm · ksa-pdpl · sdaia · qatar-pdppl · bahrain-pdpl · oman-pdpl |
| Asia-Pacific | singapore-agentic · singapore-pdpa · india-dpdp · china-genai · korea-ai-basic · japan-appi · vietnam-ai |
| Americas | us-nist-rmf · colorado-adm · texas-traiga · ccpa-admt |
| Standards | iso-42001 · owasp-agentic |
What it enforces, in the loop, before anything happens:
- Residency, per call and per person. Each model call is checked
after the fallback chain has picked the backend, so the check sees the real
destination: the Bedrock region, the
eu.inference profile, the Vertex location, the endpoint host. An EU customer's data never reaches a US model. The call moves on to the next model in the chain instead. A Saudi customer on the same harness is held to Saudi rules (Trace(tags={"jurisdiction": "sa"})). - Least privilege that survives delegation. A sub-agent acts under the intersection of its own identity and every identity above it. It can never reach a tool its parent could not.
- Purpose limitation and minimisation. A run carries a purpose. Data classes
that purpose does not allow are redacted before the call leaves. They can also
be pseudonymised: the model sees
⟨email:3f9a1c0b2d4e⟩, and the tool gets the real address back. - Exact data classification. Card numbers are Luhn-checked. National IDs are validated by their own check digits: Emirates ID, Saudi ID and Iqama, Aadhaar (Verhoeff), Singapore NRIC, the Chinese resident ID; EU VAT numbers by country prefix. Special-category data (health, religion and so on) is detected by keyword screens, reported at lower confidence, and labelled as such; redacting it removes the whole sentence that carries it. A data-class name a policy misspells is refused at load time instead of never matching.
- Human oversight that fails closed.
require_approvalrules need a quorum of distinct approvers. The person the agent acts for cannot approve their own request. A timeout, an error or a missing approver all count as no. Pass anapprover=callable, or leave it out and answer the queue from your own UI (gov.oversight.pending(),.approve(id, by=...));approval_notify=tells your team when something is waiting. - Agent-specific threats. A tool result carrying prompt-injection signals marks the run untrusted (the flag spreads to the parent run). After that, payments and other sensitive tools wait for a person, and nothing is written to memory. Loops, tool storms and runaway delegation are stopped, a tool that keeps failing has its breaker opened for the rest of the run, and an incident is opened once per cause.
- No way round it. Context summaries go through the same egress check as
the loop's own calls. A sub-agent on a harness without governance cannot be
delegated to, and
Agent.as_tool()nesting obeys the same depth limit. A governance failure refuses the action rather than crashing the run — or letting it through. - Supply chain. Every tool's schema is fingerprinted.
inventory.pin()freezes them. A tool that changes afterwards (an MCP server swapping what a tool does) is reported, or refused withtool_drift: deny. - Prohibited uses are refused at build time. Declaring
domains=["social_scoring"]raises an error instead of producing an agent.
And what it records:
- Every decision, signed. Each decision names the policy version (sha256)
that made it.
signing_key=adds HMAC signatures.Ed25519Signer(pip install agent-harness-adk[governance]) adds public-key signatures, so auditors can verify the trail without being able to write it. - Erasure that keeps the audit intact. Nothing personal enters the trail
in the clear: detected values are tokens under a per-person key, and task
text — which can name anyone — is kept only as a token
(
audit_args="tokens"does the same for every tool argument).await gov.rights.erase(trace)deletes their memory, sessions, deliverables, journal entries and trace spans, and destroys that key. The chain still verifies, but nothing in it can be linked back to them. You get a signed receipt;gov.rights.access(trace)exports everything held about the person. - Transparency.
result.disclosureholds the AI notice in the languages the packs need.result.provenanceis a signed, machine-readable manifest (EU AI Act Art 50, China GB 45438). Visible labels are applied only where the law asks for them. - Incidents with the regulator's deadlines. Opening one works out every clock that applies: GDPR 72 h, DORA 4 h / 72 h / 1 month, EU AI Act Art 73, NIS2, SDAIA, PDPC and the rest.
- Evidence.
gov.report(pack)shows each requirement as met, partial, gap or manual. It is worked out from the live configuration and the audit trail, with the next action for every gap. It also produces an AI inventory (to_cyclonedx()), a DORA third-party register, and draft DPIA / FRIA text (gov.impact_assessment(agent)).
# governance.yaml
packs: [eu-ai-act, gdpr, owasp-agentic]
home: eu
signing_key_env: AUDIT_KEY
policy:
purposes:
customer_support: [contact, financial] # nothing else reaches the model
rules:
- id: refunds-need-a-human
on: tool
match: {tags: [payments]}
when: "args.amount > 100" # a small, safe expression language
effect: require_approval
approvers: 2
timeout: 15m
Runnable end to end with no key: examples/07_governance.py (an EU bank also
serving Saudi customers), 08_governance_saudi_government.py (in-Kingdom
residency, four-eyes permits, SDAIA breach clock, DPIA draft) and
09_governance_singapore_fintech.py (monitor-then-enforce rollout, prompt
injection, inherited authority, tool drift).
mode="monitor" makes and records every decision without enforcing any of them.
That is the way to roll governance out on a live system. Without governance
attached the loop pays nothing, and import agent_harness does not load it.
agent-harness governance packs # what ships
agent-harness governance pack ksa-pdpl # requirements, rules, residency, clocks
agent-harness governance check --config governance.yaml # validate, with warnings
agent-harness governance report --config governance.yaml --pack gdpr
agent-harness governance inventory --config governance.yaml --agents agents.yaml
agent-harness governance verify --state .harness --key-env AUDIT_KEY
agent-harness governance dsar erase --user alice --tenant acme --state .harness \
--config governance.yaml
This layer supplies technical controls and the evidence for them. It is not legal advice and does not make anyone compliant by itself. The reports say so, and list what is left for people: the DPO, the EU database registration, the signed DPIA.
Providers
Agent("a", model="claude-opus-5") # → Anthropic
Agent("b", model="gpt-4.1") # → OpenAI
Agent("c", model="gemini-2.5-pro") # → Gemini
Agent("d", model="grok-4") # → xAI
Agent("e", model="llama3.2", provider="ollama")
The provider is inferred from the model id, or named. They live in
agent_harness.llm_providers: the three direct APIs, Bedrock, Vertex (Claude
and Gemini), Azure OpenAI and Azure AI Foundry, and every OpenAI-compatible
vendor as a preset — OpenRouterProvider, GroqProvider, TogetherProvider,
DeepSeekProvider, MistralProvider, XAIProvider, FireworksProvider,
CerebrasProvider, OllamaProvider, LMStudioProvider, VLLMProvider — plus
OpenAICompatibleProvider(base_url=...) for any gateway or proxy.
register_provider("name", MyProvider) adds your own.
What does each provider need?
from agent_harness import list_llm_providers, describe_llm_provider, check_llm_provider
for spec in list_llm_providers(): # every provider, one ProviderSpec each
print(spec.name, spec.configured, [f.name for f in spec.required_fields])
print(describe_llm_provider("azure").render())
Azure OpenAI — provider='azure' (aliases: azure-openai)
OpenAI models deployed in your own Azure OpenAI resource.
status needs setup — missing endpoint (or set $AZURE_OPENAI_ENDPOINT), one of: api_key ($AZURE_OPENAI_API_KEY) | credential
auth api_key or Entra ID
can embeddings, json_schema, streaming, thinking, tools, vision
fields
endpoint required url ← $AZURE_OPENAI_ENDPOINT
The resource endpoint from the Azure portal.
deployment optional str ← $AZURE_OPENAI_DEPLOYMENT
The deployment name. Defaults to the model id.
...
example get_provider('azure', endpoint='https://my-resource.openai.azure.com', deployment='gpt-4.1-prod')
Every provider declares its connection fields as ProviderFields — which are
required, which are secret, which are alternatives to each other, and which
environment variables they are read from — so you can see what a backend needs
before you touch it. list_llm_providers(configured_only=True) shows what is
ready to use now; capability="embeddings" filters by what a provider can do.
check_llm_provider("bedrock", region="eu-west-1") answers "would this
connect?" offline, as a ProviderCheck: what is missing, where each setting
was found (argument or $ENV_VAR — never the value), and any argument that is
not a setting at all. await ping_llm_provider("openai") connects for real,
using the free model-listing endpoint where there is one, and returns a result
rather than raising. From a terminal:
agent-harness providers # the table: status and what each needs
agent-harness providers bedrock # every field, env var, capability, an example
agent-harness providers openai --ping # are these credentials any good?
When a call fails
Every provider sends through one transport, so every one gets the same behaviour:
- Retries on 408, 409, 425, 429, 5xx and Anthropic's 529, on timeouts, and on dropped connections — exponential back-off with jitter, so a hundred throttled sub-agents do not retry in lockstep.
- Every Retry-After hint is honoured:
Retry-Afterin seconds or as a date,retry-after-ms, OpenAI'sx-ratelimit-reset-*, Gemini'sRetryInfo, and thex-should-retryverdict OpenAI and Anthropic send. A server asking for longer thanmax_retry_afterfails the call at once, so a fallback model can take over instead of the run sleeping for ten minutes. - A shared cool-down: when one call is told to wait, the calls running alongside it on the same provider wait too.
- Streams recover — a stream that fails before its first token (Anthropic's
mid-stream
overloaded_error, a throttled Bedrock stream) is restarted; one that fails after is raised, because replaying it would repeat text you have already shown. - A circuit breaker: after five consecutive failures a provider is not
called for 30 seconds — calls fail at once with
ProviderUnavailableError, and the model router moves to the next model immediately. A 400 is the caller's fault and does not count. - Stale credentials are refreshed once on a 401 — Vertex and Entra ID tokens, and Bedrock credentials from the AWS chain — and every Bedrock retry is signed afresh.
from agent_harness import AnthropicProvider, RetryPolicy, CircuitBreaker
provider = AnthropicProvider(
retry=RetryPolicy(max_retries=5, initial_delay=1.0, max_delay=30,
max_retry_after=60, max_elapsed=180),
circuit_breaker=CircuitBreaker(failure_threshold=3, reset_timeout=60),
max_concurrency=16, # cap on in-flight requests
timeout=120, connect_timeout=5, # per attempt; CompletionRequest.timeout per call
on_retry=lambda e: print(f"{e.provider} retry {e.attempt} in {e.delay:.1f}s: {e.reason}"),
)
provider.health() # requests, retries, rate_limited, timeouts, last_request_id, circuit state
on_retry receives a RetryEvent; retries are also logged on the
agent_harness.llm_providers logger. What finally fails is typed, so you can
handle the cases differently — all are ProviderErrors, and each carries
status, retryable, retry_after, attempts and the vendor's request_id:
| Error | Meaning | Retried |
|---|---|---|
RateLimitError |
429 or throttling | yes |
QuotaExceededError |
out of credit or quota — waiting won't help | no (falls back) |
ProviderTimeoutError |
no answer in time | yes |
ProviderConnectionError |
DNS, TLS, a dropped connection | yes |
ProviderUnavailableError |
5xx, 529 overloaded, or the circuit is open | yes |
AuthenticationError |
key, token or signature refused | no |
ContextWindowExceededError |
the prompt is longer than the model reads | no |
ModelNotFoundError |
no such model for this account | no |
InvalidRequestError |
anything else the request got wrong | no |
The same models, on your cloud
from agent_harness import BedrockProvider, VertexProvider, AzureOpenAIProvider
# AWS Bedrock — SigV4 signed, no API key
Agent("support", model="anthropic.claude-opus-5",
provider=BedrockProvider(region="eu-west-1"))
# Google Vertex AI
Agent("support", model="claude-opus-5",
provider=VertexProvider(project="my-project", region="europe-west1"))
Agent("support", model="gemini-2.5-pro",
provider=VertexGeminiProvider(project="my-project"))
# Azure OpenAI, and Azure AI Foundry
Agent("support", provider=AzureOpenAIProvider(
endpoint="https://my-resource.openai.azure.com",
deployment="gpt-4.1-prod", api_version="2024-10-21"))
Agent("support", provider=AzureFoundryProvider(
endpoint="https://my-project.services.ai.azure.com",
deployment="claude-opus-5"))
Credentials follow each platform's own conventions: AWS_* environment
variables, a Bedrock API key (AWS_BEARER_TOKEN_BEDROCK), or the botocore chain
(instance roles, SSO) if boto3 happens to be installed; google-auth or gcloud auth print-access-token; an api-key or
credential= from azure-identity for Entra ID. None of those libraries are
required — SigV4 is implemented against AWS's published test vectors using only
the standard library.
Platform model ids are normalised, so anthropic.claude-opus-5 on Bedrock,
us.anthropic.claude-opus-5 on a cross-region profile and claude-opus-5@20260401
on Vertex are all recognised as the same model and priced the same — cost
attribution keeps working wherever a model is served from.
Every generation parameter
Agent("deep",
model="claude-opus-5",
effort="high", # none · minimal · low · medium · high · xhigh · max
thinking=True,
thinking_budget=16_000, # for models that take a budget, not a level
temperature=0.2, top_p=0.9, top_k=40, min_p=0.05,
frequency_penalty=0.3, presence_penalty=0.1, repetition_penalty=1.05,
seed=7, max_tokens=4096, stop=["</answer>"],
cache=True, # prompt caching where the provider has it
user="customer-42", # for abuse tracing
model_options={"speed": "fast", "safety_settings": [...]})
The same names work on AgentVersion, SubAgentSpec, blueprint YAML and
CompletionRequest. They are validated when the agent is built —
temperature=5 or top_p=1.5 raises a ConfigurationError naming every bad
value and its allowed range, instead of a 400 halfway through a run.
validate_parameters(...) runs the same check on its own.
Each provider then turns them into what it accepts:
- A parameter a provider does not have is dropped, not translated. Anthropic
has no
seedor penalties; OpenAI has notop_k;min_pandrepetition_penaltyreach only the hosts that take them (vLLM, Together, Fireworks, OpenRouter).describe_llm_provider(name).parameterslists them. - A value a model would reject is fitted, not sent to fail. Claude takes
temperature0–1 and one oftemperature/top_p; with extended thinking it refusestemperatureandtop_kand needstop_p≥ 0.95, a budget of at least 1024, andmax_tokensabove the budget. Thinking-only models (Opus 5, Sonnet 5 — under any Bedrock or Vertex id) get no sampling at all. OpenAI's reasoning models refuse sampling andstop. Grok's reasoning models refuse penalties andstop. Gemini 2.5 Pro cannot switch thinking off. effortis the portable way to ask for reasoning, mapped to the nearest level each model has:
| Provider | effort becomes |
|---|---|
| Anthropic, Bedrock, Vertex | output_config.effort (low–max); none switches thinking off |
| OpenAI o-series / GPT-5 | reasoning_effort (low–high; GPT-5 also minimal) |
| Gemini 2.5 | a thinking budget: 0 / 512 / 1024 / 8192 / 16384 / 24576 / 32768, fitted to the model's range |
| Gemini 3 | thinkingLevel (Pro: low, high; Flash: minimal–high) |
| OpenRouter | reasoning.effort, or reasoning.max_tokens for a budget |
| gpt-oss on Groq, Together, Fireworks, Cerebras, Ollama, vLLM | reasoning_effort |
| grok-3-mini | reasoning_effort (low, high) |
Nothing is dropped or changed silently. Every decision is recorded in a
ParameterPlan, logged once on agent_harness.llm_providers, and shown by
explain — without sending anything:
OpenAIProvider().explain(CompletionRequest(model="o3", messages=[...],
temperature=0.3, effort="max"))
# {"sent": {"effort": "high", "max_tokens": 8192},
# "dropped": {"temperature": "o3 is a reasoning model and rejects sampling settings"},
# "adjusted": {"effort": "'max' → 'high': o3 takes low, medium, high"},
# "payload": {...exactly what would go on the wire...}}
extra={...} is merged into the payload verbatim for anything not covered —
beta headers, new fields, a provider feature that shipped this morning.
Adapters normalise everything the loop depends on: tool calls, tool results,
thinking blocks (with their signatures, so extended thinking survives a tool
loop — including Gemini's thought signatures and Anthropic's redacted
thinking), cache tokens, stop reasons and refusals. All of them stream,
including Bedrock (its binary event stream is decoded and CRC-checked) and
Vertex. Cost is computed per
call from a built-in price table (register_model to extend it), so
result.cost_usd is real money, not an estimate.
Stopping, reproducing, and proving it got better
A human is always in charge:
harness.stop("the customer withdrew the request") # drains; nothing new starts
harness.control.abort("pull the plug") # cancels what is in flight
The loop checks between steps and before every tool, so a stop lands at a safe boundary and the work already done is kept.
Reproduce a failure before you fix it:
from agent_harness import RecordingProvider, ReplayProvider
agent = Agent("support", provider=RecordingProvider(AnthropicProvider(), "run.jsonl"))
await agent.run("...") # once, against the real model
replay = ReplayProvider("run.jsonl") # then as often as you like: no network, no spend
twin = Agent("support", provider=replay)
assert (await twin.run("...")).output == original.output
Or travel back to any step and try it differently:
print(await harness.replayer.timeline(result.run_id))
again = await harness.replayer.resume(agent, result.run_id, step=3,
task="Give the figure, not a summary.")
And prove a change actually helped:
from agent_harness import Evaluator, Expect, GoldenTask
suite = Evaluator([
GoldenTask(id="refund-window", input="Can I refund a 40-day-old order?",
expect=Expect(contains=["30-day"], not_contains=["yes, of course"])),
GoldenTask(id="uses-lookup", input="Where is order 4182?",
expect=Expect(tool_called="order_status", max_steps=4)),
])
baseline = await suite.run(agent, label="before"); baseline.save("baseline.json")
# ... change the prompt ...
after = await suite.run(agent, label="after")
print(after.compare(baseline).render())
# REGRESSED: 100.00% → 50.00% (-50.00%)
# REGRESSED: uses-lookup
Expectations can check the text, the tools that were called, the structured
output, the step count or the cost. llm_judge is there for genuinely
open-ended answers — reach for it last; it costs money and it can be wrong.
Streaming and structured output
async for event in agent.stream("Summarise the incident"):
if event.type == "text":
print(event.text, end="", flush=True)
elif event.type == "tool_result":
print(f"\n· {event.data['tool']}")
elif event.type == "run_end":
result = event.data["result"]
from pydantic import BaseModel
class Ticket(BaseModel):
id: str
priority: int
summary: str
agent = Agent("triage", output_type=Ticket)
result = await agent.run("Customer cannot log in since the deploy")
result.data.priority # a validated Ticket, retried if the model got it wrong
Testing your agents
from agent_harness import Agent, FakeProvider, Harness, tool_call
provider = FakeProvider([tool_call("order_status", order_id="4182"),
"It ships Thursday."])
agent = Agent("support", provider=provider, harness=Harness.testing(provider),
tools=[order_status])
result = await agent.run("Where is order 4182?")
assert result.output == "It ships Thursday."
assert provider.requests[0].system.startswith("You are support")
No network, no keys, no recorded cassettes. Script strings, tool calls, whole messages, exceptions, or a callable that inspects the request and answers accordingly. The harness's own suite is 296 tests and runs in half a second.
CLI
agent-harness run "summarise this incident" --tools --stream --state .harness
agent-harness chat --skills ./skills --state .harness --approve
agent-harness models
agent-harness sessions --state .harness
agent-harness journal --state .harness
agent-harness mcp npx -y @modelcontextprotocol/server-filesystem /data
Design notes
- Async core, sync wrapper. Parallel sub-agents, MCP and the concurrency cap
all need it.
run_synccovers scripts. - Compaction never orphans a tool call. Fat tool results are hollowed out
first, and the summarise-the-head fallback moves its cut forward until no
tool_resultis left without itstool_use. Naive trimming corrupts a conversation; this does not. - Least privilege by default. Sub-agents get an explicit tool allowlist,
delegation does not cascade,
shellis opt-in, and a tool's own permission can only tighten the policy. - Everything is optional. An
Agentwith no memory, no skills and no sub-agents is a tightwhileloop around one model call.
Contributing
git clone https://github.com/MuhammadHusnainAli/agent-harness-adk
cd agent-harness-adk
uv sync --extra dev
uv run pytest -q # no API key needed — everything runs on FakeProvider
uv run ruff check src tests examples
Every push to main runs the suite on Python 3.10, 3.11, 3.12, 3.13 and 3.14,
lints, builds the wheel, installs it into a clean environment and smoke-tests
it, and runs every example without an API key.
CONTRIBUTING.md has the details — including the three things this project is picky about: the dependency count, the import time, and the 3.10 floor.
| Report a bug or ask for a feature | Issues |
| Ask how to do something | Discussions · SUPPORT.md |
| Report a vulnerability | Privately — SECURITY.md |
| Community standards | CODE_OF_CONDUCT.md |
| Cut a release (maintainers) | RELEASING.md |
Running agents safely — tool allowlists, policy gates, workspace isolation and what this library does not defend against — is covered in SECURITY.md. Worth reading before you give an agent a tool that writes, spends or sends.
Status
0.1.0 — the first release. The public API above is what we intend to keep. Changes are recorded in CHANGELOG.md.
Not in this release: a vector-database backend (the built-in index is exact brute force, fine to ~50k records) and provider-side batch APIs. OCR, PDF and DOCX parsing work through an optional install each rather than shipping in the default dependency set.
Licence
MIT — see LICENSE.
Release files for agent-harness-adk 0.1.4
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| agent_harness_adk-0.1.4.tar.gz | 316.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| agent_harness_adk-0.1.4-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 689.2 kB
Release files / agent_harness_adk-0.1.4.tar.gz
| Download URL | agent_harness_adk-0.1.4.tar.gz |
|---|---|
| Size | 316.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
065e1ed4049e8ca0807a9dc212043c75afeac21f00edc60388d5b37b9c1c86be
|
|
BLAKE2b-256 checksum How to use checksums |
edb7cfb0f7cee753875f9bbf3b98b0d72ffd9a563fb89a0afe0c1e8ea778788e
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 24, 2026.
Transparency logRelease files / agent_harness_adk-0.1.4-py3-none-any.whl
| Download URL | agent_harness_adk-0.1.4-py3-none-any.whl |
|---|---|
| Size | 373.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
338c0c38551cec99655ce03d45c7704751690220967231f08914442398b57d25
|
|
BLAKE2b-256 checksum How to use checksums |
e24986c151e0e3e279aea03eb30e0e8f20f7644ce1a3fd61b7a24abc52afa3ad
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 24, 2026.
Transparency log