Temporal Agent Harness
Build durable, composable AI agents with a rich tool-approval policy engine that can seamlessly elevate to a human with built-in human-in-the-loop.
⚠️ Experimental. An early, fast-moving project from Temporal Technologies. APIs will change.
The Temporal Agent Harness gives your agents capabilities that are painful and error-prone to build yourself:
- agents survive crashes and resume mid-turn, exactly where they left off (no wasted tokens!);
- a rich tool-approval policy engine decides exactly when a tool call needs human sign-off — handing control to a human only then, and resuming the moment it's granted —
- agents compose programmatically, through real typed contracts (not limited to just text in, text out);
- Code Mode — one tool that runs a Python script over your toolset — so a single turn orchestrates many tool calls with real control flow and concurrency;
- callback tools let an agent invoke a tool that runs on the client — reading a file on a user's laptop, capturing a photo on their phone — even though the agent runs on a remote worker;
- agents are fully observable — a standardized, full-lifecycle event stream lets you watch them live or replay exactly what they did;
all while you write the actual turn logic with the AI SDKs you already know.
Every agent is a durable Temporal workflow at its core, giving agents access to the full breadth of Temporal primitives. Tools as activities or workflow functions, every turn a streamed, replayable history. The harness packages those primitives into a toolkit built for first-class agent development, so you get the power without hand-rolling the orchestration.
Installation
There are two ways in, depending on what you want to do. Both pin to a tagged release — see Versioning and stability for why that matters.
Try it — run the example agents
Clone at a release tag. No Node/pnpm needed: the browser UI ships prebuilt.
git clone --branch 0.3.0 https://github.com/temporal-community/temporal-agent-harness.git
cd temporal-agent-harness
Git will note that you're in "detached HEAD" — that's expected, it just means you're sitting on
the tag rather than on a branch. Later, move to a newer release with
git fetch --tags && git checkout <version>, or see what changed between two of them with
git diff 0.1.0 0.2.0.
Then jump to Run the examples. You'll need uv, just, and the Temporal cli (or Temporal Cloud).
Build with it — add the harness to your own project
Add it as a git dependency pinned to a tag. In a uv-managed
project:
# core harness — define and run agent workflows
uv add "temporal-agent-harness @ git+https://github.com/temporal-community/temporal-agent-harness.git@0.3.0"
Or declare it in pyproject.toml — depend on the package (with any extras you need) and point
its source at the tag:
[project]
dependencies = [
"temporal-agent-harness[ui]",
]
[tool.uv.sources]
temporal-agent-harness = { git = "https://github.com/temporal-community/temporal-agent-harness.git", tag = "0.3.0" }
Then run uv sync.
Extras:
ui— the reusable FastAPI server and packaged browser UI (pulls infastapi[standard], including Uvicorn). The built Svelte assets are always in the artifact; only the server runtime dependencies are gated behind this extra, so core agent-worker installs stay smaller.code-mode— for workers that host Code Mode agents; pulls inpydantic-monty, the sandbox the scripts run in. (The workflow-sideagent.code_mode_toolfactory itself needs nothing extra, so importing it never requires this dependency.)
Combine extras in the dependency spec, e.g. "temporal-agent-harness[ui,code-mode]".
Agent authors use the harness runtime from temporal_agent_harness.harness. Applications that
want the built-in session manager and UI use temporal_agent_harness.web:
from temporalio.client import Client
from temporalio.envconfig import ClientConfig
from temporal_agent_harness.plugin import AgentHarnessPlugin
from temporal_agent_harness.web import (
create_agent_harness_app,
create_session_manager_worker,
)
async def run_session_manager() -> None:
connect_config = ClientConfig.load_client_connect_config()
client = await Client.connect(**connect_config, plugins=[AgentHarnessPlugin()])
worker = create_session_manager_worker(client)
await worker.run()
app = create_agent_harness_app(registry_path="agents.toml")
Then serve the app with Uvicorn:
uvicorn my_app.web:app --host 0.0.0.0 --port 8000
The registry lists the launchable agents the UI can create:
[[agents]]
key = "my-agent"
workflow_type = "MyAgent"
task_queue = "my-agent-task-queue"
label = "My Agent"
description = "A short description shown in the UI."
The app factory serves both /api/* and the packaged Svelte UI. The helper
create_session_manager_worker only registers the packaged session-manager
workflow; run your own agent workflows on their own workers and task queues.
Versioning and stability
Install from a tagged release. Every release is listed on the releases page, and each tag marks a commit that is known-good at the moment it was cut: the tests passed, and the prebuilt browser UI matches the source it was built from.
This project is very early and experimental - main may break at any time.
Releases are cut manually, when the maintainers judge the state stable. This project is experimental and pre-1.0: APIs will change between releases, without warning. Pinning to a tag is what keeps that churn from reaching you unannounced — so pin, and upgrade deliberately when you want to try out new harness capabilities.
What you get
🛡️ Durable execution — the foundation
Every agent is a Temporal workflow, so durability isn't a feature you add, it's the ground you build on. A worker can crash, redeploy, or restart mid-turn and the agent resumes precisely where it was — no lost state, no double-run tool calls. Model and tool calls retry by policy; an agent can wait minutes or days for an external event without holding a process open; and every turn, tool call, and decision is recorded and replayable.
📊 Standardized agents, fully observable
The interface to agents built on this harness is carefully standardized. Every agent takes the same configuration contract and emits the same structured event stream: a protocol (under active development) that captures an agent's entire lifecycle — every turn, model interaction, tool call (start / end / error), reply token, citation, approval decision, subagent hand-off, and token-usage tally.
That one standardized stream is the foundation of observability and analytics for every agent you build. Watch an agent live as it works, or replay exactly what happened afterward — what it decided, which tools it ran, what it cost, and where a human stepped in. You instrument once; every agent on the harness gets it.
🙋 Human-in-the-loop, solved
Tool approvals are built in and safe-by-default: any tool call can require human sign-off, and a gated call pauses inside the workflow and resumes durably whenever a decision arrives (no matter how long that takes) — there's no approval queue, state machine, or callback plumbing for you to build. The policy engine is sophisticated out of the box: layered rules, inherently-safe auto-approval, per-tool allow-lists, "approve and stop asking," per-session overrides, runtime policy updates, and custom predicates.
🔌 Bring your own AI SDK
Write turn logic with the SDK you already know. The harness's integrations turn each SDK call into a durable Temporal activity — so retries and credentials never leak into your workflow code. Support is growing across the Python AI SDKs and agent frameworks Temporal integrates with:
| AI SDK | Status | Notes |
|---|---|---|
| Google Gemini | ✅ Available now | Ships in this repo and is experimental - Python SDK has a fully-supported non-harness integration. |
| OpenAI Agents SDK | ✅ Available now | Ships in this repo and is experimental - Python SDK has a fully-supported non-harness integration. |
| Pydantic AI | ✅ Available now | Directly uses Pydantic's Temporal plugin |
| Google ADK | 🟡 Planned | - |
| Strands Agents | 🟡 Planned | - |
| LangGraph | 🟡 Planned | - |
📡 Durable and inline tools
Tools come in two on-worker flavors — durable, activity-backed tools (@agent.activity_tool_defn)
that run as retried, observable Temporal activities, and inline workflow tools
(@agent.tool_defn). Each publishes its own start/end lifecycle events onto the agent's
standardized event stream. (A third flavor — callback tools — runs on an attached client
instead of the worker; see below.)
📞 Callback tools — let the client run the tool
An agent running on a Temporal worker often needs to act somewhere it can't reach — a file on the
user's laptop, a photo from their phone, a device on a private network. A callback tool
(@agent.callback_tool_defn) has no worker-side body: the agent pauses inside the workflow,
publishes the call, and an attached client executes it on its own machine and sends the result
back. You declare only the tool's typed contract (its ... body is enforced) — the harness
supplies the single generic implementation. Because it's dispatched like any other tool, a callback
tool inherits the same approval policy, tool_start/tool_end events, and durable pause/resume:
the workflow simply waits (seconds or days) until a result arrives, and that result is validated
against the tool's declared output type before the turn continues. See
examples/callback_tools/wiki_agent — a cloud-shaped agent
that organizes a Markdown wiki on your local disk through a thin terminal client.
🧩 Agents that are more than chatbots
Most frameworks treat an agent as a single text-in / text-out function. Here, an agent exposes a
strongly-typed interface: declare named operations with @agent.accepts over pydantic models
(plain text is just one shape). The agent advertises its own callable surface — operation
names and input/output schemas — so it's self-describing and ready to be driven programmatically,
by your code or by another agent.
🔗 Composable by construction
Because agents have typed, self-describing interfaces, any harness agent can become a strongly-typed toolset another agent drives — start it, call its operations, stop it. Multi-agent systems compose through real contracts, not strings pasted between prompts.
🧑💻 Code Mode — give the model code, not just a call menu
Hand a model one tool that runs a Python script over your existing tools, instead of a long
menu of individual calls. agent.code_mode_tool([...tools...]) turns any set of harness tools
into a single run-a-script tool: the model writes Python that calls them as async host functions,
with real control flow — loops, conditionals, min/max, and asyncio.gather concurrency — so
one turn orchestrates many calls. Each host call is still dispatched through the runner as a
durable, approval-gated, observable activity, and the script is statically type-checked against
your tools' signatures before it runs. And since a subagent toolset is just a list of tools,
Code Mode composes over subagents for free.
A taste
from datetime import timedelta
from pydantic import BaseModel
from temporalio import workflow
from temporalio.contrib.workflow_streams import WorkflowStream
from temporalio.workflow import ActivityConfig
from temporal_agent_harness.harness import AgentWorkflowRunner, agent, slash_commands
from temporal_agent_harness.harness.agent_protocol import AgentConfig, ToolApprovalPolicy
# A durable, activity-backed tool: runs as a retried, observable Temporal activity and
# publishes its own tool_start/tool_end events on the turn stream.
@agent.activity_tool_defn(
activity_config=ActivityConfig(start_to_close_timeout=timedelta(seconds=30)),
)
async def search_flights(origin: str, destination: str, date: str) -> str:
...
# Strongly-typed messages — an agent operation is more than a string in and a string out.
class PlanTrip(BaseModel):
destination: str
nights: int
class Itinerary(BaseModel):
summary: str
total_usd: float
@agent.defn
class TravelAgent:
@workflow.init
def __init__(self, config: AgentConfig) -> None:
# Tool approvals are safe-by-default; here, auto-approve only tools that
# statically declare themselves inherently safe.
self._runner = AgentWorkflowRunner(
config,
stream=WorkflowStream(),
approval_policy_default=ToolApprovalPolicy.allow_inherently_safe(),
slash_commands=slash_commands.default_commands(),
)
@workflow.run
async def run(self, config: AgentConfig) -> None:
await self._runner.run(self)
# A typed, self-describing operation. The agent advertises this signature, so callers —
# your code or another agent — can drive it programmatically. Your turn logic goes here:
# call your AI SDK, run tools through the runner, and return the typed reply.
# (See examples/monty for a complete, model-in-the-loop agent on the Gemini integration.)
@agent.accepts
async def plan_trip(self, request: PlanTrip) -> Itinerary:
...
Running a worker — one plugin
An agent is a Temporal workflow, so it runs on a Temporal worker. AgentHarnessPlugin is the
single registration that wires that worker (and its client) for the harness — you never
assemble the harness's activity list or its data converter by hand:
from temporalio.client import Client
from temporalio.envconfig import ClientConfig
from temporalio.worker import Worker
from temporal_agent_harness.plugin import AgentHarnessPlugin
client = await Client.connect(
**ClientConfig.load_client_connect_config(),
# Your AI SDK's plugin first, the harness plugin last.
plugins=[OpenAIAgentsPlugin(model_params=...), AgentHarnessPlugin(tools=MY_TOOLS)],
)
worker = Worker(client, task_queue="my-agent", workflows=[TravelAgent])
await worker.run()
The worker declares only its workflows. Adding the plugin brings:
- the harness's data converter (Pydantic + large-payload offload) — add the plugin to every client, worker, and server in the deployment so they all agree on it;
- the durable activity body of each
@agent.activity_tool_defntool intools=(pass your whole toolset — tools with no worker-side body are skipped); - the subagent and Code Mode activities.
Order it last, after any AI SDK's plugin, so that SDK's payload converter wins. Registering it on the client is enough — Temporal applies a client's plugins to workers built from it.
Code Mode
Most agents call tools one at a time — a round-trip per call. Code Mode hands the model a
single tool that runs a Python script over your tools, so one turn can search, filter, branch,
and act across many calls with ordinary control flow and asyncio.gather concurrency.
agent.code_mode_tool(tools, name=...) takes any list of harness tools and returns one inline
tool. Its generated description tells the model the sandbox contract and every host function's
signature + result shape — derived from your tools, so you never hand-write or maintain it. Hand
the returned tool to your model's tool-calling loop like any other tool.
from temporal_agent_harness.harness import agent
# Any @agent.activity_tool_defn / @agent.tool_defn tools — including a subagent toolset.
run_code = agent.code_mode_tool(
[search_flights, search_hotels, book_flight, get_trip_summary],
name="run_travel_code",
)
# The model then writes, e.g., a script like this and calls run_travel_code with it:
#
# import asyncio
# async def main():
# flights, hotels = await asyncio.gather( # independent calls run concurrently
# search_flights({"origin": "SFO", "destination": "JFK", "date": "2026-07-01"}),
# search_hotels({"city": "New York", "check_in": "2026-07-01", "check_out": "2026-07-05"}),
# )
# cheapest = min(flights["flights"], key=lambda f: f["price_usd"])
# return await book_flight({"flight_id": cheapest["flight_id"], "passenger_name": "Ada Lovelace"})
# asyncio.run(main())
- Durable, gated, and observable per call. The script runs in a sandbox; each host call is
dispatched back through the runner as its own durable activity — keeping that tool's approval
policy and
tool_start/tool_endevents. Writing the script is inert; only the host calls act. - Type-checked before it runs. Code Mode generates static type-check stubs from your tools' signatures, so a wrong argument or an unknown result key comes back as an error to fix rather than a bad run.
- Composes over subagents.
agent.subagent_toolset(...)returns a list of tools, so drop it straight intocode_mode_tool([...])— the model's script can drive subagents too. - Several per agent. Give one agent multiple
code_mode_tools (distinctnames) over disjoint or overlapping tool sets.
A worker that hosts a Code Mode agent needs the two sandbox-stepping activities and the durable
bodies of any activity-backed host tools. Both come from
AgentHarnessPlugin — the stepping activities as soon as the
code-mode extra (which pulls in pydantic-monty,
the sandbox the scripts run in) is installed:
client = await Client.connect(..., plugins=[AgentHarnessPlugin(tools=my_tools)])
worker = Worker(client, task_queue=..., workflows=[MyAgent])
See examples/monty for three agents all built on Code Mode: a no-model script
runner, a conversational agent that writes its own scripts, and a subagent-driven variant.
Slash Commands
Agents can expose human/operator slash commands through a small library of
workflow-safe command definitions. A command bundles the UI metadata returned by
the operator_interface query with the deterministic handler that runs inside
the workflow.
If slash_commands is omitted, AgentWorkflowRunner enables the packaged
defaults:
| Command | Effect |
|---|---|
/approvals strict|safe|skip |
Change the live tool-approval policy. |
/allow-tools tool_name |
Auto-approve one or more named tools for this session. |
/status |
Show the current harness status. |
/stop |
Stop the agent workflow. |
Configure exactly the packaged commands you want in one place:
from temporal_agent_harness.harness import slash_commands
self._runner = AgentWorkflowRunner(
config,
stream=WorkflowStream(),
approval_policy_default=ToolApprovalPolicy.always_require_approvals(),
slash_commands=slash_commands.commands("approvals", "status", "stop"),
)
Pass an empty list to disable packaged slash commands:
slash_commands=[]
Custom commands use the same registry. For example, a model selector can share
one implementation across the first-class operator update path and the normal
slash turn path:
SUPPORTED_MODELS = ("gemini-3.5-flash", "gemini-3.1-flash-lite")
self._runner = AgentWorkflowRunner(
config,
stream=WorkflowStream(),
approval_policy_default=ToolApprovalPolicy.always_require_approvals(),
slash_commands=[
*slash_commands.default_commands(),
slash_commands.model_selector(
choices=SUPPORTED_MODELS,
set_model=lambda model: setattr(self, "_model", model),
description="Set the model for this session.",
),
],
)
Requirements
- Python 3.11+
- uv for dependency management
- just for the example recipes
- A Temporal service.
just temporalstarts a local dev server if you have thetemporalCLI installed.
pnpm is not required to run anything: the browser UI ships prebuilt in
temporal_agent_harness/ui/dist, both in release archives and in the repo. You only need it to
change the UI — see UI development.
Run the examples
These run from a checkout or an unpacked release archive — the steps are identical.
One .env.local at the project root serves every example. Create it once:
cp .env.example .env.local
Set the creds for whichever agents you'll run: OPENAI_API_KEY (react_agent, openai_hello,
pydantic_ai_hello) and/or GEMINI_API_KEY (monty, wiki, coding). The default committed
temporal.local.toml profile points at a local Temporal dev server.
One example, standalone
Each example runs on its own from its directory — examples/monty (a
conversational Code Mode travel agent + a subagent variant) is the best starting point:
cd examples/monty
just temporal # local Temporal dev server; skip if you bring your own
just session-manager # worker hosting the packaged SessionManagerWorkflow
just server # serves the Svelte UI + /api on :8000 (this example's agents only)
just worker # this example's agent worker
Open http://localhost:8000 and pick an agent. Every example follows the same recipe set
(temporal / session-manager / server / worker, plus client where noted). just server
serves the prebuilt UI from temporal_agent_harness/ui/dist, so it needs no Node/pnpm. If you're
changing the UI, use just dev-server (rebuild + serve) instead.
All examples behind one UI
The root justfile runs every example agent at once so the UI lists them all. From the project root, each in its own terminal:
just temporal # start FRESH (or `just reset-manager` first — see the gotcha)
just session-manager # shared session-manager worker
just server # serves the MERGED registry (all agents) on http://localhost:8000
just workers # co-launch all six agent workers (Ctrl-C stops them; or run `just worker-<name>` each)
Then create a session for any agent in the UI. A few need extra setup or a client:
| Agent | Needs |
|---|---|
| OpenAI Hello · Pydantic AI Hello | OPENAI_API_KEY; chat directly in the UI |
| Monty (both) | GEMINI_API_KEY; chat directly in the UI |
| ReAct Agent | OPENAI_API_KEY; the F1 MCP server at F1_MCP_SERVER_HOME (setup); just react-client to answer its ask_user (chat alone works in the UI) |
| Wiki (callback) | GEMINI_API_KEY; just wiki-client --wiki-dir ./wiki — required, or its tool calls hang |
| Coding (callback) | GEMINI_API_KEY; just coding-shim <dir> + the OpenCode TUI — required |
Gotcha — the session manager caches its registry. The server seeds the session-manager
workflow with the registry on first start and reuses the existing one after that. So when you switch
between a single-example server and the all-agents server (or change the set), run
just reset-manager before the next just server, or start a fresh Temporal dev server. Also: an
agent whose worker isn't running will accept a created session but never progress (it parks) — start
its worker.
Status & docs
This is experimental and under active development; expect breaking changes — see
Versioning and stability for what that means for installs. Deeper
design documentation — the agent protocol, the streaming model, human-in-the-loop approvals, and
agents-as-subagents — lives under docs/internal/. Contributor setup
(repository layout, the root justfile, UI development, and packaging) is in
docs/internal/development.md.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file temporal_agent_harness-0.3.0.tar.gz.
File metadata
- Download URL: temporal_agent_harness-0.3.0.tar.gz
- Upload date:
- Size: 474.1 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
715b370b2bd498122e19cebc7d2e7690fe5d4e93f99d40914e199ebb49bd1611
|
|
| MD5 |
9fc73afd5be4b359d203e3e1ec15d230
|
|
| BLAKE2b-256 |
8b97f04a0a026a9d2dd4034c94d50ecbbf6b33ab7b3c2ee7a4bed406040b4ddc
|
Provenance
The following attestation bundles were made for temporal_agent_harness-0.3.0.tar.gz:
Publisher:
publish.yml on temporal-community/temporal-agent-harness
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
temporal_agent_harness-0.3.0.tar.gz -
Subject digest:
715b370b2bd498122e19cebc7d2e7690fe5d4e93f99d40914e199ebb49bd1611 - Sigstore transparency entry: 2791528155
- Sigstore integration time:
-
Permalink:
temporal-community/temporal-agent-harness@3e6ff46cf0b7534a0ed1553ff7198b1979284e3b -
Branch / Tag:
refs/tags/0.3.0 - Owner: https://github.com/temporal-community
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@3e6ff46cf0b7534a0ed1553ff7198b1979284e3b -
Trigger Event:
workflow_dispatch
-
Statement type:
File details
Details for the file temporal_agent_harness-0.3.0-py3-none-any.whl.
File metadata
- Download URL: temporal_agent_harness-0.3.0-py3-none-any.whl
- Upload date:
- Size: 512.8 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
7cf4cf383c7432dcc2658bfef15e96e1c7a75440a933345a5f70ac52dbf45f83
|
|
| MD5 |
0b105bbd95cf798750d9734d07dac8dd
|
|
| BLAKE2b-256 |
ad6b3efc6438865e74533ac0a29710d6af1c91549a314d118830825cc42e9fe3
|
Provenance
The following attestation bundles were made for temporal_agent_harness-0.3.0-py3-none-any.whl:
Publisher:
publish.yml on temporal-community/temporal-agent-harness
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
temporal_agent_harness-0.3.0-py3-none-any.whl -
Subject digest:
7cf4cf383c7432dcc2658bfef15e96e1c7a75440a933345a5f70ac52dbf45f83 - Sigstore transparency entry: 2791528191
- Sigstore integration time:
-
Permalink:
temporal-community/temporal-agent-harness@3e6ff46cf0b7534a0ed1553ff7198b1979284e3b -
Branch / Tag:
refs/tags/0.3.0 - Owner: https://github.com/temporal-community
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@3e6ff46cf0b7534a0ed1553ff7198b1979284e3b -
Trigger Event:
workflow_dispatch
-
Statement type: