workhorse
A fail-soft runner for agent workflows written as Python state machines — drives an agent CLI (Claude, Codex, Copilot, Cline or OpenCode) unattended for days.
A workflow is a Python package: its states are methods that return the next state,
its nodes are plain functions. workhorse drives the machine, renders Jinja2
prompts, invokes the agent CLI, validates JSON replies into typed models,
checkpoints after every transition, and writes run artifacts.
The PyPI distribution is
workhorse-agentand the import package isworkhorse. Workflow distributions provide the commands users run.
Why
workhorse exists to run long, multi-step agent workflows unattended — the
design target is a single run that survives for a week without a human babysitting
it. That goal drives the two defining properties of the tool:
- Resilience is the default, not a mode. A single flaky node (an empty agent response, a rate limit, a spending cap, an unparseable output) must never crash the whole run. The runner retries transient failures, reframes the prompt, and finally defaults a turn's outputs so the machine advances to its next state rather than aborting. See docs/GUARDRAILS.md for the full recovery ladder and its tuning knobs.
- Reproducibility and resume. Every step is recorded as a run artifact and the driver checkpoints after each transition, so a run resumes from exactly where it left off after a crash or reboot.
It is repository-agnostic: repository setup belongs to the workflow or to the container supervisor, while the engine only drives the state machine. A containerized harness for isolated, unattended runs lives in the source repo; see docs/DOCKER.md (not shipped in the PyPI package).
Python workflows
A workflow is an ordinary Python distribution. Constants, loops, and branches use
native Python; values crossing state boundaries are typed; and dependencies live in
the workflow's [project.dependencies]. Workhorse remains one dependency of that
distribution rather than owning its application-specific tools.
State transitions remain statically inspectable, so workhorse-<name> dot can draw
the machine. Code inside a state stays ordinary Python and is tested through the
workflow's node and agent substitution seams.
Install
pip install workhorse-agent # or: uv add workhorse-agent
This installs the library, not a command: workhorse drives no executable of its own. What you run is the console script the workflow's distribution binds — see Running a workflow — so in practice you install that distribution, and workhorse comes along as one of its dependencies.
You also need the agent CLI you intend to
drive on your PATH and authenticated — by default the Claude
CLI (claude), authenticated via a
Claude subscription or claude setup-token. codex, copilot, cline and
opencode are also supported (see
Choosing the agent CLI backend).
Requires Python ≥ 3.12.
Quick start
Install a workflow distribution and run one of the commands it brings. You need the
agent CLI (claude by default) installed and authenticated:
workhorse-hello-world run
workhorse-coder run qa --params '{"story":"CASE-1234","target_env":"dev"}'
Key flags (run workhorse-<name> --help for the full list):
| Flag | Purpose |
|---|---|
--runs-dir <dir> |
Where to write run artifacts (default: <cwd>/.agents/runs) |
--run-id <id> |
Name the stable run dir (<workflow>-<id>); default: a digest of --params, else default |
--cli {claude,codex,copilot,cline,opencode} |
Which agent CLI drives the run (or AGENT_CLI, else the config's default_cli, else claude) |
--profile <name> |
Which named [profiles.<name>] set of models this run uses (see Naming a set of models) |
--config <path> |
Use this config file instead of the discovered one, for this run and everything it spawns |
--params '<json>' / --params-file <path> |
Set the workflow's declared inputs on a fresh start |
--dry-run |
Check the workflow and exit without running a node (see Checking and diagramming a workflow) |
--resume-run <path-or-id> / --resume-latest |
Manually resume a checkpointed run |
Running a workflow (workhorse-<name> run)
Every workflow brings its own command. run is the default subcommand, so it can be
left out:
workhorse-research run # the research workflow's entry flow
workhorse-coder run qa # coder's QA flow standalone
workhorse-coder run qa --params '{"story":"CASE-1234"}'
workhorse-coder qa # same as `run qa`
There is no workhorse executable and no resolution by name. The command is bound
inside the workflow's own distribution, which hands its Registry straight to the CLI:
# myworkflows/research/workflow.py
main = console_script(workflow.entry_point(Research))
[project.scripts]
workhorse-research = "myworkflows.research.workflow:main"
So the workflow that runs is the one whose command you typed — nothing is looked up, and
a name that has no script simply has no command, which you notice at install time rather
than at resolution time. The script must point at what console_script(...) returns
([project.scripts] targets are called after import, so a module-level main = … is the
shape).
The package must be installed unpacked (any pip/uv wheel is): the prompt renderer is a filesystem template loader rooted at the workflow's own directory, so a zip-imported package is refused at startup rather than failing later as a missing template.
The five subcommands each command carries are run, dot,
control, inbox and version — what the operator of a workflow
needs: run it, draw it, steer the live process, read or answer the messages a run left at
its operator gates, and say which engine version did it.
A workflow's node functions run in the workflow distribution's environment, so every
imported tool must be declared in that distribution's [project.dependencies], not
merely installed somewhere else on PATH.
The skill and prompt content those prompts reference is separate, and separately configured — see Initial setup.
The skill and prompt references those prompts make are checked before the first state, and
the ones that will not resolve are printed with the fix. It is a warning, not an error — an
unresolved reference degrades a prompt rather than stopping the run — and --dry-run is
what turns it into an exit code. Which references count as required, and the Jinja
helpers (find_by_tags, instruction_refs, skill_load_ref) a workflow shipping to
unknown repos uses to ask for a skill without demanding it, are in
docs/CHECKING.md.
Running unattended in a container? The source repo ships a Docker harness (image + compose) for fully isolated, week-long runs with credential seeding and persistent volumes. It is not part of the PyPI package — see docs/DOCKER.md.
Checking and diagramming a workflow
--dry-run checks a workflow and exits without running a node — 0 when it is clean, 1
on the first problem, so CI can read it. The failure it exists to catch is a typo found at
hour 30 of an unattended run. dot renders the same workflow to
Graphviz DOT, read off the states rather than a separate diagram,
so it cannot go stale.
workhorse-research run --dry-run
workhorse-coder dot -o wf.dot && dot -Tsvg wf.dot -o wf.svg # render (needs graphviz)
The dry run does a static pass over the states' source — every rendered prompt path exists,
every state is reachable, something can return Done, no transition names a non-state —
and then drives the machine for real over a substituted node index, where every node body
and agent reply is a stand-in. What each half catches, what a fail terminal means with and
without stub_agents({...}), the dry-run run dir, and dot's flags and rendering rules
are in docs/CHECKING.md.
Choosing the agent CLI backend
The controller drives one agent CLI per run, behind a backend port
(workhorse/runner/backends/, one module per CLI). The CLI is chosen per-run; the
model is per-node:
workhorse-<name> run # the configured default_cli, else claude
workhorse-<name> run --cli codex # or copilot, cline, opencode
# Equivalently, set the AGENT_CLI={claude,codex,copilot,cline,opencode} env var.
The unnamed case is configurable, so a machine set up for one CLI does not have to
name it on every run — put default_cli in the shared config (see
BACKENDS.md):
default_cli = "opencode"
| Backend | CLI | Default model | In-place compaction |
|---|---|---|---|
claude |
claude -p (stream-json) |
sonnet |
yes (/compact) |
codex |
codex exec --json |
CLI default | no — ladder reframes on overflow |
copilot |
copilot -p --output-format json |
CLI default | no — ladder reframes on overflow |
cline |
cline --json |
— (node names it) | no — ladder reframes on overflow |
opencode |
opencode run --format json |
— (node names it) | no — ladder reframes on overflow |
JSONL provider error events and logs that identify a transient failure are aborted immediately and retried by workhorse's bounded backoff instead of being left to a CLI's opaque internal retry loop.
A workflow does not name a model. An agent turn asks for an abstract power tier
(high / medium / low) and your user-wide config — one file shared with farrier,
at ~/.config/stablemate/config.toml — maps that tier to a concrete model and effort for
the active backend. Turns with no power=, and tiers with no mapping, fall through to
AGENT_MODEL, then to a per-backend [default.<backend>] table, then to the harness's
own default:
[power.high.claude]
model = "opus"
effort = "high"
Naming a set of models (profiles)
Editing those tables moves every run on the machine, including the six-day one already
going. A [profiles.<name>] table instead holds its own power, default and
default_cli under one name, picked per run with --profile cheap (or a whole other file
with --config ./experiment.toml). A profile replaces the top-level tables rather than
layering over them, and the one a run chose is recorded in its run.json and stamped on its
root span as workhorse.profile — so a flagless --resume-run re-applies the same set, and
a finished run can still say which models it bought.
The full reference — the power and [default.<backend>] tables, per-harness
environment variables, where the config file lives and how its schema version keeps
workhorse and farrier in step, initial setup, codex config profiles, and running
OpenRouter models on cline/opencode (where pinning the upstream endpoint is the
largest cost lever on a long run), and the full [profiles.<name>] reference — is in
docs/BACKENDS.md.
The resilience and timeout knobs are env vars, documented in
docs/GUARDRAILS.md.
Resuming and run identity
The controller is auto-resume-in-place by default. Each (workflow, run-id) pair maps
to one stable run dir (<workflow>-<run-id>), and when you don't pass --run-id the id
defaults to a short digest of --params (e.g. okf-builder-p1c7e4b2a). Re-running the
same params re-derives the same id, so a crash, a reboot or a plain re-run resumes the
existing checkpoint — which is what lets an unattended run survive either. Distinct params
get distinct dirs, so one target never silently resumes another's checkpoint and drops its
own. To start over, delete the run dir; to keep independent runs side by side, pass
distinct --run-ids.
Ctrl-C is recorded rather than silent: the run pauses the way a crash does and stamps
run.json, so a stopped run and a wedged one are not byte-identical on disk.
The resume branches in full, the --resume-run / --resume-latest / --params flag
table, and why a fresh start empties the dir first are in
docs/RUNS.md.
Reaching a run that is already going
A healthy run can be spending real money on a flow you have already fixed on disk. It does not need stopping — it needs the pushed code. A control channel on the run dir carries commands into the live process, answered even from inside a multi-day cap sleep:
workhorse-coder control --run <id> reload # cut the turn, re-enter on the pushed code
workhorse-coder control --run <id> reload --at-boundary # let the turn land first
workhorse-coder control --run <id> reload --core # …and replace workhorse itself
workhorse-coder control --run <id> status # is this process still serving this run dir
workhorse-coder control --run <id> questions # what is this run asking an operator?
workhorse-coder control --run <id> answer --text "go" # answer the gate it is parked on
workhorse-coder control --run <id> switch-cli claude # move it onto another agent CLI
workhorse-coder control --run <id> switch-profile cheap # move it onto another set of models
A reload cuts the turn within about a second, closes its span with the usage it really
accrued (stamped workhorse.cut=reload, so groom does not read it as churn), spends no
recovery budget, and re-enters the checkpoint the state wrote on entry — same process, same
pid, same root span, same run dir, same wall-clock budget. Not a new run. --core is the
one that costs a process image, because workhorse's own modules are on the frame doing the
reload.
Full mechanics — what --at-boundary is for, which packages a reload replaces and which it
deliberately leaves alone, what the spans are stamped with, and what each sibling command
does to a run mid-flight — are in docs/RELOAD.md.
Run artifacts
Each run writes a directory under --runs-dir (default <cwd>/.agents/runs), holding
run.json and the final context.json, the checkpoint.json a resume restarts from and
the events.jsonl log of every state transition, one <step-id>/ per step with the
rendered prompt.md and extracted output.json, a turns/ dir keeping every earlier
visit a looping node overwrote, sessions.jsonl mapping each turn to its agent-CLI
session, and transcripts/ capturing the turns themselves — from a native session store,
OpenCode's full JSON session export, or a redacted stream fallback — because the CLI's own
session data lives on one host and is pruned whenever it likes.
For argv-based CLIs, prompts above 96 KiB are delivered through that persisted
prompt.md instead of being placed in one subprocess argument; stdin-native CLIs keep
their existing transport.
runs/<workflow>-<run-id>/{run.json,launch.json,checkpoint.json,context.json,events.jsonl,sessions.jsonl,turns/,transcripts/,<step-id>/}
launch.json is the one written for whoever outlives the process. A run killed outright —
OOM, a sweep, a hard crash — cannot report its own death, so something outside has to notice
and act, and until this file the directory said everything about the run except how to start
it again. It holds two argvs, and the difference is load-bearing:
resume_argvis the command. Run it from the recordedcwd. Because a resume lands on the stable run dir in place, that is the whole of what a supervisor has to do.argvis forensics — what this process was actually exec'd with — and must never be executed. It can carry--no-cache, which deletes the run directory before starting, and a--params-filethat has since moved on from what the checkpoint holds. Replaying it is not a resume; it is how the run gets lost.
container: true means every path in the record is namespace-local, so a host-side reader
must refuse it rather than re-spawn coordinates that mean something else outside. The
environment is deliberately not recorded: it is read at the process boundary and would put
secrets on disk.
The full tree, what prompt.md does and does not capture, how a transcript capture records
which source it came from, and the WORKHORSE_CAPTURE_TRANSCRIPTS /
WORKHORSE_TRANSCRIPT_MAX_BYTES bounds are in
docs/RUNS.md.
The Docker harness redirects artifacts to a persistent volume instead — see
docs/DOCKER.md.
Telemetry (automatic when a collector is reachable)
For away-from-keyboard monitoring of long runs, workhorse streams OpenTelemetry spans,
metrics and log records to a local OTLP collector — by default
groom, which
stores them in SQLite and pages you (ntfy/webhook + browser) on stall/stuck/churn:
groom serve # now every run is observed
The OTel SDK is a required dependency, so any install of workhorse can export. Telemetry still fails soft when no collector is available, allowing the workflow to continue without monitoring becoming a runtime dependency.
Live gauges distinguish an open agent turn, an explicit operator/cap/retry wait, and ordinary deterministic node work. Turn idle and elapsed values are cleared when the turn closes, so a later operator wait cannot inherit stale evidence that an agent is streaming.
Enablement is auto by default. With WORKHORSE_OTEL unset, start_run opens one
short TCP connection to the endpoint and enables telemetry only if something answers — so
a machine running a collector gets spans with no env var, and one without stays a complete
no-op. WORKHORSE_OTEL=1 forces it on, 0 forces it off. That is the default because the
runs most worth observing are the unattended week-long ones, which are exactly the runs
nobody remembers to export a variable before launching. Auto stays off inside a test
process regardless of the probe — a suite is not a run anyone revisits, and left on it
buries the ones that are.
What is emitted, how turn spans are normalized so cost and tokens are comparable across
harnesses, the cap-wait heartbeat that distinguishes a spending-cap sleep from a hang, and
how to tag spans with your own unit of work are in
docs/TELEMETRY.md.
There is also a wall-clock ceiling, WORKHORSE_MAX_RUNTIME_S — see
docs/GUARDRAILS.md.
Repository isolation
workhorse never assumes a repository layout. A workflow receives repository paths
as declared parameters and performs repository-specific preparation in its Python
nodes. The source checkout's Docker harness prepares a dedicated Git worktree before
launch and passes its location to the workflow, allowing concurrent runs to share the
host repository's object store without sharing a working tree. See
docs/DOCKER.md.
Writing a workflow
A workflow is a Python package: workflow.py holds the Registry and the Workflow
subclasses whose methods are its states, nodes.py holds the @blueprint.node
functions that are its nodes, and prompts/ holds the Jinja2 templates an agent turn
renders. Once a workflow grows more than one machine, each one takes a directory of its
own — dev/flow.py beside the nodes/ and prompts/ only it renders — workflow.py
shrinks to the Registry alone, and prompt paths are written from the package root down
("dev/prompts/implement-plan.md"), which is what Registry(name, package=__package__)
declares. Control flow is ordinary Python — if, for, a counter that is just a counter —
and each state returns the next one:
class Build(Workflow):
subject: str # inputs — filled from --params
def start(self):
reading = self.call(measure, self.subject) # a node
return Continue(None, self.review, count=reading.count)
def review(self, count: int): # state parameters — one hop only
verdict = self.agent("prompts/review.md", returns=Verdict, args={"n": count})
if verdict.ok:
return Done(verdict)
return Continue(None, self.review, count=count + 1)
An agent prompt must output JSON matching the model its turn declared in returns=, and
— because runs go unattended for days — a state must be ready for a reply whose fields
came back empty: after transient retries and reframing, the runner defaults a turn's
declared outputs and lets the machine advance rather than crashing the run.
The authoring reference — the package layout, the worked example end to end, the three
tiers of state and why there is no fourth, where a turn runs (cwd / add_dirs), the
transition table, checkpoints and the aliases= that survive a rename, the node index
that tests substitute through instead of patching, and the labels() that tell a
collector what a run is working on — is in
docs/AUTHORING.md.
Measuring something outside an agent turn
A benchmark, a training run, an evaluation sweep — anything whose value is a number —
does not belong inside an agent turn, whose budget is a budget for thinking and which
kills and re-enters a command that outruns it. workhorse.job submits such a command
detached, under a supervisor that outlives the node, and records what it cost in a file
the command itself cannot write: exit code, peak RSS, wall time, kill reason, and the
containment tier the machine actually delivered. The workflow parks on an Await and a
later state classifies the two artifacts with no model call.
The manifest keys, the three containment tiers, why time is advisory while memory is hard, and how a job is polled, adopted on resume and killed are in docs/JOBS.md.
Development
Working on the controller itself — not on a workflow — starts from a clone of the
stablemate repo rather than a PyPI install.
make help lists the tasks; make test runs the suite, which is dependency-free (each
file in tests/ also runs standalone under uv run python tests/test_x.py).
The project layout, how the driver's loop works, why every agent turn gets a clean session, where to put a test and which of the two styles it is, where docs go, and the container build are in docs/DEVELOPMENT.md. The Docker harness for isolated unattended runs — not shipped in the PyPI package — is in docs/DOCKER.md.
Metadata
Release files for workhorse-agent 3.0.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| workhorse_agent-3.0.0.tar.gz | 355.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| workhorse_agent-3.0.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 694.6 kB
Release files / workhorse_agent-3.0.0.tar.gz
| Download URL | workhorse_agent-3.0.0.tar.gz |
|---|---|
| Size | 355.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
c84872951cdd834d9fbdb3874af790c458cf65957f26df8570614c900d4db905
|
|
BLAKE2b-256 checksum How to use checksums |
9dd0cead921e3db9057c23ef11a20aae7c52d9e83f7283344b27a4a864bb3120
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 11, 2026.
Transparency logRelease files / workhorse_agent-3.0.0-py3-none-any.whl
| Download URL | workhorse_agent-3.0.0-py3-none-any.whl |
|---|---|
| Size | 338.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
b4cbe6f473268b9a78180c659f5fdbdb81bd5bde10d13f966b29fd48e9790c32
|
|
BLAKE2b-256 checksum How to use checksums |
a5ee117347e099e2d4037a1cbc498dc3caf17aad668bef55c55982e984b11e14
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 11, 2026.
Transparency log