Skip to main content

workhorse

PyPI

A fail-soft runner for agent workflows written as Python state machines — drives an agent CLI (Claude, Codex, Copilot, Cline or OpenCode) unattended for days.

A workflow is a Python package: its states are methods that return the next state, its nodes are plain functions. workhorse drives the machine, renders Jinja2 prompts, invokes the agent CLI, validates JSON replies into typed models, checkpoints after every transition, and writes run artifacts.

The PyPI distribution is workhorse-agent and the import package is workhorse. Workflow distributions provide the commands users run.

Why

workhorse exists to run long, multi-step agent workflows unattended — the design target is a single run that survives for a week without a human babysitting it. That goal drives the two defining properties of the tool:

  • Resilience is the default, not a mode. A single flaky node (an empty agent response, a rate limit, a spending cap, an unparseable output) must never crash the whole run. The runner retries transient failures, reframes the prompt, and finally defaults a turn's outputs so the machine advances to its next state rather than aborting. See docs/GUARDRAILS.md for the full recovery ladder and its tuning knobs.
  • Reproducibility and resume. Every step is recorded as a run artifact and the driver checkpoints after each transition, so a run resumes from exactly where it left off after a crash or reboot.

It is repository-agnostic: repository setup belongs to the workflow or to the container supervisor, while the engine only drives the state machine. A containerized harness for isolated, unattended runs lives in the source repo; see docs/DOCKER.md (not shipped in the PyPI package).

Python workflows

A workflow is an ordinary Python distribution. Constants, loops, and branches use native Python; values crossing state boundaries are typed; and dependencies live in the workflow's [project.dependencies]. Workhorse remains one dependency of that distribution rather than owning its application-specific tools.

State transitions remain statically inspectable, so workhorse-<name> dot can draw the machine. Code inside a state stays ordinary Python and is tested through the workflow's node and agent substitution seams.

Install

pip install workhorse-agent     # or: uv add workhorse-agent

This installs the library, not a command: workhorse drives no executable of its own. What you run is the console script the workflow's distribution binds — see Running a workflow — so in practice you install that distribution, and workhorse comes along as one of its dependencies.

You also need the agent CLI you intend to drive on your PATH and authenticated — by default the Claude CLI (claude), authenticated via a Claude subscription or claude setup-token. codex, copilot, cline and opencode are also supported (see Choosing the agent CLI backend).

Requires Python ≥ 3.12.

Quick start

Install a workflow distribution and run one of the commands it brings. You need the agent CLI (claude by default) installed and authenticated:

workhorse-hello-world run
workhorse-coder run qa --params '{"story":"CASE-1234","target_env":"dev"}'

Key flags (run workhorse-<name> --help for the full list):

Flag Purpose
--runs-dir <dir> Where to write run artifacts (default: <cwd>/.agents/runs)
--run-id <id> Name the stable run dir (<workflow>-<id>); default: a digest of --params, else default
--cli {claude,codex,copilot,cline,opencode} Which agent CLI drives the run (or AGENT_CLI, else the config's default_cli, else claude)
--profile <name> Which named [profiles.<name>] set of models this run uses (see Naming a set of models)
--config <path> Use this config file instead of the discovered one, for this run and everything it spawns
--params '<json>' / --params-file <path> Set the workflow's declared inputs on a fresh start
--dry-run Check the workflow and exit without running a node (see Checking and diagramming a workflow)
--resume-run <path-or-id> / --resume-latest Manually resume a checkpointed run

Running a workflow (workhorse-<name> run)

Every workflow brings its own command. run is the default subcommand, so it can be left out:

workhorse-research run                              # the research workflow's entry flow
workhorse-coder run qa                              # coder's QA flow standalone
workhorse-coder run qa --params '{"story":"CASE-1234"}'
workhorse-coder qa                                  # same as `run qa`

There is no workhorse executable and no resolution by name. The command is bound inside the workflow's own distribution, which hands its Registry straight to the CLI:

# myworkflows/research/workflow.py
main = console_script(workflow.entry_point(Research))
[project.scripts]
workhorse-research = "myworkflows.research.workflow:main"

So the workflow that runs is the one whose command you typed — nothing is looked up, and a name that has no script simply has no command, which you notice at install time rather than at resolution time. The script must point at what console_script(...) returns ([project.scripts] targets are called after import, so a module-level main = … is the shape).

The package must be installed unpacked (any pip/uv wheel is): the prompt renderer is a filesystem template loader rooted at the workflow's own directory, so a zip-imported package is refused at startup rather than failing later as a missing template.

The five subcommands each command carries are run, dot, control, inbox and version — what the operator of a workflow needs: run it, draw it, steer the live process, read or answer the messages a run left at its operator gates, and say which engine version did it.

A workflow's node functions run in the workflow distribution's environment, so every imported tool must be declared in that distribution's [project.dependencies], not merely installed somewhere else on PATH.

The skill and prompt content those prompts reference is separate, and separately configured — see Initial setup.

The skill and prompt references those prompts make are checked before the first state, and the ones that will not resolve are printed with the fix. It is a warning, not an error — an unresolved reference degrades a prompt rather than stopping the run — and --dry-run is what turns it into an exit code. Which references count as required, and the Jinja helpers (find_by_tags, instruction_refs, skill_load_ref) a workflow shipping to unknown repos uses to ask for a skill without demanding it, are in docs/CHECKING.md.

Running unattended in a container? The source repo ships a Docker harness (image + compose) for fully isolated, week-long runs with credential seeding and persistent volumes. It is not part of the PyPI package — see docs/DOCKER.md.

Checking and diagramming a workflow

--dry-run checks a workflow and exits without running a node — 0 when it is clean, 1 on the first problem, so CI can read it. The failure it exists to catch is a typo found at hour 30 of an unattended run. dot renders the same workflow to Graphviz DOT, read off the states rather than a separate diagram, so it cannot go stale.

workhorse-research run --dry-run
workhorse-coder dot -o wf.dot && dot -Tsvg wf.dot -o wf.svg   # render (needs graphviz)

The dry run does a static pass over the states' source — every rendered prompt path exists, every state is reachable, something can return Done, no transition names a non-state — and then drives the machine for real over a substituted node index, where every node body and agent reply is a stand-in. What each half catches, what a fail terminal means with and without stub_agents({...}), the dry-run run dir, and dot's flags and rendering rules are in docs/CHECKING.md.

Choosing the agent CLI backend

The controller drives one agent CLI per run, behind a backend port (workhorse/runner/backends/, one module per CLI). The CLI is chosen per-run; the model is per-node:

workhorse-<name> run                      # the configured default_cli, else claude
workhorse-<name> run --cli codex          # or copilot, cline, opencode
# Equivalently, set the AGENT_CLI={claude,codex,copilot,cline,opencode} env var.

The unnamed case is configurable, so a machine set up for one CLI does not have to name it on every run — put default_cli in the shared config (see BACKENDS.md):

default_cli = "opencode"
Backend CLI Default model In-place compaction
claude claude -p (stream-json) sonnet yes (/compact)
codex codex exec --json CLI default no — ladder reframes on overflow
copilot copilot -p --output-format json CLI default no — ladder reframes on overflow
cline cline --json — (node names it) no — ladder reframes on overflow
opencode opencode run --format json — (node names it) no — ladder reframes on overflow

JSONL provider error events and logs that identify a transient failure are aborted immediately and retried by workhorse's bounded backoff instead of being left to a CLI's opaque internal retry loop.

A workflow does not name a model. An agent turn asks for an abstract power tier (high / medium / low) and your user-wide config — one file shared with farrier, at ~/.config/stablemate/config.toml — maps that tier to a concrete model and effort for the active backend. Turns with no power=, and tiers with no mapping, fall through to AGENT_MODEL, then to a per-backend [default.<backend>] table, then to the harness's own default:

[power.high.claude]
model = "opus"
effort = "high"

Naming a set of models (profiles)

Editing those tables moves every run on the machine, including the six-day one already going. A [profiles.<name>] table instead holds its own power, default and default_cli under one name, picked per run with --profile cheap (or a whole other file with --config ./experiment.toml). A profile replaces the top-level tables rather than layering over them, and the one a run chose is recorded in its run.json and stamped on its root span as workhorse.profile — so a flagless --resume-run re-applies the same set, and a finished run can still say which models it bought.

The full reference — the power and [default.<backend>] tables, per-harness environment variables, where the config file lives and how its schema version keeps workhorse and farrier in step, initial setup, codex config profiles, and running OpenRouter models on cline/opencode (where pinning the upstream endpoint is the largest cost lever on a long run), and the full [profiles.<name>] reference — is in docs/BACKENDS.md. The resilience and timeout knobs are env vars, documented in docs/GUARDRAILS.md.

Resuming and run identity

The controller is auto-resume-in-place by default. Each (workflow, run-id) pair maps to one stable run dir (<workflow>-<run-id>), and when you don't pass --run-id the id defaults to a short digest of --params (e.g. okf-builder-p1c7e4b2a). Re-running the same params re-derives the same id, so a crash, a reboot or a plain re-run resumes the existing checkpoint — which is what lets an unattended run survive either. Distinct params get distinct dirs, so one target never silently resumes another's checkpoint and drops its own. To start over, delete the run dir; to keep independent runs side by side, pass distinct --run-ids.

Ctrl-C is recorded rather than silent: the run pauses the way a crash does and stamps run.json, so a stopped run and a wedged one are not byte-identical on disk.

The resume branches in full, the --resume-run / --resume-latest / --params flag table, and why a fresh start empties the dir first are in docs/RUNS.md.

Reaching a run that is already going

A healthy run can be spending real money on a flow you have already fixed on disk. It does not need stopping — it needs the pushed code. A control channel on the run dir carries commands into the live process, answered even from inside a multi-day cap sleep:

workhorse-coder control --run <id> reload                 # cut the turn, re-enter on the pushed code
workhorse-coder control --run <id> reload --at-boundary   # let the turn land first
workhorse-coder control --run <id> reload --core          # …and replace workhorse itself
workhorse-coder control --run <id> status                 # is this process still serving this run dir
workhorse-coder control --run <id> questions              # what is this run asking an operator?
workhorse-coder control --run <id> answer --text "go"     # answer the gate it is parked on
workhorse-coder control --run <id> switch-cli claude      # move it onto another agent CLI
workhorse-coder control --run <id> switch-profile cheap   # move it onto another set of models

A reload cuts the turn within about a second, closes its span with the usage it really accrued (stamped workhorse.cut=reload, so groom does not read it as churn), spends no recovery budget, and re-enters the checkpoint the state wrote on entry — same process, same pid, same root span, same run dir, same wall-clock budget. Not a new run. --core is the one that costs a process image, because workhorse's own modules are on the frame doing the reload.

Full mechanics — what --at-boundary is for, which packages a reload replaces and which it deliberately leaves alone, what the spans are stamped with, and what each sibling command does to a run mid-flight — are in docs/RELOAD.md.

Run artifacts

Each run writes a directory under --runs-dir (default <cwd>/.agents/runs), holding run.json and the final context.json, the checkpoint.json a resume restarts from and the events.jsonl log of every state transition, one <step-id>/ per step with the rendered prompt.md and extracted output.json, a turns/ dir keeping every earlier visit a looping node overwrote, sessions.jsonl mapping each turn to its agent-CLI session, and transcripts/ capturing the turns themselves — from a native session store, OpenCode's full JSON session export, or a redacted stream fallback — because the CLI's own session data lives on one host and is pruned whenever it likes. For argv-based CLIs, prompts above 96 KiB are delivered through that persisted prompt.md instead of being placed in one subprocess argument; stdin-native CLIs keep their existing transport.

runs/<workflow>-<run-id>/{run.json,launch.json,checkpoint.json,context.json,events.jsonl,sessions.jsonl,turns/,transcripts/,<step-id>/}

launch.json is the one written for whoever outlives the process. A run killed outright — OOM, a sweep, a hard crash — cannot report its own death, so something outside has to notice and act, and until this file the directory said everything about the run except how to start it again. It holds two argvs, and the difference is load-bearing:

  • resume_argv is the command. Run it from the recorded cwd. Because a resume lands on the stable run dir in place, that is the whole of what a supervisor has to do.
  • argv is forensics — what this process was actually exec'd with — and must never be executed. It can carry --no-cache, which deletes the run directory before starting, and a --params-file that has since moved on from what the checkpoint holds. Replaying it is not a resume; it is how the run gets lost.

container: true means every path in the record is namespace-local, so a host-side reader must refuse it rather than re-spawn coordinates that mean something else outside. The environment is deliberately not recorded: it is read at the process boundary and would put secrets on disk.

The full tree, what prompt.md does and does not capture, how a transcript capture records which source it came from, and the WORKHORSE_CAPTURE_TRANSCRIPTS / WORKHORSE_TRANSCRIPT_MAX_BYTES bounds are in docs/RUNS.md. The Docker harness redirects artifacts to a persistent volume instead — see docs/DOCKER.md.

Telemetry (automatic when a collector is reachable)

For away-from-keyboard monitoring of long runs, workhorse streams OpenTelemetry spans, metrics and log records to a local OTLP collector — by default groom, which stores them in SQLite and pages you (ntfy/webhook + browser) on stall/stuck/churn:

groom serve                                                # now every run is observed

The OTel SDK is a required dependency, so any install of workhorse can export. Telemetry still fails soft when no collector is available, allowing the workflow to continue without monitoring becoming a runtime dependency.

Live gauges distinguish an open agent turn, an explicit operator/cap/retry wait, and ordinary deterministic node work. Turn idle and elapsed values are cleared when the turn closes, so a later operator wait cannot inherit stale evidence that an agent is streaming.

Enablement is auto by default. With WORKHORSE_OTEL unset, start_run opens one short TCP connection to the endpoint and enables telemetry only if something answers — so a machine running a collector gets spans with no env var, and one without stays a complete no-op. WORKHORSE_OTEL=1 forces it on, 0 forces it off. That is the default because the runs most worth observing are the unattended week-long ones, which are exactly the runs nobody remembers to export a variable before launching. Auto stays off inside a test process regardless of the probe — a suite is not a run anyone revisits, and left on it buries the ones that are.

What is emitted, how turn spans are normalized so cost and tokens are comparable across harnesses, the cap-wait heartbeat that distinguishes a spending-cap sleep from a hang, and how to tag spans with your own unit of work are in docs/TELEMETRY.md. There is also a wall-clock ceiling, WORKHORSE_MAX_RUNTIME_S — see docs/GUARDRAILS.md.

Repository isolation

workhorse never assumes a repository layout. A workflow receives repository paths as declared parameters and performs repository-specific preparation in its Python nodes. The source checkout's Docker harness prepares a dedicated Git worktree before launch and passes its location to the workflow, allowing concurrent runs to share the host repository's object store without sharing a working tree. See docs/DOCKER.md.

Writing a workflow

A workflow is a Python package: workflow.py holds the Registry and the Workflow subclasses whose methods are its states, nodes.py holds the @blueprint.node functions that are its nodes, and prompts/ holds the Jinja2 templates an agent turn renders. Once a workflow grows more than one machine, each one takes a directory of its own — dev/flow.py beside the nodes/ and prompts/ only it renders — workflow.py shrinks to the Registry alone, and prompt paths are written from the package root down ("dev/prompts/implement-plan.md"), which is what Registry(name, package=__package__) declares. Control flow is ordinary Python — if, for, a counter that is just a counter — and each state returns the next one:

class Build(Workflow):
    subject: str                                   # inputs — filled from --params

    def start(self):
        reading = self.call(measure, self.subject)          # a node
        return Continue(None, self.review, count=reading.count)

    def review(self, count: int):                  # state parameters — one hop only
        verdict = self.agent("prompts/review.md", returns=Verdict, args={"n": count})
        if verdict.ok:
            return Done(verdict)
        return Continue(None, self.review, count=count + 1)

An agent prompt must output JSON matching the model its turn declared in returns=, and — because runs go unattended for days — a state must be ready for a reply whose fields came back empty: after transient retries and reframing, the runner defaults a turn's declared outputs and lets the machine advance rather than crashing the run.

The authoring reference — the package layout, the worked example end to end, the three tiers of state and why there is no fourth, where a turn runs (cwd / add_dirs), the transition table, checkpoints and the aliases= that survive a rename, the node index that tests substitute through instead of patching, and the labels() that tell a collector what a run is working on — is in docs/AUTHORING.md.

Measuring something outside an agent turn

A benchmark, a training run, an evaluation sweep — anything whose value is a number — does not belong inside an agent turn, whose budget is a budget for thinking and which kills and re-enters a command that outruns it. workhorse.job submits such a command detached, under a supervisor that outlives the node, and records what it cost in a file the command itself cannot write: exit code, peak RSS, wall time, kill reason, and the containment tier the machine actually delivered. The workflow parks on an Await and a later state classifies the two artifacts with no model call.

The manifest keys, the three containment tiers, why time is advisory while memory is hard, and how a job is polled, adopted on resume and killed are in docs/JOBS.md.

Development

Working on the controller itself — not on a workflow — starts from a clone of the stablemate repo rather than a PyPI install. make help lists the tasks; make test runs the suite, which is dependency-free (each file in tests/ also runs standalone under uv run python tests/test_x.py).

The project layout, how the driver's loop works, why every agent turn gets a clean session, where to put a test and which of the two styles it is, where docs go, and the container build are in docs/DEVELOPMENT.md. The Docker harness for isolated unattended runs — not shipped in the PyPI package — is in docs/DOCKER.md.

Metadata

Release files for workhorse-agent 3.0.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for workhorse-agent 3.0.0
File Size Uploaded
workhorse_agent-3.0.0.tar.gz 355.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for workhorse-agent 3.0.0
File Interpreter ABI Platform
workhorse_agent-3.0.0-py3-none-any.whl Python 3 none any Details

Total release size: 694.6 kB

Release files / workhorse_agent-3.0.0.tar.gz

Download URL workhorse_agent-3.0.0.tar.gz
Size 355.8 kB
Tags Source
SHA-256 checksum
How to use checksums
c84872951cdd834d9fbdb3874af790c458cf65957f26df8570614c900d4db905
BLAKE2b-256 checksum
How to use checksums
9dd0cead921e3db9057c23ef11a20aae7c52d9e83f7283344b27a4a864bb3120
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 11, 2026.

Transparency log

Release files / workhorse_agent-3.0.0-py3-none-any.whl

Download URL workhorse_agent-3.0.0-py3-none-any.whl
Size 338.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
b4cbe6f473268b9a78180c659f5fdbdb81bd5bde10d13f966b29fd48e9790c32
BLAKE2b-256 checksum
How to use checksums
a5ee117347e099e2d4037a1cbc498dc3caf17aad668bef55c55982e984b11e14
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 11, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

3.0.0 This release

2 release files

2.1.0

2 release files

2.0.0

2 release files

1.0.0

2 release files

0.8.0

2 release files

0.7.5

2 release files

0.7.4

2 release files

0.7.3

2 release files

0.7.2

2 release files

0.7.1

2 release files

0.7.0

2 release files

0.6.0

2 release files

0.5.0

2 release files

0.4.2

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.1

2 release files

0.3.0

2 release files

0.2.5

2 release files

0.2.4

2 release files

0.2.3

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page