Run a fleet of AI coding agents in parallel, in isolated git worktrees, with per-run cost you can trust — measured where the provider reports it, marked unavailable where it does not, and never a fake $0.
Proof
One task, four routing strategies, measured cost and latency from the ledger. Not a guess.
| Strategy | Backend | Model | Status | Cost | Source | Duration | Tokens (in / out) |
|---|---|---|---|---|---|---|---|
| deepseek | opencode | opencode-go/deepseek-v4-flash | exited_clean | $0.0029 | native | 81.8 s | 11,740 / 1,977 |
| claude | claude-code | claude-sonnet-4-6 | exited_clean | $0.3374 | native | 121.4 s | 17 / 6,837 |
| cmdcode | command-code | zai-org/GLM-5.2 | exited_clean | unavailable |
unavailable | 252.6 s | 0 / 0 |
| codex-glm | codex | z-ai/glm-5.1 (via EastRouter) | exited_clean | unavailable |
unavailable | 283.0 s | 231,075 / 7,812 |
cheapest: deepseek ($0.0029) · fastest: deepseek (81.8 s)
Each produced solution's tests: deepseek, claude, and cmdcode passed 6/6. deepseek was cheapest, fastest, and correct, for ~1/115th of claude's cost. Full methodology and tables: examples/benchmark-output.md and docs/nerds.md.
Install
You need Python 3.11+, uv, git, and at least one backend CLI
installed and logged in. Marshal drives agents; it does not ship one. marshal doctor tells you
which are ready. (The Claude Code plugin launches the MCP server through uv, so it is required on
that path too.)
Claude Code plugin. The fastest path: Skills and the MCP server in one step.
/plugin marketplace add chiruu12/marshal
/plugin install marshal@marshal
From PyPI, for the CLI and the Python library:
uv tool install "MarshalFleet[mcp]"
# or
pipx install "MarshalFleet[mcp]"
The [mcp] extra is what makes marshal mcp work. Without it the server exits before a host can
connect.
To track unreleased work on main:
uv tool install "marshalfleet[mcp] @ git+https://github.com/chiruu12/marshal"
Backend CLI auth and MCP wiring: SETUP.md.
60 seconds
1. Configure a fleet. From your project repo, scaffold a starter fleet.config.yaml:
marshal init # scaffolds fleet.config.yaml in the current repo
The scaffold ships every client commented out, so uncomment at least one and save. Using the
codex example:
clients:
codex:
backend: codex
model: gpt-5.6-luna
2. Check the fleet is ready. doctor verifies auth, not just that a CLI is on $PATH. A backend you cannot actually run fails here rather than 3 seconds into a job.
$ marshal doctor
✓ repo: /path/to/your-repo (branch main)
✓ config: fleet.config.yaml (1 client)
✓ backend:codex: available
✓ plan:codex: logged-in
3. Dispatch a job. It returns immediately with a run id. The agent works in its own git worktree on its own branch. Your checkout is never touched.
$ marshal spawn --client codex --goal "Add a docstring to hello()"
hello-docstring.codex.d56489fe codex/gpt-5.6-luna running (poll: marshal status)
4. Watch the fleet. Any number of agents, any mix of providers, each with its own cost line.
$ marshal status
hello-docstring.codex.d56489fe codex exited_clean unavailable ~/.marshal/worktrees/myrepo-a1b2c3d4e5f6/hello-docstring.codex.d56489fe
5. Review, then merge. From a driver agent over MCP: collect_run("<run_id>") returns the diff read-only, and integrate("<run_id>", message="...") merges it. exited_clean means the process exited cleanly. It does not mean the code is correct, so the diff review is not optional.
6. Record what you didn't merge. integrate records integrated for you. Nothing records the rest — so a run you reviewed and threw away looks identical to one nobody read.
marshal outcome <run_id> rejected --note "wrong approach"
marshal routing # which client's work actually got kept, per task kind
This is what turns the ledger into routing evidence. routing rates are computed over judged
runs only, so skipping this step makes every client read 100% — flattering and useless. Tag work at
spawn (--task-kind refactor) and the answer comes back per kind of task, with n_judged beside
every rate and null rather than a guess when nothing has been judged.
MCP tool reference: docs/mcp-tools.md. Orientation for drivers: call marshal_quickstart() first.
Demo
A real run, start to finish: doctor verifies the fleet, spawn dispatches, status reports, and
the diff lands in the agent's own worktree. The waiting is cut; nothing else is.
This one ran on OpenCode's free deepseek-v4-flash-free, which bills nothing and so reports no
per-run cost — Marshal prints unavailable rather than inventing a number, while still recording
the tokens. On a paid backend that reports cost, that column carries the real figure; see
Proof. Recorded with VHS from
assets/demo.tape.
Why Marshal
- One base class, many backends. Backend choice is a per-call parameter, never global.
- Parallel by default. Each agent runs in its own git worktree until you explicitly integrate.
- Per-provider usage tracking. Every run's cost is tagged by provenance; unknown cost is
unavailable, never a fake $0. - Routing from evidence, not vibes. Record which runs you kept, and
routingranks clients by what actually survived review — per kind of task, with the sample size attached, and never ranked on a cost it could not measure. - Robust headless execution: hard timeouts, process-group kill, and no prompting modes that deadlock without stdin.
How it compares
| Worktree isolation | Real parallelism | Per-provider cost accounting | Human merge review | Multi-repo | |
|---|---|---|---|---|---|
| Marshal | yes (one worktree per run) | yes (capped run_many) |
yes (native / admin-api / unavailable) | yes (collect_run then integrate) |
yes (one MCP server, many workspaces) |
| Hand-rolled git worktree scripts | yes (if you build it) | possible (you manage threads/processes) | no (unless you wire it yourself) | possible (you own the merge step) | no (one repo per script unless you extend it) |
| Subagents inside a single coding harness | partial (often same checkout or sandbox) | limited (shared process/context) | partial (whatever the host exposes) | varies (some hosts auto-apply) | no (bound to one session/repo) |
| Generic workflow orchestrators (CI, Airflow, Temporal) | no (not their job) | yes (at the workflow layer) | no (not agent-token aware) | yes (human gates are the point) | yes (but not coding-agent native) |
What you can drive it with
Model is set per client in fleet.config.yaml (or ad-hoc via --model / MCP model=). marshal backends lists what this build ships and marshal doctor says which are actually usable on your machine.
| Backend | Model flag | Cost provenance |
|---|---|---|
cursor |
optional (defaults to CLI default) | unavailable (individual plans expose no per-run cost) |
opencode |
optional (defaults to opencode-go/glm-5.2, the Go subscription) |
native (tokens + cost from the CLI) |
codex |
provider/model (e.g. gpt-5.6-luna) |
admin-api via EastRouter usage_api; else unavailable |
claude-code |
claude-* (e.g. claude-sonnet-4-6) |
native (total_cost_usd + tokens) |
command-code |
provider/model (e.g. zai-org/GLM-5.2) |
unavailable (hosted account, no token/cost in stdout) |
goose |
provider/model or bare model (e.g. cursor-agent/auto) |
native when the provider reports positive cost; else unavailable (best-effort stream-json) |
antigravity (experimental) |
gemini-*, claude-*, etc. |
unavailable (text-only output) |
Routing playbook: docs/model-playbook.md. Verification matrix: docs/status.md.
Architecture
Marshal is the infrastructure layer between a driver agent and a fleet of headless coding CLIs. The driver plans; Marshal creates worktrees, runs backends with timeouts, records usage to an immutable ledger, and returns diffs for review. A future end-user product (Chauffeur) will sit on top — see docs/chauffeur-future.md.
Full design: docs/design.md.
Documentation
SETUP.md: clone-to-first-run setup.docs/usage.md: configure a fleet; drive it via MCP, CLI, or library.docs/mcp-tools.md: MCP tool reference.docs/config.md: everyfleet.config.yamlkey.docs/model-playbook.md: which client to route a task to.docs/status.md: what's built and verified.docs/design.md: architecture and backend cheat sheets.docs/nerds.md: numbers, methodology, and the stuff we argue about.docs/sources.md: primary sources.examples/: runnable copy-paste examples (library_quickstart, pipelinedthenreview,read_paths, adversarial teams, multi-workspace, per-clientenv:); seeexamples/README.md.
Contributing
CONTRIBUTING.md: dev setup, quality gate, adding a backend.SECURITY.md: security model and private disclosure.CHANGELOG.md·CODE_OF_CONDUCT.md
License
MIT
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file marshalfleet-0.2.3.tar.gz.
File metadata
- Download URL: marshalfleet-0.2.3.tar.gz
- Upload date:
- Size: 2.1 MB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
3556adf56867ed9783625374d12046f8dad34fe322eeeb57f73b710cce134a93
|
|
| MD5 |
c7d551c5d4e2606cefa795a57426a6ab
|
|
| BLAKE2b-256 |
832aeeac8f6d826db576c1600908dc0ee04a9b60a47b02e6f359526f0af42a23
|
Provenance
The following attestation bundles were made for marshalfleet-0.2.3.tar.gz:
Publisher:
release.yml on chiruu12/marshal
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
marshalfleet-0.2.3.tar.gz -
Subject digest:
3556adf56867ed9783625374d12046f8dad34fe322eeeb57f73b710cce134a93 - Sigstore transparency entry: 2473570859
- Sigstore integration time:
-
Permalink:
chiruu12/marshal@5dcbd6f70fd807d800f7aa042f1183f1893efef7 -
Branch / Tag:
refs/tags/v0.2.3 - Owner: https://github.com/chiruu12
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@5dcbd6f70fd807d800f7aa042f1183f1893efef7 -
Trigger Event:
release
-
Statement type:
File details
Details for the file marshalfleet-0.2.3-py3-none-any.whl.
File metadata
- Download URL: marshalfleet-0.2.3-py3-none-any.whl
- Upload date:
- Size: 313.9 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
8fb02637d12935c050001e3aca2e357057ecd80feb9403b3cb913bb768898f48
|
|
| MD5 |
e38e35f0726d76d7293cc23f66688f0c
|
|
| BLAKE2b-256 |
c213d4e2f4dce42c5192c021ea4b275df80379d41bfacd216fa8f6e172fb148d
|
Provenance
The following attestation bundles were made for marshalfleet-0.2.3-py3-none-any.whl:
Publisher:
release.yml on chiruu12/marshal
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
marshalfleet-0.2.3-py3-none-any.whl -
Subject digest:
8fb02637d12935c050001e3aca2e357057ecd80feb9403b3cb913bb768898f48 - Sigstore transparency entry: 2473570875
- Sigstore integration time:
-
Permalink:
chiruu12/marshal@5dcbd6f70fd807d800f7aa042f1183f1893efef7 -
Branch / Tag:
refs/tags/v0.2.3 - Owner: https://github.com/chiruu12
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@5dcbd6f70fd807d800f7aa042f1183f1893efef7 -
Trigger Event:
release
-
Statement type: