agent-loop
A checkpoint-gated developer/reviewer loop. One coding-agent CLI plays developer and writes the code for the current checkpoint; another plays reviewer and approves or sends it back. The loop commits per checkpoint, refuses to run on main/master, and halts after a configurable number of failed review attempts.
Originally extracted from a Flutter project's run_loop.sh. Project-agnostic: build / test / lint commands are declared per project in plan_checkpoints.json.
Install
The PyPI distribution is agent-loop-tool; the installed command is
agent-loop. Until the first PyPI release, install from the repository
(requires repository access):
pipx install git+https://github.com/dcurioustech/agent-loop.git
# or
uv tool install git+https://github.com/dcurioustech/agent-loop.git
After the first PyPI release, use pipx install agent-loop-tool or
uv tool install agent-loop-tool.
Requires at least one of these CLIs on PATH, depending on the roles you pick:
| Provider | Install |
|---|---|
| claude | https://docs.claude.com/claude-code |
| codex | https://github.com/openai/codex |
| grok | https://docs.x.ai/docs/grok-cli |
| antigravity | https://antigravity.google/docs/cli-overview (binary: agy) |
Usage
In a consumer repo containing plan_checkpoints.json:
agent-loop validate # schema-check the state file
agent-loop status # one-line summary per checkpoint
agent-loop run # developer=claude, reviewer=codex (defaults)
agent-loop run --developer codex --reviewer claude
agent-loop run --developer grok --reviewer antigravity
agent-loop run --developer antigravity --reviewer claude
agent-loop run --developer-model claude-opus-4-8 --reviewer-model gpt-5-codex
Choosing the model
--developer / --reviewer pick the CLI; the model is resolved per role with this
precedence (first match wins):
--developer-model/--reviewer-modelon the command line- a
modelsblock inplan_checkpoints.json(see schema below) - whatever default the CLI itself resolves (its config file / env / built-in)
So if you set nothing, each CLI keeps using its own default model — the loop never
overrides it. Model names are provider-specific, so a value pinned for one role only
makes sense for the CLI you assigned to that role. Under the hood the model is passed as
--model <name> (claude/grok) or -m <name> (codex/antigravity).
Safety envs (per provider, off by default):
ALLOW_AUTO_MODE_CLAUDE=1 # claude --dangerously-skip-permissions
ALLOW_AUTO_MODE_CODEX=1 # codex exec --dangerously-bypass-approvals-and-sandbox
ALLOW_AUTO_MODE_GROK=1 # grok --always-approve
ALLOW_AUTO_MODE_ANTIGRAVITY=1 # agy --dangerously-skip-permissions
The loop refuses to start unless the assigned provider's auto mode gate is set, so unattended runs cannot stall on a permission prompt.
Run logs and the audit trail
Every run writes a log to loop_<YYYYMMDD>_<HHMMSS>.log. Where that lands is resolved in this order:
--log-dir <path>— explicit CLI flag, wins over everything.LOG_DIR=<path>— environment override.<repo-root>/logs— the default, resolved from the git root regardless of your current working directory (falls back to a relativelogs/outside a git repo).
Logs are committed as audit artifacts
Logs under the repository are tracked and committed, not ignored. The loop commits them alongside code at each checkpoint (built, revision, approved) and on every exit path (audit: halted, audit: run failed, audit: completed run), so a run's history survives in git even when it fails.
Two consequences worth knowing:
- Each commit holds a partial log. The log is still being appended to while the loop commits it, so a checkpoint commit captures the log as of that moment. The closing
audit:commit flushes the tail. Bytes written after that land in the next run's first commit — partial by design, never lost. - A modified log does not block the next run. The preflight worktree check ignores changes under the log directory (and only there); real source changes still refuse to start the loop.
Point --log-dir outside the repository and the loop warns that logs will not be committed, then runs normally.
What the log contains
Agent stdout and stderr are streamed into the log as well as to your terminal, so the log can hold the agents' actual working output, not just the loop's own bookkeeping — subject to --audit-level, below. Alongside it the loop emits structured, one-line JSON events that are easy to grep or parse:
| Event | Emitted when |
|---|---|
AGENT_TRACE |
A developer/reviewer invocation starts, its prompt, and its result (action is start, prompt, or finish) |
REVIEW_COMMENT |
The reviewer returns, carrying its notes and resulting status |
APPROVAL_COMMENT |
A checkpoint is approved, or skipped because it already was |
COMMIT_STATEMENT |
A commit is made (with hash) or found unnecessary |
RUN_OUTCOME |
The run ends: completed, halted, or failed |
$ grep RUN_OUTCOME logs/loop_20260816_101500.log
[agent-loop] RUN_OUTCOME {"branch": "feature-x", "event": "RUN_OUTCOME", "outcome": "completed"}
--audit-level: how much of that ends up in git
Because logs are committed, anything an agent prints — or writes into a checkpoint's review_notes — becomes part of permanent git history the moment its commit is made. --audit-level controls how much of that actually reaches the log:
| Level | Raw agent stdout/stderr | Prompts, review notes, comments | Structured events |
|---|---|---|---|
full |
logged as-is | logged as-is | logged |
redacted |
scrubbed for known secret shapes first | scrubbed first | logged |
off (default) |
not logged at all | replaced with a <suppressed: N chars> placeholder |
logged |
Resolution order: --audit-level <level> flag, then AGENT_LOOP_AUDIT_LEVEL env var, then off.
agent-loop run --audit-level full # everything, unfiltered
agent-loop run --audit-level redacted # best-effort secret scrub
agent-loop run # off — structured log events only (default)
At every level, agents still receive the real, unredacted prompt — --audit-level only changes what gets logged, never what an agent is told to do. The state file retains review_notes as written by the agents and is committed with checkpoint changes, regardless of audit level.
redacted is a best-effort net, not a guarantee. It catches known secret shapes — AWS access keys, GitHub/Slack/OpenAI tokens, bearer tokens, PEM private-key blocks, key: value-style assignments — via regex over live, line-streamed subprocess output. It cannot catch a project's own custom secret formats, and a secret split across two flushed writes can slip through. Treat committed logs as something a human should skim before pushing, not as pre-cleared for a public remote. off suppresses raw content in the log, but does not redact the state file.
The --log-dir worktree exemption above only ever tolerates changes to log files; it has no bearing on what those files contain — that's entirely --audit-level's job.
Generating a plan (agent-loop init)
agent-loop init turns a feature description into a schema-valid plan_checkpoints.json
by asking a coding-agent CLI to break it into ordered, independently reviewable
checkpoints. It runs the provider once, non-interactively, captures its output, and
only writes the state file after the result parses as JSON and passes the same
validation load_state applies — nothing is written on a provider error, a timeout,
malformed output, or a schema failure.
Plain-English input:
agent-loop init "Add a login page with email/password auth and a logout button"
Markdown input (e.g. an existing design doc):
agent-loop init --feature-file docs/feature.md
Either form accepts:
| Flag | Default |
|---|---|
--state |
plan_checkpoints.json |
--branch |
sanitized feature/<slug> derived from the feature text (or the --feature-file filename) |
--plan-file |
docs/implementation_plan.md for plain-English input, or the --feature-file path for Markdown input |
--provider |
claude — any provider from agent-loop run's table can generate the plan |
--model |
the provider CLI's own default |
--timeout |
1800 seconds |
--force |
off — refuses to overwrite an existing --state file |
agent-loop init "Add CSV export to the reports page" \
--provider codex --model gpt-5-codex --timeout 600 --branch feature/csv-export
By default init refuses to touch an existing state file so you don't accidentally
clobber checkpoint progress; pass --force to regenerate and overwrite it. A target
branch of main/master is rejected before the provider is ever invoked, same as
agent-loop run. Generated checkpoints always start pending with attempts: 0 and
empty review_notes, regardless of what the provider returned for those fields.
A freshly generated file is immediately usable:
agent-loop init "Add a login page" && agent-loop validate && agent-loop status
Bootstrapping a new consumer repo
Download the example as your starting point and edit branch, plan_file, and the project block to match your stack:
curl -L https://raw.githubusercontent.com/dcurioustech/agent-loop/main/examples/plan_checkpoints.example.json \
-o plan_checkpoints.json
agent-loop validate
The example covers all three checkpoint statuses (pending / built / approved) and shows a per-checkpoint test_cmd override on top of the project-level default.
plan_checkpoints.json schema
{
"plan_file": "docs/implementation_plan.md",
"branch": "feature-branch",
"project": {
"build_cmd": "flutter build web --release",
"test_cmd": "flutter test",
"lint_cmd": "flutter analyze",
"verify_in_review": true
},
"models": { // optional; per-role model, overridden by CLI flags
"developer": "claude-opus-4-8",
"reviewer": "gpt-5-codex"
},
"checkpoints": [
{
"id": "phase0",
"name": "...",
"status": "pending", // pending | built | approved
"scope": "...",
"exit_criteria": ["..."], // non-empty; each entry a non-empty string
"attempts": 0, // non-negative integer
"review_notes": ""
}
]
}
Per-checkpoint build_cmd / test_cmd / lint_cmd overrides are also supported. The reviewer prompt includes these commands so the agent knows how to verify.
Metadata
Release files for agent-loop-tool 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| agent_loop_tool-0.1.0.tar.gz | 53.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| agent_loop_tool-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 85.4 kB
Release files / agent_loop_tool-0.1.0.tar.gz
| Download URL | agent_loop_tool-0.1.0.tar.gz |
|---|---|
| Size | 53.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
c0299a0095fbcc5362cb003c71b75ec112a049429c31b62f1d5cbb4723fdd9dc
|
|
BLAKE2b-256 checksum How to use checksums |
8b909303d8a03f142b2dde1c603a2bae67b7152714cdcf382ebb2322f7872df2
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 27, 2026.
Transparency logRelease files / agent_loop_tool-0.1.0-py3-none-any.whl
| Download URL | agent_loop_tool-0.1.0-py3-none-any.whl |
|---|---|
| Size | 31.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
061f1fb7570079346531b40d84a8bd2d76e99953d1cc1cba4f74d75e7f4ffb52
|
|
BLAKE2b-256 checksum How to use checksums |
3cdc84c5c8a412378d31283f46e3d4c48bdb435cbaabf9ac291f0b6681732e0d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 27, 2026.
Transparency log