Skip to main content

redgear

Your coding agent says it's done. Nobody checked.

That's the failure redgear exists to close. Coding agents are good at writing code and bad at grading their own work. They report success whether or not the tests pass, whether or not they touched files they weren't asked to, whether or not the change does what it claims. Someone has to independently verify each turn before the next one starts, and today that someone is a human reading diffs at 11pm.

redgear is that someone, automated. It's an orchestrator, not an agent: it never writes code itself. It decides what to work on next, writes the prompt that says so, hands it to a coding agent CLI (Claude Code is the reference implementation), and then, before trusting a word of the agent's report, independently re-derives the truth: runs the real test suite itself, re-hashes every file the agent was told not to touch, recomputes the real git diff instead of reading the agent's claimed file list. If the agent says "done" and the evidence disagrees, the run doesn't advance. If the agent says "I'm blocked," that costs nothing: an honest stop should always be cheaper than a lie.

Install

pipx install redgear

Needs Python 3.12+, git, and a Claude Code install. See Requirements below: the Claude Code binary resolution has one gotcha worth reading before your first run.

A worked example

cd my-project
redgear init

Scaffolds .redgear/. The whole audit trail lives here and is committed to your repo like any other file.

redgear plan --from docs/PRD.md

Dispatches a read-only agent turn (Read, Glob, Grep only, it cannot edit a single file) to turn your requirements doc into a task graph: which pieces of work, in what order, with what acceptance criteria. It lands in .redgear/task_graph.json in state draft and redgear run refuses to touch a draft plan. The plan defines what "correct" means for every task that follows. A bad plan produces confidently verified wrong software, and no amount of gate rigor downstream catches that. A human has to look at it first. There is no flag that skips this.

redgear status
┌────────┬───────────────┬─────────┬─────┬────────────┐
│ task   │ type          │ state   │ att │ blocked by │
├────────┼───────────────┼─────────┼─────┼────────────┤
│ T-0001 │ test_authoring│ ready   │ 0/3 │ -          │
│ T-0002 │ implementation│ blocked │ 0/3 │ T-0001     │
└────────┴───────────────┴─────────┴─────┴────────────┘

Read the plan, then approve it explicitly. This records who approved which version of the spec:

redgear approve --by "your name"
redgear run --dry-run

Composes and prints every prompt the loop would send, dispatching nothing. Costs nothing. Use it constantly: it's the fastest way to catch a badly-scoped task before it costs a real agent turn.

redgear run
redgear run
  stop with: redgear stop  (or create .redgear/STOP)

  agent CLI: claude (2.1.229)

complete: 6 iteration(s), 6 verified, 0 escalated

That's the whole interface for a clean run: a banner naming the brake, the resolved agent CLI, and one summary line at the end. Everything else lives in the audit trail, because the point isn't a chatty console. It's a record you can actually check.

What the audit trail shows when something goes wrong

A task that fails a gate doesn't die. It goes back into the queue, and the next prompt for it carries the actual failure excerpt, so the retry is corrective rather than a blind repeat. redgear status shows this as an attempt count climbing against the cap:

│ T-0004 │ implementation│ rejected│ 1/3 │ -          │

If it exhausts its attempts, or the agent honestly reports itself blocked or under-scoped, the run stops there rather than pushing forward on an assumption:

│ T-0007 │ implementation│ escalated│ 2/3 │ -         │

escalated: T-0007 (needs a human)

Reporting "blocked" costs an agent nothing: no attempt is consumed. Claiming completion it can't support does; verification runs independently either way. Every one of these transitions is one line in redgear log, redacted, readable, and reconstructible from .redgear/events.jsonl alone. That file, not the console output, is the actual source of truth.

The seven guarantees

Every design decision in this project traces back to one of these:

  1. It runs the tests itself. No field the agent reports ever decides a verdict. Only real exit codes and a real git diff, recomputed by redgear after the agent's process has already exited.
  2. Tests are frozen during implementation, and code is frozen during test authoring. SHA-256-enforced. An agent that could edit both the tests and the code they check would be grading its own homework.
  3. It's free to say "I'm stuck." An agent whose only options are "pass" or "fail and retry" is structurally pushed toward faking a pass. Reporting blocked or under-scoped costs nothing.
  4. Every verdict has a receipt. .redgear/events.jsonl is an append-only log; every other state file is a projection that can be rebuilt from it byte-for-byte. Nothing is asserted that isn't reconstructible.
  5. redgear holds no API key and makes no outbound network call. All inference is delegated to your own agent CLI subprocess, authenticated with whatever you already configured. redgear never touches a credential, never calls a model API directly, and adds zero egress of its own. All spend belongs to your agent CLI session, not to redgear.
  6. Every run is bounded, and interruptible. Hard caps on iterations, wall-clock time, and consecutive failures; a stop file honored between iterations; a process-tree kill on timeout. It commits verified work to your local repository and touches git in no other way: it never pushes, rebases, resets, or rewrites history. One commit per verified task, each carrying the proof that justifies it. A task that fails gets its working tree restored before the retry; a task that escalates keeps its failure state, untouched, for you to read.
  7. Untrusted text is never treated as an instruction. Test output, diffs, and source documents are explicitly delimited in every prompt as data to diagnose, not commands to follow.

Requirements

  • Python 3.12+, a floor, not a target. Nothing in redgear needs a newer interpreter; this just keeps the requirement honest against what's actually tested.
  • git, with a clean working tree before every run. Without a clean baseline, the diff audit redgear runs is fiction.
  • Claude Code, the reference agent CLI adapter. Other conforming CLIs are architecturally supported but untested.

If Claude Code is installed as the Desktop app (Windows, MSIX-packaged): the claude binary is deliberately not on PATH. redgear run and redgear plan will fail to find it by default. Point redgear at it directly:

redgear run --executable "C:\Users\<you>\AppData\Local\Packages\<PackageFamilyName>\LocalCache\Roaming\Claude\claude-code\<version>\claude.exe"

or persist it once in .redgear/config.json:

{ "runner": { "executable": "C:\\...\\claude.exe" } }

redgear doctor reports whichever one is actually configured, and whether it resolves. Run it first if a run fails with "not installed or not on PATH."

What redgear does to your git repository

Worth knowing before the first run, because two of these will surprise you:

  • It commits each verified task, to your local repository only, using your own git identity — no signature, no co-author trailer. The message carries the task id, the gates that passed, the spec hash, and a path to the proof.
  • It commits with --no-verify, bypassing your pre-commit hooks. A hook that reformats would change the tree after redgear computed the proof, so the commit would carry code no gate ever checked. redgear has already run your configured lint and test commands itself, independently, as part of verification.
  • It restores the working tree after a failed attempt, so the retry starts clean instead of on top of half-finished work. It never touches anything outside the failed turn: the tree is checked clean before every task, .redgear/ is never reverted, and ignored files (your .venv, your build output) are never removed.
  • It leaves an escalated task's tree exactly as the agent left it. That failure state is the evidence you need, so nothing is committed and nothing is discarded. redgear status tells you how to clear it when you're done.
  • It never pushes, rebases, resets, or rewrites history. That stays yours.

The plan

The build plan this project is executing on itself (task graph, spec, architectural contract) lives in .redgear/task_graph.json and .redgear/spec/spec.json. The contract itself is CLAUDE.md; read it before touching any code here.

Development

python -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"

ruff format --check .
ruff check .
mypy
pytest -q

Release files for redgear 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for redgear 0.2.0
File Size Uploaded
redgear-0.2.0.tar.gz 352.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for redgear 0.2.0
File Interpreter ABI Platform
redgear-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 476.1 kB

Release files / redgear-0.2.0.tar.gz

Download URL redgear-0.2.0.tar.gz
Size 352.6 kB
Tags Source
SHA-256 checksum
How to use checksums
b7fb84b506edbe1547eca67cbb381b1739bfe2a40ba2b6148b6326ecfd8ea23f
BLAKE2b-256 checksum
How to use checksums
b936de918472b82c6ee7a33b93ca3082c689faee666a7cb68fb4a2dc2b52ee42
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 24, 2026.

Transparency log

Release files / redgear-0.2.0-py3-none-any.whl

Download URL redgear-0.2.0-py3-none-any.whl
Size 123.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
053490bfb1624cd4633045f519a0f4c9b132412d3d6f3e828771bd70b21ab6f1
BLAKE2b-256 checksum
How to use checksums
5c26087ce7e4c8af5c64f62c0d0011f21928b2a8fa466b6ae359a22b2d131319
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 24, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.2.0 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page