Skip to main content

rocket-review

CI License: Apache 2.0 Python 3.13+ PyPI

rr — a second opinion on your code, from a model that didn't write it.

Pre-1.0. Interfaces and flags may change between minor versions. Read Security & data flow before pointing it at anything sensitive.

One small CLI that sends your plan, diff, commit, or PR to an agentic reviewer (Codex CLI, Claude Code, or opencode) that explores your project before judging — then gives you prose or structured JSON you can gate CI on.

Single-model local review is built into the vendor CLIs now (codex review, Claude Code's /code-review) — rr exists for what a single vendor can't give you: a second opinion from a different vendor's model, in one command, from any shell, editor, agent, or CI.

rr is deliberately not a PR bot. It reviews before you push — the point is that the issues get fixed before a PR exists. It posts nothing anywhere; if you want PR comments, pipe the --json envelope into whatever posts them.

What a review looks like

Reviewing an uncommitted diff that quietly breaks a billing invariant (rr --diff, abridged):

[HIGH] billing.py — `split_evenly` no longer preserves the input total
> The documented contract says the returned list "always sums to exactly
> `total_cents`," but the new implementation rounds one share and repeats it.
> For inputs like `split_evenly(100, 3)` it returns `[33, 33, 33]`, summing to 99.
> Suggested fix:
    base = total_cents // n
    parts = [base] * n
    for i in range(total_cents - base * n):
        parts[i] += 1
    return parts

- Change Assessment: Do not merge
- Top Issues: The changed implementation loses cents and breaks the documented
  remainder allocation contract.

The reviewer read the function's docstring and reasoned about its contract before flagging the regression — not a pattern match on the diff.

Why

  • Cross-vendor review — the model that wrote the code shouldn't be the only one grading it: a model reviewing its own output inherits its own blind spots, and a vendor's built-in reviewer is always the same family that wrote the code. Implement with Claude, review with GPT; implement with Codex, review with Claude; or --backend codex,claude fans out to both and shows you where they disagree.
  • Plans are reviewable too — rr plan.md stress-tests a design doc before you build it. Most review tools only understand diffs.
  • Standards-aware — --docs auto-discovers llms.txt, AGENTS.md, or CLAUDE.md (or takes explicit paths) and the reviewer flags deviations from your documented rules, not generic style opinions.
  • CI-gateable — --json --fail-on high exits 2 when a high-severity finding lands. Pipe the envelope to jq or your bot of choice.

Install

pipx install rocket-review
rr init && rr doctor

rr init writes the default config to ~/.config/rocket-review/config.toml, and rr doctor tells you what is still missing — a backend CLI, a login, a model pin.

Or with Homebrew, which brings its own Python:

brew install ledger-rocket/tap/rocket-review

Or the latest from source:

pipx install git+https://github.com/ledger-rocket/rocket-review.git

Requires Python 3.13+ — the Homebrew formula vendors it, so that floor only applies to the pipx and from-source routes. Either way you need at least one backend, which no install method provides:

  • Codex CLI (default for plan reviews)
  • Claude Code (default for code and diff reviews)
  • opencode (--backend opencode — any provider, including local models; experimental, see below)
  • or none of the above — --backend api (shorthand --api) calls the OpenAI API directly, with no agent CLI. It needs three things:
    • OPENAI_API_KEY in your environment
    • the SDK extra: pipx install 'rocket-review[api]', or pipx inject rocket-review openai into an existing install
    • a pipx install — the Homebrew formula ships the base package only

--pr also needs the gh CLI.

Usage

rr plan.md                        # stress-test a plan before building
rr --diff                         # review uncommitted changes
rr --staged                       # review staged changes only
rr --commit abc1234               # review a commit
rr --pr 123                       # review a GitHub PR (number, URL, or branch)
rr --pr 123 --repo acme/api-server  # ...from outside that repo's checkout
git diff HEAD~3 | rr              # pipe anything
rr src/auth.py --docs             # review files against your documented standards
rr --diff --no-config             # ignore the config files (hermetic run)
rr init                           # write the default user config file
rr doctor                         # check this host: config, backends, logins
rr --version                      # print the installed version

rr init and rr doctor

rr init writes the default user config — effort, the per-mode [backends] table, and a [models] pin per backend — to ~/.config/rocket-review/config.toml (or $XDG_CONFIG_HOME/rocket-review/config.toml), creating the directory if it is not there. It never overwrites: run it again and it prints the path and leaves the file alone, --force replaces it. Edit the file afterwards — every key is documented inline and mirrors a flag, exactly as Config file describes.

rr doctor reports what rr would do on this host, one line per fact, each opening with ok, missing, unknown, or failed:

ok       rr 0.4.0 (pipx)
ok       user config: /Users/you/.config/rocket-review/config.toml
missing  project config: no .rocket-review.toml found from /Users/you/src/api (optional)
ok       setting effort = medium (/Users/you/.config/rocket-review/config.toml)
ok       setting backends.diff = claude (built-in default)
ok       backend codex · cli /usr/local/bin/codex · auth ok · model gpt-6-astra (pinned)
missing  backend claude · cli not on PATH — npm install -g @anthropic-ai/claude-code · ...

fixes:
  - install claude: npm install -g @anthropic-ai/claude-code (https://claude.com/claude-code)

It never prompts and never reviews anything: each backend is asked only its own status command (codex login status, claude auth status), with a few seconds' timeout. A CLI that cannot answer — no status command, too old to have one, a timeout — reads unknown, which is never counted as broken.

  • Exit code — 0 when every backend this host depends on is installed and not refusing a login; 1 when one of them is missing or refusing, when the config file is invalid, or when it holds a combination rr refuses to review with (fail_on or full without json, effort with opencode); 2 on a usage error or an internal one. Only the backends your modes actually use are checked; --backend codex,claude checks that list instead — and a mode whose backend is missing still runs, on the substitute rr announces, so a 1 here means the reviewer you configured is not the one you would get.

  • --quiet prints nothing and only sets the exit code, for a pre-push hook that soft-passes on non-zero:

    rr doctor --quiet || echo "rr is not set up here; skipping the review" >&2
    

Default backends by mode

With no --backend, the reviewer follows the review mode:

Mode Default Why
plan codex focused plan reviews, the fastest runs, schema-enforced JSON
code claude deeper findings on source files, with fewer false alarms
diff claude deeper findings on changes, with fewer false alarms

These come from measuring the backends against each other on rocket-review's own eval corpus, not from a preference between vendors — and they are only defaults.

  • Overriding — an explicit --backend always wins, in any mode, single (--backend codex) or as a list (--backend codex,claude). --mode is applied before the backend is chosen, so rr plan.md --mode code reviews as code and uses the code default. To change the table itself for a project or for yourself, set [backends] in a config file.
  • Missing backend — if the mode's default isn't available, rr uses the next available one and says so in a single line on stderr; pass --backend to choose explicitly and silence it. The order is the other agentic CLI first, then opencode (skipped when --effort is set, which it doesn't support), then api — and only when both OPENAI_API_KEY and the SDK extra are present, since a substitute that can't run would defeat the point. The fallback is never silent. If nothing is available, rr errors with the install hint for the backend the mode wanted.
  • Two opinions at once — --backend codex,claude runs both and prints each review under its own heading. Some teams do this on pre-merge changes and read the disagreements; others keep the single per-mode default and spend the time elsewhere. Both are reasonable — it is a choice about your pipeline, not a recommendation.

Pick your reviewer

rr --diff --backend claude                     # Claude reviews (read-only sandbox)
rr --diff --backend opencode:ollama/qwen3      # fully local (opencode → Ollama)
rr --diff --backend codex,claude               # both, side by side
rr --diff --backend codex:gpt-5.6-sol,claude:claude-opus-4-8
rr --diff --effort high                        # more reasoning effort (per-backend flag)

Model names.

  • codex passes no -m, so it honors your own codex default (model in ~/.codex/config.toml). On ChatGPT plans use gpt-5.6-sol, the ChatGPT-account-accessible 5.6 variant — codex signed into a ChatGPT account rejects both the bare gpt-5.6 alias and gpt-5.6-codex.
  • api (API-key auth) defaults to gpt-5.6-terra, balanced on cost and quality. --model gpt-5.6-sol buys max quality at flagship pricing; gpt-5.6-luna is the cheapest tier.
  • rr always names models explicitly, with the suffix, and never relies on the bare gpt-5.6 alias — it points at the flagship today but OpenAI can remap it.

Reasoning effort and timeouts.

  • --effort sets reasoning effort, and the accepted values differ by backend: codex and api take minimal|low|medium|high, claude takes low|medium|high|xhigh|max. An unsupported value fails loudly at the backend rather than being silently ignored.
  • opencode has no effort flag at all, so --effort errors if opencode is among the selected backends.
  • Heavy --effort high reviews — especially on reasoning models — can outrun the default 900s (15 min) subprocess timeout. Raise it with --timeout 1800.

The Codex and Claude backends run agentically in read-only mode: they navigate your project — imports, tests, related files — before judging. That context is what makes the review worth reading. The api backend is the exception — it calls the OpenAI API directly on the supplied content plus any files it references, without navigating your project.

opencode is experimental. The integration works, but end-to-end review reliability depends on the provider you have configured, and non-interactive opencode run can restrict the read-only plan agent's tools. rr materializes the diff and feeds the prompt to opencode on stdin so it always reviews the real change, but for a gated CI check prefer codex or claude. The local-model value prop stands — point opencode at Ollama to keep everything on your machine — just verify its output before trusting it as a gate.

Structured output

rr --diff --json | jq '.findings[] | {severity, title, backend}'
rr --staged --json --fail-on high && git commit   # block the commit on high+ findings
  • Findings each carry severity, title, file, line, why, fix, backend, model.
  • A summary block leads the envelope — findings_total, per-severity counts (with explicit zeros for absent severities), worst_severity, per-backend verdicts, and the gate result when --fail-on is set. An agent gets the counts and the gate answer without parsing the findings array.
  • schema_version (currently the string "1") tags the envelope shape. It bumps only on a breaking change — a field removed, or a type or meaning changed — so new fields can appear without one. Match it exactly against versions you know, treat anything else as unsupported, and ignore keys you don't recognise. No numeric ordering is implied, which is why it is a string.
  • Long output is truncated at 4000 chars: raw keeps the head plus a marker naming the full length. This bounds the envelope and keeps review text — which may quote proprietary code — off disk. --full inlines the untruncated output instead.
  • Failures fail the gate closed, both parse failures and backend errors.

Review modes

  • plan — auto-detected for .md/.txt/.plan files: completeness, ordering, risks, over-engineering.
  • code — source files: correctness, security, performance, maintainability.
  • diff — for --diff/--staged/--commit/--pr/stdin: bugs introduced, missing changes, contract breaks.

Override with --mode, add focus with --prompt "check the locking". The mode also picks the reviewer — see Default backends by mode.

Project standards (--docs)

Point the reviewer at your project's standards docs — it flags deviations from your documented rules:

rr src/auth.py --docs                     # auto-discovers llms.txt / AGENTS.md / CLAUDE.md
rr src/auth.py --docs docs/standards.md docs/smells.md

Relative markdown links inside the docs are followed one level, so an index file (like llms.txt) pulls in everything it references.

When you ask for docs and there are none, that is an error — except as a standing preference. The three ways to ask differ only in that:

  • --docs (bare) — errors if none of llms.txt / AGENTS.md / CLAUDE.md is in the current directory. Pass explicit paths when your standards live elsewhere.
  • --llms [PATH] — a compatibility alias for --docs [PATH], identical in both forms: bare it takes the repository's llms.txt and errors if there isn't one; with a path it reads what you name.
  • docs = true in a config file — a standing preference rather than a request, so it stays silent when a project has no standards doc.

Inside a git repository, an auto-discovered doc and every link followed out of any doc must be a file the repository tracks — the repo, not you, chose those, and a standards doc is copied into the prompt verbatim. Name a path directly (--docs CLAUDE.md) to read one the repository does not carry. The full rule is under What rr will read as a standards doc.

Config file

Two optional TOML files hold the defaults you would otherwise retype:

  • Project — .rocket-review.toml, found by walking up from the current directory to the git root (and no further; outside a repo only the current directory is read).
  • User — ~/.config/rocket-review/config.toml, or $XDG_CONFIG_HOME/rocket-review/config.toml.

Precedence: CLI flag > project file > user file > built-in default, settled per key. A project file that sets only [backends] leaves your user file's timeout in force, and the same holds inside [backends] and [models].

timeout = 1800          # --timeout
effort = "high"         # --effort
fail_on = "high"        # --fail-on (needs json = true, exactly as the flag needs --json)
json = false            # --json
full = false            # --full
docs = true             # --docs with no path (auto-discovery); or a list of paths
codex_sandbox = "read-only"  # --codex-sandbox; user file only, see below

[backends]              # per-mode default backend, overriding the built-in table
plan = "codex"
code = "claude"
diff = "claude"
default = "claude"      # optional: covers any mode you leave out above

[models]                # per backend, i.e. what `--backend name:model` pins
codex = "gpt-5.6-sol"
claude = "claude-opus-5"

Every key mirrors a flag — a config file changes what rr does by default, never what it can do. codex_sandbox is the one key a project file may not set: it is codex exec -s <mode>, read-only by default, and the code under review must not choose how much of the reviewer's machine the reviewer may touch. Set workspace-write or danger-full-access in your user file, or pass --codex-sandbox, only on a host where codex's own sandbox cannot start (a container that may not create user namespaces fails with bwrap: No permissions to create a new namespace) and the surrounding environment is the boundary. What to review (--diff, --pr, files, --mode, --prompt) stays on the command line, where it is visible in the invocation.

Anything else is an error, named rather than ignored: an unknown key or mode, a backend rr doesn't have, a non-integer timeout, a fail_on that isn't a severity, malformed TOML (reported with the file and the parse position). Config is validated before any git, gh, or backend work starts.

--no-config ignores both files. It is also the way out of a json = true, full = true, or docs you don't want on this run: rr has no --no-json-style inverse flags, so a lower layer can only be turned off by a higher config file (json = false in the project file) or by --no-config. Flags always win where they can say anything at all.

In a gated CI job, pass --no-config (or at least an explicit --backend). The config a job picks up comes from the branch under test, which on a fork's PR is the contributor's: [backends]/[models] decide which model reviews the code, so a config change alone can downgrade the reviewer your gate depends on. --fail-on is the one direction that is safe either way — a file can only tighten a gate the job did not set, never loosen the one it did, since your flag outranks it.

docs in a config file

docs = true means "use this project's standards doc if it has one" — llms.txt / AGENTS.md / CLAUDE.md, and a project without one is not an error the way a typed --docs is. Which project depends on which file asked:

  • Project file — looks in the directory holding the .rocket-review.toml, so everyone gets the same standards wherever in the repo they run from.
  • User file — looks in the checkout you are in (its git root, or the current directory outside a repo), never beside ~/.config/rocket-review/config.toml, which is nobody's project.

Explicit docs paths are relative to the config file that names them.

What rr will read as a standards doc

One rule decides every docs path, whether a config named it, discovery found it, or a markdown link inside another doc points at it:

A doc the repository chose is read only if the repository tracks it (at HEAD), it resolves inside the directory it came from, and it is not inside .git. A doc you name is read as-is.

"The repository chose it" covers three cases:

  • a path in a project's .rocket-review.toml — that file is repository content, and on a fork's PR it is a contributor's;
  • anything auto-discovered (docs = true, or --docs / --llms with no path) — the repo decides which file answers the pattern, whoever asked for the pattern;
  • every markdown link followed out of a doc that resolves inside a repository — the link was written by whoever wrote that doc. This holds even for a doc you named yourself: naming --docs STANDARDS.md vouches for that file, not for what it points at.

Everything is decided on the resolved path, so a tracked symlink is judged by the file it actually opens — the repository carries the link, not its target. Untracked and ignored files are yours, not the project's: a .env, a key, a private note. A named path that is refused stops the run and says so; a discovered or linked one is skipped with a warning, since discovery is a standing "if there is one".

Outside a git repository nothing is tracked, so what remains is the directory confinement and the .git guard: a config's docs must stay inside the config's own directory, and a link inside the doc's.

What you name is yours — --docs PATH, --llms PATH, or a path in your own ~/.config/rocket-review/config.toml — and is read wherever it points, including files no repository tracks. The single exception is .git, which rr never reads into a review, as a footgun check.

--docs and --llms sit on the same boundary in both forms, which is what makes them the alias this README calls them: bare, each takes what the repository offers; with a path, each reads what you named.

How it works

rr assembles a review prompt — a mode-specific rubric, your standards docs, and the plan or diff under review — and hands it to an agentic CLI (Codex, Claude Code, or opencode) running read-only inside your project. Because the reviewer runs in your checkout, it can open imports, tests, and related files to understand context before it judges, rather than reasoning from the diff alone. It then returns either prose or a parsed findings envelope (--json). There is no rocket-review server in the loop: the only thing that leaves your machine is the review request — the diff or plan, your standards docs, and any files the reviewer opens — sent to whichever backend and provider you chose. Point rr at a local opencode/Ollama model to keep everything on your machine.

Security & data flow

rr runs the reviewer in a read-only sandbox (no writes to your files), but read-only is not the same as safe:

  • Your code leaves your machine. Each review sends the diff/plan, your standards docs, and any files the reviewer opens to that backend's provider — codex/api → OpenAI, claude → Anthropic, opencode → whichever provider you configured (point it at a local Ollama model to keep everything on your machine).
  • Read-only stops writes, not reads. An agent can still read any secret your shell can (.env, ~/.aws, tokens) and send it upstream.
  • Untrusted input can prompt-inject the reviewer — a hostile PR body, diff, comment, or AGENTS.md can try to steer an agentic backend. Be especially careful with --pr on a dev machine.
  • A repo's .rocket-review.toml configures the runs you make inside it — including which backend, and therefore which provider, gets your code, and which model at what cost. The docs it can name are bounded (see docs in a config file) and it can only select backends you have installed, but it is repo-supplied input: read it like any other tooling config a clone brings with it, or use --no-config.

Don't run agentic backends against untrusted repos or PRs on a machine where readable secrets exist. See SECURITY.md for the full threat model and how to report a vulnerability.

Requirements

  • Python ≥ 3.13
  • OS — macOS or Linux
  • A backend CLI, installed and authenticated — you only need the one(s) you use:
    • codex — Codex CLI, signed in with your ChatGPT/OpenAI account
    • claude — Claude Code, on a Claude subscription or API key. Needs a version supporting --permission-mode manual (Claude Code 2.1.x+); older CLIs fail the review closed with a usage error. Check with claude --help | grep -A3 permission-mode.
    • opencode — opencode, configured for any provider (including a local Ollama model)
    • api — no CLI, but needs the OpenAI SDK (pipx install 'rocket-review[api]', or pipx inject rocket-review openai); set OPENAI_API_KEY and rr calls the OpenAI API directly
  • gh CLI, authenticated, for --pr

No install method brings a backend: each one above is a separate install and its own login. rr doctor says which of them this host has and which it does not.

Agent integration

Drop into your CLAUDE.md / AGENTS.md:

Before pushing non-trivial changes, run `rr --diff --docs` and address the findings.
For plans, run `rr plan.md --docs` before implementing. Use a 900000ms timeout.

Notes

  • Every backend runs in a read-only sandbox on your project — no writes: Codex runs with -s read-only, Claude Code with a read-only tool allowlist under --permission-mode manual, and opencode with its built-in read-only plan agent (edit/write denied at the tool level). Read-only stops writes; it does not stop the agent reading readable secrets and sending them to the backend's provider — see Security & data flow.
  • Every agentic prompt carries the backend's own description of what its sandbox allows, and asks the reviewer to say which checks it ran and which it could not run. A verdict from a Claude review, whose sandbox denies tests and linters, says so instead of implying they passed.
  • --fail-on requires --json — including when a config file is what set it; the error names the file.
  • Exit codes: 0 no gate tripped · 1 operational error (or every backend failed) · 2 findings at/above --fail-on. A partial backend failure warns on stderr but still exits 0 — gate CI with --json --fail-on to fail closed.

Contributing

Issues and PRs are welcome. To set up a dev environment:

python3 -m venv .venv
.venv/bin/pip install -e ".[dev]"
.venv/bin/pytest -q                  # run the tests
.venv/bin/ruff check .               # lint
.venv/bin/mypy rocket_review/        # type-check
.venv/bin/yamllint .                 # yaml lint

CI gates all four plus a package build — run them before opening a PR.

License

Apache-2.0 licensed. See LICENSE.

Release files for rocket-review 0.4.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for rocket-review 0.4.1
File Size Uploaded
rocket_review-0.4.1.tar.gz 134.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for rocket-review 0.4.1
File Interpreter ABI Platform
rocket_review-0.4.1-py3-none-any.whl Python 3 none any Details

Total release size: 208.9 kB

Release files / rocket_review-0.4.1.tar.gz

Download URL rocket_review-0.4.1.tar.gz
Size 134.4 kB
Tags Source
SHA-256 checksum
How to use checksums
61e1b415feaf426735dbc9069ef87a8bacab4c46903ca9433a56628fa6d8d9c0
BLAKE2b-256 checksum
How to use checksums
f733c921fe9f3caf95f06f5ed681d5ba15995bed6b88daf7044d9b036462edc0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.13

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 10, 2026.

Transparency log

Release files / rocket_review-0.4.1-py3-none-any.whl

Download URL rocket_review-0.4.1-py3-none-any.whl
Size 74.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
6abbb16ff9dc248415f77e7f57b0086b3d4b775aa8ae231ae9b920b112a70cb0
BLAKE2b-256 checksum
How to use checksums
0dc3e55cf6b4e96a2d1fe429983e73f121d9e5088bf81b7132172029911c5a2d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.13

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 10, 2026.

Transparency log

Release history Release notifications | RSS feed

0.5.0

2 release files

This release

0.4.1 This release

2 release files

0.4.0

2 release files

0.3.2

2 release files

0.3.1

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page