code-forge
A 5-step code review pipeline for AI coding assistants. Treats review as a state machine: three independent passes per cycle, three consecutive clean cycles required, any finding resets the counter. The minimum path to a commit is 9 static review passes plus a runtime smoke test.
Why
AI coding assistants ship code that compiles, runs, and looks right. Single-pass review (Copilot, Cursor, CodeRabbit, etc.) catches the obvious defects but misses two failure modes:
- Author and reviewer collapse. When the same model writes and reviews the change, it inherits its own blind spots. code-forge runs three independent review perspectives (qodo, expert, adversarial) and treats their findings as untrusted claims that must be reproduced before any fix.
- Self-claimed completion. Hooks that gate on "I finished" markers are
bypassable by any agent that can write a string. code-forge gates on
actual state: a real
pre-commithook running the test suite, a mutation runner proving the tests catch regressions, and a coverage heuristic detecting drift across components.
Quick start
pip install code-review-forge
code-forge install-skill
For MCP server support (IDE integration):
pip install code-review-forge[mcp]
The first command installs the CLI (Python >=3.12). The second copies the
6 review skills into ~/.claude/skills/. Then in Claude Code, run the
full pipeline:
/code-forge
Or invoke individual passes:
/qodo-review # change-aware pre-review (Pass 1)
/code-review-expert # SOLID, architecture, security (Pass 2)
/adversarial-qe # red-team QE, 12 attack dimensions (Pass 3)
/kernel-fp-verify # false-positive verification (Step 3.5)
/smoke-test # runtime verification (Step 4)
Other agent targets:
code-forge install-skill --target vscode # <cwd>/.claude/skills/
code-forge install-skill --target universal # <cwd>/.agents/skills/
code-forge install-skill --dest /path/to/dir # explicit location
code-forge install-skill --skill code-forge # one skill only
code-forge install-skill --force # overwrite existing
Getting Started
A complete walkthrough from install to your first gated commit.
1. Install and initialize
pip install code-review-forge
cd your-repo
code-forge init
init creates .code-forge/gate.yaml with commented examples. It does not
configure a backend or a test runner -- both are your responsibility.
2. Configure a backend
Open .code-forge/gate.yaml and uncomment one backend block. For the Anthropic
API:
backends:
claude-api:
type: api
format: anthropic
base_url: https://api.anthropic.com
api_key_env: ANTHROPIC_API_KEY
default: true
api_key_env is the name of an environment variable, not the key itself.
Set the variable in your shell:
export ANTHROPIC_API_KEY=sk-ant-...
Then trust the configuration (required once per gate.yaml change):
code-forge trust
Review refuses to run until at least one backend is configured and trusted.
3. Run your first review
code-forge requires a git worktree for review isolation. Create one and make a change:
git worktree add .worktrees/work -b my-feature
cd .worktrees/work
# edit some code, then stage it
git add -A
code-forge review
The review reads the staged (index) diff -- you do not need to commit first.
Simple single-branch repos: if worktrees are overkill for your workflow, bypass the check:
code-forge review --allow-main
# or permanently:
FORGE_ALLOW_MAIN=1 code-forge review
4. Install the commit gate (optional but recommended)
The commit gate runs code-forge verify (receipt tamper check) and
code-forge gate-check (test suite) on every code commit. It requires a
test section in gate.yaml:
test:
command: [pytest, -q]
timeout_seconds: 900
Then install:
code-forge install-hooks
What this means for your workflow:
- Code commits (
.py,.go,.rs, etc.) require a passing review first. The hook runscode-forge verifywhich checks that a full review has been recorded for the staged diff. Without a prior review, the commit is blocked. - Doc/config commits (
.md,.yaml,.toml,LICENSE, etc.) skip the gate automatically -- no review or--no-verifyneeded.
Bootstrapping a new repo: if you install hooks before running your first
review, code commits will be blocked. Either run a successful code-forge review
first, or set FORGE_ALLOW_NO_BACKEND=1 temporarily to bypass the receipt gate
while you get your backend working.
5. Commit with the gate active
# 1. Stage your changes
git add -A
# 2. Run a review (must pass before committing code)
code-forge review
# 3. Commit -- the pre-commit hook verifies the review receipt
git commit -m "your message"
If the review found issues, fix them and re-run code-forge review until it
passes. The commit gate checks that the staged diff has a clean review receipt.
6. Diagnostics
code-forge doctor # check backend reachability, config health
code-forge verify # check review receipt status
Backend configuration
Backends can also be set once for all projects in
~/.config/code-forge/config.yaml (override the location with
FORGE_CONFIG_DIR); a backend with the same name in a project's
gate.yaml wins. code-forge doctor prints the resolved user-level
config path every run.
By default, code-forge uses the claude CLI in your PATH with the session
model (no model pin). Three environment variables control the backend:
| Variable | Purpose | Default |
|---|---|---|
FORGE_BACKEND |
Select a named backend from gate.yaml |
session-default |
FORGE_OUTLET |
Force outlet: subprocess | inline | subagent |
auto-detected |
FORGE_LLM_MODEL |
Override model for CLI backends | claude-sonnet-4-6 |
Quick examples:
# Use the default (claude CLI, session model)
code-forge review
# Pin a specific model for this run
FORGE_LLM_MODEL=claude-opus-4-5 code-forge review
# Use a named API backend from gate.yaml
FORGE_BACKEND=claude-api code-forge review
# Force inline outlet (no subprocess)
FORGE_OUTLET=inline code-forge review
Named backends (optional) are defined in the backends: key of
.code-forge/gate.yaml (created by code-forge init):
backends:
claude-api:
type: api
format: anthropic
base_url: https://api.anthropic.com
api_key_env: ANTHROPIC_API_KEY
default: true
openai-compatible:
type: api
format: openai
base_url: https://api.openai.com/v1
api_key_env: OPENAI_API_KEY
local-claude:
type: cli
model: claude-opus-4-5
command: claude
Full reference: docs/configuration.md
Editor setup guides:
- Claude Code: docs/setup-claude-code.md
- VS Code: docs/setup-vscode.md
- Cursor: docs/setup-cursor.md
- PyCharm: docs/setup-pycharm.md
MCP server (IDE integration)
code-forge-mcp is a local stdio MCP server that exposes forge as tools
callable from any MCP-capable editor (Claude Code, VS Code Copilot, Cursor,
PyCharm AI Assistant). Reviews route to the configured CN backend -- the
calling model never reviews its own code.
| Tool | Purpose |
|---|---|
forge_review |
Review the current git diff (inline if fast, job_id if slow) |
forge_gate_check |
Pre-commit gate on staged changes |
forge_resolve_outlet |
Show which backend forge will use (read-only) |
forge_job_status |
Poll a long-running review by job_id |
forge_init |
Create .code-forge/ in the workspace |
forge_trust |
Trust the gate.yaml backends |
Prerequisite: a configured backend with its API key in the server
environment. Without it, forge_review fails closed (same as the CLI).
One server instance serves one project -- for multi-project setups, use
per-project FORGE_PROJECT_DIR entries (see docs/setup-mcp.md).
Claude Code:
claude mcp add forge -- code-forge-mcp
Launch claude from the repo root so the server finds .code-forge/gate.yaml.
VS Code (1.102+, .vscode/mcp.json):
{
"servers": {
"forge": {
"type": "stdio",
"command": "code-forge-mcp",
"cwd": "${workspaceFolder}"
}
}
}
Gotcha: GUI editors do not inherit your shell environment. Either wrap
code-forge-mcp in a script that exports the API key, or set env in the
server config. See docs/setup-mcp.md for a pass-based
wrapper example and the per-editor MCP configs.
Verify: call forge_resolve_outlet -- it should name a backend, not
"key not set". Then call forge_review on a real diff.
Troubleshooting: stale server processes
If old code-forge-mcp processes accumulate (visible as high memory or
multiple PIDs), clean them up manually:
pgrep -af code-forge-mcp # list survivors
pkill -TERM -f code-forge-mcp # graceful shutdown
sleep 3
pkill -KILL -f code-forge-mcp # force-kill any that remain
After a server restart, job IDs from the previous instance become invalid.
Completed reviews leave receipts under .code-forge/ regardless.
The pipeline
Code Change
|
v
[Step 0] Syntax (0a) + Lint (0b) + Non-ASCII (0c)
|
v
[Cycle 1] Pass 1: qodo-review
Pass 2: code-review-expert
Pass 3: adversarial-qe
|
| zero findings -> counter += 1
| any finding -> fix, counter = 0, restart Cycle 1
v
[Cycle 2] (same 3 passes)
|
v
[Cycle 3] (same 3 passes)
| counter = 3
v
[Step 3.5] kernel-fp-verify (if fixes were applied during cycles)
|
v
[Step 4] smoke-test (runtime verification)
|
v
[COMMIT GATE] # post-review-c3
What ships
| Skill | Step | Purpose |
|---|---|---|
| code-forge | Orchestrator | Runs the full 5-step pipeline |
| qodo-review | Pass 1 | Change-aware pre-review with feature-grouped walkthrough |
| code-review-expert | Pass 2 | SOLID, architecture, security analysis |
| adversarial-qe | Pass 3 | Red-team QE with 12 attack dimensions |
| kernel-fp-verify | Step 3.5 | 10-step false-positive verification protocol |
| smoke-test | Step 4 | Runtime verification with bash assertion primitives |
What code-forge does that others don't
- Multi-pass convergence. Three consecutive clean cycles from three independent perspectives. Any finding resets the counter to zero. Copilot, CodeRabbit, Cursor, and Devin are single-pass.
- Anti-hallucination gates. code-forge treats LLM review output as untrusted claims. Parser-deterministic findings auto-confirm; LLM findings require falsification before disposition; Step 4 runs the actual code. Prompt-only mitigations cap at 15% hallucination reduction; tool grounding reaches 65-80% (CodeAnt and Suprmind data, 2026).
- Real commit gate (R1). A real
.git/hooks/pre-committhat runs the test suite and blocks on NEW failures vs a baseline. Gates on diff content and test results, not a self-claimed marker. Closes the terminal-and-IDE bypass that PreToolUse hooks cannot reach. - Mutation-gated review (R2). Diff-scoped mutation runs after static review and before the verdict. Each mutant introduced into the changed code is run against the test suite; a surviving mutant flags tests that cannot catch the change. Toothless tests block the same cycle that finds the defect.
- Cross-component coverage heuristic (R3). Detects diffs that span multiple source areas with a changed function signature. An opt-in components mapping raises an uncertain finding when a hub and a dependent both change in the same diff and no integration test under the dependent's paths matches the configured test patterns.
- Execution before the verdict. The reviewed diff is run, not just read. A declared environment is verified against its lockfile by sha256; the run happens in a disposable directory so the reviewed tree stays read-only; a timeout kills the whole descendant process group. The evidence is asymmetric on purpose -- a test failing before the fix is a verdict input, a test passing after it is recorded but cannot confirm a finding. When the environment cannot be grounded, the report says so instead of reasoning against a version the build never used.
- Large diffs do not silently review as clean. Past a token budget a diff is split along def-use lines and each group reviewed separately, with the shared prompt prefix placed ahead of the per-pass role sentence so the backend can cache it. Measured on a 14-file diff that previously lost all three passes to truncation and returned zero findings: 15/15 passes completed, 16-22 findings per cycle, billed input per pass down from 65,748 to 21,987 tokens. Under-budget diffs keep the byte-identical single-pass path. A reply truncated inside its own reasoning is salvaged only where the verdict is already complete, never into a clean round.
Honest limitations
- No cross-repo impact. code-forge reviews a single repository.
Multi-repo dependency analysis requires CodeRabbit-style tooling or
Chromium's
Cq-Depend. - No feedback learning. code-forge does not adapt to dismissed findings or developer preferences. Each review is independent.
- No long-term maintainability scoring. code-forge does not assess technical debt accumulation. SonarQube's tech-debt tracking is the closest automated approximation.
- No performance regression suite. No benchmark harness equivalent to
Rust's
perf.rust-lang.org. - R3 is artifact-presence, not coverage proof. The cross-component check confirms an integration test file exists under the expected path; it does not verify that the test exercises the specific code that changed. A present-but-stale test passes the gate.
Static review (3-cycle convergence) is one layer. code-forge learned from its own Phase 2 experience where 9 static passes and 639 mock tests missed 3 bugs that dynamic verification caught. Verification grounding (test suite + mutation + e2e coverage check) is the thesis -- not a passes count.
Requirements
- Python 3.12 or newer
jqfor the bash smoke primitives- Claude Code or a compatible AI coding assistant for skill invocation
mcpPython package (optional, forcode-forge-mcp):pip install code-review-forge[mcp]
Installation alternatives
git clone
git clone https://github.com/HouMinXi/forge.git
cd forge
./install.sh
Symlinks each of the 6 skills from ~/.claude/skills/<name> to this
repo's skills/<name>. Hook installation is manual -- see
hooks/README.md and hooks/settings-snippet.json.
Enabling the commit gate (R1)
install-skill and ./install.sh install the review skills only -- they
do not set up enforcement. The R1 pre-commit gate -- the un-fakeable layer that
runs the test suite on every commit and blocks on new failures, regardless of
what the in-editor review claims -- is a separate, manual step:
For running this gate in CI rather than as a local hook, see docs/setup-ci.md.
-
Add a
test:section to.code-forge/gate.yaml. Without it,gate-checkexits withgate.yaml must have a 'test' section:test: command: [pytest, -q] timeout_seconds: 900
command[0]must be a known runner (python3,python,pytest,cargo,go,make,npm,npx,node); no shell metacharacters are allowed. -
Install the hook:
code-forge install-hooksThis writes
.git/hooks/pre-committhat runscode-forge verify(a receipt tamper check) and thencode-forge gate-check(the test gate). -
If
git config core.hooksPathis set,install-hooksrefuses to write to a custom hooks path and prints a manual fallback. Add these two lines to your existing pre-commit hook by hand:code-forge verify --quiet 2>/dev/null || exit 1 exec code-forge gate-check
The skills give you the review passes; this gate is what makes a green verdict
mean the tests actually pass. Without it, an in-editor review that never ran can
still reach a commit. Commits that stage only non-code files (docs, config,
metadata such as .md, .yaml, .toml, LICENSE, README) are detected by
the hook and skip the gate automatically -- no receipts and no --no-verify
needed. Any staged file outside that set, including unknown extensions, re-arms
the gate for the whole commit.
Hooks (reference implementations)
| Hook | Trigger | Purpose |
|---|---|---|
check_worktree.sh |
PreToolUse Edit/Write | Block edits in main worktree |
check_non_ascii.sh |
PreToolUse Write/Edit | Non-ASCII character detection |
check_read_before_edit.sh |
PreToolUse Edit | 1:1 read-before-edit ratio |
check_review_tracker.sh |
PostToolUse Bash | Review cycle state machine |
check_git_commit_review.sh |
PreToolUse Bash | Block unreviewed commits |
check_git_push_review.sh |
PreToolUse Bash | Block unreviewed pushes |
Some hooks contain environment-specific logic (Kerberos auth, pattern
matching) you will need to adapt. See hooks/README.md.
Bash smoke primitives
skills/smoke-test/test-library/shell/ ships 19 reusable bash assertion
functions with no dependencies beyond jq:
run_and_capture,run_concurrent,concurrent_waitassert_success,assert_failure,assert_exit_codeassert_output_contains,assert_output_not_containsassert_stderr_contains,assert_stderr_emptyassert_file_exists,assert_file_not_exists,assert_file_containsassert_json_validassert_no_zombie,assert_temp_cleanassert_no_command_exec,assert_no_command_exec_json,assert_no_path_traversal
A backward-compatible symlink at test-library/ points to
skills/smoke-test/test-library/ for users migrating from
bash-smoke-primitives.
Documentation
evidence/cross-model-complementarity.md-- why 3 different review passesevidence/design-iterations.md-- how the pipeline evolvedevidence/ground-truth-verification.md-- why smoke tests must inject bugsevidence/shell-assertion-footguns.md-- 5 bash-specific trapsevidence/v9-model-coverage-matrix.md-- 4-model coverage datahooks/README.md-- hook installation and adaptation guide
Contributing
Issues and discussion: https://github.com/HouMinXi/forge/issues.
License
Apache-2.0
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file code_review_forge-2.9.0.tar.gz.
File metadata
- Download URL: code_review_forge-2.9.0.tar.gz
- Upload date:
- Size: 968.8 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/6.2.0 CPython/3.14.6
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
a6845732fc102a013203752b0d3fa35783524aa4061eae9d992d82b1770d115c
|
|
| MD5 |
9bbb6013a0ffa40827cd073056ba93f6
|
|
| BLAKE2b-256 |
a5bfde41284ed8da56561895cd397348e2d7b25c38b6e1569307efa880d49535
|
File details
Details for the file code_review_forge-2.9.0-py3-none-any.whl.
File metadata
- Download URL: code_review_forge-2.9.0-py3-none-any.whl
- Upload date:
- Size: 516.3 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/6.2.0 CPython/3.14.6
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
e8883e124f61e4096ffb1de14880533cc790b88fc489fe5db43971b2a5a8e724
|
|
| MD5 |
588ee882ec9b85611018356956334e60
|
|
| BLAKE2b-256 |
96f90062ca855286bec995594ab9615d6aba8c2f23349fe3a5801f178be5cbc4
|