Agentic Pull Request Reviewer
A LangGraph workflow that reviews a Git diff, optionally runs deterministic checks (pytest / ruff), proposes structured findings with an LLM, critiques those findings with a verifier, and prints a Markdown PR review. It fails closed: model or verifier failures are reported as incomplete, never as a clean bill of health.
Install
Isolated CLI install (recommended — does not pollute a project's environment):
pipx install agentic-pr-reviewer
# or
uv tool install agentic-pr-reviewer
Or into an environment:
pip install agentic-pr-reviewer
In security-sensitive workflows, pin the version (e.g. agentic-pr-reviewer==0.2.0).
Set one provider API key (in a .env in the directory you run from, or exported in your shell). Run the command from your repository root so its .env is picked up. See Providers for non-OpenAI options.
Usage
# Review the current repo against the previous commit
agentic-pr-reviewer --base HEAD~1
# Review a specific repo / base
agentic-pr-reviewer --repo ./my-project --base main
# Review a saved unified diff and save the report
agentic-pr-reviewer --diff changes.diff --output review.md
# Pick a provider/model and allow more reviewer<->verifier retries
agentic-pr-reviewer --provider anthropic --model claude-3-5-sonnet-latest --max-retries 3
--repo and --diff are mutually exclusive; with neither, the current directory is used. Deterministic checks are opt-in via --run-checks (see Security).
Providers
The reviewer auto-detects the provider from whichever API key is set, so you only have to export a key:
| Provider | Install | API key | Default model |
|---|---|---|---|
| OpenAI | (bundled) | OPENAI_API_KEY |
gpt-4o-mini |
| Anthropic | pip install "agentic-pr-reviewer[anthropic]" |
ANTHROPIC_API_KEY |
claude-3-5-sonnet-latest |
| Google Gemini | pip install "agentic-pr-reviewer[google]" |
GOOGLE_API_KEY |
gemini-1.5-flash |
| Groq | pip install "agentic-pr-reviewer[groq]" |
GROQ_API_KEY |
llama-3.3-70b-versatile |
| Mistral | pip install "agentic-pr-reviewer[mistral]" |
MISTRAL_API_KEY |
mistral-small-latest |
Install every provider with pip install "agentic-pr-reviewer[all]".
- Auto-detection: if exactly one key is set, it is used. If several are set, choose with
--provider(orPR_REVIEWER_PROVIDER). If none are set, the command exits with a clear error. - Model: override the default with
--model(orPR_REVIEWER_MODEL;OPENAI_MODELstill works for OpenAI).
Retries
The reviewer critiques its own findings and can retry with feedback. --max-retries controls the budget (default 1):
--max-retries 0— single pass, no retry.--max-retries N— up to N retries.--max-retries unlimited— retry until the verifier is satisfied, still bounded by an internal safety cap (LangGraph requires a finite recursion limit).
Why LangGraph
LangGraph models the review as an explicit state machine:
- State is the shared clipboard (diff, findings, status, retry count, report).
- Nodes are single-purpose functions (load, check, review, verify, report).
- Edges (including conditional ones) decide the next step — including skipping verification when the reviewer fails and a bounded reviewer -> verifier retry loop.
That makes the agentic loop inspectable and testable, instead of burying control flow inside one prompt.
Architecture
flowchart TD
START([START]) --> loadDiff[load_diff]
loadDiff --> identifyFiles[identify_files]
identifyFiles --> runChecks[run_checks]
runChecks --> reviewCode[review_code]
reviewCode --> routeReview{route_after_review}
routeReview -->|reviewer_failed| generateReport[generate_report]
routeReview -->|ok| verifyFindings[verify_findings]
verifyFindings --> route{route_after_verification}
route -->|retry max 1| prepareRetry[prepare_retry]
prepareRetry --> reviewCode
route -->|finish| generateReport
generateReport --> END([END])
| Node | Role |
|---|---|
load_diff |
Validate / load the patch; reject empty or oversized diffs (input_error) |
identify_files |
Parse changed file paths from the unified diff |
run_checks |
Run ruff and pytest, only when --run-checks is set |
review_code |
LLM review focused on real defects; drops findings outside the diff; fails closed on model errors |
verify_findings |
LLM critique; rejects unsupported findings; never promotes unverified candidates |
prepare_retry |
Increment retry_count (budget: 1) |
generate_report |
Emit status-aware Markdown + stats |
Review status
Every run ends with a status that the report and CLI exit code reflect:
| Status | Meaning | Exit code |
|---|---|---|
success |
Completed; every surviving finding was verified | 0 |
partial |
Completed with warnings (malformed/out-of-diff findings dropped, or checks unavailable) | 0 |
input_error |
Empty, oversized, or unparseable diff | 2 |
reviewer_failed |
Reviewer model errored or returned only malformed output | 1 |
verifier_failed |
Verifier errored; candidates surfaced as unverified, never confirmed | 1 |
A partial or failed run is never reported as "No verified defects found".
Security model
- Secret-free CI (.github/workflows/ci.yml): runs ruff + pytest on PR code. It holds no secrets, so executing untrusted PR code is safe.
- Trusted review (.github/workflows/pr-review.yml): uses
pull_request_target(which has secrets), so it never checks out or executes PR head code. It installs the trusted reviewer from the base branch and reads the PR diff as data via the GitHub API, then posts a sticky comment. - Local checks are opt-in:
--run-checksruns the target repo's pytest/ruff, which executes that repo's code. It is off by default and prints a warning; only enable it for repositories you trust. - Prompt-injection: the diff, comments, filenames, and tool output are treated as untrusted data in both system prompts; instructions embedded in them are ignored.
Privacy
The diff (and, with --run-checks, test/lint output) is sent to the configured model provider (OpenAI by default). Do not run it on diffs you cannot share with that provider. Secrets and environment variables are never sent to the model.
GitHub integration
| Workflow | Trigger | Output | Secrets |
|---|---|---|---|
ci.yml |
PR + push to main | check status | none |
pr-review.yml |
pull_request_target (opened/synchronize/reopened) |
sticky PR comment | OPENAI_API_KEY |
commit-review.yml |
push to main |
job summary + artifact | OPENAI_API_KEY |
publish.yml |
tag v* |
PyPI release (Trusted Publishing) | none (OIDC) |
One-time setup: add OPENAI_API_KEY under repo Settings > Secrets and variables > Actions. Optionally set an OPENAI_MODEL variable (default gpt-4o-mini).
Caveat: GitHub does not expose secrets to pull requests from forks, so those PRs post a "skipped" note.
Deterministic and trajectory tests
Automated tests (LLM calls mocked) cover input validation, routing, fail-closed statuses, out-of-diff/path-normalized filtering, and CLI exit codes:
pip install -e ".[checks]"
pytest -q
These verify control flow and failure behavior — not model accuracy. A labeled evaluation of real review quality is future work; no accuracy percentages are claimed until measured.
Publishing (maintainers)
Releases go out via publish.yml on a v* tag using PyPI Trusted Publishing (OIDC, no stored token).
One-time PyPI setup (cannot be automated from CI): at pypi.org, create/verify the account, then Publishing -> add a pending Trusted Publisher with:
| Field | Value |
|---|---|
| PyPI Project Name | agentic-pr-reviewer |
| Owner | BitBucket0 |
| Repository | agentic-pr-reviewer |
| Workflow | publish.yml |
| Environment | pypi |
Then cut a release:
git tag v0.2.0 && git push origin v0.2.0 # triggers publish.yml
Local rehearsal before tagging:
rm -rf build dist && python -m build && python -m twine check dist/*
Limitations
- Python-focused (ruff/pytest checks). Other languages get diff-only LLM review.
- No measured detection accuracy yet (see tests note above).
- Configurable retry loop, but not a fully autonomous multi-tool agent.
Project layout
reviewer/ # LangGraph package (state, schemas, prompts, nodes, routes, graph, cli)
examples/ # Sample diffs + tiny failing pytest package
tests/ # Unit, graph trajectory, and CLI tests
.github/workflows/ # ci, pr-review, commit-review, publish
pyproject.toml # packaging + dependencies
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file agentic_pr_reviewer-0.2.0.tar.gz.
File metadata
- Download URL: agentic_pr_reviewer-0.2.0.tar.gz
- Upload date:
- Size: 22.4 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
c53fade56c344a85e6246ab9fcf6ceaf5b850a0a92def61f5802139c0692cfae
|
|
| MD5 |
caf7746a62154cb7c115c1417843e3e8
|
|
| BLAKE2b-256 |
0bc4f247220babbeb819d463ae47e0541e53cf917b7a92e5dcdfbaa6e08cf65c
|
Provenance
The following attestation bundles were made for agentic_pr_reviewer-0.2.0.tar.gz:
Publisher:
publish.yml on BitBucket0/agentic-pr-reviewer
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
agentic_pr_reviewer-0.2.0.tar.gz -
Subject digest:
c53fade56c344a85e6246ab9fcf6ceaf5b850a0a92def61f5802139c0692cfae - Sigstore transparency entry: 2265778701
- Sigstore integration time:
-
Permalink:
BitBucket0/agentic-pr-reviewer@ec5b249526977c715fd9b657b875d85e7fcbd777 -
Branch / Tag:
refs/tags/v0.2.0 - Owner: https://github.com/BitBucket0
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@ec5b249526977c715fd9b657b875d85e7fcbd777 -
Trigger Event:
push
-
Statement type:
File details
Details for the file agentic_pr_reviewer-0.2.0-py3-none-any.whl.
File metadata
- Download URL: agentic_pr_reviewer-0.2.0-py3-none-any.whl
- Upload date:
- Size: 21.8 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
4bf347608b013ebda0977f44e1bc2c95aeaeccc22fb91c44d7a6fb7d837780fa
|
|
| MD5 |
283312719a60dcb62c109070b2ac4085
|
|
| BLAKE2b-256 |
5ece63b42ca460f5a54e39241988222e1588634ca3a594b94e7858cca048c6eb
|
Provenance
The following attestation bundles were made for agentic_pr_reviewer-0.2.0-py3-none-any.whl:
Publisher:
publish.yml on BitBucket0/agentic-pr-reviewer
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
agentic_pr_reviewer-0.2.0-py3-none-any.whl -
Subject digest:
4bf347608b013ebda0977f44e1bc2c95aeaeccc22fb91c44d7a6fb7d837780fa - Sigstore transparency entry: 2265778812
- Sigstore integration time:
-
Permalink:
BitBucket0/agentic-pr-reviewer@ec5b249526977c715fd9b657b875d85e7fcbd777 -
Branch / Tag:
refs/tags/v0.2.0 - Owner: https://github.com/BitBucket0
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@ec5b249526977c715fd9b657b875d85e7fcbd777 -
Trigger Event:
push
-
Statement type: