ADRA · Adversarial Dev Review Agent
A client-agnostic, deterministic-first, adversarial-validation engine that supports the software lifecycle: it reviews PRs, designs and runs validation/refutation experiments, writes documentation back, and escalates to a human exactly where a senior engineer would. Governed by LLM-as-Judge + a blocking adversarial critic, grounded by deterministic tools, with immutable provenance. Runs offline with no API key.
pip install adra · Python ≥ 3.11 · Apache-2.0 · status: v0.04.001
What's different
The AI-code-review market splits in two, and both miss the same spot:
- Reviewers (CodeRabbit, Greptile, Qodo, Korbit, Sourcery, Bito) feed linters into an LLM, but the model's prose is the verdict: hallucinated and "consistently-stated- but-false" findings leak through; the deterministic signals are inputs, never the gate.
- Autonomous coders (Devin, OpenHands, SWE-agent, Sweep) write code and treat "tests pass" as success rather than adversarially trying to prove the change wrong.
ADRA occupies the gap: a deterministic spine (git / CI / static analysis / SQL probes) that grounds a blocking adversarial critic whose job is to refute, not bless, each artifact: every finding carrying its evidence, with disciplined human escalation when nothing deterministic backs the verdict. Existing tools generate opinions; ADRA generates proofs and refutations, and escalates when it can't.
The six capabilities
| Skill | What it does |
|---|---|
code_review |
Review a diff: language/leak scan + test-discoverability + exact CI command + semantic findings |
pr_eval |
Evaluate a PR: merge-base health → bundle validate → conformance → verdict + PR body |
experiment |
Hypothesis-driven validation experiment: SQL-warehouse probes + synthesis |
improve |
Minimum-functional improvement proposal (prune filler, smallest safe diff) |
document |
Turn a run record into a PR page / experiment page / methodology-history row |
decide |
Route analysis: candidate routes + trade-offs + recommendation · human-owned |
Each skill is the same loop, differing only by its domain prompt and deterministic tools.
Use it as a Claude Code skill
ADRA also ships as a portable Claude Code Agent Skill at
skills/adra-applied/: the same deterministic-first method
delivered into an interactive coding agent. It has no model runtime of its own (the
agent harness is the runtime) and reaches systems through the already-authenticated
CLIs (gh, az, databricks, git), read-only by default. Install it by copying
the folder into ~/.claude/skills/, or enable this repo as a plugin. See
skills/README.md.
Why deterministic-first
Tools (git, the exact CI command, bundle validate, language scan, SQL probe) run
first and become both the grounding the model may not contradict and the evidence in
the provenance log. Because the deterministic floor carries the verdict, the whole loop
runs, and the test suite passes, offline with no API key. Connecting a real provider
adds the semantic layer on top.
Architecture
intake ─▶ plan ─▶ ground (deterministic tools) ─▶ generate ─▶ CRITIC ─┐
▲ revise ◀──────┘
└── accepted / escalate ─▶ artifacts + run record
adra/state.py· the typed domain model (Severity,Finding,ToolResult,CriticVerdict,RunState). One contract end to end.adra/rubric.py· the shared adversarial rubric (criteria as typed data); drives both the deterministic critic and the critic prompt, so "what we check" never drifts.adra/orchestrator.py· the hand-rolled, framework-free state machine.adra/critic.py· deterministic red-team pass (rubric-driven) + LLM semantic attacks.adra/judge.py· rubric scoring with swap-and-average + reference anchoring.adra/llm.py· the tiny ADRA-ownedChatModelseam:mock(offline) | any real provider via pydantic-ai (provider:model, config-only). No LangChain/LangGraph.adra/tools/· each returns aToolResult(git / CI / bundle / lang / discovery / sql).adra/skills/· theSkillbase + the six skills.adra/clients/· client governance suites (the bundled fictional Northwind Data Platform); selectable viaADRA_CLIENT_DIR.adra/provenance.py· the immutable run record (the deep change-history layer).cli/· theadracommand.refs/· annotated bibliography + papers.docs/· deep docs + engine ADRs (docs/adr/).
Quickstart (offline, no key)
python -m venv .venv && . .venv/Scripts/activate # Windows: .venv\Scripts\activate
pip install -e ".[dev]"
pytest -q # 11 passing, fully offline
python scripts/demo_offline.py # end-to-end demo, all six skills
Expected: the stale-base PR is blocked + escalated (12 commits behind, a notebook
deletion, a .yml → .yml.t rename dropping a bundle resource, bundle validate failing);
the language/leak review is blocked; the clean experiment / improve / document / decide
runs are accepted.
adra review path/to.diff --ci-command 'python -m coverage run -m unittest discover -s . -p "test*.py"'
adra decide "Raise the refresh cadence" "edit the shared CI template" "change it in the owning repo"
Enable a real provider
pip install -e ".[llm]" # pydantic-ai; the seam is provider-agnostic
export ANTHROPIC_API_KEY=... # or put it in .env (any provider's key works)
export ADRA_PROVIDER=anthropic # adding a provider is config only: ADRA_PROVIDER / ADRA_MODEL / ADRA_MODEL_<ROLE>
adra pr-eval --source task/123/x --repo /path/to/repo --external
--external (or ADRA_ALLOW_EXTERNAL=1) lets the tools actually run git / the CI command.
Default is dry-run / read-only.
Client-agnostic grounding
A client = a governance suite (conventions, ADRs, CI standards, glossary, incident cases)
the engine grounds on. ADRA ships a complete, fictional client, Northwind Data
Platform, under adra/clients/synthetic/northwind/. Point ADRA at any client:
export ADRA_CLIENT_DIR=/path/to/your/standards # or Settings(client_dir=...)
The rubric references the suite by id and the prompts cite it; the engine code does not change per client.
Connectors & emulator
The engine grounds through one connector Protocol so the same skills run against a real
platform or a synthetic one:
- Real: GitHub (PRs/reviews/issues/contents via a thin
httpxREST v3 client), Azure DevOps (REST 7.1), Databricks (databricks-sdk+ bundles CLI), Azure (azure-identity+ monitor / health). - Emulator: a self-contained platform (synthetic git repos + PRs + wiki + boards + CI + a SQLite warehouse) so the full flow runs offline.
The GitHub, Azure DevOps, Databricks, and Azure connectors are implemented; the offline emulator runs the full flow with no external calls. Each is enabled by its
pip install adra[...]extra and exercised read-only by default.
Security model
Deterministic floor (tools are ground truth; the LLM cannot overturn a blocker) · read-only
by default (writes require --external and explicit human confirmation) · human gates on
PR create / push / merge and any risk claim · English-only + AI-authorship-leak scan on
anything written to disk · immutable provenance for every run. The agent reads untrusted
repo/PR/issue content; a dual-LLM / capability split + sandboxed, egress-filtered execution are
planned for the connector phase (not yet implemented; OWASP LLM/Agentic Top-10).
Two-repo layout
ADRA is the public-destined OSS engine (this repo): no secrets, ever. A separate private ADRA Console (a private web app + backend) consumes this engine for experiments and real connections behind access control. The engine is the serious tool you can run anywhere with your own tokens; the console is the connected instance.
Extending
- New criterion: add a
RubricItemtoadra/rubric.py(it shows up in the critic prompt and, ifkind="deterministic", wire its check incritic.py). - New capability: add a
Skillsubclass + aprompts/<skill>.md, register it inadra/skills/__init__.py, add aNode. - New tool: a function returning a
ToolResult; call it from a skill'sground. - New provider: config only; set
ADRA_PROVIDER/ADRA_MODEL/ADRA_MODEL_<ROLE>(pydantic-ai resolves theprovider:model); no new code.
Quality gates
python scripts/check_content.py is the content guard (ADR-0067): it fails on any em-dash or
emoji in the tracked text surfaces and lists directional arrows found in Markdown prose so a
reviewer can confirm they are notation; tests/test_content_guard.py runs it inside pytest -q.
Status
v0.04.001: the engine is complete and green offline, and the GitHub / Azure DevOps / Databricks
/ Azure connectors plus the offline emulator are implemented. Multi-industry synthetic clients and
the web console are the next phases. Stays 0.x while connectors are partly untested-live.
License
Apache-2.0; see LICENSE.
References
refs/ holds an annotated bibliography (refs/README.md) + BibTeX (refs/references.bib)
- the core papers (ReAct, Reflexion, Self-Refine, Constitutional AI, LLM-as-judge bias, agentic security / CaMeL, NIST AI RMF, OWASP LLM/Agentic, provenance, code-review agents).
Metadata
Release files for adra 0.6.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| adra-0.6.0.tar.gz | 82.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| adra-0.6.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 177.2 kB
Release files / adra-0.6.0.tar.gz
| Download URL | adra-0.6.0.tar.gz |
|---|---|
| Size | 82.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
35247c03f9e6f4e5cc7894a8363c8e53d67173011aee73800d5591a380bc256e
|
|
BLAKE2b-256 checksum How to use checksums |
b427c5e66b75ac2b8e818eb6fdcff722f049ffe69bdb7fdae124b64c368279ce
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 12, 2026.
Transparency logRelease files / adra-0.6.0-py3-none-any.whl
| Download URL | adra-0.6.0-py3-none-any.whl |
|---|---|
| Size | 95.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
3c2d846aa69d4c50b221f50e515a46a2c46a9dd4df308b6dd5afd4068e876880
|
|
BLAKE2b-256 checksum How to use checksums |
00de95817208a3bced9ed24c0c6bde88835330a5508fa0c5bc7b9f817b516062
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 12, 2026.
Transparency log