Skip to main content

ADRA · Adversarial Dev Review Agent

License Version

A client-agnostic, deterministic-first, adversarial-validation engine that supports the software lifecycle: it reviews PRs, designs and runs validation/refutation experiments, writes documentation back, and escalates to a human exactly where a senior engineer would. Governed by LLM-as-Judge + a blocking adversarial critic, grounded by deterministic tools, with immutable provenance. Runs offline with no API key.

pip install adra · Python ≥ 3.11 · Apache-2.0 · status: v0.04.001


What's different

The AI-code-review market splits in two, and both miss the same spot:

  • Reviewers (CodeRabbit, Greptile, Qodo, Korbit, Sourcery, Bito) feed linters into an LLM, but the model's prose is the verdict: hallucinated and "consistently-stated- but-false" findings leak through; the deterministic signals are inputs, never the gate.
  • Autonomous coders (Devin, OpenHands, SWE-agent, Sweep) write code and treat "tests pass" as success rather than adversarially trying to prove the change wrong.

ADRA occupies the gap: a deterministic spine (git / CI / static analysis / SQL probes) that grounds a blocking adversarial critic whose job is to refute, not bless, each artifact: every finding carrying its evidence, with disciplined human escalation when nothing deterministic backs the verdict. Existing tools generate opinions; ADRA generates proofs and refutations, and escalates when it can't.

The six capabilities

Skill What it does
code_review Review a diff: language/leak scan + test-discoverability + exact CI command + semantic findings
pr_eval Evaluate a PR: merge-base health → bundle validate → conformance → verdict + PR body
experiment Hypothesis-driven validation experiment: SQL-warehouse probes + synthesis
improve Minimum-functional improvement proposal (prune filler, smallest safe diff)
document Turn a run record into a PR page / experiment page / methodology-history row
decide Route analysis: candidate routes + trade-offs + recommendation · human-owned

Each skill is the same loop, differing only by its domain prompt and deterministic tools.

Use it as a Claude Code skill

ADRA also ships as a portable Claude Code Agent Skill at skills/adra-applied/: the same deterministic-first method delivered into an interactive coding agent. It has no model runtime of its own (the agent harness is the runtime) and reaches systems through the already-authenticated CLIs (gh, az, databricks, git), read-only by default. Install it by copying the folder into ~/.claude/skills/, or enable this repo as a plugin. See skills/README.md.

Why deterministic-first

Tools (git, the exact CI command, bundle validate, language scan, SQL probe) run first and become both the grounding the model may not contradict and the evidence in the provenance log. Because the deterministic floor carries the verdict, the whole loop runs, and the test suite passes, offline with no API key. Connecting a real provider adds the semantic layer on top.

Architecture

intake ─▶ plan ─▶ ground (deterministic tools) ─▶ generate ─▶ CRITIC ─┐
                                                      ▲   revise ◀──────┘
                                                      └── accepted / escalate ─▶ artifacts + run record
  • adra/state.py · the typed domain model (Severity, Finding, ToolResult, CriticVerdict, RunState). One contract end to end.
  • adra/rubric.py · the shared adversarial rubric (criteria as typed data); drives both the deterministic critic and the critic prompt, so "what we check" never drifts.
  • adra/orchestrator.py · the hand-rolled, framework-free state machine.
  • adra/critic.py · deterministic red-team pass (rubric-driven) + LLM semantic attacks.
  • adra/judge.py · rubric scoring with swap-and-average + reference anchoring.
  • adra/llm.py · the tiny ADRA-owned ChatModel seam: mock (offline) | any real provider via pydantic-ai (provider:model, config-only). No LangChain/LangGraph.
  • adra/tools/ · each returns a ToolResult (git / CI / bundle / lang / discovery / sql).
  • adra/skills/ · the Skill base + the six skills.
  • adra/clients/ · client governance suites (the bundled fictional Northwind Data Platform); selectable via ADRA_CLIENT_DIR.
  • adra/provenance.py · the immutable run record (the deep change-history layer).
  • cli/ · the adra command. refs/ · annotated bibliography + papers. docs/ · deep docs + engine ADRs (docs/adr/).

Quickstart (offline, no key)

python -m venv .venv && . .venv/Scripts/activate     # Windows: .venv\Scripts\activate
pip install -e ".[dev]"
pytest -q                                            # 11 passing, fully offline
python scripts/demo_offline.py                       # end-to-end demo, all six skills

Expected: the stale-base PR is blocked + escalated (12 commits behind, a notebook deletion, a .yml → .yml.t rename dropping a bundle resource, bundle validate failing); the language/leak review is blocked; the clean experiment / improve / document / decide runs are accepted.

adra review path/to.diff --ci-command 'python -m coverage run -m unittest discover -s . -p "test*.py"'
adra decide "Raise the refresh cadence" "edit the shared CI template" "change it in the owning repo"

Enable a real provider

pip install -e ".[llm]"                  # pydantic-ai; the seam is provider-agnostic
export ANTHROPIC_API_KEY=...             # or put it in .env (any provider's key works)
export ADRA_PROVIDER=anthropic           # adding a provider is config only: ADRA_PROVIDER / ADRA_MODEL / ADRA_MODEL_<ROLE>
adra pr-eval --source task/123/x --repo /path/to/repo --external

--external (or ADRA_ALLOW_EXTERNAL=1) lets the tools actually run git / the CI command. Default is dry-run / read-only.

Client-agnostic grounding

A client = a governance suite (conventions, ADRs, CI standards, glossary, incident cases) the engine grounds on. ADRA ships a complete, fictional client, Northwind Data Platform, under adra/clients/synthetic/northwind/. Point ADRA at any client:

export ADRA_CLIENT_DIR=/path/to/your/standards   # or Settings(client_dir=...)

The rubric references the suite by id and the prompts cite it; the engine code does not change per client.

Connectors & emulator

The engine grounds through one connector Protocol so the same skills run against a real platform or a synthetic one:

  • Real: GitHub (PRs/reviews/issues/contents via a thin httpx REST v3 client), Azure DevOps (REST 7.1), Databricks (databricks-sdk + bundles CLI), Azure (azure-identity + monitor / health).
  • Emulator: a self-contained platform (synthetic git repos + PRs + wiki + boards + CI + a SQLite warehouse) so the full flow runs offline.

The GitHub, Azure DevOps, Databricks, and Azure connectors are implemented; the offline emulator runs the full flow with no external calls. Each is enabled by its pip install adra[...] extra and exercised read-only by default.

Security model

Deterministic floor (tools are ground truth; the LLM cannot overturn a blocker) · read-only by default (writes require --external and explicit human confirmation) · human gates on PR create / push / merge and any risk claim · English-only + AI-authorship-leak scan on anything written to disk · immutable provenance for every run. The agent reads untrusted repo/PR/issue content; a dual-LLM / capability split + sandboxed, egress-filtered execution are planned for the connector phase (not yet implemented; OWASP LLM/Agentic Top-10).

Two-repo layout

ADRA is the public-destined OSS engine (this repo): no secrets, ever. A separate private ADRA Console (a private web app + backend) consumes this engine for experiments and real connections behind access control. The engine is the serious tool you can run anywhere with your own tokens; the console is the connected instance.

Extending

  • New criterion: add a RubricItem to adra/rubric.py (it shows up in the critic prompt and, if kind="deterministic", wire its check in critic.py).
  • New capability: add a Skill subclass + a prompts/<skill>.md, register it in adra/skills/__init__.py, add a Node.
  • New tool: a function returning a ToolResult; call it from a skill's ground.
  • New provider: config only; set ADRA_PROVIDER / ADRA_MODEL / ADRA_MODEL_<ROLE> (pydantic-ai resolves the provider:model); no new code.

Quality gates

python scripts/check_content.py is the content guard (ADR-0067): it fails on any em-dash or emoji in the tracked text surfaces and lists directional arrows found in Markdown prose so a reviewer can confirm they are notation; tests/test_content_guard.py runs it inside pytest -q.

Status

v0.04.001: the engine is complete and green offline, and the GitHub / Azure DevOps / Databricks / Azure connectors plus the offline emulator are implemented. Multi-industry synthetic clients and the web console are the next phases. Stays 0.x while connectors are partly untested-live.

License

Apache-2.0; see LICENSE.

References

refs/ holds an annotated bibliography (refs/README.md) + BibTeX (refs/references.bib)

  • the core papers (ReAct, Reflexion, Self-Refine, Constitutional AI, LLM-as-judge bias, agentic security / CaMeL, NIST AI RMF, OWASP LLM/Agentic, provenance, code-review agents).

Metadata

Release files for adra 0.6.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for adra 0.6.0
File Size Uploaded
adra-0.6.0.tar.gz 82.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for adra 0.6.0
File Interpreter ABI Platform
adra-0.6.0-py3-none-any.whl Python 3 none any Details

Total release size: 177.2 kB

Release files / adra-0.6.0.tar.gz

Download URL adra-0.6.0.tar.gz
Size 82.1 kB
Tags Source
SHA-256 checksum
How to use checksums
35247c03f9e6f4e5cc7894a8363c8e53d67173011aee73800d5591a380bc256e
BLAKE2b-256 checksum
How to use checksums
b427c5e66b75ac2b8e818eb6fdcff722f049ffe69bdb7fdae124b64c368279ce
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 12, 2026.

Transparency log

Release files / adra-0.6.0-py3-none-any.whl

Download URL adra-0.6.0-py3-none-any.whl
Size 95.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
3c2d846aa69d4c50b221f50e515a46a2c46a9dd4df308b6dd5afd4068e876880
BLAKE2b-256 checksum
How to use checksums
00de95817208a3bced9ed24c0c6bde88835330a5508fa0c5bc7b9f817b516062
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 12, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.6.0 This release

2 release files

0.4.1

2 release files

0.4.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page