Adversarial Friends
Hand your spec, plan, or review to other agent CLIs —
claude,codex,agy,opencode— as independent adversarial reviewers, then merge their critiques into one ranked findings report.
It automates a workflow you may already do by hand: run a review, paste the findings into a different model, ask whether they hold up, carry the argument back. Doing that manually means holding a claim ledger in your head. This keeps the ledger on disk.
📋 Contents
- Why more than one model
- Install
- Quickstart
- How it works
- Lenses
- What you get back
- What's implemented
- Documentation
- Development
🎯 Why more than one model
A single reviewer produces confident prose. Several reviewers produce claims that can be compared — and the disagreements are where the real problems are.
This tool's own design spec was built exactly this way:
| Reviewer | Result |
|---|---|
codex |
17 findings |
claude |
15 findings, plus one marked unproven — "lens leaks attribution" |
agy |
independently reproduced two of claude's findings, and caught a shared-worktree race neither of the other two flagged |
That unproven claim was later confirmed and fixed. No single reviewer's pass
would have surfaced all of it — see the revision history in the design
spec
for the full account.
📦 Install
Requires Python 3.11+ and at least one agent CLI besides the one you're running under. The runner itself is stdlib-only — zero runtime dependencies.
uv tool install adversarial-friends
Other install methods
# From git, for a version that is not yet released
uv tool install git+https://github.com/livingstaccato/adversarial-friends
# From a local checkout
git clone https://github.com/livingstaccato/adversarial-friends
cd adversarial-friends
uv tool install .
# Without installing at all
python -m adversarial_friends doctor
Then confirm what's actually available:
afriend doctor
agy found schema=True readonly=True effort=native /Users/you/.local/bin/agy
claude found schema=True readonly=True effort=native /Users/you/.local/bin/claude
codex found schema=True readonly=True effort=native /opt/homebrew/bin/codex
ollama found schema=False readonly=False effort=none http://127.0.0.1:11434/api/generate
opencode found schema=False readonly=False effort=unverified /Users/you/.opencode/bin/opencode
For ollama, found means a reachable endpoint rather than a binary on
PATH; it shows unreachable when no server is listening.
doctor reports what each friend can genuinely enforce — schema
validation, a real read-only mode, a verifiable effort level — rather than
what it claims to support. opencode showing readonly=False is not a bug;
it has no read-only mode, so the tool says so instead of pretending.
🚀 Quickstart
afriend run docs/my-design.md --mode report
Stdout carries one thing — the run directory. Read report.md inside it:
cat "$(afriend run docs/my-design.md --mode report)/report.md"
Pick your reviewers and lenses explicitly:
afriend run spec.md --friend codex:security --friend claude:ops
A third slot picks the model — required for ollama, which has no default:
afriend run spec.md --friend ollama:security:qwen3:0.6b
⚠️
--friendreplaces discovery rather than adding to it. One--friendflag means a one-friend run — which cannot cross-examine anything. The tool records that as a downgrade inrun.jsonandreport.mdrather than letting it look like a full review.
A run takes minutes, not seconds — a friend is a whole agent CLI reading a
document. Progress goes to stderr: a line per friend as it finishes, and
a heartbeat naming whatever is still outstanding, so a quiet run is
distinguishable from a hung one. --no-progress silences it for a caller
that captures both streams together.
⚙️ How it works
Every friend gets its own prompt built from its own lens, runs in its own isolated directory, in its own process group:
| Stage | What happens |
|---|---|
| 🔍 Resolve | Discover agent CLIs on PATH, round-robin a lens to each |
| ✍️ Prompt | Build a per-friend prompt: shared contract header + that friend's lens prose + the artifact |
| 🔒 Isolate | Friends with a real read-only mode get a private git worktree from one shared snapshot. A CLI with no read-only mode is confined by the OS instead (sandbox-exec / bwrap) — or refused |
| ⚡ Dispatch | Parallel, one thread per friend, each in its own process group with a kill deadline of --timeout + 60s |
| 🧩 Normalize | Unwrap the CLI's own JSON envelope, strip ANSI, recover the payload, validate against the claim schema |
| 🔗 Merge | Exact-merge identical claims into aliases — accumulating origins so corroboration survives |
| 📄 Report | Rank findings, render report.md, write the append-only ledger |
The snapshot includes untracked files (git stash create omits them), and
the working tree is never touched — a friend reviewing your repo can't see a
half-staged index or scribble on your checkout.
Full run flow, step by step
Corroboration is the point
Two friends independently reaching the same conclusion is the strongest signal this tool produces, so deduplication is built to never destroy it:
Dedup is deliberately exact-match — whitespace and case only. Two friends describing one defect in different words produce two claims, which costs a round. Guessing at equivalence would corrupt the ledger, which is worse.
Claim states in cross-examination
Every claim --mode crossexam produces ends in one of eight states. Two of
them — deadlocked and settled-upheld — deliberately need a human, and the
report says so rather than quietly resolving them.
The gate loop
--mode gate is the one that fails a build, and clearing it is a
back-and-forth rather than a single command:
🔬 Lenses
A lens is prose, not a config string. Its text is injected into that friend's prompt, so it shapes what the friend actually looks for.
| Lens | Default scope | Requires a failure scenario |
|---|---|---|
assumptions |
doc | ✅ |
security |
repo | ✅ |
ops |
repo | ✅ |
testability |
repo | ✅ |
spec-vs-reality |
repo | ✅ |
scope |
doc | ❌ — advisory only |
scope is the one lens that doesn't demand a concrete failure scenario;
"this is more than you need" is a legitimate finding without one. Claims from
it are marked (advisory) in the report so they never carry the same weight
as a reproducible defect.
📂 What you get back
<run-dir>/
├── report.md ← ranked findings, corroboration, downgrades
├── run.json ← machine-readable: friends, statuses, downgrades
├── claims.jsonl ← append-only ledger: claims, aliases
├── artifact/ ← frozen copy of what was reviewed, hashed
└── round-1/
├── <friend>.prompt ← exactly what this friend was asked
├── <friend>.raw ← its unmodified stdout
├── <friend>.err ← its stderr (always present, even when empty)
├── <friend>.meta ← argv, exit code, duration, timeout, orphan status
└── <friend>.sandbox ← the OS policy it ran under, when one was applied
Runs land under ${XDG_STATE_HOME:-~/.local/state}/adversarial-friends/runs/,
or wherever --out points.
Everything a friend was asked and everything it said is on disk. When a run comes back thin, that's what you read — not a guess.
✅ What's implemented
All four modes run.
| Mode | What it does |
|---|---|
report |
One round. Every friend critiques in parallel; claims merge into one ranked report. |
crossexam |
Then friends judge the claims they did not write, blind, until each settles or deadlocks. |
gate |
Then every non-advisory claim that did not clear needs an explicit resolution — this is the one that fails a build. |
loop |
Repeats until two consecutive rounds surface nothing new. |
afriend run docs/design.md --mode crossexam
afriend run docs/design.md --mode gate # exit 1 while anything blocks
afriend resolve <run-id> --claim c-0001@1 \
--disposition fixed --evidence src/auth.py:38
Disagreement is the output rather than a problem: two judges who still
disagree at --max-rounds leave the claim deadlocked, and the report
quotes both sides verbatim instead of resolving it by majority.
A resolution is an attestation, and the tool says so. It cannot know a
defect is gone — only whether the location you named actually changed since
the run started. A fix that landed outside the reviewed artifact is fine.
--disposition fixed requires a verifiably changed location; unchanged or
unverifiable evidence is refused. Use accepted-risk when verification is
intentionally unavailable.
Deduplication is judgment the runner declines to fake. --merge exact
(the default) merges only identical claims and always finishes unaided;
--merge orchestrator stops with exit 10, writes the claims to
REQUEST.json, and waits for you to say which are duplicates:
afriend run docs/design.md --merge orchestrator # exit 10, writes REQUEST.json
# ...fill in the merges, save as RESPONSE.json...
afriend run --resume <run-id> # round 1 is not re-run
Tired of --friend flags? afriend init writes a roster from what is
actually installed, and ~/.config/adversarial-friends/roster.toml is picked
up automatically. A repo-local roster never is — a cloned repo does not get
to choose who reviews it (§13).
The same halt serves unparseable output (§14.2): repair is a pure transformation with no model call, so when it fails the runner asks you to read the raw text rather than discarding whatever the friend found.
There is no --max-spend-usd. A dollar cap needs per-CLI cost reporting
nobody has captured, and a flag that silently never fires is worse than none
— you would set it and believe you were protected. Use --max-calls, which
is derived from your roster and actually enforced.
| Friend | Status |
|---|---|
claude |
✅ ships |
codex |
✅ ships |
agy |
✅ ships |
opencode |
✅ ships — no read-only mode, reported honestly |
ollama |
✅ ships — local models over HTTP, no schema/read-only to enforce; needs an explicit model |
There is no gemini adapter: the gemini CLI returns an ineligible-tier
error on the individual free tier, and Google's own supported path from there
is Antigravity — which is agy.
Exit codes
| Code | Meaning |
|---|---|
0 |
the run reached terminal states with nothing blocked |
1 |
a gate still has claims needing a resolution, or every dispatched friend failed |
2 |
usage error — bad flag, unknown CLI, missing artifact |
3 |
no usable friends found at all |
10 |
--merge orchestrator is waiting for you to adjudicate merges |
11 |
a ceiling was hit — the run was truncated, not decided |
12 |
--require-friends N was set and fewer than N friends answered |
128+N |
aborted by signal N — isolation torn down, friends killed |
A ceiling outranks everything below it, so a CI wrapper can read 11 as
"retry" and 1 as "block" without ambiguity.
📚 Documentation
| Where | What |
|---|---|
| docs/ | Documentation index |
| SKILL.md | The skill itself — when it fires, how to read its output |
| modes.md | report, crossexam, gate, loop — and which are real |
| ledger.md | Claim, verdict, alias, and resolution records |
| troubleshooting.md | Verified CLI traps, empty reports, timeouts |
| architecture/ | Diagrams and their sources |
| design spec | The full design, including the adversarial review that produced it |
Using it as a skill or plugin
The skill payload ships inside the wheel as package data, and is mirrored
under plugins/ for loaders that can't install a Python package:
# Claude Code
/plugin marketplace add /path/to/adversarial-friends/plugins
The skill invokes afriend, so the package must be installed for it to work —
afriend doctor is the check.
🛠 Development
make install # uv sync
make test # pytest
make quality # every portable CI gate, wheel checks, and tests
make diagrams # re-render docs/architecture/*.puml
make quality runs every portable CI gate, including wheel construction and
isolated installation. Linux CI additionally installs bubblewrap and requires
the real OS-confinement tests to execute; macOS cannot reproduce that Linux-
specific assertion locally. Use make act-ci for the closest local Linux run.
Two gates catch drift that is otherwise silent:
plugin-sync—src/adversarial_friends/assets/is canonical; theplugins/tree is a byte-identical mirror. Edit assets, thenmake plugin-sync-copy.version-sync—VERSIONmust match theversionfield in every plugin manifest.
See AGENTS.md for repository layout and conventions.
📄 License
MIT — see LICENSE.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file adversarial_friends-0.2.0.tar.gz.
File metadata
- Download URL: adversarial_friends-0.2.0.tar.gz
- Upload date:
- Size: 374.2 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
391c025127425bee078aad493366d2a8a9f1405cb7ce1312958d63b51aaf155c
|
|
| MD5 |
68378ad496a07690add39ad74a210e4e
|
|
| BLAKE2b-256 |
ef566d09d9c6edf4867c8e171727840d675717a683da4c37f17312da79296786
|
Provenance
The following attestation bundles were made for adversarial_friends-0.2.0.tar.gz:
Publisher:
release.yml on livingstaccato/adversarial-friends
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
adversarial_friends-0.2.0.tar.gz -
Subject digest:
391c025127425bee078aad493366d2a8a9f1405cb7ce1312958d63b51aaf155c - Sigstore transparency entry: 2648591740
- Sigstore integration time:
-
Permalink:
livingstaccato/adversarial-friends@21e990f64eb7b0e7dc2c2865d845a850e24e406f -
Branch / Tag:
refs/tags/v0.2.0 - Owner: https://github.com/livingstaccato
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@21e990f64eb7b0e7dc2c2865d845a850e24e406f -
Trigger Event:
push
-
Statement type:
File details
Details for the file adversarial_friends-0.2.0-py3-none-any.whl.
File metadata
- Download URL: adversarial_friends-0.2.0-py3-none-any.whl
- Upload date:
- Size: 241.1 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
c9b8f6ce6ec8a03a02558d38d286ffa792bf2ab457c745abecdeb2fe81c63c27
|
|
| MD5 |
69a190995985e0f336e885bdb2402a0d
|
|
| BLAKE2b-256 |
e79c74d706f72a29055cf058d8e766cecb8bb5867eecf89abde08b8706e5be97
|
Provenance
The following attestation bundles were made for adversarial_friends-0.2.0-py3-none-any.whl:
Publisher:
release.yml on livingstaccato/adversarial-friends
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
adversarial_friends-0.2.0-py3-none-any.whl -
Subject digest:
c9b8f6ce6ec8a03a02558d38d286ffa792bf2ab457c745abecdeb2fe81c63c27 - Sigstore transparency entry: 2648591897
- Sigstore integration time:
-
Permalink:
livingstaccato/adversarial-friends@21e990f64eb7b0e7dc2c2865d845a850e24e406f -
Branch / Tag:
refs/tags/v0.2.0 - Owner: https://github.com/livingstaccato
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@21e990f64eb7b0e7dc2c2865d845a850e24e406f -
Trigger Event:
push
-
Statement type: