Skip to main content

adversarial-friends

Adversarial Friends

Hand your spec, plan, or review to other agent CLIs — claude, codex, agy, opencode — as independent adversarial reviewers, then merge their critiques into one ranked findings report.

Python Dependencies License Tests

It automates a workflow you may already do by hand: run a review, paste the findings into a different model, ask whether they hold up, carry the argument back. Doing that manually means holding a claim ledger in your head. This keeps the ledger on disk.


📋 Contents


🎯 Why more than one model

A single reviewer produces confident prose. Several reviewers produce claims that can be compared — and the disagreements are where the real problems are.

This tool's own design spec was built exactly this way:

Reviewer Result
codex 17 findings
claude 15 findings, plus one marked unproven"lens leaks attribution"
agy independently reproduced two of claude's findings, and caught a shared-worktree race neither of the other two flagged

That unproven claim was later confirmed and fixed. No single reviewer's pass would have surfaced all of it — see the revision history in the design spec for the full account.


📦 Install

Requires Python 3.11+ and at least one agent CLI besides the one you're running under. The runner itself is stdlib-only — zero runtime dependencies.

uv tool install adversarial-friends
Other install methods
# From git, for a version that is not yet released
uv tool install git+https://github.com/livingstaccato/adversarial-friends

# From a local checkout
git clone https://github.com/livingstaccato/adversarial-friends
cd adversarial-friends
uv tool install .

# Without installing at all
python -m adversarial_friends doctor

Then confirm what's actually available:

afriend doctor
agy        found    schema=True  readonly=True  effort=native      /Users/you/.local/bin/agy
claude     found    schema=True  readonly=True  effort=native      /Users/you/.local/bin/claude
codex      found    schema=True  readonly=True  effort=native      /opt/homebrew/bin/codex
opencode   found    schema=False readonly=False effort=unverified  /Users/you/.opencode/bin/opencode
ollama     found    schema=False readonly=False effort=none       http://127.0.0.1:11434/api/generate

For ollama, found means a reachable endpoint rather than a binary on PATH; it shows unreachable when no server is listening.

doctor reports what each friend can genuinely enforce — schema validation, a real read-only mode, a verifiable effort level — rather than what it claims to support. opencode showing readonly=False is not a bug; it has no read-only mode, so the tool says so instead of pretending.


🚀 Quickstart

afriend run docs/my-design.md --mode report

It prints one thing — the run directory. Read report.md inside it:

cat "$(afriend run docs/my-design.md --mode report)/report.md"

Pick your reviewers and lenses explicitly:

afriend run spec.md --friend codex:security --friend claude:ops

A third slot picks the model — required for ollama, which has no default:

afriend run spec.md --friend ollama:security:qwen3:0.6b

⚠️ --friend replaces discovery rather than adding to it. One --friend flag means a one-friend run — which cannot cross-examine anything. The tool records that as a downgrade in run.json and report.md rather than letting it look like a full review.


⚙️ How it works

module architecture

Every friend gets its own prompt built from its own lens, runs in its own isolated directory, in its own process group:

Stage What happens
🔍 Resolve Discover agent CLIs on PATH, round-robin a lens to each
✍️ Prompt Build a per-friend prompt: shared contract header + that friend's lens prose + the artifact
🔒 Isolate Friends with a real read-only mode get a private git worktree from one shared snapshot. A CLI with no read-only mode is confined by the OS instead (sandbox-exec / bwrap) — or refused
Dispatch Parallel, one thread per friend, each in its own process group with a kill deadline of --timeout + 60s
🧩 Normalize Unwrap the CLI's own JSON envelope, strip ANSI, recover the payload, validate against the claim schema
🔗 Merge Exact-merge identical claims into aliases — accumulating origins so corroboration survives
📄 Report Rank findings, render report.md, write the append-only ledger

The snapshot includes untracked files (git stash create omits them), and the working tree is never touched — a friend reviewing your repo can't see a half-staged index or scribble on your checkout.

Full run flow, step by step

run flow

Corroboration is the point

Two friends independently reaching the same conclusion is the strongest signal this tool produces, so deduplication is built to never destroy it:

claim lifecycle

Dedup is deliberately exact-match — whitespace and case only. Two friends describing one defect in different words produce two claims, which costs a round. Guessing at equivalence would corrupt the ledger, which is worse.

Claim states in cross-examination

Every claim --mode crossexam produces ends in one of eight states. Two of them — deadlocked and settled-upheld — deliberately need a human, and the report says so rather than quietly resolving them.

crossexam states

The gate loop

--mode gate is the one that fails a build, and clearing it is a back-and-forth rather than a single command:

gate workflow


🔬 Lenses

A lens is prose, not a config string. Its text is injected into that friend's prompt, so it shapes what the friend actually looks for.

Lens Default scope Requires a failure scenario
assumptions doc
security repo
ops repo
testability repo
spec-vs-reality repo
scope doc ❌ — advisory only

scope is the one lens that doesn't demand a concrete failure scenario; "this is more than you need" is a legitimate finding without one. Claims from it are marked (advisory) in the report so they never carry the same weight as a reproducible defect.


📂 What you get back

<run-dir>/
├── report.md          ← ranked findings, corroboration, downgrades
├── run.json           ← machine-readable: friends, statuses, downgrades
├── claims.jsonl       ← append-only ledger: claims, aliases
├── artifact/          ← frozen copy of what was reviewed, hashed
└── round-1/
    ├── <friend>.prompt ← exactly what this friend was asked
    ├── <friend>.raw    ← its unmodified stdout
    ├── <friend>.err    ← its stderr (always present, even when empty)
    └── <friend>.meta   ← argv, exit code, duration, timeout, orphan status

Runs land under ${XDG_STATE_HOME:-~/.local/state}/adversarial-friends/runs/, or wherever --out points.

Everything a friend was asked and everything it said is on disk. When a run comes back thin, that's what you read — not a guess.


✅ What's implemented

All four modes run.

Mode What it does
report One round. Every friend critiques in parallel; claims merge into one ranked report.
crossexam Then friends judge the claims they did not write, blind, until each settles or deadlocks.
gate Then every non-advisory claim that did not clear needs an explicit resolution — this is the one that fails a build.
loop Repeats until two consecutive rounds surface nothing new.
afriend run docs/design.md --mode crossexam
afriend run docs/design.md --mode gate       # exit 1 while anything blocks
afriend resolve <run-id> --claim c-0001@1 \
    --disposition fixed --evidence src/auth.py:38

Disagreement is the output rather than a problem: two judges who still disagree at --max-rounds leave the claim deadlocked, and the report quotes both sides verbatim instead of resolving it by majority.

A resolution is an attestation, and the tool says so. It cannot know a defect is gone — only whether the location you named actually changed since the run started. A fix that landed outside the reviewed artifact is fine; a location it cannot reconstruct is recorded as unverifiable rather than waved through. The one thing it refuses is --disposition fixed naming a location that did not change.

Deduplication is judgment the runner declines to fake. --merge exact (the default) merges only identical claims and always finishes unaided; --merge orchestrator stops with exit 10, writes the claims to REQUEST.json, and waits for you to say which are duplicates:

afriend run docs/design.md --merge orchestrator   # exit 10, writes REQUEST.json
# ...fill in the merges, save as RESPONSE.json...
afriend run --resume <run-id>                     # round 1 is not re-run

Tired of --friend flags? afriend init writes a roster from what is actually installed, and ~/.config/adversarial-friends/roster.toml is picked up automatically. A repo-local roster never is — a cloned repo does not get to choose who reviews it (§13).

The same halt serves unparseable output (§14.2): repair is a pure transformation with no model call, so when it fails the runner asks you to read the raw text rather than discarding whatever the friend found.

There is no --max-spend-usd. A dollar cap needs per-CLI cost reporting nobody has captured, and a flag that silently never fires is worse than none — you would set it and believe you were protected. Use --max-calls, which is derived from your roster and actually enforced.

Friend Status
claude ✅ ships
codex ✅ ships
agy ✅ ships
opencode ✅ ships — no read-only mode, reported honestly
ollama ✅ ships — local models over HTTP, no schema/read-only to enforce; needs an explicit model

There is no gemini adapter: the gemini CLI returns an ineligible-tier error on the individual free tier, and Google's own supported path from there is Antigravity — which is agy.

Exit codes

Code Meaning
0 at least one friend produced a usable critique
1 ran, but every dispatched friend failed
2 usage error — bad flag, unknown CLI, unimplemented mode
3 no usable friends found at all
128+N aborted by signal N — isolation torn down, friends killed

📚 Documentation

Where What
docs/ Documentation index
SKILL.md The skill itself — when it fires, how to read its output
modes.md report, crossexam, gate, loop — and which are real
ledger.md Claim, verdict, alias, and resolution records
troubleshooting.md Verified CLI traps, empty reports, timeouts
architecture/ Diagrams and their sources
design spec The full design, including the adversarial review that produced it

Using it as a skill or plugin

The skill payload ships inside the wheel as package data, and is mirrored under plugins/ for loaders that can't install a Python package:

# Claude Code
/plugin marketplace add /path/to/adversarial-friends/plugins

The skill invokes afriend, so the package must be installed for it to work — afriend doctor is the check.


🛠 Development

make install    # uv sync
make test       # pytest — 365 tests
make quality    # lint + type-check + every sync gate + tests
make diagrams   # re-render docs/architecture/*.puml

make quality runs exactly what CI runs. Two gates catch drift that is otherwise silent:

  • plugin-syncsrc/adversarial_friends/assets/ is canonical; the plugins/ tree is a byte-identical mirror. Edit assets, then make plugin-sync-copy.
  • version-syncVERSION must match the version field in every plugin manifest.

See AGENTS.md for repository layout and conventions.


📄 License

MIT — see LICENSE.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

adversarial_friends-0.1.5.tar.gz (327.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

adversarial_friends-0.1.5-py3-none-any.whl (214.4 kB view details)

Uploaded Python 3

File details

Details for the file adversarial_friends-0.1.5.tar.gz.

File metadata

  • Download URL: adversarial_friends-0.1.5.tar.gz
  • Upload date:
  • Size: 327.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.13.3

File hashes

Hashes for adversarial_friends-0.1.5.tar.gz
Algorithm Hash digest
SHA256 d2b6f6155dcea3f674c9c8d972dc9ac4f6a9b2ec846353512ab6bf1912bd5169
MD5 111952198a9760d47083e17995949fbb
BLAKE2b-256 a3a325ae8f2e5a1148174acc1ae5975397bad4016ac3f48e1ac13ae6664b0e43

See more details on using hashes here.

File details

Details for the file adversarial_friends-0.1.5-py3-none-any.whl.

File metadata

File hashes

Hashes for adversarial_friends-0.1.5-py3-none-any.whl
Algorithm Hash digest
SHA256 e55c71249d0f86c879cbe5ff785fc47d98588067d99694ba9da44b29d9bd0dd4
MD5 7d6cf04f107b095723307a6b3966ac28
BLAKE2b-256 4fd85b2f1b7a5fab84ef45b1d60e4b88095cd5529ae4efb4333927b50d292179

See more details on using hashes here.

Release history Release notifications | RSS feed

0.7.0

2 files

0.6.2

2 files

0.6.1

2 files

0.5.0

2 files

0.4.1

2 files

0.4.0

2 files

0.3.0

2 files

0.2.1

2 files

0.2.0

2 files

0.1.8

2 files

0.1.7

2 files

0.1.6

2 files

This release

0.1.5 This release

2 files

0.1.3

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page