Skip to main content

hostile-facilitator

Retry-Safety

This battery is the conformance suite of the draft MCP retry-safety proposal (SEP working draft, discussion #3188) — ported into its tools/call fixture and used to verify the reference implementation, where it caught two real double-executions before scoring it 7/7.

Free to be listed. Submit any implementation for grading — yours or someone else's — and it gets read and put on the public Retry-Safety Index at no cost. If a finding is confirmed you are counted, never named, until you ship a fix; when you do, the row goes up with credit and your time-to-fix. Only a full battery run against a live system is paid, and nothing above requires it.

An AI agent that pays twice for one order is a refund storm, a chargeback, and a trust problem — and it happens on a dropped connection, not a bug you'd catch in review. This tells you in 60 seconds whether your agent does it.

Here's the trap. A payment settles on-chain, and then the connection drops — a timeout, a 502, whatever. Your client reads that as "failed," retries, and sends a fresh payment. Both go through. Your customer paid twice, and every log on your side shows one clean payment after one transient error. Nobody notices until the refunds start.

The happy path and the clean-failure path both get tested. The settled-but-looks-failed path almost never does — because you need a facilitator that misbehaves on cue. This is that facilitator: it does the worst-moment things real ones do, and counts how many times your agent actually paid for one order. One is safe. Two is the money you're about to lose.

hostile-facilitator is the adversary. It stands in for the facilitator, deliberately produces each ambiguous failure, and — because every settle passes through it — counts how many distinct payments your client actually made for one purchase. One is safe. Two is a real double-charge, caught.

No keys. No chain. No real money. Just your client's retry behaviour, which is where the bug lives.

Sign the result

test --json writes the run as data, so it can be signed and re-checked instead of trusted:

hostile-facilitator test --json result.json -- ./make-one-purchase.sh
coherence conformance result.json --out session.json
coherence attest --session session.json --key <key> --anchor rekor

A mode where the client settled twice is recorded as unfinished work, never as a pass — so a failing run cannot be signed as clean, by you or by us. Worked example, including the attempt to launder a failure into a pass: coherence/examples/conformance.

60 seconds

pip install hostile-facilitator
# or straight from source:
# pip install "git+https://github.com/aurumflux20/hostile-facilitator@v0.2.2"

# prove the instrument is honest (catches a broken client, clears a safe one):
hostile-facilitator selftest

# test YOUR client in one command — it runs the whole battery for you.
# Give it a command that makes ONE purchase and reads the facilitator URL
# from an env var (default FACILITATOR_URL):
hostile-facilitator test -- your-client --pay-once
#   → 10/10 safe, or a FAIL row per ambiguous failure your client double-pays on.

# or drive it by hand against one failure mode:
hostile-facilitator serve --mode accept_then_timeout

Proof on a real chain

The battery above models settlement in memory — fast and honest, but a finding written from it carries a caveat: not a live reproduction. proof removes the caveat. It starts a local anvil node, deploys an EIP-3009 token with the same transferWithAuthorization / authorizationState surface real USDC exposes, and settles for real — counting payments from Transfer logs on the chain rather than from its own bookkeeping.

hostile-facilitator proof     # needs foundry: curl -L https://foundry.paradigm.xyz | bash && foundryup
  hostile-facilitator - ON-CHAIN proof (payments counted from Transfer logs)

    naive (known-broken): 3/7 safe
      [FAIL] accept_then_timeout    2 REAL transfers for one purchase - DOUBLE PAY
      [FAIL] 5xx_after_settle       2 REAL transfers for one purchase - DOUBLE PAY
      [FAIL] double_402             2 REAL transfers for one purchase - DOUBLE PAY
      [FAIL] reconcile_unavailable  2 REAL transfers for one purchase - DOUBLE PAY

    safe (known-correct): 10/10 safe

  instrument valid on-chain: True

Two transfers for one purchase is no longer an argument about control flow: it is two entries in a ledger anyone can re-read. tests/test_chain.py asserts both directions and skips itself when foundry is absent.

The failure modes

Each one leaves the world in the same true state — the payment settled — and hands your client a signal that's easy to misread as "it failed, try again":

mode what it does
accept_then_timeout settles, then hangs past your client's timeout
5xx_after_settle settles, then returns 502
double_402 re-challenges a request that already paid
slow_answer settles, answers just under the wire
reconcile_unavailable settles ambiguously, and the "did it land?" read also fails — loudly, with a 503
reconcile_soft_404 the read answers HTTP 200 with prose: "not available at this time"
reconcile_oversized the record is in the body, past a buffer a naive read won't survive
reconcile_truncated the body is cut mid-record, so the parse yields nothing
declared_safe the tool declares replay is safe — checks you're not over-refusing
clean control: settles, answers 200

declared_safe matters because a retry gate fails in two directions. Everything else here asks "did you fire twice?" — it asks "did you refuse work that was safe?" Bricking legitimate work to avoid an impossible duplicate is a real bug, not a conservative virtue.

The four reconcile_* modes ask the hardest question: when the effect may have landed and the read that would tell you is broken, do you hold? "Could not determine" is terminal; a client that reads it as "didn't happen" and retries has reintroduced the exact double-pay the read exists to prevent.

Only the first of those four announces itself. The other three are the half that actually ships, and they come from solim on the SEP discussion, out of two production bugs: a grep -q gate whose SIGPIPE became exit 141 under pipefail and was read as "no match" — but only once the body grew past the 64KB pipe buffer, so a present record reported as absent precisely when there was more evidence to read; and a grep -c counter, where "genuinely zero" and "the command failed" arrive as the same value with the same status. Neither is exotic. Both are the default idiom in the language they were written in.

The generalisation, and the rule the battery now tests: a reconcile read must be able to return three values — settled, not settled, and read failed — and "read failed" must not be collapsed into "not settled" by the client's own plumbing. That defect lives below the protocol, which is why nothing upstream catches it.

The instrument check

reconcile_controls() exercises your reconciliation checker against a case that must answer settled and one that must answer not-settled. solim's rule: a checker validated only against the negative is indistinguishable from a function that returns a constant — which is what both of theirs were. Until a checker has been shown capable of both answers, every "absent" it has ever returned is unevidenced.

Holding is not the same as proving

Over HTTP the reconcile modes are scored with a third state, not_exercised. A client that takes the ambiguous 504 and simply gives up has not double-paid — but it also never performed the read, so the run proves nothing about its read path. Reporting that as "safe" would be a verdict the evidence cannot support, which is the same false-pass shape this battery exists to catch. It is reported as its own state instead, and never counted as a pass.

A run that never purchased is not a score

test runs the clean control first. If one ordinary purchase does not settle exactly once, the command is not buying through the hostile facilitator — a wrong env var, a seller that rejected our replies — and "no double payment" in every other mode would be true for the wrong reason. The run stops there, prints HARNESS INVALID, writes "verdict": "INVALID" to --json, and exits 2.

Testing a real x402 seller

The facilitator speaks the x402 v2 wire format: GET /supported, and /verify / /settle replies the official SDK models parse (success, transaction, network, payer). The battery's timeout is exported as HOSTILE_CLIENT_TIMEOUT; set your seller's facilitator timeout from it, or the timeout modes never time out (the x402 Python SDK defaults to 90s).

A complete harness for the official Python SDK — FastAPI seller, httpx buyer, throwaway key, no chain — is in examples/x402-python/pay_once.py:

pip install "x402[fastapi,evm,clients]" uvicorn
hostile-facilitator test --client-timeout 3 -- python examples/x402-python/pay_once.py

How it recognises a payment

Your client's stable payment identity — an EIP-3009 authorization.nonce, a top-level nonce, or an Idempotency-Key header — is how the facilitator knows a re-presented payment from a brand-new one. A client that sends no stable identity can't be safe under retries, and the tool says so.

The fix, when it fails

Treat an ambiguous outcome as unknown, never as failed. On retry, re-present the same authorization; let the facilitator settle it once. Don't mint a fresh nonce for a payment you already sent.

Gate every PR (GitHub Action)

Keep a client retry-safe forever: drop this into .github/workflows/retry-safety.yml and every PR that makes your payment client double-pay fails the check.

name: Retry-Safety
on: [pull_request]
jobs:
  check:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with: { python-version: "3.11" }
      - uses: aurumflux20/hostile-facilitator@v0.2.2
        with:
          client-command: "python scripts/pay_once.py"   # your one-purchase client

It posts a sticky PR comment with the scorecard and fails the build on a double-pay. Earn the badge for your README once it's green:

[![Retry-Safety](https://img.shields.io/badge/retry--safety-checked-2ea44f?logo=shieldsdotio)](https://github.com/aurumflux20/hostile-facilitator)

Conformance (x402 §5.3.6)

§5.3.6 of the x402 settlement-status amendment — proposed text in an open PR (zjzJoez/x402#1 toward x402-foundation/x402#3325), not yet ratified — makes a conformance claim checkable instead of declarative: name the battery you ran, count settlements actually recorded rather than response bodies, and carry a mutation control — evidence the same battery fails against an implementation known to be unsafe. The clause that does the work: a self-administered pass reported without a control is a declaration, not a verification. (Committed in zjzJoez/x402#1 against x402-foundation/x402#3325; under review, not yet ratified.)

This battery satisfies the first two requirements and ships its own mutation control (hostile-facilitator selftest), so you can run it yourself and see exactly where you stand — free, MIT, no signup. That part never costs anything.

What you cannot self-issue is the third-party half. We run the battery against your live facilitator or client and issue a signed conformance result you can publish: findings within five business days with any failing case reproduced in full, and a clean run signed within 24 hours.

Self-run, signed — $99, first ten servers (checkout). You run this battery against your own server and send us result.json and session.json; we sign and anchor the result in Sigstore so a reader can verify it without trusting either of us. No call, no scheduling. A failing mode is recorded open, never proven — a failing run cannot be signed green by anyone, including us. This is the cheapest rung and the only one that needs nothing from us but a signature.

Founding rate — the first three implementations to carry a public result: $300 (checkout), on one condition: the result is published on the Retry-Safety Index, because a verification nobody can check isn't one. Standard rate after the founding three is $1,200.

If it fails, and you want the whole path checked

This tool tests one purchase against the ambiguous-failure battery. It won't tell you whether the rest of your money path holds — the reservation lifecycle, every settle-timeout branch, or whether what you believe you spent matches what actually settled.

Two ways to take it further, both written-only, no calls:

  • Attestation run — $1,200. We run the full battery against your live endpoint and issue a signed result you can publish. Findings within five business days; a clean run signed within 24 hours. Checkout
  • Money-path review — $12,000, fixed scope. Every path that moves or counts money, each one graded, the unsafe ones with a failing case that reproduces it — file and line, 7–10 days. If we can't show you a real double-fire on a path you actually run, there is no invoice. Details

It's the same reading that found the bugs behind this tool — four projects have shipped fixes from it, including a company running 1M+ paid API calls a month.


MIT © AurumFlux AI, Inc — part of the Seal work on exactly-once for agents that move money.

Metadata

Release files for hostile-facilitator 0.2.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for hostile-facilitator 0.2.2
File Size Uploaded
hostile_facilitator-0.2.2.tar.gz 58.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for hostile-facilitator 0.2.2
File Interpreter ABI Platform
hostile_facilitator-0.2.2-py3-none-any.whl Python 3 none any Details

Total release size: 96.9 kB

Release files / hostile_facilitator-0.2.2.tar.gz

Download URL hostile_facilitator-0.2.2.tar.gz
Size 58.5 kB
Tags Source
SHA-256 checksum
How to use checksums
8cf6ec1ee1d7c15637a9bd391ee676cffee8eb69f3cf732d6368eb4f84f2099b
BLAKE2b-256 checksum
How to use checksums
e36d8d6793f1fe799bd320b1553a770bfa50647d2d24d58f0a7ef373ad792a41
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.

Transparency log

Release files / hostile_facilitator-0.2.2-py3-none-any.whl

Download URL hostile_facilitator-0.2.2-py3-none-any.whl
Size 38.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
961b36bb146ca086a9d68c2aa83aa6c664b5e98f5153a8db5d27d009305cc5fd
BLAKE2b-256 checksum
How to use checksums
ff545e41755e9adcaf435b9a10d6bb0387d01937a525ecae3c28eb413f539eec
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.2.2 This release

2 release files

0.2.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page