hostile-facilitator
This battery is the conformance suite of the draft MCP retry-safety proposal (SEP working draft, discussion #3188) — ported into its
tools/callfixture and used to verify the reference implementation, where it caught two real double-executions before scoring it 7/7.
Free to be listed. Submit any implementation for grading — yours or someone else's — and it gets read and put on the public Retry-Safety Index at no cost. If a finding is confirmed you are counted, never named, until you ship a fix; when you do, the row goes up with credit and your time-to-fix. Only a full battery run against a live system is paid, and nothing above requires it.
An AI agent that pays twice for one order is a refund storm, a chargeback, and a trust problem — and it happens on a dropped connection, not a bug you'd catch in review. This tells you in 60 seconds whether your agent does it.
Here's the trap. A payment settles on-chain, and then the connection drops — a timeout, a 502, whatever. Your client reads that as "failed," retries, and sends a fresh payment. Both go through. Your customer paid twice, and every log on your side shows one clean payment after one transient error. Nobody notices until the refunds start.
The happy path and the clean-failure path both get tested. The settled-but-looks-failed path almost never does — because you need a facilitator that misbehaves on cue. This is that facilitator: it does the worst-moment things real ones do, and counts how many times your agent actually paid for one order. One is safe. Two is the money you're about to lose.
hostile-facilitator is the adversary. It stands in for the facilitator, deliberately produces each ambiguous failure, and — because every settle passes through it — counts how many distinct payments your client actually made for one purchase. One is safe. Two is a real double-charge, caught.
No keys. No chain. No real money. Just your client's retry behaviour, which is where the bug lives.
Sign the result
test --json writes the run as data, so it can be signed and re-checked
instead of trusted:
hostile-facilitator test --json result.json -- ./make-one-purchase.sh
coherence conformance result.json --out session.json
coherence attest --session session.json --key <key> --anchor rekor
A mode where the client settled twice is recorded as unfinished work, never as a pass — so a failing run cannot be signed as clean, by you or by us. Worked example, including the attempt to launder a failure into a pass: coherence/examples/conformance.
60 seconds
pip install hostile-facilitator
# or straight from source:
# pip install "git+https://github.com/aurumflux20/hostile-facilitator@v0.2.2"
# prove the instrument is honest (catches a broken client, clears a safe one):
hostile-facilitator selftest
# test YOUR client in one command — it runs the whole battery for you.
# Give it a command that makes ONE purchase and reads the facilitator URL
# from an env var (default FACILITATOR_URL):
hostile-facilitator test -- your-client --pay-once
# → 10/10 safe, or a FAIL row per ambiguous failure your client double-pays on.
# or drive it by hand against one failure mode:
hostile-facilitator serve --mode accept_then_timeout
Proof on a real chain
The battery above models settlement in memory — fast and honest, but a finding
written from it carries a caveat: not a live reproduction. proof removes the
caveat. It starts a local anvil node, deploys an EIP-3009
token with the same transferWithAuthorization / authorizationState surface real
USDC exposes, and settles for real — counting payments from Transfer logs on the
chain rather than from its own bookkeeping.
hostile-facilitator proof # needs foundry: curl -L https://foundry.paradigm.xyz | bash && foundryup
hostile-facilitator - ON-CHAIN proof (payments counted from Transfer logs)
naive (known-broken): 3/7 safe
[FAIL] accept_then_timeout 2 REAL transfers for one purchase - DOUBLE PAY
[FAIL] 5xx_after_settle 2 REAL transfers for one purchase - DOUBLE PAY
[FAIL] double_402 2 REAL transfers for one purchase - DOUBLE PAY
[FAIL] reconcile_unavailable 2 REAL transfers for one purchase - DOUBLE PAY
safe (known-correct): 10/10 safe
instrument valid on-chain: True
Two transfers for one purchase is no longer an argument about control flow: it is
two entries in a ledger anyone can re-read. tests/test_chain.py asserts both
directions and skips itself when foundry is absent.
The failure modes
Each one leaves the world in the same true state — the payment settled — and hands your client a signal that's easy to misread as "it failed, try again":
| mode | what it does |
|---|---|
accept_then_timeout |
settles, then hangs past your client's timeout |
5xx_after_settle |
settles, then returns 502 |
double_402 |
re-challenges a request that already paid |
slow_answer |
settles, answers just under the wire |
reconcile_unavailable |
settles ambiguously, and the "did it land?" read also fails — loudly, with a 503 |
reconcile_soft_404 |
the read answers HTTP 200 with prose: "not available at this time" |
reconcile_oversized |
the record is in the body, past a buffer a naive read won't survive |
reconcile_truncated |
the body is cut mid-record, so the parse yields nothing |
declared_safe |
the tool declares replay is safe — checks you're not over-refusing |
clean |
control: settles, answers 200 |
declared_safe matters because a retry gate fails in two directions. Everything
else here asks "did you fire twice?" — it asks "did you refuse work that was
safe?" Bricking legitimate work to avoid an impossible duplicate is a real bug,
not a conservative virtue.
The four reconcile_* modes ask the hardest question: when the effect may have
landed and the read that would tell you is broken, do you hold? "Could not
determine" is terminal; a client that reads it as "didn't happen" and retries has
reintroduced the exact double-pay the read exists to prevent.
Only the first of those four announces itself. The other three are the half that
actually ships, and they come from solim
on the SEP discussion, out of two production bugs: a grep -q gate whose SIGPIPE
became exit 141 under pipefail and was read as "no match" — but only once the
body grew past the 64KB pipe buffer, so a present record reported as absent
precisely when there was more evidence to read; and a grep -c counter, where
"genuinely zero" and "the command failed" arrive as the same value with the same
status. Neither is exotic. Both are the default idiom in the language they were
written in.
The generalisation, and the rule the battery now tests: a reconcile read must be able to return three values — settled, not settled, and read failed — and "read failed" must not be collapsed into "not settled" by the client's own plumbing. That defect lives below the protocol, which is why nothing upstream catches it.
The instrument check
reconcile_controls() exercises your reconciliation checker against a case that
must answer settled and one that must answer not-settled. solim's rule:
a checker validated only against the negative is indistinguishable from a
function that returns a constant — which is what both of theirs were. Until a
checker has been shown capable of both answers, every "absent" it has ever
returned is unevidenced.
Holding is not the same as proving
Over HTTP the reconcile modes are scored with a third state, not_exercised. A
client that takes the ambiguous 504 and simply gives up has not double-paid — but
it also never performed the read, so the run proves nothing about its read path.
Reporting that as "safe" would be a verdict the evidence cannot support, which is
the same false-pass shape this battery exists to catch. It is reported as its own
state instead, and never counted as a pass.
A run that never purchased is not a score
test runs the clean control first. If one ordinary purchase does not settle
exactly once, the command is not buying through the hostile facilitator — a
wrong env var, a seller that rejected our replies — and "no double payment" in
every other mode would be true for the wrong reason. The run stops there,
prints HARNESS INVALID, writes "verdict": "INVALID" to --json, and exits 2.
Testing a real x402 seller
The facilitator speaks the x402 v2 wire format: GET /supported, and
/verify / /settle replies the official SDK models parse (success,
transaction, network, payer). The battery's timeout is exported as
HOSTILE_CLIENT_TIMEOUT; set your seller's facilitator timeout from it, or the
timeout modes never time out (the x402 Python SDK defaults to 90s).
A complete harness for the official Python SDK — FastAPI seller, httpx buyer,
throwaway key, no chain — is in
examples/x402-python/pay_once.py:
pip install "x402[fastapi,evm,clients]" uvicorn
hostile-facilitator test --client-timeout 3 -- python examples/x402-python/pay_once.py
How it recognises a payment
Your client's stable payment identity — an EIP-3009 authorization.nonce, a top-level nonce, or an Idempotency-Key header — is how the facilitator knows a re-presented payment from a brand-new one. A client that sends no stable identity can't be safe under retries, and the tool says so.
The fix, when it fails
Treat an ambiguous outcome as unknown, never as failed. On retry, re-present the same authorization; let the facilitator settle it once. Don't mint a fresh nonce for a payment you already sent.
Gate every PR (GitHub Action)
Keep a client retry-safe forever: drop this into .github/workflows/retry-safety.yml
and every PR that makes your payment client double-pay fails the check.
name: Retry-Safety
on: [pull_request]
jobs:
check:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with: { python-version: "3.11" }
- uses: aurumflux20/hostile-facilitator@v0.2.2
with:
client-command: "python scripts/pay_once.py" # your one-purchase client
It posts a sticky PR comment with the scorecard and fails the build on a double-pay. Earn the badge for your README once it's green:
[](https://github.com/aurumflux20/hostile-facilitator)
Conformance (x402 §5.3.6)
§5.3.6 of the x402 settlement-status amendment — proposed text in an open PR (zjzJoez/x402#1 toward x402-foundation/x402#3325), not yet ratified — makes a conformance claim checkable instead of declarative: name the battery you ran, count settlements actually recorded rather than response bodies, and carry a mutation control — evidence the same battery fails against an implementation known to be unsafe. The clause that does the work: a self-administered pass reported without a control is a declaration, not a verification. (Committed in zjzJoez/x402#1 against x402-foundation/x402#3325; under review, not yet ratified.)
This battery satisfies the first two requirements and ships its own mutation control (hostile-facilitator selftest), so you can run it yourself and see exactly where you stand — free, MIT, no signup. That part never costs anything.
What you cannot self-issue is the third-party half. We run the battery against your live facilitator or client and issue a signed conformance result you can publish: findings within five business days with any failing case reproduced in full, and a clean run signed within 24 hours.
Self-run, signed — $99, first ten servers (checkout). You run this battery against your own server and send us result.json and session.json; we sign and anchor the result in Sigstore so a reader can verify it without trusting either of us. No call, no scheduling. A failing mode is recorded open, never proven — a failing run cannot be signed green by anyone, including us. This is the cheapest rung and the only one that needs nothing from us but a signature.
Founding rate — the first three implementations to carry a public result: $300 (checkout), on one condition: the result is published on the Retry-Safety Index, because a verification nobody can check isn't one. Standard rate after the founding three is $1,200.
If it fails, and you want the whole path checked
This tool tests one purchase against the ambiguous-failure battery. It won't tell you whether the rest of your money path holds — the reservation lifecycle, every settle-timeout branch, or whether what you believe you spent matches what actually settled.
Two ways to take it further, both written-only, no calls:
- Attestation run — $1,200. We run the full battery against your live endpoint and issue a signed result you can publish. Findings within five business days; a clean run signed within 24 hours. Checkout
- Money-path review — $12,000, fixed scope. Every path that moves or counts money, each one graded, the unsafe ones with a failing case that reproduces it — file and line, 7–10 days. If we can't show you a real double-fire on a path you actually run, there is no invoice. Details
It's the same reading that found the bugs behind this tool — four projects have shipped fixes from it, including a company running 1M+ paid API calls a month.
MIT © AurumFlux AI, Inc — part of the Seal work on exactly-once for agents that move money.
Metadata
Release files for hostile-facilitator 0.2.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| hostile_facilitator-0.2.2.tar.gz | 58.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| hostile_facilitator-0.2.2-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 96.9 kB
Release files / hostile_facilitator-0.2.2.tar.gz
| Download URL | hostile_facilitator-0.2.2.tar.gz |
|---|---|
| Size | 58.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
8cf6ec1ee1d7c15637a9bd391ee676cffee8eb69f3cf732d6368eb4f84f2099b
|
|
BLAKE2b-256 checksum How to use checksums |
e36d8d6793f1fe799bd320b1553a770bfa50647d2d24d58f0a7ef373ad792a41
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.
Transparency logRelease files / hostile_facilitator-0.2.2-py3-none-any.whl
| Download URL | hostile_facilitator-0.2.2-py3-none-any.whl |
|---|---|
| Size | 38.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
961b36bb146ca086a9d68c2aa83aa6c664b5e98f5153a8db5d27d009305cc5fd
|
|
BLAKE2b-256 checksum How to use checksums |
ff545e41755e9adcaf435b9a10d6bb0387d01937a525ecae3c28eb413f539eec
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.
Transparency log