uacp-interop
What does your agent definition lose when it crosses a format boundary?
Every agent framework invented its own idea of what a tool is. MCP has
tools/list. OpenAI function-calling has a parameters block with a narrow
schema subset. UACP has a capability schema that uses JSON Schema properly. The
differences between them are invisible until an agent fails in production, and
nobody has tooling for measuring them.
This measures them. It takes a format-native agent description, runs the resulting artifact through a real JSON Schema validator, and reports every piece of information that does not survive — with evidence, and with a confidence tier saying how much the finding is worth.
uvx uacp-interop
No install, no clone, no API key. A run takes under 200 ms (median 139 ms over 20 runs).
It found eleven real defects in its own protocol
We built this to grade wippa-uacp, the protocol we wrote. The first run reported five blockers. All five were real, and all five are fixed:
| Was | Defect | Status |
|---|---|---|
| 🔴 2 blockers | metadata.auth existed in the TypeScript model and not the Python one, so an authenticated message decoded as anonymous, with no error |
fixed |
| 🟠 3 majors | capability was in the schema's global required list, so heartbeat and register had to invent a placeholder — and both implementations sent the literal "_internal" |
fixed |
| 🟠 1 major | flowControl was specified in the prose, modelled by both implementations, and declared by no schema |
fixed |
| 🟠 1 major | async defaulted to true, so a plain request/response capability was documented as streaming — the consumer waits for a stream.end that never comes, and sees a hang, not an error |
fixed |
| 🔴 1 blocker | the audit trail persisted bearer tokens in plaintext | fixed |
Seven of those are against the protocol; the eighth is in our own reference implementation, which is the more useful direction to be wrong in. Three more came from measuring live servers rather than reading the schema, and are written up below because they are a different kind of defect: two are naming rules that made real tool and server names unrepresentable, the other a compatibility failure that no amount of self-consistent testing would have found.
It also found a defect in the naming rule, by measuring two live MCP servers
rather than by reading the schema. UACP rejects hyphens in capability names, so
resolve-library-id and query-docs — both accepted by OpenAI's own
function-name grammar — cannot be named without renaming, and a renamed tool is
a different tool to any client that refers to it by name. The two grammars are
not subsets of each other, and that runs both ways: UACP's required dot-namespace
is rejected by OpenAI too. Remedy in PR.
The same measurement found the identical rule applied to agent ids, which is
what turned the corpus entry's own tool-caller into a reported blocker. The
agent id pattern was the same ^[a-zA-Z0-9_]{1,128}$, so a server announcing
inside-ads or adoraads-beauty-gateway could not describe itself at all — and
here the corpus entry was right and the protocol was wrong.
Remedy in PR.
It found a security control that was never wired
The most useful thing it has caught, and the one nothing else here could have.
SPEC.md said a bus "can enforce capability allow-lists per caller". Both
implementations shipped a correct check_authorization, and both had it
unit-tested. Bus.send never called it. A caller the policy denied reached the
protected capability, and the app's own check returned denied while the message
went through anyway.
A schema-versus-prose comparison cannot see this: both were fine, because the prose under-committed. "Can" reads as a capability, not an obligation. Neither can a cross-language field diff. Only reading the code can.
It found it by a person reading the code, and the check that now guards it is
narrower than that. It probes the source for the call — (path, regex) per
claim, verdict derived — so it catches the wiring being removed, in either
language. It does not execute the bus, and it is not behavioural testing. An
earlier version recorded enforced: true by hand, which meant it was asserting
the author's own reading back at him; had the call been deleted later it would
have reported clean forever, and the green test count would have been a true
statement about nothing. A source probe is strictly better than that and still not
the same thing as running the code.
So check_spec_claims_are_implemented does exactly that, and reports a normative
claim the implementation does not honour at UNVERIFIABLE_CLAIM. It is pinned by
a test that flips the recorded evidence and asserts the check fires, because a
check that cannot fail is decoration. The claim now reads as enforced, so the
check is quiet — wippa-uacp#6
has the fix.
That is the part worth caring about. A conformance checklist proves your implementation matches your own reading of a spec. It says nothing about whether your reading and your code agree with each other, or with anyone else's. That is what this measures instead, and every finding it produced here was one a checklist would have passed.
What it actually checks
- Schema versus prose. Build a message the spec requires an implementation to accept, validate it against the shipped schema, report the rejection.
- Schema versus the implementations. Check each reference implementation can represent what the schema declares, and that they model the same envelope.
- Cross-format divergence, both directions. Map a framework-native agent description into a target descriptor and validate the artifact actually produced; separately, test whether a capability survives being expressed in a format's native tool schema.
- Target-format rules JSON Schema cannot express. An
OpenAI-compatible
parametersblock must carryproperties, so a capability publishing{"type": "object"}gets the whole tool list rejected with a 400. The capability schema accepts that shape perfectly well, so the constraint only exists in the target format — which is the same class as a schema keyword with no representation there.
Losses are computed, not asserted. SCHEMA_VIOLATION findings come from
running the real validator over the artifact an adapter produced — a hand-copied
regex predicting the same thing was removed in 0.1.0 because it double-counted
every finding. Divergences in the other direction are found by walking the source
JSON Schema and collecting keywords outside the target format's supported
subset.
Every finding carries evidence and a confidence tier
This is the part most tooling in this space skips, and it is the reason to trust the output:
| Tier | Meaning |
|---|---|
observed-serialization |
we read a real serialized artifact |
documented-api |
from official documentation of the public API |
inferred |
our modelling choice, not a documented shape |
The report prints the tier for every entry and says which findings rest on
weaker evidence. The bundled AutoGen entry models a documented constructor
surface rather than an observed serialization, so it is labelled as the weakest
in the corpus. An entry that cannot be sourced honestly is worse than a missing
one, and the loader refuses to guess: a missing or misspelt confidence is an
error, not a silent downgrade to the weakest tier.
Add your format
A worked example, captured from a live server rather than synthesised:
uvx uacp-interop --corpus examples/deepwiki-mcp.entry.json
examples/deepwiki-mcp.capture.json and examples/context7.capture.json are raw
tools/list results from two live servers — DeepWiki 2.14.3 and Context7 4.1.1 —
each taken over a real MCP handshake, with the matching .entry.json a
projection of it.
scripts/capture_mcp.py does both halves: it performs the handshake and then
projects the result, and it runs the fidelity check before writing an entry, so a
projection that disagrees with its own artifact is refused rather than published.
That check exists because an outputSchema was dropped in silence during
exactly this conversion, by hand, and the harness then reported the omission as a
major finding on all three tools — as if the server had left it out. It had not.
A field carried from a real artifact must either appear in the entry or be
declared in dropped_source_fields, and a carried field must equal the source
value.
git clone https://github.com/wippa-studios/uacp-interop && cd uacp-interop
python scripts/capture_mcp.py https://mcp.deepwiki.com/mcp --out deepwiki-mcp
The scripts are repository tooling rather than installed console entry points, so run them from a clone.
A corpus entry is pure data and the checks read nothing else, so measuring a format needs no Python and no pull request:
uacp-interop --corpus my-format.json
{
"framework": "my-framework",
"agent_id": "my_agent",
"description": "What this is.",
"confidence": "documented",
"provenance": "Where you read the shape, and when",
"capabilities": [
{ "name": "web.search", "description": "Search the web.",
"parameters": { "type": "object",
"properties": { "query": { "type": "string" } } } }
]
}
It gets its own section in the report, computed by the same checks as everything
else. --corpus also takes a directory and is repeatable. To make it permanent,
append a ForeignAgent to uacp_interop/corpus.py — about twenty minutes, and
CONTRIBUTING.md has the recipe and the provenance rules.
Install
uvx uacp-interop # or: pipx run uacp-interop
pip install -e . # from a checkout
uacp-interop
| Flag | |
|---|---|
--scorecard |
the per-section table on its own |
--markdown |
the full report as Markdown, for an issue or a PR |
--json |
machine-readable |
--badge |
status-badge JSON for the worst finding against the spec |
--corpus PATH |
measure a format that is not bundled |
--fail-on-blocker |
exit 1 if any blocker is found |
Every renderer derives its counts from the same findings, and a test asserts they
cannot disagree — a renderer that tallied severities independently is the
0.1.0 double-counting bug in a new disguise. --badge reports only the spec
and implementation audit, because a blocker in a foreign format is a finding
about that format and reddening the badge would misattribute the fault.
It cannot go stale quietly
The report is only worth anything if it is current, so:
- the vendored schemas are pinned to an upstream commit in
uacp_interop/schemas/PROVENANCE.json, and CI fails on drift - a scheduled workflow re-runs the probe weekly and opens an issue if the result stops matching the committed artifact
docs/scorecard.svgin this README is generated from a live run, and CI fails if it goes stale
A conformance report that silently tests an old protocol is worse than no report.
Known limits
- One external
$refneeds a local registry to resolve. The schemas claim canonical URLs on a host that does not resolve, so a plain offline validator fetchingagent-descriptorwill attempt a network call. This harness builds areferencing.Registryfor exactly that reason, and the finding is reported rather than hidden. - No behavioural half. Everything here is static: these two definitions cannot both be true, and, for the security claims, this code path does or does not call this function. Nothing executes an implementation. The claim that found the unwired authorization was a person reading the code; the check that guards it afterwards is a source probe, not a behavioural test, and the difference matters if you are deciding how much to trust a clean report.
- An absent protocol checkout makes the code checks unverified, not passing.
Set
UACP_SOURCE_ROOTto a checkout to have them run. Silently skipping them and reporting clean would be the failure this harness exists to catch. - The corpus is small, because the honest entry is one you can source. Three ship; a LangGraph entry is open.
- Not yet format-to-format. It measures loss crossing into a target descriptor and loss expressing one in a native tool schema, not directly between two foreign formats. That is the obvious next thing and it is not built.
A compatibility bug, which is a different kind
Six live MCP servers were measured. Two of them announce a name containing a
hyphen — inside-ads and adoraads-beauty-gateway — and under UACP's agent-id
rule neither could supply a conformant id at all. The protocol could not describe
a third of the ecosystem it exists to interoperate with.
The rule had been tightened deliberately, on the argument that an agent is an identity rather than a namespace. That argument does not support the conclusion: the separator is the dot, not the hyphen, and nothing technical turned on excluding it. It was tidied-up symmetry with the capability-name rule, which is the specific failure the test guarding that rule predicted and could not prevent.
This one is worth separating from the rest. The other eight were defects in something this project wrote. This was an incompatibility with other people's conventions, discovered only by measuring them, and it would not have shown up in any amount of self-consistent testing.
The gap that is not a bug
The spec is silent where it should not be, and no implementation can be blamed.
Two conforming agents, each invoking the other's declared capability, ran at roughly 20,000 round trips per second. The only bound was the host language's recursion limit — tightening it from 1000 to 200 scaled the round trips from 124 to 24, proportionally — and the initiating caller received an ordinary success response with nothing indicating the pipeline had been cut short.
That silence is the finding. It has the same shape as the streaming-default defect, where a consumer observed a hang rather than an error: in both cases the caller cannot distinguish a completed run from a truncated one. Rate limiting is specified as a MAY, scoped per-agent, and nothing bounds the sum across a pipeline.
wippa-uacp §15.8 records it as unspecified, and this harness reports it as a
major UNVERIFIABLE_CLAIM that stays reported even against a perfect
implementation — it is marked a gap, so a future change cannot quietly promote it
to "implemented" and close it by assertion.
Verifying it
pip install -e ".[dev]"
pytest -q
96 tests. The suite is mostly meta-tests: each mutates a copy of the input so the defect it looks for is absent, and asserts the check goes quiet. A check that quietly stopped working fails the suite rather than passing silently — which has twice caught a guard in this repo that had quietly stopped guarding anything, both times because relaxing a rule invalidated the fixture it used to trigger on.
MIT.
Release files for uacp-interop 0.17.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| uacp_interop-0.17.0.tar.gz | 89.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| uacp_interop-0.17.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 135.5 kB
Release files / uacp_interop-0.17.0.tar.gz
| Download URL | uacp_interop-0.17.0.tar.gz |
|---|---|
| Size | 89.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
134260263f4b29b1556f20586e91e687205788b9ed16571ebfacf9777db44421
|
|
BLAKE2b-256 checksum How to use checksums |
dfa738039cdccb36f966e22ce677cf9a5dffd48ecb1efb030b63c7c964fee898
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.
Transparency logRelease files / uacp_interop-0.17.0-py3-none-any.whl
| Download URL | uacp_interop-0.17.0-py3-none-any.whl |
|---|---|
| Size | 45.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
431a83cd6f7c4d63d63609b76ca39f88c4f1fd561afb9d33e958a8e9b635e17f
|
|
BLAKE2b-256 checksum How to use checksums |
b79e21ea362be28cc7c7db10f443414db011014e79222ca18c8e1ae4de432558
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.
Transparency log