uacp-interop
Interop conformance probe for wippa-uacp.
UACP's own L0 to L3 conformance levels are a feature checklist. Passing proves an implementation matches its author's reading of the spec, and says nothing about whether two implementations built by different people interoperate, which is the premise the project rests on.
This harness tests the three things the checklist cannot:
- Schema versus prose. Build a message the spec requires an implementation to accept, validate it against the shipped schema, report the rejection.
- Schema versus the implementations. Check that each reference implementation can represent what the schema declares, and that the two implementations model the same envelope. This class found a silent credential downgrade between the TypeScript and Python implementations.
- Cross-format divergence, both directions. Map a framework-native agent description into a UACP descriptor, validate the artifact actually produced, and separately test whether a UACP capability survives being expressed in a framework's native tool format.
Results
See FINDINGS.md. Against wippa-uacp@b3fe7a1f: 2 schema
conflicts and 12 cross-format findings (3 blocker, 7 major, 4 minor).
UACP itself is at zero blockers: its schemas now agree with its normative
text, and the TypeScript and Python reference implementations model the same
envelope. All three remaining blockers are measured loss in a foreign format,
which is what the harness is for — a UACP capability using const and oneOf
that an OpenAI function-calling block cannot express, and a corpus entry whose
agent id violates UACP's identifier pattern.
Every blocker this harness once reported against its own protocol has been fixed upstream and is retained here as a regression guard, so a regression re-opens it rather than passing silently. See Fixed upstream.
Install
uvx uacp-interop # or: pipx run uacp-interop
Nothing to clone and no virtualenv to build. From a checkout:
pip install -e . # runtime only
uacp-interop # the report
Quick start
uacp-interop # scorecard, then the full report
uacp-interop --scorecard # the one-glance per-section table
uacp-interop --markdown # full report as Markdown, for an issue or a PR
uacp-interop --json # machine-readable
uacp-interop --badge # status-badge JSON for the worst UACP finding
uacp-interop --fail-on-blocker # exit 1 if any blocker is found
uacp-interop --corpus my-format.json # measure a format nobody has contributed
pip install -e ".[dev]" && pytest -q # only to run the tests
Every renderer derives its counts from the same findings via report.tally(),
and a test asserts they cannot disagree — a renderer that tallied severities
independently would be the double-counting defect from 0.1.0 in a new disguise.
--badge deliberately reports only the spec and implementation audit: a
blocker in a foreign format is a finding about that format, and reddening the
badge for it would misattribute the fault to UACP.
Measuring a format that is not in the package
A corpus entry is pure data and the checks read nothing else, so any format can be measured from a JSON file:
uacp-interop --corpus my-format.json
It gets its own report section, computed by the same checks as the bundled
entries. --corpus takes a directory too, and is repeatable. The loader rejects
a missing or misspelt confidence rather than defaulting it, because a finding
with no declared provenance would look as trustworthy as an observed one. See
CONTRIBUTING.md.
The scorecard, which is the shareable view:
uacp-interop 0.2.0
section findings blocker major minor
----------------------- -------- ------- ----- -----
uacp-spec 2 0 2 0
autogen-agentchat 6 0 4 2
openai-function-calling 4 1 1 2
uacp-native 2 2 0 0
----------------------- -------- ------- ----- -----
TOTAL 14 3 7 4
A captured run of the current corpus is committed at docs/sample-report.txt, so you can see the output before installing anything.
Method
Losses are computed, not asserted. SCHEMA_VIOLATION findings are observed by
running the real validator over the artifact an adapter produced, and it is the
only source of them: the adapters do not predict schema violations from
hand-copied patterns, because a copied pattern is a second source of truth that
can drift from the schema, and because predicting as well as observing counted
each such defect twice. Divergences in the reverse direction are found by
walking the source JSON Schema and collecting keywords outside the target
format's supported subset.
Each field yields at most one finding. When both reference implementations
disagree the same way about a field, that is one defect with two witnesses and
both are cited in the evidence line, not two findings. duplicate_findings()
enforces this over the whole report, and the suite asserts it holds for the real
corpus and that the guard is able to see a duplicate when one is injected.
Fixed upstream
The harness was written against its own protocol, so its first reports were mostly defects in UACP. All of them are fixed, and each is now a test:
| Was | Defect | Fixed in |
|---|---|---|
| 2 blockers | metadata.auth modelled in TypeScript, absent in Python — an authenticated message decoded silently as anonymous |
wippa-uacp@b3fe7a1f |
| 3 majors | capability in the schema's global required, so heartbeat, register and discover had to invent a placeholder. Both implementations sent the literal _internal |
wippa-uacp@b3fe7a1f |
| 1 major | flowControl specified in SPEC.md and modelled by both implementations, declared by no schema |
wippa-uacp@b3fe7a1f |
| 1 blocker | audit-trail credential persistence in plaintext | wippa-uacp@e0463b4 |
That sequence is the argument for the tool: it found real defects in the protocol it ships alongside, including a silent security downgrade, and the fixes are verified by the same harness that reported them.
Every corpus entry declares a confidence:
observed-serialization: a real serialized artifact was readdocumented-api: taken from official documentation of the public APIinferred: our modelling choice, not a documented shape
Findings derived from weaker entries are weaker evidence, and the CLI says so. The AutoGen entry currently models a documented constructor surface rather than an observed serialization, which makes it the weakest evidence in the corpus.
Each check is backed by a meta-test that mutates a copy of the schema and confirms the check fires, so a check that quietly stopped working fails the suite rather than passing silently.
Vendored schemas
uacp_interop/schemas/ holds a copy of UACP's JSON Schemas so the harness can
validate offline. A conformance report that silently tests an old protocol is
worse than no report, so the upstream commit is pinned in
uacp_interop/schemas/PROVENANCE.json and CI fails on drift:
python scripts/refresh_schemas.py # refresh and re-pin
python scripts/refresh_schemas.py --check # verify only, exit 1 on drift
Known limits
- No behavioural half. Everything here is structural: "these two definitions cannot both be true", not "these two implementations produced different messages". That is the harder half and the one not yet built.
- No LangGraph, CrewAI or OpenAI Assistants entries. The LangGraph docs did not yield a sourceable serialization format, and an entry that cannot be sourced honestly is worse than a missing one.
- The implementation audit records observed field sets as data with provenance rather than parsing the dataclass and TypeScript interface at runtime. Re-reading them is a manual step, so they can drift.
License
MIT.
Release files for uacp-interop 0.4.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| uacp_interop-0.4.0.tar.gz | 54.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| uacp_interop-0.4.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 93.2 kB
Release files / uacp_interop-0.4.0.tar.gz
| Download URL | uacp_interop-0.4.0.tar.gz |
|---|---|
| Size | 54.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
d6f18aba7733d3f1f473d6bb1f23d976f724d3692e9c81eccd98f9dc1e8869e7
|
|
BLAKE2b-256 checksum How to use checksums |
c6b5296fd4090bc3ae994b10936fd1dfa5718a9ef7df2e37bfbd0b32496f1518
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 27, 2026.
Transparency logRelease files / uacp_interop-0.4.0-py3-none-any.whl
| Download URL | uacp_interop-0.4.0-py3-none-any.whl |
|---|---|
| Size | 38.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
020982458cca5326ce817588cb992b76fff47d13dd03103ca547fa13a426d09f
|
|
BLAKE2b-256 checksum How to use checksums |
c556bc1ca012fc9e1064798cb3ccbb1b69b27fbd72284d0eca0288bd4e67e8ab
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 27, 2026.
Transparency log