Skip to main content

uacp-interop

Interop conformance probe for wippa-uacp.

UACP's own L0 to L3 conformance levels are a feature checklist. Passing proves an implementation matches its author's reading of the spec, and says nothing about whether two implementations built by different people interoperate, which is the premise the project rests on.

This harness tests the three things the checklist cannot:

  1. Schema versus prose. Build a message the spec requires an implementation to accept, validate it against the shipped schema, report the rejection.
  2. Schema versus the implementations. Check that each reference implementation can represent what the schema declares, and that the two implementations model the same envelope. This class found a silent credential downgrade between the TypeScript and Python implementations.
  3. Cross-format divergence, both directions. Map a framework-native agent description into a UACP descriptor, validate the artifact actually produced, and separately test whether a UACP capability survives being expressed in a framework's native tool format.

Results

See FINDINGS.md. Against wippa-uacp@b3fe7a1f: 2 schema conflicts and 12 cross-format findings (3 blocker, 7 major, 4 minor).

UACP itself is at zero blockers: its schemas now agree with its normative text, and the TypeScript and Python reference implementations model the same envelope. All three remaining blockers are measured loss in a foreign format, which is what the harness is for — a UACP capability using const and oneOf that an OpenAI function-calling block cannot express, and a corpus entry whose agent id violates UACP's identifier pattern.

Every blocker this harness once reported against its own protocol has been fixed upstream and is retained here as a regression guard, so a regression re-opens it rather than passing silently. See Fixed upstream.

Install

uvx uacp-interop                       # or: pipx run uacp-interop

Nothing to clone and no virtualenv to build. From a checkout:

pip install -e .                       # runtime only
uacp-interop                           # the report

Quick start

uacp-interop                            # scorecard, then the full report
uacp-interop --scorecard                # the one-glance per-section table
uacp-interop --markdown                 # full report as Markdown, for an issue or a PR
uacp-interop --json                     # machine-readable
uacp-interop --badge                    # status-badge JSON for the worst UACP finding
uacp-interop --fail-on-blocker          # exit 1 if any blocker is found
uacp-interop --corpus my-format.json    # measure a format nobody has contributed

pip install -e ".[dev]" && pytest -q    # only to run the tests

Every renderer derives its counts from the same findings via report.tally(), and a test asserts they cannot disagree — a renderer that tallied severities independently would be the double-counting defect from 0.1.0 in a new disguise. --badge deliberately reports only the spec and implementation audit: a blocker in a foreign format is a finding about that format, and reddening the badge for it would misattribute the fault to UACP.

Measuring a format that is not in the package

A corpus entry is pure data and the checks read nothing else, so any format can be measured from a JSON file:

uacp-interop --corpus my-format.json

It gets its own report section, computed by the same checks as the bundled entries. --corpus takes a directory too, and is repeatable. The loader rejects a missing or misspelt confidence rather than defaulting it, because a finding with no declared provenance would look as trustworthy as an observed one. See CONTRIBUTING.md.

The scorecard, which is the shareable view:

uacp-interop 0.2.0

  section                  findings  blocker  major  minor
  -----------------------  --------  -------  -----  -----
  uacp-spec                       2        0      2      0
  autogen-agentchat               6        0      4      2
  openai-function-calling         4        1      1      2
  uacp-native                     2        2      0      0
  -----------------------  --------  -------  -----  -----
  TOTAL                          14        3      7      4

A captured run of the current corpus is committed at docs/sample-report.txt, so you can see the output before installing anything.

Method

Losses are computed, not asserted. SCHEMA_VIOLATION findings are observed by running the real validator over the artifact an adapter produced, and it is the only source of them: the adapters do not predict schema violations from hand-copied patterns, because a copied pattern is a second source of truth that can drift from the schema, and because predicting as well as observing counted each such defect twice. Divergences in the reverse direction are found by walking the source JSON Schema and collecting keywords outside the target format's supported subset.

Each field yields at most one finding. When both reference implementations disagree the same way about a field, that is one defect with two witnesses and both are cited in the evidence line, not two findings. duplicate_findings() enforces this over the whole report, and the suite asserts it holds for the real corpus and that the guard is able to see a duplicate when one is injected.

Fixed upstream

The harness was written against its own protocol, so its first reports were mostly defects in UACP. All of them are fixed, and each is now a test:

Was Defect Fixed in
2 blockers metadata.auth modelled in TypeScript, absent in Python — an authenticated message decoded silently as anonymous wippa-uacp@b3fe7a1f
3 majors capability in the schema's global required, so heartbeat, register and discover had to invent a placeholder. Both implementations sent the literal _internal wippa-uacp@b3fe7a1f
1 major flowControl specified in SPEC.md and modelled by both implementations, declared by no schema wippa-uacp@b3fe7a1f
1 blocker audit-trail credential persistence in plaintext wippa-uacp@e0463b4

That sequence is the argument for the tool: it found real defects in the protocol it ships alongside, including a silent security downgrade, and the fixes are verified by the same harness that reported them.

Every corpus entry declares a confidence:

  • observed-serialization: a real serialized artifact was read
  • documented-api: taken from official documentation of the public API
  • inferred: our modelling choice, not a documented shape

Findings derived from weaker entries are weaker evidence, and the CLI says so. The AutoGen entry currently models a documented constructor surface rather than an observed serialization, which makes it the weakest evidence in the corpus.

Each check is backed by a meta-test that mutates a copy of the schema and confirms the check fires, so a check that quietly stopped working fails the suite rather than passing silently.

Vendored schemas

uacp_interop/schemas/ holds a copy of UACP's JSON Schemas so the harness can validate offline. A conformance report that silently tests an old protocol is worse than no report, so the upstream commit is pinned in uacp_interop/schemas/PROVENANCE.json and CI fails on drift:

python scripts/refresh_schemas.py            # refresh and re-pin
python scripts/refresh_schemas.py --check    # verify only, exit 1 on drift

Known limits

  • No behavioural half. Everything here is structural: "these two definitions cannot both be true", not "these two implementations produced different messages". That is the harder half and the one not yet built.
  • No LangGraph, CrewAI or OpenAI Assistants entries. The LangGraph docs did not yield a sourceable serialization format, and an entry that cannot be sourced honestly is worse than a missing one.
  • The implementation audit records observed field sets as data with provenance rather than parsing the dataclass and TypeScript interface at runtime. Re-reading them is a manual step, so they can drift.

License

MIT.

Release files for uacp-interop 0.4.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for uacp-interop 0.4.0
File Size Uploaded
uacp_interop-0.4.0.tar.gz 54.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for uacp-interop 0.4.0
File Interpreter ABI Platform
uacp_interop-0.4.0-py3-none-any.whl Python 3 none any Details

Total release size: 93.2 kB

Release files / uacp_interop-0.4.0.tar.gz

Download URL uacp_interop-0.4.0.tar.gz
Size 54.6 kB
Tags Source
SHA-256 checksum
How to use checksums
d6f18aba7733d3f1f473d6bb1f23d976f724d3692e9c81eccd98f9dc1e8869e7
BLAKE2b-256 checksum
How to use checksums
c6b5296fd4090bc3ae994b10936fd1dfa5718a9ef7df2e37bfbd0b32496f1518
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 27, 2026.

Transparency log

Release files / uacp_interop-0.4.0-py3-none-any.whl

Download URL uacp_interop-0.4.0-py3-none-any.whl
Size 38.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
020982458cca5326ce817588cb992b76fff47d13dd03103ca547fa13a426d09f
BLAKE2b-256 checksum
How to use checksums
c556bc1ca012fc9e1064798cb3ccbb1b69b27fbd72284d0eca0288bd4e67e8ab
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 27, 2026.

Transparency log

Release history Release notifications | RSS feed

0.17.0

2 release files

0.16.0

2 release files

0.15.0

2 release files

0.13.0

2 release files

0.12.0

2 release files

0.11.0

2 release files

0.10.0

2 release files

0.9.1

2 release files

0.9.0

2 release files

0.8.1

2 release files

0.8.0

2 release files

0.7.0

2 release files

0.6.0

2 release files

This release

0.4.0 This release

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page