Skip to main content

uacp-interop

uacp spec PyPI Python

What does your agent definition lose when it crosses a format boundary?

uacp-interop report

Every agent framework invented its own idea of what a tool is. MCP has tools/list. OpenAI function-calling has a parameters block with a narrow schema subset. UACP has a capability schema that uses JSON Schema properly. The differences between them are invisible until an agent fails in production, and nobody has tooling for measuring them.

This measures them. It takes a format-native agent description, runs the resulting artifact through a real JSON Schema validator, and reports every piece of information that does not survive — with evidence, and with a confidence tier saying how much the finding is worth.

uvx uacp-interop

No install, no clone, no API key. A run takes under 200 ms (median 139 ms over 20 runs).

It found eight real defects in its own protocol

We built this to grade wippa-uacp, the protocol we wrote. The first run reported five blockers. All five were real, and all five are fixed:

Was Defect Status
🔴 2 blockers metadata.auth existed in the TypeScript model and not the Python one, so an authenticated message decoded as anonymous, with no error fixed
🟠 3 majors capability was in the schema's global required list, so heartbeat and register had to invent a placeholder — and both implementations sent the literal "_internal" fixed
🟠 1 major flowControl was specified in the prose, modelled by both implementations, and declared by no schema fixed
🟠 1 major async defaulted to true, so a plain request/response capability was documented as streaming — the consumer waits for a stream.end that never comes, and sees a hang, not an error fixed
🔴 1 blocker the audit trail persisted bearer tokens in plaintext fixed

Seven of those are against the protocol; the eighth is in our own reference implementation, which is the more useful direction to be wrong in.

It also found a defect in the naming rule, by measuring two live MCP servers rather than by reading the schema. UACP rejects hyphens in capability names, so resolve-library-id and query-docs — both accepted by OpenAI's own function-name grammar — cannot be named without renaming, and a renamed tool is a different tool to any client that refers to it by name. The two grammars are not subsets of each other, and that runs both ways: UACP's required dot-namespace is rejected by OpenAI too. Remedy in PR.

It found a security control that was never wired

The most useful thing it has caught, and the one nothing else here could have.

SPEC.md said a bus "can enforce capability allow-lists per caller". Both implementations shipped a correct check_authorization, and both had it unit-tested. Bus.send never called it. A caller the policy denied reached the protected capability, and the app's own check returned denied while the message went through anyway.

A schema-versus-prose comparison cannot see this: both were fine, because the prose under-committed. "Can" reads as a capability, not an obligation. Neither can a cross-language field diff. Only reading the routing path can.

So check_spec_claims_are_implemented does exactly that, and reports a normative claim the implementation does not honour at UNVERIFIABLE_CLAIM. It is pinned by a test that flips the recorded evidence and asserts the check fires, because a check that cannot fail is decoration. The claim now reads as enforced, so the check is quiet — wippa-uacp#6 has the fix.

That is the part worth caring about. A conformance checklist proves your implementation matches your own reading of a spec. It says nothing about whether your reading and your code agree with each other, or with anyone else's. That is what this measures instead, and every finding it produced here was one a checklist would have passed.

What it actually checks

  1. Schema versus prose. Build a message the spec requires an implementation to accept, validate it against the shipped schema, report the rejection.
  2. Schema versus the implementations. Check each reference implementation can represent what the schema declares, and that they model the same envelope.
  3. Cross-format divergence, both directions. Map a framework-native agent description into a target descriptor and validate the artifact actually produced; separately, test whether a capability survives being expressed in a format's native tool schema.
  4. Target-format rules JSON Schema cannot express. An OpenAI-compatible parameters block must carry properties, so a capability publishing {"type": "object"} gets the whole tool list rejected with a 400. The capability schema accepts that shape perfectly well, so the constraint only exists in the target format — which is the same class as a schema keyword with no representation there.

Losses are computed, not asserted. SCHEMA_VIOLATION findings come from running the real validator over the artifact an adapter produced — a hand-copied regex predicting the same thing was removed in 0.1.0 because it double-counted every finding. Divergences in the other direction are found by walking the source JSON Schema and collecting keywords outside the target format's supported subset.

Every finding carries evidence and a confidence tier

This is the part most tooling in this space skips, and it is the reason to trust the output:

Tier Meaning
observed-serialization we read a real serialized artifact
documented-api from official documentation of the public API
inferred our modelling choice, not a documented shape

The report prints the tier for every entry and says which findings rest on weaker evidence. The bundled AutoGen entry models a documented constructor surface rather than an observed serialization, so it is labelled as the weakest in the corpus. An entry that cannot be sourced honestly is worse than a missing one, and the loader refuses to guess: a missing or misspelt confidence is an error, not a silent downgrade to the weakest tier.

Add your format

A worked example, captured from a live server rather than synthesised:

uvx uacp-interop --corpus examples/deepwiki-mcp.entry.json

examples/deepwiki-mcp.capture.json and examples/context7.capture.json are raw tools/list results from two live servers — DeepWiki 2.14.3 and Context7 4.1.1 — each taken over a real MCP handshake, with the matching .entry.json a projection of it.

scripts/capture_mcp.py does both halves: it performs the handshake and then projects the result, and it runs the fidelity check before writing an entry, so a projection that disagrees with its own artifact is refused rather than published. That check exists because an outputSchema was dropped in silence during exactly this conversion, by hand, and the harness then reported the omission as a major finding on all three tools — as if the server had left it out. It had not. A field carried from a real artifact must either appear in the entry or be declared in dropped_source_fields, and a carried field must equal the source value.

git clone https://github.com/wippa-studios/uacp-interop && cd uacp-interop
python scripts/capture_mcp.py https://mcp.deepwiki.com/mcp --out deepwiki-mcp

The scripts are repository tooling rather than installed console entry points, so run them from a clone.

A corpus entry is pure data and the checks read nothing else, so measuring a format needs no Python and no pull request:

uacp-interop --corpus my-format.json
{
  "framework": "my-framework",
  "agent_id": "my_agent",
  "description": "What this is.",
  "confidence": "documented",
  "provenance": "Where you read the shape, and when",
  "capabilities": [
    { "name": "web.search", "description": "Search the web.",
      "parameters": { "type": "object",
                      "properties": { "query": { "type": "string" } } } }
  ]
}

It gets its own section in the report, computed by the same checks as everything else. --corpus also takes a directory and is repeatable. To make it permanent, append a ForeignAgent to uacp_interop/corpus.py — about twenty minutes, and CONTRIBUTING.md has the recipe and the provenance rules.

Install

uvx uacp-interop                       # or: pipx run uacp-interop
pip install -e .                       # from a checkout
uacp-interop
Flag
--scorecard the per-section table on its own
--markdown the full report as Markdown, for an issue or a PR
--json machine-readable
--badge status-badge JSON for the worst finding against the spec
--corpus PATH measure a format that is not bundled
--fail-on-blocker exit 1 if any blocker is found

Every renderer derives its counts from the same findings, and a test asserts they cannot disagree — a renderer that tallied severities independently is the 0.1.0 double-counting bug in a new disguise. --badge reports only the spec and implementation audit, because a blocker in a foreign format is a finding about that format and reddening the badge would misattribute the fault.

It cannot go stale quietly

The report is only worth anything if it is current, so:

  • the vendored schemas are pinned to an upstream commit in uacp_interop/schemas/PROVENANCE.json, and CI fails on drift
  • a scheduled workflow re-runs the probe weekly and opens an issue if the result stops matching the committed artifact
  • docs/scorecard.svg in this README is generated from a live run, and CI fails if it goes stale

A conformance report that silently tests an old protocol is worse than no report.

Known limits

  • One external $ref needs a local registry to resolve. The schemas claim canonical URLs on a host that does not resolve, so a plain offline validator fetching agent-descriptor will attempt a network call. This harness builds a referencing.Registry for exactly that reason, and the finding is reported rather than hidden.
  • No behavioural half, still. Almost everything here is structural: these two definitions cannot both be true, not these two implementations produced different messages. The spec-claims check is the one exception — it reads the routing path rather than a data structure, and it is what caught the unwired authorization. Extending that to actually executing an implementation's behaviour is not built.
  • The corpus is small, because the honest entry is one you can source. Three ship; a LangGraph entry is open.
  • Not yet format-to-format. It measures loss crossing into a target descriptor and loss expressing one in a native tool schema, not directly between two foreign formats. That is the obvious next thing and it is not built.

Verifying it

pip install -e ".[dev]"
pytest -q

The suite is mostly meta-tests: each mutates a copy of the input so the defect it looks for is absent, and asserts the check goes quiet. A check that quietly stopped working fails the suite rather than passing silently — which has twice caught a guard in this repo that had quietly stopped guarding anything.

MIT.

Release files for uacp-interop 0.15.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for uacp-interop 0.15.0
File Size Uploaded
uacp_interop-0.15.0.tar.gz 86.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for uacp-interop 0.15.0
File Interpreter ABI Platform
uacp_interop-0.15.0-py3-none-any.whl Python 3 none any Details

Total release size: 130.2 kB

Release files / uacp_interop-0.15.0.tar.gz

Download URL uacp_interop-0.15.0.tar.gz
Size 86.3 kB
Tags Source
SHA-256 checksum
How to use checksums
56429f809d399e40cee4f5f3e3c3239fe8d025ccffaaf94e483811c89c662a20
BLAKE2b-256 checksum
How to use checksums
d2351f6baad1178d4c6482821eed2ffb075c79dfcf742a1ec581faccc293b084
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 27, 2026.

Transparency log

Release files / uacp_interop-0.15.0-py3-none-any.whl

Download URL uacp_interop-0.15.0-py3-none-any.whl
Size 43.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
07024b343a63b12b0f4e245391d7e1dae83903a68380315d8ad9967fb8cb6adc
BLAKE2b-256 checksum
How to use checksums
a5bbe8d6e64e08f0d5747262bdb3781a599e57849545f5dc09a336ab6971bf3b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 27, 2026.

Transparency log

Release history Release notifications | RSS feed

0.17.0

2 release files

0.16.0

2 release files

This release

0.15.0 This release

2 release files

0.13.0

2 release files

0.12.0

2 release files

0.11.0

2 release files

0.10.0

2 release files

0.9.1

2 release files

0.9.0

2 release files

0.8.1

2 release files

0.8.0

2 release files

0.7.0

2 release files

0.6.0

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page