Skip to main content

agent-loss-map

PyPI Python CI

What does your agent definition lose when it crosses a format boundary?

agent-loss-map report

Every agent framework invented its own idea of what a tool is, and the differences are invisible until an agent fails in production. This measures them. It projects a format-native tool definition into another format, runs the artifact it actually produced, and reports every piece of information that did not survive — with evidence, and a confidence tier saying how much the finding is worth.

uvx agent-loss-map --matrix

No install, no clone, no API key. The bundled run makes no network call at all — --matrix, --chain and --behavioural are computed entirely from the shipped corpus and the vendored schemas. 16 formats in 7 classes, 208 crossings in the matrix, 325 tests. The whole matrix computes in about 400 ms — though that is lopsided: the Stripe entry alone is 361 ms of it, because its 20-property schemas are walked against fifteen dialects. The other twelve sources together are 48 ms. (--mcp is the one flag that reaches the network, because its job is to measure a server you have not read yet.)


Some formats cannot hold a tool contract at all

A2A's protocol data model has no JSON Schema in it. message AgentSkill in the normative source — specification/a2a.proto at tag v1.0.1 — has eight fields: id, name, description, tags, examples, input_modes, output_modes, security_requirements. The words schema and json_schema do not appear anywhere in that file's 811 lines. A skill is described by prose and by media-type strings.

So a real MCP server crossing into A2A loses its entire structured contract, and the reverse direction cannot recover what was never there:

$ agent-loss-map --from deepwiki --to a2a

  source    findings  blocker  major  minor
  ---------  --------  -------  -----  -----
  deepwiki         6        6      0      0

Six blockers, zero majors, zero minors — because there is no degraded path to report. Every schema is destroyed rather than downgraded, and a bridge that papers over this by inventing an inputSchema has written a private convention and called it a protocol.

This is a property of the format, not a criticism of it: A2A is opaque-by-design about internals, which is a legitimate goal for agent-to-agent exchange. It does mean any inputSchema a bridge attaches is a private convention, not part of the protocol. Section 1.4 of specification.md makes the proto the single authoritative definition and describes the generated JSON artifact as a non-normative build artifact.


The formats are not one kind of thing

formats what a schema crossing does
can carry a schema 7 — mcp, openai-fc, anthropic, gemini, bedrock-converse, bedrock-agent, langchain the schema crosses, minus whatever keywords the target's dialect refuses
contract in pieces, no join rule 2 — openapi, stripe the parameters exist; nothing specifies how to flatten them into one object
a different type system 2 — shopify, langgraph GraphQL arguments and state channels are not a JSON Schema dialect
media types only 1 — a2a nothing survives
taxonomy references 1 — oasf skills are {name, id} vocabulary refs, not contracts
prose 1 — agent-skills five frontmatter fields and Markdown
no machine-readable artifact 2 — anp, emvco nothing to project into

Seven classes, and the split is the finding rather than a category list: two formats that look interchangeable in a README destroy and downlevel the same schema by completely different mechanisms. Where a format's inputs are several located declarations with no join rule — OpenAPI's parameters[], and Stripe's — no specification says how to flatten them into one object, and three real implementations produce three incompatible answers, so the harness reports the ambiguity rather than inventing a mapping and calling the crossing lossless. DESIGN.md has the registry, the guard that refuses a self-contradictory record, and OpenAPI/Stripe in full.


Multi-hop loss is not additive, and no single crossing shows it

A single crossing answers "what does this source lose entering that format". A bridge asks a third question: what survives the whole path. --chain walks it, feeding each hop the artifact the previous hop actually produced:

agent-loss-map --from deepwiki --chain mcp,a2a,anthropic,openai-fc
hop format findings cumulative intact lost
1 mcp 0 100% 0
2 a2a 6 blockers 46% 7
3 anthropic 0 46% 7
4 openai-fc 3 blockers 46% 7

Cumulative intact does not rise at hop 3 or hop 4: the fields were already gone, and no later hop invented any back. Two things in that table are worth more than the whole of --matrix.

Hop 3 reporting zero findings is not a clean hop. Anthropic has a schema field, and it reported nothing at all — because A2A's AgentSkill left it nothing to report about. Summing per-hop counts renders that as a clean second hop. It was a hop that ran on an empty document.

No pair of single-hop crossings reveals hop 4's blockers. Three blockers appear only at the end of the chain; the same two hops measured independently produce nothing. And they land on a {"type": "object"} the chain itself fabricated at hop 3, because A2A had already destroyed the real contracts — so Anthropic was handed nothing and emitted an empty schema, and the target's own rule that a parameters block must carry properties then refused it. The finding is real and the attribution belongs to the chain, not to Anthropic or OpenAI. That is the compound failure this whole feature exists to name, and a per-crossing matrix scores it clean.

The intermediate is this harness's own re-read, not a client's parse: hop N is fed the artifact hop N−1 actually produced, converted back into a capability list by agent_loss_map.chain.reread. intact means the artifact carries that value, not that any client took it. The chain report prints that caveat with the result and --json carries it in the payload. DESIGN.md covers field fate, terminal states, and the re-read's known distortions.


The executed half: third-party code, actually run

Everything above reasons. --behavioural does not — it imports client libraries, calls their own code with the artifacts this harness produced, and writes down what came back. Its vocabulary has five words and none of them is a pass by another name: accepted, REFUSED, altered, DIVERGED, NOT RUN.

Only three formats have a client library to run, and you have to name one of them. openai-fc, anthropic and mcp are bound to the vendors' own SDKs. The other 13 report NOT RUN and say why — and the reason is not always "no library": some have no client binding, and some have no input contract to execute at all.

The default run measures against UACP, which has no client library, so:

$ agent-loss-map --behavioural
  accepted=0  refused=0  altered=0  diverged=0  not run=5

Nothing executed, at exit 0, because that is an honest result and not a crash. To see real client code run:

agent-loss-map --from mcp --to mcp --behavioural        # all four probes
agent-loss-map --matrix --behavioural                    # every crossing

Three decisions keep it from inflating the result:

  • It is a separate type with no severity and no code, so it cannot enter the finding totals, the scorecard, the badge or the gate. An executed result is not a Divergence; concatenating the two raises, and a test asserts it. The default run, --matrix, --chain and --gate are byte-identical without --behavioural.
  • NOT RUN is never a pass. A machine that cannot run a check and a machine whose check passed would otherwise produce identical bytes. clean, pass, passed, ok and success are not in the vocabulary and a test asserts they cannot get in.
  • A local type check is not a service acceptance. Every executed result says which of the two it is, and the test suite asserts the wording.

--behavioural is refused with --chain (exit 2) rather than ignored. What each of the four probes runs, which callable, and which pinned version: DESIGN.md.


Who this is for

Bridge and gateway authors. If you are converting between any of these formats, your bridge is losing information at the boundary and the usual answer is a hand-written lossy dict. This computes it — and for a multi-hop path, it computes the compounding the dict cannot show.

Anyone choosing a format. The matrix is the comparison: which targets a source crosses into intact, and which destroy it. A format that can hold a schema and one that cannot look identical in a README and behave nothing alike.

Protocol and spec authors. If your spec has normative text and JSON Schemas, this machine-checks whether they agree with each other — a class of defect a feature checklist cannot see, because a checklist proves your code matches your own reading of the spec, not that your reading and your code agree. It found eleven such defects in our own protocol: eight in the schema-versus-prose ledger, one security control that was correct, unit-tested, and never called, and two naming-rule incompatibilities found by measuring servers we did not build. Every one is fixed upstream and retained as a regression guard. DESIGN.md has the ledger, its arithmetic, and the defect-by-defect account.

Framework maintainers. If your framework publishes agent definitions, this tells you what a consumer in another format will not be able to represent.


Install

uvx agent-loss-map                       # or: pipx run agent-loss-map
pip install -e .                       # from a checkout
agent-loss-map

--behavioural needs three client libraries, which are an optional extra rather than a default dependency:

pip install 'agent-loss-map[behavioural]'

The client libraries are pinned exactly (openai==1.109.1, anthropic==1.11.0, mcp==1.20.0) rather than with a floor, because every executed result names the installed version in its evidence — "the client's types accept this" is worth nothing without knowing whose types they are.

--mcp measures a live server: the harness performs a real handshake, projects tools/list faithfully, refuses to report anything if its own projection dropped a field, and reports the era it spoke. Three MCP captures ship with the package, so agent-loss-map --from mcp --to openai-fc works with nothing else installed.

Flag
--matrix every source against every target in one table
--chain F,F,F measure a multi-hop path, feeding each hop the previous hop's artifact
--behavioural also run the executed half against real client-library code
--scorecard the per-section table on its own
--markdown the full report as Markdown, for an issue or a PR
--json machine-readable
--badge status-badge JSON for the worst finding against the spec
--mcp URL capture tools/list from a live server and measure it now
--from SOURCE only measure sources matching a framework id or format family
--to TARGET project into a target format and validate the artifact produced
--corpus PATH measure a format that is not bundled
--fail-on-blocker exit 1 if any blocker is found
--gate exit 1 only on a blocker against UACP itself (the release gate)

Modes answer different questions and refuse to be combined: --chain rejects --to, --matrix and --behavioural with exit 2 and says why. --gate is not evaluated on a chain run, because a chain cannot answer whether UACP's own schemas agree with its own normative text, and answering it "clean" would be a false green.

A report that silently tests an old protocol is worse than no report, so two workflows guard that: CI fails a build on real schema drift and on the test suite, while a scheduled probe re-runs the measurement and opens a PR when the committed report, badge or scorecard differ. Staleness becomes visible in review rather than breaking a build — DESIGN.md has both workflows and the pinned vendored-schema provenance.


Add your format

A corpus entry is pure data and the checks read nothing else, so measuring a format needs no Python and no pull request:

agent-loss-map --corpus my-format.json
{
  "framework": "my-framework",
  "agent_id": "my_agent",
  "description": "What this is.",
  "confidence": "documented",
  "provenance": "Where you read the shape, and when",
  "capabilities": [
    { "name": "web.search", "description": "Search the web.",
      "parameters": { "type": "object",
                      "properties": { "query": { "type": "string" } } } }
  ]
}

It gets its own section in the report, computed by the same checks as everything else. --corpus also takes a directory and is repeatable. Every entry declares how well-sourced it is, and the loader refuses to guess — a missing or misspelt confidence is an error, not a silent downgrade:

Tier Meaning
observed-serialization we read a real serialized artifact
documented-api from official documentation of the public API
inferred our modelling choice, not a documented shape

The corpus is 13 entries: 5 observed-serialization (three MCP captures, a UACP descriptor quoted from its spec, and the Stripe subset derived by script) and 8 documented-api. The default run reports 44 findings — 4 spec-audit (2 major, 2 minor) and 40 cross-format (1 blocker, 13 major, 26 minor). Today's run against wippa-uacp reports zero blockers in the spec audit; the run's one blocker is loss measured in a foreign format, which is what the harness is for. Captured in docs/sample-report.txt.

To make a format permanent, append a ForeignAgent to agent_loss_map/corpus.py — about twenty minutes. CONTRIBUTING.md has the recipe and the provenance rules; DESIGN.md has the registry guards a new format must satisfy.


Known limits

  • The security claims are still guarded by a source probe, not by execution, and that is the limit that matters most. The unwired authorization was found by a person reading the code and is now guarded by a regex over the routing path. --behavioural does not change that: no probe runs UACP's own Bus, because the package must work without a checkout of the protocol repo. A source probe can be satisfied by code that is never called, and this one is still open to that.
  • A chain is fed forward through this harness's own re-read, not a client's parse. --chain mcp,openai-fc,a2a measures the path, but hop 2 consumes the artifact hop 1 actually produced, converted back into a capability list by agent_loss_map/chain.py. No OpenAI server, A2A client or Anthropic API was involved, and a real client may refuse a shape the re-read accepts. intact means the artifact carries that value, not that a client took it. The re-read's known distortions are enumerated per hop and printed with the chain, and the intermediate is labelled inferred because a real system did not serialize it — this package did. --matrix still computes each crossing independently; a chain is the third question, not a substitute for the other two.
  • Per-hop finding counts in a chain are not additive, and the report will not pretend they are. The headline is the cumulative share of the source's own declared fields the artifact still carries, and a hop that reported nothing because it had nothing left to lose is called out by name.
  • The executed half is real, and much smaller than the static one. --behavioural covers 3 formats with a maintained Python client; the other 13 report NOT RUN and say why. unavailable is never rendered as a pass, and clean/pass are not statuses the vocabulary can produce. A local type check is not a service acceptance: when the openai model's pydantic validation accepts a function block, no request was built, no key was needed, and nothing is known about what the API would do.
  • A format with no published keyword subset cannot produce a keyword finding. mcp, langchain, openapi, shopify and stripe declare no accepted subset, so a clean cell for them means "the rules they publish were satisfied", not "fully supported". A constraint on what a format accepts is not a thing it has documented.
  • Two records rest on mirror-sourced or disputed facts, and say so. The Anthropic keyword subset was read from a mirror of the structured-outputs limitations page, so it accepts rather than reports the keywords in doubt; the Anthropic and Gemini name-length limits are disputed between two official surfaces, so no length finding is reported for either.
  • The bundled MCP captures are not all the same protocol era, and each one says which it is. context7 is 2026-07-28 — modern, captured over server/discover, with per-request _meta and no session. deepwiki-mcp and adoraads-beauty are 2025-06-18 — legacy, captured over an initialize handshake. Every capture was probed for server/discover and labelled by what the server answered, not inferred from the corpus. The Tool data type — what the harness actually measures — has not changed in the ways that matter, but the two eras negotiate the transport differently, and a legacy client against a modern server fails on the current revision's own compatibility matrix.
  • A standards-compliant validator still cannot resolve the UACP schemas offline. The $ids claim canonical URLs on a host that does not resolve, so loading agent-descriptor and following its relative capability.schema.json ref attempts a network call. This harness resolves those refs locally by filename as well as by $id, and a test proves the resolution is offline rather than incidentally working. The finding stays major because the defect is in the published artifact, not in the workaround.
  • An absent protocol checkout makes the code checks unverified, not passing. Set UACP_SOURCE_ROOT to a checkout to have them run. Silently skipping them and reporting clean would be the failure this harness exists to catch.
  • Release notes are per version, in docs/. 0.21.0 is the executed half's self-comparison, six keyword absences, and unreachable version negotiation. 0.20.0 closes an SSRF in --behavioural — if you are on 0.19.0 or earlier, upgrade; agent-loss-map --version. 0.19.0 adds --chain and --behavioural.
  • FINDINGS.md covers the spec audit, not the whole run. It explains what the schema-versus-prose and schema-versus-implementation checks found in wippa-uacp, which is the durable part. For current counts across all 13 corpus sources, read docs/sample-report.txt, which the scheduled probe regenerates from a live run.

Verifying it

pip install -e ".[dev]"
pytest -q

325 tests, 1 skipped. The suite is mostly meta-tests: each mutates a copy of the input so the defect it looks for is absent, and asserts the check goes quiet. A check that quietly stopped working fails the suite rather than passing silently — which has twice caught a guard in this repo that had quietly stopped guarding anything, both times because relaxing a rule invalidated the fixture it used to trigger on.

MIT.

Metadata

Release files for agent-loss-map 0.21.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for agent-loss-map 0.21.0
File Size Uploaded
agent_loss_map-0.21.0.tar.gz 333.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for agent-loss-map 0.21.0
File Interpreter ABI Platform
agent_loss_map-0.21.0-py3-none-any.whl Python 3 none any Details

Total release size: 541.5 kB

Release files / agent_loss_map-0.21.0.tar.gz

Download URL agent_loss_map-0.21.0.tar.gz
Size 333.6 kB
Tags Source
SHA-256 checksum
How to use checksums
d111cec80d94ffe5fbe50c086c6b2b8f3ac7e0b73f31810c4e6ecfe7fb5955f5
BLAKE2b-256 checksum
How to use checksums
86e22356ddb07b5caa588b684b88378b328b7e4fb00a3b77c714dc096bfaa177
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.

Transparency log

Release files / agent_loss_map-0.21.0-py3-none-any.whl

Download URL agent_loss_map-0.21.0-py3-none-any.whl
Size 207.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
3cbe3528ee899dc647d784e5cccef7913dcecf0ebfd76c740a2dfa20a2ec8be3
BLAKE2b-256 checksum
How to use checksums
283b2a8eb7fdffa11fb5294e78549f548e3935d5ada763a2c2d4d88fc11e2b12
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.21.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page