Skip to main content

agent-loss-map

PyPI Python CI

What does your agent definition lose when it crosses a format boundary?

agent-loss-map report

Every agent framework invented its own idea of what a tool is, and the differences are invisible until an agent fails in production. This measures them. It projects a format-native tool definition into another format, runs the artifact it actually produced, and reports every piece of information that did not survive — with evidence, and a confidence tier saying how much the finding is worth.

uvx agent-loss-map --matrix

No install, no clone, no API key. The bundled run makes no network call at all — --matrix, --chain and --behavioural are computed entirely from the shipped corpus and the vendored schemas. 16 formats in 7 classes, 208 crossings in the matrix (about 400 ms for all of them, ~37 ms for one), 316 tests. (--mcp is the one flag that reaches the network, because its job is to measure a server you have not read yet.)


Some formats cannot hold a tool contract at all

A2A's protocol data model has no JSON Schema in it. message AgentSkill in the normative source — specification/a2a.proto at tag v1.0.1 — has eight fields: id, name, description, tags, examples, input_modes, output_modes, security_requirements. The words schema and json_schema do not appear anywhere in that file's 811 lines. A skill is described by prose and by media-type strings.

So a real MCP server crossing into A2A loses its entire structured contract, and the reverse direction cannot recover what was never there:

$ agent-loss-map --from deepwiki --to a2a

  source    findings  blocker  major  minor
  ---------  --------  -------  -----  -----
  deepwiki         6        6      0      0

Six blockers, zero majors, zero minors — because there is no degraded path to report. Every schema is destroyed rather than downgraded, and a bridge that papers over this by inventing an inputSchema has written a private convention and called it a protocol.

This is a property of the format, not a criticism of it: A2A is opaque-by-design about internals, which is a legitimate goal for agent-to-agent exchange. It does mean any inputSchema a bridge attaches is a private convention, not part of the protocol. Section 1.4 of specification.md makes the proto the single authoritative definition and describes the generated JSON artifact as a non-normative build artifact.


The formats are not one kind of thing

formats what a schema crossing does
can carry a schema 7 — mcp, openai-fc, anthropic, gemini, bedrock-converse, bedrock-agent, langchain the schema crosses, minus whatever keywords the target's dialect refuses
contract in pieces, no join rule 2 — openapi, stripe the parameters exist; nothing specifies how to flatten them into one object
a different type system 2 — shopify, langgraph GraphQL arguments and state channels are not a JSON Schema dialect
media types only 1 — a2a nothing survives
taxonomy references 1 — oasf skills are {name, id} vocabulary refs, not contracts
prose 1 — agent-skills five frontmatter fields and Markdown
no machine-readable artifact 2 — anp, emvco nothing to project into

Seven classes, and the split is the finding rather than a category list: two formats that look interchangeable in a README destroy and downlevel the same schema by completely different mechanisms. Where a format's inputs are several located declarations with no join rule — OpenAPI's parameters[], and Stripe's — no specification says how to flatten them into one object, and three real implementations produce three incompatible answers, so the harness reports the ambiguity rather than inventing a mapping and calling the crossing lossless. DESIGN.md has the registry, the guard that refuses a self-contradictory record, and OpenAPI/Stripe in full.


Multi-hop loss is not additive, and no single crossing shows it

A single crossing answers "what does this source lose entering that format". A bridge asks a third question: what survives the whole path. --chain walks it, feeding each hop the artifact the previous hop actually produced:

agent-loss-map --from deepwiki --chain mcp,a2a,anthropic,openai-fc
hop format findings cumulative intact lost
1 mcp 0 100% 0
2 a2a 6 blockers 46% 7
3 anthropic 0 46% 7
4 openai-fc 3 blockers 46% 7

Cumulative intact does not rise at hop 3 or hop 4: the fields were already gone, and no later hop invented any back. Two things in that table are worth more than the whole of --matrix.

Hop 3 reporting zero findings is not a clean hop. Anthropic has a schema field, and it reported nothing at all — because A2A's AgentSkill left it nothing to report about. Summing per-hop counts renders that as a clean second hop. It was a hop that ran on an empty document.

No pair of single-hop crossings reveals hop 4's blockers. Three blockers appear only at the end of the chain; the same two hops measured independently produce nothing. And they land on a {"type": "object"} the chain itself fabricated at hop 3, because A2A had already destroyed the real contracts — so Anthropic was handed nothing and emitted an empty schema, and the target's own rule that a parameters block must carry properties then refused it. The finding is real and the attribution belongs to the chain, not to Anthropic or OpenAI. That is the compound failure this whole feature exists to name, and a per-crossing matrix scores it clean.

The intermediate is this harness's own re-read, not a client's parse: hop N is fed the artifact hop N−1 actually produced, converted back into a capability list by agent_loss_map.chain.reread. intact means the artifact carries that value, not that any client took it. The chain report prints that caveat with the result and --json carries it in the payload. DESIGN.md covers field fate, terminal states, and the re-read's known distortions.


The executed half: third-party code, actually run

Everything above reasons. --behavioural does not — it imports client libraries, calls their own code with the artifacts this harness produced, and writes down what came back. Its vocabulary has five words and none of them is a pass by another name: accepted, REFUSED, altered, DIVERGED, NOT RUN. Three formats have a maintained Python client: openai-fc, anthropic, mcp. The other 13 report NOT RUN and say why.

Three decisions keep it from inflating the result:

  • It is a separate type with no severity and no code, so it cannot enter the finding totals, the scorecard, the badge or the gate. An executed result is not a Divergence; concatenating the two raises, and a test asserts it. The default run, --matrix, --chain and --gate are byte-identical without --behavioural.
  • NOT RUN is never a pass. A machine that cannot run a check and a machine whose check passed would otherwise produce identical bytes. clean, pass, passed, ok and success are not in the vocabulary and a test asserts they cannot get in.
  • A local type check is not a service acceptance. Every executed result says which of the two it is, and the test suite asserts the wording.

--behavioural is refused with --chain (exit 2) rather than ignored. What each of the four probes runs, which callable, and which pinned version: DESIGN.md.


Who this is for

Bridge and gateway authors. If you are converting between any of these formats, your bridge is losing information at the boundary and the usual answer is a hand-written lossy dict. This computes it — and for a multi-hop path, it computes the compounding the dict cannot show.

Anyone choosing a format. The matrix is the comparison: which targets a source crosses into intact, and which destroy it. A format that can hold a schema and one that cannot look identical in a README and behave nothing alike.

Protocol and spec authors. If your spec has normative text and JSON Schemas, this machine-checks whether they agree with each other — a class of defect a feature checklist cannot see, because a checklist proves your code matches your own reading of the spec, not that your reading and your code agree. It found eleven such defects in our own protocol: eight in the schema-versus-prose ledger, one security control that was correct, unit-tested, and never called, and two naming-rule incompatibilities found by measuring servers we did not build. Every one is fixed upstream and retained as a regression guard. DESIGN.md has the ledger, its arithmetic, and the defect-by-defect account.

Framework maintainers. If your framework publishes agent definitions, this tells you what a consumer in another format will not be able to represent.


Install

uvx agent-loss-map                       # or: pipx run agent-loss-map
pip install -e .                       # from a checkout
agent-loss-map

--behavioural needs three client libraries, which are an optional extra rather than a default dependency:

pip install 'agent-loss-map[behavioural]'

The client libraries are pinned exactly (openai==1.109.1, anthropic==1.11.0, mcp==1.20.0) rather than with a floor, because every executed result names the installed version in its evidence — "the client's types accept this" is worth nothing without knowing whose types they are.

--mcp measures a live server: the harness performs a real handshake, projects tools/list faithfully, refuses to report anything if its own projection dropped a field, and reports the era it spoke. Three MCP captures ship with the package, so agent-loss-map --from mcp --to openai-fc works with nothing else installed.

Flag
--matrix every source against every target in one table
--chain F,F,F measure a multi-hop path, feeding each hop the previous hop's artifact
--behavioural also run the executed half against real client-library code
--scorecard the per-section table on its own
--markdown the full report as Markdown, for an issue or a PR
--json machine-readable
--badge status-badge JSON for the worst finding against the spec
--mcp URL capture tools/list from a live server and measure it now
--from SOURCE only measure sources matching a framework id or format family
--to TARGET project into a target format and validate the artifact produced
--corpus PATH measure a format that is not bundled
--fail-on-blocker exit 1 if any blocker is found
--gate exit 1 only on a blocker against UACP itself (the release gate)

Modes answer different questions and refuse to be combined: --chain rejects --to, --matrix and --behavioural with exit 2 and says why. --gate is not evaluated on a chain run, because a chain cannot answer whether UACP's own schemas agree with its own normative text, and answering it "clean" would be a false green.

A report that silently tests an old protocol is worse than no report, so two workflows guard that: CI fails a build on real schema drift and on the test suite, while a scheduled probe re-runs the measurement and opens a PR when the committed report, badge or scorecard differ. Staleness becomes visible in review rather than breaking a build — DESIGN.md has both workflows and the pinned vendored-schema provenance.


Add your format

A corpus entry is pure data and the checks read nothing else, so measuring a format needs no Python and no pull request:

agent-loss-map --corpus my-format.json
{
  "framework": "my-framework",
  "agent_id": "my_agent",
  "description": "What this is.",
  "confidence": "documented",
  "provenance": "Where you read the shape, and when",
  "capabilities": [
    { "name": "web.search", "description": "Search the web.",
      "parameters": { "type": "object",
                      "properties": { "query": { "type": "string" } } } }
  ]
}

It gets its own section in the report, computed by the same checks as everything else. --corpus also takes a directory and is repeatable. Every entry declares how well-sourced it is, and the loader refuses to guess — a missing or misspelt confidence is an error, not a silent downgrade:

Tier Meaning
observed-serialization we read a real serialized artifact
documented-api from official documentation of the public API
inferred our modelling choice, not a documented shape

The corpus is 13 entries: 5 observed-serialization (three MCP captures, a UACP descriptor quoted from its spec, and the Stripe subset derived by script) and 8 documented-api. The default run reports 44 findings — 4 spec-audit (2 major, 2 minor) and 40 cross-format (1 blocker, 13 major, 26 minor). Today's run against wippa-uacp reports zero blockers in the spec audit; the run's one blocker is loss measured in a foreign format, which is what the harness is for. Captured in docs/sample-report.txt.

To make a format permanent, append a ForeignAgent to agent_loss_map/corpus.py — about twenty minutes. CONTRIBUTING.md has the recipe and the provenance rules; DESIGN.md has the registry guards a new format must satisfy.


Known limits

  • The security claims are still guarded by a source probe, not by execution, and that is the limit that matters most. The unwired authorization was found by a person reading the code and is now guarded by a regex over the routing path. --behavioural does not change that: no probe runs UACP's own Bus, because the package must work without a checkout of the protocol repo. A source probe can be satisfied by code that is never called, and this one is still open to that.
  • A chain is fed forward through this harness's own re-read, not a client's parse. --chain mcp,openai-fc,a2a measures the path, but hop 2 consumes the artifact hop 1 actually produced, converted back into a capability list by agent_loss_map/chain.py. No OpenAI server, A2A client or Anthropic API was involved, and a real client may refuse a shape the re-read accepts. intact means the artifact carries that value, not that a client took it. The re-read's known distortions are enumerated per hop and printed with the chain, and the intermediate is labelled inferred because a real system did not serialize it — this package did. --matrix still computes each crossing independently; a chain is the third question, not a substitute for the other two.
  • Per-hop finding counts in a chain are not additive, and the report will not pretend they are. The headline is the cumulative share of the source's own declared fields the artifact still carries, and a hop that reported nothing because it had nothing left to lose is called out by name.
  • The executed half is real, and much smaller than the static one. --behavioural covers 3 formats with a maintained Python client; the other 13 report NOT RUN and say why. unavailable is never rendered as a pass, and clean/pass are not statuses the vocabulary can produce. A local type check is not a service acceptance: when the openai model's pydantic validation accepts a function block, no request was built, no key was needed, and nothing is known about what the API would do.
  • A format with no published keyword subset cannot produce a keyword finding. mcp, langchain, openapi, shopify and stripe declare no accepted subset, so a clean cell for them means "the rules they publish were satisfied", not "fully supported". A constraint on what a format accepts is not a thing it has documented.
  • Two records rest on mirror-sourced or disputed facts, and say so. The Anthropic keyword subset was read from a mirror of the structured-outputs limitations page, so it accepts rather than reports the keywords in doubt; the Anthropic and Gemini name-length limits are disputed between two official surfaces, so no length finding is reported for either.
  • The bundled MCP captures are not all the same protocol era, and each one says which it is. context7 is 2026-07-28 — modern, captured over server/discover, with per-request _meta and no session. deepwiki-mcp and adoraads-beauty are 2025-06-18 — legacy, captured over an initialize handshake. Every capture was probed for server/discover and labelled by what the server answered, not inferred from the corpus. The Tool data type — what the harness actually measures — has not changed in the ways that matter, but the two eras negotiate the transport differently, and a legacy client against a modern server fails on the current revision's own compatibility matrix.
  • A standards-compliant validator still cannot resolve the UACP schemas offline. The $ids claim canonical URLs on a host that does not resolve, so loading agent-descriptor and following its relative capability.schema.json ref attempts a network call. This harness resolves those refs locally by filename as well as by $id, and a test proves the resolution is offline rather than incidentally working. The finding stays major because the defect is in the published artifact, not in the workaround.
  • An absent protocol checkout makes the code checks unverified, not passing. Set UACP_SOURCE_ROOT to a checkout to have them run. Silently skipping them and reporting clean would be the failure this harness exists to catch.
  • FINDINGS.md covers the spec audit, not the whole run. It explains what the schema-versus-prose and schema-versus-implementation checks found in wippa-uacp, which is the durable part. For current counts across all 13 corpus sources, read docs/sample-report.txt, which the scheduled probe regenerates from a live run.

Verifying it

pip install -e ".[dev]"
pytest -q

316 tests, 1 skipped. The suite is mostly meta-tests: each mutates a copy of the input so the defect it looks for is absent, and asserts the check goes quiet. A check that quietly stopped working fails the suite rather than passing silently — which has twice caught a guard in this repo that had quietly stopped guarding anything, both times because relaxing a rule invalidated the fixture it used to trigger on.

MIT.

Metadata

Release files for agent-loss-map 0.20.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for agent-loss-map 0.20.0
File Size Uploaded
agent_loss_map-0.20.0.tar.gz 318.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for agent-loss-map 0.20.0
File Interpreter ABI Platform
agent_loss_map-0.20.0-py3-none-any.whl Python 3 none any Details

Total release size: 523.5 kB

Release files / agent_loss_map-0.20.0.tar.gz

Download URL agent_loss_map-0.20.0.tar.gz
Size 318.9 kB
Tags Source
SHA-256 checksum
How to use checksums
23a59d498b2439c339425986d09b80948018cee9eb3fbe78cc4b6714c8786144
BLAKE2b-256 checksum
How to use checksums
b566f62a7ca81d7cb3377c2a9d097480b43d45ad559624adf475c0a82459a1ee
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.

Transparency log

Release files / agent_loss_map-0.20.0-py3-none-any.whl

Download URL agent_loss_map-0.20.0-py3-none-any.whl
Size 204.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
e496bf9d0b25d1aa2000fe1b23f266068108323b6d67bebc6cb8e610e4947c9c
BLAKE2b-256 checksum
How to use checksums
e094019c02207b5f4e15da642629f67c0e57c1a9c8ab3051a0a964b0b4b75624
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.20.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page