agent-ready
The practical implementation of The Silent Revolution: How AI Agents Are Rewriting the Rules of API Growth — turning "audit your agent readiness" from advice into a command you can run.
Audit any OpenAPI spec for AI-agent readiness — then generate an MCP server scaffold for the endpoints that pass.
Your API works fine for human developers. That doesn't mean an AI agent can use
it reliably. agent-ready scores a spec against seven categories of
agent-specific failure, tells you exactly which endpoints will cause trouble and
why, and can gate a CI pipeline so specs don't regress.
pip install openapi-agent-ready
agent-ready https://example.com/openapi.yaml
Generic Booking Platform API (v1.0)
AI-readiness score: 54.4/100
Endpoints: 4 Fails: 6 Warnings: 12
Category scores:
Description Clarity 37.5/100
Side Effects 70.0/100
Error Responses 62.5/100
Auth Clarity 50.0/100
Ambiguity 100.0/100
Parameter Explanation 16.7/100
Usage Guidelines 50.0/100
Why this exists
A human developer who hits a vague 404 shrugs and checks the docs. An agent
dead-ends. A human who sees two similar endpoints picks the right one from
context. An agent picks whichever description sounds closer.
This tool is the practical implementation of an argument I made in The Silent Revolution: How AI Agents Are Rewriting the Rules of API Growth. That article's first recommendation was to audit your agent readiness — check OpenAPI spec completeness, authentication friction, and error message machine-readability, and score yourself honestly. This is that audit, made runnable.
The three characteristics of an agent-ready API identified in the article map directly onto the rubric:
| Article | Rubric categories |
|---|---|
| Machine-readable schemas | 1 (description clarity), 6 (parameters), 7 (usage guidelines) |
| Frictionless authentication | 4 (auth/scope clarity) |
| Structured error handling | 3 (actionable errors) |
Categories 2 (side effects) and 5 (ambiguity) were added during implementation — they turned out to matter as much as the original three, which is the usual reward for actually building the thing you wrote about.
The gap is measurable, and enormous. A 2026 study of 856 tools across 103 MCP servers found that 97.1% of tool descriptions contained at least one quality defect, and 56% failed to state their purpose clearly — with no significant difference between official and community-maintained servers.
agent-ready catches those defects at the spec level, before they become
agent failures in production.
Install
pip install openapi-agent-ready # core
pip install "openapi-agent-ready[server]" # + deps to RUN a generated scaffold
From source:
git clone https://github.com/sfaisal/agent-ready
cd agent-ready
pip install -e ".[dev]"
Usage
# Audit a local spec
agent-ready openapi.yaml
# Audit a remote spec
agent-ready https://api.example.com/openapi.json
# Full markdown report
agent-ready openapi.yaml --format markdown --out report.md
# Machine-readable output
agent-ready openapi.yaml --format json
# Generate an MCP server for the endpoints that pass
agent-ready openapi.yaml --generate-mcp server.py
# Generating only writes a file — running it needs the server deps:
pip install "openapi-agent-ready[server]"
python server.py
Use it as a CI gate
This is the point. Add it to your pipeline so agent-readiness can't silently regress when someone adds an endpoint:
- name: Check API is agent-ready
run: |
pip install openapi-agent-ready
agent-ready openapi.yaml --min-score 60 --max-fails 0
60 is a reasonable starting threshold — below it, an API is likely to lose
agent-driven integrations to a better-documented competitor. Raise it as you
close gaps; a spec that scores 60 has real problems, it just doesn't have
catastrophic ones.
Exit codes:
| Code | Meaning |
|---|---|
0 |
All configured gates passed |
1 |
A gate failed (--min-score / --max-fails / --max-new) |
2 |
The spec could not be loaded or parsed |
Adopting on an existing API
A mature API audited for the first time will produce hundreds of findings. No team can fix all of them before merging anything, so a score gate is unusable on day one and gets switched off within a week.
Record the current state instead, and block only what gets worse:
# once, when you adopt the tool — commit this file
agent-ready openapi.yaml --write-baseline .agent-ready-baseline.json
# in CI, from then on
agent-ready openapi.yaml --baseline .agent-ready-baseline.json --max-new 0
Existing findings stay visible in the report but don't fail the build. A pull request that introduces a new one does, and the output names exactly which findings it added:
Compared against .agent-ready-baseline.json: 5 new, 0 fixed, 11 unchanged
NEW FAIL GET /invoices — Description is a bare label (12 chars) ...
NEW FAIL GET /invoices — No error responses documented ...
Findings are matched on category, endpoint, and severity — not on message text, which contains counts and parameter names that shift when unrelated things change. A warning escalating to a fail counts as new; rewording a finding in a future release does not invalidate your baseline.
As the team fixes the backlog, regenerate the baseline to lock in the progress.
Real-world results
Running it against Stripe's production spec (587 endpoints), as of v0.1.0. These figures move whenever the rubric changes — rerun the tool for current numbers:
Stripe API (v2026-06-24.dahlia)
AI-readiness score: 62.9/100
Description Clarity 68.3/100
Side Effects 64.7/100
Error Responses 100.0/100
Auth Clarity 50.0/100
Ambiguity 50.0/100
Parameter Explanation 43.0/100
Usage Guidelines 51.2/100
Two things worth drawing out.
The first run scored 47.9 with error handling at 0/100 — and that was a bug
in this tool, not in Stripe. Stripe documents errors via OpenAPI's default
response rather than enumerating explicit 4xx/5xx codes. Perfectly valid, and
arguably cleaner. The check only looked for 4/5 prefixes, so it failed all
587 endpoints. Fixed, regression-tested, and the reason the README says to run
this against a real spec before trusting it.
Stripe is widely and correctly regarded as agent-ready, yet its raw OpenAPI
spec scores 62.9. That isn't a contradiction — it's the whole point. Stripe's
agent-readiness comes from the Agent Toolkit layered on top: curated tool
definitions, an MCP server, managed OAuth. The raw spec is the input to that
layer, not the layer itself. Which is exactly why this tool ships with
--generate-mcp: the audit tells you what to fix, and the scaffold is where the
curation happens.
Genuine findings that survived scrutiny: GET /v1/account and
GET /v1/accounts/{account} have 100% identical descriptions, and the spec
uses bearer/basic API keys rather than scoped OAuth2, so there's no way to
express "this agent may only read."
Python API
from agent_ready import load_spec, run_audit
spec = load_spec("openapi.yaml")
result = run_audit(spec)
print(result["overall_score"])
for finding in result["gaps"]:
print(finding.severity, finding.endpoint, finding.message)
The eight categories
| # | Category | What it catches | Published anchor |
|---|---|---|---|
| 1 | Description clarity | Endpoints a model can't decide whether to call | "Unclear Purpose" smell — 56% prevalence |
| 2 | Side-effect signaling | Write endpoints that don't say they're writes; missing idempotency | MCP tool annotations (destructive / open-world) |
| 3 | Actionable errors | Failures an agent can't recover from | Adjacent to "Limitations" — 89.8% prevalence |
| 4 | Auth/scope clarity | No way to tell delegated from autonomous agent access | Not covered by existing rubrics |
| 5 | Endpoint ambiguity | Two endpoints a model could confuse | 73% of servers had repeated tool names |
| 6 | Parameter explanation | Undocumented or format-ambiguous parameters | "Opaque Parameters" — 84.3% prevalence |
| 7 | Usage guidelines | States what it does, but not when to call it | "Missing Usage Guidelines" — 89.3% prevalence |
| 8 | Tool surface & schema strictness | Too many tools to select from; schemas incompatible with strict mode | OpenAI function-calling guidance |
A worked example of why category 6 matters
From the Hasan et al. paper: a Yahoo Finance MCP tool referred to "start" and
"end" without naming them as explicit parameters or specifying a date format.
The model couldn't construct a bounded range, so it fell back to a broad
period parameter and pulled multi-year windows to answer narrow questions —
inflating payload size, latency, and token cost. As the authors put it, that is
a specification problem in the description, not a model bug.
agent-ready flags exactly this: a date-like parameter with no format,
pattern, enum, or format hint in its description.
Provenance
Categories 1–5 were derived independently, from how MCP tool selection works plus standard API design principles reframed for a non-human caller. Categories 6–7 were added after validating the rubric against published empirical research:
- Hasan, Li, Rajbahadur, Adams & Hassan (Queen's University), "Model Context Protocol (MCP) Tool Descriptions Are Smelly! Towards Improving AI Agent Efficiency with Augmented MCP Tool Descriptions" — arXiv:2602.14878
- Wang et al., "From Docs to Descriptions: Smell-Aware Evaluation of MCP Server Descriptions" — arXiv:2602.18914
- Anthropic, Writing effective tools for agents
- OpenAI's function-calling guidance — source for category 8. OpenAI recommends keeping the initially-available tool count small (they suggest under 20) and notes that function definitions count against the model's context limit and are billed as input tokens.
The score is a diagnostic, not a target
Hasan et al. also found that enriching all description components improved task success by a median 5.85 percentage points — but increased execution steps by 67.46% and regressed performance in 16.67% of cases, because richer descriptions consume context window and raise cost. Their ablation study found that shorter, targeted descriptions often performed equivalently.
Fix the fails. Don't gold-plate the passes. The right operating point depends on your domain and your model, and choosing it is a product decision rather than an engineering one.
Known limitations
Named up front, because a tool that hides its weaknesses is worse than one that doesn't:
- The ambiguity check uses string similarity, so it misses endpoints that are semantically similar but differently worded. Embedding-based similarity would be stronger. (#1)
- Category weights are judgement calls, not measured effect sizes. Description clarity is weighted highest on the prior that wrong tool selection is the costliest failure — that prior is untested here.
- Static analysis can't catch descriptions that are fluent but wrong. Only running a real agent against the real API catches semantic inaccuracy.
- OpenAPI 3.x only. Swagger 2.0 and AsyncAPI aren't supported yet.
Related writing
- The Silent Revolution: How AI Agents Are Rewriting the Rules of API Growth — the argument this tool implements: why agent discoverability is displacing documentation quality as the competitive moat for API platforms, and what to measure instead of DAU and NPS.
Roadmap
- Embedding-based ambiguity detection
- Swagger 2.0 support
- Pagination and rate-limit documentation checks
- LLM-assisted rewriting of flagged descriptions (
--fix) - Live tool-selection testing against a generated MCP server
- GitHub Action wrapper
- Baseline mode (
--baseline) so CI only fails on new regressions
Contributing
Contributions welcome — see CONTRIBUTING.md. New rubric categories should come with a rationale for why the gap causes agent-specific failure, ideally with a citation or a reproducible example.
License
MIT — see LICENSE.
Metadata
Release files for openapi-agent-ready 0.4.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| openapi_agent_ready-0.4.0.tar.gz | 109.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| openapi_agent_ready-0.4.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 141.3 kB
Release files / openapi_agent_ready-0.4.0.tar.gz
| Download URL | openapi_agent_ready-0.4.0.tar.gz |
|---|---|
| Size | 109.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
a0930834685a8947f9eee0a323bd2f27612f99a39c3ff7a4b6b59a2e75cfdfde
|
|
BLAKE2b-256 checksum How to use checksums |
682e6f1487eca211e96ad328b00d5c44ea3a972e28186883eefb1729c21d3fcb
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 7, 2026.
Transparency logRelease files / openapi_agent_ready-0.4.0-py3-none-any.whl
| Download URL | openapi_agent_ready-0.4.0-py3-none-any.whl |
|---|---|
| Size | 31.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
c08c8b6f6ce2b831415a0f38cdc5f606e24ac43c245763acd36aafb09121e459
|
|
BLAKE2b-256 checksum How to use checksums |
2f202b180268de3f5da6fe60022df450dd20562ea4b3eaaadb11df4c2bb60883
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 7, 2026.
Transparency log