Skip to main content

AMI survey client

Measures what a completed agent workflow cost to run, by reading the runtime's own session log, and submits the result to the AMI survey at survey.agentbenchmark.dev.

The cost of a single API call is easy to obtain. The cost of one finished piece of work is not: one triaged ticket, one screened CV, one drafted reply, across every call, retry and tool round-trip the agent made getting there. This measures that figure.

Every number is read from the runtime's records rather than reported by the agent. An agent asked how many tokens it has just used will estimate, and will present the estimate with confidence.

What comes back

A real run of six support tickets, triaged and answered by Claude Opus 5 in Claude Code:

Maturity Index      85.0  Strong          (observability 40%, evidence 30%, quality 30%)
Performance         78.13 Strong          confidence Very High

  quality           80.0   graded Good on ami-quality-v2
  cost              72.73  $0.123226 per ticket   ($0.739355 for the run)
  speed             84.47  20.90s per ticket
  evidence          70.0   measured, on a self-issued token
  observability    100.0

findings
  weakness  Cost is the weakest pillar at 72.73; speed is strongest at 84.47.
            $0.123226 per unit against a $0.01 reference. A cheaper model, or
            fewer calls, moves this; check calls[] for where the tokens went.

  note      Cost and speed were scored against a provisional reference, which is
            a placeholder rather than a measurement. Do not quote them as settled
            yet. The Maturity Index does not use the reference and is unaffected.

$0.12 per ticket, 21 seconds per ticket. That is the figure this exists to produce, and it is the one most teams cannot currently state about their own work.

The findings record where the scorecard's own numbers are soft. A cost reference that is still a placeholder is reported as a placeholder rather than folded into the score.

Install

There are two routes. They differ in whether anything can read your runtime's logs, which determines whether the result is recorded as measured or unmeasured.

Measured

Add this to your agent's MCP configuration and restart it:

{ "mcpServers": { "ami-survey": { "command": "uvx", "args": ["ami-survey"] } } }

Then, once the agent finishes a piece of work, ask it:

Take the AMI survey regarding the ticket triage you just did

There is nothing to clone and nothing to keep updated. No token needs to be supplied: the first call that requires one registers the machine and stores the token at ~/.ami-survey/token.

uvx is part of uv, the tool the MCP documentation uses for Python servers. If uv is not installed, GETTING-STARTED.md gives a pipx form and a route that requires neither.

If you already hold a token, set it as AMI_API_TOKEN in that block's env and it will be used instead of registering a new one.

Unmeasured

In claude.ai, open Settings, then Connectors, then Add custom connector, and supply:

https://survey.agentbenchmark.dev/mcp

Nothing is installed and no token is required. These runs are recorded as unmeasured and are never compared against measured ones: a remote server cannot read your runtime's logs, so token counts and cost are absent rather than estimated.

Licence

An evaluation licence. This is not open source.

Permitted: installing and running the software on machines you control, for the purpose of evaluating it and submitting survey responses; and redistributing verbatim, unmodified copies with the licence intact.

Not permitted: modification beyond what is needed to run it for that purpose, derivative works, sublicensing, and sale.

Full terms in LICENSE.

What leaves your computer

Token counts, timings, model names, the stage names the workflow declared, and the grade. Not your files, not your prompts, not your shell commands. GETTING-STARTED.md sets this out in full.

Submissions go to survey.agentbenchmark.dev and nowhere else. That destination is a constant in the source rather than a setting, so a stale environment variable cannot redirect a submission onto your own disk. That is the one failure which would make a run appear successful while collecting nothing.

Requirements

Python 3.9 or newer. No dependencies; the standard library only.

Further reading

  • GETTING-STARTED.md assumes no prior setup, covers macOS, Linux and Windows, and explains what to do when the tools do not appear.
  • COMMANDS.md documents the clone-and-run route, which benchmarks one workflow across several models on your own API key. It is not required in order to take part.
  • MAKE-IT-MEASURABLE.md explains how to structure a workflow so that there is something worth measuring.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

ami_survey-1.0.1.tar.gz (70.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

ami_survey-1.0.1-py3-none-any.whl (81.1 kB view details)

Uploaded Python 3

File details

Details for the file ami_survey-1.0.1.tar.gz.

File metadata

  • Download URL: ami_survey-1.0.1.tar.gz
  • Upload date:
  • Size: 70.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.14.0

File hashes

Hashes for ami_survey-1.0.1.tar.gz
Algorithm Hash digest
SHA256 d6345ee93e66293f3c2522a40f18c056137f85be880f0c1de8f4cc7eaa373a1b
MD5 5c2c7ce44bc35c2d86f758b1de42d6c7
BLAKE2b-256 cd87070fa5e49263d404f5d13f64e5aee56e37a1666fd1e14da321b938bf917b

See more details on using hashes here.

File details

Details for the file ami_survey-1.0.1-py3-none-any.whl.

File metadata

  • Download URL: ami_survey-1.0.1-py3-none-any.whl
  • Upload date:
  • Size: 81.1 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.14.0

File hashes

Hashes for ami_survey-1.0.1-py3-none-any.whl
Algorithm Hash digest
SHA256 23ad23bc6572c6644dd226976a12a121eef200c85309797de111cf8d74bdae1e
MD5 3866fa5c16afb065ebb4c724b545ca27
BLAKE2b-256 329eaac1579ef9fa490e21d9ee95c342e4141fdd7d72a85cb2d3f5f57d0aa6ed

See more details on using hashes here.

Release history Release notifications | RSS feed

1.2.0

2 files

1.1.0

2 files

This release

1.0.1 This release

2 files

1.0.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page