Skip to main content

maf-cachebench

PyPI Python License

Experimental. Releases before 1.0 may change or remove APIs. Importing this package emits MafCachebenchExperimentalWarning.

A benchmark for Microsoft Agent Framework compaction strategies. It measures what each strategy costs in provider prompt-cache hits and what it destroys while doing it: the framework's own strategies and the ones in maf-compaction, twenty in all, on the same conversation, priced at the provider's cached and uncached rates.

This is an independent package. It is not affiliated with or endorsed by Microsoft.

What it measures

Provider prompt caches match on an exact prefix, and every compaction strategy rewrites history inside that prefix, so compaction breaks the cache by construction. The question is what that costs and what it saves. The benchmark plants facts in a conversation's tool results, runs the conversation under each strategy, and then asks for the facts back from the compacted context. Each strategy gets a cost, a cache hit rate, and a count of facts it kept, and the strategies are ranked on cost behind an accuracy bar.

Two harnesses answer two questions:

  • cachebench replays a scripted transcript. Every provider and strategy replays the same scripted structure, with a fixed-width per-cell salt isolating its cache prefix. These controlled workloads support comparisons across providers. It answers how a prompt caches.
  • cachebench_live drives a real agent, with real replies and real tool calls, so compaction acts on the history an agent would accumulate. Its numbers compare strategies within one model. It answers what a strategy costs an agent.

Install

pip install maf-cachebench[openai]        # Azure OpenAI, OpenRouter and the Azure Responses route
pip install maf-cachebench[foundry]       # Microsoft Foundry project endpoints
pip install maf-cachebench[mistral]
pip install maf-cachebench[ollama]
pip install maf-cachebench[tiktoken]      # then pass --tokenizer tiktoken for exact counts

Five commands land on the path: cachebench, cachebench_live, cachebench_advise, cachebench_recall and cachebench_summary. Each prints its own --help.

Pointing it at a provider

A provider is selected as provider or provider:model, so two models on one provider can sit in one run. Credentials come from the environment:

Provider Variables
azure AZURE_OPENAI_ENDPOINT, AZURE_OPENAI_API_KEY, AZURE_OPENAI_CHAT_COMPLETION_MODEL
azure-responses AZURE_OPENAI_ENDPOINT, AZURE_OPENAI_RESPONSES_MODEL, and a working DefaultAzureCredential (Entra authentication)
foundry FOUNDRY_PROJECT_ENDPOINT, FOUNDRY_MODEL, and a working DefaultAzureCredential
openrouter OPENROUTER_API_KEY, OPENROUTER_MODEL; pin OPENROUTER_PROVIDER_ORDER, or you measure the router
mistral MISTRAL_API_KEY, MISTRAL_CHAT_MODEL
ollama OLLAMA_MODEL, optionally OLLAMA_HOST and OLLAMA_API_KEY; this provider caches and never reports it

Prices are passed as --price-input, --price-cached and --price-output per million tokens, except on OpenRouter where they are discovered.

A first run

# see the plan and the prompt sizes without spending anything
cachebench_live foundry:gpt-5.6-luna --dry-run --strategies none,anchored,tool_summary_anchored

# a live cell: a 120K window, a conversation sized to 1.5 times it, five seeds, records kept
cachebench_live foundry:gpt-5.6-luna --agent harness --context-window 120000 --fill 1.5 --repeats 5 \
  --strategies none,tool_summary_anchored,tool_and_user_summary_anchored --summarizer-provider foundry:gpt-5.6-luna \
  --price-input 0.20 --price-cached 0.02 --price-output 1.20 --results-jsonl runs/cell.jsonl

# render the table again from the records, calling nothing
cachebench_live --from-jsonl runs/cell.jsonl

--results-jsonl is not optional in practice: a cell runs for hours, and the file is appended one record per seed the moment it is scored, so a stopped run keeps what it paid for. --from-jsonl rebuilds the table and the verdict from the records with no provider configured, and files that measured the same cell merge into one table.

Reading the table

Rows that keep at least --min-correctness of the control's accuracy come first, cheapest first; the rest follow below a line. The columns to read, in order:

  • seed$ is what the conversation cost, the strategy's own summariser calls included and the probes left out. The ranking and the verdict are on it.
  • seed$+- is the spread between the cheapest and dearest seed. A gap smaller than this is not a result, and the verdict line says NOT SUPPORTED when that happens.
  • seed hit% is the share of the conversation's input the provider served from its cache, probes excluded. This is what compaction did to the cache.
  • facts and lost are the planted facts that survived into the compacted context, and the ones compaction removed.
  • acc1 and acc2 are the model's accuracy on the scoped and the combined closing questions, asked from the same restored snapshot.
  • dq marks a row that sent a prompt larger than the window it stood in for. A cell that disqualifies at all leaves the ranking.
  • flags is where a row says it is not measuring what its name claims. Read it before the money columns.

The full option reference, the flags legend and the measurement design are in the benchmark documentation, and the findings in the compaction overview.

What it depends on

agent-framework-core, pinned to one minor because the benchmark reads the framework's private _compaction helpers; maf-compaction for the strategies it was built to measure; and httpx for the advisor's price lookup. The provider clients are extras.

Licence

MIT. See the LICENSE.

All commands default to the dependency-free estimator. Install the tiktoken extra and pass --tokenizer tiktoken for BPE counts. Explicit price rates must be finite and non-negative.

Metadata

Release files for maf-cachebench 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for maf-cachebench 0.1.0
File Size Uploaded
maf_cachebench-0.1.0.tar.gz 200.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for maf-cachebench 0.1.0
File Interpreter ABI Platform
maf_cachebench-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 417.3 kB

Release files / maf_cachebench-0.1.0.tar.gz

Download URL maf_cachebench-0.1.0.tar.gz
Size 200.9 kB
Tags Source
SHA-256 checksum
How to use checksums
a390b453bde9779830afd425642ce653823d671b27d4f5113341d43f00617862
BLAKE2b-256 checksum
How to use checksums
7e6b8c22f0208755fd4189176bd5a8c5c1b93e980c88810f9ae4ca0a565eb79f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 7, 2026.

Transparency log

Release files / maf_cachebench-0.1.0-py3-none-any.whl

Download URL maf_cachebench-0.1.0-py3-none-any.whl
Size 216.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
19a417fce78f5a4505903943503115cf50555c619a143dc350e6f24864bceac2
BLAKE2b-256 checksum
How to use checksums
c45b7e7d813f8954c7093340a1396fe747d4908ff2d2af68649372920704f92c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 7, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page