Skip to main content

Cage — a flux: deterministic attribution ledger for LLM token traffic and tool savings ($0, stdlib).

Project description

Cage

Cost dashboards tell you what your AI stack spent. Cage tells you what each tool actually saved you — and what a human would have cost instead.

PyPI Python Dependencies License

You're paying for an agent, a graph tool, a rules engine, maybe Copilot. At the end of the month someone asks "is any of this worth it?" — and the honest answer is a shrug and a Slack thread. Cage meters every LLM call, collects a savings receipt from each tool in the stack, and turns the raw stream into an attribution ledger: what you spent, what each tool saved you, what every other combination of tools would have cost, and how much money and time the agent saved versus a person doing the same task. $0, deterministic, zero dependencies, no model in the maintenance path.

Named after John Cage. · Python ≥ 3.11 · stdlib only · MIT · sits beside fux, graphify, bach, wagner, orff.

▶ Demo GIF coming soon.

The story

Another README story. Yeah. Because nobody ever walked out of a meeting humming a feature table, and you will not remember mine. So forget the table. Here's ninety seconds about a conference room, a pile of money, and a bunch of people who have no idea what they're talking about. One of them is you. — Arpit

You ever notice how everybody's saving money now? Everybody. The agent's saving money. The graph tool's saving money. Copilot's saving money. Two tools you built over a weekend — saving money. Add it all up and you should be getting a check in the mail. Funny thing about that. The bill went up.

Here's the con. Nobody — and I mean nobody — can show you the number. They got slides. They got a roadmap. They got a guy named Kevin who "feels like it's a game-changer." What they don't got is one honest figure that says this tool saved this much on this task, and here's what it would've cost to do the boring old way, by hand. Ask for that number and watch the room go quiet and somebody suggest we "circle back."

And the kicker — you built half of it. So when finance points at you and says "is this worth it," you, the expert, the one who's supposed to know — you got a screenshot and a feeling. You're not in trouble for spending the money, folks. You're in trouble because you bought the same fog everybody else did.

Cage is the thing that ruins the fog. It's the itemized receipt nobody asks for and everybody needs: graphify saved 27,000 tokens here, fux saved 6,400, the agent did in four minutes what a person does in two hours — plus every other combo you could've run, priced out, each number stamped so you know which ones are real and which ones are some computer's best guess. It doesn't do synergy. It does arithmetic.

See it

$ cage matrix --task fix-handover-bug
Counterfactual matrix · task 'fix-handover-bug' · anthropic/claude-opus-4-8
  base 2,000 tok + output 1,500 tok held constant

graphify  fux  compressor   input tok    cost    source
   ✗       ✗       ✗           50,000   $0.1725   modeled
   ✓       ✗       ✗           23,000   $0.0915   modeled
   ✓       ✓       ✗           16,600   $0.0723   modeled
   ✓       ✓       ✓            8,600   $0.0483   measured   ← the run you actually made

  full stack vs all-off: 72% cheaper ($0.1725 → $0.0483)

Per-tool savings any meter can attempt. The part no cost dashboard does is the rest of that table — what each stack you didn't run would have cost — and the source column, so you always know which row is an invoice and which is a reconstruction. Only the configuration you actually ran is measured; no projection ever masquerades as an invoice. That discipline is the whole product.

Quickstart

pip install cage-flux           # the CLI, zero third-party deps
cd your-project
cage init                       # scaffold .cage/ (policy + gitignored ledger)
cage adopt                      # wire all four agents + the graphify interceptor
cage demo                       # seed the worked example
cage matrix                     # the counterfactual permutation table
cage human                      # agent-vs-human: $ and hours saved

Adopting into a project is one idempotent command — cage adopt wires Claude Code / Codex / Copilot / Kiro onto one ledger and drops a transparent bin/graphify interceptor. Pass --claude (etc.) for a subset, --no-graphify / --no-hooks to skip parts.

Metering from your own code is the library adapter — it targets the protocol, not any named client, and is fail-open (a metering error never breaks your call):

import cage

with cage.meter("code-edit", task="fix-bug") as m:
    resp = client.messages.create(...)            # any Anthropic/OpenAI client
    m.usage(provider="anthropic", model="claude-opus-4-8",
            tokens_in=8600, tokens_out=1500, cached_in=3200)

# A tool that shrank the context files a receipt for what it spared you:
cage.record_receipt(tool="fux", raw_alternative=8000, actual=1600,
                    call=m.call_id, task="fix-bug", method="modeled")

Explain it like I'm five

You and a robot helper did the chores. At the end of the day someone wants to know: did the robot actually help, or did it just look busy?

Cage is the chart on the fridge. It writes down how long each chore took with the robot, and how long it would have taken if you'd done it yourself — so you can see, in real minutes and real dollars, which helper earned its place and which one just made noise. And it's careful to mark which numbers it actually timed and which ones are its best guess, so nobody gets fooled by a confident-looking total. It does all of this for free, without ever phoning a friend for the answer.

Why it's different

It's not another cost dashboard. The difference is a set of properties, not features:

  • Deterministic. Every derived view — report, attribution, the counterfactual matrix, ROI, the human axis — is pure parse/arithmetic over an append-only log. Same ledger + same policy ⇒ identical tables, every time. The numbers never drift because nothing guesses.
  • Honest by construction. Every figure carries a method: measured (a real invoice), modeled (a reconstructed counterfactual), or estimated (a human/labor guess). A projection can never read as an invoice — the one property a "trust me, it paid off" slide can't offer.
  • $0 and zero-dependency. Stdlib-only Python, dependencies = []. Heavy ML is an opt-in, off-by-default tier ([embeddings], [ml]), never on the default path. Portable as a tarball, auditable line by line.
  • Agent-native. Every read command takes --json; the ledger is served over MCP. Built so an agent can pull its own cost numbers and verify them, not just read a chart.

The "so what" chain: deterministic → so the numbers never hallucinate → so each one carries a defensible method → so you can put the savings claim in front of finance, or an auditor. That last clause is the one a dashboard can't say.

Honest attribution — the part that survives the room

Anyone can sum a bill. Cage's job is to divide credit without lying about it, and it does that with three rules:

  • Marginal-by-fixed-order. Each tool's receipt reports the saving it produced given the tools upstream of it in the canonical pipeline. The marginals sum exactly to the total — no overlap, no double-counting, $0 to compute, and defensible because the order is fixed and visible (not a black-box Shapley pass; that's a deferred opt-in audit mode).
  • The counterfactual matrix. For a task whose tools each shrank a slice of context, Cage enumerates the 2ⁿ on/off permutations and prices each at the task's model — so "what would graphify-off + fux-on have cost?" is a row, not a hand-wave. Only the configuration actually run is measured; every reconstructed cell is modeled (or estimated if it leans on an estimate).
  • Tier-1 — agent vs human. Beyond tool-vs-tool, Cage models the whole-task counterfactual: what a person would have cost in time and money. A human receipt is just a receipt whose tool is "human", priced in minutes → money at a configured rate ([human] in policy.toml, or CAGE_HUMAN_RATE). It is estimated unless you supply a real timesheet, and carries a confidence so round task-type guesses read as low-credibility instead of masquerading as precise. The time metric can go negative — if the agent thrashed longer than a human would have, the table says so.
Agent vs human · 14 tasks · rate source: policy ($80/hr)

agent     tasks   human $    agent $    saved $   saved hrs   conf   method
claude       9    $1,140.00    $4.12    $1,135.88     13.2     0.51   estimated
codex        3      $260.00    $1.55      $258.45      3.1     0.50   estimated
TOTAL       14    $1,530.00    $6.55    $1,523.45     17.9     0.51

The savings are anchored to the commit they produced — Cage snapshots a git-aware task record (SHA, branch, diff size, wall-clock) at task close, so a number can always be traced back to the change that earned it.

How it works

One append-only log in, every view derived from it for $0:

record_call / record_receipt  →  .cage/ledger/{calls,receipts,tasks}.jsonl  (append-only)
        (meter, fail-open)                    │
                                              ▼  derive ($0, no model)
   policy.toml (prices/order/budgets/rates) → report · attrib · matrix · roi
                                             · human · trend · budget · why

You meter at the provider boundary (library adapter, a reverse proxy for clients you can't edit, or by parsing a Claude Code / Codex transcript). Everything downstream is a deterministic projection. The ledger carries token counts, never prompt bodies — PII-safe by construction; point CAGE_LEDGER at a private store to keep even the counts off-disk.

A tool earns rows in attrib/matrix/roi by filing a savings receipt, and there are two ways in, by who owns the tool:

  • In-tool (you own it) — e.g. fux carries a fail-open cage_receipt.py and emits its own tool="fux" receipt. Cage stays optional; fux runs unchanged with cage absent.
  • External adapter (third-party) — e.g. graphify: cage graphify -- graphify query "…" runs graphify unmodified, passes its output through byte-for-byte, and files a tool="graphify" receipt by parsing the cited source_files. graphify is never edited; a metering error never alters its result.
The full command surface (ledger · attribution · human axis · ops · agents)
cage init                      # scaffold .cage/ (policy + gitignored ledger)
cage adopt [--no-graphify]     # per-project setup: wire 4 agents + graphify interceptor
cage doctor --json             # verify this project's setup is correct (non-zero on failure)
cage report --by model         # ledger rollup: spend by route / model / day / agent
cage attrib --task ID          # per-tool marginal savings (sum of marginals = total)
cage matrix --task ID          # the counterfactual permutation table (2ⁿ on/off)
cage matrix --task ID --human  # …with a human anchor row + vs-human columns
cage roi --since 30d           # saved $ per tool vs its own cost + added latency
cage human [--agent claude]    # agent-vs-human: $ AND hours saved, per agent
cage human-record --task ID --type feature   # record a Tier-1 human alternative
cage trend --by week --metric both           # cost + time savings as a time-series
cage why <call-id>             # full provenance: a call + every receipt against it
cage quality / cage outcome ID # cost per *successful* task (cost is honest with outcome)
cage regression                # alert when cost-per-call drifts up
cage recommend                 # cheapest-path: which tools to enable / skip
cage forecast                  # project monthly spend vs the budget
cage graphify -- graphify     # meter a third-party graphify call (transparent passthrough)
cage setup                     # install /cage + /cage-doctor into every agent home
cage proxy --port 8788         # the universal meter for clients you can't edit
cage mcp                       # serve the ledger to agents over MCP (stdio)
cage serve                     # local dashboard over the ledger
cage demo                      # seed the worked example that proves the thesis

Every read command takes --json for the agent-as-user (machine-readable, typed).

Works with any agent — target the protocol, not the tool

Cage meters whatever speaks the wire format and reads the ledger over MCP, so all four agents share one ledger contract:

Agent Meter its spend Read the ledger
Claude Code SessionEnd hook parses the transcript (proxy-free) cage MCP (.mcp.json)
Codex cage meter -- codex … / cage import-codex cage MCP (~/.codex/config.toml)
Copilot cage proxy (point its base URL at it) cage MCP + instructions (.vscode/)
Kiro cage proxy cage MCP + steering (.kiro/)
Your code / Orff cage.meter() library adapter cage CLI / MCP

cage setup installs two assets into every agent home, idiomatically per agent: /cage (read the ledger) and /cage-doctor (verify the wiring works). Any agent can then report Cage health on request, in any project.

The $0 guarantee

Every derived view is parse / arithmetic over the log — no LLM call, ever, on the read or maintenance path. The only model spend is whatever your agent already does; Cage just meters it. The semantic cache and learned compressor ship behind opt-in [embeddings] / [ml] extras; the default install is model-free and dependency-free. 92 tests passing; cage demo reproduces the worked attribution example against a real ledger.

Honest limits. Cage doesn't decide your human rate — it prices minutes at a blended rate you set, and labels the result estimated so it never pretends to be a timesheet. Marginal-by-fixed-order is defensible and $0, but it is an ordering convention, not a Shapley value (that's a deferred audit mode). And a counterfactual cell is an honest reconstruction, never an invoice — the method column says so on every row, on purpose.

What's new

  • v0.3.0 — the Tier-1 human axis. cage human / cage trend price agent-vs-human in dollars and hours, anchored to a git-aware task record; a minutes unit, a [human] rate table with confidence laddering, and CAGE_HUMAN_RATE. Third-party tools join via the external adapter (cage graphify).
  • v0.2.0 — attribution + the counterfactual matrix. Marginal-by-fixed-order attribution, the 2ⁿ permutation table, ROI per tool, and the measured/modeled/estimated discipline — the differentiator.
  • v0.1.0 — substrate + meter. The call/receipt contract, the append-only ledger, policy.toml, and cage report.

The name

Named after John Cage, whose 4′33″ framed four and a half minutes of "silence" so an audience would finally hear the ambient cost they'd been ignoring. Cage the tool does the same to your AI stack: it takes the spend and the savings everyone assumed were free or unknowable, and makes them something you can actually account for. It's the third in a family of deterministic substrate → derived views tools — graphify (code → graph), fux (decisions → rules) — and now Cage (LLM traffic + receipts → ledger). The names are deliberate, and they sit beside bach, wagner, and orff.


If you've ever been the one in the room with no numbers, ★ star the repo and run cage demopip install cage-flux.

License

MIT — see LICENSE.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

cage_flux-0.4.0.tar.gz (77.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

cage_flux-0.4.0-py3-none-any.whl (77.1 kB view details)

Uploaded Python 3

File details

Details for the file cage_flux-0.4.0.tar.gz.

File metadata

  • Download URL: cage_flux-0.4.0.tar.gz
  • Upload date:
  • Size: 77.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for cage_flux-0.4.0.tar.gz
Algorithm Hash digest
SHA256 2e850a6ead0e333f3abab09d7b873768461f6472cba1ad793343d26875d71d64
MD5 2613c96ae92a390ffa06ffbe5ac9ceae
BLAKE2b-256 c29797afe5b6347654e71285ff796126c9dc1c8d4eb6c328985348791d6474ee

See more details on using hashes here.

Provenance

The following attestation bundles were made for cage_flux-0.4.0.tar.gz:

Publisher: publish.yml on arpitarya/cage

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file cage_flux-0.4.0-py3-none-any.whl.

File metadata

  • Download URL: cage_flux-0.4.0-py3-none-any.whl
  • Upload date:
  • Size: 77.1 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for cage_flux-0.4.0-py3-none-any.whl
Algorithm Hash digest
SHA256 fdbdc54b485f60a875930daf2ba0e63d36c18bf3fa602b870b1658f24297a8cd
MD5 5021d523b9686d0ee8033a7d61316357
BLAKE2b-256 b4f6cbf75f781e56a1dffd15ef553f58979ec445e14c014ad9a3e4f499062148

See more details on using hashes here.

Provenance

The following attestation bundles were made for cage_flux-0.4.0-py3-none-any.whl:

Publisher: publish.yml on arpitarya/cage

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page