Skip to main content

Cage — a flux: deterministic attribution ledger for LLM token traffic and tool savings ($0, stdlib).

Project description

Cage — a flux

Cage is a flux: a deterministic engine for the flow of tokens and calls through an AI tool stack. It meters every LLM call, collects a savings receipt from each tool in the stack, and turns the raw stream into an attribution ledger — what you spent, what each tool saved you, what any other combination of tools would have cost, and how much money and time the agent saved vs a person doing the task (anchored to the commit it produced). $0, stdlib-only, deterministic, and independent of any single AI tool.

Cage is the third in a family of deterministic substrate → derived views tools: graphify (code → graph), fux (decisions → rules/memory), and now Cage (LLM traffic + savings receipts → ledger, attribution, counterfactuals).

The full design of record is in docs/cage-plan.md.


Status — v0.3 (Tier-1 human axis + tool-savings receipts)

Build-order step (plan §9) Status
1. Substrate contract (call record, receipt, policy.toml)
2. Tier-0 meter + ledger (record_call, cage report)
3. Receipt emitters (record_receipt, compressor, response-cache)
4. Attribution + matrix (cage attrib, cage matrix)
5. Adapters — library meter() · proxy · transcript hooks · cage meter
6. Plugin — cage mcp · /cage skill · agent hooks/wiring
7. Tier-0 savings (compressor, exact-match cache) + §8 features

The attribution engine (§4, the differentiator) reproduces the plan's worked example against a real ledger — cage demo. 78 tests passing. The optional [embeddings]/[ml] tiers stay off by default (semantic cache + learned compressor are pluggable adapters over the same receipt shape).

Tier-1 — agent vs human. Beyond tool-vs-tool savings, Cage models the whole-task counterfactual: what a person would have cost in time and money. A human receipt is just a receipt whose tool is "human", priced in minutes → money at a configured rate ([human] in policy.toml, or CAGE_HUMAN_RATE). Every figure is estimated (never measured unless you supply a real timesheet) and carries a confidence so round task-type guesses read as low-credibility:

Agent vs human · 14 tasks · rate source: policy ($80/hr)
agent     tasks   human $    agent $    saved $   saved hrs   conf   method
claude       9    $1,140.00    $4.12    $1,135.88     13.2     0.51   estimated

Quickstart

./install.sh                 # editable install → the `cage` binary ($0, stdlib only)
cd your-project && cage init # scaffold .cage/ (policy + gitignored ledger)

cage demo                    # seed the plan's §4.4 worked example
cage attrib                  # per-tool marginal savings (the §4.2 table)
cage matrix                  # the counterfactual permutation table (§4.4)
cage report --by model       # ledger rollup: spend by model

Metering from your code (the library adapter)

The adapter targets the protocol, not any named tool — you call it, it doesn't wrap you, and it is fail-open (a metering error never breaks your call):

import cage

with cage.meter("code-edit", task="fix-bug") as m:
    resp = client.messages.create(...)            # any Anthropic/OpenAI client
    m.usage(provider="anthropic", model="claude-opus-4-8",
            tokens_in=8600, tokens_out=1500, cached_in=3200)

# A tool that shrank the context files a receipt for what it spared you:
cage.record_receipt(tool="fux", raw_alternative=8000, actual=1600,
                    call=m.call_id, task="fix-bug", method="modeled")

What cage demo proves

The §4.4 worked example — one task, three deterministic tools each shrinking a different slice of context — reproduced against a real ledger:

Marginal attribution · task 'fix-handover-bug' · anthropic/claude-opus-4-8
tool        saved tok  saved $  method
graphify       27,000  $0.0810  modeled
fux             6,400  $0.0192  modeled
compressor      8,000  $0.0240  measured
TOTAL          41,400  $0.1242

Counterfactual matrix … full stack vs all-off: 72% cheaper ($0.1725 → $0.0483)

Every cell is tagged measured / modeled / estimated — you always know which numbers are invoices and which are projections. Only the configuration you actually ran is measured; no projection masquerades as an invoice (plan §4.1).

CLI

Command What it does
cage init scaffold .cage/ (policy + gitignored ledger)
cage report [--by route|model|day|agent] [--since 7d] ledger rollup
cage attrib [--task ID] per-tool marginal savings (§4.2)
cage matrix [--task ID] counterfactual permutation table (§4.4)
cage budget [--session ID] session/day spend vs policy.toml ceilings
cage roi [--since 30d] saved $ per tool vs its own cost + latency
cage human [--since|--task|--agent] [--html] agent-vs-human: $ and hours saved per agent (§4.1)
cage human-record --task ID (--type T|--minutes N|--usd N) record the Tier-1 human alternative for a task (§5)
cage matrix --human the §4.4 matrix with a human anchor row + vs-human columns
cage trend [--by week|month] [--metric cost|time|both] cost+time savings as a time-series (§5b.4)
cage why <call-id> full provenance: a call + every receipt against it
cage quality / cage outcome <task> cost per successful task (§8.2)
cage regression alert when cost-per-call drifts up (§8.3)
cage recommend cheapest-path: which tools to enable/skip (§8.4)
cage forecast project monthly spend vs the budget (§8.5)
cage graphify -- graphify <query|path|explain> … meter a third-party graphify call (transparent passthrough; files a savings receipt)
cage serve local dashboard over the ledger
cage demo seed the §4.4 worked example

Every read command takes --json for the agent-as-user (machine-readable, typed).

Works with any agent — target the protocol, not the tool

Cage meters whatever speaks the wire format and reads the ledger over MCP, so all four agents share one ledger contract:

Agent Meter its spend Read the ledger
Claude Code SessionEnd hook parses the transcript (proxy-free) /cage skill + cage MCP
Codex cage meter -- codex … / cage import-codex cage MCP (~/.codex/config.toml)
Copilot cage proxy (point its base URL at it) cage MCP (.vscode/mcp.json) + instructions
Kiro cage proxy cage MCP (.kiro/settings/mcp.json) + steering
Your code / Orff cage.meter() library adapter cage CLI / MCP
cage setup                 # install a global /cage asset into all four agent homes
cage hooks install         # wire claude/codex/copilot/kiro in this project
cage hooks install --claude   # or one surface at a time
cage proxy --port 8788     # the universal meter for clients you can't edit
cage meter -- codex exec   # run any agent under the proxy for one shot

Tool-savings receipts — owned vs third-party

A tool earns rows in attrib/matrix/roi by filing a savings receipt. Two strategies, by who owns the tool (see docs/agents.md and the receipt contract):

  • In-tool (you own it) — e.g. fux carries a fail-open cage_receipt.py and emits its own tool="fux" receipt; cage stays optional (fux runs unchanged with cage absent).
  • External adapter (third-party) — e.g. graphify: cage graphify -- graphify query "…" runs graphify unmodified, passes its output through byte-for-byte, and files a tool="graphify" receipt by parsing the cited source_files. graphify is never edited; a metering error never alters its result.

Design constitution

  • $0, stdlib-only, deterministic — no model in the maintenance path. Heavy ML is an optional, off-by-default tier ([embeddings], [ml]), never required.
  • Target the wire protocol, never the tool — Cage speaks the message format and the receipt schema. Anything that speaks them works; nothing is named.
  • PII-safe by construction — the ledger stores token counts, never prompt bodies. Point CAGE_LEDGER at a private store to keep even the counts off-disk.
  • Honest attribution — marginal-by-fixed-order ($0, defensible); every number carries its method. Shapley is a deferred opt-in audit mode (plan §9).

MIT licensed.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

cage_flux-0.3.0.tar.gz (65.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

cage_flux-0.3.0-py3-none-any.whl (66.3 kB view details)

Uploaded Python 3

File details

Details for the file cage_flux-0.3.0.tar.gz.

File metadata

  • Download URL: cage_flux-0.3.0.tar.gz
  • Upload date:
  • Size: 65.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for cage_flux-0.3.0.tar.gz
Algorithm Hash digest
SHA256 eced40cd2ea5576a3d519d21eaff409f95c7cc7af59ec0e6cceb4e50a10684da
MD5 5ffcaf91ed3cf82598739f76897c3cd6
BLAKE2b-256 a1868a950e8f396390a9eba7f7f88aefdfd71a134f653ffd37650a12ddb991c2

See more details on using hashes here.

Provenance

The following attestation bundles were made for cage_flux-0.3.0.tar.gz:

Publisher: publish.yml on arpitarya/cage

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file cage_flux-0.3.0-py3-none-any.whl.

File metadata

  • Download URL: cage_flux-0.3.0-py3-none-any.whl
  • Upload date:
  • Size: 66.3 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for cage_flux-0.3.0-py3-none-any.whl
Algorithm Hash digest
SHA256 c32e4d055d3bfb07210d2af65875769167f88f875e7de217bbc2ff9e610cf723
MD5 ace9c8ce5224887da4ca06d71341433c
BLAKE2b-256 f115d391b4e892eddf1888685aed63e1d74bf7b08bbdc0e7e8c6c056db9ed40f

See more details on using hashes here.

Provenance

The following attestation bundles were made for cage_flux-0.3.0-py3-none-any.whl:

Publisher: publish.yml on arpitarya/cage

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page