TokenAtlas
Formerly claude-code-energy-monitor.
TokenAtlas is a local, private token and cost atlas for AI coding agents (Claude Code, Codex, OpenCode, Pi): a durable history CLI (tokenatlas), an offline HTML report, list-price costing, session trees and fixed-context overhead. The statusline and companion scripts below are the original energy monitor and keep working from a checkout.
A statusline script for Claude Code that shows real-time token usage and order-of-magnitude energy estimates. It tracks daily, weekly, and monthly totals, distinguishes cheap cached tokens from expensive fresh tokens, and logs history automatically.
The repo also includes companion monitors for Codex CLI (codex_status.py) and the Pi coding harness (pi_status.py). Both read their harness's session JSONL and print the same style of one-line summary for a prompt, tmux status, or sidecar terminal.
Opus 4.8 | Ctx:42% | 5h:29% 7d:52% | D:2.0M ~2kWh | W:45.3M ~20kWh | M:412M ~50kWh
Reading left to right:
| Segment | Meaning |
|---|---|
Opus 4.8 |
Active model |
Ctx:42% |
Context-window utilization (current session) |
5h:29% 7d:52% |
API quota consumption (5-hour and 7-day rolling windows) |
D:2.0M ~2kWh |
Daily total tokens and energy estimate |
W:45.3M ~20kWh |
Weekly total (rolling 7 days) |
M:412M ~50kWh |
Monthly total (rolling 30 days) |
Energy is model-weighted (Haiku ×0.3, Sonnet ×0.6, Opus ×1.0; order-of-magnitude) for today and going forward; older history days are model-agnostic. Quota prefers the statusline payload's rate_limits (no API call), falling back to the OAuth usage endpoint when absent.
Explain a spike across Claude, Codex, Pi, and OpenCode
why.py reads retained Claude Code transcripts, Codex rollouts, Pi sessions, and OpenCode's local
usage store directly and ranks the calls that consumed a time window. It groups by project, main
thread vs subagent, entrypoint, agent, exact model, effort, and session. Codex Desktop calls are
included when they are present in the shared ~/.codex/sessions/ directory.
# All supported harnesses, trailing 24 hours
python3 why.py
# A short burst, or one local calendar day
python3 why.py --hours 0.5
python3 why.py --date 2026-09-03
# Narrow to one harness or emit structured output
python3 why.py --harness codex --limit 10
python3 why.py --hours 6 --json
The command is stateless: it writes no ledger or cache. Claude requests are deduplicated by
requestId using the maximum of each streamed token field; Codex uses each turn's
last_token_usage delta and discards repeated cumulative snapshots. Pi deduplicates copied fork
history by provider response ID and preserves the original session attribution. OpenCode transiently
parses local message records, then retains only allowlisted usage and attribution fields. Resumed sessions land on the
day when each call actually happened. For a full Claude calendar day the command also compares
transcript totals with the older statusline counter and reports the latter's coverage. Use
--harness both for the legacy Claude+Codex pair, or select claude, codex, pi, or opencode.
Durable local usage history
tokenatlas imports Claude, Codex, Pi, and OpenCode observations into a local SQLite
history. It needs Python 3.10+ and the standard library, makes no network or LLM
calls, and does not change your installed statusline.
# Default store: ~/.local/state/tokenatlas/history.sqlite3 (or XDG_STATE_HOME)
tokenatlas refresh --harness claude
tokenatlas refresh --harness codex
tokenatlas refresh --harness pi
tokenatlas refresh --harness opencode
tokenatlas refresh --all # every harness from its default roots (Claude incl. Cowork)
tokenatlas doctor
# Stored history, local time buckets; timestamps must include an offset
tokenatlas report --start 2026-09-01T00:00:00+02:00 --end 2026-10-01T00:00:00+02:00 --granularity hour
tokenatlas report --granularity minute --session SESSION_ID --records
# Isolated input/state, useful for a demonstration or fixture
tokenatlas --db /private/tmp/tokenatlas-demo/history.sqlite3 refresh --harness claude --root /path/to/transcripts
All commands print JSON. refresh imports only changed files; running it twice
or importing a copied transcript does not add the same usage again. Recorded
history survives source deletion. Schedule refresh separately if desired;
there is no background service, and already-deleted logs cannot be recovered.
refresh exits 2 for missing sources or partial imports and records diagnostics.
report supports day/hour/minute, --timezone, --harness, --project (full
identity), --session (including displayed HARNESS:SESSION_ID names), and --turn. --records includes source-file references
and observed/derived turn links. The output is private local data: project
paths are visible, although prompt content and tool inputs are never stored.
Missing counters remain null; known_tokens is only the identified observed subtotal.
Ambiguous identities are excluded and reported separately because they may overlap.
A complete normalized observation does not mean complete account coverage or
verified billing. Nontrivial Claude iterations are retained and flagged rather
than silently ignored or added twice. Claude subagent transcripts often lack the final
usage row of a request, so their output can be a lower bound (warning output_not_final). See accounting and storage decisions.
Keeping the report fresh
A report is a snapshot. To keep one current without manual steps, refresh on a schedule and rebuild the
report only when the data changed and the file is old enough. The history keeps a revision counter
(shown by doctor) that grows whenever a refresh or import stores something new. The report records a
state fingerprint in a <meta name="tokenatlas-state"> tag; it covers the data revision, the report options
(privacy, timezone, granularity, filters) and the coverage, so changing any of them rebuilds the report.
# Every harness from its default roots; a harness that is not installed here is reported as "absent"
tokenatlas refresh --all
# Rebuild only if data changed (--if-changed) and the report is older than 1h (--max-age: 90s, 30m, 1h, 2d)
tokenatlas report --html ~/.local/state/tokenatlas/report.html --private --if-changed --max-age 1h
refresh --all prints one JSON document with a status (ok, partial or missing, the worst of the
present harnesses) and one entry per harness. It exits 0 when every present harness is ok and 2
otherwise; absent harnesses never fail it, so a machine without Claude Code still exits 0. A skipped
report prints {"html": ..., "skipped": true, "reason": "unchanged"|"too recent"}, exits 0 and does not
touch the file. Without --if-changed and --max-age, report --html always rebuilds. A refresh that
finds nothing new leaves the revision alone; a re-import of identical rows with known sources does too.
tokenatlas open reuses a background report built with
report --html $XDG_STATE_HOME/tokenatlas/report.html --private when nothing changed, and only opens it.
Cron, every 30 minutes (use the full path; cron has a short PATH):
*/30 * * * * ~/.local/bin/tokenatlas refresh --all && ~/.local/bin/tokenatlas report --html ~/.local/state/tokenatlas/report.html --private --if-changed --max-age 1h
On macOS prefer launchd: it runs a missed job after sleep, cron does not. Save as
~/Library/LaunchAgents/com.tokenatlas.report.plist and run launchctl load on it:
<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<plist version="1.0"><dict>
<key>Label</key><string>com.tokenatlas.report</string>
<key>StartInterval</key><integer>1800</integer>
<key>ProgramArguments</key><array>
<string>/bin/sh</string><string>-c</string>
<string>$HOME/.local/bin/tokenatlas refresh --all && $HOME/.local/bin/tokenatlas report --html $HOME/.local/state/tokenatlas/report.html --private --if-changed --max-age 1h</string>
</array>
</dict></plist>
On a laptop that sleeps, the background report can be hours old after a wake. Run tokenatlas open when
you want an exact view: it runs refresh --all (continuing if a harness is partial), writes a private
report to $XDG_STATE_HOME/tokenatlas/report.html (or --html PATH; --shared pseudonymizes it) and
opens it in your default browser. --no-refresh skips the refresh. report itself keeps the shared,
pseudonymized default unless you pass --private.
Session trees and outcomes
# Tree of one root session: subagents, workflows, child sessions, per-model totals
tokenatlas session <session-id> [--json] [--no-infer]
# List thread keys, then rate a unit of work (appends to outcomes.jsonl next to the database)
tokenatlas rate <session-id>
tokenatlas rate <session-id> --unit review --thread claude:agent:<id> --outcome pass --note "ok"
Costs are API-equivalent list prices, not what was paid; unknown cost prints n/a. Once outcomes exist,
session also shows cost per passed result per model. Details and limits: docs/sessions.md.
Top prompts
tokenatlas top [-n 5] [--by cost|tokens] [--harness H] [--project P] [--start ISO] [--end ISO] [--json]
tokenatlas top --keep-text -n 5 # opt in: remember the text of the current top 5 prompts
tokenatlas top --forget-text # delete the stored text
tokenatlas top --json --with-text # include stored text in JSON
A prompt is one user turn plus every request it caused, including subagent work (rolled up by time and
parent session). Prompts are ranked by list-price cost (unpriced last, and any non-USD price counts as unpriced) or by tokens; cost is API-equivalent,
n/a when unpriced and ≥ when only partly priced. The HTML report has a matching "Dyraste prompterna" card
that follows the report filters.
Prompt text is never stored in the history database and never in snapshots or imports. Only if you run
top --keep-text does TokenAtlas write top-prompts.json next to the history (mode 0600, written atomically).
It holds, for the current global top prompts only, the harness, session, turn id, capture time and a sanitized
preview (whitespace collapsed, secret-like strings masked, at most about 200 characters) read from this
machine's own logs; prompts from other machines get no text. A prompt that falls out of the top is removed
from the file on the next --keep-text. Shared (pseudonymized, default) reports never contain prompt text;
private reports (report --html --private, open without --shared) show the previews of the prompts that are in the
current global top (by the recorded -n and ranking) when the file exists. Previews whose prompt has fallen out of the top
are kept only until the next --keep-text. Prompts whose text could not be read are retried on each --keep-text.
On POSIX the store is only read when it is a regular file owned by you with mode 0600 and no other hard links; a symlink,
a foreign owner, group/other access or extra links are refused with a warning, and --keep-text replaces such a file
with a fresh 0600 one (never following a symlink). --forget-text unlinks the file itself and warns if other hard links still
hold the text. On Windows these owner and mode checks are unavailable: the file relies on the user profile's ACLs.
Fixed context overhead
tokenatlas overhead --refresh reports the floor tokens of a session's first request, the sizes of
instruction files, skills lists and skill bodies, and an API-equivalent estimate of re-reading that floor on
every call. Only sizes, names, counts and token counts are stored, never content. Component tokens are
character counts divided by 4 (an estimate). A comparison table at the top gives usage-normalized measures per
harness (fixed share of input, per 1k output, share of cost, calls per session), sorted most efficient first. See docs/overhead.md.
Install the history and TokenAtlas CLI
Python 3.10 or newer is required. pipx keeps the command isolated from other Python tools:
pipx install git+https://github.com/Magnus-Gille/tokenatlas
tokenatlas --version
For a checkout, use pipx install /path/to/tokenatlas. On systems with an externally managed Python (for example Raspberry Pi OS), use the venv recipe in docs/remote-machines.md. The old energy-monitor command remains as a deprecated alias that prints a one-line notice.
Upgrade or remove the command with:
pipx upgrade tokenatlas
pipx uninstall tokenatlas
Uninstalling the command intentionally preserves the local history database at
~/.local/state/tokenatlas/history.sqlite3 (or $XDG_STATE_HOME/tokenatlas/history.sqlite3). Back it up
or remove it separately according to your own data-retention policy. Running python3 -m tokenatlas from
a source checkout remains supported. On first run the former agentmon data directory is moved to tokenatlas automatically (never with an explicit --db).
Install the Claude Code statusline
30-second setup — paste this into Claude Code:
Please set up a custom statusline for me. Do the following:
- Download
statusline.pyfrom https://github.com/Magnus-Gille/tokenatlas and save it to~/.claude/statusline.py- Run
chmod +x ~/.claude/statusline.py- Run
claude config set --global statusline "python3 ~/.claude/statusline.py"to enable it
Or do it manually:
# Download the script
curl -o ~/.claude/statusline.py https://raw.githubusercontent.com/Magnus-Gille/tokenatlas/master/statusline.py
# Make it executable
chmod +x ~/.claude/statusline.py
# Register it with Claude Code
claude config set --global statusline "python3 ~/.claude/statusline.py"
That's it. The statusline appears the next time you start a Claude Code session.
Codex CLI
Codex support lives in codex_status.py. It is not injected into the Codex TUI itself; Codex currently writes rollout logs instead of calling an external statusline command.
# From this repo
chmod +x ./codex_status.py
# Print one status line for the latest Codex rollout
python3 ./codex_status.py
# Continuously refresh it in a side terminal
python3 ./codex_status.py --watch
Example output:
gpt-5.4 | Win:258k | 5h:36% 7d:46% | D:1.8M ~100Wh | W:27.4M ~5kWh | M:75.6M ~10kWh
Useful integrations:
# zsh right prompt
setopt PROMPT_SUBST
RPROMPT='$(python3 /path/to/codex_status.py 2>/dev/null)'
# tmux status-right
set -g status-right "#(python3 /path/to/codex_status.py 2>/dev/null)"
Notes:
- Codex rollout logs expose total input, cached input, output, reasoning output, context window, and 5h/7d rate-limit usage.
- Current Codex rollouts expose cache-write tokens.
why.pyreports them; the older energy status scripts do not yet apply a distinct cache-write energy constant. - Parsed rollout summaries are cached in
~/.codex/statusline_rollout_cache.jsonto keep prompt-time execution reasonably fast.
Codex step counter
For copy/pasteable summaries similar to stepcount.py, use codex_stepcount.py:
python3 ./codex_stepcount.py
python3 ./codex_stepcount.py -d
python3 ./codex_stepcount.py -w
python3 ./codex_stepcount.py -m
python3 ./codex_stepcount.py -t
python3 ./codex_stepcount.py --rough-energy-estimate
python3 ./codex_stepcount.py --copy
Example output:
⚡ Codex
Today 5.1M tokens · 5 sessions
Week 30.8M tokens · 1395 sessions
Month 79.0M tokens · 2129 sessions
Print after Codex exits
Codex still has no native Stop hook, so the practical equivalent is a wrapper that runs Codex and then prints the summary after the process exits.
This repo includes codex_with_summary.py:
python3 ./codex_with_summary.py
python3 ./codex_with_summary.py resume --last
python3 ./codex_with_summary.py exec "fix the failing tests"
If you want this behavior on your normal codex command, add a shell function to your ~/.zshrc:
codex() {
python3 /path/to/codex_with_summary.py "$@"
}
Optional:
- Set
CODEX_SUMMARY_ARGS="--rough-energy-estimate"to change the summary flags. - The wrapper intentionally skips printing a summary for
--help,--version,completion,features,login,logout,mcp, anddebug.
Pi coding harness
Pi support lives in pi_status.py. Pi already displays live-session token and cache usage in its footer; this companion adds daily, weekly, and monthly totals with the repository's energy proxy by reading ~/.pi/agent/sessions/.
# Print one status line from saved Pi sessions
python3 ./pi_status.py
# Continuously refresh it in a side terminal
python3 ./pi_status.py --watch
# Emit machine-readable totals
python3 ./pi_status.py --json
Example output:
gpt-5.6-sol | D:512k ~200Wh | W:1.8M ~500Wh | M:4.2M ~1kWh
For shareable summaries:
python3 ./pi_stepcount.py
python3 ./pi_stepcount.py -d
python3 ./pi_stepcount.py -w
python3 ./pi_stepcount.py -m
python3 ./pi_stepcount.py -t
python3 ./pi_stepcount.py --rough-energy-estimate
python3 ./pi_stepcount.py --copy
Notes:
- Pi records normalized
input,output,cacheRead,cacheWrite, andreasoningusage on every assistant response. Reasoning is a subset of output and is not added twice. - Calls are assigned to the date of each assistant response, including sessions continued on later days.
- Pi forks and clones can copy earlier entries into a new session file. The monitor deduplicates copied responses while retaining genuinely new calls on each branch.
- Pi's internal compaction and branch-summary model requests are persisted without usage fields, so their tokens and energy cannot be recovered from session JSONL and are omitted.
--no-sessionruns are not persisted and therefore cannot be counted afterward.- Parsed responses are cached by file size and modification time in
~/.pi/agent/statusline_session_cache.jsonfor prompt, tmux, and watch performance. - Pi can use many providers. The energy values remain the same documented, provider-agnostic order-of-magnitude proxy; they are not suitable for comparing providers or models.
pi_scanner.pyis unrelated: its name refers to a Raspberry Pi machine running headless Claude Code jobs.
Useful integrations are the same as for Codex:
# zsh right prompt
setopt PROMPT_SUBST
RPROMPT='$(python3 /path/to/pi_status.py 2>/dev/null)'
# tmux status-right
set -g status-right "#(python3 /path/to/pi_status.py 2>/dev/null)"
How Claude Code monitoring works
- Claude Code calls the script on every status update (after each API call, during streaming). It pipes a JSON object with session data into stdin.
- The script reads the data, computes energy estimates, updates daily totals, and prints one line to stdout — which Claude Code renders in the status bar.
- Daily totals persist across sessions in
~/.claude/statusline_daily.json. Multiple concurrent sessions are handled safely with file locking. - At midnight (or rather, on the first prompt of a new day), the previous day's totals are archived to
~/.claude/statusline_history.jsonland the daily counter resets.
No cron jobs, no daemons, no manual intervention. You just use Claude Code and the data accumulates.
Data files
| File | Purpose | Format |
|---|---|---|
~/.claude/statusline_daily.json |
Today's running totals | JSON object with per-session and daily aggregates |
~/.claude/statusline_history.jsonl |
Historical daily log | One JSON line per day, appended automatically |
~/.claude/statusline_quota_cache.json |
Cached API quota data | JSON with 5-minute TTL |
All files are created with owner-only permissions (0600). Example history entry:
{"date": "2026-02-18", "input": 2797805, "output": 693769, "cache_read": 1901548, "cache_write": 312000, "sessions": 12}
Step counter (shareable summaries)
stepcount.py generates copy-pasteable usage summaries from the accumulated history — like a fitness tracker for AI coding.
python3 stepcount.py # today + week + month stacked (default)
python3 stepcount.py -d # today only
python3 stepcount.py -w # last 7 days only
python3 stepcount.py -m # last 30 days only
python3 stepcount.py -t # today + week + month as ASCII table
python3 stepcount.py --copy # copy output to clipboard
Add --rough-energy-estimate to any view to include the order-of-magnitude energy guess.
Week and Month rows only appear once you have 7 and 30 days of data respectively.
Optional: auto-print after each session. Add a Stop hook to ~/.claude/settings.json:
"hooks": {
"Stop": [
{
"hooks": [
{
"type": "command",
"command": "python3 '/path/to/stepcount.py'"
}
]
}
]
}
Example output (default):
⚡ Claude Code
Today 18M tokens · 9 sessions
Week 563M tokens · 106 sessions
Month 917M tokens · 156 sessions
Example output (-t):
⚡ Claude Code
┌───────┬────────┬──────┬────────────┐
│ │ tokens │ sess │ │
├───────┼────────┼──────┼────────────┤
│ Today │ 18M │ 9 │ █░░░░░░░░░ │
│ Week │ 563M │ 106 │ ██████░░░░ │
│ Month │ 917M │ 156 │ ██████████ │
└───────┴────────┴──────┴────────────┘
Results & Claims
This section separates what we can measure with high confidence from what we can only estimate at order-of-magnitude level.
What we measure (data source)
The script reads Claude Code's statusbar JSON payload, piped to stdin on every status update. This payload contains:
- Per-call snapshot (
current_usage):input_tokens,output_tokens,cache_read_input_tokens,cache_creation_input_tokens— these reflect the most recent API call. The monitor accumulates these across calls (detecting call boundaries) to build daily totals. - Current-context totals:
total_input_tokens(=input + cache_creation + cache_readof the most recent response) andtotal_output_tokens(that response's output). As of Claude Code v2.1.122 these are current-context snapshots, not cumulative session counters. Earlier builds of this monitor treated them as cumulative; the accumulation now sums per-callcurrent_usagefields and derives fresh (uncached) input astotal_input − cache_read − cache_creation.
Daily totals are accumulated across sessions via a locked JSON file. The statusbar payload and Claude Code's JSONL transcripts are complementary local sources whose coverage depends on the Claude Code version and workload. In a later v2.1.226 comparison, deduplicated transcript output was within 1% of the statusline output for a full observed day, while the statusline missed subagent and non-interactive calls. Neither source should be treated as universally complete (see Known limitations below).
Token accounting claims (high confidence, validated)
These claims are supported by a validation harness that logged raw statusbar payloads (Feb 2026: 31 calls / 3 sessions; re-validated May 2026 on CC v2.1.157), plus direct API billing reconciliation:
- No double-counting. Energy applies separate constants to four non-overlapping token types: fresh (uncached) input, cache reads, cache creation, and output. Fresh input is derived per call as
total_input − cache_read − cache_creation(≈current_usage.input_tokens). - Thinking visibility is version-dependent. In the February 2026 captures, the statusbar
output was roughly 3x the deduplicated JSONL output, and those JSONL records did not expose
thinking in
usage.output_tokens. This was a dated observation, not a universal multiplier or rule: on Claude Code v2.1.226, deduplicated transcript output was within 1% of statusline output for a full observed day. - Cache metrics are accurate. Per-call
cache_read_input_tokensandcache_creation_input_tokensmatch API billing to the token across 4 direct API test calls, and are stable within a call (the monitor sums them once per call). - Most prefill work shows up as cache creation. In a heavily-cached workload, truly-fresh input is tiny (~1% of energy); the bulk of prefill work is
cache_creation. This is whyE_CWis treated as roughly prefill-cost.
⚠️ Schema change (CC v2.1.122):
total_input_tokens/total_output_tokensbecame current-context snapshots rather than cumulative session counters. Builds of this monitor before the 2026-05 fix mis-counted input (≈53× over) and output (under) as a result. See docs/energy-constants.md for the corrected accounting and re-validation.
Full evidence in FINDINGS.md. To collect your own validation data, set ENERGY_DEBUG=1 as an env var in the statusline command, then run python3 analyze_tokens.py after a session.
Energy estimate claims (order-of-magnitude proxy)
The energy numbers shown in the statusbar are order-of-magnitude estimates, not measurements. They use per-token energy constants derived from published research and refined via adversarial debate (see methodology below), applied to each token type:
Energy = (fresh_input × 0.39) + (output × 1.40) + (cache_read × 0.015) + (cache_write × 0.49) Wh per 1k tokens
The display snaps to order-of-magnitude steps (1, 2, 5, 10, 20, 50, ...) because the real uncertainty is at least ±3x in each direction. This is intentionally coarse — it reflects genuine uncertainty, not imprecision in the token counting.
Which token type dominates energy depends heavily on the workload. For long interactive sessions (e.g. a Feb 2026 Opus month, ~757M tokens, ~48 kWh mid) output dominated (~46% of energy from ~3% of tokens, since decode is ~3.6x more expensive per token than parallel prefill). For many short automated/headless sessions (the more recent pattern) cache creation dominates (~34–40%), because each session writes a fresh cache that is read few times before expiring. Don't assume one fixed breakdown — the statusline reflects whatever mix you actually run.
What the energy estimate does NOT include
- Full datacenter overhead. The estimates cover GPU/accelerator compute only — not cooling, networking, CPU/RAM, storage, idle power, or PUE. Google reported a 2.4x multiplier from chip-active to full-stack for Gemini queries. The real operational energy is likely 1.5–3x higher than what this script shows.
- Training energy. Training a frontier model costs tens of gigawatt-hours, but that's a one-time cost amortized across millions of users.
- Embodied energy. Manufacturing GPUs, building datacenters, networking infrastructure.
- Your own hardware. Your laptop and monitor also consume energy while you wait for responses.
Known limitations
-
Pricing ≠ energy. Anthropic's pricing ratios were the original basis for relative energy cost between token types. We've since revised the output and cache read constants using physics-derived cross-checks (FLOP-based estimates, AI Energy Score benchmarks, Google's measured per-query energy). The fresh input and cache write constants still inherit from pricing. Pricing reflects margin, competitive positioning, and demand management — not just energy.
-
Per-model constants are approximate. The current implementation uses rough order-of-magnitude weights — Haiku ×0.3, Sonnet ×0.6, Opus ×1.0 — as a model for this monitor, based on the input-price ratio (1:3:5) discounted for sub-linear parameter-to-energy scaling. These are project assumptions, not universal energy multipliers or measurements. Anthropic discloses no parameter counts. (Earlier versions used the same constant for all three tiers.)
-
Context-length decode scaling. The formula uses a fixed per-output-token constant regardless of context length. With very long cached contexts (25M+ tokens observed in practice), decode cost increases due to larger KV-cache attention. The formula underestimates in exactly these long-context sessions.
-
Infrastructure variability. We don't know Anthropic's hardware (GPU types, cluster config), batch sizes, scheduling strategies, model sizes, or datacenter locations. Inference efficiency is a rapidly moving target — Google reported a 33x improvement in a single year.
-
Source coverage and fields vary by version and workload. In the February 2026 captures, JSONL had streaming placeholder values for
usage.input_tokens(75% were ≤1 in that dataset) and did not expose thinking inusage.output_tokens. On Claude Code v2.1.226, a full observed day of deduplicated transcript output was within 1% of statusline output, but statusline collection missed subagent and non-interactive calls. Neither source is a universal billing ledger. Streaming records must be deduplicated byrequestId, taking the maximum of each token field. These comparisons measure local source coverage; they are not API billing reconciliation. See GitHub issue #28197.
How to use these numbers responsibly
Appropriate uses:
- Awareness — understanding the general scale of compute behind AI-assisted coding
- Relative comparison — "today was a heavier compute day than yesterday"
- Order-of-magnitude budgeting — "our team's AI usage is in the X kWh/day range"
- Motivating efficiency — choosing smaller models for simple tasks, being mindful of long-context sessions
Not appropriate:
- Precise carbon accounting or ESG reporting (the uncertainty is too large)
- Comparing energy efficiency between AI providers (the constants are derived from one provider's data)
- Claiming exact energy figures without stating the ±3x uncertainty range
Energy estimation methodology
The problem
No one outside Anthropic knows the actual energy per token for Claude models. There are no published measurements. Rather than pretending to have precise numbers, the script shows a range with ~10x total uncertainty.
Mid estimates
The constants use a hybrid approach: Couch's (2026) base estimates from Epoch AI's GPT-4o research, revised via adversarial debate (Claude vs Codex, Feb 2026) using physics-derived cross-checks.
| Token type | Mid estimate | Derivation |
|---|---|---|
| Fresh input (prefill) | 390 mWh/1k tokens | Epoch AI long-context anchor (unchanged from Couch) |
| Output (decode) | 1,400 mWh/1k tokens | Reduced from 1,950; cross-checks cluster 600–1,800 |
| Cached input (cache read) | 15 mWh/1k tokens | Reduced from 39; physics-derived ~26x discount vs input |
| Cache creation (write) | 490 mWh/1k tokens | Prefill + write overhead, ~1.25x fresh input (unchanged) |
The display shows order-of-magnitude estimates (snapping to 1/2/5 per decade) because the real uncertainty is at least ±3x in each direction.
Why output tokens cost ~3.6x more than input
Input tokens are processed in parallel (prefill), while output tokens are generated one at a time (autoregressive decode). This serial generation is inherently less efficient. The original 5:1 ratio (from Anthropic's pricing) was revised down to ~3.6:1 based on FLOP-based estimates, AI Energy Score benchmarks, and measured Llama 405B inference data, which cluster around 1,000–1,800 mWh/1k output tokens.
Why cached tokens are ~26x cheaper
Claude Code aggressively caches conversation context. When tokens are read from cache, they skip the expensive prefill computation entirely — the cost is primarily loading pre-computed KV pairs from memory. Anthropic's pricing gives a 10x discount, but physics analysis shows the real compute saving is much larger (potentially 100–1,000x). The 26x discount is a conservative compromise: it accounts for the near-zero compute cost of cache loading plus the ongoing attention cost during decode over cached context.
Why cache creation costs ~1.25x fresh input
When tokens are written to cache for the first time, they require the same prefill computation as fresh input plus the overhead of writing to cache storage. Anthropic charges a 25% surcharge for cache creation, which Couch uses as a proxy for the additional energy cost.
Cross-checks against published measurements
The estimates were sanity-checked against multiple independent data points:
- Google measured 0.24 Wh per median Gemini query (August 2025) — comprehensive, including idle capacity and PUE
- OpenAI reported 0.34 Wh per average ChatGPT query (June 2025)
- Couch derived 41 Wh per median Claude Code session (January 2026) — based on JSONL logs which undercount by ~2.8x (see FINDINGS.md)
- AI Energy Score benchmarks: ~600 mWh/1k output tokens for 70B models, scaled to ~1,200 for 200B+
- FLOP-based estimate: 750–1,500 mWh/1k output tokens for a 200B-class model with datacenter overhead
- Llama 405B measured (batched): ~2,800 mWh/1k output tokens including overhead
Rationale for each constant in docs/energy-constants.md. Full debate transcripts are in debate/ (gitignored due to size).
Why the uncertainty is so large
We don't know:
- Anthropic's hardware (GPU types, cluster configuration)
- Batch sizes and scheduling strategies
- Model size (parameter count)
- Inference optimization stack
- Geographic location of datacenters
Inference efficiency is also a rapidly moving target. Google reported a 33x improvement in a single year.
Everyday comparisons
To put the numbers in context:
| Activity | Energy |
|---|---|
| One Google search | ~0.3 Wh |
| One ChatGPT/Gemini query | ~0.3 Wh |
| Charging a smartphone | ~15 Wh |
| LED bulb for 1 hour | 10 Wh |
| Electric oven for 1 minute | 50 Wh |
| Driving an EV 1 km | ~150 Wh |
| Swedish household daily use (with electric heating) | ~40 kWh |
A typical day of AI-assisted coding likely falls in the 1–5 kWh range (mid estimate). A heavy month (119 sessions on Opus 4.6) estimated at ~48 kWh, roughly equivalent to running a fridge for a month. See FINDINGS.md for detailed analysis.
Platform support
| Feature | macOS | Linux | Windows/WSL |
|---|---|---|---|
| Token tracking | Yes | Yes | Yes |
| Energy estimates | Yes | Yes | Yes |
| Daily history | Yes | Yes | Yes |
| Prompt cache tracking | Yes | Yes | Yes |
| Durable history and TokenAtlas | Yes | Yes | Yes** |
| API quota display | Yes | Yes* | Yes* |
* Quota now comes primarily from the statusline payload's rate_limits fields (Claude Code v2.1.80+, Pro/Max), which work on every platform with no API call. The legacy fallback reads the OAuth token from the macOS Keychain via the security command and is macOS-only; on Linux/Windows it's simply skipped. So if your build provides rate_limits, quota shows everywhere; otherwise it's macOS-only.
** The packaged CLI installs tzdata on native Windows so named IANA timezones such as
Europe/Stockholm remain available. WSL normally uses the distribution's timezone database.
Note: The fallback quota path uses an undocumented Anthropic beta API endpoint (/api/oauth/usage with anthropic-beta: oauth-2025-04-20), which may change or disappear without notice. The token and energy features do not depend on it.
Dependencies
The statusline script uses only the Python 3 standard library (json, os, sys, subprocess,
time, fcntl, pathlib, datetime). The packaged history CLI is also standard-library-only on
macOS, Linux, and WSL; native Windows installs the small tzdata package for named timezones.
fcntl note: The file locking uses fcntl.flock, which is available on macOS and Linux. On Windows (outside WSL), this would need to be replaced with an alternative locking mechanism.
Security
- The OAuth token is read from macOS Keychain and sent only to
api.anthropic.com. It is never written to disk. - On POSIX systems, data files are created with
0600permissions (owner read/write only). Native Windows uses the current user profile's filesystem ACLs because POSIX owner/mode checks are unavailable. - The script makes no network calls other than the optional quota fetch to Anthropic.
- No telemetry, no third-party services, no analytics.
Risk: If someone modifies ~/.claude/statusline.py, they get code execution in your user context on every Claude Code update. Same threat model as a shell alias or git hook. Keep the file owner-only writable.
Token counting: validated semantics
We built a validation harness that logs raw statusbar payloads and analyzes token-counting behavior across API calls (Feb 2026: 31 calls / 3 sessions; re-validated May 2026 on CC v2.1.157). Key findings:
total_input_tokens=input + cache_creation + cache_readof the most recent response (current-context, not cumulative, since CC v2.1.122). Fresh (uncached) input is recovered astotal_input − cache_read − cache_creation.total_output_tokens= the most recent response's output (per-call, not cumulative since v2.1.122). The February 2026 captures showed a roughly 3x statusbar-to-JSONL output ratio, but that ratio and the associated thinking-token interpretation are version/date-specific; v2.1.226 matched within 1% for a full observed day after requestId deduplication. The monitor accumulatescurrent_usage.output_tokensper call.- Cache fields are stable within a call and summed once per detected call boundary; no double-counting across the four token types.
To collect your own validation data, set ENERGY_DEBUG=1 as an env var in the statusline command, then run python3 analyze_tokens.py after a session.
Full investigation details in FINDINGS.md.
References
Primary sources for energy estimates
| Source | Year | Description |
|---|---|---|
| Couch, "Claude Code's Environmental Impact" | 2026 | Derived per-token energy for Claude from Epoch AI data and Anthropic pricing ratios. Base for fresh input and cache write constants. |
| Epoch AI, "How much energy does ChatGPT use?" | 2025 | Empirical analysis of GPT-4o inference energy |
| Google, "Measuring the environmental impact of AI inference" | 2025 | Google's own measurements: 0.24 Wh per median Gemini query, 33x efficiency improvement in one year |
| AI Energy Score v2 | 2025 | Standardized inference energy benchmarks on H100 hardware; used to cross-check output constant |
| Altman, "The Gentle Singularity" | 2025 | OpenAI's reported 0.34 Wh per average ChatGPT query |
Additional academic references
| Source | Year | Description |
|---|---|---|
| Luccioni et al., "Power Hungry Processing" | 2024 | Systematic measurement of inference energy across model sizes and tasks |
| Husom et al., "The Price of Prompting" | 2024 | Analysis of energy costs for LLM prompting strategies |
| AI Energy Score v2 | 2025 | Documents 150-700x energy overhead for reasoning/chain-of-thought modes |
License
Author
Magnus Gille — gille.ai
Built collaboratively with Claude Opus 4.6, with a 2026-05 accuracy audit (token-accounting fix for CC v2.1.122, per-model multipliers, refreshed energy literature) by Claude Opus 4.8. Energy estimates, comparisons, and arithmetic independently verified by OpenAI Codex against DOE, ENERGY STAR, IEA, and Swedish Energy Agency sources.
TokenAtlas: standalone offline dashboard
Create an interactive Swedish report from the retained local history (Python standard library only):
tokenatlas --db /path/to/history.sqlite3 report --html tokenatlas.html
tokenatlas --db /path/to/history.sqlite3 report --html private.html --private
open -a Safari private.html
The default HTML pseudonymizes projects, sessions, turns and agents. --private retains short unique
project labels and session references. Neither mode embeds prompts, tool text, source paths, raw usage
objects or machine IDs. Exact timestamps and model metadata remain in shared reports: pseudonymization
is not guaranteed anonymity. Shared reports show a model name verbatim only when it is in a public model family under a public provider; a custom endpoint that reuses a public provider id with a family-like model name is shown, so review a shared report before sending it. JSON from the original --records command remains a private diagnostic
export and has different privacy semantics.
The report works without a server, network, CDN or runtime model calls. Filters cover dates, harness, provider, model, effort, project, session, agent and thread type. Click a timeline bar to zoom from day to hour/minute, or a ranking to filter. Expand sessions and turns for individual observations. CSV/JSON exports preserve the current selection, and SVG/print exports support presentations. Top rankings retain an Other subtotal. Empty intervals are not claimed to be measured zeros. Large timelines group consecutive observed buckets with an explicit note; the horizontal spacing is categorical. Cache-read share comparisons show the highest and lowest comparable session, harness, model, and project with observation count, input volume, and explicit excluded-group coverage.
Unknown counters remain unknown, ambiguous identities are excluded from the known subtotal, and reasoning is normalized as a subset of output even when a harness stores it additively. No invoice, quota, wall time or quality is inferred. Source coverage is separate from arithmetic completeness. Collector revisions now invalidate file checkpoints so existing retained source files are reparsed after parser fixes; lost source files cannot be recovered.
Checks: python3 -m unittest discover -v and python3 test_interactive_export.py. The release suite
also builds and installs a wheel in an isolated environment. Optional local browser
check: PLAYWRIGHT_MODULE=/path/to/@playwright/test node test_report_browser.cjs /absolute/report.html.
Use an existing Playwright installation and its bundled browser; no browser dependency is installed
by these commands. The browser test blocks network access and checks filters, totals, zoom, drilldown,
exports and mobile overflow.
Demo data and screenshots
scripts/demo.py regenerates the product-page screenshots from entirely fictional data:
python3 scripts/demo.py OUTDIR [--seed N] [--no-screens] [--shared]
It builds a temporary HOME with synthetic Claude Code, Codex, Pi and OpenCode logs (about 30 days in September 2026, three fictional projects under /Users/demo/code, dummy lorem-style content), runs python3 -m tokenatlas against it (refresh, overhead --refresh, session demo-orchestrated, report --html) and writes demo-report.html, session.txt, overhead.txt, demo-summary.json and, unless --no-screens, overview.png, session.png and overhead.png. Real logs and state are never read. The HTML report keeps the fictional project labels by default; --shared redacts them. The screenshots need Playwright (PLAYWRIGHT_MODULE, else a cached npx copy) via scripts/demo_screens.cjs.
Metadata
Release files for tokenatlas 1.3.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| tokenatlas-1.3.0.tar.gz | 133.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| tokenatlas-1.3.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 237.8 kB
Release files / tokenatlas-1.3.0.tar.gz
| Download URL | tokenatlas-1.3.0.tar.gz |
|---|---|
| Size | 133.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
c132c7092ee3be6f8394936043a9174da55eb7723bdf4a018f496ceee5c82008
|
|
BLAKE2b-256 checksum How to use checksums |
1fe9df71ad607a2f53b7c607baf2eeb4d3f714d54b11572fee2f87131428e7f4
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.
Transparency logRelease files / tokenatlas-1.3.0-py3-none-any.whl
| Download URL | tokenatlas-1.3.0-py3-none-any.whl |
|---|---|
| Size | 104.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
5d364f6beef6c8c3fddcf2f012a2b975fe701d104cf677f88269c2b776fb3e75
|
|
BLAKE2b-256 checksum How to use checksums |
0c57bdc9422fb85e0b71218e57bffdd69fc248a76c9e80ab3fa6b77eb0d70f76
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.
Transparency log