ctxprof
A context profiler for Claude Code sessions (package ctxprof, command ctx) — a small, dependency-free CLI
(ctx) for understanding where a session's context and cache budget actually
goes. Like a CPU/memory profiler, but for tokens. Several lenses:
- Live —
ctx dashboard(all sessions) /ctx watch(one session): how much, right now. - Inventory —
ctx sessions: pick the right session id before runningshow,watch,explain,audit, orguard. - Per-turn —
ctx steps/ctx digest/ctx rates: what each turn cost and where the money went. - Cross-session —
ctx compare: cache-reuse quality across many sessions (best/worst/average), aggregate idle-rebuild waste, and spend by repo over a window — a productivity lens, not just one session's diagnostics. - Offline —
ctx audit: why one session burned context/cache, after the fact. - Explain —
ctx explain: ranks a session's biggest avoidable costs (idle cache rebuilds, repeated reads, large inserts) with $ impact and a fix for each, splitting measured (real billed rebuilds) from estimated waste. - Guard —
ctx guard: scans a session's trust boundary for secrets that entered context (.env/keys/tokens) and dangerous shell commands that ran (rm -rf /,curl … | bash,sudo). Findings are redacted — a secret is shown only as a prefix + length + sha1, never in the clear.
All read the same per-session heartbeat files that a statusline writes. The
statusline is the only piece that runs inside Claude Code; everything else is an
ordinary command you run in any terminal (or via !ctx … / /ctx inside Claude).
Claude Code session
│ (each statusline refresh)
▼
heartbeat statusline ──writes──► ~/.claude/session-status/<session>.json
~/.claude/session-status/steps/<session>.jsonl
▲ │
│ one stdout line ├──► ctx dashboard (live, all sessions)
in-app status bar └──► ctx audit (offline, one session)
Requires python3 (3.9+). The heartbeat statusline also needs jq + awk
(standard on macOS/Linux).
Install
# Recommended: isolated global install with pipx
pipx install ctxprof # from PyPI
# or straight from the git repo:
pipx install "git+https://github.com/kraftaa/ctxprof"
# Or plain pip (ideally in a venv)
pip install ctxprof
# Or straight from a clone, editable
pip install -e .
This installs one command, ctx, an umbrella with subcommands:
dashboard, sessions, watch, explain, show, steps, digest, compare, rates,
audit, guard, attack-path, tui, install (plus statusline/hook
plumbing). Run ctx with no arguments to see them all.
No clone needed after install. To run without installing, use
python3 -m claude_context_tools.cli <subcommand> from this directory.
When to run what
| I want to… | Run |
|---|---|
| See what all my open sessions cost right now | ctx dashboard |
| Pick the right session id before drilling in | ctx sessions |
| Find which session is the money sink | ctx digest (ranks sessions by cost) |
| Compare cache reuse / waste across many sessions | ctx compare --since 30d |
| Know why a session got expensive + how to fix it | ctx explain <id> |
| Watch one session live while I work (split pane) | ctx watch <id> |
| See per-turn costs / catch a spike as it happens | ctx steps |
| Deep offline breakdown of one session's transcript | ctx audit <id> |
| Scan a session for secrets in context / dangerous commands | ctx guard <id> |
| See which MCP servers are risky / the attack surface | ctx guard --mcp · ctx attack-path |
| Check prices / keep-warm-vs-rebuild math | ctx rates |
Typical flow: ctx digest to see which session is costing the most →
ctx explain <id> on that one to see the avoidable costs and the fix.
What ctxprof does not try to solve
ctxprof is a context/cache profiler, not a full agent-judgment engine. It
shows what entered context, what repeated, what rebuilt cache, and what it cost.
A higher-level analyzer could consume ctxprof's signals to answer different
questions: whether the agent strategy made sense, whether it looped, and whether
the session achieved the user's objective:
-
Decision analysis — why did Claude read 70 files, run
pytest12 times, or choose one tool instead of another? Which actions were evidence-gathering, which were implementation, and which were avoidable churn? -
Agent loop detection — recognize repeated action sequences such as:
read A grep B pytest read A grep B pytest read A grep B pytest
and flag them as cycles rather than just reporting each read or test run as an isolated cost.
-
Goal completion analysis — summarize the session as work against an objective, for example:
Task: Fix failing test Time: 25 min Useful work: 2 edits Exploration: 85% Result: test still failing
These questions are deliberately outside ctxprof's scope. The profiler supplies
the observability signals: tool calls, read/grep/test repetition, per-turn costs,
cache behavior, and transcript evidence. A decision/goal analyzer would use those
signals to judge strategy, progress, loops, and outcome quality.
Which session do the single-session commands act on? With no id they pick the
newest (most recently active) session, and the header line shows its name
so you can confirm. To target another, pass a session id or prefix — get the
ids from ctx sessions (or the ID column in ctx dashboard) and run ctx show <id> to see a
session's repo/cwd. No time window: ctx explain/audit cover the whole
session; ctx digest --since 2h and ctx steps --limit N are the windowed ones.
Use it
# Live table of all active sessions (Ctrl-C to quit):
ctx dashboard
# Print once and exit (good for tmux panes / scripts):
ctx dashboard --refresh 0
# Include sessions idle past the stale timeout (shown dimmed, with `!`):
ctx dashboard --include-stale
# List sessions for choosing a target id:
ctx sessions
ctx sessions --include-stale
ctx sessions --json
# Details + a ready-to-run audit command for one session (or the newest):
ctx show
ctx show <session-id-or-prefix>
# Live single-session panel + recommendations (the "sidebar"; run in a split pane):
ctx watch
ctx watch <session-id>
# Rank a session's biggest avoidable costs (the "profiler" view):
ctx explain
ctx explain <session-id> --top 5
ctx explain looks like:
ctx explain Make Claude workflow utility installable … — top avoidable costs
1. $ 51.29 Cache rebuilds after idle >5min (×15)
→ rebuilt ~7.2M cache tok — /clear before breaks; keep sessions short
2. $ 0.02 ~est Large Read result (~35.2k tok)
→ summarize/cap big output before it enters context
3. $ 0.01 ~est Re-reading dashboard.py 27×
→ ~12.5k redundant tok — read targeted ranges / keep a summary
Total avoidable: ~$51.32 ($51.29 measured + $0.03 estimated)
# Recent per-turn cost feed across all sessions (catches expensive turns):
ctx steps
ctx steps --limit 40
ctx steps --json --limit 200 # machine-readable (feed to Claude to analyze)
# Rollup summary: totals, top-cost turns, biggest cache writes, longest gaps:
ctx digest --since 2h
# Price table + keep-warm-vs-rebuild economics for the current context:
ctx rates
ctx rates --context 100000 --model "Opus 4.8"
# (edit prices without touching code via ~/.claude/ctx-pricing.json)
# Interactive browser (scroll/filter/drill into a turn):
ctx tui
Inside Claude Code
Prefix any command with ! in the Claude prompt to run it and drop the output
into the conversation, or use the bundled /ctx slash command:
!ctx digest --since 2h
/ctx steps
/ctx steps --json --limit 200 # then ask Claude to summarize the big turns
# Cross-session learning (cache reuse, rebuild waste, spend by repo):
ctx compare --since 30d
ctx compare --since 7d --deep # also scan transcripts for top wasted files (slower)
ctx compare --json
# Offline "why did it burn?" audit:
ctx audit --latest
ctx audit --session <session-id>
ctx audit --transcript ~/.claude/projects/<proj>/<session>.jsonl
ctx audit --latest --json # machine-readable
# Security: secrets in context / dangerous commands (redacted output):
ctx guard --latest
ctx guard --session <session-id> --json
ctx guard --latest --strict # exit non-zero if anything is found (CI gating)
ctx guard --latest --mcp # MCP server risk surface (catalog + observed)
ctx attack-path --latest # potential injection→action→exfil reachability
Turn on the heartbeat (one-time)
The dashboard and audit need a statusline that writes heartbeat files. Claude
Code supports exactly one statusLine, so this never silently replaces an
existing one:
ctx install --statusline # safe: refuses if a statusLine exists
ctx install --statusline --force # replace it (writes a timestamped backup)
ctx install --print-path # just print the heartbeat script path
If you already have a statusline you like, keep it — just make it also write the
~/.claude/session-status/<session>.json and steps/<session>.jsonl shapes (see
claude_context_tools/data/claude-statusline-heartbeat.sh for the exact fields)
and these tools will read it. To wire it by hand instead:
{
"statusLine": {
"type": "command",
"command": "bash /ABSOLUTE/PATH/TO/claude-statusline-heartbeat.sh"
}
}
Global vs per-repo
- Global (all sessions): set the heartbeat in
~/.claude/settings.json. This is whatctx install --statuslinedoes. Usual choice. - Per-repo: set
statusLinein a project's.claude/settings.json; it wins for sessions in that repo. Use an absolute path to the heartbeat script.
What the audit reports
- What burned context — per-category token estimates (assistant/user text, thinking, tool inputs, per-tool results, attachments).
- Attachment breakdown — system-injected context (skill listings, hook output, deferred tool lists) totalled separately.
- Repeated file reads / commands / duplicate output blobs — content paid for more than once.
- Loaded but never referenced again — files read into context whose path never reappears afterward (a lower bound on dead context: the model can use a file's content without naming its path, so this under-reports, never over-claims).
- Large pasted logs / large tool results and subagent report weight.
- Cache telemetry — per-turn behavior, separating expected initial warm-up from likely mid-session invalidation (cache was warm, then largely rewritten).
- Recommendations derived from the findings.
Honest limits
- Token counts in
ctx auditare estimates (~4 chars/token, scaled by the model's tokenizer — ×1.35 for Opus 4.7+; override with--token-factor). The live per-turn numbers (statusline,ctx steps, COST) are Claude Code's own exact usage fields, not estimates. - Calibration.
ctx auditmeasures the real chars/token ratio from the cache-free output side (generated chars ÷ output tokens) and prints it next to the default. When the session used subagents (their output bills tokens but isn't in the transcript) or redacted thinking, the ratio is confounded — so it says unavailable and why, rather than print a misleading number. - Cache analysis is heuristic and token-level — the heartbeat exposes per-turn token counts, not cache-block boundaries, so it flags likely invalidation, not exact cache-block attribution.
- Unknown is not zero. With no step file, the audit says per-turn cache
behavior is unknown rather than reporting
0.
What ctx guard reports
A session trust-boundary scan (Tier 1), not a repo secret scanner — it looks at what actually crossed into this conversation:
- Secrets in context — known-format keys/tokens (AWS
AKIA…, private-key blocks, GitHub/Slack/Google/Stripe tokens, bearer tokens) andkey = valuesecret assignments found in tool results, file reads, pasted text, or tool inputs. - Sensitive file reads —
.env,*.pem/*.key, SSH private keys, cloudcredentials/.netrc/.npmrc, kube/service-account configs. - Dangerous commands —
curl … | bash,rm -rf /(or~), fork bombs,chmod 777, writes to block devices,sudo(LOW). Severity-ranked HIGH/MED/LOW. Pure-data carriers (echo/grep/printf) are suppressed — a dangerous string quoted as their argument isn't an action. - Potential taint — untrusted content (a
WebFetch/WebSearch, or a file read outside the repo) followed within a few turns by a shell command or file edit. Low/medium confidence by design (see limits). - Prompt-injection phrases — jailbreak/override phrases ("ignore previous instructions", role reassignment, model control tokens, "do not tell the user") found in ingested content only (tool results — files read, pages fetched, command output), never the user's own messages. LOW severity, deduped.
Honest limits
- Redacted by design. A matched secret is never printed — only a 3-char
prefix, its length, and a sha1 fingerprint (so duplicates collapse) — so the
report can't become the leak it warns about. Placeholder-looking values
(
your_password,<token>,example…) are suppressed. - Pattern-based, so it has both misses and false positives. It finds known secret formats and named dangerous commands; a novel token shape or an obfuscated command can slip through, and a real-looking string in sample data can over-trigger. Treat findings as leads to verify, not verdicts.
- Taint is correlation, not proof. It only knows untrusted content arrived before an action, not that the action was influenced by it — so taint findings are LOW/MED and show the turn distance, as leads to review. Reads are treated as untrusted only when the repo root is known; otherwise taint stays silent on reads rather than guess.
- Injection detection is intentionally noisy and LOW. These phrases occur in
legitimate security docs and prompt-engineering material (including this repo's
own
guard.py), so a hit means "a human should glance at this," not "you were attacked." It scans ingested content only, and dedupes repeats.
MCP risk + attack path
ctx guard --mcp— the MCP server risk surface. Servers come from~/.claude.json(user + project scope) and.mcp.json; capabilities are reported on two clearly-labelled bases:[catalog](what a known server can do — an assumption from a built-in list) and[observed](classified from themcp__server__toolcalls actually made this session). A server that's neither known nor used reports riskunknown, never a guessed score. Remote/claude.ai connectors not in local config still appear if their tools were called.ctx attack-path— chains the detected signals (injection → untrusted input → capable MCP → shell use → credentials in context) into a potential reachability surface with an overall HIGH/MED/LOW. Every node is a real detection from the lenses above; it is a surface map, not proof of an exploit.
What ctx compare reports
Cross-session learning over a window (--since 30d, or all recorded sessions):
- Cache reuse — both token-weighted (read ÷ cacheable across all turns) and the per-session average, plus the best and worst sessions by name. A gap between the two means your big sessions reuse cache worse than your small ones.
- Idle cache rebuilds — turns that rewrote a large window with little/no read
(the same
is_rebuildruleexplain/watchuse), totalled with their cost — your biggest avoidable cross-session pattern. - Lowest-reuse sessions and spend by repo — where to focus.
--deepadditionally scans each session's transcript for the top wasted files (repeated reads + loaded-but-unused) aggregated across sessions.
Honest limits
- Reuse is summed from per-turn deltas (same basis as
ctx audit), so a session with no step file is absent, not counted as 0% — unknown is not zero. The rated-session count is reported separately. --deeptoken figures are~chars/4estimates; it can only scan transcripts still present on disk (it reports how many it skipped).
Dashboard notes
- LABEL is the session name/task (from
session_name) so you identify sessions by what they're doing, not by UUID. ID is a short id prefix (enough forctx show <prefix>). - CONTEXT is a fill bar + percent; RECENT is a per-session sparkline of the last few turns' token volume, colored by cache-read share (green = cheap cache hits, cyan = fresh input / rebuild) so big-write rebuild turns show as expensive, not green.
- The table is responsive: on narrow terminals lower-priority columns (MODE, AGENT, 5H/7D, IN/OUT, API, DUR, CHG) drop out, always keeping REPO, the CONTEXT bar, COST, RECENT and LABEL.
ctx stepsis a separate cross-session per-turn feed (STEP, CTX, tokens, CACHE R/W, COST, APIΔ, GAP) — costs over $0.25/$1.00 a turn are highlighted.- Stale sessions are hidden by default; the header shows
stale-hidden:Nso you know they exist.--include-staleshows them dimmed with an!. The cutoff isCLAUDE_STATUS_STALE_AFTER_SECONDS(default 900s). An open session that's idle (not generating turns) stops writing heartbeats and goes stale. - COST and TIME are cumulative session totals; live-ctx / IN / OUT are the current context, not lifetime. An old session legitimately shows a large frozen cost/time — its lifetime total at the last heartbeat, not ongoing spend.
Files written
| Path | Written by | Used by |
|---|---|---|
~/.claude/session-status/<session>.json |
heartbeat statusline | dashboard, audit |
~/.claude/session-status/steps/<session>.jsonl |
heartbeat statusline | audit (cache telemetry) |
Claude transcript JSONL (under ~/.claude/projects/) |
Claude Code | audit |
Override the state dir with CLAUDE_STATUS_STATE_DIR. Nothing is deleted
automatically. These tools only ever read your transcripts — they never
modify or upload them.
Tests
Fixture-based smoke tests, stdlib only (synthetic data, generated by
tests/make_fixtures.py — no private content):
python3 tests/test_claude_context_tools.py
python3 tests/make_fixtures.py # regenerate fixtures after changing shapes
See ../03-harnesses.md → "Status line" for how this fits
the broader harness guidance.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file ctxprof-0.2.0.tar.gz.
File metadata
- Download URL: ctxprof-0.2.0.tar.gz
- Upload date:
- Size: 66.7 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
a5fa633e2c849dc65aa6156480940caa6c0c79d9f0ffbe2743192c6ef73fe0c8
|
|
| MD5 |
828e4a2941bcd7273b1fdc847d1723f7
|
|
| BLAKE2b-256 |
5bc9e3675742bcd18e12768e9ba5e6bdfd9bcfc9b71af7537c4d9f32a8af9b24
|
Provenance
The following attestation bundles were made for ctxprof-0.2.0.tar.gz:
Publisher:
release.yml on kraftaa/ctxprof
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
ctxprof-0.2.0.tar.gz -
Subject digest:
a5fa633e2c849dc65aa6156480940caa6c0c79d9f0ffbe2743192c6ef73fe0c8 - Sigstore transparency entry: 2138318254
- Sigstore integration time:
-
Permalink:
kraftaa/ctxprof@bcf5dcb887978ce5936c53d05a2c8052629f1b29 -
Branch / Tag:
refs/tags/v0.2.0 - Owner: https://github.com/kraftaa
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@bcf5dcb887978ce5936c53d05a2c8052629f1b29 -
Trigger Event:
push
-
Statement type:
File details
Details for the file ctxprof-0.2.0-py3-none-any.whl.
File metadata
- Download URL: ctxprof-0.2.0-py3-none-any.whl
- Upload date:
- Size: 62.4 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
42d68a2067e8d737f7df0164ca7e413c3e8ff2cf3f35eda82d892e1eb48bf389
|
|
| MD5 |
3f7b855a33b7060c9b509f36f120fd00
|
|
| BLAKE2b-256 |
52e8ed33350657dea8677d425a09b38e70163933a932174811a80662a8a9e7c7
|
Provenance
The following attestation bundles were made for ctxprof-0.2.0-py3-none-any.whl:
Publisher:
release.yml on kraftaa/ctxprof
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
ctxprof-0.2.0-py3-none-any.whl -
Subject digest:
42d68a2067e8d737f7df0164ca7e413c3e8ff2cf3f35eda82d892e1eb48bf389 - Sigstore transparency entry: 2138319232
- Sigstore integration time:
-
Permalink:
kraftaa/ctxprof@bcf5dcb887978ce5936c53d05a2c8052629f1b29 -
Branch / Tag:
refs/tags/v0.2.0 - Owner: https://github.com/kraftaa
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@bcf5dcb887978ce5936c53d05a2c8052629f1b29 -
Trigger Event:
push
-
Statement type: