Skip to main content

TokenMon

A fast, zero-dependency command-line monitor for local AI coding agents.

It measures how fast models actually generate tokens (Tokens Per Second, TPS) by tracking pure generation time—separating thinking and output from tool runs, file edits, and idle waiting.

Recent generation speeds and session activity


Screenshots

Expand a screenshot below. Click the image to view it at full size.

stats — generation speed and token throughput

Generation statistics

ps — sessions, activity, and token usage

Session overview

logs — chronological session timeline

Session timeline


Features

  • Generation Speed: Measures TPS from recorded output boundaries and labels timing fallbacks, including Claude prompt-to-response estimates.
  • Zero Extra Dependencies: Runs on standard Python 3.10+ without installing third-party packages.
  • Strictly Read-Only: Safely opens local files and SQLite databases in read-only mode (?mode=ro). Never locks or changes your logs.
  • Meaningful Averages: Calculates true weighted speed ($\frac{\text{total tokens}}{\text{total time}}$), median speed, and min/max ranges over rolling time windows (30m, 1d, 7d, 30d, all).
  • Supports Popular Agents: Auto-detects Codex (~/.codex), Claude Code (~/.claude/projects/), and Antigravity (~/.gemini/antigravity-cli).
  • Session Timelines: Step-by-step history of user prompts, thinking, assistant responses, and tool calls.
  • JSON Ready: Add --json to pipe clean data into jq or external dashboards.

Installation

Requires Python 3.10+.

Try it immediately with uv:

uvx tokenmon
uvx tokenmon ps
uvx tokenmon logs

For regular use, install it as an isolated tool:

uv tool install tokenmon
tokenmon

If tokenmon is not on your PATH, run uv tool update-shell and restart your shell. Upgrade with uv tool upgrade tokenmon.

Or install with pip in your Python environment:

pip install tokenmon
tokenmon

Usage

TokenMon uses simple Docker-style subcommands: stats (default), ps (sessions), logs (timelines), and interactive (shell).

Running tokenmon by itself defaults directly to stats. The examples below use the installed command; with uvx, use uvx tokenmon in its place.

Generation Speed & Metrics (stats, top, default)

View token throughput (TPS), rolling averages, and recent generation outputs:

# Auto-detect local agents and show stats (default)
tokenmon
# or explicitly
tokenmon stats

# Filter to a specific time window (30m, 1d, 7d, 30d, all)
tokenmon stats --window 1d

# Refresh the full dashboard every 2 seconds (Ctrl+C to stop)
tokenmon stats --watch
tokenmon stats -w
tokenmon stats --watch --window 1d

# Customize the refresh interval
tokenmon stats --watch --window 1d --interval 5

# Include full history beyond the default 30-day cutoff
tokenmon stats --all

# Target a specific agent or custom folder
tokenmon stats codex
tokenmon stats claude --home ~/.claude

# Compact layout for narrow panes or wide full-detail table
tokenmon stats --compact
tokenmon stats --wide

Watch mode refreshes throughput metrics, recent generation streams, and recent session cards together. It redraws the screen in a terminal; redirected output appends complete snapshots. Agent selection, --home, --tasks, --recent, --compact, and --wide also work with --watch.

-w is a shortcut for --watch on stats. Time filtering uses the explicit --window option across commands; replace the former -w 1d syntax with --window 1d in existing commands and scripts. Combine them as tokenmon stats -w --window 1d.

The wide stats table formats total generation time as hours, minutes, and seconds (for example, 13h 53m 46s). This sums durations across streams and sessions, so overlapping sessions can produce a total longer than the window's wall-clock time. JSON retains numeric duration values in seconds.

Recent streams, session lists/cards, and timelines show explicit reasoning effort and speed metadata when available:

gpt-6.1-sol medium fast : 45.9 TPS (937 tokens in 20.43s) [stream-log]

Codex reads effort from turn context and settings, and tier from recorded thread settings or response usage. Recorded priority/fast tiers display as fast, following the OpenAI fast-mode tier names. Settings describe the recorded request mode; they do not independently confirm the server's delivered tier. Claude reads explicit effort/configuration and usage tier/speed fields. Antigravity reads explicit named settings in generation metadata when present; current logs may not record them. Missing metadata is omitted in human output and becomes null in JSON's reasoning_effort, service_tier, speed, and normalized speed_mode fields. Model names and measured TPS never determine effort or fast mode.

Range is the minimum and maximum individual valid stream TPS, rather than an interval around the weighted average. A rate above 200 TPS can be valid when its tokens and timing agree. Streams shorter than one second, above 400 TPS, or with unconfirmed boundaries are excluded. Codex repeated cumulative usage snapshots are ignored so old tokens cannot inflate a later reasoning item's rate. Claude's single-record turn-span fallback includes prompt-to-response latency and should be treated as an estimate, not pure streaming speed.

All JSON views include the same optional fields: stats recent_streams and recent_sessions, ps entries, timeline session objects, and individual events (including window exports). Session fields describe the latest recorded settings; events preserve their own settings. Each stats summary and model window includes configurations, the distinct settings of valid streams in that window. Its singular metadata fields are populated only when all configurations agree on that field; mixed or missing values are null.

For example:

{
  "model": "gpt-6.1-sol",
  "reasoning_effort": "medium",
  "service_tier": "priority",
  "speed": null,
  "speed_mode": "fast"
}

Active & Recent Sessions (sessions, ps, ls)

See all recent sessions, message counts, tool runs, and idle status:

# List recent sessions
tokenmon ps

# Filter sessions within a time window
tokenmon ps --window 1d

# Filter to a specific agent
tokenmon ps codex

Event Timelines (timeline, logs, log)

See the chronological step-by-step history of prompts, model thoughts, responses, and tool calls:

# View timeline of the latest session
tokenmon logs

# View timeline of a specific session ID or prefix
tokenmon logs 01a10275

# Follow only new events from the session that is latest at startup
tokenmon logs -f

# Follow a specific session ID/prefix, or an agent's latest session
tokenmon logs 01a10275 --follow
tokenmon logs codex -f

# Customize the polling interval (default: 2 seconds)
tokenmon logs -f --interval 1

# Export all session timelines from today in JSON format
tokenmon timeline --window 1d --json

Follow mode starts at the current end: it prints a short session header, then only newly observed timeline events. It keeps the initially selected session even when another session becomes newer, and appends output without clearing the screen or replaying history. Updates to an existing event's usage or timestamp do not print that event again. Press Ctrl+C to stop.

-f and --follow work with logs, log, and timeline. Follow selects one session, so it cannot be combined with the batch export option --window; use --all to select a session older than the default 30-day cutoff. The interval must be a positive, finite number. Polling reads the selected session again rather than following raw file bytes, which also supports SQLite-backed Antigravity sessions.

Interactive Shell (interactive, repl, shell, -i)

Open an interactive terminal shell with live auto-refresh and tab completion:

tokenmon interactive
# or
tokenmon -i

Available commands inside the shell:

(tokenmon) summary 30m        # view 30m throughput table
(tokenmon) sessions           # list active & inactive sessions
(tokenmon) timeline latest    # view step-by-step event timeline
(tokenmon) recent 15          # view latest generation speeds
(tokenmon) watch 2.0 1d       # metrics, recent streams, and session cards (Ctrl+C to stop)
(tokenmon) help               # list all commands

Machine-Readable JSON Output

Stats, sessions, and historical timelines support --json for easy scripting. Live --watch and --follow modes require human output and reject --json:

tokenmon stats --json | jq .
tokenmon ps --json | jq .
tokenmon logs 01a10275 --json | jq .
tokenmon timeline --window 1d --json | jq .

Development

Clone the repository to work on TokenMon:

git clone https://github.com/quanhua92/tokenmon.git
cd tokenmon
uv run tokenmon
uv run python -m unittest discover -s tests

Adding New Adapters

TokenMon uses a base class in src/tokenmon/adapters/base.py:

class BaseAdapter(ABC):
    @property
    @abstractmethod
    def name(self) -> str: ...

    @abstractmethod
    def detect(self) -> bool: ...

    @abstractmethod
    def collect(self, max_sessions: int = 64, min_timestamp: float | None = None) -> list[GenerationSpan]: ...

    @abstractmethod
    def collect_sessions(self, max_sessions: int = 32, min_timestamp: float | None = None) -> list[SessionTimeline]: ...

To add support for a new agent (e.g. OpenCode):

  1. Create src/tokenmon/adapters/opencode.py subclassing BaseAdapter.
  2. Implement discovery (detect()), stream parsing (collect()), and timelines (collect_sessions()).
  3. Register it in src/tokenmon/adapters/__init__.py.

License

MIT

Metadata

Release files for tokenmon 0.1.3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for tokenmon 0.1.3
File Size Uploaded
tokenmon-0.1.3.tar.gz 1.8 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for tokenmon 0.1.3
File Interpreter ABI Platform
tokenmon-0.1.3-py3-none-any.whl Python 3 none any Details

Total release size: 1.8 MB

Release files / tokenmon-0.1.3.tar.gz

Download URL tokenmon-0.1.3.tar.gz
Size 1.8 MB
Tags Source
SHA-256 checksum
How to use checksums
ca1f75b14c11f5959bbbd26eb78c590b55e76b9a99a2a2135efe4e801c2cdc7c
BLAKE2b-256 checksum
How to use checksums
fe9581c012b17110570d817eeac5d13557ab63313dcad9fd86b2a1c90457deab
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.

Transparency log

Release files / tokenmon-0.1.3-py3-none-any.whl

Download URL tokenmon-0.1.3-py3-none-any.whl
Size 39.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
e95030174a8d8460126419f9a979a2fa41926b8d791280d10debee0856d91a4c
BLAKE2b-256 checksum
How to use checksums
608d66ca78719e7c29964272d60bb84e82af1bc26845564f567b4e6d396e59f2
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.

Transparency log

Release history Release notifications | RSS feed

0.2.2

2 release files

0.2.1

2 release files

0.1.6

2 release files

0.1.5

2 release files

0.1.4

2 release files

This release

0.1.3 This release

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page