Skip to main content

trainmeter

Where your training FLOPs go, and how well your GPUs are used.

Run tm train.py instead of python train.py. trainmeter starts your command, watches the node from outside, and serves a live dashboard on localhost. The dashboard shows what a training run needs to be judged: training and validation loss, FLOPs spent and required, MFU with its FLOPs convention spelled out, and how busy the GPUs really are, from GPU-Util down to MFU, next to arithmetic intensity. Your training code stays as it is.

The trainmeter dashboard: progress, throughput, MFU and OFU, goodput and the loss curve of a run on four H100s

The dashboard on a synthetic run.

Status: early scaffold. tm runs any command untouched, records the run, and turns what a PyTorch loop does into tokens, FLOPs spent and required, ETA, MFU (given a device) and the losses you emit(). GPU counters (NVML and GPM: the utilization ladder, OFU, observed FLOPs, arithmetic intensity) and tm doctor are built and tested against fakes, not yet on a real GPU. Series your loop already logs to wandb, TensorBoard or MLflow are picked up without a code change (--no-tap turns it off), and DataLoader waits, checkpoint saves and eval spans feed a "where the time goes" view. One real step is also counted with FlopCounterMode (shown beside MFU as the counted convention), and layers and width are read from a model config when there is one. Each run ends with a run passport (passport.md and passport.json) for a model card. The dashboard, one page of progress, tiles and charts with the run details and model card folded away, is served on localhost while the run goes; it has been checked in a browser on synthetic data, not on a real run. Not yet used on a real run. To try it, examples/ has a tiny GPT to run under tm on a laptop and a run planner that needs no GPU.

Install

trainmeter is on PyPI:

# into the training environment, next to torch: this is what lets tm see inside the program
uv add "trainmeter[gpu]"
# or as a standalone tool, enough for tm view, tm doctor and tm passport
uv tool install trainmeter

The gpu extra pulls in nvidia-ml-py for the GPU counters and can be left out on a machine with no NVIDIA GPU. Without uv, pip install "trainmeter[gpu]" does the same.

How it will be used

tm train.py                                 # instead of: python train.py
tm --budget-tokens 20B train.py             # declare the plan: FLOPs required, progress, ETA
tm -- torchrun --nproc_per_node=8 train.py  # any command, one grammar
tm doctor                                   # what tm can see on this node
tm passport                                 # model-card passport of the latest run
tm export wandb --project p                 # upload a finished run to W&B
tm view                                     # replay the latest finished run in the dashboard
tm attach 12345                             # watch a program that is already running, from outside

tm prints one line with the dashboard URL and, at the end, one summary line. The dashboard binds to loopback on the node, so a remote node needs an SSH tunnel (ssh -L 8765:localhost:8765 node). --port N picks the first port to try and --no-web turns the dashboard off.

How it works

  • From outside. tm starts your command as a child process, reads the GPU counters (tensor, DRAM and SM activity via NVML GPM on Hopper and newer), and keeps the run log.
  • From inside, passively. A small agent that tm injects into the training interpreter sees steps, batch shapes, parameter counts and data waits. It never synchronizes a device and never raises into your loop.
  • From your loggers. The loss and any other series your loop already sends to wandb, TensorBoard or MLflow are picked up as they are. A loop with no tracker adds one emit(train_loss=...) line.
  • Physics, not prices. The run log holds time, tokens, FLOPs, counters and the series your loop reports. Money is an optional overlay computed at view time, and a price is never an input.

docs/architecture.md has the design, and docs/spec.md defines every metric.

Why the convention matters

On a 1.04B model with d 2048, 18 layers and context 1024, FLOPs per token are 5.83 G by the Hugging Face count, 6.06 G by Megatron, 6.28 G by TorchTitan and nanochat, and 6.69 G when embeddings are counted as matmuls. Same run, MFU differing by 15%. trainmeter always says which one it used.

Develop

uv sync
uv run pytest
uv run ruff check .
uv run ruff format --check .

See docs/spec.md for what trainmeter shows and records, docs/architecture.md for how it is built, and CLAUDE.md for the invariants.

License

Apache-2.0.

Metadata

Release files for trainmeter 0.0.3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for trainmeter 0.0.3
File Size Uploaded
trainmeter-0.0.3.tar.gz 693.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for trainmeter 0.0.3
File Interpreter ABI Platform
trainmeter-0.0.3-py3-none-any.whl Python 3 none any Details

Total release size: 784.9 kB

Release files / trainmeter-0.0.3.tar.gz

Download URL trainmeter-0.0.3.tar.gz
Size 693.6 kB
Tags Source
SHA-256 checksum
How to use checksums
87abddb2d3eb612d822a68fc172c05a1d16ecd2d28420206937fe8ec486efe7c
BLAKE2b-256 checksum
How to use checksums
1e63bebc4e1dcfd094f3c2b9133c78728d1cade188e7eb96626889a502995fb9
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.21 {"installer":{"name":"uv","version":"0.12.21","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release files / trainmeter-0.0.3-py3-none-any.whl

Download URL trainmeter-0.0.3-py3-none-any.whl
Size 91.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
82efa29618ef4df1f1d2f7244991171a7bacf58ead170b5d473ef00baba800a3
BLAKE2b-256 checksum
How to use checksums
43607e447c3444412af07677fd7e935c36ebd078b2f7ec116df50ae54eb987a0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.21 {"installer":{"name":"uv","version":"0.12.21","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release history Release notifications | RSS feed

0.0.6

2 release files

0.0.5

2 release files

0.0.4

2 release files

This release

0.0.3 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page