trainmeter
Where your training FLOPs go, and how well your GPUs are used.
Run tm train.py instead of python train.py.
trainmeter starts your command, watches the node from outside, and serves a live dashboard on localhost.
The dashboard shows what a training run needs to be judged: training and validation loss, FLOPs spent and required, MFU with its FLOPs convention spelled out, and how busy the GPUs really are, from GPU-Util down to MFU, next to arithmetic intensity.
Your training code stays as it is.
The dashboard on a synthetic run.
Status: early scaffold.
tm runs any command untouched, records the run, and turns what a PyTorch loop does into tokens, FLOPs spent and required, ETA, MFU (given a device) and the losses you emit().
GPU counters (NVML and GPM: the utilization ladder, OFU, observed FLOPs, arithmetic intensity) and tm doctor are built and tested against fakes, not yet on a real GPU.
Series your loop already logs to wandb, TensorBoard or MLflow are picked up without a code change (--no-tap turns it off), and DataLoader waits, checkpoint saves and eval spans feed a "where the time goes" view.
One real step is also counted with FlopCounterMode (shown beside MFU as the counted convention), and layers and width are read from a model config when there is one.
Each run ends with a run passport (passport.md and passport.json) for a model card.
The dashboard, one page of progress, tiles and charts with the run details and model card folded away, is served on localhost while the run goes; it has been checked in a browser on synthetic data, not on a real run.
Not yet used on a real run.
To try it, examples/ has a tiny GPT to run under tm on a laptop and a run planner that needs no GPU.
Install
trainmeter is on PyPI:
# into the training environment, next to torch: this is what lets tm see inside the program
uv add "trainmeter[gpu]"
# or as a standalone tool, enough for tm view, tm doctor and tm passport
uv tool install trainmeter
The gpu extra pulls in nvidia-ml-py for the GPU counters and can be left out on a machine with no NVIDIA GPU.
Without uv, pip install "trainmeter[gpu]" does the same.
How it will be used
tm train.py # instead of: python train.py
tm --budget-tokens 20B train.py # declare the plan: FLOPs required, progress, ETA
tm -- torchrun --nproc_per_node=8 train.py # any command, one grammar
tm doctor # what tm can see on this node
tm passport # model-card passport of the latest run
tm export wandb --project p # upload a finished run to W&B
tm view # replay the latest finished run in the dashboard
tm attach 12345 # watch a program that is already running, from outside
tm prints one line with the dashboard URL and, at the end, one summary line.
The dashboard binds to loopback on the node, so a remote node needs an SSH tunnel (ssh -L 8765:localhost:8765 node).
--port N picks the first port to try and --no-web turns the dashboard off.
How it works
- From outside.
tmstarts your command as a child process, reads the GPU counters (tensor, DRAM and SM activity via NVML GPM on Hopper and newer), and keeps the run log. - From inside, passively.
A small agent that
tminjects into the training interpreter sees steps, batch shapes, parameter counts and data waits. It never synchronizes a device and never raises into your loop. - From your loggers.
The loss and any other series your loop already sends to wandb, TensorBoard or MLflow are picked up as they are.
A loop with no tracker adds one
emit(train_loss=...)line. - Physics, not prices. The run log holds time, tokens, FLOPs, counters and the series your loop reports. Money is an optional overlay computed at view time, and a price is never an input.
docs/architecture.md has the design, and docs/spec.md defines every metric.
Why the convention matters
On a 1.04B model with d 2048, 18 layers and context 1024, FLOPs per token are 5.83 G by the Hugging Face count, 6.06 G by Megatron, 6.28 G by TorchTitan and nanochat, and 6.69 G when embeddings are counted as matmuls. Same run, MFU differing by 15%. trainmeter always says which one it used.
Develop
uv sync
uv run pytest
uv run ruff check .
uv run ruff format --check .
See docs/spec.md for what trainmeter shows and records, docs/architecture.md for how it is built, and CLAUDE.md for the invariants.
License
Apache-2.0.
Metadata
Release files for trainmeter 0.0.3
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| trainmeter-0.0.3.tar.gz | 693.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| trainmeter-0.0.3-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 784.9 kB
Release files / trainmeter-0.0.3.tar.gz
| Download URL | trainmeter-0.0.3.tar.gz |
|---|---|
| Size | 693.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
87abddb2d3eb612d822a68fc172c05a1d16ecd2d28420206937fe8ec486efe7c
|
|
BLAKE2b-256 checksum How to use checksums |
1e63bebc4e1dcfd094f3c2b9133c78728d1cade188e7eb96626889a502995fb9
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.12.21 {"installer":{"name":"uv","version":"0.12.21","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|
Release files / trainmeter-0.0.3-py3-none-any.whl
| Download URL | trainmeter-0.0.3-py3-none-any.whl |
|---|---|
| Size | 91.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
82efa29618ef4df1f1d2f7244991171a7bacf58ead170b5d473ef00baba800a3
|
|
BLAKE2b-256 checksum How to use checksums |
43607e447c3444412af07677fd7e935c36ebd078b2f7ec116df50ae54eb987a0
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.12.21 {"installer":{"name":"uv","version":"0.12.21","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|