trainmeter
Where your training FLOPs go, and how well your GPUs are used.
You rented a GPU node, your loss is going down, and nvidia-smi says 100%.
That number does not tell you whether the run is any good.
A loop at 100% GPU-Util can sit at 20% MFU, the MFU you computed by hand can be off by 15% depending on what you counted, and a 2:4 sparse peak quietly halves it.
trainmeter tells you the real number while the run goes.
Run tm train.py instead of python train.py, open the dashboard, and see progress, ETA, MFU with its convention spelled out, how busy the tensor cores really are, and where the time goes.
Your training code stays as it is.
A real run, replayed with tm view: a 111M transformer trained on 2.27B tokens in 1.26 h on one H100 SXM, at 533k tokens per second and 92% of the wall time in training steps.
Quick start
Install trainmeter into the same environment as your training code, next to torch:
uv add "trainmeter[gpu]" # or: pip install "trainmeter[gpu]"
Then launch your run through tm:
tm train.py # instead of: python train.py
tm --budget-tokens 20B train.py # tell it the plan, get progress and ETA
tm -- torchrun --nproc_per_node=8 train.py # any command works after --
tm prints one line with the dashboard URL, http://127.0.0.1:8765/ by default.
The dashboard listens on the node only, so from your laptop open a tunnel first:
ssh -L 8765:localhost:8765 your-gpu-node
When the run ends, tm prints one summary line and leaves the full run in .trainmeter/runs/.
No GPU at hand?
examples/ has a tiny GPT that trains under tm on a laptop in half a minute, and a run planner that needs no GPU at all.
What you see
- Progress and ETA. Tokens and FLOPs done against the budget, and when the run will finish at the current pace.
- MFU with its convention. Model FLOPs utilization under a named FLOPs convention, divided by the dense peak of your card, with the datasheet that peak comes from.
- OFU next to MFU. What the tensor cores actually did, read from the GPU's own counters (NVML GPM). A gap between the two is a signal: recompute, padding, or a wrong model shape.
- The utilization ladder. From
GPU-Utildown through SM activity and tensor activity to MFU, so you see at which rung the GPU stops doing useful work. - Where the time goes. Data loader waits, checkpoint saves and eval, as a share of the run, and goodput.
- Your loss curves. Whatever your loop already logs to W&B, TensorBoard or MLflow is picked up as is. A loop with no tracker adds one line:
emit(train_loss=loss.item()). - The node. CPU, memory, disks, network and temperatures, so a starving data loader or a full disk does not go unnoticed.
- A run passport. At exit,
passport.mdandpassport.jsonwith the numbers a model card needs.
The same run one screen down: the GPU was busy 95% of the time, the tensor pipes 44% of it, and MFU under the titan convention is 34% of the dense peak.
Every number can be traced to where it came from: declared by you, emitted by the loop, picked up from a tracker, detected, measured or assumed. The button beside a number opens how it was computed, the formula and every input with its value and provenance, so you can redo MFU by hand or copy the derivation into a notebook. A number that lacks an input is not shown, and the dashboard says which input would enable it.
Why another tool
The answer needs three sources joined, and today each tool sees only one of them.
| Tool | What it gives | What it misses |
|---|---|---|
nvidia-smi, nvtop |
GPU-Util, memory, power | 100% Util is possible at 20% MFU |
| W&B, TensorBoard | The loss and what the loop logs | No MFU, no FLOPs, no GPU counters |
| A notebook calculation | MFU from the 6N formula | The convention is implicit, the peak is often the sparse one |
torch.profiler, Nsight |
Every kernel of a few steps | Heavy, and not the whole run |
| DCGM, Prometheus | Counters on a cluster | Needs root and infrastructure of your own |
trainmeter joins the training loop (loss, tokens, budget), the model (FLOPs per token under a named convention) and the GPU counters (tensor, DRAM and SM activity) into one picture, with no edit to train.py, no root and no network.
Why the convention matters
On a 1.04B model with d 2048, 18 layers and context 1024, FLOPs per token are 5.83 G by the Hugging Face count, 6.06 G by Megatron, 6.28 G by TorchTitan and nanochat, and 6.69 G when embeddings are counted as matmuls. The same run reads as anything from 36.8% to 42.2% MFU depending on which one you pick. trainmeter always says which convention it used, and when you have not chosen one it shows the range instead of guessing.
Commands
tm train.py # run and watch
tm --name baseline --budget-steps 5000 --convention titan train.py
tm attach 12345 # watch a program that is already running, from outside
tm view # replay the latest finished run in the dashboard
tm passport # the model-card passport of the latest run
tm doctor # what tm can see on this node
tm doctor --dump node.tar.gz # pack what it reads, for a bug report
tm doctor --python PATH # can the agent join the Python you train with?
python -m trainmeter.measure_peak # measure the dense BF16 peak where no datasheet gives one
tm --help lists every flag, including the model shape (--layers, --d-model, --seq-len) for when the agent cannot see it, and --plan-mfu and --plan-hours to compare the run with your plan.
Supported hardware
MFU needs the dense BF16 peak of the card, and trainmeter only uses peaks it can cite.
It ships spec files for the A100 SXM 80GB, H100 SXM and PCIe, H200, B200, GB200, and GB10 (DGX Spark), each next to the datasheet lines its figures come from.
On a card it does not know, MFU is not shown, rather than borrowed from another card.
Where no datasheet gives a dense figure, as on the GB10, measure it once with python -m trainmeter.measure_peak or declare one with --peak-tflops and --peak-source, and MFU is labeled as such.
The GPU counters behind OFU and the utilization ladder need NVML GPM, which NVIDIA provides on Hopper and newer data center GPUs. Everything else works on any machine, including a Mac or a CPU-only box, where the dashboard says why each GPU number is missing.
How it stays out of your way
- No code change.
tmstarts your command as a child process and injects a small passive agent into the training interpreter, which sees steps, batch shapes, parameter counts and data waits. - It never breaks your run. Nothing trainmeter puts inside the training process raises into it; a monitoring bug warns once and goes quiet.
- It never changes your numbers. The agent never synchronizes a device, never allocates on a GPU, and
tmnever creates a CUDA context. A test trains a model with and withouttmand compares the loss curves exactly. - Physics, not prices. The run log holds time, tokens, FLOPs and counters. Dollars and energy are computed at view time, and a price is never an input.
Status
trainmeter is young, version 0.0.x, and is used on the author's own training runs first.
The formulas are pinned by tests against hand-computed values, and the GPU code is tested against fakes.
It runs on a DGX Spark, and its first H100 run, on 2026-10-05, went through most of the checks that only a real node can make.
What that run confirmed and the seven things it found to fix are in docs/hardware-checklist.md.
What comes next, and why, is in docs/roadmap.md.
Bug reports are welcome, and tm doctor --dump packs what trainmeter reads on your node so the report can become a test.
Learn more
docs/spec.mddefines every metric and its formula.docs/architecture.mdexplains how the parts fit together.docs/users.mddescribes who trainmeter is for.
Develop
uv sync
uv run pytest
uv run ruff check . && uv run ruff format --check .
CLAUDE.md lists the invariants every change has to keep.
License
Apache-2.0.
Metadata
Release files for trainmeter 0.0.6
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| trainmeter-0.0.6.tar.gz | 1.6 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| trainmeter-0.0.6-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 1.8 MB
Release files / trainmeter-0.0.6.tar.gz
| Download URL | trainmeter-0.0.6.tar.gz |
|---|---|
| Size | 1.6 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
8ee6a7a7119ac31b9a44eb834ca4fccb90cb20a9d2076201447bca739a2aac89
|
|
BLAKE2b-256 checksum How to use checksums |
fd10b6b8c5472c4d852102c336aaa488aad5d48f239db883fff3b5b4ddc8b0b7
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.12.23 {"installer":{"name":"uv","version":"0.12.23","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|
Release files / trainmeter-0.0.6-py3-none-any.whl
| Download URL | trainmeter-0.0.6-py3-none-any.whl |
|---|---|
| Size | 198.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
8ead649010651e7c43ba5f6923b128e479b2e17288146a504a0ae9ada7af5b75
|
|
BLAKE2b-256 checksum How to use checksums |
549c064da3220a0447b83a8183d3fa7be25713f90004a051eae813eb805da8f8
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.12.23 {"installer":{"name":"uv","version":"0.12.23","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|