Skip to main content

trainmeter

Where your training FLOPs go, and how well your GPUs are used.

You rented a GPU node, your loss is going down, and nvidia-smi says 100%. That number does not tell you whether the run is any good. A loop at 100% GPU-Util can sit at 20% MFU, the MFU you computed by hand can be off by 15% depending on what you counted, and a 2:4 sparse peak quietly halves it.

trainmeter tells you the real number while the run goes. Run tm train.py instead of python train.py, open the dashboard, and see progress, ETA, MFU with its convention spelled out, how busy the tensor cores really are, and where the time goes. Your training code stays as it is.

The trainmeter dashboard on a real run: progress, throughput, MFU and OFU, goodput and the loss curve of a 111M model trained on one H100

A real run, replayed with tm view: a 111M transformer trained on 2.27B tokens in 1.26 h on one H100 SXM, at 533k tokens per second and 92% of the wall time in training steps.

Quick start

Install trainmeter into the same environment as your training code, next to torch:

uv add "trainmeter[gpu]"          # or: pip install "trainmeter[gpu]"

Then launch your run through tm:

tm train.py                                # instead of: python train.py
tm --budget-tokens 20B train.py            # tell it the plan, get progress and ETA
tm -- torchrun --nproc_per_node=8 train.py # any command works after --

tm prints one line with the dashboard URL, http://127.0.0.1:8765/ by default. The dashboard listens on the node only, so from your laptop open a tunnel first:

ssh -L 8765:localhost:8765 your-gpu-node

When the run ends, tm prints one summary line and leaves the full run in .trainmeter/runs/.

No GPU at hand? examples/ has a tiny GPT that trains under tm on a laptop in half a minute, and a run planner that needs no GPU at all.

What you see

  • Progress and ETA. Tokens and FLOPs done against the budget, and when the run will finish at the current pace.
  • MFU with its convention. Model FLOPs utilization under a named FLOPs convention, divided by the dense peak of your card, with the datasheet that peak comes from.
  • OFU next to MFU. What the tensor cores actually did, read from the GPU's own counters (NVML GPM). A gap between the two is a signal: recompute, padding, or a wrong model shape.
  • The utilization ladder. From GPU-Util down through SM activity and tensor activity to MFU, so you see at which rung the GPU stops doing useful work.
  • Where the time goes. Data loader waits, checkpoint saves and eval, as a share of the run, and goodput.
  • Your loss curves. Whatever your loop already logs to W&B, TensorBoard or MLflow is picked up as is. A loop with no tracker adds one line: emit(train_loss=loss.item()).
  • The node. CPU, memory, disks, network and temperatures, so a starving data loader or a full disk does not go unnoticed.
  • A run passport. At exit, passport.md and passport.json with the numbers a model card needs.
The utilization ladder from GPU-Util down to MFU, the ladder over time, the roofline and the arithmetic intensity of the same H100 run

The same run one screen down: the GPU was busy 95% of the time, the tensor pipes 44% of it, and MFU under the titan convention is 34% of the dense peak.

Every number can be traced to where it came from: declared by you, emitted by the loop, picked up from a tracker, detected, measured or assumed. The button beside a number opens how it was computed, the formula and every input with its value and provenance, so you can redo MFU by hand or copy the derivation into a notebook. A number that lacks an input is not shown, and the dashboard says which input would enable it.

Why another tool

The answer needs three sources joined, and today each tool sees only one of them.

Tool What it gives What it misses
nvidia-smi, nvtop GPU-Util, memory, power 100% Util is possible at 20% MFU
W&B, TensorBoard The loss and what the loop logs No MFU, no FLOPs, no GPU counters
A notebook calculation MFU from the 6N formula The convention is implicit, the peak is often the sparse one
torch.profiler, Nsight Every kernel of a few steps Heavy, and not the whole run
DCGM, Prometheus Counters on a cluster Needs root and infrastructure of your own

trainmeter joins the training loop (loss, tokens, budget), the model (FLOPs per token under a named convention) and the GPU counters (tensor, DRAM and SM activity) into one picture, with no edit to train.py, no root and no network.

Why the convention matters

On a 1.04B model with d 2048, 18 layers and context 1024, FLOPs per token are 5.83 G by the Hugging Face count, 6.06 G by Megatron, 6.28 G by TorchTitan and nanochat, and 6.69 G when embeddings are counted as matmuls. The same run reads as anything from 36.8% to 42.2% MFU depending on which one you pick. trainmeter always says which convention it used, and when you have not chosen one it shows the range instead of guessing.

Commands

tm train.py                        # run and watch
tm --name baseline --budget-steps 5000 --convention titan train.py
tm attach 12345                    # watch a program that is already running, from outside
tm view                            # replay the latest finished run in the dashboard
tm passport                        # the model-card passport of the latest run
tm doctor                          # what tm can see on this node
tm doctor --dump node.tar.gz       # pack what it reads, for a bug report
tm doctor --python PATH            # can the agent join the Python you train with?
python -m trainmeter.measure_peak  # measure the dense BF16 peak where no datasheet gives one

tm --help lists every flag, including the model shape (--layers, --d-model, --seq-len) for when the agent cannot see it, and --plan-mfu and --plan-hours to compare the run with your plan.

Supported hardware

MFU needs the dense BF16 peak of the card, and trainmeter only uses peaks it can cite. It ships spec files for the A100 SXM 80GB, H100 SXM and PCIe, H200, B200, GB200, and GB10 (DGX Spark), each next to the datasheet lines its figures come from. On a card it does not know, MFU is not shown, rather than borrowed from another card. Where no datasheet gives a dense figure, as on the GB10, measure it once with python -m trainmeter.measure_peak or declare one with --peak-tflops and --peak-source, and MFU is labeled as such.

The GPU counters behind OFU and the utilization ladder need NVML GPM, which NVIDIA provides on Hopper and newer data center GPUs. Everything else works on any machine, including a Mac or a CPU-only box, where the dashboard says why each GPU number is missing.

How it stays out of your way

  • No code change. tm starts your command as a child process and injects a small passive agent into the training interpreter, which sees steps, batch shapes, parameter counts and data waits.
  • It never breaks your run. Nothing trainmeter puts inside the training process raises into it; a monitoring bug warns once and goes quiet.
  • It never changes your numbers. The agent never synchronizes a device, never allocates on a GPU, and tm never creates a CUDA context. A test trains a model with and without tm and compares the loss curves exactly.
  • Physics, not prices. The run log holds time, tokens, FLOPs and counters. Dollars and energy are computed at view time, and a price is never an input.

Status

trainmeter is young, version 0.0.x, and is used on the author's own training runs first. The formulas are pinned by tests against hand-computed values, and the GPU code is tested against fakes. It runs on a DGX Spark, and its first H100 run, on 2026-10-05, went through most of the checks that only a real node can make. What that run confirmed and the seven things it found to fix are in docs/hardware-checklist.md. What comes next, and why, is in docs/roadmap.md. Bug reports are welcome, and tm doctor --dump packs what trainmeter reads on your node so the report can become a test.

Learn more

Develop

uv sync
uv run pytest
uv run ruff check . && uv run ruff format --check .

CLAUDE.md lists the invariants every change has to keep.

License

Apache-2.0.

Metadata

Release files for trainmeter 0.0.6

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for trainmeter 0.0.6
File Size Uploaded
trainmeter-0.0.6.tar.gz 1.6 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for trainmeter 0.0.6
File Interpreter ABI Platform
trainmeter-0.0.6-py3-none-any.whl Python 3 none any Details

Total release size: 1.8 MB

Release files / trainmeter-0.0.6.tar.gz

Download URL trainmeter-0.0.6.tar.gz
Size 1.6 MB
Tags Source
SHA-256 checksum
How to use checksums
8ee6a7a7119ac31b9a44eb834ca4fccb90cb20a9d2076201447bca739a2aac89
BLAKE2b-256 checksum
How to use checksums
fd10b6b8c5472c4d852102c336aaa488aad5d48f239db883fff3b5b4ddc8b0b7
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.23 {"installer":{"name":"uv","version":"0.12.23","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release files / trainmeter-0.0.6-py3-none-any.whl

Download URL trainmeter-0.0.6-py3-none-any.whl
Size 198.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
8ead649010651e7c43ba5f6923b128e479b2e17288146a504a0ae9ada7af5b75
BLAKE2b-256 checksum
How to use checksums
549c064da3220a0447b83a8183d3fa7be25713f90004a051eae813eb805da8f8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.23 {"installer":{"name":"uv","version":"0.12.23","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release history Release notifications | RSS feed

This release

0.0.6 This release

2 release files

0.0.5

2 release files

0.0.4

2 release files

0.0.3

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page