Skip to main content

trainmeter

Where your training FLOPs go, and how well your GPUs are used.

Run tm train.py instead of python train.py. trainmeter starts your command, watches the node from outside, and serves a live dashboard on localhost. The dashboard shows what a training run needs to be judged: training and validation loss, FLOPs spent and required, MFU with its FLOPs convention spelled out, and how busy the GPUs really are, from GPU-Util down to MFU, next to arithmetic intensity. Your training code stays as it is.

The trainmeter dashboard: progress, throughput, MFU and OFU, goodput and the loss curve of a run on four H100s

The dashboard on a synthetic run.

Status: early scaffold. tm runs any command untouched, records the run, and turns what a PyTorch loop does into tokens, FLOPs spent and required, ETA, MFU (given a device) and the losses you emit(). GPU counters (NVML and GPM: the utilization ladder, OFU, observed FLOPs, arithmetic intensity) and tm doctor are built and tested against fakes, not yet on a real GPU. Series your loop already logs to wandb, TensorBoard or MLflow are picked up without a code change (--no-tap turns it off), and DataLoader waits, checkpoint saves and eval spans feed a "where the time goes" view. One real step is also counted with FlopCounterMode (shown beside MFU as the counted convention), and layers and width are read from a model config when there is one. The node around the run is shown the way beszel shows a system: CPU by kind and per core, load, memory with its cache and swap, filesystems, disk and network throughput, temperatures and fans, read from /proc and /sys and checked on a cloud VM. Each run ends with a run passport (passport.md and passport.json) for a model card. The dashboard, one page of progress, tiles and charts with the run details and model card folded away, is served on localhost while the run goes; it has been checked in a browser on synthetic data, not on a real run. Not yet used on a real run. To try it, examples/ has a tiny GPT to run under tm on a laptop and a run planner that needs no GPU.

Install

trainmeter is on PyPI:

# into the training environment, next to torch: this is what lets tm see inside the program
uv add "trainmeter[gpu]"
# or as a standalone tool, enough for tm view, tm doctor and tm passport
uv tool install trainmeter

The gpu extra pulls in nvidia-ml-py for the GPU counters and can be left out on a machine with no NVIDIA GPU. Without uv, pip install "trainmeter[gpu]" does the same.

How it will be used

tm train.py                                 # instead of: python train.py
tm --budget-tokens 20B train.py             # declare the plan: FLOPs required, progress, ETA
tm -- torchrun --nproc_per_node=8 train.py  # any command, one grammar
tm doctor                                   # what tm can see on this node
tm doctor --dump node.tar.gz                # and pack what it reads, for a bug report
tm passport                                 # model-card passport of the latest run
tm export wandb --project p                 # upload a finished run to W&B
tm view                                     # replay the latest finished run in the dashboard
tm attach 12345                             # watch a program that is already running, from outside

tm prints one line with the dashboard URL and, at the end, one summary line. The dashboard binds to loopback on the node, so a remote node needs an SSH tunnel (ssh -L 8765:localhost:8765 node). --port N picks the first port to try and --no-web turns the dashboard off.

How it works

  • From outside. tm starts your command as a child process, reads the GPU counters (tensor, DRAM and SM activity via NVML GPM on Hopper and newer), and keeps the run log.
  • From inside, passively. A small agent that tm injects into the training interpreter sees steps, batch shapes, parameter counts and data waits. It never synchronizes a device and never raises into your loop.
  • From your loggers. The loss and any other series your loop already sends to wandb, TensorBoard or MLflow are picked up as they are. A loop with no tracker adds one emit(train_loss=...) line.
  • Physics, not prices. The run log holds time, tokens, FLOPs, counters and the series your loop reports. Money is an optional overlay computed at view time, and a price is never an input.

docs/architecture.md has the design, and docs/spec.md defines every metric.

Why the convention matters

On a 1.04B model with d 2048, 18 layers and context 1024, FLOPs per token are 5.83 G by the Hugging Face count, 6.06 G by Megatron, 6.28 G by TorchTitan and nanochat, and 6.69 G when embeddings are counted as matmuls. Same run, MFU differing by 15%. trainmeter always says which one it used.

Develop

uv sync
uv run pytest
uv run ruff check .
uv run ruff format --check .

See docs/spec.md for what trainmeter shows and records, docs/architecture.md for how it is built, and CLAUDE.md for the invariants.

License

Apache-2.0.

Metadata

Release files for trainmeter 0.0.5

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for trainmeter 0.0.5
File Size Uploaded
trainmeter-0.0.5.tar.gz 671.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for trainmeter 0.0.5
File Interpreter ABI Platform
trainmeter-0.0.5-py3-none-any.whl Python 3 none any Details

Total release size: 804.3 kB

Release files / trainmeter-0.0.5.tar.gz

Download URL trainmeter-0.0.5.tar.gz
Size 671.8 kB
Tags Source
SHA-256 checksum
How to use checksums
f38f6e78e718a9866beabba4c57ac9c18453bcfd5e6ae4cefe9919cd05e7f811
BLAKE2b-256 checksum
How to use checksums
64dca3bf5cd6d21828fe8d258b01ea5d084ac397e176d742c95791f35a435cfd
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.22 {"installer":{"name":"uv","version":"0.12.22","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release files / trainmeter-0.0.5-py3-none-any.whl

Download URL trainmeter-0.0.5-py3-none-any.whl
Size 132.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
c079395b0831483d9a340a8af563f9bfc213b31c8942bb42182e0d3b429dae7d
BLAKE2b-256 checksum
How to use checksums
59b091363780b0c4da72f2e2269c40a5b9fdf62823b0a64f2a72c4f42e505e77
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.22 {"installer":{"name":"uv","version":"0.12.22","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release history Release notifications | RSS feed

0.0.6

2 release files

This release

0.0.5 This release

2 release files

0.0.4

2 release files

0.0.3

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page