gpumanager
gpumanager is a lightweight Python CLI that samples NVIDIA GPU utilization, stores one CSV snapshot per sample, averages utilization over a reporting window, and posts GPU-wise summaries to Slack through an incoming webhook.
One sampler feeds any number of report schedules. A realtime ping every 10 minutes, a daily average, and a quarterly summary can all run side by side from the same collected data.
Features
- Samples NVIDIA GPU utilization with
nvidia-smi - Stores one CSV file per sample, aggregated by GPU UUID
- Any number of report schedules, each with its own cron time and aggregation window
- Rolling windows (
last 7d) or calendar-anchored windows (since the start of this quarter) - Sends reports to Slack via incoming webhook
- Installs and reconciles system-wide
systemdservices and timers - Interactive configuration, minimal dependencies, close to the standard library
Requirements
- Linux with
systemd - Python 3.8 or newer
- NVIDIA GPU with
nvidia-smiinPATH - A Slack incoming webhook URL — see the Slack documentation
Python 3.8 and 3.9 pull in small compatibility dependencies automatically (tomli, backports.zoneinfo).
Installation
pipx is recommended: it keeps gpumanager in its own virtualenv while exposing the command globally.
pipx install gpumanager
pipx ensurepath # run once if the command is not found
source ~/.bashrc
With plain pip:
pip install --user gpumanager
From a local checkout, to test a build before publishing it:
python3 -m build
pipx install --force dist/gpumanager-<version>-py3-none-any.whl
Verify which build is actually on PATH:
gpumanager --help
python3 -c "import gpumanager; print(gpumanager.__version__)"
If a project virtualenv is active, its own copy shadows the pipx one. Run deactivate first, or call ~/.local/bin/gpumanager directly.
Quick Start
gpumanager init # answer the prompts, add one or more reports
gpumanager test-sample # collect one sample now
gpumanager test-report # send it to Slack now
gpumanager install-systemd --enable-now
gpumanager status
init shows the current server time and cron examples while asking for each report's schedule. If timers are already installed, it rewrites and reloads them so changes take effect immediately.
On a multi-user machine, pass the account the services should run as:
gpumanager install-systemd --enable-now --run-user "$USER"
How It Works
gpumanager-sample.timer ──▶ nvidia-smi ──▶ one CSV per sample in csv_dir
│
┌───────────────────────────┴───────────────────────────┐
▼ ▼
gpumanager-report-daily.timer gpumanager-report-quarterly.timer
averages the last 1d ──▶ Slack averages since Jul 1 ──▶ Slack
Sampling and reporting are separate timers. Reports only read the CSV files, so adding a report costs no extra sampling and no extra disk space. Disabling sampling leaves every report with nothing to aggregate.
Configuration
Configuration is searched in this order:
- Path passed with
--config GPUMANAGER_CONFIG~/.config/gpumanager/config.toml/etc/gpumanager/config.toml
[slack]
webhook_url = "https://hooks.slack.com/services/..."
[storage]
csv_dir = "/var/lib/gpumanager"
[sample]
interval = "1m"
[[report]]
name = "daily"
report_time = "0 9 * * *"
interval = "1d"
[[report]]
name = "quarterly"
report_time = "0 9 1 1,4,7,10 *"
interval = "since:quarter"
[general]
timezone = "Asia/Seoul"
server_name = "AICA_H100"
| Key | Meaning |
|---|---|
slack.webhook_url |
Slack incoming webhook, shared by every report |
storage.csv_dir |
Where samples are written; must be writable by the service user |
sample.interval |
How often a sample is taken: 7s, 30s, 2m, 15m, 1h |
report.name |
Unique schedule name, also the systemd unit name |
report.report_time |
When the report is sent, as a 5-field cron string |
report.interval |
How far back the average reaches |
general.timezone |
IANA name; report times and since: boundaries resolve against it |
general.server_name |
Shown in the Slack message header |
Report schedules
Each [[report]] block (note the double brackets) is one schedule, and any number of them can be configured.
name is required and must be unique. It may contain lowercase letters, digits, - and _ only, up to 32 characters, because it becomes part of a systemd unit name installed with sudo (gpumanager-report-<name>.timer).
A single [report] table from an older version is still read correctly, as one report named default.
report_time — when to send
A 5-field cron string: minute, hour, day-of-month, month, day-of-week.
| Schedule | Cron |
|---|---|
| Every day at 09:00 | 0 9 * * * |
| Every hour | 0 * * * * |
| Every 10 minutes | */10 * * * * |
| Every Monday at 09:00 | 0 9 * * 1 |
| First day of each quarter at 09:00 | 0 9 1 1,4,7,10 * |
interval — how far back to average
A rolling duration, counted back from the moment the report is sent:
| Value | Window |
|---|---|
30m, 1h, 12h |
The last 30 minutes / 1 hour / 12 hours |
1d, 7d |
The last 24 hours / 7 days |
Or a calendar anchor with the since: prefix, starting at 00:00 on the boundary day:
| Value | Window |
|---|---|
since:day |
Since midnight today |
since:week |
Since Monday of this week |
since:month |
Since the 1st of this month |
since:quarter |
Since the most recent Jan 1 / Apr 1 / Jul 1 / Oct 1 |
since:year |
Since January 1st |
since:2026-01-15 |
Since a fixed date |
The difference matters at boundaries: on September 21st, 7d reaches back into the previous quarter, while since:quarter stops cleanly at July 1st. The resolved start is printed in the Slack message, so the message itself says what was counted.
Managing Report Schedules
Add a schedule. The new timer is installed and started; existing schedules are untouched.
gpumanager add-report realtime --report-time "*/10 * * * *" --interval 1h
gpumanager add-report # prompts for name, cron and window
Remove a schedule. The block is dropped from the config, the timer is disabled and its unit files are deleted, so the report also disappears from gpumanager status.
gpumanager remove-report realtime
gpumanager remove-report realtime --yes # skip the confirmation
The last remaining report cannot be removed — use uninstall-systemd to remove everything instead.
Stop a report without deleting it. The [[report]] block stays, and reload will not bring the timer back.
gpumanager disable-report --report weekly
sudo systemctl enable --now gpumanager-report-weekly.timer # to resume
Editing the config by hand works too — the config file is the single source of truth:
$EDITOR ~/.config/gpumanager/config.toml
gpumanager reload
reload reconciles everything: new reports get their timers installed and started, deleted reports get their timers disabled and removed.
Send a report immediately:
gpumanager test-report # every report; asks first when several exist
gpumanager test-report --report daily # just one
gpumanager test-report --yes # every report, no question (for scripts)
Commands that act on every report at once ask for confirmation only when more than one report is configured and the terminal is interactive. The installed timers always target a single report by name, so they never wait for an answer.
Automatic Scheduling
gpumanager does not collect anything in the background on its own. Install the system timers:
gpumanager install-systemd --enable-now
This writes to /etc/systemd/system/ (via sudo):
gpumanager-sample.serviceandgpumanager-sample.timergpumanager-report-<name>.serviceandgpumanager-report-<name>.timer, one pair per[[report]]block
add-report, remove-report, install-systemd and reload all reconcile the installed units against the config file:
- timers for newly added reports are enabled and started
- units for reports no longer in the config are disabled and deleted
- the single unnamed
gpumanager-report.{service,timer}pair from versions before 0.3.0 is replaced bygpumanager-report-default.*
Only newly added timers are enabled, so a schedule stopped with disable-report is not silently switched back on.
Stop sampling, which stops the data every report depends on:
gpumanager disable-sample
sudo systemctl enable --now gpumanager-sample.timer # to resume
Remove every unit:
gpumanager uninstall-systemd
Checking Status
gpumanager status
{
"config_path": "/home/master/.config/gpumanager/config.toml",
"csv_dir": "/var/lib/gpumanager",
"csv_dir_exists": true,
"sample.interval": "1m",
"reports": [
{
"name": "daily",
"report_time": "0 9 * * *",
"on_calendar": "*-*-* 09:00:00 Asia/Seoul",
"interval": "1d",
"timer": "gpumanager-report-daily.timer",
"timer_installed": true,
"next_trigger": "Tue 2026-03-24 09:00:00 KST; 22h left"
}
],
"sample_timer_installed": true,
"sample.next_trigger": "Tue 2026-03-24 14:41:35 KST; 9s left"
}
next_trigger values are read by parsing the Trigger: line of systemctl status <timer unit>.
Sampling and Report Format
Each sample writes one CSV named after its timestamp:
2026-03-22T16-21-00.csv
timestamp,gpu_index,gpu_uuid,gpu_name,util_gpu
2026-03-22T16:21:00+09:00,0,GPU-aaa,NVIDIA A100,35
2026-03-22T16:21:00+09:00,1,GPU-bbb,NVIDIA A100,2
Reports average by GPU UUID over the window and round to two decimal places. Missing samples are ignored. The header is [server_name/report_name]; a report named default shows only the server name.
[AICA_H100/quarterly] 2026.10.01 09:00:00 KST
Window: since 2026-07-01 00:00
GPU 0: 31.38%
GPU 1: 29.39%
GPU 2: 31.57%
GPU 3: 56.36%
Old CSV files can be deleted by range:
gpumanager delete-csv --start "2026-01-01 00:00:00" --end "2026-03-31 23:59:59"
Command Reference
| Command | Purpose |
|---|---|
gpumanager init |
Interactive setup; walks through every setting and report |
gpumanager add-report [NAME] [--report-time CRON] [--interval WINDOW] |
Add one report schedule and install its timer |
gpumanager remove-report NAME [--yes] |
Delete one report schedule and its timer |
gpumanager test-sample |
Collect and store one sample now |
gpumanager test-report [--report NAME] [--yes] |
Aggregate and send to Slack now |
gpumanager status |
Configuration, installed timers and next run times, as JSON |
gpumanager install-systemd [--enable-now] [--run-user USER] |
Install and reconcile the system units |
gpumanager reload |
Re-apply the config to the installed units |
gpumanager disable-sample |
Stop the sampling timer |
gpumanager disable-report [--report NAME] [--yes] |
Stop report timers, keeping their config |
gpumanager uninstall-systemd |
Disable and delete every unit |
gpumanager delete-csv [--start S] [--end E] [--yes] |
Delete stored CSV files in a datetime range |
Every command accepts --config PATH to target a specific configuration file.
Recommended Setups
Realtime activity ping:
[[report]]
name = "realtime"
report_time = "*/10 * * * *"
interval = "1h"
Daily average at 09:00:
[[report]]
name = "daily"
report_time = "0 9 * * *"
interval = "1d"
Weekly summary every Monday:
[[report]]
name = "weekly"
report_time = "0 9 * * 1"
interval = "7d"
Quarterly summary that starts exactly at the quarter boundary:
[[report]]
name = "quarterly"
report_time = "0 9 1 1,4,7,10 *"
interval = "since:quarter"
Troubleshooting
The test report arrives, but nothing comes at the scheduled time.
The timers are probably not installed. Run gpumanager status and check sample_timer_installed and each report's timer_installed. Install them with gpumanager install-systemd --enable-now.
Reports say No GPU samples found in the selected window.
Either sampling is not running (gpumanager status → sample.next_trigger), or the service user cannot write to csv_dir, or the window is shorter than the sampling interval.
A schedule was deleted from the config but still fires.
Run gpumanager reload, which removes units for reports that are no longer configured.
Upgrading
pipx upgrade gpumanager # or: pip install --user --upgrade gpumanager
gpumanager reload
gpumanager status
Existing configuration files keep working across upgrades. A pre-0.3.0 single [report] table is read as one report named default, and its unit pair is replaced by gpumanager-report-default.* the next time units are installed. The config file itself is only rewritten when a command that changes it runs (init, add-report, remove-report).
Version History
0.3.4
- Rewritten README: installation, configuration, schedule management, systemd behaviour, command reference and troubleshooting, plus this version history.
0.3.3
remove-report NAMEdeletes one report from the config and removes its timer, so it also disappears fromstatus. The last remaining report is protected.intervalaccepts calendar-anchored windows:since:day,since:week,since:month,since:quarter,since:yearandsince:YYYY-MM-DD, in addition to rolling durations such as7d. Useful for quarterly or month-to-date reports that must start on a boundary rather than N days ago.- The Slack
Window:line shows the resolved start (since 2026-07-01 00:00) instead of only the raw value. - Window validation happens in one place, so an invalid
intervalis rejected when the config is read rather than when the report is sent.
0.3.2
add-report [NAME] [--report-time CRON] [--interval WINDOW]adds a schedule without rerunninginit, and installs and starts its timer.test-reportanddisable-reportask for confirmation before acting on every report at once, with--yesto skip. Non-interactive runs, including the systemd services, are never blocked by the prompt.- Only newly added timers are enabled during a reconcile, so
disable-reportis not undone by a laterreload. - If the systemd step fails after the config was written, the error explains that
gpumanager reloadfinishes the job.
0.3.0
- Multiple report schedules.
[[report]]blocks replace the single[report]table; each has aname,report_timeandinterval. - One systemd unit pair per report (
gpumanager-report-<name>.{service,timer}), reconciled against the config byinstall-systemdandreload: new timers installed, obsolete ones disabled and deleted. test-report --report NAMEanddisable-report --report NAMEact on a single schedule.- Slack messages include the report name in the header:
[AICA_H100/weekly]. statusreports areportsarray with per-reporton_calendar,timer_installedandnext_trigger, replacing the flatreport.*keys.- Configs from earlier versions load unchanged as a single report named
default, and the old unnamed unit pair is replaced automatically.
0.2.x and earlier
- Single report schedule,
nvidia-smisampling into per-sample CSV files, Slack webhook delivery, systemd sample and report timers,--run-userfor the service account.
Notes
sample.intervalcontrols collection;report.intervalcontrols only how far back each report averagessince:boundaries and cron times both resolve againstgeneral.timezone; weeks start on Monday- The README is used as the package long description, so this guide also appears on the PyPI project page
Metadata
Release files for gpumanager 0.3.4
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| gpumanager-0.3.4.tar.gz | 31.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| gpumanager-0.3.4-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 56.7 kB
Release files / gpumanager-0.3.4.tar.gz
| Download URL | gpumanager-0.3.4.tar.gz |
|---|---|
| Size | 31.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
52df577e6a824853a2ee4c11c7f260f548791d57efb95f87746f369ec43d324f
|
|
BLAKE2b-256 checksum How to use checksums |
61d3cf4772537a9c4f3214be21469d57d3dd357787052b6cccaf49a3e5c02418
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.1.0 CPython/3.8.10
|
Release files / gpumanager-0.3.4-py3-none-any.whl
| Download URL | gpumanager-0.3.4-py3-none-any.whl |
|---|---|
| Size | 25.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
162a63c654204b366e4b75681e1f79153c63429c12834d8e14c0482b4637ec8e
|
|
BLAKE2b-256 checksum How to use checksums |
dc4fffae49c680b8aa60cd7c26fe00354b2a9783c8a49a639cba5ddbb6aa5b5e
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.1.0 CPython/3.8.10
|