Skip to main content

gpumanager

gpumanager is a lightweight Python CLI that samples NVIDIA GPU utilization, stores one CSV snapshot per sample, averages utilization over a reporting window, and posts GPU-wise summaries to Slack through an incoming webhook.

One sampler feeds any number of report schedules. A realtime ping every 10 minutes, a daily average, and a quarterly summary can all run side by side from the same collected data.

Features

  • Samples NVIDIA GPU utilization with nvidia-smi
  • Stores one CSV file per sample, aggregated by GPU UUID
  • Any number of report schedules, each with its own cron time and aggregation window
  • Rolling windows (last 7d) or calendar-anchored windows (since the start of this quarter)
  • Sends reports to Slack via incoming webhook
  • Installs and reconciles system-wide systemd services and timers
  • Interactive configuration, minimal dependencies, close to the standard library

Requirements

  • Linux with systemd
  • Python 3.8 or newer
  • NVIDIA GPU with nvidia-smi in PATH
  • A Slack incoming webhook URL — see the Slack documentation

Python 3.8 and 3.9 pull in small compatibility dependencies automatically (tomli, backports.zoneinfo).

Installation

pipx is recommended: it keeps gpumanager in its own virtualenv while exposing the command globally.

pipx install gpumanager
pipx ensurepath        # run once if the command is not found
source ~/.bashrc

With plain pip:

pip install --user gpumanager

From a local checkout, to test a build before publishing it:

python3 -m build
pipx install --force dist/gpumanager-<version>-py3-none-any.whl

Verify which build is actually on PATH:

gpumanager --help
python3 -c "import gpumanager; print(gpumanager.__version__)"

If a project virtualenv is active, its own copy shadows the pipx one. Run deactivate first, or call ~/.local/bin/gpumanager directly.

Quick Start

gpumanager init                      # answer the prompts, add one or more reports
gpumanager test-sample               # collect one sample now
gpumanager test-report               # send it to Slack now
gpumanager install-systemd --enable-now
gpumanager status

init shows the current server time and cron examples while asking for each report's schedule. If timers are already installed, it rewrites and reloads them so changes take effect immediately.

On a multi-user machine, pass the account the services should run as:

gpumanager install-systemd --enable-now --run-user "$USER"

How It Works

gpumanager-sample.timer  ──▶  nvidia-smi  ──▶  one CSV per sample in csv_dir
                                                      │
                          ┌───────────────────────────┴───────────────────────────┐
                          ▼                                                       ▼
        gpumanager-report-daily.timer                        gpumanager-report-quarterly.timer
        averages the last 1d  ──▶ Slack                      averages since Jul 1 ──▶ Slack

Sampling and reporting are separate timers. Reports only read the CSV files, so adding a report costs no extra sampling and no extra disk space. Disabling sampling leaves every report with nothing to aggregate.

Configuration

Configuration is searched in this order:

  1. Path passed with --config
  2. GPUMANAGER_CONFIG
  3. ~/.config/gpumanager/config.toml
  4. /etc/gpumanager/config.toml
[slack]
webhook_url = "https://hooks.slack.com/services/..."

[storage]
csv_dir = "/var/lib/gpumanager"

[sample]
interval = "1m"

[[report]]
name = "daily"
report_time = "0 9 * * *"
interval = "1d"

[[report]]
name = "quarterly"
report_time = "0 9 1 1,4,7,10 *"
interval = "since:quarter"

[general]
timezone = "Asia/Seoul"
server_name = "AICA_H100"
Key Meaning
slack.webhook_url Slack incoming webhook, shared by every report
storage.csv_dir Where samples are written; must be writable by the service user
sample.interval How often a sample is taken: 7s, 30s, 2m, 15m, 1h
report.name Unique schedule name, also the systemd unit name
report.report_time When the report is sent, as a 5-field cron string
report.interval How far back the average reaches
general.timezone IANA name; report times and since: boundaries resolve against it
general.server_name Shown in the Slack message header

Report schedules

Each [[report]] block (note the double brackets) is one schedule, and any number of them can be configured.

name is required and must be unique. It may contain lowercase letters, digits, - and _ only, up to 32 characters, because it becomes part of a systemd unit name installed with sudo (gpumanager-report-<name>.timer).

A single [report] table from an older version is still read correctly, as one report named default.

report_time — when to send

A 5-field cron string: minute, hour, day-of-month, month, day-of-week.

Schedule Cron
Every day at 09:00 0 9 * * *
Every hour 0 * * * *
Every 10 minutes */10 * * * *
Every Monday at 09:00 0 9 * * 1
First day of each quarter at 09:00 0 9 1 1,4,7,10 *

interval — how far back to average

A rolling duration, counted back from the moment the report is sent:

Value Window
30m, 1h, 12h The last 30 minutes / 1 hour / 12 hours
1d, 7d The last 24 hours / 7 days

Or a calendar anchor with the since: prefix, starting at 00:00 on the boundary day:

Value Window
since:day Since midnight today
since:week Since Monday of this week
since:month Since the 1st of this month
since:quarter Since the most recent Jan 1 / Apr 1 / Jul 1 / Oct 1
since:year Since January 1st
since:2026-01-15 Since a fixed date

The difference matters at boundaries: on September 21st, 7d reaches back into the previous quarter, while since:quarter stops cleanly at July 1st. The resolved start is printed in the Slack message, so the message itself says what was counted.

Managing Report Schedules

Add a schedule. The new timer is installed and started; existing schedules are untouched.

gpumanager add-report realtime --report-time "*/10 * * * *" --interval 1h
gpumanager add-report                   # prompts for name, cron and window

Remove a schedule. The block is dropped from the config, the timer is disabled and its unit files are deleted, so the report also disappears from gpumanager status.

gpumanager remove-report realtime
gpumanager remove-report realtime --yes # skip the confirmation

The last remaining report cannot be removed — use uninstall-systemd to remove everything instead.

Stop a report without deleting it. The [[report]] block stays, and reload will not bring the timer back.

gpumanager disable-report --report weekly
sudo systemctl enable --now gpumanager-report-weekly.timer   # to resume

Editing the config by hand works too — the config file is the single source of truth:

$EDITOR ~/.config/gpumanager/config.toml
gpumanager reload

reload reconciles everything: new reports get their timers installed and started, deleted reports get their timers disabled and removed.

Send a report immediately:

gpumanager test-report                  # every report; asks first when several exist
gpumanager test-report --report daily   # just one
gpumanager test-report --yes            # every report, no question (for scripts)

Commands that act on every report at once ask for confirmation only when more than one report is configured and the terminal is interactive. The installed timers always target a single report by name, so they never wait for an answer.

Automatic Scheduling

gpumanager does not collect anything in the background on its own. Install the system timers:

gpumanager install-systemd --enable-now

This writes to /etc/systemd/system/ (via sudo):

  • gpumanager-sample.service and gpumanager-sample.timer
  • gpumanager-report-<name>.service and gpumanager-report-<name>.timer, one pair per [[report]] block

add-report, remove-report, install-systemd and reload all reconcile the installed units against the config file:

  • timers for newly added reports are enabled and started
  • units for reports no longer in the config are disabled and deleted
  • the single unnamed gpumanager-report.{service,timer} pair from versions before 0.3.0 is replaced by gpumanager-report-default.*

Only newly added timers are enabled, so a schedule stopped with disable-report is not silently switched back on.

Stop sampling, which stops the data every report depends on:

gpumanager disable-sample
sudo systemctl enable --now gpumanager-sample.timer   # to resume

Remove every unit:

gpumanager uninstall-systemd

Checking Status

gpumanager status
{
  "config_path": "/home/master/.config/gpumanager/config.toml",
  "csv_dir": "/var/lib/gpumanager",
  "csv_dir_exists": true,
  "sample.interval": "1m",
  "reports": [
    {
      "name": "daily",
      "report_time": "0 9 * * *",
      "on_calendar": "*-*-* 09:00:00 Asia/Seoul",
      "interval": "1d",
      "timer": "gpumanager-report-daily.timer",
      "timer_installed": true,
      "next_trigger": "Tue 2026-03-24 09:00:00 KST; 22h left"
    }
  ],
  "sample_timer_installed": true,
  "sample.next_trigger": "Tue 2026-03-24 14:41:35 KST; 9s left"
}

next_trigger values are read by parsing the Trigger: line of systemctl status <timer unit>.

Sampling and Report Format

Each sample writes one CSV named after its timestamp:

2026-03-22T16-21-00.csv
timestamp,gpu_index,gpu_uuid,gpu_name,util_gpu
2026-03-22T16:21:00+09:00,0,GPU-aaa,NVIDIA A100,35
2026-03-22T16:21:00+09:00,1,GPU-bbb,NVIDIA A100,2

Reports average by GPU UUID over the window and round to two decimal places. Missing samples are ignored. The header is [server_name/report_name]; a report named default shows only the server name.

[AICA_H100/quarterly] 2026.10.01 09:00:00 KST
Window: since 2026-07-01 00:00
GPU 0: 31.38%
GPU 1: 29.39%
GPU 2: 31.57%
GPU 3: 56.36%

Old CSV files can be deleted by range:

gpumanager delete-csv --start "2026-01-01 00:00:00" --end "2026-03-31 23:59:59"

Command Reference

Command Purpose
gpumanager init Interactive setup; walks through every setting and report
gpumanager add-report [NAME] [--report-time CRON] [--interval WINDOW] Add one report schedule and install its timer
gpumanager remove-report NAME [--yes] Delete one report schedule and its timer
gpumanager test-sample Collect and store one sample now
gpumanager test-report [--report NAME] [--yes] Aggregate and send to Slack now
gpumanager status Configuration, installed timers and next run times, as JSON
gpumanager install-systemd [--enable-now] [--run-user USER] Install and reconcile the system units
gpumanager reload Re-apply the config to the installed units
gpumanager disable-sample Stop the sampling timer
gpumanager disable-report [--report NAME] [--yes] Stop report timers, keeping their config
gpumanager uninstall-systemd Disable and delete every unit
gpumanager delete-csv [--start S] [--end E] [--yes] Delete stored CSV files in a datetime range

Every command accepts --config PATH to target a specific configuration file.

Realtime activity ping:

[[report]]
name = "realtime"
report_time = "*/10 * * * *"
interval = "1h"

Daily average at 09:00:

[[report]]
name = "daily"
report_time = "0 9 * * *"
interval = "1d"

Weekly summary every Monday:

[[report]]
name = "weekly"
report_time = "0 9 * * 1"
interval = "7d"

Quarterly summary that starts exactly at the quarter boundary:

[[report]]
name = "quarterly"
report_time = "0 9 1 1,4,7,10 *"
interval = "since:quarter"

Troubleshooting

The test report arrives, but nothing comes at the scheduled time. The timers are probably not installed. Run gpumanager status and check sample_timer_installed and each report's timer_installed. Install them with gpumanager install-systemd --enable-now.

Reports say No GPU samples found in the selected window. Either sampling is not running (gpumanager status → sample.next_trigger), or the service user cannot write to csv_dir, or the window is shorter than the sampling interval.

A schedule was deleted from the config but still fires. Run gpumanager reload, which removes units for reports that are no longer configured.

Upgrading

pipx upgrade gpumanager          # or: pip install --user --upgrade gpumanager
gpumanager reload
gpumanager status

Existing configuration files keep working across upgrades. A pre-0.3.0 single [report] table is read as one report named default, and its unit pair is replaced by gpumanager-report-default.* the next time units are installed. The config file itself is only rewritten when a command that changes it runs (init, add-report, remove-report).

Version History

0.3.4

  • Rewritten README: installation, configuration, schedule management, systemd behaviour, command reference and troubleshooting, plus this version history.

0.3.3

  • remove-report NAME deletes one report from the config and removes its timer, so it also disappears from status. The last remaining report is protected.
  • interval accepts calendar-anchored windows: since:day, since:week, since:month, since:quarter, since:year and since:YYYY-MM-DD, in addition to rolling durations such as 7d. Useful for quarterly or month-to-date reports that must start on a boundary rather than N days ago.
  • The Slack Window: line shows the resolved start (since 2026-07-01 00:00) instead of only the raw value.
  • Window validation happens in one place, so an invalid interval is rejected when the config is read rather than when the report is sent.

0.3.2

  • add-report [NAME] [--report-time CRON] [--interval WINDOW] adds a schedule without rerunning init, and installs and starts its timer.
  • test-report and disable-report ask for confirmation before acting on every report at once, with --yes to skip. Non-interactive runs, including the systemd services, are never blocked by the prompt.
  • Only newly added timers are enabled during a reconcile, so disable-report is not undone by a later reload.
  • If the systemd step fails after the config was written, the error explains that gpumanager reload finishes the job.

0.3.0

  • Multiple report schedules. [[report]] blocks replace the single [report] table; each has a name, report_time and interval.
  • One systemd unit pair per report (gpumanager-report-<name>.{service,timer}), reconciled against the config by install-systemd and reload: new timers installed, obsolete ones disabled and deleted.
  • test-report --report NAME and disable-report --report NAME act on a single schedule.
  • Slack messages include the report name in the header: [AICA_H100/weekly].
  • status reports a reports array with per-report on_calendar, timer_installed and next_trigger, replacing the flat report.* keys.
  • Configs from earlier versions load unchanged as a single report named default, and the old unnamed unit pair is replaced automatically.

0.2.x and earlier

  • Single report schedule, nvidia-smi sampling into per-sample CSV files, Slack webhook delivery, systemd sample and report timers, --run-user for the service account.

Notes

  • sample.interval controls collection; report.interval controls only how far back each report averages
  • since: boundaries and cron times both resolve against general.timezone; weeks start on Monday
  • The README is used as the package long description, so this guide also appears on the PyPI project page

Metadata

Release files for gpumanager 0.3.4

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for gpumanager 0.3.4
File Size Uploaded
gpumanager-0.3.4.tar.gz 31.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for gpumanager 0.3.4
File Interpreter ABI Platform
gpumanager-0.3.4-py3-none-any.whl Python 3 none any Details

Total release size: 56.7 kB

Release files / gpumanager-0.3.4.tar.gz

Download URL gpumanager-0.3.4.tar.gz
Size 31.1 kB
Tags Source
SHA-256 checksum
How to use checksums
52df577e6a824853a2ee4c11c7f260f548791d57efb95f87746f369ec43d324f
BLAKE2b-256 checksum
How to use checksums
61d3cf4772537a9c4f3214be21469d57d3dd357787052b6cccaf49a3e5c02418
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.8.10

Release files / gpumanager-0.3.4-py3-none-any.whl

Download URL gpumanager-0.3.4-py3-none-any.whl
Size 25.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
162a63c654204b366e4b75681e1f79153c63429c12834d8e14c0482b4637ec8e
BLAKE2b-256 checksum
How to use checksums
dc4fffae49c680b8aa60cd7c26fe00354b2a9783c8a49a639cba5ddbb6aa5b5e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.8.10

Release history Release notifications | RSS feed

This release

0.3.4 This release

2 release files

0.3.3

2 release files

0.3.2

2 release files

0.3.0

2 release files

0.2.4

2 release files

0.2.3

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page