Skip to main content

gpumanager

gpumanager is a lightweight Python CLI tool for sampling NVIDIA GPU utilization, storing minute-by-minute CSV snapshots, aggregating utilization over a reporting window, and sending GPU-wise summaries to Slack. It is designed to work together with a Slack incoming webhook for notifications.

The installable Python distribution is named gpumanager. The CLI entrypoint is gpumanager.

Website: https://happilee12.github.io/gpu-util-webhook/ Pip Page: https://pypi.org/project/gpumanager/

Features

  • Samples NVIDIA GPU utilization with nvidia-smi
  • Stores one CSV file per sample
  • Aggregates average utilization by GPU UUID
  • Sends reports to Slack via webhook
  • Supports interactive configuration
  • Installs system-wide systemd services and timers
  • Uses minimal dependencies and stays close to the standard library

Requirements

  • Linux
  • Python 3.8+
  • NVIDIA GPU
  • nvidia-smi in PATH
  • systemd recommended

Installation

Python 3.8 support uses small compatibility dependencies installed automatically by pip:

  • tomli on Python < 3.11
  • backports.zoneinfo on Python < 3.9
pip install .
# or
pipx install .

If you install with pipx, make sure the pipx binary path is added to your shell:

pipx ensurepath
source ~/.bashrc

After publishing:

pip install gpumanager
# or
pipx install gpumanager

After a published pipx install, run this once if needed:

pipx ensurepath
source ~/.bashrc

Quick Start

gpumanager init
gpumanager install-systemd --enable-now

During init, the CLI shows the current server time and a few common cron examples so it is easier to enter report.report_time.

If you edit the config file manually after timers are installed, run gpumanager reload to apply the updated systemd timer settings.

Troubleshooting

4. Test

After finishing the configuration, send a test report.

gpumanager test-sample
gpumanager test-report

gpumanager test-report sends every configured report. Use gpumanager test-report --report weekly to send just one.

If the Slack message arrives normally, the setup is working.

If the message is delivered here but does not arrive at the scheduled time, gpumanager install-systemd may not have been run yet. In that case, run gpumanager status and check sample_timer_installed, sample.next_trigger, and the timer_installed / next_trigger fields of each entry in reports. The next scheduled runs are visible directly in status output:

"sample.next_trigger": "Tue 2026-03-24 14:41:35 KST; 9s left",
"reports": [
  {
    "name": "daily",
    "timer": "gpumanager-report-daily.timer",
    "timer_installed": true,
    "next_trigger": "Tue 2026-03-24 14:42:00 KST; 33s left"
  }
]

These values are read by parsing the Trigger: line from systemctl status <timer unit>.

If systemd timers are already installed, gpumanager init automatically rewrites and reloads the installed timer files so schedule changes take effect immediately. If you edit the config file manually later, run gpumanager reload.

Commands

  • gpumanager init
  • gpumanager test-sample
  • gpumanager test-report [--report NAME]
  • gpumanager delete-csv
  • gpumanager status
  • gpumanager install-systemd
  • gpumanager uninstall-systemd
  • gpumanager disable-sample
  • gpumanager disable-report [--report NAME]
  • gpumanager reload

Configuration

The tool searches for configuration in this order:

  1. Path passed with --config
  2. GPUMANAGER_CONFIG
  3. ~/.config/gpumanager/config.toml
  4. /etc/gpumanager/config.toml

Example:

[slack]
webhook_url = "https://hooks.slack.com/services/..."

[storage]
csv_dir = "/var/lib/gpumanager"

[sample]
interval = "1m"

[[report]]
name = "daily"
report_time = "0 9 * * *"
interval = "1d"

[[report]]
name = "weekly"
report_time = "0 9 * * 1"
interval = "7d"

[general]
timezone = "Asia/Seoul"
server_name = "AICA_H100"

Each [[report]] block is one schedule, and any number of them can be configured. name is required, must be unique, and may contain lowercase letters, digits, - and _ only, because it becomes part of the installed systemd unit name (gpumanager-report-<name>.timer).

A single [report] table from older versions is still accepted and is read as one report named default.

To delete a schedule, remove its [[report]] block and run gpumanager reload; the matching timer is disabled and its unit files are removed. To keep the block but stop the notification, use gpumanager disable-report --report <name>.

Common report_time examples:

  • Every day at 09:00: 0 9 * * *
  • Every hour: 0 * * * *
  • Every 10 minutes: */10 * * * *

Sampling examples:

  • Every 7 seconds: 7s
  • Every 30 seconds: 30s
  • Every 2 minutes: 2m
  • Every 15 minutes: 15m
  • Every hour: 1h

1. Realtime report

Check near-realtime GPU activity every 10 minutes.

[sample]
interval = "1m"

[[report]]
name = "realtime"
report_time = "*/10 * * * *"
interval = "1m"

2. Daily Average report

This matches the current default-style daily setup.

[sample]
interval = "1m"

[[report]]
name = "daily"
report_time = "0 9 * * *"
interval = "1d"

3. Weekly report

Send one summary per week and aggregate the last 7 days.

[sample]
interval = "1m"

[[report]]
name = "weekly"
report_time = "0 9 * * 1"
interval = "7d"

4. Several reports at once

Reports are independent, so a realtime ping and a daily summary can run side by side. Sampling is shared: one sampler feeds every report.

[sample]
interval = "1m"

[[report]]
name = "realtime"
report_time = "*/10 * * * *"
interval = "1m"

[[report]]
name = "daily"
report_time = "0 9 * * *"
interval = "1d"

Before Running Reports

A few things must be prepared by the user before gpumanager can collect data and send Slack notifications through a Slack incoming webhook:

  • nvidia-smi must work on the server
  • A valid Slack incoming webhook URL must be configured
  • The CSV storage directory must be writable
  • If you want automatic collection and reporting, the system-wide systemd timers must be enabled

Slack incoming webhook setup reference:

Quick manual verification:

nvidia-smi
gpumanager status
gpumanager test-sample
gpumanager test-report

Automatic Scheduling

gpumanager does not start background collection on its own. To run sampling every minute and reporting on the configured cron-style schedule, install and enable the system timers.

Install the unit files:

gpumanager install-systemd

Then enable the timers:

sudo systemctl enable --now gpumanager-sample.timer gpumanager-report-daily.timer

gpumanager install-systemd --enable-now enables the sample timer and every configured report timer without typing the unit names.

Check timer status or reload installed timers:

gpumanager status
gpumanager reload

gpumanager status shows the next scheduled sample time in sample.next_trigger and the next run of every report in reports[].next_trigger. These values are read by parsing the Trigger: line from systemctl status <timer unit>.

Disable only sampling:

gpumanager disable-sample

Disable reporting (all reports, or one by name):

gpumanager disable-report
gpumanager disable-report --report weekly

Sampling

Each sample creates a CSV file named like:

2026-03-22T16-21-00.csv

Each CSV contains one row per GPU:

timestamp,gpu_index,gpu_uuid,gpu_name,util_gpu
2026-03-22T16:21:00+09:00,0,GPU-aaa,NVIDIA A100,35
2026-03-22T16:21:00+09:00,1,GPU-bbb,NVIDIA A100,2

Report Format

Reports use the configured general.server_name as the bracketed name prefix, followed by the report name as [server/report]. A report named default (what an older single-[report] config becomes) shows only the server name. Average GPU utilization is rounded to two decimal places.

Example:

[AICA_H100/weekly] 2025.09.06 16:49:32 KST
Window: last 1h
GPU 0: 31.38%
GPU 1: 29.39%
GPU 2: 31.57%
GPU 3: 56.36%
GPU 4: 61.25%
GPU 5: 61.52%
GPU 6: 59.88%
GPU 7: 63.93%

systemd

gpumanager install-systemd installs system services into /etc/systemd/system/:

  • gpumanager-sample.service
  • gpumanager-sample.timer
  • gpumanager-report-<name>.service and gpumanager-report-<name>.timer, one pair per [[report]] block

install-systemd and reload reconcile the installed units with the config file: units for reports that are no longer configured are disabled and deleted, and the single unnamed gpumanager-report.{service,timer} pair from older versions is replaced by gpumanager-report-default.*.

Notes

  • Each [[report]] entry needs a unique name; it is used as the systemd unit name
  • report_time uses a 5-field cron string such as 0 9 * * *
  • sample.interval controls how often GPU utilization is sampled and saved
  • interval controls the aggregation window shown as Window: last ... and supports minute-based values such as 1m
  • Missing samples are ignored during aggregation
  • The README content is used as the package long description, so this setup guide will also be visible on package index web pages after publishing

Metadata

Release files for gpumanager 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for gpumanager 0.3.0
File Size Uploaded
gpumanager-0.3.0.tar.gz 23.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for gpumanager 0.3.0
File Interpreter ABI Platform
gpumanager-0.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 44.3 kB

Release files / gpumanager-0.3.0.tar.gz

Download URL gpumanager-0.3.0.tar.gz
Size 23.1 kB
Tags Source
SHA-256 checksum
How to use checksums
5e27c90907a19a5b656c3eede6c166624af8987eadb5883ab439a0df4c8e141a
BLAKE2b-256 checksum
How to use checksums
ea50183f5ddac25e87ded5be21f416dcf4be9408ac8250081f7a0968a6aba34a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.8.10

Release files / gpumanager-0.3.0-py3-none-any.whl

Download URL gpumanager-0.3.0-py3-none-any.whl
Size 21.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
f4de0aa801c240f5e0846d6860c422a3041fb43756d821e4f1a584306e477bbc
BLAKE2b-256 checksum
How to use checksums
393950c7f8ab1064aa4cdac0cf74a674c7e090b958a57fa9e0328a71047ff56e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.8.10

Release history Release notifications | RSS feed

0.3.4

2 release files

0.3.3

2 release files

0.3.2

2 release files

This release

0.3.0 This release

2 release files

0.2.4

2 release files

0.2.3

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page