clustop
One terminal window that shows what every GPU server in your cluster is doing right now.
Think htop + nvitop, but for a whole cluster at once (any set of SSH-reachable GPU machines, saved as profiles): CPU, RAM, every
GPU, who is using it, and whether the job came from Slurm or someone running it by hand.
Most importantly, it tells you which GPUs are free.
clustop 3/4 hosts · 12 GPUs · avg util 48% · VRAM 212G/566G · free: node1:0, node2:3 q quit · space pause · +/- 2s
╭──────────────────────── node1 10.0.0.1 1/4 GPU free ────────────────────────╮
│ CPU ██░░░░░░░░░░░░░░░░░░░░░░ 6% ld 9.1/128c │
│ RAM ███░░░░░░░░░░░░░░░░░░░░░ 9% 23.0G/250.9G │
│ │
│ GPU Name Util Memory T Who (S=slurm M=manual) │
│ 0 A30 ░░░░░░░░░░ 0% ░░░░░░░░░░ 1.0G/24.0G 28° FREE │
│ 1 A30 ███░░░░░░░ 33% ░░░░░░░░░░ 508M/24.0G 25° 1p alice:M │
│ 3 A30 ██████████ 100% ██░░░░░░░░ 3.5G/24.0G 49° 3p hmslati:S │
╰─────────────────────────────────────────────────────────────────────────────────────╯
GPU processes (by memory)
Host GPU User PID GPU mem Src Time Command
node4 1 awhite 23879 83.3G slurm 57381 3-04:30:58 python .../train.py
Quick start
You need Python 3.11+ and pipx (brew install pipx on a Mac).
pipx install clustop # or: pip install clustop
clustop profile add mycluster 10.0.0.1 10.0.0.2 10.0.0.3 --user YOUR_USER --default
clustop
(prefer the latest development version? pipx install git+https://github.com/aipyth/clustop)
(or skip the profile for a quick look: clustop 10.0.0.1 10.0.0.2)
On first start it asks for your SSH password once per server. After that it keeps the connections open for 8 hours, so the next launches start without asking again.
Want to skip passwords entirely? Copy your SSH key to each server once:
ssh-copy-id YOUR_USER@10.0.0.1 # repeat for the other servers
Make sure the right username is used: put this in ~/.ssh/config:
Host 10.0.0.1 10.0.0.2 10.0.0.3 10.0.0.4
User YOUR_USER
What you're looking at
Header: how many servers answered, total GPUs, average GPU utilization, total
VRAM in use, and a list of every free GPU as host:index.
One panel per server
| Part | Meaning |
|---|---|
CPU |
Overall CPU use and the 1-minute load average next to the number of cores (ld 9.1/128c). |
RAM |
Memory in use (not counting reclaimable cache) out of total. |
Util |
How busy the GPU's compute units are. |
Memory |
GPU memory used / total. |
T |
GPU temperature. |
Who |
Number of processes on this GPU and their users. :S = started by Slurm, :M = started manually (ssh, Jupyter, a screen session...). |
Bars turn green → yellow → red as things fill up.
GPU status in the Who column
FREE(green badge): no process on it, under 5% utilization and under 1 GB of memory in use. Safe to grab.3p alice:S bob:M: three processes, from alice (Slurm) and bob (manual). The GPU is shared.busy, proc not visible: the GPU is in use but no process shows up, which can happen with some containers. Treat it as taken.
Process table (bottom): every process using a GPU, across all servers, biggest
memory user first. Src says slurm <jobid> or manual.
Keys and options
| Key | Action |
|---|---|
q |
quit |
space |
pause / resume updating |
+ / - |
slower / faster refresh (0.5 s to 60 s) |
clustop # your default profile, refresh every 2 s
clustop -i 5 # refresh every 5 s
clustop -p lab # another profile (see Profiles)
clustop 10.0.0.1 10.0.0.4 # only these servers
Profiles (multiple clusters)
A profile is a named list of servers. Create one per cluster, make one the default,
and plain clustop opens it:
clustop profile add lab 10.0.0.1 10.0.0.2 --user bob # create a profile
clustop profile add lab 10.0.0.1 10.0.0.2 -u bob --default # ...and make it the default
clustop profile list # * marks the default
clustop profile default lab # change the default
clustop profile show lab
clustop profile remove lab
clustop profile path # where the config file lives
clustop # runs the default profile (set one up first, see below)
clustop -p lab # runs a specific profile
clustop -p lab -u alice # same, with another ssh user
clustop 10.1.1.5 10.1.1.6 # ad-hoc hosts, no profile needed
Options per profile: hosts (required), user (ssh user; otherwise your
~/.ssh/config decides) and interval (refresh seconds; -i on the command
line overrides it). The active profile name is shown in the header.
Profiles are stored in ~/.config/clustop/config.toml (created on first run), so you
can also edit it by hand:
default = "gpu-lab"
[profiles.gpu-lab]
hosts = ["10.0.0.1", "10.0.0.2", "10.0.0.3", "10.0.0.4"]
[profiles.lab]
hosts = ["10.0.0.1", "10.0.0.2"]
user = "bob"
interval = 5.0
Set CLUSTOP_CONFIG=/path/to/file.toml to use a different config file, for example
one shared by your team in a git repo. To switch profiles, quit and start it again with -p.
How it works (and what it touches)
- It runs plain
sshfrom your laptop. Nothing is installed on the servers and nothing is written there. - Every refresh it runs one small read-only script on each server: it reads
/proc(CPU, memory), callsnvidia-smi(GPUs and their processes) andps(who owns them). - Slurm jobs are recognized from each process's cgroup path
(
.../slurmstepd.scope/job_12345/...). Thesqueuecommand isn't needed. - If
nvidia-smihangs on a node (it happens when a driver is unhappy), it is cut off after 8 seconds and you still see that server's CPU and RAM, with a red note instead of GPU rows. - The server list comes from your profile (see Profiles above).
Troubleshooting
| Problem | Fix |
|---|---|
could not connect (wrong password / unreachable) |
Check your VPN/network and try ssh YOUR_USER@10.0.0.1 by hand. Then restart clustop. |
A server shows Permission denied after a while |
The saved connection expired. Quit and start clustop again. |
A server is red with timed out |
The machine or its GPU driver is unresponsive. Others keep updating. |
clustop: command not found |
Run pipx ensurepath and open a new terminal. |
| Colors/bars look odd | Use a modern terminal (iTerm2, Terminal.app, Ghostty, Kitty) with a UTF-8 font. |
| Want to drop all saved connections | rm ~/.ssh/clustop-* |
Updating or uninstalling
pipx upgrade clustop # update to the newest release
pipx uninstall clustop # remove
Project layout
clustop/
├── pyproject.toml # package + the `clustop` command
├── README.md
└── clustop/
├── app.py # ssh polling, parsing, rendering
└── config.py # profiles and the `clustop profile` command
Requires: Python 3.11+, rich (installed
automatically), and ssh on your PATH. Works on macOS and Linux.
License
MIT, see LICENSE.
Metadata
Release files for clustop 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| clustop-0.1.0.tar.gz | 16.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| clustop-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 30.1 kB
Release files / clustop-0.1.0.tar.gz
| Download URL | clustop-0.1.0.tar.gz |
|---|---|
| Size | 16.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
12dfaae04e517a6e35bbcb7fb7d912a73b611e86bc516e717bc13d27058c22a4
|
|
BLAKE2b-256 checksum How to use checksums |
38390beb8c106d1517e81d6455299c256ab129103f0abd7beb709eedc05454e4
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.0
|
Release files / clustop-0.1.0-py3-none-any.whl
| Download URL | clustop-0.1.0-py3-none-any.whl |
|---|---|
| Size | 13.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
4c12453625053426d0f9c60895ed8d0cb2925e15158912b0967f0b6794f12d41
|
|
BLAKE2b-256 checksum How to use checksums |
07c923d7c5733fad03b5e20f5e0d0b9ef3ebc1e8919aa850176064f56cef7fde
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.0
|