clust — where should I submit this job?
clust reads SLURM directly and answers the two questions that actually decide
where a job goes: how long will it survive there, and when will it start.
It groups the QOSes you can submit to by how long a job is guaranteed to run
before anything can preempt it, shows how much eviction pressure is really queued
behind each one, and will tell you whether a given job shape can start right now.
A full refresh takes well under a second on a cluster with thousands of nodes and jobs. All examples below are rendered from a fictional cluster.
Install
pip install slurmclust # installs the `clust` command
No third-party Python dependencies. Needs scontrol and squeue on PATH;
sacctmgr too for quotas and preemption. Tested against SLURM 25.05; the
squeue --Format fields it relies on (tres-alloc, PendingTime, Reason,
Account) have been available since SLURM 20.11.
Usage
clust # your QOSes, grouped by safe duration; fit for 8 GPUs / 1h
clust --fit 32x2d # test a different job shape
clust --no-fit # drop the START column
clust -q gpu_shared # per-user breakdown of one QOS — who do I ping?
clust -a # whole cluster: partitions, free capacity, top users, queue
clust -p gpu,gpu-big # cluster view restricted to some partitions
clust -A vision # focus on an account instead of your own
clust -u someone # resolve QOS access as if you were another user
clust -g a100 # restrict to one GPU type (case-insensitive substring)
clust -w 10 # refresh every 10s
clust --ascii # no non-ASCII characters
Default view — grouped by how long you are safe
CLUSTER mycluster 96 nodes · 672 GPUs · 79.6% used 20:34:39
535 used · 8 blocked · 18 drain · 8 down · 103 free · 9 running · 6 queued
YOUR QOS user alice
SAFE FOR ANY DURATION never preempted
QOS PART QUOTA USED %QUOTA PEND MAXWALL TOP CONSUMER
gpu_dev - 64 0 0% - 1d00h -
gpu_vision_high gpu-big 256 224 88% 96 – alice (128)
SAFE FOR 1h00m, THEN PREEMPTIBLE requeue, 10m grace
QOS PART QUOTA USED %QUOTA PEND PREEMPTORS THREAT TOP CONSUMER
gpu_shared gpu 512 128 25% 16 2 2j / 96g dave (80)
gpu_scavenge gpu 672 16 2% 64 3 3j / 112g dave (16)
Reading it:
- The group header is the headline. A QOS no other QOS is allowed to preempt
is safe for as long as you like (or until
MAXWALL). Everything else is safe only for thePreemptExemptTimewindow, and then fair game. PREEMPTORSis how many QOSes are allowed to evict you.THREATis what those QOSes have queued and able to run —2j / 96gmeans 2 jobs wanting 96 GPUs are waiting on QOSes that can take your nodes. Held, dependency-blocked, deferred and invalid jobs are excluded, because they are not going to evict anyone. That matters more than it sounds: on the cluster this was built against, three quarters of all pending GPU demand was parked like that, so counting it made every QOS look far more dangerous than it was.SHIELDis how many running GPUs the same preemptors would evict before reaching you — work in a lower-priority QOS that shares your preemptors.you firstmeans you are the front line. Treat it as a tendency, not a promise: see the caveats below.- Within a group the best place to submit is the top row.
Which job actually gets preempted
With PreemptType=preempt/qos, SLURM raids the lowest-priority preemptable QOS
first — that is what SHIELD measures. Beyond that, ordering is not something
to plan around:
- Node overlap dominates. Only jobs sitting on nodes the incoming job wants are candidates at all. A job elsewhere in the partition is untouched.
- Unless the cluster sets
PreemptParameters=preempt_strict_order, SLURM may deviate from strict order to find a cheaper set of victims.preempt_youngest_firstinverts it to newest-first where configured. - The exempt window is the one hard guarantee.
PreemptExemptTimeis enforced; nothing touches a job younger than that.
If you want the empirical answer for your own cluster, sacct records it:
sacct -X -a -S now-2days -o JobID,User,QOS,Elapsed,AllocTRES,State | grep PREEMPTED
So: a 30-minute job can go anywhere. A 3-day job that cannot checkpoint belongs
in the top group. A long job that can checkpoint can take gpu_shared, and
THREAT tells you how often you should expect to be requeued.
--fit — could this job start right now?
$ clust --fit 8x4h
YOUR QOS user alice · can a 8-GPU / 4:00:00 job start?
SAFE FOR ANY DURATION never preempted
QOS PART QUOTA USED %QUOTA PEND MAXWALL TOP CONSUMER START
gpu_dev - 64 0 0% - 1d00h - now
gpu_vision_high gpu-big 256 224 88% 96 – alice (128) now
SAFE FOR 1h00m, THEN PREEMPTIBLE requeue, 10m grace
QOS PART QUOTA USED %QUOTA PEND PREEMPTORS THREAT TOP CONSUMER START
gpu_shared gpu 512 128 25% 16 2 2j / 96g dave (80) no space
gpu_scavenge gpu 672 16 2% 64 3 3j / 112g dave (16) no space
best bet · start now in gpu_dev
--fit is on by default at 8x1h, since it costs nothing. SPEC is
<gpus>x<time> — 8x4h, 32x2d, 1x30m, or a raw SLURM duration
(4x2-00:00:00). Rows that can start now sort first. --no-fit drops it.
The walltime is not decoration: it is checked against the QOS MaxWall and the
partition MaxTime, so --fit 8x2d reports too long for a QOS that caps jobs
at a day. It also drives the warning when your job cannot finish inside the
preemption-safe window.
START names the binding constraint, so you know what to change:
| verdict | meaning |
|---|---|
now |
quota, per-user cap and free nodes all allow it this instant |
quota full |
the QOS is at its GrpTRES gres/gpu limit |
your cap |
you are at the QOS MaxTRESPU limit — someone else could run |
job too big |
exceeds the QOS MaxTRESPerJob limit at any time |
no space |
limits are fine, but no node has that many GPUs free right now |
too long |
walltime exceeds the QOS MaxWall or the partition MaxTime |
no gpu |
the QOS forbids GPUs (GrpTRES gres/gpu=0) |
This is computed from cluster state, not asked of the scheduler. clust used to
call sbatch --test-only for a start time, and that was removed because it is
not trustworthy on a busy cluster: it queues the hypothetical job behind every
pending job in the QOS, including the held and dependency-blocked ones that will
never run. On the cluster this was developed against it returned the same
far-future timestamp for a 1-GPU job as for a 64-GPU one, and the same for a
10-minute job as a 7-day one, while 1-GPU jobs in fact started immediately.
There is deliberately no time estimate for the no space case. A number there
would have to guess when running jobs release their GPUs, and jobs almost never
run to their full TimeLimit.
Drill-down — who is holding the GPUs
$ clust -q gpu_shared
gpu_shared 128 / 512 GPUs (25%) +16 pending
safe 1h00m · 2 preemptors holding 2 jobs / 96 GPUs pending · 10m grace · per-user cap 128
USER GPUS SHARE JOBS NODES OLDEST PEND WAITING
dave 80 62% 1 10 1d04h 16 10m
alice 48 38% 1 6 7h05m - -
OLDEST is the age of the user's longest-running job — what makes a ping
defensible. PEND/WAITING show who is about to take more.
Cluster view
$ clust -a
PARTITIONS
PARTITION NODES GPUS USED FREE BLOCKED DRAIN DOWN UTIL
gpu-big 60 480 387 67 0 18 8 81%
gpu 24 192 148 36 8 0 0 77%
cpu 12 – – – – – – cpu only
FREE CAPACITY GPUs obtainable in whole chunks of N per node
PARTITION ≥1 GPU ≥2 GPU ≥4 GPU ≥8 GPU
gpu-big 67 60 60 32
gpu 36 36 24 0
TOP USERS by allocated GPUs · 6 active
USER ACCOUNT GPUS SHARE JOBS NODES OLDEST PEND
alice vision 176 36.1% 2 22 2d06h 32
dave speech 96 19.7% 2 12 1d04h 80
QUEUE 6 pending jobs
PARTITION JOBS GPUS REQ USERS LONGEST WAIT TOP WAITER
gpu-big 4 272 4 3d03h erin (128)
gpu 2 80 1 10m dave (80)
Columns that are zero everywhere are hidden, so a healthy cluster shows a narrower table.
How capacity is counted
Every GPU falls into exactly one bucket, so the columns always add up:
| bucket | meaning |
|---|---|
used |
allocated to a running job |
free |
actually obtainable right now |
blocked |
idle GPU on a healthy node whose CPUs are exhausted |
drain |
idle GPU on a draining node |
down |
GPU on a node that is down or unreachable |
blocked is the "virtual GPU" idea this tool was built around: a job that takes
many CPUs and no GPUs still makes the rest of the node unusable. A node can only
offer floor(free_cpus / cpus_per_gpu) GPUs, whatever its idle GPU count says.
Capacity comes from node state only; per-user, per-QOS and per-account
attribution comes from jobs only. QOS and account usage is in raw allocated
GPUs, because that is what SLURM's GrpTRES limits are enforced against.
In quota columns, – means no GPU ceiling while no gpu means GPUs
forbidden (GrpTRES gres/gpu=0) — opposite meanings that are easy to conflate.
Preemption analysis assumes PreemptType=preempt/qos. Under any other scheme the
QOS Preempt= lists do not govern preemption, so clust says the safe duration
is unknown rather than guessing.
Performance
clust deliberately avoids SLURM's --json output, which serialises the entire
object graph. Measured on a 3,464-node / 4,265-job cluster:
| data | via --json |
what clust uses |
|---|---|---|
| node state | 20 s, 166 MB | scontrol show node -o — 0.07 s |
| job list | 41 s, 45 MB | squeue -O — 0.4 s |
| node list expansion per job | 16 ms/job | pure Python — 0 forks |
All five SLURM calls are issued in parallel, so wall clock is the slowest one,
and --fit adds no calls at all. Use --debug-timing for the breakdown.
Handy commands
scontrol show partition -a | grep PartitionName= # list partitions
sinfo -hNo '%f' | tr -d ' ' | tr ',' '\n' | sort -u # list node features
sacctmgr show qos format=Name,GrpTRES,Preempt # quotas and preemption
scontrol show config | grep -i preempt # the cluster's preempt policy
Development
pip install -e .
python -m pytest tests/ -q # 154 tests, no SLURM required
python -m flake8 clust.py tests/
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file slurmclust-0.5.1.tar.gz.
File metadata
- Download URL: slurmclust-0.5.1.tar.gz
- Upload date:
- Size: 41.4 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.13.5
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
5e3d8fee98b338132dbe743dba14392334d1a40c40f600637e4aba020f1c8f53
|
|
| MD5 |
1a985fab8a9f26eb082f57d6b9a404d2
|
|
| BLAKE2b-256 |
3597ac11a071bd6d6b73760774d2e35d576d9290b8f1723e50dadb9a47c3b441
|
File details
Details for the file slurmclust-0.5.1-py3-none-any.whl.
File metadata
- Download URL: slurmclust-0.5.1-py3-none-any.whl
- Upload date:
- Size: 28.0 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.13.5
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
f7acd7d5ab9e3729bdcc26dc51123c2cc7017da85beb6ebc26f3543bdd7cba53
|
|
| MD5 |
d0ea52ae8a449d3a0ea9d72c0802c358
|
|
| BLAKE2b-256 |
b1e3e600b15004776e3a63cafaef7edb6aaaa3b091720c674cfde1dd6114be26
|