slurm-workflows: HPC workflow helpers for Slurm clusters
slurm-workflows lets you run Python functions on a Slurm cluster
without sbatch scripts written by hand.
It provides an interface inspired by
concurrent.futures.
The interface launches long-lived workers inside pilot jobs.
It then dispatches tasks to those workers.
You pay Slurm's scheduling latency once per pilot job, not once per task.
Use it in three cases:
- You have many Python tasks to run on one cluster allocation.
- An exploration or a search has to spread across a pool of workers.
- Per-worker state is expensive, and you want it to stay warm between tasks.
Installation
A run needs:
- Python >= 3.12
- Access to a Slurm cluster (
sbatch,squeue,scancelonPATH) - A running
ds-serviceserver
pip install -U slurm-workflows
To set up on UVA's Rivanna cluster, read How to install slurm-workflows on Rivanna.
Usage
from ds_service_client import DsServiceServer
from slurm_workflows import SlurmPilotExecutor
def square(x):
return x * x
SETUP_SCRIPT = """
module load gcc/14.2.0
conda activate my-env
"""
with DsServiceServer(interface="ib0") as ds_service:
ds_service.wait_until_ready()
with SlurmPilotExecutor("my-run", ds_service.address) as executor:
# 1. Describe a job group. This submits nothing.
executor.define_job_group(
name="cpu",
sbatch_args=["-A my_alloc", "-p standard", "-t 01:00:00"],
setup_script=SETUP_SCRIPT,
)
# 2. Launch 4 pilot jobs of that job group.
executor.scale_jobs("cpu", 4)
# 3. Submit tasks to a named queue. That job group's workers claim them.
tasks = [executor.submit("cpu", square, i) for i in range(100)]
# 4. Block until every result is in.
executor.wait(tasks, desc="squaring")
# The executor canceled every pilot job at the end of the block.
print(sum(task.output for task in tasks))
328350
You can submit tasks before the workers exist. They wait on the queue until a worker starts and claims them.
Documentation
| Document | What it covers |
|---|---|
| Computing pi on a Slurm cluster | The main features of slurm-workflows, by creating a pool of workers to compute $\pi$. |
| Computing pi with a Sobol' QMC exploration | Using ExploreSpaceSobolQMC to create a space filling design and evaluate it. |
| Optimizing Himmelblau's function | Using OptimizeSpaceBotorch to run a batch Bayesian search. |
| How to install slurm-workflows on Rivanna | Installing the package and the ds-service binary on Rivanna. |
How to run the ds-service server |
Starting a ds-service server from the driver and binding it where workers can reach it. |
| How to keep per-worker state with actors | Loading an expensive model or connection once per worker instead of once per task. |
| How to fold results across workers | Using mapreduce to run one function over a whole collection and bring back a single value. |
How to watch a run with swtop |
Following a live run from another shell, and keeping a record of one. |
| How to troubleshoot a failing run | Finding the right log, and what each RuntimeError means. |
| How to resume a search | Carrying a search on across a Slurm time limit. |
SlurmPilotExecutor |
The executor, Task, RaiseOnError, and the job group options. |
mapreduce |
Mapping an iterable across the pool, the fold contract, and what the call creates on the server. |
| What a run publishes | The environment a task sees, the keys and series a run writes, the worker entry point, and the logs. |
ExploreSpaceSobolQMC |
The Sobol' exploration, its study fields, and the results file. |
OptimizeSpaceBotorch |
The batch Bayesian search, its study fields, and its stopping rule. |
| Search spaces | IntRange, FloatRange and CategoricalRange, the objective contract, and what a failed evaluation does to a run. |
swtop |
The CLI, the blocks on screen, and what the host and job readings measure. |
| About the pilot-job model | Why pilot jobs, the three processes, where the driver runs, and which class to reach for. |
| About batch Bayesian optimization | Why a search has rounds, where the fit runs, and when it is worth the overhead. |
| About what a run publishes | Why a run is observable from outside itself, and the limits of that. |
For contributors
| Document | What it covers |
|---|---|
| Developer notes | The layout of the code, where each kind of documentation goes, and the conventions a change is held to. |
| Terminology | The word this project uses for each concept, in prose and in identifiers, and where each word comes from. |
| How to run the tests | Running the suite, what it mocks, and what it runs for real. |
License
MIT. See LICENSE.
Release files for slurm-workflows 3.0.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| slurm_workflows-3.0.0.tar.gz | 46.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| slurm_workflows-3.0.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 100.3 kB
Release files / slurm_workflows-3.0.0.tar.gz
| Download URL | slurm_workflows-3.0.0.tar.gz |
|---|---|
| Size | 46.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
1fddd17bf30108f9536b52dd42884ccafe030c0657a593730486f036ee30a92c
|
|
BLAKE2b-256 checksum How to use checksums |
ac2695aa83678a0879336e847f6b553ca40e7e72038eebbf10491859778d65c8
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.15
|
Release files / slurm_workflows-3.0.0-py3-none-any.whl
| Download URL | slurm_workflows-3.0.0-py3-none-any.whl |
|---|---|
| Size | 53.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
a91cd7216330ae88ee87ea0d25ece3ed209e1c941c5c97af13fef531f6f4e15f
|
|
BLAKE2b-256 checksum How to use checksums |
81bfe57abbef36ccde27339eecc0c5b53b67ca0d696d7d53b3a3aa297611605f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.15
|