Unified HPC job submission across multiple schedulers
Project description
hpc-runner
Unified HPC job submission across multiple schedulers
Write your jobs once, run them on any cluster - SGE, Slurm, PBS, or locally for testing.
Features
- Unified CLI - Same commands work across SGE, Slurm, PBS
- Python API - Programmatic job submission with dependencies and pipelines
- Auto-detection - Automatically finds your cluster's scheduler
- Interactive TUI - Monitor jobs with a terminal dashboard
- Job Dependencies - Chain jobs with afterok, afterany, afternotok
- Array Jobs - Batch processing with throttling support
- Virtual Environment Handling - Automatic venv activation on compute nodes
- Module Integration - Load environment modules in job scripts
- Dry-run Mode - Preview generated scripts before submission
Installation
pip install hpc-runner
Or with uv:
uv pip install hpc-runner
Quick Start
CLI
# Basic job submission
hpc run python train.py
# With resources
hpc run --cpu 4 --mem 16G --time 4:00:00 "python train.py"
# GPU job
hpc run --queue gpu --cpu 4 --mem 32G "python train.py --epochs 100"
# Preview without submitting
hpc run --dry-run --cpu 8 "make -j8"
# Interactive session
hpc run --interactive bash
# Array job
hpc run --array 1-100 "python process.py --task-id \$SGE_TASK_ID"
# Wait for completion
hpc run --wait python long_job.py
Python API
from hpc_runner import Job
# Create and submit a job
job = Job(
command="python train.py",
cpu=4,
mem="16G",
time="4:00:00",
queue="gpu",
)
result = job.submit()
# Wait for completion
status = result.wait()
print(f"Exit code: {result.returncode}")
# Read output
print(result.read_stdout())
Job Dependencies
from hpc_runner import Job
# First job
preprocess = Job(command="python preprocess.py", cpu=8, mem="32G")
result1 = preprocess.submit()
# Second job runs after first succeeds
train = Job(command="python train.py", cpu=4, mem="48G", queue="gpu")
train.after(result1, type="afterok")
result2 = train.submit()
Pipelines
from hpc_runner import Pipeline
with Pipeline("ml_workflow") as p:
p.add("python preprocess.py", name="preprocess", cpu=8)
p.add("python train.py", name="train", depends_on=["preprocess"], queue="gpu")
p.add("python evaluate.py", name="evaluate", depends_on=["train"])
results = p.submit()
p.wait()
Scheduler Support
| Scheduler | Status | Notes |
|---|---|---|
| SGE | Fully implemented | qsub, qstat, qdel, qrsh |
| Local | Fully implemented | Run as subprocess (for testing) |
| Slurm | Planned | sbatch, squeue, scancel |
| PBS | Planned | qsub, qstat, qdel |
Auto-detection Priority
HPC_SCHEDULERenvironment variable- SGE (
SGE_ROOTorqstatavailable) - Slurm (
sbatchavailable) - PBS (
qsubwith PBS) - Local fallback
Configuration
hpc-runner uses TOML configuration files. Location priority:
--config /path/to/config.toml./hpc-runner.toml./pyproject.tomlunder[tool.hpc-runner]- Git repository root
hpc-runner.toml ~/.config/hpc-runner/config.toml- Package defaults
Example Configuration
[defaults]
cpu = 1
mem = "4G"
time = "1:00:00"
inherit_env = true
[schedulers.sge]
parallel_environment = "smp"
memory_resource = "mem_free"
purge_modules = true
[types.gpu]
queue = "gpu"
resources = [{name = "gpu", value = 1}]
[types.interactive]
queue = "interactive"
time = "8:00:00"
Use named job types:
hpc run --job-type gpu "python train.py"
SGE Configuration
SGE clusters vary widely in how resources are named. The [schedulers.sge]
section lets you match your site's conventions without touching job definitions.
How job fields map to SGE flags:
| Job Field | SGE Flag | Configurable Via |
|---|---|---|
cpu |
-pe <pe_name> <slots> |
parallel_environment |
mem |
-l <resource>=<value> |
memory_resource |
time |
-l <resource>=<value> |
time_resource |
queue |
-q <queue> |
direct |
resources |
-l <name>=<value> |
direct |
Full [schedulers.sge] reference:
[schedulers.sge]
# Resource naming -- these must match your site's SGE configuration
parallel_environment = "smp" # PE name for CPU slots (some sites use "mpi", "threaded", etc.)
memory_resource = "mem_free" # Memory resource name (common alternatives: "h_vmem", "virtual_free")
time_resource = "h_rt" # Time limit resource name (commonly "h_rt")
# Output handling
merge_output = true # Merge stderr into stdout (-j y)
# Module system
purge_modules = true # Run 'module purge' before loading job modules
silent_modules = false # Suppress module command output (-s flag)
module_init_script = "" # Path to module init script (auto-detected if empty)
# Environment
expand_makeflags = true # Expand $NSLOTS in MAKEFLAGS for parallel make
unset_vars = [] # Environment variables to unset in jobs
# e.g. ["https_proxy", "http_proxy"]
Fully populated config example:
[defaults]
scheduler = "auto"
cpu = 1
mem = "4G"
time = "1:00:00"
queue = "batch.q"
use_cwd = true
inherit_env = true
stdout = "hpc.%N.%J.out"
modules = ["gcc/12.2", "python/3.11"]
resources = [
{ name = "scratch", value = "20G" }
]
[schedulers.sge]
parallel_environment = "smp"
memory_resource = "mem_free"
time_resource = "h_rt"
merge_output = true
purge_modules = true
silent_modules = false
expand_makeflags = true
unset_vars = ["https_proxy", "http_proxy"]
[tools.python]
cpu = 4
mem = "16G"
time = "4:00:00"
queue = "short.q"
modules = ["-", "python/3.11"] # leading "-" replaces the list instead of merging
resources = [
{ name = "tmpfs", value = "8G" }
]
[types.interactive]
queue = "interactive.q"
time = "8:00:00"
cpu = 2
mem = "8G"
[types.gpu]
queue = "gpu.q"
cpu = 8
mem = "64G"
time = "12:00:00"
resources = [
{ name = "gpu", value = 1 }
]
This config can also be embedded in pyproject.toml under [tool.hpc-runner].
TUI Monitor
Launch the interactive job monitor:
hpc monitor
Key bindings:
q- Quitr- Refreshu- Toggle user filter (my jobs / all)/- SearchEnter- View job detailsTab- Switch tabs
CLI Reference
hpc run [OPTIONS] COMMAND
Options:
--job-name TEXT Job name
--cpu INTEGER Number of CPUs
--mem TEXT Memory (e.g., 16G, 4096M)
--time TEXT Time limit (e.g., 4:00:00)
--queue TEXT Queue/partition name
--directory PATH Working directory
--module TEXT Module to load (repeatable)
--array TEXT Array spec (e.g., 1-100, 1-100%5)
--depend TEXT Job dependencies
--inherit-env Inherit environment (default: true)
--no-inherit-env Don't inherit environment
--interactive Run interactively (qrsh/srun)
--local Run locally (no scheduler)
--dry-run Show script without submitting
--wait Wait for completion
--keep-script Keep job script for debugging
-h, --help Show help
Other commands:
hpc status [JOB_ID] Check job status
hpc cancel JOB_ID Cancel a job
hpc monitor Interactive TUI
hpc config show Show active configuration
Development
# Setup environment
source sourceme
source sourceme --clean # Clean rebuild
# Run tests
pytest
pytest -v
pytest -k "test_job"
# Type checking
mypy src/hpc_runner
# Linting
ruff check src/hpc_runner
ruff format src/hpc_runner
License
MIT License - see LICENSE file for details.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file hpc_runner-0.5.0.tar.gz.
File metadata
- Download URL: hpc_runner-0.5.0.tar.gz
- Upload date:
- Size: 103.9 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
e3502e095eedd8ad24814b5836d0e2cb7c679fda9bcac970c33fd051b7e50182
|
|
| MD5 |
df9415830feb7ad848509215de618548
|
|
| BLAKE2b-256 |
a7f8874639eb0f955b61726b3c6b1a001113814d125d0e0a9da8c28832cd72cf
|
Provenance
The following attestation bundles were made for hpc_runner-0.5.0.tar.gz:
Publisher:
publish.yml on sjalloq/hpc-runner
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
hpc_runner-0.5.0.tar.gz -
Subject digest:
e3502e095eedd8ad24814b5836d0e2cb7c679fda9bcac970c33fd051b7e50182 - Sigstore transparency entry: 953621028
- Sigstore integration time:
-
Permalink:
sjalloq/hpc-runner@71263b46dbc6710901b99bc914841d80b588feae -
Branch / Tag:
refs/tags/v0.5.0 - Owner: https://github.com/sjalloq
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@71263b46dbc6710901b99bc914841d80b588feae -
Trigger Event:
push
-
Statement type:
File details
Details for the file hpc_runner-0.5.0-py3-none-any.whl.
File metadata
- Download URL: hpc_runner-0.5.0-py3-none-any.whl
- Upload date:
- Size: 83.0 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
fcfb7ade28b90bbc100a60a7dec60ae69ff553c16a69a3ae2991c8f24c61c65d
|
|
| MD5 |
cdc348e4a423b91062b4e39f2ddceafb
|
|
| BLAKE2b-256 |
318248812a749d56677686b72194bcfe8aa2764d8645d83a23fc62eed8e96e4d
|
Provenance
The following attestation bundles were made for hpc_runner-0.5.0-py3-none-any.whl:
Publisher:
publish.yml on sjalloq/hpc-runner
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
hpc_runner-0.5.0-py3-none-any.whl -
Subject digest:
fcfb7ade28b90bbc100a60a7dec60ae69ff553c16a69a3ae2991c8f24c61c65d - Sigstore transparency entry: 953621030
- Sigstore integration time:
-
Permalink:
sjalloq/hpc-runner@71263b46dbc6710901b99bc914841d80b588feae -
Branch / Tag:
refs/tags/v0.5.0 - Owner: https://github.com/sjalloq
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@71263b46dbc6710901b99bc914841d80b588feae -
Trigger Event:
push
-
Statement type: