Skip to main content

Render + policy library for SLURM JobSpecs (Mila + DRAC clusters)

Project description

salvo

A SLURM job is a shell script with a few #SBATCH headers. The catch: getting the headers right takes runs. Too little memory and the job dies after hours of compute. Too little time and it requeues. The wrong partition and it sits in the queue for a day. Each correction is a manual edit, push, resubmit.

Salvo replaces that loop with a typed JobSpec, a tiny DSL for what to do on OOM or preempt, a memory estimator that reads sacct history, and an sbatch-shelling submitter that records each hop. No SSH, no async.

Install

pip install pysalvo
salvo doctor                                 # pre-flight checks
salvo render spec.yaml --cluster mila        # JobSpec YAML to sbatch text

Render

One source of truth for sbatch text; same inputs, same bytes every run.

from salvo import JobSpec, render

spec = JobSpec(
    name="train", cmd=["python", "train.py"],
    gpus=1, cpus=8, mem="32G", time="2h",
    on_oom=["bump_mem(1.5x, max=128G)", "fail"],
)

sbatch_text = render(spec, cluster_id="mila", account="mila", partition="unkillable")

Leave account/partition off and salvo picks them via salvo.dispatch (login-node only, see below). on_oom is a list of DSL strings, validated when policy runs.

Submit

salvo.submit shells out to sbatch, writes the rendered script plus spec.json and cluster.json into ~/.salvo/runs/<job_id>/, and returns a JobHandle.

from salvo import submit

handle = submit(spec)                                  # cluster auto-detected
handle.status().state                                  # PENDING | RUNNING | COMPLETED | OUT_OF_MEMORY | ...
for line in handle.logs(stream="stdout"):
    print(line)
handle.hops()                                          # list[Hop] read from hops.jsonl
handle.cancel()                                        # scancel

submit(spec, hop="3/5", parent_artifact_dir=...) is used by salvo.runtime.epilog for resubmit chains; user code rarely sets those directly.

Decorator

@cluster.submit(...) wraps a module-level Python function as a remote-callable job. The decorated object is still callable locally; .submit(**kwargs) enqueues it on SLURM.

from salvo import cluster

@cluster.submit(gpus=2, mem="32G", time="1h")
def train(seed: int, lr: float) -> None:
    ...

handle = train.submit(seed=0, lr=1e-3)

Signature is preserved through the wrapper (mypy --strict-friendly via ParamSpec). .submit() validates kwargs against the function signature, then JSON-round-trips them before constructing the JobSpec. Closures, lambdas, and <locals>-defined functions are rejected at decoration time (they cannot be re-imported on a worker).

OOM policy

Declare the recovery strategy once instead of hand-editing --mem after every failure.

from salvo.policy import parse, apply_oom, OomContext

steps = parse(["bump_mem(1.5x, max=128G)", "escalate_partition", "fail"])
new_spec, action = apply_oom(prev_spec, OomContext(kind="cpu", max_rss_mb=33_500))

DSL steps, applied in order: bump_mem(<f>x, max=<size>), escalate_partition, fail.

Memory estimator

Past sacct rows for the same (script, commit, args) triple become a P95 + safety estimate, so the next submit doesn't have to guess.

from salvo.history import spec_key, estimate_mem, JobRecord

key = spec_key("train.sh", git_commit="cd1a0b4", program_args=("--seed", "0"))
records: list[JobRecord] = ...                         # cluv or your own cache supplies these
est = estimate_mem(records, safety=1.2, window=20, min_samples=3)

if est.mem_mb is not None:
    spec = spec.model_copy(update={"mem": f"{est.mem_mb}M"})

spec_key is a 12-byte blake2s; a code change resets history. estimate_mem returns MemEstimate(mem_mb, confidence, n_samples, p95_mb, growth_slope_mb_per_run, rationale). COMPLETED jobs with MaxRSS < 5% of ReqMem are treated as degenerate sacct sampling and fall back to ReqMem. Salvo owns the math; the caller owns where the records live.

Preempt

End-to-end resubmit. When spec.on_preempt == "resubmit", render emits #SBATCH --signal=SIGUSR1@90 and a bash trap that invokes python -m salvo.runtime.epilog --artifact-dir "$SALVO_ARTIFACT_DIR" --reason preempt. The epilog reads the parent's spec.json + cluster.json, bumps SALVO_HOP, calls salvo.submit for the child, and appends a Hop to hops.jsonl.

from salvo.job.preempt import next_hop, should_resubmit, strip_account_suffix

new_hop, max_exceeded = next_hop("2/5")                # ("3/5", False)
strip_account_suffix("rrg-bengioy-ad_gpu")             # "rrg-bengioy-ad"

max_hops (default 5, capped at 20) bounds the chain.

Topology and dispatch

Cluster knowledge as data: each cluster's accounts, partitions, capacity rules, and login constraints live in one YAML.

from salvo.topology import load_preset, list_presets
from salvo.dispatch import pick_account, pick_partition, CapsTracker

cluster_obj = load_preset("mila")
snap = CapsTracker(cluster_id="mila", user="wietze").snapshot()
account = pick_account(spec, cluster_obj, snap)
partition = pick_partition(spec, cluster_obj, account)

Five presets ship: mila, rorqual, narval, beluga, cedar. Adding one is a single YAML. Dispatch is login-node only (shells out to squeue); library callers should pass account/partition to render() directly.

Python entrypoint

The rendered sbatch script for a PythonEntrypoint cmd invokes:

python -m salvo.runtime.entrypoint '{"target":"pkg.mod:fn","kwargs":{...}}'

Stdlib-only shim. Exit 2 on payload error; user exceptions propagate naturally so SLURM marks the job FAILED.

Doctor

salvo doctor runs pre-flight checks (cluster detected, ssh alias not an FQDN, manifest fresh) and prints OK / WARN / FAIL with a one-line fix per check.

With cluv

cluv handles SSH and sbatch from a laptop; salvo handles policy, memory, and on-cluster submit. Opt in via pyproject.toml:

[tool.cluv.retry]
on_oom = ["bump_mem(1.5x, max=128G)", "fail"]
max_hops = 5

[tool.cluv.estimate]
enabled = true

Two example wirings under examples/ cover cluv (policy + render) and xgenius (policy-only).

License

MIT. See LICENSE.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

pysalvo-0.4.0.tar.gz (111.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

pysalvo-0.4.0-py3-none-any.whl (47.0 kB view details)

Uploaded Python 3

File details

Details for the file pysalvo-0.4.0.tar.gz.

File metadata

  • Download URL: pysalvo-0.4.0.tar.gz
  • Upload date:
  • Size: 111.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for pysalvo-0.4.0.tar.gz
Algorithm Hash digest
SHA256 6b4529142b50d5063ff77f1c70fd2168e7862d1930d61133063cb0650ff6bd87
MD5 d778bd8834d2b7d2d7d3eeb36cfd3140
BLAKE2b-256 3489b8139c0a44c76463aea51e37c5d6a7b8719008eba592e4b6a0ab39d431ce

See more details on using hashes here.

Provenance

The following attestation bundles were made for pysalvo-0.4.0.tar.gz:

Publisher: release.yml on wietzesuijker/salvo

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file pysalvo-0.4.0-py3-none-any.whl.

File metadata

  • Download URL: pysalvo-0.4.0-py3-none-any.whl
  • Upload date:
  • Size: 47.0 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for pysalvo-0.4.0-py3-none-any.whl
Algorithm Hash digest
SHA256 fc7a2e3ee1b6dce18a06a3b925c2b1c7ce66252a5f08056da4b45739c8d47226
MD5 b7a6f08e9f8b1689ea3d6184f0d08c10
BLAKE2b-256 19b89dc136acaf72d2d149afa2a604f366849f819fed789c61e7d706d0e65c61

See more details on using hashes here.

Provenance

The following attestation bundles were made for pysalvo-0.4.0-py3-none-any.whl:

Publisher: release.yml on wietzesuijker/salvo

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page