Skip to main content

gpulander

Opportunistic AWS GPU grabber. Point it at the GPU you want; it polls run-instances across regions and availability zones (spot by default), and the instant capacity appears it lands one box, writes launched.json, and exits 0. That's the whole trick: the exit is a signal. Run it under an agent's background-task runner and the exit wakes the agent to SSH in, provision, and train — so a multi-hour "wait for a scarce GPU" costs zero model tokens until there's actually something to do.

New Blackwell/Hopper SKUs (g7e, H200) are routinely InsufficientInstanceCapacity on both spot and on-demand across every US region. gpulander exists to wait that out without you babysitting the console.

pip install gpulander

gpulander --list                              # the GPU SKU catalog (what you can hunt)
gpulander accounts                            # your AWS profiles -> which account each maps to
gpulander check --gpu rtxpro6000              # read-only: placement score, offered AZs, live spot $/hr
gpulander grab  --gpu rtxpro6000 --name job   # poll + grab (spot); writes launched.json; exit 0 on win

Needs the AWS CLI configured (a profile, env creds, or an instance role) and bash. Nothing is launched by --list or check.

Pick a GPU

Three ways to say what you want (choose one); the resolved list is tried cheapest / fewest-GPU first:

selector example meaning
--gpu NAME --gpu rtxpro6000 by GPU family (see table below)
--min-vram GB --min-vram 48 cheapest single-GPU type with ≥ that per-GPU VRAM
--instance T[,T2] --instance g7e.2xlarge,g7e.4xlarge exact instance type(s), in priority order

--gpu names: rtxpro6000 · l40s · l4 · a10g · h100 · h200 · a100-80 · a100-40 · t4 · v100.

SKU catalog (gpulander --list)

instance GPU VRAM #GPU ~$OD/hr note
g7e.2xlarge RTX PRO 6000 Blackwell 96 GB 1 3.36 cheapest single 96 GB
g6e.2xlarge L40S 48 GB 1 2.24 great mid-size
g6.2xlarge L4 24 GB 1 0.98 cheap
g5.2xlarge A10G 24 GB 1 1.21 cheap
g4dn.xlarge T4 16 GB 1 0.53 tiny/cheap
p5.48xlarge H100 80 GB 8 98.3 8×H100, 640 GB total
p5e.48xlarge H200 141 GB 8 ~110 8×H200, 1128 GB total — the only H200 SKU on AWS
p4de.24xlarge A100-80 80 GB 8 40.9 8×A100-80

There is no single-GPU H200 on AWS — --gpu h200 resolves to the 8×H200 p5e.48xlarge (big and pricey; ~$25/hr on spot, ~$110/hr on-demand). For a single big GPU, --gpu rtxpro6000 (96 GB) is it.

Commands

gpulander accounts — which accounts can I hunt in?

gpulander never embeds an account — it uses the standard AWS CLI credential chain, where each profile in ~/.aws (SSO or access keys) maps to one account/role. accounts enumerates your configured profiles and, for each, prints the 12-digit account id and whether its creds currently work:

$ gpulander accounts
Configured AWS profiles (from ~/.aws/config & ~/.aws/credentials):
  PROFILE                    ACCOUNT        STATUS
  dev                        1111....       ok   arn:aws:sts::1111...:assumed-role/...
  prod                       2222....       ok   arn:aws:sts::2222...:assumed-role/...
  sandbox                    -              NEEDS AUTH (Token has expired)

Poll one or several:  gpulander grab --gpu rtxpro6000 --profile p1,p2,p3
Refresh SSO creds:    aws sso login --profile <name>

Add a profile with aws configure --profile NAME (keys) or aws configure sso (SSO).

Multiple accounts — poll several at once

Point --profile at one account, or a comma list to sweep several every round. The first account to land a box wins, and launched.json records which one:

gpulander check --gpu rtxpro6000 --profile dev,prod            # compare availability across accounts
gpulander grab  --gpu rtxpro6000 --profile dev,prod,sandbox    # hunt across 3 accounts simultaneously

Per account, the per-region keypair/security-group/AMI are created and cached independently (keyed by profile), so one run can safely span accounts. With no --profile, gpulander uses $AWS_PROFILE / $GPULANDER_PROFILE / the default chain (a single account).

gpulander check — read-only availability

For each resolved instance type: the spot placement score per region/AZ (1 = low … 10 = high chance of actually getting spot), which AZs offer the type, and the latest spot $/hr (cheapest first). Launches nothing, spends nothing. Use it to pick market + region before you grab.

gpulander grab — poll until one lands, then exit

Loops run-instances across the chosen regions/AZs for the resolved type(s). On the first success it writes $GPULANDER_HOME/runs/<name>/launched.json and exits 0. Every launched box carries a user-data self-terminate (--cap-hours, default 4h) and instance-initiated-shutdown-behavior=terminate, so it can never run forever even if nothing tears it down.

gpulander grab --gpu rtxpro6000 --name gemma --spot --deadline 240 --cap-hours 3
gpulander grab --gpu h200 --regions us-east-2 --name ornith --spot      # ~$25/hr 8×H200
gpulander grab --min-vram 48 --either --max-price 3.00 --name midsize
gpulander grab --instance g7e.2xlarge --dry-run                         # see the plan, launch nothing

Options (grab)

flag default meaning
--spot / --on-demand / --either --spot market (--either tries spot then on-demand per AZ)
--max-price X 1.25× the catalog on-demand hint spot ceiling $/hr
--regions r1,r2 us-east-1,us-east-2,us-west-2 where to hunt
--profile p1,p2 $AWS_PROFILE / $GPULANDER_PROFILE / default chain AWS profile(s) = account(s); comma list sweeps several
--name TAG grab instance Name tag + runtime dir
--deadline MINS 240 give up after N minutes → exit 7
--cap-hours H 4 box self-terminates after H hours (cost guard)
--dry-run off resolve + print the plan; launch nothing

The background-exit pattern (why grab exits instead of training)

grab deliberately does only the grab, then exits — it does not run the training. That keeps the wait dumb and token-free, and makes the win an event:

Bash(run_in_background):  gpulander grab --gpu rtxpro6000 --name gemma --spot
   └─ detached shell polls across turns, spending no model tokens
   └─ on the first i-… it writes launched.json and `exit 0`
   └─ the EXIT re-invokes the agent (same stream as a task notification)
   └─ the agent reads launched.json and does the smart part: SSH, provision, train, hand off

Exit codes double as typed wake reasons so the woken agent can branch without parsing logs:

code meaning
0 grabbed — see launched.json ({id, region, az, instance, market, key})
7 deadline reached, no capacity
2 bad args / setup error

Agent skill

The package ships a SKILL.md card (trigger-phrase description + the recipe) for Claude Code / Codex / other agent runtimes. Install or inspect it straight from the CLI:

gpulander --skill            # print the SKILL.md card
gpulander --skill list       # one-line name + summary
gpulander --skill export     # tar of a gpulander/ skill dir (SKILL.md + the script) to stdout
gpulander --skill install    # install into ~/.claude/skills, ~/.codex/skills, ~/.agents/skills

install is hash-guarded: if you've edited an installed SKILL.md, an upgrade leaves it alone.

Environment

var default meaning
GPULANDER_HOME ~/.gpulander runtime/state root (runs/<name>/ holds the SSH key, launched.json, status)
GPULANDER_PROFILE — default AWS profile(s) when --profile isn't given (comma list ok)

Per-run state (an ed25519 keypair, launched.json, status) lives in $GPULANDER_HOME/runs/<name>/. The keypair and SG (gpulander-<name>) are created per region on demand; SSH ingress is opened only from the caller's current public IP (/32).

Requirements

  • AWS CLI v2, configured with credentials that can ec2 run-instances / describe-* / get-spot-placement-scores and ssm get-parameter (for the Deep Learning base AMI).
  • bash and ssh-keygen on PATH.
  • Enough vCPU service quota for G/P spot or on-demand in your target regions (one *.2xlarge = 8 vCPU).

License

MIT.

watch (read-only observatory)

gpulander watch run --hours 24 --interval 600 --profile a,b,c records spot placement scores, current spot prices and 24h Capacity Block offers (p5/p5e/p5en/p6-b200/b300/trn/p4) for 18 GPU types across 8 regions into SQLite (~/.gpulander/watch/availability.sqlite). It never launches, reserves or buys. gpulander watch report prints a summary. On-demand capacity has no read-only API (run-instances --dry-run always says it would succeed), so it is deliberately not probed.

Metadata

Release files for gpulander 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for gpulander 0.1.0
File Size Uploaded
gpulander-0.1.0.tar.gz 20.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for gpulander 0.1.0
File Interpreter ABI Platform
gpulander-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 42.8 kB

Release files / gpulander-0.1.0.tar.gz

Download URL gpulander-0.1.0.tar.gz
Size 20.4 kB
Tags Source
SHA-256 checksum
How to use checksums
b20f34febc7fa7a913199aa0742419cb63e9aff128a9cb68f656f0abd7a1ccf0
BLAKE2b-256 checksum
How to use checksums
97799c5f7f472c47128dd942e685d90cc57298b2c907eb0e6e886ce5bfe0f39e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.7

Release files / gpulander-0.1.0-py3-none-any.whl

Download URL gpulander-0.1.0-py3-none-any.whl
Size 22.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
bed8350ee3bd515503f80d32cdd0cfdb8df0525f8dfad88ab1cbbb8a39e607cb
BLAKE2b-256 checksum
How to use checksums
f97ebcda0bc0111e680b771e4be92d9ac12854cc8bd19597b8fbb3484a585863
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.7

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page