gpulander
Opportunistic AWS GPU grabber. Point it at the GPU you want; it polls run-instances across
regions and availability zones (spot by default), and the instant capacity appears it lands one box,
writes launched.json, and exits 0. That's the whole trick: the exit is a signal. Run it under an
agent's background-task runner and the exit wakes the agent to SSH in, provision, and train — so a
multi-hour "wait for a scarce GPU" costs zero model tokens until there's actually something to do.
New Blackwell/Hopper SKUs (g7e, H200) are routinely InsufficientInstanceCapacity on both spot and
on-demand across every US region. gpulander exists to wait that out without you babysitting the console.
pip install gpulander
gpulander --list # the GPU SKU catalog (what you can hunt)
gpulander accounts # your AWS profiles -> which account each maps to
gpulander check --gpu rtxpro6000 # read-only: placement score, offered AZs, live spot $/hr
gpulander grab --gpu rtxpro6000 --name job # poll + grab (spot); writes launched.json; exit 0 on win
Needs the AWS CLI configured (a profile, env creds, or an instance role) and bash. Nothing is
launched by --list or check.
Pick a GPU
Three ways to say what you want (choose one); the resolved list is tried cheapest / fewest-GPU first:
| selector | example | meaning |
|---|---|---|
--gpu NAME |
--gpu rtxpro6000 |
by GPU family (see table below) |
--min-vram GB |
--min-vram 48 |
cheapest single-GPU type with ≥ that per-GPU VRAM |
--instance T[,T2] |
--instance g7e.2xlarge,g7e.4xlarge |
exact instance type(s), in priority order |
--gpu names: rtxpro6000 · l40s · l4 · a10g · h100 · h200 · a100-80 · a100-40 · t4 · v100.
SKU catalog (gpulander --list)
| instance | GPU | VRAM | #GPU | ~$OD/hr | note |
|---|---|---|---|---|---|
| g7e.2xlarge | RTX PRO 6000 Blackwell | 96 GB | 1 | 3.36 | cheapest single 96 GB |
| g6e.2xlarge | L40S | 48 GB | 1 | 2.24 | great mid-size |
| g6.2xlarge | L4 | 24 GB | 1 | 0.98 | cheap |
| g5.2xlarge | A10G | 24 GB | 1 | 1.21 | cheap |
| g4dn.xlarge | T4 | 16 GB | 1 | 0.53 | tiny/cheap |
| p5.48xlarge | H100 | 80 GB | 8 | 98.3 | 8×H100, 640 GB total |
| p5e.48xlarge | H200 | 141 GB | 8 | ~110 | 8×H200, 1128 GB total — the only H200 SKU on AWS |
| p4de.24xlarge | A100-80 | 80 GB | 8 | 40.9 | 8×A100-80 |
There is no single-GPU H200 on AWS —
--gpu h200resolves to the 8×H200p5e.48xlarge(big and pricey; ~$25/hr on spot, ~$110/hr on-demand). For a single big GPU,--gpu rtxpro6000(96 GB) is it.
Commands
gpulander accounts — which accounts can I hunt in?
gpulander never embeds an account — it uses the standard AWS CLI credential chain, where each
profile in ~/.aws (SSO or access keys) maps to one account/role. accounts enumerates your
configured profiles and, for each, prints the 12-digit account id and whether its creds currently work:
$ gpulander accounts
Configured AWS profiles (from ~/.aws/config & ~/.aws/credentials):
PROFILE ACCOUNT STATUS
dev 1111.... ok arn:aws:sts::1111...:assumed-role/...
prod 2222.... ok arn:aws:sts::2222...:assumed-role/...
sandbox - NEEDS AUTH (Token has expired)
Poll one or several: gpulander grab --gpu rtxpro6000 --profile p1,p2,p3
Refresh SSO creds: aws sso login --profile <name>
Add a profile with aws configure --profile NAME (keys) or aws configure sso (SSO).
Multiple accounts — poll several at once
Point --profile at one account, or a comma list to sweep several every round. The first account
to land a box wins, and launched.json records which one:
gpulander check --gpu rtxpro6000 --profile dev,prod # compare availability across accounts
gpulander grab --gpu rtxpro6000 --profile dev,prod,sandbox # hunt across 3 accounts simultaneously
Per account, the per-region keypair/security-group/AMI are created and cached independently (keyed by
profile), so one run can safely span accounts. With no --profile, gpulander uses $AWS_PROFILE /
$GPULANDER_PROFILE / the default chain (a single account).
gpulander check — read-only availability
For each resolved instance type: the spot placement score per region/AZ (1 = low … 10 = high chance
of actually getting spot), which AZs offer the type, and the latest spot $/hr (cheapest first).
Launches nothing, spends nothing. Use it to pick market + region before you grab.
gpulander grab — poll until one lands, then exit
Loops run-instances across the chosen regions/AZs for the resolved type(s). On the first success it
writes $GPULANDER_HOME/runs/<name>/launched.json and exits 0. Every launched box carries a
user-data self-terminate (--cap-hours, default 4h) and instance-initiated-shutdown-behavior=terminate,
so it can never run forever even if nothing tears it down.
gpulander grab --gpu rtxpro6000 --name gemma --spot --deadline 240 --cap-hours 3
gpulander grab --gpu h200 --regions us-east-2 --name ornith --spot # ~$25/hr 8×H200
gpulander grab --min-vram 48 --either --max-price 3.00 --name midsize
gpulander grab --instance g7e.2xlarge --dry-run # see the plan, launch nothing
Options (grab)
| flag | default | meaning |
|---|---|---|
--spot / --on-demand / --either |
--spot |
market (--either tries spot then on-demand per AZ) |
--max-price X |
1.25× the catalog on-demand hint | spot ceiling $/hr |
--regions r1,r2 |
us-east-1,us-east-2,us-west-2 |
where to hunt |
--profile p1,p2 |
$AWS_PROFILE / $GPULANDER_PROFILE / default chain |
AWS profile(s) = account(s); comma list sweeps several |
--name TAG |
grab |
instance Name tag + runtime dir |
--deadline MINS |
240 |
give up after N minutes → exit 7 |
--cap-hours H |
4 |
box self-terminates after H hours (cost guard) |
--dry-run |
off | resolve + print the plan; launch nothing |
The background-exit pattern (why grab exits instead of training)
grab deliberately does only the grab, then exits — it does not run the training. That keeps the
wait dumb and token-free, and makes the win an event:
Bash(run_in_background): gpulander grab --gpu rtxpro6000 --name gemma --spot
└─ detached shell polls across turns, spending no model tokens
└─ on the first i-… it writes launched.json and `exit 0`
└─ the EXIT re-invokes the agent (same stream as a task notification)
└─ the agent reads launched.json and does the smart part: SSH, provision, train, hand off
Exit codes double as typed wake reasons so the woken agent can branch without parsing logs:
| code | meaning |
|---|---|
0 |
grabbed — see launched.json ({id, region, az, instance, market, key}) |
7 |
deadline reached, no capacity |
2 |
bad args / setup error |
Agent skill
The package ships a SKILL.md card (trigger-phrase description + the recipe) for Claude Code / Codex /
other agent runtimes. Install or inspect it straight from the CLI:
gpulander --skill # print the SKILL.md card
gpulander --skill list # one-line name + summary
gpulander --skill export # tar of a gpulander/ skill dir (SKILL.md + the script) to stdout
gpulander --skill install # install into ~/.claude/skills, ~/.codex/skills, ~/.agents/skills
install is hash-guarded: if you've edited an installed SKILL.md, an upgrade leaves it alone.
Environment
| var | default | meaning |
|---|---|---|
GPULANDER_HOME |
~/.gpulander |
runtime/state root (runs/<name>/ holds the SSH key, launched.json, status) |
GPULANDER_PROFILE |
— | default AWS profile(s) when --profile isn't given (comma list ok) |
Per-run state (an ed25519 keypair, launched.json, status) lives in $GPULANDER_HOME/runs/<name>/.
The keypair and SG (gpulander-<name>) are created per region on demand; SSH ingress is opened only from
the caller's current public IP (/32).
Requirements
- AWS CLI v2, configured with credentials that can
ec2 run-instances/describe-*/get-spot-placement-scoresandssm get-parameter(for the Deep Learning base AMI). - bash and ssh-keygen on
PATH. - Enough vCPU service quota for G/P spot or on-demand in your target regions (one
*.2xlarge= 8 vCPU).
License
MIT.
watch (read-only observatory)
gpulander watch run --hours 24 --interval 600 --profile a,b,c records spot placement scores, current spot prices and 24h Capacity Block offers (p5/p5e/p5en/p6-b200/b300/trn/p4) for 18 GPU types across 8 regions into SQLite (~/.gpulander/watch/availability.sqlite). It never launches, reserves or buys. gpulander watch report prints a summary. On-demand capacity has no read-only API (run-instances --dry-run always says it would succeed), so it is deliberately not probed.
Metadata
Release files for gpulander 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| gpulander-0.1.0.tar.gz | 20.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| gpulander-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 42.8 kB
Release files / gpulander-0.1.0.tar.gz
| Download URL | gpulander-0.1.0.tar.gz |
|---|---|
| Size | 20.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
b20f34febc7fa7a913199aa0742419cb63e9aff128a9cb68f656f0abd7a1ccf0
|
|
BLAKE2b-256 checksum How to use checksums |
97799c5f7f472c47128dd942e685d90cc57298b2c907eb0e6e886ce5bfe0f39e
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.7
|
Release files / gpulander-0.1.0-py3-none-any.whl
| Download URL | gpulander-0.1.0-py3-none-any.whl |
|---|---|
| Size | 22.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
bed8350ee3bd515503f80d32cdd0cfdb8df0525f8dfad88ab1cbbb8a39e607cb
|
|
BLAKE2b-256 checksum How to use checksums |
f97ebcda0bc0111e680b771e4be92d9ac12854cc8bd19597b8fbb3484a585863
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.7
|