Skip to main content

AI compute arbitrage CLI — move GPU training jobs between clouds automatically

Project description

VaultLayer

Run AI training jobs on managed GPU capacity with checkpointing, log streaming, and provider failover.

pip install -U vaultlayer
vl init
vl run train.py
Job submitted
Training on vast_ai
Training output:
...
Training completed successfully.

What It Does

VaultLayer sits between your training script and the cloud. It:

  • Checkpoints automatically — syncs model weights + optimizer state to a zero-egress R2 Vault on every save
  • Detects interruptions — intercepts AWS/GCP/Azure termination signals before your job dies
  • Migrates instantly — provisions a replacement node on the cheapest available provider and resumes from last checkpoint
  • Tracks savings — shows real-time cost vs what you would have paid on AWS On-Demand

No changes to your PyTorch or JAX code. No YAML configs. No PhD-level infra knowledge required.

Commands

# Training
vl run train.py
vl ps
vl logs <job-id> --follow
vl stop <job-id>

# Dataset storage (no S3 required)
vl sync ./data --dataset-id my-dataset
vl run --data r2://my-dataset train.py
vl datasets

Supported Providers

Provider Type Status
Vast.ai Marketplace Production-included
RunPod Neocloud Production-included
Lambda Labs Neocloud Production-included
AWS Spot Hyperscaler Production-included for validated failover paths
AWS On-Demand Hyperscaler Internal testing
GCP, CoreWeave, Crusoe, Nebius, Voltage Park, Hyperstack, Azure Mixed Pending validation

Current provider status lives in docs/provider_test_matrix.md and docs/provider_testing_matrix.md.


Model Size Support

Model Size Method Checkpoint Size Status
1B QLoRA small Validated smoke path
3B QLoRA small Validated matrix path
7B QLoRA medium Validated matrix path
72B QLoRA large Routed to 96GB+ high-VRAM capacity
Full fine-tune / multi-GPU varies varies Future work

Tech Stack

Layer Technology Cost
Code + Docs GitHub (this repo) Free
CI/CD GitHub Actions Free (2k min/mo)
Vault / Storage Cloudflare R2 Free up to 10GB
Agent Runtime Railway Free $5/mo credit
Webhooks Cloudflare Workers Free 100k req/day
Agent Message Queue Upstash Redis Free 10k cmd/day

Repository Structure

vaultlayer/
├── README.md
├── docs/
│   ├── PRD.md              # Full product requirements
│   ├── ARCHITECTURE.md     # System design + agent topology
│   └── AGENTS.md           # Agent specs + build order
├── dashboard/
│   └── index.html          # Savings dashboard prototype
└── src/
    ├── cli/
    │   ├── main.py
    │   ├── run.py
    │   ├── checkpoint_template.py
    │   └── init.py
    ├── vaultlayer/
    │   └── _resume_hook.py
    ├── agents/
    │   ├── orchestration/
    │   ├── pricing/
    │   ├── watchdog/
    │   │   └── signals.py
    │   ├── vault/
    │   ├── broker/
    │   ├── finops/
    │   └── namespace/
    └── shared/

SLA

VaultLayer tracks job completion, checkpoint persistence, and resume behavior. Public SLA numbers are not committed during beta; see docs/SLA_SLI.md for definitions.


Dataset Storage (No S3 Required)

VaultLayer's Neutral Zone (Cloudflare R2) is a first-class storage provider. Users with no AWS or cloud storage account can upload training data directly and train from it on any provider.

# Upload from your laptop / on-prem server
vl sync ./training-data --dataset-id my-dataset

# Train — data is mounted at /mnt/vaultlayer on every provisioned node
vl run --data r2://my-dataset train.py

# See what you're storing and the monthly cost
vl datasets

Pricing:

Action Cost
Upload (local → R2) Free
Storage $0.020 / GB / month ($0.0195 — 30% markup over Cloudflare R2 base rate)
Read (R2 → training node) $0.00 (zero egress within Cloudflare network)
S3 mirror (one-time) AWS egress charge (~$0.09/GB, first 100 GB/month free)

Storage quotas by plan:

Plan Storage limit
Free 10 GB
Pro 500 GB
Enterprise Unlimited

Datasets are soft-deleted with vl datasets --delete <id> — billing stops immediately, R2 objects are purged within 24 hours.

Getting Started

pip install -U vaultlayer
vl init
vl run train.py

License

Private — © 2026 VaultLayer

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

vaultlayer-0.1.42.tar.gz (1.0 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

vaultlayer-0.1.42-py3-none-any.whl (683.0 kB view details)

Uploaded Python 3

File details

Details for the file vaultlayer-0.1.42.tar.gz.

File metadata

  • Download URL: vaultlayer-0.1.42.tar.gz
  • Upload date:
  • Size: 1.0 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for vaultlayer-0.1.42.tar.gz
Algorithm Hash digest
SHA256 2a5d9a33d5905e736433bd791691e2a0c0cd37e8779c156b3c1a0331a895bcc6
MD5 d92576851c8efd5d89b7ae6bfc4e8958
BLAKE2b-256 8db72fa88cc785db2497cac0cb681af0c2ef0232e82bee78f934df7d0e577521

See more details on using hashes here.

Provenance

The following attestation bundles were made for vaultlayer-0.1.42.tar.gz:

Publisher: publish.yml on hector25/vaultlayer

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file vaultlayer-0.1.42-py3-none-any.whl.

File metadata

  • Download URL: vaultlayer-0.1.42-py3-none-any.whl
  • Upload date:
  • Size: 683.0 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for vaultlayer-0.1.42-py3-none-any.whl
Algorithm Hash digest
SHA256 365214bec9db5ae328ca5aaf54029832bc0a7500afec16572b4e2645444d7176
MD5 8d7d90784b6f44c851520b09e0e53a76
BLAKE2b-256 8b80de170c15df1a273075171a7e6f5de5c64a3b1d521dca3fb9bd5f6ed2134f

See more details on using hashes here.

Provenance

The following attestation bundles were made for vaultlayer-0.1.42-py3-none-any.whl:

Publisher: publish.yml on hector25/vaultlayer

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page