Skip to main content

open-lease

Make GPU infrastructure programmable. Provision GPUs, deploy open-source LLMs, manage their lifecycle with a reconcile loop, and serve inference through an OpenAI-compatible API. One orchestration core, thin interfaces over it. RunPod is Provider #1.

The product is the orchestration layer, not the provider: two seams (Provider, Runtime), one facade (Orchestrator), one vocabulary (DeploymentState), one contract (models.py). A deployment is driven by a reconcile loop comparing desired vs observed state, so interruption and crash recovery are free, and the cost-safety invariant (no orphaned pods burning money) holds by construction.

Status: v0.3.0 on PyPI, beta. The engine is validated against real RunPod (deploy, kill-and-recover, crash-resume, orphan sweep, concurrent deploys, runtime-crash cap), and drives from the CLI, REST API, MCP server, or the visual workbench. Single provider and the §18 24h soak remain. See What's not done.

Quickstart

The base install is the engine, CLI, and the OpenAI proxy. The REST API and MCP server are optional extras (open-lease[api], open-lease[mcp], or open-lease[all]) so the core stays lean.

pip install open-lease                      # base: CLI + OpenAI proxy
# add the REST API and MCP server:  pip install 'open-lease[all]'

# credentials (see docs/configuration.md)
export RUNPOD_API_KEY=...                    # and HF_TOKEN for gated models; or use a .env

gpu models                                  # the model catalog
gpu availability qwen3-0.6b                 # which data centers can run it right now
gpu deploy qwen3-0.6b --wait                # provision + wait for READY
gpu status                                  # id, state, endpoint, uptime, accrued $
gpu stop <id>                               # tear down; verify with `gpu status`

First deploy of a model is download-bound: the vLLM image and the weights are pulled onto an ephemeral disk, so a small model is ready in a few minutes and a large one takes longer. An opt-in model cache (cache_volume_enabled) makes warm redeploys fast. gpu status shows a percent or an elapsed/budget ETA so a cold start never looks stuck.

Talk to a deployed model

gpu proxy                                   # OpenAI-compatible proxy on :8080 (another terminal)
curl localhost:8080/v1/chat/completions \
  -H 'content-type: application/json' \
  -d '{"model":"qwen3-0.6b","messages":[{"role":"user","content":"hi"}]}'

The proxy routes by the request model field (catalog id or HF repo) to the matching READY deployment. Or hit deployment.endpoint_url from gpu status directly.

Run it in the background

gpu up            # start the daemon (reconcile/health/orphan sweep) and proxy, detached
gpu deploy qwen3-0.6b     # non-blocking; the daemon drives it to READY
gpu down          # stop both

How it works

  • Core (core/): orchestrator.py (the §7.1 facade), reconciler.py (the reconcile loop), health.py, costs.py, catalog.py, daemon.py.
  • Provider seam (providers/): provisions compute, knows nothing about LLMs. RunPod + an in-memory mock; new providers are an ABC + a dict entry, verified by one contract suite.
  • Runtime seam (runtimes/): serves a model on compute, knows nothing about providers. vLLM.
  • Interfaces: the gpu CLI, a REST API (gpu serve, routes mirroring the Orchestrator, the OpenAI proxy mounted at /v1/*, auto-docs at /docs), and an MCP server (gpu-mcp, agent-facing tools over the same core). All three reach the same surface, capacity controls included, so an agent or an HTTP client can set a ceiling and not just spend. A Swamp extension is specified for later and consumes the REST API.

See docs/architecture.md for the full picture, and requirements/gpu-orchestrator-requirements.md for the authoritative spec.

Any vLLM-servable model

The catalog holds curated, GPU-tuned recipes and marks which are validated on real hardware, but the engine is model-neutral. To run a model that is not in the catalog, pass its HF repo directly:

gpu deploy --hf-repo Qwen/Qwen3-14B --gpu A100-80GB --wait   # no catalog entry needed

--gpu is required (an ad-hoc model has no recommended GPU); --context, --image, and --disk tune the rest, and --set passes vLLM flags. The deployment carries its own repo id, so it needs no catalog lookup to reconcile or route. Agents get the same via the deploy_hf_model MCP tool. Adding a curated, validated entry to the catalog is a small TOML edit (see CONTRIBUTING.md).

Docs

What's not done

  • RunPod is the only real provider; the seam is proven but no second provider yet.
  • Only the three qwen models are validated on real hardware; llama-3.1-8b is unvalidated (Meta gating, HF access pending). This is validation coverage, not a limit: any vLLM-servable model runs today via gpu deploy --hf-repo (see "Any vLLM-servable model").
  • Gauntlet §18: the 24h soak is not yet run. OOM/terminal-failed is closed (a runtime-crash cap drives a persistently-failing deploy to terminal FAILED; covered by an offline test).
  • The warm-cache speedup is proven mechanically but its timing is capacity-pending.

Development

uv sync --extra dev
uv run python -m pytest tests/ -q
uv run ruff check src/ tests/ && uv run ruff format --check src/ tests/

The gpu CLI and gpu_orchestrator import package keep their names for now; the distribution is open-lease. See CONTRIBUTING.md for the dev setup, architecture constraints, and how to add a provider; build order and non-negotiable rules are in CLAUDE.md.

License

Apache-2.0. See LICENSE.

Metadata

Release files for open-lease 0.5.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for open-lease 0.5.1
File Size Uploaded
open_lease-0.5.1.tar.gz 1.9 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for open-lease 0.5.1
File Interpreter ABI Platform
open_lease-0.5.1-py3-none-any.whl Python 3 none any Details

Total release size: 2.6 MB

Release files / open_lease-0.5.1.tar.gz

Download URL open_lease-0.5.1.tar.gz
Size 1.9 MB
Tags Source
SHA-256 checksum
How to use checksums
676016bd3a62234eebd0f29818e41815d124d07103563adfb02095f2e1b3dda4
BLAKE2b-256 checksum
How to use checksums
de1d4e6ce9c7bc6e2c0b8f18e9498bd196ec4a3a29be500920b9087a54bb2702
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 25, 2026.

Transparency log

Release files / open_lease-0.5.1-py3-none-any.whl

Download URL open_lease-0.5.1-py3-none-any.whl
Size 665.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
5cf80718706b3ed946f75e29b6aeb45f051012aa5dacd58d03ec33a8d94e8da4
BLAKE2b-256 checksum
How to use checksums
39e8cb57eb8dad54bdc663b044eaad078775730306c0bef5949543ab8adea96b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 25, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.5.1 This release

2 release files

0.5.0

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page