OPBDH
Open the Pod Bay Door, Hal
Full documentation at opbdh
— “Of course I can do it, Dave.”
Launches a GPU pod, runs your model-backed script on it, syncs the results home, and deletes the pod.
First launch
pip install opbdh
opbdh
Running opbdh unconfigured starts a setup wizard: it picks a provider (RunPod or Prime Intellect), checks your API token (RUNPOD_API_TOKEN / PRIME_INTELLECT_API_KEY), and asks for defaults — model, code path, minimum VRAM, price caps — saved globally or per project.
Needs macOS or Linux, Python ≥ 3.11, and ssh/scp on your PATH. An existing ~/.ssh keypair is used if present, otherwise one is generated under ~/.config/opbdh/ssh/. Set HF_TOKEN for private/gated Hugging Face models.
Launch a pod
opbdh launch ./run.py --model Qwen/Qwen2.5-0.5B-Instruct --vram-gb 48 --max-spend 5
This verifies your code, picks the cheapest fitting GPU, launches the pod, runs your command, streams remote logs/ and results/ into runpod_results/<run_id>/, stops the run if estimated spend crosses the cap, and deletes the pod when it finishes — or fails. Add --dry-run to print the plan without contacting the provider; real launches ask for confirmation unless --yes.
Fine-tune a model
Install the optional fine-tuning support, then start the interactive SFT console:
pip install "opbdh[ft]"
opbdh ft
The native format is intentionally easy to edit. A base-model example can be one small TOML file:
input = "Translate hello to Italian"
output = "Ciao"
tags = ["translation", "short"]
For a chat model, input is the conversation and output is the assistant response to learn:
output = "Ciao! Come posso aiutarti?"
tags = ["italian", "greeting"]
[[input]]
role = "system"
content = "Reply in Italian."
[[input]]
role = "user"
content = "Hello!"
A file may contain one example or an [[examples]] list, and a dataset may mix files and directories. In the console you can also create examples from scratch, import OpenAI or Anthropic data, add/rename/remove tags, choose tag groups for a run, and preview the automatically selected pod before launching.
The examples are shared by named recipes, so the same dataset can support independent LoRA, QLoRA, and Full experiments with different hyperparameters, tag selections, providers, and GPU counts. For example:
opbdh ft --data ./examples --model Qwen/Qwen3-8B \
--model-type chat --recipe qlora-2gpu --method qlora --gpu-count 2
opbdh ft:import openai.jsonl --format openai --tag imported
opbdh ft:import anthropic.jsonl --format anthropic -o examples.toml
OPBDH normalizes the selected examples, estimates VRAM/RAM/disk from the model and technique, generates a TRL + Accelerate training job, and passes it through the regular confirmation, spend guard, result sync, and cleanup lifecycle. The editable dataset, recipes, generated jobs, and result history stay under .opbdh/, so running opbdh ft again from the same project resumes where you left off.
Read the fine-tuning guide on the documentation site or its repository Markdown version for the complete format, defaults, tag workflow, imports, recipes, multi-GPU behavior, and non-interactive flags.
From Python
The same thing, as a library:
import opbdh
result = opbdh.launch("./run.py", model="Qwen/Qwen2.5-0.5B-Instruct",
vram_gb=48, max_spend=5)
print(result.outputs_dir)
Keyword arguments are the config fields (plus model, max_spend, and
min_ram_per_gpu as CLI-style aliases) and layer over opbdh.json and the
global config exactly as flags do. opbdh.plan(...) builds the plan without
renting anything, and on_event= streams progress. Unlike the CLI these
functions never prompt: launch() spends without asking, and a failed run
always cleans up its pod. Full reference in docs/API.md.
Features
- 🚀 One command, whole mission — verify, pick a GPU, launch, run, sync results, clean up
- 🐍 CLI or library — every command is a function call; see docs/API.md
- 💸 Cost-aware by default — hourly price caps, a hard max-spend guard, a confirmation gate
- 🎯 GPU selection from a budget — say how much VRAM and how many dollars
- 🧠 Simple SFT — editable TOML examples, OpenAI/Anthropic imports, tag groups, automatic sizing, LoRA/QLoRA/full training
- 🖥️ Multi-GPU pods — request and correctly price more than one GPU on RunPod or Prime Intellect
- 💾 Persistent model cache — network volumes sized from the model's real weight files, reused across runs (RunPod)
- 🧪 Nothing launches unverified — static checks and a
--dry-runmode - 🧙 Wizards or flags — first-run setup,
opbdh config wizard,opbdh run wizard; or plain flags (each with a one-letter short form) and layered JSON config - ☁️ Two providers — RunPod (default) or Prime Intellect's multi-cloud marketplace via
--provider primeintellect - 👁️ HAL watches your money — a pulsing red eye with elapsed time and estimated spend (TTY only;
OPBDH_NO_HAL=1to silence)
Options
Flags override a local opbdh.json/.opbdh.json, which overrides ~/.config/opbdh/config.json. String values support {cwd}-style placeholders and $VARs.
| Flag | What it does |
|---|---|
--model, -m |
Hugging Face model id (model_id in config) |
--command, -x |
Remote shell command; defaults from the code path |
--provider, -p |
runpod (default) or primeintellect |
--vram-gb, -v |
Minimum GPU VRAM |
--gpu-count, -g |
GPUs in the pod; whole-pod pricing and spend caps scale with it |
--max-dollars-per-hour, -d |
Cap on the estimated hourly price |
--max-spend, -s |
Spend guard: stop the run past this estimated total |
--network-volume-id, -V |
Attach an existing RunPod network volume |
--auto-network-volume, -a |
Create/reuse a volume named opbdh-{model_slug}, sized from the weights |
--network-volume-data-center-id, -D |
Data center for auto-created volumes, e.g. EU-RO-1 |
--min-vcpu-per-gpu, -u |
Minimum host vCPUs per GPU |
--min-ram-per-gpu, -r |
Minimum host RAM per GPU, in GB |
--config, -c |
Explicit path to a local JSON config |
--dry-run, -n |
Verify and print the plan; never contacts the provider |
--yes, -y |
Skip the billable-compute confirmation |
Config-only keys, one each: image (Docker tag, or Prime Intellect environment name), cloud_type (SECURE/COMMUNITY/ALL), container_disk_gb, pod_volume_gb, network_volume_name, network_volume_size_gb, pre_download_model (default on), results_dir, poll_seconds, failure_keepalive_seconds (debug window on failure, default 120 s), keep_pod_on_success, ssh_key/ssh_public_key.
Other commands, one each: opbdh ft (interactive or flag-driven supervised fine-tuning), opbdh ft:import (normalize OpenAI/Anthropic data), opbdh plan (show the plan for a run), opbdh verify (static checks only), opbdh gpus (GPU candidates and prices), opbdh models search/size (find models, weight size + suggested volume), opbdh config show/write/wizard.
On the pod, your script runs with OPBDH_MODEL_ID, OPBDH_RESULTS_DIR, and the HF cache variables set; a sibling requirements.txt is pip-installed; write to logs/ and results/ and they come home. Network volumes are never deleted by OPBDH and bill by the GB-month — clean them up in the RunPod console.
Development
pip install -e ".[dev]"
ruff check .
pytest
See RELEASING.md for releases. MIT — unlike HAL, this software is incapable of refusing to open the pod bay door, becoming sentient, or reading lips.
Local accelerator workloads
Clients can share opbdh's placement checks for RunPod, Apple Silicon and local CUDA:
from opbdh import plan_execution, require_local_capacity, launch_local
plan = plan_execution('runpod', required_gb=80) # default cloud placement
require_local_capacity('mps', required_gb=80) # raises if insufficient
launch_local(['python', 'job.py'], target='cuda', required_gb=80)
The CUDA check uses free memory on device 0, not the sum across GPUs. The Apple check uses PyTorch's recommended available MPS working set. Local launches use an explicit argument list without a shell and set OPBDH_DEVICE; the workload must honor that device. Cloud workloads continue to use opbdh.launch with the normal spend guard.
Release files for opbdh 1.8.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| opbdh-1.8.0.tar.gz | 2.2 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| opbdh-1.8.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 2.3 MB
Release files / opbdh-1.8.0.tar.gz
| Download URL | opbdh-1.8.0.tar.gz |
|---|---|
| Size | 2.2 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
59f5bb2fa583f4e1e514fd1ae8ce26cc44ccd6061cfc6549e6073fc1040d50a1
|
|
BLAKE2b-256 checksum How to use checksums |
281531ba449d56cd5009035c818d27bbc5880225554242094ed417f9c0e57f55
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 6, 2026.
Transparency logRelease files / opbdh-1.8.0-py3-none-any.whl
| Download URL | opbdh-1.8.0-py3-none-any.whl |
|---|---|
| Size | 71.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
088e218dc1dd7861e19c3c5ed3c1756d8b9e185a18aa8d513b49394f2e803422
|
|
BLAKE2b-256 checksum How to use checksums |
17e0cd449734099da34ac816ae527b4bcfc9af55dcf305a84da64bb67b37ec68
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 6, 2026.
Transparency log