ComputeFence
Pre-flight safety gate for ML training jobs on rented GPU compute. Built specifically for RunPod, Vast.ai, Lambda Labs, and similar bare metal providers.
Install
pip install computefence
Usage
computefence doctor
With a dataset:
computefence doctor --dataset train.csv --input-column text --label-column label
Example output
ComputeFence v0.1.2 — Pre-flight diagnostic ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 2 WARNINGS · 0 BLOCKERS · 3 PASSED
Environment ✓ Python 3.11.4 ✓ PyTorch 2.1.0 detected ✓ CUDA available — NVIDIA A40
Storage ⚠ HF_HOME is not set. HuggingFace will use default local cache. Fix: Set HF_HOME to a persistent volume path before training ⚠ No mounted persistent volumes detected at /workspace, /runpod-volume, or /vast Fix: Attach a network volume before training to persist checkpoints and cache
Dataset ✓ No dataset path provided — skipping dataset checks
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 2 warning(s) found. Review before launching.
What it checks
- CUDA and GPU visibility — confirms PyTorch can see the GPU and training will not silently fall back to CPU
- HuggingFace cache path — confirms model files go to persistent storage not ephemeral disk
- Accelerate GPU count — confirms your distributed training config matches the GPUs actually on the instance
- Dataset integrity — optional scan for duplicates, missing values, and conflicting labels
What it does not check
- Training script correctness
- Model architecture compatibility
- Learning rate or hyperparameter safety
- Runtime monitoring during the job
- Slow dataloader or data pipeline throughput
Why this exists
I burned approximately £1,000 on GPU training runs that failed silently. CUDA fell back to CPU with no error. Class weights caused loss collapse to 0.693 immediately. My dataset had 28,432 duplicate rows and 312 conflicting labels I only found after the run.
Nothing existed that caught these before the job started. So I built it.
Real operator results
David at Neuralic ran ComputeFence on a RunPod A100. It caught HF_HOME writing to /root/.cache and an Accelerate GPU count mismatch. He fixed both before launch.
GitHub
github.com/Francisco-Booth/ComputeFence
PyPI
pypi.org/project/computefence
Metadata
Release files for computefence 0.2.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| computefence-0.2.1.tar.gz | 10.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| computefence-0.2.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 21.9 kB
Release files / computefence-0.2.1.tar.gz
| Download URL | computefence-0.2.1.tar.gz |
|---|---|
| Size | 10.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
f17a436e6ffaf2eec6836d9e013e77c63105c2c4ea66e5c27d747ff7c8483097
|
|
BLAKE2b-256 checksum How to use checksums |
257f22a3e0621721a3dc798cdee4e63c031a07feb2fc9e207fdcdcf7808989c4
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.3
|
Release files / computefence-0.2.1-py3-none-any.whl
| Download URL | computefence-0.2.1-py3-none-any.whl |
|---|---|
| Size | 11.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
495a402e3ca76d4133b85f1f81278fac9fafe3c9d8363e25e6a063fd9648c1e3
|
|
BLAKE2b-256 checksum How to use checksums |
06f61283d242b2dbd583fbee3ab837cdb8d2ccd557ade4a9dd23b245544405c4
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.3
|