Skip to main content

ComputeFence

Pre-flight safety gate for ML training jobs on rented GPU compute. Built for RunPod, Vast.ai, Lambda Labs, and similar bare metal providers.

Install

pip install computefence

Or with UV:

uvx computefence doctor

Usage

computefence doctor

With a dataset:

computefence doctor --dataset train.csv --input-column text --label-column label

Example output

ComputeFence v0.2.0 — Pre-flight diagnostic ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 2 WARNINGS · 0 BLOCKERS · 3 PASSED

Environment ✓ Python 3.11.4 ✓ PyTorch 2.1.0 detected ✓ CUDA available — NVIDIA A40

Storage ⚠ HF_HOME is not set. HuggingFace will use default local cache. Fix: export HF_HOME=/workspace/.cache/huggingface ⚠ Root disk (/) — 14.3 GB free of 460.4 GB (below 20 GB) Fix: Free up disk space or attach a larger volume before launching

Dataset ✓ No dataset path provided — skipping dataset checks

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 2 warning(s) found. Review before launching.

What it checks

  • CUDA and GPU visibility — confirms PyTorch can see the GPU and training will not silently fall back to CPU
  • HuggingFace cache path — confirms model files go to persistent storage not ephemeral disk
  • Accelerate GPU count — confirms your distributed training config matches the GPUs actually on the instance
  • Disk space headroom — checks free space on workspace volumes and root disk. Warns below 20 GB, blocks below 5 GB
  • Dataset integrity — optional scan for duplicates, missing values, and conflicting labels

What it does not check

  • Training script correctness
  • Model architecture compatibility
  • Learning rate or hyperparameter safety
  • Runtime monitoring during the job
  • Slow dataloader or data pipeline throughput
  • Dataloader bottleneck causing low GPU utilisation

Why this exists

I burned approximately £1,000 on GPU training runs that failed silently. CUDA fell back to CPU with no error. 24 seconds per iteration instead of 0.4. Class weights caused loss collapse to 0.693 immediately. My dataset had 28,432 duplicate rows and 312 conflicting labels I only found after the run.

Nothing existed that caught these before the job started. So I built it.

Real operator results

David at Neuralic ran ComputeFence on a RunPod A100. It caught HF_HOME writing to /root/.cache and an Accelerate GPU count mismatch. He fixed both before launch.

Shahzeb Ali, a computer vision engineer running client training jobs on RunPod, confirmed the storage warning matches real pod behaviour and would not have caught the HF_HOME issue explicitly without the tool.

The problem it solves

A healthy GPU does not mean you are training the right job.

Docker makes environments reproducible. It does not check whether your HuggingFace cache is writing to ephemeral storage, whether your Accelerate config matches the GPUs actually on the instance, or whether your disk has enough headroom for checkpoints. ComputeFence addresses the job layer not the environment layer.

Five independent operators confirmed this independently. Docker does not solve job-specific configuration mistakes.

GitHub

github.com/Francisco-Booth/ComputeFence

PyPI

pypi.org/project/computefence

Metadata

Release files for computefence 0.2.4

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for computefence 0.2.4
File Size Uploaded
computefence-0.2.4.tar.gz 11.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for computefence 0.2.4
File Interpreter ABI Platform
computefence-0.2.4-py3-none-any.whl Python 3 none any Details

Total release size: 23.2 kB

Release files / computefence-0.2.4.tar.gz

Download URL computefence-0.2.4.tar.gz
Size 11.3 kB
Tags Source
SHA-256 checksum
How to use checksums
8253ce51c341edc6b8672bc444f8f52a897b15feedfec47e94788fd3b3747b68
BLAKE2b-256 checksum
How to use checksums
88a9918b3d789591ae4a014eca18179596149dd249c11b4ff7b452d8de78fcf1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.3

Release files / computefence-0.2.4-py3-none-any.whl

Download URL computefence-0.2.4-py3-none-any.whl
Size 11.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
f7f04a7756928d3b7886ec0f449997fab6b38ab13fb2ec476d33d6463dd36765
BLAKE2b-256 checksum
How to use checksums
c71372a2b3b64ff49fc9294a45025b6d4368ce3f3614e2a32eb1b00eb9799acc
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.3

Release history Release notifications | RSS feed

0.2.6

2 release files

0.2.5

2 release files

This release

0.2.4 This release

2 release files

0.2.3

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page