GPU thermal-power forensics agent. Computes R_theta = ΔT/P in real time.
Project description
Theta
GPU thermal-power forensics agent. Computes R_θ = ΔT / P in real time from your existing DCGM telemetry. That ratio is the only signal that separates a busy-hot GPU from a failing-hot one — and no incumbent computes it.
theta_gpu_rtheta_cwatt{gpu_index="3"} 2.104 # zombie recovery — CUDA context stuck
theta_gpu_rtheta_cwatt{gpu_index="3"} 0.724 # under load — healthy
theta_gpu_rtheta_cwatt{gpu_index="3"} 1.281 # clean idle — normal
The problem
A GPU at 82°C could be:
- Busy and healthy — running a job at thermal equilibrium
- Cooling path failing — ambient temperature up, heatsink degrading
- CUDA zombie — process exited but context retained, drawing 31W at 0% utilization
nvidia-smi, DCGM, and Mission Control all expose T and P as separate fields. None of them divide the two. Theta does.
Quick start
pip (single node, free forever)
pip install runtheta
theta setup # interactive wizard — 90 seconds to first R_θ reading
theta monitor # start monitoring
Docker
docker run --gpus all -p 9101:9101 theta/agent:latest
Docker Compose (agent + Prometheus + Grafana)
git clone https://github.com/Asomisetty27/theta
cd theta
docker compose --profile metrics up
Open http://localhost:3000 — Grafana dashboard pre-provisioned, no setup required.
Login: admin / theta
How it works
GPU (pynvml)
→ T_junction, P_GPU, util, P-state every 5s
→ R_θ = (T_junction − T_ref) / P_GPU
→ 15s steady-state window (σ < 0.03 C/W)
→ Decision Tree classifier → {under_load, clean_idle, zombie_recovery, child_exit_recovery}
→ Rolling baseline + k·σ drift detector
→ Alert (stdout / Slack webhook / JSONL / Prometheus)
Virtual ambient — T_ref is derived from the GPU's own stable idle windows. No thermocouple, no rack modification, no extra hardware.
Steady-state filter — classification only runs on stable windows. This takes Naive Bayes accuracy from 84% → 99.8% and eliminates transient false positives.
Classifier — Decision Tree trained on 4,570 rows of Stage 1 Tesla T4 data. 100% 5-fold CV accuracy on steady-state samples. Rules are human-readable and publishable:
IF R_θ ≤ 0.87 → under_load (n=963, conf=1.00)
IF R_θ > 0.87, P0 → zombie_recovery (n=584, conf=1.00) ← CUDA zombie
IF R_θ > 1.50, P8 → child_exit_recovery (n=696, conf=1.00)
ELSE → clean_idle / early recovery
CLI reference
theta setup Interactive wizard (run this first)
theta monitor Run agent — blocks until Ctrl+C
theta monitor --interval 2 Sample every 2s
theta monitor --gpus 0,1,3 Monitor specific GPUs
theta monitor --webhook <url> Send alerts to Slack / generic webhook
theta monitor --log alerts.jsonl Append alerts to JSONL file
theta monitor --port 9101 Prometheus metrics port (0 = disabled)
theta monitor --nb Use Naive Bayes instead of Decision Tree
theta baseline --gpu 0 Lock virtual ambient T_ref from idle window
theta baseline --gpu 0 --manual 24 Set T_ref manually (°C)
theta classify Snapshot classify all GPUs right now
theta serve --port 9101 Metrics export only (no stdout alerts)
theta train /path/data.csv Retrain bundled models from new data
Prometheus metrics
| Metric | Type | Description |
|---|---|---|
theta_gpu_rtheta_cwatt |
gauge | R_θ (C/W) — the core signal |
theta_gpu_state_info |
gauge | Current classified state (label: state) |
theta_gpu_drift_sigma |
gauge | Deviation from baseline in σ units |
theta_gpu_temperature_celsius |
gauge | Junction temperature |
theta_gpu_power_watts |
gauge | GPU power consumption |
theta_gpu_utilization_ratio |
gauge | 0–1 utilization |
theta_gpu_perf_state |
gauge | P-state (0=max, 8=idle) |
theta_gpu_baseline_tref_celsius |
gauge | Virtual ambient T_ref |
theta_gpu_window_rtheta_std |
gauge | Steady-state window σ |
theta_gpu_alerts_total |
counter | Alerts (labels: severity, state) |
All metrics include a gpu_index label.
Alert payload (webhook / JSONL)
Every alert includes full forensic context:
{
"source": "theta",
"severity": "critical",
"gpu_index": 3,
"state": "zombie_recovery",
"prev_state": "under_load",
"rtheta": 1.541,
"rtheta_baseline": 0.724,
"drift_sigma": 4.2,
"confidence": 1.0,
"message": "[CRITICAL] GPU 3 — CUDA zombie detected. R_θ=1.541 at 0% utilisation. Action: release CUDA context.",
"context": {
"severity": "critical",
"duration_prev": 3842.1,
"history": [
{ "ts": 1748995200.1, "state": "under_load", "r": 0.721, "conf": 0.99 }
]
}
}
Why not DCGM / Mission Control / Phaidra?
| Capability | DCGM | Mission Control | Phaidra | Theta |
|---|---|---|---|---|
| Computes R_θ | ✗ | ✗ | ✗ | ✓ |
| Separates busy-hot vs failing-hot | ✗ | ✗ | ✗ | ✓ |
| CUDA zombie detection | ✗ | ✗ | ✗ | ✓ |
| Drift detection (baseline + k·σ) | ✗ | ✗ | ◐ | ✓ |
| Virtual ambient (no hardware) | ✗ | ✗ | ✗ | ✓ |
| Serves neocloud / mixed fleets | ✓ | ✗ | ✗ | ✓ |
| Open-source agent | ✓ | ✗ | ✗ | ✓ |
Mission Control ships only on Blackwell DGX/GB200. Theta runs on any NVIDIA GPU reachable by pynvml.
Requirements
- Python 3.10+
- NVIDIA GPU with driver ≥ 450 (for pynvml)
- No DCGM required — pynvml only
For Docker: nvidia-container-toolkit
Retrain on your own data
theta train /path/to/measurements.csv
CSV schema: phase, trial_second, rtheta_cwatt, power_w, util_pct, perf_state, ...
Research basis
- F1 — R_θ separates idle (1.28 C/W) from load (0.72 C/W) with 77.9% margin, Tesla T4
- F2 — Ambient sensitivity: 7.1%/°C at idle vs 2.0%/°C at load (3.5× difference)
- F6 — CUDA zombie: same-process exit leaves GPU at P0 (~31W), invisible to utilization
Stage 1: 4,570 rows · Tesla T4 · E001–E004 · 9 child-exit trials
Stage 2 (in progress): Cal Poly DGX B200 AI Factory · E005–E008
License
MIT — free forever for single-node use.
Built at Cal Poly SLO · asomisetty27@gmail.com
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file runtheta-0.1.10.tar.gz.
File metadata
- Download URL: runtheta-0.1.10.tar.gz
- Upload date:
- Size: 118.7 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.14.5
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
d9f73eb58c315c6d9b7f12fe30ec3c938d8497769974d05634a186def7938ac1
|
|
| MD5 |
89f1438e19e47c56c750a39c92fee733
|
|
| BLAKE2b-256 |
8b080f21908c95b15be7983e193b7b0f9342a47c100e9b8043c681dbb6cbc39d
|
File details
Details for the file runtheta-0.1.10-py3-none-any.whl.
File metadata
- Download URL: runtheta-0.1.10-py3-none-any.whl
- Upload date:
- Size: 86.1 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.14.5
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
b2879351a20ed7e95c46f7ac9194733f25f7c281e83da05535bf0619317b7fbc
|
|
| MD5 |
cb40c506b583ed15aea2fca43f433711
|
|
| BLAKE2b-256 |
49a85bf61d093cb74b4f43cb3020758316a6af64da20a50dce3a3dad40333e04
|