cuda-doctor
Read-only local diagnostics for CUDA / PyTorch development environments.
Detect → Analyze → Explain → Recommend. Never modifies your system.
cuda-doctor inspects your machine the way a colleague would before helping you
debug a CUDA problem: it looks at the GPUs, the NVIDIA driver, the CUDA toolkit(s)
on disk and on PATH, the host compiler, and your PyTorch install — then tells
you what looks wrong, why, and what to do about it.
It is inspired by flutter doctor: findings are explained in plain language,
and a completed run always exits 0 (only internal tool failures exit non-zero).
Quick start
Install from PyPI (Python 3.10+):
# Linux / macOS
python -m pip install cuda-doctor
# Windows
py -m pip install cuda-doctor
Then run it:
cuda-doctor # full diagnosis, terminal report
Install it where your PyTorch lives. cuda-doctor imports PyTorch in-process from the Python environment it runs in. To diagnose the PyTorch in one of your conda/venv environments, install and run it inside that environment:
conda activate my-torch-env python -m pip install cuda-doctor cuda-doctorFor the same reason, prefer
pip installover pipx: a pipx install runs in its own isolated environment and cannot see the PyTorch installed in your conda/venv environment — it would report "torch not installed" even when it is.
CUDA Doctor v0.1.2
System
OS Linux Ubuntu 22.04.5 LTS (kernel 5.15.0-91-generic)
...
NVIDIA Driver
Version 580.126.09
Reported CUDA 13.0
...
Summary
Status: HEALTHY (2 info)
Commands
| Command | What it does |
|---|---|
cuda-doctor |
Same as cuda-doctor diagnose |
cuda-doctor diagnose |
Full diagnosis, terminal report |
cuda-doctor diagnose --verbose |
Also show info-level notes and probe errors |
cuda-doctor diagnose --format json |
Machine-readable report (schema v1) |
cuda-doctor diagnose --format markdown --output report.md |
Write a shareable report |
cuda-doctor info |
Tool metadata: version, checks, issue codes |
cuda-doctor --version |
Version |
python -m cuda_doctor … works identically to the cuda-doctor entry point.
Exit codes: 0 the diagnosis ran (regardless of findings) · 2 internal
fatal error. Findings never change the exit code — this is a diagnostician, not
a CI gate.
What it checks (v0.1)
| Area | Codes | Highlights |
|---|---|---|
| GPU | GPU001–GPU003 |
nvidia-smi missing, no GPUs, smi execution failure |
| Driver | DRV001–DRV002 |
unknown driver version; driver below the documented minimum for the toolkit's CUDA generation |
| CUDA | CUDA001–CUDA006 |
no nvcc; CUDA_HOME unset/invalid; multiple toolkits; multiple CUDA bins in PATH; nvcc ≠ CUDA_HOME |
| PyTorch | TORCH001–TORCH006 |
not installed; import failures (with known-signature advice); is_available() False; CPU-only wheel; runtime vs toolkit; runtime from a newer CUDA generation than the driver |
| Compiler | CMP001–CMP002 |
no host compiler; gcc/Visual Studio outside the toolkit's supported range |
| Environment | ENV001–ENV004 |
stale CUDA paths; duplicates; conflicting LD_LIBRARY_PATH; older CUDA shadowing newer in PATH |
Every finding carries a stable issue code, evidence, and concrete
recommendations. Example reports: examples/sample_report.md
· examples/sample_report.json.
The nuance that matters: three different "CUDA versions"
A machine reports three unrelated CUDA versions, and most false alarms in CUDA debugging tools come from comparing the wrong pair:
- The local toolkit (
nvcc --version) — what you compile with. - PyTorch's bundled runtime (
torch.version.cuda) — the CUDA libraries shipped inside the wheel; official wheels bundle their own runtime, so it does not have to match the local toolkit (TORCH004stays INFO). - The driver's CUDA UMD version (the
CUDA Version:line innvidia-smi) — the toolkit generation the driver was validated with. It is not a hard ceiling.
That third point is the one most tools get wrong. Since CUDA 11, NVIDIA
supports CUDA minor-version compatibility: within a CUDA major family
(e.g. any CUDA 12.x), applications built with a newer minor release run on
older drivers of the same generation, as long as the driver meets the
documented family minimum (Linux 525.60.13 / Windows 528.33 for CUDA 12.x;
for CUDA 13.x the documented rule is the R580 driver branch, i.e. >= 580).
So "toolkit 12.6, nvidia-smi says 12.2" is at most an informational note
(DRV002 INFO), not an error — with the caveats that newer
driver-dependent features and newer PTX may still need a driver update.
What is a genuine error is a generation gap — a CUDA 13 toolkit or
PyTorch runtime on a CUDA 12-generation driver (DRV002/TORCH006) — and
even then only when CUDA is not observed working: if
torch.cuda.is_available() is True, observed runtime success always
overrides the static version comparison.
Platforms
- Linux — first-class; developed and continuously smoke-tested on Ubuntu 22.04 with NVIDIA H20 GPUs and 4 parallel CUDA toolkit installations.
- Windows — supported (separate code paths for
CUDA_PATH*variables, Visual Studio detection via vswhere,CREATE_NO_WINDOWon subprocesses). Automated tests cover the Windows code paths with simulated environments; real-hardware validation is ongoing. - macOS — runs and reports "no NVIDIA driver" as informational (Apple Silicon is out of scope for v0.1).
Privacy
Reports are designed to be shareable:
- your home directory is rewritten to
~, leftover usernames to<user>; - hostname and username are never collected in the first place;
PATH/LD_LIBRARY_PATHare shown only as the CUDA-relevant subset, never dumped in full;- nvidia-smi output excerpts are truncated.
Still, a JSON report describes your hardware and software stack — read it before posting publicly.
Architecture
CLI (Typer) ─ CollectionRunner ─▶ Collectors (facts only)
│ subprocess (no shell, timeouts),
│ filesystem, env, torch import
▼
EnvironmentSnapshot
(dataclasses; absence = None)
│
DiagnosisEngine (24 checks)
+ bundled compatibility JSON
▼
Issues + Summary
▼
Reporters: terminal / JSON / Markdown
Key invariants (see AGENTS.md):
- Collectors never judge; checks never collect. Rules are pure functions of the snapshot (+ compatibility data), which makes them unit-testable without hardware.
- Safe failure over crashes. Every external probe is guarded; a failing collector surfaces as data, not an exception.
- Compatibility knowledge is data, versioned JSON inside the package
(
src/cuda_doctor/data/), updatable without touching logic. - Read-only, always. v0.1 executes no command that modifies the system.
Limitations (v0.1)
- Single-user, single-machine scope; no container/WSL-specific detection.
- Compatibility decisions are deliberately conservative: driver minimums follow NVIDIA's documented CUDA minor-version-compatibility baselines per major family, and anything the bundled knowledge does not cover — a future CUDA major (e.g. CUDA 14 before the data ships), an unlisted toolkit minor for compiler rules, or an unreported driver version — is reported as unknown, never guessed from older versions.
- Compiler compatibility tables are coarse and worded as potential issues; the definitive source is always the CUDA Installation Guide for your toolkit version.
- No conda-environment awareness beyond what
PATH/env vars imply. - nvidia-smi is the only GPU source;
NVML/lspcifallbacks are future work. - Chinese localization is planned (the maintainers are bilingual); v0.1 output is English.
Development
bash scripts/dev_install.sh # Linux/macOS (pip install -e ".[dev]")
# or: pwsh scripts/dev_install.ps1 (Windows)
pytest # 310 tests: unit + integration (hermetic)
ruff check src tests # lint
mypy # types (strict-ish: disallow_untyped_defs)
CI runs the same commands on Ubuntu and Windows across Python 3.10–3.13
(.github/workflows/ci.yml) — no GPU, driver, CUDA, or PyTorch required.
The test suite needs no GPU and no CUDA: collectors run against an injected
fake command runner, and real parser fixtures captured from actual H20 /
RTX 4090 machines live under tests/fixtures/.
Repository layout
src/cuda_doctor/
├── cli.py # Typer app: diagnose / info / --version
├── core/ # models, enums, CollectionRunner, exceptions
├── collectors/ # one module per fact source
├── checks/ # 24 diagnostic rules (pure)
├── diagnosis/ # engine, Issue/Summary, recommendations
├── compatibility/ # driver/compiler/torch knowledge + JSON loading
├── reporters/ # terminal (Rich), JSON, Markdown
├── data/ # bundled compatibility tables
└── utils/ # commands, parsing, paths, redaction, versions
Roadmap (v0.2+)
- NVML and
lspcifallbacks when nvidia-smi is absent or broken. - Conda-aware environment inspection.
--json-schemaself-description and a stable JSON schema contract test.- WSL detection and WSL-specific driver mismatch explanations.
- Chinese output (
--lang zh). cuda-doctor fix --dry-runadvisory mode (still no modifications in v0.2).
License
MIT — see LICENSE.
Metadata
Release files for cuda-doctor 0.1.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| cuda_doctor-0.1.2.tar.gz | 53.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| cuda_doctor-0.1.2-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 121.2 kB
Release files / cuda_doctor-0.1.2.tar.gz
| Download URL | cuda_doctor-0.1.2.tar.gz |
|---|---|
| Size | 53.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
543d0ac23bd2f62aa1f039e0f8d23cff510a8402f455f926705e0aa861b7a164
|
|
BLAKE2b-256 checksum How to use checksums |
8366c04a8deb5287d55c9f68ba70f3d09b55414ac5c148f003685a33ee265982
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 9, 2026.
Transparency logRelease files / cuda_doctor-0.1.2-py3-none-any.whl
| Download URL | cuda_doctor-0.1.2-py3-none-any.whl |
|---|---|
| Size | 67.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
8fcf385c458d9a99a00e4f1db8c254d47e9857f249d88e7448c474af9b889741
|
|
BLAKE2b-256 checksum How to use checksums |
bba1a6702e764ce26748ae7662477bf73b88d3d5c45dcbde4ffd29b47a128452
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 9, 2026.
Transparency log