Skip to main content

cuda-doctor

CI

Read-only local diagnostics for CUDA / PyTorch development environments.

Detect → Analyze → Explain → Recommend. Never modifies your system.

cuda-doctor inspects your machine the way a colleague would before helping you debug a CUDA problem: it looks at the GPUs, the NVIDIA driver, the CUDA toolkit(s) on disk and on PATH, the host compiler, and your PyTorch install — then tells you what looks wrong, why, and what to do about it.

It is inspired by flutter doctor: findings are explained in plain language, and a completed run always exits 0 (only internal tool failures exit non-zero).

Quick start

Install from PyPI (Python 3.10+):

# Linux / macOS
python -m pip install cuda-doctor

# Windows
py -m pip install cuda-doctor

Then run it:

cuda-doctor                  # full diagnosis, terminal report

Install it where your PyTorch lives. cuda-doctor imports PyTorch in-process from the Python environment it runs in. To diagnose the PyTorch in one of your conda/venv environments, install and run it inside that environment:

conda activate my-torch-env
python -m pip install cuda-doctor
cuda-doctor

For the same reason, prefer pip install over pipx: a pipx install runs in its own isolated environment and cannot see the PyTorch installed in your conda/venv environment — it would report "torch not installed" even when it is.

CUDA Doctor v0.1.2

System
  OS              Linux Ubuntu 22.04.5 LTS (kernel 5.15.0-91-generic)
  ...
NVIDIA Driver
  Version         580.126.09
  Reported CUDA   13.0
...
Summary
  Status: HEALTHY  (2 info)

Commands

Command What it does
cuda-doctor Same as cuda-doctor diagnose
cuda-doctor diagnose Full diagnosis, terminal report
cuda-doctor diagnose --verbose Also show info-level notes and probe errors
cuda-doctor diagnose --format json Machine-readable report (schema v1)
cuda-doctor diagnose --format markdown --output report.md Write a shareable report
cuda-doctor info Tool metadata: version, checks, issue codes
cuda-doctor --version Version

python -m cuda_doctor … works identically to the cuda-doctor entry point.

Exit codes: 0 the diagnosis ran (regardless of findings) · 2 internal fatal error. Findings never change the exit code — this is a diagnostician, not a CI gate.

What it checks (v0.1)

Area Codes Highlights
GPU GPU001–GPU003 nvidia-smi missing, no GPUs, smi execution failure
Driver DRV001–DRV002 unknown driver version; driver below the documented minimum for the toolkit's CUDA generation
CUDA CUDA001–CUDA006 no nvcc; CUDA_HOME unset/invalid; multiple toolkits; multiple CUDA bins in PATH; nvcc ≠ CUDA_HOME
PyTorch TORCH001–TORCH006 not installed; import failures (with known-signature advice); is_available() False; CPU-only wheel; runtime vs toolkit; runtime from a newer CUDA generation than the driver
Compiler CMP001–CMP002 no host compiler; gcc/Visual Studio outside the toolkit's supported range
Environment ENV001–ENV004 stale CUDA paths; duplicates; conflicting LD_LIBRARY_PATH; older CUDA shadowing newer in PATH

Every finding carries a stable issue code, evidence, and concrete recommendations. Example reports: examples/sample_report.md · examples/sample_report.json.

The nuance that matters: three different "CUDA versions"

A machine reports three unrelated CUDA versions, and most false alarms in CUDA debugging tools come from comparing the wrong pair:

  1. The local toolkit (nvcc --version) — what you compile with.
  2. PyTorch's bundled runtime (torch.version.cuda) — the CUDA libraries shipped inside the wheel; official wheels bundle their own runtime, so it does not have to match the local toolkit (TORCH004 stays INFO).
  3. The driver's CUDA UMD version (the CUDA Version: line in nvidia-smi) — the toolkit generation the driver was validated with. It is not a hard ceiling.

That third point is the one most tools get wrong. Since CUDA 11, NVIDIA supports CUDA minor-version compatibility: within a CUDA major family (e.g. any CUDA 12.x), applications built with a newer minor release run on older drivers of the same generation, as long as the driver meets the documented family minimum (Linux 525.60.13 / Windows 528.33 for CUDA 12.x; for CUDA 13.x the documented rule is the R580 driver branch, i.e. >= 580). So "toolkit 12.6, nvidia-smi says 12.2" is at most an informational note (DRV002 INFO), not an error — with the caveats that newer driver-dependent features and newer PTX may still need a driver update.

What is a genuine error is a generation gap — a CUDA 13 toolkit or PyTorch runtime on a CUDA 12-generation driver (DRV002/TORCH006) — and even then only when CUDA is not observed working: if torch.cuda.is_available() is True, observed runtime success always overrides the static version comparison.

Platforms

  • Linux — first-class; developed and continuously smoke-tested on Ubuntu 22.04 with NVIDIA H20 GPUs and 4 parallel CUDA toolkit installations.
  • Windows — supported (separate code paths for CUDA_PATH* variables, Visual Studio detection via vswhere, CREATE_NO_WINDOW on subprocesses). Automated tests cover the Windows code paths with simulated environments; real-hardware validation is ongoing.
  • macOS — runs and reports "no NVIDIA driver" as informational (Apple Silicon is out of scope for v0.1).

Privacy

Reports are designed to be shareable:

  • your home directory is rewritten to ~, leftover usernames to <user>;
  • hostname and username are never collected in the first place;
  • PATH/LD_LIBRARY_PATH are shown only as the CUDA-relevant subset, never dumped in full;
  • nvidia-smi output excerpts are truncated.

Still, a JSON report describes your hardware and software stack — read it before posting publicly.

Architecture

CLI (Typer) ─ CollectionRunner ─▶ Collectors (facts only)
                                     │  subprocess (no shell, timeouts),
                                     │  filesystem, env, torch import
                                     ▼
                               EnvironmentSnapshot
                               (dataclasses; absence = None)
                                     │
                          DiagnosisEngine (24 checks)
                             + bundled compatibility JSON
                                     ▼
                              Issues + Summary
                                     ▼
                    Reporters: terminal / JSON / Markdown

Key invariants (see AGENTS.md):

  • Collectors never judge; checks never collect. Rules are pure functions of the snapshot (+ compatibility data), which makes them unit-testable without hardware.
  • Safe failure over crashes. Every external probe is guarded; a failing collector surfaces as data, not an exception.
  • Compatibility knowledge is data, versioned JSON inside the package (src/cuda_doctor/data/), updatable without touching logic.
  • Read-only, always. v0.1 executes no command that modifies the system.

Limitations (v0.1)

  • Single-user, single-machine scope; no container/WSL-specific detection.
  • Compatibility decisions are deliberately conservative: driver minimums follow NVIDIA's documented CUDA minor-version-compatibility baselines per major family, and anything the bundled knowledge does not cover — a future CUDA major (e.g. CUDA 14 before the data ships), an unlisted toolkit minor for compiler rules, or an unreported driver version — is reported as unknown, never guessed from older versions.
  • Compiler compatibility tables are coarse and worded as potential issues; the definitive source is always the CUDA Installation Guide for your toolkit version.
  • No conda-environment awareness beyond what PATH/env vars imply.
  • nvidia-smi is the only GPU source; NVML/lspci fallbacks are future work.
  • Chinese localization is planned (the maintainers are bilingual); v0.1 output is English.

Development

bash scripts/dev_install.sh      # Linux/macOS   (pip install -e ".[dev]")
# or: pwsh scripts/dev_install.ps1   (Windows)

pytest                           # 310 tests: unit + integration (hermetic)
ruff check src tests             # lint
mypy                             # types (strict-ish: disallow_untyped_defs)

CI runs the same commands on Ubuntu and Windows across Python 3.10–3.13 (.github/workflows/ci.yml) — no GPU, driver, CUDA, or PyTorch required.

The test suite needs no GPU and no CUDA: collectors run against an injected fake command runner, and real parser fixtures captured from actual H20 / RTX 4090 machines live under tests/fixtures/.

Repository layout

src/cuda_doctor/
├── cli.py                  # Typer app: diagnose / info / --version
├── core/                   # models, enums, CollectionRunner, exceptions
├── collectors/             # one module per fact source
├── checks/                 # 24 diagnostic rules (pure)
├── diagnosis/              # engine, Issue/Summary, recommendations
├── compatibility/          # driver/compiler/torch knowledge + JSON loading
├── reporters/              # terminal (Rich), JSON, Markdown
├── data/                   # bundled compatibility tables
└── utils/                  # commands, parsing, paths, redaction, versions

Roadmap (v0.2+)

  • NVML and lspci fallbacks when nvidia-smi is absent or broken.
  • Conda-aware environment inspection.
  • --json-schema self-description and a stable JSON schema contract test.
  • WSL detection and WSL-specific driver mismatch explanations.
  • Chinese output (--lang zh).
  • cuda-doctor fix --dry-run advisory mode (still no modifications in v0.2).

License

MIT — see LICENSE.

Metadata

Release files for cuda-doctor 0.1.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for cuda-doctor 0.1.2
File Size Uploaded
cuda_doctor-0.1.2.tar.gz 53.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for cuda-doctor 0.1.2
File Interpreter ABI Platform
cuda_doctor-0.1.2-py3-none-any.whl Python 3 none any Details

Total release size: 121.2 kB

Release files / cuda_doctor-0.1.2.tar.gz

Download URL cuda_doctor-0.1.2.tar.gz
Size 53.4 kB
Tags Source
SHA-256 checksum
How to use checksums
543d0ac23bd2f62aa1f039e0f8d23cff510a8402f455f926705e0aa861b7a164
BLAKE2b-256 checksum
How to use checksums
8366c04a8deb5287d55c9f68ba70f3d09b55414ac5c148f003685a33ee265982
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 9, 2026.

Transparency log

Release files / cuda_doctor-0.1.2-py3-none-any.whl

Download URL cuda_doctor-0.1.2-py3-none-any.whl
Size 67.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
8fcf385c458d9a99a00e4f1db8c254d47e9857f249d88e7448c474af9b889741
BLAKE2b-256 checksum
How to use checksums
bba1a6702e764ce26748ae7662477bf73b88d3d5c45dcbde4ffd29b47a128452
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 9, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.2 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page