Skip to main content

comprisk

PyPI version CI docs DOI arXiv

A Python toolkit for competing-risks survival analysis: a scalable, scikit-learn-compatible competing-risks random survival forest plus the canonical regression / non-parametric methods — Fine-Gray, Aalen-Johansen CIF, cause-specific Cox — so applied researchers can drop the Python → R round-trip.

Status: alpha — API may change before v1.0. Renamed from crforest in 0.3.1 (pip install comprisk; from comprisk import CompetingRiskForest).

Highlights

  • Four canonical CR methods, native Python — Fine-Gray (+ penalized), cause-specific Cox, Aalen-Johansen CIF, Gray's test — each validated to floating-point tolerance against cmprsk / crrp / survival.
  • The only native-Python CR forest — composite & cause-specific CR log-rank splitting, AJ CIF, Nelson-Aalen CHF, Wolbers + Uno IPCW concordance, OOB Breiman VIMP, Ishwaran minimal depth, exact TreeSHAP.
  • CR-aware evaluation — score_cr (IPCW time-dependent AUC/Brier + bootstrap CIs) and calibration_cr, replacing the CR-mode riskRegression::Score() block; concordance_index_ci gives a closed-form (no-bootstrap) CI for the Uno C-index and a paired model-comparison concordance_index_delta_ci.
  • Fast — 10–22× vs randomForestSRC on real EHR, 16.6–544× vs scikit-survival (n = 5k → 50k), n = 10⁶ in 63 s — at matched C ≈ 0.85. Benchmarks →
  • Reproducible — equivalence="rfsrc" reproduces rfSRC's per-tree mtry/nsplit RNG stream bit-for-bit. Methodology →

Install

pip install comprisk          # or:  uv add comprisk
pip install "comprisk[gpu]"   # CUDA 12 preview (faster only at low p today)

Python ≥ 3.10. Core deps: numpy, scipy, pandas, joblib, numba, scikit-learn.

Quickstart

from comprisk import CompetingRiskForest

# event: 0 = censored, k≥1 = cause-k event. Defaults: 100 trees, logrankCR, n_jobs=-1.
forest = CompetingRiskForest(n_estimators=200, random_state=42).fit(X, time, event)

cif  = forest.predict_cif(X[:5])          # (5, n_causes, n_times) — Aalen-Johansen
print(forest.oob_score(cause=1))          # honest out-of-bag C-index (no holdout split)
shap, base = forest.shap_values(X[:10])   # exact TreeSHAP (n, p, n_times, n_causes)

Prediction shapes, scoring, cross-validation, VIMP, minimal depth, GPU, and rfSRC migration — all with runnable code — are in the quickstart. Every estimator here is a real sklearn estimator — get_params / set_params / clone, and cross_val_score / Pipeline / GridSearchCV without a wrapper. Fit on a DataFrame and feature_names_in_ carries your column names through to VIMP output.

Like every survival library, these estimators cannot pass sklearn's full check_estimator suite: it supplies a plain numeric y, while competing-risks data needs both a time and an event per subject (use Surv.from_arrays). Everything outside that constraint — and sparse X, which is unsupported — is covered by tests/test_sklearn_compat.py.

Regression & non-parametric models

from comprisk import FineGrayRegression

fg = FineGrayRegression(cause=1, robust_se=True).fit(X, time=time, event=event)
print(fg.coef_, fg.se_)                    # log subdistribution-HRs
Estimator Estimates R parity
FineGrayRegression subdistribution-hazard ratios cmprsk::crr() (β̂ to fp noise)
PenalizedFineGrayRegression LASSO / ridge / EN / MCP / SCAD path crrp::crrp() to ~1e-6
CauseSpecificCox cause-specific hazard ratios survival::coxph() to 1e-9
CumulativeIncidence non-parametric Aalen-Johansen CIF cmprsk::cuminc()
gray_test K-sample test for equal CIFs cmprsk::cuminc()$Tests to 1e-14

Worked code for every row is in examples/02_regression_models.ipynb.

comprisk vs alternatives

comprisk randomForestSRC scikit-survival
Language Python R Python
Native competing risks ✓ ✓ ✗ (single-event)
Aalen–Johansen CIF output ✓ ✓ n/a
Cumulative hazard at scale ✓ ✓ ✗ (low-memory only)
OOB permutation VIMP ✓ ✓ ✗
Bit-identical reproducibility mode ✓ (equivalence="rfsrc") — n/a
Scales to n = 10⁶ ✓ (63 s on i7) memory-bound ✗ / OOM
GPU preview ✓ (CUDA 12) ✗ ✗

scikit-survival's CHF/survival outputs and scaling caveats are detailed in the benchmarks.

Benchmarks

Matched-pair, real EHR data (full tables + methodology in docs/benchmarks.md):

Cohort n × p comprisk rfSRC (OMP-on) Speedup
CHF (cardio) 75k × 58 5.6–9.4 s 84.8–207.3 s 14–22×
SEER breast 238k × 17 7.0 s 81.6 s 11.6×

Both fit similarly well (C ≈ 0.85); the band tracks feature count. Also 16.6–544× vs scikit-survival (n = 5k → 50k) and n = 10⁶ in 63 s on a consumer i7.

Roadmap

comprisk is intentionally CR-focused — for non-CR survival (general Cox, AFT, deep-survival), use lifelines or scikit-survival.

  • Shipped (v0.3–0.6): CR forest, Fine-Gray (+ penalized), cause-specific Cox, Aalen-Johansen CIF, Gray's test, score_cr / calibration_cr.
  • v1.0 (planned): API freeze + JMLR MLOSS submission.
  • v1.1 (planned): full GPU rewrite.

Documentation

📖 Full documentation site — searchable, autogenerated API reference.

Examples

Runnable notebooks in examples/ (rendered on GitHub; open in Colab to run):

Development

Requires uv.

uv venv && uv pip install -e ".[dev]"
uv run pre-commit install
uv run pytest && uv run ruff check .

License & citation

Apache-2.0 (LICENSE, NOTICE). If you use comprisk in your research, please cite the paper:

@article{yang_comprisk_2026,
  author        = {Yang, Sunny and Zhao, Weiyan and Zhao, Wanqi},
  title         = {{comprisk: A scikit-learn-compatible Python toolkit for competing-risks survival analysis}},
  year          = {2026},
  eprint        = {2607.09431},
  archivePrefix = {arXiv},
  primaryClass  = {stat.CO},
  url           = {https://arxiv.org/abs/2607.09431},
}

To cite a specific release, use the archived Zenodo DOI (concept-level 10.5281/zenodo.19876282, resolves to latest) or GitHub's "Cite this repository" button (CITATION.cff).

Metadata

Release files for comprisk 0.8.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for comprisk 0.8.0
File Size Uploaded
comprisk-0.8.0.tar.gz 337.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for comprisk 0.8.0
File Interpreter ABI Platform
comprisk-0.8.0-py3-none-any.whl Python 3 none any Details

Total release size: 491.9 kB

Release files / comprisk-0.8.0.tar.gz

Download URL comprisk-0.8.0.tar.gz
Size 337.2 kB
Tags Source
SHA-256 checksum
How to use checksums
731ed2880cbe0dd60be86c94ca9d549703645c3c4fff00ca1cd13f2f0c0a222d
BLAKE2b-256 checksum
How to use checksums
dc16e8804c9aafd22f7859ac24238c9d975ccaa576d8bdb6c8c2816d3e8b97cb
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.9

Release files / comprisk-0.8.0-py3-none-any.whl

Download URL comprisk-0.8.0-py3-none-any.whl
Size 154.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
3b4cb03434028dafd7a4ea0f74e665eb057cf0a7bd07d97dfa8008bc381c2e6b
BLAKE2b-256 checksum
How to use checksums
46d5130be09d592e09be7034eafae478c440ed90908f0c712a144f4efbf83681
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.9

Release history Release notifications | RSS feed

This release

0.8.0 This release

2 release files

0.7.0

2 release files

0.6.0

2 release files

0.5.0

2 release files

0.4.0

2 release files

0.3.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page