causarray
Advances in single-cell sequencing and CRISPR technologies have enabled detailed case-control comparisons and experimental perturbations at single-cell resolution. However, uncovering causal relationships in observational genomic data remains challenging due to selection bias and inadequate adjustment for unmeasured confounders, particularly in heterogeneous datasets. To address these challenges, we introduce causarray [Du26], a doubly robust causal inference framework for analyzing array-based genomic data at both bulk-cell and single-cell levels. causarray integrates a generalized confounder adjustment method to account for unmeasured confounders and employs semiparametric inference with flexible machine learning techniques to ensure robust statistical estimation of treatment effects.
Usage
We recommend using causarray in a conda environment:
# create a new conda environment and install the necessary packages
conda create -n causarray python=3.12 -y
# activate the environment
conda activate causarray
On an Apple Silicon Mac, make sure the environment is native: python -c "import platform; print(platform.machine())" should print arm64. A conda started from
a terminal running under Rosetta creates Intel (x86_64) environments, in which
causarray runs several times slower; causarray warns about this at import.
Create a native one with CONDA_SUBDIR=osx-arm64 conda create -n causarray python=3.12 -y, then conda config --env --set subdir osx-arm64 inside it.
The module can be installed via PyPI:
pip install causarray
For optimal parallel performance, we recommend installing llvm-openmp if using conda:
conda install -c conda-forge llvm-openmp
For R users, reticulate can be used to call causarray from R while
keeping NumPy >2 in the Python environment. Create the separate R environment
with a current NumPy-2-compatible reticulate build:
conda env create -f environment-r.yaml
The R tutorial runs from causarray-r and connects to the Python package in
the causarray environment. The two must share an architecture: on an Apple
Silicon Mac, create both natively (CONDA_SUBDIR=osx-arm64 conda env create -f environment-r.yaml), since reticulate cannot load an arm64 Python into an
Intel R.
The documentation and tutorials using both Python and R are available at causarray.readthedocs.io.
Tutorials
| Tutorial | Language | Description | Link |
|---|---|---|---|
| Perturb-seq [Jin20] | Python | CRISPR screen analysis on excitatory neurons | Notebook |
| Perturb-seq [Jin20] | R | Same analysis using reticulate |
Notebook |
| Genome-wide CRISPRi screen [Replogle22] | Python | Batch fitting on 200 perturbations from a K562 genome-wide CRISPRi screen | Notebook |
| Case-control: SEA-AD [Gabitto24] | Python | Causal inference on observational single-cell data (Alzheimer's disease) | Notebook |
Batch fitting API
For screens with hundreds to thousands of perturbations, use gcate_lfc_batch so
that peak memory is bounded by one batch at a time:
from causarray import gcate_lfc_batch
df_res = gcate_lfc_batch(
Y, X, A, r,
batch_size=10, # perturbations per batch (or use n_batches= for a fixed count)
max_cells=2000, # max pert cells per batch (ctrl added on top)
n_ctrl=2000, # fixed ctrl subsample shared across batches
cache_path='results.h5', # resume if interrupted
verbose=True,
)
See the Replogle-E-K562 tutorial for a demonstration on 200 perturbations from a genome-wide CRISPRi screen.
Diagnostic masks without refitting effects
Treatment-by-gene support or quality-control rules can be aligned to an existing causarray result table by label, even when its rows are reordered:
from causarray import align_test_mask
keep = align_test_mask(
df_res,
support_mask, # treatments × genes, Boolean
treatment_names=perturbations,
gene_names=genes,
)
df_res_flagged = df_res.assign(support_keep=keep)
This operation only annotates or subsets existing results. It does not refit the causarray LFC or change standard errors and p-values. If a diagnostic rule was selected after inspecting the outcomes, retain the original adjusted p-values rather than redefining the multiple-testing family post hoc. The Replogle tutorial compares several expression-support rules with marginal Wilcoxon results.
Changelog
See CHANGELOG for a full version history.
References
[Du26] Jin-Hong Du, Maya Shen, Hansruedi Mathys, and Kathryn Roeder. "Uncovering causal relationships in single cell omic studies with causarray". In: Briefings in Bioinformatics (2026).
[Gabitto24] Mariano I. Gabitto et al. "Integrated multimodal cell atlas of Alzheimer's disease". In: Nature Neuroscience (2024).
[Jin20] Xin Jin et al. "In vivo Perturb-seq reveals neuronal and glial abnormalities associated with autism risk genes". In: Nature Neuroscience (2020).
[Replogle22] Joseph M. Replogle et al. "Mapping information-rich genotype-phenotype landscapes with genome-scale Perturb-seq". In: Cell (2022).
Metadata
Release files for causarray 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| causarray-0.1.1.tar.gz | 74.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| causarray-0.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 154.0 kB
Release files / causarray-0.1.1.tar.gz
| Download URL | causarray-0.1.1.tar.gz |
|---|---|
| Size | 74.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
bef85073cfee6fc04ecf22b762f30e805e1c7a369ffe3d68319b887beb6cb6f2
|
|
BLAKE2b-256 checksum How to use checksums |
5cecf23aad9e81f66d08e878dd28d7aa719b61e844551479443c931eb7d1c9ad
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.
Transparency logRelease files / causarray-0.1.1-py3-none-any.whl
| Download URL | causarray-0.1.1-py3-none-any.whl |
|---|---|
| Size | 79.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
14728cb45e9f6cc3cc024c8ec7aec99264ed3c9fbfad92295ca706a949b17427
|
|
BLAKE2b-256 checksum How to use checksums |
91bdd23649c4256d95f44cbf0eee23cda3f1d3fb77e69fcbbc14156e146e95ee
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.
Transparency log