InSituCNV + CopyKAT backend
⚠️ Attribution — this is a fork
This repository is a fork of Moldia/InSituCNV. All credit for InSituCNV — the method, the package, and the original code — belongs to the InSituCNV authors (Jensen et al., Moldia lab). This fork adds only an optional CopyKAT backend as an alternative to
infercnvpy; everything else is their work, unchanged.
- Original repository: https://github.com/Moldia/InSituCNV
- Original paper: Jensen et al., In Situ inference of copy number variations in image-based spatial transcriptomics, bioRxiv (2025), https://doi.org/10.1101/2025.07.02.662761
- CopyKAT: Gao et al., navinlabcode/copykat
If you use this work, please cite the original InSituCNV paper (and CopyKAT if you use that backend). Please direct questions about the core method to the upstream project.
Reusable notebook workflow for inferring copy-number variation (CNV) profiles from image-based spatial transcriptomics data.
This repository contains the general package version of the workflow used in In Situ Inference of Copy Number Variations in Image-Based Spatial Transcriptomics by Jensen et al. The manuscript reproduction code is kept in a separate repository: https://github.com/Moldia/InSituCNV-manuscript
CopyKAT backend (this fork)
This fork adds an alternative CNV-calling engine based on the R package
CopyKAT as a drop-in replacement for
infercnvpy. Everything else in the pipeline — loading, spatial smoothing,
CNV-space clustering, chromosome heatmaps and spatial plots — is shared between
the two backends, because run_copykat writes CopyKAT's output into the same
adata.obsm["X_cnv"] / adata.uns["cnv"]["chr_pos"] keys that infercnvpy uses.
Select the backend with the backend argument:
import insitucnv as icv
# inferCNV (default)
icv.run_insitucnv(adata, output_dir="out_infercnv", reference_key="cell_type")
# CopyKAT
icv.run_insitucnv(
adata,
output_dir="out_copykat",
reference_key="cell_type",
backend="copykat",
copykat_kwargs={"genome": "hg20", "min_gene_per_cell": 5}, # e.g. use_raw_counts=True
)
By default CopyKAT is fed InSituCNV's spatially-smoothed layer ("M"); pass
copykat_kwargs={"use_raw_counts": True} to use raw counts instead. Reference
cells (reference_categories) are handed to CopyKAT as norm.cell.names to seed
its diploid baseline. Note: CopyKAT was built for scRNA-seq with thousands of
genes; its min.gene.per.cell default (200) is relaxed here for small image-based
panels. See insitucnv/tl/copykat.py for details.
Runtime. CopyKAT's cost is ~O(cells²) (repeated cell×cell parallelDist +
hclust), e.g. ~2 h for ~10k cells. run_copykat skips CopyKAT's redundant internal
heatmap by default (skip_internal_plots=True, ~30% faster, identical results). Using
fewer cells is the strongest lever; to still label every cell from a fast subsampled
run, insitucnv.tl.propagate_cnv_labels (or the comparison script's --propagate
flag) copies the calls to all cells by transcriptomic cluster or kNN — a copied label,
not per-cell inferred CNV.
Environment and install (CopyKAT backend)
conda env create -f insitucnv_copykat.yml
conda activate insitucnv_copykat
R -e 'remotes::install_github("navinlabcode/copykat", dependencies=FALSE)'
pip install -e .
From the command line:
python scripts/run_insitucnv_copykat.py \
--input-dir /path/to/xenium_output \
--output-dir results/sample1 \
--reference-key cell_type
References
InSituCNV builds upon several excellent packages in the single-cell and spatial transcriptomics ecosystem:
- infercnvpy: Core CNV inference logic.
- scVelo: Transcriptomic smoothing and analysis.
- Scanpy: General single-cell analysis framework.
- AnnData: Data structures for single-cell data.
If you use InSituCNV in your research, please cite:
Jensen et al. In Situ inference of copy number variations in image-based spatial transcriptomics. bioRxiv (2025). https://doi.org/10.1101/2025.07.02.662761
Installation
This fork is published on PyPI as insitucnv-copykat, because insitucnv is
upstream's name. The import name is unchanged:
pip install insitucnv-copykat # then: import insitucnv
Install upstream's own release with pip install insitucnv instead if you do not need
the CopyKAT backend or the relaxed anndata/pandas bounds.
Clone the repository and install the package in editable mode:
git clone https://github.com/Moldia/InSituCNV.git
cd InSituCNV
conda env create -f insitucnv.yml
conda activate insitucnv_env
pip install -e ".[dev,docs]"
jupyter lab
Open the notebooks from the notebooks/ directory in JupyterLab.
Development Setup
Install the package and development tools into your active environment:
pip install -e ".[dev,docs]"
Notebooks and examples should import the installed package directly:
import insitucnv as icv
No sys.path edits are required when the editable install is active.
Running Tests
pytest
Building Docs Locally
sphinx-build -b html docs docs/_build/html
The documentation uses Sphinx, the Read the Docs theme, autodoc/autosummary API
pages, and myst-nb for notebook tutorials.
Releasing Package
Build the source distribution and wheel locally:
python -m build
Publish TestPyPI releases through the publish-testpypi.yml GitHub Actions
workflow after configuring Trusted Publishing in TestPyPI. Do not store PyPI or
TestPyPI API tokens in the repository.
What You Need Before Running
Prepare an .h5ad file with the information needed by the notebooks:
- raw counts, preferably in
adata.layers["raw_counts"]; if raw counts are inadata.X, the first notebook can copy them intoadata.layers["raw_counts"]; - spatial coordinates in
adata.obsm["spatial"]; - a nearest-neighbor graph for smoothing, usually from
scanpy.pp.neighbors; - an
adata.obscolumn that identifies normal or healthy reference cells forinfercnvpy; - the exact category names in that reference column that should be used as the normal reference, for example immune, stromal, or other non-tumor cell types;
- gene names that can be matched to genomic coordinates. The package includes
the default
infercnvpygene coordinate table, and the notebooks show where to adjust this if your gene annotation differs.
Do not use normalized or log-transformed values as raw counts. The CNV workflow normalizes, smooths, and log-transforms the raw counts itself.
Notebook Workflow
Run the notebooks in order for your own dataset, editing the setup cells at the top of each notebook.
1. Run InSituCNV
notebooks/run_insitucnv.ipynb
Use this notebook to run the full CNV analysis step by step:
- load your spatial transcriptomics
.h5adfile; - choose the raw count layer name with
RAW_LAYER; - choose the normal/reference annotation with
REFERENCE_KEYandREFERENCE_CATEGORIES; - normalize raw counts;
- smooth normalized counts over the neighbor graph;
- add genomic positions;
- run
infercnvpy; - cluster CNV profiles across selected Leiden resolutions;
- visualize chromosome heatmaps and spatial CNV cluster plots.
After reviewing the heatmap and spatial plot, manually edit:
NORMAL_CLUSTERS;TUMOR_CLUSTERS, if you want to specify tumor clusters directly;TUMOR_CLONE_CLUSTERS, if some tumor clusters have distinct enough CNV profiles to report as separate tumor clones.
No cluster is marked as normal automatically.
Outputs
The notebooks write results under their configured output directories, usually
results/... or outputs/.... Typical outputs include:
- checked input
.h5adfiles; - AnnData objects with CNV values and CNV cluster labels;
- CNV chromosome heatmaps;
- spatial CNV cluster plots;
- optional manually annotated tumor/normal and tumor-clone cell group tables;
- optional mean CNV profiles for manually selected tumor cells.
Metadata
Release files for insitucnv-copykat 0.3.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| insitucnv_copykat-0.3.0.tar.gz | 44.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| insitucnv_copykat-0.3.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 78.7 kB
Release files / insitucnv_copykat-0.3.0.tar.gz
| Download URL | insitucnv_copykat-0.3.0.tar.gz |
|---|---|
| Size | 44.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
e9e2f095ecdcf291168c8eb3450ea1a98e70faf2492e90ac2e9a60b8d0357ee4
|
|
BLAKE2b-256 checksum How to use checksums |
37afd4d5750400fed9844cdfd09f9296bac27c090cd81eec97c87f56498ad435
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.10.12
|
Release files / insitucnv_copykat-0.3.0-py3-none-any.whl
| Download URL | insitucnv_copykat-0.3.0-py3-none-any.whl |
|---|---|
| Size | 34.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
82fd93282541034c0d54787e53e2e583c03f579ac2c83c37a758bb0646eac4ad
|
|
BLAKE2b-256 checksum How to use checksums |
e311f36e22e34bbef60feaa34be2d080657fafc3bca929fcb68103311e09fe0b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.10.12
|