Skip to main content

scAce: an adaptive embedding and clustering method for scRNA-seq data

Overview

Overview

scAce is consisted of three major steps, a pre-training step based on a variational autoencoder, a cluster initialization step to obtain initial cluster labels, and an adaptive cluster merging step to iteratively update cluster labels and cell embeddings. In the pre-training step, scAce takes the single-cell gene expression matrix as its input to train a VAE network. For each gene, the VAE learns and outputs three parameters of a ZINB distribution (mean, dispersion, and proportion of zero). In the cluster initialization step, scAce offeres two manners. With de novo initialization, Leiden is used to obtain initial cluster labels; with clustering enhancement, initial cluster labels are obtained by applying a cluster splitting approach to a set of existing clustering results. In the adaptive cluster merging step, given the pre-trained VAE network and the initial cluster labels, the network parameters, cell embeddings, cluster labels and centroids are iteratively updated by alternately performing network update and cluster merging steps. The final results of cell embeddings and cluster labels are output by scAce after the iteration process stops.

Installation

Please install scAce from pypi with:

pip install scace

Or clone this repository and use

pip install -e .

in the root of this repository.

Quick start

Load the data to be analyzed:

import scanpy as sc

adata = sc.AnnData(data)

Perform data pre-processing:

# Basic filtering
sc.pp.filter_genes(adata, min_cells=3)
sc.pp.filter_cells(adata, min_genes=200)

adata.raw = adata.copy()

# Total-count normlize, logarithmize and scale the data  
sc.pp.normalize_per_cell(adata)
adata.obs['scale_factor'] = adata.obs.n_counts / adata.obs.n_counts.median()

sc.pp.log1p(adata)
sc.pp.scale(adata)

Run the scAce method:

from scace import run_scace
adata = run_scace(adata)

The output adata contains cluster labels in adata.obs['scace_cluster'] and the cell embeddings in adata.obsm['scace_emb']. The embeddings can be used as input of other downstream analyses.

Please refer to tutorial.ipynb for a detailed description of scAce's usage.

Release files for scace 0.1.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for scace 0.1.2
File Size Uploaded
scace-0.1.2.tar.gz 12.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for scace 0.1.2
File Interpreter ABI Platform
scace-0.1.2-py3-none-any.whl Python 3 none any Details

Total release size: 25.5 kB

Release files / scace-0.1.2.tar.gz

Download URL scace-0.1.2.tar.gz
Size 12.1 kB
Tags Source
SHA-256 checksum
How to use checksums
5ebd434902a3614f4a57d9cf36844e4932845f81c7fb1415ea177f69e653a53d
BLAKE2b-256 checksum
How to use checksums
eb8ed9f1b28346b9837dac34c91a9611de2e451c4542a43081f186d59a251c9e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.2 CPython/3.8.8

Release files / scace-0.1.2-py3-none-any.whl

Download URL scace-0.1.2-py3-none-any.whl
Size 13.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
3020024a7a1f4dc32288fc358436c574053a4bf61c956c7f97700f403e8487ae
BLAKE2b-256 checksum
How to use checksums
6a347c32d3428231ef94f58b4422b3770dfca47788cff8041338675f82ec408c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.2 CPython/3.8.8

Release history Release notifications | RSS feed

This release

0.1.2 This release

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page