A Python package for batch correction (CALM) and activity scoring (CARZ, TAS) of gene expression data.
Project description
calmcarz
A Python package for batch correction (CALM) and activity scoring (CARZ, TAS) of perturbed expression data
calmcarz is a lightweight and easy-to-use Python package for normalizing high-throughput gene expression data from perturbation screens. It provides a simple, function-based workflow to remove technical batch effects and calculate robust, interpretable scores for transcriptional activity.
The pipeline consists of three main steps:
- CALM (Control-Anchored Linear Model): Removes batch effects by fitting a linear model on control samples.
- CARZ (Control-Anchored Robust Z-score): Standardizes gene expression relative to the control distribution using robust statistics.
- TAS (Transcriptional Activity Score): Summarizes the overall degree of perturbation for each cell based on the number of significantly altered genes.
Installation
You can install calmcarz directly from PyPI using pip:
pip install calmcarz
Ensure you have the necessary dependencies (or it will be installed automatically: numpy, pandas, anndata, statsmodels, scipy, and tqdm)
Usage
The library is designed to work seamlessly with AnnData objects.
import anndata as ad
from calmcarz import calm_correction, calculate_carz, calculate_tas
# 1. Load your data into an AnnData object
# This assumes your data is in a log-transformed layer (e.g., 'log1p_counts')
# and adata.obs contains 'batch' and 'treatment' columns.
adata = ad.read_h5ad("path/to/your/data.h5ad")
# 2. Apply CALM for batch correction
# The corrected data will be stored in a new layer, 'calm_corrected'.
calm_correction(
adata,
plate_column='rna_plate',
treatment_column='cmap_name',
control_treatment_value='DMSO',
corrected_layer_name='X_CALM' # the CALM-corrected expression will be store in layer 'X_CALM'
)
# 3. Calculate CARZ scores
# CARZ scores will be stored in a new layer, 'carz_scores'.
calculate_carz(
adata,
plate_key="rna_plate",
layer="X_CALM", # calculate using CALM-corrected expression
pert_type_key="pert_type",
control_categories=["ctl_vehicle"], # calculate using plate controls
key_added="X_CARZ", # the CARZ matrix will be stored in layer 'X_CARZ'
copy=False,
)
# 4. Calculate Transcriptional Activity Score (TAS)
# The final TAS will be added to adata.obs.
calculate_tas(
adata,
layer='X_CARZ',
tas_obs_key='TAS_score',
alpha=2.5 # Set the significance threshold for CARZ scores
)
# View the results
print(adata.obs[['cmap_name', 'TAS_score']].head())
API Overview
calm_correction(adata, ...)
Corrects for batch effects using a control-anchored linear model.
adata: AnnData object. Assumes expression data in adata.X (cells x genes). adata.obs must contain theplate_columnandtreatment_column.plate_column: Name of the column in adata.obs that identifies the plate.treatment_column: Name of the column in adata.obs that identifies the treatment.control_treatment_value: Value intreatment_columnthat identifies control samples.corrected_layer_name: Name of the layer in adata to store the corrected expression data. Original adata.X is not modified.min_control_samples_per_gene_for_fitting: Minimum number of total control samples (across all plates) required for a gene to attempt model fitting.min_control_variance_for_fitting: Minimum variance of expression in control samples for a gene to attempt model fitting.verbose: If True, prints progress information.
calculate_carz(adata, ...)
Calculates robust Z-scores relative to the control distribution.
adata: Annotated data matrix (cells x genes).plate_key: Key inadata.obsfor plate identifiers.pert_type_key: Key inadata.obsindicating perturbation type.control_categories: List of strings inadata.obs[pert_type_key]that identify control samples.layer: Layer inadata.layersto use as input. If None,adata.X.key_added: Name of the layer to store results. If None, overwrites input.copy: Whether to modify adata inplace (False) or return a copy (True).
calculate_tas(adata, ...)
Calculates transcriptional activity score (TAS).
adata: AnnData object with Z-scored expression data inadata.X.adata.Xshould have cells as rows and genes as columns.layer: Layer inadata.layersto use as input. If None,adata.X.alpha: The threshold for the absolute Z-score.tas_obs_key: The key under which the TAS will be stored inadata.obs. Defaults to "TAS_score".gene_subset: Optional list of gene names (strings). If provided, TAS will be calculated only for these genes. If None, all genes inadata.Xwill be used. Defaults to None.
Contributing
Contributions are welcome! If you have suggestions for improvements or find a bug, please feel free to open an issue or submit a pull request on the GitHub repository.
License
This project is licensed under the MIT License. See the LICENSE file for details.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file calmcarz-0.1.0.tar.gz.
File metadata
- Download URL: calmcarz-0.1.0.tar.gz
- Upload date:
- Size: 12.2 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.1.0 CPython/3.11.11
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
68a591d01da116cbd27e73e65e46bdf75a3adc226a61b53b237cdd8d779a8e42
|
|
| MD5 |
103b44db3edac32e3e84ee13d2bcf220
|
|
| BLAKE2b-256 |
1eac3f23c5ebe87b95030599f9dacd813d01232a1f0cdd3b7b7744bde98a2fbe
|
File details
Details for the file calmcarz-0.1.0-py3-none-any.whl.
File metadata
- Download URL: calmcarz-0.1.0-py3-none-any.whl
- Upload date:
- Size: 11.6 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.1.0 CPython/3.11.11
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
9d37e7cd96da2e2637374795f421c87aab86c11e0f1143392aa2ab630e7a164f
|
|
| MD5 |
dbbd14946c91d7c8a0b2249d5d8ce72d
|
|
| BLAKE2b-256 |
2542e45f172644709b8f95ecbb609b0ac751d0ab2f4bafb05644b7e504ece6ca
|