Skip to main content

CI codecov Docs Python 3.10-3.13 License: MIT

CloneTrast

Revealing pan-cancer clonal niches of T cells from single-cell RNA sequencing using contrastive learning

CloneTrast maps gene expression alone to a latent space where cells from the same T-cell clone cluster together. At inference you only need scRNA-seq UMI counts, without paired scTCR-seq. Public pre-trained model (hosted on Figshare) is downloaded automatically on first use, with an additional complementary model used to predict clone size from gene expression alone. Following, you can visualize your data and further explore the clonal functional structure for downstream analysis.

Graphical description

CloneTrast figure 1 overview

Clone labels vs. TCR sequences

CloneTrast separates how models are trained from how they are applied. This distinction is central to what the embedding represents.

During training: categorical clone labels - not TCR sequence identity

Training requires a per-cell clone_id column: a categorical label that groups cells from the same T-cell clone. Those labels are inferred from scTCR-seq (please refer to our manuscript for more details). Importantly, CloneTrast is not trained on TCR sequences. Supervision comes only from the clone_id category.

At inference: scRNA-seq only - no scTCR-seq

When you embed new cells (with the public pre-trained model), CloneTrast reads gene expression only. You do not need scTCR-seq, CDR3 sequences, or clone_id labels. The model maps each cell’s transcriptome to a fixed-dimensional embedding (X_clonetrast).

What this means for the embedding space

During training, every cell with the same clone_id is a positive pair: the contrastive loss rewards them for being located next to each other. At inference, the encoder applies that learned mapping - cells from the same clone are expected to cluster in X_clonetrast when their expression reflects shared clonal identity.

This specific modeling approach results in a distinct embedding space capturing multiple components of clonal organization:

(1) Different clones, similar functional states: Two clones defined by different TCR sequences can be positioned near one another if they share similar functional status.

(2) Same clone, divergent states: T cells belonging to the same clone are pulled toward one another despite having diverse differentiation states due to intra-clonal transcriptional heterogeneity.

(3) Same TCR sequence, different functional states: Two clones with identical TCR sequences but distinct clonal identities, because they originate from different patients, can be positioned far apart if they exhibit different functional states.

Training and inference description

CloneTrast figure 2 overview

Installation

For complete installation instructions including prerequisites, package installation, and development setup, please see the Installation Guide.

Quick start - inference with pre-trained models

You only need gene expression (scRNA-seq). CloneTrast does not use scTCR-seq or TCR sequences at inference - see Clone labels vs. TCR sequences above. Optional clone_id labels are for coloring plots or computing metrics only.

Pre-trained models

The pre-trained models expect Ensembl gene IDs in adata.var_names (e.g. ENSG00000156234). Input genes are automatically aligned to the training gene list (11,950 genes): missing genes are set to zero (thus mimicking sequencing drop-outs), extra genes are dropped.

Data format (inference)

  • AnnData with gene expression (cells × genes), e.g. raw UMI counts in adata.X. Raw counts require subsequent row-normalization and log-transformation, as described below.
  • adata.var_names: Ensembl IDs for compatibility with the public checkpoints.
  • Optional: adata.obs['clone_id'] for visualization or evaluation (not required to run the model).
  • Importantly: the model was trained using T cells alone. It is therefore required to ensure your data includes only T cells as input.

Python API

import scanpy as sc
import clonetrast as ct

# These are raw UMI counts
adata = sc.read_h5ad("your_data.h5ad")

# Raw counts should be subsequently normalized
sc.pp.normalize_total(adata, target_sum = 1e4)
sc.pp.log1p(adata)

# Applying the contrastive model and creating a two-dimensional representation with UMAP
ct.tl.embed(
    adata,
    use_pretrained=True,          # default; downloads contrastive model on first run
    device="cuda",                # or "cpu"
)
ct.tl.compute_umap(adata, obsm_key="X_clonetrast", umap_key="X_umap_clonetrast")

# Optional complementary model to predict log clone size per cell
ct.tl.predict_clone_size(adata, use_pretrained=True, device="cuda")

# Visualize
sc.pl.embedding(adata, basis="umap_clonetrast", color=["predicted_clone_size"])

Embeddings are stored in adata.obsm['X_clonetrast']; UMAP coordinates in adata.obsm['X_umap_clonetrast'].

Tutorial notebooks

Two step-by-step tutorials:

  1. Data preparation - prepare paired scRNA/TCR-seq data for training (QC, T-cell subtyping, clone labeling, export)
  2. Applying CloneTrast - apply CloneTrast to gene expression alone (embed, predict clone size, visualize and compare to a standard UMAP)

Documentation

Full documentation (includes the contrastive loss reference and tutorials).

License

This project is licensed under the MIT License - see the LICENSE file for more details.

Citation

Reference will be added here when available.

Acknowledgements

CloneTrast's logo, the project's graphical description, and the graphical description of the training and inference, were created with BioRender.com using a paid license.


This project was created in favor of the scientific community worldwide, with a special dedication to the cancer research community.

We hope you'll find this repository helpful, and we warmly welcome any requests or suggestions - please don't hesitate to reach out!

Visitor Map

Metadata

Release files for clonetrast 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for clonetrast 0.1.1
File Size Uploaded
clonetrast-0.1.1.tar.gz 5.7 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for clonetrast 0.1.1
File Interpreter ABI Platform
clonetrast-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 5.7 MB

Release files / clonetrast-0.1.1.tar.gz

Download URL clonetrast-0.1.1.tar.gz
Size 5.7 MB
Tags Source
SHA-256 checksum
How to use checksums
27d6d71600c7783b77a282be880ecfb20e81a55b7bb3d740cb243f03d489cb9e
BLAKE2b-256 checksum
How to use checksums
cfefb954d8293f00c9a0c07a61b69433061afa67a1a590be61f5a22e9bc54dd2
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 1, 2026.

Transparency log

Release files / clonetrast-0.1.1-py3-none-any.whl

Download URL clonetrast-0.1.1-py3-none-any.whl
Size 41.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
50e9f64f2705964b69b2af14656f11bb57aa8a84a0f1e28c9bdc5d7a7dbdf4c8
BLAKE2b-256 checksum
How to use checksums
5d50829bd73b995ad8742b109251fd975299b36461e0b4442c823a41c856cabd
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 1, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.1 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page