Revealing pan-cancer clonal niches of T cells from single-cell RNA sequencing using contrastive learning
CloneTrast maps gene expression alone to a latent space where cells from the same T-cell clone cluster together. At inference you only need scRNA-seq UMI counts, without paired scTCR-seq. Public pre-trained model (hosted on Figshare) is downloaded automatically on first use, with an additional complementary model used to predict clone size from gene expression alone. Following, you can visualize your data and further explore the clonal functional structure for downstream analysis.
Graphical description
Clone labels vs. TCR sequences
CloneTrast separates how models are trained from how they are applied. This distinction is central to what the embedding represents.
During training: categorical clone labels - not TCR sequence identity
Training requires a per-cell clone_id column: a categorical label that groups cells from the same T-cell clone. Those labels are inferred from scTCR-seq (please refer to our manuscript for more details). Importantly, CloneTrast is not trained on TCR sequences. Supervision comes only from the clone_id category.
At inference: scRNA-seq only - no scTCR-seq
When you embed new cells (with the public pre-trained model), CloneTrast reads gene expression only. You do not need scTCR-seq, CDR3 sequences, or clone_id labels. The model maps each cell’s transcriptome to a fixed-dimensional embedding (X_clonetrast).
What this means for the embedding space
During training, every cell with the same clone_id is a positive pair: the contrastive loss rewards them for being located next to each other. At inference, the encoder applies that learned mapping - cells from the same clone are expected to cluster in X_clonetrast when their expression reflects shared clonal identity.
This specific modeling approach results in a distinct embedding space capturing multiple components of clonal organization:
(1) Different clones, similar functional states: Two clones defined by different TCR sequences can be positioned near one another if they share similar functional status.
(2) Same clone, divergent states: T cells belonging to the same clone are pulled toward one another despite having diverse differentiation states due to intra-clonal transcriptional heterogeneity.
(3) Same TCR sequence, different functional states: Two clones with identical TCR sequences but distinct clonal identities, because they originate from different patients, can be positioned far apart if they exhibit different functional states.
Training and inference description
Installation
For complete installation instructions including prerequisites, package installation, and development setup, please see the Installation Guide.
Quick start - inference with pre-trained models
You only need gene expression (scRNA-seq). CloneTrast does not use scTCR-seq or TCR sequences at inference - see Clone labels vs. TCR sequences above. Optional clone_id labels are for coloring plots or computing metrics only.
Pre-trained models
The pre-trained models expect Ensembl gene IDs in adata.var_names (e.g. ENSG00000156234). Input genes are automatically aligned to the training gene list (11,950 genes): missing genes are set to zero (thus mimicking sequencing drop-outs), extra genes are dropped.
Data format (inference)
- AnnData with gene expression (cells × genes), e.g. raw UMI counts in
adata.X. Raw counts require subsequent row-normalization and log-transformation, as described below. adata.var_names: Ensembl IDs for compatibility with the public checkpoints.- Optional:
adata.obs['clone_id']for visualization or evaluation (not required to run the model). - Importantly: the model was trained using T cells alone. It is therefore required to ensure your data includes only T cells as input.
Python API
import scanpy as sc
import clonetrast as ct
# These are raw UMI counts
adata = sc.read_h5ad("your_data.h5ad")
# Raw counts should be subsequently normalized
sc.pp.normalize_total(adata, target_sum = 1e4)
sc.pp.log1p(adata)
# Applying the contrastive model and creating a two-dimensional representation with UMAP
ct.tl.embed(
adata,
use_pretrained=True, # default; downloads contrastive model on first run
device="cuda", # or "cpu"
)
ct.tl.compute_umap(adata, obsm_key="X_clonetrast", umap_key="X_umap_clonetrast")
# Optional complementary model to predict log clone size per cell
ct.tl.predict_clone_size(adata, use_pretrained=True, device="cuda")
# Visualize
sc.pl.embedding(adata, basis="umap_clonetrast", color=["predicted_clone_size"])
Embeddings are stored in adata.obsm['X_clonetrast']; UMAP coordinates in adata.obsm['X_umap_clonetrast'].
Tutorial notebooks
Two step-by-step tutorials:
- Data preparation - prepare paired scRNA/TCR-seq data for training (QC, T-cell subtyping, clone labeling, export)
- Applying CloneTrast - apply CloneTrast to gene expression alone (embed, predict clone size, visualize and compare to a standard UMAP)
Documentation
Full documentation (includes the contrastive loss reference and tutorials).
License
This project is licensed under the MIT License - see the LICENSE file for more details.
Citation
Reference will be added here when available.
Acknowledgements
CloneTrast's logo, the project's graphical description, and the graphical description of the training and inference, were created with BioRender.com using a paid license.
Metadata
Release files for clonetrast 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| clonetrast-0.1.1.tar.gz | 5.7 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| clonetrast-0.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 5.7 MB
Release files / clonetrast-0.1.1.tar.gz
| Download URL | clonetrast-0.1.1.tar.gz |
|---|---|
| Size | 5.7 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
27d6d71600c7783b77a282be880ecfb20e81a55b7bb3d740cb243f03d489cb9e
|
|
BLAKE2b-256 checksum How to use checksums |
cfefb954d8293f00c9a0c07a61b69433061afa67a1a590be61f5a22e9bc54dd2
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 1, 2026.
Transparency logRelease files / clonetrast-0.1.1-py3-none-any.whl
| Download URL | clonetrast-0.1.1-py3-none-any.whl |
|---|---|
| Size | 41.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
50e9f64f2705964b69b2af14656f11bb57aa8a84a0f1e28c9bdc5d7a7dbdf4c8
|
|
BLAKE2b-256 checksum How to use checksums |
5d50829bd73b995ad8742b109251fd975299b36461e0b4442c823a41c856cabd
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 1, 2026.
Transparency log