Skip to main content

gbcms

Complete orientation-aware counting system for genomic variants

Tests Python 3.10+ Ask DeepWiki

Features

  • 🚀 High Performance: Rust-powered core engine with multi-threading
  • 🧬 Complete Variant Support: SNP, MNP, insertion, deletion, and complex variants (DelIns, SNP+Indel)
  • 🧪 WFA + PairHMM Phase 3: Pangenomic fast-path WFA alignment with PairHMM fallback for complex multi-allelic classification
  • 📊 Orientation-Aware: Forward and reverse strand analysis with fragment counting
  • 📏 mFSD (Mutant Fragment Size Distribution): Per-allele cfDNA fragment size profiling with KS test and log-likelihood ratio
  • 🔬 Statistical Analysis: Fisher's exact test for strand bias (read-level and fragment-level)
  • 📁 Flexible I/O: BAM and CRAM input; VCF and MAF variant input/output formats
  • 🎯 Quality Filters: 8 configurable read and quality filtering options with heuristic BAQ
  • 🧬 RNA Mode: Transcriptome-aware counting with strandedness, splice detection, and A-to-I editing
  • 🔗 UMI Support: Molecule-level deduplication with UMI-aware fragment grouping
  • 🔬 Per-Molecule Observations: Export which molecule carried which allele at each variant — the layer beneath the counts — enabling read-backed phasing and allelic imbalance
  • 🔧 Normalize Command: Standalone variant normalization (left-align + REF validation) without counting

Installation

Quick install:

pip install gbcms

From source (requires Rust):

git clone https://github.com/msk-access/gbcms.git
cd gbcms
pip install .

Docker:

docker pull ghcr.io/msk-access/gbcms:X.Y.Z  # Replace X.Y.Z with latest from PyPI

💡 Find the latest version on PyPI or GHCR.

📖 Full documentation: https://msk-access.github.io/gbcms/


Usage

gbcms can be used in two ways:

🔧 Option 1: Standalone CLI (1-10 samples)

Best for: Quick analysis, local processing, direct control

gbcms dna \
    --variants variants.vcf \
    --bam sample1.bam \
    --fasta reference.fa \
    --output-dir results/

Output: results/sample1.vcf

Learn more:


🔄 Option 2: Nextflow Workflow (10+ samples, HPC)

Best for: Many samples, HPC clusters (SLURM), reproducible pipelines

nextflow run nextflow/main.nf \
    --input samplesheet.csv \
    --variants variants.vcf \
    --fasta reference.fa \
    --mode dna \
    -profile slurm

Features:

  • ✅ Automatic parallelization across samples
  • ✅ SLURM/HPC integration
  • ✅ Container support (Docker/Singularity)
  • ✅ Resume failed runs

Learn more:


Which Should I Use?

Scenario Recommendation
1-10 samples, local machine CLI
10+ samples, HPC cluster Nextflow
Quick ad-hoc analysis CLI
Production pipeline Nextflow
Need auto-parallelization Nextflow
Full manual control CLI

Quick Examples

CLI: DNA Single Sample

gbcms dna \
    --variants variants.vcf \
    --bam tumor.bam \
    --fasta hg19.fa \
    --output-dir results/ \
    --threads 4

CLI: RNA-seq

gbcms rna \
    --variants variants.vcf \
    --bam rna_sample:aligned.bam \
    --fasta hg19.fa \
    --rna-editing-db TABLE1_hg38.txt.gz \
    --output-dir results/

CLI: Normalize Variants

gbcms normalize \
    --variants variants.vcf \
    --fasta hg19.fa \
    --output results/normalized.tsv

CLI: Multiple Samples (Sequential)

gbcms dna \
    --variants variants.vcf \
    --bam-list samples.txt \
    --fasta hg19.fa \
    --output-dir results/

Nextflow: Many Samples (Parallel)

# samplesheet.csv:
# sample,bam,bai
# tumor1,/path/to/tumor1.bam,
# tumor2,/path/to/tumor2.bam,

nextflow run nextflow/main.nf \
    --input samplesheet.csv \
    --variants variants.vcf \
    --fasta hg19.fa \
    --mode dna \
    --outdir results \
    -profile slurm

Documentation

📚 Full Documentation: https://msk-access.github.io/gbcms/

Quick Links:


Contributing

See CONTRIBUTING.md for development guidelines.

To contribute to documentation, see the gh-pages branch.


Citation

If you use gbcms in your research, please cite:

Shah, R. et al. (2026). gbcms: A high-performance orientation-aware genotype counting system for genomic variants. Available at: https://github.com/msk-access/gbcms

BibTeX:

@software{gbcms,
  author       = {Shah, Ronak and contributors},
  title        = {gbcms: A high-performance orientation-aware genotype counting system for genomic variants},
  year         = {2026},
  url          = {https://github.com/msk-access/gbcms},
  note         = {GitHub repository}
}

License

AGPL-3.0 - see LICENSE for details.


Support

Release files for gbcms 6.4.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for gbcms 6.4.0
File Size Uploaded
gbcms-6.4.0.tar.gz 299.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for gbcms 6.4.0
File Interpreter ABI Platform
gbcms-6.4.0-cp311-cp311-manylinux_2_34_x86_64.whl CPython 3.11 CPython 3.11 Linux glibc 2.34+ x86-64 Details

Total release size: 7.3 MB

Release files / gbcms-6.4.0.tar.gz

Download URL gbcms-6.4.0.tar.gz
Size 299.7 kB
Tags Source
SHA-256 checksum
How to use checksums
aabdc7938728c2fc67ab611d9f0da9a1e7037f822649a9324251147c018c2fbe
BLAKE2b-256 checksum
How to use checksums
947f01e7cc4fcdae0774edf8eed863cb65889d9a0c72d3d9c47cb3e67bdcaed2
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 23, 2026.

Transparency log

Release files / gbcms-6.4.0-cp311-cp311-manylinux_2_34_x86_64.whl

Download URL gbcms-6.4.0-cp311-cp311-manylinux_2_34_x86_64.whl
Size 7.0 MB
Tags CPython 3.11 Linux glibc 2.34+ x86-64
SHA-256 checksum
How to use checksums
a11a401dae1cca48cd6be0caf55a8ac1f0a50766ba187a633be59188869c5740
BLAKE2b-256 checksum
How to use checksums
34ea87e809c339c84c024569838f5f18d9a9008e0f11425fca192cca2b0b314c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 23, 2026.

Transparency log

Release history Release notifications | RSS feed

6.5.0

2 release files

This release

6.4.0 This release

2 release files

6.3.1

2 release files

6.3.0

2 release files

6.2.0

2 release files

6.1.0

2 release files

6.0.0

2 release files

5.3.0

2 release files

5.2.0

2 release files

5.1.0

2 release files

5.0.0

2 release files

4.2.0

2 release files

4.1.0

2 release files

4.0.1

2 release files

4.0.0

2 release files

3.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page