Skip to main content

splicekit: comprehensive toolkit for splicing analysis from short-read RNA-seq

Project description

splicekit: an integrative toolkit for splicing analysis from short-read RNA-seq

splicekit is a modular platform for splicing analysis from short-read RNA-seq datasets. The platform also integrates an JBrowse2 instance, pybio for genomic operations and scanRBP for RNA-protein binding studies. The whole analysis is self-contained (one single folder) and the platform is written in Python, in a modular way.

Check a short video presentation about splicekit (poster) at ECCB 2023 on Youtube:

Quick start

Since version 0.7, splicekit is a Snakemake pipeline, and there is also a Conda environment yaml file.

git clone git@github.com:bedapub/splicekit.git    # clone rep
cd splicekit                                      # change working directory

micromamba -y create -f splicekit.yaml            # create conda env
micromamba activate splicekit                     # activate env
./install.sh                                      # install dependencies
pip install .                                     # install splicekit

cd datasets/GSE126543                             # move to sample folder
./1_download.sh                                   # download sample FASTQs
pybio homo_sapiens                                # human genome

./run_snakemake_local.sh --configfile config.yaml # run snakemake (local)
./run_snakemake_slurm.sh --configfile config.yaml # OR run snakemake SLURM

After snakemake finishes, you can explore results interactively running splicekit web and follow instructions on how to open the html reports in your browser.

Installing splicekit directly from the GitHub repository
pip install git+https://github.com/bedapub/splicekit.git@main
If you already have aligned reads in BAM files

All you need is samples.tab (note that this is a TAB delimited file) and splicekit.config in one folder (check datasets for examples).

You can easily download and prepare the reference genome (e.g. $ pybio genome homo_sapiens).

Finally run ./run_snakemake_[local/slurm].sh --configfile config.yaml (inside the folder with samples.tab and splicekit.config).

Easiest is to check datasets examples to see how the above files look like and also to check scripts if you need to map reads from FASTQ files with pybio.

Documentation

Changelog

v0.8.1: released in July 2026

  • added bam_file column support in samples.tab for per-sample BAM paths (subfolder layouts)
  • new get_bam_path(sample_id) helper in splicekit/core/annotation.py — falls back to {bam_path}/{sample_id}.bam when no per-sample path is given
  • exons.py, genes.py, anchors.py: replaced os.listdir() BAM discovery with annotation.samples list + get_bam_path() (enables subfolder BAMs, no dir scan needed)
  • junctions.py, jbrowse2.py: same get_bam_path() adoption
  • default bam_column = "bam_file" added to config

v0.8: released in October 2025

  • removed platform config option (now snakemake submits jobs to the cluster)
  • pandas and other minor improvements

v0.7: released in February 2025

Past change notes (click to view)
v0.6: released in April 2024
  • updated reports
  • JUNE analysis (junction-events to classify skipped and mutually exclusive exons)

v0.4.9: released in November 2023

  • added rMATS analysis for splicing events
  • added Docker container that can be directly imported to singularity via ghcr.io
  • fixed dependencies
  • other small fixes

v0.4: released in May 2023

  • added singularity container with all dependencies
  • added local integrated JBrowse2
  • cluster or desktop runs
  • scanRBP and bootstrap analysis of RNA-protein binding
  • further development and integration with pybio
  • extended documentation of concepts, analysis and results

v0.3: released in January 2023 (click to show details)

  • re-coded junction analysis
    • independent junctions parsing from provided bam files
    • master table of all junctions in the samples of the analyzed project, including novel junctions (refseq/ensembl non-annotated)
  • clustering by logFC of pairwise-comparisons with dendrogram: junction, exon and gene levels (clusterlogfc module)
  • added first_exon annotation for junctions touching annotated first exons of transcripts
  • extended documentation of concepts, analysis and results

v0.2: released in October 2022

  • software architecture restructure with python modules
  • filtering of lowly expressed features by edgeR
  • DonJuan analysis (junction-anchor analysis)
  • more advanced motif analysis with DREME
  • filtering regulated junctions with regulated donors

v0.1: released in July 2022

  • initial version of splicekit
  • parsing of junction and exon counts
  • computing edgeR analysis from count tables and producing a results file with direct links to JBrowse2
  • basic motif analysis

Citing and Contact

If you find splicekit useful in your work and research, please cite:

Rot, G., Wehling, A., Schmucki, R., Berntenis, N., Zhang, J. D., & Ebeling, M. (2024)
splicekit : an integrative toolkit for splicing analysis from short-read RNA-seq
Bioinformatics Advances, 4(1). https://doi.org/10.1093/bioadv/vbae121

In case of questions, issues and other ideas, please use the GitHub Issues or write directly to Gregor Rot.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

splicekit-0.8.2.tar.gz (75.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

splicekit-0.8.2-py3-none-any.whl (81.8 kB view details)

Uploaded Python 3

File details

Details for the file splicekit-0.8.2.tar.gz.

File metadata

  • Download URL: splicekit-0.8.2.tar.gz
  • Upload date:
  • Size: 75.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.11

File hashes

Hashes for splicekit-0.8.2.tar.gz
Algorithm Hash digest
SHA256 6533ae5c95e81d46e5ccd268cd5f824fb655acf5946da94c3771f7fd9e36c112
MD5 ceeeade8c16c31da053d3f997e8bda73
BLAKE2b-256 b92315f3c66def7f00f0ee46fa00c9a5ea449e58480b8f40a5ccc624a6f2c07c

See more details on using hashes here.

File details

Details for the file splicekit-0.8.2-py3-none-any.whl.

File metadata

  • Download URL: splicekit-0.8.2-py3-none-any.whl
  • Upload date:
  • Size: 81.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.11

File hashes

Hashes for splicekit-0.8.2-py3-none-any.whl
Algorithm Hash digest
SHA256 6746ebc68186ee9cce8df4ddaf75eff131ecba7720d5c1ed15842eb577816486
MD5 f0b751864417146a6d55d6c837c8c4a6
BLAKE2b-256 094da00a759c9ba95ee6297856761d7fb5a4f88d51434f9dfbd9d558502312b8

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page