Coralsnake
Conventional genomics tools are built around DNA — a linear, double-stranded
reference where each locus lines up 1:1 with the sequence — and often lack
explicit consideration of RNA's structural properties: an abundance
hierarchy (ribosomal rRNA is orders of magnitude more abundant than
mRNA), strand orientation (sense vs. antisense), and splicing (mRNAs
are assembled from exons, so a transcript does not line up with its genomic
locus). Coralsnake is an exon-aware RNA analysis pipeline built around
exactly these properties: it turns a GTF/GFF into spliced
transcript references (prepare), splices and joins reads between
transcript and genome coordinates in both
directions (liftover), and runs the analyses you do on the results —
annotate places sites/variants on the RNA hierarchy (5'UTR / CDS / 3'UTR /
intronic / intergenic) and calls the variant effect, metagene profiles how
sites distribute across 5'UTR / CDS / 3'UTR, motif fetches the strand-aware
reference motif around each site, and logo renders a DNA/RNA sequence logo.
Installation
pip install coralsnake
The visualization commands (metagene plot, sequence logo) need the lightweight
plot extra, which only pulls in matplotlib when you need it:
pip install "coralsnake[plot]"
Commands
| Command | What it does |
|---|---|
prepare |
Extract the spliced primary transcript reference from GTF/GFF. |
liftover |
Splice/join reads between genome/transcript BAMs (-d t2g default, -d g2t inverts). |
annotate |
Unified site/variant annotation: region on the RNA hierarchy + gene/transcript + variant effect. |
metagene |
Exon-aware metagene profiling across 5'UTR/CDS/3'UTR (profile plot needs coralsnake[plot]). |
motif |
Strand-aware genomic motif fetch around variant sites. |
coordinate |
Map chromosome names between coordinate systems (UCSC↔Ensembl). |
group |
Group genes and build a consensus sequence. |
logo |
Plot a DNA/RNA sequence logo (needs coralsnake[plot]). |
annotateis the single annotation tool — one command, one schema, two input modes:--reference-gtf(region + gene/transcript, and the full variant effect when given a genome FASTA + ref/alt) or--annotation <table>(fast precomputed-table site labeling).
Quick example
A typical end-to-end run (read alignment is done by any external mapper, e.g.
bwa / prismalign):
# 1. Build the spliced transcript reference from the annotation
coralsnake prepare -g annotation.gtf -f genome.fa -o transcript.fa --with-txpos
# 2. Align reads to transcript.fa with an external mapper → tx.bam
# 3. Splice the transcript-aligned BAM back to genome coordinates
coralsnake liftover -d t2g -i tx.bam -o genome.bam -a annotation.tsv -f genome.fai
# 4. Annotate sites to genes/transcripts (RNA hierarchy + variant effect)
coralsnake annotate -i sites.tsv -o annotated.tsv \
--reference-gtf annotation.gtf \
--reference-transcript genome.fa -s -a
# 5. Exon-aware metagene profile across 5'UTR / CDS / 3'UTR
coralsnake metagene -i sites.tsv -g annotation.gtf -o profile.tsv -p profile.png
Subcommands
prepare — spliced transcript reference
Build the spliced transcript reference (the target that reads are aligned to) from a GTF/GFF and a genome FASTA:
coralsnake prepare -g annotation.gtf -f genome.fa -o transcript.fa \
--with-codon --with-genename --filter-biotype protein_coding
liftover — splice-aware BAM conversion
prepare builds the transcript reference; liftover round-trips a BAM between
transcript and genome coordinates, splicing reads at exon boundaries:
coralsnake liftover -d t2g(default) — transcript BAM → genome BAM (splices reads at exon boundaries, inserts introns).coralsnake liftover -d g2t— genome BAM → transcript BAM (clips to exons, joins spliced reads contiguously on the transcript).
annotate — exon-aware annotation
Label a site (chrom,pos,strand) with gene/transcript/position + region
(5'UTR/CDS/3'UTR/intronic/intergenic); add a genome FASTA + ref/alt to get the
full variant effect (codon/AA + mut_type). Fast precomputed-table mode via
--annotation <table>.
coralsnake annotate -i sites.tsv -o out.tsv --reference-gtf annotation.gtf -c 1,2,3
coralsnake annotate -i variants.tsv -o effects.tsv \
--reference-gtf annotation.gtf \
--reference-transcript genome.fa -s -a
metagene — exon-aware metagene profiling
Built on the high-performance polars + ruranges stack. Computes the
distribution of sites relative to gene regions (5'UTR, CDS, 3'UTR) and can emit
binned statistics and a publication-ready profile plot. The -p profile plot
needs the plot extra; the tabular outputs (-o, -s) work without it.
# Using a built-in reference (GRCh38) or a custom GTF:
coralsnake metagene -i sites.tsv.gz -r GRCh38 -H -m 1,2,3 -w 5 \
-o output.tsv -s scores.tsv -p plot.png
coralsnake metagene -i sites.bed -g custom.gtf.gz -m 1,2,3 -w 5 \
-o output.tsv -s scores.tsv -p plot.png
List or download the built-in references:
coralsnake metagene --list
coralsnake metagene --download GRCh38
Python API — the functions are importable from the flat modules:
from coralsnake.io import load_sites, load_reference
from coralsnake.gtf import load_gtf
from coralsnake.annotation import map_to_transcripts, normalize_positions
from coralsnake.map_to_local import map_to_local
from coralsnake.plotting import plot_profile
sites = load_sites("sites.tsv.gz", with_header=True, meta_col_index=[0, 1, 2])
ref = load_reference("GRCh38") # or load_gtf("custom.gtf.gz")
annotated = map_to_transcripts(sites, ref)
gene_bins, gene_stats, gene_splits = normalize_positions(
annotated, split_strategy="median", bin_number=100
)
plot_profile(gene_bins, gene_splits, "metagene_plot.png")
# Map global coordinates to local transcript coordinates (strand-aware):
local = map_to_local(sites, ref, ref_id_col="transcript_id")
motif & coordinate
A migration of the standalone variant package — fused into the top-level CLI
with the old pyfaidx / urllib3 dependencies removed and coralsnake's
pysam + ruranges stack used instead. Naming and output format are unchanged.
# Motif fetch (strand-aware, padded with N)
coralsnake motif -i sites.tsv -o motifs.tsv -f genome.fa -n 2,3 -w
# Chromosome-name mapping (UCSC ↔ Ensembl)
coralsnake coordinate -i sites.tsv -o mapped.tsv -M U2E
from coralsnake.motif import get_motif
from coralsnake.coordinate import run_coordinate
from coralsnake.annotate import run_annotate
group — gene clustering & consensus
Group related genes and build a consensus sequence:
coralsnake group -f genes.fa -g genes.gtf -o grouped.tsv \
--output-consensus consensus.fa --threads 8
logo — sequence logo
Plots a DNA/RNA sequence logo from a set of motif sequences. The scoring engine
is pure numpy; the renderer needs matplotlib (plot extra).
coralsnake logo -m ACGT -m ACGG -m CCGT -o logo.png
# or with per-motif weights from a file (seq\tcount)
coralsnake logo -i motifs.tsv -o logo.svg
from coralsnake import Mlogo
m = Mlogo(motifs=["ACGT", "ACGG", "CCGT"], to2bit=True)
m.plot(ax) # requires matplotlib (plot extra)
Performance
The core is built on the vectorized polars + ruranges stack, using Rust-backed
ruranges primitives instead of slow per-group Python applies:
map_to_transcriptspicks the best transcript per gene with a vectorized sort +group_by().first()(wasgroup_by().map_groups()python apply) — ~20× faster on realistic inputs.map_to_localusesruranges.numpy.group_cumsumfor strand-aware cumulative transcript offsets (was a hand-rolledmap_groupsapply) — ~7× faster.Mlogo(sequence logo) builds its score matrix with vectorizednumpy(bincount+ codepoint lookup) — ~1.5× faster and fixes a0·log2(0)NaN edge case.
Documentation
- Architecture & Design — package layout and design decisions.
- Full docs site: https://coralsnake.yech.science/ (see
docs/).
Mapping is out of scope. Nucleotide-conversion (two-color / three-color) mapping is not part of coralsnake any more: it lives in the dedicated
prismalignpackage (pluggable backends: bwamem, minimap2/mappy, pure-Python), on top of the lightweightbwamemBWA-MEM binding.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distributions
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file coralsnake-0.1.1.tar.gz.
File metadata
- Download URL: coralsnake-0.1.1.tar.gz
- Upload date:
- Size: 113.1 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
73ed4c20d62c0472f37766a27a72cdfbcaabc04394ffb5be95bad3a0bb2a4808
|
|
| MD5 |
9f22369ffb589eda81026f252b69c716
|
|
| BLAKE2b-256 |
fa8c3d561e46458071dde6436cbfe39bb366aed31dc8efbc5e79de85d6f58046
|
File details
Details for the file coralsnake-0.1.1-cp313-cp313-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.
File metadata
- Download URL: coralsnake-0.1.1-cp313-cp313-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
- Upload date:
- Size: 104.1 kB
- Tags: CPython 3.13, manylinux: glibc 2.17+ x86-64
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
e99b3488bb695a56230b83d97301fc4824784cb2d14626fd65140f02b128fa8f
|
|
| MD5 |
a5606f6d73c13243f40df7a4aad59b52
|
|
| BLAKE2b-256 |
5e529e779ce490713b2c60a8987926f339626ea9299475762a9916f97a471419
|
File details
Details for the file coralsnake-0.1.1-cp312-cp312-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.
File metadata
- Download URL: coralsnake-0.1.1-cp312-cp312-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
- Upload date:
- Size: 104.1 kB
- Tags: CPython 3.12, manylinux: glibc 2.17+ x86-64
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
c86be4dc64fd1e4611e56a61466c3f19aab26dcf8ba68b0bb9ece22f68f697f3
|
|
| MD5 |
8fef1a8ce995b5ab90a1dfece81f6f63
|
|
| BLAKE2b-256 |
403bc1da021f456e8919dcc2809495bc90f561a2a8b572f66e027a9c75261aac
|