Prismalign
N-color nucleotide-conversion alignment engine with pluggable backends.
Prismalign maps sequencing reads from any nucleotide-conversion chemistry
(bisulfite-seq C→T, SLAM-seq T→C, m6A / A-to-I A→G, MK/KM dual-base,
or a custom 3rd channel) using a HISAT-3N-style strategy:
- build a converted reference index (
scheme.ref_from → ref_to) - transform each read per color channel and align it to the converted index via a pluggable backend (bwamem by default; WFA2-lib; minimap2/mappy; minibwa, all optional)
- re-score every hit against the original reference so that real
conversions are rewarded (not counted as mismatches), emitting a
color-correct
MDplus per-channelY/Zcounts in BAM tags.
All per-read heavy kernels are native C (BWA-MEM / WFA2-lib / minimap2 / minibwa); the Python layer is a thin, friendly wrapper.
Install
pip install -e . # bwamem + built-in WFA2 C backends
pip install -e "./[mappy]" # + minimap2 backend
Usage — Python (clean wrapper)
import prismalign as ps
# one-shot mapping -> BAM (builds indexes, maps, cleans up)
ps.map_reads("reads.fq", "ref.fa", "out.bam",
scheme="MK", backend="bwamem", threads=4)
# object API / reuse
with ps.NColorMapper(scheme=ps.BS, backend="bwamem") as mapper:
mapper.map_file("reads.fq", ref_files=["ref.fa"], output_files=["bs.bam"])
Usage — CLI
# classic two-color (MK: A->G + C->T); --backend auto picks the fastest
# importable backend (minibwa > mappy > bwamem)
prismalign map -s MK -r ref.fa -o out.bam reads.fq
# force the fastest native backend (bwa-mem2 speed tier)
prismalign map -s MK --backend minibwa -r ref.fa -o out.bam reads.fq
# bisulfite-seq (3-nt single channel C->T)
prismalign map -s BS -r genome.fa -o bs.bam --index-dir idx reads.fq
# parallel (2 copies of the reads, byte-identical output to -t 1)
prismalign map -s MK -r ref.fa -o out.bam -t 4 reads.fq
# list built-in schemes
prismalign schemes
Schemes
| name | reference index | channels | use case |
|---|---|---|---|
MK |
AC→GT |
2 | dual-base conversion A→G + C→T (classic two-color) |
KM |
GT→AC |
2 | reverse of MK |
BS |
C→T |
1 | bisulfite-seq (3-nt) |
SLAM |
T→C |
1 | SLAM-seq |
A2G |
A→G |
1 | m6A / A-to-I editing |
THREE |
AC→GT |
3 | three-color demo (add your 3rd base pair in schemes.py) |
Python API
from prismalign import NColorMapper, BS
mapper = NColorMapper(scheme=BS, backend="bwamem", index_dir="idx")
mapper.map_file(r1_file="reads.fq", ref_files=["genome.fa"],
output_files=["out.bam"])
Backends vs Adapters — two integration layers
Prismalign plugs in aligners at two layers:
- Backends — in-process (compiled C / a Python binding). The engine's per-read
map_onecallsbackend.align(seq)directly. Selectable via--backend. - Adapters — subprocess wrappers for external command-line aligners. They
map reads at three granularities (
map_read/map_batch/map_file), unified on theCliAdapterbase; not--backend-selectable.
Backends (in-process, --backend)
| backend | engine | notes |
|---|---|---|
bwamem |
BWA-MEM via the bwamem package |
default, fast C backend (SE + PE) |
minibwa |
lh3/minibwa (bwa-mem successor) via PyO3 pip binding minibwa (fg-labs) |
~2-3x faster than bwa-mem; pip install minibwa (SE + PE) |
mappy |
minimap2 via mappy |
official minimap2 Python binding (SE + PE) |
bwamem2 |
bwa-mem2 native in-process via the bwamem2 Cython binding |
correct, but per-read; pip install 'prismalign[bwamem2]' (standard, non-free-threaded CPython) |
wfa2 |
WFA2-lib (vendored v2.3.6, MIT) compiled in-process | exact gapped (indel-aware) wavefront alignment; SE only |
Adapters (subprocess, prismalign.adapters)
All built on the shared CliAdapter base, which maps reads at three granularities:
| method | granularity | notes |
|---|---|---|
map_read(seq) |
1 read | compat / debug (slow: one subprocess per read) |
map_batch(reads) |
1 batch ((name,seq,qual) list) |
one tool invocation |
map_file(fastq, batch_size, threads, workers) |
whole FASTQ (plain/.gz) | path prismalign reads + batches itself (default batch_size=8192 = 8×bwa-mem2's internal 512); workers>1 runs batches concurrently (each an independent tool subprocess — the real throughput lever, since a tool's own threads saturate on a shared index) |
| adapter | tool | index |
|---|---|---|
BwaMemAdapter |
bwa mem |
bwa index |
BwaMem2Adapter |
bwa-mem2 mem |
bwa-mem2 index |
Bowtie2Adapter |
bowtie2 |
bowtie2-build |
Minimap2Adapter |
minimap2 -a |
none (reads FASTA directly) |
Hisat2Adapter |
hisat2 |
hisat2-build |
StrobealignAdapter |
strobealign |
.sti |
SamAdapter |
any SAM mapper | (generic; you supply the command template) |
Each adapter needs its binary on PATH (or an env var: BWA_BIN, BWA_MEM2_BIN,
BOWTIE2_BIN, MINIMAP2_BIN, HISAT2_BIN, STROBEALIGN_BIN).
Throughput tip: for a large job use map_file(path, batch_size=8192, workers=cpu/2..cpu, threads=1) — batching at 8192 (a multiple of the tool's
internal 512) amortizes the subprocess startup, and workers parallelizes
batches across separate tool processes (each with its own memory bandwidth;
the tool's own -t saturates, so parallelize via processes). .gz input is
read with an internal fast library (isal/xopen/rapidgzip) when available.
bwa-mem2 appears at BOTH layers (intentional)
BwaMem2Backend(backend) = bwa-mem2 in-process, per-read — for embedding /--backend bwamem2.BwaMem2Adapter(adapter) = bwa-mem2 batched CLI (map_file) — for throughput.
For large references, bwa-mem2's ~2-3x (from SIMD FM-index search) shows up; build it cleanly for AVX2/AVX-512 (stale object files cause SIGILL). The in-process backend is per-read (batch the calls, or use the adapter, for throughput).
Full inventory — including where each wrapper lives — in docs/backends.md.
Speed & IO
- Auto backend:
--backend auto(explicit) picks the fastest importable native backend —minibwa(~2-3x BWA-MEM, the bwa-mem2 speed tier), elsemappy(in-process), elsebwamem. (Default isbwamem.) The chosen backend is printed at startup ([prismalign] backend auto=minibwa). - Process parallelism is the lever:
-t/--threads Nmaps reads in an ordered fork+COW process pool (any backend); batches are drained in read order so the BAM is byte-identical tothreads=1. Processes (each with its own index) scale; bwa-mem2's internal-tthreads do not (they share one index and are memory-bound). So scale processes/workers, not bwa-mem2 threads. - Reduced repeated IO: references are copy+converted once even when reused across layers (cache keyed by path+scheme); per-hit reference fetch is cached in memory for small contigs (RNA/transcript references), so only one indexed read per contig.
Throughput ceiling. The heavy per-read work (BWA-MEM / minibwa / minimap2 / WFA2 kernels) is native C. With BWA-MEM the runtime is essentially that kernel's throughput — prismalign's glue (conversion, re-scoring, tag emission) adds only a small fraction. For the largest runs, use
minibwa(runs on standard CPython; no free-threaded-3.14 wheel) ormappy; both are far faster than BWA-MEM.
Limitations (v0.2.x)
- paired-end is supported natively by the
bwamem,mappyandminibwabackends;wfa2is single-end for now (the subprocesssam/strobealign/bowtie2/bwa_mem2adapters are also SE). - hierarchical (layered) mapping uses the
PLAINidentity scheme for non-converted short-RNA references.minibwais the fastest native backend but needs standard (non-free-threaded) CPython ≤ 3.13, so on 3.14tmappyis the fastest available.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
File details
Details for the file prismalign-0.2.9.tar.gz.
File metadata
- Download URL: prismalign-0.2.9.tar.gz
- Upload date:
- Size: 178.9 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
134bc75ab4965b1c9d6af2d2402a8117b5be09ced5240d7145cfde948dcde9a8
|
|
| MD5 |
40a9c6d55a61ba88e1057739d442af9e
|
|
| BLAKE2b-256 |
3f58870200ec260a7c83ec49569de282e29bfb32840590d3b3f4d647c143aac3
|
Provenance
The following attestation bundles were made for prismalign-0.2.9.tar.gz:
Publisher:
workflow.yml on y9c/prismalign
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
prismalign-0.2.9.tar.gz -
Subject digest:
134bc75ab4965b1c9d6af2d2402a8117b5be09ced5240d7145cfde948dcde9a8 - Sigstore transparency entry: 2713208820
- Sigstore integration time:
-
Permalink:
y9c/prismalign@376520023be1a80b2c6ec81568fbe24d7d2a27f9 -
Branch / Tag:
refs/heads/main - Owner: https://github.com/y9c
-
Access:
private
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
workflow.yml@376520023be1a80b2c6ec81568fbe24d7d2a27f9 -
Trigger Event:
push
-
Statement type: