Skip to main content

Demultiplex multi-reference BAM files into per-reference buckets and call consensus

Project description

midsplit

Demultiplex a multi-reference BAM file into per-reference buckets and call a consensus sequence for each non-empty bucket.

What it does

When reads are aligned to multiple reference sequences in a single BAM file, midsplit assigns each read (or read pair) to the reference(s) it matches best, then produces a separate BAM, consensus FASTA, and per-site statistics file for each reference that received at least one read.

The classification uses the NM (edit-distance) tag to compute a percent identity for each alignment. A read is assigned to a reference if its percent identity is at least --threshold times the best percent identity that read achieves across all references (default 0.95). Paired-end reads are treated as a unit using an overlap-aware combined percent identity, so both mates always land in the same bucket(s).

Output

For each reference that receives reads, midsplit writes:

File Contents
{ID}.bam / {ID}.bam.bai Sorted, indexed per-reference BAM
{ID}-consensus.fasta Consensus sequence called by ivar
{ID}-per-site.tsv Per-position depth, A/C/G/T counts, ref base, and consensus base
summary.txt Run-level statistics and consensus-vs-reference comparison

The per-site TSV has columns: site, ref_base, consensus_base, depth, A, C, G, T. When --align is used, the consensus base is mapped back to the correct reference position even when ivar has inserted or deleted bases relative to the reference.

Usage

midsplit [options] INPUT_BAM

Options

Option Default Description
--output-dir DIR . Directory for all output files (created if absent)
--threshold FLOAT 0.95 Minimum fraction of best PID to assign a read
--reference FASTA Multi-reference FASTA; enables ref_base column and consensus comparison
--align off Align consensus to reference before comparison (recommended when lengths differ)
--aligner mafft Aligner for --align: mafft, needle, or edlib
--aligner-options OPTIONS Extra options forwarded to the aligner (implies --align)
--consensus-quality INT 20 Minimum base quality passed to ivar (-q)
--consensus-frequency-threshold FLOAT 0.0 Minimum frequency for ivar to call a base (-t)
--consensus-low-coverage INT 0 Depth below which ivar masks with N (-m)
--consensus-id TEMPLATE ID for the consensus sequence; use {ID} to embed the reference name

Example

midsplit \
  --reference references/multi.fasta \
  --output-dir results/ \
  --align \
  --threshold 0.95 \
  alignments/reads-vs-multi.bam

Requirements

  • Python 3.11+
  • samtools and ivar on PATH
  • Python dependencies are managed with uv; run uv sync to install them

Notes

  • Only primary and secondary alignments are used; supplementary alignments (chimeric/split reads) are skipped.
  • Reads aligned with Bowtie2 --all or -k N emit non-best hits as secondary alignments; midsplit includes these in classification so that all alignment evidence is used.
  • For circular genomes (e.g. HBV) mapped against a linearised reference, local alignment (bowtie2 --very-sensitive-local) is strongly recommended over end-to-end alignment. End-to-end mode cannot soft-clip reads that span the linearisation junction, which introduces artefactual bases near position 1 of the reference and can corrupt the consensus at those positions.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

midsplit-0.1.0.tar.gz (126.1 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

midsplit-0.1.0-py3-none-any.whl (17.3 kB view details)

Uploaded Python 3

File details

Details for the file midsplit-0.1.0.tar.gz.

File metadata

  • Download URL: midsplit-0.1.0.tar.gz
  • Upload date:
  • Size: 126.1 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.9.6

File hashes

Hashes for midsplit-0.1.0.tar.gz
Algorithm Hash digest
SHA256 761578d34b2e7e4871f7ce95599de06d4701cae02ae5b4e28219123fdaeec5db
MD5 660756655f57d50e9131394e2852156c
BLAKE2b-256 396a71c69b9def4a809fc1aa929dc605ead2da79b7f743bfc0bd4062c5b877da

See more details on using hashes here.

File details

Details for the file midsplit-0.1.0-py3-none-any.whl.

File metadata

  • Download URL: midsplit-0.1.0-py3-none-any.whl
  • Upload date:
  • Size: 17.3 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.9.6

File hashes

Hashes for midsplit-0.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 1818ca7f791679640b4f763b833d8242956ab9091a8932f9af30b165e110e186
MD5 c53f9035e2490b1564308cfe872a2212
BLAKE2b-256 da6c1c87d4b983ad2b14b5d6a4ca9dd44835a4d45dd4cb9b539519d4d85b3358

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page