Skip to main content

bam2bw

Downloads Tests

A command-line tool for converting SAM/BAM files of reads, or .tsv/tsv.gz files of fragments, into either stranded or unstraded basepair resolution bigWig files. By default, only the 5' end of reads are mapped (not the full span of the read) and these bigWig file(s) contain the integer count of reads mapping to each basepair. Optionally, both the 3' and 5' of the entry can be mapped if they correspond to fragments, such as from ATAC-seq experiments. As a convenience, the starts and ends can be shifted (e.g., to account for Tn5 bias), a scaling factor can be used to multiply the mapped counts at each basepair, and read depth normalization can be applied to make the sum across the bigWigs be equal to 1. When a scaling factor and read depth normalization are used together, the sum across the two bigWigs is equal to the scaling factor.

bam2bw does not produce any intermediary files and can even stream SAM/BAM files remotely (but not .tsv/.tsv.gz). This means that you can go directly from finding a SAM/BAM file somewhere on the internet to the bigWig files used to train ML programs without several time-consuming steps. v0.4.0 allows parallel processing of files, even if they are remote, reducing the time needed to process inputs to just the time needed to process the biggest one.

usage: bam2bw [-h] -s SIZES [-u] [-f | -3p] [-ps POS_SHIFT] [-ns NEG_SHIFT] [-mp] [--rna5 {read1,read2}] [--opposite_strand] [-sf SCALE_FACTOR] [-r] [-p PARALLEL] -n NAME [-z ZOOMS] [-v] filename [filename ...]

This tool will convert BAM files to bigwig files without an intermediate.

positional arguments:
  filename              The SAM/BAM or tsv/tsv.gz file to be processed.

options:
  -h, --help            show this help message and exit
  -s SIZES, --sizes SIZES
                        A chrom_sizes, .fai, or FASTA file. Only the first two
                        columns of a chrom_sizes/.fai file are read. A
                        compressed FASTA must be BGZF, not gzip.
  -u, --unstranded      Have only one, unstranded, output.
  -f, --fragments       The data is fragments and so both ends should be recorded.
  -3p, --three_prime    Record the 3' end of each read instead of the 5' end.
  -ps POS_SHIFT, --pos_shift POS_SHIFT
                        A shift to apply to positive strand reads.
  -ns NEG_SHIFT, --neg_shift NEG_SHIFT
                        A shift to apply to negative strand reads.
  -mp, --mate_pairs     Treat paired-end BAM/SAM reads as a single RNA/fragment tag instead of
                        counting each mate independently (see --rna5/--opposite_strand).
  --rna5 {read1,read2}  Which mate carries the 5' end of the RNA/fragment. Only used with
                        --mate_pairs. Default: read1.
  --opposite_strand     Report the strand of the mate opposite the one chosen by --rna5.
                        Only used with --mate_pairs.
  -sf SCALE_FACTOR, --scale_factor SCALE_FACTOR
                        A scaling factor to multiply each position by.
  -r, --read_depth      Whether to divide through by total (pre-scaled) read depth.
  -p PARALLEL, --parallel PARALLEL
                        The number of jobs to use, max of one per input file.
  -n NAME, --name NAME
  -z ZOOMS, --zooms ZOOMS
                        The number of zooms to store in the bigwig.
  -v, --verbose

Installation

pip install bam2bw

Timings

These timings involve the processing of https://www.encodeproject.org/files/ENCFF638WXQ/ which has slightly over 70M reads. Local means applied to a file that was already downloaded, and remote means including the downloading time.

bam2bw (local): 2m10s
bam2bw (remote): 4m50s
existing pipeline (local): 18m5s

On my compute server, I can usually get over 1.5M records/second when reading BAM files locally and have been able to stream three BAMs from the ENCODE Portal at ~800k records/second, processing the almost 945M records in just over 8 minutes (ENCFF337YBN.bam, ENCFF981FXV.bam, ENCFF144GBU.bam). Admittedly, the first setting will be influenced by the quality of your hard drive, and the second setting by the quality of your internet connection (and whether the speed is throttled).

Usage

(1) On a local file:

bam2bw my.bam -s hg38.chrom.sizes -n test-run -v

(2) On several local files:

bam2bw my1.bam my2.bam my3.bam -s hg38.chrom.sizes -n test-run -v

(3) On a remote file:

bam2bw https://path/to/my.bam -s hg38.chrom.sizes -n test-run -v

(4) On several remote files:

bam2bw https://path/to/my1.bam https://path/to/my2.bam https://path/to/my3.bam -s hg38.chrom.sizes -n test-run -v

Each will return two bigWig files: test-run.+.bw and test-run.-.bw. When multiple files are passed in their reads are concatenated without the need to produce an intermediary file of concatenated reads.

(5) When wanting a single unstranded bigWig:

bam2bw my.bam -s hg38.chrom.sizes -n test-run -v -u

(6) When wanting to map fragments (and get a single unstranded bigWig):

bam2bw fragments.tsv.gz -s hg38.chrom.sizes -n test-run -v -f -u

(7) With a FASTA instead of a .chrom.sizes:

bam2bw my.bam -s hg38.fa -n test-run -v

(8) When wanting to normalize by read depth such that the sum across both bigWigs is equal to 1.

bam2bw my.bam -s hg38.chrom.sizes -n test-run -v -r

(9) When wanting to normalize by read depth such that the sum across both bigWigs is equal to 1,000,000.

bam2bw my.bam -s hg38.chrom.sizes -n test-run -v -r -sf 1000000

(10) When wanting to record 3' ends instead of 5' ends:

bam2bw my.bam -s hg38.chrom.sizes -n test-run -v -3p

(11) When a BAM has paired-end reads that jointly represent a single RNA/fragment tag (e.g. PRO-seq/PRO-cap), rather than two independent events (e.g. the two Tn5 cut sites of an ATAC-seq fragment): use --mate_pairs so each pair contributes exactly one position instead of one from each mate. --rna5 picks which mate carries the RNA's 5' end (the other mate's own 5' end is used as the RNA's 3' end), and -3p/--opposite_strand behave as before but are applied to the jointly-determined position/strand:

bam2bw my.bam -s hg38.chrom.sizes -n test-run -v -mp --rna5 read2 -3p

A note on paired-end BAMs

By default, bam2bw (like bedtools genomecov) counts every mapped alignment record independently, including both mates of a pair. For ATAC-seq/DNase-seq/ChIP-seq this is correct: each mate's end is its own real cut/fragment-boundary event. For assays where a fragment carries exactly one meaningful tag position determined jointly by both mates (PRO-seq, PRO-cap, and similar run-on/CAGE-style protocols), counting both mates independently will roughly double the signal and scatter it across the wrong positions (each mate maps to a different point in the fragment). Use --mate_pairs (see example 11) for that case instead.

Existing Pipeline

This tool is meant specifically to replace the following pipeline which produces several large intermediary files:

wget https://path/to/my.bam -O my.bam
samtools sort my.bam -o my.sorted.bam

bedtools genomecov -5 -bg -strand + -ibam my.sorted.bam | sort -k1,1 -k2,2n > my.+.bedGraph
bedtools genomecov -5 -bg -strand - -ibam my.sorted.bam | sort -k1,1 -k2,2n > my.-.bedGraph

bedGraphToBigWig my.+.bedGraph hg38.chrom.sizes my.+.bw
bedGraphToBigWig my.-.bedGraph hg38.chrom.sizes my.-.bw

Testing

The test suite runs bam2bw end-to-end on small synthetic BAM, BED, and tsv files built on the fly, and checks the values inside the bigWigs it writes.

pip install -e .[test]
pytest

or, with uv,

uv sync --extra test
uv run pytest

The same suite runs on GitHub Actions on every push and pull request to main, across Python 3.10 through 3.13.

Version Log

v0.5.1
======

  Fixes
  - Blank lines and # comments in a BED/tsv input file are skipped instead of raising. Only the chrom_sizes reader skipped them, so a 10x CellRanger fragments file -- which opens with # id=, # description= and # pipeline_version= -- failed with "expected at least three columns" after peak callers and other tools had read the same file without complaint. A comment carrying three or more whitespace-separated fields was worse than the error: "# a note" has three, so it was read as an interval on a chromosome named "#" and silently discarded rather than reported.

  Packaging
  - v0.5.0 was tagged and released on GitHub but never published to PyPI. The last release on PyPI was v0.4.1, so this release supersedes v0.4.2, v0.4.3 and v0.5.0 there.
  - `project.license` is now the SPDX string "MIT" with `license-files`, rather than a TOML table. setuptools deprecated the table form, which warned on every build and stops being supported on 2027-Feb-18. The build requirement rises from setuptools>=64 to >=77, the first version to accept the new fields.
  - MANIFEST.in prunes tests/__pycache__ instead of listing __pycache__ in a global-exclude. global-exclude matches files, and __pycache__ is a directory, so the directive never matched and any non-.pyc file inside a __pycache__ directory was shipped in the sdist.

v0.5.0
======

  New
  - Added -3p/--three_prime to record the 3' end of each read or interval instead of the 5' end.
  - Added -mp/--mate_pairs, --rna5, and --opposite_strand to jointly count paired-end
    reads as a single RNA/fragment tag (e.g. for PRO-seq/PRO-cap) instead of counting
    each mate independently.

  Fixes
  - .sam input is now actually read. It was listed as an accepted format but no branch handled it, so a SAM file exited 0 after writing bigWigs with no entries, and a nonexistent .sam path did the same.
  - Entries falling outside the chromosome sizes given to -s are now discarded explicitly, with a count reported, instead of being dropped by pyBigWig without notice.
  - -r now normalizes over the entries actually written rather than over every counted read, so a sizes file that disagrees with the input no longer yields a track summing to less than the scale factor.
  - -r with -v reports the read depth that was divided through.
  - A samtools .fai index, or any sizes file with more than two columns, can now be passed to -s; only the first two fields of each line are read.
  - Blank lines and # comments in a chrom_sizes file are skipped instead of raising.
  - Malformed BED/tsv lines and mapped reads with no CIGAR now report the file and the offending line or read rather than raising a bare unpacking error.
  - A gzip-compressed FASTA passed to -s now says that BGZF is required and names the file. Only BGZF has ever worked; the v0.4.2 note about .gz variants refers to BGZF.

  Packaging
  - Packaging moved from setup.py to pyproject.toml. `pip install bam2bw` and `pip install -e .[test]` are unchanged.
  - Added an end-to-end test suite, run on GitHub Actions against Python 3.10 through 3.13 on every push and pull request.
  - v0.4.2 and v0.4.3 were version-logged but never published to PyPI, so this release supersedes both. The last release on PyPI was v0.4.1.

v0.4.2
======

  - Recognize additional FASTA extensions (.fasta, .fna, .fas) and their .gz variants when passed to -s.
  - Fixed tqdm progress bars clashing when processing input files in parallel.
  - Added biopython as an explicit dependency.

v0.4.1
======

  - Oops removed print statement.

v0.4.0
======

  - Added in a -p/--parallel option to read from input files in parallel using joblib. Max 1 process per input file
  - Moved the opening of bigWig objects until AFTER the input files are completely read to avoid deleting everything immediately if you accidentally re-run the command
  - Allow reading from .bed and .bed.gz files.
  - Create one progress bar per input file that can update in parallel with the description being the file name.


v0.3.2
======

  - Added -sf which is a scale factor that multiplies each position.
  - Added -r which performs read-depth normalization, dividing each position by the (pre-scaled) sum.
  - You can use -sf and -r together to make the bigWigs sum to your desired value.


v0.3.1
======

  - Fixed a minor bug

v0.3.0
======

  - A FASTA can be passed in to -s instead of a chrom_sizes file, and the chromosomes and their sizes automatically extracted
  - Entries not mapping to a chromosome in the chrom_sizes/FASTA file will be ignored and a warning will be raised if -v is set
  - .tsv and .tsv.gz files can processed now using the same arguments.
  - The -f argument has been added which treats entires as fragments where both the 3' and 5' ends should be added instead of just the 5' one

Metadata

Release files for bam2bw 0.5.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for bam2bw 0.5.1
File Size Uploaded
bam2bw-0.5.1.tar.gz 34.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for bam2bw 0.5.1
File Interpreter ABI Platform
bam2bw-0.5.1-py3-none-any.whl Python 3 none any Details

Total release size: 47.9 kB

Release files / bam2bw-0.5.1.tar.gz

Download URL bam2bw-0.5.1.tar.gz
Size 34.2 kB
Tags Source
SHA-256 checksum
How to use checksums
fc0e4ecf085dd740865d85a3a5593e346d51b85e74b7746bfab7087449cdf021
BLAKE2b-256 checksum
How to use checksums
1980f7a6b4e5c5b895bb5683feaefae9e16d430439514fbdc654d323208d0092
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.5

Release files / bam2bw-0.5.1-py3-none-any.whl

Download URL bam2bw-0.5.1-py3-none-any.whl
Size 13.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
de62943ed0ae72cd044d23748c1afbf411683718768af7345d39ddaf5d6fb38b
BLAKE2b-256 checksum
How to use checksums
abc231eab54fdc237ec7968a1e5ed50aac7a417d8bc2b7194b4850c8da261013
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.5

Release history Release notifications | RSS feed

This release

0.5.1 This release

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.2

2 release files

0.3.1

2 release files

0.3.0

2 release files

0.2.2

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page