An extension of breseq to determine copy number variations from coverage data
Project description
CNery
breseq copy-number-variation extension. CNery reads the sequencing coverage output from breseq and predicts copy-number variation (CNV) across the genome. Predictions are corrected for coverage biases introduced by sequencing chemistry (GC-content bias) and prokaryotic replication state during DNA isolation (origin-to-terminus / OTR bias).
Recent updates (latest commits):
- Multi-genome CNV analysis —
CNerynow processes all reference sequences found in the breseq BAM/FASTA in one pass. Each reference (chromosome, plasmid, contig, etc.) is preprocessed per genome, pooled for a shared LOWESS GC-bias fit, and then bias-corrected and CN-called independently. - Input/output flexibility — inputs default to
<input>/data/reference.bamand<input>/data/reference.fasta; output prefix defaults to<input>/CNV_out/. Output subfolders (CNV_plt/,CNV_csv/,GC_bias/,OTR_corr/) are created automatically. - Modular bias correction — the
--biasflag lets you chooseall(GC + OTR),gc,otr, ornone. - Pip-installable package —
requirements.txtand a fixedpyproject.tomlallow install directly from GitHub viapip install git+....
Installation
Recommended: create a conda/mamba environment from the provided spec.
mamba env create -f environment.yml
mamba activate CNery
Install CNery (a.k.a. breseq-ext-cnv) from GitHub:
pip install git+https://github.com/barricklab/breseq-ext-cnv.git
Quick start
Run CNery inside a breseq output folder that contains the data/ and output/ subfolders:
CNery [-o <output folder>] [-w <window>] [-s <step size>] [-f <fragment length>]
To run from a different working directory, point -i at the breseq output folder (or supply -ref and the BAM path manually):
CNery -i <breseq output folder> \
-ref <reference.fasta> \
-o <output folder> \
-w <window> \
-s <step size> \
-f <fragment length>
Usage examples
Calculate coverage with a 500 bp window sliding in 250 bp steps; sequencing fragment length is 300 bp:
CNery -o <output folder> -w 500 -s 250 -f 300
Analyze coverage across the whole genome, but restrict CNV/coverage plots to a specific genomic segment:
CNery -o <output folder> --region 3497890-3955678 -w 1000 -s 500
The --region argument accepts open intervals too (-reg 3497890- from a start to end of genome, -reg -3955678 from start of genome to an end position).
Control which bias correction is applied before CN prediction:
# Both GC + OTR corrections (default)
CNery -o <output folder> -w 500 -s 250 --bias all
# Only correct OTR (replication) bias
CNery -o <output folder> -w 500 -s 250 --bias otr
# Only correct GC-content bias
CNery -o <output folder> -w 500 -s 250 --bias gc
# No bias correction before CN prediction
CNery -o <output folder> -w 500 -s 250 --bias none
When OTR correction is applied, the origin and terminus of replication are automatically inferred from the coverage profile — no manual coordinates are required.
Outputs
Given an output folder CNV_out/, CNery writes:
CNV_out/CNV_plt/— per-reference CNV prediction plots.CNV_out/CNV_csv/— per-window coverage + CN calls as CSV.CNV_out/GC_bias/— pooled LOWESS GC-bias diagnostic plot.CNV_out/OTR_corr/— per-reference OTR bias plots and a JSON summary (*_otr_results.json) containing the inferred origin window, terminus window, normalized coverage at each, and the origin-to-terminus ratio.
Each reference sequence in the BAM/FASTA produces its own set of outputs, named with the reference / genome identifier.
All command-line options
$ CNery -h
usage: CNery [-h] [-i I] [-ref REF] [-reg REG] [-o O] [-w W] [-s S] [-f F] [-e E]
[--bias {all,none,gc,otr}]
CNery is a Python package extension to breseq that analyzes the sequencing
coverage across the genome to predict copy number variation (CNV).
options:
-h, --help show this help message and exit
-i, --input I input folder path (the breseq output folder with
'data' and 'output' folders). Defaults to the current
folder.
-ref REF select the reference file used for breseq. Defaults
to data/reference.fasta.
-reg REG select the region of the genome to evaluate
(format: START-END, e.g. 1000-50000).
-o, --output O output file prefix / storage location. Defaults to
the 'CNV_out' folder in the current dir.
-w, --window W Window length used to parse the genome and compute
coverage and GC statistics. Default: 200.
-s, --step-size S Step size (<= window size) for each progression of
the window across the genome. Set step-size = window
size for non-overlapping windows. Default: 100.
-f, --frag_size F Average fragment size of the sequencing reads.
Default: 500.
-e, --error-rate E Approximate error rate in sequencing read coverage /
reference alignment. Default: 0.05.
--bias {all,none,gc,otr}
Select which bias correction to apply before CN
prediction. 'all' applies GC + OTR, 'gc' or 'otr'
applies only that one, 'none' skips bias correction.
Default: all.
Run this script in the breseq output folder that contains 'data' and 'output'
folders.
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file cnery-1.1.0.tar.gz.
File metadata
- Download URL: cnery-1.1.0.tar.gz
- Upload date:
- Size: 24.7 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.8
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
b1f51b43fc0f952929a29be3da65d929b9b5e7946a7e667a8057eb968552ec47
|
|
| MD5 |
755c5df3d3cf2d9792118cb713058c0c
|
|
| BLAKE2b-256 |
893cc4bb852b61ead23b76d115d92a713f6178343340c0ab9ee93795aa01af72
|
File details
Details for the file cnery-1.1.0-py3-none-any.whl.
File metadata
- Download URL: cnery-1.1.0-py3-none-any.whl
- Upload date:
- Size: 23.0 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.8
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
26e55d56d559166d44eccb2f680aa6a319bf5cf76dd9660c105666422b26f143
|
|
| MD5 |
5ad216ef9c8dbec76fe1647cc7d5011e
|
|
| BLAKE2b-256 |
ad142e502f48faaa70a52d86c3776c8da570698fe60897643475a3c00bd58bdc
|