Skip to main content

cuteSV

PyPI version Anaconda-Server Badge Anaconda-Server Badge Anaconda-Server Badge Anaconda-Server Badge


Getting Start

                                               __________    __       __
                                              |   ____   |  |  |     |  |
                          _                   |  |    |__|  |  |     |  |
 _______    _     _   ___| |___     ______    |  |          |  |     |  |
|  ___  |  | |   | | |___   ___|   / ____ \   |  |_______   |  |     |  |
| |   |_|  | |   | |     | |      / /____\ \  |_______   |  |  |     |  |
| |        | |   | |     | |      | _______|   __     |  |  \  \     /  /
| |    _   | |   | |     | |  _   | |     _   |  |    |  |   \  \   /  /
| |___| |  | |___| |     | |_| |  \ \____/ |  |  |____|  |    \  \_/  /
|_______|  |_______|     |_____|   \______/   |__________|     \_____/

Installation

$ pip install cuteSV
or
$ conda install -c bioconda cutesv
or 
$ git clone https://github.com/tjiangHIT/cuteSV.git && cd cuteSV/ && python setup.py install 

Introduction

Long-read sequencing enables the comprehensive discovery of structural variations (SVs). However, it is still non-trivial to achieve high sensitivity and performance simultaneously due to the complex SV characteristics implied by noisy long reads. Therefore, we propose cuteSV, a sensitive, fast and scalable long-read-based SV detection approach. cuteSV uses tailored methods to collect the signatures of various types of SVs and employs a clustering-and-refinement method to analyze the signatures to implement sensitive SV detection. Benchmarks on real Pacific Biosciences (PacBio) and Oxford Nanopore Technology (ONT) datasets demonstrate that cuteSV has better yields and scalability than state-of-the-art tools.

The benchmark results of cuteSV on the HG002 human sample are below:

BTW, we used Truvari to calculate the recall, precision, and f-measure. For more detailed implementation of SV benchmarks, we show an example here.


Dependence

1. python3
2. pysam
3. Biopython
4. cigar
5. numpy

Usage

cuteSV <sorted.bam> <output.vcf> <work_dir>

Suggestions

> For PacBio CLR/ONT data:
	--max_cluster_bias_INS		100
	--diff_ratio_merging_INS	0.2
	--diff_ratio_filtering_INS	0.6
	--diff_ratio_filtering_DEL	0.7
> For PacBio CCS(HIFI) data:
	--max_cluster_bias_INS		200
	--diff_ratio_merging_INS	0.65
	--diff_ratio_filtering_INS	0.65
	--diff_ratio_filtering_DEL	0.35
Parameter Description Default
--threads Number of threads to use. 16
--batches Batch of genome segmentation interval. 10,000,000
--sample Sample name/id NULL
--retain_work_dir Enable to retain temporary folder and files. False
--max_split_parts Maximum number of split segments a read may be aligned before it is ignored. 7
--min_mapq Minimum mapping quality value of alignment to be taken into account. 20
--min_read_len Ignores reads that only report alignments with not longer than bp. 500
--merge_del_threshold Maximum distance of deletion signals to be merged. 0
--merge_ins_threshold Maximum distance of insertion signals to be merged. 100
--min_support Minimum number of reads that support a SV to be reported. 10
--min_size Minimum length of SV to be reported. 30
--max_size Minimum length of SV to be reported. 100000
--genotype Enable to generate genotypes. False
--gt_round Maximum round of iteration for alignments searching if perform genotyping. 500
--max_cluster_bias_INS Maximum distance to cluster read together for insertion. 100
--diff_ratio_merging_INS Do not merge breakpoints with basepair identity more than the ratio of default for insertion. 0.2
--diff_ratio_filtering_INS Filter breakpoints with basepair identity less than the ratio of default for insertion. 0.6
--max_cluster_bias_DEL Maximum distance to cluster read together for deletion. 200
--diff_ratio_merging_DEL Do not merge breakpoints with basepair identity more than the ratio of default for deletion. 0.3
--diff_ratio_filtering_DEL Filter breakpoints with basepair identity less than the ratio of default for deletion. 0.7
--max_cluster_bias_INV Maximum distance to cluster read together for inversion. 500
--max_cluster_bias_DUP Maximum distance to cluster read together for duplication. 500
--max_cluster_bias_TRA Maximum distance to cluster read together for translocation. 50
--diff_ratio_filtering_TRA Filter breakpoints with basepair identity less than the ratio of default for translocation. 0.6

Datasets generated from cuteSV

We provided the SV callsets of the HG002 human sample produced by cuteSV form three different long-read sequencing platforms (i.e. PacBio CLR, PacBio CCS, and ONT PromethION).

You can download them at: DOI

Please cite the manuscript of cuteSV before using these callsets.


Changelog

cuteSV (v1.0.7):
1. Add read name list for each SV call.
2. Fix several descriptions in VCF header field.

cuteSV (v1.0.6):
1.Improvement of genotyping by calculation of likelihood.
2.Add variant quality value, phred-scaled genotype likelihood and genotype quality in order to filter false positive SV or quality control.
3.Add --gt_round parameter to control the number of read scans.
4.Add variant strand of DEL/DUP/INV.
5.Fix several bugs.

cuteSV (v1.0.5):
1.Add new options for specificly setting the threshold of deletion/insertion signals merging in the same read. The default parameters are 0 bp for deletion and 100 bp for insertion.
2.Remove parameter --merge_threshold.
3.Fix bugs in inversion and translocation calling.
4.Add new option for specificly setting the maximum size of SV to be discovered. The default value is 100,000 bp. 


cuteSV (v1.0.4):
1.Add a new option for specificly setting the threshold of SV signals merging in the same read. The default parameter is 500 bp. You can reduce it for high-quality sequencing datasets like PacBio HiFi (CCS).
2.Make the genotyping function optional.
3.Enable users to set the threshold of SV allele frequency of homozygous/heterozygous.
4.Update the description of recommendation parameters in processing ONT data.

cuteSV (v1.0.3):
1.Refine the genotyping model.
2.Adjust the threshold value of heterozygosis alleles.

cuteSV (v1.0.2):
1.Improve the genotyping performance and enable it to be default option.
2.Make the description of parameters better.
3.Modify the header description of vcf file.
4.Add two new indicators, i.e., BREAKPOINT_STD and SVLEN_STD, to further characterise deletion and insertion.
5.Remove a few redundant functions which will reduce code readability.

Citation

Jiang T et al. Long-read-based human genomic structural variation detection with cuteSV. Genome Biol 21, 189 (2020). https://doi.org/10.1186/s13059-020-02107-y


Contact

For advising, bug reporting and requiring help, please post on Github Issue or contact tjiang@hit.edu.cn.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

cuteSV-1.0.7.tar.gz (32.3 kB view details)

Uploaded Source

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

cuteSV-1.0.7-py3.6.egg (75.1 kB view details)

Uploaded Egg

cuteSV-1.0.7-py3-none-any.whl (38.9 kB view details)

Uploaded Python 3

File details

Details for the file cuteSV-1.0.7.tar.gz.

File metadata

  • Download URL: cuteSV-1.0.7.tar.gz
  • Upload date:
  • Size: 32.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/3.2.0 pkginfo/1.5.0.1 requests/2.22.0 setuptools/49.3.1 requests-toolbelt/0.9.1 tqdm/4.36.1 CPython/3.6.7

File hashes

Hashes for cuteSV-1.0.7.tar.gz
Algorithm Hash digest
SHA256 05b88c5a61ae98615974f45f1d384d3ffb365d30a297388c1c67e7aa13f2660b
MD5 b25d5d4a71cc47757d967d0943c3410b
BLAKE2b-256 96ee9190a65eadbd18c5bb9149fb1e42ae66166bed0c819b481e3e52677533a1

See more details on using hashes here.

File details

Details for the file cuteSV-1.0.7-py3.6.egg.

File metadata

  • Download URL: cuteSV-1.0.7-py3.6.egg
  • Upload date:
  • Size: 75.1 kB
  • Tags: Egg
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/3.2.0 pkginfo/1.5.0.1 requests/2.22.0 setuptools/49.3.1 requests-toolbelt/0.9.1 tqdm/4.36.1 CPython/3.6.7

File hashes

Hashes for cuteSV-1.0.7-py3.6.egg
Algorithm Hash digest
SHA256 58c19bf85f5d9d56297a7ade74e0208052f8dceee813e2d2e10e1eb1ae8bb961
MD5 f8ad52451c4cfe191bafea6d7d405359
BLAKE2b-256 146c39467f8a82f008a95c670270ecb2395740e459587fd61afbbd8f6107c1f9

See more details on using hashes here.

File details

Details for the file cuteSV-1.0.7-py3-none-any.whl.

File metadata

  • Download URL: cuteSV-1.0.7-py3-none-any.whl
  • Upload date:
  • Size: 38.9 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/3.2.0 pkginfo/1.5.0.1 requests/2.22.0 setuptools/49.3.1 requests-toolbelt/0.9.1 tqdm/4.36.1 CPython/3.6.7

File hashes

Hashes for cuteSV-1.0.7-py3-none-any.whl
Algorithm Hash digest
SHA256 916804e7484daddd792ec5de146a6d9193bdec49f60bd8c20a5959d4bb461759
MD5 616faf1370dc62e72884036932f991d2
BLAKE2b-256 3f311cd5a63b17fb6ff227f0aa7e853d2ae16b1088e55abd97ced9a853c06873

See more details on using hashes here.

Release history Release notifications | RSS feed

2.1.4

1 file

2.1.3

1 file

2.1.2

1 file

2.1.1

1 file

2.1.0

1 file

2.0.3

1 file

2.0.2

2 files

2.0.1

1 file

2.0.0

2 files

1.0.13

3 files

1.0.12

2 files

1.0.11

3 files

1.0.10

3 files

1.0.9

3 files

This release

1.0.7 This release

3 files

1.0.6

2 files

1.0.5

2 files

1.0.4

2 files

1.0.3

2 files

1.0.2.1

1 file

1.0.1

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page