Neural network sequence error correction.

Project description

Oxford Nanopore Technologies logo

Medaka

medaka is a tool to create a consensus sequence from nanopore sequencing data. This task is performed using neural networks applied from a pileup of individual sequencing reads against a draft assembly. It outperforms graph-based methods operating on basecalled data, and can be competitive with state-of-the-art signal-based methods, whilst being much faster.

Features

Requires only basecalled data. (.fasta or .fastq)
Improved accurary over graph-based methods (e.g. Racon).
50X faster than Nanopolish (and can run on GPUs).
Methylation aggregation from Guppy .fast5 files.
Benchmarks are provided here.
Includes extras for implementing and training bespoke correction networks.
Works on Linux and MacOS.
Open source (Mozilla Public License 2.0).

Tools to enable the creation of draft assemblies can be found in a sister project pomoxis.

Documentation can be found at https://nanoporetech.github.io/medaka/.

Installation

Medaka can be installed in one of several ways.

Installation with conda

Perhaps the simplest way to start using medaka on both Linux and MacOS is through conda; medaka is available via the bioconda channel:

conda create -n medaka -c conda-forge -c bioconda medaka

Installation with pip

For those who prefer python's native pacakage manager, medaka is also available on pypi and can be installed using pip:

pip install medaka

On Linux platforms this will install a precompiled binary, on MacOS (and other) platforms this will fetch and compile a source distribution.

We recommend using medaka within a virtual environment, viz.:

virtualenv medaka --python=python3 --prompt "(medaka) "
. medaka/bin/activate
pip install medaka

Using this method requires the user to provide several binaries:

and place these within the PATH. samtools/bgzip/tabix version 1.9 and minimap2 version 2.17 are recommended as these are those used in development of medaka.

Installation from source

Medaka can be installed from its source quite easily on most systems.

Before installing medaka it may be required to install some prerequisite libraries, best installed by a package manager. On Ubuntu theses are:

bzip2 g++ zlib1g-dev libbz2-dev liblzma-dev libffi-dev libncurses5-dev
libcurl4-gnutls-dev libssl-dev curl make cmake wget python3-all-dev
python-virtualenv

In addition it is required to install and set up git LFS before cloning the repository.

A Makefile is provided to fetch, compile and install all direct dependencies into a python virtual environment. To set-up the environment run:

# Note: certain files are stored in git-lfs, https://git-lfs.github.com/,
#       which must therefore be installed first.
git clone https://github.com/nanoporetech/medaka.git
cd medaka
make install
. ./venv/bin/activate

Using this method both samtools and minimap2 are built from source and need not be provided by the user.

Using a GPU

All installation methods will allow medaka to be used with CPU resource only. To enable the use of GPU resource it is necessary to install the tensorflow-gpu package. Unfortunately depending on your python version it may be necessary to modify the requirements of the medaka package for it to run without complaining. Using the source code from github a working GPU-powered medaka can be configured with:

# Note: certain files are stored in git-lfs, https://git-lfs.github.com/,
#       which must therefore be installed first.
git clone https://github.com/nanoporetech/medaka.git
cd medaka
sed -i 's/tensorflow/tensorflow-gpu/' requirements.txt
make install

However, note that The tensorflow-gpu GPU package is compiled against specific versions of the NVIDIA CUDA and cuDNN libraries; users are directed to the tensorflow installation pages for further information. cuDNN can be obtained from the cuDNN Archive, whilst CUDA from the CUDA Toolkit Archive.

Depending on your GPU, medaka may show out of memory errors when running. To avoid these the inference batch size can be reduced from the default value by setting the -b option when running medaka_consensus. A value -b 100 is suitable for 11Gb GPUs.

For users with RTX series GPUs it may be required to additionally set an environment variable to have medaka run without failure:

export TF_FORCE_GPU_ALLOW_GROWTH=true

In this situation a further reduction in batch size may be required.

Usage

medaka can be run using its default settings through the medaka_consensus program. An assembly in .fasta format and basecalls in .fasta or .fastq formats are required. The program uses both samtools and minimap2. If medaka has been installed using the from-source method these will be present within the medaka environment, otherwise they will need to be provided by the user.

source ${MEDAKA}  # i.e. medaka/venv/bin/activate
NPROC=$(nproc)
BASECALLS=basecalls.fa
DRAFT=draft_assm/assm_final.fa
OUTDIR=medaka_consensus
medaka_consensus -i ${BASECALLS} -d ${DRAFT} -o ${OUTDIR} -t ${NPROC} -m r941_min_high_g303

The variables BASECALLS, DRAFT, and OUTDIR in the above should be set appropriately. For the selection of the model (-m r941_min_high_g303 in the example above) see the Model section following.

When medaka_consensus has finished running, the consensus will be saved to ${OUTDIR}/consensus.fasta.

Models

For best results it is important to specify the correct model, -m in the above, according to the basecaller used. Allowed values can be found by running medaka tools list\_models.

Medaka models are named to indicate i) the pore type, ii) the sequencing device (MinION or PromethION), iii) the basecaller variant, and iv) the basecaller version, with the format:

{pore}_{device}_{caller variant}_{caller version}

For example the model named r941_min_fast_g303 should be used with data from MinION (or GridION) R9.4.1 flowcells using the fast Guppy basecaller version 3.0.3. By contrast the model r941_prom_hac_g303 should be used with PromethION data and the high accuracy basecaller (termed "hac" in Guppy configuration files). Where a version of Guppy has been used without an exactly corresponding medaka model, the medaka model with the highest version equal to or less than the guppy version should be selected.

Methylation Calling

medaka includes a basic workflow for aggregating Guppy basecalling results for Dcm, Dam, and CpG methylation. The workflow is currently very preliminary and subject to change and improvement.

Aggregating the information from Guppy outputs is a two stage process, first the basecalling results are extracted .fast5 files and placed in a .bam file:

FAST5PATH=guppy/workspace
REFERENCE=grch38.fasta
OUTBAM=meth.bam
medaka methylation guppy2sam ${FAST5PATH} ${REFERENCE} \
    --workers 74 --recursive \
    | samtools sort -@ 8 | samtools view -b -@ 8 > ${OUTBAM}
samtools sort ${OUTBAM}

This program will extract both the basecall sequence and methylation scores, align the basecall to the reference, and store results in a standard format. In this preliminary workflow the methylation scores are stored in two SAM tags, 'MC' and 'MA', one each for 5mC and 6mA respectively. The tags are 8bit integer array-values, one value per basecall position. This is a different form to that proposed in the current hts-specs proposition, but allows for more trivial parsing.

The second step is to aggregate the reference-aligned information to produce a simple tabular summary of read methylation counts:

BAM=meth.bam
REFERENCE=grch38.fasta
REGION=chr20:500000-1000000
OUTPUT=meth.tsv
medaka methylation call --meth cpg ${BAM} ${REFERENCE} ${REGION} ${OUTPUT}

Here the option --meth cpg indicates that loci containing the sequence motif CG should be examined for 5mC presence. Other choices are dcm for which the motifs CCAGG and CCTGG are examined for 5mC and dam (GATC) for 6mA.

The output file is a simple tab-delimited text file with columns: 'ref.name', 'position', 'motif', 'fwd.meth.count', 'rev.meth.count', 'fwd.canon.count', and 'rev.canon.count'. Here fwd./ref. indicate counts on the two DNA strands and meth./canon. indicate counts for methylated and canonical bases. Note that the position field records the position of the first base in the motif recorded.

Origin of the draft sequence

Medaka has been trained to correct draft sequences processed through racon, specifically racon run four times iteratively with:

racon -m 8 -x -6 -g -8 -w 500 ...

Processing a draft sequence from alternative sources (e.g. the output of canu or wtdbg2) may lead to different results.

The documentation provides a discussion and some guidance on how to obtain a draft sequence.

Acknowledgements

We thank Joanna Pineda and Jared Simpson for providing htslib code samples which aided greatly development of the optimised feature generation code, and for testing the version 0.4 release candidates.

We thank Devin Drown for working through use of medaka with his RTX 2080 GPU.

Help

Licence and Copyright

medaka is distributed under the terms of the Mozilla Public License 2.0.

Research Release

Research releases are provided as technology demonstrators to provide early access to features or stimulate Community development of tools. Support for this software will be minimal and is only provided directly by the developers. Feature requests, improvements, and discussions are welcome and can be implemented by forking and pull requests. However much as we would like to rectify every issue and piece of feedback users may have, the developers may have limited resource for support of this software. Research releases may be unstable and subject to rapid iteration by Oxford Nanopore Technologies.

Project details

Release history Release notifications | RSS feed

2.0.1

Oct 11, 2024

2.0.0

Sep 11, 2024

2.0.0a2 pre-release

Aug 20, 2024

2.0.0a1 pre-release

Jul 19, 2024

1.12.1

Jul 12, 2024

1.12.0

May 20, 2024

1.11.3

Dec 6, 2023

1.11.2

Nov 29, 2023

1.11.1

Oct 24, 2023

1.11.0

Oct 23, 2023

1.10.0

Oct 16, 2023

1.9.1

Aug 15, 2023

1.8.2

Aug 8, 2023

1.7.3

Feb 13, 2023

1.6.1

Jun 20, 2022

1.5.0

Dec 13, 2021

1.4.4

Sep 23, 2021

1.3.4

May 19, 2021

1.2.6

Apr 1, 2021

1.1.3

Oct 15, 2020

1.0.3

Jun 10, 2020

This version

1.0.0

May 4, 2020

0.12.1

Apr 3, 2020

0.11.5

Jan 24, 2020

0.10.1

Nov 11, 2019

0.9.2

Sep 27, 2019

0.8.1

Jul 29, 2019

0.7.1

May 20, 2019

0.6.5

May 1, 2019

0.5.2

Feb 7, 2019

0.4.3

Dec 21, 2018

0.4.1

Dec 14, 2018

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

medaka-1.0.0.tar.gz (39.0 MB view hashes)

Uploaded May 4, 2020 Source

Built Distributions

medaka-1.0.0-cp36-cp36m-manylinux1_x86_64.whl (40.7 MB view hashes)

Uploaded May 4, 2020 CPython 3.6m

medaka-1.0.0-cp35-cp35m-manylinux1_x86_64.whl (40.7 MB view hashes)

Uploaded May 4, 2020 CPython 3.5m

Hashes for medaka-1.0.0.tar.gz

Hashes for medaka-1.0.0.tar.gz
Algorithm	Hash digest
SHA256	`531f4770a298c690345996f10d54dc3c1eed7cef835cd155a2a02211b4f51b43`
MD5	`14cba4802e5608f7e5f4a6219020e337`
BLAKE2b-256	`4bf7a2970846fbd59ad860d7a6f9caabaf1df63079d90da94e6912cf0a989e55`

Hashes for medaka-1.0.0-cp36-cp36m-manylinux1_x86_64.whl

Hashes for medaka-1.0.0-cp36-cp36m-manylinux1_x86_64.whl
Algorithm	Hash digest
SHA256	`c06e8f4a534bafefe1b2568057fea791fd93e0fdcb6fd9403053476c0ca2b585`
MD5	`0949e6acf1d783faa53dbde82f50112a`
BLAKE2b-256	`f4020cda43f4585386c4f9d724623c4aef10ef715835130781cb6f87753e89e9`

Hashes for medaka-1.0.0-cp35-cp35m-manylinux1_x86_64.whl

Hashes for medaka-1.0.0-cp35-cp35m-manylinux1_x86_64.whl
Algorithm	Hash digest
SHA256	`e45c626ae2db4d822695e722a3ffd586906f38ec739e7943a69a1c54f7af7573`
MD5	`d6a1f624f8777179fee4b5dd65248983`
BLAKE2b-256	`c1ba5b289ec8d06f4549a36cd8e869fb0de4d9a3dedebe026aced7dbfda5773d`