hundo

Amplicon processing protocol

Project description

Performs quality control based on quality, can trim adapters, and remove sequences matching a contaminant database
Handles paired-end read merging
Integrates de novo and reference-based chimera filtering
Clusters sequences and annotates using databases that are downloaded as needed
Generates standard outputs for these data like a newick tree, a tabular OTU table with taxonomy, and .biom.

This workflow is built using Snakemake and makes use of Bioconda to install its dependencies.

Documentation

For complete documentation and install instructions, see:

https://hundo.readthedocs.io

Install

This protocol leverages the work of Bioconda and depends on conda. For complete setup of these, please see:

https://bioconda.github.io/#using-bioconda

Really, you just need to make sure conda is executable and you’ve set up your channels (numbers 1 and 2). Then:

conda install python>=3.6 click \
    pyyaml snakemake>=5.1.4 biopython
pip install hundo

Usage

Running samples through annotation requires that input FASTQs be paired-end, named in a semi-conventional style starting sample ID, contain “_R1” (or “_r1”) and “_R2” (or “_r2”) index identifiers, and have an extension “.fastq” or “.fq”. The files may be gzipped and end with “.gz”. By default, both R1 and R2 need to be larger than 10K in size. This cutoff is arbitrary and can be set using --prefilter-file-size.

Using the example data of the mothur SOP located in our tests directory, we can annotate across SILVA using:

cd example
hundo annotate \
    --filter-adapters qc_references/adapters.fa.gz \
    --filter-contaminants qc_references/phix174.fa.gz \
    --out-dir mothur_sop_silva \
    --database-dir annotation_references \
    --reference-database silva \
    mothur_sop_data

Dependencies are installed by default in the results directory defined on the command line as --out-dir. If you want to re-use dependencies across many analyses and not have to re-install each time you update the output directory, use Snakemake’s --conda-prefix:

hundo annotate \
    --out-dir mothur_sop_silva \
    --database-dir annotation_references \
    --reference-database silva \
    mothur_sop_data \
    --conda-prefix /Users/brow015/devel/hundo/example/conda

Output

OTU.biom

Biom table with raw counts per sample and their associated taxonomic assignment formatted to be compatible with downstream tools like phyloseq.

OTU.fasta

Representative DNA sequences of each OTU.

OTU.tree

Newick tree representation of aligned OTU sequences.

OTU.txt

Tab-delimited text table with columns OTU ID, a column for each sample, and taxonomy assignment in the final column as a comma delimited list.

OTU_aligned.fasta

OTU sequences after alignment using Clustal Omega.

all-sequences.fasta

Quality-controlled, dereplicated DNA sequences of all samples. The header of each record identifies the sample of origin and the count resulting from dereplication.

blast-hits.txt

The BLAST assignments per OTU sequence.

summary.html

Captures and summarizes data of the experimental dataset. Things like sequence quality, counts per sample at varying stages of pre-processing, and summarized taxonomic composition per sample across phylum, class, and order.

Project details

Release history Release notifications | RSS feed

1.2.8

Aug 5, 2019

1.2.7

Aug 5, 2019

1.2.6

Jul 8, 2019

1.2.5

Mar 13, 2019

1.2.4

Nov 1, 2018

1.2.3

Oct 31, 2018

1.2.2

Oct 31, 2018

1.2.1

Sep 21, 2018

This version

1.2.0

Sep 21, 2018

1.1.21

Aug 14, 2018

1.1.20

Jul 20, 2018

1.1.19

Jul 20, 2018

1.1.18

Jun 25, 2018

1.1.17

Jun 25, 2018

1.1.16

Jun 22, 2018

1.1.15

Jun 18, 2018

1.1.14

Jun 16, 2018

1.1.13

Jun 6, 2018

1.1.12

Jun 5, 2018

1.1.11

Jun 5, 2018

1.1.10

Jun 1, 2018

1.1.9

May 24, 2018

1.1.8

May 15, 2018

1.1.7

May 7, 2018

1.1.6

Mar 2, 2018

1.1.5

Jan 8, 2018

1.1.4

Dec 6, 2017

1.1.3

Nov 8, 2017

1.1.2

Oct 17, 2017

1.1.1

Oct 16, 2017

1.1.0

Oct 14, 2017

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distributions

No source distribution files available for this release.See tutorial on generating distribution archives.

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

The dropdown lists show the available interpreters, ABIs, and platforms. Enable javascript to be able to filter the list of wheel files.

hundo-1.2.0-py3.6.egg (55.4 kB view details)

Uploaded Sep 21, 2018 Egg

File details

Details for the file hundo-1.2.0-py3.6.egg.

File metadata

Download URL: hundo-1.2.0-py3.6.egg
Upload date: Sep 21, 2018
Size: 55.4 kB
Tags: Egg
Uploaded using Trusted Publishing? No
Uploaded via: twine/1.10.0 pkginfo/1.4.1 requests/2.18.4 setuptools/38.4.0 requests-toolbelt/0.8.0 tqdm/4.23.1 CPython/3.6.4

File hashes

Hashes for hundo-1.2.0-py3.6.egg
Algorithm	Hash digest
SHA256	`d8e38dc3c4574c45899d0c32a1e44b63053969d22c3fbba8e5ba73213e26ddf8`
MD5	`3d6ba3ee64d72547cdbd8f116764b6c8`
BLAKE2b-256	`24e0b23d75ca9f6a0009717fdfa4b5fe8e100a771e9d13f01664d7ca4e34ca18`

See more details on using hashes here.

hundo 1.2.0

Navigation

Verified details

Maintainers

Unverified details

Project links

Meta