Skip to main content

Metagenomics toolkit

Project description

Build Status PyPi package Downloads

Metagenomics toolkit enables scientists to download all of the sample metadata for a given study or sequence to a single csv file.

Install metagenomics toolkit

pip install -U mg-toolkit

Usage

$ mg-toolkit -h
usage: mg-toolkit [-h] [-V] [-d]
                  {original_metadata,sequence_search,bulk_download} ...

Metagenomics toolkit
--------------------

positional arguments:
  {original_metadata,sequence_search,bulk_download}
    original_metadata   Download original metadata.
    sequence_search     Search non-redundant protein database using HMMER
    bulk_download       Download result files in bulks for an entire study.

optional arguments:
  -h, --help            show this help message and exit
  -V, --version         print version information
  -d, --debug           print debugging information

Examples

Download metadata:

$ mg-toolkit original_metadata -a ERP001736

Search non-redundant protein database using HMMER and fetch metadata:

$ mg-toolkit sequence_search -seq test.fasta -db full evalue -incE 0.02

Databases:
- full - Full length sequences (default)
- all - All sequences
- partial - Partial sequences

How to bulk download result files for an entire study?

$ mg-toolkit bulk_download -h
usage: mg-toolkit bulk_download [-h] -a ACCESSION [-o OUTPUT_PATH]
                                  [-p {1.0,2.0,3.0,4.0,4.1}]
                                  [-g {sequence_data,functional_analysis,taxonomic_analysis,taxonomic_analysis_ssu,taxonomic_analysis_lsu,stats,non_coding_rna}]

optional arguments:
  -h, --help            show this help message and exit
  -a ACCESSION, --accession ACCESSION
                        Provide the study/project accession of your interest,
                        e.g. ERP001736, SRP000319. The study must be publicly
                        available in MGnify.
  -o OUTPUT_PATH, --output_path OUTPUT_PATH
                        Location of the output directory, where the
                        downloadable files are written to. DEFAULT: CWD
  -p {1.0,2.0,3.0,4.0,4.1}, --pipeline {1.0,2.0,3.0,4.0,4.1}
                        Specify the version of the pipeline you are interested
                        in. Lets say your study of interest has been analysed
                        with multiple version, but you are only interested in
                        a particular version then used this option to filter
                        down the results by the version you interested in.
                        DEFAULT: Downloads all versions
  -g {sequence_data,functional_annotations,taxonomic_annotations,taxonomic_annot_ssu,taxonomic_annot_lsu,stats,non_coding_rna}, --result_group {sequence_data,functional_annotations,taxonomic_annotations,taxonomic_annot_ssu,taxonomic_annot_lsu,stats,non_coding_rna}
                        Provide a single result group if needed. Supported
                        result groups are: [sequence_data (all version),
                        functional_annotations (all version),
                        taxonomic_annotations (1.0-3.0), taxonomic_annot_ssu
                        (>=4.0), taxonomic_annot_lsu (>=4.0), stats,
                        non_coding_rna (>=4.0) DEFAULT: Downloads all result
                        groups if not provided. (default: None).

How to download all files for a given study accession?

$ mg-toolkit -d bulk_download -a ERP009703

How to download results of a specific version for given study accession?

$ mg-toolkit -d bulk_download -a ERP009703 -v 4.0

How to download specific result file groups (e.g. functional annotations only) for given study accession?

$ mg-toolkit -d bulk_download -a ERP009703 -g functional_annotations

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

mg-toolkit-0.6.5.tar.gz (15.4 kB view details)

Uploaded Source

Built Distribution

mg_toolkit-0.6.5-py3-none-any.whl (17.1 kB view details)

Uploaded Python 3

File details

Details for the file mg-toolkit-0.6.5.tar.gz.

File metadata

  • Download URL: mg-toolkit-0.6.5.tar.gz
  • Upload date:
  • Size: 15.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/3.1.1 pkginfo/1.5.0.1 requests/2.23.0 setuptools/46.1.3.post20200325 requests-toolbelt/0.9.1 tqdm/4.45.0 CPython/3.8.2

File hashes

Hashes for mg-toolkit-0.6.5.tar.gz
Algorithm Hash digest
SHA256 fca893fcea80a886fb259a32005c3d60fcd9a1a41cdd21d7bd2367a9c7f34a21
MD5 f4203bd0378c804d81f2cf6b7f7e8244
BLAKE2b-256 92eeb65c088473aa0802399979000a6f53d4c4fb5d773fda2300cd2b114cea4f

See more details on using hashes here.

File details

Details for the file mg_toolkit-0.6.5-py3-none-any.whl.

File metadata

  • Download URL: mg_toolkit-0.6.5-py3-none-any.whl
  • Upload date:
  • Size: 17.1 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/3.1.1 pkginfo/1.5.0.1 requests/2.23.0 setuptools/46.1.3.post20200325 requests-toolbelt/0.9.1 tqdm/4.45.0 CPython/3.8.2

File hashes

Hashes for mg_toolkit-0.6.5-py3-none-any.whl
Algorithm Hash digest
SHA256 fea14df0316ad69d9164281cfc0c70b53959c6b83b03f22ac2f22c5738995be2
MD5 2937b236e81d9ac82d20503289edd149
BLAKE2b-256 ae17991bdb9b87e449f5254fe70d9c5eae8378db846ef927851987f583ddad61

See more details on using hashes here.

Supported by

AWS AWS Cloud computing and Security Sponsor Datadog Datadog Monitoring Fastly Fastly CDN Google Google Download Analytics Microsoft Microsoft PSF Sponsor Pingdom Pingdom Monitoring Sentry Sentry Error logging StatusPage StatusPage Status page