Skip to main content

MetaSnake: A metagenomic Bioinformatics Workflow Powered by Snakemake

Project description

Metasnake

Introduction

An easy-to-use wrapper for a robust snakemake pipeline designed for metagenomic assembly, gene prediction, genomic binning, taxonomy and functional annotations analysis of biogeochemical cycles. The pipeline features a series of flexible, interdependent modules optimized for fast execution. Each module can also be executed independently as a single pipeline, depending on the provided input. Interestingly, Diting 2.0v can be able to modify and add new modules based on snakemake workflow.

Workflow

This workflow defines data analysis in terms of rules that are specified in the Snakefile. workflow

How to use

This workflow provides a series of .yaml files (including config.yaml and dependent environment configuration file, like kegg.yaml), module files () and one Snakefile. You can run by editing the parameter file config.yaml and the main file Snakefile.

  • config.yaml: including defines of input and output folders, thread and cpu configuration for specific rules. Before starting, you can specify or define the input and output folder, the number of threads and the cpu for specific running rules.
    # Input and output folder
    READS_SUF: ".fastq"
    READS_DIR: "cleanreads"
    ASSEMBLY_DIR: "Assembly"
    PRODIGAL_DIR: "Prodigal"  
    CDHIT_DIR: "CD_hit"  
    FILTERED_DIR: "Filtered"
    BBMAP_DIR: "BBMap"  
    BWA_INDEX: "BBMap/bwa_index"
    MAPPING: "BBMap/mapping"
    PILEUP: "BBMap/pileup"
    GENE_ABUN_DIR: "Abundance" 
    KODB_DIR: "kofam_database"        
    KEGG_DIR: "KEGG_annotation"       
    OUT_DIR: "final_output"
    TABLE: "table" 
    # Thread configuration for specific rules
    threads:
      megahit: 8
      prodigal: 4
      cdhit: 4
      bwa: 8
      metawrap_binning: 8
      metawrap_refinement: 8
      quantify_bins: 8
    # CPU configuration for specific rules
    cpu:
      kegg_annotation: 4
      gtdb_classification: 4
    
  • Snakefile: including rule all and a series of modules.
# Define the main workflow
rule all:
    input:
        expand(os.path.join(config["GENE_ABUN_DIR"], "{basename}.abundance"), basename=config["BASENAMES"]),
        os.path.join(config["KEGG_DIR"], 'pathways_relative_abundance_gene_level.tab'),
        os.path.join(config["OUT_DIR"], 'carbon_cycle_sketch.png'),
        os.path.join(config["OUT_DIR"], 'carbon_cycle_heatmap.pdf'),
        os.path.join(config["BIN_ABUNDANCE_DIR"], "bin_abundance_table.tab")

# input modules
module Assembly_gene_prediction:
    snakefile: "./Assembly_gene_prediction.smk"
    config: config

module Gene_abundance:
    snakefile: "./Gene_abundance.smk"
    config: config

module Function_annotation:
    snakefile: "./Function_annotation.smk"
    config: config

module Pathway_abundance:
    snakefile: "./Pathway_abundance.smk"
    config: config

module Visualization:
    snakefile: "./Visualization.smk"
    config: config

module Binning:
    snakefile: "./Binning.smk"
    config: config
    
# execute modules
use rule * from Assembly_gene_prediction as *
use rule * from Gene_abundance as *
use rule * from Function_annotation as *
use rule * from Pathway_abundance as *
use rule * from Visualization as *
use rule * from Binning as *

Running

1. Download databases

# At the home directory of this program
  mkdir kofam_database/
  cd kofam_database/
  wget -c ftp://ftp.genome.jp/pub/db/kofam/ko_list.gz 
  wget -c ftp://ftp.genome.jp/pub/db/kofam/profiles.tar.gz 
  gzip -d ko_list.gz
  tar zxvf profiles.tar.gz

2. Create input folder contain your data

# At the home directory of this program
  mkdir cleanreads/ 

3. Activate snakemake environment

  conda activate snakemake

4. After defining the threads or cpu of specific rules at config.yaml file, and choosing the rule all inmodules of the specified rule, then run:

  snakemake --use-conda --core 8

Dependencies

  • metagenome.yaml: configuration environment of metagenomic analysis.
name: metagenome_env
channels:
 - bioconda
 - conda-forge
dependencies:
 - megahit >= 1.1.3
 - Prodigal >= 2.6.3
 - cd-hit >= 4.8.1
 - bwa >= 0.7.17
 - python>=3.8
 - hmmer
 - parallel
 - kofamscan
  • genomic.yaml: configuration environment of genomic binning.
name: genomic_env
channels:
 - ursky
 - bioconda
 - conda-forge
 - defaults
dependencies:
 - python=2.7
 - bwa
 - samtools
 - metabat2
 - concoct
 - MaxBin2
 - checkm-genome
 - prodigal
 - pplacer
 - numpy
 - matplotlib
 - pysam
 - hmmer
 - SPAdes
 - salmon
 - seaborn
 - metawrap-mg

Download databases

Installation

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

metasnake-1.0.0.tar.gz (15.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

metasnake-1.0.0-py3-none-any.whl (15.5 kB view details)

Uploaded Python 3

File details

Details for the file metasnake-1.0.0.tar.gz.

File metadata

  • Download URL: metasnake-1.0.0.tar.gz
  • Upload date:
  • Size: 15.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.0.1 CPython/3.10.13

File hashes

Hashes for metasnake-1.0.0.tar.gz
Algorithm Hash digest
SHA256 d3ee6d048b37a80c85a5e190e9c7ad6542d33d6bc55003bf05b4f7f8633c8e9e
MD5 20ec0789890c4f83857169e491cb1624
BLAKE2b-256 b130899d42aafae02f0004abd73f9fc67bd1da3aac42c85989e32efc4fe91e98

See more details on using hashes here.

File details

Details for the file metasnake-1.0.0-py3-none-any.whl.

File metadata

  • Download URL: metasnake-1.0.0-py3-none-any.whl
  • Upload date:
  • Size: 15.5 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.0.1 CPython/3.10.13

File hashes

Hashes for metasnake-1.0.0-py3-none-any.whl
Algorithm Hash digest
SHA256 eb32263df869acdc63bf55a96fa56a7637561f3cda22fc6149d493fd49ae7714
MD5 b5966760db5b3edee4c278a52b79b14b
BLAKE2b-256 89a419565140983b832f6d59cf4b0fc4772a929f6f3467782353b9e542b76724

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page