Skip to main content

cinful: A fully automated pipeline to identify microcinswith associated immunity proteins and export machinery

Project description

cinful

A fully automated pipeline to identify microcins along with their associated immunity proteins and export machinery

Installation

First, make sure to clone this repository:

git clone https://github.com/wilkelab/cinful.git

All software dependencies needed to run cinful are available through conda and are specified in cinful_conda.yml, the following helper script can be used to generate the cinful conda environment scripts/build_conda_env.sh, to run this script, you will need to have conda installed, as well as mamba (which helps speed up installation). To install mamba, use the following command:

conda install mamba -c conda-forge

To build the environment, run

bash scripts/build_conda_env.sh

Once setup is complete, you can activate the environment with

conda activate cinful

How to use

cinful takes a directory containing genome assemblies as input. All assemblies in the directory must end in .fna, if they end in a different extension, cinful will ignore them.

Snakemake is the core workflow management used by cinful, the main snakefile is located under cinful/Snakefile, which issues subroutines located in cinful/rules.

If installed properly, running python cinful.py -h will produce the following output.

cinful

optional arguments:
  -h, --help            show this help message and exit
  -d DIRECTORY, --directory DIRECTORY
                        Must be a directory containing uncompressed FASTA formatted genome assemblies with
                        .fna extension. Files within nested directories are fine
  -o OUTDIR, --outDir OUTDIR
                        This directory will contain all output files. It will be nested under the input
                        directory.
  -t THREADS, --threads THREADS
                        This specifies how many threads to allow snakemake to have access to for
                        parallelization

Example usage

There is a test dataset with an E. coli genome assembly to test cinful on under test/colcinV_Ecoli, you can run cinful on this dataset by running the following from the initial cinful directory:

python cinful/cinful.py -d test/colcinV_Ecoli -o <output_directory> -t <threads>

Workflow

The following workflow will be executed. cinful

Three output directories will be generated in your assembly_directory under a directory called cinfulOut.

  • 00_dbs
    • This is the initial location of the databases of verified microcins, CvaB, and immunity proteins.
  • 01_orf_homology
    • Prodigal will generate Open Reading Frame (ORF) predictions for the input assemblies
    • Those ORFs will be searched against the previously mentioned databases
  • 02_homology_results
    • The results from all the homology searches will be merged here
  • 03_best_hits
    • The top hits from the homology results will be placed here

Contributing

cinful currently exists as a wrapper to a series of snakemake subroutines, so adding functionality to it is as simple as adding additional subroutines. If there are any subroutines that you see are needed, feel free to raise an issue, and I will be glad to guide you through the process of making a pull request to add that feature.

Additionally, since cinful primarily works through snakemake, it can also be used by simply running the snakefiles separately, so if additional configuration is needed, in terms of the types of input files, this can probably be achieved that way.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

cinful-1.1.5.tar.gz (1.6 MB view details)

Uploaded Source

Built Distribution

cinful-1.1.5-py3-none-any.whl (1.6 MB view details)

Uploaded Python 3

File details

Details for the file cinful-1.1.5.tar.gz.

File metadata

  • Download URL: cinful-1.1.5.tar.gz
  • Upload date:
  • Size: 1.6 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/4.0.1 CPython/3.9.13

File hashes

Hashes for cinful-1.1.5.tar.gz
Algorithm Hash digest
SHA256 ef6c3342c591f58f7ebd2cf6e6d252b225e151171def6884f98e3db66ab26bbd
MD5 9a7b638c70b55779d1f4784df2f7d263
BLAKE2b-256 92b584c10cb7655d6d4d10e8427871831d3a732c753d2ccc4a7f5bc37e7077bf

See more details on using hashes here.

File details

Details for the file cinful-1.1.5-py3-none-any.whl.

File metadata

  • Download URL: cinful-1.1.5-py3-none-any.whl
  • Upload date:
  • Size: 1.6 MB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/4.0.1 CPython/3.9.13

File hashes

Hashes for cinful-1.1.5-py3-none-any.whl
Algorithm Hash digest
SHA256 da8ddd824b12b2929b198dde3c66aebb75d2bf34168c98333931b54631c8ee79
MD5 4487f10ea31fccfa2c915836878d029d
BLAKE2b-256 6605b40897f25b97ef78fc00a5d9cb8d264e69efd84d2d794e90e6c21705c7c9

See more details on using hashes here.

Supported by

AWS AWS Cloud computing and Security Sponsor Datadog Datadog Monitoring Fastly Fastly CDN Google Google Download Analytics Microsoft Microsoft PSF Sponsor Pingdom Pingdom Monitoring Sentry Sentry Error logging StatusPage StatusPage Status page