Skip to main content

Publish to PyPI

EnrichM is a set of comparative genomics tools for large sets of metagenome assembled genomes (MAGs). The current functionality includes:

  1. A basic annotation pipeline for MAGs.
  2. A pipeline to determine the metabolic pathways that are encoded by MAGs, using KEGG modules as a reference (although custom pathways can be specified).
  3. A pipeline to identify genes or metabolic pathways that are enriched within and between user-defined groups of genomes (groups can be genomes that are related functionally, phylogenetically, recovered from different environments, etc).
  4. Construct random forest machine learning models from the functional composition of MAGs, metagenomes or transcriptomes.
  5. Apply random forest models to classify new MAGs or metagenomes.

EnrichM is under active development, so there is no guarantee that master is stable. It is recommended to install from a tagged release (see below).

Installation

Dependencies

EnrichM is written in Python 3 and requires >= 3.8. EnrichM requires the following non-Python dependencies:

conda (recommended)

Clone the repository and create the conda environment:

git clone https://github.com/geronimp/enrichM.git
cd enrichM
conda env create -f environment.yml
conda activate enrichm
pip install .

PyPI

pip install enrichm

Note: non-Python dependencies (hmmer, diamond, prodigal, parallel, mmseqs2, and mcl if using --annotate_ortholog) must be installed separately when using PyPI.

After installation, you'll need to download the back-end databases.

Setup

Loading EnrichM's database

The database contains Pfam-A HMMs, TIGRfam HMMs, dbCAN HMMs, and KoFamKOALA HMMs. By default it is installed in ~/enrichm_data. Build it using:

enrichm data --create

To store the database in a custom location:

enrichm data --create --db_path /path/to/database/

To uninstall:

enrichm data --uninstall

Using an existing database

If the database was built in a custom location, set the ENRICHM_DB environment variable so EnrichM can find it:

export ENRICHM_DB=/path/to/database/

Add this to your .bashrc or conda activate.d script to avoid setting it each session.

Subcommands

annotate

Annotate population genomes with KO HMMs, Pfam, TIGRfam, and CAZymes using dbCAN. The result is a GFF file for each genome and a frequency matrix for each annotation type (annotation IDs as rows, genomes as columns).

classify

Reads KO annotations in the form of a matrix and determines which KEGG modules are complete. Annotation matrices can be generated using annotate.

enrichment

Enrichment reads an annotation matrix (IDs as rows, genomes as columns) and a metadata file separating genomes into groups, and runs statistical tests (Mann-Whitney U, Fisher's exact, Kruskal-Wallis) to identify enriched annotations between groups. Outputs include effect sizes, fold changes, and FDR-corrected p-values. Additional features include:

  • Synteny analysis: identifies conserved gene blocks (operons) among enriched genes using intergenic distance thresholds
  • Mobile element proximity: flags enriched genes located near transposases or insertion sequences
  • IndVal (Indicator Value) analysis: automatically computed for every run; scores each annotation for specificity and fidelity to each group with permutation-based FDR-corrected p-values (indval_results.tsv)
  • NMF decomposition (--decompose): factorises the annotation matrix into latent functional components and tests component scores between groups via Mann-Whitney U; outputs loadings, per-genome scores, and a heatmap. Use --n_components to fix the number of components or --select_components to select automatically via cophenetic correlation
  • Phylogenetic correction (--tree): Scoary1-style paired comparison that tests whether pairwise enrichment signals survive correction for shared ancestry; requires a Newick tree with tip labels matching genome names
  • Multi-type batch mode (--all): when used with --annotate_output, runs enrichment for every annotation type present and writes results to per-type subdirectories under --output
  • Accepts output from annotate, or external tools including DRAM, eggNOG-mapper

generate

Trains a random forest classifier or regressor from an annotation matrix and a metadata file of labels. Performs automated hyperparameter tuning via RandomizedSearchCV (with optional GridSearchCV refinement). Outputs the trained model, feature importances, and accuracy summary.

predict

Applies a trained model (from generate) to a new annotation matrix and outputs per-sample predictions and class probabilities.

Contact

If you have any feedback about EnrichM, drop an email to the SupportM public help forum. Software by Joel A. Boyd (@geronimp) at the Australian Centre for Ecogenomics (ACE).

License

EnrichM is licensed under the GNU GPL v3+. See LICENSE.txt for further details.

Contributing

I want EnrichM to be as useful as possible, so please feel free to leave feature requests and bug reports.

Citation

If you find EnrichM useful and use it in your work, please cite it as follows:

Comparative genomics using EnrichM. Joel A Boyd, Ben J Woodcroft, Gene W Tyson. In preparation.

Release files for enrichm 0.6.10

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for enrichm 0.6.10
File Size Uploaded
enrichm-0.6.10.tar.gz 95.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for enrichm 0.6.10
File Interpreter ABI Platform
enrichm-0.6.10-py3-none-any.whl Python 3 none any Details

Total release size: 177.6 kB

Release files / enrichm-0.6.10.tar.gz

Download URL enrichm-0.6.10.tar.gz
Size 95.9 kB
Tags Source
SHA-256 checksum
How to use checksums
041d46880495b740e6024da0cd752bab0196913275d3c8f7365be1f86c9f9b79
BLAKE2b-256 checksum
How to use checksums
38177fa2f7a015e1f42b21049cd77d5f0c89cf27decd0c1a8fc77b29bc48182a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 30, 2026.

Transparency log

Release files / enrichm-0.6.10-py3-none-any.whl

Download URL enrichm-0.6.10-py3-none-any.whl
Size 81.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
1aa7ef165ce78df982b222930c0fe73af22752b4206cb57df02209591e11a5c5
BLAKE2b-256 checksum
How to use checksums
edef4a9de24329038c50f42e45708b0362ae9ae1131f437484a2b6135de0f655
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 30, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.6.10 This release

2 release files

0.6.9

2 release files

0.6.8

2 release files

0.6.6

2 release files

0.6.5

1 release file

0.6.4

1 release file

0.6.3

1 release file

0.6.2

1 release file

0.5.0

1 release file

0.4.15

1 release file

0.4.14

1 release file

0.4.13

1 release file

0.4.12

1 release file

0.4.11

1 release file

0.4.10

1 release file

0.4.9

1 release file

0.4.8

1 release file

0.4.7

1 release file

0.4.6

1 release file

0.4.5

1 release file

0.4.4

1 release file

0.4.3

2 release files

0.4.2

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.2.0

1 release file

0.1.8

2 release files

0.1.7

2 release files

0.1.6

2 release files

0.1.5

2 release files

0.1.3

2 release files

0.1.2

1 release file

0.1.1

2 release files

0.1.0

2 release files

0.0.11

1 release file

0.0.10

1 release file

0.0.9

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page