Skip to main content

A project to process enzymatic reactions

Project description

EVODEX

EVODEX is a Python package that provides tools for the prediction of mechanistically plausible reaction products, validation of reactions, and mass spectrometry interpretation. It can be installed via PyPI for immediate use in Python projects. Alternatively, users can clone the repository to run the full mining pipeline and generate a customized dataset or website.

Current Release

This is the EVODEX.1 collection. All IDs start with 'EVODEX.1' and are immutable, ensuring they can be externally referenced without collisions or missing references. Future distributions will be numbered EVODEX.2, EVODEX.3, etc., and may not have reverse compatibility with previous EVODEX.0 IDs. For example, EVODEX.1-E2 may not represent the same SMIRKS as EVODEX.0-E2.

Table of Contents

  1. Installation
  2. Usage
  3. Website and Dataset Access
  4. Building and Running the Mining Pipeline
  5. Citing EVODEX
  6. License

Running EVODEX via PyPi

Installation

To use EVODEX in a python project, use the following command:

pip install evodex

Usage

The following Jupyter notebooks demonstrate the usage of the PyPI distribution for the three primary use cases:

These notebooks are also available in the notebooks section of the repository.

Synthesis

The synthesis module provides tools for predicting reaction products using reaction operators. Below is an example usage:

from evodex.synthesis import project_reaction_operator, project_evodex_operator, project_synthesis_operators

# Specify propanol as the substrate as SMILES
substrate = "CCCO"

# Representation of alcohol oxidation as SMIRKS:
smirks = "[H][C:8]([C:7])([O:9][H])[H:19]>>[C:7][C:8](=[O:9])[H:19]"

# Project the oxidation operator on propanol:
result = project_reaction_operator(smirks, substrate)
print("Direct projection: ", result)

# Specify the dehydrogenase reaction by its EVODEX ID:
evodex_id = "EVODEX.0-E2"

# Apply the dehydrogenase operator to propanol
result = project_evodex_operator(evodex_id, substrate)
print("Referenced EVODEX projection: ", result)

# Project All Synthesis Subset EVODEX-E operators on propanol
result = project_synthesis_operators(substrate)
print("All Synthesis Subset projection: ", result)

For more detailed usage, refer to the EVODEX Synthesis Demo.

Evaluation

The evaluation module provides tools for evaluating reaction operators and synthesis results. Below is an example usage:

from evodex.evaluation import assign_evodex_F, match_operators

# Define reaction as oxidation of propanol
reaction = "CCCO>>CCC=O"

# Assign EVODEX-F IDs
assign_results = assign_evodex_F(reaction)
print(assign_results)

# Match reaction operators of type 'E' (or C or N)
match_results = match_operators(reaction, 'E')
print(match_results)

For more detailed usage, refer to the EVODEX Evaluation Demo.

Mass Spectrometry

The mass spectrometry module provides tools for predicting masses and identifying reaction operators. Below is an example usage:

from evodex.mass_spec import calculate_mass, find_evodex_m, get_reaction_operators, predict_products

# Calculate exact mass of the compound cortisol as an [M+H]+ ion
cortisol_M_plus_H = "O=C4\C=C2/[C@]([C@H]1[C@@H](O)C[C@@]3([C@@](O)(C(=O)CO)CC[C@H]3[C@@H]1CC2)C)(C)CC4.[H+]"
mass = calculate_mass(cortisol_M_plus_H)

# Define observed masses
substrate_mass = 363.2166 # The expected mass for cortisol's ion
potential_product_mass = 377.2323 # A mass of unknown identity
mass_diff = potential_product_mass - substrate_mass

# Find matching EVODEX-M entries
matching_evodex_m = find_evodex_m(mass_diff, 0.01)
print(matching_evodex_m)

# Get reaction operators
matching_operators = get_reaction_operators(mass_diff, 0.01)
print(matching_operators)

# Predict product structures
predicted_products = predict_products(cortisol_M_plus_H, mass_diff, 0.01)
print(predicted_products)

For more detailed usage, refer to the EVODEX Mass Spec Demo.

Website and Dataset Access

A static website for exploring the EVODEX.1 dataset is available here:

🔗 https://ucb-bioe-anderson-lab.github.io/evodex-1-site/

This site provides a browsable, hyperlinked index of all EVODEX.1 operators and includes links to Colab demos and operator definitions.

If you prefer to work offline or programmatically, you can download the full set of operator tables (in CSV format) directly from this directory:

📂 https://github.com/UCB-BioE-Anderson-Lab/EVODEX/tree/main/evodex/data

Building and Running the Mining Pipeline

The mining pipeline allows you to reproduce the full EVODEX operator set from raw data. This is the process we use to generate the operators that ship with the PyPI distribution. Running the mining pipeline is only required if you want to modify the data, change the curation process, or experiment with new reactions.

⚠️ This project depends on a narrow window of compatible package versions. It is strongly recommended to use Python 3.11.13 and avoid newer Python versions unless you are familiar with rebuilding the RDKit dependency stack using Conda.

Build Environment

  • This release is tested and works with Python 3.11.13.
  • Using other Python versions (especially ≥ 3.12) may cause installation issues with key dependencies like rdkit-pypi, scipy, and pandas.
  • Do not attempt to upgrade individual packages or change the Python version unless you're using Conda, which is not recommended for this build.
  • For consistency and stability, we recommend using this exact Python version and installing via venv.
/opt/homebrew/bin/python3.11 -m venv venv
source venv/bin/activate
pip install --upgrade pip
pip install -r requirements.txt

Running the Pipeline

You can run the entire pipeline automatically by running:

python run_pipeline.py

Alternatively, you can run the pipeline step-by-step. First, download and prepare the raw data file, then run the modules sequentially.

To download and prepare the raw data file:

import requests
import gzip

# Download the file
url = "https://github.com/hesther/enzymemap/blob/main/data/processed_reactions.csv.gz?raw=true"
r = requests.get(url)
with open("/content/processed_reactions.csv.gz", "wb") as f:
    f.write(r.content)

# Decompress the file
with gzip.open("/content/processed_reactions.csv.gz", "rt") as f_in:
    with open("/content/processed_reactions.csv", "wt") as f_out:
        f_out.write(f_in.read())

Then run the following modules sequentially:

python -m pipeline.phase1_data_preparation
python -m pipeline.phase2_formula_pruning
python -m pipeline.phase3_ero_mining
python -m pipeline.phase3a_ero_pruning
python -m pipeline.phase3b_ero_trimming
python -m pipeline.phase3c_ero_publishing
python -m pipeline.phase4_operator_completion
python -m pipeline.phase5_mass_subset
python -m pipeline.phase6_synthesis_subset
python -m pipeline.phase7_website

Citing EVODEX

If you use EVODEX, please cite our preprint: https://www.biorxiv.org/content/10.1101/2025.06.19.660615v1.full

License

EVODEX is released under the MIT License.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

evodex-1.0.5.tar.gz (4.6 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

evodex-1.0.5-py3-none-any.whl (4.7 MB view details)

Uploaded Python 3

File details

Details for the file evodex-1.0.5.tar.gz.

File metadata

  • Download URL: evodex-1.0.5.tar.gz
  • Upload date:
  • Size: 4.6 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.11.13

File hashes

Hashes for evodex-1.0.5.tar.gz
Algorithm Hash digest
SHA256 bbc133780f91b95e2501ad64f8f8074f122c1d53f1db77c17e2c8f1b52a2dd08
MD5 80149de68a515b09a4259b6a522f0c6b
BLAKE2b-256 c8f43a135fea748e5114b65b6e27cbdc4a2160e5fb6b3d78cbe6481fda65c8ce

See more details on using hashes here.

File details

Details for the file evodex-1.0.5-py3-none-any.whl.

File metadata

  • Download URL: evodex-1.0.5-py3-none-any.whl
  • Upload date:
  • Size: 4.7 MB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.11.13

File hashes

Hashes for evodex-1.0.5-py3-none-any.whl
Algorithm Hash digest
SHA256 3bd5a79888f754d93b1dfe83ebb470b31a56293574718322c8b95d812d8a6dee
MD5 3b70b852f99d13842ad325f3f54e8c31
BLAKE2b-256 f162ed1b3a33565239bd7960ce0a9264da7d5171953fc9e68dafc3ceecffab63

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page