Cypstrate
CYPstrate consists of a collection of machine learning classifiers (random forest and support vector machines) for the prediction of substrates and non-substrates of the nine most important human CYP isozymes in the metabolism of xenobiotics (i.e. CYPs 1A2, 2A6, 2B6, 2C8, 2C9, 2C19, 2D6, 2E1 and 3A4). The models are trained on a high-quality data set of 1831 substrates and non-substrates compiled from public sources.
Installation
# requires Python 3.8
pip install -U cypstrate
Usage
CYPstrate can be called from the command line. Examples:
# input in SMILES format
cypstrate "CCOC(=O)N1CCN(CC1)C2=C(C(=O)C2=O)N3CCN(CC3)C4=CC=C(C=C4)OC"
# prediction is one of "best_performance" (default) or "full_coverage"
cypstrate --prediction-mode full_coverage "CCN(C)C(=O)OC1=CC=CC(=C1)C(C)N(C)C"
# input can be a file
cypstrate molecules.sdf > result.csv
# output format can be specified
cypstrate --output sdf molecules.smiles > result.sdf
# more information via --help
cypstrate --help
The model can be used in Python. Calling the predict function of the
CypstrateModel class results in a pandas DataFrame containing the prediction
results for each input molecule.
from cypstrate import CypstrateModel
model = CypstrateModel()
# "predict" method accepts a list of SMILES representations
df_predictions = model.predict(['CCN(C)C(=O)OC1=CC=CC(=C1)C(C)N(C)C'])
# ... or a list of file paths
df_predictions = model.predict(['part1.sdf', 'part2.sdf'])
The result DataFrame contains the columns:
- mol_id: unique number identifying the input molecule
- input: the raw representation provided as input (e.g. OCCCCC)
- input_type: the representation type of the input (e.g. smiles)
- source: the input source (e.g. my_molecules.sdf)
- name: the name of the input molecule (if provided in the input)
- input_mol: the RDKit molecule parsed from the input representation
- preprocessed_mol: the RDKit molecule after preprocessing
- errors: a list of errors that occured during reading or preprocessing the input
- prediction_1a2, prediction_2a6, prediction_2b6, prediction_2c8, prediction_2c9, prediction_2c19, prediction_2d6, prediction_2e1, prediction_3a4: probability (between 0 and 1) of being a substrate of the given CYP isozyme
- neighbor_1a2, neighbor_2a6, neighbor_2b6, neighbor_2c8, neighbor_2c9, neighbor_2c19, neighbor_2d6,neighbor_2e1,neighbor_3a4: similarity to the most similar molecule in the corresponding training set
Contribute
conda env create -f environment.yml
conda activate cypstrate
pip install -e .[dev,test]
ptw
Contributors
- Malte Holmer
- Steffen Hirte
- Axinya Tokareva
Metadata
Release files for cypstrate 0.1.6
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| cypstrate-0.1.6.tar.gz | 97.9 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| cypstrate-0.1.6-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 196.9 MB
Release files / cypstrate-0.1.6.tar.gz
| Download URL | cypstrate-0.1.6.tar.gz |
|---|---|
| Size | 97.9 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
28c24199d1018be9609ede4fd4f33ac4e39fe5ae02843df97568ea8dad2724d4
|
|
BLAKE2b-256 checksum How to use checksums |
c74a3d96cf2a139627470eb24d08ac498ff5ffcc86b1103751062e1bc6123783
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/4.0.2 CPython/3.8.15
|
Release files / cypstrate-0.1.6-py3-none-any.whl
| Download URL | cypstrate-0.1.6-py3-none-any.whl |
|---|---|
| Size | 99.1 MB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
7346a1a8cd1d7cd0f3114766dd59b03f2d495d773b24aaee8956952fed92efec
|
|
BLAKE2b-256 checksum How to use checksums |
028e802c35baa7b10d562073b9760c87b7c76e1cf7b44fdeb40cd2e5fddc126b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/4.0.2 CPython/3.8.15
|