📜 Papyrus Structure Pipeline
A small, opinionated molecule standardization pipeline built on top of the ChEMBL Structure Pipeline and RDKit. It turns arbitrary input structures into a consistent parent form suitable for bioactivity datasets: salts and metals stripped, charges neutralized, tautomers canonicalized, and mixtures/inorganics/out-of-range molecular weights filtered out. First used to curate the Papyrus bioactivity dataset.
✨ Features
- 🧪 ChEMBL-based standardization — wraps the ChEMBL Structure Pipeline's parent-structure extraction and normalization, run twice (before and after tautomer canonicalization) for a consistent final form.
- 🧂 Configurable salt & metal stripping — removes a user-extensible set of salts/metals beyond what the ChEMBL pipeline already handles.
- ⚛️ Charge neutralization — uncharges ionizable groups and zwitterions.
- 🔀 Bounded tautomer canonicalization — deterministic canonical tautomer picking with a capped search, so runtime stays predictable even for polyphenol-like molecules with many tautomerizable sites.
- 🧬 Organic/inorganic filtering — configurable allowed-atom list and C-C bond requirement.
- 🧩 Mixture & size filtering — drop multi-fragment mixtures and molecules outside a molecular-weight window.
- 🔍 Rich failure reporting —
return_type=Truereports why a molecule was rejected (mixture, inorganic, too small/large, standardization error) instead of just returningNone. - 🏷️ Typed API — ships
py.typed, fully type-hinted, mypy-clean.
Citing
If you use this package, please cite the Papyrus dataset it was first used in:
📦 Installation
pip install papyrus-structure-pipeline
Or from source:
git clone https://github.com/OlivierBeq/Papyrus_structure_pipeline.git
pip install ./Papyrus_structure_pipeline
🛠️ Requirements
- Python 3.11+
- RDKit
- ChEMBL Structure Pipeline (installed automatically)
💡 Usage
Standardize a compound
Comparison to the ChEMBL Structure Pipeline:
from rdkit import Chem
from chembl_structure_pipeline import standardizer as ChEMBL_standardizer
from papyrus_structure_pipeline import standardizer as Papyrus_standardizer
# CHEMBL1560279
smiles = "CCN(CC)C(=O)[n+]1ccc(OC)cc1.c1ccc([B-](c2ccccc2)(c2ccccc2)c2ccccc2)cc1"
mol = Chem.MolFromSmiles(smiles)
out1 = ChEMBL_standardizer.standardize_mol(mol)
out2 = Papyrus_standardizer.standardize(mol)
print(Chem.MolToSmiles(out1))
# CCN(CC)C(=O)[n+]1ccc(OC)cc1.c1ccc([B-](c2ccccc2)(c2ccccc2)c2ccccc2)cc1
print(Chem.MolToSmiles(out2))
# CCN(CC)C(=O)[n+]1ccc(OC)cc1
Get details on the standardization to identify why it fails for some molecules:
smiles_list = [
# erlotinib
"n1cnc(c2cc(c(cc12)OCCOC)OCCOC)Nc1cc(ccc1)C#C",
# midecamycin
"CCC(=O)O[C@@H]1CC(=O)O[C@@H](C/C=C/C=C/[C@@H]([C@@H](C[C@@H]([C@@H]([C@H]1OC)O[C@H]2[C@@H]([C@H]([C@@H]([C@H](O2)C)O[C@H]3C[C@@]([C@H]([C@@H](O3)C)OC(=O)CC)(C)O)N(C)C)O)CC=O)C)O)C",
# selenofolate
"C1=CC(=CC=C1C(=O)NC(CCC(=O)OCC[Se]C#N)C(=O)O)NCC2=CN=C3C(=N2)C(=O)NC(=N3)N",
# cisplatin
"N.N.Cl[Pt]Cl",
]
for smiles in smiles_list:
mol = Chem.MolFromSmiles(smiles)
print(Papyrus_standardizer.standardize(mol, return_type=True))
# (<rdkit.Chem.rdchem.Mol object at 0x000000946F99B580>, <StandardizationResult.CORRECT_MOLECULE: 1>)
# (None, <StandardizationResult.NON_SMALL_MOLECULE: 2>)
# (None, <StandardizationResult.INORGANIC_MOLECULE: 3>)
# (None, <StandardizationResult.MIXTURE_MOLECULE: 4>)
Allow other atoms to be considered organic:
smiles = "CCN(CC)C(=O)C1=CC=C(S1)C2=C3C=CC(=[N+](C)C)C=C3[Se]C4=C2C=CC(=C4)N(C)C.F[P-](F)(F)(F)(F)F"
mol = Chem.MolFromSmiles(smiles)
print(Papyrus_standardizer.standardize(mol, return_type=True))
# (None, <StandardizationResult.INORGANIC_MOLECULE: 3>)
Papyrus_standardizer.ORGANIC_ATOMS.append('Se')
print(Papyrus_standardizer.standardize(mol, return_type=True))
# (<rdkit.Chem.rdchem.Mol object at 0x0000009F24D15F90>, <StandardizationResult.CORRECT_MOLECULE: 1>)
Papyrus_standardizer.ORGANIC_ATOMS = Papyrus_standardizer.ORGANIC_ATOMS[:-1]
print(Papyrus_standardizer.standardize(mol, return_type=True))
# (None, <StandardizationResult.INORGANIC_MOLECULE: 3>)
Add custom substructures to be removed as salts:
# lomitapide
smiles = "C1CN(CCC1NC(=O)C2=CC=CC=C2C3=CC=C(C=C3)C(F)(F)F)CCCCC4(C5=CC=CC=C5C6=CC=CC=C64)C(=O)NCC(F)(F)F.c1ccccc1"
mol = Chem.MolFromSmiles(smiles)
print(Papyrus_standardizer.standardize(mol, return_type=True))
# (None, <StandardizationResult.MIXTURE_MOLECULE: 4>)
Papyrus_standardizer.SALTS.append('c1ccccc1')
print(Papyrus_standardizer.standardize(mol, return_type=True))
# (<rdkit.Chem.rdchem.Mol object at 0x0000009F24D15F90>, <StandardizationResult.CORRECT_MOLECULE: 1>)
Papyrus_standardizer.SALTS = Papyrus_standardizer.SALTS[:-1]
print(Papyrus_standardizer.standardize(mol, return_type=True))
# (None, <StandardizationResult.MIXTURE_MOLECULE: 4>)
📄 License
This project is licensed under the MIT License - see the LICENSE file for details.
📚 API Documentation
def standardize(mol,
remove_additional_salts=True, remove_additional_metals=True,
filter_mixtures=True, filter_inorganic=True, filter_non_small_molecule=True,
canonicalize_tautomer=True, small_molecule_min_mw=200, small_molecule_max_mw=800,
tautomer_allow_stereo_removal=True, tautomer_max_tautomers=50, return_type=False,
raise_error=True
) -> Chem.Mol | None:
Standardizes a molecule: ChEMBL parent extraction, salt/metal stripping, mixture/inorganic/size filtering, charge neutralization, tautomer canonicalization, then ChEMBL parent extraction again.
Parameters
- mol : Chem.Mol RDKit molecule object to standardize.
- remove_additional_salts : bool Removes a custom set of fragments if present in the molecule object.
- remove_additional_metals : bool
Removes metal fragments if present in the molecule object. Ignored if
remove_additional_saltsis set toFalse. - filter_mixtures : bool
Return
Noneif the molecule is a mixture. - filter_inorganic : bool
Return
Noneif the molecule is inorganic. - filter_non_small_molecule : bool
Return
Noneif the molecule is not a small molecule. - canonicalize_tautomer : bool Canonicalize the tautomeric state of the molecule.
- small_molecule_min_mw : float Molecular weight under which a molecule is considered too small.
- small_molecule_max_mw : float Molecular weight above which a molecule is considered too big.
- tautomer_allow_stereo_removal : bool Allow the tautomer search algorithm to remove stereocenters.
- tautomer_max_tautomers : int Maximum number of tautomers to consider by the tautomer search algorithm. Bounds worst-case runtime for molecules with many tautomerizable sites.
- return_type : bool
Add a
StandardizationResultto the return value. - raise_error : bool
Raise an exception upon failure, otherwise return
None(or(None, StandardizationResult.STANDARDIZATION_ERROR)ifreturn_typeis also set).
def is_organic(mol, return_type=False) -> bool:
Returns whether the RDKit molecule is organic: no ChEMBL exclusion flag, at least one C-C bond,
and made up only of atoms listed in ORGANIC_ATOMS.
Parameters
- mol : Chem.Mol RDKit molecule object to check the organic nature of.
- return_type : bool
Add an
InorganicSubtypeto the return value.
def is_small_molecule(mol, min_molwt=200, max_molwt=800) -> bool:
Returns whether the RDKit molecule has a molecular weight within min_molwt and max_molwt.
Parameters
- mol : Chem.Mol RDKit molecule object to check the molecular weight of.
- min_molwt : float Molecular weight under which a molecule is considered too small.
- max_molwt : float Molecular weight above which a molecule is considered too big.
def is_mixture(mol) -> bool:
Returns whether the RDKit molecule is composed of multiple fragments.
Parameters
- mol : Chem.Mol RDKit molecule object to check the fragment count of.
Result types
- StandardizationResult — outcome of
standardize():CORRECT_MOLECULE,NON_SMALL_MOLECULE,INORGANIC_MOLECULE,MIXTURE_MOLECULE,STANDARDIZATION_ERROR. - InorganicSubtype — reason
is_organic()returnedFalse:NO_CC_BOND,NOT_CHONPSFClIBrB,EXCLUSION_FLAG_SET. - SaltStrippingResult — outcome of the internal salt-stripping step:
STRIPPED_MOLECULE,EMPTY_MOLECULE.
Module-level configuration
- SALTS : list[str] — extra salt SMILES stripped beyond the ChEMBL pipeline's own list; user-extensible.
- METALS : list[str] — metal SMILES stripped when
remove_additional_metals=True. - ORGANIC_ATOMS : list[str] — atom symbols allowed for a molecule to be considered organic.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file papyrus_structure_pipeline-1.0.0.tar.gz.
File metadata
- Download URL: papyrus_structure_pipeline-1.0.0.tar.gz
- Upload date:
- Size: 19.2 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/6.1.0 CPython/3.12.3
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
c612813fb4f6ae2614025278581a06a7a8225a014b08db90b51bd88cc025626c
|
|
| MD5 |
f4d2376240505495b6f73ffcaa433bd5
|
|
| BLAKE2b-256 |
f6ad751e3116bbb62bd31a4df7e436b03621c2fad2fed17d69e02140c5463a8a
|
File details
Details for the file papyrus_structure_pipeline-1.0.0-py3-none-any.whl.
File metadata
- Download URL: papyrus_structure_pipeline-1.0.0-py3-none-any.whl
- Upload date:
- Size: 14.2 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/6.1.0 CPython/3.12.3
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
1cdab3c0777ff6a11f9cc9ce77a206d5324e56b17b2fcbf083928b0d14fe1fb4
|
|
| MD5 |
8795735126392bfca88f6def193dad96
|
|
| BLAKE2b-256 |
3b4c0b5d21f2e3ef7ebcde9c64412f3ab2afc08adede3bd60e58a452fdd6581b
|