CLI tool to group SMILES by Bemis–Murcko scaffold, with deterministic 3D conformer generation and on-disk caching (RDKit-based).
Project description
bm-group
CLI tool to group SMILES by Bemis–Murcko scaffold, with deterministic 3D conformer generation and on-disk caching (RDKit-based).
Summary
Given an input text file containing one SMILES per line (extra columns are ignored), bm_group.py:
- parses and canonicalizes each SMILES (isomeric SMILES),
- computes its Bemis–Murcko scaffold,
- (optionally) generates multiple 3D conformers (ETKDG v3) and optimizes them (MMFF if requested and available; otherwise UFF),
- writes a JSON that groups canonical SMILES by scaffold.
To scale to large datasets, grouping is implemented via an external sort: it streams scaffold<TAB>canonical_smiles pairs to a temporary file, sorts them in chunks, merges the sorted chunks, then writes groups.
Features
- Scaffold grouping via RDKit Murcko scaffolds.
- Deterministic 3D cache keyed by canonical SMILES hash (conformers + energies stored).
- External sort (chunk sort + k-way merge) for large inputs.
- Multiprocessing (configurable start method and worker count).
- Atomic writes for cache entries to reduce corruption risk.
Requirements
- Python ≥ 3.10
- RDKit (tested with modern RDKit builds)
- tqdm
Installation
This project is a single script. Create an environment with RDKit, then run it.
Conda (recommended)
conda create -n bm-group -c conda-forge python=3.11 rdkit tqdm
conda activate bm-group
pip
pip install tqdm rdkit-pypi bm-group
Quick start
Input file (one SMILES per line; other columns are ignored):
CC[C@H](F)C(=O)O some_id_1
c1ccccc1 benzene
Run:
python bm_group.py --input example_set.smiles --output example_set.json
With a custom cache dir and fewer workers:
python bm_group.py --input example_set.smiles --output example_set.json --cache-dir .cache_bm --nprocs 4
Prefer MMFF when available (otherwise fallback to UFF):
python bm_group.py --input example_set.smiles --output example_set.json --prefer-mmff
CLI
python bm_group.py --input INPUT --output OUTPUT [options]
Arguments
--input(required): input text file (one SMILES per line; extra columns ignored).--output(required): output JSON path.
Options
--cache-dir: cache directory for metadata and conformers (default:.cache_bm).--tmp-dir: temp directory for external sort files (default: auto-created).--num-confs: number of 3D conformers per molecule (default:20).--max-attempts: ETKDG attempts/iterations (RDKit-version dependent) (default:20).--prefer-mmff: if set, use MMFF optimization when possible; otherwise UFF.--prune-rms: prune conformers closer than this RMS threshold in Å (default: none).--chunk-lines: lines per chunk for external sorting (default:300000).--max-open: maximum open files during merge (default:64).--nprocs: number of worker processes (0= all available cores) (default:0).--start-method: multiprocessing start method:auto,fork,spawn(default:auto).--keep-temp: keep temporary sort/merge files (useful for debugging).
Output format
The output is a JSON object:
- keys are stringified incremental group IDs (
"1","2", ...), - each value is a list of canonical isomeric SMILES belonging to the same scaffold group.
Example:
{
"1": [
"c1ccccc1"
],
"2": [
"CC[C@H](F)C(=O)O",
"CCC(F)C(=O)O"
]
}
Notes:
- Group IDs are assigned in the lexicographic order of scaffold SMILES (because pairs are sorted by scaffold).
- Within each group, molecules are written in sorted order; exact duplicates are skipped during writing.
Cache
--cache-dir stores two things per canonical SMILES (keyed by SHA1 hash):
molecules/<pfx>/<hash>.json— metadata:cache_version,rdkit_version,canonical_smiles,scaffold_smiles,num_confs,max_attempts,prefer_mmff,prune_rms,conformers_path(relative path to the gzipped payload)
conformers/<pfx>/<hash>.json.gz— gzipped JSON:atoms: list of atomic symbols (with explicit H),conformers: list of conformers, each as[[x,y,z], ...],energies: per-conformer energies (NaN if optimization failed),method:"MMFF"or"UFF".
The cache is automatically reused only when the stored parameters match the current run (RDKit version and key settings).
How it works
- Read non-empty lines, take the first token as SMILES.
- Canonicalize with RDKit (
isomericSmiles=True). - Compute scaffold (RDKit Murcko scaffold).
- 3D embedding + optimization (optional but enabled by default because conformers are generated for the cache):
- seed derived from canonical SMILES SHA1 (deterministic),
- ETKDG v3 embedding,
- MMFF (if
--prefer-mmffand parameters available) or UFF.
- Stream pairs
scaffold<TAB>canonicalto a temp file. - External sort: sort in chunks, then k-way merge.
- Write groups as JSON.
Troubleshooting
Windows / multiprocessing:
- On Windows, the start method is typically
spawn. If you hit multiprocessing or RDKit-related worker issues, try:--nprocs 1(single process), or- explicitly
--start-method spawn.
No valid SMILES:
- If every line is invalid, the script writes
{}and prints a message on stderr.
License
Released under the MIT License. See LICENSE.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file bm_group-0.1.1.tar.gz.
File metadata
- Download URL: bm_group-0.1.1.tar.gz
- Upload date:
- Size: 10.0 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
40c60fbb436f53bcc4a5ee750c65a29929574bae604f705c4e4f37065183e19f
|
|
| MD5 |
03ff3e06e5308da79e4b6e111072a689
|
|
| BLAKE2b-256 |
6dcdfdfb712e04077a4da5457dae5f989cffcf307451a321b563727bb4e6e29d
|
Provenance
The following attestation bundles were made for bm_group-0.1.1.tar.gz:
Publisher:
publish.yml on mmassuoli/bm_group
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
bm_group-0.1.1.tar.gz -
Subject digest:
40c60fbb436f53bcc4a5ee750c65a29929574bae604f705c4e4f37065183e19f - Sigstore transparency entry: 850114198
- Sigstore integration time:
-
Permalink:
mmassuoli/bm_group@2d48582d2eb07832f6244966fba85fda5d30a1bd -
Branch / Tag:
refs/tags/v0.1.1 - Owner: https://github.com/mmassuoli
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@2d48582d2eb07832f6244966fba85fda5d30a1bd -
Trigger Event:
push
-
Statement type:
File details
Details for the file bm_group-0.1.1-py3-none-any.whl.
File metadata
- Download URL: bm_group-0.1.1-py3-none-any.whl
- Upload date:
- Size: 10.7 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
4a3b515dbdf181608ff2aa02894d93f92a3c4a90ff7630786af355e0863788dd
|
|
| MD5 |
6fcfeab9eadf8cbecbce3316e1f12c7c
|
|
| BLAKE2b-256 |
4d403f0aba1a7266afcec145ae4b52f0f71ea02d85d44546ab24ddd163825b01
|
Provenance
The following attestation bundles were made for bm_group-0.1.1-py3-none-any.whl:
Publisher:
publish.yml on mmassuoli/bm_group
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
bm_group-0.1.1-py3-none-any.whl -
Subject digest:
4a3b515dbdf181608ff2aa02894d93f92a3c4a90ff7630786af355e0863788dd - Sigstore transparency entry: 850114200
- Sigstore integration time:
-
Permalink:
mmassuoli/bm_group@2d48582d2eb07832f6244966fba85fda5d30a1bd -
Branch / Tag:
refs/tags/v0.1.1 - Owner: https://github.com/mmassuoli
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@2d48582d2eb07832f6244966fba85fda5d30a1bd -
Trigger Event:
push
-
Statement type: