Retrosynthesis toolkit: single-step prediction, multi-step route planning, synthesizability scoring
Project description
synomega
A retrosynthesis toolkit that turns a target molecule into synthesis routes and a continuous synthesizability score. Three decoupled layers:
synthesizability is this target reachable from purchasable material, in N steps?
↑
search Retro* / MCTS / best-first over an AND-OR graph
↑
single-step product SMILES -> ranked reactant candidates
The layers meet at a deliberately narrow interface — a single-step backend only
implements predict(smiles, top_k) -> [Prediction] — so the planner and the
scorer do not care whether predictions come from a graph neural network, a
transformer, or plain template matching.
Installation
pip install synomega # core: rdkit + numpy
pip install "synomega[gnn]" # + the D-MPNN neural single-step backend (torch)
The neural backend is an optional extra on purpose: the template-rule backend runs anywhere, with no GPU and no PyTorch. Requires Python ≥ 3.10.
Zero-config quickstart
pip install synomega ships only code. The first time you ask for the default
model or stock, synomega downloads them (a few hundred MB) into
~/.cache/synomega — the same way spaCy and HuggingFace fetch models. So this
works out of the box:
import synomega
planner = synomega.load_default_planner() # downloads model + stock once
print(planner.plan("CC(=O)Nc1ccccc1O").best_route.describe())
from synomega import SynthesizabilityScorer
score = SynthesizabilityScorer(planner).score("CC(=O)Nc1ccccc1O", max_steps=5)
print(score.bb_coverage, score.min_steps)
Or pre-fetch from the command line, then use the CLI with no --model/--stock:
synomega download # cache the default assets
synomega plan --target "CC(=O)Nc1ccccc1O" --max-steps 5
Download mirrors. Assets are hosted on both a USTC GitLab registry (fast in
China) and (soon) GitHub. synomega auto-selects the faster reachable one by
latency; override with SYNOMEGA_MIRROR=ustc|github or point
SYNOMEGA_ASSETS_BASE=<url> at your own mirror. Change the cache location with
SYNOMEGA_CACHE.
Quick start
from synomega import Planner, SynthesizabilityScorer
from synomega.singlestep import TemplateGNN
from synomega.stock import InMemoryStock
model = TemplateGNN.from_pretrained("path/to/model_run") # a trained checkpoint
stock = InMemoryStock.from_keys_file("building_blocks.keys.gz")
planner = Planner(model, stock, algorithm="retrostar")
result = planner.plan("CC(=O)Nc1ccccc1", max_depth=5, time_limit=60)
print(result.solved)
print(result.best_route.describe())
target: CC(=O)Nc1ccccc1
solved: True steps: 2 depth: 2 bb_coverage: 1.00
[1] CC(=O)O.Nc1ccccc1>>CC(=O)Nc1ccccc1 (score=0.4348)
[2] O=[N+]([O-])c1ccccc1>>Nc1ccccc1 (score=0.2174)
Synthesizability scoring
scorer = SynthesizabilityScorer(planner)
r = scorer.score("CC(=O)Nc1ccccc1", max_steps=5)
r.solved # True — a complete route to purchasable material exists
r.bb_coverage # 1.0 — fraction of the best route's leaves that are buyable
r.min_steps # 2 — reactions in the shortest solved route
r.min_route_depth # 2 — longest linear sequence of that route
report = scorer.score_batch(targets, max_steps=5)
report.solve_rate # fraction of targets solved
report.mean_bb_coverage
report.to_dataframe()
Excluding the target from stock. A molecule that is itself a catalogue item
is otherwise "solved" in zero steps. Pass exclude_target=True to force a real
disconnection — the target is treated as not purchasable, while its intermediates
still are. Available on both planning and scoring (default off):
planner.plan("CC(=O)Nc1ccccc1O", exclude_target=True)
scorer.score("CC(=O)Nc1ccccc1O", max_steps=5, exclude_target=True)
Two synthesizability metrics
These are conflated in the literature; synomega keeps them apart because they answer different questions.
| Metric | Meaning | Use it for |
|---|---|---|
solved@N / solve_rate |
Binary — does a route of depth ≤ N exist whose leaves are all purchasable? | Comparing against published numbers |
bb_coverage@N |
Continuous — fraction of the best route's leaves that are purchasable | Ranking molecules by how close they are |
bb_coverage matters because most targets are unsolved at realistic step limits.
A 5-step route with 4 of 5 leaves buyable scores 0.8, not 0 — so a near-miss is
distinguishable from a total failure, and a set of molecules can be ranked
rather than merely split into solved/unsolved.
Search algorithms
| Algorithm | Character | When to use |
|---|---|---|
retrostar |
Expands the frontier molecule with the lowest estimated total route cost (Chen et al. 2020) | Default |
mcts |
UCT with greedy rollouts; tolerant of an unreliable top-1 | Weak single-step model |
bfs |
Best-first on g + h |
Baseline / debugging |
All three share the AND-OR graph, the budget, and the route extractor, so their results are directly comparable.
Command line
# one-time: precompute building-block InChIKeys so later loads take seconds
synomega build-stock --catalogue catalogue.smi.gz --out building_blocks.keys.gz
synomega plan --target "CC(=O)Nc1ccccc1" --model path/to/model_run \
--stock building_blocks.keys.gz --stock-is-keys --max-steps 5
synomega score --targets targets.smi --model path/to/model_run \
--stock building_blocks.keys.gz --stock-is-keys \
--max-steps 5 --out report.json
Bring your own single-step model
Any object implementing the SingleStepModel interface plugs into the planner:
from synomega.singlestep import SingleStepModel, Prediction
class MyModel(SingleStepModel):
name = "my-model"
def predict(self, smiles: str, top_k: int = 50) -> list[Prediction]:
# return candidate disconnections, best first
return [Prediction(reactants=("CCO", "CC(=O)O"), score=0.9)]
planner = Planner(MyModel(), stock, algorithm="retrostar")
Built-in backends: TemplateGNN (D-MPNN template classifier, needs [gnn]) and
TemplateRuleModel (pure template matching, no PyTorch).
Design notes
- AND-OR graph, not a tree. A molecule is solved if it is in stock or any of its reactions is solved; a reaction is solved if all its reactants are. Molecules are interned by InChIKey, so an intermediate reached down two branches is one node, expanded once. Cycles are rejected at edge creation.
- Batched expansion. The search pulls a batch of frontier molecules and
issues one
predict_batch, so a GPU-backed model is not left idle. - Caching.
Planner(cache=True)(default) memoizes expansions;cache_path=persists them to SQLite across runs. - Stock membership is by InChIKey, matching a vendor catalogue written by a different toolkit. There is deliberately no Bloom-filter backend — false positives would inflate solve-rate and break comparability with published numbers.
Development
git clone https://github.com/zbc0315/synomega
cd synomega
pip install -e ".[gnn,dev]"
pytest
License
MIT — see LICENSE.
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file synomega-0.3.0.tar.gz.
File metadata
- Download URL: synomega-0.3.0.tar.gz
- Upload date:
- Size: 50.0 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.9.6
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
b65221b7e4ba88e104d9f2dc798bc611f6fbceb00e10d644ec4a8e0281ca3311
|
|
| MD5 |
73e051f0e2638b65a851690b9d19c2c8
|
|
| BLAKE2b-256 |
da8e3a17be4bc9ede1880d861002acfdd815176dbf49e817f89b9c24c3559a92
|
File details
Details for the file synomega-0.3.0-py3-none-any.whl.
File metadata
- Download URL: synomega-0.3.0-py3-none-any.whl
- Upload date:
- Size: 56.9 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.9.6
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
f09ca6bb25193f1ac7753cc088e0e8212926954ab93eb2d58d484f679bb31395
|
|
| MD5 |
b2239aa4e195f7401f94b92a672c0dbb
|
|
| BLAKE2b-256 |
8fa12436d78d2aa5a80ae0531db38af91c4370cab44e04d4321ab9b166f4ce8a
|