Skip to main content

SynOmega

PyPI Python License: MIT

A retrosynthesis toolkit that turns a target molecule into synthesis routes and a continuous synthesizability score. Three decoupled layers:

synthesizability   is this target reachable from purchasable material, in N steps?
     ↑
search             Retro* / MCTS / best-first over an AND-OR graph
     ↑
single-step        product SMILES -> ranked reactant candidates

The layers meet at a deliberately narrow interface — a single-step backend only implements predict(smiles, top_k) -> [Prediction] — so the planner and the scorer do not care whether predictions come from a graph neural network, a transformer, or plain template matching.

Installation

pip install synomega           # core: rdkit + numpy
pip install "synomega[gnn]"    # + the D-MPNN neural single-step backend (torch)

The neural backend is an optional extra on purpose: the template-rule backend runs anywhere, with no GPU and no PyTorch. Requires Python ≥ 3.10.

Zero-config quickstart

pip install synomega ships only code. The first time you ask for the default model or stock, synomega downloads them (a few hundred MB) into ~/.cache/synomega — the same way spaCy and HuggingFace fetch models. So this works out of the box:

import synomega

planner = synomega.load_default_planner()          # downloads model + stock once
print(planner.plan("CC(=O)Nc1ccccc1O").best_route.describe())

from synomega import SynthesizabilityScorer
score = SynthesizabilityScorer(planner).score("CC(=O)Nc1ccccc1O", max_steps=5)
print(score.bb_coverage, score.min_steps)

Or pre-fetch from the command line, then use the CLI with no --model/--stock:

synomega download                                   # cache the default assets
synomega plan --target "CC(=O)Nc1ccccc1O" --max-steps 5

Download mirrors. Assets are hosted on both a USTC GitLab registry (fast in China) and (soon) GitHub. SynOmega auto-selects the faster reachable one by latency; override with SYNOMEGA_MIRROR=ustc|github or point SYNOMEGA_ASSETS_BASE=<url> at your own mirror. Change the cache location with SYNOMEGA_CACHE.

Quick start

from synomega import Planner, SynthesizabilityScorer
from synomega.singlestep import TemplateGNN
from synomega.stock import InMemoryStock

model   = TemplateGNN.from_pretrained("path/to/model_run")   # a trained checkpoint
stock   = InMemoryStock.from_keys_file("building_blocks.keys.gz")
planner = Planner(model, stock, algorithm="retrostar")

result = planner.plan("CC(=O)Nc1ccccc1", max_depth=5, time_limit=60)
print(result.solved)
print(result.best_route.describe())
target: CC(=O)Nc1ccccc1
solved: True  steps: 2  depth: 2  bb_coverage: 1.00
  [1] CC(=O)O.Nc1ccccc1>>CC(=O)Nc1ccccc1  (score=0.4348)
  [2] O=[N+]([O-])c1ccccc1>>Nc1ccccc1     (score=0.2174)

Synthesizability scoring

scorer = SynthesizabilityScorer(planner)

r = scorer.score("CC(=O)Nc1ccccc1", max_steps=5)
r.solved            # True — a complete route to purchasable material exists
r.bb_coverage       # 1.0 — fraction of the best route's leaves that are buyable
r.min_steps         # 2  — reactions in the shortest solved route
r.min_route_depth   # 2  — longest linear sequence of that route

report = scorer.score_batch(targets, max_steps=5)
report.solve_rate         # fraction of targets solved
report.mean_bb_coverage
report.to_dataframe()

Excluding the target from stock. A molecule that is itself a catalogue item is otherwise "solved" in zero steps. Pass exclude_target=True to force a real disconnection — the target is treated as not purchasable, while its intermediates still are. Available on both planning and scoring (default off):

planner.plan("CC(=O)Nc1ccccc1O", exclude_target=True)
scorer.score("CC(=O)Nc1ccccc1O", max_steps=5, exclude_target=True)

Reaction-plausibility filtering

Every single-step prediction is screened by a mapping-free dual-tower reaction- plausibility model: two shared-encoder D-MPNN towers embed the candidate reactants and the target product separately (no atom mapping needed), and score how likely those reactants actually give the product. Candidates below a threshold are dropped — the filter only removes wrong disconnections, it never re-ranks the survivors. Because search and synthesizability both expand through the single-step model, this screens every single-step prediction in the system.

It is off by default: benchmarks (scripts/BENCHMARKS.md) show it does not improve top-k retrieval of the recorded reaction (it is marginally negative, −0.2…−0.9 pp) and adds latency (×1.7 GPU / ×4.6 CPU). Enable it when you want the candidate list pruned of implausible disconnections rather than maximal recall.

# off by default:
planner = synomega.load_default_planner()

# enable / tune:
planner = synomega.load_default_planner(plausibility=True)
planner = synomega.load_default_planner(plausibility=True, plausibility_threshold=0.5)

# bring your own single-step model + explicit scorer:
from synomega.plausibility import PlausibilityScorer
scorer = PlausibilityScorer.default(device="cpu")
planner = Planner(model, stock, plausibility=scorer, plausibility_threshold=0.4)

When enabled, each surviving prediction carries its raw plausibility in prediction.meta["plausibility"].

Two synthesizability metrics

These are conflated in the literature; SynOmega keeps them apart because they answer different questions.

Metric Meaning Use it for
solved@N / solve_rate Binary — does a route of depth ≤ N exist whose leaves are all purchasable? Comparing against published numbers
bb_coverage@N Continuous — fraction of the best route's leaves that are purchasable Ranking molecules by how close they are

bb_coverage matters because most targets are unsolved at realistic step limits. A 5-step route with 4 of 5 leaves buyable scores 0.8, not 0 — so a near-miss is distinguishable from a total failure, and a set of molecules can be ranked rather than merely split into solved/unsolved.

Search algorithms

Algorithm Character When to use
retrostar Expands the frontier molecule with the lowest estimated total route cost (Chen et al. 2020) Default
mcts UCT with greedy rollouts; tolerant of an unreliable top-1 Weak single-step model
bfs Best-first on g + h Baseline / debugging

All three share the AND-OR graph, the budget, and the route extractor, so their results are directly comparable.

Command line

# one-time: precompute building-block InChIKeys so later loads take seconds
synomega build-stock --catalogue catalogue.smi.gz --out building_blocks.keys.gz

synomega plan  --target "CC(=O)Nc1ccccc1" --model path/to/model_run \
               --stock building_blocks.keys.gz --stock-is-keys --max-steps 5

synomega score --targets targets.smi --model path/to/model_run \
               --stock building_blocks.keys.gz --stock-is-keys \
               --max-steps 5 --out report.json

Bring your own single-step model

Any object implementing the SingleStepModel interface plugs into the planner:

from synomega.singlestep import SingleStepModel, Prediction

class MyModel(SingleStepModel):
    name = "my-model"
    def predict(self, smiles: str, top_k: int = 50) -> list[Prediction]:
        # return candidate disconnections, best first
        return [Prediction(reactants=("CCO", "CC(=O)O"), score=0.9)]

planner = Planner(MyModel(), stock, algorithm="retrostar")

Built-in backends: TemplateGNN (D-MPNN template classifier, needs [gnn]) and TemplateRuleModel (pure template matching, no PyTorch).

Design notes

  • AND-OR graph, not a tree. A molecule is solved if it is in stock or any of its reactions is solved; a reaction is solved if all its reactants are. Molecules are interned by InChIKey, so an intermediate reached down two branches is one node, expanded once. Cycles are rejected at edge creation.
  • Batched expansion. The search pulls a batch of frontier molecules and issues one predict_batch, so a GPU-backed model is not left idle.
  • Caching. Planner(cache=True) (default) memoizes expansions; cache_path= persists them to SQLite across runs.
  • Stock membership is by InChIKey, matching a vendor catalogue written by a different toolkit. There is deliberately no Bloom-filter backend — false positives would inflate solve-rate and break comparability with published numbers.

Development

git clone https://github.com/zbc0315/synomega
cd synomega
pip install -e ".[gnn,dev]"
pytest

License

MIT — see LICENSE.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

synomega-0.4.1.tar.gz (56.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

synomega-0.4.1-py3-none-any.whl (64.8 kB view details)

Uploaded Python 3

File details

Details for the file synomega-0.4.1.tar.gz.

File metadata

  • Download URL: synomega-0.4.1.tar.gz
  • Upload date:
  • Size: 56.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.9.6

File hashes

Hashes for synomega-0.4.1.tar.gz
Algorithm Hash digest
SHA256 df4b4611e1cc8bb1b9a2deac710934e4b8dd048dca2a3372222c7b5732c3e690
MD5 69830c44e522c31a3a2c4a3e2afb2935
BLAKE2b-256 5c3289a848571229d2537e1676d0f495b16fbc1daebfb3ebef759f5a04fb9394

See more details on using hashes here.

File details

Details for the file synomega-0.4.1-py3-none-any.whl.

File metadata

  • Download URL: synomega-0.4.1-py3-none-any.whl
  • Upload date:
  • Size: 64.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.9.6

File hashes

Hashes for synomega-0.4.1-py3-none-any.whl
Algorithm Hash digest
SHA256 c3706d97f1e40d1376899874a30e11d52d7840b6e35829c60dd239ca7ab4d71d
MD5 871c4be2ceb04b775242e4c7b07d4985
BLAKE2b-256 c1f01a65e05722b2fad41c599e10e309da36e284d1a491328c613a61cabbee50

See more details on using hashes here.

Release history Release notifications | RSS feed

0.10.0

2 files

0.9.4

2 files

0.9.3

2 files

0.9.2

2 files

0.9.1

2 files

0.9.0

2 files

0.8.0

2 files

0.7.0

2 files

0.6.0

2 files

0.5.0

2 files

This release

0.4.1 This release

2 files

0.4.0

2 files

0.3.0

2 files

0.2.0

2 files

0.1.1

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page