Skip to main content

:safety_vest: SAFE

Sequential Attachment-based Fragment Embedding (SAFE) is a novel molecular line notation that represents molecules as an unordered sequence of fragment blocks to improve molecule design using generative models.



Paper | Docs | 🤗 Model | 🤗 Training Dataset



PyPI Conda PyPI - Downloads Conda Code license Data License GitHub Repo stars GitHub Repo stars arXiv

test release code-check doc

Overview of SAFE

SAFE is the deep learning molecular representation. It's an encoding leveraging a peculiarity in the decoding schemes of SMILES, to allow representation of molecules as a contiguous sequence of connected fragments. SAFE strings are valid SMILES strings, and thus are able to preserve the same amount of information. The intuitive representation of molecules as an ordered sequence of connected fragments greatly simplifies the following tasks often encountered in molecular design:

  • de novo design
  • superstructure generation
  • scaffold decoration
  • motif extension
  • linker generation
  • scaffold morphing.

The construction of a SAFE strings requires defining a molecular fragmentation algorithm. By default, we use [BRICS], but any other fragmentation algorithm can be used. The image below illustrates the process of building a SAFE string. The resulting string is a valid SMILES that can be read by datamol or RDKit.


News 🚀

💥 2026/09/03 💥

  1. SAFE 0.2.0 release. A maintenance-focused release. It preserves E/Z and atom stereochemistry across fragmentation, makes strict and permissive decoding behaviour explicit, supports extended ring closures, and updates SAFE-GPT to Transformers 5 without changing the established seeded generation paths. The core notation package is lightweight, while model, training, visualization and Weights & Biases support are independent extras. Sampling gains an optional try_hard quality pass and deterministic handling of linker and pattern constraints. See the complete changelog and the migration guide.

💥 2024/01/15 💥

  1. @IanAWatson has a C++ implementation of SAFE in LillyMol that is quite fast and use a custom fragmentation algorithm. Follow the installation instruction on the repo and checkout the docs of the CLI here: docs/Molecule_Tools/SAFE.md

Installation

SAFE 0.2.0 supports Python 3.11 through 3.14. Add it to a uv-managed project:

uv add safe-mol

Pip and conda-forge remain supported:

pip install safe-mol
mamba install -c conda-forge safe-mol

SAFE's core install contains only encoding, decoding and notation splitting. It supports Mac Intel without PyTorch. Model and training extras require PyTorch 2.5+; official Mac Intel wheels stop at 2.2, so use Linux, Windows or Apple Silicon for that stack. Add safe-mol[model] for SAFETokenizer and SAFEDesign, safe-mol[train] for the model stack plus safe-train, or safe-mol[all] to install every maintained feature. Model APIs retain their top-level imports but load their dependencies only when used. For example:

uv add "safe-mol[model]"
# or: python -m pip install "safe-mol[model]"

Visualization and Weights & Biases remain independently available through safe-mol[viz] and safe-mol[wandb]. SAFE's optional model stack uses Transformers 5. SAFE maintains random, greedy, beam and beam-sampling paths. The constrained beam backend required by model-only linker generation is loaded lazily from a reviewed, commit-pinned Hugging Face repository. RDKit 2026.03 is excluded because of an upstream stereochemistry regression; RDKit 2024.09 through 2025.09 are covered by CI. See the migration guide for details.

For GPU workloads, install the PyTorch build matching your CUDA driver before installing SAFE. You can verify the resulting environment with:

import torch

print(torch.cuda.is_available())

Datasets and Models

Type Name Infos Size Comment
Model datamol-io/safe-gpt 87M params 350M Default model
Training Dataset datamol-io/safe-gpt 1.1B rows 250GB Training dataset
Drug Benchmark Dataset datamol-io/safe-drugs 26 rows 20 kB Benchmarking dataset

Usage

Please refer to the documentation, which contains tutorials for getting started with safe and detailed descriptions of the functions provided, as well as an example of how to get started with SAFE-GPT.

API

We summarize some key functions provided by the safe package below.

Function Description
safe.encode Translates a SMILES string into its corresponding SAFE string.
safe.decode Translates a SAFE string into its corresponding SMILES string. The SAFE decoder just augment RDKit's Chem.MolFromSmiles with an optional correction argument to take care of missing hydrogen bonds.
safe.split Tokenizes a SAFE string to build a generative model.

Examples

Translation between SAFE and SMILES representations

import safe

ibuprofen = "CC(Cc1ccc(cc1)C(C(=O)O)C)C"

# SMILES -> SAFE -> SMILES translation
try:
    ibuprofen_sf = safe.encode(ibuprofen)  # c12ccc3cc1.C3(C)C(=O)O.CC(C)C2
    ibuprofen_smi = safe.decode(ibuprofen_sf, canonical=True)  # CC(C)Cc1ccc(C(C)C(=O)O)cc1
except safe.SAFEEncodeError:
    pass
except safe.SAFEDecodeError:
    pass

ibuprofen_tokens = list(safe.split(ibuprofen_sf))

Training/Finetuning a (new) model

A command line interface is available to train a new model, please run safe-train --help. You can also provide an existing checkpoint to continue training or finetune on you own dataset.

For example:

safe-train --config <path to config> \
    --model-path <path to model> \
    --tokenizer  <path to tokenizer> \
    --dataset <path to dataset> \
    --num_labels 9 \
    --torch_compile True \
    --optim "adamw_torch" \
    --learning_rate 1e-5 \
    --prop_loss_coeff 1e-3 \
    --gradient_accumulation_steps 1 \
    --output_dir "<path to outputdir>" \
    --max_steps 5

References

If you use this repository, please cite the following related paper:

@misc{noutahi2023gotta,
      title={Gotta be SAFE: A New Framework for Molecular Design},
      author={Emmanuel Noutahi and Cristian Gabellini and Michael Craig and Jonathan S. C Lim and Prudencio Tossou},
      year={2023},
      eprint={2310.10773},
      archivePrefix={arXiv},
      primaryClass={cs.LG}
}

License

The training dataset is licensed under CC BY 4.0. See DATA_LICENSE for details. This code base is licensed under the Apache-2.0 license. See LICENSE for details.

Note that the model weights of SAFE-GPT are exclusively licensed for research purposes (CC BY-NC 4.0).

These licences apply to separate materials. The Python package does not redistribute the model weights.

Development lifecycle

Setup dev environment

uv sync --all-extras

This creates an isolated .venv with the training, visualisation, reporting, test, documentation and development extras. env.yml remains available when a Conda environment is required.

Tests

You can run tests locally with:

uv run python -m pytest -m "not integration"
uv run python -m pytest -m integration --no-cov

The integration command validates the published SAFE-GPT model and executes the maintained tutorials. GitHub Actions runs the same command. Use uv run python -m pytest -m notebook --no-cov when iterating on tutorials only.

Releasing

Release maintainers: see the manual release guide.

Metadata

Release files for safe-mol 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for safe-mol 0.2.0
File Size Uploaded
safe_mol-0.2.0.tar.gz 534.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for safe-mol 0.2.0
File Interpreter ABI Platform
safe_mol-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 605.6 kB

Release files / safe_mol-0.2.0.tar.gz

Download URL safe_mol-0.2.0.tar.gz
Size 534.7 kB
Tags Source
SHA-256 checksum
How to use checksums
f529f66a6a10ff872b1519208cf4b95928425704e4815fd091120879f692e209
BLAKE2b-256 checksum
How to use checksums
ad206b90c7e45140f7df74f4ff195d197b13afa6916a77de7f28b5afb73f1a2f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.8 {"installer":{"name":"uv","version":"0.12.8","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 3, 2026.

Transparency log

Release files / safe_mol-0.2.0-py3-none-any.whl

Download URL safe_mol-0.2.0-py3-none-any.whl
Size 70.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
e99e10a7c8829139ebfe632875df5e705e5dabbb00f98897572f4f38644444c4
BLAKE2b-256 checksum
How to use checksums
3921213d1afb290f7684a586d82fac0c0d7d69bd798d2c86ed607647604f5978
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.8 {"installer":{"name":"uv","version":"0.12.8","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 3, 2026.

Transparency log

Release history Release notifications | RSS feed

0.2.1

2 release files

This release

0.2.0 This release

2 release files

0.1.14

2 release files

0.1.13

2 release files

0.1.12

2 release files

0.1.11

2 release files

0.1.10

2 release files

0.1.9

2 release files

0.1.8

2 release files

0.1.7

2 release files

0.1.6

2 release files

0.1.5

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page