Skip to main content

HARMONSMILE: Harmonize SMILES Strings for Cheminformatics and Machine Learning

License: LGPL v3 Version PyPI Python Docs


Description

HARMONSMILE solves a common problem in cheminformatics: SMILES strings for the same molecule look different depending on the source (PubChem, ChEMBL, COCONUT, in-house databases). This inconsistency breaks comparisons, deduplication, and machine learning pipelines that expect a uniform molecular representation.

It is intended for computational chemists, cheminformatics researchers, ML practitioners preparing molecular datasets, and maintainers integrating PubChem, ChEMBL, and in-house sources.


Purpose

The primary objective of HARMONSMILE is to automate the preparation of molecular datasets for cheminformatics workflows and phase 1 machine learning applications within the computational drug discovery pipeline.

The platform enables:

  • Data Harmonization: Standardizes SMILES strings to a consistent format - canonical + isomeric + Kekulized - ensuring that the same molecule is represented identically across different datasets and sources. It follows the RDKit convention for canonicalization, which is widely adopted in the cheminformatics community.

SMILES Column Contract

HARMONSMILE v0.3.0 exposes three universal SMILES representations in pipeline outputs:

  • SMILES: source/input SMILES. This is the value used as input for RDKit canonicalization and lab harmonization.
  • SMILES_RDKit: RDKit canonical + isomeric + Kekulized representation, produced by RDKitStandardizer.to_iso_kek(SMILES). This column is preserved for compatibility with v0.2.5. It is not a full chemical harmonization layer and does not intentionally desalt, neutralize, reionize, or canonicalize tautomers.
  • SMILES_Harmonized: v0.3.0 lab-harmonized representation, produced by RDKitStandardizer.to_lab_harmonized(SMILES). It uses an RDKit-native MolStandardize policy: normalization, largest-fragment selection, allowed-elements validation, uncharging, reionization, optional tautomer canonicalization, and final canonical/isomeric/Kekule serialization.

SMILES_Harmonized is intended for database harmonization, deduplication, and cross-source matching. It is not guaranteed to represent the most stable, most abundant, pH-specific, or biologically active tautomer. HARMONSMILE does not export tautomer ensembles; it stores one canonical harmonized representation per input row. Source traceability is preserved through SMILES and SMILES_RDKit.

Rows are not dropped when harmonization fails. Instead, the result is carried in:

  • SMILES_Harmonization_Status: status returned by the harmonization engine. Expected values include ok, missing_smiles, invalid_smiles, disallowed_elements, tautomer_limit_exceeded when available in the installed RDKit, and harmonization_failed.
  • SMILES_Harmonization_Error: short auditable error message when harmonization fails; empty/None for successful rows.

Some sources may also provide source-specific SMILES columns. For example, ConnectivitySMILES is a PubChem-provided connectivity SMILES column preserved when available; it is not generated by HARMONSMILE and is not present for all sources.


Installation

For package users:

Create and activate a Python environment:

conda create -n harmonsmile_env python=3.11
conda activate harmonsmile_env

Install HARMONSMILE from PyPI:

pip install harmonsmile

For contributors/developers:

Clone the repository:

git clone https://github.com/NanoBiostructuresRG/harmonsmile.git
cd harmonsmile

Create and activate the development environment:

conda env create -f environment.yml
conda activate harmonsmile_env

Install HARMONSMILE in editable mode with development dependencies:

python -m pip install -e .[dev]

RDKit is a required runtime dependency (rdkit>=2022.09). For package users, it is declared in pyproject.toml and installed through the package dependency resolver. For contributors, environment.yml preinstalls RDKit from conda-forge for a stable local scientific stack.


Quick Start

Python API

Standardize a single SMILES string:

from harmonsmile import RDKitStandardizer

std = RDKitStandardizer()
print(std.to_iso_kek("c1ccccc1"))    # canonical + isomeric + Kekulized
print(std.to_conn_kek("c1ccccc1"))   # canonical + connectivity-only + Kekulized

Fetch properties from PubChem and harmonize:

from harmonsmile import PubChemIngest, PubChemConfig, save_table

cfg = PubChemConfig(
    input_path="examples/example_pubchem.csv",   # requires: id, PubChem CID
)
df = PubChemIngest(cfg).run()
save_table(df, "results/example_pubchem_harmonized.csv")

Fetch properties from ChEMBL and harmonize:

from harmonsmile import ChEMBLIngest, ChEMBLConfig, save_table

cfg = ChEMBLConfig(
    input_path="examples/example_chembl.csv",    # requires: id, ChEMBL ID
)
df = ChEMBLIngest(cfg).run()
save_table(df, "results/example_chembl_harmonized.csv")

Harmonize any file with a SMILES column (COCONUT, in-house, etc.):

from harmonsmile import SMILESPrep, SMILESConfig, save_table

cfg = SMILESConfig(
    input_path="examples/example_smiles.csv",
    smiles_col="SMILES",                      # any column name
)
df = SMILESPrep(cfg).run()
save_table(df, "results/example_smiles_harmonized.csv")

Command-Line Interface

# PubChem pipeline
harmonsmile --pubchem-in examples/database1.csv --pubchem-out results/database1_harmonized.csv

# SMILES pipeline (COCONUT, independent, etc.)
harmonsmile --smiles-in examples/database2.csv --smiles-col canonical_smiles \
            --smiles-out results/database2_harmonized.csv

# Both pipelines in one run
harmonsmile \
  --pubchem-in examples/database1.csv --pubchem-out results/database1_harmonized.csv \
  --smiles-in  examples/database2.csv --smiles-col  canonical_smiles \
  --smiles-out results/database2_harmonized.csv

# Single Entry - fetch one compound by ID
harmonsmile --pubchem-cid 2723949
harmonsmile --chembl-id CHEMBL294199

# Check version
harmonsmile --version

Also available as a Python module:

python -m harmonsmile --pubchem-in examples/database1.csv --pubchem-out results/out.csv

Pipelines

Pipeline Config Source Input API
PubChemIngest PubChemConfig PubChem CSV with PubChem CID column REST (public)
ChEMBLIngest ChEMBLConfig ChEMBL CSV with ChEMBL ID column REST (public)
SMILESPrep SMILESConfig Any CSV/Excel with any SMILES column Local file

All pipelines preserve the source SMILES, append SMILES_RDKit, and append the lab harmonization columns SMILES_Harmonized, SMILES_Harmonization_Status, and SMILES_Harmonization_Error. Pipeline .run() methods return a pandas.DataFrame and do not write files. Use save_table(df, path) to persist results from Python, or use the CLI --*-out options.


Input Format

Pipeline Required columns
PubChemIngest id (optional), PubChem CID
ChEMBLIngest id (optional), ChEMBL ID
SMILESPrep id (optional), <smiles_col> (any name)

Supported file formats: CSV, TSV, XLSX, XLS.


Roadmap

  • v0.3.0 - Lab-harmonized SMILES contract/layer for database harmonization, deduplication, and cross-source matching.

Development

Project Structure

HARMONSMILE/
|-- harmonsmile/
|   |-- __init__.py        # Public API
|   |-- __main__.py        # python -m harmonsmile entry point
|   |-- _cli.py            # CLI implementation
|   |-- chembl.py          # ChEMBL REST client
|   |-- config.py          # PubChemConfig, ChEMBLConfig, SMILESConfig dataclasses
|   |-- io.py              # Table I/O utilities
|   |-- pipelines.py       # PubChemIngest, ChEMBLIngest, SMILESPrep
|   |-- pubchem.py         # PubChem REST client
|   |-- standardize.py     # RDKitStandardizer
|   `-- version.py         # Package version metadata
|-- tests/                 # Unit test suite (pytest) - 146 tests
|-- examples/              # Example scripts and datasets
|-- pyproject.toml
|-- environment.yml
|-- mkdocs.yml
|-- requirements-dev.txt
|-- CHANGELOG.md
|-- CITATION.cff
|-- CODE_OF_CONDUCT.md
|-- CONTRIBUTING.md
|-- COPYING
|-- COPYING.LESSER
|-- LICENSE
`-- README.md

Running Tests

python -m pytest tests -p no:cacheprovider --basetemp .pytest_tmp

Contributing

Contributions are welcome. Please open an issue before submitting a pull request. Follow the existing code style: NumPy-style docstrings, type hints, and SPDX license headers in all source files.

See CONTRIBUTING.md for full guidelines. Please also read our Code of Conduct.


Citation

If you use HARMONSMILE in your research, please cite it using the metadata in CITATION.cff or the format below:

Contreras-Torres, F. F. (2026). HARMONSMILE: Harmonize SMILES Strings for
Cheminformatics and Machine Learning. Zenodo. https://doi.org/10.5281/zenodo.20275498

Author

Developed by Flavio F. Contreras-Torres (Tecnologico de Monterrey) Monterrey, Mexico - May 2026


License

This project is licensed under the terms of the GNU Lesser General Public License v3.0 or later. SPDX identifier: LGPL-3.0-or-later.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

harmonsmile-0.3.0.tar.gz (52.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

harmonsmile-0.3.0-py3-none-any.whl (38.8 kB view details)

Uploaded Python 3

File details

Details for the file harmonsmile-0.3.0.tar.gz.

File metadata

  • Download URL: harmonsmile-0.3.0.tar.gz
  • Upload date:
  • Size: 52.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for harmonsmile-0.3.0.tar.gz
Algorithm Hash digest
SHA256 1009b4b4ee166f0646c80d87232cfec50152c3f0457363f151e6501994cf451e
MD5 8803e66d43d4a902ec2b5873a31cd792
BLAKE2b-256 5287f98c670e4a611996322df9c8f74534b41efb64775550aa98a313dccff291

See more details on using hashes here.

Provenance

The following attestation bundles were made for harmonsmile-0.3.0.tar.gz:

Publisher: publish-to-pypi.yml on NanoBiostructuresRG/harmonsmile

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file harmonsmile-0.3.0-py3-none-any.whl.

File metadata

  • Download URL: harmonsmile-0.3.0-py3-none-any.whl
  • Upload date:
  • Size: 38.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for harmonsmile-0.3.0-py3-none-any.whl
Algorithm Hash digest
SHA256 384987dfa7a19fd6c60d602f5ce96a2481a743f2e385c46842a13826273935ea
MD5 56c43984f6cf0b2c8ae8d1828be24810
BLAKE2b-256 006a62e6fb56f6998c9b360ecbfacee7acc4f485b0068ddcc7a0a859a325570e

See more details on using hashes here.

Provenance

The following attestation bundles were made for harmonsmile-0.3.0-py3-none-any.whl:

Publisher: publish-to-pypi.yml on NanoBiostructuresRG/harmonsmile

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.3.3

2 files

0.3.2

2 files

0.3.1

2 files

This release

0.3.0 This release

2 files

0.2.5

2 files

0.2.4

2 files

0.2.3

2 files

0.2.2

2 files

0.2.1

2 files

0.2.0

2 files

0.1.3

2 files

0.1.2

2 files

0.1.1

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page