ChemDataCheck
Quality checks for molecular datasets — like pytest for the data behind chemistry and molecular-machine-learning projects.
One command to sanity-check a chemistry dataset before you train, publish, or trust it:
python -m pip install chemdatacheck
chemdatacheck dataset.csv
For datasets with labels and predefined train/test splits:
chemdatacheck dataset.csv --label-col activity --split-col split \
--format html --output chemdatacheck-report.html
It checks for:
- valid molecules
- duplicates
- stereochemical collisions
- salt/tautomer
- ambiguity
- split leakage
- target shift
- chemical-space bias
- etc.
Every warning comes with row IDs, evidence, and an actionable recommendation, plus a 0–100 dataset quality score. See an example HTML report, the scientific validation, and the interpretation limits. Questions and field reports are welcome in the field-report form.
Install
ChemDataCheck requires Python 3.10 or newer and is tested with Python 3.10–3.12. Install it from PyPI in an isolated environment. On macOS or Linux:
python3.11 -m venv .venv
source .venv/bin/activate
python -m pip install chemdatacheck
On Windows PowerShell, create and activate the environment with:
py -3.11 -m venv .venv
.venv\Scripts\Activate.ps1
python -m pip install chemdatacheck
To add colored terminal output, install chemdatacheck[pretty] instead.
Required dependencies—including RDKit, pandas, SciPy, PyArrow, and
openpyxl—are declared in pyproject.toml and installed automatically by pip.
If a compatible RDKit wheel is not available for your Python version or platform, use a Conda environment instead:
conda create -n chemdatacheck -c conda-forge python=3.11 rdkit pip
conda activate chemdatacheck
python -m pip install chemdatacheck
Contributors who need an editable installation, tests, coverage reporting, or
the notebooks should follow CONTRIBUTING.md.
Inputs
CSV/TSV, SDF/SD, Parquet, Excel, JSON-lines. The SMILES column is autodetected
(smiles, canonical_smiles, ...); alternatively, you can pass --smiles-col. Splits (training set / test set)
can be passed via a column (--split-col) or by passing two files (train.csv test.csv).
Exit codes (pytest-like)
0clean1warnings2errors
What it checks
- Chemical integrity: invalid SMILES, valence errors, aromaticity/kekulization, impossible charges, disconnected components (salts/mixtures), isotopes, radicals, unspecified stereocenters, tautomer ambiguity.
- Dataset duplicates: exact, canonical, stereochemical collisions, salt/solvate duplicates, tautomer duplicates, near-duplicates (Morgan fingerprints, similarity measured with Tanimoto distance).
- Split leakage: 2D-identity leakage, analog leakage (Tc ≥ 0.6), near-duplicates across splits, scaffold overlap, suspiciously-easy-split heuristic.
- ML readiness: duplicated/conflicting measurements, target distribution shift, split-predicts-label leakage suspects, label outliers.
- Chemical space: rare elements, rare/promiscuous functional groups, unusual ring systems, representation bias, applicability-domain gaps.
Every finding carries row IDs, evidence examples, and a recommendation, e.g.
train_row ↔ test_row pairs with Tanimoto and shared scaffold for leakage.
To run leakage checks, ChemDataCheck must know which rows are training data and
which are held out for testing or validation. Common split values such as
train, test, and valid are recognized automatically. If your dataset uses
different names, map them explicitly; for example:
chemdatacheck dataset.csv --split-col partition \
--train-value development --test-value external
For numbered cross-validation folds, you only need to identify the held-out
fold. This treats fold 0 as test data and every other observed fold as
training data:
chemdatacheck dataset.csv --split-col fold --test-value 0
If ChemDataCheck cannot form both a non-empty training group and a non-empty test group, it reports a warning and does not run the leakage checks. If only some split values are mapped, it warns that the remaining rows were excluded from those checks.
The quality score is a prioritization heuristic, not a validated scientific metric.
The quality score starts from 100, and points are subtracted for each bad quality finding.
Related findings are overlap-capped so the same underlying
invalid structure, duplicate, or leakage issue is not fully deducted multiple
times. JSON reports record tool/RDKit versions, settings, resolved columns, and
SHA-256 input hashes for reproducibility.
When a large-dataset check uses sampling or a bounded candidate search, JSON
reports expose it in meta.approximations and in the finding's metadata.
What ChemDataCheck does not do
ChemDataCheck is an evidence-producing audit and triage tool, not an automatic certification or data-cleaning system. It does not silently rewrite structures, prove that a model will generalize, or replace assay review and domain expertise. A flagged salt, tautomer, scaffold overlap, or repeated measurement may be scientifically appropriate; the row-level evidence is more important than the summary score.
Learn
- Tutorial:
docs/tutorial.md— from install to CI-gated audits in ~20 minutes. - Scientific validation:
docs/validation.md— public ML datasets plus fixed ChEMBL/PubChem samples, seeded defects, threshold sensitivity, false-positive guidance, and runtime through 40k+ molecules. - Examples:
examples/— three executed notebooks, all data generated inline (no downloads):01_quickstart.ipynb— CLI + Python API on a deliberately dirty dataset.02_leakage_splits.ipynb— identity vs analog vs scaffold leakage, and fixes.03_curation_case_study.ipynb— full curation loop (39 → 80) with HTML report.
Metadata
Release files for chemdatacheck 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| chemdatacheck-0.1.0.tar.gz | 41.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| chemdatacheck-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 81.8 kB
Release files / chemdatacheck-0.1.0.tar.gz
| Download URL | chemdatacheck-0.1.0.tar.gz |
|---|---|
| Size | 41.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
5175abfe59bff73d2076f1adee80ea86249525d7c62aba1dbdd5cc5c32956493
|
|
BLAKE2b-256 checksum How to use checksums |
ef839cc80f96d8b26fc0da897c196e634728f83a69ad8b2f7c0fb959effcc5f1
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.
Transparency logRelease files / chemdatacheck-0.1.0-py3-none-any.whl
| Download URL | chemdatacheck-0.1.0-py3-none-any.whl |
|---|---|
| Size | 40.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
93480d6f7e54c2692f9ee1737fc594dd9c80ab8ddf8ddd6dcff87fd3cded7ba2
|
|
BLAKE2b-256 checksum How to use checksums |
34fb2df509d748dac05e18c856abd40828629b6ba101274cb232b32179967a69
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.
Transparency log