A research-grade pipeline linking PubChem compounds with ClinicalTrials.gov study data
Project description
clinical-data-pipeline
A research-grade pipeline for collecting, normalizing, and linking clinical compound data from PubChem with clinical trial documents from ClinicalTrials.gov.
This repository focuses on reproducible, API-based data collection. It prioritizes official APIs and uses web-derived fallbacks only when PubChem REST payloads do not expose trial IDs for specific compounds.
This package is intended for research data collection and analysis. It does not provide medical advice, clinical recommendations, or regulatory conclusions.
Beginner Quickstart (Copy & Paste)
If this is your first run, use this section first.
Option A: Conda (recommended if you already use conda)
conda create -n clinical-pipeline python=3.11 -y
conda activate clinical-pipeline
pip install clinpipe
Run a small smoke test (first CID + first NCT):
clinpipe collect-ctgov \
--hnid 1856916 \
--limit 1 \
--out out_ctgov_smoke
Option B: uv (recommended if you prefer fast Python tooling)
uv venv .venv --python 3.11
source .venv/bin/activate
uv pip install clinpipe
Run the same smoke test:
clinpipe collect-ctgov \
--hnid 1856916 \
--limit 1 \
--out out_ctgov_smoke
Expected output files:
out_ctgov_smoke/
├─ cid_nct_links.jsonl
├─ compounds.jsonl
└─ studies.jsonl
What this project does
This project provides a minimal but extensible pipeline to:
-
Retrieve clinical-trial–related compounds from PubChem
- Uses the official PubChem PUG REST Classification Nodes API
- Retrieves compound lists (CIDs) from specific classification nodes (HNIDs), such as Clinical Trials and ClinicalTrials.gov
-
Collect compound metadata from PubChem
- Canonical SMILES, InChIKey, IUPAC name
- Synonyms and identifiers via PUG REST
-
Link PubChem compounds to ClinicalTrials.gov
- Extracts NCT IDs from PubChem annotations (PUG-View)
- Uses fallback sources when needed (PUG-View heading lookup, PubChem web clinicaltrials endpoint, optional CT.gov term linking)
- Retrieves full clinical trial documents via the ClinicalTrials.gov v2 API
-
Export analysis-ready datasets
- JSONL outputs for compounds, links, and clinical trial documents
- Designed to be consumed by downstream analysis, modeling, or visualization pipelines
Key design principles
-
Official APIs first
- PubChem PUG REST (Classification Nodes, PUG-View)
- ClinicalTrials.gov v2 API
- PubChem web clinicaltrials endpoint fallback (
/sdq/sphinxql.cgi) when REST payload is incomplete
-
No Selenium/browser automation
-
Reproducible
- Classification nodes (HNID) are stable identifiers
-
Modular
- PubChem-related functionality is organized as a self-contained subpackage
PubChem classification nodes (HNID)
PubChem provides an official API to retrieve identifiers associated with classification nodes:
https://pubchem.ncbi.nlm.nih.gov/rest/pug/classification/hnid/{HNID}/{id_type}/{format}
This project currently supports compound (CID) retrieval from clinical-trial–related nodes.
Clinical trial–related HNIDs used
| HNID | Description |
|---|---|
| 1856916 | Clinical Trials (all sources) |
| 3647573 | ClinicalTrials.gov |
| 3647574 | EU Clinical Trials Register |
| 3647575 | NIPH Clinical Trials Search of Japan |
Quick examples
1) Download PubChem CIDs for clinical trials (HNID-based)
from clinpipe.pubchem.clinical_trials_nodes import download_clinical_trials_cids
results = download_clinical_trials_cids(out_dir="out_hnid")
print("Clinical Trials (all):", len(results["clinical_trials"]))
print("ClinicalTrials.gov only:", len(results["clinicaltrials_gov"]))
This will create files such as:
out_hnid/
├─ clinical_trials_cids.txt
├─ clinicaltrials_gov_cids.txt
├─ eu_register_cids.txt
└─ japan_niph_cids.txt
2) From HNID → CID → ClinicalTrials.gov documents
The example below shows the full pipeline:
- download clinical-trial–related CIDs from PubChem (HNID)
- extract NCT IDs from PubChem annotations (PUG-View)
- retrieve full trial documents from ClinicalTrials.gov
from clinpipe.pubchem import (
PubChemClient,
PubChemClassificationClient,
PubChemPugViewClient,
)
from clinpipe.ctgov import CTGovClient
# Clinical Trials HNID
HNID = 1856916
pubchem = PubChemClient()
class_nodes = PubChemClassificationClient()
pug_view = PubChemPugViewClient()
ctgov = CTGovClient()
# Step 1: HNID → CID list
cids = class_nodes.get_cids(HNID)
print("Total CIDs:", len(cids))
# (optional) limit for a quick test
cids = cids[:10]
# Step 2–3: CID → NCT → CTGov study document
for cid in cids:
nct_ids = pug_view.nct_ids_for_cid(cid)
for nct in nct_ids:
study = ctgov.get_study(nct)
print(cid, nct, study.get("protocolSection", {}).get("identificationModule", {}).get("briefTitle"))
This example demonstrates how the individual modules can be composed into a reproducible, end-to-end data collection pipeline.
This will create files such as:
out_hnid/
├─ clinical_trials_cids.txt
├─ clinicaltrials_gov_cids.txt
├─ eu_register_cids.txt
└─ japan_niph_cids.txt
Package structure
src/clinpipe/
├─ pubchem/
│ ├─ client.py # PUG REST: CID, properties, synonyms
│ ├─ classification_nodes.py # HNID → CID (Classification Nodes API)
│ ├─ clinical_trials_nodes.py # Clinical-trial–related HNID helpers
│ └─ pug_view.py # PUG-View: NCT ID extraction
│ └─ web_fallback/ # Web clinicaltrials endpoint/HTML fallback for NCT IDs
│
├─ ctgov/
│ └─ client.py # ClinicalTrials.gov v2 API
│
├─ pipeline/
│ └─ ... # Dataset builders and linkers
Installation
User Install (PyPI)
pip install clinpipe
or with uv:
uv pip install clinpipe
Quick smoke:
clinpipe collect-ctgov \
--hnid 1856916 \
--limit 1 \
--out out_ctgov_smoke
Expected output:
out_ctgov_smoke/
├─ cid_nct_links.jsonl
├─ compounds.jsonl
└─ studies.jsonl
Development Setup (Repository)
Choose one setup method.
uv (Recommended for fast local setup)
Create environment and install:
uv venv .venv --python 3.11
source .venv/bin/activate
uv pip install -e .
Optional development dependencies:
uv pip install -e ".[dev]"
Run without activating shell state (optional):
uv run python scripts/run_mvp_pipeline.py --hnid 3647573 --out-dir out_mvp --use-ctgov-fallback --resume
Conda (Recommended if you already use conda)
conda create -n clinical-pipeline python=3.11 -y
conda activate clinical-pipeline
pip install -e .
Optional development dependencies:
pip install -e ".[dev]"
venv (Alternative)
python -m venv .venv
source .venv/bin/activate
pip install -e .
Optional development dependencies:
pip install -e ".[dev]"
Documentation
Project documentation is in the docs/ folder.
- English:
- docs/overview.md
- docs/ctgov.md
- docs/pubchem.md
- docs/pipeline.md
- docs/cli.md
- Korean:
- docs/overview.ko.md
- docs/ctgov.ko.md
- docs/pubchem.ko.md
- docs/pipeline.ko.md
- docs/cli.ko.md
CLI usage
A minimal command-line interface is provided for quick, reproducible runs without writing Python code.
Show help
clinpipe --help
Legacy alias (still supported): clinical-data-analyzer
Import alias (transition path): clinpipe (while clinical_data_analyzer remains supported)
Script usage (MVP)
For staged execution:
python scripts/fetch_cids.py --hnid 3647573 --out-dir out_mvp
python scripts/map_cid_to_nct.py --cids-file out_mvp/cids.txt --out-dir out_mvp --use-ctgov-fallback
python scripts/fetch_ctgov_docs.py --links-file out_mvp/cid_nct_links.jsonl --out-path out_mvp/studies.jsonl --resume
python scripts/build_clinical_dataset.py --links-file out_mvp/cid_nct_links.jsonl --studies-file out_mvp/studies.jsonl --out-dir out_mvp/final
One-shot:
python scripts/run_mvp_pipeline.py --hnid 3647573 --out-dir out_mvp --use-ctgov-fallback --resume
Step1-3 only (streaming CID -> NCT -> CTGov docs, progress-friendly):
PYTHONUNBUFFERED=1 conda run -n clinical-pipeline python -u scripts/collect_ctgov_docs.py \
--hnid 3647573 \
--folder-name ctgov_docs_run1 \
--out-root out \
--use-ctgov-fallback \
--resume \
--show-progress \
--progress-every 1
Quick smoke (first CID + first NCT):
PYTHONUNBUFFERED=1 conda run -n clinical-pipeline python -u scripts/collect_ctgov_docs.py \
--hnid 3647573 \
--limit-cids 1 \
--limit-ncts 1 \
--folder-name ctgov_docs_first1 \
--out-root out \
--use-ctgov-fallback \
--show-progress \
--progress-every 1
Scheduled Automation (GitHub Actions)
The repository includes a scheduled workflow:
.github/workflows/ctgov_collect.yml.github/workflows/clinical_compound_table_pages.yml
What it does on each run:
- collect/refresh CTGov docs (
collect_ctgov_docs.py --resume) - build normalized dataset (
build_clinical_dataset.py) - build static table page (
build_studies_table.py) - update persistent data snapshots in repo:
data/ctgov/studies.jsonl(latest)data/ctgov/history/studies_*.jsonl(history, only when changed)data/ctgov/collection_state.json(last collected/changed metadata)
- deploy table page from
docs/datato GitHub Pages
Manual run (Actions UI) supports optional:
hnidlimit_cidslimit_nctsshard_size(500default,0disables shard mode)
Recommended workflow presets:
- smoke check:
limit_cids=200shard_size=200image_size=400x400
- production run:
limit_cids=(empty)shard_size=500(start here; tune by runtime)image_size=400x400
PubChem workflow snapshot outputs:
snapshots/clinical_trials/latest/trials.json(latest)snapshots/clinical_trials/latest/compounds.json(CID-level compound cache)snapshots/clinical_trials/latest/trials_compact.json(trial-only compact rows)snapshots/clinical_trials/history/trials_*.json(timestamped history)snapshots/clinical_trials/history/compounds_*.json(timestamped compound history)snapshots/clinical_trials/history/trials_compact_*.json(timestamped compact history)snapshots/clinical_trials/collection_state.json(last collected/changed metadata, includessource: pubchem)
Local snapshot update after collecting dataset files:
python scripts/update_pubchem_trials_history.py \
--trials-file out/pubchem_trials_dataset_check_v2/trials.json \
--compounds-file out/pubchem_trials_dataset_check_v2/compounds.json \
--trials-compact-file out/pubchem_trials_dataset_check_v2/trials_compact.json \
--state-file snapshots/clinical_trials/collection_state.json \
--latest-file snapshots/clinical_trials/latest/trials.json \
--latest-compounds-file snapshots/clinical_trials/latest/compounds.json \
--latest-trials-compact-file snapshots/clinical_trials/latest/trials_compact.json \
--history-dir snapshots/clinical_trials/history \
--retention-days 365
Shard collection and merge (recommended for large runs):
# shard 1 (first 500 CIDs)
python scripts/export_pubchem_trials_dataset.py \
--hnid 1856916 \
--cid-offset 0 \
--cid-count 500 \
--resume \
--out-dir out/pubchem_trials_shards/s1
# shard 2 (next 500 CIDs)
python scripts/export_pubchem_trials_dataset.py \
--hnid 1856916 \
--cid-offset 500 \
--cid-count 500 \
--resume \
--out-dir out/pubchem_trials_shards/s2
# merge shards
python scripts/merge_pubchem_trials_shards.py \
--shard-dirs out/pubchem_trials_shards/s1,out/pubchem_trials_shards/s2 \
--out-dir out/pubchem_trials_merged
Package CI and Publishing
- CI workflow:
.github/workflows/package_ci.yml- runs focused tests
- builds package (
python -m build) - validates artifacts (
twine check)
- Publish workflow:
.github/workflows/publish_pypi.yml(manual dispatch)repository=testpypiusesTEST_PYPI_API_TOKENrepository=pypiusesPYPI_API_TOKEN
TestPyPI install smoke command:
python -m pip install \
--index-url https://test.pypi.org/simple/ \
--extra-index-url https://pypi.org/simple \
clinpipe
Clinical Compound Table Pages (Tabulator Default)
The repository also includes a dedicated Pages workflow for the latest PubChem clinical trial snapshot:
.github/workflows/clinical_compound_table_pages.yml
What it does on each run:
- read the latest snapshot files from
snapshots/clinical_trials/latest/ - build static HTML table (
scripts/build_pubchem_trials_table.py) - publish
index.html,trials.json,trials.csv,summary.json,cids.txtto GitHub Pages
Manual run (Actions UI) input:
table_mode(default:tabulator, also supportsvanilla,datatables)
Download clinical-trial–related CIDs (HNID)
Download PubChem compound IDs associated with the Clinical Trials classification node:
clinpipe hnid-cids \
--hnid 1856916 \
--out out_hnid/clinical_trials_cids.txt
This uses the official PubChem Classification Nodes API:
/rest/pug/classification/hnid/{HNID}/cids/TXT
End-to-end example: HNID → CID → ClinicalTrials.gov documents
Run a small end-to-end collection for a quick sanity check:
clinpipe collect-ctgov \
--hnid 1856916 \
--limit 10 \
--out out_ctgov
This will:
- retrieve CIDs from the given HNID
- extract NCT IDs from PubChem annotations (PUG-View)
- download full trial documents from ClinicalTrials.gov
Generated files:
out_ctgov/
├─ compounds.jsonl
├─ links.jsonl
└─ studies.jsonl
Contributors
- Young-Mook Kang (Korea Research Institute of Chemical Technology, KRICT)
License
MIT License
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file clinpipe-0.6.2.tar.gz.
File metadata
- Download URL: clinpipe-0.6.2.tar.gz
- Upload date:
- Size: 42.3 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.11.15
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
dd8ed2c2894433e60d6c8d09068944517cdf9bb2203570e215b9581857970d14
|
|
| MD5 |
96c42a927df576b933c2f6b04c160054
|
|
| BLAKE2b-256 |
0c9caf4d5e27b51a6a7100813e8ac652ca89f1e5bdea1bd4ff055f7e6a2e7d76
|
File details
Details for the file clinpipe-0.6.2-py3-none-any.whl.
File metadata
- Download URL: clinpipe-0.6.2-py3-none-any.whl
- Upload date:
- Size: 39.9 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.11.15
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
c933b82decbf9263e851918b8a48a6b8d491a451cc1bf1f8e08c99aaf1ffe97e
|
|
| MD5 |
2d58da584d3816dff50c50c824e2f9a9
|
|
| BLAKE2b-256 |
95ab813c46af50ede34b3903821738ed2b8e80d1dcf79767735bf50cf88d1975
|