Skip to main content

hpo-drift

ci DOI drift check coverage python HPO releases License: MIT

What did a new HPO release change for your phenotype terms — and by how much did your similarity scores move?

The Human Phenotype Ontology is updated regularly. Terms get renamed, obsoleted and merged — and, the part nobody notices, the hierarchy gets new edges. New edges change information content, and information content is what Resnik and Lin similarity are made of. So the same patient set, scored against two HPO releases, gives different numbers even if none of your terms were touched. hpo-drift shows you exactly that, for the term list you actually use, in about three seconds.

The headline result

Feb 2026 → Jun 2026. The example is the HPO-annotated phenotype profile of familial isolated hypoparathyroidism (OMIM:146200): 11 phenotypic-abnormality terms straight from phenotype.hpoa. I chose it because the mechanism is a single edit that any clinician can judge. Then, as a sanity check, hpo-drift cohort ran over the complete phenotype.hpoa corpus — 12 935 disease profiles, no size cutoff: 11 947 have at least one informative pair, 705 are single-term, 283 have root-only pairs, all of them stay in the table with a status — and hpo-drift rank ordered the rankable ones by mean |ΔLin|. This profile is 20th of 11 947 (median profile 0.0013, 99th percentile 0.043); the same edit drives four of the top five (isolated hypoparathyroidism ×2, parathyroid agenesis, pseudohypoparathyroidism type 2).

Lin similarity drift

What happened to this profile, computed with Seco intrinsic IC on the is_a graph under Phenotypic abnormality:

  • One term lost one parent. Hypocalcemic seizures (HP:0002199) was is_a Hypocalcemia (HP:0002901) and is_a Symptomatic seizures; in v2026-06-23 the Hypocalcemia edge is gone (so is the same edge under Hypocalcemic tetany, HP:0003472). No label changed, nothing was obsoleted, the other 10 terms were not touched.
  • Consequence: Hypocalcemic seizuresHypocalcemia went 0.94 → 0.00 — their most-informative common ancestor is now the root. Hypocalcemic seizuresHyperphosphatemia 0.59 → 0.00, ↔ Decreased circulating PTH level 0.21 → 0.00. Hypocalcemia itself, having lost its last child, became a leaf: IC 0.888 → 1.000, which nudges its pairs with Hyperphosphatemia (0.63 → 0.59) and PTH level.
  • 14 informative pairs, all 14 moved, 3 of them by more than 0.1; 41 pairs share only the root (ROOT_ONLY, 0 → 0). IC moved by more than 0.01 for 1 of 11 terms. Mean |ΔLin| 0.129; the median disease profile: 0.0013.

Information-content drift

Whether the edit is right is an ontology-design question (a seizure caused by hypocalcemia is arguably not a kind of hypocalcemia). What is not a matter of opinion: a phenotype-similarity score between two terms that share half their name went from near-identical to zero, without changing the input phenotype profile. Pin the release tag in Methods, match on IDs, and report how much the numbers depend on the release. hpo-drift gives you that sentence with real figures; hpo-drift cohort tells you where your disease of interest sits.

Drift across all disease profiles

rank disease profile retained terms informative pairs changed mean |ΔLin|
1 Hypoparathyroidism, familial isolated 2 (OMIM:618883) 4 6 6 0.298
2 Deafness, autosomal recessive 37 (OMIM:607821) 4 2 2 0.257
3 Cortisone reductase deficiency 2 (OMIM:614662) 7 4 4 0.243
4 Familial isolated hypoparathyroidism due to agenesis of parathyroid gland (ORPHA:2239) 8 12 12 0.214
5 Pseudohypoparathyroidism type 2 (ORPHA:94090) 12 20 20 0.209
6 Cone-Rod dystrophy, X-linked, 2 (OMIM:300085) 2 1 1 0.207
7 Immunodeficiency 65, susceptibility to viral infections (OMIM:618648) 5 6 6 0.175
8 ACTH-independent macronodular adrenal hyperplasia 2 (OMIM:615954) 12 8 8 0.172
20 Hypoparathyroidism, familial isolated (OMIM:146200) 11 14 14 0.129
50 Chronic granulomatous disease (ORPHA:379) 22 95 95 0.067
median of 11 947 0.0013

Reproduce (the annotation file is pinned: phenotype.hpoa from HPO release v2026-06-23, #version: 2026-06-23, SHA-256 89004f85b253f980ffe84218d2c080665cbf67a57bbb322111d6a2db5eb31dff; the script prints the version and hash of the file it was given):

curl -LO https://github.com/obophenotype/human-phenotype-ontology/releases/download/v2026-06-23/phenotype.hpoa
hpo-drift cohort --hpoa phenotype.hpoa --old v2026-02-16 --new v2026-06-23 --out all_profiles.csv   # every profile, ~30 s
hpo-drift rank all_profiles.csv --metric mean_abs_dlin > ranked.csv                                  # optional

Output as run on 2026-09-03: examples/cohort-v2026-02-16_v2026-06-23.csv (all 12 935 profiles with status, plus a .meta.json recording the annotation file's version and hash) and examples/ranked-v2026-02-16_v2026-06-23.csv. Where familiar syndromes sit: X-linked agammaglobulinemia 0.033, autosomal dominant hyper-IgE syndrome 0.029, cystic fibrosis 0.023, Wiskott–Aldrich 0.020, Kabuki / Noonan / Marfan about 0.001.

30-second start

pip install hpo-drift

# one term per line — HP IDs (recommended) or labels
hpo-drift report --old v2026-02-16 --new v2026-06-23 --terms examples/fih_omim146200_terms.txt
IC: Seco 2004 intrinsic, on the is_a graph under root HP:0000118 (Phenotypic abnormality); N = 18690 → 19120
active terms: 19389 → 19836 (added 469, obsoleted 22, renamed 266)
is_a edges:   +886 / −185

term        label                  status     parents        IC old → new
HP:0002199  Hypocalcemic seizures  unchanged  −HP:0002901    1.000 → 1.000
HP:0002901  Hypocalcemia           unchanged  =              0.888 → 1.000
…
pair                                      Lin old → new   Δ       MICA
Hypocalcemic seizures ↔ Hypocalcemia      0.941 → 0.000   −0.941  HP:0002901 → HP:0000118
Hypocalcemic seizures ↔ Hyperphosphatemia 0.594 → 0.000   −0.594  HP:0003111 → HP:0000118
…

Add --json for a machine-readable report. Releases are pulled from the official GitHub assets of obophenotype/human-phenotype-ontology; any tag like v2026-06-23 works and is cached under ~/.cache/hpo-drift.

Four things it does

1 · report — the drift itself

Per term: status (unchanged / renamed / obsoleted → replacement / merged / missing), label change, parents added or removed, IC before and after. Per pair: Resnik and Lin in both releases, the delta, and the most-informative common ancestor — so you can see why a pair moved. Plus the ontology-wide counts.

2 · lint — hygiene for a term list

hpo-drift lint --release v2026-06-23 --terms my_terms.txt
⚠️ Arthritis: matched by LABEL — labels get renamed; store the ID → HP:0001369
❌ Recurrent infection: label not found (exact match on names/synonyms; the term is 'Recurrent infections')
❌ HP:0002961: OBSOLETE term → replaced_by HP:0010701

Exit code 1 on errors, so it works as a CI gate for a phenotype spreadsheet. Matching is exact on purpose: it reproduces the failure mode of a pipeline that matches by label. Store IDs, not labels: 266 labels changed in this interval alone, and fuzzy resolution (roadmap) must suggest, never auto-map.

3 · cohort and rank — the whole annotation corpus, no cutoffs

hpo-drift cohort --hpoa phenotype.hpoa --old v2026-02-16 --new v2026-06-23 --out all_profiles.csv
hpo-drift rank all_profiles.csv --metric mean_abs_dlin --top 20

cohort computes the drift summary for every disease with at least one positive phenotypic-abnormality annotation in phenotype.hpoa — a profile is the unique HP ids of a disease's rows with aspect P and no NOT qualifier (in v2026-06-23: 12 956 disease ids in the file, 12 935 profiles, 21 ids have only negated or non-P rows; the sidecar .meta.json records both counts). There is no minimum or maximum size: a profile that cannot support pairwise analysis stays in the table with a status — NO_USABLE_TERMS, TERM_ONLY (IC drift only), NO_INFORMATIVE_PAIRS (every pair shares only the root) or RANKABLE. Columns: a mutually exclusive disposition of every raw term (retained, unknown, new_only, missing_new, merged_or_alt, obsolete, out_of_domain — they add up to n_raw_terms), IC changes, pairs split into informative and root-only, pairs changed and pairs moved by > 0.01 / > 0.1, mean and max |ΔLin|. rank is a separate, optional step over that complete table. report follows the same rules for any list you give it: 0, 1 or N terms, root-only pairs marked ROOT_ONLY, terms outside the root marked OUT_OF_DOMAIN with a pointer to --root. IDs and labels are resolved against both releases: a term that exists only in the new release is reported as new-in-new (not silently dropped), a label the two releases map to different ids as ambiguous, an unresolvable token as unknown.

Provenance. Every hp.obo is downloaded atomically, hashed, checked against the SHA-256 the official HPO GitHub release publishes for the asset, and cached only after that check; reports, JSON and cohort sidecars record the SHA-256 of both ontology inputs and of the annotation file.

4 · a monthly GitHub Action

.github/workflows/drift.yml fetches the latest release, compares it with your pinned one (PINNED_HPO) and uploads the report. Add a threshold on lin_delta from the JSON and it becomes a failing check.

How it works

flowchart LR
  A["release tag<br/>v2026-02-16"] -->|hp.obo| C[parse: terms · is_a · alt_id · obsolete]
  B["release tag<br/>v2026-06-23"] -->|hp.obo| C
  C --> D["intrinsic IC<br/>Seco 2004"]
  T[your term list] --> E[resolve IDs / labels]
  E --> F[per-term status · parents · IC Δ]
  D --> G[Resnik / Lin per pair · MICA · Δ]
  F --> R[report · JSON · lint]
  G --> R

IC is intrinsic (Seco et al. 2004: 1 − log(descendants+1)/log(N)), computed on the is_a graph only, with N and descendant counts taken inside the closure of a root — HP:0000118 Phenotypic abnormality by default (--root to change; inheritance, frequency and modifier branches are excluded, and the root's IC is exactly 0). Every report states the method, the root and N. Because this IC depends only on the graph, the drift measured here is caused purely by ontology edits, which is the effect this tool isolates. Annotation-based IC (from phenotype.hpoa) adds a second, independent source of drift and is the next option on the roadmap — then the two can be shown side by side.

Companion tools

Roadmap

--ic annotations · fuzzy label suggestions in lint · Phenopackets v2 export · --pairs file for patient × disease scoring · JOSS paper.

Cite

DOI (all versions): 10.5281/zenodo.22286170 — each release has its own version DOI on that Zenodo page; cite the version you ran (hpo-drift --version).

Soloshenko M. hpo-drift: quantifying the effect of HPO release changes on phenotype-similarity results. 2026. MIT License.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

hpo_drift-0.1.7.tar.gz (27.1 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

hpo_drift-0.1.7-py3-none-any.whl (19.6 kB view details)

Uploaded Python 3

File details

Details for the file hpo_drift-0.1.7.tar.gz.

File metadata

  • Download URL: hpo_drift-0.1.7.tar.gz
  • Upload date:
  • Size: 27.1 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for hpo_drift-0.1.7.tar.gz
Algorithm Hash digest
SHA256 063862c56a504006ccb20b0293d27df12479944913be3230eafb3ea38cbc29a9
MD5 710f28a8db6fcb81107a70e6f64df5dd
BLAKE2b-256 d4becd9c0f9dc35b5820dd236f47ca966c433840ab6897498b373d8ca0cd6eec

See more details on using hashes here.

Provenance

The following attestation bundles were made for hpo_drift-0.1.7.tar.gz:

Publisher: publish.yml on MargoSolo/hpo-drift

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file hpo_drift-0.1.7-py3-none-any.whl.

File metadata

  • Download URL: hpo_drift-0.1.7-py3-none-any.whl
  • Upload date:
  • Size: 19.6 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for hpo_drift-0.1.7-py3-none-any.whl
Algorithm Hash digest
SHA256 cab6e2e95c21451f84dc70ef9f1f9febc5170f4cb5ff182d48aab3fa4d2ceac3
MD5 27cba9c0c60dcc333a8e15edbf8987ab
BLAKE2b-256 817c7d38c8842abe5ef27fcecca3d6cff295163632d2ee945d8e4fe9cec0940d

See more details on using hashes here.

Provenance

The following attestation bundles were made for hpo_drift-0.1.7-py3-none-any.whl:

Publisher: publish.yml on MargoSolo/hpo-drift

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.2.1

2 files

0.2.0

2 files

This release

0.1.7 This release

2 files

0.1.6

2 files

0.1.5

2 files

0.1.4

2 files

0.1.3

2 files

0.1.2

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page