Crosswalk ICD-10-GM codes across annual editions using official BfArM transition tables.
Project description
icd10gm-crosswalk
Map ICD-10-GM diagnosis codes across annual editions using the official BfArM Umsteiger (transition) tables — chained over any year range, with the automatic-vs-manual transition flag preserved and the ambiguity from chained splits and merges reported honestly rather than swept under the rug.
ICD-10-GM (the German modification of ICD-10) is re-issued every year. A code valid in 2017 may be unchanged, renamed 1:1, split into finer children, merged into a broader code, or deleted by 2024. Getting this right matters whenever you compare codings made under different editions — e.g. scoring a 2017-annotated corpus against a 2024 prediction. The official source of truth is BfArM's yearly Umsteiger tables; there is no direct multi-year table, so 2017→2024 must be chained 2017→2018→…→2024, and a naive verbatim diff silently misses splits whose parent survives as a category string. This library does the chaining.
Python counterpart to the R
ICD10gmpackage. The R package is excellent and more comprehensive; this one is focused on the multi-year crosswalk problem, ships zero runtime dependencies, and — importantly — never redistributes BfArM data (see Data & licensing).
Features
- Multi-year chaining —
map("J45.0", 2017, 2024)walks every intermediate edition, not a single before/after diff. - Honest classification — every mapping is
identity,one_to_one,split(1:n),merge(n:1), ordeleted. - Per-step provenance — the automatic-vs-manual flag from each Umsteiger row
is preserved and aggregated into a single
needs_manual_reviewsignal. - Ambiguity, surfaced — chained splits that never re-merge are flagged
(
ambiguous), and arecommend()helper picks a single representative code (parent fallback) when you must collapse a 1:n mapping. - Reads BfArM's real layout — loose
.txt, a year ZIP, or a whole directory, including the nested-ZIP packaging used since the 2022 edition. - Combined codes & Kreuz-Stern notation — maps the components of a compound
diagnosis and carries the Kreuz (
†/+), Stern (*), and!role markers through the mapping, optionally validating them against the ClaML systematik. - Zero runtime dependencies, fully typed (
py.typed), tested, with a CLI.
Installation
This repo uses uv. Until it is published to PyPI, install from source:
git clone https://github.com/JohannKaspar/icd10gm-crosswalk
cd icd10gm-crosswalk
uv sync
Once released, it will be installable directly:
pip install icd10gm-crosswalk # or: uv add icd10gm-crosswalk (after PyPI release)
Quick start
from icd10gm_crosswalk import Crosswalk
# Point at a directory of BfArM year ZIPs (see "Data & licensing" below).
cw = Crosswalk.from_source("~/icd10gm-zips")
res = cw.map("J45.0", 2017, 2024)
res.kind # MappingKind.SPLIT
res.targets # ('J45.00', 'J45.01', 'J45.02', 'J45.03', 'J45.04', 'J45.05', 'J45.09')
res.ambiguous # True -> needs the original mention to disambiguate
res.needs_manual_review # True
cw.recommend(res) # 'J45.0' -> longest shared prefix (a lossy single-code pick)
cw.map("A00.0", 2017, 2024).kind # MappingKind.IDENTITY
cw.map("K55.88", 2017, 2024) # one_to_one -> ('K55.8',)
# Merges are detected via the inverse table — even when a code also splits.
# M79.67 splits, but one of its new targets (G90.71) also absorbs sibling codes,
# so the co-merge is surfaced rather than lost:
m = cw.map("M79.67", 2017, 2024)
m.kind # MappingKind.SPLIT
m.is_merge # True
m.merged_from # ('M79.60', 'M79.65', 'M79.66') -> the codes it co-merged with
Inspecting the chain
res = cw.map("J45.0", 2017, 2024)
for step in res.steps:
print(step.from_year, "→", step.to_year, step.code, step.targets, step.kind.value)
Combined codes & Kreuz-Stern notation
ICD-10-GM marks the roles in a combined diagnosis with a Kreuz/dagger († or +,
the underlying disease), a Stern (*, the manifestation), and an !
(Ausrufezeichen, a secondary code), and sources often join the components into one
string. BfArM's transition tables key on bare codes, so map() strips a trailing
marker for the lookup and re-applies it to the result — and map_components()
handles a whole compound:
cw.map("E10.30†", 2017, 2024).targets # ('E10.30†',) -> marker preserved
cw.map("B18.1†", 2017, 2024).targets # ('B18.11†', 'B18.12†', 'B18.14†', 'B18.19†')
# A compound diagnosis: one result per component, each keeping its role marker.
for r in cw.map_components("A41.9,R65.1!", 2017, 2024):
print(r.code, "→", r.targets) # A41.9 → (...); R65.1! → ('R65.1!',)
cw.map("A41.9,R65.1!", 2017, 2024) # ValueError: use map_components()
Helpers strip_markers, split_marker, and split_components are exported if you
need them directly.
Validating markers against the systematik (optional)
By default the marker is re-applied syntactically — it preserves the source
annotation's role but isn't checked. Pass a Systematik (parsed from BfArM's
ClaML) and the crosswalk will verify each marker against the code's real role,
warning on a contradiction (and warning that it can't verify when no systematik
is loaded):
from icd10gm_crosswalk import Crosswalk, Systematik
syst = Systematik.from_source("~/icd10gm2024syst-claml.zip") # one ClaML file
cw = Crosswalk.from_source("~/icd10gm-zips", systematik=syst)
cw.map("H36.0*", 2017, 2024) # fine — H36.0 really is a Stern code
cw.map("A09.9+", 2017, 2024) # MarkerValidationWarning: A09.9 is not a dagger code
cw.map("H36.0+", 2017, 2024) # MarkerValidationWarning: H36.0 is star, not dagger
ClaML is the single source of truth here: it carries the exact dagger/aster
designation (the flat metadata lumps dagger in with ordinary primary codes — of
~13k primary codes only ~131 are truly dagger), and the modifier-expanded blocks
(diabetes E10–E14, etc.) are resolved too, so E10.30† validates correctly.
Filter the notices with Python's warnings module (category
MarkerValidationWarning) if you don't want them.
Command line
# What editions does a data source cover?
icd10gm-crosswalk info --data ~/icd10gm-zips
# Map a single code, with the per-year trace
icd10gm-crosswalk map J45.0 --from 2017 --to 2024 --trace --data ~/icd10gm-zips
# Which BfArM files do I need to download for a given range?
icd10gm-crosswalk urls --from 2017 --to 2024
Data & licensing
This library deliberately does not bundle or redistribute BfArM data. BfArM's download conditions make the files free but copyrighted: you accept a usage agreement on download, you may not pass the files on "in the acquired format", but you may build and distribute derived "value-added products" (such as the crosswalk this library produces). Bundling the tables would violate the first clause — so instead the library reads data you obtain.
Download once via your browser (one click, accept the terms) and point the
library at the files. You need the Überleitung (transition) package for each
edition, e.g. icd10gm2024syst-ueberl_zip.html from the
BfArM download page.
Drop the year ZIPs in a folder and pass that folder to from_source.
To save you hunting BfArM's site, the library tells you exactly which files a given range needs — and only that. It does not fetch them: BfArM's portal sits behind an anti-bot gate that rejects scripted requests, so an automated downloader would be unreliable, and the files may not be redistributed anyway.
icd10gm-crosswalk urls --from 2017 --to 2024
from icd10gm_crosswalk import download_instructions, transition_zip_urls
print(download_instructions(2017, 2024)) # copy-pasteable recipe + terms link
transition_zip_urls(2017, 2024) # [(2018, url), ..., (2024, url)]
When you publish a crosswalk produced with this tool, attribute the source ("ICD-10-GM, © BfArM") as the terms require.
How it works
- Parse each yearly Umsteiger table — a UTF-8, semicolon-separated file of
CodePrev ; CodeCur ; A(uto-forward) ; A(uto-backward)rows. - Index each step forward (
prev → rows) and inverse (cur → rows); the inverse index is what makes n:1 merge detection possible. - Chain the steps: starting from your code, take the union of successors at each step (a code absent from a table is carried forward unchanged), tracking where a non-automatic row or a split was introduced.
- Classify the net result by forward cardinality and surface merges and
ambiguity separately, so the forward mapping stays faithful to the official
tables. See the design note in
crosswalk.py.
Correctness
The synthetic test suite (tests/data/) exercises every code path — identity,
1:1, split, merge, deletion, chained split-then-merge re-convergence, and
manual-flag propagation — offline, with no BfArM data.
A separate golden regression (tests/test_bronco_regression.py) reproduces a
real research crosswalk (BRONCO150, ICD-10-GM 2017→2024): 26 splits + 2 remaps,
0 deletions, each with its exact target set and manual flag. Every code in it is
a public ICD-10-GM code and every mapping a public BfArM transition. It runs when
ICD10GM_DATA_DIR points at a directory of BfArM year ZIPs, and is skipped
otherwise.
Development
uv sync
uv run ruff check . # lint
uv run ruff format --check . # format
uv run ty check # type-check
uv run pytest # test (synthetic fixtures; no BfArM data needed)
# To also run the real-data regression:
ICD10GM_DATA_DIR=~/icd10gm-zips uv run pytest
License
MIT. Note that the BfArM ICD-10-GM data this tool consumes is under its own separate terms — see Data & licensing.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file icd10gm_crosswalk-0.1.0.tar.gz.
File metadata
- Download URL: icd10gm_crosswalk-0.1.0.tar.gz
- Upload date:
- Size: 38.8 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
18915d06e5abd08932693f81d7bdcd18563bb5f5bb0809385b89f44f5e33d774
|
|
| MD5 |
fb35a47b2492251d01b7f9d57e947534
|
|
| BLAKE2b-256 |
24b9392218f8ae14646884210e590ed7608934d44491503917bc540aa7a83f3b
|
Provenance
The following attestation bundles were made for icd10gm_crosswalk-0.1.0.tar.gz:
Publisher:
release.yml on JohannKaspar/icd10gm-crosswalk
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
icd10gm_crosswalk-0.1.0.tar.gz -
Subject digest:
18915d06e5abd08932693f81d7bdcd18563bb5f5bb0809385b89f44f5e33d774 - Sigstore transparency entry: 1826220722
- Sigstore integration time:
-
Permalink:
JohannKaspar/icd10gm-crosswalk@d2354017125fc437371f6424c2591e4ffe166167 -
Branch / Tag:
refs/tags/v0.1.0 - Owner: https://github.com/JohannKaspar
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@d2354017125fc437371f6424c2591e4ffe166167 -
Trigger Event:
release
-
Statement type:
File details
Details for the file icd10gm_crosswalk-0.1.0-py3-none-any.whl.
File metadata
- Download URL: icd10gm_crosswalk-0.1.0-py3-none-any.whl
- Upload date:
- Size: 27.0 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
7702df1b380e285a66a26e036c92d6b6f124ed39a5728bd834f1d29f80d7aeb3
|
|
| MD5 |
759eb5c7a521b265cc8360709a5fe4b0
|
|
| BLAKE2b-256 |
925bac8bb5314679b954f460380ebc68643573064d192cdcef283a521fdfcd58
|
Provenance
The following attestation bundles were made for icd10gm_crosswalk-0.1.0-py3-none-any.whl:
Publisher:
release.yml on JohannKaspar/icd10gm-crosswalk
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
icd10gm_crosswalk-0.1.0-py3-none-any.whl -
Subject digest:
7702df1b380e285a66a26e036c92d6b6f124ed39a5728bd834f1d29f80d7aeb3 - Sigstore transparency entry: 1826220856
- Sigstore integration time:
-
Permalink:
JohannKaspar/icd10gm-crosswalk@d2354017125fc437371f6424c2591e4ffe166167 -
Branch / Tag:
refs/tags/v0.1.0 - Owner: https://github.com/JohannKaspar
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@d2354017125fc437371f6424c2591e4ffe166167 -
Trigger Event:
release
-
Statement type: