WeavEHR
WeavEHR is an open-source Python framework for extracting and harmonising intensive care unit (ICU) data from heterogeneous sources — such as MIMIC-IV, eICU-CRD, AmsterdamUMCdb (via OMOP CDM), and your own institutional or OMOP CDM database exports — into the standardised MEDS (Medical Event Data Standard) format.
Instead of writing one-off SQL or pandas scripts for every dataset and every study, you describe what to extract in declarative YAML configurations and let WeavEHR handle the how: typed reading, joins, timestamp reconstruction, unit-aware event codes, and cross-dataset concept harmonisation. WeavEHR ships with curated configurations for public ICU datasets and a growing dictionary of clinical concepts (vital signs, laboratory values, vasopressors, ventilation, and more), so the same study code can run against any supported dataset.
WeavEHR is the spiritual successor to dataset-harmonisation tools like ricu (R), rebuilt in Python on a modern stack: Polars for fast, out-of-core streaming processing, Pydantic for validated configuration, and MEDS for an interoperable output format that plugs directly into the wider MEDS ecosystem of modelling and evaluation tools.
Source code: https://github.com/aidh-ms/WeavEHR · Issues: https://github.com/aidh-ms/WeavEHR/issues
Why WeavEHR?
- One concept, many datasets. Define a clinical concept (e.g. heart rate, norepinephrine rate, antibiotics) once; map it per dataset with small YAML files. Extraction code stays identical across MIMIC-IV, eICU-CRD, NWICU, and custom sources.
- MEDS-native output. All output is written as MEDS-compliant Parquet (
subject_id,time,code,numeric_value,text_value, plus configurable extension columns) with full metadata (dataset.json,codes.parquet) — ready for MEDS-compatible downstream tooling. - Declarative, versioned, reproducible. Every table, event, and concept is a versioned YAML config with a stable identifier. The exact merged configuration used for a run is snapshotted into the output project, so results can be traced and reproduced.
- Versions and variants as diffs. New dataset versions and variants (like the eICU demo)
extenda reference version and state only their differences — no copied configs to keep in sync. The entire eICU demo configuration is two small files. The same mechanism works across datasets: the AmsterdamUMCdb config simplyextends the shared OMOP CDM model, inheriting all of its table configs. - Fully offline. Designed for sensitive medical data: no network access required, nothing leaves your secure perimeter.
- Scales down and up. Polars lazy streaming lets you process full MIMIC-IV on a 16–32 GB laptop, without a database server or cluster.
- Extensible. Add new datasets by writing YAML (no Python required); add new transformation callbacks or complex concept logic in Python where YAML isn't enough.
How it works
WeavEHR organises processing as a pipeline of steps that operate inside a project directory:
flowchart LR
A[Raw source data<br>Parquet / CSV / CSV.GZ] -->|ExtractionStep| B[Dataset-level MEDS events<br>e.g. mimic-iv//chartevents//220045//Heart Rate//bpm]
B -->|ConceptStep| C[Harmonised concepts<br>e.g. heart_rate//bpm]
C --> D[Your analysis /<br>MEDS ecosystem tools]
- Extraction step — reads the raw tables of each dataset (typed Parquet/CSV scan, lookup-table joins, timestamp reconstruction) and emits one MEDS event stream per configured event. Event codes preserve full provenance:
dataset//table//itemid//label//unit. - Concept step — maps dataset-specific event codes onto shared, dataset-agnostic clinical concepts using pattern matching (simple concepts), computation over other concepts (derived concepts), or custom Python transformers (complex concepts). Concept dependencies are resolved automatically.
- Sharding step (in development) — re-partitions harmonised data into analysis-ready shards.
Each step reads its inputs and writes its outputs as a self-contained MEDS dataset inside the project, so intermediate results are inspectable and individual steps can be re-run in isolation.
Supported datasets
WeavEHR ships with ready-to-use extraction and concept configurations under configs/:
| Dataset | Version | Extraction configs | Concept mappings |
|---|---|---|---|
| MIMIC-IV | 3.1, 2.2 | 19 tables (ICU + hosp) | ~80 concepts |
| MIMIC-IV demo | 2.2 | inherited from MIMIC-IV via extends |
inherited |
| eICU-CRD | 2.0 | 14 tables | in progress |
| eICU-CRD demo | 2.0 | inherited from eICU-CRD via extends |
inherited |
| NWICU | 0.1.0 | 9 tables | in progress |
| OMOP CDM | 5.4 | 11 tables (reusable model reference, read from Parquet) | — (added per OMOP dataset) |
| AmsterdamUMCdb (OMOP, via AMSTEL) | 1.5.0 | inherited from OMOP CDM via extends |
~60 concepts (in progress) |
| HiRID | 1.1.1 | 3 tables (native long format) | ~60 concepts (in progress) |
| SICdb | 1.0.6 | 4 tables (native long format) | ~55 concepts (in progress) |
The OMOP CDM 5.4 entry is a reusable model configuration rather than a single dataset: any data exported to the OMOP Common Data Model inherits its table configs via extends and adds only concept mappings. AmsterdamUMCdb is the first such dataset, read through its AMSTEL OMOP CDM 5.4 export, with concept mappings keyed on the OMOP concept_ids that AMSTEL assigns. HiRID is read in its native long format (observations/pharma/general), with concept mappings keyed on HiRID variableids and unit conversions ported from ricu. SICdb is read in its native long format, with concept mappings keyed on SICdb LaboratoryID/DataIDs and unit conversions ported from ricu.
The shared concept dictionary in configs/concepts/ currently covers ~90 concepts across vital signs, blood gas, clinical chemistry, hematology, medications (incl. vasopressors and antibiotics), neurological scores, respiratory parameters, fluid output, and demographics.
Installation
WeavEHR requires Python 3.13+. You can install WeavEHR via PyPI or from source.
With PyPI
# with pip
pip install open-icu
# with uv
uv add open-icu
With GitHub
# with pip
pip install git+https://github.com/aidh-ms/WeavEHR
# with uv
uv add git+https://github.com/aidh-ms/WeavEHR
For development, clone the repository and use the included dev container or uv:
git clone https://github.com/aidh-ms/WeavEHR.git
cd WeavEHR
uv sync --all-groups
Quickstart
A pipeline is two small YAML files plus a few lines of Python.
1. Configure the extraction step (config/extraction.yml) — point WeavEHR at the bundled dataset configs and your local data:
name: Extraction
version: 1.0.0
config:
data:
- name: mimic-iv
version: "3.1"
path: /path/to/physionet.org/files/mimiciv/3.1
2. Configure the concept step (config/concept.yml) — select the concept dictionary and per-dataset mappings:
name: Concept
version: 1.0.0
config:
extraction_step: Extraction
mapping_configs:
- name: mimic-iv
version: "3.1"
3. Run the pipeline:
from pathlib import Path
from weavehr import WeavEHRProject, ExtractionStep, ConceptStep
config_path = Path.cwd() / "config"
project_path = Path.cwd() / "output" / "project"
with WeavEHRProject(project_path) as project:
extraction_step = ExtractionStep.load(project, config_path / "extraction.yml")
extraction_step.run()
concept_step = ConceptStep.load(project, config_path / "concept.yml")
concept_step.run()
4. Use the results. The project directory now contains one MEDS dataset per step:
output/project/
├── configs/ # snapshot of every config used (reproducibility)
├── workspace/ # intermediate per-step files
└── datasets/
├── extraction/
│ ├── data/mimic-iv/3.1/<table>/<EVENT>.parquet
│ └── metadata/{dataset.json, codes.parquet}
└── concept/
├── data/<concept>/<version>/<dataset>.parquet
└── metadata/{dataset.json, codes.parquet}
Read it with any Parquet-capable tool, e.g.:
import polars as pl
heart_rate = pl.scan_parquet(
project_path / "datasets" / "concept" / "data" / "heart_rate" / "1.0.0" / "mimic-iv.parquet"
).collect()
A complete runnable example lives in example/pipeline.ipynb.
Configuration at a glance
WeavEHR is configured at three levels — see the documentation for the full reference.
Dataset/table configs (configs/datasets/<dataset>/<version>/tables/*.yml) describe how to turn one raw table into MEDS events: typed columns, joins, and event definitions.
path: icu/chartevents.csv.gz
columns:
- name: subject_id
type: int64
- name: charttime
type: datetime
params: { format: "%Y-%m-%d %H:%M:%S" }
- name: itemid
type: int64
- name: value
type: string
- name: valuenum
type: float32
- name: valueuom
type: string
join:
- path: icu/d_items.csv.gz
both_on: [itemid]
columns: [{ name: itemid, type: int64 }, { name: label, type: string }]
events:
- name: CHART
columns:
subject_id: col(subject_id)
time: col(charttime)
code: [col(itemid), col(label), col(valueuom)]
numeric_value: col(valuenum)
text_value: col(value)
name is a technical event identifier used by the extraction step. By default it is included in the generated MEDS code; include_event_name_in_code can disable it globally, with an optional per-table override. The remaining code is built from the automatic dataset/table prefix, optional code_prefix, configured columns.code components, and optional code_suffix.
Concept configs (configs/concepts/<category>/*.yml) define dataset-agnostic clinical concepts:
name: heart_rate
version: 1.0.0
unit: bpm
Concept mappings (configs/datasets/<dataset>/<version>/mappings/*.yml) connect a concept to dataset-specific codes:
type: simple
mappings:
- pattern:
table: chartevents
event: CHART # technical extraction event identifier; part of `code` by default
code: (220045//Heart Rate)
columns:
numeric_value: col(numeric_value)
Here, event selects the technical extraction event stream. It is part of the MEDS code by default, but can be omitted through include_event_name_in_code.
Wherever values are computed, configs use a small expression language with composable callbacks — e.g. reconstructing absolute timestamps from eICU's relative offsets:
callbacks:
- to_datetime(hospitaldischargeyear, 1, 1, hospitaladmittime24, hospitaladmitoffset, "minutes", output=admission_timestamp)
post_join_callbacks:
- add_offset(admission_timestamp, labresultoffset, output=event_timestamp)
Arithmetic, comparisons, and boolean logic work as ordinary expressions (col(weight) / (col(height) * col(height))), and custom callbacks can be registered from Python.
Documentation
- Package overview
- Installation
- Basic usage
- User guide: pipeline & projects · extraction configs · dataset versions & variants · concept configs · expression language
- Contributing
- Architecture documentation following arc42 in
docs/arc/
Project status & roadmap
WeavEHR is in active development (pre-1.0); configuration formats may still change between minor versions. Current focus areas:
-
Completing concept mappings for eICU-CRD, NWICU, AmsterdamUMCdb, HiRID, and SICdb
-
The sharding step for analysis-ready data partitioning
-
Additional datasets and broader OMOP CDM coverage
Contributions of dataset configurations and concept definitions are especially welcome — they are pure YAML and require no changes to the framework itself.
Contributing
We welcome contributions! Please use the dev container for a ready-to-go environment, follow Conventional Commits (feat:, fix:, docs:, …), and make sure checks pass before opening a PR:
uv run ruff format && uv run ruff check # format & lint
uv run ty check . # type checking
uv run pytest # tests
See the contributing guide for details, and our Code of Conduct.
Citation
If you use WeavEHR in your research, please cite it via the metadata in CITATION.cff (use the "Cite this repository" button on GitHub).
Related projects
- MEDS — the Medical Event Data Standard that WeavEHR targets
- ricu — R package for ICU data harmonisation that inspired WeavEHR's concept dictionary
- YAIB — Yet Another ICU Benchmark, a complementary benchmarking framework
License
WeavEHR is released under the MIT License. Developed by the AIDH MS team at the University of Münster and the Medical University of Innsbruck.
Metadata
Release files for weavehr 0.0.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| weavehr-0.0.1.tar.gz | 121.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| weavehr-0.0.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 509.5 kB
Release files / weavehr-0.0.1.tar.gz
| Download URL | weavehr-0.0.1.tar.gz |
|---|---|
| Size | 121.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
558e0c0bef12cb1f28ea4638c772f5238df32d2d2a7a3beaadc450d8b271a452
|
|
BLAKE2b-256 checksum How to use checksums |
2a014b544562547ae141f0ad24211732dd44c8ad2a19a28b112d6b617c147327
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.12.22 {"installer":{"name":"uv","version":"0.12.22","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|
Release files / weavehr-0.0.1-py3-none-any.whl
| Download URL | weavehr-0.0.1-py3-none-any.whl |
|---|---|
| Size | 388.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
1b8b96668def0e2415583458074c76d2bcf48d33203b21d59364cba1cf3a1089
|
|
BLAKE2b-256 checksum How to use checksums |
f51e2a24f4f0745038de907ee35bcafe9d0bbe6536840c1d7ca1ed8ee94b39df
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.12.22 {"installer":{"name":"uv","version":"0.12.22","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|