Skip to main content

WeavEHR

CI Coverage Status Python License: MIT

WeavEHR is an open-source Python framework for extracting and harmonising intensive care unit (ICU) data from heterogeneous sources — such as MIMIC-IV, eICU-CRD, AmsterdamUMCdb (via OMOP CDM), and your own institutional or OMOP CDM database exports — into the standardised MEDS (Medical Event Data Standard) format.

Instead of writing one-off SQL or pandas scripts for every dataset and every study, you describe what to extract in declarative YAML configurations and let WeavEHR handle the how: typed reading, joins, timestamp reconstruction, unit-aware event codes, and cross-dataset concept harmonisation. WeavEHR ships with curated configurations for public ICU datasets and a growing dictionary of clinical concepts (vital signs, laboratory values, vasopressors, ventilation, and more), so the same study code can run against any supported dataset.

WeavEHR is the spiritual successor to dataset-harmonisation tools like ricu (R), rebuilt in Python on a modern stack: Polars for fast, out-of-core streaming processing, Pydantic for validated configuration, and MEDS for an interoperable output format that plugs directly into the wider MEDS ecosystem of modelling and evaluation tools.


Source code: https://github.com/aidh-ms/WeavEHR · Issues: https://github.com/aidh-ms/WeavEHR/issues


Why WeavEHR?

  • One concept, many datasets. Define a clinical concept (e.g. heart rate, norepinephrine rate, antibiotics) once; map it per dataset with small YAML files. Extraction code stays identical across MIMIC-IV, eICU-CRD, NWICU, and custom sources.
  • MEDS-native output. All output is written as MEDS-compliant Parquet (subject_id, time, code, numeric_value, text_value, plus configurable extension columns) with full metadata (dataset.json, codes.parquet) — ready for MEDS-compatible downstream tooling.
  • Declarative, versioned, reproducible. Every table, event, and concept is a versioned YAML config with a stable identifier. The exact merged configuration used for a run is snapshotted into the output project, so results can be traced and reproduced.
  • Versions and variants as diffs. New dataset versions and variants (like the eICU demo) extend a reference version and state only their differences — no copied configs to keep in sync. The entire eICU demo configuration is two small files. The same mechanism works across datasets: the AmsterdamUMCdb config simply extends the shared OMOP CDM model, inheriting all of its table configs.
  • Fully offline. Designed for sensitive medical data: no network access required, nothing leaves your secure perimeter.
  • Scales down and up. Polars lazy streaming lets you process full MIMIC-IV on a 16–32 GB laptop, without a database server or cluster.
  • Extensible. Add new datasets by writing YAML (no Python required); add new transformation callbacks or complex concept logic in Python where YAML isn't enough.

How it works

WeavEHR organises processing as a pipeline of steps that operate inside a project directory:

flowchart LR
    A[Raw source data<br>Parquet / CSV / CSV.GZ] -->|ExtractionStep| B[Dataset-level MEDS events<br>e.g. mimic-iv//chartevents//220045//Heart Rate//bpm]
    B -->|ConceptStep| C[Harmonised concepts<br>e.g. heart_rate//bpm]
    C --> D[Your analysis /<br>MEDS ecosystem tools]
  1. Extraction step — reads the raw tables of each dataset (typed Parquet/CSV scan, lookup-table joins, timestamp reconstruction) and emits one MEDS event stream per configured event. Event codes preserve full provenance: dataset//table//itemid//label//unit.
  2. Concept step — maps dataset-specific event codes onto shared, dataset-agnostic clinical concepts using pattern matching (simple concepts), computation over other concepts (derived concepts), or custom Python transformers (complex concepts). Concept dependencies are resolved automatically.
  3. Sharding step (in development) — re-partitions harmonised data into analysis-ready shards.

Each step reads its inputs and writes its outputs as a self-contained MEDS dataset inside the project, so intermediate results are inspectable and individual steps can be re-run in isolation.

Supported datasets

WeavEHR ships with ready-to-use extraction and concept configurations under configs/:

Dataset Version Extraction configs Concept mappings
MIMIC-IV 3.1, 2.2 19 tables (ICU + hosp) ~80 concepts
MIMIC-IV demo 2.2 inherited from MIMIC-IV via extends inherited
eICU-CRD 2.0 14 tables in progress
eICU-CRD demo 2.0 inherited from eICU-CRD via extends inherited
NWICU 0.1.0 9 tables in progress
OMOP CDM 5.4 11 tables (reusable model reference, read from Parquet) — (added per OMOP dataset)
AmsterdamUMCdb (OMOP, via AMSTEL) 1.5.0 inherited from OMOP CDM via extends ~60 concepts (in progress)
HiRID 1.1.1 3 tables (native long format) ~60 concepts (in progress)
SICdb 1.0.6 4 tables (native long format) ~55 concepts (in progress)

The OMOP CDM 5.4 entry is a reusable model configuration rather than a single dataset: any data exported to the OMOP Common Data Model inherits its table configs via extends and adds only concept mappings. AmsterdamUMCdb is the first such dataset, read through its AMSTEL OMOP CDM 5.4 export, with concept mappings keyed on the OMOP concept_ids that AMSTEL assigns. HiRID is read in its native long format (observations/pharma/general), with concept mappings keyed on HiRID variableids and unit conversions ported from ricu. SICdb is read in its native long format, with concept mappings keyed on SICdb LaboratoryID/DataIDs and unit conversions ported from ricu.

The shared concept dictionary in configs/concepts/ currently covers ~90 concepts across vital signs, blood gas, clinical chemistry, hematology, medications (incl. vasopressors and antibiotics), neurological scores, respiratory parameters, fluid output, and demographics.

Installation

WeavEHR requires Python 3.13+. You can install WeavEHR via PyPI or from source.

With PyPI

# with pip
pip install open-icu

# with uv
uv add open-icu

With GitHub

# with pip
pip install git+https://github.com/aidh-ms/WeavEHR

# with uv
uv add git+https://github.com/aidh-ms/WeavEHR

For development, clone the repository and use the included dev container or uv:

git clone https://github.com/aidh-ms/WeavEHR.git
cd WeavEHR
uv sync --all-groups

Quickstart

A pipeline is two small YAML files plus a few lines of Python.

1. Configure the extraction step (config/extraction.yml) — point WeavEHR at the bundled dataset configs and your local data:

name: Extraction
version: 1.0.0

config:
  data:
    - name: mimic-iv
      version: "3.1"
      path: /path/to/physionet.org/files/mimiciv/3.1

2. Configure the concept step (config/concept.yml) — select the concept dictionary and per-dataset mappings:

name: Concept
version: 1.0.0

config:
  extraction_step: Extraction
  mapping_configs:
    - name: mimic-iv
      version: "3.1"

3. Run the pipeline:

from pathlib import Path

from weavehr import WeavEHRProject, ExtractionStep, ConceptStep

config_path = Path.cwd() / "config"
project_path = Path.cwd() / "output" / "project"

with WeavEHRProject(project_path) as project:
    extraction_step = ExtractionStep.load(project, config_path / "extraction.yml")
    extraction_step.run()

    concept_step = ConceptStep.load(project, config_path / "concept.yml")
    concept_step.run()

4. Use the results. The project directory now contains one MEDS dataset per step:

output/project/
├── configs/        # snapshot of every config used (reproducibility)
├── workspace/      # intermediate per-step files
└── datasets/
    ├── extraction/
    │   ├── data/mimic-iv/3.1/<table>/<EVENT>.parquet
    │   └── metadata/{dataset.json, codes.parquet}
    └── concept/
        ├── data/<concept>/<version>/<dataset>.parquet
        └── metadata/{dataset.json, codes.parquet}

Read it with any Parquet-capable tool, e.g.:

import polars as pl

heart_rate = pl.scan_parquet(
    project_path / "datasets" / "concept" / "data" / "heart_rate" / "1.0.0" / "mimic-iv.parquet"
).collect()

A complete runnable example lives in example/pipeline.ipynb.

Configuration at a glance

WeavEHR is configured at three levels — see the documentation for the full reference.

Dataset/table configs (configs/datasets/<dataset>/<version>/tables/*.yml) describe how to turn one raw table into MEDS events: typed columns, joins, and event definitions.

path: icu/chartevents.csv.gz

columns:
  - name: subject_id
    type: int64
  - name: charttime
    type: datetime
    params: { format: "%Y-%m-%d %H:%M:%S" }
  - name: itemid
    type: int64
  - name: value
    type: string
  - name: valuenum
    type: float32
  - name: valueuom
    type: string

join:
  - path: icu/d_items.csv.gz
    both_on: [itemid]
    columns: [{ name: itemid, type: int64 }, { name: label, type: string }]

events:
  - name: CHART
    columns:
      subject_id: col(subject_id)
      time: col(charttime)
      code: [col(itemid), col(label), col(valueuom)]
      numeric_value: col(valuenum)
      text_value: col(value)

name is a technical event identifier used by the extraction step. By default it is included in the generated MEDS code; include_event_name_in_code can disable it globally, with an optional per-table override. The remaining code is built from the automatic dataset/table prefix, optional code_prefix, configured columns.code components, and optional code_suffix.

Concept configs (configs/concepts/<category>/*.yml) define dataset-agnostic clinical concepts:

name: heart_rate
version: 1.0.0
unit: bpm

Concept mappings (configs/datasets/<dataset>/<version>/mappings/*.yml) connect a concept to dataset-specific codes:

type: simple
mappings:
  - pattern:
      table: chartevents
      event: CHART  # technical extraction event identifier; part of `code` by default
      code: (220045//Heart Rate)
    columns:
      numeric_value: col(numeric_value)

Here, event selects the technical extraction event stream. It is part of the MEDS code by default, but can be omitted through include_event_name_in_code.

Wherever values are computed, configs use a small expression language with composable callbacks — e.g. reconstructing absolute timestamps from eICU's relative offsets:

callbacks:
  - to_datetime(hospitaldischargeyear, 1, 1, hospitaladmittime24, hospitaladmitoffset, "minutes", output=admission_timestamp)
post_join_callbacks:
  - add_offset(admission_timestamp, labresultoffset, output=event_timestamp)

Arithmetic, comparisons, and boolean logic work as ordinary expressions (col(weight) / (col(height) * col(height))), and custom callbacks can be registered from Python.

Documentation

Project status & roadmap

WeavEHR is in active development (pre-1.0); configuration formats may still change between minor versions. Current focus areas:

  • Completing concept mappings for eICU-CRD, NWICU, AmsterdamUMCdb, HiRID, and SICdb

  • The sharding step for analysis-ready data partitioning

  • Additional datasets and broader OMOP CDM coverage

Contributions of dataset configurations and concept definitions are especially welcome — they are pure YAML and require no changes to the framework itself.

Contributing

We welcome contributions! Please use the dev container for a ready-to-go environment, follow Conventional Commits (feat:, fix:, docs:, …), and make sure checks pass before opening a PR:

uv run ruff format && uv run ruff check    # format & lint
uv run ty check .                          # type checking
uv run pytest                              # tests

See the contributing guide for details, and our Code of Conduct.

Citation

If you use WeavEHR in your research, please cite it via the metadata in CITATION.cff (use the "Cite this repository" button on GitHub).

  • MEDS — the Medical Event Data Standard that WeavEHR targets
  • ricu — R package for ICU data harmonisation that inspired WeavEHR's concept dictionary
  • YAIB — Yet Another ICU Benchmark, a complementary benchmarking framework

License

WeavEHR is released under the MIT License. Developed by the AIDH MS team at the University of Münster and the Medical University of Innsbruck.

Metadata

Release files for weavehr 0.0.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for weavehr 0.0.1
File Size Uploaded
weavehr-0.0.1.tar.gz 121.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for weavehr 0.0.1
File Interpreter ABI Platform
weavehr-0.0.1-py3-none-any.whl Python 3 none any Details

Total release size: 509.5 kB

Release files / weavehr-0.0.1.tar.gz

Download URL weavehr-0.0.1.tar.gz
Size 121.1 kB
Tags Source
SHA-256 checksum
How to use checksums
558e0c0bef12cb1f28ea4638c772f5238df32d2d2a7a3beaadc450d8b271a452
BLAKE2b-256 checksum
How to use checksums
2a014b544562547ae141f0ad24211732dd44c8ad2a19a28b112d6b617c147327
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.22 {"installer":{"name":"uv","version":"0.12.22","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release files / weavehr-0.0.1-py3-none-any.whl

Download URL weavehr-0.0.1-py3-none-any.whl
Size 388.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
1b8b96668def0e2415583458074c76d2bcf48d33203b21d59364cba1cf3a1089
BLAKE2b-256 checksum
How to use checksums
f51e2a24f4f0745038de907ee35bcafe9d0bbe6536840c1d7ca1ed8ee94b39df
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.22 {"installer":{"name":"uv","version":"0.12.22","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release history Release notifications | RSS feed

This release

0.0.1 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page