medh5
One medical imaging sample — a subject, at every timepoint, with all of its ground truth — in a single self-describing HDF5 file.
Multi-modality images, segmentation in five encodings, detection boxes, keypoints, contours, meshes, classification, registration between visits, provenance and quality records, and per-object integrity digests. Format version 1.0, with a normative specification and a 117-case conformance suite any implementation can run.
import medh5
with medh5.open("case_0001.medh5") as s:
s.identity.subject_id # "BRATS-GLI-01234"
s.at("tp1").images["CT_tp1"].read(physical=True) # HU, not raw counts
s.annotations["organs"].dense(["liver", "spleen"]) # any encoding, one API
s.transform_between("tp0", "tp1") # resolved via frames
s.tracks("lesion") # lesions joined across visits
Install
pip install medh5
pip install "medh5[torch,nifti,dicom]"
Reading and writing needs only h5py, hdf5plugin and numpy. Extras:
torch, monai, nifti, dicom, dicomseg, itk, schema, interp.
Documentation
medh5.readthedocs.io — tutorials, how-to guides, the Python and CLI reference, and the normative specification.
Write your first sample · How-to guides · Python API · CLI · Specification
What the format is for
- One file per subject, not per scan — every visit in one place, so longitudinal work has a referent and splitting by file cannot leak a patient.
- Geometry is stated once and never guessed — declared grids, boxes at voxel edges, and converters that refuse rather than invent.
- Absence is not silence — a class examined and not found is recorded as such, which is a different training signal from one nobody examined.
- Every claim is checkable — per-object digests, a Merkle
content_idthat survives recompression, a stable diagnostic-code table, and a 117-case conformance corpus. - Reading a patch is fast — a 64³ multi-class patch in ~4 ms, and O(1)
foreground sampling once
build_index()has run.
Write a sample
import numpy as np
import medh5
from medh5 import LabelClass, LabelSet
labels = LabelSet("demo-v1", version="1.0.0", classes=[
LabelClass(1, "liver", "Liver", category="organ"),
LabelClass(2, "spleen", "Spleen", category="organ"),
LabelClass(3, "lesion", "Lesion", parents=[1], category="lesion"),
])
with medh5.create("case_0001.medh5", sample_id="case_0001",
subject_id="DEMO-0001") as w:
w.label_set(labels)
w.add_timepoint("tp0", label="baseline", days_from_baseline=0)
w.add_grid("ct", shape=ct.shape, spacing=(2.0, 0.8, 0.8),
origin=(-64.0, -38.4, -38.4), timepoint="tp0")
w.add_image("CT", ct, grid="ct", modality="CT",
value_type="quantitative", value_units="HU")
w.add_segmentation("organs", grid="ct",
masks={"liver": liver, "lesion": lesion},
annotated_classes=["liver", "spleen", "lesion"])
w.build_index() # optional; foreground sampling is O(1) only with it
annotated_classes names the spleen although there is no spleen mask: that
records "we looked and found none". The encoding is chosen by measuring the
class overlap graph — liver and lesion overlap, so it picks one that can
represent that — and the write is atomic.
Train on it
from torch.utils.data import DataLoader
from medh5.torch import PatchDataset, collate, worker_init_fn
from medh5.sampling import PatchSampler
sampler = PatchSampler((96, 96, 96), strategy="balanced",
foreground_classes=["liver", "lesion"])
dataset = PatchDataset(paths, sampler, images=["CT"],
annotations={"organs": ["liver", "lesion"]},
samples_per_volume=8)
loader = DataLoader(dataset, batch_size=2, num_workers=8,
worker_init_fn=worker_init_fn, collate_fn=collate)
worker_init_fn drops handles inherited across a fork. It is recommended
rather than required: the handle cache is PID-keyed and re-checks ownership on
every access, so a forked worker abandons the parent's handles on first use
rather than reading through or closing them.
Command line
medh5 info case.medh5 # grids, images, annotations, coverage
medh5 validate case.medh5 --level strict
medh5 verify case.medh5 # digests and content_id
medh5 timeline case.medh5 # visits and intervals
medh5 track case.medh5 --class lesion # per-lesion volumes across visits
medh5 dataset index studies/ -o cohort.json
medh5 dataset split cohort.json --group-by group_id --stratify-by site_id
medh5 dataset stats cohort.json --partition train --workers 8
medh5 dataset check cohort.json --deep
medh5 convert from-dicom /studies out/ # one sample per patient, all visits
medh5 convert from-nifti case.medh5 --image CT=ct.nii.gz
medh5 convert from-rtstruct plan.dcm case.medh5 --rasterize
medh5 migrate old/*.medh5 -o new/ --group-by subject
medh5 scrub out/*.medh5 --apply --date-shift-days -117
medh5 pack cohort/*.medh5 -o shard.medh5c
medh5 recompress cohort/*.medh5 --profile training
medh5 bench # reproduce the performance targets
medh5 conformance publish suite/ # the suite, for another implementation
Interoperability
| Format | |
|---|---|
| NIfTI | affine and voxels bit-identical on round trip; RAS↔LPS is a sign flip, never a resample |
| DICOM | slices ordered by geometry, spacing measured between origins, modality LUT stored not applied, tags on an explicit allow-list; slices that disagree about orientation, spacing or rescale are refused rather than read off the first one |
| DICOM SEG | frames placed by geometry; segments matched by label, not number; import preserves overlap and FRACTIONAL, export writes BINARY |
| RTSTRUCT | contours stay contours; rasterisation is opt-in and recorded in provenance |
| nnU-Net v2 | class ids kept; region labels become label-set DAG parents; dataset.json round-trips |
| MONAI | to_metatensor gives a MetaTensor with the correct affine |
| 0.x | medh5 migrate, reporting every decision and every guess |
Every conversion returns a report distinguishing what it decided from the data and where it guessed — the encoding chosen, the class ids minted, a half-voxel convention changed, a timepoint order inferred rather than read.
COCO is deliberately unsupported: it has no world geometry, spacing or frame of reference, so importing means inventing a grid and exporting means discarding the geometry that makes a medical annotation reproducible.
Reading it without medh5
import h5py, json, hdf5plugin # hdf5plugin only for blosc2 profiles
with h5py.File("case_0001.medh5") as f:
doc = json.loads(f["meta"][()])
doc["identity"]["subject_id"]
dict(f["grids"]["ct"].attrs) # spacing, origin, direction
f["images"]["CT"][10:20]
medh5 recompress --profile portable writes gzip, readable by any HDF5 build.
Versioning
The format is 1.0. A minor version may add optional objects, profiles, encodings and diagnostic codes; it may not change what an existing one means (spec §16). The package follows semantic versioning from 1.0.0.
0.x files are not readable by 1.0 and are not meant to be — medh5 migrate
converts them once. See Converters.
License
MIT
Release files for medh5 1.4.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| medh5-1.4.1.tar.gz | 732.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| medh5-1.4.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 1.2 MB
Release files / medh5-1.4.1.tar.gz
| Download URL | medh5-1.4.1.tar.gz |
|---|---|
| Size | 732.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
770db52ef44cb38adf17a5f9699eb4a2a989ab821414dd193b091384fc805d28
|
|
BLAKE2b-256 checksum How to use checksums |
c84d5b761dc2e646a656269378fe4417e62d78c995834a4e536e627131363bd1
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 26, 2026.
Transparency logRelease files / medh5-1.4.1-py3-none-any.whl
| Download URL | medh5-1.4.1-py3-none-any.whl |
|---|---|
| Size | 422.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
7f6c26edd4ab01edcf8d736c12c266246804520acee7dd018f1ac1dcccad162d
|
|
BLAKE2b-256 checksum How to use checksums |
5d737600fea8d9eae8f0f42aea4cdcd9521b77fba6c82e0c9db648167baa97cc
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 26, 2026.
Transparency log