medh5
One medical imaging sample — a subject, at every timepoint, with all of its ground truth — in a single self-describing HDF5 file.
Multi-modality images, segmentation in five encodings, detection boxes, keypoints, contours, meshes, classification, registration between visits, provenance and quality records, and per-object integrity digests. Format version 1.0, with a normative specification and a 103-case conformance suite any implementation can run.
import medh5
with medh5.open("case_0001.medh5") as s:
s.identity.subject_id # "BRATS-GLI-01234"
s.at("tp1").images["CT_tp1"].read(physical=True) # HU, not raw counts
s.annotations["organs"].dense(["liver", "spleen"]) # any encoding, one API
s.transform_between("tp0", "tp1") # resolved via frames
s.tracks("lesion") # lesions joined across visits
Install
pip install medh5
pip install "medh5[torch,nifti,dicom]"
Reading and writing needs only h5py, hdf5plugin and numpy. Extras:
torch, monai, nifti, dicom, dicomseg, itk, schema, interp.
Documentation
docs/ · Getting started · Concepts · Specification
Python API · CLI · Annotations · Longitudinal · Training · Converters · Curation · Cohorts · Storage · Conformance
What the format is for
One file per subject, not per scan. A sample is a subject, and a subject has visits. Longitudinal work — change detection, response assessment, lesion tracking, follow-up registration — lives inside one file, which also means assigning whole files to train and test cannot leak a patient between them.
Geometry is stated once and never guessed. Every array is bound to a declared grid with spacing, origin and direction. A box sits at voxel edges and an integer index is a voxel centre, both written down. Converting to NIfTI or DICOM moves numbers between conventions explicitly, and refuses when it cannot.
Absence is not silence. class_ids says what an annotation contains;
annotated_class_ids says what was looked for. A class searched for and not
found is a usable negative example; a class nobody examined is not. Collapsing
the two is how a model learns a site's scans have no spleens.
Every claim is checkable. Per-object SHA-256 over decompressed content and
a Merkle content_id that survives recompression; a validator with a stable
diagnostic-code table; and a conformance corpus with one case per code.
Reading a patch is fast. A 64³ multi-class patch reads in 4 ms against
117 ms in 0.x, and foreground sampling is O(1) in the volume. Reproduce it with
medh5 bench.
Write a sample
import numpy as np
import medh5
from medh5 import LabelClass, LabelSet
labels = LabelSet("demo-v1", version="1.0.0", classes=[
LabelClass(1, "liver", "Liver", category="organ"),
LabelClass(2, "spleen", "Spleen", category="organ"),
LabelClass(3, "lesion", "Lesion", parents=[1], category="lesion"),
])
with medh5.create("case_0001.medh5", sample_id="case_0001",
subject_id="DEMO-0001") as w:
w.label_set(labels)
w.add_timepoint("tp0", label="baseline", days_from_baseline=0)
w.add_grid("ct", shape=ct.shape, spacing=(2.0, 0.8, 0.8),
origin=(-64.0, -38.4, -38.4), timepoint="tp0")
w.add_image("CT", ct, grid="ct", modality="CT",
value_type="quantitative", value_units="HU")
w.add_segmentation("organs", grid="ct",
masks={"liver": liver, "lesion": lesion},
annotated_classes=["liver", "spleen", "lesion"])
annotated_classes names the spleen although there is no spleen mask: that
records "we looked and found none". The encoding is chosen by measuring the
class overlap graph — liver and lesion overlap, so it picks one that can
represent that — and the write is atomic.
Train on it
from torch.utils.data import DataLoader
from medh5.torch import PatchDataset, collate, worker_init_fn
from medh5.sampling import PatchSampler
sampler = PatchSampler((96, 96, 96), strategy="balanced",
foreground_classes=["liver", "lesion"])
dataset = PatchDataset(paths, sampler, images=["CT"],
annotations={"organs": ["liver", "lesion"]},
samples_per_volume=8)
loader = DataLoader(dataset, batch_size=2, num_workers=8,
worker_init_fn=worker_init_fn, collate_fn=collate)
worker_init_fn is required for num_workers > 0: HDF5 handles must not cross
a fork, so the cache is PID-keyed and a forked child abandons the parent's
handles rather than closing descriptors the parent still owns.
Command line
medh5 info case.medh5 # grids, images, annotations, coverage
medh5 validate case.medh5 --level strict
medh5 verify case.medh5 # digests and content_id
medh5 timeline case.medh5 # visits and intervals
medh5 track case.medh5 --class lesion # per-lesion volumes across visits
medh5 dataset index studies/ -o cohort.json
medh5 dataset split cohort.json --group-by group_id --stratify-by site_id
medh5 dataset stats cohort.json --partition train --workers 8
medh5 dataset check cohort.json --deep
medh5 convert from-dicom /studies out/ # one sample per patient, all visits
medh5 convert from-nifti case.medh5 --image CT=ct.nii.gz
medh5 convert from-rtstruct plan.dcm case.medh5 --rasterize
medh5 migrate old/*.medh5 -o new/ --group-by subject
medh5 scrub out/*.medh5 --apply --date-shift-days -117
medh5 pack cohort/*.medh5 -o shard.medh5c
medh5 recompress cohort/*.medh5 --profile training
medh5 bench # reproduce the performance targets
medh5 conformance publish suite/ # the suite, for another implementation
Interoperability
| Format | |
|---|---|
| NIfTI | affine and voxels bit-identical on round trip; RAS↔LPS is a sign flip, never a resample |
| DICOM | slices ordered by geometry, spacing measured between origins, modality LUT stored not applied, tags on an explicit allow-list |
| DICOM SEG | frames placed by geometry; overlap and FRACTIONAL survive; segments matched by label, not number |
| RTSTRUCT | contours stay contours; rasterisation is opt-in and recorded in provenance |
| nnU-Net v2 | class ids kept; region labels become label-set DAG parents; dataset.json round-trips |
| MONAI | to_metatensor gives a MetaTensor with the correct affine |
| 0.x | medh5 migrate, reporting every decision and every guess |
Every conversion returns a report distinguishing what it decided from the data and where it guessed — the encoding chosen, the class ids minted, a half-voxel convention changed, a timepoint order inferred rather than read.
COCO is deliberately unsupported: it has no world geometry, spacing or frame of reference, so importing means inventing a grid and exporting means discarding the geometry that makes a medical annotation reproducible.
Reading it without medh5
import h5py, json, hdf5plugin # hdf5plugin only for blosc2 profiles
with h5py.File("case_0001.medh5") as f:
doc = json.loads(f["meta"][()])
doc["identity"]["subject_id"]
dict(f["grids"]["ct"].attrs) # spacing, origin, direction
f["images"]["CT"][10:20]
medh5 recompress --profile portable writes gzip, readable by any HDF5 build.
Versioning
The format is 1.0. A minor version may add optional objects, profiles, encodings and diagnostic codes; it may not change what an existing one means (spec §16). The package follows semantic versioning from 1.0.0.
0.x files are not readable by 1.0 and are not meant to be — medh5 migrate
converts them once. See Converters.
License
MIT
Release files for medh5 1.0.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| medh5-1.0.0.tar.gz | 510.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| medh5-1.0.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 853.2 kB
Release files / medh5-1.0.0.tar.gz
| Download URL | medh5-1.0.0.tar.gz |
|---|---|
| Size | 510.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
75c6358e3432ecd649d79f60e1f87fa5b1d129fd7bb1e096984a9c67464f5afc
|
|
BLAKE2b-256 checksum How to use checksums |
9bb39ea77b4142cdfaaba941663805e23e1c11cd6c06bc524cd8d44cd160b965
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 18, 2026.
Transparency logRelease files / medh5-1.0.0-py3-none-any.whl
| Download URL | medh5-1.0.0-py3-none-any.whl |
|---|---|
| Size | 342.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
036d9caef72ab39404dd723ccf6b6c7e04368d6577c85cb6875e56a21885bd1e
|
|
BLAKE2b-256 checksum How to use checksums |
29f31834917d4fb7c7cb8bfb2ab14f69b0ee79a25aba49227532536f17000581
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 18, 2026.
Transparency log