Skip to main content

Process DCM

Maintenance GitHub GitHub release (latest by date) GitHub Release PyPI Python uv Ruff ty

About The Project

Python library and app to extract images from DCM files with metadata in a JSON-based standard format.

It targets ophthalmic DICOMs (OCT, fundus photography, fluorescein angiography, SLO, Optomap, ...), writes one folder per acquisition with the extracted frames as PNG/JPG/WEBP plus a metadata.json, and can anonymise patient identifiers while keeping a study_id -> patient_id mapping.

Installation and Usage

pip install process-dcm
# or, without touching your environment
uvx process-dcm --help
 Usage: process-dcm [OPTIONS] {input_path}

 Process DICOM files in subfolders, extract images and metadata.

 Version: 0.10.0

╭─ Arguments ───────────────────────────────────────────────────────────────────────────────────────────────────────────────╮
│ *    input_path      <path>  Input path to either a DCM file or a folder containing DICOM files. [required]               │
╰───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯
╭─ Options ─────────────────────────────────────────────────────────────────────────────────────────────────────────────────╮
│ --image_format        -f      <str>    Image format for extracted images (png, jpg, webp). [default: png]                 │
│ --output_dir          -o      <path>   Output directory for extracted images and metadata. [default: exported_data]       │
│ --group               -g               Re-group DICOM files in a given folder by AcquisitionDateTime.                     │
│ --tol                 -t      <float>  Tolerance in seconds for grouping DICOM files by AcquisitionDateTime. Only used    │
│                                        when --group is set.                                                               │
│ --n_jobs              -j      <int>    Number of parallel jobs. [default: 1]                                              │
│ --mapping             -m      <str>    Path to CSV containing patient_id to study_id mapping. If not provided and         │
│                                        patient_id is anonymised, a 'study_2_patient.csv' file will be generated.          │
│ --keep                -k      <str>    Keep the specified fields (p: patient_key, n: names, d: date_of_birth, D:          │
│                                        year-only DOB, g: gender)                                                          │
│ --preserve_folder_structure  -p        Mirror the input folder structure under the output directory instead of the flat   │
│                                        '{patient}_{date}_{hash}_{eye}_{modality}.DCM' folders. Not compatible with        │
│                                        --group or --reset.                                                                │
│ --keep_dcm_name_as_folder / --no_keep_dcm_name_as_folder                                                                  │
│                                        With --preserve_folder_structure, write each DICOM's images into a folder named    │
│                                        after the file. Disable to write all acquisitions of an input folder into one      │
│                                        output folder. [default: keep_dcm_name_as_folder]                                  │
│ --relative_source_file                 Write metadata 'source_file' relative to INPUT_PATH instead of the current working │
│                                        directory.                                                                         │
│ --overwrite           -w               Overwrite existing images if found.                                                │
│ --reset               -r               Reset the output directory if it exists.                                           │
│ --quiet               -q               Silence verbosity.                                                                  │
│ --version             -V               Prints app version.                                                                │
│ --install-completion                   Install completion for the current shell.                                          │
│ --show-completion                      Show completion for the current shell, to copy it or customize the installation.   │
│ --help                -h               Show this message and exit.                                                        │
╰───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯

Output layout

By default every acquisition group (DICOMs sharing a FrameOfReferenceUID, or an acquisition time with --group) is written to one flat, self-describing folder under the output directory:

exported_data/
└── {PatientKey}_{Date}_{Time}_{Hash}_{Laterality}_{Modality}.DCM/
    ├── {Modality}-{GroupID}_{FrameIndex}.{png|jpg|webp}
    └── metadata.json

--preserve_folder_structure (-p) mirrors the input tree instead. The group is written under the relative folder of its first DICOM, in a leaf folder named after that file (drop the leaf with --no_keep_dcm_name_as_folder):

exported_data/
└── {relative input folder}/{dicom file stem}/
    ├── {Modality}-{GroupID}_{FrameIndex}.{png|jpg|webp}
    └── metadata.json

This layout cannot be combined with --group or --reset. In both layouts each image entry in metadata.json records its source_file relative to the current working directory; --relative_source_file makes it relative to INPUT_PATH instead. The metadata format is versioned by parser_version (currently 1.7.0). Each image entry carries its own laterality and scan_datetime, because a DICOM group may hold scans of both eyes taken at different times; the series-level laterality is B and the exam-level scan_datetime is the earliest one in that case.

Project structure

process-dcm/
├── .github/
│   └── workflows/
│       ├── ci.yml              # Calls reusable CI from e2g-workflows
│       └── release.yml         # Semantic release + publish on main
├── .kiro/
│   └── steering/
│       ├── commits.md          # Conventional commit rules for Kiro
│       ├── python.md           # Python coding standards for Kiro
│       └── safety.md           # Command safety guardrails for Kiro
├── .vscode/
│   ├── extensions.json         # Recommended VS Code/Kiro extensions
│   └── settings.json           # Editor settings (Ruff, pytest, etc.)
├── src/
│   └── process_dcm/
│       ├── __init__.py         # Package version via importlib.metadata
│       ├── __main__.py         # python -m support
│       ├── main.py             # Typer CLI entry point (`process-dcm`)
│       ├── const.py            # ImageModality / ModalityFlag enums
│       ├── py.typed            # PEP 561 typing marker
│       └── utils.py            # DICOM processing, grouping, metadata
├── tests/
│   ├── conftest.py             # Shared pytest fixtures and warning filters
│   ├── test_*.py
│   └── <sample DICOMs>         # Small fixtures committed to git
├── .cruft.json                 # Links project to e2g-pypkg template
├── .editorconfig               # Editor-agnostic formatting rules
├── .gitignore
├── .python-version             # Pin Python version for uv
├── CLAUDE.md                   # AI rules for Claude CLI/VS Code
├── README.md
├── justfile                    # Command runner (just qa, just test, etc.)
└── pyproject.toml              # Project metadata, deps, tool config

Key files:

  • .github/workflows/ — CI calls the shared e2g-workflows reusable workflow. CI logic is centralised there.
  • .kiro/steering/ — AI steering rules loaded automatically by Kiro. Enforces team coding standards.
  • CLAUDE.md — Same AI rules for Claude CLI/VS Code users, plus project-specific notes.
  • .vscode/ — Shared editor settings (Ruff format-on-save, pytest discovery, recommended extensions). Works in VS Code and Kiro out of the box.
  • justfile — Single interface for all dev commands. CI uses the same recipes, so local and CI behaviour match.
  • .cruft.json — Template tracking. Run cruft update to pull improvements from e2g-pypkg.

Development

Requirements: uv. just is installed into the environment by uv sync (rust-just), so uv run just ... always works; a system-wide just is optional.

# Clone the repo
git clone git@github.com:eye2Gene/process-dcm.git
cd process-dcm

# Create the virtual environment (uses .python-version -> 3.12)
uv venv

# Activate the virtual environment
source .venv/bin/activate

# Install the project (editable) and all dev dependency groups
uv sync

Python 3.11 is the minimum supported version; development and CI default to 3.12, and CI also tests 3.11 and 3.13.

Run tests (parallel, with coverage):

just test             # or: uv run pytest
just test -k optomap  # pass any pytest args through
just pdb              # single process, drop into the debugger on failure

Run quality checks (format, lint, type check with ty, dependency audit, then tests):

just qa       # fixes what it can
just ci       # check only, what CI runs
just qa-all   # qa + tests

Test data

Most fixtures are small DICOMs committed under tests/. The larger sample set tests/example_dir (about 625 MB) is not in git (see .gitignore); the tests that depend on it are skipped automatically when the folder is absent, both locally and in CI. If you need them, ask the maintainers for a copy and unpack it at tests/example_dir/.

Maintainers may also keep a tests_local/ folder, ignored by git, with smoke tests that run process-dcm on vendor-conversion outputs: Topcon FDA and Heidelberg E2E files exported to DICOM with OCT-Converter, plus a Heidelberg DICOMDIR export. That data cannot be shared, so the folder exists only on maintainers' machines and is not part of just test; run it explicitly with uv run pytest tests_local. Its own README explains where the data comes from and how to regenerate it.

Several tests compare MD5 hashes of generated PNGs. PNG bytes depend on the Pillow encoder version even when the pixels are identical, so the expected hash lists accept one value per known encoder. If a dependency bump changes a hash, verify the pixels are unchanged before adding the new value.

How releases work

This project uses conventional commits and python-semantic-release:

  1. Develop on a feature branch with conventional commit messages (feat:, fix:, docs:, etc.)
  2. Open a PR and merge to main when tests pass
  3. On merge, GitHub Actions automatically:
    • Determines the next version from commit messages
    • Updates the changelog
    • Creates a git tag and GitHub Release
    • Builds and publishes the package

Never edit the version number by hand: it lives only in pyproject.toml and is read at runtime via importlib.metadata (process_dcm.__version__).

Publishing to PyPI

This project publishes to PyPI.org using Trusted Publishing (OIDC). No tokens needed — the repo is configured as a trusted publisher on PyPI (owner eye2Gene, repository process-dcm, workflow release.yml).

AI-assisted development

This project includes configuration for AI coding assistants:

  • Kiro: .kiro/steering/ contains rules for Python style, commit messages, and command safety
  • Claude: CLAUDE.md contains the same rules in Claude's format

These are committed to git so all contributors get consistent AI behaviour. No personal setup required — just open the project in Kiro or Claude and the rules apply automatically.

Keeping up with template updates

This project follows the e2g-pypkg template and uses cruft to stay in sync with template improvements.

Check if there are template updates available:

uv run cruft check

See what would change:

uv run cruft diff

Apply template updates to your project:

uv run cruft update

If there are merge conflicts, cruft will create .rej files showing the rejected changes. Resolve them manually, then commit.

Tip: Run cruft update on a clean branch so you can review the changes in a PR.

Author

Process DCM was created in 2024 by Alan Wilter at the Moorfields Ophthalmic Reading Centre & Clinical AI Lab and is maintained by Eye2Gene. Licensed under the MIT License.

Aligned with the eye2Gene/e2g-pypkg project template.

Metadata

Release files for process-dcm 1.0.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for process-dcm 1.0.0
File Size Uploaded
process_dcm-1.0.0.tar.gz 24.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for process-dcm 1.0.0
File Interpreter ABI Platform
process_dcm-1.0.0-py3-none-any.whl Python 3 none any Details

Total release size: 48.9 kB

Release files / process_dcm-1.0.0.tar.gz

Download URL process_dcm-1.0.0.tar.gz
Size 24.2 kB
Tags Source
SHA-256 checksum
How to use checksums
0cfec914d5394c1ac428161271c4cdb847c8fbefb4d4dea790d5cb1f2e10fcf5
BLAKE2b-256 checksum
How to use checksums
c7cc71ec1bdcd85e789ec8d4cf50f9616ad2c808f16a9f8cea9b35d9ad72dbaf
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 6, 2026.

Transparency log

Release files / process_dcm-1.0.0-py3-none-any.whl

Download URL process_dcm-1.0.0-py3-none-any.whl
Size 24.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
e00115d12203f221aa1d6603762849d30ba132330f6a24e4d644aa4cd00887e5
BLAKE2b-256 checksum
How to use checksums
12ecd10375bf4954650cae71e315e3baa79521b72138c50b375413bac7d1c9fa
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 6, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

1.0.0 This release

2 release files

0.9.0

2 release files

0.8.0

2 release files

0.7.0

2 release files

0.6.1

2 release files

0.6.0

2 release files

0.5.0

2 release files

0.4.9

2 release files

0.4.8

2 release files

0.4.7

2 release files

0.4.6

2 release files

0.4.5

2 release files

0.4.4

2 release files

0.4.3

2 release files

0.4.2

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.1

2 release files

0.1.0

2 release files

0.0.2

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page