Skip to main content

HistoPrep

Preprocessing large medical images for machine learning made easy!

Description • Installation • Usage • API Documentation • Citation

Description

HistoPrep makes is easy to prepare your histological slide images for deep learning models. You can easily cut large slide images into smaller tiles and then preprocess those tiles (remove tiles with shitty tissue, finger marks etc).

Installation

Install OpenSlide on your system and then install histoprep with pip!

pip install histoprep

Usage

Typical workflow for training deep learning models with histological images is the following:

  1. Cut each slide image into smaller tile images.
  2. Preprocess smaller tile images by removing tiles with bad tissue, staining artifacts.
  3. Overfit a pretrained ResNet50 model, report 100% validation accuracy and publish it in Nature like everyone else.

With HistoPrep, steps 1. and 2. are as easy as accidentally drinking too much at the research group christmas party and proceeding to work remotely until June.

Let's start by cutting a slide from the PANDA kaggle challenge into small tiles.

from histoprep import SlideReader

# Read slide image.
reader = SlideReader("./slides/slide_with_ink.jpeg")
# Detect tissue.
threshold, tissue_mask = reader.get_tissue_mask(level=-1)
# Extract overlapping tile coordinates with less than 50% background.
tile_coordinates = reader.get_tile_coordinates(
    tissue_mask, width=512, overlap=0.5, max_background=0.5
)
# Save tile images with image metrics for preprocessing.
tile_metadata = reader.save_regions(
    "./train_tiles/", tile_coordinates, threshold=threshold, save_metrics=True
)
slide_with_ink: 100%|██████████| 390/390 [00:01<00:00, 295.90it/s]

Let's take a look at the output and visualise the thumbnails.

jopo666@~$ tree train_tiles
train_tiles
└── slide_with_ink
    ├── metadata.parquet       # tile metadata
    ├── properties.json        # tile properties
    ├── thumbnail.jpeg         # thumbnail image
    ├── thumbnail_tiles.jpeg   # thumbnail with tiles
    ├── thumbnail_tissue.jpeg  # thumbnail of the tissue mask
    └── tiles [390 entries exceeds filelimit, not opening dir]

Prostate biopsy sample Tissue mask Thumbnail with tiles

That was easy, but it can be annoying to whip up a new python script every time you want to cut slides, and thus it is recommended to use the HistoPrep CLI program!

# Repeat the above code for all images in the PANDA dataset!
jopo666@~$ HistoPrep --input './train_images/*.tiff' --output ./tiles --width 512 --overlap 0.5 --max-background 0.5

As we can see from the above images, histological slide images often contain areas that we would not like to include into our training data. Might seem like a daunting task but let's try it out!

from histoprep.utils import OutlierDetector

# Let's wrap the tile metadata with a helper class.
detector = OutlierDetector(tile_metadata)
# Cluster tiles based on image metrics.
clusters = detector.cluster_kmeans(num_clusters=4, random_state=666)
# Visualise first cluster.
reader.get_annotated_thumbnail(
    image=reader.read_level(-1), coordinates=detector.coordinates[clusters == 0]
)

Tiles in cluster 0

I said it was gonna be easy! Now we can mark tiles in cluster 0 as outliers and start overfitting our neural network! This was a simple example but the same code can be used to cluster all several million tiles extracted from the PANDA dataset and discard outliers simultaneously!

Citation

If you use HistoPrep to process the images for your publication, please cite the github repository.

@misc{histoprep,
  author = {Pohjonen, Joona and Ariotta, Valeria},
  title = {HistoPrep: Preprocessing large medical images for machine learning made easy!},
  year = {2022},
  publisher = {GitHub},
  journal = {GitHub repository},
  howpublished = {https://github.com/jopo666/HistoPrep},
}

Release files for histoprep 2.0.5

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for histoprep 2.0.5
File Size Uploaded
histoprep-2.0.5.tar.gz 35.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for histoprep 2.0.5
File Interpreter ABI Platform
histoprep-2.0.5-py3-none-any.whl Python 3 none any Details

Total release size: 77.9 kB

Release files / histoprep-2.0.5.tar.gz

Download URL histoprep-2.0.5.tar.gz
Size 35.4 kB
Tags Source
SHA-256 checksum
How to use checksums
a3b5495db53e0701d911adef79427473e141ff131fc79247e65b41114450be76
BLAKE2b-256 checksum
How to use checksums
92e93078708714503b8e222e312c7393a16876d0f0f76ffc9123b7fc0201acbe
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via poetry/1.4.2 CPython/3.11.3 Linux/5.15.0-72-generic

Release files / histoprep-2.0.5-py3-none-any.whl

Download URL histoprep-2.0.5-py3-none-any.whl
Size 42.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
19c0de3090878425fcdcc4a8edce81fb826848d76d3e315a7b2b4ca128c6bc9a
BLAKE2b-256 checksum
How to use checksums
4f597d2e1bf243ec17fbb75c373d3d4b50fa7d088a9c2caa2710167d92dd2b62
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via poetry/1.4.2 CPython/3.11.3 Linux/5.15.0-72-generic

Release history Release notifications | RSS feed

This release

2.0.5 This release

2 release files

2.0.4

2 release files

2.0.3

2 release files

2.0.2

2 release files

2.0.1

2 release files

1.0.8

2 release files

1.0.7

2 release files

1.0.6

2 release files

1.0.5

2 release files

1.0.4

2 release files

1.0.3

2 release files

1.0.2

2 release files

1.0.1

2 release files

1.0.0

2 release files

0.0.2.0

1 release file

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page