Skip to main content

MultiClean

PyPI Conda Python 3.9+ License: MIT Tutorials

MultiClean is a Python library for morphological cleaning of multiclass 2D numpy arrays (segmentation masks and classification rasters). It provides efficient tools for edge smoothing and small-island removal across multiple classes, then fills gaps using the nearest valid class.

Visual Example

Below: Land Use before/after cleaning (smoothed edges, small-island removal, nearest-class gap fill).

Land Use before/after

Installation

pip install multiclean

or

uv add multiclean

or

conda install -c conda-forge multiclean

Quick Start

import numpy as np
from multiclean import clean_array

# Create a sample classification array with classes 0, 1, 2, 3
array = np.random.randint(0, 4, (1000, 1000), dtype=np.int32)

# Clean with default parameters
cleaned = clean_array(array)

# Custom parameters
cleaned = clean_array(
    array,
    class_values=[0, 1, 2, 3],
    smooth_edge_size=2,     # kernel width, larger value increases smoothness
    min_island_size=100,    # remove components with area < 100
    connectivity=8,         # 4 or 8
    max_workers=4,
    fill_nan=False          # enable/disable the filling of nan values in input array
)

Use Cases

MultiClean is designed for cleaning segmentation outputs from:

  • Remote sensing: Land cover classification, crop mapping
  • Computer vision: Semantic segmentation post-processing
  • Geospatial analysis: Raster classification cleaning
  • Machine learning: Neural network output refinement

Key Features

  • Multi-class processing: Clean all classes in one pass
  • Edge smoothing: Morphological opening to reduce jagged boundaries
  • Island removal: Remove small connected components per class
  • Gap filling: Fill invalids via nearest valid class (distance transform)
  • Fast: NumPy + OpenCV with parallelism

How It Works

MultiClean uses morphological operations to clean classification arrays:

  1. Edge smoothing (per class): Morphological opening with a circular kernel.
  2. Island removal (per class): Find connected components (OpenCV) and mark components with area < min_island_size as invalid.
  3. Gap filling: Compute a distance transform to copy the nearest valid class into invalid pixels.

Classes are processed together and the result maintains a valid label at every pixel.

API Reference

clean_array

from multiclean import clean_array

out = clean_array(
    array: np.ndarray,
    class_values: int | list[int] | None = None,
    smooth_edge_size: int = 2,
    min_island_size: int = 100,
    connectivity: int = 4,
    max_workers: int | None = None,
    fill_nan: bool = False
)
  • array: 2D numpy array of class labels (int or float). For float arrays, NaN is treated as nodata and will remain NaN unless fill_nan is set to True.
  • class_values: Classes to consider. If None, inferred from array (ignores NaN for floats). An int restricts cleaning to a single class.
  • smooth_edge_size: Kernel size (pixels) for morphological opening. Use 0 to disable.
  • min_island_size: Remove components with area strictly < min_island_size. Use 1 to keep single pixels.
  • connectivity: Pixel connectivity for components, 4 or 8.
  • max_workers: Parallelism for per-class operations (None lets the executor choose).
  • fill_nan: If True will fill NAN values from input array with nearest valid value.

Returns a numpy array matching the input shape and dtype. Float arrays with NaN are supported and can be filled or remain as NaN.

Examples

Cleaning Land Cover Data

from multiclean import clean_array
import rasterio

# Read land cover classification
with rasterio.open('landcover.tif') as src:
    landcover = src.read(1)

# Clean with appropriate parameters for satellite data
cleaned = clean_array(
    landcover,
    class_values=[0, 1, 2, 3, 4],  # forest, water, urban, crop, other
    smooth_edge_size=1,
    min_island_size=25,
    connectivity=8,
    fill_nan=False
)

Cleaning Neural Network Segmentation Output

from multiclean import clean_array

# Model produces logits; convert to class predictions
np_pred = np_model_logits.argmax(axis=0)  # shape: (H, W)

# Clean the segmentation
cleaned = clean_array(
    np_pred,
    smooth_edge_size=2,
    min_island_size=100,
    connectivity=4,
)

Notebooks

See the notebooks folder for end-to-end examples:

Try in Colab

Colab_Button

Changelog

Release notes and the full version history are kept in CHANGELOG.md.

Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

Maintainers: see RELEASING.md for how to cut a release.

License

This project is licensed under the MIT License - see the LICENSE file for details.

Metadata

Release files for multiclean 0.5.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for multiclean 0.5.0
File Size Uploaded
multiclean-0.5.0.tar.gz 4.1 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for multiclean 0.5.0
File Interpreter ABI Platform
multiclean-0.5.0-py3-none-any.whl Python 3 none any Details

Total release size: 4.1 MB

Release files / multiclean-0.5.0.tar.gz

Download URL multiclean-0.5.0.tar.gz
Size 4.1 MB
Tags Source
SHA-256 checksum
How to use checksums
64f88200e28cd257da721256ea5b8a99f66fd7d42c20d5a04ae5b7ae80416b57
BLAKE2b-256 checksum
How to use checksums
c8e3db37f1a6a459c160e7ddfd60ad89729ed2b39be181ef21e2d66f762af597
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 14, 2026.

Transparency log

Release files / multiclean-0.5.0-py3-none-any.whl

Download URL multiclean-0.5.0-py3-none-any.whl
Size 11.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
8485d44226d916f32d48ae9cf106a50a63b704d83b389d2ba1a609484c8182f0
BLAKE2b-256 checksum
How to use checksums
a5a2a74853aec68b5a2541b1d8a7fc18ce206480314d5d8b7d97423150abba42
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 14, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.5.0 This release

2 release files

0.4.0

2 release files

0.3.1

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page