Skip to main content

MAITE Datasets

MAITE Datasets are a collection of public datasets wrapped in a MAITE compliant format.

Installation

To install and use maite-datasets you can use pip:

pip install maite-datasets

Extras

Optional features are installed as extras, e.g. pip install maite-datasets[tqdm].

Extra Enables
tqdm Status bar indicators when downloading.
hf Downloading datasets hosted on Hugging Face Hub and decoding their TIFF images.
av Decoding video into frames for multi-object tracking datasets.
datamaite Exporting and reading datasets via datamaite wire formats on disk.
all All of the above.

Available Downloadable Datasets

Task Dataset Description
Classification CIFAR10 CIFAR10 dataset.
Classification MilitaryAircraft 88 classes of military aircraft, cropped to one object per datum.
Classification MilitaryVehicles 24 classes of military vehicles. Requires the hf extra.
Classification MNIST A dataset of hand-written digits.
Classification Ships A dataset that focuses on identifying ships from satellite images.
Detection AntiUAVDetection A UAV detection dataset in natural images with varying backgrounds.
Detection DroneSwarm A synthetic multi-drone swarm dataset. Requires the hf extra.
Detection DroneVehicle Paired RGB-infrared drone imagery for vehicle detection.
Detection M3FD Paired RGB-infrared imagery for people and vehicle detection.
Detection MILCO A side-scan sonar dataset focused on mine-like object detection.
Detection MilitaryAircraft The detection version of the 88 class military aircraft dataset.
Detection SeaDrone A UAV dataset focused on open water object detection.
Detection SkySeaLand Satellite imagery of air, sea and land transportation.
Detection VOCDetection Pascal VOC dataset.

The paired RGB-infrared datasets (DroneVehicle, M3FD) stack the infrared channel onto the RGB image. MilitaryAircraft and MilitaryVehicles expose a hierarchy attribute describing their class ontology. Every dataset accepts lazy=True to defer image decoding until the image is actually read, which keeps metadata-only passes over large datasets cheap.

Usage

Here is an example of how to import MNIST for usage with your workflow.

>>> from maite_datasets.image_classification import MNIST

>>> mnist = MNIST(root="data", download=True)
>>> print(mnist)
MNIST Dataset
-------------
    Corruption: None
    Transforms: []
    Image_set: train
    Metadata: {'id': 'MNIST_train', 'index2label': {0: 'zero', 1: 'one', 2: 'two', 3: 'three', 4: 'four', 5: 'five', 6: 'six', 7: 'seven', 8: 'eight', 9: 'nine'}, 'split': 'train'}
    Path: /home/user/maite-datasets/data/mnist
    Size: 60000

>>> print("tuple("+", ".join([str(type(t)) for t in mnist[0]])+")")
tuple(<class 'numpy.ndarray'>, <class 'numpy.ndarray'>, <class 'dict'>)

datamaite Interoperability

With the datamaite extra installed, any dataset can be written to disk in a datamaite wire format (COCO or YOLO) and read back as a datamaite dataset.

>>> from maite_datasets.object_detection import MILCO

# Export an existing dataset, then load it back through datamaite
>>> milco = MILCO(root="data", download=True)
>>> dm_milco = milco.to_datamaite(dest="exports/milco")

# Or download, convert and load in one step
>>> dm_milco = MILCO(root="data", download=True, as_datamaite=True)

The format follows from the task and is not selectable, because each task only has one that works. Object detection writes COCO — the only one of the two that records a datum's own metadata, a dataset's telemetry as per-image columns and its per-object values as detection attributes. Image classification writes YOLO, because datamaite's COCO writer is task-closed and refuses an IC dataset outright.

Reading is more permissive than writing: an existing directory in either format is sniffed and loaded, so a YOLO dataset already on disk still works with as_datamaite=True.

as_datamaite=True writes to <root>/<dataset>_datamaite, kept separate from the <root>/<dataset> folder raw downloads use so the two can coexist under one root. An existing export there is reused; otherwise a raw dataset already under root is converted in place, and only a genuinely missing dataset is downloaded (to a temporary directory that is discarded after conversion).

Two conversion details are worth knowing. Splits are folded onto train/val/test, because those are the only ones the wire formats read back: an image_set such as operational is exported as train, with a warning, and its original name survives in each sample's metadata, which object detection keeps and image classification cannot. And datasets carrying transforms export re-encoded pixels rather than referencing the original files, so the written images always match the written annotations.

Dataset Wrappers

Wrappers provide a way to convert datasets to allow usage of tools within specific backend frameworks.

Torchvision

TorchvisionWrapper is a convenience class that wraps any of the datasets and provides the capability to apply torchvision transforms to the dataset.

NOTE: TorchvisionWrapper requires torch and torchvision to be installed.

>>> from maite_datasets.object_detection import MILCO

>>> milco = MILCO(root="data", download=True)
>>> print(milco)
MILCO Dataset
-------------
    Transforms: []
    Image Set: train
    Metadata: {'id': 'MILCO_train', 'index2label': {0: 'MILCO', 1: 'NOMBO'}, 'split': 'train'}
    Path: /home/user/maite-datasets/data/milco
    Size: 261

>>> print(f"type={milco[0][0].__class__.__name__}, shape={milco[0][0].shape}")
type=ndarray, shape=(3, 1024, 1024)

>>> print(milco[0][1].boxes[0])
[ 75. 217. 130. 247.]

>>> from maite_datasets.wrappers import TorchvisionWrapper
>>> from torchvision.transforms.v2 import Resize

>>> milco_torch = TorchvisionWrapper(milco, transforms=Resize(224))
>>> print(milco_torch)
Torchvision Wrapped MILCO Dataset
---------------------------
    Transforms: Resize(size=[224], interpolation=InterpolationMode.BILINEAR, antialias=True)

MILCO Dataset
-------------
    Transforms: []
    Image Set: train
    Metadata: {'id': 'MILCO_train', 'index2label': {0: 'MILCO', 1: 'NOMBO'}, 'split': 'train'}
    Path: /home/user/maite-datasets/data/milco
    Size: 261

>>> print(f"type={milco_torch[0][0].__class__.__name__}, shape={milco_torch[0][0].shape}")
type=Image, shape=torch.Size([3, 224, 224])

>>> print(milco_torch[0][1].boxes[0])
tensor([16.4062, 47.4688, 28.4375, 54.0312], dtype=torch.float64)

Dataset Adapters

Adapters provide a way to read in datasets from other popular formats.

Huggingface

Hugging face datasets can be adapted into MAITE compliant format using the from_huggingface adapter.

>>> from datasets import load_dataset
>>> from maite_datasets.adapters import from_huggingface

>>> cppe5 = load_dataset("cppe-5")
>>> m_cppe5 = from_huggingface(cppe5["train"])
>>> print(m_cppe5)
HFObjectDetection Dataset
-------------------------
    Source: Dataset({
    features: ['image_id', 'image', 'width', 'height', 'objects'],
    num_rows: 1000
})
    Metadata: {'id': 'cppe-5', 'index2label': {0: 'Coverall', 1: 'Face_Shield', 2: 'Gloves', 3: 'Goggles', 4: 'Mask'}, 'description': '', 'citation': '', 'homepage': '', 'license': '', 'features': {'image_id': Value('int64'), 'image': Image(mode=None, decode=True), 'width': Value('int32'), 'height': Value('int32'), 'objects': {'id': List(Value('int64')), 'area': List(Value('int64')), 'bbox': List(List(Value('float32'), length=4)), 'category': List(ClassLabel(names=['Coverall', 'Face_Shield', 'Gloves', 'Goggles', 'Mask']))}}, 'post_processed': None, 'supervised_keys': None, 'builder_name': 'parquet', 'dataset_name': 'cppe-5', 'config_name': 'default', 'version': 0.0.0, 'splits': {'train': SplitInfo(name='train', num_bytes=240478590, num_examples=1000, shard_lengths=None, dataset_name='cppe-5'), 'test': SplitInfo(name='test', num_bytes=4172706, num_examples=29, shard_lengths=None, dataset_name='cppe-5')}, 'download_checksums': {'hf://datasets/cppe-5@66f6a5efd474e35bd7cb94bf15dea27d4c6ad3f8/data/train-00000-of-00001.parquet': {'num_bytes': 237015519, 'checksum': None}, 'hf://datasets/cppe-5@66f6a5efd474e35bd7cb94bf15dea27d4c6ad3f8/data/test-00000-of-00001.parquet': {'num_bytes': 4137134, 'checksum': None}}, 'download_size': 241152653, 'post_processing_size': None, 'dataset_size': 244651296, 'size_in_bytes': 485803949}

>>> image = m_cppe5[0][0]
>>> print(f"type={image.__class__.__name__}, shape={image.shape}")
type=ndarray, shape=(3, 663, 943)

>>> target = m_cppe5[0][1]
>>> print(f"box={target.boxes[0]}, label={target.labels[0]}")
box=[302.0, 109.0, 73.0, 52.0], label=4

>>> print(m_cppe5[0][2])
{'id': [114, 115, 116, 117], 'image_id': 15, 'width': 943, 'height': 663, 'area': [3796, 1596, 152768, 81002]}

Additional Information

For more information on the MAITE protocol, check out their documentation.

Acknowledgement

CDAO Funding Acknowledgement

This material is based upon work supported by the Chief Digital and Artificial Intelligence Office under Contract No. W519TC-23-9-2033. The views and conclusions contained herein are those of the author(s) and should not be interpreted as necessarily representing the official policies or endorsements, either expressed or implied, of the U.S. Government.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

maite_datasets-0.0.22.tar.gz (136.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

maite_datasets-0.0.22-py3-none-any.whl (165.8 kB view details)

Uploaded Python 3

File details

Details for the file maite_datasets-0.0.22.tar.gz.

File metadata

  • Download URL: maite_datasets-0.0.22.tar.gz
  • Upload date:
  • Size: 136.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: uv/0.12.10 {"installer":{"name":"uv","version":"0.12.10","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

File hashes

Hashes for maite_datasets-0.0.22.tar.gz
Algorithm Hash digest
SHA256 765ea871d7da61dd717b0d81896cdd1fcef0210d01ca27adde5f763bf549b26e
MD5 bdd80555e7644dd3905f901ca821598e
BLAKE2b-256 565e49adf68c4b8c1e65895703c302d459df235fc66429a539c200fa0440d1ed

See more details on using hashes here.

File details

Details for the file maite_datasets-0.0.22-py3-none-any.whl.

File metadata

  • Download URL: maite_datasets-0.0.22-py3-none-any.whl
  • Upload date:
  • Size: 165.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: uv/0.12.10 {"installer":{"name":"uv","version":"0.12.10","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

File hashes

Hashes for maite_datasets-0.0.22-py3-none-any.whl
Algorithm Hash digest
SHA256 180cac0d0980e4c7e9df637d39dff9ac89360d2fc880699751561c5a3b0927be
MD5 38592f40b159dd396a55a9d38d012c74
BLAKE2b-256 9c248e4c0b8ea3a1d16de1bccc262d8aec3c6edef665daba804a56d01005043f

See more details on using hashes here.

Release history Release notifications | RSS feed

This release

0.0.22 This release

2 files

0.0.21

2 files

0.0.20

2 files

0.0.19

2 files

0.0.18

2 files

0.0.17

2 files

0.0.16

2 files

0.0.15

2 files

0.0.14

2 files

0.0.13

2 files

0.0.12

2 files

0.0.11

2 files

0.0.10

2 files

0.0.9

2 files

0.0.8

2 files

0.0.7

2 files

0.0.6

2 files

0.0.5

2 files

0.0.4

2 files

0.0.3

2 files

0.0.2

2 files

0.0.1

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page