Introduction
Full documentation can be found at http://datamaestro.rtfd.io
This projects aims at grouping utilities to deal with the numerous and heterogenous datasets present on the Web. It aims at being
- a reference for available resources, listing datasets
- a tool to automatically download and process resources (when freely available)
- integration with the experimaestro experiment manager.
- (planned) a tool that allows to copy data from one computer to another
Each datasets is uniquely identified by a qualified name such as com.lecun.mnist, which is usually the inversed path to the domain name of the website associated with the dataset.
The main repository only deals with very generic processing (downloading, basic pre-processing and data types). Plugins can then be registered that provide access to domain specific datasets.
List of repositories
-
NLP datasets
Natural Language Processing (e.g. Sentiment101) and Information access (e.g. TREC) datasets -
image-related dataset
Image related datasets (e.g. MNIST)
-
machine learning
Generic machine learning datasets
Command line interface (CLI)
The command line interface allows to interact with the datasets. The commands are listed below, help can be found by typing datamaestro COMMAND --help:
searchsearch dataset by name, tags and/or tasksdownloaddownload files (if accessible on Internet) or ask for download path otherwisepreparedownload dataset files and outputs a JSON containing path and other dataset informationrepositorieslist the available repositoriesorphanslist data directories that do no correspond to any registered dataset (and allows to clean them up)create-datasetcreates a dataset definition
Example (CLI)
Retrieve and download
The commmand line interface allows to download automatically the different resources. Datamaestro extensions can provide additional processing tools.
$ datamaestro search tag:image
[image] com.lecun.mnist
$ datamaestro prepare com.lecun.mnist
INFO:root:Materializing 4 resources
INFO:root:Downloading https://ossci-datasets.s3.amazonaws.com/mnist/train-images-idx3-ubyte.gz into .../datamaestro/store/com/lecun/train_images.idx
INFO:root:Downloading https://ossci-datasets.s3.amazonaws.com/mnist/t10k-images-idx3-ubyte.gz into .../datamaestro/store/com/lecun/test_images.idx
INFO:root:Downloading https://ossci-datasets.s3.amazonaws.com/mnist/t10k-labels-idx1-ubyte.gz into .../datamaestro/store/com/lecun/test_labels.idx
The previous command also returns a JSON on standard output
{
"train": {
"images": {
"path": ".../data/image/com/lecun/mnist/train_images.idx"
},
"labels": {
"path": ".../data/image/com/lecun/mnist/train_labels.idx"
}
},
"test": {
"images": {
"path": ".../data/image/com/lecun/mnist/test_images.idx"
},
"labels": {
"path": ".../data/image/com/lecun/mnist/test_labels.idx"
}
},
"id": "com.lecun.mnist"
}
For those using Python, this is even better since the IDX format is supported
In [1]: from datamaestro import prepare_dataset
In [2]: ds = prepare_dataset("com.lecun.mnist")
In [3]: ds.train.images.data().dtype, ds.train.images.data().shape
Out[3]: (dtype('uint8'), (60000, 28, 28))
Python definition of datasets
Datasets are defined as Python classes with resource attributes that describe how to download and process data. The framework automatically builds a dependency graph and handles downloads with two-path safety and state tracking.
from datamaestro_image.data import ImageClassification, LabelledImages
from datamaestro.data.tensor import IDX
from datamaestro.download.single import FileDownloader
from datamaestro.definitions import Dataset, dataset
@dataset(url="http://yann.lecun.com/exdb/mnist/")
class MNIST(Dataset):
"""The MNIST database of handwritten digits."""
TRAIN_IMAGES = FileDownloader(
"train_images.idx",
"http://yann.lecun.com/exdb/mnist/train-images-idx3-ubyte.gz",
)
TRAIN_LABELS = FileDownloader(
"train_labels.idx",
"http://yann.lecun.com/exdb/mnist/train-labels-idx1-ubyte.gz",
)
TEST_IMAGES = FileDownloader(
"test_images.idx",
"http://yann.lecun.com/exdb/mnist/t10k-images-idx3-ubyte.gz",
)
TEST_LABELS = FileDownloader(
"test_labels.idx",
"http://yann.lecun.com/exdb/mnist/t10k-labels-idx1-ubyte.gz",
)
def config(self) -> ImageClassification:
return ImageClassification.C(
train=LabelledImages(
images=IDX(path=self.TRAIN_IMAGES.path),
labels=IDX(path=self.TRAIN_LABELS.path),
),
test=LabelledImages(
images=IDX(path=self.TEST_IMAGES.path),
labels=IDX(path=self.TEST_LABELS.path),
),
)
Its syntax is described in the documentation.
Release files for datamaestro 1.15.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| datamaestro-1.15.2.tar.gz | 303.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| datamaestro-1.15.2-py3-none-any.whl | Python 3 | none | any | Details |
Total release size:448.1 kB
Release files / datamaestro-1.15.2.tar.gz
| Download URL | datamaestro-1.15.2.tar.gz |
|---|---|
| Size | 303.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
2f43fe1a2e268d7b4daf79b0c5cef5a4379457f6fda78c069dbaf009144b5baa
|
|
BLAKE2b-256 checksum How to use checksums |
a67c51723d8a2dff520e3c29252a037a3776f635f41c8476fddff54175d7b5c7
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jul 29, 2026.
Transparency logRelease files / datamaestro-1.15.2-py3-none-any.whl
| Download URL | datamaestro-1.15.2-py3-none-any.whl |
|---|---|
| Size | 144.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
ade4856e2b77298b1c1f796668074405182a5a769718d6fb3c770820fdcc8032
|
|
BLAKE2b-256 checksum How to use checksums |
8e016b16679bf39dcc8b4f0a467d0da879ec812bf6a0a0d6c69886db545c70a6
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jul 29, 2026.
Transparency log