Skip to main content

Project logo

Published Docs Tests

Mozilla Data Collective Python Client Library

The official Python SDK for accessing and contributing to the Mozilla Data Collective platform.

[!WARNING] Our platform is evolving rapidly. Expect breaking changes while the Python SDK is on 0.X.X versions. Please ensure you are always on the latest version available.

Installation

pip install datacollective

Quick Start

IMPORTANT NOTE: Before trying to access any dataset, make sure you have thoroughly read and agreed to the specific dataset's conditions & licensing terms.

  1. Get your API key from the Mozilla Data Collective dashboard

  2. Set the API key in your environment variable:

Option A: Run this command in your terminal (replace your-api-key-here with your actual API key):

export MDC_API_KEY=your-api-key-here

Option B: Create a .env file in your project directory and add this line:

MDC_API_KEY=your-api-key-here
  1. Get your dataset ID and/or slug from the dataset's page at the MDC website.

[!TIP] You can find the dataset's ID and/or slug from its page on the MDC platform. Click the Download button and select API / Python Access, then switch to the Python Library tab on the pop-up modal. The Dataset ID and Dataset Slug appear at the bottom of the modal. Both can be used interchangeably in the Python library, but the slug is a more user-friendly identifier.

  1. Save a dataset locally:
from datacollective import download_dataset

dataset_path = download_dataset("your-dataset-id")

[!NOTE] download_dataset was previously called save_dataset_to_disk. The old name still works for backward compatibility, but it is deprecated and new code should use download_dataset.

[!TIP] Automatic Resume: If a download is interrupted (e.g., due to a network error or it gets stopped it manually), the next time you try download the same dataset at the same folder location, we will automatically resume from where the download left off!

[!TIP] Set enable_logging=True to emit detailed SDK logs to the console and a local log file at ~/.mozdata/datacollective.log with timestamped entries, a per-session id, and retention of 5 backup files at 10 MB each.

  1. Get information & metadata about a dataset:
from datacollective import get_dataset_details

details = get_dataset_details("your-dataset-id")

[!TIP] Don't know the dataset ID yet? Browse or search the public catalog straight from Python, no API key needed:

from datacollective import list_datasets, list_dataset_filters

filters = list_dataset_filters()      # available tasks, locales, licenses and formats
page = list_datasets("swahili speech", task="ASR", results_per_page=10)
print(page.total)                      # matches across all pages
for dataset in page:
    print(dataset.id, dataset.slug, dataset.name)
  1. Load the dataset into a pandas DataFrame (Alpha version: Only certain MDC datasets are supported right now):
from datacollective import load_dataset

dataset = load_dataset("your-dataset-id")
  1. Or load it as a HuggingFace Dataset object (requires the optional hf extra: pip install / uv add "datacollective[hf]"):
from datacollective import load_dataset

dataset = load_dataset("your-dataset-id", return_format="hf")

Returns a Dataset, or a DatasetDict keyed by split name for datasets with multiple splits. See our docs for more details, including how to lazily decode audio with the Audio() feature.

Programmatic submissions and uploads

[!ΝΟΤΕ] In order to be able to upload datasets in the MDC platform you will first need to Request Access to Upload by navigating to your profile under the Upload tab. Only the credentials (API keys) created after your request has been approved will be able to upload datasets. Any credentials created before your request was approved will not be able to upload datasets.

You can create dataset submissions and upload files with resumable uploads into the MDC platform programmatically using our Python SDK:

from datacollective import DatasetSubmission, License, Task, create_submission_with_upload

submission = DatasetSubmission(
    name="Dataset Name",
    longDescription="A detailed description of the dataset.",
    shortDescription="A brief description of the dataset.",
    locale="en-US",
    task=Task.ASR,
    format="TSV",
    licenseAbbreviation=License.CC_BY_4_0,
    other="This text should provide a detailed description of the dataset, "
          "including its contents, structure, and any relevant information "
          "that would help users understand what the dataset is about "
          "and how it can be used.",
    restrictions="Any restrictions you want to impose on the dataset",
    forbiddenUsage="Use cases that are not allowed with this dataset",
    additionalConditions="Any additional conditions for using the dataset",
    pointOfContactFullName="Jane Doe",
    pointOfContactEmail="jane@example.com",
    fundedByFullName="Funder Name",
    fundedByEmail="funder@example.com",
    legalContactFullName="Legal Name",
    legalContactEmail="legal@example.com",
    createdByFullName="Creator Name",
    createdByEmail="creator@example.com",
    intendedUsage="Describe the intended usage of the dataset, including "
                  "potential applications and use cases.",
    ethicalReviewProcess="Describe the ethical review process that was "
                         "followed for this dataset, including any approvals "
                         "or considerations related to data collection and usage.",
    isPaid=False,  # True = the dataset is compensated and requires `basePriceCents`,
                   # False (default) = the dataset is free to access
    exclusivityOptOut=False,  # True = This dataset is non-exclusive to Mozilla Data Collective, 
                              # False = Dataset is exclusively hosted in Mozilla Data Collective
    agreeToSubmit=True,  # True = You confirm that you have the right to submit this dataset and 
                         # that all information provided in the datasheet is accurate. 
                         # Required to be True to complete the submission process
)

response = create_submission_with_upload(
    file_path="/path/to/dataset.tar.gz",
    submission=submission
)

print(response)

For predefined licenses, pass licenseAbbreviation=License.<VALUE> and leave licenseUrl and license unset. For custom licenses, pass a custom string to license and optionally include licenseUrl and licenseAbbreviation.

To publish a compensated dataset, set isPaid=True and a basePriceCents price in USD cents (US Dollars), e.g. basePriceCents=100_000 for $1,000.00.

[!TIP] To also attach an optional sample of your dataset, pass sample_file_path="/path/to/dataset-sample.tar.gz" to create_submission_with_upload, or upload it separately with upload_sample_file(file_path=..., submission_id=...).

[!TIP] To upload a new .tar.gz version to an already approved dataset, call upload_dataset_file(file_path=..., submission_id=...) directly. Find the submission under Profile → Uploads, open the approved dataset, and copy the value after /profile/submissions/ in the URL. Note that this value is the submission ID, which is different from the public dataset ID.

For more details, visit our docs

License

This project is released under MPL (Mozilla Public License) 2.0.

Metadata

Release files for datacollective 0.6.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for datacollective 0.6.2
File Size Uploaded
datacollective-0.6.2.tar.gz 51.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for datacollective 0.6.2
File Interpreter ABI Platform
datacollective-0.6.2-py3-none-any.whl Python 3 none any Details

Total release size: 116.8 kB

Release files / datacollective-0.6.2.tar.gz

Download URL datacollective-0.6.2.tar.gz
Size 51.3 kB
Tags Source
SHA-256 checksum
How to use checksums
b1539aba596426fb0321c84bfb8497c44248ee88c180538aea0dfcb46d3e68ce
BLAKE2b-256 checksum
How to use checksums
eee899c0e5d41629c98a1c185e8f0cf912c6d3968dd6a44b55216053bd1990b8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / datacollective-0.6.2-py3-none-any.whl

Download URL datacollective-0.6.2-py3-none-any.whl
Size 65.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
851bca8c9f29817386f4f98588dd3549b25d632fa7c4780023817370289af91c
BLAKE2b-256 checksum
How to use checksums
973f0915a63f4e90ed738dcccbeb050fe8209d61b21b87a2603bfdd153598638
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release history Release notifications | RSS feed

0.6.4

2 release files

0.6.3

2 release files

This release

0.6.2 This release

2 release files

0.6.1

2 release files

0.6.0

2 release files

0.5.7

2 release files

0.5.6

2 release files

0.5.5

2 release files

0.5.4

2 release files

0.5.3

2 release files

0.5.2

2 release files

0.5.1

2 release files

0.5.0

2 release files

0.4.5

2 release files

0.4.4

2 release files

0.4.3

2 release files

0.4.2

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.0

2 release files

0.0.34

2 release files

0.0.33

2 release files

0.0.32

2 release files

0.0.27

2 release files

0.0.23

2 release files

0.0.16

2 release files

0.0.11

2 release files

0.0.8

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page