Skip to main content

jump-image-datasets

Tests PyPI version Publish to PyPI

jump-image-datasets provides packaged JUMP pilot metadata and utilities for downloading image files from metadata tables.

Install

Install from PyPI

pip install jump-image-datasets

Install from PyPI for stable, versioned releases.

Local development with uv

uv venv
uv sync --group test

Editable install

uv pip install -e .

Install from the GitHub repo with pip

pip install "git+https://github.com/WayScience/jump_image_data_downloader.git"

Install from GitHub if you want the latest unreleased changes.

AWS CLI requirement for CPG0016 analysis CSV downloads

Bulk downloading CPG0016 workspace/analysis CSVs with CPG0016AnalysisCSVDownloader requires the AWS CLI to be installed and available as aws on your PATH. The downloader uses the AWS CLI transfer manager for faster recursive CSV downloads from the public S3 bucket.

Usage

from jump_image_datasets.jump_pilot import image_downloader, image_metadata

# Load packaged metadata parquet as a DataFrame.
metadata_df = image_metadata.load_metadata()

# Download a small subset.
summary = image_downloader.download_images_with_metadata(
    df=metadata_df.head(10),
    url_column="Metadata_FileUrl",
    default_output_dir="downloaded_jump_pilot_images",
    parallel=True,
    workers=8,
)
print(summary)
from jump_image_datasets.cpg0016 import CPG0016LoadDataWithIllumDownloader

downloader = CPG0016LoadDataWithIllumDownloader(
    csv_download_dir="downloaded_cpg0016_csvs",
)

metadata_df = downloader.get_dataframe()

filtered_df = metadata_df.iloc[:10].copy()
filtered_df["OutputDir"] = (
    filtered_df["Metadata_Plate"].astype(str).radd("downloaded_cpg0016_images/")
)

downloader.download_illumination_files(
    dataframe=filtered_df,
    output_dir_column="OutputDir",
)

downloader.download_files_from_column(
    dataframe=filtered_df,
    column_name="URL_OrigDNA",
    output_dir_column="OutputDir",
)
from jump_image_datasets.cpg0016 import CPG0016AnalysisCSVDownloader

# Discover all CPG0016 analysis CSVs from S3 and download them into an
# organized local directory tree.
downloader = CPG0016AnalysisCSVDownloader(
    output_dir="downloaded_cpg0016_profiles",
    parallel=True,
    max_concurrent_requests=50,
)

# Omit csv_names to download all standard profile CSVs.
summary = downloader.download_all_csv_profiles()
print(summary)

# Or download only a selected subset.
nuclei_and_image_summary = downloader.download_all_csv_profiles(
    csv_names=["nuclei", "image"],
)
print(nuclei_and_image_summary)

# Later, reuse the existing local CSV tree without checking S3.
local_only_downloader = CPG0016AnalysisCSVDownloader(
    output_dir="downloaded_cpg0016_profiles",
    use_existing_csvs_without_s3_check=True,
)

# Iterate through one analysis folder at a time and choose your own operations,
# such as reading Image.csv and Nuclei.csv and merging them on ImageNumber.
for csv_set in local_only_downloader.iter_analysis_csv_sets():
    print(csv_set.folder_local_path)
    print(csv_set.image_local_path)
    print(csv_set.nuclei_local_path)
    break

For full runnable examples, see docs/download_images_examples.ipynb, docs/download_cpg0016_examples.ipynb, and docs/download_cpg0016_profiles_examples.ipynb.

Packaged metadata provenance

This repository ships a packaged metadata table at:

  • src/jump_image_datasets/jump_pilot/data/2020_11_04_CPJUMP1_all_plates.parquet

Why this file exists

The file is included so users can immediately load a stable JUMP pilot metadata table (via jump_image_datasets.jump_pilot.image_metadata) without requiring a separate data-fetch or preprocessing step.

How it was created

This parquet was generated from the JUMP Cell Painting Gallery using:

Upstream source pattern used by that notebook:

  • s3://cellpainting-gallery/cpg0000-jump-pilot/source_4/workspace/load_data_csv/2020_11_04_CPJUMP1/*/load_data.csv

Transform summary

The generation workflow in 2.download_image_metadata.ipynb:

  • Lists all per-plate load_data.csv files for run 2020_11_04_CPJUMP1 (51 files in the captured run) from public S3 (anon=True).
  • Reads each plate CSV, appends provenance columns:
    • source_plate (plate ID parsed from path)
    • source_s3_path (full S3 CSV path)
  • Concatenates all plate tables into one DataFrame.
  • Reshapes channel URL columns from wide to long using melt:
    • URL columns become Metadata_ChannelURLName
    • URL values become Metadata_FileUrl
  • Adds normalized channel/stain annotations by mapping URL column names:
    • Metadata_ChannelName: ER, AGP, Mito, DNA, RNA, BF, HZ_BF, LZ_BF
    • Metadata_StainName: corresponding stain labels (or NA for brightfield channels)
  • Derives Metadata_Filename from the final path component of Metadata_FileUrl.
  • Writes parquet with index=False as data/2020_11_04_CPJUMP1_all_plates.parquet (captured shape: (1495400, 32)).

Release files for jump-image-datasets 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for jump-image-datasets 0.3.0
File Size Uploaded
jump_image_datasets-0.3.0.tar.gz 14.4 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for jump-image-datasets 0.3.0
File Interpreter ABI Platform
jump_image_datasets-0.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 28.8 MB

Release files / jump_image_datasets-0.3.0.tar.gz

Download URL jump_image_datasets-0.3.0.tar.gz
Size 14.4 MB
Tags Source
SHA-256 checksum
How to use checksums
9933cc17ba3bb14751a87740ceaf71d244f5df032055df31794370380f4e7ccb
BLAKE2b-256 checksum
How to use checksums
1388690adc59750e3aed861641f71efe3a83e5131e852c3129dbc39f37865058
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 10, 2026.

Transparency log

Release files / jump_image_datasets-0.3.0-py3-none-any.whl

Download URL jump_image_datasets-0.3.0-py3-none-any.whl
Size 14.3 MB
Tags Python 3
SHA-256 checksum
How to use checksums
2e65742fe1bc12566aeb8b56baaa263db17c66d93a58a30991594e66488ae13f
BLAKE2b-256 checksum
How to use checksums
a09ac4686297b2cb1328e0797996e744ac1f9271595990950e121e7151e99dc2
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 10, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.3.0 This release

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page