Skip to main content

scigantic-cryoet

CI PyPI PyPI - Python Version License

Search the CZ CryoET Data Portal from Python.

import scigantic_cryoet as cryoet

cat = cryoet.CryoetCatalog()
cat.search("Legionella", has_annotations=True)

Installation

$ pip install scigantic-cryoet

Why this exists

s3://cryoet-data-portal-public is public and anonymous, but its top level is nothing but numeric dataset ids. There is no search: finding "a tomogram of a ribosome in a mammalian neuron" otherwise means opening dataset_metadata.json by hand, once per dataset, across all ~370 of them. This package is a catalog and search layer on top of that bucket, the same move scigantic-empiar and scigantic-emdb made for their own archives.

It does not read tomogram pixel data or reimplement OME-Zarr access. The portal already ships a thumbnail.gif/snapshot.gif per dataset (unlike EMPIAR, which ships no previews at all), and the official cryoet-data-portal client already reads the zarr volumes well. This package's job ends at "which dataset, and where."

The catalog

CryoetCatalog loads a pre-built index so search/filter across every dataset is instant, no live reads. Measured against the full portal, 2026-09-02 (370/370 datasets, 0 failures, 129s to build):

field fill rate scope
title, description, organism, sample_type, disease, assay 100% every dataset
n_runs, authors, release/deposition dates 100% every dataset
organism (named), tissue, cell_type 97% / 96% / 94% every dataset
voxel_spacings, reconstruction_method, annotation_objects 98% / 98% / 92% one run per dataset
emdb_ids / empiar_ids present 12% / 6.5% cross-reference, when deposited

Read catalog_meta before trusting a fill rate. Fields drawn from dataset_metadata.json (title, organism, sample_type, disease, assay, cross-references, ...) are COMPLETE: every dataset has one. Fields that live one level down at the run/tomogram level (voxel_spacings, reconstruction_method, ctf_corrected, annotation_objects, ...) describe one run, named in sampled_run: walking every run of every dataset is thousands of listings for detail that barely varies within a dataset. catalog_meta names exactly which fields are sampled, so this can't be missed the way an earlier internal EMPIAR catalog once advertised "all ~3,000 entries" while actually holding 12.

cat = cryoet.CryoetCatalog()
cat.catalog_meta        # {"catalog_entries": 370, "sampled_fields": [...], ...}

cat.search(organism="Homo sapiens", has_annotations=True, sort="runs")
cat.search(has_emdb=True, sort="runs")   # datasets cross-referenced to EMDB, most runs first
cat.gallery(cat.search("spike"))         # HTML gallery, portal thumbnails

The index isn't published yet. CryoetCatalog() with no argument points at where the monorepo's onboarding batch job will land it. Until then, point it at a local file built with the catalog script:

cat = cryoet.CryoetCatalog(url="/path/to/cryoet-catalog.json")
# or: export SCIGANTIC_CRYOET_CATALOG=/path/to/cryoet-catalog.json

Cross-referencing EMDB and EMPIAR

Many CryoET Data Portal datasets cite an EMDB structure or an EMPIAR raw-data deposit in their own cross-references. with_emdb()/with_empiar() join on those ids directly, no separate download or lookup table:

$ pip install "scigantic-cryoet[bridge]"   # pulls in scigantic-emdb / scigantic-empiar
cat.with_emdb()      # every (dataset, EMDB structure, resolution) pair the portal cites
cat.with_empiar()    # every (dataset, EMPIAR raw deposit, size, method) pair

Measured against the full portal, 2026-09-02: 160 dataset-to-EMDB pairs across 45 datasets, 27 dataset-to-EMPIAR pairs across 24 datasets. Both raise RuntimeError if the sibling package isn't installed, rather than returning an empty (and misleadingly "no cross-references") result.

Live reads

CryoetClient hits the portal's bucket directly, for a dataset newer than whatever snapshot is loaded, or to double-check a stale field:

client = cryoet.CryoetClient()
client.dataset_metadata(10000)   # full dataset_metadata.json, live
client.runs(10000)               # run directory names, COMPLETE (not sampled)

Data license

The portal describes its data as "publicly available" / "open access" (checked 2026-09-02, portal homepage and Terms of Use), but does not state a specific license designation (CC0, CC-BY, or otherwise) on either page. Check a dataset's own citation/attribution requirements (via dataset_metadata.json's authors/publications fields, or the portal's own dataset page) before redistributing.

What this doesn't do (yet)

No synonym/alias expansion in search() (a plain "GPCR" won't match "G protein-coupled receptor"); scigantic-empiar's query expansion is the template once this has real query traffic to tune against. No walk of every run per dataset (see the fill-rate table above); if you need every run's metadata for a specific dataset of interest, use CryoetClient.runs() plus the official cryoet-data-portal client.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

scigantic_cryoet-0.1.2.tar.gz (16.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

scigantic_cryoet-0.1.2-py3-none-any.whl (12.5 kB view details)

Uploaded Python 3

File details

Details for the file scigantic_cryoet-0.1.2.tar.gz.

File metadata

  • Download URL: scigantic_cryoet-0.1.2.tar.gz
  • Upload date:
  • Size: 16.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.11.16

File hashes

Hashes for scigantic_cryoet-0.1.2.tar.gz
Algorithm Hash digest
SHA256 8a1a5ba02cdf3b9ca572218ae5d43c745b10b7fcba8c87dcf828c128c9fb6aaf
MD5 b65304dbfac742b7683d92d9b159ee35
BLAKE2b-256 3f135db594ff427aefbd0a8483786de389daccc7f13e83fdbf43a3ae54e54df4

See more details on using hashes here.

File details

Details for the file scigantic_cryoet-0.1.2-py3-none-any.whl.

File metadata

File hashes

Hashes for scigantic_cryoet-0.1.2-py3-none-any.whl
Algorithm Hash digest
SHA256 d1702996ef2d444f84bfa74502f94e38991693ad820eb25ca5d3c3d909859825
MD5 b5908c58e7f31ebf5e23eb2eeec2fb42
BLAKE2b-256 d86e4b59b68f3d6ff7dc261ed2cb6695e12b8cb100d1016b1a3efc52ccfc8516

See more details on using hashes here.

Release history Release notifications | RSS feed

0.1.3

2 files

This release

0.1.2 This release

2 files

0.1.1

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page