scigantic-cryoet
Search the CZ CryoET Data Portal from Python.
import scigantic_cryoet as cryoet
cat = cryoet.CryoetCatalog()
cat.search("Legionella", has_annotations=True)
Installation
$ pip install scigantic-cryoet
Why this exists
s3://cryoet-data-portal-public is public and anonymous, but its top level is nothing but numeric dataset ids. There is no search: finding "a tomogram of a ribosome in a mammalian neuron" otherwise means opening dataset_metadata.json by hand, once per dataset, across all ~370 of them. This package is a catalog and search layer on top of that bucket, the same move scigantic-empiar and scigantic-emdb made for their own archives.
It does not read tomogram pixel data or reimplement OME-Zarr access. The portal already ships a thumbnail.gif/snapshot.gif per dataset (unlike EMPIAR, which ships no previews at all), and the official cryoet-data-portal client already reads the zarr volumes well. This package's job ends at "which dataset, and where."
The catalog
CryoetCatalog loads a pre-built index so search/filter across every dataset is instant, no live reads. Measured against the full portal, 2026-09-02 (370/370 datasets, 0 failures, 129s to build):
| field | fill rate | scope |
|---|---|---|
| title, description, organism, sample_type, disease, assay | 100% | every dataset |
| n_runs, authors, release/deposition dates | 100% | every dataset |
| organism (named), tissue, cell_type | 97% / 96% / 94% | every dataset |
| voxel_spacings, reconstruction_method, annotation_objects | 98% / 98% / 92% | one run per dataset |
| emdb_ids / empiar_ids present | 12% / 6.5% | cross-reference, when deposited |
Read catalog_meta before trusting a fill rate. Fields drawn from dataset_metadata.json (title, organism, sample_type, disease, assay, cross-references, ...) are COMPLETE — every dataset has one. Fields that live one level down at the run/tomogram level (voxel_spacings, reconstruction_method, ctf_corrected, annotation_objects, ...) describe one run, named in sampled_run — walking every run of every dataset is thousands of listings for detail that barely varies within a dataset. catalog_meta names exactly which fields are sampled, so this can't be missed the way an earlier internal EMPIAR catalog once advertised "all ~3,000 entries" while actually holding 12.
cat = cryoet.CryoetCatalog()
cat.catalog_meta # {"catalog_entries": 370, "sampled_fields": [...], ...}
cat.search(organism="Homo sapiens", has_annotations=True, sort="runs")
cat.search(has_emdb=True, sort="runs") # datasets cross-referenced to EMDB, most runs first
cat.gallery(cat.search("spike")) # HTML gallery, portal thumbnails
The index isn't published yet — CryoetCatalog() with no argument points at where the monorepo's onboarding batch job will land it. Until then, point it at a local file built with the catalog script:
cat = cryoet.CryoetCatalog(url="/path/to/cryoet-catalog.json")
# or: export SCIGANTIC_CRYOET_CATALOG=/path/to/cryoet-catalog.json
Live reads
CryoetClient hits the portal's bucket directly, for a dataset newer than whatever snapshot is loaded, or to double-check a stale field:
client = cryoet.CryoetClient()
client.dataset_metadata(10000) # full dataset_metadata.json, live
client.runs(10000) # run directory names, COMPLETE (not sampled)
Data license
The portal describes its data as "publicly available" / "open access" (checked 2026-09-02, portal homepage and Terms of Use), but does not state a specific license designation (CC0, CC-BY, or otherwise) on either page. Check a dataset's own citation/attribution requirements — via dataset_metadata.json's authors/publications fields, or the portal's own dataset page — before redistributing.
What this doesn't do (yet)
No synonym/alias expansion in search() (a plain "GPCR" won't match "G protein-coupled receptor") — scigantic-empiar's query expansion is the template once this has real query traffic to tune against. No walk of every run per dataset (see the fill-rate table above); if you need every run's metadata for a specific dataset of interest, use CryoetClient.runs() plus the official cryoet-data-portal client.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file scigantic_cryoet-0.1.0.tar.gz.
File metadata
- Download URL: scigantic_cryoet-0.1.0.tar.gz
- Upload date:
- Size: 13.9 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.11.16
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
39779ce2cafb59195076dcab341018fe5b345cf27a01089b46f4c31c12e4c987
|
|
| MD5 |
35ec5a8c80362274b7d78bb804df6d45
|
|
| BLAKE2b-256 |
bb3194cd84bf2c7c9be09f658ac979301bf87a682a57f2f8fddbaef22a486a1e
|
File details
Details for the file scigantic_cryoet-0.1.0-py3-none-any.whl.
File metadata
- Download URL: scigantic_cryoet-0.1.0-py3-none-any.whl
- Upload date:
- Size: 11.0 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.11.16
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
eb96fbbed1e95e45ec9c75f2c84ebef3c964786cf03bef96ce40f1f5375e895b
|
|
| MD5 |
03e6faa7af5c7ff60e64e16b74b9bc95
|
|
| BLAKE2b-256 |
68dbaf9bf02267c5eb308537984b666fa6cffa1efc381d9a353d56dee61ae6a0
|