Skip to main content

scigantic-emdb

Search every structure in EMDB, EMBL-EBI's public archive of 3D cryo-EM density maps, and read one in two calls.

pip install "scigantic-emdb[maps]"
from scigantic_emdb import EmdbCatalog, load_map, slices

cat = EmdbCatalog()
cat.search("GPCR", organism="Homo sapiens", max_res=3.0, sort="resolution")

vol, meta = load_map("EMD-22962")   # decompresses, caches, opens
slices(vol)                          # central XY / XZ / YZ sections

Why this exists

Reading an EMDB map is easy. A map is 20 to 60 MB and mrcfile opens it. What EMDB lacks is a way to answer "which of these 60,895 structures do I want" without already knowing the accession. This is that index.

Searching

cat.search("spliceosome", max_res=3.5)
cat.search("GPCR", max_chain_kda=100, has_half_maps=True)
cat.search("protease", ligand="ATP", has_model=True)
cat.search("capsid", microscope="KRIOS", min_year=2023)
cat.search(min_proteins=2, has_ligand=True)          # a complex with something bound
cat.search("ribosome", has_raw_data=True, max_raw_gb=2000)

Filters: max_res, min_res, organism, method, microscope, ligand, has_pdb, has_uniprot, has_model, has_half_maps, has_mask, has_ligand, min_proteins, max_proteins, max_chain_kda, complex_kda_max, has_raw_data, max_raw_gb, year, min_year, max_year. Sort by relevance, resolution, year, raw_size or id.

max_chain_kda is the largest single protein chain, which is what you want for "a receptor under 100 kDa". The assembled complex carries the G protein and often a nanobody, so it is almost always heavier than the molecule of interest. Measured across the catalog, entries with a chain at or under 100 kDa have a median complex weight of 240 kDa.

Coverage

cat.coverage()
# {'catalog_entries': 60895, 'ftp_released_entries': 60895, 'coverage_pct': 100.0, ...}

Every released EMDB entry is indexed. That took two sources: EBI Search covers only 46,900 of them (76.9%), so the remaining 14,042 come from the per-entry REST API.

Molecular fields depend on what each group chose to deposit, so they are not universal:

field fill
microscope, box, map_mb 100%
image 99.0%
contour_level 95.2%
has_half_maps 67.2%
complex_kda 59.1%
max_chain_kda 58.5%
has_mask 26.4%
has_raw_data 8.2%

A record missing the field you filter on is excluded, never silently kept. So "12 structures match" means twelve among those that deposited a weight, not twelve in EMDB. Read the real numbers from cat.coverage()["enrichment"]["field_fill_pct"].

Images

EBI publishes a rendered isosurface for about 97% of entries, so a gallery reads no maps and copies no pixels. The catalog stores the filename and each card points at EBI's public URL.

cat.gallery(cat.search("spliceosome", max_res=4.0).head(8))

Raw data, and which structures you can reprocess

EMDB holds the reconstruction. EMPIAR holds the raw movies it was computed from, when the authors deposited them.

cat.search("GPCR", max_res=3.0, has_raw_data=True, sort="raw_size")
       id  resolution_a empiar_ids  raw_size_gb
EMD-24900          2.60    [10852]       1331.2
EMD-24898          2.90    [10854]       1638.4

Two things to know before planning a reprocessing run. Only 8.2% of EMDB (5,007 of 60,895) has public raw data, because most groups never deposit it. And those datasets are large: of the sub-3 Å GPCR structures with public movies, none is under 1 TB and the cheapest is 1,331 GB. max_raw_gb is the difference between reprocessable and reprocessable by you.

Only EMPIAR records this link. EMDB's own API exposes citations and PDB references but never mentions EMPIAR, so the cross-reference is built from EMPIAR's side.

cat.with_empiar(hits) returns the full pairing when you want the EMPIAR entries themselves.

Half-maps

67% of entries deposit half-maps: two reconstructions built from independent halves of the data, kept separate so the agreement between them measures how much detail is real. With them you can recompute an FSC rather than trusting the reported resolution. Filter with has_half_maps=True.

How search works

Free text runs over title, sample name, organism, method and accessions, with a small cryo-EM synonym vocabulary. What you literally typed always ranks first.

Expansion is a fallback, not a rewrite. Spelling variants (cryoet to cryo-et, ribosome to ribosomal) always apply. Family expansion (GPCR to its 30+ member receptors) engages only when the literal query is thin, because at this scale an unconditional rewrite turned search("rhodopsin") into 1,189 hits of which 40 mentioned rhodopsin. cat.last_query_expanded tells you which happened.

Ligand abbreviations resolve to the deposited chemical names, so ligand="ATP" matches the 965 entries filed as ADENOSINE-5'-TRIPHOSPHATE rather than the 4 that spell it "ATP".

has_ligand=True counts only notable molecules: nucleotides, cofactors, inhibitors, drugs. It excludes ions, water, glycosylation sugars and the lipids and detergents used in sample prep, which are 34,700 of the 56,321 deposited ligand mentions. Counting those would report 37.2% of EMDB as ligand-bound; the honest figure is 21.9%. Use ligand="..." to match the full list by name, including ions.

Relationship to scigantic-empiar

The query layer is imported from scigantic-empiar, never copied, so both archives share one implementation and its fixes cannot diverge. See SYNC.md. It is also why with_empiar() works with no extra setup.

Notes

  • The catalog is a prebuilt index fetched over HTTPS, about 13 MB gzipped, loading in roughly two seconds. Nothing else is downloaded until you read a map.
  • load_map() works with or without the archive mounted. Off-mount it fetches from EBI.
  • entry_files() reports only what an entry actually deposited. Half-maps, masks and FSC curves are per-deposition, so check rather than assume.
  • Accessions are zero-padded to four digits. acc(339) gives EMD-0339, which is the entry that exists.

MIT licensed. EMDB data is CC0. Please cite EMDB.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

scigantic_emdb-0.4.2.tar.gz (30.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

scigantic_emdb-0.4.2-py3-none-any.whl (22.0 kB view details)

Uploaded Python 3

File details

Details for the file scigantic_emdb-0.4.2.tar.gz.

File metadata

  • Download URL: scigantic_emdb-0.4.2.tar.gz
  • Upload date:
  • Size: 30.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for scigantic_emdb-0.4.2.tar.gz
Algorithm Hash digest
SHA256 fd758fb8eb3d497169c31fc8d5b1e17a60806b9624981ba207ff7b2b4987206a
MD5 333117fc24d6a2777155e1088c04c833
BLAKE2b-256 73c76fe6aa599e1e70a1eccc2ce9416b508577fd87cae724c4920e4ef7fa7530

See more details on using hashes here.

File details

Details for the file scigantic_emdb-0.4.2-py3-none-any.whl.

File metadata

  • Download URL: scigantic_emdb-0.4.2-py3-none-any.whl
  • Upload date:
  • Size: 22.0 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for scigantic_emdb-0.4.2-py3-none-any.whl
Algorithm Hash digest
SHA256 6f76e6eb0b8148a61dd92ca2d1b2ec86fba0830a958ff94c1642db75f0b12e00
MD5 e09bb9d0fb104695c6c6a5808643a74d
BLAKE2b-256 c8864749ddc4073b63efab0a0ddd74aef6628c26402e5c58de88700ba9e0c64d

See more details on using hashes here.

Release history Release notifications | RSS feed

0.4.3

2 files

This release

0.4.2 This release

2 files

0.4.1

2 files

0.4.0

2 files

0.3.0

2 files

0.2.0

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page