scigantic-emdb
Search every structure in EMDB, EMBL-EBI's public archive of 3D cryo-EM density maps, and read one in two calls.
pip install "scigantic-emdb[maps]"
from scigantic_emdb import EmdbCatalog, load_map, slices
cat = EmdbCatalog()
cat.search("GPCR", organism="Homo sapiens", max_res=3.0, sort="resolution")
vol, meta = load_map("EMD-22962") # decompresses, caches, opens
slices(vol) # central XY / XZ / YZ sections
Why this exists
Reading an EMDB map is easy. A map is 20 to 60 MB and mrcfile opens it. What EMDB lacks is a way to answer "which of these 60,895 structures do I want" without already knowing the accession. This is that index.
Searching
cat.search("spliceosome", max_res=3.5)
cat.search("GPCR", max_chain_kda=100, has_half_maps=True)
cat.search("protease", ligand="ATP", has_model=True)
cat.search("capsid", microscope="KRIOS", min_year=2023)
cat.search(min_proteins=2, has_ligand=True) # a complex with something bound
cat.search("ribosome", has_raw_data=True, max_raw_gb=2000)
Filters: max_res, min_res, organism, method, microscope, ligand, has_pdb, has_uniprot, has_model, has_half_maps, has_mask, has_ligand, min_proteins, max_proteins, max_chain_kda, complex_kda_max, has_raw_data, max_raw_gb, year, min_year, max_year. Sort by relevance, resolution, year, raw_size or id.
max_chain_kda is the largest single protein chain, which is what you want for "a receptor under 100 kDa". The assembled complex carries the G protein and often a nanobody, so it is almost always heavier than the molecule of interest. Measured across the catalog, entries with a chain at or under 100 kDa have a median complex weight of 240 kDa.
Coverage
cat.coverage()
# {'catalog_entries': 60895, 'ftp_released_entries': 60895, 'coverage_pct': 100.0, ...}
Every released EMDB entry is indexed. That took two sources: EBI Search covers only 46,900 of them (76.9%), so the remaining 14,042 come from the per-entry REST API.
Molecular fields depend on what each group chose to deposit, so they are not universal:
| field | fill |
|---|---|
microscope, box, map_mb |
100% |
image |
99.0% |
contour_level |
95.2% |
has_half_maps |
67.2% |
complex_kda |
59.1% |
max_chain_kda |
58.5% |
has_mask |
26.4% |
has_raw_data |
8.2% |
A record missing the field you filter on is excluded, never silently kept. So "12 structures match" means twelve among those that deposited a weight, not twelve in EMDB. Read the real numbers from cat.coverage()["enrichment"]["field_fill_pct"].
Images
EBI publishes a rendered isosurface for about 97% of entries, so a gallery reads no maps and copies no pixels. The catalog stores the filename and each card points at EBI's public URL.
cat.gallery(cat.search("spliceosome", max_res=4.0).head(8))
Raw data, and which structures you can reprocess
EMDB holds the reconstruction. EMPIAR holds the raw movies it was computed from, when the authors deposited them.
cat.search("GPCR", max_res=3.0, has_raw_data=True, sort="raw_size")
id resolution_a empiar_ids raw_size_gb
EMD-24900 2.60 [10852] 1331.2
EMD-24898 2.90 [10854] 1638.4
Two things to know before planning a reprocessing run. Only 8.2% of EMDB (5,007 of 60,895) has public raw data, because most groups never deposit it. And those datasets are large: of the sub-3 Å GPCR structures with public movies, none is under 1 TB and the cheapest is 1,331 GB. max_raw_gb is the difference between reprocessable and reprocessable by you.
Only EMPIAR records this link. EMDB's own API exposes citations and PDB references but never mentions EMPIAR, so the cross-reference is built from EMPIAR's side.
cat.with_empiar(hits) returns the full pairing when you want the EMPIAR entries themselves.
Half-maps
67% of entries deposit half-maps: two reconstructions built from independent halves of the data, kept separate so the agreement between them measures how much detail is real. With them you can recompute an FSC rather than trusting the reported resolution. Filter with has_half_maps=True.
How search works
Free text runs over title, sample name, organism, method and accessions, with a small cryo-EM synonym vocabulary. What you literally typed always ranks first.
Expansion is a fallback, not a rewrite. Spelling variants (cryoet to cryo-et, ribosome to ribosomal) always apply. Family expansion (GPCR to its 30+ member receptors) engages only when the literal query is thin, because at this scale an unconditional rewrite turned search("rhodopsin") into 1,189 hits of which 40 mentioned rhodopsin. cat.last_query_expanded tells you which happened.
Ligand abbreviations resolve to the deposited chemical names, so ligand="ATP" matches the 965 entries filed as ADENOSINE-5'-TRIPHOSPHATE rather than the 4 that spell it "ATP".
has_ligand=True counts only notable molecules: nucleotides, cofactors, inhibitors, drugs. It excludes ions, water, glycosylation sugars and the lipids and detergents used in sample prep, which are 34,700 of the 56,321 deposited ligand mentions. Counting those would report 37.2% of EMDB as ligand-bound; the honest figure is 21.9%. Use ligand="..." to match the full list by name, including ions.
Relationship to scigantic-empiar
The query layer is imported from scigantic-empiar, never copied, so both archives share one implementation and its fixes cannot diverge. See SYNC.md. It is also why with_empiar() works with no extra setup.
Notes
- The catalog is a prebuilt index fetched over HTTPS, about 13 MB gzipped, loading in roughly two seconds. Nothing else is downloaded until you read a map.
load_map()works with or without the archive mounted. Off-mount it fetches from EBI.entry_files()reports only what an entry actually deposited. Half-maps, masks and FSC curves are per-deposition, so check rather than assume.- Accessions are zero-padded to four digits.
acc(339)givesEMD-0339, which is the entry that exists.
MIT licensed. EMDB data is CC0. Please cite EMDB.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file scigantic_emdb-0.4.2.tar.gz.
File metadata
- Download URL: scigantic_emdb-0.4.2.tar.gz
- Upload date:
- Size: 30.9 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
fd758fb8eb3d497169c31fc8d5b1e17a60806b9624981ba207ff7b2b4987206a
|
|
| MD5 |
333117fc24d6a2777155e1088c04c833
|
|
| BLAKE2b-256 |
73c76fe6aa599e1e70a1eccc2ce9416b508577fd87cae724c4920e4ef7fa7530
|
File details
Details for the file scigantic_emdb-0.4.2-py3-none-any.whl.
File metadata
- Download URL: scigantic_emdb-0.4.2-py3-none-any.whl
- Upload date:
- Size: 22.0 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
6f76e6eb0b8148a61dd92ca2d1b2ec86fba0830a958ff94c1642db75f0b12e00
|
|
| MD5 |
e09bb9d0fb104695c6c6a5808643a74d
|
|
| BLAKE2b-256 |
c8864749ddc4073b63efab0a0ddd74aef6628c26402e5c58de88700ba9e0c64d
|