Skip to main content

EMPIARreader

Python package to access any EMPIAR dataset using its entry number. EMPIARReader provides utilities to lazily load into a machine-learning-friendly dataset format or to locally download the files. The lazy-loading utility allows use of EMPIAR data without the local storage overhead of downloading data permanently. The local download functionality is available via a simple command line interface which allows the user to download EMPIAR data without requiring a user account or proprietary software. Command line utilities are also provided for searching for files within an EMPIAR entry.

Background

EMPIAR is the biggest online archive for cryo-electron microscopy associated raw data. Usually, with each experimental paper there is an associated EMPIAR dataset uploaded. While there is some structure on the database, it is cumbersome for someone without experience in the field to find and access the data. Particularly, it is often necessary the installation of different software. The idea behind EMPIARReader is to provide a package that is easily installable using Python libraries, in order to quickly access the data.

Installation

For Users

EMPIARReader can be installed as a pypi package using Python >=3.8 via:

pip install empiarreader

Otherwise, installation can be done with:

pip install git+https://github.com/alan-turing-institute/empiarreader/

For Developers

For easier installation and dependency handling, EMPIAR reader is also packaged with Poetry

git clone https://github.com/alan-turing-institute/empiarreader/
cd empiarreader
poetry install

Usage

EMPIARReader has an application programming interface (API) and a command line interface (CLI). The API can be used to lazily load EMPIAR datasets into a machine learning compatible format. The command line interface can be used to search the EMPIAR archive and to download files from the archive.

EMPIARReader API

Data from .star format metadata files and .mrc format image files are currently supported for lazy loading into a machine learning compatible format via EMPIARReader.

To retrieve a dataset from an Empiar entry, use the following code:

from empiarreader import EmpiarSource

dataset = EmpiarSource(
            number,
            directory=directory,
            filename_regexp=pattern,
        )

where number is the entry number, directory is the folder path and filename_regexp the file pattern with which to search. For example, if the user wants only the mrc files from the entry number 10943 from a specific folder, the code would be:

ds = EmpiarSource(
            10943,
            directory="data/MotionCorr/job003/Tiff/EER/Images-Disc1/GridSquare_11149304/Data",
            filename_regexp=".*EER\\.mrc",
        )

An example of usage of this package can be found in the notebook available in examples\run_empiarreader.ipynb.

EMPIARReader CLI

Search EMPIAR Entry For Files

To search a particular entry in the EMPIAR archive for files, the empiarreader search utility can be used:

empiarreader search --entry 10934  --select "*"

where --entry is the EMPIAR entry number, --select is the path to use to search for files in this EMPIAR entry (which supports bash-style wildcards). Please enclose this string in quotation marks ("").

Once you know the directory you want to search, you can provide the --dir argument, for example:

empiarreader search --entry 10934  --dir "data" --select "*"

To save the file paths output by the search in a text file the --save_search argument can be supplied:

empiarreader search --entry 10934  --dir "data/CL44-1_20201106_111915/Images-Disc1/GridSquare_6089277/Data" --select "*fractions.tiff.bz2" --save_search saved_search.txt

It is possible to use regex instead of bash-style wildcards to specify files using the --regex argument. To increase the interpretability of the terminal output you can use the --verbose argument. This numbers the matching files and separates files from subdirectories.

Download EMPIAR Files

To download files, first save a list of files to download with the empiarreader search utility. For example,

empiarreader search --entry 10934  --dir "data/CL44-1_20201106_111915/Images-Disc1/GridSquare_6089277/Data" --select "*gain.tiff.bz2" --save_search saved_search.txt

This will contain file paths from a given directory of the EMPIAR entry. You can then download these entries (currently via HTTPS) to a local directory with:

empiarreader download --download saved_search.txt --save_dir new_dir --verbose

Component Description

  • EmpiarCatalog (an Intake catalog, representing entries in the EMPIAR catalog)
  • EmpiarSource (Intake driver for loading from EMPIAR)
  • MrcSource (Intake driver for loading from a file in mrc format)
  • StarSource (Intake driver for starfiles)

Documentation

You can find more documentation including a description of the python api here.

Issues and Feature Requests

If you run into an issue, or if you find a workaround for an existing issue, we would very much appreciate it if you could post your question or code as a GitHub issue.

Contributions

If you would like to help contribute to EMPIARReader, please read our contribution guide and code of conduct.

Metadata

Release files for empiarreader 0.0.17

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for empiarreader 0.0.17
File Size Uploaded
empiarreader-0.0.17.tar.gz 15.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for empiarreader 0.0.17
File Interpreter ABI Platform
empiarreader-0.0.17-py3-none-any.whl Python 3 none any Details

Total release size: 29.5 kB

Release files / empiarreader-0.0.17.tar.gz

Download URL empiarreader-0.0.17.tar.gz
Size 15.8 kB
Tags Source
SHA-256 checksum
How to use checksums
1aee101efd47eb72dfa8cab1b1d856e7c2114312d7f3246a7262a660dbd41e4a
BLAKE2b-256 checksum
How to use checksums
8c03bee71bd167a6008966b42a44c9fad89e815aef1a6f89ff04dfd89302bf3c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.12.9

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jun 18, 2025.

Transparency log

Release files / empiarreader-0.0.17-py3-none-any.whl

Download URL empiarreader-0.0.17-py3-none-any.whl
Size 13.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
3c67ba5c9d76bc372a94bec8c15f04e1a0349d7883781a59862a2bd0f815f2f1
BLAKE2b-256 checksum
How to use checksums
9f9620ca3e553c7ea0fcd4f43b29265879b5ac727aff67fb287f53d96b526dea
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.12.9

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jun 18, 2025.

Transparency log

Release history Release notifications | RSS feed

This release

0.0.17 This release

2 release files

0.0.16

2 release files

0.0.15

2 release files

0.0.14

2 release files

0.0.12

2 release files

0.0.4

2 release files

0.0.1

2 release files

0.0.0

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page