Skip to main content

sdrbench

The SDRBench scientific datasets as numpy arrays, served from the Hugging Face mirror (byte-for-byte copies of the original archives, pinned and sha256-verified), with automatic fallback to the original Globus archives at Argonne.

pip install sdrbench            # or "sdrbench[sz3]" to also get pysz (SZ3)
import sdrbench

sdrbench.list()                          # ['basilisk-turbulence', 'cesm-atm', 'exaalt', ...]
sdrbench.list(variants=True)             # ['basilisk-turbulence/small', ..., 'nyx/original', 'nyx/log', ...]

nyx = sdrbench.dataset("nyx")            # default variant; sdrbench.dataset("nyx", "log")
nyx.fields                               # ['baryon_density', 'dark_matter_density', 'temperature', ...]
t = nyx["temperature"]                   # numpy memmap, float32, shape (512, 512, 512), C order
sdrbench.load("cesm-atm", "CLDHGH")      # one-liner

# time series: fields are "var/step"
hur = sdrbench.dataset("hurricane-isabel", "P")
hur.steps                                # ['01', ..., '48']
p = hur.series("P")                      # ordered sequence of fields
x = hur["P", "07"]                       # one step
block = p.stack(slice(0, 4))             # (4, 100, 500, 500) copy

# stacked files are split into zero-copy views per variable
T = sdrbench.dataset("s3d")["T/1.1000E-03"]      # (500, 500, 500), component 6 of the stored file

# raw files to a directory (original SDRBench names, downloaded in parallel)
nyx.download("data/")                    # -> data/512x512x512/*.f32
sdrbench.dataset("nyx", root="data/")    # later: use local files first (also $SDRBENCH_DATA)

print(nyx.citation())                    # BibTeX + provider request + exact data version

Compress with SZ3 through pysz:

import numpy as np
from pysz import sz, szConfig, szErrorBoundMode

x = sdrbench.dataset("scale-letkf")["T"]
conf = szConfig()
conf.errorBoundMode = szErrorBoundMode.REL
conf.relErrorBound = 1e-3
compressed, ratio = sz.compress(np.ascontiguousarray(x), conf)
y, _ = sz.decompress(compressed, x.dtype.type, x.shape)
print(ratio, sz.verify(np.asarray(x), y))  # ratio, (max error, PSNR, NRMSE)

Command line:

sdrbench list                                   # every dataset/variant with field count and size
sdrbench info hurricane-isabel/P                # description, fields, dtype, shape, files
sdrbench download nyx temperature -o data/      # plain files under data/<repo path>, -j parallel
sdrbench path cesm-atm CLDHGH                   # local path (downloads if needed), e.g. for `sz3 -i`
sdrbench cite qmcpack

How it works

  • Datasets, variants, fields. A dataset (nyx) has variants with descriptive names (original, log; cesm-atm: 2d, 2d-cleared, 3d; hacc: medium, big, sbig, region1...; hurricane-isabel: snapshot, P, QCLOUD, ...). The first is the default. The SDRBench / Hugging Face folder name also works as a variant name (dataset("nyx", "512x512x512")).
  • Field names are physical variables (CLDHGH, T, temperature), or var/step for time series and slabs (P/07, xx/00042, density/31); ds.variables, ds.steps and ds.series(var) expose the structure. The original SDRBench file name also works (ds["CLDHGH_1_1800_3600.f32"]), lookups are case-insensitive, and errors suggest close matches. Non-array files are in ds.extra_files.
  • dtype and shape come from the catalog, never from names. Shapes are C order (slowest first; the convention of the SZ3 test-suite table): np.fromfile(path, dtype).reshape(field.stored_shape) reads any stored file. field.shape_fastest_first is the order command-line compressors expect.
  • Views, no copies of the data. QMCPACK's default variant preconditioned (288 x 115 x 69 x 69, the layout of the SDRBench examples) is computed on load by transposing the stored native file (variant original); S3D's stacked files are split into per-variable memmap views. field.save(path) writes the array exactly as loaded.
  • Downloads come from Hugging Face, pinned to the commit this release's catalog was built from and checked against the catalog sha256, and are cached (SDRBENCH_CACHE moves the cache). Local copies are used first when root= / $SDRBENCH_DATA is given (an untarred SDRBench directory works). If Hugging Face cannot deliver, the package falls back to the original archive on Globus (whole archive, md5 and sha256 verified, size shown in a warning; resumes dropped connections). source="globus" forces it.
  • Arrays are read-only memmaps; np.array(x) (or load(mmap=False)) gives a writable copy, e.g. for torch.from_numpy.

Datasets

Dataset Hugging Face Variants Source
CESM-ATM sdrbench/cesm-atm 2d, 2d-cleared, 3d climate (SNL)
EXAALT sdrbench/exaalt small, copper-1/2, helium-1/2 molecular dynamics
Hurricane ISABEL sdrbench/hurricane-isabel snapshot, P, U, ..., CLOUD_log10, ... weather (NCAR, IEEE Vis 2004)
EXAFEL sdrbench/exafel small, large, assembled LCLS X-ray images
HACC sdrbench/hacc medium, big, sbig, region1-6 cosmology particles
NYX sdrbench/nyx original, log cosmology
NWChem sdrbench/nwchem default, f32 quantum chemistry
SCALE-LETKF sdrbench/scale-letkf original, log weather (RIKEN)
QMCPACK sdrbench/qmcpack preconditioned, original quantum Monte Carlo
Miranda sdrbench/miranda small, big turbulence (LLNL)
S3D sdrbench/s3d default combustion (SNL)
Basilisk-Turbulence sdrbench/basilisk-turbulence small, large 2D turbulence
XGC sdrbench/xgc raw, adios fusion (PPPL)

sdrbench list and sdrbench info show fields and sizes. NSTX GPI is not mirrored (its owner asks to be contacted before results are published); HACC region 5 is not available on Globus (HTTP 404).

Please cite SDRBench and the data provider named on each dataset card (sdrbench cite <dataset>), and report the sdrbench version you used: a release always reads the same data.

Keeping the mirror in sync

Globus remains the source of truth. .github/workflows/release.yml runs weekly:

  1. sync/sync.py check HEADs every archive linked from the SDRBench page and compares size, ETag and Last-Modified with sync/state.json; archives that appear on, or disappear from, the page open an issue.
  2. Changed archives are re-mirrored in parallel jobs (sync.py mirror). Archives are streamed: each file is unpacked, hashed and committed to sdrbench/<dataset> in batches and then deleted, so a runner only needs room for the largest single file (~17 GB); the job frees disk like the SZ3 CI. Files removed from an archive are removed from the repo, but only if the whole archive arrived and at most 25% of a variant goes (--allow-delete overrides). The same command works on any machine: HF_TOKEN=... python sync/sync.py mirror --only <dataset>.
  3. sync.py catalog regenerates src/sdrbench/catalog.json (pinned to the new Hugging Face revisions), sync.py verify checks dtypes, element counts and dimension order against the data, sync.py cards refreshes the dataset cards and file tables, and a new patch version is tagged and published to PyPI.

Adding a dataset

  1. Add an entry to sync/datasets.json (see its _comment): title, provider, science description, and per variant the Hugging Face folder, Globus archive, optional select/exclude, and rules mapping file globs to dtype, C-order shape and var/step name templates.
  2. python sync/sync.py mirror --only <dataset> --dry --workdir /tmp/w downloads and hashes without uploading (writes /tmp/w/state.dry.json); fix rules until sync.py catalog builds without errors (it refuses files without a rule or with a size that does not match the shape).
  3. Mirror for real, then sync.py catalog, sync.py verify --only <dataset>, sync.py cards --only <dataset>, and python sync/selftest.py --only <dataset> --workdir /big/tmp (every field through the package).

Development

pip install -e ".[test,sz3]"
pytest                       # offline tests
pytest -m online             # end-to-end against Hugging Face and Globus
python sync/sync.py verify   # data-level check of every dtype/shape (network)
python sync/selftest.py --workdir /big/tmp   # every field of every dataset through the package

Metadata

Release files for sdrbench 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for sdrbench 0.1.1
File Size Uploaded
sdrbench-0.1.1.tar.gz 198.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for sdrbench 0.1.1
File Interpreter ABI Platform
sdrbench-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 390.8 kB

Release files / sdrbench-0.1.1.tar.gz

Download URL sdrbench-0.1.1.tar.gz
Size 198.1 kB
Tags Source
SHA-256 checksum
How to use checksums
9ad78bb7f274645bcd5de9285779702aa5cc576327cd91225c911031ae8617c7
BLAKE2b-256 checksum
How to use checksums
3310017165ef493c8656837f78914f91b2c60557e36ef81fdf96a4e6b46555e6
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 9, 2026.

Transparency log

Release files / sdrbench-0.1.1-py3-none-any.whl

Download URL sdrbench-0.1.1-py3-none-any.whl
Size 192.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
c56a21dd0322b6ef40d9847f351bcc5cdedf5f7a063564c442f6855e91d03769
BLAKE2b-256 checksum
How to use checksums
c8cf79089d404edf09037e9328fad90d363d1b3cb792f00c3d10c2807ddf8bcd
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 9, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.1 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page