esgpull-plus
API and processing extension to esgf-download: YAML-based download config, fast downloads, CDO regridding, and surface/seafloor subsetting.
Contents
- Installation and set-up
- File structure
- Dependencies
- Keeping up with upstream
- Git configuration
- Searching for data
- CDO regridding pipeline
- Works in progress
- License
Installation and set-up
1. Install the package (in a conda env if you need CDO regridding):
pip install esgpull-plus
# or from a clone:
rye sync
After install, use the esgplus CLI for search and download:
esgplus download # search + download (reads search.yaml by default)
esgplus search # search only
esgplus search-analysis # availability analysis + optional plots
From a dev checkout before rye sync completes, the same commands work via:
python -m esgpull.esgpullplus.cli download
2. Optional – search analysis plots (matplotlib + seaborn):
pip install "esgpull-plus[plotting]"
# or, from a clone with rye:
rye sync --features plotting
3. Optional – CDO regridding / post-processing (cdo-toolkit):
pip install "esgpull-plus[processing]"
# CDO binary (conda recommended):
conda install -c conda-forge cdo
4. Base esgpull:
esgpull self install
See esgf-download installation.
File structure
esgf-download/
├── esgpull/ # Original esgpull
│ └── esgpullplus/ # Download API, file watcher, load_data, …
├── update-from-upstream.sh
Regridding lives in the separate cdo-toolkit package (optional [processing] extra).
Dependencies
- Base: from
pyproject.toml(httpx, click, rich, sqlalchemy, pydantic, etc.). - esgpullplus: pandas, numpy, requests, watchdog, xarray; geospatial via xesmf and
python-cdo(conda). - Optional
[plotting]:matplotlib,seaborn— foresgplus search-analysisandSearchResults.visualize_*(not required for search/download). - Optional
[processing]:cdo-toolkit— for post-download regridding and the async file watcher (also requires the CDO binary).
Keeping up with upstream
Recommended:
./update-from-upstream.sh
Manual:
git fetch upstream && git merge upstream/main
# Then reinstall (conda-aware): conda install -c conda-forge pandas xarray numpy; pip install xesmf cdo-python watchdog orjson
Git configuration
git remote -v
# origin https://github.com/orlando-code/esgpull-plus/ (fetch/push)
# upstream https://github.com/ESGF/esgf-download.git (fetch/push)
If upstream is missing: git remote add upstream https://github.com/ESGF/esgf-download.git
Searching for data
Main search
Populate the search.yaml file (in the repo root) with your ESGF facets and meta options:
search_criteria:
project: CMIP6
table_id: Omon
experiment_id: historical,ssp585
variable: uo,vo
filter:
top_n: 3 # top N datasets to keep
limit: 10 # max results per sub-search
meta_criteria:
test: false
data_dir: /path/to/data
max_workers: 4
Run the search + download pipeline (uses search.yaml automatically):
esgplus download
esgplus download --symmetrical # only sources with both historical + SSP experiments
esgplus download --config path/to/search.yaml
Search only (no download):
esgplus search
esgplus search --config path/to/search.yaml
Legacy equivalents:
python -m esgpull.esgpullplus.api
esgpullplus-download
- Symmetry: in
--symmetricalmode the tool first analyses all experiments and then only downloads datasets from sources that have both historical and SSP-style experiments (e.g.ssp*), so historical/SSP are matched. - Sorting by resolution: search results are converted to a DataFrame and sorted by parsed nominal horizontal resolution, then by
dataset_id, so you always get a consistent “highest resolution first” ordering. - Stable IDs: multi-value facets like
variable: uo,voare normalised (split, trimmed, sorted) so the order you write them insearch.yamldoes not affect the generated search IDs or caching.
Inputs (YAML keys):
| Key | Description |
|---|---|
search_criteria.* |
ESGF facets (project, table_id, experiment_id, variable/variable_id, frequency, etc.). |
search_criteria.filter.top_n |
Number of top grouped datasets to keep per variable and experiment. |
search_criteria.filter.limit |
Maximum number of results per sub-search (useful for debugging). |
meta_criteria.test |
If true (default), downloads go flat to test_downloads/ in the repo. Set false to use data_dir with CMIP6 directory layout. |
meta_criteria.data_dir |
Base directory for downloaded data and cached search results (used when test: false). |
meta_criteria.output_dir |
Optional flat output directory when test: false (overrides data_dir layout). |
meta_criteria.max_workers |
Worker count used for any post-download regridding. |
meta_criteria.regrid_variables |
Optional list of CMIP variable prefixes to regrid after download (e.g. tos). |
meta_criteria.find_alternatives |
If true (default), retry failed downloads from other ESGF data nodes. |
meta_criteria.start_year / end_year |
Optional global year filter applied to all files (filename _YYYYMM-YYYYMM overlap). |
meta_criteria.historic_start_year / historic_end_year |
Year filter for historical experiment files (use with --symmetrical and mixed experiment_id). |
meta_criteria.future_start_year / future_end_year |
Year filter for SSP experiment files (ssp*). When set, each experiment type uses its own range; unset bounds fall back to start_year / end_year. |
search_criteria.member_id |
Ensemble member in YAML; sent to ESGF as variant_label (metagrid/CMIP6 standard). |
meta_criteria.cache_negative_searches |
If true, reuse empty cached subsearches (skip re-querying ESGF). Default false so failed searches are retried. |
Download errors are appended to a single session log under logs/download_errors_<timestamp>.log (path printed at startup and in batch summaries when failures occur).
Search analysis
esgplus search-analysis runs an ESGF search from search.yaml, analyzes source availability (which sources have both historical and SSP experiments, resolutions, ensemble counts), and optionally writes an analysis_df.csv plus PNG plots. It ignores filter.top_n and filter.limit so the analysis uses all matching results.
Run:
esgplus search-analysis
esgplus search-analysis --output-dir notebooks/plots --no-show-plots
Legacy: python run_search_analysis.py [OPTIONS]
| Option | Default | Description |
|---|---|---|
--config / --config-path |
search.yaml |
Path to search config YAML. |
--output-dir |
plots/ (repo) |
Directory for analysis_df.csv and plot PNGs. |
--save-plots |
True | Save plot images (source availability heatmap, ensemble counts, resolution distribution, summary table). |
--show-plots |
True | Display plots interactively; pass --show-plots to disable. |
--require-both |
True | Only include sources that have both historical and SSP experiments. |
Outputs: analysis_df.csv plus, when --save-plots is on, source_availability_heatmap.png, ensemble_counts.png, resolution_distribution.png, source_summary_table.png in the output directory. Plotting requires the optional extra: pip install "esgpull-plus[plotting]".
CDO regridding pipeline
Regridding uses the standalone cdo-toolkit package (PyPI). Install with pip install "esgpull-plus[processing]" (or pip install cdo-toolkit). It supports general NetCDF files, with optional CMIP6 filename helpers. Supports surface (top level) and seafloor extraction: each writes a file next to the original (*_top_level.nc, *_seafloor.nc) and that file is regridded like any other.
You also need the CDO binary (e.g. conda install -c conda-forge cdo).
Command line
# Directory: surface only, tos variable only
cdo-toolkit /path/to/dir -o /path/to/out -r 1.0 1.0 --extract-surface --variable tos
# Directory: seafloor only
cdo-toolkit /path/to/dir -o /path/to/out --extract-seafloor --max-workers 2
# Both surface and seafloor per file
cdo-toolkit /path/to/dir --extreme-levels
# Single file
cdo-toolkit /path/to/file.nc -o /path/to/out.nc --extract-seafloor
Options:
| Option | Default | Description |
|---|---|---|
input (positional) |
required | Input file or directory. |
-o, --output |
same as input dir | Output file or directory; if omitted, writes next to input. |
-r, --resolution lon lat |
1.0 1.0 |
Target output resolution (lon_res, lat_res). |
-p, --pattern |
"*.nc" |
File pattern when input is a directory. |
--variable, -V |
all | CMIP variable prefix(es) to regrid (e.g. tos or tos uo). |
--include-subdirectories |
True |
Include subdirectories when walking a directory. |
--extract-surface |
False |
Extract and regrid only the top level (surface). |
--extract-seafloor |
False |
Extract and regrid only seafloor values. |
--extreme-levels |
False |
Regrid both surface and seafloor for each file. |
--no-regrid-cache |
False |
Disable reuse of CDO weight files. |
--no-seafloor-cache |
False |
Disable reuse of seafloor depth index cache. |
-w, --max-workers |
4 |
Maximum parallel workers. |
--chunk-size-gb |
2.0 |
Maximum time-chunk size in GB. |
--max-memory-gb |
8.0 |
Soft cap for memory-aware chunking. |
--no-parallel |
False |
Process files sequentially. |
--no-chunking |
False |
Disable time chunking (process files in one go). |
-v, --verbose |
True |
Verbose progress UI. |
--verbose-max |
False |
Extra diagnostics (grid type, size, large file messages). |
--quiet |
False |
Disable verbose output. |
--use-ui |
True |
Use the rich progress UI. |
--unlink-unprocessed |
False |
Remove any files that could not be processed. |
--overwrite |
False |
Overwrite existing output files. |
N.B. if --output is not specified, new files will be written to the same directory as the inputs.
File watcher regridding
Continuously watch a directory for new NetCDF files and regrid them as they arrive, using the same CDO pipeline. This is helpful when downloading files and wanting them to be processed directly:
python -m esgpull.esgpullplus.file_watcher /path/to/watch \
-r 1.0 1.0 \
--extract-surface \
--use-regrid-cache \
--process-existing # also process files that are already present
Options:
| Option | Default | Description |
|---|---|---|
watch_dir (positional) |
required | Directory to watch for new NetCDF files. |
-r, --target-resolution lon lat |
1.0 1.0 |
Target output resolution (lon_res, lat_res). |
--target-grid |
"lonlat" |
CDO target grid type. |
--weight-cache-dir |
None |
Directory to store/reuse CDO weight files. |
--max-workers |
4 |
Maximum parallel workers. |
--batch-size |
10 |
Maximum files to accumulate before triggering a batch regrid. |
--batch-timeout |
30.0 |
Maximum seconds to wait before processing a partial batch. |
--extract-surface |
False |
Extract and regrid only the top level (surface). |
--extract-seafloor |
False |
Extract and regrid only seafloor values. |
--use-regrid-cache |
False |
Enable reuse of CDO weight files. |
--use-seafloor-cache |
False |
Enable reuse of seafloor depth index cache. |
--file-settle-seconds |
10.0 |
Wait time to ensure files are no longer being written before processing. |
--validate-can-open |
True |
Validate that files can be opened before scheduling regridding. |
--overwrite |
False |
Overwrite existing regridded outputs. |
--delete-original |
False |
Delete original files after successful regridding. |
--process-existing |
True |
Process files already present in watch_dir on startup. |
Python API
from pathlib import Path
from cdo_toolkit import regrid_directory, regrid_single_file, CDORegridPipeline
# Directory
results = regrid_directory(
Path("data/input"),
output_dir=Path("data/output"),
target_resolution=(1.0, 1.0),
extract_surface=True,
extract_seafloor=False,
max_workers=4,
)
# results["successful"], results["failed"], results["skipped"]
# Single file
ok = regrid_single_file(
Path("data/file.nc"),
output_dir=Path("data/output"),
target_resolution=(1.0, 1.0),
extract_seafloor=True,
)
Features
- Surface/seafloor: Writes
*_top_level.ncor*_seafloor.ncbeside the original, then regrids that file (same CDO path). - Weight reuse: Weights cached per directory (e.g.
cdo_weights/); shared when grids match. - Chunking: Large files split by time; optional
--chunk-size-gb,--max-memory-gb. - Parallel: Per-file locking;
--max-workers;--no-parallelto disable. - Grids: Structured, curvilinear, unstructured (e.g.
ncells); multi-level and time series.
Works in Progress
- There's a fair bit of functionality here! Time to get a proper documentation site in order...
- Merge as much of this functionality as is welcome/useful into the original
esgpullrepository
I am more than happy to take suggestions/contributions from anyone. Just get in touch via email: rt582@cam.ac.uk
License
Same license terms as the esgpull project.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file esgpull_plus-1.1.0.tar.gz.
File metadata
- Download URL: esgpull_plus-1.1.0.tar.gz
- Upload date:
- Size: 343.9 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/6.2.0 CPython/3.14.0
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
8b2962887d522f65440b38ec938d533d80ca5d937482f550bc45358594d159c3
|
|
| MD5 |
dc1149e845533fdb9a4f9b9306993786
|
|
| BLAKE2b-256 |
8ba8452a25715d8f50b9632ede54838e5d8c3d98067d379518adb5d3ae2445fb
|
File details
Details for the file esgpull_plus-1.1.0-py3-none-any.whl.
File metadata
- Download URL: esgpull_plus-1.1.0-py3-none-any.whl
- Upload date:
- Size: 179.7 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/6.2.0 CPython/3.14.0
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
993a383e62f63dbe3a46e40da7cbab79a754e9dc36a3bb708e472878a06a5fc7
|
|
| MD5 |
ce199b066efb6e65047b2e25061ad589
|
|
| BLAKE2b-256 |
1d66dce25a8b0e1d37663c52cef12563ed4d0c5e3e959994af5640ad3af51fe6
|