Copy a zarr dataset while changing chunking - variable by variable
Project description
zarr-copy
Copy a zarr dataset while optionally rechunking variables — variable by variable.
Installation
pip install -e .
Or with uv:
uv sync
Usage
zarr-copy [-h] [-v VAR [VAR ...]] [--rechunk DIM=SIZE [DIM=SIZE ...]]
[--double-remap DIM] [--tmp-path DIR] [--transform FILE]
PATH [PATH ...] DST
Positional arguments
| Argument | Description |
|---|---|
PATH [PATH …] |
One or more source paths or S3 URIs. Multiple sources require --transform. |
DST |
Path or S3 URI of the destination zarr group (always the last positional argument). |
Optional arguments
| Flag | Description |
|---|---|
-v VAR … |
Variable names to copy. Copies all variables when omitted. |
--rechunk DIM=SIZE … |
Rechunk one or more dimensions, e.g. --rechunk time=1 lat=256. |
--double-remap DIM |
Use the two-pass copy strategy for the given dimension to reduce peak memory usage. |
--tmp-path DIR |
Directory for temporary files used by --double-remap. Defaults to the system temp dir. |
--transform FILE |
Python config file defining output variable names and how to compute them (see below). |
Examples
Copy all variables without rechunking:
zarr-copy /data/input.zarr /data/output.zarr
Copy selected variables and rechunk the time dimension:
zarr-copy /data/input.zarr /data/output.zarr -v temperature salinity --rechunk time=1
Copy from S3, rechunk multiple dimensions, and use the double-remap strategy:
zarr-copy s3://my-bucket/input.zarr /data/output.zarr \
--rechunk time=1 lat=256 lon=256 \
--double-remap time \
--tmp-path /scratch
Write to S3 from a local source:
zarr-copy /data/input.zarr s3://my-bucket/output.zarr \
--rechunk lat=256 lon=256
Increase spatial chunk sizes for better read performance along lat/lon:
zarr-copy /data/input.zarr /data/output.zarr --rechunk lat=256 lon=256
Optimise for time-series access by making the time dimension unchunked (one chunk per step):
zarr-copy /data/input.zarr /data/output.zarr --rechunk time=1
Rechunk a large dataset where the time axis is being heavily reordered, using
/scratch for intermediate storage to cap memory use:
zarr-copy /data/input.zarr /data/output.zarr \
-v u v w temperature \
--rechunk time=1 level=1 lat=180 lon=360 \
--double-remap time \
--tmp-path /scratch
Notes
- Rechunking requires
_ARRAY_DIMENSIONSmetadata on each array (standard for xarray/CF-convention zarr stores). - S3 access is handled transparently via s3fs.
- The
--double-remapstrategy writes data to a temporary intermediate zarr (lz4-compressed) to avoid holding large in-memory buffers during rechunking.
Transform: rename and compute output variables
The --transform flag accepts a Python file that defines a module-level
transform dict. Each key is the output variable name; each value is
either:
- a string — rename from the named source variable (simple copy), or
- a callable — compute the output from one or more source variables whose names are taken from the callable's parameter names.
When --transform is provided the -v flag is ignored; all output variables
are determined by the spec.
Transform config file example
# transform.py
transform = {
"time": "time", # rename only
"lon": "lon",
"lat": "lat",
"rlut": lambda ttr: ttr / -3600, # single source
"rsut": lambda tisr, tsr: (tisr - tsr) / 3600, # two sources
"pr": lambda tp: tp * (1000 / 3600),
"psl": "msl", # rename
"ts": "skt",
"uas": "10u",
}
Usage
zarr-copy /data/era5.zarr /data/cmip.zarr --transform transform.py
Combine with --rechunk to rename, compute, and rechunk in one pass:
zarr-copy /data/era5.zarr /data/cmip.zarr \
--transform transform.py \
--rechunk time=1 lat=256 lon=256
Merge variables from two source datasets into one output using a transform. The last positional argument is always the destination; sources are searched in order and the first match wins:
zarr-copy /data/era5_atm.zarr /data/era5_sfc.zarr /data/cmip.zarr \
--transform transform.py \
--rechunk time=1 lat=256 lon=256
SLURM: parallel rechunking on HPC
For large datasets, zarr-copy-slurm submits one SLURM job per variable so
all variables are rechunked concurrently. It accepts the same rechunking
flags as zarr-copy and forwards them to each job.
zarr-copy-slurm [-h] [-v VAR [VAR ...]] [--rechunk DIM=SIZE [DIM=SIZE ...]]
[--double-remap DIM] [--tmp-path DIR] [--transform FILE]
[--cpus N] [--mem MEM] [--walltime HH:MM:SS] [--log-dir DIR]
[--dry-run]
PATH [PATH ...] DST
Examples
Preview the generated job scripts without submitting (useful for checking before committing to a long run):
zarr-copy-slurm /data/input.zarr /data/output.zarr \
--rechunk time=1 lat=256 lon=256 \
--dry-run
Submit one job per variable, rechunking the time dimension into single steps:
zarr-copy-slurm /data/input.zarr /data/output.zarr --rechunk time=1
Submit jobs for selected variables only, with custom resource requests:
zarr-copy-slurm /data/input.zarr /data/output.zarr \
-v temperature salinity u v \
--rechunk time=1 lat=256 lon=256 \
--cpus 8 --mem 64G --walltime 08:00:00
Large dataset with the double-remap strategy, temporaries on scratch:
zarr-copy-slurm /data/input.zarr /data/output.zarr \
--rechunk time=1 level=1 lat=180 lon=360 \
--double-remap time \
--tmp-path /scratch/$USER \
--log-dir /data/logs
Submit one job per output variable defined by a transform spec:
zarr-copy-slurm /data/era5.zarr /data/cmip.zarr \
--transform transform.py \
--rechunk time=1 lat=256 lon=256
Without installing the package, run directly with uvx:
uvx --from zarr-copy zarr-copy-slurm /data/input.zarr /data/output.zarr \
--rechunk time=1 --dry-run
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file zarr_copy-0.4.0.tar.gz.
File metadata
- Download URL: zarr_copy-0.4.0.tar.gz
- Upload date:
- Size: 188.3 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: uv/0.11.1 {"installer":{"name":"uv","version":"0.11.1","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Debian GNU/Linux","version":"13","id":"trixie","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
326af16194a2f123cb5e741290a5042964a5c9e98982a95f72fe8531fd42c914
|
|
| MD5 |
4a8f0b01303891d708f265d1f46dc0da
|
|
| BLAKE2b-256 |
029155e92bdef0a18f4d0d8478035a1409c0694be6ef438f8df6fcdc4b0351d7
|
File details
Details for the file zarr_copy-0.4.0-py3-none-any.whl.
File metadata
- Download URL: zarr_copy-0.4.0-py3-none-any.whl
- Upload date:
- Size: 17.3 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: uv/0.11.1 {"installer":{"name":"uv","version":"0.11.1","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Debian GNU/Linux","version":"13","id":"trixie","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
ba3bd76e6c5b0e38cea1fafff4a53ae2dc50fbb53fd2ce5b5127bd8801c047d5
|
|
| MD5 |
38f051a456f784d26db7db0d96fbdfe3
|
|
| BLAKE2b-256 |
22a006620cedd5a077327ae1fbd1f99077929ec4916d169047f740073ac2893d
|