dyna-zarr
A lightweight, dask-free Python library for lazy, memory-bounded operations on large Zarr (and TIFF) arrays, with an optional GPU path.
Overview
dyna-zarr is a thin, pull-based array layer over Zarr. Instead of building a task graph, every operation is a lazy transform whose read(key) maps an output slice back to a bounded input read, ending at a direct zarr/TensorStore read. Slicing a result pulls only that region through the whole operation chain, so no intermediates are materialized.
The practical consequence is memory-boundedness. When you stream a result to disk with io.write, the array is processed region by region, so peak RAM is a function of the region and worker budget rather than the array size. This makes it possible to read, transform, and write arrays far larger than memory.
Memory-boundedness
There are two ways to run a lazy result, with different memory behavior:
io.write(result, path)streams the result to disk region by region. Peak RAM is roughlyregion_size_mb * max_workers, independent of the array size. This is the memory-bounded path.result.compute()returns a single in-memory NumPy array. It materializes the whole result by design (mirroringdask.array.compute), so it is not memory-bounded. Use it only for results that fit in RAM.
Every operation is memory-bounded on the io.write path except median, argmin, and argmax, which are flagged in the operations catalog below.
Features
- Pull-based and lazy. Operations defer until
.compute()(materialize) orio.write(stream to disk). - Memory-bounded streaming. Region-wise
io.writewith per-worker memory and worker-count knobs. Even reshape, flatten, and rechunk of incompatibly-chunked data stay bounded, by staging through disk. - NumPy-like. Operator overloads, array methods (
.astype,.clip,.round), and the NumPy ufunc protocol (np.sqrt(a),np.add(a, 2)) all work on aDynamicArray. - Rich op set. About 90 operations: pointwise ufuncs, streaming reductions, neighborhood (halo) filters, structural reshaping, differences, and array creation.
- Multi-format I/O. Read TIFF, Zarr v2, and Zarr v3 (local, S3/GCS, HTTP); write Zarr v2/v3 with optional sharding.
- Optional GPU. Run an op chain on CUDA via CuPy, with a single host-to-device transfer per region.
Installation
pip install dyna-zarr
Optional GPU support (pick the extra matching your CUDA toolkit from nvidia-smi):
pip install "dyna-zarr[gpu-cu12]" # CUDA 12.x ([gpu] is an alias for this)
pip install "dyna-zarr[gpu-cu11]" # CUDA 11.x
pip install "dyna-zarr[gpu-cu13]" # CUDA 13.x (e.g. Blackwell)
Quick start
Read
from dyna_zarr import io
arr = io.read("image.tiff") # TIFF via tifffile's zarr bridge
arr = io.read("array_v2.zarr") # Zarr v2
arr = io.read("array_v3.zarr") # Zarr v3 (also s3://, gs://, http://)
print(arr.shape, arr.dtype, arr.chunks)
data = arr.compute() # materialize the whole array
region = arr[10:20, 50:150, 100:200].compute() # pull just this region
Write (memory-bounded streaming)
from dyna_zarr import io, Codecs
io.write(arr, "out_v3.zarr", zarr_format=3)
io.write(arr, "out.zarr", chunks=(64, 64, 64), zarr_format=3)
io.write(arr, "out.zarr", dtype="float32", zarr_format=3) # cast on write
io.write(arr, "out.zarr", compressor=Codecs(compressor="zstd", clevel=5), zarr_format=3)
# memory and parallelism controls (peak RAM is roughly region_size_mb * max_workers)
io.write(arr, "out.zarr", region_size_mb=64, max_workers=4)
Lazy operation chains
from dyna_zarr import io, operations as ops
arr = io.read("input.zarr")
result = ops.sqrt(ops.clip(ops.abs(arr), 0, 1)) # nothing computed yet
io.write(result, "output.zarr", zarr_format=3) # streamed, region by region
# ...or result.compute() to materialize
NumPy-like interface
A DynamicArray behaves like a NumPy or dask array. Operators, methods, and ufuncs are all lazy:
import numpy as np
masked = (arr > 3) & (arr < 100) # elementwise operators build a lazy mask
scaled = (arr.astype("float32") / 255).clip(0, 1)
out = np.sqrt(np.abs(arr)) # NumPy ufunc protocol dispatches to lazy ops
Neighborhood filters
Neighborhood (halo) filters wrap scipy.ndimage. Each read pulls its own halo, so results are chunk-invariant and exact, and stay memory-bounded when streamed.
import numpy as np
from dyna_zarr import io, operations as ops
img = io.read("volume.zarr") # e.g. (z, y, x)
# LoG filtering
log = ops.gaussian_laplace(img, sigma=2)
io.write(log, "log.zarr", zarr_format=3) # halo handled per region
# median denoise
denoised = ops.median_filter(img, size=3)
io.write(denoised, "denoised.zarr")
# a custom per-plane kernel
kernel = np.ones((1, 3, 3), dtype="float32") / 9 # 3x3 mean within each z-plane
blurred = ops.convolve(img, kernel)
io.write(blurred, "blurred.zarr")
Operations catalog
Every operation is lazy, and memory-bounded on the io.write path except median, argmin, and argmax (see Memory-boundedness). All are available flat on dyna_zarr.operations, and also grouped by category submodule.
- Pointwise / ufuncs.
abs,negative,sign,sqrt,square,exp,log,log2,log10,floor,ceil,reciprocal,round,clip,astype; binaryadd,subtract,multiply,divide,floor_divide,mod,power,maximum,minimum; comparisonsgreater(_equal),less(_equal),equal,not_equal; logicaland,or,xor,not;where,isin,digitize. - Reductions. Streaming and memory-bounded:
min,max,sum,prod,mean,any,all,var,std,histogram(withaxis=andkeepdims=). Not fully bounded (hold the full reduced axis):median,argmin,argmax. - Neighborhood (halo/overlap).
gaussian_filter,uniform_filter,median_filter,minimum_filter,maximum_filter,grey_erosion,grey_dilation,convolve,correlate,laplace,gaussian_laplace,gaussian_gradient_magnitude. - Structural.
concatenate,stack,transpose,swap_axes,reshape,flatten,squeeze,expand_dims,pad,tile,roll,flip,rot90,slice_array. - Differences.
diff,gradient. - Scan (prefix, along one axis).
cumsum,cumprod,cummax,cummin. Streamed with a bounded carry on theio.writepath, so memory-bounded despite the sequential dependency. - Creation.
zeros,ones,full,empty,random(and the*_likevariants).randomis position-deterministic, so the result is independent of chunking. - Primitives.
map_blocks(pointwise),map_overlap(neighborhood with a halo),reduce(streaming). Use these to build your own ops.
Memory-bounded reshape, flatten, and rechunk
C-order reshape and flatten conflict with n-dimensional chunk layout, so a naive implementation blows up. dyna-zarr stages these through disk (a Rechunker-style two-phase, read-once/write-once copy), so peak RAM stays a function of the per-worker budget rather than the array size. When you io.write an outermost reshape or flatten, this path is used automatically:
from dyna_zarr import io, operations as ops
arr = io.read("big_4d.zarr") # e.g. 5 GB, awkward chunks
io.write(ops.flatten(arr), "flat.zarr", region_size_mb=128, max_workers=2)
io.write(ops.reshape(arr, (a, b)), "reshaped.zarr") # (a, b) is any target shape of the same size
GPU (optional)
With a CuPy install, run a chain on the GPU. Setting device='cuda' on a terminal call (compute or io.write) makes device-inheriting ops run on the GPU. A single host-to-device transfer happens at the first CUDA op and the data stays resident up the chain. Results are returned or written from the host.
result = ops.gaussian_filter(arr, sigma=3)
out = result.compute(device="cuda") # whole chain on the GPU
io.write(result, "out.zarr", device="cuda") # per-region GPU compute, streamed write
Relationship to dask
dyna-zarr is not a general replacement for dask.array. It targets one job: memory-bounded read, transform, and write of large Zarr/TIFF arrays.
The core idea is to drop the task graph. Because every operation is a pull-based chain, where each output slice maps back to a bounded input read, there is no graph to build and no scheduler to run it. That keeps the engine small, keeps peak RAM bounded by region_size_mb * max_workers on the io.write path, and avoids scheduling overhead, which makes the read-transform-write pipeline efficient.
The tradeoff is that only operations that fit this slice-pushdown model belong in the chain: pointwise math, neighborhood/halo filters, streaming reductions, and structural reshaping. These are operations that are commonly used in image processing, which is what dyna-zarr is mainly built for. Operations that would need a global, data-dependent graph do not fit directly, and a few that do (such as non-associative reductions) trade extra reads or memory to stay correct.
Two more differences worth knowing:
- Single machine, for now. Parallelism today is threaded I/O within one process, plus the optional GPU path. There is no cluster or distributed execution yet; better and process-based parallelism is a possible future direction.
- Narrower surface. About 90 operations today, extended where the slice-pushdown model permits. Binary ops also need equal-shaped operands (no general broadcasting between differently shaped lazy arrays yet).
Core components
io.read(source)reads TIFF, Zarr v2, or Zarr v3 (local or remote) into aDynamicArray.io.write(array, path, ...)streams aDynamicArrayto Zarr v2/v3 (chunks, sharding, compression, dtype cast,region_size_mb,max_workers,device).operationsis the lazy op set above.DynamicArrayis the pull-based lazy array (slicing,.compute(), operators,.astype/.clip/.round, ufunc protocol).Codecsis the compression configuration for Zarr v2 and v3.
Requirements
- Python 3.11 or newer
- zarr 3.0.0+, numpy 1.20+, scipy 1.6+, tensorstore, tifffile
- Optional: CuPy (via the
gpu-cuXXextras) for the GPU path
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file dyna_zarr-0.0.6.tar.gz.
File metadata
- Download URL: dyna_zarr-0.0.6.tar.gz
- Upload date:
- Size: 101.3 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
682b683c8b795df68b4a3dcb12ea855f3286c2e0f879d63929e45e1e78c4169a
|
|
| MD5 |
04ecb6a1311e5e644d22aedfb1afdb08
|
|
| BLAKE2b-256 |
cc51cbe4bf80068a38484a958ef82c25de0c58400264843e54d5b6fea58700d5
|
Provenance
The following attestation bundles were made for dyna_zarr-0.0.6.tar.gz:
Publisher:
publish.yml on bugraoezdemir/dyna_zarr
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
dyna_zarr-0.0.6.tar.gz -
Subject digest:
682b683c8b795df68b4a3dcb12ea855f3286c2e0f879d63929e45e1e78c4169a - Sigstore transparency entry: 2521032051
- Sigstore integration time:
-
Permalink:
bugraoezdemir/dyna_zarr@7c9fd33362ecacdb6d0e3b8f20c73ee14f23a3ae -
Branch / Tag:
refs/tags/v0.0.6 - Owner: https://github.com/bugraoezdemir
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@7c9fd33362ecacdb6d0e3b8f20c73ee14f23a3ae -
Trigger Event:
release
-
Statement type:
File details
Details for the file dyna_zarr-0.0.6-py3-none-any.whl.
File metadata
- Download URL: dyna_zarr-0.0.6-py3-none-any.whl
- Upload date:
- Size: 77.7 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
58fff3f0bead34d8508a0f277bea19e18023a205536d540993297c63045a681c
|
|
| MD5 |
b17d3621559e000538df31a19b497f6b
|
|
| BLAKE2b-256 |
ce8d6e0ed4cacde513da6fddfd47cca259ccd3a1827fbf5cdb2dc9927e0381f3
|
Provenance
The following attestation bundles were made for dyna_zarr-0.0.6-py3-none-any.whl:
Publisher:
publish.yml on bugraoezdemir/dyna_zarr
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
dyna_zarr-0.0.6-py3-none-any.whl -
Subject digest:
58fff3f0bead34d8508a0f277bea19e18023a205536d540993297c63045a681c - Sigstore transparency entry: 2521032064
- Sigstore integration time:
-
Permalink:
bugraoezdemir/dyna_zarr@7c9fd33362ecacdb6d0e3b8f20c73ee14f23a3ae -
Branch / Tag:
refs/tags/v0.0.6 - Owner: https://github.com/bugraoezdemir
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@7c9fd33362ecacdb6d0e3b8f20c73ee14f23a3ae -
Trigger Event:
release
-
Statement type: