AnniZarr
Convert to, append, sort anndata zarr storages efficiently with parallel and lazy operations for all commands and api usage. Also, enable versioned anndata storage with Icechunk ease of use.
Install
annizarr and the shorter anz are the same command. The icechunk extra adds versioned stores and
remote (s3://, gs://) targets.
pip install annizarr
pip install "annizarr[icechunk]"
To work on the code instead, use the pixi dev environment:
pixi install
Quickstart usage
Convert
The input format is sniffed from the HDF5 contents (h5ad vs 10x); --from h5ad|10x overrides it. Configuration can enable chunking, sharding, format ...
Two or more inputs are concatenated. --ic writes an Icechunk repository instead of plain zarr.
annizarr convert sample.h5ad -o sample.zarr
annizarr convert a.h5ad b.h5ad -o merged.zarr
annizarr convert sample.h5ad -o repo.icechunk --ic -m "initial import"
Zarr operations and edits
rechunk rewrites one matrix with new chunking and streams everything else through unchanged.
sort physically orders rows by obs column(s) into a new store.
append adds cells to an existing zarr/ic store; it prompts before dropping obsm/obsp/layers (--drop-derived consents up front)
add-expr adds a log-normalized layer (layers/gexp) derived from X/. Csc by default for fast column reads
annizarr rechunk merged.zarr -o rechunked.zarr --row-chunk 2000
annizarr sort merged.zarr -o sorted.zarr --by cell_type
annizarr append sorted.zarr more_cells.zarr --drop-derived
annizarr add-expr merged.zarr --layout csr
annizarr add-expr repo.icechunk --layout csr --branch dev -m "add lognorm layer"
Icechunk history, branches, cherry-picks live in the Python API annizarr.Repo below
Configuration and zarr stores
Every flag is documented in annizarr <command> --help; docs/cli.md has the same
reference in one place. Inputs stream band by band by default, so memory stays bounded at any store
size (--eager loads the whole input first, faster for small files), and matrix writes use every
core unless --cpus says otherwise.
Some of the main options; the docs have all of them:
--cpus N— parallel band workers (default: all cores). Every command.--layout csr|csc|dense— on-disk layout of the matrix being written. Defaultcsrfor X onconvert,cscfor theadd-exprlayer. CSR suits row-wise access, CSC and dense column queries.--row-chunk N,--col-chunk N— chunk shape. Exact rows and columns for dense; for sparse, about N cells (csr) or N genes (csc) per chunk, sized from the average nonzeros. Defaults: 2048 for dense, about 9,000,000 nonzeros for sparse.convert,rechunk,add-expr.--auto-shard— shard the 1-D sparse arrays and anndata-written elements with zarr's automatic shard shape (default: off). Every command butappend.--ic— write through an Icechunk repository;--branch Band-m MSGpick the branch and commit message.appendandadd-exprdetect an existing repo on their own.convert,rechunk,sort.--sort-by COL…onconvert/--by COL…onsort— physically order rows by obs columns, primary key first
What gets written. Every store is anndata-readable zarr v3 with encoding-type/encoding-version
attrs matching anndata 0.13's on-disk spec. X and layers may be dense, CSR or CSC on input (h5ad or
10x, mixed across inputs) and are written in whatever --layout asks for; X defaults to CSR.
convert/rechunk/sort write to a sibling temp store, verify it opens, then rename it onto the
target, so a killed run never leaves a partial store. On Icechunk every op is exactly one commit,
and remote (s3://, gs://) outputs require it. More in docs/stores.md.
Python API for Icechunk usage
The full Python API, including the config dataclasses and the ops functions, is in docs/python-api.md.
import annizarr as az
repo = az.Repo("repo.icechunk")
repo.branches()
repo.log()
repo.tree()
root = repo.open_zarr("w")
root.attrs["step"] = "lognorm"
repo.commit("normalize")
repo.checkout("experiment", create=True) # switches this object only; nothing is persisted
old = repo.open_zarr("r", snapshot_id=repo.log()[1].id)
repo.copy("/work/store") # clone with every branch and snapshot id intact, good for readonly cases
OpResult(path, n_obs, n_vars, snapshot_id) — snapshot_id is None for plain zarr,
the committed Icechunk snapshot id otherwise.
Development
pixi install
pixi run -e default pytest # excludes -m slow by default
pixi run -e default pytest -m slow # large synthetic-store memory-ceiling test
pixi run -e default ruff check .
pixi run -e default mypy --strict src/annizarr
pre-commit run --all-files
Golden stores for regression tests live in tests/golden/.
License
MIT — see LICENSE.
Metadata
Release files for annizarr 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| annizarr-0.1.0.tar.gz | 455.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| annizarr-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 539.4 kB
Release files / annizarr-0.1.0.tar.gz
| Download URL | annizarr-0.1.0.tar.gz |
|---|---|
| Size | 455.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
23418df5d8574e9fad583a16094b8c81d5535774691f62d2572b70bd66781562
|
|
BLAKE2b-256 checksum How to use checksums |
c9a8c5ad285fc7365e1b55b2af97586e19ae55bcea6068eff83f1073f7caa296
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 7, 2026.
Transparency logRelease files / annizarr-0.1.0-py3-none-any.whl
| Download URL | annizarr-0.1.0-py3-none-any.whl |
|---|---|
| Size | 83.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
0e1780bfef2d66eefd9ae535a349253695b2e45b4056f8a02d9a0e851121b3c9
|
|
BLAKE2b-256 checksum How to use checksums |
94cf6fc9f423ce21f1f34de33dc2f70f6db8d3e82e912a7470f2ac82b59fbad6
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 7, 2026.
Transparency log