Skip to main content

AnniZarr

CI PyPI

Convert to, append, sort anndata zarr storages efficiently with parallel and lazy operations for all commands and api usage. Also, enable versioned anndata storage with Icechunk ease of use.

Install

annizarr and the shorter anz are the same command. The icechunk extra adds versioned stores and remote (s3://, gs://) targets.

pip install annizarr
pip install "annizarr[icechunk]"

To work on the code instead, use the pixi dev environment:

pixi install

Quickstart usage

Convert

The input format is sniffed from the HDF5 contents (h5ad vs 10x); --from h5ad|10x overrides it. Configuration can enable chunking, sharding, format ... Two or more inputs are concatenated. --ic writes an Icechunk repository instead of plain zarr.

annizarr convert sample.h5ad -o sample.zarr

annizarr convert a.h5ad b.h5ad -o merged.zarr
annizarr convert sample.h5ad -o repo.icechunk --ic -m "initial import"

Zarr operations and edits

rechunk rewrites one matrix with new chunking and streams everything else through unchanged.
sort physically orders rows by obs column(s) into a new store.
append adds cells to an existing zarr/ic store; it prompts before dropping obsm/obsp/layers (--drop-derived consents up front)
add-expr adds a log-normalized layer (layers/gexp) derived from X/. Csc by default for fast column reads

annizarr rechunk merged.zarr -o rechunked.zarr --row-chunk 2000

annizarr sort merged.zarr -o sorted.zarr --by cell_type
annizarr append sorted.zarr more_cells.zarr --drop-derived

annizarr add-expr merged.zarr --layout csr
annizarr add-expr repo.icechunk --layout csr --branch dev -m "add lognorm layer"

Icechunk history, branches, cherry-picks live in the Python API annizarr.Repo below

Configuration and zarr stores

Every flag is documented in annizarr <command> --help; docs/cli.md has the same reference in one place. Inputs stream band by band by default, so memory stays bounded at any store size (--eager loads the whole input first, faster for small files), and matrix writes use every core unless --cpus says otherwise.

Some of the main options; the docs have all of them:

  • --cpus N — parallel band workers (default: all cores). Every command.
  • --layout csr|csc|dense — on-disk layout of the matrix being written. Default csr for X on convert, csc for the add-expr layer. CSR suits row-wise access, CSC and dense column queries.
  • --row-chunk N, --col-chunk N — chunk shape. Exact rows and columns for dense; for sparse, about N cells (csr) or N genes (csc) per chunk, sized from the average nonzeros. Defaults: 2048 for dense, about 9,000,000 nonzeros for sparse. convert, rechunk, add-expr.
  • --auto-shard — shard the 1-D sparse arrays and anndata-written elements with zarr's automatic shard shape (default: off). Every command but append.
  • --ic — write through an Icechunk repository; --branch B and -m MSG pick the branch and commit message. append and add-expr detect an existing repo on their own. convert, rechunk, sort.
  • --sort-by COL… on convert / --by COL… on sort — physically order rows by obs columns, primary key first

What gets written. Every store is anndata-readable zarr v3 with encoding-type/encoding-version attrs matching anndata 0.13's on-disk spec. X and layers may be dense, CSR or CSC on input (h5ad or 10x, mixed across inputs) and are written in whatever --layout asks for; X defaults to CSR. convert/rechunk/sort write to a sibling temp store, verify it opens, then rename it onto the target, so a killed run never leaves a partial store. On Icechunk every op is exactly one commit, and remote (s3://, gs://) outputs require it. More in docs/stores.md.

Python API for Icechunk usage

The full Python API, including the config dataclasses and the ops functions, is in docs/python-api.md.

import annizarr as az

repo = az.Repo("repo.icechunk")

repo.branches()
repo.log()
repo.tree()

root = repo.open_zarr("w")
root.attrs["step"] = "lognorm"
repo.commit("normalize")

repo.checkout("experiment", create=True)  # switches this object only; nothing is persisted
old = repo.open_zarr("r", snapshot_id=repo.log()[1].id)

repo.copy("/work/store")  # clone with every branch and snapshot id intact, good for readonly cases

OpResult(path, n_obs, n_vars, snapshot_id) — snapshot_id is None for plain zarr, the committed Icechunk snapshot id otherwise.

Development

pixi install
pixi run -e default pytest              # excludes -m slow by default
pixi run -e default pytest -m slow      # large synthetic-store memory-ceiling test
pixi run -e default ruff check .
pixi run -e default mypy --strict src/annizarr
pre-commit run --all-files

Golden stores for regression tests live in tests/golden/.

License

MIT — see LICENSE.

Metadata

Release files for annizarr 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for annizarr 0.1.0
File Size Uploaded
annizarr-0.1.0.tar.gz 455.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for annizarr 0.1.0
File Interpreter ABI Platform
annizarr-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 539.4 kB

Release files / annizarr-0.1.0.tar.gz

Download URL annizarr-0.1.0.tar.gz
Size 455.5 kB
Tags Source
SHA-256 checksum
How to use checksums
23418df5d8574e9fad583a16094b8c81d5535774691f62d2572b70bd66781562
BLAKE2b-256 checksum
How to use checksums
c9a8c5ad285fc7365e1b55b2af97586e19ae55bcea6068eff83f1073f7caa296
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 7, 2026.

Transparency log

Release files / annizarr-0.1.0-py3-none-any.whl

Download URL annizarr-0.1.0-py3-none-any.whl
Size 83.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
0e1780bfef2d66eefd9ae535a349253695b2e45b4056f8a02d9a0e851121b3c9
BLAKE2b-256 checksum
How to use checksums
94cf6fc9f423ce21f1f34de33dc2f70f6db8d3e82e912a7470f2ac82b59fbad6
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 7, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page