A minibatch loader for AnnData stores

These details have been verified by PyPI

Project links

GitHub Statistics

Maintainers

ilan-gold

These details have not been verified by PyPI

Project links

Documentation

Project description

annbatch

[!IMPORTANT] This package will now only make breaking changes on the minor version release until its major release.

A data loader and io utilities for mini-batched data loading of on-disk AnnData files, co-developed by Lamin Labs and scverse

Getting started

Please refer to the documentation, in particular, the API documentation.

Installation

You need to have Python 3.12 or newer installed on your system. If you don't have Python installed, we recommend installing uv.

To install the latest release of annbatch from PyPI:

pip install "annbatch[zarrs]"

We provide extras for torch, cupy-cuda12, cupy-cuda13, and zarrs-python. cupy provides accelerated handling of the data via preload_to_gpu once it has been read off disk and does not need to be used in conjunction with torch.

[!IMPORTANT] zarrs-python gives the necessary performance boost for the sharded data produced by our preprocessing functions to be useful when loading data off a local filesystem.

To install all optional dependencies::

pip install "annbatch[zarrs,torch,cupy-cuda13]"

(Note: Replace cupy-cuda13 with the extra matching your local CUDA version)

Detailed tutorial

For a detailed tutorial, please see the in-depth section of our docs

Basic usage example

Basic preprocessing:

from annbatch import DatasetCollection

import zarr
from pathlib import Path

# Using zarrs is necessary for local filesystem performance.
# Ensure you installed it using our `[zarrs]` extra i.e., `pip install "annbatch[zarrs]"` to get the right version.
zarr.config.set(
    {"codec_pipeline.path": "zarrs.ZarrsCodecPipeline"}
)

# Create a collection at the given path. The subgroups will all be anndata stores.
collection = DatasetCollection("path/to/output/collection.zarr")
collection.add_adatas(
    adata_paths=[
        "path/to/your/file1.h5ad",
        "path/to/your/file2.h5ad"
    ],
    shuffle=True,  # shuffling is needed if you want to use chunked access, but is the default
)

Data loading:

[!IMPORTANT] Without custom loading via {meth}annbatch.Loader.use_collection or load_adata{s} or load_dataset{s}, all columns of the (obs) {class}pandas.DataFrame will be loaded and yielded potentially degrading performance.

from pathlib import Path

from annbatch import DatasetCollection, Loader
import anndata as ad
import zarr

# Using zarrs is necessary for local filesystem performance.
# Ensure you installed it using our `[zarrs]` extra i.e., `pip install "annbatch[zarrs]"` to get the right version.
zarr.config.set(
    {"codec_pipeline.path": "zarrs.ZarrsCodecPipeline"}
)


# WARNING: Without custom loading *all* obs columns will be loaded and yielded potentially degrading performance.
def custom_load_func(g: zarr.Group) -> ad.AnnData:
    return ad.AnnData(
        X=ad.io.sparse_dataset(g["layers"]["counts"]),
        obs=ad.io.read_elem(g["obs"])[some_subset_of_columns_useful_for_training]
    )


# A non empty collection
collection = DatasetCollection("path/to/output/collection.zarr")
# This settings override ensures that you don't lose/alter your categorical codes when reading the data in!
with ad.settings.override(remove_unused_categories=False):
    ds = Loader(
        batch_size=4096,
        chunk_size=32,
        preload_nchunks=256,
        to_torch=True
    )
    # `use_collection` automatically uses the on-disk `X` and full `obs` in the `Loader`
    # but the `load_adata` arg can override this behavior
    # (see `custom_load_func` above for an example of customization).
    ds = ds.use_collection(collection, load_adata=custom_load_func)

# Iterate over dataloader (plugin replacement for torch.utils.DataLoader)
for batch in ds:
    x, obs = batch["X"], batch["obs"]
    # Important: For performance reasons convert to dense on GPU
    x = x.cuda().to_dense()

[!IMPORTANT] For usage of our loader inside of torch, please see this note for more info. At the minimum, be aware that deadlocking will occur on linux unless you pass multiprocessing_context="spawn" to the torch.utils.data.DataLoader class.

Release notes

See the changelog.

Contact

For questions and help requests, you can reach out in the scverse discourse. If you found a bug, please use the issue tracker.

Citation

If you use annbatch in your work, please cite the annbatch publication as follows:

annbatch unlocks terabyte-scale training of biological data in anndata

Gold, I., Fischer, F., Arnoldt, L., Wolf, F. A., & Theis, F. J. (2026b). annbatch unlocks terabyte-scale training of biological data in anndata. arXiv (Cornell University). https://doi.org/10.48550/arxiv.2604.01949

Project details

These details have been verified by PyPI

Project links

GitHub Statistics

Maintainers

ilan-gold

These details have not been verified by PyPI

Project links

Documentation

Release history Release notifications | RSS feed

0.1.5

May 7, 2026

This version

0.1.4

May 4, 2026

0.1.3

Apr 15, 2026

0.1.2

Mar 26, 2026

0.1.1

Mar 25, 2026

0.1.0

Mar 18, 2026

0.0.8

Feb 12, 2026

0.0.7

Feb 2, 2026

0.0.6

Jan 26, 2026

0.0.5

Jan 23, 2026

0.0.4

Jan 21, 2026

0.0.3

Jan 19, 2026

0.0.2

Jan 16, 2026

0.0.1

Oct 30, 2025

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

annbatch-0.1.4.tar.gz (262.5 kB view details)

Uploaded May 4, 2026 Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

The dropdown lists show the available interpreters, ABIs, and platforms. Enable javascript to be able to filter the list of wheel files.

annbatch-0.1.4-py3-none-any.whl (42.5 kB view details)

Uploaded May 4, 2026 Python 3

File details

Details for the file annbatch-0.1.4.tar.gz.

File metadata

Download URL: annbatch-0.1.4.tar.gz
Upload date: May 4, 2026
Size: 262.5 kB
Tags: Source
Uploaded using Trusted Publishing? Yes
Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for annbatch-0.1.4.tar.gz
Algorithm	Hash digest
SHA256	`9eeee827b2eae4979f93560fc4c9886a05b29ba8147e46b3f643a910ef19bd49`
MD5	`6ed39e964d02975a3efa6ab068e98c4e`
BLAKE2b-256	`1963f4d7faaf95f9be195abc0920ef49c684f2848230222746bf29931643fd8f`

See more details on using hashes here.

Provenance

The following attestation bundles were made for annbatch-0.1.4.tar.gz:

Publisher: release.yaml on scverse/annbatch

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Statement:
- Statement type: https://in-toto.io/Statement/v1
- Predicate type: https://docs.pypi.org/attestations/publish/v1
- Subject name: annbatch-0.1.4.tar.gz
- Subject digest: 9eeee827b2eae4979f93560fc4c9886a05b29ba8147e46b3f643a910ef19bd49
- Sigstore transparency entry: 1437439627
- Sigstore integration time: May 4, 2026
Source repository:
- Permalink: scverse/annbatch@a146f50e79a4dbb9ba68044f9e929c9931ae2c25
- Branch / Tag: refs/tags/v0.1.4
- Owner: https://github.com/scverse
- Access: public
Publication detail:
- Token Issuer: https://token.actions.githubusercontent.com
- Runner Environment: github-hosted
- Publication workflow: release.yaml@a146f50e79a4dbb9ba68044f9e929c9931ae2c25
- Trigger Event: release

File details

Details for the file annbatch-0.1.4-py3-none-any.whl.

File metadata

Download URL: annbatch-0.1.4-py3-none-any.whl
Upload date: May 4, 2026
Size: 42.5 kB
Tags: Python 3
Uploaded using Trusted Publishing? Yes
Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for annbatch-0.1.4-py3-none-any.whl
Algorithm	Hash digest
SHA256	`d80d36ab39a5611e8760b4146674e919b11a5492839d7931378bd723fa005650`
MD5	`49e3d4e740725af6682e677687c50138`
BLAKE2b-256	`3a4b9bf4546f404f96fa8b58050bb3465a275344edb8656705f13a782b006fa1`

See more details on using hashes here.

Provenance

The following attestation bundles were made for annbatch-0.1.4-py3-none-any.whl:

Publisher: release.yaml on scverse/annbatch

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Statement:
- Statement type: https://in-toto.io/Statement/v1
- Predicate type: https://docs.pypi.org/attestations/publish/v1
- Subject name: annbatch-0.1.4-py3-none-any.whl
- Subject digest: d80d36ab39a5611e8760b4146674e919b11a5492839d7931378bd723fa005650
- Sigstore transparency entry: 1437439694
- Sigstore integration time: May 4, 2026
Source repository:
- Permalink: scverse/annbatch@a146f50e79a4dbb9ba68044f9e929c9931ae2c25
- Branch / Tag: refs/tags/v0.1.4
- Owner: https://github.com/scverse
- Access: public
Publication detail:
- Token Issuer: https://token.actions.githubusercontent.com
- Runner Environment: github-hosted
- Publication workflow: release.yaml@a146f50e79a4dbb9ba68044f9e929c9931ae2c25
- Trigger Event: release

annbatch 0.1.4

Navigation

Verified details

Project links

GitHub Statistics

Maintainers

Unverified details

Project links

Meta

Classifiers

Project description

annbatch

Getting started

Installation

Detailed tutorial

Basic usage example

Release notes

Contact

Citation

Project details

Verified details

Project links

GitHub Statistics

Maintainers

Unverified details

Project links

Meta

Classifiers

Release history Release notifications | RSS feed

Download files

Source Distribution

Built Distribution

File details

File metadata

File hashes

Provenance

File details

File metadata

File hashes

Provenance