Skip to main content

PBZ (Per-Base Zarr) — a Zarr v3 convention for per-base genomic data

Project description

pbzarr

A Zarr v3 convention for storing per-base resolution genomic data — read depths, methylation, boolean masks, and other cohort-shaped per-base values. Built as an alternative to D4 and bigWig that compresses cleanly across samples and integrates with the xarray / zarr ecosystem.

This repo provides:

  • pbzarr (Rust crate) — store layout, metadata, region I/O. Delegates array storage and compression to zarrs. This is the only crate published to crates.io.
  • pbzarr-readers (Rust crate, unpublished) — input-format readers (currently d4). Kept separate so the core crate stays free of git dependencies and remains publishable. d4 import is reachable from the Python wheel or when building from this repo, not from the crates.io pbzarr crate.
  • pbzarr (Python wheel) — PyO3 binding for import_d4 plus pure-Python create_store / create_track over zarr-python. Read API is an xarray accessor.

The on-disk format is the same; both libraries write .pbz stores that the other can read.

Install — Rust

[dependencies]
pbzarr = "0.1"

Install — Python

pip install pbzarr            # once published on PyPI
# or from source: pixi run install-wheel

The Python wheel pulls in zarr>=3, xarray>=2024.10, numpy>=2.

Quickstart — Python

import pbzarr

# 1. Create the store
pbzarr.create_store(
    "out.pbz",
    contigs=["chr1", "chr2"],
    contig_lengths=[248_956_422, 242_193_529],
    coordinate_space="GRCh38",
)

# 2. Register a 1D scalar track
pbzarr.create_track("out.pbz", track="mask", dtype="bool")

# 2b. Or a 2D cohort track
pbzarr.create_track(
    "out.pbz",
    track="depth",
    dtype="int32",
    columns=["A", "B", "C"],
    column_dim="sample",
)

# 3a. Bulk-import from d4 (PyO3 -> Rust)
pbzarr.import_d4(
    "out.pbz",
    track="depth",
    sources=[("/data/A.d4", "A"), ("/data/B.d4", "B"), ("/data/C.d4", "C")],
)

# 3b. Or write arbitrary numpy data via zarr-python
import zarr, numpy as np
g = zarr.open_group("out.pbz", mode="r+")
g["chr1/mask"][:] = np.random.rand(248_956_422) > 0.5

Read with xarray

import pbzarr

dt = pbzarr.open("out.pbz")                 # xr.DataTree
dt.pbz.tracks                                # ['depth', 'mask']
dt.pbz.region("chr1:1000-2000")              # xr.Dataset (one contig, sliced)
dt.pbz.region("chr1:1000-2000", track="depth")             # xr.DataArray
dt.pbz.region("chr1:1000-2000", track="depth", column="A") # 1D DataArray

The .pbz accessor on xr.DataTree is registered when you import pbzarr. Regions use 0-based, half-open coordinates (chr1:1000-2000 is [1000, 2000)).

Note: Python create_store / create_track consolidate metadata after each call. Stores written by the Rust crate don't consolidate yet, so pbzarr.open(...) will emit a benign RuntimeWarning for those; run zarr.consolidate_metadata(path) once to silence.

Quickstart — Rust

use ndarray::Array2;
use pbzarr::io::Dtype;
use pbzarr::import::Config;
use pbzarr::{Contig, Genome, PbzStore, Region, TrackConfig};
// d4 import lives in the unpublished pbzarr-readers crate (it carries a git
// dependency on d4); depend on this repo by path or git to use it.
use pbzarr_readers::d4::{from_d4, D4Source};

fn main() -> Result<(), Box<dyn std::error::Error>> {
    let genome = Genome::new(vec![
        Contig { name: "chr1".into(), length: 248_956_422 },
        Contig { name: "chr2".into(), length: 242_193_529 },
    ])?;
    let mut store = PbzStore::create("out.pbz", genome, Some("GRCh38".into()))?;

    // 1D scalar track
    store.create_track("mask", TrackConfig::new(Dtype::Bool))?;

    // 2D cohort track; d4 import requires int32
    store.create_track(
        "depth",
        TrackConfig::new(Dtype::I32)
            .columns(vec!["A".into(), "B".into(), "C".into()])
            .column_dim("sample"),
    )?;

    // Import from d4
    from_d4(
        &store,
        "depth",
        &[
            D4Source { path: "/data/A.d4".into(), sample_label: Some("A".into()) },
            D4Source { path: "/data/B.d4".into(), sample_label: Some("B".into()) },
            D4Source { path: "/data/C.d4".into(), sample_label: Some("C".into()) },
        ],
        Config::default(),
    )?;

    // Read a region
    let chr1 = store.genome().id("chr1").unwrap();
    let region = Region { contig: chr1, start: 1_000, end: 2_000 };
    let data = store.track("depth").unwrap().read_region::<i32>(&region)?;
    let arr2: Array2<i32> = data.into_dimensionality::<ndarray::Ix2>()?;
    let _ = arr2;
    Ok(())
}

Format at a glance

  • Layout: contig-major Zarr v3 store with <contig>/<track> arrays. Position is the first axis; an optional column dim (default name "column", often overridden to "sample") is the second.
  • Tracks: 1D for scalar (e.g., masks), 2D for cohort (e.g., per-sample depths). Rank-faithful on disk.
  • Coordinates: 0-based, half-open.
  • Compression: Blosc(zstd-5, byte-shuffle) on every data array.
  • Coord arrays: cohort tracks write per-contig 1D string arrays at <contig>/<column_dim> listing the column labels; xarray promotes them to coordinates automatically.

For the full design see docs/DESIGN.md.

Links

Development note

The per-base Zarr format that pbzarr standardizes was first prototyped by hand in clam, where the initial concepts (the contig-major layout, cohort-shaped tracks, and the zarr/ndarray I/O path) were fleshed out before AI tooling was introduced. pbzarr lifts those concepts into a dedicated, spec-driven library.

From that point, development of pbzarr was heavily assisted by Claude (Anthropic), accelerating the library implementation, d4 import, tests, and documentation. The architecture, domain knowledge, and direction remain the author's own; Claude was used as an accelerant, not an author.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distributions

No source distribution files available for this release.See tutorial on generating distribution archives.

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

pbzarr-0.2.0-cp311-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (3.3 MB view details)

Uploaded CPython 3.11+manylinux: glibc 2.17+ x86-64

pbzarr-0.2.0-cp311-abi3-macosx_10_12_x86_64.macosx_11_0_arm64.macosx_10_12_universal2.whl (5.7 MB view details)

Uploaded CPython 3.11+macOS 10.12+ universal2 (ARM64, x86-64)macOS 10.12+ x86-64macOS 11.0+ ARM64

File details

Details for the file pbzarr-0.2.0-cp311-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.

File metadata

File hashes

Hashes for pbzarr-0.2.0-cp311-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Algorithm Hash digest
SHA256 c3a4d854a8909dffc7cb6d989c570e148b53229d076e511ce443ff54509cab91
MD5 d0603574d29c9c09f26cd37982f255fe
BLAKE2b-256 48d8907f4239753cd937c9b03ff06a2c7a07d8353c5a62174a947bb24fff1b98

See more details on using hashes here.

Provenance

The following attestation bundles were made for pbzarr-0.2.0-cp311-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl:

Publisher: release.yml on pbzarr/pbzarr

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file pbzarr-0.2.0-cp311-abi3-macosx_10_12_x86_64.macosx_11_0_arm64.macosx_10_12_universal2.whl.

File metadata

File hashes

Hashes for pbzarr-0.2.0-cp311-abi3-macosx_10_12_x86_64.macosx_11_0_arm64.macosx_10_12_universal2.whl
Algorithm Hash digest
SHA256 4b19151872201b1a869cb1c04b85c26e21362287dcc90924b931b28d8029760a
MD5 822b2dccb80b1992b5688ff00ba27774
BLAKE2b-256 120b8b1e778055d74cc3117fe1ba806922c50915088a0f6a965806bdb2dcbb46

See more details on using hashes here.

Provenance

The following attestation bundles were made for pbzarr-0.2.0-cp311-abi3-macosx_10_12_x86_64.macosx_11_0_arm64.macosx_10_12_universal2.whl:

Publisher: release.yml on pbzarr/pbzarr

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page