Loaderx
Zrecord is a rebuildable, typed, ordered record sequence built from authoritative
source data and scripts. A creator consumes records in append order, close
publishes one immutable container, and readers support both sequential slicing
and indexed gather without changing record identity. To change content or order,
rebuild it at a new path.
Zrecord is the typed on-disk container; Loaderx is the sampler and prefetch loader that consumes Zrecord streams. They currently ship together while both layers mature, but their public responsibilities remain separate.
pip install loaderx
Wheels are published for glibc Linux (x86-64 and ARM64), Apple Silicon macOS,
and Windows AMD64. Three domain Cython extensions separate Store, Sampler and
Ragged packing; Store and Sampler each statically link their Zig library.
cp310-abi3 serves regular CPython 3.10+, while free-threaded CPython 3.14 uses
cp314-cp314t.
Design Philosophy
loaderx is built around several core principles:
- A pragmatic approach that prioritizes minimal memory overhead and minimal dependencies.
- A strong focus on single-machine training workflows.
- We implement based on NumPy semantics, persisted by the private native store engine.
- An immortal (endless) step-based data loader, rather than the traditional epoch-based design—better aligned with modern ML training practices.
- Dense and ragged are separate contracts, and the loader serves both. A dense stream stacks into one array per batch; a ragged one comes back as a list. Neither is padded, and equal length is never treated as a special case of variable length.
- Logical IDs are stable sequence positions. Native chunks may complete in
any physical order, but append input order defines
0..N-1and a published container never deletes, compacts, updates, or renumbers those records.
设计文档
Quick Start
import numpy as np
from loaderx.zrecord import Dense
from loaderx.dataloader import DataLoader
data = np.load('data.npy', mmap_mode='r')
label = np.load('label.npy', mmap_mode='r')
with Dense.create('train_data', data.dtype, data.shape[1:], codec='zstd') as ds:
ds.append(data)
with Dense.create('train_label', label.dtype, label.shape[1:], codec='zstd') as ds:
ds.append(label)
data_store = Dense.open('train_data')
label_store = Dense.open('train_label')
loader = DataLoader({'data': data_store, 'label': label_store},
transform=lambda batch: batch)
for i, batch in enumerate(loader):
if i >= 256:
break
print(batch['data'].shape)
print(batch['label'].shape)
loader.close()
data_store.close()
label_store.close()
A batch is a dict {name: values}: each value is the stacked
(batch_size, *item_shape) array for that stream. Every stream is gathered
at the same indices, so record i lines up across them. The transform
callback is the collate step — reshape, cast, stack — where values is the
plain dense batch ready for the model.
Offline Hugging Face conversion
The optional converter uses Hugging Face Datasets for remote discovery,
download, caching, revision handling and source-format decoding. It then writes
explicit typed Zrecord streams, so training needs neither datasets nor Arrow:
pip install 'loaderx[converter]'
from loaderx.converter import convert, huggingface
dataset = huggingface("ylecun/mnist", revision="main", token=None)
convert(
dataset,
"mnist",
codec="zstd",
)
huggingface resolves the requested branch, tag or commit to an immutable
snapshot, then downloads and returns the repository's DatasetDict. Repositories
without one default config require config=.... convert is the independent
persistence stage: it accepts that DatasetDict or one already normalized
Dataset. Conversion batches are bounded internally by decoded numeric bytes,
not a public row-count parameter; they flush at 32K records or near 64 MiB,
whichever comes first.
For private or gated repositories, pass token="hf_..."; None uses the
standard HF environment, local login, or anonymous access.
Static or regularly batched numeric arrays become Dense stores, while dynamic
arrays become Ragged stores; source dtype and record shape are preserved.
Zrecord persists numeric tensors rather than Python objects. A standard HF feature that cannot be represented losslessly fails with its column name; conversion does not filter, pad, tokenize, reshape, or run a transform stage. Arrow dtypes, including float64 and uint64, are preserved. Hugging Face formats only bounded decode slices; the converter retains their numeric results only until the byte-aware append boundary and never formats the complete dataset.
The result groups aligned streams under one published root:
mnist/
train/
image/
label/
test/
image/
label/
Conversion happens under a temporary sibling directory. The root is renamed into place only after all stores close and their record counts agree; a failed build that raises a Python exception is discarded. A process kill may leave the hidden temporary sibling for manual cleanup. No collection manifest is added; the output is only split directories containing ordinary column Zrecord stores.
Creating a dense store
import numpy as np
from loaderx.zrecord import Dense
data = np.load('data.npy', mmap_mode='r')
with Dense.create('train_data', data.dtype, data.shape[1:], codec='zstd') as ds:
ds.append(data)
One record per slice along axis 0; a 1-D array (the usual shape of a label set)
becomes a store of scalar records. The caller controls each append batch and can
slice a large or mmapped array to set its own memory bound. Creation and opening
are explicit: create requires a new path and returns an append-only writer;
leaving its context calls close(), which publishes the result.
open requires an existing container produced by a successful build and returns a read-only
reader. Open the result with Dense.open:
ds = Dense.open('train_data')
batch = ds[:]
ds.close()
Python defines the exact record schema: both geometries store dtype and
ndim; Dense additionally stores item_shape. The MsgPack bytes live opaquely in the
static page at the front of meta.zr; Zig persists them but never interprets
them. The schema accepts no user metadata. Python selects the geometry and gives
the private native engine only the runtime record boundaries it needs. Each
Ragged record carries exactly ndim inline little-endian u64 dimensions.
One native physical
engine consumes the trusted Dense stride or Ragged offsets. Append inputs are
strictly NumPy arrays: Dense takes one batched ndarray and Ragged takes an
iterable of ndarrays. Raw bytes and pre-encoded images are made explicit with
np.frombuffer(raw, dtype=np.uint8) and stored in a
Ragged.create(path, dtype=np.uint8, ndim=1, codec="zstd") rather than creating a second
public storage API.
Records
One persistent format, two native execution contracts. Dense is the dense contract
where every record is exactly one row of the recorded item_shape; reads
allocate a fixed-stride destination whose batch shape follows from the schema.
Physical records still use the shared RecordLoc[logical_id] -> payload
pipeline, so compressed completion order never becomes a second Dense layout:
import numpy as np
from loaderx.zrecord import Dense
data = np.arange(64, dtype=np.float32).reshape(8, 2, 4)
with Dense.create('data', data.dtype, data.shape[1:], codec='zstd') as ds:
ds.append(data)
ds = Dense.open('data')
ds[0, 5, 2] # (3, 2, 4) — shape from the persisted schema
ds.close()
A store is an ordered sequence of records, not an ndarray, so ds[0, 5, 2]
selects sequence positions 0, 5 and 2 — never ds[0][5][2]. A scalar selects
one record; indices must be in 0..len(ds)-1. The ragged example below reads
the same way.
Ragged is the ragged contract for variable-length records. It is a
separate contract: :class:Ragged hands back a list of arrays, so a
loader never has to carry row_splits around. dtype and ndim are
unified and explicit; each record keeps its own dimension lengths, recorded
per record and restored exactly on read. Every axis may vary, but every record
has the schema rank; nothing is inferred from the source. Ragged requires at
least one axis; scalar records use Dense with item_shape=().
zero-byte arrays are rejected because physical records are nonempty. Densifying a list into a dense
batch is the model's call — a plain numpy loop, wherever you need it:
from loaderx.zrecord import Ragged
seqs = [np.arange(L, dtype=np.int32) for L in (3, 1, 4, 1, 5)]
with Ragged.create('tokens', np.int32, ndim=1, codec='zstd') as rs:
rs.append(seqs) # dtype/rank fixed; lengths remain per-record
rs = Ragged.open('tokens')
records = rs[0, 2, 4] # list of ndarray — one per record, exact shapes
rs.close()
lengths = np.array([len(r) for r in records])
padded = np.zeros((len(records), lengths.max()), dtype=records[0].dtype)
for i, r in enumerate(records):
padded[i, :len(r)] = r # (B, max_len) — your policy, your loop
Ordered sequences and build streams
Logical ID is the stable sequence position. One append preserves every
record in its input order; successive calls from one producer extend that
sequence. Native lanes may reserve and write payload chunks in a different
physical completion order, but each RecordLoc is installed in its original
logical slot, so physical scheduling never changes ds[i]. After close,
the sequence is immutable: there is no delete, compact, update, or reopen-append
operation that can renumber it.
This makes a creator a finite build stream and its published result an immutable sequence. It is not a live log: readers do not tail a writer, and an unbounded producer must choose a finite publication boundary. A large sequence can be consumed in bounded ordered batches with ordinary slices:
with Dense.open("events") as events:
for start in range(0, len(events), 1024):
batch = events[start:start + 1024]
consume(batch)
Ragged creators may consume a one-pass iterable, while Dense creators take explicit ndarray batches. Dense concurrent append calls are serialized by the native writer lock, and their batch order is lock-acquisition order. Ragged packing is synchronous and each batch enters the same native mutation boundary before append returns. Sequence builders that require an external temporal order should still use one producer or order batches before append.
A DataLoader dynamically composes a dict of dense and ragged streams.
Collation is the transform — a batch dict in, a batch dict out:
def collate(batch):
return {'input_ids': batch['tokens'], 'label': batch['label']}
loader = DataLoader({'tokens': dense_tokens, 'label': labelset}, batch_size=32,
transform=collate)
batch = next(loader)
loader.close()
The transform runs once for each gathered batch on a loader transform worker. Dense values are independent writable contiguous arrays; Ragged records share one backing allocation per batch. The return value is handed to the consumer unchanged. Calls may run concurrently and complete out of order, so the callback must be thread-safe; exceptions are propagated to the consumer. Keep shared mutable state and nested parallel runtimes out of the callback.
Numba can optionally accelerate a CPU-heavy Dense transform while releasing the GIL. Compile it before timing, then call it from the ordinary transform:
import numba
import numpy as np
@numba.njit(nogil=True, parallel=False)
def normalize_u8(x):
out = np.empty(x.shape, dtype=np.float32)
for i in range(x.size):
out.flat[i] = x.flat[i] / 255.0
return out
normalize_u8(np.zeros((1, 3, 224, 224), dtype=np.uint8)) # compile warmup
def transform(batch):
batch["image"] = normalize_u8(batch["image"])
return batch
Numba is neither a loaderx dependency nor a loader backend. Its compilation warmup belongs outside build or loader benchmark timing.
Creating containers
Dense.create and Ragged.create return append-only ordered-sequence
builders. The codec keyword is required: callers must explicitly select
"raw", "zstd", or "zstd_dict". append is explicit — one input
batch extends the logical sequence without exposing native physical completion
order.
Dense append is synchronous and borrows an already-contiguous ndarray without a
snapshot copy. Ragged consumes its iterable once and validates the complete
batch before native mutation. Ragged snapshots each shape and retains the
normalized payload owners while Zig materializes the trusted byte spans in
parallel, either into raw shard scratch or lane-local compression input. It
completes native append before returning, without a writer thread, delayed
errors, a full-batch snapshot, or a second completion contract. Borrowed inputs
must remain quiescent until the synchronous call returns. Nothing is inferred.
from loaderx.zrecord import Dense, Ragged
ds = Dense.create('mnist/x', dtype=np.uint8, item_shape=(28, 28),
codec='zstd', data_shards=4)
ds.append(images[i:i + 1024]) # synchronous native batch; returns None
ds.append(single_image[None]) # one sample is batch_size 1 — add the axis yourself
ds.close() # publish before opening
tok = Ragged.create('tokens', dtype=np.int32, ndim=1, codec='zstd')
tok.append([seq_a, seq_b, seq_c]) # synchronous validation and native commit
tok.close() # publish
with Dense.open('mnist/x') as ds:
first_four = ds[:4] # opened containers are read-only
data_shards is the keyword-only write-parallelism setting. It is persisted in
the native Header, must be in 1..255, and defaults to four. Values near the
writer lane count spread payload writes across independent files; opening a
completed container discovers the value automatically.
Dense and Ragged append report validation and native errors in the current call.
A failed build is discarded and rebuilt rather than recovered. A writer cannot
be read, and a reader cannot be appended to.
close() on a writer publishes the container Header;
close() on a reader releases it. Dense, Ragged, and DataLoader support
with for scoped lifetimes. Content is never changed in place: rerun the
authoritative build at a new path, validate it, then switch consumers to it.
Loaderx optimizes for that conventional lifecycle. Use-after-close, closing an object from its own input/transform callback, concurrent close, fabricated instances, private-attribute mutation, and direct calls into private Cython or C ABIs are outside the contract and receive no dedicated validation or diagnostic. This does not relax public dtype/shape/index validation or native checks for persisted data, arithmetic and buffer bounds, I/O, codec results, locking, and commit ordering.
The exact schema is declared at creation and encoded by Python as MsgPack. Both
schemas contain dtype and ndim; Dense additionally contains
item_shape and requires its length to equal ndim. Ragged requires every
appended record to have that rank, while every dimension length may vary.
Structured, subarray, object, metadata-bearing, and zero-itemsize dtypes are not
supported: their semantics do not round-trip through one canonical NumPy dtype
string. The encoded schema has 2040 bytes available in the fixed 4096-byte
metadata page. Ragged schema size is fixed; Dense schema size grows
only with the integer item_shape, so the physical limit is far above any
practical NumPy array rank.
Unexpected fields are rejected. append validates dtype and shape in Python,
then passes the derived byte width to the private Dense operation. It does not
coerce Python lists, tuples, bytes, bytearrays, or other array-like objects;
callers convert them with np.asarray or np.frombuffer first.
Dictionary training stays explicit and separate; creating a
zstd_dict store requires the completed dictionary, so no valid store is
published without one.
Codec notes
"zstd" compresses each record independently with plain zstd (level 3). Use it
for general-purpose compression; it is fast and must be selected explicitly.
"zstd_dict" trains a shared dictionary on a sample of the data before writing
any record, then compresses every record against it at level 3. The dictionary
captures structure shared across records that per-record compression cannot see —
a useful fit for large corpora of small, similar records such as token sequences
and image tiles. Plain "zstd" or "raw" remains the better default for large
records and small corpora.
Dictionary geometry is fixed: train_dict trains a dictionary of at most 32 KiB
from a strided sample bounded at 4 MiB. Real WikiText and MNIST profiling found
that larger 128 KiB and 1 MiB dictionaries added training and cache cost without
a stable ratio or gather advantage. There is no dictionary tuning API.
"zstd_dict" records can only be read from a store that has the dictionary
(dict.zr). The dictionary is loaded on open and shared, lock-free, across all
reader threads.
A dictionary must train on the settled, complete data. :func:train_dict is the
standalone, manual training step — a numpy array, an iterable of records, or
a typed container all train the same fixed contract — and its bytes are handed to
a container-writing path via dict_bytes. zstd_dict never trains by itself: a
write without a dictionary is an error. A stream cannot train its own
dictionary, but it can write with one trained on the settled data:
from loaderx.utils import train_dict
from loaderx.zrecord import Dense, Ragged
# train once on the settled data — a standalone, reusable artifact
d = train_dict(settled_array)
# then any new store can install it and append explicitly
with Ragged.create('tokens', np.int32, ndim=1,
codec='zstd_dict', dict_bytes=d) as ds:
ds.append(token_generator)
with Dense.create('data', data.dtype, data.shape[1:],
codec='zstd_dict', dict_bytes=d) as ds:
ds.append(data)
Changing an existing store's codec is a rewrite, not a store mutation: read the
records as numpy values and append them to a new store with the new codec. A
whole-store slice preserves index order, which keeps multi-stream alignment.
There is no dedicated recode or store-to-store path because the ordinary
read and append contracts already express the operation:
from loaderx.utils import train_dict
from loaderx.zrecord import Dense
with Dense.open("src") as s, \
Dense.create("dst", dtype=s.dtype, item_shape=s.item_shape,
codec="zstd_dict", dict_bytes=train_dict(s)) as d:
d.append(s[:]) # ndarray -> dense append
Ragged is the same shape: s[:] returns list[np.ndarray], which
is exactly its append input; a destination is created with s.dtype and
s.ndim. The
native compression path bounds its own working memory; there is no public chunk
parameter. dst must not already hold a store.
Important: Train the dictionary from settled authoritative input before building the container. Training it before preprocessing is complete wastes compression and does not describe the final records.
Multi-stream stores
A zrecord store is one stream; a training sample is usually several named
streams (skeleton + label + id, tokens + label, ...). Composition is a plain
Python dict passed to DataLoader. There is no persistent wrapper,
manifest, directory convention or bundle mutation API. DataLoader verifies
that all streams have the same length, then gathers every stream with the same
indices.
from loaderx.zrecord import Dense, Ragged
from loaderx.dataloader import DataLoader
root = "xsub/train"
with Dense.create(root + "/joint", joint.dtype, joint.shape[1:], codec="zstd") as s:
s.append(joint)
with Dense.create(root + "/label", label.dtype, label.shape[1:], codec="zstd") as s:
s.append(label)
with Ragged.create(root + "/token", np.int32, ndim=1, codec="zstd") as s:
s.append(seqs)
streams = {
"joint": Dense.open(root + "/joint"),
"label": Dense.open(root + "/label"),
"token": Ragged.open(root + "/token"),
}
streams["joint"][0, 5, 2] # each stream keeps its own index API
loader = DataLoader(streams, batch_size=256)
batch = next(loader) # {name: values}, index-aligned
loader.close()
for stream in streams.values():
stream.close()
Equal-length and variable-length are both just stores — Dense (one
fixed-shape record per sample) and Ragged (variable row count per
record). The loader does not interpret either: it fetches by index and packs a
dict, so a dense stream's batch value is the stacked (B, *item_shape) array
and a ragged stream's is a list of per-record arrays. No padding is imposed —
densify to a fixed shape however the model needs (a plain numpy loop), or
reshape/stack in the transform collate.
CPU → GPU transfer
loaderx hands over CPU batches; getting them to the accelerator is the
transform's job — the one place your framework is already imported. The batch
dict is a plain {name: numpy array}, zero-copy on the way out, so a device
transfer is one call per stream:
import torch
device = "cuda:0"
def to_device(batch):
return {k: torch.from_numpy(v).to(device, non_blocking=True)
for k, v in batch.items()}
streams = {
"joint": Dense.open(root + "/joint"),
"label": Dense.open(root + "/label"),
"token": Ragged.open(root + "/token"),
}
loader = DataLoader(streams, transform=to_device)
for batch in loader:
model(batch) # already on device
Call loader.close() externally when the training loop exits. Do not close the
loader or its streams from transform; the streams remain caller-owned and
should be closed at the application lifecycle boundary after the loader stops.
A non_blocking=True copy is genuinely asynchronous only when its source is
pinned. loaderx does not pin memory for you — pinning is framework-owned
(torch's .pin_memory(), CUDA's cudaHostAlloc), and a vendor-free core stops
exactly at the CPU batch. Pin in the transform what you copy:
def to_device(batch):
return {k: torch.from_numpy(v).pin_memory().to(device, non_blocking=True)
for k, v in batch.items()}
JAX is the same shape — jax.device_put is already an asynchronous handoff on
GPU:
import jax
def to_device(batch):
return {k: jax.device_put(v) for k, v in batch.items()}
The transfer runs on the transform stage and never touches loaderx internals:
the copy overlaps the next batch's gather/transform, and any pinned pool is the
caller's to own and reuse. This is the entire H2D answer — there is no pin=
hook or device backend, because the only unified thing a multi-framework loader
can own is the CPU batch.
For a practical JAX/Flax integration, see the MNIST layer-representation example.
Benchmarks
Dense and Ragged are measured separately because they expose different
contracts, but every store table uses the same columns. scripts/bench_dense.py
measures fixed-shape random gather, scripts/bench_ragged.py measures
variable-shape records, and scripts/bench.py covers machine, sampler, and the
end-to-end loader comparison. Every path runs through the public Python binding,
so Cython dispatch, NumPy allocation, and Ragged list/shape reconstruction are timed.
Methodology
The comparison matrix began as one complete run on a warm page cache. All four
Store matrices, the sampler, and the primary loader matrix were refreshed on
2.4.2 after Cython became the direct Store/Sampler binding, with the default
data_shards=4. Every workload uses a prepared real source.
These are not three-run medians: an unexpected result is traced separately
instead of being hidden by repeated aggregation. Within each store workload every
backend receives identical source records. Correctness and timing use independent
deterministic zsampler IID streams. Loader backends receive
the same source and seed but use their own shipped samplers, so their exact
permutations differ. Before gather timing, the store benchmark sweeps every record and
validates a fixed number of IID batches for exact
dtype, shape, order, and values. logical write and logical gather divide
uncompressed NumPy payload bytes by elapsed time; they measure bytes accepted or
returned by the public API, not physical storage bandwidth. Writable containers
are created before timing. logical write starts when the already-generated
source enters the backend, includes byte encoding, packing, key construction and
Arrow array construction, and ends after logical commit/finalization returns. No
backend requests fsync, LMDB env.sync(), or another stable-media durability
operation. Reusable source preparation is outside that timer; in particular,
zstd_dict trains its standalone dictionary first.
The measured Store paths are explicit:
write: source is ready -> writer setup [untimed]
-> start -> encode/pack -> append/write -> logical finalize -> return -> stop
-> resource-only cleanup [untimed]
read: open -> full warm sweep -> IID correctness stream(seed) [untimed]
-> IID timing stream(seed + 1): sampler.next() [untimed]
-> start one gather -> public return -> stop
-> destroy returned batch [untimed]
-> repeat timed calls until their accumulated time is at least 2 seconds
Logical finalize means Zrecord Header publication, LMDB transaction commit, or
an Arrow/Parquet footer; none of these paths requests stable-media sync.
Disk size is allocated blocks, not sparse apparent size. krecords/s is gather
record throughput and p95 is the 95th-percentile latency of one random gather
batch. Results are comparable within one workload table, not across payload
distributions or geometries.
The finalized output is opened read-only before timing. After the warm sweep and
correctness stream, an independent IID stream draws fresh indices until timed
gather calls accumulate at least two seconds. IID sampling is uniform with
replacement, and sampler time is excluded. Output allocation, reads, decompression and
reconstruction are included, while open, close and destruction after return are
not.
Every backend name states its actual codec; the full default set is required
rather than silently skipped when a package is missing.
The vision source is the Oxford-IIIT Pet train split prepared by
scripts/prepare_vision.py. Dense uses RGB photographs resized on the short side
to 256 and center-cropped to (3, 224, 224); Ragged uses the same ordered source
at native RGB resolution. Both are mmap-loaded uint8 CHW records.
Machine — one local workstation (AMD Ryzen AI 9 HX PRO 370, 12 cores / 24 threads):
| machine | value |
|---|---|
| CPU | AMD Ryzen AI 9 HX PRO 370 w/ Radeon 890M, 1 socket, 12 cores / 24 threads |
| frequency | 605–5158 MHz |
| caches | L1d 576 KiB, L1i 384 KiB, L2 12 MiB, L3 24 MiB |
| NUMA | 1 node |
| memory | 31 GiB (not limited by cgroup) |
| shared memory | 16 GiB /dev/shm |
| OS | Debian GNU/Linux forky/sid, kernel 7.1.8+deb13-amd64, x86_64 |
| python | CPython 3.14.7 (standard GIL build), numpy 2.5.2 |
The benchmark process sees all 24 threads and is not memory-limited by cgroup. The 16 GiB shared-memory mount accommodates the four-worker, 36.8 MiB-batch torch pipeline. Store reads run on the ordinary page cache.
Direct-binding A/B. To separate the 2.4.2 CFFI removal from cross-pass machine
state, revision e8cddf4 was rebuilt in a separate worktree and run in the same
session with the same interpreter and corpus. These focused diagnostic passes are
not substituted into the complete tables below: Dense vision raw append moved
from 4866 to 5978 MiB/s (+23%), Ragged vision raw append from 1990 to 2200 MiB/s
(+11%), Dense token raw gather from 5667 to 6078 MiB/s (+7%), and Ragged token
gather moved from 481/583/750 to 528/642/827 MiB/s (+10% for each codec). The raw
end-to-end loader moved from 180.5 to 202.1 batches/s (+12%). Compression writes
did not show a consistent binding gain because codec work dominates; sampler
figures at sub-microsecond batch sizes also fluctuate too much for a standalone
percentage claim. The evidence supports lower boundary overhead, not a uniform
speedup of the Zig engine.
Large Vision Records
Fixed-Shape Dense
Zrecord against array-store alternatives: random batch gather over the first
2,500 Oxford-IIIT Pet train images, batch 256. Every prepared image is
uint8[3,224,224]; metadata.json records the source revision, transform,
decoder versions and logical SHA-256.
Fixed-resolution vision records — 147 KiB per record, 36.8 MiB per batch:
| backend | logical write | logical gather | krecords/s | p95 | disk | ratio |
|---|---|---|---|---|---|---|
| zrecord-raw | 5749 MiB/s | 15231 MiB/s | 106.1 | 2.74 ms | 358.9 MiB | 1.00x |
| npy-mmap-raw | 2101 MiB/s | 5143 MiB/s | 35.8 | 8.11 ms | 358.9 MiB | 1.00x |
| hdf5-raw | 2488 MiB/s | 2036 MiB/s | 14.2 | 22.12 ms | 359.0 MiB | 1.00x |
| lmdb-raw | 1627 MiB/s | 4313 MiB/s | 30.0 | 9.94 ms | 361.4 MiB | 0.99x |
| arrow-ipc-raw | 1061 MiB/s | 3769 MiB/s | 26.3 | 11.59 ms | 358.9 MiB | 1.00x |
| parquet-raw | 1396 MiB/s | 466 MiB/s | 3.2 | 83.08 ms | 358.9 MiB | 1.00x |
| arrayrecord-raw | 1857 MiB/s | 3293 MiB/s | 22.9 | 12.71 ms | 359.2 MiB | 1.00x |
| zrecord-zstd | 1603 MiB/s | 6812 MiB/s | 47.5 | 6.13 ms | 311.1 MiB | 1.15x |
| zrecord-zstdict | 1713 MiB/s | 8143 MiB/s | 56.7 | 5.35 ms | 320.1 MiB | 1.12x |
| hdf5-gzip | 49 MiB/s | 245 MiB/s | 1.7 | 165.36 ms | 301.8 MiB | 1.19x |
| arrow-ipc-zstd | 253 MiB/s | 103 MiB/s | 0.7 | 378.58 ms | 305.9 MiB | 1.17x |
| parquet-zstd | 244 MiB/s | 93 MiB/s | 0.6 | 421.82 ms | 305.9 MiB | 1.17x |
| arrayrecord-zstd | 233 MiB/s | 1866 MiB/s | 13.0 | 23.17 ms | 320.1 MiB | 1.12x |
At 147 KiB per record, Zrecord-raw reaches 14.9 GiB/s and is 3.0x npy-mmap-raw. Decoded photographs have little remaining redundancy: plain zstd reduces them only 1.15x, while dictionary mode falls to 1.12x. This is why large real images should use plain zstd or raw. LMDB and Arrow IPC are competitive raw record stores, while codecs tied to whole IPC batches or Parquet row groups pay read amplification on random gathers. Dense demonstrates that typed record ownership and per-record compression do not turn fixed tensors into an object-store slow path.
Native-Resolution Ragged
This workload contains the first 2,500 native-resolution Oxford-IIIT Pet CHW RGB
images. Height is 108..2606 (median 375) and width is 117..3264 (median 500),
totaling 1252.3 MiB of logical uint8 payload. Each of 50 random
batches contains 256 records. Every backend persists payload plus exact shape
and must return an ordered list[np.ndarray] of (3, H, W) arrays; a flat byte
list or a one-dimensional variable-length abstraction is not enough.
| backend | logical write | logical gather | krecords/s | p95 | disk | ratio |
|---|---|---|---|---|---|---|
| zrecord-raw | 5171 MiB/s | 13921 MiB/s | 27.9 | 13.03 ms | 1252.4 MiB | 1.00x |
| hdf5-raw | 1752 MiB/s | 2122 MiB/s | 4.2 | 75.72 ms | 1253.1 MiB | 1.00x |
| lmdb-raw | 1675 MiB/s | 8875 MiB/s | 17.7 | 21.59 ms | 1257.0 MiB | 1.00x |
| arrow-ipc-raw | 1242 MiB/s | 4552 MiB/s | 9.1 | 39.58 ms | 1252.4 MiB | 1.00x |
| parquet-raw | 694 MiB/s | 523 MiB/s | 1.0 | 273.64 ms | 1252.4 MiB | 1.00x |
| arrayrecord-raw | 1729 MiB/s | 2819 MiB/s | 5.6 | 66.91 ms | 1253.0 MiB | 1.00x |
| zrecord-zstd | 1371 MiB/s | 5104 MiB/s | 10.2 | 46.26 ms | 1032.2 MiB | 1.21x |
| zrecord-zstdict | 897 MiB/s | 4975 MiB/s | 9.9 | 47.31 ms | 1032.9 MiB | 1.21x |
| arrayrecord-zstd | 195 MiB/s | 1983 MiB/s | 4.0 | 100.94 ms | 1054.6 MiB | 1.19x |
Zrecord-raw is 1.6x LMDB and 3.1x Arrow IPC in logical gather. Plain and dictionary zstd both reach 1.21x, confirming that dictionary mode is not useful for these large photographs. The Ragged matrix is intentionally asymmetric: HDF5 gzip, Arrow IPC zstd and Parquet zstd adapters are not implemented because the real 256-record batch took 1.1–1.7 seconds; TileDB variable queries took 8.6–9.6 seconds and the backend was removed entirely. ArrayRecord zstd remains because its batch p95 is 92 ms.
Small Token Records
Both token workloads come from the same real corpus: WikiText-103 raw train,
tokenized with GPT-2 and stored as int32 IDs. Preparation is outside every
measurement. scripts/prepare_tokens.py preserves nonempty text boundaries,
combines fragments shorter than 16 tokens, and splits records at 512 tokens into
tokens.npy plus offsets.npy; both benchmark scripts mmap those files.
Fixed Token Blocks
The Dense workload ignores text boundaries and packs the stream into 200,000
fixed int32[512] records: 2 KiB per record and 390.6 MiB logical payload.
Each IID batch gathers 256 records; fresh draws continue until timed gathers
accumulate at least two seconds.
| backend | logical write | logical gather | krecords/s | p95 | disk | ratio |
|---|---|---|---|---|---|---|
| zrecord-raw | 5526 MiB/s | 7852 MiB/s | 4020.2 | 0.08 ms | 393.7 MiB | 0.99x |
| npy-mmap-raw | 2026 MiB/s | 10689 MiB/s | 5472.6 | 0.07 ms | 390.6 MiB | 1.00x |
| lmdb-raw | 680 MiB/s | 1241 MiB/s | 635.5 | 0.47 ms | 786.3 MiB | 0.50x |
| arrow-ipc-raw | 1183 MiB/s | 237 MiB/s | 121.5 | 2.45 ms | 390.8 MiB | 1.00x |
| arrayrecord-raw | 876 MiB/s | 168 MiB/s | 86.0 | 4.75 ms | 401.4 MiB | 0.97x |
| zrecord-zstd | 1378 MiB/s | 2413 MiB/s | 1235.4 | 0.26 ms | 193.2 MiB | 2.02x |
| zrecord-zstdict | 1823 MiB/s | 2693 MiB/s | 1378.9 | 0.24 ms | 168.5 MiB | 2.32x |
| arrayrecord-zstd | 131 MiB/s | 143 MiB/s | 73.2 | 4.27 ms | 201.3 MiB | 1.94x |
The contiguous NumPy baseline is strongest for gather when the whole corpus is one fixed typed matrix. Zrecord-raw reaches 4.02 Mrecords/s while retaining independent record semantics and writes 2.7x faster than npy-mmap-raw; dictionary zstd writes 32% faster than plain zstd, uses 13% less disk and gathers 12% faster in this pass. LMDB's B-tree/page overhead is visible in both throughput and disk.
Variable Token Sequences
The Ragged workload keeps 200,000 real text records of 16..512 tokens: p10 28,
median 129, mean 138.8, p90 254, totaling 105.9 MiB. It uses the same 256-record,
random plan and every backend must return ordered list[np.ndarray]
with exact int32 values and original one-dimensional shapes.
| backend | logical write | logical gather | krecords/s | p95 | disk | ratio |
|---|---|---|---|---|---|---|
| zrecord-raw | 2144 MiB/s | 966 MiB/s | 1826.2 | 0.18 ms | 110.5 MiB | 0.96x |
| lmdb-raw | 349 MiB/s | 115 MiB/s | 217.6 | 1.42 ms | 151.8 MiB | 0.70x |
| arrow-ipc-raw | 317 MiB/s | 44 MiB/s | 84.0 | 3.52 ms | 109.1 MiB | 0.97x |
| arrayrecord-raw | 296 MiB/s | 25 MiB/s | 47.6 | 8.52 ms | 118.0 MiB | 0.90x |
| zrecord-zstd | 459 MiB/s | 617 MiB/s | 1166.4 | 0.28 ms | 67.4 MiB | 1.57x |
| zrecord-zstdict | 827 MiB/s | 771 MiB/s | 1457.0 | 0.22 ms | 53.0 MiB | 2.00x |
| arrayrecord-zstd | 60 MiB/s | 25 MiB/s | 47.8 | 6.71 ms | 76.0 MiB | 1.39x |
Here the record contract, not bulk byte bandwidth, is the useful scale. Zrecord's three codecs return 1.17–1.83 Mrecords/s with 0.18–0.28 ms p95. Dictionary zstd writes 80% faster than plain zstd, uses 21% less disk and gathers 25% faster in this pass. Raw reaches 2144 MiB/s write after Cython compiles the validated records into a borrowed native plan.
Sampler
Index generation on its own, IID (with replacement), 1M index space, against NumPy's modern API. The µs-scale figures fluctuate with box load; this pass shows a 2.33–9.55x margin across batch sizes.
| sampler | batch | per batch | vs default_rng |
|---|---|---|---|
| numpy default_rng | 256 | 5.6 µs | 1.00x |
| zsampler | 256 | 0.6 µs | 9.55x |
| numpy default_rng | 1024 | 6.9 µs | 1.00x |
| zsampler | 1024 | 1.2 µs | 5.77x |
| numpy default_rng | 8192 | 16.8 µs | 1.00x |
| zsampler | 8192 | 7.2 µs | 2.33x |
End-to-End DataLoader
The full input pipeline comparison (sample, fetch, collate, hand over a batch)
uses exactly the Dense vision source above: 2,500 (3,224,224) uint8
Oxford-IIIT Pet records from the prepared mmap corpus, 4 high-level workers, and batch 256
(36.8 MiB). After warmup, throughput and memory are collected during one
200-batch qualitative pass. peak PSS sums proportional
set size across the process tree, apportioning mapped shared and copy-on-write
pages instead of counting each once per worker. It does not include ordinary
kernel page-cache pages used by pread, while resident mmap pages are attributed
to the mapping process, so it is a process-mapping diagnostic rather than total
pipeline physical memory. aggregate RSS deliberately sums
each process's full resident set: on Linux it double-counts shared/COW pages,
which explains process-tree RSS inflation but is neither physical memory nor a
projection of Windows committed memory. The explicit spawn row is the relevant
no-fork control; Windows itself still requires a native run. Torch fork is kept
because it is the normal Linux mode, while spawn exposes the ownership model
used on platforms without fork.
Grain setup remains optional through --only grain; the published command
selects it explicitly in the same workload matrix.
storage is the actual backing store used by each pipeline. This is an
end-to-end systems comparison, not a scheduler-only comparison over one shared
storage layer: Torch reads read-only NumPy mmap files, Loaderx reads Zrecord,
and Grain reads ArrayRecord.
| loader | model | storage | batches/s | p95 | steady PSS | peak PSS | peak RSS |
|---|---|---|---|---|---|---|---|
| loaderx | threads | zrecord-zstd | 115.5 | 13.12 ms | 641 MiB | 665 MiB | 669 MiB |
| loaderx-raw | threads | zrecord-raw | 186.2 | 10.93 ms | 667 MiB | 668 MiB | 671 MiB |
| torch | fork | npy-mmap-raw | 82.9 | 40.11 ms | 1397 MiB | 1601 MiB | 3885 MiB |
| torch-spawn | spawn | npy-mmap-raw | 111.3 | 30.71 ms | 2625 MiB | 2740 MiB | 4717 MiB |
| grain | processes | arrayrecord-zstd | 45.3 | 88.48 ms | 1701 MiB | 1757 MiB | 1879 MiB |
At 36.8 MiB per batch the per-batch gather dominates the tiny sampler cost, and
the transform threads overlap Python-side collation with the next gather. The
memory is the source, Zrecord container and bounded in-flight batches. loaderx prefetches in
threads inside one process, so workers share one interpreter, one NumPy runtime
and one set of gather buffers. With source geometry and entropy held constant,
raw is 1.61x compressed loaderx; compressed loaderx is 1.39x Torch fork,
1.04x Torch spawn and 2.55x Grain, while raw is 2.25x, 1.67x and 4.11x faster.
Torch's aggregate RSS is high because
Linux fork mappings are counted repeatedly; it is not a total-memory ratio
against Zrecord's unaccounted page cache. The explicit torch-spawn row removes
fork/COW dependence. Because this Dataset keeps only mmap paths, spawn does
not copy the full corpus into every worker; a Windows Dataset holding Python
lists or in-memory arrays would be a different, deliberately harsher workload.
The Torch-only worker sweep runs each count once. Worker 0 is an in-process baseline (68.4 and 69.6 batches/s with 1411/1414 MiB peak PSS in the two equivalent rows), so the process-context comparison starts at one worker:
| workers | fork batches/s | fork peak PSS | fork peak RSS | spawn batches/s | spawn peak PSS | spawn peak RSS |
|---|---|---|---|---|---|---|
| 1 | 45.0 | 1430 MiB | 1730 MiB | 46.8 | 1708 MiB | 1946 MiB |
| 2 | 68.5 | 1510 MiB | 2477 MiB | 73.4 | 2068 MiB | 2890 MiB |
| 4 | 100.4 | 1615 MiB | 3902 MiB | 110.0 | 2709 MiB | 4689 MiB |
| 8 | 116.6 | 1815 MiB | 6607 MiB | 122.0 | 4001 MiB | 8188 MiB |
Spawn peak PSS grows from 1708 to 4001 MiB as workers rise from one to eight, while fork grows from 1430 to 1815 MiB because it retains COW sharing. At eight workers spawn uses 2.20x fork's peak PSS and throughput has already flattened. This demonstrates no-fork memory pressure; it is not labeled OOM because this 31 GiB machine completed the run. Fork aggregate RSS grows faster because Linux counts shared/COW mappings in every process, so RSS is diagnostic rather than physical memory.
CPU-heavy transform solutions. The same Python per-sample transform exposes
the GIL bottleneck on standard CPython. Numba is an explicit solution, not the
default: --transform numba-nogil compiles the equivalent batch transform with
nogil=True. The other explicit solution runs the unchanged Python transform on
free-threaded CPython. Numba compilation is warmed before timing; both
interpreters use NumPy 2.4.6 and the same prepared real corpus.
| loader | GIL Python | GIL + Numba nogil | free-threaded Python | Numba gain | free-threaded gain |
|---|---|---|---|---|---|
| loaderx | 50.4 batches/s | 103.5 batches/s | 97.0 batches/s | 2.05x | 1.92x |
| loaderx-raw | 55.9 batches/s | 133.6 batches/s | 102.0 batches/s | 2.39x | 1.82x |
Peak PSS for compressed/raw was 876/887 MiB with GIL Python, 940/986 MiB with Numba, and 883/875 MiB with free-threaded Python. These are two deployment solutions to the transform bottleneck, not claims that Store itself was optimized for either runtime.
Conclusion — why the numbers look like this.
Every hot path is batched natively. Zsampler draws a whole batch of indices; Dense gathers and decompresses a whole fixed-shape batch in one Cython call; Ragged obtains lengths, validates cumulative ends, allocates and gathers the encoded batch in one Store call before Cython reconstructs the exact arrays. The speedup is not bought with sampling shortcuts: the IID draw is unbiased like NumPy's (Lemire with rejection, so uniformity costs nothing over a real index space).
The layouts match what a training loader does. Zrecord preserves an ordered record sequence while supporting indexed access: Dense gathers fixed-width records directly into one ndarray; Ragged restores independently shaped records from inline shape/payload entries. The benchmark deliberately stresses random gather rather than claiming that ordering is absent. Array stores are built primarily for contiguous scans, so a scattered batch fights their layout. A dense raw gather fans out across the shared Executor budget where NumPy fancy indexing is one thread; Ragged instead trades some raw specialization for compression and a complete variable-shape persistence model.
Compression is in the storage kernel, and there is one codec. zrecord-zstd is not
"storage plus a codec": the layout, the multi-core decompress and the GIL-free
copy are one path, so turning compression on costs part of a margin, not an
order of magnitude. The ratio is the data, not the
format: in the current real-photograph Dense workload, plain zstd reaches only
1.15x and the fixed dictionary reaches 1.12x. Dictionary mode is for repeated
small records, not a universal higher-ratio setting.
Loader results combine architecture and storage. loaderx uses threads and
never ends an epoch, so a step pays no IPC and never waits on an epoch boundary;
torch uses finite shuffled epochs, worker processes and shared-memory handoff.
Here compressed loaderx is 1.39x Torch fork, 1.04x Torch spawn and 2.55x
Grain; raw loaderx is 2.25x, 1.67x and 4.11x faster, respectively.
Storage also differs per loader — each reads from what it was
built for — so the loader table is a different comparison from either store
table, not a rerun.
The one crack in the thread model is a CPU-heavy Python transform, which the GIL
serializes. Numba nogil=True and free-threaded Python are measured as two
explicit solutions rather than silently changing the default transform.
What these numbers do not claim. Everything runs with a warm page cache: this
measures the access path, not cold storage or disk. disk is allocated
blocks, and zrecord files grow to their written frontier. logical write
measures source adaptation through logical finalization after writer setup; it
does not benchmark durability, stronger transactional guarantees, or reusable
dictionary training. Arrow IPC and Parquet use 256-record groups, so random
batches pay their real group-level read amplification. Ragged Zrecord arrays are
views into one batch allocation, while most byte-store adapters return
independent copies. This is a single benchmark run, while each Store read path
accumulates at least two timed seconds —
the µs-scale sampler timings and the loader batches/s fluctuate with box load
(on these 12 cores the compressed loader trails the raw one materially), so
treat the absolute numbers as ballpark and the cross-backend margins as the
signal.
Real-data verification
Production dataset-specific preprocessing and verification remain in the DataPipe repository. The built-in converter covers standardized Hugging Face datasets; DataPipe handles sources without a common remote protocol and implements derived modalities as loader transforms without NumPy dump intermediates. Oxford-IIIT Pet is only the shared benchmark fixture.
Current Limitations
- Single-host only; multi-host training is not supported.
- A single sample must be at most 2 GiB (2^31 bytes). There is no fixed record
count:
lengthis a u64 and the record table grows on demand. Practical store size is bounded by disk and the platform's positional file-offset range. - Metadata is read and written as the host's struct layout, so a store carries the host's byte order and is not portable to a machine of the opposite endianness. Every published platform is little-endian, so this only matters if you build for one yourself.
Build
python3 setup.py build_ext --inplace # three Cython extensions + static Zig libraries
zig build test # native store suite, in both Debug and ReleaseFast
python3 scripts/test_loaderx.py # Python integration suite against the real build
The Zig side is tested for behaviour only; throughput is measured from Python, through the binding a client actually uses. Optional benchmark contenders are skipped when not installed. See Benchmarks.
Publishing
The release is intentionally fixed: one Linux x86-64 machine uses Zig to
cross-compile the Store, Sampler and Ragged Cython extensions for four platforms
and both CPython ABIs. uv supplies the pinned CPython toolchains:
python3 scripts/build_release.py
This always produces and validates eight wheels: cp310-abi3 and
cp314-cp314t for glibc Linux on x86-64 and ARM64, macOS on Apple Silicon, and
Windows on AMD64. Alpine/musl, Intel macOS, and Windows ARM64 are not published:
their JAX/PyTorch ecosystem or active hardware coverage does not justify an
untested binary-support claim. Partial matrices and alternate native toolchains
are not supported.
The dist directories are named after their Python wheel platform tag, so the tag
mapping lives in exactly one place (dist_targets in build.zig). glibc and
macOS minimums are pinned in the target triple, which is what makes
manylinux_2_17 and macosx_11_0 honest rather than aspirational. Pinned
CPython and NumPy target headers are checksum-verified before zig cc; assembly
then refuses a missing ABI artifact and checks that every wheel carries exactly
the three domain extensions for that platform. Host extensions are never
retagged for a foreign platform.
Source distributions are intentionally not published. The wheel matrix covers the platforms we build and test; avoiding an automatic source-build fallback makes unsupported environments fail clearly. If your platform is not covered, please open an issue. Developers working from a source checkout can still use the local build commands above, which require Zig.
Zsampler
Index Generator: a high-performance sampler implemented in Zig. Every mode is a
pure function of (seed, step), so a run resumes exactly by seeking to a step —
there is no epoch to track, in keeping with the endless step-based loader.
from loaderx.zsampler import Sampler
sampler = Sampler(1_000_000, 256, Sampler.Mode.IID, seed=42)
indices = sampler.next() # borrowed until this sampler's next draw
saved = indices.copy() # retain across draws only when needed
next() and iteration return a view of one reusable uint64 batch buffer.
The contents stay unchanged until the next explicit draw from that Sampler;
copy only plans that must outlive it. DataLoader consumes each view synchronously
before drawing again.
- Sequential — traverse the index space in order through a fixed-size sliding window, treating the space as a circular queue so the tail never truncates.
- IID — draw each index uniformly at random with replacement. Unbiased (Lemire with rejection), matching NumPy. Simplest, but coverage is uneven over any short run.
- Cyclic — without replacement, round-robin. Each cycle traverses a fresh
permutation of the whole index space, so within a cycle every record appears
exactly once and no batch repeats an index — coverage is even by construction,
which keeps how often each sample is seen uniform. The permutation is a
stateless bijection (a small Feistel network over the index space, brought
into range by cycle-walking), so a million-record shuffle materializes nothing
the size of the dataset and reshuffling each cycle is free. Every batch is
exactly
batch_size: an endless step-based loader has no final partial batch to special-case, so whenbatch_sizedoes not divide the length the cycle's remainder is dropped — a different remainder each cycle, since the permutation changes, so every record is still reached over time.
Zrecord
Zrecord is loaderx's rebuildable typed record container. Its private native
runtime is the byte-oriented engine beneath the public Dense and Ragged
contracts.
Three private Cython extensions preserve the native domain boundaries:
loaderx/_store.pyx owns Store and Executor handles, typed NumPy buffers and
native storage errors, and compiles validated append geometry into byte-bounded
shard chunk commands; loaderx/_sampler.pyx owns the independent Sampler
handle; loaderx/_ragged.pyx owns Ragged shape planning and ndarray view
materialization. Store and Sampler each statically link their corresponding Zig
library, while Ragged has no Zig dependency. There is no runtime CFFI or
separately loaded native library. Python owns all schema semantics; Zig stores
the opaque schema bytes in meta.zr. zrecord.py remains the unified
Python-facing record API.
Trust and usage boundary. Python and Zig are one zrecord implementation, not two independently supported products. The three Cython modules, the Store and Sampler C ABIs, and the native handles are private implementation details; no defensive-validation contract is provided to code that calls them directly. The same applies to fabricated public objects, mutation of private Python state, use-after-close, callback-initiated lifecycle changes, and concurrent close. Public inputs are normalized and validated once, in whichever half can express the rule most simply, and the internal Python→Cython→Zig call then trusts that contract instead of repeating it at every layer. Python owns the creator/reader capability model and chooses append, gather, creator Header sync, or reader handle release; the Zig engine stores no writable/reader mode. Native create/open still select read-write/exclusive or read-only/shared file handles, because those are OS access and locking mechanics rather than API capabilities. This is not a license to trust storage or the operating system: native code still validates persisted addresses and lengths, buffer bounds and integer overflow, short or failed I/O, codec output, locking, and commit ordering. Those checks protect normal I/O behavior, basic malformed-store rejection and native memory safety; checks that only defend against bypassing the public Python API do not belong in zrecord.
RecordEnginestores N logically ordered records. Payload recordibelongs todata_{i % data_shards}.zr; logical ID remains the stable append position andRecordLoc[ID]preserves its shard-local offset. Index and slice operations are implemented as ordered gathers over those positions.- It hands the container layer a dense sequence space: records are exactly
0..N-1. Named streams are composed dynamically by a plain Python dict;DataLoadervalidates that the independent containers have equal lengths. - The engine reads and writes byte ranges. Python owns the record schema and
selects Dense or Ragged geometry; Cython normalizes append input to a fixed
byte buffer or trusted Ragged header/payload spans,
then supplies complete window and chunk plans.
Zig dynamically executes those commands and owns compression, shard-local
allocation, I/O,
RecordLocconstruction and publication. Dense stores persist one fixed-width physical record per logical record. Ragged stores persist one variable-width physical record per logical record:[u64le dim] * schema.ndim + [payload]. Shape and payload therefore share one location, codec frame, and append publication. - The IO model (
append | read) is batch-oriented and shape-agnostic. Cython's per-call command plan supplies record boundaries and source windows; the engine has one append operation and carries no Dense, Ragged, dtype, or array-shape semantics. A single-record operation is just thebatch_size == 1case. - The engine owns its temporary memory internally — allocation and release are explicit.
- A store has exactly one codec in its native physical header, fixed at creation and immutable afterwards. Every record is compressed and decompressed independently:
| tag | name | algorithm |
|-------|-----------|----------------------------------------|
| 0 | raw | none |
| 1 | zstd | zstd (plain, level 3) |
| 2 | zstdict | zstd with a trained dictionary (level 3) |
- Compression is transparent to the client:
- Compression runs concurrently across the shared Executor budget. A compressed store never
falls back to raw: each record is stored as the codec's output, even when
an incompressible record's frame is larger than its input — write
rawif the data does not compress. - Decompression writes straight into the caller's destination memory
(
gather), with no intermediate buffer and no extra copy. - zstd is the one transparent codec — faster than Deflate at both ends and a better ratio, so there is no reason to carry a second. It is vendored C, built for every platform by Zig, so it does not expand the platform/ABI wheel matrix.
zstd_dictadditionally trains one dictionary on a sample of the data (stored asdict.zr) and compresses every record against it. Because each record is still independent, random access is unchanged — but the dictionary carries the structure shared across records, which per-record compression cannot see. On many small, similar records (image tiles, token sequences) this can be a large win. Decoded real photographs are a useful counterexample where plain zstd or raw is the better fit. The 32 KiB dictionary is loaded once on open and shared, lock-free, across all reader threads.- A
zstd_dictstore needs its dictionary to read every record; araworzstdstore rejects an unexpected dictionary as malformed state.
- Compression runs concurrently across the shared Executor budget. A compressed store never
falls back to raw: each record is stored as the codec's output, even when
an incompressible record's frame is larger than its input — write
Persistence format
The current format is the settled internal baseline for implementation work: optimizations keep
the fixed files, Header/schema page, contiguous RecordLoc table and independent
record payloads unless the product boundary is deliberately reopened. "Settled"
does not promise cross-version persistence compatibility: there is no compatibility
layer, migration, version dispatch, checksum, or recovery facility. Zrecord is not
the authority for irreplaceable data. Keep authoritative source data and reproducible
build scripts; after an interrupted build, storage failure, incompatible implementation
change, or content change, rebuild a complete container at a new path.
Native storage uses one metadata file and a create-time-fixed payload file set:
store/
├── meta.zr 4096-byte Header/schema/tails page + RecordLoc table
├── data_0.zr payload records where ID % data_shards == 0
├── ...
├── data_{N-1}.zr final static payload shard
└── dict.zr zstd dictionary (only in dict stores)
data_shards is a write-performance parameter in 1..255, fixed by create
and recovered automatically by open. The default is four; practical values are
usually 2, 4, 8 or 16, near the writer lane count. More files spread positional
writes across payload inodes but consume one descriptor each. This physical
striping does not change record IDs, order, codec, or read results.
Metadata (meta.zr)
Files are read and written positionally — pread/pwrite at computed offsets,
no mmap. meta.zr starts with one fixed 4096-byte static page: a naturally
aligned 16-byte Header, 255 shard-local u64 tails at bytes 16..2055, then up to
2040 bytes of opaque MsgPack schema at bytes 2056..4095.
An array of 16-byte RecordLocs starts at offset 4096. Record i is one
pread/pwrite at 4096 + i * 16; there is no variable table base, segment
mapping, or rollover fd table.
1. Python schema — bytes 2056..2056+schema_length are exactly one immutable
MsgPack object. Both stores contain dtype and ndim; Dense additionally
contains item_shape. Native create persists these bytes together with
the physical container but does not decode them. Open acquires the native lifetime
lock before copying the schema to Python for validation, so schema and physical
metadata are one locked snapshot. Dense record width is derived once from
dtype/item_shape and passed to the native handle as runtime geometry; it is not
independently persisted as a second authority. There is no format version or
legacy kind dispatch.
2. Physical header — the first 16 bytes of meta.zr. The format
deliberately carries no payload or metadata checksum.
codecis the store's one compression method, stamped at creation and immutable — there is no per-record tag anywhere.length(u64) is the physical record count; it equals logical length for both dense stores and inline ragged stores.schema_length(u16) is the occupied prefix of the static schema area and must be in1..2040.data_shards(u8) is the static payload file count and must be in1..255.
const Codec = enum(u8) { raw = 0, zstd = 1, zstdict = 2, _ };
const Header = extern struct {
length: u64,
schema_length: u16,
data_shards: u8,
codec: u8,
reserved: [4]u8,
};
3. Shard frontiers — tail slot s at 16 + s * 8 is the committed
end of data_s.zr. Unused slots among the 255 fixed u64 entries are zero.
Open requires every data file to be at least its persisted tail; locations may
not cross that shard-local frontier.
4. Record table — contiguous 16-byte entries start at offset 4096 in
meta.zr and grow as location windows are written. offset is local to
data_{ID % data_shards}.zr;
phys_length/logic_length are the stored and original sizes. The
codec is not here: it is the header's, so a record is stored exactly the way the
store is declared.
const RecordLoc = extern struct {
offset: u64,
phys_length: u32,
logic_length: u32,
};
There is no liveness flag. Every entry below length is a record.
5. No fixed record-count cap. The table and payload streams grow naturally. The practical bounds are the u64 count, supported positional file offsets, 2 GiB per record, descriptor budget, and disk.
Executor
1. Write. Writes are append-only; everything else is offset redirection.
The codec is immutable store state, so append and gather dispatch once at
their entry points into raw, zstd, or zstd-dictionary execution. Window and
chunk scheduling are shared; codec contexts and record execution remain
compile-time-specialized.
Geometry is equally explicit across the whole stack: Python derives Dense
record width from its schema, while Cython turns fixed stride or validated
Ragged spans into explicit append windows and shard-strided
chunks. The engine has one append and one gather operation. Its shared opaque
handle remains private and carries no typed-store geometry.
- Append validates the complete call before physical I/O, then processes bounded
Cython-planned logical windows. Within a window, record positions are divided by
ID % data_shardsand planned as bounded shard-local chunks. Fixed Executor lanes dynamically claim those chunks, encode or pack them, briefly lock only the selected shard's tail reservation, then issue positional payload writes directly. Multiple lanes may write non-overlapping ranges of one shard; the static file set spreads that pressure across inodes without limiting codec concurrency to the shard count. Workers fill disjoint entries in one contiguous location buffer. A payload barrier precedes one contiguousmeta.zrlocation write. All windows must succeed before the in-process length and per-shard tails advance. Writerclose()truncates each data file to its committed tail and publishes the 4096-byte static page. - Compressed shard tasks lease process-bounded
ExecutionSlotscratch and write independent frames in bounded subchunks. Python budgets the process-wide executor at three quarters of the logical CPUs available to the process, leaving headroom for packing, transforms, and the caller without encoding a platform-specific thread count. Each producer configures its CCtx or shared immutable CDict once, then starts every independent record frame withZSTD_compress2. - Raw records for one shard are strided in the source. Each shard task packs a bounded subchunk into its reusable scratch, materializing trusted Ragged header/payload spans when present, and performs one contiguous positional write. A contiguous single record larger than the normal subchunk budget is written directly; borrowed Ragged spans are materialized once to join their header and payload. There is no per-record syscall or platform-specific vectored path. The extra memory copy is the deliberate cost paid to remove concentrated single-inode writes. 2. Read. Fill the destination memory concurrently, in place from the Python side (executed on async threads).
- Committed records are immutable and
lengthis published through an atomic. The fixed metadata and static data handles require no rollover synchronization. - Every record first selects
data_{ID % data_shards}.zr, then reads the shard-local offset its table entry records — the record table is addressed by pure arithmetic, so random access is one pread for the location and one for the bytes, with no batching assumptions about layout. Each lane reads one location and immediately reads/decompresses that record; there is no separate metadata phase or sequential-run special case. Compressed records use a per-lane staging buffer and decode in place into the destination.
3. Internal fan-out. Io.Group.async fans work out up to the executor lane
budget; lanes for which the runtime cannot reserve concurrency run inline on the
calling thread. Python configures the process-level budget as
max(physical cores, logical cores * 3 / 4), using platform topology where
available.
- Gather lanes receive contiguous request blocks. Append lanes dynamically claim bounded chunks of strided shard ranges; per-shard atomic reservation cursors keep offsets disjoint while positional writes remain concurrent.
- Each lane creates one zstd context (
ZSTD_CCtxto write,ZSTD_DCtxto read) and reuses it across every record it handles, rather than paying that setup per record. The dictionary (ZSTD_CDict/ZSTD_DDict) is immutable, so all lanes share one, lock-free. - Decompression writes straight into the caller's destination buffer, so there is no intermediate copy.
4. File access.
- Metadata: one naturally growing
meta.zr, containing the fixed Header/schema /tails page and loc table. It is intentionally not sharded because measured write pressure is in payload I/O; one coordinator writes each loc window. - Payload: a static list of naturally growing
data_<shard>.zrfiles, accessed throughreadPositionalAll/writePositionalAll. No path depends on sparse files.
Execution model. Opened readers are immutable, so calls on the same reader may
gather concurrently. Creator appends are synchronous and native Storage
serialization protects their physical commit. close() requires a quiescent
handle; it is not concurrent with append or gather and is not called from an
input or transform callback. Each native append or gather
fans out internally across the shared Executor. A creator holds a lifetime, nonblocking
exclusive lock on meta.zr; opened readers hold shared locks, so multiple
handles and processes may consume one completed container concurrently.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distributions
Built Distributions
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file loaderx-2.4.7-cp314-cp314t-win_amd64.whl.
File metadata
- Download URL: loaderx-2.4.7-cp314-cp314t-win_amd64.whl
- Upload date:
- Size: 933.1 kB
- Tags: CPython 3.14t, Windows x86-64
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.14.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
40c2fc4806577f292fa8d9b3a419e853b10d8e8470538091dda60079f7dea179
|
|
| MD5 |
0b53d6b668cda1c47c19f9aba6ad851d
|
|
| BLAKE2b-256 |
b7d8472e44b4686ad93899981626259dffcbdd75e83796b57dd48eac4133312f
|
File details
Details for the file loaderx-2.4.7-cp314-cp314t-manylinux_2_17_x86_64.whl.
File metadata
- Download URL: loaderx-2.4.7-cp314-cp314t-manylinux_2_17_x86_64.whl
- Upload date:
- Size: 2.8 MB
- Tags: CPython 3.14t, manylinux: glibc 2.17+ x86-64
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.14.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
9aff7d790ab61fc645c0fee5214a9df5baf3d5a2305d75778671d1e2fc38c171
|
|
| MD5 |
d8245cf69f82d3987274839a3bb79f34
|
|
| BLAKE2b-256 |
212f848354138a4f1cd342f43fe5b17b6460dcb20c31070d3801aa3fc91c37cf
|
File details
Details for the file loaderx-2.4.7-cp314-cp314t-manylinux_2_17_aarch64.whl.
File metadata
- Download URL: loaderx-2.4.7-cp314-cp314t-manylinux_2_17_aarch64.whl
- Upload date:
- Size: 2.6 MB
- Tags: CPython 3.14t, manylinux: glibc 2.17+ ARM64
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.14.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
dc86135f2b738cfb65d5352fd83a6bb47e1e88535d0069ce45ea4fba62960d9f
|
|
| MD5 |
b51344563ec6357843c61e0b301fc164
|
|
| BLAKE2b-256 |
a1157632e447bb46000748d95b52022a8a83f94c5753a4926f03bc85c2893d82
|
File details
Details for the file loaderx-2.4.7-cp314-cp314t-macosx_11_0_arm64.whl.
File metadata
- Download URL: loaderx-2.4.7-cp314-cp314t-macosx_11_0_arm64.whl
- Upload date:
- Size: 573.1 kB
- Tags: CPython 3.14t, macOS 11.0+ ARM64
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.14.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
e6333811554056599a86f4ffa0be8629554388ab285c19d9ead03efd096887af
|
|
| MD5 |
ccdc8b9de84f98ce7d62272875f033c7
|
|
| BLAKE2b-256 |
498566e3002507e86159fa9045b87f45bead3eb15ecd68d0acb9203a0a4b30ce
|
File details
Details for the file loaderx-2.4.7-cp310-abi3-win_amd64.whl.
File metadata
- Download URL: loaderx-2.4.7-cp310-abi3-win_amd64.whl
- Upload date:
- Size: 923.5 kB
- Tags: CPython 3.10+, Windows x86-64
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.14.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
e491bd83f1d02596394ecc45e9ac79f97744245792f208d7561e8260ab5c6a5e
|
|
| MD5 |
e6728d2413a705a23cfe55c86951b858
|
|
| BLAKE2b-256 |
a3c5fdc3eab73c457d2ffb93d9964322ce75026e32d4211c874e4f133609e91d
|
File details
Details for the file loaderx-2.4.7-cp310-abi3-manylinux_2_17_x86_64.whl.
File metadata
- Download URL: loaderx-2.4.7-cp310-abi3-manylinux_2_17_x86_64.whl
- Upload date:
- Size: 2.7 MB
- Tags: CPython 3.10+, manylinux: glibc 2.17+ x86-64
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.14.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
45903c259026fdf1d03248ce89dfcea04154698a85d8f022c63b6c0088decbee
|
|
| MD5 |
e88c46c800823211a93849c03a63bb29
|
|
| BLAKE2b-256 |
2bf46656fec554314b8a1f96c51f11204c2c5c4d69d25dd728f7d552875e9e34
|
File details
Details for the file loaderx-2.4.7-cp310-abi3-manylinux_2_17_aarch64.whl.
File metadata
- Download URL: loaderx-2.4.7-cp310-abi3-manylinux_2_17_aarch64.whl
- Upload date:
- Size: 2.5 MB
- Tags: CPython 3.10+, manylinux: glibc 2.17+ ARM64
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.14.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
f6ef633bfd5dc5cfd5b7cde375af4c7cdd0ffb3fed26f0d2d9d3771da2c2f339
|
|
| MD5 |
a613cadf97987775f2ecd82789ee1bde
|
|
| BLAKE2b-256 |
fbaaa66976fafa92570302225c535ab7e767466c55c01aaaadde1969059094fd
|
File details
Details for the file loaderx-2.4.7-cp310-abi3-macosx_11_0_arm64.whl.
File metadata
- Download URL: loaderx-2.4.7-cp310-abi3-macosx_11_0_arm64.whl
- Upload date:
- Size: 562.5 kB
- Tags: CPython 3.10+, macOS 11.0+ ARM64
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.14.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
88faedc276e27a777ef27ee8561861b579e997dc2a602047f94fde085db3c83d
|
|
| MD5 |
03e2f7bdf280e3951421c1a7cd97a3b7
|
|
| BLAKE2b-256 |
60a22511f64ac88507594fa25339d4e0fd4040b08048f905dd5b0b47a7f5da5f
|