Loaderx
A compact, high-performance persistent record store with zero-copy batch gathering and transparent per-record compression, designed for single-machine AI training and serving pipelines
Zrecord is the typed persistent store; Loaderx is the sampler and prefetch loader that consumes Zrecord streams. They currently ship together while both layers mature, but their public responsibilities remain separate.
pip install loaderx
Wheels are published for Linux (glibc ≥ 2.17 and musl, x86-64 and arm64), macOS (≥ 11.0, Intel and Apple Silicon) and Windows (x64 and arm64). The bindings use cffi in ABI mode, so nothing links against the CPython ABI and one wheel per platform serves every supported Python.
Design Philosophy
loaderx is built around several core principles:
- A pragmatic approach that prioritizes minimal memory overhead and minimal dependencies.
- A strong focus on single-machine training workflows.
- We implement based on NumPy semantics, persisted by the private native store engine.
- An immortal (endless) step-based data loader, rather than the traditional epoch-based design—better aligned with modern ML training practices.
- Dense and ragged are separate contracts, and the loader serves both. A dense stream stacks into one array per batch; a ragged one comes back as a list. Neither is padded, and equal length is never treated as a special case of variable length.
Quick Start
from loaderx.zrecord import DenseStore
from loaderx.dataloader import DataLoader
data = np.load('data.npy', mmap_mode='r')
label = np.load('label.npy', mmap_mode='r')
with DenseStore('train_data', data.dtype, data.shape[1:]) as ds:
ds.append(data)
with DenseStore('train_label', label.dtype, label.shape[1:]) as ds:
ds.append(label)
loader = DataLoader({'data': DenseStore('train_data'), 'label': DenseStore('train_label')},
transform=lambda batch: batch)
for i, batch in enumerate(loader):
if i >= 256:
break
print(batch['data'].shape)
print(batch['label'].shape)
A batch is a dict {name: values}: each value is the stacked
(batch_size, *item_shape) array for that stream. Every stream is gathered
at the same indices, so record i lines up across them. The transform
callback is the collate step — reshape, cast, stack — where values is the
plain dense batch ready for the model.
Creating a dense store
import numpy as np
from loaderx.zrecord import DenseStore
data = np.load('data.npy', mmap_mode='r')
with DenseStore('train_data', data.dtype, data.shape[1:]) as ds:
ds.append(data)
One record per slice along axis 0; a 1-D array (the usual shape of a label set)
becomes a store of scalar records. The caller controls each append batch and can
slice a large or mmapped array to set its own memory bound. Open the result with
DenseStore:
ds = DenseStore('train_data')
The native owner persists an exact Python type identity: dtype plus the
dense item_shape, or only dtype for ragged stores. It accepts no user
metadata. Ragged shapes are native per-record records, not part of that static
identity. Distinct native DenseStore and
RaggedStore owners embed a shared physical record engine; that engine is an
implementation detail. Raw bytes and pre-encoded images use a
RaggedStore(path, dtype=np.uint8) rather than a second public storage API.
Records
One persistent format, two native execution contracts. DenseStore is the dense contract
where every record is exactly one row of the recorded item_shape; reads
are fixed-stride gathers and the batch shape follows from the type identity, so no
per-record metadata is touched:
import numpy as np
from loaderx.zrecord import DenseStore
data = np.arange(64, dtype=np.float32).reshape(8, 2, 4)
with DenseStore('data', data.dtype, data.shape[1:]) as ds:
ds.append(data)
ds = DenseStore('data')
ds[0, 5, 2] # (3, 2, 4) — shape from the persisted identity
A store is a collection of records, not an ndarray, so ds[0, 5, 2] is a
record set — never ds[0][5][2]. A scalar selects one record; negatives wrap
from the end. The ragged example below reads the same way.
RaggedStore is the ragged contract for variable-length records. It is a
separate contract: :class:RaggedStore hands back a list of arrays, so a
loader never has to carry row_splits around. dtype is unified and
explicit; each record keeps its own shape, recorded per record as it is
written and restored exactly on read — so records may differ in shape
arbitrarily, and nothing is ever inferred from the source (an iterator can't
tell you what its later records look like). Scalar shape == () is preserved;
zero-byte arrays are rejected because physical records are nonempty. Densifying a list into a dense
batch is the model's call — a plain numpy loop, wherever you need it:
from loaderx.zrecord import RaggedStore
seqs = [np.arange(L, dtype=np.int32) for L in (3, 1, 4, 1, 5)]
with RaggedStore('tokens', np.int32) as rs:
rs.append(seqs) # dtype explicit; each record keeps its shape
rs = RaggedStore('tokens')
records = rs[0, 2, 4] # list of ndarray — one per record, exact shapes
lengths = np.array([len(r) for r in records])
padded = np.zeros((len(records), lengths.max()), dtype=records[0].dtype)
for i, r in enumerate(records):
padded[i, :len(r)] = r # (B, max_len) — your policy, your loop
A DataLoader dynamically composes a dict of dense and ragged streams.
Collation is the transform — a batch dict in, a batch dict out:
def collate(batch):
return {'input_ids': batch['tokens'], 'label': batch['label']}
loader = DataLoader({'tokens': dense_tokens, 'label': labelset}, batch_size=32,
transform=collate)
Writable stores
The Store classes are loaderx's public persistence API: they are
read-write, and append is explicit — one container of records, one batch,
one native call. Nothing is buffered, inferred, or compressed for you.
from loaderx.zrecord import DenseStore, RaggedStore
ds = DenseStore('mnist/x', dtype=np.uint8, item_shape=(28, 28)) # type identity committed now
ds.append(images[i:i + 1024]) # (B, 28, 28) → one batch, returns first index
ds.append(single_image[None]) # one sample is batch_size 1 — add the axis yourself
ds[:4] # read path is unchanged
tok = RaggedStore('tokens', dtype=np.int32) # each record keeps its own shape
tok.append([seq_a, seq_b, seq_c]) # list in, list out — `tok[0, 2]` returns a list
The exact type identity is declared at construction, encoded by Python as
msgpack, and persisted opaquely by the native owner. Dense identity contains
only dtype and item_shape; ragged identity contains only dtype.
Unexpected fields are rejected. append validates dtype and shape in Python,
while native DenseStore independently enforces the persisted byte width.
Dictionary training stays explicit and separate; creating a
zstd_dict store requires the completed dictionary, so no valid store is
published without one.
The store's maintenance pass-throughs are on the same object, normalized like reads (scalar, slice, array, bool mask):
ds.delete(np.arange(0, len(ds), 2)) # swap-last: index space stays dense, survivors reindex
ds.stats() # {'records', 'live_bytes', 'chunk_bytes', 'reclaimable'}
ds.compact() # reclaim the deleted bytes, in place — offline only
Deletion and compaction move records, so a multi-stream index you keep yourself
must tolerate it — the guarantee is only that the live records are exactly
0..len(ds).
Codec notes
"zstd" compresses each record independently with plain zstd (level 3). Use it
for general-purpose compression — it is fast and the default.
"zstd_dict" trains a shared dictionary on a sample of the data before writing
any record, then compresses every record against it at level 19. The dictionary
captures structure shared across records that per-record compression cannot see —
a large win for many small, similar records (image tiles, token sequences).
The dictionary's cost is a cache footprint: every record's decompression
references the shared dictionary window, so a larger dictionary means more
cache misses on gather — the path a loader pays forever. Three named tiers
(loaderx.zrecord.DICT_TIERS) preset the whole tradeoff — the dictionary size
and how much data trains it — so a caller picks a tier, never a number:
| tier | dict | sample | tradeoff |
|---|---|---|---|
"fast" |
32 KiB | 4 MiB | fastest gather and training; ratio barely above zstd |
"balanced" |
128 KiB | 16 MiB | default — most of the ratio at a fraction of the gather cost |
"max" |
1 MiB | 64 MiB | best ratio; slowest gather and training |
The sample is the byte budget the dictionary trains on (a strided subset of the
records), so each tier costs the same training time whatever the record size.
At realistic image scale (768 KiB records) the tiers converge — on the earlier
measurement box's structured data "balanced" and "max" both gather
~1.6 GiB/s at a 1.74x ratio — because a dictionary is a small fraction of a
large frame. The tiers still matter at small record sizes, where the dict is
most of a record and "max" trades gather throughput for ratio.
"zstd_dict" records can only be read from a store that has the dictionary
(dict.zr). The dictionary is loaded on open and shared, lock-free, across all
reader threads.
A dictionary must train on the settled, complete data. :func:train_dict is the
standalone, manual training step — a numpy array, an iterable of records, or
a typed store all train the same way, sized by a tier — and its bytes are handed to
a store-writing path via dict_bytes. zstd_dict never trains by itself: a
write without a dictionary is an error. A stream cannot train its own
dictionary, but it can write with one trained on the settled data:
from loaderx.utils import train_dict
from loaderx.zrecord import DenseStore, RaggedStore
# train once on the settled data — a standalone, reusable artifact
d = train_dict(settled_array, tier="balanced")
# then any new store can install it and append explicitly
with RaggedStore('tokens', np.int32, codec='zstd_dict', dict_bytes=d) as ds:
ds.append(token_generator)
with DenseStore('data', data.dtype, data.shape[1:],
codec='zstd_dict', dict_bytes=d) as ds:
ds.append(data)
Changing an existing store's codec is a rewrite, not a store mutation: create a
new store with the new codec and copy the records in index order. Index order is
what keeps multi-stream alignment, and the copy loop is a few lines with the
Store layer — which is why no dedicated recode path exists:
from loaderx.utils import train_dict
from loaderx.zrecord import DenseStore
CHUNK = 1 << 16
with DenseStore("src") as s, \
DenseStore("dst", dtype=s.dtype, item_shape=s.item_shape,
codec="zstd_dict", dict_bytes=train_dict(s)) as d:
for start in range(0, len(s), CHUNK):
d.append(s[start:start + CHUNK]) # records copied in index order
RaggedStore is the same shape — swap the constructors and dtype is all
the destination needs. dst must not already hold a store.
Important: Do not use "zstd_dict" while data is still changing (through
deletes or compaction below the Python layer). The dictionary captures a snapshot
of the data; training it before the data settles wastes compression. Train the
dictionary once preprocessing is complete and the content is final.
Multi-stream stores
A zrecord store is one stream; a training sample is usually several named
streams (skeleton + label + id, tokens + label, ...). Composition is a plain
Python dict passed to DataLoader. There is no persistent wrapper,
manifest, directory convention or bundle mutation API. DataLoader verifies
that all streams have the same length, then gathers every stream with the same
indices.
from loaderx.zrecord import DenseStore, RaggedStore
from loaderx.dataloader import DataLoader
root = "xsub/train"
with DenseStore(root + "/joint", joint.dtype, joint.shape[1:]) as s:
s.append(joint)
with DenseStore(root + "/label", label.dtype, label.shape[1:]) as s:
s.append(label)
with RaggedStore(root + "/token", np.int32) as s:
s.append(seqs)
streams = {
"joint": DenseStore(root + "/joint"),
"label": DenseStore(root + "/label"),
"token": RaggedStore(root + "/token"),
}
streams["joint"][0, 5, 2] # each stream keeps its own index API
loader = DataLoader(streams, batch_size=256)
batch = next(loader) # {name: values}, index-aligned
Equal-length and variable-length are both just stores — DenseStore (one
fixed-shape record per sample) and RaggedStore (variable row count per
record). The loader does not interpret either: it fetches by index and packs a
dict, so a dense stream's batch value is the stacked (B, *item_shape) array
and a ragged stream's is a list of per-record arrays. No padding is imposed —
densify to a fixed shape however the model needs (a plain numpy loop), or
reshape/stack in the transform collate.
CPU → GPU transfer
loaderx hands over CPU batches; getting them to the accelerator is the
transform's job — the one place your framework is already imported. The batch
dict is a plain {name: numpy array}, zero-copy on the way out, so a device
transfer is one call per stream:
import torch
device = "cuda:0"
def to_device(batch):
return {k: torch.from_numpy(v).to(device, non_blocking=True)
for k, v in batch.items()}
loader = DataLoader(streams, transform=to_device)
for batch in loader:
model(batch) # already on device
A non_blocking=True copy is genuinely asynchronous only when its source is
pinned. loaderx does not pin memory for you — pinning is framework-owned
(torch's .pin_memory(), CUDA's cudaHostAlloc), and a vendor-free core stops
exactly at the CPU batch. Pin in the transform what you copy:
def to_device(batch):
return {k: torch.from_numpy(v).pin_memory().to(device, non_blocking=True)
for k, v in batch.items()}
JAX is the same shape — jax.device_put is already an asynchronous handoff on
GPU:
import jax
def to_device(batch):
return {k: jax.device_put(v) for k, v in batch.items()}
The transfer runs on the transform stage and never touches loaderx internals:
the copy overlaps the next batch's gather/transform, and any pinned pool is the
caller's to own and reuse. This is the entire H2D answer — there is no pin=
hook or device backend, because the only unified thing a multi-framework loader
can own is the CPU batch.
For practical integration examples, please refer to the Data2Latent repository
Benchmarks
Dense and Ragged are measured separately because they expose different contracts
and have different useful metrics. scripts/bench_dense.py measures fixed-shape
random gather, scripts/bench_ragged.py measures variable-length records, and
scripts/bench.py covers machine, sampler, and the original dense loader
comparison. Every path runs through the public Python binding, so CFFI, NumPy
allocation, and Ragged list/shape reconstruction are inside the timings.
Methodology
These 1.7.2 store results are one qualitative pass on a warm page cache. Each
backend receives identical source records and random index plans. Before timing,
the benchmark sweeps every record and validates every planned result for exact
dtype, shape, order, and values. write is page-cache ingestion with durability
deferred consistently across backends; call sync() explicitly when measuring
Zrecord durability. Disk size is allocated blocks, not sparse apparent size.
Machine — one local workstation (AMD Ryzen AI 9 HX PRO 370, 12 cores / 24 threads):
| machine | value |
|---|---|
| CPU | AMD Ryzen AI 9 HX PRO 370 w/ Radeon 890M, 1 socket, 12 cores / 24 threads |
| frequency | 605–5158 MHz |
| caches | L1d 576 KiB, L1i 384 KiB, L2 12 MiB, L3 24 MiB |
| NUMA | 1 node |
| memory | 31 GiB (not limited by cgroup) |
| shared memory | 16 GiB /dev/shm |
| OS | Debian GNU/Linux forky/sid, kernel 7.1.3+deb13-amd64, x86_64 |
| python | 3.14.7, numpy 2.5.2 |
The benchmark process sees all 24 threads and is not memory-limited by cgroup. The 16 GiB shared-memory mount accommodates the four-worker, 64 MiB-batch torch pipeline. Store reads run on the ordinary page cache.
Sampler
Index generation on its own, IID (with replacement), 1M index space, against NumPy's modern API. The µs-scale figures fluctuate with box load; the stable signal is the ~1.9x margin, roughly flat across batch sizes.
| sampler | batch | per batch | vs default_rng |
|---|---|---|---|
| numpy default_rng | 256 | 3.8 µs | 1.00x |
| zsampler | 256 | 2.0 µs | 1.92x |
| numpy default_rng | 1024 | 5.1 µs | 1.00x |
| zsampler | 1024 | 2.8 µs | 1.84x |
| numpy default_rng | 8192 | 20.4 µs | 1.00x |
| zsampler | 8192 | 10.0 µs | 2.03x |
DenseStore: Fixed-Shape Records
Zrecord against array-store alternatives: random batch gather, 2500
records, batch 256, structured (mildly compressible) data. A single 256 KiB
sample (512×512 single-channel, a realistic image record) is the horizontal
comparison; other record sizes are a --shape away.
Realistic records — 256 KiB per record, 64 MiB per batch:
| store | write | gather | on disk | ratio |
|---|---|---|---|---|
| zrecord-zstd | 911 MiB/s | 3987 MiB/s | 555 MiB | 1.13x |
| zrecord-zstdict | 39 MiB/s | 1531 MiB/s | 388 MiB | 1.61x |
| zrecord-raw | 2531 MiB/s | 7894 MiB/s | 625 MiB | 1.00x |
| npy-mmap | 2403 MiB/s | 3418 MiB/s | 625 MiB | 1.00x |
| hdf5 | 2517 MiB/s | 1871 MiB/s | 625 MiB | 1.00x |
| blosc2 | 63 MiB/s | 543 MiB/s | 536 MiB | 1.17x |
| tensorstore | 345 MiB/s | 268 MiB/s | 462 MiB | 1.35x |
At 256 KiB per record, a 64 MiB batch exceeds L3 and the raw paths are primarily DRAM-bandwidth-bound. Zrecord-raw reaches 7.7 GiB/s and is 2.3x npy-mmap here; plain zstd gathers at 3.9 GiB/s. DenseStore demonstrates that typed record ownership and per-record compression do not turn fixed tensors into an object- store slow path.
RaggedStore: Variable-Length Records
This workload contains 100,000 int32 token sequences with lognormal lengths
clipped to 16..2048 (observed median 253), totaling 142.7 MiB of logical payload.
Each of 100 random batches contains 256 records. Every backend must return an
ordered list[np.ndarray]; returning only flat values and offsets is not enough.
| backend | write | gather | krecords/s | p95 | disk | ratio |
|---|---|---|---|---|---|---|
| zrecord-zstd | 127 MiB/s | 317 MiB/s | 221.0 | 1.48 ms | 28.1 MiB | 5.08x |
| zrecord-zstdict | 23 MiB/s | 314 MiB/s | 219.0 | 1.48 ms | 17.1 MiB | 8.32x |
| zrecord-raw | 499 MiB/s | 313 MiB/s | 218.2 | 1.72 ms | 147.2 MiB | 0.97x |
| packed-offsets | 1507 MiB/s | 976 MiB/s | 681.3 | 0.53 ms | 143.4 MiB | 0.99x |
| hdf5-vlen | 36 MiB/s | 27 MiB/s | 18.9 | 20.01 ms | 146.4 MiB | 0.97x |
packed-offsets is the useful raw lower bound: a specialized payload file plus
one offsets array for one-dimensional tokens. It is not a feature-equivalent
store: it has no compression, arbitrary-rank shape records, native typed owner,
paired append publication, swap-delete, or compact. Zrecord pays for those
semantics, while still gathering about 11.6x as many records/s as HDF5 vlen. The
compressed paths gather at essentially raw Zrecord speed for this small-record
workload; zstdict reduces 142.7 MiB to 17.1 MiB (8.32x) at 1.48 ms p95. Disk
includes shape records, both loc entries, codec frames, dictionaries, and global
headers. Arrow IPC was unavailable on this machine and was skipped.
DataLoader: Fixed-Shape Pipeline
The original full input pipeline comparison (sample, fetch, collate, hand
over a batch), same workload, 4 workers, batch 256 (64 MiB). The peak RSS
column (added with the uniform-codec rewrite) is the highest resident set size
of the whole process tree while batches are flowing. These are retained
published loader results; current defaults keep the very slow Grain setup
optional through --only grain.
| loader | batches/s | peak RSS |
|---|---|---|
| loaderx | 72.2 | 2794 MiB |
| loaderx-raw | 136.7 | 2805 MiB |
| torch | 66.7 | 16047 MiB |
| grain | 16.9 | 4803 MiB |
At 64 MiB per batch the loader is DRAM-bound, not sampler-bound — the per-batch gather (zstd ~4.6 GiB/s, raw ~8.6 GiB/s) is the whole story, and the prefetch threads keep it at store-gather speed while the Python side collates. The memory is the prefetch buffers plus the two Store handles (raw + label stores over the same backing data). loaderx prefetches in threads inside one process, so workers share one interpreter, one numpy and one set of gather buffers; torch and grain run a worker process per prefetch thread, which is most of their RSS (torch's peak also includes the shared-memory collated batches). Compressed loaderx is slightly ahead of torch; loaderx-raw is 2.0x faster than torch and 8.1x faster than grain.
Free-threaded Python. loaderx targets free-threaded builds (no GIL), and
the key sections are also measured on 3.14t — the question is whether anything
gains when nothing is GIL-limited. These are opt-in runs (a second interpreter
and --transform augment); they are not part of the default all:
-
Sampler — zsampler draws ~1.9–2x faster than free-threaded numpy (the 256-row run was a noisy 2.46x), the same broad margin as on the GIL build: a batch is one Zig call either way.
-
Store — the free-threaded numbers track the GIL table closely.
zrecord-rawgathers 8.9 vs 8.8 GiB/s andzrecord-zstd4.7 vs 4.6 GiB/s. The no-GIL table was run without blosc2 so its import could not change the process mode. blosc2 was measured separately at 678 MiB/s; its extension emits a warning and re-enables the GIL, so that cell is not a no-GIL result. ArrayRecord has no usable 3.14t extension. Random gather, 256 KiB records:store CPython 3.14 (GIL) free-threaded 3.14t zrecord-raw 8803 MiB/s 8878 MiB/s zrecord-zstd 4603 MiB/s 4707 MiB/s npy-mmap 4408 MiB/s 3974 MiB/s hdf5 2288 MiB/s 2210 MiB/s blosc2 646 MiB/s 678 MiB/s (GIL enabled) tensorstore 299 MiB/s 323 MiB/s -
Loader, identity — unchanged on both interpreters at 64 MiB per batch:
loader CPython 3.14 (GIL) free-threaded 3.14t loaderx 72.2 batches/s 71.1 batches/s loaderx-raw 136.7 batches/s 136.0 batches/s -
Loader, CPU-heavy transform — the one place the free-threaded build matters. A transform runs on the prefetch threads, so the GIL serializes it on standard CPython — the case where worker processes win, and the reason for the free-threaded build.
--transform augment(a Python-bound per-sample loop), 4 workers. Torch and Grain have no free-threaded rows, so their GIL worker processes are shown only as context:loader CPython 3.14 (GIL) free-threaded 3.14t loaderx 33.1 batches/s 41.9 batches/s loaderx-raw 36.4 batches/s 51.9 batches/s torch 35.9 batches/s — grain 10.2 batches/s — Free-threaded wins because the transform parallelizes across the prefetch threads: 1.27x on loaderx and 1.43x on raw at 256 KiB. The margin is bounded by the 64 MiB gather, which remains DRAM-bandwidth-heavy.
Conclusion — why the numbers look like this.
Every hot path is batched natively. Zsampler draws a whole batch of indices; DenseStore gathers and decompresses a whole fixed-shape batch in one CFFI call; RaggedStore uses batched native calls for shape lengths, shape records, and payload records before Python reconstructs the exact arrays. The speedup is not bought with sampling shortcuts: the IID draw is unbiased like NumPy's (Lemire with rejection, so uniformity costs nothing over a real index space).
The layouts match what a training loader does. Zrecord is built for random record access: DenseStore gathers fixed-width records directly into one ndarray; RaggedStore restores independently shaped records from paired payload/shape entries. Array stores are built primarily for contiguous scans, so a scattered batch fights their layout. A dense raw gather fans out across every core where NumPy fancy indexing is one thread; RaggedStore instead trades some raw specialization for compression and a complete variable-shape persistence model.
Compression is in the kernel, and there is one codec. zrecord-zstd is not
"storage plus a codec": the layout, the multi-core decompress and the GIL-free
copy are one path, so turning compression on costs part of a margin, not an
order of magnitude — ~3.9 GiB/s here, ~7.3x ahead of the next compressed store
at a similar ratio. The modest ratio is the data, not the
format: these samples are barely compressible, and on a smooth image set plain
zstd reaches 7.6x and the dictionary (at the "max" tier) 16.5x.
The loader gap is architecture, not storage. loaderx uses threads and never ends an epoch, so a step pays no IPC and never waits on an epoch boundary; torch restarts per epoch with worker processes. Here compressed loaderx is 1.1x faster than torch and 4.3x faster than grain; raw loaderx is 2.0x and 8.1x faster, respectively. Storage also differs per loader — each reads from what it was built for — so the loader table is a different comparison from either store table, not a rerun. The one crack in the thread model is a CPU-heavy transform, which the GIL serializes — that is exactly what free-threaded Python removes, so loaderx is developed and benchmarked against free-threaded builds first.
What these numbers do not claim. Everything runs with a warm page cache: this
measures the access path, not cold storage or disk. on disk is allocated
blocks; zrecord files grow to their written frontier, while TensorStore's
sharding carries its own layout overhead. write is
page-cache ingestion with durability deferred, matching the other backends —
zrecord's own durability (sync/close) is a separate, explicit cost that
none of the store tables pay. This is a single qualitative pass —
the µs-scale sampler timings and the loader batches/s fluctuate with box load
(on these 12 cores the compressed loader trails the raw one substantially), so
treat the absolute numbers as ballpark and the cross-backend margins as the
signal.
Real-data verification: NTU RGB-D skeletons
The tables above are synthetic workloads. As a historical ground-truth check,
loaderx was run end to end on NTU RGB-D skeleton data — 114,480 raw .skeleton files,
120 action classes, 25 joints — processed into the ST-GCN N C T V M layout
(per-sample (3, 300, 25, 2) float32) for the xsub/xview protocols. The
.npy outputs of the standard preprocessing pipeline were treated as ground
truth. This verification was not rerun with the synthetic benchmarks above;
its throughput is retained as a separate historical 12-core result.
Correctness — the read path is bit-exact against the ground truth:
- Full scan of all 228,356 records (joint float32 + label int64, all four
splits) through
DenseStore: byte-for-byte identical to the reference npy. - A
DataLoaderover joint + label + an index stream, run under all three sampler modes (sequential,iid,cyclic): every received batch is bit-exact to the ground truth at its own declared indices, and the streams stay index-aligned. - Sampler semantics hold on real index spaces:
sequentialwalks in order,cyclicdraws a full cycle without replacement,iidis deterministic per seed.
Storage — zstd on this data:
| store | on disk | ratio |
|---|---|---|
| npy (raw float32) | 6.4 GB | 1.00x |
| zrecord raw | 6.86 GB | 1.00x |
| zrecord zstd | 0.79 GB | 8.66x |
| zrecord zstd_dict | 0.75 GB | 9.13x |
Sizes above are for one split (xview/val, 38,132 records); across all four
splits the zstd joint stores total 4.97 GB against 41 GB of raw npy (~8x).
Throughput (180 KB per record, warm page cache, 12 physical cores):
| path | throughput |
|---|---|
| random-batch gather, zstd store | 4.1–4.5 GiB/s |
| same, npy-mmap fancy indexing | 0.6–1.3 GiB/s |
| DataLoader, 4 prefetch threads | 5.7–6.6 GiB/s (123–144 batches/s) |
zstd decompression reads ~8x fewer bytes than raw storage, so the compressed store gathers faster than the raw one (zstd 4587 MiB/s vs raw 1792 MiB/s on the same split).
The npy intermediate is optional. Parse the skeleton files in parallel and
feed DenseStore.append fixed-shape chunks directly. This writes the store in
one pass with no npy staging or second read; the caller chooses chunk size as an
explicit memory bound while zrecord independently caps filesystem writes.
Current Limitations
- Single-host only; multi-host training is not supported.
- A single sample must be at most 2 GiB (2^31 bytes). There is no fixed record
count:
lengthis a u64 and the record table grows on demand, so how many records a store holds is bounded by its total chunk capacity (up to 2^64 bytes) divided by the average record size — e.g. roughly 2^44 records at 1 MiB each, 2^33 at 2 GiB each. - Metadata is read and written as the host's struct layout, so a store carries the host's byte order and is not portable to a machine of the opposite endianness. Every published platform is little-endian, so this only matters if you build for one yourself.
Build
zig build # host shared objects, into loaderx/lib/
zig build test # native store suite, in both Debug and ReleaseFast
python3 scripts/test_loaderx.py # Python integration suite against the real build
python3 scripts/bench_dense.py # fixed-shape store comparison
python3 scripts/bench_ragged.py # variable-length store comparison
python3 scripts/bench.py # machine, sampler, and dense loader layers
The Zig side is tested for behaviour only; throughput is measured from Python, through the binding a client actually uses. Optional benchmark contenders are skipped when not installed. See Benchmarks.
Publishing
Zig cross-compiles every target from one machine, so releases need no CI matrix:
zig build dist # every platform, into zig-out/dist/<wheel tag>/
python3 scripts/build_wheels.py # one wheel per platform, plus the sdist
The dist directories are named after their Python wheel platform tag, so the tag
mapping lives in exactly one place (dist_targets in build.zig). glibc and
macOS minimums are pinned in the target triple, which is what makes
manylinux_2_17 and macosx_11_0 honest rather than aspirational. Each wheel is
checked after packing: it must carry this platform's libraries and no others.
The sdist ships sources only. Installing from it runs zig build through
setup.py, so it needs the Zig compiler; wheel users never hit that path.
Zsampler
Index Generator: a high-performance sampler implemented in Zig. Every mode is a
pure function of (seed, step), so a run resumes exactly by seeking to a step —
there is no epoch to track, in keeping with the endless step-based loader.
from loaderx.zsampler import Sampler
sampler = Sampler(1_000_000, 256, Sampler.Mode.IID, seed=42)
indices = sampler.next()
- Sequential — traverse the index space in order through a fixed-size sliding window, treating the space as a circular queue so the tail never truncates.
- IID — draw each index uniformly at random with replacement. Unbiased (Lemire with rejection), matching NumPy. Simplest, but coverage is uneven over any short run.
- Cyclic — without replacement, round-robin. Each cycle traverses a fresh
permutation of the whole index space, so within a cycle every record appears
exactly once and no batch repeats an index — coverage is even by construction,
which keeps how often each sample is seen uniform. The permutation is a
stateless bijection (a small Feistel network over the index space, brought
into range by cycle-walking), so a million-record shuffle materializes nothing
the size of the dataset and reshuffling each cycle is free. Every batch is
exactly
batch_size: an endless step-based loader has no final partial batch to special-case, so whenbatch_sizedoes not divide the length the cycle's remainder is dropped — a different remainder each cycle, since the permutation changes, so every record is still reached over time.
Zrecord
Zrecord is loaderx's typed persistent record store. Its private native runtime is
the byte-oriented engine beneath the public DenseStore and RaggedStore contracts.
The private Python CFFI surface is kept together in loaderx/_store.py because
both handles share one libstore, error model and lifecycle. In Zig,
src/store/dense.zig and src/store/ragged.zig own their respective typed semantics,
src/store/common.zig owns the typed header/type-identity rules, and the top-level
src/store.zig is the sole C ABI export/composition root. zrecord.py remains
the unified Python-facing API.
RecordEngineis an unordered physical store made of N records. Records are independent and carry no ordering, so every index and slice operation is equivalent to a gather.- It hands the Store layer a dense index space: the live records are always
exactly
0..N. Deletion preserves that by swapping the tail into the hole, which means an index is stable only until something is deleted. Named streams are composed dynamically by a plain Python dict;DataLoadervalidates that the independent stores have equal lengths. - The engine reads and writes byte ranges. Native
DenseStoreandRaggedStoreown kind, type identity, geometry and transaction semantics. Dense stores persist one record width. Ragged stores place payload and opaque shape metadata in alternating physical records (2*i,2*i+1). One engine append publishes both with one header-length commit; grouped delete swaps both locations and publishes one shortened length. - The IO model (
append | read | delete) is batch-oriented and shape-agnostic. Fixed, variable and paired operations describe only the physical buffer geometry; they do not carry dtype or array-shape semantics. A single-record operation is just thebatch_size == 1case. - The engine owns its temporary memory internally — allocation and release are explicit.
- A store has exactly one codec in its native physical header, fixed at creation and immutable afterwards. Every record is compressed and decompressed independently:
| tag | name | algorithm |
|-------|-----------|----------------------------------------|
| 0 | raw | none |
| 1 | zstd | zstd (plain, level 3) |
| 2 | zstdict | zstd with a trained dictionary (level 19) |
- Compression is transparent to the client:
- Compression runs concurrently across all cores. A compressed store never
falls back to raw: each record is stored as the codec's output, even when
an incompressible record's frame is larger than its input — write
rawif the data does not compress. - Decompression writes straight into the caller's destination memory
(
gather), with no intermediate buffer and no extra copy. - zstd is the one transparent codec — faster than Deflate at both ends and a better ratio, so there is no reason to carry a second. It is vendored C, built for every platform by Zig, so the one-wheel-per-platform story is unchanged.
zstd_dictadditionally trains one dictionary on a sample of the data (stored asdict.zr) and compresses every record against it. Because each record is still independent, random access is unchanged — but the dictionary carries the structure shared across records, which per-record compression cannot see. On many small, similar records (image tiles, token sequences) this is a large win: a smooth-image set that plain zstd takes to 7.6x compresses 16.5x with a large dictionary. The dictionary is loaded once on open and shared, lock-free, across all reader threads. The dictionary size is chosen from theDICT_TIERSpresets (see Codec notes).- A
zstd_dictstore needs its dictionary to read every record; araworzstdstore rejects an unexpected dictionary as malformed state.
- Compression runs concurrently across all cores. A compressed store never
falls back to raw: each record is stored as the codec's output, even when
an incompressible record's frame is larger than its input — write
Persistence format
Native storage is metadata plus chunked data. The extension is the type:
.zr files are store-global singletons, .loc files are record-table segments,
.chunk files are record data:
store/
├── store.zr immutable typed header + exact Python type identity
├── meta.zr header (global state)
├── dict.zr zstd dictionary (only in dict stores)
├── 0.loc record table segment [0, 2^28)
├── 1.loc record table segment [2^28, 2^29)
├── 0.chunk record data
└── 1.chunk
Metadata (store.zr + meta.zr + {id}.loc)
Files are read and written positionally — pread/pwrite at computed offsets,
no mmap. The header is a bit-packed struct and a .loc segment an array of
16-byte RecordLocs, exactly as wide as they declare, so a location is one
pread/pwrite of 16 bytes at a computed offset and there is no serializer
anywhere in the code. Every header field is byte-aligned (no bit fields cross a
byte), so the packed header reads as plain memory.
The record table is partitioned into {id}.loc segments so it can grow by
appending a segment instead of reserving the maximum. The id→segment mapping
is pure arithmetic — seg = idx >> 28, off = (idx & (2^28−1)) × 16 — so a
segment needs no per-record bookkeeping. Committed entries are immutable; one
shared fd-table lock keeps the in-memory ArrayList stable while readers index it
during the rare append that adds another segment.
1. Typed header — store.zr contains a 32-byte typed-header version 1 followed by
the immutable encoded type identity. It records native kind and dense record
width; codec lives only in the physical header. It is written only after the
engine and required dictionary are durable, and typed open validates it before
exposing the store.
2. Physical header — 32 bytes, the whole of meta.zr.
magic(ZREC) andversionare ordinary fields, so opening a directory that is not a loaderx store fails immediately instead of decoding garbage.codecis the store's one compression method, stamped at creation and immutable — there is no per-record tag anywhere.length(u64) is the physical record count; it equals logical length for dense stores and twice logical length for ragged payload/shape pairs.tail_chunk/tail_offsetmark the last write position.- There is no chunk count. Chunks are created in order and the frontier is
always in the last one, so the store holds exactly chunks
0..=tail_chunk— a count would be a second copy of that fact to keep in sync.
const Codec = enum(u8) { raw = 0, zstd = 1, zstdict = 2, _ };
const Header = packed struct {
magic: u32, version: u8, codec: Codec,
tail_chunk: u32, tail_offset: u32, length: u64,
_reserved: u80,
};
3. Record table — a .loc segment addresses up to 2^28 entries of 16 bytes
(4 GiB), indexed directly. The file starts empty and grows as contiguous loc
runs are written; 4 GiB is its address boundary, not its initial file size.
Mapping an index to a physical address is what makes random access efficient.
chunk_idis the containing chunk |offsetis the position within it |phys_length/logic_lengthare the stored and original sizes. The codec is not here: it is the header's, so a record is stored exactly the way the store is declared.
const RecordLoc = extern struct {
offset: u32, phys_length: u32, logic_length: u32,
chunk_id: u32,
};
There is no liveness flag. Every entry below length is live, because deletion
swaps the tail into the hole rather than tombstoning.
4. No maximum length. The table grows a .loc segment at a time and the
data grows a chunk at a time, so there is no static record-count cap to size
against. The real bounds are the field widths — 2^32 chunks of 2^32 bytes
(2^64 bytes total), 2^31 (2 GiB) per record — and disk.
Executor
1. Write. Writes are append-only; everything else is offset redirection.
The codec is immutable store state, so append and gather dispatch once at
their entry points into separate raw, zstd, or zstd-dictionary implementations.
Their contexts and workers are deliberately not unified: only validation,
location bounds, locking, and the final metadata transaction are shared.
Geometry is equally explicit across the whole stack: DenseStore calls its
fixed-stride ABI without constructing an offset array or resupplying record
width, while RaggedStore supplies payload and shape offsets to its own ABI.
There is no generic append/gather handle or ABI.
- Append: compression workers claim individual records, compress into independent
slots, reserve physical offsets in completion order through a short frontier
lock, and issue positional payload writes in parallel. Logical IDs remain in
the loc table, so physical completion order does not change random gather.
The caller waits for every payload before committing metadata. Append publishes
through the page cache and is intentionally lazy;
sync()makes the committed frontier durable in payload → record-table → header order and propagates any sync failure. A record never straddles two chunks; one that would not fit rolls over to a fresh chunk. Chunk files start empty and positional writes extend them to the current frontier; the 4 GiB u32 offset range is a logical capacity, not a sparse preallocation requirement. Compression fans out to the executor's CPU count without a separate append-specific worker cap. Each producer configures its CCtx or shared immutable CDict once, then starts every independent record frame withZSTD_compress2. - Raw flush granularity is the store's, not the caller's. The batch size a
client passes to
appendbounds memory only; the raw flush stage packs contiguous records intopwritevs capped at 256 KiB (flush_bytes). This decouples syscall size from batch and record size — a huge batch and a tiny one write identically, which matters because per-syscall writes above a few hundred KiB land on a slow regime on some filesystems/hosts (~1.4 GiB/s vs ~2.1 GiB/s at 256 KiB writes on the reference box). On POSIX the gather is a realpwritevwith up to IOV_MAX (1024) iovecs. Windows has no vectored file write —WriteFileGatheris async-only, OVERLAPPED, page-aligned — so it falls back to one positional write per record (correct, if not gathered). If Windows grows a synchronous vectored/IOCP file-write story, this is the one place writes there can be pulled up to parity; the byte-budget grouping is already in place, only the syscall differs. - Delete: swap the last table entry into the deleted slot and drop the length by
one. A batch is applied in descending index order, so each swap pulls from a
slot no later target refers to. The index space stays dense — which is what
the sampler needs, since it draws uniformly from
0..Nand would otherwise keep hitting holes. The deleted record's bytes become garbage.
2. Read. Fill the destination memory concurrently, in place from the Python side (executed on async threads).
- Committed records are immutable and
lengthis published through an atomic. A gather takes one shared fd-table lock so a rare chunk/segment ArrayList growth cannot invalidate its file handles; record I/O itself stays lock free. - Every record is read at the offset its table entry records — the record table is addressed by pure arithmetic, so random access is one pread for the location and one for the bytes, with no batching assumptions about layout. Each shard reads one location and immediately reads/decompresses that record; there is no separate metadata phase or sequential-run special case. Compressed records use a per-shard staging buffer and decode in place into the destination.
3. Concurrency model. Io.Group.async shards work by CPU count, and shards
beyond the limit run inline on the calling thread.
- Shards receive contiguous blocks rather than a strided subset, keeping each worker's reads and writes sequential.
- Each shard creates one zstd context (
ZSTD_CCtxto write,ZSTD_DCtxto read) and reuses it across every record it handles, rather than paying that setup per record. The dictionary (ZSTD_CDict/ZSTD_DDict) is immutable, so all shards share one, lock-free. - Decompression writes straight into the caller's destination buffer, so there is no intermediate copy.
4. Garbage collection. Deletion leaves the record's bytes stranded, so space
is reclaimed by an offline compact that rewrites the live records in place.
- Records are visited in physical order — which after swap-last deletes no longer matches table order — and repacked densely in that same order. A record therefore never moves to a higher address than it already had.
- Two things follow. Writing a record can never land on the bytes of a record not yet moved, so the rewrite is safe in place with no scratch copy of the store. And each table entry can be updated the instant its record lands, so the store is consistent at every point: an interrupted compaction leaves some records moved and the rest where they were, and re-running finishes the job.
- Records that are already in the right place are skipped, so a store with a small amount of garbage near the end is cheap to compact.
- Chunk files past the new frontier are closed and deleted; the final retained chunk is truncated to its new tail offset, releasing its old physical tail.
- Delete has already made the live loc table a dense
0..lengthprefix, so compact also deletes loc segments beyond that prefix and truncates its final segment to exactlylength * 16bytes. stats()reportslive_bytesagainstchunk_bytesso callers can decide when it is worth running. Note the difference is an upper bound: a record never straddles a chunk boundary, so up to one record's worth per chunk is slack that compaction cannot remove.
5. File access.
- Metadata:
meta.zr(32 bytes) plus naturally growing.locsegments, each with a 4 GiB maximum address range. - Chunk data: naturally growing files with a 4 GiB logical capacity, accessed
concurrently through
readPositionalAll/writePositionalAll. No path depends on filesystem sparse-file support.
Concurrency contract. One native handle owns a store at a time through a
lifetime, nonblocking exclusive lock on meta.zr; a second handle or process
fails with StoreBusy. Within that handle, gather and append are safe to
call concurrently from many threads. Append remains single-writer; only rare chunk/segment fd-table
growth briefly waits for active gathers. delete and compact mutate the table
in ways a reader would observe half-applied, so they require exclusive access to
the store.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distributions
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file loaderx-1.7.2.tar.gz.
File metadata
- Download URL: loaderx-1.7.2.tar.gz
- Upload date:
- Size: 614.4 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/7.0.0 CPython/3.14.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
ef8c3c29d6e453a689140a497fdc8001140d31119460aca89374c3770cc8c788
|
|
| MD5 |
9a81e32f619bc33d7c5437bd63d1c0f9
|
|
| BLAKE2b-256 |
b8b2da9dc4befc047986167f6e50f3280a9968a6347b2483626366ea6508c0d1
|
File details
Details for the file loaderx-1.7.2-py3-none-win_arm64.whl.
File metadata
- Download URL: loaderx-1.7.2-py3-none-win_arm64.whl
- Upload date:
- Size: 394.8 kB
- Tags: Python 3, Windows ARM64
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/7.0.0 CPython/3.14.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
cf329d5b70e3a4730d2ef23754f1e2b95bea61315029c3e12b9a54d6434940e0
|
|
| MD5 |
81b0a73f149cc9a02e7f9239abb5c656
|
|
| BLAKE2b-256 |
aaf23b78858395348c1e0c905b55c324e4d60fde3fcb2a042500adda580e6a69
|
File details
Details for the file loaderx-1.7.2-py3-none-win_amd64.whl.
File metadata
- Download URL: loaderx-1.7.2-py3-none-win_amd64.whl
- Upload date:
- Size: 539.4 kB
- Tags: Python 3, Windows x86-64
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/7.0.0 CPython/3.14.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
cd2bf89cf783833f66a4f1c278eed8f0c5ac097cc491ab1b20a70bd3e9f37854
|
|
| MD5 |
25cdf3c2ca7d0f5e5cab67b582b111ea
|
|
| BLAKE2b-256 |
0f6d24b9192f31f6a7ed3ed0f4c42dc5aed5235b6c0ed473489146a04819918e
|
File details
Details for the file loaderx-1.7.2-py3-none-musllinux_1_2_x86_64.whl.
File metadata
- Download URL: loaderx-1.7.2-py3-none-musllinux_1_2_x86_64.whl
- Upload date:
- Size: 469.2 kB
- Tags: Python 3, musllinux: musl 1.2+ x86-64
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/7.0.0 CPython/3.14.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
cea57c0f960086691312f83d241def865f3dc9e8a2c51b86b958a8858f168bbd
|
|
| MD5 |
438722db97454db22ef561e820e22de4
|
|
| BLAKE2b-256 |
4b47fae3047057529a924f4f02204517fe62252ae941f901197089289b353f91
|
File details
Details for the file loaderx-1.7.2-py3-none-musllinux_1_2_aarch64.whl.
File metadata
- Download URL: loaderx-1.7.2-py3-none-musllinux_1_2_aarch64.whl
- Upload date:
- Size: 395.9 kB
- Tags: Python 3, musllinux: musl 1.2+ ARM64
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/7.0.0 CPython/3.14.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
b3c48a5c2e4365a45f4f6984deaa2d0236f3872083a7b375f8c6bf990aebf7fb
|
|
| MD5 |
f083e966fc1a2d9f664581c1692c8a17
|
|
| BLAKE2b-256 |
e3c7b3e6ec3766e863acd6668fd73de862cb18518b04b040fca2531eb9df9183
|
File details
Details for the file loaderx-1.7.2-py3-none-manylinux_2_17_x86_64.whl.
File metadata
- Download URL: loaderx-1.7.2-py3-none-manylinux_2_17_x86_64.whl
- Upload date:
- Size: 458.6 kB
- Tags: Python 3, manylinux: glibc 2.17+ x86-64
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/7.0.0 CPython/3.14.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
e7234b571be2a86113914cf2195b6e422c06c5d280ae13e3eb496df617e77449
|
|
| MD5 |
d556a13f06b9580ae8509266da4dbfd9
|
|
| BLAKE2b-256 |
b4a1749bf5295c1b020570cf272b7a5565f374e079ad73b1f6189e1d94b7e6f2
|
File details
Details for the file loaderx-1.7.2-py3-none-manylinux_2_17_aarch64.whl.
File metadata
- Download URL: loaderx-1.7.2-py3-none-manylinux_2_17_aarch64.whl
- Upload date:
- Size: 387.0 kB
- Tags: Python 3, manylinux: glibc 2.17+ ARM64
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/7.0.0 CPython/3.14.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
7502073d6d96fb63b4f9806684f5b814286f39974b462c34ff37a04c2b699066
|
|
| MD5 |
0346e7e55926d43be81ac98d14954c4a
|
|
| BLAKE2b-256 |
4c9d965f4e7f0ef31fd24be84332dc701d854a7966d9419a0208d8286c46f748
|
File details
Details for the file loaderx-1.7.2-py3-none-macosx_11_0_x86_64.whl.
File metadata
- Download URL: loaderx-1.7.2-py3-none-macosx_11_0_x86_64.whl
- Upload date:
- Size: 434.9 kB
- Tags: Python 3, macOS 11.0+ x86-64
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/7.0.0 CPython/3.14.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
0c9d232d9e934f59170eba58589383bb361a9caf73c64aee0b4afd254b4c795e
|
|
| MD5 |
300dd3bdf46ae10cd5f948a780d4a412
|
|
| BLAKE2b-256 |
82ce6c42bb0c268a409e760355a6f05c94156f800fff933368ac3ea4ad038f9b
|
File details
Details for the file loaderx-1.7.2-py3-none-macosx_11_0_arm64.whl.
File metadata
- Download URL: loaderx-1.7.2-py3-none-macosx_11_0_arm64.whl
- Upload date:
- Size: 371.9 kB
- Tags: Python 3, macOS 11.0+ ARM64
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/7.0.0 CPython/3.14.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
4d321b6515c5d93c7487039623837ff4086bd3307dedd2ee5b399acf5fa05c4e
|
|
| MD5 |
9e734e021e32d151df77cc11cb6f9628
|
|
| BLAKE2b-256 |
f89a4b76d177abd0e3deabf181cc546ace51cf5537998b942a1d65681f491191
|