Skip to main content

Loaderx

A compact, high-performance persistent record store with zero-copy batch gathering and transparent per-record compression, designed for single-machine AI training and serving pipelines

Zrecord is the typed persistent store; Loaderx is the sampler and prefetch loader that consumes Zrecord streams. They currently ship together while both layers mature, but their public responsibilities remain separate.

pip install loaderx

Wheels are published for Linux (glibc ≥ 2.17 and musl, x86-64 and arm64), macOS (≥ 11.0, Intel and Apple Silicon) and Windows (x64 and arm64). The bindings use cffi in ABI mode, so nothing links against the CPython ABI and one wheel per platform serves every supported Python.

Design Philosophy

loaderx is built around several core principles:

  1. A pragmatic approach that prioritizes minimal memory overhead and minimal dependencies.
  2. A strong focus on single-machine training workflows.
  3. We implement based on NumPy semantics, persisted through the Zrecord storage runtime.
  4. An immortal (endless) step-based data loader, rather than the traditional epoch-based design—better aligned with modern ML training practices.
  5. Dense and ragged are separate contracts, and the loader serves both. A dense stream stacks into one array per batch; a ragged one comes back as a list. Neither is padded, and equal length is never treated as a special case of variable length.

Quick Start

from loaderx.utils import from_numpy
from loaderx.zrecord import DenseStore
from loaderx.dataloader import DataLoader

from_numpy('train_data', np.load('data.npy', mmap_mode='r'))
from_numpy('train_label', np.load('label.npy', mmap_mode='r'))

loader = DataLoader({'data': DenseStore('train_data'), 'label': DenseStore('train_label')},
                    transform=lambda batch: batch)

for i, batch in enumerate(loader):
    if i >= 256:
        break

print(batch['data'].shape)
print(batch['label'].shape)

A batch is a dict {name: values}: each value is the stacked (batch_size, *item_shape) array for that stream. Every stream is gathered at the same indices, so record i lines up across them. The transform callback is the collate step — reshape, cast, stack — where values is the plain dense batch ready for the model.

Converting a NumPy tensor

import numpy as np
from loaderx.utils import from_numpy
from loaderx.zrecord import DenseStore

from_numpy('train_data', np.load('data.npy', mmap_mode='r'))
from_numpy('train_label', np.load('label.npy', mmap_mode='r'))

# many small, similar records compress far better against a trained dictionary:
from_numpy('train_data', images, codec='zstd_dict')

One record per slice along axis 0; a 1-D array (the usual shape of a label set) becomes a store of scalar records. The conversion streams in bounded chunks, so an mmapped array is never fully materialized. Open the result with :class:DenseStore:

ds = DenseStore('train_data')

The store carries a single schema.msgpack recording the kind ("dense" or "ragged"), the per-sample dtype, and the shape contract — one item_shape for a dense store, one shape per record for a ragged store. That is the only loaderx-level metadata. The native zrecord engine beneath it is an implementation detail; raw bytes and pre-encoded images use a RaggedStore(path, dtype=np.uint8) rather than a second public storage API.

Records

One store engine, two kinds of view. from_numpy writes a dense store where every record is exactly one row of the recorded item_shape; reads are fixed-stride gathers and the batch shape follows from the schema, so no per-record metadata is touched:

import numpy as np
from loaderx.utils import from_numpy
from loaderx.zrecord import DenseStore

from_numpy('data', np.arange(64, dtype=np.float32).reshape(8, 2, 4))
ds = DenseStore('data')
ds[0, 5, 2]                          # (3, 2, 4) — a record set, shape straight from the schema

A store is a collection of records, not an ndarray, so ds[0, 5, 2] is a record set — never ds[0][5][2]. A scalar selects one record; negatives wrap from the end. The ragged example below reads the same way.

from_iterator writes a ragged store of variable-length records. It is a separate contract: :class:RaggedStore hands back a list of arrays, so a loader never has to carry row_splits around. dtype is unified and explicit; each record keeps its own shape, recorded per record as it is written and restored exactly on read — so records may differ in shape arbitrarily, and nothing is ever inferred from the source (an iterator can't tell you what its later records look like). Densifying a list into a dense batch is the model's call — a plain numpy loop, wherever you need it:

from loaderx.utils import from_iterator
from loaderx.zrecord import RaggedStore

seqs = [np.arange(L, dtype=np.int32) for L in (3, 1, 4, 1, 5)]
from_iterator('tokens', seqs, np.int32)   # dtype explicit; each record keeps its shape
rs = RaggedStore('tokens')

records = rs[0, 2, 4]                # list of ndarray — one per record, exact shapes
lengths = np.array([len(r) for r in records])
padded = np.zeros((len(records), lengths.max()), dtype=records[0].dtype)
for i, r in enumerate(records):
    padded[i, :len(r)] = r             # (B, max_len) — your policy, your loop

A DataLoader works with dense streams, so collation is just the transform — a batch dict in, a batch dict out:

def collate(batch):
    return {'input_ids': batch['tokens'], 'label': batch['label']}

loader = DataLoader({'tokens': dense_tokens, 'label': labelset}, batch_size=32,
                    transform=collate)

Writable stores

The Store classes are loaderx's public persistence API: they are read-write, and append is explicit — one container of records, one batch, one native call. Nothing is buffered, inferred, or compressed for you.

from loaderx.zrecord import DenseStore, RaggedStore

ds = DenseStore('mnist/x', dtype=np.uint8, item_shape=(28, 28))  # schema written now
ds.append(images[i:i + 1024])        # (B, 28, 28) → one batch, returns first index
ds.append(single_image[None])        # one sample is batch_size 1 — add the axis yourself
ds[:4]                               # read path is unchanged

tok = RaggedStore('tokens', dtype=np.int32)     # each record keeps its own shape
tok.append([seq_a, seq_b, seq_c])    # list in, list out — `tok[0, 2]` returns a list

The schema is declared at construction and written to schema.msgpack immediately; append validates every record against it and errors loudly. The dictionary is the one thing kept out of append: train_dict is explicit and separate, and a zstd_dict store refuses to append until it is called.

The store's maintenance pass-throughs are on the same object, normalized like reads (scalar, slice, array, bool mask):

ds.delete(np.arange(0, len(ds), 2))   # swap-last: index space stays dense, survivors reindex
ds.stats()                            # {'records', 'live_bytes', 'chunk_bytes', 'reclaimable'}
ds.compact()                          # reclaim the deleted bytes, in place — offline only

Deletion and compaction move records, so a multi-stream index you keep yourself must tolerate it — the guarantee is only that the live records are exactly 0..len(ds).

Codec notes

"zstd" compresses each record independently with plain zstd (level 3). Use it for general-purpose compression — it is fast and the default.

"zstd_dict" trains a shared dictionary on a sample of the data before writing any record, then compresses every record against it at level 19. The dictionary captures structure shared across records that per-record compression cannot see — a large win for many small, similar records (image tiles, token sequences).

The dictionary's cost is a cache footprint: every record's decompression references the shared dictionary window, so a larger dictionary means more cache misses on gather — the path a loader pays forever. Three named tiers (loaderx.zrecord.DICT_TIERS) preset the whole tradeoff — the dictionary size and how much data trains it — so a caller picks a tier, never a number:

tier dict sample tradeoff
"fast" 32 KiB 4 MiB fastest gather and training; ratio barely above zstd
"balanced" 128 KiB 16 MiB default — most of the ratio at a fraction of the gather cost
"max" 1 MiB 64 MiB best ratio; slowest gather and training

The sample is the byte budget the dictionary trains on (a strided subset of the records), so each tier costs the same training time whatever the record size. At realistic image scale (768 KiB records) the tiers converge — on the earlier measurement box's structured data "balanced" and "max" both gather ~1.6 GiB/s at a 1.74x ratio — because a dictionary is a small fraction of a large frame. The tiers still matter at small record sizes, where the dict is most of a record and "max" trades gather throughput for ratio.

"zstd_dict" records can only be read from a store that has the dictionary (dict.zr). The dictionary is loaded on open and shared, lock-free, across all reader threads.

A dictionary must train on the settled, complete data. :func:train_dict is the standalone, manual training step — a numpy array, an iterable of records, or a typed store all train the same way, sized by a tier — and its bytes are handed to a store-writing path via dict_bytes. zstd_dict never trains by itself: a write without a dictionary is an error. A stream cannot train its own dictionary, but it can write with one trained on the settled data:

from loaderx.utils import train_dict, from_iterator, from_numpy

# train once on the settled data — a standalone, reusable artifact
d = train_dict(settled_array, tier="balanced")

# then any store can write zstd_dict with it, including a stream
from_iterator('tokens', token_generator, np.int32, codec='zstd_dict', dict_bytes=d)
from_numpy('data', data, codec='zstd_dict', dict_bytes=d)

Changing an existing store's codec is a rewrite, not a store mutation: create a new store with the new codec and copy the records in index order. Index order is what keeps multi-stream alignment, and the copy loop is a few lines with the Store layer — which is why no dedicated recode path exists:

from loaderx.utils import train_dict
from loaderx.zrecord import DenseStore

CHUNK = 1 << 16
with DenseStore("src") as s, \
     DenseStore("dst", dtype=s.dtype, item_shape=s.item_shape,
                  codec="zstd_dict") as d:
    if not d.has_dict():                       # zstd_dict needs a dictionary
        d._store.install_dict(train_dict(s))   # train one on the source store
    for start in range(0, len(s), CHUNK):
        d.append(s[start:start + CHUNK])       # records copied in index order

RaggedStore is the same shape — swap the constructors and dtype is all the destination needs. dst must not already hold a store.

Important: Do not use "zstd_dict" while data is still changing (through deletes or compaction below the Python layer). The dictionary captures a snapshot of the data; training it before the data settles wastes compression. Train the dictionary once preprocessing is complete and the content is final.

Multi-stream stores (MultiStore)

A zrecord store is one stream; a training sample is usually several streams (skeleton + label + id, tokens + label, ...). A multi-store is just a directory whose immediate children are stores. MultiStore wraps them and guarantees the one thing that matters: index alignment — record i lines up across every stream, append writes every stream at the same indices, delete removes the same indices from every stream, so alignment survives. No new storage format, no manifest, no bundle-level batch API.

from loaderx.utils import from_numpy, from_iterator
from loaderx.zrecord import MultiStore
from loaderx.dataloader import DataLoader

root = "xsub/train"
from_numpy(root + "/joint", joint)     # the streams are ordinary zrecord stores
from_numpy(root + "/label", label)
from_iterator(root + "/token", iter(seqs), np.int32)

ds = MultiStore(root)                   # wrap + verify they hold the same count
ds["joint"][0, 5, 2]                 # the stream's own index forms apply
ds.append({"joint": b, "label": l})    # same batch size, all streams, once
ds.delete([0, 5])                      # same indices, all streams, still aligned

loader = DataLoader(ds.streams, batch_size=256)   # hand the streams to a loader
batch = next(loader)                      # {name: values}, index-aligned

Equal-length and variable-length are both just stores — DenseStore (one fixed-shape record per sample) and RaggedStore (variable row count per record). The loader does not interpret either: it fetches by index and packs a dict, so a dense stream's batch value is the stacked (B, *item_shape) array and a ragged stream's is a list of per-record arrays. No padding is imposed — densify to a fixed shape however the model needs (a plain numpy loop), or reshape/stack in the transform collate.

CPU → GPU transfer

loaderx hands over CPU batches; getting them to the accelerator is the transform's job — the one place your framework is already imported. The batch dict is a plain {name: numpy array}, zero-copy on the way out, so a device transfer is one call per stream:

import torch

device = "cuda:0"
def to_device(batch):
    return {k: torch.from_numpy(v).to(device, non_blocking=True)
            for k, v in batch.items()}

loader = DataLoader(streams, transform=to_device)
for batch in loader:
    model(batch)                        # already on device

A non_blocking=True copy is genuinely asynchronous only when its source is pinned. loaderx does not pin memory for you — pinning is framework-owned (torch's .pin_memory(), CUDA's cudaHostAlloc), and a vendor-free core stops exactly at the CPU batch. Pin in the transform what you copy:

def to_device(batch):
    return {k: torch.from_numpy(v).pin_memory().to(device, non_blocking=True)
            for k, v in batch.items()}

JAX is the same shape — jax.device_put is already an asynchronous handoff on GPU:

import jax

def to_device(batch):
    return {k: jax.device_put(v) for k, v in batch.items()}

The transfer runs on the transform stage and never touches loaderx internals: the copy overlaps the next batch's gather/transform, and any pinned pool is the caller's to own and reuse. This is the entire H2D answer — there is no pin= hook or device backend, because the only unified thing a multi-framework loader can own is the CPU batch.

For practical integration examples, please refer to the Data2Latent repository

Benchmarks

scripts/bench.py measures zrecord and loaderx against the alternatives in the index sampler, the record store, and the full data loader — each through the binding a client actually uses, so cffi, the GIL and the NumPy allocation are all inside the timings. This is a qualitative horizontal comparison, sized to a realistic image workload (625 MiB of 256 KiB records). python3 scripts/bench.py all prints the machine block first and then runs the default tables. Warm page cache.

Machine — one local workstation (AMD Ryzen AI 9 HX PRO 370, 12 cores / 24 threads):

machine value
CPU AMD Ryzen AI 9 HX PRO 370 w/ Radeon 890M, 1 socket, 12 cores / 24 threads
frequency 605–5158 MHz
caches L1d 576 KiB, L1i 384 KiB, L2 12 MiB, L3 24 MiB
NUMA 1 node
memory 31 GiB (not limited by cgroup)
shared memory 16 GiB /dev/shm
OS Debian GNU/Linux forky/sid, kernel 7.1.3+deb13-amd64, x86_64
python 3.14.7, numpy 2.5.2

The benchmark process sees all 24 threads and is not memory-limited by cgroup. The 16 GiB shared-memory mount accommodates the four-worker, 64 MiB-batch torch pipeline. Store reads run on the ordinary page cache.

1. Sampler — index generation on its own, IID (with replacement), 1M index space, against NumPy's modern API. The µs-scale figures fluctuate with box load; the stable signal is the ~1.9x margin, roughly flat across batch sizes.

sampler batch per batch vs default_rng
numpy default_rng 256 3.8 µs 1.00x
zsampler 256 2.0 µs 1.92x
numpy default_rng 1024 5.1 µs 1.00x
zsampler 1024 2.8 µs 1.84x
numpy default_rng 8192 20.4 µs 1.00x
zsampler 8192 10.0 µs 2.03x

2. Store — Zrecord against the alternatives: random batch gather, 2500 records, batch 256, structured (mildly compressible) data. A single 256 KiB sample (512×512 single-channel, a realistic image record) is the horizontal comparison; other record sizes are a --shape away. hdf5-gzip, arrayrecord, torch and grain are installed and run in the default tables.

Realistic records — 256 KiB per record, 64 MiB per batch:

store write gather on disk ratio
zrecord-zstd 816 MiB/s 4603 MiB/s 555 MiB 1.13x
zrecord-zstdict 43 MiB/s 2017 MiB/s 388 MiB 1.61x
zrecord-raw 1990 MiB/s 8803 MiB/s 625 MiB 1.00x
npy-mmap 2686 MiB/s 4408 MiB/s 625 MiB 1.00x
hdf5 2559 MiB/s 2288 MiB/s 625 MiB 1.00x
hdf5-gzip 51 MiB/s 237 MiB/s 445 MiB 1.40x
blosc2 73 MiB/s 646 MiB/s 536 MiB 1.17x
tensorstore 397 MiB/s 299 MiB/s 462 MiB 1.35x
arrayrecord 518 MiB/s 1602 MiB/s 567 MiB 1.10x

write is page-cache ingestion — the store is built and left open, exactly like np.save and the other backends, which defer durability to the kernel. A zrecord is append-only, so this is one pass over the frontier; at 256 KiB records the 16-byte table entry per record is negligible. Durability is the explicit sync()/close — earlier write numbers included zrecord's close-time fsync while the competition did not, which compared durable writes against lazy ones. The absolute write figures fluctuate with machine load and should be treated as ballpark figures.

gather is fully page-cache-warmed — the benchmark sweeps every record once before timing, so it measures the pure access path (memory), not first-touch page faults or disk. At 256 KiB records the batch (64 MiB) exceeds L3, so these numbers are DRAM-bandwidth-bound for every backend. Fully warm, zrecord-raw leads npy-mmap by 2.0x on this box — its gather fans out across every core while numpy's fancy index is single-threaded. Plain zstd reaches 4.6 GiB/s, 7.1x the next compressed contender at a similar ratio (blosc2).

3. Loader — the full input pipeline end to end (sample, fetch, collate, hand over a batch), same workload, 4 workers, batch 256 (64 MiB). The peak RSS column (added with the uniform-codec rewrite) is the highest resident set size of the whole process tree while batches are flowing. torch, grain and arrayrecord are installed and wired into the defaults, and all four loaders run at the full batch.

loader batches/s peak RSS
loaderx 72.2 2794 MiB
loaderx-raw 136.7 2805 MiB
torch 66.7 16047 MiB
grain 16.9 4803 MiB

At 64 MiB per batch the loader is DRAM-bound, not sampler-bound — the per-batch gather (zstd ~4.6 GiB/s, raw ~8.6 GiB/s) is the whole story, and the prefetch threads keep it at store-gather speed while the Python side collates. The memory is the prefetch buffers plus the two Store handles (raw + label stores over the same backing data). loaderx prefetches in threads inside one process, so workers share one interpreter, one numpy and one set of gather buffers; torch and grain run a worker process per prefetch thread, which is most of their RSS (torch's peak also includes the shared-memory collated batches). Compressed loaderx is slightly ahead of torch; loaderx-raw is 2.0x faster than torch and 8.1x faster than grain.

Free-threaded Python. loaderx targets free-threaded builds (no GIL), and the key sections are also measured on 3.14t — the question is whether anything gains when nothing is GIL-limited. These are opt-in runs (a second interpreter and --transform augment); they are not part of the default all:

  • Sampler — zsampler draws ~1.9–2x faster than free-threaded numpy (the 256-row run was a noisy 2.46x), the same broad margin as on the GIL build: a batch is one Zig call either way.

  • Store — the free-threaded numbers track the GIL table closely. zrecord-raw gathers 8.9 vs 8.8 GiB/s and zrecord-zstd 4.7 vs 4.6 GiB/s. The no-GIL table was run without blosc2 so its import could not change the process mode. blosc2 was measured separately at 678 MiB/s; its extension emits a warning and re-enables the GIL, so that cell is not a no-GIL result. ArrayRecord has no usable 3.14t extension. Random gather, 256 KiB records:

    store CPython 3.14 (GIL) free-threaded 3.14t
    zrecord-raw 8803 MiB/s 8878 MiB/s
    zrecord-zstd 4603 MiB/s 4707 MiB/s
    npy-mmap 4408 MiB/s 3974 MiB/s
    hdf5 2288 MiB/s 2210 MiB/s
    blosc2 646 MiB/s 678 MiB/s (GIL enabled)
    tensorstore 299 MiB/s 323 MiB/s
  • Loader, identity — unchanged on both interpreters at 64 MiB per batch:

    loader CPython 3.14 (GIL) free-threaded 3.14t
    loaderx 72.2 batches/s 71.1 batches/s
    loaderx-raw 136.7 batches/s 136.0 batches/s
  • Loader, CPU-heavy transform — the one place the free-threaded build matters. A transform runs on the prefetch threads, so the GIL serializes it on standard CPython — the case where worker processes win, and the reason for the free-threaded build. --transform augment (a Python-bound per-sample loop), 4 workers. Torch and Grain have no free-threaded rows, so their GIL worker processes are shown only as context:

    loader CPython 3.14 (GIL) free-threaded 3.14t
    loaderx 33.1 batches/s 41.9 batches/s
    loaderx-raw 36.4 batches/s 51.9 batches/s
    torch 35.9 batches/s
    grain 10.2 batches/s

    Free-threaded wins because the transform parallelizes across the prefetch threads: 1.27x on loaderx and 1.43x on raw at 256 KiB. The margin is bounded by the 64 MiB gather, which remains DRAM-bandwidth-heavy.

Conclusion — why the numbers look like this.

Every hot path is one native call. Zsampler draws a whole batch of indices, zrecord gathers a whole batch and decompresses it, in a single cffi call into Zig with the GIL released and the batch copied straight into its destination buffer. The contenders do the same work one record at a time from Python. That one fact runs through all three tables: the sampler's margin is roughly flat across batch sizes (a batch costs one call either way, so the per-index work is what divides them), and the store gather column is where one-record-per-call costs the most. The speedup is not bought with distribution shortcuts: the IID draw is unbiased like NumPy's (Lemire with rejection, so uniformity costs nothing over a real index space).

The layout matches what a training loader does. zrecord is built for random record access — the record table is indexed in O(1), a dense store gathers at a fixed stride. Array stores are built for contiguous scans, so a scattered batch — exactly what a loader reads — fights their layout. A raw gather is a fan-out across every core, where NumPy's fancy index is one thread; on the multi-core memory subsystem that closes the gap to an uncompressed .npy which does no per-record work at all.

Compression is in the kernel, and there is one codec. zrecord-zstd is not "storage plus a codec": the layout, the multi-core decompress and the GIL-free copy are one path, so turning compression on costs part of a margin, not an order of magnitude — ~4.6 GiB/s here, ~7.1x ahead of the next compressed store at a similar ratio. The modest ratio is the data, not the format: these samples are barely compressible, and on a smooth image set plain zstd reaches 7.6x and the dictionary (at the "max" tier) 16.5x.

The loader gap is architecture, not storage. loaderx uses threads and never ends an epoch, so a step pays no IPC and never waits on an epoch boundary; torch restarts per epoch with worker processes. Here compressed loaderx is 1.1x faster than torch and 4.3x faster than grain; raw loaderx is 2.0x and 8.1x faster, respectively. Storage also differs per loader — each reads from what it was built for — so the loader table is a different comparison than the store table, not a rerun of it. The one crack in the thread model is a CPU-heavy transform, which the GIL serializes — that is exactly what free-threaded Python removes, so loaderx is developed and benchmarked against free-threaded builds first.

What these numbers do not claim. Everything runs with a warm page cache: this measures the access path, not cold storage or disk. on disk is allocated blocks, so zrecord's sparse preallocation is not charged to it while TensorStore's sharding is; both are layout, not compression. write is page-cache ingestion with durability deferred, matching the other backends — zrecord's own durability (sync/close) is a separate, explicit cost that neither this table nor the competition pays. This is a single qualitative pass — the µs-scale sampler timings and the loader batches/s fluctuate with box load (on these 12 cores the compressed loader trails the raw one substantially), so treat the absolute numbers as ballpark and the cross-backend margins as the signal.

Real-data verification: NTU RGB-D skeletons

The tables above are synthetic workloads. As a historical ground-truth check, loaderx was run end to end on NTU RGB-D skeleton data — 114,480 raw .skeleton files, 120 action classes, 25 joints — processed into the ST-GCN N C T V M layout (per-sample (3, 300, 25, 2) float32) for the xsub/xview protocols. The .npy outputs of the standard preprocessing pipeline were treated as ground truth. This verification was not rerun with the synthetic benchmarks above; its throughput is retained as a separate historical 12-core result.

Correctness — the read path is bit-exact against the ground truth:

  • Full scan of all 228,356 records (joint float32 + label int64, all four splits) through DenseStore: byte-for-byte identical to the reference npy.
  • A DataLoader over joint + label + an index stream, run under all three sampler modes (sequential, iid, cyclic): every received batch is bit-exact to the ground truth at its own declared indices, and the streams stay index-aligned.
  • Sampler semantics hold on real index spaces: sequential walks in order, cyclic draws a full cycle without replacement, iid is deterministic per seed.

Storage — zstd on this data:

store on disk ratio
npy (raw float32) 6.4 GB 1.00x
zrecord raw 6.86 GB 1.00x
zrecord zstd 0.79 GB 8.66x
zrecord zstd_dict 0.75 GB 9.13x

Sizes above are for one split (xview/val, 38,132 records); across all four splits the zstd joint stores total 4.97 GB against 41 GB of raw npy (~8x).

Throughput (180 KB per record, warm page cache, 12 physical cores):

path throughput
random-batch gather, zstd store 4.1–4.5 GiB/s
same, npy-mmap fancy indexing 0.6–1.3 GiB/s
DataLoader, 4 prefetch threads 5.7–6.6 GiB/s (123–144 batches/s)

zstd decompression reads ~8x fewer bytes than raw storage, so the compressed store gathers faster than the raw one (zstd 4587 MiB/s vs raw 1792 MiB/s on the same split).

The npy intermediate is optional. from_numpy is chunked append under the hood, so the whole npy staging step can be skipped: parse the skeleton files in parallel and feed DenseStore.append the fixed-shape chunks directly. This writes the store in one pass (no npy, no second read), and the resulting store is byte-for-byte the same size as the from_numpy equivalent — verified record by record against the npy ground truth.

Current Limitations

  • Single-host only; multi-host training is not supported.
  • A single sample must be at most 2 GiB (2^31 bytes). There is no fixed record count: length is a u64 and the record table grows on demand, so how many records a store holds is bounded by its total chunk capacity (up to 2^64 bytes) divided by the average record size — e.g. roughly 2^44 records at 1 MiB each, 2^33 at 2 GiB each.
  • Metadata is read and written as the host's struct layout, so a store carries the host's byte order and is not portable to a machine of the opposite endianness. Every published platform is little-endian, so this only matters if you build for one yourself.

Build

zig build                       # host shared objects, into loaderx/lib/
zig build test                  # Zrecord suite, in both Debug and ReleaseFast
python3 tests/test_loaderx.py   # Python layer, against whichever build is importable
python3 scripts/bench.py        # throughput, against other stores

The Zig side is tested for behaviour only; throughput is measured from Python, through the binding a client actually uses. scripts/bench.py runs in three layers (sampler, store, loader, or all) and skips contenders that are not installed. See Benchmarks.

Publishing

Zig cross-compiles every target from one machine, so releases need no CI matrix:

zig build dist                    # every platform, into zig-out/dist/<wheel tag>/
python3 scripts/build_wheels.py   # one wheel per platform, plus the sdist

The dist directories are named after their Python wheel platform tag, so the tag mapping lives in exactly one place (dist_targets in build.zig). glibc and macOS minimums are pinned in the target triple, which is what makes manylinux_2_17 and macosx_11_0 honest rather than aspirational. Each wheel is checked after packing: it must carry this platform's libraries and no others.

The sdist ships sources only. Installing from it runs zig build through setup.py, so it needs the Zig compiler; wheel users never hit that path.


Zsampler

Index Generator: a high-performance sampler implemented in Zig. Every mode is a pure function of (seed, step), so a run resumes exactly by seeking to a step — there is no epoch to track, in keeping with the endless step-based loader.

from loaderx.zsampler import Sampler

sampler = Sampler(1_000_000, 256, Sampler.Mode.IID, seed=42)
indices = sampler.next()
  1. Sequential — traverse the index space in order through a fixed-size sliding window, treating the space as a circular queue so the tail never truncates.
  2. IID — draw each index uniformly at random with replacement. Unbiased (Lemire with rejection), matching NumPy. Simplest, but coverage is uneven over any short run.
  3. Cyclic — without replacement, round-robin. Each cycle traverses a fresh permutation of the whole index space, so within a cycle every record appears exactly once and no batch repeats an index — coverage is even by construction, which keeps how often each sample is seen uniform. The permutation is a stateless bijection (a small Feistel network over the index space, brought into range by cycle-walking), so a million-record shuffle materializes nothing the size of the dataset and reshuffling each cycle is free. Every batch is exactly batch_size: an endless step-based loader has no final partial batch to special-case, so when batch_size does not divide the length the cycle's remainder is dropped — a different remainder each cycle, since the permutation changes, so every record is still reached over time.

Zrecord

Zrecord is loaderx's typed persistent record store. Its private native runtime is the byte-oriented engine beneath the public DenseStore and RaggedStore contracts.

  1. The engine is an unordered physical store made of N records. Records are independent and carry no ordering, so every index and slice operation is equivalent to a gather.
  2. It hands the Store layer a dense index space: the live records are always exactly 0..N. Deletion preserves that by swapping the tail into the hole, which means an index is stable only until something is deleted. A multi-stream store is purely a Python-layer notion; MultiStore keeps its own stream index table pointing at records.
  3. The engine reads and writes byte ranges. DenseStore and RaggedStore own dtype, shape, allocation and return-value semantics.
  4. The IO model (append | read | delete) is batch-oriented and shape-agnostic. A record is a packed byte range located by an offset; there is no notion of fixed vs variable length, and a single-record operation is just the batch_size == 1 case. Equal length and single records are special cases, not parallel code paths.
  5. The engine owns its temporary memory internally — allocation and release are explicit.
  6. A store has exactly one codec, fixed at creation and immutable afterwards — matching the Store semantics above it, where the schema records a single codec. Every record is compressed and decompressed independently:
| tag   |  name     | algorithm                              |
|-------|-----------|----------------------------------------|
|   0   |  raw      | none                                   |
|   1   |  zstd     | zstd (plain, level 3)                  |
|   2   |  zstdict  | zstd with a trained dictionary (level 19) |
  1. Compression is transparent to the client:
    • Compression runs concurrently across all cores. A compressed store never falls back to raw: each record is stored as the codec's output, even when an incompressible record's frame is larger than its input — write raw if the data does not compress.
    • Decompression writes straight into the caller's destination memory (gather), with no intermediate buffer and no extra copy.
    • zstd is the one transparent codec — faster than Deflate at both ends and a better ratio, so there is no reason to carry a second. It is vendored C, built for every platform by Zig, so the one-wheel-per-platform story is unchanged.
    • zstd_dict additionally trains one dictionary on a sample of the data (stored as dict.zr) and compresses every record against it. Because each record is still independent, random access is unchanged — but the dictionary carries the structure shared across records, which per-record compression cannot see. On many small, similar records (image tiles, token sequences) this is a large win: a smooth-image set that plain zstd takes to 7.6x compresses 16.5x with a large dictionary. The dictionary is loaded once on open and shared, lock-free, across all reader threads. The dictionary size is chosen from the DICT_TIERS presets (see Codec notes).
    • A zstd_dict store needs its dictionary to read every record; a raw or zstd store never touches a dictionary even if one is present.

Persistence format

Native storage is metadata plus chunked data. The extension is the type: .zr files are store-global singletons, .loc files are record-table segments, .chunk files are record data:

zrecord/
  ├── meta.zr      header (global state)
  ├── dict.zr      zstd dictionary (only in dict stores)
  ├── 0.loc        record table segment [0, 2^28)
  ├── 1.loc        record table segment [2^28, 2^29)
  ├── 0.chunk      record data
  └── 1.chunk

Metadata (meta.zr + {id}.loc)

Files are read and written positionally — pread/pwrite at computed offsets, no mmap. The header is a bit-packed struct and a .loc segment an array of 16-byte RecordLocs, exactly as wide as they declare, so a location is one pread/pwrite of 16 bytes at a computed offset and there is no serializer anywhere in the code. Every header field is byte-aligned (no bit fields cross a byte), so the packed header reads as plain memory.

The record table is partitioned into {id}.loc segments so it can grow by appending a segment instead of reserving the maximum. The id→segment mapping is pure arithmetic — seg = idx >> 28, off = (idx & (2^28−1)) × 16 — so a segment needs no per-record bookkeeping. Committed entries are immutable; one shared fd-table lock keeps the in-memory ArrayList stable while readers index it during the rare append that adds another segment.

1. Header — 32 bytes, the whole of meta.zr.

  • magic (ZREC) and version are ordinary fields, so opening a directory that is not a loaderx store fails immediately instead of decoding garbage.
  • codec is the store's one compression method, stamped at creation and immutable — there is no per-record tag anywhere.
  • length (u64) is the total number of records | tail_chunk/tail_offset mark the last write position.
  • There is no chunk count. Chunks are created in order and the frontier is always in the last one, so the store holds exactly chunks 0..=tail_chunk — a count would be a second copy of that fact to keep in sync.
const Codec = enum(u8) { raw = 0, zstd = 1, zstdict = 2, _ };
const Header = packed struct {
    magic: u32, version: u8, codec: Codec,
    tail_chunk: u32, tail_offset: u32, length: u64,
    _reserved: u80,
};

2. Record table — a .loc segment is 2^28 entries of 16 bytes (4 GiB, sparse), indexed directly. Mapping an index to a physical address is what makes random access efficient.

  • chunk_id is the containing chunk | offset is the position within it | phys_length/logic_length are the stored and original sizes. The codec is not here: it is the header's, so a record is stored exactly the way the store is declared.
const RecordLoc = extern struct {
    offset: u32, phys_length: u32, logic_length: u32,
    chunk_id: u32,
};

There is no liveness flag. Every entry below length is live, because deletion swaps the tail into the hole rather than tombstoning.

3. No maximum length. The table grows a .loc segment at a time and the data grows a chunk at a time, so there is no static record-count cap to size against. The real bounds are the field widths — 2^32 chunks of 2^32 bytes (2^64 bytes total), 2^31 (2 GiB) per record — and disk.

Executor

1. Write. Writes are append-only; everything else is offset redirection.

  • Append: compress concurrently → assign physical locations serially → flush concurrently → commit metadata. Append publishes through the page cache and is intentionally lazy; sync() makes the committed frontier durable in payload → record-table → header order and propagates any sync failure. A record never straddles two chunks; one that would not fit rolls over to a fresh chunk.
  • Flush granularity is the store's, not the caller's. The batch size a client passes to append bounds memory only; zrecord's flush stage packs a shard's contiguous records into pwritevs capped at 256 KiB (flush_bytes). This decouples syscall size from batch and record size — a huge batch and a tiny one write identically, which matters because per-syscall writes above a few hundred KiB land on a slow regime on some filesystems/hosts (~1.4 GiB/s vs ~2.1 GiB/s at 256 KiB writes on the reference box). On POSIX the gather is a real pwritev with up to IOV_MAX (1024) iovecs. Windows has no vectored file writeWriteFileGather is async-only, OVERLAPPED, page-aligned — so it falls back to one positional write per record (correct, if not gathered). If Windows grows a synchronous vectored/IOCP file-write story, this is the one place writes there can be pulled up to parity; the byte-budget grouping is already in place, only the syscall differs.
  • Delete: swap the last table entry into the deleted slot and drop the length by one. A batch is applied in descending index order, so each swap pulls from a slot no later target refers to. The index space stays dense — which is what the sampler needs, since it draws uniformly from 0..N and would otherwise keep hitting holes. The deleted record's bytes become garbage.

2. Read. Fill the destination memory concurrently, in place from the Python side (executed on async threads).

  • Committed records are immutable and length is published through an atomic. A gather takes one shared fd-table lock so a rare chunk/segment ArrayList growth cannot invalidate its file handles; record I/O itself stays lock free.
  • Every record is read at the offset its table entry records — the record table is addressed by pure arithmetic, so random access is one pread for the location and one for the bytes, with no batching assumptions about layout. Each shard reads one location and immediately reads/decompresses that record; there is no separate metadata phase or sequential-run special case. Compressed records use a per-shard staging buffer and decode in place into the destination.

3. Concurrency model. Io.Group.async shards work by CPU count, and shards beyond the limit run inline on the calling thread.

  • Shards receive contiguous blocks rather than a strided subset, keeping each worker's reads and writes sequential.
  • Each shard creates one zstd context (ZSTD_CCtx to write, ZSTD_DCtx to read) and reuses it across every record it handles, rather than paying that setup per record. The dictionary (ZSTD_CDict/ZSTD_DDict) is immutable, so all shards share one, lock-free.
  • Decompression writes straight into the caller's destination buffer, so there is no intermediate copy.

4. Garbage collection. Deletion leaves the record's bytes stranded, so space is reclaimed by an offline compact that rewrites the live records in place.

  • Records are visited in physical order — which after swap-last deletes no longer matches table order — and repacked densely in that same order. A record therefore never moves to a higher address than it already had.
  • Two things follow. Writing a record can never land on the bytes of a record not yet moved, so the rewrite is safe in place with no scratch copy of the store. And each table entry can be updated the instant its record lands, so the store is consistent at every point: an interrupted compaction leaves some records moved and the rest where they were, and re-running finishes the job.
  • Records that are already in the right place are skipped, so a store with a small amount of garbage near the end is cheap to compact.
  • Chunk files past the new frontier are closed and deleted.
  • stats() reports live_bytes against chunk_bytes so callers can decide when it is worth running. Note the difference is an upper bound: a record never straddles a chunk boundary, so up to one record's worth per chunk is slack that compaction cannot remove.

5. File access.

  • Metadata: meta.zr (32 bytes) plus one file per .loc segment, each 4 GiB (sparse), opened when the segment is created.
  • Chunk data: 4 GiB files created at init, accessed concurrently through readPositionalAll/writePositionalAll.

Concurrency contract. gather and append are safe to call concurrently from many threads. Append remains single-writer; only rare chunk/segment fd-table growth briefly waits for active gathers. delete and compact mutate the table in ways a reader would observe half-applied, so they require exclusive access to the store.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

loaderx-1.6.2.tar.gz (602.0 kB view details)

Uploaded Source

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

loaderx-1.6.2-py3-none-win_arm64.whl (384.3 kB view details)

Uploaded Python 3Windows ARM64

loaderx-1.6.2-py3-none-win_amd64.whl (527.8 kB view details)

Uploaded Python 3Windows x86-64

loaderx-1.6.2-py3-none-musllinux_1_2_x86_64.whl (458.2 kB view details)

Uploaded Python 3musllinux: musl 1.2+ x86-64

loaderx-1.6.2-py3-none-musllinux_1_2_aarch64.whl (384.9 kB view details)

Uploaded Python 3musllinux: musl 1.2+ ARM64

loaderx-1.6.2-py3-none-manylinux_2_17_x86_64.whl (448.2 kB view details)

Uploaded Python 3manylinux: glibc 2.17+ x86-64

loaderx-1.6.2-py3-none-manylinux_2_17_aarch64.whl (376.8 kB view details)

Uploaded Python 3manylinux: glibc 2.17+ ARM64

loaderx-1.6.2-py3-none-macosx_11_0_x86_64.whl (425.3 kB view details)

Uploaded Python 3macOS 11.0+ x86-64

loaderx-1.6.2-py3-none-macosx_11_0_arm64.whl (363.2 kB view details)

Uploaded Python 3macOS 11.0+ ARM64

File details

Details for the file loaderx-1.6.2.tar.gz.

File metadata

  • Download URL: loaderx-1.6.2.tar.gz
  • Upload date:
  • Size: 602.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.14.7

File hashes

Hashes for loaderx-1.6.2.tar.gz
Algorithm Hash digest
SHA256 f90e031c1c5680591701792e4de692d82f794f7d26c22dfb6fbc1067cb56f0d0
MD5 e5d0c275b3081eda096fefb4fcb6e1d5
BLAKE2b-256 cc74694c3e44556880b7d7d4201b441f16a7fb815b0e8b62b9de5e284d7cd96d

See more details on using hashes here.

File details

Details for the file loaderx-1.6.2-py3-none-win_arm64.whl.

File metadata

  • Download URL: loaderx-1.6.2-py3-none-win_arm64.whl
  • Upload date:
  • Size: 384.3 kB
  • Tags: Python 3, Windows ARM64
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.14.7

File hashes

Hashes for loaderx-1.6.2-py3-none-win_arm64.whl
Algorithm Hash digest
SHA256 21f047c68502fcc2cada8ab2990d5c2ac85039619d0bfe6d8bf74c343d9849f1
MD5 4f9fbc8cce451a6b185cdefd9b5bb554
BLAKE2b-256 c4de307c7aa173cffa2f68813e5b1b59c184b62d42632c4fd632a3c4bd7e50ce

See more details on using hashes here.

File details

Details for the file loaderx-1.6.2-py3-none-win_amd64.whl.

File metadata

  • Download URL: loaderx-1.6.2-py3-none-win_amd64.whl
  • Upload date:
  • Size: 527.8 kB
  • Tags: Python 3, Windows x86-64
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.14.7

File hashes

Hashes for loaderx-1.6.2-py3-none-win_amd64.whl
Algorithm Hash digest
SHA256 37dc1f6b0e0b6b8722efb8203192dd949e2390ce7a36fd54b1df31f4f060dbe5
MD5 aef67fb56fd9ee9bb8e88922e59b0190
BLAKE2b-256 8b3c3e9fc7f5fa2823c74a9aeb031ea7fc048860f721cf3d702ea56b710b7b4f

See more details on using hashes here.

File details

Details for the file loaderx-1.6.2-py3-none-musllinux_1_2_x86_64.whl.

File metadata

File hashes

Hashes for loaderx-1.6.2-py3-none-musllinux_1_2_x86_64.whl
Algorithm Hash digest
SHA256 d31e068e1481fc072686a0538b833422df86be19721a1d2af37a6377aacf2996
MD5 487d45af23b524da70e8e01e43fb9d72
BLAKE2b-256 e5d2257d3b0aff76a4401d5579e8d647ee45afb9db9450a0d942c1dd67dcf2f4

See more details on using hashes here.

File details

Details for the file loaderx-1.6.2-py3-none-musllinux_1_2_aarch64.whl.

File metadata

File hashes

Hashes for loaderx-1.6.2-py3-none-musllinux_1_2_aarch64.whl
Algorithm Hash digest
SHA256 31bdfb89c89d90d205b295aa01708c8bcce668003fec522e9ba0a2c5749b0a6c
MD5 589f05a18fe002d03649e62247a39aa0
BLAKE2b-256 d5758928673ede626834fd29aa8f1fc38ae984d803ca4d2d963c4264c52e30fd

See more details on using hashes here.

File details

Details for the file loaderx-1.6.2-py3-none-manylinux_2_17_x86_64.whl.

File metadata

File hashes

Hashes for loaderx-1.6.2-py3-none-manylinux_2_17_x86_64.whl
Algorithm Hash digest
SHA256 70b92b2d7ba554145a74e381cc073ec05c1747c06317b79defc60230c496c5ea
MD5 3ec782a49863a8a3c2082e15c50bc79c
BLAKE2b-256 adfa18cbc518419b1864977c91eee5e847aed6b66c25d7cca810a434182caf97

See more details on using hashes here.

File details

Details for the file loaderx-1.6.2-py3-none-manylinux_2_17_aarch64.whl.

File metadata

File hashes

Hashes for loaderx-1.6.2-py3-none-manylinux_2_17_aarch64.whl
Algorithm Hash digest
SHA256 c33d4a097d78f5caa7f3e80b6604fb1b31730dbf393a7abc3fcfd87286596653
MD5 09fdd4cb266edb2f83e3d72bf503e2d7
BLAKE2b-256 0145de2a015b787af3490c52b6e4214301ca61bec9ddf7f48008098e6ce16f80

See more details on using hashes here.

File details

Details for the file loaderx-1.6.2-py3-none-macosx_11_0_x86_64.whl.

File metadata

File hashes

Hashes for loaderx-1.6.2-py3-none-macosx_11_0_x86_64.whl
Algorithm Hash digest
SHA256 cf06837d350fe1dd9304c8af4458a8beddeb0902d7858c99c654ff17a75455a2
MD5 4255a2589392371f0002b86c952b358a
BLAKE2b-256 161da6280b0728b57d108f03f0ea0fccd7d6d37fbcc0242c1c03eebab7a2b80d

See more details on using hashes here.

File details

Details for the file loaderx-1.6.2-py3-none-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for loaderx-1.6.2-py3-none-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 b85e4ece4c0e0005d734dd0b79f23379b092f02fe5c1820fdb6c6931486691f6
MD5 d5bd215877538a405c4a898f969f3940
BLAKE2b-256 afa07c16ef60214d2daf50672a41cc7e3a1a4fb786df75c123510241a8995bc5

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page