DatasetRT
DatasetRT is a correctness-first dataset cache for ML training loops.
It gives you a deterministic, immutable cache on disk, backed by a Rust runtime and exposed through a small Python API. You keep your model code in PyTorch, JAX, TensorFlow, NumPy, or plain Python; DatasetRT handles cache integrity, metadata, sampling weights, and repeatable iteration without becoming another framework.
Authorship
Created by Vadym Stupakov vadim.stupakov@gmail.com.
Why ML Users Need This
Dataset bugs are expensive. A silent shuffle change, corrupt shard, mismatched metadata row, or weight vector applied to the wrong sample can waste training runs and make experiments impossible to reproduce.
DatasetRT is built around one rule:
If it affects correctness, Rust owns it.
Rust owns:
- immutable cache publication
- manifest and checksum validation
- metadata schema validation
- shard offsets and index generation
- deterministic weighted sampling
- iterator state
- bounded reader/writer prefetch and worker pools
- weight table validation
Python stays thin and ergonomic. It describes your source data and receives bytes plus metadata back.
Quickstart
from pathlib import Path
import polars as pl
from dataset_rt import (
CacheInput,
CachedDataset,
ReaderConfig,
ShardCompression,
WriterConfig,
)
class Images:
name = "train_images"
def __iter__(self):
for sample_id, image_bytes, label in load_my_images():
yield CacheInput(
data=image_bytes,
metadata={"sample_id": sample_id, "label": label},
)
dataset = CachedDataset.from_cache_sources(
Images(),
Path("cache"),
reader_config=ReaderConfig(seed=42, prefetch_size=64, num_workers=4, shuffle=True),
writer_config=WriterConfig(
prefetch_size=64,
num_threads=4,
shard_compression=ShardCompression(algo="none", ratio=1.0),
),
)
for sample in dataset:
image = decode_image(sample.data) # domain decoding stays in Python
label = sample.metadata["label"]
CachedDataset.from_cache_sources creates missing caches, reuses valid existing caches, and returns a ready-to-iterate dataset. The cache directory argument is always a base cache directory; Rust writes each source under base_cache_dir / name_hash.
PyTorch
When PyTorch is installed, turn the same DatasetRT object into a sized IterableDataset:
torch_dataset = dataset.to_torch_iterable_dataset()
loader = torch.utils.data.DataLoader(torch_dataset, batch_size=None)
for sample in loader:
image = decode_image(sample.data)
label = sample.metadata["label"]
The adapter does not decode payloads or add a Torch dependency to DatasetRT. It yields CachedSample values and reports len(torch_dataset).
Metadata-Aware Weights
Weights are not a loose list that can drift out of alignment. DatasetRT exposes them as a Polars table with sample identity and metadata:
weights = dataset.weight_table()
rare = weights.with_columns(
pl.when(pl.col("label") == "rare_class")
.then(5.0)
.otherwise(1.0)
.alias("weight")
)
dataset.set_weight_table(rare)
The table contains:
cache_id | sample_id | <metadata columns...> | weight
Rust validates that every physical (cache_id, sample_id) appears exactly once and that every weight is positive and finite.
Multiple Sources
dataset = CachedDataset.from_cache_sources(
[TrainImages(), SyntheticImages(), HardNegatives()],
Path("cache"),
reader_config=ReaderConfig(seed=123),
writer_config=WriterConfig(prefetch_size=128, num_threads=8),
)
If one source fails during a multi-source write, DatasetRT cleans up caches created earlier in that call. A loaded dataset means every cache was validated from its manifest.
Storage Layout
cache/
train_images_<hash>/
manifest.json
metadata.arrow
index.bin
shards/
000000.bin
000001.bin
Metadata is stored separately from payload bytes. This keeps sampling, filtering, auditing, and weight editing independent of domain payload decoding.
What DatasetRT Does Not Do
DatasetRT does not decode JPEGs, PNGs, tensors, or framework-specific objects in the Rust core.
The core returns payload bytes. Your Python code or optional adapters can decode those bytes into tensors, arrays, images, token sequences, or any other domain object.
DatasetRT v0.1 also intentionally supports only:
- bytes-like payloads:
bytes,bytearray,memoryview - primitive metadata:
bool,int,float,str ShardCompression(algo="none", ratio=1.0)
The common Rust zstd crate uses C bindings, so zstd compression is not enabled for the first stable version.
Documentation
- Architecture
- Python API
- Storage Format
- Runtime Model
- Determinism
- Serialization Boundary
- Development
- Release
- Changelog
Release Builds
Release wheels are built locally with just release-all. Linux wheels use Zig cross-compilation; macOS arm64 and x86_64 wheels build on the local host. GitHub Actions stays checks-only.
Wheels use Python's stable ABI (cp310-abi3) and support Python 3.10 through 3.13.
Status
DatasetRT is at foundational v0.1 architecture. The core cache lifecycle, immutable storage, metadata, deterministic weighted sampling, Rust-owned reader/writer prefetching, and Polars weight table are in place.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distributions
Built Distributions
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file dataset_rt-0.1.1-cp310-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl.
File metadata
- Download URL: dataset_rt-0.1.1-cp310-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
- Upload date:
- Size: 880.0 kB
- Tags: CPython 3.10+, manylinux: glibc 2.17+ ARM64
- Uploaded using Trusted Publishing? No
- Uploaded via:
uv/0.11.17 {"installer":{"name":"uv","version":"0.11.17","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
b9c6376aa25728b32240855a99c2df0dea8686f93926767289f6fa3c08bd775c
|
|
| MD5 |
d4f956f86e486d73f2b2b4e32a0e96a8
|
|
| BLAKE2b-256 |
dd5f8913478c16dcb5eee433204d140c978c4c7bac4c59f0a40963c3483e3370
|
File details
Details for the file dataset_rt-0.1.1-cp310-abi3-macosx_11_0_arm64.whl.
File metadata
- Download URL: dataset_rt-0.1.1-cp310-abi3-macosx_11_0_arm64.whl
- Upload date:
- Size: 904.8 kB
- Tags: CPython 3.10+, macOS 11.0+ ARM64
- Uploaded using Trusted Publishing? No
- Uploaded via:
uv/0.11.17 {"installer":{"name":"uv","version":"0.11.17","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
613a9c16425a97a2536cbf74467878933dab0b55e8605d7e148f133818d9cc15
|
|
| MD5 |
7574283d888ed776d9b85c8709434dea
|
|
| BLAKE2b-256 |
3bcca76af4600a3d77c625e3d9471e738269cb2aee538167ce30d332a4dc70aa
|