Fast, flexible, elastically deterministic data loading for foundation model training
zephon.io · User Guide · Launch blog post · Paper · PyPI
Zephon is a high-performance data loading library that separates what you train on from how data is loaded and transformed. It supports online processing (tokenization, sequence packing, filtering, dynamic mixtures) rather than requiring expensive offline preprocessing, while guaranteeing deterministic, reproducible output regardless of parallelism or hardware topology.
We built Zephon at DatologyAI because running data experiments is a lot of what we do, and an experiment that compares two datasets is only useful if nothing else about the two training runs changes. In practice, the data loader is often the part that doesn't hold still: if a job gets preempted and comes back on fewer GPUs (which happens more often than any of us would like), many loaders will either refuse to resume or quietly start feeding it different data. That's a manageable problem for datasets you can index into directly. It's much harder for pipelines that pack sequences, since what ends up in each packed sequence depends on every sample that came before it. Zephon was designed from the start to keep that kind of pipeline deterministic, so the number of GPUs you happen to have on a given day is one less thing to worry about.
Installation
Zephon requires Python 3.10 or newer.
uv pip install zephon
Most of what Zephon can read is behind optional extras, so you only install what you use:
| Extra | What it adds |
|---|---|
zephon[cloud] |
Reading from S3, GCS, and Azure (s3://, gs://, az:// paths) |
zephon[hf] |
Hugging Face datasets (hf:// paths), plus Parquet |
zephon[parquet] |
Parquet shards |
zephon[streaming] |
MosaicML Streaming / MDS shards |
zephon[litdata] |
LitData shards |
zephon[vortex] |
Vortex shards (Python 3.11+) |
zephon[ray] |
The experimental Ray runner |
For example, to train on Parquet data in S3: uv pip install "zephon[cloud,parquet]".
Tokenization uses Hugging Face tokenizers, so you'll also want uv pip install transformers
if you're following the example below.
A First Pipeline
A Zephon pipeline has three parts: a Dataset that describes your data, a WorkSource
that decides what to train on and in what order, and a Pipeline of operators that turns
samples into training batches.
from zephon import Pipeline
from zephon.io import Dataset
from zephon.work import MixtureSpec, StaticMixtureWorkSource
# 1. Describe the data: a directory of JSONL, Parquet, MDS, LitData, or Vortex shards
dataset = Dataset.from_path("train", "/data/train/")
# 2. Decide what to train on, and in what order
ws = StaticMixtureWorkSource(
datasets=[dataset],
mixture=MixtureSpec({"train": 1.0}),
seed=42,
)
# 3. Turn samples into training batches
pipeline = (
Pipeline(ws)
.decode_text()
.tokenize(tokenizer_id="gpt2", field="text", padding=True)
.batch(microbatch_size=2)
)
for batch in pipeline:
train_step(batch.to_training())
Mixing datasets is a matter of adding them to the WorkSource with the proportions you
want:
ws = StaticMixtureWorkSource(
datasets=[
Dataset.from_path("fineweb", "s3://my-bucket/fineweb/"),
Dataset.from_path("dclm", "s3://my-bucket/dclm/"),
],
mixture=MixtureSpec({"fineweb": 0.7, "dclm": 0.3}),
seed=42,
)
A Pipeline is a plain Python iterable, so it works with whatever training loop you
have. It will also run inside a PyTorch DataLoader or torchdata StatefulDataLoader if
your framework insists on one, although we recommend iterating over it directly.
Checkpointing and Elastic Resume
You save and restore a pipeline's position alongside your model checkpoint:
state = pipeline.checkpoint() # store this with the model and optimizer state
pipeline.restore(state) # on resume, before iterating
for batch in pipeline:
...
Resuming on a different number of GPUs comes down to one setting that you pick at the
start of the run: the number of lanes (canonical_replicas) that Zephon divides the
data into. Each data-parallel group owns an equal share of the lanes, so you can run or
resume on any data-parallel size that divides the lane count evenly, as long as each
optimizer step consumes a multiple of canonical_replicas batches. For example, a
lane count of 32 lets you run on any of 8, 16, or 32 data-parallel groups. Only the data-parallel degree
counts here; tensor, pipeline, and context parallelism don't affect it.
What Zephon guarantees is that every global batch contains the same data after a resume. That's not quite the same as promising identical losses or weights, which also depend on things like the order of reductions and which kernels get selected, and those are outside of anything a data loader can control.
Training Framework Integrations
We maintain reference integrations for two training frameworks, each in its own fork with a README, launch commands, and smoke tests:
- TorchTitan: datologyai/torchtitan-zephon
- Megatron-LM: datologyai/Megatron-LM-zephon
The Training Integrations chapter of the User Guide explains the decisions behind them, which is the place to start if you want to connect Zephon to a different framework.
How It Works
Under the hood, Zephon is organized into three layers:
WorkSource what to train on (datasets, mixtures, shuffling)
│ emits lightweight pointers: (dataset, shard, sample)
▼
Pipeline how to process it (decode, tokenize, pack, batch, ...)
│ a chain of operators
▼
Engine where and when to run it (threads, processes, queues)
deterministic scheduling with backpressure
The WorkSource produces a deterministic sequence of sample pointers, working mostly from
metadata like shard listings and sample counts rather than the samples themselves.
The Engine compiles the Pipeline into concurrent stages connected by bounded queues, so
fetching, processing, and your training step all overlap, and it does this without letting
the amount of parallelism change the order of the output. If you're curious what it came
up with, pipeline.explain() will show you the compiled plan.
Documentation
The Zephon User Guide covers all of this in much more detail:
- Why Zephon?: elastic determinism and why it matters for experiments
- Basic Concepts: Datasets, WorkSources, and Pipelines
- Working with Datasets: shard formats, storage backends, and the shard cache
- WorkSources: mixing, shuffling, and repeating samples
- Pipelines: operators, packing, distributed training, and checkpointing
- Training Integrations: the TorchTitan and Megatron-LM reference integrations
- API Reference
Status
Zephon is beta software, and the Ray runner in particular is experimental. If something doesn't work the way this README or the User Guide says it should, please open an issue.
Citing Zephon
If you use Zephon in your research, please cite our paper:
@misc{boether2026zephon,
title = {Zephon: Elastic Determinism for Online, Stateful Foundation Model Data Loading Pipelines},
author = {B{\"o}ther, Maximilian and Wills, Josh and Robroek, Ties and Xu, Sonnet and
Burstein, Paul and Zayas, Daniel and Blakeney, Cody and Joshi, Siddharth and
Yin, Haoli and Adiga, Rishabh and Mongstad, Haakon and Merrick, Luke and
Maini, Pratyush and Morcos, Ari and Leavitt, Matthew and Klimovic, Ana and
Gaza, Bogdan},
year = {2026},
eprint = {2610.03087},
archivePrefix = {arXiv},
primaryClass = {cs.LG},
url = {https://arxiv.org/abs/2610.03087}
}
License
Zephon is released under the Apache 2.0 License.
Metadata
Release files for zephon 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| zephon-0.1.0.tar.gz | 2.6 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| zephon-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 3.2 MB
Release files / zephon-0.1.0.tar.gz
| Download URL | zephon-0.1.0.tar.gz |
|---|---|
| Size | 2.6 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
67e353e0b2d02438e1e5287ebe7e37eb2165f82f48a67124cd4ae2575ee8c714
|
|
BLAKE2b-256 checksum How to use checksums |
c4a0e41b0dea5167085dbed3a3b85baeae59b2562af6caa81a343505d9eab280
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 7, 2026.
Transparency logRelease files / zephon-0.1.0-py3-none-any.whl
| Download URL | zephon-0.1.0-py3-none-any.whl |
|---|---|
| Size | 618.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
5d70412b9f218a3de5e849ab64d266914dd396db0a52712166eae0b082be1ac5
|
|
BLAKE2b-256 checksum How to use checksums |
4eefa659edb3c8a20de33f95693f6db9a408ddaba0125fc1e4e65e7da99bd0c6
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 7, 2026.
Transparency log