Skip to main content

tokbin

Pretraining data for language models. tokbin tokenizes a text corpus once into binary shards, reads random windows through numpy.memmap without loading the corpus into memory, mixes sources by weight, packs datasets for transfer and verifies their integrity.

tokbin supports pretraining only: plain text packed into a continuous token stream. It is not a tool for fine-tuning, SFT or chat data.

Status: alpha. The API and the on-disk format may still change before 1.0.

Is tokbin for you?

tokbin is a narrow tool: it prepares pretraining data for language models. It was built for one workflow: tokenize a large text corpus ahead of time on a cheap machine, move it to a rented GPU server, and start training at once, without paying for idle GPUs.

Use it if you pretrain or continue pretraining on plain text that is too large for memory, prepare data on one machine and train on another, need writes that survive crashes, or mix several sources by weight.

Skip it if you fine-tune on chats or instructions (no chat templates, loss masks or padding), need attention masks at document borders, use a tokenizer that is not a Hugging Face tokenizer.json, work with non-text data, need distributed writing or streaming from object storage, or your data fits in memory. A 30-line numpy.memmap script may be all you need.

The details, a comparison with a hand-written script and the measured costs are in Is tokbin for you?

Documentation

The full documentation, with runnable examples and their output, is in docs/:

Installation

pip install tokbin          # full: write and read
pip install tokbin-core     # lightweight: read, verify, unpack

tokbin-core depends only on numpy. The tokbin meta-package adds tokenizers, which is required for writing datasets.

Interrupted writes

A write goes to <name>.partial/ and becomes visible only when it is complete. It saves a checkpoint after every closed shard; after a crash or Ctrl+C it continues from there:

ds.write("web", docs(), tokenizer, resume=True)

docs() must yield the same documents in the same order as before. At most one shard of work is lost; the result is identical to an uninterrupted write.

Mixtures

from tokbin import Mixture

mix = Mixture.from_config("corpus")    # weights from corpus/mix.json
x = mix.batch(batch_size=32, block_size=1024, rng=np.random.default_rng(0))

For torch, tokbin.adapters.torch has WindowDataset and MixtureDataset (pip install 'tokbin-core[torch]').

Command line

tokbin build corpus/web --from-jsonl data/web --tokenizer gpt2/tokenizer.json
tokbin build corpus/web --resume            # after Ctrl+C or a crash: same settings
tokbin ls corpus                            # sources, states and sizes
tokbin info corpus                          # tokens, tokenizer, mixture, remarks
tokbin info corpus/web                      # dtype, splits, special tokens of one source
tokbin status corpus/web                    # write state, unfinished writes included
tokbin verify corpus/web                    # sha256 of every shard and all index files
tokbin clean corpus/web                     # remove an unfinished write (corpus/web.partial)
tokbin pack corpus/web --tar                # corpus/web.tbpack.tar: compressed, with sha256 manifest
tokbin unpack web.tbpack.tar --into corpus  # unpack, verify, publish
tokbin migrate corpus/web                   # convert to the current format schema (a copy)
tokbin rm corpus/web --yes                  # delete a source
tokbin doctor                               # version, mode, optional packages

Every command accepts --json (machine-readable output), --no-color and --strict (warnings fail with exit code 1). Exit codes: 0 success, 1 error, 2 wrong usage, 3 integrity check failed, 4 missing dependency, 130 interrupted. Error codes are listed in docs/errors.md.

Development

uv sync --all-packages
git config core.hooksPath scripts/hooks   # enable the pre-commit quality gate

The pre-commit hook runs ruff, mypy, the test suite and a guard against non-English text in staged files. The full OS x Python matrix runs in GitHub Actions on demand and on release tags.

License

MIT

Metadata

Release files for tokbin 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for tokbin 0.1.0
File Size Uploaded
tokbin-0.1.0.tar.gz 3.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for tokbin 0.1.0
File Interpreter ABI Platform
tokbin-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 7.9 kB

Release files / tokbin-0.1.0.tar.gz

Download URL tokbin-0.1.0.tar.gz
Size 3.9 kB
Tags Source
SHA-256 checksum
How to use checksums
5d32bd58ddcde2f5f9e3774bab6f72322c62ca4fd6d2e4b4d663211ceee506e5
BLAKE2b-256 checksum
How to use checksums
5e36c7269c22ec39fffd2efb7680a4eee9e491a159ea2f0487bb09d142d53a1f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.

Transparency log

Release files / tokbin-0.1.0-py3-none-any.whl

Download URL tokbin-0.1.0-py3-none-any.whl
Size 4.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
cb975da607e9b10ea9216f29fcd0c2a1c45b82c3c5fe64912a271ea0f448259c
BLAKE2b-256 checksum
How to use checksums
82f1f4e257de22c012b1b81ceb5e6cb7388f5ccb733b6bffd623a4869b30e22d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page