Skip to main content

tokbin

Pretraining data for language models. tokbin tokenizes a text corpus once into binary shards, reads random windows through numpy.memmap without loading the corpus into memory, mixes sources by weight, packs datasets for transfer and verifies their integrity.

tokbin supports pretraining only: plain text packed into a continuous token stream. It is not a tool for fine-tuning, SFT or chat data.

Status: alpha. The API and the on-disk format may still change before 1.0.

Is tokbin for you?

tokbin is a narrow tool: it prepares pretraining data for language models. It was built for one workflow: tokenize a large text corpus ahead of time on a cheap machine, move it to a rented GPU server, and start training at once, without paying for idle GPUs.

Use it if you pretrain or continue pretraining on plain text that is too large for memory, prepare data on one machine and train on another, need writes that survive crashes, or mix several sources by weight.

Skip it if you fine-tune on chats or instructions (no chat templates, loss masks or padding), need attention masks at document borders, use a tokenizer that is not a Hugging Face tokenizer.json, work with non-text data, need distributed writing or streaming from object storage, or your data fits in memory. A 30-line numpy.memmap script may be all you need.

The details, a comparison with a hand-written script and the measured costs are in Is tokbin for you?

Documentation

The full documentation, with runnable examples and their output, is in docs/:

Installation

pip install tokbin          # full: write and read
pip install tokbin-core     # lightweight: read, verify, unpack

tokbin-core depends only on numpy. The tokbin meta-package adds tokenizers, which is required for writing datasets.

Interrupted writes

A write goes to <name>.partial/ and becomes visible only when it is complete. It saves a checkpoint after every closed shard; after a crash or Ctrl+C it continues from there:

ds.write("web", docs(), tokenizer, resume=True)

docs() must yield the same documents in the same order as before. At most one shard of work is lost; the result is identical to an uninterrupted write.

Mixtures

from tokbin import Mixture

mix = Mixture.from_config("corpus")    # weights from corpus/mix.json
x = mix.batch(batch_size=32, block_size=1024, rng=np.random.default_rng(0))

For torch, tokbin.adapters.torch has WindowDataset and MixtureDataset (pip install 'tokbin-core[torch]').

Command line

tokbin build corpus/web --from-jsonl data/web --tokenizer gpt2/tokenizer.json
tokbin build corpus/web --resume            # after Ctrl+C or a crash: same settings
tokbin ls corpus                            # sources, states and sizes
tokbin info corpus                          # tokens, tokenizer, mixture, remarks
tokbin info corpus/web                      # dtype, splits, special tokens of one source
tokbin status corpus/web                    # write state, unfinished writes included
tokbin verify corpus/web                    # sha256 of every shard and all index files
tokbin clean corpus/web                     # remove an unfinished write (corpus/web.partial)
tokbin pack corpus/web --tar                # corpus/web.tbpack.tar: compressed, with sha256 manifest
tokbin unpack web.tbpack.tar --into corpus  # unpack, verify, publish
tokbin migrate corpus/web                   # convert to the current format schema (a copy)
tokbin rm corpus/web --yes                  # delete a source
tokbin doctor                               # version, mode, optional packages

Every command accepts --json (machine-readable output), --no-color and --strict (warnings fail with exit code 1). Exit codes: 0 success, 1 error, 2 wrong usage, 3 integrity check failed, 4 missing dependency, 130 interrupted. Error codes are listed in docs/errors.md.

Development

uv sync --all-packages
git config core.hooksPath scripts/hooks   # enable the pre-commit quality gate

The pre-commit hook runs ruff, mypy, the test suite and a guard against non-English text in staged files. The full OS x Python matrix runs in GitHub Actions on demand and on release tags.

License

MIT

Metadata

Release files for tokbin-core 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for tokbin-core 0.1.0
File Size Uploaded
tokbin_core-0.1.0.tar.gz 159.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for tokbin-core 0.1.0
File Interpreter ABI Platform
tokbin_core-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 297.9 kB

Release files / tokbin_core-0.1.0.tar.gz

Download URL tokbin_core-0.1.0.tar.gz
Size 159.2 kB
Tags Source
SHA-256 checksum
How to use checksums
130f4eafb9e941165070315e0eca21226dcb612e8e579f904887375153ae5987
BLAKE2b-256 checksum
How to use checksums
a1fed61fd99ea9854eb1a53d4cb70d574dd81bb5d3a1a89d1e5a3f071ecc94f3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.

Transparency log

Release files / tokbin_core-0.1.0-py3-none-any.whl

Download URL tokbin_core-0.1.0-py3-none-any.whl
Size 138.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
3f25f911ebd2a58b5c9b71e9f38010e15cb6e9e9cb974e7a9a28c51c28d3ebe8
BLAKE2b-256 checksum
How to use checksums
8551c9541d3b7b1c4de34149bd486f48147c666d9d4170c46f6dd06f673ec821
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page