2 projects
tokbin
Pretraining data for language models: tokenize text once into resumable, verifiable binary shards (full install: write and read). Pretraining only; not for SFT or chat data.
tokbin-core
Pretraining data for language models: tokenize text once into resumable, verifiable binary shards and read random windows through memmap. Pretraining only; not for SFT or chat data.