Skip to main content

Tokenize your documents at GB/s

Project description

Gigatoken

~300-1000x faster than HuggingFace's tokenizers, drop-in replacement.

Tokenize your text data at GB/s!

GPT-2 Speedup

Note that both HF tokenizers and tiktoken are already running multithreaded Rust!

What is Gigatoken?

Gigatoken is the fastest tokenizer for language modeling. It supports a wide range of CPU hardware, and nearly all commonly used tokenizers.

Installation

pip install gigatoken

Usage

Gigatoken can be used with its own API, or in compatibility mode with HuggingFace Tokenizers or Tiktoken.

Compatibility Mode (Easiest)

import gigatoken as gt

# Minimum change from existing HuggingFace tokenizers usage (compatibility mode)
hf_tokenizer = ...
tokenizer = gt.Tokenizer(hf_tokenizer).as_hf()

# tokenizer can be used in the same contexts as hf_tokenizer
tokens = tokenizer.encode_batch(["This is a test string", "And here is another"])

A substantial amount of effort has been put into making sure the outputs match exactly with what you would get with HuggingFace Tokenizers in this setting, but this is at a non-negligible cost to performance. You can still expect way faster performance across the board, but not quite the 1000x you will get with the Gigatoken API.

Gigatoken API (Fastest)

import gigatoken as gt

tokenizer = gt.Tokenizer("Qwen/Qwen3-8B")  # Accepts HF model names
file_source = gt.TextFileSource(["owt_train.txt"], separator=b"<|endoftext|>")
tokens = tokenizer.encode_files(file_source)

Using the Gigatoken API lets the Rust implementation read data directly, and skips as much overhead as possible while allowing for maximum parallelism. Keep in mind that passing Python data structures through this API still incurs the overhead of reading from Python.

FAQ

Q: Did you just way over-optimize for a specific CPU and tokenizer? How is it so fast?

No, I way over-optimized for every combination of these! The results are very consistent across CPUs (modern x86 and ARM), and across specific tokenizers.

The major improvements are in optimizing heavily an implementation that usually is outsourced to a Regex engine (pretokenization) using SIMD and other tricks, as well as heavily optimizing caching of pretoken mappings (if a word has been seen before, look it up its encoded tokens efficiently). In addition, interactions with Python are minimized, and threads are minimally interacting with each other.

Q: How can I quickly check if my tokenizer is supported?

You can try it out without installing anything! The following command will validate and time tokenization for a given HuggingFace model repo:

# Download your data
wget https://huggingface.co/datasets/stanford-cs336/owt-sample/resolve/main/owt_train.txt.gz  # Just an example!
gunzip owt_train.txt.gz
uvx --with tokenizers gigatoken bench 'openai-community/gpt2' owt_train.txt \
    --validate --doc-separator "<|endoftext|>"
      cpu: Apple M4 Max, 16 cores
gigatoken:    1.432 s |   11920.51 MB at  8327.05 MB/s |  2701.65 Mtok at 1887.23 Mtok/s
       hf:   16.250 s |     100.00 MB at     6.15 MB/s |    22.76 Mtok at    1.40 Mtok/s
gigatoken is 1353.13x faster than hf
validation OK: 20401 documents match
      cpu: AMD EPYC 9565 72-Core Processor, 144 cores, 2 sockets
gigatoken:    0.584 s |   11920.51 MB at 20412.35 MB/s |  2701.65 Mtok at 4626.23 Mtok/s
       hf:    3.738 s |     100.00 MB at    26.75 MB/s |    22.76 Mtok at    6.09 Mtok/s
gigatoken is 763.08x faster than hf
validation OK: 20401 documents match

This example uses the train sample from this dataset. You can see help for these flags with uvx gigatoken bench --help. You might need to run these twice on macOS to get a good reading, since the first run will always perform a security scan.

At the rates we see on the EPYC CPU, you could tokenize the entirety of Common Crawl (often considered to be the entire internet, 130 trillion tokens) in just under 8 hours!

Q: I've found a mismatch/slow use-case, is this expected?

Most likely not! Despite reasonably wide testing I don't have every use-case on hand, so please report anything you find in a GitHub Issue so I can address it as soon as possible.

Benchmarks

Encoding throughput on owt_train.txt (11.9 GB) — Apple M4 Max (16 cores)

Best of 3 interleaved rounds, one fresh process per measurement, all libraries with parallelism enabled. gigatoken encodes the whole file un-split; HuggingFace tokenizers (encode_batch_fast) gets the first 100 MB and tiktoken (encode_ordinary_batch) the first 1 GB, both presplit on <|endoftext|>. tiktoken rows exist only for tokenizers with official support.

Tokenizer gigatoken HF tokenizers tiktoken vs HF vs tiktoken
GPT-2 8.79 GB/s 6.9 MB/s 62.8 MB/s 1,268× 140×
Nemotron 3 7.82 GB/s 10.9 MB/s 715×
Phi-4 7.76 GB/s 7.7 MB/s 1,012×
Llama 3 / 3.1 / 3.2 7.60 GB/s 11.2 MB/s 676×
OLMo 2 / 3 7.56 GB/s 5.8 MB/s 1,299×
Llama 3.3 7.50 GB/s 15.7 MB/s 479×
Phi-4-mini 6.97 GB/s 7.2 MB/s 964×
Kimi K2 6.88 GB/s
Llama 4 6.81 GB/s 11.6 MB/s 590×
Qwen 2 / 2.5 6.37 GB/s 5.8 MB/s 1,105×
Qwen 3 6.36 GB/s 6.9 MB/s 918×
Qwen 3.5 / 3.6 6.31 GB/s 6.3 MB/s 994×
GPT-OSS 6.20 GB/s 20.2 MB/s 87.2 MB/s 306× 71×
GLM 4 6.17 GB/s 15.8 MB/s 392×
DeepSeek V3 / R1 / V4 5.68 GB/s 7.2 MB/s 788×
GLM 5 5.55 GB/s 12.2 MB/s 456×
ModernBERT 2.64 GB/s 5.8 MB/s 452×
Mistral 7B v0.3 1.99 GB/s 95.1 MB/s 21×
Gemma 4 1.82 GB/s 85.2 MB/s 21×
CodeLlama 1.73 GB/s 80.2 MB/s 22×
TinyLlama / Phi-3 (Llama 2) 1.69 GB/s 80.1 MB/s 21×
Gemma 1 1.42 GB/s 85.7 MB/s 17×
Gemma 3 1.38 GB/s 82.2 MB/s 17×

The slowest rows are the SentencePiece-based tokenizers (Mistral 7B and below), which remain more expensive to encode than byte-level BPE even with gigatoken's internal SP parallelism; ModernBERT is byte-level BPE with a heavier pretokenizer than the GPT-2 family.

Each row is one distinct tokenizer (identical vocab/merges/pretokenizer), measured on a representative repo. Rows whose tokenizer is shared beyond their own name (verified by matching tokenizer definitions across the local HF model cache) cover:

  • Nemotron 3 — Nemotron 3 Nano, Super, and Ultra
  • Llama 3 / 3.1 / 3.2 — Llama 3 / 3.1 / 3.2, DeepSeek-R1-Distill-Llama, Hermes 3, Saiga, and other Llama-3 finetunes
  • Llama 3.3 — Llama 3.3, Llama-3.1-Nemotron-Nano-VL, SmolLM3, Kanana 1.5, jina-embeddings-v5, Ultravox
  • Phi-4-mini — Phi-4-mini and Phi-4-multimodal
  • Kimi K2 — Kimi K2 / K2.5 / K2.6 / K2.7, Kimi-Linear, Kimi-VL, Moonlight
  • Qwen 2 / 2.5 — Qwen 2 and 2.5 (incl. Coder and VL), Qwen3-Coder, Qwen3-VL, DeepSeek-R1 Qwen distills, MiMo V2.5, MiniCPM-o 2.6, InternVL3
  • Qwen 3 — Qwen 3 (incl. Embedding and Reranker), Qwen2.5-Omni, Qwen3-VL-Embedding, MiMo V2.5 Pro, jina-reranker-m0, pplx-embed, MOSS-TTS, Zeta
  • GLM 4 — GLM 4.1V, 4.5, and 4.7
  • DeepSeek V3 / R1 / V4 — DeepSeek V3 / V3.1 / V3.2, R1, V4 Flash and Pro, DeepSeek-VL2
  • GLM 5 — GLM 5 / 5.2 and GLM-4.7-Flash
  • Gemma 4 — Gemma 4 (dense, MoE, and E-series) and DiffusionGemma
  • TinyLlama / Phi-3 (Llama 2) — TinyLlama, Phi-3-mini, Phi-3.5-mini and Phi-3.5-vision (the Llama 2 vocab)
  • Gemma 3 — Gemma 3 (270M–27B) and EmbeddingGemma

Citation

If you use Gigatoken in your research, please cite it as:

@software{roed2026gigatoken,
  author = {Marcel R{\o}d},
  title = {{G}igatoken: SIMD and Cache Hierarchies for 1000x Faster Byte-Pair Encoding Tokenization on Modern CPUs},
  url = {https://github.com/marcelroed/gigatoken},
  year = {2026},
}

AI Use Disclosure A majority of this code base was crafted by hand without any use of AI (which can be seen from the project's Git history). In the final stages of the project, AI was used to assist:
  • Implementing the user-facing API
  • Widening of compatibility, for instance generalizing and porting the pretokenizer implementations to support more tokenizers, less interesting features like padding/truncation/unicode normalization
  • Porting SIMD strategies between AVX512/AVX2/NEON
  • Final profiling stages and the last ~4x worth of performance from eliminating branching and improving the pretoken cache hierarchy
  • Refactoring and code reuse

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

gigatoken-0.8.0.tar.gz (1.3 MB view details)

Uploaded Source

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

gigatoken-0.8.0-cp310-abi3-win_arm64.whl (4.4 MB view details)

Uploaded CPython 3.10+Windows ARM64

gigatoken-0.8.0-cp310-abi3-win_amd64.whl (5.4 MB view details)

Uploaded CPython 3.10+Windows x86-64

gigatoken-0.8.0-cp310-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (5.4 MB view details)

Uploaded CPython 3.10+manylinux: glibc 2.17+ x86-64

gigatoken-0.8.0-cp310-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl (4.5 MB view details)

Uploaded CPython 3.10+manylinux: glibc 2.17+ ARM64

gigatoken-0.8.0-cp310-abi3-macosx_11_0_arm64.whl (4.2 MB view details)

Uploaded CPython 3.10+macOS 11.0+ ARM64

gigatoken-0.8.0-cp310-abi3-macosx_10_12_x86_64.whl (5.2 MB view details)

Uploaded CPython 3.10+macOS 10.12+ x86-64

File details

Details for the file gigatoken-0.8.0.tar.gz.

File metadata

  • Download URL: gigatoken-0.8.0.tar.gz
  • Upload date:
  • Size: 1.3 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.11.29 {"installer":{"name":"uv","version":"0.11.29","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for gigatoken-0.8.0.tar.gz
Algorithm Hash digest
SHA256 12f48ad3691c6979628e1bc72f9a0209c8fe04ba5fd4948898f5d77f3d3d2c9a
MD5 d933f0acdaea81be2f4e292029939ccd
BLAKE2b-256 dcf6938c80e04617d9a83a5770c8008761e0696cc1f82469c748ed4adbe4d0cf

See more details on using hashes here.

File details

Details for the file gigatoken-0.8.0-cp310-abi3-win_arm64.whl.

File metadata

  • Download URL: gigatoken-0.8.0-cp310-abi3-win_arm64.whl
  • Upload date:
  • Size: 4.4 MB
  • Tags: CPython 3.10+, Windows ARM64
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.11.29 {"installer":{"name":"uv","version":"0.11.29","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for gigatoken-0.8.0-cp310-abi3-win_arm64.whl
Algorithm Hash digest
SHA256 e09ccd157678cc6af2db146b9724b95b7aee8bb791f56ad0a3f6a3987d62d70e
MD5 f37c088fb6fe70562028137ae0cf80a5
BLAKE2b-256 18361d24344c45b59fc2bbb085562b4be2851435d99fd3c256cb90b93ad55dd1

See more details on using hashes here.

File details

Details for the file gigatoken-0.8.0-cp310-abi3-win_amd64.whl.

File metadata

  • Download URL: gigatoken-0.8.0-cp310-abi3-win_amd64.whl
  • Upload date:
  • Size: 5.4 MB
  • Tags: CPython 3.10+, Windows x86-64
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.11.29 {"installer":{"name":"uv","version":"0.11.29","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for gigatoken-0.8.0-cp310-abi3-win_amd64.whl
Algorithm Hash digest
SHA256 2dd4b17a92e113f233a47454ce99057727f325114781634f5e97b7ea954f8d00
MD5 f04b010c45fb83c840893a4a9ed7cd07
BLAKE2b-256 eb1e4fb5c723671fbe607e7a89443875fa4d9500cd0a4ed9f0df4de9ebf7a6f3

See more details on using hashes here.

File details

Details for the file gigatoken-0.8.0-cp310-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.

File metadata

  • Download URL: gigatoken-0.8.0-cp310-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
  • Upload date:
  • Size: 5.4 MB
  • Tags: CPython 3.10+, manylinux: glibc 2.17+ x86-64
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.11.29 {"installer":{"name":"uv","version":"0.11.29","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for gigatoken-0.8.0-cp310-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Algorithm Hash digest
SHA256 0848f92822be5d570bf4449fd9b487907d7ecf9492e81ecdad95abfb4dd60dd6
MD5 a6dc57404435e48c015874289c6e0117
BLAKE2b-256 126b38aa32777fa8320c2bcf6efc99181ecaeaaec84026776dfed33ac31ebdc4

See more details on using hashes here.

File details

Details for the file gigatoken-0.8.0-cp310-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl.

File metadata

  • Download URL: gigatoken-0.8.0-cp310-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
  • Upload date:
  • Size: 4.5 MB
  • Tags: CPython 3.10+, manylinux: glibc 2.17+ ARM64
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.11.29 {"installer":{"name":"uv","version":"0.11.29","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for gigatoken-0.8.0-cp310-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
Algorithm Hash digest
SHA256 435586a55db0c7d45dc6611283787f40d0a82229a1209eb0ef87ec71031ebd84
MD5 e054c3c94e88446f7c755b011ca69397
BLAKE2b-256 3dc3dd81401f0723da1bd0b99632074f368c84f87b6f9151ecedd77c9e29a9b0

See more details on using hashes here.

File details

Details for the file gigatoken-0.8.0-cp310-abi3-macosx_11_0_arm64.whl.

File metadata

  • Download URL: gigatoken-0.8.0-cp310-abi3-macosx_11_0_arm64.whl
  • Upload date:
  • Size: 4.2 MB
  • Tags: CPython 3.10+, macOS 11.0+ ARM64
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.11.29 {"installer":{"name":"uv","version":"0.11.29","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for gigatoken-0.8.0-cp310-abi3-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 4da76701d31c2f3b40d3d6c80b279cb4c4aaf1ff94b845e304c933e0aa28b85d
MD5 ec2d9bf5ae2210db48a94886d24b547b
BLAKE2b-256 b440f1deeb874f7ddf2b9f668dc66748876d113c52f32246a8d2faf3156736f7

See more details on using hashes here.

File details

Details for the file gigatoken-0.8.0-cp310-abi3-macosx_10_12_x86_64.whl.

File metadata

  • Download URL: gigatoken-0.8.0-cp310-abi3-macosx_10_12_x86_64.whl
  • Upload date:
  • Size: 5.2 MB
  • Tags: CPython 3.10+, macOS 10.12+ x86-64
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.11.29 {"installer":{"name":"uv","version":"0.11.29","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for gigatoken-0.8.0-cp310-abi3-macosx_10_12_x86_64.whl
Algorithm Hash digest
SHA256 1e5c9177054b40a6514e6f572d280eeb5d86accd1e2c11c6df98ec07c6afa7ae
MD5 b36430747c8eb307350fd99515ce41d4
BLAKE2b-256 7886eaad37b965e06e568d11a029e4697074366fea8b3cc18e8440af90097078

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page