Skip to main content

📄 README.md

🔣 tokeniser-py-lite

Imp Links: PyPI Library | PyPI Main Library (tokeniser-py) | Main Library GitHub (tokeniser-py) | Demo (HF Spaces) | Complete repo (unchunked) - HF | Complete repo (chunked) - GitHub | Imp Files Github

A high-performance, fully custom tokeniser built from scratch — no BPE, no existing NLP tokenisation scheme. This tokeniser is based on a unique algorithm developed independently and trained on over 1 billion tokens from the SlimPajama dataset (Val + Test), providing an efficient, interpretable, and extendable tokenisation pipeline.

🚀 What This Library Offers

  • Tokeniser built on a vocabulary of 131,072 tokens
  • Two versions of vocab:
    • 0.5B: Validation-only data
    • 1B: Validation + Test data
  • Token vocab built via a custom algorithm — no Byte Pair Encoding (BPE)
  • Tokenisation logic includes:
    • Token lookup from pre-generated token map
    • Dynamic programming-based segmentation for out-of-vocab tokens
    • One-hot encoding (NumPy or PyTorch)
    • Visualisation utilities for tokens and token IDs
  • Lightweight JSON format for token maps & token count maps
  • Ready for integration into any LLM pre-tokenisation pipeline

Note: Files (chunked less than 2GB) are stored on Hugging Face instead of GitHub due to LFS file size constraints. On GitHub (files chunked below 100MB) are available.

📦 Installation

pip install tokeniser-py-lite

🛠 Usage

from tokeniser import Tokeniser

t = Tokeniser()
tokens, count = t.tokenise("Your input text here.")
token_ids = t.token_ids(tokens)

Use t.one_hot_tokens(token_ids) for NumPy-based one-hot encoding, or op='torch' for PyTorch.

📚 Data Sources

All token maps and token counts are generated from the SlimPajama dataset by Cerebras.

📁 Vocab Files

  • ordered_tokenizer_1b_val_test_data.json — Ordered tokens (1B data)
  • unordered_tokenizer_1b_val_test_data.json — Unordered tokens (1B)
  • count_tokenizer_1b_val_test_data.json — Token counts (1B)
  • (Similar structure for 0.5B val-only version)

📌 Design Philosophy

This tokeniser is built from scratch before learning existing algorithms like BPE. It is designed with the intent to understand, innovate, and compare with existing solutions from first principles.

Some parts may overlap with BPE/WordPiece in spirit — but the core algorithm was independently designed.

🤝 Contributions

Feel free to contribute anything via GitHub.

📖 License

MIT License

📄 CHANGELOG

📦 Changelog

[0.1.0] - 2025-04-04

Added

  • Initial release of custom tokeniser library
  • Light-weight version of the original python library tokeniser-py
  • Tokeniser class with support for:
    • tokenise() using DP segmentation
    • Custom token map and count map loading
    • One-hot encoding support (NumPy & PyTorch)
    • Token and token ID visualisation functions
    • token_map(), token_count_map(), max_token_length() accessors
  • Full support for:
    • 0.5B val-only vocab
    • 1B val + test vocab
  • JSON-based token and count maps from SlimPajama corpus

[0.1.1] - 2025-04-04

Added

  • Corrected dates in Changelog
  • Updated Readme

Notes

  • Built on top of a custom token creation algorithm not based on any standard BPE/WordPiece method
  • SlimPajama dataset used for vocab extraction
  • Token count files are optimized to stay under 2GB for compatibility with Git LFS (and Hugging Face storage)

Metadata

Release files for tokeniser-py-lite 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for tokeniser-py-lite 0.1.1
File Size Uploaded
tokeniser-py-lite-0.1.1.tar.gz 4.9 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for tokeniser-py-lite 0.1.1
File Interpreter ABI Platform
tokeniser_py_lite-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 9.9 MB

Release files / tokeniser-py-lite-0.1.1.tar.gz

Download URL tokeniser-py-lite-0.1.1.tar.gz
Size 4.9 MB
Tags Source
SHA-256 checksum
How to use checksums
0d1eee98e86aea682fcdf8fa528e29968668ad7f7e7013b1ad9a07111b4d09cf
BLAKE2b-256 checksum
How to use checksums
a141b89dfde55be033a88038003deebd2d27cfe37a4124386215a1fd0e5ebe3c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.0.1 CPython/3.10.8

Release files / tokeniser_py_lite-0.1.1-py3-none-any.whl

Download URL tokeniser_py_lite-0.1.1-py3-none-any.whl
Size 4.9 MB
Tags Python 3
SHA-256 checksum
How to use checksums
f5b500b9c29ab20b3c7a44623a1aea5b450093cfb5631673b87cf52f5c7cd739
BLAKE2b-256 checksum
How to use checksums
b1ad569ef8f3cd4fa88f7a06755883e337717df14b45c5f0da4335277dff09e6
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.0.1 CPython/3.10.8

Release history Release notifications | RSS feed

This release

0.1.1 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page