Skip to main content

Indic-BPE logo

Indic-BPE

A lightweight Byte Pair Encoding (BPE) tokenizer built from scratch for Hindi written in Devanagari.

Simple implementation • Efficient BPE training • Small Python API

Features

  • Hindi / Devanagari BPE tokenizer
  • Train BPE merges from a text corpus
  • Heap-based most-frequent pair selection
  • Linked-list based merge operations
  • Incremental local pair updates during training
  • Encode and decode support
  • Vocabulary and merge serialization
  • Simple Python package API
  • Streamlit tokenizer playground
  • Automated test suite
  • Training and tokenizer benchmarks

Architecture

Indic-BPE Architecture

Installation

From PyPI

pip install indic-bpe

Development Installation

git clone https://github.com/Rktim/indic-bpe.git
cd indic-bpe
uv sync

Quick Start

from indic_bpe import BPETokenizer

tokenizer = BPETokenizer()

tokenizer.load(
    "corpus/processed/hindi_vocab.json",
    "corpus/processed/hindi_merges.json",
)

text = "यह एक हिंदी भाषा का परीक्षण है।"

tokens = tokenizer.encode(text)
print(tokens)

decoded = tokenizer.decode(tokens)
print(decoded)

Example:

Tokens:
['य', 'ह ', 'ए', 'क ', 'ह', 'ि', 'ंद', 'ी ', 'भ', 'ा', 'ष', 'ा क', 'ा ',
 'प', 'र', 'ी', 'क्ष', 'ण', ' ', 'है', '।']

Decoded:
यह एक हिंदी भाषा का परीक्षण है।

Training

A tokenizer can be trained directly from a Hindi text corpus:

from indic_bpe import BPETokenizer

tokenizer = BPETokenizer()

tokenizer.train(
    loader,
    num_merges=1000,
)

The trainer repeatedly selects the most frequent adjacent symbol pair and merges it into a new token.

The implementation maintains pair statistics incrementally rather than rebuilding all pair frequencies after every merge.

Save and Load

Save a trained tokenizer:

tokenizer.save(
    "hindi_vocab.json",
    "hindi_merges.json",
)

Load it later:

tokenizer.load(
    "hindi_vocab.json",
    "hindi_merges.json",
)

The tokenizer stores:

  • Vocabulary in vocab.json
  • Merge rules in merges.json

API

train()

tokenizer.train(loader, num_merges=1000)

Train the BPE tokenizer from a corpus.

encode()

tokens = tokenizer.encode(text)

Convert Hindi text into BPE tokens.

decode()

text = tokenizer.decode(tokens)

Convert tokens back into the original text.

save()

tokenizer.save(vocab_path, merges_path)

Save the vocabulary and merge rules.

load()

tokenizer.load(vocab_path, merges_path)

Load a previously trained tokenizer.

Tokenizer Evaluation

Example:

यह एक हिंदी भाषा का परीक्षण है।

Current evaluation:

Characters: 31
Tokens: 21

Example tokens:

य | ह  | ए | क  | ह | ि | ंद | ी  | भ | ा | ष | ा क | ा  | प | र | ी | क्ष | ण |   | है | ।

The evaluation verifies that:

decoded == original_text

Streamlit Demo

Indic-BPE includes a small Hindi tokenizer playground.

Run:

uv run streamlit run scripts/tokenizer_app.py

The demo provides:

  • Hindi text input
  • Character count
  • Token count
  • Colored token visualization
  • Token text view
  • Token ID view
  • Token breakdown

Package Validation

The 0.0.1 package has been validated by:

  1. Building a wheel and source distribution.
  2. Installing the wheel into a clean virtual environment.
  3. Importing BPETokenizer from the installed package.
  4. Running a Hindi encode/decode round-trip.
  5. Loading trained Hindi vocabulary and merge files.
  6. Running a trained-tokenizer Hindi round-trip.

Example:

from indic_bpe import BPETokenizer

tokenizer = BPETokenizer()

text = "यह एक हिंदी भाषा का परीक्षण है।"

tokens = tokenizer.encode(text)
decoded = tokenizer.decode(tokens)

assert decoded == text

Project Structure

indic-bpe/
├── indic_bpe/
│   ├── __init__.py
│   ├── corpus.py
│   ├── decoder.py
│   ├── encoder.py
│   ├── merges.py
│   ├── serialization.py
│   ├── tokenizer.py
│   ├── trainer.py
│   ├── utils.py
│   ├── version.py
│   └── vocabulary.py
│
├── corpus/
│   ├── raw/
│   └── processed/
│
├── scripts/
│   ├── benchmark_tokenizer.py
│   ├── evaluate_tokenizer.py
│   └── tokenizer_app.py
│
├── tests/
├── assets/
├── README.md
├── LICENSE
├── pyproject.toml
└── .gitignore

Current Status

Version: 0.0.1

Indic-BPE 0.0.1 is an early release focused on Hindi / Devanagari BPE tokenization.

Completed:

  • BPE trainer
  • Optimized pair tracking
  • Encoder
  • Decoder
  • Vocabulary handling
  • Serialization
  • Tokenizer API
  • Test suite
  • Benchmark scripts
  • Streamlit demo
  • Python package build
  • Clean wheel installation validation

License

MIT License. See LICENSE.

Metadata

Release files for indic-bpe 0.0.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for indic-bpe 0.0.1
File Size Uploaded
indic_bpe-0.0.1.tar.gz 7.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for indic-bpe 0.0.1
File Interpreter ABI Platform
indic_bpe-0.0.1-py3-none-any.whl Python 3 none any Details

Total release size: 18.1 kB

Release files / indic_bpe-0.0.1.tar.gz

Download URL indic_bpe-0.0.1.tar.gz
Size 7.7 kB
Tags Source
SHA-256 checksum
How to use checksums
7f7859cbfa6d68c7fe147e4128706daaccf00b57a6e3713a408094c8df69939b
BLAKE2b-256 checksum
How to use checksums
3254981228941e9c1a8e1f2eff3c3a6bbdce3b61c0c9efecc80f08abd67de6d2
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.3

Release files / indic_bpe-0.0.1-py3-none-any.whl

Download URL indic_bpe-0.0.1-py3-none-any.whl
Size 10.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
dc40b9f672be72ff89706c204574aa4d9674a7f93fa415eab1eeb099c5570b96
BLAKE2b-256 checksum
How to use checksums
db2d551512c806f926cbacc48366a3b6ea4d6c155d85080b089f5a2d4c7377f3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.3

Release history Release notifications | RSS feed

This release

0.0.1 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page