A Byte Pair Encoding (BPE) library for the Bengali language.
Project description
Bengali BPE
bengali_bpe is a Python library for Byte Pair Encoding (BPE) specifically designed for the Bengali language.
It enables you to train BPE models on Bengali text, encode words and sentences into subword units, and decode them back.
This helps improve NLP model performance for Bengali text processing, tokenization, and embedding preparation.
✨ Features
- 🧠 Train a Byte Pair Encoding model on Bengali text corpus
- 🔠 Encode Bengali sentences or words into subword tokens
- 🔁 Decode subword tokens back into full Bengali words
- ⚙️ Simple, lightweight, and easy to integrate into your NLP pipelines
- 🪶 Supports Bengali Unicode normalization
📦 Installation
Install directly from PyPI:
pip install bengali_bpe
Usage Examples
Train a BPE Model and Encode Sentences
from bengali_bpe import BengaliBPE
from bengali_bpe.utils import normalize_bengali_text
# Sample Bengali corpus
corpus = [
"বাংলা ভাষা সুন্দর",
"আমি বাংলা পড়ি",
"বাংলা ভয়ানক নয়"
]
# Normalize text
corpus = [normalize_bengali_text(sentence) for sentence in corpus]
# Initialize and train the model
bpe = BengaliBPE(num_merges=10)
bpe.train(corpus)
# Encode a sentence
sentence = "বাংলা ভাষা সুন্দর"
encoded = bpe.encode(sentence)
print("Encoded:", encoded)
# Decode back
decoded = bpe.decode(encoded)
print("Decoded:", decoded)
Output
Encoded: [['বা', 'ংলা'], ['ভা', 'ষা'], ['সু', 'ন্', 'দর']]
Decoded: বাংলা ভাষা সুন্দর
Encode and Decode a Single Word
from bengali_bpe import BengaliBPE
bpe = BengaliBPE(num_merges=5)
bpe.train(["বাংলা ভাষা সুন্দর"])
encoded_word = bpe.encode_word("বাংলা")
print("Encoded Word:", encoded_word)
decoded_word = bpe.decode([encoded_word])
print("Decoded Word:", decoded_word)
Output
Encoded Word: ['বা', 'ংলা']
Decoded Word: বাংলা
Normalize Bengali Text
from bengali_bpe.utils import normalize_bengali_text
text = "বাংলা ভাষা সুন্দর।।"
print(normalize_bengali_text(text))
Output
বাংলা ভাষা সুন্দর।।
Example: Training and Applying BPE on a Bengali Paragraph
from bengali_bpe import BengaliBPE
from bengali_bpe.utils import normalize_bengali_text
text = """বাংলা একটি মধুর ভাষা। এটি বিশ্বের অন্যতম প্রাচীন ও সমৃদ্ধ ভাষাগুলোর একটি।
বাংলা ভাষার ইতিহাস ও ঐতিহ্য হাজার বছরের পুরোনো।"""
corpus = [normalize_bengali_text(text)]
bpe = BengaliBPE(num_merges=15)
bpe.train(corpus)
encoded = bpe.encode("বাংলা একটি মধুর ভাষা")
print("Encoded:", encoded)
decoded = bpe.decode(encoded)
print("Decoded:", decoded)
Full Example: Combine All Steps
from bengali_bpe import BengaliBPE
from bengali_bpe.utils import normalize_bengali_text
corpus = [
"আমি বাংলা ভাষা ভালোবাসি",
"বাংলা একটি সুন্দর ভাষা",
"বাংলা আমাদের মাতৃভাষা"
]
corpus = [normalize_bengali_text(c) for c in corpus]
bpe = BengaliBPE(num_merges=12)
bpe.train(corpus)
sentence = "আমি বাংলা ভালোবাসি"
encoded = bpe.encode(sentence)
decoded = bpe.decode(encoded)
print("Original:", sentence)
print("Encoded:", encoded)
print("Decoded:", decoded)
Output
Original: আমি বাংলা ভালোবাসি
Encoded: [['আ', 'মি'], ['বা', 'ংলা'], ['ভা', 'লো', 'বা', 'সি']]
Decoded: আমি বাংলা ভালোবাসি
Example Use Cases
| Use Case | Description |
|---|---|
| 🔤 Subword Tokenization | Split Bengali words into meaningful subword units for NLP models |
| 🧩 Embedding Preparation | Generate stable subword tokens for embedding or transformer-based models |
| 🧠 Text Compression | Apply BPE for efficient text representation |
| 📚 Data Preprocessing | Clean and normalize Bengali text before training models |
API References
| Function | Description |
|---|---|
train(corpus) |
Train the BPE model on a list of Bengali sentences |
encode(text) |
Encode an entire Bengali sentence into subword tokens |
encode_word(word) |
Encode a single Bengali word |
decode(encoded_words) |
Decode BPE tokens back to full Bengali text |
normalize_bengali_text(text) |
Normalize and clean Bengali text (NFC normalization) |
Project Structure
bengali-bpe/
├─ README.md
├─ LICENSE
├─ pyproject.toml
├─ src/
│ └─ bengali_bpe/
│ ├─ __init__.py
│ ├─ encoder.py
│ └─ utils.py
└─ tests/
└─ test_import.py
Developer
Firoj Ahmmed Patwary
BSc & MSc in Statistics, Jagannath University
MSc in Data Science, Freie Universität Berlin
Researcher in Data Science, Machine Learning, NLP, and Explainable AI
Contact:
🌐 Website: www.firoj.net
📧 Email: firoj.stat@gmail.com
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file bengali_bpe-0.2.0.tar.gz.
File metadata
- Download URL: bengali_bpe-0.2.0.tar.gz
- Upload date:
- Size: 5.6 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
92fe3f49fe30b18d4c41b87565d9ed1a9b44c31fe72e3e2493e4e72dd74e1a33
|
|
| MD5 |
e81b63139105a2f50087f0656f7a6e6b
|
|
| BLAKE2b-256 |
9ddab6b2d466468682f00824fa5fbe56689333356d1b9c5b570865a9ea0bd928
|
File details
Details for the file bengali_bpe-0.2.0-py3-none-any.whl.
File metadata
- Download URL: bengali_bpe-0.2.0-py3-none-any.whl
- Upload date:
- Size: 5.6 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
eabc25fba2cca4c0d0bbff55bab129071769bf96c451dfdaf0ba894d14b3f0fe
|
|
| MD5 |
0b94eaed9ea168b52adf2aaa71c1bd68
|
|
| BLAKE2b-256 |
b73834f3f43f5ff2b068f9890899d97f2786ff8c904fe9f7aed219814a5df252
|