A Dzongkha tokenizer and segmenter (32k Subword & Tseg)
Project description
dzoseg 🇧🇹
dzoseg is a lightweight Python library for Dzongkha text segmentation. It offers a hybrid approach, allowing users to choose between high-performance AI subword tokenization (using SentencePiece) and traditional Tseg-based segmentation.
🚀 Features
- Subword Segmentation: Uses a pre-trained Unigram model with a 32,000 vocabulary size for deep learning applications.
- Tseg Segmentation: A rule-based approach to split text into syllables based on the traditional Dzongkha Tseg (་).
- Byte-Fallback: Automatically handles non-Dzongkha characters (English, numbers, etc.) without crashing.
📦 Installation
pip install dzoseg
🛠 Usage
1. Subword (AI-Based) Segmentation
This method is recommended for Machine Learning tasks like machine translation or sentiment analysis.
from dzoseg import DzongkhaTokenizer
# Initialize the segmenter
ds = DzongkhaTokenizer()
text = "འབྲུག་རྒྱལ་གཞུང་གཙུག་ལག་སློབ་སྡེའི་འོག་ལུ་ཡོད་པའི་ཚན་རིག་དང་འཕྲུལ་རིག་མཐོ་རིམ་སློབ་གྲྭའི་གློག་རིག་དང་འཕྲུལ་རིག་ལས་ཁུངས།"
# Get subword tokens
tokens = ds.segment_subwords(text)
print(f"Subwords: {tokens}")
# Expected Output:
# Subwords: [' འབྲུག་', 'རྒྱལ་གཞུང་', 'གཙུག་ལག་སློབ་སྡེ', 'འི་', 'འོག་ལུ་ཡོད་པའི་', 'ཚན་རིག་', 'དང་', 'འཕྲུལ་རིག་', 'མཐོ་རིམ་སློབ་གྲྭ', 'འི་', 'གློག་རིག་', 'དང་', 'འཕྲུལ་རིག་', 'ལས་ཁུངས།']
2. Tseg-based Segmentation
This method splits the text strictly based on the Dzongkha syllable delimiter (Tseg).
# Get syllable-level tokens
syllables = ds.segment_tseg(text)
print(f"Syllables: {syllables}")
# Expected Output:
# Syllables: ['འབྲུག','རྒྱལ','གཞུང', 'ཙུག', 'ལག', 'སློབ', 'སྡེའི', 'འོག', 'ལུ', 'ཡོད', 'པའི', 'ཚན', 'རིག', 'དང', 'འཕྲུལ', 'རིག', 'མཐོ', 'རིམ', 'སློབ', 'གྲྭའི', 'གློག', 'རིག', 'དང', 'འཕྲུལ', 'རིག', 'ལས', 'ཁུངས།']
📊 Model Details
Model Type: Unigram (SentencePiece)
Vocabulary Size: 32,000
Character Coverage: 100%
Training Data: Cleaned Dzongkha Web & Literary Corpus
📜 License
This project is licensed under the MIT License.
Maintained by: Karma Wangchuk
Contact: karma.cst@rub.edu.bt
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
dzoseg-0.1.5.tar.gz
(742.9 kB
view details)
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
dzoseg-0.1.5-py3-none-any.whl
(768.3 kB
view details)
File details
Details for the file dzoseg-0.1.5.tar.gz.
File metadata
- Download URL: dzoseg-0.1.5.tar.gz
- Upload date:
- Size: 742.9 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.11.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
d09dd9feac1ae1fcfb1e947191705c58e39efecb928c194137d2531febc16305
|
|
| MD5 |
79f747ca4553cb4b16568855fefd2b26
|
|
| BLAKE2b-256 |
e2094ba041059c503e73d541b080eae19dda4bc9a8b3f35f0992b7554c1512f6
|
File details
Details for the file dzoseg-0.1.5-py3-none-any.whl.
File metadata
- Download URL: dzoseg-0.1.5-py3-none-any.whl
- Upload date:
- Size: 768.3 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.11.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
0ece05fb1330c21ad7a3caeec64379114aaa976d8fdc4a0df70654c1187216e1
|
|
| MD5 |
5920e0c2db3656bc31db893018a7d2a9
|
|
| BLAKE2b-256 |
5b84d2b37dcd28926520bb629002732e8c41a3c9a8acc4126fd4e6dfbd1d7692
|