A Dzongkha tokenizer and segmenter (32k Subword & Tseg)
Project description
dzoseg 🇧🇹
dzoseg is a lightweight Python library for Dzongkha text segmentation. It offers a hybrid approach, allowing users to choose between high-performance AI subword tokenization (using SentencePiece) and traditional Tseg-based segmentation.
🚀 Features
- Subword Segmentation: Uses a pre-trained Unigram model with a 32,000 vocabulary size for deep learning applications.
- Tseg Segmentation: A rule-based approach to split text into syllables based on the traditional Dzongkha Tseg (་).
- Byte-Fallback: Automatically handles non-Dzongkha characters (English, numbers, etc.) without crashing.
📦 Installation
pip install dzoseg
🛠 Usage
1. Subword (AI-Based) Segmentation
This method is recommended for Machine Learning tasks like machine translation or sentiment analysis.
from dzoseg import DzongkhaTokenizer
# Initialize the segmenter
ds = DzongkhaTokenizer()
text = "འབྲུག་རྒྱལ་གཞུང་གཙུག་ལག་སློབ་སྡེའི་འོག་ལུ་ཡོད་པའི་ཚན་རིག་དང་འཕྲུལ་རིག་མཐོ་རིམ་སློབ་གྲྭའི་གློག་རིག་དང་འཕྲུལ་རིག་ལས་ཁུངས།"
# Get subword tokens
tokens = ds.segment_subwords(text)
print(f"Subwords: {tokens}")
# Expected Output:
# Subwords: [' འབྲུག་', 'རྒྱལ་གཞུང་', 'གཙུག་ལག་སློབ་སྡེ', 'འི་', 'འོག་ལུ་ཡོད་པའི་', 'ཚན་རིག་', 'དང་', 'འཕྲུལ་རིག་', 'མཐོ་རིམ་སློབ་གྲྭ', 'འི་', 'གློག་རིག་', 'དང་', 'འཕྲུལ་རིག་', 'ལས་ཁུངས།']
2. Tseg-based Segmentation
This method splits the text strictly based on the Dzongkha syllable delimiter (Tseg).
# Get syllable-level tokens
syllables = ds.segment_tseg(text)
print(f"Syllables: {syllables}")
# Expected Output:
# Syllables: ['འབྲུག','རྒྱལ','གཞུང', 'ཙུག', 'ལག', 'སློབ', 'སྡེའི', 'འོག', 'ལུ', 'ཡོད', 'པའི', 'ཚན', 'རིག', 'དང', 'འཕྲུལ', 'རིག', 'མཐོ', 'རིམ', 'སློབ', 'གྲྭའི', 'གློག', 'རིག', 'དང', 'འཕྲུལ', 'རིག', 'ལས', 'ཁུངས།']
📊 Model Details
Model Type: Unigram (SentencePiece)
Vocabulary Size: 32,000
Character Coverage: 100%
Training Data: Cleaned Dzongkha Web & Literary Corpus
📜 License
This project is licensed under the MIT License.
Maintained by: Karma Wangchuk
Contact: karma.cst@rub.edu.bt
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
dzoseg-0.1.4.tar.gz
(743.1 kB
view details)
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
dzoseg-0.1.4-py3-none-any.whl
(768.5 kB
view details)
File details
Details for the file dzoseg-0.1.4.tar.gz.
File metadata
- Download URL: dzoseg-0.1.4.tar.gz
- Upload date:
- Size: 743.1 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.11.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
a862bb96949c182ac54910711ddb9137e4de40df3f5b38cf1944bcf25707551d
|
|
| MD5 |
6c4aee606f84f5c7fd7ceb22e461e6f2
|
|
| BLAKE2b-256 |
81febcbffdb230ca8cda7da1bffa6c1e572397b5165152e926b79a6df9d37d33
|
File details
Details for the file dzoseg-0.1.4-py3-none-any.whl.
File metadata
- Download URL: dzoseg-0.1.4-py3-none-any.whl
- Upload date:
- Size: 768.5 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.11.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
0fc534913aed59fe8ba74d8ada7e446c909afe4d26b8ea220e364ce0faad41b4
|
|
| MD5 |
7f591a55da8f32ed649e1b966bec2660
|
|
| BLAKE2b-256 |
4a7a6ccd81c86fe6d3f8d488088ec0849ad2118348399b6bf6df6f67b16a2599
|