Khmer Segment
A Khmer word segmentation tool built for NIPTICT (now CADT) Khmer Word Segmentation CRF model.
[!IMPORTANT]
km-5tag-seg-modelis required for this script to work. This library doesn't provide the model file.
Usage
pip install khmersegment
from khmersegment import Segmenter
segmenter = Segmenter("-m km-5tag-seg-model")
print(segmenter("Hello មិនដឹងប្រាប់អ្នកណាទេ?", deep=False))
# => ['Hello', ' ', 'មិន', 'ដឹង', 'ប្រាប់', 'អ្នកណា', 'ទេ', '?']
print(segmenter("Hello មិនដឹងប្រាប់អ្នកណាទេ?", deep=True))
# => ['Hello', ' ', 'មិន', 'ដឹង', 'ប្រាប់', 'អ្នក', 'ណា', 'ទេ', '?']
License
Apache-2.0
Related
- pycrfpp Python binding for CRF++
Metadata
Release files for khmersegment 0.1.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| khmersegment-0.1.2.tar.gz | 7.0 kB | Details |
Release files / khmersegment-0.1.2.tar.gz
| Download URL | khmersegment-0.1.2.tar.gz |
|---|---|
| Size | 7.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
66779cc5b220fe1099dc8a763431cf1f896f3609bf47a23a9370b39a1934291d
|
|
BLAKE2b-256 checksum How to use checksums |
1dc3974d7d091c78db5da3c1d4c63dbbedb977099ff023c971d9683c1db81535
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/5.1.0 CPython/3.8.19
|