Khmer Segmenter
A lightweight Khmer word segmentation package for Python ≥ 3.10, designed for simple and efficient tokenization.
This project is adapted and simplified from the original khnlp package, with the following goals:
- Support modern Python versions (≥ 3.10)
- Simplified installation
- Focus on a single task: Khmer word segmentation
- Lightweight and easy to integrate into NLP pipelines
This package is intended as a small academic contribution to support Khmer NLP research and practical applications.
Installation
pip install khmer-segmenter
Requires:
- Python >= 3.10
Usage
from khmer_segmenter import Tokenizer
tokenizer = Tokenizer(seg_type="com")
print(tokenizer.tokenize("សួស្ដីអ្នកទាំងអស់យើង"))
Segmentation Modes
The tokenizer supports two segmentation strategies:
1. Compound-based (seg_type="com")
- Segments text into compound words
- Suitable for general word-level NLP tasks
- Recommended for downstream applications such as:
- Text classification
- Named Entity Recognition
- Information retrieval
2. Morpheme-based (seg_type="mor")
- Performs finer-grained segmentation
- Splits text into smaller morphological units
- Useful for:
- Linguistic analysis
- Subword modeling
- Research-focused NLP tasks
Motivation
Khmer is a low-resource language with no explicit word boundary markers (spaces are not consistently used to separate words). This creates challenges for:
- Automatic Speech Recognition (ASR)
- Language Modeling
- Machine Translation
- Information Extraction
Existing tools for Khmer segmentation often have:
- Limited Python version support
- Heavy dependencies
- Broader NLP scope than necessary
This project provides a focused, minimal, and modern alternative dedicated solely to segmentation.
Metadata
Release files for khmer-segmenter 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| khmer_segmenter-0.1.0.tar.gz | 6.0 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| khmer_segmenter-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 11.9 MB
Release files / khmer_segmenter-0.1.0.tar.gz
| Download URL | khmer_segmenter-0.1.0.tar.gz |
|---|---|
| Size | 6.0 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
b993ecc9d2a38b9721b28d03ca6efbfaa1490870908986278d0d703d3c68c545
|
|
BLAKE2b-256 checksum How to use checksums |
b6b855df39844282dfcfb5b8ae87bee6864e356273e0b12cb187ebdbe7fef592
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.11.13
|
Release files / khmer_segmenter-0.1.0-py3-none-any.whl
| Download URL | khmer_segmenter-0.1.0-py3-none-any.whl |
|---|---|
| Size | 5.9 MB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
23c3824fe547b92f5fea370a0325bc2953de77d3a1b3cbb41dce7b892308a70d
|
|
BLAKE2b-256 checksum How to use checksums |
954a2f9e8034ec976275ad24182900870349c66f6e4036c28683e2eb6b5f7dbe
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.11.13
|