Skip to main content

Khmer Segmenter

A deterministic, dictionary-based Khmer word segmenter for NLP preprocessing, search, formal documents, and embedded applications. It combines Khmer Unicode normalization, frequency-weighted Viterbi decoding, linguistic rules, and unknown-word recovery without a runtime machine-learning model.

Documentation · Data preparation · C port · Rust port · Live demo

[!IMPORTANT] The package includes an attributed Khmer dictionary and derived runtime data for noncommercial use. Project code is MIT licensed; the bundled linguistic data has separate terms in DATA_LICENSE.md.

Features

  • Deterministic segmentation for the same code and local data
  • Khmer Unicode normalization
  • Frequency-weighted dictionary decisions
  • Layered curated and supplemental segmentation lexicons
  • A small, explicitly identified author-curated vocabulary for signature terms and names
  • RAC-curated spelling correction and autocomplete vocabulary
  • Word spelling checks through Python and the CLI
  • Whole-span typo diagnostics with Khmer-aware ranked suggestions
  • Unknown-span preservation
  • Typed token metadata with offsets and lexical POS candidates
  • Safe Khmer word-break opportunities for layout engines
  • Python API and khmer-segment CLI
  • Shared KDIC format for C and Rust applications
  • Rust/WASM segmentation and experimental spelling APIs for browser applications

This is a lexical segmenter, not a semantic parser or contextual POS tagger. pos_candidates are possibilities found in optional local lexical data.

Install for development

Python 3.10 or newer is required.

git clone https://github.com/Sovichea/khmer_segmenter.git
cd khmer_segmenter
python -m venv .venv

Activate the environment and install the src-layout package:

# Linux/macOS
source .venv/bin/activate
python -m pip install -e .
# Windows PowerShell
.venv\Scripts\Activate.ps1
python -m pip install -e .

After its first release, the distribution will install with:

pip install khmer-viterbi-segmenter

The import package remains khmer_segmenter. The bundled runtime data works immediately; no separate download is required.

Dictionary source and optional replacement

The original dictionary is published by Seanghay Hay (seanghay) on Hugging Face and was extracted from the Khmer Dictionary 2022 of the National Council of Khmer Language, Royal Academy of Cambodia:

https://huggingface.co/datasets/seanghay/khmer-dictionary-44k

The dataset may be redistributed for noncommercial use with attribution. The bundled normalized lexicons, RAC-only frequencies, and lexical POS candidates retain that credit and restriction. See the linguistic data notice.

For an exact model rebuild, download the structured RAC CSV directly from the original publisher:

mkdir -p dataset
curl -L \
  "https://huggingface.co/datasets/seanghay/khmer-dictionary-44k/resolve/525c0171894465cba920a9181387a032c11610d3/RAC-Khmer-Dict-2022.csv?download=true" \
  -o dataset/RAC-Khmer-Dict-2022.csv

Windows PowerShell:

New-Item -ItemType Directory -Force dataset | Out-Null
Invoke-WebRequest `
  -Uri "https://huggingface.co/datasets/seanghay/khmer-dictionary-44k/resolve/525c0171894465cba920a9181387a032c11610d3/RAC-Khmer-Dict-2022.csv?download=true" `
  -OutFile "dataset/RAC-Khmer-Dict-2022.csv"

Rebuild the authoritative RAC runtime artifacts deterministically:

python scripts/rebuild_rac_model.py \
  --rac-csv dataset/RAC-Khmer-Dict-2022.csv \
  --output-dir build/rac

khmer-segment data prepare --rac-tsv PATH remains available for simple custom 0.1-style dictionary overrides; it does not reproduce the strict RAC model.

The installed layered model additionally contains conservative supplemental segmentation chunks. Supplemental entries can preserve names, newer vocabulary, and known typo spans as single tokens, but they never become valid spellings or autocomplete candidates. Curated words keep their normal costs; supplemental words receive a cost penalty. Recreate that layer from a reviewed legacy list:

python scripts/prepare_supplemental_lexicon.py path/to/legacy_words.txt \
  --audit build/supplemental_audit.tsv
python scripts/build_dictionary_kdict.py

The compiler now produces one unified KDIC v2 language pack containing segmentation costs, spelling-validity flags, autocomplete eligibility, and approved typo corrections. Developers can replace any source list and deploy only the resulting .kdict file. See Unified KDIC v2 Language Packs.

Application developers can compile a single editable KLEX source using the installed CLI, without cloning the repository scripts:

khmer-segment data compile custom.klex.json --output custom.kdict

Extend an official pack without rebuilding it from source:

khmer-segment data compile local.klex.json \
  --base official.kdict --output application.kdict

The generated pack is standalone and works with Python, Rust, and WASM. See Unified KDIC v2 Language Packs.

For independently replaceable RAC, official terminology, user, and community packs, use the layered lexicon workflow. Strict mode loads RAC + reviewed official lexicons + user dictionaries; inclusive mode also loads explicitly supplied community evidence.

Corpus discovery and frequency curation are documented separately in the community corpus workflow. Community frequency can improve segmentation in inclusive mode, but it never grants spelling or autocomplete authority.

The audit records every curated match, retained chunk, and rejected fragment.

python scripts/validate_findings.py \
  --rac-csv dataset/RAC-Khmer-Dict-2022.csv

The Python resolver checks these locations in order:

  1. data_dir= or CLI --data-dir
  2. KHMER_SEGMENTER_DATA_DIR
  3. The user data directory for the operating system
  4. The data bundled with the installed package
  5. khmer_segmenter/dictionary_data/ in a development checkout

Check the resolved files:

khmer-segment data status
khmer-segment data sources
khmer-segment data prepare --rac-tsv dataset/rac_dictionary_2022_pairs.tsv

See Prepare Dictionaries for Python, C, and Rust for frequency generation and KDIC compilation.

Python API

from khmer_segmenter import KhmerSegmenter

segmenter = KhmerSegmenter()

tokens = segmenter.segment("ខ្ញុំស្រឡាញ់ប្រទេសកម្ពុជា")
print(tokens)

To use a replacement dictionary, set KHMER_SEGMENTER_DATA_DIR or pass data_dir= explicitly:

segmenter = KhmerSegmenter(data_dir="/path/to/replacement-data")

Load a unified custom pack instead:

segmenter = KhmerSegmenter.from_kdict("/path/to/custom.kdict")

The equivalent CLI option is --kdict custom.kdict.

Typed analysis results include normalized offsets and optional lexical data:

for token in segmenter.analyze("ខ្ញុំសរសេរឯកសារ"):
    print(token.text, token.start, token.end, token.known)
    print(token.frequency, token.pos_candidates)
    print(token.spelling_valid)

Check whole words independently of segmentation:

segmenter.is_spelling_valid("នីមួយៗ")
segmenter.check_spelling(["នីមួយៗ", "ពាក្យមិនស្គាល់"])

Detect probable typos in continuous text:

diagnostics = segmenter.detect_typos("សម្បត្ត")

for diagnostic in diagnostics:
    print(diagnostic.text, diagnostic.start, diagnostic.end)
    for suggestion in diagnostic.suggestions:
        print(suggestion.text, suggestion.edit_cost, suggestion.edits)

For an explicit editor lookup, treat the complete input as one word rather than relying on its initial segmentation:

suggestions = segmenter.suggest_spelling("សសេរ")
print(suggestions[0].text)  # សរសេរ

This reports the whole input span សម្បត្ត, suggests សម្បត្តិ, and records an insertion of at offset 7. Diagnostics are separate from segmentation tokens, so typo recovery does not silently alter segment() output. Offsets refer to normalized text by default; use normalize=False when the caller has already normalized the input.

Typo detection searches only near invalid Khmer tokens and uses weighted edits: dependent vowels and signs cost less than consonant substitutions. Results are probable corrections, not automatic replacements. Proper names, dialectal forms, and historical spellings still require application-level review.

Use a named spellcheck profile for application integration:

from khmer_segmenter import SpellcheckProfile

# Live editor underlines: strict confidence filtering and low latency.
diagnostics = segmenter.check_text(text, profile=SpellcheckProfile.TYPING)

# One segmentation pass with diagnostics and original-source offsets.
analysis = segmenter.analyze_text(text, profile=SpellcheckProfile.TYPING)

# Explicit "Check document": broader OOV correction search.
diagnostics = segmenter.check_text(text, profile="document")

Spelling accuracy is independent of the detection profile. The default, lexical, requires the exact curated spelling. Use visual when an editor should treat the legacy COENG DA/TA forms as equivalent. The forms ស្ដាប់ and ស្តាប់ then both pass validation, while completion and correction suggestions continue to show only the curated spelling.

from khmer_segmenter import SpellingAccuracy

segmenter.is_spelling_valid("ស្តាប់")  # False when RAC stores ស្ដាប់
segmenter.is_spelling_valid("ស្តាប់", accuracy=SpellingAccuracy.VISUAL)  # True

The CLI exposes the same combined result without making applications run segmentation twice:

khmer-segment analyze --profile typing --format json "សួរស្តី"
khmer_segmenter analyze --profile typing "សួរស្តី"

typing is the production default. document allows a wider edit distance but still avoids scanning every valid dictionary fragment. high-recall examines valid fragments and is intentionally experimental because it can produce many false positives. The old include_valid_fragments option remains as a low-level compatibility override.

Reviewed exact typo pairs live in dictionary_data/khmer_typo_corrections.tsv. Only approved rows affect spellcheck; proposed additions remain pending until reviewed. Run python scripts/sync_typo_corrections.py after editing the canonical Python copy so Rust and WASM consume the same list.

The legacy dictionary result remains available as segment_with_metadata(text).

For text layout, request legal break positions without changing the source:

offsets = segmenter.word_break_opportunities("ខ្មែរស្រឡាញ់ខ្មែរ")
text_with_breaks = segmenter.insert_word_breaks("ខ្មែរស្រឡាញ់ខ្មែរ")

Offsets refer to the original Python string. Breaks are offered only between adjacent known Khmer words, never inside a dictionary word. Layout engines can choose which opportunity to use; simpler consumers can use the insertion helper.

CLI

Segment positional text, a file, or standard input:

khmer-segment segment "ខ្ញុំស្រឡាញ់ប្រទេសកម្ពុជា"
khmer-segment segment --input input.txt --output segmented.txt
cat input.txt | khmer-segment segment

Machine-readable output:

khmer-segment segment "ខ្ញុំសរសេរឯកសារ" --format json
khmer-segment analyze "ខ្ញុំសរសេរឯកសារ" --format json
khmer-segment spellcheck "នីមួយៗ ពាក្យមិនស្គាល់"
khmer-segment diagnose "សម្បត្ត" --format json
khmer-segment diagnose --profile document --input manuscript.txt --format json
khmer-segment diagnose "រស់ជាតិ" --profile high-recall --format json

analyze reports lexical candidates; it does not claim contextual POS tagging.

Word-break opportunities and benchmarking:

khmer-segment word-breaks "ខ្មែរស្រឡាញ់ខ្មែរ" --format json
khmer-segment benchmark --input dataset/my_corpus.txt --limit 1000

Use a non-default local data directory with the global option before the subcommand:

khmer-segment --data-dir /path/to/local/data segment "អត្ថបទខ្មែរ"

Build the Python distribution

python -m pip install --upgrade build twine
python -m build
python -m twine check dist/*

The wheel contains code plus the attributed runtime data and its reproducibility manifest. Tests audit the archive to reject corpora, backups, provenance payloads, and unapproved linguistic artifacts.

Documentation

Data policy

The four runtime files listed in DATA_LICENSE.md are redistributed with attribution for noncommercial use. Source downloads, evaluation corpora, backups, provenance payloads, intermediate tables, and native build artifacts remain ignored and local.

Removing files from the current Git tree does not remove copies from old Git history. See the data policy before publishing or rewriting repository history.

License and acknowledgements

Project code is licensed under the MIT License. Bundled linguistic data is subject to the separate attribution and noncommercial notice.

Original data authors, authorities, corpus creators, and annotators are listed in Data Sources, Attribution, and Provenance.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

khmer_viterbi_segmenter-0.2.0.tar.gz (1.0 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

khmer_viterbi_segmenter-0.2.0-py3-none-any.whl (1.0 MB view details)

Uploaded Python 3

File details

Details for the file khmer_viterbi_segmenter-0.2.0.tar.gz.

File metadata

  • Download URL: khmer_viterbi_segmenter-0.2.0.tar.gz
  • Upload date:
  • Size: 1.0 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for khmer_viterbi_segmenter-0.2.0.tar.gz
Algorithm Hash digest
SHA256 b4b53b1c6752fe5bfd190968a0ada21f0300449b45e0f51387b62ba76acf378f
MD5 6e9fbeddec5f830fb347db86b8b4f027
BLAKE2b-256 4f44d3d50d16707bf56ea03073dfdbc559c7dc8d9f57f6078faee32bbf4a8377

See more details on using hashes here.

Provenance

The following attestation bundles were made for khmer_viterbi_segmenter-0.2.0.tar.gz:

Publisher: publish-to-pypi.yml on Sovichea/khmer_segmenter

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file khmer_viterbi_segmenter-0.2.0-py3-none-any.whl.

File metadata

File hashes

Hashes for khmer_viterbi_segmenter-0.2.0-py3-none-any.whl
Algorithm Hash digest
SHA256 9ce66c8de2cda060db1faec246162910549e145897956926703a67188576e32c
MD5 574814c35a5934582f7b91163f483c61
BLAKE2b-256 0bd9b1608a617db2c06da9aab48c136deddb443634c3a711247fb91750027218

See more details on using hashes here.

Provenance

The following attestation bundles were made for khmer_viterbi_segmenter-0.2.0-py3-none-any.whl:

Publisher: publish-to-pypi.yml on Sovichea/khmer_segmenter

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

This release

0.2.0 This release

2 files

0.1.1

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page