hmaraniam
Zero-dependency language identification for Hmar.
"Hmar a ni am?" ("Is it Hmar?")
hmaraniam is a Python library that identifies Hmar text and tells it apart from English and the related Kuki-Chin / Zo languages (Mizo, Paite, Thadou, Vaiphei, Gangte, Zou).
Maintained by the Hmar Heritage Foundation as part of the Hmar Heritage Archival Project.
Features
- Frequency-weighted detection: Scores each token using corpus log-frequencies (
log(1 + count)) built from 2 Hmar Bibles and 583 verified web articles, covering 45,042 unigrams. Core structural words likechu,chun, andanscore higher than rare or loanword tokens. - Dual diacritic scoring: Reports
casual_hmar_ratio(ASCII-normalized for QWERTY typing) andformal_hmar_ratio(exact diacritic matches). - Sibling Zo language resolution: Separates Mizo, Paite, Thadou, Gangte, Zou, and Vaiphei from Hmar using dialect-exclusive particles and per-language vocabulary lists.
- Separate confidence scores:
hmar_confidenceanswers "how Hmar is this text?" independently ofdetected_language_confidence, which rates the overall classification call. - Consistent output shape: Every call returns the same dictionary structure, including word counts, frequency-weighted ratios, sibling scores, and diacritic breakdowns.
- Raw text cleanup: Strips HTML tags, Markdown syntax (bold, italic, links, code blocks), URLs, and email addresses from pasted input before scoring.
- Custom vocabulary: Pass custom unigram sets, extra domain words, or custom stopword lists directly to the
Detector. - Offline first: Ships with a bundled shard so it works without a network call. CDN sync via jsDelivr is available when you need the latest data.
- Zero dependencies: Pure Python standard library. No PyTorch, TensorFlow, NumPy, or spaCy.
Design
hmaraniam answers one question: "Is this text Hmar?" It does not correct spelling or modify the input.
Diacritic normalization. Mobile keyboards produce inconsistent accent codepoints. casual_hmar_ratio strips diacritics before matching so ṭha and tha both score against the same vocabulary entry.
Token boundaries. When you need precise control over how a token is defined (for example, whether mithiem-hai counts as one word or two), pass a pre-tokenized list. The library will not re-split it. For plain strings, it tokenizes by word boundary.
Vocabulary, not grammar. The score reflects dictionary overlap, not sentence structure. A list of valid Hmar words scores the same as a grammatical sentence with the same words.
Installation
pip install hmaraniam
Output schema
{
"language": "hmar",
"hmar_confidence": 0.9842,
"detected_language_confidence": 0.9842,
"sibling_heuristic": false,
"mode": "basic",
"scores": {
"casual_hmar_ratio": 0.9524,
"weighted_hmar_ratio": 0.9103,
"formal_hmar_ratio": 0.8095,
"english_stopword_ratio": 0.0000,
"sibling_zo_stopword_ratio": 0.0000,
"hmar_stopword_ratio": 0.1429,
"unknown_words_ratio": 0.0476,
"total_words": 21,
"hmar_words_count": 20,
"non_hmar_words_count": 1,
"unknown_words_count": 1,
"english_stopwords_count": 0,
"sibling_zo_stopwords_count": 0,
"hmar_stopwords_count": 3,
"sibling_lang_scores": {},
"hmar_diacritic_words_count": 17,
"non_hmar_diacritic_words_count": 0,
"total_diacritic_words_count": 17
}
}
Usage
Quick start
import hmaraniam
# Authentic text quote from L. Keivom archive (Coleman Factor, 2002)
sample_text = "Khawvel fe dan phung ei en chun, ram le hnam damna thuruk chu lien lema intel le insung khawm, zai khat le trong khata luong khawm a nih."
result = hmaraniam.detect(sample_text)
print(result)
Pre-tokenized inputs
If you need to define token boundaries yourself (for example, to treat mithiem-hai as a single token rather than two words), pass a list, JSON file, CSV, or line-delimited TXT. The library scores each entry as-is without re-splitting.
Supported formats
-
JSON array (
tokens.json):[ "khawvel", "fe", "dan", "mithiem-hai", "pathien", "hnenah" ]
Usage:
hmaraniam.detect("tokens.json")or CLIhmaraniam tokens.json -
CSV (
tokens.csv):token khawvel fe dan mithiem-hai pathien hnenah
Usage:
hmaraniam.detect("tokens.csv")or CLIhmaraniam tokens.csv -
Line-delimited TXT (
tokens.txt, 1 word per line):khawvel fe dan mithiem-hai pathien hnenah
Usage:
hmaraniam.detect("tokens.txt")or CLIhmaraniam tokens.txt -
Python list:
tokens = ["mithiem-hai", "pathien", "hnenah", "khawvel"] result = hmaraniam.detect(tokens)
Raw text
Pass a plain string or a .txt file path and the library tokenizes it automatically. HTML, Markdown, and URLs are stripped before scoring.
# Raw text string
result = hmaraniam.detect("Khawvel fe dan phung ei en chun, ram le hnam damna thuruk...")
# Raw text file
result = hmaraniam.detect("path/to/article.txt")
Custom unigrams and stopwords
from hmaraniam import Detector
# Provide custom unigrams or extra domain vocabulary
detector = Detector(
mode="basic",
extra_unigrams=["customworda", "customwordb"],
custom_stopwords=["and", "the", "with"],
disable_default_stopwords=False
)
result = detector.detect("Khawvel fe dan phung...")
Modes
In v0.2+, the entire 45,042 frequency-weighted vocabulary is bundled directly into the library, operating 100% offline with zero network latency.
from hmaraniam import Detector
# Default / Basic mode (45k unified unigrams, offline)
detector = Detector()
# High mode (retained for backward compatibility, uses unified dataset)
high_detector = Detector(mode="high")
# Explicit offline flag
offline_detector = Detector(offline_only=True)
Datasets
- Hmar Unigrams (
unigrams): 58,983 verified Hmar surface words and active loanwords. - Corpus Archive (
corpus-archive): Archival text corpus of Hmar literature and lexicons.
Error handling
hmaraniam raises standard Python exceptions:
import hmaraniam
# Raises ValueError for unsupported modes
try:
hmaraniam.detect("Text", mode="ultra")
except ValueError as e:
print(e)
# Raises TypeError for non-string input
try:
hmaraniam.detect(12345)
except TypeError as e:
print(e)
License
MIT License. Published by the Hmar Heritage Foundation.
Citation
@software{hmaraniam_2026,
author = {Hmar Heritage Foundation},
title = {hmaraniam: Zero-dependency language identification library for Hmar},
year = {2026},
publisher = {PyPI / GitHub},
iso_code = {hmr},
glottolog = {hmar1241},
clade = {Zo Languages},
howpublished = {\url{https://github.com/hmar-heritage-org/hmaraniam}}
}
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file hmaraniam-0.2.1.tar.gz.
File metadata
- Download URL: hmaraniam-0.2.1.tar.gz
- Upload date:
- Size: 499.3 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
97cc9dcff8d16058e428f212b54a1b6e6070ba0bfe6287d1d071e81f3b2cd202
|
|
| MD5 |
a657953d8a382526c7ea32b1b00e4d31
|
|
| BLAKE2b-256 |
cb8ed420ce2e6fcd6d7c4f92929436741b4c1e15964172f9447c835c468a82b9
|
File details
Details for the file hmaraniam-0.2.1-py3-none-any.whl.
File metadata
- Download URL: hmaraniam-0.2.1-py3-none-any.whl
- Upload date:
- Size: 498.0 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
a89bbefda10e11d2f223390e221d488c8a047e509dec0bc01fb8c8527b540e68
|
|
| MD5 |
a050682d8c68aaaac1236a06d3050c24
|
|
| BLAKE2b-256 |
d2cc8a117b74d8c283205c81f729e3301f4a707fcc36a45fbe6d737e211da9a5
|