Skip to main content

hmaraniam 🇲z

Zero-dependency language identification library for Hmar.

"Hmar a ni am?""Is it Hmar?"

hmaraniam is a lightweight, zero-dependency Python library that identifies Hmar text and cleanly distinguishes it from English and related Kuki-Chin / Zo languages (Mizo, Paite, Thadou, Vaiphei, Gangte, Zou).

Maintained by the Hmar Heritage Foundation as part of the Hmar Heritage Archival Project.


Features

  • Fast dictionary lookups: Uses $O(1)$ set matching backed by 37,093 verified pure Hmar unigrams with no machine learning dependencies (PyTorch and TensorFlow free).
  • Dual diacritic scoring: Reports casual_hmar_ratio (ASCII-normalized for standard QWERTY typing) and formal_hmar_ratio (exact diacritic matches for formal text).
  • Specific Sibling Language Resolution: Distinguishes sibling Zo languages (mizo, paite, thadou, gangte, zou, vaiphei) using dialect-exclusive particles and exclusive vocabulary sets.
  • Separate confidence scores: Separates overall classification confidence (detected_language_confidence) from Hmar-specific confidence (hmar_confidence).
  • Consistent JSON output: Returns the same dictionary structure for every call, including word counts, sibling scores, and diacritic breakdowns.
  • Custom unigrams & stopwords: Pass custom unigram sets, extra domain vocabulary, or custom stopword lists.
  • Offline & CDN dataset loading: Syncs unigram sets via jsDelivr CDN with local disk caching and bundled offline fallbacks.

Design Principles

  • Language ID vs. Spell Correction: hmaraniam measures vocabulary identity ("Is this text Hmar?"). It is not a spell checker and does not modify typos, character variants (acute á, grave à, circumflex â), or mobile keyboard codepoints ( vs ţ).
  • ASCII Normalization (casual_hmar_ratio): Mobile keyboards produce varying accent codepoints. Stripping diacritics (strip_diacritics) allows consistent vocabulary evaluation across devices.
  • 1-Token-Per-Row Boundaries: To evaluate hyphenated (mithiem-hai), spaced (mithiem hai), or compound (mithiemhai) terms directly, hmaraniam accepts 1-token-per-row inputs (JSON, CSV, TXT, Python lists) without re-tokenizing.
  • Vocabulary Identity vs. Grammar: hmaraniam measures dictionary presence and token overlap, not syntax or semantics. A random sequence of valid Hmar words yields a high vocabulary score regardless of grammatical structure.

Installation

pip install hmaraniam

Output Schema

{
  "language": "hmar",
  "hmar_confidence": 0.9842,
  "detected_language_confidence": 0.9842,
  "sibling_heuristic": false,
  "mode": "basic",
  "scores": {
    "casual_hmar_ratio": 0.9524,
    "formal_hmar_ratio": 0.8095,
    "english_stopword_ratio": 0.0000,
    "sibling_zo_stopword_ratio": 0.0000,
    "unknown_words_ratio": 0.0476,
    "total_words": 21,
    "hmar_words_count": 20,
    "non_hmar_words_count": 1,
    "unknown_words_count": 1,
    "english_stopwords_count": 0,
    "sibling_zo_stopwords_count": 0,
    "sibling_lang_scores": {},
    "hmar_diacritic_words_count": 17,
    "non_hmar_diacritic_words_count": 0,
    "total_diacritic_words_count": 17
  }
}

Usage

Quick Start

import hmaraniam

# Authentic text quote from L. Keivom archive (Coleman Factor, 2002)
sample_text = "Khawvel fe dan phung ei en chun, ram le hnam damna thuruk chu lien lema intel le insung khawm, zai khat le trong khata luong khawm a nih."

result = hmaraniam.detect(sample_text)
print(result)

1-Token-Per-Row Inputs

When token boundaries are pre-defined (such as distinguishing "mithiem-hai" vs "mithiem hai" vs "mithiemhai"), hmaraniam evaluates 1-token-per-row inputs without internal re-tokenization:

Expected File Formats & Code Examples

  1. JSON Array File (tokens.json):

    [
      "khawvel",
      "fe",
      "dan",
      "mithiem-hai",
      "pathien",
      "hnenah"
    ]
    

    Usage: hmaraniam.detect("tokens.json") or CLI hmaraniam tokens.json

  2. CSV File (tokens.csv):

    token
    khawvel
    fe
    dan
    mithiem-hai
    pathien
    hnenah
    

    Usage: hmaraniam.detect("tokens.csv") or CLI hmaraniam tokens.csv

  3. Line-Delimited TXT File (tokens.txt - 1 word per line):

    khawvel
    fe
    dan
    mithiem-hai
    pathien
    hnenah
    

    Usage: hmaraniam.detect("tokens.txt") or CLI hmaraniam tokens.txt

  4. Python List (List[str]):

    tokens = ["mithiem-hai", "pathien", "hnenah", "khawvel"]
    result = hmaraniam.detect(tokens)
    

Un-tokenized Raw Text Documents

For raw text files or strings (article.txt, raw text string, or stdin pipe), hmaraniam extracts word tokens using word-boundary regex matching:

# Raw text string evaluation
result = hmaraniam.detect("Khawvel fe dan phung ei en chun, ram le hnam damna thuruk...")

# Raw text article file evaluation
result = hmaraniam.detect("path/to/article.txt")

Custom Unigrams & Stopwords

from hmaraniam import Detector

# Provide custom unigrams or extra domain vocabulary
detector = Detector(
    mode="basic",
    extra_unigrams=["customworda", "customwordb"],
    custom_stopwords=["and", "the", "with"],
    disable_default_stopwords=False
)

result = detector.detect("Khawvel fe dan phung...")

Modes & Advanced Options

from hmaraniam import Detector

# Basic Mode (Default ~30k core unigrams)
basic_detector = Detector(mode="basic")

# High Mode (Loads extended unigram shards, falling back to basic if unavailable)
high_detector = Detector(mode="high")

# Offline-only mode (uses cached or bundled dataset without network calls)
offline_detector = Detector(offline_only=True)

Benchmarks & Reports

Check out the reports/ folder for full test results and breakdown (if you're reading this on PyPI, head over to our GitHub repo):

  • Parallel Zo Bible Test (reports/parallel_zo_bibles.md): Tested across 10 Bible translations in 8 Zo languages + English. Hits 100% accuracy on Hmar (CLB & OV), Mizo, Paite, Thadou, and English.
  • Web Archives Test (reports/web_archives.md): Tested on ~1,300 web posts across 5 site archives (Keivom, Inpui, HSA, Hmarram, Virthli).
  • Text Length Tests (reports/length_sensitivity.md): Checks how well the detector handles short vs long text snippets.

Datasets & Repositories


Error Handling

hmaraniam raises standard Python exceptions:

import hmaraniam

# Raises ValueError for unsupported modes
try:
    hmaraniam.detect("Text", mode="ultra")
except ValueError as e:
    print(e)

# Raises TypeError for non-string input
try:
    hmaraniam.detect(12345)
except TypeError as e:
    print(e)

License

Published under the MIT License by the Hmar Heritage Foundation.


Citation & Attribution

If you use this software in your research or tools, please cite:

@software{hmaraniam_2026,
  author       = {Hmar Heritage Foundation},
  title        = {hmaraniam: Zero-dependency language identification library for Hmar},
  year         = {2026},
  publisher    = {PyPI / GitHub},
  iso_code     = {hmr},
  glottolog    = {hmar1241},
  clade        = {Zo Languages},
  howpublished = {\url{https://github.com/hmar-heritage-org/hmaraniam}}
}

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

hmaraniam-0.1.6.tar.gz (403.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

hmaraniam-0.1.6-py3-none-any.whl (398.8 kB view details)

Uploaded Python 3

File details

Details for the file hmaraniam-0.1.6.tar.gz.

File metadata

  • Download URL: hmaraniam-0.1.6.tar.gz
  • Upload date:
  • Size: 403.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for hmaraniam-0.1.6.tar.gz
Algorithm Hash digest
SHA256 cce1ef9e85fecb1c6a054e81edf2eb6f8b02abbad20f99e4c6a71778e874ae7c
MD5 97d7c9584965f0f483cb6323f99c66ed
BLAKE2b-256 b8fac1649095cdb83ca86f098625b5607b7d31ef2a08b5fbb66cdf8b00db3319

See more details on using hashes here.

File details

Details for the file hmaraniam-0.1.6-py3-none-any.whl.

File metadata

  • Download URL: hmaraniam-0.1.6-py3-none-any.whl
  • Upload date:
  • Size: 398.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for hmaraniam-0.1.6-py3-none-any.whl
Algorithm Hash digest
SHA256 10c8e82ce5201b2d844817b0a0d81b0e1c0320686ea08d4addf43ccd016c2b8f
MD5 58df1a8e75f50a04622ddb209ab6ed33
BLAKE2b-256 4aae0f3cfd99ea49b733a14665319c34964252c44ecba967fb1c3ccd5a6d0b4c

See more details on using hashes here.

Release history Release notifications | RSS feed

0.2.2

2 files

0.2.1

2 files

0.2.0

2 files

0.1.7

2 files

This release

0.1.6 This release

2 files

0.1.2

2 files

0.1.1

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page