Skip to main content

babelcode

Normalize any language identifier — ISO 639-1 (en), ISO 639-3 (eng), BCP-47 (zh-Hans), NLLB-style (eng_Latn), WikiPron filenames (wp_eng_latn_us), CHILDES corpus names (EnglishNA), or plain English (German) — into a single canonical form:

{iso_639_3}_{iso_15924}

For example: eng_Latn, arb_Arab, cmn_Hans, jpn_Jpan.

Review the mapping decisions used in this project.

Installation

pip install babelcode

Quick start

from babelcode import BabelCode

bc = BabelCode()

bc.normalize("en")          # → "eng_Latn"
bc.normalize("zh-Hans")     # → "cmn_Hans"
bc.normalize("ar")          # → "arb_Arab"  (macrolanguage → preferred individual)
bc.normalize("wp_deu_latn") # → "deu_Latn"  (WikiPron format)
bc.normalize("Farsi")       # → "pes_Arab"  (English name)
bc.normalize("EnglishNA")   # → "eng_Latn"  (CHILDES corpus name)

Singleton shortcut

from babelcode import get_instance

bc = get_instance()  # cached singleton — same instance every call

Accessors

bc.iso639_3("eng_Latn")        # → "eng"
bc.script("eng_Latn")          # → "Latn"
bc.bcp47("eng_Latn")           # → "en"
bc.name("arb")                 # → "Standard Arabic"
bc.scripts("srp")              # → ["Cyrl", "Latn"]
bc.is_macrolanguage("ara")     # → True
bc.macro_members("ara")        # → ["arb", "arz", ...]

Script detection from text

from babelcode import detect_script

detect_script("مرحبا")     # → "Arab"
detect_script("Привет")    # → "Cyrl"
detect_script("こんにちは")  # → "Jpan"

Batch operations

bc.normalize_list(["en", "de", "unknown", "fr"])
# → ["eng_Latn", "deu_Latn", None, "fra_Latn"]

bc.build_mapping(
    source_codes=["en", "de"],
    target_codes=["eng_Latn", "deu_Latn", "fra_Latn"],
)
# → {"en": "eng_Latn", "de": "deu_Latn"}

GlotLID Compatibility Mode

To follow the same decisions made by the GlotLID project use glotlid=True when calling normalize.

from babelcode import BabelCode

bc = BabelCode()

bc.normalize("Farsi", glotlid=True)             # → "fas_Arab"
bc.normalize("ktu_Latn", glotlid=True)          # → "kng_Latn"
bc.normalize("tgl_Latn", glotlid=True)          # → "fil_Latn"

Why iso3_Script instead of BCP-47?

BCP-47 tags (en, zh-Hans) are familiar but they conflate macrolanguages with individual languages, have variable length, and omit the script when it is "obvious" — which is ambiguous for multi-script languages like Serbian.

The {iso_639_3}_{iso_15924} canonical form:

  • Always has exactly two components — easy to split and compare.
  • Uses ISO 639-3 — one code per individual language, no macrolanguage ambiguity.
  • Always includes the ISO 15924 script — no guessing for Serbian (srp_Cyrl vs srp_Latn) or Chinese (cmn_Hans vs cmn_Hant).

Data sources & methodology

babelcode is pre-built from LinguaMeta, a comprehensive open dataset from Google Research that aggregates:

Source What it provides
ISO 639-3 (SIL) Three-letter codes for 7 800+ languages
ISO 15924 (Unicode) Four-letter script codes (Latn, Arab, …)
IETF BCP 47 Standard language tags (en, zh-Hans, …)
LinguaMeta JSON files (~7 500) Canonical script, macrolanguage membership, English names

Build process

  1. Raw data: ~7500 per-language JSON files from LinguaMeta, each containing script associations, name data, and macrolanguage membership.
  2. Cache compilation (babelcode-build-cache): Reads all JSON files and produces a single linguameta_cache.json (~550 KB) with pre-resolved BCP-47 → ISO 639-3 mappings, canonical scripts, English names, and macrolanguage membership tables.
  3. Runtime: BabelCode loads the cache once and resolves any input format through a cascade of regex matchers and lookup tables.

The cache is distributed inside the package — no network calls at runtime, zero dependencies.

Macrolanguage resolution

By default, macrolanguages resolve to their preferred individual language (e.g. ar → arb Standard Arabic, zh → cmn Mandarin). This can be disabled with resolve_macro=False.

Script inference

When a script is not explicit in the input, babelcode uses this cascade:

  1. Script from the input tag itself (e.g. zh-Hans → Hans)
  2. text_hint parameter — runs Unicode-range script detection on sample text
  3. Canonical script from the LinguaMeta cache
  4. Fallback to Latn

Development

git clone https://github.com/omneity-labs/babelcode
cd babelcode
pip install -e ".[dev]"
pytest

Rebuilding the cache

If you update the LinguaMeta data files under src/babelcode/data/_url_nlp_repo/, rebuild the cache:

babelcode-build-cache

License

MIT — Omar Kamali / Omneity Labs

Metadata

Release files for babelcode 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for babelcode 0.1.1
File Size Uploaded
babelcode-0.1.1.tar.gz 406.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for babelcode 0.1.1
File Interpreter ABI Platform
babelcode-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 812.1 kB

Release files / babelcode-0.1.1.tar.gz

Download URL babelcode-0.1.1.tar.gz
Size 406.8 kB
Tags Source
SHA-256 checksum
How to use checksums
331bb0bb92c93d0de116a39a04c8c78412298595283d322ce3cae1e22f16268c
BLAKE2b-256 checksum
How to use checksums
39b7d90797b7d4c3d192da18855e152e9c744ff82c8632ca30f5e7710af82493
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.7

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Apr 4, 2026.

Transparency log

Release files / babelcode-0.1.1-py3-none-any.whl

Download URL babelcode-0.1.1-py3-none-any.whl
Size 405.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
ed98fb66f0a52dcf626c431becf36dceec773f2ff2d8aec951b982683f6479da
BLAKE2b-256 checksum
How to use checksums
a540a953107d2282af9638f0c0978d49bddf97df52be81400d8f9791d2c63656
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.7

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Apr 4, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.1 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page