Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

arbtok

Rule-based Arabic (MSA) text→IPA with tashkeel diacritization — a downstream Arabic engine built on orthography2ipa.

Word phonology is built on the orthography2ipa shared lattice: the language-agnostic grapheme tokenizer (PhonetokTokenizer) over the ar spec grapheme table produces a per-position candidate lattice. The ar engine (orthography2ipa ≥ 1.70) handles the segment-local phonology natively — gemination (shadda ّ, glides included), lam-alif / presentation ligatures (ﻻ → laː), onset glides (يَ → ja), a hamza carrier's bare /ʔ/ before an explicit harakah, a fatḥa + standalone alif maksūra as one long vowel (حَتَّى → ħattaː), a sukūn-final coda glide (ظَبْي → ðˤabj, رَمْي → ramj, while فِي stays fiː), and pausal tāʾ marbūṭa. The last two were once patched by arbtok's own MaterLectionisRescorer / GlideCodaRescorer; orthography2ipa 1.70 (upstream #251) fixed them at source, so those rescorers are gone. The Arabic morpho-phonology that the shared table still cannot express is layered on as composable LatticeRescorers (arbtok/lattice.py) rather than a private tokenizer fork:

  • sun-letter assimilation (idghām ash-shamsiyya) — the lām of the definite article ⟨ال⟩ assimilates into a following coronal (sun) letter (al-šamsaš-šams); moon letters keep the lām (al-qamar);
  • hamzat al-waṣl elision — a word-initial prosthetic alif is silent, its harakah carrying the vowel (istiqbāl);
  • accusative-alif silencing after tanwīn al-fatḥ (marħaban), and the bare glottal stop of a hamza carrier before a sukūn or word edge (taʔθīr).

Emphatic (pharyngealization) spreading rides on the ar spec's own B8 allophone_rules. Cross-word sandhi — clitic joining, cross-word waṣl elision, tanwīn pausal forms, tāʾ marbūṭa, and idgham/iqlab nasal assimilation — is orthogonal to the word lattice and lives in the sentence-level orchestration. orthography2ipa 1.70 also added a shared sentence-context seam (orthography2ipa.sentence: SentenceLattice + SentenceRescorer with prev_word/next_word edge slots and is_phrase_final), the sanctioned home for that cross-word layer; arbtok's migration of its space-boundary waṣl elision and tanwīn pausal forms onto the seam is in progress (see docs/ and the tracking notes). Bare (undiacritized) text is diacritized first via text2tashkeel — a model picker over bundled ONNX diacritization models.

Honesty note: the gold IPA reference set was LLM-generated and has not been validated by a native MSA speaker. If you speak MSA, pull requests are very welcome.

Installation

pip install arbtok

Usage

arbtok is built on orthography2ipa (spec data and the shared G2PPlugin/WordContext base types) and owns the Arabic pipeline — orthography2ipa stays the language-agnostic base library.

Engine class

from arbtok.tokenizer import Sentence

Sentence("اَلسَّلَامُ عَلَيْكُمْ").ipa

An isolated MSA word transcribes on the shared lattice directly:

from arbtok.lattice import word_ipa

word_ipa("الشَّمْس")   # 'aʃʃams' — sun-letter assimilation as a rescorer
word_ipa("الْقَمَر")   # 'alqamar' — moon-letter control (lām kept)

Bare text is handled by diacritizing first:

from arbtok.plugin import ArbtokG2PPlugin

plugin = ArbtokG2PPlugin()
plugin.transcribe("كتاب جميل")    # auto-tashkeel + IPA

Diacritization only

from arbtok.tashkeel import TashkeelDiacritizer   # wraps text2tashkeel

TashkeelDiacritizer().diacritize("كتاب جميل")

Quality benchmarks

The test suite pins a gold sentence set (CER target ≤ 5% against the reference transcriptions) and benchmarks against espeak-ng. See tests/test_ipa_fuzzy.py and docs/ for details.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

arbtok-0.0.0a12.tar.gz (130.1 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

arbtok-0.0.0a12-py3-none-any.whl (117.7 kB view details)

Uploaded Python 3

File details

Details for the file arbtok-0.0.0a12.tar.gz.

File metadata

  • Download URL: arbtok-0.0.0a12.tar.gz
  • Upload date:
  • Size: 130.1 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for arbtok-0.0.0a12.tar.gz
Algorithm Hash digest
SHA256 433634a49c3d21a75a5df9bc7427c3e9f255c3ebf8ec37350e40411dbbfdfdda
MD5 7b50973e79e672b8a89df71bbc9f19e3
BLAKE2b-256 dd1ac374279f968651bf16eda5e55176c2689f783327bbc4970961730bc39a38

See more details on using hashes here.

File details

Details for the file arbtok-0.0.0a12-py3-none-any.whl.

File metadata

  • Download URL: arbtok-0.0.0a12-py3-none-any.whl
  • Upload date:
  • Size: 117.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for arbtok-0.0.0a12-py3-none-any.whl
Algorithm Hash digest
SHA256 da30ecce7451c9768b614c1baf9041f465f797f48e5c5c77f45381bc9af25554
MD5 9e8b4764199fa506e354c6e291cdb1cd
BLAKE2b-256 4fd9ba17e952d33bbd9af317dce21268f7def328bef8716621cf32ec9905f0f4

See more details on using hashes here.

Release history Release notifications | RSS feed

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page