Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

arbtok

Rule-based Arabic (MSA) text→IPA with tashkeel diacritization — a downstream Arabic engine built on orthography2ipa.

Word phonology is built on the orthography2ipa shared lattice: the language-agnostic grapheme tokenizer (PhonetokTokenizer) over the ar spec grapheme table produces a per-position candidate lattice. The ar engine (orthography2ipa ≥ 1.70) handles the segment-local phonology natively — gemination (shadda ّ, glides included), lam-alif / presentation ligatures (ﻻ → laː), onset glides (يَ → ja), a hamza carrier's bare /ʔ/ before an explicit harakah, a fatḥa + standalone alif maksūra as one long vowel (حَتَّى → ħattaː), a sukūn-final coda glide (ظَبْي → ðˤabj, رَمْي → ramj, while فِي stays fiː), and pausal tāʾ marbūṭa. The last two were once patched by arbtok's own MaterLectionisRescorer / GlideCodaRescorer; orthography2ipa 1.70 (upstream #251) fixed them at source, so those rescorers are gone. The Arabic morpho-phonology that the shared table still cannot express is layered on as composable LatticeRescorers (arbtok/lattice.py) rather than a private tokenizer fork:

  • sun-letter assimilation (idghām ash-shamsiyya) — the lām of the definite article ⟨ال⟩ assimilates into a following coronal (sun) letter (al-šamsaš-šams); moon letters keep the lām (al-qamar);
  • hamzat al-waṣl elision — a word-initial prosthetic alif is silent, its harakah carrying the vowel (istiqbāl);
  • accusative-alif silencing after tanwīn al-fatḥ (marħaban), and the bare glottal stop of a hamza carrier before a sukūn or word edge (taʔθīr).

Emphatic (pharyngealization) spreading rides on the ar spec's own B8 allophone_rules. Cross-word sandhi — clitic joining, cross-word waṣl elision, tanwīn pausal forms, tāʾ marbūṭa, and idgham/iqlab nasal assimilation — is orthogonal to the word lattice and lives in the sentence-level orchestration. orthography2ipa 1.70 also added a shared sentence-context seam (orthography2ipa.sentence: SentenceLattice + SentenceRescorer with prev_word/next_word edge slots and is_phrase_final), the sanctioned home for that cross-word layer; arbtok's migration of its space-boundary waṣl elision and tanwīn pausal forms onto the seam is in progress (see docs/ and the tracking notes). Bare (undiacritized) text is diacritized first via text2tashkeel — a model picker over bundled ONNX diacritization models.

Honesty note: the gold IPA reference set was LLM-generated and has not been validated by a native MSA speaker. If you speak MSA, pull requests are very welcome.

Installation

pip install arbtok

Usage

arbtok is built on orthography2ipa (spec data and the shared G2PPlugin/WordContext base types) and owns the Arabic pipeline — orthography2ipa stays the language-agnostic base library.

Engine class

from arbtok.tokenizer import Sentence

Sentence("اَلسَّلَامُ عَلَيْكُمْ").ipa

An isolated MSA word transcribes on the shared lattice directly:

from arbtok.lattice import word_ipa

word_ipa("الشَّمْس")   # 'aʃʃams' — sun-letter assimilation as a rescorer
word_ipa("الْقَمَر")   # 'alqamar' — moon-letter control (lām kept)

Bare text is handled by diacritizing first:

from arbtok.plugin import ArbtokG2PPlugin

plugin = ArbtokG2PPlugin()
plugin.transcribe("كتاب جميل")    # auto-tashkeel + IPA

Varieties

Pass a spec code as lang= to phonemize a variety; arbtok.supported_lects() lists every code it resolves to, with the orthography2ipa quality tier of each. Bare (undiacritized) input is restored on MSA orthography before dialect allophony applies — the diacritizer and stem lexicon are MSA artifacts. See docs/dialects.md for the resolution rules, the supported list, and the pinned pipeline order.

import arbtok
from arbtok.plugin import ArbtokG2PPlugin

arbtok.supported_lects()[:2]                                   # [Lect('ar', 'research'), …]
ArbtokG2PPlugin(lang="ar-SA-x-najd").transcribe_word("قَهْوَة")  # 'ˈɡahawa'

Diacritization only

from arbtok.tashkeel import TashkeelDiacritizer   # wraps text2tashkeel

TashkeelDiacritizer().diacritize("كتاب جميل")

Quality benchmarks

The test suite pins a gold sentence set (CER target ≤ 5% against the reference transcriptions) and benchmarks against espeak-ng. See tests/test_ipa_fuzzy.py and docs/ for details.

For per-lect scoring — every resolvable variety against the orthography2ipa Arabic TTS gold, diacritized and bare, next to espeak-ng — run python scripts/benchmark_stack.py --lect and see docs/benchmarks.md, which carries the full table and the honesty note on why those figures are engine-similarity to cited-rule o2i output rather than native-validated truth.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

arbtok-0.0.0a20.tar.gz (155.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

arbtok-0.0.0a20-py3-none-any.whl (137.3 kB view details)

Uploaded Python 3

File details

Details for the file arbtok-0.0.0a20.tar.gz.

File metadata

  • Download URL: arbtok-0.0.0a20.tar.gz
  • Upload date:
  • Size: 155.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for arbtok-0.0.0a20.tar.gz
Algorithm Hash digest
SHA256 b51d712e9253b39966bb838c1b9542e417a91a55e46c4db0e0f27afdcc46ee34
MD5 cf73137563dd8fdddd51b2b99730eb4f
BLAKE2b-256 efe3833c374a122aba726187c0d1fc53c6fb82241a64527ae18d9ad0b2a4f449

See more details on using hashes here.

File details

Details for the file arbtok-0.0.0a20-py3-none-any.whl.

File metadata

  • Download URL: arbtok-0.0.0a20-py3-none-any.whl
  • Upload date:
  • Size: 137.3 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for arbtok-0.0.0a20-py3-none-any.whl
Algorithm Hash digest
SHA256 a7c170fd82a318676367ed491f1ff8315844ae53e5772fe558bc6c8ff0577932
MD5 22395c4b7775e28fac7ddbf1299905a2
BLAKE2b-256 2a3bd28272cd2ecac0e1729a26f24d1a9961eaaba3fbd1b2f9e590be42263555

See more details on using hashes here.

Release history Release notifications | RSS feed

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page