This release is a pre-release and may not be stable for production use.
arbtok
Rule-based Arabic (MSA) text→IPA with tashkeel diacritization — a downstream Arabic engine built on orthography2ipa.
Word phonology is built on the orthography2ipa shared lattice: the
language-agnostic grapheme tokenizer (PhonetokTokenizer) over the ar
spec grapheme table produces a per-position candidate lattice. The ar
engine (orthography2ipa ≥ 1.70) handles the segment-local phonology
natively — gemination (shadda ّ, glides included), lam-alif / presentation
ligatures (ﻻ → laː), onset glides (يَ → ja), a hamza carrier's bare
/ʔ/ before an explicit harakah, a fatḥa + standalone alif maksūra as one
long vowel (حَتَّى → ħattaː), a sukūn-final coda glide (ظَبْي → ðˤabj,
رَمْي → ramj, while فِي stays fiː), and pausal tāʾ marbūṭa. The last two
were once patched by arbtok's own MaterLectionisRescorer /
GlideCodaRescorer; orthography2ipa 1.70 (upstream #251) fixed them at
source, so those rescorers are gone. The Arabic morpho-phonology that the
shared table still cannot express is layered on as composable
LatticeRescorers (arbtok/lattice.py) rather than a private tokenizer
fork:
- sun-letter assimilation (idghām ash-shamsiyya) — the lām of the
definite article ⟨ال⟩ assimilates into a following coronal (sun) letter
(
al-šams→aš-šams); moon letters keep the lām (al-qamar); - hamzat al-waṣl elision — a word-initial prosthetic alif is silent,
its harakah carrying the vowel (
istiqbāl); - accusative-alif silencing after tanwīn al-fatḥ (
marħaban), and the bare glottal stop of a hamza carrier before a sukūn or word edge (taʔθīr).
Emphatic (pharyngealization) spreading rides on the ar spec's own B8
allophone_rules. Cross-word sandhi — clitic joining, cross-word waṣl
elision, tanwīn pausal forms, tāʾ marbūṭa, and idgham/iqlab nasal
assimilation — is orthogonal to the word lattice and lives in the
sentence-level orchestration. orthography2ipa 1.70 also added a shared
sentence-context seam (orthography2ipa.sentence: SentenceLattice +
SentenceRescorer with prev_word/next_word edge slots and
is_phrase_final), the sanctioned home for that cross-word layer; arbtok's
migration of its space-boundary waṣl elision and tanwīn pausal forms onto
the seam is in progress (see docs/ and the tracking notes). Bare
(undiacritized) text is diacritized
first via text2tashkeel —
a model picker over bundled ONNX diacritization models.
Honesty note: the gold IPA reference set was LLM-generated and has not been validated by a native MSA speaker. If you speak MSA, pull requests are very welcome.
Installation
pip install arbtok
Usage
arbtok is built on orthography2ipa
(spec data and the shared G2PPlugin/WordContext base types) and owns the
Arabic pipeline — orthography2ipa stays the language-agnostic base library.
Engine class
from arbtok.tokenizer import Sentence
Sentence("اَلسَّلَامُ عَلَيْكُمْ").ipa
An isolated MSA word transcribes on the shared lattice directly:
from arbtok.lattice import word_ipa
word_ipa("الشَّمْس") # 'aʃʃams' — sun-letter assimilation as a rescorer
word_ipa("الْقَمَر") # 'alqamar' — moon-letter control (lām kept)
Bare text is handled by diacritizing first:
from arbtok.plugin import ArbtokG2PPlugin
plugin = ArbtokG2PPlugin()
plugin.transcribe("كتاب جميل") # auto-tashkeel + IPA
Diacritization only
from arbtok.tashkeel import TashkeelDiacritizer # wraps text2tashkeel
TashkeelDiacritizer().diacritize("كتاب جميل")
Quality benchmarks
The test suite pins a gold sentence set (CER target ≤ 5% against the
reference transcriptions) and benchmarks against espeak-ng. See
tests/test_ipa_fuzzy.py and docs/ for details.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file arbtok-0.0.0a18.tar.gz.
File metadata
- Download URL: arbtok-0.0.0a18.tar.gz
- Upload date:
- Size: 150.0 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
f3808a566694a6c83b8d6d94d72eb925d1b034c7e41d20c7a2f68ffba43b6dae
|
|
| MD5 |
e6f6288265175d38150b99ffadee6b5a
|
|
| BLAKE2b-256 |
631b6afc2793666c7d07cd7400b8c038b66e16c91a3e0eb9c6beb2ff2fba8e9d
|
File details
Details for the file arbtok-0.0.0a18-py3-none-any.whl.
File metadata
- Download URL: arbtok-0.0.0a18-py3-none-any.whl
- Upload date:
- Size: 135.9 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
fe0c079b838e86f72e50439606baeafc9e28434d59d0945f5d0ffb2420f68c87
|
|
| MD5 |
cdd544130613259fda954f29efff47c5
|
|
| BLAKE2b-256 |
316dd7474219f8eae9e74058aa442c2e21906345d09302c9137da5f5e6c6fb1d
|