Dialect-aware Basque (Euskara) text-to-IPA phonemizer over the orthography2ipa lattice
Project description
euskaphone
Dialect-aware Basque (Euskara) text-to-IPA phonemizer — a TTS frontend built on the shared orthography2ipa pronunciation lattice. euskaphone is the Basque sibling of tugaphone (Portuguese) and mwl_phonemizer (Mirandese), and is offered to the Basque research and speech-technology community as an open, inspectable, source-cited pipeline.
A dialect is an orthography2ipa Basque lect spec: the spec's grapheme table,
allophony and cross-word sandhi produce the dialect's phonology directly — the
apical/laminal sibilant contrast, the affricates, Souletin aspiration and /y/.
euskaphone adds only the stages orthography2ipa leaves to the caller.
Install
pip install euskaphone
Quickstart
from euskaphone import EuskaPhonemizer
ph = EuskaPhonemizer()
ph.phonemize_sentence("Zazpi katu zuri ikusi ditut.") # Batua (default)
# 's̻as̻pi katu s̻uɾi ikus̺i ditut'
ph.phonemize_sentence("Hotza egiten du gaur mendian.", "souletin")
# 'hots̻a eɡiten dy ɡauɾ mendian' (Souletin: aspiration, /y/)
Dialects
Eight lects, selectable by BCP-47 code or by a human-readable alias:
| code | dialect | aliases |
|---|---|---|
eu |
Standard Basque | batua, standard |
eu-x-bizkaiera |
Biscayan | biscayan, western |
eu-x-gipuzkera |
Gipuzkoan | guipuzcoan, central |
eu-x-lapurtera |
Lapurdian | labourdin |
eu-x-nafarra-garaia |
High Navarrese | high-navarrese |
eu-x-nafarra-beherea |
Low Navarrese | low-navarrese |
eu-x-zuberera |
Souletin | souletin, xiberotarra |
eu-x-erronkariera |
Roncalese (†) | roncalese |
The three things euskaphone adds
1. Vigesimal number normalization (the flagship)
Basque cardinals are base-20: hogei (20), berrogei (40 = 2×20),
hirurogei (60), laurogei (80); 30/50/70/90 are the preceding twenty plus
ten; and the copulative eta ("and") contracts to -ta on the twenties.
from euskaphone.number_utils import BasqueNumberParser, normalize_numbers
p = BasqueNumberParser("eu")
p.cardinal(31) # 'hogeita hamaika' (20-and-11)
p.cardinal(1936) # 'mila bederatziehun eta hogeita hamasei'
p.ordinal(16) # 'hamaseigarren'
# case-suffixed numerals and clock time in running text (declension on the
# last element of the numeral phrase)
normalize_numbers("1936ko gerra", "eu") # '… hogeita hamaseiko gerra'
normalize_numbers("20:00etan", "eu") # 'hogeietan'
Every base form is cited to Euskaltzaindia Araua 7 (cardinals) and
Araua 18 (ordinals), and each is tagged attested-in-source vs derived-by-rule
in euskaphone.number_utils (the ATTESTED / DERIVED sets). One notable
correction the sources force: Araua 7 records bost (not bortz) for Zuberoa
and most of Low Navarre, so euskaphone keeps bost everywhere by default and
offers the bortz series only for Lapurdian — null beats a wrong dialect swap.
See docs/numbers.md.
2. Code-switch handling
Real Basque text embeds Spanish (Hegoalde), French (Iparralde) and — in the
tech/music/media register — English. Word-level detection identifies the embedded
material, classifies each contact word per word among es/fr/en, routes
it through the orthography2ipa es-ES/fr-FR/en-US lattice, and nativizes
the result onto the Basque inventory (project every foreign phone onto its
nearest Basque phone — never drop, always project). The contact parameter is
auto (per-word detection), es, fr, en, or none.
Detection prefers a small bundled char-Markov language detector (Basque /
Spanish / French / English, ~180 KB total, scored by markovonnx; install with
euskaphone[langdetect]) and falls back to an orthographic heuristic when the
models are unavailable. Basque is the in-language default: a word is routed out
of Basque only when a foreign model beats the Basque model by a clear margin (and
short, plausibly-native words need orthographic corroboration), so a weak signal
never misroutes a native word.
ph.phonemize_sentence("Madrilen Plaza Mayor ikusi dut.", "eu")
# '… plas̻a majoɾ …' (Spanish routed + nativized: no θ, no stress)
ph.phonemize_sentence("Streaming plataforma berria erabili dugu.")
# 'strimin platafoɾma …' (English routed + nativized onto the 5-vowel system)
ph.phonemize_sentence("Maison Rouge etxean.", "souletin")
# continental dialect → French contact by default
See docs/codeswitch.md.
3. Lexicon hook + shipped lexicons
register_lexicon(dialect, source) registers a proper-name/loan pronunciation
table (the orthography2ipa lexicon contract: a word<TAB>ipa path, URL or
hf:// id); a covered word bypasses the lattice.
Two lexicons ship built-in: a curated Euskaltzaindia toponym seed (on by
default, EuskaPhonemizer(toponyms=False) to opt out) and an opt-in HiTZ
proper-noun overlay (EuskaPhonemizer(lexicon="hitz")) that carries a
benchmark circularity caveat. A caller's own lexicon always wins.
See docs/lexicons.md.
4. Orthographic normalization
Before the lattice runs, everything written as a symbol, abbreviation, Roman
numeral or numeric date is spelled into Basque words
(euskaphone.normalize.normalize_text, wired in as the orthography2ipa
normalize stage):
from euskaphone.normalize import normalize_text
normalize_text("XX. mendean %20 igo da", "eu")
# 'hogeigarren mendean ehuneko hogei igo da'
normalize_text("Luis XIV.a errege zen", "eu") # '… hamalaugarrena …'
normalize_text("Sarrera €5eko da", "eu") # '… bost euroko …'
normalize_text("15/01/2024", "eu") # 'bi mila eta hogeita lauko urtarrilaren hamabostean'
Roman-numeral ordinals (XX. → hogeigarren) and monarch numerals
(Luis XIV.a → hamalaugarrena), a cited abbreviation list
(etab. → eta abar), unit/currency/percent symbols with Basque word order and
declension (%5 and 5% both → ehuneko bost), and numeric dates — each cited
to Euskaltzaindia (Araua 18/37/196/197) and the EIMA Ortotipografia guide. The
abbreviation and unit tables are expandable data files
(euskaphone/data/*.tsv). See docs/normalization.md.
No homograph subsystem — by design
Basque orthography is near-phonemic: a grapheme's reading is essentially fixed
and sense-independent. Unlike tugaphone's Portuguese (which needs bifonia to
resolve heterophonic homographs), euskaphone deliberately ships no homograph
stage. This is a scope decision, documented so it is not mistaken for a gap.
OVOS plugin
from euskaphone.plugin import EuskaphoneG2PPlugin
EuskaphoneG2PPlugin("souletin").transcribe("kaixo mundua")
Registered under the opm.g2p entry-point group as euskaphone.
How it measures up
euskaphone drives the shared orthography2ipa lattice, so it is only as accurate
as that lattice — these numbers measure the lattice, honestly labelled. Run them
with python scripts/benchmark.py [--hitz].
| gold | kind | figure |
|---|---|---|
| orthography2ipa eu gold (166 rows, 8 lects) | regression fixture (same lattice made it) | PER 0.0000 — proves the lattice is untouched, not accuracy |
WikiPron eus_latn broad |
independent word gold | PER 0.0749, word-acc 0.4563 (full set, 20,115 words) |
| HiTZ/EHU Wikipedia G2P | independent cross-engine (Univ. of the Basque Country) | PER 0.1384 folded (3k-sentence sample) |
The WikiPron and HiTZ figures are the honest ones: a pure lattice with no
lexicon, on proper-name-heavy and raw-Wikipedia text, where transcription
conventions and multiple valid pronunciations account for much of the gap. That
gap is exactly what the lexicon hook exists to close. Measured against the
comparable open tools on the same full WikiPron set (segment-level protocol,
stress stripped — see docs/benchmarks.md): euskaphone PER 0.1009 / word-acc
0.4298, espeak-ng eu 0.1704 / 0.1750, ahotts-g2p 0.3360 / 0.0003
(ahotts's figure is dominated by its non-IPA symbol conventions). AhoTTS
covers Standard Batua only; espeak-ng's
eu voice is single-dialect and rule-thin — euskaphone's contribution is the
eight-lect coverage, the sourced vigesimal numerals, and the code-switch
nativization, over a lattice shared across the whole Iberian family.
License
Apache-2.0.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file euskaphone-0.2.0a2.tar.gz.
File metadata
- Download URL: euskaphone-0.2.0a2.tar.gz
- Upload date:
- Size: 252.5 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
ad13f81ce98196235952f98d594696179b16095f1aab72ea40582b6ce806fe9f
|
|
| MD5 |
cf3cc57f20334ebf1c5458a6cbc16efc
|
|
| BLAKE2b-256 |
36f3da1be4b0aaaffd4fa5d48a765eaf8a191eab204edc1c1a93603982b7c928
|
File details
Details for the file euskaphone-0.2.0a2-py3-none-any.whl.
File metadata
- Download URL: euskaphone-0.2.0a2-py3-none-any.whl
- Upload date:
- Size: 238.7 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
deb84755606b9fe20fd493088a1a9a9ee1efba26d8b1a0cfd55c079344e68db6
|
|
| MD5 |
6da5b7865cbd8afcea80af3247a8e2c8
|
|
| BLAKE2b-256 |
a9949e6a93f8fad24c2df7426b326b28db3452a728991280193f4c3ebc908ab9
|