Skip to main content

Dialect-aware Basque (Euskara) text-to-IPA phonemizer over the orthography2ipa lattice

Project description

euskaphone

Dialect-aware Basque (Euskara) text-to-IPA phonemizer — a TTS frontend built on the shared orthography2ipa pronunciation lattice. euskaphone is the Basque sibling of tugaphone (Portuguese) and mwl_phonemizer (Mirandese), and is offered to the Basque research and speech-technology community as an open, inspectable, source-cited pipeline.

A dialect is an orthography2ipa Basque lect spec: the spec's grapheme table, allophony and cross-word sandhi produce the dialect's phonology directly — the apical/laminal sibilant contrast, the affricates, Souletin aspiration and /y/. euskaphone adds only the stages orthography2ipa leaves to the caller.

Install

pip install euskaphone

Quickstart

from euskaphone import EuskaPhonemizer

ph = EuskaPhonemizer()
ph.phonemize_sentence("Zazpi katu zuri ikusi ditut.")          # Batua (default)
# 's̻as̻pi katu s̻uɾi ikus̺i ditut'

ph.phonemize_sentence("Hotza egiten du gaur mendian.", "souletin")
# 'hots̻a eɡiten dy ɡauɾ mendian'   (Souletin: aspiration, /y/)

Dialects

Eight lects, selectable by BCP-47 code or by a human-readable alias:

code dialect aliases
eu Standard Basque batua, standard
eu-x-bizkaiera Biscayan biscayan, western
eu-x-gipuzkera Gipuzkoan guipuzcoan, central
eu-x-lapurtera Lapurdian labourdin
eu-x-nafarra-garaia High Navarrese high-navarrese
eu-x-nafarra-beherea Low Navarrese low-navarrese
eu-x-zuberera Souletin souletin, xiberotarra
eu-x-erronkariera Roncalese (†) roncalese

The three things euskaphone adds

1. Vigesimal number normalization (the flagship)

Basque cardinals are base-20: hogei (20), berrogei (40 = 2×20), hirurogei (60), laurogei (80); 30/50/70/90 are the preceding twenty plus ten; and the copulative eta ("and") contracts to -ta on the twenties.

from euskaphone.number_utils import BasqueNumberParser, normalize_numbers

p = BasqueNumberParser("eu")
p.cardinal(31)     # 'hogeita hamaika'   (20-and-11)
p.cardinal(1936)   # 'mila bederatziehun eta hogeita hamasei'
p.ordinal(16)      # 'hamaseigarren'

# case-suffixed numerals and clock time in running text (declension on the
# last element of the numeral phrase)
normalize_numbers("1936ko gerra", "eu")   # '… hogeita hamaseiko gerra'
normalize_numbers("20:00etan", "eu")      # 'hogeietan'

Every base form is cited to Euskaltzaindia Araua 7 (cardinals) and Araua 18 (ordinals), and each is tagged attested-in-source vs derived-by-rule in euskaphone.number_utils (the ATTESTED / DERIVED sets). One notable correction the sources force: Araua 7 records bost (not bortz) for Zuberoa and most of Low Navarre, so euskaphone keeps bost everywhere by default and offers the bortz series only for Lapurdian — null beats a wrong dialect swap. See docs/numbers.md.

2. Code-switch handling

Real Basque text embeds Spanish (Hegoalde), French (Iparralde) and — in the tech/music/media register — English. Word-level detection identifies the embedded material, classifies each contact word per word among es/fr/en, routes it through the orthography2ipa es-ES/fr-FR/en-US lattice, and nativizes the result onto the Basque inventory (project every foreign phone onto its nearest Basque phone — never drop, always project). The contact parameter is auto (per-word detection), es, fr, en, or none.

Detection prefers a small bundled char-Markov language detector (Basque / Spanish / French / English, ~180 KB total, scored by markovonnx; install with euskaphone[langdetect]) and falls back to an orthographic heuristic when the models are unavailable. Basque is the in-language default: a word is routed out of Basque only when a foreign model beats the Basque model by a clear margin (and short, plausibly-native words need orthographic corroboration), so a weak signal never misroutes a native word.

ph.phonemize_sentence("Madrilen Plaza Mayor ikusi dut.", "eu")
# '… plas̻a majoɾ …'  (Spanish routed + nativized: no θ, no stress)

ph.phonemize_sentence("Streaming plataforma berria erabili dugu.")
# 'strimin platafoɾma …'  (English routed + nativized onto the 5-vowel system)

ph.phonemize_sentence("Maison Rouge etxean.", "souletin")
# continental dialect → French contact by default

See docs/codeswitch.md.

3. Lexicon hook + shipped lexicons

register_lexicon(dialect, source) registers a proper-name/loan pronunciation table (the orthography2ipa lexicon contract: a word<TAB>ipa path, URL or hf:// id); a covered word bypasses the lattice.

Two lexicons ship built-in: a curated Euskaltzaindia toponym seed (on by default, EuskaPhonemizer(toponyms=False) to opt out) and an opt-in HiTZ proper-noun overlay (EuskaPhonemizer(lexicon="hitz")) that carries a benchmark circularity caveat. A caller's own lexicon always wins. See docs/lexicons.md.

4. Orthographic normalization

Before the lattice runs, everything written as a symbol, abbreviation, Roman numeral or numeric date is spelled into Basque words (euskaphone.normalize.normalize_text, wired in as the orthography2ipa normalize stage):

from euskaphone.normalize import normalize_text

normalize_text("XX. mendean %20 igo da", "eu")
# 'hogeigarren mendean ehuneko hogei igo da'
normalize_text("Luis XIV.a errege zen", "eu")     # '… hamalaugarrena …'
normalize_text("Sarrera €5eko da", "eu")          # '… bost euroko …'
normalize_text("15/01/2024", "eu")   # 'bi mila eta hogeita lauko urtarrilaren hamabostean'

Roman-numeral ordinals (XX.hogeigarren) and monarch numerals (Luis XIV.ahamalaugarrena), a cited abbreviation list (etab.eta abar), unit/currency/percent symbols with Basque word order and declension (%5 and 5% both → ehuneko bost), and numeric dates — each cited to Euskaltzaindia (Araua 18/37/196/197) and the EIMA Ortotipografia guide. The abbreviation and unit tables are expandable data files (euskaphone/data/*.tsv). See docs/normalization.md.

No homograph subsystem — by design

Basque orthography is near-phonemic: a grapheme's reading is essentially fixed and sense-independent. Unlike tugaphone's Portuguese (which needs bifonia to resolve heterophonic homographs), euskaphone deliberately ships no homograph stage. This is a scope decision, documented so it is not mistaken for a gap.

OVOS plugin

from euskaphone.plugin import EuskaphoneG2PPlugin
EuskaphoneG2PPlugin("souletin").transcribe("kaixo mundua")

Registered under the opm.g2p entry-point group as euskaphone.

How it measures up

euskaphone drives the shared orthography2ipa lattice, so it is only as accurate as that lattice — these numbers measure the lattice, honestly labelled. Run them with python scripts/benchmark.py [--hitz].

gold kind figure
orthography2ipa eu gold (166 rows, 8 lects) regression fixture (same lattice made it) PER 0.0000 — proves the lattice is untouched, not accuracy
WikiPron eus_latn broad independent word gold PER 0.0749, word-acc 0.4563 (full set, 20,115 words)
HiTZ/EHU Wikipedia G2P independent cross-engine (Univ. of the Basque Country) PER 0.1384 folded (3k-sentence sample)

The WikiPron and HiTZ figures are the honest ones: a pure lattice with no lexicon, on proper-name-heavy and raw-Wikipedia text, where transcription conventions and multiple valid pronunciations account for much of the gap. That gap is exactly what the lexicon hook exists to close. Measured against the comparable open tools on the same full WikiPron set (segment-level protocol, stress stripped — see docs/benchmarks.md): euskaphone PER 0.1009 / word-acc 0.4298, espeak-ng eu 0.1704 / 0.1750 (espyak, the pure-Python espeak port, scores identically), ahotts-g2p 0.3360 / 0.0003 (ahotts's figure is dominated by its non-IPA symbol conventions). AhoTTS covers Standard Batua only; espeak-ng's eu voice is single-dialect and rule-thin — euskaphone's contribution is the eight-lect coverage, the sourced vigesimal numerals, and the code-switch nativization, over a lattice shared across the whole Iberian family.

License

Apache-2.0.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

euskaphone-0.2.0a3.tar.gz (252.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

euskaphone-0.2.0a3-py3-none-any.whl (238.7 kB view details)

Uploaded Python 3

File details

Details for the file euskaphone-0.2.0a3.tar.gz.

File metadata

  • Download URL: euskaphone-0.2.0a3.tar.gz
  • Upload date:
  • Size: 252.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for euskaphone-0.2.0a3.tar.gz
Algorithm Hash digest
SHA256 847050d0c554c843a5034797fc1fdf289118f7fa41664862ebe41b55e0464b2c
MD5 5d363eb990d9dbe94ea69be169c74288
BLAKE2b-256 79043f5c3c69f7cc32203eedb8d1bc91c1b9e9c7f01e4b350149771ed9b96ae4

See more details on using hashes here.

File details

Details for the file euskaphone-0.2.0a3-py3-none-any.whl.

File metadata

  • Download URL: euskaphone-0.2.0a3-py3-none-any.whl
  • Upload date:
  • Size: 238.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for euskaphone-0.2.0a3-py3-none-any.whl
Algorithm Hash digest
SHA256 75ce64bca5e7aaedff0dd158b43635bfbf4a1ab47047fb20b6f9771a6eb137f7
MD5 12743a25ad66b570a37a2151a0bc1ec8
BLAKE2b-256 b03382eaad10bfae25b3688046d47939a01cde930f3582583cc9a11cb72c964f

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page