Skip to main content

Dialect-aware Basque (Euskara) text-to-IPA phonemizer over the orthography2ipa lattice

Project description

euskaphone

Dialect-aware Basque (Euskara) text-to-IPA phonemizer — a TTS frontend built on the shared orthography2ipa pronunciation lattice. euskaphone is the Basque sibling of tugaphone (Portuguese) and mwl_phonemizer (Mirandese), and is offered to the Basque research and speech-technology community as an open, inspectable, source-cited pipeline.

A dialect is an orthography2ipa Basque lect spec: the spec's grapheme table, allophony and cross-word sandhi produce the dialect's phonology directly — the apical/laminal sibilant contrast, the affricates, Souletin aspiration and /y/. euskaphone adds only the stages orthography2ipa leaves to the caller.

Install

pip install euskaphone

Quickstart

from euskaphone import EuskaPhonemizer

ph = EuskaPhonemizer()
ph.phonemize_sentence("Zazpi katu zuri ikusi ditut.")          # Batua (default)
# 's̻as̻pi katu s̻uɾi ikus̺i ditut'

ph.phonemize_sentence("Hotza egiten du gaur mendian.", "souletin")
# 'hots̻a eɡiten dy ɡauɾ mendian'   (Souletin: aspiration, /y/)

Dialects

Eight lects, selectable by BCP-47 code or by a human-readable alias:

code dialect aliases
eu Standard Basque batua, standard
eu-x-bizkaiera Biscayan biscayan, western
eu-x-gipuzkera Gipuzkoan guipuzcoan, central
eu-x-lapurtera Lapurdian labourdin
eu-x-nafarra-garaia High Navarrese high-navarrese
eu-x-nafarra-beherea Low Navarrese low-navarrese
eu-x-zuberera Souletin souletin, xiberotarra
eu-x-erronkariera Roncalese (†) roncalese

The three things euskaphone adds

1. Vigesimal number normalization (the flagship)

Basque cardinals are base-20: hogei (20), berrogei (40 = 2×20), hirurogei (60), laurogei (80); 30/50/70/90 are the preceding twenty plus ten; and the copulative eta ("and") contracts to -ta on the twenties.

from euskaphone.number_utils import BasqueNumberParser, normalize_numbers

p = BasqueNumberParser("eu")
p.cardinal(31)     # 'hogeita hamaika'   (20-and-11)
p.cardinal(1936)   # 'mila bederatziehun eta hogeita hamasei'
p.ordinal(16)      # 'hamaseigarren'

# case-suffixed numerals and clock time in running text (declension on the
# last element of the numeral phrase)
normalize_numbers("1936ko gerra", "eu")   # '… hogeita hamaseiko gerra'
normalize_numbers("20:00etan", "eu")      # 'hogeietan'

Every base form is cited to Euskaltzaindia Araua 7 (cardinals) and Araua 18 (ordinals), and each is tagged attested-in-source vs derived-by-rule in euskaphone.number_utils (the ATTESTED / DERIVED sets). One notable correction the sources force: Araua 7 records bost (not bortz) for Zuberoa and most of Low Navarre, so euskaphone keeps bost everywhere by default and offers the bortz series only for Lapurdian — null beats a wrong dialect swap. See docs/numbers.md.

2. Code-switch handling

Real Basque text embeds Spanish (Hegoalde), French (Iparralde) and — in the tech/music/media register — English. Word-level detection identifies the embedded material, classifies each contact word per word among es/fr/en, routes it through the orthography2ipa es-ES/fr-FR/en-US lattice, and nativizes the result onto the Basque inventory (project every foreign phone onto its nearest Basque phone — never drop, always project). The contact parameter is auto (per-word detection), es, fr, en, or none.

Detection prefers a small bundled char-Markov language detector (Basque / Spanish / French / English, ~180 KB total, scored by markovonnx; install with euskaphone[langdetect]) and falls back to an orthographic heuristic when the models are unavailable. Basque is the in-language default: a word is routed out of Basque only when a foreign model beats the Basque model by a clear margin (and short, plausibly-native words need orthographic corroboration), so a weak signal never misroutes a native word.

ph.phonemize_sentence("Madrilen Plaza Mayor ikusi dut.", "eu")
# '… plas̻a majoɾ …'  (Spanish routed + nativized: no θ, no stress)

ph.phonemize_sentence("Streaming plataforma berria erabili dugu.")
# 'strimin platafoɾma …'  (English routed + nativized onto the 5-vowel system)

ph.phonemize_sentence("Maison Rouge etxean.", "souletin")
# continental dialect → French contact by default

See docs/codeswitch.md.

3. Lexicon hook + shipped lexicons

register_lexicon(dialect, source) registers a proper-name/loan pronunciation table (the orthography2ipa lexicon contract: a word<TAB>ipa path, URL or hf:// id); a covered word bypasses the lattice.

Two lexicons ship built-in: a curated Euskaltzaindia toponym seed (on by default, EuskaPhonemizer(toponyms=False) to opt out) and an opt-in HiTZ proper-noun overlay (EuskaPhonemizer(lexicon="hitz")) that carries a benchmark circularity caveat. A caller's own lexicon always wins. See docs/lexicons.md.

4. Orthographic normalization

Before the lattice runs, everything written as a symbol, abbreviation, Roman numeral or numeric date is spelled into Basque words (euskaphone.normalize.normalize_text, wired in as the orthography2ipa normalize stage):

from euskaphone.normalize import normalize_text

normalize_text("XX. mendean %20 igo da", "eu")
# 'hogeigarren mendean ehuneko hogei igo da'
normalize_text("Luis XIV.a errege zen", "eu")     # '… hamalaugarrena …'
normalize_text("Sarrera €5eko da", "eu")          # '… bost euroko …'
normalize_text("15/01/2024", "eu")   # 'bi mila eta hogeita lauko urtarrilaren hamabostean'

Roman-numeral ordinals (XX.hogeigarren) and monarch numerals (Luis XIV.ahamalaugarrena), a cited abbreviation list (etab.eta abar), unit/currency/percent symbols with Basque word order and declension (%5 and 5% both → ehuneko bost), and numeric dates — each cited to Euskaltzaindia (Araua 18/37/196/197) and the EIMA Ortotipografia guide. The abbreviation and unit tables are expandable data files (euskaphone/data/*.tsv). See docs/normalization.md.

No homograph subsystem — by design

Basque orthography is near-phonemic: a grapheme's reading is essentially fixed and sense-independent. Unlike tugaphone's Portuguese (which needs bifonia to resolve heterophonic homographs), euskaphone deliberately ships no homograph stage. This is a scope decision, documented so it is not mistaken for a gap.

OVOS plugin

from euskaphone.plugin import EuskaphoneG2PPlugin
EuskaphoneG2PPlugin("souletin").transcribe("kaixo mundua")

Registered under the opm.g2p entry-point group as euskaphone.

How it measures up

euskaphone drives the shared orthography2ipa lattice, so it is only as accurate as that lattice — these numbers measure the lattice, honestly labelled. Run them with python scripts/benchmark.py [--hitz].

gold kind figure
orthography2ipa eu gold (166 rows, 8 lects) regression fixture (same lattice made it) PER 0.0000 — proves the lattice is untouched, not accuracy
WikiPron eus_latn broad independent word gold PER 0.0749, word-acc 0.4563 (full set, 20,115 words)
HiTZ/EHU Wikipedia G2P independent cross-engine (Univ. of the Basque Country) PER 0.1384 folded (3k-sentence sample)

The WikiPron and HiTZ figures are the honest ones: a pure lattice with no lexicon, on proper-name-heavy and raw-Wikipedia text, where transcription conventions and multiple valid pronunciations account for much of the gap. That gap is exactly what the lexicon hook exists to close. Comparable open tools: AhoTTS (the EHU TTS front end) covers Standard Batua only; espeak-ng's eu voice is single-dialect and rule-thin — euskaphone's contribution is the eight-lect coverage, the sourced vigesimal numerals, and the code-switch nativization, over a lattice shared across the whole Iberian family.

License

Apache-2.0.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

euskaphone-0.2.0a1.tar.gz (252.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

euskaphone-0.2.0a1-py3-none-any.whl (238.6 kB view details)

Uploaded Python 3

File details

Details for the file euskaphone-0.2.0a1.tar.gz.

File metadata

  • Download URL: euskaphone-0.2.0a1.tar.gz
  • Upload date:
  • Size: 252.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for euskaphone-0.2.0a1.tar.gz
Algorithm Hash digest
SHA256 819de4d9f8ac2e34204fd577eb5638007b80842b7b84c2ad2825c00baf1cfbce
MD5 11c632e43bf8b0f24129391b5b881a0b
BLAKE2b-256 59fc441c4e1af4c9f206cddc888fada1e87d4b4f8bbc585ae9beda9f3026271c

See more details on using hashes here.

File details

Details for the file euskaphone-0.2.0a1-py3-none-any.whl.

File metadata

  • Download URL: euskaphone-0.2.0a1-py3-none-any.whl
  • Upload date:
  • Size: 238.6 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for euskaphone-0.2.0a1-py3-none-any.whl
Algorithm Hash digest
SHA256 8032cffcea1fd1c3694e92667d97fcb08b38cf517906373e7ab459a3fb0f6764
MD5 f5bd4f7c309b470dfcbd93067f8d58f6
BLAKE2b-256 90de8469bc4ce793daac4abeeea79470bbc08974f8a4fc03cf0d167605456109

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page