Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

Mirandese Phonemizer

Grapheme-to-phoneme (G2P) conversion for Mirandese (mwl), the Asturleonese language of Terra de Miranda, Portugal. It takes text and returns IPA, with cross-word sandhi, allophony, and stress.

from mwl_phonemizer import phonemize

phonemize("Falo la lhéngua mirandesa.")   # 'ˈfalu lɐ ˈʎɛŋɡwa miɾɐˈndez̺ɐ.'

Install

pip install mwl_phonemizer

This pulls in orthography2ipa, which carries the Mirandese language specs and gold data.

Usage

One-shot

from mwl_phonemizer import phonemize

phonemize("lhéngua")                          # 'ˈʎɛŋɡwa'
phonemize("fuogo", dialect="mwl-x-sendim")    # Sendinese variety

phonemize caches one phonemizer per dialect, so repeated calls are cheap.

Reusable instance

from mwl_phonemizer import MirandesePhonemizer

pho = MirandesePhonemizer(dialect="mwl")
pho.phonemize("Buonos dies, cumo stás?")   # full text, punctuation preserved
pho.phonemize_word("amportante")           # a single word -> 'ɐ̃puˈɾtɐ̃tɨ'

transcribe / transcribe_word are aliases of phonemize / phonemize_word. language_codes reports the BCP-47 codes the instance covers, the surface that downstream engines call.

Dialects

The dialect argument is an orthography2ipa Mirandese spec code:

code variety
mwl Central Mirandese (default)
mwl-x-sendim Sendinese, depalatalizes lh/initial l to [l]
mwl-x-ifanes Ifanês / Raiano (northern)
MirandesePhonemizer("mwl-x-sendim").phonemize("lhobo")   # 'ˈloβu', not 'ˈʎobu'

Numbers

Digits carry no orthography the lattice can read, so numeric tokens are spelled out into Mirandese words before transcription, in the normalizer stage. This is on by default in phonemize. Pass expand_numbers=False to leave digits as-is.

from mwl_phonemizer import MirandesePhonemizer
from mwl_phonemizer.number_utils import normalize_numbers, MirandeseNumberParser

pho = MirandesePhonemizer("mwl")
pho.phonemize("tengo 2 gatos")                 # 2 -> 'dous', then transcribed
pho.phonemize("tengo 2 gatos", expand_numbers=False)   # digit left untouched

# spell numbers to text without phonemizing
normalize_numbers("tengo 21 anhos")            # 'tengo binte i un anhos'
normalize_numbers("la casa 5ª")                # 'la casa quinta'  (º masc / ª fem)

# the verbaliser directly
p = MirandeseNumberParser("mwl")
p.cardinal(256)                # 'duzentos i cincoenta i seis'
p.cardinal(2, "feminine")      # 'dues'
p.ordinal(1, "feminine")       # 'purmeira'
p.pronounce_token("3,5")       # 'trés bírgula cinco'
MirandeseNumberParser("mwl-x-sendim").cardinal(7)   # 'site'  (central 'siête')

Numeral groups join with the copulative i ("and"). The number words come from a source-cited table. Cardinals through 500 and the tens 50-90 are attested in Leite de Vasconcelos, Estudos de Philologia Mirandesa vol. I §189 (pp. 347-351). The hundreds 600-900 follow the periphrastic cardinal + -cientos rule Vasconcelos states for that range (p. 349). The decimal word bírgula and the 6th/10th ordinals come from the regular v→b / final-vowel adaptation. number_utils.ATTESTED and number_utils.DERIVED list which words fall in each group.

How it works

The transcription is the orthography2ipa Mirandese pronunciation lattice. That engine owns the phonology: grapheme rules, allophony, cross-word sandhi, and stress, for all three lects. This library is a thin Mirandese-facing wrapper that adds dialect selection, punctuation-preserving text handling, and two opt-in layers:

  • Lexicon overlay (lookup=True) is a bundled native-speaker word dictionary (mwl_phonemizer.gold, from the TigreGotico/mirandese_g2p dataset). Words present in it are returned verbatim. Its transcription convention is finer-grained (marking, for example, vowel centralization) and differs from the sentence gold below, so it is off by default.

    pho.phonemize("lhéngua")               # 'ˈʎɛŋɡwa'   (lattice)
    pho.phonemize("lhéngua", lookup=True)  # 'ˈʎɛ̃ɡwɐ'   (dictionary)
    
  • CRF correction (use_crf=True) is a linear-chain CRF over the engine's per-grapheme feature export, trained on that same word dictionary. It is tuned to the dictionary's convention and moves output away from the sentence gold, so it too is off by default. It is kept for callers whose target matches that convention.

    MirandesePhonemizer("mwl", use_crf=True).phonemize_word("amportante")
    

Accuracy

Phoneme Error Rate (PER = character edit distance / gold length), gold lookup disabled so the numbers reflect the model.

Human gold: the only accuracy measurement (primary)

The only human-authored Mirandese gold is the 219-word native-speaker dictionary TigreGotico/mirandese_g2p (central 206, sendinese 11, raiano 2, rows routed to mwl / mwl-x-sendim / mwl-x-ifanes by their dialect tag). Everything below is scored against it, full dataset, no caps. Three normalizations are reported:

  • strict: only structural markers (syllable dots, optional-phoneme parentheses) are removed. Stress and every diacritic count.
  • folded: also folds three documented notation conventions: stress marks (ˈ ˌ), length (ː), and tie-bars (t͡ʃ).
  • broad: also folds the documented broad-vs-narrow gap. The human gold is a narrow transcription, and this engine is broad-phonemic. This folds centralized ʉ ʊu, spirant ðd, dark ɫl, lowered e, and the apical/laminal sibilant diacritics (s̺ s̻s). What remains is the residual true phonemic error, not convention distance.
system strict folded broad
pure lattice (deployed default) 20.68% 18.01% 11.62%
+ CRF, 5-fold cross-validated (honest OOD) 21.18% 18.16% 12.73%
+ CRF, fit to dictionary (circular upper bound) 8.79% 3.83% 2.59%
lexicon lookup (lookup=True, memorisation) 0.25% 0.28% 0.30%

Reading the table honestly:

  • Lexicon lookup ≈ 0% is pure memorization: the lexicon is this gold, so every word is returned verbatim. It is not an accuracy signal, and it mixes the narrow lexicon convention into otherwise-broad sentences, which is why it is off by default.
  • CRF fit-to-dictionary (3.83% folded) is trained and scored on the same words, a circular upper bound, not accuracy.
  • CRF 5-fold CV (18.16% folded) is the honest out-of-dictionary estimate. It edges the lattice by about 1.3pp folded.

On the convention-neutral broad basis, the gap collapses to 0.37pp (12.73% vs 11.62%). Almost all of the CRF's apparent gain comes from matching the lexicon's narrow convention, not from fixing real errors, and it couples every output to that convention. The pure lattice remains the default: it is convention-neutral, deterministic, untrained, and statistically tied with the CRF on true phonemic error.

Reproduce with:

python -m mwl_phonemizer.evaluate                # dialect mwl
python -m mwl_phonemizer.evaluate mwl-x-sendim

Engine sentence set: a consistency check, NOT an accuracy claim

orthography2ipa also ships 20-sentence sets per lect (mwl/mwl-x-sendim/mwl-x-ifanes). The lattice reproduces them at about 0% PER. Those sentences were authored to match this engine's own output (they are engine-pinned), so that 0% is an internal consistency check, and not a measure of accuracy against human ground truth. Earlier versions of this README, and two downstream dataset cards, presented that 0% as the primary accuracy figure. That was circular. The human-gold table above is the real measurement.

Related projects

License

Apache-2.0

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

mwl_phonemizer-2.1.0a4.tar.gz (33.1 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

mwl_phonemizer-2.1.0a4-py3-none-any.whl (27.4 kB view details)

Uploaded Python 3

File details

Details for the file mwl_phonemizer-2.1.0a4.tar.gz.

File metadata

  • Download URL: mwl_phonemizer-2.1.0a4.tar.gz
  • Upload date:
  • Size: 33.1 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for mwl_phonemizer-2.1.0a4.tar.gz
Algorithm Hash digest
SHA256 2a49ab95db74368bb25f7c3cbb284c296e594ee9497c64af502e9d5b1852f30c
MD5 b5a03c7597297533e52511001c4fdd29
BLAKE2b-256 51343a1797ad069b46fa12f7ec9f83bcd6516353da498b9ea4e4feb26433695e

See more details on using hashes here.

File details

Details for the file mwl_phonemizer-2.1.0a4-py3-none-any.whl.

File metadata

File hashes

Hashes for mwl_phonemizer-2.1.0a4-py3-none-any.whl
Algorithm Hash digest
SHA256 900bf4ca2b1301f30584b89a0cc9255b3d81e776d14eaee315a5cb61f2c1f6fc
MD5 30cddfd0ba62ae7ac3407b27efbc1d51
BLAKE2b-256 305bfa1c02db1f0fa7416efe0f2cccb6e511b5a930780a8c1b9ece7fb1adb489

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page