Skip to main content

Dialect-aware Portuguese (Lusophone) text-to-IPA phonemizer

Project description

tugaphone — dialect-aware Portuguese phonemizer

tugaphone converts Portuguese text to IPA across all five Lusophone dialect groups. It combines a curated phonetic lexicon, part-of-speech tagging for homograph disambiguation, meaning-based heterophone resolution via bifonia, and a scientifically-grounded regional-accent layer.

O gato dorme.
pt-PT → ˈu gˈa·tu ˈdoɾ·mɨ ˈ···
pt-BR → ˈu gˈa·tʊ ˈdoɾ·mɪ ˈ···
pt-AO → ˈu gˈa·tʊ ˈdoɾ·me ˈ···
pt-MZ → ˈu gˈa·tu ˈdoɾ·me ˈ···
pt-TL → ˈu gˈa·tʊ ˈdoɾ·me ˈ···

Install

pip install tugaphone

30-second quick start

from tugaphone import TugaPhonemizer

ph = TugaPhonemizer()
print(ph.phonemize_sentence("O gato dorme.", "pt-PT"))
# ˈu gˈa·tu ˈdoɾ·mɨ ˈ···

TugaPhonemizer() loads the lexicon and POS tagger once; then call phonemize_sentence(text, lang) as many times as you like. Output is a space-separated phoneme string — one token per word — with ˈ marking primary stress and · marking syllable boundaries.


Features

Five dialect inventories

Code Region
pt-PT European Portuguese — heavy vowel reduction, post-alveolar fricatives, uvular /ʁ/
pt-BR Brazilian Portuguese — fuller vowels, /t d/ palatalisation, l-vocalisation
pt-AO Angolan Portuguese — moderate reduction, alveolar trill, Bantu substrate
pt-MZ Mozambican Portuguese — similar to European with regional variation
pt-TL Timorese Portuguese — conservative pronunciation, Tetum substrate
for code in ["pt-PT", "pt-BR", "pt-AO", "pt-MZ", "pt-TL"]:
    print(code, "→", ph.phonemize_sentence("Choveu muito ontem.", code))
# pt-PT → ʃu·ˈvew mˈũj·tu ˈõ·tẽ ˈ···
# pt-BR → ʃo·ˈvew mwˈĩ·tʊ ˈõ·tẽ ˈ···
# pt-AO → ʃo·ˈvew mˈũjn·tʊ ˈõ·tẽ ˈ···
# pt-MZ → ʃu·ˈvew mˈũj·tu ˈõ·tẽ ˈ···
# pt-TL → ʃo·ˈvew mˈuj·tʊ ˈõ·tẽ ˈ···

Homograph disambiguation

Heterophonic homographs are resolved at two levels:

  1. Meaning-based (via bifonia): sede thirst vs HQ, forma mould vs shape.
  2. POS-based: gosto noun /ˈgoʃtu/ vs verb /ˈgɔʃtu/, para preposition vs verb.
print(ph.phonemize_sentence("Eu gosto de música."))   # verb → ˈgɔʃ·tu
print(ph.phonemize_sentence("Tenho bom gosto."))      # noun → ˈgoʃ·tu

Sub-regional accents

RegionalTransforms presets layer phonological rules on top of any dialect. Rules are grounded in published phonology (Cintra 1971; ALEPG):

from tugaphone.regional import PortoDialect, AzoresDialect

# Porto: stressed /o/ → [uo] (rising diphthong)
print(ph.phonemize_sentence("O vinho é muito bom.", "pt-PT", regional_dialect=PortoDialect))
# ˈu bˈi·ɲu ˈɛ mˈũj·tu bˈuõ ˈ···

# Açores: stressed /u/ → [y], l-palatalisation
print(ph.phonemize_sentence("O vinho é muito bom.", "pt-PT", regional_dialect=AzoresDialect))
# ˈy vˈi·ɲu ˈɛ mˈỹj·tu bˈõ ˈ···

Available presets: NorthernDialect, PortoDialect, MinhoDialect, BragaDialect, FamalicaoDialect, FafeDialect, TrasMontanoDialect, CoimbraDialect, AlentejoDialect, AlgarveDialect, MadeiraDialect, AzoresDialect.

Number normalization

Digits are spelled out with gender agreement and long/short scale per dialect:

from tugaphone.number_utils import normalize_numbers

normalize_numbers("vou comprar 1 casa")   # 'vou comprar uma casa'
normalize_numbers("vou adotar 1 cão")    # 'vou adotar um cão'
normalize_numbers("comprei 2 casas")     # 'comprei duas casas'

Syllabification and stress

Syllabification is handled by silabificador, registered as an orthography2ipa syllabifier plugin. Stress detection delegates to orthography2ipa's declarative StressRules.

Rules-only mode

Pass an IRREGULAR_WORDS-emptied dialect inventory to bypass the lexicon and use only grapheme rules — useful for testing rule coverage or synthesising unknown words.

orthography2ipa plugin interface

TugaphoneG2PPlugin implements orthography2ipa's G2PPlugin interface; SilabificadorSyllabifier implements its SyllabifierPlugin interface and is registered at the orthography2ipa.syllabify entry point.

from tugaphone.plugin import TugaphoneG2PPlugin

p = TugaphoneG2PPlugin(lang="pt-BR")
print(p.transcribe("o gato dorme"))   # ˈu gˈa·tʊ ˈdoɾ·mɪ

Sibling libraries

tugaphone is part of the TigreGotico Portuguese NLP stack:

Library Role
tugalex Phonetic lexicon
tugatagger POS tagger
silabificador Syllabifier
bifonia Heterophone sense disambiguation
orthography2ipa G2P plugin base + stress rules

Documentation


License

Apache License 2.0. See LICENSE.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

tugaphone-0.5.0a2.tar.gz (79.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

tugaphone-0.5.0a2-py3-none-any.whl (69.3 kB view details)

Uploaded Python 3

File details

Details for the file tugaphone-0.5.0a2.tar.gz.

File metadata

  • Download URL: tugaphone-0.5.0a2.tar.gz
  • Upload date:
  • Size: 79.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for tugaphone-0.5.0a2.tar.gz
Algorithm Hash digest
SHA256 7659770397aa51fb1f6d0e375f7de7c8c757a2101590057e68f76dbddd5f43a2
MD5 822e78c947aa0a69be88b5b80fe77a36
BLAKE2b-256 562a9c7ca24ca7b8203826a4a0ab3e6ccfc98a742809dfab181595433af08e15

See more details on using hashes here.

File details

Details for the file tugaphone-0.5.0a2-py3-none-any.whl.

File metadata

  • Download URL: tugaphone-0.5.0a2-py3-none-any.whl
  • Upload date:
  • Size: 69.3 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for tugaphone-0.5.0a2-py3-none-any.whl
Algorithm Hash digest
SHA256 5c186970cdd5866ae464070107f1e01995ab76b49f9803f507178d3b17435708
MD5 a316e43e6c6157e6e7d9b41feeeafea5
BLAKE2b-256 a3a2df658e6f1c46de958495412ec67141d2dffb7aff7829c68a93dfd1512772

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page