Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

TugaLex 🗣️

TugaLex is a comprehensive lexicon handler and linguistic utility for Portuguese dialects. It provides deep data on phoneme transcription (IPA), syllable segmentation, and historical/modern orthographic rules.

Designed for NLP pipelines and linguistics research, it allows you to handle the complexities of Portuguese across different regions and historical agreements with a unified API.


🌍 Supported Dialects

TugaLex maps standard ISO codes to internal regional datasets:

ISO Code Internal Code Region
pt-PT lbx Portugal
pt-BR rjx Brazil
pt-AO lda Angola
pt-MZ mpx Mozambique
pt-TL dli Timor-Leste

✨ Key Features

1. Phonemes & Syllables

Retrieve IPA transcriptions and syllable breaks based on the word's Part-of-Speech (POS) and specific region.

from tugalex import TugaLexicon

lex = TugaLexicon()

# Get both syllables and phonemes
info = lex.get("acordo", pos="NOUN", region="lbx")
# Output: {'syllables': ['a', 'cor', 'do'], 'phonemes': 'ɐˈkoɾdu'}

# POS matters!
verb_phonemes = lex.get_phonemes("acordo", pos="VERB")
# Output: 'ɐˈkɔɾdu'
verb_phonemes = lex.get_phonemes("acordo", pos="NOUN")
# Output: 'ɐˈkoɾdu'

2. Orthographic Agreement (AO1990)

Effortlessly convert text between pre-agreement and post-agreement (AO1990) standards for both Portugal and Brazil.

  • Normalize: Old spelling → Modern spelling.
  • Reverse: Modern spelling → Old regional spelling (PT or BR).
normalized = lex.normalize_ao1900(sentence)

3. Linguistic Insights

TugaLex identifies specific linguistic phenomena programmatically:

  • Homographs: Words that change pronunciation based on POS (e.g., sede (thirst) vs sede (headquarters)).
  • Archaic Words: Mapping 19th-century etymological spellings to modern ones.
  • Silent Letters: Identifying words with silent 'p' or 'c' common before the 1990 agreement.
  • Voiced 'u': Detecting words where 'u' is pronounced in 'gue/gui/que/qui' clusters (formerly marked with a trema ü).

📂 Datasets

TugaLex ships the following datasets which contain over 100,000 entries sourced from the Portal da Língua Portuguesa.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

tugalex-0.2.0a1.tar.gz (6.4 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

tugalex-0.2.0a1-py3-none-any.whl (6.4 MB view details)

Uploaded Python 3

File details

Details for the file tugalex-0.2.0a1.tar.gz.

File metadata

  • Download URL: tugalex-0.2.0a1.tar.gz
  • Upload date:
  • Size: 6.4 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for tugalex-0.2.0a1.tar.gz
Algorithm Hash digest
SHA256 2fa79c9a14a468ea869eec54d983afddbcdb686609f36b24d7dfd627d32299fd
MD5 61a1b6e9e5d97446d590e9ddd1ab5d20
BLAKE2b-256 32b5ccea318de7271b3077bed97f2e39d3af1c019c90cf3117f74a36296326d1

See more details on using hashes here.

File details

Details for the file tugalex-0.2.0a1-py3-none-any.whl.

File metadata

  • Download URL: tugalex-0.2.0a1-py3-none-any.whl
  • Upload date:
  • Size: 6.4 MB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for tugalex-0.2.0a1-py3-none-any.whl
Algorithm Hash digest
SHA256 fce3cff7b143691f91b2976fb7e5a86d329c4bfa97e7fb99e9c93494c121b18b
MD5 478f20ce3a9e1ff281967686fc318e0e
BLAKE2b-256 0db37e5718fc5812b39060b6e07004dc55510ae7916f8e56403918bda128e184

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page