Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

TugaLex 🗣️

TugaLex is a comprehensive lexicon handler and linguistic utility for Portuguese dialects. It provides deep data on phoneme transcription (IPA), syllable segmentation, and historical/modern orthographic rules.

Designed for NLP pipelines and linguistics research, it allows you to handle the complexities of Portuguese across different regions and historical agreements with a unified API.


🌍 Supported Dialects

TugaLex maps standard ISO codes to internal regional datasets:

ISO Code Internal Code Region
pt-PT lbx Portugal
pt-BR rjx Brazil
pt-AO lda Angola
pt-MZ mpx Mozambique
pt-TL dli Timor-Leste

✨ Key Features

1. Phonemes & Syllables

Retrieve IPA transcriptions and syllable breaks based on the word's Part-of-Speech (POS) and specific region.

from tugalex import TugaLexicon

lex = TugaLexicon()

# Get both syllables and phonemes
info = lex.get("acordo", pos="NOUN", region="lbx")
# Output: {'syllables': ['a', 'cor', 'do'], 'phonemes': 'ɐˈkoɾdu'}

# POS matters!
verb_phonemes = lex.get_phonemes("acordo", pos="VERB")
# Output: 'ɐˈkɔɾdu'
verb_phonemes = lex.get_phonemes("acordo", pos="NOUN")
# Output: 'ɐˈkoɾdu'

2. Orthographic Agreement (AO1990)

Effortlessly convert text between pre-agreement and post-agreement (AO1990) standards for both Portugal and Brazil.

  • Normalize: Old spelling → Modern spelling.
  • Reverse: Modern spelling → Old regional spelling (PT or BR).
normalized = lex.normalize_ao1900(sentence)

3. Linguistic Insights

TugaLex identifies specific linguistic phenomena programmatically:

  • Homographs: Words that change pronunciation based on POS (e.g., sede (thirst) vs sede (headquarters)).
  • Archaic Words: Mapping 19th-century etymological spellings to modern ones.
  • Silent Letters: Identifying words with silent 'p' or 'c' common before the 1990 agreement.
  • Voiced 'u': Detecting words where 'u' is pronounced in 'gue/gui/que/qui' clusters (formerly marked with a trema ü).

📂 Datasets

TugaLex ships the following datasets which contain over 100,000 entries sourced from the Portal da Língua Portuguesa.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

tugalex-0.3.0a1.tar.gz (6.4 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

tugalex-0.3.0a1-py3-none-any.whl (6.4 MB view details)

Uploaded Python 3

File details

Details for the file tugalex-0.3.0a1.tar.gz.

File metadata

  • Download URL: tugalex-0.3.0a1.tar.gz
  • Upload date:
  • Size: 6.4 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for tugalex-0.3.0a1.tar.gz
Algorithm Hash digest
SHA256 263d43b268dfa8f9d68a0f54d4e7e7f2e3ce39d6ce4bd8f0388454cc3a8dc8ca
MD5 a57754384cd7aed24fb67371f3b6efc8
BLAKE2b-256 de272e19cf662de89d47b05768fa471881bf7021305660c26de82b4ea64bd7b2

See more details on using hashes here.

File details

Details for the file tugalex-0.3.0a1-py3-none-any.whl.

File metadata

  • Download URL: tugalex-0.3.0a1-py3-none-any.whl
  • Upload date:
  • Size: 6.4 MB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for tugalex-0.3.0a1-py3-none-any.whl
Algorithm Hash digest
SHA256 c871d6aa4be253bf23e603abe81c8ffb6c152bf48fd4cbd0f2ed222c7ed64bc2
MD5 ac58a7c59dc913f73c778a99dbdeefbb
BLAKE2b-256 f05ff667be21a408536bbaa3349826b9905871e1a539ad5ae2b4197f5a0be50a

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page