This release is a pre-release and may not be stable for production use.
TugaLex
TugaLex is a lexicon handler and linguistic utility for Portuguese dialects. It provides data on phoneme transcription (IPA), syllable segmentation, and historical and modern orthographic rules.
It gives NLP pipelines and linguistics research a single API for the differences between Portuguese regions and historical spelling agreements.
Supported Dialects
TugaLex maps standard ISO codes to internal regional datasets:
| ISO Code | Internal Code | Region |
|---|---|---|
pt-PT |
lbx |
Portugal (Lisbon) |
pt-BR |
rjx |
Brazil (Rio de Janeiro) |
| n/a | spx |
Brazil (São Paulo) |
pt-AO |
lda |
Angola (Luanda) |
pt-MZ |
mpx |
Mozambique (Maputo) |
pt-TL |
dli |
Timor-Leste (Díli) |
The regional dictionary ships as a slim gzip extract (word, POS, phones,
syllables, region) of the unified Portuguese pronunciation lexicon.
Regenerate it with python scripts/build_regional_dict.py. Words whose
pronunciation does not vary by part of speech carry a single POS-invariant
entry that answers every POS query.
Key Features
1. Phonemes & Syllables
Retrieve IPA transcriptions and syllable breaks based on the word's Part-of-Speech (POS) and specific region.
from tugalex import TugaLexicon
lex = TugaLexicon()
# Get both syllables and phonemes
info = lex.get("acordo", pos="NOUN", region="lbx")
# Output: {'syllables': ['a', 'cor', 'do'], 'phonemes': 'ɐˈkoɾdu'}
# POS matters!
verb_phonemes = lex.get_phonemes("acordo", pos="VERB")
# Output: 'ɐˈkɔɾdu'
verb_phonemes = lex.get_phonemes("acordo", pos="NOUN")
# Output: 'ɐˈkoɾdu'
2. Orthographic Agreement (AO1990)
Convert text between pre-agreement and post-agreement (AO1990) spelling for Portugal and Brazil.
- Normalize: Old spelling to modern spelling.
- Reverse: Modern spelling to old regional spelling (PT or BR).
normalized = lex.normalize_ao1900(sentence)
3. Linguistic Insights
TugaLex identifies specific linguistic phenomena programmatically:
- Homographs: Words that change pronunciation based on POS (e.g., sede (thirst) vs sede (headquarters)).
- Archaic Words: Mapping 19th-century etymological spellings to modern ones.
- Silent Letters: Identifying words with silent 'p' or 'c' common before the 1990 agreement.
- Voiced 'u': Detecting words where 'u' is pronounced in 'gue/gui/que/qui' clusters (formerly marked with a trema
ü).
Datasets
TugaLex ships the following datasets. Together they contain over 100,000 entries sourced from the Portal da Língua Portuguesa.
regional_dict.csv: Phoneme and syllable mappings.heterophonic_homographs.csv: words pronounced differently depending on postag.acordo_ortografico_pt_PT.csv: Portugal old orthographic spellings.acordo_ortografico_pt_BR.csv: Brazil old orthographic spellings.archaisms.csv: normalized words from before the 20th century.
Install
pip install -e .
TugaLex has no third-party runtime dependencies. All datasets ship inside the package under tugalex/data/.
Related Projects
TugaLex is one piece of the TigreGotico Portuguese NLP stack:
- tugaphone: phonemizes arbitrary Portuguese text across dialects, using TugaLex plus a rule-based fallback.
- tugamorph: rule-based morphological analyzer for Portuguese.
- tugatagger: Part-of-Speech tagging wrapper for Portuguese.
- desacordo_ortografico: detects and converts between Portuguese orthographies, including the AO1990 reform.
License
TugaLex is released under the Apache License 2.0.
Metadata
Release files for tugalex 2.0.2a1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| tugalex-2.0.2a1.tar.gz | 2.6 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| tugalex-2.0.2a1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 5.2 MB
Release files / tugalex-2.0.2a1.tar.gz
| Download URL | tugalex-2.0.2a1.tar.gz |
|---|---|
| Size | 2.6 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
411c9a00d5133660a47606a3b4c1adf15108b96d1bd179fd3509e92d4b77edbe
|
|
BLAKE2b-256 checksum How to use checksums |
71cd0f393a633f0ddc8b5daf4c81373778b311e8a8deb28b2687fd240b25c8d8
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Release files / tugalex-2.0.2a1-py3-none-any.whl
| Download URL | tugalex-2.0.2a1-py3-none-any.whl |
|---|---|
| Size | 2.6 MB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
32b3eaae505b2035ad763fc52eecfe4afbd80b794f6f6a936bd49422818de460
|
|
BLAKE2b-256 checksum How to use checksums |
b73f906a688d282f5ea2bae1d5dabba5e96b4f37fc0fa395f6053d818896fb9a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|