Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

bifonia

Pronunciation disambiguation for European-Portuguese heterophonic homographs — words spelled identically whose pronunciation depends on meaning, not just part of speech. Zero runtime dependencies, pure standard library.

sede is thirst (ˈsedɨ, closed e) or a headquarters (ˈsɛdɨ, open e); forma is a mould (ˈfoɾmɐ) or a shape (ˈfɔɾmɐ); molho is sauce (ˈmoʎu) or a bundle (ˈmɔʎu). A text-to-speech front-end that guesses wrong says the wrong word out loud. bifonia picks the right reading — and therefore the right IPA — from context.

from bifonia import tokenize, guess_sense, disambiguate

words = tokenize("Tinha tanta sede que bebi a garrafa toda.")
i = words.index("sede")
guess_sense(words, i)    # 'thirst'
disambiguate(words, i)   # 'ˈsedɨ'   (closed e)

words = tokenize("A sede da empresa fica em Lisboa.")
i = words.index("sede")
disambiguate(words, i)   # 'ˈsɛdɨ'   (open e — same spelling, different word)

Why this is hard (and why POS-tagging isn't enough)

The obvious approach — tag the part of speech and pick the pronunciation from it — cannot work when two readings share a POS. sede thirst and seat are both nouns; forma mould and shape are both nouns; corte cut and court are both nominal. A part-of-speech tagger labels them identically and is wrong on the minority reading by construction.

It is easy to look like you've solved this and not have. Encyclopedic text (Wikipedia) is noun-skewed — one reading dominates and the minority readings barely occur — so a POS-tagger that just predicts the majority sense scores well. The honest test is a sense-balanced set where the minority readings actually appear.

bifonia ships that test: bifonia-pt-homographs-gold — 2,346 real sentences (OpenSubtitles, Wikipedia, web), ~40% of them on POS-ambiguous readings, labelled independently of the engine. To our knowledge it is the first real-text heterophonic-homograph benchmark for European Portuguese.

Accuracy on the balanced real benchmark

approach all POS-ambiguous (same-POS)
most-common (majority sense) 48.2% 72.0%
cue-only (curated wordlists, no POS, no learning) 53.2% 87.3%
Yarowsky decision list (trained) 88.6% 86.5%
spaCy pt_core_news_lg (POS → sense) 85.8% 74.0%
bifonia rules (zero-dependency) 95.8% 97.6%

Accuracy by approach on the balanced real benchmark

The headline: a strong neural POS-tagger drops to chance-like 74% on the POS-ambiguous half — it has no signal to separate two senses that share a part of speech — while a zero-dependency, meaning-keyed rule engine holds 97.6%. On the POS-separable half a tagger does fine (~93%); on the half that needs meaning, only meaning works.

Rules vs spaCy by subset

A second finding worth its own line: curated word-lists alone reach 87.3% on the hard cases, before any POS reasoning or learning — meaning lives in collocations.

Two engines, one ensemble

engine needs a corpus? how it decides
rules no context rules over part-of-speech + meaning cues in .voc wordlists
learned yes per-word Naive-Bayes / averaged perceptron over context features

The rule engine is self-contained and needs no training data — the right fit for a low-resource language. The learned models are trained from the labelled corpus and serve as a full-roster baseline and a corpus-QC engine. guess_sense is a per-word ensemble: each word is served by whichever engine wins on held-out behavioural data, so the combined system never does worse than the rules.

A consistent result across this project: models trained on the corpus do not beat the rules out-of-distribution. Because the corpus is rule-labelled, a model can at best mimic the rules in-distribution and generalises worse on real text — true for Naive-Bayes, the averaged perceptron, and the classic Yarowsky decision list alike. The rules remain the predictor; an external POS tag can be fused on top:

# hybrid: trust an external tagger only where POS uniquely picks a sense, else rules
guess_sense(words, i, postag="NOUN")

On the encyclopedic wild set this hybrid ensemble reaches 95.3% (rules 90.7%, spaCy 91.6%) — the tagger helps on POS-separable words, the rules carry the rest.

Install

pip install bifonia          # zero dependencies, pure standard library

API

from bifonia import (tokenize, is_ambiguous, guess_sense, guess_pos,
                     disambiguate, add_extra_diacritics)

sentence = "Resolveu o problema desta forma simples."
words = tokenize(sentence)
i = words.index("forma")

guess_sense(words, i)                    # 'shape'
disambiguate(words, i)                   # 'ˈfɔɾmɐ'
disambiguate(words, i, sense="mould")    # 'ˈfoɾmɐ'  (override)
add_extra_diacritics(sentence)           # '...desta fórma simples.'  (acute = open vowel)

add_extra_diacritics rewrites each homograph with a disambiguating diacritic (acute → open vowel, circumflex → closed) that a downstream grapheme-to-phoneme stage can read directly.

Datasets

All on the Hugging Face Hub, schema {word, sense, pos, ipa, sentence}:

Architecture

Layered and dependency-free, so a cue behaves identically in the rules and in the model features:

bifonia/text.py      strip / accent-fold / proximity-weighted cue scoring   (shared, stdlib)
bifonia/cues.py      SENSE_CUES + FEATURE_CUES — the cue registry            (single source of truth)
bifonia/scoring.py   context rules + one generic cue resolver
bifonia/features.py  language-agnostic features, incl. CUE:<sense> from the registry
bifonia/locale/<lang>/*.voc   editable context + meaning wordlists
scripts/baselines.py         a Baseline protocol + zero-dep baselines for benchmarking

The meaning cues that disambiguate same-spelling readings live in one declarative registry (bifonia/cues.py) read by both the rule resolver and the model feature extractor. Adding a homograph is data-only: drop the .voc wordlist(s) and add one registry line — no new code. Porting to a related language means supplying .voc files and (optionally) a corpus; the algorithm carries no hardcoded Portuguese.

Coverage

131 homographs, 265 sense readings. The disambiguation axis is vowel quality (open ɔ/ɛ vs closed o/e). The hard core is the same-POS pairs a tagger cannot touch: sede, corte, forma, molho, bola, cor, lobo, polo, tola. Per-word IPA, senses, and diacritized forms are in docs/word_references.md.

See also

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

bifonia-0.2.0a3.tar.gz (5.5 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

bifonia-0.2.0a3-py3-none-any.whl (5.6 MB view details)

Uploaded Python 3

File details

Details for the file bifonia-0.2.0a3.tar.gz.

File metadata

  • Download URL: bifonia-0.2.0a3.tar.gz
  • Upload date:
  • Size: 5.5 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for bifonia-0.2.0a3.tar.gz
Algorithm Hash digest
SHA256 b352339bd98a21bb16da8657c2ae354516cbd8ccda8ee574e10c4eca1ee1e0c5
MD5 18c955d2afd358d0056baec75dfc9b84
BLAKE2b-256 4266ecfa8b72ecf795f838706708193745ecf2077e49464afcd685a563e41f38

See more details on using hashes here.

File details

Details for the file bifonia-0.2.0a3-py3-none-any.whl.

File metadata

  • Download URL: bifonia-0.2.0a3-py3-none-any.whl
  • Upload date:
  • Size: 5.6 MB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for bifonia-0.2.0a3-py3-none-any.whl
Algorithm Hash digest
SHA256 e2fea58de8b3d418528a52f00c6f0e82d6375a0ff954f3fa31a9898a8a8c5c7f
MD5 14580ffe0fb692e29cb257d196c34b0d
BLAKE2b-256 859cbec32fdf6037bb32095f95d26cbd9beb8ee82fa590a4c945dfacdc6d321f

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page