abbr2words
Multilingual, context-aware abbreviation expansion for text normalization and speech.
This standalone package was extracted from the abbreviation framework and language
registries in kokorog2p. It has no runtime dependencies and uses a flat package
layout (no src/ directory).
Supported languages
- Czech (
cs) - German (
de) - English (
en) - Spanish (
es) - French (
fr) - Italian (
it) - Dutch (
nl) - Polish (
pl) - Portuguese (
pt) - Russian (
ru) - Swedish (
sv) - Turkish (
tr)
Locale forms such as de-DE, en_GB, and pt-BR are accepted and currently map
to their base-language registry.
Dutch, Polish, Russian, Swedish, and Turkish currently provide conservative abbreviation and reviewed numeric-unit registries. Their multilingual examples are abbreviation-only; full speech-number normalization remains limited to the optional examples configured for the original scenario languages.
Installation
python -m pip install abbr2words
For development:
python -m pip install -e ".[dev]"
python -m build
pytest
API
from abbr2words import abbr2words
text = "Prof. Klein kommt ggf. am Fr."
print(abbr2words(text, lang="de"))
# Professor Klein kommt gegebenenfalls am Freitag
Context can be disabled:
abbr2words("Fr. Klein", lang="de", context=False)
# Freitag Klein
External linguistic annotations
abbr2words remains dependency-free. Applications that already tokenize and
tag text can pass provider-neutral TokenAnnotation objects with character
offsets and optional POS labels. spaCy is not installed or imported by
abbr2words; see the external POS annotation guide
and examples/spacy_pos.py.
Bundled registries do not currently require POS labels; annotations are used by
custom entries configured with POS guards. The provider-specific tag value is
retained as metadata but is not currently evaluated.
Use an isolated mutable registry for project-specific entries:
from abbr2words import Expander
expander = Expander("de")
expander.add("KI", "Künstliche Intelligenz", case_sensitive=True)
print(expander("KI hilft."))
Consumers that need the shared language registry can use get_shared_expander() and
reset_expanders(). Expander and get_expander() remain isolated mutable registries.
Command line
python -m abbr2words --lang de "Prof. Klein kommt ggf."
printf 'Prof. Klein kommt ggf.' | abbr2words --lang de
Scope
abbr2words expands registered abbreviations and a reviewed set of unit symbols
when they occur after numeric quantities. It preserves numeric values and does
not spell ordinary numbers, dates, or times, and does not perform unit conversion
or currency realization. Unit support is not universal UCUM support. Use the
public iter_unit_matches() API when a downstream semantic normalizer needs the
original numeric lexeme, source span, and stable canonical quantity identity.
abbr2words recognizes and identifies quantity symbols; it does not decide how
a complete numeric quantity is spoken. Number words, grammatical number,
currency major/minor decomposition, and locale-specific spoken decimal policy
belong to the consuming speech normalizer.
Structured currency identities are available in the reviewed quantity registry
for Czech, English, French, Italian, Portuguese, and Spanish. Czech recognizes
Kč/CZK as currency-czech-koruna; Portuguese also recognizes
R$/BRL as currency-brazilian-real; English, French, Italian, and Spanish
recognize the shared currency-euro, currency-us-dollar, and
currency-pound-sterling identities for EUR/USD/GBP. The other listed
languages do not currently expose structured currency identities.
These identities are recognized when a numeric value is adjacent in either
prefix or suffix position:
from abbr2words import iter_unit_matches
match = next(iter_unit_matches("12,80 EUR", "it"))
match.value # "12,80"
match.canonical_id # "currency-euro"
match.canonical_symbol # "€"
The match preserves the written numeric lexeme, symbol, and source-relative offsets. Currency names, number wording, singular/plural agreement, gender, cents, decimal realization, and arithmetic remain the responsibility of the downstream speech normalizer; standalone currency symbols and codes are not lexical rewrites. The reviewed shared inventory is limited to EUR/USD/GBP.
abbr2words("500 g", lang="en")
# "500 gram"
abbr2words("section g", lang="en")
# "section g"
Examples
The repository includes runnable examples for abbreviation-only expansion and
for composing abbr2words with num2words:
python -m pip install "abbr2words[examples]"
python examples/abbreviations.py
python examples/full_text_demo.py --sample german
abbr2words itself expands abbreviations only. The optional examples show how
to combine it with num2words for broader speech-text normalization, including
output such as 500 g -> five hundred grams. The full-text demo is example code,
not part of the stable public API. See
examples/README.md for the complete command reference.
Versioning
The package version is derived from Git tags by setuptools-scm. Use tags in the
form v0.2.3, v0.2.4, and so on; the corresponding package version is generated
automatically during builds. A checkout without tags falls back to 0.2.3 for
this release.
For a release, commit the changes, create an annotated tag, and build from that tag:
git tag -a v0.2.3 -m "Release 0.2.3"
python -m build
License
Apache License 2.0. See LICENSE and NOTICE.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file abbr2words-0.2.4.tar.gz.
File metadata
- Download URL: abbr2words-0.2.4.tar.gz
- Upload date:
- Size: 126.7 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/6.1.0 CPython/3.13.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
059effcc103d0aba0f89f50b7ea22f92766d7f769606793f6cb808036513e32f
|
|
| MD5 |
b8164bf13faf77798e9d799e1cea77bb
|
|
| BLAKE2b-256 |
f946b3ff62e2454d43e4b9de670b7634d4bd2089fcff9070f8521a73d09b4027
|
File details
Details for the file abbr2words-0.2.4-py3-none-any.whl.
File metadata
- Download URL: abbr2words-0.2.4-py3-none-any.whl
- Upload date:
- Size: 58.3 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/6.1.0 CPython/3.13.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
43d83e7da6d663f64ddb3c552194d49eb75c7eaaf75bb2b156ccaa1271520c3f
|
|
| MD5 |
c7bc20403406b967fe6b6053a08fb9a6
|
|
| BLAKE2b-256 |
2d3361642c86907367dc422d62652009344d8da6b6109a4803fbfd1ddae99ed6
|