alifbe
Robust text normalization for real-world Uzbek text — the eight problems that quietly corrupt Uzbek data in most pipelines.
| Problem | What alifbe does |
|---|---|
8+ different "apostrophe" characters (' ‘ ’ ` ´ ...) all meaning oʻ / gʻ / tutuq belgisi |
Normalizes all of them to the two correct Unicode letters, based on context |
"Isʼhoq" → "Işoq" |
Detects s+h morpheme boundaries (via apostrophe and a known-word list) so sh isn't wrongly merged into one letter |
ş vs ș (cedilla vs. comma-below — different code points, identical glyph) |
Detects and normalizes confusable characters so search/dedup actually works |
The "Turkish I" bug (.upper()/.lower() under some locales turns i into İ) |
Locale-independent, explicit-table casing that never produces İ/ı |
Cyrillic е, ц, ё are position/origin-dependent |
Transliteration always returns warnings for ambiguous letters instead of silently guessing (and can raise instead, if you'd rather fail loudly) |
| The Sept 2026 alphabet reform (sh→ş, ch→ç, oʻ→ö, gʻ→ğ) | to_new_latin() / to_old_latin() convert between the two orthographies, with brand-name/URL/code protection |
| "Which script is this text even in?" | detect_alphabet() — cyrillic / old-latin / new-latin / mixed / unknown |
"oʻzbek", "özbek", and "ўзбек" are the same word, but == doesn't think so |
fold_search_key() gives every spelling the same canonical key |
Install
Not yet published to PyPI — the name alifbe is free, checked via
pip download alifbe returning no match, but publishing itself is a
manual step (PyPI account + 2FA + twine upload). Install locally for now:
pip install /path/to/alifbe # normal install
pip install -e /path/to/alifbe # editable, for development
This also registers an alifbe command-line tool (see below).
Quick tour
import alifbe as uz
# 1. Apostrophes: every variant collapses to the correct, same string
uz.normalize_apostrophes("o'zbek") # -> "oʻzbek" (U+02BB, turned comma)
uz.normalize_apostrophes("o‘zbek") # -> "oʻzbek" (same result)
uz.normalize_apostrophes("san'at") # -> "sanʼat" (U+02BC — different rule, not after o/g)
# 2. "Isʼhoq" stays "Isʼhoq" -- sh is not wrongly merged into ш
uz.latin_to_cyrillic("Isʼhoq").text # -> "Исъҳоқ" (not "Ишоқ")
uz.latin_to_cyrillic("Ishoq").warnings # non-empty: flags this as a known
# s+h boundary word even without
# the apostrophe
# 3. ş vs ș
uz.find_confusables("Kraiova munșasi") # -> [Confusable(char='ș', codepoint='U+0219', ...)]
uz.normalize_confusables("munşa") == uz.normalize_confusables("munșa") # -> True
# 4. Turkish I bug
uz.uz_upper("olib") # -> "OLIB" (never "OLİB")
uz.find_turkish_i_corruption("OLİB") # -> flags the İ as corruption
uz.fix_turkish_i_corruption("OLİB") # -> "OLIB"
# 5. Ambiguous Cyrillic letters warn instead of silently guessing
result = uz.cyrillic_to_latin("центр")
result.text # -> "tsentr" (best guess)
result.warnings # -> non-empty, explains ц is ambiguous
uz.cyrillic_to_latin("ёлғон", on_ambiguous="raise") # -> raises instead
# 6. The Sept 2026 alphabet reform
uz.to_new_latin("Shahzoda Oʻzbekistonda choy ichdi").text
# -> "Şahzoda Özbekistonda çoy içdi"
uz.to_old_latin("Şahzoda Özbekistonda çoy içdi").text
# -> "Shahzoda Oʻzbekistonda choy ichdi"
# ... with brand names / URLs / code protected from conversion
uz.to_new_latin("MyShop: sotib oling", protected_terms=["MyShop"]).text
# -> "MyShop: sotib oling" (MyShop untouched, rest still converts if applicable)
uz.to_new_latin("See https://x.com/shahar for info").text
# -> "See https://x.com/shahar for info" (URL untouched)
# 7. What script is this?
uz.detect_alphabet("Shahzoda") # -> AlphabetDetection(alphabet='old-latin', confidence=0.65)
uz.detect_alphabet("Şahzoda") # -> AlphabetDetection(alphabet='new-latin', confidence=0.7)
uz.detect_alphabet("Шаҳзода") # -> AlphabetDetection(alphabet='cyrillic', confidence=1.0)
# 8. One search key regardless of script or apostrophe style
uz.fold_search_key("o'zbek") == uz.fold_search_key("özbek") == uz.fold_search_key("ўзбек")
# -> True (all fold to "özbek")
Command line
echo "o'zbek" | alifbe normalize-apostrophes # oʻzbek
alifbe to-new-latin "Shahzoda choy ichdi" # Şahzoda çoy içdi
alifbe to-cyrillic "Ishoq" # Исъҳоқ (+ warning on stderr)
alifbe detect "Шаҳзода" # cyrillic (confidence: 1.0)
alifbe fold-key "oʻzbek" # özbek
alifbe check "OLİB" # corruption: 'İ' (U+0130) at index 2
Every subcommand reads from the positional argument if given, or stdin otherwise — so it pipes cleanly. Conversion warnings go to stderr, so stdout stays clean for piping the result onward.
Design principle
Every function that could plausibly get something wrong either:
- makes the correct choice deterministically (apostrophes, casing, alphabet-reform digraphs), or
- tells you it's not sure, instead of guessing silently (Cyrillic е/ц/ё, the sh/s+h boundary when no apostrophe survives).
That second category is deliberate: a library that always looks
confident is more dangerous than one that sometimes says "I'm not sure,
here's my best guess and why." All conversion functions return the same
Result(text, warnings) shape — warnings is empty when the library is
confident, and populated (with a machine-readable rule plus a
human-readable message) when it made a judgment call.
On the Sept 2026 alphabet reform
Uzbekistan's Senate approved a bill on 10 September 2026 replacing the
sh/ch/oʻ/gʻ digraphs with single letters ş/ç/ö/ğ. As of this writing the
bill has been sent to the president and is not yet in force — school
materials are expected to transition starting 2027. to_new_latin() is a
forward-looking convenience, not a claim about which spelling is
currently mandatory. Check lex.uz or the Ministry of Education for the
authoritative status before treating conversion as required.
Known simplifications (read before relying on this for anything critical)
SH_BOUNDARY_EXCEPTIONS(inexceptions_data.py) is a small starting list, not a linguistic corpus — extend it viaextra=for your data, and get a native speaker to review it before production use.- Cyrillic
цhas no reliable positional rule (can be "s" or "ts" depending on the word's origin); alifbe defaults to "ts" and always warns. Cyrillicеuses a positional heuristic (ye at word start/after a vowel, else e) which is usually right but isn't a certainty for borrowed words. latin_to_cyrillic("c")maps to "к" (since old-orthography Uzbek Latin has no standalone "c" outside the "ch" digraph) — for text with loanwords spelled with a bare "c", double-check the output.fold_search_key_loose()is intentionally lossy (strips ö/ğ/ş/ç diacritics and apostrophes) — never use it as a unique key, only for fuzzy "did you mean" style matching.
Extending the sh-boundary word list
uz.latin_to_cyrillic("Asʼhad", extra_sh_exceptions={"ashad"})
Tests
pip install -e ".[dev]"
pytest tests/ -v
102 tests across normalization, transliteration, the alphabet-reform converter, script detection, search-key folding, and the CLI.
Release files for alifbe 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| alifbe-0.2.0.tar.gz | 29.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| alifbe-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 55.3 kB
Release files / alifbe-0.2.0.tar.gz
| Download URL | alifbe-0.2.0.tar.gz |
|---|---|
| Size | 29.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
0d3c2aa9fcbcf7e5cc2cf03ecf1fbfb1ddf4419b150b3f5353f808e2f5edc25f
|
|
BLAKE2b-256 checksum How to use checksums |
8acf3bbed4a722f7002183dbfb2f1b5b283a7745d301324dbf7c6c55abb7668e
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.3
|
Release files / alifbe-0.2.0-py3-none-any.whl
| Download URL | alifbe-0.2.0-py3-none-any.whl |
|---|---|
| Size | 25.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
293d740c268b36a23942d9e8085092c35ab5a25e6002a4a2fbb24619615b4414
|
|
BLAKE2b-256 checksum How to use checksums |
efe85a53187ef0abdc545b7d65386e0c6a36482335ede40d583da5823f318a13
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.3
|