An open-source Python library for Transformer-based part-of-speech tagging of the Uzbek language using a fine-tuned Tahrirchi-BERT model.
Project description
UzbekTaggerBERT
UzbekTaggerBERT is an open-source Python library for Transformer-based part-of-speech tagging of the Uzbek language. It is built on a fine-tuned Tahrirchi-BERT model and is designed to assign Universal POS tags to Uzbek words in context.
- Current software version: v1.0.1
- Previous PyPI release: v0.1.2 (and v1.0.0)
- PyPI package: uzbek-tagger-bert
- GitHub repository: MaksudSharipov/uzbek-tagger-bert
- Hugging Face model: MaksudSharipov/Uzbek-POS-Tagger-TahrirchiBERT
- License: Apache License 2.0
The library addresses key challenges of Uzbek NLP, including agglutinative word formation, contextual homonym disambiguation, apostrophe normalization, and subword-to-word alignment in Transformer-based token classification.
Main Features
- Transformer-based POS tagging for Uzbek texts.
- Fine-tuned Tahrirchi-BERT model for contextual POS prediction.
- Support for 16 Universal POS tags.
- Contextual homonym disambiguation, for example distinguishing yuz as a noun or numeral depending on context.
- Smart apostrophe normalization for Uzbek Latin texts.
- Overlap-based offset mapping for aligning subword tokens with original words.
- Python API for integration into Uzbek NLP pipelines.
- Command-line interface for quick text tagging.
- Public distribution through PyPI and Hugging Face Hub.
Installation
Install the package from PyPI:
pip install uzbek-tagger-bert
The package requires Python 3.7 or higher. The main dependencies are torch, transformers, and numpy.
For development installation from GitHub:
git clone https://github.com/MaksudSharipov/uzbek-tagger-bert.git
cd uzbek-tagger-bert
pip install -e .
Quick Start
from uzbek_tagger_bert import UzbekTaggerBERT
tagger = UzbekTaggerBERT()
text = (
"U yuz burib ketdi va yuz soat ishladi. "
"Alisher sariq olma olib menga qizil olma olmading dedi."
)
result = tagger(text)
print(result)
Expected output:
U/PRON yuz/NOUN burib/VERB ketdi/VERB va/CCONJ yuz/NUM soat/NOUN ishladi/VERB ./PUNCT
Alisher/PROPN sariq/ADJ olma/NOUN olib/VERB menga/PRON qizil/ADJ olma/NOUN olmading/VERB dedi/VERB ./PUNCT
Command-Line Usage
After installation, the package provides the uzbek-tagger command:
uzbek-tagger "U yuz burib ketdi va yuz soat ishladi."
JSON output:
uzbek-tagger "Bugun talabalar yangi mavzuni o‘rgandilar." --format json
CoNLL-like output:
uzbek-tagger "Bozordan olma olib kel." --format conll
POS Tagset
UzbekTaggerBERT supports 16 Universal POS tags:
| Tag | Description | Uzbek explanation |
|---|---|---|
| ADJ | Adjective | Sifat |
| ADP | Adposition | Ko‘makchi |
| ADV | Adverb | Ravish |
| AUX | Auxiliary | Yordamchi fe’l |
| CCONJ | Coordinating conjunction | Teng bog‘lovchi |
| INTJ | Interjection | Undov |
| MOD | Modal word | Modal so‘z |
| NOUN | Noun | Ot |
| NUM | Numeral | Son |
| PART | Particle | Yuklama |
| PRON | Pronoun | Olmosh |
| PROPN | Proper noun | Atoqli ot |
| PUNCT | Punctuation | Tinish belgisi |
| SCONJ | Subordinating conjunction | Ergash bog‘lovchi |
| SYM | Symbol | Simvol |
| VERB | Verb | Fe’l |
More information is available in docs/tagset.md.
Model
The core POS tagging model is available on Hugging Face Hub:
MaksudSharipov/Uzbek-POS-Tagger-TahrirchiBERT
Model page:
https://huggingface.co/MaksudSharipov/Uzbek-POS-Tagger-TahrirchiBERT
Model Performance
The current public version reports the following evaluation results on the UzbekPOS test dataset:
| Metric | Score |
|---|---|
| Accuracy | 0.9810 |
| Weighted F1 | 0.9811 |
Additional benchmark results, per-tag scores, baseline comparisons, and inference-speed tests are stored in the benchmark/ directory.
Methodology
Uzbek is an agglutinative Turkic language in which grammatical and derivational meanings are widely expressed through suffixation. This creates a large number of surface word forms and increases the complexity of token-level linguistic analysis.
UzbekTaggerBERT handles this problem by using a Transformer-based contextual encoder and a token-level POS prediction layer. The package also includes practical preprocessing and alignment mechanisms:
- Text normalization: standardizes common apostrophe variants in Uzbek Latin texts.
- Subword tokenization: uses the Transformer tokenizer associated with the fine-tuned Tahrirchi-BERT model.
- Offset alignment: maps subword-level model predictions back to original word boundaries.
- Contextual POS prediction: assigns POS tags based on surrounding context.
- Punctuation post-processing: punctuation tokens are deterministically tagged as
PUNCT(e.g.,.,,,!,?, etc. are always mapped toPUNCT).
Use Cases
UzbekTaggerBERT can be used in:
- Uzbek corpus annotation;
- morphological analysis pipelines;
- syntactic analysis pipelines;
- educational NLP systems;
- linguistic feature extraction;
- retrieval-augmented generation systems for Uzbek grammar;
- low-resource NLP research;
- downstream Uzbek text processing applications.
Repository Structure
The repository is organized as follows:
uzbek-tagger-bert/
├── README.md
├── LICENSE
├── CITATION.cff
├── pyproject.toml
├── repo/
│ └── src/
│ └── uzbek_tagger_bert/
│ ├── __init__.py
│ ├── tagger.py
│ ├── model_loader.py
│ ├── tokenizer_utils.py
│ ├── label_map.py
│ └── cli.py
├── examples/
│ ├── basic_usage.py
│ ├── batch_tagging.py
│ ├── export_conll.py
│ └── cli_usage.py
├── tests/
│ ├── test_import.py
│ ├── test_label_map.py
│ ├── test_tokenizer_utils.py
│ ├── test_output_format.py
│ └── test_model_integration_optional.py
├── benchmark/
│ ├── evaluate.py
│ ├── speed_test.py
│ ├── results.csv
│ └── README.md
└── docs/
├── installation.md
├── usage.md
└── tagset.md
Documentation
- Installation guide:
docs/installation.md - Usage guide:
docs/usage.md - POS tagset:
docs/tagset.md - Benchmark scripts:
benchmark/README.md
Examples
Run the basic example:
python examples/basic_usage.py
Run batch tagging:
python examples/batch_tagging.py
Export results in CoNLL-like format:
python examples/export_conll.py
Testing
Install development dependencies:
pip install pytest scikit-learn
Run tests:
pytest tests
Optional real-model integration test:
RUN_MODEL_TESTS=1 pytest tests/test_model_integration_optional.py
Benchmarking
Evaluate on a CoNLL-like test file:
python benchmark/evaluate.py --test-file path/to/test.conll --output benchmark/results.csv
Run inference-speed test:
python benchmark/speed_test.py
SoftwareX Article
This repository accompanies the SoftwareX manuscript:
UzbekTaggerBERT: An Open-Source Python Library for Transformer-Based Part-of-Speech Tagging of the Uzbek Language
The repository provides the source code, usage examples, documentation, benchmark scripts, and reproducibility materials required for the software publication.
Citation
If you use this library, model, or dataset in academic research, please cite the repository using the Cite this repository button on GitHub.
Related dataset publication:
Sharipov, M., Kuriyozov, E., & Vičič, J. (2026). UzbekPOS: A multi-domain dataset for Uzbek part-of-speech tagging. Data in Brief, 112640. https://doi.org/10.1016/j.dib.2026.112640
@article{sharipov2026uzbekpos,
title={UzbekPOS: A multi-domain dataset for Uzbek part-of-speech tagging},
author={Sharipov, Maksud and Kuriyozov, Elmurod and Vičič, Jernej},
journal={Data in Brief},
pages={112640},
year={2026},
publisher={Elsevier},
doi={10.1016/j.dib.2026.112640}
}
Author
Maksud Sharipov
- GitHub: https://github.com/MaksudSharipov
- Hugging Face: https://huggingface.co/MaksudSharipov
- PyPI: https://pypi.org/project/uzbek-tagger-bert/
License
This project is licensed under the Apache License 2.0. See the LICENSE file for details.
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file uzbek_tagger_bert-1.0.1.tar.gz.
File metadata
- Download URL: uzbek_tagger_bert-1.0.1.tar.gz
- Upload date:
- Size: 19.3 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.14.4
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
cf77f06f4467d8e322ff318e5780c93e97cc2312b3f6ae359ba4b19e421234e7
|
|
| MD5 |
a60733e479ff92aca3a05695950a9c91
|
|
| BLAKE2b-256 |
2ffa63036e998b18fab87f3d5fd972a66b23f0cbb938b28358f7f07c26296fb9
|
File details
Details for the file uzbek_tagger_bert-1.0.1-py3-none-any.whl.
File metadata
- Download URL: uzbek_tagger_bert-1.0.1-py3-none-any.whl
- Upload date:
- Size: 15.1 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.14.4
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
84791a280b273c1165ca426e86283c8d58a53d5e7edb160c4a1e32a53c179fc9
|
|
| MD5 |
b6294d1363e832b6f83b1f53a74c9de6
|
|
| BLAKE2b-256 |
42af5639c9e391be7c6f3b740e8f03d02932442390680d9d034a2993f806ac36
|