Skip to main content

tokenize-uk

Ukrainian tokenization: paragraphs → sentences → words. Small, fast and robust, now with LanguageTool-grade engines.

Since 2.0 the default engines are ports of what LanguageTool uses for Ukrainian, both verified byte-identical to their Java originals on multi-million-token real-world corpora:

  • words: a port of LanguageTool's UkrainianWordTokenizer (Andriy Rysin's abbreviation-aware tokenizer) — verified on 3.87M tokens across four corpora, zero differences, faster than the Java original;
  • sentences: choppa-srx, the Python port of the Java segment SRX library with LanguageTool's rules.

The original 2016 regex engines are preserved and one argument away.

Quick Start

pip install tokenize_uk
import tokenize_uk

tokenize_uk.tokenize_words("Це проф. Артюхов.")
# ['Це', 'проф.', 'Артюхов', '.']

tokenize_uk.tokenize_sents("Це проф. Артюхов. Він приїхав у м. Київ.")
# ['Це проф. Артюхов.', 'Він приїхав у м. Київ.']

tokenize_uk.tokenize_text("Перший абзац. Ще речення.\nДругий абзац.")
# [[['Перший', 'абзац', '.'], ['Ще', 'речення', '.']], [['Другий', 'абзац', '.']]]

Command line:

echo "Це проф. Артюхов. Він приїхав у м. Київ." | tokenize-uk -l sents

The two engines

Every function keeps the pre-2.0 signature and return shape. If the new behavior breaks your pipeline, fall back per call:

tokenize_uk.tokenize_words(text, legacy=True)   # 2016 regex engine

or import the old engines directly from tokenize_uk.legacy. The legacy code is preserved byte-for-byte (regexes included) and covered by the original test fixtures.

Differences you will notice with the default engine:

  • abbreviations keep their dots and don't split (проф., т., зв., чл.-кор.), and don't end sentences where they shouldn't;
  • dates (12.03.2022), times (15:30), decimals (10,5) and web entities (Цензор.НЕТ) stay single tokens;
  • sentence splitting follows LanguageTool's segment.srx rules instead of a punctuation-plus-uppercase heuristic.

For raw LanguageTool-compatible word tokens (including whitespace tokens), use the class directly:

from tokenize_uk import UkrainianWordTokenizer
UkrainianWordTokenizer().tokenize("а б")
# ['а', ' ', 'б']

Using with spaCy

pip install tokenize_uk[spacy]
from tokenize_uk.spacy import blank_pipeline

nlp = blank_pipeline()
doc = nlp("Це проф. Артюхов. Він приїхав у м. Київ.")
[t.text for t in doc]              # LT-grade word tokens
[sent.text for sent in doc.sents]  # LT-grade sentences

doc.text round-trips the input exactly, token offsets are exact, and token text is taken verbatim from the input (unlike tokenize_words, which canonicalizes a few apostrophe/quote characters). nlp.to_disk()/spacy.load() work (the tokenizer is a registered spaCy factory). You can also drop UkrainianTokenizer or the tokenize_uk_sentencizer component into an existing pipeline — with the caveat that pretrained statistical components were trained on spaCy's own tokenization.

Verification and performance

The word tokenizer is compared byte-for-byte against LanguageTool master's UkrainianWordTokenizer (compiled from source; harness in scripts/java-harness/):

corpus lines tokens Java tokenize-uk 2.0 output
Militarny news 100,004 1,736,362 7.4 s 5.4 s identical
uanews.dp.ua 133,120 1,973,203 8.9 s 7.8 s identical
Liga.net 5,982 69,695 0.5 s 0.2 s identical

The sentence layer's own verification (byte-identity with the Java segment library on ~136k segments) is documented in choppa's README.

Reproduce with scripts/benchmark.py (see docs/design.md).

Documentation

  • API — functions, shapes, engines, CLI
  • Design — how the LT port works and how it's verified
  • History — 2016 origins, the LanguageTool ports

Copyrights and kudos

  • Vsevolod Dyomkin, Dmytro Chaplynskyi — original library, lang-uk project
  • Andriy Rysin and the LanguageTool team — the original UkrainianWordTokenizer and the Ukrainian SRX rules
  • Jarek Lipski — the segment library behind the sentence layer

MIT licensed.

Release files for tokenize-uk 2.0.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for tokenize-uk 2.0.0
File Size Uploaded
tokenize_uk-2.0.0.tar.gz 188.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for tokenize-uk 2.0.0
File Interpreter ABI Platform
tokenize_uk-2.0.0-py3-none-any.whl Python 3 none any Details

Total release size: 207.7 kB

Release files / tokenize_uk-2.0.0.tar.gz

Download URL tokenize_uk-2.0.0.tar.gz
Size 188.1 kB
Tags Source
SHA-256 checksum
How to use checksums
0577442efcb44838f5b0ff199892a1c2189f9dc429c4dde26aacffce8141388d
BLAKE2b-256 checksum
How to use checksums
1ea01d60d87a93824a245e1823de320c410691c43e24e22747d0632f5e1e4565
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 5, 2026.

Transparency log

Release files / tokenize_uk-2.0.0-py3-none-any.whl

Download URL tokenize_uk-2.0.0-py3-none-any.whl
Size 19.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
c8d5b7ca34b42ca670dfdaca38824d41d0780592f1d11d27e88c67dd9395246d
BLAKE2b-256 checksum
How to use checksums
5e21300022a22f06f838e9c8da48be6ed4099c8fa514a9112985c62ae7d49e78
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 5, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

2.0.0 This release

2 release files

0.2.0

1 release file

0.1.5

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page