Skip to main content

belarusian-verse

License PyPI CI Dataset Data licence Contributors

Belarusian stress, rhyme, rhythm, agreement and spelling — the checks you need to write or judge a line of Belarusian verse, as a small Python library.

The tables behind it — about 2 million word forms — are published separately as 🤗 YauhenBichel/belarusian-verse and downloaded on first use.

from belarusian_verse import mark_stress, rhyme, suggest_rhymes, check_agreement

mark_stress("Побач ты, і добра мне")      # 'По́бач ты, і до́бра мне'
rhyme("вадзе", "ідзе")                    # ('rich', 1.0)
rhyme("ідзе", "знайдзе")                  # ('none', 0.0)  — зна́йдзе is stressed at the start
check_agreement("тваіх вачам")["errors"]  # 1  — no shared case
check_agreement("Яна ідуць")["errors"]    # 1  — singular subject, plural verb
suggest_rhymes("вадзе")[:4]               # what *would* rhyme, commonest first

Why

Belarusian has good dictionaries and good speech models, and little in between:

  • no rhyme dictionary in software — only Minkin's printed one;
  • LanguageTool's Belarusian module has ~66 rules and almost no morphology, so it cannot see that «тваіх вачам» is ungrammatical;
  • a spell checker accepts every real word, so «ідзе́ / зна́йдзе» passes as a rhyme even though the two words are stressed on different syllables.

Everything needed was already inside the Belarusian Grammar Database: stress marks and full morphological tags for millions of forms. This library is those tables plus the rules that make them answer musical questions.

Install

pip install belarusian-verse            # tables download on first use, then cached
pip install "belarusian-verse[spelling]"  # adds the spell checker (spylls)

Set BELARUSIAN_VERSE_DATA to a folder holding the tables to work offline. Put BelVoice's stresses-stat.json there too for homograph stress; without it homographs keep GrammarDB's order.

What it does

Function Question it answers
mark_stress(text) Where does the stress fall? (по́бач, not паба́ч)
rhyme(a, b) Do these rhyme, and how well: rich / exact / near / none
analyse_verse(text, "AABB") Rhyme and rhythm scores for a whole verse, line by line
check_agreement(text) Does each adjective match its noun in gender, case and number?
check_spelling(text, dictionary=…, forms=…) Real Belarusian words, correct orthography, syllable counts
respell(text, "uk"|"ru") Rewrite Belarusian so a Ukrainian- or Russian-trained model pronounces it, syllable count unchanged
suggest_rhymes(word) Words that rhyme with it, commonest first — rhyme as help, not just a verdict

Each is also a command: be-stress, be-poetry, be-grammar, be-spelling, be-respell, be-rhymes.

$ be-poetry verse.txt --rhyme AABB
 1 100111001    rhythm 1.00  Белы туман па лесе плыве,
 2 101000101    rhythm 0.44  Свежасць ранішняя тут жыве.
rhyme 1-2: rich (1.0)
rhyme 0.95  rhythm 0.78

What it checks that a spell checker cannot

  • Agreement. An adjective must match its noun in gender, case and number, on either side of it («цёплы вечар», «свежасць ранішняя»); a verb must match its subject in person and number, or in gender for the past tense («яна спявала», «яны спявалі»).
  • The у/ў rule between words. у becomes ў after a vowel — «Я іду ў школу» — and stays у after a consonant. The decision lives in the gap between two words, so no word-by-word checker can see it.
  • Rhyme and rhythm, from the stressed vowel, not from the final letters.
  • Spelling, with the dictionary's gaps marked as gaps. The bundled Hunspell dictionary is GrammarDB's own 2008-spelling export, which leaves out forms whose only source is Піскуноў 2012 («слухаўка») and adjectives from place names («смаленская»). Pass forms=load_index() and a word the dictionary rejects but GrammarDB lists is a warning (NOT_IN_2008_LIST) instead of an UNKNOWN_WORD error. By default forms is not loaded and the old rule holds; extra-words.txt still accepts its words outright.

How rhyme is judged

A Belarusian rhyme matches from the last stressed vowel to the end of the line, so the stress table does the real work. маёй and спакой rhyme (iotated vowels are folded: ёй = ой); ідзе́ and зна́йдзе do not, because the stress sits elsewhere. Final consonants are devoiced before comparison, as they are when sung.

A scheme shorter than the verse repeats for each stanza: ABCB over eight lines pairs 2–4 and 6–8, and AABB pairs 1–2, 3–4, 5–6, 7–8.

Homograph stress

GrammarDB lists the stresses of a homograph in no particular order. Where one reading must be chosen (pick_first=True, and rhyme and rhythm in analyse_verse), the library puts first the most frequent stress from BelVoice's table of common homographs, when that stress is one GrammarDB lists. On the Belarusian Homographs Stress Benchmark (python -m belarusian_verse.stress_benchmark), share of homograph tokens stressed right:

Set GrammarDB order BelVoice first
Common Voice sentences 48.9 % 82.9 %
«Засценак Малінаўка» (literary) 24.9 % 68.8 %
10×10 (every stress equally often) 50.0 % 50.0 %

10×10 is built so that frequency alone cannot win; only context could. If BelVoice's table cannot be downloaded, homographs silently keep GrammarDB's order. load_lexicon(homographs="") asks for that order explicitly.

Two Belarusian rules this encodes

  • ё is always stressed, so a word containing it needs no lookup.
  • A word starting with ў is the same word as the one starting with у (ўдваіх = удваіх), but ў is not a vowel, so every stress index shifts by one. Getting this wrong moves the stress a syllable — it is the kind of bug that only shows up when a singer sings it.

Limits

  • Homographs are not disambiguated by context: 20,754 forms carry more than one stress, and mark_stress leaves them unmarked. Pass your own overrides file, or pick_first=True to accept the most frequent reading — wrong about one time in six in ordinary sentences, and about one in three in literary text.
  • Agreement checking looks at adjacent adjective/pronoun and noun pairs on either side, not at full syntax; it finds the common errors, not every possible one.
  • Meaning is not checked at all. A line can pass every check here and still say nothing.
  • Preposition government (да wants the genitive, у the accusative or locative) is not checked.
  • Rhyme suggestions are ranked by Wikipedia frequency, which leans encyclopaedic: place names are commoner there than in song.

Data and licence

The code is Apache-2.0. The tables it downloads are CC BY-SA 4.0, derived from the Belarusian Grammar Database by Aleś Bułojčyk and Uładzimir Koščanka — if you redistribute the data or a derivative of it, keep that licence and the attribution: https://huggingface.co/datasets/YauhenBichel/belarusian-verse

Two more sources are downloaded at run time and never shipped with the package:

  • BelVoice's stresses-stat.json (most frequent stress of common homographs), by Aleś Bułojčyk and contributors, LGPL-3.0-or-later, fetched from a pinned commit and checked against its SHA-256.
  • The Belarusian Homographs Stress Benchmark, used only by stress_benchmark, created for BelVoice by Aleś Bułojčyk and contributors, CC BY-SA 4.0: https://huggingface.co/datasets/alex73/benchmarks-stress-bel (revision 94bb5a8).

Tests

pip install -e ".[spelling]" && python -m unittest discover tests

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

belarusian_verse-0.3.0.tar.gz (38.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

belarusian_verse-0.3.0-py3-none-any.whl (34.4 kB view details)

Uploaded Python 3

File details

Details for the file belarusian_verse-0.3.0.tar.gz.

File metadata

  • Download URL: belarusian_verse-0.3.0.tar.gz
  • Upload date:
  • Size: 38.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for belarusian_verse-0.3.0.tar.gz
Algorithm Hash digest
SHA256 311e754185665eb6b45f907398652bc17619e9e8f5d5a5be728da3d8a181040a
MD5 ea38d8be2eee16ca06afe26a2af49e32
BLAKE2b-256 47b72507651fd15adbba337357185a07c7e130cb9ae2d3dd3fe50c85fca048e4

See more details on using hashes here.

Provenance

The following attestation bundles were made for belarusian_verse-0.3.0.tar.gz:

Publisher: release.yml on YauhenBichel/belarusian-verse

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file belarusian_verse-0.3.0-py3-none-any.whl.

File metadata

File hashes

Hashes for belarusian_verse-0.3.0-py3-none-any.whl
Algorithm Hash digest
SHA256 a39e07a72ec3b87c867daa454632ad2bd29f2aeb214310cb665398c1131cb180
MD5 16b38842820f79577adf47033415fb65
BLAKE2b-256 44830c139213c6de9a116da8a0bf71c00ba5b8ae73d1dc8c362ff1cd1f48155d

See more details on using hashes here.

Provenance

The following attestation bundles were made for belarusian_verse-0.3.0-py3-none-any.whl:

Publisher: release.yml on YauhenBichel/belarusian-verse

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

This release

0.3.0 This release

2 files

0.2.0

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page