nbsplitter
A tool for splitting Japanese text into graphemes (i.e. the smallest unit of written text that preserves pronunciation).
Usage
>>> from nbsplitter import split_graphemes
>>> graphemes = split_graphemes("東大和市")
>>> print(graphemes.surface())
['東', '大和', '市']
>>> print(graphemes.reading_form())
['ヒガシ', 'ヤマト', 'シ']
Context
Note the use of the phrase "preserves pronunciation" in the description above; it is impossible to define a "grapheme" without first explaining the very act of splitting text into graphemes to begin with. In summary,
Splitting graphemes is the act of dividing written text into the smallest units possible while ensuring each unit bears a valid pronunciation (in the case of kanji, a valid reading) that reflects its actual pronunciation in the broader string of text. That is to say, by examining an individual unit, it should be evident how exactly that unit is pronounced within the original text.
Technically speaking, this falls outside most definitions of a "grapheme" as it excludes a number of characters that affect pronunciation, but the term most closely coincides with what the intention of this package is.
However, splitting Japanese graphemes as described above isn't exactly a straightforward task. For instance, each individual kanji has several possible readings, certain kanji groupings must be considered unique graphemes because their pronunciations aren't obtainable by merely combining individual readings, and there exist numerous whole-character kana modifiers and dependent characters that don't bear individual pronunciations at all. This package elegantly handles the vast majority of such exceptions under the hood and exposes a single interface for splitting graphemes as desired.
Known Limitations
Inability to differentiate between rendaku and arbitrary unvoiced-to-voiced consonant changes
As noted in the documentation below, the splitter algorithm cannot definitively distinguish between examples of true rendaku and multi-kanji graphemes whose second component appears to be read as the voiced equivalent of one of its standalone readings. For example, consider the output below:
graphemes = split_graphemes("富士", split_rendaku=True)
print(graphemes.reading_form())
# Output: ['フ', 'ジ'] <-- WRONG: should be ['フジ']
This occurs because the kanji 士 can be read as シ, so the algorithm interprets the ジ found in the grapheme's actual reading as an intentional change in voicing (i.e. rendaku) when in practice the grapheme 富士 cannot be split (although this may reveal a thing or two about historical changes in pronunciation, I wouldn't say it's particularly useful for parsing modern Japanese).
If this behavior is undesirable, simply disable the option to split graphemes affected by rendaku, at the cost of compound words such as 船橋 being treated as a single grapheme. May be fixed in a future update.
API Reference
Grapheme
A single grapheme.
Represents the smallest unit of written text that maintains its intended pronunciation. Can either be a single character or a multi-character compound with a distinct pronunciation.
Grapheme.reading_form()
The reading form of this grapheme (in katakana).
Grapheme.surface()
The original Japanese form of this grapheme.
GraphemeList
A list of graphemes.
GraphemeList.reading_form()
A list containing every grapheme's reading form (in katakana).
GraphemeList.surface()
A list containing every grapheme's original Japanese form.
split_graphemes()
Splits Japanese text into graphemes.
Args:
- japanese: The text to be split.
- split_rendaku: NOT RECOMMENDED: Use only if intending on verifying graphemes later on. This option may interpret compounds whose latter parts happen to be the voiced equivalents of unvoiced counterparts as examples of rendaku when they should not be considered as such. If True latter parts of a multi-kanji compound affected by rendaku (see https://en.wikipedia.org/wiki/Rendaku) are treated as separate graphemes.
Returns:
- A GraphemeList representing the split text.
Acknowledgements
This package relies on KANJIDIC dictionary files. These files are property of the Electronic Dictionary Research and Development Group (EDRDG) and are used in accordance with the Group's license.
Release files for nbsplitter 1.0.5
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| nbsplitter-1.0.5.tar.gz | 1.5 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| nbsplitter-1.0.5-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 3.1 MB
Release files / nbsplitter-1.0.5.tar.gz
| Download URL | nbsplitter-1.0.5.tar.gz |
|---|---|
| Size | 1.5 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
afa5278dcc0531f4b8baa0355c80623b40236cc8ac287960cf6f97ec61a17eb4
|
|
BLAKE2b-256 checksum How to use checksums |
26e7ea222326c137ed03369adfe640905893e982f749358762f7ab151068e31f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 15, 2026.
Transparency logRelease files / nbsplitter-1.0.5-py3-none-any.whl
| Download URL | nbsplitter-1.0.5-py3-none-any.whl |
|---|---|
| Size | 1.6 MB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
4e303c36d2bdb472925621011999bf60b0e917fb3b9decd9c8366934432824b0
|
|
BLAKE2b-256 checksum How to use checksums |
041b04452c55dee514a1a20d27af094d155a317800eabbaa515ef527deac134a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 15, 2026.
Transparency log