Skip to main content

Text readability calculator for Japanese learners 🇯🇵


jReadability allows python developers to calculate the readability of Japanese text using the model developed by Jae-ho Lee and Yoichiro Hasebe in Introducing a readability evaluation system for Japanese language education and Readability measurement of Japanese texts based on levelled corpora. Note that this is not an official implementation.

Demo

You can play with an interactive demo here.

Installation

pip install jreadability

Quickstart

from jreadability import compute_readability

# "Good morning! The weather is nice today."
text = 'おはようございます!今日は天気がいいですね。' 

score = compute_readability(text)

print(score) # 6.438000000000001

Readability scores

Level Readability score range
Upper-advanced [0.5, 1.5)
Lower-advanced [1.5, 2.5)
Upper-intermediate [2.5, 3.5)
Lower-intermediate [3.5, 4.5)
Upper-elementary [4.5, 5.5)
Lower-elementary [5.5, 6.5)

Note that this readability calculator is specifically for non-native speakers learning to read Japanese. This is not to be confused with something like grade level or other readability scores meant for native speakers.

Model

readability = {mean number of words per sentence} * -0.056
            + {percentage of kango} * -0.126
            + {percentage of wago} * -0.042
            + {percentage of verbs} * -0.145
            + {percentage of particles} * -0.044
            + 11.724

* "kango" (漢語) means Japanese word of Chinese origin while "wago" (和語) means native Japanese word.

Note on model consistency

The readability scores produced by this python package tend to differ slightly from the scores produced on the official jreadability website. This is likely due to the version difference in UniDic between these two implementations as this package uses UniDic 2.1.2 while theirs uses UniDic 2.2.0. This issue may be resolved in the future.

Batch processing

jreadability makes use of fugashi's tagger under the hood and initializes a new tagger everytime compute_readability is invoked. If you are processing a large number of texts, it is recommended to initialize the tagger first on your own, then pass it as an argument to each subsequent compute_readability call.

from fugashi import Tagger

texts = [...]

tagger = Tagger()

for text in texts:
    
    score = compute_readability(text, tagger) # fast :D
    #score = compute_readability(text) # slow :'(
    ...

Documentation

You can find this repo's (very minimal) documentation here.

Other implementations

The official jReadability implementation can be found on jreadability.net

A node.js implementation can also be found here.

Metadata

Release files for jreadability 1.1.5

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for jreadability 1.1.5
File Size Uploaded
jreadability-1.1.5.tar.gz 12.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for jreadability 1.1.5
File Interpreter ABI Platform
jreadability-1.1.5-py3-none-any.whl Python 3 none any Details

Total release size: 18.4 kB

Release files / jreadability-1.1.5.tar.gz

Download URL jreadability-1.1.5.tar.gz
Size 12.1 kB
Tags Source
SHA-256 checksum
How to use checksums
173341dd9a9b6dac16b901757629762237733b7f1823824f92d20b26bd1a40e5
BLAKE2b-256 checksum
How to use checksums
6312f42bc3b9b72bbd1263879fed20a927041a028aecbfa4c7b11b9993eebf5a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.9.24

Release files / jreadability-1.1.5-py3-none-any.whl

Download URL jreadability-1.1.5-py3-none-any.whl
Size 6.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
56d0500282c88697de878dc6f1d11619510b7e6d7e4f5a8cb1372988e35e5a1d
BLAKE2b-256 checksum
How to use checksums
48a2d431f3218557b43970a65036e849d33ce1011245b1c276a1b5ad3bce059a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.9.24

Release history Release notifications | RSS feed

This release

1.1.5 This release

2 release files

1.1.4

2 release files

1.1.3

2 release files

1.1.2

2 release files

1.1.1

2 release files

1.1.0

2 release files

1.0.1

2 release files

1.0.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page