Calculate readability scores for Japanese texts.
Project description
jreadability allows python developers to calculate the readability of Japanese text using the model developed by Jae-ho Lee and Yoichiro Hasebe in "Readability measurement of Japanese texts based on levelled corpora." Note that this is not an official implementation.
Installation
pip install jreadability
Quickstart
from jreadability import compute_readability
# "Good morning! The weather is nice today."
text = 'おはようございます!今日は天気がいいですね。'
score = compute_readability(text)
print(score) # 5.596333333333334
Readability scores
Level | Readability score range |
---|---|
Upper-advanced | 0.5-1.4 |
Lower-advanced | 1.5 - 2.4 |
Upper-intermediate | 2.5 - 3.4 |
Lower-intermediate | 3.5 - 4.4 |
Upper-elementary | 4.5 - 5.4 |
Lower-elementary | 5.5 - 6.4 |
Note that this readability calculator is specifically for non-native speakers learning to read Japanese. This is not to be confused with something like grade level or other readability scores meant for native speakers.
Equation
$$ \begin{align*} \textrm{readability}~= &\ (\textrm{mean number of words per sentence}) \cdot -0.056 \ &\ + (\textrm{proportion of kango}) \cdot -0.126 \ &\ + (\textrm{proportion of wago}) \cdot -0.042 \ &\ + (\textrm{proportion of verbs}) \cdot -0.145 \ &\ + (\textrm{proportion of auxiliary verbs}) \cdot -0.044 \ &\ + 11.724 \end{align*} $$
* "kango" (漢語) means Japanese word of Chinese origin while "wago" (和語) means native Japanese word.
Note on model consistency
The readability scores produced by this python package tend to differ slightly from the scores produced on the official jreadability website. This is likely due to the version difference in UniDic between these two implementations as this package uses UniDic 2.1.2 while theirs uses UniDic 2.2.0. This issue will hopefully be resolved in the future.
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Hashes for jreadability-1.0.0-py3-none-any.whl
Algorithm | Hash digest | |
---|---|---|
SHA256 | 5f395096b4205eb9aaf5bba66a76fc728d68497004da67edefffe7e98fa9e7ef |
|
MD5 | 52471662974937e1ced90db6ff5965b6 |
|
BLAKE2b-256 | 8f890a05c2f88b09f813276836f5078ec0a4ca5f73604ea93eb4654071b81e39 |