No project description provided
Project description
kuro2sudachi
kuro2sudachi lets you to convert kuromoji user dict to sudachi user dict.
Usage
$ pip install kuro2sudachi
# prepase riwirte.def
# https://github.com/WorksApplications/Sudachi/blob/develop/src/main/resources/rewrite.def
$ ls
rewiite.def
$ kuro2sudachi kuromoji_dict.txt -o sudachi_user_dict.txt
Develop
test kuro2sudachi
$ poetry install
$ poetry run pytest
exec kuro2sudachi command
$ poetry run kuro2sudachi tests/kuromoji_dict_test.txt -o sudachi_user_dict.txt
Custom pos convert dict
you can overwrite convert setting with setting json file.
{
"固有名詞": {
"sudachi_pos": "名詞,固有名詞,一般,*,*,*",
"left_id": 4786,
"right_id": 4786,
"cost": 5000
},
"名詞": {
"sudachi_pos": "名詞,普通名詞,一般,*,*,*",
"left_id": 5146,
"right_id": 5146,
"cost": 5000
}
}
$ kuro2sudachi kuromoji_dict.txt -o sudachi_user_dict.txt -s convert_setting.json
if you want to ignore unsupported pos error & invalid format, use --ignore
flag.
Splitting Words
Currently, the CLI does not support word splitting. Therefore, the split representation of kuromoji is ignored.
中咽頭ガン,中咽頭 ガン,チュウイントウ ガン,カスタム名詞
↓
中咽頭ガン,4786,4786,7000,中咽頭ガン,名詞,固有名詞,一般,*,*,*,チュウイントウガン,中咽頭ガン,*,*,*,*,*
TODO
- split mode
- change connection cost
- supports many pos
- supports custom dict converts pos
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
kuro2sudachi-0.2.4.tar.gz
(4.0 kB
view hashes)
Built Distribution
Close
Hashes for kuro2sudachi-0.2.4-py3-none-any.whl
Algorithm | Hash digest | |
---|---|---|
SHA256 | ec1f31902fdcd4a49517eaee295f5b283926d21a32f37fd93040a2f6dab15e38 |
|
MD5 | fcc9c11f9c1b2c42446872e816624a5f |
|
BLAKE2b-256 | 31cfbc45bd8f7c7f3760658aac1eadbada2c954474defd3d01965944b3768d93 |