Skip to main content

PyKoTokenizer

PyKoTokenizer is a Korean text tokenizer for Korean Natural Language Processing tasks. It includes deep learning (RNN) model-based word tokenizers as well as morphological analyzer based word tokenizers for Korean language.

Segmentation of Korean Words

Written Korean texts do employ white space characters. However, more often than not, Korean words occur in a text concatenated immediately to adjacent words without an intervening space character. This low degree of separation of words from each other in writing is due somewhat to an abundance of what linguists call "endoclitics" in the language.

As the language has been subjected to principled and rigorous study for a few decades, the issue of which strings of sounds, or letters, are words and which are not, has been settled among a small group of selected linguists. This kind of advancement has not been propagated to the general public yet, and nlp engineers working on Korean cannot but make do with whatever inconsistent grammars they happen to have access to. Thus, a major source of difficulty in developing competent Korean text processors has been, and still is, the notion of a word as the smallest syntactic unit.

How to install

Before using this package please make sure you have the following dependencies installed in your system.

  • Python >= 3.6
  • numpy >= 1.19.0
  • pandas >= 1.1.5
  • tensorflow >= 2.6.2
  • h5py >= 3.1.0
  • konlpy >= 0.5.2

Use the following command to install the package:

pip install pykotokenizer

How to Use

Model-based Tokenizers

Below, we show examples of using model-based tokenizers.

Using KoTokenizer

from pykotokenizer import KoTokenizer

tokenizer = KoTokenizer()

korean_text = "김형호영화시장분석가는'1987'의네이버영화정보네티즌10점평에서언급된단어들을지난해12월27일부터올해1월10일까지통계프로그램R과KoNLP패키지로텍스트마이닝하여분석했다."

tokenizer(korean_text)

Output:

"김 형호 영화 시장 분석가 는 ' 1987 ' 의 네이버 영화 정보 네티즌 10 점평 에서 언급 된 단어 들 을 지난 해 12 월 27 일 부터 올해 1 월 10 일 까지 통계 프로그램 R 과 KoNLP 패키지 로 텍스트 마이닝 하여 분석 했다 ."

Using KoSpacing

from pykotokenizer import KoSpacing

spacing = KoSpacing()

korean_text = "김형호영화시장분석가는'1987'의네이버영화정보네티즌10점평에서언급된단어들을지난해12월27일부터올해1월10일까지통계프로그램R과KoNLP패키지로텍스트마이닝하여분석했다."

spacing(korean_text)

Output:

"김형호 영화시장 분석가는 '1987'의 네이버 영화 정보 네티즌 10점 평에서 언급된 단어들을 지난해 12월 27일부터 올해 1월 10일까지 통계 프로그램 R과 KoNLP 패키지로 텍스트마이닝하여 분석했다."

Morphological analyzer based Tokenizers

Below, we show examples of using morphological analyzer based tokenizers. These tokenizers has dependency on KoNLPy. So, please install KoNLPy before using these. To install KoNLPy please visit this link - https://konlpy.org/en/latest/install/ and follow the procedure. KoNLPy requires Java in your system.

Using KoKkma

from pykotokenizer import KoKkma

kokkma = KoKkma()

korean_text = "김형호영화시장분석가는'1987'의네이버영화정보네티즌10점평에서언급된단어들을지난해12월27일부터올해1월10일까지통계프로그램R과KoNLP패키지로텍스트마이닝하여분석했다."

kokkma(korean_text)

Output:

"김 형 호 영화 시장 분석가 는 ' 1987 ' 의 네이버 영화 정보 네티즌 10 점 평 에서 언급 되 ㄴ 단어 들 을 지난해 12 월 27 일 부터 올해 1 월 10 일 까지 통계 프로그램 R 과 KoNLP 패키지 로 텍스트 마이닝 하 여 분석 하 었 다 ."

Using KoKomoran

from pykotokenizer import KoKomoran

kokomoran = KoKomoran()

korean_text = "김형호영화시장분석가는'1987'의네이버영화정보네티즌10점평에서언급된단어들을지난해12월27일부터올해1월10일까지통계프로그램R과KoNLP패키지로텍스트마이닝하여분석했다."

kokomoran(korean_text)

Output:

"김형호 영화 시장 분석가 는 ' 1987 ' 의 네이버 영화 정보 네티즌 10 점 평 에서 언급 되 ㄴ 단어 들 을 지난해 12월 27 일 부터 올해 1월 10 일 까지 통계 프로그램 R 과 KoNLP 패키지 로 텍스트 마 이닝 하 아 분석 하 았 다 ."

Credits

This package is a revamped and customized version of the following two sources:

Release files for pykotokenizer 0.0.3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for pykotokenizer 0.0.3
File Size Uploaded
pykotokenizer-0.0.3.tar.gz 11.3 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for pykotokenizer 0.0.3
File Interpreter ABI Platform
pykotokenizer-0.0.3-py3-none-any.whl Python 3 none any Details

Total release size: 22.6 MB

Release files / pykotokenizer-0.0.3.tar.gz

Download URL pykotokenizer-0.0.3.tar.gz
Size 11.3 MB
Tags Source
SHA-256 checksum
How to use checksums
da787f7be36c50b459b6735a3fc76a63bbb68f093ddbf25b534bce0ec5efddd1
BLAKE2b-256 checksum
How to use checksums
3b2e121126022e2f5f857609dcb5963290ccbb3844f6fc6c083a071a9c3ed305
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/3.7.1 importlib_metadata/4.8.3 pkginfo/1.8.2 requests/2.23.0 requests-toolbelt/0.9.1 tqdm/4.62.3 CPython/3.6.9

Release files / pykotokenizer-0.0.3-py3-none-any.whl

Download URL pykotokenizer-0.0.3-py3-none-any.whl
Size 11.3 MB
Tags Python 3
SHA-256 checksum
How to use checksums
4e20e6d0afe168530102b005973f0e78a4385abd4c9d6bc9a4713baadd0a7dc1
BLAKE2b-256 checksum
How to use checksums
3b51c94b1251fe8786644242a5b766724620778c6d0e9d6355095676e6468a00
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/3.7.1 importlib_metadata/4.8.3 pkginfo/1.8.2 requests/2.23.0 requests-toolbelt/0.9.1 tqdm/4.62.3 CPython/3.6.9

Release history Release notifications | RSS feed

This release

0.0.3 This release

2 release files

0.0.2

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page