Skip to main content

Thai Word Segmentation using TCC + Bidirectional RNNs

Project description

NokCut

Thai Word Segmentation using TCC + Bidirectional RNNs

Credit code from A Beginner's Guide to Deep NLP with PyTorch - Dr. Prachya Boonkwan

Colab Notebook : https://colab.research.google.com/drive/1WS08VsjlZGAmCGsoI7AlRm-Do3zo-b-g

Train by BEST I Corpus Training set. (90% training , 10% test)

ep 6
loss: 0.017879242024514966
f1 : 98.47012481095481

F1 From BEST I Corpus Test set

F-measure: 96.94929
Recall: 122271.00000/125850.00000 = 97.15614

Precision: 122271.00000/126387.00000 = 96.74333

Number of incorrect : 3579.00000 words

Mr. Wannaphong Phatthiyaphaibun wannaphong@kkumail.com

Project details


Release history Release notifications | RSS feed

This version

0.4

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Files for nokcut, version 0.4
Filename, size File type Python version Upload date Hashes
Filename, size nokcut-0.4-py3-none-any.whl (5.9 MB) File type Wheel Python version py3 Upload date Hashes View

Supported by

Pingdom Pingdom Monitoring Google Google Object Storage and Download Analytics Sentry Sentry Error logging AWS AWS Cloud computing DataDog DataDog Monitoring Fastly Fastly CDN DigiCert DigiCert EV certificate StatusPage StatusPage Status page