Skip to main content

ko-hand-ocr

Reads one line of handwritten Korean + English. A 41M model trained from scratch on synthetic data — no handwriting dataset, no inherited terms.

PyPI License Python Params Weights Hugging Face CPU Data

한국어 · English

photo  ->  line cutting  ->  ko-hand-ocr  ->  "부서 : 포테토뭉부서"

ko-hand-ocr scoreboard

The top table puts us next to six other public OCR models on the same test sheets, on the same PC at the same moment (tools/vs.py — who was run and how is below). Below it are our two models item by item, from tools/verify.py at about 960 lines per font. The single model reads with one 41M model. The ensemble runs three models and picks the answer they agree on — more accurate, seven times slower.


At a glance

What it reads One line of handwritten Korean + English, from a photo
Model ViT-Small encoder (ImageNet, Apache-2.0) + 8-layer TrOCR decoder trained from scratch
Vocabulary 229 jamo tokens — not 11,172 syllables, so unseen characters are still writable
Parameters 41M — one fifth of ko-trocr (213.7M)
Weights 156 MB single model · 450 MB ensemble (3 models) (downloaded separately — GitHub Releases · Hugging Face)
Package 119 KB — the code only
Plugs into a server that speaks the OpenAI and Ollama APIs (ko-hand-ocr-serve) · an MCP tool for LM Studio (ko-hand-ocr-mcp)
Speed 0.18 s per line · 1.04 s for a whole photo — CPU only, no GPU
Memory 1.07 GB single model · 1.51 GB ensemble
Accuracy 95.6% mean on six handwriting fonts never seen in training (95.6% ensemble) · worst font 90.4% (90.7% ensemble) · 11 handwriting photos 94.9% (96.6% ensemble) · 11 of 17 items over 95% (ensemble 12)
Against others On the same sheets, ahead of the best of six public OCR models by +17.7 points on the six fonts (ko-trocr) and +15.7 points on the photos (ko-trocr) — single model
Training data Synthetic — drawn on the fly from OFL fonts. Nothing is stored on disk
License Apache-2.0, weights included — no dataset terms inherited
Python 3.10+ · PyTorch 2.5+

What it does — and does not

Does

  • Reads a whole photo: finds the lines, cuts them, reads each one (read_photo)
  • Reads pre-cut line images, batched (read)
  • Constrained decoding — restrict the output to a candidate list, and report whether the free and constrained readings agree (both)
  • Ensemble reading — run several models and take the answer they agree on
  • Runs fully offline on CPU. Nothing is sent anywhere
  • Plugs into LM Studio, Ollama and vLLM tooling — a server speaking their API, and an MCP tool ("Use it with LM Studio, Ollama and vLLM" below)

Does not

  • Split a table cell into label and value — read_photo only cuts down to lines
  • Handle vertical writing or merged cells
  • Stay silent on an empty cell — a cell that is pure scribble still produces something

Why it exists

The only public model aimed at Korean handwriting that we could find is ddobokki/ko-trocr, and its training data comes from AI Hub, which places conditions on purpose of use and on redistribution. That blocks it from being embedded in an in-house tool or shipped as a public package. General OCR engines (PaddleOCR, EasyOCR, Tesseract) and large VLMs fall well short on a line of handwriting — on the same sheets the best of them reads 77.6% of the six fonts and 78.5% of the photos (below). So a model for this job was built from scratch — with a provenance chain that can be audited part by part.

Measured against other OCR on the same ruler

There are six rivals — one from each kind of thing people actually reach for to read Korean.

Rival Kind Size How it was run
ddobokki/ko-trocr Korean TrOCR (trained on AI Hub) 214M transformers, float32, beam 5, length cap raised 16 -> 64
PaddleOCR PP-OCRv5 line recogniser (korean_PP-OCRv5_mobile_rec) 3.3M paddleocr 3.7.0 / paddle 3.4.0, TextRecognition
EasyOCR line recogniser (korean_g2) 4.0M easyocr 1.7.2, Reader(["ko", "en"]).recognize()
Tesseract 5 line recogniser (LSTM, no parameter count published) 6MB file tesseract 5.5.3, --oem 1 --psm 7 -l kor+eng
PaddleOCR-VL-1.6 OCR-specific VLM 906M transformers 5.16.1, bfloat16, prompt "OCR:" (model card)
Qwen3-VL-2B-Instruct general VLM 2.1B transformers 5.16.1, bfloat16, Korean prompt — the best of three we tried

Every rival runs exactly as its model card says, each in its own environment (tools/rival_worker.py), and only the time the engine itself reports is counted (handing images across is not).

A comparison only holds if the model is the only thing that differs. Photos were cut once with our line cutter (kohandocr.page) and the same cells were handed to everyone. Font test sheets were handed over raw, exactly as synth.render drew them — putting our preprocessing (64×640 letterbox) on them would squash the others twice. Metric, GPU and the moment of measurement are all shared. What was measured and how is spelled out in the header of tools/vs.py, and every number below comes from the runs/VS.json it leaves behind.

Accuracy

Real handwriting photos · 34 lines, 209 characters

Jamo similarity Character error rate Exact line match Worst photo
ko-hand-ocr ensemble (118M) 96.52% 8.61% 67.6% 90.28%
ko-hand-ocr single model (41M) 94.24% 13.40% 61.8% 83.03%
ddobokki/ko-trocr (214M) 78.50% 39.23% 26.5% 58.56%
PaddleOCR PP-OCRv5 (3.3M) 71.75% 39.71% 8.8% 34.81%
Qwen3-VL-2B (2.1B) 69.86% 38.28% 8.8% 46.86%
PaddleOCR-VL-1.6 (906M) 61.71% 44.02% 2.9% 18.52%
EasyOCR (4.0M) 66.88% 55.02% 2.9% 46.19%
Tesseract 5 (6MB) 13.21% 98.56% 0.0% 0.00%

Six unseen handwriting fonts · 670 lines, 4708 characters

Jamo similarity Character error rate Exact line match Worst font
ko-hand-ocr ensemble (118M) 95.65% 8.28% 73.0% 89.83%
ko-hand-ocr single model (41M) 95.36% 8.45% 72.2% 89.02%
ddobokki/ko-trocr (214M) 77.62% 41.67% 29.4% 65.24%
PaddleOCR PP-OCRv5 (3.3M) 72.28% 40.40% 28.4% 52.43%
Qwen3-VL-2B (2.1B) 66.02% 79.89% 19.9% 50.02%
PaddleOCR-VL-1.6 (906M) 62.97% 62.34% 18.1% 43.05%
EasyOCR (4.0M) 56.35% 64.21% 9.7% 36.95%
Tesseract 5 (6MB) 17.42% 98.13% 0.9% 7.29%

Breadth — 24 more fonts · 648 lines, 3768 characters

Jamo similarity Character error rate Exact line match Worst font
ko-hand-ocr ensemble (118M) 97.71% 3.42% 85.5% 92.46%
ko-hand-ocr single model (41M) 97.20% 3.98% 84.7% 89.14%
ddobokki/ko-trocr (214M) 79.94% 32.38% 42.0% 46.52%
PaddleOCR PP-OCRv5 (3.3M) 75.27% 30.02% 39.5% 18.66%
Qwen3-VL-2B (2.1B) 67.81% 62.98% 26.1% 24.46%
PaddleOCR-VL-1.6 (906M) 66.69% 45.94% 24.8% 27.33%
EasyOCR (4.0M) 64.38% 48.91% 19.1% 19.28%
Tesseract 5 (6MB) 19.98% 95.12% 2.2% 3.07%

All 480 fonts

All 480 fonts · 12 lines each, 4320 lines, 22560 characters

held out (6)
6
breadth (24)
24
seen in training
450
All
480
fonts ≥ 95% fonts < 80%
ko-hand-ocr ensemble (118M) 99.15% 99.00% 98.22% 98.27% 436 5
ko-hand-ocr single model (41M) 98.49% 98.58% 98.06% 98.09% 434 4
ddobokki/ko-trocr (214M) 85.29% 83.29% 82.48% 82.56% 46 158
PaddleOCR PP-OCRv5 (3.3M) 80.07% 78.48% 78.02% 78.07% 86 198
Qwen3-VL-2B (2.1B) 65.17% 67.35% 69.32% 69.17% 50 289
PaddleOCR-VL-1.6 (906M) — — — — — —
EasyOCR (4.0M) 67.73% 68.17% 67.15% 67.21% 9 319
Tesseract 5 (6MB) 21.56% 19.28% 20.50% 20.45% 0 480
  • Even the best of the six rivals is far behind. On the six unseen fonts the best rival is ko-trocr at 77.62%; our single model (41M) reads 95.36%. On the twenty-four it is ko-trocr 79.94% against 97.20%.
  • Nobody has ever seen the eleven photos. That is the cleanest ground here: our ensemble reads 96.52% and the single model 94.24%, against 78.50% for the best rival (ko-trocr). Lines read exactly right: 67.6% against 26.5% (the best rival value).
  • A large VLM does not read a handwritten line well just by being large. Qwen3-VL-2B (2.1B) reads 66.02% of the six fonts and PaddleOCR-VL-1.6 (906M) 62.97%. They are twenty to fifty times our size but were not trained for this job. Reading whole documents is another matter (not measured).
  • Tesseract barely reads handwriting (17.42% on the fonts, 13.21% on the photos). It is a printed-text engine, as expected. It stays as a baseline.
  • Run every single font in the folder — all 480 of them — and we read 98.27% (ensemble) and 98.09% (single) against 82.56% for the best rival (ko-trocr). The 24 fonts we have never seen (99.00%) score higher than the 450 used in training (98.22%) — so this is not a number propped up by memorisation. PaddleOCR-VL takes about a second per line at its cheapest batch, so it was not run on all 480 (blank in the table).
  • Some fonts still collapse. The worst of the 480 sits at 47.31% for our single model and 63.21% for the ensemble. Six fonts would never have shown that. With only 12 lines per font a single font's score swings hard, so read the herd totals and the distribution, not one row.
  • The single model has a lower photo floor than 0.3.0. Its worst photo is hand-06 at 83.03% — one line (전자금융TF서약) of a three-line photo. In the ensemble another model carries it and it stays at 90.28%.
  • The 0.4.1 single model gained on fonts only. Against 0.4.0 (v83) it moves from 95.20 to 95.36% on the six fonts, 96.87 to 97.20% on the twenty-four and 98.00 to 98.09% on all 480 — up on all three rulers. The photo row in this table went from 94.95 to 94.24%: 22 -> 21 of 34 lines read exactly, one line's worth (tools/verify.py gives 95.0 vs 94.9%).

Where the gap opens

CER by line content

CER by line content — lower is better

Latin mixed in
66 lines
form label
60 lines
digits mixed in
84 lines
Hangul word / name
304 lines
Hangul sentence
156 lines
ko-hand-ocr ensemble (118M) 14.1% 8.3% 9.6% 4.7% 8.1%
ko-hand-ocr single model (41M) 14.5% 7.6% 8.7% 5.5% 8.6%
ddobokki/ko-trocr (214M) 59.5% 30.3% 49.4% 25.0% 46.1%
PaddleOCR PP-OCRv5 (3.3M) 53.7% 32.1% 43.7% 33.5% 40.9%
Qwen3-VL-2B (2.1B) 58.7% 96.5% 49.1% 104.5% 82.1%
PaddleOCR-VL-1.6 (906M) 75.7% 65.9% 45.6% 81.7% 51.4%
EasyOCR (4.0M) 68.5% 55.6% 64.8% 64.7% 64.1%
Tesseract 5 (6MB) 90.8% 103.3% 82.3% 114.5% 96.0%

CER by line length

CER by line length — lower is better

1-5 chars
286 lines
6-10 chars
270 lines
11-20 chars
108 lines
21+ chars
6 lines
ko-hand-ocr ensemble (118M) 6.1% 8.4% 8.3% 17.6%
ko-hand-ocr single model (41M) 6.5% 8.9% 8.1% 15.7%
ddobokki/ko-trocr (214M) 30.7% 35.0% 50.7% 92.1%
PaddleOCR PP-OCRv5 (3.3M) 42.9% 33.8% 40.1% 92.1%
Qwen3-VL-2B (2.1B) 101.2% 84.0% 59.0% 85.6%
PaddleOCR-VL-1.6 (906M) 96.3% 50.8% 55.5% 58.8%
EasyOCR (4.0M) 70.4% 58.5% 63.0% 97.2%
Tesseract 5 (6MB) 114.6% 98.9% 86.0% 98.6%
  • Latin abbreviations split them. Character error rate is 14.5% for our single model against 53.7% for the best rival there (PaddleOCR PP-OCRv5). The Latin that turns up on forms is mostly abbreviations (TF, OCR, Codex); reading Hangul and Latin together on one line is where they part.
  • Longer lines split them further. At 21+ characters: 15.7% against 58.8% for the best rival (PaddleOCR-VL-1.6). ko-trocr's encoder is a 384×384 square, so a long line is squashed whole into it; ours is 64×640. The 21+ bucket holds only 6 lines, though — read it as a direction, not a result.

What it costs

Speed and size

Speed and size

Parameters Download lines/s (GPU) lines/s (CPU) 30-line page (GPU) 30-line page (CPU) Peak VRAM
ko-hand-ocr ensemble (118M) 118M 450MB 6.9 0.66 4.76s 45.8s 1772MB
ko-hand-ocr single model (41M) 41M 156MB 34.2 3.08 1.31s 10.2s 1052MB
ddobokki/ko-trocr (214M) 214M 408MB 5.3 0.52 6.09s 58.1s 1798MB
PaddleOCR PP-OCRv5 (3.3M) 3.3M 13MB 336.5 14.68 0.52s 2.5s 111MB
Qwen3-VL-2B (2.1B) 2.1B 4058MB 6.5 0.47 5.02s 64.1s 5724MB
PaddleOCR-VL-1.6 (906M) 906M 1828MB 0.9 0.04 32.80s 680.7s 1937MB
EasyOCR (4.0M) 4.0M 15MB 32.5 14.88 1.36s 2.5s 70MB
Tesseract 5 (6MB) — 6MB — 22.25 — 1.8s —
  • We are not the fastest. On CPU, Tesseract 5 (22.2 lines/s), EasyOCR (14.9 lines/s), PaddleOCR PP-OCRv5 (14.7 lines/s) beat our single model (3.1 lines/s). They read Tesseract 5 17.4%, EasyOCR 56.4%, PaddleOCR PP-OCRv5 72.3% of the six fonts. Read accuracy and speed together — on a form a wrong answer is worse than a blank.
  • Against a model of similar purpose, ko-trocr (214M), we are 5.9× faster on CPU (a thirty-line page in 10.2s against 58.1s). Whether it fits on an office PC with no GPU is decided here.
  • CPU figures ride hard on this laptop's state at the time. Same day, same cells, same method: the single model ran at 4.6 lines/s in the morning, 3.1 in this run and 2.6 in the evening, and its lead over ko-trocr moved through 6.3×, 5.9× and 4.4×. That is why the 0.18 s per line under "Speed and footprint" below (bench.py, measured separately) does not match. Read the order within one run rather than lines/s itself, and the lead as 4–6× — being faster than ko-trocr held all three times.
  • The large VLMs need a GPU. On CPU Qwen3-VL-2B does 0.47 lines/s and PaddleOCR-VL-1.6 0.04, so a thirty-line page takes 64s and 681s. Even on GPU they run at 6.5 and 0.9 lines/s, slower than our single model (34.2). On CPU they take 2s and 23s per line, so only the one-line batch was timed, once.
  • Batch size is measured, not guessed. Each model reads at its own cheapest batch. A batch that does not fit this card is not timed at all — ko-trocr's 32-line batch is one, and the tool's own reason reads: "앞 묶음 4642MB(가중치 848MB)에서 활성값을 두 배 하면 8437MB 로 카드 8151MB 를 넘는다". A number that was not measured is not written down as if it were.

What is not equal — and which side it favours

What How Favours
Test fonts The six and the twenty-four were held out of our training only. What the rivals were trained on is unknown to us rivals
Line cutting Our cutter was lent to every rival (they all receive one line image) rivals
Length cap ko-trocr ships max_length: 16 (character tokenizer). Like the two VLMs it was given 64 tokens. Every line where Qwen3-VL hit the cap was a runaway repeat (11 11 11 …); no long answer was cut short (re-reading the 670 font lines, all 24 that hit the cap were like that) ko-trocr
Prompt Qwen3-VL has no fixed OCR prompt; three were tried and the one it reads best with was used Qwen3-VL
Precision The two VLMs run in their released precision (bfloat16), the rest in float32. ko-trocr's float16 speed was also timed —
VLM on CPU Seconds to tens of seconds per line, so only the one-line batch was timed, once. An 8-line batch might be faster ko-hand-ocr
All 480 fonts PaddleOCR-VL takes about a second per line, so it was not run on all 480 (blank) —
Design intent The three line recognisers target mainly printed and scene text, the VLMs whole documents, ko-trocr AI Hub's handwriting and administrative documents. Only this model targets a single handwritten line ko-hand-ocr
Number of models Our ensemble is three models (one 36M + two 41M, 118M). For a size-matched view read the "single model" row ko-hand-ocr
Training fonts 450 of the 480 fonts are ones we have seen. That is why the herds are reported apart ko-hand-ocr

None of this says the rivals are bad models. Models built for printed documents, scene text or whole pages were handed a line of handwriting, and on that job this one does better. The other direction — whole printed documents, tables, scene text — was not measured, and if it were, this model would probably lose.

Measured on: NVIDIA GeForce RTX 5060 Laptop GPU · torch 2.14.0+cu130 · 2026-10-07 17:32. Accuracy for our two models and ko-trocr matches the same morning's run to the character — same models, same ruler. Speed rides on the machine's state, so compare only figures taken at the same moment.

Install

pip install ko-hand-ocr

Weights ship separately — they are too large for the repository, so they are attached to Releases. The same files are on Hugging Face — the server and the MCP tool below fetch them from there.

Weights are versioned separately from the package. The current weights are on the v0.4.1 release (the package on PyPI is 0.5.0 — 0.5.0 adds the server and the MCP tool below; the reading code and the weights are unchanged). Earlier weights stay where they are — the previous single model (v83) is on v0.4.0, and the 6-layer v62 on v0.3.0 is faster (about 40% when measured the same day) and has a higher photo floor (worst photo 90.3%); the 4-layer weights on v0.2.0 are faster still.

Download Size What it is
ko-hand-ocr-single.zip 144MB Single model. Enough for most uses. 0.18 s/line
ko-hand-ocr-ensemble.zip 417MB Ensemble (3 models). Raises the floor, 7x slower
curl -LO https://github.com/jysvai/ko-hand-ocr/releases/download/v0.4.1/ko-hand-ocr-single.zip

Pass the unzipped folder straight to Reader(). It must contain config.json, vocab.json and model.safetensors.

Quick start

from kohandocr.reader import Reader

reader = Reader("ko-hand-ocr-single", device="cpu")   # the unzipped folder

# A whole photo. Line cutting is included.
reader.read_photo(open("scan.jpg", "rb").read())
# -> ['보안점검표', '부서 : 포테토뭉부서', '이름 : 감자밭']

# If the lines are already cut, hand the images in directly. Batch them for speed.
reader.read([cell_a, cell_b, cell_c])

# Given a candidate list, the output can be constrained to it.
reader.both([cell_b], options=["포테토뭉부서", "감자밭", "김클로드"])
# -> [{'text': '부서 : 포테토뭄부서', 'constrained': '포테토뭉부서', 'agrees': True}]

Do not trust the constrained result on its own. Because it is constrained, the model must return something from the list. If the true answer is not in the list you get a confident wrong answer and nothing looks off. When agrees is false, hand it to a human.

Ensemble reading (several models)

Different models fail on different cells. Train a little longer on the same data and some cells survive while others die — which tells you those cells were being read by a thin margin. So keep several models and take the answer they agree on.

unzipped/
  config.json  vocab.json  model.safetensors
  also/
    v20b/  config.json  vocab.json  model.safetensors
    v19/   ...

If also/ is present, Reader uses it automatically. The call site does not change.

Combining by confidence was measured and it did not work — the model is sometimes more confident about a wrong answer. So selection is by how much the outputs agree, not by confidence.

Use it with LM Studio, Ollama and vLLM

It cannot be loaded into them as a model. All three are engines for LLMs that continue text: LM Studio and Ollama run llama.cpp GGUF files, and vLLM runs the architectures it knows. This model is a TrOCR-style image encoder with a cross-attention decoder, plus its own jamo vocabulary and line cutter, so it neither converts to GGUF nor fits vLLM's list. Instead the package ships a server that speaks the same APIs and an MCP tool.

pip install "ko-hand-ocr[mcp]"      # drop [mcp] if you only need the server

Server — the OpenAI and Ollama APIs

ko-hand-ocr-serve                   # fetches the single model (156MB) from Hugging Face, serves 127.0.0.1:8765
ko-hand-ocr-serve --ensemble        # the ensemble (450MB)
Caller How
openai client, code written for vLLM set base_url="http://127.0.0.1:8765/v1"; send the image as base64 in image_url
ollama CLI set OLLAMA_HOST=127.0.0.1:8765, then ollama run ko-hand-ocr "C:\scan.jpg" · ollama list
curl curl --data-binary @scan.jpg http://127.0.0.1:8765/read
import base64
from openai import OpenAI

client = OpenAI(base_url="http://127.0.0.1:8765/v1", api_key="none")
picture = "data:image/jpeg;base64," + base64.b64encode(open("scan.jpg", "rb").read()).decode()
reply = client.chat.completions.create(model="ko-hand-ocr", messages=[
    {"role": "user", "content": [{"type": "image_url", "image_url": {"url": picture}}]}])
print(reply.choices[0].message.content)     # one line per handwritten line
  • Text in the prompt is ignored. The image is cut into lines and each line comes back as one line of text. It cannot summarise or tidy up — that is the chat model's job (MCP, below).
  • Image URLs (http://…) are not fetched. Send base64. Nothing leaves the machine.
  • By default it listens on this PC only (127.0.0.1). There is no password, so use --host 0.0.0.0 only on a network you trust.

LM Studio — as an MCP tool

Since 0.3.17, chat models in LM Studio can call outside tools (MCP). In ~/.lmstudio/mcp.json:

{"mcpServers": {"ko-hand-ocr": {"command": "ko-hand-ocr-mcp"}}}

If ko-hand-ocr-mcp is not on PATH, give the full path (Scripts\ko-hand-ocr-mcp.exe in the virtual environment). Then ask a chat model that can call tools to "read C:\scan.jpg and put it in a table": it calls read_handwriting for the text and does the tidying itself. This model recognises the characters; the chat model organises them. An image pasted into the chat window goes to the chat model, not the tool, so pass a file path. The first call downloads the model (156MB, 20–30 s here); after that a photo takes about a second.

What was checked and what was not. The ollama CLI 0.34 (list, show, run), the official openai client (whole and streamed replies) and an MCP SDK 2.3 client were all run against it. A full run inside the LM Studio chat window or Open WebUI has not been done yet.

Putting it inside vLLM itself would mean writing the architecture as a plugin. For a 41M model a GPU server buys little, so that was not done — the server above exposes the same OpenAI address vLLM does.

How it is built

  64 x 640 line image
          |
   ViT-Small encoder          facebook/deit-small-patch16-224 (ImageNet-1k, Apache-2.0)
   patch 16                   pretrained weights, continued at a lower learning rate
          |
   TrOCR decoder              8 layers, 6 heads — trained from scratch
          |
   229 jamo tokens            ㄱ ㅏ ㅁ ... assembled into 감
          |
      "부서 : 감자밭"

Two decisions carry most of the result.

Jamo, not syllables. Memorising all 11,172 Hangul syllables needs a large output layer and still fails on anything rare. 229 jamo compose into any syllable, so a character the model never saw in training is still writable — which matters for names.

The encoder is not frozen, but it is not shaken either. It already learned to see on ImageNet, so it continues at a lower learning rate while the blank decoder learns fast. Training both at the same rate destroys what the encoder knew.

Accuracy

Measured two ways. Both matter. All figures below come from python tools/verify.py (about 960 lines per font, 11 photos).

unseen fonts (mean) worst font 11 handwriting photos worst photo items over 95%
single model (0.4.1) 95.6% 90.4% 94.9% 84.2% 11 / 17
ensemble (3 models) 95.6% 90.7% 96.6% 90.3% 12 / 17
0.4.0 single model (v83) 95.5% 90.1% 95.0% 83.0% 10 / 17
0.3.0 single model (v62) 94.6% 88.5% 95.4% 90.3%
0.3.0 ensemble 95.3% 90.3% 95.8% 90.3% 10 / 17

Fonts are measured at about 960 lines each (0.4.0 shipped with 380). Italic rows are earlier releases: the 0.4.0 single model was re-measured with the same 960-line ruler, while the 0.3.0 rows used the old 380-line ruler, so part of those differences is the ruler.

0.4.1 changes the single model only. It was trained to follow the ensemble's next-character probabilities (distillation), then averaged half-and-half with 0.4.0's v83. Size and architecture are unchanged, so speed is too. Items over 95% go from 10 to 11 of 17 (hand-07 94.4 -> 95.8%), all six fonts move up a little (worst font 90.1 -> 90.4%), and the worst photo, hand-06, goes from 83.0 to 84.2%. In exchange hand-08 drops from 100 to 96.7%. This is a small step, not a jump — the photo mean is the same (95.0 vs 94.9%). Re-running the same method at other strengths (0.25, 1.0), and trying dropout, weight EMA and SAM, all failed to beat v83.

What 0.4.0 changed over 0.3.0. The single model (v83) moved from 94.6 to 95.5% on the font mean and from 88.5 to 90.1% on the worst font. In exchange its photo floor dropped (hand-06 90.3 -> 83.0%, one line of a three-line photo). The ensemble holds the floor and moves the photo mean from 95.8 to 96.6%. Of the 17 items (6 fonts + 11 photos), the ensemble clears 95% on 12, up from 10.

Similarity is measured per jamo; on the photo side each photo gets one vote.

Do not take the photo number at face value. It is measured on 11 photos (35 cells) taken by the author, so it swings ±5%p. The figure closer to what you would see on someone else's handwriting is the unseen-fonts column — tested only on six font families never used in training, where the swing is half as wide (±2-3%p).

What the ensemble buys is the floor, not the mean. The font mean is now the same (95.6 vs 95.6), but the worst font goes from 90.4% to 90.7%, and a photo the single model dropped to 84.2% (hand-06) comes back at 90.3%. Where one model reads a cell by a hair, another carries it.

Not every page improves, though. A page the single model cleared at 95.2% can come down to 92.9% (hand-10) — when two models agree on the wrong answer, the one that was right is outvoted. And it is seven times slower. Use the ensemble where a single page reaching a human is costly, the single model where you need to sweep a lot of pages.

Per font (about 960 lines each)

Accuracy by handwriting font

Look only at the average and the font that collapses is hidden. The real spread is close to 9%p.

Font single model (0.4.1) ensemble 0.4.0 single model (v83)
EastSeaDokdo (동해독도) 90.4% 90.7% 90.1%
KirangHaerang (기랑해랑) 94.7% 94.7% 94.7%
HiMelody (하이멜로디) 95.6% 95.4% 95.4%
NanumPenScript (나눔펜스크립트) 95.9% 95.9% 95.7%
GamjaFlower (감자꽃) 98.0% 98.0% 97.9%
SingleDay (싱글데이) 99.0% 98.9% 98.9%

What has to clear the bar is not the mean but the lowest font. Once all six passed 90% (2026-09-18) the goal was raised to 95%. Four are over it now; EastSeaDokdo and KirangHaerang are not. Both are hands where strokes merge and get dropped.

All six are fonts never used in training, and all are SIL Open Font License.

python tools/holdout.py <checkpoint-dir> --per-font --sample sample.png

What the test sheet actually looks like

unseen-font test sheet

Three lines per font, with the ground truth, what the model read, and the similarity. The pale band on the right is the leftover space from fitting into a 64x640 frame — this is exactly what the model sees.

Every one of the six gets these short form lines right. When the same picture was produced on 2026-09-02, EastSeaDokdo missed 안전사무국; it does not any more. Where the fonts pull apart is not lines like these but long sentences and lines no language knowledge helps with — a third of the sheet is meaningless syllables strung together (put there on purpose, to train reading from strokes alone), and there a single wrong jamo has nothing to fall back on.

By shape, SingleDay has well-separated characters and intact strokes, close to print, while EastSeaDokdo and KirangHaerang smear and drop them. That is the difference the table above measures.

This is an image, not a font file. Drawing glyphs with a font is what the OFL permits; what is not redistributed is the .ttf files themselves (PROVENANCE.md).

Speed and footprint (measured)

No GPU needed. Emitting jamo one at a time is the bottleneck, so more cores do not help much — and by the same token it does not get slower on a weak machine.

CPU only, median over 11 photos (35 cells). Time is split into cutting and reading — shrinking the model does not shrink cutting, and when a photo holds only three or four lines, cutting is more than half the cost.

cut read one photo (3.2 lines) per line 30-line page
single model, beam 5 0.48s 0.57s 1.04s 0.18s 6.9s
single model, greedy 0.48s 0.30s 0.78s 0.09s 2.0s
ensemble, beam 5 0.48s 4.36s 4.83s 1.37s 47.1s

Measured on 2026-10-07. The same day, 0.4.0's single model (v83) took 1.00 s per photo and 0.17 s per line — size and architecture are identical, so 0.4.1 is neither heavier nor slower (differences under 5% are day-to-day noise). The machine itself varies by day: on 2026-09-22 the same v83 measured 0.25 s per line. Only compare figures taken on the same day.

Memory, as the peak while reading the 11 photos, is about 1.07GB for the single model and about 1.51GB for the ensemble. Measured on 2026-09-22 (the 0.4.1 single model has the same architecture).

Eight layers are not free. On 2026-09-22, re-measuring 0.3.0's 6-layer v62 with the same harness gave 0.18 s per line (beam 5), 0.09 s (greedy) and 1.03GB, against 0.25 s for the 8-layer model that day. The 8-layer model is about 40% slower at beam 5. If speed comes first, take v62 from v0.3.0.

The right-hand column is not the photo figure multiplied out. Our test photos hold three or four lines each, while the "page" other people quote is a thirty-line document. Quoting our seconds-per-page as if it were theirs flatters us by roughly ten times. So the batch size is raised step by step to measure how far seconds-per-line falls, and that value is what gets converted (python tools/bench.py --throughput).

Produced with python tools/bench.py --all --device cpu. That harness walks exactly the path the application walks — a speed measured along a different path is not the speed the user gets.

Where it runs

Built and tested on Windows. Training and inference were both run on Windows 11 with Python 3.13/3.14.

Windows Where it was built. Training and inference both verified
macOS Untested. Inference is pure PyTorch + PIL so it should run, but it has not been confirmed
Linux Not tried yet

tools/train.ps1 is PowerShell, so it is Windows-only. Elsewhere, call python -m kohandocr.train directly — all that script does is wait for the previous run to release the GPU, launch it, and watch the first few steps.

Training

No data is baked ahead of time. A corpus line is generated, drawn with a handwriting font, and fed in on the spot. Nothing is left on disk, and a change to the distribution takes effect from the next step.

python tools/fetch_fonts.py                       # fetch OFL handwriting fonts
python -m kohandocr.train --out runs/v1 --layers 8 \
       --steps 20000 --batch 40 --workers 5 --letters 128 \
       --strike 0.15 --warp 0.5 --digits 0.04

The last three change the training data only. Test sheets are always drawn with the defaults, so switching them on or off leaves the ruler alone (switched off, they do not even draw a random number — a test holds that).

  • --strike — share of lines that get a crossed-out mark. Four of the 37 photo cells have one.
  • --warp — share of lines whose strokes are bent inside each character (synth._warp). Every earlier distortion moved whole characters, so jamo kept the font's exact shape — yet every photo miss was a single jamo (금->글, 감->갈, 피->회).
  • --digits — share of lines replaced by digit strings that are not dates (corpus.digit_line). The corpus had twice as many date lines as bare digit strings, so a faint digit string got filled in as a date ('816883539' -> '2025-05-19').

Fonts are looked up in this order. No path is hard-coded.

  1. whatever --fonts was given
  2. the KOHAND_FONTS environment variable
  3. ~/.cache/ko-hand-ocr/fonts

Without --layers you get a 4-layer decoder. The single model that ships has 8 layers (the ensemble mixes in one 6-layer model). At 4 layers, thirteen rounds of changing the recipe left the font mean stuck between 92.8 and 93.4%, and 6 layers was the first to clear that band. At 6 layers the lowest item then sat at 90.1% for thirteen rounds; the ensemble only started changing again with 8 layers plus in-character stroke warping. Going from 6 to 8 layers costs about 40% on CPU (measured on the same ruler). The encoder and the input size were left alone.

fetch_fonts.py fetches 482 handwriting fonts. Only 452 of them are trained on; six are the test above and twenty-four are held back for "does it read widely". A font with the same design as a test font would quietly inflate the score, so they are compared as images, not by name, and anything scoring 0.7 or above is set aside.

The fonts are not in this repository and are not redistributed. The 109 Nanum handwriting fonts come from clova.ai/handwriting. Details in PROVENANCE.md.

How closely the synthetic images match real handwriting is checked against the measured distributions recorded in kohandocr/measured.json — nine of them (ink coverage, density, aspect ratio, margins and so on).

python tools/match.py        # target vs current synthesis, side by side

Running longer does not make it better. After changing the synthesis distribution, run short, keep intermediate checkpoints, and pick between them (measured: 15,000 steps was the peak and 30,000 steps was worse).

python tools/pick.py runs/v1-15000 runs/v1-30000    # score the kept checkpoints and pick

How this was built

The engineering behind this model is written up as a 12-page document — why it exists, the four decisions that shaped it, how the training data was produced at zero labelling cost, how the evaluation was designed, and what is still missing.

Engineering write-up — PDF, 12 pages  ·  한국어판

It includes the parts that are usually left out: the measurement that was wrong, the approach that was tried and failed, and the numbers that have not been measured yet.

License

Apache-2.0, weights included. The licenses of the fonts used in training do not attach to the weights — the weights are not a derivative of the glyph artwork, they are values learned from images. The reasoning is recorded part by part in PROVENANCE.md.

Metadata

Release files for ko-hand-ocr 0.5.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for ko-hand-ocr 0.5.0
File Size Uploaded
ko_hand_ocr-0.5.0.tar.gz 164.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for ko-hand-ocr 0.5.0
File Interpreter ABI Platform
ko_hand_ocr-0.5.0-py3-none-any.whl Python 3 none any Details

Total release size: 283.2 kB

Release files / ko_hand_ocr-0.5.0.tar.gz

Download URL ko_hand_ocr-0.5.0.tar.gz
Size 164.3 kB
Tags Source
SHA-256 checksum
How to use checksums
cd1a09ef7fde13a18ed015e07cd6c7435f6fcbc918bd10ed82cca3fb429b2a3a
BLAKE2b-256 checksum
How to use checksums
24120046a8797e49a8775496d4c37239b5828403f495031c46906d2eb76bc805
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.7

Release files / ko_hand_ocr-0.5.0-py3-none-any.whl

Download URL ko_hand_ocr-0.5.0-py3-none-any.whl
Size 118.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
ec2f9e20640c58133270a902fc2ae401a33df7fc1a90824c7618d07cb551f4a6
BLAKE2b-256 checksum
How to use checksums
4712728acf282542e37cdbf4392a174b81a54dd4fe6471d691aeae78ae1df396
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.7

Release history Release notifications | RSS feed

This release

0.5.0 This release

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.2.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page