Skip to main content

Python tools for Tesseract OCR training

Training tools for Tesseract OCR.

Installation

Install using pip:

pip install pytesstrain

This will also install Python packages pytesseract (used for running Tesseract) and editdistance (used for calculation of error rates).

Getting started

This package contains tools for specific problems:

text2image is crashing (issue #1781 @ Tesseract OCR)

The text2image tool crashes, if text lines are too long. As stated in the issue above, rewrapping text lines to smaller length is the official workaround for this problem. For example, to reduce line length to 35 characters at most, run

rewrap corpus.txt corpus-35.txt 35

Creating dictionary data from corpus file

In case you do not have a dictionary file for the training language, you might want to create one from the corpus file. To create dictionary file for the language lang, run

create_dictdata -l lang -i corpus.txt -d ./langdata/lang

This tool creates following files:

  • lang.training_text (copy of the corpus file)
  • lang.wordlist (dictionary)
  • lang.word.bigrams (word bigrams)
  • lang.training_text.bigram_freqs (character bigram frequencies)
  • lang.training_text.unigram_freqs (character frequencies)

The file lang.wordlist.freq is usually created by training tools, such as tesstrain.sh and the likewise, so there is no need to create it with create_dictdata.

Language metrics

The tool language_metrics runs Tesseract OCR over images of random word sequences, which are created out of the supplied wordlist, and calculates median metrics (currently CER and WER) from the results. It enables you to assess the quality of your .traineddata file.

To calculate metrics for the language lang with fonts Arial and Courier using wordlist file lang.wordlist, run

language_metrics -l lang -w lang.wordlist --fonts Arial,Courier

Creating unicharambigs file

There are two tools in this package, which enable automatic creation of an unicharambigs file.

The first tool, collect_ambiguities, compares the recognised text with the reference text and extracts smallest possible differences as error and correction pairs, and stores them sorted by frequency of occurrence in a JSON file. You may look at the ambiguities by yourself before converting them to unicharambigs file with the second tool.

The second tool, json2unicharambigs, takes the intermediate JSON file and puts the ambiguities into the unicharambigs file. The resulting file has v2 format. You may limit the ambiguities, which go into the unicharambigs file, with additional command-line switches.

To create the file lang.unicharambigs for the language lang using wordlist file lang.wordlist, run

collect_ambiguities -l lang -w lang.wordlist --fonts Arial,Courier -o ambigs.json
json2unicharambigs --mode safe --mandatory_only ambigs.json lang.unicharambigs

Creating ground truth files

To help with training of Tesseract>=4, the tool create_ground_truth creates single-line ground truth files either from an input file or from a directory with .txt files (the tool searches for the latter ones recursively).

To create ground truth files from a directory corpora in the directory ground-truth, run

create_ground_truth --fonts Arial,Courier corpora ground-truth

API Reference

The main workhorse is the function pytesstrain.train.run_test. There is also a parallel version, pytesstrain.train.run_tests, which uses a pool of threads to run the former function on multiple processors simultaneously (using threads instead of processes for parallelisation is possible, because the run_test function starts processes itself and is thus I/O-bound).

The subpackage tesseract simply imports the package pytesseract. The subpackage text2image imitates the former one, but for the text2image tool instead of tesseract.

The subpackage metrics contains implementation of evaluation metrics, such as CER and WER. The subpackage utils has often-used, miscellaneous functions and the subpackage ambigs contains ambiguity processing functions.

Finally, the subpackage cli contains the console scripts.

License

Pytesstrain is released under Apache License 2.0.

Metadata

Release files for pytesstrain 0.1.16

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for pytesstrain 0.1.16
File Size Uploaded
pytesstrain-0.1.16.tar.gz 18.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for pytesstrain 0.1.16
File Interpreter ABI Platform
pytesstrain-0.1.16-py3-none-any.whl Python 3 none any Details

Total release size: 42.7 kB

Release files / pytesstrain-0.1.16.tar.gz

Download URL pytesstrain-0.1.16.tar.gz
Size 18.7 kB
Tags Source
SHA-256 checksum
How to use checksums
d207d520b1985e6d7c6e4fcdade60a9303f385c3180c80efb6baff9373dedc9c
BLAKE2b-256 checksum
How to use checksums
461d3553329f01fa7a209103f306a32209ef6cddd01a7e6d279ba387a2044a88
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.0 CPython/3.9.12

Release files / pytesstrain-0.1.16-py3-none-any.whl

Download URL pytesstrain-0.1.16-py3-none-any.whl
Size 24.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
6085d80bb2c3b9140abc805af0202aa1f819d85cb0912a5a54828371907c419c
BLAKE2b-256 checksum
How to use checksums
db7a130b6959770092c6266ff032d0847c98ea2fb01e351b7ae201da3d0b295c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.0 CPython/3.9.12

Release history Release notifications | RSS feed

This release

0.1.16 This release

2 release files

0.1.12

2 release files

0.1.11

2 release files

0.1.10

2 release files

0.1.9

2 release files

0.1.8

2 release files

0.1.7

2 release files

0.1.6

2 release files

0.1.5

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page