Skip to main content

Opt-in glyphive OCR model fine-tuned for Liberation Mono (base16c channel).

Project description

glyphive-ocrmodel-libmono

An opt-in OCR model for glyphive, fine-tuned for Liberation Mono renderings of the base16c alphabet.

pip install glyphive-ocrmodel-libmono
glyphive extract -f scan/ --from-images --ocr-engine tesseract-glyphive-libmono -C out

Installing it registers a tesseract-glyphive-libmono OCR provider through the glyphive.ocr_providers entry point. If the model file or Tesseract is not present the provider reports itself unavailable and glyphive falls back to a core engine — installing this package never breaks the stock path.

Why

On the held-out training sweep (benchmarks/results/ocr-training-sweep-20260718.json in the core repo) the fine-tuned model reads the base16c channel at 0.000% CER clean and blurred, versus ~4.6% for stock eng. The core document-wide Reed-Solomon + per-line CRC already make the stock path restore correctly, so this model is a robustness upgrade for marginal scans, not a requirement.

The model file (not in source control)

The trained glyphivelibmono.traineddata (~15 MB) is not committed. It is produced by the reproducible VM recipe in the core repo (benchmarks/training/measure_sweep.py / train_ocr_models.py) and added as package data only when a release wheel is built:

# uncomment in pyproject.toml [tool.hatch.build.targets.wheel] once present:
artifacts = ["src/glyphive_ocrmodel_libmono/glyphivelibmono.traineddata"]

Licensing

  • Base model: a fine-tune of Tesseract tessdata_best/eng (Apache-2.0, redistributable). Exact base used: version 4.00.00alpha:eng:synth20170629 (tessdata_best; trained with Tesseract 5.4.1), SHA-256 8280aed0782fe27257a68ea10fe7ef324ca0f8d85bd2fd145d1c2b560bcb66ba.
  • Training corpus: synthetic — randomly generated base16c-alphabet lines rendered to images. Liberation Mono under SIL OFL 1.1 (Reserved Font Name 'Liberation'). Only synthetic renders were used for training; no font file ships in this wheel.
  • Engine compatibility: the LSTM .traineddata format is shared across the Tesseract 4.x/5.x line, so one model file serves both (verified).

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distributions

No source distribution files available for this release.See tutorial on generating distribution archives.

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

glyphive_ocrmodel_libmono-0.1.0-py3-none-any.whl (8.0 MB view details)

Uploaded Python 3

File details

Details for the file glyphive_ocrmodel_libmono-0.1.0-py3-none-any.whl.

File metadata

File hashes

Hashes for glyphive_ocrmodel_libmono-0.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 23870de269ad0b75aa414bec3a0cf649522cb576aa16377a5f17e7cc095ce59c
MD5 e09c7503671fbdbc89f51b0231029869
BLAKE2b-256 30b66b506485b04fbe99311aba06b3ceefa14ac8f4620d644c5f47f0f3921e76

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page