Skip to main content

Compact CRNN OCR for printed text (Portuguese charset) — pure PyTorch, CPU-friendly, ~360k parameters

Project description

Legere

Compact CRNN OCR for printed text (Portuguese charset with accents, digits, punctuation), built from scratch with PyTorch and trained entirely on synthetic pages generated with PIL. No pretrained models, no OCR libraries.

  • ~360k parameters (hard limit enforced: 1M) — bundled weights are 1.4 MB
  • CRNN: MobileNet-style depthwise separable CNN + 1 BiGRU(128) + linear head, CTC loss
  • Classic (non-neural) deskew + projection-based line segmentation
  • CPU-friendly: full-HD page in well under a second; INT8/TorchScript exports
  • Trained weights ship inside the package — install and read pages immediately

Install

pip install legere              # inference only
pip install legere[train]       # + training extras (rich, psutil)

From a checkout:

pip install -e .[train]

Library usage

from legere import Legere

ocr = Legere()                      # bundled weights, CPU
result = ocr.read("page.png")       # path, or a numpy array (gray or BGR)

print(result.text)                  # full text in reading order
print(result.skew_angle)            # estimated page skew (degrees)
for line in result.lines:
    print(line.text, line.bbox)     # per-line text + (x0, y0, x1, y1)

Options: Legere(model="checkpoints/best.pt") for your own checkpoint, Legere(model="exports/model_int8.ts.pt", torchscript=True) for the INT8 export, beam_width=8 for CTC beam search, threads=N to cap CPU threads.

CLI

legere read page.png                    # text on stdout
legere page.png                         # same (shortcut)
legere read page.png --json out.json    # + per-line text and bboxes
legere read page.png --beam 8           # CTC beam search
legere read page.png --model exports/model_int8.ts.pt --torchscript

legere train                            # train on synthetic pages
legere export                           # TorchScript fp32 + INT8 exports
legere benchmark                        # CPU latency + page CER report

Pipeline of read: fit to 1920x1080 → adaptive binarization + deskew → line segmentation by horizontal projection (rules removed via morphology) → height-32 line batch through the CRNN → CTC greedy decode → text in reading order.

Project layout

src/legere/
├── charset.py        # character set + encode/decode helpers
├── model.py          # CompactCRNN, parameter budget, CTC greedy/beam decode
├── segment.py        # classic deskew + projection line segmentation
├── lines.py          # line-crop normalization and batching
├── inference.py      # Legere engine + full-page pipeline
├── metrics.py        # edit distance, CER, WER
├── cli.py            # `legere` command
├── data/model.pt     # bundled trained weights (state dict + charset)
└── training/         # synthetic pagegen, dataset stream, train/export/benchmark

Training

Training data is generated on the fly: synthetic pages with headers, paragraphs, key/value lines, tables, Portuguese-like words, dates, monetary values and CPF/CNPJ, degraded with noise/blur/JPEG artifacts. Fonts are read from C:\Windows\Fonts (training currently expects Windows).

legere train                          # live Rich dashboard
legere train --no-dashboard           # plain log lines
legere train --resume checkpoints/last.pt

Checkpoints go to checkpoints/ (best.pt by validation CER, last.pt every eval); metrics append to runs/metrics.jsonl. Training stops early once CER < 1% holds with no further improvement, or on a clear plateau. Useful flags: --threads N, --width-budget N (batch size via padded-width budget; lower it to reduce RAM).

After training, refresh the packaged weights:

legere export --checkpoint checkpoints/best.pt --bundle src/legere/data/model.pt

Export & benchmark

legere export        # writes exports/model_fp32.ts.pt and exports/model_int8.ts.pt
legere benchmark     # line/page CPU latency + full-page CER per variant

Results

Bundled weights: validation CER 0.044% (line level, 2000-line fixed set) at step 26k. legere benchmark reproduces line/page latency and full-page CER on synthetic pages.

License

MIT

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

legere-0.1.1.tar.gz (1.4 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

legere-0.1.1-py3-none-any.whl (1.4 MB view details)

Uploaded Python 3

File details

Details for the file legere-0.1.1.tar.gz.

File metadata

  • Download URL: legere-0.1.1.tar.gz
  • Upload date:
  • Size: 1.4 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.10

File hashes

Hashes for legere-0.1.1.tar.gz
Algorithm Hash digest
SHA256 7615722b5ed67121f980b8d3ca3f95ac3dffc12b46f920d672eaf709150c6c3b
MD5 3f199f10956c04f69746f6684043fa59
BLAKE2b-256 4a77ecfead0ff17b01223611311dd53b6863f08fefcbc93cf2ea7ff17b22a560

See more details on using hashes here.

File details

Details for the file legere-0.1.1-py3-none-any.whl.

File metadata

  • Download URL: legere-0.1.1-py3-none-any.whl
  • Upload date:
  • Size: 1.4 MB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.10

File hashes

Hashes for legere-0.1.1-py3-none-any.whl
Algorithm Hash digest
SHA256 a5b9a7b855d16027b434cafad66ecfe5499b148b1e978ed1564697b6ecec713b
MD5 653f4dd6965e1db669e7d5caa802cfd0
BLAKE2b-256 c6797e3f6eab6191e34aba8ba8d9442fbb8e84a103a089ecfa4fa75b7c497d2c

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page