Skip to main content

pagelens-ocr

OCR for scanned documents that runs on a normal CPU - no GPU, no cloud. Give it a PDF, JPG, JPEG or PNG (the type is detected automatically) and it prints the text of every page, keeping headings, side-by-side fields and tables in place.

Install

Python 3.10 or newer, then:

pip install pagelens-ocr

The first run downloads the OCR models once (~250 MB, needs internet); after that it works offline.

Use from the command line

pagelens-ocr scan.pdf
pagelens-ocr photo.jpg
pagelens-ocr scan.pdf > scan.txt        # save the text to a file

Output - the file name and page number, then that page's text:

scan.pdf - Page 1
# INVOICE
Invoice No   : 1234                    Date : 01/01/2025
| Item     | Qty | Amount |
|----------|-----|--------|
| Paper A4 | 2   | 500.00 |

scan.pdf - Page 2
...

Use from Python

from pagelens_ocr import extract, extract_text

for page in extract("scan.pdf"):          # one result per page
    print(page["file"], page["page"], page["seconds"])
    print(page["text"])

text = extract_text("photo.jpg")          # all pages as one string

What it does

Per page: crop scanner borders, straighten tilted scans, remove punch holes -> detect and read every text line (PP-OCRv6; a fast model reads every line, a stronger one re-reads uncertain lines) -> find headings, text blocks, tables and logos (PP-DocLayout v2) -> rebuild table rows and columns (SLANet+) -> print in reading order. Typically 4-9 seconds per page on a laptop CPU.

Handwriting is not supported. Always check important values (numbers, dates, IDs) against the original.

License

Apache-2.0

Metadata

Release files for pagelens-ocr 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for pagelens-ocr 0.1.0
File Size Uploaded
pagelens_ocr-0.1.0.tar.gz 19.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for pagelens-ocr 0.1.0
File Interpreter ABI Platform
pagelens_ocr-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 38.5 kB

Release files / pagelens_ocr-0.1.0.tar.gz

Download URL pagelens_ocr-0.1.0.tar.gz
Size 19.4 kB
Tags Source
SHA-256 checksum
How to use checksums
97e078f2548211b2e8943be37f25c0576beda762f844336b242a0c467b9643df
BLAKE2b-256 checksum
How to use checksums
a384c6b3f7a275b0891b9502ddec10e21d856cd1a7e96b9231f7ef62643d658e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / pagelens_ocr-0.1.0-py3-none-any.whl

Download URL pagelens_ocr-0.1.0-py3-none-any.whl
Size 19.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
faf2d7f04e1375d3e1ed2f58d04665c55d046ce40b67566c17acd122f48b6806
BLAKE2b-256 checksum
How to use checksums
0c2a55c7e1e46fe8920c422a11de37c67c0d495117cbfdf1322a46e347d9877d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page