pagelens-ocr
OCR for scanned documents that runs on a normal CPU - no GPU, no cloud. Give it a PDF, JPG, JPEG or PNG (the type is detected automatically) and it prints the text of every page, keeping headings, side-by-side fields and tables in place.
Install
Python 3.10 or newer, then:
pip install pagelens-ocr
The first run downloads the OCR models once (~250 MB, needs internet); after that it works offline.
Use from the command line
pagelens-ocr scan.pdf
pagelens-ocr photo.jpg
pagelens-ocr scan.pdf > scan.txt # save the text to a file
Output - the file name and page number, then that page's text:
scan.pdf - Page 1
# INVOICE
Invoice No : 1234 Date : 01/01/2025
| Item | Qty | Amount |
|----------|-----|--------|
| Paper A4 | 2 | 500.00 |
scan.pdf - Page 2
...
Use from Python
from pagelens_ocr import extract, extract_text
for page in extract("scan.pdf"): # one result per page
print(page["file"], page["page"], page["seconds"])
print(page["text"])
text = extract_text("photo.jpg") # all pages as one string
What it does
Per page: crop scanner borders, straighten tilted scans, remove punch holes -> detect and read every text line (PP-OCRv6; a fast model reads every line, a stronger one re-reads uncertain lines) -> find headings, text blocks, tables and logos (PP-DocLayout v2) -> rebuild table rows and columns (SLANet+) -> print in reading order. Typically 4-9 seconds per page on a laptop CPU.
Handwriting is not supported. Always check important values (numbers, dates, IDs) against the original.
License
Apache-2.0
Metadata
Release files for pagelens-ocr 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| pagelens_ocr-0.1.0.tar.gz | 19.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| pagelens_ocr-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 38.5 kB
Release files / pagelens_ocr-0.1.0.tar.gz
| Download URL | pagelens_ocr-0.1.0.tar.gz |
|---|---|
| Size | 19.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
97e078f2548211b2e8943be37f25c0576beda762f844336b242a0c467b9643df
|
|
BLAKE2b-256 checksum How to use checksums |
a384c6b3f7a275b0891b9502ddec10e21d856cd1a7e96b9231f7ef62643d658e
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Release files / pagelens_ocr-0.1.0-py3-none-any.whl
| Download URL | pagelens_ocr-0.1.0-py3-none-any.whl |
|---|---|
| Size | 19.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
faf2d7f04e1375d3e1ed2f58d04665c55d046ce40b67566c17acd122f48b6806
|
|
BLAKE2b-256 checksum How to use checksums |
0c2a55c7e1e46fe8920c422a11de37c67c0d495117cbfdf1322a46e347d9877d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|