One-call PaddleOCR workflow: extract text/numbers/symbols, draw bounding boxes, search for a word/phrase, locate tap coordinates (click_by_ocr), and verify text presence/absence (verify_by_ocr) in an image.
Project description
ocr-locator
A tiny Python package that wraps a full PaddleOCR workflow — engine setup, text/number/symbol extraction, bounding-box drawing, and word search — behind a single class or one CLI command, so you don't have to re-copy notebook cells every time.
Install
pip install ocr-locator
(Or, from a local checkout: pip install .)
First run downloads PaddleOCR's model weights automatically — no manual setup needed.
Usage
Python API
from ocr_locator import OCRLocator
ocr = OCRLocator() # loads the PaddleOCR engine once (put outside any loop)
# 1. Extract all text/numbers/symbols with boxes
extracted = ocr.extract("screenshot.png")
# -> [{"text": "Login", "score": 0.98, "box": [120, 40, 210, 70]}, ...]
# 2. Draw a box + label around every detection
annotated_path = ocr.annotate("screenshot.png", extracted, output_path="out.png")
# 3. Search for a word/phrase; matches are highlighted in a separate image
result = ocr.search("screenshot.png", extracted, "Login")
result["found"] # True/False
result["matches"] # matching detection dicts
result["annotated_path"] # path to the highlighted image, if found
# 4. Locate text and get tap/click coordinates for the best match
click = ocr.click_by_ocr("screenshot.png", "Login")
click["found"] # True/False
click["coordinates"] # {"x": 165, "y": 55} -- center of the best match
click["box"] # [x1, y1, x2, y2] of the best match
click["annotated_image"] # path to an image with all matches highlighted
# 5. Verify a word/phrase is present (or absent)
verify = ocr.verify_by_ocr("screenshot.png", "Error", presence_flag=False)
verify["found"] # was it actually detected?
verify["success"] # found == presence_flag
verify["message"] # human-readable summary
click_by_ocr and verify_by_ocr are also available as module-level
functions that share one lazily-created, reused OCRLocator engine --
handy for a web server handling many uploaded images without reloading
model weights per request:
from ocr_locator import click_by_ocr, verify_by_ocr
click_by_ocr(image_path, "Login")
verify_by_ocr(image_path, "Error", presence_flag=False)
Both work for any image path/format PaddleOCR + Pillow can open
(png/jpg/jpeg/bmp/webp/...) -- there's no dependency on a prior call, a
cached engine, or a specific file extension, so they behave the same
regardless of which image was just uploaded. Repeated calls never
overwrite each other's annotated output: each result image gets a unique
filename. Pass results_dir= (module functions) / output_dir= (methods)
to control where annotated results are written.
Command line
Runs the whole workflow — extract, annotate, search — in one shot:
ocr-locate screenshot.png --search "Login" --output detected.png --json-output detections.json
Omit --search and you'll be prompted interactively; pass --no-prompt to
skip search entirely and just get the all-boxes annotated image.
Get tap coordinates or check presence/absence directly from the CLI:
ocr-locate screenshot.png --click "Login"
ocr-locate screenshot.png --verify "Error" --expect-absent
--results-dir controls where --click/--verify save their annotated
output image (defaults to alongside the input image). --verify exits with
status code 1 if the expectation isn't met, so it's script-friendly.
What it handles for you
- Dependency management —
paddlepaddle,paddleocr,paddlex, andPillowdeclared inpyproject.toml. - Engine initialization —
PaddleOCR(...)created once perOCRLocatorinstance, withenable_mkldnn=Falseset by default to avoid a known oneDNN runtime crash. - Box normalization — handles rectangles, 4-point polygons, and flattened 8-number quads, whatever shape PaddleOCR returns.
- Drawing — labeled bounding boxes for every detection, with separate colors for "all detections" vs. "search matches".
- Search — case-insensitive substring or exact match, with match details printed and a highlighted image saved automatically.
- Click —
click_by_ocr()locates text and returns tap coordinates (center point) for the highest-confidence match, plus every match found. - Verify —
verify_by_ocr()checks whether text is present or absent (viapresence_flag) and reports whether that expectation held. - Works for any uploaded image — no hardcoded extension checks; OCR and annotation both operate on whatever path is passed in, and output filenames are unique per call so results never collide.
Package layout
ocr_locator/
├── __init__.py # public API: OCRLocator
├── config.py # default confidence, colors, font, PaddleOCR init kwargs
├── core.py # OCRLocator class: extract(), annotate(), search(),
│ # click_by_ocr(), verify_by_ocr() + module-level wrappers
├── utils.py # to_rect() box-shape normalizer
└── cli.py # `ocr-locate` command line entry point
Notes
- Unlike the original notebook, this package does not force an interactive Colab file upload/download — you pass a file path in and get a file path out, so it works the same in a script, a notebook, or a server.
ocr.extract()is a pure function-ish call you can run once and reuse for bothannotate()and multiplesearch()calls without re-running OCR.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file ocr_locator-0.2.0.tar.gz.
File metadata
- Download URL: ocr_locator-0.2.0.tar.gz
- Upload date:
- Size: 12.8 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.11.8
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
7e544d2e24fe2945e11913d92a5b750fd5818e899a5d23563e400495374a66bb
|
|
| MD5 |
9667aa756d139f12f6ea01b762102a1f
|
|
| BLAKE2b-256 |
e011548d19ccbb07fdd4e1601fc84288e509d5989748b752624bc0b1c96b4b50
|
File details
Details for the file ocr_locator-0.2.0-py3-none-any.whl.
File metadata
- Download URL: ocr_locator-0.2.0-py3-none-any.whl
- Upload date:
- Size: 12.2 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.11.8
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
5455eccab807d9deff0f8008c326e5e63673ec8d5f1fc6fc1d466be010400314
|
|
| MD5 |
bbc48a3bc1b7e9262decf0f1393c672e
|
|
| BLAKE2b-256 |
37a41f7dd09d93885168c07609e763b8d77664dc91a924a2ca587c5c1cd067f3
|