ocr-locator
A tiny Python package that wraps a full PaddleOCR workflow — engine setup, text/number/symbol extraction, bounding-box drawing, and word search — behind a single class or one CLI command, so you don't have to re-copy notebook cells every time.
Install
pip install ocr-locator
(Or, from a local checkout: pip install .)
First run downloads PaddleOCR's model weights automatically — no manual setup needed.
Usage
Python API
from ocr_locator import OCRLocator
ocr = OCRLocator() # loads the PaddleOCR engine once (put outside any loop)
# 1. Extract all text/numbers/symbols with boxes
extracted = ocr.extract("screenshot.png")
# -> [{"text": "Login", "score": 0.98, "box": [120, 40, 210, 70]}, ...]
# 2. Draw a box + label around every detection
annotated_path = ocr.annotate("screenshot.png", extracted, output_path="out.png")
# 3. Search for a word/phrase; matches are highlighted in a separate image
result = ocr.search("screenshot.png", extracted, "Login")
result["found"] # True/False
result["matches"] # matching detection dicts
result["annotated_path"] # path to the highlighted image, if found
# 4. Locate text and get tap/click coordinates for the best match
click = ocr.click_by_ocr("screenshot.png", "Login")
click["found"] # True/False
click["coordinates"] # {"x": 165, "y": 55} -- center of the best match
click["box"] # [x1, y1, x2, y2] of the best match
click["annotated_image"] # path to an image with all matches highlighted
# 5. Verify a word/phrase is present (or absent)
verify = ocr.verify_by_ocr("screenshot.png", "Error", presence_flag=False)
verify["found"] # was it actually detected?
verify["success"] # found == presence_flag
verify["message"] # human-readable summary
click_by_ocr and verify_by_ocr are also available as module-level
functions that share one lazily-created, reused OCRLocator engine --
handy for a web server handling many uploaded images without reloading
model weights per request:
from ocr_locator import click_by_ocr, verify_by_ocr
click_by_ocr(image_path, "Login")
verify_by_ocr(image_path, "Error", presence_flag=False)
Both work for any image path/format PaddleOCR + Pillow can open
(png/jpg/jpeg/bmp/webp/...) -- there's no dependency on a prior call, a
cached engine, or a specific file extension, so they behave the same
regardless of which image was just uploaded. Repeated calls never
overwrite each other's annotated output: each result image gets a unique
filename. Pass results_dir= (module functions) / output_dir= (methods)
to control where annotated results are written.
Command line
Runs the whole workflow — extract, annotate, search — in one shot:
ocr-locate screenshot.png --search "Login" --output detected.png --json-output detections.json
Omit --search and you'll be prompted interactively; pass --no-prompt to
skip search entirely and just get the all-boxes annotated image.
Get tap coordinates or check presence/absence directly from the CLI:
ocr-locate screenshot.png --click "Login"
ocr-locate screenshot.png --verify "Error" --expect-absent
--results-dir controls where --click/--verify save their annotated
output image (defaults to alongside the input image). --verify exits with
status code 1 if the expectation isn't met, so it's script-friendly.
What it handles for you
- Dependency management —
paddlepaddle,paddleocr,paddlex, andPillowdeclared inpyproject.toml. - Engine initialization —
PaddleOCR(...)created once perOCRLocatorinstance, withenable_mkldnn=Falseset by default to avoid a known oneDNN runtime crash. - Box normalization — handles rectangles, 4-point polygons, and flattened 8-number quads, whatever shape PaddleOCR returns.
- Drawing — labeled bounding boxes for every detection, with separate colors for "all detections" vs. "search matches".
- Search — case-insensitive substring or exact match, with match details printed and a highlighted image saved automatically.
- Click —
click_by_ocr()locates text and returns tap coordinates (center point) for the highest-confidence match, plus every match found. - Verify —
verify_by_ocr()checks whether text is present or absent (viapresence_flag) and reports whether that expectation held. - Works for any uploaded image — no hardcoded extension checks; OCR and annotation both operate on whatever path is passed in, and output filenames are unique per call so results never collide.
Package layout
ocr_locator/
├── __init__.py # public API: OCRLocator
├── config.py # default confidence, colors, font, PaddleOCR init kwargs
├── core.py # OCRLocator class: extract(), annotate(), search(),
│ # click_by_ocr(), verify_by_ocr() + module-level wrappers
├── utils.py # to_rect() box-shape normalizer
└── cli.py # `ocr-locate` command line entry point
Notes
- Unlike the original notebook, this package does not force an interactive Colab file upload/download — you pass a file path in and get a file path out, so it works the same in a script, a notebook, or a server.
ocr.extract()is a pure function-ish call you can run once and reuse for bothannotate()and multiplesearch()calls without re-running OCR.
Release files for ocr-locator 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| ocr_locator-0.2.0.tar.gz | 12.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| ocr_locator-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 25.0 kB
Release files / ocr_locator-0.2.0.tar.gz
| Download URL | ocr_locator-0.2.0.tar.gz |
|---|---|
| Size | 12.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
7e544d2e24fe2945e11913d92a5b750fd5818e899a5d23563e400495374a66bb
|
|
BLAKE2b-256 checksum How to use checksums |
e011548d19ccbb07fdd4e1601fc84288e509d5989748b752624bc0b1c96b4b50
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.11.8
|
Release files / ocr_locator-0.2.0-py3-none-any.whl
| Download URL | ocr_locator-0.2.0-py3-none-any.whl |
|---|---|
| Size | 12.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
5455eccab807d9deff0f8008c326e5e63673ec8d5f1fc6fc1d466be010400314
|
|
BLAKE2b-256 checksum How to use checksums |
37a41f7dd09d93885168c07609e763b8d77664dc91a924a2ca587c5c1cd067f3
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.11.8
|