VisionParse
VisionParse is a small, practical toolkit for turning messy image-based documents into useful text and data. It wraps the pieces that usually end up scattered across notebooks: OCR, image preprocessing, YOLO/object detection, price extraction, and optional LLM cleanup.
It started life as a set of computer-vision experiments. This package gives those ideas a proper home: import-safe modules, a CLI, tests, PyPI metadata, and GitHub Actions publishing.
The heart of the project is still research-minded: use free/local OCR first, keep text localized with bounding boxes, preserve the page/menu layout as much as possible, and only bring in heavier YOLO or LLM tools when they genuinely help.
pip install visionparse-free-ocr
The PyPI distribution is named visionparse-free-ocr. The Python import stays short:
import visionparse
What it does
- Runs OCR with Tesseract, EasyOCR, Keras OCR, or Google Vision.
- Preprocesses images before OCR: resize, grayscale, denoise, threshold, contrast, crop.
- Runs YOLO detections and returns clean bounding boxes.
- Groups localized OCR words into lines and blocks so aligned text stays aligned.
- Extracts prices from noisy OCR text.
- Turns menu-like OCR into lightweight structured items.
- Optionally asks an LLM/LangChain flow to clean up the structure.
- Provides one document pipeline and one CLI so the pieces fit together.
Installation
The base install is intentionally light:
pip install visionparse-free-ocr
For OCR with Tesseract:
pip install "visionparse-free-ocr[ocr]"
You still need the Tesseract system binary installed. On Windows, install Tesseract and either add it to PATH or pass the path when you create the engine.
For YOLO detection:
pip install "visionparse-free-ocr[yolo]"
For the full kitchen sink:
pip install "visionparse-free-ocr[all]"
Extras are split this way because object detection, EasyOCR, Keras OCR, and Google Vision pull in heavier dependencies. Most projects do not need all of them at once.
Quick start
Extract prices from text
from visionparse import extract_prices
text = "Chicken Biryani £8.99\nFamily Platter 24.50\nMango Lassi Rs. 450"
for price in extract_prices(text):
print(price.raw, price.amount, price.currency)
Parse menu-like OCR text
from visionparse import extract_menu_items
ocr_text = """
Starters
Samosa £3.50
Chicken Pakora £5.99
Mains
Lamb Karahi £12.95
"""
items = extract_menu_items(ocr_text)
for item in items:
print(item.name, item.prices, item.category)
OCR an image with Tesseract
from visionparse.ocr.engine import TesseractOCR
ocr = TesseractOCR(
languages="eng",
config="--oem 3 --psm 6",
# tesseract_cmd=r"C:\Program Files\Tesseract-OCR\tesseract.exe",
)
result = ocr.read("menu.jpg")
print(result.text)
Run the document pipeline
from visionparse.pipelines.document_pipeline import DocumentPipeline
pipeline = DocumentPipeline(ocr_engine="tesseract")
result = pipeline.run("menu.jpg")
print(result.text)
print(result.layout_text) # layout-preserving text when OCR boxes are available
print([price.raw for price in result.prices])
print([item.to_dict() for item in result.items])
Preserve layout from localized OCR
from visionparse.ocr.localization import TextToken, group_tokens_into_lines, render_aligned_text
tokens = [
TextToken("Burger", (10, 10, 70, 25)),
TextToken("£7.99", (180, 10, 230, 25)),
TextToken("Fries", (10, 45, 55, 60)),
TextToken("£2.50", (180, 45, 230, 60)),
]
lines = group_tokens_into_lines(tokens)
print(render_aligned_text(lines, char_width=10))
Use YOLO regions before OCR
from visionparse.detection.yolo import YoloDetector
from visionparse.pipelines.document_pipeline import DocumentPipeline
detector = YoloDetector("models/menu-sections.pt", confidence=0.25)
pipeline = DocumentPipeline(ocr_engine="tesseract", detector=detector)
result = pipeline.run("menu.jpg")
for region in result.regions:
print(region.box, region.text[:120])
Model weights are not bundled. Keep them outside the package or in visionparse/models/ locally, but do not commit them.
The original research code referenced fine-tuned YOLO weights such as best (1).pt and best (2).pt. Those binary files were not present in this workspace when the public package was prepared. The repo does include the safe model reference assets that were present, including yolov3.cfg, coco.names, and COCO/TensorFlow config files. See docs/model-card.md.
Darknet/OpenCV YOLO is also supported:
from visionparse.detection.yolo import OpenCVDarknetYoloDetector
detector = OpenCVDarknetYoloDetector(
weights_path="models/yolov3.weights",
# config_path and names_path default to the packaged yolov3.cfg/coco.names
)
detections = detector.detect("menu.jpg")
Command line
After installation, the visionparse command is available.
OCR:
visionparse ocr menu.jpg --engine tesseract --lang eng --pretty
Extract prices from a string:
visionparse prices "Burger £7.99 Fries 2.50" --pretty
Extract prices from a file:
visionparse prices --file ocr-output.txt --pretty
Run the full parser:
visionparse parse menu.jpg --engine tesseract --pretty
Run the parser with YOLO regions:
visionparse parse menu.jpg --engine tesseract --yolo-model models/menu-sections.pt --pretty
Run YOLO only:
visionparse detect menu.jpg --model models/menu-sections.pt --pretty
Run Darknet YOLO with the packaged config/labels and your local weights:
visionparse detect menu.jpg --backend darknet --model models/yolov3.weights --pretty
Save an annotated detection image:
visionparse detect menu.jpg --model models/menu-sections.pt --output annotated.jpg
OCR engines
Tesseract
Good default when you want a local, lightweight OCR engine. Install the Python extra and the system binary:
pip install "visionparse-free-ocr[ocr]"
from visionparse.ocr.engine import TesseractOCR
ocr = TesseractOCR(languages="eng+ara", config="--oem 3 --psm 6")
print(ocr.read("receipt.jpg").text)
If Tesseract is installed in a custom location:
ocr = TesseractOCR(tesseract_cmd=r"C:\Program Files\Tesseract-OCR\tesseract.exe")
You can also set:
set TESSERACT_CMD=C:\Program Files\Tesseract-OCR\tesseract.exe
EasyOCR
pip install "visionparse-free-ocr[easyocr]"
from visionparse.ocr.engine import EasyOCR
ocr = EasyOCR(languages=("en",))
result = ocr.read("shop-sign.jpg")
Keras OCR
pip install "visionparse-free-ocr[keras]"
from visionparse.ocr.engine import KerasOCR
ocr = KerasOCR()
result = ocr.read("menu.jpg")
Google Vision
pip install "visionparse-free-ocr[google]"
Use Application Default Credentials, or pass a service-account file at runtime. Do not commit the JSON file.
from visionparse.ocr.engine import GoogleVisionOCR
ocr = GoogleVisionOCR(credentials_path="local-only-service-account.json")
print(ocr.read("invoice.jpg").text)
LLM/LangChain cleanup
The regular parser is deterministic and does not need an API key. If you want LLM cleanup, install the LLM extra and use an environment variable:
pip install "visionparse-free-ocr[llm]"
set OPENAI_API_KEY=your-key-here
from visionparse.extraction.structured_text import structure_with_llm
cleaned = structure_with_llm(raw_ocr_text, model="gpt-4o-mini")
print(cleaned)
No OpenAI key is stored in the package. The code reads from OPENAI_API_KEY at runtime.
Research notes, examples, and benchmarks
The repo includes:
docs/research.md— project findings and outcomes from the OCR/layout experiments.docs/model-card.md— how the fine-tuned YOLO model should be handled.examples/— free OCR and YOLO+OCR usage scripts.benchmarks/— lightweight text/layout benchmarks plus an optional local image OCR runner.
The public package does not commit the old generated images, notebooks, OCR outputs, or service-account files. If you want to benchmark the legacy images locally, keep them in .visionparse_private_legacy/ or another local folder:
python benchmarks/run_benchmarks.py --images .visionparse_private_legacy
Package layout
visionparse/
├── detection/
│ └── yolo.py
├── ocr/
│ ├── engine.py
│ └── preprocessing.py
├── extraction/
│ ├── prices.py
│ └── structured_text.py
├── pipelines/
│ └── document_pipeline.py
├── models/
├── cli.py
└── tests/
Publishing to PyPI from GitHub Actions
This repo includes .github/workflows/publish.yml.
To publish:
- Create a PyPI API token.
- Add it to the GitHub repository secrets as
PYPI_API_TOKEN. - Push a version tag:
git tag v0.1.1
git push origin v0.1.1
The workflow builds the source distribution and wheel, checks them with Twine, and publishes to PyPI using the secret.
You can also run the publish workflow manually from GitHub Actions.
Security notes
This package should not contain:
- OpenAI keys
- AWS keys
- Google service-account JSON files
- YOLO weights
- generated OCR output files
- test images or notebook outputs
Use environment variables or local-only files instead:
set OPENAI_API_KEY=...
set GOOGLE_APPLICATION_CREDENTIALS=C:\path\to\service-account.json
The .gitignore is set up to keep the common mistakes out of the repo.
Development
python -m pip install -e ".[dev]"
python -m pytest
python -m build
twine check dist/*
The tests avoid heavyweight OCR/model dependencies. They check the parser, price extraction, and import safety first; model-specific tests can be added later with fixtures.
License
MIT.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file visionparse_free_ocr-0.1.1.tar.gz.
File metadata
- Download URL: visionparse_free_ocr-0.1.1.tar.gz
- Upload date:
- Size: 40.4 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
d162f87a0893146a5868b9ad94b16bc754c808210ed805c19c303c9ae5d32265
|
|
| MD5 |
19dfb49bea284e90cc5333a1062ba45e
|
|
| BLAKE2b-256 |
a0b8e2ffd4862da33c3540365c8dbb98233f34b8ea1ee45cec4aa3da54dec6e5
|
File details
Details for the file visionparse_free_ocr-0.1.1-py3-none-any.whl.
File metadata
- Download URL: visionparse_free_ocr-0.1.1-py3-none-any.whl
- Upload date:
- Size: 34.6 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
74a0d442750ca369a667905c41e54930eb5d114d445b620ec8d858094739cb30
|
|
| MD5 |
23f6044a327b41dcc284a17c604ef37c
|
|
| BLAKE2b-256 |
043902dbd352f86937bcca393d46b67c978cebb494d52b33cadb5ac2eb6a067a
|