Skip to main content

llama_ocr

A Python package for document text extraction using Ollama Vision models, with a focus on OCR (Optical Character Recognition) capabilities.

Features

  • 🔍 Text extraction from images using Ollama Vision models
  • 🖼️ Advanced image preprocessing for better OCR results
  • 🛠️ Configurable settings and parameters
  • 🔧 Extensible architecture with dependency injection
  • 📝 Comprehensive logging and error handling

Installation

pip install llama_ocr

Or install from source:

git clone https://github.com/princexoleo/llama_ocr.git
cd llama_ocr
pip install -e .

Quick Start

from llama_ocr import DocumentExtractor

# Initialize extractor
extractor = DocumentExtractor()

# Extract text from an image
result = extractor.extract_from_image(
    "path/to/image.jpg",
    prompt="Extract all text from this document"
)

# Print extracted text
print(result["text"])

Advanced Usage

Custom Configuration

from llama_ocr import DocumentExtractor, OCRConfig

config = OCRConfig(
    model_name="llama3.2-vision",
    preprocess_images=True,
    optimize_for_ocr=True,
    log_level="INFO"
)

extractor = DocumentExtractor(config=config)

Image Processing Options

# Extract with OCR optimization
result = extractor.extract_from_image(
    "path/to/image.jpg",
    save_processed=True  # Save processed image for inspection
)

Architecture

The package follows SOLID principles and uses dependency injection for flexibility:

  • DocumentExtractor: Main class orchestrating the extraction process
  • VisionClient: Abstract base class for vision model interactions
  • ImagePreprocessor: Abstract base class for image processing
  • OCRConfig: Configuration management using dataclasses

Components

  1. Core Module

    • extractor.py: Main document extraction logic
    • vision_client.py: Ollama Vision API integration
    • image_processor.py: Image preprocessing utilities
    • base.py: Abstract base classes and interfaces
  2. Configuration

    • config.py: Configuration management using dataclasses

Dependencies

  • ollama>=0.1.27
  • Pillow>=10.1.0
  • python-dotenv>=1.0.0
  • opencv-python>=4.8.1.78
  • numpy>=1.21.0

Development

Setting up Development Environment

# Create virtual environment
python -m venv venv
source venv/bin/activate  # Linux/Mac
# or
.\venv\Scripts\activate  # Windows

# Install development dependencies
pip install -e ".[dev]"

Running Tests

python -m unittest discover -s tests

Contributing

  1. Fork the repository
  2. Create your feature branch (git checkout -b feature/amazing-feature)
  3. Commit your changes (git commit -m 'Add some amazing feature')
  4. Push to the branch (git push origin feature/amazing-feature)
  5. Open a Pull Request

License

This project is licensed under the MIT License - see the LICENSE file for details.

Acknowledgments

  • Thanks to the Ollama team for providing the vision models
  • Built with ❤️ by Mazharul Islam Leon

Metadata

Release files for llama-ocr-py 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for llama-ocr-py 0.1.0
File Size Uploaded
llama_ocr_py-0.1.0.tar.gz 11.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for llama-ocr-py 0.1.0
File Interpreter ABI Platform
llama_ocr_py-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 24.8 kB

Release files / llama_ocr_py-0.1.0.tar.gz

Download URL llama_ocr_py-0.1.0.tar.gz
Size 11.7 kB
Tags Source
SHA-256 checksum
How to use checksums
3f75195157565e7f557f2e0bb293e0b64f21fd310d21c7f4a40e2065a4c1a161
BLAKE2b-256 checksum
How to use checksums
8bd86127f444ca4c9759bb8ed2bcb44f869cf017f65b7e00614342ff9bae0a09
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.0.1 CPython/3.11.3

Release files / llama_ocr_py-0.1.0-py3-none-any.whl

Download URL llama_ocr_py-0.1.0-py3-none-any.whl
Size 13.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
7fec80f86967950c700330f7483a07463b4b69df2030cee245ce9655c53264e9
BLAKE2b-256 checksum
How to use checksums
36173b133acb394df9865654bfb026fe4bd85c1d6b9d71c41f006c7bd6c4078a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.0.1 CPython/3.11.3

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page