Skip to main content

Retrieval Augmented Generation with Image Snippets from PDFs

Project description

SnipRAG: Retrieval Augmented Generation with Image Snippets

License: MIT Python 3.8+ PyPI version

SnipRAG is a specialized Retrieval Augmented Generation (RAG) system that not only finds semantically relevant text in PDF documents but also extracts precise image snippets from the areas containing the matching text.

SnipRAG Architecture

Key Features

  • Semantic PDF Search: Find information in PDF documents using natural language queries
  • Image Snippet Extraction: Get visual context from the exact regions containing relevant information
  • Multiple Extraction Strategies: Choose between semantic text extraction or OCR-based processing
  • Precise Coordinate Mapping: Maps text matches to their exact visual location in the document
  • Customizable Snippet Size: Adjust padding around text regions to control snippet size
  • S3 Integration: Process documents stored in Amazon S3
  • Flexible Filtering: Filter search results by document, page, or custom metadata

Installation

From PyPI

pip install sniprag

For visualization support (recommended for demos):

pip install sniprag[viz]

For OCR support:

pip install sniprag[ocr]

For all features:

pip install sniprag[all]

From Source

git clone https://github.com/ishandikshit/SnipRAG.git
cd SnipRAG
pip install -e .

From GitHub

You can install directly from GitHub using pip:

pip install git+https://github.com/ishandikshit/SnipRAG.git

For visualization and OCR support:

pip install "git+https://github.com/ishandikshit/SnipRAG.git#egg=sniprag[all]"

Quick Start

from sniprag import create_engine

# Initialize the engine with semantic strategy (default)
engine = create_engine("semantic", num_blocks=20, block_overlap=0.2)

# Or use OCR-based strategy
# engine = create_engine("ocr", num_slices=10, tesseract_cmd="/path/to/tesseract")

# Process a PDF document
engine.process_pdf("path/to/document.pdf", "document-id")

# Search with image snippets
results = engine.search_with_snippets("your search query", top_k=3)

# Access results
for result in results:
    print(f"Text: {result['text']}")
    print(f"Page: {result['metadata']['page_number']}")
    print(f"Score: {result['score']}")
    
    # The image snippet is available as base64 data that can be displayed or saved
    if "image_data" in result:
        image_base64 = result["image_data"]
        # Use this to display or save the image

Extraction Strategies

SnipRAG offers two different text extraction strategies:

Semantic Strategy

The semantic strategy uses PyMuPDF's built-in text extraction:

  • Divides each page into configurable horizontal blocks (default: 20)
  • Extracts text directly from the PDF structure
  • Works best with native digital PDFs
  • More precise for well-structured documents
engine = create_engine("semantic", num_blocks=20, block_overlap=0.2)

OCR Strategy

The OCR strategy uses Tesseract OCR:

  • Divides each page into configurable horizontal slices (default: 10)
  • Performs OCR on each slice to extract text
  • Better for scanned documents or images
  • More resilient to poor quality documents
engine = create_engine("ocr", num_slices=10, tesseract_cmd="/path/to/tesseract")

Example Snippets

Here are some examples of SnipRAG in action, showing how it extracts image snippets from PDF documents based on semantic search queries:

Structured Table Extraction

Query: "quarterly financial performance table"

Financial Table Snippet

SnipRAG preserves the entire table structure, making it possible to understand relationships between rows and columns that would be lost in text-only extraction.

Specific Table Cell Data

Query: "Q2 2022 revenue"

Q2 Revenue Snippet

When searching for specific data points within tables, SnipRAG extracts not just the matching cell but also the surrounding context, showing related row and column data.

Financial Data with Context

Query: "total profit"

Total Profit Snippet

For financial documents, seeing the numbers in their original tabular format provides critical context that would be lost in pure text extraction.

Technical Comparison Charts

Query: "technical components comparison"

Technical Comparison Snippet

Complex comparison tables maintain their structure in the extracted snippets, making it easier to understand the relationships between different items.

Note: To generate these snippets yourself, run the basic demo with a sample PDF as shown in the Demos section below.

Demos

SnipRAG includes several demo applications:

Strategy Demo

Compare different extraction strategies with a test document:

python demo_strategies.py --strategy semantic  # Default strategy
python demo_strategies.py --strategy ocr --tesseract /path/to/tesseract

Basic Demo

Process a local PDF file and search for information with image snippets:

python examples/basic_demo.py --pdf path/to/document.pdf

S3 Demo

Process a PDF stored in Amazon S3:

python examples/s3_demo.py --s3-uri s3://bucket/path/to/document.pdf --aws-profile your-profile

Jupyter Notebook Example

For those working in Jupyter environments, there's also a notebook example available:

# View the notebook example
cat examples/example_notebook.md

This markdown file contains code snippets you can use in a Jupyter notebook to process PDFs and visualize search results with image snippets.

How It Works

SnipRAG combines semantic search with coordinate mapping to provide visual context:

  1. PDF Processing:

    • Extracts text using the selected strategy (semantic or OCR)
    • Renders page images at high resolution
    • Creates text embeddings for semantic search
  2. Search Process:

    • User submits a natural language query
    • System finds semantically similar text using embeddings
    • For each match, it identifies the exact location in the PDF
    • It extracts an image snippet from that location
  3. Result Delivery:

    • Returns the matching text
    • Provides a visual snippet of the area containing the text
    • Includes metadata (page number, coordinates, etc.)

API Reference

Factory Function

The main entry point for creating SnipRAG engines.

from sniprag import create_engine

engine = create_engine(
    strategy="semantic",  # "semantic" or "ocr"
    **kwargs  # Strategy-specific parameters
)

BaseSnipRAGEngine

Base class for all SnipRAG engines.

SemanticSnipRAGEngine

Engine that uses PyMuPDF's text extraction.

engine = create_engine(
    "semantic",
    num_blocks=20,  # Number of horizontal blocks per page
    block_overlap=0.2,  # Overlap between blocks (0.0-1.0)
    embedding_model_name="all-MiniLM-L6-v2",  # Model for text embeddings
    aws_credentials=None  # Optional AWS credentials for S3 access
)

OCRSnipRAGEngine

Engine that uses Tesseract OCR for text extraction.

engine = create_engine(
    "ocr",
    num_slices=10,  # Number of horizontal slices per page
    tesseract_cmd="/path/to/tesseract",  # Path to Tesseract executable
    embedding_model_name="all-MiniLM-L6-v2",  # Model for text embeddings
    aws_credentials=None  # Optional AWS credentials for S3 access
)

Common Methods

  • process_pdf(pdf_path, document_id): Process a local PDF file
  • process_document_from_s3(s3_uri, document_id): Process a PDF from S3
  • search(query, top_k=5, filter_metadata=None): Search for text matches
  • search_with_snippets(query, top_k=5, filter_metadata=None, include_snippets=True, snippet_padding=None): Search with image snippets
  • get_image_snippet(result_idx, padding=None): Get an image snippet for a specific result
  • clear_index(): Clear the search index and stored documents

Use Cases

SnipRAG is particularly valuable for:

  • Financial Document Analysis: Extract specific items from invoices or financial statements
  • Legal Document Review: Find and visualize specific clauses in contracts
  • Technical Documentation: Locate diagrams, tables, and code snippets
  • Research Papers: Find equations, figures, and important text
  • Medical Records: Identify specific sections, charts, or results
  • Scanned Documents: Process historical or legacy documents with OCR capabilities

Requirements

  • Python 3.8+
  • Required packages:
    • pymupdf (PyMuPDF)
    • sentence-transformers
    • faiss-cpu
    • pillow
    • numpy
    • boto3 (for S3 integration)
    • langchain (for text splitting)
    • pytesseract (for OCR support)
    • matplotlib (for visualization, optional)

Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

  1. Fork the repository
  2. Create your feature branch (git checkout -b feature/AmazingFeature)
  3. Commit your changes (git commit -m 'Add some AmazingFeature')
  4. Push to the branch (git push origin feature/AmazingFeature)
  5. Open a Pull Request

License

This project is licensed under the MIT License - see the LICENSE file for details.

Acknowledgements

  • This project was inspired by the need for more precise visual context in RAG systems
  • Thanks to the developers of PyMuPDF, sentence-transformers, FAISS, and Tesseract for their excellent libraries

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

sniprag-0.2.0.tar.gz (18.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

sniprag-0.2.0-py3-none-any.whl (19.3 kB view details)

Uploaded Python 3

File details

Details for the file sniprag-0.2.0.tar.gz.

File metadata

  • Download URL: sniprag-0.2.0.tar.gz
  • Upload date:
  • Size: 18.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.4

File hashes

Hashes for sniprag-0.2.0.tar.gz
Algorithm Hash digest
SHA256 9dea7e8e75f3d66e7051426444d4e3c69107f0c42acd9f9d3ef7bf3096e3df9c
MD5 a715004df0f34e52fa28e324e1e8b792
BLAKE2b-256 5db5e0560634d3ecdae9180e40701ecd9c4ed4886f33d5c8ba4bd30c3d0f9636

See more details on using hashes here.

File details

Details for the file sniprag-0.2.0-py3-none-any.whl.

File metadata

  • Download URL: sniprag-0.2.0-py3-none-any.whl
  • Upload date:
  • Size: 19.3 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.4

File hashes

Hashes for sniprag-0.2.0-py3-none-any.whl
Algorithm Hash digest
SHA256 882b20c7572aaad7f6fe79ac4f704fbda7866d17852aed1ff0ea2527d48c50b8
MD5 69dc4bfdc050716693bc38cbc15a2d9b
BLAKE2b-256 35a785fab7958d44ef5a4333d1d6cdf2078bc4f8ddf93714a9e466e74332e33a

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page