Retrieval Augmented Generation with Image Snippets from PDFs
Project description
SnipRAG: Retrieval Augmented Generation with Image Snippets
SnipRAG is a specialized Retrieval Augmented Generation (RAG) system that not only finds semantically relevant text in PDF documents but also extracts precise image snippets from the areas containing the matching text.
Key Features
- Semantic PDF Search: Find information in PDF documents using natural language queries
- Image Snippet Extraction: Get visual context from the exact regions containing relevant information
- Precise Coordinate Mapping: Maps text matches to their exact visual location in the document
- Customizable Snippet Size: Adjust padding around text regions to control snippet size
- S3 Integration: Process documents stored in Amazon S3
- Flexible Filtering: Filter search results by document, page, or custom metadata
Installation
From PyPI
⚠️ Note: This package is not yet available on PyPI. Please use the source installation method below.
In the future, once published to PyPI:
pip install sniprag
From Source
git clone https://github.com/ishandikshit/SnipRAG.git
cd SnipRAG
pip install -e .
For visualization support (recommended for demos):
pip install -e ".[viz]"
Quick Start
from sniprag import SnipRAGEngine
# Initialize the engine
engine = SnipRAGEngine()
# Process a PDF document
engine.process_pdf("path/to/document.pdf", "document-id")
# Search with image snippets
results = engine.search_with_snippets("your search query", top_k=3)
# Access results
for result in results:
print(f"Text: {result['text']}")
print(f"Page: {result['metadata']['page_number']}")
print(f"Score: {result['score']}")
# The image snippet is available as base64 data that can be displayed or saved
if "image_data" in result:
image_base64 = result["image_data"]
# Use this to display or save the image
Example Snippets
Here are some examples of SnipRAG in action, showing how it extracts image snippets from PDF documents based on semantic search queries:
Financial Data Extraction
Query: "What is the total revenue?"
SnipRAG extracts the exact region containing revenue information, providing visual context alongside the text match.
Technical Specification Extraction
Query: "How does the system implement semantic search?"
When searching for technical details, SnipRAG locates and extracts the relevant section, preserving formatting and visual context.
Document Navigation
Query: "Show me the introduction section"
SnipRAG can help navigate to specific sections of a document based on semantic understanding of the content.
Note: To generate these snippets yourself, run the basic demo with a sample PDF as shown in the Demos section below.
Demos
SnipRAG includes two demo applications:
Basic Demo
Process a local PDF file and search for information with image snippets:
python examples/basic_demo.py --pdf path/to/document.pdf
S3 Demo
Process a PDF stored in Amazon S3:
python examples/s3_demo.py --s3-uri s3://bucket/path/to/document.pdf --aws-profile your-profile
Jupyter Notebook Example
For those working in Jupyter environments, there's also a notebook example available:
# View the notebook example
cat examples/example_notebook.md
This markdown file contains code snippets you can use in a Jupyter notebook to process PDFs and visualize search results with image snippets.
How It Works
SnipRAG combines semantic search with coordinate mapping to provide visual context:
-
PDF Processing:
- Extracts text blocks with their coordinates
- Renders page images at high resolution
- Creates text embeddings for semantic search
-
Search Process:
- User submits a natural language query
- System finds semantically similar text using embeddings
- For each match, it identifies the exact location in the PDF
- It extracts an image snippet from that location
-
Result Delivery:
- Returns the matching text
- Provides a visual snippet of the area containing the text
- Includes metadata (page number, coordinates, etc.)
API Reference
SnipRAGEngine
Main class for the SnipRAG engine.
engine = SnipRAGEngine(
embedding_model_name="all-MiniLM-L6-v2", # Model for text embeddings
aws_credentials=None # Optional AWS credentials for S3 access
)
Methods
process_pdf(pdf_path, document_id): Process a local PDF fileprocess_document_from_s3(s3_uri, document_id): Process a PDF from S3search(query, top_k=5, filter_metadata=None): Search for text matchessearch_with_snippets(query, top_k=5, filter_metadata=None, include_snippets=True, snippet_padding=None): Search with image snippetsget_image_snippet(result_idx, padding=None): Get an image snippet for a specific resultclear_index(): Clear the search index and stored documents
Use Cases
SnipRAG is particularly valuable for:
- Financial Document Analysis: Extract specific items from invoices or financial statements
- Legal Document Review: Find and visualize specific clauses in contracts
- Technical Documentation: Locate diagrams, tables, and code snippets
- Research Papers: Find equations, figures, and important text
- Medical Records: Identify specific sections, charts, or results
Requirements
- Python 3.8+
- Required packages:
- pymupdf (PyMuPDF)
- sentence-transformers
- faiss-cpu
- pillow
- numpy
- boto3 (for S3 integration)
- langchain (for text splitting)
- matplotlib (for visualization, optional)
Contributing
Contributions are welcome! Please feel free to submit a Pull Request.
- Fork the repository
- Create your feature branch (
git checkout -b feature/AmazingFeature) - Commit your changes (
git commit -m 'Add some AmazingFeature') - Push to the branch (
git push origin feature/AmazingFeature) - Open a Pull Request
License
This project is licensed under the MIT License - see the LICENSE file for details.
Acknowledgements
- This project was inspired by the need for more precise visual context in RAG systems
- Thanks to the developers of PyMuPDF, sentence-transformers, and FAISS for their excellent libraries
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file sniprag-0.1.0.tar.gz.
File metadata
- Download URL: sniprag-0.1.0.tar.gz
- Upload date:
- Size: 13.5 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.1.0 CPython/3.13.1
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
1cab25c64e5838977329b5a50b2de28f5cbfd4528c4ae3b524749beee87a4bb6
|
|
| MD5 |
e3a3867eec3e28be53da20f6968d3f40
|
|
| BLAKE2b-256 |
23aaea13774f6e8dee31cf7443ac374a1feb8c72a3c1d1961cb1074ae7e6c501
|
File details
Details for the file sniprag-0.1.0-py3-none-any.whl.
File metadata
- Download URL: sniprag-0.1.0-py3-none-any.whl
- Upload date:
- Size: 10.0 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.1.0 CPython/3.13.1
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
a0808d018e054dfc0b246a4347558386f11236f98d9b6bea8437e14007738671
|
|
| MD5 |
1f35707b9def8cdaf02161f7956c3bea
|
|
| BLAKE2b-256 |
37467cb6a0c64fc10877909b13cea9488a90cb71a386eda976969ea11187fdd5
|