Skip to main content

Document structuring library for RAG and vector databases

Project description

pydocstruct

PyPI version License: MIT

pydocstruct is a Python library designed to convert various file formats (PDF, DOCX, Excel, Markdown, CSV, JSON, etc.) into structured data suitable for RAG (Retrieval-Augmented Generation) systems and vector databases.

It provides features optimized for RAG, such as token-based chunking, Markdown table conversion, and PII redaction.

Features

  • 🌟 Unified Interface: Support for all formats via a single pydocstruct.load() function.
  • 📄 Diverse Formats: PDF, Excel (XLSX), Word (DOCX), Markdown, Text, HTML, XML, JSON, CSV.
  • 🤖 RAG Optimization:
    • Advanced Chunking: Supports Recursive Character Splitter and Token-based (TokenChunker) splitting.
    • Enhanced Representation: Extracts CSV/Excel as Markdown tables; preserves HTML structure.
    • Advanced PDF Handling: Layout analysis and OCR support for scanned PDFs.
  • 🛡️ Data Cleaning:
    • PII Redaction: Automatically masks personal information like emails and phone numbers.
    • Noise Reduction: Removes HTML headers, footers, advertisements, etc.

Installation

pip install pydocstruct

Installing Dependencies

To use all features (OCR, Excel, table conversion, etc.), we recommend installing the additional dependencies:

pip install "pydocstruct[all]"

(Note: If installing manually: pandas, openpyxl, tabulate, pypdf, python-docx, beautifulsoup4, markdownify, tiktoken, pdf2image, pytesseract)

⚠️ Requirements for OCR: To use the OCR feature (use_ocr=True), you must separately install the following system tools and ensure they are in your PATH:

  • Tesseract-OCR
  • Poppler (for pdf2image)

Quick Start

Basic Usage

from pydocstruct import load

# Load a PDF file
documents = load("paper.pdf")

for doc in documents:
    print(f"File: {doc.source}")
    print(f"Page: {doc.page_number}")
    print(f"Content: {doc.content[:100]}...")

The load function returns a list[Document].

Supported Formats & Details

Format Extension Description Key Options
PDF .pdf Documents per page. Supports OCR/Layout analysis. use_ocr=True, use_layout=True
Excel .xlsx Load row-by-row or as Markdown table. output_format="markdown"
CSV .csv Load row-by-row or as Markdown table. output_format="markdown"
Word .docx Loads entire document. -
HTML .html Extracts text or preserves structure as Markdown. preserve_structure=True
Markdown .md Supports splitting by headers. split_by_headers=True
Text .txt Loads as UTF-8 text. -

Advanced Usage

1. Markdown Table Conversion (CSV/Excel)

For RAG, Markdown tables are often better understood by LLMs than flattened key-value pairs.

# Load Excel as a Markdown table
docs = load("data.xlsx", output_format="markdown")
print(docs[0].content)
# Output Example:
# | Product | Price |
# |:--------|------:|
# | Apple   |   100 |

2. PDF OCR and Layout Analysis

Handles scanned PDFs or complex multi-column layouts.

# Extract while preserving layout (better for multi-column)
docs = load("paper.pdf", use_layout=True)

# Run OCR on pages with no text (Requires Tesseract/Poppler)
docs = load("scanned.pdf", use_ocr=True, ocr_lang="eng")

3. Advanced Chunking

Split text intelligently to fit LLM context windows.

from pydocstruct.core.chunker import RecursiveCharacterChunker, TokenChunker

# Recursive splitting (Try Paragraph -> Sentence -> Word)
recursive_chunker = RecursiveCharacterChunker(chunk_size=1000, chunk_overlap=200)
chunked_docs = recursive_chunker.split_documents(documents)

# Token-based splitting (using tiktoken, optimized for GPT models)
token_chunker = TokenChunker(chunk_size=500, model_name="gpt-4")
chunked_docs = token_chunker.split_documents(documents)

4. Data Cleaning (Processors)

Remove PII or HTML noise.

from pydocstruct.processors import PiiRedactor, HtmlNoiseCleaner

text = "Contact: 090-1234-5678, email@example.com"

# Redact PII
cleaned_text = PiiRedactor.redact(text)
print(cleaned_text) 
# -> "Contact: [REDACTED], [REDACTED]"

# Remove HTML noise (headers, footers, ads)
clean_html_text = HtmlNoiseCleaner.clean(html_content)

License

MIT License

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

pydocstruct-0.1.6.tar.gz (19.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

pydocstruct-0.1.6-py3-none-any.whl (24.1 kB view details)

Uploaded Python 3

File details

Details for the file pydocstruct-0.1.6.tar.gz.

File metadata

  • Download URL: pydocstruct-0.1.6.tar.gz
  • Upload date:
  • Size: 19.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.10.11

File hashes

Hashes for pydocstruct-0.1.6.tar.gz
Algorithm Hash digest
SHA256 1218cc3c322972e04b7a32b9a51975120672411ae11d3091f5a56c957a5a6866
MD5 441987c3038bd7ce6991469b0c85208d
BLAKE2b-256 efe3bab0447505e47c767d740bdfbf4342db8ac2ef4c5e774eb2f38c7c0b273d

See more details on using hashes here.

File details

Details for the file pydocstruct-0.1.6-py3-none-any.whl.

File metadata

  • Download URL: pydocstruct-0.1.6-py3-none-any.whl
  • Upload date:
  • Size: 24.1 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.10.11

File hashes

Hashes for pydocstruct-0.1.6-py3-none-any.whl
Algorithm Hash digest
SHA256 b1151260677d5c7d08ab7a982961b0dc67078944214f12ab46eaa08f7065caef
MD5 f0a01a83f228beddc4873f27bb89f587
BLAKE2b-256 ecc4930edc31e729975dbd775e8728f00a6a9fb96bd1988da29bb0a809c0dccc

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page