Skip to main content

structx

Advanced structured data extraction from any document using LLMs with multimodal support.

Documentation PyPI GitHub Actions

See the project roadmap for planned portable business rules and semantic validation work.

structx is a powerful Python library for extracting structured data from text, tables, and documents using Large Language Models (LLMs). It passes existing PDFs directly to vision-capable models; the optional docs extra converts other document formats to PDF first.

🔔 Package rename notice (PyPI)

The PyPI distribution has been renamed from structx-llm to structx (September 2025).

  • Imports are unchanged: continue using import structx
  • Document processing now lives in the optional docs extra
  • Please update your environments and requirement files to use the new name

Upgrade commands:

pip uninstall -y structx-llm
pip install -U structx

If you previously pinned structx-llm in requirements or lock files, replace it with structx. Install structx[docs] for non-PDF document conversion.

✨ Key Features

🎯 Advanced Document Processing

  • � Multimodal PDF Pipeline: Passes PDFs directly to vision-capable models and converts supported non-PDF documents to PDF
  • 🖼️ Vision-Enabled Extraction: Native instructor multimodal support for PDFs and images
  • 🔄 Smart Format Detection: Automatic processing mode selection for best results
  • 📊 Flexible File Support: CSV, Excel, JSON, Parquet, raw text, and existing PDFs in the base install; DOCX, PPTX, images, and more via structx[docs]

🚀 Intelligent Data Extraction

  • 🔄 Dynamic Model Generation: Create type-safe Pydantic models from natural language queries
  • 🎯 Automatic Schema Inference: Intelligent schema generation and refinement
  • 📊 Complex Data Structures: Support for nested and hierarchical data
  • 🔄 Natural Language Refinement: Improve models with conversational instructions

Performance & Reliability

  • 🚀 High-Performance Processing: Threaded sync and native async row requests
  • 🔄 Robust Error Handling: Automatic retry mechanism with exponential backoff
  • 📈 Token Usage Tracking: Detailed step-by-step metrics for cost monitoring
  • Flexible Configuration: Model settings from arguments, YAML, environment variables, dotenv files, and secrets through Pydantic Settings
  • 🔌 Multiple LLM Providers: Support through litellm integration

Installation

pip install structx

For converting DOCX, PowerPoint, OpenDocument, markup, image, and other non-PDF document formats:

pip install "structx[docs]"

🔧 What The Package Provides

  • Structured readers for CSV, Excel, JSON, Parquet, and Feather
  • Instructor multimodal vision support
  • Optional Docling document parsing with CPU-only PyTorch resolution for uv on Linux
  • Optional WeasyPrint PDF rendering for non-PDF document formats

Quick Start

Basic Text Extraction

from structx import Extractor

# Initialize extractor
extractor = Extractor.from_litellm(
    model="gpt-4o",
    api_key="your-api-key",
    max_retries=3,      # Automatically retry on transient errors
    min_wait=1,         # Start with 1 second wait
    max_wait=10         # Maximum 10 seconds between retries
)

# Extract from text
result = extractor.extract(
    data="System check on 2024-01-15 detected high CPU usage (92%) on server-01.",
    query="extract incident date and details"
)

# Access results
print(f"Successful rows: {result.success_count}")
print(result.data[0].model_dump_json(indent=2))

📄 Document Processing with Multimodal Support

Install structx[docs] before using non-PDF document formats. Existing PDFs can be passed directly through the multimodal path with the base installation.

# Process a PDF invoice through the multimodal pipeline
result = extractor.extract(
    data="scripts/example_input/S0305SampleInvoice.pdf",
    query="extract the invoice number, total amount, and line items"
)

# Convert a DOCX contract and process with multimodal support
result = extractor.extract(
    data="scripts/example_input/free-consultancy-agreement.docx",
    query="extract parties, effective date, and payment terms"
)

📊 Token Usage Monitoring

# Check token usage for cost monitoring
usage = result.usage
if usage:
    print(f"Total tokens: {usage.total_tokens}")
    for step, calls in usage.steps.items():
        print(step.value, [call.total_tokens for call in calls])

🚀 Why Multimodal PDF Processing?

The innovative multimodal approach provides significant advantages over traditional text-based extraction:

  • 📄 Context Preservation: Full document layout and structure are maintained
  • 🎯 Higher Accuracy: Vision models can interpret tables, charts, and complex layouts
  • 🔄 No Chunking Issues: Eliminates problems with information split across chunks
  • 📊 Universal Format: Existing PDFs are passed through directly; supported non-PDF documents become processable through PDF conversion
  • 🖼️ Visual Understanding: Handles documents with visual elements, formatting, and structure

📚 Documentation

For comprehensive documentation, examples, and guides, visit our documentation site.

Examples

Check out our example gallery for real-world use cases,

📁 Supported File Formats

📊 Structured Data (Direct Processing)

  • CSV: Comma-separated values with custom delimiters
  • Excel: .xlsx/.xls with sheet selection and custom options
  • JSON: JavaScript Object Notation with nested support
  • Parquet: Columnar storage format for large datasets
  • Feather: Fast binary format for data frames

📄 Unstructured Documents (Multimodal Pipeline)

Format Extensions Processing Method Quality
PDF .pdf PDF → Multimodal ⭐⭐⭐⭐⭐
Word .docx, .doc Docling → HTML → PDF → Multimodal ⭐⭐⭐⭐⭐
PowerPoint .pptx, .ppt Docling → HTML → PDF → Multimodal ⭐⭐⭐⭐
Text .txt, .md, .py, .log, .xml, .html Docling → HTML → PDF → Multimodal ⭐⭐⭐⭐

🔄 Processing Pipeline

  • PDF passthrough: Existing PDFs are sent directly to multimodal extraction
  • Docling parsing: Reads non-PDF document-like inputs into a structured document model
  • WeasyPrint rendering: Converts Docling HTML to a temporary PDF for non-PDF inputs
  • Multimodal extraction: Sends the rendered PDF to instructor's multimodal API

Contributing

Contributions are welcome! Please read our Contributing Guidelines for details.

License

This project is licensed under the MIT License - see the LICENSE file for details.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

structx-0.6.3.tar.gz (54.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

structx-0.6.3-py3-none-any.whl (45.4 kB view details)

Uploaded Python 3

File details

Details for the file structx-0.6.3.tar.gz.

File metadata

  • Download URL: structx-0.6.3.tar.gz
  • Upload date:
  • Size: 54.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for structx-0.6.3.tar.gz
Algorithm Hash digest
SHA256 d109d7e1c7af183f7253e3b60b04e34ee094bf3aee81beedd871d90dbc3812bb
MD5 e6fb42177658b8fb7c502dd48e2d83b8
BLAKE2b-256 4f90e65fd3bd98c09a7c30565104ca4b882e12879be2669754f7f7ac2fd2aae5

See more details on using hashes here.

Provenance

The following attestation bundles were made for structx-0.6.3.tar.gz:

Publisher: publish.yml on Blacksuan19/structx

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file structx-0.6.3-py3-none-any.whl.

File metadata

  • Download URL: structx-0.6.3-py3-none-any.whl
  • Upload date:
  • Size: 45.4 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for structx-0.6.3-py3-none-any.whl
Algorithm Hash digest
SHA256 62b74df4a1ecbef1233c53fb9603281b20f42388513b57e7aad24a633831daba
MD5 762b01a08543dcffba30f1a155b57074
BLAKE2b-256 160f14c44793faa8a7119a3d2c4b2cdad73bb155ad273ec9993d64e14f8ddf0f

See more details on using hashes here.

Provenance

The following attestation bundles were made for structx-0.6.3-py3-none-any.whl:

Publisher: publish.yml on Blacksuan19/structx

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

This release

0.6.3 This release

2 files

0.6.2

2 files

0.6.1

2 files

0.6.0

2 files

0.5.1

2 files

0.5.0

2 files

0.4.11

2 files

0.4.10

2 files

0.4.8

2 files

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page