Skip to main content

langchain-xparse

LangChain integration with xParse Parse API for intelligent document parsing. Converts unstructured documents (PDF, images, Word, Excel, PPT, etc.) into AI-friendly structured data (JSON, Markdown) with rich metadata.

Installation

From PyPI:

pip install langchain-xparse

Configuration

Set your TextIn credentials (from Textin Workspace):

export XPARSE_APP_ID="your-app-id"
export XPARSE_SECRET_CODE="your-secret-code"

Or pass them when creating the loader:

loader = XParseLoader(
    file_path="doc.pdf",
    app_id="your-app-id",
    secret_code="your-secret-code",
)

Usage

Basic Usage

from langchain_xparse import XParseLoader

loader = XParseLoader(file_path="example.pdf")
docs = loader.load()
print(docs[0].page_content[:200])
print(docs[0].metadata)  # source, category, element_id, filename, page_number

Lazy Load

for doc in loader.lazy_load():
    # process each document
    print(doc.page_content[:100])

Async Load

async for doc in loader.alazy_load():
    # process each document asynchronously
    print(doc.page_content[:100])

Custom Parse Configuration

Customize parsing behavior using the config parameter. See Parse Config Documentation for details.

loader = XParseLoader(
    file_path="doc.pdf",
    config={
        "document": {
            "password": "pdf-password"  # For encrypted PDFs
        },
        "capabilities": {
            "include_hierarchy": True,         # Include parent-child relationships
            "include_inline_objects": True,    # Extract formulas, handwriting, etc.
            "include_table_structure": True,   # Detailed table structure
            "include_char_details": True,      # Character-level details
            "include_image_data": True,        # Image URLs and data
            "pages": True,                     # Page metadata
            "title_tree": True,                # Document outline/TOC
            "table_view": "html"               # Table format: "html" or "markdown"
        },
        "scope": {
            "page_range": "1-10"               # Process specific pages
        },
        "config": {
            "force_engine": "textin",          # Engine selection (expert mode)
            "engine_params": {
                "formula_level": 0,
                "image_output_type": "url"
            }
        }
    }
)
docs = loader.load()

Multiple Files

loader = XParseLoader(file_path=["a.pdf", "b.pdf", "c.docx"])
for doc in loader.lazy_load():
    print(f"{doc.metadata.get('source')}: {doc.page_content[:50]}")

File-like Object

When passing a file-like object instead of a path, you must set metadata_filename:

with open("doc.pdf", "rb") as f:
    loader = XParseLoader(file=f, metadata_filename="doc.pdf")
    docs = loader.load()

Document Metadata

Each loaded document includes rich metadata:

  • source: File path or filename
  • category: Element type (Title, NarrativeText, Table, Image, Formula, etc.)
  • element_id: Unique element identifier
  • filename: Original filename
  • page_number: Page number (if available)
  • parent_id: Parent element ID (with include_hierarchy)
  • children_ids: Child element IDs (with include_hierarchy)
  • Additional element-specific metadata

References

Metadata

Release files for langchain-xparse 1.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for langchain-xparse 1.2.0
File Size Uploaded
langchain_xparse-1.2.0.tar.gz 6.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for langchain-xparse 1.2.0
File Interpreter ABI Platform
langchain_xparse-1.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 14.4 kB

Release files / langchain_xparse-1.2.0.tar.gz

Download URL langchain_xparse-1.2.0.tar.gz
Size 6.9 kB
Tags Source
SHA-256 checksum
How to use checksums
4dd99c93f1dcb004b64e01a9eac40632cd5149c98ce21c0e3ab05972b950fddc
BLAKE2b-256 checksum
How to use checksums
c85fdf3bd5bc3651e14f00048c37dd8d9189e2a0bd89fb1f7d6ed8be5d09baad
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.10.19

Release files / langchain_xparse-1.2.0-py3-none-any.whl

Download URL langchain_xparse-1.2.0-py3-none-any.whl
Size 7.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
83031d7e7fb4eeb49fa7b4e17e1c0e2418f1ce7abb3e4560148dc0be98f4ad83
BLAKE2b-256 checksum
How to use checksums
2d117010217e589958fa78987995c5690d62d67d4cdc76d33a1d87bbd432ee76
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.10.19

Release history Release notifications | RSS feed

This release

1.2.0 This release

2 release files

1.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page