langchain-xparse
LangChain integration with xParse Parse API for intelligent document parsing. Converts unstructured documents (PDF, images, Word, Excel, PPT, etc.) into AI-friendly structured data (JSON, Markdown) with rich metadata.
Installation
From PyPI:
pip install langchain-xparse
Configuration
Set your TextIn credentials (from Textin Workspace):
export XPARSE_APP_ID="your-app-id"
export XPARSE_SECRET_CODE="your-secret-code"
Or pass them when creating the loader:
loader = XParseLoader(
file_path="doc.pdf",
app_id="your-app-id",
secret_code="your-secret-code",
)
Usage
Basic Usage
from langchain_xparse import XParseLoader
loader = XParseLoader(file_path="example.pdf")
docs = loader.load()
print(docs[0].page_content[:200])
print(docs[0].metadata) # source, category, element_id, filename, page_number
Lazy Load
for doc in loader.lazy_load():
# process each document
print(doc.page_content[:100])
Async Load
async for doc in loader.alazy_load():
# process each document asynchronously
print(doc.page_content[:100])
Custom Parse Configuration
Customize parsing behavior using the config parameter. See Parse Config Documentation for details.
loader = XParseLoader(
file_path="doc.pdf",
config={
"document": {
"password": "pdf-password" # For encrypted PDFs
},
"capabilities": {
"include_hierarchy": True, # Include parent-child relationships
"include_inline_objects": True, # Extract formulas, handwriting, etc.
"include_table_structure": True, # Detailed table structure
"include_char_details": True, # Character-level details
"include_image_data": True, # Image URLs and data
"pages": True, # Page metadata
"title_tree": True, # Document outline/TOC
"table_view": "html" # Table format: "html" or "markdown"
},
"scope": {
"page_range": "1-10" # Process specific pages
},
"config": {
"force_engine": "textin", # Engine selection (expert mode)
"engine_params": {
"formula_level": 0,
"image_output_type": "url"
}
}
}
)
docs = loader.load()
Multiple Files
loader = XParseLoader(file_path=["a.pdf", "b.pdf", "c.docx"])
for doc in loader.lazy_load():
print(f"{doc.metadata.get('source')}: {doc.page_content[:50]}")
File-like Object
When passing a file-like object instead of a path, you must set metadata_filename:
with open("doc.pdf", "rb") as f:
loader = XParseLoader(file=f, metadata_filename="doc.pdf")
docs = loader.load()
Document Metadata
Each loaded document includes rich metadata:
source: File path or filenamecategory: Element type (Title, NarrativeText, Table, Image, Formula, etc.)element_id: Unique element identifierfilename: Original filenamepage_number: Page number (if available)parent_id: Parent element ID (withinclude_hierarchy)children_ids: Child element IDs (withinclude_hierarchy)- Additional element-specific metadata
References
- xParse Parse API - API endpoint documentation
- Parse Config - Configuration parameters
- Parse Response - Response structure and fields
Metadata
Release files for langchain-xparse 1.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| langchain_xparse-1.2.0.tar.gz | 6.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| langchain_xparse-1.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 14.4 kB
Release files / langchain_xparse-1.2.0.tar.gz
| Download URL | langchain_xparse-1.2.0.tar.gz |
|---|---|
| Size | 6.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
4dd99c93f1dcb004b64e01a9eac40632cd5149c98ce21c0e3ab05972b950fddc
|
|
BLAKE2b-256 checksum How to use checksums |
c85fdf3bd5bc3651e14f00048c37dd8d9189e2a0bd89fb1f7d6ed8be5d09baad
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.10.19
|
Release files / langchain_xparse-1.2.0-py3-none-any.whl
| Download URL | langchain_xparse-1.2.0-py3-none-any.whl |
|---|---|
| Size | 7.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
83031d7e7fb4eeb49fa7b4e17e1c0e2418f1ce7abb3e4560148dc0be98f4ad83
|
|
BLAKE2b-256 checksum How to use checksums |
2d117010217e589958fa78987995c5690d62d67d4cdc76d33a1d87bbd432ee76
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.10.19
|