LangChain document loader for Xberg — extract 88+ file formats into Documents with async, batching, chunking, and rich metadata
Project description
langchain-xberg
A LangChain document loader backed by Xberg. XbergLoader extracts text and metadata from 88+ formats — running OCR where needed — and returns LangChain Document objects. Extraction is async at the core; multiple sources go through Xberg's extract_batch in a single native call, so concurrency happens Rust-side.
Install
pip install langchain-xberg
Requires Python 3.10+.
Load
Pass a path, a list of paths, a directory, or raw bytes. One source becomes one Document.
from langchain_xberg import XbergLoader
# Single file
docs = XbergLoader(file_path="report.pdf").load()
print(docs[0].page_content) # extracted markdown
print(docs[0].metadata["title"]) # source, mime_type, title, authors, detected_languages, page_count, ...
# Multiple files — one batched extraction
docs = XbergLoader(file_path=["report.pdf", "notes.docx"]).load()
# A directory with a glob
docs = XbergLoader(file_path="./corpus/", glob="**/*.pdf").load()
# Raw bytes (mime_type required)
docs = XbergLoader(data=raw_bytes, mime_type="application/pdf").load()
Chunk for retrieval
Enable Xberg's native chunking to emit one Document per chunk, sized for embedding. Each chunk carries chunk_index, total_chunks, heading_path, page, and token_count in its metadata.
from langchain_xberg import XbergLoader
from xberg import ChunkingConfig, ExtractionConfig
config = ExtractionConfig(chunking=ChunkingConfig(max_characters=1000, overlap=200))
docs = XbergLoader(file_path="report.pdf", config=config).load() # one Document per chunk
To split by page instead, pass pages=PageConfig(extract_pages=True); each Document then gets a 0-indexed page.
Configure extraction
Pass any Xberg ExtractionConfig to control OCR, output format, and batch concurrency.
from xberg import ExtractionConfig, OcrConfig
config = ExtractionConfig(
output_format="markdown",
ocr=OcrConfig(backend="tesseract"),
force_ocr=True,
max_concurrent_extractions=8,
)
docs = XbergLoader(file_path="./corpus/", config=config).load()
Async
Inside an event loop, use the async API — await loader.aload() or async for doc in loader.alazy_load(). These use Xberg's native async extraction end to end. The synchronous load / lazy_load bridge to it and must not run inside a running loop.
loader = XbergLoader(file_path="report.pdf")
docs = await loader.aload()
Errors
A per-input failure raises xberg.XbergError with the offending source and message. In batch mode the failure comes from ExtractionResult.errors; a single load surfaces the raised exception directly.
For the full API, see the Xberg documentation.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distributions
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file langchain_xberg-1.0.0rc34-py3-none-any.whl.
File metadata
- Download URL: langchain_xberg-1.0.0rc34-py3-none-any.whl
- Upload date:
- Size: 8.8 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: uv/0.11.32 {"installer":{"name":"uv","version":"0.11.32","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
75a0295245d608de27fbc58cfc8eb86169821834c2f5a5631cc7ed19b2ac1a16
|
|
| MD5 |
41d22f4fb740c11c8b2dea86796b92d7
|
|
| BLAKE2b-256 |
eeddb9fb29fbb9787191d3cdda4365b74a4dcb5c10398419d9437a2a456e2dee
|