Skip to main content

Xberg

langchain-xberg

A LangChain document loader backed by Xberg. XbergLoader extracts text and metadata from 107 formats — running OCR where needed — and returns LangChain Document objects. Extraction is async at the core; multiple sources go through Xberg's extract_batch in a single native call, so concurrency happens Rust-side.

Install

pip install langchain-xberg

Requires Python 3.10+.

Load

Pass a path, a list of paths, a directory, or raw bytes. One source becomes one Document.

from langchain_xberg import XbergLoader

# Single file
docs = XbergLoader(file_path="report.pdf").load()
print(docs[0].page_content)           # extracted markdown
print(docs[0].metadata["title"])      # source, mime_type, title, authors, detected_languages, page_count, ...

# Multiple files — one batched extraction
docs = XbergLoader(file_path=["report.pdf", "notes.docx"]).load()

# A directory with a glob
docs = XbergLoader(file_path="./corpus/", glob="**/*.pdf").load()

# Raw bytes (mime_type required)
docs = XbergLoader(data=raw_bytes, mime_type="application/pdf").load()

Chunk for retrieval

Enable Xberg's native chunking to emit one Document per chunk, sized for embedding. Each chunk carries chunk_index, total_chunks, heading_path, page, and token_count in its metadata.

from langchain_xberg import XbergLoader
from xberg import ChunkingConfig, ExtractionConfig

config = ExtractionConfig(chunking=ChunkingConfig(max_characters=1000, overlap=200))
docs = XbergLoader(file_path="report.pdf", config=config).load()  # one Document per chunk

To split by page instead, pass pages=PageConfig(extract_pages=True); each Document then gets a 0-indexed page.

Configure extraction

Pass any Xberg ExtractionConfig to control OCR, output format, and batch concurrency.

from xberg import ExtractionConfig, OcrConfig

config = ExtractionConfig(
    output_format="markdown",
    ocr=OcrConfig(backend="tesseract"),
    force_ocr=True,
    max_concurrent_extractions=8,
)
docs = XbergLoader(file_path="./corpus/", config=config).load()

Async

Inside an event loop, use the async API — await loader.aload() or async for doc in loader.alazy_load(). These use Xberg's native async extraction end to end. The synchronous load / lazy_load bridge to it and must not run inside a running loop.

loader = XbergLoader(file_path="report.pdf")
docs = await loader.aload()

Errors

A per-input failure raises xberg.XbergError with the offending source and message. In batch mode the failure comes from ExtractionResult.errors; a single load surfaces the raised exception directly.

For the full API, see the Xberg documentation.

Release files for langchain-xberg 1.2.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for langchain-xberg 1.2.2
File Size Uploaded
langchain_xberg-1.2.2.tar.gz 12.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for langchain-xberg 1.2.2
File Interpreter ABI Platform
langchain_xberg-1.2.2-py3-none-any.whl Python 3 none any Details

Total release size: 21.3 kB

Release files / langchain_xberg-1.2.2.tar.gz

Download URL langchain_xberg-1.2.2.tar.gz
Size 12.3 kB
Tags Source
SHA-256 checksum
How to use checksums
0c5fb8fccea6017fa3173801bcb6796809d83ccd10c38dcb95fb4e7ef6895a47
BLAKE2b-256 checksum
How to use checksums
def5a1ce65589ec6f64266e664f0ab1740bbed372fcd143583188010d407792d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.15 {"installer":{"name":"uv","version":"0.12.15","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release files / langchain_xberg-1.2.2-py3-none-any.whl

Download URL langchain_xberg-1.2.2-py3-none-any.whl
Size 9.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
8ecd0d5812b00ec20fc7081ddf369dd0419be0bcfe86c2570ff7bcf1827ed483
BLAKE2b-256 checksum
How to use checksums
dc122495ffcd93a9126f24ac717c3f3487c0634ea3166805dbee12d6d83bf4f4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.15 {"installer":{"name":"uv","version":"0.12.15","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page