Skip to main content

Xberg

langchain-xberg

A LangChain document loader backed by Xberg. XbergLoader extracts text and metadata from 101 formats — running OCR where needed — and returns LangChain Document objects. Extraction is async at the core; multiple sources go through Xberg's extract_batch in a single native call, so concurrency happens Rust-side.

Install

pip install langchain-xberg

Requires Python 3.10+.

Load

Pass a path, a list of paths, a directory, or raw bytes. One source becomes one Document.

from langchain_xberg import XbergLoader

# Single file
docs = XbergLoader(file_path="report.pdf").load()
print(docs[0].page_content)           # extracted markdown
print(docs[0].metadata["title"])      # source, mime_type, title, authors, detected_languages, page_count, ...

# Multiple files — one batched extraction
docs = XbergLoader(file_path=["report.pdf", "notes.docx"]).load()

# A directory with a glob
docs = XbergLoader(file_path="./corpus/", glob="**/*.pdf").load()

# Raw bytes (mime_type required)
docs = XbergLoader(data=raw_bytes, mime_type="application/pdf").load()

Chunk for retrieval

Enable Xberg's native chunking to emit one Document per chunk, sized for embedding. Each chunk carries chunk_index, total_chunks, heading_path, page, and token_count in its metadata.

from langchain_xberg import XbergLoader
from xberg import ChunkingConfig, ExtractionConfig

config = ExtractionConfig(chunking=ChunkingConfig(max_characters=1000, overlap=200))
docs = XbergLoader(file_path="report.pdf", config=config).load()  # one Document per chunk

To split by page instead, pass pages=PageConfig(extract_pages=True); each Document then gets a 0-indexed page.

Configure extraction

Pass any Xberg ExtractionConfig to control OCR, output format, and batch concurrency.

from xberg import ExtractionConfig, OcrConfig

config = ExtractionConfig(
    output_format="markdown",
    ocr=OcrConfig(backend="tesseract"),
    force_ocr=True,
    max_concurrent_extractions=8,
)
docs = XbergLoader(file_path="./corpus/", config=config).load()

Async

Inside an event loop, use the async API — await loader.aload() or async for doc in loader.alazy_load(). These use Xberg's native async extraction end to end. The synchronous load / lazy_load bridge to it and must not run inside a running loop.

loader = XbergLoader(file_path="report.pdf")
docs = await loader.aload()

Errors

A per-input failure raises xberg.XbergError with the offending source and message. In batch mode the failure comes from ExtractionResult.errors; a single load surfaces the raised exception directly.

For the full API, see the Xberg documentation.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

langchain_xberg-1.0.8.tar.gz (11.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

langchain_xberg-1.0.8-py3-none-any.whl (8.8 kB view details)

Uploaded Python 3

File details

Details for the file langchain_xberg-1.0.8.tar.gz.

File metadata

  • Download URL: langchain_xberg-1.0.8.tar.gz
  • Upload date:
  • Size: 11.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: uv/0.12.1 {"installer":{"name":"uv","version":"0.12.1","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

File hashes

Hashes for langchain_xberg-1.0.8.tar.gz
Algorithm Hash digest
SHA256 4934e22e6d13e7cb81864e5f8057ba5f711f34fd1dd088bab7fa95b045bfbb4c
MD5 7b2e37774f05812acbf3e55a82f36849
BLAKE2b-256 0bbb2d703b3afdb25500aa4eb3640e583b5ce6a07233abf478b9eea199299fc0

See more details on using hashes here.

File details

Details for the file langchain_xberg-1.0.8-py3-none-any.whl.

File metadata

  • Download URL: langchain_xberg-1.0.8-py3-none-any.whl
  • Upload date:
  • Size: 8.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: uv/0.12.1 {"installer":{"name":"uv","version":"0.12.1","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

File hashes

Hashes for langchain_xberg-1.0.8-py3-none-any.whl
Algorithm Hash digest
SHA256 45a86e3e41809c83ffa598d8ba17db2a97eba1435a75665016836ac600ef3c2a
MD5 e3a643ccb99b5b35ed6a9969420c8cb3
BLAKE2b-256 3ac377d6c9838c2209fd23e3dce51014202f576b42bab08d9d4c61ff345d2fe4

See more details on using hashes here.

Release history Release notifications | RSS feed

1.2.1

2 files

1.1.5

2 files

1.1.3

2 files

1.1.2

2 files

1.1.1

2 files

1.1.0

2 files

1.0.14

2 files

1.0.12

2 files

1.0.11

2 files

1.0.10

2 files

1.0.9

2 files

This release

1.0.8 This release

2 files

1.0.7

2 files

1.0.5

2 files

1.0.3

2 files

1.0.1

2 files

1.0.0

2 files

0.0.1

1 file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page