Skip to main content

LlamaIndex Readers Xberg

Xberg Banner

LlamaIndex reader for 101 file formats powered by xberg's Rust extraction engine.

Installation

pip install llama-index-readers-xberg

Requires Python ≥3.10, xberg>=1.0.0, and llama-index-core>=0.14.23,<0.15.

Features

  • 101 file formats -- PDF, DOCX, PPTX, XLSX, HTML, images, emails, archives, and more (full list)
  • Rich metadata -- quality scores, language detection, keywords, annotations
  • Native chunking -- xberg semantic chunks with heading path and page span, forwarded to XbergNodeParser
  • Element extraction -- structural elements for structure-aware RAG pipelines
  • Image extraction -- base64-encoded image data with position, format, and OCR metadata
  • Per-page splitting -- one Document per page for fine-grained retrieval
  • Batch processing -- multiple files in a single call
  • Raw bytes input -- extract from in-memory bytes with a MIME type
  • Native async -- true async via xberg's Rust tokio runtime
  • Error tolerance -- skip failed files with warnings, or raise on failure
  • Full serialization -- custom ExtractionConfig round-trips through to_dict()/from_dict() for pipeline caching

Usage

Basic Extraction

from llama_index.readers.xberg import XbergReader

reader = XbergReader()
documents = reader.load_data("report.pdf")

# Each document carries rich metadata
print(documents[0].metadata["file_name"])       # "report.pdf"
print(documents[0].metadata["file_type"])        # "application/pdf"
print(documents[0].metadata["total_pages"])      # 12
print(documents[0].metadata["quality_score"])    # 0.95
print(documents[0].metadata["detected_languages"])  # ["en"]

OCR Configuration

force_ocr is a top-level ExtractionConfig option. Language and backend are set on OcrConfig.

from xberg import ExtractionConfig, OcrConfig

reader = XbergReader(
    extraction_config=ExtractionConfig(
        force_ocr=True,
        ocr=OcrConfig(language="deu", backend="tesseract"),
    )
)
documents = reader.load_data("scanned.pdf")

Per-Page Splitting

PageConfig is nested inside ExtractionConfig. Each page becomes its own Document with a page_number metadata field.

from xberg import ExtractionConfig, PageConfig

reader = XbergReader(
    extraction_config=ExtractionConfig(
        pages=PageConfig(extract_pages=True),
    )
)
documents = reader.load_data("multi_page.pdf")  # One Document per page

for doc in documents:
    print(f"Page {doc.metadata['page_number']}: {doc.text[:80]}...")

Native Chunking

Enable xberg's native chunker to attach semantic chunks — each carrying its heading path and page span — under _xberg_chunks. Feed the documents into XbergNodeParser (from llama-index-node-parser-xberg) to turn each chunk into a node. This is the preferred input for RAG.

from xberg import ChunkingConfig, ExtractionConfig

reader = XbergReader(
    extraction_config=ExtractionConfig(
        chunking=ChunkingConfig(max_characters=1000, overlap=200, prepend_heading_context=True),
    )
)
documents = reader.load_data("report.pdf")

chunks = documents[0].metadata["_xberg_chunks"]

Element Extraction

Setting result_format="element_based" populates _xberg_elements in document metadata for structure-aware processing. XbergNodeParser uses these when no chunks are present.

from xberg import ExtractionConfig

reader = XbergReader(
    extraction_config=ExtractionConfig(result_format="element_based")
)
documents = reader.load_data("report.pdf")

# Structural elements available for downstream node parsers
elements = documents[0].metadata["_xberg_elements"]

Batch Processing

Pass a list of file paths to extract multiple files in one call.

reader = XbergReader()
documents = reader.load_data(["report.pdf", "slides.pptx", "data.xlsx"])

Raw Bytes

Use data= and mime_type= keyword arguments to extract from in-memory bytes.

reader = XbergReader()

# Single bytes input
documents = reader.load_data(data=pdf_bytes, mime_type="application/pdf")

# Batch bytes input -- parallel lists of data and MIME types
documents = reader.load_data(
    data=[pdf_bytes, docx_bytes],
    mime_type=["application/pdf", "application/vnd.openxmlformats-officedocument.wordprocessingml.document"],
)

Async

aload_data provides native async extraction backed by xberg's Rust runtime.

documents = await reader.aload_data(["file1.pdf", "file2.pdf"])

SimpleDirectoryReader Integration

Register XbergReader as a file extractor for any supported extension.

from llama_index.core import SimpleDirectoryReader

reader = XbergReader()
sdr = SimpleDirectoryReader(
    input_dir="./documents",
    file_extractor={".pdf": reader, ".docx": reader, ".html": reader},
)
documents = sdr.load_data()

# Async variant works too
documents = await sdr.aload_data()

Behavior Notes

  • Error tolerance: By default, raise_on_error=False -- the reader logs warnings and skips files that fail extraction.
  • Strict mode: Set raise_on_error=True to propagate extraction exceptions immediately.
  • Deterministic IDs: Document IDs are SHA-256 hashes of the file path (or byte content) and page number, enabling stable deduplication across pipeline runs.
  • Metadata exclusion: Large metadata fields (_xberg_chunks, _xberg_elements, images) are automatically excluded from LLM and embedding metadata keys to keep prompt sizes manageable.
  • Table handling: Tables extracted by xberg are appended as markdown to the document text when they are not already present in the content.
  • Serialization: The reader fully supports to_dict()/from_dict() round-tripping, including ExtractionConfig with nested OcrConfig and PageConfig. This enables pipeline caching and persistence with IngestionPipeline.

Metadata Reference

Each Document produced by XbergReader includes these metadata fields (when available from the source document):

Field Type Description
file_name str Source file name or "bytes" for raw bytes input
file_path str Absolute path to the source file
file_type str MIME type of the source document
total_pages int Total page count of the source document
page_number int Page number (present only with per-page splitting)
quality_score float Extraction quality score (0.0 -- 1.0)
detected_languages list[str] ISO language codes detected in the text
output_format str Format of the extracted content ("text", "markdown", etc.)
extracted_keywords list[dict] Keywords with text, score, and algorithm
annotations list[dict] Document annotations (comments, highlights)
processing_warnings list[dict] Warnings encountered during extraction
_xberg_chunks list[dict] Native semantic chunks with heading path and page span (with chunking=...)
_xberg_elements list Structural elements (with result_format="element_based")
images list[dict] Base64-encoded images with position and format metadata

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

llama_index_readers_xberg-1.0.11.tar.gz (13.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

llama_index_readers_xberg-1.0.11-py3-none-any.whl (12.7 kB view details)

Uploaded Python 3

File details

Details for the file llama_index_readers_xberg-1.0.11.tar.gz.

File metadata

  • Download URL: llama_index_readers_xberg-1.0.11.tar.gz
  • Upload date:
  • Size: 13.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: uv/0.12.1 {"installer":{"name":"uv","version":"0.12.1","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

File hashes

Hashes for llama_index_readers_xberg-1.0.11.tar.gz
Algorithm Hash digest
SHA256 51476f9e52ba5829338181c332577bc931bea3222cd179d08e9a3f0b4086d201
MD5 06b7cafaff06e99c9a14077bfb17b317
BLAKE2b-256 5b3bc7fff84dbbde63688d8c845eab7aad51e0c0c6e640d00b1cc3b826681695

See more details on using hashes here.

File details

Details for the file llama_index_readers_xberg-1.0.11-py3-none-any.whl.

File metadata

  • Download URL: llama_index_readers_xberg-1.0.11-py3-none-any.whl
  • Upload date:
  • Size: 12.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: uv/0.12.1 {"installer":{"name":"uv","version":"0.12.1","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

File hashes

Hashes for llama_index_readers_xberg-1.0.11-py3-none-any.whl
Algorithm Hash digest
SHA256 66f2c8b6f26619da0549e8d40ebc8af3d24662bedb62240aec97e13d0a69fd34
MD5 2122ec7efbbb0dfdd7fc41c19b672663
BLAKE2b-256 2378581f784602b3888938c5b68319eee7df9cb96cab9dcd85aef6100054e2ce

See more details on using hashes here.

Release history Release notifications | RSS feed

1.2.1

2 files

1.1.5

2 files

1.1.3

2 files

1.1.2

2 files

1.1.1

2 files

1.1.0

2 files

1.0.14

2 files

1.0.12

2 files

This release

1.0.11 This release

2 files

1.0.10

2 files

1.0.9

2 files

1.0.8

2 files

1.0.7

2 files

1.0.5

2 files

1.0.3

2 files

1.0.1

2 files

1.0.0

2 files

0.0.1

1 file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page