Skip to main content

Structure-aware LlamaIndex node parser that turns xberg native chunks and elements into nodes

Project description

LlamaIndex Node Parser Xberg

Xberg Banner

Structure-aware LlamaIndex node parser for xberg-extracted documents. It turns xberg's native chunks into nodes, and falls back to structural elements when chunks are absent.

Installation

pip install llama-index-node-parser-xberg

Requires llama-index-core>=0.14.23,<0.15. This package does not depend on xberg directly — xberg is a dependency of the reader (llama-index-readers-xberg), which produces the documents this parser splits.

Prerequisites

This parser requires documents with _xberg_chunks or _xberg_elements metadata. These are produced by XbergReader. Prefer native chunking; use element-based extraction when you want one node per structural element. Documents carrying neither pass through unchanged with a warning.

from xberg import ChunkingConfig, ExtractionConfig
from llama_index.readers.xberg import XbergReader

# Preferred: native semantic chunks with heading path and page span.
reader = XbergReader(
    extraction_config=ExtractionConfig(chunking=ChunkingConfig(max_characters=1000, overlap=200))
)
documents = reader.load_data("report.pdf")

Features

  • Chunk-aware splitting — each xberg native chunk becomes a node, carrying chunk_type, heading_path, and page span
  • Element fallback — when no chunks are present, headings, paragraphs, tables, and code blocks each become a node
  • Source and prev/next relationships tracked via NodeRelationship
  • Graceful degradation — documents without chunk or element metadata pass through with a warning
  • Composes with other transformations (e.g., SentenceSplitter)
  • Async support via aget_nodes_from_documents
  • Serialization support (to_dict / from_dict)

Usage

Basic

Full reader-to-nodes flow:

from xberg import ChunkingConfig, ExtractionConfig
from llama_index.readers.xberg import XbergReader
from llama_index.node_parser.xberg import XbergNodeParser

reader = XbergReader(
    extraction_config=ExtractionConfig(chunking=ChunkingConfig(max_characters=1000, overlap=200))
)
documents = reader.load_data("report.pdf")

parser = XbergNodeParser()
nodes = parser.get_nodes_from_documents(documents)

IngestionPipeline

Chain with SentenceSplitter to further split any oversized nodes:

from llama_index.core.ingestion import IngestionPipeline
from llama_index.core.node_parser import SentenceSplitter

pipeline = IngestionPipeline(
    transformations=[
        XbergNodeParser(),
        SentenceSplitter(chunk_size=512),  # Further split large nodes
    ]
)
nodes = pipeline.run(documents=documents)

VectorStoreIndex

Using the transformations parameter:

from llama_index.core import VectorStoreIndex

index = VectorStoreIndex.from_documents(
    documents,
    transformations=[XbergNodeParser()],
)

Async

nodes = await parser.aget_nodes_from_documents(documents)

Behavior Notes

  • Chunks take priority over elements. When a document carries both _xberg_chunks and _xberg_elements, the parser splits on chunks.
  • Documents without either metadata key pass through unchanged with a warning. This is intentional — silently falling back would hide that you are not getting structure-aware splitting.
  • Empty or whitespace-only chunks and elements are automatically skipped.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

llama_index_node_parser_xberg-1.0.0rc37.tar.gz (7.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

File details

Details for the file llama_index_node_parser_xberg-1.0.0rc37.tar.gz.

File metadata

  • Download URL: llama_index_node_parser_xberg-1.0.0rc37.tar.gz
  • Upload date:
  • Size: 7.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: uv/0.11.32 {"installer":{"name":"uv","version":"0.11.32","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

File hashes

Hashes for llama_index_node_parser_xberg-1.0.0rc37.tar.gz
Algorithm Hash digest
SHA256 fe8b5ff2764c516422ed00ae244e4bf58fad7e7df5d44c05094b66a7419e8f27
MD5 2ef7b1d9649c3e5153b9961f752b630f
BLAKE2b-256 1eba3957df303c8fa6abc1df6951dd31747a8140e0fe277f3a04ae8be6b3b101

See more details on using hashes here.

File details

Details for the file llama_index_node_parser_xberg-1.0.0rc37-py3-none-any.whl.

File metadata

  • Download URL: llama_index_node_parser_xberg-1.0.0rc37-py3-none-any.whl
  • Upload date:
  • Size: 5.6 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: uv/0.11.32 {"installer":{"name":"uv","version":"0.11.32","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

File hashes

Hashes for llama_index_node_parser_xberg-1.0.0rc37-py3-none-any.whl
Algorithm Hash digest
SHA256 0100c10b07d6064ef94505e153bbd05b5662cbea1db00988d2e645ce78b444cf
MD5 7d141329dd42d7e691fa758ee3075022
BLAKE2b-256 2591557035339bc5298ccdf5a69361a2afeaa3909a8925b4f39db0d15cd58154

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page