Skip to main content

Structure-aware LlamaIndex node parser that turns xberg native chunks and elements into nodes

Project description

LlamaIndex Node Parser Xberg

Xberg Banner

Structure-aware LlamaIndex node parser for xberg-extracted documents. It turns xberg's native chunks into nodes, and falls back to structural elements when chunks are absent.

Installation

pip install llama-index-node-parser-xberg

Requires llama-index-core>=0.14.23,<0.15. This package does not depend on xberg directly — xberg is a dependency of the reader (llama-index-readers-xberg), which produces the documents this parser splits.

Prerequisites

This parser requires documents with _xberg_chunks or _xberg_elements metadata. These are produced by XbergReader. Prefer native chunking; use element-based extraction when you want one node per structural element. Documents carrying neither pass through unchanged with a warning.

from xberg import ChunkingConfig, ExtractionConfig
from llama_index.readers.xberg import XbergReader

# Preferred: native semantic chunks with heading path and page span.
reader = XbergReader(
    extraction_config=ExtractionConfig(chunking=ChunkingConfig(max_characters=1000, overlap=200))
)
documents = reader.load_data("report.pdf")

Features

  • Chunk-aware splitting — each xberg native chunk becomes a node, carrying chunk_type, heading_path, and page span
  • Element fallback — when no chunks are present, headings, paragraphs, tables, and code blocks each become a node
  • Source and prev/next relationships tracked via NodeRelationship
  • Graceful degradation — documents without chunk or element metadata pass through with a warning
  • Composes with other transformations (e.g., SentenceSplitter)
  • Async support via aget_nodes_from_documents
  • Serialization support (to_dict / from_dict)

Usage

Basic

Full reader-to-nodes flow:

from xberg import ChunkingConfig, ExtractionConfig
from llama_index.readers.xberg import XbergReader
from llama_index.node_parser.xberg import XbergNodeParser

reader = XbergReader(
    extraction_config=ExtractionConfig(chunking=ChunkingConfig(max_characters=1000, overlap=200))
)
documents = reader.load_data("report.pdf")

parser = XbergNodeParser()
nodes = parser.get_nodes_from_documents(documents)

IngestionPipeline

Chain with SentenceSplitter to further split any oversized nodes:

from llama_index.core.ingestion import IngestionPipeline
from llama_index.core.node_parser import SentenceSplitter

pipeline = IngestionPipeline(
    transformations=[
        XbergNodeParser(),
        SentenceSplitter(chunk_size=512),  # Further split large nodes
    ]
)
nodes = pipeline.run(documents=documents)

VectorStoreIndex

Using the transformations parameter:

from llama_index.core import VectorStoreIndex

index = VectorStoreIndex.from_documents(
    documents,
    transformations=[XbergNodeParser()],
)

Async

nodes = await parser.aget_nodes_from_documents(documents)

Behavior Notes

  • Chunks take priority over elements. When a document carries both _xberg_chunks and _xberg_elements, the parser splits on chunks.
  • Documents without either metadata key pass through unchanged with a warning. This is intentional — silently falling back would hide that you are not getting structure-aware splitting.
  • Empty or whitespace-only chunks and elements are automatically skipped.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distributions

No source distribution files available for this release.See tutorial on generating distribution archives.

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

File details

Details for the file llama_index_node_parser_xberg-1.0.0rc36-py3-none-any.whl.

File metadata

  • Download URL: llama_index_node_parser_xberg-1.0.0rc36-py3-none-any.whl
  • Upload date:
  • Size: 5.6 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: uv/0.11.32 {"installer":{"name":"uv","version":"0.11.32","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

File hashes

Hashes for llama_index_node_parser_xberg-1.0.0rc36-py3-none-any.whl
Algorithm Hash digest
SHA256 623caced3da36573fd6e3bfeaf4d9d25797ce3c8ab40a5a96cc2dfb4c10646dc
MD5 e2e391611b417d6c58368b44b59ef307
BLAKE2b-256 b790fe5643a956791dbf7cfde5736f9e3e943f5571478e95d90ad4ebe3a516c6

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page