Skip to main content

Structure-aware LlamaIndex node parser that turns xberg native chunks and elements into nodes

Project description

LlamaIndex Node Parser Xberg

Xberg Banner

Structure-aware LlamaIndex node parser for xberg-extracted documents. It turns xberg's native chunks into nodes, and falls back to structural elements when chunks are absent.

Installation

pip install llama-index-node-parser-xberg

Requires llama-index-core>=0.14.23,<0.15. This package does not depend on xberg directly — xberg is a dependency of the reader (llama-index-readers-xberg), which produces the documents this parser splits.

Prerequisites

This parser requires documents with _xberg_chunks or _xberg_elements metadata. These are produced by XbergReader. Prefer native chunking; use element-based extraction when you want one node per structural element. Documents carrying neither pass through unchanged with a warning.

from xberg import ChunkingConfig, ExtractionConfig
from llama_index.readers.xberg import XbergReader

# Preferred: native semantic chunks with heading path and page span.
reader = XbergReader(
    extraction_config=ExtractionConfig(chunking=ChunkingConfig(max_characters=1000, overlap=200))
)
documents = reader.load_data("report.pdf")

Features

  • Chunk-aware splitting — each xberg native chunk becomes a node, carrying chunk_type, heading_path, and page span
  • Element fallback — when no chunks are present, headings, paragraphs, tables, and code blocks each become a node
  • Source and prev/next relationships tracked via NodeRelationship
  • Graceful degradation — documents without chunk or element metadata pass through with a warning
  • Composes with other transformations (e.g., SentenceSplitter)
  • Async support via aget_nodes_from_documents
  • Serialization support (to_dict / from_dict)

Usage

Basic

Full reader-to-nodes flow:

from xberg import ChunkingConfig, ExtractionConfig
from llama_index.readers.xberg import XbergReader
from llama_index.node_parser.xberg import XbergNodeParser

reader = XbergReader(
    extraction_config=ExtractionConfig(chunking=ChunkingConfig(max_characters=1000, overlap=200))
)
documents = reader.load_data("report.pdf")

parser = XbergNodeParser()
nodes = parser.get_nodes_from_documents(documents)

IngestionPipeline

Chain with SentenceSplitter to further split any oversized nodes:

from llama_index.core.ingestion import IngestionPipeline
from llama_index.core.node_parser import SentenceSplitter

pipeline = IngestionPipeline(
    transformations=[
        XbergNodeParser(),
        SentenceSplitter(chunk_size=512),  # Further split large nodes
    ]
)
nodes = pipeline.run(documents=documents)

VectorStoreIndex

Using the transformations parameter:

from llama_index.core import VectorStoreIndex

index = VectorStoreIndex.from_documents(
    documents,
    transformations=[XbergNodeParser()],
)

Async

nodes = await parser.aget_nodes_from_documents(documents)

Behavior Notes

  • Chunks take priority over elements. When a document carries both _xberg_chunks and _xberg_elements, the parser splits on chunks.
  • Documents without either metadata key pass through unchanged with a warning. This is intentional — silently falling back would hide that you are not getting structure-aware splitting.
  • Empty or whitespace-only chunks and elements are automatically skipped.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

llama_index_node_parser_xberg-1.0.0rc39.tar.gz (7.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

File details

Details for the file llama_index_node_parser_xberg-1.0.0rc39.tar.gz.

File metadata

  • Download URL: llama_index_node_parser_xberg-1.0.0rc39.tar.gz
  • Upload date:
  • Size: 7.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: uv/0.11.32 {"installer":{"name":"uv","version":"0.11.32","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

File hashes

Hashes for llama_index_node_parser_xberg-1.0.0rc39.tar.gz
Algorithm Hash digest
SHA256 61c33e2f93f5e055a401d52a4f6df989f6fc484228698afb2e5252babc47fad9
MD5 9115aaf893b6f12867803b54d3532b8a
BLAKE2b-256 651c74d09ce9e2319422e53340e2bc51555a15e0c95c11e470696db04919885d

See more details on using hashes here.

File details

Details for the file llama_index_node_parser_xberg-1.0.0rc39-py3-none-any.whl.

File metadata

  • Download URL: llama_index_node_parser_xberg-1.0.0rc39-py3-none-any.whl
  • Upload date:
  • Size: 5.6 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: uv/0.11.32 {"installer":{"name":"uv","version":"0.11.32","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

File hashes

Hashes for llama_index_node_parser_xberg-1.0.0rc39-py3-none-any.whl
Algorithm Hash digest
SHA256 39dc96e3b06d0537ffa61a4dfc969ffefab34d6131990e80f4544ebeea5d3f95
MD5 7114e5e78a52296b2d6f001d4320cd9a
BLAKE2b-256 af47e42922639302001175973a75c1470093d97ee5411a90b776e8c84ee84456

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page