Skip to main content

LlamaIndex Node Parser Kreuzberg

Kreuzberg Banner

Element-aware LlamaIndex node parser for kreuzberg-extracted documents.

Installation

pip install llama-index-node-parser-kreuzberg

Requires llama-index-core>=0.13.0,<0.15. This package does not depend on kreuzberg directly — the kreuzberg package is a dependency of the reader (llama-index-readers-kreuzberg), which is needed for producing documents with element metadata.

Prerequisites

This parser requires documents with _kreuzberg_elements metadata. These are produced by KreuzbergReader configured with element-based extraction. Install llama-index-readers-kreuzberg (which brings in kreuzberg) to use the full workflow.

from kreuzberg import ExtractionConfig
from llama_index.readers.kreuzberg import KreuzbergReader

reader = KreuzbergReader(
    extraction_config=ExtractionConfig(result_format="element_based")
)
documents = reader.load_data("report.pdf")

Features

  • Element-aware splitting — headings, paragraphs, tables, and code blocks each become a node
  • Element type metadata preserved on each node (element_type, page_number, element_index)
  • Source document relationships tracked via NodeRelationship.SOURCE
  • Graceful degradation — documents without elements pass through with a warning
  • Composes with other transformations (e.g., SentenceSplitter for further chunking)
  • Async support via aget_nodes_from_documents
  • Serialization support (to_dict / from_dict)

Usage

Basic

Full reader-to-nodes flow:

from kreuzberg import ExtractionConfig
from llama_index.readers.kreuzberg import KreuzbergReader
from llama_index.node_parser.kreuzberg import KreuzbergNodeParser

reader = KreuzbergReader(
    extraction_config=ExtractionConfig(result_format="element_based")
)
documents = reader.load_data("report.pdf")

parser = KreuzbergNodeParser()
nodes = parser.get_nodes_from_documents(documents)

IngestionPipeline

Chain with SentenceSplitter for further chunking of large elements:

from llama_index.core.ingestion import IngestionPipeline
from llama_index.core.node_parser import SentenceSplitter

pipeline = IngestionPipeline(
    transformations=[
        KreuzbergNodeParser(),
        SentenceSplitter(chunk_size=512),  # Further split large elements
    ]
)
nodes = pipeline.run(documents=documents)

VectorStoreIndex

Using the transformations parameter:

from llama_index.core import VectorStoreIndex

index = VectorStoreIndex.from_documents(
    documents,
    transformations=[KreuzbergNodeParser()],
)

Async

nodes = await parser.aget_nodes_from_documents(documents)

Behavior Notes

  • Documents without _kreuzberg_elements metadata pass through unchanged with a warning. This is intentional — silently falling back would prevent users from noticing they are not getting element-aware splitting.
  • Empty or whitespace-only elements are automatically skipped.

Metadata

Release files for llama-index-node-parser-kreuzberg 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for llama-index-node-parser-kreuzberg 0.1.0
File Size Uploaded
llama_index_node_parser_kreuzberg-0.1.0.tar.gz 4.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for llama-index-node-parser-kreuzberg 0.1.0
File Interpreter ABI Platform
llama_index_node_parser_kreuzberg-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 9.3 kB

Release files / llama_index_node_parser_kreuzberg-0.1.0.tar.gz

Download URL llama_index_node_parser_kreuzberg-0.1.0.tar.gz
Size 4.4 kB
Tags Source
SHA-256 checksum
How to use checksums
6941aacb7e44ff8119ab6c8566984623b0eb5e1abff8b47b975bf5050ac39b86
BLAKE2b-256 checksum
How to use checksums
1be8debc70a91133ffaf5218a5c439c7a65188c5063134631ce94a79d88e50ed
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.7

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Mar 21, 2026.

Transparency log

Release files / llama_index_node_parser_kreuzberg-0.1.0-py3-none-any.whl

Download URL llama_index_node_parser_kreuzberg-0.1.0-py3-none-any.whl
Size 4.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
3b768fe93b9fd0f456a780dedb733e89c01441753e815ed55017af981367a8fa
BLAKE2b-256 checksum
How to use checksums
e2b2584f46cbafd0abaa99fbeb83559d9266398a6ad11543745ec903fbcaa84a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.7

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Mar 21, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page