LlamaIndex Node Parser Kreuzberg
Element-aware LlamaIndex node parser for kreuzberg-extracted documents.
Installation
pip install llama-index-node-parser-kreuzberg
Requires llama-index-core>=0.13.0,<0.15. This package does not depend on
kreuzberg directly — the kreuzberg package is a dependency of the reader
(llama-index-readers-kreuzberg), which is needed for producing documents with
element metadata.
Prerequisites
This parser requires documents with
_kreuzberg_elementsmetadata. These are produced byKreuzbergReaderconfigured with element-based extraction. Installllama-index-readers-kreuzberg(which brings inkreuzberg) to use the full workflow.
from kreuzberg import ExtractionConfig
from llama_index.readers.kreuzberg import KreuzbergReader
reader = KreuzbergReader(
extraction_config=ExtractionConfig(result_format="element_based")
)
documents = reader.load_data("report.pdf")
Features
- Element-aware splitting — headings, paragraphs, tables, and code blocks each become a node
- Element type metadata preserved on each node (
element_type,page_number,element_index) - Source document relationships tracked via
NodeRelationship.SOURCE - Graceful degradation — documents without elements pass through with a warning
- Composes with other transformations (e.g.,
SentenceSplitterfor further chunking) - Async support via
aget_nodes_from_documents - Serialization support (
to_dict/from_dict)
Usage
Basic
Full reader-to-nodes flow:
from kreuzberg import ExtractionConfig
from llama_index.readers.kreuzberg import KreuzbergReader
from llama_index.node_parser.kreuzberg import KreuzbergNodeParser
reader = KreuzbergReader(
extraction_config=ExtractionConfig(result_format="element_based")
)
documents = reader.load_data("report.pdf")
parser = KreuzbergNodeParser()
nodes = parser.get_nodes_from_documents(documents)
IngestionPipeline
Chain with SentenceSplitter for further chunking of large elements:
from llama_index.core.ingestion import IngestionPipeline
from llama_index.core.node_parser import SentenceSplitter
pipeline = IngestionPipeline(
transformations=[
KreuzbergNodeParser(),
SentenceSplitter(chunk_size=512), # Further split large elements
]
)
nodes = pipeline.run(documents=documents)
VectorStoreIndex
Using the transformations parameter:
from llama_index.core import VectorStoreIndex
index = VectorStoreIndex.from_documents(
documents,
transformations=[KreuzbergNodeParser()],
)
Async
nodes = await parser.aget_nodes_from_documents(documents)
Behavior Notes
- Documents without
_kreuzberg_elementsmetadata pass through unchanged with a warning. This is intentional — silently falling back would prevent users from noticing they are not getting element-aware splitting. - Empty or whitespace-only elements are automatically skipped.
Metadata
Release files for llama-index-node-parser-kreuzberg 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| llama_index_node_parser_kreuzberg-0.1.0.tar.gz | 4.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| llama_index_node_parser_kreuzberg-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 9.3 kB
Release files / llama_index_node_parser_kreuzberg-0.1.0.tar.gz
| Download URL | llama_index_node_parser_kreuzberg-0.1.0.tar.gz |
|---|---|
| Size | 4.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
6941aacb7e44ff8119ab6c8566984623b0eb5e1abff8b47b975bf5050ac39b86
|
|
BLAKE2b-256 checksum How to use checksums |
1be8debc70a91133ffaf5218a5c439c7a65188c5063134631ce94a79d88e50ed
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.7
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Mar 21, 2026.
Transparency logRelease files / llama_index_node_parser_kreuzberg-0.1.0-py3-none-any.whl
| Download URL | llama_index_node_parser_kreuzberg-0.1.0-py3-none-any.whl |
|---|---|
| Size | 4.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
3b768fe93b9fd0f456a780dedb733e89c01441753e815ed55017af981367a8fa
|
|
BLAKE2b-256 checksum How to use checksums |
e2b2584f46cbafd0abaa99fbeb83559d9266398a6ad11543745ec903fbcaa84a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.7
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Mar 21, 2026.
Transparency log