LlamaIndex Readers Kreuzberg
LlamaIndex reader for 88+ document formats powered by kreuzberg's Rust extraction engine.
Installation
pip install llama-index-readers-kreuzberg
Requires kreuzberg>=4.4.6 and llama-index-core>=0.13.0,<0.15.
Features
- 88+ formats -- PDF, DOCX, PPTX, XLSX, HTML, images, and more (full list)
- Rich metadata -- quality scores, language detection, keywords, annotations
- Element extraction -- structural elements for structure-aware RAG pipelines
- Image extraction -- base64-encoded image data with position, format, and OCR metadata
- Per-page splitting -- one
Documentper page for fine-grained retrieval - Batch processing -- multiple files in a single call
- Raw bytes input -- extract from in-memory bytes with a MIME type
- Native async -- true async via kreuzberg's Rust tokio runtime
- Error tolerance -- skip failed files with warnings, or raise on failure
- Full serialization -- custom
ExtractionConfiground-trips throughto_dict()/from_dict()for pipeline caching
Usage
Basic Extraction
from llama_index.readers.kreuzberg import KreuzbergReader
reader = KreuzbergReader()
documents = reader.load_data("report.pdf")
# Each document carries rich metadata
print(documents[0].metadata["file_name"]) # "report.pdf"
print(documents[0].metadata["file_type"]) # "application/pdf"
print(documents[0].metadata["total_pages"]) # 12
print(documents[0].metadata["quality_score"]) # 0.95
print(documents[0].metadata["detected_languages"]) # ["en"]
OCR Configuration
force_ocr is a top-level ExtractionConfig option. Language and backend are set on OcrConfig.
from kreuzberg import ExtractionConfig, OcrConfig
reader = KreuzbergReader(
extraction_config=ExtractionConfig(
force_ocr=True,
ocr=OcrConfig(language="deu", backend="tesseract"),
)
)
documents = reader.load_data("scanned.pdf")
Per-Page Splitting
PageConfig is nested inside ExtractionConfig. Each page becomes its own Document
with a page_number metadata field.
from kreuzberg import ExtractionConfig, PageConfig
reader = KreuzbergReader(
extraction_config=ExtractionConfig(
pages=PageConfig(extract_pages=True),
)
)
documents = reader.load_data("multi_page.pdf") # One Document per page
for doc in documents:
print(f"Page {doc.metadata['page_number']}: {doc.text[:80]}...")
Element Extraction
Setting result_format="element_based" populates _kreuzberg_elements in document
metadata for structure-aware processing.
from kreuzberg import ExtractionConfig
reader = KreuzbergReader(
extraction_config=ExtractionConfig(result_format="element_based")
)
documents = reader.load_data("report.pdf")
# Structural elements available for downstream node parsers
elements = documents[0].metadata["_kreuzberg_elements"]
Batch Processing
Pass a list of file paths to extract multiple files in one call.
reader = KreuzbergReader()
documents = reader.load_data(["report.pdf", "slides.pptx", "data.xlsx"])
Raw Bytes
Use data= and mime_type= keyword arguments to extract from in-memory bytes.
reader = KreuzbergReader()
# Single bytes input
documents = reader.load_data(data=pdf_bytes, mime_type="application/pdf")
# Batch bytes input -- parallel lists of data and MIME types
documents = reader.load_data(
data=[pdf_bytes, docx_bytes],
mime_type=["application/pdf", "application/vnd.openxmlformats-officedocument.wordprocessingml.document"],
)
Async
aload_data provides native async extraction backed by kreuzberg's Rust runtime.
documents = await reader.aload_data(["file1.pdf", "file2.pdf"])
SimpleDirectoryReader Integration
Register KreuzbergReader as a file extractor for any supported extension.
from llama_index.core import SimpleDirectoryReader
reader = KreuzbergReader()
sdr = SimpleDirectoryReader(
input_dir="./documents",
file_extractor={".pdf": reader, ".docx": reader, ".html": reader},
)
documents = sdr.load_data()
# Async variant works too
documents = await sdr.aload_data()
Behavior Notes
- Error tolerance: By default,
raise_on_error=False-- the reader logs warnings and skips files that fail extraction. - Strict mode: Set
raise_on_error=Trueto propagate extraction exceptions immediately. - Deterministic IDs: Document IDs are SHA-256 hashes of the file path (or byte content) and page number, enabling stable deduplication across pipeline runs.
- Metadata exclusion: Large metadata fields (
_kreuzberg_elements,images) are automatically excluded from LLM and embedding metadata keys to keep prompt sizes manageable. - Table handling: Tables extracted by kreuzberg are appended as markdown to the document text when they are not already present in the content.
- Serialization: The reader fully supports
to_dict()/from_dict()round-tripping, includingExtractionConfigwith nestedOcrConfigandPageConfig. This enables pipeline caching and persistence withIngestionPipeline.
Metadata Reference
Each Document produced by KreuzbergReader includes these metadata fields (when available from the source document):
| Field | Type | Description |
|---|---|---|
file_name |
str |
Source file name or "bytes" for raw bytes input |
file_path |
str |
Absolute path to the source file |
file_type |
str |
MIME type of the source document |
total_pages |
int |
Total page count of the source document |
page_number |
int |
Page number (present only with per-page splitting) |
quality_score |
float |
Extraction quality score (0.0 -- 1.0) |
detected_languages |
list[str] |
ISO language codes detected in the text |
output_format |
str |
Format of the extracted content ("text", "markdown", etc.) |
extracted_keywords |
list[dict] |
Keywords with text, score, and algorithm |
annotations |
list[dict] |
Document annotations (comments, highlights) |
processing_warnings |
list[dict] |
Warnings encountered during extraction |
_kreuzberg_elements |
list |
Structural elements (with result_format="element_based") |
images |
list[dict] |
Base64-encoded images with position and format metadata |
Metadata
Release files for llama-index-readers-kreuzberg 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| llama_index_readers_kreuzberg-0.1.1.tar.gz | 10.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| llama_index_readers_kreuzberg-0.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 23.2 kB
Release files / llama_index_readers_kreuzberg-0.1.1.tar.gz
| Download URL | llama_index_readers_kreuzberg-0.1.1.tar.gz |
|---|---|
| Size | 10.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
bce80572687cde7fc1f286a1737a9ed1e7d83955d6ba0cf359b5a8293255c03d
|
|
BLAKE2b-256 checksum How to use checksums |
6b158354e9a4fe40d43a92a73c18b9161af4e2b8b905f3ae6a6b10256dc20303
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.12
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Apr 25, 2026.
Transparency logRelease files / llama_index_readers_kreuzberg-0.1.1-py3-none-any.whl
| Download URL | llama_index_readers_kreuzberg-0.1.1-py3-none-any.whl |
|---|---|
| Size | 12.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
3072d750eb512c91a0ac605e0371a88222862e4fc319686224d3f6fbd5e109a2
|
|
BLAKE2b-256 checksum How to use checksums |
1c9070326293bb01e1a073e22018f965e6558400280be0711c358c73d127c1cc
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.12
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Apr 25, 2026.
Transparency log