kreuzberg-txtai
A Kreuzberg-backed document extraction pipeline for txtai and any Python framework built around the __call__ convention.
KreuzbergPipeline replaces txtai's built-in Textractor (Apache Tika-based) with Kreuzberg's Rust-powered extraction stack, turning document paths into a list[dict] with content and metadata fields — surfacing title, MIME type, and page count that Tika flattens away.
Installation
pip install kreuzberg-txtai
For the txtai integration examples below:
pip install "kreuzberg-txtai[txtai]"
Requires Python 3.10+.
Quick Start
from kreuzberg_txtai import KreuzbergPipeline
pipeline = KreuzbergPipeline()
docs = pipeline(["doc1.pdf", "doc2.docx", "doc3.html"])
for doc in docs:
print(doc["metadata"]["source"], "->", len(doc["content"]), "chars")
Each element in docs looks like:
{
"content": "# Sample Document\n\nExtracted text...",
"metadata": {
"source": "doc1.pdf",
"mime_type": "application/pdf",
"title": "Sample Document",
"page_count": 5,
},
}
Features
- 88+ file formats — PDF, DOCX, PPTX, XLSX, images, HTML, Markdown, plain text, and more via Kreuzberg
- Stable dict contract — every extraction returns
content+metadatawith the same four keys, regardless of source format - Rich metadata — source path, MIME type, title, and page count surface directly
- Batch support — pass a single path or a
list[str]; output is alwayslist[dict]in input order - Full Kreuzberg control — pass an
ExtractionConfigto drive output format, OCR backend/language,force_ocr, and every other Kreuzberg knob - Framework-agnostic — txtai is an optional extra, not a hard dependency; the pipeline works in any framework that accepts a callable
- Typed — ships with a
py.typedmarker; full mypy strict compatibility
Usage Examples
RAG ingestion with txtai.Embeddings
The dominant real-world pattern — extract, index, search:
from kreuzberg_txtai import KreuzbergPipeline
from txtai import Embeddings
pipeline = KreuzbergPipeline()
docs = pipeline(["doc1.pdf", "doc2.docx", "doc3.html"])
embeddings = Embeddings({
"path": "sentence-transformers/all-MiniLM-L6-v2",
"content": True,
})
embeddings.index([(i, doc["content"], None) for i, doc in enumerate(docs)])
results = embeddings.search("query", limit=5)
Inside a txtai.workflow.Task
Task accepts any callable, so KreuzbergPipeline drops in without wrappers. Because the pipeline returns list[dict], downstream tasks that expect strings need a one-line adapter:
from txtai.workflow import Task, Workflow
from kreuzberg_txtai import KreuzbergPipeline
extract = KreuzbergPipeline()
wf = Workflow([
Task(extract),
Task(lambda docs: [d["content"] for d in docs]), # flatten dicts -> strings
])
list(wf(["doc1.pdf", "doc2.pdf"]))
Framework-free loop
from kreuzberg import ExtractionConfig
from kreuzberg_txtai import KreuzbergPipeline
pipeline = KreuzbergPipeline(config=ExtractionConfig(output_format="plain"))
for doc in pipeline(["scan1.pdf", "scan2.pdf"]):
print(doc["metadata"]["source"], "->", len(doc["content"]), "chars")
No txtai needed — the class works on just the core kreuzberg dependency.
Tuning extraction with ExtractionConfig
Every Kreuzberg knob — output format, OCR backend and language, force_ocr, chunking, custom mime handling — lives on ExtractionConfig. Build one and hand it to the pipeline:
from kreuzberg import ExtractionConfig, OcrConfig
from kreuzberg_txtai import KreuzbergPipeline
custom = ExtractionConfig(
output_format="markdown",
ocr=OcrConfig(backend="tesseract", language="eng+deu"),
force_ocr=True,
)
pipeline = KreuzbergPipeline(config=custom)
docs = pipeline("scanned_report.pdf")
See the Kreuzberg docs for the full list of ExtractionConfig and OcrConfig fields.
Constructor
| Parameter | Type | Default | Notes |
|---|---|---|---|
config |
ExtractionConfig | None |
None |
Drives output format, OCR settings, force_ocr, and every other Kreuzberg option. None falls back to Kreuzberg's defaults. |
Return Shape
__call__ always returns list[dict] — a single-path input still returns a length-1 list. Each dict has exactly two top-level keys:
content— the extracted text in the format set byconfig.output_format(Kreuzberg's default when no config is passed)metadata— a dict with exactly four keys:source,mime_type,title,page_count
Missing metadata fields are None (rather than omitted) to keep the dict shape stable across document types.
Related Projects
- kreuzberg — the extraction engine powering this package
- langchain-kreuzberg — Kreuzberg document loader for LangChain
- llama-index-kreuzberg — LlamaIndex reader and node parser
- kreuzberg-crewai — CrewAI agent tool
- kreuzberg-surrealdb — SurrealDB ingestion connector
License
MIT — see LICENSE.
Metadata
Release files for kreuzberg-txtai 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| kreuzberg_txtai-0.1.0.tar.gz | 6.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| kreuzberg_txtai-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 13.1 kB
Release files / kreuzberg_txtai-0.1.0.tar.gz
| Download URL | kreuzberg_txtai-0.1.0.tar.gz |
|---|---|
| Size | 6.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
886c194e4762205c90353d382d9f7725123e701b84169c7e7a88f97c7ef7e4c1
|
|
BLAKE2b-256 checksum How to use checksums |
a4d7a7de0a2f3b89bb72156ad008646018394aa50f2e519dceb858a1030c7dea
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.12
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Apr 16, 2026.
Transparency logRelease files / kreuzberg_txtai-0.1.0-py3-none-any.whl
| Download URL | kreuzberg_txtai-0.1.0-py3-none-any.whl |
|---|---|
| Size | 6.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
8e91f6a85c204d3845ae3961ecb45c3167f91b4280ba90b40d91e0603883e7b7
|
|
BLAKE2b-256 checksum How to use checksums |
c1526432f2cea53b4daa130bacb8eb58de749fe034197b25a62bcef225add1ae
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.12
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Apr 16, 2026.
Transparency log