Skip to main content

kreuzberg-txtai

Kreuzberg Banner

A Kreuzberg-backed document extraction pipeline for txtai and any Python framework built around the __call__ convention.

KreuzbergPipeline replaces txtai's built-in Textractor (Apache Tika-based) with Kreuzberg's Rust-powered extraction stack, turning document paths into a list[dict] with content and metadata fields — surfacing title, MIME type, and page count that Tika flattens away.

Installation

pip install kreuzberg-txtai

For the txtai integration examples below:

pip install "kreuzberg-txtai[txtai]"

Requires Python 3.10+.

Quick Start

from kreuzberg_txtai import KreuzbergPipeline

pipeline = KreuzbergPipeline()
docs = pipeline(["doc1.pdf", "doc2.docx", "doc3.html"])

for doc in docs:
    print(doc["metadata"]["source"], "->", len(doc["content"]), "chars")

Each element in docs looks like:

{
    "content": "# Sample Document\n\nExtracted text...",
    "metadata": {
        "source": "doc1.pdf",
        "mime_type": "application/pdf",
        "title": "Sample Document",
        "page_count": 5,
    },
}

Features

  • 88+ file formats — PDF, DOCX, PPTX, XLSX, images, HTML, Markdown, plain text, and more via Kreuzberg
  • Stable dict contract — every extraction returns content + metadata with the same four keys, regardless of source format
  • Rich metadata — source path, MIME type, title, and page count surface directly
  • Batch support — pass a single path or a list[str]; output is always list[dict] in input order
  • Full Kreuzberg control — pass an ExtractionConfig to drive output format, OCR backend/language, force_ocr, and every other Kreuzberg knob
  • Framework-agnostic — txtai is an optional extra, not a hard dependency; the pipeline works in any framework that accepts a callable
  • Typed — ships with a py.typed marker; full mypy strict compatibility

Usage Examples

RAG ingestion with txtai.Embeddings

The dominant real-world pattern — extract, index, search:

from kreuzberg_txtai import KreuzbergPipeline
from txtai import Embeddings

pipeline = KreuzbergPipeline()
docs = pipeline(["doc1.pdf", "doc2.docx", "doc3.html"])

embeddings = Embeddings({
    "path": "sentence-transformers/all-MiniLM-L6-v2",
    "content": True,
})
embeddings.index([(i, doc["content"], None) for i, doc in enumerate(docs)])

results = embeddings.search("query", limit=5)

Inside a txtai.workflow.Task

Task accepts any callable, so KreuzbergPipeline drops in without wrappers. Because the pipeline returns list[dict], downstream tasks that expect strings need a one-line adapter:

from txtai.workflow import Task, Workflow
from kreuzberg_txtai import KreuzbergPipeline

extract = KreuzbergPipeline()

wf = Workflow([
    Task(extract),
    Task(lambda docs: [d["content"] for d in docs]),  # flatten dicts -> strings
])

list(wf(["doc1.pdf", "doc2.pdf"]))

Framework-free loop

from kreuzberg import ExtractionConfig
from kreuzberg_txtai import KreuzbergPipeline

pipeline = KreuzbergPipeline(config=ExtractionConfig(output_format="plain"))
for doc in pipeline(["scan1.pdf", "scan2.pdf"]):
    print(doc["metadata"]["source"], "->", len(doc["content"]), "chars")

No txtai needed — the class works on just the core kreuzberg dependency.

Tuning extraction with ExtractionConfig

Every Kreuzberg knob — output format, OCR backend and language, force_ocr, chunking, custom mime handling — lives on ExtractionConfig. Build one and hand it to the pipeline:

from kreuzberg import ExtractionConfig, OcrConfig
from kreuzberg_txtai import KreuzbergPipeline

custom = ExtractionConfig(
    output_format="markdown",
    ocr=OcrConfig(backend="tesseract", language="eng+deu"),
    force_ocr=True,
)

pipeline = KreuzbergPipeline(config=custom)
docs = pipeline("scanned_report.pdf")

See the Kreuzberg docs for the full list of ExtractionConfig and OcrConfig fields.

Constructor

Parameter Type Default Notes
config ExtractionConfig | None None Drives output format, OCR settings, force_ocr, and every other Kreuzberg option. None falls back to Kreuzberg's defaults.

Return Shape

__call__ always returns list[dict] — a single-path input still returns a length-1 list. Each dict has exactly two top-level keys:

  • content — the extracted text in the format set by config.output_format (Kreuzberg's default when no config is passed)
  • metadata — a dict with exactly four keys: source, mime_type, title, page_count

Missing metadata fields are None (rather than omitted) to keep the dict shape stable across document types.

Related Projects

License

MIT — see LICENSE.

Metadata

Release files for kreuzberg-txtai 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for kreuzberg-txtai 0.1.0
File Size Uploaded
kreuzberg_txtai-0.1.0.tar.gz 6.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for kreuzberg-txtai 0.1.0
File Interpreter ABI Platform
kreuzberg_txtai-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 13.1 kB

Release files / kreuzberg_txtai-0.1.0.tar.gz

Download URL kreuzberg_txtai-0.1.0.tar.gz
Size 6.5 kB
Tags Source
SHA-256 checksum
How to use checksums
886c194e4762205c90353d382d9f7725123e701b84169c7e7a88f97c7ef7e4c1
BLAKE2b-256 checksum
How to use checksums
a4d7a7de0a2f3b89bb72156ad008646018394aa50f2e519dceb858a1030c7dea
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Apr 16, 2026.

Transparency log

Release files / kreuzberg_txtai-0.1.0-py3-none-any.whl

Download URL kreuzberg_txtai-0.1.0-py3-none-any.whl
Size 6.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
8e91f6a85c204d3845ae3961ecb45c3167f91b4280ba90b40d91e0603883e7b7
BLAKE2b-256 checksum
How to use checksums
c1526432f2cea53b4daa130bacb8eb58de749fe034197b25a62bcef225add1ae
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Apr 16, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page