vera-ingest-docling
Optional Docling ingest pipeline for VERA. Registers the docling provider
(default variant hybrid) under the vera.ingest_pipelines entry-point group.
The pipeline uses Docling's DocumentConverter and HybridChunker. PDFs get
layout-aware parsing, RapidOCR, and page-level recovery. DOCX, PPTX, XLSX, and
HTML are converted for search only (no PDF layout models or highlight overlay).
Readable chunk text is stored for keyword search; contextualized text from
HybridChunker.contextualize() is used for embeddings.
Install
python -m pip install "vera-ingest-docling>=0.3.0"
From a repository checkout with uv (workspace .venv for CLI and tests):
uv sync --extra docling
Non-desktop users can install the CLI extra:
pip install "vera[docling]>=0.3.0"
Python 3.10 or newer is required. The package depends on Docling's rapidocr
extra so RapidOCR and onnxruntime are installed for OCR. RapidOCR ONNX
weights come with that extra; first PDF conversion may still download Docling
layout models (about 380 MB: Heron ONNX + TableFormer accurate). Set
DOCLING_ARTIFACTS_PATH to a local cache (or prefetch layout models offline)
for air-gapped runs. Incomplete caches are not treated as ready. Hub progress
is visible on CLI stderr.
This extra is not bundled in the 0.3.0 desktop installer and is not listed in
Convert.
Usage
vera convert "manual.pdf" "manual.vera" --parser docling
vera convert "memo.docx" "memo.vera" --parser docling
vera convert "notes.html" "notes.vera" # Docling is selected from the extension
# or explicitly:
vera convert "manual.pdf" "manual.vera" --parser docling:hybrid
from vera_ingest import convert
convert("manual.pdf", "manual.vera", parser="docling")
convert("memo.docx", "memo.vera", parser="docling")
convert("notes.html", "notes.vera") # parser inferred when the extra is installed
Notes
- Defaults:
chunk_size=500whitespace tokens (not LLM subword tokens),ocr_mode=auto,ocr_language=en,pdf_backend=docling_parse. Docling does not advertiseoverlaporocr_dpi, so those legacy convert/CLI aliases are not forwarded. The Tesseract--ocr-languagealias is also not forwarded; Docling keepsen. - Prefer
pipeline_options=/--pipeline-option KEY=VALUEfor provider-owned settings;--chunk-sizeand--ocr*remain compatibility aliases for pipelines that accept them. - OCR modes map to Docling/RapidOCR:
off,auto(default), andforce(full-page OCR).ocr_languageexpects a RapidOCR-native code (en,fr,cyrillic, ...) — Tesseract-style codes such asengare not translated. The shared--ocr-languageCLI default (eng) is not forwarded; pass--pipeline-option ocr_language=fr(or another RapidOCR code) when you need a non-default language. - Torch model compilation is disabled so Windows does not need MSVC
cl.exe. - Picture crops are stored as figure attachments and linked onto a nearby
same-page chunk so search
--figurescan return them. Docling's HybridChunker omits pictures from chunk text. - On PDF page-level memory errors (
bad_alloc), VERA retries failed pages then falls back to whole-documentpypdfium2, then page-batchpypdfium2if that still raises. Force the backend with--pipeline-option pdf_backend=pypdfium2. Conversion rejects only when recovery is exhausted. Failures include the underlying exception and print it on sidecar stderr. Office/HTML conversions do not use this PDF recovery path.
See the vera-ingest-docling documentation and conversion guide.
Metadata
Release files for vera-ingest-docling 0.3.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| vera_ingest_docling-0.3.2.tar.gz | 33.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| vera_ingest_docling-0.3.2-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 58.8 kB
Release files / vera_ingest_docling-0.3.2.tar.gz
| Download URL | vera_ingest_docling-0.3.2.tar.gz |
|---|---|
| Size | 33.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
bc42b4b5454cadbc04d7820ec5407706223ce0e5e42b9c97796a889703c8bbae
|
|
BLAKE2b-256 checksum How to use checksums |
8b62d324bd1a09b90a12160dac13e8fe84274ecbca1de41cca867edf4bb1e4d9
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 11, 2026.
Transparency logRelease files / vera_ingest_docling-0.3.2-py3-none-any.whl
| Download URL | vera_ingest_docling-0.3.2-py3-none-any.whl |
|---|---|
| Size | 25.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
f1c76a2b2ac49143859fa1be187597de4d52ecd4768ec4620fb862741ecf8fd9
|
|
BLAKE2b-256 checksum How to use checksums |
5fb22072282a6c2989324ba95f96ad9488faabb91d628e76042fbeca9362b955
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 11, 2026.
Transparency log