Skip to main content

DocVortex overview: native document inputs flow through a unified document model to Markdown, HTML, LaTeX, DOCX, EPUB, PDF and structured content.

DocVortex

A fast, multi-format document parsing and conversion engine.

DocVortex provides a complete, standalone document pipeline:

Document -> ModelJson + Assets -> MiddleJson + Assets -> Render / Export

Native inputs include text PDFs, DOC/DOCX, PPT/PPTX, XLS/XLSX, RTF, ODT/ODS/ODP, EPUB, HTML, OFD and CSV. Output formats include Markdown, HTML, LaTeX, DOCX, EPUB, PDF and structured content. Content List V1/V2 are provided by MinerU.

Native parsing runs without OCR or VLM inference services. PDF classification is an explicit document operation; native analysis does not silently classify the document or select another inference backend.

Install

pip install docvortex
docvortex convert report.pdf --format markdown --output output/report.md
docvortex classify report.pdf

Python 3.10–3.14 is supported. Native parsing does not require OCR/VLM inference services. PDF access uses pypdfium2>=5.10.1,<6; the compatibility matrix also exercises 5.13.0.

Parse once, export many times

import docvortex

result = docvortex.parse("report.pdf", keep_model_json=True)
result.export("output/report.md", output_format="markdown")
result.export("output/report.docx", output_format="docx")
result.export("output/report.epub", output_format="epub")
result.save_bundle("output/report.bundle")

# This works after the source document and its parsing process are gone.
restored = docvortex.load_bundle("output/report.bundle")
restored.export("output/report.pdf", output_format="pdf")

Bundles contain manifest.json, middle.json, optional model.json, and image assets. The loader verifies asset hashes. Missing external assets must be supplied before saving a portable bundle. Existing files are protected unless the caller explicitly sets overwrite=True.

Stage APIs

from docvortex.api import analyze, postprocess, render

analysis = analyze("report.pdf", page_range="1-5")
result = postprocess(analysis)
artifact = render(result.middle_json, "docx", assets=result.assets)
artifact.write("output/report.docx")

The stage API lives in docvortex.api. Root-level conveniences include parse, analyze, convert, postprocess_document, and render_artifact. The docvortex.render package also exposes the low-level renderers and their original string, bytes, dictionary, or list return values.

PDF page selections use 1-5, r1 and all; other native formats are parsed as whole documents. A caller-owned PDFDocument can be passed to analyze or parse and remains open afterward.

Explicit PDF classification

from docvortex.document.pdf import PDFDocument

with PDFDocument("report.pdf") as document:
    mode = document.classify()  # "txt" or "ocr"; no inference is started
    if mode == "txt":
        result = docvortex.parse(document)

Native analysis trusts the caller's choice and does not classify automatically. Applications can use the classification result to select their own OCR or inference service when a document requires it.

DocVortex JSON uses schema identity docvortex.model or docvortex.middle, schema version 2.0, and required metadata.file_suffix / metadata.producer. Definitions are in schemas/. Application-specific metadata belongs in extensions. See the shared JSON protocol and migration guide and the compatibility guide for existing application integrations and historical data formats. The HTML protocol describes DocVortex markers and semantic round trips.

Scope and development

PDF output is a semantic reflow of the document, not a lossless reproduction of the original page drawing instructions. Input support for PPTX/XLSX does not imply PPTX/XLSX output support. Rust implementation work is a future stage behind these public data and processing boundaries.

uv venv
uv pip install -e ".[test,dev]"
uv run --no-project python -m pytest -q
uv run --no-project ruff check src
uv run --no-project ruff format --check src
uv build

DocVortex project code is licensed under the MIT License.

Dependencies include pydantic>=2.12.5,<3 and numpy>=1.21.6; the installer selects versions compatible with the active Python interpreter. On Apple Silicon, Python 3.14 installation requires macOS 14 or newer because of the ONNX Runtime dependency used by file-type detection.

See the standalone example for native parsing and portable result bundles, and the validation record for test coverage.

PDFium uses a bundled, pinned CJK fallback font for non-embedded CJK fonts; no system font installation is required. See PDF font policy for initialization, diagnostics and replacement boundaries.

PDF output normalizes fullwidth Latin letters, digits and selected technical symbols in natural-language text and table cells, preserving Chinese punctuation, formulas, code and link targets. See PDF text normalization for scope and API usage.

See the DocVortex upgrade guide for package, protocol, PDF text rules and publishing configuration changes.

See rendering ownership for the seven engine targets, MinerU Content List integration and public fragment helpers.

See refactor validation for the staged internal refactoring, compatibility checks, corpus comparisons and measured performance.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

docvortex-0.2.1.tar.gz (27.9 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

docvortex-0.2.1-py3-none-any.whl (4.3 MB view details)

Uploaded Python 3

File details

Details for the file docvortex-0.2.1.tar.gz.

File metadata

  • Download URL: docvortex-0.2.1.tar.gz
  • Upload date:
  • Size: 27.9 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for docvortex-0.2.1.tar.gz
Algorithm Hash digest
SHA256 979fd35138e945c07b7664da3fba7f7df79bee7e57a44f0d4d62ecfd12e1f1b9
MD5 01de45d5e3726cf2e9223e1102427ec6
BLAKE2b-256 6f19d13d400cb4aa3abe3432252f7478b0faccdecec46693879cc4ec56166220

See more details on using hashes here.

Provenance

The following attestation bundles were made for docvortex-0.2.1.tar.gz:

Publisher: publish.yml on myhloli/DocVortex

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file docvortex-0.2.1-py3-none-any.whl.

File metadata

  • Download URL: docvortex-0.2.1-py3-none-any.whl
  • Upload date:
  • Size: 4.3 MB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for docvortex-0.2.1-py3-none-any.whl
Algorithm Hash digest
SHA256 61fd78d4126c537277a7a44e6eece5bcd1736ef72180996dd1f483fe7d8d247d
MD5 f7b22521698a38521e4253a7b573ad45
BLAKE2b-256 fcbd7ccd61015cf3fc1aec0a01d6deb79be4e8e0dceab05b7573f7e89d43e89c

See more details on using hashes here.

Provenance

The following attestation bundles were made for docvortex-0.2.1-py3-none-any.whl:

Publisher: publish.yml on myhloli/DocVortex

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.4.7

2 files

0.4.6

2 files

0.4.5

2 files

0.4.4

2 files

0.4.3

2 files

0.4.2

2 files

0.4.1

2 files

0.4.0

2 files

0.3.0

2 files

0.2.6

2 files

0.2.4

2 files

0.2.3

2 files

0.2.2

2 files

This release

0.2.1 This release

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page