Skip to main content

document2md

test Documentation Status

Converts a PDF, or a set of scanned page images, into Markdown — any document, such as an edition of Mexico's official gazette (DOF, Diario Oficial de la Federación) — optionally cropped down to a single note. It's a wrapper around mineru for the OCR/layout analysis itself; document2md's own contribution is:

  • Keeping mineru's mineru-api server warm across a batch of documents, instead of paying its startup (and model-loading) cost once per document.
  • Stitching the OCR of a list of page images (several scanned pages of the same note) into one continuous Markdown document.
  • Rewriting the raw HTML tables mineru falls back to (rowspan/colspan) into Markdown tables, so the output is Markdown all the way through.
  • Cropping the result down to a single note, by locating its title and the next note's title in the OCR'd text — useful because a scanned page usually holds the tail of one note and the head of the next.

It was extracted from the LegalIA monorepo at commit e1f258c, where every commit of its earlier history (as packages/document2md, and as packages/dof2md before the rename) can still be read.

Install

pip install document2md

For development, from a clone of this repository:

pip install -e ".[test]"

Usage

CLI

document2md takes exactly one input source — a local PDF or a set of local page images — and converts it to Markdown. It never downloads anything itself; get the PDF first (e.g. dofjson.download_edicion_pdf for a whole DOF edition by date and edition, see the dofjson README), then convert it:

document2md --pdf edicion.pdf   # a local PDF

document2md --images pagina-1.jpg pagina-2.jpg \
    --filename out.md      # scanned pages, in order

--filename sets the output Markdown's name; with --pdf it defaults to the PDF's own name (edicion.pdf → edicion.md), but with --images it's required, since a set of images has no single name to derive one from. --outdir sets the output directory (default: output/).

Since one edition's PDF holds every note published that day, --titulo/--titulo-siguiente crop the resulting Markdown down to just one note — its own title, and the next note's title, as they appear in the gazette's own index:

document2md --pdf edicion.pdf \
    --titulo "ACUERDO por el que se..." \
    --titulo-siguiente "DECRETO por el que se..."

Title matching is fuzzy (OCR text rarely matches an index title exactly), so a match below --min-confidence (default 0.6) is treated as not found and the crop falls back to keeping more text rather than dropping content. Other flags:

  • --keep-pages — also keep the uncropped Markdown, as <outdir>/<pdf stem>.full.md.
  • --keep-mineru-output — keep mineru's own raw output (layout/model JSON, rendered PDFs...) in <outdir>/<pdf stem>_mineru/ instead of discarding it; useful when a conversion looks wrong and mineru's own read of the page is the first thing worth inspecting.

Python: batch conversion

Converting many documents in one run is where mineru's startup cost starts to matter. BatchConverter keeps a single mineru-api server warm across the whole batch instead of restarting it per document:

from document2md import BatchConverter

jobs = [
    ("a.pdf", "output", "a.md"),
    (["b-p1.jpg", "b-p2.jpg"], "output", "b.md"),
]

with BatchConverter() as convert:
    for path_or_paths, outdir, filename in jobs:
        convert(path_or_paths, outdir, filename)

Each call takes a single PDF path, or a list of image paths for a document spanning several scanned pages, and writes the result to outdir/filename. The same titulo/titulo_siguiente, min_confidence, keep_pages and keep_mineru_output options the CLI exposes are also its keyword arguments — see BatchConverter.__call__'s docstring for the full signature.

nota2md.legal_provisions accepts an already-__enter__'d BatchConverter as its own converter parameter, so a batch of DOF legal provisions can share the same warm server too.

Tests

pytest -v

Metadata

Release files for document2md 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for document2md 0.3.0
File Size Uploaded
document2md-0.3.0.tar.gz 28.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for document2md 0.3.0
File Interpreter ABI Platform
document2md-0.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 50.2 kB

Release files / document2md-0.3.0.tar.gz

Download URL document2md-0.3.0.tar.gz
Size 28.2 kB
Tags Source
SHA-256 checksum
How to use checksums
83cfc87c13d78e81da8c8ea8de2d03f24655e6a7f4763f6085e675d378838ed8
BLAKE2b-256 checksum
How to use checksums
122a6045b204da0fc8efb3833cd2cf2f05d71e8274808360fb2cb142329544a7
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.14

Release files / document2md-0.3.0-py3-none-any.whl

Download URL document2md-0.3.0-py3-none-any.whl
Size 22.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
7a38423bfa4e574de412719a7150204e351dd26d4634f1ec8e0d967f57e40ce7
BLAKE2b-256 checksum
How to use checksums
e86474d35da9bb6a4d45296e848747cbc8692f32771ecbea4a57ee32f6a7eb0e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.14

Release history Release notifications | RSS feed

This release

0.3.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page