Skip to main content

llm-markdownify

PyPI License CI

Turn PDFs, scans and images into clean Markdown with the vision model you already pay for. Built for RAG pipelines and AI agents.

pip install llm-markdownify
markdownify report.pdf -o report.md --model gpt-5.4-mini

Each page goes to a vision LLM with a prompt we measure on a public benchmark, so the Markdown comes back the way a retriever or an agent wants it:

  • the text exactly as printed, in reading order, including multi-column layouts
  • no running headers, footers or page numbers polluting your chunks
  • math as LaTeX ($...$, $$...$$)
  • tables as Markdown, or as HTML when they have merged cells, and stitched back together when they break across pages
  • charts as a description plus the data values, diagrams as Mermaid
  • old scans and handwriting transcribed, not summarized

It works with any provider: OpenAI, Anthropic, Gemini, DeepSeek, Azure, OpenRouter, or a model running on your own machine through Ollama, LM Studio or vLLM.

Handwritten notes converted to Markdown

How good is it?

We measure every change on olmOCR-Bench, the benchmark most PDF-to-Markdown tools publish. It checks 1,403 real pages with about 7,000 pass/fail tests: is this sentence present, is the page footer gone, is paragraph A before paragraph B, is this table cell next to that one, does this equation render the same.

llm-markdownify 0.5 with gpt-6-luna scores 81.6 ± 1.0 on the full benchmark.

System olmOCR-Bench Source
Chandra 2 85.9 Datalab (self-reported)
Mistral OCR 4 85.2 Mistral (self-reported)
olmOCR 2 (7B model trained for this benchmark) 82.4 Ai2
llm-markdownify 0.5 + gpt-6-luna 81.6 measured with evals/olmocr_bench
Marker 2 (balanced) 76.0 Datalab (self-reported)
Docling 50.3 Marker's benchmark

By category: tables 89.1, headers/footers 86.6, old scans with math 82.5, multi-column 81.4, arXiv math 79.4, tiny text 88.7, old scans 45.1. Old, faded scans are the weak spot.

What the library adds on top of the model, measured on a fixed 105-page stratified subset:

Same pages Score Median time per page
gpt-6-luna with a bare "convert this page to Markdown" prompt 72.6 11.9 s
llm-markdownify 0.4 + gpt-6-luna 72.5 (18 pages failed) 14.8 s
llm-markdownify 0.5 + gpt-5.4-mini 80.4 4.8 s
llm-markdownify 0.5 + gpt-6-luna 84.9 14.6 s

The biggest single difference is headers and footers. With the bare prompt, 24% of the checks that running headers, footers and page numbers are gone pass (88% with llm-markdownify). Left in, that furniture ends up in every RAG chunk. Other projects' numbers are what they published; ours come from running the benchmark's own scorer on our output.

Reproduce it yourself with evals/olmocr_bench. The harness scores the Markdown this library writes, through the same convert() call you would use.

Quickstart

export OPENAI_API_KEY="sk-..."

markdownify input.pdf -o output.md --model gpt-5.4-mini   # PDF
markdownify scan.png -o scan.md --model gpt-5.4-mini      # PNG, JPG, WEBP, TIFF (multi-page), BMP, GIF

From Python:

from llm_markdownify import convert

convert("input.pdf", "output.md", model="gpt-5.4-mini")

Use any provider

The model name decides where the request goes (via LiteLLM, 100+ providers). Set that provider's key and pick a vision-capable model.

Provider Key Example --model
OpenAI OPENAI_API_KEY gpt-5.4-mini
Anthropic ANTHROPIC_API_KEY anthropic/claude-sonnet-5
Google Gemini GEMINI_API_KEY gemini/gemini-2.5-flash
DeepSeek DEEPSEEK_API_KEY deepseek/deepseek-flash
OpenRouter OPENROUTER_API_KEY openrouter/anthropic/claude-sonnet-5
Azure OpenAI AZURE_API_KEY, AZURE_API_BASE, AZURE_API_VERSION azure/<deployment>

Anything that speaks the OpenAI API works too, including local servers with no key at all:

# Ollama, LM Studio, vLLM, llama.cpp, or a hosted OpenAI-compatible provider
markdownify input.pdf -o output.md --model openai/<model-name> --api-base http://localhost:11434/v1

Built for production

  • Safe to call from threads. Use convert() from a web server or a worker pool. (Retry, cache and rate-limit settings are currently process-wide, so give concurrent calls the same settings.)
  • Rate limits are waited out, real errors fail fast. Rate limits, timeouts and 5xx responses are retried with backoff. A bad key or a bad request fails on the first attempt with one clear line, and never prints your API key.
  • Oversized pages are handled. Page images are capped at 2048 px, below provider limits, so big scans don't get rejected.
  • Re-runs are free with --cache: responses are cached on disk and keyed on the exact request.
  • Throughput and cost controls: --concurrency, --rate-limit, --max-image-px, --reasoning-effort.

Options

Flag Default What it does
--model gpt-4.1-mini or $LLM_MARKDOWNIFY_MODEL any LiteLLM model name
--profile generic prompt profile: generic, contracts, or a JSON file with your own prompts
--dpi 200 PDF render resolution (ignored for images)
--max-image-px 2048 longest side of each page image sent to the model
--max-group-pages 3 max pages merged when a table or chart continues onto the next page
--no-grouping skip cross-page detection (one fewer model call per page)
--temperature, --max-tokens, --reasoning-effort provider defaults, 16000 generation settings
--api-base any OpenAI-compatible endpoint
--concurrency, --grouping-concurrency, --rate-limit 4, same, none throughput
--cache, --cache-dir off, ~/.cache/llm-markdownify response cache
-q, -v, --version quiet, verbose, version

Custom prompts: copy a built-in profile from prompt_profiles.py into a JSON file with the fields name, continuation_system, continuation_user, markdown_system and markdown_user, then pass --profile my_profile.json.

DOCX input goes through Microsoft Word (macOS/Windows): pip install "llm-markdownify[docx]" and pass --allow-docx. Exporting to PDF yourself is more reliable.

Examples

The gallery has about 80 inputs with their Markdown output: receipts, charts, handwriting, formulas, forms, screenshots, scene text.

Roadmap

Next up: an in-memory API that returns per-page results, stdout output and page ranges for agents, an MCP server, a hybrid mode that uses the PDF's own text layer to cut cost, and presets for small local models. Details and current numbers are in docs/ROADMAP.md.

Markdownify Cloud

To run Markdownify in production on your own infrastructure, with your own LLMs, or tuned for your documents, see markdownify.xyz. The cloud version adds features beyond the open-source library and comes with hands-on integration support.

Contributing

uv sync --all-extras --dev
uv run pytest
uv run ruff check src tests

See CONTRIBUTING.md. Maintainers and coding agents: start with AGENTS.md.

License

Apache 2.0. If you distribute this project, keep the LICENSE and NOTICE files intact, crediting the original author, Sethu Pavan Venkata Reddy Pastula.

Metadata

Release files for llm-markdownify 0.5.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for llm-markdownify 0.5.0
File Size Uploaded
llm_markdownify-0.5.0.tar.gz 27.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for llm-markdownify 0.5.0
File Interpreter ABI Platform
llm_markdownify-0.5.0-py3-none-any.whl Python 3 none any Details

Total release size: 61.2 kB

Release files / llm_markdownify-0.5.0.tar.gz

Download URL llm_markdownify-0.5.0.tar.gz
Size 27.1 kB
Tags Source
SHA-256 checksum
How to use checksums
3e022b60b6f583de754b1b94c27536c367b7d72757c918e47476cb3a2de2bdaa
BLAKE2b-256 checksum
How to use checksums
2f287be19c90b13464eb0dcf1bf669315168ba5661737e032fac6c6dedc69df7
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / llm_markdownify-0.5.0-py3-none-any.whl

Download URL llm_markdownify-0.5.0-py3-none-any.whl
Size 34.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
4dce758c8d4107a88b497247c7c971a48800231b29d7f3047e5e3b30c23f32f6
BLAKE2b-256 checksum
How to use checksums
7d08259b7fc1957e0aa8f97f18a2fae947ad35a35e467ccb965f355253a87cf4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release history Release notifications | RSS feed

1.0.0

2 release files

0.6.0

2 release files

This release

0.5.0 This release

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page