Skip to main content

llm-markdownify

PyPI License CI

Turn PDFs, scans and images into clean Markdown with the vision model you already pay for. Built for RAG pipelines and AI agents.

pip install llm-markdownify
markdownify report.pdf -o report.md --model gpt-5.4-mini

Each page goes to a vision LLM with a prompt we measure on a public benchmark, so the Markdown comes back the way a retriever or an agent wants it:

  • the text exactly as printed, in reading order, including multi-column layouts
  • no running headers, footers or page numbers polluting your chunks
  • math as LaTeX ($...$, $$...$$)
  • tables as Markdown, or as HTML when they have merged cells, and stitched back together when they break across pages
  • charts as a description plus the data values, diagrams as Mermaid
  • old scans and handwriting transcribed, not summarized

It works with any provider: OpenAI, Anthropic, Gemini, DeepSeek, Azure, OpenRouter, or a model running on your own machine through Ollama, LM Studio or vLLM.

Handwritten notes converted to Markdown

How good is it?

We measure every change on olmOCR-Bench, the benchmark most PDF-to-Markdown tools publish. It checks 1,403 real pages with about 7,000 pass/fail tests: is this sentence present, is the page footer gone, is paragraph A before paragraph B, is this table cell next to that one, does this equation render the same.

llm-markdownify 0.5 with gpt-6-luna scores 81.6 ± 1.0 on the full benchmark.

System olmOCR-Bench Source
Chandra 2 85.9 Datalab (self-reported)
Mistral OCR 4 85.2 Mistral (self-reported)
olmOCR 2 (7B model trained for this benchmark) 82.4 Ai2
llm-markdownify 0.5 + gpt-6-luna 81.6 measured with evals/olmocr_bench
Marker 2 (balanced) 76.0 Datalab (self-reported)
Docling 50.3 Marker's benchmark

By category: tables 89.1, headers/footers 86.6, old scans with math 82.5, multi-column 81.4, arXiv math 79.4, tiny text 88.7, old scans 45.1. Old, faded scans are the weak spot.

What the library adds on top of the model, measured on a fixed 105-page stratified subset:

Same pages Score Median time per page
gpt-6-luna with a bare "convert this page to Markdown" prompt 72.6 11.9 s
llm-markdownify 0.4 + gpt-6-luna 72.5 (18 pages failed) 14.8 s
llm-markdownify 0.5 + gpt-5.4-mini 80.4 4.8 s
llm-markdownify 0.5 + gpt-6-luna 84.9 14.6 s

The biggest single difference is headers and footers. With the bare prompt, 24% of the checks that running headers, footers and page numbers are gone pass (88% with llm-markdownify). Left in, that furniture ends up in every RAG chunk. Other projects' numbers are what they published; ours come from running the benchmark's own scorer on our output.

Reproduce it yourself with evals/olmocr_bench. The harness scores the Markdown this library writes, through the same convert() call you would use.

Quickstart

export OPENAI_API_KEY="sk-..."

markdownify input.pdf -o output.md --model gpt-5.4-mini   # PDF
markdownify scan.png -o scan.md --model gpt-5.4-mini      # PNG, JPG, WEBP, TIFF (multi-page), BMP, GIF

From Python:

from llm_markdownify import convert

convert("input.pdf", "output.md", model="gpt-5.4-mini")

Use any provider

The model name decides where the request goes (via LiteLLM, 100+ providers). Set that provider's key and pick a vision-capable model.

Provider Key Example --model
OpenAI OPENAI_API_KEY gpt-5.4-mini
Anthropic ANTHROPIC_API_KEY anthropic/claude-sonnet-5
Google Gemini GEMINI_API_KEY gemini/gemini-2.5-flash
DeepSeek DEEPSEEK_API_KEY deepseek/deepseek-flash
OpenRouter OPENROUTER_API_KEY openrouter/anthropic/claude-sonnet-5
Azure OpenAI AZURE_API_KEY, AZURE_API_BASE, AZURE_API_VERSION azure/<deployment>

Anything that speaks the OpenAI API works too, including local servers with no key at all:

# Ollama, LM Studio, vLLM, llama.cpp, or a hosted OpenAI-compatible provider
markdownify input.pdf -o output.md --model openai/<model-name> --api-base http://localhost:11434/v1

Large jobs, low latency

Thousands of documents and no one waiting? markdownify-batch sends them through the OpenAI Batch API or Anthropic Message Batches at about half the price, finished within 24 hours:

markdownify-batch submit ./archive -o ./archive-md --model gpt-5.4-mini   # or anthropic/claude-opus-5
markdownify-batch status ./archive-md
markdownify-batch collect ./archive-md --retry-failed

Someone waiting on the answer? Use a fast model and skip the cross-page check:

markdownify doc.pdf -o doc.md --model gpt-5.4-mini --no-grouping --concurrency 16   # ~5 s per page

No API at all? Serve a vision model locally with LM Studio, Ollama, vLLM or llama.cpp and pass --api-base. Measured speed and quality trade-offs, and the local setup guide, are in docs/speed-and-cost.md.

Built for production

  • Safe to call from threads. Use convert() from a web server or a worker pool. (Retry, cache and rate-limit settings are currently process-wide, so give concurrent calls the same settings.)
  • Rate limits are waited out, real errors fail fast. Rate limits, timeouts and 5xx responses are retried with backoff. A bad key or a bad request fails on the first attempt with one clear line, and never prints your API key.
  • Oversized pages are handled. Page images are capped at 2048 px, below provider limits, so big scans don't get rejected.
  • Re-runs are free with --cache: responses are cached on disk and keyed on the exact request.
  • Throughput and cost controls: --concurrency, --rate-limit, --max-image-px, --reasoning-effort.

Options

Flag Default What it does
--model gpt-4.1-mini or $LLM_MARKDOWNIFY_MODEL any LiteLLM model name
--profile generic prompt profile: generic, contracts, or a JSON file with your own prompts
--dpi 200 PDF render resolution (ignored for images)
--max-image-px 2048 longest side of each page image sent to the model
--max-group-pages 3 max pages merged when a table or chart continues onto the next page
--no-grouping skip cross-page detection (one fewer model call per page)
--temperature, --max-tokens, --reasoning-effort provider defaults, 16000 generation settings
--api-base any OpenAI-compatible endpoint
--concurrency, --grouping-concurrency, --rate-limit 4, same, none throughput
--cache, --cache-dir off, ~/.cache/llm-markdownify response cache
-q, -v, --version quiet, verbose, version

Custom prompts: copy a built-in profile from prompt_profiles.py into a JSON file with the fields name, continuation_system, continuation_user, markdown_system and markdown_user, then pass --profile my_profile.json.

DOCX input goes through Microsoft Word (macOS/Windows): pip install "llm-markdownify[docx]" and pass --allow-docx. Exporting to PDF yourself is more reliable.

Examples

The gallery has about 80 inputs with their Markdown output: receipts, charts, handwriting, formulas, forms, screenshots, scene text.

Roadmap

Next up: an in-memory API that returns per-page results, stdout output and page ranges for agents, an MCP server, a hybrid mode that uses the PDF's own text layer to cut cost, and presets for small local models. Details and current numbers are in docs/ROADMAP.md.

Markdownify Cloud

To run Markdownify in production on your own infrastructure, with your own LLMs, or tuned for your documents, see markdownify.xyz. The cloud version adds features beyond the open-source library and comes with hands-on integration support.

Contributing

uv sync --all-extras --dev
uv run pytest
uv run ruff check src tests

See CONTRIBUTING.md. Maintainers and coding agents: start with AGENTS.md.

License

Apache 2.0. If you distribute this project, keep the LICENSE and NOTICE files intact, crediting the original author, Sethu Pavan Venkata Reddy Pastula.

Metadata

Release files for llm-markdownify 0.6.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for llm-markdownify 0.6.0
File Size Uploaded
llm_markdownify-0.6.0.tar.gz 37.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for llm-markdownify 0.6.0
File Interpreter ABI Platform
llm_markdownify-0.6.0-py3-none-any.whl Python 3 none any Details

Total release size: 83.3 kB

Release files / llm_markdownify-0.6.0.tar.gz

Download URL llm_markdownify-0.6.0.tar.gz
Size 37.5 kB
Tags Source
SHA-256 checksum
How to use checksums
f4ba317f4402e22d0f6a9c4c64ac1191a623ee8c25399d7dea565309a085af8a
BLAKE2b-256 checksum
How to use checksums
5c07cf3f2e4a891bd6bd7d391d88eae35289b69d92b738e46a7c715869decb5e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / llm_markdownify-0.6.0-py3-none-any.whl

Download URL llm_markdownify-0.6.0-py3-none-any.whl
Size 45.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
ea1a9250f5cc550706de67a68b154e3b09f00a6a4186653366eabd317ef7b706
BLAKE2b-256 checksum
How to use checksums
5fe4b79c4827a3b247ba881fdc694622af6ff82b722947c1426c7f35dcaa2aee
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release history Release notifications | RSS feed

1.0.0

2 release files

This release

0.6.0 This release

2 release files

0.5.0

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page