Skip to main content

pagewise-pdf-extractor

Page-wise PDF to Markdown extraction with text extraction, OCR, LLM fallback, and progress metadata.

pagewise-pdf-extractor is a Python package and CLI for converting PDFs into deterministic page-level Markdown files. It routes each page through embedded-text extraction, scanned-page OCR, and optional vision-model fallback, then returns structured results for RAG and document-processing pipelines.

What It Does

  • Extracts text-native PDF pages with PyMuPDF.
  • Extracts scanned/image pages with Marker.
  • Rejects corrupt or structurally unreliable embedded text before provider routing.
  • Detects and splits two-up scanned pages before OCR.
  • Falls back to Ollama vision OCR when configured OCR fails.
  • Writes one UTF-8 Markdown file per page.
  • Writes atomic progress.json with provider attempts, status, config hash, source hash, and page metadata.
  • Exposes a library API for applications and a CLI for operators.
  • Keeps local processing as the default; remote services are only used if explicitly configured.

Status

v0.3.0 is the current documented release. The public API is intended for early downstream use by applications that need page-wise PDF extraction, but the project is still pre-1.0.

Install

From PyPI after publication:

python -m pip install pagewise-pdf-extractor

Pinned Git dependency:

pagewise-pdf-extractor @ git+https://github.com/ebmurha/pagewise-pdf-extractor.git@v0.3.0

Local development:

python -m pip install -e D:\Developer\Projects\pagewise-pdf-extractor

Runtime dependencies are declared in pyproject.toml. OCR providers also require local binaries:

  • marker_single for Marker OCR
  • ollama for Ollama fallback
  • pdftoppm for rendering pages passed to Ollama

Check the local environment:

pagewise-pdf-extractor --validate-environment

Quickstart

CLI:

pagewise-pdf-extractor document.pdf --output-root output

Python:

from pathlib import Path

from pagewise_pdf_extractor import ExtractionConfig, process_pdf, validate_environment

config = ExtractionConfig(
    text_provider="pymupdf",
    ocr_provider="marker",
    fallback_provider="ollama",
    fallback_enabled=True,
    ollama_model="deepseek-ocr",
)

report = validate_environment(config)
if report.has_fatal_errors:
    raise RuntimeError(report.summary)

result = process_pdf(
    input_pdf=Path("document.pdf"),
    output_root=Path("output"),
    config=config,
)

Public import contract:

from pagewise_pdf_extractor import (
    ExtractionConfig,
    ExtractionResult,
    process_pdf,
    validate_environment,
)

Output

Default layout:

output/
  <input_sha256>/
    page_0001.md
    page_0002.md
    progress.json

Successful page:

# Page N

<provider markdown content>

Failed page:

# Page N

OCR FAILED

Error: <error_message>

Provider Routing

Default page-level routing:

  1. Try embedded text extraction with PyMuPDF.
  2. Validate embedded text for corrupt glyphs, abnormal spacing, and table-heavy visual layout.
  3. Accept embedded text when it meets configured length and structural quality thresholds.
  4. Render Marker input at 350 DPI and split confidently detected two-up pages into logical pages.
  5. Use Marker OCR when embedded text is absent, low quality, or force_ocr=True.
  6. Use Ollama fallback when Marker fails or returns unusable output and fallback is enabled.
  7. Write failure Markdown if all configured providers fail.

Documentation

Tests

python -m unittest discover -s tests -p "test_*.py"
python -c "from pagewise_pdf_extractor import ExtractionConfig, process_pdf, validate_environment; print('ok')"
pagewise-pdf-extractor --help

Future Goals

Future work to preserve the public API, keep provider behavior explicit, and add new providers or extraction quality improvements behind documented configuration.

Release files for pagewise-pdf-extractor 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for pagewise-pdf-extractor 0.3.0
File Size Uploaded
pagewise_pdf_extractor-0.3.0.tar.gz 30.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for pagewise-pdf-extractor 0.3.0
File Interpreter ABI Platform
pagewise_pdf_extractor-0.3.0-py3-none-any.whl Python 3 none any Details

Total release size:62.2 kB

Release files / pagewise_pdf_extractor-0.3.0.tar.gz

Download URL pagewise_pdf_extractor-0.3.0.tar.gz
Size 30.6 kB
Tags Source
SHA-256 checksum
How to use checksums
1e2b26e4213e53c5b9ed2bd6565fafa250f396c3e5cc5c1411d5739c146b1e5c
BLAKE2b-256 checksum
How to use checksums
74a95eaafc708875e5c29a4dfbb019ac2133d2f88fcf399f0c90332fa62c1633
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.13.12

Release files / pagewise_pdf_extractor-0.3.0-py3-none-any.whl

Download URL pagewise_pdf_extractor-0.3.0-py3-none-any.whl
Size 31.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
b972d4144372a12b12c110dce2e84c21fe88cf8f629c0bc8654788955629bdf9
BLAKE2b-256 checksum
How to use checksums
b474ecdf8b731736d45e8c5c0bb17daa584ebfc9ea4c7c44ba7da1ffd80baf76
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.13.12

Release history Release notifications | RSS feed

This release

0.3.0 This release

2 release files

0.2.0

2 release files

0.1.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page