Skip to main content

doc-page-extractor

Document page extraction tool that converts page images into text layouts with pixel coordinates.

The package provides local Hugging Face OCR backends and vendor OCR adapters that all return the same page layout shape.

Installation

Default installation supports vendor OCR backends and the common extraction pipeline without installing Hugging Face model runtime dependencies:

pip install doc-page-extractor

Local Hugging Face OCR backends require the local extra and CUDA PyTorch. Install PyTorch for your CUDA version first, then install the local runtime:

# CUDA 12.1
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu121

# CUDA 11.8
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu118

# CUDA 12.6
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu126

pip install "doc-page-extractor[local]"

The package does not declare a CUDA-specific PyTorch wheel. Vendor-only users, including macOS users, should use the default install.

Backends

DeepSeek OCR Local

Use this backend for local DeepSeek OCR or DeepSeek OCR 2 inference:

from doc_page_extractor import create_deepseek_ocr_page_extractor

extractor = create_deepseek_ocr_page_extractor(
    ocr_model="deepseek-ocr",
    model_path="models-cache",
    local_only=True,
)

extractor2 = create_deepseek_ocr_page_extractor(
    ocr_model="deepseek-ocr2",
    model_path="models-cache",
    local_only=True,
)

Install the local runtime dependencies before using this backend. See Installation.

Check CUDA with:

nvidia-smi
python -c "import torch; print(torch.cuda.is_available())"

Unlimited OCR Local

Use this backend for local Unlimited OCR Transformers inference:

from doc_page_extractor import create_unlimited_ocr_page_extractor

extractor = create_unlimited_ocr_page_extractor(
    model_path="models-cache",
    local_only=True,
)

The local Unlimited OCR backend uses the Hugging Face model baidu/Unlimited-OCR. Single-page local inference supports the base and gundam size presets.

DeepSeek OCR Vendor

Use this backend when DeepSeek OCR is exposed through an OpenAI-compatible endpoint:

from doc_page_extractor import (
    DeepSeekOCRVendorConfig,
    create_deepseek_ocr_vendor_page_extractor,
)

extractor = create_deepseek_ocr_vendor_page_extractor(
    DeepSeekOCRVendorConfig(
        base_url="https://example.test/openai",
        api_key="...",
        model="deepseek-ocr",
    )
)

The package does not read environment variables automatically. .env.template is only for local debugging scripts.

DEEPSEEK_OCR_BASE_URL=
DEEPSEEK_OCR_API_KEY=
DEEPSEEK_OCR_MODEL=deepseek-ocr
DEEPSEEK_OCR_TEMPERATURE=0.0
DEEPSEEK_OCR_TOP_P=0.7
DEEPSEEK_OCR_MAX_TOKENS=8000
DEEPSEEK_OCR_TIMEOUT_SECONDS=180

DeepSeek OCR 2 Vendor

Use this backend for DeepSeek OCR 2 through an OpenAI-compatible endpoint:

from doc_page_extractor import (
    DeepSeekOCR2VendorConfig,
    create_deepseek_ocr2_vendor_page_extractor,
)

extractor = create_deepseek_ocr2_vendor_page_extractor(
    DeepSeekOCR2VendorConfig(
        base_url="https://example.test/openai",
        api_key="...",
        model="deepseek-ocr2",
    )
)

The package does not read environment variables automatically. .env.template is only for local debugging scripts.

DEEPSEEK_OCR2_BASE_URL=
DEEPSEEK_OCR2_API_KEY=
DEEPSEEK_OCR2_MODEL=
DEEPSEEK_OCR2_TEMPERATURE=0.0
DEEPSEEK_OCR2_TOP_P=0.7
DEEPSEEK_OCR2_MAX_TOKENS=8000
DEEPSEEK_OCR2_TIMEOUT_SECONDS=180

Unlimited OCR Vendor

Use this backend for Baidu Cloud Unlimited OCR:

from doc_page_extractor import (
    UnlimitedOCRVendorConfig,
    create_unlimited_ocr_vendor_page_extractor,
)

extractor = create_unlimited_ocr_vendor_page_extractor(
    UnlimitedOCRVendorConfig(
        ak="...",
        sk="...",
    )
)

The package does not read environment variables automatically. .env.template is only for local debugging scripts.

UNLIMITED_OCR_ACCESS_KEY=
UNLIMITED_OCR_SECRET_KEY=
UNLIMITED_OCR_BASE_URL=https://aip.baidubce.com
UNLIMITED_OCR_POLL_INTERVAL_SECONDS=2
UNLIMITED_OCR_TIMEOUT_SECONDS=180

Unlimited OCR Vendor images with a side longer than 8192 px are resized proportionally before upload. Returned layout coordinates are mapped back to the original image size.

Vendor request errors

Vendor HTTP, transport, and provider response errors are raised as VendorOCRRequestError. The original requests exception is preserved as the exception's __cause__, including its response when one is available. The package does not classify or retry these errors; callers can inspect the cause and apply their own retry and fallback policy.

import requests

from doc_page_extractor import VendorOCRRequestError

try:
    next(extractor.extract_page_results(image, size="gundam"))
except VendorOCRRequestError as error:
    cause = error.__cause__
    if isinstance(cause, requests.RequestException) and cause.response is not None:
        print(cause.response.status_code, cause.response.text)

Extraction

All backends return the same PageExtractor shape:

from PIL import Image
from doc_page_extractor import ExtractionContext

context = ExtractionContext(check_aborted=lambda: False)

for page_image, result in extractor.extract_page_results(
    image=Image.open("page.png"),
    size="gundam",
    stages=1,
    context=context,
):
    for layout in result.layouts:
        print(layout.kind, layout.det, layout.text)

Layout.kind is the stable layout semantic. Adapter metadata remains available through optional fields such as type, polygon, html, source, and raw.

Structured page blocks are available on each OCRPageResult:

from doc_page_extractor import LayoutKind

for page_image, result in extractor.extract_page_results(
    image=Image.open("page.png"),
    size="gundam",
    stages=1,
    context=context,
):
    if result.structured is None:
        continue
    for block in result.structured.blocks:
        if block.kind == LayoutKind.TABLE:
            print(block.html)

The structured model groups asset captions with images, tables, and equations when possible. DeepSeek output is structured from flat OCR tags; DeepSeek OCR 2 output is structured from line blocks; Unlimited OCR local output is structured from local detection tags; Unlimited OCR Vendor output is normalized from richer layout JSON into the same public kinds.

Unlimited OCR extracts footnotes directly. If stages > 1 is requested with an Unlimited OCR adapter, the extractor emits a warning and runs a single stage because DeepSeek-style multi-stage redaction can erase footnote regions.

Development

For contributors and developers, see Development Guide.

Useful local commands:

poetry run python test.py
poetry run pylint --disable=import-error doc_page_extractor
poetry run python scripts/ocr_sample.py --adapter unlimited-ocr-vendor --image tests/images/friendly-title.png

scripts/ocr_sample.py also supports deepseek-ocr-local, deepseek-ocr2-local, unlimited-ocr-local, and all; all includes the local CUDA-backed modes.

Requirements

  • Python >= 3.10, < 3.14
  • CUDA-capable NVIDIA GPU only when using local Hugging Face OCR backends
  • Remote OCR credentials only when using vendor OCR backends

Dependencies & Licenses

This project is licensed under the MIT License. Local Hugging Face OCR backends depend on their upstream model code and runtime dependencies. The DeepSeek-OCR model uses easydict (LGPLv3) for configuration management.

Metadata

Release files for doc-page-extractor 1.2.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for doc-page-extractor 1.2.1
File Size Uploaded
doc_page_extractor-1.2.1.tar.gz 24.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for doc-page-extractor 1.2.1
File Interpreter ABI Platform
doc_page_extractor-1.2.1-py3-none-any.whl Python 3 none any Details

Total release size: 53.0 kB

Release files / doc_page_extractor-1.2.1.tar.gz

Download URL doc_page_extractor-1.2.1.tar.gz
Size 24.2 kB
Tags Source
SHA-256 checksum
How to use checksums
326463c9a42602e99990fc0d2aa43179d34f9147c18a5194248555f492ad0071
BLAKE2b-256 checksum
How to use checksums
a6db689893b7f46f3e62789f1ad7620d6d40ef94ea79f372193c8693489034c5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via poetry/2.1.3 CPython/3.14.6 Darwin/25.6.0

Release files / doc_page_extractor-1.2.1-py3-none-any.whl

Download URL doc_page_extractor-1.2.1-py3-none-any.whl
Size 28.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
ada8a1326b28c652018604fadf0621c8f416e51811b7c0c0725fc9e34fa2955e
BLAKE2b-256 checksum
How to use checksums
cc061bc71d1d238931797b2ae4d73fbfc4a0c660160da933c848992bde378f34
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via poetry/2.1.3 CPython/3.14.6 Darwin/25.6.0

Release history Release notifications | RSS feed

This release

1.2.1 This release

2 release files

1.2.0

2 release files

1.1.1

2 release files

1.1.0

2 release files

1.0.12

2 release files

1.0.11

2 release files

1.0.10

2 release files

1.0.9

2 release files

1.0.8

2 release files

1.0.7

2 release files

1.0.6

2 release files

1.0.5

2 release files

1.0.4

2 release files

1.0.3

2 release files

1.0.2

2 release files

1.0.1

2 release files

1.0.0

2 release files

0.2.4

2 release files

0.2.3

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

0.0.10

2 release files

0.0.9

2 release files

0.0.8

2 release files

0.0.7

2 release files

0.0.6

2 release files

0.0.5

2 release files

0.0.2

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page