Skip to main content

pdf-inspector

Fast PDF classification and text extraction. Detects whether a PDF is text-based or scanned, extracts text with position awareness, and converts to clean Markdown — all without OCR. Python bindings via PyO3 for the pdf-inspector Rust library.

Built by Firecrawl to handle text-based PDFs locally in under 200ms, skipping expensive OCR services for the ~54% of PDFs that don't need them.

Features

  • Smart classificationtext_based / scanned / image_based / mixed in ~10–50ms, with a confidence score and per-page OCR routing.
  • Markdown conversion — headings, lists, code blocks, bold/italic, URL linking, and dual-mode table detection (PDF drawing ops + text-alignment heuristics).
  • Layout-aware extraction — multi-column reading order, position and font info per text item, RTL support.
  • Robust text decoding — CID/Type0 fonts via ToUnicode CMaps, plus automatic flagging of broken encodings so callers can fall back to OCR.
  • Lightweight — native Rust core, no ML models, no external services; ships type stubs.

Benchmark

opendataloader-bench corpus (200 PDFs), local engines without model-based PDF parsing; OCR disabled. Scores 0–1, higher is better:

Engine Overall Reading order Tables (TEDS) Headings Speed
pdf-inspector 0.875 0.915 0.814 0.788 0.470s
liteparse 0.873 0.913 0.693 0.811 0.750s
opendataloader 0.831 0.902 0.489 0.739 2.569s
pymupdf4llm 0.735 0.886 0.401 0.424 17.117s
markitdown 0.589 0.844 0.273 0.000 16.165s

Refreshed July 31, 2026, on Apple M4 Pro; speed is the median of five complete corpus runs after an excluded warm-up. Full methodology and versions are in the repo README, with raw timings and artifacts in the results branch.

Install

pip install pdf-inspector

Prebuilt wheels cover CPython ≥3.8 on Linux (x86_64, aarch64), macOS (Intel, Apple Silicon), and Windows (x64). Other platforms build from source, which requires a Rust toolchain. For local development in a repo checkout:

pip install maturin
maturin develop --release

Usage

import pdf_inspector

# Full processing: detect + extract + convert to Markdown
result = pdf_inspector.process_pdf("document.pdf")
print(result.pdf_type)      # "text_based", "scanned", "image_based", "mixed"
print(result.confidence)     # 0.0 - 1.0
print(result.page_count)     # number of pages
print(result.markdown)       # Markdown string or None

# Process specific pages only
result = pdf_inspector.process_pdf("document.pdf", pages=[1, 3, 5])

# Process from bytes (no filesystem needed)
with open("document.pdf", "rb") as f:
    result = pdf_inspector.process_pdf_bytes(f.read())

# Fast detection only (no text extraction)
result = pdf_inspector.detect_pdf("document.pdf")
if result.pdf_type == "text_based":
    print("Can extract locally!")
else:
    print(f"Pages needing OCR: {result.pages_needing_ocr}")

# Plain text extraction
text = pdf_inspector.extract_text("document.pdf")

# Positioned text items with font info
items = pdf_inspector.extract_text_with_positions("document.pdf")
for item in items[:5]:
    print(f"'{item.text}' at ({item.x:.0f}, {item.y:.0f}) size={item.font_size}")

# Per-page markdown (one Markdown string per page, plus layout metadata)
result = pdf_inspector.extract_pages_markdown("document.pdf")
for page in result.pages:
    print(f"Page {page.page}: {len(page.markdown)} chars, needs_ocr={page.needs_ocr}")

# Restrict to specific 0-indexed pages (preserves caller order)
result = pdf_inspector.extract_pages_markdown("document.pdf", pages=[0, 2])

API reference

Function Description
process_pdf(path, pages=None) Full processing (detect + extract + markdown)
process_pdf_bytes(data, pages=None) Full processing from bytes
detect_pdf(path) Fast detection only (returns PdfResult)
detect_pdf_bytes(data) Fast detection from bytes
classify_pdf(path) Lightweight classification (returns PdfClassification)
classify_pdf_bytes(data) Lightweight classification from bytes
extract_text(path) Plain text extraction
extract_text_bytes(data) Plain text extraction from bytes
extract_text_with_positions(path, pages=None) Text with X/Y coords and font info
extract_text_with_positions_bytes(data, pages=None) Text with positions from bytes
extract_text_in_regions(path, page_regions) Extract text in bounding-box regions
extract_text_in_regions_bytes(data, page_regions) Region extraction from bytes
extract_pages_markdown(path, pages=None) Per-page Markdown + layout metadata (all pages by default)
extract_pages_markdown_bytes(data, pages=None) Per-page Markdown from bytes

Types

Type stubs (pdf_inspector.pyi) ship with the package. Result types at a glance:

class PdfResult:                     # process_pdf / detect_pdf
    pdf_type: str                    # "text_based" | "scanned" | "image_based" | "mixed"
    markdown: str | None             # extracted Markdown (None for detect_pdf)
    page_count: int
    processing_time_ms: int
    pages_needing_ocr: list[int]     # 1-indexed
    ocr_reasons_by_page: list[PageOcrReasons]
    title: str | None
    confidence: float                # 0.0 - 1.0
    is_complex_layout: bool
    pages_with_tables: list[int]
    pages_with_columns: list[int]
    has_encoding_issues: bool        # broken font encodings — consider OCR fallback

class PageOcrReasons:                # per-page OCR diagnostics
    page: int                        # 1-indexed
    reasons: list[str]               # machine-readable reason identifiers

class PdfClassification:             # classify_pdf
    pdf_type: str
    page_count: int
    pages_needing_ocr: list[int]     # 0-indexed
    confidence: float

class TextItem:                      # extract_text_with_positions
    text: str
    x: float
    y: float
    width: float
    height: float
    font: str
    font_size: float
    page: int
    is_bold: bool
    is_italic: bool
    is_underline: bool
    is_strikeout: bool
    item_type: str

class RegionText:                    # extract_text_in_regions
    text: str
    needs_ocr: bool
    ocr_reason: str | None           # machine-readable OCR reason

class PageRegionTexts:               # extract_text_in_regions
    page: int                        # 0-indexed
    regions: list[RegionText]

class PagesExtractionResult:         # extract_pages_markdown
    pages: list[PageMarkdown]        # PageMarkdown: page (0-indexed), markdown, needs_ocr, ocr_reason
    pages_with_tables: list[int]     # 1-indexed
    pages_with_columns: list[int]    # 1-indexed
    pages_needing_ocr: list[int]     # 1-indexed
    ocr_reasons_by_page: list[PageOcrReasons]
    is_complex: bool                 # any page has tables or multi-column layout

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

pdf_inspector-1.14.0.tar.gz (1.6 MB view details)

Uploaded Source

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

pdf_inspector-1.14.0-cp38-abi3-win_amd64.whl (2.7 MB view details)

Uploaded CPython 3.8+Windows x86-64

pdf_inspector-1.14.0-cp38-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (3.0 MB view details)

Uploaded CPython 3.8+manylinux: glibc 2.17+ x86-64

pdf_inspector-1.14.0-cp38-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl (2.9 MB view details)

Uploaded CPython 3.8+manylinux: glibc 2.17+ ARM64

pdf_inspector-1.14.0-cp38-abi3-macosx_11_0_arm64.whl (2.8 MB view details)

Uploaded CPython 3.8+macOS 11.0+ ARM64

pdf_inspector-1.14.0-cp38-abi3-macosx_10_12_x86_64.whl (2.9 MB view details)

Uploaded CPython 3.8+macOS 10.12+ x86-64

File details

Details for the file pdf_inspector-1.14.0.tar.gz.

File metadata

  • Download URL: pdf_inspector-1.14.0.tar.gz
  • Upload date:
  • Size: 1.6 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for pdf_inspector-1.14.0.tar.gz
Algorithm Hash digest
SHA256 edbb5133d9d91c44247724c89e26c76b24b21610f664218b571d697d763758e9
MD5 6d0e4b342b25b306e2eae0330a96e6b4
BLAKE2b-256 039791f67dc709342d06bd5a44a3c8edafb5d8063c7a697bf0d20e685088f170

See more details on using hashes here.

Provenance

The following attestation bundles were made for pdf_inspector-1.14.0.tar.gz:

Publisher: publish-pypi.yml on firecrawl/pdf-inspector

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file pdf_inspector-1.14.0-cp38-abi3-win_amd64.whl.

File metadata

File hashes

Hashes for pdf_inspector-1.14.0-cp38-abi3-win_amd64.whl
Algorithm Hash digest
SHA256 6722bb57bd9b6ae8f299508f7f711262785b693f3f41edb962a1d8235f8352d1
MD5 e6ba825da180a243327e8a0a319f47c0
BLAKE2b-256 ff9d21934abcca2e82ac6771e919e13f4487d6dfefaf7c26f88e7fed909989f6

See more details on using hashes here.

Provenance

The following attestation bundles were made for pdf_inspector-1.14.0-cp38-abi3-win_amd64.whl:

Publisher: publish-pypi.yml on firecrawl/pdf-inspector

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file pdf_inspector-1.14.0-cp38-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.

File metadata

File hashes

Hashes for pdf_inspector-1.14.0-cp38-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Algorithm Hash digest
SHA256 df0470b38081d52bb110887cce7e0df3ef9f6af0389d47ea397386150e879c94
MD5 1a1e8195bd0b3e5c3928f1d2cfa2ef5f
BLAKE2b-256 83ad5410b76e53d664648fa37ccd4ee6f1487606f7a39e974f6344501a78e3a3

See more details on using hashes here.

Provenance

The following attestation bundles were made for pdf_inspector-1.14.0-cp38-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl:

Publisher: publish-pypi.yml on firecrawl/pdf-inspector

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file pdf_inspector-1.14.0-cp38-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl.

File metadata

File hashes

Hashes for pdf_inspector-1.14.0-cp38-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
Algorithm Hash digest
SHA256 80e1aef1d8c511c0a29c5739f6c5ab19f084c1c87bd629eb85ea1104f34101e8
MD5 fd7942247907e633d8e7208a55f8fa54
BLAKE2b-256 d1e0c488f6d186e4b4e8dd7de705027d42e73ad574208496e15d5f3428ad8493

See more details on using hashes here.

Provenance

The following attestation bundles were made for pdf_inspector-1.14.0-cp38-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl:

Publisher: publish-pypi.yml on firecrawl/pdf-inspector

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file pdf_inspector-1.14.0-cp38-abi3-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for pdf_inspector-1.14.0-cp38-abi3-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 3d6734bcbad3a94e4ee79e5526c305cf71ca6f25f42fd2eddc28a81298cc86cd
MD5 9b47057dde8780d976548f59593e6cd6
BLAKE2b-256 794af8316888c948c7bbbbb04ccd97bc5d9f70dc2d87d8f0a20b5c68f32af5b0

See more details on using hashes here.

Provenance

The following attestation bundles were made for pdf_inspector-1.14.0-cp38-abi3-macosx_11_0_arm64.whl:

Publisher: publish-pypi.yml on firecrawl/pdf-inspector

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file pdf_inspector-1.14.0-cp38-abi3-macosx_10_12_x86_64.whl.

File metadata

File hashes

Hashes for pdf_inspector-1.14.0-cp38-abi3-macosx_10_12_x86_64.whl
Algorithm Hash digest
SHA256 7c586fb727bae496e5333f7c64e6ae203a46cbc3b46bf334bf4952f77a9e8603
MD5 1517f9f5b119f7b0290cc6f6dd9632b5
BLAKE2b-256 b7381326dff5ef40431aa83391a81ffa53c05dcac0f81587289c5f635e185827

See more details on using hashes here.

Provenance

The following attestation bundles were made for pdf_inspector-1.14.0-cp38-abi3-macosx_10_12_x86_64.whl:

Publisher: publish-pypi.yml on firecrawl/pdf-inspector

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page