Skip to main content

pdf-inspector

Fast PDF classification and text extraction. Detects whether a PDF is text-based or scanned, extracts text with position awareness, and converts to clean Markdown — all without OCR. Python bindings via PyO3 for the pdf-inspector Rust library.

Built by Firecrawl to handle text-based PDFs locally in under 200ms, skipping expensive OCR services for the ~54% of PDFs that don't need them.

Features

  • Smart classificationtext_based / scanned / image_based / mixed in ~10–50ms, with a confidence score and per-page OCR routing.
  • Markdown conversion — headings, lists, code blocks, bold/italic, URL linking, and dual-mode table detection (PDF drawing ops + text-alignment heuristics).
  • Layout-aware extraction — multi-column reading order, position and font info per text item, RTL support.
  • Robust text decoding — CID/Type0 fonts via ToUnicode CMaps, plus automatic flagging of broken encodings so callers can fall back to OCR.
  • Lightweight — native Rust core, no ML models, no external services; ships type stubs.

Benchmark

opendataloader-bench corpus (200 PDFs), local engines without model-based PDF parsing; OCR disabled. Scores 0–1, higher is better:

Engine Overall Reading order Tables (TEDS) Headings Speed
pdf-inspector 0.875 0.915 0.814 0.788 0.470s
liteparse 0.873 0.913 0.693 0.811 0.750s
opendataloader 0.831 0.902 0.489 0.739 2.569s
pymupdf4llm 0.735 0.886 0.401 0.424 17.117s
markitdown 0.589 0.844 0.273 0.000 16.165s

Refreshed July 31, 2026, on Apple M4 Pro; speed is the median of five complete corpus runs after an excluded warm-up. Full methodology and versions are in the repo README, with raw timings and artifacts in the results branch.

Install

pip install pdf-inspector

Prebuilt wheels cover CPython ≥3.8 on Linux (x86_64, aarch64), macOS (Intel, Apple Silicon), and Windows (x64). Other platforms build from source, which requires a Rust toolchain. For local development in a repo checkout:

pip install maturin
maturin develop --release

Usage

import pdf_inspector

# Full processing: detect + extract + convert to Markdown
result = pdf_inspector.process_pdf("document.pdf")
print(result.pdf_type)      # "text_based", "scanned", "image_based", "mixed"
print(result.confidence)     # 0.0 - 1.0
print(result.page_count)     # number of pages
print(result.markdown)       # Markdown string or None

# Process specific pages only
result = pdf_inspector.process_pdf("document.pdf", pages=[1, 3, 5])

# Process from bytes (no filesystem needed)
with open("document.pdf", "rb") as f:
    result = pdf_inspector.process_pdf_bytes(f.read())

# Fast detection only (no text extraction)
result = pdf_inspector.detect_pdf("document.pdf")
if result.pdf_type == "text_based":
    print("Can extract locally!")
else:
    print(f"Pages needing OCR: {result.pages_needing_ocr}")

# Plain text extraction
text = pdf_inspector.extract_text("document.pdf")

# Positioned text items with font info
items = pdf_inspector.extract_text_with_positions("document.pdf")
for item in items[:5]:
    print(f"'{item.text}' at ({item.x:.0f}, {item.y:.0f}) size={item.font_size}")

# Per-page markdown (one Markdown string per page, plus layout metadata)
result = pdf_inspector.extract_pages_markdown("document.pdf")
for page in result.pages:
    print(f"Page {page.page}: {len(page.markdown)} chars, needs_ocr={page.needs_ocr}")

# Restrict to specific 0-indexed pages (preserves caller order)
result = pdf_inspector.extract_pages_markdown("document.pdf", pages=[0, 2])

# Structure-tree elements from tagged PDFs (empty list when untagged).
# Pages are 1-indexed to match TextItem.page, so (page, mcid) joins directly
# against extract_text_with_positions — e.g. to recover real heading levels:
elements = pdf_inspector.extract_structure_elements("tagged.pdf")
roles = {(e.page, e.mcid): e.role for e in elements}
headings = [
    item.text
    for item in pdf_inspector.extract_text_with_positions("tagged.pdf")
    if item.mcid is not None and roles.get((item.page, item.mcid), "").startswith("H")
]

API reference

Function Description
process_pdf(path, pages=None) Full processing (detect + extract + markdown)
process_pdf_bytes(data, pages=None) Full processing from bytes
detect_pdf(path) Fast detection only (returns PdfResult)
detect_pdf_bytes(data) Fast detection from bytes
classify_pdf(path) Lightweight classification (returns PdfClassification)
classify_pdf_bytes(data) Lightweight classification from bytes
extract_text(path) Plain text extraction
extract_text_bytes(data) Plain text extraction from bytes
extract_text_with_positions(path, pages=None) Text with X/Y coords and font info
extract_text_with_positions_bytes(data, pages=None) Text with positions from bytes
extract_text_in_regions(path, page_regions) Extract text in bounding-box regions
extract_text_in_regions_bytes(data, page_regions) Region extraction from bytes
extract_pages_markdown(path, pages=None) Per-page Markdown + layout metadata (all pages by default)
extract_pages_markdown_bytes(data, pages=None) Per-page Markdown from bytes
extract_structure_elements(path, pages=None) Structure-tree elements from tagged PDFs (page, mcid, role)
extract_structure_elements_bytes(data, pages=None) Structure-tree elements from bytes

Types

Type stubs (pdf_inspector.pyi) ship with the package. Result types at a glance:

class PdfResult:                     # process_pdf / detect_pdf
    pdf_type: str                    # "text_based" | "scanned" | "image_based" | "mixed"
    markdown: str | None             # extracted Markdown (None for detect_pdf)
    page_count: int
    processing_time_ms: int
    pages_needing_ocr: list[int]     # 1-indexed
    ocr_reasons_by_page: list[PageOcrReasons]
    title: str | None
    confidence: float                # 0.0 - 1.0
    is_complex_layout: bool
    pages_with_tables: list[int]
    pages_with_columns: list[int]
    has_encoding_issues: bool        # broken font encodings — consider OCR fallback

class PageOcrReasons:                # per-page OCR diagnostics
    page: int                        # 1-indexed
    reasons: list[str]               # machine-readable reason identifiers

class PdfClassification:             # classify_pdf
    pdf_type: str
    page_count: int
    pages_needing_ocr: list[int]     # 0-indexed
    confidence: float

class TextItem:                      # extract_text_with_positions
    text: str
    x: float
    y: float
    width: float
    height: float
    font: str
    font_size: float
    page: int
    is_bold: bool
    is_italic: bool
    is_underline: bool
    is_strikeout: bool
    item_type: str
    mcid: int | None                 # marked-content ID for tagged PDFs (None otherwise)

class StructureElement:              # extract_structure_elements
    page: int                        # 1-indexed (matches TextItem.page)
    mcid: int
    role: str                        # "H1".."H6", "P", "Table", ... (resolved via /RoleMap)

class RegionText:                    # extract_text_in_regions
    text: str
    needs_ocr: bool
    ocr_reason: str | None           # machine-readable OCR reason

class PageRegionTexts:               # extract_text_in_regions
    page: int                        # 0-indexed
    regions: list[RegionText]

class PagesExtractionResult:         # extract_pages_markdown
    pages: list[PageMarkdown]        # PageMarkdown: page (0-indexed), markdown, needs_ocr, ocr_reason
    pages_with_tables: list[int]     # 1-indexed
    pages_with_columns: list[int]    # 1-indexed
    pages_needing_ocr: list[int]     # 1-indexed
    ocr_reasons_by_page: list[PageOcrReasons]
    is_complex: bool                 # any page has tables or multi-column layout

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

pdf_inspector-1.14.1.tar.gz (1.6 MB view details)

Uploaded Source

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

pdf_inspector-1.14.1-cp38-abi3-win_amd64.whl (2.7 MB view details)

Uploaded CPython 3.8+Windows x86-64

pdf_inspector-1.14.1-cp38-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (3.0 MB view details)

Uploaded CPython 3.8+manylinux: glibc 2.17+ x86-64

pdf_inspector-1.14.1-cp38-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl (2.9 MB view details)

Uploaded CPython 3.8+manylinux: glibc 2.17+ ARM64

pdf_inspector-1.14.1-cp38-abi3-macosx_11_0_arm64.whl (2.8 MB view details)

Uploaded CPython 3.8+macOS 11.0+ ARM64

pdf_inspector-1.14.1-cp38-abi3-macosx_10_12_x86_64.whl (2.9 MB view details)

Uploaded CPython 3.8+macOS 10.12+ x86-64

File details

Details for the file pdf_inspector-1.14.1.tar.gz.

File metadata

  • Download URL: pdf_inspector-1.14.1.tar.gz
  • Upload date:
  • Size: 1.6 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for pdf_inspector-1.14.1.tar.gz
Algorithm Hash digest
SHA256 e9d0fc1669ca77950d28488c8450d1aa28018e839ea260308d87253dfb8bec03
MD5 32ca25c77eb8783c9c1b2bff2ae1e581
BLAKE2b-256 cc7d29ac7865f4c7332d28638d08b73aa6b7bf2f1325c383094bb6717a1a816d

See more details on using hashes here.

Provenance

The following attestation bundles were made for pdf_inspector-1.14.1.tar.gz:

Publisher: publish-pypi.yml on firecrawl/pdf-inspector

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file pdf_inspector-1.14.1-cp38-abi3-win_amd64.whl.

File metadata

File hashes

Hashes for pdf_inspector-1.14.1-cp38-abi3-win_amd64.whl
Algorithm Hash digest
SHA256 d4a2455c316e15eac745c3836f72b641e1c33207afb84c250abff2843209e527
MD5 aa1b798bc9f7557c58f0ed8a8db3de11
BLAKE2b-256 247c4c84916a97cfd85d2e72dc3aaf6988cd19aff6e050d6a1b79d20f97e4566

See more details on using hashes here.

Provenance

The following attestation bundles were made for pdf_inspector-1.14.1-cp38-abi3-win_amd64.whl:

Publisher: publish-pypi.yml on firecrawl/pdf-inspector

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file pdf_inspector-1.14.1-cp38-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.

File metadata

File hashes

Hashes for pdf_inspector-1.14.1-cp38-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Algorithm Hash digest
SHA256 62e761a32ffd9c2b0c803dbefa2b0ad6f98bc26a75847571d99c3f7a65a7cea3
MD5 0f37d7ce17243a4338d50f812434c1bb
BLAKE2b-256 37cc589c355fed7b50098ddb9d09116c81d966c7b4ff78741183e600c1b4de6f

See more details on using hashes here.

Provenance

The following attestation bundles were made for pdf_inspector-1.14.1-cp38-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl:

Publisher: publish-pypi.yml on firecrawl/pdf-inspector

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file pdf_inspector-1.14.1-cp38-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl.

File metadata

File hashes

Hashes for pdf_inspector-1.14.1-cp38-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
Algorithm Hash digest
SHA256 501b8a6ffcd2c16ea74d71b9d7d70503c945c7d38c81151121790fd278021f0b
MD5 c4e0edb5e24255134ad3c3aae45474f0
BLAKE2b-256 cd365b332c80858407fb19b34e6f18f6c99022d04324caf98b8aaf4547044871

See more details on using hashes here.

Provenance

The following attestation bundles were made for pdf_inspector-1.14.1-cp38-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl:

Publisher: publish-pypi.yml on firecrawl/pdf-inspector

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file pdf_inspector-1.14.1-cp38-abi3-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for pdf_inspector-1.14.1-cp38-abi3-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 54cb32691640e5684832d21fbb58d2b30c068cbc055e0157da37774c1e8b9218
MD5 b74f39bf72cb337f510c91d11181d88a
BLAKE2b-256 982400594b8ae250c0b8a87631a86067896416920879c43c8cdcdfeb50586d49

See more details on using hashes here.

Provenance

The following attestation bundles were made for pdf_inspector-1.14.1-cp38-abi3-macosx_11_0_arm64.whl:

Publisher: publish-pypi.yml on firecrawl/pdf-inspector

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file pdf_inspector-1.14.1-cp38-abi3-macosx_10_12_x86_64.whl.

File metadata

File hashes

Hashes for pdf_inspector-1.14.1-cp38-abi3-macosx_10_12_x86_64.whl
Algorithm Hash digest
SHA256 b2159efa28f132df0442129a9b3eaf70c3f9915eb8202efcaaf759d85a01993f
MD5 00e464bbc9a3d2da87a0d8125dd8cf43
BLAKE2b-256 96477ddacbcac7a04ec60014e2a03b943c4643a8bd794ed67588f5ebadfbeb15

See more details on using hashes here.

Provenance

The following attestation bundles were made for pdf_inspector-1.14.1-cp38-abi3-macosx_10_12_x86_64.whl:

Publisher: publish-pypi.yml on firecrawl/pdf-inspector

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page