Fast PDF inspection, classification, and text extraction with smart scanned vs text-based detection
Project description
pdf-inspector
Fast PDF classification and text extraction. Detects whether a PDF is text-based or scanned, extracts text with position awareness, and converts to clean Markdown — all without OCR. Python bindings via PyO3 for the pdf-inspector Rust library.
Built by Firecrawl to handle text-based PDFs locally in under 200ms, skipping expensive OCR services for the ~54% of PDFs that don't need them.
Features
- Smart classification —
text_based/scanned/image_based/mixedin ~10–50ms, with a confidence score and per-page OCR routing. - Markdown conversion — headings, lists, code blocks, bold/italic, URL linking, and dual-mode table detection (PDF drawing ops + text-alignment heuristics).
- Layout-aware extraction — multi-column reading order, position and font info per text item, RTL support.
- Robust text decoding — CID/Type0 fonts via ToUnicode CMaps, plus automatic flagging of broken encodings so callers can fall back to OCR.
- Lightweight — native Rust core, no ML models, no external services; ships type stubs.
Benchmark
opendataloader-bench corpus (200 PDFs), local engines without model-based PDF parsing; OCR disabled. Scores 0–1, higher is better:
| Engine | Overall | Reading order | Tables (TEDS) | Headings | Speed |
|---|---|---|---|---|---|
| pdf-inspector | 0.875 | 0.915 | 0.814 | 0.788 | 2.8s |
| liteparse | 0.870 | 0.908 | 0.693 | 0.811 | 13.9s |
| opendataloader | 0.843 | 0.912 | 0.489 | 0.760 | 9.8s |
| pymupdf4llm | 0.735 | 0.886 | 0.401 | 0.424 | 15.5s |
| markitdown | 0.583 | 0.879 | 0.000 | 0.000 | 6.7s |
Refreshed July 16, 2026, on Apple M4 Pro; speed is the median of three complete corpus runs. Full methodology and versions are in the repo README.
Install
pip install pdf-inspector
Prebuilt wheels cover CPython ≥3.8 on Linux (x86_64, aarch64), macOS (Intel, Apple Silicon), and Windows (x64). Other platforms build from source, which requires a Rust toolchain. For local development in a repo checkout:
pip install maturin
maturin develop --release
Usage
import pdf_inspector
# Full processing: detect + extract + convert to Markdown
result = pdf_inspector.process_pdf("document.pdf")
print(result.pdf_type) # "text_based", "scanned", "image_based", "mixed"
print(result.confidence) # 0.0 - 1.0
print(result.page_count) # number of pages
print(result.markdown) # Markdown string or None
# Process specific pages only
result = pdf_inspector.process_pdf("document.pdf", pages=[1, 3, 5])
# Process from bytes (no filesystem needed)
with open("document.pdf", "rb") as f:
result = pdf_inspector.process_pdf_bytes(f.read())
# Fast detection only (no text extraction)
result = pdf_inspector.detect_pdf("document.pdf")
if result.pdf_type == "text_based":
print("Can extract locally!")
else:
print(f"Pages needing OCR: {result.pages_needing_ocr}")
# Plain text extraction
text = pdf_inspector.extract_text("document.pdf")
# Positioned text items with font info
items = pdf_inspector.extract_text_with_positions("document.pdf")
for item in items[:5]:
print(f"'{item.text}' at ({item.x:.0f}, {item.y:.0f}) size={item.font_size}")
# Per-page markdown (one Markdown string per page, plus layout metadata)
result = pdf_inspector.extract_pages_markdown("document.pdf")
for page in result.pages:
print(f"Page {page.page}: {len(page.markdown)} chars, needs_ocr={page.needs_ocr}")
# Restrict to specific 0-indexed pages (preserves caller order)
result = pdf_inspector.extract_pages_markdown("document.pdf", pages=[0, 2])
API reference
| Function | Description |
|---|---|
process_pdf(path, pages=None) |
Full processing (detect + extract + markdown) |
process_pdf_bytes(data, pages=None) |
Full processing from bytes |
detect_pdf(path) |
Fast detection only (returns PdfResult) |
detect_pdf_bytes(data) |
Fast detection from bytes |
classify_pdf(path) |
Lightweight classification (returns PdfClassification) |
classify_pdf_bytes(data) |
Lightweight classification from bytes |
extract_text(path) |
Plain text extraction |
extract_text_bytes(data) |
Plain text extraction from bytes |
extract_text_with_positions(path, pages=None) |
Text with X/Y coords and font info |
extract_text_with_positions_bytes(data, pages=None) |
Text with positions from bytes |
extract_text_in_regions(path, page_regions) |
Extract text in bounding-box regions |
extract_text_in_regions_bytes(data, page_regions) |
Region extraction from bytes |
extract_pages_markdown(path, pages=None) |
Per-page Markdown + layout metadata (all pages by default) |
extract_pages_markdown_bytes(data, pages=None) |
Per-page Markdown from bytes |
Types
Type stubs (pdf_inspector.pyi) ship with the package. Result types at a glance:
class PdfResult: # process_pdf / detect_pdf
pdf_type: str # "text_based" | "scanned" | "image_based" | "mixed"
markdown: str | None # extracted Markdown (None for detect_pdf)
page_count: int
processing_time_ms: int
pages_needing_ocr: list[int]
title: str | None
confidence: float # 0.0 - 1.0
is_complex_layout: bool
pages_with_tables: list[int]
pages_with_columns: list[int]
has_encoding_issues: bool # broken font encodings — consider OCR fallback
class PdfClassification: # classify_pdf
pdf_type: str
page_count: int
pages_needing_ocr: list[int] # 0-indexed
confidence: float
class TextItem: # extract_text_with_positions
text: str
x: float
y: float
width: float
height: float
font: str
font_size: float
page: int
is_bold: bool
is_italic: bool
is_underline: bool
is_strikeout: bool
item_type: str
class PageRegionTexts: # extract_text_in_regions
page: int # 0-indexed
regions: list[RegionText] # RegionText: text: str, needs_ocr: bool
class PagesExtractionResult: # extract_pages_markdown
pages: list[PageMarkdown] # PageMarkdown: page (0-indexed), markdown, needs_ocr
pages_with_tables: list[int] # 1-indexed
pages_with_columns: list[int] # 1-indexed
pages_needing_ocr: list[int] # 1-indexed
is_complex: bool # any page has tables or multi-column layout
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distributions
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file pdf_inspector-0.2.6.tar.gz.
File metadata
- Download URL: pdf_inspector-0.2.6.tar.gz
- Upload date:
- Size: 1.5 MB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
5bb387f39bf7a93b02b49188b670b9798f8ccc7e58f68eee5883a512aa05ceb2
|
|
| MD5 |
63480e474f7dc86c5bbab265777ba370
|
|
| BLAKE2b-256 |
8d7a525ee06cad5c46d8aa7357c15b8f72c60e803246f3b7e24523e602f01f6d
|
Provenance
The following attestation bundles were made for pdf_inspector-0.2.6.tar.gz:
Publisher:
publish-pypi.yml on firecrawl/pdf-inspector
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
pdf_inspector-0.2.6.tar.gz -
Subject digest:
5bb387f39bf7a93b02b49188b670b9798f8ccc7e58f68eee5883a512aa05ceb2 - Sigstore transparency entry: 2308181725
- Sigstore integration time:
-
Permalink:
firecrawl/pdf-inspector@7b7960ee73c64aafa39bbffd8e7c97e4c59a448f -
Branch / Tag:
refs/heads/main - Owner: https://github.com/firecrawl
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish-pypi.yml@7b7960ee73c64aafa39bbffd8e7c97e4c59a448f -
Trigger Event:
push
-
Statement type:
File details
Details for the file pdf_inspector-0.2.6-cp38-abi3-win_amd64.whl.
File metadata
- Download URL: pdf_inspector-0.2.6-cp38-abi3-win_amd64.whl
- Upload date:
- Size: 2.6 MB
- Tags: CPython 3.8+, Windows x86-64
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
89814af887c5ff90f013702ba285ec8a4c5ae4e2ccbdb948f3e421f7c4b4f89d
|
|
| MD5 |
581e7ff2f4178b31ac85a8cf854084f9
|
|
| BLAKE2b-256 |
848a22e2bf7413444036138f55a0c9f2b6420f7ef54b6a044755f4dbb3a7395b
|
Provenance
The following attestation bundles were made for pdf_inspector-0.2.6-cp38-abi3-win_amd64.whl:
Publisher:
publish-pypi.yml on firecrawl/pdf-inspector
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
pdf_inspector-0.2.6-cp38-abi3-win_amd64.whl -
Subject digest:
89814af887c5ff90f013702ba285ec8a4c5ae4e2ccbdb948f3e421f7c4b4f89d - Sigstore transparency entry: 2308181866
- Sigstore integration time:
-
Permalink:
firecrawl/pdf-inspector@7b7960ee73c64aafa39bbffd8e7c97e4c59a448f -
Branch / Tag:
refs/heads/main - Owner: https://github.com/firecrawl
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish-pypi.yml@7b7960ee73c64aafa39bbffd8e7c97e4c59a448f -
Trigger Event:
push
-
Statement type:
File details
Details for the file pdf_inspector-0.2.6-cp38-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.
File metadata
- Download URL: pdf_inspector-0.2.6-cp38-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
- Upload date:
- Size: 2.9 MB
- Tags: CPython 3.8+, manylinux: glibc 2.17+ x86-64
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
df76dd100504b705ce92ef2c668f152f05277d942abd94a7d3f252cb5448d56d
|
|
| MD5 |
9cfad0f7bd0cc990a99c361d4e57f5e7
|
|
| BLAKE2b-256 |
59afbe72ab2bd310b6532e4311cf06175eb722802ccea29dbd9281f66c645bd8
|
Provenance
The following attestation bundles were made for pdf_inspector-0.2.6-cp38-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl:
Publisher:
publish-pypi.yml on firecrawl/pdf-inspector
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
pdf_inspector-0.2.6-cp38-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl -
Subject digest:
df76dd100504b705ce92ef2c668f152f05277d942abd94a7d3f252cb5448d56d - Sigstore transparency entry: 2308182193
- Sigstore integration time:
-
Permalink:
firecrawl/pdf-inspector@7b7960ee73c64aafa39bbffd8e7c97e4c59a448f -
Branch / Tag:
refs/heads/main - Owner: https://github.com/firecrawl
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish-pypi.yml@7b7960ee73c64aafa39bbffd8e7c97e4c59a448f -
Trigger Event:
push
-
Statement type:
File details
Details for the file pdf_inspector-0.2.6-cp38-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl.
File metadata
- Download URL: pdf_inspector-0.2.6-cp38-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
- Upload date:
- Size: 2.8 MB
- Tags: CPython 3.8+, manylinux: glibc 2.17+ ARM64
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
79ba33b224029b68edbd30e7f6dc034f730663749ba0bea12e5873b31435ca28
|
|
| MD5 |
420b8bcc7ced06cf0034ce02750a4788
|
|
| BLAKE2b-256 |
45f3be670e07df2eb57a1adaeb63b1011d17061fd3036370d9f903984eda94e4
|
Provenance
The following attestation bundles were made for pdf_inspector-0.2.6-cp38-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl:
Publisher:
publish-pypi.yml on firecrawl/pdf-inspector
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
pdf_inspector-0.2.6-cp38-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl -
Subject digest:
79ba33b224029b68edbd30e7f6dc034f730663749ba0bea12e5873b31435ca28 - Sigstore transparency entry: 2308182474
- Sigstore integration time:
-
Permalink:
firecrawl/pdf-inspector@7b7960ee73c64aafa39bbffd8e7c97e4c59a448f -
Branch / Tag:
refs/heads/main - Owner: https://github.com/firecrawl
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish-pypi.yml@7b7960ee73c64aafa39bbffd8e7c97e4c59a448f -
Trigger Event:
push
-
Statement type:
File details
Details for the file pdf_inspector-0.2.6-cp38-abi3-macosx_11_0_arm64.whl.
File metadata
- Download URL: pdf_inspector-0.2.6-cp38-abi3-macosx_11_0_arm64.whl
- Upload date:
- Size: 2.6 MB
- Tags: CPython 3.8+, macOS 11.0+ ARM64
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
d2b2aaa95b242da38630bbd0644ffe9a929466f9c4e6406d6f1957b59b413d08
|
|
| MD5 |
f3394fb12cb5c1c15621cd6d4a76d59f
|
|
| BLAKE2b-256 |
72702049766fec20c2ee6c01776cc2b7b8599fc9cf2d48e8940419600f65d4a0
|
Provenance
The following attestation bundles were made for pdf_inspector-0.2.6-cp38-abi3-macosx_11_0_arm64.whl:
Publisher:
publish-pypi.yml on firecrawl/pdf-inspector
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
pdf_inspector-0.2.6-cp38-abi3-macosx_11_0_arm64.whl -
Subject digest:
d2b2aaa95b242da38630bbd0644ffe9a929466f9c4e6406d6f1957b59b413d08 - Sigstore transparency entry: 2308182032
- Sigstore integration time:
-
Permalink:
firecrawl/pdf-inspector@7b7960ee73c64aafa39bbffd8e7c97e4c59a448f -
Branch / Tag:
refs/heads/main - Owner: https://github.com/firecrawl
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish-pypi.yml@7b7960ee73c64aafa39bbffd8e7c97e4c59a448f -
Trigger Event:
push
-
Statement type:
File details
Details for the file pdf_inspector-0.2.6-cp38-abi3-macosx_10_12_x86_64.whl.
File metadata
- Download URL: pdf_inspector-0.2.6-cp38-abi3-macosx_10_12_x86_64.whl
- Upload date:
- Size: 2.7 MB
- Tags: CPython 3.8+, macOS 10.12+ x86-64
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
c935a354facbfb935e88cb4f3f60fd158402f2af661c1a25a3da328eb21f19fb
|
|
| MD5 |
bdeb84ee0a83383bbb4a668ecf814c98
|
|
| BLAKE2b-256 |
8cdc9963c292151dd82678dddb63e1c640ec420e37d400f304dca2780c6d11a3
|
Provenance
The following attestation bundles were made for pdf_inspector-0.2.6-cp38-abi3-macosx_10_12_x86_64.whl:
Publisher:
publish-pypi.yml on firecrawl/pdf-inspector
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
pdf_inspector-0.2.6-cp38-abi3-macosx_10_12_x86_64.whl -
Subject digest:
c935a354facbfb935e88cb4f3f60fd158402f2af661c1a25a3da328eb21f19fb - Sigstore transparency entry: 2308182334
- Sigstore integration time:
-
Permalink:
firecrawl/pdf-inspector@7b7960ee73c64aafa39bbffd8e7c97e4c59a448f -
Branch / Tag:
refs/heads/main - Owner: https://github.com/firecrawl
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish-pypi.yml@7b7960ee73c64aafa39bbffd8e7c97e4c59a448f -
Trigger Event:
push
-
Statement type: