Skip to main content

Natural PDF

CI

A friendly library for working with PDFs, built on top of pdfplumber.

Natural PDF lets you find and extract content from PDFs using simple code that makes sense.

Installation

pip install natural-pdf

Need OCR, semantic search, export, or AI-powered extraction? Install what you need:

pip install "natural-pdf[all]"      # Recommended feature-complete install
pip install "natural-pdf[export]"   # Export helpers only
pip install rapidocr                # Default OCR backend
pip install "natural-pdf[paddle]"   # PaddleOCR stack
pip install python-doctr            # Doctr OCR engine

More details in the installation guide.

natural-pdf[all] is the recommended feature-complete runtime bundle for core features: the default RapidOCR engine, sentence-transformers-based semantic search, QA/extraction dependencies, YOLO layout detection, and export support. It does not install every optional backend. Extra engines such as PaddleOCR and Doctr stay opt-in, and Natural PDF will tell you what to install when you try to use something that is missing.

Check your local setup with:

npdf doctor

Quick Start

from natural_pdf import PDF

# Open a PDF
pdf = PDF('https://github.com/jsoma/natural-pdf/raw/refs/heads/main/pdfs/01-practice.pdf')
page = pdf.pages[0]

# Extract all of the text on the page
page.extract_text()

# Find elements using CSS-like selectors
heading = page.find('text:contains("Summary"):bold')

# Extract content below the heading
content = heading.below().extract_text()

# Examine all the bold text on the page
page.find_all('text:bold').show()

# Exclude parts of the page from selectors/extractors
header = page.find('text:contains("CONFIDENTIAL")').above()
footer = page.find_all('line')[-1].below()
page.add_exclusion(header)
page.add_exclusion(footer)

# Extract clean text from the page ignoring exclusions
clean_text = page.extract_text()

And as a fun bonus, page.viewer() will provide an interactive method to explore the PDF.

Key Features

Natural PDF offers a range of features for working with PDFs:

  • CSS-like Selectors: Find elements using intuitive query strings (page.find('text:bold')).
  • Spatial Navigation: Select content relative to other elements (heading.below(), element.select_until(...)).
  • Text & Table Extraction: Get clean text or structured table data, automatically handling exclusions.
  • OCR Integration: Extract text from scanned documents with RapidOCR by default, plus opt-in engines like PaddleOCR or Doctr.
  • Layout Analysis: Detect document structures (titles, paragraphs, tables) using various engines (e.g., YOLO, Paddle, LLM via API).
  • Document QA: Ask natural language questions about your document's content.
  • Semantic Search: Rank pages within a PDF by semantic similarity using sentence-transformer embeddings.
  • Visual Debugging: Highlight elements and use an interactive viewer or save images to understand your selections.

Learn More

Dive deeper into the features and explore advanced usage in the Complete Documentation.

Extending Natural PDF

Natural PDF now exposes its pluggable engines through small helper functions so you rarely have to touch the core registry directly. Two handy entry points:

from natural_pdf.tables import register_table_function

def table_delim(region, *, context=None, **kwargs):
    # return a TableResult or list-of-lists
    ...

register_table_function("table_delim", table_delim)
from natural_pdf.selectors import register_selector_engine

class DebugSelectorEngine:
    def query(self, *, context, selector, options):
        ...

register_selector_engine("debug", lambda **_: DebugSelectorEngine())

Best friends

Natural PDF sits on top of a lot of fantastic tools and models, some of which are:

Metadata

Release files for natural-pdf 0.7.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for natural-pdf 0.7.0
File Size Uploaded
natural_pdf-0.7.0.tar.gz 963.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for natural-pdf 0.7.0
File Interpreter ABI Platform
natural_pdf-0.7.0-py3-none-any.whl Python 3 none any Details

Total release size: 1.8 MB

Release files / natural_pdf-0.7.0.tar.gz

Download URL natural_pdf-0.7.0.tar.gz
Size 963.9 kB
Tags Source
SHA-256 checksum
How to use checksums
f0aae27eae65eff9f76d388920558b1cf1b04c3ad6ddfb1b752d397f5f317f97
BLAKE2b-256 checksum
How to use checksums
67a29c9f99aff0b97f9595452d991d7b3ba1e01c16419c59616910526eb96d78
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.13.5

Release files / natural_pdf-0.7.0-py3-none-any.whl

Download URL natural_pdf-0.7.0-py3-none-any.whl
Size 841.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
b6b55fe28df6ed0c18363d97343b084be259341156ccaf41e631802c21fbd984
BLAKE2b-256 checksum
How to use checksums
8420c7804a6b2ad33c2e8da5a8e99262f4f2e231f4ef61a79ff1602b7bc56b43
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.13.5

Release history Release notifications | RSS feed

This release

0.7.0 This release

2 release files

0.6.5

2 release files

0.6.4

2 release files

0.6.3

2 release files

0.6.2

2 release files

0.6.1

2 release files

0.6.0

2 release files

0.5.4

2 release files

0.5.3

2 release files

0.5.2

2 release files

0.5.1

2 release files

0.5.0

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.3

2 release files

0.3.2

2 release files

0.3.1

2 release files

0.3.0

2 release files

0.2.22

2 release files

0.2.21

2 release files

0.2.16

2 release files

0.2.15

2 release files

0.2.13

2 release files

0.2.12

2 release files

0.2.9

2 release files

0.2.8

2 release files

0.2.6

2 release files

0.2.5

2 release files

0.2.4

2 release files

0.2.3

2 release files

0.2.2

2 release files

0.2.0

1 release file

0.1.40

2 release files

0.1.38

2 release files

0.1.37

2 release files

0.1.36

2 release files

0.1.35

2 release files

0.1.34

2 release files

0.1.33

2 release files

0.1.32

2 release files

0.1.31

2 release files

0.1.30

2 release files

0.1.28

2 release files

0.1.27

2 release files

0.1.24

2 release files

0.1.23

2 release files

0.1.22

2 release files

0.1.21

2 release files

0.1.20

2 release files

0.1.19

2 release files

0.1.18

2 release files

0.1.17

2 release files

0.1.16

2 release files

0.1.15

2 release files

0.1.12

2 release files

0.1.9

2 release files

0.1.8

2 release files

0.1.7

2 release files

0.1.6

2 release files

0.1.5

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page