PdfTokenizer
A Python library for extracting text from PDFs with automatic OCR detection.
Features
- 🔍 Smart OCR Detection: Automatically determines if OCR is needed by analyzing text extractability
- 🔄 Dual Extraction Methods: Uses PdfPlumber for native PDFs and Tesseract for scanned documents
- 🪟 Windows Support: Automatic Poppler download and setup for Windows users
Installation
pip install pdftokenizer
Quick Start
from pdftokenizer import extract_tokens_from_pdf
# Read your PDF file
with open("document.pdf", "rb") as f:
pdf_bytes = f.read()
# Extract tokens - OCR will be used automatically if needed
pages = extract_tokens_from_pdf(pdf_bytes)
# Force OCR if desired
pages_ocr = extract_tokens_from_pdf(pdf_bytes, force_ocr=True)
How It Works
The library automatically determines whether to use OCR based on text extractability:
- Attempts to extract text from the PDF using PyPDF
- If the extracted text contains fewer than 10 characters (configurable threshold), the PDF is considered to need OCR
- Based on this detection:
- Text-based PDFs: Processed using PdfPlumber for efficient extraction
- Scanned/Image PDFs: Processed using Tesseract OCR
Requirements
Poppler
PDF processing backend:
- Windows: Automatically downloaded and configured
- Linux:
apt-get install poppler-utils - macOS:
brew install poppler
Tesseract
Required for OCR functionality:
- Windows: Download from UB Mannheim
- Linux:
apt-get install tesseract-ocr - macOS:
brew install tesseract
License
pdftokenizer is distributed under the terms of the MIT license.
Metadata
Release files for pdftokenizer 0.0.4
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| pdftokenizer-0.0.4.tar.gz | 562.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| pdftokenizer-0.0.4-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 570.5 kB
Release files / pdftokenizer-0.0.4.tar.gz
| Download URL | pdftokenizer-0.0.4.tar.gz |
|---|---|
| Size | 562.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
3a96765251c72db93785d2b944db4aed8dfc6de158432c876f2646fccfad322b
|
|
BLAKE2b-256 checksum How to use checksums |
1e241f21bae8c1e275ed979c6c0767ca9db35c5b1636d4ca708a7a39fd1ff267
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
python-httpx/0.27.0
|
Release files / pdftokenizer-0.0.4-py3-none-any.whl
| Download URL | pdftokenizer-0.0.4-py3-none-any.whl |
|---|---|
| Size | 8.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
22dae8bde95c1d5d44d3b5193abb1e4b344496b18edc500362eb8a71f082e4a5
|
|
BLAKE2b-256 checksum How to use checksums |
6884c6af22d290ff10e03a9dfdb399d75fd5f5659e4f1a8017388e08923ca51f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
python-httpx/0.27.0
|