A local CLI tool for PDF and image manipulation - merge PDFs, split by page/range, convert images to PDF with enhancement
Project description
PdfX
PdfX is a local CLI tool for PDF and image manipulation that keeps your documents private on your own machine. No uploading to online services — everything happens locally! The "X" stands for utility functions: merge PDFs, split by pages or ranges, convert images to PDFs, and enhance scanned documents.
This README explains how to install and use the tool, with clear examples and available command-line options.
Why use PdfX?
- Privacy first: Keep your documents private — everything happens locally on your machine
- Comprehensive features: Merge PDFs, split by pages/ranges, convert and enhance images
- Scanned document enhancement: Improve image and PDF quality with color enhancement filters
- Lightweight: Minimal dependencies for basic features, optional advanced features available
- Script-friendly: Perfect for automation and batch processing workflows
Features
- Merge PDFs: Combine multiple PDF files from a directory into a single PDF document
- Split PDF into pages: Break a PDF into individual page files
- Split PDF by page range: Extract specific pages or page ranges (e.g., pages 1-3, 5, 7-10)
- Convert image to PDF: Turn a single image (JPEG, PNG, BMP, GIF, TIFF, WebP) into a PDF
- Merge images to PDF: Combine multiple images from a directory into a single PDF document
- Image to PDF with enhancement: Convert images to PDF with color enhancement for scanned document quality
- Enhance PDF colors: Improve existing PDFs with scanned document color enhancement
- Organized outputs: Results saved to
Merged_Doc/andSplitted_Docs/directories for easy management
Installation 🔧
Install from PyPI (Recommended)
For basic PDF merge/split functionality:
pip install pdfx-tool
For full functionality (including color filtering, OCR, image processing):
pip install pdfx-tool[full]
For development:
pip install pdfx-tool[dev]
Install from source
Clone the repository and install:
git clone https://github.com/yourusername/pdfx.git
cd pdfx
pip install -e .
Or with full dependencies:
pip install -e .[full]
System dependencies (optional)
For advanced features, you may need additional system tools:
- Tesseract OCR (for
--ocr):brew install tesseract(macOS) or install from https://github.com/tesseract-ocr/tesseract - Poppler (for some PDF imaging backends used by tools like
pdf2image):brew install poppler(macOS)
Note: Basic merge/split functionality works without these extras.
Quick start ✨
After installation, use the pdfx command:
pdfx --help
Or run as a Python module:
python -m pdfx --help
Command-line options (friendly explanations)
-
-m, --merge-dir PATH
Directory containing PDFs to merge. If you omit this, the tool looks forfilesToMergenext tomain.py. -
-o, --output NAME
Output filename for the merged PDF (default:Merged Document.pdf). The file will be written to<merge_dir_parent>/Merged_Doc/. -
--split
Split each PDF located in the source directory into single-page PDFs. By default, the source directory used isfilesToSplitnext tomain.py. -
--split-dir PATH
Use this directory as the source for--splitinstead of the default. -
--split-file PATH
Split a single PDF file into pages (or a page range). By default the output goes to../Splitted_Docs/<file_stem>/. -
--split-out PATH
Write split pages to this directory (overrides the default Splitted_Docs location). -
--page-range "RANGE"
When splitting, extract only the pages you want. Example:--page-range "1-3,5". Ranges are 1-based and multiple segments are allowed. -
--filter-file PATH
Filter a single PDF by text color and write a new PDF containing only text spans that match the color (requirespymupdf). -
--color-filter COLOR
Target color to filter. Accepts multiple formats:- Hex:
#RRGGBB(e.g.,#FF0000) - RGB:
R,G,B(e.g.,255,0,0) - Named:
red,blue,green,black,white,yellow,cyan,magenta,orange,purple,pink,brown,gray,lime,navy,teal,silver,maroon,olive
- Hex:
-
--color-tolerance FLOAT
Color distance tolerance (default0.0). A larger value allows fuzzy color matches. -
--image-filter NAME
Apply an image-level filter to every page in a PDF. Choices:enhance,bw,grayscale,invert,auto. -
--image-strength FLOAT
Strength or threshold for image filters. Meaning depends on the filter: forenhanceit's a contrast factor (1.0 = no-change), forbwit's the threshold (0–255, default 128). -
--image-to-pdf
Convert image file(s) to PDF. Use with--image-filefor a single image or--image-dirfor a directory of images. -
--image-file PATH
Single image file to convert to PDF (requires--image-to-pdf). -
--image-dir PATH
Directory containing images to convert to a single PDF. All supported image formats (JPEG, PNG, BMP, GIF, TIFF, WebP) are processed in sorted order (requires--image-to-pdf). -
--image-out PATH
Output path for the PDF generated from images (optional; defaults to<image_name>.pdfor<dir_name>_to_pdf.pdf). -
--image-enhance TYPE
Apply color enhancement to images during PDF conversion. Choices:enhance(contrast),brightness,color(saturation),sharpness,auto(autocontrast). -
--image-enhance-strength FLOAT
Strength of image enhancement. For contrast/brightness/color/sharpness: 1.0 = no change, < 1.0 = decrease, > 1.0 = increase. Default: 1.5. Forautoenhancement, this parameter is ignored.
Image filter mapping (what each filter does) 🔧
- enhance — increases page contrast (useful for scanned docs to make text crisper). Default strength 1.5.
- bw — converts to pure black & white using a threshold; pass threshold via
--image-strength(0–255). - grayscale — converts pages to grayscale.
- invert — inverts colors (useful for turning light-on-dark into dark-on-light).
- auto — applies autocontrast.
Advanced features (recoloring, OCR, image-to-PDF, auto-enhancement, diagnostics) ✨
-
Auto PDF Enhancement (Smart Scanned Document Enhancement) —
--pdf-enhanceautomatically enhances scanned PDFs with intelligent defaults. Requires--filter-filefor input. Choose enhancement mode with--pdf-enhance-mode:smart(default, recommended): Applies contrast enhancement (1.5x) followed by autocontrast for optimal claritycontrast: Standard contrast enhancement (1.5x) for moderately faded documentsstrong: Heavy contrast enhancement (2.0x) for heavily faded or low-quality scansauto: Autocontrast only for fine-tuning brightness levels- Example:
pdfx --filter-file scanned.pdf --pdf-enhance # Creates: scanned_enhanced.pdf with smart auto-enhancement
- Example (strong mode for heavily faded documents):
pdfx --filter-file faded_scan.pdf --pdf-enhance --pdf-enhance-mode strong # Creates: faded_scan_enhanced.pdf with 2.0x contrast boost
-
Image-to-PDF conversion with color enhancement —
--image-to-pdfconverts one or more images into a PDF document with optional color enhancement. Use--image-filefor a single image or--image-dirto batch-convert all images in a directory. Apply enhancements with--image-enhance(contrast, brightness, color saturation, sharpness, or auto) and adjust strength with--image-enhance-strength. Perfect for scanned documents that need enhanced clarity.- Example (single image with contrast enhancement):
pdfx --image-to-pdf --image-file scan.jpg --image-enhance enhance --image-enhance-strength 1.8 --image-out scan.pdf
- Example (directory of scans with auto-enhancement):
pdfx --image-to-pdf --image-dir /path/to/scans --image-enhance auto # Creates: /path/to/scans_to_pdf.pdf (with autocontrast applied to each image)
- Example (increase color saturation):
pdfx --image-to-pdf --image-file photo.jpg --image-enhance color --image-enhance-strength 1.3 --image-out photo.pdf
- Example (single image with contrast enhancement):
-
Shorthand flags — Use
--enhance,--bw,--grayscale,--invert, or--autoas convenient aliases for--image-filter <name>. When using--bwa sensible default threshold is set if you don't pass--image-strength. -
Recolor vector text —
--recolor COLORrecolors matched vector text spans (requires--color-filterand--filter-file).--recolor-tolerance FLOATto allow fuzzy matches.--recolor-replacewill attempt to replace underlying text; use cautiously as it may affect layout.- Example:
pdfx --filter-file filesToMerge/1.pdf --color-filter '255,0,0' --recolor '#0000FF' --filter-out out_recolor.pdf
-
OCR (make searchable) —
--ocrproduces a searchable PDF by OCR-ing each page (requires Tesseract +pytesseract). Use--ocr-langto set language(s).- Example:
pdfx --filter-file filesToMerge/1.pdf --ocr --ocr-lang eng --filter-out out_ocr.pdf
- Example:
-
Diagnostics —
--dump-colorsprints detected text-span colors and samples to help you pick a--color-filter.- Example:
pdfx --filter-file filesToMerge/1.pdf --dump-colors
- Example:
Note: Recolor overlays recolored text to preserve layout; image filters rasterize pages and may make text unsearchable. Pick the approach that suits your needs.
Examples
Image to PDF Conversion
-
Convert a single image to PDF:
pdfx --image-to-pdf --image-file photo.jpg --image-out photo.pdf
-
Convert and enhance a scanned document (increase contrast):
pdfx --image-to-pdf --image-file scan.jpg --image-enhance enhance --image-enhance-strength 1.8 --image-out scan_enhanced.pdf
-
Auto-enhance multiple scans in a folder:
pdfx --image-to-pdf --image-dir /path/to/scans --image-enhance auto # Creates: /path/to/scans_to_pdf.pdf with autocontrast applied
-
Restore color vibrancy to a faded photo:
pdfx --image-to-pdf --image-file old_photo.jpg --image-enhance color --image-enhance-strength 1.3 --image-out photo_restored.pdf
-
Increase sharpness for blurry images:
pdfx --image-to-pdf --image-file blurry.jpg --image-enhance sharpness --image-enhance-strength 1.6 --image-out sharp.pdf
-
Brighten dark/underexposed images:
pdfx --image-to-pdf --image-file dark_photo.jpg --image-enhance brightness --image-enhance-strength 1.2 --image-out bright.pdf
-
Convert all images from a folder (no enhancement):
pdfx --image-to-pdf --image-dir /path/to/images # Creates: /path/to/images_to_pdf.pdf
Merging PDFs
-
Merge all PDFs in a directory into one file:
pdfx -m /path/to/pdfs -o "Final Report.pdf" # Creates: /path/to/Merged_Doc/Final Report.pdf
-
Merge using default
filesToMergedirectory:pdfx # Creates: Merged_Doc/Merged Document.pdf -
Merge and specify custom output name:
pdfx -m documents/ -o "Combined.pdf"
Splitting PDFs
-
Split all PDFs in a directory into single-page files:
pdfx --split # Creates: Splitted_Docs/<pdf_name>/page_1.pdf, page_2.pdf, ...
-
Split a single PDF file:
pdfx --split-file report.pdf # Creates: Splitted_Docs/report/report_page_1.pdf, report_page_2.pdf, ...
-
Extract specific pages (pages 1-3 and 5):
pdfx --split-file document.pdf --page-range "1-3,5" # Creates: Splitted_Docs/document/document_page_1.pdf ... document_page_5.pdf
-
Extract single page:
pdfx --split-file document.pdf --page-range 10 # Creates: Splitted_Docs/document/document_page_10.pdf
-
Split and save to custom directory:
pdfx --split-file document.pdf --split-out /custom/output/path # Creates: /custom/output/path/document_page_1.pdf, ...
Color Filtering (Extract Colored Text)
-
Extract only red text from a PDF:
pdfx --filter-file document.pdf --color-filter "#FF0000" --filter-out red_text_only.pdf # Or use RGB: --color-filter "255,0,0"
-
Extract text with fuzzy color matching (tolerance):
pdfx --filter-file document.pdf --color-filter "#FF0000" --color-tolerance 30 --filter-out similar_red.pdf
-
Find what colors are in your PDF:
pdfx --filter-file document.pdf --dump-colors # Prints all detected text colors and samples
Recoloring Text
-
Recolor all red text to blue (preserves layout):
pdfx --filter-file document.pdf --color-filter "#FF0000" --recolor "#0000FF" --filter-out recolored.pdf
-
Recolor with fuzzy matching and replacement:
pdfx --filter-file document.pdf --color-filter "255,0,0" --recolor "0,0,255" --recolor-tolerance 20 --recolor-replace --filter-out recolored.pdf
Page-Level Image Filters (applies to entire page raster)
Automatic PDF Enhancement (Smart Scanned Document Enhancement)
-
Auto-enhance a scanned PDF with smart defaults (recommended):
pdfx --filter-file scanned.pdf --pdf-enhance # Creates: scanned_enhanced.pdf # Applies: contrast boost (1.5x) + autocontrast for optimal clarity
-
Auto-enhance with strong mode (for heavily faded documents):
pdfx --filter-file heavily_faded.pdf --pdf-enhance --pdf-enhance-mode strong # Creates: heavily_faded_enhanced.pdf with 2.0x contrast boost
-
Auto-enhance with standard contrast (moderate fading):
pdfx --filter-file scanned.pdf --pdf-enhance --pdf-enhance-mode contrast # Creates: scanned_enhanced.pdf with 1.5x contrast only
-
Auto-enhance with autocontrast only (fine-tuning brightness):
pdfx --filter-file scanned.pdf --pdf-enhance --pdf-enhance-mode auto # Creates: scanned_enhanced.pdf with automatic brightness adjustment
-
Complete workflow: Auto-enhance then make searchable:
# Step 1: Auto-enhance the scanned PDF pdfx --filter-file raw_scan.pdf --pdf-enhance --filter-out enhanced.pdf # Step 2: Make it searchable with OCR pdfx --filter-file enhanced.pdf --ocr --filter-out searchable.pdf # Creates fully enhanced and searchable PDF
PDF Enhancement & Quality Improvements (Manual Control)
-
Increase contrast for scanned PDFs (light enhancement):
pdfx --filter-file scanned.pdf --enhance --image-strength 1.3 --filter-out enhanced_light.pdf # 1.3 = 30% contrast increase; good for slightly faded scans
-
Increase contrast for scanned PDFs (medium enhancement):
pdfx --filter-file scanned.pdf --enhance --image-strength 1.6 --filter-out enhanced_medium.pdf # 1.6 = 60% contrast increase; standard for most scanned documents
-
Increase contrast for heavily faded documents (strong enhancement):
pdfx --filter-file faded.pdf --enhance --image-strength 2.0 --filter-out enhanced_strong.pdf # 2.0 = double the contrast; use for very faded or low-quality scans
-
Apply auto-contrast (automatic optimization):
pdfx --filter-file scanned.pdf --auto --filter-out auto_enhanced.pdf # Automatically adjusts contrast for optimal visibility
Color Processing
-
Convert to grayscale (save space, remove colors):
pdfx --filter-file colored.pdf --grayscale --filter-out gray.pdf # Useful for reducing file size and printing in black & white
-
Convert PDF pages to black & white (with threshold):
pdfx --filter-file document.pdf --bw --image-strength 128 --filter-out bw.pdf # Threshold 128 (middle gray) - adjust 0-255 for darker/lighter cutoff
-
Convert with darker threshold (more text retained):
pdfx --filter-file document.pdf --bw --image-strength 100 --filter-out bw_darker.pdf # Lower threshold (e.g., 100) keeps more dark areas
-
Convert with lighter threshold (cleaner appearance):
pdfx --filter-file document.pdf --bw --image-strength 150 --filter-out bw_lighter.pdf # Higher threshold (e.g., 150) produces cleaner white backgrounds
-
Invert colors (light → dark, dark → light):
pdfx --filter-file document.pdf --invert --filter-out inverted.pdf # Useful for converting dark background documents to light backgrounds
Enhancement Workflow Examples
-
Complete scanning workflow (enhance + auto-optimize):
# Step 1: Enhance contrast pdfx --filter-file raw_scan.pdf --enhance --image-strength 1.8 --filter-out scan_enhanced.pdf # Step 2: Make searchable with OCR pdfx --filter-file scan_enhanced.pdf --ocr --filter-out scan_searchable.pdf
-
Batch enhance all PDFs in a folder:
# Create a script to enhance multiple PDFs for pdf in ./documents/*.pdf; do pdfx --filter-file "$pdf" --enhance --image-strength 1.5 --filter-out "${pdf%.pdf}_enhanced.pdf" done
-
Compare different enhancement strengths:
# Light pdfx --filter-file scan.pdf --enhance --image-strength 1.2 --filter-out scan_light.pdf # Medium pdfx --filter-file scan.pdf --enhance --image-strength 1.5 --filter-out scan_medium.pdf # Strong pdfx --filter-file scan.pdf --enhance --image-strength 1.8 --filter-out scan_strong.pdf # Then compare the three versions to find the best result
-
Convert faded document to clean black & white:
pdfx --filter-file faded_doc.pdf --bw --image-strength 120 --filter-out clean.pdf # Removes gray backgrounds and produces crisp text
OCR (Make Searchable)
-
Make a scanned PDF searchable:
pdfx --filter-file scanned.pdf --ocr --filter-out searchable.pdf # Requires: Tesseract installed
-
OCR with specific language:
pdfx --filter-file french_scan.pdf --ocr --ocr-lang fra --filter-out searchable_fra.pdf
-
OCR with multiple languages:
pdfx --filter-file multilang.pdf --ocr --ocr-lang "eng+fra+deu" --filter-out searchable_multi.pdf
Complex Workflows
-
Scan → Enhance → Extract Pages → Merge:
# Step 1: Convert scans to PDF with enhancement pdfx --image-to-pdf --image-dir ./scans --image-enhance enhance --image-enhance-strength 1.8 --image-out scans.pdf # Step 2: Extract pages 1-10 pdfx --split-file scans.pdf --page-range 1-10 --split-out ./extracted_pages # Step 3: Merge extracted pages with other PDFs pdfx -m ./extracted_pages -o "Final.pdf"
-
Process batch of documents (merge → enhance → OCR):
# Merge all project PDFs pdfx -m ./documents -o "Project.pdf" # Enhance for better readability pdfx --filter-file Merged_Doc/Project.pdf --enhance --image-strength 1.5 --filter-out Project_Enhanced.pdf # Make searchable pdfx --filter-file Project_Enhanced.pdf --ocr --filter-out Project_Searchable.pdf
Exit codes
0— success (merge or split completed)1— directory or file not found2— invalid input or processing error
Troubleshooting & tips
- If you see a pypdf error about AES, install
cryptographyin your environment (see Installation section). - The tool will create
Merged_DocandSplitted_Docsdirectories as needed. - Filenames with spaces are supported; use quotes or escape spaces in shell commands.
Contributing & Feedback 🤝
Contributions, bug reports, and improvements are very welcome. Please open an issue or submit a pull request with tests and a short description of your change.
License
This project is provided as-is. Feel free to use and adapt it for your personal or internal projects.
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file pdfx_tool-0.2.0.tar.gz.
File metadata
- Download URL: pdfx_tool-0.2.0.tar.gz
- Upload date:
- Size: 26.7 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.9.6
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
1adf2486dad55ce513de3a063497d8aa07168ec88e5efe36b9db47fb8f26531a
|
|
| MD5 |
6d01a34134f185d6379ece2cd6fd7352
|
|
| BLAKE2b-256 |
0c10682ebadff814795f588fc6722916b25bc47bbd7b6522dcec103a7a3da775
|
File details
Details for the file pdfx_tool-0.2.0-py3-none-any.whl.
File metadata
- Download URL: pdfx_tool-0.2.0-py3-none-any.whl
- Upload date:
- Size: 19.8 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.9.6
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
baa95cd0c4a693c51c9e24259b629abf24bcd8d496f005080fdb2b58137ac646
|
|
| MD5 |
6f75b6ab5fca8f4166108c834b594acb
|
|
| BLAKE2b-256 |
f614f4ba40331b0c5b5f2e6ae01687f8f29ae078e19b0ade0dd58bc1cf4798c6
|