PDF-Helper
A simple python package that helps with doing simple stuff with PDFs.
Features
- Bundle: Bundle multiple files into one PDF
- PDF inputs
- Image inputs (e.g. PNG, JPG, etc.)
- Markdown inputs
- Merge PDFs: Merge multiple PDFs into one PDF
- Split PDFs: Split a PDF into multiple PDFs, each containing a range of pages
- Export as image: Export designated pages from a PDF as image files
- Remove pages: Remove designated pages from a PDF
- Extract text: Export text from a PDF file and optionally save it to a text file
- Recipe system: Chain multiple operations together using YAML recipe files
- Add watermark: Overlay selectable vector text or an image on PDF pages
- Encrypt a PDF
- Decrypt a PDF
- Extract images from a PDF
- Extract links from a PDF
- Set PDF metadata (title, author, etc.)
If you want any other feature to be added, feel free to open an issue or fork the repo and make a merge request after adding your contribution.
Usage
Installation
You can install PDF-Helper via pip:
pip install pdf-helper
# Or use uv to install the tool
uv tool install pdf-helper
And run it using the command line:
pdf-helper <command> [options]
Or you can use uvx to run the package without installing it in a specific python environment:
uvx pdf-helper <command> [options]
You can also clone the repository and use uv run:
git clone https://gitlab.com/CodeWriter21/pdf-helper.git
cd pdf-helper
uv run pdf-helper <command> [options]
Bundle PDFs
Bundle multiple files into one PDF:
pdf-helper bundle <input_file_1> <input_file_2>... <input_file_n> <output_file>
# E.g. Bundle PDFs 1, 2 and 3 into a new PDF
pdf-helper bundle 1.pdf 2.pdf 3.pdf new.pdf
# E.g. Take 1.png, 2.jpg, and 3.png and create a PDF named 123.pdf and override
# if already exists
pdf-helper bundle 1.png 2.jpg 3.png 123.pdf -f
# E.g. Take part1.pdf, image1.png, ending.pdf and bundle them into a PDF named final.pdf
pdf-helper bundle part1.pdf image1.png ending.pdf final.pdf -v
Split PDFs
Split a PDF into multiple PDFs, each containing a range of pages:
pdf-helper split <input_file> <output_folder> -s <split_point_1>,<split_point_2>
# E.g. Split a PDF into three PDFs, one with pages 1-10, the second with pages 11-20 and
# the third with pages 21-end
pdf-helper split my-pdf.pdf my-split-pdfs -s 10,20
# E.g. Split a PDF into PDFs each containing one page
pdf-helper split my-pdf.pdf my-split-pdfs # No need to specify split points
Export PDF pages as image files
Export PDF pages as image files:
pdf-helper to-image <input_file> <output_folder> \
-p <page_number_1>,<page_number_2>,...,<page_number_n> -s <scale_factor>
# E.g. Export pages 1, 2, 3 and 6 from a PDF with scale factor 1
pdf-helper to-image 1.pdf images -p 1:3,6 -s 1
# E.g. Export the last three pages with scale factor 1
pdf-helper to-image 1.pdf images -p -3:-1 -s 1
# E.g. Export all pages from a PDF with scale 2
pdf-helper to-image my-pdf.pdf my-images
Remove pages from a PDF
Remove pages from a PDF:
pdf-helper remove-pages <input_file> <output_file> <page_number_1>,<page_number_2>,...,<page_number_n>
# E.g. Remove pages 1, 2, 3 and 6 from a PDF
pdf-helper remove-pages 1.pdf new.pdf 1:3,6
Add watermark to a PDF
Overlay selectable vector text or an image file on PDF pages:
# E.g. Stamp diagonal DRAFT text across all pages
pdf-helper add-watermark my-pdf.pdf watermarked.pdf DRAFT
# E.g. Overlay a logo in the bottom-right corner of pages 1-3
pdf-helper add-watermark my-pdf.pdf watermarked.pdf --watermark-image logo.png \
--position bottom-right --image-scale 0.5 --pages 1:3
# E.g. Stamp text with a custom font (file path or installed name like Arial)
pdf-helper add-watermark my-pdf.pdf watermarked.pdf "DRAFT" --font ./fonts/MyFont.ttf
Not sure which fonts are available? Query them (name or path both work
with --font):
pdf-helper list-fonts
pdf-helper list-fonts arial
See examples/recipes/watermark-showcase.yaml for a recipe
that chains a text watermark and an image watermark in one run.
Manage PDF metadata
Read, edit, or clear the metadata of a PDF file:
# E.g. Show all set metadata fields (accepts multiple files)
pdf-helper metadata show my-pdf.pdf other.pdf
# E.g. Set the title and author
pdf-helper metadata edit my-pdf.pdf tagged.pdf --title "Report" --author "Me"
# E.g. Remove all metadata (Info dictionary and XMP packet)
pdf-helper metadata clear my-pdf.pdf clean.pdf
Settable fields: title, author, subject, keywords, creator,
producer, creationdate, moddate. Dates accept PDF date strings, ISO
dates (2026-09-04, 2026-09-04 12:30), or datetimes, and show prints
them in ISO format. Writes keep the XMP packet synchronized with the Info
dictionary (and show prefers XMP values when present). Every file
pdf-helper writes gets
Producer (PDF-Helper <version>, unless already set), a CreationDate
if missing, and a fresh ModDate automatically.
Export text from a PDF
To extract text from a PDF file and export them to text files you can do as follows:
pdf-helper extract-text <input_file> -o <output_file_name>
# E.g. Extract text from a PDF named my-pdf.pdf and save it to my-text.txt
pdf-helper extract-text my-pdf.pdf -o my-text.txt
Run Recipes
The recipe system lets you chain multiple PDF operations together in a single run using a YAML file. This unlocks features not available through individual CLI commands (e.g. selecting specific pages per file when bundling).
pdf-helper run-recipe <recipe_file.yaml>
# E.g. Run a simple recipe
pdf-helper run-recipe remove-pages.yaml
# E.g. Run with force overwrite and verbose logging
pdf-helper run-recipe bundle-workflow.yaml --force --verbose
# E.g. Run a built-in recipe template instead of a file
pdf-helper run-recipe --builtin remove-pages --input report.pdf --output cleaned.pdf
# E.g. Run a custom recipe from ~/.local/pdf-helper/recipes by name
pdf-helper run-recipe --custom my-clean --input report.pdf --output cleaned.pdf
Custom recipes are plain YAML recipe files stored in
~/.local/pdf-helper/recipes (override with the PDF_HELPER_RECIPES_DIR
environment variable) and run with the same --input/--output/--defines
bindings. generate-recipe lists them alongside the built-in templates.
Parameterized Recipes
Recipes can take their main input/output paths (and any extra values) from
the command line instead of hardcoding them, so one recipe works on many
files. Write {input}, {output}, or any {name} placeholder in the YAML
and bind it at runtime:
name: "Remove specific pages"
version: "1.0"
parameters: [input, output]
steps:
- id: clean
operation: remove_pages
input: "{input}"
pages_to_remove: [2, 4, 6]
output: "{output}"
pdf-helper run-recipe clean.yaml --input report.pdf --output cleaned.pdf
pdf-helper run-recipe clean.yaml --input a.pdf --output a-clean.pdf \
--defines pages=1:3 --defines mode=strict
--input/--output bind {input}/{output}; repeatable --defines KEY=VALUE binds any other {name}. The optional parameters: list
declares what a recipe expects — bare names are required, while
name: default mappings provide defaults that CLI flags override:
parameters:
- input
- output
- label: PDF-Helper
- pages: "1,-1"
Path parts are available as {input.stem}, {input.name},
{input.suffix}, and {input.parent}, so outputs can be named after the
input (e.g. output: "{input.stem}-pages"). Defaults may reference other
placeholders, so - output: "{input.stem}-stamped-pages" works as an
overridable default.
Page selections use comma-separated items where start:end selects an
inclusive range: 1:3,6 means pages 1, 2, 3 and 6. Numbers are 1-based
(0 is rejected), and negatives count back from the last page, so 5:-2
means from page 5 to one page before the last. Values starting with -
must use the --option=value form (--pages="-2:-1") so the CLI does not
mistake them for flags. Built-in templates (generate-recipe lists them)
accept bindings the same way via --builtin <name>.
Recipe File Format
A recipe is a YAML file with a steps list. Each step has an id, an
operation, input/output paths, and operation-specific options. Steps can
reference each other's outputs using { step: step_id }.
name: "Remove specific pages"
description: "Removes pages 2, 4, 6 from a PDF."
version: "1.0"
steps:
- id: clean
operation: remove_pages
input: document.pdf
pages_to_remove: [2, 4, 6]
output: cleaned.pdf
Supported Operations
| Operation | Status | Description |
|---|---|---|
bundle |
Available | Bundle files with optional per-file page selection |
remove_pages |
Available | Remove pages by 1-based index |
split_pdf |
Available | Split at given page boundaries |
pdf_to_image |
Available | Render pages as PNG images |
extract_text |
Available | Extract text content |
watermark |
Available | Overlay vector text (text) or an image (image) |
encrypt |
Planned | Password-protect PDF (graceful fallback) |
metadata |
Available | Set title/author/keywords (+creator/dates) |
clear_metadata |
Available | Remove all metadata (Info + XMP) |
Operations marked Planned are not yet implemented — the recipe runner logs a warning and copies the input file through, so pipelines don't break.
Watermark text is inserted as selectable vector objects; image watermarks are
raster overlays. Both support position, opacity, rotation, and pages.
Advanced Example: Multi-step Pipeline
# yaml-language-server: $schema=https://gitlab.com/CodeWriter21/pdf-helper/-/raw/master/schemas/recipe-schema.json
name: "Split, Convert, and Extract Pipeline"
version: "1.0"
settings:
temp_dir: "./.recipe-tmp"
steps:
# Step 1: Split the PDF at pages 5 and 10
- id: split
operation: split_pdf
input: report.pdf
split_points: [5, 10]
output_dir: .
output_prefix: "report_part_"
# Step 2: Convert the second chunk to images
- id: to_images
operation: pdf_to_image
input:
step: split
file: report_part_2.pdf
pages: "1:3"
scale: 3
output: ./output/images
# Step 3: Extract text from the first chunk
- id: extract
operation: extract_text
input:
step: split
file: report_part_1.pdf
pages: "1:4"
max_characters: 5000
reverse_lines: true
output: ./output/chapter-1-text.txt
Recipe Settings
| Setting | Default | Description |
|---|---|---|
temp_dir |
./.recipe-tmp |
Directory for intermediate files |
overwrite |
false |
Overwrite existing output files |
cleanup_temp |
false |
Remove temp directory after completion |
Input values (e.g. passwords) can be sourced from environment variables or prompted at runtime:
inputs:
password:
env: PDF_PASSWORD
prompt: "Enter output PDF password"
See examples/recipes/ for more example recipe files.
About
Author: CodeWriter21
GitLab: CodeWriter21/pdf-helper
Donations
Your donations are very welcome: nowpayments.io
You can also consider donating a Star to the repo.
License
This project is licensed under the MIT License.
See the LICENSE
References
Release files for pdf-helper 0.6.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| pdf_helper-0.6.0.tar.gz | 28.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| pdf_helper-0.6.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 59.0 kB
Release files / pdf_helper-0.6.0.tar.gz
| Download URL | pdf_helper-0.6.0.tar.gz |
|---|---|
| Size | 28.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
7a363bb46453aadd0bbc6b52e8de9f7e4fd992124a891680167fe59a3fdfb97f
|
|
BLAKE2b-256 checksum How to use checksums |
d53ad954c08d491adc795df6b30ef6edf994f66e41430abbea9cca36c5b9b973
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.12.10 {"installer":{"name":"uv","version":"0.12.10","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Debian GNU/Linux","version":"13","id":"trixie","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|
Release files / pdf_helper-0.6.0-py3-none-any.whl
| Download URL | pdf_helper-0.6.0-py3-none-any.whl |
|---|---|
| Size | 30.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
38f1e00db6e7351ef54ab29189292fd517acce3dca8e6af6da6a2f9b25f361e6
|
|
BLAKE2b-256 checksum How to use checksums |
f7ef68b85010e0659167e43c1fee836b091bc671d9120250d8abb1f490b9bea8
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.12.10 {"installer":{"name":"uv","version":"0.12.10","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Debian GNU/Linux","version":"13","id":"trixie","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|