Skip to main content

PDF-Helper

A simple python package that helps with doing simple stuff with PDFs.

Features

  • Bundle: Bundle multiple files into one PDF
    • PDF inputs
    • Image inputs (e.g. PNG, JPG, etc.)
    • Markdown inputs
  • Merge PDFs: Merge multiple PDFs into one PDF
  • Split PDFs: Split a PDF into multiple PDFs, each containing a range of pages
  • Export as image: Export designated pages from a PDF as image files
  • Remove pages: Remove designated pages from a PDF
  • Extract text: Export text from a PDF file and optionally save it to a text file
  • Recipe system: Chain multiple operations together using YAML recipe files
  • Add watermark: Overlay selectable vector text or an image on PDF pages
  • Encrypt a PDF (user/owner passwords, AES-256/128, RC4)
  • Decrypt a PDF (user or owner password)
  • Extract images from a PDF
  • Extract links from a PDF
  • Set PDF metadata (title, author, etc.)

If you want any other feature to be added, feel free to open an issue or fork the repo and make a merge request after adding your contribution.

Usage

Installation

You can install PDF-Helper via pip:

pip install pdf-helper

# Or use uv to install the tool
uv tool install pdf-helper

Document operations (metadata editing, encryption, and the automatic metadata stamping) need the optional pikepdf dependency:

pip install "pdf-helper[pikepdf]"

Image operations (rendering pages to images, image watermarks, bundling or converting image files) need the optional Pillow dependency:

pip install "pdf-helper[pillow]"

Or everything at once:

pip install "pdf-helper[all]"

Everything else — bundling PDFs, splitting, text watermarking, text extraction, and recipes not using those operations — works without them. If a feature needs the missing extra, the command fails with the install instruction above instead of a traceback.

And run it using the command line:

pdf-helper <command> [options]

Or you can use uvx to run the package without installing it in a specific python environment:

uvx pdf-helper <command> [options]

You can also clone the repository and use uv run:

git clone https://gitlab.com/CodeWriter21/pdf-helper.git
cd pdf-helper
uv run pdf-helper <command> [options]

Bundle PDFs

Bundle multiple files into one PDF:

pdf-helper bundle <input_file_1> <input_file_2>... <input_file_n> <output_file>

# E.g. Bundle PDFs 1, 2 and 3 into a new PDF
pdf-helper bundle 1.pdf 2.pdf 3.pdf new.pdf

# E.g. Take 1.png, 2.jpg, and 3.png and create a PDF named 123.pdf and override
# if already exists
pdf-helper bundle 1.png 2.jpg 3.png 123.pdf -f

# E.g. Take part1.pdf, image1.png, ending.pdf and bundle them into a PDF named final.pdf
pdf-helper bundle part1.pdf image1.png ending.pdf final.pdf -v

Split PDFs

Split a PDF into multiple PDFs, each containing a range of pages:

pdf-helper split <input_file> <output_folder> -s <split_point_1>,<split_point_2>

# E.g. Split a PDF into three PDFs, one with pages 1-10, the second with pages 11-20 and
# the third with pages 21-end
pdf-helper split my-pdf.pdf my-split-pdfs -s 10,20

# E.g. Split a PDF into PDFs each containing one page
pdf-helper split my-pdf.pdf my-split-pdfs  # No need to specify split points

Export PDF pages as image files

Export PDF pages as image files:

pdf-helper to-image <input_file> <output_folder> \
        -p <page_number_1>,<page_number_2>,...,<page_number_n> -s <scale_factor>

# E.g. Export pages 1, 2, 3 and 6 from a PDF with scale factor 1
pdf-helper to-image 1.pdf images -p 1:3,6 -s 1

# E.g. Export the last three pages with scale factor 1
pdf-helper to-image 1.pdf images -p -3:-1 -s 1

# E.g. Export all pages from a PDF with scale 2
pdf-helper to-image my-pdf.pdf my-images

Remove pages from a PDF

Remove pages from a PDF:

pdf-helper remove-pages <input_file> <output_file> <page_number_1>,<page_number_2>,...,<page_number_n>

# E.g. Remove pages 1, 2, 3 and 6 from a PDF
pdf-helper remove-pages 1.pdf new.pdf 1:3,6

Add watermark to a PDF

Overlay selectable vector text or an image file on PDF pages:

# E.g. Stamp diagonal DRAFT text across all pages
pdf-helper watermark add my-pdf.pdf watermarked.pdf DRAFT

# E.g. Overlay a logo in the bottom-right corner of pages 1:3
pdf-helper watermark add my-pdf.pdf watermarked.pdf --watermark-image logo.png \
    --position bottom-right --image-scale 0.5 --pages 1:3

# E.g. Stamp 5% above the bottom edge, at the left edge
pdf-helper watermark add my-pdf.pdf watermarked.pdf "DRAFT" --position "0 -5%"

# E.g. Stamp text with a custom font (file path or installed name like Arial)
pdf-helper watermark add my-pdf.pdf watermarked.pdf "DRAFT" --font ./fonts/MyFont.ttf

# E.g. Stamp red text (CSS name, hex, or RGB tuple)
pdf-helper watermark add my-pdf.pdf watermarked.pdf "DRAFT" --color red

# E.g. Scale the text with the page: '50%' spans half the page width
pdf-helper watermark add my-pdf.pdf watermarked.pdf "DRAFT" --font-size "50%" --rotation 0

# E.g. Remove watermarks added by pdf-helper
pdf-helper watermark clear watermarked.pdf clean.pdf

# E.g. Swap one watermark for another in a single run
pdf-helper watermark replace watermarked.pdf new.pdf NEWMARK --remove-text DRAFT

Watermarks placed by pdf-helper carry an invisible mark, so clear removes exactly those; with text, matching text objects are stripped too (best effort).

--font-size accepts points (48) or a percent ("50%", quoted so the shell passes it through): a percent scales the text so its width spans that fraction of the page extent along the watermark direction — '50%' with --rotation 0 is a horizontal text half as wide as the page. Percents resolve per page, keeping mixed portrait/landscape documents proportional. Recipes accept the same values in font_size.

Text watermarks use Arial when it is installed (covering Latin, Greek, Cyrillic, Arabic and more), falling back to built-in Helvetica (Latin only) otherwise. RTL text (Persian/Arabic) is shaped into visual order automatically:

# E.g. Stamp a Persian watermark with an installed font
pdf-helper watermark add my-pdf.pdf watermarked.pdf "سلام" --font "Vazirmatn Regular"

Not sure which fonts are available? Query them (name or path both work with --font):

pdf-helper list-fonts
pdf-helper list-fonts arial

See examples/recipes/watermark-showcase.yaml for a recipe that chains a text watermark and an image watermark in one run.

Manage PDF metadata

Read, edit, or clear the metadata of a PDF file:

# E.g. Show all set metadata fields (accepts multiple files)
pdf-helper metadata show my-pdf.pdf other.pdf

# E.g. Set the title and author
pdf-helper metadata edit my-pdf.pdf tagged.pdf --title "Report" --author "Me"

# E.g. Remove all metadata (Info dictionary and XMP packet)
pdf-helper metadata clear my-pdf.pdf clean.pdf

Settable fields: title, author, subject, keywords, creator, producer, creationdate, moddate. Dates accept PDF date strings, ISO dates (2026-09-04, 2026-09-04 12:30), or datetimes, and show prints them in ISO format. Writes keep the XMP packet synchronized with the Info dictionary (and show prefers XMP values when present). Every file pdf-helper writes gets Producer (PDF-Helper <version>, unless already set), a CreationDate if missing, and a fresh ModDate automatically.

Encrypt and decrypt PDFs

Protect a PDF with a password, or remove the protection again:

# E.g. Require a password to open the file
pdf-helper encrypt my-pdf.pdf secure.pdf --password s3cret

# E.g. Acrobat-style lock: anyone may open, only the owner may edit
pdf-helper encrypt my-pdf.pdf secure.pdf --password "" --owner-password admin

# E.g. Decrypt with either password
pdf-helper decrypt secure.pdf open.pdf --password s3cret

The user password opens the file; the owner password (defaults to the user password) additionally grants full control. Passwords can also come from the PDF_HELPER_PASSWORD / PDF_HELPER_OWNER_PASSWORD environment variables, and are never echoed back. Algorithms: AES-256 (default), AES-128, RC4 (recipes additionally support per-operation permissions).

Every command that reads PDFs accepts --password (or PDF_HELPER_PASSWORD) for encrypted inputs — including metadata edit/clear, watermark, extract-text, and friends. In recipes, add a password: key (plain value or {input: name} reference) to any step that opens an encrypted file.

Export text from a PDF

To extract text from a PDF file and export them to text files you can do as follows:

pdf-helper extract-text <input_file> -o <output_file_name>

# E.g. Extract text from a PDF named my-pdf.pdf and save it to my-text.txt
pdf-helper extract-text my-pdf.pdf -o my-text.txt

Run Recipes

The recipe system lets you chain multiple PDF operations together in a single run using a YAML file. This unlocks features not available through individual CLI commands (e.g. selecting specific pages per file when bundling).

pdf-helper recipe run <recipe_file.yaml>

# E.g. Run a simple recipe
pdf-helper recipe run remove-pages.yaml

# E.g. Run with force overwrite and verbose logging
pdf-helper recipe run bundle-workflow.yaml --force --verbose

# E.g. Run a built-in recipe template instead of a file
pdf-helper recipe run --builtin remove-pages --input report.pdf --output cleaned.pdf

# E.g. Run a custom recipe from ~/.local/pdf-helper/recipes by name
pdf-helper recipe run --custom my-clean --input report.pdf --output cleaned.pdf

Custom recipes are plain YAML recipe files stored in ~/.local/pdf-helper/recipes (override with the PDF_HELPER_RECIPES_DIR environment variable) and run with the same --input/--output/--defines bindings. recipe generate lists them alongside the built-in templates.

Parameterized Recipes

Recipes can take their main input/output paths (and any extra values) from the command line instead of hardcoding them, so one recipe works on many files. Write {input}, {output}, or any {name} placeholder in the YAML and bind it at runtime:

name: "Remove specific pages"
version: "1.0"
parameters: [input, output]

steps:
  - id: clean
    operation: remove_pages
    input: "{input}"
    pages_to_remove: [2, 4, 6]
    output: "{output}"
pdf-helper recipe run clean.yaml --input report.pdf --output cleaned.pdf
pdf-helper recipe run clean.yaml --input a.pdf --output a-clean.pdf \
    --defines pages=1:3 --defines mode=strict

--input/--output bind {input}/{output}; repeatable --defines KEY=VALUE binds any other {name}. The optional parameters: list declares what a recipe expects — bare names are required, while name: default mappings provide defaults that CLI flags override:

parameters:
  - input
  - output
  - label: PDF-Helper
  - pages: "1,-1"

Path parts are available as {input.stem}, {input.name}, {input.suffix}, and {input.parent}, so outputs can be named after the input (e.g. output: "{input.stem}-pages"). Defaults may reference other placeholders, so - output: "{input.stem}-stamped-pages" works as an overridable default.

Repeat --input to pass multiple files as one array input: a bare {input} item inside a step list (e.g. bundle inputs: ["{input}"]) expands to every file. Array inputs need scalar outputs (path attributes like {input.stem} and embedding in longer strings are rejected with a clear error):

# E.g. Bundle two images, then stamp a label on every page
pdf-helper recipe run stamp-on-images.yaml --input a.png --input b.png \
    --output stamped.pdf --defines label=DRAFT

See examples/recipes/stamp-on-images.yaml for the full recipe.

Page selections use comma-separated items where start:end selects an inclusive range: 1:3,6 means pages 1, 2, 3 and 6. Numbers are 1-based (0 is rejected), and negatives count back from the last page, so 5:-2 means from page 5 to one page before the last. Values starting with - must use the --option=value form (--pages="-2:-1") so the CLI does not mistake them for flags. Built-in templates (recipe generate lists them) accept bindings the same way via --builtin <name>.

Recipe File Format

A recipe is a YAML file with a steps list. Each step has an id, an operation, input/output paths, and operation-specific options. Steps can reference each other's outputs using { step: step_id }.

name: "Remove specific pages"
description: "Removes pages 2, 4, 6 from a PDF."
version: "1.0"

steps:
  - id: clean
    operation: remove_pages
    input: document.pdf
    pages_to_remove: [2, 4, 6]
    output: cleaned.pdf

Supported Operations

Operation Status Description
bundle Available Bundle files with optional per-file page selection
remove_pages Available Remove pages by 1-based index
split_pdf Available Split at given page boundaries
pdf_to_image Available Render pages as PNG images
extract_text Available Extract text content
watermark Available Overlay vector text (text) or an image (image)
clear_watermark Available Remove pdf-helper watermarks (plus text matches)
replace_watermark Available Clear old watermarks and add new ones in one run
encrypt Available Password-protect PDF (AES-256/128, RC4)
decrypt Available Remove password protection
metadata Available Set title/author/keywords (+creator/dates)
clear_metadata Available Remove all metadata (Info + XMP)

Operations marked Planned are not yet implemented — the recipe runner logs a warning and copies the input file through, so pipelines don't break.

Watermark text is inserted as selectable vector objects; image watermarks are raster overlays. Both support position, opacity, rotation, and pages.

Advanced Example: Multi-step Pipeline

# yaml-language-server: $schema=https://gitlab.com/CodeWriter21/pdf-helper/-/raw/master/schemas/recipe-schema.json
name: "Split, Convert, and Extract Pipeline"
version: "1.0"

settings:
  temp_dir: "./.recipe-tmp"

steps:
  # Step 1: Split the PDF at pages 5 and 10
  - id: split
    operation: split_pdf
    input: report.pdf
    split_points: [5, 10]
    output_dir: .
    output_prefix: "report_part_"

  # Step 2: Convert the second chunk to images
  - id: to_images
    operation: pdf_to_image
    input:
      step: split
      file: report_part_2.pdf
    pages: "1:3"
    scale: 3
    output: ./output/images

  # Step 3: Extract text from the first chunk
  - id: extract
    operation: extract_text
    input:
      step: split
      file: report_part_1.pdf
    pages: "1:4"
    max_characters: 5000
    reverse_lines: true
    output: ./output/chapter-1-text.txt

Recipe Settings

Setting Default Description
temp_dir ./.recipe-tmp Directory for intermediate files
overwrite false Overwrite existing output files
cleanup_temp false Remove temp directory after completion

Input values (e.g. passwords) can be sourced from environment variables or prompted at runtime:

inputs:
  password:
    env: PDF_PASSWORD
    prompt: "Enter output PDF password"

See examples/recipes/ for more example recipe files.

About

Author: CodeWriter21

GitLab: CodeWriter21/pdf-helper

Donations

Your donations are very welcome: nowpayments.io

You can also consider donating a Star to the repo.

License

This project is licensed under the MIT License.

See the LICENSE

References

Release files for pdf-helper 0.9.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for pdf-helper 0.9.0
File Size Uploaded
pdf_helper-0.9.0.tar.gz 38.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for pdf-helper 0.9.0
File Interpreter ABI Platform
pdf_helper-0.9.0-py3-none-any.whl Python 3 none any Details

Total release size: 78.3 kB

Release files / pdf_helper-0.9.0.tar.gz

Download URL pdf_helper-0.9.0.tar.gz
Size 38.1 kB
Tags Source
SHA-256 checksum
How to use checksums
ec4a8c89c4088311985aed321b1f4397c7f9c79b6cc41d86073440cbdf50327b
BLAKE2b-256 checksum
How to use checksums
5cb4afaae9d79691d5220116298c169301fe3039c518e2754aa9eaa1af4a999f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.19 {"installer":{"name":"uv","version":"0.12.19","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Debian GNU/Linux","version":"12","id":"bookworm","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release files / pdf_helper-0.9.0-py3-none-any.whl

Download URL pdf_helper-0.9.0-py3-none-any.whl
Size 40.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
3bbf3be3d0560e56af8e55a4274ebd2d9efd24f94239b65ad84395ad212c8252
BLAKE2b-256 checksum
How to use checksums
6a7464572bc28b95fa5477368b3e0e46eb74df563693d59353c5fe6ae5b25849
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.19 {"installer":{"name":"uv","version":"0.12.19","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Debian GNU/Linux","version":"12","id":"bookworm","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release history Release notifications | RSS feed

This release

0.9.0 This release

2 release files

0.8.0

2 release files

0.7.0

2 release files

0.6.0

2 release files

0.5.0

2 release files

0.4.0

2 release files

0.3.2

2 release files

0.3.1

2 release files

0.3.0

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page