pdftopdfa
pdftopdfa is a free and open-source alternative to Ghostscript-based PDF/A converters. Ghostscript uses a dual license (AGPL/commercial) that makes it difficult to use in commercial products without purchasing a license. pdftopdfa uses MPL-2.0-or-later, a file-level copyleft license. It permits commercial use and combination with proprietary code, provided its terms are followed. For non-OCR conversions, pdftopdfa modifies the PDF structure directly using pikepdf (based on QPDF), avoiding full-document re-rendering and preserving the original content, fonts, and layout where possible.
Highlights
- No Ghostscript required -- direct PDF manipulation via pikepdf/QPDF
- PDF/A-2a, 2b, 2u, 3a, 3b, 3u -- supports modern PDF/A levels (ISO 19005-2 and ISO 19005-3), including Tagged PDF output for scanned documents
- Automatic font embedding -- uses policy-approved Windows system fonts or bundled replacements
- Font subsetting -- reduces file size by removing unused glyphs
- CJK support -- embeds Noto Sans CJK for Chinese, Japanese, and Korean text
- ICC color profiles -- automatically embeds sRGB, CMYK, and grayscale profiles
- XRechnung metadata -- adds canonical Factur-X XMP metadata for recognized, unambiguous embedded XRechnung 3.0 invoices in PDF/A-3 output
- Batch processing -- converts entire directories, optionally recursive
- Integrated validation -- checks conformance via veraPDF
- OCR support -- optional PP-OCRv6 Medium text recognition on the CPU or through DirectML in the supported Windows 11 configuration, with external offline text-model directories, a bundled page-orientation model, and no runtime model downloads
- Layout-aware OCR -- optional column-based reading order for multi-column documents without an additional model or OCR pass
- Table recognition -- recognizes already-cropped bordered ("wired") and borderless ("wireless") tables as typed cells and HTML using only explicitly supplied local ONNX models
- Simple API -- usable as CLI tool or Python library
How It Works
pdftopdfa applies a multi-step conversion pipeline to make a PDF compliant with the PDF/A standard:
- Pre-check -- copies encrypted and, by default, digitally signed PDFs
unchanged; otherwise, detects if the PDF is already a valid PDF/A file
(skips conversion if the existing conformance level meets or exceeds the
target within the same PDF/A part; optionally skips any veraPDF-compliant
PDF/A via
--skip-any-pdfa; see the Usage Guide for details) - OCR (optional) -- optionally orients pages with the bundled PP-LCNet document-orientation model, straightens only scan-like raster-dominant pages, and recognizes text with externally supplied PP-OCRv6 Medium models; OCRmyPDF rasterizes OCR target pages and creates the searchable text layer
- Font compliance -- analyzes all fonts, embeds missing ones, adds ToUnicode mappings, subsets embedded fonts, and fixes encoding issues
- Sanitization -- removes or fixes non-compliant elements (JavaScript, non-standard actions, transparency groups, annotations, optional content, etc.)
- Metadata -- synchronizes XMP metadata with the document info dictionary and sets the PDF/A conformance level
- Color profiles -- detects color spaces and embeds the required ICC profiles (sRGB, CMYK/FOGRA39, sGray)
- Logical structure -- for level A, preserves an existing Tagged PDF structure or creates a structure tree from the final page and annotation order, including OCR-processed scans
- Save -- writes the output with the correct PDF version header
Installation
Prerequisites
- Python 3.12, 3.13, or 3.14
- macOS 14 or later on Apple Silicon, Linux, or Windows
Intel-based Macs are not supported. CPU OCR on macOS requires the ARM64 wheels provided by ONNX Runtime for Apple Silicon.
python -m pip install pdftopdfa
If the pdftopdfa console script is not on PATH, use
python -m pdftopdfa in the examples below.
Optional: PDF/A validation
Validation uses the external veraPDF application, which
is not bundled. Install it and make its launcher available on PATH, or set
VERAPDF_PATH to the executable or its parent directory, before using
--validate or validate=True.
Optional: OCR support
Install exactly one OCR runtime. CPU inference is the default:
python -m pip install "pdftopdfa[ocr]"
For the supported DirectML configuration on Windows 11:
python -m pip install "pdftopdfa[directml]"
Do not install both extras in the same Python installation: onnxruntime and
onnxruntime-directml provide overlapping runtime files. pdftopdfa supports
DirectML on Windows 11 with a DirectX 12-capable integrated or dedicated Intel,
AMD, or NVIDIA GPU and a current graphics driver.
OCR uses PaddleOCR 3.7 with the selected ONNX Runtime provider. Installing the
DirectML extra does not select it automatically; use
--ocr-execution-provider directml or
ocr_execution_provider="directml". CPU remains the default. If DirectML is
requested but unavailable, processing stops with an error instead of falling
back to the CPU.
On a machine with several GPUs, directml:<index> passes a raw DXGI adapter
index, for example --ocr-execution-provider directml:1. Plain directml uses
DirectML's default adapter. The internal diagnostic helper
pdftopdfa._ocr_runtime.list_directml_devices() lists the available adapters
and their raw indices; as part of a private module, it has no public API
stability guarantee. The indices may have gaps because software adapters are
omitted, and repeated DXGI entries with the same PCI identity are listed once
using their lowest index. Use the reported index, not its position in the
filtered list.
The page-orientation model is bundled. PP-OCRv6 text-recognition and table
models are external and are never downloaded at runtime. Pass their local
directories to each top-level conversion or recognition call; an OCRSession
instead receives the PP-OCRv6 text-model pair once when it is created and
reuses it across its image-recognition calls. CPU and DirectML use the same
FP32 ONNX model files.
See the OCR guide for the recognize_table() model contract and
typed result. Cell text and grid structure come from the table-structure model,
while bounding boxes come from the separate cell-detection model; if the two
models report different cell counts, cells are returned without
bounding_box and confidence instead of failing.
PP-OCRv6 model setup
The following model revisions are tested and recommended:
- Detection:
PP-OCRv6_medium_det_onnxat6132380 - Recognition:
PP-OCRv6_medium_rec_onnxat50c7eac
Each model directory must contain exactly inference.onnx and
inference.yml. Before initialization, pdftopdfa performs a quick
structural check that rejects missing or extra entries, non-regular files, and
symbolic links. PaddleOCR then loads the model files and checks that they are
compatible detection and recognition models. pdftopdfa does not verify the
repository revision or model-file hashes.
The models are not included in the source distribution or wheel. Keep them in
deployment-managed, read-only directories. Both
--ocr-detection-model-dir and --ocr-recognition-model-dir are required
together; supplying the pair enables OCR without an additional --ocr flag.
Conversely, --ocr, --ocr-force, --deskew, --rotate-pages,
--ocr-layout, and a non-CPU --ocr-execution-provider value (directml or
directml:INDEX) are rejected unless both model options are present.
--ocr-lang defaults to en. Use de for German and de+en for mixed
German/English recognition. Latin-script languages restrict decoding to Latin
letters while retaining numbers, punctuation, and symbols, which prevents
Chinese-character output on German scans. The accepted PaddleOCR 3.7 codes are:
af, az, bs, ca, ch, chinese_cht, cs, cy, da, de, en,
es, et, eu, fi, fr, french, ga, german, gl, hr, hu,
id, is, it, japan, ku, la, lb, lt, lv, mi, ms, mt,
nl, no, oc, pl, pt, qu, rm, ro, rs_latin, sk, sl,
sq, sv, sw, tl, tr, uz, vi.
Legacy codes such as eng and deu are not accepted. See the
PaddleOCR language documentation
for the language families represented by these codes.
Quick Start
# Simple conversion (creates document_pdfa.pdf)
pdftopdfa document.pdf
# Specific PDF/A level
pdftopdfa -l 2b document.pdf
# Accessible PDF/A output
pdftopdfa -l 2a document.pdf
# With validation (note: -v = --validate, not verbose; use --verbose for logs)
pdftopdfa -v document.pdf
# Skip any existing veraPDF-compliant PDF/A
pdftopdfa --skip-any-pdfa document.pdf
# Explicitly convert a signed PDF, invalidating its digital signatures
pdftopdfa --allow-signature-invalidation document.pdf
# Convert an entire directory
pdftopdfa -r ./documents/ ./output/
# The OCR examples below use the externally managed model directories
DET_MODEL=/opt/pdftopdfa/models/PP-OCRv6_medium_det_onnx
REC_MODEL=/opt/pdftopdfa/models/PP-OCRv6_medium_rec_onnx
# OCR a German/English scan to tagged PDF/A-2a
pdftopdfa -l 2a --ocr-lang de+en \
--ocr-detection-model-dir "$DET_MODEL" \
--ocr-recognition-model-dir "$REC_MODEL" \
document.pdf
# Order OCR lines by detected columns for a cleaner reading order
pdftopdfa --ocr-layout \
--ocr-detection-model-dir "$DET_MODEL" \
--ocr-recognition-model-dir "$REC_MODEL" \
document.pdf
# Use the same models through DirectML on Windows 11
pdftopdfa --ocr-execution-provider directml \
--ocr-detection-model-dir "$DET_MODEL" \
--ocr-recognition-model-dir "$REC_MODEL" \
document.pdf
# Automatically orient pages without deskewing them
pdftopdfa --rotate-pages \
--ocr-detection-model-dir "$DET_MODEL" \
--ocr-recognition-model-dir "$REC_MODEL" \
document.pdf
# Deskew pages without changing their 90-degree orientation
pdftopdfa --deskew \
--ocr-detection-model-dir "$DET_MODEL" \
--ocr-recognition-model-dir "$REC_MODEL" \
document.pdf
# Deskew and orient pages without converting the result to PDF/A
# (creates document_processed.pdf)
pdftopdfa --no-pdfa --deskew --rotate-pages \
--ocr-detection-model-dir "$DET_MODEL" \
--ocr-recognition-model-dir "$REC_MODEL" \
document.pdf
# Preserve known proprietary stamps as PDF Stamp annotations
pdftopdfa --preserve-stamps document.pdf
The OCR examples above use POSIX shell syntax. For DirectML in PowerShell on Windows 11, for example:
$DET_MODEL = "C:\models\PP-OCRv6_medium_det_onnx"
$REC_MODEL = "C:\models\PP-OCRv6_medium_rec_onnx"
pdftopdfa --ocr-execution-provider directml `
--ocr-detection-model-dir "$DET_MODEL" `
--ocr-recognition-model-dir "$REC_MODEL" document.pdf
from pathlib import Path
from pdftopdfa import convert_to_pdfa
result = convert_to_pdfa(
input_path=Path("input.pdf"),
output_path=Path("output.pdf"),
level="2b",
)
ocr_result = convert_to_pdfa(
input_path=Path("scan.pdf"),
output_path=Path("scan_pdfa.pdf"),
level="2b",
ocr_languages=["de", "en"],
ocr_detection_model_dir=Path(
"/opt/pdftopdfa/models/PP-OCRv6_medium_det_onnx"
),
ocr_recognition_model_dir=Path(
"/opt/pdftopdfa/models/PP-OCRv6_medium_rec_onnx"
),
ocr_execution_provider="cpu",
)
Supplying both model directories enables OCR in convert_to_pdfa(),
convert_files(), and convert_directory(). Supplying only one directory, or
requesting OCR through ocr_languages, ocr_force, ocr_deskew,
ocr_rotate_pages, ocr_layout=True, or a non-CPU execution provider without
both directories, raises ValueError before processing starts. Set
ocr_execution_provider="directml" to use the supported DirectML configuration
on Windows 11, or ocr_execution_provider="directml:1" to pass a specific raw
DXGI adapter index.
Set pdfa=False to apply only the requested OCR processing. This skips font
embedding, PDF/A sanitization, metadata synchronization, color-profile
embedding, and PDF/A validation. The result is not validated or guaranteed to
be PDF/A compliant.
See the Usage Guide
for the full CLI reference, conversion API documentation, and examples. The
OCR guide covers
image, table, and reusable OCRSession APIs.
Limitations
- No PDF/A-1 support -- only PDF/A-2 and PDF/A-3 levels are supported
- Automatic level A semantics -- generated tags follow page and PDF content-stream order and include annotations. They provide the structural basis required by PDF/A-2a and PDF/A-3a, but automatic conversion cannot infer authorial semantics such as heading levels, table relationships, or alternative descriptions. PDF/A level A does not imply PDF/UA conformance.
- Encrypted PDFs -- password-protected PDFs cannot be converted and are
copied unchanged. With an automatically generated output name, the unchanged
copy still receives the
_pdfa.pdfsuffix; it is not a converted PDF/A file - Digitally signed PDFs -- signed PDFs are copied unchanged by default because conversion would invalidate their signatures; use
--allow-signature-invalidationonly when an unsigned archival copy is intentional - Font replacement -- fonts without a suitable metrically compatible replacement produce a warning; the resulting file may not be fully compliant
- Non-embedded CIDFonts (Identity encoding) -- content streams reference glyph IDs of the original font; after replacement with a substitute font the same glyph IDs point to different or missing glyphs, so the affected text may render incorrectly or invisibly. Text extraction and copy/paste stay correct because the original ToUnicode mapping is preserved. A warning is emitted for each replaced CIDFont
Font Sourcing
- On Windows,
pdftopdfamay automatically embed a conservative fixed allowlist of local fonts from%WINDIR%\Fonts. - A Windows system font is only used when the installed file lives under
%WINDIR%\Fonts, its actual PostScript name is allowlisted, and its OpenTypefsTypepermits outline embedding. - On macOS and Linux, system fonts are never auto-embedded; bundled replacement fonts are used instead.
fsTypechecks are a technical safeguard only and do not replace the font vendor's EULA or other license terms.- For auditable deployments, keep the allowlist tied to reviewed target systems or golden images.
Development
python -m pip install -e ".[dev,ocr]"
Running Tests
python -m pytest
The test suite covers fonts, color profiles, metadata, sanitization, OCR, and end-to-end conversion.
Code Quality
ruff check src/ tests/ # Linting
ruff format src/ tests/ # Formatting
Documentation
Additional documentation is available in the docs/ folder:
Contributing
Contributions are welcome! Please open an issue to report bugs or suggest features, or submit a pull request.
Dependencies
Core:
- pikepdf -- PDF manipulation (based on QPDF)
- lxml -- XMP metadata processing
- fonttools -- Font analysis, subsetting, and embedding
- pdfminer.six -- CMap decoding for font-to-Unicode mappings
- NumPy -- Array processing for OCR and table recognition
- click -- CLI framework
- colorama -- Colored terminal output
- tqdm -- Progress bars
- PaddleOCR -- document orientation and PP-OCRv6 text recognition
Optional:
- OCRmyPDF -- PDF rasterization, text-layer generation, and page merging for optional OCR
- ONNX Runtime -- CPU or DirectML inference for Paddle models
- PaddleX -- local-model OCR and table-recognition pipelines
- pypdfium2 -- PDF page rasterizer for OCR
- veraPDF -- external application for ISO-compliant PDF/A validation
Acknowledgments
This project bundles the following resources:
- Liberation Fonts -- metrically compatible replacements for the PDF Standard 14 fonts (SIL OFL 1.1)
- PP-LCNet_x1_0_doc_ori -- bundled document-orientation model (Apache-2.0)
- Noto Sans CJK -- CJK font coverage (SIL OFL 1.1)
- Noto Sans Symbols 2 -- symbol font replacement (SIL OFL 1.1)
- STIX Two Math -- math font replacement (SIL OFL 1.1)
- sRGB2014.icc -- ICC sRGB profile (ICC)
- ISOcoated_v2_300_bas.icc -- ICC CMYK profile, FOGRA39 (zlib/libpng license)
- sGray -- compact grayscale ICC profile (CC0-1.0)
- Adobe cmap-resources -- CID-to-Unicode mapping data (BSD 3-Clause)
License
This project is licensed under the Mozilla Public License 2.0 or later (MPL-2.0+) -- see LICENSE for details.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file pdftopdfa-0.9.3.tar.gz.
File metadata
- Download URL: pdftopdfa-0.9.3.tar.gz
- Upload date:
- Size: 27.4 MB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
91e97f2c6ccddf9ec7165d3577e9fc756706e256b4b49ed0e7ba1e8860da487b
|
|
| MD5 |
31983d27a2c75a893821ed51dfe66999
|
|
| BLAKE2b-256 |
f9f3c68bd02c0371a1636933a4862adce06eefeebfc64865610df1985f8eba3e
|
Provenance
The following attestation bundles were made for pdftopdfa-0.9.3.tar.gz:
Publisher:
publish.yml on iRedPaul/pdftopdfa
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
pdftopdfa-0.9.3.tar.gz -
Subject digest:
91e97f2c6ccddf9ec7165d3577e9fc756706e256b4b49ed0e7ba1e8860da487b - Sigstore transparency entry: 2497399705
- Sigstore integration time:
-
Permalink:
iRedPaul/pdftopdfa@9dbd900d69b7bb97234b8568cd8230bed82b01a4 -
Branch / Tag:
refs/tags/v0.9.3 - Owner: https://github.com/iRedPaul
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@9dbd900d69b7bb97234b8568cd8230bed82b01a4 -
Trigger Event:
release
-
Statement type:
File details
Details for the file pdftopdfa-0.9.3-py3-none-any.whl.
File metadata
- Download URL: pdftopdfa-0.9.3-py3-none-any.whl
- Upload date:
- Size: 27.0 MB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
513ce215530c1e4aba4d9182896268201910bbaefaa0cf423dabc89a71b88372
|
|
| MD5 |
92fb51468513780484ef81d86ec13338
|
|
| BLAKE2b-256 |
9ab1202b4f62c787a2e0d07955ff54999176a09444b8cb31122f345f488c1a74
|
Provenance
The following attestation bundles were made for pdftopdfa-0.9.3-py3-none-any.whl:
Publisher:
publish.yml on iRedPaul/pdftopdfa
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
pdftopdfa-0.9.3-py3-none-any.whl -
Subject digest:
513ce215530c1e4aba4d9182896268201910bbaefaa0cf423dabc89a71b88372 - Sigstore transparency entry: 2497400128
- Sigstore integration time:
-
Permalink:
iRedPaul/pdftopdfa@9dbd900d69b7bb97234b8568cd8230bed82b01a4 -
Branch / Tag:
refs/tags/v0.9.3 - Owner: https://github.com/iRedPaul
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@9dbd900d69b7bb97234b8568cd8230bed82b01a4 -
Trigger Event:
release
-
Statement type: