pdfua
Check the machine-checkable subset of PDF/UA-1 and WCAG 2.1 for PDFs — a small
Python package built on pikepdf (QPDF), with no JVM.
pip install pdfua
pdfua check report.pdf --format sarif
This implements 12 checks covering 16 of the 106 machine-checkable rules in the PDF/UA-1 profile — 15%. Read What this does NOT check before you rely on it. It is a fast triage gate, not a conformance certifier. For a conformance claim, run veraPDF.
The problem
There is a dated legal forcing function and no tool that fits a CI pipeline.
The U.S. Department of Justice's Title II rule requires state and local
governments to meet WCAG 2.1 Level AA for their web content and mobile apps.
PDFs are explicitly named as covered content. The compliance dates, from
ada.gov:
| State and local government size | Compliance date |
|---|---|
| 50,000 or more persons | April 26, 2027 |
| 0 to 49,999 persons | April 26, 2028 |
| Special district governments | April 26, 2028 |
"UPDATE: On April 20, 2026, the Department published an Interim Final Rule (IFR) extending the compliance date for State and local government entities with a total population of 50,000 or more to April 26, 2027. The compliance date for public entities with a total population of less than 50,000, or any special district government, is extended to April 26, 2028."
"The documents are word processing, presentation, PDF, or spreadsheet files" — https://www.ada.gov/resources/2024-03-08-web-rule/
So there is a hard deadline, a large population of PDFs that must comply, and consequently a need to check PDFs automatically.
Why the existing options do not fit
| Option | Why it does not work in CI |
|---|---|
| veraPDF — the authoritative validator | Java. Requires a JVM on every build agent. GPL-3.0 (dual MPLv2+). |
| opendataloader-pdf | The PyPI package is "a Python wrapper for the opendataloader-pdf Java CLI". Its free tier is auto-tagging (producing a Tagged PDF); PDF/UA export is an enterprise add-on. It is a remediator, not an auditor. |
| pdf-a11y | A remediation pipeline whose validator.py is a subprocess wrapper around the veraPDF binary. It needs the JVM too. |
| wcag_pdf_pytest | Regex-scans raw PDF bytes for /StructTreeRoot and returns hardcoded passes for most criteria ("Heuristic pass: … not applicable"). It reports PASS on criteria it never examined. |
| Commercial services | Per-document pricing; not a build gate. |
pdfua occupies the empty slot: a pip install, no JVM, structured
output, exit codes. It checks the failure classes that actually dominate real government
PDFs, and it says plainly what it did not check.
What the evidence showed
This library was built only after a falsification pass against veraPDF 1.30.2 on
16 real government PDF files from govinfo.gov, irs.gov and ada.gov (15
distinct documents — see below). The full
evidence is in docs/FALSIFICATION.md. The two findings
that shaped the design:
- veraPDF reports 0 of 16 compliant. All 16 fail PDF/UA-1.
- 99.55% of the 643,471 failed checks are two rules —
7.1-3("content shall be marked as Artifact or tagged as real content") and7.2-34("natural language for text in page content shall be determined"). Both are reachable in Python without a JVM, and both are implemented here. A prototype's untagged-text counter tracked veraPDF's7.1-3at 92–99.9% agreement (e.g. 86,118 vs 86,131 on one 1,039-page document).
Those totals cover 16 validated files, but only 15 distinct documents: two
corpus files (p08.pdf and 7ef8534308.pdf) are byte-identical (sha256
9f5806…a853), so the 643,471 figure double-counts 172,257 checks. Across the 15
distinct documents the de-duplicated total is 471,214, of which the same two
rules account for 468,314 — 99.38%, down from 99.55%. The concentration claim
holds either way, which is what the design argument rests on.
That is the honest case for this tool: the failures that dominate real documents are cheap to detect and currently require a JVM to detect.
Install
pip install pdfua
# or with uv:
uv pip install pdfua
To track main instead of a release:
pip install "git+https://github.com/beduldul/pdfua.git"
Requires Python 3.10+ and pikepdf. No JVM, no Java, no external binary.
Why pikepdf
A PDF/UA check has to walk the COS object graph: /StructTreeRoot, /K child
chains, indirect references, content streams. pikepdf binds QPDF, ships binary
wheels for Linux/macOS/Windows (so pip install needs no compiler), and exposes
references rather than flattening them. pypdf was the alternative; it hides
indirect references, which PDF/UA rules care about. You are not expected to
write a PDF parser, and this package does not contain one.
Usage
CLI
pdfua check document.pdf
pdfua check *.pdf --format sarif > results.sarif
pdfua check document.pdf --rules UA-01,UA-13-004
pdfua check document.pdf --min-severity warning --certain-only
pdfua rules # what is implemented, and what is not
Exit codes
| Code | Meaning |
|---|---|
0 |
No findings |
1 |
Findings at INFO severity |
2 |
Findings at WARNING severity |
3 |
Findings at ERROR severity |
4 |
The file could not be read as a PDF |
With --min-severity set, any surviving finding exits 1. Without it, the exit
code is ordered by severity, so the obvious thing works:
pdfua check report.pdf || echo "PDF/UA problems found"
Library
from pdfua import validate, Severity
report = validate("document.pdf")
print(report.exit_code) # 0 clean, 1 info, 2 warning, 3 error
print(report.worst_severity()) # Severity.ERROR or None
for finding in report.findings:
print(finding.rule_id, finding.severity.name, finding.location.describe())
print(" ", finding.message)
print(" fix:", finding.remediation)
Filtering, and the coverage figure:
from pdfua import coverage, describe_rules
from pdfua.model import filter_findings, Confidence
cov = coverage()
print(cov.summary()) # "16 of 106 PDF/UA-1 rules (15%) across 12 implemented checks"
certain = filter_findings(report.findings, min_confidence=Confidence.CERTAIN)
What this checks
12 checks, covering 16 PDF/UA-1 rule identifiers. Each is decidable from the PDF
alone, with no interpretation. pdfua rules prints this table from the registry,
so the numbers cannot drift from the code.
| Rule | PDF/UA-1 clause | WCAG | Severity | What it decides |
|---|---|---|---|---|
UA-01-005 |
§7.1 | 1.3.1 | error | /StructTreeRoot is present |
UA-01-002 |
§6.2 | 1.3.1 | error | /MarkInfo /Marked is true |
UA-01-003 |
§7.1 | 1.3.1 | error | Page content is tagged or marked as an artifact |
UA-06-001 |
§7.2 | 3.1.1 | error | Catalog /Lang is present and well-formed |
UA-07-001 |
§7.1 | 2.4.2 | warning | A document /Title exists |
UA-07-002 |
§7.1 | 2.4.2 | warning | DisplayDocTitle is true |
UA-13-004 |
§7.3 | 1.1.1 | error | Every /Figure has non-empty /Alt or /ActualText |
UA-14-001 |
§7.4 | 1.3.1, 2.4.6 | warning | No empty heading elements |
UA-15-003 |
§7.5 | 1.3.1 | error | Table structure is determinable |
UA-18-001 |
§7.18.1 | 1.1.1, 4.1.2 | warning | Annotations have descriptions |
UA-18-005 |
§7.18.5 | 2.4.4 | warning | Link annotations have descriptions |
UA-28-004 |
§5 | — | warning | PDF/UA identification in XMP |
What this does NOT check
A validator that overstates its coverage is worse than no validator. This section is the most important one in this file.
It implements 12 checks, covering 16 of 106 rules
The public veraPDF PDF/UA-1 profile enumerates 106 machine-checkable rules for
ISO 14289-1. This tool implements 12 checks, which decide 16 of those
rules. pdfua rules prints both counts from the registry, so neither can drift.
The other 90 rules are not checked. A clean run means "these 16 rules found nothing", never "this document conforms".
Coverage is counted at rule granularity, not per section, and that distinction
is the point. Clause 7.2 alone contains 41 rules; implementing one /Lang
check does not make the other forty checked. A tool that reported "7.2: 41/41"
because of a single check would be producing exactly the false comfort this
project exists to avoid.
Notably not checked:
- Font embedding,
CIDSet, glyph widths, font descriptors (§7.21). These are the rules veraPDF uses to catch professionally-remediated files —ada.gov's own web-rule PDF is105/106clean and fails only7.21.4.2-2(CIDSet). This tool will report that document as clean, because it does not look at fonts. - Reading order correctness (§7.2). The tool confirms content is marked; it cannot judge whether the order is meaningful.
- Colour spaces, output intents, transparency, ICC profiles.
- WTPDF 1.0 and PDF/UA-2 (ISO 14289-2 / ISO 32005).
- PDF/A of any flavour.
- Encrypted documents. They raise, rather than guess.
It cannot judge quality, only presence
This is the fundamental line, and the tool never crosses it:
| Machine-checkable (this tool) | Requires a human |
|---|---|
A /Figure has an /Alt entry |
Whether the alt text is accurate |
The /Alt entry is non-empty |
Whether it is appropriate in context |
A table has /TH cells |
Whether the header association is correct |
| Content is marked | Whether the reading order is sensible |
| A title exists | Whether the title is descriptive |
A figure whose alt text is " " (a single space) is structurally present and
semantically worthless. The tool flags the empty case (UA-13-004) but cannot
tell you that "image" or "figure 3" is a bad description. Automated checks
reduce manual review; they do not replace it.
Findings marked heuristic are questions, not verdicts
Three finding sites carry confidence: heuristic, and in each case it is the
site, not the whole rule, that is a judgement call:
UA-07-001(DocumentTitleRule) — every finding it emits is heuristic (src/pdfua/rules/language.py:60sets the class-level confidence).UA-13-004(FigureAltTextRule) — only the empty or whitespace-only/Altbranch, which is stricter than the letter of §7.3 and that veraPDF passes (src/pdfua/rules/structure.py:73). A figure with no/Altat all is a certain finding.UA-15-003(TableHeaderRule) — only the association branch, where header cells exist but none carry/Scopeor/Headers(src/pdfua/rules/structure.py:144). The rule itself has class-level severityERROR, and its other branches are certain.
They narrow the question for a human. Use --certain-only to see only findings
the tool is certain about.
A mistake this project made, recorded on purpose
An early version reported "TH cells lack Scope" on documents veraPDF passes.
That was wrong about the standard: ISO 14289-1 §7.5 requires Scope only if
the structure is not determinable via Headers and IDs. It is a fallback, not
a universal requirement. The IRS W-9 has /TH cells with neither, and is valid.
The rule was removed. It is documented in docs/FALSIFICATION.md and in the
structure.py module docstring because this is exactly the failure mode that
makes accessibility tooling distrusted, and the credibility of this tool depends
on not repeating it.
For a conformance claim, use veraPDF
pdfua and veraPDF are complements, not competitors.
# Every commit: fast, no JVM
pdfua check report.pdf --format sarif || true
# Before you claim conformance: the authority
verapdf --flavour ua1 --format json report.pdf
pdfua gives you a JVM-free gate that produces SARIF for code scanning. Runtime
scales with document size and structure, so treat any single figure with care:
measured on this project's corpus (median of five runs) it is ~2–3 ms for a
184 KB document and ~1.4 s for the 3.9 MB ada.gov web-rule PDF — fast on
typical documents, seconds on very large ones, and always far below the cost of
starting a JVM. veraPDF remains the reference implementation. If your CI has a JVM and
you need a conformance claim, use veraPDF and do not install this. The tool
exists for the pipelines where a JVM is the reason nothing is checked at all.
Development
uv venv && uv pip install -e ".[dev]"
.venv/bin/pytest
.venv/bin/mypy src/pdfua
.venv/bin/ruff check .
Test fixtures that violate exactly one rule each are generated in code by
tests/fixtures/build.py using pikepdf; no binary PDFs are committed. See
tests/fixtures/README.md.
License
MIT. See LICENSE.
The bundled pdfua1_rules.json contains rule identifiers only (ISO clause
number, test number, validation-object name) — no descriptive text is reproduced
from any validation profile, because that text belongs to its authors.
Metadata
Release files for pdfua 0.1.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| pdfua-0.1.2.tar.gz | 51.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| pdfua-0.1.2-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 93.6 kB
Release files / pdfua-0.1.2.tar.gz
| Download URL | pdfua-0.1.2.tar.gz |
|---|---|
| Size | 51.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
8885129ff3125cacc99acf6daa3f344b688e88b6f523a4114ed1f7387304f47f
|
|
BLAKE2b-256 checksum How to use checksums |
d941fb938b34b4cf2e35968a6685b3242a8744f3eedd2e9438bf610554ccfdd5
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 4, 2026.
Transparency logRelease files / pdfua-0.1.2-py3-none-any.whl
| Download URL | pdfua-0.1.2-py3-none-any.whl |
|---|---|
| Size | 41.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
f849fd9af358fe2207f702ee9b552e0087c7ca3d1bd18df2bc5006424fb36330
|
|
BLAKE2b-256 checksum How to use checksums |
22dc4c0336c0e87d142535ec940c8a981b77e8087e6104f3d2ecbae0f9786938
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 4, 2026.
Transparency log