A Python library to inspect and modify the internal structure of a PDF file

These details have been verified by PyPI

Project links

GitHub Statistics

Maintainers

desgeeko

These details have not been verified by PyPI

Development Status
- 4 - Beta
Intended Audience
- Developers
License
- OSI Approved :: MIT License
Operating System
- OS Independent
Programming Language
- Python :: 3
Topic
- Software Development :: Libraries
- Utilities

Project description

PDFSyntax

A Python library to inspect and transform the internal structure of PDF files

Introduction

The project is focused on chapter 7 ("Syntax") of the Portable Document Format (PDF) Specification. It implements all the detailed document structure management down to the byte level for inspection and transformation use cases (access to metadata, rotation,...).

Internal functions are being exposed as an API toolkit for PDF read/write operations,
Some specific functions are additionally exposed as a command line interface for use in a terminal or a browser.

PDFSyntax is lightweight (no dependencies) and written from scratch in pure Python, with a focus on simplicity and immutability.

It favors non-destructive edits allowed by the PDF Specification: by default incremental updates are added at the end of the original file (you may rewind or squash all revisions into a single one).

Project status

WORK IN PROGRESS! This is BETA quality software. The API may change anytime. Next on TO-DO list:

Cut & append pages
Lossless compression
More filters
Improve text extraction
Augment text extraction with layout detection

Installation

You can install from PyPI:

pip install pdfsyntax

CLI overview

Please refer to the CLI README for details.

The general form of the CLI usage is:

pdfsyntax COMMAND FILE

Or this longer form if you installed from source:

python3 -m pdfsyntax COMMAND FILE

You can get quick insights on a PDF file with these commands:

overview outputs text data about the structure and the metadata.
disasm outputs a dump of the file structure on the terminal.
text outputs extracted text spatially, as if it was a kind of scan.
fonts outputs list of fonts used.
browse outputs static html data that lets you browse the internal structure of the PDF file: the PDF source is pretty-printed and augmented with hyperlinks.

API overview

Please refer to the API README for details.

PDFSyntax is mostly made of simple functions. Example:

>>> from pdfsyntax import readfile, metadata
>>> doc = readfile("samples/simple_text_string.pdf")
>>> metadata(doc) #returns a Python dict whose keys are 'Title', 'Author', etc...

The Doc object is probably the only dedicated class you will need to handle. It is a black box that stores all the internal states of a document:

content that is cached/memoized from an original file,
modifications that add/modifiy/delete content and that are tracked as incremental updates.

>>> doc
<PDF Doc in revision 1 with 0 modified object(s)>

This object exposes as a method the same metadata function, therefore you can get the same result with:

>>> doc.metadata() #returns a Python dict whose keys are 'Title', 'Author', etc...

Low-level functions like get_object or update_object allow you to directly access and manipulate the inner objects of the document structure. You may also use higher-level functions like rotate:

>>> from pdfsyntax import rotate, writefile
>>> doc180 = rotate(doc, 180) #rotate pages by 180°

The original object is unchanged and a new object is created with an incremental update (revision 2) that encloses the ongoing orientation modification:

>>> doc180
<PDF Doc in revision 1 with 1 modified object(s)>

You then can write the modified PDF to disk. Note that the resulting file contains a new section appended to the original content. You may cut this section to revert the change.

>>> writefile(doc180, "rotated_doc.pdf")

Open-Source, not Open-Contribution yet

PDFSyntax is MIT licensed but is currently closed to contributions.

Personal note: this is a pet projet of mine and my time is limited. First I need to focus on my roadmap (new features and refactoring) and then I will happily accept contributions when everything is a little more stabilised.

Project details

These details have been verified by PyPI

Project links

GitHub Statistics

Maintainers

desgeeko

These details have not been verified by PyPI

Development Status
- 4 - Beta
Intended Audience
- Developers
License
- OSI Approved :: MIT License
Operating System
- OS Independent
Programming Language
- Python :: 3
Topic
- Software Development :: Libraries
- Utilities

Release history Release notifications | RSS feed

This version

0.1.6

Jun 1, 2025

0.1.5

May 11, 2025

0.1.4

Feb 10, 2025

0.1.3

Jan 19, 2025

0.1.2

Sep 7, 2024

0.1.1

May 11, 2024

0.1.0

Apr 14, 2024

0.0.8

Jan 21, 2024

0.0.7

Sep 2, 2023

0.0.6

Apr 29, 2023

0.0.5

Nov 18, 2022

0.0.4

Sep 24, 2022

0.0.3

Sep 17, 2022

0.0.2

Jul 28, 2021

0.0.1

Jul 17, 2021

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

pdfsyntax-0.1.6.tar.gz (45.2 kB view details)

Uploaded Jun 1, 2025 Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

The dropdown lists show the available interpreters, ABIs, and platforms. Enable javascript to be able to filter the list of wheel files.

pdfsyntax-0.1.6-py3-none-any.whl (45.0 kB view details)

Uploaded Jun 1, 2025 Python 3

File details

Details for the file pdfsyntax-0.1.6.tar.gz.

File metadata

Download URL: pdfsyntax-0.1.6.tar.gz
Upload date: Jun 1, 2025
Size: 45.2 kB
Tags: Source
Uploaded using Trusted Publishing? Yes
Uploaded via: twine/6.1.0 CPython/3.12.9

File hashes

Hashes for pdfsyntax-0.1.6.tar.gz
Algorithm	Hash digest
SHA256	`b4d8823a83e77c4617bbcc96caabd1157e1ab8ae73c4af085c4397fb9a29f320`
MD5	`15c648019bc69db5d9c117d7136686e6`
BLAKE2b-256	`675eab7a10970258b28f569de8981b4ca511936ece713b6b3f495d51e34c76ef`

See more details on using hashes here.

Provenance

The following attestation bundles were made for pdfsyntax-0.1.6.tar.gz:

Publisher: python-publish.yml on desgeeko/pdfsyntax

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Statement:
- Statement type: https://in-toto.io/Statement/v1
- Predicate type: https://docs.pypi.org/attestations/publish/v1
- Subject name: pdfsyntax-0.1.6.tar.gz
- Subject digest: b4d8823a83e77c4617bbcc96caabd1157e1ab8ae73c4af085c4397fb9a29f320
- Sigstore transparency entry: 226926469
- Sigstore integration time: Jun 1, 2025
Source repository:
- Permalink: desgeeko/pdfsyntax@d3e395508f18ecc6d4fa4037e98b961390421eeb
- Branch / Tag: refs/tags/0.1.6
- Owner: https://github.com/desgeeko
- Access: public
Publication detail:
- Token Issuer: https://token.actions.githubusercontent.com
- Runner Environment: github-hosted
- Publication workflow: python-publish.yml@d3e395508f18ecc6d4fa4037e98b961390421eeb
- Trigger Event: release

File details

Details for the file pdfsyntax-0.1.6-py3-none-any.whl.

File metadata

Download URL: pdfsyntax-0.1.6-py3-none-any.whl
Upload date: Jun 1, 2025
Size: 45.0 kB
Tags: Python 3
Uploaded using Trusted Publishing? Yes
Uploaded via: twine/6.1.0 CPython/3.12.9

File hashes

Hashes for pdfsyntax-0.1.6-py3-none-any.whl
Algorithm	Hash digest
SHA256	`2f99cccd0cf3ace8400cd6046743fc0db83ebcd286bc59cd69c81d95d8944de2`
MD5	`fce3bffff068befcdfbb11dd1233bcfb`
BLAKE2b-256	`c5e46291d8c552bb49c5473f346651b596478a20cc5126f748172cfcf1bdae89`

See more details on using hashes here.

Provenance

The following attestation bundles were made for pdfsyntax-0.1.6-py3-none-any.whl:

Publisher: python-publish.yml on desgeeko/pdfsyntax

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Statement:
- Statement type: https://in-toto.io/Statement/v1
- Predicate type: https://docs.pypi.org/attestations/publish/v1
- Subject name: pdfsyntax-0.1.6-py3-none-any.whl
- Subject digest: 2f99cccd0cf3ace8400cd6046743fc0db83ebcd286bc59cd69c81d95d8944de2
- Sigstore transparency entry: 226926470
- Sigstore integration time: Jun 1, 2025
Source repository:
- Permalink: desgeeko/pdfsyntax@d3e395508f18ecc6d4fa4037e98b961390421eeb
- Branch / Tag: refs/tags/0.1.6
- Owner: https://github.com/desgeeko
- Access: public
Publication detail:
- Token Issuer: https://token.actions.githubusercontent.com
- Runner Environment: github-hosted
- Publication workflow: python-publish.yml@d3e395508f18ecc6d4fa4037e98b961390421eeb
- Trigger Event: release

pdfsyntax 0.1.6

Navigation

Verified details

Project links

GitHub Statistics

Maintainers

Unverified details

Meta

Classifiers

Project description

PDFSyntax

Introduction

Project status

Installation

CLI overview

API overview

Open-Source, not Open-Contribution yet

Project details

Verified details

Project links

GitHub Statistics

Maintainers

Unverified details

Meta

Classifiers

Release history Release notifications | RSS feed

Download files

Source Distribution

Built Distribution

File details

File metadata

File hashes

Provenance

File details

File metadata

File hashes

Provenance