Skip to main content

## Tools for processing pdf files

This is a light-weighted library for processing pdf files in python. One of the use-cases might be the extraction of pdf-annotations for ML / NLP.

Support for

  • obtaining textual and vizual content of pdf files

  • locating positions of words

  • fetching pdf annotations

  • adding a digital layer to image-pdfs

  • re-creating a clean pdf file with annotations removed

## Dependencies

Main tools for reading pdf files are the PyPDF2 library. Non-python dependencies are

To install Poppler, see the guide in the [pdf2image readme](https://pypi.org/project/pdf2image/).

## How to

Some examples of usage are shown in the [notebook](./notebook/Demo.ipynb).

## Todo

Release files for pdf-utils 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for pdf-utils 0.1.1
File Size Uploaded
pdf-utils-0.1.1.tar.gz 16.7 kB Details

Release files / pdf-utils-0.1.1.tar.gz

Download URL pdf-utils-0.1.1.tar.gz
Size 16.7 kB
Tags Source
SHA-256 checksum
How to use checksums
208bf612970ae01ab81df0637539c0d52a1f9a9f15759ee6deef3402e3924eb5
BLAKE2b-256 checksum
How to use checksums
f7e618173ef2985b5ae6d707ad050876d35436a7e1340077fe8e6bfd680c96e3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/3.2.0 pkginfo/1.5.0.1 requests/2.24.0 setuptools/49.3.1 requests-toolbelt/0.9.1 tqdm/4.48.2 CPython/3.8.5

Release history Release notifications | RSS feed

This release

0.1.1 This release

1 release file

0.0.0

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page