pypdf-lib
A (maybe) Better PDF Parsing for Python focused on textual extraction. WIP.
This library is a Python Wrapper built around PdfAct, which is built using Java.
Pre-requisites
Java
# Linux
apt-get update && apt-get install -y default-jre # openjdk-8-jre-headless / openjdk-11-jdk / openjdk-11-jre-headless
# Mac
brew install java
# Windows
# idk
Installation
!pip install --upgrade git+https://github.com/trisongz/pypdf-lib.git
Usage
from pypdf import PyPDF
from fileio import File
base_dir = '/content/output'
File.mkdirs(base_dir)
# Using a remap function expects the file extension to be mapped properly - i.e. if 'txt' is selected, .txt file extension should be returned.
def remap_fnames(fname):
fname = File.base(fname).replace('- ', '').replace(' ', '_').strip().replace('.pdf', '.json')
return File.join(base_dir, fname)
converter = PyPDF(input_dir='/content/inputs', output_dir='/content/output', units=['paragraphs', 'blocks'], visualize=True)
# remap_funct is optional.
for res in converter.extract(remap_funct=remap_fnames):
print(res)
# > /content/output/your_json_file_1.json
converter.extracted
'''
{'/content/inputs/input_1.pdf': '/content/output/your_json_file_1.json',
'/content/inputs/input_2.pdf': '/content/output/your_json_file_2.json',
'/content/inputs/input_3.pdf': '/content/output/your_json_file_3.json',
'params': {'exclude_roles': None,
'format': 'json',
'include_roles': ['title',
'body',
'appendix',
'keywords',
'heading',
'general_terms',
'toc',
'caption',
'table',
'other',
'categories',
'keywords',
'page_header'],
'units': ['paragraphs', 'blocks'],
'visualize': True,
'with_control_characters': False}}
'''
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
pypdf-lib-0.0.1.tar.gz
(4.5 kB
view details)
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file pypdf-lib-0.0.1.tar.gz.
File metadata
- Download URL: pypdf-lib-0.0.1.tar.gz
- Upload date:
- Size: 4.5 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/3.4.1 importlib_metadata/4.0.1 pkginfo/1.7.0 requests/2.25.1 requests-toolbelt/0.9.1 tqdm/4.46.0 CPython/3.7.10
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
c677fc257814b3e633e3949c9bf83cd525fee24a67e15ef974098ef2b9d19046
|
|
| MD5 |
476daae57d0da97ee9701bf2cdc47deb
|
|
| BLAKE2b-256 |
c0c349e783d6c6c0574719003ebec6fc93dfbd6810d6b0b690d31359351ef8e6
|
File details
Details for the file pypdf_lib-0.0.1-py3-none-any.whl.
File metadata
- Download URL: pypdf_lib-0.0.1-py3-none-any.whl
- Upload date:
- Size: 6.0 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/3.4.1 importlib_metadata/4.0.1 pkginfo/1.7.0 requests/2.25.1 requests-toolbelt/0.9.1 tqdm/4.46.0 CPython/3.7.10
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
45164ef666eab3e73f1e6b72f572d24afe54478bad8e0138ea07d498bd670d88
|
|
| MD5 |
1b12d58dde1c4bda997c3da31865991f
|
|
| BLAKE2b-256 |
696e7ebbb9b51914fa2410b19f13a4715d5df3042ada17c6034f2bc4904beb40
|