Scale2Pdf
A library made at LAION to scale the parsing of PDFs on CPUs. We tested our pipeline on 44-page pdf and on cheap 2 thread CPU with 12GB ram, it took us 3 mins 22 seconds to parse the pdf, save its content both structure, bulk and its images. We provide following results through our framework
- Table extraction
- Equation extraction
- Image Captions
- Page extraction
- Keyword extraction
- Section extraction
- Authors extraction
- Bibliography extraction
- Paragraph extraction
- Image extraction
- Abstract extraction
Features
- added support for ray for scalability
Installation
pip install scale2pdf ray
then install
sudo apt install poppler-utils
from scale2pdf import scalablepdf
from scale2pdf import extractimages
pdf_path = "/content/2408.06257v3.pdf"
scalablepdf(pdf_path, extract_images=True) # folder is automatically created and results are saved
# if you want to process a folder of pdfs with ray then
scalable_ray("example_folder", extract_images=True, num_cpus=4)
extractimages("2408.06257v3.pdf", "/path/to/output/folder")
Ray caution:
If you don't specify the CPU numbers then 4 CPU cores will be used at a time. You can increase it to the highest number of CPU cores available.
Speedup depends entirely on the CPU and resources available. I had used on a cheap CPU and it was bad since I had only two threads akin 2 CORE. (although threads here means core not threads themselves like in Computer Hardware)
CRAP CPU (NO GPU): 3 min 22 seconds to finish parsing and saving it to JSON.
A Sleeping AI framework made for friends at LAION AI.
Release files for scale2pdf 0.0.7
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| scale2pdf-0.0.7.tar.gz | 3.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| scale2pdf-0.0.7-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 9.3 kB
Release files / scale2pdf-0.0.7.tar.gz
| Download URL | scale2pdf-0.0.7.tar.gz |
|---|---|
| Size | 3.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
2556a1fc89459fd6a77e644a78b742a84cf0d8a261b0b8a7bd0317564097b809
|
|
BLAKE2b-256 checksum How to use checksums |
6872f5d5fe83e5f2134cb2b65b217dae978ba2deffe1ae25901a1b10632439df
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/5.1.1 CPython/3.10.8
|
Release files / scale2pdf-0.0.7-py3-none-any.whl
| Download URL | scale2pdf-0.0.7-py3-none-any.whl |
|---|---|
| Size | 5.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
ca0d73ef1f941c12a8eba4311b9755ec86f41a1799b2602e5cc269839a660916
|
|
BLAKE2b-256 checksum How to use checksums |
2937d4df5966bfd07d31b85e8ff9115410ee2a434499b31e06a5f1f5e13fa9f7
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/5.1.1 CPython/3.10.8
|