Skip to main content

Scale2Pdf

A library made at LAION to scale the parsing of PDFs on CPUs. We tested our pipeline on 44-page pdf and on cheap 2 thread CPU with 12GB ram, it took us 3 mins 22 seconds to parse the pdf, save its content both structure, bulk and its images. We provide following results through our framework

  1. Table extraction
  2. Equation extraction
  3. Image Captions
  4. Page extraction
  5. Keyword extraction
  6. Section extraction
  7. Authors extraction
  8. Bibliography extraction
  9. Paragraph extraction
  10. Image extraction
  11. Abstract extraction

Features

  1. added support for ray for scalability

Installation

pip install scale2pdf ray

then install

sudo apt install poppler-utils

from scale2pdf import scalablepdf 
from scale2pdf import extractimages

pdf_path = "/content/2408.06257v3.pdf"
scalablepdf(pdf_path, extract_images=True) # folder is automatically created and results are saved
# if you want to process a folder of pdfs with ray then
scalable_ray("example_folder", extract_images=True, num_cpus=4)
extractimages("2408.06257v3.pdf", "/path/to/output/folder")
Ray caution:

If you don't specify the CPU numbers then 4 CPU cores will be used at a time. You can increase it to the highest number of CPU cores available.

Speedup depends entirely on the CPU and resources available. I had used on a cheap CPU and it was bad since I had only two threads akin 2 CORE. (although threads here means core not threads themselves like in Computer Hardware)

CRAP CPU (NO GPU): 3 min 22 seconds to finish parsing and saving it to JSON.

A Sleeping AI framework made for friends at LAION AI.

Release files for scale2pdf 0.0.7

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for scale2pdf 0.0.7
File Size Uploaded
scale2pdf-0.0.7.tar.gz 3.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for scale2pdf 0.0.7
File Interpreter ABI Platform
scale2pdf-0.0.7-py3-none-any.whl Python 3 none any Details

Total release size: 9.3 kB

Release files / scale2pdf-0.0.7.tar.gz

Download URL scale2pdf-0.0.7.tar.gz
Size 3.7 kB
Tags Source
SHA-256 checksum
How to use checksums
2556a1fc89459fd6a77e644a78b742a84cf0d8a261b0b8a7bd0317564097b809
BLAKE2b-256 checksum
How to use checksums
6872f5d5fe83e5f2134cb2b65b217dae978ba2deffe1ae25901a1b10632439df
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/5.1.1 CPython/3.10.8

Release files / scale2pdf-0.0.7-py3-none-any.whl

Download URL scale2pdf-0.0.7-py3-none-any.whl
Size 5.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
ca0d73ef1f941c12a8eba4311b9755ec86f41a1799b2602e5cc269839a660916
BLAKE2b-256 checksum
How to use checksums
2937d4df5966bfd07d31b85e8ff9115410ee2a434499b31e06a5f1f5e13fa9f7
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/5.1.1 CPython/3.10.8

Release history Release notifications | RSS feed

This release

0.0.7 This release

2 release files

0.0.6

2 release files

0.0.5

2 release files

0.0.4

2 release files

0.0.3

2 release files

0.0.2

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page