EDS-PDF
EDS-PDF provides a modular framework to extract text information from PDF documents.
You can use it out-of-the-box, or extend it to fit your specific use case. We provide a pipeline system and various utilities for visualizing and processing PDFs, as well as multiple components to build complex models:complex models:
- 📄 Extractors to parse PDFs (based on pdfminer, mupdf or poppler)
- 🎯 Classifiers to perform text box classification, in order to segment PDFs
- 🧩 Aggregators to produce an aggregated output from the detected text boxes
- 🧠 Trainable layers to incorporate machine learning in your pipeline (e.g., embedding building blocks or a trainable classifier)
Visit the :book: documentation for more information!
Getting started
Installation
Install the library with pip:
pip install edspdf
Extracting text
Let's build a simple PDF extractor that uses a rule-based classifier. There are two ways to do this, either by using the configuration system or by using the pipeline API.
Create a configuration file:
config.cfg
[pipeline]
pipeline = ["extractor", "classifier", "aggregator"]
[components.extractor]
@factory = "pdfminer-extractor"
[components.classifier]
@factory = "mask-classifier"
x0 = 0.2
x1 = 0.9
y0 = 0.3
y1 = 0.6
threshold = 0.1
[components.aggregator]
@factory = "simple-aggregator"
and load it from Python:
import edspdf
from pathlib import Path
model = edspdf.load("config.cfg") # (1)
Or create a pipeline directly from Python:
from edspdf import Pipeline
model = Pipeline()
model.add_pipe("pdfminer-extractor")
model.add_pipe(
"mask-classifier",
config=dict(
x0=0.2,
x1=0.9,
y0=0.3,
y1=0.6,
threshold=0.1,
),
)
model.add_pipe("simple-aggregator")
This pipeline can then be applied (for instance with this PDF):
# Get a PDF
pdf = Path("/Users/perceval/Development/edspdf/tests/resources/letter.pdf").read_bytes()
pdf = model(pdf)
body = pdf.aggregated_texts["body"]
text, style = body.text, body.properties
See the rule-based recipe for a step-by-step explanation of what is happening.
Citation
If you use EDS-PDF, please cite us as below.
@software{edspdf,
author = {Dura, Basile and Wajsburt, Perceval and Calliger, Alice and Gérardin, Christel and Bey, Romain},
doi = {10.5281/zenodo.6902977},
license = {BSD-3-Clause},
title = {{EDS-PDF: Smart text extraction from PDF documents}},
url = {https://github.com/aphp/edspdf}
}
Acknowledgement
We would like to thank Assistance Publique – Hôpitaux de Paris and AP-HP Foundation for funding this project.
Metadata
Release files for edspdf 0.10.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| edspdf-0.10.0.tar.gz | 1.7 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| edspdf-0.10.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 1.8 MB
Release files / edspdf-0.10.0.tar.gz
| Download URL | edspdf-0.10.0.tar.gz |
|---|---|
| Size | 1.7 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
67708cf34269527f69d881f1338ffb1935eb0044cf2f1af2806ad22ca67b62a1
|
|
BLAKE2b-256 checksum How to use checksums |
4e3b29afc5e765f42edcddc1c3e377cb6c6c4b7b18dd4477ce5fd82acf4a4733
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.12.8
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Feb 12, 2025.
Transparency logRelease files / edspdf-0.10.0-py3-none-any.whl
| Download URL | edspdf-0.10.0-py3-none-any.whl |
|---|---|
| Size | 100.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
5225655888d9fb5af894757b775ce3b9f4e06bc92f5ebcbee4df0316b48b44a6
|
|
BLAKE2b-256 checksum How to use checksums |
017abcd85613c55c9586954aa40ce4ca51f8100292fbe07d44e49f3b1d64cd30
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.12.8
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Feb 12, 2025.
Transparency log