GrobidArticleExtractor is a Python package designed to extract and organize content from scientific papers in PDF format.
Project description
GrobidArticleExtractor
This Python tool extracts content from PDF files using GROBID and organizes it by sections. It provides a structured way to extract both metadata and content from academic papers and other structured documents.
Features
- Direct PDF processing using GROBID API
- Metadata extraction (title, authors, abstract, publication date)
- Hierarchical section organization with subsections
Prerequisites
-
Install GROBID:
# Using Docker (recommended) docker pull lfoppiano/grobid:0.7.3 docker run -t --rm -p 8070:8070 lfoppiano/grobid:0.7.3
-
Install Python dependencies:
pip install poetry
poetry install
Usage
Command Line Interface
The tool provides a user-friendly command-line interface for batch processing PDF files:
# Basic usage (processes PDFs from 'pdfs' directory)
python cli.py
# Process PDFs from a specific directory
python cli.py path/to/pdfs
# Specify custom output directory
python cli.py path/to/pdfs -o path/to/output
# Use custom GROBID server and disable content preview
python cli.py path/to/pdfs --grobid-url http://custom:8070 --no-preview
Available options:
$ python cli.py --help
Usage: cli.py [OPTIONS] [INPUT_FOLDER]
Process PDF files from INPUT_FOLDER and extract their content using GROBID.
The extracted content is saved as JSON files in the output directory.
Each JSON file is named after its source PDF file.
Options:
-o, --output-dir PATH Directory to save extracted JSON files (default: output)
-g, --grobid-url TEXT GROBID service URL (default: http://localhost:8070)
--preview / --no-preview
Show preview of extracted content (default: True)
--help Show this message and exit.
Example:
python cli.py path/to/pdfs -o path/to/output
Python API Usage
You can also use the tool programmatically in your Python code:
from GrobidArticleExtractor import GrobidArticleExtractor
# Initialize extractor (default GROBID URL: http://localhost:8070)
extractor = GrobidArticleExtractor()
# Process a PDF file
xml_content = extractor.process_pdf("path/to/your/paper.pdf")
if xml_content:
# Extract and organize content
result = extractor.extract_content(xml_content)
# Access metadata
print(result['metadata'])
# Access sections
for section in result['sections']:
print(section['heading'])
if 'content' in section:
print(section['content'])
Custom GROBID server:
extractor = GrobidArticleExtractor(grobid_url="http://your-grobid-server:8070")
Output Structure
The extracted content is organized as follows:
{
'metadata': {
'title': 'Paper Title',
'authors': ['Author 1', 'Author 2'],
'abstract': 'Paper abstract...',
'publication_date': '2023'
},
'sections': [
{
'heading': 'Introduction',
'content': ['Paragraph 1...', 'Paragraph 2...'],
'subsections': [
{
'heading': 'Background',
'content': ['Subsection content...']
}
]
}
# More sections...
]
}
Project Structure
The project is organized into two main files:
app.py- Contains the coreGrobidArticleExtractorclass with all the PDF processing and content extraction functionalitycli.py- Contains the command-line interface implementation using Click
Error Handling
The tool includes comprehensive error handling for common scenarios:
- PDF file not found
- GROBID service unavailable
- XML parsing errors
- Invalid content structure
All errors are logged with appropriate messages using Python's logging module.
Contributing
Feel free to submit issues and enhancement requests!
License
MIT License
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file grobidarticleextractor-0.2.0.tar.gz.
File metadata
- Download URL: grobidarticleextractor-0.2.0.tar.gz
- Upload date:
- Size: 8.2 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: poetry/1.8.3 CPython/3.11.7 Darwin/23.6.0
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
733a6d321f58c42d83f572ceb1bb927d8437727cb46b120e45904c4bb3e5acfb
|
|
| MD5 |
7a5a4088815d907f57a8bb39dd5aa39d
|
|
| BLAKE2b-256 |
f89074fd45afeefacfd517de89cfb233eda5e0d11cfcca031920ba0c5d9add95
|
File details
Details for the file grobidarticleextractor-0.2.0-py3-none-any.whl.
File metadata
- Download URL: grobidarticleextractor-0.2.0-py3-none-any.whl
- Upload date:
- Size: 9.5 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: poetry/1.8.3 CPython/3.11.7 Darwin/23.6.0
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
96f3ad4575295c87e40bb59a36f45b26db48e9c7b4c42800a13ad3b8a89315eb
|
|
| MD5 |
856f8bb36bfa983f92a56060bcf5afdf
|
|
| BLAKE2b-256 |
ccebedf13ab04f68c5645b3ec30b04474a155fe981105e9d15f2a1668a586307
|