Papercutter
Turn your PDF collection into a dataset you can actually use.
For researchers doing systematic reviews, meta-analyses, or literature surveys who have PDFs piling up but need structured data for analysis.
Requires Python 3.10+ and an OpenAI or Anthropic API key.
Installation
pip install papercutter[full]
Optional extras:
pip install papercutter[docling] # PDF processing only
pip install papercutter[llm] # LLM extraction only
pip install papercutter[report] # Report generation only
Usage
1. Ingest
Convert PDFs to Markdown and extract tables.
papercutter ingest ./pdfs/
Output: markdown/, tables/, figures/, inventory.json
2. Configure
Generate an extraction schema from paper abstracts.
papercutter configure
Output: columns.yaml
columns:
- key: sample_size
description: "Total observations (N)"
type: integer
- key: method
description: "Estimation strategy (DiD, RDD, OLS, etc.)"
type: string
- key: effect
description: "Main treatment coefficient"
type: float
3. Extract
Extract data from all papers using LLM.
papercutter extract
Output: extractions.json
4. Report
Generate analysis outputs.
papercutter report # matrix.csv + review.pdf
papercutter report --condensed # appendix format
Book Pipeline
Process books or handbooks with chapter-level summaries.
papercutter book index ./book.pdf # Detect chapters
papercutter book extract # Extract text
papercutter book summarize # Summarize chapters
papercutter book report # Generate PDF
Output: output/book_summary.pdf
Output Files
| File | Description |
|---|---|
inventory.json |
Processing status for each PDF |
columns.yaml |
Extraction schema definition |
extractions.json |
Extracted data per paper |
matrix.csv |
Flattened dataset for R/Stata/Python |
review.pdf |
LaTeX report with structured summaries |
License
MIT
Metadata
Release files for papercutter 3.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| papercutter-3.1.0.tar.gz | 258.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| papercutter-3.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 297.0 kB
Release files / papercutter-3.1.0.tar.gz
| Download URL | papercutter-3.1.0.tar.gz |
|---|---|
| Size | 258.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
2ee877fa9ba6143ed6ff81bd774b8998d0c10127c7527140e631a302ef1f994a
|
|
BLAKE2b-256 checksum How to use checksums |
bed4144d65872f0d876bef68fc6db21133f8c7ea499476b3cd51e13fa9c9fc65
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.11.6
|
Release files / papercutter-3.1.0-py3-none-any.whl
| Download URL | papercutter-3.1.0-py3-none-any.whl |
|---|---|
| Size | 38.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
bdf0dc1d8fc17a13899d82290c674331a3e33ce2bce79a6990deca32646a97d8
|
|
BLAKE2b-256 checksum How to use checksums |
152853910dcf65426a9b2e9606d32a2fcc3bfac132cabb214c7e0df0ffc63390
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.11.6
|