Skip to main content

Papercutter

Turn your PDF collection into a dataset you can actually use.

For researchers doing systematic reviews, meta-analyses, or literature surveys who have PDFs piling up but need structured data for analysis.

Requires Python 3.10+ and an OpenAI or Anthropic API key.

Installation

pip install papercutter[full]

Optional extras:

pip install papercutter[docling]  # PDF processing only
pip install papercutter[llm]      # LLM extraction only
pip install papercutter[report]   # Report generation only

Usage

1. Ingest

Convert PDFs to Markdown and extract tables.

papercutter ingest ./pdfs/

Output: markdown/, tables/, figures/, inventory.json

2. Configure

Generate an extraction schema from paper abstracts.

papercutter configure

Output: columns.yaml

columns:
  - key: sample_size
    description: "Total observations (N)"
    type: integer
  - key: method
    description: "Estimation strategy (DiD, RDD, OLS, etc.)"
    type: string
  - key: effect
    description: "Main treatment coefficient"
    type: float

3. Extract

Extract data from all papers using LLM.

papercutter extract

Output: extractions.json

4. Report

Generate analysis outputs.

papercutter report            # matrix.csv + review.pdf
papercutter report --condensed  # appendix format

Book Pipeline

Process books or handbooks with chapter-level summaries.

papercutter book index ./book.pdf   # Detect chapters
papercutter book extract            # Extract text
papercutter book summarize          # Summarize chapters
papercutter book report             # Generate PDF

Output: output/book_summary.pdf

Output Files

File Description
inventory.json Processing status for each PDF
columns.yaml Extraction schema definition
extractions.json Extracted data per paper
matrix.csv Flattened dataset for R/Stata/Python
review.pdf LaTeX report with structured summaries

License

MIT

Metadata

Release files for papercutter 3.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for papercutter 3.1.0
File Size Uploaded
papercutter-3.1.0.tar.gz 258.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for papercutter 3.1.0
File Interpreter ABI Platform
papercutter-3.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 297.0 kB

Release files / papercutter-3.1.0.tar.gz

Download URL papercutter-3.1.0.tar.gz
Size 258.8 kB
Tags Source
SHA-256 checksum
How to use checksums
2ee877fa9ba6143ed6ff81bd774b8998d0c10127c7527140e631a302ef1f994a
BLAKE2b-256 checksum
How to use checksums
bed4144d65872f0d876bef68fc6db21133f8c7ea499476b3cd51e13fa9c9fc65
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.11.6

Release files / papercutter-3.1.0-py3-none-any.whl

Download URL papercutter-3.1.0-py3-none-any.whl
Size 38.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
bdf0dc1d8fc17a13899d82290c674331a3e33ce2bce79a6990deca32646a97d8
BLAKE2b-256 checksum
How to use checksums
152853910dcf65426a9b2e9606d32a2fcc3bfac132cabb214c7e0df0ffc63390
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.11.6

Release history Release notifications | RSS feed

This release

3.1.0 This release

2 release files

3.0.2

2 release files

3.0.1

2 release files

3.0.0

2 release files

2.0.2

2 release files

2.0.1

2 release files

2.0.0

2 release files

1.2.0

2 release files

1.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page