Skip to main content

PaperFlow

PaperFlow is an open-source post-processing layer for PDF-to-Markdown workflows.

It does not try to be the parser. It takes raw markdown from a parser you choose and upgrades it into structured, knowledge-ready markdown with:

  • normalized LaTeX delimiters
  • linked citations as standard footnotes
  • figure and table jump links
  • YAML frontmatter
  • cleaned repeated headers and footers

PyPI package: https://pypi.org/project/paperflow-postprocess/
GitHub: https://github.com/TylerMorrison21/paperflow
Project page: https://www.paperflowing.com

Install

pip install paperflow-postprocess

Quick Start

from paperflow_postprocess import enhance

raw_markdown = """
Text with citation [1].

## References

[1] Example Author. Example Paper.
"""

markdown = enhance(
    raw_markdown=raw_markdown,
    images={},
    metadata={
        "title": "Example Paper",
        "authors": ["Example Author"],
        "source": "https://example.com/paper",
        "date": "2026-03-11",
    },
)

print(markdown)

PDF Parsers - bring your own

PaperFlow is a post-processing layer. It enhances raw Markdown from any upstream PDF parser. You need to choose a parser:

Option 1: Datalab Marker API (recommended, easiest)

  • Sign up at datalab.to - $25/month free credits
  • Cloud API, no GPU needed
  • Set DATALAB_API_KEY in your .env
  • PaperFlow's built-in api/services/marker.py calls this automatically

Option 2: Marker (self-hosted, free)

  • github.com/datalab-to/marker
  • Run locally with GPU (CUDA) or CPU
  • Free for orgs under $5M revenue
  • You'll need to modify api/services/marker.py to call your local endpoint

Option 3: MinerU (self-hosted, free)

  • github.com/opendatalab/MinerU
  • Strong on Chinese docs, scientific papers, complex tables
  • Outputs Markdown + LaTeX - compatible with PaperFlow's postprocess
  • Needs GPU (~6GB VRAM minimum)
  • Replace marker.py with a MinerU client

Option 4: Docling, PyMuPDF4LLM, or any other parser

  • Any tool that outputs Markdown will work
  • Feed the raw Markdown into paperflow_postprocess.enhance() or save it and run it through the API pipeline

Using the pip package with any parser

from paperflow_postprocess import enhance

# Get raw markdown from ANY parser
raw_md = my_parser.convert("paper.pdf")

# Enhance with PaperFlow
result = enhance(raw_md, images={}, metadata={"title": "My Paper"})

# result has footnotes, fixed LaTeX, figure links, YAML frontmatter

The whole point of PaperFlow is that parsing is commoditized. Marker, MinerU, LlamaParse, Docling, and PyMuPDF4LLM can all produce decent raw output. The post-processing layer is where the value is, and that is what PaperFlow does.

Visual Comparison

Source PDF Generic converter output PaperFlow output
Source PDF Generic converter output PaperFlow output

Additional example:

Calligraphy comparison

What enhance() does

enhance() upgrades parser output into structured markdown with:

  • standard footnotes [^N] from inline citations like [1], [1, 2], or [1-3]
  • normalized LaTeX delimiters using $...$ and $$...$$
  • figure links like [[#^fig-3|Fig. 3]]
  • table links like [[#^tab-2|Table 2]]
  • YAML frontmatter with title, authors, source, date, and extraction hash
  • cleaned repeated headers, footers, and page number lines

API

  • enhance(raw_markdown, images=None, metadata=None)
  • postprocess(raw_markdown, images=None, metadata=None) for backward compatibility
  • fix_latex_delimiters(md)
  • clean_headers_footers(md)
  • convert_to_footnotes(md)
  • linkify_figures(md)
  • linkify_tables(md)
  • inject_frontmatter(md, metadata)

Release files for paperflow-postprocess 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for paperflow-postprocess 0.1.1
File Size Uploaded
paperflow_postprocess-0.1.1.tar.gz 7.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for paperflow-postprocess 0.1.1
File Interpreter ABI Platform
paperflow_postprocess-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 15.5 kB

Release files / paperflow_postprocess-0.1.1.tar.gz

Download URL paperflow_postprocess-0.1.1.tar.gz
Size 7.9 kB
Tags Source
SHA-256 checksum
How to use checksums
3a7ee567cb635b53540dffe600937c1fefcc237d01f2d729b04f5ce7aa59fd47
BLAKE2b-256 checksum
How to use checksums
70551836c9191ea42ca9e2df21986d42ffa99c8940dad1b3ac2b03b9863ea058
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.10.11

Release files / paperflow_postprocess-0.1.1-py3-none-any.whl

Download URL paperflow_postprocess-0.1.1-py3-none-any.whl
Size 7.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
e8db982540efc0ceaaa8cf612048f9238859d1e8affd3c9e44a6d9f9bf274361
BLAKE2b-256 checksum
How to use checksums
aaf33c1405e7234d2530c51efaf81422d7e93c96c4f6eacf1da1ed21e440da8a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.10.11

Release history Release notifications | RSS feed

This release

0.1.1 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page