Skip to main content

pdf-book2md

pdf-book2md turns PDF books and manuals into cleaned Markdown organized as chapter files. It is built for long-form, structured PDFs where you want a repeatable output tree for reading, search, indexing, or downstream processing.

Quick Start

Install from PyPI:

python -m pip install pdf-book2md

Extract one book:

pdf-book2md extract ./manual.pdf --output-dir ./out/manual/chapters

Batch-process a directory:

pdf-book2md batch ./pdfs --output-root ./out

What It Produces

When chapter headings are detected, pdf-book2md writes numbered chapter files and preserves the source text as full_document.md:

out/
  manual/
    chapters/
      00_frontmatter.md
      01_Chapter_1_Overview.md
      02_Chapter_2_Setup.md
      03_Appendix_A_Reference.md
      full_document.md

If no chapter split is detected, the command writes a single full_document.md.

Features

  • Extract one PDF into cleaned Markdown.
  • Batch-process PDFs into per-document output directories.
  • Split book/manual-style documents into deterministic chapter files.
  • Preserve frontmatter separately when introductory material exists before the first detected chapter.
  • Skip already-processed batch outputs unless --force is supplied.
  • Print JSON summaries for automation with --json.
  • Lazy-load pymupdf4llm, so tests and CLI help work without the runtime extraction dependency installed.

Usage

Single PDF

pdf-book2md extract ./manual.pdf --output-dir ./out/manual/chapters

Use --skip-full-document when you only want chapter files and do not need the combined Markdown copy:

pdf-book2md extract ./manual.pdf --output-dir ./out/manual/chapters --skip-full-document

Batch Directory

pdf-book2md batch ./pdfs --output-root ./out

Batch mode writes each input PDF under OUTPUT_ROOT/<pdf-stem>/chapters/. Existing outputs are skipped when the target chapters/ directory already contains Markdown files.

Reprocess existing outputs:

pdf-book2md batch ./pdfs --output-root ./out --force

Select PDFs with a custom glob:

pdf-book2md batch ./pdfs --output-root ./out --pattern "*guide*.pdf"

JSON Output

pdf-book2md batch ./pdfs --output-root ./out --json

The JSON summary includes absolute input and output paths, processed documents, and skipped documents:

{
  "input_dir": "/path/to/pdfs",
  "output_root": "/path/to/out",
  "processed": [],
  "skipped": ["/path/to/pdfs/manual.pdf"]
}

Common runtime failures are printed to stderr as error: ... and exit with status code 1. Argument parsing errors exit with status code 2.

Library API

from pathlib import Path

from pdf_extract_cli import extract_pdf

result = extract_pdf(Path("manual.pdf"), Path("out/manual/chapters"))
print(result.chapter_count)

Development

Install from source:

python3 -m pip install -e . --no-deps

Run tests:

python3 -m unittest discover -s tests -v

Smoke-test the CLI:

pdf-book2md --help

Releases are managed with Python Semantic Release. Conventional commits on main determine whether a release is created; release builds publish to PyPI using trusted publishing for .github/workflows/release.yml on the main branch.

Metadata

Release files for pdf-book2md 1.0.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for pdf-book2md 1.0.0
File Size Uploaded
pdf_book2md-1.0.0.tar.gz 9.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for pdf-book2md 1.0.0
File Interpreter ABI Platform
pdf_book2md-1.0.0-py3-none-any.whl Python 3 none any Details

Total release size: 19.1 kB

Release files / pdf_book2md-1.0.0.tar.gz

Download URL pdf_book2md-1.0.0.tar.gz
Size 9.5 kB
Tags Source
SHA-256 checksum
How to use checksums
a6d0b00cf3ed3531d289eddd255c1c5bca2e2d1eb5c9be2871ad64f33dca664b
BLAKE2b-256 checksum
How to use checksums
b3d88436a79ad3d6d3baf9db800f7be799d2c00a5cc97dd606c4a2a65ed8d940
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on May 26, 2026.

Transparency log

Release files / pdf_book2md-1.0.0-py3-none-any.whl

Download URL pdf_book2md-1.0.0-py3-none-any.whl
Size 9.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
c7fa7dd96623f501b8b1bacbbc9300a7d753a3be9b08aac9fa964ddb023efe5d
BLAKE2b-256 checksum
How to use checksums
272dca6a6ffd8c9d83eded7c5699486717dc82dcdeb298c3e85782148e9c930f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on May 26, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

1.0.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page