Skip to main content

PDF Craft

ci pip install pdf-craft pypi pdf-craft python versions Ask DeepWiki license

oomol-lab%2Fpdf-craft | Trendshift

English | 中文

What is pdf-craft?

pdf-craft is a PDF-centered conversion library. It turns PDFs into Markdown or EPUB, and can translate the converted content or write a translated result back to PDF. It is especially useful for scanned documents: pages that are otherwise only readable as images become searchable, editable Markdown or EPUB.

The pipeline is designed for books and academic or technical documents, including body text, tables of contents, footnotes, tables, formulas, and images. OCR can run entirely on a compatible local GPU, or use a vendor service that supplies remote compute. Translation uses a separate text LLM.

Online Version

Want to try the workflow before installing anything? Open PDF Craft Online, the online version of the same core experience. Upload a PDF in your browser and see the main workflow in action.

PDF Craft Online Version

Installation

If you are getting started, use the standard installation:

pip install pdf-craft

This includes vendor OCR, Markdown/EPUB rendering, and PDF translation. Vendor OCR uses remote compute, so your machine does not need CUDA; you provide the service URL, model name, and API key in the OCR configuration.

Only install the local extra when you explicitly want to run OCR models on your own NVIDIA GPU. If you are unsure, use the standard installation above:

pip install "pdf-craft[local]"

Local OCR also requires a CUDA-compatible PyTorch build, model storage, and enough GPU memory. Before processing PDFs, install Poppler; translated-PDF patching also needs Ghostscript and suitable local fonts. Optional vector inline-formula rendering additionally uses Matplotlib and a local TeX installation. See the Installation Guide for the supported Python versions and complete system setup. If something goes wrong, start with the Troubleshooting Guide.

Quick Start

The following example converts a scanned PDF into a Markdown file. Replace the example OCR endpoint, model name, and API key with your own service configuration.

from pdf_craft import DeepSeekOCRVendorConfig, PDFCraft, PDFOptions

craft = PDFCraft(pdf=PDFOptions(ocr=DeepSeekOCRVendorConfig(
    base_url="https://example.com/v1",
    api_key="your-api-key",
    model="deepseek-ocr",
)))
craft.convert_pdf_to_markdown(
    "input.pdf", "output.md",
)

PDF to Markdown example

The conversion uses a temporary analysis workspace automatically and removes it when the conversion finishes or fails. Pass analysing_path to retain diagnostics, or extraction_path="book.pcex" to retain the portable intermediate extraction.

For the complete PDF conversion workflow and customization options, see the PDF Translation Guide and API Reference. For the field-level contract of the portable intermediate format, see the PDFCraftExtraction (.pcex) Format Reference.

Advanced Features

PDF → EPUB

To produce an EPUB instead of Markdown, call convert_pdf_to_epub. This complete example also shows how to set the book title and author metadata:

from pdf_craft import BookMeta, DeepSeekOCRVendorConfig, PDFCraft, PDFOptions

ocr_config = DeepSeekOCRVendorConfig(
    base_url="https://example.com/v1",
    api_key="your-api-key",
    model="deepseek-ocr",
)
craft = PDFCraft(pdf=PDFOptions(ocr=ocr_config))
craft.convert_pdf_to_epub(
    "input.pdf", "output.epub",
    book_meta=BookMeta(title="Book title", authors=["Author"]),
)

PDF to EPUB example

If book_meta is omitted, pdf-craft uses metadata stored in the PDFCraftExtraction manifest when the PDF was extracted.

PDF → translated Markdown or EPUB

To translate while converting, pass one chapter translator to either conversion method. The translator sends chapter text to your text LLM and returns the translated chapter.

craft.convert_pdf_to_markdown(
    "input.pdf", "translated.md", translator=translator,
)
craft.convert_pdf_to_epub(
    "input.pdf", "translated.epub", translator=translator,
)

PDF → translated PDF

Use the PDF translation workflow when you want to keep the original PDF layout. It extracts the page content, translates it, and writes the result back into the matching source pages. OCR and translation use separate configurations.

from pdf_craft import DeepSeekOCRVendorConfig, PDFCraft, PDFOptions

craft = PDFCraft(pdf=PDFOptions(ocr=DeepSeekOCRVendorConfig(
    base_url="https://example.com/v1",
    api_key="your-ocr-api-key",
    model="deepseek-ocr",
)))

# Placeholder only: replace this with your text LLM call.
def translator(text: str) -> str:
    return text

extraction = craft.extract_pdf("input.pdf", "work/book.pcex")
craft.translate_pdf("input.pdf", extraction, "translated.pdf", translator)

EPUB → translated EPUB

If you already have an EPUB, translate it directly by providing the target language and a text LLM:

from pdf_craft import LLM, PDFCraft, SubmitKind

llm = LLM(
    key="your-api-key",
    url="https://api.openai.com/v1",
    model="gpt-4.1-mini",
    token_encoding="o200k_base",
)

PDFCraft().translate_epub(
    "input.epub", "translated.epub",
    target_language="zh", submit=SubmitKind.REPLACE, llm=llm,
)

REPLACE creates a target-language-only edition. Use APPEND_BLOCK to keep the original and append the translation as a separate block, or APPEND_TEXT to place the translation directly after the original text. See the EPUB translation guide for prompts, retries, concurrency, caching, progress callbacks, and failure handling.

OCR Backends and Model Cache

OCR turns page images into text. pdf-craft offers six backends; choose the runtime location first, then choose the model family:

  • No CUDA or minimal local setup: choose vendor OCR. Pages are sent to a remote service and processed with its compute resources, so you need network access, a service URL, and credentials.
  • A compatible NVIDIA GPU and local execution: choose local OCR. Models are cached locally and run on your GPU, which keeps processing on your machine but requires CUDA, VRAM, and model files.

DeepSeek OCR and DeepSeek OCR 2 are from DeepSeek; Unlimited OCR is from Baidu. Each model family has local and vendor configurations:

Backend Owner Runs on Choose it when You need
DeepSeekOCRLocalConfig DeepSeek Local GPU You want local DeepSeek OCR CUDA, VRAM, model cache
DeepSeekOCR2LocalConfig DeepSeek Local GPU You want local DeepSeek OCR 2 CUDA, VRAM, model cache; base is the verified preset
UnlimitedOCRLocalConfig Baidu Local GPU You want local Unlimited OCR CUDA, VRAM, model cache
DeepSeekOCRVendorConfig DeepSeek Remote service You do not have CUDA or prefer remote DeepSeek OCR URL, model, API key, network
DeepSeekOCR2VendorConfig DeepSeek Remote service You prefer remote DeepSeek OCR 2 URL, model, API key, network
UnlimitedOCRVendorConfig Baidu Remote service You prefer remote Unlimited OCR URL, credentials, network

If you simply want to get the workflow running, start with the vendor you already have credentials for. Choose local OCR when you specifically want local execution. The library accepts these configuration objects through PDFOptions(ocr=...) and does not read environment variables. See the OCR Backend Guide for detailed configuration examples.

Unlimited OCR local supports base and gundam. DeepSeek OCR 2 local is verified with base; an explicit tiny selection fails early with a clear message.

Model Cache and Common Parameters

Local OCR models are downloaded from Hugging Face by default. You can pre-download one into a chosen cache directory and then run with local_only=True:

from pdf_craft import DeepSeekOCRLocalConfig, predownload_models

predownload_models(
    ocr=DeepSeekOCRLocalConfig(models_cache_path="models"),
    revision=None,
)

ocr_size supports tiny, small, base, large, and gundam, although presets vary by backend. Markdown defaults to toc_assumed=False; EPUB defaults to toc_assumed=True. Complex chapter hierarchies can use an optional toc_llm.

Related Projects

  • Wiki Graph: turn a converted EPUB or Markdown book into structured summaries, chapter topology, and a knowledge graph.

License

This project is licensed under the MIT License. See the LICENSE file for details.

Since v1.0.0, pdf-craft has used DeepSeek OCR under the MIT license and removed the previous AGPL-3.0 dependency. The project still receives easydict transitively through the OCR stack under the LGPLv3 license. Thanks to the community for their support and contributions.

Acknowledgments

Metadata

Release files for pdf-craft 2.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for pdf-craft 2.2.0
File Size Uploaded
pdf_craft-2.2.0.tar.gz 174.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for pdf-craft 2.2.0
File Interpreter ABI Platform
pdf_craft-2.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 399.3 kB

Release files / pdf_craft-2.2.0.tar.gz

Download URL pdf_craft-2.2.0.tar.gz
Size 174.5 kB
Tags Source
SHA-256 checksum
How to use checksums
1b04a93ac2497f206a2e27d319cc11e1d002c77da65f3e152b8476e1aafaba47
BLAKE2b-256 checksum
How to use checksums
3b5b0770c95734e7394cfe22f76e42fc95e7350e084597484dc914deb01f8ef8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 9, 2026.

Transparency log

Release files / pdf_craft-2.2.0-py3-none-any.whl

Download URL pdf_craft-2.2.0-py3-none-any.whl
Size 224.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
f3c183679124061655bf323e617a33fcfe87bbdbb2d0bd2c07a642061fbf54c3
BLAKE2b-256 checksum
How to use checksums
5a26609c05c95bac112a66ffe5280c700bb07323374efd31217ea560cd864fe6
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 9, 2026.

Transparency log

Release history Release notifications | RSS feed

2.4.1

2 release files

2.4.0

2 release files

2.3.1

2 release files

2.3.0

2 release files

2.2.1

2 release files

This release

2.2.0 This release

2 release files

2.1.0

2 release files

2.0.2

2 release files

2.0.1

2 release files

2.0.0

2 release files

1.0.14

2 release files

1.0.13

2 release files

1.0.9

2 release files

1.0.8

2 release files

1.0.7

2 release files

1.0.6

2 release files

1.0.5

2 release files

1.0.4

2 release files

1.0.3

2 release files

1.0.2

2 release files

1.0.1

2 release files

1.0.0

2 release files

0.2.8

2 release files

0.2.7

2 release files

0.2.6

2 release files

0.2.5

2 release files

0.2.4

2 release files

0.2.3

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

0.0.19

2 release files

0.0.15

2 release files

0.0.14

2 release files

0.0.12

2 release files

0.0.10

2 release files

0.0.9

2 release files

0.0.8

2 release files

0.0.7

2 release files

0.0.6

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page