Skip to main content

OCRmyPDF PaddleOCR Plus

A maintained fork of the PaddleOCR engine plugin for OCRmyPDF.

This fork is based on the batched single-model implementation from upstream PR #6 and adds practical multilingual model selection, quieter CPU inference, runtime tuning, tests, CI, and a PyPI release workflow.

Highlights

  • One PaddleOCR model per process with batched page inference instead of recreating models for every page.
  • Native English + Hindi + Sanskrit support through PaddleOCR's official PP-OCRv5 Devanagari recognition model.
  • PP-OCRv6 for normal English OCR by default.
  • Quiet CPU defaults: oneDNN/MKLDNN is disabled unless explicitly enabled, avoiding the repeated ReduceMeanCheckIfOneDNNSupport console output seen with some PaddlePaddle versions.
  • Suppresses Paddle's non-actionable No ccache found inference warning.
  • CPU thread, batch size, model family, detector, and recognizer overrides.
  • Optional PaddleOCR-VL support.
  • Prepared for PyPI Trusted Publishing.

Installation

OCRmyPDF still needs its normal operating-system dependencies. On Debian/Ubuntu or WSL, installing the distro package is a convenient way to get them:

sudo apt update
sudo apt install ocrmypdf

PyPI

The fork is published under a separate distribution name because ocrmypdf-paddleocr is already owned by another PyPI maintainer.

pip install ocrmypdf-paddleocr-plus

For an isolated global OCRmyPDF CLI with uv:

uv tool install --python 3.11 \
  --with ocrmypdf-paddleocr-plus \
  ocrmypdf
hash -r

The Python import/plugin name intentionally stays unchanged:

ocrmypdf_paddleocr

Install directly from GitHub

Before the first PyPI release, or when testing a branch:

uv tool install --force --python 3.11 \
  --with 'ocrmypdf-paddleocr-plus @ git+https://github.com/MarkFazekas/ocrmypdf-paddleocr.git' \
  ocrmypdf
hash -r

Basic usage

ocrmypdf \
  --plugin ocrmypdf_paddleocr \
  --language eng \
  --jobs 2 \
  --paddle-batch-size 2 \
  --output-type pdf \
  --optimize 0 \
  input.pdf output.pdf

English, Hindi and Sanskrit

OCRmyPDF uses Tesseract-style language codes. This fork understands the language list instead of silently using only the first entry.

For a document containing English, Hindi and Sanskrit:

ocrmypdf \
  --plugin ocrmypdf_paddleocr \
  --language eng+hin+san \
  --jobs 2 \
  --paddle-batch-size 2 \
  --output-type pdf \
  --optimize 0 \
  input.pdf output.pdf

When any supported Devanagari language is requested, the plugin automatically uses PaddleOCR's official PP-OCRv5 Devanagari pipeline:

PP-OCRv5_server_det
devanagari_PP-OCRv5_mobile_rec

That recognition model is multilingual and includes English plus Devanagari languages, so mixed English/Hindi/Sanskrit pages can be recognized in one pass.

Common OCRmyPDF codes:

OCRmyPDF code Paddle code Language
eng en English
hin hi Hindi
san sa Sanskrit
mar mr Marathi
nep ne Nepali
bho bho Bhojpuri
mai mai Maithili

English-only OCR keeps PaddleOCR's normal/current English model selection (PP-OCRv6 on supported PaddleOCR versions).

The automatic Devanagari profile accepts English plus Devanagari languages. An unrelated mixed-script request such as hin+fra is rejected instead of silently selecting the wrong recognizer. Advanced users can override the models explicitly.

Performance and batching

The plugin shares one PaddleOCR instance across OCRmyPDF worker threads. Requests are collected into batches controlled by --paddle-batch-size.

A good CPU starting point is:

--jobs 2 --paddle-batch-size 2

Paddle itself uses multiple CPU threads, so a high OCRmyPDF --jobs value is usually counterproductive.

You can tune Paddle's CPU thread count:

--paddle-cpu-threads 8

Quiet CPU mode and oneDNN

On CPU, this fork disables oneDNN/MKLDNN by default. This avoids noisy output such as:

ReduceMeanCheckIfOneDNNSupport

and avoids the oneDNN compatibility path that has caused failures with some PaddlePaddle releases.

If oneDNN is stable on your machine and you want to benchmark its speedup:

--paddle-enable-mkldnn

Internal Paddle/PaddleX INFO logging is also hidden by default. To show it:

--paddle-show-log

Paddle-specific options

Option Meaning
`--paddle-engine classic vl`
--paddle-batch-size N Maximum pages sent to one Paddle batch
--paddle-cpu-threads N Paddle inference threads on CPU
--paddle-enable-mkldnn Opt in to oneDNN/MKLDNN on CPU
--paddle-show-log Show Paddle/PaddleX internal INFO logs
--paddle-use-gpu Use a GPU-enabled PaddlePaddle installation
--paddle-ocr-version VERSION Override OCR family, e.g. PP-OCRv5
--paddle-det-model NAME Override detector model name
--paddle-rec-model NAME Override recognition model name
--paddle-det-model-dir DIR Use a local detector model directory
--paddle-rec-model-dir DIR Use a local recognizer model directory
--paddle-cls-model-dir DIR Use a local orientation model directory

Example explicit model selection:

ocrmypdf --plugin ocrmypdf_paddleocr \
  --paddle-det-model PP-OCRv5_server_det \
  --paddle-rec-model devanagari_PP-OCRv5_mobile_rec \
  input.pdf output.pdf

Batch a directory of PDFs

Example using an og/ input directory and ocr/ output directory:

mkdir -p ocr

find og -maxdepth 1 -type f -iname '*.pdf' -print0 |
while IFS= read -r -d '' f; do
    out="ocr/$(basename "$f")"

    if [ -f "$out" ]; then
        echo "SKIP: $out already exists"
        continue
    fi

    echo "OCR: $f -> $out"
    ocrmypdf \
        --plugin ocrmypdf_paddleocr \
        --language eng \
        --jobs 2 \
        --paddle-batch-size 2 \
        --output-type pdf \
        --optimize 0 \
        "$f" "$out"
done

PaddleOCR-VL

Install the optional dependencies:

pip install 'ocrmypdf-paddleocr-plus[vl]'

Then:

ocrmypdf --plugin ocrmypdf_paddleocr \
  --paddle-engine vl input.pdf output.pdf

VL is substantially heavier than the classic OCR pipeline.

Development

python -m pip install -e '.[dev]'
pytest -q
python -m build
python -m twine check dist/*

CI validates the lightweight language/model-selection tests and package build. Full Paddle inference is intentionally an integration test because the runtime and model downloads are large.

PyPI release

Releases are manual but fully automated: open Actions -> Release -> Run workflow, select the main branch, enter a version such as 0.2.0, and run it. The workflow tests the code, creates the Git tag, builds and validates the package, publishes through PyPI Trusted Publishing, creates the GitHub Release, and attaches the wheel/sdist.

See RELEASING.md for the one-time main default-branch and PyPI Trusted Publisher setup.

The distribution name is:

ocrmypdf-paddleocr-plus

while the import and plugin name remains:

ocrmypdf_paddleocr

Troubleshooting

The system OCRmyPDF runs instead of the uv-installed one

Check:

type -a ocrmypdf

If the shell cached /usr/bin/ocrmypdf after an uv tool install, refresh the command cache:

hash -r

PaddlePaddle version

The PyPI package currently pins PaddlePaddle to the 3.2.x line because that is the stable CPU combination used by this fork. The plugin disables MKLDNN by default to avoid the noisy/problematic oneDNN path.

Credits

This fork builds on:

License

MPL-2.0. See LICENSE.

Release files for ocrmypdf-paddleocr-plus 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for ocrmypdf-paddleocr-plus 0.2.0
File Size Uploaded
ocrmypdf_paddleocr_plus-0.2.0.tar.gz 39.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for ocrmypdf-paddleocr-plus 0.2.0
File Interpreter ABI Platform
ocrmypdf_paddleocr_plus-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 63.0 kB

Release files / ocrmypdf_paddleocr_plus-0.2.0.tar.gz

Download URL ocrmypdf_paddleocr_plus-0.2.0.tar.gz
Size 39.0 kB
Tags Source
SHA-256 checksum
How to use checksums
35a0a690d903f5d3d4fb19c58cce7888cccf1fe4a75905340b1efd1ff678b38b
BLAKE2b-256 checksum
How to use checksums
67d06a4c9b543d892aa415a4eae32f741b17fe3bc89457b63fc441baf164b1c0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.

Transparency log

Release files / ocrmypdf_paddleocr_plus-0.2.0-py3-none-any.whl

Download URL ocrmypdf_paddleocr_plus-0.2.0-py3-none-any.whl
Size 24.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
f5ebec03124b960a6f3936501a3458f030404f3f08ad7977e717265488c1f704
BLAKE2b-256 checksum
How to use checksums
fb25e274f848f85f315b5080edf4c6b768585de73719042a1b6dc1cccaa05782
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.2.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page