Skip to main content

faster-paddle

Fast, CPU-only OCR in Rust with Python bindings — a self-contained reimplementation of PaddleOCR's PP-OCRv6 detection + recognition pipeline powered by ONNX Runtime.

  • ⚡ CPU latency optimizations: fused image transforms, compact DB components, dynamic recognition scheduling, and model-specific input widths.
  • 📦 Self-contained — the tiny + small ONNX models are bundled inside the wheel. No paddlepaddle, no model downloads for tiny/small.
  • 🎚️ Three model sizes: tiny (default, fastest), small, and medium (higher accuracy; downloaded once on first use and cached).
  • 🦀 Pure-Rust pre/post-processing (detection DB decode, minAreaRect, perspective crop, CTC decode, reading-order text reconstruction). No OpenCV.
  • 🖥️ Prebuilt wheels for Linux, Windows, macOS (x86-64 + arm64).

Defaults automatically respect available physical cores, Linux CPU affinity, and container CPU limits. Reuse an engine to amortize model/session loading. See performance results and tradeoffs and the benchmark harnesses for reproducible measurements.


Install

pip install faster-paddle

Usage

import faster_paddle

# One-shot, using a shared default engine (lazily initialized):
with open("document.jpg", "rb") as f:
    result = faster_paddle.ocr(f.read())

print(result["text"])              # reading-order reconstructed text
for idx, b in result["bounds"].items():
    print(idx, b["text"], b["confidence"], b["topLeftCoord"], b["bottomRightCoord"])

Reuse an explicit engine (recommended for servers — load the models once):

from faster_paddle import OcrEngine

# model_size: "tiny" (default), "small", or "medium"
engine = OcrEngine(model_size="tiny", threads=None, det_max_side=1600)

result = engine.ocr(image_bytes)                 # raw jpeg/png/webp/bmp/tiff/gif bytes
result = engine.ocr_base64(b64_string)           # base64-encoded image

Optional preprocessing

ocr / ocr_base64 take four optional flags (all default False), applied — when enabled — in the optimal order, all in fast parallel Rust:

result = engine.ocr(
    image_bytes,
    resize=True,     # 1. downscale source to ≤ 2100×3000 (aspect preserved) if larger
    denoise=True,    # 2. fast Non-Local-Means denoise (grayscale)
    deskew=True,     # 3. detect skew (Canny + Hough) and rotate to straighten
    binarize=True,   # 4. Sauvola adaptive thresholding (clean black/white)
)

Order rationale: resize first (everything downstream is then faster), denoise before angle detection and thresholding, deskew on the cleaned image, binarize last to produce the final B/W. resize can reduce source/crop work, but may add an extra resampling pass when detection already hits its size cap; benchmark it on your input. Any of denoise/deskew/binarize converts the image to grayscale.

Returned bounds are always in the original image's coordinate space — even when resize or deskew changes the working image, the boxes are mapped back so they line up with your input.

Preprocess only (no OCR)

prepare runs the same preprocessing in one pass and returns the prepared image as PNG bytes (grayscale once any of denoise/deskew/binarize is on, else color). If every option is False the original bytes are returned unchanged.

prepared = engine.prepare(image_bytes, resize=True, denoise=True, deskew=True, binarize=False)
# or module-level:  faster_paddle.prepare(image_bytes, resize=True, ...)

with open("prepared.png", "wb") as f:
    f.write(prepared)
# you can also feed it straight back in:
result = engine.ocr(prepared)

Model sizes

size bundled det+rec notes
tiny ✅ yes ~6 MB default, fastest, lightweight
small ✅ yes ~31 MB better accuracy
medium ⬇️ on demand ~138 MB best accuracy; downloaded once from the GitHub release and cached under your user cache dir

tiny and small are embedded in the wheel (offline). medium exceeds PyPI's file-size limit, so the first OcrEngine(model_size="medium") downloads it once (needs network that time only) and caches it for subsequent runs.

Result shape

{
  "text": "full reconstructed text...",
  "structured_text": "layout-preserving text (see below)",
  "bounds": {
     0: {
        "topLeftCoord":     (x1, y1),
        "bottomRightCoord": (x2, y2),
        "text":             "line text",
        "confidence":       0.97,
     },
     1: { ... },
  }
}

text and bounds match the JSON contract of the original paddle-ocr-api service, so it is a drop-in replacement.

structured_text

A spatial reconstruction that reads left-to-right, top-to-bottom while preserving the visual layout: vertical whitespace gaps split the page into columns/panes (each read fully before the next), and within each one the rows are laid out as a monospace grid, so indentation (tree nesting) and aligned sub-columns (key/value tables) are kept. Single-glyph UI icon noise is dropped.

Use structured_text for screenshots, forms, table/tree UIs, and code — anything where spatial structure carries meaning. Use text for dense multi-column prose: there the absolute pixel spacing of structured_text produces very wide lines, so the column-merging text reconstruction reads better. Both are always returned, so you can pick per use case.

Example structured_text for a two-pane file-tree + settings UI:

Project
 src (14)
   main.rs
   parser.rs
   utils.rs
 tests
 docs

Setting                                            Value
        max_connections                            128
        request_timeout_seconds                    30
        cache_size_mb                              512

API

faster_paddle.ocr(image, resize=False, denoise=False, deskew=False, binarize=False) -> dict OCR encoded image bytes (shared default engine).
faster_paddle.ocr_base64(image_base64, resize=False, denoise=False, deskew=False, binarize=False) -> dict OCR a base64 image string.
OcrEngine(model_size="tiny", threads=None, rec_batch=None, det_max_side=None, *, det_min_side=None, rec_min_width=None) Construct a reusable engine.
OcrEngine.ocr(image, resize=False, denoise=False, deskew=False, binarize=False) -> dict OCR encoded image bytes.
OcrEngine.ocr_base64(image_base64, resize=False, denoise=False, deskew=False, binarize=False) -> dict OCR a base64 image string.
faster_paddle.prepare(image, resize=False, denoise=False, deskew=False, binarize=False) -> bytes Preprocess only; returns PNG bytes (no OCR).
OcrEngine.prepare(image, resize=False, denoise=False, deskew=False, binarize=False) -> bytes Preprocess only; returns PNG bytes (no OCR).
  • resize/denoise/deskew/binarize: optional preprocessing (see above).
  • model_size: "tiny" (default), "small", or "medium".
  • threads: total CPU budget; defaults to available physical cores, limited by affinity and OS/container quotas. An explicit positive value overrides discovery.
  • rec_batch: actual maximum crops per recognition tensor (default 1). Independent crops share the worker pool. Larger caps are available for tuning.
  • det_max_side: detector long-side limit (default 1600); recognition crops still come from the original source. Lower values can miss small text.
  • det_min_side: minimum detector short side (default 736). Set 0 to disable minimum-side upscaling. Dimensions are still rounded to multiples of 32.
  • rec_min_width: recognition padding floor, automatically 64 for tiny, 96 for small, and 320 for medium. Set 320 for the reference padding behavior. Shorter padding changes context and can change text/confidence; validate your languages and documents when migrating.
  • engine.config: dictionary of resolved thread counts, pool size and input limits, for diagnostics and reproducible benchmarks.

Calls release the GIL. Inference on a shared engine is serialized; decode and layout can overlap. Independent engines have independent CPU budgets, so set threads explicitly when running several engines concurrently.

Multiple images

engine = OcrEngine(model_size="small")
results = engine.ocr_batch([image_a_bytes, image_b_bytes, image_c_bytes])
# results[i] corresponds to input i and has the same shape as engine.ocr(...).
# Module-level faster_paddle.ocr_batch([...]) uses the shared tiny engine.

ocr_batch(images, *, batch_size=4, resize=False, denoise=False, deskew=False, binarize=False) processes bounded windows of images. It decodes/preprocesses in parallel and schedules recognition crops across pages. Each page keeps its own detector resolution and coordinate transform. batch_size controls the number of decoded pages resident in a window; it is independent of rec_batch. Input bytes are borrowed when possible. An empty list returns []; invalid images raise an error with their zero-based input index.

Batching helps fill recognition workers on sparse pages and can improve throughput. It is not guaranteed to accelerate dense pages, and the caller waits for the complete result list. Use ocr for minimum time to the first page's result. Minor floating-point confidence differences can occur between the single-line and multi-page execution paths.

Parallelism

  • The detector uses up to 8 threads within the CPU budget.
  • Recognition uses up to one worker per available physical core, capped at 32 (8 for medium) to limit model memory. Jobs are scheduled dynamically, largest estimated jobs first. Sparse long-line jobs use a separate pool of up to four sessions with up to four threads each, within the same CPU budget. Session construction is parallel in bounded groups to reduce startup latency.
  • An engine-local Rayon pool handles image and geometry work within the same CPU budget. ONNX workers spin during inference to reduce wakeup latency, and stop spinning immediately when their inference call finishes.
  • Input tensor buffers are reused. Recognition budgets include actual padding.

Advanced overrides (read at engine construction): OCR_THREADS, OCR_DET_THREADS, REC_POOL, REC_BUDGET, RAYON_NUM_THREADS, OCR_REC_MIN_WIDTH, OCR_DET_MIN_SIDE, OCR_MEMPAT, OCR_PREPACK, OCR_DET_SPIN, OCR_REC_SPIN, OCR_APPROX_GELU. Approximate GELU stays off: its measured end-to-end benefit was inconsistent and it changed a detection. Constructor arguments take precedence over their corresponding environment values; detector/worker/Rayon counts are capped by the total CPU budget. The CPU-count defaults are measured heuristics, not online autotuning; benchmark other CPU architectures with benchmarks/cpu_latency.py before overriding them.


How it works

The pipeline faithfully mirrors PaddleOCR's lightweight path:

  1. Detection — resize (min-side 736, cap the longer side at det_max_side = 1600 by default vs PaddleOCR's 4000, round to ×32), normalize (BGR mean/std), run the DB detector. Lower detector resolution reduces work but can miss small text; recognition still crops from the full-res image.
  2. DB post-process — threshold 0.2, connected components, minAreaRect, box score ≥ 0.4, unclip ratio 1.4, rescale to source coordinates.
  3. Sort boxes top-to-bottom / left-to-right; crop each via perspective warp.
  4. Recognition — resize each crop to H=48, normalize, batch, run the CTC recognizer (6,906 output classes for tiny, 18,710 for small), greedy CTC decode.
  5. Reconstruct reading-order text with dynamic column/line detection.

Detection matches PaddlePaddle at 96 % IoU>0.5 with 0.93 character-level similarity on the recognized text; the residual difference is ONNX-Runtime vs PaddlePaddle floating-point numerics, not the algorithm.

The bundled tiny and small PP-OCRv6 models were exported with paddle2onnx.

Building from source

pip install maturin
maturin develop --release      # build + install into the current environment
# or
maturin build --release        # produce a wheel in target/wheels/

Requires a Rust toolchain. ONNX Runtime is fetched automatically by the ort crate at build time and linked into the extension.

Tests

cargo test --release                 # Rust unit tests (geometry, resize, CTC)
maturin develop --release            # then the Python integration tests:
python -m pytest tests -q

The tests cover geometry equivalence, fused resize/padding, recognition caps, known text, both small models, batch ordering and transforms, buffer reuse, concurrent calls, error recovery, CPU affinity/budgets, and preprocessing. CI runs Rust and Python tests before publishing. Install pytest and pillow for the Python suite. Performance comparisons live under benchmarks/; the small bundled test corpus is not a comprehensive OCR accuracy benchmark.

License

MIT

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

faster_paddle-1.0.2.tar.gz (32.4 MB view details)

Uploaded Source

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

faster_paddle-1.0.2-cp38-abi3-win_amd64.whl (40.7 MB view details)

Uploaded CPython 3.8+Windows x86-64

faster_paddle-1.0.2-cp38-abi3-manylinux_2_34_x86_64.whl (41.8 MB view details)

Uploaded CPython 3.8+manylinux: glibc 2.34+ x86-64

faster_paddle-1.0.2-cp38-abi3-manylinux_2_28_aarch64.whl (42.7 MB view details)

Uploaded CPython 3.8+manylinux: glibc 2.28+ ARM64

faster_paddle-1.0.2-cp38-abi3-macosx_11_0_arm64.whl (40.6 MB view details)

Uploaded CPython 3.8+macOS 11.0+ ARM64

File details

Details for the file faster_paddle-1.0.2.tar.gz.

File metadata

  • Download URL: faster_paddle-1.0.2.tar.gz
  • Upload date:
  • Size: 32.4 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: maturin/1.15.0

File hashes

Hashes for faster_paddle-1.0.2.tar.gz
Algorithm Hash digest
SHA256 f33abc11985030a506956fc6fb2ebc120d46e14b11624801e93f02e285f7ece5
MD5 3d8ff5dd67f12f31f51e19caf81591ce
BLAKE2b-256 94642e807ba0e0e609c6426901594db4e55b2cb5c0c3df41832f5d15509b6a82

See more details on using hashes here.

File details

Details for the file faster_paddle-1.0.2-cp38-abi3-win_amd64.whl.

File metadata

File hashes

Hashes for faster_paddle-1.0.2-cp38-abi3-win_amd64.whl
Algorithm Hash digest
SHA256 a67c3d8cef7b021373d5fca745a61c723ee14add3fc8a84fc6668349d74469a1
MD5 bea031869ab7d9acfe763811e810bccd
BLAKE2b-256 e7784e3b701c276cfb280137718676fdbfd030f7975dd3c7f762626ed1d30277

See more details on using hashes here.

File details

Details for the file faster_paddle-1.0.2-cp38-abi3-manylinux_2_34_x86_64.whl.

File metadata

File hashes

Hashes for faster_paddle-1.0.2-cp38-abi3-manylinux_2_34_x86_64.whl
Algorithm Hash digest
SHA256 dd5812d6c97c85f7b767fb2209d538668cbf964542d5618605e8352b4db0b855
MD5 cb7cca9273b1285079ee368759b47953
BLAKE2b-256 ae35158db96e33239de475871b1ad7366d672006025666c27f13d0d1150f8c00

See more details on using hashes here.

File details

Details for the file faster_paddle-1.0.2-cp38-abi3-manylinux_2_28_aarch64.whl.

File metadata

File hashes

Hashes for faster_paddle-1.0.2-cp38-abi3-manylinux_2_28_aarch64.whl
Algorithm Hash digest
SHA256 b76be0826955424c76b5e27aab8a307133f48646c8b306fb749f47ff2e3a29c1
MD5 0f231c935787c82c01c46e9ece351e5b
BLAKE2b-256 69fc4970428ecd869fe0856a7f1fceb3a8611abdc42f3145e2e959d35107767e

See more details on using hashes here.

File details

Details for the file faster_paddle-1.0.2-cp38-abi3-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for faster_paddle-1.0.2-cp38-abi3-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 f81e680a6e84342a504ac7a473eb1bfbf4b8cf34fcc3af4e56fe41d1d48654e7
MD5 eceb0ded799698977706a49d99030b1b
BLAKE2b-256 78f91a0cc1bcbfa21d58f5e3617877b214b8c446918ecf314591475bc4fe600d

See more details on using hashes here.

Release history Release notifications | RSS feed

1.0.4

5 files

1.0.3

5 files

This release

1.0.2 This release

5 files

1.0.1

5 files

1.0.0

5 files

0.0.14

5 files

0.0.13

5 files

0.0.12

5 files

0.0.11

5 files

0.0.10

5 files

0.0.9

5 files

0.0.8

5 files

0.0.7

5 files

0.0.6

5 files

0.0.5

5 files

0.0.4

4 files

0.0.3

4 files

0.0.2

5 files

0.0.1

5 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page