faster-paddle
Fast, CPU-only OCR in Rust with Python bindings — a self-contained reimplementation of PaddleOCR's PP-OCRv6 detection + recognition pipeline powered by ONNX Runtime.
- ⚡ CPU latency optimizations: fused image transforms, compact DB components, dynamic recognition scheduling, and model-specific input widths.
- 🗂️ Multi-image OCR with
ocr_batch: ordered results and shared recognition work across pages. See batch usage. - 📦 Self-contained — the tiny + small ONNX models are bundled inside the
wheel. No
paddlepaddle, no model downloads for tiny/small. - 🎚️ Three model sizes:
tiny(default, fastest),small, andmedium(higher accuracy; downloaded once on first use and cached). - 🦀 Pure-Rust pre/post-processing (detection DB decode,
minAreaRect, perspective crop, CTC decode, reading-order text reconstruction). No OpenCV. - 🖥️ Prebuilt wheels for Linux, Windows, macOS (x86-64 + arm64).
Defaults automatically respect available physical cores, Linux CPU affinity, and container CPU limits. Reuse an engine to amortize model/session loading. See performance results and tradeoffs and the benchmark harnesses for reproducible measurements.
Install
pip install faster-paddle
v1.0.4 upgrades the bundled native ONNX Runtime from 1.24.2 to 1.28.0.
This runtime is linked into the Rust extension; no Python onnxruntime package
is required. Inspect it with faster_paddle.__runtime_build__.
See the full OCR upgrade measurements.
Usage
import faster_paddle
# One-shot, using a shared default engine (lazily initialized):
with open("document.jpg", "rb") as f:
result = faster_paddle.ocr(f.read())
print(result["text"]) # reading-order reconstructed text
for idx, b in result["bounds"].items():
print(idx, b["text"], b["confidence"], b["topLeftCoord"], b["bottomRightCoord"])
Reuse an explicit engine (recommended for servers — load the models once):
from faster_paddle import OcrEngine
# model_size: "tiny" (default), "small", or "medium"
engine = OcrEngine(model_size="tiny", threads=None, det_max_side=1600)
result = engine.ocr(image_bytes) # raw jpeg/png/webp/bmp/tiff/gif bytes
result = engine.ocr_base64(b64_string) # base64-encoded image
Multiple images
Available in v1.0.2 and later. Pass a list of encoded image bytes and get one result dictionary per image, in the same order:
from pathlib import Path
from faster_paddle import OcrEngine
paths = [Path("page1.png"), Path("page2.jpg"), Path("page3.png")]
images = [path.read_bytes() for path in paths]
engine = OcrEngine(model_size="small")
results = engine.ocr_batch(images, batch_size=4)
for path, result in zip(paths, results):
print(path.name, result["text"])
The shared default tiny engine also supports batches:
import faster_paddle
results = faster_paddle.ocr_batch(images, batch_size=4, resize=True)
ocr_batch(images, *, batch_size=4, resize=False, denoise=False, deskew=False, binarize=False) processes bounded windows of images. It decodes/preprocesses
in parallel and schedules recognition crops across pages. Each page keeps its
own detector resolution and coordinate transform. batch_size controls the
number of decoded pages resident in a window; it is independent of rec_batch.
Input bytes are borrowed when possible. An empty list returns []; invalid
images raise an error with their zero-based input index.
Batching helps fill recognition workers on sparse pages and can improve
throughput. It is not guaranteed to accelerate dense pages, and the caller waits
for the complete result list. Use ocr for minimum time to the first page's
result. Minor floating-point confidence differences can occur between the
single-line and multi-page execution paths.
Optional preprocessing
ocr, ocr_base64, and ocr_batch take four optional flags (all default False), applied —
when enabled — in the optimal order, all in fast parallel Rust:
result = engine.ocr(
image_bytes,
resize=True, # 1. downscale source to ≤ 2100×3000 (aspect preserved) if larger
denoise=True, # 2. fast Non-Local-Means denoise (grayscale)
deskew=True, # 3. detect skew (Canny + Hough) and rotate to straighten
binarize=True, # 4. Sauvola adaptive thresholding (clean black/white)
)
Order rationale: resize first (everything downstream is then faster), denoise
before angle detection and thresholding, deskew on the cleaned image, binarize
last to produce the final B/W. resize can reduce source/crop work, but may add an extra resampling pass when
detection already hits its size cap; benchmark it on your input. Any of denoise/deskew/binarize converts the
image to grayscale.
Returned bounds are always in the original image's coordinate space — even
when resize or deskew changes the working image, the boxes are mapped back so
they line up with your input.
Preprocess only (no OCR)
prepare runs the same preprocessing in one pass and returns the prepared image
as PNG bytes (grayscale once any of denoise/deskew/binarize is on, else
color). If every option is False the original bytes are returned unchanged.
prepared = engine.prepare(image_bytes, resize=True, denoise=True, deskew=True, binarize=False)
# or module-level: faster_paddle.prepare(image_bytes, resize=True, ...)
with open("prepared.png", "wb") as f:
f.write(prepared)
# you can also feed it straight back in:
result = engine.ocr(prepared)
Model sizes
| size | bundled | det+rec | notes |
|---|---|---|---|
tiny |
✅ yes | ~6 MB | default, fastest, lightweight |
small |
✅ yes | ~31 MB | better accuracy |
medium |
⬇️ on demand | ~138 MB | best accuracy; downloaded once from the GitHub release and cached under your user cache dir |
tiny and small are embedded in the wheel (offline). medium exceeds PyPI's
file-size limit, so the first OcrEngine(model_size="medium") downloads it once
(needs network that time only) and caches it for subsequent runs.
Result shape
{
"text": "full reconstructed text...",
"structured_text": "layout-preserving text (see below)",
"bounds": {
0: {
"topLeftCoord": (x1, y1),
"bottomRightCoord": (x2, y2),
"text": "line text",
"confidence": 0.97,
},
1: { ... },
}
}
text and bounds match the JSON contract of the original paddle-ocr-api
service, so it is a drop-in replacement.
structured_text
A spatial reconstruction that reads left-to-right, top-to-bottom while preserving the visual layout: vertical whitespace gaps split the page into columns/panes (each read fully before the next), and within each one the rows are laid out as a monospace grid, so indentation (tree nesting) and aligned sub-columns (key/value tables) are kept. Single-glyph UI icon noise is dropped.
Use structured_text for screenshots, forms, table/tree UIs, and code —
anything where spatial structure carries meaning. Use text for dense
multi-column prose: there the absolute pixel spacing of structured_text
produces very wide lines, so the column-merging text reconstruction reads
better. Both are always returned, so you can pick per use case.
Example structured_text for a two-pane file-tree + settings UI:
Project
src (14)
main.rs
parser.rs
utils.rs
tests
docs
Setting Value
max_connections 128
request_timeout_seconds 30
cache_size_mb 512
API
faster_paddle.ocr(image, resize=False, denoise=False, deskew=False, binarize=False) -> dict |
OCR encoded image bytes (shared default engine). |
faster_paddle.ocr_batch(images, *, batch_size=4, resize=False, denoise=False, deskew=False, binarize=False) -> list[dict] |
OCR multiple encoded images; results preserve input order. |
faster_paddle.ocr_base64(image_base64, resize=False, denoise=False, deskew=False, binarize=False) -> dict |
OCR a base64 image string. |
OcrEngine(model_size="tiny", threads=None, rec_batch=None, det_max_side=None, *, det_min_side=None, rec_min_width=None) |
Construct a reusable engine. |
OcrEngine.ocr(image, resize=False, denoise=False, deskew=False, binarize=False) -> dict |
OCR encoded image bytes. |
OcrEngine.ocr_batch(images, *, batch_size=4, resize=False, denoise=False, deskew=False, binarize=False) -> list[dict] |
OCR multiple encoded images; results preserve input order. |
OcrEngine.ocr_base64(image_base64, resize=False, denoise=False, deskew=False, binarize=False) -> dict |
OCR a base64 image string. |
faster_paddle.prepare(image, resize=False, denoise=False, deskew=False, binarize=False) -> bytes |
Preprocess only; returns PNG bytes (no OCR). |
OcrEngine.prepare(image, resize=False, denoise=False, deskew=False, binarize=False) -> bytes |
Preprocess only; returns PNG bytes (no OCR). |
resize/denoise/deskew/binarize: optional preprocessing (see above).model_size:"tiny"(default),"small", or"medium".threads: total CPU budget; defaults to available physical cores, limited by affinity and OS/container quotas. An explicit positive value overrides discovery.batch_size: maximum decoded images per batch window (default 4, must be positive).rec_batch: actual maximum crops per recognition tensor (default 1). Independent crops share the worker pool. Larger caps are available for tuning.det_max_side: detector long-side limit (default 1600); recognition crops still come from the original source. Lower values can miss small text.det_min_side: minimum detector short side (default 736). Set 0 to disable minimum-side upscaling. Dimensions are still rounded to multiples of 32.rec_min_width: recognition padding floor, automatically 64 for tiny, 96 for small, and 320 for medium. Set 320 for the reference padding behavior. Shorter padding changes context and can change text/confidence; validate your languages and documents when migrating.engine.config: dictionary of resolved thread counts, pool size and input limits, for diagnostics and reproducible benchmarks.
Calls release the GIL. Inference on a shared engine is serialized; decode and
layout can overlap. Independent engines have independent CPU budgets, so set
threads explicitly when running several engines concurrently.
Parallelism
- The detector uses up to 8 threads within the CPU budget.
- Recognition uses up to one worker per available physical core, capped at 32 (8 for medium) to limit model memory. Jobs are scheduled dynamically, largest estimated jobs first. Sparse long-line jobs use a separate pool of up to four sessions with up to four threads each, within the same CPU budget. Session construction is parallel in bounded groups to reduce startup latency.
- An engine-local Rayon pool handles image and geometry work within the same CPU budget. ONNX workers spin during inference to reduce wakeup latency, and stop spinning immediately when their inference call finishes.
- Input tensor buffers are reused. Recognition budgets include actual padding.
Advanced overrides (read at engine construction): OCR_THREADS,
OCR_DET_THREADS, REC_POOL, REC_BUDGET, RAYON_NUM_THREADS, OCR_REC_MIN_WIDTH,
OCR_DET_MIN_SIDE, OCR_MEMPAT, OCR_PREPACK, OCR_DET_SPIN, OCR_REC_SPIN,
OCR_APPROX_GELU. Approximate GELU stays off: its measured end-to-end benefit
was inconsistent and it changed a detection.
Constructor arguments take precedence over their corresponding environment
values; detector/worker/Rayon counts are capped by the total CPU budget.
The CPU-count defaults are measured heuristics, not online autotuning; benchmark
other CPU architectures with benchmarks/cpu_latency.py before overriding them.
How it works
The pipeline faithfully mirrors PaddleOCR's lightweight path:
- Detection — resize (min-side 736, cap the longer side at
det_max_side= 1600 by default vs PaddleOCR's 4000, round to ×32), normalize (BGR mean/std), run the DB detector. Lower detector resolution reduces work but can miss small text; recognition still crops from the full-res image. - DB post-process — threshold 0.2, connected components,
minAreaRect, box score ≥ 0.4,unclipratio 1.4, rescale to source coordinates. - Sort boxes top-to-bottom / left-to-right; crop each via perspective warp.
- Recognition — resize each crop to H=48, normalize, batch, run the CTC recognizer (6,906 output classes for tiny, 18,710 for small), greedy CTC decode.
- Reconstruct reading-order text with dynamic column/line detection.
Detection matches PaddlePaddle at 96 % IoU>0.5 with 0.93 character-level similarity on the recognized text; the residual difference is ONNX-Runtime vs PaddlePaddle floating-point numerics, not the algorithm.
The bundled tiny and small PP-OCRv6 models were exported with paddle2onnx.
Building from source
pip install maturin
maturin develop --release # build + install into the current environment
# or
maturin build --release # produce a wheel in target/wheels/
Requires a Rust toolchain. ONNX Runtime is fetched automatically by the ort
crate at build time and linked into the extension.
Tests
cargo test --release # Rust unit tests (geometry, resize, CTC)
maturin develop --release # then the Python integration tests:
python -m pytest tests -q
The tests cover geometry equivalence, fused resize/padding, recognition caps,
known text, both small models, batch ordering and transforms, buffer reuse,
concurrent calls, error recovery, CPU affinity/budgets, and preprocessing. CI
runs Rust and Python tests before publishing. Install pytest and pillow
for the Python suite. Performance comparisons live under benchmarks/; the
small bundled test corpus is not a comprehensive OCR accuracy benchmark.
License
MIT
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distributions
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file faster_paddle-1.0.4.tar.gz.
File metadata
- Download URL: faster_paddle-1.0.4.tar.gz
- Upload date:
- Size: 32.4 MB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
maturin/1.15.0
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
585aa6ab1b29fe939db4055909058134e1881f677889cec209a59a50d2fa359d
|
|
| MD5 |
a9041f31f35a5c433b9e88b67001443c
|
|
| BLAKE2b-256 |
a317acf26453cb87703fc2625f40a027f06810d7c4896659c1c45c6eed396784
|
File details
Details for the file faster_paddle-1.0.4-cp38-abi3-win_amd64.whl.
File metadata
- Download URL: faster_paddle-1.0.4-cp38-abi3-win_amd64.whl
- Upload date:
- Size: 41.2 MB
- Tags: CPython 3.8+, Windows x86-64
- Uploaded using Trusted Publishing? No
- Uploaded via:
maturin/1.15.0
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
3a8a8e5a3fd7cf5c0a9c70e30d0bc7254ccdb9c1a588b2d60166ea885f95c871
|
|
| MD5 |
1a978896933f94c1687f9bd5985512d0
|
|
| BLAKE2b-256 |
c2fd745aeaa0549d52f9cbdeb87f807afa9ce11583ff7945c791b0b6bb8d8f80
|
File details
Details for the file faster_paddle-1.0.4-cp38-abi3-manylinux_2_34_x86_64.whl.
File metadata
- Download URL: faster_paddle-1.0.4-cp38-abi3-manylinux_2_34_x86_64.whl
- Upload date:
- Size: 43.1 MB
- Tags: CPython 3.8+, manylinux: glibc 2.34+ x86-64
- Uploaded using Trusted Publishing? No
- Uploaded via:
maturin/1.15.0
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
cf6218a183da869c6fb8e7ddceabc224b18361c1f1bc6763977171bb19fbb177
|
|
| MD5 |
6d355ae23fa23f1f51d28968d03785f0
|
|
| BLAKE2b-256 |
281fbe0964ba5beba10610e11c73f0e3f0a9ad799254997450c49a0444e233bc
|
File details
Details for the file faster_paddle-1.0.4-cp38-abi3-manylinux_2_28_aarch64.whl.
File metadata
- Download URL: faster_paddle-1.0.4-cp38-abi3-manylinux_2_28_aarch64.whl
- Upload date:
- Size: 43.7 MB
- Tags: CPython 3.8+, manylinux: glibc 2.28+ ARM64
- Uploaded using Trusted Publishing? No
- Uploaded via:
maturin/1.15.0
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
ec1aedbd5fba1cdf552d879565d1fb7d7d5b395bf19a06e9e2958b7ece538a7e
|
|
| MD5 |
9283398fb83408b3469c577a15ddff90
|
|
| BLAKE2b-256 |
231a2e9f864211753c1f6a2076db810d09eae6b843d08e152e6fcb55739a9609
|
File details
Details for the file faster_paddle-1.0.4-cp38-abi3-macosx_11_0_arm64.whl.
File metadata
- Download URL: faster_paddle-1.0.4-cp38-abi3-macosx_11_0_arm64.whl
- Upload date:
- Size: 41.3 MB
- Tags: CPython 3.8+, macOS 11.0+ ARM64
- Uploaded using Trusted Publishing? No
- Uploaded via:
maturin/1.15.0
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
a90a90da1fb4610c8d817afaca2158463b85af2ef6cd8011b5ca9dedeebb0f95
|
|
| MD5 |
f583990af147697a4c6711408c487310
|
|
| BLAKE2b-256 |
98c6fd9dc0a3b042c9dd50c49c7911e5b5183b394c34ccba79b0bc19fd77120f
|