Awesome OCR toolkits based on PaddlePaddle （8.6M ultra-lightweight pre-trained model, support training and deployment among server, mobile, embeded and IoT devices

These details have not been verified by PyPI

Project links

Project description

Note from Maintainer:

This is a fork of the fork of PaddleOCR which makes it compatible with paddle 2.6.1 which enables you to use PaddleOCR with unstructured with a 4x performance increase over using Tesseract.

This is a fork of the PaddleOCR repository, created with the purpose of complying with an Apache license. The original repo contains at least one dependency that is not Apache compliant so this fork was created to remove any non-compliant dependencies.

The original documentation at the time of the fork is found below.

Introduction

PaddleOCR aims to create multilingual, awesome, leading, and practical OCR tools that help users train better models and apply them into practice.

📣 Recent updates

🔨2022.11 Add implementation of 4 cutting-edge algorithms：Text Detection DRRG, Text Recognition RFL, Image Super-Resolution Text Telescope，Handwritten Mathematical Expression Recognition CAN
2022.10 Release optimized JS version PP-OCRv3 model with 4.3M model size, 8x faster inference time, and a ready-to-use web demo
💥 Live Playback: Introduction to PP-StructureV2 optimization strategy. Scan the QR code below using WeChat, follow the PaddlePaddle official account and fill out the questionnaire to join the WeChat group, get the live link and 20G OCR learning materials (including PDF2Word application, 10 models in vertical scenarios, etc.)
🔥2022.8.24 Release PaddleOCR release/2.6
- Release PP-StructureV2，with functions and performance fully upgraded, adapted to Chinese scenes, and new support for Layout Recovery and one line command to convert PDF to Word;
- Layout Analysis optimization: model storage reduced by 95%, while speed increased by 11 times, and the average CPU time-cost is only 41ms;
- Table Recognition optimization: 3 optimization strategies are designed, and the model accuracy is improved by 6% under comparable time consumption;
- Key Information Extraction optimization：a visual-independent model structure is designed, the accuracy of semantic entity recognition is increased by 2.8%, and the accuracy of relation extraction is increased by 9.1%.
🔥2022.8 Release OCR scene application collection
- Release 9 vertical models such as digital tube, LCD screen, license plate, handwriting recognition model, high-precision SVTR model, etc, covering the main OCR vertical applications in general, manufacturing, finance, and transportation industries.
2022.8 Add implementation of 8 cutting-edge algorithms
- Text Detection: FCENet, DB++
- Text Recognition: ViTSTR, ABINet, VisionLAN, SPIN, RobustScanner
- Table Recognition: TableMaster
2022.5.9 Release PaddleOCR release/2.5
- Release PP-OCRv3: With comparable speed, the effect of Chinese scene is further improved by 5% compared with PP-OCRv2, the effect of English scene is improved by 11%, and the average recognition accuracy of 80 language multilingual models is improved by more than 5%.
- Release PPOCRLabelv2: Add the annotation function for table recognition task, key information extraction task and irregular text image.
- Release interactive e-book "Dive into OCR", covers the cutting-edge theory and code practice of OCR full stack technology.
more

🌟 Features

PaddleOCR support a variety of cutting-edge algorithms related to OCR, and developed industrial featured models/solution PP-OCR and PP-Structure on this basis, and get through the whole process of data production, model training, compression, inference and deployment.

It is recommended to start with the “quick experience” in the document tutorial

⚡ Quick Experience

Web online experience for the ultra-lightweight OCR: Online Experience
Mobile DEMO experience (based on EasyEdge and Paddle-Lite, supports iOS and Android systems): Sign in to the website to obtain the QR code for installing the App
One line of code quick use: Quick Start

📚 E-book: Dive Into OCR

Dive Into OCR

👫 Community

For international developers, we regard PaddleOCR Discussions as our international community platform. All ideas and questions can be discussed here in English.
For Chinese develops, Scan the QR code below with your Wechat, you can join the official technical discussion group. For richer community content, please refer to 中文README, looking forward to your participation.

🛠️ PP-OCR Series Model List（Update on September 8th）

Model introduction	Model name	Recommended scene	Detection model	Direction classifier	Recognition model
Chinese and English ultra-lightweight PP-OCRv3 model（16.2M）	ch_PP-OCRv3_xx	Mobile & Server	inference model / trained model	inference model / trained model	inference model / trained model
English ultra-lightweight PP-OCRv3 model（13.4M）	en_PP-OCRv3_xx	Mobile & Server	inference model / trained model	inference model / trained model	inference model / trained model
Chinese and English ultra-lightweight PP-OCRv2 model（11.6M）	ch_PP-OCRv2_xx	Mobile & Server	inference model / trained model	inference model / trained model	inference model / trained model
Chinese and English ultra-lightweight PP-OCR model (9.4M)	ch_ppocr_mobile_v2.0_xx	Mobile & server	inference model / trained model	inference model / trained model	inference model / trained model
Chinese and English general PP-OCR model (143.4M)	ch_ppocr_server_v2.0_xx	Server	inference model / trained model	inference model / trained model	inference model / trained model

For more model downloads (including multiple languages), please refer to PP-OCR series model downloads.
For a new language request, please refer to Guideline for new language_requests.
For structural document analysis models, please refer to PP-Structure models.

📖 Tutorials

👀 Visualization more

PP-OCRv3 Chinese model

PP-OCRv3 English model

PP-OCRv3 Multilingual model

PP-StructureV2

layout analysis + table recognition

SER (Semantic entity recognition)

RE (Relation Extraction)

🇺🇳 Guideline for New Language Requests

If you want to request a new language support, a PR with 1 following files are needed：

In folder ppocr/utils/dict, it is necessary to submit the dict text to this path and name it with {language}_dict.txt that contains a list of all characters. Please see the format example from other files in that folder.

If your language has unique elements, please tell me in advance within any way, such as useful links, wikipedia and so on.

More details, please refer to Multilingual OCR Development Plan.

📄 License

This project is released under Apache 2.0 license

Project details

These details have not been verified by PyPI

Project links

Release history Release notifications | RSS feed

This version

2.6.1.3.post2

Jun 12, 2024

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

vaaale_paddleocr-2.6.1.3.post2.tar.gz (719.4 kB view details)

Uploaded Jun 12, 2024 Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

The dropdown lists show the available interpreters, ABIs, and platforms. Enable javascript to be able to filter the list of wheel files.

vaaale_paddleocr-2.6.1.3.post2-py3-none-any.whl (919.0 kB view details)

Uploaded Jun 12, 2024 Python 3

File details

Details for the file vaaale_paddleocr-2.6.1.3.post2.tar.gz.

File metadata

Download URL: vaaale_paddleocr-2.6.1.3.post2.tar.gz
Upload date: Jun 12, 2024
Size: 719.4 kB
Tags: Source
Uploaded using Trusted Publishing? No
Uploaded via: twine/5.1.0 CPython/3.10.14

File hashes

Hashes for vaaale_paddleocr-2.6.1.3.post2.tar.gz
Algorithm	Hash digest
SHA256	`2846fe35c22399171c11a57dc695045d91aa04675967d1eb02f7a89f37c6a0b9`
MD5	`b84b3f6d3bca074df933d7460e997705`
BLAKE2b-256	`5eca41c697aea32b0150f20d5ff0ce826ccebf57765b9b379c4bfade34e6db09`

See more details on using hashes here.

File details

Details for the file vaaale_paddleocr-2.6.1.3.post2-py3-none-any.whl.

File metadata

Download URL: vaaale_paddleocr-2.6.1.3.post2-py3-none-any.whl
Upload date: Jun 12, 2024
Size: 919.0 kB
Tags: Python 3
Uploaded using Trusted Publishing? No
Uploaded via: twine/5.1.0 CPython/3.10.14

File hashes

Hashes for vaaale_paddleocr-2.6.1.3.post2-py3-none-any.whl
Algorithm	Hash digest
SHA256	`1e8618b134b3a3a9be07ee58e46034669c71af2f015b2abb5acb1196873779fe`
MD5	`a481efcc40c2636c29cc18c7363edccb`
BLAKE2b-256	`cff214ff8913716f94670a07343ef387f34356c0e0d87becd6835b2da6c90c7f`

See more details on using hashes here.

vaaale-paddleocr 2.6.1.3.post2

Navigation

Verified details

Maintainers

Unverified details

Project links

Meta

Classifiers

Project description

Note from Maintainer:

Introduction

📣 Recent updates

🌟 Features

⚡ Quick Experience

📚 E-book: Dive Into OCR

👫 Community

🛠️ PP-OCR Series Model List（Update on September 8th）

📖 Tutorials

👀 Visualization more

🇺🇳 Guideline for New Language Requests

📄 License

Project details

Verified details

Maintainers

Unverified details

Project links

Meta

Classifiers

Release history Release notifications | RSS feed

Download files

Source Distribution

Built Distribution

File details

File metadata

File hashes

File details

File metadata

File hashes