Skip to main content

Optical Character Recognition for images, Pdfs, zip files, tif files.

What you can expect from this repository:

  • Efficient ways to get textual information from your documents like images, pdfs, zip files.

Quick Tour

Get text from documents and save results in JSON.

Installation

Developer mode

pip install python-ocr

For tesseractOcr process

storage_type='local/aws' #currently only local and aws supported. local storage_path='Desired path of your OS where you want to store the output' # for local storage. local storage_path='S3 bucket' # for AWS storage (CASE SENSTIVE).

e.g. for Storing output to AWS

config={'storage_type':'AWS','storage_path:'your-bucket-name'}
from ocr import TesseractOcrProcessor
process=TesseractOcrProcessor(config)

For EasyOcr process

from ocr import EasyOcrProcessor
process=EasyOcrProcessor(config)

storage_type: type of storage local or aws.

storage_path: storage path is path where user wants to store the output result.

# Path of file
PATH=''

# reading image files
process.process_image(PATH)

# reading pdf files
process.process_pdf(PATH)

# reading zip files
process.process_zip(PATH)

Documentation:

The full package documentation is available here.

First of all, you have to create dict of storage_type and storage_path.

  1. storage_type: storage type is type of storage where the user wants to store the output result. It may be local or aws.

  2. storage_path: storage path is path where the user wants to store the output result.

    • if you want to store the file in local system than give the path of folder where user wants to store the result as storage_path.

    • if user wants to store the result in aws than in storage_path you have to give the bucket name.

config={'storage_type':'','storage_path':''}

Now create the object of EasyOcrProcessor which take the config as a object parameter.

process = EasyOcrProcessor(config)

Image process:

To read the text from image user have to call the process_image method of EasyOcrProcessor and pass the path of image file as a parameter in it. process_image method store the output at the storage_path.

process.process_image(PATH)

Pdf process:

To read the text from pdf file user have to call the process_pdf method of EasyOcrProcessor and pass the path of pdf file as a parameter in it. process_pdf method convert each page of pdf into images and create the result of each page and store the result at the storage_path.

process.process_pdf(PATH)

Zip process:

To read the text from zip file user have to call the process_zip method of EasyOcrProcessor and pass the path of zip file as a parameter in it. Zip should contain only files with valid extensions. process_zip method extract each file of zip one by one and save the result at the storage path.

process.process_zip(PATH)

Result output:

[{
        "left": 125,
        "top": 141,
        "right": 259,
        "bottom": 161,
        "text": "Folin MGA-5875",
        "confidence": 0.3961432168382489
    },
    {
        "left": 1115,
        "top": 140,
        "right": 1272,
        "bottom": 161,
        "text": "OM8 N0 : 2126-0006",
        "confidence": 0.41482855467690777
    },
    {
        "left": 1281,
        "top": 139,
        "right": 1498,
        "bottom": 165,
        "text": "Epiration Datc 12/31/2024",
        "confidence": 0.40780972855935615
}]

Release files for python-ocr 0.1.5

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for python-ocr 0.1.5
File Size Uploaded
python_ocr-0.1.5.tar.gz 6.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for python-ocr 0.1.5
File Interpreter ABI Platform
python_ocr-0.1.5-py3-none-any.whl Python 3 none any Details

Total release size: 13.9 kB

Release files / python_ocr-0.1.5.tar.gz

Download URL python_ocr-0.1.5.tar.gz
Size 6.8 kB
Tags Source
SHA-256 checksum
How to use checksums
c59c783ae8bc0137bba73cd3e04eb36994b087615e8b259ca198cb57a8baf6ff
BLAKE2b-256 checksum
How to use checksums
e9ec8b678905965ad8e97e1fdcdc6895ad8355c3d0ab5fa0216a5395322ff33d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.1 CPython/3.8.10

Release files / python_ocr-0.1.5-py3-none-any.whl

Download URL python_ocr-0.1.5-py3-none-any.whl
Size 7.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
57e811b09b69951e26693c2c9c00d6cc9630536d439d7a9416a8fd0f2ef84c80
BLAKE2b-256 checksum
How to use checksums
ea110fe8c7493552db408e4bd08dab7691971d0da7e73b59622f57b7babec1cf
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.1 CPython/3.8.10

Release history Release notifications | RSS feed

This release

0.1.5 This release

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page