Khmer Document Parser v0.3.0
khmerdocparser is a smart, all-in-one command-line tool to extract Khmer text from both PDF and image files.
It intelligently handles PDFs by first attempting a fast, direct text extraction. If that fails (as with a scanned document), it automatically falls back to a powerful OCR engine with image preprocessing to ensure the best possible results.
Features
- Universal Support: Handles both PDF and common image files (
.png,.jpg, etc.). - Smart PDF Parsing: Uses
pdfplumberfor native PDFs and falls back to Tesseract OCR for scanned PDFs. - Advanced OCR: Applies image preprocessing (Grayscaling, Binarization, Noise Removal) for high accuracy on scanned documents.
- User-Friendly: Provides progress bars and detailed logging.
Prerequisites
This package requires two crucial external dependencies: Poppler and Tesseract OCR.
1. Tesseract OCR Installation
You must install the Tesseract engine and the Khmer (khm) language pack.
- Windows: Download and run the installer from UB-Mannheim's GitHub. Ensure you select the Khmer language pack during installation. Add Tesseract to your system's PATH.
- macOS:
brew install tesseract tesseract-lang - Linux (Ubuntu/Debian):
sudo apt-get install tesseract-ocr tesseract-ocr-khm
2. Poppler Installation
- Windows: Download the latest binary from here, extract it, and add the
binfolder to your system's PATH. - macOS:
brew install poppler - Linux (Ubuntu/Debian):
sudo apt-get install poppler-utils
Installation
Once Poppler and Tesseract are installed, you can install or upgrade the package from PyPI:
pip install --upgrade khmerdocparser
Usage
The command is the same for any supported file type.
Extract from a PDF or Image
# Process a PDF
khmerdocparser /path/to/your/document.pdf
# Process an image
khmerdocparser /path/to/your/scanned_image.png
Save Output to a File
This is the recommended way to view Khmer text correctly.
khmerdocparser my_document.pdf -o my_document_text.txt
Specifying Paths Manually (if not in PATH)
khmerdocparser doc.pdf --tesseract_path "C:\Tesseract\tesseract.exe" --poppler_path "C:\Poppler\bin"
Metadata
Release files for khmerdocparser 0.3.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| khmerdocparser-0.3.0.tar.gz | 5.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| khmerdocparser-0.3.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 11.6 kB
Release files / khmerdocparser-0.3.0.tar.gz
| Download URL | khmerdocparser-0.3.0.tar.gz |
|---|---|
| Size | 5.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
46c9ffcbbd0bada4d28def318c48aba6e358defb4e06269ba42d8304de526070
|
|
BLAKE2b-256 checksum How to use checksums |
597fbf59ff53214c2012111e01cd602e30158b2ef2b1ff2c0acda41b343803fb
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.1.0 CPython/3.10.16
|
Release files / khmerdocparser-0.3.0-py3-none-any.whl
| Download URL | khmerdocparser-0.3.0-py3-none-any.whl |
|---|---|
| Size | 6.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
2235183efa6f3f3f6674820b1c30901f278c420cd41d25671e8eaaf0ef69c195
|
|
BLAKE2b-256 checksum How to use checksums |
c1bc2b64bc95eda7b6db264f3197dabbbdee7c9e70a0ab3e11f816b7d0d5cec8
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.1.0 CPython/3.10.16
|