pdf2odt
pdf2odt is a Python command-line utility and library that converts PDF documents into LibreOffice Writer (.odt) documents.
Key Features
- 100% Pure Python & Self-Contained: No need to install external system tools such as Poppler, Tesseract, or LibreOffice. Everything is installed via
pip. - High-Quality Page Rendering: Uses PyMuPDF (MuPDF) for fast and pixel-perfect rendering to PNG images with configurable DPI resolution.
- Smart Hybrid Text Extraction & OCR:
- Extracts native vector text directly from digital PDFs with 100% accuracy and near-zero latency.
- Automatically runs RapidOCR (ONNX Runtime) on scanned images or bitmap pages to recognize text.
- Character-Anchored Images (
as-char): Page images are embedded in the ODT document using odfdo with proportional dimensions and anchored as characters, ensuring consistent layout in LibreOffice Writer. - Fast & Multi-Threaded: Uses Python thread pooling to process pages concurrently across all available CPU cores.
Installation
Install pdf2odt using pip:
pip install pdf2odt
Or using Poetry:
poetry add pdf2odt
Command-Line Usage
1. Basic Conversion
Convert a PDF into an ODT document (renders pages at default 300 DPI):
pdf2odt --pdf document.pdf output.odt
2. Conversion with Text Extraction / OCR
Extract native text and perform OCR on images, inserting the text below each page image in the ODT document:
pdf2odt --pdf document.pdf --ocr output.odt
3. Custom Image Resolution
Set a custom image resolution in DPI (default is 300 DPI; use lower values like 150 DPI for smaller file sizes):
pdf2odt --pdf document.pdf --resolution 150 output.odt
4. Full Options Reference
usage: pdf2odt [-h] [--version] --pdf PDF [--resolution RESOLUTION] [--ocr] output
Converts a pdf to a LibreOffice Writer document with pages as images
positional arguments:
output Output odt file
options:
-h, --help show this help message and exit
--version show program's version number and exit
--pdf PDF PDF file to convert
--resolution RESOLUTION
Sets DPI image resolution. Default is 300
--ocr Extracts text with page.get_text() or OCR and inserts result after image in ODT document
Python API Usage
You can also use pdf2odt directly in your Python applications:
from pdf2odt.core import main_command
# Convert a PDF to ODT with 300 DPI and OCR enabled
main_command(
pdf="path/to/document.pdf",
resolution=300,
ocr=True,
output="path/to/output.odt"
)
Dependencies
pdf2odt relies on the following Python packages:
- PyMuPDF: High-performance PDF rendering and native text extraction.
- odfdo: Pure Python OpenDocument (.odt) document generator.
- rapidocr-onnxruntime: Lightweight ONNX-powered OCR engine.
- Pillow: Image dimension and format processing.
- tqdm: Console progress bar.
- colorama: Colored terminal output.
Development & Testing
This project uses Poetry and Poe the Poet for development tasks.
Run Tests
poetry run poe test
Run Coverage Report
poetry run poe coverage
Update Translations
poetry run poe translate
License
Distributed under the GPL-3.0 License.
Metadata
Release files for pdf2odt 1.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| pdf2odt-1.1.0.tar.gz | 19.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| pdf2odt-1.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 44.2 kB
Release files / pdf2odt-1.1.0.tar.gz
| Download URL | pdf2odt-1.1.0.tar.gz |
|---|---|
| Size | 19.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
00e604caa60371eb918f1968d7f7bcb83567e613d851c6c5b4e4b515600912ad
|
|
BLAKE2b-256 checksum How to use checksums |
6f9e6f80a184e7d400dfbac8058b397c2831f8ef828a46131411e2132f50f980
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
poetry/2.4.2 CPython/3.14.7 Linux/7.2.6-gentoo
|
Release files / pdf2odt-1.1.0-py3-none-any.whl
| Download URL | pdf2odt-1.1.0-py3-none-any.whl |
|---|---|
| Size | 24.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
2f824b27774a668d1ed8abb3bc4619670a2347fe088f38046eef60e8180cf9c9
|
|
BLAKE2b-256 checksum How to use checksums |
a3308e35525aee478543e5b6c29594a1bc422cdf00287bddbaee254abe4ebd22
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
poetry/2.4.2 CPython/3.14.7 Linux/7.2.6-gentoo
|