Skip to main content

LiteParse Python

Python bindings for LiteParse — fast, lightweight PDF and document parsing with spatial text extraction.

Installation

pip install liteparse

This also installs the lit CLI command.

Quick Start

from liteparse import LiteParse

parser = LiteParse()
result = parser.parse("document.pdf")
print(result.text)
print(f"Source document pages: {result.total_pages}")

# Access structured data
for page in result.pages:
    print(f"Page {page.page_num}: {len(page.text_items)} text items")

Markdown Output

LiteParse can render documents directly to Markdown including headings, tables, lists, images, and links reconstructed from the spatial layout. Great for feeding LLMs and RAG pipelines. The rendered Markdown is returned on result.text:

parser = LiteParse(
    output_format="markdown",   # "json" | "text" | "markdown"
    image_mode="placeholder",   # "placeholder" | "off" | "embed"
    extract_links=True,         # render [text](url) link syntax (default: True)
)
result = parser.parse("document.pdf")
print(result.text)  # rendered Markdown

Reconstruction quality varies with document complexity.

Configuration

All options are passed to the constructor:

parser = LiteParse(
    ocr_enabled=True,              # Enable OCR (default: True)
    ocr_language="eng",            # Tesseract language code
    ocr_server_url=None,           # HTTP OCR server URL (optional)
    tessdata_path=None,            # Path to tessdata directory (optional)
    max_pages=1000,                # Max pages to parse
    target_pages="1-5,10",         # Specific pages (optional)
    extract_screenshots=False,      # Return parsed pages as PNG bytes
    continue_on_page_error=False,   # Skip broken pages and return page_errors
    dpi=150,                       # Rendering DPI
    output_format="json",          # "json" | "text" | "markdown"
    image_mode="placeholder",      # Markdown image handling: "placeholder" | "off" | "embed"
    extract_images=True,           # Extract image bytes + metadata (default: False)
    image_output_dir="./images",   # Write images and return name/path metadata (optional)
    extract_links=True,            # Render [text](url) links in markdown output
    keep_headers_footers=False,    # Keep running headers/footers in markdown output
    extract_vector_graphics=False, # Opt-in shapes + merged H/V lines per page
    extract_annotations=False,     # Include page annotations in structured output
    extract_form_fields=False,      # Include AcroForm widget fields and values
    extract_structure_tree=False,   # Include tagged-PDF logical structure
    preserve_very_small_text=False, # Keep tiny text
    extract_text_metadata=False,    # Opt in to MCID, font metrics, colors, char codes, and trailing_space_generated
    password=None,                 # Password for protected documents
    quiet=False,                   # Suppress progress output
    num_workers=4,                 # Concurrent OCR workers
)

When extract_images=True, image extraction is enabled. image_output_dir requires that explicit opt-in and writes the extracted bytes to disk. Each result.images entry includes its page bbox, intrinsic pixel dimensions, rotation, format, name, and path. Valid source JPEGs are preserved, exact duplicates reuse one file, and JSON CLI output contains metadata only (no base64 image data). image_mode controls Markdown presentation only and does not imply extraction. With extract_images=False, lightweight Markdown placement refs are still collected and result.images stays empty.

When extract_annotations is enabled, each parsed page has an annotations list containing the subtype, contents, author/title, PDF date strings, viewport-space rectangle and quadpoint rectangles, and URI for external link annotations. It is independent of extract_links, which controls Markdown link rendering. The field is None when extraction is disabled.

When extract_structure_tree=True, each page has a structure_tree containing all tagged-PDF roots and recursive elements with type, ID, actual/alternate text, title, typed attributes, MCIDs, children, and referenced link annotations. Untagged pages have an empty roots list; the field is None when disabled.

Every result also carries creator/producer from the PDF /Info dictionary. With extract_document_metadata=True, result.doc_meta adds a provenance object with dates, PDF version/security, signature state, incremental-save markers, trailer ID comparison, the catalog's XMP packet (capped at 64 KiB; skipped for sources over 16 MiB), and source size. It is off by default because it streams the whole source file, and it is None for inputs converted from a non-PDF format. These document fields are API-only and do not alter default CLI JSON.

Parsing from Bytes

Pass raw PDF bytes directly — useful for web uploads or downloaded files:

with open("document.pdf", "rb") as f:
    result = parser.parse(f.read())
print(result.text)

Worker Pool and Hard Timeouts

PDFium is not thread-safe, so in-process parses serialize on a process-global lock. For high-throughput services, or cases where you need to enforce a timeout, run parses in a pool of persistent worker processes instead:

from liteparse import LiteParse, ParseTimeoutError

parser = LiteParse(pool_size=4, parse_timeout=15)
parser.warm_up()  # optional: pre-initialize workers (~60ms total)

try:
    result = parser.parse("document.pdf")
except ParseTimeoutError as e:
    print(f"killed rogue document: {e.source} (deadline {e.timeout}s)")

parser.close()  # or use `with LiteParse(...) as parser:`

Screenshots

Generate PNG screenshots of document pages:

screenshots = parser.screenshot("document.pdf", page_numbers=[1, 2, 3])
for s in screenshots:
    print(f"Page {s.page_num}: {s.width}x{s.height}")
    with open(f"page_{s.page_num}.png", "wb") as f:
        f.write(s.image_bytes)

Document Complexity

Before committing to a full parse, check whether a document needs OCR or heavier processing. is_complex is a cheap, text-layer-only pass that returns one entry per page with a needs_ocr verdict and the signals behind it — useful for routing documents to different pipelines, rejecting ones you can't handle, or estimating cost.

parser = LiteParse()
pages = parser.is_complex("document.pdf")

if any(p.needs_ocr for p in pages):
    # Route to the OCR-enabled pipeline
    result = parser.parse("document.pdf")
else:
    # Cheap path — skip OCR entirely
    result = LiteParse(ocr_enabled=False).parse("document.pdf")

# Inspect why specific pages were flagged
for page in pages:
    if page.needs_ocr:
        print(f"Page {page.page_number}: {', '.join(page.reasons)}")

reasons is one of "scanned", "no-text", "sparse-text", "embedded-images", "garbled", "vector-text", or "annotation-text". Raw bytes work here too.

Supported Formats

  • PDF (.pdf)
  • Microsoft Office (.docx, .xlsx, .pptx, etc.) — requires LibreOffice
  • OpenDocument (.odt, .ods, .odp) — requires LibreOffice
  • Images (.png, .jpg, .tiff, etc.)
  • And more!

CLI

The Python package includes the lit CLI:

lit parse document.pdf
lit parse document.pdf --format json -o output.json
lit parse document.pdf --format json --extract-annotations
lit parse document.pdf --format json --extract-form-fields
lit screenshot document.pdf -o ./screenshots
lit batch-parse ./input ./output
lit is-complex document.pdf

Metadata

Release files for liteparse 2.15.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for liteparse 2.15.1
File Size Uploaded
liteparse-2.15.1.tar.gz 564.3 kB Details

Built distributions (wheels)

Table of built distributions (wheels) for liteparse 2.15.1
File
liteparse-2.15.1-cp310-abi3-win_arm64.whl CPython 3.10 abi3 Windows ARM64 Details
liteparse-2.15.1-cp310-abi3-win_amd64.whl CPython 3.10 abi3 Windows x86-64 Details
liteparse-2.15.1-cp310-abi3-musllinux_1_2_x86_64.whl CPython 3.10 abi3 Linux musl 1.2+ x86-64 Details
liteparse-2.15.1-cp310-abi3-manylinux_2_28_x86_64.whl CPython 3.10 abi3 Linux glibc 2.28+ x86-64 Details
liteparse-2.15.1-cp310-abi3-manylinux_2_28_aarch64.whl CPython 3.10 abi3 Linux glibc 2.28+ ARM64 Details
liteparse-2.15.1-cp310-abi3-macosx_11_0_arm64.whl CPython 3.10 abi3 macOS 11.0+ ARM64 Details
liteparse-2.15.1-cp310-abi3-macosx_10_12_x86_64.whl CPython 3.10 abi3 macOS 10.12+ x86-64 Details

Total release size: 93.8 MB

Release files / liteparse-2.15.1.tar.gz

Download URL liteparse-2.15.1.tar.gz
Size 564.3 kB
Tags Source
SHA-256 checksum
How to use checksums
4a2dab9d8d89a517f989d0282f53a608f983c0a8c10e1b24c759b9e6ea6e4144
BLAKE2b-256 checksum
How to use checksums
eec7b840a00a3b4f34087db00dd360ec309188adc6f7dbb5c68c334da4ad3ad4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / liteparse-2.15.1-cp310-abi3-win_arm64.whl

Download URL liteparse-2.15.1-cp310-abi3-win_arm64.whl
Size 11.5 MB
Tags CPython 3.10 Windows ARM64 abi3
SHA-256 checksum
How to use checksums
2ba9ee3f104f96eb5ef4ddb7ed5fdc8fee2125846753d9a82e79b56f55d9790d
BLAKE2b-256 checksum
How to use checksums
dcc658b6698cb2f8b2f175a6d4ad629a5f4f8858f26ccea2eececdd99ddc27a8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / liteparse-2.15.1-cp310-abi3-win_amd64.whl

Download URL liteparse-2.15.1-cp310-abi3-win_amd64.whl
Size 12.1 MB
Tags CPython 3.10 Windows x86-64 abi3
SHA-256 checksum
How to use checksums
698a332ec1b6f932f1f065df11fcb52bd1f95acaec7056333070501173ba8cac
BLAKE2b-256 checksum
How to use checksums
d7ad3bc535c88fe9b613c3481aa1ff1edcc7a30a5b8766d1d8f30674f6f803ab
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / liteparse-2.15.1-cp310-abi3-musllinux_1_2_x86_64.whl

Download URL liteparse-2.15.1-cp310-abi3-musllinux_1_2_x86_64.whl
Size 17.1 MB
Tags CPython 3.10 Linux musl 1.2+ x86-64 abi3
SHA-256 checksum
How to use checksums
78b7b660c666093b1a0c69fc9f66da6394ec5c2f4d3486411b00faa262a24a41
BLAKE2b-256 checksum
How to use checksums
e02b28a39f20166dc942bcb7865817061410da9bc0582ab864678d23c446b18f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / liteparse-2.15.1-cp310-abi3-manylinux_2_28_x86_64.whl

Download URL liteparse-2.15.1-cp310-abi3-manylinux_2_28_x86_64.whl
Size 14.0 MB
Tags CPython 3.10 Linux glibc 2.28+ x86-64 abi3
SHA-256 checksum
How to use checksums
c7d3f3588241aa654f5f75b6781b0cc5c31e26d88bf9dba1b7c8cd23931d829c
BLAKE2b-256 checksum
How to use checksums
392c116959772cc8b08da025295169edceb20e6a7cf61979c86718dca5c7247a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / liteparse-2.15.1-cp310-abi3-manylinux_2_28_aarch64.whl

Download URL liteparse-2.15.1-cp310-abi3-manylinux_2_28_aarch64.whl
Size 13.7 MB
Tags CPython 3.10 Linux glibc 2.28+ ARM64 abi3
SHA-256 checksum
How to use checksums
ad27a27a00d7658d9ee5f6aa117f3ce4762d9b09fb68376352e117d0540f17c4
BLAKE2b-256 checksum
How to use checksums
171610fa7dfb1315fa873790f85f9bfae0517a97797adbe9d19699742f14d313
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / liteparse-2.15.1-cp310-abi3-macosx_11_0_arm64.whl

Download URL liteparse-2.15.1-cp310-abi3-macosx_11_0_arm64.whl
Size 12.1 MB
Tags CPython 3.10 abi3 macOS 11.0+ ARM64
SHA-256 checksum
How to use checksums
48ed7426d995d40a4834130d2cafe96d21281cbe75a943222506f66dfdbb884c
BLAKE2b-256 checksum
How to use checksums
72f1146b61a8543a5bb80e1178278d8269070c2ec0fbea4270575b6f9ec6e361
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / liteparse-2.15.1-cp310-abi3-macosx_10_12_x86_64.whl

Download URL liteparse-2.15.1-cp310-abi3-macosx_10_12_x86_64.whl
Size 12.7 MB
Tags CPython 3.10 abi3 macOS 10.12+ x86-64
SHA-256 checksum
How to use checksums
77d3ab1b1daf4d4f67c971b69309b4a7fcca940c28633e37121727981a31e571
BLAKE2b-256 checksum
How to use checksums
3aac198f1ee575afa24574ee1d99e349e8f078258f36bfb874229358e52cd06b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release history Release notifications | RSS feed

This release

2.15.1 This release

8 release files

2.15.0

8 release files

2.14.7

8 release files

2.14.6

8 release files

2.14.5

8 release files

2.14.2

8 release files

2.14.1

8 release files

2.14.0

8 release files

2.9.0

36 release files

2.8.1

36 release files

2.8.0

36 release files

2.7.0

36 release files

2.6.0

36 release files

2.5.1

36 release files

2.4.0

36 release files

2.3.0

36 release files

2.2.1

36 release files

2.2.0

36 release files

2.1.1

36 release files

2.1.0

36 release files

2.0.8

32 release files

2.0.4

24 release files

2.0.3

24 release files

2.0.1

24 release files

2.0.0

24 release files

1.2.1

2 release files

1.2.0

2 release files

1.1.0

2 release files

1.0.1

2 release files

1.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page