Skip to main content

PDF Parser Light Logo

PDF Parser Light

GitHub Release PyPI Package License: MIT Python 3.9+

A lightweight desktop and CLI tool that converts PDFs into clean Markdown using the Google Gemini API. Formats math into LaTeX equations and preserves complex tables without requiring a paid subscription or third-party service.

PDF Parser Light Demo


Why PDF Parser Light?

Most PDF extraction tools fall into two categories:

  1. Traditional OCR & PDF libraries (like Tesseract, PyPDF, or pdfplumber): Fast and local, but often mangle multi-column layouts, tables, and mathematical formulas into broken plain text.
  2. Commercial OCR / Document AI APIs (like Mathpix or AWS Textract): High accuracy, but lock you into monthly subscriptions, proprietary dashboards, or per-page billing.

PDF Parser Light takes a different approach:

  • Direct API access: Calls Gemini's vision models directly using your own API key. No middlemen, no tracking, and no added subscription fees.
  • Works within Gemini's free tier: Handles daily document conversion within Google's free quota (20 requests/day on Flash, falling back to Flash Lite for higher volume).
  • Multimodal extraction: Gemini understands full page layouts in one shot, converting formulas into clean inline/block LaTeX ($...$ / $$...$$) and multi-column data into Markdown/HTML tables.
  • Both GUI & CLI: Use the drag-and-drop desktop app or automate large batches via terminal.

Features

  • LaTeX math & table preservation: Extracts formulas into LaTeX and structured data into Markdown/HTML tables without summarizing.
  • Desktop GUI & CLI: CustomTkinter desktop interface with live progress, plus a CLI for batch processing directories.
  • Automatic page chunking: Automatically splits documents larger than 20 pages into chunks to prevent timeouts and context overflows.
  • Quota tracking & model fallback: Tracks daily free tier usage locally and automatically falls back to Lite models if primary quotas run low.
  • Resume support: Resumes processing from partial output files if a job is interrupted mid-way.
  • Cross-platform: Pre-built binaries available for macOS, Windows, and Linux.

Installation

Option 1: Standalone App (No Python required)

Download the pre-built binary for your OS from the Releases page:

  • macOS: Download pdf-parser-light-macos.zip, extract, and move PDF Parser Light.app to /Applications.
  • Windows: Download pdf-parser-light-windows.zip, extract, and run PDF Parser Light.exe.
  • Linux: Download pdf-parser-light-linux.tar.gz, extract, and run PDF Parser Light. (Requires Tk: sudo apt install python3-tk).

Opening Unsigned Binaries (First-Time Setup)

Because these builds are open-source releases built via GitHub Actions without commercial developer certificates, your OS will block them by default on first launch:

  • macOS (Gatekeeper):
    • Finder / System Settings: Right-click (or Control-click) PDF Parser Light.app in Finder, click Open, and confirm Open. On macOS Sequoia (15+), if it doesn't open, go to System Settings → Privacy & Security, scroll down to the Security section, and click Open Anyway.
    • Terminal: Alternatively, remove the quarantine attribute directly:
      xattr -d com.apple.quarantine "/Applications/PDF Parser Light.app"
      
  • Windows (SmartScreen):
    • Click More info on the popup, then click Run anyway.

Option 2: Install via pip

pip install pdf-parser-light

Launch the GUI or CLI:

pdf-parser-light-gui   # Launches Desktop GUI
pdf-parser-light       # Runs CLI

Usage

Desktop Application

PDF Parser Light Tutorial

  1. Launch the app and enter your Gemini API Key ([https://aistudio.google.com/api-keys]) (check Remember API Key to save locally).
  2. Drag and drop a PDF file (or click Browse).
  3. Click Process.
  4. Copy the result or save it directly as .md or .txt.

Command Line (CLI)

Set your Gemini API key:

export GEMINI_API_KEY="your_api_key_here"

Examples

# Convert a single PDF
pdf-parser-light document.pdf --output output.md

# Batch process an entire directory of PDFs
pdf-parser-light ./pdf_folder/ --output ./markdown_output/

# Process a specific page range (e.g. pages 1 to 50)
pdf-parser-light document.pdf --pages 1-50

# Resume an interrupted parsing job
pdf-parser-light document.pdf --resume

# Check remaining free requests for today
pdf-parser-light --usage

# Reset local quota counter
pdf-parser-light --reset-quota

# Custom transcription prompt
pdf-parser-light document.pdf --prompt "Transcribe equations only into LaTeX."

Options

Option Description
input_path Path to a .pdf file or a directory containing PDFs.
--output, -o Output file (single PDF) or output directory (batch mode).
--pages, -p Page range to parse (e.g. 1-50, 40-120, or 10).
--resume, -r Resume parsing from existing partial output.
--prompt Custom instructions for the model.
--api-key Pass API key explicitly (overrides GEMINI_API_KEY).
--usage Print remaining daily free requests and exit.
--reset-quota Reset local request counter to 0.
--force Bypass local quota check and run anyway.

Daily Quota & Rate Limits

  • Free Tier Limits: By default, the app tracks 20 free requests/day locally for primary Flash models (configurable via GEMINI_FREE_LIMIT).
  • Model Fallbacks: When your primary quota runs out, requests cascade to Flash Lite models (gemini-3.5-flash-lite, gemini-3.1-flash-lite), which offer significantly higher capacity (up to 500 free requests/day per model on Google AI Studio).
  • Retries: Automatically retries with exponential backoff on HTTP 429 rate limits or transient errors.

Privacy & Security

  • Direct Connections: All requests and file uploads go directly to Google's Gemini API endpoints using the official google-genai SDK. No third-party servers, analytics, or telemetry proxies are involved.
  • Data Sensitivity Notice: Do not upload confidential, sensitive, or restricted documents. Files are processed remotely on Google's servers in accordance with Google Gemini API terms.
  • File Retention & Cleanup: Files are uploaded temporarily to Gemini API storage for transcription and deleted immediately upon completion. If a process terminates unexpectedly, files remain until Google's temporary storage retention limit expires.
  • Key Storage: Saved API keys are stored locally on your machine in standard user configuration directories (~/Library/Application Support/pdf_parser_light/ on macOS, %LOCALAPPDATA%\pdf_parser_light\ on Windows, ~/.config/pdf_parser_light/ on Linux).

Development

git clone https://github.com/jtaroreh/pdf-parser-light.git
cd pdf-parser-light

python3 -m venv venv
source venv/bin/activate
pip install -e ".[test]"

pytest -v

Build App Bundle (macOS)

./build.sh

Output is created in dist/PDF Parser Light.app.

Metadata

Release files for pdf-parser-light 0.1.4

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for pdf-parser-light 0.1.4
File Size Uploaded
pdf_parser_light-0.1.4.tar.gz 2.5 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for pdf-parser-light 0.1.4
File Interpreter ABI Platform
pdf_parser_light-0.1.4-py3-none-any.whl Python 3 none any Details

Total release size: 2.5 MB

Release files / pdf_parser_light-0.1.4.tar.gz

Download URL pdf_parser_light-0.1.4.tar.gz
Size 2.5 MB
Tags Source
SHA-256 checksum
How to use checksums
328a757253f9c889be3c84a66c8012b9abc17e508cc4afe83dca4fc5562aa269
BLAKE2b-256 checksum
How to use checksums
2694b3ac83b5086d0efbae8ad7dced505d25938b4cc35e4d8a6d7696a280ee3f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 9, 2026.

Transparency log

Release files / pdf_parser_light-0.1.4-py3-none-any.whl

Download URL pdf_parser_light-0.1.4-py3-none-any.whl
Size 26.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
8ac15af7cdc95c72fa43807d88c2ad8b2630b63c292eadc1f85dc16938e734e7
BLAKE2b-256 checksum
How to use checksums
a3b4b41cde9ca37477177de82e77470274aec96f5882c7e9425990be9f661c33
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 9, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.4 This release

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page