Skip to main content

Legacy Font PDF Translator - Translate PDFs with legacy Indian font encodings to English

Project description

LegacyLipi

Legacy Font PDF Translator - Translate PDF documents with legacy Indian font encodings to English.

Installation

From PyPI (Recommended)

pip install legacylipi

Or with uv:

uv tool install legacylipi

From Source

git clone https://github.com/biswasbiplob/legacylipi.git
cd legacylipi
uv sync

Frontend (for development)

cd frontend
npm install

Usage

# CLI translation
legacylipi translate input.pdf -o output.txt

# Launch React web UI (production build served by FastAPI)
legacylipi api

# Launch legacy NiceGUI web UI (deprecated)
legacylipi ui

Problem

Millions of government documents, legal papers, and archival materials in Indian regional languages (Marathi, Hindi, Tamil, etc.) were created using legacy font encoding systems (Shree-Lipi, Kruti Dev, APS, Chanakya, etc.). These fonts map Devanagari/regional script glyphs to ASCII/Latin code points, making them unreadable by standard translation tools.

Example:

  • What the PDF displays: महाराष्ट्र राजभाषा अधिनियम
  • What text extraction produces: ´ÖÆüÖ¸üÖ™Òü ¸üÖ•Ö³ÖÖÂÖÖ †×¬Ö×®ÖμÖ´Ö
  • What Google Translate sees: Gibberish

Solution

LegacyLipi:

  1. Detects the font encoding scheme used in a PDF (legacy or Unicode)
  2. Converts legacy-encoded text to proper Unicode
  3. Alternatively, uses OCR to extract text from scanned PDFs
  4. Translates the Unicode text to the target language
  5. Outputs translated text in various formats (text, markdown, PDF)

Installation

# Clone and install
git clone https://github.com/biswasbiplob/legacylipi.git
cd legacylipi
uv sync

# With all optional backends
uv sync --all-extras

OCR Support (Optional)

LegacyLipi supports multiple OCR backends:

Backend Description GPU Support
Tesseract Local, free, most language packs CPU only
Google Vision Cloud, paid, best accuracy N/A
EasyOCR Local, free, good for Indian languages CUDA, MPS (Apple Silicon)

Tesseract (default):

# Ubuntu/Debian
sudo apt-get install tesseract-ocr tesseract-ocr-mar tesseract-ocr-hin

# macOS
brew install tesseract tesseract-lang

EasyOCR with GPU (optional):

# Install with EasyOCR support
uv sync --extra easyocr

# For GPU acceleration, install PyTorch with CUDA or MPS support

Google Vision (optional):

uv sync --extra vision
# Requires GCP credentials (GOOGLE_APPLICATION_CREDENTIALS)

See docs/cli-reference.md for detailed OCR options and language codes.

Quick Start

# Basic translation
uv run legacylipi translate input.pdf -o output.txt

# Output as PDF (preserves layout)
uv run legacylipi translate input.pdf -o output.pdf --format pdf

# OCR for scanned documents
uv run legacylipi translate input.pdf --use-ocr -o output.txt

# Use local LLM (requires Ollama)
uv run legacylipi translate input.pdf --translator ollama --model llama3.2

# Detect encoding only
uv run legacylipi detect input.pdf

See docs/cli-reference.md for complete CLI documentation.

Web UI

LegacyLipi includes a modern React-based web interface backed by a FastAPI REST API.

Production (single command)

# Serves the built React frontend + API on one port
uv run legacylipi api
# or
uv run legacylipi-web

Open http://localhost:8000 in your browser.

Development (hot-reload)

# Start both FastAPI backend and Vite dev server
./scripts/dev.sh

This runs:

  • Backend at http://localhost:8000 (FastAPI with auto-reload)
  • Frontend at http://localhost:5173 (Vite dev server with HMR, proxies /api to backend)

Legacy NiceGUI UI (deprecated)

The original NiceGUI-based UI is still available but deprecated:

uv run legacylipi ui
# Open http://localhost:8080

Workflow Modes:

  • Scanned Copy - Create image-based PDF copy (adjust DPI, color, quality)
  • Convert to Unicode - OCR + Unicode conversion without translation
  • Full Translation - Complete pipeline with OCR, conversion, and translation

Features:

  • Drag-and-drop PDF upload
  • Workflow-based UI with mode selection
  • Multiple translation backends (Translate-Shell, Google, Ollama, OpenAI, etc.)
  • OCR support with engine and language selection
  • Structure-preserving or flowing text modes
  • Real-time SSE progress streaming
  • Direct download of translated files
  • Responsive dark-theme design

Translation Backends

Backend Description Setup
trans translate-shell CLI (recommended) brew install translate-shell
google Google Translate (free API) Works out of the box
mymemory MyMemory API (free) Works out of the box
ollama Local LLM via Ollama Ollama required
openai OpenAI GPT models Set OPENAI_API_KEY
gcp_cloud Google Cloud Translation GCP project + credentials

See docs/translation-backends.md for detailed setup guides.

Supported Encodings

Encoding Font Family Language Status
shree-lipi Shree-Lipi, Shree-Dev-0714 Marathi ✅ Built-in
kruti-dev Kruti Dev Hindi ✅ Built-in
aps-dv APS-DV Hindi 🔄 Detection only
chanakya Chanakya Hindi 🔄 Detection only
dvb-tt DVB-TT, DV-TTYogesh Hindi 🔄 Detection only
walkman-chanakya Walkman Chanakya Hindi 🔄 Detection only
shusha Shusha Hindi 🔄 Detection only

CLI Commands

Command Description
api Launch the React web UI + FastAPI REST API
translate Full pipeline: parse → detect → convert → translate → output
convert Convert legacy encoding to Unicode (no translation)
extract Extract text from PDF (OCR or font-based)
detect Analyze PDF and report detected encoding
scan-copy Create an image-based scanned copy of a PDF
encodings List supported font encodings
usage Show API usage statistics
ui Launch legacy NiceGUI web interface (deprecated)

See docs/cli-reference.md for full command reference.

Development

See docs/development.md for setup instructions, running tests, project structure, and adding new encodings.

Architecture

┌─────────────────────────────────────────────────────────────────────────┐
│                              LegacyLipi                                 │
├─────────────────────┬───────────────────────────────────────────────────┤
│   React Frontend    │                  FastAPI Backend                   │
│  (Vite + TS + TW)   │                                                   │
│                     │   ┌──────────────────────────────────────────┐    │
│  FileUploader       │   │              REST API                    │    │
│  WorkflowSelector   │   │  /api/v1/config/*     GET config         │    │
│  Settings panels    │◄─▶│  /api/v1/sessions/*   Upload/delete      │    │
│  StatusPanel (SSE)  │   │  /api/v1/sessions/*/  Start pipeline     │    │
│  DownloadButton     │   │  /api/v1/sessions/*/progress  SSE stream │    │
│                     │   │  /api/v1/sessions/*/download  Get result │    │
│                     │   └────────────────────┬─────────────────────┘    │
├─────────────────────┘                        │                          │
│                                              ▼                          │
│  ┌──────────────────────────────────────────────────────────────────┐   │
│  │                      Core Pipeline                               │   │
│  │                                                                  │   │
│  │  PDF Parser / OCR Parser                                         │   │
│  │       │                                                          │   │
│  │  Encoding Detector → Unicode Converter                           │   │
│  │       │                                                          │   │
│  │  Translation Engine (trans, Google, Ollama, OpenAI, GCP, ...)    │   │
│  │       │                                                          │   │
│  │  Output Generator (.txt, .md, .pdf)                              │   │
│  └──────────────────────────────────────────────────────────────────┘   │
└─────────────────────────────────────────────────────────────────────────┘

Pipeline Flow:

  1. Parse PDF → Extract text with PDF parser or OCR
  2. Detect Encoding → Identify legacy encoding scheme
  3. Convert to Unicode → Transform legacy text to Unicode
  4. Translate → Use translation backend
  5. Generate Output → Create PDF/text/markdown

License

MIT

Contributing

Contributions are welcome! Please:

  1. Fork the repository
  2. Create a feature branch (git checkout -b feature/amazing-feature)
  3. Make your changes
  4. Run tests (uv run pytest)
  5. Commit and push
  6. Open a Pull Request

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

legacylipi-0.8.0.tar.gz (705.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

legacylipi-0.8.0-py3-none-any.whl (99.6 kB view details)

Uploaded Python 3

File details

Details for the file legacylipi-0.8.0.tar.gz.

File metadata

  • Download URL: legacylipi-0.8.0.tar.gz
  • Upload date:
  • Size: 705.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.7

File hashes

Hashes for legacylipi-0.8.0.tar.gz
Algorithm Hash digest
SHA256 502e1cde15c7f3c81a2c9077050a8ec444cc020fe75f982d4991ea7de818416b
MD5 3f6857cad83645156e2d0424bb24b8e7
BLAKE2b-256 15b8231b4b871b4d6e737cee09dbf1ef4a03b59346846be6ca6cc521fa12b955

See more details on using hashes here.

File details

Details for the file legacylipi-0.8.0-py3-none-any.whl.

File metadata

  • Download URL: legacylipi-0.8.0-py3-none-any.whl
  • Upload date:
  • Size: 99.6 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.7

File hashes

Hashes for legacylipi-0.8.0-py3-none-any.whl
Algorithm Hash digest
SHA256 5181e9ebfeee6646dfd6fef0ebfe4c5f82ef4da77b240e97439d587a66b297bb
MD5 35bdc31c89f6e208a573150fac964560
BLAKE2b-256 9615c1476cf1cb78858786be924ab784c977e9b30240c3d1c558456e33c3896e

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page