Skip to main content

Legacy Font PDF Translator - Translate PDFs with legacy Indian font encodings to English

Project description

LegacyLipi

Legacy Font PDF Translator - Translate PDF documents with legacy Indian font encodings to English.

Installation

From PyPI (Recommended)

pip install legacylipi

Or with uv:

uv tool install legacylipi

From Source

git clone https://github.com/biswasbiplob/legacylipi.git
cd legacylipi
uv sync

Usage

# CLI translation
legacylipi translate input.pdf -o output.txt

# Launch web UI
legacylipi ui

# Launch UI on custom port
legacylipi ui --port 3000

Problem

Millions of government documents, legal papers, and archival materials in Indian regional languages (Marathi, Hindi, Tamil, etc.) were created using legacy font encoding systems (Shree-Lipi, Kruti Dev, APS, Chanakya, etc.). These fonts map Devanagari/regional script glyphs to ASCII/Latin code points, making them unreadable by standard translation tools.

Example:

  • What the PDF displays: महाराष्ट्र राजभाषा अधिनियम
  • What text extraction produces: ´ÖÆüÖ¸üÖ™Òü ¸üÖ•Ö³ÖÖÂÖÖ †×¬Ö×®ÖμÖ´Ö
  • What Google Translate sees: Gibberish

Solution

LegacyLipi:

  1. Detects the font encoding scheme used in a PDF (legacy or Unicode)
  2. Converts legacy-encoded text to proper Unicode
  3. Alternatively, uses OCR to extract text from scanned PDFs
  4. Translates the Unicode text to the target language
  5. Outputs translated text in various formats (text, markdown, PDF)

Installation

# Clone and install
git clone https://github.com/biswasbiplob/legacylipi.git
cd legacylipi
uv sync

# With all optional backends
uv sync --all-extras

OCR Support (Optional)

LegacyLipi supports multiple OCR backends:

Backend Description GPU Support
Tesseract Local, free, most language packs CPU only
Google Vision Cloud, paid, best accuracy N/A
EasyOCR Local, free, good for Indian languages CUDA, MPS (Apple Silicon)

Tesseract (default):

# Ubuntu/Debian
sudo apt-get install tesseract-ocr tesseract-ocr-mar tesseract-ocr-hin

# macOS
brew install tesseract tesseract-lang

EasyOCR with GPU (optional):

# Install with EasyOCR support
uv sync --extra easyocr

# For GPU acceleration, install PyTorch with CUDA or MPS support

Google Vision (optional):

uv sync --extra vision
# Requires GCP credentials (GOOGLE_APPLICATION_CREDENTIALS)

See docs/cli-reference.md for detailed OCR options and language codes.

Quick Start

# Basic translation
uv run legacylipi translate input.pdf -o output.txt

# Output as PDF (preserves layout)
uv run legacylipi translate input.pdf -o output.pdf --format pdf

# OCR for scanned documents
uv run legacylipi translate input.pdf --use-ocr -o output.txt

# Use local LLM (requires Ollama)
uv run legacylipi translate input.pdf --translator ollama --model llama3.2

# Detect encoding only
uv run legacylipi detect input.pdf

See docs/cli-reference.md for complete CLI documentation.

Web UI

LegacyLipi includes a web interface for easy PDF translation without command-line usage.

uv run legacylipi-ui

Open http://localhost:8080 in your browser.

LegacyLipi Web UI

Features:

  • Drag-and-drop PDF upload
  • Multiple translation backends
  • OCR support with language selection
  • Structure-preserving or flowing text modes
  • Real-time progress tracking
  • Direct download of translated files

Translation Backends

Backend Description Setup
trans translate-shell CLI (recommended) brew install translate-shell
google Google Translate (free API) Works out of the box
mymemory MyMemory API (free) Works out of the box
ollama Local LLM via Ollama Ollama required
openai OpenAI GPT models Set OPENAI_API_KEY
gcp_cloud Google Cloud Translation GCP project + credentials

See docs/translation-backends.md for detailed setup guides.

Supported Encodings

Encoding Font Family Language Status
shree-lipi Shree-Lipi, Shree-Dev-0714 Marathi ✅ Built-in
kruti-dev Kruti Dev Hindi ✅ Built-in
aps-dv APS-DV Hindi 🔄 Detection only
chanakya Chanakya Hindi 🔄 Detection only
dvb-tt DVB-TT, DV-TTYogesh Hindi 🔄 Detection only
walkman-chanakya Walkman Chanakya Hindi 🔄 Detection only
shusha Shusha Hindi 🔄 Detection only

CLI Commands

Command Description
translate Full pipeline: parse → detect → convert → translate → output
convert Convert legacy encoding to Unicode (no translation)
extract Extract text from PDF (OCR or font-based)
detect Analyze PDF and report detected encoding
encodings List supported font encodings
usage Show API usage statistics

See docs/cli-reference.md for full command reference.

Development

See docs/development.md for setup instructions, running tests, project structure, and adding new encodings.

Architecture

┌─────────────────────────────────────────────────────────────────────────┐
│                              LegacyLipi                                 │
├─────────────────────────────────────────────────────────────────────────┤
│                                                                         │
│  ┌──────────────────────────────────────────────────────────────────┐   │
│  │                      Text Extraction                             │   │
│  │  ┌──────────────┐              ┌──────────────┐                  │   │
│  │  │   PDF        │    OR        │   OCR        │                  │   │
│  │  │   Parser     │              │   Parser     │                  │   │
│  │  │ (font-based) │              │ (Tesseract)  │                  │   │
│  │  └──────────────┘              └──────────────┘                  │   │
│  └──────────────────────────────────────────────────────────────────┘   │
│         │                                │                              │
│         ▼                                ▼                              │
│  ┌──────────────┐    ┌──────────────┐                                   │
│  │   Encoding   │───▶│   Unicode    │◀──── (OCR output is               │
│  │   Detector   │    │   Converter  │       already Unicode)            │
│  └──────────────┘    └──────────────┘                                   │
│                             │                                           │
│                             ▼                                           │
│         ┌───────────────────────────────────────────────────────────┐   │
│         │                 Translation Engine                        │   │
│         │  ┌────────┬────────┬──────────┬────────┬────────┬─────┐   │   │
│         │  │ trans  │ Google │ MyMemory │ Ollama │ OpenAI │ GCP │   │   │
│         │  │ (CLI)  │ Trans. │  (API)   │(Local) │ (API)  │Cloud│   │   │
│         │  └────────┴────────┴──────────┴────────┴────────┴─────┘   │   │
│         └───────────────────────────────────────────────────────────┘   │
│                             │                                           │
│                             ▼                                           │
│         ┌───────────────────────────────────────────────┐               │
│         │            Output Generator                   │               │
│         │  ┌──────┬────────┬───────┐                    │               │
│         │  │ .txt │  .md   │ .pdf  │                    │               │
│         │  └──────┴────────┴───────┘                    │               │
│         └───────────────────────────────────────────────┘               │
│                                                                         │
└─────────────────────────────────────────────────────────────────────────┘

Pipeline Flow:

  1. Parse PDF → Extract text with PDF parser or OCR
  2. Detect Encoding → Identify legacy encoding scheme
  3. Convert to Unicode → Transform legacy text to Unicode
  4. Translate → Use translation backend
  5. Generate Output → Create PDF/text/markdown

License

MIT

Contributing

Contributions are welcome! Please:

  1. Fork the repository
  2. Create a feature branch (git checkout -b feature/amazing-feature)
  3. Make your changes
  4. Run tests (uv run pytest)
  5. Commit and push
  6. Open a Pull Request

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

legacylipi-0.6.0.tar.gz (608.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

legacylipi-0.6.0-py3-none-any.whl (86.2 kB view details)

Uploaded Python 3

File details

Details for the file legacylipi-0.6.0.tar.gz.

File metadata

  • Download URL: legacylipi-0.6.0.tar.gz
  • Upload date:
  • Size: 608.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.7

File hashes

Hashes for legacylipi-0.6.0.tar.gz
Algorithm Hash digest
SHA256 5df885a410e7c884b0f34a72ece005a167b67e400360049cb31092772cfcadba
MD5 42732816744883e7b35ca4d94078eef2
BLAKE2b-256 621bf843867ec32419196feba4476791c3b18b80ac30138706348cf459b32041

See more details on using hashes here.

File details

Details for the file legacylipi-0.6.0-py3-none-any.whl.

File metadata

  • Download URL: legacylipi-0.6.0-py3-none-any.whl
  • Upload date:
  • Size: 86.2 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.7

File hashes

Hashes for legacylipi-0.6.0-py3-none-any.whl
Algorithm Hash digest
SHA256 45f0981c9d4f7235a19c9497edb6aeecccfb2ee0e54ea571bd178010fcd24432
MD5 c1e22712e88f977e85dc0ef57da4b7b0
BLAKE2b-256 fb31c3f5c0706e1c24ec93efc10118a30cd2658b0c685d9cf331bebca3bd5d98

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page