Legacy Font PDF Translator - Translate PDFs with legacy Indian font encodings to English
Project description
LegacyLipi
Legacy Font PDF Translator - Translate PDF documents with legacy Indian font encodings to English.
Installation
From PyPI (Recommended)
pip install legacylipi
Or with uv:
uv tool install legacylipi
From Source
git clone https://github.com/biswasbiplob/legacylipi.git
cd legacylipi
uv sync
Usage
# CLI translation
legacylipi translate input.pdf -o output.txt
# Launch web UI
legacylipi ui
# Launch UI on custom port
legacylipi ui --port 3000
Problem
Millions of government documents, legal papers, and archival materials in Indian regional languages (Marathi, Hindi, Tamil, etc.) were created using legacy font encoding systems (Shree-Lipi, Kruti Dev, APS, Chanakya, etc.). These fonts map Devanagari/regional script glyphs to ASCII/Latin code points, making them unreadable by standard translation tools.
Example:
- What the PDF displays: महाराष्ट्र राजभाषा अधिनियम
- What text extraction produces:
´ÖÆüÖ¸üÖ™Òü ¸üÖ•Ö³ÖÖÂÖÖ †×¬Ö×®ÖμÖ´Ö - What Google Translate sees: Gibberish
Solution
LegacyLipi:
- Detects the font encoding scheme used in a PDF (legacy or Unicode)
- Converts legacy-encoded text to proper Unicode
- Alternatively, uses OCR to extract text from scanned PDFs
- Translates the Unicode text to the target language
- Outputs translated text in various formats (text, markdown, PDF)
Installation
# Clone and install
git clone https://github.com/biswasbiplob/legacylipi.git
cd legacylipi
uv sync
# With all optional backends
uv sync --all-extras
OCR Support (Optional)
LegacyLipi supports multiple OCR backends:
| Backend | Description | GPU Support |
|---|---|---|
| Tesseract | Local, free, most language packs | CPU only |
| Google Vision | Cloud, paid, best accuracy | N/A |
| EasyOCR | Local, free, good for Indian languages | CUDA, MPS (Apple Silicon) |
Tesseract (default):
# Ubuntu/Debian
sudo apt-get install tesseract-ocr tesseract-ocr-mar tesseract-ocr-hin
# macOS
brew install tesseract tesseract-lang
EasyOCR with GPU (optional):
# Install with EasyOCR support
uv sync --extra easyocr
# For GPU acceleration, install PyTorch with CUDA or MPS support
Google Vision (optional):
uv sync --extra vision
# Requires GCP credentials (GOOGLE_APPLICATION_CREDENTIALS)
See docs/cli-reference.md for detailed OCR options and language codes.
Quick Start
# Basic translation
uv run legacylipi translate input.pdf -o output.txt
# Output as PDF (preserves layout)
uv run legacylipi translate input.pdf -o output.pdf --format pdf
# OCR for scanned documents
uv run legacylipi translate input.pdf --use-ocr -o output.txt
# Use local LLM (requires Ollama)
uv run legacylipi translate input.pdf --translator ollama --model llama3.2
# Detect encoding only
uv run legacylipi detect input.pdf
See docs/cli-reference.md for complete CLI documentation.
Web UI
LegacyLipi includes a web interface for easy PDF translation without command-line usage.
uv run legacylipi-ui
Open http://localhost:8080 in your browser.
Features:
- Drag-and-drop PDF upload
- Multiple translation backends
- OCR support with language selection
- Structure-preserving or flowing text modes
- Real-time progress tracking
- Direct download of translated files
Translation Backends
| Backend | Description | Setup |
|---|---|---|
trans |
translate-shell CLI (recommended) | brew install translate-shell |
google |
Google Translate (free API) | Works out of the box |
mymemory |
MyMemory API (free) | Works out of the box |
ollama |
Local LLM via Ollama | Ollama required |
openai |
OpenAI GPT models | Set OPENAI_API_KEY |
gcp_cloud |
Google Cloud Translation | GCP project + credentials |
See docs/translation-backends.md for detailed setup guides.
Supported Encodings
| Encoding | Font Family | Language | Status |
|---|---|---|---|
| shree-lipi | Shree-Lipi, Shree-Dev-0714 | Marathi | ✅ Built-in |
| kruti-dev | Kruti Dev | Hindi | ✅ Built-in |
| aps-dv | APS-DV | Hindi | 🔄 Detection only |
| chanakya | Chanakya | Hindi | 🔄 Detection only |
| dvb-tt | DVB-TT, DV-TTYogesh | Hindi | 🔄 Detection only |
| walkman-chanakya | Walkman Chanakya | Hindi | 🔄 Detection only |
| shusha | Shusha | Hindi | 🔄 Detection only |
CLI Commands
| Command | Description |
|---|---|
translate |
Full pipeline: parse → detect → convert → translate → output |
convert |
Convert legacy encoding to Unicode (no translation) |
extract |
Extract text from PDF (OCR or font-based) |
detect |
Analyze PDF and report detected encoding |
encodings |
List supported font encodings |
usage |
Show API usage statistics |
See docs/cli-reference.md for full command reference.
Development
See docs/development.md for setup instructions, running tests, project structure, and adding new encodings.
Architecture
┌─────────────────────────────────────────────────────────────────────────┐
│ LegacyLipi │
├─────────────────────────────────────────────────────────────────────────┤
│ │
│ ┌──────────────────────────────────────────────────────────────────┐ │
│ │ Text Extraction │ │
│ │ ┌──────────────┐ ┌──────────────┐ │ │
│ │ │ PDF │ OR │ OCR │ │ │
│ │ │ Parser │ │ Parser │ │ │
│ │ │ (font-based) │ │ (Tesseract) │ │ │
│ │ └──────────────┘ └──────────────┘ │ │
│ └──────────────────────────────────────────────────────────────────┘ │
│ │ │ │
│ ▼ ▼ │
│ ┌──────────────┐ ┌──────────────┐ │
│ │ Encoding │───▶│ Unicode │◀──── (OCR output is │
│ │ Detector │ │ Converter │ already Unicode) │
│ └──────────────┘ └──────────────┘ │
│ │ │
│ ▼ │
│ ┌───────────────────────────────────────────────────────────┐ │
│ │ Translation Engine │ │
│ │ ┌────────┬────────┬──────────┬────────┬────────┬─────┐ │ │
│ │ │ trans │ Google │ MyMemory │ Ollama │ OpenAI │ GCP │ │ │
│ │ │ (CLI) │ Trans. │ (API) │(Local) │ (API) │Cloud│ │ │
│ │ └────────┴────────┴──────────┴────────┴────────┴─────┘ │ │
│ └───────────────────────────────────────────────────────────┘ │
│ │ │
│ ▼ │
│ ┌───────────────────────────────────────────────┐ │
│ │ Output Generator │ │
│ │ ┌──────┬────────┬───────┐ │ │
│ │ │ .txt │ .md │ .pdf │ │ │
│ │ └──────┴────────┴───────┘ │ │
│ └───────────────────────────────────────────────┘ │
│ │
└─────────────────────────────────────────────────────────────────────────┘
Pipeline Flow:
- Parse PDF → Extract text with PDF parser or OCR
- Detect Encoding → Identify legacy encoding scheme
- Convert to Unicode → Transform legacy text to Unicode
- Translate → Use translation backend
- Generate Output → Create PDF/text/markdown
License
MIT
Contributing
Contributions are welcome! Please:
- Fork the repository
- Create a feature branch (
git checkout -b feature/amazing-feature) - Make your changes
- Run tests (
uv run pytest) - Commit and push
- Open a Pull Request
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file legacylipi-0.7.0.tar.gz.
File metadata
- Download URL: legacylipi-0.7.0.tar.gz
- Upload date:
- Size: 609.6 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.1.0 CPython/3.13.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
28e9bafea98c06092d7113df777b6a0d9256024a3e56fae19a8c06ccd167bd58
|
|
| MD5 |
d6a14e2fb91cc5af9e7c21e7278526bd
|
|
| BLAKE2b-256 |
8489394328b5bc2e6f9f963eee936d439755d20d71a32f5415a8dc8b0161be7b
|
File details
Details for the file legacylipi-0.7.0-py3-none-any.whl.
File metadata
- Download URL: legacylipi-0.7.0-py3-none-any.whl
- Upload date:
- Size: 87.3 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.1.0 CPython/3.13.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
614b01f696ab8dee522d1e8dbf75fc5dc66d7f26c20fccfca4aa2830c2c3cdb6
|
|
| MD5 |
612d13610dd1283e6d26b7b11919b54f
|
|
| BLAKE2b-256 |
2ddc069a61519de5b3534fcd150912f512ac33503622687d03a3ed993a0715d0
|