indic-language-utils
Provider-neutral foundations for Indian language operations in Python.
indic-language-utils standardizes language detection, neural machine translation, script identification, and text processing across Indian languages. It abstracts concrete cloud and local engines behind unified, resilient interfaces, allowing applications to start with local or low-cost providers and switch routing configuration for production without rewriting business logic.
Key Features
- Provider Neutrality: Code against high-level capability interfaces. Swap, configure, or chain providers without changing text processing or domain code.
- Text Translation: Translate plain text or complex Markdown documents across 22 scheduled Indian languages and English.
- Text Language Detection: Identify languages using offline FastText classification (
lid.176.ftz) or cloud inference pipelines via Bhashini. - Transliteration: Convert between Roman script and native Indic scripts via Bhashini or AI4Bharat IndicXlit.
- Speech to text: Transcribe audio with Bhashini, Sarvam, local Faster-Whisper, or keyless Google Speech.
- Text to speech: Generate audio with Bhashini and pass model-specific voice settings.
- Script Identification: Fast, zero-dependency Unicode script identification across 12+ Indic scripts and Latin.
- Document and Code Protection: Structural pre-processors and post-processors protect headings, bullet markers, inline code spans, URLs, and code blocks from neural translation corruption.
- Resilient Execution: Automatic multi-provider fallback routing, bounded concurrency limits per provider, and exponential backoff retries with jitter.
- High-Performance Caching: In-memory LRU and multi-process SQLite caches with write-ahead logging (WAL mode) and stampede protection.
- Canonical Normalization: Shared language registry recognizing all 22 Eighth Schedule Indian languages plus English, mapping aliases and regional codes to BCP 47.
- Localization Catalogs: Match reviewed human translations for critical UI strings before dispatching to neural engines.
Getting started
Install the core package with pip install indic-language-utils. Provider selection is explicit,
and local detection requires an optional extra. The
installation and quick start guide covers a credential-free setup,
Bhashini configuration, and the first detection and translation calls.
Supported Providers
| Provider | Capability | Mode | Prerequisites |
|---|---|---|---|
| Aksharamukha | Transliteration (120+ scripts) | Offline / Local | [local-transliteration] extra |
| Bhashini | Translation, Detection, Transliteration, STT, TTS | Cloud API | API key, Endpoint, Service ID |
| Sarvam AI | Translation, Detection, STT, TTS | Cloud API | API key (SARVAM_API_KEY) |
| AI4Bharat IndicXlit | Transliteration | Offline / Local | [neural-transliteration] extra |
| FastText | Text Language Detection | Offline / Local | [local-tld] extra |
| Faster-Whisper | Speech to text | Offline / Local | [stt-whisper] extra |
| Google Free STT | Speech to text | Cloud (unofficial) | [stt-google-free] extra |
| Google Translate | Translation | Cloud (unofficial) | [googletrans] extra |
| Microsoft Edge TTS | Text to speech | Cloud (unofficial) | [tts-edge] extra |
Configuration
The library uses a tiered configuration system combining project TOML files (.indic-language-utils.toml), environment variables, and programmatic overrides.
Example .indic-language-utils.toml:
[cache]
enabled = true
backend = "sqlite"
path = ".cache/translations.sqlite3"
max_entries = 50000
ttl_seconds = 86400
[retry]
max_attempts = 3
base_delay_seconds = 0.25
max_delay_seconds = 5.0
[providers.bhashini]
endpoint = "https://dhruva-api.bhashini.gov.in/services/inference/pipeline"
translation_service_id = "default-translation-model-id"
detection_service_id = "default-tld-model-id"
transliteration_service_id = "default-transliteration-model-id"
max_concurrency = 8
[providers.sarvam]
endpoint = "https://api.sarvam.ai"
model = "sarvam-translate:v1"
stt_model_id = "saaras:v4"
tts_model_id = "bulbul:v3"
max_concurrency = 8
[routes]
translation = ["sarvam", "bhashini", "googletrans"]
text_language_detection = ["sarvam", "bhashini", "fasttext"]
transliteration = ["bhashini", "indicxlit"]
speech_to_text = ["bhashini", "sarvam", "google_free", "faster_whisper"]
text_to_speech = ["bhashini", "sarvam", "edge_tts"]
Supply credentials securely through environment variables:
export BHASHINI_API_KEY="your-bhashini-api-key"
export SARVAM_API_KEY="your-sarvam-api-key"
Documentation
Comprehensive guides are available in the documentation site:
- Installation and Quick Start: Provider setup and first calls.
- User Guide: Architecture overview, core capabilities, and usage styles.
- Translation Guide: Synchronous and asynchronous translation, Markdown preservation, and catalogs.
- Detection Guide: Local FastText and cloud Bhashini detection, script analysis, and candidate scoring.
- Speech to text guide: Bhashini and Sarvam transcription and model IDs.
- STT example: Transcribe audio with Bhashini or Sarvam.
- Text to speech guide: Bhashini and Sarvam synthesis and voice options.
- TTS example: Generate audio with Bhashini or Sarvam.
- Configuration Reference: Project TOML file formats, precedence rules, and environment variables.
- Processor Pipelines: Structural processors, segment processors, and custom pipeline authoring.
- Architecture Proposal: Design philosophy and ecosystem review.
- Changelog: Release history adhering to Keep a Changelog.
Development
Install the locked development environment with uv:
uv sync --dev
uv run pre-commit install
Run test suite, code quality checks, and local documentation server:
uv run ruff format --check .
uv run ruff check .
uv run mypy
uv run pytest
uv build
uv run mkdocs serve
uv run mkdocs build --strict
Interactive Testing Workbench
A local evaluation workbench and REST API server is included for testing translation, language detection, and script identification interactively:
# Build the web interface
cd web && pnpm install && pnpm build && cd ..
# Launch the FastAPI server with static UI mounted at http://127.0.0.1:8000
uv run indic-server
Explore interactive API documentation at http://127.0.0.1:8000/docs.
License
This project is licensed under the MIT License.
Release files for indic-language-utils 0.4.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| indic_language_utils-0.4.0.tar.gz | 617.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| indic_language_utils-0.4.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 738.3 kB
Release files / indic_language_utils-0.4.0.tar.gz
| Download URL | indic_language_utils-0.4.0.tar.gz |
|---|---|
| Size | 617.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
ebb3e01d09315cd2bee6e84afd44908c4fa16e30127dff4e49f222cb7759ef9d
|
|
BLAKE2b-256 checksum How to use checksums |
a120e98af043c4c0d5f71e2175208687071d368cc9c3c580b62b1c7298c551fc
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 23, 2026.
Transparency logRelease files / indic_language_utils-0.4.0-py3-none-any.whl
| Download URL | indic_language_utils-0.4.0-py3-none-any.whl |
|---|---|
| Size | 121.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
07b07644e808b94d37e04d594c08842041c7d2fc177a3cd452d35db473bc2c10
|
|
BLAKE2b-256 checksum How to use checksums |
b18de73437260f38807a7e74fc5cb5148625683e078b9aa227431a4ea065aa2b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 23, 2026.
Transparency log