Skip to main content

Model GitHub Release License Python 3.9+ PyTorch ONNX Hindi Tamil Telugu Bengali Marathi

🎙️ PolyWhisper v9 — Efficient Multilingual Indic ASR

by Eulogik — Frontier Edge AI · Vernacular Intelligence · eulogik.com

TL;DR: PolyWhisper v9 is a production-ready automatic speech recognition (ASR) system for Hindi, Tamil, Telugu, Bengali, and Marathi. It pairs a frozen OpenAI Whisper-Small backbone (244M params) with tiny per-language LoRA adapters (~14MB each). Bengali WER drops −34.5% and Marathi −43.2% versus the no-augmentation baseline — at roughly 1% of the storage cost of full fine-tuning.

✨ Why PolyWhisper?

Full fine-tune (per language) PolyWhisper v9
Storage per language ~1.5 GB ~14 MB (100× smaller)
Backbone retrained each time frozen once, shared by all 5
Bengali (bn) FLEURS WER 198.8 (baseline) 130.2 (−34.5%)
Marathi (mr) FLEURS WER 170.1 (baseline) 96.7 (−43.2%)
Telugu (te) FLEURS WER 105.9 (baseline) 100.1 (−5.5%)
Hindi (hi) FLEURS WER 43.0 (baseline) 46.3
Tamil (ta) FLEURS WER 68.2 (baseline) 70.1
CPU deployment heavy ONNX INT8, no GPU needed

WER = word error rate (lower is better). FLEURS test set, beam=1, punctuation-normalized scoring.

📊 Benchmarks (FLEURS, beam=1, normalized WER)

Language Code Script v7 (no augment) v9 final Δ vs v7
Hindi hi Devanagari 43.0 46.3 +7.7%
Tamil ta Tamil 68.2 70.1 +2.8%
Telugu te Telugu 105.9 100.1 ✅ −5.5%
Bengali bn Bengali 198.8 130.2 ✅ −34.5%
Marathi mr Devanagari 170.1 96.7 ✅ −43.2%

🧪 The v9 finding: augment per language, not globally

Training with SpecAugment + speed perturbation on all languages damaged Hindi/Tamil (token-loop degeneration) while massively helping Bengali/Marathi. The v9 recipe augments only bn/mr and trains hi/ta clean:

Language Augmentation Result
Hindi, Tamil none (clean) matches no-augment baseline
Telugu, Bengali, Marathi SpecAugment + 0.9×/1.1× speed perturb large gains on hard languages

📦 Which adapter should I use?

Language Adapter file Backbone WER
Hindi (hi) polywhisper_output_hi/adapters_v3/hi_best_clean.pt openai/whisper-small 46.3
Tamil (ta) polywhisper_output_ta/adapters_v3/ta_best_clean.pt openai/whisper-small 70.1
Telugu (te) polywhisper_output_gpu0/adapters_v3/te_best_prod.pt openai/whisper-small 100.1
Bengali (bn) polywhisper_output_gpu0/adapters_v3/bn_best_prod.pt openai/whisper-small 130.2
Marathi (mr) polywhisper_output_gpu1/adapters_v3/mr_best_prod.pt openai/whisper-small 96.7

All adapters are rank-16 LoRA (decoder + encoder attention), ~14MB each. Backbone weights are not included — they load from openai/whisper-small at runtime.

🚀 Quickstart

pip install -e .
# Hindi speech to text
polywhisper transcribe audio.wav --lang hi

# Tamil with JSON output
polywhisper transcribe audio.wav --lang ta --format json

# Auto-detect language, SRT subtitles
polywhisper transcribe audio.wav --format srt > subs.srt

# Batch a folder
polywhisper batch ./audio_folder/ --lang bn --output results.json
from polywhisper import transcribe

result = transcribe("audio.wav", lang="mr")
print(result.text)
print(result.segments)  # timestamped segments

🖥️ CPU-only inference (ONNX Runtime)

Export INT8-quantized ONNX graphs (no PyTorch, no GPU needed at inference):

polywhisper export --lang hi --variant prod --int8

Pre-exported v9 graphs live under export/onnx/ on the Hub — per language, fp32 + INT8:

Lang Encoder (fp32 / INT8) Decoder (fp32 / INT8)
hi 358MB / 97MB 784MB / 204MB
ta 358MB / 97MB 784MB / 204MB
te 358MB / 97MB 784MB / 204MB
bn 358MB / 97MB 784MB / 204MB
mr 358MB / 97MB 784MB / 204MB

Files are named {lang}_{lang}_best_prod_{encoder,decoder}{,_int8}.onnx. INT8 is ~4× smaller.

Verification: fp32 ONNX vs PyTorch max diff < 1e-3 on all five languages (encoder + decoder). End-to-end greedy spot-checks (FLEURS audio, beam=1):

Lang torch WER ONNX INT8 WER
hi (10 samples) 43.4% 48.3%
ta (5 samples) 100.0% 100.0%
te (5 samples) 100.0% 101.6%
bn (5 samples) 104.9% 118.7%
mr (5 samples) 82.9% 89.4%

Spot-checks are tiny (5–10 utterances) so single-sentence flips move the numbers; fp32 ONNX is at parity with torch. INT8 trades a few points for 4× smaller files.

🏋️ Training recipe (reproducible)

  • Data: IndicVoices-ST (~19–20k clips/language) · Eval: FLEURS
  • Backbone: openai/whisper-small, frozen · Adapters: LoRA rank-16, encoder + decoder attention
  • Schedule: 3–5 epochs/language, batch 4, AdamW, cosine LR (peak 1e-4), 2× NVIDIA T4
  • Augmentation (v9): SpecAugment + speed perturb for bn/mr only; hi/ta/te clean
  • Selection: WER-gated checkpoints (*_best_*.pt) on FLEURS dev slices
  • Code: train_v3.py · orchestrator kaggle_train_resumable.py · scoring normalize_ortho.py

❓ FAQ

What is PolyWhisper? PolyWhisper is an open-source Indic ASR toolkit: one frozen Whisper-Small backbone plus five small per-language LoRA adapters covering Hindi, Tamil, Telugu, Bengali, and Marathi.

How is it different from fine-tuning Whisper? Full fine-tuning rewrites ~244M–1.5B weights per language. PolyWhisper freezes the backbone and trains ~3.5M LoRA parameters per language (~14MB), so five languages ship for the storage cost of a rounding error.

Which languages are production-ready? All five ship working adapters. Hindi (46.3 WER) and Tamil (70.1) are strongest; Bengali and Marathi improved dramatically in v9 (−34.5% / −43.2% vs baseline) but remain the hardest languages.

Can I run it on CPU? Yes — export to ONNX INT8 and run with ONNX Runtime, no GPU required.

Can I run it on a Mac? Yes — PyTorch MPS is supported (Device: mps), plus CPU via ONNX.

What data was it trained/evaluated on? Trained on IndicVoices-ST conversational speech, evaluated on FLEURS read speech with punctuation-normalized, script-aware scoring.

⚠️ Limitations

  • Absolute WER on Telugu/Bengali/Marathi is still high — usable for assistive/search/subtitle-draft workflows, not verbatim legal/medical transcription.
  • Evaluated on read speech (FLEURS); spontaneous conversational accuracy will differ.
  • Beam=1 numbers above; beam=5 decoding improves results at higher latency.

📄 License & citation

MIT. Whisper weights © OpenAI. Training data: IndicVoices-ST (CC-BY) · Eval: FLEURS (CC-BY).

@misc{polywhisper2026,
  title  = {PolyWhisper: Efficient Multilingual Indic ASR via Frozen Backbone + Per-Language LoRA},
  author = {Eulogik},
  year   = {2026},
  publisher = {HuggingFace},
  url    = {https://huggingface.co/eulogik/polywhisper}
}

🔗 Links

Metadata

Release files for polywhisper 0.2.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for polywhisper 0.2.1
File Size Uploaded
polywhisper-0.2.1.tar.gz 15.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for polywhisper 0.2.1
File Interpreter ABI Platform
polywhisper-0.2.1-py3-none-any.whl Python 3 none any Details

Total release size: 34.7 kB

Release files / polywhisper-0.2.1.tar.gz

Download URL polywhisper-0.2.1.tar.gz
Size 15.6 kB
Tags Source
SHA-256 checksum
How to use checksums
8e6046b79bcc9f6a38a7aa9333a4677d6ad7a1e26e9a62d931b166c1a38a9494
BLAKE2b-256 checksum
How to use checksums
d1420e0da3a9e775ee9a71b39ded159fb55f3eaa45db9b726f97677a5487f6be
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.14

Release files / polywhisper-0.2.1-py3-none-any.whl

Download URL polywhisper-0.2.1-py3-none-any.whl
Size 19.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
dd852c4588aa8d7a7244ac25b12033f2dc4c67a7e85c3106af0a2f3b907d06c0
BLAKE2b-256 checksum
How to use checksums
d66b03e614d23004f67162f4e64bc5c03b07e30ef914cfbc93d3ac63638b4588
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.14

Release history Release notifications | RSS feed

0.2.2

2 release files

This release

0.2.1 This release

2 release files

0.2.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page