Skip to main content

language:

  • hi
  • ta
  • te
  • bn
  • mr license: mit library_name: transformers pipeline_tag: automatic-speech-recognition base_model: openai/whisper-small tags:
  • polywhisper
  • indic-asr
  • hindi-asr
  • tamil-speech-recognition
  • telugu-stt
  • bengali-asr
  • marathi-speech-to-text
  • speech-recognition
  • multilingual
  • lora
  • whisper
  • hindi
  • tamil
  • telugu
  • bengali
  • marathi
  • indic-languages
  • indian-languages
  • automatic-speech-recognition
  • speech-to-text
  • low-resource-asr
  • fleurs
  • indicvoices
  • onnx
  • quantized
  • efficient-asr
  • edge-asr
  • peft datasets:
  • ai4bharat/indicvoices-st
  • google/fleurs model-index:
  • name: PolyWhisper v9 (Whisper-Small + Per-Language LoRA) results:
    • task: type: automatic-speech-recognition name: Hindi Speech Recognition dataset: name: FLEURS Hindi (hi_in) type: google/fleurs metrics:
      • type: wer value: 46.3 name: WER (beam=1, normalized)
    • task: type: automatic-speech-recognition name: Tamil Speech Recognition dataset: name: FLEURS Tamil (ta_in) type: google/fleurs metrics:
      • type: wer value: 70.1 name: WER (beam=1, normalized)
    • task: type: automatic-speech-recognition name: Telugu Speech Recognition dataset: name: FLEURS Telugu (te_in) type: google/fleurs metrics:
      • type: wer value: 100.1 name: WER (beam=1, normalized)
    • task: type: automatic-speech-recognition name: Bengali Speech Recognition dataset: name: FLEURS Bengali (bn_in) type: google/fleurs metrics:
      • type: wer value: 130.2 name: WER (beam=1, normalized)
    • task: type: automatic-speech-recognition name: Marathi Speech Recognition dataset: name: FLEURS Marathi (mr_in) type: google/fleurs metrics:
      • type: wer value: 96.7 name: WER (beam=1, normalized)

Model GitHub Release License Python 3.9+ PyTorch ONNX Hindi Tamil Telugu Bengali Marathi

🎙️ PolyWhisper v9 — Efficient Multilingual Indic ASR

TL;DR: PolyWhisper v9 is a production-ready automatic speech recognition (ASR) system for Hindi, Tamil, Telugu, Bengali, and Marathi. It pairs a frozen OpenAI Whisper-Small backbone (244M params) with tiny per-language LoRA adapters (~14MB each). Bengali WER drops −34.5% and Marathi −43.2% versus the no-augmentation baseline — at roughly 1% of the storage cost of full fine-tuning.

✨ Why PolyWhisper?

Full fine-tune (per language) PolyWhisper v9
Storage per language ~1.5 GB ~14 MB (100× smaller)
Backbone retrained each time frozen once, shared by all 5
Bengali (bn) FLEURS WER 198.8 (baseline) 130.2 (−34.5%)
Marathi (mr) FLEURS WER 170.1 (baseline) 96.7 (−43.2%)
Telugu (te) FLEURS WER 105.9 (baseline) 100.1 (−5.5%)
Hindi (hi) FLEURS WER 43.0 (baseline) 46.3
Tamil (ta) FLEURS WER 68.2 (baseline) 70.1
CPU deployment heavy ONNX INT8, no GPU needed

WER = word error rate (lower is better). FLEURS test set, beam=1, punctuation-normalized scoring.

📊 Benchmarks (FLEURS, beam=1, normalized WER)

Language Code Script v7 (no augment) v9 final Δ vs v7
Hindi hi Devanagari 43.0 46.3 +7.7%
Tamil ta Tamil 68.2 70.1 +2.8%
Telugu te Telugu 105.9 100.1 ✅ −5.5%
Bengali bn Bengali 198.8 130.2 ✅ −34.5%
Marathi mr Devanagari 170.1 96.7 ✅ −43.2%

🧪 The v9 finding: augment per language, not globally

Training with SpecAugment + speed perturbation on all languages damaged Hindi/Tamil (token-loop degeneration) while massively helping Bengali/Marathi. The v9 recipe augments only bn/mr and trains hi/ta clean:

Language Augmentation Result
Hindi, Tamil none (clean) matches no-augment baseline
Telugu, Bengali, Marathi SpecAugment + 0.9×/1.1× speed perturb large gains on hard languages

📦 Which adapter should I use?

Language Adapter file Backbone WER
Hindi (hi) polywhisper_output_hi/adapters_v3/hi_best_clean.pt openai/whisper-small 46.3
Tamil (ta) polywhisper_output_ta/adapters_v3/ta_best_clean.pt openai/whisper-small 70.1
Telugu (te) polywhisper_output_gpu0/adapters_v3/te_best_prod.pt openai/whisper-small 100.1
Bengali (bn) polywhisper_output_gpu0/adapters_v3/bn_best_prod.pt openai/whisper-small 130.2
Marathi (mr) polywhisper_output_gpu1/adapters_v3/mr_best_prod.pt openai/whisper-small 96.7

All adapters are rank-16 LoRA (decoder + encoder attention), ~14MB each. Backbone weights are not included — they load from openai/whisper-small at runtime.

🚀 Quickstart

pip install -e .
# Hindi speech to text
polywhisper transcribe audio.wav --lang hi

# Tamil with JSON output
polywhisper transcribe audio.wav --lang ta --format json

# Auto-detect language, SRT subtitles
polywhisper transcribe audio.wav --format srt > subs.srt

# Batch a folder
polywhisper batch ./audio_folder/ --lang bn --output results.json
from polywhisper import transcribe

result = transcribe("audio.wav", lang="mr")
print(result.text)
print(result.segments)  # timestamped segments

🖥️ CPU-only inference (ONNX Runtime)

Export INT8-quantized ONNX graphs (no PyTorch, no GPU needed at inference):

polywhisper export --lang hi --variant prod --int8

Pre-exported v9 graphs live under export/onnx/ on the Hub — per language, fp32 + INT8:

Lang Encoder (fp32 / INT8) Decoder (fp32 / INT8)
hi 358MB / 97MB 784MB / 204MB
ta 358MB / 97MB 784MB / 204MB
te 358MB / 97MB 784MB / 204MB
bn 358MB / 97MB 784MB / 204MB
mr 358MB / 97MB 784MB / 204MB

Files are named {lang}_{lang}_best_prod_{encoder,decoder}{,_int8}.onnx. INT8 is ~4× smaller.

Verification: fp32 ONNX vs PyTorch max diff < 1e-3 on all five languages (encoder + decoder). End-to-end greedy spot-checks (FLEURS audio, beam=1):

Lang torch WER ONNX INT8 WER
hi (10 samples) 43.4% 48.3%
ta (5 samples) 100.0% 100.0%
te (5 samples) 100.0% 101.6%
bn (5 samples) 104.9% 118.7%
mr (5 samples) 82.9% 89.4%

Spot-checks are tiny (5–10 utterances) so single-sentence flips move the numbers; fp32 ONNX is at parity with torch. INT8 trades a few points for 4× smaller files.

🏋️ Training recipe (reproducible)

  • Data: IndicVoices-ST (~19–20k clips/language) · Eval: FLEURS
  • Backbone: openai/whisper-small, frozen · Adapters: LoRA rank-16, encoder + decoder attention
  • Schedule: 3–5 epochs/language, batch 4, AdamW, cosine LR (peak 1e-4), 2× NVIDIA T4
  • Augmentation (v9): SpecAugment + speed perturb for bn/mr only; hi/ta/te clean
  • Selection: WER-gated checkpoints (*_best_*.pt) on FLEURS dev slices
  • Code: train_v3.py · orchestrator kaggle_train_resumable.py · scoring normalize_ortho.py

❓ FAQ

What is PolyWhisper? PolyWhisper is an open-source Indic ASR toolkit: one frozen Whisper-Small backbone plus five small per-language LoRA adapters covering Hindi, Tamil, Telugu, Bengali, and Marathi.

How is it different from fine-tuning Whisper? Full fine-tuning rewrites ~244M–1.5B weights per language. PolyWhisper freezes the backbone and trains ~3.5M LoRA parameters per language (~14MB), so five languages ship for the storage cost of a rounding error.

Which languages are production-ready? All five ship working adapters. Hindi (46.3 WER) and Tamil (70.1) are strongest; Bengali and Marathi improved dramatically in v9 (−34.5% / −43.2% vs baseline) but remain the hardest languages.

Can I run it on CPU? Yes — export to ONNX INT8 and run with ONNX Runtime, no GPU required.

Can I run it on a Mac? Yes — PyTorch MPS is supported (Device: mps), plus CPU via ONNX.

What data was it trained/evaluated on? Trained on IndicVoices-ST conversational speech, evaluated on FLEURS read speech with punctuation-normalized, script-aware scoring.

⚠️ Limitations

  • Absolute WER on Telugu/Bengali/Marathi is still high — usable for assistive/search/subtitle-draft workflows, not verbatim legal/medical transcription.
  • Evaluated on read speech (FLEURS); spontaneous conversational accuracy will differ.
  • Beam=1 numbers above; beam=5 decoding improves results at higher latency.

📄 License & citation

Apache 2.0. Whisper weights © OpenAI. Training data: IndicVoices-ST (CC-BY) · Eval: FLEURS (CC-BY).

@misc{polywhisper2026,
  title  = {PolyWhisper: Efficient Multilingual Indic ASR via Frozen Backbone + Per-Language LoRA},
  author = {Eulogik},
  year   = {2026},
  publisher = {HuggingFace},
  url    = {https://huggingface.co/eulogik/polywhisper}
}

🔗 Links

Metadata

Release files for polywhisper 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for polywhisper 0.2.0
File Size Uploaded
polywhisper-0.2.0.tar.gz 15.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for polywhisper 0.2.0
File Interpreter ABI Platform
polywhisper-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 35.4 kB

Release files / polywhisper-0.2.0.tar.gz

Download URL polywhisper-0.2.0.tar.gz
Size 15.9 kB
Tags Source
SHA-256 checksum
How to use checksums
ce32aa17151e143ff78145af6590f436d58d06fb2aee98fe339c1a3257af7f19
BLAKE2b-256 checksum
How to use checksums
8269b34bbdc5821201857ad9f0357f027f28ad1c2da4f649bac3975d1ba952ad
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.14

Release files / polywhisper-0.2.0-py3-none-any.whl

Download URL polywhisper-0.2.0-py3-none-any.whl
Size 19.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
7b1ed7cd06360f4e71a6a3dc915d8efb8cd4f814e7c622184693ec03adeb6778
BLAKE2b-256 checksum
How to use checksums
dbb0a97fb234cd5ff1fa866af3cc559b404d536ac1a2e7bde26cb7b7c367e0df
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.14

Release history Release notifications | RSS feed

0.2.2

2 release files

0.2.1

2 release files

This release

0.2.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page