Skip to main content

QwenCleo-ASR โ€” Egyptian Arabic & code-switching speech recognition, built on Qwen3-ASR.

Project description

๐ŸŽ™๏ธ QwenCleo-ASR

The best open-source model for Egyptian Arabic & code-switching speech recognition

Built on Qwen3-ASR-1.7B, fine-tuned for Egyptian dialect and Arabic โ†” English code-switching.

๐Ÿค— Model PyPI GitHub Base model License Open In Colab

QwenCleo-ASR

QwenCleo โ€” the name carries three meanings: Qwen, the powerful base model it is built on; Queen, signalling a model that reigns over its domain; and Cleo, for Cleopatra, the queen of Egypt โ€” because this model is tailored for Egyptian Arabic. ๐Ÿ‘‘๐Ÿบ

QwenCleo-ASR is, to our knowledge, the best open-source ASR model for Egyptian Arabic and Arabic/English code-switching. It cuts the error rate of the strong Qwen3-ASR base roughly in half, and correctly keeps embedded English tech/loan words in Latin script (engineering, download, React, at least) instead of mangling them into broken Arabic.

  • ๐ŸŽฏ Egyptian dialect โ€” tuned on hundreds of hours of Egyptian podcast speech.
  • ๐Ÿ”€ Code-switching โ€” keeps English terms in code-script, Arabic in Arabic.
  • ๐Ÿฅ‡ State-of-the-art (open) โ€” beats Qwen3-ASR base, NVIDIA Nemotron, Cohere, and every Whisper variant on our Egyptian + CS test set.
  • ๐Ÿ“ฆ pip install qwencleo-asr โ€” inference & chunked long-audio transcription in three lines.
  • โšก Real streaming โ€” token-by-token via vLLM (asr.stream(...)), plus a mic web demo.
  • ๐Ÿš€ Serving โ€” FastAPI server, Gradio demo, OpenAI-compatible vLLM endpoint.

๐Ÿ“Š Results

WER / CER (%) on an Egyptian-Arabic + code-switching test set (3,699 utterances). Lower is better. All models scored with the same Egyptian-aware normalization.

Benchmark overview

Rank Model Params WER all CER all WER ยท AR CER ยท AR WER ยท CS CER ยท CS
๐Ÿฅ‡ QwenCleo-ASR 1.7B 19.85 10.52 19.08 10.30 20.75 10.81
๐Ÿฅˆ Arabic-Whisper-CodeSwitching 1.55B 37.98 17.86 38.48 18.71 36.92 16.22
๐Ÿฅ‰ NVIDIA Nemotron-3.5 0.6B 38.04 19.72 36.23 16.43 41.44 25.66
4 Qwen3-ASR-1.7B (base) 1.7B 41.51 20.86 40.59 18.52 43.20 25.04
5 Whisper Large-v3 Turbo (FT) 0.81B 50.83 22.86 48.37 18.42 55.08 37.84
6 Cohere Transcribe 03-2026 2.0B 53.20 39.12 47.95 33.45 63.76 49.66
7 MasriSwitch-Gemma3n 8B 57.30 30.19 60.38 32.91 51.32 25.13
8 Whisper Large-v3 1.54B 63.94 39.76 49.25 22.76 59.32 31.52
9 Whisper Large-v2 1.54B 72.34 48.73 60.75 33.21 66.85 40.75
10 Whisper Large-v3 Turbo 0.81B 73.83 46.86 59.37 29.42 66.08 37.84
11 Whisper Medium 0.76B 80.46 53.19 74.77 41.76 74.15 44.90
12 Whisper Small 0.24B 89.99 60.34 77.42 42.53 87.09 55.22
13 Whisper Tiny 0.04B 124.68 89.42 116.02 77.74 110.67 74.57

๐Ÿ—ฃ๏ธ Sample outputs

Real transcriptions from the test set, scored after Egyptian-aware normalization. Examples span short, medium, and long clips in both Egyptian Arabic and code-switching. These are picked to be honest โ€” on some clips other models also do well; QwenCleo's edge is consistency, dialect fidelity, and keeping English terms in Latin script.

โœ… = matches ground truth ยท โš ๏ธ = minor slips ยท โŒ = clear errors. Bold marks the wrong spans.

๐Ÿ”€ Code-switching

Short

โœ… Ground truth ุงุงู‡ ุจุญุณ ุณุชุงูŠู„ู‡ ุญู„ูˆ ูˆุดูƒู„ู‡ ุญู„ูˆ ูƒุงุฑูŠุฒู…ุง ูˆู‡ูˆ ุจูŠู„ุนุจ

Model Output
๐Ÿฅ‡ QwenCleo ุจุญุณ style ู‡ ุญู„ูˆ ูˆุดูƒู„ู‡ ุญู„ูˆ ูƒุงุฑูŠุฒู…ุง ูˆู‡ูˆ ุจูŠู„ุนุจ โš ๏ธ
Qwen3-ASR (base) ุฃู‡ ุจุญุณ ุทุงูŠู„ู‡ ุญู„ูˆ ูˆุดูƒู„ู‡ ุญู„ูˆ ูˆูƒุฑุฒู…ู‡ ูˆู‡ูˆ ูŠู„ุนุจ โŒ
Cohere ุงุงุงุงุงุงุงุง ุจุญุณ ุณุชุงูŠู„ู‡ ุญู„ูˆ ูˆุดูƒู„ู‡ ุญู„ูˆ ูˆูƒุงุฑูŠุฒู…ุง ูˆู‡ูˆ ุจูŠู„ุนุจ โœ…
Nemotron ุขู‡ ุจุญุณ ุณุชุงูŠู„ู‡ ุญู„ูˆ ูˆุดูƒู„ู‡ ุญู„ูˆ ูƒุงุฑูŠุฒู…ุง ู‡ูˆ ุจูŠู„ุนุจ โœ…
Arabic-Whisper-CS ุจุญุณ style ูˆ ุญู„ูˆ ูˆุดูƒู„ ูˆุญู„ูˆ ูˆ charisma ูˆู‡ูˆ ุจูŠู„ุนุจ โš ๏ธ
MasriSwitch ุจุญุณ ุณุชุงูŠู„ู‡ ุญู„ูˆ ูˆ ุดูƒู„ู‡ ุญู„ูˆ ูˆ ูƒุงุฑูŠุฒู…ุง ูˆู‡ูˆ ุจูŠุนู„ู‚ โŒ

Medium

โœ… Ground truth โ€” ุฏู‡ ุญุตู„ ุงุฒุงูŠ ู‚ุฑุงุฑ ุงุตู„ุง ุงู†ูƒ ุชุนู…ู„ ู‚ู†ุงุฉ YouTube

Model Output
๐Ÿฅ‡ QwenCleo ุฏู‡ ุญุตู„ ุงุฒุงูŠ ู‚ุฑุงุฑ ุงุตู„ุง ุงู†ูƒ ุชุนู…ู„ ู‚ู†ุงุฉ YouTube โœ…
Qwen3-ASR (base) ุฏู‡ ุญุตู„ ุฅุฒุงูŠ ู‚ุฑุงุฑ ุฃุตู„ุง ุฅู†ูƒ ุชุนู…ู„ ู‚ู†ุงุฉ ูŠูˆุชูŠูˆุจ โš ๏ธ
Cohere ุฏู‡ ุญุตู„ ุงุฒุงูŠ ู‚ุฑุงุฑ ุงุตู„ุง ุงู†ูƒ ุชุนู…ู„ ู‚ู†ุงู‡ ูŠูˆุชูŠูˆุจ โš ๏ธ
Nemotron ุฏู‡ ุญุตู„ ุฅุฒุงูŠ ู‚ุฑุงุฑ ุฃุตู„ุงู‹ ุฅู†ูƒ ุชุนู…ู„ ู‚ู†ุงุฉ ูŠูˆุชูŠูˆุจ โš ๏ธ
Arabic-Whisper-CS ุฏุง ุญุตู„ ุงุฒุงูŠ ู‚ุฑุงุฑ ุงุตู„ุง ุงู†ูƒ ุชุนู…ู„ ู‚ู†ุงุฉ ูŠูˆุชูŠูˆุจ โš ๏ธ
MasriSwitch ุฏู‡ ุญุตู„ ุงุฒุงูŠ ู‚ุฑุงุฑ ุฃุตู„ุง ุฅู†ูƒ ุชุนู…ู„ ู‚ู†ุงุฉ YouTube โœ…

โœ… Ground truth โ€” ุฎู„ูŠู†ูŠ ุงูˆุถุญ ุจุณ ู‡ุฏูู‡ุง ุงู† ุงู„ุทุงู„ุจ ูŠุฌูŠ ู‡ูˆ already

Model Output
๐Ÿฅ‡ QwenCleo ุฎู„ูŠู†ูŠ ุงูˆุถุญ ุจุณ ู‡ุฏูู‡ุง ุงู† ุงู„ุทุงู„ุจ ูŠุฌูŠ ู‡ูˆ already โœ…
Qwen3-ASR (base) ุฎู„ูŠู†ุง ุฃูˆุถุญ ุจุณ ู‡ุฏูู‡ ุฅู† ุงู„ุทุงู„ุจ ูŠุฌูŠ ู‡ูˆ ุฃุฑุฑูŠุฏูŠ โŒ
Cohere ุฎู„ูŠู†ูŠ ุงูˆุถุญ ุจุณ ู‡ุฏูู‡ุง ุงู† ุงู„ุทุงู„ุจ ูŠุฌูŠ ู‡ูˆ ุงูˆุฑุฏูŠ โŒ
Nemotron ุฎู„ูŠู†ุง ูˆุถุญ ุจุณ ู‡ุฏูู‡ุง ุฅู† ุงู„ุทุงู„ุจ ูŠุฌูŠ ู‡ูˆ ุฃูˆุฑุฏูŠ โŒ
Arabic-Whisper-CS ุฎู„ูŠู†ูŠ ุงูˆุถุญ ุจุณ ู‡ุฏูู‡ุง ุงู† ุงู„ุทุงู„ุจ ูŠูŠุฌูŠ ู‡ูˆ already โœ…
MasriSwitch ุฎู„ูŠู†ูŠ ุฃูˆุถุญ ุจุณ ุงู„ุฃู‡ุฏุงู ุฅู† ุงู„ุทุงู„ุจ ุจูŠุฌูŠ ู‡ูˆ already โš ๏ธ

Long

โœ… Ground truth โ€” ู…ุง ู‡ูˆ ุงู„ููƒุฑุฉ ุฌุงูŠุฉ ููŠู† ุงู„ู†ู‡ุงุฑุฏู‡ average ุงู„ุนุฑุจูŠุฉ ุงู„ู„ูŠ ู…ู…ูƒู† ุชุฌูŠุจู‡ุง ู…ู† ุฃูˆุฑูˆุจุง ู…ุง ุจูŠู† 23 ู„ 28000 dollars ุชู‚ุฑูŠุจุง

Model Output
๐Ÿฅ‡ QwenCleo ู…ุง ู‡ูˆ ุงู„ููƒุฑู‡ ุฌุงูŠู‡ ููŠู† ุงู„ู†ู‡ุงุฑุฏู‡ average ุงู„ุนุฑุจูŠู‡ ุงู„ู„ูŠ ู…ู…ูƒู† ุชุฌูŠุจู‡ุง ู…ู† ุงูˆุฑูˆุจุง ู…ุง ุจูŠู† 23 ู„ 28000 dollars โš ๏ธ
Qwen3-ASR (base) ู…ุง ู‡ูˆ ุงู„ููƒุฑุฉ ุฌุงูŠุฉ ููŠู† ุงู„ู†ู‡ุงุฑุฏุฉ ุฃูุฑุฌ ุงู„ุนุฑุจูŠุฉ ุงู„ู„ูŠ ู…ู…ูƒู† ุชุฌุงุจู‡ุง ู…ู† ุฃูˆุฑูˆุจุง ู…ุง ุจูŠู† 23 ู„ 28000 ุฏูˆู„ุงุฑ โš ๏ธ
Cohere ู…ุง ู‡ูˆ ุงู„ููƒุฑู‡ ุฌุงูŠู‡ ููŠู† ุงู„ู†ู‡ุงุฑุฏู‡ ุงูุฑูŠู‚ูŠุง ุงู„ุนุฑุจูŠู‡ ุงู„ู„ูŠ ู…ู…ูƒู† ุชุฌูŠุจู‡ุง ู…ู† ุงูˆุฑูˆุจุง ู…ุง ุจูŠู† 23 ู„ 28000 ุฏูˆู„ุงุฑ โŒ
Nemotron ู…ุง ู‡ูˆ ุงู„ููƒุฑุฉ ุฌุงูŠุฉ ููŠู† ุงู„ู†ู‡ุงุฑ ุฏู‡ ุฃูุฑูŠุฏ ุงู„ุนุฑุจูŠุฉ ุงู„ู„ูŠ ู…ู…ูƒู† ุชุฌูŠุจู‡ุง ู…ู† ุฃูˆุฑูˆุจุง ู…ุง ุจูŠู† 23 ู„ 28000 ุฏูˆู„ุงุฑ ุซู„ุงุซุฉ โŒ
Arabic-Whisper-CS ู…ุง ู‡ูˆ ุงู„ููƒุฑุฉ ุฌุงูŠุฉ ููŠู† ุงู„ู†ู‡ุงุฑ ุฏุง average ุงู„ุนุฑุจูŠุฉ ุงู„ู„ูŠ ู…ู…ูƒู† ุชุฌูŠุจู‡ุง ู…ู† ุฃูˆุฑูˆุจุง ู…ุง ุจูŠู† 23 ู„ 28 ุฃู„ู ุฏูˆู„ุงุฑ ุชู‚ุฑูŠุจุง โš ๏ธ
MasriSwitch ู…ุง ู‡ูˆ ุงู„ููƒุฑุฉ ุฌุงูŠุฉ ููŠู† ุงู„ู†ู‡ุงุฑุฏุฉ average ุงู„ุนุฑุจูŠุฉ ุงู„ู„ูŠ ู…ู…ูƒู† ุชุฌูŠุจู‡ุง ู…ู† ุฃูˆุฑูˆุจุง ู…ุง ุจูŠู† 23 ู„ู€28 ุฃู„ู ุฏูˆู„ุงุฑ ุชู‚ุฑูŠุจู‹ุง โš ๏ธ

โœ… Ground truth โ€” ูุจุชุฏูŠูƒ ูƒู„ ุงู„soft skills ุจุดูƒู„ indirect ุญุฑููŠุง ุจุชุฏูŠู‡ุงู„ูƒ ุจุงู„ู…ุนู„ู‚ู‡ ู„ูŠู‡ ุจุชุชุนู„ู… ุชุดุชุบู„ ุนู„ู‰ ูƒู„ ุงู„ุญุงุฌุงุช ุจุชุงุนู‡ Microsoft Excel PowerPoint Word whatever ุจุนุฏูŠู† ูƒู…ุงู† ุจุชุชุนู„ู… ุงู† ุงู†ุช ุชุชูƒู„ู… ู…ุน ุงู„ู†ุงุณ public speaking

Model Output
๐Ÿฅ‡ QwenCleo ูุจุชุฏูŠูƒ ูƒู„ ุงู„soft skills ุจุดูƒู„ indirect ุญุฑููŠุง ุจุชุฏูŠู‡ุงู„ูƒ ุจุงู„ู…ุนู„ู‚ู‡ ู„ูŠู‡ ุจุชุชุนู„ู… ุชุดุชุบู„ ุนู„ู‰ ูƒู„ ุงู„ุญุงุฌุงุช ุจุชุงุนู‡ Microsoft Excel, PowerPoint, Word ูˆ whatever ุจุนุฏูŠู† ูƒู…ุงู† ุจุชุชุนู„ู… ุงู† ุงู†ุช ุชุชูƒู„ู… ู…ุน ุงู„ู†ุงุณ public speaking โœ…
Arabic-Whisper-CS ูุจุชุฏูŠูƒ ูƒู„ ุงู„ soft skills ุจุดูƒู„ indirect ุญุฑููŠุง ุจุชุฏูŠู‡ุงู„ูƒ ุจุงู„ู…ุนู„ู‚ุฉ ู„ูŠู‡ ุจุชุชุนู„ู… ุชุดุชุบู„ ุนู„ู‰ ูƒู„ ุงู„ุญุงุฌุงุช ุจุชุงุนุฉ Microsoft, Excel, PowerPoint, Word whatever ุจุนุฏูŠู† ูƒู…ุงู† ุจุชุชุนู„ู… ุฃู† ุฃู†ุช โš ๏ธ
MasriSwitch ูุจุชุฏูŠูƒ ูƒู„ ุงู„ soft skills ุจุดูƒู„ indirect ุญุฑููŠุง ุจุชุฏูŠู‡ุงู„ูƒ ุจุงู„ู…ุนู„ุงุฉ ู„ูŠู‡ ุจุชุชุนู„ู… ุชุดุชุบู„ ุนู„ู‰ ูƒู„ ุงู„ุญุงุฌุงุช ุจุชุงุนุฉ Microsoft Excel, PowerPoint, Word whatever ุจุนุฏูŠู† ูƒู…ุงู† ุจุชุชุนู„ู… ุฃู† ุฃู†ุช โš ๏ธ

๐Ÿ‡ช๐Ÿ‡ฌ Egyptian Arabic

Pure-Arabic clips (no English). After normalization these reflect genuine word-level accuracy, not spelling/number-format differences.

Short

โœ… Ground truth โ€” ุงู„ุดู‡ุฑุฉ ู…ุฑุช ุจู…ุฑุญู„ุชูŠู† ู…ุนุงูƒ

Model Output
๐Ÿฅ‡ QwenCleo ุงู„ุดู‡ุฑุฉ ู…ุฑุช ุจู…ุฑุญู„ุชูŠู† ู…ุนุงูƒ โœ…
Qwen3-ASR (base) ุงู„ุดู‡ุฑุฉ ู…ุฑุช ุจู…ุฑุญู„ุชูŠู† ู…ุนูƒ โš ๏ธ
Cohere ุงู„ุดู‡ุฑู‡ ู…ุฑุช ุจู…ุฑุญู„ุชูŠู† ู…ุนุงูƒ โœ…
Nemotron ุงู„ุดู‡ุฑุฉ ู…ุฑุช ุจู…ุฑุญู„ุชูŠู† ู…ุนูƒ โš ๏ธ
Arabic-Whisper-CS ุงู„ุดู‡ุฑุฉ ู…ุฑุช ุจู…ุฑุญู„ุชูŠู† ู…ุนุงูƒ โœ…
MasriSwitch ุงู„ุดู‡ุฑุฉ ู…ุฑุช ุจู…ุฑุญู„ุชูŠู† ู…ุนุงูƒ โœ…

Medium

โœ… Ground truth โ€” ู„ุง ุฑูƒุฒ ุนุดุงู† ุงู†ุช ุฏู„ูˆู‚ุชูŠ ู‡ุชุชุญุงุณุจ ุนู„ู‰ ุงู„ูƒู„ุงู… ุฏู‡

Model Output
๐Ÿฅ‡ QwenCleo ู„ุง ุฑูƒุฒ ุนุดุงู† ุงู†ุช ุฏู„ูˆู‚ุชูŠ ู‡ุชุชุญุงุณุจ ุนู„ู‰ ุงู„ูƒู„ุงู… ุฏู‡ โœ…
Qwen3-ASR (base) ู„ุง ุฑูƒุฒ ุนุดุงู† ุฃู†ุช ุฏู„ูˆู‚ุชูŠ ู‡ุชุญุณุจ ุนู„ู‰ ูƒู„ุงู… ุฏู‡ โš ๏ธ
Cohere ู„ุง ุฑูƒุฒ ุนุดุงู† ุงู†ุช ุฏู„ูˆู‚ุชูŠ ู‡ุชุชุญุงุณุจ ุนุงู„ูƒู„ุงู… ุฏู‡ โœ…
Nemotron ู„ุฃ ุฑูƒุฒ ุนู„ู‰ ุดุงู† ุฃู†ุช ุฏูŠ ุงู„ูˆู‚ุช ู‡ุชุชุญุณุจ ุนู„ู‰ ุงู„ูƒู„ุงู… ุฏู‡ โŒ
Arabic-Whisper-CS ู„ุฃ ุฑูƒุฒ ุนุดุงู† ุงู†ุช ุฏู„ูˆู‚ุช ู‡ุชุชุญุณุจ ุนู„ู‰ ุงู„ูƒู„ุงู… ุฏู‡ โš ๏ธ
MasriSwitch ู„ุฃ ุฑูƒุฒ ุนุดุงู† ุงู†ุช ุฏู„ูˆู‚ุชูŠ ู‡ุชุชุญุงุณุจ ุนู„ู‰ ุงู„ูƒู„ุงู… ุฏู‡ โœ…

โœ… Ground truth โ€” ุจุตุฑุงุญุฉ ูƒุงู†ุช ู…ู† ุฃุณุนุฏ ุฃูŠุงู… ุญูŠุงุชูŠ ู„ู…ุง ุดูุช ุงู„ู„ูŠ ู‡ู…ุง ูุฑุญุงู†ูŠู†

Model Output
๐Ÿฅ‡ QwenCleo ุจุตุฑุงุญุฉ ูƒุงู†ุช ู…ู† ุฃุณุนุฏ ุฃูŠุงู… ุญูŠุงุชูŠ ู„ู…ุง ุดูุช ุงู„ู„ูŠ ู‡ู…ุง ูุฑุญุงู†ูŠู† โœ…
Qwen3-ASR (base) ุจุตุฑุงุญุฉ ูƒู†ุช ู…ู† ุฃุณุนุฏ ุฃูŠุงู… ุญูŠุงุชูŠ ู„ู…ุง ุดูˆูุช ุงู„ู„ูŠ ู‡ู… ูุฑุญุงู†ูŠ โŒ
Cohere ุจุตุฑุงุญู‡ ูƒุงู†ุช ู‚ุจู„ ุงุณุนุฏ ู‚ูŠู… ุญูŠุงุชูŠ ู„ู…ุง ุดูˆูุช ุงู„ู„ูŠ ู‡ู… ูุฑุญุงู†ูŠู† โŒ
Nemotron ุจุตุฑุงุญุฉ ูƒุงู†ุช ู…ู† ุฃุณุนุฏ ู‚ุงู… ุญุงุชูŠ ู„ู…ุง ุดูุช ุงู„ู„ูŠ ู‡ู… ูุฑุญูŠู† โŒ
Arabic-Whisper-CS ุจุตุฑุงุญุฉ ูƒุงู†ุช ุฃุจู†ุฉ ุฃุณุนุฏ ุฃู‚ุงู… ุญูŠุงุชูŠ ู„ู…ุง ุดูˆูุช ุงู„ู„ูŠ ู‡ู… ูุฑุญุงู†ูŠู† โŒ
MasriSwitch ุจุตุฑุงุญุฉ ูƒุงู†ุช ู…ู† ุฃุณุนุฏ ุฃูŠุงู… ุญูŠุงุชูŠ ู„ู…ุง ุดูุช ุงู„ู„ูŠ ู‡ู… ูุฑุญุงู†ูŠู† โœ…

Long

โœ… Ground truth โ€” ุงู†ุง ุงู„ู†ู‡ุงุฑุฏุฉ ู…ุซู„ุง ุณุงูƒู† ููŠ ู…ุฏูŠู†ุฉ ู†ุตุฑ ูˆุดุบู„ูŠ ููŠ ุงู„ู…ุนุงุฏูŠ ูุงู†ุง ู…ู† ุงู„ุจูŠุช ู„ู„ุดุบู„ ุจู‚ุทุน 20 ูƒูŠู„ูˆ ููŠ ุงู„ูŠูˆู…

Model Output
๐Ÿฅ‡ QwenCleo ุงู†ุง ุงู„ู†ู‡ุงุฑุฏู‡ ู…ุซู„ุง ุณุงูƒู† ููŠ ู…ุฏูŠู†ุฉ ู†ุตุฑ ูˆุดุบู„ูŠ ููŠ ุงู„ู…ุนุงุฏูŠ ูุงู†ุง ู…ู† ุงู„ุจูŠุช ู„ู„ุดุบู„ ุจู‚ุทุน 20 ูƒูŠู„ูˆ ููŠ ุงู„ูŠูˆู… โœ…
Qwen3-ASR (base) ุฃู†ุง ุงู„ู†ู‡ุงุฑุฏุฉ ู…ุซู„ุงู‹ ุณุงูƒู† ููŠ ู…ุฏูŠู†ุฉ ู†ุตุฑ ูˆุดุบู„ูŠ ููŠ ุงู„ู…ุนุงุฏูŠ ูุฃู†ุง ู…ู† ุงู„ุจูŠุช ู„ู„ุดุบู„ ู„ุจู‚ูŠุช 20 ูƒูŠู„ูˆ ููŠ ุงู„ูŠูˆู… โš ๏ธ
Cohere ุงู†ุง ุงู„ู†ู‡ุงุฑุฏู‡ ู…ุซู„ุง ุณุงูƒู† ููŠ ู…ุฏูŠู†ู‡ ู…ุตุฑ ูˆุดุบู„ูŠ ููŠ ุงู„ู…ุนุงุฏูŠ ูุงู†ุง ู…ู† ุงู„ุจูŠุช ู„ู„ุดุบู„ ุจู‚ุทุน 20 ูƒูŠู„ูˆ ููŠ ุงู„ูŠูˆู… โš ๏ธ
Nemotron ุฃู†ุง ุงู„ู†ู‡ุงุฑ ุฏู‡ ู…ุซู„ุงู‹ ุณุงูƒู† ููŠ ู…ุฏูŠู†ุฉ ู†ุตุฑ ูˆุดุบู„ูŠ ููŠ ุงู„ู…ุนุงุฏูŠ ูุฃู†ุง ู…ู† ุงู„ุจูŠุช ู„ู„ุดุบู„ ุจู‚ุทุน 20 ูƒูŠู„ูˆ ููŠ ุงู„ูŠูˆู… โœ…
Arabic-Whisper-CS ุฃู†ุง ุงู„ู†ู‡ุงุฑุฏุฉ ู…ุซู„ุง ุณุงูƒู† ููŠ ู…ุฏูŠู†ุฉ ู†ุตุฑ ูˆ ุดุบู„ูŠ ููŠ ุงู„ู…ุนุงุฏูŠ ูุงู†ุง ู…ู† ุงู„ุจูŠุช ู„ู„ุดุบู„ ุจู‚ุทุน 20 ูƒูŠู„ูˆ ููŠ ุงู„ูŠูˆู… โœ…
MasriSwitch ุฃู†ุง ุงู„ู†ู‡ุงุฑุฏุฉ ู…ุซู„ุง ุณุงูƒู† ููŠ ู…ุฏูŠู†ุฉ ู†ุตุฑ ูˆุดุบู„ูŠ ููŠ ุงู„ู…ุนุงุฏูŠ ูุฃู†ุง ู…ู† ุงู„ุจูŠุช ู„ู„ุดุบู„ ุจู‚ุทุน 20 ูƒูŠู„ูˆ ููŠ ุงู„ูŠูˆู… โœ…

โœ… Ground truth โ€” ุงุดุชุบู„ุช ูƒุงู† ููŠู‡ ู…ุนุงู†ุงุฉ ููŠ ุชุญุถูŠุฑ ุงู„ู…ู†ู‡ุฌ

Model Output
๐Ÿฅ‡ QwenCleo ุงุดุชุบู„ุช ูƒุงู† ููŠู‡ ู…ุนุงู†ุงุฉ ููŠ ุชุญุถูŠุฑ ุงู„ู…ู†ู‡ุฌ โœ…
Qwen3-ASR (base) ุงุดุชุบู„ุช ูƒุงู† ููŠ ู…ุนุงู†ุงุฉ ููŠ ุชุญุถูŠุฑ ุงู„ู…ู†ุงูƒ โŒ
Cohere ุงุดุชุบู„ุช ูƒุงู† ููŠ ู…ุนุงู†ุงู‡ ููŠ ุชุญุถูŠุฑ ุงู„ู…ู†ุนูƒุณ โŒ
Nemotron ุขู‡ ุงุณุชู‚ุงู„ุช ูƒุงู† ููŠ ู…ุนู†ุงู‡ ููŠ ุชุญุถูŠุฑ ุงู„ู…ู†ุงูƒ โŒ
Arabic-Whisper-CS ุงุดุชุบู„ุช ูƒุงู† ููŠ ู…ุนู†ุงู‡ ููŠ ุชุญุถูŠุฑ ุงู„ู…ู†ู‡ุฌ โœ…
MasriSwitch ุฅุดุชุบู„ุช ูƒุงู† ููŠ ู…ุนู†ูŠุฉ ููŠ ุชุญุถูŠุฑ ุงู„ู…ู†ู‡ุฌ โŒ

๐Ÿ“ฆ Installation

Install the right torch first. A plain pip install pulls the newest torch (built for the latest CUDA), which fails on older drivers with "NVIDIA driver too old". Install a torch build matching your driver before the package, then add QwenCleo with --no-deps so torch is never reinstalled.

Pick the wheel index for your CUDA driver โ€” cu121 (driver โ‰ฅ 12.1, e.g. CUDA 12.2), cu118 (driver โ‰ฅ 11.8), or cpu. Check yours with nvidia-smi.

For inference & chunked transcription (PyPI)

conda create -n qwencleo-asr python=3.12 -y
conda activate qwencleo-asr

# 1) torch matching your driver (cu121 shown โ€” change the index for yours)
pip install torch==2.5.1 torchaudio==2.5.1 \
  --index-url https://download.pytorch.org/whl/cu121

# 2) QwenCleo without touching torch, then its remaining deps
pip install qwencleo-asr --no-deps
pip install "qwen-asr>=0.0.6" numpy soundfile huggingface_hub

That's all you need for the Python API and the qwencleo CLI.

For serving / Gradio / vLLM (clone the repo)

conda create -n qwencleo-asr python=3.12 -y
conda activate qwencleo-asr

# 1) torch matching your driver, first
pip install torch==2.5.1 torchaudio==2.5.1 \
  --index-url https://download.pytorch.org/whl/cu121

# 2) the repo (without re-resolving torch) + serving deps
git clone https://github.com/MohammedAly22/qwencleo-asr.git
cd qwencleo-asr
pip install -e . --no-deps
pip install "qwen-asr>=0.0.6" numpy soundfile huggingface_hub
pip install -r requirements-serving.txt

Verify torch sees the GPU before running:

python -c "import torch; print(torch.__version__, torch.cuda.is_available())"
# -> 2.5.1+cu121 True

๐Ÿš€ Usage

Python โ€” basic transcription

from qwencleo_asr import QwenCleoASR

asr = QwenCleoASR()                       # loads mohammedaly22/QwenCleo-ASR
result = asr.transcribe("clip.wav")       # language defaults to "Arabic"
print(result.text)

Batch, auto-detect language, and Egyptian normalization:

results = asr.transcribe(["a.wav", "b.wav"], language=None)   # auto-detect
clean   = asr.transcribe("clip.wav", normalize=True)          # normalized text

Python โ€” chunked transcription of long audio / mic

from qwencleo_asr import QwenCleoASR, stream_file

asr = QwenCleoASR()
for chunk in stream_file(asr, "long_podcast.wav", chunk_s=20, overlap_s=2):
    print(f"[{chunk.start:.0f}-{chunk.end:.0f}s] {chunk.text}")

โ„น๏ธ This is chunked transcription, not true streaming. It splits long/live audio into overlapping windows and transcribes each โ€” convenient for captioning without a server, but latency is per-window. For genuine token-by-token streaming, use the vLLM path below.

Python โ€” true streaming (vLLM)

QwenCleo inherits Qwen3-ASR's real token-by-token streaming via vLLM. Two one-time setup commands, then stream from Python.

qwencleo install-vllm     # installs the vLLM nightly (cu129) โ€” the only build
                          # with Qwen3-ASR support (not on PyPI; uses uv)
qwencleo serve            # launches the server (sets the right flags for you:
                          # VLLM_USE_FLASHINFER_SAMPLER=0, --gpu-memory-utilization 0.8)

Needs an Ampere-or-newer GPU (L4 / A100 / H100). See server/vllm_serve.md for details and manual flags.

Then stream straight off the model object โ€” deltas arrive as they're generated (the language X<asr_text> prefix is stripped for you):

from qwencleo_asr import QwenCleoASR

asr = QwenCleoASR()
for delta in asr.stream("clip.wav", port=8000):   # talks to the vLLM server
    print(delta, end="", flush=True)

Or use the helpers directly:

from qwencleo_asr import stream_vllm, transcribe_vllm, VLLMOffline

for delta in stream_vllm("clip.wav", port=8000, language="Arabic"):
    print(delta, end="", flush=True)

print(transcribe_vllm("clip.wav", port=8000))   # one-shot via the server
print(VLLMOffline().transcribe("clip.wav"))      # in-process, no server

From the shell:

qwencleo stream-vllm clip.wav --port 8000        # token-by-token to stdout

CLI

qwencleo transcribe clip.wav
qwencleo transcribe a.wav b.wav --language None --normalize
qwencleo stream long_podcast.wav --chunk-s 20 --overlap-s 2

๐ŸŒ Serving

FastAPI server

QWENCLEO_MODEL=mohammedaly22/QwenCleo-ASR \
uvicorn server.app:app --host 0.0.0.0 --port 8000

curl -X POST http://localhost:8000/v1/transcribe -F file=@clip.wav -F language=Arabic

Gradio demo

python app/gradio_app.py        # http://localhost:7860  (mic + file upload)

vLLM โ€” serving, streaming & OpenAI-compatible API

Full guide in server/vllm_serve.md. In short:

qwencleo install-vllm      # vLLM nightly (cu129) โ€” the only build with Qwen3-ASR support
qwencleo serve             # OpenAI-compatible server on :8000

OpenAI-compatible transcription:

from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
print(client.audio.transcriptions.create(
    model="mohammedaly22/QwenCleo-ASR", file=open("clip.wav","rb").read()).text)

Streaming mic web demo

Live browser-mic transcription via the upstream Flask demo:

qwen-asr-demo-streaming \
  --asr-model-path mohammedaly22/QwenCleo-ASR \
  --host 0.0.0.0 --port 8000 --gpu-memory-utilization 0.9
# open http://<your-ip>:8000

๐Ÿ““ Examples (Colab)

Runnable notebooks in examples/ โ€” open one, set the runtime to GPU (Runtime โ†’ Change runtime type โ†’ GPU), and run the cells top to bottom.

Notebook What it shows Open in Colab
Quick Start Install, transcribe, batch, CLI Open In Colab
Chunked transcription Long-audio windowing + mic-style frames Open In Colab
FastAPI server Run the server, call it over HTTP Open In Colab
Gradio demo Browser UI with a public share link Open In Colab
vLLM streaming True token-by-token streaming + OpenAI API Open In Colab

๐Ÿ”— Links


๐Ÿ“œ License & citation

Apache-2.0, inheriting the Qwen3-ASR license terms.

@misc{qwencleo_asr_2026,
  title  = {QwenCleo-ASR: The Best Open-Source Egyptian Arabic and Code-Switching Speech Recognition Model},
  author = {Mohammed Aly},
  year   = {2026},
  howpublished = {\url{https://huggingface.co/mohammedaly22/QwenCleo-ASR}},
  note   = {Fine-tuned from Qwen3-ASR-1.7B}
}

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

qwencleo_asr-0.2.1.tar.gz (27.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

qwencleo_asr-0.2.1-py3-none-any.whl (23.2 kB view details)

Uploaded Python 3

File details

Details for the file qwencleo_asr-0.2.1.tar.gz.

File metadata

  • Download URL: qwencleo_asr-0.2.1.tar.gz
  • Upload date:
  • Size: 27.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.13

File hashes

Hashes for qwencleo_asr-0.2.1.tar.gz
Algorithm Hash digest
SHA256 8a04d16fcb148a8494d6916777a9cd63313e9f8aa4652c5ab0e41857a8269e4f
MD5 d2fd7a3537ac31ac780ec452585b6f43
BLAKE2b-256 fc3571b39154c6a7b96f2af489b90356a238662fc6a1ab7493e3523e06666013

See more details on using hashes here.

File details

Details for the file qwencleo_asr-0.2.1-py3-none-any.whl.

File metadata

  • Download URL: qwencleo_asr-0.2.1-py3-none-any.whl
  • Upload date:
  • Size: 23.2 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.13

File hashes

Hashes for qwencleo_asr-0.2.1-py3-none-any.whl
Algorithm Hash digest
SHA256 b2ba72a23a01abad81974c125857e65800062225e94ddb5b29e2b1f0c314d979
MD5 d43633a2a0dfb7f31fa59360180a4ffa
BLAKE2b-256 0c1db2be522c7aeaa45bd909e14b0c60d12fb28c0b015c84b42a480cd700aafb

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page