QwenCleo-ASR โ Egyptian Arabic & code-switching speech recognition, built on Qwen3-ASR.
Project description
๐๏ธ QwenCleo-ASR
The best open-source model for Egyptian Arabic & code-switching speech recognition
Built on Qwen3-ASR-1.7B, fine-tuned for Egyptian dialect and Arabic โ English code-switching.
QwenCleo โ the name carries three meanings: Qwen, the powerful base model it is built on; Queen, signalling a model that reigns over its domain; and Cleo, for Cleopatra, the queen of Egypt โ because this model is tailored for Egyptian Arabic. ๐๐บ
QwenCleo-ASR is, to our knowledge, the best open-source ASR model for Egyptian Arabic and Arabic/English code-switching. It cuts the error rate of the strong Qwen3-ASR base roughly in half, and correctly keeps embedded English tech/loan words in Latin script (engineering, download, React, at least) instead of mangling them into broken Arabic.
- ๐ฏ Egyptian dialect โ tuned on hundreds of hours of Egyptian podcast speech.
- ๐ Code-switching โ keeps English terms in
code-script, Arabic in Arabic. - ๐ฅ State-of-the-art (open) โ beats Qwen3-ASR base, NVIDIA Nemotron, Cohere, and every Whisper variant on our Egyptian + CS test set.
- ๐ฆ
pip install qwencleo-asrโ inference & chunked long-audio transcription in three lines. - โก Real streaming โ token-by-token via vLLM (
asr.stream(...)), plus a mic web demo. - ๐ Serving โ FastAPI server, Gradio demo, OpenAI-compatible vLLM endpoint.
๐ Results
WER / CER (%) on an Egyptian-Arabic + code-switching test set (3,699 utterances). Lower is better. All models scored with the same Egyptian-aware normalization.
| Rank | Model | Params | WER all | CER all | WER ยท AR | CER ยท AR | WER ยท CS | CER ยท CS |
|---|---|---|---|---|---|---|---|---|
| ๐ฅ | QwenCleo-ASR | 1.7B | 19.85 | 10.52 | 19.08 | 10.30 | 20.75 | 10.81 |
| ๐ฅ | Arabic-Whisper-CodeSwitching | 1.55B | 37.98 | 17.86 | 38.48 | 18.71 | 36.92 | 16.22 |
| ๐ฅ | NVIDIA Nemotron-3.5 | 0.6B | 38.04 | 19.72 | 36.23 | 16.43 | 41.44 | 25.66 |
| 4 | Qwen3-ASR-1.7B (base) | 1.7B | 41.51 | 20.86 | 40.59 | 18.52 | 43.20 | 25.04 |
| 5 | Whisper Large-v3 Turbo (FT) | 0.81B | 50.83 | 22.86 | 48.37 | 18.42 | 55.08 | 37.84 |
| 6 | Cohere Transcribe 03-2026 | 2.0B | 53.20 | 39.12 | 47.95 | 33.45 | 63.76 | 49.66 |
| 7 | MasriSwitch-Gemma3n | 8B | 57.30 | 30.19 | 60.38 | 32.91 | 51.32 | 25.13 |
| 8 | Whisper Large-v3 | 1.54B | 63.94 | 39.76 | 49.25 | 22.76 | 59.32 | 31.52 |
| 9 | Whisper Large-v2 | 1.54B | 72.34 | 48.73 | 60.75 | 33.21 | 66.85 | 40.75 |
| 10 | Whisper Large-v3 Turbo | 0.81B | 73.83 | 46.86 | 59.37 | 29.42 | 66.08 | 37.84 |
| 11 | Whisper Medium | 0.76B | 80.46 | 53.19 | 74.77 | 41.76 | 74.15 | 44.90 |
| 12 | Whisper Small | 0.24B | 89.99 | 60.34 | 77.42 | 42.53 | 87.09 | 55.22 |
| 13 | Whisper Tiny | 0.04B | 124.68 | 89.42 | 116.02 | 77.74 | 110.67 | 74.57 |
๐ฃ๏ธ Sample outputs
Real transcriptions from the test set, scored after Egyptian-aware normalization. Examples span short, medium, and long clips in both Egyptian Arabic and code-switching. These are picked to be honest โ on some clips other models also do well; QwenCleo's edge is consistency, dialect fidelity, and keeping English terms in Latin script.
โ = matches ground truth ยท โ ๏ธ = minor slips ยท โ = clear errors. Bold marks the wrong spans.
๐ Code-switching
Short
โ Ground truth ุงุงู ุจุญุณ ุณุชุงููู ุญูู ูุดููู ุญูู ูุงุฑูุฒู ุง ููู ุจููุนุจ
| Model | Output |
|---|---|
| ๐ฅ QwenCleo | ุจุญุณ style ู ุญูู ูุดููู ุญูู ูุงุฑูุฒู ุง ููู ุจููุนุจ โ ๏ธ |
| Qwen3-ASR (base) | ุฃู ุจุญุณ ุทุงููู ุญูู ูุดููู ุญูู ููุฑุฒู ู ููู ููุนุจ โ |
| Cohere | ุงุงุงุงุงุงุงุง ุจุญุณ ุณุชุงููู ุญูู ูุดููู ุญูู ููุงุฑูุฒู ุง ููู ุจููุนุจ โ |
| Nemotron | ุขู ุจุญุณ ุณุชุงููู ุญูู ูุดููู ุญูู ูุงุฑูุฒู ุง ูู ุจููุนุจ โ |
| Arabic-Whisper-CS | ุจุญุณ style ู ุญูู ูุดูู ูุญูู ู charisma ููู ุจููุนุจ โ ๏ธ |
| MasriSwitch | ุจุญุณ ุณุชุงููู ุญูู ู ุดููู ุญูู ู ูุงุฑูุฒู ุง ููู ุจูุนูู โ |
Medium
โ Ground truth โ ุฏู ุญุตู ุงุฒุงู ูุฑุงุฑ ุงุตูุง ุงูู ุชุนู ู ููุงุฉ
YouTube
| Model | Output |
|---|---|
| ๐ฅ QwenCleo | ุฏู ุญุตู ุงุฒุงู ูุฑุงุฑ ุงุตูุง ุงูู ุชุนู
ู ููุงุฉ YouTube โ
|
| Qwen3-ASR (base) | ุฏู ุญุตู ุฅุฒุงู ูุฑุงุฑ ุฃุตูุง ุฅูู ุชุนู ู ููุงุฉ ููุชููุจ โ ๏ธ |
| Cohere | ุฏู ุญุตู ุงุฒุงู ูุฑุงุฑ ุงุตูุง ุงูู ุชุนู ู ููุงู ููุชููุจ โ ๏ธ |
| Nemotron | ุฏู ุญุตู ุฅุฒุงู ูุฑุงุฑ ุฃุตูุงู ุฅูู ุชุนู ู ููุงุฉ ููุชููุจ โ ๏ธ |
| Arabic-Whisper-CS | ุฏุง ุญุตู ุงุฒุงู ูุฑุงุฑ ุงุตูุง ุงูู ุชุนู ู ููุงุฉ ููุชููุจ โ ๏ธ |
| MasriSwitch | ุฏู ุญุตู ุงุฒุงู ูุฑุงุฑ ุฃุตูุง ุฅูู ุชุนู
ู ููุงุฉ YouTube โ
|
โ Ground truth โ ุฎูููู ุงูุถุญ ุจุณ ูุฏููุง ุงู ุงูุทุงูุจ ูุฌู ูู
already
| Model | Output |
|---|---|
| ๐ฅ QwenCleo | ุฎูููู ุงูุถุญ ุจุณ ูุฏููุง ุงู ุงูุทุงูุจ ูุฌู ูู already โ
|
| Qwen3-ASR (base) | ุฎูููุง ุฃูุถุญ ุจุณ ูุฏูู ุฅู ุงูุทุงูุจ ูุฌู ูู ุฃุฑุฑูุฏู โ |
| Cohere | ุฎูููู ุงูุถุญ ุจุณ ูุฏููุง ุงู ุงูุทุงูุจ ูุฌู ูู ุงูุฑุฏู โ |
| Nemotron | ุฎูููุง ูุถุญ ุจุณ ูุฏููุง ุฅู ุงูุทุงูุจ ูุฌู ูู ุฃูุฑุฏู โ |
| Arabic-Whisper-CS | ุฎูููู ุงูุถุญ ุจุณ ูุฏููุง ุงู ุงูุทุงูุจ ููุฌู ูู already โ
|
| MasriSwitch | ุฎูููู ุฃูุถุญ ุจุณ ุงูุฃูุฏุงู ุฅู ุงูุทุงูุจ ุจูุฌู ูู already โ ๏ธ |
Long
โ Ground truth โ ู ุง ูู ุงูููุฑุฉ ุฌุงูุฉ ููู ุงูููุงุฑุฏู
averageุงูุนุฑุจูุฉ ุงููู ู ู ูู ุชุฌูุจูุง ู ู ุฃูุฑูุจุง ู ุง ุจูู 23 ู 28000dollarsุชูุฑูุจุง
| Model | Output |
|---|---|
| ๐ฅ QwenCleo | ู
ุง ูู ุงูููุฑู ุฌุงูู ููู ุงูููุงุฑุฏู average ุงูุนุฑุจูู ุงููู ู
ู
ูู ุชุฌูุจูุง ู
ู ุงูุฑูุจุง ู
ุง ุจูู 23 ู 28000 dollars โ ๏ธ |
| Qwen3-ASR (base) | ู ุง ูู ุงูููุฑุฉ ุฌุงูุฉ ููู ุงูููุงุฑุฏุฉ ุฃูุฑุฌ ุงูุนุฑุจูุฉ ุงููู ู ู ูู ุชุฌุงุจูุง ู ู ุฃูุฑูุจุง ู ุง ุจูู 23 ู 28000 ุฏููุงุฑ โ ๏ธ |
| Cohere | ู ุง ูู ุงูููุฑู ุฌุงูู ููู ุงูููุงุฑุฏู ุงูุฑูููุง ุงูุนุฑุจูู ุงููู ู ู ูู ุชุฌูุจูุง ู ู ุงูุฑูุจุง ู ุง ุจูู 23 ู 28000 ุฏููุงุฑ โ |
| Nemotron | ู ุง ูู ุงูููุฑุฉ ุฌุงูุฉ ููู ุงูููุงุฑ ุฏู ุฃูุฑูุฏ ุงูุนุฑุจูุฉ ุงููู ู ู ูู ุชุฌูุจูุง ู ู ุฃูุฑูุจุง ู ุง ุจูู 23 ู 28000 ุฏููุงุฑ ุซูุงุซุฉ โ |
| Arabic-Whisper-CS | ู
ุง ูู ุงูููุฑุฉ ุฌุงูุฉ ููู ุงูููุงุฑ ุฏุง average ุงูุนุฑุจูุฉ ุงููู ู
ู
ูู ุชุฌูุจูุง ู
ู ุฃูุฑูุจุง ู
ุง ุจูู 23 ู 28 ุฃูู ุฏููุงุฑ ุชูุฑูุจุง โ ๏ธ |
| MasriSwitch | ู
ุง ูู ุงูููุฑุฉ ุฌุงูุฉ ููู ุงูููุงุฑุฏุฉ average ุงูุนุฑุจูุฉ ุงููู ู
ู
ูู ุชุฌูุจูุง ู
ู ุฃูุฑูุจุง ู
ุง ุจูู 23 ูู28 ุฃูู ุฏููุงุฑ ุชูุฑูุจูุง โ ๏ธ |
โ Ground truth โ ูุจุชุฏูู ูู ุงู
soft skillsุจุดููindirectุญุฑููุง ุจุชุฏููุงูู ุจุงูู ุนููู ููู ุจุชุชุนูู ุชุดุชุบู ุนูู ูู ุงูุญุงุฌุงุช ุจุชุงุนูMicrosoft Excel PowerPoint Word whateverุจุนุฏูู ูู ุงู ุจุชุชุนูู ุงู ุงูุช ุชุชููู ู ุน ุงููุงุณpublic speaking
| Model | Output |
|---|---|
| ๐ฅ QwenCleo | ูุจุชุฏูู ูู ุงูsoft skills ุจุดูู indirect ุญุฑููุง ุจุชุฏููุงูู ุจุงูู
ุนููู ููู ุจุชุชุนูู
ุชุดุชุบู ุนูู ูู ุงูุญุงุฌุงุช ุจุชุงุนู Microsoft Excel, PowerPoint, Word ู whatever ุจุนุฏูู ูู
ุงู ุจุชุชุนูู
ุงู ุงูุช ุชุชููู
ู
ุน ุงููุงุณ public speaking โ
|
| Arabic-Whisper-CS | ูุจุชุฏูู ูู ุงู soft skills ุจุดูู indirect ุญุฑููุง ุจุชุฏููุงูู ุจุงูู
ุนููุฉ ููู ุจุชุชุนูู
ุชุดุชุบู ุนูู ูู ุงูุญุงุฌุงุช ุจุชุงุนุฉ Microsoft, Excel, PowerPoint, Word whatever ุจุนุฏูู ูู
ุงู ุจุชุชุนูู
ุฃู ุฃูุช โ ๏ธ |
| MasriSwitch | ูุจุชุฏูู ูู ุงู soft skills ุจุดูู indirect ุญุฑููุง ุจุชุฏููุงูู ุจุงูู
ุนูุงุฉ ููู ุจุชุชุนูู
ุชุดุชุบู ุนูู ูู ุงูุญุงุฌุงุช ุจุชุงุนุฉ Microsoft Excel, PowerPoint, Word whatever ุจุนุฏูู ูู
ุงู ุจุชุชุนูู
ุฃู ุฃูุช โ ๏ธ |
๐ช๐ฌ Egyptian Arabic
Pure-Arabic clips (no English). After normalization these reflect genuine word-level accuracy, not spelling/number-format differences.
Short
โ Ground truth โ ุงูุดูุฑุฉ ู ุฑุช ุจู ุฑุญูุชูู ู ุนุงู
| Model | Output |
|---|---|
| ๐ฅ QwenCleo | ุงูุดูุฑุฉ ู ุฑุช ุจู ุฑุญูุชูู ู ุนุงู โ |
| Qwen3-ASR (base) | ุงูุดูุฑุฉ ู ุฑุช ุจู ุฑุญูุชูู ู ุนู โ ๏ธ |
| Cohere | ุงูุดูุฑู ู ุฑุช ุจู ุฑุญูุชูู ู ุนุงู โ |
| Nemotron | ุงูุดูุฑุฉ ู ุฑุช ุจู ุฑุญูุชูู ู ุนู โ ๏ธ |
| Arabic-Whisper-CS | ุงูุดูุฑุฉ ู ุฑุช ุจู ุฑุญูุชูู ู ุนุงู โ |
| MasriSwitch | ุงูุดูุฑุฉ ู ุฑุช ุจู ุฑุญูุชูู ู ุนุงู โ |
Medium
โ Ground truth โ ูุง ุฑูุฒ ุนุดุงู ุงูุช ุฏูููุชู ูุชุชุญุงุณุจ ุนูู ุงูููุงู ุฏู
| Model | Output |
|---|---|
| ๐ฅ QwenCleo | ูุง ุฑูุฒ ุนุดุงู ุงูุช ุฏูููุชู ูุชุชุญุงุณุจ ุนูู ุงูููุงู ุฏู โ |
| Qwen3-ASR (base) | ูุง ุฑูุฒ ุนุดุงู ุฃูุช ุฏูููุชู ูุชุญุณุจ ุนูู ููุงู ุฏู โ ๏ธ |
| Cohere | ูุง ุฑูุฒ ุนุดุงู ุงูุช ุฏูููุชู ูุชุชุญุงุณุจ ุนุงูููุงู ุฏู โ |
| Nemotron | ูุฃ ุฑูุฒ ุนูู ุดุงู ุฃูุช ุฏู ุงูููุช ูุชุชุญุณุจ ุนูู ุงูููุงู ุฏู โ |
| Arabic-Whisper-CS | ูุฃ ุฑูุฒ ุนุดุงู ุงูุช ุฏูููุช ูุชุชุญุณุจ ุนูู ุงูููุงู ุฏู โ ๏ธ |
| MasriSwitch | ูุฃ ุฑูุฒ ุนุดุงู ุงูุช ุฏูููุชู ูุชุชุญุงุณุจ ุนูู ุงูููุงู ุฏู โ |
โ Ground truth โ ุจุตุฑุงุญุฉ ูุงูุช ู ู ุฃุณุนุฏ ุฃูุงู ุญูุงุชู ูู ุง ุดูุช ุงููู ูู ุง ูุฑุญุงููู
| Model | Output |
|---|---|
| ๐ฅ QwenCleo | ุจุตุฑุงุญุฉ ูุงูุช ู ู ุฃุณุนุฏ ุฃูุงู ุญูุงุชู ูู ุง ุดูุช ุงููู ูู ุง ูุฑุญุงููู โ |
| Qwen3-ASR (base) | ุจุตุฑุงุญุฉ ููุช ู ู ุฃุณุนุฏ ุฃูุงู ุญูุงุชู ูู ุง ุดููุช ุงููู ูู ูุฑุญุงูู โ |
| Cohere | ุจุตุฑุงุญู ูุงูุช ูุจู ุงุณุนุฏ ููู ุญูุงุชู ูู ุง ุดููุช ุงููู ูู ูุฑุญุงููู โ |
| Nemotron | ุจุตุฑุงุญุฉ ูุงูุช ู ู ุฃุณุนุฏ ูุงู ุญุงุชู ูู ุง ุดูุช ุงููู ูู ูุฑุญูู โ |
| Arabic-Whisper-CS | ุจุตุฑุงุญุฉ ูุงูุช ุฃุจูุฉ ุฃุณุนุฏ ุฃูุงู ุญูุงุชู ูู ุง ุดููุช ุงููู ูู ูุฑุญุงููู โ |
| MasriSwitch | ุจุตุฑุงุญุฉ ูุงูุช ู ู ุฃุณุนุฏ ุฃูุงู ุญูุงุชู ูู ุง ุดูุช ุงููู ูู ูุฑุญุงููู โ |
Long
โ Ground truth โ ุงูุง ุงูููุงุฑุฏุฉ ู ุซูุง ุณุงูู ูู ู ุฏููุฉ ูุตุฑ ูุดุบูู ูู ุงูู ุนุงุฏู ูุงูุง ู ู ุงูุจูุช ููุดุบู ุจูุทุน 20 ูููู ูู ุงูููู
| Model | Output |
|---|---|
| ๐ฅ QwenCleo | ุงูุง ุงูููุงุฑุฏู ู ุซูุง ุณุงูู ูู ู ุฏููุฉ ูุตุฑ ูุดุบูู ูู ุงูู ุนุงุฏู ูุงูุง ู ู ุงูุจูุช ููุดุบู ุจูุทุน 20 ูููู ูู ุงูููู โ |
| Qwen3-ASR (base) | ุฃูุง ุงูููุงุฑุฏุฉ ู ุซูุงู ุณุงูู ูู ู ุฏููุฉ ูุตุฑ ูุดุบูู ูู ุงูู ุนุงุฏู ูุฃูุง ู ู ุงูุจูุช ููุดุบู ูุจููุช 20 ูููู ูู ุงูููู โ ๏ธ |
| Cohere | ุงูุง ุงูููุงุฑุฏู ู ุซูุง ุณุงูู ูู ู ุฏููู ู ุตุฑ ูุดุบูู ูู ุงูู ุนุงุฏู ูุงูุง ู ู ุงูุจูุช ููุดุบู ุจูุทุน 20 ูููู ูู ุงูููู โ ๏ธ |
| Nemotron | ุฃูุง ุงูููุงุฑ ุฏู ู ุซูุงู ุณุงูู ูู ู ุฏููุฉ ูุตุฑ ูุดุบูู ูู ุงูู ุนุงุฏู ูุฃูุง ู ู ุงูุจูุช ููุดุบู ุจูุทุน 20 ูููู ูู ุงูููู โ |
| Arabic-Whisper-CS | ุฃูุง ุงูููุงุฑุฏุฉ ู ุซูุง ุณุงูู ูู ู ุฏููุฉ ูุตุฑ ู ุดุบูู ูู ุงูู ุนุงุฏู ูุงูุง ู ู ุงูุจูุช ููุดุบู ุจูุทุน 20 ูููู ูู ุงูููู โ |
| MasriSwitch | ุฃูุง ุงูููุงุฑุฏุฉ ู ุซูุง ุณุงูู ูู ู ุฏููุฉ ูุตุฑ ูุดุบูู ูู ุงูู ุนุงุฏู ูุฃูุง ู ู ุงูุจูุช ููุดุบู ุจูุทุน 20 ูููู ูู ุงูููู โ |
โ Ground truth โ ุงุดุชุบูุช ูุงู ููู ู ุนุงูุงุฉ ูู ุชุญุถูุฑ ุงูู ููุฌ
| Model | Output |
|---|---|
| ๐ฅ QwenCleo | ุงุดุชุบูุช ูุงู ููู ู ุนุงูุงุฉ ูู ุชุญุถูุฑ ุงูู ููุฌ โ |
| Qwen3-ASR (base) | ุงุดุชุบูุช ูุงู ูู ู ุนุงูุงุฉ ูู ุชุญุถูุฑ ุงูู ูุงู โ |
| Cohere | ุงุดุชุบูุช ูุงู ูู ู ุนุงูุงู ูู ุชุญุถูุฑ ุงูู ูุนูุณ โ |
| Nemotron | ุขู ุงุณุชูุงูุช ูุงู ูู ู ุนูุงู ูู ุชุญุถูุฑ ุงูู ูุงู โ |
| Arabic-Whisper-CS | ุงุดุชุบูุช ูุงู ูู ู ุนูุงู ูู ุชุญุถูุฑ ุงูู ููุฌ โ |
| MasriSwitch | ุฅุดุชุบูุช ูุงู ูู ู ุนููุฉ ูู ุชุญุถูุฑ ุงูู ููุฌ โ |
๐ฆ Installation
Install the right torch first. A plain
pip installpulls the newest torch (built for the latest CUDA), which fails on older drivers with "NVIDIA driver too old". Install a torch build matching your driver before the package, then add QwenCleo with--no-depsso torch is never reinstalled.Pick the wheel index for your CUDA driver โ
cu121(driver โฅ 12.1, e.g. CUDA 12.2),cu118(driver โฅ 11.8), orcpu. Check yours withnvidia-smi.
For inference & chunked transcription (PyPI)
conda create -n qwencleo-asr python=3.12 -y
conda activate qwencleo-asr
# 1) torch matching your driver (cu121 shown โ change the index for yours)
pip install torch==2.5.1 torchaudio==2.5.1 \
--index-url https://download.pytorch.org/whl/cu121
# 2) QwenCleo without touching torch, then its remaining deps
pip install qwencleo-asr --no-deps
pip install "qwen-asr>=0.0.6" numpy soundfile huggingface_hub
That's all you need for the Python API and the qwencleo CLI.
For serving / Gradio / vLLM (clone the repo)
conda create -n qwencleo-asr python=3.12 -y
conda activate qwencleo-asr
# 1) torch matching your driver, first
pip install torch==2.5.1 torchaudio==2.5.1 \
--index-url https://download.pytorch.org/whl/cu121
# 2) the repo (without re-resolving torch) + serving deps
git clone https://github.com/MohammedAly22/qwencleo-asr.git
cd qwencleo-asr
pip install -e . --no-deps
pip install "qwen-asr>=0.0.6" numpy soundfile huggingface_hub
pip install -r requirements-serving.txt
Verify torch sees the GPU before running:
python -c "import torch; print(torch.__version__, torch.cuda.is_available())"
# -> 2.5.1+cu121 True
๐ Usage
Python โ basic transcription
from qwencleo_asr import QwenCleoASR
asr = QwenCleoASR() # loads mohammedaly22/QwenCleo-ASR
result = asr.transcribe("clip.wav") # language defaults to "Arabic"
print(result.text)
Batch, auto-detect language, and Egyptian normalization:
results = asr.transcribe(["a.wav", "b.wav"], language=None) # auto-detect
clean = asr.transcribe("clip.wav", normalize=True) # normalized text
Python โ chunked transcription of long audio / mic
from qwencleo_asr import QwenCleoASR, stream_file
asr = QwenCleoASR()
for chunk in stream_file(asr, "long_podcast.wav", chunk_s=20, overlap_s=2):
print(f"[{chunk.start:.0f}-{chunk.end:.0f}s] {chunk.text}")
โน๏ธ This is chunked transcription, not true streaming. It splits long/live audio into overlapping windows and transcribes each โ convenient for captioning without a server, but latency is per-window. For genuine token-by-token streaming, use the vLLM path below.
Python โ true streaming (vLLM)
QwenCleo inherits Qwen3-ASR's real token-by-token streaming via vLLM. Two one-time setup commands, then stream from Python.
qwencleo install-vllm # installs the vLLM nightly (cu129) โ the only build
# with Qwen3-ASR support (not on PyPI; uses uv)
qwencleo serve # launches the server (sets the right flags for you:
# VLLM_USE_FLASHINFER_SAMPLER=0, --gpu-memory-utilization 0.8)
Needs an Ampere-or-newer GPU (L4 / A100 / H100). See
server/vllm_serve.mdfor details and manual flags.
Then stream straight off the model object โ deltas arrive as they're generated
(the language X<asr_text> prefix is stripped for you):
from qwencleo_asr import QwenCleoASR
asr = QwenCleoASR()
for delta in asr.stream("clip.wav", port=8000): # talks to the vLLM server
print(delta, end="", flush=True)
Or use the helpers directly:
from qwencleo_asr import stream_vllm, transcribe_vllm, VLLMOffline
for delta in stream_vllm("clip.wav", port=8000, language="Arabic"):
print(delta, end="", flush=True)
print(transcribe_vllm("clip.wav", port=8000)) # one-shot via the server
print(VLLMOffline().transcribe("clip.wav")) # in-process, no server
From the shell:
qwencleo stream-vllm clip.wav --port 8000 # token-by-token to stdout
CLI
qwencleo transcribe clip.wav
qwencleo transcribe a.wav b.wav --language None --normalize
qwencleo stream long_podcast.wav --chunk-s 20 --overlap-s 2
๐ Serving
FastAPI server
QWENCLEO_MODEL=mohammedaly22/QwenCleo-ASR \
uvicorn server.app:app --host 0.0.0.0 --port 8000
curl -X POST http://localhost:8000/v1/transcribe -F file=@clip.wav -F language=Arabic
Gradio demo
python app/gradio_app.py # http://localhost:7860 (mic + file upload)
vLLM โ serving, streaming & OpenAI-compatible API
Full guide in server/vllm_serve.md. In short:
qwencleo install-vllm # vLLM nightly (cu129) โ the only build with Qwen3-ASR support
qwencleo serve # OpenAI-compatible server on :8000
OpenAI-compatible transcription:
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
print(client.audio.transcriptions.create(
model="mohammedaly22/QwenCleo-ASR", file=open("clip.wav","rb").read()).text)
Streaming mic web demo
Live browser-mic transcription via the upstream Flask demo:
qwen-asr-demo-streaming \
--asr-model-path mohammedaly22/QwenCleo-ASR \
--host 0.0.0.0 --port 8000 --gpu-memory-utilization 0.9
# open http://<your-ip>:8000
๐ Examples (Colab)
Runnable notebooks in examples/ โ open one, set the runtime to
GPU (Runtime โ Change runtime type โ GPU), and run the cells top to bottom.
๐ Links
- ๐ค Model card:
mohammedaly22/QwenCleo-ASR - ๐ฆ PyPI:
qwencleo-asr - ๐งฑ Base model:
Qwen/Qwen3-ASR-1.7Bยท Qwen3-ASR repo - Languages: Egyptian Arabic, Modern Standard Arabic, ArabicโEnglish code-switching
- Recommended
languagehint:"Arabic"(orNoneto auto-detect)
๐ License & citation
Apache-2.0, inheriting the Qwen3-ASR license terms.
@misc{qwencleo_asr_2026,
title = {QwenCleo-ASR: The Best Open-Source Egyptian Arabic and Code-Switching Speech Recognition Model},
author = {Mohammed Aly},
year = {2026},
howpublished = {\url{https://huggingface.co/mohammedaly22/QwenCleo-ASR}},
note = {Fine-tuned from Qwen3-ASR-1.7B}
}
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file qwencleo_asr-0.2.1.tar.gz.
File metadata
- Download URL: qwencleo_asr-0.2.1.tar.gz
- Upload date:
- Size: 27.8 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.13
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
8a04d16fcb148a8494d6916777a9cd63313e9f8aa4652c5ab0e41857a8269e4f
|
|
| MD5 |
d2fd7a3537ac31ac780ec452585b6f43
|
|
| BLAKE2b-256 |
fc3571b39154c6a7b96f2af489b90356a238662fc6a1ab7493e3523e06666013
|
File details
Details for the file qwencleo_asr-0.2.1-py3-none-any.whl.
File metadata
- Download URL: qwencleo_asr-0.2.1-py3-none-any.whl
- Upload date:
- Size: 23.2 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.13
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
b2ba72a23a01abad81974c125857e65800062225e94ddb5b29e2b1f0c314d979
|
|
| MD5 |
d43633a2a0dfb7f31fa59360180a4ffa
|
|
| BLAKE2b-256 |
0c1db2be522c7aeaa45bd909e14b0c60d12fb28c0b015c84b42a480cd700aafb
|