Local voice dictation for macOS (Apple Silicon). Tap a key, speak, tap again — your words are typed into whatever window has focus. 100% on-device: audio never leaves your machine.
Features
- Fully local & offline — MLX-accelerated Qwen3-ASR on the Apple Neural stack; works on a plane, nothing is uploaded, no API keys.
- Tap-to-dictate — single tap of right ⌥ Option starts/stops recording (the same pattern superwhisper and Wispr Flow use). No always-on mic.
- Types anywhere — text is delivered as native keyboard events (Quartz CGEvent unicode), so it works in any app, editor, terminal, or chat box.
- 30 languages — English by default; switch instantly (e.g.
--language Turkish) with no model reload. - Voice commands — "period", "comma", "question mark", "new line",
"send" (presses Enter). A Turkish command set activates with
--language Turkish. - Cleanup & polish — vocalized fillers ("um", "uh", "eee") are removed
automatically; optional
--polishruns a small local LLM that drops contextual fillers ("you know", "yani"), stutters and false starts, and fixes punctuation — guarded so it can only edit your words, never answer them. - Menu bar app —
(idle · recording · paused) with start/stop, language switcher, and start-at-login support.
- Self-healing — audio watchdog survives sleep/wake and device changes; inference errors never kill the session.
- Streaming mode (optional) — words appear as you speak, stabilized with LocalAgreement-2 and append-only commits (typed text is never retracted).
Requirements
- macOS on Apple Silicon
uvinstalled- ~2 GB disk for the default ASR model (downloaded once from Hugging Face)
Install
One-liner (installs uv if needed, puts
parlando on your PATH):
curl -fsSL https://raw.githubusercontent.com/furkanc/parlando/main/install.sh | sh
Then:
parlando # terminal
parlando-menubar # menu bar app
Click the window you want to type into → tap right ⌥ Option → speak → tap again. Your words appear at the cursor. The speech model (~1-2 GB) downloads once on first run; after that everything is offline.
Run from source / uninstall
git clone https://github.com/furkanc/parlando && cd parlando
uv run parlando # run without installing
uv tool install . # or install the commands from the checkout
# uninstall
uv tool uninstall parlando
Permissions (one-time)
macOS will ask for two permissions, granted to the app that runs parlando (Terminal, iTerm, VS Code, ...):
| Permission | Why | Where |
|---|---|---|
| Microphone | hear you | Settings → Privacy & Security → Microphone |
| Accessibility | type keystrokes + global hotkey | Settings → Privacy & Security → Accessibility |
parlando detects missing permissions and tells you explicitly instead of failing silently.
Usage
parlando --language Turkish # dictate in another language
parlando --polish # LLM cleanup (fillers, false starts, punctuation)
parlando --enter # press Enter after each utterance
parlando --mode stream # live word-by-word streaming
parlando --pipe # print to stdout (scriptable)
parlando --hotkey cmd_r # tap right Command instead
parlando --list-devices # list microphones
parlando-menubar --install-login # start at login
Voice commands
| Say (English) | Say (Turkish) | Result |
|---|---|---|
| period / comma | nokta / virgül | . , appended to previous word |
| question mark | soru işareti | ? appended |
| new line / new paragraph | yeni satır / yeni paragraf | line break |
| send | gönder | presses Enter |
Disable with --no-commands.
Models
Speech recognition (--model, MLX Qwen3-ASR — 30 languages):
| Model | Size | Notes |
|---|---|---|
mlx-community/Qwen3-ASR-1.7B-8bit |
~2 GB | default — best accuracy, recommended |
mlx-community/Qwen3-ASR-0.6B-8bit |
~700 MB | good balance for low-RAM machines |
mlx-community/Qwen3-ASR-0.6B-4bit |
~400 MB | fastest, lowest accuracy |
mlx-community/Qwen3-ASR-0.6B-bf16 |
~1.3 GB | higher fidelity 0.6B, more RAM |
Polish LLM (--polish-model, used only with --polish; scores from
scripts/eval_polish.py, an 11-case cleanup benchmark):
| Model | Size | Eval | Latency | Notes |
|---|---|---|---|---|
mlx-community/Qwen3-1.7B-4bit |
~1 GB | 10/11 | ~0.4 s | default — fast, safe |
mlx-community/Qwen3-4B-4bit |
~2.3 GB | 11/11 | ~0.9 s | best quality; use on 16 GB+ Macs |
mlx-community/Qwen3-0.6B-4bit |
~350 MB | — | ~0.2 s | minimal RAM, weakest cleanup |
parlando --polish --polish-model mlx-community/Qwen3-4B-4bit
All models download once from Hugging Face and are cached in
~/.cache/huggingface; everything runs on-device.
The two modes
record (default) — tap, speak freely (pauses are fine), tap again; the
whole recording is transcribed once and typed in one go. Recordings longer
than 28 s are split at natural silence points (2 min cap). This mode has no
VAD guessing, so it is the most robust.
stream — always listening; an energy VAD (hysteresis, adaptive noise
floor, optional --silero hybrid) segments utterances and words appear as
you speak. More "live", more sensitive to room noise.
Troubleshooting
| Symptom | Fix |
|---|---|
| Nothing typed, "only silence" warning | Grant microphone permission, restart |
| Nothing typed, no warning | Grant Accessibility permission (see startup warning) |
| Stalls after sleep | The watchdog reopens the stream within ~5 s automatically |
| Ghost text while silent (stream mode) | --energy-floor 0.008 or --silero |
| Missing soft speech (stream mode) | --energy-floor 0.002 |
| Wrong words | Move closer to the mic; prefer the built-in mic over AirPods (Bluetooth input drops to a low-quality codec) |
| Detailed trace | ~/Library/Logs/parlando.log |
Note: --device indices shift when devices connect/disconnect, and virtual
devices ("Microsoft Teams Audio", "ZoomAudioDevice") are not microphones.
Prefer the default device.
Architecture
microphone ── 30 ms frames ──▶ hotkey-gated recorder (record mode)
or energy VAD w/ hysteresis + adaptive floor,
optional Silero hybrid, 300 ms pre-roll (stream)
│ segment boundaries
▼
segment audio only ──▶ Qwen3-ASR (MLX, on-device GPU)
│ readings
▼
LocalAgreement-2 ──▶ append-only word commits ──▶ CGEvent unicode
(commit on 2-reading (typed text is never keyboard events
agreement) retracted)
Key decisions (each one earned by a real failure during development):
- Segment-based ASR, not a rolling window — re-decoding a growing window revises earlier words (visible flicker) and hallucinates on silence.
- Append-only, word-count-based commits — exact-prefix matching stalls permanently when the ASR reshapes punctuation on committed words.
- Normalized agreement + stall safety valve — punctuation/case flapping between readings must not freeze the stream.
- Asymmetric noise-floor learning — a symmetric EMA lets speech scraps poison the floor and the app goes progressively deaf.
- Every loop iteration guarded — a single ASR exception must not silently kill the pipeline.
- Chord-aware hotkey — ⌥+Q-style character chords never toggle dictation.
Development
uvx --with numpy pytest tests -q # unit tests; no model/mic needed
uv run scripts/eval_polish.py # polish quality benchmark (real local LLM)
uv build # build the wheel/sdist
Package layout: src/parlando/ (engine, menubar, ASR engine), entry points
parlando and parlando-menubar. CI runs the test suite and a packaging
build on macOS via GitHub Actions.
Roadmap
- One-line installer (
install.sh) withparlandolauncher command - Proper Python package (
uv tool install, entry points) — PyPI publication pending, after whichuvx parlandowill work - Publish to PyPI
- Homebrew tap (
brew install parlando) - End-to-end regression tests with recorded WAV fixtures
- Custom vocabulary / context biasing
License & credits
MIT. The Qwen3-ASR MLX engine (stt.py) is from
dictate.sh by Marc Puig (MIT), with
local modifications (partial streaming, energy gating, error hardening).
Turn-taking research: Whisper-Streaming /
LocalAgreement-2. VAD:
Silero VAD.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file parlando-0.4.0.tar.gz.
File metadata
- Download URL: parlando-0.4.0.tar.gz
- Upload date:
- Size: 197.4 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
70dae955cd60ddcd8237b61ac40af809703b310dab9a34b56a20f0d12c0327f1
|
|
| MD5 |
bd44030f2e6247bd0b811d44c94784fb
|
|
| BLAKE2b-256 |
e7ee35cb71077814ed5b65cbd16d79c37e5eb4480178901e5d7bcad7e489507d
|
Provenance
The following attestation bundles were made for parlando-0.4.0.tar.gz:
Publisher:
release.yml on furkanc/parlando
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
parlando-0.4.0.tar.gz -
Subject digest:
70dae955cd60ddcd8237b61ac40af809703b310dab9a34b56a20f0d12c0327f1 - Sigstore transparency entry: 2768734851
- Sigstore integration time:
-
Permalink:
furkanc/parlando@47a95440f19795e8cd8d0dce0dcb7833465d8276 -
Branch / Tag:
refs/tags/v0.4.0 - Owner: https://github.com/furkanc
-
Access:
private
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@47a95440f19795e8cd8d0dce0dcb7833465d8276 -
Trigger Event:
release
-
Statement type:
File details
Details for the file parlando-0.4.0-py3-none-any.whl.
File metadata
- Download URL: parlando-0.4.0-py3-none-any.whl
- Upload date:
- Size: 44.9 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
ec3281199cddd59e707ddf285dd2b427f0969662883de15cd700fe728034f298
|
|
| MD5 |
9a7ea2064dfc2ff7cdfd7f933798656d
|
|
| BLAKE2b-256 |
94fb54d072ef9097a05959e0b8a26ad135711d1d973afa45d3827ecc6c0b5f5c
|
Provenance
The following attestation bundles were made for parlando-0.4.0-py3-none-any.whl:
Publisher:
release.yml on furkanc/parlando
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
parlando-0.4.0-py3-none-any.whl -
Subject digest:
ec3281199cddd59e707ddf285dd2b427f0969662883de15cd700fe728034f298 - Sigstore transparency entry: 2768734916
- Sigstore integration time:
-
Permalink:
furkanc/parlando@47a95440f19795e8cd8d0dce0dcb7833465d8276 -
Branch / Tag:
refs/tags/v0.4.0 - Owner: https://github.com/furkanc
-
Access:
private
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@47a95440f19795e8cd8d0dce0dcb7833465d8276 -
Trigger Event:
release
-
Statement type: