Skip to main content

parlando — local voice dictation for macOS

Local voice dictation for macOS (Apple Silicon). Tap a key, speak, tap again — your words are typed into whatever window has focus. 100% on-device: audio never leaves your machine.

Tests

Features

  • Fully local & offline — MLX-accelerated Qwen3-ASR on the Apple Neural stack; works on a plane, nothing is uploaded, no API keys.
  • Tap-to-dictate — single tap of right ⌥ Option starts/stops recording (the same pattern superwhisper and Wispr Flow use). No always-on mic.
  • Types anywhere — text is delivered as native keyboard events (Quartz CGEvent unicode), so it works in any app, editor, terminal, or chat box.
  • 30 languages — English by default; switch instantly (e.g. --language Turkish) with no model reload.
  • Voice commands — "period", "comma", "question mark", "new line", "send" (presses Enter). A Turkish command set activates with --language Turkish.
  • Cleanup & polish — vocalized fillers ("um", "uh", "eee") are removed automatically; optional --polish runs a small local LLM that drops contextual fillers ("you know", "yani"), stutters and false starts, and fixes punctuation — guarded so it can only edit your words, never answer them.
  • Menu bar appmenu bar icon states: idle, recording, paused (idle · recording · paused) with start/stop, language switcher, and start-at-login support.
  • Self-healing — audio watchdog survives sleep/wake and device changes; inference errors never kill the session.
  • Streaming mode (optional) — words appear as you speak, stabilized with LocalAgreement-2 and append-only commits (typed text is never retracted).

Requirements

  • macOS on Apple Silicon
  • uv installed
  • ~2 GB disk for the default ASR model (downloaded once from Hugging Face)

Install

One-liner (installs uv if needed, puts parlando on your PATH):

curl -fsSL https://raw.githubusercontent.com/furkanc/parlando/main/install.sh | sh

Then:

parlando            # terminal
parlando-menubar    # menu bar app

Click the window you want to type into → tap right ⌥ Option → speak → tap again. Your words appear at the cursor. The speech model (~1-2 GB) downloads once on first run; after that everything is offline.

Run from source / uninstall
git clone https://github.com/furkanc/parlando && cd parlando
uv run parlando              # run without installing
uv tool install .           # or install the commands from the checkout

# uninstall
uv tool uninstall parlando

Permissions (one-time)

macOS will ask for two permissions, granted to the app that runs parlando (Terminal, iTerm, VS Code, ...):

Permission Why Where
Microphone hear you Settings → Privacy & Security → Microphone
Accessibility type keystrokes + global hotkey Settings → Privacy & Security → Accessibility

parlando detects missing permissions and tells you explicitly instead of failing silently.

Usage

parlando --language Turkish   # dictate in another language
parlando --polish             # LLM cleanup (fillers, false starts, punctuation)
parlando --enter              # press Enter after each utterance
parlando --mode stream        # live word-by-word streaming
parlando --pipe               # print to stdout (scriptable)
parlando --hotkey cmd_r       # tap right Command instead
parlando --list-devices       # list microphones
parlando-menubar --install-login   # start at login

Voice commands

Say (English) Say (Turkish) Result
period / comma nokta / virgül . , appended to previous word
question mark soru işareti ? appended
new line / new paragraph yeni satır / yeni paragraf line break
send gönder presses Enter

Disable with --no-commands.

Models

Speech recognition (--model, MLX Qwen3-ASR — 30 languages):

Model Size Notes
mlx-community/Qwen3-ASR-1.7B-8bit ~2 GB default — best accuracy, recommended
mlx-community/Qwen3-ASR-0.6B-8bit ~700 MB good balance for low-RAM machines
mlx-community/Qwen3-ASR-0.6B-4bit ~400 MB fastest, lowest accuracy
mlx-community/Qwen3-ASR-0.6B-bf16 ~1.3 GB higher fidelity 0.6B, more RAM

Polish LLM (--polish-model, used only with --polish; scores from scripts/eval_polish.py, an 11-case cleanup benchmark):

Model Size Eval Latency Notes
mlx-community/Qwen3-1.7B-4bit ~1 GB 10/11 ~0.4 s default — fast, safe
mlx-community/Qwen3-4B-4bit ~2.3 GB 11/11 ~0.9 s best quality; use on 16 GB+ Macs
mlx-community/Qwen3-0.6B-4bit ~350 MB ~0.2 s minimal RAM, weakest cleanup
parlando --polish --polish-model mlx-community/Qwen3-4B-4bit

All models download once from Hugging Face and are cached in ~/.cache/huggingface; everything runs on-device.

The two modes

record (default) — tap, speak freely (pauses are fine), tap again; the whole recording is transcribed once and typed in one go. Recordings longer than 28 s are split at natural silence points (2 min cap). This mode has no VAD guessing, so it is the most robust.

stream — always listening; an energy VAD (hysteresis, adaptive noise floor, optional --silero hybrid) segments utterances and words appear as you speak. More "live", more sensitive to room noise.

Troubleshooting

Symptom Fix
Nothing typed, "only silence" warning Grant microphone permission, restart
Nothing typed, no warning Grant Accessibility permission (see startup warning)
Stalls after sleep The watchdog reopens the stream within ~5 s automatically
Ghost text while silent (stream mode) --energy-floor 0.008 or --silero
Missing soft speech (stream mode) --energy-floor 0.002
Wrong words Move closer to the mic; prefer the built-in mic over AirPods (Bluetooth input drops to a low-quality codec)
Detailed trace ~/Library/Logs/parlando.log

Note: --device indices shift when devices connect/disconnect, and virtual devices ("Microsoft Teams Audio", "ZoomAudioDevice") are not microphones. Prefer the default device.

Architecture

microphone ── 30 ms frames ──▶ hotkey-gated recorder (record mode)
                              or energy VAD w/ hysteresis + adaptive floor,
                              optional Silero hybrid, 300 ms pre-roll (stream)
                │ segment boundaries
                ▼
        segment audio only ──▶ Qwen3-ASR (MLX, on-device GPU)
                │ readings
                ▼
        LocalAgreement-2 ──▶ append-only word commits ──▶ CGEvent unicode
        (commit on 2-reading    (typed text is never        keyboard events
         agreement)              retracted)

Key decisions (each one earned by a real failure during development):

  • Segment-based ASR, not a rolling window — re-decoding a growing window revises earlier words (visible flicker) and hallucinates on silence.
  • Append-only, word-count-based commits — exact-prefix matching stalls permanently when the ASR reshapes punctuation on committed words.
  • Normalized agreement + stall safety valve — punctuation/case flapping between readings must not freeze the stream.
  • Asymmetric noise-floor learning — a symmetric EMA lets speech scraps poison the floor and the app goes progressively deaf.
  • Every loop iteration guarded — a single ASR exception must not silently kill the pipeline.
  • Chord-aware hotkey — ⌥+Q-style character chords never toggle dictation.

Development

uvx --with numpy pytest tests -q   # unit tests; no model/mic needed
uv run scripts/eval_polish.py      # polish quality benchmark (real local LLM)
uv build                           # build the wheel/sdist

Package layout: src/parlando/ (engine, menubar, ASR engine), entry points parlando and parlando-menubar. CI runs the test suite and a packaging build on macOS via GitHub Actions.

Roadmap

  • One-line installer (install.sh) with parlando launcher command
  • Proper Python package (uv tool install, entry points) — PyPI publication pending, after which uvx parlando will work
  • Publish to PyPI
  • Homebrew tap (brew install parlando)
  • End-to-end regression tests with recorded WAV fixtures
  • Custom vocabulary / context biasing

License & credits

MIT. The Qwen3-ASR MLX engine (stt.py) is from dictate.sh by Marc Puig (MIT), with local modifications (partial streaming, energy gating, error hardening). Turn-taking research: Whisper-Streaming / LocalAgreement-2. VAD: Silero VAD.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

parlando-0.4.0.tar.gz (197.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

parlando-0.4.0-py3-none-any.whl (44.9 kB view details)

Uploaded Python 3

File details

Details for the file parlando-0.4.0.tar.gz.

File metadata

  • Download URL: parlando-0.4.0.tar.gz
  • Upload date:
  • Size: 197.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for parlando-0.4.0.tar.gz
Algorithm Hash digest
SHA256 70dae955cd60ddcd8237b61ac40af809703b310dab9a34b56a20f0d12c0327f1
MD5 bd44030f2e6247bd0b811d44c94784fb
BLAKE2b-256 e7ee35cb71077814ed5b65cbd16d79c37e5eb4480178901e5d7bcad7e489507d

See more details on using hashes here.

Provenance

The following attestation bundles were made for parlando-0.4.0.tar.gz:

Publisher: release.yml on furkanc/parlando

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file parlando-0.4.0-py3-none-any.whl.

File metadata

  • Download URL: parlando-0.4.0-py3-none-any.whl
  • Upload date:
  • Size: 44.9 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for parlando-0.4.0-py3-none-any.whl
Algorithm Hash digest
SHA256 ec3281199cddd59e707ddf285dd2b427f0969662883de15cd700fe728034f298
MD5 9a7ea2064dfc2ff7cdfd7f933798656d
BLAKE2b-256 94fb54d072ef9097a05959e0b8a26ad135711d1d973afa45d3827ecc6c0b5f5c

See more details on using hashes here.

Provenance

The following attestation bundles were made for parlando-0.4.0-py3-none-any.whl:

Publisher: release.yml on furkanc/parlando

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.6.1

2 files

0.6.0

2 files

0.5.0

2 files

0.4.2

2 files

0.4.1

2 files

This release

0.4.0 This release

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page