Skip to main content

Beseda

Voice conversations with your coding agent. Say “Alice, …” (in Russian, “Вика, …”) and talk to an AI agent in English or Russian, right from the terminal: speech recognition and synthesis run locally on your Mac, the agent works in the current folder.

Русская версия

Beseda demo: wake word, a question answered with shell commands, a hold phrase, a follow-up, stop

▶ The same conversation with sound (MP4) (Russian). The user's phrases are spoken by a second Silero voice instead of a microphone; recognition, the pi agent and the answers are the real app. Long waits in the GIF are shortened.

microphone → Silero VAD + whisper.cpp (local) → brain: pi agent or DeepSeek → TTS: Silero (local) → speakers

Status: alpha. macOS on Apple Silicon only. Russian (default) and English; more languages are a TOML file away.

Features

  • Wake word, like a smart speaker. Only phrases addressed to «Вика» reach the model; after an answer you can keep talking without the wake word, «стоп» ends the conversation, «подожди» gives you time to think.
  • Local speech. whisper.cpp on the Mac GPU (Metal) for recognition, Silero for synthesis: from the first word of the answer to sound in 0.1–0.25 s.
  • Pluggable brains. The pi coding agent (reads and edits files, runs commands) or a plain DeepSeek chat; adding Claude or Codex is one class.
  • Observability. Every session writes a log with a per-turn latency timeline: recognition, first token, first audio, tool calls.
  • Dialog recording. The whole conversation as one WAV, with the real pauses.

Requirements

  • macOS on Apple Silicon, Python 3.12+, uv, brew install portaudio
  • For --tts edge: brew install ffmpeg
  • For the default brain: npm install -g @earendil-works/pi-coding-agent with a DeepSeek key configured in pi. For --brain deepseek: the DEEPSEEK_API_KEY environment variable.

Install

brew install portaudio
uv tool install beseda

Models download on first launch into ~/.beseda/models/ (Whisper small ~490 MB, Silero ~145 MB, VAD ~1 MB).

Use

cd ~/some/project   # the agent works in the current folder
beseda                      # Russian
beseda --language en        # English

The table shows the Russian phrases; the English pack has its own (“Alice, …”, “hold on”, “that's all”).

You say What happens
«Вика, какая погода?» Conversation starts (Tink sound); «какая погода?» goes to the model
«Вика» A beep, then it waits for the request
anything, within --follow-up s after an answer (8 s) Goes to the model without the wake word; the status line counts down
«Подожди», «дай подумать», «секунду» Not sent to the model; the wait extends to --hold s (2 min)
«Стоп», «хватит», «спасибо, всё», or silence Conversation ends (Bottle sound); the wake word is needed again
anything else without the wake word Ignored, shown dimmed

Keys: Space turns the microphone on/off, Esc interrupts the answer, q quits.

The microphone is off while the assistant speaks, so it never hears itself; interrupting by voice isn't supported yet. The macOS microphone indicator stays on while Beseda runs: the stream must stay open to hear the wake word.

Configuration

Any option can be set in ~/.beseda/config.toml; command-line flags win.

language = "ru"           # ru | en | your own pack
brain = "pi"              # pi | deepseek
model = "deepseek/deepseek-v4-flash"
whisper = "small"         # small | turbo
vocabulary = ["JavaScript", "DeepSeek"]
tts = "silero"            # silero | edge | say
voice = "baya"
wake-word = "вика"        # default comes from the language pack; "" answers everything
follow-up = 8
hold = 120
record = true             # or a file path
log-days = 14

See beseda --help for the full list.

Languages

Everything that depends on the spoken language lives in a language pack, a TOML file: Whisper's language and style prompt, the voice prompt for the model, the wake word and its grammatical endings, stop and hold phrases, default TTS voices and the Silero model, and the terminal UI strings.

Built-in packs: ru (default) and en. To change a pack or add a language, put a file into ~/.beseda/languages/: ru.toml there overrides the built-in one, de.toml adds German (beseda --language de). Copy a built-in pack as a starting point; every key is required.

Speech recognition

Techniques carried over from Vadic and checked on real dialogs: loudness normalization before recognition, Whisper's own VAD (without it a cough becomes «Спасибо.», which is a stop phrase), a style prompt for punctuation plus your vocabulary for terms, greedy decoding at temperature 0.

whisper Time per phrase (M1) Notes
small (default) ~0.5 s Occasional wrong words
turbo (large-v3-turbo-q5_0) ~2.1 s Far more accurate; 574 MB

Models other apps already downloaded (VoiceInk, Vadic) are reused.

Speech synthesis

tts Russian voices English voices Runs Per sentence (ru) CPU per second of speech RAM
silero (default) xenia, baya, kseniya, aidar, eugene en_0 … en_4 locally 0.04 s (en: ~0.35 s) 16 ms ~760 MB
say Milena Daniel locally (macOS) 0.6 s 140 ms ~40 MB
edge (experimental) ru-RU-SvetlanaNeural, ru-RU-DmitryNeural en-US-AriaNeural, en-US-GuyNeural Microsoft cloud 1–2 s 60 ms ~55 MB

Silero's Russian model skips digits and Latin letters, so the Russian voice prompt asks the model to write numbers and names in Russian words. Compare the voices by ear: beseda-samples [--language en] writes the same phrase in every voice to ~/Downloads/beseda-tts-samples/.

Logs and recordings

Each session logs to ~/.beseda/logs/beseda-<time>.log; logs older than log-days are deleted on start.

grep summary ~/.beseda/logs/*.log | tail      # latency per turn: stt, llm_first_token, tts_first_audio, voice_to_voice
grep -E 'ERROR|WARNING' ~/.beseda/logs/*.log   # incidents

--debug adds every pi event. --record saves the dialog to ~/Downloads/beseda-<time>.wav on exit.

Security

With the default pi brain the agent has full access: it reads and writes files and runs shell commands in the current folder, by voice. Recognition makes mistakes. Run Beseda only in folders where that is acceptable, and keep them under version control.

Privacy

  • Recognition, synthesis with silero or say, logs and recordings stay on your Mac.
  • What you say to the assistant (after the wake word) is sent to the LLM provider: DeepSeek by default.
  • With tts = "edge" the answers are sent to Microsoft.
  • Logs contain the full text of your dialogs, including phrases that weren't addressed to the assistant; they are kept for log-days days.

Model licenses

Beseda's code is MIT, and no models are bundled: they download to your machine on first use.

  • Silero TTS (default voice): CC BY-NC-SA 4.0, non-commercial use only. For commercial use pick say or edge, or obtain a license from Silero.
  • Whisper models and the Silero VAD used by whisper.cpp: MIT.
  • Edge TTS uses an unofficial endpoint of the Microsoft Edge read-aloud service and may stop working.

Extending

  • A brain (beseda/brains.py) yields ("text", …), ("tool", …), ("error", …) events and supports abort(); register it in BRAINS.
  • A TTS engine (beseda/tts.py) subclasses PcmEngine with render(text) -> PCM; register it in TTS_ENGINES and list its voices under [voices] in each language pack.
  • A language is a TOML file, see Languages.

Development

uv sync
uv run pytest
uv tool install --editable .   # the beseda command picks up code changes
uv run python scripts/demo.py  # re-record docs/demo-ru.gif and .mp4; --language en for the English one (needs brew install agg ffmpeg)
git tag v0.1.0 && git push origin v0.1.0   # release: CI tests and publishes to PyPI (the tag must match the version in pyproject.toml)

License

MIT

Metadata

Release files for beseda 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for beseda 0.1.0
File Size Uploaded
beseda-0.1.0.tar.gz 25.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for beseda 0.1.0
File Interpreter ABI Platform
beseda-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 55.2 kB

Release files / beseda-0.1.0.tar.gz

Download URL beseda-0.1.0.tar.gz
Size 25.0 kB
Tags Source
SHA-256 checksum
How to use checksums
65c009337b2d23811eb096e65f054c1c3a72d54c61d431d28aeb62409df21e61
BLAKE2b-256 checksum
How to use checksums
f1cf6c7dccaac48731d2d637eb4c296b4127526e9b32f4a5bc1981181ab5cf52
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.23 {"installer":{"name":"uv","version":"0.12.23","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release files / beseda-0.1.0-py3-none-any.whl

Download URL beseda-0.1.0-py3-none-any.whl
Size 30.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
b9f7f7297d5ed6f9bf5bd6ed56003c5cb09d0eeefd47a999f0c3d3e2794e982a
BLAKE2b-256 checksum
How to use checksums
71e1d771573fbc6aa3067103a5178b9e2695ea24731f06d90b412b20e3fc0a6a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.23 {"installer":{"name":"uv","version":"0.12.23","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release history Release notifications | RSS feed

0.3.0

2 release files

0.2.0

2 release files

0.1.1

2 release files

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page