Beseda
Voice conversations with your coding agent. Say “Alice, …” (in Russian, “Вика, …”) and talk to an AI agent in English or Russian, right from the terminal: speech recognition and synthesis run locally on your Mac, the agent works in the current folder.
▶ The same conversation with sound (MP4) (Russian). The user's phrases are spoken by a second Silero voice instead of a microphone; recognition, the pi agent and the answers are the real app. Long waits in the GIF are shortened.
microphone → Silero VAD + GigaAM or whisper.cpp (local) → brain: pi agent, Codex or DeepSeek → TTS: Silero (local) → speakers
Status: alpha. macOS on Apple Silicon only. Russian (default) and English; more languages are a TOML file away.
Features
- Wake word, like a smart speaker. Only phrases addressed to «Вика» reach the model; after an answer you can keep talking without the wake word, «стоп» ends the conversation, «подожди» gives you time to think.
- Local speech. GigaAM for Russian (0.1–0.3 s per phrase) or whisper.cpp on the Mac GPU for recognition, Silero for synthesis: from the first word of the answer to sound in 0.1–0.25 s.
- Pluggable brains. The pi coding agent (reads and edits files, runs commands), Codex or a plain DeepSeek chat; adding another agent is one class.
- Works with music or a video playing. macOS echo cancellation (the one FaceTime uses) removes the sound of the speakers from the microphone; other apps keep their volume.
- Observability. Every session writes a log with a per-turn latency timeline: recognition, first token, first audio, tool calls.
- Dialog recording. The whole conversation as one WAV, with the real pauses.
Requirements
- macOS on Apple Silicon, Python 3.12+, uv,
brew install portaudio - For
--tts edge:brew install ffmpeg - For the default brain:
npm install -g @earendil-works/pi-coding-agentwith a DeepSeek key configured in pi. For--brain codex: the Codex CLI signed in withcodex login; the model comes from~/.codex/config.tomlunless--modelis set. For--brain deepseek: theDEEPSEEK_API_KEYenvironment variable.
Install
brew install portaudio
uv tool install beseda
Models download on first launch into ~/.beseda/models/ (GigaAM ~890 MB for Russian or Whisper small ~490 MB for English, Silero ~145 MB, VAD ~1 MB).
Use
cd ~/some/project # the agent works in the current folder
beseda # Russian
beseda --language en # English
The table shows the Russian phrases; the English pack has its own (“Alice, …”, “hold on”, “that's all”).
| You say | What happens |
|---|---|
| «Вика, какая погода?» | Conversation starts (Tink sound); «какая погода?» goes to the model |
| «Вика» | A beep, then it waits for the request |
anything, within --follow-up s after an answer (8 s) |
Goes to the model without the wake word; the status line counts down |
| «Подожди», «дай подумать», «секунду» | Not sent to the model; the wait extends to --hold s (2 min) |
| «Стоп», «хватит», «спасибо, всё», or silence | Conversation ends (Bottle sound); the wake word is needed again |
| anything else without the wake word | Ignored, shown dimmed |
Keys: Space turns the microphone on/off, Esc interrupts the answer, q quits.
The microphone is off while the assistant speaks, so it never hears itself; interrupting by voice isn't supported yet. The macOS microphone indicator stays on while Beseda runs: the stream must stay open to hear the wake word.
Configuration
Any option can be set in ~/.beseda/config.toml; command-line flags win.
language = "ru" # ru | en | your own pack
brain = "pi" # pi | codex | deepseek
model = "deepseek/deepseek-v4-flash"
stt = "gigaam" # gigaam (Russian) | small | turbo
vocabulary = ["JavaScript", "DeepSeek"]
tts = "silero" # silero | edge | say
voice = "baya"
wake-word = "вика" # default comes from the language pack; "" answers everything
follow-up = 8
hold = 120
record = true # or a file path
log-days = 14
See beseda --help for the full list.
Languages
Everything that depends on the spoken language lives in a language pack, a TOML file: Whisper's language and style prompt, the voice prompt for the model, the wake word and its grammatical endings, stop and hold phrases, default TTS voices and the Silero model, and the terminal UI strings.
Built-in packs: ru (default) and en. To change a pack
or add a language, put a file into ~/.beseda/languages/: ru.toml there overrides the built-in one, de.toml
adds German (beseda --language de). Copy a built-in pack as a starting point; every key is required.
Speech recognition
Each language pack picks a default stt: GigaAM for Russian, Whisper small for English.
stt |
Time per phrase (M1) | Notes |
|---|---|---|
gigaam (default for ru) |
0.1–0.3 s, CPU | GigaAM v3 by Sber: Russian only, punctuation and capitals; 890 MB |
small (default for en) |
~0.5 s, Metal | Occasional wrong words |
turbo (large-v3-turbo-q5_0) |
~2.1 s, Metal | Far more accurate than small; 574 MB |
Whisper keeps the techniques from Vadic, checked on real dialogs:
loudness normalization before recognition, Whisper's own VAD (without it a cough becomes «Спасибо.», which is a
stop phrase), a style prompt for punctuation plus your vocabulary for terms, greedy decoding at temperature 0.
GigaAM takes no prompt, so vocabulary applies to Whisper only.
Models other apps already downloaded (VoiceInk, Vadic) are reused.
Speech synthesis
tts |
Russian voices | English voices | Runs | Per sentence (ru) | CPU per second of speech | RAM |
|---|---|---|---|---|---|---|
silero (default) |
xenia, baya, kseniya, aidar, eugene | en_0 … en_4 | locally | 0.04 s (en: ~0.35 s) | 16 ms | ~760 MB |
say |
Milena | Daniel | locally (macOS) | 0.6 s | 140 ms | ~40 MB |
edge (experimental) |
ru-RU-SvetlanaNeural, ru-RU-DmitryNeural | en-US-AriaNeural, en-US-GuyNeural | Microsoft cloud | 1–2 s | 60 ms | ~55 MB |
Silero's Russian model skips digits and Latin letters, so the Russian voice prompt asks the model to write numbers
and names in Russian words. Compare the voices by ear: beseda-samples [--language en] writes the same phrase in
every voice to ~/Downloads/beseda-tts-samples/.
Logs and recordings
Each session logs to ~/.beseda/logs/beseda-<time>.log; logs older than log-days are deleted on start.
grep summary ~/.beseda/logs/*.log | tail # latency per turn: stt, llm_first_token, tts_first_audio, voice_to_voice
grep -E 'ERROR|WARNING' ~/.beseda/logs/*.log # incidents
--debug adds every pi event. --record saves the dialog to ~/Downloads/beseda-<time>.wav on exit.
Security
With the default pi brain the agent has full access: it reads and writes files and runs shell commands in
the current folder, by voice. With codex it runs commands without asking in Codex's workspace-write sandbox:
it can change the current folder but not the rest of the disk, and commands have no network. Recognition makes mistakes. Run Beseda only in folders where that is acceptable,
and keep them under version control.
Privacy
- Recognition, synthesis with
sileroorsay, logs and recordings stay on your Mac. - What you say to the assistant (after the wake word) is sent to the LLM provider: DeepSeek by default, OpenAI with
codex. - With
tts = "edge"the answers are sent to Microsoft. - Logs contain the full text of your dialogs, including phrases that weren't addressed to the assistant;
they are kept for
log-daysdays.
Model licenses
Beseda's code is MIT, and no models are bundled: they download to your machine on first use.
- Silero TTS (default voice): CC BY-NC-SA 4.0,
non-commercial use only. For commercial use pick
sayoredge, or obtain a license from Silero. - GigaAM, Whisper models and the Silero VAD used by whisper.cpp: MIT.
- Edge TTS uses an unofficial endpoint of the Microsoft Edge read-aloud service and may stop working.
Extending
- A brain (
beseda/brains.py) yields("text", …),("tool", …),("error", …)events and supportsabort(); register it inBRAINS. - A TTS engine (
beseda/tts.py) subclassesPcmEnginewithrender(text) -> PCM; register it inTTS_ENGINESand list its voices under[voices]in each language pack. - A language is a TOML file, see Languages.
Development
uv sync
uv run pytest
uv tool install --editable . # the beseda command picks up code changes; add --force after dependency changes
uv run python scripts/demo.py # re-record docs/demo-ru.gif and .mp4; --language en for the English one (needs brew install agg ffmpeg)
git tag vX.Y.Z && git push origin vX.Y.Z # release: CI tests and publishes to PyPI (the tag must match the version in pyproject.toml)
License
Metadata
Release files for beseda 0.3.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| beseda-0.3.0.tar.gz | 29.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| beseda-0.3.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 65.6 kB
Release files / beseda-0.3.0.tar.gz
| Download URL | beseda-0.3.0.tar.gz |
|---|---|
| Size | 29.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
7e42ee0b1427da83d04bddbc5681288d89d080a14e54e0544b9646dd619e5e67
|
|
BLAKE2b-256 checksum How to use checksums |
bd80edfcadc0369c73daedb4bef76db5e8bdd510d544fc6cc0ea4728b9e4dacd
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.12.23 {"installer":{"name":"uv","version":"0.12.23","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|
Release files / beseda-0.3.0-py3-none-any.whl
| Download URL | beseda-0.3.0-py3-none-any.whl |
|---|---|
| Size | 35.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
ea0c73419e77402d88c6769f9b5ca1df401d20ba15102e2851d5522578d8b2bb
|
|
BLAKE2b-256 checksum How to use checksums |
8953c47b75a5c3fa7d1fd6a46d2940ca50ff4370dd4f7c1c7ae1bfa9d460e705
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.12.23 {"installer":{"name":"uv","version":"0.12.23","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|