Skip to main content

MCP Voice Summary Server

CI License: MIT Python

A local MCP server that reads out loud a summary of the actions an AI assistant just performed on your code. Built as an accessibility layer: you hear what the agent did without having to read the full response.

Speaks in the user's language automatically, with female and male voices, and works fully offline if you want it to.

Quick start

pip install mcp-voice-summary

Then point your MCP client at the installed command. For Claude Desktop, Cursor or Cline, in claude_desktop_config.json:

{
  "mcpServers": {
    "voice-summary": {
      "command": "mcp-voice-summary",
      "env": {
        "VOICE_ENGINE": "edge",
        "VOICE_LANGUAGE": "es",
        "VOICE_GENDER": "male"
      }
    }
  }
}

No env block is needed at all: the server picks the voice from your system language. The values above only pin it.

With uvx, without installing anything:

{
  "mcpServers": {
    "voice-summary": {
      "command": "uvx",
      "args": ["mcp-voice-summary"],
      "env": { "VOICE_ENGINE": "edge" }
    }
  }
}

Then tell your assistant to use it, for example: "After you finish a task, read a short summary out loud."

Platform support

Platform Recommended engine Voice source Extra install
Windows edge for quality, sapi5 for offline System voices, or neural pip install -e ".[sapi5]" for offline
macOS edge for quality, sapi5 for offline System voices, or neural pip install pyttsx3 for offline
Linux edge Neural voices pip install ffmpeg for playback
Linux offline sapi5 espeak-ng sudo apt install espeak-ng libespeak-ng1

The edge engine needs an internet connection on every utterance, because the audio is generated by Microsoft's servers. The sapi5 engine is fully offline.

On Linux the audio player is detected automatically; install ffmpeg (which provides ffplay) or mpg123 if nothing is found.

Higher quality Windows voices, offline

Windows ships better voices than the three exposed by default, but hides them under a registry key SAPI5 does not read. Run this once as administrator:

powershell -Command "Start-Process powershell -Verb RunAs -ArgumentList '-ExecutionPolicy Bypass -File .\register_voices_onecore.ps1'"

Afterwards Microsoft Laura and Microsoft Pablo become available to sapi5.

What it does

Exposes an MCP tool, speak_summary, which the assistant calls after modifying code, creating files or running commands.

The summary length is decided by the assistant based on how much work it did: a short sentence for a small change, or a fuller summary when the task was large or had several steps. See Controlling the summary length.

Playback is asynchronous: the tool queues the text and returns immediately, so speaking never blocks the assistant.

Requirements

  • Python 3.10 or newer.
  • Windows, Linux or macOS.

Only the sapi5 engine needs a system dependency (the native synthesizer). The edge engine needs nothing beyond Python.

Checking that it works

mcp-voice-summary

The server speaks over stdio, so you will see nothing in the console. That is correct: anything printed to stdout would break the protocol. Ctrl+C to stop.

For the offline engine, install its extra first:

pip install -e ".[sapi5]"

Configuration

Everything is controlled through environment variables.

Variable Values Default Description
VOICE_ENGINE edge or sapi5 sapi5 Synthesis engine
VOICE_NAME voice id or name auto Exact voice, overrides everything
VOICE_LANGUAGE es, en, pt-BR, auto system language Voice language
VOICE_GENDER female, male language default Voice gender
VOICE_RATE 100, 110, -15% 100 Speed
VOICE_VOLUME 100 100 Volume
VOICE_PLAYER path or name auto-detected Forced MP3 player
MAX_SUMMARY_WORDS integer 400 Safety cap on words per summary
MAX_QUEUE_SIZE integer 20 Maximum summaries waiting to be spoken
PLAYBACK_TIMEOUT seconds 30 Gives up on playback that never returns
VOICE_MUTE 1, true, yes unset Silent mode: still queues, plays nothing
VOICE_REDACT 0, false, no on Set to 0 to stop stripping credentials
VOICE_RATE_LIMIT notifications per minute 30 Stops a client flooding the queue
MAX_SUMMARY_AGE seconds 120 Queued summaries older than this are dropped
VOICE_MODE full, short full short speaks a fixed generic message

VOICE_RATE and VOICE_VOLUME accept both an absolute notation (the SAPI5 scale, where 100 is normal) and a relative one (+10%, -15%). The server translates automatically into the format each engine requires.

After changing the configuration, restart the MCP client.

Languages

Nothing needs configuring for this to work in your language. The server picks the voice on its own. The precedence is:

  1. VOICE_NAME, if set: it always wins as a manual override.
  2. set_language, if the assistant called it during the session.
  3. VOICE_LANGUAGE, if set in the configuration.
  4. The operating system language.
  5. English, as a last resort.

Changing language and gender

Permanently, in the configuration:

"environment": { "VOICE_LANGUAGE": "fr", "VOICE_GENDER": "male" }

Or at runtime, without editing anything. The user can ask in natural language and the assistant calls the tool:

set_language("en")             -> Language English (en), gender default
set_language("pt-BR", "male")  -> Language Portuguese (pt), gender male
set_language("", "f")           -> keep the language, switch to a female voice
set_language("system")         -> back to the OS language
set_language("", "any")         -> back to the language default gender

A runtime change applies to the following utterances and lasts until the server restarts or the tool is called again with a different value.

Gender only works with the edge engine. SAPI5 does not expose the gender of its voices, so with sapi5 the tool says so instead of pretending. Picking a gender requires neural voices.

Automatic language detection

VOICE_LANGUAGE=auto makes the server infer the language from each summary's text.

Use this with caution. Language detection is not reliable on short text, and short summaries are the main use case of this project. Measured with real summaries:

Text Actual language Detected
"Done." Spanish Czech (with 100 % confidence)
"Listo" Spanish German
"Test 123" anything French
"I updated the login endpoint and fixed the dependencies" English English

Detectors return high confidence even when they are wrong, so the errors cannot be filtered out by probability. In short: detection works on long sentences and fails on short ones. For a fixed language, VOICE_LANGUAGE is always more reliable.

If langdetect is not installed, auto logs a warning and falls back to English. The other modes work without that dependency.

Coverage

The edge engine exposes 142 locales across 322 voices. These 34 languages have both a female and a male voice assigned:

Language Female Male Language Female Male
es Ximena Álvaro da Christel Jeppe
en Ava Andrew fi Noora Harri
fr Vivienne Rémy nl Colette Maarten
de Seraphina Florian pl Zofia Marek
it Elsa Giuseppe ru Svetlana Dmitry
pt Thalita Antonio uk Polina Ostap
ca Joana Enric cs Vlasta Antonín
gl Sabela Roi sk Viktoria Lukáš
hu Noémi Tamás ro Alina Emil
bg Kalina Borislav el Athina Nestoras
sv Sofie Mattias tr Emel Ahmet
nb Pernille Finn ar Salma Shakir
ja Nanami Keita he Hila Avri
ko Sun-Hi Hyunsu hi Swara Madhur
zh Xiaoxiao Yunxian th Premwadee Niwat
vi HoaiMy NamMinh id Gadis Ardi
ms Yasmin Osman

(Short names; the exact identifiers carry the Neural suffix and can be seen with list_voices or python -m edge_tts --list-voices.)

Of the 142 locales, only 2 have a single gender: the Chinese dialects zh-CN-liaoning and zh-CN-shaanxi. If you ask for a gender that does not exist in the chosen locale, the server logs a warning and uses the available voice instead of going silent.

Languages without curated voices still work: the server searches the 322 voices for the first one matching the requested language and gender. For example, sw (Swahili) resolves to sw-KE-RafikiNeural.

If you request a specific regional variant it is respected, by gender too:

You ask Female Male
en-GB en-GB-LibbyNeural en-GB-RyanNeural
pt-PT pt-PT-RaquelNeural pt-PT-DuarteNeural
es-MX es-MX-DaliaNeural es-MX-JorgeNeural
zh-TW zh-TW-HsiaoChenNeural zh-TW-YunJheNeural
fr-CA fr-CA-SylvieNeural fr-CA-ThierryNeural

The sapi5 engine can only use voices installed on the system, so its language coverage is whatever your operating system provides, and it cannot pick gender. Windows ships with Spanish and English; other languages require adding voices.

Voices

Neural voices (edge engine)

Full catalog of all 322 voices:

python -m edge_tts --list-voices

The 45 Spanish voices include:

Voice Region
es-ES-AlvaroNeural Spain, male
es-ES-XimenaNeural Spain, female
es-ES-ElviraNeural Spain, female
es-MX-DaliaNeural Mexico, female
es-MX-JorgeNeural Mexico, male
es-US-PalomaNeural United States, female

Native Windows voices (sapi5 engine)

Out of the box only three very basic voices are visible. Windows already ships better ones, such as Microsoft Laura and Microsoft Pablo, but registers them under the registry key Speech_OneCore, which SAPI5 does not read.

register_voices_onecore.ps1 copies those keys into the branch SAPI5 does read. Run it once from a PowerShell prompt as administrator:

powershell -Command "Start-Process powershell -Verb RunAs -ArgumentList '-ExecutionPolicy Bypass -File .\register_voices_onecore.ps1'"

It is a read-only copy: no existing voice is deleted or overwritten, and keys that already exist are skipped. Then restart the MCP client and call list_voices to see them. From that point they are available offline as Laura and Pablo.

Integration with MCP clients

OpenCode

Add it globally to use it in every project:

opencode mcp add voice-summary -- "C:\path\mcp-voice-summary\.venv\Scripts\python.exe" "C:\path\mcp-voice-summary\mcp_voice_summary.py"

Check the connection with opencode mcp list.

To pin engine, voice and language, edit ~/.config/opencode/opencode.json:

{
  "$schema": "https://opencode.ai/config.json",
  "mcp": {
    "servers": {
      "voice-summary": {
        "type": "local",
        "command": ["mcp-voice-summary"],
        "environment": {
          "VOICE_ENGINE": "edge",
          "VOICE_LANGUAGE": "en",
          "VOICE_RATE": "120"
        }
      }
    }
  }
}

Note: setting VOICE_NAME pins one exact voice and disables automatic language and gender selection. To keep those working, use VOICE_LANGUAGE and VOICE_GENDER instead.

Claude Desktop

claude_desktop_config.json, in %APPDATA%\Claude\:

{
  "mcpServers": {
    "voice-summary": {
      "command": "mcp-voice-summary",
      "env": {
        "VOICE_ENGINE": "edge",
        "VOICE_LANGUAGE": "en",
        "VOICE_RATE": "120"
      }
    }
  }
}

Cursor

.cursor/mcp.json in the project:

{
  "mcpServers": {
    "voice-summary": {
      "command": "mcp-voice-summary",
      "env": { "VOICE_ENGINE": "edge", "VOICE_LANGUAGE": "en" }
    }
  }
}

Any other client

It is a standard stdio MCP server, so declaring the command, the arguments and the environment variables is enough.

Tools

speak_summary(text)

Reads the text out loud through the system speakers. The length is free: the assistant sends a short or a detailed summary depending on the size of the work. If the text exceeds the safety cap, it is truncated and the response includes a notice.

list_voices()

Shows which language and voice are in use right now, plus the voices installed on the system.

set_language(language="", gender="")

Sets the voice language and gender without editing the configuration.

  • language: an ISO 639-1 code (es, en, fr), a regional variant (pt-BR, en-GB), "system" to follow the OS, or "auto" to detect it per summary. Empty keeps the current language.
  • gender: "f" or "m", or "any" to follow the language default. Empty keeps the current gender.

Controlling the summary length

The length is not imposed by the server: the assistant decides it based on how much it did. There are two levels of control, and it is worth understanding the difference.

The assistant chooses the size based on the task. You do this through the instruction you give your assistant, and it is what to use most of the time, because it adapts to the work on its own.

A balanced example:

When calling `speak_summary`, match the summary length to the size of the work:

- Small or single change: one short sentence, 10 to 20 words.
- Medium task or several files: two or three sentences, 30 to 60 words.
- Large task or long project: a summary of 80 to 150 words that walks through
  the main actions performed.

Always write in the first person, without reading out literal code.

Prefer a more conversational tone? Raise the numbers. Prefer concise? Lower them. There is no single correct value: it depends on how long your tasks are and on whether you also read the full response.

2. The server safety cap

MAX_SUMMARY_WORDS is not a length recommendation, it is an emergency brake. It stops an oversized text from turning into a multi-minute announcement. If it is exceeded, the server truncates the summary and says so in the response.

"environment": { "MAX_SUMMARY_WORDS": "400" }
Value Roughly When to use it
0 No truncation Only if you want unlimited announcements
150 ~1 minute You prefer short summaries even on big tasks
400 ~2-3 minutes Default, balanced
800 ~5 minutes Very long work, and waiting does not bother you

A neural voice in English speaks roughly 2.5 words per second, so 100 words is about 40 seconds. Short sentences and pauses matter more for comprehension than the exact word count.

If you need more detail than a long summary allows, the best option is to split the task into several calls instead of raising the cap: that way you hear each phase as it happens, rather than one long block at the end.

3. Adjusting the reading speed

If the summary feels too slow, adjust the speed rather than the length:

"environment": { "VOICE_RATE": "120" }  // 20% faster

Wiring up the behaviour rule

To make the assistant use this automatically, add this instruction to your client's rules. In OpenCode it goes in ~/.config/opencode/AGENTS.md:

## Voice accessibility rule

You have access to the `voice-summary` MCP server with the `speak_summary`
tool. You MUST use it right after you finish modifying code, creating files or
running commands.

When calling it, match the summary length to the size of the work:

- Small or single change: one short sentence, 10 to 20 words.
- Medium task or several files: two or three sentences, 30 to 60 words.
- Large task or long project: a summary of 80 to 150 words that walks through
  the main actions performed.

Always write in the first person, without reading out literal code.

If the user asks to speak another language, use another voice, or a male or female voice, call set_language. The language defaults to the system one, so it does not need configuring.

Privacy and security

This server is meant to run locally on your own machine, started by your MCP client as a child process. It opens no ports and listens on nothing.

What leaves your machine with the edge engine

The edge engine sends the summary text to Microsoft's servers to render the audio. Only the text of the summary is sent, never your code, but keep in mind that a summary can mention file names, function names or error messages that you would rather keep private.

If that matters for your setup, use the sapi5 engine instead: it is fully offline and nothing ever leaves the machine.

Engine Network Audio quality
edge Summary text sent to Microsoft High
sapi5 Nothing leaves the machine Basic, few voices

Hardening already in place

  • No shell. The external audio player is invoked with an argument list, never with shell=True. VOICE_PLAYER must resolve to a real executable through shutil.which.
  • Child processes cannot touch stdin. On a stdio MCP server, stdin and stdout carry the JSON-RPC stream. A player started without stdin=DEVNULL can drain protocol messages meant for the server; this was verified with a reader child and is now blocked, along with timeout and piped output.
  • stdout is never used for logging. Anything printed while the synthesizer is imported is redirected to stderr.
  • Timeouts everywhere. Both the network synthesis call and the player process give up after PLAYBACK_TIMEOUT instead of hanging the worker forever.
  • Input is sanitized. Control characters and markup are stripped, so the text is only ever spoken. edge-tts escapes for SSML itself, but sapi5 hands the text straight to the OS synthesizer.
  • Summaries are never written to the log. On failure the log records the word count and a short sha256 digest, so lines can be correlated without persisting file names, error text or secrets.
  • Credentials are stripped before the text is spoken or sent anywhere. Private key blocks, provider tokens (OpenAI, GitHub, Slack, Google, AWS), JWTs, key=value secrets and email addresses are replaced with [redacted]. A summary spoken in an open office, or sent to a cloud TTS, is a leak channel. Replacements go through placeholders so one pattern cannot mangle another's output. Turn it off with VOICE_REDACT=0 only if summaries never carry anything sensitive.
  • Redaction is a single pass over merged spans, not a chain of substitutions. Matching every pattern against the original text and replacing once means no pattern can re-match another one's output, and the function is idempotent: redacting twice gives the same string.
  • Abuse is bounded on three axes. The queue size stops a burst, the rate limit stops sustained spam that would otherwise slip past a size check, and consecutive identical summaries are skipped. All three run under one lock, so concurrent handlers cannot exceed the limits.
  • Counters, not more tools. list_voices reports accepted, rate-limited, queue-full, deduplicated, truncated, redacted, stale-dropped and error counts. Diagnostics do not deserve a fourth tool when every tool costs context on every turn.
  • Stale notifications are dropped. A summary that waited more than MAX_SUMMARY_AGE is discarded instead of being read out minutes after the work finished.
  • Orphan processes are handled by the standard library. subprocess.run kills the child when the timeout expires, verified with a process that writes its own PID.
  • Voice names are validated against the live catalog before anything is queued, so a typo in VOICE_NAME fails immediately with a helpful message instead of after a wasted network round trip.
  • The queue is bounded by MAX_QUEUE_SIZE. Without a cap, a client calling speak_summary faster than playback could grow the queue without limit and exhaust memory. Requests over the cap are rejected, not silently dropped.
  • Tool annotations declare the side effects: speak_summary is not read-only, not idempotent and touches the outside world.
  • No dynamic code execution. There is no eval, exec, pickle, or os.system anywhere in the source.

Silent mode

Set VOICE_MUTE=1 to keep the server running without making any sound. Summaries are still validated, queued and acknowledged, which makes it safe to leave the server enabled in an office or a meeting. speak_summary answers "Muted." so you can tell the difference from a real playback.

Generic message mode

Set VOICE_MODE=short and the spoken text becomes a fixed "Task completed." whatever the assistant submitted. Nothing from the summary reaches the speakers or the cloud engine.

Use it when summaries may mention confidential details, or when you would rather not hear the raw text at all. The trade-off is that you lose the information: you get "something finished", not what.

Before every release

.venv\Scripts\python.exe pre_release_check.py

Runs the test suite, pip-audit, a stdout purity check, a scan for dangerous calls, the tool catalog budget, and a check that the counters expose no text. Exits non-zero on failure so it can gate a release.

When you add a secret pattern

The redaction tests are driven by one table at the top of tests/test_server.py. Adding a pattern means adding a row to SECRET_CASES (a positive case plus the fragment that must not survive) and, when it could plausibly match legitimate text, a row to BENIGN_CASES. Idempotency, marker safety, multiline handling and spacing or case variants are then checked automatically for every case.

Tests

The suite covers input handling, queue limits, voice resolution, the platform fallbacks, the privacy of the logs and the token budget. It never plays real audio, so it runs offline in a couple of seconds:

python -m unittest discover -s tests

Known limitations

  • Redaction is defensive, not exhaustive. It catches private key blocks, the common provider token formats, JWTs, key=value secrets and email addresses. It will not catch a base64 blob with no marker, a secret spelled with Unicode lookalike characters, or an unusual internal format. Do not rely on it as your only control. Also, because spans are replaced whole, the surrounding context of a match is dropped too.
  • Voice and language are process-global. set_language changes the state of the whole server, so if two MCP sessions talk to the same server process, one can change the voice for the other. This does not matter for the intended single-user local setup, but it is not per-session isolation.
  • Unexpected tool arguments are ignored, not rejected. The SDK derives the schema from the function signature and offers no way to set additionalProperties: false. Passing tts_endpoint or api_key to speak_summary is silently dropped and has no effect, verified by calling the tool with extra fields. Ignored is safe here because no argument the model can pass reaches the network, the filesystem or a command line.
  • Sync tool bodies run on a worker thread, not the event loop: the SDK dispatches them through anyio.to_thread.run_sync, so the redaction regexes cannot stall the server.
  • VOICE_PLAYER executes a program. That is its purpose, and it is only read from your own configuration, but do not build it from untrusted input.
  • langdetect accuracy. Covered in detail in Automatic language detection.

Implementation notes

Details that are not obvious and worth knowing before modifying this:

  • pyttsx3 blocks forever if the engine is created in one thread and used in another. COM is apartment-threaded. That is why the engine is initialized lazily and always used from the same worker thread.
  • A single worker thread consumes a queue, so two consecutive utterances never overlap or cut each other off.
  • SAPI5's runAndWait can return before the utterance finishes, so there is an active wait using isBusy(). The edge engine uses a blocking MCI playback, so its timing is exact.
  • The worker thread is a daemon: speech never prevents process shutdown.
  • The pyttsx3, edge_tts and langdetect imports are lazy, so the server starts even if any of them is missing.
  • Voice selection by language has three fallback levels: the curated voice for the language, then any voice of its preferred regional variant, then any voice of that language. That is why a language without a curated voice still works.
  • Any audio failure is caught and logged. The MCP server never goes down because it could not speak.

Troubleshooting

The tool does not appear in my client. Restart the MCP client after editing its configuration; most clients read the config only at startup. Confirm the command resolves: run mcp-voice-summary in a terminal and you should see no output at all, since it waits silently on stdio.

Nothing is spoken. Check the system volume and the default output device. With the edge engine, verify there is an internet connection. To separate configuration from playback, call list_voices and check the voice it reports: if that looks right, the problem is audio output rather than the server.

A secret is still spoken. Check that VOICE_REDACT is not set to 0. Redaction is defensive and cannot catch every format, so for genuinely sensitive material prefer VOICE_MODE=short, which replaces the text with a fixed message before anything is spoken or sent.

The voice gets cut off or messages overlap. Check that only one instance of the MCP server is running.

No audio player was found. Only affects Linux and macOS with the edge engine. Install ffmpeg, mpg123 or vlc, or set VOICE_PLAYER. This does not happen on Windows, which uses MCI.

The sapi5 engine finds no voices on Linux. Install the system synthesizer: sudo apt install espeak-ng libespeak-ng1. Bear in mind that espeak voices are far inferior to edge-tts.

ImportError: No module named mcp.server.fastmcp. That is v1 code. Install mcp>=2.0, where FastMCP was renamed to MCPServer. This project needs 2.x.

It speaks in an unexpected language. If the text comes out with a strange accent, VOICE_LANGUAGE=auto most likely misdetected it. Detection on very short text is known to be unreliable: pin it with set_language("fr") or with VOICE_LANGUAGE. See the table of detection failures.

Windows voices sound basic. Run register_voices_onecore.ps1 as administrator once, then use VOICE_NAME=Laura or VOICE_NAME=Pablo.

License

MIT. See LICENSE.

Metadata

Release files for mcp-voice-summary 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for mcp-voice-summary 0.1.0
File Size Uploaded
mcp_voice_summary-0.1.0.tar.gz 42.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for mcp-voice-summary 0.1.0
File Interpreter ABI Platform
mcp_voice_summary-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 68.7 kB

Release files / mcp_voice_summary-0.1.0.tar.gz

Download URL mcp_voice_summary-0.1.0.tar.gz
Size 42.0 kB
Tags Source
SHA-256 checksum
How to use checksums
8e75440affb05e105d00dafbb794d8471f481c8cdd68d1dc81b13d92878aa5b3
BLAKE2b-256 checksum
How to use checksums
490632ba7269dc699dbeafd362b987d5928a923b9bb82450a03dff1a7256cdbf
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.3

Release files / mcp_voice_summary-0.1.0-py3-none-any.whl

Download URL mcp_voice_summary-0.1.0-py3-none-any.whl
Size 26.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
8eed4293f8ed95186490a32fa8e306cc352b2ac3342ebbfd37cd99770998cb4b
BLAKE2b-256 checksum
How to use checksums
b53c7093ec009a5c2c3bd2efa4921fe94a8b63170058b6387669e3fc598d9ba9
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.3

Release history Release notifications | RSS feed

0.1.1

2 release files

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page