Skip to main content

vocalize

CI

A command-line tool that turns text, markdown files, or piped stdin into natural-sounding speech using the ElevenLabs API — plus a hook that wires it directly into Claude Code, so Claude's responses get read aloud automatically in your terminal or IDE.

Why this exists

Text-to-speech readers are good at voices and bad at structure. Point one at a markdown report and it reads a table cell-by-cell, left to right, with no sense of which row or column you're in — "Q1. 4.2 million. Q2. 5.1 million" instead of "for Q1, revenue is 4.2 million." Headings, bullet lists, and inline code fare the same way: read exactly as typed, syntax and all.

vocalize fixes the part of that problem that's actually fixable without a vision model: a preprocessing pass (vocalize/preprocess.py) rewrites markdown into short, declarative sentences before it ever reaches the TTS API — tables become "for X, Y is Z" sentences, bullets become "First, ... Second, ...", links keep their text and drop the URL, and fenced code blocks are replaced with a spoken placeholder instead of being read character by character. It's a text transform, so it's fully unit tested without any API key or network access (see tests/test_preprocess.py).

Install

pipx install vocalize-cli

(or uvx --from vocalize-cli vocalize for a one-off run without installing anything). The package is published on PyPI as vocalize-cli; the command it installs is still vocalize.

For a from-source or dev install:

git clone <this-repo>
cd vocalize
pip install -e .

Get a free ElevenLabs API key at elevenlabs.io/app/settings/api-keys (free tier: 10,000 characters/month, API access included, no commercial license). Then either:

export ELEVENLABS_API_KEY=your-key-here

or copy .env.example to .env and fill it in (requires the optional python-dotenv extra: pip install -e ".[dotenv]").

Usage

# Speak a string directly
vocalize speak "Hello, this is a test."

# Speak a markdown file — tables and formatting get flattened first
vocalize speak-file report.md

# Pipe anything in
cat notes.md | vocalize speak-file -

# List available voices and grab an ID
vocalize voices

# Use a specific voice/model, save without playing
vocalize speak-file report.md --voice <voice-id> --model eleven_flash_v2_5 \
  --output out.mp3 --no-play

# Cap how much gets sent (handy for free-tier character budgets)
vocalize speak-file long-report.md --max-chars 2000

# Skip the markdown flattening entirely
vocalize speak "raw **markdown** stays raw" --raw

Every synthesis result is cached on disk under ~/.cache/vocalize/, keyed by a hash of (text, voice, model, format) — re-running the same command twice doesn't burn API quota twice.

Claude Code integration

The hook scripts ship in the git repository, not the PyPI package — clone the repo to install the hook (it shells out to the vocalize command, so a pipx-installed CLI plus a cloned repo works fine together).

hooks/claude_stop_hook.py is a Claude Code Stop hook: a script Claude Code runs every time it finishes a response. This one reads the transcript, pulls out Claude's last message, and pipes it through the same vocalize CLI — so it works identically whether Claude Code is running in a bare terminal or inside an IDE's integrated terminal (VS Code, Cursor, etc.), since both use the same ~/.claude/settings.json hook config.

On-demand mode. If you'd rather trigger speech yourself than have every response spoken, skip the install and run the script with --latest. It finds your most recent Claude Code response — in any session — and speaks that one:

python3 hooks/claude_stop_hook.py --latest

Combine it with VOCALIZE_MAX_CHARS to control how much gets read.

To install it as an automatic hook instead:

python3 hooks/install_hook.py

This merges a Stop hook entry into ~/.claude/settings.json (backing up the existing file first) rather than overwriting your other hooks. Every Claude Code response after that gets spoken aloud automatically. Uninstall by removing the vocalize entry from the Stop array in that file.

By default the hook truncates each response to 500 characters before speaking it (DEFAULT_MAX_CHARS in claude_stop_hook.py) — a Stop hook fires after every turn, so a long response would burn through the ElevenLabs free-tier quota fast. Override with VOCALIZE_MAX_CHARS in the environment.

The hook looks up the vocalize binary on PATH, but Claude Code hooks run in Claude Code's own environment, not your interactive shell — if vocalize was installed into a virtualenv that isn't on that PATH, set VOCALIZE_BIN to the full path (e.g. /path/to/.venv/bin/vocalize) to point the hook at it directly.

How it's built

Four decisions shaped the design:

  • The markdown flattener is a pure function. The hardest logic in the project — deciding what a table, list, or code block should sound like — takes a string and returns a string. No I/O, no client, no key. That's why it has the deepest test coverage in the repo, including the edge cases that bit during review: prose containing a stray |, single-dash GFM separators, ragged rows, duplicate column names.
  • One code path for humans and hooks. The Claude Code hook doesn't reimplement synthesis; it shells out to the same vocalize CLI you'd type by hand (with an -- argv guard so a response starting with a bullet isn't parsed as a flag). Anything the hook can do, you can reproduce and debug from your own terminal.
  • The hook may fail; the session may not. Every failure path in the Stop hook logs one line to stderr and exits 0. A dead API key or a hung request costs you the audio, never the coding session.
  • The cache is an optimization, never a failure source. Synthesis results are content-addressed on disk; an unreadable or unwritable cache degrades to a fresh API call instead of an error.

Architecture

vocalize/
  __init__.py     # package version
  __main__.py     # python -m vocalize entry point
  preprocess.py   # markdown -> speakable text (pure function, fully unit tested)
  config.py       # API key resolution: --api-key > $ELEVENLABS_API_KEY > .env
  exceptions.py   # VocalizeError / TTSRequestError
  tts.py          # ElevenLabs API wrapper + disk cache (client is injected, so
                   # it's mockable in tests without hitting the network)
  audio.py        # save to disk + play via the OS's native player
                   # (afplay / mpg123 / ffplay / PowerShell, whichever exists)
  cli.py          # click-based CLI wiring the above together
hooks/
  claude_stop_hook.py  # Claude Code Stop hook -> calls the vocalize CLI
  install_hook.py      # safely merges the hook into ~/.claude/settings.json
tests/                  # pytest, all mocked — no API key needed to run these

Testing

pip install -e ".[dev]"
pytest

All tests run offline: the ElevenLabs client is dependency-injected into tts.py, so tests pass in a fake client instead of hitting the real API.

Known limitations

  • Charts and images aren't described. Flattening markdown tables is a text problem; a rendered chart is an image, and describing it well needs a vision model in the loop, not a text transform. Out of scope for this project, but a natural next step — pipe the image through a vision-capable model first, feed its description into vocalize in place of the chart.
  • Free tier is 10,000 characters/month — plenty for reading a handful of documents aloud, not for continuous use. --max-chars and the disk cache both help stretch it.
  • Table flattening handles standard GFM pipe tables; it doesn't attempt to handle merged cells or nested tables (rare enough in practice that it wasn't worth the complexity).
  • Windows playback is untested. The PowerShell SoundPlayer fallback only plays WAV, so the mp3 files this tool generates likely won't play there. Use --no-play and open the saved file with whatever's on hand.
  • The disk cache under ~/.cache/vocalize grows unbounded. It's content-addressed (keyed by a hash of text, voice, model, and format), so it's always safe to delete some or all of it — nothing will break, you'll just re-pay for a re-synthesized clip.
  • --api-key on the command line is visible to other local processes (anything that can run ps). Prefer the ELEVENLABS_API_KEY environment variable or a .env file instead.
  • vocalize voices lists only the first page of results from the ElevenLabs API.

License

MIT

Release files for vocalize-cli 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for vocalize-cli 0.1.1
File Size Uploaded
vocalize_cli-0.1.1.tar.gz 23.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for vocalize-cli 0.1.1
File Interpreter ABI Platform
vocalize_cli-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size:39.0 kB

Release files / vocalize_cli-0.1.1.tar.gz

Download URL vocalize_cli-0.1.1.tar.gz
Size 23.0 kB
Tags Source
SHA-256 checksum
How to use checksums
2ef4a2f5a15b3bf880fe4144d7561db81bee39a0dd969498d541fbfe6f8a94f2
BLAKE2b-256 checksum
How to use checksums
7a1c6fafe61a07576cd335c63b3158b5ae86bb2dabeea7c7165d32df366283a4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.7

Release files / vocalize_cli-0.1.1-py3-none-any.whl

Download URL vocalize_cli-0.1.1-py3-none-any.whl
Size 16.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
692c49c5498d64a9322fe1476217ab9a64a5c7dbda7088680aa7ba6d6880fe48
BLAKE2b-256 checksum
How to use checksums
f8468153533690c2801a127aa3d02bb33a1dbfeae2060734e21671b5dcb1c987
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.7

Release history Release notifications | RSS feed

0.9.1

2 release files

0.9.0

2 release files

0.8.1

2 release files

0.8.0

2 release files

0.7.1

2 release files

0.7.0

2 release files

0.6.0

2 release files

0.5.0

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.1

2 release files

0.2.0

2 release files

This release

0.1.1 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page