Skip to main content

vocalize

CI

A command-line tool that turns text, markdown files, or piped stdin into natural-sounding speech using the ElevenLabs API — plus a hook that wires it directly into Claude Code, so Claude's responses get read aloud automatically in your terminal or IDE.

Why this exists

Text-to-speech readers are good at voices and bad at structure. Point one at a markdown report and it reads a table cell-by-cell, left to right, with no sense of which row or column you're in — "Q1. 4.2 million. Q2. 5.1 million" instead of "for Q1, revenue is 4.2 million." Headings, bullet lists, and inline code fare the same way: read exactly as typed, syntax and all.

vocalize fixes the part of that problem that's actually fixable without a vision model: a preprocessing pass (vocalize/preprocess.py) rewrites markdown into short, declarative sentences before it ever reaches the TTS API — tables become "for X, Y is Z" sentences, bullets become "First, ... Second, ...", links keep their text and drop the URL, and fenced code blocks are replaced with a spoken placeholder instead of being read character by character. It's a text transform, so it's fully unit tested without any API key or network access (see tests/test_preprocess.py).

Install

pipx install vocalize-cli

(or uvx --from vocalize-cli vocalize for a one-off run without installing anything). The package is published on PyPI as vocalize-cli; the command it installs is still vocalize.

For a from-source or dev install:

git clone <this-repo>
cd vocalize
pip install -e .

Get a free ElevenLabs API key at elevenlabs.io/app/settings/api-keys (free tier: 10,000 characters/month, API access included, no commercial license). Then either:

export ELEVENLABS_API_KEY=your-key-here

or copy .env.example to .env and fill it in (requires the optional python-dotenv extra: pip install -e ".[dotenv]").

Usage

# Speak a string directly
vocalize speak "Hello, this is a test."

# Speak a markdown file — tables and formatting get flattened first
vocalize speak-file report.md

# Pipe anything in
cat notes.md | vocalize speak-file -

# List available voices and grab an ID
vocalize voices

# Use a specific voice/model, save without playing
vocalize speak-file report.md --voice <voice-id> --model eleven_flash_v2_5 \
  --output out.mp3 --no-play

# Cap how much gets sent (handy for free-tier character budgets)
vocalize speak-file long-report.md --max-chars 2000

# Slow it down a little
vocalize speak-file report.md --speed 0.9

# Skip the markdown flattening entirely
vocalize speak "raw **markdown** stays raw" --raw

Every synthesis result is cached on disk under ~/.cache/vocalize/, keyed by a hash of (text, voice, model, format, speed) — re-running the same command twice doesn't burn API quota twice.

Configuration

Each setting is resolved on its own, taking the first source that supplies it: CLI flag, then environment variable, then config file, then the built-in default.

There's an interactive way to set the file up, if you'd rather not write TOML by hand. It walks through three lists — voice (with a live preview of the highlighted one), model, and speed — shows you a summary, and writes the config file below. Unrecognised top-level keys already in that file are carried through; comments and layout are not preserved. A file containing a TOML table or array is left alone entirely, with a message saying to edit it by hand.

vocalize config

Hotkeys: / or k/j move, Enter selects, p previews the highlighted voice, m types a value by hand, q or Esc cancels without writing anything.

Setting Flag Env var Config file key Default
Voice ID --voice VOCALIZE_VOICE voice 21m00Tcm4TlvDq8ikWAM ("Rachel")
Model ID --model VOCALIZE_MODEL model eleven_multilingual_v2
Speed --speed VOCALIZE_SPEED speed unset — the API's own 1.0
Max characters --max-chars VOCALIZE_MAX_CHARS (hook only) not read from the config file unset on the CLI; 500 in the hook
Hook binary VOCALIZE_BIN not read from the config file vocalize as found on PATH

The config file is TOML at $XDG_CONFIG_HOME/vocalize/config.toml, falling back to ~/.config/vocalize/config.toml. Flat keys, no sections:

voice = "21m00Tcm4TlvDq8ikWAM"
model = "eleven_flash_v2_5"
speed = 0.95

Not having a config file is normal and silent. A file that isn't valid TOML is an error naming the file; a key that isn't recognised is a warning on stderr, so a typo doesn't pass unnoticed but doesn't stop the run either. speed must be a number between 0.7 and 1.2 — anything else is a one-line error naming the source it came from.

The API key is separate and never read from this file: use --api-key, ELEVENLABS_API_KEY, or a .env file.

Claude Code integration

The hook scripts ship in the git repository, not the PyPI package — clone the repo to install the hook (it shells out to the vocalize command, so a pipx-installed CLI plus a cloned repo works fine together).

hooks/claude_stop_hook.py is a Claude Code Stop hook: a script Claude Code runs every time it finishes a response. This one reads the transcript, pulls out Claude's last message, and pipes it through the same vocalize CLI — so it works identically whether Claude Code is running in a bare terminal or inside an IDE's integrated terminal (VS Code, Cursor, etc.), since both use the same ~/.claude/settings.json hook config.

On-demand mode. If you'd rather trigger speech yourself than have every response spoken, skip the install and run the script with --latest. It finds your most recent Claude Code response — in any session — and speaks that one:

python3 hooks/claude_stop_hook.py --latest

Combine it with VOCALIZE_MAX_CHARS to control how much gets read.

To install it as an automatic hook instead:

python3 hooks/install_hook.py

This merges a Stop hook entry into ~/.claude/settings.json (backing up the existing file first) rather than overwriting your other hooks. Every Claude Code response after that gets spoken aloud automatically. Uninstall by removing the vocalize entry from the Stop array in that file.

By default the hook truncates each response to 500 characters before speaking it (DEFAULT_MAX_CHARS in claude_stop_hook.py) — a Stop hook fires after every turn, so a long response would burn through the ElevenLabs free-tier quota fast. Override with VOCALIZE_MAX_CHARS in the environment.

The hook looks up the vocalize binary on PATH, but Claude Code hooks run in Claude Code's own environment, not your interactive shell — if vocalize was installed into a virtualenv that isn't on that PATH, set VOCALIZE_BIN to the full path (e.g. /path/to/.venv/bin/vocalize) to point the hook at it directly.

How it's built

Four decisions shaped the design:

  • The markdown flattener is a pure function. The hardest logic in the project — deciding what a table, list, or code block should sound like — takes a string and returns a string. No I/O, no client, no key. That's why it has the deepest test coverage in the repo, including the edge cases that bit during review: prose containing a stray |, single-dash GFM separators, ragged rows, duplicate column names.
  • One code path for humans and hooks. The Claude Code hook doesn't reimplement synthesis; it shells out to the same vocalize CLI you'd type by hand (with an -- argv guard so a response starting with a bullet isn't parsed as a flag). Anything the hook can do, you can reproduce and debug from your own terminal.
  • The hook may fail; the session may not. Every failure path in the Stop hook logs one line to stderr and exits 0. A dead API key or a hung request costs you the audio, never the coding session.
  • The cache is an optimization, never a failure source. Synthesis results are content-addressed on disk; an unreadable or unwritable cache degrades to a fresh API call instead of an error.

Architecture

vocalize/
  __init__.py     # package version
  __main__.py     # python -m vocalize entry point
  preprocess.py   # markdown -> speakable text (pure function, fully unit tested)
  config.py       # API key resolution + settings: flag > env > config.toml > default
  exceptions.py   # VocalizeError / TTSRequestError
  tts.py          # ElevenLabs API wrapper + disk cache (client is injected, so
                   # it's mockable in tests without hitting the network)
  audio.py        # save to disk + play via the OS's native player
                   # (afplay / mpg123 / ffplay / PowerShell, whichever exists)
  cli.py          # click-based CLI wiring the above together
hooks/
  claude_stop_hook.py  # Claude Code Stop hook -> calls the vocalize CLI
  install_hook.py      # safely merges the hook into ~/.claude/settings.json
tests/                  # pytest, all mocked — no API key needed to run these

Testing

pip install -e ".[dev]"
pytest

All tests run offline: the ElevenLabs client is dependency-injected into tts.py, so tests pass in a fake client instead of hitting the real API.

Known limitations

  • Charts and images aren't described. Flattening markdown tables is a text problem; a rendered chart is an image, and describing it well needs a vision model in the loop, not a text transform. Out of scope for this project, but a natural next step — pipe the image through a vision-capable model first, feed its description into vocalize in place of the chart.
  • Free tier is 10,000 characters/month — plenty for reading a handful of documents aloud, not for continuous use. --max-chars and the disk cache both help stretch it.
  • Table flattening handles standard GFM pipe tables; it doesn't attempt to handle merged cells or nested tables (rare enough in practice that it wasn't worth the complexity).
  • Windows playback is untested. The PowerShell SoundPlayer fallback only plays WAV, so the mp3 files this tool generates likely won't play there. Use --no-play and open the saved file with whatever's on hand.
  • The disk cache under ~/.cache/vocalize grows unbounded. It's content-addressed (keyed by a hash of text, voice, model, format, and speed), so it's always safe to delete some or all of it — nothing will break, you'll just re-pay for a re-synthesized clip.
  • --api-key on the command line is visible to other local processes (anything that can run ps). Prefer the ELEVENLABS_API_KEY environment variable or a .env file instead.
  • vocalize voices lists only the first page of results from the ElevenLabs API.

License

MIT

Release files for vocalize-cli 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for vocalize-cli 0.2.0
File Size Uploaded
vocalize_cli-0.2.0.tar.gz 32.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for vocalize-cli 0.2.0
File Interpreter ABI Platform
vocalize_cli-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size:56.0 kB

Release files / vocalize_cli-0.2.0.tar.gz

Download URL vocalize_cli-0.2.0.tar.gz
Size 32.9 kB
Tags Source
SHA-256 checksum
How to use checksums
70a0027f5dc4a2756fa1795adc7265d9c8802d0b05b748049b930019faae4dc5
BLAKE2b-256 checksum
How to use checksums
217561585df31f50b9a9a9e80e4ed2bf14ae120f7bcd3c2cc443c8cf95a2b964
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.7

Release files / vocalize_cli-0.2.0-py3-none-any.whl

Download URL vocalize_cli-0.2.0-py3-none-any.whl
Size 23.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
1ae0d226475ab00d234529d992eb7a161cc398db8f570d2897ea66eb07f7cf5e
BLAKE2b-256 checksum
How to use checksums
ddef6bb4d4c2f16e98790e22479a4f18df83f670cc9b700027019d380877636c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.7

Release history Release notifications | RSS feed

0.9.1

2 release files

0.9.0

2 release files

0.8.1

2 release files

0.8.0

2 release files

0.7.1

2 release files

0.7.0

2 release files

0.6.0

2 release files

0.5.0

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.1

2 release files

This release

0.2.0 This release

2 release files

0.1.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page