Skip to main content

gensay

PyPI - Version PyPI - Python Version

A multi-provider text-to-speech (TTS) tool that implements the Apple macOS /usr/bin/say command interface while supporting multiple TTS backends including Chatterbox (local AI), OpenAI, ElevenLabs, Deepgram, and Amazon Polly.

Features

  • macOS say Compatible: Drop-in replacement for the macOS say command with identical CLI interface
  • Multiple TTS Providers: Extensible provider system with support for:
  • Smart Text Chunking: Intelligently splits long text for optimal TTS processing
  • Audio Caching: Automatic caching with LRU eviction to speed up repeated synthesis
  • Progress Tracking: Built-in progress bars with tqdm and customizable callbacks
  • Multiple Audio Formats: Support for AIFF, WAV, M4A, MP3, CAF, FLAC, AAC, OGG
  • Background Pre-caching: Queue and cache audio chunks in the background (Chatterbox only)
  • Interactive REPL Mode: Start an interactive session with provider initialized once for repeated use
  • Warm Inference Daemon: Keep local AI models loaded in a background process; ad-hoc gensay calls reuse the warm model via a Unix socket
  • Offline Resilience: Cloud providers (ElevenLabs, Deepgram, OpenAI, Polly) automatically fall back to macOS say when the network is unreachable

Table of Contents

Installation

It's 2026, use uv

gensay is intended to be used as a CLI tool that is a drop-in replacement to the macOS say CLI.

System Dependencies (ElevenLabs provider only)

PortAudio is required if you plan to use the ElevenLabs provider. The pyaudio dependency needs the PortAudio C library to compile successfully.

Other providers (macOS, OpenAI, Amazon Polly, Chatterbox, Deepgram) do not require PortAudio.

Homebrew (macOS):

brew install portaudio

Nix:

nix-env -iA nixpkgs.portaudio

Install gensay

# Install as a tool
uv tool install gensay

# With extras: ElevenLabs provider (requires PortAudio, see above)
uv tool install 'gensay[elevenlabs]'

# Deepgram provider (Flux TTS / Aura) ships in the core install; this extra
# only adds keyring support for `gensay config set deepgram.api_key`
uv tool install 'gensay[deepgram]'

# With extras: Chatterbox provider (local Text-to-Speech model, ~2GB PyTorch dependencies)
uv tool install 'gensay[chatterbox]' \
  --with git+https://github.com/anthonywu/chatterbox.git@allow-dep-updates

# Or add to your project
uv add gensay

# From source (with automatic PortAudio path configuration)
git clone https://github.com/anthonywu/gensay
cd gensay
just setup

Optional Dependencies

# Audio format conversion (for non-native formats like MP3, OGG, FLAC)
# Requires ffmpeg installed on system
pip install 'gensay[audio-formats]'

# Install all optional dependencies
pip install 'gensay[all]'

DInstallation Help:

For developer/maintainer installation, just setup automatically configures PortAudio and FFmpeg paths for both Nix and Homebrew.

Developer/Maintainer Build Dependencies

PortAudio Paths (for ElevenLabs)

Homebrew:

export C_INCLUDE_PATH="$(brew --prefix portaudio)/include:$C_INCLUDE_PATH"
export LIBRARY_PATH="$(brew --prefix portaudio)/lib:$LIBRARY_PATH"

Nix:

export C_INCLUDE_PATH="$(nix-build '<nixpkgs>' -A portaudio --no-out-link)/include:$C_INCLUDE_PATH"
export LIBRARY_PATH="$(nix-build '<nixpkgs>' -A portaudio --no-out-link)/lib:$LIBRARY_PATH"

Then install into local venv:

uv sync --all-extras
# temporarily, we have to use a special release of chatterbox library to allow for dependency resolution
uv pip install git+https://github.com/anthonywu/chatterbox.git@allow-dep-updates

FFmpeg Library Path (for Chatterbox on macOS)

Chatterbox uses TorchCodec which requires FFmpeg libraries at process start. Since 0.5.0, gensay self-heals: if FFmpeg is detectable (Nix store or Homebrew), it re-executes itself once with DYLD_LIBRARY_PATH set correctly — no manual export needed for gensay -p chatterbox, the daemon, or daemon start children.

If FFmpeg can't be auto-detected, set it manually:

Homebrew:

export DYLD_LIBRARY_PATH="$(brew --prefix ffmpeg)/lib:$DYLD_LIBRARY_PATH"
gensay --provider chatterbox "Hello"

Nix:

# Find the ffmpeg-lib output in the Nix store
FFMPEG_LIB=$(nix-store -qR "$(which ffmpeg)" | grep 'ffmpeg.*-lib$')
export DYLD_LIBRARY_PATH="$FFMPEG_LIB/lib:$DYLD_LIBRARY_PATH"
gensay --provider chatterbox "Hello"

Note: DYLD_LIBRARY_PATH must be set before the Python process starts; it cannot be set from within Python.

Quick Start

# Basic usage - speaks the text
gensay "Hello, world!"

# Use specific voice
gensay -v Samantha "Hello from Samantha"

# Save to audio file
gensay -o greeting.m4a "Welcome to gensay"

# List available voices (two ways)
gensay -v '?'
gensay --list-voices

Command Line Usage

Basic Options

# Speak text
gensay "Hello, world!"

# Read from file
gensay -f document.txt

# Read from stdin
echo "Hello from pipe" | gensay -f -

# Specify voice
gensay -v Alex "Hello from Alex"

# Adjust speech rate (words per minute)
gensay -r 200 "Speaking faster"

# Save to file
gensay -o output.m4a "Save this speech"

# Specify audio format
gensay -o output.wav --format wav "Different format"

Provider Selection

# Use macOS native say command
gensay --provider macos "Using system TTS"

# List voices for specific provider
gensay --provider macos --list-voices
gensay --provider mock --list-voices

# Use mock provider for testing
gensay --provider mock "Testing without real TTS"

# Use Chatterbox explicitly
gensay --provider chatterbox "Local AI voice"

# Default provider depends on platform
gensay "Hello"  # Uses 'macos' on macOS, 'chatterbox' on other platforms

Advanced Options

# Show progress bar
gensay --progress "Long text with progress tracking"

# Pre-cache audio chunks in background
gensay --provider chatterbox --cache-ahead "Pre-process this text"

# Adjust chunk size
gensay --chunk-size 1000 "Process in larger chunks"

# Cache management
gensay --cache-stats     # Show cache statistics
gensay --clear-cache     # Clear all cached audio
gensay --no-cache "Text" # Disable cache for this run

Interactive Modes and Performance Optimization

Warm Inference Daemon (recommended for Chatterbox)

Local AI providers (Chatterbox) pay multi-second model load on every process start. The daemon keeps the provider resident; subsequent gensay invocations are cheap RPCs over a user-local Unix socket.

Cloud providers (ElevenLabs, OpenAI, Polly) are intentionally not hostable in the daemon — their client init is milliseconds, so there is nothing worth keeping warm.

Defaults: daemon vs user config

When user config and the daemon meet, resolution is:

Situation Result
daemon start without -p Provider from daemon.provider config key → top-level provider config (if daemon-hostable) → chatterbox. A cloud speak-default (e.g. provider = "elevenlabs") is ignored here.
Bare gensay "…" with a cloud default provider Cloud provider speaks directly (cold path); warm routing only engages for warm-eligible providers.
--via-daemon / warm routing without explicit -p The running daemon decides. If your configured provider default differs, gensay prints a warning and forwards the request with no provider assertion.
--via-daemon with explicit -p The provider is asserted; a daemon hosting a different one answers provider_mismatch (the fail-loud path for "you really meant it").

Rule of thumb: config picks your defaults, -p makes a claim, the daemon's resident model wins unless you make a claim.

# Start once per session (preloads model, detaches)
gensay daemon start -p chatterbox

# Ad-hoc speak — auto-routes to the daemon when provider is warm-eligible
gensay -p chatterbox "Build finished"
gensay -p chatterbox "Need your input"

# Force / forbid daemon routing
gensay --via-daemon -p chatterbox "must use daemon"
gensay --no-daemon -p chatterbox "cold path this time"
gensay --auto-daemon -p chatterbox "start daemon if missing"

# Lifecycle
gensay daemon status
gensay daemon status --json
gensay daemon stop

# Foreground (launchd / debugging)
gensay daemon run -p chatterbox

Environment knobs (twelve-factor):

Env Meaning
GENSAY_RUNTIME_DIR Directory for socket + pidfile
GENSAY_SOCKET Explicit socket path
GENSAY_VIA_DAEMON 1 = require daemon
GENSAY_NO_DAEMON 1 = force cold path
GENSAY_AUTO_DAEMON 1 = auto-start when missing
GENSAY_DAEMON_IDLE_UNLOAD_S Unload model after idle seconds (0 = never)
GENSAY_DAEMON_IDLE_EXIT_S Exit process after idle seconds (0 = never)

Socket location defaults to platformdirs.user_runtime_dir("gensay") (not /tmp).

User config (per-user defaults)

Bare gensay "hello" can pick up preferred flags from a TOML file in the platform config dir (XDG on Linux):

Platform Default path
Linux ~/.config/gensay/config.toml ($XDG_CONFIG_HOME/gensay/…)
macOS ~/Library/Application Support/gensay/config.toml
Override GENSAY_CONFIG=/path/to/config.toml

Precedence: CLI flags > GENSAY_* env > config file > built-ins.

# Scaffold an annotated example, or set keys directly
gensay config init
gensay config path
gensay config keys
gensay config set provider chatterbox
gensay config set auto_daemon true
gensay config set daemon.provider chatterbox
gensay config get provider
gensay config get auto_daemon
gensay config unset voice
gensay config show
gensay config show --json

Example config.toml:

provider = "chatterbox"
voice = "default"
rate = 150
auto_daemon = true

[daemon]
provider = "chatterbox"
idle_unload_s = 0

After that, gensay "Build finished" uses chatterbox (and auto-starts the warm daemon if configured) without repeating flags.

Provider API keys

<provider>.api_key keys are secrets: config set stores them in the OS keychain (via keyring), never in the plaintext TOML file.

gensay config set elevenlabs.api_key '<your-key>'   # → Keychain/Secret Service
gensay config show    # prints "elevenlabs.api_key = (stored in OS keychain)"
gensay config unset elevenlabs.api_key              # removes from keychain

Runtime precedence: provider env var (ELEVENLABS_API_KEY, also via .env) > OS keychain.

REPL Mode

Start an interactive session where the provider is initialized once and reused for each prompt (in-process; no daemon required).

# Start REPL mode (--repl, --interactive, and -i are all equivalent)
gensay --repl
gensay --interactive
gensay -i

# With a specific provider and voice
gensay --provider openai -v nova --repl

# Chatterbox with REPL (keeps model loaded in this terminal)
gensay -p chatterbox -i

In REPL mode:

  • Type text and press Enter to speak it
  • Type exit or quit to exit
  • Press Ctrl+C or Ctrl+D to exit

Python API

Basic Usage

from gensay import ChatterboxProvider, TTSConfig, AudioFormat

# Create provider
provider = ChatterboxProvider()

# Speak text
provider.speak("Hello from Python")

# Save to file
provider.save_to_file("Save this", "output.m4a")

# List voices
voices = provider.list_voices()
for voice in voices:
    print(f"{voice['id']}: {voice['name']}")

Advanced Configuration

from gensay import ChatterboxProvider, TTSConfig, AudioFormat

# Configure TTS
config = TTSConfig(
    voice="default",
    rate=150,
    format=AudioFormat.M4A,
    cache_enabled=True,
    extra={
        'show_progress': True,
        'chunk_size': 500
    }
)

# Create provider with config
provider = ChatterboxProvider(config)

# Add progress callback
def on_progress(progress: float, message: str):
    print(f"Progress: {progress:.0%} - {message}")

config.progress_callback = on_progress

# Use the configured provider
provider.speak("Text with all options configured")

Text Chunking

from gensay import chunk_text_for_tts, TextChunker

# Simple chunking
chunks = chunk_text_for_tts(long_text, max_chunk_size=500)

# Advanced chunking with custom strategy
chunker = TextChunker(
    max_chunk_size=1000,
    strategy="paragraph",  # or "sentence", "word", "character"
    overlap_size=50
)
chunks = chunker.chunk_text(document)

Provider Configurations

ElevenLabs

  1. Install the optional dependency (requires PortAudio):
    pip install 'gensay[elevenlabs]'
    
  2. Get an API key from ElevenLabs
  3. Set the environment variable:
    export ELEVENLABS_API_KEY="your-api-key"
    
# List ElevenLabs voices
gensay --provider elevenlabs --list-voices

# Use a specific ElevenLabs voice
gensay --provider elevenlabs -v Rachel "Hello from ElevenLabs"

# Save to file with high quality
gensay --provider elevenlabs -o speech.mp3 "High quality AI speech"

OpenAI TTS

  1. Get an API key from OpenAI Platform
  2. Set the environment variable:
    export OPENAI_API_KEY="sk-..."
    
# List OpenAI voices
gensay --provider openai --list-voices

# Use a specific voice (alloy, ash, ballad, coral, echo, fable, onyx, nova, sage, shimmer)
gensay --provider openai -v nova "Hello from OpenAI"

# Save to file
gensay --provider openai -o speech.mp3 "OpenAI TTS output"

OpenAI offers two models via config.extra['model']:

  • tts-1 (default): Faster, lower latency
  • tts-1-hd: Higher quality audio

Amazon Polly

Option A - Environment variables:

  1. Sign in to AWS Console
  2. Go to IAMUsersCreate user
  3. Attach the AmazonPollyReadOnlyAccess policy
  4. Create access keys under Security credentialsAccess keys
  5. Configure credentials (choose one method):
export AWS_ACCESS_KEY_ID="AKIA..."
export AWS_SECRET_ACCESS_KEY="..."
export AWS_DEFAULT_REGION="us-west-2"

Option B - AWS CLI v2:

This easy lets you sign in through the AWS Command Line Interface

export AWS_DEFAULT_REGION=us-west-2
# on your desktop with a browser
aws login --region
# in an env without a browser
aws login --region --remote
# List Polly voices (60+ voices in many languages)
gensay --provider polly --list-voices

# Use a specific voice
gensay --provider polly -v Joanna "Hello from Amazon Polly"

# Save to file
gensay --provider polly -o speech.mp3 "Polly TTS output"

Polly supports multiple engines via config.extra['engine']:

  • neural (default): Higher quality, natural-sounding
  • standard: Lower cost, available for all voices

Deepgram

Deepgram's Flux TTS (/v2/speak) and Aura/Aura-2 (/v1/speak) batch REST APIs. The voice is embedded in the model string (e.g. flux-haley-en, aura-2-thalia-en) — pass either a short voice name or a full model string with -v.

Default model: Flux (flux-haley-en). Which model speaks is resolved in this order:

  1. -v <model string> — full model passthrough, e.g. -v flux-kit-en, -v aura-2-thalia-en
  2. -v <short name> — resolved from the voice catalog, newest family wins: Flux > Aura-2 > Aura (e.g. -v asteriaaura-2-asteria-en; use the full aura-asteria-en string for the legacy Aura voice)
  3. gensay config set deepgram.model <model string> — your per-user provider default
  4. Built-in fallback — flux-haley-en

Rate mapping (-r WPM, ~150 WPM = 1.0x): Flux accepts only the discrete ladder {0.85, 0.9, ..., 1.15} and gensay snaps to it (out-of-range values clamp); Aura accepts a continuous multiplier, clamped to 0.5–2.0.

  1. (Optional) install the extra — only needed to store the API key in the OS keychain via config set:
    pip install 'gensay[deepgram]'
    
  2. Get an API key from Deepgram Console
  3. Set the environment variable (or use the OS keychain, see below):
    export DEEPGRAM_API_KEY="your-api-key"
    
# List Deepgram voices (Flux + Aura-2 + Aura catalog)
gensay --provider deepgram --list-voices

# No voice flags → Flux default (flux-haley-en)
gensay --provider deepgram "Hello from Deepgram Flux"

# Use a short voice name or a full model string
gensay --provider deepgram -v kit "British Flux voice"
gensay --provider deepgram -v aura-2-thalia-en "Aura-2 voice"

# Save to file
gensay --provider deepgram -o speech.mp3 "Deepgram Flux TTS output"

Provider-specific config keys:

gensay config set deepgram.api_key '<your-key>'    # → Keychain/Secret Service
gensay config set deepgram.model aura-2-thalia-en  # override the Flux default

Advanced Features

Caching System

The caching system automatically stores generated audio to speed up repeated synthesis:

from gensay import TTSCache

# Create cache instance
cache = TTSCache(
    enabled=True,
    max_size_mb=10000,
    max_items=1000
)

# Get cache statistics
stats = cache.get_stats()
print(f"Cache size: {stats['size_mb']:.2f} MB")
print(f"Cached items: {stats['items']}")

# Clear cache
cache.clear()

Cache Location

Cache files are stored in platform-specific user cache directories:

  • macOS: ~/Library/Caches/gensay
  • Linux: ~/.cache/gensay
  • Windows: %LOCALAPPDATA%\gensay\gensay\Cache

Managing Cache

# Show cache statistics
gensay --cache-stats

# Clear all cached audio
gensay --clear-cache

# Disable caching for a specific command
gensay --no-cache "Text to synthesize without caching"

Manual Deletion

To manually delete the cache, remove the cache directory:

# macOS/Linux
rm -rf ~/Library/Caches/gensay  # macOS
rm -rf ~/.cache/gensay          # Linux

# Windows (PowerShell)
Remove-Item -Recurse -Force $env:LOCALAPPDATA\gensay\gensay\Cache

Creating Custom Providers

from gensay.providers import TTSProvider, TTSConfig, AudioFormat
from typing import Optional, Union, Any
from pathlib import Path

class MyCustomProvider(TTSProvider):
    def speak(self, text: str, voice: Optional[str] = None,
              rate: Optional[int] = None) -> None:
        # Your implementation
        self.update_progress(0.5, "Halfway done")
        # ... generate and play audio ...
        self.update_progress(1.0, "Complete")

    def save_to_file(self, text: str, output_path: Union[str, Path],
                     voice: Optional[str] = None, rate: Optional[int] = None,
                     format: Optional[AudioFormat] = None) -> Path:
        # Your implementation
        return Path(output_path)

    def list_voices(self) -> list[dict[str, Any]]:
        return [
            {'id': 'voice1', 'name': 'Voice One', 'language': 'en-US'}
        ]

    def get_supported_formats(self) -> list[AudioFormat]:
        return [AudioFormat.WAV, AudioFormat.MP3]

Async Support

All providers support async operations:

import asyncio
from gensay import ChatterboxProvider

async def main():
    provider = ChatterboxProvider()

    # Async speak
    await provider.speak_async("Async speech")

    # Async save
    await provider.save_to_file_async("Async save", "output.m4a")

asyncio.run(main())

Development

This project uses just for common development tasks. First, install just:

# macOS (using Nix which you already have)
nix-env -iA nixpkgs.just

# Or using Homebrew
brew install just

# Or using cargo
cargo install just

Getting Started

# Setup development environment
just setup

# Run tests
just test

# Run all quality checks
just check

# See all available commands
just

Common Development Commands

Testing

# Run all tests
just test

# Run tests with coverage
just test-cov

# Run specific test
just test-specific tests/test_providers.py::test_mock_provider_speak

# Quick test (mock provider only)
just quick-test

Python version support

just test runs the nox matrix across Python 3.11–3.15.

  • 3.15 is best-effort: verified against a pre-release interpreter (uv-managed 3.15.0a3); expect sharper edges until the final 3.15 release, and avoid claiming stable-grade 3.15 support to end users.
  • Pre-release Pythons need modern Rust: on interpreters without published wheels, deps build from source and Rust-backed sdists (e.g. jiter via OpenAI) require rustup stable ≥ 1.88 (rustup update stable).

Code Quality

# Run linter
just lint

# Auto-fix linting issues
just lint-fix

# Format code
just format

# Type checking
just typecheck

# Run all checks (lint, format, typecheck)
just check

# Pre-commit checks (format, lint, test)
just pre-commit

Running the CLI

# Run with mock provider
just run-mock "Hello, world!"
just run-mock -v '?'

# Run with macOS provider
just run-macos "Hello from macOS"

# Cache management
just cache-stats
just cache-clear

Development Utilities

# Run example script
just demo

# Clean build artifacts
just clean

# Build package
just build

Manual Setup (without just)

If you prefer not to use just, here are the equivalent commands:

# Setup
uv venv
uv pip install -e ".[dev]"

# Testing
uv run pytest -v
uv run pytest --cov=gensay --cov-report=term-missing

# Linting and formatting
uv run ruff check src tests
uv run ruff format src tests

# Type checking
uvx ty check src

Project Structure

gensay/
├── src/gensay/
│   ├── __init__.py
│   ├── main.py              # CLI entry point
│   ├── providers/           # TTS provider implementations
│   │   ├── base.py         # Abstract base provider
│   │   ├── chatterbox.py   # Chatterbox provider
│   │   ├── macos_say.py    # macOS say wrapper
│   │   └── ...            # Other providers
│   ├── cache.py            # Caching system
│   └── text_chunker.py     # Text chunking logic
├── tests/                  # Test suite
├── examples/               # Example scripts
├── justfile                # Development commands
└── README.md

Code Style Guide

  • Python 3.11+ with type hints
  • Follow PEP8 and Google Python Style Guide
  • Use ruff for linting and formatting
  • Keep docstrings concise but informative
  • Prefer pathlib.Path over os.path
  • Use pytest for testing

License

gensay is distributed under the terms of the MIT license.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

gensay-0.6.0.tar.gz (65.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

gensay-0.6.0-py3-none-any.whl (69.5 kB view details)

Uploaded Python 3

File details

Details for the file gensay-0.6.0.tar.gz.

File metadata

  • Download URL: gensay-0.6.0.tar.gz
  • Upload date:
  • Size: 65.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.12.0 {"installer":{"name":"uv","version":"0.12.0","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for gensay-0.6.0.tar.gz
Algorithm Hash digest
SHA256 f482a80e0b3ec52d5b97bb89cd4901174a75b3d5bf524d0477a403cb85cb6b2d
MD5 12e38e2ae9655091a8f0f39fa1a11302
BLAKE2b-256 f502bd9c4086b5e3183edd423089691b2ccb79436f6f072fa36f6315594875fc

See more details on using hashes here.

File details

Details for the file gensay-0.6.0-py3-none-any.whl.

File metadata

  • Download URL: gensay-0.6.0-py3-none-any.whl
  • Upload date:
  • Size: 69.5 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.12.0 {"installer":{"name":"uv","version":"0.12.0","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for gensay-0.6.0-py3-none-any.whl
Algorithm Hash digest
SHA256 b5ef4cd966d029b2330674a752178692868948dd66cee9f74ac32d65e85ada6f
MD5 871886336eb24ef952071bf2dbe4cb16
BLAKE2b-256 5550d3096876a6a293416be4b7c52c479b679d3a9f1d56f3a8649ca6b1b4def0

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page