Skip to main content

piper-tts-server

Paced, streaming Piper TTS for real-time voice pipelines — Asterisk AudioSocket, RTP, or a browser WebSocket. Optional ElevenLabs cloud voices with automatic local fallback.

CI PyPI License: MIT

If you have tried to bolt Piper onto a phone call and hit dead air before the greeting, the first few words clipped, or the caller only hearing the tail of each sentence — this library is the fix. It packages the three things that make Piper work in a real-time loop:

  1. Process-wide voice cache. PiperVoice.load() costs 2.5–5.5 s. Loading per request means that much dead air before every utterance. Voices load once and are reused.
  2. A single synthesis lock. Piper phonemizes via espeak-ng, whose C API is not thread-safe. All synthesis is serialized behind one lock so concurrent calls don't corrupt each other.
  3. Deadline-based frame pacing. Real-time transports forward each frame the instant it arrives; bursting a whole utterance overruns the far end's jitter buffer. PacedWriter releases exactly one frame per interval on a monotonic deadline — and re-clamps every frame so a synthesis stall never turns into a catch-up burst that clips the next words.

Install

pip install "piper-tts-server[all]"     # library + server + piper + elevenlabs
# or minimal:
pip install piper-tts-server            # library core (numpy only)
pip install "piper-tts-server[piper,server]"

Download a Piper voice (.onnx + .onnx.json) into a voices/ directory — see the Piper voices list.

Run the service

PIPER_TTS_VOICES_DIR=voices PIPER_TTS_DEFAULT_VOICE=en_US-amy-medium \
  piper-tts-server            # listens on 127.0.0.1:8080
# one-shot WAV
curl -X POST localhost:8080/synthesize \
  -H 'Content-Type: application/json' \
  -d '{"text":"Hello from Piper."}' --output hello.wav

# paced streaming slin16 frames (feed straight to AudioSocket/RTP)
curl -N -X POST localhost:8080/synthesize/stream \
  -H 'Content-Type: application/json' \
  -d '{"text":"Your call is being connected.","paced":true}' --output frames.slin
Endpoint Purpose
GET /health status + cached voices
POST /warm {voice_id} preload a voice into the cache
POST /synthesize {text, voice_id, provider, format} one-shot wav (default) or raw pcm
POST /synthesize/stream {text, …, paced} chunked frames; paced:true = deadline-clocked

Use as a library (no HTTP in the media path)

from piper_tts_server import StreamingTTS, PacedWriter, tts_sanitize

tts = StreamingTTS(voices_dir="voices", default_voice_path="voices/en_US-amy-medium.onnx")
await tts.start()

writer = PacedWriter(my_audiosocket_write, frame_ms=20)   # sink: async fn(bytes)
async for frame in tts.synthesise(tts_sanitize(text)):
    await writer.write(frame)                              # 320 bytes / 20 ms, paced

See examples/asterisk_audiosocket.md for the full AudioSocket wiring and streaming-LLM sentence chunking.

Configuration

All via env (or ServerConfig): PIPER_TTS_HOST, PIPER_TTS_PORT, PIPER_TTS_VOICES_DIR, PIPER_TTS_DEFAULT_VOICE, PIPER_TTS_PROVIDER (piper|elevenlabs), PIPER_TTS_ELEVENLABS_API_KEY, PIPER_TTS_SAMPLE_RATE (default 8000), PIPER_TTS_FRAME_MS (default 20), PIPER_TTS_WARM_ON_START.

Default output is 8 kHz slin16 / 20 ms frames (telephony). Set PIPER_TTS_SAMPLE_RATE=16000 for wideband.

ElevenLabs with local fallback

Set PIPER_TTS_PROVIDER=elevenlabs + PIPER_TTS_ELEVENLABS_API_KEY and pass a voice_id. Any failure falls back to the local Piper default voice so the stream never goes silent; a 401/403 disables further cloud attempts for that engine.

Development

pip install "piper-tts-server[dev]"
pytest          # pure-DSP, sanitize, and pacing tests run without Piper installed

Licensing

This package is MIT licensed and contains no Piper source code. Piper is an optional extra, imported lazily only when the local engine is used — the core library installs with numpy alone, and the test suite runs without Piper present.

Piper itself is a separate project under a different licence. The maintained release (OHF-Voice/piper1-gpl, PyPI piper-tts) is GPL-3.0-or-later; the older MIT rhasspy/piper is archived. If you install the [piper] or [all] extra, the resulting combined installation is subject to GPL-3.0 terms. The ElevenLabs provider, and the pacing and sanitising helpers on their own, pull in no GPL code.

Not affiliated with, or endorsed by, the Open Home Foundation or the Piper project. This is an independent server and library that drives Piper as one of several backends.

Provenance & credits

Extracted and generalized from the TTS layer of ICTContact's AI Voice Agent, by ICT Innovations / ICT Vision. Author: Tahir Almas. Companion project: asterisk-ai-voice-agent.

MIT licensed — see LICENSE.

Metadata

Release files for piper-tts-server 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for piper-tts-server 0.1.0
File Size Uploaded
piper_tts_server-0.1.0.tar.gz 17.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for piper-tts-server 0.1.0
File Interpreter ABI Platform
piper_tts_server-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 32.8 kB

Release files / piper_tts_server-0.1.0.tar.gz

Download URL piper_tts_server-0.1.0.tar.gz
Size 17.1 kB
Tags Source
SHA-256 checksum
How to use checksums
7ae9b0413d9bda8697883aeb13c3fd5ad92f62184a39f6ebc3c7bd63d8839783
BLAKE2b-256 checksum
How to use checksums
946a02debb0dcc9187ee43edf58a190cb41f4d6981aa197b58b396cf7ec7c5df
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 13, 2026.

Transparency log

Release files / piper_tts_server-0.1.0-py3-none-any.whl

Download URL piper_tts_server-0.1.0-py3-none-any.whl
Size 15.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
16b08ab7b1f86185ef568f52434acc5a96d4d0ee0cfba4f08e916ecb4bfc2c67
BLAKE2b-256 checksum
How to use checksums
d39877b98dd3ee1f72aceb8ac957b030f541890f320c6f5998e5d6b6a739f67b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 13, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page