svara-voice
Python SDK for Svara, Kenpath Labs' text-to-speech API: 82 languages with automatic code-switching, 320 voices, streaming over HTTP and WebSocket, and G.711 µ-law/A-law at 8 kHz for telephony.
- Docs: https://docs.kenpathlabs.com · package reference: api-reference.md · changelog
- API base:
https://api.kenpathlabs.com· keys: https://platform.kenpathlabs.com - Requires Python 3.9+. Depends on
httpxandwebsocketsonly.
pip install svara-voice
The distribution is svara-voice; the import is svara.
Quickstart
from svara import Svara
client = Svara(api_key="sk_live_...") # or set SVARA_API_KEY; keys come from the console
audio = client.speech.create(
input="नमस्ते! Welcome to Svara.", # any language, mixed scripts are fine
voice="sv_enhdbrj5", # any id from client.voices.list()
response_format="mp3",
)
audio.save("hello.mp3") # it is bytes, with headers attached
Hear it instead (needs ffplay, which ships with FFmpeg):
from svara import play
play(client.speech.stream(input="नमस्ते!", voice="sv_enhdbrj5")) # starts at the first chunk
Stream while it generates
for chunk in client.speech.stream(input="...", voice="sv_enhdbrj5"):
player.write(chunk) # pcm by default: 24 kHz, 16-bit, mono; first audio ≈ 200 ms
Speak an LLM's tokens as they arrive
The lowest-latency path for voice agents: first audio 427 ms after the call on
a fresh socket, 132 ms on a prepared one. Feed the token stream in; Svara starts speaking eight words in by default (chunk_words=4)
and keeps prosody continuous across the whole reply.
import asyncio
from svara import AsyncSvara
async def speak(llm_token_stream):
async with AsyncSvara() as client:
async for audio in client.speech.stream_input(llm_token_stream, voice="sv_enhdbrj5"):
player.write(audio)
asyncio.run(speak(my_llm_tokens()))
Open the socket before the text exists and first audio lands ~300 ms sooner:
async def turn(client: AsyncSvara, get_llm_tokens):
prepared = await client.speech.prepare(voice="sv_enhdbrj5") # call while the user is still talking
tokens = await get_llm_tokens() # the LLM request goes out here
async for audio in prepared.stream(tokens):
player.write(audio)
There is a blocking twin, Svara().speech.stream_input(...), for code without an
event loop. Clients own a connection pool: use them as context managers, or
call close() / aclose() when done.
Telephony
ulaw = client.speech.create(input="...", voice="sv_enhdbrj5", response_format="ulaw", sample_rate=8000)
Always pass sample_rate=8000 for a phone leg: the API renders every format at
24 kHz unless told otherwise, G.711 included, and 24 kHz µ-law on an 8 kHz leg
plays at three times speed. The SDK warns when the rate is missing.
Timestamps
r = client.speech.create_with_timestamps(input="...", voice="sv_enhdbrj5")
r.audio, r.alignment.words() # [(word, start_s, end_s), ...] for subtitles or karaoke
Voices, languages, usage
from svara import PronunciationRule
client.voices.list(language="hi", gender="female") # filtered client-side
client.voices.search("tamil male") # any words from name, accent, language, labels
client.voices.preview("sv_enhdbrj5") # a sample clip, audio/mpeg
client.languages.list() # 82 languages and the codes `language=` accepts
client.usage.get().characters_remaining # plan, month-to-date, balance
client.pronunciation_dictionaries.create_from_rules(name="brand", rules=[PronunciationRule("SQL", "sequel")])
Errors
from svara import SvaraError, RateLimitError, QuotaExceededError
try:
client.speech.create(input="...", voice="sv_enhdbrj5")
except QuotaExceededError: # 429 insufficient_quota — terminal until the month resets
...
except RateLimitError as e: # 429 — already retried with backoff; e.retry_after in seconds
...
except SvaraError as e:
print(e.status_code, e.code, e.message)
HTTP calls are retried twice on connection errors, 429 and 5xx, with jittered
backoff and Retry-After honoured; stream() only until its first byte, so a
retry never replays audio the caller is already playing. The WebSocket paths
(stream_input, prepare) are never retried: a refused connection raises at
once.
Voice-agent frameworks
LiveKit Agents — pip install "svara-voice[livekit]"
from svara.livekit import TTS
session = AgentSession(tts=TTS(voice="sv_enhdbrj5"), stt=..., llm=...)
Pipecat — pip install "svara-voice[pipecat]"
from svara.pipecat import SvaraTTSService
pipeline = Pipeline([transport.input(), stt, llm, SvaraTTSService(voice="sv_enhdbrj5"), transport.output()])
The LiveKit plugin streams 24 kHz PCM over the input-streaming WebSocket and keeps a socket prewarmed between turns; any telephony downsampling is left to LiveKit. The Pipecat service asks the server for the transport's own rate, so nothing is resampled in the pipeline.
Using the OpenAI or ElevenLabs SDKs instead
Svara is request-compatible with both. Point base_url at Svara and they work
unmodified, including ElevenLabs' realtime WebSocket client. The other
direction is as short: code written for the OpenAI SDK
(client.audio.speech.create(...), .write_to_file(),
with_streaming_response) runs on a Svara client unchanged, provided it does
not pass instructions= or stream_format= (Svara has no equivalent; they
raise TypeError).
See compatibility.md for the base URLs and a field-by-field mapping. Only this SDK exposes the native input-streaming socket: first audio arrives 0.6 s sooner than over the ElevenLabs realtime protocol on a fresh connection (427 ms vs 1047 ms), 0.9 s sooner on a prepared one (132 ms).
CLI
svara say "नमस्ते दुनिया" --voice sv_enhdbrj5 --out hello.mp3
svara say "Your call is important" -v sv_enhdbrj5 -f ulaw -r 8000 -o prompt.ulaw
svara voices --language hi
svara languages
svara usage
svara doctor # connectivity + key check with timings
Formats
mp3 · opus · aac · flac · wav · pcm (16-bit LE mono) · ulaw / alaw (G.711).
sample_rate ∈ 8000, 16000, 22050, 24000 (default), 32000, 44100, 48000.
bitrate_kbps for the lossy three. ElevenLabs-style names translate with
output_format("mp3_44100_128").
Documentation
docs/ covers installation, streaming and latency (with the measurements behind the defaults), voices, the full API reference, compatibility with other SDKs, troubleshooting, and deployment guides for local, Docker, cloud, LiveKit + SIP and raw WebSocket telephony. MEASUREMENTS.md is the lab notebook. docs/llms.txt is the same material condensed for coding assistants. Release notes: CHANGELOG.md.
Svara may be shared across threads (httpx's pool is thread-safe); an
AsyncSvara belongs to one event loop; a stream object has one consumer.
License
Proprietary © Kenpath Labs. See LICENSE.
Metadata
Release files for svara-voice 0.2.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| svara_voice-0.2.1.tar.gz | 106.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| svara_voice-0.2.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 161.8 kB
Release files / svara_voice-0.2.1.tar.gz
| Download URL | svara_voice-0.2.1.tar.gz |
|---|---|
| Size | 106.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
860d8c2634a08248b41e4cc850a423cf9c29eb98f7e4425d56ff9942c9171281
|
|
BLAKE2b-256 checksum How to use checksums |
13c2f68ff03c34d36c406e782e900af27e8315ebe4751d5526fc4c604e8c2382
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.14
|
Release files / svara_voice-0.2.1-py3-none-any.whl
| Download URL | svara_voice-0.2.1-py3-none-any.whl |
|---|---|
| Size | 55.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
4140a034b5146fa9ed7928375c1b99fae17f60ea47164e992cf4a3ae61abea37
|
|
BLAKE2b-256 checksum How to use checksums |
a55ffe8cca03ec4301b95d826dcd0ba8384f18cc81d81c6bad04be3ae64df4be
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.14
|