Skip to main content

KugelAudio Python SDK

Official Python SDK for the KugelAudio Text-to-Speech API.

Installation

pip install kugelaudio

Or with uv:

uv add kugelaudio

Quick Start

from kugelaudio import KugelAudio

# Initialize the client - just needs an API key!
client = KugelAudio(api_key="your_api_key")

# Generate speech
audio = client.tts.generate(
    text="Hello, world!",
    model_id="kugel-1-turbo",
)

# Save to file
audio.save("output.wav")

Local CPU Turn Detection

The optional turn-detection runtime downloads the private, version-pinned ONNX bundle from Hugging Face, verifies every declared SHA-256 checksum, and then runs fully locally on CPU. ONNX Runtime performs model inference; the extra also loads the CPU PyTorch runtime to reproduce the native numerical environment used for the published quality gates. It requires Python 3.11 or newer.

pip install "kugelaudio[turn-detection]"
hf auth login  # required while the model repository is private

Load one model per process and create cheap state for each conversation:

from kugelaudio.turn import TurnDecisionReason, TurnDetector

detector = TurnDetector.from_pretrained(cpu_threads=4)
turn = detector.create_session("en")

# Feed the current user's audio and interim ASR transcript continuously.
turn.push_pcm16(pcm_chunk, sample_rate=16_000)
turn.update_transcript("I think we should probably")

# Call with the current duration whenever VAD reports real silence.
decision = turn.observe_silence(duration_ms=200)

if decision.reason is TurnDecisionReason.MODEL_ACTION_DELAY:
    # The model predicts completion, but the measured policy is still waiting.
    decision = turn.observe_silence(duration_ms=600)

if decision.end_turn:
    start_assistant_response()
    turn.reset_turn()

For a latency-first application, commit a confident completion immediately when the first 200 ms score returns while retaining the calibrated threshold and incomplete-turn timeout:

turn = detector.create_session("en", action_delay_ms=200, timeout_ms=10_000)

The incomplete-turn fallback defaults to ten seconds. Pass a different timeout_ms only when the surrounding product has a stricter wait budget.

When VAD detects speech before the endpoint action, cancel the pending decision:

turn.speech_resumed()

Important input and lifecycle rules:

  • Audio must be mono 16 kHz. push_audio accepts normalized float32; use push_pcm16 for explicit little-endian signed PCM16 conversion.
  • The transcript must be the current interim user transcript. Do not wait for a post-endpoint final transcript.
  • Each silence episode is scored once from an audio/transcript snapshot. Later interim or final transcript revisions do not discard or repeat that inference; resumed speech opens a new episode with the latest transcript.
  • Call observe_silence with monotonically increasing durations for one silence span. Call speech_resumed before starting a new span.
  • Only P(complete) can trigger a model endpoint. incomplete, backchannel, and wait keep listening until speech resumes or the measured timeout fires.
  • One TurnDetector owns the roughly 1.55 GiB model runtime. Share it across sessions instead of loading one copy per conversation.
  • No network is used after the immutable model revision is cached. Pass local_files_only=True to enforce offline startup.

The default policies cover de, en, es, it, and nl. Unsupported languages, missing private-repository access, corrupt bundles, wrong sample rates, and invalid session ordering raise typed TurnDetectionError subclasses instead of silently falling back.

LiveKit Agents

Install both optional integrations and LiveKit's VAD plugin:

pip install "kugelaudio[livekit,turn-detection]" "livekit-agents[silero]"

KugelTurnBridge keeps LiveKit's normal STT pipeline intact while teeing its audio into KugelTurn. Configure LiveKit for manual endpointing so two detectors cannot commit the same user turn:

import logging
from collections.abc import AsyncIterable

from livekit import rtc
from livekit.agents import Agent, AgentSession, ModelSettings
from livekit.plugins import silero

from kugelaudio.livekit import KugelTurnBridge
from kugelaudio.turn import TurnDetector

logger = logging.getLogger(__name__)
detector = TurnDetector.from_pretrained(cpu_threads=4)  # once per worker


def observe_decision(decision):
    logger.info(
        "turn decision",
        extra={
            "reason": decision.reason.value,
            "end_turn": decision.end_turn,
            "silence_ms": decision.silence_ms,
            "inference_ms": decision.inference_ms,
            "probabilities": decision.probabilities,
        },
    )


bridge = KugelTurnBridge(
    detector,
    language="en",
    vad_silence_ms=200,
    action_delay_ms=200,  # omit to use the calibrated language delay
    timeout_ms=10_000,
    on_decision=observe_decision,
)


class VoiceAgent(Agent):
    def stt_node(
        self,
        audio: AsyncIterable[rtc.AudioFrame],
        model_settings: ModelSettings,
    ):
        return bridge.stt_node(self, audio, model_settings)


session = AgentSession(
    turn_detection="manual",
    vad=silero.VAD.load(min_silence_duration=0.2),
    stt=stt,
    llm=llm,
    tts=tts,
)

try:
    await session.start(agent=VoiceAgent(instructions="..."), room=ctx.room)
finally:
    await bridge.aclose()

On LiveKit Agents 1.5+, the equivalent non-deprecated configuration is turn_handling=TurnHandlingOptions(turn_detection="manual"). The bridge resamples mono LiveKit input to 16 kHz, accumulates interim/final transcripts, runs ONNX inference outside the event loop, cancels pending decisions when speech resumes, and calls commit_user_turn() when the measured policy ends.

Keep vad_silence_ms equal to LiveKit's min_silence_duration in milliseconds. The validated configuration is 200 ms.

on_decision runs for the first model score and each meaningful policy-state change, including the final model_complete or timeout. Its TurnDecision contains the transcript, threshold, four class probabilities, silence duration, and inference_ms when a new model score was computed. Keep the callback non-blocking; enqueue network or storage work instead of performing it inline.

Pipecat

Install the Pipecat and turn-detection extras:

pip install "kugelaudio[pipecat,turn-detection]"

Use KugelTurnStopStrategy as the sole Pipecat user-turn stop strategy:

from pipecat.audio.vad.silero import SileroVADAnalyzer
from pipecat.audio.vad.vad_analyzer import VADParams
from pipecat.processors.aggregators.llm_response_universal import (
    LLMContextAggregatorPair,
    LLMUserAggregatorParams,
)
from pipecat.turns.user_turn_strategies import UserTurnStrategies

from kugelaudio.pipecat import KugelTurnStopStrategy
from kugelaudio.turn import TurnDetector

detector = TurnDetector.from_pretrained(cpu_threads=4)  # once per process
turn_strategy = KugelTurnStopStrategy(
    detector,
    language="en",
    vad_silence_ms=200,
    action_delay_ms=200,  # omit to use the calibrated language delay
    timeout_ms=10_000,
)

user_aggregator, assistant_aggregator = LLMContextAggregatorPair(
    context,
    user_params=LLMUserAggregatorParams(
        vad_analyzer=SileroVADAnalyzer(
            params=VADParams(stop_secs=0.2),
        ),
        user_turn_strategies=UserTurnStrategies(
            stop=[turn_strategy],
        ),
    ),
)

The strategy supports Pipecat 0.0.101+ and 1.x turn-management APIs. It consumes native InputAudioRawFrame, transcription, and VAD frames, explicitly rejects multichannel input, resamples mono PCM to 16 kHz, and emits Pipecat's standard on_user_turn_stopped event. For Pipecat 0.x releases whose VAD stop frame does not carry its own duration, keep vad_silence_ms equal to VADParams.stop_secs.

Client Configuration

from kugelaudio import KugelAudio

# Simple setup - single URL handles everything
client = KugelAudio(api_key="your_api_key")

# Or with custom options
client = KugelAudio(
    api_key="your_api_key",           # Required: Your API key
    api_url="https://api.kugelaudio.com",  # Optional: API base URL (default)
    timeout=60.0,                      # Optional: Request timeout in seconds
)

Region Selection

By default, KugelAudio uses the canonical geo-routed API endpoint. You can select the direct EU endpoint when you need to pin traffic to Europe.

Region hint Endpoint
default api.kugelaudio.com (geo-routed)
eu api.eu.kugelaudio.com

Option 1 — API key prefix (simplest, works with env vars):

client = KugelAudio(api_key="eu-ka_your_api_key")       # → EU
client = KugelAudio(api_key="ka_your_api_key")          # → canonical geo-routed API

Option 2 — region parameter:

client = KugelAudio(api_key="ka_your_api_key", region="eu")

The prefix is always stripped before authentication. Priority: api_url > region > key prefix > default.

Single URL Architecture

The SDK uses a single URL for both REST API and WebSocket streaming. The TTS server provides both REST endpoints (/v1/models, /v1/voices) and WebSocket (/ws/tts) - no proxy needed, minimal latency.

Local Development

For local development, point directly to your TTS server:

client = KugelAudio(
    api_key="your_api_key",
    api_url="http://localhost:8000",   # TTS server handles everything
)

Or if you have separate backend and TTS servers:

client = KugelAudio(
    api_key="your_api_key",
    api_url="http://localhost:8001",   # Backend for REST API
    tts_url="http://localhost:8000",   # TTS server for WebSocket streaming
)

Available Models

Model ID Name Description
kugel-1-turbo Kugel 1 Turbo Fast, low-latency model for real-time applications
kugel-1 Kugel 1 Premium quality model for pre-recorded content

List Available Models

models = client.models.list()

for model in models:
    print(f"{model.id}: {model.name}")
    print(f"  Description: {model.description}")
    print(f"  Max Input: {model.max_input_length} characters")
    print(f"  Sample Rate: {model.sample_rate} Hz")

Voices

List Available Voices

# List all available voices (paginated)
result = client.voices.list()

for voice in result.voices:
    print(f"{voice.id}: {voice.name}")
    print(f"  Category: {voice.category}")
    print(f"  Languages: {', '.join(voice.supported_languages)}")
print(f"Showing {len(result.voices)} of {result.total} voices")

# Filter by language
result = client.voices.list(language="de")

# Get only public voices
result = client.voices.list(include_public=True)

# Paginate through results
page1 = client.voices.list(limit=10, offset=0)
page2 = client.voices.list(limit=10, offset=10)

Get a Specific Voice

voice = client.voices.get(voice_id=123)
print(f"Voice: {voice.name}")
print(f"Sample text: {voice.sample_text}")

Text-to-Speech Generation

Basic Generation (Non-Streaming)

Generate complete audio and receive it all at once:

audio = client.tts.generate(
    text="Hello, this is a test of the KugelAudio text-to-speech system.",
    model_id="kugel-1-turbo",  # 'kugel-1-turbo' (fast) or 'kugel-1' (quality)
    voice_id=123,              # Optional: specific voice ID
    cfg_scale=2.0,             # Guidance scale (1.2-2.5)
    max_new_tokens=2048,       # Maximum tokens to generate
    sample_rate=24000,         # Output sample rate
    output_format=None,        # Optional: 'pcm_24000', 'ulaw_8000', 'alaw_8000', ...
    normalize=True,            # Enable text normalization (see below)
    language="en",             # Language for normalization
)

# Audio properties
print(f"Duration: {audio.duration_seconds:.2f}s")
print(f"Samples: {audio.samples}")
print(f"Sample rate: {audio.sample_rate} Hz")
print(f"Generation time: {audio.generation_ms:.0f}ms")
print(f"RTF: {audio.rtf:.2f}")  # Real-time factor

# Save to WAV file
audio.save("output.wav")

# Get raw PCM bytes
pcm_data = audio.audio

# Get WAV bytes (with header)
wav_bytes = audio.to_wav_bytes()

Streaming Audio Output

Receive audio chunks as they are generated for lower latency:

# Synchronous streaming
for item in client.tts.stream(
    text="Hello, this is streaming audio.",
    model_id="kugel-1-turbo",
):
    if hasattr(item, 'audio'):  # AudioChunk
        # Process audio chunk immediately
        print(f"Chunk {item.index}: {len(item.audio)} bytes, {item.samples} samples")
        # play_audio(item.audio)
    elif isinstance(item, dict) and item.get('final'):
        # Final stats
        print(f"Total duration: {item.get('dur_ms', 0):.0f}ms")
        print(f"Generation time: {item.get('gen_ms', 0):.0f}ms")

Async Streaming

For async applications:

import asyncio

async def generate_speech():
    async for item in client.tts.stream_async(
        text="Async streaming example.",
        model_id="kugel-1-turbo",
    ):
        if hasattr(item, 'audio'):
            # Process chunk
            pass

asyncio.run(generate_speech())

Async Generation

import asyncio

async def main():
    audio = await client.tts.generate_async(
        text="Async generation example.",
        model_id="kugel-1-turbo",
    )
    audio.save("async_output.wav")

asyncio.run(main())

Text Normalization

Text normalization converts numbers, dates, times, and other non-verbal text into spoken words. For example:

  • "I have 3 apples" → "I have three apples"
  • "The meeting is at 2:30 PM" → "The meeting is at two thirty PM"
  • "€50.99" → "fifty euros and ninety-nine cents"

Usage

# With explicit language (recommended - fastest)
audio = client.tts.generate(
    text="I bought 3 items for €50.99 on 01/15/2024.",
    normalize=True,
    language="en",  # Specify language for best performance
)

# With auto-detection (adds ~150ms latency)
audio = client.tts.generate(
    text="Ich habe 3 Artikel für 50,99€ gekauft.",
    normalize=True,
    # language not specified - will auto-detect
)

Supported Languages

Code Language Code Language
de German nl Dutch
en English pl Polish
fr French sv Swedish
es Spanish da Danish
it Italian no Norwegian
pt Portuguese fi Finnish
cs Czech hu Hungarian
ro Romanian el Greek
uk Ukrainian bg Bulgarian
tr Turkish vi Vietnamese
ar Arabic hi Hindi
zh Chinese ja Japanese
ko Korean

Performance Warning

⚠️ Latency Warning: Using normalize=True without specifying language adds approximately 150ms latency for language auto-detection. For best performance in latency-sensitive applications, always specify the language parameter.

LLM Integration: Streaming Text Input

For real-time TTS when streaming text from an LLM (GPT-4, Claude, etc.), use a StreamingSession. Forward LLM tokens directly to session.send() without flush=True — the server accumulates them and starts generation at natural sentence boundaries. Flush exactly once at the end of the assistant turn.

⚠️ Do not call session.send(text, flush=True) between sentences or words. Each explicit flush is a separate TTS request that pays the full model time-to-first-audio (TTFA) again and produces an audible gap. See Chunking & per-segment latency for the full rationale and ElevenLabs migration notes.

Async Streaming Session

import asyncio

async def speak_turn(llm_token_stream):
    async with client.tts.streaming_session(
        voice_id=123,
        model_id="kugel-1-turbo",
        language="en",
    ) as session:
        # Forward every LLM token directly. No flush=True per token —
        # the server's text buffer chunks at sentence boundaries.
        async for token in llm_token_stream:
            async for chunk in session.send(token):
                play_audio(chunk.audio)

        # Single flush at turn end emits any trailing text.
        async for chunk in session.flush():
            play_audio(chunk.audio)

        # Per-session usage — bill your own customers per conversation.
        # cost_cents is the actual charge in EUR cents (None if undetermined).
        usage = session.last_usage
        if usage:
            print(f"audio: {usage.audio_seconds}s, cost: {usage.cost_cents} ct")

asyncio.run(speak_turn(my_llm_stream()))

Synchronous Streaming Session

with client.tts.streaming_session_sync(
    voice_id=123,
    model_id="kugel-1-turbo",
    language="en",
) as session:
    for token in llm_token_stream:
        for chunk in session.send(token):  # no flush per token
            play_audio(chunk.audio)

    for chunk in session.flush():  # single flush at turn end
        play_audio(chunk.audio)

Error Handling

from kugelaudio import KugelAudio
from kugelaudio.exceptions import (
    KugelAudioError,
    AuthenticationError,
    RateLimitError,
    InsufficientCreditsError,
    ValidationError,
    NotFoundError,
)

try:
    audio = client.tts.generate(text="Hello!")
except AuthenticationError:
    print("Invalid API key")
except RateLimitError:
    print("Rate limit exceeded, please wait")
except InsufficientCreditsError:
    print("Not enough credits, please top up")
except ValidationError as e:
    print(f"Invalid request: {e}")
except NotFoundError as e:
    print(f"Resource not found (e.g. unknown voice_id): {e}")
except KugelAudioError as e:
    print(f"API error: {e}")

Data Models

AudioChunk

Represents a single audio chunk from streaming:

class AudioChunk:
    audio: bytes          # Raw PCM16 audio data
    encoding: str         # 'pcm_s16le' | 'mulaw' | 'alaw' (G.711 when output_format set)
    index: int           # Chunk index (0-based)
    sample_rate: int     # Sample rate (24000)
    samples: int         # Number of samples in chunk
    
    @property
    def duration_seconds(self) -> float:
        """Duration of this chunk in seconds."""

AudioResponse

Complete audio response from generation:

class AudioResponse:
    audio: bytes              # Complete PCM16 audio
    sample_rate: int          # Sample rate (24000)
    samples: int              # Total samples
    duration_ms: float        # Duration in milliseconds
    generation_ms: float      # Generation time in milliseconds
    rtf: float               # Real-time factor
    
    @property
    def duration_seconds(self) -> float:
        """Duration in seconds."""
    
    def save(self, path: str) -> None:
        """Save as WAV file."""
    
    def to_wav_bytes(self) -> bytes:
        """Get WAV file as bytes."""

Model

TTS model information:

class Model:
    id: str                   # 'kugel-1-turbo' or 'kugel-1'
    name: str                 # Human-readable name
    description: str          # Model description
    max_input_length: int     # Maximum input characters
    sample_rate: int          # Output sample rate

Supported native output format tokens are pcm_8000, pcm_16000, pcm_22050, pcm_24000, ulaw_8000, and alaw_8000.

Voice

Voice information:

class Voice:
    id: int                          # Voice ID
    name: str                        # Voice name
    description: Optional[str]       # Description
    category: Optional[VoiceCategory]  # 'premade', 'cloned', 'generated'
    sex: Optional[VoiceSex]          # 'male', 'female', 'neutral'
    age: Optional[VoiceAge]          # 'young', 'middle_aged', 'old'
    supported_languages: List[str]   # ['en', 'de', ...]
    sample_text: Optional[str]       # Sample text for preview
    avatar_url: Optional[str]        # Avatar image URL
    sample_url: Optional[str]        # Sample audio URL
    is_public: bool                  # Whether voice is public
    verified: bool                   # Whether voice is verified

Complete Example

from kugelaudio import KugelAudio

# Initialize client
client = KugelAudio(api_key="your_api_key")

# List available models
print("Available Models:")
for model in client.models.list():
    print(f"  - {model.id}: {model.name}")

# List available voices
print("\nAvailable Voices:")
for voice in client.voices.list(limit=5).voices:
    print(f"  - {voice.id}: {voice.name}")

# Generate audio
print("\nGenerating audio...")
audio = client.tts.generate(
    text="Welcome to KugelAudio. This is an example of high-quality text-to-speech synthesis.",
    model_id="kugel-1-turbo",
)

print(f"Generated {audio.duration_seconds:.2f}s of audio in {audio.generation_ms:.0f}ms")
print(f"Real-time factor: {audio.rtf:.2f}x")

# Save to file
audio.save("example.wav")
print("Saved to example.wav")

# Close client
client.close()

License

MIT

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

kugelaudio-1.7.0.tar.gz (95.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

kugelaudio-1.7.0-py3-none-any.whl (92.3 kB view details)

Uploaded Python 3

File details

Details for the file kugelaudio-1.7.0.tar.gz.

File metadata

  • Download URL: kugelaudio-1.7.0.tar.gz
  • Upload date:
  • Size: 95.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.11.30 {"installer":{"name":"uv","version":"0.11.30","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

File hashes

Hashes for kugelaudio-1.7.0.tar.gz
Algorithm Hash digest
SHA256 a549dc9275b41380768359ee62138ce6fc86aefad8f0b6d01576e75e7c248c35
MD5 707d32e0c522d3dad32caffba8e40743
BLAKE2b-256 acc5fe27b5f60384514ecba298f6d715a030bac0034cc9b5b10cb928d9691a95

See more details on using hashes here.

File details

Details for the file kugelaudio-1.7.0-py3-none-any.whl.

File metadata

  • Download URL: kugelaudio-1.7.0-py3-none-any.whl
  • Upload date:
  • Size: 92.3 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.11.30 {"installer":{"name":"uv","version":"0.11.30","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

File hashes

Hashes for kugelaudio-1.7.0-py3-none-any.whl
Algorithm Hash digest
SHA256 21c00428178786faff885a3fdb618f3875ef54dfffa1ea596351c50c289a6c78
MD5 30f852d748c97c316353cacb94df8c15
BLAKE2b-256 9e98ac24de52b8416e04bc91a6723454847311b4def7bebbc4caac6d05e0fae6

See more details on using hashes here.

Release history Release notifications | RSS feed

1.11.0

2 files

1.10.0

2 files

1.9.0

2 files

1.8.0

2 files

This release

1.7.0 This release

2 files

1.6.0

2 files

1.5.0

2 files

1.4.4

2 files

1.4.3

2 files

1.4.2

2 files

1.4.1

2 files

1.4.0

2 files

1.3.2

2 files

1.3.1

2 files

1.3.0

2 files

1.2.3

2 files

1.2.2

2 files

1.2.1

2 files

1.2.0

2 files

1.1.2

2 files

1.1.1

2 files

1.1.0

2 files

1.0.0

2 files

0.5.0

2 files

0.4.0

2 files

0.3.1

2 files

0.3.0

2 files

0.2.3

2 files

0.2.2

2 files

0.2.1

2 files

0.2.0

2 files

0.1.15

2 files

0.1.14

2 files

0.1.13

2 files

0.1.12

2 files

0.1.11

2 files

0.1.9

2 files

0.1.7

2 files

0.1.5

2 files

0.1.2

2 files

0.1.1

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page