Skip to main content

VoiceKit — Python SDK

Official Python wrapper for VoiceKit: neural speech synthesis, transcription, sentiment analysis, and batch operations.

Install

pip install voicekit-client

Quick start

from voicekit import VoiceKitClient, b64

client = VoiceKitClient(api_key="YOUR_KEY")

# Synthesis → raw audio bytes
audio = client.synthesize("Привет! Это синтез русской речи.", voice="preset_anna", format="mp3")
with open("speech.mp3", "wb") as f:
    f.write(audio)

# Streaming (Pro/Business)
for chunk in client.synthesize_stream("Первое предложение. Второе."):
    pass  # write chunks to a file or socket

# Transcription (async → poll)
job = client.transcribe("audio.wav", keyterms=["диагноз"])
result = client.get_transcription_job(job["job_id"])
while result["status"] not in ("completed", "failed"):
    result = client.get_transcription_job(job["job_id"])

# Short-file sync transcription
transcript = client.transcribe_sync("audio.wav")

# Analysis (sentiment + keywords + entities)
analysis = client.analyze_sync("audio.wav")

# Text intelligence
lang = client.detect_language("Как дела?")
topics = client.topics("Нейросети и алгоритмы")
summary = client.summarize("Длинный текст для резюме.", max_sentences=3)
moderation = client.moderate("Это оскорбительное сообщение.")

# Batches
batch = client.batch_synthesize([
    {"text": "Первый текст", "voice": "preset_anna"},
    {"text": "Второй текст", "voice": "dmitri"},
])
status = client.get_batch(batch["batch_id"])

analysis_batch = client.batch_analyze([
    {"audio": b64("a.wav"), "language": "ru"},
    {"audio": b64("b.wav"), "language": "ru"},
])

# Voice cloning (Pro/Business)
clone = client.create_clone_voice(
    name="My voice",
    prompt_text="Точный текст образца.",
    samples="reference.wav",
)
print(client.list_clone_voices())
client.delete_clone_voice(clone["id"])

# VAD (speech segments)
segments = client.vad("audio.wav")

# Account
usage = client.usage()
balance = client.billing_balance()

Audio effects (Pro/Business)

# Inline during synthesis — the chain is applied to the synthesized audio
audio = client.synthesize(
    "Привет!",
    effects='[{"type":"reverb","room_size":0.5},{"type":"pitch","semitones":2}]',
)

# Async processing of an existing file
import time

job = client.apply_audio_effects("voice.mp3", [{"type": "compressor", "ratio": 3}])
while client.get_audio_effects_job(job["job_id"])["status"] not in ("completed", "failed"):
    time.sleep(1)
result = client.download_audio_effects(job["job_id"])
open("voice_fx.mp3", "wb").write(result)

WebSocket streaming (Pro/Business)

Требует опциональной зависимости:

pip install "voicekit-client[streaming]"
import asyncio

async def main():
    client = VoiceKitClient(api_key="YOUR_KEY")

    stream = await client.transcribe_stream(language="ru", keyterms=["диагноз"])
    await stream.send_audio(pcm16_chunk_1)   # raw PCM16, 16 kHz mono
    await stream.send_audio(pcm16_chunk_2)
    await stream.stop()                       # finalize the utterance
    async for event in stream:                # session / vad / partial / final / error
        print(event["type"], event)
    await stream.close()

    vad = await client.vad_stream()           # VAD events only (speech_started/ended)
    await vad.send_audio(pcm16_chunk)
    await vad.stop()
    async for event in vad:
        print(event["type"], event)
    await vad.close()

asyncio.run(main())

Configuration

Option Default Description
api_key — API key (required)
base_url https://ttsapi.ru API base URL (e.g. http://localhost:5080 for local dev)
timeout 120.0 Per-request timeout in seconds

Errors raise VoiceKitError with .status (HTTP status), .code (machine-readable code) and .message.

Metadata

Release files for voicekit-client 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for voicekit-client 0.1.0
File Size Uploaded
voicekit_client-0.1.0.tar.gz 10.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for voicekit-client 0.1.0
File Interpreter ABI Platform
voicekit_client-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 19.3 kB

Release files / voicekit_client-0.1.0.tar.gz

Download URL voicekit_client-0.1.0.tar.gz
Size 10.0 kB
Tags Source
SHA-256 checksum
How to use checksums
71146d83cba23296c88ce07f3b52972a09fbaeb6b217bbf2a6f79e27a5095484
BLAKE2b-256 checksum
How to use checksums
50f818dae2aadd0d7d20fb47f95d0ff6a4d6fd2960c6b161983583e1f5a62396
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.3

Release files / voicekit_client-0.1.0-py3-none-any.whl

Download URL voicekit_client-0.1.0-py3-none-any.whl
Size 9.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
e90766ef9722c6932b08c55f2f98d0daa495af45c97991f7dbdb80a6f63369d3
BLAKE2b-256 checksum
How to use checksums
1ed4d02043556e5670eb0360ccf01bb70434fbb0a7cd09d61f1821b24cba19cc
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.3

Release history Release notifications | RSS feed

0.4.0

2 release files

0.3.0

2 release files

0.2.0

2 release files

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page