Skip to main content

VoiceKit — Python SDK

PyPI version Python License: MIT

Official Python wrapper for VoiceKit — the REST API for Russian speech: neural speech synthesis (TTS), transcription (STT) with diarization and timestamps, sentiment analysis, voice cloning, voice biometrics, audio effects and batch operations.

Links: Website · Documentation · API reference · Pricing · Blog

Install

pip install voicekit-client

Streaming (WebSocket) requires an optional extra:

pip install "voicekit-client[streaming]"

Quick start

from voicekit import VoiceKitClient, b64

client = VoiceKitClient(api_key="YOUR_KEY")

# Synthesis → raw audio bytes
audio = client.synthesize("Привет! Это синтез русской речи.", voice="preset_anna", format="mp3")
with open("speech.mp3", "wb") as f:
    f.write(audio)

# Streaming (Pro/Business)
for chunk in client.synthesize_stream("Первое предложение. Второе."):
    pass  # write chunks to a file or socket

# Transcription (async → poll)
job = client.transcribe("audio.wav", keyterms=["диагноз"])
result = client.get_transcription_job(job["job_id"])
while result["status"] not in ("completed", "failed"):
    result = client.get_transcription_job(job["job_id"])

# Short-file sync transcription
transcript = client.transcribe_sync("audio.wav")

# Analysis (sentiment + keywords + entities)
analysis = client.analyze_sync("audio.wav")

# Text intelligence
lang = client.detect_language("Как дела?")
topics = client.topics("Нейросети и алгоритмы")
summary = client.summarize("Длинный текст для резюме.", max_sentences=3)
moderation = client.moderate("Это оскорбительное сообщение.")

# Batches
batch = client.batch_synthesize([
    {"text": "Первый текст", "voice": "preset_anna"},
    {"text": "Второй текст", "voice": "dmitri"},
])
status = client.get_batch(batch["batch_id"])

analysis_batch = client.batch_analyze([
    {"audio": b64("a.wav"), "language": "ru"},
    {"audio": b64("b.wav"), "language": "ru"},
])

# Voice cloning (Pro/Business)
clone = client.create_clone_voice(
    name="My voice",
    prompt_text="Точный текст образца.",
    samples="reference.wav",
)
print(client.list_clone_voices())
client.delete_clone_voice(clone["id"])

# VAD (speech segments)
segments = client.vad("audio.wav")

# Account
usage = client.usage()
balance = client.billing_balance()

Audio effects (Pro/Business)

# Inline during synthesis — the chain is applied to the synthesized audio
audio = client.synthesize(
    "Привет!",
    effects='[{"type":"reverb","room_size":0.5},{"type":"pitch","semitones":2}]',
)

# Async processing of an existing file
import time

job = client.apply_audio_effects("voice.mp3", [{"type": "compressor", "ratio": 3}])
while client.get_audio_effects_job(job["job_id"])["status"] not in ("completed", "failed"):
    time.sleep(1)
result = client.download_audio_effects(job["job_id"])
open("voice_fx.mp3", "wb").write(result)

Audio cleaning (Pro/Business)

# Denoise + normalize an existing file as a background job
import time

job = client.clean_audio("noisy.wav")            # one-click preset
# or with options:
# job = client.clean_audio("noisy.wav", options={"denoise": {"strength": 0.8}})

while client.get_audio_cleaning_job(job["job_id"])["status"] not in ("completed", "failed"):
    time.sleep(1)

clean = client.download_audio_cleaning(job["job_id"])
open("voice_clean.wav", "wb").write(clean)

WebSocket streaming (Pro/Business)

import asyncio

async def main():
    client = VoiceKitClient(api_key="YOUR_KEY")

    stream = await client.transcribe_stream(language="ru", keyterms=["диагноз"])
    await stream.send_audio(pcm16_chunk_1)   # raw PCM16, 16 kHz mono
    await stream.send_audio(pcm16_chunk_2)
    await stream.stop()                       # finalize the utterance
    async for event in stream:                # session / vad / partial / final / error
        print(event["type"], event)
    await stream.close()

    vad = await client.vad_stream()           # VAD events only (speech_started/ended)
    await vad.send_audio(pcm16_chunk)
    await vad.stop()
    async for event in vad:
        print(event["type"], event)
    await vad.close()

asyncio.run(main())

Voice ID (Pro/Business)

# Voice passport: language, gender, age, emotion, speaker embedding, AI-vs-human
passport = client.voice_id("recording.wav")

# Voice biometrics on your own profiles
profile = client.enroll_voice("speaker.wav", name="Alice")    # → {"profile_id": "voice_…"}

check = client.verify_voice("check.wav", profile["profile_id"])
# → {"profile_id": "voice_…", "similarity": 0.81, "verified": true, "threshold": 0.7}

match = client.identify_voice("check.wav")                    # 1:N across your profiles
# → {"best_match": {…}, "matches": […], "threshold": 0.7}

profiles = client.list_voice_profiles()
client.delete_voice_profile(profile["profile_id"])

Examples

Runnable scripts live in examples/: synthesis, streaming, transcription, analysis, voice cloning and batches. Each script reads the VOICEKIT_API_KEY environment variable.

Configuration

Option Default Description
api_key — API key (required)
base_url https://ttsapi.ru API base URL (e.g. http://localhost:5080 for local dev)
timeout 120.0 Per-request timeout in seconds

Errors raise VoiceKitError with .status (HTTP status), .code (machine-readable code) and .message.

Documentation

Full API reference and guides: ttsapi.ru/docs.

License

MIT

Metadata

Release files for voicekit-client 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for voicekit-client 0.3.0
File Size Uploaded
voicekit_client-0.3.0.tar.gz 13.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for voicekit-client 0.3.0
File Interpreter ABI Platform
voicekit_client-0.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 24.6 kB

Release files / voicekit_client-0.3.0.tar.gz

Download URL voicekit_client-0.3.0.tar.gz
Size 13.0 kB
Tags Source
SHA-256 checksum
How to use checksums
04e94a0e5f1652719a15835b07fb9364efcbca5a6cabb2b625cef32c009c534b
BLAKE2b-256 checksum
How to use checksums
ff1902b973847b1e4429f808f4d20ea152307fd7df50c725583686e0d55e6f77
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / voicekit_client-0.3.0-py3-none-any.whl

Download URL voicekit_client-0.3.0-py3-none-any.whl
Size 11.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
26fbc7cb83ca4c791e59fe57c95db7b2da96ffab814f83847ca936a8d2a2f42a
BLAKE2b-256 checksum
How to use checksums
b36b1ce55668433cbb7cc2150835b119efb68b22f66a7cbfca291462ec119e39
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release history Release notifications | RSS feed

0.4.0

2 release files

This release

0.3.0 This release

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page