Skip to main content

VoiceKit — Python SDK

PyPI version Python License: MIT

Official Python wrapper for VoiceKit — the REST API for Russian speech: neural speech synthesis (TTS), transcription (STT) with diarization and timestamps, sentiment analysis, voice cloning, voice biometrics, audio effects and batch operations.

Links: Website · Documentation · API reference · Pricing · Blog

Install

pip install voicekit-client

Streaming (WebSocket) requires an optional extra:

pip install "voicekit-client[streaming]"

Quick start

from voicekit import VoiceKitClient, b64

client = VoiceKitClient(api_key="YOUR_KEY")

# Synthesis → raw audio bytes
audio = client.synthesize("Привет! Это синтез русской речи.", voice="preset_anna", format="mp3")
with open("speech.mp3", "wb") as f:
    f.write(audio)

# Streaming (Pro/Business)
for chunk in client.synthesize_stream("Первое предложение. Второе."):
    pass  # write chunks to a file or socket

# Transcription (async → poll)
job = client.transcribe("audio.wav", keyterms=["диагноз"])
result = client.get_transcription_job(job["job_id"])
while result["status"] not in ("completed", "failed"):
    result = client.get_transcription_job(job["job_id"])

# Short-file sync transcription
transcript = client.transcribe_sync("audio.wav")

# Analysis (sentiment + keywords + entities)
analysis = client.analyze_sync("audio.wav")

# Text intelligence
lang = client.detect_language("Как дела?")
topics = client.topics("Нейросети и алгоритмы")
summary = client.summarize("Длинный текст для резюме.", max_sentences=3)
moderation = client.moderate("Это оскорбительное сообщение.")

# Batches
batch = client.batch_synthesize([
    {"text": "Первый текст", "voice": "preset_anna"},
    {"text": "Второй текст", "voice": "dmitri"},
])
status = client.get_batch(batch["batch_id"])

analysis_batch = client.batch_analyze([
    {"audio": b64("a.wav"), "language": "ru"},
    {"audio": b64("b.wav"), "language": "ru"},
])

# Voice cloning (Pro/Business)
clone = client.create_clone_voice(
    name="My voice",
    prompt_text="Точный текст образца.",
    samples="reference.wav",
)
print(client.list_clone_voices())
client.delete_clone_voice(clone["id"])

# VAD (speech segments)
segments = client.vad("audio.wav")

# Account
usage = client.usage()
balance = client.billing_balance()

Audio effects (Pro/Business)

# Inline during synthesis — the chain is applied to the synthesized audio
audio = client.synthesize(
    "Привет!",
    effects='[{"type":"reverb","room_size":0.5},{"type":"pitch","semitones":2}]',
)

# Async processing of an existing file
import time

job = client.apply_audio_effects("voice.mp3", [{"type": "compressor", "ratio": 3}])
while client.get_audio_effects_job(job["job_id"])["status"] not in ("completed", "failed"):
    time.sleep(1)
result = client.download_audio_effects(job["job_id"])
open("voice_fx.mp3", "wb").write(result)

Audio cleaning (Pro/Business)

# Denoise + normalize an existing file as a background job
import time

job = client.clean_audio("noisy.wav")            # one-click preset
# or with options:
# job = client.clean_audio("noisy.wav", options={"denoise": {"strength": 0.8}})

while client.get_audio_cleaning_job(job["job_id"])["status"] not in ("completed", "failed"):
    time.sleep(1)

clean = client.download_audio_cleaning(job["job_id"])
open("voice_clean.wav", "wb").write(clean)

Search & Q&A (Pro/Business)

# Hybrid semantic/full-text search over your recordings
hits = client.search("почему клиент отказался?", limit=5,
                     keywords="дорого", source="upload")
for h in hits["hits"]:
    print(h["score"], h["start"], h["text"])

# RAG question with verbatim citations
answer = client.ask("почему клиент отказался от Pro?")
print(answer["answer"])
for c in answer["citations"]:
    print(c["recording_id"], c["start"], c["quote"])

WebSocket streaming (Pro/Business)

import asyncio

async def main():
    client = VoiceKitClient(api_key="YOUR_KEY")

    stream = await client.transcribe_stream(language="ru", keyterms=["диагноз"])
    await stream.send_audio(pcm16_chunk_1)   # raw PCM16, 16 kHz mono
    await stream.send_audio(pcm16_chunk_2)
    await stream.stop()                       # finalize the utterance
    async for event in stream:                # session / vad / partial / final / error
        print(event["type"], event)
    await stream.close()

    vad = await client.vad_stream()           # VAD events only (speech_started/ended)
    await vad.send_audio(pcm16_chunk)
    await vad.stop()
    async for event in vad:
        print(event["type"], event)
    await vad.close()

asyncio.run(main())

Voice ID (Pro/Business)

# Voice passport: language, gender, age, emotion, speaker embedding, AI-vs-human
passport = client.voice_id("recording.wav")

# Voice biometrics on your own profiles
profile = client.enroll_voice("speaker.wav", name="Alice")    # → {"profile_id": "voice_…"}

check = client.verify_voice("check.wav", profile["profile_id"])
# → {"profile_id": "voice_…", "similarity": 0.81, "verified": true, "threshold": 0.7}

match = client.identify_voice("check.wav")                    # 1:N across your profiles
# → {"best_match": {…}, "matches": […], "threshold": 0.7}

profiles = client.list_voice_profiles()
client.delete_voice_profile(profile["profile_id"])

Recordings, QA & meeting intelligence (Pro/Business)

# Recordings library (list / link channel / speakers)
recordings = client.list_recordings(limit=10)
job = client.recording_from_link("https://example.com/call.mp3")   # Link channel
speakers = client.recording_speakers(recording_id)
client.update_speaker(recording_id, "SPEAKER_00", display_name="Alice", role="operator")

# Call QA
qa = client.qa_evaluate(recording_id, [{"id": "greeting", "kind": "required", "description": "..."}])
trend = client.qa_analytics(days=30)
csv = client.qa_export(format="csv")

# Meeting protocol
protocol = client.meeting_protocol(recording_id, template="standup")

# Translation & speech evaluation
translated = client.translate_transcript(job_id, target_language="en")
wer = client.evaluate("audio.wav", reference="Ожидаемый текст")

Examples

Runnable scripts live in examples/: synthesis, streaming, transcription, analysis, voice cloning and batches. Each script reads the VOICEKIT_API_KEY environment variable.

Configuration

Option Default Description
api_key — API key (required)
base_url https://ttsapi.ru API base URL (e.g. http://localhost:5080 for local dev)
timeout 120.0 Per-request timeout in seconds

Errors raise VoiceKitError with .status (HTTP status), .code (machine-readable code) and .message.

Documentation

Full API reference and guides: ttsapi.ru/docs.

License

MIT

Metadata

Release files for voicekit-client 0.4.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for voicekit-client 0.4.0
File Size Uploaded
voicekit_client-0.4.0.tar.gz 16.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for voicekit-client 0.4.0
File Interpreter ABI Platform
voicekit_client-0.4.0-py3-none-any.whl Python 3 none any Details

Total release size: 30.7 kB

Release files / voicekit_client-0.4.0.tar.gz

Download URL voicekit_client-0.4.0.tar.gz
Size 16.4 kB
Tags Source
SHA-256 checksum
How to use checksums
2d8b2dd92689cac4ce0c01ddcb94b0d44752c469f68046af1906503cf26e0f04
BLAKE2b-256 checksum
How to use checksums
8cce2ef7e8f07010ce14451ccd4cb41407b9861c6324cd999bca63cb3d8b7cf0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / voicekit_client-0.4.0-py3-none-any.whl

Download URL voicekit_client-0.4.0-py3-none-any.whl
Size 14.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
5fe89d9690dc4e90aead84c64ac71b5eb6076b141db1c30918ae0aadf6f24157
BLAKE2b-256 checksum
How to use checksums
3dee2f1a227726d0eda9a2f5ac8340d331724aed8c36371131bbcc2df45058d8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release history Release notifications | RSS feed

This release

0.4.0 This release

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page