VoiceKit — Python SDK
Official Python wrapper for VoiceKit — the REST API for Russian speech: neural speech synthesis (TTS), transcription (STT) with diarization and timestamps, sentiment analysis, voice cloning, voice biometrics, audio effects and batch operations.
Links: Website · Documentation · API reference · Pricing · Blog
Install
pip install voicekit-client
Streaming (WebSocket) requires an optional extra:
pip install "voicekit-client[streaming]"
Quick start
from voicekit import VoiceKitClient, b64
client = VoiceKitClient(api_key="YOUR_KEY")
# Synthesis → raw audio bytes
audio = client.synthesize("Привет! Это синтез русской речи.", voice="preset_anna", format="mp3")
with open("speech.mp3", "wb") as f:
f.write(audio)
# Streaming (Pro/Business)
for chunk in client.synthesize_stream("Первое предложение. Второе."):
pass # write chunks to a file or socket
# Transcription (async → poll)
job = client.transcribe("audio.wav", keyterms=["диагноз"])
result = client.get_transcription_job(job["job_id"])
while result["status"] not in ("completed", "failed"):
result = client.get_transcription_job(job["job_id"])
# Short-file sync transcription
transcript = client.transcribe_sync("audio.wav")
# Analysis (sentiment + keywords + entities)
analysis = client.analyze_sync("audio.wav")
# Text intelligence
lang = client.detect_language("Как дела?")
topics = client.topics("Нейросети и алгоритмы")
summary = client.summarize("Длинный текст для резюме.", max_sentences=3)
moderation = client.moderate("Это оскорбительное сообщение.")
# Batches
batch = client.batch_synthesize([
{"text": "Первый текст", "voice": "preset_anna"},
{"text": "Второй текст", "voice": "dmitri"},
])
status = client.get_batch(batch["batch_id"])
analysis_batch = client.batch_analyze([
{"audio": b64("a.wav"), "language": "ru"},
{"audio": b64("b.wav"), "language": "ru"},
])
# Voice cloning (Pro/Business)
clone = client.create_clone_voice(
name="My voice",
prompt_text="Точный текст образца.",
samples="reference.wav",
)
print(client.list_clone_voices())
client.delete_clone_voice(clone["id"])
# VAD (speech segments)
segments = client.vad("audio.wav")
# Account
usage = client.usage()
balance = client.billing_balance()
Audio effects (Pro/Business)
# Inline during synthesis — the chain is applied to the synthesized audio
audio = client.synthesize(
"Привет!",
effects='[{"type":"reverb","room_size":0.5},{"type":"pitch","semitones":2}]',
)
# Async processing of an existing file
import time
job = client.apply_audio_effects("voice.mp3", [{"type": "compressor", "ratio": 3}])
while client.get_audio_effects_job(job["job_id"])["status"] not in ("completed", "failed"):
time.sleep(1)
result = client.download_audio_effects(job["job_id"])
open("voice_fx.mp3", "wb").write(result)
Audio cleaning (Pro/Business)
# Denoise + normalize an existing file as a background job
import time
job = client.clean_audio("noisy.wav") # one-click preset
# or with options:
# job = client.clean_audio("noisy.wav", options={"denoise": {"strength": 0.8}})
while client.get_audio_cleaning_job(job["job_id"])["status"] not in ("completed", "failed"):
time.sleep(1)
clean = client.download_audio_cleaning(job["job_id"])
open("voice_clean.wav", "wb").write(clean)
Search & Q&A (Pro/Business)
# Hybrid semantic/full-text search over your recordings
hits = client.search("почему клиент отказался?", limit=5,
keywords="дорого", source="upload")
for h in hits["hits"]:
print(h["score"], h["start"], h["text"])
# RAG question with verbatim citations
answer = client.ask("почему клиент отказался от Pro?")
print(answer["answer"])
for c in answer["citations"]:
print(c["recording_id"], c["start"], c["quote"])
WebSocket streaming (Pro/Business)
import asyncio
async def main():
client = VoiceKitClient(api_key="YOUR_KEY")
stream = await client.transcribe_stream(language="ru", keyterms=["диагноз"])
await stream.send_audio(pcm16_chunk_1) # raw PCM16, 16 kHz mono
await stream.send_audio(pcm16_chunk_2)
await stream.stop() # finalize the utterance
async for event in stream: # session / vad / partial / final / error
print(event["type"], event)
await stream.close()
vad = await client.vad_stream() # VAD events only (speech_started/ended)
await vad.send_audio(pcm16_chunk)
await vad.stop()
async for event in vad:
print(event["type"], event)
await vad.close()
asyncio.run(main())
Voice ID (Pro/Business)
# Voice passport: language, gender, age, emotion, speaker embedding, AI-vs-human
passport = client.voice_id("recording.wav")
# Voice biometrics on your own profiles
profile = client.enroll_voice("speaker.wav", name="Alice") # → {"profile_id": "voice_…"}
check = client.verify_voice("check.wav", profile["profile_id"])
# → {"profile_id": "voice_…", "similarity": 0.81, "verified": true, "threshold": 0.7}
match = client.identify_voice("check.wav") # 1:N across your profiles
# → {"best_match": {…}, "matches": […], "threshold": 0.7}
profiles = client.list_voice_profiles()
client.delete_voice_profile(profile["profile_id"])
Recordings, QA & meeting intelligence (Pro/Business)
# Recordings library (list / link channel / speakers)
recordings = client.list_recordings(limit=10)
job = client.recording_from_link("https://example.com/call.mp3") # Link channel
speakers = client.recording_speakers(recording_id)
client.update_speaker(recording_id, "SPEAKER_00", display_name="Alice", role="operator")
# Call QA
qa = client.qa_evaluate(recording_id, [{"id": "greeting", "kind": "required", "description": "..."}])
trend = client.qa_analytics(days=30)
csv = client.qa_export(format="csv")
# Meeting protocol
protocol = client.meeting_protocol(recording_id, template="standup")
# Translation & speech evaluation
translated = client.translate_transcript(job_id, target_language="en")
wer = client.evaluate("audio.wav", reference="Ожидаемый текст")
Examples
Runnable scripts live in examples/: synthesis, streaming, transcription,
analysis, voice cloning and batches. Each script reads the VOICEKIT_API_KEY
environment variable.
Configuration
| Option | Default | Description |
|---|---|---|
api_key |
— | API key (required) |
base_url |
https://ttsapi.ru |
API base URL (e.g. http://localhost:5080 for local dev) |
timeout |
120.0 |
Per-request timeout in seconds |
Errors raise VoiceKitError with .status (HTTP status), .code (machine-readable
code) and .message.
Documentation
Full API reference and guides: ttsapi.ru/docs.
License
Metadata
Release files for voicekit-client 0.4.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| voicekit_client-0.4.0.tar.gz | 16.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| voicekit_client-0.4.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 30.7 kB
Release files / voicekit_client-0.4.0.tar.gz
| Download URL | voicekit_client-0.4.0.tar.gz |
|---|---|
| Size | 16.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
2d8b2dd92689cac4ce0c01ddcb94b0d44752c469f68046af1906503cf26e0f04
|
|
BLAKE2b-256 checksum How to use checksums |
8cce2ef7e8f07010ce14451ccd4cb41407b9861c6324cd999bca63cb3d8b7cf0
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Release files / voicekit_client-0.4.0-py3-none-any.whl
| Download URL | voicekit_client-0.4.0-py3-none-any.whl |
|---|---|
| Size | 14.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
5fe89d9690dc4e90aead84c64ac71b5eb6076b141db1c30918ae0aadf6f24157
|
|
BLAKE2b-256 checksum How to use checksums |
3dee2f1a227726d0eda9a2f5ac8340d331724aed8c36371131bbcc2df45058d8
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|