Python SDK for ConvoZen TTS and STT services
Project description
ConvoZen Voice SDK
Text-to-Speech (TTS) and Speech-to-Text (STT) for Python.
One API key for both TTS and STT.
Install
pip install convozen
Requires Python 3.9+
TTS — Convert Text to Speech
Copy and run:
import convozen
client = convozen.Client(api_key="your-api-key")
audio = client.tts.synthesize("नमस्ते… मुझे आपके सर्विस के बारे में थोड़ी जानकारी चाहिए।", language="hi")
with open("output.wav", "wb") as f:
f.write(audio)
With options:
import convozen
client = convozen.Client(api_key="your-api-key")
audio = client.tts.synthesize(
text="Welcome to Playground.",
language="en", # default: "en"
voice="roohi", # default: "roohi"
model="ragini-v1", # default: "ragini-v1"
speed=1.2, # default: 1.0
)
with open("output.wav", "wb") as f:
f.write(audio)
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
text |
str |
required | Text to convert to speech |
language |
str |
"en" |
Language code (see supported languages below) |
voice |
str |
"roohi" |
Voice ID for synthesis (see available voices below) |
model |
str |
"ragini-v1" |
TTS model to use. Currently only ragini-v1 is available |
speed |
float |
1.0 |
Speech speed. 0.5 = half speed, 2.0 = double speed |
stream |
bool |
False |
If True, returns audio chunks as they are generated |
format |
str |
"wav" |
Audio output format. Currently only wav is supported |
Streaming
import convozen
client = convozen.Client(api_key="your-api-key")
for chunk in client.tts.synthesize("Long text here...", language="en", stream=True):
audio_player.write(chunk)
Available Voices
| Voice | ID | Default |
|---|---|---|
| Roohi | roohi |
Yes |
| Amaya | amaya |
|
| Kiyansh | kiyansh |
|
| Neeraj | neeraj |
|
| Manya | manya |
|
| Nidhi | nidhi |
|
| Ira | ira |
|
| Trisha | trisha |
|
| Charvi | charvi |
audio = client.tts.synthesize("Hello", language="en", voice="roohi")
STT — Convert Speech to Text
Copy and run:
import convozen
client = convozen.Client(api_key="your-api-key")
result = client.stt.transcribe("recording.wav")
print(result.text) # "hello how can I help you"
print(result.score) # -2.45 (confidence — closer to 0 is better)
By default this transcribes mono, single-speaker audio and returns an STTResponse.
Word-level timestamps
Timestamps are opt-in — pass word_time_stamps=True:
result = client.stt.transcribe("recording.wav", word_time_stamps=True)
for w in result.word_timestamps:
print(f"{w.word} {w.start_s:.2f}s – {w.end_s:.2f}s")
Works on both the transcript and diarization paths (per-turn timestamps on DiarizeTurn.word_timestamps).
Speaker Diarization (mono)
For a mono recording with multiple speakers, set diarize=True to get a per-speaker breakdown (DiarizeResponse):
result = client.stt.transcribe("call.wav", diarize=True)
for turn in result.turns:
print(f"[{turn.speaker_id}] {turn.start_sec:.1f}s – {turn.end_sec:.1f}s")
print(f" {turn.transcript}")
# [s0] 0.0s – 3.2s
# Hello, this is support. How can I help you?
# [s1] 4.0s – 7.8s
# Hi, I'd like to reschedule my appointment.
If you know how many speakers to expect, pass n_speakers (an integer 2–4) as a hint. Leave it as None (default) to auto-detect:
result = client.stt.transcribe("interview.wav", diarize=True, n_speakers=2)
Response fields (DiarizeResponse):
| Field | Type | Description |
|---|---|---|
turns |
list[DiarizeTurn] |
Speaker turns in order |
num_speakers |
int |
Number of distinct speakers detected |
total_duration_sec |
float |
Total audio duration in seconds |
Each DiarizeTurn has speaker_id, start_sec, end_sec, transcript, and word_timestamps (populated when word_time_stamps=True).
Stereo audio
If each speaker is on a separate channel (e.g. a 2-channel call recording), set audio_channels="stereo". Each channel is treated as one speaker, and you get a DiarizeResponse:
result = client.stt.transcribe("stereo_call.wav", audio_channels="stereo")
For stereo, keep
diarize=False(the channels already are the speakers) and leaven_speakers=None. The declaredaudio_channelsmust match the file — a mono file declared"stereo"(or vice-versa) is rejected with a400.
Denoising
For noisy multi-speaker audio, enable denoise=True. It applies on the diarization path (diarize=True or audio_channels="stereo"):
result = client.stt.transcribe("noisy_call.wav", diarize=True, denoise=True)
Specify Languages
If you know which languages are spoken in the audio, pass them in lang_tags. This improves accuracy — especially for multilingual audio like Hindi + English call recordings:
result = client.stt.transcribe(
"recording.wav",
lang_tags=["hi", "en"],
)
print(result.text) # "हां मुझे appointment reschedule करना है"
Keyword Boosting
Have domain-specific words the model keeps getting wrong? Pass them in keywords and the model will bias towards recognizing them:
# Without keyword boosting: "can you tell me about conversion"
# With keyword boosting: "can you tell me about convozen"
result = client.stt.transcribe(
"recording.wav",
keywords=["convozen", "akshara"],
)
Choose a Model
Use model to select which ASR model processes the audio:
result = client.stt.transcribe("recording.wav", model="akshara-pro") # default
result = client.stt.transcribe("recording.wav", model="akshara-base")
| Model | Best for |
|---|---|
akshara-pro |
Maximum accuracy (default) |
akshara-base |
Lighter / faster |
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
audio |
str | os.PathLike |
required | Path to the audio file |
model |
str |
"akshara-pro" |
ASR model: "akshara-pro" or "akshara-base" |
audio_channels |
str |
"mono" |
"mono" or "stereo". Must match the actual file |
diarize |
bool |
False |
True to return per-speaker turns (mono only) |
n_speakers |
int |
None |
Expected speaker count (2–4) for mono diarization. None → auto-detect |
denoise |
bool |
False |
Apply denoising on the diarization path — helps on noisy audio |
lang_tags |
list[str] |
None |
Languages spoken in the audio — improves accuracy when known |
word_time_stamps |
bool |
False |
True to return per-word timestamps |
blank_penalty |
float |
None |
Penalizes silence/blank tokens; increase to reduce empty gaps. None → derived from lang_tags |
keywords |
list[str] |
None |
Domain words the model should bias towards recognizing |
When to use what
audio_channels |
diarize |
n_speakers |
Result |
|---|---|---|---|
"mono" |
False |
None |
Plain transcript → STTResponse |
"mono" |
True |
None |
Auto speaker count → DiarizeResponse |
"mono" |
True |
2–4 |
Hinted speaker count → DiarizeResponse |
"stereo" |
False |
None |
Channel-split diarization → DiarizeResponse |
Anything outside these rows raises a ValueError (e.g. n_speakers without diarize, n_speakers outside 2–4, diarize=True with "stereo").
Supported Languages
| Code | Language | TTS | STT |
|---|---|---|---|
en |
English | Yes | Yes |
hi |
Hindi | Yes | Yes |
ta |
Tamil | Yes | Yes |
te |
Telugu | Yes | Yes |
kn |
Kannada | Yes | Yes |
mr |
Marathi | — | Yes |
bn |
Bengali | — | Yes |
gu |
Gujarati | — | Yes |
ml |
Malayalam | — | Yes |
Check Credits
import convozen
client = convozen.Client(api_key="your-api-key")
info = client.account.info()
print(info.credits.tts.balance)
print(info.credits.stt.balance)
Error Handling
import convozen
from convozen import AuthenticationError, RateLimitError, APIError
client = convozen.Client(api_key="your-api-key")
try:
audio = client.tts.synthesize("Hello")
except AuthenticationError:
print("Invalid API key")
except RateLimitError:
print("Too many requests")
except APIError as e:
print(f"Server error: {e}")
Standalone Clients
from convozen import TTS, STT
tts = TTS(api_key="your-api-key")
audio = tts.synthesize("Hello", language="hi")
stt = STT(api_key="your-api-key")
# Flat transcript
result = stt.transcribe("recording.wav")
print(result.text)
# Per-speaker diarization (mono, multiple speakers)
result = stt.transcribe("call.wav", diarize=True)
for turn in result.turns:
print(f"[{turn.speaker_id}] {turn.start_sec:.2f}s – {turn.end_sec:.2f}s : {turn.transcript}")
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file convozen-0.3.8.tar.gz.
File metadata
- Download URL: convozen-0.3.8.tar.gz
- Upload date:
- Size: 15.9 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.11
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
ce637397080fc8d4ac903792e1d97f3e89b7682dde2d11a185de0a1d9cbdec1a
|
|
| MD5 |
13600fbc08083e409fca89fe4cbcde2c
|
|
| BLAKE2b-256 |
24c04db05c57dcd079f8955e5829fc241f0424b69ee6e832296e880caf7fdbcf
|
File details
Details for the file convozen-0.3.8-py3-none-any.whl.
File metadata
- Download URL: convozen-0.3.8-py3-none-any.whl
- Upload date:
- Size: 20.8 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.11
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
5fede1f4ce2e78e16f740b4e12f73b31dab5cb020ae74dcb291ed2ac478270d4
|
|
| MD5 |
4f42606f80c15e47dc066ba36015db84
|
|
| BLAKE2b-256 |
3dbba64915587102e8e0afa9f3e4e4b3be9918ebfef850af1dfaa7d5b9d3c840
|