Skip to main content

Zeli Avatar SDK (Python)

Real-time, lip-synced talking-head avatars streamed to your app over WebRTC. Connect to a Zeli avatar server, receive synchronized audio + video frames, and drive the avatar with text, either through the conversational model or straight to text-to-speech.

📖 Documentation: this README covers the whole surface, and runnable programs for each capability ship in examples/. A hosted documentation site is coming; it is deliberately not linked here yet, because a link in PyPI metadata cannot be corrected after a version is published.

pip install zeli-avatar
# with display helpers (OpenCV / sounddevice) for the examples:
pip install "zeli-avatar[display]"

Requires Python 3.10+. You also need a running Zeli avatar server, and an API key to authenticate against it.

client = ZeliClient(
    api_key="zsk_live_...",                       # or os.environ["ZELI_API_KEY"]
    avatar_config=AvatarConfig(avatar_id="01-presenter-male__confident"),
    options=ClientOptions(server_url="https://your-zeli-host"),
)

The key is sent as an X-Api-Key header on every request. A rejected key raises AuthenticationError.

Which box am I talking to

There is no default base URL, and that is deliberate. This product is deployed per environment, so a client that picks one for you picks an environment for you, and the only way to find out which is to look at the traffic. Name it one of two ways:

# In the code, when the app knows its own target:
client = ZeliClient(api_key=..., options=ClientOptions(server_url="https://your-box.example.com"))

# Or from the environment, when the deployment knows and the code should not:
#   export ZELI_SERVER_URL=https://your-box.example.com
client = ZeliClient(api_key=...)

Say neither and ZeliClient(...) raises ConfigurationError naming both options. Say both and they must agree: two different URLs raise rather than one silently beating the other, because the developer who exported ZELI_SERVER_URL to aim a script somewhere else would otherwise be overruled with nothing printed. The URL must start with http:// or https://; a bare hostname is refused at construction rather than becoming an InvalidURL deep inside a request, or a silent ws:// downgrade on the control channel.

The variable name is exported as zeli.SERVER_URL_ENV_VAR, and the JavaScript SDK reads the same one.

Quickstart

import asyncio, os
from zeli import ZeliClient, AvatarConfig, ClientOptions, ZeliEvent

client = ZeliClient(
    api_key=os.environ["ZELI_API_KEY"],
    avatar_config=AvatarConfig(avatar_id="01-presenter-male__confident", voice_id="your-voice-id"),
    options=ClientOptions(server_url="http://your-server:8080"),
)

@client.on(ZeliEvent.MESSAGE_RECEIVED)
async def on_message(message):
    print(f"{message.role.value}: {message.content}")

# send_message() and talk() hand work to the server and return immediately, so a
# server-side refusal ("TTS not configured", unknown voice) arrives here and
# nowhere else. Without this handler the script prints nothing and looks hung.
@client.on(ZeliEvent.ERROR)
async def on_error(err):
    print("server error:", err)

async def main():
    async with client.connect() as session:
        await session.send_message("Hi! Introduce yourself in one sentence.")
        await session.wait_until_closed()

asyncio.run(main())

Core concepts

Object What it is
ZeliClient Entry point. Holds config + event handlers; opens sessions.
ManagementApi client.management: avatars, voices, settings, session tokens. Needs a full API key, so it belongs on your server.
AvatarConfig Which avatar/voice/persona to bring to life.
ClientOptions Where the server is (server_url), ICE servers, timeouts.
Session One live connection: media streams + control methods.
TalkStream Push text to the avatar's voice incrementally.
ZeliEvent The events you can subscribe to with @client.on(...).

Connecting

# Use an already-prepared avatar by id (uploaded files are keyed by filename stem):
client = ZeliClient(
    api_key=os.environ["ZELI_API_KEY"],
    avatar_id="01-presenter-male__confident",
    options=ClientOptions(server_url="https://your-box.example.com"),
)

# Or pass the config object, which also carries the voice:
client = ZeliClient(
    api_key=os.environ["ZELI_API_KEY"],
    avatar_config=AvatarConfig(
        avatar_id="01-presenter-male__confident",
        voice_id="your-voice-id",
    ),
    options=ClientOptions(server_url="https://your-box.example.com"),
)

async with client.connect() as session:
    print(session.session_id, session.avatar)
    await session.wait_until_closed()

An avatar id the server cannot prepare is not an error there. It answers 200, echoes your id back, and streams a clip it does have instead. The SDK turns that into a SERVER_WARNING and exposes the substitute as session.substituted_avatar (None when you got what you asked for), because session.avatar only reflects what you requested. Check ids against await client.list_avatars().

Receiving media

Video and audio arrive as PyAV frames you can turn into NumPy arrays.

async with client.connect() as session:
    async def show_video():
        async for frame in session.video_frames():
            img = frame.to_ndarray(format="rgb24")   # (H, W, 3) uint8

    async def play_audio():
        async for frame in session.audio_frames():
            samples = frame.to_ndarray()              # int16 PCM, 48 kHz stereo

    await asyncio.gather(show_video(), play_audio())

Driving the avatar

# Through the conversational model (LLM -> reply -> TTS -> avatar):
await session.send_message("What's the weather like on Mars?")

# Straight to text-to-speech, bypassing the model:
await session.talk("This line is spoken immediately.")

# Incrementally, as text becomes available:
async with session.create_talk_stream() as talk:
    await talk.send("Streaming ")
    await talk.send("this ")
    await talk.send("out loud.", end_of_speech=True)

# Cut the avatar off mid-sentence (barge-in):
await session.interrupt()

# Start a fresh conversation without dropping the session. The avatar stays on
# screen and connected; it just stops remembering what was said.
conversation_id = await session.new_conversation()

new_conversation() erases nothing. The box keeps the conversation you closed, so its turns stay browsable, and opens a new one beside it. The persona and the system prompt survive, because only the turn list rotates: the avatar keeps its character and loses its memory.

It returns the new conversation id, or None from a box old enough not to report one. None is not a failure. The conversation still rotated; the box simply did not say which one replaced it.

It also clears the local transcript, so client.get_message_history() agrees with the box rather than continuing to show a conversation the avatar has forgotten. Handlers registered on MESSAGE_HISTORY_UPDATED receive the empty history, and the clear happens only after the box confirms: a refused reset leaves your view intact, because what is on screen is still the truth.

Scoped to your credential. A box is shared, and the conversation id is never taken from the caller, so this cannot reach anybody else's history.

Events

Register handlers as decorators or with add_listener. Handlers may be sync or async.

from zeli import ZeliEvent

@client.on(ZeliEvent.SESSION_READY)
async def on_ready(info):
    print("ready:", info.session_id)

@client.on(ZeliEvent.MESSAGE_STREAM_EVENT_RECEIVED)
async def on_chunk(event):
    print(event.content, end="", flush=True)   # transcript, clause by clause

@client.on(ZeliEvent.AVATAR_SPEECH_STARTED)
async def on_speaking(correlation_id):
    print("avatar started speaking")
Event Fires when
CONNECTION_ESTABLISHED The media connection is up.
SESSION_READY The control channel handshake completed.
MESSAGE_RECEIVED A user or assistant message is finalized.
MESSAGE_STREAM_EVENT_RECEIVED An incremental transcript chunk arrives.
MESSAGE_HISTORY_UPDATED The transcript changed.
AVATAR_SPEECH_STARTED / AVATAR_SPEECH_ENDED The avatar starts / stops speaking.
TALK_STREAM_INTERRUPTED A barge-in interrupted playback.
CONNECTION_CLOSED The session ended. Carries a ConnectionCloseCode: normal when you closed it, server_closed when the server did, webrtc_failure when the transport broke.
ERROR / SERVER_WARNING A server-side error / warning.
EMOTION_DETECTED, USER_SPEECH_STARTED, USER_SPEECH_ENDED Reserved, not currently emitted. No Zeli server sends the frames these map to.

SESSION_READY hands you a SessionInfo carrying session_id, the avatar the server bound the session to, and that avatar's tones, the only place to learn which values session.talk(tone=...) will accept on this server.

The box holds one live session at a time, so anyone else connecting takes your slot. That arrives as CONNECTION_CLOSED with server_closed, and session.is_active goes False; handle it if your process is long-lived.

Read the running transcript any time with client.get_message_history().

Server-side management

Everything that is not the live conversation lives on client.management: creating avatars and voices, reading and writing the box's configuration, and minting the short-lived token a browser streams with.

It is a separate namespace because most of it needs a full API key. A zsk_live_ key can create avatars, upload voices and mint tokens, so it belongs on your server and must never reach page JavaScript. The session half of the SDK is what runs beside a viewer, on a token your server mints. Keeping the privileged calls behind one name makes that boundary visible.

client = ZeliClient(
    api_key=os.environ["ZELI_API_KEY"],
    options=ClientOptions(server_url="https://your-box.example.com"),
)

# Mint the credential a browser streams with. Your backend does this, never the page.
token = await client.management.create_session_token(expires_in_seconds=600)
print(token.token, token.seconds_remaining(now=time.time()))

# Build an avatar from a portrait or a video, then wait for it.
status = await client.management.create_avatar("alice.png", name="alice", tones=["confident"])
while not avatar_is_ready(await client.management.list_avatars(), status.avatar_id):
    await asyncio.sleep(2)

Keeping a session alive past its token

A session token lives for minutes, and a conversation can easily outlive it. When it runs out the avatar simply stops answering: the next request is refused and the caller is told the API key is wrong, which is both unhelpful and untrue.

The SDK cannot mint a replacement, and that is by design rather than an omission. Minting needs the full API key, and the whole point of a session token is that the full key never reaches the client holding it. So the renewal has to come from a server, and token_provider is how it gets there:

async def fetch_token() -> str:
    async with aiohttp.ClientSession() as http:          # YOUR endpoint
        async with http.get("https://your-app.example.com/api/zeli-token") as resp:
            return (await resp.json())["token"]

client = ZeliClient(
    session_token=first_token,
    token_provider=fetch_token,
    options=ClientOptions(server_url="https://your-box.example.com"),
)

That endpoint is yours: it calls create_session_token() with the full key, exactly as above, and returns the token. The SDK then calls the provider when the token it holds is within fifteen seconds of expiring, and again if the box refuses it anyway. Both paths are needed: the second covers a client clock and a box clock that disagree, which would otherwise strand a session this SDK believed still had time left.

Four things are worth knowing:

  • It is optional. Leave it out and the client behaves exactly as it always has.
  • It is never called twice at once. Two requests refused at the same moment share one refresh, and both then proceed. Your endpoint is not hammered, and you do not end up with a spare token nobody used.
  • A failure is named as yours. If your endpoint is down the SDK raises TokenProviderError, not AuthenticationError, so nobody goes looking at an API key that was never part of the call. The original failure is the __cause__.
  • It is refused beside api_key. A full key does not expire, so there is nothing to refresh, and refreshing one would mean treating the full key and the browser token as the same kind of credential. That pairing raises a ConfigurationError.

The provider may return a SessionToken rather than a bare string, which is what create_session_token() already hands back. Worth doing: a bare string cannot say when it expires, so the SDK can only react to a refusal, while the expiry keeps the renewal ahead of the failure for every token after the first.

client.credential exposes seconds_remaining, presented and refresh() for a caller who wants to drive it, for instance to renew before something slow rather than during it.

Operation What it does
create_session_token(expires_in_seconds=None) Mint the browser credential. Returns a SessionToken; the box clamps the TTL, currently to 600s.
get_settings() / update_settings(patch) Read and merge the box's configuration for this caller. A partial patch: omitted fields keep their value.
get_status() Whether the language model and the voice engine are both ready.
clear_conversation() Forget this caller's history. Also interrupts a reply in progress.
list_voices() Every voice the engine offers. An id from here goes into AvatarConfig.voice_id.
preview_voice(voice_id=None) Audition a voice. Returns a BinaryAsset (WAV bytes plus content type).
create_voice(file_path, name=...) Clone a voice from a reference clip of at least 6 seconds. Minutes of work.
prepare_voice(voice_id) Ask the engine to make a voice ready to speak.
list_avatars() Every avatar, with its tone variants and what is still preparing.
create_avatar(file_path, name=None, tones=None, consent_token=None, framing=None, gesture_amplitude=None) Create from a portrait or a video. Answers immediately with preparing.
create_photo_avatar(photo_path, avatar_id=...) Build from a single portrait and wait for every clip. Gated by a feature flag and a consent check.
delete_avatar(avatar_id) Delete an uploaded avatar. Never a stock one.
list_avatar_clips(avatar_id) How each emotional variant is getting on. Sparse: a missing tone was never requested.
get_avatar_clip(avatar_id, tone=None) One rendered clip, as MP4 bytes.
retry_avatar_clip(avatar_id, tone) Regenerate one tone. Answers 202; a 409 means one is already running.
get_avatar_preview(avatar_id) A still frame, as image bytes.
upload_audio(file_path, filename=None) Upload an audio file for the box to play.

The JavaScript SDK offers the same eighteen operations under the same names, spelled client.management.createSessionToken(...) and so on: every name here is that name in snake_case, with no exceptions, so you can predict one from the other rather than look it up. Results are the box's own JSON with the box's own field names in both SDKs. The two signature differences are deliberate: this SDK takes a filesystem path where the browser one takes a Blob, because a Python caller is the server.

create_avatar, create_avatar_from_photo, create_session_token and list_avatars are also still reachable directly on ZeliClient. They keep working and delegate to client.management, but they are deprecated. Note the rename on the way across: client.create_avatar_from_photo(...) is client.management.create_photo_avatar(...), which is the name the other SDK uses.

Configuration reference

ClientOptions

Field Default Meaning
server_url none. ZELI_SERVER_URL, or an error. Base URL of the avatar server.
connect_path /connect Persistent WebRTC offer/answer path.
ice_servers Google STUN ICE servers for NAT traversal.
connect_timeout 30.0 Seconds to wait for the media connection.

ZeliClient credential arguments, all keyword only:

Argument Meaning
api_key A full API key (zsk_live_...). Server side only: it can create avatars, upload voices and mint tokens.
session_token A short lived zsk_temp_... credential, as a string or as the SessionToken the mint route returns. Wins over api_key when both are set, because the only reason to have both is a copy paste mid migration and the safer one should win.
token_provider An async callable that asks YOUR backend for a fresh session token. Called as expiry approaches and on a refusal, never twice at once. Optional; omitting it leaves today's behaviour unchanged. Refused beside api_key. See above.

AvatarConfig: only avatar_id and voice_id do anything. avatar_id is sent on POST /connect; voice_id rides along with every send_message / talk / talkstream. The remaining fields (name, system_prompt, llm_id, language_code, emotion_responsive, enhance, enhance_strength, loop_mode, max_session_length_seconds) have no representation on the wire protocol: the SDK never transmits them and no Zeli server reads them. Setting one logs a warning on the zeli logger rather than silently doing nothing. Their server-side equivalents (system prompt, model, TTS provider and voice) are configured on the box itself, per box rather than per session, and this SDK can read and write them: await client.management.get_settings() and await client.management.update_settings({...}).

Error handling

from zeli import (
    ZeliError, ConnectionError, ConsentError, SessionError,
    AuthenticationError, TokenProviderError,
)

try:
    async with client.connect() as session:
        await session.wait_until_closed()
except AuthenticationError:
    ...   # server rejected the connection
except TimeoutError:
    ...   # a transport didn't come up in time; also a ConnectionError
except ConnectionError:
    ...   # couldn't reach the server / WebRTC failed
except SessionError:
    ...   # server refused or a session operation failed
except ConsentError:
    ...   # a photo avatar was refused because consent isn't verified
except TokenProviderError:
    ...   # YOUR token endpoint failed; the box and the API key were not involved
except ZeliError as e:
    print(f"[{e.code}] {e.message}")

ConsentError is deliberately not an AuthenticationError. The key was fine; the consent record was not, and reporting a rejected key would send you to rotate a credential that was never the problem. The JavaScript SDK has no equivalent yet and reports that refusal as an authentication failure, so if you are moving code across, catch it here and check the message there.

Compatibility

The SDK connects over the server's POST /connect (media) and, when present, the /api/session/ws control gateway (talk, incremental transcripts, streaming events). Against a server without the gateway it automatically falls back to POST /api/chat for send_message. talk/create_talk_stream then require an upgraded server.

Examples

See examples/: quickstart.py, save_video.py, talk_stream.py. All three take their API key from ZELI_API_KEY and let the SDK itself read ZELI_SERVER_URL, so none of them names a box in code:

export ZELI_SERVER_URL=https://your-box.example.com
export ZELI_API_KEY=zsk_live_...
export ZELI_AVATAR_ID=01-presenter-male__confident   # optional
python examples/quickstart.py

License

Apache-2.0.

Release files for zeli-avatar 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for zeli-avatar 0.1.0
File Size Uploaded
zeli_avatar-0.1.0.tar.gz 61.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for zeli-avatar 0.1.0
File Interpreter ABI Platform
zeli_avatar-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 131.0 kB

Release files / zeli_avatar-0.1.0.tar.gz

Download URL zeli_avatar-0.1.0.tar.gz
Size 61.0 kB
Tags Source
SHA-256 checksum
How to use checksums
dbce54714518e1ad20a18ba9d484d0eda73d902effb34b20499652b2eae9a998
BLAKE2b-256 checksum
How to use checksums
23a706a31ee021706c2a411bddefb9b47b536804bfa88a7f8235fc30e3a66932
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 17, 2026.

Transparency log

Release files / zeli_avatar-0.1.0-py3-none-any.whl

Download URL zeli_avatar-0.1.0-py3-none-any.whl
Size 70.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
7fa470958b7193c24b5c1865c3d80e9d025ef7e905b7dc01a9edc28e53180f1d
BLAKE2b-256 checksum
How to use checksums
cf52d2ab5a6395ee4d8a7c16d09a13f00fd654edbf3b819ec5721a43dc8a3056
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 17, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page