Skip to main content

voqalize-avatar

The pipecat half of voqalize/avatar — a 2-D talking head for AI voice calls that renders in the browser, not in a video track.

The widget is a state machine wearing a face: it renders a state enum, an emotion, a gaze target, interjection and hand-gesture ids, and a stream of timed viseme letters, and it decides none of them. This package is the half that decides. It reads your pipeline's frames, infers what the avatar should be doing, and pushes the result to the client as RTVI server-messages over the data channel you already have.

There is no video track, no per-minute avatar vendor, and no second media path. The face is dependency-free JavaScript on the other end — about 75 KB gzipped for the widget plus the one rig you mount.

pip install voqalize-avatar

The browser half is @voqalize/avatar.

Drop it in

AvatarProcessor goes between your TTS service and the transport's output — the seat where it can see the audio that is about to be spoken.

from voqalize_avatar import AvatarProcessor

pipeline = Pipeline([
    transport.input(),
    stt,
    context_aggregator.user(),
    llm,
    tts,
    AvatarProcessor(),           # <-- here
    transport.output(),
    context_aggregator.assistant(),
])

No arguments, no binaries to install, no environment variables — that is the whole integration, and it needs no other application code. From stock pipecat frames the state machine delivers IDLE, LISTENING, THINKING, SPEAKING, TAKING_FLOOR, WAITING_FOR_USER, YIELDED, DEGRADED and OFFLINE, plus the turn-clock anchor the client splices cues onto and the user-speaking truth the listening engine times backchannels off.

Mouth shapes

Lipsync is the headline feature, and there is nothing to wire up: the processor starts its viseme engine when StartFrame arrives, at the sample rate that frame declares, and drives it from the same karaoke frames pipecat already pushes for word-level captions.

The wheel carries its own aligner — avatarsync, our fork of Rhubarb Lip Sync, which emits the A–H+X mouth-shape alphabet the wire format is built on — along with the 56 MB acoustic model it needs. That is why the wheel is ~44 MB and why it is platform-specific. No path, no environment variable, no separate artifact to ship into your image.

platform wheel
Linux x86-64 / aarch64 manylinux_2_25 — RHEL 8+, Debian 10+, Ubuntu 18.04+
macOS arm64 macosx_11_0_arm64 — macOS 11+

Intel macOS is not on that list, and the reason is upstream: pipecat-ai requires onnxruntime, which publishes no macOS x86-64 wheel, so nothing that depends on pipecat installs there at all.

Anything else installs the sdist, which carries no binary. So does an explicit --no-binary. Both are fine, and the two APIs differ here on purpose. The internal one, build_viseme_engine(), is a library call and fails fast: it raises RhubarbUnavailableError naming the paths it looked in. AvatarProcessor is the layer that decides a missing aligner is survivable — it catches, logs once, and runs the session state-channel only. The face still listens, thinks, claims the floor and yields it; its mouth does not move while it speaks. Worse, not broken.

A source checkout of this repo is found by walking up to native/avatarsync, so the tests and the demo run against a locally built binary with no configuration either.

The engine runs three legs and the client splices between them: a fast leg that predicts the timeline from text before the audio exists (~0.4 ms), an accurate leg that recognises phones from the rendered PCM (~15 ms), and an early-prefix leg for the first sentence, where latency is most visible.

Saying what the pipeline cannot infer

Some states need to know what your application is doing — TYPING, SEARCHING_SCREEN, CANT_HEAR, a deliberate interjection, a hand gesture. No amount of frame-watching infers those correctly, and a library that guessed would nod at the wrong moment. Two seams, in order of reach for:

Push an AvatarControlFrame from anywhere in your pipeline:

from voqalize_avatar import AvatarControlFrame, AvatarMessage, Interjection

await self.push_frame(AvatarControlFrame(message=AvatarMessage.interject(Interjection.MM_HMM)))

Hand gestures ride the same seam and are never inferred — a hand in frame is an application's decision:

from voqalize_avatar import HandGesture

await self.push_frame(AvatarControlFrame(message=AvatarMessage.gesture(HandGesture.HI)))

Or subclass AvatarStateMachine (from voqalize_avatar.state_machine) when your application's frames are simply its own spelling of something the library already models — an LLM that runs out of process, say, whose tool calls never appear as pipecat function-call frames:

from voqalize_avatar import AvatarProcessor
from voqalize_avatar.state_machine import AvatarStateMachine

class MyStateMachine(AvatarStateMachine):
    def on_frame(self, frame):
        if isinstance(frame, MyToolStartedFrame):
            return self.tool_started(frame.call_id)
        if isinstance(frame, MyToolResultFrame):
            return self.tool_finished(frame.call_id)
        return super().on_frame(frame)

class MyAvatarProcessor(AvatarProcessor):
    STATE_MACHINE = MyStateMachine

STATE_MACHINE is a class attribute rather than a constructor argument deliberately: the front door takes no arguments and must keep taking none, and a second door that costs a class statement is not one you walk through by accident.

tool_started / tool_finished are public for exactly this: you inherit the call-id dedup and the parallel-call hold — a turn with three tools settles on one THINKING instead of flickering — rather than re-implementing them approximately.

A tool call shows THINKING, and only that. A tool_states={"search_web": ...} map existed and was removed in 0.2 — an application that knows its tool is searching says so in one AvatarControlFrame. See docs/removed.md.

What this package will not do

It never decides what the agent says or when. The server is the source of truth and the client only looks right while rendering it; a heuristic here that guessed at call content would be a bug, not a feature. See docs/contract-protocol.md, which is binding for both halves.

Compatibility

pipecat-ai>=1.4,<2, Python 3.12+. The floor is where FunctionCallsStartedFrame and UserTurnInferenceCompletedFrame exist; the test suite runs at the floor as well as at the resolved version, so "we support 1.4" is a claim something actually checks. Base pipecat only — no transport, STT or TTS extras, because this package sits in somebody else's pipeline and must not have an opinion about which services they chose.

License

AGPL-3.0-only. LICENSE here is a copy of the repository's, kept beside the package because a wheel carries its own license file. Commercial licensing: open an issue.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

voqalize_avatar-0.2.0.tar.gz (371.2 kB view details)

Uploaded Source

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

voqalize_avatar-0.2.0-py3-none-manylinux_2_25_x86_64.whl (47.2 MB view details)

Uploaded Python 3manylinux: glibc 2.25+ x86-64

voqalize_avatar-0.2.0-py3-none-manylinux_2_25_aarch64.whl (47.2 MB view details)

Uploaded Python 3manylinux: glibc 2.25+ ARM64

voqalize_avatar-0.2.0-py3-none-macosx_11_0_arm64.whl (46.6 MB view details)

Uploaded Python 3macOS 11.0+ ARM64

File details

Details for the file voqalize_avatar-0.2.0.tar.gz.

File metadata

  • Download URL: voqalize_avatar-0.2.0.tar.gz
  • Upload date:
  • Size: 371.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for voqalize_avatar-0.2.0.tar.gz
Algorithm Hash digest
SHA256 dab0359964dc5cad6768e43e6bb1be1af040661dff613aad7613c9d97bdd01c1
MD5 6826dc73cb8e4753066933ca2552a2f8
BLAKE2b-256 c266a7a9ccb8c2243571df14977b74d12f7e751df9aa1148fd7697c587c6d56a

See more details on using hashes here.

Provenance

The following attestation bundles were made for voqalize_avatar-0.2.0.tar.gz:

Publisher: release.yml on voqalize/avatar

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file voqalize_avatar-0.2.0-py3-none-manylinux_2_25_x86_64.whl.

File metadata

File hashes

Hashes for voqalize_avatar-0.2.0-py3-none-manylinux_2_25_x86_64.whl
Algorithm Hash digest
SHA256 42210a223a7335b3f2b37b22832a0da5aa689f618885d1154c1dc56a3c5b635d
MD5 5cd31fe7974ef09293390c28206f0412
BLAKE2b-256 98c987a275058d1888b9e68f998740468ea59a452b330f1af48f8e2968886d8d

See more details on using hashes here.

Provenance

The following attestation bundles were made for voqalize_avatar-0.2.0-py3-none-manylinux_2_25_x86_64.whl:

Publisher: release.yml on voqalize/avatar

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file voqalize_avatar-0.2.0-py3-none-manylinux_2_25_aarch64.whl.

File metadata

File hashes

Hashes for voqalize_avatar-0.2.0-py3-none-manylinux_2_25_aarch64.whl
Algorithm Hash digest
SHA256 48b59bbd479458603d7bf01941e2627e1735f9eaf5bf1e74aad7f485d5070589
MD5 0dfab6edce57023198e45569989f7c4f
BLAKE2b-256 e832990c28f93205454de9fd37bc89b6629941a55e294536e295b11a90d01b81

See more details on using hashes here.

Provenance

The following attestation bundles were made for voqalize_avatar-0.2.0-py3-none-manylinux_2_25_aarch64.whl:

Publisher: release.yml on voqalize/avatar

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file voqalize_avatar-0.2.0-py3-none-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for voqalize_avatar-0.2.0-py3-none-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 a1d7cfbb4725581f284a822eb1c5374606b8507f85842ffb0150b1db9b1a5a76
MD5 e3e129f88a7d3dff1c4fb34d763c07c8
BLAKE2b-256 808ca2746aca6459f3a384d7e7d659fd6f437928346b9051cf4e02dcce553dce

See more details on using hashes here.

Provenance

The following attestation bundles were made for voqalize_avatar-0.2.0-py3-none-macosx_11_0_arm64.whl:

Publisher: release.yml on voqalize/avatar

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page