voqalize-avatar
The pipecat half of voqalize/avatar — a 2-D talking head for AI voice calls that renders in the browser, not in a video track.
The avatar is lip-synced to your TTS audio and state aware: it knows when it has been interrupted, when the user is talking versus idle, when a tool call started and stopped. This package is the half that decides. It reads your pipeline's frames, infers what the avatar should be doing, and pushes the result to the client as RTVI server-messages over the data channel you already have.
No video track, no per-minute avatar vendor, no second media path. The face is
dependency-free JavaScript on the other end —
@voqalize/avatar, about
75 KB gzipped plus the one rig you mount.
pip install voqalize-avatar
Drop it in
AvatarProcessor goes between your TTS service and the transport's output —
the seat where it can see the audio that is about to be spoken, at generation
speed.
from voqalize_avatar import AvatarProcessor
pipeline = Pipeline([
transport.input(),
stt,
context_aggregator.user(),
llm,
tts,
AvatarProcessor(), # <-- here
transport.output(),
context_aggregator.assistant(),
])
No arguments, no binaries to install, no environment variables — that is the
whole integration, and it needs no other application code. Pipecat's JavaScript
client projects the standard lifecycle locally; this processor supplies the
TTS-context-correlated viseme cues it cannot reconstruct, plus explicit intent
you push yourself. Why this seat and not an observer, and why the frames go
downstream: the module docstring in processor.py.
Mouth shapes
Lipsync is the headline feature and there is nothing to wire up. The processor
starts its viseme engine on StartFrame, at the sample rate that frame
declares, and drives it from the same karaoke frames pipecat already pushes for
word-level captions.
The wheel carries its own aligner —
avatarsync,
our fork of Rhubarb Lip Sync,
emitting the A–H+X mouth-shape alphabet the wire format is built on — along with
the 56 MB acoustic model it needs. That is why the wheel is ~44 MB and
platform-specific. No path, no environment variable, no separate artifact to
ship into your image.
Two legs, and the server splices between them. The fast leg predicts the
whole timeline from the sentence's text before any audio exists (~0.15 ms, on
the event loop) so the mouth is already moving when the first sample plays. The
accurate leg then recognises phones from the rendered PCM as it streams, off
the loop, and overwrites the prediction from the point it has reached. The
client never chooses: a cues message carries from_ms, and everything queued
at or after it is discarded. The reasoning and the constants are in
visemes.py, next to the numbers they explain.
| platform | wheel |
|---|---|
| Linux x86-64 / aarch64 | manylinux_2_25 — RHEL 8+, Debian 10+, Ubuntu 18.04+ |
| macOS arm64 | macosx_11_0_arm64 — macOS 11+ |
That is the installer's view. The tags are derived from the compiled binary
rather than declared, and .github/workflows/wheels.yml is the canonical
statement of what gets built
(RELEASING.md);
if this table
and a published wheel ever disagree, the wheel is right.
Intel macOS is absent for an upstream reason: pipecat-ai requires
onnxruntime, which publishes no macOS x86-64 wheel, so nothing depending on
pipecat installs there at all.
Anything else installs the sdist, which carries no binary. So does an explicit
--no-binary. Both are fine, and an install with no aligner is an ordinary
condition rather than a failure: AvatarProcessor catches, logs once, and
runs the session state-channel only. The face still listens, thinks, claims the
floor and yields it; its mouth does not move while it speaks. Worse, not broken.
(The internal build_viseme_engine() fails fast instead, raising
AvatarsyncUnavailableError — a caller who asked for an engine and silently did
not get one has been lied to.)
A source checkout of the repo is found by walking up to native/avatarsync, so
the tests and the demo run against a locally built library with no configuration
either. voqalize-avatar info says which one was found and proves it answers.
Saying what the pipeline cannot infer
Some behavior needs application knowledge — a deliberate acknowledgement, a hand gesture, a tool call that should read as reviewing the screen rather than thinking. No amount of frame-watching infers those correctly, and a library that guessed would nod at the wrong moment.
Push an AvatarControlFrame from anywhere in your pipeline:
from voqalize_avatar import AvatarAction, AvatarControlFrame, AvatarMessage
await self.push_frame(AvatarControlFrame(message=AvatarMessage.action(AvatarAction.ACK_RECEIVE)))
await self.push_frame(AvatarControlFrame(message=AvatarMessage.action(AvatarAction.GESTURE_GREET)))
Or subclass AvatarStateMachine (from voqalize_avatar.state_machine) when
your application's frames are simply its own spelling of something the library
already models — an LLM running out of process, say, whose tool calls never
appear as pipecat function-call frames:
from voqalize_avatar import AvatarProcessor
from voqalize_avatar.state_machine import AvatarStateMachine
class MyStateMachine(AvatarStateMachine):
def on_frame(self, frame):
if isinstance(frame, MyToolStartedFrame):
return self.tool_started(frame.call_id)
if isinstance(frame, MyToolResultFrame):
return self.tool_finished(frame.call_id)
return super().on_frame(frame)
class MyAvatarProcessor(AvatarProcessor):
STATE_MACHINE = MyStateMachine
STATE_MACHINE is a class attribute rather than a constructor argument
deliberately: the front door takes no arguments and must keep taking none, and a
second door that costs a class statement is not one you walk through by
accident. tool_started / tool_finished are public for exactly this — you
inherit the call-id dedup and the parallel-call hold, so a turn with three tools
settles on one THINKING instead of flickering.
These two seams are the whole extension surface. There is no per-tool state map:
an application that knows its tool is searching says so in one
AvatarControlFrame.
What this package will not do
It never decides what the agent says or when. Pipecat owns the facts, the server owns intent, the rig only renders — a heuristic here that guessed at call content would be a bug, not a feature. Binding for both halves: contract-wire.md and pipecat-lifecycle-protocol.md.
Compatibility
pipecat-ai>=1.4,<2, Python 3.12+. The floor is where
FunctionCallsStartedFrame and UserTurnInferenceCompletedFrame exist; the
test suite runs at the floor as well as at the resolved version, so "we support
1.4" is a claim something actually checks. Base pipecat only — no transport, STT
or TTS extras, because this package sits in somebody else's pipeline and must
not have an opinion about which services they chose.
License
MIT. LICENSE here is a copy of the repository's, kept beside the package
because a wheel carries its own license file.
The wheel also carries the aligner and its acoustic model, whose upstream
notices are native/avatarsync/UPSTREAM-LICENSE.md — all permissive (MIT, BSD,
Boost), and they have to travel with the binary.
Metadata
Release files for voqalize-avatar 0.3.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| voqalize_avatar-0.3.1.tar.gz | 473.1 kB | Details |
Built distributions (wheels)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| voqalize_avatar-0.3.1-py3-none-manylinux_2_25_x86_64.whl | Python 3 | none | Linux glibc 2.25+ x86-64 | Details |
| voqalize_avatar-0.3.1-py3-none-manylinux_2_25_aarch64.whl | Python 3 | none | Linux glibc 2.25+ ARM64 | Details |
| voqalize_avatar-0.3.1-py3-none-macosx_11_0_arm64.whl | Python 3 | none | macOS 11.0+ ARM64 | Details |
Total release size: 141.1 MB
Release files / voqalize_avatar-0.3.1.tar.gz
| Download URL | voqalize_avatar-0.3.1.tar.gz |
|---|---|
| Size | 473.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
01c9b01ff03f52b3bb4917a1a05ff2873e3455000eb2a8fa77d12a2fd7fa7bd2
|
|
BLAKE2b-256 checksum How to use checksums |
1ac76915982457b784e19b1e0c29b94b659a7bee4bfc9d8be70bd33b3f9c688a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 11, 2026.
Transparency logRelease files / voqalize_avatar-0.3.1-py3-none-manylinux_2_25_x86_64.whl
| Download URL | voqalize_avatar-0.3.1-py3-none-manylinux_2_25_x86_64.whl |
|---|---|
| Size | 47.1 MB |
| Tags | Linux glibc 2.25+ x86-64 Python 3 |
|
SHA-256 checksum How to use checksums |
131ce5a77a2945d7d17e878794ba1eda8f0cccf3412c634ffab282efbbbebeb8
|
|
BLAKE2b-256 checksum How to use checksums |
e120473006914e2b1554e88c867c60bb9db77fc1efd25e2928a5d6f22d8810a7
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 11, 2026.
Transparency logRelease files / voqalize_avatar-0.3.1-py3-none-manylinux_2_25_aarch64.whl
| Download URL | voqalize_avatar-0.3.1-py3-none-manylinux_2_25_aarch64.whl |
|---|---|
| Size | 47.1 MB |
| Tags | Linux glibc 2.25+ ARM64 Python 3 |
|
SHA-256 checksum How to use checksums |
0f87a6acd572cda45201a6c83a294346ea7a399fc3627f4324a9f092a0f2c0c8
|
|
BLAKE2b-256 checksum How to use checksums |
d70a58145c972e8d6ce8f25adde43d393dba0dd0a5c6f226817d93dceb2078ca
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 11, 2026.
Transparency logRelease files / voqalize_avatar-0.3.1-py3-none-macosx_11_0_arm64.whl
| Download URL | voqalize_avatar-0.3.1-py3-none-macosx_11_0_arm64.whl |
|---|---|
| Size | 46.4 MB |
| Tags | Python 3 macOS 11.0+ ARM64 |
|
SHA-256 checksum How to use checksums |
9e2beed1dc8d0cb8652d482597bcb359fb1a9800bbab51e1bb5ddc3ee098cb7c
|
|
BLAKE2b-256 checksum How to use checksums |
e1005047d1704abbf495d9afaf030798614c5009fadd0dec5154545f5f58924b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 11, 2026.
Transparency log