Skip to main content

voqalize-avatar

The pipecat half of voqalize/avatar — a 2-D talking head for AI voice calls that renders in the browser, not in a video track.

The avatar is lip-synced to your TTS audio and state aware: it knows when it has been interrupted, when the user is talking versus idle, when a tool call started and stopped. This package is the half that decides. It reads your pipeline's frames, infers what the avatar should be doing, and pushes the result to the client as RTVI server-messages over the data channel you already have.

No video track, no per-minute avatar vendor, no second media path. The face is dependency-free JavaScript on the other end — @voqalize/avatar, about 75 KB gzipped plus the one rig you mount.

pip install voqalize-avatar

Drop it in

AvatarProcessor goes between your TTS service and the transport's output — the seat where it can see the audio that is about to be spoken, at generation speed.

from voqalize_avatar import AvatarProcessor

pipeline = Pipeline([
    transport.input(),
    stt,
    context_aggregator.user(),
    llm,
    tts,
    AvatarProcessor(),           # <-- here
    transport.output(),
    context_aggregator.assistant(),
])

No arguments, no binaries to install, no environment variables — that is the whole integration, and it needs no other application code. Pipecat's JavaScript client projects the standard lifecycle locally; this processor supplies the TTS-context-correlated viseme cues it cannot reconstruct, plus explicit intent you push yourself. Why this seat and not an observer, and why the frames go downstream: the module docstring in processor.py.

Mouth shapes

Lipsync is the headline feature and there is nothing to wire up. The processor starts its viseme engine on StartFrame, at the sample rate that frame declares, and drives it from the same karaoke frames pipecat already pushes for word-level captions.

The wheel carries its own aligner — avatarsync, our fork of Rhubarb Lip Sync, emitting the A–H+X mouth-shape alphabet the wire format is built on — along with the 56 MB acoustic model it needs. That is why the wheel is ~44 MB and platform-specific. No path, no environment variable, no separate artifact to ship into your image.

Two legs, and the server splices between them. The fast leg predicts the whole timeline from the sentence's text before any audio exists (~0.15 ms, on the event loop) so the mouth is already moving when the first sample plays. The accurate leg then recognises phones from the rendered PCM as it streams, off the loop, and overwrites the prediction from the point it has reached. The client never chooses: a cues message carries from_ms, and everything queued at or after it is discarded. The reasoning and the constants are in visemes.py, next to the numbers they explain.

platform wheel
Linux x86-64 / aarch64 manylinux_2_25 — RHEL 8+, Debian 10+, Ubuntu 18.04+
macOS arm64 macosx_11_0_arm64 — macOS 11+

That is the installer's view. The tags are derived from the compiled binary rather than declared, and .github/workflows/wheels.yml is the canonical statement of what gets built (RELEASING.md); if this table and a published wheel ever disagree, the wheel is right.

Intel macOS is absent for an upstream reason: pipecat-ai requires onnxruntime, which publishes no macOS x86-64 wheel, so nothing depending on pipecat installs there at all.

Anything else installs the sdist, which carries no binary. So does an explicit --no-binary. Both are fine, and an install with no aligner is an ordinary condition rather than a failure: AvatarProcessor catches, logs once, and runs the session state-channel only. The face still listens, thinks, claims the floor and yields it; its mouth does not move while it speaks. Worse, not broken. (The internal build_viseme_engine() fails fast instead, raising AvatarsyncUnavailableError — a caller who asked for an engine and silently did not get one has been lied to.)

A source checkout of the repo is found by walking up to native/avatarsync, so the tests and the demo run against a locally built library with no configuration either. voqalize-avatar info says which one was found and proves it answers.

Saying what the pipeline cannot infer

Some behavior needs application knowledge — a deliberate acknowledgement, a hand gesture, a tool call that should read as reviewing the screen rather than thinking. No amount of frame-watching infers those correctly, and a library that guessed would nod at the wrong moment.

Push an AvatarControlFrame from anywhere in your pipeline:

from voqalize_avatar import AvatarAction, AvatarControlFrame, AvatarMessage

await self.push_frame(AvatarControlFrame(message=AvatarMessage.action(AvatarAction.ACK_RECEIVE)))
await self.push_frame(AvatarControlFrame(message=AvatarMessage.action(AvatarAction.GESTURE_GREET)))

Or subclass AvatarStateMachine (from voqalize_avatar.state_machine) when your application's frames are simply its own spelling of something the library already models — an LLM running out of process, say, whose tool calls never appear as pipecat function-call frames:

from voqalize_avatar import AvatarProcessor
from voqalize_avatar.state_machine import AvatarStateMachine

class MyStateMachine(AvatarStateMachine):
    def on_frame(self, frame):
        if isinstance(frame, MyToolStartedFrame):
            return self.tool_started(frame.call_id)
        if isinstance(frame, MyToolResultFrame):
            return self.tool_finished(frame.call_id)
        return super().on_frame(frame)

class MyAvatarProcessor(AvatarProcessor):
    STATE_MACHINE = MyStateMachine

STATE_MACHINE is a class attribute rather than a constructor argument deliberately: the front door takes no arguments and must keep taking none, and a second door that costs a class statement is not one you walk through by accident. tool_started / tool_finished are public for exactly this — you inherit the call-id dedup and the parallel-call hold, so a turn with three tools settles on one THINKING instead of flickering.

These two seams are the whole extension surface. There is no per-tool state map: an application that knows its tool is searching says so in one AvatarControlFrame.

What this package will not do

It never decides what the agent says or when. Pipecat owns the facts, the server owns intent, the rig only renders — a heuristic here that guessed at call content would be a bug, not a feature. Binding for both halves: contract-wire.md and pipecat-lifecycle-protocol.md.

Compatibility

pipecat-ai>=1.4,<2, Python 3.12+. The floor is where FunctionCallsStartedFrame and UserTurnInferenceCompletedFrame exist; the test suite runs at the floor as well as at the resolved version, so "we support 1.4" is a claim something actually checks. Base pipecat only — no transport, STT or TTS extras, because this package sits in somebody else's pipeline and must not have an opinion about which services they chose.

License

MIT. LICENSE here is a copy of the repository's, kept beside the package because a wheel carries its own license file.

The wheel also carries the aligner and its acoustic model, whose upstream notices are native/avatarsync/UPSTREAM-LICENSE.md — all permissive (MIT, BSD, Boost), and they have to travel with the binary.

Metadata

Release files for voqalize-avatar 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for voqalize-avatar 0.3.0
File Size Uploaded
voqalize_avatar-0.3.0.tar.gz 471.4 kB Details

Built distributions (wheels)

Table of built distributions (wheels) for voqalize-avatar 0.3.0
File Interpreter ABI Platform
voqalize_avatar-0.3.0-py3-none-manylinux_2_25_x86_64.whl Python 3 none Linux glibc 2.25+ x86-64 Details
voqalize_avatar-0.3.0-py3-none-manylinux_2_25_aarch64.whl Python 3 none Linux glibc 2.25+ ARM64 Details
voqalize_avatar-0.3.0-py3-none-macosx_11_0_arm64.whl Python 3 none macOS 11.0+ ARM64 Details

Total release size: 141.1 MB

Release files / voqalize_avatar-0.3.0.tar.gz

Download URL voqalize_avatar-0.3.0.tar.gz
Size 471.4 kB
Tags Source
SHA-256 checksum
How to use checksums
55a1e5e1eaf3acd38916118f8047f4a8519cb0576edb40dcc95c50b515c34ad7
BLAKE2b-256 checksum
How to use checksums
81cbef9eae33eb65c810467e540727d77df733b201743effd071e00bced36706
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 27, 2026.

Transparency log

Release files / voqalize_avatar-0.3.0-py3-none-manylinux_2_25_x86_64.whl

Download URL voqalize_avatar-0.3.0-py3-none-manylinux_2_25_x86_64.whl
Size 47.1 MB
Tags Linux glibc 2.25+ x86-64 Python 3
SHA-256 checksum
How to use checksums
8b9a4eae3914a393e57fbcab706c08b345be9f4b9ea90cba22d2e16e2b3c56f0
BLAKE2b-256 checksum
How to use checksums
755ade5e6dfb6546287e1169ccf4e1907ac865bfd7925e81c6a94fa263ccb784
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 27, 2026.

Transparency log

Release files / voqalize_avatar-0.3.0-py3-none-manylinux_2_25_aarch64.whl

Download URL voqalize_avatar-0.3.0-py3-none-manylinux_2_25_aarch64.whl
Size 47.1 MB
Tags Linux glibc 2.25+ ARM64 Python 3
SHA-256 checksum
How to use checksums
9445cc7ec58c798d33cf2d0c5dc111030ae9adac6b7da6adba3bfd6ff521499f
BLAKE2b-256 checksum
How to use checksums
dc48af374d4aa5e85ca19821b6b34f01949ae7043ee7e60dea2bc26180bc4a30
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 27, 2026.

Transparency log

Release files / voqalize_avatar-0.3.0-py3-none-macosx_11_0_arm64.whl

Download URL voqalize_avatar-0.3.0-py3-none-macosx_11_0_arm64.whl
Size 46.4 MB
Tags Python 3 macOS 11.0+ ARM64
SHA-256 checksum
How to use checksums
4f04351cab6bc5257f69debc3f32b33a011653e135fa64f4b3502ee7c956e9ec
BLAKE2b-256 checksum
How to use checksums
b381b6014fc5678b6562eabf74dc2bae4f37da14c7f0bc8e6f95cc4d5d92cb51
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 27, 2026.

Transparency log

Release history Release notifications | RSS feed

0.4.0

4 release files

0.3.1

4 release files

This release

0.3.0 This release

4 release files

0.2.2

4 release files

0.2.1

4 release files

0.2.0

4 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page