Skip to main content

voqalize-avatar

The pipecat half of voqalize/avatar — a 2-D talking head for AI voice calls that renders in the browser, not in a video track.

The avatar is lip-synced to your TTS audio and state aware: it knows when it has been interrupted, when the user is talking versus idle, when a tool call started and stopped. This package is the half that decides. It reads your pipeline's frames, infers what the avatar should be doing, and pushes the result to the client as RTVI server-messages over the data channel you already have.

No video track, no per-minute avatar vendor, no second media path. The face is dependency-free JavaScript on the other end — @voqalize/avatar, about 75 KB gzipped plus the one rig you mount.

pip install voqalize-avatar

Drop it in

AvatarProcessor goes between your TTS service and the transport's output — the seat where it can see the audio that is about to be spoken, at generation speed.

from voqalize_avatar import AvatarProcessor

pipeline = Pipeline([
    transport.input(),
    stt,
    context_aggregator.user(),
    llm,
    tts,
    AvatarProcessor(),           # <-- here
    transport.output(),
    context_aggregator.assistant(),
])

No arguments, no binaries to install, no environment variables — that is the whole integration, and it needs no other application code. Pipecat's JavaScript client projects the standard lifecycle locally; this processor supplies the TTS-context-correlated viseme cues it cannot reconstruct, plus explicit intent you push yourself. Why this seat and not an observer, and why the frames go downstream: the module docstring in processor.py.

Mouth shapes

Lipsync is the headline feature and there is nothing to wire up. The processor starts its viseme engine on StartFrame, at the sample rate that frame declares, and drives it from the same karaoke frames pipecat already pushes for word-level captions.

The wheel carries its own aligner — avatarsync, our fork of Rhubarb Lip Sync, emitting the A–H+X mouth-shape alphabet the wire format is built on — along with the 56 MB acoustic model it needs. That is why the wheel is ~44 MB and platform-specific. No path, no environment variable, no separate artifact to ship into your image.

Two legs, and the server splices between them. The fast leg predicts the whole timeline from the sentence's text before any audio exists (~0.15 ms, on the event loop) so the mouth is already moving when the first sample plays. The accurate leg then recognises phones from the rendered PCM as it streams, off the loop, and overwrites the prediction from the point it has reached. The client never chooses: a cues message carries from_ms, and everything queued at or after it is discarded. The reasoning and the constants are in visemes.py, next to the numbers they explain.

platform wheel
Linux x86-64 / aarch64 manylinux_2_25 — RHEL 8+, Debian 10+, Ubuntu 18.04+
macOS arm64 macosx_11_0_arm64 — macOS 11+

That is the installer's view. The tags are derived from the compiled binary rather than declared, and .github/workflows/wheels.yml is the canonical statement of what gets built (RELEASING.md); if this table and a published wheel ever disagree, the wheel is right.

Intel macOS is absent for an upstream reason: pipecat-ai requires onnxruntime, which publishes no macOS x86-64 wheel, so nothing depending on pipecat installs there at all.

Anything else installs the sdist, which carries no binary. So does an explicit --no-binary. Both are fine, and an install with no aligner is an ordinary condition rather than a failure: AvatarProcessor catches, logs once, and runs the session state-channel only. The face still listens, thinks, claims the floor and yields it; its mouth does not move while it speaks. Worse, not broken. (The internal build_viseme_engine() fails fast instead, raising AvatarsyncUnavailableError — a caller who asked for an engine and silently did not get one has been lied to.)

A source checkout of the repo is found by walking up to native/avatarsync, so the tests and the demo run against a locally built library with no configuration either. voqalize-avatar info says which one was found and proves it answers.

Saying what the pipeline cannot infer

Some behavior needs application knowledge — a deliberate acknowledgement, a hand gesture, a tool call that should read as reviewing the screen rather than thinking. No amount of frame-watching infers those correctly, and a library that guessed would nod at the wrong moment.

Push an AvatarControlFrame from anywhere in your pipeline:

from voqalize_avatar import AvatarAction, AvatarControlFrame, AvatarMessage

await self.push_frame(AvatarControlFrame(message=AvatarMessage.action(AvatarAction.ACK_RECEIVE)))
await self.push_frame(AvatarControlFrame(message=AvatarMessage.action(AvatarAction.GESTURE_GREET)))

Or subclass AvatarStateMachine (from voqalize_avatar.state_machine) when your application's frames are simply its own spelling of something the library already models — an LLM running out of process, say, whose tool calls never appear as pipecat function-call frames:

from voqalize_avatar import AvatarProcessor
from voqalize_avatar.state_machine import AvatarStateMachine

class MyStateMachine(AvatarStateMachine):
    def on_frame(self, frame):
        if isinstance(frame, MyToolStartedFrame):
            return self.tool_started(frame.call_id)
        if isinstance(frame, MyToolResultFrame):
            return self.tool_finished(frame.call_id)
        return super().on_frame(frame)

class MyAvatarProcessor(AvatarProcessor):
    STATE_MACHINE = MyStateMachine

STATE_MACHINE is a class attribute rather than a constructor argument deliberately: the front door takes no arguments and must keep taking none, and a second door that costs a class statement is not one you walk through by accident. tool_started / tool_finished are public for exactly this — you inherit the call-id dedup and the parallel-call hold, so a turn with three tools settles on one THINKING instead of flickering.

These two seams are the whole extension surface. There is no per-tool state map: an application that knows its tool is searching says so in one AvatarControlFrame.

What this package will not do

It never decides what the agent says or when. Pipecat owns the facts, the server owns intent, the rig only renders — a heuristic here that guessed at call content would be a bug, not a feature. Binding for both halves: contract-wire.md and pipecat-lifecycle-protocol.md.

Compatibility

pipecat-ai>=1.4,<2, Python 3.12+. The floor is where FunctionCallsStartedFrame and UserTurnInferenceCompletedFrame exist; the test suite runs at the floor as well as at the resolved version, so "we support 1.4" is a claim something actually checks. Base pipecat only — no transport, STT or TTS extras, because this package sits in somebody else's pipeline and must not have an opinion about which services they chose.

License

MIT. LICENSE here is a copy of the repository's, kept beside the package because a wheel carries its own license file.

The wheel also carries the aligner and its acoustic model, whose upstream notices are native/avatarsync/UPSTREAM-LICENSE.md — all permissive (MIT, BSD, Boost), and they have to travel with the binary.

Metadata

Release files for voqalize-avatar 0.4.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for voqalize-avatar 0.4.0
File Size Uploaded
voqalize_avatar-0.4.0.tar.gz 480.6 kB Details

Built distributions (wheels)

Table of built distributions (wheels) for voqalize-avatar 0.4.0
File Interpreter ABI Platform
voqalize_avatar-0.4.0-py3-none-manylinux_2_25_x86_64.whl Python 3 none Linux glibc 2.25+ x86-64 Details
voqalize_avatar-0.4.0-py3-none-manylinux_2_25_aarch64.whl Python 3 none Linux glibc 2.25+ ARM64 Details
voqalize_avatar-0.4.0-py3-none-macosx_11_0_arm64.whl Python 3 none macOS 11.0+ ARM64 Details

Total release size: 141.1 MB

Release files / voqalize_avatar-0.4.0.tar.gz

Download URL voqalize_avatar-0.4.0.tar.gz
Size 480.6 kB
Tags Source
SHA-256 checksum
How to use checksums
df996151475ce44a7b2ea6b2311e751288a61b932b8dba65d5b014d6ef126142
BLAKE2b-256 checksum
How to use checksums
97aa833fa898c65a4cdbe8a12d8cf0f86048c7e77186c8506d4246402823347f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 19, 2026.

Transparency log

Release files / voqalize_avatar-0.4.0-py3-none-manylinux_2_25_x86_64.whl

Download URL voqalize_avatar-0.4.0-py3-none-manylinux_2_25_x86_64.whl
Size 47.1 MB
Tags Linux glibc 2.25+ x86-64 Python 3
SHA-256 checksum
How to use checksums
4be2ea85fe083279fe65080f2d2fce31dcaad1ea7d665e20df385d0add2038bc
BLAKE2b-256 checksum
How to use checksums
21a66cf5487188959243cbf6f6a04f93e47820ffc64be7f735f6fdc905db1776
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 19, 2026.

Transparency log

Release files / voqalize_avatar-0.4.0-py3-none-manylinux_2_25_aarch64.whl

Download URL voqalize_avatar-0.4.0-py3-none-manylinux_2_25_aarch64.whl
Size 47.1 MB
Tags Linux glibc 2.25+ ARM64 Python 3
SHA-256 checksum
How to use checksums
bf061d108629406d9dc2223d50a1c3a245622dfd6ebbd9881bfbab538fd15ae4
BLAKE2b-256 checksum
How to use checksums
05ef0beb0b142d1c09a25ad08a0982f8fbb7e44d61978851f33bf71fdd388445
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 19, 2026.

Transparency log

Release files / voqalize_avatar-0.4.0-py3-none-macosx_11_0_arm64.whl

Download URL voqalize_avatar-0.4.0-py3-none-macosx_11_0_arm64.whl
Size 46.5 MB
Tags Python 3 macOS 11.0+ ARM64
SHA-256 checksum
How to use checksums
f52475a40869ff6c3dbc240bfb4443fdecccd018fa49f130783390b5b1386a18
BLAKE2b-256 checksum
How to use checksums
d3e794905e33a2997355f44a3a553fa982eb94f796101059e0e7d561479c24c3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 19, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.4.0 This release

4 release files

0.3.1

4 release files

0.3.0

4 release files

0.2.2

4 release files

0.2.1

4 release files

0.2.0

4 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page