Make voice AI conversations feel live — Pipecat agents that listen back with natural "mhm"s and "yeah"s while the user is still talking.
Project description
A community integration for Pipecat
pipecat-backchannel
Your Pipecat agent listens back: quiet "mhm"s and "yeah"s while the user is still talking.
Why
Listen to two people on a phone call. While one tells a story, the other keeps feeding back small signals: "mhm", "yeah". Linguists call these backchannels. They mean keep going, I'm with you, and the speaker talks straight through them.
Talk to a voice agent and you get silence until you finish. Silence on a call reads as a dropped connection, so you hesitate, trail off, or ask "hello?".
This library adds the listening sounds. The agent keeps the floor with you the whole time: a backchannel takes no turn and leaves no trace in your LLM context.
Install
uv add pipecat-backchannel
Usage
from pipecat_backchannel import Backchannel
backchannel = Backchannel()
pipeline = Pipeline(backchannel([
transport.input(),
stt,
context_aggregator.user(),
llm,
tts,
transport.output(),
context_aggregator.assistant(),
]))
One wrapper call is the whole setup. The default configuration works without an API key, a voice ID, or a sample rate, and places every processor for you.
The first run records the clips through your own TTS, so they come out in
your bot's voice, and caches them in .clip_cache/. Every later run starts
from the cache. Delete that directory after changing voice.
Record before the first call
The first session waits a few seconds while the recorder fills the clip cache. Move that wait to app startup if you want to avoid delay for first caller:
backchannel = Backchannel() # one per process
async def run_bot(...):
pipeline = Pipeline(backchannel([... tts ...]))
if __name__ == "__main__":
asyncio.run(backchannel.prewarm(tts=make_tts())) # before the server opens
main()
Tuning
Backchannel(
volume=0.6, # how loud, relative to the recording
params=BackchannelParams(
fire_probability=0.8, # chance of speaking up at a good moment
cooldown_s=2.5, # minimum gap between backchannels
min_speech_before_eligible_s=0.7, # ignore pauses right after a turn starts
),
)
fire_probability and cooldown_s set how present the listener feels. Raise
the first and lower the second for a more active one. Keep some slack: a bot
that reacts to each eligible pause sounds mechanical, because human listeners
skip some of them.
volume sits below 1.0 by default so clips play under the user's voice.
Set loguru to DEBUG to see each fire/skip decision and its reason.
Going quiet mid-session
Some stretches of a call want a silent listener: the user dictating a card number, or the bot walking through a form where each pause belongs to a field, and an "mhm" would read as confirmation.
The component that decides when to fire is the BackchannelProcessor, one
per pipeline. backchannel([...]) returns your processor list with it
inserted, so keep that list, build the Pipeline from it, and pull the gate
out of it:
from pipecat_backchannel.processor import BackchannelProcessor
processors = backchannel([transport.input(), stt, llm, tts, transport.output()])
pipeline = Pipeline(processors)
gate = next(p for p in processors if isinstance(p, BackchannelProcessor))
Then flip it from any event handler, as many times as the call needs:
gate.enabled = False # stop firing; every frame still passes through
...
gate.enabled = True # resume
enabled is a plain attribute, safe to set at any moment. While it is
False the processor keeps passing audio through and keeps tracking speech,
so nothing downstream notices and re-enabling picks up mid-turn. It fires no
clips and runs no end-of-turn classification in that state. The pipeline, the
LLM context, and the other sessions' gates stay untouched.
Custom clips
The library groups clips by conversational function: continuer, agreement, hesitation, surprise.
Backchannel(clip_groups={"continue": ["Mhm.", "Right."], "affirm": ["Yeah.", "Yep."]})
Two rules for a custom inventory:
- Give each group at least two clips. A one-clip group repeats itself.
- Get variety from punctuation, and keep the vocabulary small. Real
backchannels draw on a handful of sounds: uh-huh, yeah, mm-hmm, right,
okay, oh, huh, hm. A written phrase like "absolutely" reads fine on the
page, then sounds like an answer when spoken over someone mid-sentence.
Punctuation is the handle your TTS gives you:
"Mhm."falls,"Mhm,"stays up,"Hm..."draws out.
Pass synthesizer= to use a different voice, or cache= to supply your own
recorded .wav files. See pipecat_backchannel.synth and .cache.
Run the example
cp examples/.env.template examples/.env # DEEPGRAM_API_KEY, CARTESIA_API_KEY, CEREBRAS_API_KEY
uv run examples/bot.py
Open http://localhost:7860/client, talk in mid-length sentences, and pause
mid-thought. A quiet "mhm" lands on clause-complete pauses. Word-search pauses
("we need the, uh...") and finished questions get silence, and the questions
get a real reply.
Limitations
- Pauses only. Humans backchannel during speech too, cued by pitch. This library reacts to pauses. The roadmap covers the rest.
- About 2× inference while the user speaks. It runs a second VAD and turn-detector on purpose: the pipeline's copies are tuned for turn-taking, and pause timing needs the opposite settings. A handful of concurrent sessions runs fine; measure before a large deployment.
- English only. The clips and the regexes are English. The underlying models handle other languages.
Roadmap
- Fire during speech. Real listeners backchannel mid-utterance, cued by pitch and rhythm rather than silence. That takes Voice Activity Projection; MaAI ships a usable model. It reads a different signal, so it slots in beside the pause gate rather than replacing it, and it is the largest quality jump left.
- Match the speaker's prosody. The selector picks a clip at random within its group. Choosing the one whose pitch and energy fit the last utterance would likely want a small prediction model at fire time, trained to map the tail of the user's speech to a clip. Playback stays free — the clips already sit on disk — so the only new cost is that one lightweight inference.
License
BSD 2-Clause © Taras Maister, matching Pipecat's own license. Not affiliated with Daily or the Pipecat core team.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file pipecat_backchannel-0.3.0.tar.gz.
File metadata
- Download URL: pipecat_backchannel-0.3.0.tar.gz
- Upload date:
- Size: 256.3 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
4169c0a905cf6d7a1e4c9d4ccd60bea3af2bf907207d9a1b1b51da97666422d5
|
|
| MD5 |
aa66bbcbe4f4e83b973ccefce3a19b79
|
|
| BLAKE2b-256 |
da09c705a88c1fb18386d066bb6db31b5eb0f56d1fa2945dbeab1c25b3ff20d1
|
Provenance
The following attestation bundles were made for pipecat_backchannel-0.3.0.tar.gz:
Publisher:
release.yml on maisterr/pipecat-backchannel
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
pipecat_backchannel-0.3.0.tar.gz -
Subject digest:
4169c0a905cf6d7a1e4c9d4ccd60bea3af2bf907207d9a1b1b51da97666422d5 - Sigstore transparency entry: 2253676630
- Sigstore integration time:
-
Permalink:
maisterr/pipecat-backchannel@7da9d1652b19c436bb5c104cb155d1d20056f6cf -
Branch / Tag:
refs/tags/v0.3.0 - Owner: https://github.com/maisterr
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@7da9d1652b19c436bb5c104cb155d1d20056f6cf -
Trigger Event:
release
-
Statement type:
File details
Details for the file pipecat_backchannel-0.3.0-py3-none-any.whl.
File metadata
- Download URL: pipecat_backchannel-0.3.0-py3-none-any.whl
- Upload date:
- Size: 35.9 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
e875f6623f018db873a59ebef377f62b48c7d24e2c55752711dc2c91915d5269
|
|
| MD5 |
521195795f8d5b4999f8331f57eb009c
|
|
| BLAKE2b-256 |
4c1cb31a0bede23fccb3c57f617ae41551dd4520f1f40999319f5b4df32eae7b
|
Provenance
The following attestation bundles were made for pipecat_backchannel-0.3.0-py3-none-any.whl:
Publisher:
release.yml on maisterr/pipecat-backchannel
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
pipecat_backchannel-0.3.0-py3-none-any.whl -
Subject digest:
e875f6623f018db873a59ebef377f62b48c7d24e2c55752711dc2c91915d5269 - Sigstore transparency entry: 2253676889
- Sigstore integration time:
-
Permalink:
maisterr/pipecat-backchannel@7da9d1652b19c436bb5c104cb155d1d20056f6cf -
Branch / Tag:
refs/tags/v0.3.0 - Owner: https://github.com/maisterr
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@7da9d1652b19c436bb5c104cb155d1d20056f6cf -
Trigger Event:
release
-
Statement type: