Pipecat realtime LLM service for the Boson Realtime API
Project description
Use Higgs Realtime with Pipecat
Build low-latency voice agents with Higgs Realtime and Pipecat.
The pipecat-boson package exposes Higgs Realtime as a Pipecat
speech-to-speech LLMService. It receives live audio or text, manages the
conversation, calls tools, and streams audio or text responses. A voice
pipeline does not need separate STT, LLM, and TTS services.
Prerequisites
- Python 3.11 or newer.
- A Boson API key.
- Access to the Higgs Realtime API.
- An existing Pipecat application with an audio transport.
Note: Keep the Boson API key on the server. Never embed it in a browser or mobile client.
Install the package
Install the package from PyPI:
uv add pipecat-boson
The equivalent pip command is pip install pipecat-boson.
The core service does not require WebRTC. Install the webrtc extra only when
you want to run the browser example or use Pipecat's WebRTC transport:
uv add "pipecat-boson[webrtc]"
To develop the package or run its included example:
git clone git@github.com:boson-ai/pipecat-boson.git
cd pipecat-boson
uv sync --extra dev
To use a local checkout from another uv project:
uv add --editable ../pipecat-boson
The package supports pipecat-ai>=1.4.0,<2.
Configure the connection
Set the API key, WebSocket endpoint, and model ID in your server environment:
export BOSON_API_KEY=bai-xxxx
export BOSON_REALTIME_URL=wss://api.boson.ai/v1/realtime/
export BOSON_REALTIME_MODEL=higgs-realtime
Create the realtime service:
import os
from pipecat_boson.realtime import BosonRealtimeLLMService
llm = BosonRealtimeLLMService(
url=os.environ["BOSON_REALTIME_URL"],
api_key=os.environ["BOSON_API_KEY"],
model=os.environ["BOSON_REALTIME_MODEL"],
voice="default",
instructions="You are a concise and helpful voice assistant.",
)
Add the service to a Pipecat pipeline
The following example assumes that transport is an existing Pipecat audio
transport:
from pipecat.pipeline.pipeline import Pipeline
from pipecat.pipeline.worker import PipelineParams, PipelineWorker
from pipecat.processors.aggregators.llm_context import LLMContext
from pipecat.processors.aggregators.llm_response_universal import (
LLMContextAggregatorPair,
)
from pipecat.workers.runner import WorkerRunner
async def run_bot(transport, llm):
context = LLMContext()
user_aggregator, assistant_aggregator = LLMContextAggregatorPair(
context,
realtime_service_mode=True,
)
pipeline = Pipeline(
[
transport.input(),
user_aggregator,
llm,
transport.output(),
assistant_aggregator,
]
)
worker = PipelineWorker(
pipeline,
params=PipelineParams(
enable_metrics=True,
),
)
runner = WorkerRunner()
await runner.add_workers(worker)
await runner.run()
realtime_service_mode=True lets the context aggregators follow the
server-driven turn lifecycle. Do not add separate STT or TTS services around
BosonRealtimeLLMService.
Call run_bot(transport, llm) from your application's async entry point.
Higgs Realtime responds after server VAD detects the end of a user turn. If the
assistant should speak first, queue an LLMRunFrame after the client is ready,
as demonstrated by the included browser example.
Run the browser example
From the repository checkout created above, copy the example environment file:
cp .env.example .env
Set BOSON_API_KEY, BOSON_REALTIME_URL, and BOSON_REALTIME_MODEL in
.env, then start the WebRTC example:
uv run --extra webrtc \
python examples/pipecat_boson_realtime_agent.py \
-t webrtc \
--host 127.0.0.1 \
--port 7860
Open http://localhost:7860 and connect your microphone.
If WebRTC ICE cannot reach the server, use the WebSocket transport:
uv run --extra webrtc \
python examples/pipecat_boson_realtime_agent.py \
-t websocket \
--host localhost \
--port 7860
Select WebSocket in the page before connecting. Both commands use the
webrtc extra because it also installs the Pipecat runner used by the browser
example.
Receive user transcripts
Set an input transcription model to receive finalized user transcripts as
Pipecat TranscriptionFrame objects:
llm = BosonRealtimeLLMService(
url=os.environ["BOSON_REALTIME_URL"],
api_key=os.environ["BOSON_API_KEY"],
model=os.environ["BOSON_REALTIME_MODEL"],
input_audio_transcription={
"model": "higgs-stt-3.1",
"language": "en",
},
)
Omitting input_audio_transcription, passing None, or passing a dictionary
without a non-empty model suppresses client-facing user transcript events.
Higgs Realtime still understands the audio and can respond.
Call Python functions
Declare an async Python function with typed arguments and return its result
through result_callback:
from pipecat.processors.aggregators.llm_context import LLMContext
from pipecat.services.llm_service import FunctionCallParams
async def get_weather(
params: FunctionCallParams,
location: str,
) -> None:
"""Get the current weather for a location.
Args:
location: City or place name.
"""
await params.result_callback(
{
"location": location,
"condition": "sunny",
"temperature_c": 22,
}
)
tools = [get_weather]
llm = BosonRealtimeLLMService(
url=os.environ["BOSON_REALTIME_URL"],
api_key=os.environ["BOSON_API_KEY"],
model=os.environ["BOSON_REALTIME_MODEL"],
instructions="Use get_weather when the user asks about weather.",
tools=tools,
)
context = LLMContext(tools=tools)
Pass the same tool list to the service and the context. The service advertises
and registers the handlers for the Higgs Realtime session, while
LLMContext keeps the tool definitions with the conversation state. After the
function completes, Higgs Realtime continues the response with its result.
Configure turn detection
Server VAD is enabled by default. It detects the end of the user's turn, creates a response, and interrupts an active response when the user starts speaking. Override its thresholds only when the default behavior does not fit the application:
turn_detection = {
"type": "server_vad",
"prefix_padding_ms": 300,
"silence_duration_ms": 500,
"threshold": 0.55,
}
Pass this dictionary as turn_detection=turn_detection when constructing the
service. For most voice agents, keep the default server VAD settings.
Higgs Realtime also supports OpenAI-compatible semantic VAD:
semantic_turn_detection = {
"type": "semantic_vad",
}
llm = BosonRealtimeLLMService(
url=os.environ["BOSON_REALTIME_URL"],
api_key=os.environ["BOSON_API_KEY"],
turn_detection=semantic_turn_detection,
)
Use text-only output
Pass output_modalities=["text"] when constructing the service. Text-only
sessions emit streamed LLMTextFrame objects and no audio frames.
The service supports exactly one session output modality: ["audio"] or
["text"]. Mixed output modalities and per-response modality overrides are not
supported.
Handle session events
Use Pipecat service event handlers to observe the Higgs Realtime session lifecycle:
def register_session_handlers(llm):
@llm.event_handler("on_session_created")
async def on_session_created(service, event):
print("Session:", event.session.id)
@llm.event_handler("on_session_terminated")
async def on_session_terminated(service, event_type, event):
print("Session terminated:", event_type)
Call register_session_handlers(llm) before starting WorkerRunner. The
integration reports terminal session events but does not close the Pipecat
transport automatically.
Keep on_session_created handlers fast. Session setup waits for this handler
to return.
Supported Higgs Realtime options
Connection options:
| Parameter | Default | Description |
|---|---|---|
url |
Required | Higgs Realtime WebSocket endpoint. |
api_key |
Required for the hosted API | Boson API key sent as a Bearer token. |
model |
"higgs-realtime" |
Realtime model ID sent when the session is configured. |
Optional session settings supported by Higgs Realtime:
| Parameter | Default | Description |
|---|---|---|
voice |
"default" |
Voice preset or voice ID used for audio output. |
instructions |
Helpful assistant prompt | System instructions used to initialize the conversation. |
output_modalities |
["audio"] |
Exactly ["audio"] or ["text"]. |
temperature |
0.7 |
Sampling temperature used for model responses. |
max_output_tokens |
"inf" |
Maximum response tokens. Numeric values are capped at 4096. |
tools |
Not set | Python functions or Pipecat-compatible tool definitions. |
tool_choice |
"auto" |
Tool selection behavior used when tools are available. |
turn_detection |
Server VAD | OpenAI-compatible server_vad or semantic_vad configuration. |
input_audio_transcription |
Not set | Transcription dictionary. A non-empty model enables client-facing user transcript events. |
input_audio_transcription_model |
"" |
Convenience option for the transcription model. |
input_audio_transcription_language |
None |
Convenience option for the transcription language. |
input_audio_noise_reduction |
Not set | OpenAI-compatible {"type": "near_field"} or {"type": "far_field"} input noise reduction setting. The corresponding type string is also accepted. |
truncation |
"auto" |
"auto" enables smart context summarization when the selected model publishes a context limit; "disabled" turns it off. |
This Pipecat integration sends and receives 24 kHz PCM audio.
Next steps
- Learn about Higgs Realtime.
- Read the Pipecat documentation.
License
BSD-2-Clause. See LICENSE.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file pipecat_boson-0.1.1.tar.gz.
File metadata
- Download URL: pipecat_boson-0.1.1.tar.gz
- Upload date:
- Size: 31.8 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: uv/0.7.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
97b395d09e5f328e5170bb35cbe2feadd66deff9ea1cead8e53b3fa06357695f
|
|
| MD5 |
9bcc7d33f82d4ca14a76041586b9d566
|
|
| BLAKE2b-256 |
74f26d7517573dae46fe65615b0538bf01561b675cdfd5d6aefdc2045165b0a0
|
File details
Details for the file pipecat_boson-0.1.1-py3-none-any.whl.
File metadata
- Download URL: pipecat_boson-0.1.1-py3-none-any.whl
- Upload date:
- Size: 18.5 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: uv/0.7.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
86ac2e0dd86bc9c8dd6eeca46daa93ff463a908fdc4e4e444958b86c015eeb42
|
|
| MD5 |
7885fcda4c0b1c9f82fe1e716da4cbed
|
|
| BLAKE2b-256 |
9e2be809fff920eaf0e62af1e78aecfa1088b0b62ca2cdd90ed5a717dfe822b8
|