Skip to main content

Pipecat realtime LLM service for the Boson Realtime API

Project description

Use Higgs Realtime with Pipecat

Build low-latency voice agents with Higgs Realtime and Pipecat.

The pipecat-boson package exposes Higgs Realtime as a Pipecat speech-to-speech LLMService. It receives live audio or text, manages the conversation, calls tools, and streams audio or text responses. A voice pipeline does not need separate STT, LLM, and TTS services.

Prerequisites

  • Python 3.11 or newer.
  • A Boson API key.
  • Access to the Higgs Realtime API.
  • An existing Pipecat application with an audio transport.

Note: Keep the Boson API key on the server. Never embed it in a browser or mobile client.

Install the package

Install the package from PyPI:

uv add pipecat-boson

The equivalent pip command is pip install pipecat-boson.

The core service does not require WebRTC. Install the webrtc extra only when you want to run the browser example or use Pipecat's WebRTC transport:

uv add "pipecat-boson[webrtc]"

To develop the package or run its included example:

git clone git@github.com:boson-ai/pipecat-boson.git
cd pipecat-boson
uv sync --extra dev

To use a local checkout from another uv project:

uv add --editable ../pipecat-boson

The package supports pipecat-ai>=1.4.0,<2.

Configure the connection

Set the API key, WebSocket endpoint, and model ID in your server environment:

export BOSON_API_KEY=bai-xxxx
export BOSON_REALTIME_URL=wss://api.boson.ai/v1/realtime/
export BOSON_REALTIME_MODEL=higgs-realtime

Create the realtime service:

import os

from pipecat_boson.realtime import BosonRealtimeLLMService

llm = BosonRealtimeLLMService(
    url=os.environ["BOSON_REALTIME_URL"],
    api_key=os.environ["BOSON_API_KEY"],
    model=os.environ["BOSON_REALTIME_MODEL"],
    voice="default",
    instructions="You are a concise and helpful voice assistant.",
)

Add the service to a Pipecat pipeline

The following example assumes that transport is an existing Pipecat audio transport:

from pipecat.pipeline.pipeline import Pipeline
from pipecat.pipeline.worker import PipelineParams, PipelineWorker
from pipecat.processors.aggregators.llm_context import LLMContext
from pipecat.processors.aggregators.llm_response_universal import (
    LLMContextAggregatorPair,
)
from pipecat.workers.runner import WorkerRunner


async def run_bot(transport, llm):
    context = LLMContext()
    user_aggregator, assistant_aggregator = LLMContextAggregatorPair(
        context,
        realtime_service_mode=True,
    )

    pipeline = Pipeline(
        [
            transport.input(),
            user_aggregator,
            llm,
            transport.output(),
            assistant_aggregator,
        ]
    )

    worker = PipelineWorker(
        pipeline,
        params=PipelineParams(
            enable_metrics=True,
        ),
    )

    runner = WorkerRunner()
    await runner.add_workers(worker)
    await runner.run()

realtime_service_mode=True lets the context aggregators follow the server-driven turn lifecycle. Do not add separate STT or TTS services around BosonRealtimeLLMService.

Call run_bot(transport, llm) from your application's async entry point.

Higgs Realtime responds after server VAD detects the end of a user turn. If the assistant should speak first, queue an LLMRunFrame after the client is ready, as demonstrated by the included browser example.

Run the browser example

From the repository checkout created above, copy the example environment file:

cp .env.example .env

Set BOSON_API_KEY, BOSON_REALTIME_URL, and BOSON_REALTIME_MODEL in .env, then start the WebRTC example:

uv run --extra webrtc \
  python examples/pipecat_boson_realtime_agent.py \
    -t webrtc \
    --host 127.0.0.1 \
    --port 7860

Open http://localhost:7860 and connect your microphone.

If WebRTC ICE cannot reach the server, use the WebSocket transport:

uv run --extra webrtc \
  python examples/pipecat_boson_realtime_agent.py \
    -t websocket \
    --host localhost \
    --port 7860

Select WebSocket in the page before connecting. Both commands use the webrtc extra because it also installs the Pipecat runner used by the browser example.

Receive user transcripts

Set an input transcription model to receive finalized user transcripts as Pipecat TranscriptionFrame objects:

llm = BosonRealtimeLLMService(
    url=os.environ["BOSON_REALTIME_URL"],
    api_key=os.environ["BOSON_API_KEY"],
    model=os.environ["BOSON_REALTIME_MODEL"],
    input_audio_transcription={
        "model": "higgs-stt-3.1",
        "language": "en",
    },
)

Omitting input_audio_transcription, passing None, or passing a dictionary without a non-empty model suppresses client-facing user transcript events. Higgs Realtime still understands the audio and can respond.

Call Python functions

Declare an async Python function with typed arguments and return its result through result_callback:

from pipecat.processors.aggregators.llm_context import LLMContext
from pipecat.services.llm_service import FunctionCallParams


async def get_weather(
    params: FunctionCallParams,
    location: str,
) -> None:
    """Get the current weather for a location.

    Args:
        location: City or place name.
    """
    await params.result_callback(
        {
            "location": location,
            "condition": "sunny",
            "temperature_c": 22,
        }
    )


tools = [get_weather]

llm = BosonRealtimeLLMService(
    url=os.environ["BOSON_REALTIME_URL"],
    api_key=os.environ["BOSON_API_KEY"],
    model=os.environ["BOSON_REALTIME_MODEL"],
    instructions="Use get_weather when the user asks about weather.",
    tools=tools,
)

context = LLMContext(tools=tools)

Pass the same tool list to the service and the context. The service advertises and registers the handlers for the Higgs Realtime session, while LLMContext keeps the tool definitions with the conversation state. After the function completes, Higgs Realtime continues the response with its result.

Configure turn detection

Server VAD is enabled by default. It detects the end of the user's turn, creates a response, and interrupts an active response when the user starts speaking. Override its thresholds only when the default behavior does not fit the application:

turn_detection = {
    "type": "server_vad",
    "prefix_padding_ms": 300,
    "silence_duration_ms": 500,
    "threshold": 0.55,
}

Pass this dictionary as turn_detection=turn_detection when constructing the service. For most voice agents, keep the default server VAD settings.

Higgs Realtime also supports OpenAI-compatible semantic VAD:

semantic_turn_detection = {
    "type": "semantic_vad",
}

llm = BosonRealtimeLLMService(
    url=os.environ["BOSON_REALTIME_URL"],
    api_key=os.environ["BOSON_API_KEY"],
    turn_detection=semantic_turn_detection,
)

Use text-only output

Pass output_modalities=["text"] when constructing the service. Text-only sessions emit streamed LLMTextFrame objects and no audio frames.

The service supports exactly one session output modality: ["audio"] or ["text"]. Mixed output modalities and per-response modality overrides are not supported.

Handle session events

Use Pipecat service event handlers to observe the Higgs Realtime session lifecycle:

def register_session_handlers(llm):
    @llm.event_handler("on_session_created")
    async def on_session_created(service, event):
        print("Session:", event.session.id)

    @llm.event_handler("on_session_terminated")
    async def on_session_terminated(service, event_type, event):
        print("Session terminated:", event_type)

Call register_session_handlers(llm) before starting WorkerRunner. The integration reports terminal session events but does not close the Pipecat transport automatically.

Keep on_session_created handlers fast. Session setup waits for this handler to return.

Supported Higgs Realtime options

Connection options:

Parameter Default Description
url Required Higgs Realtime WebSocket endpoint.
api_key Required for the hosted API Boson API key sent as a Bearer token.
model "higgs-realtime" Realtime model ID sent when the session is configured.

Optional session settings supported by Higgs Realtime:

Parameter Default Description
voice "default" Voice preset or voice ID used for audio output.
instructions Helpful assistant prompt System instructions used to initialize the conversation.
output_modalities ["audio"] Exactly ["audio"] or ["text"].
temperature 0.7 Sampling temperature used for model responses.
max_output_tokens "inf" Maximum response tokens. Numeric values are capped at 4096.
tools Not set Python functions or Pipecat-compatible tool definitions.
tool_choice "auto" Tool selection behavior used when tools are available.
turn_detection Server VAD OpenAI-compatible server_vad or semantic_vad configuration.
input_audio_transcription Not set Transcription dictionary. A non-empty model enables client-facing user transcript events.
input_audio_transcription_model "" Convenience option for the transcription model.
input_audio_transcription_language None Convenience option for the transcription language.
input_audio_noise_reduction Not set OpenAI-compatible {"type": "near_field"} or {"type": "far_field"} input noise reduction setting. The corresponding type string is also accepted.
truncation "auto" "auto" enables smart context summarization when the selected model publishes a context limit; "disabled" turns it off.

This Pipecat integration sends and receives 24 kHz PCM audio.

Next steps

License

BSD-2-Clause. See LICENSE.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

pipecat_boson-0.1.3.tar.gz (32.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

pipecat_boson-0.1.3-py3-none-any.whl (18.6 kB view details)

Uploaded Python 3

File details

Details for the file pipecat_boson-0.1.3.tar.gz.

File metadata

  • Download URL: pipecat_boson-0.1.3.tar.gz
  • Upload date:
  • Size: 32.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.7.14

File hashes

Hashes for pipecat_boson-0.1.3.tar.gz
Algorithm Hash digest
SHA256 e214f61cbf5eb72b6e2d6a47d9c68b93ff324200e93491f301fb3fe23fbf58f8
MD5 835eee7465134551c021de4f0c91ee37
BLAKE2b-256 86ac41d5d4d11dc590d4dd9652953a672ab29ae9f6e98925c97ef83d2c580523

See more details on using hashes here.

File details

Details for the file pipecat_boson-0.1.3-py3-none-any.whl.

File metadata

File hashes

Hashes for pipecat_boson-0.1.3-py3-none-any.whl
Algorithm Hash digest
SHA256 1b6548ecf92dcbaec1f0cb678013cb36d9c833969c1750a7c1d1284670d17cef
MD5 afcbffd6bf117bad7aac5809cc81d0dd
BLAKE2b-256 123c2c1c7a59a6d0a1715488937e9c1d1375e969d009d1ed0266d117fbec76f0

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page