Skip to main content

Gemini Live Speech-to-Speech Plugin

Google Gemini Live Speech-to-Speech (STS) plugin for GetStream. It connects a realtime Gemini Live session to a Stream video call so your assistant can speak and listen in the same call.

Installation

uv add "vision-agents[gemini]"
# or directly
uv add vision-agents-plugins-gemini

Requirements

  • Python: 3.10+
  • Dependencies: getstream[webrtc"], getstream-plugins-common, google-genai>=1.51.0
  • API key: GOOGLE_API_KEY or GEMINI_API_KEY set in your environment

Quick Start

Below is a minimal example that attaches the Gemini Live output audio track to a Stream call and streams microphone audio into Gemini. The assistant will speak back into the call, and you can also send text messages to the assistant.

from dotenv import load_dotenv
from vision_agents.core import Agent, Runner, User
from vision_agents.core.agents import AgentLauncher
from vision_agents.plugins import gemini, getstream

load_dotenv()


async def create_agent(**kwargs) -> Agent:
    agent = Agent(
        edge=getstream.Edge(),
        agent_user=User(name="AI coach"),
        instructions="Read @coaching.md",
        llm=gemini.Realtime(model="gemini-3.1-flash-live-preview"),
        processors=[],
    )
    return agent


async def join_call(agent: Agent, call_type: str, call_id: str, **kwargs) -> None:
    call = await agent.create_call(call_type, call_id)

    async with agent.join(call):
        await agent.simple_response(
            text="Say hi. After the user joins ask them about their day"
        )
        await agent.finish()


if __name__ == "__main__":
    Runner(AgentLauncher(create_agent=create_agent, join_call=join_call)).cli()

Video frames from remote participants are forwarded to Gemini automatically when fps is set and the model supports it:

llm=gemini.Realtime(fps=3)  # forward video at 3 frames per second

The Agent subscribes to track events internally, so no manual wiring is needed. For a full runnable example, see examples/02_golf_coach_example/golf_coach_example.py.

Gemini Vision (VLM)

Use Gemini 3 vision models with the Agent API (video frames are forwarded automatically when the call has active video).

from vision_agents.core import Agent, Runner, User
from vision_agents.core.agents import AgentLauncher
from vision_agents.plugins import deepgram, elevenlabs, gemini, getstream


async def create_agent(**kwargs) -> Agent:
    vlm = gemini.VLM(model="gemini-3-flash-preview")
    return Agent(
        edge=getstream.Edge(),
        agent_user=User(name="Gemini Vision Agent", id="gemini-vision-agent"),
        instructions="Describe what you see in one sentence.",
        llm=vlm,
        stt=deepgram.STT(),
        tts=elevenlabs.TTS(),
    )


async def join_call(agent: Agent, call_type: str, call_id: str, **kwargs) -> None:
    call = await agent.create_call(call_type, call_id)
    async with agent.join(call):
        await agent.finish()


Runner(AgentLauncher(create_agent=create_agent, join_call=join_call)).cli()

Key configuration knobs for GeminiVLM: fps, frame_buffer_seconds, thinking_level, media_resolution. For a full example, see plugins/gemini/example/gemini_vlm_agent_example.py.

Features

  • Bidirectional audio: Streams microphone PCM to Gemini, and plays Gemini speech into the call using output_track.
  • Video frame forwarding: Sends remote participant video frames to Gemini Live for multimodal understanding. Use start_video_sender with a remote MediaStreamTrack.
  • Text messages: Use send_text to add text turns directly to the conversation.
  • **Barge-in (interruptions) **: When the user starts speaking, current playback is interrupted so Gemini can focus on the new input. Playback automatically resumes after brief silence.
  • Auto resampling: send_audio_pcm will resample input frames to the target rate when needed.
  • Events: Subscribe to "audio" for synthesized audio chunks and "text" for assistant text.

API Overview

  • GeminiLive(api_key: str | None = None, model: str = "gemini-live-2.5-flash-preview", config: LiveConnectConfigDict | None = None): Create a new Gemini Live session. If api_key is not provided, the plugin reads GOOGLE_API_KEY or GEMINI_API_KEY from the environment.
  • **GeminiVLM(model: str = "gemini-3-flash-preview", fps: int = 1, frame_buffer_seconds: int = 10, ...) **: Vision-language model that buffers video frames and sends them with prompts.
  • output_track: An AudioStreamTrack you can publish in your call via add_tracks(audio=...).
  • await send_text(text: str): Send a user text message to the current turn.
  • await send_audio_pcm(pcm: PcmData, target_rate: int = 48000): Stream PCM frames to Gemini. Frames are converted to the required format and resampled if necessary.
  • await wait_until_ready(timeout: float | None = None) -> bool: Wait until the underlying live session is connected.
  • await interrupt_playback() / resume_playback(): Manually stop or resume synthesized audio playback. Useful if you want to manage barge-in behavior yourself.
  • await start_video_sender(track: MediaStreamTrack, fps: int = 1): Start forwarding video frames from a remote MediaStreamTrack to Gemini Live at the given frame rate.
  • await stop_video_sender(): Stop the background video sender task, if running.
  • await close(): Close the session and background tasks.

Environment Variables

  • GOOGLE_API_KEY / GEMINI_API_KEY: Gemini API key. One must be set.
  • GEMINI_LIVE_MODEL: Optional override for the model name if you need a different variant.

Troubleshooting

  • No audio playback: Ensure you publish output_track to your call and the call is subscribed to the assistant's audio.
  • No responses: Verify GOOGLE_API_KEY/GEMINI_API_KEY is set and has access to the chosen model. Try a different model via model=.
  • Sample-rate issues: Use send_audio_pcm(..., target_rate=48000) to normalize input frames.

Migration from Gemini 2.5

When migrating to Gemini 3:

  • Thinking: If you were using complex prompt engineering (like Chain-of-thought) with Gemini 2.5, try Gemini 3 with thinking_level="high" and simplified prompts.
  • Temperature: If your code explicitly sets temperature to low values, consider removing it and using the Gemini 3 default (1.0) to avoid potential looping issues.
  • PDF & Document Understanding: Default OCR resolution for PDFs has changed. Test with media_resolution="high" if you need dense document parsing.
  • Token Consumption: Gemini 3 defaults may increase token usage for PDFs but decrease for video. If requests exceed context limits, explicitly reduce media_resolution.

Metadata

Release files for vision-agents-plugins-gemini 0.6.9

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for vision-agents-plugins-gemini 0.6.9
File Size Uploaded
vision_agents_plugins_gemini-0.6.9.tar.gz 23.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for vision-agents-plugins-gemini 0.6.9
File Interpreter ABI Platform
vision_agents_plugins_gemini-0.6.9-py3-none-any.whl Python 3 none any Details

Total release size: 51.0 kB

Release files / vision_agents_plugins_gemini-0.6.9.tar.gz

Download URL vision_agents_plugins_gemini-0.6.9.tar.gz
Size 23.5 kB
Tags Source
SHA-256 checksum
How to use checksums
cd9442033d85a8254d1866a9fab4b09b72defd52dd9e00d3f151f3b4646f98d2
BLAKE2b-256 checksum
How to use checksums
46239012d90dbf9d4a8672e7c14555ef0c23ad2540d9d12c5e27aaed964b46c1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.10.10 {"installer":{"name":"uv","version":"0.10.10","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release files / vision_agents_plugins_gemini-0.6.9-py3-none-any.whl

Download URL vision_agents_plugins_gemini-0.6.9-py3-none-any.whl
Size 27.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
391930c3c73e858d6ae11befd8ad25fd132cfb43fe39235d66597826bcb36473
BLAKE2b-256 checksum
How to use checksums
d6776dbedd1fd968b57a44dcfe0846b3d097484819617b5d8d2c2ca69f177a45
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.10.10 {"installer":{"name":"uv","version":"0.10.10","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release history Release notifications | RSS feed

This release

0.6.9 This release

2 release files

0.6.8

2 release files

0.6.7

2 release files

0.6.6

2 release files

0.6.5

2 release files

0.6.4

2 release files

0.6.3

2 release files

0.6.2

2 release files

0.6.1

2 release files

0.6.0

2 release files

0.5.9

2 release files

0.5.8

2 release files

0.5.7

2 release files

0.5.6

2 release files

0.5.5

2 release files

0.5.4

2 release files

0.5.3

2 release files

0.5.2

2 release files

0.5.1

2 release files

0.5.0

2 release files

0.4.7

2 release files

0.4.6

2 release files

0.4.5

2 release files

0.4.4

2 release files

0.4.3

2 release files

0.4.2

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.8

2 release files

0.3.7

2 release files

0.3.6

2 release files

0.3.5

2 release files

0.3.4

2 release files

0.3.3

2 release files

0.3.2

2 release files

0.3.1

2 release files

0.3.0

2 release files

0.2.10

2 release files

0.2.9

2 release files

0.2.8

2 release files

0.2.7

2 release files

0.2.6

2 release files

0.2.5

2 release files

0.2.4

2 release files

0.2.3

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.14

2 release files

0.1.12

2 release files

0.1.11

2 release files

0.1.9

2 release files

0.1.8

2 release files

0.1.7

2 release files

0.1.6

2 release files

0.1.5

2 release files

0.1.3

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page