Skip to main content

Gemini MCP server

A Python FastMCP server using stdio, the official Google Gen AI SDK, and the Google Interactions API for generation, not models.generate_content. No HTTP server or listening port is started. Stdout is reserved for MCP messages; diagnostics go to stderr.

MCP configuration

Add this to your MCP client's mcpServers configuration (also available as mcp-config.example.json). It needs only uv; uvx fetches and runs the server from PyPI:

{
  "mcpServers": {
    "gemini": {
      "command": "uvx",
      "args": [
        "aio-gemini-mcp@latest"
      ],
      "env": {
        "GEMINI_API_KEY": "YOUR_GEMINI_API_KEY"
      }
    }
  }
}

Optional env entries: GEMINI_MODEL (default text model, gemini-3.8-flash) and GEMINI_OUTPUT_DIR (absolute path for saved media). Client configuration formats can vary.

Setup

This machine already has uv and Python installed. For a fresh checkout:

cd /home/factory-user/gemini-mcp
uv sync --frozen

uv installs the Python version in .python-version if needed. uv.lock pins all dependencies for reproducible installs. Stable versions selected at setup:

Component Version
uv 0.12.22
Python 3.14.8
FastMCP 4.0.10
google-genai 2.28.0

uv is installed in /home/factory-user/.local/bin. Open a new terminal to pick up the updated PATH, or use that absolute path.

Run

Supply your Gemini Developer API key through the MCP client's environment. GEMINI_API_KEY takes precedence over GOOGLE_API_KEY. No key is stored in this project. .env files are not automatically loaded.

uv --directory /home/factory-user/gemini-mcp run --frozen aio-gemini-mcp

Also supported: uv run python -m aio_gemini from the project directory. The server waits for MCP messages on stdin; it is not an interactive terminal app. An MCP client should launch it as a subprocess, as in the MCP configuration above.

Tools

Tool Purpose
generate_text Create a Gemini interaction and return its ID, status, and text
list_models List models available to your Google API key
list_media_models Current documented media defaults and supported model choices; no key needed
generate_image Nano Banana image generation and editing, with optional Google Search
generate_video / generate_omni Omni video generation, editing, extension, and first/last-frame interpolation
transcribe_audio Dedicated speech recognition, smart/verbatim modes, diarization, and word timestamps
generate_speech Single- or two-speaker TTS with voice, language, and delivery style controls
generate_music Lyria songs, instrumental music, clips, lyrics, and image-inspired music
analyze_media Understand images, audio, video (static/agentic), and PDFs
get_interaction Retrieve or poll stored interactions and save inline media
cancel_interaction / delete_interaction Explicitly cancel background work or delete stored Google interactions
upload_file / get_file / list_files / delete_file Manage reusable Google Files inputs and processing readiness
download_file Stream an ACTIVE generated Google file to a unique local file
get_prompt_guide Official prompting guide by guide: image, video, speech, music, transcription, or analysis

generate_text calls client.aio.interactions.create. It accepts prompt, optional model, optional system_instruction, max_output_tokens (default 4096, allowed range 1–65536), optional previous_interaction_id, and store (default false). Model-specific limits still apply. Set GEMINI_MODEL to change the default, gemini-3.8-flash, or pass model per request. Use list_models to check your account's access.

Generation now returns a structured object, not the old bare text string:

{"id": "interaction-id", "status": "completed", "text": "Hello!"}

To continue a conversation, call generate_text with store: true, then pass the returned id as previous_interaction_id on your next call. Keep store: true on each turn you want to continue later. Storage is opt-in: store: false requests that Google not save this interaction for later retrieval or continuation; it does not bypass Google's other data policies. System instructions and generation options are supplied on each call.

generate_text remains non-streaming and foreground-only. The returned status is preserved even if text is empty, rather than reporting an incomplete or blocked response as successful text. list_models still uses the SDK's model-discovery endpoint.

Text, discovery, and metadata requests use a 60-second timeout. Media creation, upload, and download default to 600 seconds, configurable with timeout_seconds (1–1800). Google requests use asynchronous I/O. Clients are closed after each tool call. Upstream error details are redacted from tool errors. Prompts are sent to Google and API use may incur charges. Missing keys do not prevent startup or tool discovery.

Official prompting guides

get_prompt_guide is read-only, offline, and free to call. It needs no API key, generates nothing, and sends nothing to Google. Call it before constructing a media request. It takes a required guide:

{"guide": "music"}

guide is one of image, video, speech, music, transcription, or analysis. For analysis, an optional media_type narrows the guide:

{"guide": "analysis", "media_type": "document"}

Use image, audio, video, or document (PDF), or omit media_type for all. The analysis guide always includes general file-prompting strategies. media_type is ignored for other guides.

The full guide text is returned in the response, not just a link. Guides return closely preserved official wording, templates, and examples, not AI-written summaries. sources includes each page's title, url, selected sections, markdown, and disclosed modifications. Results also include retrieved_on, related tools, supported models, Google attribution, the CC BY 4.0 license link, and separate mcp_notes explaining how the guidance maps to this server.

This is a bundled 2026-10-03 documentation snapshot, not a live lookup. Unrelated API code and illustrative media are omitted. The music guide preserves the batch-generation sections of Google's Lyria prompt guide, not its separate RealTime API examples. Transcription uses configuration rather than a free-form prompt, so its guide preserves official configuration guidance instead of inventing prompting instructions. Separate voice APIs mentioned in the TTS guide are not implemented by this MCP.

See docs/prompt-guides.md for sources, attribution, snapshot boundaries, and refresh instructions.

Media models and API boundaries

Model IDs and request formats were verified on 2026-10-03 using the retrieving-developer-knowledge skill and Google Developer Knowledge MCP. The official model catalog and task-specific guides are the source of truth, not remembered model names.

Capability Default Other current choices
Text and media analysis gemini-3.8-flash Text retains its explicit model/environment override
Images gemini-3.1-flash-image gemini-3.1-flash-lite-image, gemini-3-pro-image
Video / Omni gemini-omni-1.1-flash None
Transcription gemini-3.5-transcribe None
TTS gemini-3.8-flash-tts gemini-3.8-flash-lite-tts
Music lyria-3.5 lyria-3-clip-preview, the current short-clip specialist

Media tool model choices are constrained to these current families, with no legacy fallback. GEMINI_MODEL affects only generate_text, not media tools. This is a dated snapshot, not automatic model discovery or a promise of account access. list_models reports account availability. Refresh the catalog from official docs before adding future models.

All generation and analysis calls use client.aio.interactions.create. There is no generateContent, Imagen, or Veo fallback. Veo, Live audio/live transcription, Lyria RealTime, and voice design/replication have separate APIs and are intentionally outside this Interactions-only suite. TTS accepts existing custom voice IDs but does not create or clone voices.

Sources and model-specific constraints are recorded in docs/media-api.md.

Media inputs, outputs, and examples

Tool arguments below are JSON for your MCP client, not terminal commands. media items require type (image, audio, video, or document), mime_type, and exactly one of:

  • path: an absolute regular file path on the machine running this server.
  • data: plain base64, without a data: prefix.
  • uri: a Google Files URI or another URI supported by the selected model. URI inputs are passed to Google, never fetched by this server.

Local/base64 inputs have a 10 MiB total server memory limit per request, independent of Google's larger file/model limits. For larger input, call upload_file with an absolute path and mime_type. Poll get_file using its name until state is ACTIVE, then use the returned uri and mime_type. PROCESSING is not ready and FAILED should not be retried as ready. Uploaded files expire after 48 hours. Uploading sends the file to Google. Deleting an uploaded file can prevent pending interactions from using it.

Media results include id, status, model, text, ordered outputs, and usage when available. All model output blocks are retained, including interleaved lyrics, multiple images, and transcription word_info annotations. Stored inputs and thought steps are not returned.

Inline binary outputs are decoded and saved in GEMINI_OUTPUT_DIR, defaulting to generated-media/ under the server's working directory. Set an absolute output_directory per call to override it. File names are unique and private (mode 0600), existing files are never overwritten, and the returned path is absolute. Actual MIME types determine extensions; WAV data is saved as returned, without adding a second WAV header. Unknown formats use .bin. Local outputs remain until you remove them.

Base64 is omitted from MCP results by default to avoid filling model context. Set include_inline_data: true if you also need it. Local output paths refer to the server machine, not necessarily the MCP client's machine.

Omni defaults to delivery: "uri" for large videos. These outputs have uri and, for recognized Google Files URIs, file_name. Poll get_file until ACTIVE, then call download_file with that file_name to stream it to disk. Downloads accept Google resource names only, not arbitrary URLs. Alternatively request delivery: "inline" to save bytes immediately. Google's current Omni docs note that get_interaction can return inline data even when creation used URI delivery.

Image generation or editing

Call generate_image:

{
  "prompt": "Create a cinematic watercolor landscape",
  "aspect_ratio": "16:9",
  "image_size": "2K",
  "store": true
}

To edit a local image, supply "media": [{"type": "image", "mime_type": "image/png", "path": "/ABSOLUTE/image.png"}]. To edit a stored result, provide its id as previous_interaction_id. Use include_text: false for image-only output.

Omni video generation and editing

Call generate_video or generate_omni:

{
  "prompt": "A slow tracking shot of waves at sunset, with ocean sounds",
  "aspect_ratio": "16:9",
  "resolution": "1080p",
  "duration_seconds": 8,
  "store": true,
  "background": true
}

Both tools use the same Omni model and controls. Supply ordered reference images for first/last frames and describe the transition in the prompt. Image, audio, and video references can be combined. Prompt for an edit or extension, or use task: "edit" / "extend" with an input video or stored previous_interaction_id. Prompt-based control is preferred; explicit tasks impose stricter constraints. 1080p and 4K are upscaled. See the source guide for extension limits and regions.

Transcription

Call transcribe_audio:

{
  "audio": {"type": "audio", "mime_type": "audio/mp3", "path": "/ABSOLUTE/audio.mp3"},
  "language_codes": ["en-US"],
  "diarization": true,
  "word_timestamps": true
}

For cleaned-up prose, use mode: "smart" without diarization/timestamps. custom_vocabulary cannot be combined with diarization or word timestamps.

TTS

Call generate_speech:

{"text": "Welcome! <short pause> Let's begin.", "voice": "Kore", "style": "warm and friendly"}

For two speakers, use structured turns instead of text:

{
  "turns": [
    {"text": "Hello!", "speaker": "Joe", "style": "cheerful"},
    {"text": "Hi Joe.", "speaker": "Jane", "style": "relaxed"}
  ],
  "speakers": [
    {"speaker": "Joe", "voice": "Puck"},
    {"speaker": "Jane", "voice": "Kore"}
  ]
}

Text is spoken verbatim. Delivery directions belong in style, which maps to speech_metadata, not inline prose. Momentary vocal tags can remain in text. Audio defaults to WAV; audio/l16, audio/mulaw, and audio/alaw are also supported through mime_type. Set sample_rate in Hz if needed.

Lyria music

Call generate_music:

{
  "prompt": "A two-minute instrumental jazz track in D minor, with piano, upright bass, and brushed drums",
  "mime_type": "audio/mp3"
}

Prompt for song duration, structure, BPM, language, and custom lyrics. Music can be inspired by up to ten image references. Use model: "lyria-3-clip-preview" for short clips. WAV output requires lyria-3.5.

Background and multi-turn work

Media tools accept background: true only with explicit store: true. Use the returned id with get_interaction until a terminal status is returned. The server does not auto-poll, silently enable storage, or auto-retry generation (which could create duplicate charges). Cancel by ID with cancel_interaction. Keep store: true on turns you want to retrieve or edit later, and provide previous_interaction_id to tools that support continuation. Options such as output format, voices, and search are scoped to each call and must be repeated. Status is preserved for blocked, incomplete, failed, queued, or cancelled work; empty output never overrides the upstream status.

Develop and validate

uv run ruff check .
uv run ruff format --check .
uv run pytest
uv build

Tests mock Google requests, verify the real SDK's Interactions HTTP transport, and exercise real stdio subprocesses with both current and legacy MCP clients. They need no credentials and make no Google API calls. Guide tests also check preserved sections/examples, snapshot hashes, offline operation, read-only annotations, and MCP output schemas. Live generation requires your own API key and is not covered by these tests.

To upgrade to newer stable dependencies intentionally:

uv lock --upgrade
uv sync --frozen
uv run pytest

Metadata

Release files for aio-gemini-mcp 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for aio-gemini-mcp 0.1.1
File Size Uploaded
aio_gemini_mcp-0.1.1.tar.gz 48.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for aio-gemini-mcp 0.1.1
File Interpreter ABI Platform
aio_gemini_mcp-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 105.3 kB

Release files / aio_gemini_mcp-0.1.1.tar.gz

Download URL aio_gemini_mcp-0.1.1.tar.gz
Size 48.9 kB
Tags Source
SHA-256 checksum
How to use checksums
791fe4613dd257e8483f09ca23bd92364f0147456a5a78ac70d1f3f5633dd832
BLAKE2b-256 checksum
How to use checksums
a15208554c3fc61f4fb88b4cf8588da50dc88d7b8128ae3d9ed560751b5ff4e1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.23 {"installer":{"name":"uv","version":"0.12.23","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release files / aio_gemini_mcp-0.1.1-py3-none-any.whl

Download URL aio_gemini_mcp-0.1.1-py3-none-any.whl
Size 56.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
36abf702ba4875631365dfe86032ea6e78e52f8c2d502e89503ca42cb3292316
BLAKE2b-256 checksum
How to use checksums
db251d9f11b0196b17958261ac51442364a515daadd04a57d263ab1a9ed52c8a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.23 {"installer":{"name":"uv","version":"0.12.23","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release history Release notifications | RSS feed

0.2.0

2 release files

0.1.3

2 release files

0.1.2

2 release files

This release

0.1.1 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page