Skip to main content

Gemini MCP server

A Python FastMCP server using stdio, the official Google Gen AI SDK, and the Google Interactions API for generation, not models.generate_content. No HTTP server or listening port is started. Stdout is reserved for MCP messages; diagnostics go to stderr.

MCP configuration

Add this to your MCP client's mcpServers configuration (also available as mcp-config.example.json). It needs only uv; uvx fetches and runs the server from PyPI:

{
  "mcpServers": {
    "gemini": {
      "command": "uvx",
      "args": [
        "aio-gemini-mcp@latest"
      ],
      "env": {
        "GEMINI_API_KEY": "YOUR_GEMINI_API_KEY"
      }
    }
  }
}

Optional env entries: GEMINI_MODEL (default text model, gemini-3.8-flash) and GEMINI_OUTPUT_DIR (absolute path for saved media). Client configuration formats can vary.

Setup

This machine already has uv and Python installed. For a fresh checkout:

cd /home/factory-user/gemini-mcp
uv sync --frozen

uv installs the Python version in .python-version if needed. uv.lock pins all dependencies for reproducible installs. Stable versions selected at setup:

Component Version
uv 0.12.22
Python 3.14.8
FastMCP 4.0.10
google-genai 2.28.0

uv is installed in /home/factory-user/.local/bin. Open a new terminal to pick up the updated PATH, or use that absolute path.

Run

Supply your Gemini Developer API key through the MCP client's environment. GEMINI_API_KEY takes precedence over GOOGLE_API_KEY. No key is stored in this project. .env files are not automatically loaded.

uv --directory /home/factory-user/gemini-mcp run --frozen aio-gemini-mcp

Also supported: uv run python -m aio_gemini from the project directory. The server waits for MCP messages on stdin; it is not an interactive terminal app. An MCP client should launch it as a subprocess, as in the MCP configuration above.

Tools

Tool Purpose
generate_text Create a Gemini interaction and return its ID, status, and text
generate_image Nano Banana image generation and editing, with optional Google Search
generate_omni Omni video generation, editing, extension, and first/last-frame interpolation
transcribe_audio Dedicated speech recognition, smart/verbatim modes, diarization, and word timestamps
generate_speech Single- or two-speaker TTS with voice, language, and delivery style controls
generate_music Lyria songs, instrumental music, clips, lyrics, and image-inspired music
analyze_media Understand images, audio, video (static/agentic), and PDFs
get_interaction Retrieve or poll stored interactions and save inline media
cancel_interaction / delete_interaction Explicitly cancel background work or delete stored Google interactions
upload_file / get_file / list_files / delete_file Manage reusable Google Files inputs and processing readiness
download_file Stream an ACTIVE generated Google file to a unique local file
get_prompt_guide Official prompting guide by guide: image, video, speech, music, transcription, or analysis

See docs/tools.md for arguments, defaults, accepted values, constraints, and return shapes for every tool.

generate_text calls client.aio.interactions.create. It accepts prompt, optional model, optional system_instruction, max_output_tokens (default 4096, allowed range 1–65536), and optional previous_interaction_id. Model-specific limits still apply. Set GEMINI_MODEL to change the default, gemini-3.8-flash, or pass model per request. Media tool schemas expose their supported model choices, but do not report account-specific access.

Generation now returns a structured object, not the old bare text string:

{"id": "interaction-id", "status": "completed", "text": "Hello!"}

All Interactions calls are stored with Google automatically. To continue a conversation, pass the returned id as previous_interaction_id on the next call. The server does not expose a storage toggle; this does not bypass Google's other data policies. System instructions and generation options are supplied on each call.

generate_text remains non-streaming and foreground-only. The returned status is preserved even if text is empty, rather than reporting an incomplete or blocked response as successful text.

Text and interaction metadata requests use a 60-second timeout. Media creation, upload, and download default to 600 seconds, configurable with timeout_seconds (1–1800). Google requests use asynchronous I/O. Clients are closed after each tool call. Upstream error details are redacted from tool errors. Prompts are sent to Google and API use may incur charges. Missing keys do not prevent startup or tool discovery.

Official prompting guides

get_prompt_guide is read-only, offline, and free to call. It needs no API key, generates nothing, and sends nothing to Google. Call it before constructing a media request. It takes a required guide:

{"guide": "music"}

guide is one of image, video, speech, music, transcription, or analysis. For analysis, an optional media_type narrows the guide:

{"guide": "analysis", "media_type": "document"}

Use image, audio, video, or document (PDF), or omit media_type for all. The analysis guide always includes general file-prompting strategies. media_type is ignored for other guides.

The full guide text is returned in the response, not just a link. Guides return closely preserved official wording, templates, and examples, not AI-written summaries. sources includes each page's title, url, selected sections, markdown, and disclosed modifications. Results also include retrieved_on, related tools, supported models, Google attribution, the CC BY 4.0 license link, and separate mcp_notes explaining how the guidance maps to this server.

This is a bundled 2026-10-03 documentation snapshot, not a live lookup. Unrelated API code and illustrative media are omitted. The music guide preserves the batch-generation sections of Google's Lyria prompt guide, not its separate RealTime API examples. Transcription uses configuration rather than a free-form prompt, so its guide preserves official configuration guidance instead of inventing prompting instructions. Separate voice APIs mentioned in the TTS guide are not implemented by this MCP.

See docs/prompt-guides.md for sources, attribution, snapshot boundaries, and refresh instructions.

Media models and API boundaries

Model IDs and request formats were verified on 2026-10-03 using the retrieving-developer-knowledge skill and Google Developer Knowledge MCP. The official model catalog and task-specific guides are the source of truth, not remembered model names.

Capability Default Other current choices
Text and media analysis gemini-3.8-flash Text retains its explicit model/environment override
Images gemini-3.1-flash-image gemini-3.1-flash-lite-image, gemini-3-pro-image
Video / Omni gemini-omni-1.1-flash None
Transcription gemini-3.5-transcribe None
TTS gemini-3.8-flash-tts gemini-3.8-flash-lite-tts
Music lyria-3.5 lyria-3-clip-preview, the current short-clip specialist

Media tool model choices are constrained to these current families, with no legacy fallback. GEMINI_MODEL affects only generate_text, not media tools. This is a dated snapshot, not automatic model discovery or a promise of account access. Model schemas expose the supported choices, not account availability. Refresh the catalog from official docs before adding future models.

All generation and analysis calls use client.aio.interactions.create. There is no generateContent, Imagen, or fallback to another video-generation API. Live audio/live transcription, Lyria RealTime, and voice design/replication use separate APIs and are intentionally outside this Interactions-only suite. TTS accepts existing custom voice IDs but does not create or clone voices.

Sources and model-specific constraints are recorded in docs/media-api.md.

Media inputs, outputs, and examples

Tool arguments below are JSON for your MCP client, not terminal commands. media items require type (image, audio, video, or document), mime_type, and exactly one of:

  • path: an absolute regular file path on the machine running this server.
  • data: plain base64, without a data: prefix.
  • uri: a Google Files URI or another URI supported by the selected model. URI inputs are passed to Google, never fetched by this server.

Local/base64 inputs have a 10 MiB total server memory limit per request, independent of Google's larger file/model limits. For larger input, call upload_file with an absolute path and mime_type. Poll get_file using its name until state is ACTIVE, then use the returned uri and mime_type. PROCESSING is not ready and FAILED should not be retried as ready. Uploaded files expire after 48 hours. Uploading sends the file to Google. Deleting an uploaded file can prevent pending interactions from using it.

Media results include id, status, model, text, ordered outputs, and usage when available. All model output blocks are retained, including interleaved lyrics, multiple images, and transcription word_info annotations. Stored inputs and thought steps are not returned.

Inline binary outputs are decoded and saved in GEMINI_OUTPUT_DIR, defaulting to generated-media/ under the server's working directory. Set an absolute output_directory per call to override it. File names are unique and private (mode 0600), existing files are never overwritten, and the returned path is absolute. Actual MIME types determine extensions; WAV data is saved as returned, without adding a second WAV header. Unknown formats use .bin. Local outputs remain until you remove them.

Base64 is omitted from MCP results by default to avoid filling model context. Set include_inline_data: true if you also need it. Local output paths refer to the server machine, not necessarily the MCP client's machine.

Omni delivery defaults to "uri", which suits large videos. All Interactions calls are stored automatically. URI outputs have uri and, for recognized Google Files URIs, file_name. Poll get_file until ACTIVE, then call download_file with that file_name to stream it to disk. Downloads accept Google resource names only, not arbitrary URLs. You can also pass delivery: "inline" to save bytes immediately. Google's current Omni docs note that get_interaction can return inline data even when creation used URI delivery.

Image generation or editing

Call generate_image:

{
  "prompt": "Create a cinematic watercolor landscape",
  "aspect_ratio": "16:9",
  "image_size": "2K"
}

To edit a local image, supply "media": [{"type": "image", "mime_type": "image/png", "path": "/ABSOLUTE/image.png"}]. To edit a stored result, provide its id as previous_interaction_id. Use include_text: false for image-only output.

Omni video generation and editing

Call generate_omni:

{
  "prompt": "A slow tracking shot of waves at sunset, with ocean sounds",
  "aspect_ratio": "16:9",
  "resolution": "1080p",
  "duration_seconds": 8,
  "background": true
}

The tool uses the Omni model and controls. Supply ordered reference images for first/last frames and describe the transition in the prompt. Image, audio, and video references can be combined. Prompt for an edit or extension, or use task: "edit" / "extend" with an input video or stored previous_interaction_id. Prompt-based control is preferred; explicit tasks impose stricter constraints. 1080p and 4K are upscaled. See the source guide for extension limits and regions.

Transcription

Call transcribe_audio:

{
  "audio": {"type": "audio", "mime_type": "audio/mp3", "path": "/ABSOLUTE/audio.mp3"},
  "language_codes": ["en-US"],
  "diarization": true,
  "word_timestamps": true
}

For cleaned-up prose, use mode: "smart" without diarization/timestamps. custom_vocabulary cannot be combined with diarization or word timestamps.

TTS

Call generate_speech:

{"text": "Welcome! <short pause> Let's begin.", "voice": "Kore", "style": "warm and friendly"}

For two speakers, use structured turns instead of text:

{
  "turns": [
    {"text": "Hello!", "speaker": "Joe", "style": "cheerful"},
    {"text": "Hi Joe.", "speaker": "Jane", "style": "relaxed"}
  ],
  "speakers": [
    {"speaker": "Joe", "voice": "Puck"},
    {"speaker": "Jane", "voice": "Kore"}
  ]
}

Text is spoken verbatim. Delivery directions belong in style, which maps to speech_metadata, not inline prose. Momentary vocal tags can remain in text. Audio defaults to WAV; audio/l16, audio/mulaw, and audio/alaw are also supported through mime_type. Set sample_rate in Hz if needed.

Lyria music

Call generate_music:

{
  "prompt": "A two-minute instrumental jazz track in D minor, with piano, upright bass, and brushed drums",
  "mime_type": "audio/mp3"
}

Prompt for song duration, structure, BPM, language, and custom lyrics. Music can be inspired by up to ten image references. Use model: "lyria-3-clip-preview" for short clips. WAV output requires lyria-3.5.

Background and multi-turn work

Media tools accept background: true; interactions are stored automatically. Use the returned id with get_interaction until a terminal status is returned. The server does not auto-poll or auto-retry generation (which could create duplicate charges). Cancel by ID with cancel_interaction. Provide previous_interaction_id to tools that support continuation. Options such as output format, voices, and search are scoped to each call and must be repeated. Status is preserved for blocked, incomplete, failed, queued, or cancelled work; empty output never overrides the upstream status.

Develop and validate

uv run ruff check .
uv run ruff format --check .
uv run pytest
uv build

Tests mock Google requests, verify the real SDK's Interactions HTTP transport, and exercise real stdio subprocesses with both current and legacy MCP clients. They need no credentials and make no Google API calls. Guide tests also check preserved sections/examples, snapshot hashes, offline operation, read-only annotations, and MCP output schemas. Live generation requires your own API key and is not covered by these tests.

To upgrade to newer stable dependencies intentionally:

uv lock --upgrade
uv sync --frozen
uv run pytest

Metadata

Release files for aio-gemini-mcp 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for aio-gemini-mcp 0.2.0
File Size Uploaded
aio_gemini_mcp-0.2.0.tar.gz 48.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for aio-gemini-mcp 0.2.0
File Interpreter ABI Platform
aio_gemini_mcp-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 104.1 kB

Release files / aio_gemini_mcp-0.2.0.tar.gz

Download URL aio_gemini_mcp-0.2.0.tar.gz
Size 48.3 kB
Tags Source
SHA-256 checksum
How to use checksums
e9457f7aa7621f3beb85235fa94b8cbd568f7885fee7c48c875c2ccc1ea79eec
BLAKE2b-256 checksum
How to use checksums
f89eedefa8cda6e61a54741ff1611ec860075aa7b079506d6eccf6e1f85c8158
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.23 {"installer":{"name":"uv","version":"0.12.23","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release files / aio_gemini_mcp-0.2.0-py3-none-any.whl

Download URL aio_gemini_mcp-0.2.0-py3-none-any.whl
Size 55.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
18112ec1217f5b8453524e2b96b6737eec462211a9990bb672f525fc1a160ac9
BLAKE2b-256 checksum
How to use checksums
cb852ac1b323c896367ba911078a4fe2f00d70718a34a3eef63980cafc66e088
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.23 {"installer":{"name":"uv","version":"0.12.23","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release history Release notifications | RSS feed

This release

0.2.0 This release

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page