Skip to main content

gemini-mcp

MCP gateway for massive-payload Gemini analysis on the AI Studio free tier. Runs on a small VPS; a local Freebuff Desktop agent sends huge bodies to it and gets back a compact consolidated result — the agent's own context stays clean.

  • Cascading Flash chain 3.8 → 3.7 → 3.6 → 3.5 → 3 via gemini-router (per key×model RPD/RPM/TPM ledger shared with tldr-digest).
  • Over-limit bodies: split into overlapping chunks, each marked [CHUNK N/M], analyzed separately, then consolidated (recursively if needed).
  • thinking_level default high, per-request override; opt-in overflow (flash-lite → gemma) when the daily quota is gone; hard stop otherwise; per-request wait policy for minute-limit resets (default: wait, countdown reported in progress).
  • Deep progress to the caller (MCP progress + logging notifications + embedded trace), rotating disk logs, gemini-mcp quota|logs|calls|doctor in the terminal.
  • Transports: SSH-stdio (primary) and streamable HTTP daemon (127.0.0.1, bearer).
  • Dual backend (0.8.0): router (fleet: shared SQLite ledger on the VPS) or direct (plain Google AI Studio: same engine, in-memory counters, no server at all). backend: auto picks router only when persistent state exists; env GEMINI_MCP_BACKEND overrides. See docs/08-dual-mode.md.

Quick start

If you're human

pnpm dlx @vernikr/gemini-mcp setup     # or: uvx --from gemini-mcp-gateway gemini-mcp setup

Answer the questions: 1 Simple (direct to Google AI Studio, keys via hidden paste/file/skip) or 2 Fleet (your VPS — the wizard clones the repos, provisions the server, pushes keys, health-checks, all you answer is host + projects folder). It finishes with a numbered checklist: paste one JSON block into your MCP client (Freebuff/Claude Desktop/Cursor — or add --write-client freebuff|claude|cursor and it merges the block in for you, with a .bak backup), restart the client, send the agent the printed verify prompt. gemini-mcp doctor re-checks keys/chain/limits; gemini-mcp uninstall --yes removes everything the wizard wrote. Full client matrix + fleet JSON forms: docs/06-consumer.md.

If you're an AI agent

One shot, no TTY, machine-readable. Exit 0 + "ok": true = done; otherwise every failure carries error/deploy_error/clients_failed[].fix naming the exact remedy — apply it, re-run the same command (idempotent). Relay to the human only the two manual actions: restart the MCP client, then send the agent the prompt from verify_prompt (tool quota_status, MCP server gemini).

# direct (no server); keys from a file — never from shell history
npx -y @vernikr/gemini-mcp setup --backend direct \
  --keys-file /abs/path/keys.txt --write-client freebuff --json

# fleet (shared-quota server); clones/provisions/health-checks by itself
npx -y @vernikr/gemini-mcp setup --backend router --vps-host <ip> \
  [--vps-user root] [--projects-dir ~/projects] [--keys-file /abs/path/keys.txt] \
  --write-client freebuff --json

Fleet prerequisites (each failure message tells you how to check): GitHub auth for the private repos (git ls-remote https://github.com/vernikr/gemini-mcp), passwordless ssh <user>@<ip>, bash+rsync, macOS/Linux. No env vars needed afterwards: the gateway auto-discovers ~/.config/gemini-mcp/.

Analyzing a FILE (agent): never read it yourself and never inline its contents into a prompt. Call file_upload_ticket(filename, size_bytes) → HTTP PUT the raw bytes to upload_url → analyze_file(upload_id, task). Extraction, chunking and quota are server-side; quota_status returns this recipe in its file_flow field, and raw RTF/PDF/DOCX payloads in text tools fail fast with a raw_document_inlined fix.

Direct mode needs nothing but API keys. Usage is tracked in a LOCAL persistent ledger (~/.config/gemini-mcp/state.db): quota counters survive restarts and are shared by all gateway processes on the machine — but not across machines (that's fleet mode: VPS daemon + backend: router). Every response says backend=direct in its report line; gemini-mcp uninstall --yes removes keys, config and ledger together. The npm shim @vernikr/gemini-mcp (pnpm-first install) ships the same core at the same version.

Status

v0.8.0 — published consumer package. Fleet side (VPS daemon 0.6.x, SSH-stdio + HTTP tunnel) runs in production; consumer side ships as PyPI gemini-mcp-gateway (+ npm shim @vernikr/gemini-mcp) with the dual-mode backend (router/direct, docs/08) and the setup wizard (M7.2). Router core: gemini-router 0.1.1 on PyPI. 80 tests green, no network needed (respx/fakes only). Next: M7.3 dual-publish CI, M7.4 clean-machine beta.

Development

uv sync --extra dev    # dev pin: gemini-router from git tag (tool.uv.sources);
                       # published metadata uses the PyPI version range — no auth needed
uv run pytest -q       # 80 tests: chunker, pipeline, tools/server, logging, direct
                       # backend (respx), setup wizard — all faked, no network
uv run ruff check src tests
GEMINI_API_KEYS=... uv run gemini-mcp setup|quota|calls|logs|prune|doctor
Stage Artifact
1 · Unpacking/reframing + Q&A docs/01-unpacking-reframing.md
2 · Functional/business requirements docs/02-requirements.md
3 · Architecture & stack docs/03-architecture.md
4 · Plan (M0–M6) docs/04-plan.md
Consumer guide (client JSON, direct mode) docs/06-consumer.md
Dual-mode backend design docs/08-dual-mode.md
Risks docs/blockers/

Planned usage (Freebuff, ~/.agents/mcp.json)

{ "mcpServers": {
    "gemini": {                       // A) SSH-stdio — primary, encrypted, no open ports
      "command": "ssh",
      "args": ["root@38.244.152.2", "/opt/apps/gemini-mcp/.venv/bin/gemini-mcp", "stdio"] },
    "gemini-http": {                  // B) HTTP via `ssh -L 8790:127.0.0.1:8790 root@38.244.152.2`
      "type": "http", "url": "http://127.0.0.1:8790/mcp",
      "headers": { "Authorization": "Bearer $GEMINI_MCP_TOKEN" } } } }

Tools: analyze_large (chunking pipeline), gemini_generate (single shot), file_upload_ticket + analyze_file (file flow: upload → server-side pandoc extraction → analysis), analyze_result (async poll), quota_status, router_logs. Raw document bytes (RTF/PDF/DOCX markers) pasted into gemini_generate/analyze_large are rejected with raw_document_inlined + the exact fix — always use the file flow.

Deployment (M4)

From the workstation (sibling checkout of gemini-router required next to this repo):

bash scripts/run_remote.sh setup    # rsync both repos → venv, .env, systemd unit, doctor
bash scripts/run_remote.sh smoke    # first real-key run on the VPS (spends ≤1 RPD)
bash scripts/run_remote.sh status|logs

Freebuff Desktop wiring (SSH-stdio and HTTP-via-tunnel snippets, verification steps): docs/freebuff-wiring.md. The daemon never leaves 127.0.0.1:8790; no firewall changes are made (server-spec compliant).

Ops (after M4)

ssh root@38.244.152.2
sudo systemctl status gemini-mcp          # daemon (HTTP)
/opt/apps/gemini-mcp/.venv/bin/gemini-mcp quota    # what's left, per key×model
/opt/apps/gemini-mcp/.venv/bin/gemini-mcp logs -f  # live app log

Conventions: English in files, Conventional Commits, SemVer, docs updated with every functional change (see AGENTS.md).

Metadata

Release files for gemini-mcp-gateway 0.12.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for gemini-mcp-gateway 0.12.0
File Size Uploaded
gemini_mcp_gateway-0.12.0.tar.gz 172.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for gemini-mcp-gateway 0.12.0
File Interpreter ABI Platform
gemini_mcp_gateway-0.12.0-py3-none-any.whl Python 3 none any Details

Total release size: 227.4 kB

Release files / gemini_mcp_gateway-0.12.0.tar.gz

Download URL gemini_mcp_gateway-0.12.0.tar.gz
Size 172.4 kB
Tags Source
SHA-256 checksum
How to use checksums
80270bfc12a33b72de0aa238852b8ee2d10bccb2d9a849a5fb0d547d5f5f5694
BLAKE2b-256 checksum
How to use checksums
688c45911b74c8a142aec22ef46d595f0872aaa185af203a9f0d05d443787b25
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 10, 2026.

Transparency log

Release files / gemini_mcp_gateway-0.12.0-py3-none-any.whl

Download URL gemini_mcp_gateway-0.12.0-py3-none-any.whl
Size 55.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
1b55b4ccd367789875a6f3ea25f05acbb9e4db9e2c0791f4126b9f76c1f71b19
BLAKE2b-256 checksum
How to use checksums
2621a3f40bc31860ae76b171a276b018827128b2cf857453e9a86f3c89dfe98b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 10, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.12.0 This release

2 release files

0.11.0

2 release files

0.10.1

2 release files

0.10.0

2 release files

0.9.3

2 release files

0.9.2

2 release files

0.9.1

2 release files

0.9.0

2 release files

0.8.2

2 release files

0.8.1

2 release files

0.8.0

2 release files

0.7.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page