Skip to main content

MCP server that bridges vision models to text-only coding models using Florence-2

Project description

videre-mcp

MCP server that bridges vision models to text-only coding models using Florence-2.

Non-vision LLMs can't see images — videre-mcp fixes that. It loads a Florence-2 vision model locally and exposes four MCP tools that convert images (including SVGs) and screenshots into structured text descriptions that any text-based model can consume.

Screenshot tool → videre-mcp (Florence-2) → Text description → Coding model

Installation

pip install videre-mcp

Or with uv:

uv pip install videre-mcp

Requires Python 3.11+ and ~300MB disk space for the Florence-2-base model weights (downloaded automatically on first use).

Usage

Add to your OpenCode configuration:

{
  "mcpServers": {
    "videre-mcp": {
      "command": "videre-mcp"
    }
  }
}

Or run directly:

videre-mcp
# or
python -m videre_mcp

Tools

describe_image

Generate a natural language description of an image.

Parameters:

  • image_path (str) — Path to the image file (supports PNG, JPEG, SVG)
  • detail_level (str, optional) — "normal" (default) for brief caption, "high" for detailed description

Example:

result = describe_image("/path/to/photo.png", detail_level="high")
# Returns:
# {
#   "description": "A sunlit meadow with wildflowers in bloom...",
#   "model": "Florence-2-base",
#   "prompt_used": "<MORE_DETAILED_CAPTION>"
# }

ocr_image

Extract text from an image using optical character recognition.

Parameters:

  • image_path (str) — Path to the image file (supports PNG, JPEG, SVG)
  • detail_level (str, optional) — "normal" (default) for plain text, "high" for text with bounding regions

Example:

result = ocr_image("/path/to/document.png", detail_level="high")
# Returns:
# {
#   "text": "Invoice Number 12345",
#   "regions": [
#     {"label": "Invoice Number 12345", "bbox": [10, 20, 30, 40, 50, 60, 70, 80]}
#   ]
# }

describe_screenshot

Describe UI regions in a screenshot — designed for coding agents that need to understand screen layouts.

Parameters:

  • image_path (str) — Path to the screenshot file (supports PNG, JPEG, SVG)
  • detail_level (str, optional) — "normal" (default) for dense region captions, "high" for per-region descriptions

Example:

result = describe_screenshot("/path/to/screenshot.png")
# Returns:
# {
#   "regions": [
#     {"bbox": [10, 20, 30, 40], "label": "search bar"},
#     {"bbox": [100, 200, 300, 250], "label": "submit button"}
#   ],
#   "model": "Florence-2-base"
# }

take_screenshot

Capture a screenshot and optionally describe it using Florence-2. Supports multi-monitor setups via the monitor parameter.

Parameters:

  • output_path (str, optional) — Path to save the screenshot PNG. If None, saves to a temp file.
  • monitor (int, optional) — Monitor index: 0 = all monitors combined, 1 = primary, etc. (default: 0)
  • describe (bool, optional) — If True, also run describe_screenshot on the captured image (default: True)

Example:

result = take_screenshot(monitor=1, describe=True)
# Returns:
# {
#   "path": "/tmp/tmpxxxxxx.png",
#   "width": 1920,
#   "height": 1080,
#   "monitor": 1,
#   "regions": [
#     {"label": "search bar", "bbox": [10, 20, 30, 40]},
#     ...
#   ]
# }

Supported Formats

Format Support
PNG ✅ Native
JPEG ✅ Native
SVG ✅ Via cairosvg rasterization

Requirements

  • Python 3.11+
  • ~300MB disk for model weights (auto-downloaded on first inference)
  • Works on CPU; GPU (CUDA) is auto-detected and used if available

Continuous Integration

The Florence-2 slow tests (real model load + inference) run on a nightly schedule via GitHub Actions. See .github/workflows/slow-tests.yml.

License

MIT — see LICENSE.

Third-party licenses

This package vendors a patched copy of Microsoft's Florence-2 processor (src/videre_mcp/_vendor/processing_florence2.py) under Microsoft's MIT license. See src/videre_mcp/_vendor/LICENSE-Microsoft-Florence-2.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

videre_mcp-0.2.0.tar.gz (42.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

videre_mcp-0.2.0-py3-none-any.whl (29.9 kB view details)

Uploaded Python 3

File details

Details for the file videre_mcp-0.2.0.tar.gz.

File metadata

  • Download URL: videre_mcp-0.2.0.tar.gz
  • Upload date:
  • Size: 42.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for videre_mcp-0.2.0.tar.gz
Algorithm Hash digest
SHA256 cf7d9d3a46c5f29b6b8261a9a10b9bc84c29906b28ec854ac32c3d6f708c6a6e
MD5 1fd96b595ca94285cba1d8696a4f75c3
BLAKE2b-256 91b3030375f036755c77070f3e835980db93ed9969876574cbea17fa475f8e34

See more details on using hashes here.

Provenance

The following attestation bundles were made for videre_mcp-0.2.0.tar.gz:

Publisher: release.yml on Veedubin/Videre-MCP

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file videre_mcp-0.2.0-py3-none-any.whl.

File metadata

  • Download URL: videre_mcp-0.2.0-py3-none-any.whl
  • Upload date:
  • Size: 29.9 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for videre_mcp-0.2.0-py3-none-any.whl
Algorithm Hash digest
SHA256 e80b68047735df141191d99d51b1d6e607dce6137afa84af45e5b17a50c46f5d
MD5 a3676fcd823d3662d324323f29f38ac0
BLAKE2b-256 bb74b49fe0e64a6b42cdabc429bc157fd1a1865d63b0bcd29eb57c87f9b216a0

See more details on using hashes here.

Provenance

The following attestation bundles were made for videre_mcp-0.2.0-py3-none-any.whl:

Publisher: release.yml on Veedubin/Videre-MCP

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page