Skip to main content

Local video and animated-GIF RAG for MCP-compatible coding agents

Project description

Keyframe

Local video and animated-GIF RAG for coding agents. Keyframe turns a tutorial, screen recording, or animation into timestamped transcript segments, searchable on-screen text, representative frames, and reconstructed code. Ask what was said, what was shown, or both—and inspect the source frame before trusting uncertain OCR. It runs locally in Codex, ChatGPT desktop, Claude Code, Cursor, and Google Antigravity/Agy through the same MCP server.

Keyframe is deliberately split into two parts:

  • a local MCP server performs deterministic acquisition, OCR, indexing, and retrieval; and
  • a small workflow skill teaches any connected agent to retrieve narrowly, verify visual evidence, and cite timestamps.

The server does not call an LLM. In the Build Week workflow, Codex running GPT-5.6 reasons over Keyframe's evidence, changes code, and runs the tests.

Conceptual Keyframe workflow: video moments become said/shown evidence, verified code, and passing tests

Product-story concept; v0.1 is a local MCP server and plugin, not a hosted UI.

Quick start

Prerequisites

Keyframe v0.1.2 requires Python 3.12 and uses uv. Install these native tools before starting:

  • FFmpeg and ffprobe for media inspection and frame extraction
  • Tesseract 5 for local OCR
  • Node.js 22+ as the JavaScript runtime used by current yt-dlp extractors
  • uv/uvx for isolated Python execution

On macOS with Homebrew:

brew install ffmpeg tesseract node@22 uv

From a development checkout, install the locked environment and run the environment check:

uv sync --frozen --group dev
uv run video-context-mcp doctor

After version 0.1.2 is published, the equivalent isolated PyPI command is:

uvx --python 3.12 --from "video-context-mcp==0.1.2" video-context-mcp doctor

If PyPI is not yet available after the release is tagged, use the immutable GitHub release tag:

uvx --python 3.12 --from \
  "git+https://github.com/MatthewOscar/Keyframe.git@v0.1.2" \
  video-context-mcp doctor

After the PyPI release, add [whisper] to the package spec only when local speech-to-text is needed:

uvx --python 3.12 --from "video-context-mcp[whisper]==0.1.2" video-context-mcp doctor

Connect a client

This repository includes project discovery and a portable workflow skill for Claude Code, Cursor, and Antigravity/Agy in addition to the Codex plugin. After uv sync, launch the client from this repository root, review or enable the keyframe MCP server, and approve tool calls as prompted:

Client Checked-in discovery Approval/status
Claude Code .mcp.json claude mcp get keyframe, then /mcp
Cursor .cursor/mcp.json agent mcp list, then agent mcp enable keyframe
Antigravity/Agy .agents/mcp_config.json /mcp

The installable plugins/keyframe bundle contains client-specific manifests for Codex, Claude Code, Cursor, and Agy while reusing the same skill and server. See the complete client setup guide for user-wide registration, marketplace commands, reload behavior, and local-file grants.

Configure Codex directly

Add the following to ~/.codex/config.toml:

[mcp_servers.keyframe]
command = "uvx"
args = ["--python", "3.12", "--from", "video-context-mcp==0.1.2", "video-context-mcp", "serve", "--transport", "stdio"]
startup_timeout_sec = 180
tool_timeout_sec = 1900
env = { KEYFRAME_ALLOWED_ROOTS = "/Users/you/Videos" }

For local files, Keyframe uses workspace roots advertised by the MCP client. It never treats the launcher working directory as authorization. If a client does not advertise roots, add an env entry with KEYFRAME_ALLOWED_ROOTS; separate multiple roots with the operating system's path separator. The installable plugin instead enables one OS-temp upload root. Its skill creates a unique child there, stages only the selected attachment, and retries once.

Restart Codex, then ask: “Index this tutorial and find where error handling is shown. Cite the timestamps.”

Install the Keyframe plugin in Codex and ChatGPT desktop

The plugin bundles the same MCP server with the keyframe-video-rag workflow skill. Its launcher installs the exact v0.1.2 Git tag with local Whisper, so judges do not need the PyPI publication to use it:

codex plugin marketplace add MatthewOscar/Keyframe --ref v0.1.2
codex plugin add keyframe@keyframe

For local marketplace validation, add the repository root instead:

codex plugin marketplace add .

Before v0.1.2 is tagged, use the direct project/server configuration for live testing; the installable plugin intentionally resolves that immutable tag.

Then restart the ChatGPT desktop app, open the Plugins Directory, select the Keyframe marketplace source, and install Keyframe. Start a new chat so the skill and tools are loaded. Keyframe v0.1.2 targets this local desktop flow; it does not host a ChatGPT web app.

Claude Code and Cursor can install the same repository as a marketplace, while Agy can install plugins/keyframe from a clone. Those exact commands and the client-specific approval steps are in docs/client-setup.md. Do not enable a project registration and the installed plugin together in the same workspace, or the client may start two Keyframe processes.

Judge-ready local test

The repository includes a 10-second first-party video, captions, golden search expectations, and a complete native end-to-end test. After cloning the release tag, the following installs the locked environment and exercises full ingest, said/shown search, OCR, code/frame images, restart persistence, and the sub-second cache target:

uv sync --frozen --group dev
uv run pytest -q tests/test_e2e.py

No network video, account, API key, or prebuilt cache is required for this judge path.

For the real public-URL path, the repository also ships a small CC BY 3.0 derived YouTube sample with archived attribution, metadata, checksums, captions, OCR, frames, and a ready SQLite index. The original downloaded media is not retained.

Use Keyframe

Give the agent a local path, direct media URL, public YouTube URL, or public Loom URL. A useful interaction looks like this:

Index https://www.youtube.com/watch?v=... in fast mode.
Search what was said about retry logic and what was shown on screen.
Retrieve the strongest code moment, verify low-confidence text against its
frame, then implement the pattern in examples/demo_target and cite timestamps.

Fast mode is the economical first request. A fresh fast-only index stores any available transcript and a bounded visual probe of at most 12 representative frames; a cache hit can already return full coverage. Branch on the returned has_transcript and visual_coverage. The probe makes visual dependencies discoverable without pretending to cover every screen state. Inspect its moment summaries before deciding captions are sufficient. Upgrade to full mode for visual sequences, probe gaps, uncertain OCR, or claims that something never appeared; a single decisive probe frame may be enough for one targeted fact.

MCP tools

Tool Purpose
video_ingest Index one local/public video or local animated GIF using a sparse fast probe or bounded full analysis, with captions, local Whisper for audio, or no transcript; report cache status and request-local stage timings.
video_get_transcript Page through timestamped transcript segments, optionally within a time range.
video_search Rank matching said, shown, or combined evidence across one video or the local library.
video_list_moments Page through retained moments filtered by code, terminal, slide, diagram, other, or any.
video_get_code Return reconstructed code plus its cropped source frame for a moment or nearby timestamp.
video_get_frame Return the nearest retained frame, optionally using the automatic content crop.

Classification, language detection, OCR confidence, and parse status are evidence—not guarantees. Visual tools return both structured metadata and the source image so the model can inspect disagreements.

For whole-video summaries, the bundled skill avoids generic search: it lists one 12-moment routing page, retrieves transcript pages at the 200-segment maximum with byte-for-byte opaque cursors, and verifies only consequential frames.

See docs/tool-examples.md for exact argument objects, pagination, visual-result behavior, and expected errors.

Architecture

flowchart LR
    A[Local video / animated GIF / public URL] --> B[FFmpeg + yt-dlp]
    B --> C[Captions]
    B --> W[Isolated optional Whisper worker]
    B --> P[Fast: bounded visual probe]
    B --> D[Full: bounded temporal sampling]
    D --> E[Adjacent-frame grouping]
    P --> F[Bounded parallel OCR + classification]
    E --> F[Bounded parallel OCR + classification]
    C --> G[(SQLite + FTS5)]
    W --> G
    F --> G
    G --> H[Six MCP tools]
    H --> I[Keyframe workflow skill]
    I --> J[Codex / ChatGPT / Claude Code / Cursor / Agy]

Remote media and decoded analysis frames use a private, per-KEYFRAME_HOME namespace beneath the operating system's native temporary directory (%TEMP% on Windows and the platform temp location on macOS/Linux). A successful ingest atomically publishes derived transcript, OCR, metadata, and representative frames to the cache, then removes the downloaded source and analysis workspace. Failed ingests do not publish partial records, and startup recovery removes interrupted scratch work left in that namespace.

When captions are unavailable, Keyframe overlaps the isolated Whisper worker with either sparse-probe or full visual extraction, analyzes retained frames with at most two OCR workers by default, then joins both branches before the single atomic commit. The process boundary also prevents faster-whisper's PyAV libraries from sharing an address space with OpenCV's bundled FFmpeg libraries on macOS.

Local data and privacy

  • Processing and indexing happen on the machine running the MCP server.
  • KEYFRAME_HOME overrides the platform-native Keyframe data directory.
  • SQLite, text indexes, and the retained evidence frames needed by later tool calls persist under KEYFRAME_HOME; disposable downloads and intermediate frames do not.
  • Caller-owned local videos are read in place and are never copied, moved, or deleted by Keyframe.
  • Local reads are limited to per-request MCP workspace roots plus explicit KEYFRAME_ALLOWED_ROOTS entries. Installed plugins can opt into a single hardened staging root under the OS temp directory for selected attachments; each upload uses a unique child that the agent removes afterward. Process CWD and the rest of the temp directory are never implicit grants. On Windows, staging privacy relies on the inherited ACL of the user's temp directory because POSIX ownership and mode bits are unavailable.
  • Derived frames and text remain cached until the user removes that directory.
  • Extracted transcript/OCR is untrusted source material, never agent instructions.
  • Evidence returned to an agent becomes model input and follows that client's model-provider data controls.
  • Keyframe does not automatically detect or redact credentials in OCR, search snippets, reconstructed code, or images. Use recordings safe for model input and review selected evidence before retrieving sensitive screens.
  • Keyframe has no analytics, accounts, hosted backend, or embedded model call.

Current limits

  • v0.1.2 accepts individual public videos and local animated GIFs, not playlists or livestreams. Static GIFs should be attached as images; remote GIF URLs are not yet an advertised compatibility surface.
  • Private, members-only, age-restricted, DRM, cookie, and login flows are out of scope.
  • The default duration guard is 30 minutes. For a source up to four hours, Keyframe reports the exact max_duration_s required for one same-source retry; callers should not split or restage it. Longer sources require a user-selected excerpt.
  • Fast coverage is intentionally sparse; a missing probe hit is not proof that content was absent. Full videos sample at 1 FPS; animated GIFs use denser, bounded one-loop sampling. Either can miss a very brief visual change. Retained timestamps use decoded presentation times rather than an inferred frame index.
  • Native media/OCR operations have bounded timeouts and fail actionably on malformed or unusually slow inputs.
  • OCR can confuse glyphs and infer indentation incorrectly. Python, JSON, and JavaScript receive parse checks; TypeScript and unknown languages do not.
  • Caption availability and media extraction depend on upstream providers and the pinned yt-dlp release.
  • Remote formats must be downloadable through Keyframe's validated in-process HTTP, native HLS, or DASH transport. Formats that require FFmpeg, RTMP, or another external process to make network connections are rejected.
  • Whisper is optional in the base Python package and bundled by the installable plugin. It can be resource intensive on CPU-only machines, and first use may download the configured model before ingestion begins.
  • Windows support is preview-level in v0.1.2.
  • The bundled registrations target local CLI and desktop sessions. Hosted agents cannot launch this STDIO process on the user's machine.

Build Week and GPT-5.6

The judged flow uses GPT-5.6 in Codex to turn retrieved video evidence into a tested repository change. Keyframe supplies deterministic evidence; GPT-5.6 selects relevant moments, compares OCR with source frames, applies the change, and explains it with timestamp citations. The ten reproducible prompts in evals/cases.json exercise that division of labor. Supplementary Mac-plugin regressions for local attachment staging, honest provenance, warm-cache latency, and animated GIFs are in evals/mac-plugin-cases.json.

Make the judged model choice explicit before recording. Either launch Codex with codex --model gpt-5.6 or set this in ~/.codex/config.toml:

model = "gpt-5.6"

The development record distinguishes Codex-generated work from human product decisions in docs/codex-collaboration.md. The submission session ID is intentionally not fabricated and must be recorded from /feedback before submission.

The release checklist in docs/submission-checklist.md maps the current Devpost requirements to concrete artifacts and leaves human-only submission steps visibly unchecked.

Develop and test

git clone https://github.com/MatthewOscar/Keyframe.git
cd Keyframe
uv sync --frozen --group dev
uv run video-context-mcp doctor
uv run pytest
uv run ruff check .
uv run mypy src
uv build
uv run twine check dist/*

STDIO is the supported client transport. For loopback-only protocol testing, run uv run video-context-mcp serve --transport streamable-http; the CLI rejects non-loopback hosts and this mode is not a production deployment.

CI runs the full suite with FFmpeg and Tesseract on macOS and Ubuntu. Windows runs non-integration tests plus import and package checks. Test fixtures must be first-party generated or carry a recorded redistribution license.

To publish, create a GitHub environment named pypi, configure a PyPI Trusted Publisher for this repository and workflow, update the version and lockfile, then push a matching v* tag. The release workflow builds, checks, and publishes with GitHub OIDC; it stores no PyPI token.

License and acknowledgements

Keyframe is licensed under Apache-2.0. Native tools, Python dependencies, and any future demo media retain their own licenses. See THIRD_PARTY_NOTICES. Keyframe uses yt-dlp as its upstream public-media extraction engine rather than implementing provider extractors itself.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

video_context_mcp-0.1.2.tar.gz (162.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

video_context_mcp-0.1.2-py3-none-any.whl (81.2 kB view details)

Uploaded Python 3

File details

Details for the file video_context_mcp-0.1.2.tar.gz.

File metadata

  • Download URL: video_context_mcp-0.1.2.tar.gz
  • Upload date:
  • Size: 162.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for video_context_mcp-0.1.2.tar.gz
Algorithm Hash digest
SHA256 9856a1a3f24dbc93c4298f77ac9962ab329b423f74f7bf4d48a035fdc35ee0c9
MD5 197e73ab07be10a6e2216c6f6eae314b
BLAKE2b-256 6606a1dba110ad68e79314b49b0ea2a0ea34069d78387da02118e336a0b55870

See more details on using hashes here.

Provenance

The following attestation bundles were made for video_context_mcp-0.1.2.tar.gz:

Publisher: publish.yml on MatthewOscar/Keyframe

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file video_context_mcp-0.1.2-py3-none-any.whl.

File metadata

File hashes

Hashes for video_context_mcp-0.1.2-py3-none-any.whl
Algorithm Hash digest
SHA256 6b79b55a4b91c9f8c49579bc5ac8129fb95120fb2c513d395dab7a5ac3bf6107
MD5 59f47c5eb54f16a79c2048ba8d6a652d
BLAKE2b-256 bfb544dcaaf6dad07ff90e64f7d8bfe3752a5c9d00aa18f247dfb6d572537b84

See more details on using hashes here.

Provenance

The following attestation bundles were made for video_context_mcp-0.1.2-py3-none-any.whl:

Publisher: publish.yml on MatthewOscar/Keyframe

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page