Skip to main content

cue-kit

Terminal session showing cue-kit producing a video report with frame timeline and transcript

CI PyPI Python 3.10+ Status: alpha Code style: ruff GitHub Stars GitHub Issues

A tactical toolkit for turning video into structured, usable text. Feed it a URL or local file and get extracted frames, a timestamped transcript, and output tailored to your goal — from quick summaries and intel briefs to training docs, lecture notes, or clean transcripts.

LLMs like Claude can read pages, run code, and browse repos — but they can’t truly watch video. Drop in a YouTube link and they’re left guessing from titles or relying on incomplete transcripts that miss the visuals.

cue-kit fixes that. It converts video into frames plus a synchronized transcript that LLMs can actually process. With the bundled Claude Code skill, /cue-kit <url> <question> becomes a one-liner: Claude analyzes every frame, follows the transcript, and answers based on what’s really happening — on screen and in audio, not guesswork.

Status: alpha. The summary and transcript modes work today; training-doc and lecture-notes are scaffolded and under active development.

What it's for

  • Watching a video so you don't have to.
  • Converting recorded training videos into formatted, navigable docs.
  • Capturing lecture-style content where someone narrates over slides — keeping the transcript aligned to each slide.
  • Producing a clean, timestamped transcript for any video.

Common workflows

Tracking suspicious vessel behavior. Feed in AIS playback clips or screen recordings of vessel movement. Cue-kit isolates key segments—course changes, loitering patterns, or AIS gaps—and surfaces the exact moments where behavior deviates from norms. Instead of scrubbing timelines manually, you get a quick breakdown of when and how a vessel started acting suspiciously.

Analyzing port activity from video feeds. Drop in CCTV or drone footage from port areas and ask what changed over time. Cue-kit extracts key frames and summarizes activity patterns—unusual docking sequences, unexpected cargo handling, or irregular vessel arrivals. Useful for quickly spotting anomalies without reviewing hours of footage.

Extracting insights from incident recordings. After an onboard incident or near miss, upload bridge recordings or monitoring footage. Cue-kit identifies critical moments leading up to the event, highlights what was visible on instruments or surroundings, and reconstructs a timeline of what likely happened—helping with faster post-incident analysis.

Modes

Mode What it produces Status
summary Source metadata, frame timeline, and timestamped transcript (default) working
transcript Just a clean, timestamped transcript working
training-doc Raw frames + transcript + a suggested LLM prompt; downstream model produces the formatted doc scaffolded
lecture-notes Transcript grouped under detected slides; falls back to ungrouped transcript while slide detection is stubbed scaffolded

Install

cue-kit runs on macOS, Linux, and Windows with Python 3.10+. It shells out to two system tools, ffmpeg (a full build, including ffprobe and the libmp3lame encoder — all standard packages qualify) and yt-dlp.

macOS (Homebrew):

brew install ffmpeg yt-dlp

Linux (Debian/Ubuntu):

sudo apt install ffmpeg
python3 -m pip install --user yt-dlp

Windows (PowerShell):

winget install Gyan.FFmpeg yt-dlp.yt-dlp

Open a new terminal afterwards so the updated PATH is picked up.

Then install cue-kit itself:

pip install cue-kit                 # latest release from PyPI
# or pin a specific release straight from GitHub:
pip install "git+https://github.com/atlas-bear/cue-kit@v0.2.0"

pipx (pipx install cue-kit) keeps it in an isolated environment and is the tidiest option for a CLI.

yt-dlp breaks periodically as video sites change. If URL downloads start failing, update it first (brew upgrade yt-dlp, pip install -U yt-dlp, or winget upgrade yt-dlp.yt-dlp).

For development from a clone:

git clone https://github.com/atlas-bear/cue-kit.git
cd cue-kit
pip install -e '.[dev]'

Configure

A Whisper API key is only needed for videos without native captions (most local files). cue-kit reads keys from, in order: process environment, the user config file, then ./.env.

OS Config file
macOS / Linux ~/.config/cue-kit/.env
Windows %APPDATA%\cue-kit\.env
# macOS / Linux
mkdir -p ~/.config/cue-kit
cp .env.example ~/.config/cue-kit/.env
chmod 600 ~/.config/cue-kit/.env
# Windows
New-Item -ItemType Directory -Force "$env:APPDATA\cue-kit"
Copy-Item .env.example "$env:APPDATA\cue-kit\.env"

cue-kit prefers Groq (whisper-large-v3 — cheaper, faster) and falls back to OpenAI (whisper-1). Set whichever you have.

Data handling: what leaves your machine

  • Local files with --no-whisper: nothing. All processing (frame extraction, caption parsing) happens locally.
  • URLs: yt-dlp contacts the hosting site to download the video and any captions.
  • Whisper fallback: only when no captions are available and an API key is configured, cue-kit extracts the audio track (mono, 16 kHz) and uploads it to Groq or OpenAI for transcription. Video frames are never uploaded by cue-kit.
  • Working files (video, frames, audio) stay in the working directory on disk until you delete them.

For sensitive material, pass --no-whisper to guarantee no audio leaves the machine. Downstream, whatever reads cue-kit's output (for example, an LLM reading the frames) is governed by that tool's own data policy.

Quick start

# Default: summary of a YouTube video
cue-kit https://youtu.be/<id>

# Just a transcript
cue-kit https://youtu.be/<id> --mode transcript

# Training video → structured doc
cue-kit ./onboarding-screencast.mp4 --mode training-doc

# Lecture with slides → transcript grouped under each detected slide
cue-kit ./talk.mp4 --mode lecture-notes

# Focus on a specific section
cue-kit https://youtu.be/<id> --start 2:15 --end 5:00

Output goes to --out-dir if specified, otherwise a fresh temp directory. The working directory is printed at the start and end of the run. cue-kit --version prints the installed version.

CLI flags

Flag Default What it does
--mode {summary,transcript,training-doc,lecture-notes} summary Output shape
--start T / --end T none Focus on a section. Accepts SS, MM:SS, or HH:MM:SS. Triggers a denser frame budget.
--max-frames N 80 Cap on frame count. Hard ceiling 100.
--resolution W 512 Frame width in pixels. Bump to 1024 if Claude needs to read on-screen text.
--fps F auto Override the auto-scaled fps. Capped at 2.0.
--out-dir PATH tmp Working directory. Defaults to a fresh temp dir.
--no-whisper off Disable the Whisper fallback. Frames-only if no native captions.
--whisper {groq,openai} auto Force a backend. Default: prefer Groq, fall back to OpenAI.

From the Claude Code skill, the same flags work alongside a question:

/cue-kit https://youtu.be/<id> --start 0:00 --end 0:30 what hook did they open with?

How it works

  1. Source. A URL (anything yt-dlp supports — YouTube, Loom, TikTok, X, Instagram, hundreds more) or a local file (.mp4, .mov, .mkv, .webm, plus a few others — full list in download.VIDEO_EXTS).
  2. Download. yt-dlp fetches into a temp working directory; local files are probed in place, no copy.
  3. Frames. ffmpeg extracts at an auto-scaled rate. The frame budget is duration-aware — up to 30s gets one frame per second (minimum 12, limited by the 2 fps cap), 30-60s gets 40, 1-3 min gets 60, and anything longer gets 80 (the default --max-frames; raise it to 100 for long videos). --start/--end ranges get a denser budget. Hard caps: 2 fps, 100 frames. JPEGs at 512px wide by default; bump with --resolution 1024 to read on-screen text.
  4. Transcript. First try: yt-dlp pulls native captions (manual or auto-generated) — free, fast, and good enough for most public videos. Fallback: extract a mono 16 kHz mp3 and ship it to Whisper — Groq's whisper-large-v3 (preferred — cheaper and faster) or OpenAI's whisper-1.
  5. Output. The mode renderer prints frame paths with t=MM:SS markers and a timestamped transcript. From the skill, Claude Reads each frame in parallel — JPEGs render directly as images in its context — and answers grounded in what's actually on screen and in the audio.
  6. Working directory. Printed at the end of the run. Not auto-cleaned today (see Roadmap) — rm -rf it manually when you're done with follow-ups.

Architecture

cue_kit/
├── cli.py             # arg parsing, mode dispatch
├── config.py          # env / .env loading, per-OS config dir
├── errors.py          # CueKitError (user-facing failures)
├── pipeline.py        # download → frames → transcript orchestrator
├── download.py        # yt-dlp wrapper, local file resolver
├── frames.py          # ffmpeg frame extraction, auto-fps budgeting
├── transcribe.py      # WebVTT parsing, dedup, range filtering
├── whisper.py         # Groq / OpenAI Whisper API clients (stdlib only)
├── slides.py          # scene-change detection + slide OCR (lecture-notes)
└── modes/
    ├── summary.py
    ├── transcript.py
    ├── training_doc.py
    └── lecture_notes.py

The pipeline is mode-agnostic: it always produces a PipelineResult (frames, transcript segments, metadata). Each mode is a renderer that turns that shared payload into its own output shape.

Use as a Claude Code skill

skill/SKILL.md is a thin wrapper that lets Claude Code drive cue-kit. Copy (or symlink) the skill/ directory to ~/.claude/skills/cue-kit/ (or a project's .claude/skills/cue-kit/) and Claude can invoke /cue-kit with the same modes. The cue-kit CLI must be installed and on PATH.

Roadmap

  • training-doc mode: section detection, step extraction, glossary
  • lecture-notes mode: scene-change keyframe selection, OCR over slides, transcript grouping
  • Optional vision-model captioning of frames (alternative to OCR)
  • Output formatters: PDF, DOCX
  • Zero-config install: detect missing ffmpeg / yt-dlp and print the exact brew / apt / winget command for the current OS
  • Skill-driven cleanup of the temp working directory after a run with no follow-ups
  • claude-video — explores enabling Claude to work directly with video inputs
  • OpenAI Whisper — a widely used foundation for speech-to-text transcription
  • Google Video AI — a broader take on extracting structured information from video
  • LangChain — tooling for building structured workflows around LLMs
  • NotebookLM — an example of turning raw content into structured, usable knowledge

cue-kit builds on similar ideas but focuses on turning video into structured, task-ready outputs for downstream use.

Contributing & security

See CONTRIBUTING.md for development setup and the release process, CHANGELOG.md for version history, and SECURITY.md to report a vulnerability privately.

License

Copyright © 2025–2026 AB//LABS.

cue-kit is dual-licensed:

  • AGPLv3 for open-source use — see LICENSE. Anyone running a modified version as a network service must release their changes under the same license.
  • Commercial license available for proprietary use, private modifications, or any case where the AGPL's terms don't fit. A starting-point template lives at COMMERCIAL-LICENSE-TEMPLATE.md. Contact AB//LABS through the GitHub organization.

The Claude Code skill wrapper in skill/ is separately licensed under the MIT License so it can be copied freely into your own Claude Code setup; it contains no cue-kit program code and only invokes the installed CLI.

Metadata

Release files for cue-kit 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for cue-kit 0.2.0
File Size Uploaded
cue_kit-0.2.0.tar.gz 44.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for cue-kit 0.2.0
File Interpreter ABI Platform
cue_kit-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 83.5 kB

Release files / cue_kit-0.2.0.tar.gz

Download URL cue_kit-0.2.0.tar.gz
Size 44.2 kB
Tags Source
SHA-256 checksum
How to use checksums
d3aa34a8519c21436ac2d0e6be323b47824f3dc3e0e9ab773d9d39740ffc11e9
BLAKE2b-256 checksum
How to use checksums
1ee9f935f6d8ca95ab64f2504eed61f2b8681ce3591034ad95c9a47ae27f7538
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 9, 2026.

Transparency log

Release files / cue_kit-0.2.0-py3-none-any.whl

Download URL cue_kit-0.2.0-py3-none-any.whl
Size 39.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
d10e361203f8296c20488053ea3e16644a51716ef2a8cabdeffecfadddbf8541
BLAKE2b-256 checksum
How to use checksums
a51a9e015216e56df01df198b8aee248b590fbfe55b5be14f2bbfacc7b67f820
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 9, 2026.

Transparency log

Release history Release notifications | RSS feed

0.3.0

2 release files

0.2.1

2 release files

This release

0.2.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page