Skip to main content
Watch Skill: a pixel-art scene of the Watch Skill mascot watching a screen. A filmstrip above shows the four stages — watch a source, remember it as OCR and transcript, resolve timestamped evidence, then run THE LOOP to critique and fix. The screen shows a video library, an evidence list with timestamps, and a capture-critique-fix-verify cycle ending in VERIFIED.

Watch Skill · DeepWatch

Give an agent eyes and ears — and a record of its work that something other than the agent wrote.

CI Workspace PyPI License


Two things, and how they fit

Watch Skill is the engine. It turns video, audio and screen activity into searchable, timestamped evidence, and it runs deterministic verification contracts whose verdicts do not come from a language model. Any agent can use it — through MCP, a CLI, or a REST API.

DeepWatch is the workspace. It composes the official DeepSeek Harness with Watch Skill so an agent's work happens inside something that watches it: every tool call gets a receipt, every path a tool declares is checked against one workspace boundary, and "it worked" is a claim you can open.

The word declares is load-bearing. A write names the file it is writing and that name is checked before the call runs. A shell command is a string this does not parse, so what is enforced there is the sandbox the Harness itself runs the shell under, and what is recorded is that a shell call happened with no path of its own. Both are in the receipt; they are not the same guarantee, and reading them as one is the mistake this paragraph exists to prevent.

One sentence each: Watch Skill is what sees and proves. DeepWatch is where the work happens. You can use either on its own.


What it actually looks like

A DeepWatch session. The agent was asked to create a file and read back its total. Write, Read and Pwsh rows are shown, each naming a workspace-relative path, and the answer confirms the file contents and the calculated total.

An ordinary request — create totals.json and tell me the sum. Nobody mentioned Watch. Every row is a receipt, every path is workspace-relative, and the total was read back from the file rather than remembered.

Then the part that matters:

A VERIFIED result card from watch_verify: two of two checks passed, one confirming the file exists and one confirming its total field equals 60, with the contract's sha256 digest.

watch_verify ran a frozen contract and Watch Core answered. The agent did not grade itself: a check either passed or it did not, the contract's digest is on screen, and the same contract run from a different directory fails.

Every image here is a photograph of a room built only from sealed artifacts, with a provider bound and Watch Core running over stdio. Nothing is seeded or retouched. Each caption on the screenshot page names the build it was taken from, because more than one candidate was photographed on the way here and saying "this release" of all of them would not have been true.


Start here

I have an agent already → Watch Skill

On PyPI, one version behind. watch-skill is published and installs today; the newest release on PyPI is 1.2.0. The 1.4.0 this page describes is published by the core-v1.4.0 release.

pip install 'watch-skill[standard]'   # frames, retrieval and the MCP server
watch-skill doctor                    # checks, and repairs what it can
watch-skill watch <video-url-or-file>
watch-skill ask <id> "what changed at 3:12?"

Take the extra seriously. A bare pip install watch-skill gives you the CLI, the verifier and the Bridge, and it cannot extract a frame: watch stops at perceive.missing_dependency on the first video. [standard] is frames, retrieval and MCP; add [ocr] to read on-screen text, [whisper] for local transcription when a source has no captions, [loop] for the browser, or take [all]. watch-skill doctor names the exact command for whatever is missing.

Wire it into any MCP client ([standard] includes this):

watch-skill serve              # stdio MCP server, 39 tools

Or install the skills into 25+ agents at once — Claude Code, Cursor, Codex, Copilot, Gemini CLI, Cline, Zed and more:

npx skills add oxbshw/watch-skill -g

Per-client setup: docs/agents/.

I want the workspace → DeepWatch

Not on npm yet. Nothing exists under the @deepwatch scope until the deepwatch-v0.1.0 release publishes it. Until then the command below resolves nothing, and getting started has the path that works from a checkout.

npm install -g @deepwatch/cli   # pending the deepwatch-v0.1.0 release
deepwatch setup                 # builds the runtime and composes the profile
deepwatch web --workspace ./my-project

deepwatch doctor reports what is installed and what is missing; deepwatch setup is what builds things, and it asks before downloading anything.

Node ≥ 22.19 and Python 3.11, 3.12 or 3.13 — the versions CI runs and the classifiers declare. Windows, macOS and Linux.


THE LOOP: observe, act, verify

The cycle in the picture at the top, on a real page:

pip install 'watch-skill[standard,loop]' && playwright install chromium

watch-skill loop start http://localhost:3000/checkout \
  "the total updates when quantity changes, and no NaN appears"
  1. Observe — a real browser records the page to video; frames are extracted and OCR'd, each with an absolute timestamp.
  2. Critique — a vision model is asked whether the capture meets the criteria you wrote. It reports issues with the timestamp each was seen at.
  3. Fix — you change the code.
  4. Verify — watch-skill loop iterate re-captures and diffs against the previous run, so "fixed" means the thing that was wrong is gone.

The critique step needs a vision-capable model. Without one, capture, frames, OCR and verification still work and the critique says it cannot judge rather than guessing. See THE LOOP.


What people use it for

Ask a video a question Index a recording once, then ask about it. Answers cite timestamps you can open. 01-watch-and-ask
Prove an agent's work A deterministic contract Core runs — file digests, JSON values, SQL, HTTP, DOM. 14-browser-verification
Fix a UI by looking at it Capture, critique, fix, re-verify. 04-ui-loop
Search across everything One index over every source you have watched. 03-cross-video-search
Work offline Local whisper and OCR, no provider, nothing leaves the machine. 15-private-offline-workflow
Watch something live A stream or a browser session, bounded and cursored. 18-live-watch

Each one is a directory you can run, with its prerequisites and expected output written next to it.

Learn the core 01 Watch and ask · 02 Focused moment · 03 Cross-video search
Build with agents 06 MCP and REST · 09 Framework adapters · 15 Private offline workflow
Understand and organise 05 Multilingual Arabic · 10 Structured extraction · 11 Batch mode · 12 Library memory · 16 Shareable viewer
Verify and improve 04 UI loop · 07 Lessons and stats · 08 Loop types · 13 Self-improvement · 14 Browser verification · 17 Freshness and offline · 20 Observer loop
Watch live 18 Live watch · 19 Live browser

That is all 20 examples; the index is examples/.


How it fits together

flowchart LR
  subgraph W["DeepWatch workspace"]
    H["DeepSeek Harness<br/>agent, tools, UI"]
    P["Watch plugins<br/>tools · library · live · memory"]
    H <--> P
  end
  P <-->|"Bridge (stdio)"| C["Watch Core<br/>Python engine"]
  C --> E[("Evidence store<br/>frames · transcripts · index")]
  C --> V["Verifier<br/>isolated subprocess"]
  V --> R[("Verification records<br/>contract · checks · verdict")]
  P --> J[("Receipt journal<br/>one per tool call")]
  A["Any other agent<br/>MCP · CLI · REST"] <--> C

Watch Core is the only thing that issues a verdict. The Host may notice, correlate, freeze a contract and ask — it may not decide the answer. That is ADR-002, and a build gate fails if anything under packages/ starts producing verdicts.

More: architecture · verification.


What works, and what it needs

Capability Out of the box Needs
Start the app, browse, read diagnostics ✅ nothing
Verification contracts, containment, receipts ✅ nothing
Video frames and scenes with [standard] ffmpeg ≥ 5.1 — watch-skill doctor installs it
Reading on-screen text with [ocr] a first-use model download (~80 MB)
Speech to text with [whisper] a first-use model download; captions are used first when a source has them
Chat with an agent — a provider you add and bind
Visual scene description — a model that can see images
Browser capture / THE LOOP with [loop] playwright install chromium
Memory off enable in Settings; the store is plaintext and says so
Desktop app not distributed run the web workspace

On providers, and three things that are not the same.

DeepWatch starts, and stays useful, with no provider configured: verification, containment, the Library and local perception are all local. What needs a provider is the agent — chat, tool use, and the critique step of THE LOOP.

The three ways a capability gets added here are genuinely different, and the product does not pretend otherwise.

  • A local dependency — ffmpeg, a JS runtime, yt-dlp — runs on your machine, costs disk, and watch-skill doctor will fetch and repair it.
  • A downloaded model — OCR weights, whisper — also runs on your machine, is a large one-time download, and is slower and less capable than a hosted model of the same kind. Nothing about your files leaves the machine.
  • A hosted provider — the agent's model, and any vision model you bind — is somebody else's service, with their latency, their price and their terms, and it sees what you send it.

An OpenAI-compatible server you run yourself (Ollama, vLLM, LM Studio, llama.cpp) is a hosted route pointed at your own hardware: it keeps the data local and keeps the caveat, because a small local model may not support tool calls or images at all, and DeepWatch will report that rather than work around it. Nothing reaches a provider until you add one, and holding a provider credential is not permission to upload a frame or a transcript — that is a separate consent.

What repairs itself, and what does not. watch-skill doctor repairs dependencies: it downloads yt-dlp and keeps it current, bootstraps a JS runtime, installs OCR language data, and fetches ffmpeg where it can. That is deliberate and it is the only thing here that fixes itself. Nothing resumes a task on its own, nothing learns between runs, and nothing is encrypted at rest in this release. Known limitations is the full list.


Measured, not asserted

Against a leading video-understanding API, same files, same scorer:

Watch Skill Baseline
Written-analysis groundedness 89.7% 27.9%
Citations per 100 words 13.23 0.12
Frame delivery on real footage 96.9% 31.2%
Cue starts within half a second 100% 25%

Method and fixtures: benchmarks/video_backends/.


Documentation

Getting started Install, first watch, first agent connection
Install and upgrade Both products, optional extras, compatibility policy
Tool reference All 39 MCP tools and their REST/CLI counterparts
Verification Contracts, the fourteen check types, assurance levels
Architecture Boundaries, data flow, extension points
Agent matrix Per-client setup and how far each is verified
Troubleshooting Dependency repair and common runtime errors
Comparison Honest trade-offs against the alternatives
Ecosystem Where this project appears, and which of it is coverage
Known limitations What this release does not do
DeepWatch workspace README · setup · releasing

Three tool counts, because they answer different questions: 39 MCP tools from watch-skill serve, 22 watch_* tools added to an agent inside DeepWatch, 47 tools that agent is offered in total.


Contributing

Issues and pull requests welcome — CONTRIBUTING.md has the twenty-minute path. Security policy: SECURITY.md.

Built on DeepSeek Harness · Powered by Watch Skill. An independent project, not affiliated with or endorsed by DeepSeek.

Metadata

Release files for watch-skill 1.4.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for watch-skill 1.4.0
File Size Uploaded
watch_skill-1.4.0.tar.gz 5.7 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for watch-skill 1.4.0
File Interpreter ABI Platform
watch_skill-1.4.0-py3-none-any.whl Python 3 none any Details

Total release size: 6.6 MB

Release files / watch_skill-1.4.0.tar.gz

Download URL watch_skill-1.4.0.tar.gz
Size 5.7 MB
Tags Source
SHA-256 checksum
How to use checksums
78eea55950bbd4a9e8208f5874fc680dfa4f9bde690306bc8a0a96998ddfeb00
BLAKE2b-256 checksum
How to use checksums
d0587fc5b6f852187d2ce117bcdb83739a5abdf21f469192b5e2a7602b0541ae
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 6, 2026.

Transparency log

Release files / watch_skill-1.4.0-py3-none-any.whl

Download URL watch_skill-1.4.0-py3-none-any.whl
Size 969.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
ce032b090ba0f8f95a4254e57e1fd2af5f324e550d9a9263af56909ff683a577
BLAKE2b-256 checksum
How to use checksums
434b29d68e03cebb7630a363caa8c02cb8e10ec4e05790604d454ccc2b757cc0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 6, 2026.

Transparency log

Release history Release notifications | RSS feed

1.4.3

2 release files

1.4.2

2 release files

1.4.1

2 release files

This release

1.4.0 This release

2 release files

1.2.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page