doppelhand
A double of you at the keyboard. doppelhand captures the Windows screen and drives the mouse and keyboard: clicking, dragging, scrolling and typing the way a person would.
Use it two ways. Give an AI agent you already have eyes and hands, so it looks at your desktop with its own model and acts through doppelhand. Or let doppelhand drive itself with Claude, using its own key.
Install
pip install doppelhand
Windows only, Python 3.10 or newer. Pillow is the only dependency; screen capture and input go straight through the Win32 API.
This program moves your real mouse and types on your real keyboard. It can click
anything you can click. Run it on a machine where that is acceptable and watch what it
does. doppelhand run waits for confirmation and stops when you hold Escape; the
single-action commands do what they are told, immediately.
Give your agent hands
Install the skill into whichever harness you use, and it learns the loop by itself:
doppelhand install-skill claude-code
doppelhand install-skill opencode
doppelhand install-skill agy # global, ~/.gemini/config/skills
doppelhand install-skill agy-workspace # this project's .agents/skills
doppelhand install-skill --dest path/to/skills/doppelhand # anything else
doppelhand install-skill --print # read it first
No API key is involved. The agent takes a screenshot, reads the PNG with its own image tool, and calls back with coordinates. Every command answers with one JSON object and exits non-zero when the action did not happen.
$ doppelhand shot
{"ok": true, "path": "C:\\Users\\you\\AppData\\Local\\doppelhand\\shot.png",
"view": [1280, 720], "display": [1920, 1080], "scale": 0.666667, "space": "view"}
$ doppelhand click 640,360
{"ok": true, "action": "left_click", "at": [640, 360]}
| Command | What it does |
|---|---|
shot [PATH] [--region X,Y,W,H] [--monitor N] |
capture, write a PNG, report the coordinate space |
screen |
list the displays and report the coordinate space |
click X,Y [--button] [--count] [--modifiers] |
click, double click, right click |
move X,Y / drag X1,Y1 X2,Y2 |
move the pointer, or press and drag |
scroll up|down|left|right [N] [--at X,Y] |
scroll |
type "text" / key "ctrl+s" [--repeat N] |
type text, press a combination |
hold "shift" 2 / wait 1.5 |
hold a key, or pause |
cursor |
where the pointer is, in both spaces |
Coordinates are the pixels of the screenshot, not of the display; doppelhand scales them back up. A point outside that view is refused rather than clamped, so a scale mistake fails loudly instead of clicking a corner. Pass --space display if you already have real display pixels.
Every display is addressable. screen lists them numbered left to right, --monitor N picks one, and --monitor all captures the lot as a single wide image. Each display has its own view space starting at 0,0, so coordinates never go negative. The display and size of the last shot are remembered, so the next command lands in the same space without repeating flags; rearranging or unplugging a monitor discards that memory rather than misplacing a click.
Screenshots include the mouse pointer, which a GDI capture leaves out by default. Turn it off with shot --no-cursor.
Real time
A fresh process costs about 300ms before it does anything, which dominates a session of many small actions. Run a server once and talk to it over loopback instead:
doppelhand serve
It prints a port and a secret, and writes both to %LOCALAPPDATA%\doppelhand\serve.json.
Reads are GET, anything that moves the pointer or types is POST, and arguments keep the
names the command line uses:
curl -sH "X-Doppelhand-Token: $TOKEN" "http://127.0.0.1:$PORT/shot?fast=1"
curl -sXPOST -H "X-Doppelhand-Token: $TOKEN" "http://127.0.0.1:$PORT/click?at=640,360"
One run of python bench/benchmark.py on a three-monitor desktop. The spawned column
moves a lot with what the machine is doing, so re-measure rather than trusting these:
| action | spawned | served | |
|---|---|---|---|
cursor |
309 ms | 5.6 ms | 55× |
screen |
262 ms | 5.3 ms | 49× |
move |
262 ms | 5.9 ms | 44× |
shot |
430 ms | 77 ms | 5.5× |
shot --fast |
378 ms | 30 ms | 12× |
Most of the screenshot gain comes from capturing through DXGI Desktop Duplication, which
hands over the frame the compositor already holds. Opening one costs more than a whole
GDI capture, so it only pays inside the server; one-shot commands keep the GDI path, and
so does any display that refuses to be duplicated. Every shot reports which it used in
source.
The server binds to loopback only, needs the secret from that file, and turns away any request that carries browser headers or names a host other than loopback.
Or let it drive itself
This is the only path that calls the Anthropic API, the only one that needs a key, and the only one that needs the SDK:
pip install "doppelhand[run]"
export ANTHROPIC_API_KEY=sk-ant-...
doppelhand run "open the calculator and work out 19% of 4,320"
This loop is covered by tests against a scripted client rather than a recorded live run, so treat it as the least proven part of the package.
run prints what it is about to do and waits for confirmation. Hold Escape at any point to stop the run — the key is checked before every action.
--max-steps N model turns before giving up (default 30)
--max-edge N long edge of the screenshots sent to the model (default 1280)
--model NAME defaults to claude-opus-5
-y skip the confirmation
-q print only the final answer
As a library:
from doppelhand.agent import Agent
print(Agent(max_steps=10).run("close the notification in the corner"))
How it works
screen.pycaptures a display, through GDI for a one-shot command or through DXGI Desktop Duplication when the server is holding one open, and hands back a Pillow image.executor.pyshrinks that image to a 1280-pixel long edge, and scales every coordinate the model returns back up to real pixels.inputs.pydrives the mouse and keyboard withSendInput. Text is typed as Unicode, so it does not depend on the active keyboard layout.cli.pyexposes each of those actions as a command, oragent.pyruns the whole loop against Claude: send the screen, execute the actions in the reply, send the result, repeat.
In the built-in loop, costs are kept down two ways: one prompt-cache breakpoint moves along with the newest tool result, and screenshots older than the last ten are dropped from the history in batches.
Limits
- Actions are run against a display as a whole; there is no per-window targeting.
- Displays are addressed one at a time.
--monitor allcaptures them together but scales the result down too far to read. - A rotated display falls back to the slower GDI capture, because a duplicated frame arrives in the screen's unrotated orientation.
- Windows blocks synthetic input to windows running as administrator, so a click on one does nothing.
- The Escape stop applies to
run, which checks it before each action. Single commands have already finished by the time you could press anything. - The built-in loop is covered by tests against a scripted client, not by a recorded live run.
Development
pip install -e ".[dev]"
pytest tests -q
python bench/benchmark.py
License
MIT
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file doppelhand-1.3.1.tar.gz.
File metadata
- Download URL: doppelhand-1.3.1.tar.gz
- Upload date:
- Size: 54.6 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.11.4
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
b2c4ad066a6b14e18b5522be77a4e31815660f516d7c87a75c4a05b036a550f8
|
|
| MD5 |
ef4da6c293be4f99743efcd72f89226e
|
|
| BLAKE2b-256 |
1dd485876079356e11db160d29787aafaf6f21483fc83dac7892b3aa6360c75b
|
File details
Details for the file doppelhand-1.3.1-py3-none-any.whl.
File metadata
- Download URL: doppelhand-1.3.1-py3-none-any.whl
- Upload date:
- Size: 41.9 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.11.4
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
642e0b6c240f0ac1c10464a30ed4ff0eb033030c76906f4e61cc3f0eed93d033
|
|
| MD5 |
b96e92595d0ee61eb01235764eb12412
|
|
| BLAKE2b-256 |
e7e3491da406a032fc339da43c9297bd13025e20b4b9f2e9f56dedf7c866ca04
|