macOS Computer Use MCP Server
146 MCP tools that give AI agents full control over macOS — screenshots, mouse, keyboard, window management, app automation, file system, OCR, and built-in app semantics (Calendar, Mail, Safari, Music, Messages, and more).
Built for Claude Code, LangGraph, Pi Agent, and any MCP-compatible agent framework.
Table of Contents
- Why
- Quick Start
- macOS Permissions (TCC)
- Usage
- AI Model Configuration
- Tool Reference
- Architecture
- Examples
- Requirements
- Development
- Troubleshooting
- Contributing
- License
Why
Most computer-use agents depend on third-party binaries (Playwright, Puppeteer) or cloud services. macos-computer-use-mcp runs directly on the OS:
| Benefit | Description |
|---|---|
| Auditable | Every tool call is plain Python + AppleScript — no black boxes |
| Extensible | Add your own tools by following the two-file pattern |
| Framework-agnostic | Works with any MCP client: Claude Code, VS Code, LangGraph, custom agents |
| macOS-native | Uses Quartz, Accessibility, IOKit, Vision — no external dependencies beyond pyobjc |
| High coverage | 146 tools across 3 layers, from raw pixels to semantic app control |
Quick Start
git clone https://github.com/yyyyyyyyiiiii/macos-computer-use-mcp.git
cd macos-computer-use-mcp
# Install with uv (recommended)
uv sync
# Or with pip
pip install -e .
Verify
# Check permissions
uv run python -c "from computer_use_mcp.darwin.tcc import check_all; print(check_all().report())"
# List all 146 tools
uv run python -c "
import asyncio
from computer_use_mcp.server import mcp
async def main():
tools = await mcp.list_tools()
print(f'{len(tools)} tools registered')
asyncio.run(main())
"
macOS Permissions (TCC)
Open System Settings → Privacy & Security and grant your terminal (or IDE):
| Permission | Required by | Why |
|---|---|---|
| Screen Recording | screenshot, region_screenshot, cursor_screenshot, display_list, screen_size |
Capture pixel data from display(s) |
| Accessibility | mouse_*, keyboard_*, window_*, ax_*, app_* |
Control mouse, keyboard, and inspect UI elements |
| Automation | calendar_*, reminders_*, notes_*, mail_*, messages_*, contacts_* |
AppleScript control of built-in apps |
| Full Disk Access | file_*, clipboard_* |
Read/write files in protected directories |
The server prints a clear status report on startup. Missing permissions do not prevent the server from running — affected tools simply return errors.
Usage
Claude Code
Add to your mcp.json or Claude Code settings:
{
"mcpServers": {
"macos-computer-use": {
"command": "uv",
"args": ["run", "macos-computer-use-mcp"]
}
}
}
Then ask Claude: "Take a screenshot, find the Safari window, and search for GitHub."
Other MCP Clients
{
"mcpServers": {
"macos-computer-use": {
"command": "python",
"args": ["-m", "computer_use_mcp"]
}
}
}
MCP Inspector
npx @anthropic-ai/mcp-inspector python -m computer_use_mcp
Opens a web UI at http://localhost:5173 where you can browse and call every tool interactively.
AI Model Configuration
This MCP server provides 146 macOS control tools — but a computer-use agent also needs an AI model to:
- See the screen (vision/GUI model) — understand what's on screen: windows, buttons, text, layout
- Decide what to do (text/action model) — plan the next step and generate the right tool call
Note: Some models (GPT-4o, Claude, Gemini) handle vision + text in a single model. Others (DeepSeek, early GPT-4) are text-only — you'll need a separate vision pipeline (screenshot → OCR → text description).
Model Options
| Provider | Model | Vision? | Set these env vars |
|---|---|---|---|
| DeepSeek | deepseek-chat (V3), deepseek-reasoner (R1) |
No (text-only) | DEEPSEEK_API_KEY |
| OpenAI | gpt-4o, gpt-4o-mini, o4-mini |
Yes (built-in) | OPENAI_API_KEY |
| Anthropic | claude-sonnet-5, claude-opus-5, claude-haiku-4-5 |
Yes (built-in) | ANTHROPIC_API_KEY |
| 智谱 GLM | glm-4v, glm-4v-flash, glm-4-plus |
Yes (built-in) | ZHIPU_API_KEY |
gemini-2.5-pro, gemini-2.5-flash |
Yes (built-in) | GOOGLE_API_KEY |
|
| Local / Ollama | llava, minicpm-v, qwen2.5-vl |
Yes | OLLAMA_HOST (optional) |
Setting API Keys
# DeepSeek — get key at https://platform.deepseek.com/api_keys
export DEEPSEEK_API_KEY="sk-your-key-here"
export DEEPSEEK_BASE_URL="https://api.deepseek.com" # or https://api.deepseek.com/v1
# OpenAI — get key at https://platform.openai.com/api-keys
export OPENAI_API_KEY="sk-your-key-here"
# Anthropic — get key at https://console.anthropic.com
export ANTHROPIC_API_KEY="sk-ant-your-key-here"
# 智谱 GLM — get key at https://open.bigmodel.cn
export ZHIPU_API_KEY="your-key-here"
# Google Gemini — get key at https://aistudio.google.com/apikey
export GOOGLE_API_KEY="your-key-here"
How to handle text-only models (DeepSeek, etc.)
If your chosen model doesn't support vision input, you need an extra pipeline step to convert screenshots into text descriptions before sending them to the model:
screenshot → OCR (built-in ocr_screenshot tool) → text description → text model → tool call
Example:
# Take screenshot + OCR to get text description of the screen
screen_text = await session.call_tool("ocr_screenshot", {})
# screen_text contains all recognized text with bounding boxes
# Send the text description to a text-only model
response = client.chat.completions.create(
model="deepseek-chat",
messages=[
{"role": "system", "content": "You control a Mac desktop."},
{"role": "user", "content": f"Screen contents:\n{screen_text}\n\nGoal: {goal}"}
]
)
Building a Complete Agent
Here's a minimal but runnable agent loop — screenshot → model → execute → repeat:
"""agent.py — minimal computer-use agent loop.
Requirements:
pip install openai mcp
Usage:
export OPENAI_API_KEY="sk-..."
python agent.py
"""
import asyncio, base64, json, os
from openai import OpenAI
from mcp.client import ClientSession
from mcp.client.stdio import stdio_client, StdioServerParameters
SYSTEM_PROMPT = """You are a macOS computer-use agent. You control a real Mac desktop.
For each step:
1. Look at the screenshot carefully — identify windows, buttons, text fields, cursors
2. Decide the SINGLE next action to progress toward the user's goal
3. Reply with ONLY a JSON tool call: {"tool": "...", "args": {...}}
Available tools (partial list — 146 total):
- screenshot, region_screenshot
- mouse_move {"x": int, "y": int}, mouse_click {}, double_click {}, right_click {}
- keyboard_type {"text": str}, hotkey {"keys": ["cmd", "c"]}
- scroll {"direction": "up"|"down", "amount": int}
- ocr_find_text {"text": str} — find text on screen, returns coordinates
- open_url {"url": str}
- safari_open_url {"url": str}, safari_search {"query": str}
- window_list, window_activate {"window_id": int}
- app_launch {"name": str}, app_quit {"name": str}
Reply {"done": true} only when the goal is fully accomplished."""
async def main():
goal = input("🎯 What should I do? ")
max_steps = int(input("Steps (default 15): ") or "15")
# Connect to MCP
server = StdioServerParameters(command="uv", args=["run", "macos-computer-use-mcp"])
client = OpenAI() # reads OPENAI_API_KEY from env
async with stdio_client(server) as (read, write):
async with ClientSession(read, write) as session:
messages = [{"role": "system", "content": SYSTEM_PROMPT}]
for step in range(max_steps):
# 1. See the screen
result = await session.call_tool("screenshot", {})
img_b64 = result.content[0].data
# 2. Ask the model
messages.append({
"role": "user",
"content": [
{"type": "text", "text": f"Step {step+1}/{max_steps}. Goal: {goal}"},
{"type": "image_url", "image_url": {"url": f"data:image/png;base64,{img_b64}"}}
]
})
response = client.chat.completions.create(
model=os.getenv("MODEL", "gpt-4o"),
messages=messages,
max_tokens=1024,
)
raw = response.choices[0].message.content.strip()
# Remove markdown fences if present
if raw.startswith("```"):
raw = raw.split("\n", 1)[1].rsplit("\n", 1)[0]
print(f"\n📍 Step {step+1}: {raw}")
# 3. Parse action
try:
action = json.loads(raw)
except json.JSONDecodeError:
print(" ⚠️ Could not parse, skipping")
messages.pop() # remove the user message, model will retry
continue
if action.get("done"):
print("✅ Done!")
break
# 4. Execute
tool_name = action["tool"]
tool_args = action.get("args", {})
result = await session.call_tool(tool_name, tool_args)
print(f" ↳ {tool_name} → ok")
messages.append({"role": "assistant", "content": raw})
print("🏁 Agent finished.")
if __name__ == "__main__":
asyncio.run(main())
Framework Integration
This MCP server works with any MCP-compatible agent framework:
LangGraph (using langchain-mcp):
from langchain_mcp import MCPToolkit
toolkit = MCPToolkit(command="uv", args=["run", "macos-computer-use-mcp"])
tools = await toolkit.get_tools()
# Use tools in your LangGraph agent graph...
Pi Agent: See Pi Agent MCP integration for setup.
Claude Code (zero-config):
// .mcp.json at project root
{
"mcpServers": {
"macos-computer-use": {
"command": "uv",
"args": ["run", "macos-computer-use-mcp"]
}
}
}
Claude Code handles the vision + tool-calling loop automatically — just ask in natural language.
Tool Reference
Layer Architecture
┌────────────────────────────────────────────────────────────┐
│ L3: App Semantics (68 tools) │
│ Calendar · Reminders · Notes · Mail · Messages · Contacts │
│ Finder · Safari · Music · Shortcuts · Settings │
├────────────────────────────────────────────────────────────┤
│ L2: Deterministic Tools (54 tools) │
│ Window Mgmt · App Lifecycle · AX Tree · OCR · Clipboard │
│ File System · System Info · Battery · WiFi · Bluetooth │
├────────────────────────────────────────────────────────────┤
│ L1: OS Primitives (24 tools) │
│ Screenshot · Mouse · Keyboard · Cursor · Input Source │
│ Timing · Display │
└────────────────────────────────────────────────────────────┘
L1 — OS Primitives
screenshot
Take a full-screen or region screenshot.
| Tool | Description |
|---|---|
screenshot |
Capture entire primary display (PNG base64) |
region_screenshot |
Capture rectangle (x, y, w, h) |
cursor_screenshot |
Capture a small region around the cursor |
screen_size |
Get display dimensions in pixels |
display_list |
Enumerate all connected displays |
mouse
Absolute-mouse positioning and buttons.
| Tool | Description |
|---|---|
mouse_move |
Move to absolute (x, y) |
mouse_click |
Left-click at current position |
double_click |
Double-click at current position |
right_click |
Right-click at current position |
mouse_drag |
Drag from current to (x, y) |
scroll |
Vertical scroll (positive = up) |
horizontal_scroll |
Horizontal scroll |
cursor
Cursor position and pixel inspection.
| Tool | Description |
|---|---|
cursor_get_position |
Get current (x, y) |
cursor_screenshot |
Screenshot the ~100×100 px area around cursor |
get_pixel_color |
Get RGB color at (x, y) |
keyboard
Text input and modifier keys.
| Tool | Description |
|---|---|
keyboard_type |
Type a string (supports Unicode) |
hotkey |
Press a key combination ("cmd+c", "cmd+shift+4") |
key_press |
Press and hold a single key |
key_release |
Release a single key |
input_source
Keyboard layout switching.
| Tool | Description |
|---|---|
input_source_get |
Get current input source |
input_source_list |
List all available input sources |
input_source_set |
Switch to a specific input source |
get_modifier_keys |
Get current modifier key states |
timing
| Tool | Description |
|---|---|
sleep |
Pause execution for N seconds |
timestamp |
Get current Unix timestamp (seconds or ms) |
L2 — Deterministic Tools
window
Window enumeration and manipulation via CGWindowList (Quartz).
| Tool | Description |
|---|---|
window_list |
List all visible windows with position/size/owner |
window_activate |
Bring a window to the foreground |
window_move |
Move a window to absolute (x, y) |
window_resize |
Resize to (width, height) |
window_close |
Close a window |
window_minimize |
Minimize a window |
get_frontmost_app |
Get the frontmost application name/bundle |
get_focused_element |
Get the currently focused UI element |
app
Application lifecycle via NSWorkspace + AppleScript.
| Tool | Description |
|---|---|
app_list_running |
List all running GUI applications |
app_launch |
Launch an application by name or bundle ID |
app_quit |
Gracefully quit an application |
app_force_quit |
Force-quit an application |
app_hide |
Hide an application (Cmd+H equivalent) |
ax_tree
Accessibility tree inspection and manipulation via AXUIElement (Quartz).
| Tool | Description |
|---|---|
ax_get_tree |
Get the AX tree for a window or the whole screen |
ax_get_element |
Get detailed attributes of a specific element |
ax_get_actions |
List available actions on an element |
ax_click_element |
Click an element via accessibility |
ax_set_value |
Set the value of a text field or slider |
ax_perform_action |
Perform a named action (e.g. "press", "confirm") |
ocr
Text recognition via macOS Vision framework.
| Tool | Description |
|---|---|
ocr_screenshot |
Take a screenshot and OCR the entire display |
ocr_region |
OCR a specific rectangle |
ocr_find_text |
Search for text on screen, return bounding boxes |
ocr_get_text_at |
Get the text at a specific pixel position |
clipboard
Clipboard read/write via pbcopy/pbpaste + NSImage (AppKit).
| Tool | Description |
|---|---|
clipboard_get |
Get clipboard content (text, image, or file list) |
clipboard_set_text |
Set clipboard to plain text |
clipboard_set_image |
Set clipboard to image from file |
clipboard_clear |
Clear all clipboard contents |
file
File-system operations via Python pathlib/shutil.
| Tool | Description |
|---|---|
file_list_dir |
List directory contents |
file_exists |
Check if a path exists |
file_read |
Read file content (text or binary base64) |
file_write |
Write content to a file |
file_delete |
Delete a file or directory (recursive) |
file_move |
Move/rename a file or directory |
file_copy |
Copy a file or directory |
file_mkdir |
Create a directory (with parents=True) |
file_get_info |
Get file metadata (size, mtime, permissions) |
file_search |
Recursive file search with glob patterns |
file_get_home_dir |
Get the user home directory path |
file_get_desktop_dir |
Get the Desktop directory path |
file_get_downloads_dir |
Get the Downloads directory path |
system
System information and control via IOKit, system_profiler, pmset, networksetup.
| Tool | Description |
|---|---|
system_info |
Hostname, OS version, CPU, memory, disk |
battery_info |
Battery percent, charging, health, cycle count |
get_volume |
Get system output volume (0–100) |
set_volume |
Set system output volume |
get_brightness |
Get built-in display brightness (0.0–1.0) |
set_brightness |
Set built-in display brightness |
get_dark_mode |
Check if dark mode is active |
wifi_info |
SSID, BSSID, channel, RSSI, IP address |
bluetooth_info |
Power state and connected devices |
sleep_display |
Put all displays to sleep |
lock_screen |
Lock the screen (password required to unlock) |
open_url |
Open a URL in the default browser or specified app |
reveal_in_finder |
Reveal a file/folder in Finder |
run_command |
Execute a shell command (local trusted sessions) |
L3 — App Semantics
All L3 tools use AppleScript targeting macOS built-in applications. Apps that are not running will be launched automatically by AppleScript.
calendar (Calendar.app)
| Tool | Description |
|---|---|
calendar_list |
List upcoming events (default 7 days) |
calendar_create |
Create a new event with title, date, location, notes |
calendar_delete |
Delete an event by UID |
reminders (Reminders.app)
| Tool | Description |
|---|---|
reminders_list |
List reminders (by list, with filters) |
reminders_create |
Create a reminder with title, due date, priority |
reminders_complete |
Mark a reminder as completed |
reminders_delete |
Delete a reminder |
notes (Notes.app)
| Tool | Description |
|---|---|
notes_list |
List notes across all folders (with search) |
notes_create |
Create a note with title and body |
notes_get |
Get full note content by ID or name |
mail (Mail.app)
| Tool | Description |
|---|---|
mail_list |
List recent emails with optional filters |
mail_send |
Compose and send an email (to, cc, bcc) |
messages (Messages.app)
| Tool | Description |
|---|---|
messages_list_conversations |
List recent conversations with unread counts |
messages_get |
Get messages from a conversation (by chat_id or contact) |
messages_send |
Send an iMessage/SMS (text and/or attachment) |
messages_search |
Search all conversations by text |
messages_mark_read |
Mark a conversation as read |
messages_delete_conversation |
Delete an entire conversation |
messages_get_attachment |
Save attachments from a conversation to disk |
contacts (Contacts.app)
| Tool | Description |
|---|---|
contacts_list |
List contacts (by group, up to 500) |
contacts_search |
Search contacts by name, email, phone, org |
contacts_get |
Get full details of a specific contact |
contacts_create |
Create a new contact (name, org, email, phone) |
contacts_update |
Update an existing contact |
contacts_delete |
Delete a contact |
contacts_export_vcard |
Export contacts as .vcf file |
finder (Finder.app)
| Tool | Description |
|---|---|
finder_get_selection |
Get currently selected items |
finder_select |
Select files/folders by path |
finder_get_windows |
List all open Finder windows with target paths |
finder_get_current_folder |
Get the frontmost Finder window's folder |
finder_navigate |
Open a folder in Finder |
finder_get_info |
Get detailed Finder metadata for a file |
finder_duplicate |
Duplicate a file/folder (Cmd+D) |
finder_make_alias |
Create a Finder alias |
finder_eject_volume |
Eject a mounted disk by name |
finder_empty_trash |
Empty the Trash (irreversible) |
finder_list_disks |
List all mounted volumes with capacity/free space |
safari (Safari.app)
| Tool | Description |
|---|---|
safari_list_tabs |
List all open tabs across all windows |
safari_get_current_tab |
Get the active tab's URL and title |
safari_open_url |
Open a URL (new tab or window) |
safari_close_tab |
Close a specific tab |
safari_search |
Search the web using the default search engine |
safari_go_back |
Navigate back |
safari_go_forward |
Navigate forward |
safari_get_bookmarks |
List all bookmarks |
safari_add_bookmark |
Add a bookmark |
safari_execute_javascript |
Execute JavaScript in the current tab |
music (Music.app)
| Tool | Description |
|---|---|
music_get_state |
Get player state + current track info |
music_play |
Start playback |
music_pause |
Pause playback |
music_playpause |
Toggle play/pause |
music_next |
Skip to next track |
music_previous |
Go to previous track |
music_search |
Search library by name/artist/album |
music_get_playlists |
List all playlists with track counts |
music_play_playlist |
Play a specific playlist by name |
music_set_volume |
Set Music.app volume (0–100) |
shortcuts (Shortcuts.app)
| Tool | Description |
|---|---|
shortcuts_list |
List all shortcuts (with folders and colors) |
shortcuts_run |
Run a shortcut by name (optional text input) |
shortcuts_run_with_input |
Run a shortcut with file or text input |
shortcuts_get_info |
Get shortcut metadata (action count, subtitle, icon) |
shortcuts_list_folders |
List shortcut folders with item counts |
settings (System Settings)
| Tool | Description |
|---|---|
settings_open_pane |
Open a specific Settings pane (WiFi, Bluetooth, etc.) |
settings_get_wallpaper |
Get current desktop wallpaper path(s) |
settings_set_wallpaper |
Set desktop wallpaper from an image file |
settings_get_display |
Get display resolution, refresh rate, scaling |
settings_get_sound |
Get audio input/output device and volume |
settings_get_general |
Get appearance, accent color, sidebar size, Handoff |
Architecture
src/computer_use_mcp/
├── server.py # MCP entry point (stdio transport)
├── __init__.py # Version, package metadata
│
├── darwin/ # macOS-specific implementations (no MCP dependency)
│ ├── cg_screen.py # CGDisplay / CGImage screenshot capture
│ ├── cg_input.py # CGEvent mouse + keyboard injection
│ ├── cg_keyboard.py # Text synthesis + key-code mapping
│ ├── ax_window.py # CGWindowList + AXUIElement window ops
│ ├── ax_tree.py # Accessibility tree walker (200+ lines)
│ ├── clipboard.py # pbcopy/pbpaste + NSImage clipboard
│ ├── ocr.py # VNRecognizeTextRequest (Vision framework)
│ ├── file_ops.py # Pure-Python pathlib/shutil file operations
│ ├── tcc.py # TCC permission checker (tccutil + osascript)
│ ├── system.py # IOKit brightness, pmset, networksetup, etc.
│ ├── calendar.py # Calendar.app AppleScript
│ ├── reminders.py # Reminders.app AppleScript
│ ├── notes.py # Notes.app AppleScript
│ ├── mail.py # Mail.app AppleScript
│ ├── messages.py # Messages.app AppleScript
│ ├── contacts.py # Contacts.app AppleScript
│ ├── finder.py # Finder.app AppleScript
│ ├── safari.py # Safari.app AppleScript
│ ├── music.py # Music.app AppleScript
│ ├── shortcuts.py # Shortcuts CLI + AppleScript
│ └── settings.py # System Settings + defaults CLI
│
├── tools/ # MCP tool registration layer (thin wrappers)
│ ├── screen.py # @mcp.tool() async def screenshot()
│ ├── mouse.py # ... 21 more modules
│ ├── cursor.py # (each module has a register(mcp) entry point)
│ ├── keyboard.py
│ ├── input_source.py
│ ├── timing.py
│ ├── window.py
│ ├── app.py
│ ├── ax_tree.py
│ ├── ocr.py
│ ├── clipboard.py
│ ├── file.py
│ ├── system.py
│ ├── calendar.py
│ ├── reminders.py
│ ├── notes.py
│ ├── mail.py
│ ├── messages.py
│ ├── contacts.py
│ ├── finder.py
│ ├── safari.py
│ ├── music.py
│ ├── shortcuts.py
│ └── settings.py
│
└── tests/ # One test file per domain (30+ files)
├── test_server.py # Verifies all 146 tools are registered
├── test_screen.py
├── test_mouse.py
└── ... (28 more)
Design principles:
-
Two-layer separation:
darwin/modules contain pure macOS logic with zero MCP dependency.tools/modules are thin MCP wrappers. This means you can reuse thedarwin/modules in a non-MCP agent. -
Each tool returns a plain
dict— MCP serializes them natively. No Pydantic models, no custom types. -
AppleScript continuation: Long lines use
¬(option-return) to stay under the 100-character line limit. -
Applescript string escaping: Backslashes →
\\\\, double-quotes →\\"before interpolation into AppleScript strings.
Examples
Screenshot + OCR → Click
# In your agent's tool-calling loop:
screenshot = await client.call_tool("screenshot")
# Feed to vision model...
text = await client.call_tool("ocr_find_text", {"text": "Submit"})
if text["found"]:
x, y = text["bounds"]["x"] + text["bounds"]["w"] // 2
text["bounds"]["y"] + text["bounds"]["h"] // 2
await client.call_tool("mouse_move", {"x": x, "y": y})
await client.call_tool("mouse_click", {})
Safari automation
await client.call_tool("safari_open_url", {"url": "https://github.com"})
await client.call_tool("safari_search", {"query": "macOS automation"})
tabs = await client.call_tool("safari_list_tabs")
# → {"tabs": [{"title": "...", "url": "...", ...}], "count": 5}
Full agent loop (pseudocode)
from mcp.client import ClientSession
async with ClientSession(stdio_transport) as session:
while True:
# 1. See the screen
screen = await session.call_tool("screenshot")
# 2. Vision model decides the next action
action = vision_model.decide(screen, goal)
# 3. Execute with MCP tools
result = await session.call_tool(action.tool, action.params)
# 4. Verify
if action.done:
break
Requirements
| Requirement | Minimum | Recommended |
|---|---|---|
| macOS | 13 Ventura | 14 Sonoma+ |
| Python | 3.12 | 3.12+ |
| RAM | 2 GB | 4 GB+ |
| Disk | ~50 MB | — |
macOS Compatibility: Core L1/L2 tools work from macOS 10.13+ (High Sierra). L3 app-semantics tools require 13+ for full System Settings support. OCR requires 10.13+ (Vision framework). See the full compatibility table.
Development
# Clone and install
git clone https://github.com/yyyyyyyyiiiii/macos-computer-use-mcp.git
cd macos-computer-use-mcp
uv sync --all-extras
# Lint
uv run ruff check src tests
# Run all tests (requires macOS + permissions)
uv run pytest -q
# Run a subset
uv run pytest tests/test_server.py tests/test_safari.py -v
# Start the server locally
uv run python -m computer_use_mcp
Test conventions
- Tests are macOS-only:
pytestmark = pytest.mark.skipif(sys.platform != "darwin", reason="...") - L1 tests (screenshot, mouse, keyboard) require Screen Recording + Accessibility permissions
- L3 tests (Calendar, Reminders, etc.) require Automation permissions
- Input validation tests (empty strings, etc.) pass without any permissions
Adding a new tool
- Implement the macOS logic in
src/computer_use_mcp/darwin/<module>.py - Register the MCP wrapper in
src/computer_use_mcp/tools/<module>.py - Add the import to
server.pyand the module to_MODULES - Write a test in
tests/test_<module>.py - Run
uv run ruff check src tests && uv run pytest -q
Troubleshooting
Safari: open_url AppleScript fails
Symptom: safari_open_url returns error -10024 ("can't create or move element into container").
Cause: Safari AppleScript permissions or window state (Stage Manager, minimized windows).
Workarounds:
- Use
open_url(L2 tool) instead — it uses theopenCLI command, which is more reliable - Use
keyboard_type+hotkey(["cmd", "l"])to type URLs directly in Safari's address bar - Close and reopen Safari, then retry
Safari: execute_javascript fails
Symptom: safari_execute_javascript returns an error about "Allow JavaScript from Apple Events".
Fix: Open Safari → Develop menu → Settings → Advanced → check "Allow JavaScript from Apple Events".
If you don't see the Develop menu: Safari → Settings → Advanced → check "Show Develop menu in menu bar".
Window resize/move returns success: false
Cause: Some windows (especially Safari in Stage Manager) reject programmatic resize/move.
Workarounds:
- Disable Stage Manager temporarily
- Use
hotkey(["cmd", "shift", "f"])to toggle fullscreen - Use accessibility (
ax_*) tools as an alternative click path
Screenshots too large for model context
Symptom: Screenshot data URLs exceed model token limits.
Solutions:
- Use
region_screenshotto capture only the relevant area - Use
cursor_screenshotfor a 60×60px region around the cursor - Use
ocr_screenshotto get text-only screen descriptions (much smaller than images) - Resize screenshots before sending:
PIL.Image.open(...).resize((1280, 800))
Model doesn't understand what's on screen
Solution: Use the built-in OCR tools before sending to the model:
# Get screen text as structured data
ocr = await session.call_tool("ocr_screenshot", {})
# Append OCR results to your model prompt for better grounding
prompt = f"Screen text visible:\n{ocr['text']}\n\nGoal: {goal}"
MCP server not found
Symptom: Claude Code shows "No MCP servers configured" or tools are unavailable.
Checklist:
.mcp.jsonmust be at the project root (not in a subfolder orsrc/)- Run
uv syncfirst to install dependencies - Verify the server starts:
uv run python -m computer_use_mcp - Restart Claude Code after creating
.mcp.json
Contributing
Contributions are welcome! See CONTRIBUTING.md for the full guide.
- Tool requests: Open an issue with the app name and desired operations
- Bug reports: Include macOS version + error output
- Pull requests: Follow the two-layer pattern, include tests
License
MIT — see LICENSE for full text.
Built with ❤️ for the macOS agent ecosystem
Metadata
Release files for macos-computer-use-mcp 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| macos_computer_use_mcp-0.2.0.tar.gz | 116.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| macos_computer_use_mcp-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 227.9 kB
Release files / macos_computer_use_mcp-0.2.0.tar.gz
| Download URL | macos_computer_use_mcp-0.2.0.tar.gz |
|---|---|
| Size | 116.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
0afc7503b4bb3ffc222b14128581ef9c3b333f65a7aad48e44f273d49cf72c29
|
|
BLAKE2b-256 checksum How to use checksums |
2c47727bf234253a73531f96a4a48813ae8aa73e0114c698bdcf6981a7fad6a7
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.11.28 {"installer":{"name":"uv","version":"0.11.28","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|
Release files / macos_computer_use_mcp-0.2.0-py3-none-any.whl
| Download URL | macos_computer_use_mcp-0.2.0-py3-none-any.whl |
|---|---|
| Size | 111.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
7dee7f9ce682b9ee328e786ffc112efb5c66deb762acefc717dd5f426e3ef682
|
|
BLAKE2b-256 checksum How to use checksums |
973f7f2751156c1fd95aec5916bc3f8a88382c1a01bc56762e8ef5915c577d70
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.11.28 {"installer":{"name":"uv","version":"0.11.28","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|