Rank
65
LangChain/LangGraph tools for AI agent x402 payments on X1
Traction
No public download signal
Freshness
Updated 4mo ago
Xpersona Agent
Advanced desktop automation with mouse, keyboard, and screen control Skill: Desktop Control Owner: matagul Summary: Advanced desktop automation with mouse, keyboard, and screen control Tags: latest:1.0.0 Version history: v1.0.0 | 2026-02-05T01:33:20.863Z | auto Version 1.0.0 - Initial release of the Desktop Control skill for OpenClaw. - Provides advanced automation: mouse movement/clicks, keyboard input, hotkeys, and typing speed control. - Supports screen capture, region-based screen
clawhub skill install publishers:matagul:desktop-controlOverall rank
#62
Adoption
57.1K downloads
Trust
Unknown
Freshness
May 31, 2026
Freshness
Last checked May 31, 2026
Best For
Desktop Control is best for general automation workflows where OpenClaw compatibility matters.
Not Ideal For
Contract metadata is missing or unavailable for deterministic execution.
Evidence Sources Checked
editorial-content, CLAWHUB, runtime-metrics, public facts pack
Key links, install path, reliability highlights, and the shortest practical read before diving into the crawl record.
Overview
Advanced desktop automation with mouse, keyboard, and screen control Skill: Desktop Control Owner: matagul Summary: Advanced desktop automation with mouse, keyboard, and screen control Tags: latest:1.0.0 Version history: v1.0.0 | 2026-02-05T01:33:20.863Z | auto Version 1.0.0 - Initial release of the Desktop Control skill for OpenClaw. - Provides advanced automation: mouse movement/clicks, keyboard input, hotkeys, and typing speed control. - Supports screen capture, region-based screen Capability contract not published. No trust telemetry is available yet. 57.1K downloads reported by the source. Last updated 5/31/2026.
Trust score
Unknown
Compatibility
OpenClaw
Freshness
May 31, 2026
Vendor
Clawhub
Artifacts
0
Benchmarks
0
Last release
1.0.0
Install & run
clawhub skill install publishers:matagul:desktop-controlSetup complexity is LOW. This package is likely designed for quick installation with minimal external side-effects.
Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.
Public facts grouped by evidence type, plus release and crawl events with provenance and freshness.
Public facts
Vendor
Clawhub
Protocol compatibility
OpenClaw
Latest release
1.0.0
Adoption signal
57.1K downloads
Handshake status
UNKNOWN
Parameters, dependencies, examples, extracted files, editorial overview, and the complete README when available.
Captured outputs
Extracted files
5
Examples
0
Snippets
0
Languages
Unknown
SKILL.md
---
description: Advanced desktop automation with mouse, keyboard, and screen control
---
# Desktop Control Skill
**The most advanced desktop automation skill for OpenClaw.** Provides pixel-perfect mouse control, lightning-fast keyboard input, screen capture, window management, and clipboard operations.
## ๐ฏ Features
### Mouse Control
- โ
**Absolute positioning** - Move to exact coordinates
- โ
**Relative movement** - Move from current position
- โ
**Smooth movement** - Natural, human-like mouse paths
- โ
**Click types** - Left, right, middle, double, triple clicks
- โ
**Drag & drop** - Drag from point A to point B
- โ
**Scroll** - Vertical and horizontal scrolling
- โ
**Position tracking** - Get current mouse coordinates
### Keyboard Control
- โ
**Text typing** - Fast, accurate text input
- โ
**Hotkeys** - Execute keyboard shortcuts (Ctrl+C, Win+R, etc.)
- โ
**Special keys** - Enter, Tab, Escape, Arrow keys, F-keys
- โ
**Key combinations** - Multi-key press combinations
- โ
**Hold & release** - Manual key state control
- โ
**Typing speed** - Configurable WPM (instant to human-like)
### Screen Operations
- โ
**Screenshot** - Capture entire screen or regions
- โ
**Image recognition** - Find elements on screen (via OpenCV)
- โ
**Color detection** - Get pixel colors at coordinates
- โ
**Multi-monitor** - Support for multiple displays
### Window Management
- โ
**Window list** - Get all open windows
- โ
**Activate window** - Bring window to front
- โ
**Window info** - Get position, size, title
- โ
**Minimize/Maximize** - Control window states
### Safety Features
- โ
**Failsafe** - Move mouse to corner to abort
- โ
**Pause control** - Emergency stop mechanism
- โ
**Approval mode** - Require confirmation for actions
- โ
**Bounds checking** - Prevent out-of-screen operations
- โ
**Logging** - Track all automation actions
---
## ๐ Quick Start
### Installation
First, install required dependencies:
```bash
pip install pyautogui pillow opencv-python pygetwindow
```
### Basic Usage
```python
from skills.desktop_control import DesktopController
# Initialize controller
dc = DesktopController(failsafe=True)
# Mouse operations
dc.move_mouse(500, 300) # Move to coordinates
dc.click() # Left click at current position
dc.click(100, 200, button="right") # Right click at position
# Keyboard operations
dc.type_text("Hello from OpenClaw!")
dc.hotkey("ctrl", "c") # Copy
dc.press("enter")
# Screen operations
screenshot = dc.screenshot()
position = dc.get_mouse_position()
```
---
## ๐ Complete API Reference
### Mouse Functions
#### `move_mouse(x, y, duration=0, smooth=True)`
Move mouse to absolute screen coordinates.
**Parameters:**
- `x` (int): X coordinate (pixels from left)
- `y` (int): Y coordinate (pixels from top)
- `duration` (float): Movement time in seconds (0 = instant, 0.5 = smooth)
- `smooth` (bool): Use bezier curve for natural movement
_meta.json
{
"ownerId": "kn7ag28ra4hhta8bx2k2j1kpv180kqbk",
"slug": "desktop-control",
"version": "1.0.0",
"publishedAt": 1770255200863
}AI_AGENT_GUIDE.md
# AI Desktop Agent - Cognitive Automation Guide
## ๐ค What Is This?
The **AI Desktop Agent** is an intelligent layer on top of the basic desktop control that **understands** what you want and figures out how to do it autonomously.
Unlike basic automation that requires exact instructions, the AI Agent:
- **Understands natural language** ("Draw a cat in Paint")
- **Plans the steps** automatically
- **Executes autonomously**
- **Adapts** based on what it sees
---
## ๐ฏ What Can It Do?
### โ
Autonomous Drawing
```python
from skills.desktop_control.ai_agent import AIDesktopAgent
agent = AIDesktopAgent()
# Just describe what you want!
agent.execute_task("Draw a circle in Paint")
agent.execute_task("Draw a star in MS Paint")
agent.execute_task("Draw a house with a sun")
```
**What it does:**
1. Opens MS Paint
2. Selects pencil tool
3. Figures out how to draw the requested shape
4. Draws it autonomously
5. Takes a screenshot of the result
### โ
Autonomous Text Entry
```python
# It figures out where to type
agent.execute_task("Type 'Hello World' in Notepad")
agent.execute_task("Write an email saying thank you")
```
**What it does:**
1. Opens Notepad (or finds active text editor)
2. Types the text naturally
3. Formats if needed
### โ
Autonomous Application Control
```python
# It knows how to open apps
agent.execute_task("Open Calculator")
agent.execute_task("Launch Microsoft Paint")
agent.execute_task("Open File Explorer")
```
### โ
Autonomous Game Playing (Advanced)
```python
# It will try to play the game!
agent.execute_task("Play Solitaire for me")
agent.execute_task("Play Minesweeper")
```
**What it does:**
1. Analyzes the game screen
2. Detects game state (cards, mines, etc.)
3. Decides best move
4. Executes the move
5. Repeats until win/lose
---
## ๐๏ธ How It Works
### Architecture
```
User Request ("Draw a cat")
โ
Natural Language Understanding
โ
Task Planning (Step-by-step plan)
โ
Step Execution Loop:
- Observe Screen (Computer Vision)
- Decide Action (AI Reasoning)
- Execute Action (Desktop Control)
- Verify Result
โ
Task Complete!
```
### Key Components
1. **Task Planner** - Breaks down high-level tasks into steps
2. **Vision System** - Understands what's on screen (screenshots, OCR, object detection)
3. **Reasoning Engine** - Decides what to do next
4. **Action Executor** - Performsthe actual mouse/keyboard actions
5. **Feedback Loop** - Verifies actions succeeded
---
## ๐ Supported Tasks (Current)
### Tier 1: Fully Automated โ
| Task Pattern | Example | Status |
|-------------|---------|---------|
| Draw shapes in Paint | "Draw a circle" | โ
Working |
| Basic text entry | "Type Hello" | โ
Working |
| Launch applications | "Open Paint" | โ
Working |
### Tier 2: Partially Automated ๐จ
| Task Pattern | Example | Status |
|-------------|---------|---------|
| Form filling | "FQUICK_REFERENCE.md
# Desktop Control - Quick Reference Card
## ๐ Instant Start
```python
from skills.desktop_control import DesktopController
dc = DesktopController()
```
## ๐ฑ๏ธ Mouse Control (Top 10)
```python
# 1. Move mouse
dc.move_mouse(500, 300, duration=0.5)
# 2. Click
dc.click(500, 300) # Left click at position
dc.click() # Click at current position
# 3. Right click
dc.right_click(500, 300)
# 4. Double click
dc.double_click(500, 300)
# 5. Drag & drop
dc.drag(100, 100, 500, 500, duration=1.0)
# 6. Scroll
dc.scroll(-5) # Scroll down 5 clicks
# 7. Get position
x, y = dc.get_mouse_position()
# 8. Move relative
dc.move_relative(100, 50) # Move 100px right, 50px down
# 9. Smooth movement
dc.move_mouse(1000, 500, duration=1.0, smooth=True)
# 10. Middle click
dc.middle_click()
```
## โจ๏ธ Keyboard Control (Top 10)
```python
# 1. Type text (instant)
dc.type_text("Hello World")
# 2. Type text (human-like, 60 WPM)
dc.type_text("Hello World", wpm=60)
# 3. Press key
dc.press('enter')
dc.press('tab')
dc.press('escape')
# 4. Hotkeys (shortcuts)
dc.hotkey('ctrl', 'c') # Copy
dc.hotkey('ctrl', 'v') # Paste
dc.hotkey('ctrl', 's') # Save
dc.hotkey('win', 'r') # Run dialog
dc.hotkey('alt', 'tab') # Switch window
# 5. Hold & release
dc.key_down('shift')
dc.type_text("hello") # Types "HELLO"
dc.key_up('shift')
# 6. Arrow keys
dc.press('up')
dc.press('down')
dc.press('left')
dc.press('right')
# 7. Function keys
dc.press('f5') # Refresh
# 8. Multiple presses
dc.press('backspace', presses=5)
# 9. Special keys
dc.press('home')
dc.press('end')
dc.press('pagedown')
dc.press('delete')
# 10. Fast combo
dc.hotkey('ctrl', 'alt', 'delete')
```
## ๐ธ Screen Operations (Top 5)
```python
# 1. Screenshot (full screen)
img = dc.screenshot()
dc.screenshot(filename="screen.png")
# 2. Screenshot (region)
img = dc.screenshot(region=(100, 100, 800, 600))
# 3. Get pixel color
r, g, b = dc.get_pixel_color(500, 300)
# 4. Find image on screen
location = dc.find_on_screen("button.png")
# 5. Get screen size
width, height = dc.get_screen_size()
```
## ๐ช Window Management (Top 5)
```python
# 1. Get all windows
windows = dc.get_all_windows()
# 2. Activate window
dc.activate_window("Chrome")
# 3. Get active window
active = dc.get_active_window()
# 4. List windows
for title in dc.get_all_windows():
print(title)
# 5. Switch to app
dc.activate_window("Visual Studio Code")
```
## ๐ Clipboard (Top 2)
```python
# 1. Copy to clipboard
dc.copy_to_clipboard("Hello!")
# 2. Get from clipboard
text = dc.get_from_clipboard()
```
## ๐ฅ Real-World Examples
### Example 1: Auto-fill Form
```python
dc.click(300, 200) # Name field
dc.type_text("John Doe", wpm=80)
dc.press('tab')
dc.type_text("[email protected]", wpm=80)
dc.press('tab')
dc.type_text("Password123", wpm=60)
dc.press('enter')
```skill-card.md
## Description: <br> Advanced desktop automation with mouse, keyboard, screen capture, window management, clipboard, and autonomous task execution. <br> This skill is ready for commercial/non-commercial use. <br> ## Publisher: <br> [matagul](https://clawhub.ai/user/matagul) <br> ### License/Terms of Use: <br> ## Use Case: <br> Developers and automation builders use this skill to let an agent operate a live desktop through mouse, keyboard, screen observation, clipboard, window management, and application-launch actions. <br> ### Deployment Geography for Use: <br> Global <br> ## Known Risks and Mitigations: <br> Risk: The skill can control the live desktop, including keyboard, mouse, windows, screenshots, clipboard, and application launching. <br> Mitigation: Install it only for intentional desktop-control use, start in an isolated or test session, and keep sensitive applications and secrets out of view. <br> Risk: Automation actions may run with limited default confirmation. <br> Mitigation: Keep failsafe enabled and prefer require_approval=True for workflows that type, click, use the clipboard, submit information, post publicly, or operate on files. <br> Risk: Autonomous workflows can capture or save screenshots and interact with visible applications. <br> Mitigation: Review each planned action before use in sensitive contexts and avoid autonomous operation when private data is visible. <br> ## Reference(s): <br> - [ClawHub skill page](https://clawhub.ai/matagul/desktop-control) <br> - [SKILL.md](artifact/SKILL.md) <br> - [AI_AGENT_GUIDE.md](artifact/AI_AGENT_GUIDE.md) <br> - [QUICK_REFERENCE.md](artifact/QUICK_REFERENCE.md) <br> ## Skill Output: <br> **Output Type(s):** [Code, Shell commands, Configuration, Guidance, Files] <br> **Output Format:** [Markdown documentation, Python code, shell commands, and runtime desktop actions] <br> **Output Parameters:** [1D] <br> **Other Properties Related to Output:** [May create screenshots or image files when screen capture workflows save output to disk.] <br> ## Skill Version(s): <br> 1.0.0 (source: server release metadata) <br> ## Ethical Considerations: <br> Users should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment. <br>
Editorial read
Docs source
CLAWHUB
Editorial quality
ready
Advanced desktop automation with mouse, keyboard, and screen control Skill: Desktop Control Owner: matagul Summary: Advanced desktop automation with mouse, keyboard, and screen control Tags: latest:1.0.0 Version history: v1.0.0 | 2026-02-05T01:33:20.863Z | auto Version 1.0.0 - Initial release of the Desktop Control skill for OpenClaw. - Provides advanced automation: mouse movement/clicks, keyboard input, hotkeys, and typing speed control. - Supports screen capture, region-based screen
Skill: Desktop Control
Owner: matagul
Summary: Advanced desktop automation with mouse, keyboard, and screen control
Tags: latest:1.0.0
Version history:
v1.0.0 | 2026-02-05T01:33:20.863Z | auto
Version 1.0.0
Archive index:
Archive v1.0.0: 8 files, 26310 bytes
Files: init.py (18805b), AI_AGENT_GUIDE.md (11051b), ai_agent.py (21462b), demo.py (7188b), QUICK_REFERENCE.md (5522b), skill-card.md (2395b), SKILL.md (14085b), _meta.json (134b)
File v1.0.0:SKILL.md
The most advanced desktop automation skill for OpenClaw. Provides pixel-perfect mouse control, lightning-fast keyboard input, screen capture, window management, and clipboard operations.
First, install required dependencies:
pip install pyautogui pillow opencv-python pygetwindow
from skills.desktop_control import DesktopController
# Initialize controller
dc = DesktopController(failsafe=True)
# Mouse operations
dc.move_mouse(500, 300) # Move to coordinates
dc.click() # Left click at current position
dc.click(100, 200, button="right") # Right click at position
# Keyboard operations
dc.type_text("Hello from OpenClaw!")
dc.hotkey("ctrl", "c") # Copy
dc.press("enter")
# Screen operations
screenshot = dc.screenshot()
position = dc.get_mouse_position()
move_mouse(x, y, duration=0, smooth=True)Move mouse to absolute screen coordinates.
Parameters:
x (int): X coordinate (pixels from left)y (int): Y coordinate (pixels from top)duration (float): Movement time in seconds (0 = instant, 0.5 = smooth)smooth (bool): Use bezier curve for natural movementExample:
# Instant movement
dc.move_mouse(1000, 500)
# Smooth 1-second movement
dc.move_mouse(1000, 500, duration=1.0)
move_relative(x_offset, y_offset, duration=0)Move mouse relative to current position.
Parameters:
x_offset (int): Pixels to move horizontally (positive = right)y_offset (int): Pixels to move vertically (positive = down)duration (float): Movement time in secondsExample:
# Move 100px right, 50px down
dc.move_relative(100, 50, duration=0.3)
click(x=None, y=None, button='left', clicks=1, interval=0.1)Perform mouse click.
Parameters:
x, y (int, optional): Coordinates to click (None = current position)button (str): 'left', 'right', 'middle'clicks (int): Number of clicks (1 = single, 2 = double)interval (float): Delay between multiple clicksExample:
# Simple left click
dc.click()
# Double-click at specific position
dc.click(500, 300, clicks=2)
# Right-click
dc.click(button='right')
drag(start_x, start_y, end_x, end_y, duration=0.5, button='left')Drag and drop operation.
Parameters:
start_x, start_y (int): Starting coordinatesend_x, end_y (int): Ending coordinatesduration (float): Drag durationbutton (str): Mouse button to useExample:
# Drag file from desktop to folder
dc.drag(100, 100, 500, 500, duration=1.0)
scroll(clicks, direction='vertical', x=None, y=None)Scroll mouse wheel.
Parameters:
clicks (int): Scroll amount (positive = up/left, negative = down/right)direction (str): 'vertical' or 'horizontal'x, y (int, optional): Position to scroll atExample:
# Scroll down 5 clicks
dc.scroll(-5)
# Scroll up 10 clicks
dc.scroll(10)
# Horizontal scroll
dc.scroll(5, direction='horizontal')
get_mouse_position()Get current mouse coordinates.
Returns: (x, y) tuple
Example:
x, y = dc.get_mouse_position()
print(f"Mouse is at: {x}, {y}")
type_text(text, interval=0, wpm=None)Type text with configurable speed.
Parameters:
text (str): Text to typeinterval (float): Delay between keystrokes (0 = instant)wpm (int, optional): Words per minute (overrides interval)Example:
# Instant typing
dc.type_text("Hello World")
# Human-like typing at 60 WPM
dc.type_text("Hello World", wpm=60)
# Slow typing with 0.1s between keys
dc.type_text("Hello World", interval=0.1)
press(key, presses=1, interval=0.1)Press and release a key.
Parameters:
key (str): Key name (see Key Names section)presses (int): Number of times to pressinterval (float): Delay between pressesExample:
# Press Enter
dc.press('enter')
# Press Space 3 times
dc.press('space', presses=3)
# Press Down arrow
dc.press('down')
hotkey(*keys, interval=0.05)Execute keyboard shortcut.
Parameters:
*keys (str): Keys to press togetherinterval (float): Delay between key pressesExample:
# Copy (Ctrl+C)
dc.hotkey('ctrl', 'c')
# Paste (Ctrl+V)
dc.hotkey('ctrl', 'v')
# Open Run dialog (Win+R)
dc.hotkey('win', 'r')
# Save (Ctrl+S)
dc.hotkey('ctrl', 's')
# Select All (Ctrl+A)
dc.hotkey('ctrl', 'a')
key_down(key) / key_up(key)Manually control key state.
Example:
# Hold Shift
dc.key_down('shift')
dc.type_text("hello") # Types "HELLO"
dc.key_up('shift')
# Hold Ctrl and click (for multi-select)
dc.key_down('ctrl')
dc.click(100, 100)
dc.click(200, 100)
dc.key_up('ctrl')
screenshot(region=None, filename=None)Capture screen or region.
Parameters:
region (tuple, optional): (left, top, width, height) for partial capturefilename (str, optional): Path to save imageReturns: PIL Image object
Example:
# Full screen
img = dc.screenshot()
# Save to file
dc.screenshot(filename="screenshot.png")
# Capture specific region
img = dc.screenshot(region=(100, 100, 500, 300))
get_pixel_color(x, y)Get color of pixel at coordinates.
Returns: RGB tuple (r, g, b)
Example:
r, g, b = dc.get_pixel_color(500, 300)
print(f"Color at (500, 300): RGB({r}, {g}, {b})")
find_on_screen(image_path, confidence=0.8)Find image on screen (requires OpenCV).
Parameters:
image_path (str): Path to template imageconfidence (float): Match threshold (0-1)Returns: (x, y, width, height) or None
Example:
# Find button on screen
location = dc.find_on_screen("button.png")
if location:
x, y, w, h = location
# Click center of found image
dc.click(x + w//2, y + h//2)
get_screen_size()Get screen resolution.
Returns: (width, height) tuple
Example:
width, height = dc.get_screen_size()
print(f"Screen: {width}x{height}")
get_all_windows()List all open windows.
Returns: List of window titles
Example:
windows = dc.get_all_windows()
for title in windows:
print(f"Window: {title}")
activate_window(title_substring)Bring window to front by title.
Parameters:
title_substring (str): Part of window title to matchExample:
# Activate Chrome
dc.activate_window("Chrome")
# Activate VS Code
dc.activate_window("Visual Studio Code")
get_active_window()Get currently focused window.
Returns: Window title (str)
Example:
active = dc.get_active_window()
print(f"Active window: {active}")
copy_to_clipboard(text)Copy text to clipboard.
Example:
dc.copy_to_clipboard("Hello from OpenClaw!")
get_from_clipboard()Get text from clipboard.
Returns: str
Example:
text = dc.get_from_clipboard()
print(f"Clipboard: {text}")
'a' through 'z'
'0' through '9'
'f1' through 'f24'
'enter' / 'return''esc' / 'escape''space' / 'spacebar''tab''backspace''delete' / 'del''insert''home''end''pageup' / 'pgup''pagedown' / 'pgdn''up' / 'down' / 'left' / 'right''ctrl' / 'control''shift''alt''win' / 'winleft' / 'winright''cmd' / 'command' (Mac)'capslock''numlock''scrolllock''.' / ',' / '?' / '!' / ';' / ':''[' / ']' / '{' / '}''(' / ')''+' / '-' / '*' / '/' / '='Move mouse to any corner of the screen to abort all automation.
# Enable failsafe (enabled by default)
dc = DesktopController(failsafe=True)
# Pause all automation for 2 seconds
dc.pause(2.0)
# Check if automation is safe to proceed
if dc.is_safe():
dc.click(500, 500)
Require user confirmation before actions:
dc = DesktopController(require_approval=True)
# This will ask for confirmation
dc.click(500, 500) # Prompt: "Allow click at (500, 500)? [y/n]"
dc = DesktopController()
# Click name field
dc.click(300, 200)
dc.type_text("John Doe", wpm=80)
# Tab to next field
dc.press('tab')
dc.type_text("[email protected]", wpm=80)
# Tab to password
dc.press('tab')
dc.type_text("SecurePassword123", wpm=60)
# Submit form
dc.press('enter')
# Capture specific area
region = (100, 100, 800, 600) # left, top, width, height
img = dc.screenshot(region=region)
# Save with timestamp
import datetime
timestamp = datetime.datetime.now().strftime("%Y%m%d_%H%M%S")
img.save(f"capture_{timestamp}.png")
# Hold Ctrl and click multiple files
dc.key_down('ctrl')
dc.click(100, 200) # First file
dc.click(100, 250) # Second file
dc.click(100, 300) # Third file
dc.key_up('ctrl')
# Copy selected files
dc.hotkey('ctrl', 'c')
# Activate Calculator
dc.activate_window("Calculator")
time.sleep(0.5)
# Type calculation
dc.type_text("5+3=", interval=0.2)
time.sleep(0.5)
# Take screenshot of result
dc.screenshot(filename="calculation_result.png")
# Drag file from source to destination
dc.drag(
start_x=200, start_y=300, # File location
end_x=800, end_y=500, # Folder location
duration=1.0 # Smooth 1-second drag
)
duration=0get_screen_size() to confirm dimensionsinterval for reliabilityDesktopController(failsafe=False)Install all:
pip install pyautogui pillow opencv-python pygetwindow
Built for OpenClaw - The ultimate desktop automation companion ๐ฆ
File v1.0.0:_meta.json
{ "ownerId": "kn7ag28ra4hhta8bx2k2j1kpv180kqbk", "slug": "desktop-control", "version": "1.0.0", "publishedAt": 1770255200863 }
File v1.0.0:AI_AGENT_GUIDE.md
The AI Desktop Agent is an intelligent layer on top of the basic desktop control that understands what you want and figures out how to do it autonomously.
Unlike basic automation that requires exact instructions, the AI Agent:
from skills.desktop_control.ai_agent import AIDesktopAgent
agent = AIDesktopAgent()
# Just describe what you want!
agent.execute_task("Draw a circle in Paint")
agent.execute_task("Draw a star in MS Paint")
agent.execute_task("Draw a house with a sun")
What it does:
# It figures out where to type
agent.execute_task("Type 'Hello World' in Notepad")
agent.execute_task("Write an email saying thank you")
What it does:
# It knows how to open apps
agent.execute_task("Open Calculator")
agent.execute_task("Launch Microsoft Paint")
agent.execute_task("Open File Explorer")
# It will try to play the game!
agent.execute_task("Play Solitaire for me")
agent.execute_task("Play Minesweeper")
What it does:
User Request ("Draw a cat")
โ
Natural Language Understanding
โ
Task Planning (Step-by-step plan)
โ
Step Execution Loop:
- Observe Screen (Computer Vision)
- Decide Action (AI Reasoning)
- Execute Action (Desktop Control)
- Verify Result
โ
Task Complete!
| Task Pattern | Example | Status | |-------------|---------|---------| | Draw shapes in Paint | "Draw a circle" | โ Working | | Basic text entry | "Type Hello" | โ Working | | Launch applications | "Open Paint" | โ Working |
| Task Pattern | Example | Status | |-------------|---------|---------| | Form filling | "Fill out this form" | ๐จ In Progress | | File operations | "Copy these files" | ๐จ In Progress | | Web navigation | "Find on Google" | ๐จ Planned |
| Task Pattern | Example | Status | |-------------|---------|---------| | Game playing | "Play Solitaire" | ๐งช Experimental | | Image editing | "Resize this photo" | ๐งช Planned | | Code editing | "Fix this bug" | ๐งช Research |
agent = AIDesktopAgent()
result = agent.execute_task("Draw a circle in Paint")
# Check result
print(f"Status: {result['status']}")
print(f"Steps taken: {len(result['steps'])}")
1. Planning Phase:
Plan generated:
Step 1: Launch MS Paint
Step 2: Wait 2s for Paint to load
Step 3: Activate Paint window
Step 4: Select pencil tool (press 'P')
Step 5: Draw circle at canvas center
Step 6: Screenshot the result
2. Execution Phase:
[โ] Launched Paint via Win+R โ mspaint
[โ] Waited 2.0s
[โ] Activated window "Paint"
[โ] Pressed 'P' to select pencil
[โ] Drew circle with 72 points
[โ] Screenshot saved: drawing_result.png
3. Result:
{
"task": "Draw a circle in Paint",
"status": "completed",
"success": True,
"steps": [... 6 steps ...],
"screenshots": [... 6 screenshots ...],
}
agent = AIDesktopAgent()
# Play a simple game
result = agent.execute_task("Play Solitaire for me")
1. Analyze screen โ Detect cards, positions
2. Identify valid moves โ Find legal plays
3. Evaluate moves โ Which is best?
4. Execute move โ Click and drag card
5. Repeat until game ends
The agent can learn patterns for:
# In ai_agent.py, add to app_knowledge:
self.app_knowledge = {
"photoshop": {
"name": "Adobe Photoshop",
"launch_command": "photoshop",
"common_actions": {
"new_layer": {"hotkey": ["ctrl", "shift", "n"]},
"brush_tool": {"hotkey": ["b"]},
"eraser": {"hotkey": ["e"]},
}
}
}
# Add a custom planning method
def _plan_photo_edit(self, task: str) -> List[Dict]:
"""Plan for photo editing tasks."""
return [
{"type": "launch_app", "app": "photoshop"},
{"type": "wait", "duration": 3.0},
{"type": "open_file", "path": extracted_path},
{"type": "apply_filter", "filter": extracted_filter},
{"type": "save_file"},
]
The agent can analyze screenshots to:
# Analyze what's on screen
analysis = agent._analyze_screen()
print(analysis)
# Output:
# {
# "active_window": "Untitled - Paint",
# "mouse_position": (640, 480),
# "detected_elements": [...],
# "text_found": [...],
# }
# Future: Use OpenClaw's LLM for reasoning
agent = AIDesktopAgent(llm_client=openclaw_llm)
# The agent can now:
# - Reason about complex tasks
# - Understand context better
# - Plan more sophisticated workflows
# - Learn from feedback
Example: Adding Excel support
# Step 1: Add to app_knowledge
"excel": {
"name": "Microsoft Excel",
"launch_command": "excel",
"common_actions": {
"new_sheet": {"hotkey": ["shift", "f11"]},
"sum_formula": {"action": "type", "text": "=SUM()"},
}
}
# Step 2: Create planner
def _plan_excel_task(self, task: str) -> List[Dict]:
return [
{"type": "launch_app", "app": "excel"},
{"type": "wait", "duration": 2.0},
# ... specific Excel steps
]
# Step 3: Hook into main planner
if "excel" in task_lower or "spreadsheet" in task_lower:
return self._plan_excel_task(task)
agent.execute_task("Fill out the job application with my resume data")
agent.execute_task("Resize all images in this folder to 800x600")
agent.execute_task("Post this image to Instagram with caption 'Beautiful sunset'")
agent.execute_task("Copy data from this PDF to Excel spreadsheet")
agent.execute_task("Test the login form with invalid credentials")
# Safe mode (default)
agent = AIDesktopAgent(failsafe=True)
# Fast mode (no failsafe)
agent = AIDesktopAgent(failsafe=False)
# Prevent infinite loops
result = agent.execute_task("Play game", max_steps=100)
# Review what the agent did
print(agent.action_history)
result = agent.execute_task("Draw a star in Paint")
for i, step in enumerate(result['steps'], 1):
print(f"Step {i}: {step['step']['description']}")
print(f" Success: {step['success']}")
if 'error' in step:
print(f" Error: {step['error']}")
# Each step captures before/after screenshots
for screenshot_pair in result['screenshots']:
before = screenshot_pair['before']
after = screenshot_pair['after']
# Display or save for analysis
before.save(f"step_{screenshot_pair['step']}_before.png")
after.save(f"step_{screenshot_pair['step']}_after.png")
Planned features:
# Execute a task
result = agent.execute_task(task: str, max_steps: int = 50)
# Analyze screen
analysis = agent._analyze_screen()
# Manual mode: Execute individual steps
step = {"type": "launch_app", "app": "paint"}
result = agent._execute_step(step)
{
"task": str, # Original task
"status": str, # "completed", "failed", "error"
"success": bool, # Overall success
"steps": List[Dict], # All steps executed
"screenshots": List[Dict], # Before/after screenshots
"failed_at_step": int, # If failed, which step
"error": str, # Error message if failed
}
๐ฆ Built for OpenClaw - The future of desktop automation!
File v1.0.0:QUICK_REFERENCE.md
from skills.desktop_control import DesktopController
dc = DesktopController()
# 1. Move mouse
dc.move_mouse(500, 300, duration=0.5)
# 2. Click
dc.click(500, 300) # Left click at position
dc.click() # Click at current position
# 3. Right click
dc.right_click(500, 300)
# 4. Double click
dc.double_click(500, 300)
# 5. Drag & drop
dc.drag(100, 100, 500, 500, duration=1.0)
# 6. Scroll
dc.scroll(-5) # Scroll down 5 clicks
# 7. Get position
x, y = dc.get_mouse_position()
# 8. Move relative
dc.move_relative(100, 50) # Move 100px right, 50px down
# 9. Smooth movement
dc.move_mouse(1000, 500, duration=1.0, smooth=True)
# 10. Middle click
dc.middle_click()
# 1. Type text (instant)
dc.type_text("Hello World")
# 2. Type text (human-like, 60 WPM)
dc.type_text("Hello World", wpm=60)
# 3. Press key
dc.press('enter')
dc.press('tab')
dc.press('escape')
# 4. Hotkeys (shortcuts)
dc.hotkey('ctrl', 'c') # Copy
dc.hotkey('ctrl', 'v') # Paste
dc.hotkey('ctrl', 's') # Save
dc.hotkey('win', 'r') # Run dialog
dc.hotkey('alt', 'tab') # Switch window
# 5. Hold & release
dc.key_down('shift')
dc.type_text("hello") # Types "HELLO"
dc.key_up('shift')
# 6. Arrow keys
dc.press('up')
dc.press('down')
dc.press('left')
dc.press('right')
# 7. Function keys
dc.press('f5') # Refresh
# 8. Multiple presses
dc.press('backspace', presses=5)
# 9. Special keys
dc.press('home')
dc.press('end')
dc.press('pagedown')
dc.press('delete')
# 10. Fast combo
dc.hotkey('ctrl', 'alt', 'delete')
# 1. Screenshot (full screen)
img = dc.screenshot()
dc.screenshot(filename="screen.png")
# 2. Screenshot (region)
img = dc.screenshot(region=(100, 100, 800, 600))
# 3. Get pixel color
r, g, b = dc.get_pixel_color(500, 300)
# 4. Find image on screen
location = dc.find_on_screen("button.png")
# 5. Get screen size
width, height = dc.get_screen_size()
# 1. Get all windows
windows = dc.get_all_windows()
# 2. Activate window
dc.activate_window("Chrome")
# 3. Get active window
active = dc.get_active_window()
# 4. List windows
for title in dc.get_all_windows():
print(title)
# 5. Switch to app
dc.activate_window("Visual Studio Code")
# 1. Copy to clipboard
dc.copy_to_clipboard("Hello!")
# 2. Get from clipboard
text = dc.get_from_clipboard()
dc.click(300, 200) # Name field
dc.type_text("John Doe", wpm=80)
dc.press('tab')
dc.type_text("[email protected]", wpm=80)
dc.press('tab')
dc.type_text("Password123", wpm=60)
dc.press('enter')
# Select all
dc.hotkey('ctrl', 'a')
# Copy
dc.hotkey('ctrl', 'c')
# Wait
dc.pause(0.5)
# Switch window
dc.hotkey('alt', 'tab')
# Paste
dc.hotkey('ctrl', 'v')
# Select multiple files
dc.key_down('ctrl')
dc.click(100, 200)
dc.click(100, 250)
dc.click(100, 300)
dc.key_up('ctrl')
# Copy
dc.hotkey('ctrl', 'c')
# Take screenshot
dc.screenshot(filename=f"capture_{time.time()}.png")
# Open in Paint
dc.hotkey('win', 'r')
dc.pause(0.5)
dc.type_text('mspaint')
dc.press('enter')
# Open Find & Replace
dc.hotkey('ctrl', 'h')
dc.pause(0.3)
# Type find text
dc.type_text("old_text")
dc.press('tab')
# Type replace text
dc.type_text("new_text")
# Replace all
dc.hotkey('alt', 'a')
# With failsafe (move to corner to abort)
dc = DesktopController(failsafe=True)
# With approval mode (ask before each action)
dc = DesktopController(require_approval=True)
# Maximum speed (no safety checks)
dc = DesktopController(failsafe=False)
# Check if safe to continue
if dc.is_safe():
dc.click(500, 500)
# Pause execution
dc.pause(2.0) # Wait 2 seconds
# Emergency abort: Move mouse to any screen corner
interval=0 or wpm=Nonewpm=60 (60 words/min)duration=0.5, smooth=Trueduration=0dc.pause(0.5) between actionsdemo.py to test featuresget_mouse_position() to find thempip install pyautogui pillow opencv-python pygetwindow pyperclip
Mouse not moving correctly?
get_mouse_position()Keyboard not working?
Failsafe triggering?
failsafe=FalseBuilt for OpenClaw ๐ฆ - Desktop automation made easy!
File v1.0.0:skill-card.md
Advanced desktop automation with mouse, keyboard, screen capture, window management, clipboard, and autonomous task execution. <br>
This skill is ready for commercial/non-commercial use. <br>
matagul <br>
Developers and automation builders use this skill to let an agent operate a live desktop through mouse, keyboard, screen observation, clipboard, window management, and application-launch actions. <br>
Global <br>
Risk: The skill can control the live desktop, including keyboard, mouse, windows, screenshots, clipboard, and application launching. <br> Mitigation: Install it only for intentional desktop-control use, start in an isolated or test session, and keep sensitive applications and secrets out of view. <br> Risk: Automation actions may run with limited default confirmation. <br> Mitigation: Keep failsafe enabled and prefer require_approval=True for workflows that type, click, use the clipboard, submit information, post publicly, or operate on files. <br> Risk: Autonomous workflows can capture or save screenshots and interact with visible applications. <br> Mitigation: Review each planned action before use in sensitive contexts and avoid autonomous operation when private data is visible. <br>
Output Type(s): [Code, Shell commands, Configuration, Guidance, Files] <br> Output Format: [Markdown documentation, Python code, shell commands, and runtime desktop actions] <br> Output Parameters: [1D] <br> Other Properties Related to Output: [May create screenshots or image files when screen capture workflows save output to disk.] <br>
1.0.0 (source: server release metadata) <br>
Users should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment. <br>
Machine endpoints, contract coverage, trust signals, runtime metrics, benchmarks, and guardrails for agent-to-agent use.
Machine interfaces
Contract coverage
Status
missing
Auth
None
Streaming
No
Data region
Unspecified
Protocol support
Requires: none
Forbidden: none
Guardrails
Operational confidence: low
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-matagul-desktop-control/snapshot"
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-matagul-desktop-control/contract"
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-matagul-desktop-control/trust"
Operational fit
Trust signals
Handshake
UNKNOWN
Confidence
unknown
Attempts 30d
unknown
Fallback rate
unknown
Runtime metrics
Observed P50
unknown
Observed P95
unknown
Rate limit
unknown
Estimated cost
unknown
Do not use if
Raw contract, invocation, trust, capability, facts, and change-event payloads for machine-side inspection.
Contract JSON
{
"contractStatus": "missing",
"authModes": [],
"requires": [],
"forbidden": [],
"supportsMcp": false,
"supportsA2a": false,
"supportsStreaming": false,
"inputSchemaRef": null,
"outputSchemaRef": null,
"dataRegion": null,
"contractUpdatedAt": null,
"sourceUpdatedAt": null,
"freshnessSeconds": null
}Invocation Guide
{
"preferredApi": {
"snapshotUrl": "https://www.xpersona.co/api/v1/agents/clawhub-matagul-desktop-control/snapshot",
"contractUrl": "https://www.xpersona.co/api/v1/agents/clawhub-matagul-desktop-control/contract",
"trustUrl": "https://www.xpersona.co/api/v1/agents/clawhub-matagul-desktop-control/trust"
},
"curlExamples": [
"curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-matagul-desktop-control/snapshot\"",
"curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-matagul-desktop-control/contract\"",
"curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-matagul-desktop-control/trust\""
],
"jsonRequestTemplate": {
"query": "summarize this repo",
"constraints": {
"maxLatencyMs": 2000,
"protocolPreference": [
"OPENCLEW"
]
}
},
"jsonResponseTemplate": {
"ok": true,
"result": {
"summary": "...",
"confidence": 0.9
},
"meta": {
"source": "CLAWHUB",
"generatedAt": "2026-10-08T22:19:26.818Z"
}
},
"retryPolicy": {
"maxAttempts": 3,
"backoffMs": [
500,
1500,
3500
],
"retryableConditions": [
"HTTP_429",
"HTTP_503",
"NETWORK_TIMEOUT"
]
}
}Trust JSON
{
"status": "unavailable",
"handshakeStatus": "UNKNOWN",
"verificationFreshnessHours": null,
"reputationScore": null,
"p95LatencyMs": null,
"successRate30d": null,
"fallbackRate": null,
"attempts30d": null,
"trustUpdatedAt": null,
"trustConfidence": "unknown",
"sourceUpdatedAt": null,
"freshnessSeconds": null
}Capability Matrix
{
"rows": [
{
"key": "OPENCLEW",
"type": "protocol",
"support": "unknown",
"confidenceSource": "profile",
"notes": "Listed on profile"
}
],
"flattenedTokens": "protocol:OPENCLEW|unknown|profile"
}Facts JSON
[
{
"factKey": "vendor",
"label": "Vendor",
"value": "Clawhub",
"category": "vendor",
"href": "https://clawhub.ai/matagul/desktop-control",
"sourceUrl": "https://clawhub.ai/matagul/desktop-control",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-05-31T06:21:26.813Z",
"isPublic": true,
"metadata": {}
},
{
"factKey": "protocols",
"label": "Protocol compatibility",
"value": "OpenClaw",
"category": "compatibility",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-matagul-desktop-control/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-matagul-desktop-control/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-05-31T06:21:26.813Z",
"isPublic": true,
"metadata": {}
},
{
"factKey": "traction",
"label": "Adoption signal",
"value": "57.1K downloads",
"category": "adoption",
"href": "https://clawhub.ai/matagul/desktop-control",
"sourceUrl": "https://clawhub.ai/matagul/desktop-control",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-05-31T06:21:26.813Z",
"isPublic": true,
"metadata": {}
},
{
"factKey": "latest_release",
"label": "Latest release",
"value": "1.0.0",
"category": "release",
"href": "https://clawhub.ai/matagul/desktop-control",
"sourceUrl": "https://clawhub.ai/matagul/desktop-control",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-02-05T01:33:20.863Z",
"isPublic": true,
"metadata": {}
},
{
"factKey": "handshake_status",
"label": "Handshake status",
"value": "UNKNOWN",
"category": "security",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-matagul-desktop-control/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-matagul-desktop-control/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true,
"metadata": {}
}
]Change Events JSON
[
{
"eventType": "release",
"title": "Release 1.0.0",
"description": "Version 1.0.0 - Initial release of the Desktop Control skill for OpenClaw. - Provides advanced automation: mouse movement/clicks, keyboard input, hotkeys, and typing speed control. - Supports screen capture, region-based screenshots, image/template matching, and pixel color detection. - Includes window management (list, activate, move, resize, minimize/maximize). - Safety features: failsafe abort, logging, approval mode, bounds checks, and emergency pause. - Detailed documentation with examples and complete API reference.",
"href": "https://clawhub.ai/matagul/desktop-control",
"sourceUrl": "https://clawhub.ai/matagul/desktop-control",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-02-05T01:33:20.863Z",
"isPublic": true,
"metadata": {}
}
]Sponsored
Ads related to Desktop Control and adjacent AI workflows.