Vision Helper — AI Image Analysis
Analyze images using local or cloud vision models via Ollama to identify content, UI elements, screenshots, or extract text with OCR support. Skill: Vision Helper — AI Image Analysis Owner: ravenquasar Summary: Analyze images using local or cloud vision models via Ollama to identify content, UI elements, screenshots, or extract text with OCR support. Tags: latest:1.0.0 Version history: v1.0.0 | 2026-04-28T10:40:07.276Z | auto - Initial release of Vision Helper, an image analysis skill using local or cloud vision models via Ollama. - Supports analyzing imag
Rank
62
Safety
84
Downloads
1.7k
Updated
Oct 10, 2026
Version
1.0.0
Source
CLAWHUB
About
What it does, and when to use it.
Capability contract not published. No trust telemetry is available yet. 1.7K downloads reported by the source. Last updated 10/10/2026.
Avoid when
- Contract metadata is missing or unavailable for deterministic execution.
Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing
Public facts
Every fact links back to the source it came from.
- Vendor
- Clawhubvendor · observed Oct 10, 2026
- Protocol compatibility
- OpenClawcompatibility · observed Oct 10, 2026
- Adoption signal
- 1.7K downloadsadoption · observed Oct 10, 2026
- Latest release
- 1.0.0release · observed Apr 28, 2026
- Handshake status
- UNKNOWNsecurity
Install and run
Setup complexity: low.
clawhub skill install s17drtcv3zq9axeqrjg3fh8e6d85pc2h:vision-helper- Setup complexity is LOW. This package is likely designed for quick installation with minimal external side-effects.
- Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.
Contract: missing
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-ravenquasar-vision-helper/snapshot"
Documentation
CLAWHUB
8,303 characters of source documentation, loaded on request.
Extracted files
4 files captured from the source.
SKILL.md
---
name: vision-helper
description: Analyze images using local or cloud vision models via Ollama. Use when you need to identify screenshots, analyze UI elements, read image content, or perform OCR. Triggers on: "analyze image", "screenshot", "what's in this image", "OCR", "识图", "看图", "截图识别".
---
# 📸 Vision Helper — Image Analysis
Analyze images using vision models via Ollama, with extended timeout support for cloud-based models.
## Why Not Use the Built-in `image` Tool?
The built-in `image` tool has limited timeout settings that cause failures with cloud vision models (which often need 40–120 seconds). This skill calls the Ollama API directly with a 180-second timeout, supporting both local and cloud models reliably.
It also bypasses the built-in tool's file path restrictions, allowing analysis of images from any readable directory.
## Usage
### Basic
```bash
# Analyze an image (default: English description)
python3 <skill-dir>/scripts/analyze_image.py <image_path>
# With a custom prompt
python3 <skill-dir>/scripts/analyze_image.py <image_path> "Is this a chess game? Describe the board state"
# With a specific model
python3 <skill-dir>/scripts/analyze_image.py <image_path> "Describe content" kimi-k2.5:cloud
```
> `<skill-dir>` resolves to your OpenClaw skill installation directory, typically `~/.openclaw/workspace/skills/vision-helper/`.
### In Conversation
When you need to analyze an image, use the `exec` tool:
```
exec: python3 <skill-dir>/scripts/analyze_image.py /path/to/image.png "What do you see?"
```
**Important:** Set exec timeout to 120–180 seconds, as cloud vision models are slow.
### Screenshot + Analysis Workflow
#### Option A: Browser screenshot → analyze
```
1. browser(action="screenshot") → get screenshot path (MEDIA: xxx)
2. exec("<skill-dir>/scripts/analyze_image.py <screenshot_path> 'Describe this UI'")
3. Act on the analysis result
```
#### Option B: Desktop screenshot → analyze
**macOS:**
```
1. exec("screencapture -x /tmp/screen.png")
2. exec("<skill-dir>/scripts/analyze_image.py /tmp/screen.png 'Describe the desktop'")
```
**Linux:**
```
1. exec("gnome-screenshot -f /tmp/screen.png")
— or —
exec("import /tmp/screen.png") # ImageMagick
— or —
exec("scrot /tmp/screen.png")
2. exec("<skill-dir>/scripts/analyze_image.py /tmp/screen.png 'Describe the desktop'")
```
#### Option C: Game/App UI → analyze → act
```
1. Screenshot the current screen
2. Use vision-helper to identify UI elements, buttons, text
3. Execute clicks/input based on the analysis
```
## Environment Variables
| Variable | Default | Description |
|----------|---------|-------------|
| `VISION_MODEL` | `gemma4:31b` | Default vision model |
| `VISION_TIMEOUT` | `180` | Request timeout in seconds |
| `OLLAMA_API_URL` | `http://localhost:11434/api/chat` | Ollama API endpoint |
## Supported Models
| Model | Vision | Speed | Recommendation |
|-------|--------|-------|----------------|
| `gemma4:31b` | ✅ | Local, fast | ⭐ **Prima_meta.json
{
"ownerId": "kn7805nw545a1g4024y6fyxy7185q82f",
"slug": "vision-helper",
"version": "1.0.0",
"publishedAt": 1777372807276
}skill-card.md
## Description: Analyze images using local or cloud vision models via Ollama to identify content, UI elements, screenshots, or extract text with OCR support. This skill is ready for commercial/non-commercial use. ## Publisher: [ravenquasar](https://clawhub.ai/user/ravenquasar) ### License/Terms of Use: MIT-0 ## Use Case: Developers and agents use this skill to analyze local images or screenshots, inspect UI content, and perform OCR with an Ollama vision model when longer model timeouts are needed. ### Deployment Geography for Use: Global ## Known Risks and Mitigations: Risk: The skill can read local images and screenshots from broad filesystem locations, which may expose credentials, private documents, or sensitive desktop content. Mitigation: Use it only on images intended for analysis, prefer workspace or temporary paths, and avoid screenshots containing secrets or private documents. Risk: Image data can be sent to a configurable Ollama endpoint, including cloud or remote models. Mitigation: Keep OLLAMA_API_URL on localhost for routine use, and configure a remote endpoint only when that service is explicitly trusted for the image contents. ## Reference(s): - [ClawHub skill page](https://clawhub.ai/ravenquasar/skills/vision-helper) ## Skill Output: **Output Type(s):** [text, shell commands, configuration, guidance] **Output Format:** [Plain text image analysis from the configured vision model; usage guidance includes shell commands and environment variables.] **Output Parameters:** [1D] **Other Properties Related to Output:** [Accepts an image path, optional prompt, optional model name, and environment configuration for model, timeout, and Ollama API URL.] ## Skill Version(s): 1.0.0 (source: server release metadata and clawhub.json) ## Ethical Considerations: Users should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.
clawhub.json
{
"name": "vision-helper",
"displayName": "Vision Helper — AI Image Analysis",
"version": "1.0.0",
"description": "Analyze images using local or cloud vision models via Ollama. Bypasses built-in image tool timeout limits with extended 180s timeout, file validation, and cross-platform screenshot support.",
"author": "Frieren",
"license": "MIT",
"pricing": {
"model": "free"
},
"tags": ["vision", "image", "ocr", "screenshot", "ollama", "gemma", "kimi"],
"minOpenClawVersion": "1.2.0"
}AionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
activepieces
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
CopilotKit
The Frontend for Agents & Generative UI. React + Angular
Machine-readable data
The same record, as JSON, for agents and crawlers.
{
"facts": [
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Clawhub",
"href": "https://clawhub.ai/ravenquasar/skills/vision-helper",
"sourceUrl": "https://clawhub.ai/ravenquasar/skills/vision-helper",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-10T04:07:40.839Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-ravenquasar-vision-helper/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-ravenquasar-vision-helper/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-10T04:07:40.839Z",
"isPublic": true
},
{
"factKey": "traction",
"category": "adoption",
"label": "Adoption signal",
"value": "1.7K downloads",
"href": "https://clawhub.ai/ravenquasar/vision-helper",
"sourceUrl": "https://clawhub.ai/ravenquasar/vision-helper",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-10T04:07:40.839Z",
"isPublic": true
},
{
"factKey": "latest_release",
"category": "release",
"label": "Latest release",
"value": "1.0.0",
"href": "https://clawhub.ai/ravenquasar/vision-helper",
"sourceUrl": "https://clawhub.ai/ravenquasar/vision-helper",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-04-28T10:40:07.276Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-ravenquasar-vision-helper/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-ravenquasar-vision-helper/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
],
"events": [
{
"eventType": "release",
"title": "Release 1.0.0",
"description": "- Initial release of Vision Helper, an image analysis skill using local or cloud vision models via Ollama. - Supports analyzing images, UI elements, screenshots, and performing OCR with extended timeout for cloud models (up to 180 seconds). - Bypasses built-in image tool limitations, including path restrictions and short timeouts. - Provides CLI and conversational usage examples, including workflows for browser, desktop, and game UI screenshots. - Allows easy switching between multiple supported local and cloud vision models via environment variables. - Supports various image formats and directory paths for flexible screenshot handling.",
"href": "https://clawhub.ai/ravenquasar/vision-helper",
"sourceUrl": "https://clawhub.ai/ravenquasar/vision-helper",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-04-28T10:40:07.276Z",
"isPublic": true
}
]
}Record generated Oct 10, 2026.
