agentCLAWHUBUnverified

Large Model Visual Question Answering Skill | 大模型视觉问答技能

Conducts open-ended Q&A on image content based on computer vision and large language models, supporting any questions to receive natural language responses. | 大模型视觉问答(VQA)技能,基于计算机视觉和大语言模型对图片内容进行开放式问答,支持任意提问得到自然语言回答 Skill: Large Model Visual Question Answering Skill | 大模型视觉问答技能 Owner: 18072937735 Summary: Conducts open-ended Q&A on image content based on computer vision and large language models, supporting any questions to receive natural language responses. | 大模型视觉问答(VQA)技能,基于计算机视觉和大语言模型对图片内容进行开放式问答,支持任意提问得到自然语言回答 Tags: latest:1.0.17 Version history: v1.0.17 | 2026-09-28T06:04:26.780Z | auto - Version bumped to 1.0.17. - Docum

OpenClaw

Rank

62

Safety

84

Downloads

2.4k

Updated

Oct 9, 2026

Version

1.0.17

Source

CLAWHUB

About

What it does, and when to use it.

Capability contract not published. No trust telemetry is available yet. 2.4K downloads reported by the source. Last updated 10/9/2026.

Avoid when

  • Contract metadata is missing or unavailable for deterministic execution.

Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing

Public facts

Every fact links back to the source it came from.

Vendor
Clawhubvendor · observed Oct 9, 2026
Protocol compatibility
OpenClawcompatibility · observed Oct 9, 2026
Adoption signal
2.4K downloadsadoption · observed Oct 9, 2026
Latest release
1.0.17release · observed Sep 28, 2026
Handshake status
UNKNOWNsecurity

Install and run

Setup complexity: low.

clawhub skill install s17f8q65zg3y98t86jdg1177g583whq8:smyx-visual-qa-analysis
  1. Setup complexity is classified as HIGH. You must provision dedicated cloud infrastructure or an isolated VM. Do not run this directly on your local workstation.
  2. Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.

Contract: missing

curl -s "https://www.xpersona.co/api/v1/agents/clawhub-18072937735-smyx-visual-qa-analysis/snapshot"

Documentation

CLAWHUB

124,759 characters of source documentation, loaded on request.

Extracted files

5 files captured from the source.

SKILL.md

---
name: "visual-qa-analysis"
description: "Conducts open-ended Q&A on image content based on computer vision and large language models, supporting any questions to receive natural language responses. | 大模型视觉问答(VQA)技能,基于计算机视觉和大语言模型对图片内容进行开放式问答,支持任意提问得到自然语言回答"
version: "1.0.17"
license: "MIT-0"
---

# ❓ Large Model Visual Question Answering Skill | 大模型视觉问答技能
> **智能分析中枢** · 图片/视频智能分析 · 结构化报告 · 历史报告云端查询

---

## 🧭 技能概览 | Overview

| 模块 | 内容 |
|---|---|
| 🏷️ 技能名称 | **大模型视觉问答技能** |
| 🎯 核心目标 | 大模型视觉问答(VQA)技能,基于计算机视觉和大语言模型对图片内容进行开放式问答,支持任意提问得到自然语言回答 |
| 🖼️ 输入类型 | 图片、视频、本地文件、网络 URL |
| 📝 输出能力 | 结构化分析报告、识别/监测结果、建议与报告链接 |
| 🧩 场景码 | `VISUAL_QA` |

Deeply integrating Computer Vision (CV) and Large Language Model (LLM) technologies, this feature constructs a
next-generation open-ended image question-answering system. Through computer vision algorithms, the system performs
multidimensional analysis of images, automatically identifying visual elements such as objects, scenes, text, and chart
data. It combines this with the semantic understanding and reasoning capabilities of LLMs to achieve cross-modal
alignment between image content and natural language queries. Users can pose open-ended questions to any image (e.g., "
What is the core trend of this chart?" or "Which period does the architectural style in the picture belong to?").
Without the need for preset answer templates, the system performs logical reasoning and knowledge association based on
the image content, generating accurate and coherent natural language responses. Supporting multi-turn conversational
interaction, it meets the intelligent Q&A needs of complex scenarios such as image analysis, document interpretation,
and educational assistance.

本功能深度融合计算机视觉(CV)与大语言模型(LLM)技术,构建了新一代开放式图片问答系统。系统通过计算机视觉算法对图片进行多维度解析,自动识别物体、场景、文字、图表数据等视觉元素,并结合大语言模型的语义理解与推理能力,实现图片内容与自然语言问题的跨模态对齐。用户可对任意图片提出开放式问题(如“这张图表的核心趋势是什么?”“图片中的建筑风格属于哪个时期?”),系统无需预设答案模板,即可基于图片内容进行逻辑推理与知识关联,生成准确、连贯的自然语言回答,支持多轮对话交互,满足图像分析、文档解读、教育辅助等复杂场景下的智能问答需求

## 🎬 技能演示 | Skill Demo

[▶️ 点击查看技能使用介绍](https://lifeemergence.com/sample.html)

---

## 🎯 任务目标 | Goals

### 1. 🧩 技能用途

通过图片结合用户问题进行大模型视觉问答,获得自然语言回答

### 2. 🛠️ 能力范围

| 序号 | 具体能力 |
|---:|---|
| 1 | 图片内容理解 |
| 2 | 开放式问答 |
| 3 | 场景描述 |
| 4 | 细节识别 |
| 5 | 知识推理 |

### 3. ⚡ 触发条件

| 触发类型 | 触发规则 |
|---|---|
| ✅ 默认触发 | **默认触发**:当用户提供图片 URL 或文件,并提出问题需要对图片进行问答时,默认触发本技能 |
| 🔎 明确分析意图 | 当用户明确需要进行视觉问答,提及 VQA、看图问答、图片问答、视觉问答等关键词,并且上传了图片 |
| 📚 历史报告查询 | 当用户提及以下关键词时,**自动触发历史问答记录查询功能** :查看历史问答记录、视觉问答历史、问答记录清单、查询历史问答,显示所有问答记录 |
| 触发规则 4 | 用户提供图片后附带问题,如"这张图片里有什么?",直接触发视觉问答 |

### 4. 🤖 自动行为

| 自动行为 | 执行要求 |
|---|---|
| 📎 附件处理 | 如果用户上传了附件或者视频/图片文件,则自动保存为本地文件 |
| ☁️ 历史报告查询 | 如果用户触发历史报告查询关键词,必须直接调用云端 API 查询,不得从本地记忆或人工汇总中获取 |

#### ⚠️ 强制数据获取规则(次高优先级)

> **橙色强约束:** 历史报告清单只允许从云端接口读取,不允许从本地记录、长期记忆或人工汇总中提取。

必须执行:

```bash
python -m scripts.visual_qa_analysis --list
```

| 类型 | 要求 |
|---|---|
| ✅ 必须 | 使用 `python -m scripts.visual_qa_analysis --list` 调用 API 查询云端的历史报告数据 |
| 🚫 严格禁止 | 从本地 `memory` 目录读取历史会话信息 |
| 🚫 严

_meta.json

{
  "ownerId": "kn7e2caqj7pnsvr9r7t8zenghs83xw7n",
  "slug": "smyx-visual-qa-analysis",
  "version": "1.0.17",
  "publishedAt": 1790575466780
}

references/api_doc.md

# API 接口文档

此处用于存放宠物健康分析 API 的接口文档,待后续补充。

## 接口规范

- 基础地址:由 smyx_common 配置统一管理
- 认证方式:API Key 鉴权
- 请求格式:支持文件上传
- 响应格式:JSON

## 主要接口

1. `/web/health-analysis/v2/start-health-analysis` - 启动健康分析任务
2. `/web/health-analysis/v2/get-health-analysis-result` - 获取分析结果
3. `/web/health-analysis/page-health-analysis-result` - 分页查询历史报告
4. `/health/order/api/getReportDetailExport?id={id}` - 导出完整报告

## 场景代码

- `OPEN_PET_HEALTH_ANALYSIS` - 开放平台宠物健康分析

skills/smyx_analysis/references/api_doc.md

# API接口文档

## 接口规范

- 基础地址:由 smyx_common 配置统一管理
- 认证方式:API Key 鉴权
- 请求格式:支持文件上传
- 响应格式:JSON

## 错误码说明

| 错误码 | 说明       |
|-----|----------|
| 400 | 请求参数错误   |
| 401 | API密钥无效  |
| 403 | 权限不足     |
| 413 | 文件过大     |
| 415 | 不支持的文件格式 |
| 500 | 服务器内部错误  |

scripts/config.yaml

{}
Github ReposUpdated 2h agoRank 70

AionUi

Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!

MCPOPENCLAW
Github ReposUpdated 6mo agoRank 70

activepieces

AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents

OPENCLAW
Github ReposUpdated 6mo agoRank 70

cherry-studio

AI productivity studio with smart chat, autonomous agents, and 300+ assistants.

MCPOPENCLAW
Github ReposUpdated 7mo agoRank 70

CopilotKit

The Frontend for Agents & Generative UI. React + Angular

OPENCLAW

Machine-readable data

The same record, as JSON, for agents and crawlers.

{
  "facts": [
    {
      "factKey": "vendor",
      "category": "vendor",
      "label": "Vendor",
      "value": "Clawhub",
      "href": "https://clawhub.ai/18072937735/skills/smyx-visual-qa-analysis",
      "sourceUrl": "https://clawhub.ai/18072937735/skills/smyx-visual-qa-analysis",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-09T15:38:45.006Z",
      "isPublic": true
    },
    {
      "factKey": "protocols",
      "category": "compatibility",
      "label": "Protocol compatibility",
      "value": "OpenClaw",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-18072937735-smyx-visual-qa-analysis/contract",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-18072937735-smyx-visual-qa-analysis/contract",
      "sourceType": "contract",
      "confidence": "medium",
      "observedAt": "2026-10-09T15:38:45.006Z",
      "isPublic": true
    },
    {
      "factKey": "traction",
      "category": "adoption",
      "label": "Adoption signal",
      "value": "2.4K downloads",
      "href": "https://clawhub.ai/18072937735/smyx-visual-qa-analysis",
      "sourceUrl": "https://clawhub.ai/18072937735/smyx-visual-qa-analysis",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-09T15:38:45.006Z",
      "isPublic": true
    },
    {
      "factKey": "latest_release",
      "category": "release",
      "label": "Latest release",
      "value": "1.0.17",
      "href": "https://clawhub.ai/18072937735/smyx-visual-qa-analysis",
      "sourceUrl": "https://clawhub.ai/18072937735/smyx-visual-qa-analysis",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-09-28T06:04:26.780Z",
      "isPublic": true
    },
    {
      "factKey": "handshake_status",
      "category": "security",
      "label": "Handshake status",
      "value": "UNKNOWN",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-18072937735-smyx-visual-qa-analysis/trust",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-18072937735-smyx-visual-qa-analysis/trust",
      "sourceType": "trust",
      "confidence": "medium",
      "observedAt": null,
      "isPublic": true
    }
  ],
  "events": [
    {
      "eventType": "release",
      "title": "Release 1.0.17",
      "description": "- Version bumped to 1.0.17. - Documentation updated in SKILL.md. - Obsolete file skill-card.md removed. - Possible adjustments to configuration file (config.yaml).",
      "href": "https://clawhub.ai/18072937735/smyx-visual-qa-analysis",
      "sourceUrl": "https://clawhub.ai/18072937735/smyx-visual-qa-analysis",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-09-28T06:04:26.780Z",
      "isPublic": true
    }
  ]
}

Record generated Oct 9, 2026.

Sponsored

Ads related to Large Model Visual Question Answering Skill | 大模型视觉问答技能 and adjacent AI workflows.