Large Model Visual Question Answering Skill | 大模型视觉问答技能
Conducts open-ended Q&A on image content based on computer vision and large language models, supporting any questions to receive natural language responses. | 大模型视觉问答(VQA)技能,基于计算机视觉和大语言模型对图片内容进行开放式问答,支持任意提问得到自然语言回答 Skill: Large Model Visual Question Answering Skill | 大模型视觉问答技能 Owner: 18072937735 Summary: Conducts open-ended Q&A on image content based on computer vision and large language models, supporting any questions to receive natural language responses. | 大模型视觉问答(VQA)技能,基于计算机视觉和大语言模型对图片内容进行开放式问答,支持任意提问得到自然语言回答 Tags: latest:1.0.17 Version history: v1.0.17 | 2026-09-28T06:04:26.780Z | auto - Version bumped to 1.0.17. - Docum
Rank
62
Safety
84
Downloads
2.4k
Updated
Oct 9, 2026
Version
1.0.17
Source
CLAWHUB
About
What it does, and when to use it.
Capability contract not published. No trust telemetry is available yet. 2.4K downloads reported by the source. Last updated 10/9/2026.
Avoid when
- Contract metadata is missing or unavailable for deterministic execution.
Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing
Public facts
Every fact links back to the source it came from.
- Vendor
- Clawhubvendor · observed Oct 9, 2026
- Protocol compatibility
- OpenClawcompatibility · observed Oct 9, 2026
- Adoption signal
- 2.4K downloadsadoption · observed Oct 9, 2026
- Latest release
- 1.0.17release · observed Sep 28, 2026
- Handshake status
- UNKNOWNsecurity
Install and run
Setup complexity: low.
clawhub skill install s17f8q65zg3y98t86jdg1177g583whq8:smyx-visual-qa-analysis- Setup complexity is classified as HIGH. You must provision dedicated cloud infrastructure or an isolated VM. Do not run this directly on your local workstation.
- Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.
Contract: missing
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-18072937735-smyx-visual-qa-analysis/snapshot"
Documentation
CLAWHUB
124,759 characters of source documentation, loaded on request.
Extracted files
5 files captured from the source.
SKILL.md
--- name: "visual-qa-analysis" description: "Conducts open-ended Q&A on image content based on computer vision and large language models, supporting any questions to receive natural language responses. | 大模型视觉问答(VQA)技能,基于计算机视觉和大语言模型对图片内容进行开放式问答,支持任意提问得到自然语言回答" version: "1.0.17" license: "MIT-0" --- # ❓ Large Model Visual Question Answering Skill | 大模型视觉问答技能 > **智能分析中枢** · 图片/视频智能分析 · 结构化报告 · 历史报告云端查询 --- ## 🧭 技能概览 | Overview | 模块 | 内容 | |---|---| | 🏷️ 技能名称 | **大模型视觉问答技能** | | 🎯 核心目标 | 大模型视觉问答(VQA)技能,基于计算机视觉和大语言模型对图片内容进行开放式问答,支持任意提问得到自然语言回答 | | 🖼️ 输入类型 | 图片、视频、本地文件、网络 URL | | 📝 输出能力 | 结构化分析报告、识别/监测结果、建议与报告链接 | | 🧩 场景码 | `VISUAL_QA` | Deeply integrating Computer Vision (CV) and Large Language Model (LLM) technologies, this feature constructs a next-generation open-ended image question-answering system. Through computer vision algorithms, the system performs multidimensional analysis of images, automatically identifying visual elements such as objects, scenes, text, and chart data. It combines this with the semantic understanding and reasoning capabilities of LLMs to achieve cross-modal alignment between image content and natural language queries. Users can pose open-ended questions to any image (e.g., " What is the core trend of this chart?" or "Which period does the architectural style in the picture belong to?"). Without the need for preset answer templates, the system performs logical reasoning and knowledge association based on the image content, generating accurate and coherent natural language responses. Supporting multi-turn conversational interaction, it meets the intelligent Q&A needs of complex scenarios such as image analysis, document interpretation, and educational assistance. 本功能深度融合计算机视觉(CV)与大语言模型(LLM)技术,构建了新一代开放式图片问答系统。系统通过计算机视觉算法对图片进行多维度解析,自动识别物体、场景、文字、图表数据等视觉元素,并结合大语言模型的语义理解与推理能力,实现图片内容与自然语言问题的跨模态对齐。用户可对任意图片提出开放式问题(如“这张图表的核心趋势是什么?”“图片中的建筑风格属于哪个时期?”),系统无需预设答案模板,即可基于图片内容进行逻辑推理与知识关联,生成准确、连贯的自然语言回答,支持多轮对话交互,满足图像分析、文档解读、教育辅助等复杂场景下的智能问答需求 ## 🎬 技能演示 | Skill Demo [▶️ 点击查看技能使用介绍](https://lifeemergence.com/sample.html) --- ## 🎯 任务目标 | Goals ### 1. 🧩 技能用途 通过图片结合用户问题进行大模型视觉问答,获得自然语言回答 ### 2. 🛠️ 能力范围 | 序号 | 具体能力 | |---:|---| | 1 | 图片内容理解 | | 2 | 开放式问答 | | 3 | 场景描述 | | 4 | 细节识别 | | 5 | 知识推理 | ### 3. ⚡ 触发条件 | 触发类型 | 触发规则 | |---|---| | ✅ 默认触发 | **默认触发**:当用户提供图片 URL 或文件,并提出问题需要对图片进行问答时,默认触发本技能 | | 🔎 明确分析意图 | 当用户明确需要进行视觉问答,提及 VQA、看图问答、图片问答、视觉问答等关键词,并且上传了图片 | | 📚 历史报告查询 | 当用户提及以下关键词时,**自动触发历史问答记录查询功能** :查看历史问答记录、视觉问答历史、问答记录清单、查询历史问答,显示所有问答记录 | | 触发规则 4 | 用户提供图片后附带问题,如"这张图片里有什么?",直接触发视觉问答 | ### 4. 🤖 自动行为 | 自动行为 | 执行要求 | |---|---| | 📎 附件处理 | 如果用户上传了附件或者视频/图片文件,则自动保存为本地文件 | | ☁️ 历史报告查询 | 如果用户触发历史报告查询关键词,必须直接调用云端 API 查询,不得从本地记忆或人工汇总中获取 | #### ⚠️ 强制数据获取规则(次高优先级) > **橙色强约束:** 历史报告清单只允许从云端接口读取,不允许从本地记录、长期记忆或人工汇总中提取。 必须执行: ```bash python -m scripts.visual_qa_analysis --list ``` | 类型 | 要求 | |---|---| | ✅ 必须 | 使用 `python -m scripts.visual_qa_analysis --list` 调用 API 查询云端的历史报告数据 | | 🚫 严格禁止 | 从本地 `memory` 目录读取历史会话信息 | | 🚫 严
_meta.json
{
"ownerId": "kn7e2caqj7pnsvr9r7t8zenghs83xw7n",
"slug": "smyx-visual-qa-analysis",
"version": "1.0.17",
"publishedAt": 1790575466780
}references/api_doc.md
# API 接口文档
此处用于存放宠物健康分析 API 的接口文档,待后续补充。
## 接口规范
- 基础地址:由 smyx_common 配置统一管理
- 认证方式:API Key 鉴权
- 请求格式:支持文件上传
- 响应格式:JSON
## 主要接口
1. `/web/health-analysis/v2/start-health-analysis` - 启动健康分析任务
2. `/web/health-analysis/v2/get-health-analysis-result` - 获取分析结果
3. `/web/health-analysis/page-health-analysis-result` - 分页查询历史报告
4. `/health/order/api/getReportDetailExport?id={id}` - 导出完整报告
## 场景代码
- `OPEN_PET_HEALTH_ANALYSIS` - 开放平台宠物健康分析skills/smyx_analysis/references/api_doc.md
# API接口文档 ## 接口规范 - 基础地址:由 smyx_common 配置统一管理 - 认证方式:API Key 鉴权 - 请求格式:支持文件上传 - 响应格式:JSON ## 错误码说明 | 错误码 | 说明 | |-----|----------| | 400 | 请求参数错误 | | 401 | API密钥无效 | | 403 | 权限不足 | | 413 | 文件过大 | | 415 | 不支持的文件格式 | | 500 | 服务器内部错误 |
scripts/config.yaml
{}activepieces
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
AionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
CopilotKit
The Frontend for Agents & Generative UI. React + Angular
Machine-readable data
The same record, as JSON, for agents and crawlers.
{
"facts": [
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Clawhub",
"href": "https://clawhub.ai/18072937735/skills/smyx-visual-qa-analysis",
"sourceUrl": "https://clawhub.ai/18072937735/skills/smyx-visual-qa-analysis",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-09T15:38:45.006Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-18072937735-smyx-visual-qa-analysis/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-18072937735-smyx-visual-qa-analysis/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-09T15:38:45.006Z",
"isPublic": true
},
{
"factKey": "traction",
"category": "adoption",
"label": "Adoption signal",
"value": "2.4K downloads",
"href": "https://clawhub.ai/18072937735/smyx-visual-qa-analysis",
"sourceUrl": "https://clawhub.ai/18072937735/smyx-visual-qa-analysis",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-09T15:38:45.006Z",
"isPublic": true
},
{
"factKey": "latest_release",
"category": "release",
"label": "Latest release",
"value": "1.0.17",
"href": "https://clawhub.ai/18072937735/smyx-visual-qa-analysis",
"sourceUrl": "https://clawhub.ai/18072937735/smyx-visual-qa-analysis",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-09-28T06:04:26.780Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-18072937735-smyx-visual-qa-analysis/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-18072937735-smyx-visual-qa-analysis/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
],
"events": [
{
"eventType": "release",
"title": "Release 1.0.17",
"description": "- Version bumped to 1.0.17. - Documentation updated in SKILL.md. - Obsolete file skill-card.md removed. - Possible adjustments to configuration file (config.yaml).",
"href": "https://clawhub.ai/18072937735/smyx-visual-qa-analysis",
"sourceUrl": "https://clawhub.ai/18072937735/smyx-visual-qa-analysis",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-09-28T06:04:26.780Z",
"isPublic": true
}
]
}Record generated Oct 9, 2026.
