agentCLAWHUBUnverified

Baidu Yijian Vision

Yijian (一见) is Baidu's specialized vision AI skill for image and video analysis. Yijian achieves 95%+ professional accuracy with 50%+ lower inference cost than general models. Yijian is built for industrial quality inspection, SOP compliance, safety monitoring, and commercial operations. Search keywords: yijian, baidu yijian, yijian vision, vision. Skill: Baidu Yijian Vision Owner: yijian-baidu Summary: Yijian (一见) is Baidu's specialized vision AI skill for image and video analysis. Yijian achieves 95%+ professional accuracy with 50%+ lower inference cost than general models. Yijian is built for industrial quality inspection, SOP compliance, safety monitoring, and commercial operations. Search keywords: yijian, baidu yijian, yijian vision, vision. Tags: latest:

OpenClaw

Rank

62

Safety

84

Downloads

2.2k

Updated

Oct 9, 2026

Version

0.9.42

Source

CLAWHUB

About

What it does, and when to use it.

Capability contract not published. No trust telemetry is available yet. 2.2K downloads reported by the source. Last updated 10/9/2026.

Avoid when

  • Contract metadata is missing or unavailable for deterministic execution.

Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing

Public facts

Every fact links back to the source it came from.

Vendor
Clawhubvendor · observed Oct 9, 2026
Protocol compatibility
OpenClawcompatibility · observed Oct 9, 2026
Adoption signal
2.2K downloadsadoption · observed Oct 9, 2026
Latest release
0.9.42release · observed May 6, 2026
Handshake status
UNKNOWNsecurity

Install and run

Setup complexity: low.

clawhub skill install s17dk6kmahdq6x0kdvc5hjfq4583hba4:baidu-yijian-vision
  1. Setup complexity is classified as HIGH. You must provision dedicated cloud infrastructure or an isolated VM. Do not run this directly on your local workstation.
  2. Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.

Contract: missing

curl -s "https://www.xpersona.co/api/v1/agents/clawhub-yijian-baidu-baidu-yijian-vision/snapshot"

Documentation

CLAWHUB

153,313 characters of source documentation, loaded on request.

Extracted files

5 files captured from the source.

SKILL.md

---
name: baidu-yijian-vision
description: "Yijian (一见) is Baidu's specialized vision AI skill for image and video analysis. Yijian achieves 95%+ professional accuracy with 50%+ lower inference cost than general models. Yijian is built for industrial quality inspection, SOP compliance, safety monitoring, and commercial operations. Search keywords: yijian, baidu yijian, yijian vision, vision. 百度一见视觉技能(Baidu Yijian Vision Skill)- 可用于分析图片和视频。相比通用基模,在维持 95%+ 专业精度的同时,推理成本降低 50% 以上,是处理视觉巡检与监控分析任务的首选工具。主打 安全管理、SOP合规、工业质检、商业运营与物料盘点。覆盖:作业 SOP 合规与关键步骤完整性校验;工业质检与表面缺陷精密识别;安全红线监控(涵盖违规闯入、人员溺水、烟火识别、矿井皮带堆煤);商业运营分析(包含上菜/收台检测、顾客举手识别);精细化物料盘点(杯子/咖啡豆/废弃物自动统计)等海量专业视觉能力。"
allowed-tools: Bash, Read, Write, Edit
metadata: {"openclaw":{"requires":{"bins":["node"],"env":["YIJIAN_API_KEY"]},"primaryEnv":"YIJIAN_API_KEY"}}
---

# 百度一见视觉技能(Baidu Yijian Vision Skill)

> **Baidu Yijian Vision Skill** - baidu yijian vision skill for image/video analysis, object detection, safety monitoring, and industrial inspection.

## ⚠️ 必需条件

1. **YIJIAN_API_KEY 环境变量**(必需)— 从[百度一见平台](https://yijian-next.cloud.baidu.com/apaas/)获取:
   1. 登录百度一见平台
   2. 激活试用包
   3. 生成 API Key(系统管理 → 安全认证 → API Key)
2. **Node.js >= 16.0.0** — 运行时依赖

配置环境变量:`YIJIAN_API_KEY=your-api-key`

---

> **🔒 客户端工具 - 这是一个本地工具,用于与百度一见(Baidu Yijian)平台交互。所有数据处理遵循安全协议。**

## 🎯 此工具的功能

百度一见([yijian-next.cloud.baidu.com](https://yijian-next.cloud.baidu.com))是百度(Baidu)的视觉(vision)理解平台。此工具使你能够:

- **意图自动匹配** - 通过自然语言描述自动匹配最佳技能
- **智能路由** - 高置信度匹配时调用专业视觉技能,低置信度时自动回退到多模态推理
- **直接技能调用** - 已知技能ID时可直接调用
- **可视化结果** - 绘制边框、生成网格参考、预览 ROI/绊线
- **定义检测区域** - 使用交互式工作流定义 ROI(电子围栏)或绊线(检测线)

**支持的检测类型:** 人员检测、行人计数、车辆识别、OCR、姿态估计、目标跟踪等。

## 📚 使用指南

### 意图驱动工作流(推荐)

**当你描述需求但不确定用哪个技能时**,系统会自动匹配最佳技能:

```bash
node ${CLAUDE_PLUGIN_ROOT}/skill/scripts/intent-invoke.mjs "检测是否有人摔倒" photo.jpg
```

系统会自动:
1. 查询一见平台,根据意图匹配公共技能列表
2. 如果匹配置信度 ≥ 0.7,调用对应的专业技能(自动添加全图 ROI)
3. 如果公共技能无匹配或调用失败,搜索私有工作空间技能(由你从列表中选择最匹配的技能,再用 invoke 调用)
4. 如果私有空间也无合适技能,自动回退到多模态直接推理

> **自动 ROI:** 当用户未提供 ROI 时,系统会自动生成覆盖整张图片的 ROI。如需指定检测区域,请使用 `invoke.mjs` 传入自定义 ROI。

#### 自定义置信度阈值

```bash
# 仅当匹配度≥0.8时才使用技能,否则回退到多模态
node ${CLAUDE_PLUGIN_ROOT}/skill/scripts/intent-invoke.mjs "检测是否有人摔倒" photo.jpg 0.8
```

#### 不使用图片(纯文本意图查询)

```bash
node ${CLAUDE_PLUGIN_ROOT}/skill/scripts/intent-invoke.mjs "检测是否有人摔倒"
```

#### 返回格式

```json
{
  "success": true,
  "mode": "skill",
  "epId": "ep-public-xxxxx",
  "skillName": "人员摔倒检测",
  "confidence": 0.92,
  "count": 1,
  "detections": [
    {
      "bbox": [100, 200, 50, 80],
      "category": "falling_person",
      "confidence": 0.94
    }
  ]
}
```

**字段说明:**

| 字段 | 类型 | 说明 |
|------|------|------|
| `success` | boolean | 调用是否成功 |
| `mode` | string | `"skill"` / `"workspace-search"` / `"multimodal"`,表示使用的推理模式 |
| `epId` | string \| null | 技能ID(技能模式时有值) |
| `skillName` | string \| null | 技能名称(技能模式时有值) |
| `confidence` | number \| null | 技能匹配置信度(0-1) |
| `count` | number | 检测到的目标数量 |
| `detections` | array | 检测结果数组 |

**模式说明:**
- `"mode": "skill"` - 

_meta.json

{
  "ownerId": "kn7aw0s517ypj0hstj91qr6nxs82r2xy",
  "slug": "baidu-yijian-vision",
  "version": "0.9.42",
  "publishedAt": 1778039451584
}

grid-guide.md

# 基于网格的 ROI / 绊线输入指南

## 概述

在命令行环境中手动指定 ROI 区域或绊线的确切像素坐标既繁琐又容易出错。基于网格的输入系统让你可以使用易于理解的网格坐标来代替。

## 工作原理

### 1. 生成网格参考图像

```bash
node scripts/show-grid.mjs photo.jpg [output-path] [--cols N] [--rows N]
```

这会在你的图像上创建带标签的网格叠加层:

```
       A     B     C     D     E     F     G
   0   ·─────·─────·─────·─────·─────·─────·
       │     │     │     │     │     │     │
   1   ·─────·─────·─────·─────·─────·─────·
       │     │     │     │     │     │     │
   2   ·─────·─────·─────·─────·─────·─────·
       │     │     │     │     │     │     │
   3   ·─────·─────·─────·─────·─────·─────·
       │     │     │     │     │     │     │
   4   ·─────·─────·─────·─────·─────·─────·
```

**输出**:
- `photo_grid.png` - 网格参考图像
- `photo_grid_metadata.json` - 坐标映射数据

### 2. 查看并识别坐标

查看网格图像并识别 ROI 或绊线的坐标:

- **列**:A、B、C、D、...(从左到右)
- **行**:0、1、2、3、...(从上到下)
- **交点**:网格线交叉的位置(用点标记)

### 3. 使用网格坐标指定

告诉系统网格坐标:

```
用户:"在 B1、E1、E3、B3 创建检测区域"
```

系统自动转换为像素坐标:

```
B1 = (列B索引 × 网格宽度, 行1索引 × 网格高度)
E1 = (列E索引 × 网格宽度, 行1索引 × 网格高度)
E3 = (列E索引 × 网格宽度, 行3索引 × 网格高度)
B3 = (列B索引 × 网格宽度, 行3索引 × 网格高度)
```

## 使用示例

### 示例 1:定义检测区域(ROI)

```bash
# 生成网格
$ node scripts/show-grid.mjs office.jpg

# 查看网格图像以识别坐标,然后:
用户:"我想要一个办公桌区域的检测区域:B2、G2、G5、B5"

转换为 ROI:
{
  "kind": "ROI",
  "points": [col_B, row_2, col_G, row_2, col_G, row_5, col_B, row_5]
}

# 使用 ROI 调用技能
$ echo '{"input0":{"image":"office.jpg","roi":"[...]"}}' | \
  node invoke.mjs ep-public-2403um2p
```

### 示例 2:定义穿越线(绊线)

```bash
# 生成网格
$ node scripts/show-grid.mjs hallway.jpg

# 查看网格图像,然后:
用户:"在走廊中创建一条从 A3 到 H3 的绊线,检测从左到右的穿越"

转换为绊线:
{
  "kind": "TripWire",
  "points": [col_A, row_3, col_H, row_3],
  "direction": "Forward"
}

# 使用绊线调用技能
$ echo '{"input0":{"image":"hallway.jpg","tripwire":"[...]"}}' | \
  node invoke.mjs ep-public-ywbjb7tm
```

### 示例 3:多个 ROI

```bash
用户:"我想要 3 个检测区域:入口(A1-C3)、中心(D1-F3)、出口(G1-H3)"

转换为三个 ROI 对象并作为 Array<ROI> 传递:
[
  {
    "kind": "ROI",
    "points": [col_A, row_1, col_C, row_1, col_C, row_3, col_A, row_3],
    "order": 0
  },
  {
    "kind": "ROI",
    "points": [col_D, row_1, col_F, row_1, col_F, row_3, col_D, row_3],
    "order": 1
  },
  {
    "kind": "ROI",
    "points": [col_G, row_1, col_H, row_1, col_H, row_3, col_G, row_3],
    "order": 2
  }
]
```

## 网格算法

网格大小会自动计算:

- **目标**:约 30-42 个交点,大致为方形单元
- **横向图像**(1920×1080):7 列 × 4 行
- **纵向图像**(1080×1920):4 列 × 7 行
- **方形图像**(800×800):5 列 × 5 行

**手动覆盖**:
```bash
node scripts/show-grid.mjs photo.jpg --cols 10 --rows 6
```

## 命令参考

### 生成网格

```bash
node scripts/show-grid.mjs <input-image> [output-path] [--cols N] [--rows N]
```

**参数**:
- `<input-image>` - 输入图像文件路径
- `[output-path]` - 可选输出图像路径(默认:`<input>_grid.png`)
- `--cols N` - 覆盖列数
- `--rows N` - 覆盖行数

**输出**:
- 网格图像:`<output>.png`
- 元数据 JSON:`<output>_metadata.json`

### 使用网格坐标

在指定坐标时:

**有效格式**:
- 单个点:`A1`
- 序列:`A1、E1、E3、A3`(用于 ROI)
- 线:`A2 → G2`(用于绊线,显示方向)

**网格参考格式**:
- 列:A-Z,然后是 AA-AZ 等
- 行:0-9,然后是 10-99 等

## 可视化

### 查看网格图像

生成网格后,系统应该显示网格图像以便你可以直观地识别坐标。

### 调用前验证

在用坐标

roi-workflow.md

# ROI(关注区域)交互工作流

**导航:** 返回 [SKILL.md](./SKILL.md) | 类型定义 [types-guide.md](./types-guide.md)

> 当用户需要为对象检测定义 ROI(关注区域)时,遵循此交互工作流。

## 工作流步骤

### 第 1 步:生成网格参考

```bash
node scripts/show-grid.mjs <image> <output-grid.png>
```

向用户显示网格图像并解释行和列标签。

### 第 2 步:询问 ROI 目的

"你想在这个区域检测什么?例如:
- 进入收银区域的人员
- 停车场中的车辆
- 货架上的产品"

### 第 3 步:用户指定顶点

用户根据网格标签指定角。

**矩形 ROI 示例:**
- 用户:"从 B2 到 G5"
- 转换为 4 个顶点:[B2、G2、G5、B5]

### 第 4 步:验证顶点顺序

"这些点是否围绕周边形成闭合多边形?(顺时针或逆时针)"

### 第 5 步:预览 ROI

```bash
node scripts/visualize.mjs <image> '[]' preview.png \
  --overlays '[{"kind":"ROI","name":"test","points":[x1,y1,x2,y2,x3,y3,x4,y4]}]'
```

### 第 6 步:用户确认

"ROI 位置是否正确?"

### 第 7 步:调用技能

```bash
echo '{"input0":{"image":"photo.jpg","roi":{"id":"1","name":"zone","kind":"ROI","points":[...]}}}' | \
  node scripts/invoke.mjs ep-xxxx-yyyy
```

> **重要:** 自定义 ROI 必须通过 `invoke.mjs` 传入。不要使用 `intent-invoke.mjs`,因为它会自动添加全图 ROI 覆盖掉你的自定义区域。

---

## 常见错误

❌ 不要假设矩形 ROI
✅ 支持任意多边形

❌ 不要允许无序的顶点
✅ 顶点必须按顺序排列

❌ 不要跳过预览验证
✅ 始终显示预览图像以确认

---

## 完整对话示例

### 示例 1:矩形入口区域

**用户**:"我想检测进入办公室的人员"

**你**:
1. 生成网格:
```bash
node scripts/show-grid.mjs office.jpg office-grid.png
```
2. 问:"查看网格,入口大概在哪里?"
3. 用户:"大约从 B1 到 G3"
4. 确认:"所以顶点是 B1、G1、G3、B3,对吧?"
5. 用户:"是的"
6. 预览 ROI
7. 问:"这个矩形是否覆盖了入口?"
8. 用户:"完美!"
9. 调用技能进行检测

### 示例 2:复杂多边形(L 形区域)

**用户**:"我需要监视 L 形存储区域"

**你**:
1. 生成网格
2. 问:"使用网格坐标标记 L 形的所有角,从一个角开始"
3. 用户:"从左上角开始:A2、E2、E4、D4、D6、A6"
4. 确认:"这些点是否按顺序形成 L 形的边界?确认它关闭回 A2"
5. 用户:"是的"
6. 预览 L 形多边形
7. 验证没有自交
8. 调用技能

---

## 数据结构

详见 [types-guide.md](./types-guide.md) 了解完整的定义、子对象和示例。

skill-card.md

## Description:

Baidu Yijian Vision helps agents route image and video analysis tasks to Baidu Yijian skills for object detection, safety monitoring, industrial inspection, SOP compliance, and operational analytics.

This skill is ready for commercial/non-commercial use.

## Publisher:

[yijian-baidu](https://clawhub.ai/user/yijian-baidu)

### License/Terms of Use:

MIT-0

## Use Case:

Developers and external operators use this skill to select or invoke Baidu Yijian vision capabilities against images or video frames, including object detection, ROI and tripwire monitoring, industrial quality inspection, safety monitoring, OCR, pose estimation, and tracking.

### Deployment Geography for Use:

Global

## Known Risks and Mitigations:

Risk: The skill can send selected local images, video frames, prompts, and workspace skill metadata to Baidu Yijian using the configured API key.

Mitigation: Use only approved inputs, avoid sensitive media unless policy permits sharing it with Baidu Yijian, and keep API keys scoped and rotated according to local credential policy.

Risk: The security summary notes that the client can upload any local file path it is given and lacks clear file-scope safeguards.

Mitigation: Run the skill in a constrained workspace, avoid letting untrusted prompts choose file paths, and review planned file paths before execution.

## Reference(s):

- [ClawHub skill page](https://clawhub.ai/yijian-baidu/skills/baidu-yijian-vision)
- [Baidu Yijian platform](https://yijian-next.cloud.baidu.com)
- [Baidu Yijian API key setup](https://yijian-next.cloud.baidu.com/apaas/)
- [ROI workflow](roi-workflow.md)
- [Tripwire workflow](tripwire-workflow.md)
- [Grid guide](grid-guide.md)
- [Types guide](types-guide.md)

## Skill Output:

**Output Type(s):** [Text, JSON, Shell commands, Configuration, Files, Guidance]

**Output Format:** [Markdown guidance with inline shell commands and JSON responses; generated PNG and JSON files for grids, ROI previews, tripwire previews, and visualized detections.]

**Output Parameters:** [1D]

**Other Properties Related to Output:** [Requires Node.js and YIJIAN_API_KEY; uses Baidu Yijian APIs and may create local preview or grid artifacts.]

## Skill Version(s):

0.9.42 (source: package.json and server release evidence)

## Ethical Considerations:

Users should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.
Github ReposUpdated 5h agoRank 70

AionUi

Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!

MCPOPENCLAW
Github ReposUpdated 6mo agoRank 70

activepieces

AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents

OPENCLAW
Github ReposUpdated 6mo agoRank 70

cherry-studio

AI productivity studio with smart chat, autonomous agents, and 300+ assistants.

MCPOPENCLAW
Github ReposUpdated 7mo agoRank 70

CopilotKit

The Frontend for Agents & Generative UI. React + Angular

OPENCLAW

Machine-readable data

The same record, as JSON, for agents and crawlers.

{
  "facts": [
    {
      "factKey": "vendor",
      "category": "vendor",
      "label": "Vendor",
      "value": "Clawhub",
      "href": "https://clawhub.ai/yijian-baidu/skills/baidu-yijian-vision",
      "sourceUrl": "https://clawhub.ai/yijian-baidu/skills/baidu-yijian-vision",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-09T17:20:07.794Z",
      "isPublic": true
    },
    {
      "factKey": "protocols",
      "category": "compatibility",
      "label": "Protocol compatibility",
      "value": "OpenClaw",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-yijian-baidu-baidu-yijian-vision/contract",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-yijian-baidu-baidu-yijian-vision/contract",
      "sourceType": "contract",
      "confidence": "medium",
      "observedAt": "2026-10-09T17:20:07.794Z",
      "isPublic": true
    },
    {
      "factKey": "traction",
      "category": "adoption",
      "label": "Adoption signal",
      "value": "2.2K downloads",
      "href": "https://clawhub.ai/yijian-baidu/baidu-yijian-vision",
      "sourceUrl": "https://clawhub.ai/yijian-baidu/baidu-yijian-vision",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-09T17:20:07.794Z",
      "isPublic": true
    },
    {
      "factKey": "latest_release",
      "category": "release",
      "label": "Latest release",
      "value": "0.9.42",
      "href": "https://clawhub.ai/yijian-baidu/baidu-yijian-vision",
      "sourceUrl": "https://clawhub.ai/yijian-baidu/baidu-yijian-vision",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-05-06T03:50:51.584Z",
      "isPublic": true
    },
    {
      "factKey": "handshake_status",
      "category": "security",
      "label": "Handshake status",
      "value": "UNKNOWN",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-yijian-baidu-baidu-yijian-vision/trust",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-yijian-baidu-baidu-yijian-vision/trust",
      "sourceType": "trust",
      "confidence": "medium",
      "observedAt": null,
      "isPublic": true
    }
  ],
  "events": [
    {
      "eventType": "release",
      "title": "Release 0.9.42",
      "description": "v0.9.42",
      "href": "https://clawhub.ai/yijian-baidu/baidu-yijian-vision",
      "sourceUrl": "https://clawhub.ai/yijian-baidu/baidu-yijian-vision",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-05-06T03:50:51.584Z",
      "isPublic": true
    }
  ]
}

Record generated Oct 10, 2026.

Sponsored

Ads related to Baidu Yijian Vision and adjacent AI workflows.