bytedance-visual-recognition
Multimodal visual recognition via Doubao-Seed (6 models) + Zhipu GLM (2 free models). Image/video to text/code, auto-fallback, batch directory processing, follow-up conversations, local media caching (Temp/), persistent history (vision_history.json), auto IAM console sync. First-run privacy notice, plaintext config.json, cross-platform. Skill: bytedance-visual-recognition Owner: etmnb Summary: Multimodal visual recognition via Doubao-Seed (6 models) + Zhipu GLM (2 free models). Image/video to text/code, auto-fallback, batch directory processing, follow-up conversations, local media caching (Temp/), persistent history (vision_history.json), auto IAM console sync. First-run privacy notice, plaintext config.json, cross-platform. Tags: latest:5.0.1 Vers
Rank
62
Safety
84
Downloads
1.8k
Updated
Oct 10, 2026
Version
5.0.1
Source
CLAWHUB
About
What it does, and when to use it.
Capability contract not published. No trust telemetry is available yet. 1.8K downloads reported by the source. Last updated 10/10/2026.
Avoid when
- Contract metadata is missing or unavailable for deterministic execution.
Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing
Public facts
Every fact links back to the source it came from.
- Vendor
- Clawhubvendor · observed Oct 10, 2026
- Protocol compatibility
- OpenClawcompatibility · observed Oct 10, 2026
- Adoption signal
- 1.8K downloadsadoption · observed Oct 10, 2026
- Latest release
- 5.0.1release · observed Aug 10, 2026
- Handshake status
- UNKNOWNsecurity
Install and run
Setup complexity: low.
clawhub skill install s17d5t69ggjwkkgy7xx496krnx83pg6h:bytedance-visual-recognition- Setup complexity is classified as HIGH. You must provision dedicated cloud infrastructure or an isolated VM. Do not run this directly on your local workstation.
- Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.
Contract: missing
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-etmnb-bytedance-visual-recognition/snapshot"
Documentation
CLAWHUB
87,356 characters of source documentation, loaded on request.
Extracted files
4 files captured from the source.
SKILL.md
---
name: bytedance-visual-recognition
description: >
Multimodal visual recognition via Doubao-Seed (6 models) + Zhipu GLM (2 free models).
Image/video to text/code, auto-fallback, batch directory processing, follow-up conversations,
local media caching (Temp/), persistent history (vision_history.json), auto IAM console sync.
First-run privacy notice, plaintext config.json, cross-platform.
summary: "Doubao-Seed + GLM visual recognition — image/video to text/code with auto-fallback, batch, follow-up, IAM sync"
tags:
vision: "5.0.1"
image-recognition: "5.0.1"
video-recognition: "5.0.1"
image-to-code: "5.0.1"
video-to-code: "5.0.1"
doubao: "5.0.1"
glm: "5.0.1"
trigger_patterns:
- "豆包识别"
- "豆包视觉识别"
- "bytedance visual recognition"
- "doubao recognize"
metadata:
openclaw:
requires:
bins:
- python
permissions:
filesystem:
read: ["config.json", "vision_history.json", ".last_response"]
write: ["Temp/", "vision_history.json", ".last_response", "config.json"]
network:
- host: "ark.cn-beijing.volces.com"
purpose: "Doubao vision model API"
- host: "open.bigmodel.cn"
purpose: "GLM vision model Chat Completions API"
- host: "open.volcengineapi.com"
purpose: "Auto IAM console usage sync"
emoji: "🔍"
homepage: https://www.volcengine.com/docs/82379/1569618
locales: ["en"]
---
# ByteDance Visual Recognition — Doubao-Seed + GLM
Doubao-Seed (6 models) + Zhipu GLM (2 free models). First run auto-generates `config.json`, fill in one API Key to start. IAM console usage syncs automatically on each recognition.
## Privacy & Data Notice
- **Network**: Selected images/videos and prompts are base64-encoded and sent to Volcengine (Doubao) or Zhipu (GLM) cloud APIs.
- **Local cache**: Media files temporarily copied to `Temp/YYYYMMDD/`, default 7-day retention (`temp_retention_days` in config.json, range 1-3650).
- **History**: Recognition history stored in `vision_history.json`, follow-up context in `.last_response`, auto-cleaned after 7 days. GLM follow-up reuses full message history including base64 media data; Doubao follow-up uses previous_response_id without re-transmitting files.
- **Credentials**: API Keys stored in plaintext `config.json`. Do not use personal keys on shared machines.
- **IAM sync**: Automatically syncs token usage from Volcengine IAM console on each recognition when IAM credentials are configured.
- **First run**: A privacy notice is displayed once. Continuing past it acknowledges data handling practices.
## Setup (pick one)
### Doubao (Volcengine)
1. Join the [Collaboration Rewards Program](https://console.volcengine.com/ark/region:cn-beijing/openManagement/rewardPlan) for free quota, then get your API Key
2. Create inference endpoints, pick from:
| config key | model | priority |
|--------|------|:---:|
| `doubao_seed_21p_id` | Doubao-Seed-2.1-Pro | primary |
| `doubao_seed_21t_id` | Doubao-Seed-_meta.json
{
"ownerId": "kn75s192k98dtwr83bft4myd0582ytyg",
"slug": "bytedance-visual-recognition",
"version": "5.0.1",
"publishedAt": 1786388451204
}proposal/PROPOSAL.md
# 3.1.1 更新提案 ## 目标 精简发布,修复安全审查问题,仅上传核心文件。 ## 变更 - 收紧触发词列表,移除过于宽泛的 pattern - 批量处理改为仅复制媒体文件(不再复制整个目录树) - 剔除非核心文件(.env, Temp, .bak, .json 等) - 仅上传核心文件:doubao_vision_recognize.py, SKILL.md, skill-card.md
skill-card.md
## Description: Multimodal visual recognition via Doubao-Seed and Zhipu GLM for image or video analysis, image or video to code, batch processing, follow-up questions, provider fallback, local media caching, persistent history, and optional IAM usage sync. This skill is ready for commercial/non-commercial use. ## Publisher: [etmnb](https://clawhub.ai/user/etmnb) ### License/Terms of Use: MIT-0 ## Use Case: Developers and external users can use this skill to send selected images or videos to configured Doubao or GLM vision models and receive scene descriptions, extracted text, structured Markdown analysis, or generated code. It also supports batch directory processing and follow-up questions over prior recognition context. ### Deployment Geography for Use: Global ## Known Risks and Mitigations: Risk: Selected images, videos, screenshots, documents, and prompts may contain private or sensitive information and are sent to third-party cloud vision APIs. Mitigation: Review media before use, avoid sending regulated or confidential content unless approved, and configure only providers whose data handling terms are acceptable for the deployment. Risk: API keys and optional IAM credentials are stored in plaintext config.json. Mitigation: Use least-privileged provider keys, avoid shared or synced folders, rotate exposed keys, and remove unused credentials from config.json. Risk: Local files such as Temp/, vision_history.json, and .last_response can retain media context or model responses after recognition. Mitigation: Set the shortest practical retention period, manually delete Temp/, vision_history.json, and .last_response after sensitive sessions, and do not rely on automatic cleanup for GLM follow-up context. ## Reference(s): - [ClawHub skill page](https://clawhub.ai/etmnb/skills/bytedance-visual-recognition) - [Volcengine Doubao visual understanding documentation](https://www.volcengine.com/docs/82379/1569618) - [Zhipu Open Platform](https://open.bigmodel.cn) ## Skill Output: **Output Type(s):** [Text, Markdown, Code, Shell commands, Configuration, Guidance] **Output Format:** [CLI text and Markdown-style analysis, with generated code when code output is requested.] **Output Parameters:** [1D] **Other Properties Related to Output:** [Supports image files up to 15 MB, video files up to 50 MB, batch directory processing, follow-up context, local cache files, and persistent history.] ## Skill Version(s): 5.0.1 (source: server release metadata; artifact script and skill tags agree) ## Ethical Considerations: Users should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.
AionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
activepieces
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
CopilotKit
The Frontend for Agents & Generative UI. React + Angular
Machine-readable data
The same record, as JSON, for agents and crawlers.
{
"facts": [
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Clawhub",
"href": "https://clawhub.ai/etmnb/skills/bytedance-visual-recognition",
"sourceUrl": "https://clawhub.ai/etmnb/skills/bytedance-visual-recognition",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-10T01:38:38.452Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-etmnb-bytedance-visual-recognition/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-etmnb-bytedance-visual-recognition/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-10T01:38:38.452Z",
"isPublic": true
},
{
"factKey": "traction",
"category": "adoption",
"label": "Adoption signal",
"value": "1.8K downloads",
"href": "https://clawhub.ai/etmnb/bytedance-visual-recognition",
"sourceUrl": "https://clawhub.ai/etmnb/bytedance-visual-recognition",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-10T01:38:38.452Z",
"isPublic": true
},
{
"factKey": "latest_release",
"category": "release",
"label": "Latest release",
"value": "5.0.1",
"href": "https://clawhub.ai/etmnb/bytedance-visual-recognition",
"sourceUrl": "https://clawhub.ai/etmnb/bytedance-visual-recognition",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-08-10T19:00:51.204Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-etmnb-bytedance-visual-recognition/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-etmnb-bytedance-visual-recognition/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
],
"events": [
{
"eventType": "release",
"title": "Release 5.0.1",
"description": "v5 major release",
"href": "https://clawhub.ai/etmnb/bytedance-visual-recognition",
"sourceUrl": "https://clawhub.ai/etmnb/bytedance-visual-recognition",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-08-10T19:00:51.204Z",
"isPublic": true
}
]
}Record generated Oct 10, 2026.
