agentCLAWHUBUnverified

china-doc-ocr

智能文档OCR识别与结构化提取。Use when the user has a complex document, PDF, scanned image, photo, invoice, receipt, ID card, table, or chart that needs to be recognized a... Skill: china-doc-ocr Owner: tobewin Summary: 智能文档OCR识别与结构化提取。Use when the user has a complex document, PDF, scanned image, photo, invoice, receipt, ID card, table, or chart that needs to be recognized a... Tags: china:1.2.0, document:1.2.0, image:1.2.0, invoice:1.2.0, latest:1.2.0, ocr:1.2.0, paddleocr:1.2.0, pdf:1.2.0, recognition:1.2.0, siliconflow:1.2.0 Version history: v1.2.0 | 2026-04-29T01:19:35.463Z | user v1.

OpenClaw

Rank

62

Safety

84

Downloads

2.3k

Updated

Oct 9, 2026

Version

1.2.0

Source

CLAWHUB

About

What it does, and when to use it.

Capability contract not published. No trust telemetry is available yet. 2.3K downloads reported by the source. Last updated 10/9/2026.

Avoid when

  • Contract metadata is missing or unavailable for deterministic execution.

Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing

Public facts

Every fact links back to the source it came from.

Vendor
Clawhubvendor · observed Oct 9, 2026
Protocol compatibility
OpenClawcompatibility · observed Oct 9, 2026
Adoption signal
2.3K downloadsadoption · observed Oct 9, 2026
Latest release
1.2.0release · observed Apr 29, 2026
Handshake status
UNKNOWNsecurity

Install and run

Setup complexity: low.

clawhub skill install s174rz3f0862tcfw7pzfh5w8kn83hv2z:china-doc-ocr
  1. Setup complexity is LOW. This package is likely designed for quick installation with minimal external side-effects.
  2. Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.

Contract: missing

curl -s "https://www.xpersona.co/api/v1/agents/clawhub-tobewin-china-doc-ocr/snapshot"

Documentation

CLAWHUB

48,745 characters of source documentation, loaded on request.

Extracted files

5 files captured from the source.

SKILL.md

---
name: china-doc-ocr
description: 智能文档OCR识别与结构化提取。Use when the user has a complex document, PDF, scanned image, photo, invoice, receipt, ID card, table, or chart that needs to be recognized and converted to text or Markdown. Uses PaddleOCR-VL-1.5 and DeepSeek-OCR. 文档OCR、发票识别、证件识别。
version: 1.2.0
license: MIT-0
metadata: {"openclaw": {"emoji": "📄", "requires": {"bins": ["python3"], "env": ["SILICONFLOW_API_KEY"]}, "primaryEnv": "SILICONFLOW_API_KEY"}}
---

# 智能文档 OCR China Doc OCR

识别并提取复杂文档内容:PDF、图片、扫描件、发票、表格、证件等。
使用硅基流动 DeepSeek-OCR / PaddleOCR-VL,国内直连,无需翻墙。

模型选择与参数说明 → `references/models.md`
各场景提示词模板 → `references/prompts.md`

## 触发时机

- "帮我识别这个PDF/图片里的内容"
- "把这张发票/收据的信息提取出来"
- "将这份扫描合同转成可编辑文字"
- "这个表格里的数据帮我提取一下"
- "帮我把这张截图的文字识别出来"
- "这份报告转成 Markdown 格式"
- "识别这张身份证/营业执照的信息"

---

## 模型选择策略(优先OCR)

```
OCR优先级:
1. PaddleOCR-VL-1.5 (免费、快速、专业OCR)
2. DeepSeek-OCR (免费、效果好)
3. Qwen2.5-VL-72B (视觉语言模型,OCR效果一般但可补充)

默认使用 PaddleOCR-VL-1.5
如果识别效果不好,降级到 DeepSeek-OCR
如果仍然不好,降级到 Qwen2.5-VL-72B
```

---

## Step 0:环境检查

```bash
# 检查 API Key
if [ -z "$SILICONFLOW_API_KEY" ]; then
  echo "缺少 SILICONFLOW_API_KEY"
  echo "配置方法:"
  echo "  1. 访问 cloud.siliconflow.cn 注册(国内直连)"
  echo "  2. 进入「API密钥」页面创建 Key"
  echo "  3. export SILICONFLOW_API_KEY='sk-xxxxxxxx'"
  exit 1
fi
```

---

## Step 1:识别内容类型,选择处理模式

```
用户提供文件路径或 URL → 判断类型:

文件扩展名/用户描述 → 处理模式:

.pdf                    → PDF 模式
.jpg/.jpeg/.png/.webp   → 图片模式
.bmp/.tiff/.gif         → 图片模式(先转换格式)
URL(http/https开头)   → URL 直接模式
用户粘贴了 base64       → 直接使用

用户意图 → 选择 Prompt 模式:

"转成文字/提取文字"     → 通用OCR
"转成Markdown/保留格式" → 文档转Markdown
"提取表格/表格数据"     → 图表解析
"发票/收据/单据"        → 发票识别
"身份证/证件/执照"      → 证件识别
"图表/图形/柱状图"      → 图表解析
未指定                  → 默认文档转Markdown
```

---

## Step 2:图片 OCR

### 本地图片文件

```bash
python3 scripts/ocr.py \
  --image "/path/to/image.jpg" \
  --prompt "Convert the document to markdown." \
  --model paddleocr
```

### 图片 URL

```bash
python3 scripts/ocr.py \
  --url "https://example.com/document.jpg" \
  --prompt "Convert the document to markdown." \
  --model deepseek
```

### 指定模型

```bash
# 使用 PaddleOCR(默认,推荐)
python3 scripts/ocr.py --image photo.jpg --model paddleocr

# 使用 DeepSeek-OCR
python3 scripts/ocr.py --image photo.jpg --model deepseek

# 使用 Qwen2.5-VL
python3 scripts/ocr.py --image photo.jpg --model qwen
```

---

## Step 3:PDF OCR

### 单页或少页 PDF

```bash
python3 scripts/ocr.py \
  --pdf "/path/to/document.pdf" \
  --prompt "Convert the document to markdown." \
  --model deepseek
```

### 多页 PDF

多页 PDF 需要分页处理。使用 Python 脚本:

1. 使用 pypdf 分页
2. 对每页分别调用 OCR
3. 合并结果

---

## Step 4:格式化输出

识别完成后根据用户需求输出:

### 文档转 Markdown(保留结构)

```
直接输出 Markdown 内容,保留:
  - 标题层级(# ## ###)
  - 列表(- * 1.)
  - 表格(| 列1 | 列2 |)
  - 代码块(```)
  - 加粗、斜体等格式
```

### 发票/证件识别(结构化输出)

```
发票识别结果
━━━━━━━━━━━━━━━━━━━━
发票类型:增值税专用发票
发票号码:XXXXXXXXXXXXXXXX
开票日期:2026年03月21日
购买方:[公司名称]
销售方:[公司名称]
商品/服务:[明细]
不含税金额:¥X,XXX.XX
税率:13%
税额:¥XXX.XX
价税合计:¥X,XXX.XX
```

### 表格数据(CSV 友好格式)

```
识别结果同时输出:
1. Markdown 表格(可

_meta.json

{
  "ownerId": "kn75z6gevjsyrznm7dg2ez6sen82h8sz",
  "slug": "china-doc-ocr",
  "version": "1.2.0",
  "publishedAt": 1777425575463
}

references/models.md

# 模型选择说明

来源:硅基流动官方文档

## 主力模型:Pro/deepseek-ai/DeepSeek-V3

```
模型名:Pro/deepseek-ai/DeepSeek-V3
特点:
  - 支持图片和 PDF 输入(base64 或 URL)
  - 中文文档识别准确率极高
  - 支持专用 OCR prompt 格式(<image>\n<|grounding|>...)
  - 支持多图对比分析
适用:所有文档类型的首选模型
```

## 备用模型:Qwen2.5-VL-72B

```
模型名:Qwen/Qwen2.5-VL-72B-Instruct
特点:
  - 超大参数量,复杂文档理解力强
  - 支持图片输入(不支持 PDF base64)
  - 中英文双语文档效果好
适用:DeepSeek-OCR 效果不佳时的备选
```

## PaddleOCR-VL(专业 OCR 场景)

```
模型名:PaddlePaddle/PaddleOCR-VL
特点:
  - 专为 OCR 任务优化
  - 支持 CLI 和 API 两种调用方式
  - 对复杂版面(多列、表格)效果好
适用:需要精确版面还原的场景
```

## 图像输入计费(detail 参数影响)

```
detail=low:
  统一压缩为 448×448,约 256 token
  适合:文字较大、布局简单的文档

detail=high(推荐):
  按实际像素计费:ceil(h/28) × ceil(w/28) token
  适合:字体较小、布局复杂、表格密集的文档

建议默认使用 detail=high,确保识别准确率
```

## 模型选择决策

```
发票/收据/证件        → DeepSeek-V3(中文场景最佳)
普通文档/报告         → DeepSeek-V3
复杂表格/多列版面     → DeepSeek-V3 或 PaddleOCR-VL
英文为主的文档        → Qwen2.5-VL-72B
DeepSeek 结果不满意  → 改用 Qwen2.5-VL-72B 重试
```

references/prompts.md

# OCR 提示词模板

来源:硅基流动官方文档 DeepSeek-OCR 专用 prompt 格式

## 使用方式

所有 prompt 在 text 字段中传入,格式固定为:
`<image>\n<|grounding|>{具体指令}`

---

## 1. 文档转 Markdown(默认推荐)

```
<image>
<|grounding|>Convert the document to markdown.
```

输出:保留标题层级、列表、表格、加粗等所有格式
适用:报告、合同、论文、书籍页面、网页截图

---

## 2. 通用 OCR(纯文字提取)

```
<image>
<|grounding|>OCR this image.
```

输出:按阅读顺序提取所有文字,保留基本换行
适用:简单图片、截图文字、无复杂格式的文档

---

## 3. 无布局 OCR(去除所有格式)

```
<image>
Free OCR.
```

注意:这个格式不加 `<|grounding|>` 前缀
输出:纯文字流,不保留任何格式信息
适用:只需要文字内容、不关心排版的场景

---

## 4. 图表/表格解析

```
<image>
<|grounding|>Parse the figure.
```

输出:结构化描述图表内容,表格转为 Markdown 格式
适用:柱状图、折线图、饼图、数据表格、流程图

---

## 5. 图片详细描述

```
<image>
<|grounding|>Describe this image in detail.
```

输出:详细描述图片内容,包括文字、图形、布局
适用:混合图文内容,需要全面理解图片的场景

---

## 6. 文字定位

```
<image>
<|grounding|>Locate <|ref|>要定位的文字<|/ref|> in the image.
```

输出:指出特定文字在图片中的位置
适用:需要找到特定字段位置的场景

---

## 场景专用 Prompt

### 发票识别

```
请识别这张发票的所有信息,以结构化格式输出:
发票类型、发票号码、开票日期、
购买方信息(名称/税号/地址)、
销售方信息(名称/税号/地址)、
商品明细(名称/数量/单价/金额)、
税率、税额、价税合计、备注。
如有字段无法识别,标注"不清晰"。
```

### 身份证识别

```
请识别这张身份证的信息:
姓名、性别、民族、出生日期、住址、公民身份号码。
输出结构化格式,隐私字段用*号部分隐藏。
```

### 营业执照识别

```
请识别这份营业执照的信息:
公司名称、统一社会信用代码、类型、法定代表人、
注册资本、成立日期、营业期限、经营范围、注册地址。
```

### 银行流水/对账单

```
请识别这份银行流水的所有交易记录,
以表格格式输出:日期、摘要、支出、收入、余额。
如有多页请按时间顺序排列。
```

### 合同关键信息提取

```
请识别这份合同的关键信息:
合同编号、签订日期、甲方、乙方、
合同金额、付款方式、合同期限、
主要条款摘要(不超过200字)、
双方签字/盖章情况。
```

### 学术论文结构化

```
<image>
<|grounding|>Convert the document to markdown.
请保留所有标题层级、公式(用LaTeX格式)、
参考文献编号和表格结构。
```

### 表格数据提取(输出 CSV)

```
请识别这个表格中的所有数据,
以 CSV 格式输出(逗号分隔),
第一行为表头,之后每行为一条数据。
如有合并单元格,请拆分填充。
```

skill-card.md

## Description:

China Doc OCR helps agents recognize and structure content from documents, PDFs, scanned images, photos, invoices, receipts, ID cards, tables, and charts using SiliconFlow-hosted OCR models.

This skill is ready for commercial/non-commercial use.

## Publisher:

[tobewin](https://clawhub.ai/user/tobewin)

### License/Terms of Use:

MIT-0

## Use Case:

Developers, operators, and external users use this skill to convert Chinese and mixed-language document images, PDFs, invoices, IDs, tables, and charts into editable text, Markdown, or structured extraction results.

### Deployment Geography for Use:

Global

## Known Risks and Mitigations:

Risk: The skill sends documents to a third-party OCR API, which can expose sensitive personal, financial, legal, or regulated business information.

Mitigation: Use it only for documents the user is permitted to upload, and redact or avoid IDs, bank statements, contracts, invoices, and regulated records unless upload is approved.

Risk: OCR results and temporary outputs may remain in the workspace after processing.

Mitigation: Delete retained workspace outputs and temporary files when the task is finished, especially after handling private documents.

Risk: The skill requires a SiliconFlow API key for network calls.

Mitigation: Keep the API key in the environment, avoid pasting it into prompts or files, and rotate it if it may have been exposed.

## Reference(s):

- [ClawHub skill page](https://clawhub.ai/tobewin/skills/china-doc-ocr)
- [Model selection reference](references/models.md)
- [OCR prompt templates](references/prompts.md)
- [SiliconFlow API endpoint](https://api.siliconflow.cn/v1/chat/completions)
- [SiliconFlow console](https://cloud.siliconflow.cn)

## Skill Output:

**Output Type(s):** [Text, Markdown, Shell commands, Configuration, Guidance]

**Output Format:** [Markdown, structured text, and shell command snippets]

**Output Parameters:** [1D]

**Other Properties Related to Output:** [May save OCR results in the workspace; the helper supports image, PDF, or URL input and defaults to 4096 response tokens.]

## Skill Version(s):

1.2.0 (source: frontmatter, release metadata)

## Ethical Considerations:

Users should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.
Github ReposUpdated 2h agoRank 70

AionUi

Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!

MCPOPENCLAW
Github ReposUpdated 6mo agoRank 70

activepieces

AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents

OPENCLAW
Github ReposUpdated 6mo agoRank 70

cherry-studio

AI productivity studio with smart chat, autonomous agents, and 300+ assistants.

MCPOPENCLAW
Github ReposUpdated 7mo agoRank 70

CopilotKit

The Frontend for Agents & Generative UI. React + Angular

OPENCLAW

Machine-readable data

The same record, as JSON, for agents and crawlers.

{
  "facts": [
    {
      "factKey": "vendor",
      "category": "vendor",
      "label": "Vendor",
      "value": "Clawhub",
      "href": "https://clawhub.ai/tobewin/skills/china-doc-ocr",
      "sourceUrl": "https://clawhub.ai/tobewin/skills/china-doc-ocr",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-09T16:13:35.034Z",
      "isPublic": true
    },
    {
      "factKey": "protocols",
      "category": "compatibility",
      "label": "Protocol compatibility",
      "value": "OpenClaw",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-tobewin-china-doc-ocr/contract",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-tobewin-china-doc-ocr/contract",
      "sourceType": "contract",
      "confidence": "medium",
      "observedAt": "2026-10-09T16:13:35.034Z",
      "isPublic": true
    },
    {
      "factKey": "traction",
      "category": "adoption",
      "label": "Adoption signal",
      "value": "2.3K downloads",
      "href": "https://clawhub.ai/tobewin/china-doc-ocr",
      "sourceUrl": "https://clawhub.ai/tobewin/china-doc-ocr",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-09T16:13:35.034Z",
      "isPublic": true
    },
    {
      "factKey": "latest_release",
      "category": "release",
      "label": "Latest release",
      "value": "1.2.0",
      "href": "https://clawhub.ai/tobewin/china-doc-ocr",
      "sourceUrl": "https://clawhub.ai/tobewin/china-doc-ocr",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-04-29T01:19:35.463Z",
      "isPublic": true
    },
    {
      "factKey": "handshake_status",
      "category": "security",
      "label": "Handshake status",
      "value": "UNKNOWN",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-tobewin-china-doc-ocr/trust",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-tobewin-china-doc-ocr/trust",
      "sourceType": "trust",
      "confidence": "medium",
      "observedAt": null,
      "isPublic": true
    }
  ],
  "events": [
    {
      "eventType": "release",
      "title": "Release 1.2.0",
      "description": "v1.2.0: Security hardening - removed curl dependency, replaced with Python script (ocr.py). Uses urllib for API calls, no external network tools required.",
      "href": "https://clawhub.ai/tobewin/china-doc-ocr",
      "sourceUrl": "https://clawhub.ai/tobewin/china-doc-ocr",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-04-29T01:19:35.463Z",
      "isPublic": true
    }
  ]
}

Record generated Oct 9, 2026.

Sponsored

Ads related to china-doc-ocr and adjacent AI workflows.