PaddleOCR
面向法律 PDF 与扫描件的 PaddleOCR 结构化解析技能。默认将本地 PDF 或图片转换为 Markdown,并在技能内部保留可追溯 archive 归档。本技能应在用户需要法律 PDF OCR、卷宗 OCR、病历 OCR、证据扫描件转 Markdown、表格识别、公式识别、版面分析、PDF 转 Mark... Skill: PaddleOCR Owner: cat-xierluo Summary: 面向法律 PDF 与扫描件的 PaddleOCR 结构化解析技能。默认将本地 PDF 或图片转换为 Markdown,并在技能内部保留可追溯 archive 归档。本技能应在用户需要法律 PDF OCR、卷宗 OCR、病历 OCR、证据扫描件转 Markdown、表格识别、公式识别、版面分析、PDF 转 Mark... Tags: latest:1.1.1 Version history: v1.1.1 | 2026-04-15T10:02:54.231Z | user 面向法律 PDF 与扫描件的 PaddleOCR 结构化解析,支持表格识别、公式识别、版面分析,保留 archive 归档 Archive index: Archive v1.1.1: 13 files, 23861 bytes Files: CHANGELOG.md (2
Rank
62
Safety
84
Downloads
1.6k
Updated
Oct 10, 2026
Version
1.1.1
Source
CLAWHUB
About
What it does, and when to use it.
Capability contract not published. No trust telemetry is available yet. 1.6K downloads reported by the source. Last updated 10/10/2026.
Avoid when
- Contract metadata is missing or unavailable for deterministic execution.
Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing
Public facts
Every fact links back to the source it came from.
- Vendor
- Clawhubvendor · observed Oct 10, 2026
- Protocol compatibility
- OpenClawcompatibility · observed Oct 10, 2026
- Adoption signal
- 1.6K downloadsadoption · observed Oct 10, 2026
- Latest release
- 1.1.1release · observed Apr 15, 2026
- Handshake status
- UNKNOWNsecurity
Install and run
Setup complexity: low.
clawhub skill install s179ghx665dhgcsrdzyss2am0x83g7nf:paddle-ocr- Setup complexity is LOW. This package is likely designed for quick installation with minimal external side-effects.
- Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.
Contract: missing
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-cat-xierluo-paddle-ocr/snapshot"
Documentation
CLAWHUB
12,452 characters of source documentation, loaded on request.
Extracted files
5 files captured from the source.
SKILL.md
--- name: paddle-ocr homepage: https://github.com/cat-xierluo/legal-skills author: 杨卫薪律师(微信ywxlaw) version: "1.1.1" license: MIT description: 面向法律 PDF 与扫描件的 PaddleOCR 结构化解析技能。默认将本地 PDF 或图片转换为 Markdown,并在技能内部保留可追溯 archive 归档。本技能应在用户需要法律 PDF OCR、卷宗 OCR、病历 OCR、证据扫描件转 Markdown、表格识别、公式识别、版面分析、PDF 转 Markdown、复杂 PDF 解析时使用。 --- # PaddleOCR 法律 PDF 转 Markdown 本技能服务于**法律材料 OCR**。默认目标不是返回一段临时文本,而是: 1. 将本地 PDF / 图片转换为可继续编辑和分析的 Markdown。 2. 在 `archive/` 下保留完整归档,便于复核、追溯和二次处理。 ## 何时使用 在以下场景使用本技能: - 需要把卷宗、病历、证据材料、法院通知、财报、票据等扫描 PDF 转成 Markdown。 - 文档包含表格、印章、页眉页脚、多栏排版、公式或复杂版面。 - 希望保留一个技能内的 archive,沉淀原文件、Markdown、结构化 JSON 和批次结果。 - 后续还要继续做法律分析、证据摘录、知识入库或 RAG 切片。 在以下场景不要优先使用本技能: - 只是快速读取一小段清晰文本,且不需要 Markdown 和归档。 - 只是截图抄字,速度比结构化质量更重要。 - 输入不是 PDF / 常见图片格式。 ## 主产出 默认主产出只有两类: - **Markdown 文件**:保存在源文件同目录,默认与原文件同名、扩展名为 `.md` - **archive 归档目录**:保存在 `paddle-ocr/archive/时间戳_文件名/` archive 默认包含: - 原始输入文件副本 - 最终 `result.md` - 最终 `result.json` - 批次级 `batches/*.json` - 提取出的图片资源 - `metadata.json` ## 依赖 ### 系统依赖 | 依赖 | 安装方式 | |------|----------| | `python3` | macOS 通常已内置 | | `uv` | macOS: `brew install uv` | ### Python 包 脚本使用 `uv run` 执行,依赖写在脚本头部,无需单独维护 `requirements.txt`。 ## 首次配置 ### 获取 API 信息 1. 打开 [PaddleOCR 官网](https://www.paddleocr.com) 2. 进入对应模型的 API 页面 3. 在示例代码中复制: - `API_URL` - `Access Token` ### 配置方式 优先编辑 `config/.env`: ```bash cd paddle-ocr/config cp .env.example .env nano .env ``` 必填项: - `PADDLEOCR_DOC_PARSING_API_URL` - `PADDLEOCR_ACCESS_TOKEN` ## 常用命令 ### 主工作流:生成 Markdown + archive 在技能根目录运行: ```bash uv run scripts/convert.py "/path/to/legal-document.pdf" ``` 或继续兼容旧入口: ```bash /usr/bin/osascript -l JavaScript scripts/convert.js "/path/to/legal-document.pdf" ``` 可选参数: ```bash uv run scripts/convert.py "/path/to/legal-document.pdf" --pages "1-20" uv run scripts/convert.py "/path/to/legal-document.pdf" --output "/tmp/output.md" uv run scripts/convert.py "/path/to/legal-document.pdf" --archive-name "某案卷宗-证据一" ``` ### 底层调试:只调用解析接口,输出结构化 JSON ```bash uv run scripts/layout_caller.py --file-path "/path/to/legal-document.pdf" --pretty uv run scripts/layout_caller.py --file-url "https://example.com/document.pdf" --stdout --pretty ``` 当你只想检查原始接口结果,或后续要自己解析表格/坐标信息时,使用这个底层脚本。 ### 自检 ```bash uv run scripts/smoke_test.py --skip-api-test uv run scripts/smoke_test.py ``` ### 拆分页码 ```bash uv run scripts/split_pdf.py input.pdf output.pdf --pages "1-5,8,10-12" ``` ## 法律 PDF 工作流 按以下顺序工作: 1. 优先使用 `scripts/convert.py`。 2. 如只需部分页码,先传 `--pages`,避免整卷上传。 3. 对大体量卷宗,脚本会按配置自动分批请求,再合并为一个 Markdown。 4. 需要复核时,到 `archive/` 查看: - `output/result.md` - `output/result.json` - `metadata.json` - `batches/*.json` ## 大文件策略 本技能为了法律材料的稳定性,默认采用**保守批次策略**: - PDF 页数超过 `PADDLEOCR_BATCH_PAGES` 时自动分批 - 预估 Base64 大小超过 `PADDLEOCR_MAX_BASE64_MB` 时自动分批 这意味着它可能比官方上限更早拆分,但通常能降低长卷宗、病历合并件和扫描质量不稳定文档的失败率。 ## 输出说明 ### Markdown - 默认保存到源文件同目录 - 如果传 `--output` 且是 `.md` 文件路径,则保存到指定路径 - 如果 `--output` 是目录,则在该目录下生成同名 `.md` ### archive 默认归档目录结构: ```text archive/ └── 2026
_meta.json
{
"ownerId": "kn7bn9h1qxa9ja48qkmaxtfjgx81ksex",
"slug": "paddle-ocr",
"version": "1.1.1",
"publishedAt": 1776247374231
}references/output_schema.md
# 输出结构说明
本技能有两层输出:
1. **底层接口层**:`scripts/layout_caller.py` 输出稳定 JSON envelope
2. **高层法律工作流**:`scripts/convert.py` 生成 Markdown,并把结构化结果写入 archive
## 一、`layout_caller.py` 输出结构
`layout_caller.py` 用于直接调用 PaddleOCR 接口,返回统一包装:
```json
{
"ok": true,
"text": "从所有页面拼接出的 Markdown 文本",
"result": { "errorCode": 0, "result": { "...": "原始接口结果" } },
"error": null
}
```
失败时:
```json
{
"ok": false,
"text": "",
"result": null,
"error": {
"code": "CONFIG_ERROR | INPUT_ERROR | API_ERROR",
"message": "可直接展示给用户的错误信息"
}
}
```
重点字段:
- `text`:由 `result.result.layoutParsingResults[*].markdown.text` 拼接而成
- `result.result.layoutParsingResults[*].markdown.images`:页面内图片资源
- `result.result.layoutParsingResults[*].prunedResult`:坐标、分类、置信度等结构化版面信息
## 二、`convert.py` 的 archive 结构
`convert.py` 是高层入口,默认生成 Markdown 并写入 `archive/`。
归档目录示例:
```text
archive/
└── 20260405_153000_某案卷宗/
├── input/
│ └── 某案卷宗.pdf
├── output/
│ ├── result.md
│ ├── result.json
│ └── images/
├── batches/
│ ├── batch_001_1-40.json
│ └── batch_002_41-67.json
└── metadata.json
```
### `output/result.json`
这是高层工作流的汇总文件,包含:
- 输入文件基本信息
- 处理模式(单次 / 自动分批)
- 提取出的全文 Markdown
- 输出图片列表
- 各批次摘要
示例:
```json
{
"ok": true,
"source": {
"path": "/path/to/file.pdf",
"name": "file.pdf",
"sha256": "..."
},
"processing": {
"mode": "batched",
"batch_count": 2,
"total_pages": 67,
"processed_pages": 67,
"selected_pages": "1-67"
},
"text": "最终合并后的 Markdown",
"images": [],
"batches": [
{
"index": 1,
"label": "1-40",
"text_length": 12345,
"image_count": 2
}
]
}
```
### `batches/*.json`
每个批次对应一个底层 envelope,便于排查:
- 哪一批 OCR 异常
- 哪一批版面错乱
- 某页的 `prunedResult` 是否需要单独读取
### `metadata.json`
记录:
- 处理时间
- provider 名称
- 关键配置
- Markdown 输出路径
- 图片目录路径
## 三、建议读取顺序
如果只是要最终文本:
1. 读取 `output/result.md`
如果需要排查 OCR 质量:
1. 读取 `output/result.json`
2. 再按需读取 `batches/*.json`
如果需要提取表格、坐标、阅读顺序:
1. 直接运行 `scripts/layout_caller.py`
2. 读取 `result.result.layoutParsingResults[*].prunedResult`CHANGELOG.md
# 变更记录 ## [1.1.1] - 2026-04-05 ### 改进 - 将对外配置字段统一收敛为官方命名:`PADDLEOCR_DOC_PARSING_API_URL` 与 `PADDLEOCR_ACCESS_TOKEN`。 - `SKILL.md` 与 `.env.example` 删除旧别名说明,避免用户在配置时产生歧义。 ### 技术优化 - `scripts/lib.py` 不再从旧别名字段读取 API 地址和 Token,配置接口与官方保持一致。 ### 文档完善 - 更新配置章节,明确只按官方字段填写 `.env`。 ## [1.1.0] - 2026-04-05 ### 新增 - 新增 `scripts/lib.py`,统一配置读取、接口调用、稳定 JSON envelope 与错误包装。 - 新增 `scripts/layout_caller.py`,支持直接调试底层 JSON 结果。 - 新增 `scripts/split_pdf.py`,支持 PDF 页码提取与自动分批。 - 新增 `scripts/smoke_test.py`,支持配置检查与 API 连通性自检。 - 新增 `scripts/optimize_file.py`,支持对扫描图片做压缩优化。 - 新增 `references/output_schema.md`,说明底层 JSON envelope 与 archive 结构。 - 新增 `TASKS.md` 与 `DECISIONS.md`,补齐技能级协作文档。 - 新增 `LICENSE.txt`,统一许可证文件名与版权信息。 ### 改进 - 将技能定位收敛为“面向法律 PDF / 扫描件的 Markdown + archive 工作流”。 - 默认输出保持为 Markdown 文件,并在技能内部保留可追溯 archive。 - 为卷宗、病历、证据材料等长文档增加自动分批逻辑,优先保障稳定性。 - `convert.js` 改为兼容层,转而调用 Python 主链路,不再内置核心 OCR 逻辑。 - `SKILL.md` 重写为以法律文档场景为中心的说明文档,并补充适用/不适用场景。 ### 技术优化 - 移除旧的固定 `test/paddle-ocr` 路径依赖,改为基于脚本位置动态推导 skill 根目录。 - 用 `pypdfium2` 替代 Ghostscript 方案,降低系统依赖。 - 统一支持新旧环境变量字段,兼容已有 `.env` 配置。 - 归档目录新增 `metadata.json`、批次 JSON 和输出结构说明,增强复核与追溯能力。 ### 文档完善 - 将配置说明更新为 `PADDLEOCR_DOC_PARSING_API_URL` / `PADDLEOCR_ACCESS_TOKEN` 主字段。 - 补充大文件策略、页码范围、主入口与底层入口的分工说明。 ### 待办事项 - 增加真实法律 PDF 样本的回归测试集。 - 评估页眉页脚、印章、批注的后处理去噪规则。 ## [1.0.0] - 2026-01-15 ### 新增 - 初始版本发布。 - 支持将 PDF 和图片转换为 Markdown。 - 集成 PaddleOCR 文档解析接口。 - 支持 OCR、表格识别、公式识别与图片提取。 - 增加基础 archive 归档能力。 ### 技术优化 - 使用 JXA 与 Python 组合实现基础转换流程。 - 支持 Base64 上传与 `fileType` 自动检测。 ### 文档完善 - 提供基础配置说明、故障排除和与 MinerU 的差异说明。
skill-card.md
## Description: PaddleOCR converts legal PDFs and scanned document images into Markdown with structured OCR outputs and a traceable local archive. This skill is ready for commercial/non-commercial use. ## Publisher: [cat-xierluo](https://clawhub.ai/user/cat-xierluo) ### License/Terms of Use: MIT ## Use Case: Developers, legal operations teams, and agents use this skill to convert legal PDFs, case files, medical records, evidence scans, tables, formulas, and complex layouts into editable Markdown and archived JSON outputs for review or downstream analysis. ### Deployment Geography for Use: Global ## Known Risks and Mitigations: Risk: Sensitive legal, medical, financial, or evidence files may be uploaded to the configured OCR service. Mitigation: Use only a trusted PaddleOCR endpoint and prefer limited page ranges for sensitive documents. Risk: The skill keeps local archives by default, which may retain sensitive source files, OCR text, JSON, and extracted images. Mitigation: Use --no-archive when retention is not appropriate and review archive storage permissions before processing sensitive files. Risk: Output image directories can be replaced during conversion. Mitigation: Avoid choosing output paths where an existing *_images directory contains important files. ## Reference(s): - [PaddleOCR official site](https://www.paddleocr.com) - [Output schema](references/output_schema.md) - [ClawHub skill page](https://clawhub.ai/cat-xierluo/skills/paddle-ocr) - [Project homepage](https://github.com/cat-xierluo/legal-skills) ## Skill Output: **Output Type(s):** [Markdown, Text, JSON, Files, Shell commands, Configuration guidance] **Output Format:** [Markdown files, JSON archive files, local file paths, and shell command guidance] **Output Parameters:** [1D] **Other Properties Related to Output:** [Creates a local archive by default; optional page ranges and output paths can narrow processing.] ## Skill Version(s): 1.1.1 (source: frontmatter and changelog, released 2026-04-05) ## Ethical Considerations: Users should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.
AionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
activepieces
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
CopilotKit
The Frontend for Agents & Generative UI. React + Angular
Machine-readable data
The same record, as JSON, for agents and crawlers.
{
"facts": [
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Clawhub",
"href": "https://clawhub.ai/cat-xierluo/skills/paddle-ocr",
"sourceUrl": "https://clawhub.ai/cat-xierluo/skills/paddle-ocr",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-10T08:16:05.771Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-cat-xierluo-paddle-ocr/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-cat-xierluo-paddle-ocr/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-10T08:16:05.771Z",
"isPublic": true
},
{
"factKey": "traction",
"category": "adoption",
"label": "Adoption signal",
"value": "1.6K downloads",
"href": "https://clawhub.ai/cat-xierluo/paddle-ocr",
"sourceUrl": "https://clawhub.ai/cat-xierluo/paddle-ocr",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-10T08:16:05.771Z",
"isPublic": true
},
{
"factKey": "latest_release",
"category": "release",
"label": "Latest release",
"value": "1.1.1",
"href": "https://clawhub.ai/cat-xierluo/paddle-ocr",
"sourceUrl": "https://clawhub.ai/cat-xierluo/paddle-ocr",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-04-15T10:02:54.231Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-cat-xierluo-paddle-ocr/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-cat-xierluo-paddle-ocr/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
],
"events": [
{
"eventType": "release",
"title": "Release 1.1.1",
"description": "面向法律 PDF 与扫描件的 PaddleOCR 结构化解析,支持表格识别、公式识别、版面分析,保留 archive 归档",
"href": "https://clawhub.ai/cat-xierluo/paddle-ocr",
"sourceUrl": "https://clawhub.ai/cat-xierluo/paddle-ocr",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-04-15T10:02:54.231Z",
"isPublic": true
}
]
}Record generated Oct 10, 2026.
