WPS PDF Processing
当用户需要对 PDF 文件进行任何操作时,使用本技能。包括:读取或提取 PDF 中的文字/表格、合并多个 PDF、拆分 PDF、旋转页面、添加水印、创建新 PDF、填写 PDF 表单、加密/解密 PDF、提取图片,以及对扫描版 PDF 进行 OCR 识别使其可搜索。只要用户提到 .pdf 文件或希望生成 PDF,... Skill: WPS PDF Processing Owner: xixihaha123123123123 Summary: 当用户需要对 PDF 文件进行任何操作时,使用本技能。包括:读取或提取 PDF 中的文字/表格、合并多个 PDF、拆分 PDF、旋转页面、添加水印、创建新 PDF、填写 PDF 表单、加密/解密 PDF、提取图片,以及对扫描版 PDF 进行 OCR 识别使其可搜索。只要用户提到 .pdf 文件或希望生成 PDF,... Tags: document:1.0.0, latest:1.0.0, ocr:1.0.0, pdf:1.0.0 Version history: v1.0.0 | 2026-04-18T12:32:27.417Z | user Initial release Archive index: Archive v1.0.0: 4 files, 7849 bytes Files: scripts
Rank
62
Safety
84
Downloads
1.3k
Updated
Oct 10, 2026
Version
1.0.0
Source
CLAWHUB
About
What it does, and when to use it.
Capability contract not published. No trust telemetry is available yet. 1.3K downloads reported by the source. Last updated 10/10/2026.
Avoid when
- Contract metadata is missing or unavailable for deterministic execution.
Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing
Public facts
Every fact links back to the source it came from.
- Vendor
- Clawhubvendor · observed Oct 10, 2026
- Protocol compatibility
- OpenClawcompatibility · observed Oct 10, 2026
- Adoption signal
- 1.3K downloadsadoption · observed Oct 10, 2026
- Latest release
- 1.0.0release · observed Apr 18, 2026
- Handshake status
- UNKNOWNsecurity
Install and run
Setup complexity: low.
clawhub skill install s172gcva854bewaha4hwd8xj4s852cyf:wps-pdf- Setup complexity is LOW. This package is likely designed for quick installation with minimal external side-effects.
- Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.
Contract: missing
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-xixihaha123123123123-wps-pdf/snapshot"
Documentation
CLAWHUB
8,852 characters of source documentation, loaded on request.
Extracted files
3 files captured from the source.
SKILL.md
---
name: pdf
description: 当用户需要对 PDF 文件进行任何操作时,使用本技能。包括:读取或提取 PDF 中的文字/表格、合并多个 PDF、拆分 PDF、旋转页面、添加水印、创建新 PDF、填写 PDF 表单、加密/解密 PDF、提取图片,以及对扫描版 PDF 进行 OCR 识别使其可搜索。只要用户提到 .pdf 文件或希望生成 PDF,即使用本技能。
---
# PDF 处理指南
## 工具速查
| 任务 | 推荐库 | 说明 |
| -------------------------------- | ---------- | --------------------------------------------- |
| 合并 / 拆分 / 旋转 / 水印 / 加密 | pypdf | 轻量,纯 Python |
| 提取文本 / 表格(结构化) | pdfplumber | 精度高,支持坐标; |
| 创建排版 PDF | reportlab | 支持段落、表格、样式 |
| 扫描版 OCR / 结构化转 Markdown | pdf-to-md | 返回图片 + markdown;文字提取失败时的兜底方案 |
---
## pypdf — 基础操作
```python
from pypdf import PdfReader, PdfWriter
# 提取文本
reader = PdfReader("doc.pdf")
text = "".join(page.extract_text() for page in reader.pages)
# 合并
writer = PdfWriter()
for path in ["a.pdf", "b.pdf"]:
for page in PdfReader(path).pages:
writer.add_page(page)
with open("merged.pdf", "wb") as f:
writer.write(f)
# 拆分(每页单独保存)
for i, page in enumerate(reader.pages):
w = PdfWriter()
w.add_page(page)
with open(f"page_{i+1}.pdf", "wb") as f:
w.write(f)
# 旋转 / 水印 / 加密 / 裁剪
page.rotate(90)
page.merge_page(PdfReader("watermark.pdf").pages[0])
writer.encrypt("user_pass", "owner_pass")
page.mediabox.left, page.mediabox.bottom, page.mediabox.right, page.mediabox.top = 50, 50, 550, 750
```
---
## pdfplumber — 文本与表格提取
```python
import pdfplumber, pandas as pd
with pdfplumber.open("doc.pdf") as pdf:
# 文本
text = pdf.pages[0].extract_text()
# 表格 → DataFrame
for t in pdf.pages[0].extract_tables():
if t:
df = pd.DataFrame(t[1:], columns=t[0])
# 按坐标区域提取(左、上、右、下)
region_text = pdf.pages[0].within_bbox((100, 100, 400, 200)).extract_text()
```
---
## 注意事项
- **中文字体**:reportlab 默认字体不含中文字形,生成含中文的 PDF 时必须先通过
`pdfmetrics.registerFont(TTFont(...))` 注册系统中文字体(如 Noto Sans
CJK、微软雅黑、文泉驿等),并在样式中指定该字体,否则中文会显示为乱码。
## reportlab — 创建 PDF
```python
from reportlab.lib.pagesizes import letter
from reportlab.platypus import SimpleDocTemplate, Paragraph, Spacer
from reportlab.lib.styles import getSampleStyleSheet
doc = SimpleDocTemplate("out.pdf", pagesize=letter)
styles = getSampleStyleSheet()
doc.build([
Paragraph("标题", styles["Title"]),
Spacer(1, 12),
Paragraph("正文内容", styles["Normal"]),
])
```
> **下标/上标**:不要用 Unicode 字符(₀¹²),改用 XML 标签:`H<sub>2</sub>O`、`x<super>2</super>`。
---
## OCR 与常见问题
### OCR:PDF → Markdown(含图片)
> **如果用 `pypdf` 或 `pdfplumber` 提取到的文字为空、极少,或出现大量乱码,必须改用此 OCR 方案。** 扫描版 PDF、拍照 PDF、图片型 PDF 均无法通过普通文本提取获得内容,OCR 是唯一可靠手段。
```python
import sys, os
sys.path.insert(0, os.path.join(os.getenv('skill_path'), 'pdf', 'scripts'))
from pdf_to_md import parse
# 结果写入 <输出目录>/content.md,图片写入 <输出目录>/images/
parse('<PDF路径>', '<输出目录>')
```
> 适用场景:扫描版 PDF、图文混排、需要保留图片、中文内容居多。
### 其他常见问题
```python
# 处理加密 PDF
from pypdf import PdfReader
read_meta.json
{
"ownerId": "kn75sj7ng3atr0ybce5905vm85853fbc",
"slug": "wps-pdf",
"version": "1.0.0",
"publishedAt": 1776515547417
}skill-card.md
## Description: Helps agents process PDF files, including text and table extraction, merging, splitting, rotation, watermarking, PDF creation, form handling, encryption, image extraction, and OCR-backed PDF-to-Markdown conversion. This skill is ready for commercial/non-commercial use. ## Publisher: [xixihaha123123123123](https://clawhub.ai/user/xixihaha123123123123) ### License/Terms of Use: MIT-0 ## Use Case: Developers and agents use this skill to manipulate PDFs, extract text and tables, create new PDFs, and fall back to OCR when normal text extraction fails. It is especially relevant for scanned, photographed, image-heavy, or Chinese-language PDFs that need Markdown output with local image files. ### Deployment Geography for Use: Global ## Known Risks and Mitigations: Risk: The OCR/PDF-to-Markdown path uploads complete PDFs to WPS using ambient WPS session credentials. Mitigation: Use the OCR path only for PDFs approved for remote WPS processing, and avoid confidential, regulated, financial, legal, or proprietary documents unless that remote processing is explicitly acceptable. Risk: The WPS_API_BASE setting and image download behavior affect where network requests go and what files are written locally. Mitigation: Review and constrain WPS_API_BASE and image download behavior before use, and run conversions in an isolated output directory. ## Reference(s): - [ClawHub Skill Page](https://clawhub.ai/xixihaha123123123123/skills/wps-pdf) - [WPS API Base](https://api.wps.cn) - [WPS KDocs Web Origin](https://365.kdocs.cn) ## Skill Output: **Output Type(s):** [text, markdown, code, shell commands, configuration, guidance] **Output Format:** [Markdown guidance with Python code examples; the OCR helper writes content.md and localized image files.] **Output Parameters:** [1D] **Other Properties Related to Output:** [The OCR path may create Markdown and image files and sends complete PDFs to WPS using local WPS session credentials.] ## Skill Version(s): 1.0.0 (source: server release evidence) ## Ethical Considerations: Users should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.
AionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
activepieces
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
CopilotKit
The Frontend for Agents & Generative UI. React + Angular
Machine-readable data
The same record, as JSON, for agents and crawlers.
{
"facts": [
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Clawhub",
"href": "https://clawhub.ai/xixihaha123123123123/skills/wps-pdf",
"sourceUrl": "https://clawhub.ai/xixihaha123123123123/skills/wps-pdf",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-10T17:24:22.690Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-xixihaha123123123123-wps-pdf/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-xixihaha123123123123-wps-pdf/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-10T17:24:22.690Z",
"isPublic": true
},
{
"factKey": "traction",
"category": "adoption",
"label": "Adoption signal",
"value": "1.3K downloads",
"href": "https://clawhub.ai/xixihaha123123123123/wps-pdf",
"sourceUrl": "https://clawhub.ai/xixihaha123123123123/wps-pdf",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-10T17:24:22.690Z",
"isPublic": true
},
{
"factKey": "latest_release",
"category": "release",
"label": "Latest release",
"value": "1.0.0",
"href": "https://clawhub.ai/xixihaha123123123123/wps-pdf",
"sourceUrl": "https://clawhub.ai/xixihaha123123123123/wps-pdf",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-04-18T12:32:27.417Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-xixihaha123123123123-wps-pdf/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-xixihaha123123123123-wps-pdf/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
],
"events": [
{
"eventType": "release",
"title": "Release 1.0.0",
"description": "Initial release",
"href": "https://clawhub.ai/xixihaha123123123123/wps-pdf",
"sourceUrl": "https://clawhub.ai/xixihaha123123123123/wps-pdf",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-04-18T12:32:27.417Z",
"isPublic": true
}
]
}Record generated Oct 10, 2026.
