{"id":"9436d866-6197-4284-a0d8-461b4a68e5ab","entityType":"agent","slug":"clawhub-pengsc1994-free-pdf-processor","name":"pdf-processor","canonicalUrl":"https://www.xpersona.co/agent/clawhub-pengsc1994-free-pdf-processor","canonicalPath":"/agent/clawhub-pengsc1994-free-pdf-processor","generatedAt":"2026-10-10T14:44:19.586Z","source":"CLAWHUB","claimStatus":"UNCLAIMED","verificationTier":"NONE","summary":{"evidence":{"source":"editorial-content","verified":true,"confidence":"high","updatedAt":"2026-10-10T12:23:24.209Z","emptyReason":null},"description":"一站式 PDF 处理技能。支持 PDF 文本/图片/表格提取、格式转换（PDF↔Word/Excel）、合并拆分、OCR 识别、批量处理、水印添加、加密解密、压缩等。使用场景： (1) 从 PDF 提取文本内容进行数据分析 (2) 将 PDF 转换为 Word/Excel 方便编辑 (3) 合并或拆分 PDF 文... Skill: pdf-processor Owner: pengsc1994 Summary: 一站式 PDF 处理技能。支持 PDF 文本/图片/表格提取、格式转换（PDF↔Word/Excel）、合并拆分、OCR 识别、批量处理、水印添加、加密解密、压缩等。使用场景： (1) 从 PDF 提取文本内容进行数据分析 (2) 将 PDF 转换为 Word/Excel 方便编辑 (3) 合并或拆分 PDF 文... Tags: latest:1.0.0 Version history: v1.0.0 | 2026-04-27T06:19:43.055Z | user Initial public release: 全面升级为多功能一站式 PDF 处理工具 - 新增 14 个独立 PDF 脚本，覆盖提取文本/图片/表格、OCR、格式转换（PDF↔Word/Excel）、合并拆分、水印、加密解密、压缩、批量处理等功能 - 支持命令行一","descriptionLabel":"Technical summary","evidenceSummary":"Capability contract not published. No trust telemetry is available yet. 1.4K downloads reported by the source. Last updated 10/10/2026.","installCommand":"clawhub skill install s17cvwkcqc8b7rw95a9mh2qq0s85mrb2:free-pdf-processor","sourceUrl":"https://clawhub.ai/pengsc1994/free-pdf-processor","homepage":"https://clawhub.ai/pengsc1994/skills/free-pdf-processor","primaryLinks":[{"label":"View on ClawHub","url":"https://clawhub.ai/pengsc1994/free-pdf-processor","kind":"source"},{"label":"Homepage","url":"https://clawhub.ai/pengsc1994/skills/free-pdf-processor","kind":"homepage"}],"safetyScore":84,"overallRank":62,"popularityScore":63,"trustScore":null,"claimedByName":null,"isOwner":false,"seoDescription":"一站式 PDF 处理技能。支持 PDF 文本/图片/表格提取、格式转换（PDF↔Word/Excel）、合并拆分、OCR 识别、批量处理、水印添加、加密解密、压缩等。使用场景： (1) 从 PDF 提取文本内容进行数据分析 (2) 将 PDF 转换为 Word/Excel 方便编辑 (3) 合并或拆分 PDF 文..."},"coverage":{"evidence":{"source":"public-profile","verified":false,"confidence":"medium","updatedAt":"2026-10-10T12:23:24.209Z","emptyReason":null},"protocols":[{"protocol":"OPENCLEW","label":"OpenClaw","status":"self-declared","notes":"Declared in the public agent profile."}],"capabilities":[],"verifiedCount":0,"selfDeclaredCount":1,"capabilityMatrix":{"rows":[{"key":"OPENCLEW","type":"protocol","support":"unknown","confidenceSource":"profile","notes":"Listed on profile"}],"flattenedTokens":"protocol:OPENCLEW|unknown|profile"}},"adoption":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T12:23:24.209Z","emptyReason":null},"stars":null,"forks":null,"downloads":1435,"packageName":null,"latestVersion":"1.0.0","tractionLabel":"1.4K downloads"},"release":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T12:23:24.209Z","emptyReason":null},"lastUpdatedAt":"2026-10-10T12:23:24.209Z","lastCrawledAt":"2026-10-10T12:23:24.209Z","lastIndexedAt":null,"nextCrawlAt":"2026-10-11T12:23:24.209Z","lastVerifiedAt":null,"highlights":[{"version":"1.0.0","createdAt":"2026-04-27T06:19:43.055Z","changelog":"Initial public release: 全面升级为多功能一站式 PDF 处理工具 - 新增 14 个独立 PDF 脚本，覆盖提取文本/图片/表格、OCR、格式转换（PDF↔Word/Excel）、合并拆分、水印、加密解密、压缩、批量处理等功能 - 支持命令行一键处理多种常见 PDF 场景（提取内容、批量加水印、加解密、格式转换等） - 移除原学术专用流程及相关文档，聚焦普适 PDF 工具化处理 - 提升模块化与扩展性，每项功能独立脚本实现，方便按需调用与集成 - 全面更新文档，新增核心功能速查表与详细使用示例","fileCount":17,"zipByteSize":19277}]},"execution":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No published capability contract is available yet."},"installCommand":"clawhub skill install s17cvwkcqc8b7rw95a9mh2qq0s85mrb2:free-pdf-processor","setupComplexity":"medium","setupSteps":["Python environment detected. Create a strict virtual environment (`python -m venv .venv`) before installing dependencies to prevent system-level package conflicts.","Setup complexity is LOW. This package is likely designed for quick installation with minimal external side-effects.","Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data."],"contract":{"contractStatus":"missing","authModes":[],"requires":[],"forbidden":[],"supportsMcp":false,"supportsA2a":false,"supportsStreaming":false,"inputSchemaRef":null,"outputSchemaRef":null,"dataRegion":null,"contractUpdatedAt":null,"sourceUpdatedAt":null,"freshnessSeconds":null},"invocationGuide":{"preferredApi":{"snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-pengsc1994-free-pdf-processor/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-pengsc1994-free-pdf-processor/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-pengsc1994-free-pdf-processor/trust"},"curlExamples":["curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-pengsc1994-free-pdf-processor/snapshot\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-pengsc1994-free-pdf-processor/contract\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-pengsc1994-free-pdf-processor/trust\""],"jsonRequestTemplate":{"query":"summarize this repo","constraints":{"maxLatencyMs":2000,"protocolPreference":["OPENCLEW"]}},"jsonResponseTemplate":{"ok":true,"result":{"summary":"...","confidence":0.9},"meta":{"source":"CLAWHUB","generatedAt":"2026-10-10T14:44:19.585Z"}},"retryPolicy":{"maxAttempts":3,"backoffMs":[500,1500,3500],"retryableConditions":["HTTP_429","HTTP_503","NETWORK_TIMEOUT"]}},"endpoints":{"dossierUrl":"https://www.xpersona.co/api/v1/agents/clawhub-pengsc1994-free-pdf-processor/dossier","snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-pengsc1994-free-pdf-processor/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-pengsc1994-free-pdf-processor/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-pengsc1994-free-pdf-processor/trust"}},"reliability":{"evidence":{"source":"runtime-metrics","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No trust, reliability, or runtime telemetry is available."},"trust":{"status":"unavailable","handshakeStatus":"UNKNOWN","verificationFreshnessHours":null,"reputationScore":null,"p95LatencyMs":null,"successRate30d":null,"fallbackRate":null,"attempts30d":null,"trustUpdatedAt":null,"trustConfidence":"unknown","sourceUpdatedAt":null,"freshnessSeconds":null},"decisionGuardrails":{"doNotUseIf":["Contract metadata is missing or unavailable for deterministic execution."],"safeUseWhen":[],"riskFlags":["missing_or_unavailable_contract","trust_data_unavailable","schema_references_missing"],"operationalConfidence":"low"},"executionMetrics":{"observedLatencyMsP50":null,"observedLatencyMsP95":null,"estimatedCostUsd":null,"uptime30d":null,"rateLimitRpm":null,"rateLimitBurst":null,"lastVerifiedAt":null,"verificationSource":null},"runtimeMetrics":{"successRate":null,"avgLatencyMs":null,"avgCostUsd":null,"hallucinationRate":null,"retryRate":null,"disputeRate":null,"p50Latency":null,"p95Latency":null,"lastUpdated":null}},"benchmarks":{"evidence":{"source":"no-benchmark-data","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No benchmark suites or observed failure patterns are available."},"suites":[],"failurePatterns":[]},"artifacts":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"high","updatedAt":"2026-10-10T12:23:24.209Z","emptyReason":null},"readme":"Skill: pdf-processor\n\nOwner: pengsc1994\n\nSummary: 一站式 PDF 处理技能。支持 PDF 文本/图片/表格提取、格式转换（PDF↔Word/Excel）、合并拆分、OCR 识别、批量处理、水印添加、加密解密、压缩等。使用场景： (1) 从 PDF 提取文本内容进行数据分析 (2) 将 PDF 转换为 Word/Excel 方便编辑 (3) 合并或拆分 PDF 文...\n\nTags: latest:1.0.0\n\nVersion history:\n\nv1.0.0 | 2026-04-27T06:19:43.055Z | user\n\nInitial public release: 全面升级为多功能一站式 PDF 处理工具\n\n- 新增 14 个独立 PDF 脚本，覆盖提取文本/图片/表格、OCR、格式转换（PDF↔Word/Excel）、合并拆分、水印、加密解密、压缩、批量处理等功能\n- 支持命令行一键处理多种常见 PDF 场景（提取内容、批量加水印、加解密、格式转换等）\n- 移除原学术专用流程及相关文档，聚焦普适 PDF 工具化处理\n- 提升模块化与扩展性，每项功能独立脚本实现，方便按需调用与集成\n- 全面更新文档，新增核心功能速查表与详细使用示例\n\nArchive index:\n\nArchive v1.0.0: 17 files, 19277 bytes\n\nFiles: requirements.txt (314b), scripts/add_watermark.py (2438b), scripts/batch_process.py (4528b), scripts/compress_pdf.py (1790b), scripts/decrypt_pdf.py (1317b), scripts/encrypt_pdf.py (1198b), scripts/extract_images.py (2466b), scripts/extract_tables.py (2334b), scripts/extract_text.py (2546b), scripts/merge_pdfs.py (1492b), scripts/ocr_pdf.py (3411b), scripts/pdf_to_excel.py (2831b), scripts/pdf_to_word.py (2684b), scripts/split_pdf.py (2867b), skill-card.md (2542b), SKILL.md (4182b), _meta.json (137b)\n\nFile v1.0.0:SKILL.md\n\n---\nname: pdf-processor\ndescription: |\n  一站式 PDF 处理技能。支持 PDF 文本/图片/表格提取、格式转换（PDF↔Word/Excel）、合并拆分、OCR 识别、批量处理、水印添加、加密解密、压缩等。使用场景：\n  (1) 从 PDF 提取文本内容进行数据分析\n  (2) 将 PDF 转换为 Word/Excel 方便编辑\n  (3) 合并或拆分 PDF 文件\n  (4) 对扫描件进行 OCR 识别提取文字\n  (5) 批量处理多个 PDF 文件\n  (6) 添加水印或加密保护 PDF\n  (7) 压缩 PDF 减小文件体积\n---\n\n# PDF 处理技能\n\n## 快速开始\n\n### 安装依赖\n```bash\ncd D:\\PDF.skill\\pdf-processor\npip install -r requirements.txt\n```\n\n### 核心功能\n\n| 功能 | 命令 | 说明 |\n|------|------|------|\n| 提取文本 | `python scripts/extract_text.py <pdf_path>` | 提取 PDF 文本内容 |\n| 提取图片 | `python scripts/extract_images.py <pdf_path> <output_dir>` | 提取 PDF 中的图片 |\n| 提取表格 | `python scripts/extract_tables.py <pdf_path>` | 提取 PDF 中的表格 |\n| PDF 转 Word | `python scripts/pdf_to_word.py <pdf_path> <output_path>` | 转换为可编辑 Word |\n| PDF 转 Excel | `python scripts/pdf_to_excel.py <pdf_path> <output_path>` | 提取表格到 Excel |\n| 合并 PDF | `python scripts/merge_pdfs.py <output_path> <file1> <file2> ...` | 合并多个 PDF |\n| 拆分 PDF | `python scripts/split_pdf.py <pdf_path> <output_dir>` | 按页拆分 PDF |\n| 添加水印 | `python scripts/add_watermark.py <pdf_path> <output_path> <text>` | 添加文字水印 |\n| OCR 识别 | `python scripts/ocr_pdf.py <pdf_path> <output_path>` | OCR 识别扫描件 |\n| 加密 PDF | `python scripts/encrypt_pdf.py <input> <output> <password>` | AES-256 加密 |\n| 解密 PDF | `python scripts/decrypt_pdf.py <input> <output> <password>` | 解密 PDF |\n| 压缩 PDF | `python scripts/compress_pdf.py <input> <output>` | 压缩 PDF 文件 |\n| 批量处理 | `python scripts/batch_process.py <input_dir> <output_dir> --operation <op>` | 批量处理 |\n\n## 功能详情\n\n### extract_text.py\n提取 PDF 文本内容，支持：\n- 纯文本提取\n- 保留段落结构\n- 提取元数据（标题、作者、创建时间）\n```bash\npython scripts/extract_text.py input.pdf -o output.txt --metadata\n```\n\n### extract_tables.py\n提取 PDF 表格数据：\n- 自动检测表格边框\n- 支持合并单元格\n- 输出为 Excel 文件\n\n### pdf_to_word.py\nPDF 转 Word 转换：\n- 保留原始格式\n- 提取图片到 Word\n- 表格转换为 Word 表格\n\n### pdf_to_excel.py\nPDF 转 Excel：\n- 提取表格到不同 Sheet\n- 保留文本内容\n\n### add_watermark.py\n水印功能：\n- 支持文字水印\n- 可设置透明度、旋转角度、字体大小\n- 支持批量添加\n\n### ocr_pdf.py\nOCR 识别（需要安装 Tesseract）：\n- 使用 Tesseract 进行中文识别\n- 支持多种语言混合识别\n- 保留原有 PDF 格式\n\n### encrypt_pdf.py / decrypt_pdf.py\n加密解密：\n- AES-256 加密\n- 支持用户密码和所有者密码\n\n### compress_pdf.py\n压缩功能：\n- 清理未使用对象\n- 压缩图片\n- 5 个压缩级别可选\n\n### batch_process.py\n批量处理：\n- 支持所有单文件操作\n- 自动处理目录中所有 PDF\n- 生成处理报告\n\n## 使用示例\n\n### 从 PDF 提取文本\n```\n用户: 帮我提取这个合同的文本内容\nAI: 使用 extract_text.py 脚本提取文本\n```\n\n### PDF 转 Word\n```\n用户: 把这个 PDF 转成 Word 文档\nAI: 使用 pdf_to_word.py 进行转换\n```\n\n### 批量加水印\n```\n用户: 给这个文件夹里所有 PDF 添加\"内部资料\"水印\nAI: 使用 batch_process.py 批量处理\n```\n\n### 加密 PDF\n```\n用户: 这个文件需要加密\nAI: 使用 encrypt_pdf.py 进行 AES-256 加密\n```\n\n## 依赖安装\n\n### 基础依赖\n```bash\npip install pymupdf pdfplumber python-docx openpyxl pillow\n```\n\n### OCR 支持（可选）\n```bash\n# 安装 Tesseract OCR\n# Windows: https://github.com/UB-Mannheim/tesseract/wiki\n# macOS: brew install tesseract\n# Linux: sudo apt install tesseract-ocr\n\npip install pytesseract\n```\n\n## 注意事项\n\n- 加密 PDF 需要提供密码\n- OCR 需要安装 Tesseract 引擎\n- 大文件处理可能需要较长时间\n- 转换效果取决于 PDF 原始质量\n\nFile v1.0.0:_meta.json\n\n{\n  \"ownerId\": \"kn7awq8t980mxmrta23czncynd85m0z7\",\n  \"slug\": \"free-pdf-processor\",\n  \"version\": \"1.0.0\",\n  \"publishedAt\": 1777270783055\n}\n\nFile v1.0.0:skill-card.md\n\n## Description:\n\nProvides command-line PDF utilities for text, image, and table extraction; PDF-to-Word or Excel conversion; OCR; merging and splitting; watermarking; encryption and decryption; compression; and batch processing.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[pengsc1994](https://clawhub.ai/user/pengsc1994)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nExternal users and developers use this skill to process local PDF files by extracting content, converting formats, applying watermarks, encrypting or decrypting files, compressing documents, and running supported batch operations.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: The release evidence reports that PDF passwords can be exposed by command-line usage and by encrypt_pdf.py printing the password.\n\nMitigation: Avoid passing real passwords on the command line, remove password printing before sensitive use, and review generated logs or console output.\n\nRisk: The release evidence reports that generated XLSX files from untrusted PDFs may be unsafe active content.\n\nMitigation: Treat spreadsheets produced from untrusted PDFs as untrusted files and open them only in a controlled environment.\n\nRisk: Batch processing can create or overwrite many output files.\n\nMitigation: Use separate output directories, keep backups of source PDFs, and review the operation before running it over large directories.\n\nRisk: The release evidence recommends extra review before installing dependencies for sensitive PDF handling.\n\nMitigation: Install dependencies in an isolated environment and pin reviewed versions before processing sensitive documents.\n\n## Reference(s):\n\n- [Tesseract OCR Windows installer reference](https://github.com/UB-Mannheim/tesseract/wiki)\n\n## Skill Output:\n\n**Output Type(s):** [text, code, shell commands, configuration, guidance]\n\n**Output Format:** [Markdown guidance with command examples and local file outputs such as TXT, PDF, DOCX, XLSX, image files, and JSON indexes.]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [May create, transform, overwrite, or batch-generate local files depending on the selected PDF operation.]\n\n## Skill Version(s):\n\n1.0.0 (source: server release evidence)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nFile v1.0.0:requirements.txt\n\n# PDF 处理技能依赖\n# 安装: pip install -r requirements.txt\n\n# 核心 PDF 处理\npymupdf>=1.23.0\npdfplumber>=0.10.0\n\n# Word/Excel 转换\npython-docx>=1.1.0\nopenpyxl>=3.1.0\n\n# 图片处理\nPillow>=10.0.0\n\n# 可选: OCR 支持\n# pytesseract>=0.3.10\n# tessdata>=4.1.0\n\n# 测试\npytest>=7.4.0\npytest-cov>=4.1.0","readmeExcerpt":"Skill: pdf-processor Owner: pengsc1994 Summary: 一站式 PDF 处理技能。支持 PDF 文本/图片/表格提取、格式转换（PDF↔Word/Excel）、合并拆分、OCR 识别、批量处理、水印添加、加密解密、压缩等。使用场景： (1) 从 PDF 提取文本内容进行数据分析 (2) 将 PDF 转换为 Word/Excel 方便编辑 (3) 合并或拆分 PDF 文... Tags: latest:1.0.0 Version history: v1.0.0 | 2026-04-27T06:19:43.055Z | user Initial public release: 全面升级为多功能一站式 PDF 处理工具 - 新增 14 个独立 PDF 脚本，覆盖提取文本/图片/表格、OCR、格式转换（PDF↔Word/Excel）、合并拆分、水印、加密解密、压缩、批量处理等功能 - 支持命令行一","codeSnippets":[],"executableExamples":[{"language":"bash","snippet":"cd D:\\PDF.skill\\pdf-processor\npip install -r requirements.txt"},{"language":"bash","snippet":"python scripts/extract_text.py input.pdf -o output.txt --metadata"},{"language":"text","snippet":"用户: 帮我提取这个合同的文本内容\nAI: 使用 extract_text.py 脚本提取文本"},{"language":"text","snippet":"用户: 把这个 PDF 转成 Word 文档\nAI: 使用 pdf_to_word.py 进行转换"},{"language":"text","snippet":"用户: 给这个文件夹里所有 PDF 添加\"内部资料\"水印\nAI: 使用 batch_process.py 批量处理"},{"language":"text","snippet":"用户: 这个文件需要加密\nAI: 使用 encrypt_pdf.py 进行 AES-256 加密"}],"parameters":null,"dependencies":[],"permissions":[],"extractedFiles":[{"path":"SKILL.md","content":"---\nname: pdf-processor\ndescription: |\n  一站式 PDF 处理技能。支持 PDF 文本/图片/表格提取、格式转换（PDF↔Word/Excel）、合并拆分、OCR 识别、批量处理、水印添加、加密解密、压缩等。使用场景：\n  (1) 从 PDF 提取文本内容进行数据分析\n  (2) 将 PDF 转换为 Word/Excel 方便编辑\n  (3) 合并或拆分 PDF 文件\n  (4) 对扫描件进行 OCR 识别提取文字\n  (5) 批量处理多个 PDF 文件\n  (6) 添加水印或加密保护 PDF\n  (7) 压缩 PDF 减小文件体积\n---\n\n# PDF 处理技能\n\n## 快速开始\n\n### 安装依赖\n```bash\ncd D:\\PDF.skill\\pdf-processor\npip install -r requirements.txt\n```\n\n### 核心功能\n\n| 功能 | 命令 | 说明 |\n|------|------|------|\n| 提取文本 | `python scripts/extract_text.py <pdf_path>` | 提取 PDF 文本内容 |\n| 提取图片 | `python scripts/extract_images.py <pdf_path> <output_dir>` | 提取 PDF 中的图片 |\n| 提取表格 | `python scripts/extract_tables.py <pdf_path>` | 提取 PDF 中的表格 |\n| PDF 转 Word | `python scripts/pdf_to_word.py <pdf_path> <output_path>` | 转换为可编辑 Word |\n| PDF 转 Excel | `python scripts/pdf_to_excel.py <pdf_path> <output_path>` | 提取表格到 Excel |\n| 合并 PDF | `python scripts/merge_pdfs.py <output_path> <file1> <file2> ...` | 合并多个 PDF |\n| 拆分 PDF | `python scripts/split_pdf.py <pdf_path> <output_dir>` | 按页拆分 PDF |\n| 添加水印 | `python scripts/add_watermark.py <pdf_path> <output_path> <text>` | 添加文字水印 |\n| OCR 识别 | `python scripts/ocr_pdf.py <pdf_path> <output_path>` | OCR 识别扫描件 |\n| 加密 PDF | `python scripts/encrypt_pdf.py <input> <output> <password>` | AES-256 加密 |\n| 解密 PDF | `python scripts/decrypt_pdf.py <input> <output> <password>` | 解密 PDF |\n| 压缩 PDF | `python scripts/compress_pdf.py <input> <output>` | 压缩 PDF 文件 |\n| 批量处理 | `python scripts/batch_process.py <input_dir> <output_dir> --operation <op>` | 批量处理 |\n\n## 功能详情\n\n### extract_text.py\n提取 PDF 文本内容，支持：\n- 纯文本提取\n- 保留段落结构\n- 提取元数据（标题、作者、创建时间）\n```bash\npython scripts/extract_text.py input.pdf -o output.txt --metadata\n```\n\n### extract_tables.py\n提取 PDF 表格数据：\n- 自动检测表格边框\n- 支持合并单元格\n- 输出为 Excel 文件\n\n### pdf_to_word.py\nPDF 转 Word 转换：\n- 保留原始格式\n- 提取图片到 Word\n- 表格转换为 Word 表格\n\n### pdf_to_excel.py\nPDF 转 Excel：\n- 提取表格到不同 Sheet\n- 保留文本内容\n\n### add_watermark.py\n水印功能：\n- 支持文字水印\n- 可设置透明度、旋转角度、字体大小\n- 支持批量添加\n\n### ocr_pdf.py\nOCR 识别（需要安装 Tesseract）：\n- 使用 Tesseract 进行中文识别\n- 支持多种语言混合识别\n- 保留原有 PDF 格式\n\n### encrypt_pdf.py / decrypt_pdf.py\n加密解密：\n- AES-256 加密\n- 支持用户密码和所有者密码\n\n### compress_pdf.py\n压缩功能：\n- 清理未使用对象\n- 压缩图片\n- 5 个压缩级别可选\n\n### batch_process.py\n批量处理：\n- 支持所有单文件操作\n- 自动处理目录中所有 PDF\n- 生成处理报告\n\n## 使用示例\n\n### 从 PDF 提取文本\n```\n用户: 帮我提取这个合同的文本内容\nAI: 使用 extract_text.py 脚本提取文本\n```\n\n### PDF 转 Word\n```\n用户: 把这个 PDF 转成 Word 文档\nAI: 使用 pdf_to_word.py 进行转换\n```\n\n### 批量加水印\n```\n用户: 给这个文件夹里所有 PDF 添加\"内部资料\"水印\nAI: 使用 batch_process.py 批量处理\n```\n\n### 加密 PDF\n```\n用户: 这个文件需要加密\nAI: 使用 encrypt_pdf.py 进行 AES-256 加密\n```\n\n## 依赖安装\n\n### 基础依赖\n```bash\npip install pymupdf pdfplumber python-docx openpyxl pillow\n```\n\n### OCR 支持（可选）\n```bash\n# 安装 Tesseract OCR\n# Windows: https://github.com/UB-Mannheim/tesseract/wiki\n# macOS: brew install tesseract\n# Linux: sudo apt install tesseract-ocr\n\npip install pytesseract\n```\n\n## 注意事项\n\n- 加密 PDF 需要提供密码\n- OCR 需要安装 Tesseract 引擎\n- 大文件处理可能需要较长时间\n- 转换效果取决于 PDF 原始质量"},{"path":"_meta.json","content":"{\n  \"ownerId\": \"kn7awq8t980mxmrta23czncynd85m0z7\",\n  \"slug\": \"free-pdf-processor\",\n  \"version\": \"1.0.0\",\n  \"publishedAt\": 1777270783055\n}"},{"path":"skill-card.md","content":"## Description:\n\nProvides command-line PDF utilities for text, image, and table extraction; PDF-to-Word or Excel conversion; OCR; merging and splitting; watermarking; encryption and decryption; compression; and batch processing.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[pengsc1994](https://clawhub.ai/user/pengsc1994)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nExternal users and developers use this skill to process local PDF files by extracting content, converting formats, applying watermarks, encrypting or decrypting files, compressing documents, and running supported batch operations.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: The release evidence reports that PDF passwords can be exposed by command-line usage and by encrypt_pdf.py printing the password.\n\nMitigation: Avoid passing real passwords on the command line, remove password printing before sensitive use, and review generated logs or console output.\n\nRisk: The release evidence reports that generated XLSX files from untrusted PDFs may be unsafe active content.\n\nMitigation: Treat spreadsheets produced from untrusted PDFs as untrusted files and open them only in a controlled environment.\n\nRisk: Batch processing can create or overwrite many output files.\n\nMitigation: Use separate output directories, keep backups of source PDFs, and review the operation before running it over large directories.\n\nRisk: The release evidence recommends extra review before installing dependencies for sensitive PDF handling.\n\nMitigation: Install dependencies in an isolated environment and pin reviewed versions before processing sensitive documents.\n\n## Reference(s):\n\n- [Tesseract OCR Windows installer reference](https://github.com/UB-Mannheim/tesseract/wiki)\n\n## Skill Output:\n\n**Output Type(s):** [text, code, shell commands, configuration, guidance]\n\n**Output Format:** [Markdown guidance with command examples and local file outputs such as TXT, PDF, DOCX, XLSX, image files, and JSON indexes.]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [May create, transform, overwrite, or batch-generate local files depending on the selected PDF operation.]\n\n## Skill Version(s):\n\n1.0.0 (source: server release evidence)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment."},{"path":"requirements.txt","content":"# PDF 处理技能依赖\n# 安装: pip install -r requirements.txt\n\n# 核心 PDF 处理\npymupdf>=1.23.0\npdfplumber>=0.10.0\n\n# Word/Excel 转换\npython-docx>=1.1.0\nopenpyxl>=3.1.0\n\n# 图片处理\nPillow>=10.0.0\n\n# 可选: OCR 支持\n# pytesseract>=0.3.10\n# tessdata>=4.1.0\n\n# 测试\npytest>=7.4.0\npytest-cov>=4.1.0"}],"languages":[],"docsSourceLabel":"CLAWHUB","editorialOverview":"一站式 PDF 处理技能。支持 PDF 文本/图片/表格提取、格式转换（PDF↔Word/Excel）、合并拆分、OCR 识别、批量处理、水印添加、加密解密、压缩等。使用场景： (1) 从 PDF 提取文本内容进行数据分析 (2) 将 PDF 转换为 Word/Excel 方便编辑 (3) 合并或拆分 PDF 文... Skill: pdf-processor Owner: pengsc1994 Summary: 一站式 PDF 处理技能。支持 PDF 文本/图片/表格提取、格式转换（PDF↔Word/Excel）、合并拆分、OCR 识别、批量处理、水印添加、加密解密、压缩等。使用场景： (1) 从 PDF 提取文本内容进行数据分析 (2) 将 PDF 转换为 Word/Excel 方便编辑 (3) 合并或拆分 PDF 文... Tags: latest:1.0.0 Version history: v1.0.0 | 2026-04-27T06:19:43.055Z | user Initial public release: 全面升级为多功能一站式 PDF 处理工具 - 新增 14 个独立 PDF 脚本，覆盖提取文本/图片/表格、OCR、格式转换（PDF↔Word/Excel）、合并拆分、水印、加密解密、压缩、批量处理等功能 - 支持命令行一","editorialQuality":{"score":100,"threshold":65,"status":"ready","wordCount":958,"uniquenessScore":49,"reasons":[]}},"media":{"evidence":{"source":"no-media","verified":false,"confidence":"low","updatedAt":"2026-10-10T12:23:24.209Z","emptyReason":"No screenshots, media assets, or demo links are available."},"primaryImageUrl":null,"mediaAssetCount":0,"assets":[],"demoUrl":null},"ownerResources":{"evidence":{"source":"unclaimed","verified":false,"confidence":"low","updatedAt":"2026-10-10T12:23:24.209Z","emptyReason":"This page has not been claimed by the agent owner."},"hasCustomPage":false,"customPageUpdatedAt":null,"customLinks":[],"structuredLinks":{"docsUrl":null,"demoUrl":null,"supportUrl":null,"pricingUrl":null,"statusUrl":null},"customPage":null},"relatedAgents":{"evidence":{"source":"protocol-neighbors","verified":false,"confidence":"medium","updatedAt":"2026-10-10T14:44:19.586Z","emptyReason":null},"items":[{"id":"8ebccd8e-3863-4187-8355-c3f14e1f9edf","entityType":"agent","canonicalPath":"/agent/iofficeai-aionui","slug":"iofficeai-aionui","name":"AionUi","description":"Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!","url":"https://github.com/iOfficeAI/AionUi","homepage":"https://www.aionui.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-10-09T19:11:12.944Z","createdAt":"2026-02-25T03:38:16.584Z","downloads":null},{"id":"b917f68a-ebff-438e-84f8-3f4b2494c0bc","entityType":"agent","canonicalPath":"/agent/activepieces-activepieces","slug":"activepieces-activepieces","name":"activepieces","description":"AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents","url":"https://github.com/activepieces/activepieces","homepage":"https://www.activepieces.com","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-15T02:22:12.426Z","createdAt":"2026-02-25T03:38:12.412Z","downloads":null},{"id":"5cb26759-3a39-483f-94cf-276a98c13bb8","entityType":"agent","canonicalPath":"/agent/cherryhq-cherry-studio","slug":"cherryhq-cherry-studio","name":"cherry-studio","description":"AI productivity studio with smart chat, autonomous agents, and 300+ assistants. Unified access to frontier LLMs","url":"https://github.com/CherryHQ/cherry-studio","homepage":"https://cherry-ai.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-11T14:38:40.986Z","createdAt":"2026-02-25T03:38:19.379Z","downloads":null},{"id":"6f6582d0-5d76-4f0f-b81d-86520247950b","entityType":"agent","canonicalPath":"/agent/copilotkit-copilotkit","slug":"copilotkit-copilotkit","name":"CopilotKit","description":"The Frontend for Agents & Generative UI. React + Angular","url":"https://github.com/CopilotKit/CopilotKit","homepage":"https://docs.copilotkit.ai","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-03-25T09:50:57.846Z","createdAt":"2026-02-25T03:39:14.617Z","downloads":null}],"links":{"hub":"/agent","source":"/agent/source/clawhub","protocols":[{"label":"OpenClaw","href":"/agent/protocol/openclew"}]}}}