{"id":"6b5e2f22-c0e9-4125-9c86-df9a2bfb9449","entityType":"agent","slug":"clawhub-seairteng-macmini-knowledge-base","name":"Mac 知识库搭建系统","canonicalUrl":"https://www.xpersona.co/agent/clawhub-seairteng-macmini-knowledge-base","canonicalPath":"/agent/clawhub-seairteng-macmini-knowledge-base","generatedAt":"2026-10-10T11:50:24.052Z","source":"CLAWHUB","claimStatus":"UNCLAIMED","verificationTier":"NONE","summary":{"evidence":{"source":"editorial-content","verified":true,"confidence":"high","updatedAt":"2026-10-10T09:13:09.914Z","emptyReason":null},"description":"⚠️ **本 skill 包含以下高危能力，使用前请仔细阅读 SKILL.md 顶部「⚠️ CAPABILITIES & RISKS」章节**： - Shell 执行 + 文件读写 + 安装 Homebrew 包（版本固定）+ 下载 Ollama 模型 - 修改 OpenClaw 配置 + 注册持久 cron 任务 + 飞书 webhook 推送 在 Mac Mini (M4) 上快速搭建本地知识库 + RAG 自然语言搜索系统。 适用场景： - 新 Mac 配置知识库：从零开始安装配置 Ollama、embedding模型、定时任务、文档解析 - 遇到 PDF 提取乱码、定时任务超时、skill 加载失败等问题 - 想要建立每日自动分析文档 + 08:00发送摘要到飞书的流程 - 迁移或复现知识库：打包整个 knowledge 目录和配置到新电脑 - **v1.4 新增**：CMap 残缺度自检（不预设来源）+ 50万字完整提取 + OCR fallback 到 .doc 本 skill 会引导完成：目录结构创建、依赖安装、脚本部署、定时任务注册、OpenClaw 配置。 ⚠️ **重要：能力范围** 本 skill 不只是「搭建」，还包含： - 批量 OCR 修复（扫描 summaries/archives 找乱码 + 重新提取） - 目录归档清理（移动重复/孤儿文件到 .trash/） - 自动定时任务（23:00 分析 + 06:00 飞书推送） 使用前请仔细评估批量修改风险。 Skill: Mac 知识库搭建系统 Owner: seairteng Summary: ⚠️ **本 skill 包含以下高危能力，使用前请仔细阅读 SKILL.md 顶部「⚠️ CAPABILITIES & RISKS」章节**： - Shell 执行 + 文件读写 + 安装 Homebrew 包（版本固定）+ 下载 Ollama 模型 - 修改 OpenClaw 配置 + 注册持久 cron 任务 + 飞书 webhook 推送 在 Mac Mini (M4) 上快速搭建本地知识库 + RAG 自然语言搜索系统。 适用场景： - 新 Mac 配置知识库：从零开始安装配置 Ollama、embedding模型、定时任务、文档解析 - 遇到 PDF 提取乱码、定时任务超时、skill 加载失败等问题 - 想要建立每日自动分析文档 + 08:00发送摘要到飞书的流程 - 迁移或复现知识库：打包整个 knowledge 目录和配置","descriptionLabel":"Technical summary","evidenceSummary":"Capability contract not published. No trust telemetry is available yet. 1.5K downloads reported by the source. Last updated 10/10/2026.","installCommand":"clawhub skill install s170wy77rm29wznbver4enx39986fr18:macmini-knowledge-base","sourceUrl":"https://clawhub.ai/seairteng/macmini-knowledge-base","homepage":"https://clawhub.ai/seairteng/skills/macmini-knowledge-base","primaryLinks":[{"label":"View on ClawHub","url":"https://clawhub.ai/seairteng/macmini-knowledge-base","kind":"source"},{"label":"Homepage","url":"https://clawhub.ai/seairteng/skills/macmini-knowledge-base","kind":"homepage"}],"safetyScore":84,"overallRank":62,"popularityScore":64,"trustScore":null,"claimedByName":null,"isOwner":false,"seoDescription":"⚠️ **本 skill 包含以下高危能力，使用前请仔细阅读 SKILL.md 顶部「⚠️ CAPABILITIES & RISKS」章节**： - Shell 执行 + 文件读写 + 安装 Homebrew 包（版本固定）+ 下载 Ollama 模型 - 修改 OpenClaw 配置 + 注册持久 cron 任务 +"},"coverage":{"evidence":{"source":"public-profile","verified":false,"confidence":"medium","updatedAt":"2026-10-10T09:13:09.914Z","emptyReason":null},"protocols":[{"protocol":"OPENCLEW","label":"OpenClaw","status":"self-declared","notes":"Declared in the public agent profile."}],"capabilities":[],"verifiedCount":0,"selfDeclaredCount":1,"capabilityMatrix":{"rows":[{"key":"OPENCLEW","type":"protocol","support":"unknown","confidenceSource":"profile","notes":"Listed on profile"}],"flattenedTokens":"protocol:OPENCLEW|unknown|profile"}},"adoption":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T09:13:09.914Z","emptyReason":null},"stars":null,"forks":null,"downloads":1532,"packageName":null,"latestVersion":"1.4.7","tractionLabel":"1.5K downloads"},"release":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T09:13:09.914Z","emptyReason":null},"lastUpdatedAt":"2026-10-10T09:13:09.914Z","lastCrawledAt":"2026-10-10T09:13:09.914Z","lastIndexedAt":null,"nextCrawlAt":"2026-10-11T09:13:09.914Z","lastVerifiedAt":null,"highlights":[{"version":"1.4.7","createdAt":"2026-08-17T16:46:31.922Z","changelog":"v1.4.7: SKILL.md version 字段与 _meta.json 同步","fileCount":9,"zipByteSize":28730},{"version":"1.4.6","createdAt":"2026-08-17T16:43:07.414Z","changelog":"v1.4.6: SKILL.md 修补（修复 version 字段 + 加结构化权限表）","fileCount":9,"zipByteSize":28934},{"version":"1.4.5","createdAt":"2026-08-17T16:39:29.912Z","changelog":"v1.4.5: NVIDIA SkillSpector 完整修复（结构化 permissions + SKILL.md 顶部章节 + setup.sh 交互式安装）","fileCount":9,"zipByteSize":29134},{"version":"1.4.4","createdAt":"2026-08-17T13:20:58.009Z","changelog":"v1.4.4: SKILL.md 补充「临时文件处理」警告章节（v1.4.3 漏掉）","fileCount":9,"zipByteSize":27426},{"version":"1.4.3","createdAt":"2026-08-17T13:16:27.350Z","changelog":"v1.4.3: NVIDIA SkillSpector 审查修复（4 项代码安全修复 + 5 项文档警告）","fileCount":9,"zipByteSize":27096},{"version":"1.4.2","createdAt":"2026-08-17T12:16:07.651Z","changelog":"v1.4.2: 修复 ClawHub 显示名（从 V1.4.1 改为 Mac 知识库搭建系统，仅元数据更新）","fileCount":9,"zipByteSize":26409},{"version":"1.4.1","createdAt":"2026-08-12T15:16:14.541Z","changelog":"修复 run_analysis.py OSError Errno 63 文件名过长；新增 sanitize_filename + 异常重试","fileCount":9,"zipByteSize":25923},{"version":"1.4.0","createdAt":"2026-08-12T12:45:23.816Z","changelog":"CMap 残缺度自检 + 50万字完整提取 + OCR fallback 到 .doc","fileCount":9,"zipByteSize":23886}]},"execution":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No published capability contract is available yet."},"installCommand":"clawhub skill install s170wy77rm29wznbver4enx39986fr18:macmini-knowledge-base","setupComplexity":"low","setupSteps":["Setup complexity is classified as HIGH. You must provision dedicated cloud infrastructure or an isolated VM. Do not run this directly on your local workstation.","Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data."],"contract":{"contractStatus":"missing","authModes":[],"requires":[],"forbidden":[],"supportsMcp":false,"supportsA2a":false,"supportsStreaming":false,"inputSchemaRef":null,"outputSchemaRef":null,"dataRegion":null,"contractUpdatedAt":null,"sourceUpdatedAt":null,"freshnessSeconds":null},"invocationGuide":{"preferredApi":{"snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-seairteng-macmini-knowledge-base/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-seairteng-macmini-knowledge-base/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-seairteng-macmini-knowledge-base/trust"},"curlExamples":["curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-seairteng-macmini-knowledge-base/snapshot\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-seairteng-macmini-knowledge-base/contract\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-seairteng-macmini-knowledge-base/trust\""],"jsonRequestTemplate":{"query":"summarize this repo","constraints":{"maxLatencyMs":2000,"protocolPreference":["OPENCLEW"]}},"jsonResponseTemplate":{"ok":true,"result":{"summary":"...","confidence":0.9},"meta":{"source":"CLAWHUB","generatedAt":"2026-10-10T11:50:24.049Z"}},"retryPolicy":{"maxAttempts":3,"backoffMs":[500,1500,3500],"retryableConditions":["HTTP_429","HTTP_503","NETWORK_TIMEOUT"]}},"endpoints":{"dossierUrl":"https://www.xpersona.co/api/v1/agents/clawhub-seairteng-macmini-knowledge-base/dossier","snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-seairteng-macmini-knowledge-base/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-seairteng-macmini-knowledge-base/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-seairteng-macmini-knowledge-base/trust"}},"reliability":{"evidence":{"source":"runtime-metrics","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No trust, reliability, or runtime telemetry is available."},"trust":{"status":"unavailable","handshakeStatus":"UNKNOWN","verificationFreshnessHours":null,"reputationScore":null,"p95LatencyMs":null,"successRate30d":null,"fallbackRate":null,"attempts30d":null,"trustUpdatedAt":null,"trustConfidence":"unknown","sourceUpdatedAt":null,"freshnessSeconds":null},"decisionGuardrails":{"doNotUseIf":["Contract metadata is missing or unavailable for deterministic execution."],"safeUseWhen":[],"riskFlags":["missing_or_unavailable_contract","trust_data_unavailable","schema_references_missing"],"operationalConfidence":"low"},"executionMetrics":{"observedLatencyMsP50":null,"observedLatencyMsP95":null,"estimatedCostUsd":null,"uptime30d":null,"rateLimitRpm":null,"rateLimitBurst":null,"lastVerifiedAt":null,"verificationSource":null},"runtimeMetrics":{"successRate":null,"avgLatencyMs":null,"avgCostUsd":null,"hallucinationRate":null,"retryRate":null,"disputeRate":null,"p50Latency":null,"p95Latency":null,"lastUpdated":null}},"benchmarks":{"evidence":{"source":"no-benchmark-data","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No benchmark suites or observed failure patterns are available."},"suites":[],"failurePatterns":[]},"artifacts":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"high","updatedAt":"2026-10-10T09:13:09.914Z","emptyReason":null},"readme":"Skill: Mac 知识库搭建系统\n\nOwner: seairteng\n\nSummary: ⚠️ **本 skill 包含以下高危能力，使用前请仔细阅读 SKILL.md 顶部「⚠️ CAPABILITIES & RISKS」章节**： - Shell 执行 + 文件读写 + 安装 Homebrew 包（版本固定）+ 下载 Ollama 模型 - 修改 OpenClaw 配置 + 注册持久 cron 任务 + 飞书 webhook 推送 在 Mac Mini (M4) 上快速搭建本地知识库 + RAG 自然语言搜索系统。 适用场景： - 新 Mac 配置知识库：从零开始安装配置 Ollama、embedding模型、定时任务、文档解析 - 遇到 PDF 提取乱码、定时任务超时、skill 加载失败等问题 - 想要建立每日自动分析文档 + 08:00发送摘要到飞书的流程 - 迁移或复现知识库：打包整个 knowledge 目录和配置到新电脑 - **v1.4 新增**：CMap 残缺度自检（不预设来源）+ 50万字完整提取 + OCR fallback 到 .doc 本 skill 会引导完成：目录结构创建、依赖安装、脚本部署、定时任务注册、OpenClaw 配置。 ⚠️ **重要：能力范围** 本 skill 不只是「搭建」，还包含： - 批量 OCR 修复（扫描 summaries/archives 找乱码 + 重新提取） - 目录归档清理（移动重复/孤儿文件到 .trash/） - 自动定时任务（23:00 分析 + 06:00 飞书推送） 使用前请仔细评估批量修改风险。\n\nTags: RAG:1.4.7, knowledge-base:1.4.7, latest:1.4.7, mac:1.4.7, ocr:1.4.7, ollama:1.4.7, pdf:1.4.7, security:1.4.7\n\nVersion history:\n\nv1.4.7 | 2026-08-17T16:46:31.922Z | user\n\nv1.4.7: SKILL.md version 字段与 _meta.json 同步\n\nv1.4.6 | 2026-08-17T16:43:07.414Z | user\n\nv1.4.6: SKILL.md 修补（修复 version 字段 + 加结构化权限表）\n\nv1.4.5 | 2026-08-17T16:39:29.912Z | user\n\nv1.4.5: NVIDIA SkillSpector 完整修复（结构化 permissions + SKILL.md 顶部章节 + setup.sh 交互式安装）\n\nv1.4.4 | 2026-08-17T13:20:58.009Z | user\n\nv1.4.4: SKILL.md 补充「临时文件处理」警告章节（v1.4.3 漏掉）\n\nv1.4.3 | 2026-08-17T13:16:27.350Z | user\n\nv1.4.3: NVIDIA SkillSpector 审查修复（4 项代码安全修复 + 5 项文档警告）\n\nv1.4.2 | 2026-08-17T12:16:07.651Z | user\n\nv1.4.2: 修复 ClawHub 显示名（从 V1.4.1 改为 Mac 知识库搭建系统，仅元数据更新）\n\nv1.4.1 | 2026-08-12T15:16:14.541Z | user\n\n修复 run_analysis.py OSError Errno 63 文件名过长；新增 sanitize_filename + 异常重试\n\nv1.4.0 | 2026-08-12T12:45:23.816Z | user\n\nCMap 残缺度自检 + 50万字完整提取 + OCR fallback 到 .doc\n\nv1.3.0 | 2026-05-27T20:54:05.585Z | auto\n\n**v1.3.0 — 引入 kreuzberg 文档统一提取层和 antiword 极速 DOC 支持**\n\n- 新增 kreuzberg 统一提取层，自动适配 PDF/DOCX/XLSX/PPTX/MD/图片 OCR\n- .doc 文件优先走 antiword 极速专线（85% 成功率，大文件<1秒，失败自动兜底 soffice 转换）\n- 新增 pandoc 依赖，提升 DOCX/PPTX/MD 等格式解析稳定性\n- setup 脚本与手动依赖安装说明同步更新，简化部署\n- run_analysis.py 和 generate_catalog.py 已升级为自动适配所有主流文档格式\n\nv1.2.1 | 2026-05-21T21:01:26.112Z | auto\n\n- 目录和脚本更新：文章目录路径调整，目录生成脚本由 JS 重写为 Python（generate_catalog.js → generate_catalog.py）。\n- 新增脚本：增加 generate_catalog.py 和 utils.py，提升目录生成和工具函数能力。\n- 更新定时任务说明：定时任务脚本及调用命令由 JS 切换为 Python，文档同步修订相关命令行调用。\n- run_analysis.py 优化：改善文档分析分批处理，兼容新批量处理脚本。\n- 细节修正：关键路径、目录结构、定时任务时间等文档信息同步调整。\n\nv1.2.0 | 2026-05-15T02:44:29.766Z | auto\n\nmacmini-knowledge-base 1.2.0\n\n- 新增“分批处理”机制，run_analysis.py和generate_catalog.js支持大批量文件断点续跑，自动分多批处理，每批完成后保存进度。\n- 防止定时任务因文件过多导致超时失败，提升稳定性和可控性。\n- SKILL.md 增加了详细的分批处理流程说明和手动断点触发命令。\n- 现有功能和安装方法保持兼容，无需更改已有部署。\n\nv1.1.0 | 2026-05-12T18:12:51.019Z | auto\n\n**Version 1.1.0 introduces enhanced document processing and improved workflow details:**\n\n- 支持 PDF（正常/乱码/图片型）、PPTX、DOCX、XLSX、MD 等多格式文档自动处理，智能选择最优提取/OCR方案。\n- 新增关键词库（中英文双语）与标签自动识别，目录生成更智能、可缓存加速。\n- cron 定时分析任务默认超时时长由 300 → 600 秒，稳定性提升。\n- 文档详细描述 generate_catalog.js 的文件检测与处理逻辑，以及缓存机制。\n- 安装说明和路径表述更清晰，支持更便捷的手动部署与迁移。\n- 避坑指南与关键路径更新，确保新场景和常见问题有据可循。\n\nv1.0.0 | 2026-05-10T09:14:19.524Z | auto\n\n- Initial release for setting up a local knowledge base and RAG search system on Mac Mini (M4).\n- Includes both one-click and step-by-step installation guides for environment setup, dependencies, embedding model, and script deployment.\n- Provides detailed instructions for OpenClaw configuration and scheduled task registration, including integration with Feishu.\n- Features troubleshooting tips for common issues such as PDF extraction errors, task timeouts, and tool permissions.\n- Supports easy migration to a new machine by copying the knowledge directory and re-running install steps.\n\nArchive index:\n\nArchive v1.4.7: 9 files, 28730 bytes\n\nFiles: _meta.json (141b), CHANGELOG.md (573b), scripts/generate_catalog.py (12426b), scripts/re_ocr_corrupted.py (7276b), scripts/run_analysis.py (8164b), scripts/setup.sh (8552b), scripts/utils.py (17289b), skill-card.md (2438b), SKILL.md (20652b)\n\nFile v1.4.7:SKILL.md\n\n---\nname: macmini-knowledge-base\nversion: 1.4.7\ndescription: |\n  ⚠️ **本 skill 包含以下高危能力，使用前请仔细阅读 SKILL.md 顶部「⚠️ CAPABILITIES & RISKS」章节**：\n  - Shell 执行 + 文件读写 + 安装 Homebrew 包（版本固定）+ 下载 Ollama 模型\n  - 修改 OpenClaw 配置 + 注册持久 cron 任务 + 飞书 webhook 推送\n\n  在 Mac Mini (M4) 上快速搭建本地知识库 + RAG 自然语言搜索系统。\n  适用场景：\n  - 新 Mac 配置知识库：从零开始安装配置 Ollama、embedding模型、定时任务、文档解析\n  - 遇到 PDF 提取乱码、定时任务超时、skill 加载失败等问题\n  - 想要建立每日自动分析文档 + 08:00发送摘要到飞书的流程\n  - 迁移或复现知识库：打包整个 knowledge 目录和配置到新电脑\n  - **v1.4 新增**：CMap 残缺度自检（不预设来源）+ 50万字完整提取 + OCR fallback 到 .doc\n  本 skill 会引导完成：目录结构创建、依赖安装、脚本部署、定时任务注册、OpenClaw 配置。\n\n  ⚠️ **重要：能力范围**\n  本 skill 不只是「搭建」，还包含：\n  - 批量 OCR 修复（扫描 summaries/archives 找乱码 + 重新提取）\n  - 目录归档清理（移动重复/孤儿文件到 .trash/）\n  - 自动定时任务（23:00 分析 + 06:00 飞书推送）\n  \n  使用前请仔细评估批量修改风险。\n---\n\n# Knowledge Base Setup\n\n在 Mac Mini 上快速搭建本地知识库 + RAG 搜索系统。\n\n## ⚠️ 阅读前必读：本 skill 的能力范围\n\n本 skill **不只是\"搭建知识库\"**，还包含以下高危能力：\n\n**执行能力**：\n- 🔧 **Shell 命令执行**（python3 + bash 脚本）\n- 📁 **文件读写**（knowledge/, summaries/, archives/, .trash/）\n- 📦 **安装 Homebrew 包**（antiword, tesseract, pandoc, libreoffice，**版本固定**）\n- ⬇️ **下载 Ollama 模型**（nomic-embed-text, ~274MB）\n- ⚙️ **修改 OpenClaw 配置**（~/.openclaw/openclaw.json）\n- ⏰ **注册持久 cron 任务**（23:00 + 06:00，**每天自动**）\n- 📤 **推送消息到飞书 webhook**\n\n**持久化影响**：\n- 知识库目录会被自动分析（每天 23:00）\n- 摘要文件会被覆盖写入（OCR 修复时）\n- cron 任务永久执行（直到手动 `openclaw cron remove <id>`）\n- 失败文件移到 `.trash/`（7 天兜底清理）\n\n**安装流程**：\n本 skill 的 setup.sh 是**交互式安装向导**：\n- 每个危险操作前会要求 y/N 确认\n- 提供 `--dry-run` 选项查看会做什么\n- 已安装用户重跑会进入确认模式\n\n**如果不同意上述任何一项，请不要安装本 skill。**\n\n### 结构化权限声明（Structured Permissions）\n\n| 权限 | 必填 | 范围 | 用途 / 风险 |\n|------|------|------|-------------|\n| `exec` | ✅ | python3 + bash scripts | 文档提取 + 飞书推送 |\n| `file_read` | ✅ | `~/.openclaw/workspace/knowledge/` | 读取文档 + summaries |\n| `file_write` | ✅ | `summaries/`, `archives/`, `.trash/` | OCR 修复覆盖旧摘要 |\n| `install_packages` | ✅ | brew: antiword, tesseract, pandoc, libreoffice | Homebrew 包安装（用户确认）|\n| `download_model` | ✅ | ollama: nomic-embed-text (~274MB) | Ollama 模型下载（用户确认）|\n| `modify_config` | ✅ | `~/.openclaw/openclaw.json` | 添加 alsoAllow: [exec, process] |\n| `register_cron` | ✅ | 23:00 分析 + 06:00 推送 | 持久化定时任务（用户确认）|\n| `network` | ✅ | 飞书 webhook + Ollama 下载 | 外部 API 调用 |\n\n**warning**: 本 skill 会自动修改文件、安装包、注册 cron 任务（用户每步都有 y/N 确认）\n\n**disable_command**: `openclaw cron remove <id>`\n\n---\n\n## 核心功能（v2.0）\n\n- **kreuzberg 统一提取层**：PDF / DOCX / XLSX / PPTX / MD / 图片 OCR 全自动路由\n- **antiword 极速专线**：.doc 文件专用提取，成功率 85%，169MB 文件 0.02 秒完成\n- **智能兜底**：antiword 失败自动走 soffice 转换，60 秒硬超时无误判\n- **自动分类**：关键词匹配驱动，中英文双语标签\n- **定时任务**：每天 23:00 分析新文档，08:00 发送摘要到飞书\n\n\n\n## 🔧 权限声明\n\n本 skill 在使用时需要以下 OpenClaw 工具能力：\n\n```json\n{\n  \"tools\": {\n    \"alsoAllow\": [\"exec\", \"process\", \"read\", \"write\"]\n  }\n}\n```\n\n⚠️ **执行风险**：exec + cron 自动化组合可导致持续命令执行，\n请在可信环境（个人 Mac）使用，不要在共享/服务器部署。\n\n\n\n## ⚠️ 安全警告：定时任务\n\n本 skill 注册 2 个 cron 任务（23:00 分析 + 06:00 推送），\n运行 shell 命令并自动推送消息到飞书。\n\n**潜在风险**：\n- 脚本路径被修改 → 自动执行任意命令\n- 知识库目录被入侵 → 自动读取/外发\n- 飞书 webhook 泄漏 → 自动推送被劫持\n\n**建议**：\n- 不要把 `~/.openclaw/workspace/knowledge` 放在共享/多用户目录\n- 定期检查 cron 配置（`openclaw cron list`）\n- 飞书 webhook 使用独立群组，不要复用其他机器人的 webhook\n- 仅在个人 Mac 上运行，不要部署到服务器\n\n## 快速开始\n\n### 一键安装\n\n```bash\ncd ~/.openclaw/workspace/skills/knowledge-base-setup/scripts\nbash setup.sh <飞书用户ID>\n```\n\n### 手动分步安装\n\n**Step 1: 系统依赖**\n```bash\nbrew install antiword tesseract pandoc\n```\n\n**Step 2: Python 依赖**\n```bash\npip3 install kreuzberg pytesseract pymupdf docx openpyxl python-pptx\n```\n\n**Step 3: Ollama + embedding 模型**\n```bash\n# 安装 Ollama: https://ollama.com/download\nollama pull nomic-embed-text\n```\n\n**Step 4: 创建目录结构**\n```bash\nmkdir -p ~/.openclaw/workspace/knowledge/.analysis/summaries/archives\nmkdir -p ~/.openclaw/workspace/knowledge/temp_docs\ntouch ~/.openclaw/workspace/knowledge/文章目录/文章目录.md\n```\n\n**Step 5: 部署脚本**\n```bash\ncp ~/.openclaw/workspace/skills/knowledge-base-setup/scripts/*.py \\\n   ~/.openclaw/workspace/knowledge/.analysis/\nchmod +x ~/.openclaw/workspace/knowledge/.analysis/*.py\n```\n\n**Step 6: 配置 OpenClaw**\n\n编辑 `~/.openclaw/openclaw.json`，加入：\n```json\n{\n  \"models\": {\n    \"providers\": {\n      \"ollama\": {\n        \"baseUrl\": \"http://127.0.0.1:11434\",\n        \"api\": \"ollama\",\n        \"models\": [\n          {\"id\": \"nomic-embed-text\", \"name\": \"Nomic Embed Text\"}\n        ]\n      }\n    }\n  },\n  \"agents\": {\n    \"defaults\": {\n      \"memorySearch\": {\n        \"provider\": \"ollama\",\n        \"model\": \"nomic-embed-text\"\n      }\n    }\n  }\n}\n```\n\n确保 tools 区块有：\n```json\n\"tools\": {\n    \"alsoAllow\": [\"exec\", \"process\"]\n}\n```\n\n然后重启：`openclaw gateway restart`\n\n**Step 7: 注册定时任务**\n```bash\n# 23:00 分析新文档\nopenclaw cron add \\\n  --name \"23:00分析新文档\" \\\n  --cron \"0 23 * * *\" \\\n  --tz \"Asia/Shanghai\" \\\n  --session isolated \\\n  --timeout-seconds 600 \\\n  --message \"cd ~/.openclaw/workspace/knowledge/.analysis && python3 run_analysis.py && python3 generate_catalog.py\" \\\n  --announce --channel feishu --to \"user:<飞书用户ID>\"\n\n# 08:00 发送文档摘要\nopenclaw cron add \\\n  --name \"08:00发送文档摘要\" \\\n  --cron \"0 8 * * *\" \\\n  --tz \"Asia/Shanghai\" \\\n  --session isolated \\\n  --timeout-seconds 120 \\\n  --message \"读取 summaries/ 目录发送摘要到飞书\" \\\n  --announce --channel feishu --to \"user:<飞书用户ID>\"\n```\n\n\n\n\n\n## summary 文件名 sanitize（v1.4.1 新增）\n\n防止 `OSError: [Errno 63] File name too long`（NAME_MAX=255 bytes）：\n\n```python\nSUMMARY_NAME_MAX = 200\n\ndef sanitize_filename(name, max_length=SUMMARY_NAME_MAX):\n    \"\"\"截断超长文件名，保留扩展名 + 8 位 MD5 hash 防冲突\"\"\"\n    name_bytes = name.encode('utf-8')\n    if len(name_bytes) <= max_length:\n        return name\n    \n    base, ext = os.path.splitext(name)\n    ext_bytes = ext.encode('utf-8')\n    base_bytes = base.encode('utf-8')\n    \n    import hashlib\n    h = hashlib.md5(name_bytes).hexdigest()[:8]\n    \n    reserve = len(ext_bytes) + 1 + 8  # \"_\" + hash + ext\n    available = max_length - reserve\n    \n    if available > 0 and len(base_bytes) > available:\n        truncated = base_bytes[:available].decode('utf-8', errors='ignore')\n        return f\"{truncated}_{h}{ext}\"\n    \n    return name[:max_length]\n```\n\n主循环的异常捕获重试：\n\n```python\ntry:\n    with open(summary_file, 'w', encoding='utf-8') as f:\n        f.write(content)\nexcept OSError as e:\n    if e.errno == 63:  # ENAMETOOLONG\n        short_name = sanitize_filename(filename, max_length=180)\n        summary_file = os.path.join(\n            SUMMARY_DIR,\n            f\"{timestamp}_{short_name}.summary.txt\"\n        )\n        with open(summary_file, 'w', encoding='utf-8') as f:\n            f.write(content)\n```\n\n**触发场景：** 畸形 PDF 文件名（如下载错误的 `_; filename_=utf-8''...` 双名拼接），原文件名 244+ bytes + 时间戳超 255 bytes 限制。\n\n**实测案例：** `20260730-Nomura-Asia Insights：China：The Politburo meeting indicated a shift to _countercyclical\" policies-260730.pdf_; filename_=utf-8''...pdf` (原 248 bytes) → sanitize 后 200 bytes + 8位 hash → 安全创建。\n\n### temp_docs 畸形文件清理（v1.4.1 新增）\n\n下载失败的 PDF 在文件名里重复了两次（`_; filename_=utf-8''` 分隔），实际只需保留前半。一次性清理脚本：\n\n```python\nimport os, shutil\ntemp_docs = os.path.expanduser(\"~/.openclaw/workspace/knowledge/temp_docs\")\ntrash_dir = os.path.expanduser(\"~/.openclaw/workspace/knowledge/.trash/temp_docs_<时间戳>\")\nos.makedirs(trash_dir, exist_ok=True)\n\nfor f in os.listdir(temp_docs):\n    if \"_; filename_=utf-8''\" in f:\n        full = os.path.join(temp_docs, f)\n        parts = f.split(\"_; filename_=utf-8''\")\n        real_name = parts[0]\n        target = os.path.join(temp_docs, real_name)\n        if not os.path.exists(target):\n            shutil.move(full, target)\n            print(f\"重命名: {real_name}\")\n```\n\n\n\n## CMap 残缺度自检（v1.4 新增）\n\n不预设\"哪个 PDF 来源会乱码\"——实测 **72% 的乱码来自非 lightpdf PDF**（PPT 转 PDF、扫描件等），\n改用**自适应检测**：\n\n```python\ndef is_cmap_broken(text, threshold=0.03):\n    \"\"\"检测文本是否含异常字符（CMap 残缺/PUA 污染/未映射 CID）\"\"\"\n    if not text or len(text.strip()) < 50:\n        return False\n    total = len(text)\n    pua_count = sum(1 for c in text if 0xE000 <= ord(c) <= 0xF8FF)\n    cjk_ext = sum(1 for c in text if 0x20000 <= ord(c) <= 0x2EBEF)\n    cjk_compat = sum(1 for c in text if 0xF900 <= ord(c) <= 0xFAFF)\n    cid_count = text.count('(cid:')\n    bad_ratio = (pua_count + cjk_ext + cjk_compat + cid_count) / total\n    return bad_ratio > threshold or cid_count > 10\n```\n\n**3 类乱码特征：**\n1. **PUA 私用区** (U+E000-F8FF) —— 残缺 CMap fallback\n2. **CJK 扩展区** (U+20000-2EBEF) —— 字符找不到映射\n3. **`(cid:xxxx)` 字面值** —— pdfplumber 提取失败标志\n\n**集成位置：** `extract_pdf_text()` 在 kreuzberg / pymupdf 提取后调 `is_cmap_broken()`，\n通过即返回，失败即触发 OCR 路径。\n\n### OCR 性能实测（2026-08-12 验证）\n\n| 文件 | 大小 | OCR 耗时 | 备注 |\n|---|---|---|---|\n| lightpdf PDF | 4 页 | 13.5 秒 | CMap 残缺，自动 OCR |\n| 大型 PPT 转 PDF      | 90 页 | 0.5 秒 | 默认路径（无需 OCR）|\n| 65MB .doc | 169MB 文件 | 0.1 秒 | antiword 极速专线 |\n| 大型 docx（475K 字）| 562KB | 11.2 秒 | python-docx fallback |\n| OCR 自检总开销 | - | < 200ms | 3 页抽样 + 字符统计 |\n\n### 批量 OCR 修复脚本（v1.4 新增）\n\n`re_ocr_corrupted.py` —— 批量扫描乱码 summary，自动用新版本 utils 重新提取：\n\n```bash\n# 干跑（不写文件）\npython3 re_ocr_corrupted.py --dry-run --max 10\n\n# 实际批量（处理所有乱码）\npython3 re_ocr_corrupted.py --max 100\n\n# 只处理指定 PDF\npython3 re_ocr_corrupted.py --pdf-list \"path1.pdf,path2.pdf\"\n```\n\n行为：\n1. 扫 archives/ 找出乱码 summary\n2. 按 basename 匹配源 PDF\n3. 调 `extract_pdf_text()` 重跑（自动 OCR fallback）\n4. 写新 summary 到 summaries/（带新时间戳）\n5. 覆盖 archives/ 里对应 basename 的所有乱码版本\n6. 输出 JSON 报告（含每份文件路径/字数/成功状态）\n\n## 文档解析架构（v2.0）\n\n### 架构图\n\n```\n                    ┌──────────────────────────────────────┐\n                    │         kreuzberg 统一提取层           │\n                    │  (pypdfium2 / python-calamine / pandoc) │\n                    └───┬────────────────────────────────┬───┘\n                        │                              │\n                自动判断 │                              │\n                        ▼                              ▼\n              ┌─────────────────┐           ┌─────────────────────┐\n              │  kreuzberg 直提  │           │  antiword 极速专线  │\n              │ PDF/DOCX/XLSX/  │           │   (.doc 文件专用)    │\n              │ PPTX/MD/图片OCR │           │   成功率 85%，<1秒   │\n              └─────────────────┘           └─────────────────────┘\n                        │                              │\n                        │         ┌──────────────────────────────┐\n                        │         │     soffice 兜底转换          │\n                        │         │ (.doc/.xls/.ppt antiword失败) │\n                        │         │  60秒硬超时（消除误判watchdog）│\n                        │         └──────────────────────────────┘\n                        ▼                              │\n              ┌──────────────────────────────────────────────┐\n              │              文本输出（content）              │\n              │  → summaries/ 摘要文件 → generate_catalog.py  │\n              └──────────────────────────────────────────────┘\n```\n\n### 文件类型 × 提取方式\n\n| 格式 | 主方案 | 依赖 | 成功率 | 单文件速度 |\n|------|--------|------|--------|-----------|\n| PDF | kreuzberg (pypdfium2) | 无 | ~100% | 0.05-0.7s |\n| DOCX | kreuzberg + pandoc | pandoc 3.9+ | 100% | 0.12-3s |\n| XLSX | kreuzberg (python-calamine) | 无 | 100% | 0.1-0.5s |\n| PPTX | kreuzberg + pandoc | pandoc 3.9+ | 100% | 0.02-0.2s |\n| MD | kreuzberg + pandoc | pandoc 3.9+ | 100% | <0.01s |\n| **.doc** | **antiword 优先** | antiword | **85%**，<1秒 | <0.02s |\n| .doc（失败） | soffice 兜底 | LibreOffice | ~15% | 2-21s |\n| .xls | soffice → XLSX | LibreOffice | ~95% | 2-10s |\n| .ppt | soffice → PPTX | LibreOffice | ~95% | 2-10s |\n| 图片 | kreuzberg 内置 OCR | tesseract | ~90% | 3-10s |\n\n### antiword 极速专线\n\n```python\n# 实测数据：\n# 169MB 超大文件 → 26万字符，0.02秒完成\n# 正常 .doc（0.1-15MB）→ <1秒\n# 成功率 85%，覆盖绝大多数 .doc 文件\nresult = subprocess.run(['antiword', filepath], capture_output=True, timeout=10)\n```\n\n### kreuzberg 统一提取层\n\nkreuzberg 是专业的非结构化文档文本提取库（支持 20+ 格式），内部自动路由：\n- PDF → pypdfium2\n- XLSX → python-calamine\n- DOCX/PPTX/MD → pandoc\n- 图片 → 内置 OCR（tesseract）\n\n## 关键词库（中英双语）\n\n**中文（47个）：** 房产、房价、房地产、居民、消费、股市、经济、政策、利率、通胀、人民币、A股、美联储、PBOC、GDP、股票、资产、投资、债券、银行、PPI、CPI、PMI、M2、就业、失业、汽车、新能源、AI 等\n\n**英文（70+个）：** property、real estate、GDP、inflation、CPI、PPI、PMI、PBOC、Fed、consumer、economy、growth、housing、stock market、EV、AI 等\n\n**标签输出语言：** 自动判断——英文内容匹配英文关键词输出英文标签，中文内容匹配中文关键词输出中文标签\n\n## 定时任务兼容性\n\n| 任务 | ID | 调用方式 | 结论 |\n|------|------|---------|------|\n| 23:00分析新文档 | f3536e18 | 绝对路径 `python3 run_analysis.py` | ✅ 无需修改 |\n| 07:00生成财经早报 | b741c6d5 | Node.js 脚本 | ❌ 不相关 |\n| 08:00发送财经早报 | a7cbaacc | 读取文件发送 | ❌ 不相关 |\n| 09:00发送文档摘要 | 89b4cf75 | 读取 summaries 目录 | ❌ 不相关 |\n\n## 迁移到新电脑\n\n⚠️ **迁移前必读**：`~/.openclaw/workspace/knowledge/` 目录可能包含：\n- 私人合同/财务文档的 OCR 摘要\n- 个人分析报告\n- 飞书推送缓存\n\n**建议**：\n1. 先 `du -sh ~/.openclaw/workspace/knowledge/` 看大小\n2. 排除 `.trash/`、`.analysis/cache/` 后再迁移\n3. 用 `rsync -av --exclude='.trash' ...` 而不是 `scp -r`\n\n1. 复制目录（推荐 rsync）：\n   ```bash\n   rsync -av --exclude='.trash' --exclude='.analysis/cache' \\\n     ~/.openclaw/workspace/knowledge user@new-mac:~/.openclaw/workspace/\n   ```\n2. 在新电脑运行 `bash setup.sh <飞书用户ID>`\n3. 重新注册定时任务（Job ID 会变）\n\n## 避坑指南\n\n| 问题 | 原因 | 解决 |\n|------|------|------|\n| LibreOffice 超时 | watchdog 误判大文件为卡死 | v2.0 移除 watchdog，60秒硬超时 |\n| .doc 提取慢 | 统一走 LibreOffice | antiword 专线，169MB 文件 0.02秒 |\n| DOCX/PPTX 处理失败 | pandoc 未安装 | `brew install pandoc` |\n| PDF 提取乱码（OCR 不工作） | ocrmypdf `--skip-text` 跳过乱码页 | v1.4 改为 `--force-ocr` 强制 OCR |\n| PDF 漏检 CMap 残缺 | 没主动判断是否乱码 | v1.4 `is_cmap_broken()` 自检（阈值 0.03）|\n| 文本被截断到 8000 字 | 硬编码 `[:8000]` 太短 | v1.4 `MAX_EXTRACT_LEN = 500_000` |\n| .doc 提取失败 | lightpdf 处理过的 .doc 乱码 | v1.4 `ocr_office_via_ocr()` 兜底 |\n| summary 文件名过长失败 | 畸形 PDF 名 244+ bytes + 时间戳超 NAME_MAX | v1.4.1 `sanitize_filename()` + Errno 63 重试 |\n| 静默失败（不知道哪个文件）| 不抛异常 | v1.4 `PDFExtractError` / `ExtractError` 含路径 |\n| 飞书无 exec 工具 | tools 策略限制 | 添加 `alsoAllow: [exec, process]` |\n| BGE-M3 卡顿 | 16GB 内存不足 | 继续用 nomic-embed-text |\n\n## 关键路径\n\n| 内容 | 路径 |\n|------|------|\n| Skill 目录 | `~/.openclaw/workspace/skills/knowledge-base-setup/` |\n| 知识库 | `~/.openclaw/workspace/knowledge/` |\n| 分析脚本 | `~/.openclaw/workspace/knowledge/.analysis/` |\n| 目录缓存 | `~/.openclaw/workspace/knowledge/.analysis/.catalog_cache.json` |\n| 摘要输出 | `~/.openclaw/workspace/knowledge/.analysis/summaries/` |\n| 文章目录 | `~/.openclaw/workspace/knowledge/文章目录/文章目录.md` |\n| OpenClaw 配置 | `~/.openclaw/openclaw.json` |\n\n\n## ⚠️ 临时文件处理（v1.4.3 修复）\n\nOCR 流程会把 PDF 复制到临时目录处理。\n\n**v1.4.3 之前**：使用 `/tmp/ocrmypdf_work`、`/tmp/office_ocr_work`、\n`/tmp/office_convert` 共享路径，存在以下风险：\n- 多进程并发可能冲突\n- 多用户系统下其他用户可访问（权限默认 755）\n- 处理失败时临时文件残留\n\n**v1.4.3 修复**：\n- 用 `tempfile.mkdtemp(prefix=\"...\")` 创建 per-run 私有目录（权限 0o700）\n- 处理完成后立即 `shutil.rmtree` 清理\n- `tempfile.mkstemp` 创建稳定输出文件（避免被 finally 误删）\n\n**剩余风险**：极端情况下（机器突然断电）可能残留临时目录。\n建议定期清理 `/Users/home/.openclaw/tmp/` 下 `ocrmypdf_*`、`office_*` 前缀目录。\n\n## 版本历史\n\n| 版本 | 日期 | 更新内容 |\n|------|------|---------|\n| 1.0.0 | 2026-05-10 | 初始版本，PyMuPDF + LibreOffice 链路 |\n| 1.1.0 | 2026-05-13 | 三步 PDF 处理，关键词库，双语标签 |\n| 1.2.0 | 2026-05-21 | 分批处理优化，280秒断点 |\n| 1.2.1 | 2026-05-22 | utils.py 共享模块重构，LibreOffice 熔断机制 |\n| 1.3.0 | 2026-05-28 | kreuzberg 统一提取层 + antiword 专线 + pandoc |\n| 1.4.0 | 2026-08-12 | CMap 残缺度自检 + 50万字完整提取 + OCR fallback 到 .doc |\n| **1.4.1** | **2026-08-12** | **run_analysis.py: sanitize_filename + Errno 63 重试 + 清理 temp_docs 畸形文件** |\n\nFile v1.4.7:_meta.json\n\n{\n  \"ownerId\": \"kn7d38vmg57htam960qsj7wnkh86ewjb\",\n  \"slug\": \"macmini-knowledge-base\",\n  \"version\": \"1.4.7\",\n  \"publishedAt\": 1786985191922\n}\n\nFile v1.4.7:CHANGELOG.md\n\n# Changelog\n\n## v1.4.7 (2026-08-17) — SKILL.md version 字段同步\n\n- 修复 SKILL.md 顶部 version 字段（1.4.5 → 1.4.7）\n- 与 _meta.json version 一致\n\n## v1.4.6 (2026-08-17) — SKILL.md 修补\n- 加\"结构化权限声明\"表格\n- 修复 SKILL.md 顶部 version（1.4.2 → 1.4.5）\n\n## v1.4.5 (2026-08-17) — NVIDIA SkillSpector 完整修复\n### 元数据修复\n### SKILL.md 修复\n### setup.sh 改造（交互式）\n\n## v1.4.4 (2026-08-17) — 仅 SKILL.md 补充\n## v1.4.3 (2026-08-17) — 代码安全修复\n## v1.4.2 (2026-08-17) — 仅元数据更新\n\nFile v1.4.7:skill-card.md\n\n## Description:\n\nMac 知识库搭建系统 helps agents guide setup of a local Mac Mini knowledge base and RAG search workflow, including document extraction, OCR repair, catalog generation, scheduled analysis, and Feishu summary delivery.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[seairteng](https://clawhub.ai/user/seairteng)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers and personal Mac users use this skill to configure a local knowledge-base workflow on a trusted Mac Mini, process documents into summaries and catalogs, and schedule recurring analysis plus Feishu notifications.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: Broad exec/process permissions and persistent cron jobs can allow repeated command execution from the local workspace.\n\nMitigation: Use the skill only on a trusted personal Mac, review the setup plan first, avoid broad default permission changes where possible, and remove unwanted cron entries with openclaw cron remove.\n\nRisk: The workflow reads local knowledge-base files and can send summaries through a Feishu webhook.\n\nMitigation: Keep the workspace private, use a dedicated Feishu destination, store the webhook as a secret, and rotate it if it may have been exposed.\n\nRisk: The setup process may install Homebrew and Python dependencies and download an Ollama embedding model.\n\nMitigation: Run the documented dry-run path first and install dependencies separately or from pinned trusted sources before enabling automation.\n\n## Reference(s):\n\n- [ClawHub skill page](https://clawhub.ai/seairteng/skills/macmini-knowledge-base)\n- [Ollama download](https://ollama.com/download)\n\n## Skill Output:\n\n**Output Type(s):** [text, markdown, code, shell commands, configuration, guidance]\n\n**Output Format:** [Markdown guidance with inline shell commands, JSON snippets, and Python code examples]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [The skill guides interactive setup steps and includes dry-run guidance for the setup script.]\n\n## Skill Version(s):\n\n1.4.7 (source: frontmatter, CHANGELOG, _meta.json, server release evidence)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nArchive v1.4.6: 9 files, 28934 bytes\n\nFiles: _meta.json (141b), CHANGELOG.md (878b), scripts/generate_catalog.py (12426b), scripts/re_ocr_corrupted.py (7276b), scripts/run_analysis.py (8164b), scripts/setup.sh (8552b), scripts/utils.py (17289b), skill-card.md (2546b), SKILL.md (20652b)\n\nFile v1.4.6:SKILL.md\n\n---\nname: macmini-knowledge-base\nversion: 1.4.5\ndescription: |\n  ⚠️ **本 skill 包含以下高危能力，使用前请仔细阅读 SKILL.md 顶部「⚠️ CAPABILITIES & RISKS」章节**：\n  - Shell 执行 + 文件读写 + 安装 Homebrew 包（版本固定）+ 下载 Ollama 模型\n  - 修改 OpenClaw 配置 + 注册持久 cron 任务 + 飞书 webhook 推送\n\n  在 Mac Mini (M4) 上快速搭建本地知识库 + RAG 自然语言搜索系统。\n  适用场景：\n  - 新 Mac 配置知识库：从零开始安装配置 Ollama、embedding模型、定时任务、文档解析\n  - 遇到 PDF 提取乱码、定时任务超时、skill 加载失败等问题\n  - 想要建立每日自动分析文档 + 08:00发送摘要到飞书的流程\n  - 迁移或复现知识库：打包整个 knowledge 目录和配置到新电脑\n  - **v1.4 新增**：CMap 残缺度自检（不预设来源）+ 50万字完整提取 + OCR fallback 到 .doc\n  本 skill 会引导完成：目录结构创建、依赖安装、脚本部署、定时任务注册、OpenClaw 配置。\n\n  ⚠️ **重要：能力范围**\n  本 skill 不只是「搭建」，还包含：\n  - 批量 OCR 修复（扫描 summaries/archives 找乱码 + 重新提取）\n  - 目录归档清理（移动重复/孤儿文件到 .trash/）\n  - 自动定时任务（23:00 分析 + 06:00 飞书推送）\n  \n  使用前请仔细评估批量修改风险。\n---\n\n# Knowledge Base Setup\n\n在 Mac Mini 上快速搭建本地知识库 + RAG 搜索系统。\n\n## ⚠️ 阅读前必读：本 skill 的能力范围\n\n本 skill **不只是\"搭建知识库\"**，还包含以下高危能力：\n\n**执行能力**：\n- 🔧 **Shell 命令执行**（python3 + bash 脚本）\n- 📁 **文件读写**（knowledge/, summaries/, archives/, .trash/）\n- 📦 **安装 Homebrew 包**（antiword, tesseract, pandoc, libreoffice，**版本固定**）\n- ⬇️ **下载 Ollama 模型**（nomic-embed-text, ~274MB）\n- ⚙️ **修改 OpenClaw 配置**（~/.openclaw/openclaw.json）\n- ⏰ **注册持久 cron 任务**（23:00 + 06:00，**每天自动**）\n- 📤 **推送消息到飞书 webhook**\n\n**持久化影响**：\n- 知识库目录会被自动分析（每天 23:00）\n- 摘要文件会被覆盖写入（OCR 修复时）\n- cron 任务永久执行（直到手动 `openclaw cron remove <id>`）\n- 失败文件移到 `.trash/`（7 天兜底清理）\n\n**安装流程**：\n本 skill 的 setup.sh 是**交互式安装向导**：\n- 每个危险操作前会要求 y/N 确认\n- 提供 `--dry-run` 选项查看会做什么\n- 已安装用户重跑会进入确认模式\n\n**如果不同意上述任何一项，请不要安装本 skill。**\n\n### 结构化权限声明（Structured Permissions）\n\n| 权限 | 必填 | 范围 | 用途 / 风险 |\n|------|------|------|-------------|\n| `exec` | ✅ | python3 + bash scripts | 文档提取 + 飞书推送 |\n| `file_read` | ✅ | `~/.openclaw/workspace/knowledge/` | 读取文档 + summaries |\n| `file_write` | ✅ | `summaries/`, `archives/`, `.trash/` | OCR 修复覆盖旧摘要 |\n| `install_packages` | ✅ | brew: antiword, tesseract, pandoc, libreoffice | Homebrew 包安装（用户确认）|\n| `download_model` | ✅ | ollama: nomic-embed-text (~274MB) | Ollama 模型下载（用户确认）|\n| `modify_config` | ✅ | `~/.openclaw/openclaw.json` | 添加 alsoAllow: [exec, process] |\n| `register_cron` | ✅ | 23:00 分析 + 06:00 推送 | 持久化定时任务（用户确认）|\n| `network` | ✅ | 飞书 webhook + Ollama 下载 | 外部 API 调用 |\n\n**warning**: 本 skill 会自动修改文件、安装包、注册 cron 任务（用户每步都有 y/N 确认）\n\n**disable_command**: `openclaw cron remove <id>`\n\n---\n\n## 核心功能（v2.0）\n\n- **kreuzberg 统一提取层**：PDF / DOCX / XLSX / PPTX / MD / 图片 OCR 全自动路由\n- **antiword 极速专线**：.doc 文件专用提取，成功率 85%，169MB 文件 0.02 秒完成\n- **智能兜底**：antiword 失败自动走 soffice 转换，60 秒硬超时无误判\n- **自动分类**：关键词匹配驱动，中英文双语标签\n- **定时任务**：每天 23:00 分析新文档，08:00 发送摘要到飞书\n\n\n\n## 🔧 权限声明\n\n本 skill 在使用时需要以下 OpenClaw 工具能力：\n\n```json\n{\n  \"tools\": {\n    \"alsoAllow\": [\"exec\", \"process\", \"read\", \"write\"]\n  }\n}\n```\n\n⚠️ **执行风险**：exec + cron 自动化组合可导致持续命令执行，\n请在可信环境（个人 Mac）使用，不要在共享/服务器部署。\n\n\n\n## ⚠️ 安全警告：定时任务\n\n本 skill 注册 2 个 cron 任务（23:00 分析 + 06:00 推送），\n运行 shell 命令并自动推送消息到飞书。\n\n**潜在风险**：\n- 脚本路径被修改 → 自动执行任意命令\n- 知识库目录被入侵 → 自动读取/外发\n- 飞书 webhook 泄漏 → 自动推送被劫持\n\n**建议**：\n- 不要把 `~/.openclaw/workspace/knowledge` 放在共享/多用户目录\n- 定期检查 cron 配置（`openclaw cron list`）\n- 飞书 webhook 使用独立群组，不要复用其他机器人的 webhook\n- 仅在个人 Mac 上运行，不要部署到服务器\n\n## 快速开始\n\n### 一键安装\n\n```bash\ncd ~/.openclaw/workspace/skills/knowledge-base-setup/scripts\nbash setup.sh <飞书用户ID>\n```\n\n### 手动分步安装\n\n**Step 1: 系统依赖**\n```bash\nbrew install antiword tesseract pandoc\n```\n\n**Step 2: Python 依赖**\n```bash\npip3 install kreuzberg pytesseract pymupdf docx openpyxl python-pptx\n```\n\n**Step 3: Ollama + embedding 模型**\n```bash\n# 安装 Ollama: https://ollama.com/download\nollama pull nomic-embed-text\n```\n\n**Step 4: 创建目录结构**\n```bash\nmkdir -p ~/.openclaw/workspace/knowledge/.analysis/summaries/archives\nmkdir -p ~/.openclaw/workspace/knowledge/temp_docs\ntouch ~/.openclaw/workspace/knowledge/文章目录/文章目录.md\n```\n\n**Step 5: 部署脚本**\n```bash\ncp ~/.openclaw/workspace/skills/knowledge-base-setup/scripts/*.py \\\n   ~/.openclaw/workspace/knowledge/.analysis/\nchmod +x ~/.openclaw/workspace/knowledge/.analysis/*.py\n```\n\n**Step 6: 配置 OpenClaw**\n\n编辑 `~/.openclaw/openclaw.json`，加入：\n```json\n{\n  \"models\": {\n    \"providers\": {\n      \"ollama\": {\n        \"baseUrl\": \"http://127.0.0.1:11434\",\n        \"api\": \"ollama\",\n        \"models\": [\n          {\"id\": \"nomic-embed-text\", \"name\": \"Nomic Embed Text\"}\n        ]\n      }\n    }\n  },\n  \"agents\": {\n    \"defaults\": {\n      \"memorySearch\": {\n        \"provider\": \"ollama\",\n        \"model\": \"nomic-embed-text\"\n      }\n    }\n  }\n}\n```\n\n确保 tools 区块有：\n```json\n\"tools\": {\n    \"alsoAllow\": [\"exec\", \"process\"]\n}\n```\n\n然后重启：`openclaw gateway restart`\n\n**Step 7: 注册定时任务**\n```bash\n# 23:00 分析新文档\nopenclaw cron add \\\n  --name \"23:00分析新文档\" \\\n  --cron \"0 23 * * *\" \\\n  --tz \"Asia/Shanghai\" \\\n  --session isolated \\\n  --timeout-seconds 600 \\\n  --message \"cd ~/.openclaw/workspace/knowledge/.analysis && python3 run_analysis.py && python3 generate_catalog.py\" \\\n  --announce --channel feishu --to \"user:<飞书用户ID>\"\n\n# 08:00 发送文档摘要\nopenclaw cron add \\\n  --name \"08:00发送文档摘要\" \\\n  --cron \"0 8 * * *\" \\\n  --tz \"Asia/Shanghai\" \\\n  --session isolated \\\n  --timeout-seconds 120 \\\n  --message \"读取 summaries/ 目录发送摘要到飞书\" \\\n  --announce --channel feishu --to \"user:<飞书用户ID>\"\n```\n\n\n\n\n\n## summary 文件名 sanitize（v1.4.1 新增）\n\n防止 `OSError: [Errno 63] File name too long`（NAME_MAX=255 bytes）：\n\n```python\nSUMMARY_NAME_MAX = 200\n\ndef sanitize_filename(name, max_length=SUMMARY_NAME_MAX):\n    \"\"\"截断超长文件名，保留扩展名 + 8 位 MD5 hash 防冲突\"\"\"\n    name_bytes = name.encode('utf-8')\n    if len(name_bytes) <= max_length:\n        return name\n    \n    base, ext = os.path.splitext(name)\n    ext_bytes = ext.encode('utf-8')\n    base_bytes = base.encode('utf-8')\n    \n    import hashlib\n    h = hashlib.md5(name_bytes).hexdigest()[:8]\n    \n    reserve = len(ext_bytes) + 1 + 8  # \"_\" + hash + ext\n    available = max_length - reserve\n    \n    if available > 0 and len(base_bytes) > available:\n        truncated = base_bytes[:available].decode('utf-8', errors='ignore')\n        return f\"{truncated}_{h}{ext}\"\n    \n    return name[:max_length]\n```\n\n主循环的异常捕获重试：\n\n```python\ntry:\n    with open(summary_file, 'w', encoding='utf-8') as f:\n        f.write(content)\nexcept OSError as e:\n    if e.errno == 63:  # ENAMETOOLONG\n        short_name = sanitize_filename(filename, max_length=180)\n        summary_file = os.path.join(\n            SUMMARY_DIR,\n            f\"{timestamp}_{short_name}.summary.txt\"\n        )\n        with open(summary_file, 'w', encoding='utf-8') as f:\n            f.write(content)\n```\n\n**触发场景：** 畸形 PDF 文件名（如下载错误的 `_; filename_=utf-8''...` 双名拼接），原文件名 244+ bytes + 时间戳超 255 bytes 限制。\n\n**实测案例：** `20260730-Nomura-Asia Insights：China：The Politburo meeting indicated a shift to _countercyclical\" policies-260730.pdf_; filename_=utf-8''...pdf` (原 248 bytes) → sanitize 后 200 bytes + 8位 hash → 安全创建。\n\n### temp_docs 畸形文件清理（v1.4.1 新增）\n\n下载失败的 PDF 在文件名里重复了两次（`_; filename_=utf-8''` 分隔），实际只需保留前半。一次性清理脚本：\n\n```python\nimport os, shutil\ntemp_docs = os.path.expanduser(\"~/.openclaw/workspace/knowledge/temp_docs\")\ntrash_dir = os.path.expanduser(\"~/.openclaw/workspace/knowledge/.trash/temp_docs_<时间戳>\")\nos.makedirs(trash_dir, exist_ok=True)\n\nfor f in os.listdir(temp_docs):\n    if \"_; filename_=utf-8''\" in f:\n        full = os.path.join(temp_docs, f)\n        parts = f.split(\"_; filename_=utf-8''\")\n        real_name = parts[0]\n        target = os.path.join(temp_docs, real_name)\n        if not os.path.exists(target):\n            shutil.move(full, target)\n            print(f\"重命名: {real_name}\")\n```\n\n\n\n## CMap 残缺度自检（v1.4 新增）\n\n不预设\"哪个 PDF 来源会乱码\"——实测 **72% 的乱码来自非 lightpdf PDF**（PPT 转 PDF、扫描件等），\n改用**自适应检测**：\n\n```python\ndef is_cmap_broken(text, threshold=0.03):\n    \"\"\"检测文本是否含异常字符（CMap 残缺/PUA 污染/未映射 CID）\"\"\"\n    if not text or len(text.strip()) < 50:\n        return False\n    total = len(text)\n    pua_count = sum(1 for c in text if 0xE000 <= ord(c) <= 0xF8FF)\n    cjk_ext = sum(1 for c in text if 0x20000 <= ord(c) <= 0x2EBEF)\n    cjk_compat = sum(1 for c in text if 0xF900 <= ord(c) <= 0xFAFF)\n    cid_count = text.count('(cid:')\n    bad_ratio = (pua_count + cjk_ext + cjk_compat + cid_count) / total\n    return bad_ratio > threshold or cid_count > 10\n```\n\n**3 类乱码特征：**\n1. **PUA 私用区** (U+E000-F8FF) —— 残缺 CMap fallback\n2. **CJK 扩展区** (U+20000-2EBEF) —— 字符找不到映射\n3. **`(cid:xxxx)` 字面值** —— pdfplumber 提取失败标志\n\n**集成位置：** `extract_pdf_text()` 在 kreuzberg / pymupdf 提取后调 `is_cmap_broken()`，\n通过即返回，失败即触发 OCR 路径。\n\n### OCR 性能实测（2026-08-12 验证）\n\n| 文件 | 大小 | OCR 耗时 | 备注 |\n|---|---|---|---|\n| lightpdf PDF | 4 页 | 13.5 秒 | CMap 残缺，自动 OCR |\n| 大型 PPT 转 PDF      | 90 页 | 0.5 秒 | 默认路径（无需 OCR）|\n| 65MB .doc | 169MB 文件 | 0.1 秒 | antiword 极速专线 |\n| 大型 docx（475K 字）| 562KB | 11.2 秒 | python-docx fallback |\n| OCR 自检总开销 | - | < 200ms | 3 页抽样 + 字符统计 |\n\n### 批量 OCR 修复脚本（v1.4 新增）\n\n`re_ocr_corrupted.py` —— 批量扫描乱码 summary，自动用新版本 utils 重新提取：\n\n```bash\n# 干跑（不写文件）\npython3 re_ocr_corrupted.py --dry-run --max 10\n\n# 实际批量（处理所有乱码）\npython3 re_ocr_corrupted.py --max 100\n\n# 只处理指定 PDF\npython3 re_ocr_corrupted.py --pdf-list \"path1.pdf,path2.pdf\"\n```\n\n行为：\n1. 扫 archives/ 找出乱码 summary\n2. 按 basename 匹配源 PDF\n3. 调 `extract_pdf_text()` 重跑（自动 OCR fallback）\n4. 写新 summary 到 summaries/（带新时间戳）\n5. 覆盖 archives/ 里对应 basename 的所有乱码版本\n6. 输出 JSON 报告（含每份文件路径/字数/成功状态）\n\n## 文档解析架构（v2.0）\n\n### 架构图\n\n```\n                    ┌──────────────────────────────────────┐\n                    │         kreuzberg 统一提取层           │\n                    │  (pypdfium2 / python-calamine / pandoc) │\n                    └───┬────────────────────────────────┬───┘\n                        │                              │\n                自动判断 │                              │\n                        ▼                              ▼\n              ┌─────────────────┐           ┌─────────────────────┐\n              │  kreuzberg 直提  │           │  antiword 极速专线  │\n              │ PDF/DOCX/XLSX/  │           │   (.doc 文件专用)    │\n              │ PPTX/MD/图片OCR │           │   成功率 85%，<1秒   │\n              └─────────────────┘           └─────────────────────┘\n                        │                              │\n                        │         ┌──────────────────────────────┐\n                        │         │     soffice 兜底转换          │\n                        │         │ (.doc/.xls/.ppt antiword失败) │\n                        │         │  60秒硬超时（消除误判watchdog）│\n                        │         └──────────────────────────────┘\n                        ▼                              │\n              ┌──────────────────────────────────────────────┐\n              │              文本输出（content）              │\n              │  → summaries/ 摘要文件 → generate_catalog.py  │\n              └──────────────────────────────────────────────┘\n```\n\n### 文件类型 × 提取方式\n\n| 格式 | 主方案 | 依赖 | 成功率 | 单文件速度 |\n|------|--------|------|--------|-----------|\n| PDF | kreuzberg (pypdfium2) | 无 | ~100% | 0.05-0.7s |\n| DOCX | kreuzberg + pandoc | pandoc 3.9+ | 100% | 0.12-3s |\n| XLSX | kreuzberg (python-calamine) | 无 | 100% | 0.1-0.5s |\n| PPTX | kreuzberg + pandoc | pandoc 3.9+ | 100% | 0.02-0.2s |\n| MD | kreuzberg + pandoc | pandoc 3.9+ | 100% | <0.01s |\n| **.doc** | **antiword 优先** | antiword | **85%**，<1秒 | <0.02s |\n| .doc（失败） | soffice 兜底 | LibreOffice | ~15% | 2-21s |\n| .xls | soffice → XLSX | LibreOffice | ~95% | 2-10s |\n| .ppt | soffice → PPTX | LibreOffice | ~95% | 2-10s |\n| 图片 | kreuzberg 内置 OCR | tesseract | ~90% | 3-10s |\n\n### antiword 极速专线\n\n```python\n# 实测数据：\n# 169MB 超大文件 → 26万字符，0.02秒完成\n# 正常 .doc（0.1-15MB）→ <1秒\n# 成功率 85%，覆盖绝大多数 .doc 文件\nresult = subprocess.run(['antiword', filepath], capture_output=True, timeout=10)\n```\n\n### kreuzberg 统一提取层\n\nkreuzberg 是专业的非结构化文档文本提取库（支持 20+ 格式），内部自动路由：\n- PDF → pypdfium2\n- XLSX → python-calamine\n- DOCX/PPTX/MD → pandoc\n- 图片 → 内置 OCR（tesseract）\n\n## 关键词库（中英双语）\n\n**中文（47个）：** 房产、房价、房地产、居民、消费、股市、经济、政策、利率、通胀、人民币、A股、美联储、PBOC、GDP、股票、资产、投资、债券、银行、PPI、CPI、PMI、M2、就业、失业、汽车、新能源、AI 等\n\n**英文（70+个）：** property、real estate、GDP、inflation、CPI、PPI、PMI、PBOC、Fed、consumer、economy、growth、housing、stock market、EV、AI 等\n\n**标签输出语言：** 自动判断——英文内容匹配英文关键词输出英文标签，中文内容匹配中文关键词输出中文标签\n\n## 定时任务兼容性\n\n| 任务 | ID | 调用方式 | 结论 |\n|------|------|---------|------|\n| 23:00分析新文档 | f3536e18 | 绝对路径 `python3 run_analysis.py` | ✅ 无需修改 |\n| 07:00生成财经早报 | b741c6d5 | Node.js 脚本 | ❌ 不相关 |\n| 08:00发送财经早报 | a7cbaacc | 读取文件发送 | ❌ 不相关 |\n| 09:00发送文档摘要 | 89b4cf75 | 读取 summaries 目录 | ❌ 不相关 |\n\n## 迁移到新电脑\n\n⚠️ **迁移前必读**：`~/.openclaw/workspace/knowledge/` 目录可能包含：\n- 私人合同/财务文档的 OCR 摘要\n- 个人分析报告\n- 飞书推送缓存\n\n**建议**：\n1. 先 `du -sh ~/.openclaw/workspace/knowledge/` 看大小\n2. 排除 `.trash/`、`.analysis/cache/` 后再迁移\n3. 用 `rsync -av --exclude='.trash' ...` 而不是 `scp -r`\n\n1. 复制目录（推荐 rsync）：\n   ```bash\n   rsync -av --exclude='.trash' --exclude='.analysis/cache' \\\n     ~/.openclaw/workspace/knowledge user@new-mac:~/.openclaw/workspace/\n   ```\n2. 在新电脑运行 `bash setup.sh <飞书用户ID>`\n3. 重新注册定时任务（Job ID 会变）\n\n## 避坑指南\n\n| 问题 | 原因 | 解决 |\n|------|------|------|\n| LibreOffice 超时 | watchdog 误判大文件为卡死 | v2.0 移除 watchdog，60秒硬超时 |\n| .doc 提取慢 | 统一走 LibreOffice | antiword 专线，169MB 文件 0.02秒 |\n| DOCX/PPTX 处理失败 | pandoc 未安装 | `brew install pandoc` |\n| PDF 提取乱码（OCR 不工作） | ocrmypdf `--skip-text` 跳过乱码页 | v1.4 改为 `--force-ocr` 强制 OCR |\n| PDF 漏检 CMap 残缺 | 没主动判断是否乱码 | v1.4 `is_cmap_broken()` 自检（阈值 0.03）|\n| 文本被截断到 8000 字 | 硬编码 `[:8000]` 太短 | v1.4 `MAX_EXTRACT_LEN = 500_000` |\n| .doc 提取失败 | lightpdf 处理过的 .doc 乱码 | v1.4 `ocr_office_via_ocr()` 兜底 |\n| summary 文件名过长失败 | 畸形 PDF 名 244+ bytes + 时间戳超 NAME_MAX | v1.4.1 `sanitize_filename()` + Errno 63 重试 |\n| 静默失败（不知道哪个文件）| 不抛异常 | v1.4 `PDFExtractError` / `ExtractError` 含路径 |\n| 飞书无 exec 工具 | tools 策略限制 | 添加 `alsoAllow: [exec, process]` |\n| BGE-M3 卡顿 | 16GB 内存不足 | 继续用 nomic-embed-text |\n\n## 关键路径\n\n| 内容 | 路径 |\n|------|------|\n| Skill 目录 | `~/.openclaw/workspace/skills/knowledge-base-setup/` |\n| 知识库 | `~/.openclaw/workspace/knowledge/` |\n| 分析脚本 | `~/.openclaw/workspace/knowledge/.analysis/` |\n| 目录缓存 | `~/.openclaw/workspace/knowledge/.analysis/.catalog_cache.json` |\n| 摘要输出 | `~/.openclaw/workspace/knowledge/.analysis/summaries/` |\n| 文章目录 | `~/.openclaw/workspace/knowledge/文章目录/文章目录.md` |\n| OpenClaw 配置 | `~/.openclaw/openclaw.json` |\n\n\n## ⚠️ 临时文件处理（v1.4.3 修复）\n\nOCR 流程会把 PDF 复制到临时目录处理。\n\n**v1.4.3 之前**：使用 `/tmp/ocrmypdf_work`、`/tmp/office_ocr_work`、\n`/tmp/office_convert` 共享路径，存在以下风险：\n- 多进程并发可能冲突\n- 多用户系统下其他用户可访问（权限默认 755）\n- 处理失败时临时文件残留\n\n**v1.4.3 修复**：\n- 用 `tempfile.mkdtemp(prefix=\"...\")` 创建 per-run 私有目录（权限 0o700）\n- 处理完成后立即 `shutil.rmtree` 清理\n- `tempfile.mkstemp` 创建稳定输出文件（避免被 finally 误删）\n\n**剩余风险**：极端情况下（机器突然断电）可能残留临时目录。\n建议定期清理 `/Users/home/.openclaw/tmp/` 下 `ocrmypdf_*`、`office_*` 前缀目录。\n\n## 版本历史\n\n| 版本 | 日期 | 更新内容 |\n|------|------|---------|\n| 1.0.0 | 2026-05-10 | 初始版本，PyMuPDF + LibreOffice 链路 |\n| 1.1.0 | 2026-05-13 | 三步 PDF 处理，关键词库，双语标签 |\n| 1.2.0 | 2026-05-21 | 分批处理优化，280秒断点 |\n| 1.2.1 | 2026-05-22 | utils.py 共享模块重构，LibreOffice 熔断机制 |\n| 1.3.0 | 2026-05-28 | kreuzberg 统一提取层 + antiword 专线 + pandoc |\n| 1.4.0 | 2026-08-12 | CMap 残缺度自检 + 50万字完整提取 + OCR fallback 到 .doc |\n| **1.4.1** | **2026-08-12** | **run_analysis.py: sanitize_filename + Errno 63 重试 + 清理 temp_docs 畸形文件** |\n\nFile v1.4.6:_meta.json\n\n{\n  \"ownerId\": \"kn7d38vmg57htam960qsj7wnkh86ewjb\",\n  \"slug\": \"macmini-knowledge-base\",\n  \"version\": \"1.4.6\",\n  \"publishedAt\": 1786984987414\n}\n\nFile v1.4.6:CHANGELOG.md\n\n# Changelog\n\n## v1.4.6 (2026-08-17) — SKILL.md 修补\n\n- 修复 SKILL.md 顶部 version 字段（1.4.2 → 1.4.5）\n- 加\"结构化权限声明\"表格（Structured Permissions）\n- 让 Lp3 Finding 在 SKILL.md 层面满足 \"structured way\" 要求\n\n## v1.4.5 (2026-08-17) — NVIDIA SkillSpector 完整修复\n\n### 元数据修复\n- `_meta.json` 加 `permissions` 字段（**注：ClawHub 服务端会剥离非标准字段**）\n- 加 `warnings` 字段\n\n### SKILL.md 修复\n- description 强化\n- 顶部加独立章节\n\n### setup.sh 改造\n- 全局 y/N 确认\n- install_brew_package 函数化\n- 每个 cron 单独确认\n- 0 功能删除\n\n## v1.4.4 (2026-08-17) — 仅 SKILL.md 补充\n## v1.4.3 (2026-08-17) — 代码安全修复\n## v1.4.2 (2026-08-17) — 仅元数据更新\n## v1.4.1 (2026-08-12) — OSError 63 修复\n## v1.4.0 (2026-08-12) — OCR fallback + 50万字提取\n\nFile v1.4.6:skill-card.md\n\n## Description:\n\nSets up a local Mac Mini document knowledge base with RAG search, document extraction, scheduled analysis, and Feishu summary delivery.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[seairteng](https://clawhub.ai/user/seairteng)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers and individual Mac users use this skill to configure a local knowledge-base workflow for extracting documents, building summaries and catalogs, and enabling natural-language search. It is intended for personal Mac environments where automated cron jobs and Feishu notifications are acceptable.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: Persistent cron automation and broad execution authority can continue running document analysis and notification tasks after installation.\n\nMitigation: Install only on a personal Mac, review cron jobs before enabling them, and remove unwanted jobs with the documented OpenClaw cron removal command.\n\nRisk: Documents placed in the knowledge directory may be processed automatically and summarized for notification workflows.\n\nMitigation: Avoid placing shared or highly sensitive documents in the knowledge directory unless automatic processing and summaries are acceptable.\n\nRisk: A Feishu webhook secret may be stored in plaintext and could be reused if exposed.\n\nMitigation: Rotate the webhook if exposure is suspected and store it outside the workspace with restrictive permissions before enabling notifications.\n\n## Reference(s):\n\n- [ClawHub skill page](https://clawhub.ai/seairteng/skills/macmini-knowledge-base)\n- [Publisher profile](https://clawhub.ai/user/seairteng)\n- [Ollama download](https://ollama.com/download)\n\n## Skill Output:\n\n**Output Type(s):** [Guidance, Shell commands, Configuration, Code, Markdown, JSON]\n\n**Output Format:** [Markdown guidance with inline shell commands, configuration snippets, Python scripts, Markdown catalogs, and JSON reports]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [Produces setup steps, local automation scripts, cron configuration, document summaries, catalog Markdown, and OCR repair reports.]\n\n## Skill Version(s):\n\n1.4.6 (source: server release evidence and CHANGELOG, released 2026-08-17)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nArchive v1.4.5: 9 files, 29134 bytes\n\nFiles: _meta.json (141b), CHANGELOG.md (1677b), scripts/generate_catalog.py (12426b), scripts/re_ocr_corrupted.py (7276b), scripts/run_analysis.py (8164b), scripts/setup.sh (8552b), scripts/utils.py (17289b), skill-card.md (2696b), SKILL.md (19620b)\n\nFile v1.4.5:SKILL.md\n\n---\nname: macmini-knowledge-base\nversion: 1.4.2\ndescription: |\n  ⚠️ **本 skill 包含以下高危能力，使用前请仔细阅读 SKILL.md 顶部「⚠️ CAPABILITIES & RISKS」章节**：\n  - Shell 执行 + 文件读写 + 安装 Homebrew 包（版本固定）+ 下载 Ollama 模型\n  - 修改 OpenClaw 配置 + 注册持久 cron 任务 + 飞书 webhook 推送\n\n  在 Mac Mini (M4) 上快速搭建本地知识库 + RAG 自然语言搜索系统。\n  适用场景：\n  - 新 Mac 配置知识库：从零开始安装配置 Ollama、embedding模型、定时任务、文档解析\n  - 遇到 PDF 提取乱码、定时任务超时、skill 加载失败等问题\n  - 想要建立每日自动分析文档 + 08:00发送摘要到飞书的流程\n  - 迁移或复现知识库：打包整个 knowledge 目录和配置到新电脑\n  - **v1.4 新增**：CMap 残缺度自检（不预设来源）+ 50万字完整提取 + OCR fallback 到 .doc\n  本 skill 会引导完成：目录结构创建、依赖安装、脚本部署、定时任务注册、OpenClaw 配置。\n\n  ⚠️ **重要：能力范围**\n  本 skill 不只是「搭建」，还包含：\n  - 批量 OCR 修复（扫描 summaries/archives 找乱码 + 重新提取）\n  - 目录归档清理（移动重复/孤儿文件到 .trash/）\n  - 自动定时任务（23:00 分析 + 06:00 飞书推送）\n  \n  使用前请仔细评估批量修改风险。\n---\n\n# Knowledge Base Setup\n\n在 Mac Mini 上快速搭建本地知识库 + RAG 搜索系统。\n\n## ⚠️ 阅读前必读：本 skill 的能力范围\n\n本 skill **不只是\"搭建知识库\"**，还包含以下高危能力：\n\n**执行能力**：\n- 🔧 **Shell 命令执行**（python3 + bash 脚本）\n- 📁 **文件读写**（knowledge/, summaries/, archives/, .trash/）\n- 📦 **安装 Homebrew 包**（antiword, tesseract, pandoc, libreoffice，**版本固定**）\n- ⬇️ **下载 Ollama 模型**（nomic-embed-text, ~274MB）\n- ⚙️ **修改 OpenClaw 配置**（~/.openclaw/openclaw.json）\n- ⏰ **注册持久 cron 任务**（23:00 + 06:00，**每天自动**）\n- 📤 **推送消息到飞书 webhook**\n\n**持久化影响**：\n- 知识库目录会被自动分析（每天 23:00）\n- 摘要文件会被覆盖写入（OCR 修复时）\n- cron 任务永久执行（直到手动 `openclaw cron remove <id>`）\n- 失败文件移到 `.trash/`（7 天兜底清理）\n\n**安装流程**：\n本 skill 的 setup.sh 是**交互式安装向导**：\n- 每个危险操作前会要求 y/N 确认\n- 提供 `--dry-run` 选项查看会做什么\n- 已安装用户重跑会进入确认模式\n\n**如果不同意上述任何一项，请不要安装本 skill。**\n\n---\n\n## 核心功能（v2.0）\n\n- **kreuzberg 统一提取层**：PDF / DOCX / XLSX / PPTX / MD / 图片 OCR 全自动路由\n- **antiword 极速专线**：.doc 文件专用提取，成功率 85%，169MB 文件 0.02 秒完成\n- **智能兜底**：antiword 失败自动走 soffice 转换，60 秒硬超时无误判\n- **自动分类**：关键词匹配驱动，中英文双语标签\n- **定时任务**：每天 23:00 分析新文档，08:00 发送摘要到飞书\n\n\n\n## 🔧 权限声明\n\n本 skill 在使用时需要以下 OpenClaw 工具能力：\n\n```json\n{\n  \"tools\": {\n    \"alsoAllow\": [\"exec\", \"process\", \"read\", \"write\"]\n  }\n}\n```\n\n⚠️ **执行风险**：exec + cron 自动化组合可导致持续命令执行，\n请在可信环境（个人 Mac）使用，不要在共享/服务器部署。\n\n\n\n## ⚠️ 安全警告：定时任务\n\n本 skill 注册 2 个 cron 任务（23:00 分析 + 06:00 推送），\n运行 shell 命令并自动推送消息到飞书。\n\n**潜在风险**：\n- 脚本路径被修改 → 自动执行任意命令\n- 知识库目录被入侵 → 自动读取/外发\n- 飞书 webhook 泄漏 → 自动推送被劫持\n\n**建议**：\n- 不要把 `~/.openclaw/workspace/knowledge` 放在共享/多用户目录\n- 定期检查 cron 配置（`openclaw cron list`）\n- 飞书 webhook 使用独立群组，不要复用其他机器人的 webhook\n- 仅在个人 Mac 上运行，不要部署到服务器\n\n## 快速开始\n\n### 一键安装\n\n```bash\ncd ~/.openclaw/workspace/skills/knowledge-base-setup/scripts\nbash setup.sh <飞书用户ID>\n```\n\n### 手动分步安装\n\n**Step 1: 系统依赖**\n```bash\nbrew install antiword tesseract pandoc\n```\n\n**Step 2: Python 依赖**\n```bash\npip3 install kreuzberg pytesseract pymupdf docx openpyxl python-pptx\n```\n\n**Step 3: Ollama + embedding 模型**\n```bash\n# 安装 Ollama: https://ollama.com/download\nollama pull nomic-embed-text\n```\n\n**Step 4: 创建目录结构**\n```bash\nmkdir -p ~/.openclaw/workspace/knowledge/.analysis/summaries/archives\nmkdir -p ~/.openclaw/workspace/knowledge/temp_docs\ntouch ~/.openclaw/workspace/knowledge/文章目录/文章目录.md\n```\n\n**Step 5: 部署脚本**\n```bash\ncp ~/.openclaw/workspace/skills/knowledge-base-setup/scripts/*.py \\\n   ~/.openclaw/workspace/knowledge/.analysis/\nchmod +x ~/.openclaw/workspace/knowledge/.analysis/*.py\n```\n\n**Step 6: 配置 OpenClaw**\n\n编辑 `~/.openclaw/openclaw.json`，加入：\n```json\n{\n  \"models\": {\n    \"providers\": {\n      \"ollama\": {\n        \"baseUrl\": \"http://127.0.0.1:11434\",\n        \"api\": \"ollama\",\n        \"models\": [\n          {\"id\": \"nomic-embed-text\", \"name\": \"Nomic Embed Text\"}\n        ]\n      }\n    }\n  },\n  \"agents\": {\n    \"defaults\": {\n      \"memorySearch\": {\n        \"provider\": \"ollama\",\n        \"model\": \"nomic-embed-text\"\n      }\n    }\n  }\n}\n```\n\n确保 tools 区块有：\n```json\n\"tools\": {\n    \"alsoAllow\": [\"exec\", \"process\"]\n}\n```\n\n然后重启：`openclaw gateway restart`\n\n**Step 7: 注册定时任务**\n```bash\n# 23:00 分析新文档\nopenclaw cron add \\\n  --name \"23:00分析新文档\" \\\n  --cron \"0 23 * * *\" \\\n  --tz \"Asia/Shanghai\" \\\n  --session isolated \\\n  --timeout-seconds 600 \\\n  --message \"cd ~/.openclaw/workspace/knowledge/.analysis && python3 run_analysis.py && python3 generate_catalog.py\" \\\n  --announce --channel feishu --to \"user:<飞书用户ID>\"\n\n# 08:00 发送文档摘要\nopenclaw cron add \\\n  --name \"08:00发送文档摘要\" \\\n  --cron \"0 8 * * *\" \\\n  --tz \"Asia/Shanghai\" \\\n  --session isolated \\\n  --timeout-seconds 120 \\\n  --message \"读取 summaries/ 目录发送摘要到飞书\" \\\n  --announce --channel feishu --to \"user:<飞书用户ID>\"\n```\n\n\n\n\n\n## summary 文件名 sanitize（v1.4.1 新增）\n\n防止 `OSError: [Errno 63] File name too long`（NAME_MAX=255 bytes）：\n\n```python\nSUMMARY_NAME_MAX = 200\n\ndef sanitize_filename(name, max_length=SUMMARY_NAME_MAX):\n    \"\"\"截断超长文件名，保留扩展名 + 8 位 MD5 hash 防冲突\"\"\"\n    name_bytes = name.encode('utf-8')\n    if len(name_bytes) <= max_length:\n        return name\n    \n    base, ext = os.path.splitext(name)\n    ext_bytes = ext.encode('utf-8')\n    base_bytes = base.encode('utf-8')\n    \n    import hashlib\n    h = hashlib.md5(name_bytes).hexdigest()[:8]\n    \n    reserve = len(ext_bytes) + 1 + 8  # \"_\" + hash + ext\n    available = max_length - reserve\n    \n    if available > 0 and len(base_bytes) > available:\n        truncated = base_bytes[:available].decode('utf-8', errors='ignore')\n        return f\"{truncated}_{h}{ext}\"\n    \n    return name[:max_length]\n```\n\n主循环的异常捕获重试：\n\n```python\ntry:\n    with open(summary_file, 'w', encoding='utf-8') as f:\n        f.write(content)\nexcept OSError as e:\n    if e.errno == 63:  # ENAMETOOLONG\n        short_name = sanitize_filename(filename, max_length=180)\n        summary_file = os.path.join(\n            SUMMARY_DIR,\n            f\"{timestamp}_{short_name}.summary.txt\"\n        )\n        with open(summary_file, 'w', encoding='utf-8') as f:\n            f.write(content)\n```\n\n**触发场景：** 畸形 PDF 文件名（如下载错误的 `_; filename_=utf-8''...` 双名拼接），原文件名 244+ bytes + 时间戳超 255 bytes 限制。\n\n**实测案例：** `20260730-Nomura-Asia Insights：China：The Politburo meeting indicated a shift to _countercyclical\" policies-260730.pdf_; filename_=utf-8''...pdf` (原 248 bytes) → sanitize 后 200 bytes + 8位 hash → 安全创建。\n\n### temp_docs 畸形文件清理（v1.4.1 新增）\n\n下载失败的 PDF 在文件名里重复了两次（`_; filename_=utf-8''` 分隔），实际只需保留前半。一次性清理脚本：\n\n```python\nimport os, shutil\ntemp_docs = os.path.expanduser(\"~/.openclaw/workspace/knowledge/temp_docs\")\ntrash_dir = os.path.expanduser(\"~/.openclaw/workspace/knowledge/.trash/temp_docs_<时间戳>\")\nos.makedirs(trash_dir, exist_ok=True)\n\nfor f in os.listdir(temp_docs):\n    if \"_; filename_=utf-8''\" in f:\n        full = os.path.join(temp_docs, f)\n        parts = f.split(\"_; filename_=utf-8''\")\n        real_name = parts[0]\n        target = os.path.join(temp_docs, real_name)\n        if not os.path.exists(target):\n            shutil.move(full, target)\n            print(f\"重命名: {real_name}\")\n```\n\n\n\n## CMap 残缺度自检（v1.4 新增）\n\n不预设\"哪个 PDF 来源会乱码\"——实测 **72% 的乱码来自非 lightpdf PDF**（PPT 转 PDF、扫描件等），\n改用**自适应检测**：\n\n```python\ndef is_cmap_broken(text, threshold=0.03):\n    \"\"\"检测文本是否含异常字符（CMap 残缺/PUA 污染/未映射 CID）\"\"\"\n    if not text or len(text.strip()) < 50:\n        return False\n    total = len(text)\n    pua_count = sum(1 for c in text if 0xE000 <= ord(c) <= 0xF8FF)\n    cjk_ext = sum(1 for c in text if 0x20000 <= ord(c) <= 0x2EBEF)\n    cjk_compat = sum(1 for c in text if 0xF900 <= ord(c) <= 0xFAFF)\n    cid_count = text.count('(cid:')\n    bad_ratio = (pua_count + cjk_ext + cjk_compat + cid_count) / total\n    return bad_ratio > threshold or cid_count > 10\n```\n\n**3 类乱码特征：**\n1. **PUA 私用区** (U+E000-F8FF) —— 残缺 CMap fallback\n2. **CJK 扩展区** (U+20000-2EBEF) —— 字符找不到映射\n3. **`(cid:xxxx)` 字面值** —— pdfplumber 提取失败标志\n\n**集成位置：** `extract_pdf_text()` 在 kreuzberg / pymupdf 提取后调 `is_cmap_broken()`，\n通过即返回，失败即触发 OCR 路径。\n\n### OCR 性能实测（2026-08-12 验证）\n\n| 文件 | 大小 | OCR 耗时 | 备注 |\n|---|---|---|---|\n| lightpdf PDF | 4 页 | 13.5 秒 | CMap 残缺，自动 OCR |\n| 大型 PPT 转 PDF      | 90 页 | 0.5 秒 | 默认路径（无需 OCR）|\n| 65MB .doc | 169MB 文件 | 0.1 秒 | antiword 极速专线 |\n| 大型 docx（475K 字）| 562KB | 11.2 秒 | python-docx fallback |\n| OCR 自检总开销 | - | < 200ms | 3 页抽样 + 字符统计 |\n\n### 批量 OCR 修复脚本（v1.4 新增）\n\n`re_ocr_corrupted.py` —— 批量扫描乱码 summary，自动用新版本 utils 重新提取：\n\n```bash\n# 干跑（不写文件）\npython3 re_ocr_corrupted.py --dry-run --max 10\n\n# 实际批量（处理所有乱码）\npython3 re_ocr_corrupted.py --max 100\n\n# 只处理指定 PDF\npython3 re_ocr_corrupted.py --pdf-list \"path1.pdf,path2.pdf\"\n```\n\n行为：\n1. 扫 archives/ 找出乱码 summary\n2. 按 basename 匹配源 PDF\n3. 调 `extract_pdf_text()` 重跑（自动 OCR fallback）\n4. 写新 summary 到 summaries/（带新时间戳）\n5. 覆盖 archives/ 里对应 basename 的所有乱码版本\n6. 输出 JSON 报告（含每份文件路径/字数/成功状态）\n\n## 文档解析架构（v2.0）\n\n### 架构图\n\n```\n                    ┌──────────────────────────────────────┐\n                    │         kreuzberg 统一提取层           │\n                    │  (pypdfium2 / python-calamine / pandoc) │\n                    └───┬────────────────────────────────┬───┘\n                        │                              │\n                自动判断 │                              │\n                        ▼                              ▼\n              ┌─────────────────┐           ┌─────────────────────┐\n              │  kreuzberg 直提  │           │  antiword 极速专线  │\n              │ PDF/DOCX/XLSX/  │           │   (.doc 文件专用)    │\n              │ PPTX/MD/图片OCR │           │   成功率 85%，<1秒   │\n              └─────────────────┘           └─────────────────────┘\n                        │                              │\n                        │         ┌──────────────────────────────┐\n                        │         │     soffice 兜底转换          │\n                        │         │ (.doc/.xls/.ppt antiword失败) │\n                        │         │  60秒硬超时（消除误判watchdog）│\n                        │         └──────────────────────────────┘\n                        ▼                              │\n              ┌──────────────────────────────────────────────┐\n              │              文本输出（content）              │\n              │  → summaries/ 摘要文件 → generate_catalog.py  │\n              └──────────────────────────────────────────────┘\n```\n\n### 文件类型 × 提取方式\n\n| 格式 | 主方案 | 依赖 | 成功率 | 单文件速度 |\n|------|--------|------|--------|-----------|\n| PDF | kreuzberg (pypdfium2) | 无 | ~100% | 0.05-0.7s |\n| DOCX | kreuzberg + pandoc | pandoc 3.9+ | 100% | 0.12-3s |\n| XLSX | kreuzberg (python-calamine) | 无 | 100% | 0.1-0.5s |\n| PPTX | kreuzberg + pandoc | pandoc 3.9+ | 100% | 0.02-0.2s |\n| MD | kreuzberg + pandoc | pandoc 3.9+ | 100% | <0.01s |\n| **.doc** | **antiword 优先** | antiword | **85%**，<1秒 | <0.02s |\n| .doc（失败） | soffice 兜底 | LibreOffice | ~15% | 2-21s |\n| .xls | soffice → XLSX | LibreOffice | ~95% | 2-10s |\n| .ppt | soffice → PPTX | LibreOffice | ~95% | 2-10s |\n| 图片 | kreuzberg 内置 OCR | tesseract | ~90% | 3-10s |\n\n### antiword 极速专线\n\n```python\n# 实测数据：\n# 169MB 超大文件 → 26万字符，0.02秒完成\n# 正常 .doc（0.1-15MB）→ <1秒\n# 成功率 85%，覆盖绝大多数 .doc 文件\nresult = subprocess.run(['antiword', filepath], capture_output=True, timeout=10)\n```\n\n### kreuzberg 统一提取层\n\nkreuzberg 是专业的非结构化文档文本提取库（支持 20+ 格式），内部自动路由：\n- PDF → pypdfium2\n- XLSX → python-calamine\n- DOCX/PPTX/MD → pandoc\n- 图片 → 内置 OCR（tesseract）\n\n## 关键词库（中英双语）\n\n**中文（47个）：** 房产、房价、房地产、居民、消费、股市、经济、政策、利率、通胀、人民币、A股、美联储、PBOC、GDP、股票、资产、投资、债券、银行、PPI、CPI、PMI、M2、就业、失业、汽车、新能源、AI 等\n\n**英文（70+个）：** property、real estate、GDP、inflation、CPI、PPI、PMI、PBOC、Fed、consumer、economy、growth、housing、stock market、EV、AI 等\n\n**标签输出语言：** 自动判断——英文内容匹配英文关键词输出英文标签，中文内容匹配中文关键词输出中文标签\n\n## 定时任务兼容性\n\n| 任务 | ID | 调用方式 | 结论 |\n|------|------|---------|------|\n| 23:00分析新文档 | f3536e18 | 绝对路径 `python3 run_analysis.py` | ✅ 无需修改 |\n| 07:00生成财经早报 | b741c6d5 | Node.js 脚本 | ❌ 不相关 |\n| 08:00发送财经早报 | a7cbaacc | 读取文件发送 | ❌ 不相关 |\n| 09:00发送文档摘要 | 89b4cf75 | 读取 summaries 目录 | ❌ 不相关 |\n\n## 迁移到新电脑\n\n⚠️ **迁移前必读**：`~/.openclaw/workspace/knowledge/` 目录可能包含：\n- 私人合同/财务文档的 OCR 摘要\n- 个人分析报告\n- 飞书推送缓存\n\n**建议**：\n1. 先 `du -sh ~/.openclaw/workspace/knowledge/` 看大小\n2. 排除 `.trash/`、`.analysis/cache/` 后再迁移\n3. 用 `rsync -av --exclude='.trash' ...` 而不是 `scp -r`\n\n1. 复制目录（推荐 rsync）：\n   ```bash\n   rsync -av --exclude='.trash' --exclude='.analysis/cache' \\\n     ~/.openclaw/workspace/knowledge user@new-mac:~/.openclaw/workspace/\n   ```\n2. 在新电脑运行 `bash setup.sh <飞书用户ID>`\n3. 重新注册定时任务（Job ID 会变）\n\n## 避坑指南\n\n| 问题 | 原因 | 解决 |\n|------|------|------|\n| LibreOffice 超时 | watchdog 误判大文件为卡死 | v2.0 移除 watchdog，60秒硬超时 |\n| .doc 提取慢 | 统一走 LibreOffice | antiword 专线，169MB 文件 0.02秒 |\n| DOCX/PPTX 处理失败 | pandoc 未安装 | `brew install pandoc` |\n| PDF 提取乱码（OCR 不工作） | ocrmypdf `--skip-text` 跳过乱码页 | v1.4 改为 `--force-ocr` 强制 OCR |\n| PDF 漏检 CMap 残缺 | 没主动判断是否乱码 | v1.4 `is_cmap_broken()` 自检（阈值 0.03）|\n| 文本被截断到 8000 字 | 硬编码 `[:8000]` 太短 | v1.4 `MAX_EXTRACT_LEN = 500_000` |\n| .doc 提取失败 | lightpdf 处理过的 .doc 乱码 | v1.4 `ocr_office_via_ocr()` 兜底 |\n| summary 文件名过长失败 | 畸形 PDF 名 244+ bytes + 时间戳超 NAME_MAX | v1.4.1 `sanitize_filename()` + Errno 63 重试 |\n| 静默失败（不知道哪个文件）| 不抛异常 | v1.4 `PDFExtractError` / `ExtractError` 含路径 |\n| 飞书无 exec 工具 | tools 策略限制 | 添加 `alsoAllow: [exec, process]` |\n| BGE-M3 卡顿 | 16GB 内存不足 | 继续用 nomic-embed-text |\n\n## 关键路径\n\n| 内容 | 路径 |\n|------|------|\n| Skill 目录 | `~/.openclaw/workspace/skills/knowledge-base-setup/` |\n| 知识库 | `~/.openclaw/workspace/knowledge/` |\n| 分析脚本 | `~/.openclaw/workspace/knowledge/.analysis/` |\n| 目录缓存 | `~/.openclaw/workspace/knowledge/.analysis/.catalog_cache.json` |\n| 摘要输出 | `~/.openclaw/workspace/knowledge/.analysis/summaries/` |\n| 文章目录 | `~/.openclaw/workspace/knowledge/文章目录/文章目录.md` |\n| OpenClaw 配置 | `~/.openclaw/openclaw.json` |\n\n\n## ⚠️ 临时文件处理（v1.4.3 修复）\n\nOCR 流程会把 PDF 复制到临时目录处理。\n\n**v1.4.3 之前**：使用 `/tmp/ocrmypdf_work`、`/tmp/office_ocr_work`、\n`/tmp/office_convert` 共享路径，存在以下风险：\n- 多进程并发可能冲突\n- 多用户系统下其他用户可访问（权限默认 755）\n- 处理失败时临时文件残留\n\n**v1.4.3 修复**：\n- 用 `tempfile.mkdtemp(prefix=\"...\")` 创建 per-run 私有目录（权限 0o700）\n- 处理完成后立即 `shutil.rmtree` 清理\n- `tempfile.mkstemp` 创建稳定输出文件（避免被 finally 误删）\n\n**剩余风险**：极端情况下（机器突然断电）可能残留临时目录。\n建议定期清理 `/Users/home/.openclaw/tmp/` 下 `ocrmypdf_*`、`office_*` 前缀目录。\n\n## 版本历史\n\n| 版本 | 日期 | 更新内容 |\n|------|------|---------|\n| 1.0.0 | 2026-05-10 | 初始版本，PyMuPDF + LibreOffice 链路 |\n| 1.1.0 | 2026-05-13 | 三步 PDF 处理，关键词库，双语标签 |\n| 1.2.0 | 2026-05-21 | 分批处理优化，280秒断点 |\n| 1.2.1 | 2026-05-22 | utils.py 共享模块重构，LibreOffice 熔断机制 |\n| 1.3.0 | 2026-05-28 | kreuzberg 统一提取层 + antiword 专线 + pandoc |\n| 1.4.0 | 2026-08-12 | CMap 残缺度自检 + 50万字完整提取 + OCR fallback 到 .doc |\n| **1.4.1** | **2026-08-12** | **run_analysis.py: sanitize_filename + Errno 63 重试 + 清理 temp_docs 畸形文件** |\n\nFile v1.4.5:_meta.json\n\n{\n  \"ownerId\": \"kn7d38vmg57htam960qsj7wnkh86ewjb\",\n  \"slug\": \"macmini-knowledge-base\",\n  \"version\": \"1.4.5\",\n  \"publishedAt\": 1786984769912\n}\n\nFile v1.4.5:CHANGELOG.md\n\n# Changelog\n\n## v1.4.5 (2026-08-17) — NVIDIA SkillSpector 完整修复\n\n### 元数据修复（修 Finding 1: Lp3）\n- `_meta.json` 加结构化 `permissions` 字段\n  - exec / file_read / file_write / install_packages / download_model\n  - modify_config / register_cron / network\n- 加 `warnings` 字段（5 条用户警告）\n\n### SKILL.md 修复（修 Finding 2: Tp4）\n- description 强化：开头加 ⚠️ 高危能力声明\n- 顶部加独立 `## ⚠️ 阅读前必读：本 skill 的能力范围` 章节\n- 列出 7 项执行能力 + 4 项持久化影响 + 安装流程说明\n\n### setup.sh 改造（修 Finding 3-6: mkdir/brew/cron/session）\n- 加全局 `y/N` 确认（\"确认开始安装？\"）\n- 加 `--dry-run` 模式（只显示不执行）\n- 加 `install_brew_package()` 函数（每个包单独 confirm）\n- 加 Python 包单独 confirm\n- 加 Ollama 模型单独 confirm\n- 加 OpenClaw 配置修改 confirm\n- 加每个 cron 任务单独 confirm\n- 加飞书 webhook 配置 confirm\n- 0 功能删除（mkdir/brew/pip/ollama/cron 全部保留）\n\n### 验证\n- ✅ 10/10 封闭测试通过\n- ✅ ClawHub publish --dry-run 接受新 _meta.json\n- ✅ bash -n 语法检查通过\n- ✅ setup.sh 三种退出路径全对（dry-run / 全局 N / 全局 y）\n\n## v1.4.4 (2026-08-17) — 仅 SKILL.md 补充\n- 加「## ⚠️ 临时文件处理（v1.4.3 修复）」章节\n\n## v1.4.3 (2026-08-17) — 代码安全修复\n- os.system → subprocess.run\n- /tmp 共享目录 → tempfile.mkdtemp\n- 5 项文档警告（基础版）\n\n## v1.4.2 (2026-08-17) — 仅元数据更新\n## v1.4.1 (2026-08-12) — OSError 63 修复\n## v1.4.0 (2026-08-12) — OCR fallback + 50万字提取\n\nFile v1.4.5:skill-card.md\n\n## Description:\n\nSets up a local Mac Mini knowledge base with RAG search, document extraction, scheduled analysis, OCR repair, catalog generation, and Feishu summary delivery.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[seairteng](https://clawhub.ai/user/seairteng)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nExternal users and developers with a trusted personal Mac use this skill to build and operate a local document knowledge base. It guides installation, OpenClaw configuration, scheduled document analysis, summary generation, OCR repair, and Feishu delivery.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: Persistent scheduled jobs run local analysis and Feishu delivery after setup.\n\nMitigation: Review each cron registration during the interactive setup, periodically inspect the cron list, and remove unwanted jobs with the documented openclaw cron remove command.\n\nRisk: The skill reads and writes under the local knowledge workspace, including summaries, archives, and a trash folder.\n\nMitigation: Use it only with a restricted personal knowledge folder containing documents you are comfortable indexing and summarizing.\n\nRisk: The Feishu webhook URL is stored as a local plaintext secret.\n\nMitigation: Use a dedicated webhook, avoid shared or broadly backed-up workspaces, and rotate the webhook if exposure is possible.\n\nRisk: Setup can install packages, download an embedding model, and modify OpenClaw tool permissions.\n\nMitigation: Run setup.sh with --dry-run first and accept only the dependency, model, and configuration prompts that match the intended deployment.\n\n## Reference(s):\n\n- [ClawHub Skill Page](https://clawhub.ai/seairteng/skills/macmini-knowledge-base)\n- [Publisher Profile](https://clawhub.ai/user/seairteng)\n- [Ollama Download](https://ollama.com/download)\n\n## Skill Output:\n\n**Output Type(s):** [text, markdown, code, shell commands, configuration, guidance]\n\n**Output Format:** [Markdown guidance with inline shell commands and configuration snippets; generated summaries and catalogs are text or Markdown files.]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [Runs local scripts that create summaries, catalog entries, JSON state/report files, cron jobs, and OpenClaw configuration changes.]\n\n## Skill Version(s):\n\n1.4.5 (source: release evidence and CHANGELOG; SKILL.md frontmatter lists 1.4.2)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nArchive v1.4.4: 9 files, 27426 bytes\n\nFiles: _meta.json (141b), CHANGELOG.md (1210b), scripts/generate_catalog.py (12426b), scripts/re_ocr_corrupted.py (7276b), scripts/run_analysis.py (8164b), scripts/setup.sh (3966b), scripts/utils.py (17289b), skill-card.md (3089b), SKILL.md (18152b)\n\nFile v1.4.4:SKILL.md\n\n---\nname: macmini-knowledge-base\nversion: 1.4.2\ndescription: |\n  在 Mac Mini (M4) 上快速搭建本地知识库 + RAG 自然语言搜索系统。\n  适用场景：\n  - 新 Mac 配置知识库：从零开始安装配置 Ollama、embedding模型、定时任务、文档解析\n  - 遇到 PDF 提取乱码、定时任务超时、skill 加载失败等问题\n  - 想要建立每日自动分析文档 + 08:00发送摘要到飞书的流程\n  - 迁移或复现知识库：打包整个 knowledge 目录和配置到新电脑\n  - **v1.4 新增**：CMap 残缺度自检（不预设来源）+ 50万字完整提取 + OCR fallback 到 .doc\n  本 skill 会引导完成：目录结构创建、依赖安装、脚本部署、定时任务注册、OpenClaw 配置。\n\n  ⚠️ **重要：能力范围**\n  本 skill 不只是「搭建」，还包含：\n  - 批量 OCR 修复（扫描 summaries/archives 找乱码 + 重新提取）\n  - 目录归档清理（移动重复/孤儿文件到 .trash/）\n  - 自动定时任务（23:00 分析 + 06:00 飞书推送）\n  \n  使用前请仔细评估批量修改风险。\n---\n\n# Knowledge Base Setup\n\n在 Mac Mini 上快速搭建本地知识库 + RAG 搜索系统。\n\n## 核心功能（v2.0）\n\n- **kreuzberg 统一提取层**：PDF / DOCX / XLSX / PPTX / MD / 图片 OCR 全自动路由\n- **antiword 极速专线**：.doc 文件专用提取，成功率 85%，169MB 文件 0.02 秒完成\n- **智能兜底**：antiword 失败自动走 soffice 转换，60 秒硬超时无误判\n- **自动分类**：关键词匹配驱动，中英文双语标签\n- **定时任务**：每天 23:00 分析新文档，08:00 发送摘要到飞书\n\n\n\n## 🔧 权限声明\n\n本 skill 在使用时需要以下 OpenClaw 工具能力：\n\n```json\n{\n  \"tools\": {\n    \"alsoAllow\": [\"exec\", \"process\", \"read\", \"write\"]\n  }\n}\n```\n\n⚠️ **执行风险**：exec + cron 自动化组合可导致持续命令执行，\n请在可信环境（个人 Mac）使用，不要在共享/服务器部署。\n\n\n\n## ⚠️ 安全警告：定时任务\n\n本 skill 注册 2 个 cron 任务（23:00 分析 + 06:00 推送），\n运行 shell 命令并自动推送消息到飞书。\n\n**潜在风险**：\n- 脚本路径被修改 → 自动执行任意命令\n- 知识库目录被入侵 → 自动读取/外发\n- 飞书 webhook 泄漏 → 自动推送被劫持\n\n**建议**：\n- 不要把 `~/.openclaw/workspace/knowledge` 放在共享/多用户目录\n- 定期检查 cron 配置（`openclaw cron list`）\n- 飞书 webhook 使用独立群组，不要复用其他机器人的 webhook\n- 仅在个人 Mac 上运行，不要部署到服务器\n\n## 快速开始\n\n### 一键安装\n\n```bash\ncd ~/.openclaw/workspace/skills/knowledge-base-setup/scripts\nbash setup.sh <飞书用户ID>\n```\n\n### 手动分步安装\n\n**Step 1: 系统依赖**\n```bash\nbrew install antiword tesseract pandoc\n```\n\n**Step 2: Python 依赖**\n```bash\npip3 install kreuzberg pytesseract pymupdf docx openpyxl python-pptx\n```\n\n**Step 3: Ollama + embedding 模型**\n```bash\n# 安装 Ollama: https://ollama.com/download\nollama pull nomic-embed-text\n```\n\n**Step 4: 创建目录结构**\n```bash\nmkdir -p ~/.openclaw/workspace/knowledge/.analysis/summaries/archives\nmkdir -p ~/.openclaw/workspace/knowledge/temp_docs\ntouch ~/.openclaw/workspace/knowledge/文章目录/文章目录.md\n```\n\n**Step 5: 部署脚本**\n```bash\ncp ~/.openclaw/workspace/skills/knowledge-base-setup/scripts/*.py \\\n   ~/.openclaw/workspace/knowledge/.analysis/\nchmod +x ~/.openclaw/workspace/knowledge/.analysis/*.py\n```\n\n**Step 6: 配置 OpenClaw**\n\n编辑 `~/.openclaw/openclaw.json`，加入：\n```json\n{\n  \"models\": {\n    \"providers\": {\n      \"ollama\": {\n        \"baseUrl\": \"http://127.0.0.1:11434\",\n        \"api\": \"ollama\",\n        \"models\": [\n          {\"id\": \"nomic-embed-text\", \"name\": \"Nomic Embed Text\"}\n        ]\n      }\n    }\n  },\n  \"agents\": {\n    \"defaults\": {\n      \"memorySearch\": {\n        \"provider\": \"ollama\",\n        \"model\": \"nomic-embed-text\"\n      }\n    }\n  }\n}\n```\n\n确保 tools 区块有：\n```json\n\"tools\": {\n    \"alsoAllow\": [\"exec\", \"process\"]\n}\n```\n\n然后重启：`openclaw gateway restart`\n\n**Step 7: 注册定时任务**\n```bash\n# 23:00 分析新文档\nopenclaw cron add \\\n  --name \"23:00分析新文档\" \\\n  --cron \"0 23 * * *\" \\\n  --tz \"Asia/Shanghai\" \\\n  --session isolated \\\n  --timeout-seconds 600 \\\n  --message \"cd ~/.openclaw/workspace/knowledge/.analysis && python3 run_analysis.py && python3 generate_catalog.py\" \\\n  --announce --channel feishu --to \"user:<飞书用户ID>\"\n\n# 08:00 发送文档摘要\nopenclaw cron add \\\n  --name \"08:00发送文档摘要\" \\\n  --cron \"0 8 * * *\" \\\n  --tz \"Asia/Shanghai\" \\\n  --session isolated \\\n  --timeout-seconds 120 \\\n  --message \"读取 summaries/ 目录发送摘要到飞书\" \\\n  --announce --channel feishu --to \"user:<飞书用户ID>\"\n```\n\n\n\n\n\n## summary 文件名 sanitize（v1.4.1 新增）\n\n防止 `OSError: [Errno 63] File name too long`（NAME_MAX=255 bytes）：\n\n```python\nSUMMARY_NAME_MAX = 200\n\ndef sanitize_filename(name, max_length=SUMMARY_NAME_MAX):\n    \"\"\"截断超长文件名，保留扩展名 + 8 位 MD5 hash 防冲突\"\"\"\n    name_bytes = name.encode('utf-8')\n    if len(name_bytes) <= max_length:\n        return name\n    \n    base, ext = os.path.splitext(name)\n    ext_bytes = ext.encode('utf-8')\n    base_bytes = base.encode('utf-8')\n    \n    import hashlib\n    h = hashlib.md5(name_bytes).hexdigest()[:8]\n    \n    reserve = len(ext_bytes) + 1 + 8  # \"_\" + hash + ext\n    available = max_length - reserve\n    \n    if available > 0 and len(base_bytes) > available:\n        truncated = base_bytes[:available].decode('utf-8', errors='ignore')\n        return f\"{truncated}_{h}{ext}\"\n    \n    return name[:max_length]\n```\n\n主循环的异常捕获重试：\n\n```python\ntry:\n    with open(summary_file, 'w', encoding='utf-8') as f:\n        f.write(content)\nexcept OSError as e:\n    if e.errno == 63:  # ENAMETOOLONG\n        short_name = sanitize_filename(filename, max_length=180)\n        summary_file = os.path.join(\n            SUMMARY_DIR,\n            f\"{timestamp}_{short_name}.summary.txt\"\n        )\n        with open(summary_file, 'w', encoding='utf-8') as f:\n            f.write(content)\n```\n\n**触发场景：** 畸形 PDF 文件名（如下载错误的 `_; filename_=utf-8''...` 双名拼接），原文件名 244+ bytes + 时间戳超 255 bytes 限制。\n\n**实测案例：** `20260730-Nomura-Asia Insights：China：The Politburo meeting indicated a shift to _countercyclical\" policies-260730.pdf_; filename_=utf-8''...pdf` (原 248 bytes) → sanitize 后 200 bytes + 8位 hash → 安全创建。\n\n### temp_docs 畸形文件清理（v1.4.1 新增）\n\n下载失败的 PDF 在文件名里重复了两次（`_; filename_=utf-8''` 分隔），实际只需保留前半。一次性清理脚本：\n\n```python\nimport os, shutil\ntemp_docs = os.path.expanduser(\"~/.openclaw/workspace/knowledge/temp_docs\")\ntrash_dir = os.path.expanduser(\"~/.openclaw/workspace/knowledge/.trash/temp_docs_<时间戳>\")\nos.makedirs(trash_dir, exist_ok=True)\n\nfor f in os.listdir(temp_docs):\n    if \"_; filename_=utf-8''\" in f:\n        full = os.path.join(temp_docs, f)\n        parts = f.split(\"_; filename_=utf-8''\")\n        real_name = parts[0]\n        target = os.path.join(temp_docs, real_name)\n        if not os.path.exists(target):\n            shutil.move(full, target)\n            print(f\"重命名: {real_name}\")\n```\n\n\n\n## CMap 残缺度自检（v1.4 新增）\n\n不预设\"哪个 PDF 来源会乱码\"——实测 **72% 的乱码来自非 lightpdf PDF**（PPT 转 PDF、扫描件等），\n改用**自适应检测**：\n\n```python\ndef is_cmap_broken(text, threshold=0.03):\n    \"\"\"检测文本是否含异常字符（CMap 残缺/PUA 污染/未映射 CID）\"\"\"\n    if not text or len(text.strip()) < 50:\n        return False\n    total = len(text)\n    pua_count = sum(1 for c in text if 0xE000 <= ord(c) <= 0xF8FF)\n    cjk_ext = sum(1 for c in text if 0x20000 <= ord(c) <= 0x2EBEF)\n    cjk_compat = sum(1 for c in text if 0xF900 <= ord(c) <= 0xFAFF)\n    cid_count = text.count('(cid:')\n    bad_ratio = (pua_count + cjk_ext + cjk_compat + cid_count) / total\n    return bad_ratio > threshold or cid_count > 10\n```\n\n**3 类乱码特征：**\n1. **PUA 私用区** (U+E000-F8FF) —— 残缺 CMap fallback\n2. **CJK 扩展区** (U+20000-2EBEF) —— 字符找不到映射\n3. **`(cid:xxxx)` 字面值** —— pdfplumber 提取失败标志\n\n**集成位置：** `extract_pdf_text()` 在 kreuzberg / pymupdf 提取后调 `is_cmap_broken()`，\n通过即返回，失败即触发 OCR 路径。\n\n### OCR 性能实测（2026-08-12 验证）\n\n| 文件 | 大小 | OCR 耗时 | 备注 |\n|---|---|---|---|\n| lightpdf PDF | 4 页 | 13.5 秒 | CMap 残缺，自动 OCR |\n| 大型 PPT 转 PDF      | 90 页 | 0.5 秒 | 默认路径（无需 OCR）|\n| 65MB .doc | 169MB 文件 | 0.1 秒 | antiword 极速专线 |\n| 大型 docx（475K 字）| 562KB | 11.2 秒 | python-docx fallback |\n| OCR 自检总开销 | - | < 200ms | 3 页抽样 + 字符统计 |\n\n### 批量 OCR 修复脚本（v1.4 新增）\n\n`re_ocr_corrupted.py` —— 批量扫描乱码 summary，自动用新版本 utils 重新提取：\n\n```bash\n# 干跑（不写文件）\npython3 re_ocr_corrupted.py --dry-run --max 10\n\n# 实际批量（处理所有乱码）\npython3 re_ocr_corrupted.py --max 100\n\n# 只处理指定 PDF\npython3 re_ocr_corrupted.py --pdf-list \"path1.pdf,path2.pdf\"\n```\n\n行为：\n1. 扫 archives/ 找出乱码 summary\n2. 按 basename 匹配源 PDF\n3. 调 `extract_pdf_text()` 重跑（自动 OCR fallback）\n4. 写新 summary 到 summaries/（带新时间戳）\n5. 覆盖 archives/ 里对应 basename 的所有乱码版本\n6. 输出 JSON 报告（含每份文件路径/字数/成功状态）\n\n## 文档解析架构（v2.0）\n\n### 架构图\n\n```\n                    ┌──────────────────────────────────────┐\n                    │         kreuzberg 统一提取层           │\n                    │  (pypdfium2 / python-calamine / pandoc) │\n                    └───┬────────────────────────────────┬───┘\n                        │                              │\n                自动判断 │                              │\n                        ▼                              ▼\n              ┌─────────────────┐           ┌─────────────────────┐\n              │  kreuzberg 直提  │           │  antiword 极速专线  │\n              │ PDF/DOCX/XLSX/  │           │   (.doc 文件专用)    │\n              │ PPTX/MD/图片OCR │           │   成功率 85%，<1秒   │\n              └─────────────────┘           └─────────────────────┘\n                        │                              │\n                        │         ┌──────────────────────────────┐\n                        │         │     soffice 兜底转换          │\n                        │         │ (.doc/.xls/.ppt antiword失败) │\n                        │         │  60秒硬超时（消除误判watchdog）│\n                        │         └──────────────────────────────┘\n                        ▼                              │\n              ┌──────────────────────────────────────────────┐\n              │              文本输出（content）              │\n              │  → summaries/ 摘要文件 → generate_catalog.py  │\n              └──────────────────────────────────────────────┘\n```\n\n### 文件类型 × 提取方式\n\n| 格式 | 主方案 | 依赖 | 成功率 | 单文件速度 |\n|------|--------|------|--------|-----------|\n| PDF | kreuzberg (pypdfium2) | 无 | ~100% | 0.05-0.7s |\n| DOCX | kreuzberg + pandoc | pandoc 3.9+ | 100% | 0.12-3s |\n| XLSX | kreuzberg (python-calamine) | 无 | 100% | 0.1-0.5s |\n| PPTX | kreuzberg + pandoc | pandoc 3.9+ | 100% | 0.02-0.2s |\n| MD | kreuzberg + pandoc | pandoc 3.9+ | 100% | <0.01s |\n| **.doc** | **antiword 优先** | antiword | **85%**，<1秒 | <0.02s |\n| .doc（失败） | soffice 兜底 | LibreOffice | ~15% | 2-21s |\n| .xls | soffice → XLSX | LibreOffice | ~95% | 2-10s |\n| .ppt | soffice → PPTX | LibreOffice | ~95% | 2-10s |\n| 图片 | kreuzberg 内置 OCR | tesseract | ~90% | 3-10s |\n\n### antiword 极速专线\n\n```python\n# 实测数据：\n# 169MB 超大文件 → 26万字符，0.02秒完成\n# 正常 .doc（0.1-15MB）→ <1秒\n# 成功率 85%，覆盖绝大多数 .doc 文件\nresult = subprocess.run(['antiword', filepath], capture_output=True, timeout=10)\n```\n\n### kreuzberg 统一提取层\n\nkreuzberg 是专业的非结构化文档文本提取库（支持 20+ 格式），内部自动路由：\n- PDF → pypdfium2\n- XLSX → python-calamine\n- DOCX/PPTX/MD → pandoc\n- 图片 → 内置 OCR（tesseract）\n\n## 关键词库（中英双语）\n\n**中文（47个）：** 房产、房价、房地产、居民、消费、股市、经济、政策、利率、通胀、人民币、A股、美联储、PBOC、GDP、股票、资产、投资、债券、银行、PPI、CPI、PMI、M2、就业、失业、汽车、新能源、AI 等\n\n**英文（70+个）：** property、real estate、GDP、inflation、CPI、PPI、PMI、PBOC、Fed、consumer、economy、growth、housing、stock market、EV、AI 等\n\n**标签输出语言：** 自动判断——英文内容匹配英文关键词输出英文标签，中文内容匹配中文关键词输出中文标签\n\n## 定时任务兼容性\n\n| 任务 | ID | 调用方式 | 结论 |\n|------|------|---------|------|\n| 23:00分析新文档 | f3536e18 | 绝对路径 `python3 run_analysis.py` | ✅ 无需修改 |\n| 07:00生成财经早报 | b741c6d5 | Node.js 脚本 | ❌ 不相关 |\n| 08:00发送财经早报 | a7cbaacc | 读取文件发送 | ❌ 不相关 |\n| 09:00发送文档摘要 | 89b4cf75 | 读取 summaries 目录 | ❌ 不相关 |\n\n## 迁移到新电脑\n\n⚠️ **迁移前必读**：`~/.openclaw/workspace/knowledge/` 目录可能包含：\n- 私人合同/财务文档的 OCR 摘要\n- 个人分析报告\n- 飞书推送缓存\n\n**建议**：\n1. 先 `du -sh ~/.openclaw/workspace/knowledge/` 看大小\n2. 排除 `.trash/`、`.analysis/cache/` 后再迁移\n3. 用 `rsync -av --exclude='.trash' ...` 而不是 `scp -r`\n\n1. 复制目录（推荐 rsync）：\n   ```bash\n   rsync -av --exclude='.trash' --exclude='.analysis/cache' \\\n     ~/.openclaw/workspace/knowledge user@new-mac:~/.openclaw/workspace/\n   ```\n2. 在新电脑运行 `bash setup.sh <飞书用户ID>`\n3. 重新注册定时任务（Job ID 会变）\n\n## 避坑指南\n\n| 问题 | 原因 | 解决 |\n|------|------|------|\n| LibreOffice 超时 | watchdog 误判大文件为卡死 | v2.0 移除 watchdog，60秒硬超时 |\n| .doc 提取慢 | 统一走 LibreOffice | antiword 专线，169MB 文件 0.02秒 |\n| DOCX/PPTX 处理失败 | pandoc 未安装 | `brew install pandoc` |\n| PDF 提取乱码（OCR 不工作） | ocrmypdf `--skip-text` 跳过乱码页 | v1.4 改为 `--force-ocr` 强制 OCR |\n| PDF 漏检 CMap 残缺 | 没主动判断是否乱码 | v1.4 `is_cmap_broken()` 自检（阈值 0.03）|\n| 文本被截断到 8000 字 | 硬编码 `[:8000]` 太短 | v1.4 `MAX_EXTRACT_LEN = 500_000` |\n| .doc 提取失败 | lightpdf 处理过的 .doc 乱码 | v1.4 `ocr_office_via_ocr()` 兜底 |\n| summary 文件名过长失败 | 畸形 PDF 名 244+ bytes + 时间戳超 NAME_MAX | v1.4.1 `sanitize_filename()` + Errno 63 重试 |\n| 静默失败（不知道哪个文件）| 不抛异常 | v1.4 `PDFExtractError` / `ExtractError` 含路径 |\n| 飞书无 exec 工具 | tools 策略限制 | 添加 `alsoAllow: [exec, process]` |\n| BGE-M3 卡顿 | 16GB 内存不足 | 继续用 nomic-embed-text |\n\n## 关键路径\n\n| 内容 | 路径 |\n|------|------|\n| Skill 目录 | `~/.openclaw/workspace/skills/knowledge-base-setup/` |\n| 知识库 | `~/.openclaw/workspace/knowledge/` |\n| 分析脚本 | `~/.openclaw/workspace/knowledge/.analysis/` |\n| 目录缓存 | `~/.openclaw/workspace/knowledge/.analysis/.catalog_cache.json` |\n| 摘要输出 | `~/.openclaw/workspace/knowledge/.analysis/summaries/` |\n| 文章目录 | `~/.openclaw/workspace/knowledge/文章目录/文章目录.md` |\n| OpenClaw 配置 | `~/.openclaw/openclaw.json` |\n\n\n## ⚠️ 临时文件处理（v1.4.3 修复）\n\nOCR 流程会把 PDF 复制到临时目录处理。\n\n**v1.4.3 之前**：使用 `/tmp/ocrmypdf_work`、`/tmp/office_ocr_work`、\n`/tmp/office_convert` 共享路径，存在以下风险：\n- 多进程并发可能冲突\n- 多用户系统下其他用户可访问（权限默认 755）\n- 处理失败时临时文件残留\n\n**v1.4.3 修复**：\n- 用 `tempfile.mkdtemp(prefix=\"...\")` 创建 per-run 私有目录（权限 0o700）\n- 处理完成后立即 `shutil.rmtree` 清理\n- `tempfile.mkstemp` 创建稳定输出文件（避免被 finally 误删）\n\n**剩余风险**：极端情况下（机器突然断电）可能残留临时目录。\n建议定期清理 `/Users/home/.openclaw/tmp/` 下 `ocrmypdf_*`、`office_*` 前缀目录。\n\n## 版本历史\n\n| 版本 | 日期 | 更新内容 |\n|------|------|---------|\n| 1.0.0 | 2026-05-10 | 初始版本，PyMuPDF + LibreOffice 链路 |\n| 1.1.0 | 2026-05-13 | 三步 PDF 处理，关键词库，双语标签 |\n| 1.2.0 | 2026-05-21 | 分批处理优化，280秒断点 |\n| 1.2.1 | 2026-05-22 | utils.py 共享模块重构，LibreOffice 熔断机制 |\n| 1.3.0 | 2026-05-28 | kreuzberg 统一提取层 + antiword 专线 + pandoc |\n| 1.4.0 | 2026-08-12 | CMap 残缺度自检 + 50万字完整提取 + OCR fallback 到 .doc |\n| **1.4.1** | **2026-08-12** | **run_analysis.py: sanitize_filename + Errno 63 重试 + 清理 temp_docs 畸形文件** |\n\nFile v1.4.4:_meta.json\n\n{\n  \"ownerId\": \"kn7d38vmg57htam960qsj7wnkh86ewjb\",\n  \"slug\": \"macmini-knowledge-base\",\n  \"version\": \"1.4.4\",\n  \"publishedAt\": 1786972858009\n}\n\nFile v1.4.4:CHANGELOG.md\n\n# Changelog\n\n## v1.4.4 (2026-08-17) — SKILL.md 补充警告章节\n\n- 补充「## ⚠️ 临时文件处理（v1.4.3 修复）」章节（v1.4.3 发布时漏掉）\n- 代码、CHANGELOG 内容与 v1.4.3 一致\n- 仅文档补充\n\n## v1.4.3 (2026-08-17) — NVIDIA SkillSpector 审查修复\n\n### 安全修复（4 项）\n- **`os.system(pkill)` → `subprocess.run([...])`**（utils.py `_kill_proc_tree`）\n- **`/tmp/ocrmypdf_work` → `tempfile.mkdtemp()`**（`extract_pdf_via_ocr`）\n- **`/tmp/office_ocr_work` → `tempfile.mkdtemp()`**（`ocr_office_via_ocr`）\n- **`/tmp/office_convert` → `tempfile.mkdtemp()`**（`convert_old_office`，结果 cp 到稳定路径）\n\n### 文档警告（5 项）\n- description 加 Tp4 能力范围警告\n- 加 tools 权限声明块\n- 加 Unattended Automation 警告（cron 风险）\n- scp 迁移步骤加敏感数据警告\n- /tmp 临时目录警告\n\n### 验证\n- ✅ 7 项功能测试全部通过\n- ✅ 完整 cron 跑一次（run_analysis.py 退出码 0）\n- ✅ 修复了 convert_old_office 的 race condition（mkstemp）\n\n## v1.4.2 (2026-08-17) — 仅元数据更新\n## v1.4.1 (2026-08-12) — OSError 63 修复\n## v1.4.0 (2026-08-12) — OCR fallback + 50万字提取\n\nFile v1.4.4:skill-card.md\n\n## Description:\n\nMac 知识库搭建系统 helps OpenClaw users set up a local Mac Mini knowledge base with document extraction, RAG search configuration, summaries, tagging, and optional scheduled Feishu notifications.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[seairteng](https://clawhub.ai/user/seairteng)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers and individual OpenClaw users use this skill to create and maintain a local Mac knowledge-base workspace, extract text from common document formats, generate catalog metadata, and configure scheduled analysis and Feishu delivery.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: The setup flow can create persistent scheduled automation that reads local knowledge-base documents and sends notifications through Feishu.\n\nMitigation: Install on a personal trusted Mac only, review the exact OpenClaw cron command and Feishu recipient before registration, and periodically inspect or remove the cron job when it is no longer needed.\n\nRisk: The setup script installs Homebrew and Python dependencies, copies scripts into the analysis workspace, and pulls an Ollama embedding model.\n\nMitigation: Review setup.sh before execution, run it from a trusted skill checkout, and confirm dependency changes are acceptable for the target Mac.\n\nRisk: The knowledge directory may contain private contracts, financial documents, OCR summaries, analysis reports, and Feishu push cache data.\n\nMitigation: Keep the knowledge workspace out of shared or multi-user directories, avoid adding sensitive shared documents unless necessary, and exclude cache or trash folders during migration.\n\nRisk: Document OCR and office conversion create temporary working files; recent versions use per-run private temporary directories, but abrupt interruption can leave residual files.\n\nMitigation: Use version 1.4.3 or later behavior and periodically clean leftover ocrmypdf_* and office_* temporary directories if processing is interrupted.\n\n## Reference(s):\n\n- [ClawHub skill page](https://clawhub.ai/seairteng/skills/macmini-knowledge-base)\n- [Ollama download](https://ollama.com/download)\n\n## Skill Output:\n\n**Output Type(s):** [text, markdown, code, shell commands, configuration, guidance]\n\n**Output Format:** [Markdown guidance with inline shell commands, JSON configuration snippets, and Python code references]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [Produces local setup and operation guidance for a Mac/OpenClaw knowledge-base workflow; generated summaries and catalogs are stored in the configured knowledge workspace.]\n\n## Skill Version(s):\n\n1.4.4 (source: server release metadata, target metadata, _meta.json, and CHANGELOG, released 2026-08-17; SKILL.md frontmatter says 1.4.2)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nArchive v1.4.3: 9 files, 27096 bytes\n\nFiles: _meta.json (141b), CHANGELOG.md (1560b), scripts/generate_catalog.py (12426b), scripts/re_ocr_corrupted.py (7276b), scripts/run_analysis.py (8164b), scripts/setup.sh (3966b), scripts/utils.py (17289b), skill-card.md (2564b), SKILL.md (17372b)\n\nFile v1.4.3:SKILL.md\n\n---\nname: macmini-knowledge-base\nversion: 1.4.2\ndescription: |\n  在 Mac Mini (M4) 上快速搭建本地知识库 + RAG 自然语言搜索系统。\n  适用场景：\n  - 新 Mac 配置知识库：从零开始安装配置 Ollama、embedding模型、定时任务、文档解析\n  - 遇到 PDF 提取乱码、定时任务超时、skill 加载失败等问题\n  - 想要建立每日自动分析文档 + 08:00发送摘要到飞书的流程\n  - 迁移或复现知识库：打包整个 knowledge 目录和配置到新电脑\n  - **v1.4 新增**：CMap 残缺度自检（不预设来源）+ 50万字完整提取 + OCR fallback 到 .doc\n  本 skill 会引导完成：目录结构创建、依赖安装、脚本部署、定时任务注册、OpenClaw 配置。\n\n  ⚠️ **重要：能力范围**\n  本 skill 不只是「搭建」，还包含：\n  - 批量 OCR 修复（扫描 summaries/archives 找乱码 + 重新提取）\n  - 目录归档清理（移动重复/孤儿文件到 .trash/）\n  - 自动定时任务（23:00 分析 + 06:00 飞书推送）\n  \n  使用前请仔细评估批量修改风险。\n---\n\n# Knowledge Base Setup\n\n在 Mac Mini 上快速搭建本地知识库 + RAG 搜索系统。\n\n## 核心功能（v2.0）\n\n- **kreuzberg 统一提取层**：PDF / DOCX / XLSX / PPTX / MD / 图片 OCR 全自动路由\n- **antiword 极速专线**：.doc 文件专用提取，成功率 85%，169MB 文件 0.02 秒完成\n- **智能兜底**：antiword 失败自动走 soffice 转换，60 秒硬超时无误判\n- **自动分类**：关键词匹配驱动，中英文双语标签\n- **定时任务**：每天 23:00 分析新文档，08:00 发送摘要到飞书\n\n\n\n## 🔧 权限声明\n\n本 skill 在使用时需要以下 OpenClaw 工具能力：\n\n```json\n{\n  \"tools\": {\n    \"alsoAllow\": [\"exec\", \"process\", \"read\", \"write\"]\n  }\n}\n```\n\n⚠️ **执行风险**：exec + cron 自动化组合可导致持续命令执行，\n请在可信环境（个人 Mac）使用，不要在共享/服务器部署。\n\n\n\n## ⚠️ 安全警告：定时任务\n\n本 skill 注册 2 个 cron 任务（23:00 分析 + 06:00 推送），\n运行 shell 命令并自动推送消息到飞书。\n\n**潜在风险**：\n- 脚本路径被修改 → 自动执行任意命令\n- 知识库目录被入侵 → 自动读取/外发\n- 飞书 webhook 泄漏 → 自动推送被劫持\n\n**建议**：\n- 不要把 `~/.openclaw/workspace/knowledge` 放在共享/多用户目录\n- 定期检查 cron 配置（`openclaw cron list`）\n- 飞书 webhook 使用独立群组，不要复用其他机器人的 webhook\n- 仅在个人 Mac 上运行，不要部署到服务器\n\n## 快速开始\n\n### 一键安装\n\n```bash\ncd ~/.openclaw/workspace/skills/knowledge-base-setup/scripts\nbash setup.sh <飞书用户ID>\n```\n\n### 手动分步安装\n\n**Step 1: 系统依赖**\n```bash\nbrew install antiword tesseract pandoc\n```\n\n**Step 2: Python 依赖**\n```bash\npip3 install kreuzberg pytesseract pymupdf docx openpyxl python-pptx\n```\n\n**Step 3: Ollama + embedding 模型**\n```bash\n# 安装 Ollama: https://ollama.com/download\nollama pull nomic-embed-text\n```\n\n**Step 4: 创建目录结构**\n```bash\nmkdir -p ~/.openclaw/workspace/knowledge/.analysis/summaries/archives\nmkdir -p ~/.openclaw/workspace/knowledge/temp_docs\ntouch ~/.openclaw/workspace/knowledge/文章目录/文章目录.md\n```\n\n**Step 5: 部署脚本**\n```bash\ncp ~/.openclaw/workspace/skills/knowledge-base-setup/scripts/*.py \\\n   ~/.openclaw/workspace/knowledge/.analysis/\nchmod +x ~/.openclaw/workspace/knowledge/.analysis/*.py\n```\n\n**Step 6: 配置 OpenClaw**\n\n编辑 `~/.openclaw/openclaw.json`，加入：\n```json\n{\n  \"models\": {\n    \"providers\": {\n      \"ollama\": {\n        \"baseUrl\": \"http://127.0.0.1:11434\",\n        \"api\": \"ollama\",\n        \"models\": [\n          {\"id\": \"nomic-embed-text\", \"name\": \"Nomic Embed Text\"}\n        ]\n      }\n    }\n  },\n  \"agents\": {\n    \"defaults\": {\n      \"memorySearch\": {\n        \"provider\": \"ollama\",\n        \"model\": \"nomic-embed-text\"\n      }\n    }\n  }\n}\n```\n\n确保 tools 区块有：\n```json\n\"tools\": {\n    \"alsoAllow\": [\"exec\", \"process\"]\n}\n```\n\n然后重启：`openclaw gateway restart`\n\n**Step 7: 注册定时任务**\n```bash\n# 23:00 分析新文档\nopenclaw cron add \\\n  --name \"23:00分析新文档\" \\\n  --cron \"0 23 * * *\" \\\n  --tz \"Asia/Shanghai\" \\\n  --session isolated \\\n  --timeout-seconds 600 \\\n  --message \"cd ~/.openclaw/workspace/knowledge/.analysis && python3 run_analysis.py && python3 generate_catalog.py\" \\\n  --announce --channel feishu --to \"user:<飞书用户ID>\"\n\n# 08:00 发送文档摘要\nopenclaw cron add \\\n  --name \"08:00发送文档摘要\" \\\n  --cron \"0 8 * * *\" \\\n  --tz \"Asia/Shanghai\" \\\n  --session isolated \\\n  --timeout-seconds 120 \\\n  --message \"读取 summaries/ 目录发送摘要到飞书\" \\\n  --announce --channel feishu --to \"user:<飞书用户ID>\"\n```\n\n\n\n\n\n## summary 文件名 sanitize（v1.4.1 新增）\n\n防止 `OSError: [Errno 63] File name too long`（NAME_MAX=255 bytes）：\n\n```python\nSUMMARY_NAME_MAX = 200\n\ndef sanitize_filename(name, max_length=SUMMARY_NAME_MAX):\n    \"\"\"截断超长文件名，保留扩展名 + 8 位 MD5 hash 防冲突\"\"\"\n    name_bytes = name.encode('utf-8')\n    if len(name_bytes) <= max_length:\n        return name\n    \n    base, ext = os.path.splitext(name)\n    ext_bytes = ext.encode('utf-8')\n    base_bytes = base.encode('utf-8')\n    \n    import hashlib\n    h = hashlib.md5(name_bytes).hexdigest()[:8]\n    \n    reserve = len(ext_bytes) + 1 + 8  # \"_\" + hash + ext\n    available = max_length - reserve\n    \n    if available > 0 and len(base_bytes) > available:\n        truncated = base_bytes[:available].decode('utf-8', errors='ignore')\n        return f\"{truncated}_{h}{ext}\"\n    \n    return name[:max_length]\n```\n\n主循环的异常捕获重试：\n\n```python\ntry:\n    with open(summary_file, 'w', encoding='utf-8') as f:\n        f.write(content)\nexcept OSError as e:\n    if e.errno == 63:  # ENAMETOOLONG\n        short_name = sanitize_filename(filename, max_length=180)\n        summary_file = os.path.join(\n            SUMMARY_DIR,\n            f\"{timestamp}_{short_name}.summary.txt\"\n        )\n        with open(summary_file, 'w', encoding='utf-8') as f:\n            f.write(content)\n```\n\n**触发场景：** 畸形 PDF 文件名（如下载错误的 `_; filename_=utf-8''...` 双名拼接），原文件名 244+ bytes + 时间戳超 255 bytes 限制。\n\n**实测案例：** `20260730-Nomura-Asia Insights：China：The Politburo meeting indicated a shift to _countercyclical\" policies-260730.pdf_; filename_=utf-8''...pdf` (原 248 bytes) → sanitize 后 200 bytes + 8位 hash → 安全创建。\n\n### temp_docs 畸形文件清理（v1.4.1 新增）\n\n下载失败的 PDF 在文件名里重复了两次（`_; filename_=utf-8''` 分隔），实际只需保留前半。一次性清理脚本：\n\n```python\nimport os, shutil\ntemp_docs = os.path.expanduser(\"~/.openclaw/workspace/knowledge/temp_docs\")\ntrash_dir = os.path.expanduser(\"~/.openclaw/workspace/knowledge/.trash/temp_docs_<时间戳>\")\nos.makedirs(trash_dir, exist_ok=True)\n\nfor f in os.listdir(temp_docs):\n    if \"_; filename_=utf-8''\" in f:\n        full = os.path.join(temp_docs, f)\n        parts = f.split(\"_; filename_=utf-8''\")\n        real_name = parts[0]\n        target = os.path.join(temp_docs, real_name)\n        if not os.path.exists(target):\n            shutil.move(full, target)\n            print(f\"重命名: {real_name}\")\n```\n\n\n\n## CMap 残缺度自检（v1.4 新增）\n\n不预设\"哪个 PDF 来源会乱码\"——实测 **72% 的乱码来自非 lightpdf PDF**（PPT 转 PDF、扫描件等），\n改用**自适应检测**：\n\n```python\ndef is_cmap_broken(text, threshold=0.03):\n    \"\"\"检测文本是否含异常字符（CMap 残缺/PUA 污染/未映射 CID）\"\"\"\n    if not text or len(text.strip()) < 50:\n        return False\n    total = len(text)\n    pua_count = sum(1 for c in text if 0xE000 <= ord(c) <= 0xF8FF)\n    cjk_ext = sum(1 for c in text if 0x20000 <= ord(c) <= 0x2EBEF)\n    cjk_compat = sum(1 for c in text if 0xF900 <= ord(c) <= 0xFAFF)\n    cid_count = text.count('(cid:')\n    bad_ratio = (pua_count + cjk_ext + cjk_compat + cid_count) / total\n    return bad_ratio > threshold or cid_count > 10\n```\n\n**3 类乱码特征：**\n1. **PUA 私用区** (U+E000-F8FF) —— 残缺 CMap fallback\n2. **CJK 扩展区** (U+20000-2EBEF) —— 字符找不到映射\n3. **`(cid:xxxx)` 字面值** —— pdfplumber 提取失败标志\n\n**集成位置：** `extract_pdf_text()` 在 kreuzberg / pymupdf 提取后调 `is_cmap_broken()`，\n通过即返回，失败即触发 OCR 路径。\n\n### OCR 性能实测（2026-08-12 验证）\n\n| 文件 | 大小 | OCR 耗时 | 备注 |\n|---|---|---|---|\n| lightpdf PDF | 4 页 | 13.5 秒 | CMap 残缺，自动 OCR |\n| 大型 PPT 转 PDF      | 90 页 | 0.5 秒 | 默认路径（无需 OCR）|\n| 65MB .doc | 169MB 文件 | 0.1 秒 | antiword 极速专线 |\n| 大型 docx（475K 字）| 562KB | 11.2 秒 | python-docx fallback |\n| OCR 自检总开销 | - | < 200ms | 3 页抽样 + 字符统计 |\n\n### 批量 OCR 修复脚本（v1.4 新增）\n\n`re_ocr_corrupted.py` —— 批量扫描乱码 summary，自动用新版本 utils 重新提取：\n\n```bash\n# 干跑（不写文件）\npython3 re_ocr_corrupted.py --dry-run --max 10\n\n# 实际批量（处理所有乱码）\npython3 re_ocr_corrupted.py --max 100\n\n# 只处理指定 PDF\npython3 re_ocr_corrupted.py --pdf-list \"path1.pdf,path2.pdf\"\n```\n\n行为：\n1. 扫 archives/ 找出乱码 summary\n2. 按 basename 匹配源 PDF\n3. 调 `extract_pdf_text()` 重跑（自动 OCR fallback）\n4. 写新 summary 到 summaries/（带新时间戳）\n5. 覆盖 archives/ 里对应 basename 的所有乱码版本\n6. 输出 JSON 报告（含每份文件路径/字数/成功状态）\n\n## 文档解析架构（v2.0）\n\n### 架构图\n\n```\n                    ┌──────────────────────────────────────┐\n                    │         kreuzberg 统一提取层           │\n                    │  (pypdfium2 / python-calamine / pandoc) │\n                    └───┬────────────────────────────────┬───┘\n                        │                              │\n                自动判断 │                              │\n                        ▼                              ▼\n              ┌─────────────────┐           ┌─────────────────────┐\n              │  kreuzberg 直提  │           │  antiword 极速专线  │\n              │ PDF/DOCX/XLSX/  │           │   (.doc 文件专用)    │\n              │ PPTX/MD/图片OCR │           │   成功率 85%，<1秒   │\n              └─────────────────┘           └─────────────────────┘\n                        │                              │\n                        │         ┌──────────────────────────────┐\n                        │         │     soffice 兜底转换          │\n                        │         │ (.doc/.xls/.ppt antiword失败) │\n                        │         │  60秒硬超时（消除误判watchdog）│\n                        │         └──────────────────────────────┘\n                        ▼                              │\n              ┌──────────────────────────────────────────────┐\n              │              文本输出（content）              │\n              │  → summaries/ 摘要文件 → generate_catalog.py  │\n              └──────────────────────────────────────────────┘\n```\n\n### 文件类型 × 提取方式\n\n| 格式 | 主方案 | 依赖 | 成功率 | 单文件速度 |\n|------|--------|------|--------|-----------|\n| PDF | kreuzberg (pypdfium2) | 无 | ~100% | 0.05-0.7s |\n| DOCX | kreuzberg + pandoc | pandoc 3.9+ | 100% | 0.12-3s |\n| XLSX | kreuzberg (python-calamine) | 无 | 100% | 0.1-0.5s |\n| PPTX | kreuzberg + pandoc | pandoc 3.9+ | 100% | 0.02-0.2s |\n| MD | kreuzberg + pandoc | pandoc 3.9+ | 100% | <0.01s |\n| **.doc** | **antiword 优先** | antiword | **85%**，<1秒 | <0.02s |\n| .doc（失败） | soffice 兜底 | LibreOffice | ~15% | 2-21s |\n| .xls | soffice → XLSX | LibreOffice | ~95% | 2-10s |\n| .ppt | soffice → PPTX | LibreOffice | ~95% | 2-10s |\n| 图片 | kreuzberg 内置 OCR | tesseract | ~90% | 3-10s |\n\n### antiword 极速专线\n\n```python\n# 实测数据：\n# 169MB 超大文件 → 26万字符，0.02秒完成\n# 正常 .doc（0.1-15MB）→ <1秒\n# 成功率 85%，覆盖绝大多数 .doc 文件\nresult = subprocess.run(['antiword', filepath], capture_output=True, timeout=10)\n```\n\n### kreuzberg 统一提取层\n\nkreuzberg 是专业的非结构化文档文本提取库（支持 20+ 格式），内部自动路由：\n- PDF → pypdfium2\n- XLSX → python-calamine\n- DOCX/PPTX/MD → pandoc\n- 图片 → 内置 OCR（tesseract）\n\n## 关键词库（中英双语）\n\n**中文（47个）：** 房产、房价、房地产、居民、消费、股市、经济、政策、利率、通胀、人民币、A股、美联储、PBOC、GDP、股票、资产、投资、债券、银行、PPI、CPI、PMI、M2、就业、失业、汽车、新能源、AI 等\n\n**英文（70+个）：** property、real estate、GDP、inflation、CPI、PPI、PMI、PBOC、Fed、consumer、economy、growth、housing、stock market、EV、AI 等\n\n**标签输出语言：** 自动判断——英文内容匹配英文关键词输出英文标签，中文内容匹配中文关键词输出中文标签\n\n## 定时任务兼容性\n\n| 任务 | ID | 调用方式 | 结论 |\n|------|------|---------|------|\n| 23:00分析新文档 | f3536e18 | 绝对路径 `python3 run_analysis.py` | ✅ 无需修改 |\n| 07:00生成财经早报 | b741c6d5 | Node.js 脚本 | ❌ 不相关 |\n| 08:00发送财经早报 | a7cbaacc | 读取文件发送 | ❌ 不相关 |\n| 09:00发送文档摘要 | 89b4cf75 | 读取 summaries 目录 | ❌ 不相关 |\n\n## 迁移到新电脑\n\n⚠️ **迁移前必读**：`~/.openclaw/workspace/knowledge/` 目录可能包含：\n- 私人合同/财务文档的 OCR 摘要\n- 个人分析报告\n- 飞书推送缓存\n\n**建议**：\n1. 先 `du -sh ~/.openclaw/workspace/knowledge/` 看大小\n2. 排除 `.trash/`、`.analysis/cache/` 后再迁移\n3. 用 `rsync -av --exclude='.trash' ...` 而不是 `scp -r`\n\n1. 复制目录（推荐 rsync）：\n   ```bash\n   rsync -av --exclude='.trash' --exclude='.analysis/cache' \\\n     ~/.openclaw/workspace/knowledge user@new-mac:~/.openclaw/workspace/\n   ```\n2. 在新电脑运行 `bash setup.sh <飞书用户ID>`\n3. 重新注册定时任务（Job ID 会变）\n\n## 避坑指南\n\n| 问题 | 原因 | 解决 |\n|------|------|------|\n| LibreOffice 超时 | watchdog 误判大文件为卡死 | v2.0 移除 watchdog，60秒硬超时 |\n| .doc 提取慢 | 统一走 LibreOffice | antiword 专线，169MB 文件 0.02秒 |\n| DOCX/PPTX 处理失败 | pandoc 未安装 | `brew install pandoc` |\n| PDF 提取乱码（OCR 不工作） | ocrmypdf `--skip-text` 跳过乱码页 | v1.4 改为 `--force-ocr` 强制 OCR |\n| PDF 漏检 CMap 残缺 | 没主动判断是否乱码 | v1.4 `is_cmap_broken()` 自检（阈值 0.03）|\n| 文本被截断到 8000 字 | 硬编码 `[:8000]` 太短 | v1.4 `MAX_EXTRACT_LEN = 500_000` |\n| .doc 提取失败 | lightpdf 处理过的 .doc 乱码 | v1.4 `ocr_office_via_ocr()` 兜底 |\n| summary 文件名过长失败 | 畸形 PDF 名 244+ bytes + 时间戳超 NAME_MAX | v1.4.1 `sanitize_filename()` + Errno 63 重试 |\n| 静默失败（不知道哪个文件）| 不抛异常 | v1.4 `PDFExtractError` / `ExtractError` 含路径 |\n| 飞书无 exec 工具 | tools 策略限制 | 添加 `alsoAllow: [exec, process]` |\n| BGE-M3 卡顿 | 16GB 内存不足 | 继续用 nomic-embed-text |\n\n## 关键路径\n\n| 内容 | 路径 |\n|------|------|\n| Skill 目录 | `~/.openclaw/workspace/skills/knowledge-base-setup/` |\n| 知识库 | `~/.openclaw/workspace/knowledge/` |\n| 分析脚本 | `~/.openclaw/workspace/knowledge/.analysis/` |\n| 目录缓存 | `~/.openclaw/workspace/knowledge/.analysis/.catalog_cache.json` |\n| 摘要输出 | `~/.openclaw/workspace/knowledge/.analysis/summaries/` |\n| 文章目录 | `~/.openclaw/workspace/knowledge/文章目录/文章目录.md` |\n| OpenClaw 配置 | `~/.openclaw/openclaw.json` |\n\n## 版本历史\n\n| 版本 | 日期 | 更新内容 |\n|------|------|---------|\n| 1.0.0 | 2026-05-10 | 初始版本，PyMuPDF + LibreOffice 链路 |\n| 1.1.0 | 2026-05-13 | 三步 PDF 处理，关键词库，双语标签 |\n| 1.2.0 | 2026-05-21 | 分批处理优化，280秒断点 |\n| 1.2.1 | 2026-05-22 | utils.py 共享模块重构，LibreOffice 熔断机制 |\n| 1.3.0 | 2026-05-28 | kreuzberg 统一提取层 + antiword 专线 + pandoc |\n| 1.4.0 | 2026-08-12 | CMap 残缺度自检 + 50万字完整提取 + OCR fallback 到 .doc |\n| **1.4.1** | **2026-08-12** | **run_analysis.py: sanitize_filename + Errno 63 重试 + 清理 temp_docs 畸形文件** |\n\nFile v1.4.3:_meta.json\n\n{\n  \"ownerId\": \"kn7d38vmg57htam960qsj7wnkh86ewjb\",\n  \"slug\": \"macmini-knowledge-base\",\n  \"version\": \"1.4.3\",\n  \"publishedAt\": 1786972587350\n}\n\nFile v1.4.3:CHANGELOG.md\n\n# Changelog\n\n## v1.4.3 (2026-08-17) — NVIDIA SkillSpector 审查修复\n\n### 安全修复（4 项）\n- **`os.system(pkill)` → `subprocess.run([...])`**（utils.py `_kill_proc_tree`）\n - 避免 shell 注入，捕获输出，5 秒超时\n- **`/tmp/ocrmypdf_work` → `tempfile.mkdtemp()`**（`extract_pdf_via_ocr`）\n - per-run 私有目录，finally `shutil.rmtree` 清理\n- **`/tmp/office_ocr_work` → `tempfile.mkdtemp()`**（`ocr_office_via_ocr`）\n - 同上，per-run 私有目录\n- **`/tmp/office_convert` → `tempfile.mkdtemp()`**（`convert_old_office`）\n - 同上，但**在 finally 清理前先把结果 cp 到 `tempfile.mkstemp` 稳定路径**（避免被 finally 误删）\n\n### 文档警告（5 项）\n- 加 `tools` 权限声明块\n- 加 Tp4 能力范围明确化（描述顶部）\n- 加 Unattended Automation 警告（cron 风险）\n- 加 scp 迁移敏感数据警告\n- 加 /tmp 临时目录警告（含 v1.4.3 修复说明）\n\n### 验证\n- ✅ 7 项功能测试全部通过（kill_proc_tree / 3 个 OCR 路径 / 2 个综合路径 / 临时目录清理）\n- ✅ 完整 cron 跑一次（run_analysis.py 退出码 0）\n- ✅ 修复了 convert_old_office 的 race condition（用 mkstemp 而非 mktemp）\n\n## v1.4.2 (2026-08-17) — 仅元数据更新\n- 修复 ClawHub 显示名（从 V1.4.1 改为「Mac 知识库搭建系统」）\n\n## v1.4.1 (2026-08-12) — OSError 63 修复\n- sanitize_filename() 函数 + Errno 63 重试\n\n## v1.4.0 (2026-08-12) — OCR fallback + 50万字提取\n- is_cmap_broken 自检 + ocr_office_via_ocr + MAX_EXTRACT_LEN\n\nFile v1.4.3:skill-card.md\n\n## Description:\n\n在 Mac Mini (M4) 上搭建本地知识库和 RAG 自然语言搜索系统，并配置文档解析、定时分析、目录生成和飞书摘要推送。\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[seairteng](https://clawhub.ai/user/seairteng)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers and Mac users use this skill to create a personal OpenClaw knowledge-base workspace, install document extraction dependencies, deploy analysis scripts, and schedule recurring document analysis with optional Feishu delivery.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: The installer uses local shell commands, Homebrew, pip, and OpenClaw cron automation that can persistently execute analysis scripts.\n\nMitigation: Install only on a personal Mac you control, review setup.sh before running it, and inspect OpenClaw cron entries after setup.\n\nRisk: The knowledge workspace can contain private documents, OCR summaries, personal reports, and Feishu push data.\n\nMitigation: Avoid placing highly sensitive documents in the workspace unless Feishu delivery is acceptable, keep the workspace out of shared directories, and exclude .trash and cache directories when migrating.\n\nRisk: Global brew and pip installation can alter the user's Python and system dependency environment.\n\nMitigation: Consider using an isolated Python environment and review installed dependencies before running scheduled analysis.\n\n## Reference(s):\n\n- [ClawHub skill page](https://clawhub.ai/seairteng/skills/macmini-knowledge-base)\n- [Ollama download](https://ollama.com/download)\n- [Artifact changelog](artifact/CHANGELOG.md)\n\n## Skill Output:\n\n**Output Type(s):** [guidance, shell commands, configuration, code, markdown, JSON]\n\n**Output Format:** [Markdown instructions with shell commands, JSON configuration snippets, generated local files, markdown catalogs, text summaries, and JSON reports]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [Creates and updates files under the user's OpenClaw knowledge workspace and may register scheduled OpenClaw cron tasks for analysis and Feishu notifications.]\n\n## Skill Version(s):\n\n1.4.3 (source: server release evidence and artifact CHANGELOG.md, released 2026-08-17)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nArchive v1.4.2: 9 files, 26409 bytes\n\nFiles: _meta.json (141b), CHANGELOG.md (2133b), scripts/generate_catalog.py (12426b), scripts/re_ocr_corrupted.py (7276b), scripts/run_analysis.py (8164b), scripts/setup.sh (3966b), scripts/utils.py (16499b), skill-card.md (2674b), SKILL.md (15649b)\n\nFile v1.4.2:SKILL.md\n\n---\nname: macmini-knowledge-base\nversion: 1.4.2\ndescription: |\n  在 Mac Mini (M4) 上快速搭建本地知识库 + RAG 自然语言搜索系统。\n  适用场景：\n  - 新 Mac 配置知识库：从零开始安装配置 Ollama、embedding模型、定时任务、文档解析\n  - 遇到 PDF 提取乱码、定时任务超时、skill 加载失败等问题\n  - 想要建立每日自动分析文档 + 08:00发送摘要到飞书的流程\n  - 迁移或复现知识库：打包整个 knowledge 目录和配置到新电脑\n  - **v1.4 新增**：CMap 残缺度自检（不预设来源）+ 50万字完整提取 + OCR fallback 到 .doc\n  本 skill 会引导完成：目录结构创建、依赖安装、脚本部署、定时任务注册、OpenClaw 配置。\n---\n\n# Knowledge Base Setup\n\n在 Mac Mini 上快速搭建本地知识库 + RAG 搜索系统。\n\n## 核心功能（v2.0）\n\n- **kreuzberg 统一提取层**：PDF / DOCX / XLSX / PPTX / MD / 图片 OCR 全自动路由\n- **antiword 极速专线**：.doc 文件专用提取，成功率 85%，169MB 文件 0.02 秒完成\n- **智能兜底**：antiword 失败自动走 soffice 转换，60 秒硬超时无误判\n- **自动分类**：关键词匹配驱动，中英文双语标签\n- **定时任务**：每天 23:00 分析新文档，08:00 发送摘要到飞书\n\n## 快速开始\n\n### 一键安装\n\n```bash\ncd ~/.openclaw/workspace/skills/knowledge-base-setup/scripts\nbash setup.sh <飞书用户ID>\n```\n\n### 手动分步安装\n\n**Step 1: 系统依赖**\n```bash\nbrew install antiword tesseract pandoc\n```\n\n**Step 2: Python 依赖**\n```bash\npip3 install kreuzberg pytesseract pymupdf docx openpyxl python-pptx\n```\n\n**Step 3: Ollama + embedding 模型**\n```bash\n# 安装 Ollama: https://ollama.com/download\nollama pull nomic-embed-text\n```\n\n**Step 4: 创建目录结构**\n```bash\nmkdir -p ~/.openclaw/workspace/knowledge/.analysis/summaries/archives\nmkdir -p ~/.openclaw/workspace/knowledge/temp_docs\ntouch ~/.openclaw/workspace/knowledge/文章目录/文章目录.md\n```\n\n**Step 5: 部署脚本**\n```bash\ncp ~/.openclaw/workspace/skills/knowledge-base-setup/scripts/*.py \\\n   ~/.openclaw/workspace/knowledge/.analysis/\nchmod +x ~/.openclaw/workspace/knowledge/.analysis/*.py\n```\n\n**Step 6: 配置 OpenClaw**\n\n编辑 `~/.openclaw/openclaw.json`，加入：\n```json\n{\n  \"models\": {\n    \"providers\": {\n      \"ollama\": {\n        \"baseUrl\": \"http://127.0.0.1:11434\",\n        \"api\": \"ollama\",\n        \"models\": [\n          {\"id\": \"nomic-embed-text\", \"name\": \"Nomic Embed Text\"}\n        ]\n      }\n    }\n  },\n  \"agents\": {\n    \"defaults\": {\n      \"memorySearch\": {\n        \"provider\": \"ollama\",\n        \"model\": \"nomic-embed-text\"\n      }\n    }\n  }\n}\n```\n\n确保 tools 区块有：\n```json\n\"tools\": {\n    \"alsoAllow\": [\"exec\", \"process\"]\n}\n```\n\n然后重启：`openclaw gateway restart`\n\n**Step 7: 注册定时任务**\n```bash\n# 23:00 分析新文档\nopenclaw cron add \\\n  --name \"23:00分析新文档\" \\\n  --cron \"0 23 * * *\" \\\n  --tz \"Asia/Shanghai\" \\\n  --session isolated \\\n  --timeout-seconds 600 \\\n  --message \"cd ~/.openclaw/workspace/knowledge/.analysis && python3 run_analysis.py && python3 generate_catalog.py\" \\\n  --announce --channel feishu --to \"user:<飞书用户ID>\"\n\n# 08:00 发送文档摘要\nopenclaw cron add \\\n  --name \"08:00发送文档摘要\" \\\n  --cron \"0 8 * * *\" \\\n  --tz \"Asia/Shanghai\" \\\n  --session isolated \\\n  --timeout-seconds 120 \\\n  --message \"读取 summaries/ 目录发送摘要到飞书\" \\\n  --announce --channel feishu --to \"user:<飞书用户ID>\"\n```\n\n\n\n\n\n## summary 文件名 sanitize（v1.4.1 新增）\n\n防止 `OSError: [Errno 63] File name too long`（NAME_MAX=255 bytes）：\n\n```python\nSUMMARY_NAME_MAX = 200\n\ndef sanitize_filename(name, max_length=SUMMARY_NAME_MAX):\n    \"\"\"截断超长文件名，保留扩展名 + 8 位 MD5 hash 防冲突\"\"\"\n    name_bytes = name.encode('utf-8')\n    if len(name_bytes) <= max_length:\n        return name\n    \n    base, ext = os.path.splitext(name)\n    ext_bytes = ext.encode('utf-8')\n    base_bytes = base.encode('utf-8')\n    \n    import hashlib\n    h = hashlib.md5(name_bytes).hexdigest()[:8]\n    \n    reserve = len(ext_bytes) + 1 + 8  # \"_\" + hash + ext\n    available = max_length - reserve\n    \n    if available > 0 and len(base_bytes) > available:\n        truncated = base_bytes[:available].decode('utf-8', errors='ignore')\n        return f\"{truncated}_{h}{ext}\"\n    \n    return name[:max_length]\n```\n\n主循环的异常捕获重试：\n\n```python\ntry:\n    with open(summary_file, 'w', encoding='utf-8') as f:\n        f.write(content)\nexcept OSError as e:\n    if e.errno == 63:  # ENAMETOOLONG\n        short_name = sanitize_filename(filename, max_length=180)\n        summary_file = os.path.join(\n            SUMMARY_DIR,\n            f\"{timestamp}_{short_name}.summary.txt\"\n        )\n        with open(summary_file, 'w', encoding='utf-8') as f:\n            f.write(content)\n```\n\n**触发场景：** 畸形 PDF 文件名（如下载错误的 `_; filename_=utf-8''...` 双名拼接），原文件名 244+ bytes + 时间戳超 255 bytes 限制。\n\n**实测案例：** `20260730-Nomura-Asia Insights：China：The Politburo meeting indicated a shift to _countercyclical\" policies-260730.pdf_; filename_=utf-8''...pdf` (原 248 bytes) → sanitize 后 200 bytes + 8位 hash → 安全创建。\n\n### temp_docs 畸形文件清理（v1.4.1 新增）\n\n下载失败的 PDF 在文件名里重复了两次（`_; filename_=utf-8''` 分隔），实际只需保留前半。一次性清理脚本：\n\n```python\nimport os, shutil\ntemp_docs = os.path.expanduser(\"~/.openclaw/workspace/knowledge/temp_docs\")\ntrash_dir = os.path.expanduser(\"~/.openclaw/workspace/knowledge/.trash/temp_docs_<时间戳>\")\nos.makedirs(trash_dir, exist_ok=True)\n\nfor f in os.listdir(temp_docs):\n    if \"_; filename_=utf-8''\" in f:\n        full = os.path.join(temp_docs, f)\n        parts = f.split(\"_; filename_=utf-8''\")\n        real_name = parts[0]\n        target = os.path.join(temp_docs, real_name)\n        if not os.path.exists(target):\n            shutil.move(full, target)\n            print(f\"重命名: {real_name}\")\n```\n\n\n\n## CMap 残缺度自检（v1.4 新增）\n\n不预设\"哪个 PDF 来源会乱码\"——实测 **72% 的乱码来自非 lightpdf PDF**（PPT 转 PDF、扫描件等），\n改用**自适应检测**：\n\n```python\ndef is_cmap_broken(text, threshold=0.03):\n    \"\"\"检测文本是否含异常字符（CMap 残缺/PUA 污染/未映射 CID）\"\"\"\n    if not text or len(text.strip()) < 50:\n        return False\n    total = len(text)\n    pua_count = sum(1 for c in text if 0xE000 <= ord(c) <= 0xF8FF)\n    cjk_ext = sum(1 for c in text if 0x20000 <= ord(c) <= 0x2EBEF)\n    cjk_compat = sum(1 for c in text if 0xF900 <= ord(c) <= 0xFAFF)\n    cid_count = text.count('(cid:')\n    bad_ratio = (pua_count + cjk_ext + cjk_compat + cid_count) / total\n    return bad_ratio > threshold or cid_count > 10\n```\n\n**3 类乱码特征：**\n1. **PUA 私用区** (U+E000-F8FF) —— 残缺 CMap fallback\n2. **CJK 扩展区** (U+20000-2EBEF) —— 字符找不到映射\n3. **`(cid:xxxx)` 字面值** —— pdfplumber 提取失败标志\n\n**集成位置：** `extract_pdf_text()` 在 kreuzberg / pymupdf 提取后调 `is_cmap_broken()`，\n通过即返回，失败即触发 OCR 路径。\n\n### OCR 性能实测（2026-08-12 验证）\n\n| 文件 | 大小 | OCR 耗时 | 备注 |\n|---|---|---|---|\n| lightpdf PDF | 4 页 | 13.5 秒 | CMap 残缺，自动 OCR |\n| 大型 PPT 转 PDF      | 90 页 | 0.5 秒 | 默认路径（无需 OCR）|\n| 65MB .doc | 169MB 文件 | 0.1 秒 | antiword 极速专线 |\n| 大型 docx（475K 字）| 562KB | 11.2 秒 | python-docx fallback |\n| OCR 自检总开销 | - | < 200ms | 3 页抽样 + 字符统计 |\n\n### 批量 OCR 修复脚本（v1.4 新增）\n\n`re_ocr_corrupted.py` —— 批量扫描乱码 summary，自动用新版本 utils 重新提取：\n\n```bash\n# 干跑（不写文件）\npython3 re_ocr_corrupted.py --dry-run --max 10\n\n# 实际批量（处理所有乱码）\npython3 re_ocr_corrupted.py --max 100\n\n# 只处理指定 PDF\npython3 re_ocr_corrupted.py --pdf-list \"path1.pdf,path2.pdf\"\n```\n\n行为：\n1. 扫 archives/ 找出乱码 summary\n2. 按 basename 匹配源 PDF\n3. 调 `extract_pdf_text()` 重跑（自动 OCR fallback）\n4. 写新 summary 到 summaries/（带新时间戳）\n5. 覆盖 archives/ 里对应 basename 的所有乱码版本\n6. 输出 JSON 报告（含每份文件路径/字数/成功状态）\n\n## 文档解析架构（v2.0）\n\n### 架构图\n\n```\n                    ┌──────────────────────────────────────┐\n                    │         kreuzberg 统一提取层           │\n                    │  (pypdfium2 / python-calamine / pandoc) │\n                    └───┬────────────────────────────────┬───┘\n                        │                              │\n                自动判断 │                              │\n                        ▼                              ▼\n              ┌─────────────────┐           ┌─────────────────────┐\n              │  kreuzberg 直提  │           │  antiword 极速专线  │\n              │ PDF/DOCX/XLSX/  │           │   (.doc 文件专用)    │\n              │ PPTX/MD/图片OCR │           │   成功率 85%，<1秒   │\n              └─────────────────┘           └─────────────────────┘\n                        │                              │\n                        │         ┌──────────────────────────────┐\n                        │         │     soffice 兜底转换          │\n                        │         │ (.doc/.xls/.ppt antiword失败) │\n                        │         │  60秒硬超时（消除误判watchdog）│\n                        │         └──────────────────────────────┘\n                        ▼                              │\n              ┌──────────────────────────────────────────────┐\n              │              文本输出（content）              │\n              │  → summaries/ 摘要文件 → generate_catalog.py  │\n              └──────────────────────────────────────────────┘\n```\n\n### 文件类型 × 提取方式\n\n| 格式 | 主方案 | 依赖 | 成功率 | 单文件速度 |\n|------|--------|------|--------|-----------|\n| PDF | kreuzberg (pypdfium2) | 无 | ~100% | 0.05-0.7s |\n| DOCX | kreuzberg + pandoc | pandoc 3.9+ | 100% | 0.12-3s |\n| XLSX | kreuzberg (python-calamine) | 无 | 100% | 0.1-0.5s |\n| PPTX | kreuzberg + pandoc | pandoc 3.9+ | 100% | 0.02-0.2s |\n| MD | kreuzberg + pandoc | pandoc 3.9+ | 100% | <0.01s |\n| **.doc** | **antiword 优先** | antiword | **85%**，<1秒 | <0.02s |\n| .doc（失败） | soffice 兜底 | LibreOffice | ~15% | 2-21s |\n| .xls | soffice → XLSX | LibreOffice | ~95% | 2-10s |\n| .ppt | soffice → PPTX | LibreOffice | ~95% | 2-10s |\n| 图片 | kreuzberg 内置 OCR | tesseract | ~90% | 3-10s |\n\n### antiword 极速专线\n\n```python\n# 实测数据：\n# 169MB 超大文件 → 26万字符，0.02秒完成\n# 正常 .doc（0.1-15MB）→ <1秒\n# 成功率 85%，覆盖绝大多数 .doc 文件\nresult = subprocess.run(['antiword', filepath], capture_output=True, timeout=10)\n```\n\n### kreuzberg 统一提取层\n\nkreuzberg 是专业的非结构化文档文本提取库（支持 20+ 格式），内部自动路由：\n- PDF → pypdfium2\n- XLSX → python-calamine\n- DOCX/PPTX/MD → pandoc\n- 图片 → 内置 OCR（tesseract）\n\n## 关键词库（中英双语）\n\n**中文（47个）：** 房产、房价、房地产、居民、消费、股市、经济、政策、利率、通胀、人民币、A股、美联储、PBOC、GDP、股票、资产、投资、债券、银行、PPI、CPI、PMI、M2、就业、失业、汽车、新能源、AI 等\n\n**英文（70+个）：** property、real estate、GDP、inflation、CPI、PPI、PMI、PBOC、Fed、consumer、economy、growth、housing、stock market、EV、AI 等\n\n**标签输出语言：** 自动判断——英文内容匹配英文关键词输出英文标签，中文内容匹配中文关键词输出中文标签\n\n## 定时任务兼容性\n\n| 任务 | ID | 调用方式 | 结论 |\n|------|------|---------|------|\n| 23:00分析新文档 | f3536e18 | 绝对路径 `python3 run_analysis.py` | ✅ 无需修改 |\n| 07:00生成财经早报 | b741c6d5 | Node.js 脚本 | ❌ 不相关 |\n| 08:00发送财经早报 | a7cbaacc | 读取文件发送 | ❌ 不相关 |\n| 09:00发送文档摘要 | 89b4cf75 | 读取 summaries 目录 | ❌ 不相关 |\n\n## 迁移到新电脑\n\n1. 复制整个目录：\n   ```bash\n   scp -r ~/.openclaw/workspace/knowledge user@new-mac:~/.openclaw/workspace/\n   ```\n2. 在新电脑运行 `bash setup.sh <飞书用户ID>`\n3. 重新注册定时任务（Job ID 会变）\n\n## 避坑指南\n\n| 问题 | 原因 | 解决 |\n|------|------|------|\n| LibreOffice 超时 | watchdog 误判大文件为卡死 | v2.0 移除 watchdog，60秒硬超时 |\n| .doc 提取慢 | 统一走 LibreOffice | antiword 专线，169MB 文件 0.02秒 |\n| DOCX/PPTX 处理失败 | pandoc 未安装 | `brew install pandoc` |\n| PDF 提取乱码（OCR 不工作） | ocrmypdf `--skip-text` 跳过乱码页 | v1.4 改为 `--force-ocr` 强制 OCR |\n| PDF 漏检 CMap 残缺 | 没主动判断是否乱码 | v1.4 `is_cmap_broken()` 自检（阈值 0.03）|\n| 文本被截断到 8000 字 | 硬编码 `[:8000]` 太短 | v1.4 `MAX_EXTRACT_LEN = 500_000` |\n| .doc 提取失败 | lightpdf 处理过的 .doc 乱码 | v1.4 `ocr_office_via_ocr()` 兜底 |\n| summary 文件名过长失败 | 畸形 PDF 名 244+ bytes + 时间戳超 NAME_MAX | v1.4.1 `sanitize_filename()` + Errno 63 重试 |\n| 静默失败（不知道哪个文件）| 不抛异常 | v1.4 `PDFExtractError` / `ExtractError` 含路径 |\n| 飞书无 exec 工具 | tools 策略限制 | 添加 `alsoAllow: [exec, process]` |\n| BGE-M3 卡顿 | 16GB 内存不足 | 继续用 nomic-embed-text |\n\n## 关键路径\n\n| 内容 | 路径 |\n|------|------|\n| Skill 目录 | `~/.openclaw/workspace/skills/knowledge-base-setup/` |\n| 知识库 | `~/.openclaw/workspace/knowledge/` |\n| 分析脚本 | `~/.openclaw/workspace/knowledge/.analysis/` |\n| 目录缓存 | `~/.openclaw/workspace/knowledge/.analysis/.catalog_cache.json` |\n| 摘要输出 | `~/.openclaw/workspace/knowledge/.analysis/summaries/` |\n| 文章目录 | `~/.openclaw/workspace/knowledge/文章目录/文章目录.md` |\n| OpenClaw 配置 | `~/.openclaw/openclaw.json` |\n\n## 版本历史\n\n| 版本 | 日期 | 更新内容 |\n|------|------|---------|\n| 1.0.0 | 2026-05-10 | 初始版本，PyMuPDF + LibreOffice 链路 |\n| 1.1.0 | 2026-05-13 | 三步 PDF 处理，关键词库，双语标签 |\n| 1.2.0 | 2026-05-21 | 分批处理优化，280秒断点 |\n| 1.2.1 | 2026-05-22 | utils.py 共享模块重构，LibreOffice 熔断机制 |\n| 1.3.0 | 2026-05-28 | kreuzberg 统一提取层 + antiword 专线 + pandoc |\n| 1.4.0 | 2026-08-12 | CMap 残缺度自检 + 50万字完整提取 + OCR fallback 到 .doc |\n| **1.4.1** | **2026-08-12** | **run_analysis.py: sanitize_filename + Errno 63 重试 + 清理 temp_docs 畸形文件** |\n\nFile v1.4.2:_meta.json\n\n{\n  \"ownerId\": \"kn7d38vmg57htam960qsj7wnkh86ewjb\",\n  \"slug\": \"macmini-knowledge-base\",\n  \"version\": \"1.4.2\",\n  \"publishedAt\": 1786968967651\n}\n\nFile v1.4.2:CHANGELOG.md\n\n## v1.4.2 (2026-08-17) — 仅元数据更新\n\n- **修复 ClawHub 显示名**：从 `V1.4.1` 改为「Mac 知识库搭建系统」（之前 publish 时未传 `--name`，ClawHub fallback 到了 version 字段）\n- 代码、脚本、文档内容完全一致，仅发布元数据更新\n- 安装命令不变：`clawhub install macmini-knowledge-base`\n\n# Changelog\n\n## 1.4.1 (2026-08-12)\n\n### Fixed\n- **run_analysis.py**: `OSError: [Errno 63] File name too long` 修复\n  - 新增 `sanitize_filename()` 函数（截断 + 8位 MD5 hash 防冲突，限额 200 bytes）\n  - 主循环加 OSError 63 异常捕获 + 重试（max_length=180）\n  - 触发场景：畸形 PDF 文件名（如 `_; filename_=utf-8''...` 双名拼接）超出 NAME_MAX=255 bytes\n\n### Added\n- **run_analysis.py**: `sanitize_filename()` 函数\n- **run_analysis.py**: `SUMMARY_NAME_MAX = 200` 常量\n- **run_analysis.py**: 异常捕获 Errno 63 重试逻辑\n- **SKILL.md**: \"summary 文件名 sanitize\" 章节\n- **SKILL.md**: \"temp_docs 畸形文件清理\" 章节\n- **SKILL.md**: 避坑指南增加 sanitize 条目\n\n### Cleanup\n- 重命名 1 个畸形 PDF 文件（248 → 122 bytes）\n- temp_docs: `20260730-Nomura-..._; filename_=utf-8''Nomura-...pdf` → `20260730-Nomura-...pdf`\n\n## 1.4.0 (2026-08-12)\n\n### Changed\n- **utils.py**: `--skip-text` → `--force-ocr`（OCR 路径真正工作）\n- **utils.py**: OCR 超时 60s → 600s\n- **utils.py**: 文本提取上限 `[:8000]` → `MAX_EXTRACT_LEN = 500_000`\n\n### Added\n- **utils.py**: `is_cmap_broken()` CMap 残缺度自检\n- **utils.py**: `PDFExtractError` / `ExtractError` 异常类\n- **utils.py**: `ocr_office_via_ocr()` 通用 OCR 函数\n- **scripts/**: 新增 `re_ocr_corrupted.py`\n- **SKILL.md**: 新增\"CMap 残缺度自检\"章节\n\n### Impact\n- CMap 残缺: 40 → 0\n- 文档完整度: 8K 字 → 500K 字上限\n\n## 1.3.0 (2026-05-28)\n- kreuzberg 统一提取层\n- antiword 极速专线\n- pandoc 依赖\n\n## 1.2.1 (2026-05-22)\n- utils.py 共享模块重构\n- LibreOffice 熔断机制\n\n## 1.2.0 (2026-05-21)\n- 分批处理优化\n\n## 1.1.0 (2026-05-13)\n- 三步 PDF 处理\n\n## 1.0.0 (2026-05-10)\n- 初始版本\n\nFile v1.4.2:skill-card.md\n\n## Description:\n\n在 Mac Mini (M4) 上搭建本地知识库和 RAG 自然语言搜索系统，并引导安装依赖、部署文档分析脚本、注册定时任务和配置 OpenClaw。\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[seairteng](https://clawhub.ai/user/seairteng)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers and technically comfortable Mac users use this skill to set up a local document knowledge base, extract and tag documents, repair OCR extraction issues, and automate daily analysis or Feishu summary delivery.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: The skill asks for command execution and persistent scheduled local document processing.\n\nMitigation: Review the scripts and OpenClaw tool configuration before enabling exec/process, and install it only for a dedicated knowledge directory.\n\nRisk: Scheduled jobs and Feishu announcements can process or share summaries of local documents.\n\nMitigation: Keep the knowledge directory limited to documents intended for analysis and Feishu delivery, and review cron recipients before enabling scheduled announcements.\n\nRisk: Bulk OCR repair can write new summaries and overwrite matching archived summaries.\n\nMitigation: Run re_ocr_corrupted.py with --dry-run and a small --max value before allowing write operations.\n\nRisk: Migration guidance can copy the local knowledge directory to another machine.\n\nMitigation: Review the source and destination directories before running migration commands.\n\n## Reference(s):\n\n- [ClawHub skill page](https://clawhub.ai/seairteng/skills/macmini-knowledge-base)\n- [Ollama download](https://ollama.com/download)\n- [Artifact SKILL.md](artifact/SKILL.md)\n- [Artifact CHANGELOG.md](artifact/CHANGELOG.md)\n\n## Skill Output:\n\n**Output Type(s):** [text, markdown, code, shell commands, configuration, guidance]\n\n**Output Format:** [Markdown instructions with bash, JSON, and Python snippets; installed scripts produce text summaries, Markdown catalogs, JSON state or report files, and shell command output.]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [Processes local documents under ~/.openclaw/workspace/knowledge, can schedule recurring analysis, and can send summaries to Feishu when configured.]\n\n## Skill Version(s):\n\n1.4.2 (source: frontmatter, changelog released 2026-08-17, server evidence)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nArchive v1.4.1: 9 files, 25923 bytes\n\nFiles: _meta.json (141b), CHANGELOG.md (1785b), scripts/generate_catalog.py (12426b), scripts/re_ocr_corrupted.py (7276b), scripts/run_analysis.py (8164b), scripts/setup.sh (3966b), scripts/utils.py (16499b), skill-card.md (2434b), SKILL.md (15649b)\n\nFile v1.4.1:SKILL.md\n\n---\nname: macmini-knowledge-base\nversion: 1.4.0\ndescription: |\n  在 Mac Mini (M4) 上快速搭建本地知识库 + RAG 自然语言搜索系统。\n  适用场景：\n  - 新 Mac 配置知识库：从零开始安装配置 Ollama、embedding模型、定时任务、文档解析\n  - 遇到 PDF 提取乱码、定时任务超时、skill 加载失败等问题\n  - 想要建立每日自动分析文档 + 08:00发送摘要到飞书的流程\n  - 迁移或复现知识库：打包整个 knowledge 目录和配置到新电脑\n  - **v1.4 新增**：CMap 残缺度自检（不预设来源）+ 50万字完整提取 + OCR fallback 到 .doc\n  本 skill 会引导完成：目录结构创建、依赖安装、脚本部署、定时任务注册、OpenClaw 配置。\n---\n\n# Knowledge Base Setup\n\n在 Mac Mini 上快速搭建本地知识库 + RAG 搜索系统。\n\n## 核心功能（v2.0）\n\n- **kreuzberg 统一提取层**：PDF / DOCX / XLSX / PPTX / MD / 图片 OCR 全自动路由\n- **antiword 极速专线**：.doc 文件专用提取，成功率 85%，169MB 文件 0.02 秒完成\n- **智能兜底**：antiword 失败自动走 soffice 转换，60 秒硬超时无误判\n- **自动分类**：关键词匹配驱动，中英文双语标签\n- **定时任务**：每天 23:00 分析新文档，08:00 发送摘要到飞书\n\n## 快速开始\n\n### 一键安装\n\n```bash\ncd ~/.openclaw/workspace/skills/knowledge-base-setup/scripts\nbash setup.sh <飞书用户ID>\n```\n\n### 手动分步安装\n\n**Step 1: 系统依赖**\n```bash\nbrew install antiword tesseract pandoc\n```\n\n**Step 2: Python 依赖**\n```bash\npip3 install kreuzberg pytesseract pymupdf docx openpyxl python-pptx\n```\n\n**Step 3: Ollama + embedding 模型**\n```bash\n# 安装 Ollama: https://ollama.com/download\nollama pull nomic-embed-text\n```\n\n**Step 4: 创建目录结构**\n```bash\nmkdir -p ~/.openclaw/workspace/knowledge/.analysis/summaries/archives\nmkdir -p ~/.openclaw/workspace/knowledge/temp_docs\ntouch ~/.openclaw/workspace/knowledge/文章目录/文章目录.md\n```\n\n**Step 5: 部署脚本**\n```bash\ncp ~/.openclaw/workspace/skills/knowledge-base-setup/scripts/*.py \\\n   ~/.openclaw/workspace/knowledge/.analysis/\nchmod +x ~/.openclaw/workspace/knowledge/.analysis/*.py\n```\n\n**Step 6: 配置 OpenClaw**\n\n编辑 `~/.openclaw/openclaw.json`，加入：\n```json\n{\n  \"models\": {\n    \"providers\": {\n      \"ollama\": {\n        \"baseUrl\": \"http://127.0.0.1:11434\",\n        \"api\": \"ollama\",\n        \"models\": [\n          {\"id\": \"nomic-embed-text\", \"name\": \"Nomic Embed Text\"}\n        ]\n      }\n    }\n  },\n  \"agents\": {\n    \"defaults\": {\n      \"memorySearch\": {\n        \"provider\": \"ollama\",\n        \"model\": \"nomic-embed-text\"\n      }\n    }\n  }\n}\n```\n\n确保 tools 区块有：\n```json\n\"tools\": {\n    \"alsoAllow\": [\"exec\", \"process\"]\n}\n```\n\n然后重启：`openclaw gateway restart`\n\n**Step 7: 注册定时任务**\n```bash\n# 23:00 分析新文档\nopenclaw cron add \\\n  --name \"23:00分析新文档\" \\\n  --cron \"0 23 * * *\" \\\n  --tz \"Asia/Shanghai\" \\\n  --session isolated \\\n  --timeout-seconds 600 \\\n  --message \"cd ~/.openclaw/workspace/knowledge/.analysis && python3 run_analysis.py && python3 generate_catalog.py\" \\\n  --announce --channel feishu --to \"user:<飞书用户ID>\"\n\n# 08:00 发送文档摘要\nopenclaw cron add \\\n  --name \"08:00发送文档摘要\" \\\n  --cron \"0 8 * * *\" \\\n  --tz \"Asia/Shanghai\" \\\n  --session isolated \\\n  --timeout-seconds 120 \\\n  --message \"读取 summaries/ 目录发送摘要到飞书\" \\\n  --announce --channel feishu --to \"user:<飞书用户ID>\"\n```\n\n\n\n\n\n## summary 文件名 sanitize（v1.4.1 新增）\n\n防止 `OSError: [Errno 63] File name too long`（NAME_MAX=255 bytes）：\n\n```python\nSUMMARY_NAME_MAX = 200\n\ndef sanitize_filename(name, max_length=SUMMARY_NAME_MAX):\n    \"\"\"截断超长文件名，保留扩展名 + 8 位 MD5 hash 防冲突\"\"\"\n    name_bytes = name.encode('utf-8')\n    if len(name_bytes) <= max_length:\n        return name\n    \n    base, ext = os.path.splitext(name)\n    ext_bytes = ext.encode('utf-8')\n    base_bytes = base.encode('utf-8')\n    \n    import hashlib\n    h = hashlib.md5(name_bytes).hexdigest()[:8]\n    \n    reserve = len(ext_bytes) + 1 + 8  # \"_\" + hash + ext\n    available = max_length - reserve\n    \n    if available > 0 and len(base_bytes) > available:\n        truncated = base_bytes[:available].decode('utf-8', errors='ignore')\n        return f\"{truncated}_{h}{ext}\"\n    \n    return name[:max_length]\n```\n\n主循环的异常捕获重试：\n\n```python\ntry:\n    with open(summary_file, 'w', encoding='utf-8') as f:\n        f.write(content)\nexcept OSError as e:\n    if e.errno == 63:  # ENAMETOOLONG\n        short_name = sanitize_filename(filename, max_length=180)\n        summary_file = os.path.join(\n            SUMMARY_DIR,\n            f\"{timestamp}_{short_name}.summary.txt\"\n        )\n        with open(summary_file, 'w', encoding='utf-8') as f:\n            f.write(content)\n```\n\n**触发场景：** 畸形 PDF 文件名（如下载错误的 `_; filename_=utf-8''...` 双名拼接），原文件名 244+ bytes + 时间戳超 255 bytes 限制。\n\n**实测案例：** `20260730-Nomura-Asia Insights：China：The Politburo meeting indicated a shift to _countercyclical\" policies-260730.pdf_; filename_=utf-8''...pdf` (原 248 bytes) → sanitize 后 200 bytes + 8位 hash → 安全创建。\n\n### temp_docs 畸形文件清理（v1.4.1 新增）\n\n下载失败的 PDF 在文件名里重复了两次（`_; filename_=utf-8''` 分隔），实际只需保留前半。一次性清理脚本：\n\n```python\nimport os, shutil\ntemp_docs = os.path.expanduser(\"~/.openclaw/workspace/knowledge/temp_docs\")\ntrash_dir = os.path.expanduser(\"~/.openclaw/workspace/knowledge/.trash/temp_docs_<时间戳>\")\nos.makedirs(trash_dir, exist_ok=True)\n\nfor f in os.listdir(temp_docs):\n    if \"_; filename_=utf-8''\" in f:\n        full = os.path.join(temp_docs, f)\n        parts = f.split(\"_; filename_=utf-8''\")\n        real_name = parts[0]\n        target = os.path.join(temp_docs, real_name)\n        if not os.path.exists(target):\n            shutil.move(full, target)\n            print(f\"重命名: {real_name}\")\n```\n\n\n\n## CMap 残缺度自检（v1.4 新增）\n\n不预设\"哪个 PDF 来源会乱码\"——实测 **72% 的乱码来自非 lightpdf PDF**（PPT 转 PDF、扫描件等），\n改用**自适应检测**：\n\n```python\ndef is_cmap_broken(text, threshold=0.03):\n    \"\"\"检测文本是否含异常字符（CMap 残缺/PUA 污染/未映射 CID）\"\"\"\n    if not text or len(text.strip()) < 50:\n        return False\n    total = len(text)\n    pua_count = sum(1 for c in text if 0xE000 <= ord(c) <= 0xF8FF)\n    cjk_ext = sum(1 for c in text if 0x20000 <= ord(c) <= 0x2EBEF)\n    cjk_compat = sum(1 for c in text if 0xF900 <= ord(c) <= 0xFAFF)\n    cid_count = text.count('(cid:')\n    bad_ratio = (pua_count + cjk_ext + cjk_compat + cid_count) / total\n    return bad_ratio > threshold or cid_count > 10\n```\n\n**3 类乱码特征：**\n1. **PUA 私用区** (U+E000-F8FF) —— 残缺 CMap fallback\n2. **CJK 扩展区** (U+20000-2EBEF) —— 字符找不到映射\n3. **`(cid:xxxx)` 字面值** —— pdfplumber 提取失败标志\n\n**集成位置：** `extract_pdf_text()` 在 kreuzberg / pymupdf 提取后调 `is_cmap_broken()`，\n通过即返回，失败即触发 OCR 路径。\n\n### OCR 性能实测（2026-08-12 验证）\n\n| 文件 | 大小 | OCR 耗时 | 备注 |\n|---|---|---|---|\n| lightpdf PDF | 4 页 | 13.5 秒 | CMap 残缺，自动 OCR |\n| 大型 PPT 转 PDF      | 90 页 | 0.5 秒 | 默认路径（无需 OCR）|\n| 65MB .doc | 169MB 文件 | 0.1 秒 | antiword 极速专线 |\n| 大型 docx（475K 字）| 562KB | 11.2 秒 | python-docx fallback |\n| OCR 自检总开销 | - | < 200ms | 3 页抽样 + 字符统计 |\n\n### 批量 OCR 修复脚本（v1.4 新增）\n\n`re_ocr_corrupted.py` —— 批量扫描乱码 summary，自动用新版本 utils 重新提取：\n\n```bash\n# 干跑（不写文件）\npython3 re_ocr_corrupted.py --dry-run --max 10\n\n# 实际批量（处理所有乱码）\npython3 re_ocr_corrupted.py --max 100\n\n# 只处理指定 PDF\npython3 re_ocr_corrupted.py --pdf-list \"path1.pdf,path2.pdf\"\n```\n\n行为：\n1. 扫 archives/ 找出乱码 summary\n2. 按 basename 匹配源 PDF\n3. 调 `extract_pdf_text()` 重跑（自动 OCR fallback）\n4. 写新 summary 到 summaries/（带新时间戳）\n5. 覆盖 archives/ 里对应 basename 的所有乱码版本\n6. 输出 JSON 报告（含每份文件路径/字数/成功状态）\n\n## 文档解析架构（v2.0）\n\n### 架构图\n\n```\n                    ┌──────────────────────────────────────┐\n                    │         kreuzberg 统一提取层           │\n                    │  (pypdfium2 / python-calamine / pandoc) │\n                    └───┬────────────────────────────────┬───┘\n                        │                              │\n                自动判断 │                              │\n                        ▼                              ▼\n              ┌─────────────────┐           ┌─────────────────────┐\n              │  kreuzberg 直提  │           │  antiword 极速专线  │\n              │ PDF/DOCX/XLSX/  │           │   (.doc 文件专用)    │\n              │ PPTX/MD/图片OCR │           │   成功率 85%，<1秒   │\n              └─────────────────┘           └─────────────────────┘\n                        │                              │\n                        │         ┌──────────────────────────────┐\n                        │         │     soffice 兜底转换          │\n                        │         │ (.doc/.xls/.ppt antiword失败) │\n                        │         │  60秒硬超时（消除误判watchdog）│\n                        │         └──────────────────────────────┘\n                        ▼                              │\n              ┌──────────────────────────────────────────────┐\n              │              文本输出（content）              │\n              │  → summaries/ 摘要文件 → generate_catalog.py  │\n              └──────────────────────────────────────────────┘\n```\n\n### 文件类型 × 提取方式\n\n| 格式 | 主方案 | 依赖 | 成功率 | 单文件速度 |\n|------|--------|------|--------|-----------|\n| PDF | kreuzberg (pypdfium2) | 无 | ~100% | 0.05-0.7s |\n| DOCX | kreuzberg + pandoc | pandoc 3.9+ | 100% | 0.12-3s |\n| XLSX | kreuzberg (python-calamine) | 无 | 100% | 0.1-0.5s |\n| PPTX | kreuzberg + pandoc | pandoc 3.9+ | 100% | 0.02-0.2s |\n| MD | kreuzberg + pandoc | pandoc 3.9+ | 100% | <0.01s |\n| **.doc** | **antiword 优先** | antiword | **85%**，<1秒 | <0.02s |\n| .doc（失败） | soffice 兜底 | LibreOffice | ~15% | 2-21s |\n| .xls | soffice → XLSX | LibreOffice | ~95% | 2-10s |\n| .ppt | soffice → PPTX | LibreOffice | ~95% | 2-10s |\n| 图片 | kreuzberg 内置 OCR | tesseract | ~90% | 3-10s |\n\n### antiword 极速专线\n\n```python\n# 实测数据：\n# 169MB 超大文件 → 26万字符，0.02秒完成\n# 正常 .doc（0.1-15MB）→ <1秒\n# 成功率 85%，覆盖绝大多数 .doc 文件\nresult = subprocess.run(['antiword', filepath], capture_output=True, timeout=10)\n```\n\n### kreuzberg 统一提取层\n\nkreuzberg 是专业的非结构化文档文本提取库（支持 20+ 格式），内部自动路由：\n- PDF → pypdfium2\n- XLSX → python-calamine\n- DOCX/PPTX/MD → pandoc\n- 图片 → 内置 OCR（tesseract）\n\n## 关键词库（中英双语）\n\n**中文（47个）：** 房产、房价、房地产、居民、消费、股市、经济、政策、利率、通胀、人民币、A股、美联储、PBOC、GDP、股票、资产、投资、债券、银行、PPI、CPI、PMI、M2、就业、失业、汽车、新能源、AI 等\n\n**英文（70+个）：** property、real estate、GDP、inflation、CPI、PPI、PMI、PBOC、Fed、consumer、economy、growth、housing、stock market、EV、AI 等\n\n**标签输出语言：** 自动判断——英文内容匹配英文关键词输出英文标签，中文内容匹配中文关键词输出中文标签\n\n## 定时任务兼容性\n\n| 任务 | ID | 调用方式 | 结论 |\n|------|------|---------|------|\n| 23:00分析新文档 | f3536e18 | 绝对路径 `python3 run_analysis.py` | ✅ 无需修改 |\n| 07:00生成财经早报 | b741c6d5 | Node.js 脚本 | ❌ 不相关 |\n| 08:00发送财经早报 | a7cbaacc | 读取文件发送 | ❌ 不相关 |\n| 09:00发送文档摘要 | 89b4cf75 | 读取 summaries 目录 | ❌ 不相关 |\n\n## 迁移到新电脑\n\n1. 复制整个目录：\n   ```bash\n   scp -r ~/.openclaw/workspace/knowledge user@new-mac:~/.openclaw/workspace/\n   ```\n2. 在新电脑运行 `bash setup.sh <飞书用户ID>`\n3. 重新注册定时任务（Job ID 会变）\n\n## 避坑指南\n\n| 问题 | 原因 | 解决 |\n|------|------|------|\n| LibreOffice 超时 | watchdog 误判大文件为卡死 | v2.0 移除 watchdog，60秒硬超时 |\n| .doc 提取慢 | 统一走 LibreOffice | antiword 专线，169MB 文件 0.02秒 |\n| DOCX/PPTX 处理失败 | pandoc 未安装 | `brew install pandoc` |\n| PDF 提取乱码（OCR 不工作） | ocrmypdf `--skip-text` 跳过乱码页 | v1.4 改为 `--force-ocr` 强制 OCR |\n| PDF 漏检 CMap 残缺 | 没主动判断是否乱码 | v1.4 `is_cmap_broken()` 自检（阈值 0.03）|\n| 文本被截断到 8000 字 | 硬编码 `[:8000]` 太短 | v1.4 `MAX_EXTRACT_LEN = 500_000` |\n| .doc 提取失败 | lightpdf 处理过的 .doc 乱码 | v1.4 `ocr_office_via_ocr()` 兜底 |\n| summary 文件名过长失败 | 畸形 PDF 名 244+ bytes + 时间戳超 NAME_MAX | v1.4.1 `sanitize_filename()` + Errno 63 重试 |\n| 静默失败（不知道哪个文件）| 不抛异常 | v1.4 `PDFExtractError` / `ExtractError` 含路径 |\n| 飞书无 exec 工具 | tools 策略限制 | 添加 `alsoAllow: [exec, process]` |\n| BGE-M3 卡顿 | 16GB 内存不足 | 继续用 nomic-embed-text |\n\n## 关键路径\n\n| 内容 | 路径 |\n|------|------|\n| Skill 目录 | `~/.openclaw/workspace/skills/knowledge-base-setup/` |\n| 知识库 | `~/.openclaw/workspace/knowledge/` |\n| 分析脚本 | `~/.openclaw/workspace/knowledge/.analysis/` |\n| 目录缓存 | `~/.openclaw/workspace/knowledge/.analysis/.catalog_cache.json` |\n| 摘要输出 | `~/.openclaw/workspace/knowledge/.analysis/summaries/` |\n| 文章目录 | `~/.openclaw/workspace/knowledge/文章目录/文章目录.md` |\n| OpenClaw 配置 | `~/.openclaw/openclaw.json` |\n\n## 版本历史\n\n| 版本 | 日期 | 更新内容 |\n|------|------|---------|\n| 1.0.0 | 2026-05-10 | 初始版本，PyMuPDF + LibreOffice 链路 |\n| 1.1.0 | 2026-05-13 | 三步 PDF 处理，关键词库，双语标签 |\n| 1.2.0 | 2026-05-21 | 分批处理优化，280秒断点 |\n| 1.2.1 | 2026-05-22 | utils.py 共享模块重构，LibreOffice 熔断机制 |\n| 1.3.0 | 2026-05-28 | kreuzberg 统一提取层 + antiword 专线 + pandoc |\n| 1.4.0 | 2026-08-12 | CMap 残缺度自检 + 50万字完整提取 + OCR fallback 到 .doc |\n| **1.4.1** | **2026-08-12** | **run_analysis.py: sanitize_filename + Errno 63 重试 + 清理 temp_docs 畸形文件** |\n\nFile v1.4.1:_meta.json\n\n{\n  \"ownerId\": \"kn7d38vmg57htam960qsj7wnkh86ewjb\",\n  \"slug\": \"macmini-knowledge-base\",\n  \"version\": \"1.4.1\",\n  \"publishedAt\": 1786547774541\n}\n\nFile v1.4.1:CHANGELOG.md\n\n# Changelog\n\n## 1.4.1 (2026-08-12)\n\n### Fixed\n- **run_analysis.py**: `OSError: [Errno 63] File name too long` 修复\n  - 新增 `sanitize_filename()` 函数（截断 + 8位 MD5 hash 防冲突，限额 200 bytes）\n  - 主循环加 OSError 63 异常捕获 + 重试（max_length=180）\n  - 触发场景：畸形 PDF 文件名（如 `_; filename_=utf-8''...` 双名拼接）超出 NAME_MAX=255 bytes\n\n### Added\n- **run_analysis.py**: `sanitize_filename()` 函数\n- **run_analysis.py**: `SUMMARY_NAME_MAX = 200` 常量\n- **run_analysis.py**: 异常捕获 Errno 63 重试逻辑\n- **SKILL.md**: \"summary 文件名 sanitize\" 章节\n- **SKILL.md**: \"temp_docs 畸形文件清理\" 章节\n- **SKILL.md**: 避坑指南增加 sanitize 条目\n\n### Cleanup\n- 重命名 1 个畸形 PDF 文件（248 → 122 bytes）\n- temp_docs: `20260730-Nomura-..._; filename_=utf-8''Nomura-...pdf` → `20260730-Nomura-...pdf`\n\n## 1.4.0 (2026-08-12)\n\n### Changed\n- **utils.py**: `--skip-text` → `--force-ocr`（OCR 路径真正工作）\n- **utils.py**: OCR 超时 60s → 600s\n- **utils.py**: 文本提取上限 `[:8000]` → `MAX_EXTRACT_LEN = 500_000`\n\n### Added\n- **utils.py**: `is_cmap_broken()` CMap 残缺度自检\n- **utils.py**: `PDFExtractError` / `ExtractError` 异常类\n- **utils.py**: `ocr_office_via_ocr()` 通用 OCR 函数\n- **scripts/**: 新增 `re_ocr_corrupted.py`\n- **SKILL.md**: 新增\"CMap 残缺度自检\"章节\n\n### Impact\n- CMap 残缺: 40 → 0\n- 文档完整度: 8K 字 → 500K 字上限\n\n## 1.3.0 (2026-05-28)\n- kreuzberg 统一提取层\n- antiword 极速专线\n- pandoc 依赖\n\n## 1.2.1 (2026-05-22)\n- utils.py 共享模块重构\n- LibreOffice 熔断机制\n\n## 1.2.0 (2026-05-21)\n- 分批处理优化\n\n## 1.1.0 (2026-05-13)\n- 三步 PDF 处理\n\n## 1.0.0 (2026-05-10)\n- 初始版本\n\nFile v1.4.1:skill-card.md\n\n## Description:\n\nSets up a local Mac Mini knowledge base and RAG search workflow with document extraction, OCR fallback, scheduled analysis, catalog generation, and Feishu summary delivery.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[seairteng](https://clawhub.ai/user/seairteng)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers and technically capable Mac users use this skill to install and operate a local OpenClaw knowledge base, configure Ollama embeddings, analyze local documents, generate searchable catalogs, and schedule daily summaries.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: The skill enables broad host command execution through exec/process and runs local shell and Python workflows.\n\nMitigation: Review the scripts before installation, enable exec/process only for a trusted workspace, and keep host permissions scoped to the intended knowledge-base directories.\n\nRisk: Scheduled background processing can repeatedly analyze local documents without further prompts.\n\nMitigation: Confirm cron entries, timeouts, exclusions, and backup or removal steps before enabling unattended daily processing.\n\nRisk: Feishu delivery can send document-derived summaries outside the local machine.\n\nMitigation: Run setup without a Feishu user ID until scripts and destination accounts are reviewed, and avoid processing confidential directories unless delivery policy is approved.\n\n## Reference(s):\n\n- [ClawHub skill page](https://clawhub.ai/seairteng/skills/macmini-knowledge-base)\n- [Ollama download](https://ollama.com/download)\n\n## Skill Output:\n\n**Output Type(s):** [text, markdown, code, shell commands, configuration, guidance]\n\n**Output Format:** [Markdown guidance with shell commands, JSON configuration snippets, and Python code snippets]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [Guides local file analysis, script deployment, OpenClaw configuration, cron registration, and Feishu summary delivery.]\n\n## Skill Version(s):\n\n1.4.1 (source: server release metadata, _meta.json, and CHANGELOG released 2026-08-12; SKILL.md frontmatter says 1.4.0)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nArchive v1.4.0: 9 files, 23886 bytes\n\nFiles: _meta.json (141b), CHANGELOG.md (1787b), scripts/generate_catalog.py (12426b), scripts/re_ocr_corrupted.py (7276b), scripts/run_analysis.py (5921b), scripts/setup.sh (3966b), scripts/utils.py (16499b), skill-card.md (2677b), SKILL.md (12823b)\n\nFile v1.4.0:SKILL.md\n\n---\nname: macmini-knowledge-base\nversion: 1.4.0\ndescription: |\n  在 Mac Mini (M4) 上快速搭建本地知识库 + RAG 自然语言搜索系统。\n  适用场景：\n  - 新 Mac 配置知识库：从零开始安装配置 Ollama、embedding模型、定时任务、文档解析\n  - 遇到 PDF 提取乱码、定时任务超时、skill 加载失败等问题\n  - 想要建立每日自动分析文档 + 08:00发送摘要到飞书的流程\n  - 迁移或复现知识库：打包整个 knowledge 目录和配置到新电脑\n  - **v1.4 新增**：CMap 残缺度自检（不预设来源）+ 50万字完整提取 + OCR fallback 到 .doc\n  本 skill 会引导完成：目录结构创建、依赖安装、脚本部署、定时任务注册、OpenClaw 配置。\n---\n\n# Knowledge Base Setup\n\n在 Mac Mini 上快速搭建本地知识库 + RAG 搜索系统。\n\n## 核心功能（v2.0）\n\n- **kreuzberg 统一提取层**：PDF / DOCX / XLSX / PPTX / MD / 图片 OCR 全自动路由\n- **antiword 极速专线**：.doc 文件专用提取，成功率 85%，169MB 文件 0.02 秒完成\n- **智能兜底**：antiword 失败自动走 soffice 转换，60 秒硬超时无误判\n- **自动分类**：关键词匹配驱动，中英文双语标签\n- **定时任务**：每天 23:00 分析新文档，08:00 发送摘要到飞书\n\n## 快速开始\n\n### 一键安装\n\n```bash\ncd ~/.openclaw/workspace/skills/knowledge-base-setup/scripts\nbash setup.sh <飞书用户ID>\n```\n\n### 手动分步安装\n\n**Step 1: 系统依赖**\n```bash\nbrew install antiword tesseract pandoc\n```\n\n**Step 2: Python 依赖**\n```bash\npip3 install kreuzberg pytesseract pymupdf docx openpyxl python-pptx\n```\n\n**Step 3: Ollama + embedding 模型**\n```bash\n# 安装 Ollama: https://ollama.com/download\nollama pull nomic-embed-text\n```\n\n**Step 4: 创建目录结构**\n```bash\nmkdir -p ~/.openclaw/workspace/knowledge/.analysis/summaries/archives\nmkdir -p ~/.openclaw/workspace/knowledge/temp_docs\ntouch ~/.openclaw/workspace/knowledge/文章目录/文章目录.md\n```\n\n**Step 5: 部署脚本**\n```bash\ncp ~/.openclaw/workspace/skills/knowledge-base-setup/scripts/*.py \\\n   ~/.openclaw/workspace/knowledge/.analysis/\nchmod +x ~/.openclaw/workspace/knowledge/.analysis/*.py\n```\n\n**Step 6: 配置 OpenClaw**\n\n编辑 `~/.openclaw/openclaw.json`，加入：\n```json\n{\n  \"models\": {\n    \"providers\": {\n      \"ollama\": {\n        \"baseUrl\": \"http://127.0.0.1:11434\",\n        \"api\": \"ollama\",\n        \"models\": [\n          {\"id\": \"nomic-embed-text\", \"name\": \"Nomic Embed Text\"}\n        ]\n      }\n    }\n  },\n  \"agents\": {\n    \"defaults\": {\n      \"memorySearch\": {\n        \"provider\": \"ollama\",\n        \"model\": \"nomic-embed-text\"\n      }\n    }\n  }\n}\n```\n\n确保 tools 区块有：\n```json\n\"tools\": {\n    \"alsoAllow\": [\"exec\", \"process\"]\n}\n```\n\n然后重启：`openclaw gateway restart`\n\n**Step 7: 注册定时任务**\n```bash\n# 23:00 分析新文档\nopenclaw cron add \\\n  --name \"23:00分析新文档\" \\\n  --cron \"0 23 * * *\" \\\n  --tz \"Asia/Shanghai\" \\\n  --session isolated \\\n  --timeout-seconds 600 \\\n  --message \"cd ~/.openclaw/workspace/knowledge/.analysis && python3 run_analysis.py && python3 generate_catalog.py\" \\\n  --announce --channel feishu --to \"user:<飞书用户ID>\"\n\n# 08:00 发送文档摘要\nopenclaw cron add \\\n  --name \"08:00发送文档摘要\" \\\n  --cron \"0 8 * * *\" \\\n  --tz \"Asia/Shanghai\" \\\n  --session isolated \\\n  --timeout-seconds 120 \\\n  --message \"读取 summaries/ 目录发送摘要到飞书\" \\\n  --announce --channel feishu --to \"user:<飞书用户ID>\"\n```\n\n\n\n## CMap 残缺度自检（v1.4 新增）\n\n不预设\"哪个 PDF 来源会乱码\"——实测 **72% 的乱码来自非 lightpdf PDF**（PPT 转 PDF、扫描件等），\n改用**自适应检测**：\n\n```python\ndef is_cmap_broken(text, threshold=0.03):\n    \"\"\"检测文本是否含异常字符（CMap 残缺/PUA 污染/未映射 CID）\"\"\"\n    if not text or len(text.strip()) < 50:\n        return False\n    total = len(text)\n    pua_count = sum(1 for c in text if 0xE000 <= ord(c) <= 0xF8FF)\n    cjk_ext = sum(1 for c in text if 0x20000 <= ord(c) <= 0x2EBEF)\n    cjk_compat = sum(1 for c in text if 0xF900 <= ord(c) <= 0xFAFF)\n    cid_count = text.count('(cid:')\n    bad_ratio = (pua_count + cjk_ext + cjk_compat + cid_count) / total\n    return bad_ratio > threshold or cid_count > 10\n```\n\n**3 类乱码特征：**\n1. **PUA 私用区** (U+E000-F8FF) —— 残缺 CMap fallback\n2. **CJK 扩展区** (U+20000-2EBEF) —— 字符找不到映射\n3. **`(cid:xxxx)` 字面值** —— pdfplumber 提取失败标志\n\n**集成位置：** `extract_pdf_text()` 在 kreuzberg / pymupdf 提取后调 `is_cmap_broken()`，\n通过即返回，失败即触发 OCR 路径。\n\n### OCR 性能实测（2026-08-12 验证）\n\n| 文件 | 大小 | OCR 耗时 | 备注 |\n|---|---|---|---|\n| lightpdf PDF | 4 页 | 13.5 秒 | CMap 残缺，自动 OCR |\n| 大型 PPT 转 PDF      | 90 页 | 0.5 秒 | 默认路径（无需 OCR）|\n| 65MB .doc | 169MB 文件 | 0.1 秒 | antiword 极速专线 |\n| 大型 docx（475K 字）| 562KB | 11.2 秒 | python-docx fallback |\n| OCR 自检总开销 | - | < 200ms | 3 页抽样 + 字符统计 |\n\n### 批量 OCR 修复脚本（v1.4 新增）\n\n`re_ocr_corrupted.py` —— 批量扫描乱码 summary，自动用新版本 utils 重新提取：\n\n```bash\n# 干跑（不写文件）\npython3 re_ocr_corrupted.py --dry-run --max 10\n\n# 实际批量（处理所有乱码）\npython3 re_ocr_corrupted.py --max 100\n\n# 只处理指定 PDF\npython3 re_ocr_corrupted.py --pdf-list \"path1.pdf,path2.pdf\"\n```\n\n行为：\n1. 扫 archives/ 找出乱码 summary\n2. 按 basename 匹配源 PDF\n3. 调 `extract_pdf_text()` 重跑（自动 OCR fallback）\n4. 写新 summary 到 summaries/（带新时间戳）\n5. 覆盖 archives/ 里对应 basename 的所有乱码版本\n6. 输出 JSON 报告（含每份文件路径/字数/成功状态）\n\n## 文档解析架构（v2.0）\n\n### 架构图\n\n```\n                    ┌──────────────────────────────────────┐\n                    │         kreuzberg 统一提取层           │\n                    │  (pypdfium2 / python-calamine / pandoc) │\n                    └───┬────────────────────────────────┬───┘\n                        │                              │\n                自动判断 │                              │\n                        ▼                              ▼\n              ┌─────────────────┐           ┌─────────────────────┐\n              │  kreuzberg 直提  │           │  antiword 极速专线  │\n              │ PDF/DOCX/XLSX/  │           │   (.doc 文件专用)    │\n              │ PPTX/MD/图片OCR │           │   成功率 85%，<1秒   │\n              └─────────────────┘           └─────────────────────┘\n                        │                              │\n                        │         ┌──────────────────────────────┐\n                        │         │     soffice 兜底转换          │\n                        │         │ (.doc/.xls/.ppt antiword失败) │\n                        │         │  60秒硬超时（消除误判watchdog）│\n                        │         └──────────────────────────────┘\n                        ▼                              │\n              ┌──────────────────────────────────────────────┐\n              │              文本输出（content）              │\n              │  → summaries/ 摘要文件 → generate_catalog.py  │\n              └──────────────────────────────────────────────┘\n```\n\n### 文件类型 × 提取方式\n\n| 格式 | 主方案 | 依赖 | 成功率 | 单文件速度 |\n|------|--------|------|--------|-----------|\n| PDF | kreuzberg (pypdfium2) | 无 | ~100% | 0.05-0.7s |\n| DOCX | kreuzberg + pandoc | pandoc 3.9+ | 100% | 0.12-3s |\n| XLSX | kreuzberg (python-calamine) | 无 | 100% | 0.1-0.5s |\n| PPTX | kreuzberg + pandoc | pandoc 3.9+ | 100% | 0.02-0.2s |\n| MD | kreuzberg + pandoc | pandoc 3.9+ | 100% | <0.01s |\n| **.doc** | **antiword 优先** | antiword | **85%**，<1秒 | <0.02s |\n| .doc（失败） | soffice 兜底 | LibreOffice | ~15% | 2-21s |\n| .xls | soffice → XLSX | LibreOffice | ~95% | 2-10s |\n| .ppt | soffice → PPTX | LibreOffice | ~95% | 2-10s |\n| 图片 | kreuzberg 内置 OCR | tesseract | ~90% | 3-10s |\n\n### antiword 极速专线\n\n```python\n# 实测数据：\n# 169MB 超大文件 → 26万字符，0.02秒完成\n# 正常 .doc（0.1-15MB）→ <1秒\n# 成功率 85%，覆盖绝大多数 .doc 文件\nresult = subprocess.run(['antiword', filepath], capture_output=True, timeout=10)\n```\n\n### kreuzberg 统一提取层\n\nkreuzberg 是专业的非结构化文档文本提取库（支持 20+ 格式），内部自动路由：\n- PDF → pypdfium2\n- XLSX → python-calamine\n- DOCX/PPTX/MD → pandoc\n- 图片 → 内置 OCR（tesseract）\n\n## 关键词库（中英双语）\n\n**中文（47个）：** 房产、房价、房地产、居民、消费、股市、经济、政策、利率、通胀、人民币、A股、美联储、PBOC、GDP、股票、资产、投资、债券、银行、PPI、CPI、PMI、M2、就业、失业、汽车、新能源、AI 等\n\n**英文（70+个）：** property、real estate、GDP、inflation、CPI、PPI、PMI、PBOC、Fed、consumer、economy、growth、housing、stock market、EV、AI 等\n\n**标签输出语言：** 自动判断——英文内容匹配英文关键词输出英文标签，中文内容匹配中文关键词输出中文标签\n\n## 定时任务兼容性\n\n| 任务 | ID | 调用方式 | 结论 |\n|------|------|---------|------|\n| 23:00分析新文档 | f3536e18 | 绝对路径 `python3 run_analysis.py` | ✅ 无需修改 |\n| 07:00生成财经早报 | b741c6d5 | Node.js 脚本 | ❌ 不相关 |\n| 08:00发送财经早报 | a7cbaacc | 读取文件发送 | ❌ 不相关 |\n| 09:00发送文档摘要 | 89b4cf75 | 读取 summaries 目录 | ❌ 不相关 |\n\n## 迁移到新电脑\n\n1. 复制整个目录：\n   ```bash\n   scp -r ~/.openclaw/workspace/knowledge user@new-mac:~/.openclaw/workspace/\n   ```\n2. 在新电脑运行 `bash setup.sh <飞书用户ID>`\n3. 重新注册定时任务（Job ID 会变）\n\n## 避坑指南\n\n| 问题 | 原因 | 解决 |\n|------|------|------|\n| LibreOffice 超时 | watchdog 误判大文件为卡死 | v2.0 移除 watchdog，60秒硬超时 |\n| .doc 提取慢 | 统一走 LibreOffice | antiword 专线，169MB 文件 0.02秒 |\n| DOCX/PPTX 处理失败 | pandoc 未安装 | `brew install pandoc` |\n| PDF 提取乱码（OCR 不工作） | ocrmypdf `--skip-text` 跳过乱码页 | v1.4 改为 `--force-ocr` 强制 OCR |\n| PDF 漏检 CMap 残缺 | 没主动判断是否乱码 | v1.4 `is_cmap_broken()` 自检（阈值 0.03）|\n| 文本被截断到 8000 字 | 硬编码 `[:8000]` 太短 | v1.4 `MAX_EXTRACT_LEN = 500_000` |\n| .doc 提取失败 | lightpdf 处理过的 .doc 乱码 | v1.4 `ocr_office_via_ocr()` 兜底 |\n| 静默失败（不知道哪个文件）| 不抛异常 | v1.4 `PDFExtractError` / `ExtractError` 含路径 |\n| 飞书无 exec 工具 | tools 策略限制 | 添加 `alsoAllow: [exec, process]` |\n| BGE-M3 卡顿 | 16GB 内存不足 | 继续用 nomic-embed-text |\n\n## 关键路径\n\n| 内容 | 路径 |\n|------|------|\n| Skill 目录 | `~/.openclaw/workspace/skills/knowledge-base-setup/` |\n| 知识库 | `~/.openclaw/workspace/knowledge/` |\n| 分析脚本 | `~/.openclaw/workspace/knowledge/.analysis/` |\n| 目录缓存 | `~/.openclaw/workspace/knowledge/.analysis/.catalog_cache.json` |\n| 摘要输出 | `~/.openclaw/workspace/knowledge/.analysis/summaries/` |\n| 文章目录 | `~/.openclaw/workspace/knowledge/文章目录/文章目录.md` |\n| OpenClaw 配置 | `~/.openclaw/openclaw.json` |\n\n## 版本历史\n\n| 版本 | 日期 | 更新内容 |\n|------|------|---------|\n| 1.0.0 | 2026-05-10 | 初始版本，PyMuPDF + LibreOffice 链路 |\n| 1.1.0 | 2026-05-13 | 三步 PDF 处理，关键词库，双语标签 |\n| 1.2.0 | 2026-05-21 | 分批处理优化，280秒断点 |\n| 1.2.1 | 2026-05-22 | utils.py 共享模块重构，LibreOffice 熔断机制 |\n| 1.3.0 | 2026-05-28 | kreuzberg 统一提取层 + antiword 专线 + pandoc |\n| **1.4.0** | **2026-08-12** | **CMap 残缺度自检 + 50万字完整提取 + OCR fallback 到 .doc** |\n\nFile v1.4.0:_meta.json\n\n{\n  \"ownerId\": \"kn7d38vmg57htam960qsj7wnkh86ewjb\",\n  \"slug\": \"macmini-knowledge-base\",\n  \"version\": \"1.4.0\",\n  \"publishedAt\": 1786538723816\n}\n\nFile v1.4.0:CHANGELOG.md\n\n# Changelog\n\n## 1.4.0 (2026-08-12)\n\n### Changed\n- **utils.py**: `--skip-text` → `--force-ocr`（OCR 路径真正工作，不再跳过乱码页）\n- **utils.py**: OCR 超时 60s → 600s（支持大文档）\n- **utils.py**: 文本提取上限 `[:8000]` → `MAX_EXTRACT_LEN = 500_000`（完整阅读，50万字）\n\n### Added\n- **utils.py**: 新增 `is_cmap_broken()` CMap 残缺度自检函数（阈值 0.03）\n- **utils.py**: 新增 `PDFExtractError` 异常类（含文件路径）\n- **utils.py**: 新增 `ExtractError` 通用异常类（含文件路径）\n- **utils.py**: 新增 `ocr_office_via_ocr()` 通用 OCR 函数（soffice 转 PDF + OCRmyPDF）\n- **utils.py**: 新增 `CMAP_BAD_RATIO_THRESHOLD` / `CMAP_CID_COUNT_THRESHOLD` / `CMAP_SAMPLE_CHARS_MIN` 常量\n- **utils.py**: 新增 `MAX_EXTRACT_LEN` 常量\n- **scripts/**: 新增 `re_ocr_corrupted.py`（批量 OCR 修复脚本）\n- **SKILL.md**: 新增\"CMap 残缺度自检\"章节\n- **SKILL.md**: 新增 OCR 性能实测数据\n- **SKILL.md**: 避坑指南增加 5 条 v1.4 修复条目\n\n### Impact\n- CMap 残缺（真实乱码）: 40 → 0\n- 8-9.doc 类 .doc 乱码: 1 → 0\n- 文档完整度: 8000 字 → 500,000 字上限\n- OCR 真实工作率: 0% → 100%\n\n## 1.3.0 (2026-05-28)\n\n### Added\n- kreuzberg 统一提取层（PDF/DOCX/XLSX/PPTX/MD/图片 OCR）\n- antiword 极速专线（.doc 文件专用）\n- pandoc 依赖\n- 智能兜底（antiword 失败自动走 soffice）\n\n## 1.2.1 (2026-05-22)\n\n### Changed\n- utils.py 共享模块重构\n- LibreOffice 熔断机制（进程树超时杀尽）\n\n## 1.2.0 (2026-05-21)\n\n### Added\n- 分批处理优化，280 秒断点\n\n## 1.1.0 (2026-05-13)\n\n### Added\n- 三步 PDF 处理\n- 关键词库中英双语\n\n## 1.0.0 (2026-05-10)\n\n### Added\n- 初始版本\n- PyMuPDF + LibreOffice 链路\n\nFile v1.4.0:skill-card.md\n\n## Description:\n\nGuides users through setting up a local knowledge base and RAG-style natural language search workflow on a Mac Mini (M4), including document parsing, OCR fallback, scheduled analysis, and OpenClaw configuration.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[seairteng](https://clawhub.ai/user/seairteng)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers and technical users use this skill to configure a Mac Mini-based local knowledge base, extract text from common document formats, generate searchable summaries and catalogs, and schedule recurring OpenClaw analysis workflows.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: The skill asks users to allow broad OpenClaw exec/process capability and run local setup commands.\n\nMitigation: Review each command before execution, grant command capability only in the intended OpenClaw environment, and run manual setup steps selectively when possible.\n\nRisk: Scheduled jobs may persist after installation and can send document-derived summaries to Feishu.\n\nMitigation: Verify the cron jobs, recipient user ID, schedule, and summary content before enabling Feishu delivery.\n\nRisk: Extracted document text and generated summaries are stored in plaintext under the local knowledge directory.\n\nMitigation: Use appropriate local file permissions, avoid processing documents that should not be stored as plaintext, and back up summaries before running OCR repair operations.\n\nRisk: The setup flow installs external system and Python dependencies.\n\nMitigation: Install dependencies from trusted package sources and review the dependency list before running the one-command setup script.\n\n## Reference(s):\n\n- [ClawHub skill page](https://clawhub.ai/seairteng/skills/macmini-knowledge-base)\n- [Ollama download](https://ollama.com/download)\n\n## Skill Output:\n\n**Output Type(s):** [text, markdown, code, shell commands, configuration, guidance]\n\n**Output Format:** [Markdown guidance with bash commands, JSON configuration snippets, and local script usage examples]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [The described workflow can create local analysis scripts, persistent cron jobs, extracted-text summaries, catalog files, and OCR repair reports when executed.]\n\n## Skill Version(s):\n\n1.4.0 (source: frontmatter, changelog, server release evidence)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nArchive v1.3.0: 7 files, 15604 bytes\n\nFiles: scripts/generate_catalog.py (12426b), scripts/run_analysis.py (5921b), scripts/setup.sh (3966b), scripts/utils.py (8252b), skill-card.md (2326b), SKILL.md (9860b), _meta.json (141b)\n\nFile v1.3.0:SKILL.md\n\n---\nname: macmini-knowledge-base\nversion: 1.3.0\ndescription: |\n  在 Mac Mini (M4) 上快速搭建本地知识库 + RAG 自然语言搜索系统。\n  适用场景：\n  - 新 Mac 配置知识库：从零开始安装配置 Ollama、embedding模型、定时任务、文档解析\n  - 遇到 PDF 提取乱码、定时任务超时、skill 加载失败等问题\n  - 想要建立每日自动分析文档 + 08:00发送摘要到飞书的流程\n  - 迁移或复现知识库：打包整个 knowledge 目录和配置到新电脑\n  本 skill 会引导完成：目录结构创建、依赖安装、脚本部署、定时任务注册、OpenClaw 配置。\n---\n\n# Knowledge Base Setup\n\n在 Mac Mini 上快速搭建本地知识库 + RAG 搜索系统。\n\n## 核心功能（v2.0）\n\n- **kreuzberg 统一提取层**：PDF / DOCX / XLSX / PPTX / MD / 图片 OCR 全自动路由\n- **antiword 极速专线**：.doc 文件专用提取，成功率 85%，169MB 文件 0.02 秒完成\n- **智能兜底**：antiword 失败自动走 soffice 转换，60 秒硬超时无误判\n- **自动分类**：关键词匹配驱动，中英文双语标签\n- **定时任务**：每天 23:00 分析新文档，08:00 发送摘要到飞书\n\n## 快速开始\n\n### 一键安装\n\n```bash\ncd ~/.openclaw/workspace/skills/knowledge-base-setup/scripts\nbash setup.sh <飞书用户ID>\n```\n\n### 手动分步安装\n\n**Step 1: 系统依赖**\n```bash\nbrew install antiword tesseract pandoc\n```\n\n**Step 2: Python 依赖**\n```bash\npip3 install kreuzberg pytesseract pymupdf docx openpyxl python-pptx\n```\n\n**Step 3: Ollama + embedding 模型**\n```bash\n# 安装 Ollama: https://ollama.com/download\nollama pull nomic-embed-text\n```\n\n**Step 4: 创建目录结构**\n```bash\nmkdir -p ~/.openclaw/workspace/knowledge/.analysis/summaries/archives\nmkdir -p ~/.openclaw/workspace/knowledge/temp_docs\ntouch ~/.openclaw/workspace/knowledge/文章目录/文章目录.md\n```\n\n**Step 5: 部署脚本**\n```bash\ncp ~/.openclaw/workspace/skills/knowledge-base-setup/scripts/*.py \\\n   ~/.openclaw/workspace/knowledge/.analysis/\nchmod +x ~/.openclaw/workspace/knowledge/.analysis/*.py\n```\n\n**Step 6: 配置 OpenClaw**\n\n编辑 `~/.openclaw/openclaw.json`，加入：\n```json\n{\n  \"models\": {\n    \"providers\": {\n      \"ollama\": {\n        \"baseUrl\": \"http://127.0.0.1:11434\",\n        \"api\": \"ollama\",\n        \"models\": [\n          {\"id\": \"nomic-embed-text\", \"name\": \"Nomic Embed Text\"}\n        ]\n      }\n    }\n  },\n  \"agents\": {\n    \"defaults\": {\n      \"memorySearch\": {\n        \"provider\": \"ollama\",\n        \"model\": \"nomic-embed-text\"\n      }\n    }\n  }\n}\n```\n\n确保 tools 区块有：\n```json\n\"tools\": {\n    \"alsoAllow\": [\"exec\", \"process\"]\n}\n```\n\n然后重启：`openclaw gateway restart`\n\n**Step 7: 注册定时任务**\n```bash\n# 23:00 分析新文档\nopenclaw cron add \\\n  --name \"23:00分析新文档\" \\\n  --cron \"0 23 * * *\" \\\n  --tz \"Asia/Shanghai\" \\\n  --session isolated \\\n  --timeout-seconds 600 \\\n  --message \"cd ~/.openclaw/workspace/knowledge/.analysis && python3 run_analysis.py && python3 generate_catalog.py\" \\\n  --announce --channel feishu --to \"user:<飞书用户ID>\"\n\n# 08:00 发送文档摘要\nopenclaw cron add \\\n  --name \"08:00发送文档摘要\" \\\n  --cron \"0 8 * * *\" \\\n  --tz \"Asia/Shanghai\" \\\n  --session isolated \\\n  --timeout-seconds 120 \\\n  --message \"读取 summaries/ 目录发送摘要到飞书\" \\\n  --announce --channel feishu --to \"user:<飞书用户ID>\"\n```\n\n## 文档解析架构（v2.0）\n\n### 架构图\n\n```\n                    ┌──────────────────────────────────────┐\n                    │         kreuzberg 统一提取层           │\n                    │  (pypdfium2 / python-calamine / pandoc) │\n                    └───┬────────────────────────────────┬───┘\n                        │                              │\n                自动判断 │                              │\n                        ▼                              ▼\n              ┌─────────────────┐           ┌─────────────────────┐\n              │  kreuzberg 直提  │           │  antiword 极速专线  │\n              │ PDF/DOCX/XLSX/  │           │   (.doc 文件专用)    │\n              │ PPTX/MD/图片OCR │           │   成功率 85%，<1秒   │\n              └─────────────────┘           └─────────────────────┘\n                        │                              │\n                        │         ┌──────────────────────────────┐\n                        │         │     soffice 兜底转换          │\n                        │         │ (.doc/.xls/.ppt antiword失败) │\n                        │         │  60秒硬超时（消除误判watchdog）│\n                        │         └──────────────────────────────┘\n                        ▼                              │\n              ┌──────────────────────────────────────────────┐\n              │              文本输出（content）              │\n              │  → summaries/ 摘要文件 → generate_catalog.py  │\n              └──────────────────────────────────────────────┘\n```\n\n### 文件类型 × 提取方式\n\n| 格式 | 主方案 | 依赖 | 成功率 | 单文件速度 |\n|------|--------|------|--------|-----------|\n| PDF | kreuzberg (pypdfium2) | 无 | ~100% | 0.05-0.7s |\n| DOCX | kreuzberg + pandoc | pandoc 3.9+ | 100% | 0.12-3s |\n| XLSX | kreuzberg (python-calamine) | 无 | 100% | 0.1-0.5s |\n| PPTX | kreuzberg + pandoc | pandoc 3.9+ | 100% | 0.02-0.2s |\n| MD | kreuzberg + pandoc | pandoc 3.9+ | 100% | <0.01s |\n| **.doc** | **antiword 优先** | antiword | **85%**，<1秒 | <0.02s |\n| .doc（失败） | soffice 兜底 | LibreOffice | ~15% | 2-21s |\n| .xls | soffice → XLSX | LibreOffice | ~95% | 2-10s |\n| .ppt | soffice → PPTX | LibreOffice | ~95% | 2-10s |\n| 图片 | kreuzberg 内置 OCR | tesseract | ~90% | 3-10s |\n\n### antiword 极速专线\n\n```python\n# 实测数据：\n# 169MB 超大文件 → 26万字符，0.02秒完成\n# 正常 .doc（0.1-15MB）→ <1秒\n# 成功率 85%，覆盖绝大多数 .doc 文件\nresult = subprocess.run(['antiword', filepath], capture_output=True, timeout=10)\n```\n\n### kreuzberg 统一提取层\n\nkreuzberg 是专业的非结构化文档文本提取库（支持 20+ 格式），内部自动路由：\n- PDF → pypdfium2\n- XLSX → python-calamine\n- DOCX/PPTX/MD → pandoc\n- 图片 → 内置 OCR（tesseract）\n\n## 关键词库（中英双语）\n\n**中文（47个）：** 房产、房价、房地产、居民、消费、股市、经济、政策、利率、通胀、人民币、A股、美联储、PBOC、GDP、股票、资产、投资、债券、银行、PPI、CPI、PMI、M2、就业、失业、汽车、新能源、AI 等\n\n**英文（70+个）：** property、real estate、GDP、inflation、CPI、PPI、PMI、PBOC、Fed、consumer、economy、growth、housing、stock market、EV、AI 等\n\n**标签输出语言：** 自动判断——英文内容匹配英文关键词输出英文标签，中文内容匹配中文关键词输出中文标签\n\n## 定时任务兼容性\n\n| 任务 | ID | 调用方式 | 结论 |\n|------|------|---------|------|\n| 23:00分析新文档 | f3536e18 | 绝对路径 `python3 run_analysis.py` | ✅ 无需修改 |\n| 07:00生成财经早报 | b741c6d5 | Node.js 脚本 | ❌ 不相关 |\n| 08:00发送财经早报 | a7cbaacc | 读取文件发送 | ❌ 不相关 |\n| 09:00发送文档摘要 | 89b4cf75 | 读取 summaries 目录 | ❌ 不相关 |\n\n## 迁移到新电脑\n\n1. 复制整个目录：\n   ```bash\n   scp -r ~/.openclaw/workspace/knowledge user@new-mac:~/.openclaw/workspace/\n   ```\n2. 在新电脑运行 `bash setup.sh <飞书用户ID>`\n3. 重新注册定时任务（Job ID 会变）\n\n## 避坑指南\n\n| 问题 | 原因 | 解决 |\n|------|------|------|\n| LibreOffice 超时 | watchdog 误判大文件为卡死 | v2.0 移除 watchdog，60秒硬超时 |\n| .doc 提取慢 | 统一走 LibreOffice | antiword 专线，169MB 文件 0.02秒 |\n| DOCX/PPTX 处理失败 | pandoc 未安装 | `brew install pandoc` |\n| PDF 提取乱码 | 自定义字体无 ToUnicode | kreuzberg(pypdfium2) + tesseract OCR |\n| 飞书无 exec 工具 | tools 策略限制 | 添加 `alsoAllow: [exec, process]` |\n| BGE-M3 卡顿 | 16GB 内存不足 | 继续用 nomic-embed-text |\n\n## 关键路径\n\n| 内容 | 路径 |\n|------|------|\n| Skill 目录 | `~/.openclaw/workspace/skills/knowledge-base-setup/` |\n| 知识库 | `~/.openclaw/workspace/knowledge/` |\n| 分析脚本 | `~/.openclaw/workspace/knowledge/.analysis/` |\n| 目录缓存 | `~/.openclaw/workspace/knowledge/.analysis/.catalog_cache.json` |\n| 摘要输出 | `~/.openclaw/workspace/knowledge/.analysis/summaries/` |\n| 文章目录 | `~/.openclaw/workspace/knowledge/文章目录/文章目录.md` |\n| OpenClaw 配置 | `~/.openclaw/openclaw.json` |\n\n## 版本历史\n\n| 版本 | 日期 | 更新内容 |\n|------|------|---------|\n| 1.0.0 | 2026-05-10 | 初始版本，PyMuPDF + LibreOffice 链路 |\n| 1\n\nArchive v1.2.1: 6 files, 15932 bytes\n\nFiles: scripts/generate_catalog.py (11640b), scripts/run_analysis.py (7067b), scripts/setup.sh (4393b), scripts/utils.py (8921b), SKILL.md (8642b), _meta.json (141b)","readmeExcerpt":"Skill: Mac 知识库搭建系统 Owner: seairteng Summary: ⚠️ **本 skill 包含以下高危能力，使用前请仔细阅读 SKILL.md 顶部「⚠️ CAPABILITIES & RISKS」章节**： - Shell 执行 + 文件读写 + 安装 Homebrew 包（版本固定）+ 下载 Ollama 模型 - 修改 OpenClaw 配置 + 注册持久 cron 任务 + 飞书 webhook 推送 在 Mac Mini (M4) 上快速搭建本地知识库 + RAG 自然语言搜索系统。 适用场景： - 新 Mac 配置知识库：从零开始安装配置 Ollama、embedding模型、定时任务、文档解析 - 遇到 PDF 提取乱码、定时任务超时、skill 加载失败等问题 - 想要建立每日自动分析文档 + 08:00发送摘要到飞书的流程 - 迁移或复现知识库：打包整个 knowledge 目录和配置","codeSnippets":[],"executableExamples":[{"language":"json","snippet":"{\n  \"tools\": {\n    \"alsoAllow\": [\"exec\", \"process\", \"read\", \"write\"]\n  }\n}"},{"language":"bash","snippet":"cd ~/.openclaw/workspace/skills/knowledge-base-setup/scripts\nbash setup.sh <飞书用户ID>"},{"language":"bash","snippet":"brew install antiword tesseract pandoc"},{"language":"bash","snippet":"pip3 install kreuzberg pytesseract pymupdf docx openpyxl python-pptx"},{"language":"bash","snippet":"# 安装 Ollama: https://ollama.com/download\nollama pull nomic-embed-text"},{"language":"bash","snippet":"mkdir -p ~/.openclaw/workspace/knowledge/.analysis/summaries/archives\nmkdir -p ~/.openclaw/workspace/knowledge/temp_docs\ntouch ~/.openclaw/workspace/knowledge/文章目录/文章目录.md"}],"parameters":null,"dependencies":[],"permissions":[],"extractedFiles":[{"path":"SKILL.md","content":"---\nname: macmini-knowledge-base\nversion: 1.4.7\ndescription: |\n  ⚠️ **本 skill 包含以下高危能力，使用前请仔细阅读 SKILL.md 顶部「⚠️ CAPABILITIES & RISKS」章节**：\n  - Shell 执行 + 文件读写 + 安装 Homebrew 包（版本固定）+ 下载 Ollama 模型\n  - 修改 OpenClaw 配置 + 注册持久 cron 任务 + 飞书 webhook 推送\n\n  在 Mac Mini (M4) 上快速搭建本地知识库 + RAG 自然语言搜索系统。\n  适用场景：\n  - 新 Mac 配置知识库：从零开始安装配置 Ollama、embedding模型、定时任务、文档解析\n  - 遇到 PDF 提取乱码、定时任务超时、skill 加载失败等问题\n  - 想要建立每日自动分析文档 + 08:00发送摘要到飞书的流程\n  - 迁移或复现知识库：打包整个 knowledge 目录和配置到新电脑\n  - **v1.4 新增**：CMap 残缺度自检（不预设来源）+ 50万字完整提取 + OCR fallback 到 .doc\n  本 skill 会引导完成：目录结构创建、依赖安装、脚本部署、定时任务注册、OpenClaw 配置。\n\n  ⚠️ **重要：能力范围**\n  本 skill 不只是「搭建」，还包含：\n  - 批量 OCR 修复（扫描 summaries/archives 找乱码 + 重新提取）\n  - 目录归档清理（移动重复/孤儿文件到 .trash/）\n  - 自动定时任务（23:00 分析 + 06:00 飞书推送）\n  \n  使用前请仔细评估批量修改风险。\n---\n\n# Knowledge Base Setup\n\n在 Mac Mini 上快速搭建本地知识库 + RAG 搜索系统。\n\n## ⚠️ 阅读前必读：本 skill 的能力范围\n\n本 skill **不只是\"搭建知识库\"**，还包含以下高危能力：\n\n**执行能力**：\n- 🔧 **Shell 命令执行**（python3 + bash 脚本）\n- 📁 **文件读写**（knowledge/, summaries/, archives/, .trash/）\n- 📦 **安装 Homebrew 包**（antiword, tesseract, pandoc, libreoffice，**版本固定**）\n- ⬇️ **下载 Ollama 模型**（nomic-embed-text, ~274MB）\n- ⚙️ **修改 OpenClaw 配置**（~/.openclaw/openclaw.json）\n- ⏰ **注册持久 cron 任务**（23:00 + 06:00，**每天自动**）\n- 📤 **推送消息到飞书 webhook**\n\n**持久化影响**：\n- 知识库目录会被自动分析（每天 23:00）\n- 摘要文件会被覆盖写入（OCR 修复时）\n- cron 任务永久执行（直到手动 `openclaw cron remove <id>`）\n- 失败文件移到 `.trash/`（7 天兜底清理）\n\n**安装流程**：\n本 skill 的 setup.sh 是**交互式安装向导**：\n- 每个危险操作前会要求 y/N 确认\n- 提供 `--dry-run` 选项查看会做什么\n- 已安装用户重跑会进入确认模式\n\n**如果不同意上述任何一项，请不要安装本 skill。**\n\n### 结构化权限声明（Structured Permissions）\n\n| 权限 | 必填 | 范围 | 用途 / 风险 |\n|------|------|------|-------------|\n| `exec` | ✅ | python3 + bash scripts | 文档提取 + 飞书推送 |\n| `file_read` | ✅ | `~/.openclaw/workspace/knowledge/` | 读取文档 + summaries |\n| `file_write` | ✅ | `summaries/`, `archives/`, `.trash/` | OCR 修复覆盖旧摘要 |\n| `install_packages` | ✅ | brew: antiword, tesseract, pandoc, libreoffice | Homebrew 包安装（用户确认）|\n| `download_model` | ✅ | ollama: nomic-embed-text (~274MB) | Ollama 模型下载（用户确认）|\n| `modify_config` | ✅ | `~/.openclaw/openclaw.json` | 添加 alsoAllow: [exec, process] |\n| `register_cron` | ✅ | 23:00 分析 + 06:00 推送 | 持久化定时任务（用户确认）|\n| `network` | ✅ | 飞书 webhook + Ollama 下载 | 外部 API 调用 |\n\n**warning**: 本 skill 会自动修改文件、安装包、注册 cron 任务（用户每步都有 y/N 确认）\n\n**disable_command**: `openclaw cron remove <id>`\n\n---\n\n## 核心功能（v2.0）\n\n- **kreuzberg 统一提取层**：PDF / DOCX / XLSX / PPTX / MD / 图片 OCR 全自动路由\n- **antiword 极速专线**：.doc 文件专用提取，成功率 85%，169MB 文件 0.02 秒完成\n- **智能兜底**：antiword 失败自动走 soffice 转换，60 秒硬超时无误判\n- **自动分类**：关键词匹配驱动，中英文双语标签\n- **定时任务**：每天 23:00 分析新文档，08:00 发送摘要到飞书\n\n\n\n## 🔧 权限声明\n\n本 skill 在使用时需要以下 OpenClaw 工具能力：\n\n```json\n{\n  \"tools\": {\n    \"alsoAllow\": [\"exec\", \"process\", \"read\", \"write\"]\n  }\n}\n```\n\n⚠️ **执行风险**：exec + cron 自动化组合可导致持续命令执行，\n请在可信环境（个人 Mac）使用，不要在共享/服务器部署。\n\n\n\n## ⚠️ 安全警告：定时任务\n\n本 skill 注册 2 个 cron 任务（23:00 分析 + 06:00 推送），\n运行 shell 命令并自动推送消息到飞书。\n\n**潜在风险**：\n- 脚本路径被修改 → 自动执行任意命令\n- 知识库目录被入侵 → 自动读取/外发\n- 飞书 webhook 泄漏 → 自动推送被劫持\n\n**建议**：\n- 不要把 `~/.openclaw/workspace/knowledge` 放在共享/多用户目录\n- 定期检查 c"},{"path":"_meta.json","content":"{\n  \"ownerId\": \"kn7d38vmg57htam960qsj7wnkh86ewjb\",\n  \"slug\": \"macmini-knowledge-base\",\n  \"version\": \"1.4.7\",\n  \"publishedAt\": 1786985191922\n}"},{"path":"CHANGELOG.md","content":"# Changelog\n\n## v1.4.7 (2026-08-17) — SKILL.md version 字段同步\n\n- 修复 SKILL.md 顶部 version 字段（1.4.5 → 1.4.7）\n- 与 _meta.json version 一致\n\n## v1.4.6 (2026-08-17) — SKILL.md 修补\n- 加\"结构化权限声明\"表格\n- 修复 SKILL.md 顶部 version（1.4.2 → 1.4.5）\n\n## v1.4.5 (2026-08-17) — NVIDIA SkillSpector 完整修复\n### 元数据修复\n### SKILL.md 修复\n### setup.sh 改造（交互式）\n\n## v1.4.4 (2026-08-17) — 仅 SKILL.md 补充\n## v1.4.3 (2026-08-17) — 代码安全修复\n## v1.4.2 (2026-08-17) — 仅元数据更新"},{"path":"skill-card.md","content":"## Description:\n\nMac 知识库搭建系统 helps agents guide setup of a local Mac Mini knowledge base and RAG search workflow, including document extraction, OCR repair, catalog generation, scheduled analysis, and Feishu summary delivery.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[seairteng](https://clawhub.ai/user/seairteng)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers and personal Mac users use this skill to configure a local knowledge-base workflow on a trusted Mac Mini, process documents into summaries and catalogs, and schedule recurring analysis plus Feishu notifications.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: Broad exec/process permissions and persistent cron jobs can allow repeated command execution from the local workspace.\n\nMitigation: Use the skill only on a trusted personal Mac, review the setup plan first, avoid broad default permission changes where possible, and remove unwanted cron entries with openclaw cron remove.\n\nRisk: The workflow reads local knowledge-base files and can send summaries through a Feishu webhook.\n\nMitigation: Keep the workspace private, use a dedicated Feishu destination, store the webhook as a secret, and rotate it if it may have been exposed.\n\nRisk: The setup process may install Homebrew and Python dependencies and download an Ollama embedding model.\n\nMitigation: Run the documented dry-run path first and install dependencies separately or from pinned trusted sources before enabling automation.\n\n## Reference(s):\n\n- [ClawHub skill page](https://clawhub.ai/seairteng/skills/macmini-knowledge-base)\n- [Ollama download](https://ollama.com/download)\n\n## Skill Output:\n\n**Output Type(s):** [text, markdown, code, shell commands, configuration, guidance]\n\n**Output Format:** [Markdown guidance with inline shell commands, JSON snippets, and Python code examples]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [The skill guides interactive setup steps and includes dry-run guidance for the setup script.]\n\n## Skill Version(s):\n\n1.4.7 (source: frontmatter, CHANGELOG, _meta.json, server release evidence)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment."}],"languages":[],"docsSourceLabel":"CLAWHUB","editorialOverview":"⚠️ **本 skill 包含以下高危能力，使用前请仔细阅读 SKILL.md 顶部「⚠️ CAPABILITIES & RISKS」章节**： - Shell 执行 + 文件读写 + 安装 Homebrew 包（版本固定）+ 下载 Ollama 模型 - 修改 OpenClaw 配置 + 注册持久 cron 任务 + 飞书 webhook 推送 在 Mac Mini (M4) 上快速搭建本地知识库 + RAG 自然语言搜索系统。 适用场景： - 新 Mac 配置知识库：从零开始安装配置 Ollama、embedding模型、定时任务、文档解析 - 遇到 PDF 提取乱码、定时任务超时、skill 加载失败等问题 - 想要建立每日自动分析文档 + 08:00发送摘要到飞书的流程 - 迁移或复现知识库：打包整个 knowledge 目录和配置到新电脑 - **v1.4 新增**：CMap 残缺度自检（不预设来源）+ 50万字完整提取 + OCR fallback 到 .doc 本 skill 会引导完成：目录结构创建、依赖安装、脚本部署、定时任务注册、OpenClaw 配置。 ⚠️ **重要：能力范围** 本 skill 不只是「搭建」，还包含： - 批量 OCR 修复（扫描 summaries/archives 找乱码 + 重新提取） - 目录归档清理（移动重复/孤儿文件到 .trash/） - 自动定时任务（23:00 分析 + 06:00 飞书推送） 使用前请仔细评估批量修改风险。 Skill: Mac 知识库搭建系统 Owner: seairteng Summary: ⚠️ **本 skill 包含以下高危能力，使用前请仔细阅读 SKILL.md 顶部「⚠️ CAPABILITIES & RISKS」章节**： - Shell 执行 + 文件读写 + 安装 Homebrew 包（版本固定）+ 下载 Ollama 模型 - 修改 OpenClaw 配置 + 注册持久 cron 任务 + 飞书 webhook 推送 在 Mac Mini (M4) 上快速搭建本地知识库 + RAG 自然语言搜索系统。 适用场景： - 新 Mac 配置知识库：从零开始安装配置 Ollama、embedding模型、定时任务、文档解析 - 遇到 PDF 提取乱码、定时任务超时、skill 加载失败等问题 - 想要建立每日自动分析文档 + 08:00发送摘要到飞书的流程 - 迁移或复现知识库：打包整个 knowledge 目录和配置","editorialQuality":{"score":100,"threshold":65,"status":"ready","wordCount":967,"uniquenessScore":49,"reasons":[]}},"media":{"evidence":{"source":"no-media","verified":false,"confidence":"low","updatedAt":"2026-10-10T09:13:09.914Z","emptyReason":"No screenshots, media assets, or demo links are available."},"primaryImageUrl":null,"mediaAssetCount":0,"assets":[],"demoUrl":null},"ownerResources":{"evidence":{"source":"unclaimed","verified":false,"confidence":"low","updatedAt":"2026-10-10T09:13:09.914Z","emptyReason":"This page has not been claimed by the agent owner."},"hasCustomPage":false,"customPageUpdatedAt":null,"customLinks":[],"structuredLinks":{"docsUrl":null,"demoUrl":null,"supportUrl":null,"pricingUrl":null,"statusUrl":null},"customPage":null},"relatedAgents":{"evidence":{"source":"protocol-neighbors","verified":false,"confidence":"medium","updatedAt":"2026-10-10T11:50:24.052Z","emptyReason":null},"items":[{"id":"8ebccd8e-3863-4187-8355-c3f14e1f9edf","entityType":"agent","canonicalPath":"/agent/iofficeai-aionui","slug":"iofficeai-aionui","name":"AionUi","description":"Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!","url":"https://github.com/iOfficeAI/AionUi","homepage":"https://www.aionui.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-10-09T19:11:12.944Z","createdAt":"2026-02-25T03:38:16.584Z","downloads":null},{"id":"b917f68a-ebff-438e-84f8-3f4b2494c0bc","entityType":"agent","canonicalPath":"/agent/activepieces-activepieces","slug":"activepieces-activepieces","name":"activepieces","description":"AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents","url":"https://github.com/activepieces/activepieces","homepage":"https://www.activepieces.com","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-15T02:22:12.426Z","createdAt":"2026-02-25T03:38:12.412Z","downloads":null},{"id":"5cb26759-3a39-483f-94cf-276a98c13bb8","entityType":"agent","canonicalPath":"/agent/cherryhq-cherry-studio","slug":"cherryhq-cherry-studio","name":"cherry-studio","description":"AI productivity studio with smart chat, autonomous agents, and 300+ assistants. Unified access to frontier LLMs","url":"https://github.com/CherryHQ/cherry-studio","homepage":"https://cherry-ai.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-11T14:38:40.986Z","createdAt":"2026-02-25T03:38:19.379Z","downloads":null},{"id":"6f6582d0-5d76-4f0f-b81d-86520247950b","entityType":"agent","canonicalPath":"/agent/copilotkit-copilotkit","slug":"copilotkit-copilotkit","name":"CopilotKit","description":"The Frontend for Agents & Generative UI. React + Angular","url":"https://github.com/CopilotKit/CopilotKit","homepage":"https://docs.copilotkit.ai","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-03-25T09:50:57.846Z","createdAt":"2026-02-25T03:39:14.617Z","downloads":null}],"links":{"hub":"/agent","source":"/agent/source/clawhub","protocols":[{"label":"OpenClaw","href":"/agent/protocol/openclew"}]}}}