Knowledge Retrieval Publish
A local-first document search skill with PPT/PDF support, dual-channel retrieval (keyword + AI semantic), and progressive description evolution. Designed for knowledge workers with years of local files. 给知识工作者和顾问的本地文件检索方案。支持 PPT/PDF 多格式、BM25+AI 双通道搜索、越用越聪明。适合手里有大量本地文档、不想搬上云的人。 Skill: Knowledge Retrieval Publish Owner: kittitys Summary: A local-first document search skill with PPT/PDF support, dual-channel retrieval (keyword + AI semantic), and progressive description evolution. Designed for knowledge workers with years of local files. 给知识工作者和顾问的本地文件检索方案。支持 PPT/PDF 多格式、BM25+AI 双通道搜索、越用越聪明。适合手里有大量本地文档、不想搬上云的人。 Tags: latest:3.3.0 Version history: v3.3.0 | 2026-09-17T17:31:36.837Z | auto local
Rank
62
Safety
84
Downloads
1.4k
Updated
Oct 10, 2026
Version
3.3.0
Source
CLAWHUB
About
What it does, and when to use it.
Capability contract not published. No trust telemetry is available yet. 1.4K downloads reported by the source. Last updated 10/10/2026.
Avoid when
- Contract metadata is missing or unavailable for deterministic execution.
Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing
Public facts
Every fact links back to the source it came from.
- Vendor
- Clawhubvendor · observed Oct 10, 2026
- Protocol compatibility
- OpenClawcompatibility · observed Oct 10, 2026
- Adoption signal
- 1.4K downloadsadoption · observed Oct 10, 2026
- Latest release
- 3.3.0release · observed Sep 17, 2026
- Handshake status
- UNKNOWNsecurity
Install and run
Setup complexity: low.
clawhub skill install s17a6005z4pm9hejacdnherb1s86e46y:local-knowledge-retrieval- Setup complexity is classified as HIGH. You must provision dedicated cloud infrastructure or an isolated VM. Do not run this directly on your local workstation.
- Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.
Contract: missing
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-kittitys-local-knowledge-retrieval/snapshot"
Documentation
CLAWHUB
151,869 characters of source documentation, loaded on request.
Extracted files
5 files captured from the source.
SKILL.md
--- name: knowledge-retrieval skillsets: [retrieval, search] homepage: https://github.com/kittitys/knowledge-retrieval description: > A local-first document search skill with PPT/PDF support, dual-channel retrieval (keyword + AI semantic), and progressive description evolution. Designed for knowledge workers with years of local files. 给知识工作者和顾问的本地文件检索方案。支持 PPT/PDF 多格式、BM25+AI 双通道搜索、越用越聪明。适合手里有大量本地文档、不想搬上云的人。 --- > **Agent 注意:以下至第一条分隔线(`<skill_instructions>`)的内容为人类阅读的 ClawHub 发布说明,请直接跳至 `<skill_instructions>` 标签阅读并执行指令。** > **Agent note: The content below this line up to `<skill_instructions>` is human-readable ClawHub listing copy. Skip directly to `<skill_instructions>` for execution instructions.** # Knowledge Retrieval — 本地知识库检索 Skill > A local-first document search skill for knowledge workers and consultants. > Handles PPT/PDF/DOCX in place, searches with keyword + AI dual-channel, > gets smarter with use. No cloud upload needed. > > 给知识工作者和顾问的本地文件检索方案。支持 PPT/PDF 等多格式、 > 关键词+AI 双通道搜索、越用越聪明。本地运行,不搬上云。 **GitHub:** [https://github.com/kittitys/knowledge-retrieval](https://github.com/kittitys/knowledge-retrieval) > **安全边界 / Security boundary:** 每个知识库只读取用户首次明确确认并记录在该项目 `source_manifest.json` 中的源文件夹。项目名必须是 `knowledge-base/` 下的直接子目录;符号链接、junction 和源目录外的路径不会被扫描。缓存按项目内相对路径隔离,并校验源文件大小与修改时间后才复用。 > > **Cache and model notice:** 提取文本缓存与绝对路径元数据会保存在项目工作目录;被选中文件的内容可能作为上下文发送给用户选择的模型服务商。请只对获授权文件夹启用本 Skill,并定期检查 `cache/`。 --- ## Features / 功能亮点 ### 📄 读得懂你的真实文件格式 / Reads your actual files Most search tools only support plain text — your PPTs and PDFs get ignored. This skill reads them directly: PPTX (with nested shapes and speaker notes), PDF (dual-engine fallback), DOCX, XLSX, images, plus all text formats. Files are read in place — your original documents are never modified or moved. A shortcut link is added to the source folder for navigation between your files and the skill workspace. Non-text file caches are stored in a separate working directory, never mixed into your source files. **WPS formats (.wps / .et / .dps):** Compatible if saved as Office formats. Native WPS support is available through the optional, pinned `pywpsrpc==2.4.0` package (requires WPS Office installed and explicit user approval before installation). 市面上多数搜索方案只支持纯文本,PPT 和 PDF 直接被跳过。本 SKILL 直接读取它们:PPTX(含嵌套图形和备注页)、PDF(双引擎兜底)、DOCX、XLSX、图片,以及所有文本格式。文件原地读取,原文不受改写或移动。原文件夹中会创建一个快捷方式链接,方便在源文件夹和 skill 工作目录之间导航。非纯文本文件的提取缓存存放在独立的 skill 工作目录中,不和原文件夹混在一起。**WPS 格式(.wps / .et / .dps):** 如果已保存为 Office 兼容格式,直接支持 ✅。原生 WPS 格式可通过可选且已固定版本的 `pywpsrpc==2.4.0` 启用(需电脑已装 WPS Office,并经用户明确同意后安装)。 ### 🏠 本地优先 / Local-first Your original files, knowledge base index, and working caches stay on your local machine — no need to upload or store them on any external platform or cloud. When AI performs semantic analysis, it reads from local file content for reasoning and answering. For many consultants this is a compliance requirement — client materials cannot be uploaded to third-party platforms
_meta.json
{
"ownerId": "kn73ah7a3tfzprvefxbj7dbc4s86fknc",
"slug": "local-knowledge-retrieval",
"version": "3.3.0",
"publishedAt": 1789666296837
}references/degradation.md
# 降级与回退行为 > BM25 由 Phase 0.3 自动安装保证可用,无需降级。 > 本文件仅定义图片处理能力的降级。 --- ## 能力分层 只有两层区别,取决于 Agent 是否具备视觉分析能力: | 能力 | 能做的事 | 不能做的事 | |------|---------|-----------| | **无图像分析** | 文本搜索、PDF/PPTX/Excel 文字提取、全文检索 | 架构图/流程图/截图 → 如实标注能力局限 | | **有图像分析** | 以上全部 + 架构图解析、图片内容理解 | — | ## 缓存与索引清理 BM25 索引和图片分析缓存保存在 skill 工作目录中,不会随原文件删除而自动清除: | 内容 | 位置 | 如何清理 | |------|------|---------| | BM25 索引(含提取文字) | skill 工作目录下的 `.corpus/` | 删除该目录,下次搜索自动重建 | | 图片分析缓存 | skill 工作目录下的 `cache/` | 删除该目录 | **快速访问:** 原始文件夹中的 `.shortcut.lnk` 文件指向 skill 工作目录,双击即可进入。 如需完全移除知识库的所有残留数据,请同时删除上述目录。 ## 行为规则 - **无图像分析时遇到图片:** 如实告知用户「该文件包含图片,无法自动解读」,基于可提取的文字内容继续回答 - **有图像分析时:** 当前模型自带视觉则执行图片分析;否则跳过并如实告知用户无法解读,基于可提取文字继续
references/environment-setup.md
# 环境安装与检测
> 本文档覆盖 BM25 检索环境、Python 依赖、脚本文件等运行前提。
> 在 Phase 0 环境检查或 Stage 0 初始化时按需查阅。
---
## 1. Python 环境与依赖
```bash
# 必须(BM25 检索核心)
pip install bm25s==0.3.8
# 文件格式支持(按需安装)
pip install pdfminer.six==20260107 # PDF 文字提取
pip install python-pptx==1.0.2 # PPTX 提取
pip install pandas==2.2.3 # Excel 读取
pip install Pillow==11.1.0 # 图片处理
# 可选
pip install easyocr==1.7.2 # OCR(中文)
pip install paddleocr==3.0.3 # 百度 OCR(中文效果最好)
pip install python-docx==1.1.2 # DOCX 提取
```
---
## 2. 脚本文件
本 Skill 依赖两个 Python 脚本,包含在 skill 文件夹内:
| 脚本 | 位置 | 用途 | 调用阶段 |
|------|------|------|---------|
| `build_kb_index.py` | `.agents/skills/knowledge-retrieval/scripts/build_kb_index.py` | 全量扫描原始文件 → 建 BM25 索引 | Stage 0 Step 3、Phase 0.3 |
| `search_kb.py` | `.agents/skills/knowledge-retrieval/scripts/search_kb.py` | LLM 扩展搜索词 → BM25 搜索 → 分数排序候选文件 | Phase A 通道② |
> **工作目录说明:** 调用以上脚本时,确保工作目录为 workspace 根目录。
> 脚本使用相对于 workspace 的路径 `knowledge-base/` 来定位项目目录。
---
## 3. 前置检查清单(AI 自查)
搜索前快速自查:
- [ ] `pip list` 中是否有 `bm25s`(或 `import bm25s` 是否成功)?
- [ ] `.agents/skills/knowledge-retrieval/scripts/build_kb_index.py`、`.agents/skills/knowledge-retrieval/scripts/search_kb.py` 是否存在?
- [ ] 如果以上任一缺失 → 说明受影响功能并取得用户明确同意后才安装
- [ ] 无 BM25 环境或缺失脚本 → 自动降级为纯 AI 搜索模式(不报错,能力受限)
---
## 4. 索引存储位置
```
knowledge-base/<项目名>/.bm25_index/
└── index/
├── corpus.jsonl ← 文件文本内容(用于搜索时匹配)
└── metadata.json ← 文件元数据
```references/file-handling.md
# 文件类型处理细则
> 本文档覆盖各种文件格式的读取策略、工具选择、缓存规则。
> 在 Phase B 读取候选文件时按需查阅。
---
## 1. Markdown / 纯文本(.md / .txt)
**工具:** `read` / `Select-String`(Mac: `grep`)
**策略:** 直接全文或部分读取。通过关键词定位相关段落,只读匹配行及其前后文。
**缓存:** 不需要(秒读)。
## 2. PDF
### 2.1 文字版 PDF
**工具:** `pdfminer.six`(Python)
**策略:**
```python
from pdfminer.high_level import extract_text
text = extract_text("file.pdf")
```
→ 在提取结果上做关键词搜索 → 提取相关段落返回
**缓存:** 不需要(< 3 秒/份)。
### 2.2 扫描件/图片版 PDF
**判定:** 先尝试 pdfminer 提取 → 提取结果 < 100 字则判定为扫描件
**工具:** PaddleOCR(优先,中文效果好)→ 备选 EasyOCR
**策略:** OCR → 写入缓存
```python
# 写入 cache/<文件名>.txt
```
**缓存:** ✅ 需要(首次 10-30 秒,缓存后秒回)。
**注意:** 中文准确率约 85-90%,数字和英文更好。
## 3. PPTX(PowerPoint)
**工具:** `python-pptx` + 缓存
**策略:**
1. 递归遍历所有 slide 及 slide 内所有形状(含 GroupShape 组合图形内的子形状)→ 提取 text_frame + notes_slide + 标题
2. 合并为平铺文本 → 写入缓存
3. 在缓存文本上搜索
4. 遇到「如下图所示」等表述 → 转图片处理流程(见第 5 节)
**性能:** 50 页 PPT ≈ 10-20 秒(首次),缓存后秒回。
**局限:** 图表(Chart)、SmartArt、嵌入图片中的文字无法提取。
**缓存:** ✅ 需要(首次慢格式)。
## 4. XLSX(Excel)
**工具:** `pandas`
**策略:**
```python
import pandas as pd
df = pd.read_excel("file.xlsx", nrows=10) # 仅预览表头+前10行
```
→ 按关键词匹配表头 → 筛选相关行
**注意:** 严格限制 `nrows=10`,绝不全表加载。
**缓存:** 不需要。
## 5. 图片处理
### 触发条件
- Phase B 定位到的段落中出现「如下图所示」「见图X」等线索
- 独立图片文件落入候选列表
### 操作(仅具备视觉能力的 Agent)
> ⚠️ 图片分析需要当前模型支持多模态视觉。如果不支持,跳过图片并如实告知用户无法解读,基于可提取的文字继续回答。
> 纯文字搜索不会触发此流程。
>
> ⚠️ Image analysis requires the active model to support multimodal vision.
> If it doesn't, skip the image, report honestly, and continue with
> extractable text. Text-only search never triggers this path.
1. 从 PDF/PPTX 中提取该页的图片资源,或直接读取图片文件
2. 执行视觉分析(仅当前模型支持多模态视觉时执行)
3. 解读结果写入 `cache/<文件名>.img-<页码>.txt`
4. 后续搜到同一页 → 直接读缓存
### 不触发条件
- 装饰性图片(封面图、图标、背景)
### 不具备视觉能力的 Agent
- 如实标注「该文件包含图片,无法自动解读」
- 继续回答基于可提取的文字内容
## 6. 缓存策略
### 核心规则:缓存只服务于慢操作
| 格式 | 是否缓存 | 原因 |
|------|---------|------|
| .md / .txt | ❌ 不缓存 | 秒读,无需转换 |
| .xlsx | ❌ 不缓存 | 只读前 10 行,秒级 |
| .pdf(文字版,≤ 15 页) | ❌ 不缓存 | pdfminer < 3 秒 |
| .pdf(文字版,> 15 页) | ✅ 缓存 | 长文档提取成本高,BM25 建索引和 Phase B 都走缓存 |
| .pdf(扫描件) | ✅ 缓存 | OCR 10-30 秒 |
| .pptx | ✅ 缓存 | 50 页 10-20 秒 |
| .docx(如安装) | ✅ 可选缓存 | 格式转换不稳定 |
| 嵌入图片解析 | ✅ 缓存 | 仅当前模型支持时执行 |
| 独立图片描述 | ✅ 缓存 | 仅当前模型支持时执行,desc 可被搜索命中 |
### 缓存路径
`knowledge-base/<项目名>/cache/`
### 有效判定
原始文件 `lastModified` <= 缓存文件 `createdAt` → 有效
原始文件 `lastModified` > 缓存文件 `createdAt` → 过期,下次读取时重建
### 缓存文件名规则
- 文本提取缓存:`<文件名>.txt`
- 图片解读缓存:`<文件名>.img-<页码>.txt`
- 图片描述缓存:`<文件名>.desc.txt`AionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
activepieces
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
CopilotKit
The Frontend for Agents & Generative UI. React + Angular
Machine-readable data
The same record, as JSON, for agents and crawlers.
{
"facts": [
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Clawhub",
"href": "https://clawhub.ai/kittitys/skills/local-knowledge-retrieval",
"sourceUrl": "https://clawhub.ai/kittitys/skills/local-knowledge-retrieval",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-10T13:12:02.010Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-kittitys-local-knowledge-retrieval/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-kittitys-local-knowledge-retrieval/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-10T13:12:02.010Z",
"isPublic": true
},
{
"factKey": "traction",
"category": "adoption",
"label": "Adoption signal",
"value": "1.4K downloads",
"href": "https://clawhub.ai/kittitys/local-knowledge-retrieval",
"sourceUrl": "https://clawhub.ai/kittitys/local-knowledge-retrieval",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-10T13:12:02.010Z",
"isPublic": true
},
{
"factKey": "latest_release",
"category": "release",
"label": "Latest release",
"value": "3.3.0",
"href": "https://clawhub.ai/kittitys/local-knowledge-retrieval",
"sourceUrl": "https://clawhub.ai/kittitys/local-knowledge-retrieval",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-09-17T17:31:36.837Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-kittitys-local-knowledge-retrieval/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-kittitys-local-knowledge-retrieval/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
],
"events": [
{
"eventType": "release",
"title": "Release 3.3.0",
"description": "local-knowledge-retrieval v3.3.0 - Improved documentation: Updated and streamlined SKILL.md for clarity, security, and user instructions. - Added a roadmap (references/roadmap.md) for clearer project direction. - Added a Windows setup script (scripts/setup.bat) for easier environment setup. - Removed outdated roadmap.md and skill-card.md files. - Updated reference guides for environment setup, conventions, and execution phases. - Refined build and search scripts to align with new documentation and setup enhancements.",
"href": "https://clawhub.ai/kittitys/local-knowledge-retrieval",
"sourceUrl": "https://clawhub.ai/kittitys/local-knowledge-retrieval",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-09-17T17:31:36.837Z",
"isPublic": true
}
]
}Record generated Oct 10, 2026.
