Office Document Assistant
Read, extract, summarize, and compare office documents including PDF, Word, Excel, and PowerPoint. Use when a user provides .pdf/.doc/.docx/.xls/.xlsx/.ppt/....
Rank
62
Safety
84
Downloads
1.7k
Updated
Oct 10, 2026
Version
0.1.1
Source
CLAWHUB
About
What it does, and when to use it.
Capability contract not published. No trust telemetry is available yet. 1.7K downloads reported by the source. Last updated 10/10/2026.
Avoid when
- Contract metadata is missing or unavailable for deterministic execution.
Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing
Public facts
Every fact links back to the source it came from.
- Vendor
- Clawhubvendor · observed Oct 10, 2026
- Protocol compatibility
- OpenClawcompatibility · observed Oct 10, 2026
- Adoption signal
- 1.7K downloadsadoption · observed Oct 10, 2026
- Latest release
- 0.1.1release · observed Mar 29, 2026
- Handshake status
- UNKNOWNsecurity
Install and run
Setup complexity: low.
clawhub skill install s17e6vrn90gyc8zcjhcspvscp583gdtc:office-document-assistant- Install using `clawhub skill install s17e6vrn90gyc8zcjhcspvscp583gdtc:office-document-assistant` in an isolated environment before connecting it to live workloads.
- No published capability contract is available yet, so validate auth and request/response behavior manually.
- Review the upstream CLAWHUB listing at https://clawhub.ai/windrunner20/office-document-assistant before using production credentials.
Contract: missing
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-windrunner20-office-document-assistant/snapshot"
Documentation
CLAWHUB
26,102 characters of source documentation, loaded on request.
Extracted files
5 files captured from the source.
SKILL.md
---
name: office-document-assistant
description: Read, extract, summarize, and compare office documents including PDF, Word, Excel, and PowerPoint. Use when a user provides .pdf/.doc/.docx/.xls/.xlsx/.ppt/.pptx files and asks for summaries, key point extraction, page-by-page outlines, field extraction, table explanation, or multi-document comparison. Prefer the bundled extraction script for deterministic text extraction; for PDFs, fall back to OCR when embedded text is missing.
---
# Office Document Assistant
Read, extract, summarize, and compare common office documents:
- PDF
- Word (`.docx`, `.doc`)
- Excel (`.xlsx`, `.xls`)
- PowerPoint (`.pptx`, `.ppt`)
Use this skill when the user wants the contents of a document explained, summarized, searched, or extracted into a simpler structure.
## When to Use
Use this skill when the user:
- uploads a `.pdf` / `.doc` / `.docx` / `.xls` / `.xlsx` / `.ppt` / `.pptx`
- asks to summarize a document
- asks to extract dates, amounts, contacts, conclusions, specifications, risks, or action items
- asks for page-by-page / slide-by-slide structure
- asks what a spreadsheet or slide deck is saying
- asks to compare two or more documents after extracting their text
## When Not to Use
Do **not** position this skill as a high-fidelity layout or visual analysis system.
It is **not** ideal for:
- precise preservation of original layout, formatting, or pagination
- detailed chart / diagram / image interpretation
- password-protected or encrypted files
- OCR-heavy image understanding beyond basic text recovery
- advanced spreadsheet analytics or formula auditing
- tracked-changes / redline reconstruction in Office documents
## Core Workflow
1. Confirm the document path.
2. Run the bundled script:
- `python3 {skill_dir}/scripts/extract_office_text.py <file> --json`
3. Inspect the JSON fields:
- `type`
- `extraction`
- `warning`
- `truncated`
- `text`
4. Separate clearly in your response:
- **directly extracted content**
- **your summary / inference based on that content**
5. If extraction is empty or weak:
- for PDF, check OCR availability first
- for legacy Office formats, check conversion tools
6. If the user asks for a summary, default to:
- one-sentence overview
- 3–8 key points
- extra sections only when clearly present (dates, people, risks, data, conclusions, contacts)
7. If the user asks for extraction, prefer structured fields over long prose.
## Supported Formats and Strategy
### PDF
- First extract embedded text with `pypdf`.
- If extracted text is too short, fall back to OCR.
- OCR prefers `chi_sim+eng`, then `chi_sim`, then `eng`.
- OCR pipeline requires both `pdftoppm` and `tesseract`.
- If an official first-class PDF tool is exposed in the environment and the task is high-value or multi-PDF, you may prefer that tool; otherwise use this skill's script.
### Word
- `.docx`: extract paragraphs and tables directly.
- `.doc`: try `antiword`, then `catdoc`, then_meta.json
{
"ownerId": "kn78kfsr4c4dkdar20apces7ws829gtx",
"slug": "office-document-assistant",
"version": "0.1.1",
"publishedAt": 1774781014442
}references/capabilities.md
# Capabilities and Boundaries This skill is a text-first office document extraction skill. ## What it does well - Reads common office document formats: PDF, DOCX, DOC, XLSX, XLS, PPTX, PPT - Extracts text from modern Office documents with deterministic libraries - Falls back to OCR for scanned PDFs when embedded text is missing - Returns a stable JSON shape that is easy for agents to summarize - Works well for summary, key-point extraction, field extraction, and rough structure recovery - Handles Chinese PDFs better when `tesseract-ocr-chi-sim` is installed ## What it does only partially ### PDF layout fidelity The skill can extract text, but it does not preserve complex visual layout exactly. Multi-column PDFs, tables, footers, and sidebars may read in imperfect order. ### Spreadsheet structure The skill reads sheet names and rows, but it does not fully reconstruct merged-cell semantics, formulas, formatting logic, pivot tables, or business meaning automatically. ### PowerPoint visuals The skill extracts slide text and notes, but it does not deeply interpret diagrams, animations, screenshots, or design intent. ### OCR accuracy OCR is a recovery path, not a guarantee. Accuracy depends on scan quality, language pack availability, rotation, contrast, and font clarity. ## What it does not aim to do - Exact page layout recreation - Rich image/chart understanding - Password cracking or decryption - Full redline / tracked-changes analysis - Spreadsheet auditing at BI / finance-tool level - Perfect table recovery from arbitrary scanned documents ## Recommended positioning Present this skill as: - a reliable document text extraction and summarization helper - a strong default for uploaded office files in chat - a practical OCR-backed PDF reader for Chinese/English content Do not present it as: - a desktop publishing parser - a visual document AI system - a forensic Office analysis tool
references/troubleshooting.md
# Troubleshooting
Use this guide when extraction is weak, empty, or fails.
## Quick dependency checks
Preferred quick check:
```bash
python3 {skill_dir}/scripts/check_deps.py
```
Manual checks if needed:
### Python packages
```bash
python3 - <<'PY'
import importlib
mods = ["pypdf", "docx", "openpyxl", "pptx"]
for m in mods:
try:
importlib.import_module(m)
print(f"OK {m}")
except Exception as e:
print(f"MISS {m}: {e}")
PY
```
### System tools
```bash
command -v pdftoppm || true
command -v tesseract || true
command -v libreoffice || true
command -v antiword || true
command -v catdoc || true
```
### Tesseract languages
```bash
tesseract --list-langs
```
Look for:
- `chi_sim`
- `eng`
## Common issues
### 1. PDF returns little or no text
Likely causes:
- scanned PDF without embedded text
- missing `pdftoppm`
- missing `tesseract`
- missing OCR language packs
What to do:
- install `poppler-utils`
- install `tesseract-ocr`
- install `tesseract-ocr-chi-sim` for Chinese PDFs
- rerun extraction
### 2. Chinese OCR quality is poor
Likely causes:
- `chi_sim` missing
- low-quality scan
- rotated pages
- low contrast or tiny text
What to do:
- verify `chi_sim` appears in `tesseract --list-langs`
- warn the user that OCR quality may be limited by scan quality
### 3. Legacy `.doc` file fails
Likely causes:
- missing `antiword`
- missing `catdoc`
- missing `libreoffice`
What to do:
- install `antiword`
- install `catdoc`
- install `libreoffice`
- rerun extraction
### 4. Legacy `.xls` or `.ppt` file fails
Likely cause:
- missing `libreoffice`
What to do:
- install `libreoffice`
- rerun extraction
### 5. Extracted tables look messy
This is expected for text-first extraction.
The script is intended to expose content, not fully reconstruct styling or merged-cell semantics.
Explain the table in words instead of pretending the layout is exact.
### 6. Output is truncated
The script supports `--max-chars`.
If the result is cut off, rerun with a larger limit when needed.
## Practical guidance for replies
When extraction succeeds but is imperfect:
- say what was extracted directly
- say what parts may be noisy or OCR-derived
- avoid overclaiming certainty
When extraction fails:
- name the missing dependency or likely cause clearly
- suggest the smallest next step
- do not hallucinate document contentsskill-card.md
## Description: Office Document Assistant helps agents read, extract, summarize, and compare PDF, Word, Excel, and PowerPoint documents using deterministic text extraction with OCR fallback for PDFs. This skill is ready for commercial/non-commercial use. ## Publisher: [windrunner20](https://clawhub.ai/user/windrunner20) ### License/Terms of Use: MIT-0 ## Use Case: Employees, external users, and developers use this skill to extract text from office documents and turn the results into summaries, key points, structured fields, page or slide outlines, and multi-document comparisons. ### Deployment Geography for Use: Global ## Known Risks and Mitigations: Risk: User-supplied documents may contain sensitive information and the skill reads the document paths provided to the agent. Mitigation: Use the skill only on files the user is comfortable giving to the agent, and avoid exposing extracted content beyond the intended session or workflow. Risk: Text-first extraction and OCR can be incomplete or noisy for scanned PDFs, legacy Office files, charts, and complex layouts. Mitigation: Preserve extraction warnings, distinguish directly extracted content from summaries or inference, and verify important facts against the source document. Risk: The skill may invoke local OCR or conversion tools on supplied documents. Mitigation: Confirm the file path and required dependencies before running extraction, and review generated text before relying on downstream summaries or comparisons. ## Reference(s): - [Capabilities and Boundaries](references/capabilities.md) - [Troubleshooting](references/troubleshooting.md) ## Skill Output: **Output Type(s):** [text, markdown, code, shell commands, configuration, guidance] **Output Format:** [Markdown responses with optional JSON extraction output and inline shell commands] **Output Parameters:** [1D] **Other Properties Related to Output:** [May include extracted document text, summaries, key fields, warnings, and truncation notes.] ## Skill Version(s): 0.1.1 (source: server release metadata) ## Ethical Considerations: Users should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.
AionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
activepieces
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
CopilotKit
The Frontend for Agents & Generative UI. React + Angular
Machine-readable data
The same record, as JSON, for agents and crawlers.
{
"facts": [
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Clawhub",
"href": "https://clawhub.ai/windrunner20/skills/office-document-assistant",
"sourceUrl": "https://clawhub.ai/windrunner20/skills/office-document-assistant",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-10T03:24:37.073Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-windrunner20-office-document-assistant/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-windrunner20-office-document-assistant/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-10T03:24:37.073Z",
"isPublic": true
},
{
"factKey": "traction",
"category": "adoption",
"label": "Adoption signal",
"value": "1.7K downloads",
"href": "https://clawhub.ai/windrunner20/office-document-assistant",
"sourceUrl": "https://clawhub.ai/windrunner20/office-document-assistant",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-10T03:24:37.073Z",
"isPublic": true
},
{
"factKey": "latest_release",
"category": "release",
"label": "Latest release",
"value": "0.1.1",
"href": "https://clawhub.ai/windrunner20/office-document-assistant",
"sourceUrl": "https://clawhub.ai/windrunner20/office-document-assistant",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-03-29T10:43:34.442Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-windrunner20-office-document-assistant/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-windrunner20-office-document-assistant/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
],
"events": [
{
"eventType": "release",
"title": "Release 0.1.1",
"description": "Add bundled dependency checker and troubleshooting updates.",
"href": "https://clawhub.ai/windrunner20/office-document-assistant",
"sourceUrl": "https://clawhub.ai/windrunner20/office-document-assistant",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-03-29T10:43:34.442Z",
"isPublic": true
}
]
}Record generated Oct 10, 2026.
