Office Document Extractor
Convert Microsoft Office documents (DOCX, XLSX, PPTX) to Markdown without any external dependencies. Use when the user needs to extract text from Word docume... Skill: Office Document Extractor Owner: michealxie001 Summary: Convert Microsoft Office documents (DOCX, XLSX, PPTX) to Markdown without any external dependencies. Use when the user needs to extract text from Word docume... Tags: converter:1.0.0, docx:1.0.0, latest:1.0.1, markdown:1.0.0, office:1.0.0, pptx:1.0.0, python:1.0.0, xlsx:1.0.0 Version history: v1.0.1 | 2026-05-04T10:08:40.701Z | user Fix: Removed pycache,
Rank
62
Safety
84
Downloads
1.0k
Updated
Oct 11, 2026
Version
1.0.1
Source
CLAWHUB
About
What it does, and when to use it.
Capability contract not published. No trust telemetry is available yet. 1K downloads reported by the source. Last updated 10/11/2026.
Avoid when
- Contract metadata is missing or unavailable for deterministic execution.
Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing
Public facts
Every fact links back to the source it came from.
- Vendor
- Clawhubvendor · observed Oct 11, 2026
- Protocol compatibility
- OpenClawcompatibility · observed Oct 11, 2026
- Adoption signal
- 1K downloadsadoption · observed Oct 11, 2026
- Latest release
- 1.0.1release · observed May 4, 2026
- Handshake status
- UNKNOWNsecurity
Install and run
Setup complexity: low.
clawhub skill install s170w0cbg2mpaewxv2748qq5cx83hwge:office-doc-extractor- Setup complexity is LOW. This package is likely designed for quick installation with minimal external side-effects.
- Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.
Contract: missing
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-michealxie001-office-doc-extractor/snapshot"
Run-check
$0.02 USD1 measured facts are behind this paywall: success rate and latency, uptime and estimated cost, when not to use it, how to call it, benchmark scores.
Agents pay $0.02 in USDC. A card payment is $0.50, the smallest a card allows.
Documentation
CLAWHUB
15,876 characters of source documentation, loaded on request.
Extracted files
3 files captured from the source.
SKILL.md
--- name: office-doc-extractor description: Convert Microsoft Office documents (DOCX, XLSX, PPTX) to Markdown without any external dependencies. Use when the user needs to extract text from Word documents, Excel spreadsheets, or PowerPoint presentations for analysis, indexing, or LLM processing. Pure Python implementation — no pip install, no subprocess calls, no network downloads required. Works offline. --- # Office Document Extractor Zero-dependency converter for Microsoft Office documents. Extracts text and structure from DOCX, XLSX, and PPTX files into clean Markdown. ## Quick Start ```bash # Single file python3 scripts/main.py report.docx -o report.md # Batch convert a directory python3 scripts/main.py ./documents --batch -o ./markdown ``` ## Supported Formats | Format | Extension | Output | |---|---|---| | Word | .docx | Headings, paragraphs | | Excel | .xlsx | Tables (one per sheet) | | PowerPoint | .pptx | Slides as sections | ## How It Works - **DOCX**: Parses the ZIP archive's XML directly using Python's `zipfile` and `xml.etree` - **XLSX**: Uses bundled `openpyxl` (pure Python, no C extensions) - **PPTX**: Parses the ZIP archive's slide XML directly No external commands, no network calls, no pip install required. ## Usage ### Single File ```bash python3 scripts/main.py <input_file> [-o <output.md>] ``` Auto-detects format from file extension. If `-o` is omitted, outputs to `<input>.md`. ### Batch Conversion ```bash python3 scripts/main.py <input_directory> --batch [-o <output_directory>] ``` Converts all `.docx`, `.xlsx`, `.pptx` files in the directory. Results saved to `markdown_output/` by default. ## Resources ### scripts/ - **main.py** — Unified CLI for single-file and batch conversion - **docx_extractor.py** — DOCX → Markdown (standard library only) - **xlsx_extractor.py** — XLSX → Markdown tables (bundled openpyxl) - **pptx_extractor.py** — PPTX → Markdown (standard library only) ### Bundled Dependencies - **openpyxl/** — Pure Python Excel library (v3.1.5) - **et_xmlfile/** — openpyxl dependency (pure Python) ## Limitations - Does not extract images or embedded objects (text only) - Does not preserve complex formatting (colors, fonts, layouts) - Does not handle encrypted/password-protected files - No OCR for scanned documents (use OpenClaw's native `pdf` tool for that) ## Why This Skill? Existing markitdown-based skills require `pip install` or external CLI tools, which triggers ClawHub security warnings. This skill is **100% self-contained** — install it and use it immediately, even offline.
_meta.json
{
"ownerId": "kn79y85j3wvymevvn3nrz452a182atxe",
"slug": "office-doc-extractor",
"version": "1.0.1",
"publishedAt": 1777889320701
}skill-card.md
## Description: Convert Microsoft Office documents (DOCX, XLSX, PPTX) to Markdown without external dependencies for analysis, indexing, or LLM processing. This skill is ready for commercial/non-commercial use. ## Publisher: [michealxie001](https://clawhub.ai/user/michealxie001) ### License/Terms of Use: MIT-0 ## Use Case: Developers, analysts, and agent operators use this skill to convert DOCX, XLSX, and PPTX files into Markdown for review, indexing, summarization, or downstream LLM workflows. ### Deployment Geography for Use: Global ## Known Risks and Mitigations: Risk: Large, malformed, or malicious Office files may cause excessive resource use during ZIP and XML parsing. Mitigation: Process untrusted or bulk document uploads in a constrained environment and apply file size, decompression, timeout, and output limits before conversion. Risk: The converter extracts text and basic structure only, so images, embedded objects, complex formatting, encrypted files, and OCR needs may be omitted. Mitigation: Confirm converted Markdown against the source document when complete fidelity is required, and use a dedicated OCR or secure document-processing workflow for unsupported content. ## Reference(s): - [ClawHub skill page](https://clawhub.ai/michealxie001/skills/office-doc-extractor) - [openpyxl documentation](https://openpyxl.readthedocs.io) - [et_xmlfile project](https://foss.heptapod.net/openpyxl/et_xmlfile) ## Skill Output: **Output Type(s):** [Text, Markdown, Files, Shell commands, Guidance] **Output Format:** [Markdown files and concise command guidance] **Output Parameters:** [1D] **Other Properties Related to Output:** [Supports single-file and batch conversion; extracts text only and does not perform OCR.] ## Skill Version(s): 1.0.1 (source: server release metadata) ## Ethical Considerations: Users should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.
AionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
activepieces
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
CopilotKit
The Frontend for Agents & Generative UI. React + Angular
Machine-readable data
The same record, as JSON, for agents and crawlers.
{
"facts": [
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Clawhub",
"href": "https://clawhub.ai/michealxie001/skills/office-doc-extractor",
"sourceUrl": "https://clawhub.ai/michealxie001/skills/office-doc-extractor",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-11T17:57:31.233Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-michealxie001-office-doc-extractor/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-michealxie001-office-doc-extractor/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-11T17:57:31.233Z",
"isPublic": true
},
{
"factKey": "traction",
"category": "adoption",
"label": "Adoption signal",
"value": "1K downloads",
"href": "https://clawhub.ai/michealxie001/office-doc-extractor",
"sourceUrl": "https://clawhub.ai/michealxie001/office-doc-extractor",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-11T17:57:31.233Z",
"isPublic": true
},
{
"factKey": "latest_release",
"category": "release",
"label": "Latest release",
"value": "1.0.1",
"href": "https://clawhub.ai/michealxie001/office-doc-extractor",
"sourceUrl": "https://clawhub.ai/michealxie001/office-doc-extractor",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-05-04T10:08:40.701Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-michealxie001-office-doc-extractor/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-michealxie001-office-doc-extractor/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
],
"events": [
{
"eventType": "release",
"title": "Release 1.0.1",
"description": "Fix: Removed pycache, repackaged clean build",
"href": "https://clawhub.ai/michealxie001/office-doc-extractor",
"sourceUrl": "https://clawhub.ai/michealxie001/office-doc-extractor",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-05-04T10:08:40.701Z",
"isPublic": true
}
]
}Record generated Oct 11, 2026.
