agentCLAWHUBUnverified

Office Document Extractor

Convert Microsoft Office documents (DOCX, XLSX, PPTX) to Markdown without any external dependencies. Use when the user needs to extract text from Word docume... Skill: Office Document Extractor Owner: michealxie001 Summary: Convert Microsoft Office documents (DOCX, XLSX, PPTX) to Markdown without any external dependencies. Use when the user needs to extract text from Word docume... Tags: converter:1.0.0, docx:1.0.0, latest:1.0.1, markdown:1.0.0, office:1.0.0, pptx:1.0.0, python:1.0.0, xlsx:1.0.0 Version history: v1.0.1 | 2026-05-04T10:08:40.701Z | user Fix: Removed pycache,

OpenClaw

Rank

62

Safety

84

Downloads

1.0k

Updated

Oct 11, 2026

Version

1.0.1

Source

CLAWHUB

About

What it does, and when to use it.

Capability contract not published. No trust telemetry is available yet. 1K downloads reported by the source. Last updated 10/11/2026.

Avoid when

  • Contract metadata is missing or unavailable for deterministic execution.

Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing

Public facts

Every fact links back to the source it came from.

Vendor
Clawhubvendor · observed Oct 11, 2026
Protocol compatibility
OpenClawcompatibility · observed Oct 11, 2026
Adoption signal
1K downloadsadoption · observed Oct 11, 2026
Latest release
1.0.1release · observed May 4, 2026
Handshake status
UNKNOWNsecurity

Install and run

Setup complexity: low.

clawhub skill install s170w0cbg2mpaewxv2748qq5cx83hwge:office-doc-extractor
  1. Setup complexity is LOW. This package is likely designed for quick installation with minimal external side-effects.
  2. Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.

Contract: missing

curl -s "https://www.xpersona.co/api/v1/agents/clawhub-michealxie001-office-doc-extractor/snapshot"

Run-check

$0.02 USD

1 measured facts are behind this paywall: success rate and latency, uptime and estimated cost, when not to use it, how to call it, benchmark scores.

Agents pay $0.02 in USDC. A card payment is $0.50, the smallest a card allows.

Documentation

CLAWHUB

15,876 characters of source documentation, loaded on request.

Extracted files

3 files captured from the source.

SKILL.md

---
name: office-doc-extractor
description: Convert Microsoft Office documents (DOCX, XLSX, PPTX) to Markdown without any external dependencies. Use when the user needs to extract text from Word documents, Excel spreadsheets, or PowerPoint presentations for analysis, indexing, or LLM processing. Pure Python implementation — no pip install, no subprocess calls, no network downloads required. Works offline.
---

# Office Document Extractor

Zero-dependency converter for Microsoft Office documents. Extracts text and structure from DOCX, XLSX, and PPTX files into clean Markdown.

## Quick Start

```bash
# Single file
python3 scripts/main.py report.docx -o report.md

# Batch convert a directory
python3 scripts/main.py ./documents --batch -o ./markdown
```

## Supported Formats

| Format | Extension | Output |
|---|---|---|
| Word | .docx | Headings, paragraphs |
| Excel | .xlsx | Tables (one per sheet) |
| PowerPoint | .pptx | Slides as sections |

## How It Works

- **DOCX**: Parses the ZIP archive's XML directly using Python's `zipfile` and `xml.etree`
- **XLSX**: Uses bundled `openpyxl` (pure Python, no C extensions)
- **PPTX**: Parses the ZIP archive's slide XML directly

No external commands, no network calls, no pip install required.

## Usage

### Single File

```bash
python3 scripts/main.py <input_file> [-o <output.md>]
```

Auto-detects format from file extension. If `-o` is omitted, outputs to `<input>.md`.

### Batch Conversion

```bash
python3 scripts/main.py <input_directory> --batch [-o <output_directory>]
```

Converts all `.docx`, `.xlsx`, `.pptx` files in the directory. Results saved to `markdown_output/` by default.

## Resources

### scripts/

- **main.py** — Unified CLI for single-file and batch conversion
- **docx_extractor.py** — DOCX → Markdown (standard library only)
- **xlsx_extractor.py** — XLSX → Markdown tables (bundled openpyxl)
- **pptx_extractor.py** — PPTX → Markdown (standard library only)

### Bundled Dependencies

- **openpyxl/** — Pure Python Excel library (v3.1.5)
- **et_xmlfile/** — openpyxl dependency (pure Python)

## Limitations

- Does not extract images or embedded objects (text only)
- Does not preserve complex formatting (colors, fonts, layouts)
- Does not handle encrypted/password-protected files
- No OCR for scanned documents (use OpenClaw's native `pdf` tool for that)

## Why This Skill?

Existing markitdown-based skills require `pip install` or external CLI tools, which triggers ClawHub security warnings. This skill is **100% self-contained** — install it and use it immediately, even offline.

_meta.json

{
  "ownerId": "kn79y85j3wvymevvn3nrz452a182atxe",
  "slug": "office-doc-extractor",
  "version": "1.0.1",
  "publishedAt": 1777889320701
}

skill-card.md

## Description:

Convert Microsoft Office documents (DOCX, XLSX, PPTX) to Markdown without external dependencies for analysis, indexing, or LLM processing.

This skill is ready for commercial/non-commercial use.

## Publisher:

[michealxie001](https://clawhub.ai/user/michealxie001)

### License/Terms of Use:

MIT-0

## Use Case:

Developers, analysts, and agent operators use this skill to convert DOCX, XLSX, and PPTX files into Markdown for review, indexing, summarization, or downstream LLM workflows.

### Deployment Geography for Use:

Global

## Known Risks and Mitigations:

Risk: Large, malformed, or malicious Office files may cause excessive resource use during ZIP and XML parsing.

Mitigation: Process untrusted or bulk document uploads in a constrained environment and apply file size, decompression, timeout, and output limits before conversion.

Risk: The converter extracts text and basic structure only, so images, embedded objects, complex formatting, encrypted files, and OCR needs may be omitted.

Mitigation: Confirm converted Markdown against the source document when complete fidelity is required, and use a dedicated OCR or secure document-processing workflow for unsupported content.

## Reference(s):

- [ClawHub skill page](https://clawhub.ai/michealxie001/skills/office-doc-extractor)
- [openpyxl documentation](https://openpyxl.readthedocs.io)
- [et_xmlfile project](https://foss.heptapod.net/openpyxl/et_xmlfile)

## Skill Output:

**Output Type(s):** [Text, Markdown, Files, Shell commands, Guidance]

**Output Format:** [Markdown files and concise command guidance]

**Output Parameters:** [1D]

**Other Properties Related to Output:** [Supports single-file and batch conversion; extracts text only and does not perform OCR.]

## Skill Version(s):

1.0.1 (source: server release metadata)

## Ethical Considerations:

Users should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.
Github ReposUpdated 2d agoRank 70

AionUi

Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!

MCPOPENCLAW
Github ReposUpdated 6mo agoRank 70

activepieces

AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents

OPENCLAW
Github ReposUpdated 6mo agoRank 70

cherry-studio

AI productivity studio with smart chat, autonomous agents, and 300+ assistants.

MCPOPENCLAW
Github ReposUpdated 7mo agoRank 70

CopilotKit

The Frontend for Agents & Generative UI. React + Angular

OPENCLAW

Machine-readable data

The same record, as JSON, for agents and crawlers.

{
  "facts": [
    {
      "factKey": "vendor",
      "category": "vendor",
      "label": "Vendor",
      "value": "Clawhub",
      "href": "https://clawhub.ai/michealxie001/skills/office-doc-extractor",
      "sourceUrl": "https://clawhub.ai/michealxie001/skills/office-doc-extractor",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-11T17:57:31.233Z",
      "isPublic": true
    },
    {
      "factKey": "protocols",
      "category": "compatibility",
      "label": "Protocol compatibility",
      "value": "OpenClaw",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-michealxie001-office-doc-extractor/contract",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-michealxie001-office-doc-extractor/contract",
      "sourceType": "contract",
      "confidence": "medium",
      "observedAt": "2026-10-11T17:57:31.233Z",
      "isPublic": true
    },
    {
      "factKey": "traction",
      "category": "adoption",
      "label": "Adoption signal",
      "value": "1K downloads",
      "href": "https://clawhub.ai/michealxie001/office-doc-extractor",
      "sourceUrl": "https://clawhub.ai/michealxie001/office-doc-extractor",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-11T17:57:31.233Z",
      "isPublic": true
    },
    {
      "factKey": "latest_release",
      "category": "release",
      "label": "Latest release",
      "value": "1.0.1",
      "href": "https://clawhub.ai/michealxie001/office-doc-extractor",
      "sourceUrl": "https://clawhub.ai/michealxie001/office-doc-extractor",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-05-04T10:08:40.701Z",
      "isPublic": true
    },
    {
      "factKey": "handshake_status",
      "category": "security",
      "label": "Handshake status",
      "value": "UNKNOWN",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-michealxie001-office-doc-extractor/trust",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-michealxie001-office-doc-extractor/trust",
      "sourceType": "trust",
      "confidence": "medium",
      "observedAt": null,
      "isPublic": true
    }
  ],
  "events": [
    {
      "eventType": "release",
      "title": "Release 1.0.1",
      "description": "Fix: Removed pycache, repackaged clean build",
      "href": "https://clawhub.ai/michealxie001/office-doc-extractor",
      "sourceUrl": "https://clawhub.ai/michealxie001/office-doc-extractor",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-05-04T10:08:40.701Z",
      "isPublic": true
    }
  ]
}

Record generated Oct 11, 2026.

Sponsored

Ads related to Office Document Extractor and adjacent AI workflows.