iwork2md
Convert Apple iWork documents (Pages .pages, Numbers .numbers, Keynote .key) into Markdown. Use whenever the user wants to read, extract, or translate the content of an iWork file into text/markdown, for example 'convert this .pages file to markdown', 'extract text from a Numbers sheet', 'read a Keynote file', or 'open a .key/.numbers/.pages and turn it into markdown'. Handles the iWork '13+ format (bundle containing Index.zip with .iwa files that wrap Snappy-framed Protobuf) with no third-party dependencies.
Rank
62
Safety
84
Downloads
2.7k
Updated
Oct 9, 2026
Version
0.1.0
Source
CLAWHUB
About
What it does, and when to use it.
Capability contract not published. No trust telemetry is available yet. 2.7K downloads reported by the source. Last updated 10/9/2026.
Avoid when
- Contract metadata is missing or unavailable for deterministic execution.
Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing
Public facts
Every fact links back to the source it came from.
- Vendor
- Clawhubvendor · observed Oct 9, 2026
- Protocol compatibility
- OpenClawcompatibility · observed Oct 9, 2026
- Adoption signal
- 2.7K downloadsadoption · observed Oct 9, 2026
- Latest release
- 0.1.0release · observed Jul 29, 2026
- Handshake status
- UNKNOWNsecurity
Install and run
Setup complexity: low.
clawhub skill install s171srzvxc5431ssbjh7vggsvs884354:iwork2md- Install using `clawhub skill install s171srzvxc5431ssbjh7vggsvs884354:iwork2md` in an isolated environment before connecting it to live workloads.
- No published capability contract is available yet, so validate auth and request/response behavior manually.
- Review the upstream CLAWHUB listing at https://clawhub.ai/slearnai/iwork2md before using production credentials.
Contract: missing
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-slearnai-iwork2md/snapshot"
Documentation
CLAWHUB
17,156 characters of source documentation, loaded on request.
Extracted files
5 files captured from the source.
SKILL.md
---
name: iwork2md
slug: iwork2md
version: 1.0.0
displayName: iWork to Markdown
description: "Convert Apple iWork documents (Pages .pages, Numbers .numbers, Keynote .key) into Markdown. Use whenever the user wants to read, extract, or translate the content of an iWork file into text/markdown, for example 'convert this .pages file to markdown', 'extract text from a Numbers sheet', 'read a Keynote file', or 'open a .key/.numbers/.pages and turn it into markdown'. Handles the iWork '13+ format (bundle containing Index.zip with .iwa files that wrap Snappy-framed Protobuf) with no third-party dependencies."
license: MIT
summary: Convert Apple iWork (.pages/.numbers/.key) documents to Markdown with a dependency-free Python parser.
tags:
- iwork
- pages
- numbers
- keynote
- markdown
- conversion
---
# iwork2md — iWork (.pages / .numbers / .key) to Markdown
Convert Apple Pages / Numbers / Keynote documents to Markdown. The parser is in
`scripts/iwa.py` (pure stdlib); the converter CLI is `scripts/iwork2md.py`.
## When to use
- User provides a `.pages`, `.numbers`, or `.key` file and wants its text,
tables, or slides as Markdown (or just to *read* the content).
- User asks to "extract text / convert / translate / open" an iWork file.
- Do NOT use for: password-protected/encrypted iWork docs (unsupported),
or for reconstructing exact visual layout (not the goal).
## How to run
```bash
# Write a .md next to the source (auto-named)
python3 scripts/iwork2md.py path/to/Doc.pages
# Explicit output path
python3 scripts/iwork2md.py Doc.numbers out.md
# Print to stdout
python3 scripts/iwork2md.py Doc.key --stdout
# Debug: dump every recovered text fragment
python3 scripts/iwork2md.py Doc.numbers --texts
# List embedded media (images/video)
python3 scripts/iwork2md.py Doc.pages --media
```
From inside a chat, invoke with `exec` (or tell the user to run it). The script
is dependency-free (Python 3.8+, stdlib only: `zipfile`, `struct`, `io`,
`plistlib`).
## What it does
1. Opens the bundle ZIP; finds `Index.zip` (or `.iwa` files directly under
`Index/`).
2. For each `.iwa`: removes the iWork **Snappy framing** (chunk type + 3-byte
LE length, no stream-id, no CRC), then raw-Snappy-decompresses the body.
3. Parses the **Protobuf container** (`varint len + ArchiveInfo {identifier,
message_infos[]}` then payloads), and generically walks every message to
collect UTF-8 string fields — recovering ~100% of readable content without
needing the app-specific schema map (TSPRegistry).
4. Renders Markdown: document title (from `Metadata/Properties.plist` or first
heading), an embedded-media list, reconstructed **Numbers tables** (rows
stored as `"a | b | c"` become proper markdown tables, deduped across
mirrored components), a body block (largest multi-line text), and remaining
text fragments.
## Key facts you need (so you don't re-derive them)
- iWork `.iwa` Snappy framing is **non-standard**: type byte `0x00`, 3-byte LEREADME.md
# iwork2md
Convert Apple iWork documents — **Pages (`.pages`)**, **Numbers (`.numbers`)**, and **Keynote (`.key`)** — into Markdown. Pure Python, **no third-party dependencies**.
`iwork2md` reads the iWork '13+ bundle format directly:
```
file.pages / file.numbers / file.key (ZIP or directory bundle)
└── Index.zip
└── *.iwa (Snappy-framed Protobuf payloads)
```
It decodes the non-standard Snappy framing and Protobuf containers, extracts text and media references, and **reconstructs modern Numbers tables** (row/column grids with strings, numbers, and formulas).
---
## Features
- **No dependencies** — pure Python 3 standard library (`zipfile`, `struct`, `re`). Works on any macOS/Linux/Windows box.
- **Text extraction** — Pages body text, Numbers cell/header text, Keynote slide text, speaker notes, and comments.
- **Numbers table reconstruction** — resolves `Table (6001)` → `DataList (6005)` linkages and rebuilds the full grid, decoding:
- **strings** (field 3)
- **numbers** (IEEE-754 doubles, field 5 → `f42`)
- **formulas** (rendered as `=formula`)
- empty cells are padded so the layout is preserved
- **Robust layout handling** — works on both ZIP-bundle files and `.pages` directory packages.
- **Noise filtering** — drops locale codes (`en_HK`), timezones, month-name lists, and number-format strings that would otherwise pollute the output.
- **SkillHub-ready** — ships with `SKILL.md` (slug/version/displayName), `LICENSE.txt` (MIT), and a `references/FORMAT.md` format note.
---
## Installation
Clone the repo:
```bash
git clone https://github.com/slearnAI/iwork2md.git
cd iwork2md
```
No `pip install` required — just run the script. (Optional: `chmod +x scripts/iwork2md.py`.)
---
## Usage
Convert a single file:
```bash
python3 scripts/iwork2md.py path/to/document.numbers
```
Write to a specific output path:
```bash
python3 scripts/iwork2md.py path/to/deck.key output.md
```
If no output path is given, Markdown is printed to **stdout**.
### As an OpenClaw skill
Drop the folder into your skills directory:
```bash
cp -r iwork2md ~/.qclaw/skills/
```
Then ask naturally: *"convert this .pages file to markdown"*, *"extract text from my Numbers sheet"*, *"read that Keynote file"* — the skill triggers and runs the bundled converter.
---
## Output format
- **Pages / Keynote** → headings, paragraphs, lists, and a `## Media` section listing referenced images/video by filename.
- **Numbers** → a `## Tables` section with one Markdown table per sheet, plus any stray text fragments under `## Other text`.
Example (Numbers):
```markdown
# My Spreadsheet
## Tables
| 南航積分 | 酒店積分 | 曼谷 |
| --- | --- | --- |
| CZ3062 | 機票 | 稅費 |
| =formula | | |
```
---
## How it works
The underlying decoder (`scripts/iwa.py`) handles three layers:
1. **Bundle** — unzip the `.iwa` files (or walk a directory package).
2. **Snappy framing** — each `.iwa` is a sequence of chunks: `1-byte type` + `3-byte little-endian l_meta.json
{
"ownerId": "kn70217mcwqf7ya07qzw4xv5rd8014q7",
"slug": "iwork2md",
"version": "0.1.0",
"publishedAt": 1785330233340
}references/FORMAT.md
# iWork File Format Reference (Pages / Numbers / Keynote)
This skill targets the **iWork '13+** format used by current Pages (.pages),
Numbers (.numbers) and Keynote (.key) documents.
## Physical layout
The document is a **bundle** = a ZIP archive containing:
```
MyDoc.pages/ (outer ZIP)
├── Index.zip (all serialized objects, see below)
├── Data/ (embedded media: images, video, etc.)
│ └── 143917994_2881x1992-small.jpg
├── Metadata/
│ ├── Properties.plist (title / metadata)
│ ├── DocumentIdentifier
│ └── BuildVersionHistory.plist
├── preview.jpg (preview thumbnails, top level)
├── preview-web.jpg
└── preview-micro.jpg
```
> Note: some iWork versions place the `.iwa` files **directly** under `Index/`
> inside the outer ZIP instead of inside a nested `Index.zip`. The parser
> (`iwa.open_iwa_sources`) handles both layouts.
## Index.zip -> .iwa
Inside `Index.zip` are many `.iwa` files (one or more per Component):
`Document.iwa`, `MasterSlide-1.iwa`, `CalculationEngine.iwa`, etc.
- The iWork ZIP writer uses **no compression** and no Zip64. Re-zipping with a
normal tool can break the document, but reading is standard ZIP.
## .iwa = Protobuf stream wrapped in non-standard Snappy framing
### Snappy framing (iWork variant — NOT the official spec)
Back-to-back chunks:
```
[1 byte type][3-byte LE chunk length][length bytes data]
```
- iWork only emits **type 0x00** (compressed).
- It **omits** the mandatory stream-identifier chunk (`0xFF "sNaPpY"`).
- It **omits** the CRC-32C checksum that the official framing prepends to
compressed data.
- For type 0x00, the chunk **data is a raw Snappy block** (starts with an
uncompressed-length varint), not an officially-framed stream.
`iwa.iwa_unframe()` implements this exact variant.
### Raw Snappy block
```
varint uncompressed_length
<LZ77 stream: literals + back-references>
```
Elements start with a tag byte; lower 2 bits = type:
- `00` literal (len in upper 6 bits, or 1–4 follow bytes for len ≥ 61)
- `01` copy, 1-byte offset (len 4–11, offset 0–2047)
- `10` copy, 2-byte offset (len 1–64, offset 0–65535)
- `11` copy, 4-byte offset (len 1–64, offset 0–2^32)
`iwa.snappy_decompress()` implements the block format.
### Protobuf container stream (after unframing)
Objects are concatenated:
```
varint archive_info_len
ArchiveInfo { # message
field 1: identifier (uint64) # unique id across the document
field 2: repeated MessageInfo
}
<for each MessageInfo, the payload bytes>
```
`MessageInfo`:
```
field 1: type (uint32) # selects the payload's protobuf schema
field 2: version (packed uint32)
field 3: length (uint32) # payload byte length
field 5: object_references (packed uint64)
field 6: data_references (packed uint64)
```
`type` -> schema mapping (the **TSPRegistry**) is embedded inside the iWork
binaries and differs per app/version. Because Protobuf is not self-describing,
perfect decodingskill-card.md
## Description: Convert Apple iWork documents (Pages .pages, Numbers .numbers, Keynote .key) into Markdown. This skill is ready for commercial/non-commercial use. ## Publisher: [slearnai](https://clawhub.ai/user/slearnai) ### License/Terms of Use: MIT ## Use Case: Developers, engineers, and other agents use this skill to read, extract, or convert Apple iWork Pages, Numbers, and Keynote documents into text or Markdown. ### Deployment Geography for Use: Global ## Known Risks and Mitigations: Risk: Malicious or very large iWork documents may cause excessive memory use or long processing time. Mitigation: Use the skill only for files you intentionally provide, avoid untrusted oversized documents, and run conversion with agent-enforced memory and time limits. Risk: Converted Markdown may be written beside the source document when no explicit destination is provided. Mitigation: Prefer --stdout or an explicit output path so the user and agent know where converted content is written. Risk: Encrypted iWork files and exact visual layout reconstruction are unsupported. Mitigation: Use the output for readable text, tables, slide text, and media inventory; require a different workflow when password-protected documents or pixel-accurate layout are needed. ## Reference(s): - [iWork file format reference](references/FORMAT.md) - [Server-resolved GitHub repository](https://github.com/slearnAI/iwork2md) - [Server-resolved GitHub commit](https://github.com/slearnAI/iwork2md/tree/de2fbb57e58d6643908e2823bfd42a8c005da357) - [ClawHub skill page](https://clawhub.ai/slearnai/skills/iwork2md) ## Skill Output: **Output Type(s):** [Text, Markdown, Files, Shell commands, Guidance] **Output Format:** [Markdown, plain text, file paths, or concise shell commands] **Output Parameters:** [1D] **Other Properties Related to Output:** [May write a Markdown file next to the source document, write to an explicit output path, or print to stdout.] ## Skill Version(s): 0.1.0 (source: server release metadata) ## Ethical Considerations: Users should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.
activepieces
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
AionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
CopilotKit
The Frontend for Agents & Generative UI. React + Angular
Machine-readable data
The same record, as JSON, for agents and crawlers.
{
"facts": [
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Clawhub",
"href": "https://clawhub.ai/slearnai/skills/iwork2md",
"sourceUrl": "https://clawhub.ai/slearnai/skills/iwork2md",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-09T11:56:41.882Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-slearnai-iwork2md/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-slearnai-iwork2md/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-09T11:56:41.882Z",
"isPublic": true
},
{
"factKey": "traction",
"category": "adoption",
"label": "Adoption signal",
"value": "2.7K downloads",
"href": "https://clawhub.ai/slearnai/iwork2md",
"sourceUrl": "https://clawhub.ai/slearnai/iwork2md",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-09T11:56:41.882Z",
"isPublic": true
},
{
"factKey": "latest_release",
"category": "release",
"label": "Latest release",
"value": "0.1.0",
"href": "https://clawhub.ai/slearnai/iwork2md",
"sourceUrl": "https://clawhub.ai/slearnai/iwork2md",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-07-29T13:03:53.340Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-slearnai-iwork2md/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-slearnai-iwork2md/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
],
"events": [
{
"eventType": "release",
"title": "Release 0.1.0",
"description": "Initial release: Convert Apple iWork files (.pages, .numbers, .key) to Markdown using a pure standard library Python parser. - Supports extraction of all readable content, tables, and slide text from iWork '13+ format documents. - Dependency-free: uses only Python 3.8+ standard library (no third-party packages required). - Handles non-standard iWork Snappy framing and Protobuf containers to recover UTF-8 text. - CLI supports file conversion, text extraction, printing to stdout, and media listing. - Does not support encrypted documents or visual layout reconstruction.",
"href": "https://clawhub.ai/slearnai/iwork2md",
"sourceUrl": "https://clawhub.ai/slearnai/iwork2md",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-07-29T13:03:53.340Z",
"isPublic": true
}
]
}Record generated Oct 9, 2026.
