agentCLAWHUBUnverified

PDF Extract

Extract PDF extracts structured data from PDFs and images, including tables, OCR text, images, and stamps, built on ComPDF data extraction and AI document ex...

OpenClaw

Rank

62

Safety

84

Downloads

1.7k

Updated

Oct 10, 2026

Version

1.2.0

Source

CLAWHUB

About

What it does, and when to use it.

Capability contract not published. No trust telemetry is available yet. 1.7K downloads reported by the source. Last updated 10/10/2026.

Avoid when

  • Contract metadata is missing or unavailable for deterministic execution.

Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing

Public facts

Every fact links back to the source it came from.

Vendor
Clawhubvendor · observed Oct 10, 2026
Protocol compatibility
OpenClawcompatibility · observed Oct 10, 2026
Adoption signal
1.7K downloadsadoption · observed Oct 10, 2026
Latest release
1.2.0release · observed Jun 24, 2026
Handshake status
UNKNOWNsecurity

Install and run

Setup complexity: low.

clawhub skill install s178t897hzzp9pvkwbav4e2kfs83g980:extract-pdf-compdf
  1. Install using `clawhub skill install s178t897hzzp9pvkwbav4e2kfs83g980:extract-pdf-compdf` in an isolated environment before connecting it to live workloads.
  2. No published capability contract is available yet, so validate auth and request/response behavior manually.
  3. Review the upstream CLAWHUB listing at https://clawhub.ai/compdf-youna/extract-pdf-compdf before using production credentials.

Contract: missing

curl -s "https://www.xpersona.co/api/v1/agents/clawhub-compdf-youna-extract-pdf-compdf/snapshot"

Documentation

CLAWHUB

117,219 characters of source documentation, loaded on request.

Extracted files

5 files captured from the source.

SKILL.md

---
name: extract-pdf-compdf
description: >
  Extract PDF extracts structured data from PDFs and images, including tables, OCR text, images, and stamps, built on ComPDF data extraction and AI document extraction capabilities with exports to JSON, XML, CSV, Excel, TXT, or HTML for automation, analytics, and LLM workflows. It suits triggers such as “extract pdf text,” “extract tables from pdf,” “pdf ocr,” “pdf to json,” “parse invoice pdf,” and “extract data from pdf.” Example queries include “Extract all tables from this PDF and export them to CSV,” “Parse this invoice PDF and return the amount, vendor, and tax ID,” and “OCR this scanned PDF and return structured JSON.”
---

# PDF Extract

Process PDF files through ComPDF Cloud REST API. Supports 50+ document processing operations.

Official documentation: <https://www.compdf.com/guides/api-reference/v2/overview?utm_source=clawhub&utm_medium=skillhub&utm_campaign=pdf_skill_pdf_extract&ref_platform_id=clawhub_skills>

## When to Run

- User requests to convert file format (e.g., "convert this PDF to Word", "convert Excel to PDF")
- User requests to edit PDF pages (e.g., "merge these two PDFs", "delete page 3", "rotate PDF")
- User requests to add or remove watermarks from PDF
- User requests to compress PDF files
- User requests OCR recognition of scanned documents or text in images
- User requests AI extraction or parsing of document content
- User requests to extract tables from images
- User requests batch processing of multiple document files
- User requests to compare differences between two PDF documents
- User mentions ComPDF, compdf, or related keywords

## Workflow

### Step 1 — Obtain API Key

Check whether `config/public_key.txt` exists and contains a non-empty value.

- **If the file exists and is non-empty**: use the stored key (trim whitespace).
- **If the file is missing or empty**: ask the user for their ComPDF API Public Key. Inform them it can be obtained at <https://www.compdf.com/compdf-portal/signin?utm_source=clawhub&utm_medium=skillhub&utm_campaign=pdf_skill_pdf_extract&ref_platform_id=clawhub_skills>. After the user provides the key, ask whether they would like to save it locally for future sessions.
  - If the user agrees, write the key to `config/public_key.txt`.
  - If the user declines, use the key for the current session only without saving.

> The key file is **not included in the published skill package**. It is created at runtime only when the user explicitly opts in. The user may delete `config/public_key.txt` at any time to revoke local storage.

### Step 2 — Confirm External Upload Intent

**Before uploading any file**, explicitly inform the user:

> ⚠️ **External Upload Confirmation Required**
>
> Your file will be uploaded to ComPDF's servers (api-server.compdf.com or api-server.compdf.cn) for processing. Please confirm that:
> 1. You consent to uploading this file to external servers.
> 2. The file does not contain highly sensitive or confidential data, or you 

_meta.json

{
  "ownerId": "kn75g40xmv3tgkb21e079d055n82v5h8",
  "slug": "extract-pdf-compdf",
  "version": "1.2.0",
  "publishedAt": 1782278899472
}

references/error-codes.md

# ComPDF API Error Code Reference

## HTTP Status Codes

| Status Code | Description |
|---|---|
| 200 | HTTP request successful (does not indicate file processing success—check `code` and `taskStatus` in JSON) |
| 401 | API Key missing or invalid |
| 413 | Request body exceeds size limit |

## Business Error Codes

### 01xxx - System Errors

| Code | Description | Troubleshooting |
|---|---|---|
| 01001 | System internal error | Retry later; if issue persists, contact [email protected] |
| 01002 | Failed to upload processed file to server | Retry the task |
| 01003 | File upload failed | Check network connection and file size |
| 01004 | File download failed | Check if download link has expired |
| 01005 | File cannot be empty | Verify the uploaded file is not empty |
| 01006 | File parameter exception | Check JSON format in `parameter` field is correct |
| 01201 | System memory issue | File may be too large; try batch processing |
| 01202 | Unknown error | Contact technical support |
| 01203 | File not found or cannot be opened | Verify file path and format are correct |
| 01204 | Unsupported security mechanism | File uses unsupported encryption method |
| 01206 | CSV file not found | Confirm valid CSV file was uploaded |
| 01207 | DocumentAI API call failed | AI service temporarily unavailable; retry later |
| 01208 | Failed to write recognition results to file | Retry the task |

### 02xxx - File Format Errors

| Code | Description | Troubleshooting |
|---|---|---|
| 02001 | File format error | Verify file format matches executeTypeUrl |
| 02002 | Unsupported conversion format | Check `references/tool-list.md` for supported conversion types |
| 02201 | File is encrypted | Provide correct password in `password` field |
| 02203 | PDF file exception | File may be corrupted; try repairing with another tool before re-uploading |
| 02204 | Unsupported file format (PDF only) | This tool accepts PDF input only |
| 02206 | File content invalid or page does not exist | Check specified page number does not exceed total pages |
| 02207 | Cannot open file: unsupported format or file is encrypted | Verify format is correct and provide password if encrypted |
| 02208 | ExcelXml initialization failed | Excel file may be corrupted |
| 02209 | Conversion timeout—file too large | Use async mode or split file and batch process |
| 02210 | File conversion failed | Check if file is corrupted; try re-uploading |
| 02211 | Converted file size abnormal | Source file may have issues; check and retry |

### 03xxx - Parameter Errors

| Code | Description | Troubleshooting |
|---|---|---|
| 03000 | Parameter validation error | Check field names and values in `parameter` JSON match documentation |

### 04xxx - File Operation Errors

| Code | Description | Troubleshooting |
|---|---|---|
| 04001 | fileKey not found | Task may have expired; resubmit the task |
| 04002 | File size is zero | File is empty; upload valid file |
| 04003 | File not found or cannot be opened | Ve

references/parameters.md

# ComPDF API Tool Parameter Details

Parameters are passed as JSON strings in the `parameter` field of the request. If not passed, default values are used.

---

## Format Conversion

### PDF to Word (`pdf/docx`)

```json
{
  "enableAiLayout": 1,
  "isContainImg": 1,
  "isContainAnnot": 1,
  "enableOcr": 0,
  "ocrRecognitionLang": "AUTO",
  "pageRanges": "",
  "pageLayoutMode": "e_Flow",
  "formulaToImage": 0,
  "ocrOption": "ALL",
  "isOutputDocumentPerPage": 0,
  "containPageBackgroundImage": 1
}
```

| Parameter | Description | Default |
|---|---|---|
| `enableAiLayout` | Enable AI layout analysis (0=off, 1=on) | 1 |
| `isContainImg` | Include images during conversion (0=off, 1=on) | 1 |
| `isContainAnnot` | Include annotations during conversion (0=off, 1=on) | 1 |
| `enableOcr` | Enable OCR (0=off, 1=on) | 0 |
| `ocrRecognitionLang` | OCR recognition language (see language list below) | AUTO |
| `pageRanges` | Page range, e.g., `"1,2,3-5"`, empty=all pages | "" |
| `pageLayoutMode` | Layout mode: `e_Box`=fixed layout, `e_Flow`=flow layout | e_Flow |
| `formulaToImage` | Convert formulas to images (0=off, 1=on), recommend on for complex formulas | 0 |
| `ocrOption` | OCR recognition scope (see options below) | ALL |
| `isOutputDocumentPerPage` | Output one document per page (0=off, 1=on) | 0 |
| `containPageBackgroundImage` | Include page background images during OCR (0=off, 1=on) | 1 |

**OCR Recognition Language (ocrRecognitionLang):**
AUTO, CHINESE (Simplified Chinese), CHINESE_TRAD (Traditional Chinese), ENGLISH, KOREAN, JAPANESE, LATIN, DEVANAGARI, CYRILLIC, ARABIC, TAMIL, TELUGU, KANNADA, THAI, GREEK, ESLAV (Slavic)

**OCR Recognition Scope (ocrOption):**
- `INVALID_CHARACTER` - Recognize invalid characters in PDF
- `SCAN_PAGE` - Recognize scanned pages in PDF
- `INVALID_CHARACTERAND_SCAN_PAGE` - Recognize invalid characters and scanned pages
- `ALL` - Recognize all characters on all pages

**Layout Mode Differences:**
- `e_Flow` (Flow layout): Suitable for editing, content can be dynamically adjusted. However, display may vary across different software versions.
- `e_Box` (Fixed layout): Maintain original PDF layout unchanged, but less convenient for editing.

---

### PDF to Image (`pdf/img`)

```json
{
  "imageType": "jpg",
  "imgDpi": "300",
  "pageRanges": ""
}
```

| Parameter | Description | Default |
|---|---|---|
| `imageType` | Output image format: `jpg`, `png` | jpg |
| `imgDpi` | Output image DPI | 300 |
| `pageRanges` | Page range, empty=all pages | "" |

---

### PDF to Excel (`pdf/xlsx`)

```json
{
  "isContainImg": 1,
  "isContainAnnot": 1,
  "pageRanges": ""
}
```

| Parameter | Description | Default |
|---|---|---|
| `isContainImg` | Include images (0/1) | 1 |
| `isContainAnnot` | Include annotations (0/1) | 1 |
| `pageRanges` | Page range, empty=all pages | "" |

---

### PDF to HTML (`pdf/html`)

```json
{
  "isContainImg": 1,
  "pageRanges": ""
}
```

| Parameter | Description | Default |
|---|---|---|
| `isContainIm

references/tool-list.md

# ComPDF API Complete Tool Type Reference (executeTypeUrl)

## ComPDF AI API

### Intelligent Applications

| Feature | executeTypeUrl | Description |
|---|---|---|
| Intelligent Document Extraction | `idp/documentExtract` | Intelligently extract structured data from documents |
| Intelligent Document Parsing | `idp/documentParsing` | Parse document content and structure |

### Intelligent Tools

| Feature | executeTypeUrl | Description |
|---|---|---|
| Text Recognition (OCR) | `documentAI/ocr` | Recognize text in images/scanned documents |
| Table Extraction | `documentAI/tableRec` | Extract table data from images |
| Stamp Detection | `documentAI/detectionStamp` | Detect stamps in documents |
| Image Distortion Correction | `documentAI/dewarp` | Correct distortion in photographed/scanned images |
| Image Enhancement | `documentAI/magicColor` | Enhance image quality |

Supported image formats: JPG, PNG, JPEG, TIFF, BMP

---

## Format Conversion API

### PDF to Other Formats

| Feature | executeTypeUrl | Description |
|---|---|---|
| PDF → Word | `pdf/docx` | Convert PDF to Word document |
| PDF → Excel | `pdf/xlsx` | Convert PDF to Excel spreadsheet |
| PDF → PPT | `pdf/pptx` | Convert PDF to PowerPoint presentation |
| PDF → HTML | `pdf/html` | Convert PDF to HTML page |
| PDF → RTF | `pdf/rtf` | Convert PDF to Rich Text Format |
| PDF → Image | `pdf/img` | Convert PDF pages to images |
| PDF → CSV | `pdf/csv` | Convert PDF tables to CSV |
| PDF → TXT | `pdf/txt` | Extract plain text from PDF |
| PDF → JSON | `pdf/json` | Convert PDF content to JSON |
| PDF → Markdown | `pdf/markdown` | Convert PDF to Markdown |

### Other Formats to PDF

| Feature | executeTypeUrl | Description |
|---|---|---|
| Word → PDF | `doc/pdf` or `docx/pdf` | Convert Word document to PDF |
| Excel → PDF | `xls/pdf` or `xlsx/pdf` | Convert Excel spreadsheet to PDF |
| PPT → PDF | `ppt/pdf` or `pptx/pdf` | Convert PowerPoint to PDF |
| TXT → PDF | `txt/pdf` | Convert plain text to PDF |
| HTML → PDF | `html/pdf` | Convert HTML page to PDF |
| RTF → PDF | `rtf/pdf` | Convert Rich Text Format to PDF |
| PNG → PDF | `png/pdf` | Convert PNG image to PDF |
| CSV → PDF | `csv/pdf` | Convert CSV file to PDF |

### Image to Other Formats

| Feature | executeTypeUrl | Description |
|---|---|---|
| Image → Word | `img/docx` | Convert image to Word document |
| Image → Excel | `img/xlsx` | Convert image to Excel spreadsheet |
| Image → PPT | `img/pptx` | Convert image to PowerPoint |
| Image → JSON | `img/json` | Convert image content to JSON |
| Image → TXT | `img/txt` | Extract text from image |
| Image → HTML | `img/html` | Convert image to HTML |
| Image → RTF | `img/rtf` | Convert image to Rich Text Format |
| Image → CSV | `img/csv` | Convert image tables to CSV |
| Image → PDF | `img/pdf` | Convert image to PDF |

Supported image formats: JPG, PNG, JPEG, TIFF, BMP

---

## PDF Page Editing API

| Feature | executeTypeUrl | Description |
|---|---|---|
| Merge PDF | `pdf/mer
Github ReposUpdated 15h agoRank 70

AionUi

Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!

MCPOPENCLAW
Github ReposUpdated 6mo agoRank 70

activepieces

AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents

OPENCLAW
Github ReposUpdated 6mo agoRank 70

cherry-studio

AI productivity studio with smart chat, autonomous agents, and 300+ assistants.

MCPOPENCLAW
Github ReposUpdated 7mo agoRank 70

CopilotKit

The Frontend for Agents & Generative UI. React + Angular

OPENCLAW

Machine-readable data

The same record, as JSON, for agents and crawlers.

{
  "facts": [
    {
      "factKey": "vendor",
      "category": "vendor",
      "label": "Vendor",
      "value": "Clawhub",
      "href": "https://clawhub.ai/compdf-youna/skills/extract-pdf-compdf",
      "sourceUrl": "https://clawhub.ai/compdf-youna/skills/extract-pdf-compdf",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-10T05:23:04.039Z",
      "isPublic": true
    },
    {
      "factKey": "protocols",
      "category": "compatibility",
      "label": "Protocol compatibility",
      "value": "OpenClaw",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-compdf-youna-extract-pdf-compdf/contract",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-compdf-youna-extract-pdf-compdf/contract",
      "sourceType": "contract",
      "confidence": "medium",
      "observedAt": "2026-10-10T05:23:04.039Z",
      "isPublic": true
    },
    {
      "factKey": "traction",
      "category": "adoption",
      "label": "Adoption signal",
      "value": "1.7K downloads",
      "href": "https://clawhub.ai/compdf-youna/extract-pdf-compdf",
      "sourceUrl": "https://clawhub.ai/compdf-youna/extract-pdf-compdf",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-10T05:23:04.039Z",
      "isPublic": true
    },
    {
      "factKey": "latest_release",
      "category": "release",
      "label": "Latest release",
      "value": "1.2.0",
      "href": "https://clawhub.ai/compdf-youna/extract-pdf-compdf",
      "sourceUrl": "https://clawhub.ai/compdf-youna/extract-pdf-compdf",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-06-24T05:28:19.472Z",
      "isPublic": true
    },
    {
      "factKey": "handshake_status",
      "category": "security",
      "label": "Handshake status",
      "value": "UNKNOWN",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-compdf-youna-extract-pdf-compdf/trust",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-compdf-youna-extract-pdf-compdf/trust",
      "sourceType": "trust",
      "confidence": "medium",
      "observedAt": null,
      "isPublic": true
    }
  ],
  "events": [
    {
      "eventType": "release",
      "title": "Release 1.2.0",
      "description": "extract-pdf-compdf 1.2.0 - Removed the file: skill-card.md - Updated documentation links in SKILL.md to include tracking parameters and point to current ComPDF documentation. - No changes to functionality or workflow.",
      "href": "https://clawhub.ai/compdf-youna/extract-pdf-compdf",
      "sourceUrl": "https://clawhub.ai/compdf-youna/extract-pdf-compdf",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-06-24T05:28:19.472Z",
      "isPublic": true
    }
  ]
}

Record generated Oct 10, 2026.

Sponsored

Ads related to PDF Extract and adjacent AI workflows.