PaddleOCR Document Parsing
Parse documents using PaddleOCR's API. Skill: PaddleOCR Document Parsing Owner: Bobholamovic Summary: Parse documents using PaddleOCR's API. Tags: latest:1.0.2 Version history: v1.0.2 | 2026-02-10T08:00:39.427Z | user - No changes detected in this release. v1.0.1 | 2026-02-09T13:29:36.477Z | user - Added a new "Resource Links" section with quick access to the official website, API documentation, and GitHub. - Updated the setup instructions to clarify algo
Rank
62
Safety
84
Downloads
1.5k
Updated
Apr 15, 2026
Version
1.0.2
Source
CLAWHUB
About
What it does, and when to use it.
Capability contract not published. No trust telemetry is available yet. 1.5K downloads reported by the source. Last updated 4/15/2026.
Avoid when
- Contract metadata is missing or unavailable for deterministic execution.
Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing
Public facts
Every fact links back to the source it came from.
- Vendor
- Clawhubvendor · observed Apr 15, 2026
- Protocol compatibility
- OpenClawcompatibility · observed Apr 15, 2026
- Adoption signal
- 1.5K downloadsadoption · observed Apr 15, 2026
- Latest release
- 1.0.2release · observed Feb 10, 2026
- Handshake status
- UNKNOWNsecurity
Install and run
Setup complexity: low.
clawhub skill install kn77zppfj1a2fc620aygaf9z9980ewfa:paddleocr-doc-parsing- Setup complexity is LOW. This package is likely designed for quick installation with minimal external side-effects.
- Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.
Contract: missing
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-bobholamovic-paddleocr-doc-parsing/snapshot"
Documentation
CLAWHUB
7,680 characters of source documentation, loaded on request.
Extracted files
2 files captured from the source.
SKILL.md
---
name: paddleocr-doc-parsing
description: Parse documents using PaddleOCR's API.
homepage: https://www.paddleocr.com
metadata:
{
"openclaw":
{
"emoji": "📄",
"os": ["darwin", "linux"],
"requires":
{
"bins": ["curl", "base64", "jq"],
"env": ["PADDLEOCR_API_URL", "PADDLEOCR_ACCESS_TOKEN"],
},
},
}
---
# PaddleOCR Document Parsing
Parse images and PDF files using PaddleOCR's API. Supports multiple document parsing algorithms with structured output.
## Resource Links
| Resource | Link |
| --------------------- | ------------------------------------------------------------------------------ |
| **Official Website** | [https://www.paddleocr.com](https://www.paddleocr.com) |
| **API Documentation** | [https://ai.baidu.com/ai-doc/AISTUDIO/Cmkz2m0ma](https://ai.baidu.com/ai-doc/AISTUDIO/Cmkz2m0ma) |
| **GitHub** | [https://github.com/PaddlePaddle/PaddleOCR](https://github.com/PaddlePaddle/PaddleOCR) |
## Key Features
- **Multi-format support**: PDF and image files (JPG, PNG, BMP, TIFF)
- **Layout analysis**: Automatic detection of text blocks, tables, formulas
- **Multi-language**: Support for 110+ languages
- **Structured output**: Markdown format with preserved document structure
## Setup
1. Obtain credentials from the [PaddleOCR official website](https://www.paddleocr.com). Click the “API” button, choose the desired algorithm (e.g., PP-StructureV3, PaddleOCR-VL-1.5), and copy the API URL and the access token.
2. Set environment variables:
```bash
export PADDLEOCR_API_URL="https://your-endpoint-here"
export PADDLEOCR_ACCESS_TOKEN="your_access_token"
```
## Usage Examples
### Run Script
```bash
# Parse local image
{baseDir}/paddleocr_parse.sh document.jpg
# Parse local PDF file
{baseDir}/paddleocr_parse.sh -t pdf document.pdf
# Parse document from URL
{baseDir}/paddleocr_parse.sh -t pdf https://example.com/document.pdf
# Output to stdout (default)
{baseDir}/paddleocr_parse.sh document.jpg
# Save output to file
{baseDir}/paddleocr_parse.sh -o result.json document.jpg
```
### Response Structure
```json
{
"logId": "unique_request_id",
"errorCode": 0,
"errorMsg": "Success",
"result": {
"layoutParsingResults": [
{
"prunedResult": [...],
"markdown": {
"text": "# Document Title\n\nParagraph content...",
"images": {}
},
"outputImages": [...],
"inputImage": "http://input-image"
}
],
"dataInfo": {...}
}
}
```
**Important Fields:**
- **`prunedResult`** - Contains detailed layout element information including positions, categories, etc.
- **`markdown`** - Stores the document content converted to Markdown format with preserved structure and formatting.
## Quota Information
See official documentation: https://ai.baidu.com/ai_meta.json
{
"ownerId": "kn77zppfj1a2fc620aygaf9z9980ewfa",
"slug": "paddleocr-doc-parsing",
"version": "1.0.2",
"publishedAt": 1770710439427
}activepieces
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
AionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
CopilotKit
The Frontend for Agents & Generative UI. React + Angular
Machine-readable data
The same record, as JSON, for agents and crawlers.
{
"facts": [
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Clawhub",
"href": "https://clawhub.ai/Bobholamovic/paddleocr-doc-parsing",
"sourceUrl": "https://clawhub.ai/Bobholamovic/paddleocr-doc-parsing",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-04-15T00:45:39.800Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-bobholamovic-paddleocr-doc-parsing/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-bobholamovic-paddleocr-doc-parsing/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-04-15T00:45:39.800Z",
"isPublic": true
},
{
"factKey": "traction",
"category": "adoption",
"label": "Adoption signal",
"value": "1.5K downloads",
"href": "https://clawhub.ai/Bobholamovic/paddleocr-doc-parsing",
"sourceUrl": "https://clawhub.ai/Bobholamovic/paddleocr-doc-parsing",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-04-15T00:45:39.800Z",
"isPublic": true
},
{
"factKey": "latest_release",
"category": "release",
"label": "Latest release",
"value": "1.0.2",
"href": "https://clawhub.ai/Bobholamovic/paddleocr-doc-parsing",
"sourceUrl": "https://clawhub.ai/Bobholamovic/paddleocr-doc-parsing",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-02-10T08:00:39.427Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-bobholamovic-paddleocr-doc-parsing/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-bobholamovic-paddleocr-doc-parsing/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
],
"events": [
{
"eventType": "release",
"title": "Release 1.0.2",
"description": "- No changes detected in this release.",
"href": "https://clawhub.ai/Bobholamovic/paddleocr-doc-parsing",
"sourceUrl": "https://clawhub.ai/Bobholamovic/paddleocr-doc-parsing",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-02-10T08:00:39.427Z",
"isPublic": true
}
]
}Record generated Oct 9, 2026.
