PaddleOCR Text Recognition
Use this skill whenever the user wants text extracted from images, photos, scans, screenshots, or scanned PDFs. Returns exact machine-readable strings with l... Skill: PaddleOCR Text Recognition Owner: bobholamovic Summary: Use this skill whenever the user wants text extracted from images, photos, scans, screenshots, or scanned PDFs. Returns exact machine-readable strings with l... Tags: latest:2.0.0 Version history: v2.0.0 | 2026-06-05T10:10:23.964Z | user - Major refactor: removed all local wrapper scripts and sample/reference files, switching to direct CLI usage. - SKILL.
Rank
62
Safety
84
Downloads
4.5k
Updated
Oct 9, 2026
Version
2.0.0
Source
CLAWHUB
About
What it does, and when to use it.
Capability contract not published. No trust telemetry is available yet. 4.5K downloads reported by the source. Last updated 10/9/2026.
Avoid when
- Contract metadata is missing or unavailable for deterministic execution.
Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing
Public facts
Every fact links back to the source it came from.
- Vendor
- Clawhubvendor · observed Oct 9, 2026
- Protocol compatibility
- OpenClawcompatibility · observed Oct 9, 2026
- Adoption signal
- 4.5K downloadsadoption · observed Oct 9, 2026
- Latest release
- 2.0.0release · observed Jun 5, 2026
- Handshake status
- UNKNOWNsecurity
Install and run
Setup complexity: low.
clawhub skill install s17brd0e37a6ay53kynxvwrt6x83gx1w:paddleocr-text-recognition- Setup complexity is classified as HIGH. You must provision dedicated cloud infrastructure or an isolated VM. Do not run this directly on your local workstation.
- Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.
Contract: missing
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-bobholamovic-paddleocr-text-recognition/snapshot"
Documentation
CLAWHUB
135,231 characters of source documentation, loaded on request.
Extracted files
3 files captured from the source.
SKILL.md
---
name: paddleocr-text-recognition
description: >-
Use this skill whenever the user wants text extracted from images, photos, scans, screenshots,
or scanned PDFs. Returns exact machine-readable strings with line-level text and optional bbox
coordinates. Strong accuracy for CJK, small print, and handwritten text.
Trigger terms: OCR, 文字识别, 图片转文字, 截图识字, 提取图中文字, 扫描识字, 识字, 纯文字,
plain text extraction, 坐标, 检测框, bbox, bounding box, image to text, screenshot, photo scan,
recognize text.
license: Apache-2.0
metadata:
openclaw:
requires:
env:
- PADDLEOCR_ACCESS_TOKEN
bins:
- paddleocr
primaryEnv: PADDLEOCR_ACCESS_TOKEN
emoji: "🔤"
install:
- kind: uv
package: paddleocr
bins: [paddleocr]
---
# PaddleOCR Text Recognition
## When to Use This Skill
**Use this skill for**:
- Extract text from images (screenshots, photos, scans)
- Extract text from PDFs or document images when the goal is **line/box-level text**
- Extract text from URLs or local files that point to images/PDFs
**Do not use for**:
- Documents with tables, formulas, charts, or complex layouts — use Document Parsing instead
## Usage
### Basic OCR
From URL:
```bash
paddleocr api \
--model_type ocr \
--file_url "https://example.com/image.png"
```
From local file:
```bash
paddleocr api \
--model_type ocr \
--file_path "./document.pdf"
```
### Common Options
```bash
# With specific model
paddleocr api \
--model_type ocr \
--model PP-OCRv5 \
--file_path "./report.pdf"
# Disable preprocessing (faster, for flat/well-oriented images)
paddleocr api \
--model_type ocr \
--file_path "./document.pdf" \
--use_doc_unwarping False \
--use_doc_orientation_classify False
# Save result to file
paddleocr api \
--model_type ocr \
--file_url "https://..." \
--output result.json
# Page ranges
paddleocr api \
--model_type ocr \
--file_path "./large.pdf" \
--page_ranges "1-5,10,15-20"
```
### Output Format
```json
{
"jobId": "job-xxx",
"pages": [
{
"prunedResult": {
"rec_texts": ["Line 1", "Line 2"],
"rec_scores": [0.98, 0.95]
},
"ocrImageUrl": "https://..."
}
]
}
```
## Important Notes
**Preprocessing options**: By default, the API enables document preprocessing (unwarping and orientation classification). For flat, well-oriented images (screenshots, properly scanned documents), you can disable preprocessing for faster results:
```bash
paddleocr api --model_type ocr --file_path "./document.pdf" --use_doc_unwarping False --use_doc_orientation_classify False
```
Keep preprocessing enabled when:
- The input is a photo of a curved or folded document
- The document has significant perspective distortion
- Orientation is uncertain (rotated 90/180/270 degrees)
**Display complete results**: Always show the full extracted content to users. Do not truncate with "..." unless content exceeds 10,000 characters. When multiple pages are processed, summa_meta.json
{
"ownerId": "kn77zppfj1a2fc620aygaf9z9980ewfa",
"slug": "paddleocr-text-recognition",
"version": "2.0.0",
"publishedAt": 1780654223964
}skill-card.md
## Description: Extracts machine-readable OCR text from images, screenshots, scans, and scanned PDFs, with line-level text and optional bounding boxes. This skill is ready for commercial/non-commercial use. ## Publisher: [bobholamovic](https://clawhub.ai/user/bobholamovic) ### License/Terms of Use: Apache-2.0 ## Use Case: Developers and agent users use this skill to extract plain text from images, screenshots, photos, scans, and scanned PDFs through the PaddleOCR CLI. It is best suited for line-level OCR output and not for tables, formulas, charts, or complex document layouts. ### Deployment Geography for Use: Global ## Known Risks and Mitigations: Risk: Local images or PDFs may be sent to PaddleOCR's external API without a clear privacy confirmation. Mitigation: Use only approved data with the external OCR service; avoid submitting secrets, regulated records, internal screenshots, or confidential documents unless that service is approved for that data. Risk: Runtime dependency behavior can change if the PaddleOCR package is installed without controls. Mitigation: Use a pinned dependency version or controlled runtime for higher-risk use. Risk: OCR output may be incomplete or unsuitable for tables, formulas, charts, and complex document layouts. Mitigation: Use a document parsing workflow for complex layouts and review OCR results before relying on them. ## Reference(s): - [PaddleOCR Official CLI Documentation](https://www.paddleocr.ai/latest/en/version3.x/inference_deployment/serving/paddleocr_official_api/cli.html) - [ClawHub Skill Page](https://clawhub.ai/bobholamovic/skills/paddleocr-text-recognition) ## Skill Output: **Output Type(s):** [text, markdown, shell commands, configuration, guidance] **Output Format:** [Markdown guidance with shell command examples and JSON OCR result descriptions] **Output Parameters:** [1D] **Other Properties Related to Output:** [Returns complete extracted content when feasible; examples include line text, recognition scores, optional bounding boxes, page ranges, and saved JSON output.] ## Skill Version(s): 2.0.0 (source: ClawHub release metadata) ## Ethical Considerations: Users should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.
activepieces
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
AionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
CopilotKit
The Frontend for Agents & Generative UI. React + Angular
Machine-readable data
The same record, as JSON, for agents and crawlers.
{
"facts": [
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Clawhub",
"href": "https://clawhub.ai/bobholamovic/skills/paddleocr-text-recognition",
"sourceUrl": "https://clawhub.ai/bobholamovic/skills/paddleocr-text-recognition",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-09T05:14:42.977Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-bobholamovic-paddleocr-text-recognition/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-bobholamovic-paddleocr-text-recognition/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-09T05:14:42.977Z",
"isPublic": true
},
{
"factKey": "traction",
"category": "adoption",
"label": "Adoption signal",
"value": "4.5K downloads",
"href": "https://clawhub.ai/bobholamovic/paddleocr-text-recognition",
"sourceUrl": "https://clawhub.ai/bobholamovic/paddleocr-text-recognition",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-09T05:14:42.977Z",
"isPublic": true
},
{
"factKey": "latest_release",
"category": "release",
"label": "Latest release",
"value": "2.0.0",
"href": "https://clawhub.ai/bobholamovic/paddleocr-text-recognition",
"sourceUrl": "https://clawhub.ai/bobholamovic/paddleocr-text-recognition",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-06-05T10:10:23.964Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-bobholamovic-paddleocr-text-recognition/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-bobholamovic-paddleocr-text-recognition/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
],
"events": [
{
"eventType": "release",
"title": "Release 2.0.0",
"description": "- Major refactor: removed all local wrapper scripts and sample/reference files, switching to direct CLI usage. - SKILL.md is now much shorter and focuses on using the upstream paddleocr CLI tool directly. - Installation and environment details are simplified—API configuration only requires PADDLEOCR_ACCESS_TOKEN. - All previous internal instructions, JSON handling tips, and error guidance for wrapper scripts have been removed. - Usage examples now demonstrate running OCR directly with the paddleocr CLI, including preprocessing and output options.",
"href": "https://clawhub.ai/bobholamovic/paddleocr-text-recognition",
"sourceUrl": "https://clawhub.ai/bobholamovic/paddleocr-text-recognition",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-06-05T10:10:23.964Z",
"isPublic": true
}
]
}Record generated Oct 9, 2026.
