OCR Locally
[macOS only] Use this skill when the user requests OCR (Optical Character Recognition), image/PDF text extraction. Uses macOS native Vision/PDFKit frameworks...
Rank
62
Safety
84
Downloads
1.3k
Updated
Oct 10, 2026
Version
1.0.0
Source
CLAWHUB
About
What it does, and when to use it.
Capability contract not published. No trust telemetry is available yet. 1.3K downloads reported by the source. Last updated 10/10/2026.
Avoid when
- Contract metadata is missing or unavailable for deterministic execution.
Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing
Public facts
Every fact links back to the source it came from.
- Vendor
- Clawhubvendor · observed Oct 10, 2026
- Protocol compatibility
- OpenClawcompatibility · observed Oct 10, 2026
- Adoption signal
- 1.3K downloadsadoption · observed Oct 10, 2026
- Latest release
- 1.0.0release · observed May 4, 2026
- Handshake status
- UNKNOWNsecurity
Install and run
Setup complexity: low.
clawhub skill install s17f15dn3gc74a13v3v4g132vs856fkf:ocr-locally- Install using `clawhub skill install s17f15dn3gc74a13v3v4g132vs856fkf:ocr-locally` in an isolated environment before connecting it to live workloads.
- No published capability contract is available yet, so validate auth and request/response behavior manually.
- Review the upstream CLAWHUB listing at https://clawhub.ai/ltryee/ocr-locally before using production credentials.
Contract: missing
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-ltryee-ocr-locally/snapshot"
Documentation
CLAWHUB
29,498 characters of source documentation, loaded on request.
Extracted files
4 files captured from the source.
SKILL.md
--- name: local-ocr description: "[macOS only] Use this skill when the user requests OCR (Optical Character Recognition), image/PDF text extraction. Uses macOS native Vision/PDFKit frameworks. Triggers: '识别图片', 'OCR', '提取图片文字', '提取PDF文字', '识别PDF', 'extract text from image', 'PDF OCR'." --- # Local OCR (macOS Only) ## Overview ⚠️ **Platform Requirement**: This skill is **macOS only**. It requires macOS 10.15+ (Catalina or later) and uses macOS native frameworks: - **Vision framework** - For OCR text recognition - **PDFKit framework** - For PDF processing - **Core Graphics** - For image rendering This skill provides OCR (Optical Character Recognition) capabilities using macOS native Vision framework. It extracts text from images and PDFs without requiring any third-party libraries or internet connection. ## Platform Requirements ⚠️ **macOS Only** - This skill cannot run on Linux, Windows, or other operating systems. **Required:** - macOS 10.15+ (Catalina or later) - Vision framework (pre-installed on macOS) - PDFKit framework (pre-installed on macOS) **Why macOS Only?** - Uses `Vision` framework for OCR (macOS/iOS only) - Uses `PDFKit` framework for PDF processing (macOS/iOS only) - Uses `AppKit`/`Core Graphics` for image handling (macOS only) ## When to Use This Skill Trigger this skill when the user: - Requests OCR or image text extraction - Mentions extracting text from images, screenshots, PDF files, or scanned documents - Uses keywords like: "识别图片", "OCR", "提取文字", "提取PDF文字", "识别PDF", "extract text from image", "PDF OCR" - Provides an image file or PDF file and asks to read or extract its content ## Core Capabilities ### 1. Text Extraction from Images Use `scripts/ocr_vision_pro.swift` for comprehensive OCR with the following features: - Multi-language support (Chinese, English, Japanese, Korean, and more) - **Two output modes** (mutually exclusive): - **Text Mode** (`-t`): Output only extracted text (default) - **JSON Mode** (`-j`): Output complete raw info including text, position, and confidence as JSON - Confidence scores for each detected text block - Bounding box information (text position in image) - Output to console or file - Precise or fast recognition modes **Basic usage:** ```bash swift scripts/ocr_vision_pro.swift <image_path> ``` **With options:** ```bash swift scripts/ocr_vision_pro.swift <image_path> -l zh-Hans,en -o output.txt -f ``` ### 2. Text Extraction from PDF Files Use `scripts/pdf_ocr.swift` to extract text from PDF files with the following features: - Extract text from specific pages or all pages - Support page range specification (e.g., `1-5`, `1,3,5`) - **Two output modes** (mutually exclusive): - **Text Mode** (`-t`): Output only extracted text (default) - **JSON Mode** (`-j`): Output complete raw info as JSON - Same multi-language support as image OCR - Precise or fast recognition modes **Basic usage (all pages):** ```bash swift scripts/pdf_ocr.swift <pdf_path> ``` **With page specificati
_meta.json
{
"ownerId": "kn7azmhecfnjv9kg1t4653y21183e0wp",
"slug": "ocr-locally",
"version": "1.0.0",
"publishedAt": 1777896845706
}references/usage.md
# Local OCR Skill - Detailed Usage Guide
> ⚠️ **macOS Only** - This skill requires macOS 10.15+ (Catalina or later). It will not work on Linux, Windows, or other operating systems.
## Introduction
This document provides comprehensive usage instructions for the local-ocr skill, which uses macOS native Vision framework to perform OCR (Optical Character Recognition) on images and PDFKit for PDF processing.
## Quick Start
### Basic OCR
To extract text from an image:
```bash
swift scripts/ocr_vision_pro.swift /path/to/image.png
```
### Save Results to File (Separated Output)
```bash
swift scripts/ocr_vision_pro.swift /path/to/image.png -o result.txt
```
This will automatically create two files:
- `result.txt` - Complete extracted text
- `result_confidence.txt` - Confidence details
## Output Modes (Mutually Exclusive)
The script supports two output modes that cannot be used simultaneously:
### Text Mode (Default, `-t`)
Outputs only the extracted text. Optionally saves to file with separate confidence file.
**Console output:**
```
[Extracted text content]
```
**With `-o` option:**
Creates two files:
- `result.txt` - Complete extracted text
- `result_confidence.txt` - Confidence details
### JSON Mode (`-j`)
Outputs complete raw information as JSON to stdout. No file output options in JSON mode.
**JSON output structure:**
```json
{
"imagePath": "/path/to/image.png",
"totalBlocks": 25,
"averageConfidence": 0.85,
"blocks": [
{
"index": 1,
"text": "recognized text",
"confidence": 0.95,
"boundingBox": {
"x": 0.10,
"y": 0.20,
"width": 0.30,
"height": 0.05
}
}
]
}
```
**JSON fields:**
- `imagePath`: Path to the processed image
- `totalBlocks`: Total number of recognized text blocks
- `averageConfidence`: Average confidence score (0.0 - 1.0)
- `blocks`: Array of recognized text blocks
- `index`: Block index (1-based)
- `text`: Recognized text content
- `confidence`: Confidence score (0.0 - 1.0)
- `boundingBox`: Normalized bounding box coordinates (0.0 - 1.0)
- `x`, `y`: Top-left corner position
- `width`, `height`: Bounding box dimensions
## New Output Format (Text Mode)
## Supported Languages
The OCR script supports the following languages:
| Language Code | Language |
|---------------|----------|
| `zh-Hans` | Simplified Chinese |
| `zh-Hant` | Traditional Chinese |
| `en` | English |
| `ja` | Japanese |
| `ko` | Korean |
| `fr` | French |
| `de` | German |
| `es` | Spanish |
| `it` | Italian |
| `pt` | Portuguese |
| `ru` | Russian |
**Default languages**: `zh-Hans,zh-Hant,en`
## Command-Line Options
### -h, --help
Display help information:
```bash
swift scripts/ocr_vision_pro.swift -h
```
### -l, --language <languages>
Specify recognition language (comma-separated):
```bash
# Chinese and English
swift scripts/ocr_vision_pro.swift image.png -l zh-Hans,en
# Multiple languages
swift scripts/ocr_vision_pro.swift image.png -l zh-Hans,enskill-card.md
## Description: OCR Locally helps an agent extract text from images and PDFs on macOS using local Vision and PDFKit-based Swift scripts. This skill is ready for commercial/non-commercial use. ## Publisher: [ltryee](https://clawhub.ai/user/ltryee) ### License/Terms of Use: MIT-0 ## Use Case: Developers and agents use this skill to perform offline OCR on user-selected images, screenshots, scanned documents, and PDFs on macOS. It supports plain-text extraction, optional confidence details, JSON output, language selection, and PDF page selection. ### Deployment Geography for Use: Global ## Known Risks and Mitigations: Risk: The skill reads local image or PDF files and may write OCR output files. Mitigation: Use only intended input files, choose explicit output paths, and avoid reusing filenames that contain important content. Risk: Invalid PDF page ranges may crash or hang an OCR run. Mitigation: Prefer simple validated page ranges such as 1-5 and use the dedicated PDF script for PDF files. Risk: OCR output can be inaccurate, especially for low-confidence blocks or complex layouts. Mitigation: Review confidence details and manually verify low-confidence or business-critical extracted text. ## Reference(s): - [Usage Guide](references/usage.md) - [Apple Vision Framework Documentation](https://developer.apple.com/documentation/vision) - [VNRecognizeTextRequest Documentation](https://developer.apple.com/documentation/vision/vnrecognizetextrequest) ## Skill Output: **Output Type(s):** [Text, JSON, Files, Shell commands, Guidance] **Output Format:** [Markdown guidance with inline shell commands; OCR results are plain text, confidence text files, or JSON.] **Output Parameters:** [1D] **Other Properties Related to Output:** [Runs locally on macOS 10.15+ and may write user-specified OCR output files.] ## Skill Version(s): 1.0.0 (source: server release metadata) ## Ethical Considerations: Users should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.
AionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
activepieces
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
CopilotKit
The Frontend for Agents & Generative UI. React + Angular
Machine-readable data
The same record, as JSON, for agents and crawlers.
{
"facts": [
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Clawhub",
"href": "https://clawhub.ai/ltryee/skills/ocr-locally",
"sourceUrl": "https://clawhub.ai/ltryee/skills/ocr-locally",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-10T18:41:03.089Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-ltryee-ocr-locally/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-ltryee-ocr-locally/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-10T18:41:03.089Z",
"isPublic": true
},
{
"factKey": "traction",
"category": "adoption",
"label": "Adoption signal",
"value": "1.3K downloads",
"href": "https://clawhub.ai/ltryee/ocr-locally",
"sourceUrl": "https://clawhub.ai/ltryee/ocr-locally",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-10T18:41:03.089Z",
"isPublic": true
},
{
"factKey": "latest_release",
"category": "release",
"label": "Latest release",
"value": "1.0.0",
"href": "https://clawhub.ai/ltryee/ocr-locally",
"sourceUrl": "https://clawhub.ai/ltryee/ocr-locally",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-05-04T12:14:05.706Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-ltryee-ocr-locally/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-ltryee-ocr-locally/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
],
"events": [
{
"eventType": "release",
"title": "Release 1.0.0",
"description": "Initial release of local-ocr, providing offline OCR capabilities on macOS. - Supports image and PDF text extraction using macOS native Vision and PDFKit frameworks (macOS 10.15+ required) - Recognizes multiple languages (Chinese, English, Japanese, Korean, and more) - Two output modes: pure text or detailed JSON with confidence and bounding boxes - Handles a variety of image formats (PNG, JPEG, TIFF, BMP) and PDFs, with support for specific page selection - Command-line interface with flexible options for language, output mode, and file paths",
"href": "https://clawhub.ai/ltryee/ocr-locally",
"sourceUrl": "https://clawhub.ai/ltryee/ocr-locally",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-05-04T12:14:05.706Z",
"isPublic": true
}
]
}Record generated Oct 10, 2026.
