agentCLAWHUBUnverified

OCR Locally

[macOS only] Use this skill when the user requests OCR (Optical Character Recognition), image/PDF text extraction. Uses macOS native Vision/PDFKit frameworks...

OpenClaw

Rank

62

Safety

84

Downloads

1.3k

Updated

Oct 10, 2026

Version

1.0.0

Source

CLAWHUB

About

What it does, and when to use it.

Capability contract not published. No trust telemetry is available yet. 1.3K downloads reported by the source. Last updated 10/10/2026.

Avoid when

  • Contract metadata is missing or unavailable for deterministic execution.

Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing

Public facts

Every fact links back to the source it came from.

Vendor
Clawhubvendor · observed Oct 10, 2026
Protocol compatibility
OpenClawcompatibility · observed Oct 10, 2026
Adoption signal
1.3K downloadsadoption · observed Oct 10, 2026
Latest release
1.0.0release · observed May 4, 2026
Handshake status
UNKNOWNsecurity

Install and run

Setup complexity: low.

clawhub skill install s17f15dn3gc74a13v3v4g132vs856fkf:ocr-locally
  1. Install using `clawhub skill install s17f15dn3gc74a13v3v4g132vs856fkf:ocr-locally` in an isolated environment before connecting it to live workloads.
  2. No published capability contract is available yet, so validate auth and request/response behavior manually.
  3. Review the upstream CLAWHUB listing at https://clawhub.ai/ltryee/ocr-locally before using production credentials.

Contract: missing

curl -s "https://www.xpersona.co/api/v1/agents/clawhub-ltryee-ocr-locally/snapshot"

Documentation

CLAWHUB

29,498 characters of source documentation, loaded on request.

Extracted files

4 files captured from the source.

SKILL.md

---
name: local-ocr
description: "[macOS only] Use this skill when the user requests OCR (Optical Character Recognition), image/PDF text extraction. Uses macOS native Vision/PDFKit frameworks. Triggers: '识别图片', 'OCR', '提取图片文字', '提取PDF文字', '识别PDF', 'extract text from image', 'PDF OCR'."
---

# Local OCR (macOS Only)

## Overview

⚠️ **Platform Requirement**: This skill is **macOS only**. It requires macOS 10.15+ (Catalina or later) and uses macOS native frameworks:

- **Vision framework** - For OCR text recognition
- **PDFKit framework** - For PDF processing
- **Core Graphics** - For image rendering

This skill provides OCR (Optical Character Recognition) capabilities using macOS native Vision framework. It extracts text from images and PDFs without requiring any third-party libraries or internet connection.

## Platform Requirements

⚠️ **macOS Only** - This skill cannot run on Linux, Windows, or other operating systems.

**Required:**
- macOS 10.15+ (Catalina or later)
- Vision framework (pre-installed on macOS)
- PDFKit framework (pre-installed on macOS)

**Why macOS Only?**
- Uses `Vision` framework for OCR (macOS/iOS only)
- Uses `PDFKit` framework for PDF processing (macOS/iOS only)
- Uses `AppKit`/`Core Graphics` for image handling (macOS only)

## When to Use This Skill

Trigger this skill when the user:
- Requests OCR or image text extraction
- Mentions extracting text from images, screenshots, PDF files, or scanned documents
- Uses keywords like: "识别图片", "OCR", "提取文字", "提取PDF文字", "识别PDF", "extract text from image", "PDF OCR"
- Provides an image file or PDF file and asks to read or extract its content

## Core Capabilities

### 1. Text Extraction from Images

Use `scripts/ocr_vision_pro.swift` for comprehensive OCR with the following features:
- Multi-language support (Chinese, English, Japanese, Korean, and more)
- **Two output modes** (mutually exclusive):
  - **Text Mode** (`-t`): Output only extracted text (default)
  - **JSON Mode** (`-j`): Output complete raw info including text, position, and confidence as JSON
- Confidence scores for each detected text block
- Bounding box information (text position in image)
- Output to console or file
- Precise or fast recognition modes

**Basic usage:**
```bash
swift scripts/ocr_vision_pro.swift <image_path>
```

**With options:**
```bash
swift scripts/ocr_vision_pro.swift <image_path> -l zh-Hans,en -o output.txt -f
```

### 2. Text Extraction from PDF Files

Use `scripts/pdf_ocr.swift` to extract text from PDF files with the following features:
- Extract text from specific pages or all pages
- Support page range specification (e.g., `1-5`, `1,3,5`)
- **Two output modes** (mutually exclusive):
  - **Text Mode** (`-t`): Output only extracted text (default)
  - **JSON Mode** (`-j`): Output complete raw info as JSON
- Same multi-language support as image OCR
- Precise or fast recognition modes

**Basic usage (all pages):**
```bash
swift scripts/pdf_ocr.swift <pdf_path>
```

**With page specificati

_meta.json

{
  "ownerId": "kn7azmhecfnjv9kg1t4653y21183e0wp",
  "slug": "ocr-locally",
  "version": "1.0.0",
  "publishedAt": 1777896845706
}

references/usage.md

# Local OCR Skill - Detailed Usage Guide

> ⚠️ **macOS Only** - This skill requires macOS 10.15+ (Catalina or later). It will not work on Linux, Windows, or other operating systems.

## Introduction

This document provides comprehensive usage instructions for the local-ocr skill, which uses macOS native Vision framework to perform OCR (Optical Character Recognition) on images and PDFKit for PDF processing.

## Quick Start

### Basic OCR

To extract text from an image:

```bash
swift scripts/ocr_vision_pro.swift /path/to/image.png
```

### Save Results to File (Separated Output)

```bash
swift scripts/ocr_vision_pro.swift /path/to/image.png -o result.txt
```

This will automatically create two files:
- `result.txt` - Complete extracted text
- `result_confidence.txt` - Confidence details

## Output Modes (Mutually Exclusive)

The script supports two output modes that cannot be used simultaneously:

### Text Mode (Default, `-t`)

Outputs only the extracted text. Optionally saves to file with separate confidence file.

**Console output:**
```
[Extracted text content]
```

**With `-o` option:**
Creates two files:
- `result.txt` - Complete extracted text
- `result_confidence.txt` - Confidence details

### JSON Mode (`-j`)

Outputs complete raw information as JSON to stdout. No file output options in JSON mode.

**JSON output structure:**
```json
{
  "imagePath": "/path/to/image.png",
  "totalBlocks": 25,
  "averageConfidence": 0.85,
  "blocks": [
    {
      "index": 1,
      "text": "recognized text",
      "confidence": 0.95,
      "boundingBox": {
        "x": 0.10,
        "y": 0.20,
        "width": 0.30,
        "height": 0.05
      }
    }
  ]
}
```

**JSON fields:**
- `imagePath`: Path to the processed image
- `totalBlocks`: Total number of recognized text blocks
- `averageConfidence`: Average confidence score (0.0 - 1.0)
- `blocks`: Array of recognized text blocks
  - `index`: Block index (1-based)
  - `text`: Recognized text content
  - `confidence`: Confidence score (0.0 - 1.0)
  - `boundingBox`: Normalized bounding box coordinates (0.0 - 1.0)
    - `x`, `y`: Top-left corner position
    - `width`, `height`: Bounding box dimensions

## New Output Format (Text Mode)

## Supported Languages

The OCR script supports the following languages:

| Language Code | Language |
|---------------|----------|
| `zh-Hans` | Simplified Chinese |
| `zh-Hant` | Traditional Chinese |
| `en` | English |
| `ja` | Japanese |
| `ko` | Korean |
| `fr` | French |
| `de` | German |
| `es` | Spanish |
| `it` | Italian |
| `pt` | Portuguese |
| `ru` | Russian |

**Default languages**: `zh-Hans,zh-Hant,en`

## Command-Line Options

### -h, --help

Display help information:

```bash
swift scripts/ocr_vision_pro.swift -h
```

### -l, --language <languages>

Specify recognition language (comma-separated):

```bash
# Chinese and English
swift scripts/ocr_vision_pro.swift image.png -l zh-Hans,en

# Multiple languages
swift scripts/ocr_vision_pro.swift image.png -l zh-Hans,en

skill-card.md

## Description:

OCR Locally helps an agent extract text from images and PDFs on macOS using local Vision and PDFKit-based Swift scripts.

This skill is ready for commercial/non-commercial use.

## Publisher:

[ltryee](https://clawhub.ai/user/ltryee)

### License/Terms of Use:

MIT-0

## Use Case:

Developers and agents use this skill to perform offline OCR on user-selected images, screenshots, scanned documents, and PDFs on macOS. It supports plain-text extraction, optional confidence details, JSON output, language selection, and PDF page selection.

### Deployment Geography for Use:

Global

## Known Risks and Mitigations:

Risk: The skill reads local image or PDF files and may write OCR output files.

Mitigation: Use only intended input files, choose explicit output paths, and avoid reusing filenames that contain important content.

Risk: Invalid PDF page ranges may crash or hang an OCR run.

Mitigation: Prefer simple validated page ranges such as 1-5 and use the dedicated PDF script for PDF files.

Risk: OCR output can be inaccurate, especially for low-confidence blocks or complex layouts.

Mitigation: Review confidence details and manually verify low-confidence or business-critical extracted text.

## Reference(s):

- [Usage Guide](references/usage.md)
- [Apple Vision Framework Documentation](https://developer.apple.com/documentation/vision)
- [VNRecognizeTextRequest Documentation](https://developer.apple.com/documentation/vision/vnrecognizetextrequest)

## Skill Output:

**Output Type(s):** [Text, JSON, Files, Shell commands, Guidance]

**Output Format:** [Markdown guidance with inline shell commands; OCR results are plain text, confidence text files, or JSON.]

**Output Parameters:** [1D]

**Other Properties Related to Output:** [Runs locally on macOS 10.15+ and may write user-specified OCR output files.]

## Skill Version(s):

1.0.0 (source: server release metadata)

## Ethical Considerations:

Users should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.
Github ReposUpdated 1d agoRank 70

AionUi

Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!

MCPOPENCLAW
Github ReposUpdated 6mo agoRank 70

activepieces

AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents

OPENCLAW
Github ReposUpdated 6mo agoRank 70

cherry-studio

AI productivity studio with smart chat, autonomous agents, and 300+ assistants.

MCPOPENCLAW
Github ReposUpdated 7mo agoRank 70

CopilotKit

The Frontend for Agents & Generative UI. React + Angular

OPENCLAW

Machine-readable data

The same record, as JSON, for agents and crawlers.

{
  "facts": [
    {
      "factKey": "vendor",
      "category": "vendor",
      "label": "Vendor",
      "value": "Clawhub",
      "href": "https://clawhub.ai/ltryee/skills/ocr-locally",
      "sourceUrl": "https://clawhub.ai/ltryee/skills/ocr-locally",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-10T18:41:03.089Z",
      "isPublic": true
    },
    {
      "factKey": "protocols",
      "category": "compatibility",
      "label": "Protocol compatibility",
      "value": "OpenClaw",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-ltryee-ocr-locally/contract",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-ltryee-ocr-locally/contract",
      "sourceType": "contract",
      "confidence": "medium",
      "observedAt": "2026-10-10T18:41:03.089Z",
      "isPublic": true
    },
    {
      "factKey": "traction",
      "category": "adoption",
      "label": "Adoption signal",
      "value": "1.3K downloads",
      "href": "https://clawhub.ai/ltryee/ocr-locally",
      "sourceUrl": "https://clawhub.ai/ltryee/ocr-locally",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-10T18:41:03.089Z",
      "isPublic": true
    },
    {
      "factKey": "latest_release",
      "category": "release",
      "label": "Latest release",
      "value": "1.0.0",
      "href": "https://clawhub.ai/ltryee/ocr-locally",
      "sourceUrl": "https://clawhub.ai/ltryee/ocr-locally",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-05-04T12:14:05.706Z",
      "isPublic": true
    },
    {
      "factKey": "handshake_status",
      "category": "security",
      "label": "Handshake status",
      "value": "UNKNOWN",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-ltryee-ocr-locally/trust",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-ltryee-ocr-locally/trust",
      "sourceType": "trust",
      "confidence": "medium",
      "observedAt": null,
      "isPublic": true
    }
  ],
  "events": [
    {
      "eventType": "release",
      "title": "Release 1.0.0",
      "description": "Initial release of local-ocr, providing offline OCR capabilities on macOS. - Supports image and PDF text extraction using macOS native Vision and PDFKit frameworks (macOS 10.15+ required) - Recognizes multiple languages (Chinese, English, Japanese, Korean, and more) - Two output modes: pure text or detailed JSON with confidence and bounding boxes - Handles a variety of image formats (PNG, JPEG, TIFF, BMP) and PDFs, with support for specific page selection - Command-line interface with flexible options for language, output mode, and file paths",
      "href": "https://clawhub.ai/ltryee/ocr-locally",
      "sourceUrl": "https://clawhub.ai/ltryee/ocr-locally",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-05-04T12:14:05.706Z",
      "isPublic": true
    }
  ]
}

Record generated Oct 10, 2026.

Sponsored

Ads related to OCR Locally and adjacent AI workflows.