agentCLAWHUBUnverified

Tesseract OCR Image Text Extraction

Extract text from images using Tesseract.js (OCR). Supports multi-language recognition including Chinese and English, region recognition, character whitelist...

OpenClaw

Rank

62

Safety

84

Downloads

1.5k

Updated

Oct 10, 2026

Version

1.0.0

Source

CLAWHUB

About

What it does, and when to use it.

Capability contract not published. No trust telemetry is available yet. 1.5K downloads reported by the source. Last updated 10/10/2026.

Avoid when

  • Contract metadata is missing or unavailable for deterministic execution.

Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing

Public facts

Every fact links back to the source it came from.

Vendor
Clawhubvendor · observed Oct 10, 2026
Protocol compatibility
OpenClawcompatibility · observed Oct 10, 2026
Adoption signal
1.5K downloadsadoption · observed Oct 10, 2026
Latest release
1.0.0release · observed May 5, 2026
Handshake status
UNKNOWNsecurity

Install and run

Setup complexity: low.

clawhub skill install s1727wv2g20pc729snzcm4nf8183hy72:tesseract-image-ocr
  1. Install using `clawhub skill install s1727wv2g20pc729snzcm4nf8183hy72:tesseract-image-ocr` in an isolated environment before connecting it to live workloads.
  2. No published capability contract is available yet, so validate auth and request/response behavior manually.
  3. Review the upstream CLAWHUB listing at https://clawhub.ai/openlark/tesseract-image-ocr before using production credentials.

Contract: missing

curl -s "https://www.xpersona.co/api/v1/agents/clawhub-openlark-tesseract-image-ocr/snapshot"

Documentation

CLAWHUB

15,124 characters of source documentation, loaded on request.

Extracted files

4 files captured from the source.

SKILL.md

---
name: tesseract-image-ocr
description: Extract text from images using Tesseract.js (OCR). Supports multi-language recognition including Chinese and English, region recognition, character whitelist filtering, text orientation detection, and can run in a Node.js environment.
---

# Tesseract OCR Image Text Extraction

Extract text content from images based on Tesseract.js (the WebAssembly port of the Tesseract OCR engine).

## Use Cases

Use when users need "image to text," "OCR recognition," "extract text from images," "screenshot character recognition," "scan to text," or "image text orientation detection."

## Core Capabilities

- Recognize text from local images or image URLs
- Support for 100+ languages, with the ability to specify multiple languages simultaneously (e.g., `['eng', 'chi_sim']`)
- Support for specifying recognition regions (`--rectangle`), character whitelists (`--whitelist`)
- Support for text orientation and script detection (`--detect`)
- Support for switching page segmentation modes (`--psm`) and OCR engine modes (`--oem`)
- Output formats: `text` (default), `hocr`, `blocks` (JSON), `tsv`

## Limitations

- Does not support PDF files
- Does not modify the Tesseract recognition model to improve accuracy
- Requires a Node.js environment (this Skill uses Node.js scripts)

---

## Workflow

### 1. Confirm Environment

```shell
node -v && npm ls tesseract.js 2>/dev/null || echo "tesseract.js not installed"
```

If not installed:

```shell
cd /root/.openclaw/workspace/skills/tesseract-ocr && npm init -y > /dev/null 2>&1 && npm install tesseract.js
```

### 2. Execute Recognition

Basic command:

```shell
node scripts/ocr.js <image-path-or-url> [--options]
```

**Parameter Descriptions:**

| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `<image>` | Required | — | Local path or HTTPS URL |
| `--lang` | string | `eng` | Language code(s), multiple joined with `+`, e.g., `eng+chi_sim` |
| `--psm` | number | — | Page segmentation mode (see PSM table below) |
| `--oem` | number | — | OCR engine mode (see OEM table below) |
| `--whitelist` | string | — | Character whitelist, e.g., `0123456789` to recognize only digits |
| `--rectangle` | string | — | Recognition region, format `top,left,width,height` |
| `--output` | string | `text` | Output format: `text` / `hocr` / `blocks` / `tsv` |
| `--detect` | flag | — | Detect text orientation and script (does not perform OCR) |
| `--dpi` | number | — | Manually specify image DPI |

**Common Examples:**

```shell
# Basic mixed Chinese-English recognition
node scripts/ocr.js photo.jpg --lang chi_sim+eng

# Recognize digits only (license plates, CAPTCHAs, etc.)
node scripts/ocr.js captcha.png --whitelist 0123456789

# Column-based recognition (suitable for vertical Chinese text)
node scripts/ocr.js scroll.jpg --lang chi_sim --psm 4

# Specify a region for recognition
node scripts/ocr.js receipt.png --rectangle 50,100,400,200

# Detect image text orient

_meta.json

{
  "ownerId": "kn75qrtb885pznwsjwsh18dvf1813bv0",
  "slug": "tesseract-image-ocr",
  "version": "1.0.0",
  "publishedAt": 1777945358362
}

references/api.md

# Tesseract.js Full API Reference

> Based on the [Tesseract.js Official API Documentation](https://github.com/naptha/tesseract.js/blob/master/docs/api.md)

---

## createWorker(langs, oem, options): Worker

Creates a Tesseract.js Worker instance. A Worker manages a single Tesseract instance within a Web Worker (browser) or Worker Thread (Node.js).

**Parameters:**

| Parameter | Type | Description |
|-----------|------|-------------|
| `langs` | string\|string[] | Language code(s), e.g., `'eng'` or `['eng', 'chi_sim']` |
| `oem` | number | OCR engine mode (see OEM table) |
| `options` | object | Custom options |

**options Fields:**

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `corePath` | string | — | Path to the tesseract.js-core directory (**Note: points to a directory, not a single .js file**) |
| `langPath` | string | — | Download path for training language data (no trailing `/`) |
| `workerPath` | string | — | Download path for worker script |
| `cachePath` | string | — | Cache path; more commonly used in Node.js |
| `cacheMethod` | string | `'write'` | Cache strategy: `write`/`readOnly`/`refresh`/`none` |
| `legacyCore` | boolean | `false` | Whether to download Legacy model support code |
| `legacyLang` | boolean | `false` | Whether to download Legacy language data |
| `workerBlobURL` | boolean | `true` | Whether to load the Worker using a Blob URL |
| `gzip` | boolean | `true` | Whether remote training data is gzip compressed |
| `logger` | function | — | Progress callback, e.g., `m => console.log(m)` |
| `errorHandler` | function | — | Worker error callback |

---

## worker.recognize(image, options, output, jobId): Promise

Core OCR recognition method.

**Parameters:**

| Parameter | Type | Description |
|-----------|------|-------------|
| `image` | string\|Buffer\|... | Image (see supported image formats) |
| `options` | object | Recognition options |
| `options.rectangle` | object | Restrict recognition region `{top, left, width, height}` |
| `output` | object | Output format toggles, e.g., `{ hocr: true, blocks: true }` |
| `jobId` | string | Optional job ID |

**Returns:** `{ jobId, data: { text, hocr?, blocks?, tsv?, ... } }`

The returned `data.text` is the plain text result. An empty result is returned even if no text is detected; an exception will not be thrown.

---

## worker.setParameters(params, jobId): Promise

Sets Tesseract engine parameters (calls SetVariable).

**Common Parameters:**

| Parameter Name | Type | Default | Description |
|----------------|------|---------|-------------|
| `tessedit_pageseg_mode` | string | `'3'` | Page segmentation mode |
| `tessedit_char_whitelist` | string | `''` | Character whitelist |
| `preserve_interword_spaces` | string | `'0'` | Preserve inter-word spaces |
| `user_defined_dpi` | string | `''` | Manually specify DPI |

> `setParameters` cannot modify `oem`; to change the OEM, use `worker.reinitialize()`.

---

## worker.reinitialize(langs, o

skill-card.md

## Description:

Extract text from images using Tesseract.js (OCR). Supports multi-language recognition including Chinese and English, region recognition, character whitelist filtering, text orientation detection, and can run in a Node.js environment.

This skill is ready for commercial/non-commercial use.

## Publisher:

[openlark](https://clawhub.ai/user/openlark)

### License/Terms of Use:

MIT-0

## Use Case:

Developers and agents use this skill to extract text from local or trusted remote images, including screenshots and scanned images, with language, region, whitelist, orientation, and output-format controls.

### Deployment Geography for Use:

Global

## Known Risks and Mitigations:

Risk: npm dependency downloads may introduce unreviewed package or supply-chain changes.

Mitigation: Install only in environments where dependency downloads are allowed, and pin tesseract.js with a reviewed lockfile before operational use.

Risk: Remote image URLs can expose private content to untrusted hosts or fetch content from untrusted locations.

Mitigation: Prefer local image files for private or regulated data, and use remote URLs only when the host is trusted.

## Reference(s):

- [Tesseract.js API Reference](references/api.md)
- [Tesseract.js Official API Documentation](https://github.com/naptha/tesseract.js/blob/master/docs/api.md)
- [Tesseract.js Language List](https://github.com/naptha/tesseract.js/blob/master/docs/tesseract_lang_list.md)

## Skill Output:

**Output Type(s):** [text, code, shell commands, configuration, guidance]

**Output Format:** [Markdown guidance with shell and JavaScript examples; OCR execution can return plain text, JSON, hOCR, or TSV.]

**Output Parameters:** [1D]

**Other Properties Related to Output:** [Supports language selection, region selection, character whitelists, page segmentation modes, OCR engine modes, DPI overrides, and orientation detection.]

## Skill Version(s):

1.0.0 (source: server release evidence)

## Ethical Considerations:

Users should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.
Github ReposUpdated 19h agoRank 70

AionUi

Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!

MCPOPENCLAW
Github ReposUpdated 6mo agoRank 70

activepieces

AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents

OPENCLAW
Github ReposUpdated 6mo agoRank 70

cherry-studio

AI productivity studio with smart chat, autonomous agents, and 300+ assistants.

MCPOPENCLAW
Github ReposUpdated 7mo agoRank 70

CopilotKit

The Frontend for Agents & Generative UI. React + Angular

OPENCLAW

Machine-readable data

The same record, as JSON, for agents and crawlers.

{
  "facts": [
    {
      "factKey": "vendor",
      "category": "vendor",
      "label": "Vendor",
      "value": "Clawhub",
      "href": "https://clawhub.ai/openlark/skills/tesseract-image-ocr",
      "sourceUrl": "https://clawhub.ai/openlark/skills/tesseract-image-ocr",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-10T11:47:45.407Z",
      "isPublic": true
    },
    {
      "factKey": "protocols",
      "category": "compatibility",
      "label": "Protocol compatibility",
      "value": "OpenClaw",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-openlark-tesseract-image-ocr/contract",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-openlark-tesseract-image-ocr/contract",
      "sourceType": "contract",
      "confidence": "medium",
      "observedAt": "2026-10-10T11:47:45.407Z",
      "isPublic": true
    },
    {
      "factKey": "traction",
      "category": "adoption",
      "label": "Adoption signal",
      "value": "1.5K downloads",
      "href": "https://clawhub.ai/openlark/tesseract-image-ocr",
      "sourceUrl": "https://clawhub.ai/openlark/tesseract-image-ocr",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-10T11:47:45.407Z",
      "isPublic": true
    },
    {
      "factKey": "latest_release",
      "category": "release",
      "label": "Latest release",
      "value": "1.0.0",
      "href": "https://clawhub.ai/openlark/tesseract-image-ocr",
      "sourceUrl": "https://clawhub.ai/openlark/tesseract-image-ocr",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-05-05T01:42:38.362Z",
      "isPublic": true
    },
    {
      "factKey": "handshake_status",
      "category": "security",
      "label": "Handshake status",
      "value": "UNKNOWN",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-openlark-tesseract-image-ocr/trust",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-openlark-tesseract-image-ocr/trust",
      "sourceType": "trust",
      "confidence": "medium",
      "observedAt": null,
      "isPublic": true
    }
  ],
  "events": [
    {
      "eventType": "release",
      "title": "Release 1.0.0",
      "description": "- Initial release of tesseract-image-ocr. - Extract text from images using Tesseract.js with support for over 100 languages, including Chinese and English. - Features include region recognition, character whitelist filtering, text orientation detection, and adjustable output formats (text, hocr, blocks, tsv). - Supports multiple recognition parameters: language selection, page segmentation modes, OCR engine modes, and DPI specification. - Requires Node.js environment and does not support PDF files.",
      "href": "https://clawhub.ai/openlark/tesseract-image-ocr",
      "sourceUrl": "https://clawhub.ai/openlark/tesseract-image-ocr",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-05-05T01:42:38.362Z",
      "isPublic": true
    }
  ]
}

Record generated Oct 10, 2026.

Sponsored

Ads related to Tesseract OCR Image Text Extraction and adjacent AI workflows.