Tesseract OCR Image Text Extraction
Extract text from images using Tesseract.js (OCR). Supports multi-language recognition including Chinese and English, region recognition, character whitelist...
Rank
62
Safety
84
Downloads
1.5k
Updated
Oct 10, 2026
Version
1.0.0
Source
CLAWHUB
About
What it does, and when to use it.
Capability contract not published. No trust telemetry is available yet. 1.5K downloads reported by the source. Last updated 10/10/2026.
Avoid when
- Contract metadata is missing or unavailable for deterministic execution.
Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing
Public facts
Every fact links back to the source it came from.
- Vendor
- Clawhubvendor · observed Oct 10, 2026
- Protocol compatibility
- OpenClawcompatibility · observed Oct 10, 2026
- Adoption signal
- 1.5K downloadsadoption · observed Oct 10, 2026
- Latest release
- 1.0.0release · observed May 5, 2026
- Handshake status
- UNKNOWNsecurity
Install and run
Setup complexity: low.
clawhub skill install s1727wv2g20pc729snzcm4nf8183hy72:tesseract-image-ocr- Install using `clawhub skill install s1727wv2g20pc729snzcm4nf8183hy72:tesseract-image-ocr` in an isolated environment before connecting it to live workloads.
- No published capability contract is available yet, so validate auth and request/response behavior manually.
- Review the upstream CLAWHUB listing at https://clawhub.ai/openlark/tesseract-image-ocr before using production credentials.
Contract: missing
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-openlark-tesseract-image-ocr/snapshot"
Documentation
CLAWHUB
15,124 characters of source documentation, loaded on request.
Extracted files
4 files captured from the source.
SKILL.md
--- name: tesseract-image-ocr description: Extract text from images using Tesseract.js (OCR). Supports multi-language recognition including Chinese and English, region recognition, character whitelist filtering, text orientation detection, and can run in a Node.js environment. --- # Tesseract OCR Image Text Extraction Extract text content from images based on Tesseract.js (the WebAssembly port of the Tesseract OCR engine). ## Use Cases Use when users need "image to text," "OCR recognition," "extract text from images," "screenshot character recognition," "scan to text," or "image text orientation detection." ## Core Capabilities - Recognize text from local images or image URLs - Support for 100+ languages, with the ability to specify multiple languages simultaneously (e.g., `['eng', 'chi_sim']`) - Support for specifying recognition regions (`--rectangle`), character whitelists (`--whitelist`) - Support for text orientation and script detection (`--detect`) - Support for switching page segmentation modes (`--psm`) and OCR engine modes (`--oem`) - Output formats: `text` (default), `hocr`, `blocks` (JSON), `tsv` ## Limitations - Does not support PDF files - Does not modify the Tesseract recognition model to improve accuracy - Requires a Node.js environment (this Skill uses Node.js scripts) --- ## Workflow ### 1. Confirm Environment ```shell node -v && npm ls tesseract.js 2>/dev/null || echo "tesseract.js not installed" ``` If not installed: ```shell cd /root/.openclaw/workspace/skills/tesseract-ocr && npm init -y > /dev/null 2>&1 && npm install tesseract.js ``` ### 2. Execute Recognition Basic command: ```shell node scripts/ocr.js <image-path-or-url> [--options] ``` **Parameter Descriptions:** | Parameter | Type | Default | Description | |-----------|------|---------|-------------| | `<image>` | Required | — | Local path or HTTPS URL | | `--lang` | string | `eng` | Language code(s), multiple joined with `+`, e.g., `eng+chi_sim` | | `--psm` | number | — | Page segmentation mode (see PSM table below) | | `--oem` | number | — | OCR engine mode (see OEM table below) | | `--whitelist` | string | — | Character whitelist, e.g., `0123456789` to recognize only digits | | `--rectangle` | string | — | Recognition region, format `top,left,width,height` | | `--output` | string | `text` | Output format: `text` / `hocr` / `blocks` / `tsv` | | `--detect` | flag | — | Detect text orientation and script (does not perform OCR) | | `--dpi` | number | — | Manually specify image DPI | **Common Examples:** ```shell # Basic mixed Chinese-English recognition node scripts/ocr.js photo.jpg --lang chi_sim+eng # Recognize digits only (license plates, CAPTCHAs, etc.) node scripts/ocr.js captcha.png --whitelist 0123456789 # Column-based recognition (suitable for vertical Chinese text) node scripts/ocr.js scroll.jpg --lang chi_sim --psm 4 # Specify a region for recognition node scripts/ocr.js receipt.png --rectangle 50,100,400,200 # Detect image text orient
_meta.json
{
"ownerId": "kn75qrtb885pznwsjwsh18dvf1813bv0",
"slug": "tesseract-image-ocr",
"version": "1.0.0",
"publishedAt": 1777945358362
}references/api.md
# Tesseract.js Full API Reference
> Based on the [Tesseract.js Official API Documentation](https://github.com/naptha/tesseract.js/blob/master/docs/api.md)
---
## createWorker(langs, oem, options): Worker
Creates a Tesseract.js Worker instance. A Worker manages a single Tesseract instance within a Web Worker (browser) or Worker Thread (Node.js).
**Parameters:**
| Parameter | Type | Description |
|-----------|------|-------------|
| `langs` | string\|string[] | Language code(s), e.g., `'eng'` or `['eng', 'chi_sim']` |
| `oem` | number | OCR engine mode (see OEM table) |
| `options` | object | Custom options |
**options Fields:**
| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `corePath` | string | — | Path to the tesseract.js-core directory (**Note: points to a directory, not a single .js file**) |
| `langPath` | string | — | Download path for training language data (no trailing `/`) |
| `workerPath` | string | — | Download path for worker script |
| `cachePath` | string | — | Cache path; more commonly used in Node.js |
| `cacheMethod` | string | `'write'` | Cache strategy: `write`/`readOnly`/`refresh`/`none` |
| `legacyCore` | boolean | `false` | Whether to download Legacy model support code |
| `legacyLang` | boolean | `false` | Whether to download Legacy language data |
| `workerBlobURL` | boolean | `true` | Whether to load the Worker using a Blob URL |
| `gzip` | boolean | `true` | Whether remote training data is gzip compressed |
| `logger` | function | — | Progress callback, e.g., `m => console.log(m)` |
| `errorHandler` | function | — | Worker error callback |
---
## worker.recognize(image, options, output, jobId): Promise
Core OCR recognition method.
**Parameters:**
| Parameter | Type | Description |
|-----------|------|-------------|
| `image` | string\|Buffer\|... | Image (see supported image formats) |
| `options` | object | Recognition options |
| `options.rectangle` | object | Restrict recognition region `{top, left, width, height}` |
| `output` | object | Output format toggles, e.g., `{ hocr: true, blocks: true }` |
| `jobId` | string | Optional job ID |
**Returns:** `{ jobId, data: { text, hocr?, blocks?, tsv?, ... } }`
The returned `data.text` is the plain text result. An empty result is returned even if no text is detected; an exception will not be thrown.
---
## worker.setParameters(params, jobId): Promise
Sets Tesseract engine parameters (calls SetVariable).
**Common Parameters:**
| Parameter Name | Type | Default | Description |
|----------------|------|---------|-------------|
| `tessedit_pageseg_mode` | string | `'3'` | Page segmentation mode |
| `tessedit_char_whitelist` | string | `''` | Character whitelist |
| `preserve_interword_spaces` | string | `'0'` | Preserve inter-word spaces |
| `user_defined_dpi` | string | `''` | Manually specify DPI |
> `setParameters` cannot modify `oem`; to change the OEM, use `worker.reinitialize()`.
---
## worker.reinitialize(langs, oskill-card.md
## Description: Extract text from images using Tesseract.js (OCR). Supports multi-language recognition including Chinese and English, region recognition, character whitelist filtering, text orientation detection, and can run in a Node.js environment. This skill is ready for commercial/non-commercial use. ## Publisher: [openlark](https://clawhub.ai/user/openlark) ### License/Terms of Use: MIT-0 ## Use Case: Developers and agents use this skill to extract text from local or trusted remote images, including screenshots and scanned images, with language, region, whitelist, orientation, and output-format controls. ### Deployment Geography for Use: Global ## Known Risks and Mitigations: Risk: npm dependency downloads may introduce unreviewed package or supply-chain changes. Mitigation: Install only in environments where dependency downloads are allowed, and pin tesseract.js with a reviewed lockfile before operational use. Risk: Remote image URLs can expose private content to untrusted hosts or fetch content from untrusted locations. Mitigation: Prefer local image files for private or regulated data, and use remote URLs only when the host is trusted. ## Reference(s): - [Tesseract.js API Reference](references/api.md) - [Tesseract.js Official API Documentation](https://github.com/naptha/tesseract.js/blob/master/docs/api.md) - [Tesseract.js Language List](https://github.com/naptha/tesseract.js/blob/master/docs/tesseract_lang_list.md) ## Skill Output: **Output Type(s):** [text, code, shell commands, configuration, guidance] **Output Format:** [Markdown guidance with shell and JavaScript examples; OCR execution can return plain text, JSON, hOCR, or TSV.] **Output Parameters:** [1D] **Other Properties Related to Output:** [Supports language selection, region selection, character whitelists, page segmentation modes, OCR engine modes, DPI overrides, and orientation detection.] ## Skill Version(s): 1.0.0 (source: server release evidence) ## Ethical Considerations: Users should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.
AionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
activepieces
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
CopilotKit
The Frontend for Agents & Generative UI. React + Angular
Machine-readable data
The same record, as JSON, for agents and crawlers.
{
"facts": [
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Clawhub",
"href": "https://clawhub.ai/openlark/skills/tesseract-image-ocr",
"sourceUrl": "https://clawhub.ai/openlark/skills/tesseract-image-ocr",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-10T11:47:45.407Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-openlark-tesseract-image-ocr/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-openlark-tesseract-image-ocr/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-10T11:47:45.407Z",
"isPublic": true
},
{
"factKey": "traction",
"category": "adoption",
"label": "Adoption signal",
"value": "1.5K downloads",
"href": "https://clawhub.ai/openlark/tesseract-image-ocr",
"sourceUrl": "https://clawhub.ai/openlark/tesseract-image-ocr",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-10T11:47:45.407Z",
"isPublic": true
},
{
"factKey": "latest_release",
"category": "release",
"label": "Latest release",
"value": "1.0.0",
"href": "https://clawhub.ai/openlark/tesseract-image-ocr",
"sourceUrl": "https://clawhub.ai/openlark/tesseract-image-ocr",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-05-05T01:42:38.362Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-openlark-tesseract-image-ocr/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-openlark-tesseract-image-ocr/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
],
"events": [
{
"eventType": "release",
"title": "Release 1.0.0",
"description": "- Initial release of tesseract-image-ocr. - Extract text from images using Tesseract.js with support for over 100 languages, including Chinese and English. - Features include region recognition, character whitelist filtering, text orientation detection, and adjustable output formats (text, hocr, blocks, tsv). - Supports multiple recognition parameters: language selection, page segmentation modes, OCR engine modes, and DPI specification. - Requires Node.js environment and does not support PDF files.",
"href": "https://clawhub.ai/openlark/tesseract-image-ocr",
"sourceUrl": "https://clawhub.ai/openlark/tesseract-image-ocr",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-05-05T01:42:38.362Z",
"isPublic": true
}
]
}Record generated Oct 10, 2026.
