Private Document AI with OpenVINO
Private local document AI for Intel hardware. Parse PDFs, invoices, screenshots, and diagrams with MinerU 2.5 on OpenVINO GenAI, keep the model warm in a loc... Skill: Private Document AI with OpenVINO Owner: zhuo-yoyowz Summary: Private local document AI for Intel hardware. Parse PDFs, invoices, screenshots, and diagrams with MinerU 2.5 on OpenVINO GenAI, keep the model warm in a loc... Tags: ai-pc:0.4.1, document-ai:0.4.1, fastapi:0.4.1, invoice:0.4.1, latest:0.4.1, latest openvino document-ai:0.1.2, latest openvino document-ai ocr invoice notebook:0.1.4, local-ai:0.4.1, l
Rank
62
Safety
84
Downloads
1.2k
Updated
Oct 11, 2026
Version
0.4.1
Source
CLAWHUB
About
What it does, and when to use it.
Capability contract not published. No trust telemetry is available yet. 1.2K downloads reported by the source. Last updated 10/11/2026.
Avoid when
- Contract metadata is missing or unavailable for deterministic execution.
Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing
Public facts
Every fact links back to the source it came from.
- Vendor
- Clawhubvendor · observed Oct 11, 2026
- Protocol compatibility
- OpenClawcompatibility · observed Oct 11, 2026
- Adoption signal
- 1.2K downloadsadoption · observed Oct 11, 2026
- Latest release
- 0.4.1release · observed Jul 20, 2026
- Handshake status
- UNKNOWNsecurity
Install and run
Setup complexity: low.
clawhub skill install s1793aef4h4fwmpc33fnt9fg998602k3:local-document-ai-openvino- Setup complexity is classified as HIGH. You must provision dedicated cloud infrastructure or an isolated VM. Do not run this directly on your local workstation.
- Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.
Contract: missing
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-zhuo-yoyowz-local-document-ai-openvino/snapshot"
Documentation
CLAWHUB
152,332 characters of source documentation, loaded on request.
Extracted files
5 files captured from the source.
SKILL.md
---
name: local-document-ai-openvino
description: Private local document AI for Intel hardware. Parse PDFs, invoices, screenshots, and diagrams with MinerU 2.5 on OpenVINO GenAI, keep the model warm in a local service, and output structured JSON/Markdown with user-defined invoice fields.
---
# Private Document AI with OpenVINO
Turn local PDFs, invoices, screenshots, and diagrams into one of two useful outcomes:
1. `to-data`: classify the document and extract structured fields, tables, and JSON, including user-requested key fields.
2. `to-code`: turn screenshots, forms, and architecture diagrams into code or Jupyter notebook scaffolds.
Everything runs locally and is built for Intel CPU/GPU acceleration with OpenVINO GenAI.
The default device is `CPU` for workshop stability. Set `MINERU_OPENVINO_DEVICE=GPU` or `AUTO` only after validating the target AI PC.
The default user experience is app-like:
- the first call auto-starts a local service
- HTTP/FastAPI service is preferred when available
- the standard-library IPC service is used as a fallback
- direct CLI parsing remains available with `--no-server`
- the model stays resident after warmup so later calls avoid repeat model loading
The default runtime path in this release is:
- MinerU 2.5 Pro
- preconverted OpenVINO INT4 model bundle
- local PDF rendering with `pypdfium2`
- no local model export step
## Why install this skill
Install this when you want one local workflow for:
- invoice and receipt extraction
- private PDF understanding
- table and key-value extraction
- architecture diagram to notebook generation
- screenshot to HTML/React scaffold generation
This skill is especially good for demos because it already includes:
- medical invoice `to-data` flows
- restaurant invoice `to-data` flows
- custom invoice field extraction such as invoice number, date, seller, and amount due
- architecture diagram `to-code -> jupyter-notebook` flows
- local HTML reports for easy review and screenshots
## 30-second start
Check the environment:
```bash
python "{baseDir}/scripts/check_env.py"
```
Install with the fastest available local installer path:
```powershell
powershell -ExecutionPolicy Bypass -File "{baseDir}/scripts/install.ps1"
```
Warm up the local document AI service:
```bash
python "{baseDir}/scripts/run_skill.py" --warmup-server
```
Or run directly from the CLI. Server mode is automatic by default:
```bash
python "{baseDir}/scripts/run_skill.py" --mode to-data --file "/absolute/path/to/invoice.pdf" --out "/absolute/path/to/artifacts/invoice_data" --extract "tables,entities,kv_pairs"
```
For invoice demos with custom key fields:
```bash
python "{baseDir}/scripts/run_skill.py" --mode to-data --file "/absolute/path/to/invoice.pdf" --out "/absolute/path/to/artifacts/invoice_data" --extract "tables,entities,kv_pairs" --fields "invoice_number,invoice_date,total_amount,vendor_name"
```
For repeated workshop demos, prefer the persistent local server mode:
```bash
python "{baseD_meta.json
{
"ownerId": "kn70rgyfb422qnxmr4ryhz1k698601ww",
"slug": "local-document-ai-openvino",
"version": "0.4.1",
"publishedAt": 1784530501244
}references/mode_guide.md
# Mode Guide This file defines how each implemented mode should behave. ## Shared Rules Always: 1. Parse first. 2. Write `parsed.json`. 3. Read from `parsed.json` for downstream work. 4. Save final outputs under `task_output/`. 5. Save a source map or traceability file for downstream modes. Do not: - generate directly from raw OCR text when `parsed.json` is available - invent facts not supported by the document - hide uncertainty or warnings that MinerU OpenVINO inference was not used ## Mode: `parse` ### Goal Create the canonical structured representation only. ### Inputs - `file` - optional `out` ### Outputs - `parsed.json` - `parsed.md` - `tables/` - `figures/` ### Return Summary Include: - file processed - page count - counts of headings, paragraphs, tables, formulas, figures, charts if available - output folder path - warnings if any ## Mode: `to-code` ### Goal Turn a document into code-oriented artifacts. ### Best-Fit Inputs - UI mockups - screenshots - forms - product specs - brochures - workflow documents ### Allowed Outputs - `component_map.json` - `field_schema.json` - `app.jsx` - `index.html` - `styles.css` - `notes.md` - `traceability.json` ### Behavior - infer sections and components from parsed structure - preserve labels, fields, buttons, lists, and tables - use placeholders when business rules are not explicit - record assumptions in `notes.md` and `traceability.json` ### Good Examples - brochure image to landing page scaffold - form screenshot to React form skeleton - admin spec PDF to HTML + JSON field schema ## Mode: `to-data` ### Goal Extract machine-readable data for automation. ### Best-Fit Inputs - invoices - reports - forms - schedules - tables - structured business documents ### Allowed Outputs - `entities.json` - `kv_pairs.json` - `normalized.json` - `requested_fields.json` - `requested_fields_record.json` - `tables.csv` - `table_index.json` - `traceability.json` ### Behavior - keep original text and normalized values when useful - preserve source block references for each record - separate extraction from interpretation - when the user provides a custom field list, generate a focused structured output for only those requested fields ### Good Examples - invoice PDF to normalized invoice JSON - invoice PDF to a custom JSON record containing only `invoice_number`, `invoice_date`, `total_amount`, and `vendor_name` - annual report to CSV tables + entity summary - application form to field-value JSON ## Mode Selection Hints Prefer: - `parse` when the user mainly wants structured OCR output - `to-code` when the user wants implementation artifacts - `to-data` when the user wants extraction/normalization If unsure: - default to `parse` - then explain which downstream modes are available next
references/output_contracts.md
# Output Contracts
This file defines the folder layout and file contracts.
## Default folder layout
```text
artifacts/<document_stem>/
├── parsed.json
├── parsed.md
├── traceability.json
├── tables/
├── figures/
└── task_output/
```
If the user passes `out=...`, use that directory instead.
## Parse outputs
### `parsed.json`
Required for every successful run.
### `parsed.md`
Required for every successful run.
Purpose:
- human-readable rendering of the parse result
### `tables/`
Optional.
Write extracted CSVs or table assets here.
### `figures/`
Optional.
Write extracted figures here.
---
## Downstream outputs
### `task_output/`
Required for non-parse modes.
Examples:
- `task_output/app.jsx`
- `task_output/index.html`
- `task_output/entities.json`
- `task_output/slide_outline.md`
### `traceability.json`
Required for non-parse modes.
Purpose:
- map generated artifacts back to source page/block IDs
- record assumptions or low-confidence derivations
Example:
```json
{
"artifact": "task_output/app.jsx",
"mappings": [
{
"generated_unit_id": "component.signup_email_field",
"generated_text": "Email input field with label and helper text",
"source_refs": [
{"page_id": "page_1", "block_id": "p1_b12"},
{"page_id": "page_1", "block_id": "p1_b13"}
],
"assumption": "Validation rule was not explicit in source."
}
]
}
```
## Failure contract
If a run fails:
- do not create empty success artifacts
- optionally write `error.json` with:
- stage
- message
- input file
- mode
- timestamp
Example:
```json
{
"stage": "parse",
"message": "Unsupported file type",
"input_file": "./docs/foo.xyz",
"mode": "parse",
"timestamp": "2026-04-08T16:00:00Z"
}
```
## Naming conventions
- use lowercase snake_case for filenames
- use stable IDs for pages, blocks, tables, and figures
- use relative paths inside JSON when files live inside the same artifact folder
## Quality notes
- prefer explicit omission over silent loss
- if tables or formulas are detected but not reconstructed, note that in `parse_info.warnings`
- if output is partially inferred, record it in `traceability.json`references/schema.md
# Canonical Document Schema
This file defines the stable intermediate representation used by this skill.
## Purpose
All downstream modes must consume the canonical schema instead of raw document text.
Benefits:
- stable contract between parse and transform stages
- better grounding
- traceability from outputs back to source blocks
- easier testing and future model replacement
## Top-level structure
```json
{
"schema_version": "1.0",
"document_id": "string",
"source": {},
"parse_info": {},
"pages": [],
"tables": [],
"figures": [],
"entities": [],
"outputs": {}
}
```
## Field definitions
### `schema_version`
Version of this schema.
Type: `string`
### `document_id`
Stable ID for the current document run.
Recommended format:
`<file_stem>-<short_hash>`
### `source`
Information about the original input.
```json
{
"input_path": "string",
"input_type": "pdf|image",
"filename": "string",
"sha256": "string|null"
}
```
### `parse_info`
Information about the parser run.
```json
{
"engine": "local-document-ai-openvino",
"engine_version": "string",
"mode": "parse|to-code|to-data",
"created_at": "ISO-8601 string",
"warnings": ["string"],
"confidence_note": "string|null"
}
```
### `pages`
Ordered list of parsed pages.
```json
[
{
"page_id": "page_1",
"page_index": 1,
"width": 2480,
"height": 3508,
"blocks": []
}
]
```
### `blocks`
Ordered list of page blocks.
```json
{
"block_id": "p1_b1",
"type": "heading|paragraph|list|table|formula|chart|figure|seal|kv_pair|footer|header|caption|unknown",
"bbox": [0, 0, 100, 50],
"reading_order": 1,
"text": "string",
"markdown": "string|null",
"latex": "string|null",
"html": "string|null",
"confidence": 0.0,
"attributes": {
"heading_level": 1,
"language": "en",
"is_rotated": false
},
"relations": {
"parent_block_id": null,
"caption_for": null,
"table_id": null,
"figure_id": null
}
}
```
#### Block rules
- `page_id + block_id` must be unique
- `reading_order` must be monotonic within a page
- `type` should be as specific as possible
- `text` is plain normalized text
- `markdown` is optional rendered text
- `latex` is only for formulas
- `html` is optional for table/structured fragments
### `tables`
Normalized structured tables.
```json
[
{
"table_id": "t1",
"page_id": "page_2",
"bbox": [10, 10, 200, 150],
"caption": "Quarterly Revenue",
"headers": ["Quarter", "Revenue"],
"rows": [
["Q1", "$1M"],
["Q2", "$1.2M"]
],
"csv_path": "tables/t1.csv",
"source_block_ids": ["p2_b8"]
}
]
```
### `figures`
Saved figure assets.
```json
[
{
"figure_id": "f1",
"page_id": "page_3",
"bbox": [20, 20, 300, 200],
"caption": "Architecture Diagram",
"asset_path": "figures/f1.png",
"source_block_ids": ["p3_b4"]
}
]
```
### `entities`
Optional normalized entities.
```json
[
{
"entity_id": "e1",
"type": "invoice_number|date|peAionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
activepieces
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
CopilotKit
The Frontend for Agents & Generative UI. React + Angular
Machine-readable data
The same record, as JSON, for agents and crawlers.
{
"facts": [
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Clawhub",
"href": "https://clawhub.ai/zhuo-yoyowz/skills/local-document-ai-openvino",
"sourceUrl": "https://clawhub.ai/zhuo-yoyowz/skills/local-document-ai-openvino",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-11T00:19:55.284Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-zhuo-yoyowz-local-document-ai-openvino/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-zhuo-yoyowz-local-document-ai-openvino/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-11T00:19:55.284Z",
"isPublic": true
},
{
"factKey": "traction",
"category": "adoption",
"label": "Adoption signal",
"value": "1.2K downloads",
"href": "https://clawhub.ai/zhuo-yoyowz/local-document-ai-openvino",
"sourceUrl": "https://clawhub.ai/zhuo-yoyowz/local-document-ai-openvino",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-11T00:19:55.284Z",
"isPublic": true
},
{
"factKey": "latest_release",
"category": "release",
"label": "Latest release",
"value": "0.4.1",
"href": "https://clawhub.ai/zhuo-yoyowz/local-document-ai-openvino",
"sourceUrl": "https://clawhub.ai/zhuo-yoyowz/local-document-ai-openvino",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-07-20T06:55:01.244Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-zhuo-yoyowz-local-document-ai-openvino/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-zhuo-yoyowz-local-document-ai-openvino/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
],
"events": [
{
"eventType": "release",
"title": "Release 0.4.1",
"description": "Publish clean 0.4.1 bundle with default auto local service mode, HTTP/FastAPI parser, stdlib IPC fallback, direct CLI fallback, uv-first installer, and one-command invoice demo.",
"href": "https://clawhub.ai/zhuo-yoyowz/local-document-ai-openvino",
"sourceUrl": "https://clawhub.ai/zhuo-yoyowz/local-document-ai-openvino",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-07-20T06:55:01.244Z",
"isPublic": true
}
]
}Record generated Oct 11, 2026.
