agentCLAWHUBUnverified

百度文档解析pipeline-parser

调用百度文档解析API解析文档。支持PDF、Word、Excel、PPT、图片等18+格式。提取文本、表格、版面分析、OCR识别及RAG文档分块。当用户需要解析文档、提取文本/表格、分析文档结构、处理扫描件时使用。触发词:文档解析、PDF解析、Word解析、表格提取、OCR、文档分析、提取文本、文档结构、扫描识别。 Skill: 百度文档解析pipeline-parser Owner: maglanyulan Summary: 调用百度文档解析API解析文档。支持PDF、Word、Excel、PPT、图片等18+格式。提取文本、表格、版面分析、OCR识别及RAG文档分块。当用户需要解析文档、提取文本/表格、分析文档结构、处理扫描件时使用。触发词:文档解析、PDF解析、Word解析、表格提取、OCR、文档分析、提取文本、文档结构、扫描识别。 Tags: latest:1.0.8 Version history: v1.0.8 | 2026-09-17T11:15:55.631Z | user - 移除 skill-card.md 文件。 - SKILL.md 文档中,调整了免费额度表:企业实名认证用户额度由 1000 页改为 200 页。 - 页面对象解析字段及部分类型补充、细化(如 page_num、text 字段描述、type/版面类型等

OpenClaw

Rank

62

Safety

84

Downloads

1.4k

Updated

Oct 10, 2026

Version

1.0.8

Source

CLAWHUB

About

What it does, and when to use it.

Capability contract not published. No trust telemetry is available yet. 1.4K downloads reported by the source. Last updated 10/10/2026.

Avoid when

  • Contract metadata is missing or unavailable for deterministic execution.

Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing

Public facts

Every fact links back to the source it came from.

Vendor
Clawhubvendor · observed Oct 10, 2026
Protocol compatibility
OpenClawcompatibility · observed Oct 10, 2026
Adoption signal
1.4K downloadsadoption · observed Oct 10, 2026
Latest release
1.0.8release · observed Sep 17, 2026
Handshake status
UNKNOWNsecurity

Install and run

Setup complexity: low.

clawhub skill install s17cm4qmnj4y8xj888f3gjms0583hg9z:baidu-doc-pipeline-parser
  1. Setup complexity is classified as HIGH. You must provision dedicated cloud infrastructure or an isolated VM. Do not run this directly on your local workstation.
  2. Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.

Contract: missing

curl -s "https://www.xpersona.co/api/v1/agents/clawhub-maglanyulan-baidu-doc-pipeline-parser/snapshot"

Documentation

CLAWHUB

145,690 characters of source documentation, loaded on request.

Extracted files

5 files captured from the source.

SKILL.md

---
name: baidu-doc-pipeline-parser 百度文档解析
description: 调用百度文档解析API解析文档。支持PDF、Word、Excel、PPT、图片等18+格式。提取文本、表格、版面分析、OCR识别及RAG文档分块。当用户需要解析文档、提取文本/表格、分析文档结构、处理扫描件时使用。触发词:文档解析、PDF解析、Word解析、表格提取、OCR、文档分析、提取文本、文档结构、扫描识别。
license: MIT
---

# 百度文档解析 Skill

基于百度智能文档分析平台 API,提供文档解析能力。

## 功能概述

- 支持对 doc、pdf、图片、xlsx 等 18 种格式文档进行解析
- 输出文档的版面、表格、阅读顺序、标题层级、旋转角度等信息
- 支持中、英、日、韩、法等 20 余种语言类型
- 可返回 Markdown 格式内容,将非结构化数据转化为易于处理的结构化数据
- 识别准确率可达 90% 以上
- 文档分块(适用于 RAG 场景)

## 适用场景

当用户需要:
- 解析 PDF、Word、Excel 等格式文档
- 从文档中提取文本内容
- 识别并提取表格数据
- 分析文档结构(标题层级、章节、版面布局)
- 对扫描件进行 OCR 文字识别
- 将文档分块用于 RAG 应用


## 免费资源领取和计费说明
[百度智能文档分析平台 领取免费测试资源](https://cloud.baidu.com/doc/OCR/s/fk3h7xu7h)


[百度智能文档分析平台计费与购买方式](https://cloud.baidu.com/doc/OCR/s/Fls06fa15#%E6%96%87%E6%A1%A3%E8%A7%A3%E6%9E%90%EF%BC%88paddleocr-vl%EF%BC%89)


| 用户类型 | 免费额度 |
|---------|---------|
| 个人实名认证用户 | **200 页** |
| 企业实名认证用户 | **200 页** |

## API 配置

### 额度获取方式

您可通过百度智能云平台获取[免费额度](https://cloud.baidu.com/doc/OCR/s/fk3h7xu7h)与[购买调用资源](https://cloud.baidu.com/doc/OCR/s/Fls06fa15#%E6%96%87%E6%A1%A3%E8%A7%A3%E6%9E%90%EF%BC%88paddleocr-vl%EF%BC%89)

### 环境变量(必须)

[百度智能文档分析平台 领取免费测试资源](https://ai.baidu.com/ai-doc/OCR/dk3iqnq51)

使用前请设置以下环境变量:

```bash
export BAIDU_DOC_AI_API_KEY="your_api_key"
export BAIDU_DOC_AI_SECRET_KEY="your_secret_key"
```

### 认证方式

通过 API Key 和 Secret Key 获取 access_token,有效期 30 天。

## 支持格式

**版式文档**:pdf, jpg, jpeg, png, bmp, tif, tiff, ofd, ppt, pptx

**流式文档**:doc, docx, txt, xls, xlsx, wps, html, mhtml

## 支持语言

CHN_ENG(中英文)、JAP(日语)、KOR(韩语)、FRE(法语)、SPA(西班牙语)、POR(葡萄牙语)、GER(德语)、ITA(意大利语)、RUS(俄语)、DAN(丹麦语)、DUT(荷兰语)、MAL(马来语)、SWE(瑞典语)、IND(印尼语)、POL(波兰语)、ROM(罗马尼亚语)、TUR(土耳其语)、GRE(希腊语)、HUN(匈牙利语)、THA(泰语)、VIE(越南语)、ARA(阿拉伯语)、HIN(印地语)

## 使用方式

```bash
python3 scripts/baidu_doc_parser.py --file_data <文件的base64编码>
python3 scripts/baidu_doc_parser.py --file_url <文件公网URL>
```

## API 接口

文档解析 API 服务为异步接口,需要先调用**提交请求接口**获取 task_id,然后调用**获取结果接口**进行结果轮询。

### 提交请求接口

- **HTTP 方法**:POST
- **请求 URL**:`https://aip.baidubce.com/rest/2.0/brain/online/v2/parser/task?access_token={token}`
- **Content-Type**:`application/x-www-form-urlencoded`

### 获取结果接口

- **HTTP 方法**:POST
- **请求 URL**:`https://aip.baidubce.com/rest/2.0/brain/online/v2/parser/task/query?access_token={token}`
- **Content-Type**:`application/x-www-form-urlencoded`
- **请求参数**:`task_id`(必填,提交请求时返回的 task_id)

## 请求参数

### 文件参数(必选,二选一)

| 参数 | 必选 | 类型 | 说明 |
|------|------|------|------|
| `file_data` | 和 file_url 二选一 | string | 文件 Base64 编码数据。版式文档:pdf, jpg, jpeg, png, bmp, tif, tiff, ofd, ppt, pptx;流式文档:doc, docx, txt, xls, xlsx, wps, html, mhtml。文档大小不超过 50M,PDF 最大支持 2000 页。**若文档大小超过 50M,须从 file_url 方式上传**。优先级:file_data > file_url |
| `file_url` | 和 file_data 二选一 | string | 文件数据 URL,长度不超过 1024 字节,支持单个 URL 传入。PDF 文档大小不超过 300MB,非 PDF 不超过 50M,PDF 最大支持 2000 页。**请注意关闭 URL 防盗链** |
| `file_name` | 是 | string | 文件名,请保证文件名后缀正确,例如 "1.pdf" |

### 核心功能参数

| 参数 | 必选 | 类型 | 可选值范围 | 说明 |
|------|------|------|----------|------|
| `recognize_formula` | 否 | bool | Tru

_meta.json

{
  "ownerId": "kn75p1w9cr0c8ycct5wzrb5ken83e9nc",
  "slug": "baidu-doc-pipeline-parser",
  "version": "1.0.8",
  "publishedAt": 1789643755631
}

references/apikey-fetch.md

# 百度文档解析 API Key 配置指南

## BAIDU_DOC_AI_API_KEY 和 BAIDU_DOC_AI_SECRET_KEY 未配置

当环境变量 `BAIDU_DOC_AI_API_KEY` 和 `BAIDU_DOC_AI_SECRET_KEY` 未设置时,按照以下步骤操作:

### 1. 获取 API Key 和 Secret Key

访问:**https://ai.baidu.com/ai-doc/REFERENCE/Ck3dwjhhu#1-获取aksk**

- 登录百度云账号
- 创建应用或查看已有的 API Key 和 Secret Key
- 复制你的 **API Key** 和 **Secret Key**

### 2. 领取免费测试资源

访问:**https://ai.baidu.com/ai-doc/OCR/dk3iqnq51**

### 3. 配置环境变量

#### 方式一:直接设置环境变量

```bash
export BAIDU_DOC_AI_API_KEY="your_actual_api_key_here"
export BAIDU_DOC_AI_SECRET_KEY="your_actual_secret_key_here"
```

#### 方式二:通过配置文件

编辑配置文件:`~/.claude/settings.json` 或项目 `.claude/settings.json`

添加以下结构:

```json
{
  "skills": {
    "entries": {
      "baidu-doc-pipeline-parser": {
        "env": {
          "BAIDU_DOC_AI_API_KEY": "your_actual_api_key_here",
          "BAIDU_DOC_AI_SECRET_KEY": "your_actual_secret_key_here"
        }
      }
    }
  }
}
```

将 `your_actual_api_key_here` 替换为你的实际 API Key,`your_actual_secret_key_here` 替换为你的实际 Secret Key。

### 4. 验证配置

```bash
# 验证 access_token 是否可正常获取
curl -X POST 'https://aip.baidubce.com/oauth/2.0/token' \
  -d 'grant_type=client_credentials' \
  -d 'client_id={your_api_key}' \
  -d 'client_secret={your_secret_key}'
```

成功返回示例:

```json
{
  "access_token": "24.xxxxx.xxxxxx.xxxxxxx-xxxxxxx",
  "expires_in": 2592000
}
```

`expires_in` 为 2592000 秒(30 天),到期后需重新获取。

### 5. 测试

```bash
python3 scripts/baidu_doc_parser.py --file_url "https://example.com/test.pdf" --file_name "test.pdf"
python3 scripts/baidu_doc_parser.py --file_data "<文件的base64编码>" --file_name "test.pdf"
```

## 常见问题

- 确保环境变量已正确设置(可通过 `echo $BAIDU_DOC_AI_API_KEY` 验证)
- 确认 API Key 有效且已开通百度智能文档分析平台服务
- 检查百度云账户余额或免费额度
- access_token 有效期 30 天,过期后会自动重新获取

## 相关链接

- [获取 AK/SK 文档](https://ai.baidu.com/ai-doc/REFERENCE/Ck3dwjhhu#1-获取aksk)
- [领取免费资源](https://ai.baidu.com/ai-doc/OCR/dk3iqnq51)
- [百度云控制台](https://console.bce.baidu.com/ai/)

references/error_codes.md

# 百度文档解析 API 错误码参考

## 通用错误

### 认证相关错误

| 错误码 | 错误信息 | 说明 | 解决方案 |
|--------|---------|------|----------|
| 1 | Unknown error | 未知错误 | 重试请求,持续出现请联系技术支持 |
| 2 | Service temporarily unavailable | 服务暂不可用 | 重试请求,持续出现请联系技术支持 |
| 3 | Unsupported openapi method | API 接口不存在 | 检查 URL 是否正确,去除非英文字符 |
| 4 | Open api request limit reached | 集群超限额 | 重试请求,持续出现请联系技术支持 |
| 6 | No permission to access data | 无 API 访问权限 | 在百度云控制台开通该 API 权限 |
| 14 | IAM Certification failed | IAM 认证失败 | 检查签名生成方式或改用 AK/SK |
| 17 | Open api daily request limit reached | 日配额超限 | 购买额度或等待次日重置 |
| 18 | Open api qps request limit reached | QPS 超限 | 降低请求频率 |
| 19 | Open api total request limit reached | 总量配额超限 | 购买额外配额 |
| 100 | Invalid parameter | access_token 无效 | 重新获取 access_token |
| 110 | Access token invalid or no longer valid | access_token 无效 | token 有效期 30 天,重新获取 |
| 111 | Access token expired | access_token 过期 | token 有效期 30 天,重新获取 |

### 文件相关错误

| 错误码 | 错误信息 | 说明 | 解决方案 |
|--------|---------|------|----------|
| 216200 | empty file or fileurl | 文件或 URL 为空 | 提供 file_data 或 file_url |
| 216201 | file format error | 文件格式不支持 | 使用支持的格式(PDF、Word、Excel 等) |
| 216202 | file size error | 文件大小超限 | 缩减文件大小(file_data ≤ 50MB,file_url PDF ≤ 300MB) |

### 任务处理错误

| 错误码 | 错误信息 | 说明 | 解决方案 |
|--------|---------|------|----------|
| 282000 | internal error | 任务处理失败 | 重试或联系技术支持 |
| 282001 | template not found | 合同类型未找到 | 检查合同类型名称 |
| 282003 | missing parameters | 缺少必要参数 | 检查必填参数 |
| 282005 | quota exceed error | 额度不足 | 申请增加配额 |
| 282006 | check user auth error | 用户权限校验失败 | 验证用户权限 |
| 282007 | task not exist, please check task id | 任务不存在 | 检查 task_id 是否正确 |
| 282018 | Service busy | 服务繁忙 | 降低请求频率 |

### URL 相关错误

| 错误码 | 错误信息 | 说明 | 解决方案 |
|--------|---------|------|----------|
| 282111 | url format illegal | URL 格式不合法 | 检查 URL 格式 |
| 282112 | url download timeout | URL 下载超时 | 检查 URL 是否可访问 |
| 282113 | url response invalid | URL 响应无效 | 检查 URL 返回内容是否正确 |
| 282114 | url size error | URL 长度超过 1024 字节 | 缩短 URL |

### 参数错误

| 错误码 | 错误信息 | 说明 | 解决方案 |
|--------|---------|------|----------|
| 283016 | parameters value error | 参数值无效 | 检查参数格式和取值 |

## 错误响应格式

```json
{
  "log_id": "13665091038742503867108513247688",
  "error_code": "282007",
  "error_msg": "task not exist, please check task id",
  "result": "null"
}
```

## 错误处理策略

### 重试策略

| 错误类型 | 错误码 | 建议处理方式 |
|---------|--------|-------------|
| 瞬时错误 | 1, 2, 4, 282000, 282018 | 指数退避重试 |
| 认证错误 | 100, 110, 111 | 重新获取 access_token |
| 配额错误 | 17, 18, 19, 282005 | 等待或购买额外配额 |
| 参数错误 | 216200, 216201, 282003, 283016 | 修正参数后重试 |
| URL 错误 | 282111, 282112, 282113, 282114 | 检查并修正 URL |

### 指数退避重试示例

```python
import time

def retry_with_backoff(func, max_retries=3):
    for i in range(max_retries):
        try:
            return func()
        except Exception as e:
            if i == max_retries - 1:
                raise
            wait_time = 2 ** i  # 1s, 2s, 4s
            time.sleep(wait_time)
```

### Token 刷新示例

```python
def ensure_valid_token(cli

references/parameters.md

# 百度文档解析 API 参数详解

## 接口概述

百度文档解析 API 支持对 doc、pdf、图片、xlsx 等 18 种格式文档进行解析,输出文档的版面、表格、阅读顺序、标题层级、旋转角度等信息,支持中、英、日、韩、法等 20 余种语言类型,识别准确率可达 90% 以上。

## API 接口地址

### 提交请求接口

```
POST https://aip.baidubce.com/rest/2.0/brain/online/v2/parser/task?access_token={access_token}
Content-Type: application/x-www-form-urlencoded
```

### 获取结果接口

```
POST https://aip.baidubce.com/rest/2.0/brain/online/v2/parser/task/query?access_token={access_token}
Content-Type: application/x-www-form-urlencoded
```

## 提交请求参数

### 文件参数(必选,二选一)

| 参数 | 必选 | 类型 | 说明 |
|------|------|------|------|
| file_data | 和 file_url 二选一 | string | 文件的 base64 编码数据。版式文档:pdf, jpg, jpeg, png, bmp, tif, tiff, ofd, ppt, pptx;流式文档:doc, docx, txt, xls, xlsx, wps, html, mhtml。文档大小不超过 50M,其中 PDF 文档最大支持 2000 页。若文档大小超过 50M,须从 file_url 方式上传。优先级:file_data > file_url,当 file_data 字段存在时,file_url 字段失效 |
| file_url | 和 file_data 二选一 | string | 文件数据 URL,URL 长度不超过 1024 字节,支持单个 URL 传入。PDF 文档大小不超过 300MB,非 PDF 文档大小不超过 50M,其中 PDF 文档最大支持 2000 页。优先级:file_data > file_url。**请注意关闭 URL 防盗链** |
| file_name | 是 | string | 文件名,请保证文件名后缀正确,例如 "1.pdf" |

### 核心功能参数

| 参数 | 必选 | 类型 | 可选值范围 | 说明 |
|------|------|------|----------|------|
| recognize_formula | 否 | bool | True/False | 是否对版式类型文档进行公式识别 |
| analysis_chart | 否 | bool | True/False | 是否对统计图表进行解析 |
| angle_adjust | 否 | bool | True/False | 是否对图片进行角度矫正 |
| parse_image_layout | 否 | bool | True/False | 是否返回文档中的图片位置信息 |

### 语言与格式参数

| 参数 | 必选 | 类型 | 默认值 | 说明 |
|------|------|------|--------|------|
| language_type | 否 | string | CHN_ENG | 识别语种类型 |
| switch_digital_width | 否 | string | auto | 是否对数字进行全半角转换。auto:不转换;half:半角输出;full:全角输出 |
| html_table_format | 否 | bool | True | 是否将识别出的表格转换为 HTML 格式返回 |

### 支持语种列表

| 代码 | 语言 | 代码 | 语言 |
|------|------|------|------|
| CHN_ENG | 中英文 | DAN | 丹麦语 |
| JAP | 日语 | DUT | 荷兰语 |
| KOR | 韩语 | MAL | 马来语 |
| FRE | 法语 | SWE | 瑞典语 |
| SPA | 西班牙语 | IND | 印尼语 |
| POR | 葡萄牙语 | POL | 波兰语 |
| GER | 德语 | ROM | 罗马尼亚语 |
| ITA | 意大利语 | TUR | 土耳其语 |
| RUS | 俄语 | GRE | 希腊语 |
| HUN | 匈牙利语 | THA | 泰语 |
| VIE | 越南语 | ARA | 阿拉伯语 |
| HIN | 印地语 | - | - |

### 文档分块参数

| 参数 | 必选 | 类型 | 默认值 | 说明 |
|------|------|------|--------|------|
| return_doc_chunks | 否 | dict | - | 是否返回文档切分后的片段数据(按语义、字数、标点) |
| + switch | 否 | bool | False | 是否进行文档内容切分 |
| + split_type | 否 | str | chunk | 切分方式。chunk:按照 chunk_size 来切;mark:按照 separators 来切 |
| + separators | 否 | list | ['。',';','!','?',';','!','?'] | 切分标点 |
| + chunk_size | 否 | int | -1 | 切分块的大小,-1 表示按照语义自动切分,不限定块的大小 |

## 获取结果请求参数

| 参数 | 必选 | 类型 | 说明 |
|------|------|------|------|
| task_id | 是 | string | 发送提交请求时返回的 task_id |

## 返回结构

### 提交请求返回

| 字段 | 类型 | 说明 |
|------|------|------|
| log_id | uint64 | 唯一的 log id,用于问题定位 |
| error_code | int | 错误码 |
| error_msg | string | 错误描述信息 |
| result | dict | 返回的结果列表 |
| + task_id | string | 该请求生成的 task_id,后续使用该 task_id 获取结果 |

成功返回示例:

```json
{
  "error_code": 0,
  "error_msg": "",
  "log_id": "10138598131137362685273585665433",
  "result": {
    "task_id": "task-3zy9Bg8CHt1M4p
Github ReposUpdated 19h agoRank 70

AionUi

Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!

MCPOPENCLAW
Github ReposUpdated 6mo agoRank 70

activepieces

AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents

OPENCLAW
Github ReposUpdated 6mo agoRank 70

cherry-studio

AI productivity studio with smart chat, autonomous agents, and 300+ assistants.

MCPOPENCLAW
Github ReposUpdated 7mo agoRank 70

CopilotKit

The Frontend for Agents & Generative UI. React + Angular

OPENCLAW

Machine-readable data

The same record, as JSON, for agents and crawlers.

{
  "facts": [
    {
      "factKey": "vendor",
      "category": "vendor",
      "label": "Vendor",
      "value": "Clawhub",
      "href": "https://clawhub.ai/maglanyulan/skills/baidu-doc-pipeline-parser",
      "sourceUrl": "https://clawhub.ai/maglanyulan/skills/baidu-doc-pipeline-parser",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-10T12:10:35.808Z",
      "isPublic": true
    },
    {
      "factKey": "protocols",
      "category": "compatibility",
      "label": "Protocol compatibility",
      "value": "OpenClaw",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-maglanyulan-baidu-doc-pipeline-parser/contract",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-maglanyulan-baidu-doc-pipeline-parser/contract",
      "sourceType": "contract",
      "confidence": "medium",
      "observedAt": "2026-10-10T12:10:35.808Z",
      "isPublic": true
    },
    {
      "factKey": "traction",
      "category": "adoption",
      "label": "Adoption signal",
      "value": "1.4K downloads",
      "href": "https://clawhub.ai/maglanyulan/baidu-doc-pipeline-parser",
      "sourceUrl": "https://clawhub.ai/maglanyulan/baidu-doc-pipeline-parser",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-10T12:10:35.808Z",
      "isPublic": true
    },
    {
      "factKey": "latest_release",
      "category": "release",
      "label": "Latest release",
      "value": "1.0.8",
      "href": "https://clawhub.ai/maglanyulan/baidu-doc-pipeline-parser",
      "sourceUrl": "https://clawhub.ai/maglanyulan/baidu-doc-pipeline-parser",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-09-17T11:15:55.631Z",
      "isPublic": true
    },
    {
      "factKey": "handshake_status",
      "category": "security",
      "label": "Handshake status",
      "value": "UNKNOWN",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-maglanyulan-baidu-doc-pipeline-parser/trust",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-maglanyulan-baidu-doc-pipeline-parser/trust",
      "sourceType": "trust",
      "confidence": "medium",
      "observedAt": null,
      "isPublic": true
    }
  ],
  "events": [
    {
      "eventType": "release",
      "title": "Release 1.0.8",
      "description": "- 移除 skill-card.md 文件。 - SKILL.md 文档中,调整了免费额度表:企业实名认证用户额度由 1000 页改为 200 页。 - 页面对象解析字段及部分类型补充、细化(如 page_num、text 字段描述、type/版面类型等)。 - 其他内容未变。",
      "href": "https://clawhub.ai/maglanyulan/baidu-doc-pipeline-parser",
      "sourceUrl": "https://clawhub.ai/maglanyulan/baidu-doc-pipeline-parser",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-09-17T11:15:55.631Z",
      "isPublic": true
    }
  ]
}

Record generated Oct 10, 2026.

Sponsored

Ads related to 百度文档解析pipeline-parser and adjacent AI workflows.