{"id":"eb1d85c8-9329-4f85-8835-a08840facf74","entityType":"agent","slug":"clawhub-maglanyulan-baidu-doc-pipeline-parser","name":"百度文档解析pipeline-parser","canonicalUrl":"https://www.xpersona.co/agent/clawhub-maglanyulan-baidu-doc-pipeline-parser","canonicalPath":"/agent/clawhub-maglanyulan-baidu-doc-pipeline-parser","generatedAt":"2026-10-10T14:46:32.358Z","source":"CLAWHUB","claimStatus":"UNCLAIMED","verificationTier":"NONE","summary":{"evidence":{"source":"editorial-content","verified":true,"confidence":"high","updatedAt":"2026-10-10T12:10:35.808Z","emptyReason":null},"description":"调用百度文档解析API解析文档。支持PDF、Word、Excel、PPT、图片等18+格式。提取文本、表格、版面分析、OCR识别及RAG文档分块。当用户需要解析文档、提取文本/表格、分析文档结构、处理扫描件时使用。触发词：文档解析、PDF解析、Word解析、表格提取、OCR、文档分析、提取文本、文档结构、扫描识别。 Skill: 百度文档解析pipeline-parser Owner: maglanyulan Summary: 调用百度文档解析API解析文档。支持PDF、Word、Excel、PPT、图片等18+格式。提取文本、表格、版面分析、OCR识别及RAG文档分块。当用户需要解析文档、提取文本/表格、分析文档结构、处理扫描件时使用。触发词：文档解析、PDF解析、Word解析、表格提取、OCR、文档分析、提取文本、文档结构、扫描识别。 Tags: latest:1.0.8 Version history: v1.0.8 | 2026-09-17T11:15:55.631Z | user - 移除 skill-card.md 文件。 - SKILL.md 文档中，调整了免费额度表：企业实名认证用户额度由 1000 页改为 200 页。 - 页面对象解析字段及部分类型补充、细化（如 page_num、text 字段描述、type/版面类型等","descriptionLabel":"Technical summary","evidenceSummary":"Capability contract not published. No trust telemetry is available yet. 1.4K downloads reported by the source. Last updated 10/10/2026.","installCommand":"clawhub skill install s17cm4qmnj4y8xj888f3gjms0583hg9z:baidu-doc-pipeline-parser","sourceUrl":"https://clawhub.ai/maglanyulan/baidu-doc-pipeline-parser","homepage":"https://clawhub.ai/maglanyulan/skills/baidu-doc-pipeline-parser","primaryLinks":[{"label":"View on ClawHub","url":"https://clawhub.ai/maglanyulan/baidu-doc-pipeline-parser","kind":"source"},{"label":"Homepage","url":"https://clawhub.ai/maglanyulan/skills/baidu-doc-pipeline-parser","kind":"homepage"}],"safetyScore":84,"overallRank":62,"popularityScore":63,"trustScore":null,"claimedByName":null,"isOwner":false,"seoDescription":"调用百度文档解析API解析文档。支持PDF、Word、Excel、PPT、图片等18+格式。提取文本、表格、版面分析、OCR识别及RAG文档分块。当用户需要解析文档、提取文本/表格、分析文档结构、处理扫描件时使用。触发词：文档解析、PDF解析、Word解析、表格提取、OCR、文档分析、提取文本、文档结构、扫描识别。 S"},"coverage":{"evidence":{"source":"public-profile","verified":false,"confidence":"medium","updatedAt":"2026-10-10T12:10:35.808Z","emptyReason":null},"protocols":[{"protocol":"OPENCLEW","label":"OpenClaw","status":"self-declared","notes":"Declared in the public agent profile."}],"capabilities":[],"verifiedCount":0,"selfDeclaredCount":1,"capabilityMatrix":{"rows":[{"key":"OPENCLEW","type":"protocol","support":"unknown","confidenceSource":"profile","notes":"Listed on profile"}],"flattenedTokens":"protocol:OPENCLEW|unknown|profile"}},"adoption":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T12:10:35.808Z","emptyReason":null},"stars":null,"forks":null,"downloads":1441,"packageName":null,"latestVersion":"1.0.8","tractionLabel":"1.4K downloads"},"release":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T12:10:35.808Z","emptyReason":null},"lastUpdatedAt":"2026-10-10T12:10:35.808Z","lastCrawledAt":"2026-10-10T12:10:35.808Z","lastIndexedAt":null,"nextCrawlAt":"2026-10-11T12:10:35.808Z","lastVerifiedAt":null,"highlights":[{"version":"1.0.8","createdAt":"2026-09-17T11:15:55.631Z","changelog":"- 移除 skill-card.md 文件。 - SKILL.md 文档中，调整了免费额度表：企业实名认证用户额度由 1000 页改为 200 页。 - 页面对象解析字段及部分类型补充、细化（如 page_num、text 字段描述、type/版面类型等）。 - 其他内容未变。","fileCount":7,"zipByteSize":18321},{"version":"1.0.7","createdAt":"2026-07-08T08:40:41.710Z","changelog":"baidu-doc-pipeline-parser 1.0.7 - 切换为标准文档解析（pipeline-parser）Skill版本，替换原多模态VLM实现 - 新增 scripts/baidu_doc_parser.py，移除 scripts/baidu_doc_vlm_parser.py - 支持18+文档格式及RAG场景分块，适用PDF、Word、Excel、图片等结构化提取和OCR - API文档与调用方式更新，详细参数、返回结构说明优化 - 适用范围扩展至通用文档解析、OCR、结构化与分块场景","fileCount":7,"zipByteSize":18266},{"version":"1.0.6","createdAt":"2026-07-08T08:34:06.829Z","changelog":"**Version 1.0.6 – Major Upgrade: Migrated skill to use Baidu PaddleOCR-VL multimodal large model for document parsing.** - Switched from the traditional pipeline parser to PaddleOCR-VL-1.6 large model API for improved recognition of complex document elements. - Added support for over 100 languages with automatic language detection; enhanced support for hand-written text, formulas, tables, seals, and irregular layouts. - All parameters and API endpoints updated for PaddleOCR-VL interface; no need to manually specify language or enable formula recognition. - Supports precise line-level (span) coordinates, polygon bounding, and improved multi-page/table merging. - Skill name, trigger words, example code, and documentation fully updated to reflect VLM capability. - Old pipeline scripts and documentation removed; new parser script for PaddleOCR-VL introduced.","fileCount":7,"zipByteSize":14624},{"version":"1.0.5","createdAt":"2026-05-09T09:01:20.416Z","changelog":"baidu-doc-pipeline-parser v1.0.2 - Initial release of the skill under the new name \"baidu-doc-pipeline-parser\" with updated metadata. - Added _meta.json file. - SKILL.md revised and localized: clearer documentation, parameter and return value tables, and example usage in Chinese. - Enhanced details for file, API, and result structure; included new formatting and examples for improved usability. - Maintains support for parsing 18+ document formats with OCR, structure extraction, and RAG document chunking as described previously.","fileCount":7,"zipByteSize":18132},{"version":"1.0.4","createdAt":"2026-05-09T08:49:40.814Z","changelog":"- No changes detected in this version. - Functionality and documentation remain the same as the previous release.","fileCount":6,"zipByteSize":16619},{"version":"1.0.3","createdAt":"2026-05-09T06:18:37.912Z","changelog":"- No file changes detected in this version. - Documentation and API details remain unchanged from previous release. - No new features, bug fixes, or updates are included in version 1.0.3.","fileCount":6,"zipByteSize":16605},{"version":"1.0.2","createdAt":"2026-05-09T06:14:44.085Z","changelog":"- 添加详细 SKILL.md 文档，全面描述支持的文档格式、核心功能、API 用法及参数说明 - 说明支持文本、表格、版面分析、OCR、RAG 分块等多种能力，涵盖 18+ 文件格式与 20 余种语言 - 增加样例用法、错误码及 API 使用限制说明 - 全面提升用户参考性，便于理解和集成百度文档解析 API","fileCount":6,"zipByteSize":16606}]},"execution":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No published capability contract is available yet."},"installCommand":"clawhub skill install s17cm4qmnj4y8xj888f3gjms0583hg9z:baidu-doc-pipeline-parser","setupComplexity":"low","setupSteps":["Setup complexity is classified as HIGH. You must provision dedicated cloud infrastructure or an isolated VM. Do not run this directly on your local workstation.","Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data."],"contract":{"contractStatus":"missing","authModes":[],"requires":[],"forbidden":[],"supportsMcp":false,"supportsA2a":false,"supportsStreaming":false,"inputSchemaRef":null,"outputSchemaRef":null,"dataRegion":null,"contractUpdatedAt":null,"sourceUpdatedAt":null,"freshnessSeconds":null},"invocationGuide":{"preferredApi":{"snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-maglanyulan-baidu-doc-pipeline-parser/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-maglanyulan-baidu-doc-pipeline-parser/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-maglanyulan-baidu-doc-pipeline-parser/trust"},"curlExamples":["curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-maglanyulan-baidu-doc-pipeline-parser/snapshot\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-maglanyulan-baidu-doc-pipeline-parser/contract\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-maglanyulan-baidu-doc-pipeline-parser/trust\""],"jsonRequestTemplate":{"query":"summarize this repo","constraints":{"maxLatencyMs":2000,"protocolPreference":["OPENCLEW"]}},"jsonResponseTemplate":{"ok":true,"result":{"summary":"...","confidence":0.9},"meta":{"source":"CLAWHUB","generatedAt":"2026-10-10T14:46:32.353Z"}},"retryPolicy":{"maxAttempts":3,"backoffMs":[500,1500,3500],"retryableConditions":["HTTP_429","HTTP_503","NETWORK_TIMEOUT"]}},"endpoints":{"dossierUrl":"https://www.xpersona.co/api/v1/agents/clawhub-maglanyulan-baidu-doc-pipeline-parser/dossier","snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-maglanyulan-baidu-doc-pipeline-parser/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-maglanyulan-baidu-doc-pipeline-parser/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-maglanyulan-baidu-doc-pipeline-parser/trust"}},"reliability":{"evidence":{"source":"runtime-metrics","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No trust, reliability, or runtime telemetry is available."},"trust":{"status":"unavailable","handshakeStatus":"UNKNOWN","verificationFreshnessHours":null,"reputationScore":null,"p95LatencyMs":null,"successRate30d":null,"fallbackRate":null,"attempts30d":null,"trustUpdatedAt":null,"trustConfidence":"unknown","sourceUpdatedAt":null,"freshnessSeconds":null},"decisionGuardrails":{"doNotUseIf":["Contract metadata is missing or unavailable for deterministic execution."],"safeUseWhen":[],"riskFlags":["missing_or_unavailable_contract","trust_data_unavailable","schema_references_missing"],"operationalConfidence":"low"},"executionMetrics":{"observedLatencyMsP50":null,"observedLatencyMsP95":null,"estimatedCostUsd":null,"uptime30d":null,"rateLimitRpm":null,"rateLimitBurst":null,"lastVerifiedAt":null,"verificationSource":null},"runtimeMetrics":{"successRate":null,"avgLatencyMs":null,"avgCostUsd":null,"hallucinationRate":null,"retryRate":null,"disputeRate":null,"p50Latency":null,"p95Latency":null,"lastUpdated":null}},"benchmarks":{"evidence":{"source":"no-benchmark-data","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No benchmark suites or observed failure patterns are available."},"suites":[],"failurePatterns":[]},"artifacts":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"high","updatedAt":"2026-10-10T12:10:35.808Z","emptyReason":null},"readme":"Skill: 百度文档解析pipeline-parser\n\nOwner: maglanyulan\n\nSummary: 调用百度文档解析API解析文档。支持PDF、Word、Excel、PPT、图片等18+格式。提取文本、表格、版面分析、OCR识别及RAG文档分块。当用户需要解析文档、提取文本/表格、分析文档结构、处理扫描件时使用。触发词：文档解析、PDF解析、Word解析、表格提取、OCR、文档分析、提取文本、文档结构、扫描识别。\n\nTags: latest:1.0.8\n\nVersion history:\n\nv1.0.8 | 2026-09-17T11:15:55.631Z | user\n\n- 移除 skill-card.md 文件。\n- SKILL.md 文档中，调整了免费额度表：企业实名认证用户额度由 1000 页改为 200 页。\n- 页面对象解析字段及部分类型补充、细化（如 page_num、text 字段描述、type/版面类型等）。\n- 其他内容未变。\n\nv1.0.7 | 2026-07-08T08:40:41.710Z | user\n\nbaidu-doc-pipeline-parser 1.0.7\n\n- 切换为标准文档解析（pipeline-parser）Skill版本，替换原多模态VLM实现\n- 新增 scripts/baidu_doc_parser.py，移除 scripts/baidu_doc_vlm_parser.py\n- 支持18+文档格式及RAG场景分块，适用PDF、Word、Excel、图片等结构化提取和OCR\n- API文档与调用方式更新，详细参数、返回结构说明优化\n- 适用范围扩展至通用文档解析、OCR、结构化与分块场景\n\nv1.0.6 | 2026-07-08T08:34:06.829Z | user\n\n**Version 1.0.6 – Major Upgrade: Migrated skill to use Baidu PaddleOCR-VL multimodal large model for document parsing.**\n\n- Switched from the traditional pipeline parser to PaddleOCR-VL-1.6 large model API for improved recognition of complex document elements.\n- Added support for over 100 languages with automatic language detection; enhanced support for hand-written text, formulas, tables, seals, and irregular layouts.\n- All parameters and API endpoints updated for PaddleOCR-VL interface; no need to manually specify language or enable formula recognition.\n- Supports precise line-level (span) coordinates, polygon bounding, and improved multi-page/table merging.\n- Skill name, trigger words, example code, and documentation fully updated to reflect VLM capability.\n- Old pipeline scripts and documentation removed; new parser script for PaddleOCR-VL introduced.\n\nv1.0.5 | 2026-05-09T09:01:20.416Z | user\n\nbaidu-doc-pipeline-parser v1.0.2\n\n- Initial release of the skill under the new name \"baidu-doc-pipeline-parser\" with updated metadata.\n- Added _meta.json file.\n- SKILL.md revised and localized: clearer documentation, parameter and return value tables, and example usage in Chinese.\n- Enhanced details for file, API, and result structure; included new formatting and examples for improved usability.\n- Maintains support for parsing 18+ document formats with OCR, structure extraction, and RAG document chunking as described previously.\n\nv1.0.4 | 2026-05-09T08:49:40.814Z | user\n\n- No changes detected in this version.\n- Functionality and documentation remain the same as the previous release.\n\nv1.0.3 | 2026-05-09T06:18:37.912Z | user\n\n- No file changes detected in this version.\n- Documentation and API details remain unchanged from previous release.\n- No new features, bug fixes, or updates are included in version 1.0.3.\n\nv1.0.2 | 2026-05-09T06:14:44.085Z | user\n\n- 添加详细 SKILL.md 文档，全面描述支持的文档格式、核心功能、API 用法及参数说明\n- 说明支持文本、表格、版面分析、OCR、RAG 分块等多种能力，涵盖 18+ 文件格式与 20 余种语言\n- 增加样例用法、错误码及 API 使用限制说明\n- 全面提升用户参考性，便于理解和集成百度文档解析 API\n\nArchive index:\n\nArchive v1.0.8: 7 files, 18321 bytes\n\nFiles: _meta.json (144b), references/apikey-fetch.md (2455b), references/error_codes.md (4611b), references/parameters.md (10839b), scripts/baidu_doc_parser.py (11067b), skill-card.md (2705b), SKILL.md (13869b)\n\nFile v1.0.8:SKILL.md\n\n---\nname: baidu-doc-pipeline-parser 百度文档解析\ndescription: 调用百度文档解析API解析文档。支持PDF、Word、Excel、PPT、图片等18+格式。提取文本、表格、版面分析、OCR识别及RAG文档分块。当用户需要解析文档、提取文本/表格、分析文档结构、处理扫描件时使用。触发词：文档解析、PDF解析、Word解析、表格提取、OCR、文档分析、提取文本、文档结构、扫描识别。\nlicense: MIT\n---\n\n# 百度文档解析 Skill\n\n基于百度智能文档分析平台 API，提供文档解析能力。\n\n## 功能概述\n\n- 支持对 doc、pdf、图片、xlsx 等 18 种格式文档进行解析\n- 输出文档的版面、表格、阅读顺序、标题层级、旋转角度等信息\n- 支持中、英、日、韩、法等 20 余种语言类型\n- 可返回 Markdown 格式内容，将非结构化数据转化为易于处理的结构化数据\n- 识别准确率可达 90% 以上\n- 文档分块（适用于 RAG 场景）\n\n## 适用场景\n\n当用户需要：\n- 解析 PDF、Word、Excel 等格式文档\n- 从文档中提取文本内容\n- 识别并提取表格数据\n- 分析文档结构（标题层级、章节、版面布局）\n- 对扫描件进行 OCR 文字识别\n- 将文档分块用于 RAG 应用\n\n\n## 免费资源领取和计费说明\n[百度智能文档分析平台 领取免费测试资源](https://cloud.baidu.com/doc/OCR/s/fk3h7xu7h)\n\n\n[百度智能文档分析平台计费与购买方式](https://cloud.baidu.com/doc/OCR/s/Fls06fa15#%E6%96%87%E6%A1%A3%E8%A7%A3%E6%9E%90%EF%BC%88paddleocr-vl%EF%BC%89)\n\n\n| 用户类型 | 免费额度 |\n|---------|---------|\n| 个人实名认证用户 | **200 页** |\n| 企业实名认证用户 | **200 页** |\n\n## API 配置\n\n### 额度获取方式\n\n您可通过百度智能云平台获取[免费额度](https://cloud.baidu.com/doc/OCR/s/fk3h7xu7h)与[购买调用资源](https://cloud.baidu.com/doc/OCR/s/Fls06fa15#%E6%96%87%E6%A1%A3%E8%A7%A3%E6%9E%90%EF%BC%88paddleocr-vl%EF%BC%89)\n\n### 环境变量（必须）\n\n[百度智能文档分析平台 领取免费测试资源](https://ai.baidu.com/ai-doc/OCR/dk3iqnq51)\n\n使用前请设置以下环境变量：\n\n```bash\nexport BAIDU_DOC_AI_API_KEY=\"your_api_key\"\nexport BAIDU_DOC_AI_SECRET_KEY=\"your_secret_key\"\n```\n\n### 认证方式\n\n通过 API Key 和 Secret Key 获取 access_token，有效期 30 天。\n\n## 支持格式\n\n**版式文档**：pdf, jpg, jpeg, png, bmp, tif, tiff, ofd, ppt, pptx\n\n**流式文档**：doc, docx, txt, xls, xlsx, wps, html, mhtml\n\n## 支持语言\n\nCHN_ENG（中英文）、JAP（日语）、KOR（韩语）、FRE（法语）、SPA（西班牙语）、POR（葡萄牙语）、GER（德语）、ITA（意大利语）、RUS（俄语）、DAN（丹麦语）、DUT（荷兰语）、MAL（马来语）、SWE（瑞典语）、IND（印尼语）、POL（波兰语）、ROM（罗马尼亚语）、TUR（土耳其语）、GRE（希腊语）、HUN（匈牙利语）、THA（泰语）、VIE（越南语）、ARA（阿拉伯语）、HIN（印地语）\n\n## 使用方式\n\n```bash\npython3 scripts/baidu_doc_parser.py --file_data <文件的base64编码>\npython3 scripts/baidu_doc_parser.py --file_url <文件公网URL>\n```\n\n## API 接口\n\n文档解析 API 服务为异步接口，需要先调用**提交请求接口**获取 task_id，然后调用**获取结果接口**进行结果轮询。\n\n### 提交请求接口\n\n- **HTTP 方法**：POST\n- **请求 URL**：`https://aip.baidubce.com/rest/2.0/brain/online/v2/parser/task?access_token={token}`\n- **Content-Type**：`application/x-www-form-urlencoded`\n\n### 获取结果接口\n\n- **HTTP 方法**：POST\n- **请求 URL**：`https://aip.baidubce.com/rest/2.0/brain/online/v2/parser/task/query?access_token={token}`\n- **Content-Type**：`application/x-www-form-urlencoded`\n- **请求参数**：`task_id`（必填，提交请求时返回的 task_id）\n\n## 请求参数\n\n### 文件参数（必选，二选一）\n\n| 参数 | 必选 | 类型 | 说明 |\n|------|------|------|------|\n| `file_data` | 和 file_url 二选一 | string | 文件 Base64 编码数据。版式文档：pdf, jpg, jpeg, png, bmp, tif, tiff, ofd, ppt, pptx；流式文档：doc, docx, txt, xls, xlsx, wps, html, mhtml。文档大小不超过 50M，PDF 最大支持 2000 页。**若文档大小超过 50M，须从 file_url 方式上传**。优先级：file_data > file_url |\n| `file_url` | 和 file_data 二选一 | string | 文件数据 URL，长度不超过 1024 字节，支持单个 URL 传入。PDF 文档大小不超过 300MB，非 PDF 不超过 50M，PDF 最大支持 2000 页。**请注意关闭 URL 防盗链** |\n| `file_name` | 是 | string | 文件名，请保证文件名后缀正确，例如 \"1.pdf\" |\n\n### 核心功能参数\n\n| 参数 | 必选 | 类型 | 可选值范围 | 说明 |\n|------|------|------|----------|------|\n| `recognize_formula` | 否 | bool | True/False | 是否对版式类型文档进行公式识别 |\n| `analysis_chart` | 否 | bool | True/False | 是否对统计图表进行解析 |\n| `angle_adjust` | 否 | bool | True/False | 是否对图片进行角度矫正 |\n| `parse_image_layout` | 否 | bool | True/False | 是否返回文档中的图片位置信息 |\n\n### 语言与格式参数\n\n| 参数 | 必选 | 类型 | 说明 |\n|------|------|------|------|\n| `language_type` | 否 | string | 识别语种类型，默认为 CHN_ENG（中英文） |\n| `switch_digital_width` | 否 | string | 是否对数字进行全半角转换，默认为 auto。可选：auto（不转换）、half（半角输出）、full（全角输出） |\n| `html_table_format` | 否 | bool | 是否将识别出的表格转换为 HTML 格式返回，**default=True** |\n\n### 文档分块参数\n\n`return_doc_chunks` 为字典类型，用于返回文档切分后的片段数据（按语义、字数、标点）：\n\n| 参数 | 必选 | 类型 | 默认值 | 说明 |\n|------|------|------|--------|------|\n| `switch` | 否 | bool | False | 是否进行文档内容切分 |\n| `split_type` | 否 | str | chunk | 切分方式：chunk（按 chunk_size 来切）/ mark（按 separators 来切） |\n| `separators` | 否 | list | ['。','；','！','？',';','!','?'] | 切分标点 |\n| `chunk_size` | 否 | int | -1 | 切分块的大小，-1 表示按照语义自动切分，不限定块的大小 |\n\n## 返回结构\n\n### 提交请求返回\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `log_id` | uint64 | 唯一的 log id，用于问题定位 |\n| `error_code` | int | 错误码 |\n| `error_msg` | string | 错误描述信息 |\n| `result.task_id` | string | 该请求生成的 task_id，后续使用该 task_id 获取审查结果 |\n\n### 获取结果返回\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `log_id` | uint64 | 唯一的 log id |\n| `error_code` | int | 错误码 |\n| `error_msg` | string | 错误描述信息 |\n| `result.task_id` | string | 任务 ID |\n| `result.status` | string | 任务状态：pending（排队中）、processing（运行中）、success（成功）、failed（失败） |\n| `result.task_error` | string | 解析报错信息，包含任务失败、额度不够 |\n| `result.markdown_url` | string | 文档解析结果的 Markdown 格式链接，**链接有效期 30 天** |\n| `result.parse_result_url` | string | 文档解析结果的 BOS 链接（JSON），**链接有效期 30 天** |\n\n### 解析结果 JSON 结构（parse_result_url）\n\n#### 顶层结构\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `file_name` | string | 文档名称 |\n| `file_id` | string | 文档 ID |\n| `pages` | list | 文件单页解析内容 |\n| `chunks` | list | 文件内容切分结果（return_doc_chunks.switch=True 时有值） |\n\n#### 页面对象（pages[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `page_id` | string | 页码 ID |\n| `page_num` | int | 页码数，从 0 开始 |\n| `text` | string | 当前页的所有纯文字内容，含表格 Markdown |\n| `layouts` | list | 页面内容版式分析的结果 |\n| `tables` | list | 页面表格解析结果 |\n| `images` | list | 页面中图片解析结果 |\n| `meta` | dict | 页元信息 |\n\n#### 页面元信息（meta）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `page_width` | int | 页面宽度 |\n| `page_height` | int | 页面高度 |\n| `is_scan` | bool | 是否扫描件 |\n| `page_angle` | int | 页面倾斜角度 |\n| `page_type` | string | 页面属性：text（正文）、contents（目录）、appendix（附录）、others（其他） |\n| `sheet_name` | string | Excel 的 sheet 名 |\n\n#### 版面元素（layouts[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `layout_id` | string | layout 元素唯一标志，格式 \"xxxxx-layout-{global_layout_index}\" |\n| `text` | string | layout 对应的文本内容。注：当 type 为 table/image 时该字段为空，需根据 type 和 layout_id 分别到 tables/images 字段里找到对应内容 |\n| `position` | list | 元素在页面中的位置 [x, y, w, h]，左上角和宽高 |\n| `type` | string | 版面元素类型（见下表） |\n| `sub_type` | string | 版面元素子类型（见下表） |\n| `parent` | string | 标题层级树中父节点的 layout_id，若为一级标题则 parent 为 \"root\" |\n| `children` | list | 标题层级树中子节点的 layout_id 列表 |\n\n**版面类型（type）**：\n\n| 类型 | 说明 |\n|------|------|\n| `text` | 段落 |\n| `table` | 表格 |\n| `image` | 文档中的插图 |\n| `head_tail` | 页面顶部（页眉/页脚） |\n| `contents` | 目录 |\n| `seal` | 印章 |\n| `title` | 标题 |\n| `formula` | 公式 |\n| `hand_sign` | 手写签名 |\n\n**子类型（sub_type）**：\n\n- **title 类**：`title_{n}`（n 级标题，如 title_2 代表二级标题）、`image_title`（图标题）、`table_title`（表标题）\n- **image 类**：`chart`（统计图表）、`figure`（普通插图）、`QR_code`（二维码）、`Bar_code`（条形码）\n\n#### 表格对象（tables[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `layout_id` | string | 与 layouts 中 type 为 table 的元素的 layout ID 对应 |\n| `markdown` | string | 表格内容的 Markdown 形式 |\n| `table_title_id` | list | 表格标题对应的 layout_id，默认为 null |\n| `position` | list | 边框数据 [x, y, w, h]（以页面坐标为原点），版式格式时有效 |\n| `cells` | list | 单元格的内版面信息，layout 类型为表格时有值 |\n| `matrix` | list | 二位数组，表示表格内布局位置信息，每个元素对应 cells 列表中元素的索引 |\n| `merge_table` | string | 跨页表格标记：begin（开始）、inner（中间，超过两页）、end（结束）；非跨页表格该字段为空 |\n\n#### 图片对象（images[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `layout_id` | string | 与 layouts 中 type 为 image 的元素的 layout ID 对应 |\n| `image_title_id` | list | 图片标题对应的 layout_id，默认为 null |\n| `position` | list | 边框数据 [x, y, w, h] |\n| `content_layouts` | list | 图片的内版面信息 |\n| `data_url` | string | 图片存储链接 |\n| `image_description` | string | 对统计图表进行内容解析和描述，输出结果为 JSON 字符串 |\n\n#### 分块对象（chunks[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `chunk_id` | string | 切片的 ID |\n| `content` | string | 切片的内容 |\n| `type` | string | 切片类型：text 或 table |\n| `meta.title` | list | chunk 所属的多级标题内容 |\n| `meta.position` | list | chunk 的位置，根据分块算法有可能 chunk 跨多个页 |\n| `meta.box` | list | chunk 的位置坐标 |\n| `meta.page_num` | int | chunk 内容所在页数 |\n\n## API 特性\n\n### 异步处理流程\n\n1. 调用提交请求接口 → 获取 `task_id`\n2. 通过 `task_id` 调用获取结果接口轮询\n\n### 轮询建议\n\n- 提交请求后 5~10 秒开始轮询\n- 轮询间隔：5 秒\n- 最大轮询时间：300 秒\n\n### QPS 限制\n\n- 提交请求接口：2 QPS\n- 获取结果接口：10 QPS\n\n## 文件限制\n\n| 限制项 | 说明 |\n|--------|------|\n| 文件大小（file_data） | ≤ 50MB，超过 50M 须使用 file_url |\n| 文件大小（file_url） | PDF ≤ 300MB，非 PDF ≤ 50MB |\n| URL 长度 | ≤ 1024 字节 |\n| 页数限制 | PDF ≤ 2000 页 |\n| 优先级 | file_data > file_url（同时存在时 file_url 字段失效） |\n\n## 错误处理\n\n常见错误码：\n\n| 错误码 | 说明 | 解决方案 |\n|--------|------|----------|\n| 110/111 | access_token 无效或过期 | 重新获取 access_token |\n| 216200 | 文件或 URL 为空 | 提供 file_data 或 file_url |\n| 216201 | 文件格式错误 | 检查文件格式是否支持 |\n| 216202 | 文件大小超限 | 缩减文件大小 |\n| 282000 | 内部错误 | 重试或联系技术支持 |\n| 282003 | 缺少必要参数 | 检查必填参数 |\n| 282007 | 任务不存在 | 检查 task_id 是否正确 |\n| 282018 | 服务繁忙 | 降低请求频率 |\n\n完整错误码参见 `references/error_codes.md`\n\n## 在线调试\n\n可在 [示例代码中心](https://console.bce.baidu.com/tools/#/api?product=AI&project=%E6%96%87%E5%AD%97%E8%AF%86%E5%88%AB&parent=%E6%99%BA%E8%83%BD%E6%96%87%E6%A1%A3%E5%88%86%E6%9E%90%E5%B9%B3%E5%8F%B0&api=rest/2.0/brain/online/v2/parser/task&method=post) 申请试该接口，可进行签名验证、查看在线调用的请求内容和返回结果、示例代码的自动生成。\n\n## 脚本\n\n- `scripts/baidu_doc_parser.py`：文档解析主程序，支持命令行快速调用\n\n## 参考文档\n\n- `references/parameters.md`：完整 API 参数与返回结构详解\n- `references/error_codes.md`：完整错误码参考\n- `references/apikey-fetch.md`：API Key 配置指南\n\n## 相关链接\n\n- [官方 API 文档](https://ai.baidu.com/ai-doc/OCR/llxst5nn0)\n- [百度云控制台](https://console.bce.baidu.com/ai/)\n- [智能文档分析平台](https://ai.baidu.com/solution/intelligent-document-analysis)\n\nFile v1.0.8:_meta.json\n\n{\n  \"ownerId\": \"kn75p1w9cr0c8ycct5wzrb5ken83e9nc\",\n  \"slug\": \"baidu-doc-pipeline-parser\",\n  \"version\": \"1.0.8\",\n  \"publishedAt\": 1789643755631\n}\n\nFile v1.0.8:references/apikey-fetch.md\n\n# 百度文档解析 API Key 配置指南\n\n## BAIDU_DOC_AI_API_KEY 和 BAIDU_DOC_AI_SECRET_KEY 未配置\n\n当环境变量 `BAIDU_DOC_AI_API_KEY` 和 `BAIDU_DOC_AI_SECRET_KEY` 未设置时，按照以下步骤操作：\n\n### 1. 获取 API Key 和 Secret Key\n\n访问：**https://ai.baidu.com/ai-doc/REFERENCE/Ck3dwjhhu#1-获取aksk**\n\n- 登录百度云账号\n- 创建应用或查看已有的 API Key 和 Secret Key\n- 复制你的 **API Key** 和 **Secret Key**\n\n### 2. 领取免费测试资源\n\n访问：**https://ai.baidu.com/ai-doc/OCR/dk3iqnq51**\n\n### 3. 配置环境变量\n\n#### 方式一：直接设置环境变量\n\n```bash\nexport BAIDU_DOC_AI_API_KEY=\"your_actual_api_key_here\"\nexport BAIDU_DOC_AI_SECRET_KEY=\"your_actual_secret_key_here\"\n```\n\n#### 方式二：通过配置文件\n\n编辑配置文件：`~/.claude/settings.json` 或项目 `.claude/settings.json`\n\n添加以下结构：\n\n```json\n{\n  \"skills\": {\n    \"entries\": {\n      \"baidu-doc-pipeline-parser\": {\n        \"env\": {\n          \"BAIDU_DOC_AI_API_KEY\": \"your_actual_api_key_here\",\n          \"BAIDU_DOC_AI_SECRET_KEY\": \"your_actual_secret_key_here\"\n        }\n      }\n    }\n  }\n}\n```\n\n将 `your_actual_api_key_here` 替换为你的实际 API Key，`your_actual_secret_key_here` 替换为你的实际 Secret Key。\n\n### 4. 验证配置\n\n```bash\n# 验证 access_token 是否可正常获取\ncurl -X POST 'https://aip.baidubce.com/oauth/2.0/token' \\\n  -d 'grant_type=client_credentials' \\\n  -d 'client_id={your_api_key}' \\\n  -d 'client_secret={your_secret_key}'\n```\n\n成功返回示例：\n\n```json\n{\n  \"access_token\": \"24.xxxxx.xxxxxx.xxxxxxx-xxxxxxx\",\n  \"expires_in\": 2592000\n}\n```\n\n`expires_in` 为 2592000 秒（30 天），到期后需重新获取。\n\n### 5. 测试\n\n```bash\npython3 scripts/baidu_doc_parser.py --file_url \"https://example.com/test.pdf\" --file_name \"test.pdf\"\npython3 scripts/baidu_doc_parser.py --file_data \"<文件的base64编码>\" --file_name \"test.pdf\"\n```\n\n## 常见问题\n\n- 确保环境变量已正确设置（可通过 `echo $BAIDU_DOC_AI_API_KEY` 验证）\n- 确认 API Key 有效且已开通百度智能文档分析平台服务\n- 检查百度云账户余额或免费额度\n- access_token 有效期 30 天，过期后会自动重新获取\n\n## 相关链接\n\n- [获取 AK/SK 文档](https://ai.baidu.com/ai-doc/REFERENCE/Ck3dwjhhu#1-获取aksk)\n- [领取免费资源](https://ai.baidu.com/ai-doc/OCR/dk3iqnq51)\n- [百度云控制台](https://console.bce.baidu.com/ai/)\n\nFile v1.0.8:references/error_codes.md\n\n# 百度文档解析 API 错误码参考\n\n## 通用错误\n\n### 认证相关错误\n\n| 错误码 | 错误信息 | 说明 | 解决方案 |\n|--------|---------|------|----------|\n| 1 | Unknown error | 未知错误 | 重试请求，持续出现请联系技术支持 |\n| 2 | Service temporarily unavailable | 服务暂不可用 | 重试请求，持续出现请联系技术支持 |\n| 3 | Unsupported openapi method | API 接口不存在 | 检查 URL 是否正确，去除非英文字符 |\n| 4 | Open api request limit reached | 集群超限额 | 重试请求，持续出现请联系技术支持 |\n| 6 | No permission to access data | 无 API 访问权限 | 在百度云控制台开通该 API 权限 |\n| 14 | IAM Certification failed | IAM 认证失败 | 检查签名生成方式或改用 AK/SK |\n| 17 | Open api daily request limit reached | 日配额超限 | 购买额度或等待次日重置 |\n| 18 | Open api qps request limit reached | QPS 超限 | 降低请求频率 |\n| 19 | Open api total request limit reached | 总量配额超限 | 购买额外配额 |\n| 100 | Invalid parameter | access_token 无效 | 重新获取 access_token |\n| 110 | Access token invalid or no longer valid | access_token 无效 | token 有效期 30 天，重新获取 |\n| 111 | Access token expired | access_token 过期 | token 有效期 30 天，重新获取 |\n\n### 文件相关错误\n\n| 错误码 | 错误信息 | 说明 | 解决方案 |\n|--------|---------|------|----------|\n| 216200 | empty file or fileurl | 文件或 URL 为空 | 提供 file_data 或 file_url |\n| 216201 | file format error | 文件格式不支持 | 使用支持的格式（PDF、Word、Excel 等） |\n| 216202 | file size error | 文件大小超限 | 缩减文件大小（file_data ≤ 50MB，file_url PDF ≤ 300MB） |\n\n### 任务处理错误\n\n| 错误码 | 错误信息 | 说明 | 解决方案 |\n|--------|---------|------|----------|\n| 282000 | internal error | 任务处理失败 | 重试或联系技术支持 |\n| 282001 | template not found | 合同类型未找到 | 检查合同类型名称 |\n| 282003 | missing parameters | 缺少必要参数 | 检查必填参数 |\n| 282005 | quota exceed error | 额度不足 | 申请增加配额 |\n| 282006 | check user auth error | 用户权限校验失败 | 验证用户权限 |\n| 282007 | task not exist, please check task id | 任务不存在 | 检查 task_id 是否正确 |\n| 282018 | Service busy | 服务繁忙 | 降低请求频率 |\n\n### URL 相关错误\n\n| 错误码 | 错误信息 | 说明 | 解决方案 |\n|--------|---------|------|----------|\n| 282111 | url format illegal | URL 格式不合法 | 检查 URL 格式 |\n| 282112 | url download timeout | URL 下载超时 | 检查 URL 是否可访问 |\n| 282113 | url response invalid | URL 响应无效 | 检查 URL 返回内容是否正确 |\n| 282114 | url size error | URL 长度超过 1024 字节 | 缩短 URL |\n\n### 参数错误\n\n| 错误码 | 错误信息 | 说明 | 解决方案 |\n|--------|---------|------|----------|\n| 283016 | parameters value error | 参数值无效 | 检查参数格式和取值 |\n\n## 错误响应格式\n\n```json\n{\n  \"log_id\": \"13665091038742503867108513247688\",\n  \"error_code\": \"282007\",\n  \"error_msg\": \"task not exist, please check task id\",\n  \"result\": \"null\"\n}\n```\n\n## 错误处理策略\n\n### 重试策略\n\n| 错误类型 | 错误码 | 建议处理方式 |\n|---------|--------|-------------|\n| 瞬时错误 | 1, 2, 4, 282000, 282018 | 指数退避重试 |\n| 认证错误 | 100, 110, 111 | 重新获取 access_token |\n| 配额错误 | 17, 18, 19, 282005 | 等待或购买额外配额 |\n| 参数错误 | 216200, 216201, 282003, 283016 | 修正参数后重试 |\n| URL 错误 | 282111, 282112, 282113, 282114 | 检查并修正 URL |\n\n### 指数退避重试示例\n\n```python\nimport time\n\ndef retry_with_backoff(func, max_retries=3):\n    for i in range(max_retries):\n        try:\n            return func()\n        except Exception as e:\n            if i == max_retries - 1:\n                raise\n            wait_time = 2 ** i  # 1s, 2s, 4s\n            time.sleep(wait_time)\n```\n\n### Token 刷新示例\n\n```python\ndef ensure_valid_token(client):\n    try:\n        client.query_task(\"test-task-id\")\n    except Exception as e:\n        if \"110\" in str(e) or \"111\" in str(e):\n            client.access_token = client._fetch_auth_credential()\n```\n\n## 获取帮助\n\n- [提交工单](https://ticket.bce.baidu.com/?_=1648086674827&fromai=1#/ticket/create~productId=96&questionId=1306&channel=2)\n- [权限诊断工具](https://console.bce.baidu.com/tools/#/aiInterfacePermissions)\n- [百度云控制台](https://console.bce.baidu.com/ai/)\n\nFile v1.0.8:references/parameters.md\n\n# 百度文档解析 API 参数详解\n\n## 接口概述\n\n百度文档解析 API 支持对 doc、pdf、图片、xlsx 等 18 种格式文档进行解析，输出文档的版面、表格、阅读顺序、标题层级、旋转角度等信息，支持中、英、日、韩、法等 20 余种语言类型，识别准确率可达 90% 以上。\n\n## API 接口地址\n\n### 提交请求接口\n\n```\nPOST https://aip.baidubce.com/rest/2.0/brain/online/v2/parser/task?access_token={access_token}\nContent-Type: application/x-www-form-urlencoded\n```\n\n### 获取结果接口\n\n```\nPOST https://aip.baidubce.com/rest/2.0/brain/online/v2/parser/task/query?access_token={access_token}\nContent-Type: application/x-www-form-urlencoded\n```\n\n## 提交请求参数\n\n### 文件参数（必选，二选一）\n\n| 参数 | 必选 | 类型 | 说明 |\n|------|------|------|------|\n| file_data | 和 file_url 二选一 | string | 文件的 base64 编码数据。版式文档：pdf, jpg, jpeg, png, bmp, tif, tiff, ofd, ppt, pptx；流式文档：doc, docx, txt, xls, xlsx, wps, html, mhtml。文档大小不超过 50M，其中 PDF 文档最大支持 2000 页。若文档大小超过 50M，须从 file_url 方式上传。优先级：file_data > file_url，当 file_data 字段存在时，file_url 字段失效 |\n| file_url | 和 file_data 二选一 | string | 文件数据 URL，URL 长度不超过 1024 字节，支持单个 URL 传入。PDF 文档大小不超过 300MB，非 PDF 文档大小不超过 50M，其中 PDF 文档最大支持 2000 页。优先级：file_data > file_url。**请注意关闭 URL 防盗链** |\n| file_name | 是 | string | 文件名，请保证文件名后缀正确，例如 \"1.pdf\" |\n\n### 核心功能参数\n\n| 参数 | 必选 | 类型 | 可选值范围 | 说明 |\n|------|------|------|----------|------|\n| recognize_formula | 否 | bool | True/False | 是否对版式类型文档进行公式识别 |\n| analysis_chart | 否 | bool | True/False | 是否对统计图表进行解析 |\n| angle_adjust | 否 | bool | True/False | 是否对图片进行角度矫正 |\n| parse_image_layout | 否 | bool | True/False | 是否返回文档中的图片位置信息 |\n\n### 语言与格式参数\n\n| 参数 | 必选 | 类型 | 默认值 | 说明 |\n|------|------|------|--------|------|\n| language_type | 否 | string | CHN_ENG | 识别语种类型 |\n| switch_digital_width | 否 | string | auto | 是否对数字进行全半角转换。auto：不转换；half：半角输出；full：全角输出 |\n| html_table_format | 否 | bool | True | 是否将识别出的表格转换为 HTML 格式返回 |\n\n### 支持语种列表\n\n| 代码 | 语言 | 代码 | 语言 |\n|------|------|------|------|\n| CHN_ENG | 中英文 | DAN | 丹麦语 |\n| JAP | 日语 | DUT | 荷兰语 |\n| KOR | 韩语 | MAL | 马来语 |\n| FRE | 法语 | SWE | 瑞典语 |\n| SPA | 西班牙语 | IND | 印尼语 |\n| POR | 葡萄牙语 | POL | 波兰语 |\n| GER | 德语 | ROM | 罗马尼亚语 |\n| ITA | 意大利语 | TUR | 土耳其语 |\n| RUS | 俄语 | GRE | 希腊语 |\n| HUN | 匈牙利语 | THA | 泰语 |\n| VIE | 越南语 | ARA | 阿拉伯语 |\n| HIN | 印地语 | - | - |\n\n### 文档分块参数\n\n| 参数 | 必选 | 类型 | 默认值 | 说明 |\n|------|------|------|--------|------|\n| return_doc_chunks | 否 | dict | - | 是否返回文档切分后的片段数据（按语义、字数、标点） |\n| + switch | 否 | bool | False | 是否进行文档内容切分 |\n| + split_type | 否 | str | chunk | 切分方式。chunk：按照 chunk_size 来切；mark：按照 separators 来切 |\n| + separators | 否 | list | ['。','；','！','？',';','!','?'] | 切分标点 |\n| + chunk_size | 否 | int | -1 | 切分块的大小，-1 表示按照语义自动切分，不限定块的大小 |\n\n## 获取结果请求参数\n\n| 参数 | 必选 | 类型 | 说明 |\n|------|------|------|------|\n| task_id | 是 | string | 发送提交请求时返回的 task_id |\n\n## 返回结构\n\n### 提交请求返回\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| log_id | uint64 | 唯一的 log id，用于问题定位 |\n| error_code | int | 错误码 |\n| error_msg | string | 错误描述信息 |\n| result | dict | 返回的结果列表 |\n| + task_id | string | 该请求生成的 task_id，后续使用该 task_id 获取结果 |\n\n成功返回示例：\n\n```json\n{\n  \"error_code\": 0,\n  \"error_msg\": \"\",\n  \"log_id\": \"10138598131137362685273585665433\",\n  \"result\": {\n    \"task_id\": \"task-3zy9Bg8CHt1M4pP0cX2q5bg28j268015\"\n  }\n}\n```\n\n### 获取结果返回\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| log_id | uint64 | 唯一的 log id，用于问题定位 |\n| error_code | int | 错误码 |\n| error_msg | string | 错误描述信息 |\n| result | dict | 返回的结果列表 |\n| + task_id | string | 任务 ID |\n| + status | string | 任务状态：pending（排队中）、processing（运行中）、success（成功）、failed（失败） |\n| + task_error | string | 解析报错信息，包含任务失败、额度不够 |\n| + markdown_url | string | 文档解析结果的 Markdown 格式链接，链接有效期 30 天 |\n| + parse_result_url | string | 文档解析结果的 BOS 链接，链接有效期 30 天 |\n\n成功返回示例：\n\n```json\n{\n  \"error_code\": 0,\n  \"error_msg\": \"\",\n  \"result\": {\n    \"task_id\": \"task-UnvGsgbYZp9pS3BZRHn11ifzjNvKzTgf\",\n    \"status\": \"success\",\n    \"task_error\": null,\n    \"duration\": 902.0,\n    \"parse_result_url\": \"https://xxxxxxxxxxxxxxxxxx\"\n  }\n}\n```\n\n### parse_result_url 返回的 JSON 结构\n\n通过 `parse_result_url` 下载解析结果的 JSON 文件，结构如下：\n\n#### 顶层结构\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| file_name | string | 文档名称 |\n| file_id | string | 文档 ID |\n| pages | list | 文件单页解析内容 |\n| chunks | list | 文件内容切分结果（return_doc_chunks 中 switch 为 True 时有值） |\n\n#### 页面对象（pages[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| page_id | string | 页码 ID |\n| page_num | int | 页码数，从 0 开始 |\n| text | string | 当前页的所有纯文字内容 |\n| layouts | list | 页面内容版式分析的结果 |\n| tables | list | 页面表格解析结果 |\n| images | list | 页面中图片解析结果 |\n| meta | dict | 页元信息 |\n\n#### 页面元信息（meta）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| page_width | int | 页面宽度 |\n| page_height | int | 页面高度 |\n| is_scan | bool | 是否扫描件 |\n| page_angle | int | 页面倾斜角度 |\n| page_type | string | 页面属性：text（正文）、contents（目录）、appendix（附录）、others（其他） |\n| sheet_name | string | Excel 的 sheet 名 |\n\n#### 版面元素（layouts[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| layout_id | string | layout 元素唯一标志，以 \"xxxxx-layout-{global_layout_index}\" 形式，global_layout_index 为 layout 元素整个文档的全局索引 |\n| text | string | layout 对应的文本内容。注：当 type 为 table、image 时该字段为空，需根据 type 和 layout_id 分别到 tables、images 字段里找到对应的内容 |\n| position | list | layout 元素在页面中的位置，[x, y, w, h] box 框，左上角和宽高 |\n| type | string | 版面元素类型（见下表） |\n| sub_type | string | 版面元素子类型（见下表） |\n| parent | string | 标题层级树中父节点的 layout_id，若当前 layout 为一级标题，其 parent 为 \"root\"。在 table 和 image 的内版面信息中暂时都为空 |\n| children | list | 标题层级树中子节点的 layout_id。在 table 和 image 的内版面信息中暂时都为空 |\n\n**版面类型（type）取值**：\n\n| 类型 | 说明 |\n|------|------|\n| text | 段落 |\n| table | 表格 |\n| image | 文档中的插图 |\n| head_tail | 页面顶部 |\n| contents | 目录 |\n| seal | 印章 |\n| title | 标题 |\n| formula | 公式 |\n| hand_sign | 手写签名 |\n\n**子类型（sub_type）取值**：\n\n当 type 为 title 或 image 时，sub_type 有值：\n\n- **title 的 sub_type**：\n  - `title_{n}`：代表 n 级标题，比如 title_2 代表二级标题\n  - `image_title`：图标题\n  - `table_title`：表标题\n\n- **image 的 sub_type**：\n  - `chart`：统计图表\n  - `figure`：普通插图\n  - `QR_code`：二维码\n  - `Bar_code`：条形码\n\n#### 表格对象（tables[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| layout_id | string | layout ID，与 layouts 中的 type 为 table 的元素的 layout ID 对应 |\n| markdown | string | 表格内容的 Markdown 形式 |\n| table_title_id | list | 表格标题对应的 layout_id，默认为 null |\n| position | list | 边框数据 [x, y, w, h]（以页面坐标为原点），版式格式时有效 |\n| cells | list | 单元格的内版面信息，layout 类型为表格时有值 |\n| matrix | list | 二位数组，表示表格内布局位置信息，每个元素对应 cells 列表中元素的索引 |\n| merge_table | string | 跨页表格标记：\"begin\"（开始）、\"inner\"（中间，表格跨页超过两页）、\"end\"（结束）；非跨页表格该字段为空 |\n\n#### 图片对象（images[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| layout_id | string | layout ID，与 layouts 中 type 为 image 的元素的 layout ID 对应 |\n| image_title_id | list | 图片标题对应的 layout_id，默认为 null |\n| position | list | 边框数据 [x, y, w, h] |\n| content_layouts | list | 图片的内版面信息 |\n| data_url | string | 图片存储链接 |\n| image_description | string | 对统计图表进行内容解析和描述，输出结果为 JSON 字符串，可通过 json.loads 结构化为 JSON 格式 |\n\n#### 分块对象（chunks[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| chunk_id | string | 切片的 ID |\n| content | string | 切片的内容 |\n| type | string | 切片类型，为 text 或者 table |\n| meta | dict | chunk 元信息 |\n| + title | list | chunk 所属的多级标题内容 |\n| + position | list | chunk 的位置，根据分块算法有可能 chunk 跨多个页 |\n| + box | list | chunk 的位置坐标 |\n| + page_num | int | chunk 内容所在页数 |\n\n## 文件限制\n\n| 限制项 | 说明 |\n|--------|------|\n| 文件大小（file_data） | ≤ 50MB，超过 50M 须使用 file_url |\n| 文件大小（file_url） | PDF ≤ 300MB，非 PDF ≤ 50MB |\n| URL 长度 | ≤ 1024 字节 |\n| 页数限制 | PDF ≤ 2000 页 |\n| 优先级 | file_data > file_url（同时存在时 file_url 字段失效） |\n\n## QPS 限制\n\n- 提交请求接口：2 QPS\n- 获取结果接口：10 QPS\n\n## 轮询建议\n\n- 提交请求后 5~10 秒开始轮询\n- 轮询间隔：5 秒\n- 最大轮询时间：300 秒\n\n## 相关文档\n\n- [官方 API 文档](https://ai.baidu.com/ai-doc/OCR/llxst5nn0)\n- [错误码参考](error_codes.md)\n- [API Key 配置指南](apikey-fetch.md)\n\nFile v1.0.8:skill-card.md\n\n## Description:\n\nParses PDFs, Office files, images, and other documents with Baidu's document parsing API to extract text, tables, layout structure, OCR output, and RAG-ready chunks.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[maglanyulan](https://clawhub.ai/user/maglanyulan)\n\n### License/Terms of Use:\n\nMIT\n\n## Use Case:\n\nDevelopers and document-processing teams use this skill to submit supported documents to Baidu's intelligent document analysis API, poll asynchronous parsing jobs, and retrieve structured JSON or Markdown results for extraction, OCR, layout analysis, table parsing, and RAG chunking workflows.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: Documents or document URLs submitted through the skill are sent to Baidu for processing.\n\nMitigation: Use the skill only for documents approved for third-party processing, and avoid confidential, regulated, or credential-containing files unless the deployment has explicit approval.\n\nRisk: Returned markdown_url and parse_result_url links remain sensitive while they are valid.\n\nMitigation: Restrict access to result links, avoid posting them in shared channels, and delete or rotate stored outputs according to local retention rules.\n\nRisk: Baidu API keys can be exposed if committed to project settings or shared configuration files.\n\nMitigation: Store BAIDU_DOC_AI_API_KEY and BAIDU_DOC_AI_SECRET_KEY in a secure secret store or local environment and keep them out of source control.\n\n## Reference(s):\n\n- [ClawHub skill page](https://clawhub.ai/maglanyulan/skills/baidu-doc-pipeline-parser)\n- [Baidu document parsing API documentation](https://ai.baidu.com/ai-doc/OCR/llxst5nn0)\n- [Baidu intelligent document analysis platform](https://ai.baidu.com/solution/intelligent-document-analysis)\n- [API key setup guide](references/apikey-fetch.md)\n- [API parameters and response structure](references/parameters.md)\n- [Error code reference](references/error_codes.md)\n\n## Skill Output:\n\n**Output Type(s):** [guidance, shell commands, code, configuration, markdown, JSON]\n\n**Output Format:** [Markdown guidance with shell commands and JSON API results]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [May return parsed document JSON, Markdown result links, and downloaded parse results from Baidu's asynchronous API.]\n\n## Skill Version(s):\n\n1.0.8 (source: server-resolved release metadata)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nArchive v1.0.7: 7 files, 18266 bytes\n\nFiles: _meta.json (144b), references/apikey-fetch.md (2455b), references/error_codes.md (4611b), references/parameters.md (10795b), scripts/baidu_doc_parser.py (11067b), skill-card.md (2713b), SKILL.md (13803b)\n\nFile v1.0.7:SKILL.md\n\n---\nname: baidu-doc-pipeline-parser 百度文档解析\ndescription: 调用百度文档解析API解析文档。支持PDF、Word、Excel、PPT、图片等18+格式。提取文本、表格、版面分析、OCR识别及RAG文档分块。当用户需要解析文档、提取文本/表格、分析文档结构、处理扫描件时使用。触发词：文档解析、PDF解析、Word解析、表格提取、OCR、文档分析、提取文本、文档结构、扫描识别。\nlicense: MIT\n---\n\n# 百度文档解析 Skill\n\n基于百度智能文档分析平台 API，提供文档解析能力。\n\n## 功能概述\n\n- 支持对 doc、pdf、图片、xlsx 等 18 种格式文档进行解析\n- 输出文档的版面、表格、阅读顺序、标题层级、旋转角度等信息\n- 支持中、英、日、韩、法等 20 余种语言类型\n- 可返回 Markdown 格式内容，将非结构化数据转化为易于处理的结构化数据\n- 识别准确率可达 90% 以上\n- 文档分块（适用于 RAG 场景）\n\n## 适用场景\n\n当用户需要：\n- 解析 PDF、Word、Excel 等格式文档\n- 从文档中提取文本内容\n- 识别并提取表格数据\n- 分析文档结构（标题层级、章节、版面布局）\n- 对扫描件进行 OCR 文字识别\n- 将文档分块用于 RAG 应用\n\n\n## 免费资源领取和计费说明\n[百度智能文档分析平台 领取免费测试资源](https://cloud.baidu.com/doc/OCR/s/fk3h7xu7h)\n\n\n[百度智能文档分析平台计费与购买方式](https://cloud.baidu.com/doc/OCR/s/Fls06fa15#%E6%96%87%E6%A1%A3%E8%A7%A3%E6%9E%90%EF%BC%88paddleocr-vl%EF%BC%89)\n\n\n| 用户类型 | 免费额度 |\n|---------|---------|\n| 个人实名认证用户 | **200 页** |\n| 企业实名认证用户 | **1000 页** |\n\n## API 配置\n\n### 额度获取方式\n\n您可通过百度智能云平台获取[免费额度](https://cloud.baidu.com/doc/OCR/s/fk3h7xu7h)与[购买调用资源](https://cloud.baidu.com/doc/OCR/s/Fls06fa15#%E6%96%87%E6%A1%A3%E8%A7%A3%E6%9E%90%EF%BC%88paddleocr-vl%EF%BC%89)\n\n### 环境变量（必须）\n\n[百度智能文档分析平台 领取免费测试资源](https://ai.baidu.com/ai-doc/OCR/dk3iqnq51)\n\n使用前请设置以下环境变量：\n\n```bash\nexport BAIDU_DOC_AI_API_KEY=\"your_api_key\"\nexport BAIDU_DOC_AI_SECRET_KEY=\"your_secret_key\"\n```\n\n### 认证方式\n\n通过 API Key 和 Secret Key 获取 access_token，有效期 30 天。\n\n## 支持格式\n\n**版式文档**：pdf, jpg, jpeg, png, bmp, tif, tiff, ofd, ppt, pptx\n\n**流式文档**：doc, docx, txt, xls, xlsx, wps, html, mhtml\n\n## 支持语言\n\nCHN_ENG（中英文）、JAP（日语）、KOR（韩语）、FRE（法语）、SPA（西班牙语）、POR（葡萄牙语）、GER（德语）、ITA（意大利语）、RUS（俄语）、DAN（丹麦语）、DUT（荷兰语）、MAL（马来语）、SWE（瑞典语）、IND（印尼语）、POL（波兰语）、ROM（罗马尼亚语）、TUR（土耳其语）、GRE（希腊语）、HUN（匈牙利语）、THA（泰语）、VIE（越南语）、ARA（阿拉伯语）、HIN（印地语）\n\n## 使用方式\n\n```bash\npython3 scripts/baidu_doc_parser.py --file_data <文件的base64编码>\npython3 scripts/baidu_doc_parser.py --file_url <文件公网URL>\n```\n\n## API 接口\n\n文档解析 API 服务为异步接口，需要先调用**提交请求接口**获取 task_id，然后调用**获取结果接口**进行结果轮询。\n\n### 提交请求接口\n\n- **HTTP 方法**：POST\n- **请求 URL**：`https://aip.baidubce.com/rest/2.0/brain/online/v2/parser/task?access_token={token}`\n- **Content-Type**：`application/x-www-form-urlencoded`\n\n### 获取结果接口\n\n- **HTTP 方法**：POST\n- **请求 URL**：`https://aip.baidubce.com/rest/2.0/brain/online/v2/parser/task/query?access_token={token}`\n- **Content-Type**：`application/x-www-form-urlencoded`\n- **请求参数**：`task_id`（必填，提交请求时返回的 task_id）\n\n## 请求参数\n\n### 文件参数（必选，二选一）\n\n| 参数 | 必选 | 类型 | 说明 |\n|------|------|------|------|\n| `file_data` | 和 file_url 二选一 | string | 文件 Base64 编码数据。版式文档：pdf, jpg, jpeg, png, bmp, tif, tiff, ofd, ppt, pptx；流式文档：doc, docx, txt, xls, xlsx, wps, html, mhtml。文档大小不超过 50M，PDF 最大支持 2000 页。**若文档大小超过 50M，须从 file_url 方式上传**。优先级：file_data > file_url |\n| `file_url` | 和 file_data 二选一 | string | 文件数据 URL，长度不超过 1024 字节，支持单个 URL 传入。PDF 文档大小不超过 300MB，非 PDF 不超过 50M，PDF 最大支持 2000 页。**请注意关闭 URL 防盗链** |\n| `file_name` | 是 | string | 文件名，请保证文件名后缀正确，例如 \"1.pdf\" |\n\n### 核心功能参数\n\n| 参数 | 必选 | 类型 | 可选值范围 | 说明 |\n|------|------|------|----------|------|\n| `recognize_formula` | 否 | bool | True/False | 是否对版式类型文档进行公式识别 |\n| `analysis_chart` | 否 | bool | True/False | 是否对统计图表进行解析 |\n| `angle_adjust` | 否 | bool | True/False | 是否对图片进行角度矫正 |\n| `parse_image_layout` | 否 | bool | True/False | 是否返回文档中的图片位置信息 |\n\n### 语言与格式参数\n\n| 参数 | 必选 | 类型 | 说明 |\n|------|------|------|------|\n| `language_type` | 否 | string | 识别语种类型，默认为 CHN_ENG（中英文） |\n| `switch_digital_width` | 否 | string | 是否对数字进行全半角转换，默认为 auto。可选：auto（不转换）、half（半角输出）、full（全角输出） |\n| `html_table_format` | 否 | bool | 是否将识别出的表格转换为 HTML 格式返回，**default=True** |\n\n### 文档分块参数\n\n`return_doc_chunks` 为字典类型，用于返回文档切分后的片段数据（按语义、字数、标点）：\n\n| 参数 | 必选 | 类型 | 默认值 | 说明 |\n|------|------|------|--------|------|\n| `switch` | 否 | bool | False | 是否进行文档内容切分 |\n| `split_type` | 否 | str | chunk | 切分方式：chunk（按 chunk_size 来切）/ mark（按 separators 来切） |\n| `separators` | 否 | list | ['。','；','！','？',';','!','?'] | 切分标点 |\n| `chunk_size` | 否 | int | -1 | 切分块的大小，-1 表示按照语义自动切分，不限定块的大小 |\n\n## 返回结构\n\n### 提交请求返回\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `log_id` | uint64 | 唯一的 log id，用于问题定位 |\n| `error_code` | int | 错误码 |\n| `error_msg` | string | 错误描述信息 |\n| `result.task_id` | string | 该请求生成的 task_id，后续使用该 task_id 获取审查结果 |\n\n### 获取结果返回\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `log_id` | uint64 | 唯一的 log id |\n| `error_code` | int | 错误码 |\n| `error_msg` | string | 错误描述信息 |\n| `result.task_id` | string | 任务 ID |\n| `result.status` | string | 任务状态：pending（排队中）、processing（运行中）、success（成功）、failed（失败） |\n| `result.task_error` | string | 解析报错信息，包含任务失败、额度不够 |\n| `result.markdown_url` | string | 文档解析结果的 Markdown 格式链接，**链接有效期 30 天** |\n| `result.parse_result_url` | string | 文档解析结果的 BOS 链接（JSON），**链接有效期 30 天** |\n\n### 解析结果 JSON 结构（parse_result_url）\n\n#### 顶层结构\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `file_name` | string | 文档名称 |\n| `file_id` | string | 文档 ID |\n| `pages` | list | 文件单页解析内容 |\n| `chunks` | list | 文件内容切分结果（return_doc_chunks.switch=True 时有值） |\n\n#### 页面对象（pages[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `page_id` | string | 页码 ID |\n| `page_num` | int | 页码数 |\n| `text` | string | 当前页的所有纯文字内容 |\n| `layouts` | list | 页面内容版式分析的结果 |\n| `tables` | list | 页面表格解析结果 |\n| `images` | list | 页面中图片解析结果 |\n| `meta` | dict | 页元信息 |\n\n#### 页面元信息（meta）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `page_width` | int | 页面宽度 |\n| `page_height` | int | 页面高度 |\n| `is_scan` | bool | 是否扫描件 |\n| `page_angle` | int | 页面倾斜角度 |\n| `page_type` | string | 页面属性：text（正文）、contents（目录）、appendix（附录）、others（其他） |\n| `sheet_name` | string | Excel 的 sheet 名 |\n\n#### 版面元素（layouts[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `layout_id` | string | layout 元素唯一标志，格式 \"xxxxx-layout-{global_layout_index}\" |\n| `text` | string | layout 对应的文本内容。注：当 type 为 table/image 时该字段为空，需根据 type 和 layout_id 分别到 tables/images 字段里找到对应内容 |\n| `position` | list | 元素在页面中的位置 [x, y, w, h]，左上角和宽高 |\n| `type` | string | 版面元素类型（见下表） |\n| `sub_type` | string | 版面元素子类型（见下表） |\n| `parent` | string | 标题层级树中父节点的 layout_id，若为一级标题则 parent 为 \"root\" |\n| `children` | list | 标题层级树中子节点的 layout_id 列表 |\n\n**版面类型（type）**：\n\n| 类型 | 说明 |\n|------|------|\n| `para` | 段落 |\n| `table` | 表格 |\n| `image` | 文档中的插图 |\n| `head_tail` | 页面顶部（页眉/页脚） |\n| `contents` | 目录 |\n| `seal` | 印章 |\n| `title` | 标题 |\n| `formula` | 公式 |\n\n**子类型（sub_type）**：\n\n- **title 类**：`title_{n}`（n 级标题，如 title_2 代表二级标题）、`image_title`（图标题）、`table_title`（表标题）\n- **image 类**：`chart`（统计图表）、`figure`（普通插图）、`QR_code`（二维码）、`Bar_code`（条形码）\n\n#### 表格对象（tables[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `layout_id` | string | 与 layouts 中 type 为 table 的元素的 layout ID 对应 |\n| `markdown` | string | 表格内容的 Markdown 形式 |\n| `table_title_id` | list | 表格标题对应的 layout_id，默认为 null |\n| `position` | list | 边框数据 [x, y, w, h]（以页面坐标为原点），版式格式时有效 |\n| `cells` | list | 单元格的内版面信息，layout 类型为表格时有值 |\n| `matrix` | list | 二位数组，表示表格内布局位置信息，每个元素对应 cells 列表中元素的索引 |\n| `merge_table` | string | 跨页表格标记：begin（开始）、inner（中间，超过两页）、end（结束）；非跨页表格该字段为空 |\n\n#### 图片对象（images[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `layout_id` | string | 与 layouts 中 type 为 image 的元素的 layout ID 对应 |\n| `image_title_id` | list | 图片标题对应的 layout_id，默认为 null |\n| `position` | list | 边框数据 [x, y, w, h] |\n| `content_layouts` | list | 图片的内版面信息 |\n| `data_url` | string | 图片存储链接 |\n| `image_description` | string | 对统计图表进行内容解析和描述，输出结果为 JSON 字符串 |\n\n#### 分块对象（chunks[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `chunk_id` | string | 切片的 ID |\n| `content` | string | 切片的内容 |\n| `type` | string | 切片类型：text 或 table |\n| `meta.title` | list | chunk 所属的多级标题内容 |\n| `meta.position` | list | chunk 的位置，根据分块算法有可能 chunk 跨多个页 |\n| `meta.box` | list | chunk 的位置坐标 |\n| `meta.page_num` | int | chunk 内容所在页数 |\n\n## API 特性\n\n### 异步处理流程\n\n1. 调用提交请求接口 → 获取 `task_id`\n2. 通过 `task_id` 调用获取结果接口轮询\n\n### 轮询建议\n\n- 提交请求后 5~10 秒开始轮询\n- 轮询间隔：5 秒\n- 最大轮询时间：300 秒\n\n### QPS 限制\n\n- 提交请求接口：2 QPS\n- 获取结果接口：10 QPS\n\n## 文件限制\n\n| 限制项 | 说明 |\n|--------|------|\n| 文件大小（file_data） | ≤ 50MB，超过 50M 须使用 file_url |\n| 文件大小（file_url） | PDF ≤ 300MB，非 PDF ≤ 50MB |\n| URL 长度 | ≤ 1024 字节 |\n| 页数限制 | PDF ≤ 2000 页 |\n| 优先级 | file_data > file_url（同时存在时 file_url 字段失效） |\n\n## 错误处理\n\n常见错误码：\n\n| 错误码 | 说明 | 解决方案 |\n|--------|------|----------|\n| 110/111 | access_token 无效或过期 | 重新获取 access_token |\n| 216200 | 文件或 URL 为空 | 提供 file_data 或 file_url |\n| 216201 | 文件格式错误 | 检查文件格式是否支持 |\n| 216202 | 文件大小超限 | 缩减文件大小 |\n| 282000 | 内部错误 | 重试或联系技术支持 |\n| 282003 | 缺少必要参数 | 检查必填参数 |\n| 282007 | 任务不存在 | 检查 task_id 是否正确 |\n| 282018 | 服务繁忙 | 降低请求频率 |\n\n完整错误码参见 `references/error_codes.md`\n\n## 在线调试\n\n可在 [示例代码中心](https://console.bce.baidu.com/tools/#/api?product=AI&project=%E6%96%87%E5%AD%97%E8%AF%86%E5%88%AB&parent=%E6%99%BA%E8%83%BD%E6%96%87%E6%A1%A3%E5%88%86%E6%9E%90%E5%B9%B3%E5%8F%B0&api=rest/2.0/brain/online/v2/parser/task&method=post) 申请试该接口，可进行签名验证、查看在线调用的请求内容和返回结果、示例代码的自动生成。\n\n## 脚本\n\n- `scripts/baidu_doc_parser.py`：文档解析主程序，支持命令行快速调用\n\n## 参考文档\n\n- `references/parameters.md`：完整 API 参数与返回结构详解\n- `references/error_codes.md`：完整错误码参考\n- `references/apikey-fetch.md`：API Key 配置指南\n\n## 相关链接\n\n- [官方 API 文档](https://ai.baidu.com/ai-doc/OCR/llxst5nn0)\n- [百度云控制台](https://console.bce.baidu.com/ai/)\n- [智能文档分析平台](https://ai.baidu.com/solution/intelligent-document-analysis)\n\nFile v1.0.7:_meta.json\n\n{\n  \"ownerId\": \"kn75p1w9cr0c8ycct5wzrb5ken83e9nc\",\n  \"slug\": \"baidu-doc-pipeline-parser\",\n  \"version\": \"1.0.7\",\n  \"publishedAt\": 1783500041710\n}\n\nFile v1.0.7:references/apikey-fetch.md\n\n# 百度文档解析 API Key 配置指南\n\n## BAIDU_DOC_AI_API_KEY 和 BAIDU_DOC_AI_SECRET_KEY 未配置\n\n当环境变量 `BAIDU_DOC_AI_API_KEY` 和 `BAIDU_DOC_AI_SECRET_KEY` 未设置时，按照以下步骤操作：\n\n### 1. 获取 API Key 和 Secret Key\n\n访问：**https://ai.baidu.com/ai-doc/REFERENCE/Ck3dwjhhu#1-获取aksk**\n\n- 登录百度云账号\n- 创建应用或查看已有的 API Key 和 Secret Key\n- 复制你的 **API Key** 和 **Secret Key**\n\n### 2. 领取免费测试资源\n\n访问：**https://ai.baidu.com/ai-doc/OCR/dk3iqnq51**\n\n### 3. 配置环境变量\n\n#### 方式一：直接设置环境变量\n\n```bash\nexport BAIDU_DOC_AI_API_KEY=\"your_actual_api_key_here\"\nexport BAIDU_DOC_AI_SECRET_KEY=\"your_actual_secret_key_here\"\n```\n\n#### 方式二：通过配置文件\n\n编辑配置文件：`~/.claude/settings.json` 或项目 `.claude/settings.json`\n\n添加以下结构：\n\n```json\n{\n  \"skills\": {\n    \"entries\": {\n      \"baidu-doc-pipeline-parser\": {\n        \"env\": {\n          \"BAIDU_DOC_AI_API_KEY\": \"your_actual_api_key_here\",\n          \"BAIDU_DOC_AI_SECRET_KEY\": \"your_actual_secret_key_here\"\n        }\n      }\n    }\n  }\n}\n```\n\n将 `your_actual_api_key_here` 替换为你的实际 API Key，`your_actual_secret_key_here` 替换为你的实际 Secret Key。\n\n### 4. 验证配置\n\n```bash\n# 验证 access_token 是否可正常获取\ncurl -X POST 'https://aip.baidubce.com/oauth/2.0/token' \\\n  -d 'grant_type=client_credentials' \\\n  -d 'client_id={your_api_key}' \\\n  -d 'client_secret={your_secret_key}'\n```\n\n成功返回示例：\n\n```json\n{\n  \"access_token\": \"24.xxxxx.xxxxxx.xxxxxxx-xxxxxxx\",\n  \"expires_in\": 2592000\n}\n```\n\n`expires_in` 为 2592000 秒（30 天），到期后需重新获取。\n\n### 5. 测试\n\n```bash\npython3 scripts/baidu_doc_parser.py --file_url \"https://example.com/test.pdf\" --file_name \"test.pdf\"\npython3 scripts/baidu_doc_parser.py --file_data \"<文件的base64编码>\" --file_name \"test.pdf\"\n```\n\n## 常见问题\n\n- 确保环境变量已正确设置（可通过 `echo $BAIDU_DOC_AI_API_KEY` 验证）\n- 确认 API Key 有效且已开通百度智能文档分析平台服务\n- 检查百度云账户余额或免费额度\n- access_token 有效期 30 天，过期后会自动重新获取\n\n## 相关链接\n\n- [获取 AK/SK 文档](https://ai.baidu.com/ai-doc/REFERENCE/Ck3dwjhhu#1-获取aksk)\n- [领取免费资源](https://ai.baidu.com/ai-doc/OCR/dk3iqnq51)\n- [百度云控制台](https://console.bce.baidu.com/ai/)\n\nFile v1.0.7:references/error_codes.md\n\n# 百度文档解析 API 错误码参考\n\n## 通用错误\n\n### 认证相关错误\n\n| 错误码 | 错误信息 | 说明 | 解决方案 |\n|--------|---------|------|----------|\n| 1 | Unknown error | 未知错误 | 重试请求，持续出现请联系技术支持 |\n| 2 | Service temporarily unavailable | 服务暂不可用 | 重试请求，持续出现请联系技术支持 |\n| 3 | Unsupported openapi method | API 接口不存在 | 检查 URL 是否正确，去除非英文字符 |\n| 4 | Open api request limit reached | 集群超限额 | 重试请求，持续出现请联系技术支持 |\n| 6 | No permission to access data | 无 API 访问权限 | 在百度云控制台开通该 API 权限 |\n| 14 | IAM Certification failed | IAM 认证失败 | 检查签名生成方式或改用 AK/SK |\n| 17 | Open api daily request limit reached | 日配额超限 | 购买额度或等待次日重置 |\n| 18 | Open api qps request limit reached | QPS 超限 | 降低请求频率 |\n| 19 | Open api total request limit reached | 总量配额超限 | 购买额外配额 |\n| 100 | Invalid parameter | access_token 无效 | 重新获取 access_token |\n| 110 | Access token invalid or no longer valid | access_token 无效 | token 有效期 30 天，重新获取 |\n| 111 | Access token expired | access_token 过期 | token 有效期 30 天，重新获取 |\n\n### 文件相关错误\n\n| 错误码 | 错误信息 | 说明 | 解决方案 |\n|--------|---------|------|----------|\n| 216200 | empty file or fileurl | 文件或 URL 为空 | 提供 file_data 或 file_url |\n| 216201 | file format error | 文件格式不支持 | 使用支持的格式（PDF、Word、Excel 等） |\n| 216202 | file size error | 文件大小超限 | 缩减文件大小（file_data ≤ 50MB，file_url PDF ≤ 300MB） |\n\n### 任务处理错误\n\n| 错误码 | 错误信息 | 说明 | 解决方案 |\n|--------|---------|------|----------|\n| 282000 | internal error | 任务处理失败 | 重试或联系技术支持 |\n| 282001 | template not found | 合同类型未找到 | 检查合同类型名称 |\n| 282003 | missing parameters | 缺少必要参数 | 检查必填参数 |\n| 282005 | quota exceed error | 额度不足 | 申请增加配额 |\n| 282006 | check user auth error | 用户权限校验失败 | 验证用户权限 |\n| 282007 | task not exist, please check task id | 任务不存在 | 检查 task_id 是否正确 |\n| 282018 | Service busy | 服务繁忙 | 降低请求频率 |\n\n### URL 相关错误\n\n| 错误码 | 错误信息 | 说明 | 解决方案 |\n|--------|---------|------|----------|\n| 282111 | url format illegal | URL 格式不合法 | 检查 URL 格式 |\n| 282112 | url download timeout | URL 下载超时 | 检查 URL 是否可访问 |\n| 282113 | url response invalid | URL 响应无效 | 检查 URL 返回内容是否正确 |\n| 282114 | url size error | URL 长度超过 1024 字节 | 缩短 URL |\n\n### 参数错误\n\n| 错误码 | 错误信息 | 说明 | 解决方案 |\n|--------|---------|------|----------|\n| 283016 | parameters value error | 参数值无效 | 检查参数格式和取值 |\n\n## 错误响应格式\n\n```json\n{\n  \"log_id\": \"13665091038742503867108513247688\",\n  \"error_code\": \"282007\",\n  \"error_msg\": \"task not exist, please check task id\",\n  \"result\": \"null\"\n}\n```\n\n## 错误处理策略\n\n### 重试策略\n\n| 错误类型 | 错误码 | 建议处理方式 |\n|---------|--------|-------------|\n| 瞬时错误 | 1, 2, 4, 282000, 282018 | 指数退避重试 |\n| 认证错误 | 100, 110, 111 | 重新获取 access_token |\n| 配额错误 | 17, 18, 19, 282005 | 等待或购买额外配额 |\n| 参数错误 | 216200, 216201, 282003, 283016 | 修正参数后重试 |\n| URL 错误 | 282111, 282112, 282113, 282114 | 检查并修正 URL |\n\n### 指数退避重试示例\n\n```python\nimport time\n\ndef retry_with_backoff(func, max_retries=3):\n    for i in range(max_retries):\n        try:\n            return func()\n        except Exception as e:\n            if i == max_retries - 1:\n                raise\n            wait_time = 2 ** i  # 1s, 2s, 4s\n            time.sleep(wait_time)\n```\n\n### Token 刷新示例\n\n```python\ndef ensure_valid_token(client):\n    try:\n        client.query_task(\"test-task-id\")\n    except Exception as e:\n        if \"110\" in str(e) or \"111\" in str(e):\n            client.access_token = client._fetch_auth_credential()\n```\n\n## 获取帮助\n\n- [提交工单](https://ticket.bce.baidu.com/?_=1648086674827&fromai=1#/ticket/create~productId=96&questionId=1306&channel=2)\n- [权限诊断工具](https://console.bce.baidu.com/tools/#/aiInterfacePermissions)\n- [百度云控制台](https://console.bce.baidu.com/ai/)\n\nFile v1.0.7:references/parameters.md\n\n# 百度文档解析 API 参数详解\n\n## 接口概述\n\n百度文档解析 API 支持对 doc、pdf、图片、xlsx 等 18 种格式文档进行解析，输出文档的版面、表格、阅读顺序、标题层级、旋转角度等信息，支持中、英、日、韩、法等 20 余种语言类型，识别准确率可达 90% 以上。\n\n## API 接口地址\n\n### 提交请求接口\n\n```\nPOST https://aip.baidubce.com/rest/2.0/brain/online/v2/parser/task?access_token={access_token}\nContent-Type: application/x-www-form-urlencoded\n```\n\n### 获取结果接口\n\n```\nPOST https://aip.baidubce.com/rest/2.0/brain/online/v2/parser/task/query?access_token={access_token}\nContent-Type: application/x-www-form-urlencoded\n```\n\n## 提交请求参数\n\n### 文件参数（必选，二选一）\n\n| 参数 | 必选 | 类型 | 说明 |\n|------|------|------|------|\n| file_data | 和 file_url 二选一 | string | 文件的 base64 编码数据。版式文档：pdf, jpg, jpeg, png, bmp, tif, tiff, ofd, ppt, pptx；流式文档：doc, docx, txt, xls, xlsx, wps, html, mhtml。文档大小不超过 50M，其中 PDF 文档最大支持 2000 页。若文档大小超过 50M，须从 file_url 方式上传。优先级：file_data > file_url，当 file_data 字段存在时，file_url 字段失效 |\n| file_url | 和 file_data 二选一 | string | 文件数据 URL，URL 长度不超过 1024 字节，支持单个 URL 传入。PDF 文档大小不超过 300MB，非 PDF 文档大小不超过 50M，其中 PDF 文档最大支持 2000 页。优先级：file_data > file_url。**请注意关闭 URL 防盗链** |\n| file_name | 是 | string | 文件名，请保证文件名后缀正确，例如 \"1.pdf\" |\n\n### 核心功能参数\n\n| 参数 | 必选 | 类型 | 可选值范围 | 说明 |\n|------|------|------|----------|------|\n| recognize_formula | 否 | bool | True/False | 是否对版式类型文档进行公式识别 |\n| analysis_chart | 否 | bool | True/False | 是否对统计图表进行解析 |\n| angle_adjust | 否 | bool | True/False | 是否对图片进行角度矫正 |\n| parse_image_layout | 否 | bool | True/False | 是否返回文档中的图片位置信息 |\n\n### 语言与格式参数\n\n| 参数 | 必选 | 类型 | 默认值 | 说明 |\n|------|------|------|--------|------|\n| language_type | 否 | string | CHN_ENG | 识别语种类型 |\n| switch_digital_width | 否 | string | auto | 是否对数字进行全半角转换。auto：不转换；half：半角输出；full：全角输出 |\n| html_table_format | 否 | bool | True | 是否将识别出的表格转换为 HTML 格式返回 |\n\n### 支持语种列表\n\n| 代码 | 语言 | 代码 | 语言 |\n|------|------|------|------|\n| CHN_ENG | 中英文 | DAN | 丹麦语 |\n| JAP | 日语 | DUT | 荷兰语 |\n| KOR | 韩语 | MAL | 马来语 |\n| FRE | 法语 | SWE | 瑞典语 |\n| SPA | 西班牙语 | IND | 印尼语 |\n| POR | 葡萄牙语 | POL | 波兰语 |\n| GER | 德语 | ROM | 罗马尼亚语 |\n| ITA | 意大利语 | TUR | 土耳其语 |\n| RUS | 俄语 | GRE | 希腊语 |\n| HUN | 匈牙利语 | THA | 泰语 |\n| VIE | 越南语 | ARA | 阿拉伯语 |\n| HIN | 印地语 | - | - |\n\n### 文档分块参数\n\n| 参数 | 必选 | 类型 | 默认值 | 说明 |\n|------|------|------|--------|------|\n| return_doc_chunks | 否 | dict | - | 是否返回文档切分后的片段数据（按语义、字数、标点） |\n| + switch | 否 | bool | False | 是否进行文档内容切分 |\n| + split_type | 否 | str | chunk | 切分方式。chunk：按照 chunk_size 来切；mark：按照 separators 来切 |\n| + separators | 否 | list | ['。','；','！','？',';','!','?'] | 切分标点 |\n| + chunk_size | 否 | int | -1 | 切分块的大小，-1 表示按照语义自动切分，不限定块的大小 |\n\n## 获取结果请求参数\n\n| 参数 | 必选 | 类型 | 说明 |\n|------|------|------|------|\n| task_id | 是 | string | 发送提交请求时返回的 task_id |\n\n## 返回结构\n\n### 提交请求返回\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| log_id | uint64 | 唯一的 log id，用于问题定位 |\n| error_code | int | 错误码 |\n| error_msg | string | 错误描述信息 |\n| result | dict | 返回的结果列表 |\n| + task_id | string | 该请求生成的 task_id，后续使用该 task_id 获取结果 |\n\n成功返回示例：\n\n```json\n{\n  \"error_code\": 0,\n  \"error_msg\": \"\",\n  \"log_id\": \"10138598131137362685273585665433\",\n  \"result\": {\n    \"task_id\": \"task-3zy9Bg8CHt1M4pP0cX2q5bg28j268015\"\n  }\n}\n```\n\n### 获取结果返回\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| log_id | uint64 | 唯一的 log id，用于问题定位 |\n| error_code | int | 错误码 |\n| error_msg | string | 错误描述信息 |\n| result | dict | 返回的结果列表 |\n| + task_id | string | 任务 ID |\n| + status | string | 任务状态：pending（排队中）、processing（运行中）、success（成功）、failed（失败） |\n| + task_error | string | 解析报错信息，包含任务失败、额度不够 |\n| + markdown_url | string | 文档解析结果的 Markdown 格式链接，链接有效期 30 天 |\n| + parse_result_url | string | 文档解析结果的 BOS 链接，链接有效期 30 天 |\n\n成功返回示例：\n\n```json\n{\n  \"error_code\": 0,\n  \"error_msg\": \"\",\n  \"result\": {\n    \"task_id\": \"task-UnvGsgbYZp9pS3BZRHn11ifzjNvKzTgf\",\n    \"status\": \"success\",\n    \"task_error\": null,\n    \"duration\": 902.0,\n    \"parse_result_url\": \"https://xxxxxxxxxxxxxxxxxx\"\n  }\n}\n```\n\n### parse_result_url 返回的 JSON 结构\n\n通过 `parse_result_url` 下载解析结果的 JSON 文件，结构如下：\n\n#### 顶层结构\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| file_name | string | 文档名称 |\n| file_id | string | 文档 ID |\n| pages | list | 文件单页解析内容 |\n| chunks | list | 文件内容切分结果（return_doc_chunks 中 switch 为 True 时有值） |\n\n#### 页面对象（pages[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| page_id | string | 页码 ID |\n| page_num | int | 页码数 |\n| text | string | 当前页的所有纯文字内容 |\n| layouts | list | 页面内容版式分析的结果 |\n| tables | list | 页面表格解析结果 |\n| images | list | 页面中图片解析结果 |\n| meta | dict | 页元信息 |\n\n#### 页面元信息（meta）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| page_width | int | 页面宽度 |\n| page_height | int | 页面高度 |\n| is_scan | bool | 是否扫描件 |\n| page_angle | int | 页面倾斜角度 |\n| page_type | string | 页面属性：text（正文）、contents（目录）、appendix（附录）、others（其他） |\n| sheet_name | string | Excel 的 sheet 名 |\n\n#### 版面元素（layouts[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| layout_id | string | layout 元素唯一标志，以 \"xxxxx-layout-{global_layout_index}\" 形式，global_layout_index 为 layout 元素整个文档的全局索引 |\n| text | string | layout 对应的文本内容。注：当 type 为 table、image 时该字段为空，需根据 type 和 layout_id 分别到 tables、images 字段里找到对应的内容 |\n| position | list | layout 元素在页面中的位置，[x, y, w, h] box 框，左上角和宽高 |\n| type | string | 版面元素类型（见下表） |\n| sub_type | string | 版面元素子类型（见下表） |\n| parent | string | 标题层级树中父节点的 layout_id，若当前 layout 为一级标题，其 parent 为 \"root\"。在 table 和 image 的内版面信息中暂时都为空 |\n| children | list | 标题层级树中子节点的 layout_id。在 table 和 image 的内版面信息中暂时都为空 |\n\n**版面类型（type）取值**：\n\n| 类型 | 说明 |\n|------|------|\n| para | 段落 |\n| table | 表格 |\n| image | 文档中的插图 |\n| head_tail | 页面顶部 |\n| contents | 目录 |\n| seal | 印章 |\n| title | 标题 |\n| formula | 公式 |\n\n**子类型（sub_type）取值**：\n\n当 type 为 title 或 image 时，sub_type 有值：\n\n- **title 的 sub_type**：\n  - `title_{n}`：代表 n 级标题，比如 title_2 代表二级标题\n  - `image_title`：图标题\n  - `table_title`：表标题\n\n- **image 的 sub_type**：\n  - `chart`：统计图表\n  - `figure`：普通插图\n  - `QR_code`：二维码\n  - `Bar_code`：条形码\n\n#### 表格对象（tables[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| layout_id | string | layout ID，与 layouts 中的 type 为 table 的元素的 layout ID 对应 |\n| markdown | string | 表格内容的 Markdown 形式 |\n| table_title_id | list | 表格标题对应的 layout_id，默认为 null |\n| position | list | 边框数据 [x, y, w, h]（以页面坐标为原点），版式格式时有效 |\n| cells | list | 单元格的内版面信息，layout 类型为表格时有值 |\n| matrix | list | 二位数组，表示表格内布局位置信息，每个元素对应 cells 列表中元素的索引 |\n| merge_table | string | 跨页表格标记：\"begin\"（开始）、\"inner\"（中间，表格跨页超过两页）、\"end\"（结束）；非跨页表格该字段为空 |\n\n#### 图片对象（images[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| layout_id | string | layout ID，与 layouts 中 type 为 image 的元素的 layout ID 对应 |\n| image_title_id | list | 图片标题对应的 layout_id，默认为 null |\n| position | list | 边框数据 [x, y, w, h] |\n| content_layouts | list | 图片的内版面信息 |\n| data_url | string | 图片存储链接 |\n| image_description | string | 对统计图表进行内容解析和描述，输出结果为 JSON 字符串，可通过 json.loads 结构化为 JSON 格式 |\n\n#### 分块对象（chunks[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| chunk_id | string | 切片的 ID |\n| content | string | 切片的内容 |\n| type | string | 切片类型，为 text 或者 table |\n| meta | dict | chunk 元信息 |\n| + title | list | chunk 所属的多级标题内容 |\n| + position | list | chunk 的位置，根据分块算法有可能 chunk 跨多个页 |\n| + box | list | chunk 的位置坐标 |\n| + page_num | int | chunk 内容所在页数 |\n\n## 文件限制\n\n| 限制项 | 说明 |\n|--------|------|\n| 文件大小（file_data） | ≤ 50MB，超过 50M 须使用 file_url |\n| 文件大小（file_url） | PDF ≤ 300MB，非 PDF ≤ 50MB |\n| URL 长度 | ≤ 1024 字节 |\n| 页数限制 | PDF ≤ 2000 页 |\n| 优先级 | file_data > file_url（同时存在时 file_url 字段失效） |\n\n## QPS 限制\n\n- 提交请求接口：2 QPS\n- 获取结果接口：10 QPS\n\n## 轮询建议\n\n- 提交请求后 5~10 秒开始轮询\n- 轮询间隔：5 秒\n- 最大轮询时间：300 秒\n\n## 相关文档\n\n- [官方 API 文档](https://ai.baidu.com/ai-doc/OCR/llxst5nn0)\n- [错误码参考](error_codes.md)\n- [API Key 配置指南](apikey-fetch.md)\n\nFile v1.0.7:skill-card.md\n\n## Description:\n\nParses PDF, Word, Excel, PowerPoint, image, and other document formats with Baidu's document parsing API to extract text, tables, layout, OCR results, and RAG-oriented chunks.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[maglanyulan](https://clawhub.ai/user/maglanyulan)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers and document-processing agents use this skill to submit documents or document URLs to Baidu's asynchronous parser, poll for results, and retrieve structured text, table, layout, OCR, Markdown, and chunk data for analysis or RAG workflows.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: Documents or document URLs are sent to Baidu for processing.\n\nMitigation: Use this skill only for documents approved for Baidu processing, and avoid confidential, regulated, or internal-only files unless explicitly authorized.\n\nRisk: Baidu API keys are required for authentication.\n\nMitigation: Store API keys outside shared project files and provide them through protected environment variables or approved secret configuration.\n\nRisk: The parser can download result content from a URL returned by the external API.\n\nMitigation: Validate result URLs before downloading and disable automatic result download when URL handling requires additional review.\n\n## Reference(s):\n\n- [ClawHub skill page](https://clawhub.ai/maglanyulan/skills/baidu-doc-pipeline-parser)\n- [Baidu document parsing API documentation](https://ai.baidu.com/ai-doc/OCR/llxst5nn0)\n- [Baidu Intelligent Document Analysis Platform](https://ai.baidu.com/solution/intelligent-document-analysis)\n- [Baidu API key configuration guide](https://ai.baidu.com/ai-doc/REFERENCE/Ck3dwjhhu#1-获取aksk)\n- [API parameters reference](references/parameters.md)\n- [Error codes reference](references/error_codes.md)\n- [API key setup reference](references/apikey-fetch.md)\n\n## Skill Output:\n\n**Output Type(s):** [Text, JSON, Markdown, Shell commands, Configuration, Guidance]\n\n**Output Format:** [JSON responses and optional downloaded parse-result JSON, with Markdown result links returned by the Baidu API.]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [Requires Baidu API credentials and either base64 document data or a public document URL; asynchronous parsing can poll for up to 300 seconds by default.]\n\n## Skill Version(s):\n\n1.0.7 (source: server release metadata)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nArchive v1.0.6: 7 files, 14624 bytes\n\nFiles: _meta.json (144b), references/apikey-fetch.md (1537b), references/error_codes.md (2855b), references/parameters.md (3491b), scripts/baidu_doc_vlm_parser.py (9656b), skill-card.md (2728b), SKILL.md (14311b)\n\nFile v1.0.6:SKILL.md\n\n---\nname: baidu-doc-vlm-parser 百度文档解析(PaddleOCR-VL)\ndescription: 调用百度PaddleOCR-VL大模型API解析文档。基于PaddleOCR-VL-1.6多模态大模型，支持PDF、Word、PPT、图片等格式，精准识别印刷文本、手写文本、表格、公式、图表、印章等复杂元素，支持100+种语言，可处理不规则布局和长文档跨页解析。触发词：文档解析、VLM解析、大模型OCR、PaddleOCR、多模态文档、手写识别、公式识别、复杂版面。\nlicense: MIT\n---\n\n# 百度文档解析（PaddleOCR-VL）Skill\n\n基于 PaddleOCR-VL-1.6 多模态大模型，提供开箱即用的文档智能解析能力。\n\n## 功能概述\n\n**PaddleOCR-VL-1.6-0.9B** 是多模态文档解析领域的 SOTA 方案，具备：\n\n- **全要素精准解析**：高效识别印刷文本、手写文本、表格、公式、图表、印章等复杂文档元素\n- **智能阅读顺序**：基于人类阅读习惯推断内容排列顺序，将零散页面信息转化为有序带标签的结构化元素序列\n- **行级别坐标**：支持精准的行级别坐标输出\n- **100+ 种语言**：覆盖中、英、日、韩、拉丁文等全球化多语种文档\n- **不规则布局定位**：攻克复杂版面解析难点\n- **长文档跨页解析**：支持跨页表格合并等企业级场景\n- **直接 Markdown/JSON 输出**：无需额外处理\n\n## 与文档解析（标准版）的区别\n\n| 特性 | PaddleOCR-VL（本 Skill） | 标准版（pipeline-parser） |\n|------|-------------------------|--------------------------|\n| 底层模型 | 多模态大模型 VLM | 传统 Pipeline |\n| 语言支持 | 100+ 种 | 20+ 种 |\n| 公式/图片识别 | 默认开启，无需配置 | 需手动开启参数 |\n| 语种识别 | 自动识别，无需指定 | 需指定 language_type |\n| 版面类型 | 24 种细粒度类型 | 8 种基础类型 |\n| 行坐标 | 支持 | 不支持 |\n| 多边形坐标 | 支持（polygon） | 仅矩形框 |\n| 文件大小 | 版式 ≤100M，PDF ≤500 页 | PDF ≤300M，≤2000 页 |\n\n## 适用场景\n\n当用户需要：\n- 解析复杂版面文档（多栏、不规则布局）\n- 精准识别手写文本、数学公式、图表\n- 处理多语种混合文档\n- 获取行级别坐标信息\n- 长文档跨页表格合并\n- 免配置自动识别文档内容\n\n## 免费资源领取和计费说明\n[百度智能文档分析平台 领取免费测试资源](https://cloud.baidu.com/doc/OCR/s/fk3h7xu7h)\n\n\n[百度智能文档分析平台计费与购买方式](https://cloud.baidu.com/doc/OCR/s/Fls06fa15#%E6%96%87%E6%A1%A3%E8%A7%A3%E6%9E%90%EF%BC%88paddleocr-vl%EF%BC%89)\n\n\n| 用户类型 | 免费额度 |\n|---------|---------|\n| 个人实名认证用户 | **200 页** |\n| 企业实名认证用户 | **1000 页** |\n\n## API 配置\n\n### 额度获取方式\n\n您可通过百度智能云平台获取[免费额度](https://cloud.baidu.com/doc/OCR/s/fk3h7xu7h)与[购买调用资源](https://cloud.baidu.com/doc/OCR/s/Fls06fa15#%E6%96%87%E6%A1%A3%E8%A7%A3%E6%9E%90%EF%BC%88paddleocr-vl%EF%BC%89)\n\n### 环境变量（必须）\n\n使用前请设置以下环境变量：\n\n```bash\nexport BAIDU_DOC_AI_API_KEY=\"your_api_key\"\nexport BAIDU_DOC_AI_SECRET_KEY=\"your_secret_key\"\n```\n\n### 认证方式\n\n通过 API Key 和 Secret Key 获取 access_token，有效期 30 天。所有接口请求需以 URL 参数 `access_token` 携带。\n\n## 支持格式\n\n**版式文档**：pdf, jpg, jpeg, png, bmp, tif, tiff, ofd（图片最长边不大于 8192px）\n\n**流式文档**：doc, docx, txt, wps, ppt, pptx\n\n## 支持语言\n\n100+ 种语言，包括中文、英文、日文、韩文、拉丁文等，**无需手动指定，大模型自动识别**。\n\n## 使用方式\n\n```bash\npython3 scripts/baidu_doc_vlm_parser.py --file_data <文件的base64编码> --file_name \"test.pdf\"\npython3 scripts/baidu_doc_vlm_parser.py --file_url <文件公网URL> --file_name \"test.pdf\"\n```\n\n## API 接口\n\n文档解析（PaddleOCR-VL）API 服务为异步接口，需要先调用**提交请求接口**获取 task_id，然后调用**获取结果接口**进行结果轮询。\n\n### 提交请求接口\n\n- **HTTP 方法**：POST\n- **请求 URL**：`https://aip.baidubce.com/rest/2.0/brain/online/v2/paddle-vl-parser/task?access_token={token}`\n- **Content-Type**：`application/x-www-form-urlencoded`\n\n#### Python 示例\n\n```python\nimport requests, os, base64\n\ndef create_task(url, file_path, file_url):\n    with open(file_path, \"rb\") as f:\n        file_data = base64.b64encode(f.read())\n    data = {\n        \"file_data\": file_data,\n        \"file_url\": file_url,\n        \"file_name\": os.path.basename(file_path)\n    }\n    headers = {'Content-Type': 'application/x-www-form-urlencoded'}\n    return requests.post(url, headers=headers, data=data)\n\nrequest_host = \"https://aip.baidubce.com/rest/2.0/brain/online/v2/paddle-vl-parser/task?access_token={token}\"\nresponse = create_task(request_host, \"./test.pdf\", \"\")\nprint(response.json())\n```\n\n**成功响应示例：**\n\n```json\n{\n  \"error_code\": 0,\n  \"error_msg\": \"\",\n  \"log_id\": \"10138598131137362685273505665433\",\n  \"result\": { \"task_id\": \"task-3zy9Bg8CHt1M4pPOcX2q5bg28j26801S\" }\n}\n```\n\n### 获取结果接口\n\n- **HTTP 方法**：POST\n- **请求 URL**：`https://aip.baidubce.com/rest/2.0/brain/online/v2/paddle-vl-parser/task/query?access_token={token}`\n- **Content-Type**：`application/x-www-form-urlencoded`\n- **请求参数**：`task_id`（必填，提交请求时返回的 task_id）\n\n#### Python 示例\n\n```python\nimport requests\n\ndef query_task(url, task_id):\n    data = {\"task_id\": task_id}\n    headers = {'Content-Type': 'application/x-www-form-urlencoded'}\n    return requests.post(url, headers=headers, data=data)\n\nrequest_host = \"https://aip.baidubce.com/rest/2.0/brain/online/v2/paddle-vl-parser/task/query?access_token={access_token}\"\nresp = query_task(request_host, \"task_id\")\nprint(resp.json())\n```\n\n## 请求参数\n\n### 文件参数（必选，二选一）\n\n| 参数 | 必选 | 类型 | 说明 |\n|------|------|------|------|\n| `file_data` | 和 file_url 二选一 | string | 文件 Base64 编码数据。版式文档：pdf, jpg, jpeg, png, bmp, tif, tiff, ofd（图片最长边不大于 8192px）；流式文档：doc, docx, txt, wps, ppt, pptx。图片不超过 10M，版式文档不超过 100M，流式文档不超过 50M，PDF 最大 500 页。超过 50M 须使用 file_url。优先级：file_data > file_url |\n| `file_url` | 和 file_data 二选一 | string | 文件数据 URL，长度不超过 1024 字节。PDF 文档不超过 100M，最大 500 页。**请注意关闭 URL 防盗链** |\n| `file_name` | 是 | string | 文件名，请保证文件名后缀正确，例如 \"1.pdf\" |\n\n### 功能参数\n\n| 参数 | 必选 | 类型 | 说明 |\n|------|------|------|------|\n| `recognize_formula` | - | bool | **无需开启**，大模型默认对版式类型文档进行公式识别 |\n| `analysis_chart` | 否 | bool | 是否对统计图表进行解析 |\n| `parse_image_layout` | - | bool | **无需开启**，大模型默认解析文档中的所有图片 |\n| `language_type` | - | string | **无需开启**，大模型默认识别语种类型 |\n| `merge_tables` | 否 | bool | 是否将跨页表格合并输出，开启后 tables 内返回跨页表格合并标识 |\n| `relevel_titles` | 否 | bool | 是否对段落标题（paragraph_title）进行分级，开启后在 sub_type 中输出标题级别 |\n| `recognize_seal` | 否 | bool | 是否识别印章内容 |\n| `return_span_boxes` | 否 | bool | 是否返回行坐标 |\n\n## 返回结构\n\n### 提交请求返回\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `log_id` | uint64 | 唯一的 log id，用于问题定位 |\n| `error_code` | int | 错误码 |\n| `error_msg` | string | 错误描述信息 |\n| `result.task_id` | string | 该请求生成的 task_id |\n\n### 获取结果返回\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `result.task_id` | string | 任务 ID |\n| `result.status` | string | 任务状态：pending（排队中）、processing（运行中）、success（成功）、failed（失败） |\n| `result.task_error` | string | 解析报错信息（任务失败、额度耗尽等） |\n| `result.markdown_url` | string | Markdown 格式结果链接，**有效期 30 天** |\n| `result.parse_result_url` | string | JSON 格式结果 BOS 链接，**有效期 30 天** |\n\n### 解析结果 JSON 结构（parse_result_url）\n\n顶层字段：`file_name`、`file_id`、`pages[]`。\n\n#### 页面对象（pages[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `page_id` | string | 页码 ID |\n| `page_num` | int | 页码数 |\n| `text` | string | 当前页所有纯文字内容 |\n| `layouts` | list | 版式分析结果 |\n| `tables` | list | 表格解析结果 |\n| `images` | list | 图片解析结果 |\n| `meta` | dict | 页面元信息（page_width, page_height） |\n\n#### 版面元素（layouts[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `layout_id` | string | 唯一标志，格式 \"xxxxx-layout-{global_layout_index}\" |\n| `text` | string | 文本内容（type 为 table/image 时为空） |\n| `position` | list | 位置 [x, y, w, h] |\n| `polygon` | list | 顶点坐标列表，可围合成多边形 |\n| `span_boxes` | list | 行信息（开启 return_span_boxes 后生效），含 text 和 location |\n| `type` | string | 版面元素类型（见下表） |\n| `sub_type` | string | 标题层级（开启 relevel_titles 后生效） |\n\n**版面类型（type）— 24 种细粒度类型**：\n\n| 类型 | 说明 | 类型 | 说明 |\n|------|------|------|------|\n| `text` | 文本 | `table` | 表格 |\n| `image` | 图片 | `chart` | 图表 |\n| `doc_title` | 文档标题 | `paragraph_title` | 段落标题 |\n| `figure_title` | 图片标题 | `display_formula` | 公式 |\n| `inline_formula` | 行内公式 | `formula_number` | 公式编号 |\n| `header` | 页眉 | `footer` | 页脚 |\n| `header_image` | 页眉图片 | `footer_image` | 页脚图片 |\n| `number` | 页码 | `abstract` | 摘要 |\n| `algorithm` | 算法 | `aside_text` | 旁注文本 |\n| `content` | 目录 | `footnote` | 脚注 |\n| `reference` | 参考文献 | `reference_content` | 参考文献内容 |\n| `seal` | 印章 | `vertical_text` | 竖排文本 |\n\n#### 表格对象（tables[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `layout_id` | string | 对应 layouts 中 type 为 table 的 layout ID |\n| `markdown` | string | 表格 Markdown 形式 |\n| `position` | list | 边框数据 [x, y, w, h] |\n| `cells` | list | 单元格内版面信息（扁平列表） |\n| `matrix` | list | 2D 单元格索引矩阵，元素为 `cells` 数组下标；相同下标重复出现表示合并单元格 |\n| `merge_table` | string | 合并标识（开启 merge_tables 后）：begin（开始）、end（结束） |\n\n#### 图片对象（images[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `layout_id` | string | 对应 layouts 中 type 为 image 的 layout ID |\n| `position` | list | 边框数据 [x, y, w, h] |\n| `data_url` | string | 图片存储链接 |\n| `image_description` | string | 统计图表内容解析（JSON 字符串，需 `json.loads` 解析） |\n\n## API 特性\n\n### 异步处理流程\n\n1. 调用提交请求接口 → 获取 `task_id`\n2. 通过 `task_id` 调用获取结果接口轮询\n\n### 轮询建议\n\n- 提交请求后 5~10 秒开始轮询\n- 轮询间隔：5 秒\n- 最大轮询时间：300 秒\n\n### QPS 限制\n\n- 提交请求接口：**2 QPS**\n- 获取结果接口：**5 QPS**\n\n## 文件限制\n\n| 限制项 | 说明 |\n|--------|------|\n| 图片大小 | ≤ 10M，最长边 ≤ 8192px |\n| 版式文档大小 | ≤ 100M |\n| 流式文档大小 | ≤ 50M |\n| PDF 页数 | ≤ 500 页 |\n| URL 长度 | ≤ 1024 字节 |\n| 优先级 | file_data > file_url |\n\n## 错误处理\n\n常见错误码示例：\n\n| 错误码 | 说明 |\n|--------|------|\n| `282003` | missing parameters（缺少必要参数） |\n| `282007` | task not exist, please check task id（任务不存在） |\n\n完整错误码参见 `references/error_codes.md` 及 [OCR 错误码文档](https://cloud.baidu.com/doc/OCR/s/dk3iqnq51)。\n\n**失败响应示例：**\n\n```json\n{\n  \"log_id\": \"13665091038742503867108513247608\",\n  \"error_code\": \"282007\",\n  \"error_msg\": \"task not exist, please check task id\",\n  \"result\": \"null\"\n}\n```\n\n## 产品计费与购买方式\n\n> 数据来源：[百度智能云 OCR 计费说明](https://cloud.baidu.com/doc/OCR/s/Fls06fa15#%E6%96%87%E6%A1%A3%E8%A7%A3%E6%9E%90%EF%BC%88paddleocr-vl%EF%BC%89)\n\n### 免费额度\n\n登录[文字识别控制台](https://console.bce.baidu.com/ai/)自动领取：\n\n| 用户类型 | 免费额度 |\n|---------|---------|\n| 个人实名认证用户 | **200 页** |\n| 企业实名认证用户 | **1000 页** |\n\n### 新人优惠\n\n限时活动价 **9 元/千页**，面向新用户。\n\n### 预付费资源包\n\n有效期 **1 年**，到期未抵扣额度作废。购买后 7 天内若未产生调用可自助退订。\n\n| 规格（页） | 原价（元） | 优惠价（元） |\n|---:|---:|---:|\n| 1,000 | 180 | 90 |\n| 5,000 | 850 | 425 |\n| 10,000 | 1,600 | 800 |\n| 50,000 | 7,500 | 3,750 |\n| 100,000 | 14,000 | 7,700 |\n| 200,000 | 26,000 | 14,300 |\n| 500,000 | 55,000 | 33,000 |\n| 1,000,000 | 90,000 | 54,000 |\n| 5,000,000 | 350,000 | 210,000 |\n\n### 按量后付费\n\n- 不限月调用量\n- 原价 **0.18 元/页**，优惠价 **0.09 元/页**\n- 仅成功调用计费，**调用失败不计费**\n\n### 购买入口\n\n- **购买资源包 / 开通后付费**：均通过[百度智能云控制台](https://console.bce.baidu.com/ai/) 文字识别产品页面进行\n- **退订**：资源包购买后 7 天内未产生调用可前往\"退订管理页面\"自助退订\n\n## 脚本\n\n- `scripts/baidu_doc_vlm_parser.py`：文档解析主程序，支持命令行快速调用\n\n## 参考文档\n\n- `references/parameters.md`：完整 API 参数与返回结构详解\n- `references/error_codes.md`：完整错误码参考\n- `references/apikey-fetch.md`：API Key 配置指南\n\n## 相关链接\n\n- [官方技术文档](https://cloud.baidu.com/doc/OCR/s/3mi73at9o)\n- [产品计费与购买](https://cloud.baidu.com/doc/OCR/s/Fls06fa15#%E6%96%87%E6%A1%A3%E8%A7%A3%E6%9E%90%EF%BC%88paddleocr-vl%EF%BC%89)\n- [百度云控制台](https://console.bce.baidu.com/ai/)\n- [智能文档分析平台](https://ai.baidu.com/solution/intelligent-document-analysis)\n\nFile v1.0.6:_meta.json\n\n{\n  \"ownerId\": \"kn75p1w9cr0c8ycct5wzrb5ken83e9nc\",\n  \"slug\": \"baidu-doc-pipeline-parser\",\n  \"version\": \"1.0.6\",\n  \"publishedAt\": 1783499646829\n}\n\nFile v1.0.6:references/apikey-fetch.md\n\n# 百度文档解析（PaddleOCR-VL）API Key 配置指南\n\n## BAIDU_DOC_AI_API_KEY 和 BAIDU_DOC_AI_SECRET_KEY 未配置\n\n当环境变量未设置时，按照以下步骤操作：\n\n### 1. 获取 API Key 和 Secret Key\n\n访问：**https://ai.baidu.com/ai-doc/REFERENCE/Ck3dwjhhu#1-获取aksk**\n\n- 登录百度云账号\n- 创建应用或查看已有的 API Key 和 Secret Key\n- 复制你的 **API Key** 和 **Secret Key**\n\n### 2. 领取免费测试资源\n\n访问：**https://ai.baidu.com/ai-doc/OCR/dk3iqnq51**\n\n### 3. 配置环境变量\n\n```bash\nexport BAIDU_DOC_AI_API_KEY=\"your_actual_api_key_here\"\nexport BAIDU_DOC_AI_SECRET_KEY=\"your_actual_secret_key_here\"\n```\n\n或通过配置文件 `~/.claude/settings.json`：\n\n```json\n{\n  \"skills\": {\n    \"entries\": {\n      \"baidu-doc-vlm-parser\": {\n        \"env\": {\n          \"BAIDU_DOC_AI_API_KEY\": \"your_actual_api_key_here\",\n          \"BAIDU_DOC_AI_SECRET_KEY\": \"your_actual_secret_key_here\"\n        }\n      }\n    }\n  }\n}\n```\n\n### 4. 验证配置\n\n```bash\ncurl -X POST 'https://aip.baidubce.com/oauth/2.0/token' \\\n  -d 'grant_type=client_credentials' \\\n  -d 'client_id={your_api_key}' \\\n  -d 'client_secret={your_secret_key}'\n```\n\n### 5. 测试\n\n```bash\npython3 scripts/baidu_doc_vlm_parser.py --file_url \"https://example.com/test.pdf\" --file_name \"test.pdf\"\n```\n\n## 相关链接\n\n- [获取 AK/SK 文档](https://ai.baidu.com/ai-doc/REFERENCE/Ck3dwjhhu#1-获取aksk)\n- [领取免费资源](https://ai.baidu.com/ai-doc/OCR/dk3iqnq51)\n- [百度云控制台](https://console.bce.baidu.com/ai/)\n\nFile v1.0.6:references/error_codes.md\n\n# 百度文档解析（PaddleOCR-VL）错误码参考\n\n## 通用错误\n\n### 认证相关错误\n\n| 错误码 | 错误信息 | 说明 | 解决方案 |\n|--------|---------|------|----------|\n| 1 | Unknown error | 未知错误 | 重试请求 |\n| 2 | Service temporarily unavailable | 服务暂不可用 | 重试请求 |\n| 3 | Unsupported openapi method | API 接口不存在 | 检查 URL |\n| 4 | Open api request limit reached | 集群超限额 | 重试请求 |\n| 6 | No permission to access data | 无 API 访问权限 | 在百度云控制台开通 |\n| 17 | Open api daily request limit reached | 日配额超限 | 购买额度或等待重置 |\n| 18 | Open api qps request limit reached | QPS 超限 | 降低请求频率 |\n| 19 | Open api total request limit reached | 总量配额超限 | 购买额外配额 |\n| 100 | Invalid parameter | access_token 无效 | 重新获取 |\n| 110 | Access token invalid or no longer valid | access_token 无效 | 重新获取（30天有效期） |\n| 111 | Access token expired | access_token 过期 | 重新获取 |\n\n### 文件相关错误\n\n| 错误码 | 错误信息 | 说明 | 解决方案 |\n|--------|---------|------|----------|\n| 216200 | empty file or fileurl | 文件或 URL 为空 | 提供 file_data 或 file_url |\n| 216201 | file format error | 文件格式不支持 | 使用支持的格式 |\n| 216202 | file size error | 文件大小超限 | 缩减文件大小 |\n\n### 任务处理错误\n\n| 错误码 | 错误信息 | 说明 | 解决方案 |\n|--------|---------|------|----------|\n| 282000 | internal error | 任务处理失败 | 重试或联系技术支持 |\n| 282003 | missing parameters | 缺少必要参数 | 检查必填参数 |\n| 282005 | quota exceed error | 额度不足 | 申请增加配额 |\n| 282007 | task not exist | 任务不存在 | 检查 task_id |\n| 282018 | Service busy | 服务繁忙 | 降低请求频率 |\n\n### URL 相关错误\n\n| 错误码 | 错误信息 | 说明 | 解决方案 |\n|--------|---------|------|----------|\n| 282111 | url format illegal | URL 格式不合法 | 检查 URL 格式 |\n| 282112 | url download timeout | URL 下载超时 | 检查 URL 可访问性 |\n| 282113 | url response invalid | URL 响应无效 | 检查 URL 返回内容 |\n| 282114 | url size error | URL 长度超 1024 字节 | 缩短 URL |\n\n## 错误处理策略\n\n| 错误类型 | 错误码 | 建议处理方式 |\n|---------|--------|-------------|\n| 瞬时错误 | 1, 2, 4, 282000, 282018 | 指数退避重试 |\n| 认证错误 | 100, 110, 111 | 重新获取 access_token |\n| 配额错误 | 17, 18, 19, 282005 | 等待或购买额外配额 |\n| 参数错误 | 216200, 216201, 282003 | 修正参数后重试 |\n\n## 获取帮助\n\n- [提交工单](https://ticket.bce.baidu.com/?_=1648086674827&fromai=1#/ticket/create~productId=96&questionId=1306&channel=2)\n- [百度云控制台](https://console.bce.baidu.com/ai/)\n\nFile v1.0.6:references/parameters.md\n\n# 百度文档解析（PaddleOCR-VL）API 参数详解\n\n## 接口概述\n\n基于 PaddleOCR-VL-1.5-0.9B 多模态大模型，具备全要素精准解析能力，支持 111 种语言，可处理不规则布局和长文档跨页解析。\n\n## API 接口地址\n\n### 提交请求接口\n\n```\nPOST https://aip.baidubce.com/rest/2.0/brain/online/v2/paddle-vl-parser/task?access_token={access_token}\nContent-Type: application/x-www-form-urlencoded\n```\n\n### 获取结果接口\n\n```\nPOST https://aip.baidubce.com/rest/2.0/brain/online/v2/paddle-vl-parser/task/query?access_token={access_token}\nContent-Type: application/x-www-form-urlencoded\n```\n\n## 提交请求参数\n\n### 文件参数（必选，二选一）\n\n| 参数 | 必选 | 类型 | 说明 |\n|------|------|------|------|\n| file_data | 和 file_url 二选一 | string | 文件 Base64 编码。版式文档：pdf, jpg, jpeg, png, bmp, tif, tiff, ofd（图片最长边≤4096px）；流式文档：doc, docx, txt, wps, ppt, pptx。图片≤10M，版式文档≤100M，流式文档≤50M，PDF≤500页。超过50M须用file_url。优先级：file_data > file_url |\n| file_url | 和 file_data 二选一 | string | 文件URL，≤1024字节。PDF≤100M，≤500页。请关闭URL防盗链 |\n| file_name | 是 | string | 文件名，后缀须正确，如 \"1.pdf\" |\n\n### 功能参数\n\n| 参数 | 必选 | 类型 | 说明 |\n|------|------|------|------|\n| recognize_formula | - | bool | **无需开启**，大模型默认识别公式 |\n| analysis_chart | 否 | bool | 是否对统计图表进行解析 |\n| parse_image_layout | - | bool | **无需开启**，大模型默认解析所有图片 |\n| language_type | - | string | **无需开启**，大模型默认识别语种 |\n| merge_tables | 否 | bool | 是否合并跨页表格 |\n| relevel_titles | 否 | bool | 是否对 paragraph_title 进行分级 |\n| recognize_seal | 否 | bool | 是否识别印章 |\n| return_span_boxes | 否 | bool | 是否返回行坐标 |\n\n### 文档分块参数\n\n| 参数 | 必选 | 类型 | 默认值 | 说明 |\n|------|------|------|--------|------|\n| return_doc_chunks.switch | 否 | bool | False | 是否切分 |\n| return_doc_chunks.chunk_size | 否 | int | -1 | 切分块大小，-1为语义自动切分 |\n\n## 返回结构\n\n### 获取结果返回\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| result.task_id | string | 任务 ID |\n| result.status | string | pending/processing/success/failed |\n| result.task_error | string | 报错信息 |\n| result.markdown_url | string | Markdown 结果链接（30天有效） |\n| result.parse_result_url | string | JSON 结果 BOS 链接（30天有效） |\n\n### 版面类型（24种）\n\n| 类型 | 说明 | 类型 | 说明 |\n|------|------|------|------|\n| text | 文本 | table | 表格 |\n| image | 图片 | chart | 图表 |\n| doc_title | 文档标题 | paragraph_title | 段落标题 |\n| figure_title | 图片标题 | display_formula | 公式 |\n| inline_formula | 行内公式 | formula_number | 公式编号 |\n| header | 页眉 | footer | 页脚 |\n| header_image | 页眉图片 | footer_image | 页脚图片 |\n| number | 页码 | abstract | 摘要 |\n| algorithm | 算法 | aside_text | 旁注文本 |\n| content | 目录 | footnote | 脚注 |\n| reference | 参考文献 | reference_content | 参考文献内容 |\n| seal | 印章 | vertical_text | 竖排文本 |\n\n## 相关文档\n\n- [官方 API 文档](https://ai.baidu.com/ai-doc/OCR/3mi73at9o)\n- [错误码参考](error_codes.md)\n- [API Key 配置指南](apikey-fetch.md)\n\nFile v1.0.6:skill-card.md\n\n## Description: <br>\n调用百度 PaddleOCR-VL 大模型 API 解析 PDF、Word、PPT、图片等文档，提取文本、表格、公式、图表、印章和版面结构，并可返回 Markdown 或 JSON 结果。 <br>\n\nThis skill is ready for commercial/non-commercial use. <br>\n\n## Publisher: <br>\n[maglanyulan](https://clawhub.ai/user/maglanyulan) <br>\n\n### License/Terms of Use: <br>\nMIT-0 <br>\n\n\n## Use Case: <br>\nDevelopers and agent users use this skill to submit documents to Baidu document parsing APIs and retrieve structured OCR, layout, table, image, formula, and Markdown/JSON outputs. It is suited for document ingestion and analysis workflows where the user is authorized to send the documents to Baidu. <br>\n\n### Deployment Geography for Use: <br>\nGlobal <br>\n\n## Known Risks and Mitigations: <br>\nRisk: The authoritative security summary reports that the runnable script ignores user file inputs and queries a hard-coded Baidu task using the user's credentials. <br>\nMitigation: Review and fix the script entry point before installation so it submits or queries only the user-selected file or task. <br>\nRisk: Documents are sent to Baidu for processing. <br>\nMitigation: Use only documents the user is allowed to send to Baidu, and avoid regulated or confidential files unless approved. <br>\n\n\n## Reference(s): <br>\n- [ClawHub skill page](https://clawhub.ai/maglanyulan/skills/baidu-doc-pipeline-parser) <br>\n- [Baidu PaddleOCR-VL API documentation](https://cloud.baidu.com/doc/OCR/s/3mi73at9o) <br>\n- [Baidu OCR pricing and purchase information](https://cloud.baidu.com/doc/OCR/s/Fls06fa15#%E6%96%87%E6%A1%A3%E8%A7%A3%E6%9E%90%EF%BC%88paddleocr-vl%EF%BC%89) <br>\n- [Baidu API key setup guide](https://ai.baidu.com/ai-doc/REFERENCE/Ck3dwjhhu#1-获取aksk) <br>\n- [references/parameters.md](references/parameters.md) <br>\n- [references/error_codes.md](references/error_codes.md) <br>\n- [references/apikey-fetch.md](references/apikey-fetch.md) <br>\n\n\n## Skill Output: <br>\n**Output Type(s):** [Text, Markdown, JSON, Shell commands, Configuration guidance] <br>\n**Output Format:** [Markdown guidance with command examples and Baidu API responses containing Markdown and JSON result URLs] <br>\n**Output Parameters:** [1D] <br>\n**Other Properties Related to Output:** [Requires Baidu API credentials; returned result URLs are documented as expiring after 30 days.] <br>\n\n## Skill Version(s): <br>\n1.0.6 (source: server release metadata) <br>\n\n## Ethical Considerations: <br>\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment. <br>\n\nArchive v1.0.5: 7 files, 18132 bytes\n\nFiles: _meta.json (144b), references/apikey-fetch.md (2455b), references/error_codes.md (4611b), references/parameters.md (10795b), scripts/baidu_doc_parser.py (11067b), skill-card.md (2762b), SKILL.md (13091b)\n\nFile v1.0.5:SKILL.md\n\n---\nname: baidu-doc-pipeline-parser 百度文档解析\ndescription: 调用百度文档解析API解析文档。支持PDF、Word、Excel、PPT、图片等18+格式。提取文本、表格、版面分析、OCR识别及RAG文档分块。当用户需要解析文档、提取文本/表格、分析文档结构、处理扫描件时使用。触发词：文档解析、PDF解析、Word解析、表格提取、OCR、文档分析、提取文本、文档结构、扫描识别。\nlicense: MIT\n---\n\n# 百度文档解析 Skill\n\n基于百度智能文档分析平台 API，提供文档解析能力。\n\n## 功能概述\n\n- 支持对 doc、pdf、图片、xlsx 等 18 种格式文档进行解析\n- 输出文档的版面、表格、阅读顺序、标题层级、旋转角度等信息\n- 支持中、英、日、韩、法等 20 余种语言类型\n- 可返回 Markdown 格式内容，将非结构化数据转化为易于处理的结构化数据\n- 识别准确率可达 90% 以上\n- 文档分块（适用于 RAG 场景）\n\n## 适用场景\n\n当用户需要：\n- 解析 PDF、Word、Excel 等格式文档\n- 从文档中提取文本内容\n- 识别并提取表格数据\n- 分析文档结构（标题层级、章节、版面布局）\n- 对扫描件进行 OCR 文字识别\n- 将文档分块用于 RAG 应用\n\n## API 配置\n\n### 环境变量（必须）\n\n[百度智能文档分析平台 领取免费测试资源](https://ai.baidu.com/ai-doc/OCR/dk3iqnq51)\n\n使用前请设置以下环境变量：\n\n```bash\nexport BAIDU_DOC_AI_API_KEY=\"your_api_key\"\nexport BAIDU_DOC_AI_SECRET_KEY=\"your_secret_key\"\n```\n\n### 认证方式\n\n通过 API Key 和 Secret Key 获取 access_token，有效期 30 天。\n\n## 支持格式\n\n**版式文档**：pdf, jpg, jpeg, png, bmp, tif, tiff, ofd, ppt, pptx\n\n**流式文档**：doc, docx, txt, xls, xlsx, wps, html, mhtml\n\n## 支持语言\n\nCHN_ENG（中英文）、JAP（日语）、KOR（韩语）、FRE（法语）、SPA（西班牙语）、POR（葡萄牙语）、GER（德语）、ITA（意大利语）、RUS（俄语）、DAN（丹麦语）、DUT（荷兰语）、MAL（马来语）、SWE（瑞典语）、IND（印尼语）、POL（波兰语）、ROM（罗马尼亚语）、TUR（土耳其语）、GRE（希腊语）、HUN（匈牙利语）、THA（泰语）、VIE（越南语）、ARA（阿拉伯语）、HIN（印地语）\n\n## 使用方式\n\n```bash\npython3 scripts/baidu_doc_parser.py --file_data <文件的base64编码>\npython3 scripts/baidu_doc_parser.py --file_url <文件公网URL>\n```\n\n## API 接口\n\n文档解析 API 服务为异步接口，需要先调用**提交请求接口**获取 task_id，然后调用**获取结果接口**进行结果轮询。\n\n### 提交请求接口\n\n- **HTTP 方法**：POST\n- **请求 URL**：`https://aip.baidubce.com/rest/2.0/brain/online/v2/parser/task?access_token={token}`\n- **Content-Type**：`application/x-www-form-urlencoded`\n\n### 获取结果接口\n\n- **HTTP 方法**：POST\n- **请求 URL**：`https://aip.baidubce.com/rest/2.0/brain/online/v2/parser/task/query?access_token={token}`\n- **Content-Type**：`application/x-www-form-urlencoded`\n- **请求参数**：`task_id`（必填，提交请求时返回的 task_id）\n\n## 请求参数\n\n### 文件参数（必选，二选一）\n\n| 参数 | 必选 | 类型 | 说明 |\n|------|------|------|------|\n| `file_data` | 和 file_url 二选一 | string | 文件 Base64 编码数据。版式文档：pdf, jpg, jpeg, png, bmp, tif, tiff, ofd, ppt, pptx；流式文档：doc, docx, txt, xls, xlsx, wps, html, mhtml。文档大小不超过 50M，PDF 最大支持 2000 页。**若文档大小超过 50M，须从 file_url 方式上传**。优先级：file_data > file_url |\n| `file_url` | 和 file_data 二选一 | string | 文件数据 URL，长度不超过 1024 字节，支持单个 URL 传入。PDF 文档大小不超过 300MB，非 PDF 不超过 50M，PDF 最大支持 2000 页。**请注意关闭 URL 防盗链** |\n| `file_name` | 是 | string | 文件名，请保证文件名后缀正确，例如 \"1.pdf\" |\n\n### 核心功能参数\n\n| 参数 | 必选 | 类型 | 可选值范围 | 说明 |\n|------|------|------|----------|------|\n| `recognize_formula` | 否 | bool | True/False | 是否对版式类型文档进行公式识别 |\n| `analysis_chart` | 否 | bool | True/False | 是否对统计图表进行解析 |\n| `angle_adjust` | 否 | bool | True/False | 是否对图片进行角度矫正 |\n| `parse_image_layout` | 否 | bool | True/False | 是否返回文档中的图片位置信息 |\n\n### 语言与格式参数\n\n| 参数 | 必选 | 类型 | 说明 |\n|------|------|------|------|\n| `language_type` | 否 | string | 识别语种类型，默认为 CHN_ENG（中英文） |\n| `switch_digital_width` | 否 | string | 是否对数字进行全半角转换，默认为 auto。可选：auto（不转换）、half（半角输出）、full（全角输出） |\n| `html_table_format` | 否 | bool | 是否将识别出的表格转换为 HTML 格式返回，**default=True** |\n\n### 文档分块参数\n\n`return_doc_chunks` 为字典类型，用于返回文档切分后的片段数据（按语义、字数、标点）：\n\n| 参数 | 必选 | 类型 | 默认值 | 说明 |\n|------|------|------|--------|------|\n| `switch` | 否 | bool | False | 是否进行文档内容切分 |\n| `split_type` | 否 | str | chunk | 切分方式：chunk（按 chunk_size 来切）/ mark（按 separators 来切） |\n| `separators` | 否 | list | ['。','；','！','？',';','!','?'] | 切分标点 |\n| `chunk_size` | 否 | int | -1 | 切分块的大小，-1 表示按照语义自动切分，不限定块的大小 |\n\n## 返回结构\n\n### 提交请求返回\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `log_id` | uint64 | 唯一的 log id，用于问题定位 |\n| `error_code` | int | 错误码 |\n| `error_msg` | string | 错误描述信息 |\n| `result.task_id` | string | 该请求生成的 task_id，后续使用该 task_id 获取审查结果 |\n\n### 获取结果返回\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `log_id` | uint64 | 唯一的 log id |\n| `error_code` | int | 错误码 |\n| `error_msg` | string | 错误描述信息 |\n| `result.task_id` | string | 任务 ID |\n| `result.status` | string | 任务状态：pending（排队中）、processing（运行中）、success（成功）、failed（失败） |\n| `result.task_error` | string | 解析报错信息，包含任务失败、额度不够 |\n| `result.markdown_url` | string | 文档解析结果的 Markdown 格式链接，**链接有效期 30 天** |\n| `result.parse_result_url` | string | 文档解析结果的 BOS 链接（JSON），**链接有效期 30 天** |\n\n### 解析结果 JSON 结构（parse_result_url）\n\n#### 顶层结构\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `file_name` | string | 文档名称 |\n| `file_id` | string | 文档 ID |\n| `pages` | list | 文件单页解析内容 |\n| `chunks` | list | 文件内容切分结果（return_doc_chunks.switch=True 时有值） |\n\n#### 页面对象（pages[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `page_id` | string | 页码 ID |\n| `page_num` | int | 页码数 |\n| `text` | string | 当前页的所有纯文字内容 |\n| `layouts` | list | 页面内容版式分析的结果 |\n| `tables` | list | 页面表格解析结果 |\n| `images` | list | 页面中图片解析结果 |\n| `meta` | dict | 页元信息 |\n\n#### 页面元信息（meta）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `page_width` | int | 页面宽度 |\n| `page_height` | int | 页面高度 |\n| `is_scan` | bool | 是否扫描件 |\n| `page_angle` | int | 页面倾斜角度 |\n| `page_type` | string | 页面属性：text（正文）、contents（目录）、appendix（附录）、others（其他） |\n| `sheet_name` | string | Excel 的 sheet 名 |\n\n#### 版面元素（layouts[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `layout_id` | string | layout 元素唯一标志，格式 \"xxxxx-layout-{global_layout_index}\" |\n| `text` | string | layout 对应的文本内容。注：当 type 为 table/image 时该字段为空，需根据 type 和 layout_id 分别到 tables/images 字段里找到对应内容 |\n| `position` | list | 元素在页面中的位置 [x, y, w, h]，左上角和宽高 |\n| `type` | string | 版面元素类型（见下表） |\n| `sub_type` | string | 版面元素子类型（见下表） |\n| `parent` | string | 标题层级树中父节点的 layout_id，若为一级标题则 parent 为 \"root\" |\n| `children` | list | 标题层级树中子节点的 layout_id 列表 |\n\n**版面类型（type）**：\n\n| 类型 | 说明 |\n|------|------|\n| `para` | 段落 |\n| `table` | 表格 |\n| `image` | 文档中的插图 |\n| `head_tail` | 页面顶部（页眉/页脚） |\n| `contents` | 目录 |\n| `seal` | 印章 |\n| `title` | 标题 |\n| `formula` | 公式 |\n\n**子类型（sub_type）**：\n\n- **title 类**：`title_{n}`（n 级标题，如 title_2 代表二级标题）、`image_title`（图标题）、`table_title`（表标题）\n- **image 类**：`chart`（统计图表）、`figure`（普通插图）、`QR_code`（二维码）、`Bar_code`（条形码）\n\n#### 表格对象（tables[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `layout_id` | string | 与 layouts 中 type 为 table 的元素的 layout ID 对应 |\n| `markdown` | string | 表格内容的 Markdown 形式 |\n| `table_title_id` | list | 表格标题对应的 layout_id，默认为 null |\n| `position` | list | 边框数据 [x, y, w, h]（以页面坐标为原点），版式格式时有效 |\n| `cells` | list | 单元格的内版面信息，layout 类型为表格时有值 |\n| `matrix` | list | 二位数组，表示表格内布局位置信息，每个元素对应 cells 列表中元素的索引 |\n| `merge_table` | string | 跨页表格标记：begin（开始）、inner（中间，超过两页）、end（结束）；非跨页表格该字段为空 |\n\n#### 图片对象（images[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `layout_id` | string | 与 layouts 中 type 为 image 的元素的 layout ID 对应 |\n| `image_title_id` | list | 图片标题对应的 layout_id，默认为 null |\n| `position` | list | 边框数据 [x, y, w, h] |\n| `content_layouts` | list | 图片的内版面信息 |\n| `data_url` | string | 图片存储链接 |\n| `image_description` | string | 对统计图表进行内容解析和描述，输出结果为 JSON 字符串 |\n\n#### 分块对象（chunks[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `chunk_id` | string | 切片的 ID |\n| `content` | string | 切片的内容 |\n| `type` | string | 切片类型：text 或 table |\n| `meta.title` | list | chunk 所属的多级标题内容 |\n| `meta.position` | list | chunk 的位置，根据分块算法有可能 chunk 跨多个页 |\n| `meta.box` | list | chunk 的位置坐标 |\n| `meta.page_num` | int | chunk 内容所在页数 |\n\n## API 特性\n\n### 异步处理流程\n\n1. 调用提交请求接口 → 获取 `task_id`\n2. 通过 `task_id` 调用获取结果接口轮询\n\n### 轮询建议\n\n- 提交请求后 5~10 秒开始轮询\n- 轮询间隔：5 秒\n- 最大轮询时间：300 秒\n\n### QPS 限制\n\n- 提交请求接口：2 QPS\n- 获取结果接口：10 QPS\n\n## 文件限制\n\n| 限制项 | 说明 |\n|--------|------|\n| 文件大小（file_data） | ≤ 50MB，超过 50M 须使用 file_url |\n| 文件大小（file_url） | PDF ≤ 300MB，非 PDF ≤ 50MB |\n| URL 长度 | ≤ 1024 字节 |\n| 页数限制 | PDF ≤ 2000 页 |\n| 优先级 | file_data > file_url（同时存在时 file_url 字段失效） |\n\n## 错误处理\n\n常见错误码：\n\n| 错误码 | 说明 | 解决方案 |\n|--------|------|----------|\n| 110/111 | access_token 无效或过期 | 重新获取 access_token |\n| 216200 | 文件或 URL 为空 | 提供 file_data 或 file_url |\n| 216201 | 文件格式错误 | 检查文件格式是否支持 |\n| 216202 | 文件大小超限 | 缩减文件大小 |\n| 282000 | 内部错误 | 重试或联系技术支持 |\n| 282003 | 缺少必要参数 | 检查必填参数 |\n| 282007 | 任务不存在 | 检查 task_id 是否正确 |\n| 282018 | 服务繁忙 | 降低请求频率 |\n\n完整错误码参见 `references/error_codes.md`\n\n## 在线调试\n\n可在 [示例代码中心](https://console.bce.baidu.com/tools/#/api?product=AI&project=%E6%96%87%E5%AD%97%E8%AF%86%E5%88%AB&parent=%E6%99%BA%E8%83%BD%E6%96%87%E6%A1%A3%E5%88%86%E6%9E%90%E5%B9%B3%E5%8F%B0&api=rest/2.0/brain/online/v2/parser/task&method=post) 申请试该接口，可进行签名验证、查看在线调用的请求内容和返回结果、示例代码的自动生成。\n\n## 脚本\n\n- `scripts/baidu_doc_parser.py`：文档解析主程序，支持命令行快速调用\n\n## 参考文档\n\n- `references/parameters.md`：完整 API 参数与返回结构详解\n- `references/error_codes.md`：完整错误码参考\n- `references/apikey-fetch.md`：API Key 配置指南\n\n## 相关链接\n\n- [官方 API 文档](https://ai.baidu.com/ai-doc/OCR/llxst5nn0)\n- [百度云控制台](https://console.bce.baidu.com/ai/)\n- [智能文档分析平台](https://ai.baidu.com/solution/intelligent-document-analysis)\n\nFile v1.0.5:_meta.json\n\n{\n  \"ownerId\": \"kn75p1w9cr0c8ycct5wzrb5ken83e9nc\",\n  \"slug\": \"baidu-doc-pipeline-parser\",\n  \"version\": \"1.0.5\",\n  \"publishedAt\": 1778317280416\n}\n\nFile v1.0.5:references/apikey-fetch.md\n\n# 百度文档解析 API Key 配置指南\n\n## BAIDU_DOC_AI_API_KEY 和 BAIDU_DOC_AI_SECRET_KEY 未配置\n\n当环境变量 `BAIDU_DOC_AI_API_KEY` 和 `BAIDU_DOC_AI_SECRET_KEY` 未设置时，按照以下步骤操作：\n\n### 1. 获取 API Key 和 Secret Key\n\n访问：**https://ai.baidu.com/ai-doc/REFERENCE/Ck3dwjhhu#1-获取aksk**\n\n- 登录百度云账号\n- 创建应用或查看已有的 API Key 和 Secret Key\n- 复制你的 **API Key** 和 **Secret Key**\n\n### 2. 领取免费测试资源\n\n访问：**https://ai.baidu.com/ai-doc/OCR/dk3iqnq51**\n\n### 3. 配置环境变量\n\n#### 方式一：直接设置环境变量\n\n```bash\nexport BAIDU_DOC_AI_API_KEY=\"your_actual_api_key_here\"\nexport BAIDU_DOC_AI_SECRET_KEY=\"your_actual_secret_key_here\"\n```\n\n#### 方式二：通过配置文件\n\n编辑配置文件：`~/.claude/settings.json` 或项目 `.claude/settings.json`\n\n添加以下结构：\n\n```json\n{\n  \"skills\": {\n    \"entries\": {\n      \"baidu-doc-pipeline-parser\": {\n        \"env\": {\n          \"BAIDU_DOC_AI_API_KEY\": \"your_actual_api_key_here\",\n          \"BAIDU_DOC_AI_SECRET_KEY\": \"your_actual_secret_key_here\"\n        }\n      }\n    }\n  }\n}\n```\n\n将 `your_actual_api_key_here` 替换为你的实际 API Key，`your_actual_secret_key_here` 替换为你的实际 Secret Key。\n\n### 4. 验证配置\n\n```bash\n# 验证 access_token 是否可正常获取\ncurl -X POST 'https://aip.baidubce.com/oauth/2.0/token' \\\n  -d 'grant_type=client_credentials' \\\n  -d 'client_id={your_api_key}' \\\n  -d 'client_secret={your_secret_key}'\n```\n\n成功返回示例：\n\n```json\n{\n  \"access_token\": \"24.xxxxx.xxxxxx.xxxxxxx-xxxxxxx\",\n  \"expires_in\": 2592000\n}\n```\n\n`expires_in` 为 2592000 秒（30 天），到期后需重新获取。\n\n### 5. 测试\n\n```bash\npython3 scripts/baidu_doc_parser.py --file_url \"https://example.com/test.pdf\" --file_name \"test.pdf\"\npython3 scripts/baidu_doc_parser.py --file_data \"<文件的base64编码>\" --file_name \"test.pdf\"\n```\n\n## 常见问题\n\n- 确保环境变量已正确设置（可通过 `echo $BAIDU_DOC_AI_API_KEY` 验证）\n- 确认 API Key 有效且已开通百度智能文档分析平台服务\n- 检查百度云账户余额或免费额度\n- access_token 有效期 30 天，过期后会自动重新获取\n\n## 相关链接\n\n- [获取 AK/SK 文档](https://ai.baidu.com/ai-doc/REFERENCE/Ck3dwjhhu#1-获取aksk)\n- [领取免费资源](https://ai.baidu.com/ai-doc/OCR/dk3iqnq51)\n- [百度云控制台](https://console.bce.baidu.com/ai/)\n\nFile v1.0.5:references/error_codes.md\n\n# 百度文档解析 API 错误码参考\n\n## 通用错误\n\n### 认证相关错误\n\n| 错误码 | 错误信息 | 说明 | 解决方案 |\n|--------|---------|------|----------|\n| 1 | Unknown error | 未知错误 | 重试请求，持续出现请联系技术支持 |\n| 2 | Service temporarily unavailable | 服务暂不可用 | 重试请求，持续出现请联系技术支持 |\n| 3 | Unsupported openapi method | API 接口不存在 | 检查 URL 是否正确，去除非英文字符 |\n| 4 | Open api request limit reached | 集群超限额 | 重试请求，持续出现请联系技术支持 |\n| 6 | No permission to access data | 无 API 访问权限 | 在百度云控制台开通该 API 权限 |\n| 14 | IAM Certification failed | IAM 认证失败 | 检查签名生成方式或改用 AK/SK |\n| 17 | Open api daily request limit reached | 日配额超限 | 购买额度或等待次日重置 |\n| 18 | Open api qps request limit reached | QPS 超限 | 降低请求频率 |\n| 19 | Open api total request limit reached | 总量配额超限 | 购买额外配额 |\n| 100 | Invalid parameter | access_token 无效 | 重新获取 access_token |\n| 110 | Access token invalid or no longer valid | access_token 无效 | token 有效期 30 天，重新获取 |\n| 111 | Access token expired | access_token 过期 | token 有效期 30 天，重新获取 |\n\n### 文件相关错误\n\n| 错误码 | 错误信息 | 说明 | 解决方案 |\n|--------|---------|------|----------|\n| 216200 | empty file or fileurl | 文件或 URL 为空 | 提供 file_data 或 file_url |\n| 216201 | file format error | 文件格式不支持 | 使用支持的格式（PDF、Word、Excel 等） |\n| 216202 | file size error | 文件大小超限 | 缩减文件大小（file_data ≤ 50MB，file_url PDF ≤ 300MB） |\n\n### 任务处理错误\n\n| 错误码 | 错误信息 | 说明 | 解决方案 |\n|--------|---------|------|----------|\n| 282000 | internal error | 任务处理失败 | 重试或联系技术支持 |\n| 282001 | template not found | 合同类型未找到 | 检查合同类型名称 |\n| 282003 | missing parameters | 缺少必要参数 | 检查必填参数 |\n| 282005 | quota exceed error | 额度不足 | 申请增加配额 |\n| 282006 | check user auth error | 用户权限校验失败 | 验证用户权限 |\n| 282007 | task not exist, please check task id | 任务不存在 | 检查 task_id 是否正确 |\n| 282018 | Service busy | 服务繁忙 | 降低请求频率 |\n\n### URL 相关错误\n\n| 错误码 | 错误信息 | 说明 | 解决方案 |\n|--------|---------|------|----------|\n| 282111 | url format illegal | URL 格式不合法 | 检查 URL 格式 |\n| 282112 | url download timeout | URL 下载超时 | 检查 URL 是否可访问 |\n| 282113 | url response invalid | URL 响应无效 | 检查 URL 返回内容是否正确 |\n| 282114 | url size error | URL 长度超过 1024 字节 | 缩短 URL |\n\n### 参数错误\n\n| 错误码 | 错误信息 | 说明 | 解决方案 |\n|--------|---------|------|----------|\n| 283016 | parameters value error | 参数值无效 | 检查参数格式和取值 |\n\n## 错误响应格式\n\n```json\n{\n  \"log_id\": \"13665091038742503867108513247688\",\n  \"error_code\": \"282007\",\n  \"error_msg\": \"task not exist, please check task id\",\n  \"result\": \"null\"\n}\n```\n\n## 错误处理策略\n\n### 重试策略\n\n| 错误类型 | 错误码 | 建议处理方式 |\n|---------|--------|-------------|\n| 瞬时错误 | 1, 2, 4, 282000, 282018 | 指数退避重试 |\n| 认证错误 | 100, 110, 111 | 重新获取 access_token |\n| 配额错误 | 17, 18, 19, 282005 | 等待或购买额外配额 |\n| 参数错误 | 216200, 216201, 282003, 283016 | 修正参数后重试 |\n| URL 错误 | 282111, 282112, 282113, 282114 | 检查并修正 URL |\n\n### 指数退避重试示例\n\n```python\nimport time\n\ndef retry_with_backoff(func, max_retries=3):\n    for i in range(max_retries):\n        try:\n            return func()\n        except Exception as e:\n            if i == max_retries - 1:\n                raise\n            wait_time = 2 ** i  # 1s, 2s, 4s\n            time.sleep(wait_time)\n```\n\n### Token 刷新示例\n\n```python\ndef ensure_valid_token(client):\n    try:\n        client.query_task(\"test-task-id\")\n    except Exception as e:\n        if \"110\" in str(e) or \"111\" in str(e):\n            client.access_token = client._fetch_auth_credential()\n```\n\n## 获取帮助\n\n- [提交工单](https://ticket.bce.baidu.com/?_=1648086674827&fromai=1#/ticket/create~productId=96&questionId=1306&channel=2)\n- [权限诊断工具](https://console.bce.baidu.com/tools/#/aiInterfacePermissions)\n- [百度云控制台](https://console.bce.baidu.com/ai/)\n\nFile v1.0.5:references/parameters.md\n\n# 百度文档解析 API 参数详解\n\n## 接口概述\n\n百度文档解析 API 支持对 doc、pdf、图片、xlsx 等 18 种格式文档进行解析，输出文档的版面、表格、阅读顺序、标题层级、旋转角度等信息，支持中、英、日、韩、法等 20 余种语言类型，识别准确率可达 90% 以上。\n\n## API 接口地址\n\n### 提交请求接口\n\n```\nPOST https://aip.baidubce.com/rest/2.0/brain/online/v2/parser/task?access_token={access_token}\nContent-Type: application/x-www-form-urlencoded\n```\n\n### 获取结果接口\n\n```\nPOST https://aip.baidubce.com/rest/2.0/brain/online/v2/parser/task/query?access_token={access_token}\nContent-Type: application/x-www-form-urlencoded\n```\n\n## 提交请求参数\n\n### 文件参数（必选，二选一）\n\n| 参数 | 必选 | 类型 | 说明 |\n|------|------|------|------|\n| file_data | 和 file_url 二选一 | string | 文件的 base64 编码数据。版式文档：pdf, jpg, jpeg, png, bmp, tif, tiff, ofd, ppt, pptx；流式文档：doc, docx, txt, xls, xlsx, wps, html, mhtml。文档大小不超过 50M，其中 PDF 文档最大支持 2000 页。若文档大小超过 50M，须从 file_url 方式上传。优先级：file_data > file_url，当 file_data 字段存在时，file_url 字段失效 |\n| file_url | 和 file_data 二选一 | string | 文件数据 URL，URL 长度不超过 1024 字节，支持单个 URL 传入。PDF 文档大小不超过 300MB，非 PDF 文档大小不超过 50M，其中 PDF 文档最大支持 2000 页。优先级：file_data > file_url。**请注意关闭 URL 防盗链** |\n| file_name | 是 | string | 文件名，请保证文件名后缀正确，例如 \"1.pdf\" |\n\n### 核心功能参数\n\n| 参数 | 必选 | 类型 | 可选值范围 | 说明 |\n|------|------|------|----------|------|\n| recognize_formula | 否 | bool | True/False | 是否对版式类型文档进行公式识别 |\n| analysis_chart | 否 | bool | True/False | 是否对统计图表进行解析 |\n| angle_adjust | 否 | bool | True/False | 是否对图片进行角度矫正 |\n| parse_image_layout | 否 | bool | True/False | 是否返回文档中的图片位置信息 |\n\n### 语言与格式参数\n\n| 参数 | 必选 | 类型 | 默认值 | 说明 |\n|------|------|------|--------|------|\n| language_type | 否 | string | CHN_ENG | 识别语种类型 |\n| switch_digital_width | 否 | string | auto | 是否对数字进行全半角转换。auto：不转换；half：半角输出；full：全角输出 |\n| html_table_format | 否 | bool | True | 是否将识别出的表格转换为 HTML 格式返回 |\n\n### 支持语种列表\n\n| 代码 | 语言 | 代码 | 语言 |\n|------|------|------|------|\n| CHN_ENG | 中英文 | DAN | 丹麦语 |\n| JAP | 日语 | DUT | 荷兰语 |\n| KOR | 韩语 | MAL | 马来语 |\n| FRE | 法语 | SWE | 瑞典语 |\n| SPA | 西班牙语 | IND | 印尼语 |\n| POR | 葡萄牙语 | POL | 波兰语 |\n| GER | 德语 | ROM | 罗马尼亚语 |\n| ITA | 意大利语 | TUR | 土耳其语 |\n| RUS | 俄语 | GRE | 希腊语 |\n| HUN | 匈牙利语 | THA | 泰语 |\n| VIE | 越南语 | ARA | 阿拉伯语 |\n| HIN | 印地语 | - | - |\n\n### 文档分块参数\n\n| 参数 | 必选 | 类型 | 默认值 | 说明 |\n|------|------|------|--------|------|\n| return_doc_chunks | 否 | dict | - | 是否返回文档切分后的片段数据（按语义、字数、标点） |\n| + switch | 否 | bool | False | 是否进行文档内容切分 |\n| + split_type | 否 | str | chunk | 切分方式。chunk：按照 chunk_size 来切；mark：按照 separators 来切 |\n| + separators | 否 | list | ['。','；','！','？',';','!','?'] | 切分标点 |\n| + chunk_size | 否 | int | -1 | 切分块的大小，-1 表示按照语义自动切分，不限定块的大小 |\n\n## 获取结果请求参数\n\n| 参数 | 必选 | 类型 | 说明 |\n|------|------|------|------|\n| task_id | 是 | string | 发送提交请求时返回的 task_id |\n\n## 返回结构\n\n### 提交请求返回\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| log_id | uint64 | 唯一的 log id，用于问题定位 |\n| error_code | int | 错误码 |\n| error_msg | string | 错误描述信息 |\n| result | dict | 返回的结果列表 |\n| + task_id | string | 该请求生成的 task_id，后续使用该 task_id 获取结果 |\n\n成功返回示例：\n\n```json\n{\n  \"error_code\": 0,\n  \"error_msg\": \"\",\n  \"log_id\": \"10138598131137362685273585665433\",\n  \"result\": {\n    \"task_id\": \"task-3zy9Bg8CHt1M4pP0cX2q5bg28j268015\"\n  }\n}\n```\n\n### 获取结果返回\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| log_id | uint64 | 唯一的 log id，用于问题定位 |\n| error_code | int | 错误码 |\n| error_msg | string | 错误描述信息 |\n| result | dict | 返回的结果列表 |\n| + task_id | string | 任务 ID |\n| + status | string | 任务状态：pending（排队中）、processing（运行中）、success（成功）、failed（失败） |\n| + task_error | string | 解析报错信息，包含任务失败、额度不够 |\n| + markdown_url | string | 文档解析结果的 Markdown 格式链接，链接有效期 30 天 |\n| + parse_result_url | string | 文档解析结果的 BOS 链接，链接有效期 30 天 |\n\n成功返回示例：\n\n```json\n{\n  \"error_code\": 0,\n  \"error_msg\": \"\",\n  \"result\": {\n    \"task_id\": \"task-UnvGsgbYZp9pS3BZRHn11ifzjNvKzTgf\",\n    \"status\": \"success\",\n    \"task_error\": null,\n    \"duration\": 902.0,\n    \"parse_result_url\": \"https://xxxxxxxxxxxxxxxxxx\"\n  }\n}\n```\n\n### parse_result_url 返回的 JSON 结构\n\n通过 `parse_result_url` 下载解析结果的 JSON 文件，结构如下：\n\n#### 顶层结构\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| file_name | string | 文档名称 |\n| file_id | string | 文档 ID |\n| pages | list | 文件单页解析内容 |\n| chunks | list | 文件内容切分结果（return_doc_chunks 中 switch 为 True 时有值） |\n\n#### 页面对象（pages[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| page_id | string | 页码 ID |\n| page_num | int | 页码数 |\n| text | string | 当前页的所有纯文字内容 |\n| layouts | list | 页面内容版式分析的结果 |\n| tables | list | 页面表格解析结果 |\n| images | list | 页面中图片解析结果 |\n| meta | dict | 页元信息 |\n\n#### 页面元信息（meta）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| page_width | int | 页面宽度 |\n| page_height | int | 页面高度 |\n| is_scan | bool | 是否扫描件 |\n| page_angle | int | 页面倾斜角度 |\n| page_type | string | 页面属性：text（正文）、contents（目录）、appendix（附录）、others（其他） |\n| sheet_name | string | Excel 的 sheet 名 |\n\n#### 版面元素（layouts[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| layout_id | string | layout 元素唯一标志，以 \"xxxxx-layout-{global_layout_index}\" 形式，global_layout_index 为 layout 元素整个文档的全局索引 |\n| text | string | layout 对应的文本内容。注：当 type 为 table、image 时该字段为空，需根据 type 和 layout_id 分别到 tables、images 字段里找到对应的内容 |\n| position | list | layout 元素在页面中的位置，[x, y, w, h] box 框，左上角和宽高 |\n| type | string | 版面元素类型（见下表） |\n| sub_type | string | 版面元素子类型（见下表） |\n| parent | string | 标题层级树中父节点的 layout_id，若当前 layout 为一级标题，其 parent 为 \"root\"。在 table 和 image 的内版面信息中暂时都为空 |\n| children | list | 标题层级树中子节点的 layout_id。在 table 和 image 的内版面信息中暂时都为空 |\n\n**版面类型（type）取值**：\n\n| 类型 | 说明 |\n|------|------|\n| para | 段落 |\n| table | 表格 |\n| image | 文档中的插图 |\n| head_tail | 页面顶部 |\n| contents | 目录 |\n| seal | 印章 |\n| title | 标题 |\n| formula | 公式 |\n\n**子类型（sub_type）取值**：\n\n当 type 为 title 或 image 时，sub_type 有值：\n\n- **title 的 sub_type**：\n  - `title_{n}`：代表 n 级标题，比如 title_2 代表二级标题\n  - `image_title`：图标题\n  - `table_title`：表标题\n\n- **image 的 sub_type**：\n  - `chart`：统计图表\n  - `figure`：普通插图\n  - `QR_code`：二维码\n  - `Bar_code`：条形码\n\n#### 表格对象（tables[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| layout_id | string | layout ID，与 layouts 中的 type 为 table 的元素的 layout ID 对应 |\n| markdown | string | 表格内容的 Markdown 形式 |\n| table_title_id | list | 表格标题对应的 layout_id，默认为 null |\n| position | list | 边框数据 [x, y, w, h]（以页面坐标为原点），版式格式时有效 |\n| cells | list | 单元格的内版面信息，layout 类型为表格时有值 |\n| matrix | list | 二位数组，表示表格内布局位置信息，每个元素对应 cells 列表中元素的索引 |\n| merge_table | string | 跨页表格标记：\"begin\"（开始）、\"inner\"（中间，表格跨页超过两页）、\"end\"（结束）；非跨页表格该字段为空 |\n\n#### 图片对象（images[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| layout_id | string | layout ID，与 layouts 中 type 为 image 的元素的 layout ID 对应 |\n| image_title_id | list | 图片标题对应的 layout_id，默认为 null |\n| position | list | 边框数据 [x, y, w, h] |\n| content_layouts | list | 图片的内版面信息 |\n| data_url | string | 图片存储链接 |\n| image_description | string | 对统计图表进行内容解析和描述，输出结果为 JSON 字符串，可通过 json.loads 结构化为 JSON 格式 |\n\n#### 分块对象（chunks[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| chunk_id | string | 切片的 ID |\n| content | string | 切片的内容 |\n| type | string | 切片类型，为 text 或者 table |\n| meta | dict | chunk 元信息 |\n| + title | list | chunk 所属的多级标题内容 |\n| + position | list | chunk 的位置，根据分块算法有可能 chunk 跨多个页 |\n| + box | list | chunk 的位置坐标 |\n| + page_num | int | chunk 内容所在页数 |\n\n## 文件限制\n\n| 限制项 | 说明 |\n|--------|------|\n| 文件大小（file_data） | ≤ 50MB，超过 50M 须使用 file_url |\n| 文件大小（file_url） | PDF ≤ 300MB，非 PDF ≤ 50MB |\n| URL 长度 | ≤ 1024 字节 |\n| 页数限制 | PDF ≤ 2000 页 |\n| 优先级 | file_data > file_url（同时存在时 file_url 字段失效） |\n\n## QPS 限制\n\n- 提交请求接口：2 QPS\n- 获取结果接口：10 QPS\n\n## 轮询建议\n\n- 提交请求后 5~10 秒开始轮询\n- 轮询间隔：5 秒\n- 最大轮询时间：300 秒\n\n## 相关文档\n\n- [官方 API 文档](https://ai.baidu.com/ai-doc/OCR/llxst5nn0)\n- [错误码参考](error_codes.md)\n- [API Key 配置指南](apikey-fetch.md)\n\nFile v1.0.5:skill-card.md\n\n## Description: <br>\n调用百度文档解析API解析文档，支持PDF、Word、Excel、PPT、图片等18+格式，并提取文本、表格、版面分析、OCR识别及RAG文档分块。 <br>\n\nThis skill is ready for commercial/non-commercial use. <br>\n\n## Publisher: <br>\n[maglanyulan](https://clawhub.ai/user/maglanyulan) <br>\n\n### License/Terms of Use: <br>\nMIT-0 <br>\n\n\n## Use Case: <br>\nExternal developers and document-processing teams use this skill to submit documents to Baidu's document parsing API for OCR, text and table extraction, layout analysis, Markdown conversion, and RAG-oriented chunking. <br>\n\n### Deployment Geography for Use: <br>\nGlobal <br>\n\n## Known Risks and Mitigations: <br>\nRisk: Documents selected for parsing are sent to Baidu's document parsing service. <br>\nMitigation: Use only with documents that policy permits Baidu to process; avoid confidential, regulated, or secret material unless approved. <br>\nRisk: The skill requires Baidu API credentials. <br>\nMitigation: Use a dedicated limited-quota API key and keep BAIDU_DOC_AI_API_KEY and BAIDU_DOC_AI_SECRET_KEY out of source control. <br>\nRisk: Returned markdown_url and parse_result_url links may expose parsed document content during their 30-day lifetime. <br>\nMitigation: Treat result links as private and avoid sharing or logging them in public channels. <br>\n\n\n## Reference(s): <br>\n- [ClawHub skill page](https://clawhub.ai/maglanyulan/baidu-doc-pipeline-parser) <br>\n- [Baidu document parsing API documentation](https://ai.baidu.com/ai-doc/OCR/llxst5nn0) <br>\n- [Baidu Intelligent Document Analysis Platform](https://ai.baidu.com/solution/intelligent-document-analysis) <br>\n- [Baidu API key setup documentation](https://ai.baidu.com/ai-doc/REFERENCE/Ck3dwjhhu#1-获取aksk) <br>\n- [API key configuration guide](references/apikey-fetch.md) <br>\n- [Parameter reference](references/parameters.md) <br>\n- [Error code reference](references/error_codes.md) <br>\n\n\n## Skill Output: <br>\n**Output Type(s):** [API calls, JSON, Markdown, Shell commands, Configuration guidance] <br>\n**Output Format:** [JSON API responses with optional downloaded parse-result JSON and Markdown result links.] <br>\n**Output Parameters:** [1D] <br>\n**Other Properties Related to Output:** [Uses asynchronous task submission and polling; returned markdown_url and parse_result_url links are documented as private and valid for 30 days.] <br>\n\n## Skill Version(s): <br>\n1.0.5 (source: server release evidence) <br>\n\n## Ethical Considerations: <br>\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment. <br>\n\nArchive v1.0.4: 6 files, 16619 bytes\n\nFiles: _meta.json (144b), references/apikey-fetch.md (2455b), references/error_codes.md (4611b), references/parameters.md (10795b), scripts/baidu_doc_parser.py (11067b), SKILL.md (13091b)\n\nFile v1.0.4:SKILL.md\n\n---\nname: baidu-doc-pipeline-parser 百度文档解析\ndescription: 调用百度文档解析API解析文档。支持PDF、Word、Excel、PPT、图片等18+格式。提取文本、表格、版面分析、OCR识别及RAG文档分块。当用户需要解析文档、提取文本/表格、分析文档结构、处理扫描件时使用。触发词：文档解析、PDF解析、Word解析、表格提取、OCR、文档分析、提取文本、文档结构、扫描识别。\nlicense: MIT\n---\n\n# 百度文档解析 Skill\n\n基于百度智能文档分析平台 API，提供文档解析能力。\n\n## 功能概述\n\n- 支持对 doc、pdf、图片、xlsx 等 18 种格式文档进行解析\n- 输出文档的版面、表格、阅读顺序、标题层级、旋转角度等信息\n- 支持中、英、日、韩、法等 20 余种语言类型\n- 可返回 Markdown 格式内容，将非结构化数据转化为易于处理的结构化数据\n- 识别准确率可达 90% 以上\n- 文档分块（适用于 RAG 场景）\n\n## 适用场景\n\n当用户需要：\n- 解析 PDF、Word、Excel 等格式文档\n- 从文档中提取文本内容\n- 识别并提取表格数据\n- 分析文档结构（标题层级、章节、版面布局）\n- 对扫描件进行 OCR 文字识别\n- 将文档分块用于 RAG 应用\n\n## API 配置\n\n### 环境变量（必须）\n\n[百度智能文档分析平台 领取免费测试资源](https://ai.baidu.com/ai-doc/OCR/dk3iqnq51)\n\n使用前请设置以下环境变量：\n\n```bash\nexport BAIDU_DOC_AI_API_KEY=\"your_api_key\"\nexport BAIDU_DOC_AI_SECRET_KEY=\"your_secret_key\"\n```\n\n### 认证方式\n\n通过 API Key 和 Secret Key 获取 access_token，有效期 30 天。\n\n## 支持格式\n\n**版式文档**：pdf, jpg, jpeg, png, bmp, tif, tiff, ofd, ppt, pptx\n\n**流式文档**：doc, docx, txt, xls, xlsx, wps, html, mhtml\n\n## 支持语言\n\nCHN_ENG（中英文）、JAP（日语）、KOR（韩语）、FRE（法语）、SPA（西班牙语）、POR（葡萄牙语）、GER（德语）、ITA（意大利语）、RUS（俄语）、DAN（丹麦语）、DUT（荷兰语）、MAL（马来语）、SWE（瑞典语）、IND（印尼语）、POL（波兰语）、ROM（罗马尼亚语）、TUR（土耳其语）、GRE（希腊语）、HUN（匈牙利语）、THA（泰语）、VIE（越南语）、ARA（阿拉伯语）、HIN（印地语）\n\n## 使用方式\n\n```bash\npython3 scripts/baidu_doc_parser.py --file_data <文件的base64编码>\npython3 scripts/baidu_doc_parser.py --file_url <文件公网URL>\n```\n\n## API 接口\n\n文档解析 API 服务为异步接口，需要先调用**提交请求接口**获取 task_id，然后调用**获取结果接口**进行结果轮询。\n\n### 提交请求接口\n\n- **HTTP 方法**：POST\n- **请求 URL**：`https://aip.baidubce.com/rest/2.0/brain/online/v2/parser/task?access_token={token}`\n- **Content-Type**：`application/x-www-form-urlencoded`\n\n### 获取结果接口\n\n- **HTTP 方法**：POST\n- **请求 URL**：`https://aip.baidubce.com/rest/2.0/brain/online/v2/parser/task/query?access_token={token}`\n- **Content-Type**：`application/x-www-form-urlencoded`\n- **请求参数**：`task_id`（必填，提交请求时返回的 task_id）\n\n## 请求参数\n\n### 文件参数（必选，二选一）\n\n| 参数 | 必选 | 类型 | 说明 |\n|------|------|------|------|\n| `file_data` | 和 file_url 二选一 | string | 文件 Base64 编码数据。版式文档：pdf, jpg, jpeg, png, bmp, tif, tiff, ofd, ppt, pptx；流式文档：doc, docx, txt, xls, xlsx, wps, html, mhtml。文档大小不超过 50M，PDF 最大支持 2000 页。**若文档大小超过 50M，须从 file_url 方式上传**。优先级：file_data > file_url |\n| `file_url` | 和 file_data 二选一 | string | 文件数据 URL，长度不超过 1024 字节，支持单个 URL 传入。PDF 文档大小不超过 300MB，非 PDF 不超过 50M，PDF 最大支持 2000 页。**请注意关闭 URL 防盗链** |\n| `file_name` | 是 | string | 文件名，请保证文件名后缀正确，例如 \"1.pdf\" |\n\n### 核心功能参数\n\n| 参数 | 必选 | 类型 | 可选值范围 | 说明 |\n|------|------|------|----------|------|\n| `recognize_formula` | 否 | bool | True/False | 是否对版式类型文档进行公式识别 |\n| `analysis_chart` | 否 | bool | True/False | 是否对统计图表进行解析 |\n| `angle_adjust` | 否 | bool | True/False | 是否对图片进行角度矫正 |\n| `parse_image_layout` | 否 | bool | True/False | 是否返回文档中的图片位置信息 |\n\n### 语言与格式参数\n\n| 参数 | 必选 | 类型 | 说明 |\n|------|------|------|------|\n| `language_type` | 否 | string | 识别语种类型，默认为 CHN_ENG（中英文） |\n| `switch_digital_width` | 否 | string | 是否对数字进行全半角转换，默认为 auto。可选：auto（不转换）、half（半角输出）、full（全角输出） |\n| `html_table_format` | 否 | bool | 是否将识别出的表格转换为 HTML 格式返回，**default=True** |\n\n### 文档分块参数\n\n`return_doc_chunks` 为字典类型，用于返回文档切分后的片段数据（按语义、字数、标点）：\n\n| 参数 | 必选 | 类型 | 默认值 | 说明 |\n|------|------|------|--------|------|\n| `switch` | 否 | bool | False | 是否进行文档内容切分 |\n| `split_type` | 否 | str | chunk | 切分方式：chunk（按 chunk_size 来切）/ mark（按 separators 来切） |\n| `separators` | 否 | list | ['。','；','！','？',';','!','?'] | 切分标点 |\n| `chunk_size` | 否 | int | -1 | 切分块的大小，-1 表示按照语义自动切分，不限定块的大小 |\n\n## 返回结构\n\n### 提交请求返回\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `log_id` | uint64 | 唯一的 log id，用于问题定位 |\n| `error_code` | int | 错误码 |\n| `error_msg` | string | 错误描述信息 |\n| `result.task_id` | string | 该请求生成的 task_id，后续使用该 task_id 获取审查结果 |\n\n### 获取结果返回\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `log_id` | uint64 | 唯一的 log id |\n| `error_code` | int | 错误码 |\n| `error_msg` | string | 错误描述信息 |\n| `result.task_id` | string | 任务 ID |\n| `result.status` | string | 任务状态：pending（排队中）、processing（运行中）、success（成功）、failed（失败） |\n| `result.task_error` | string | 解析报错信息，包含任务失败、额度不够 |\n| `result.markdown_url` | string | 文档解析结果的 Markdown 格式链接，**链接有效期 30 天** |\n| `result.parse_result_url` | string | 文档解析结果的 BOS 链接（JSON），**链接有效期 30 天** |\n\n### 解析结果 JSON 结构（parse_result_url）\n\n#### 顶层结构\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `file_name` | string | 文档名称 |\n| `file_id` | string | 文档 ID |\n| `pages` | list | 文件单页解析内容 |\n| `chunks` | list | 文件内容切分结果（return_doc_chunks.switch=True 时有值） |\n\n#### 页面对象（pages[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `page_id` | string | 页码 ID |\n| `page_num` | int | 页码数 |\n| `text` | string | 当前页的所有纯文字内容 |\n| `layouts` | list | 页面内容版式分析的结果 |\n| `tables` | list | 页面表格解析结果 |\n| `images` | list | 页面中图片解析结果 |\n| `meta` | dict | 页元信息 |\n\n#### 页面元信息（meta）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `page_width` | int | 页面宽度 |\n| `page_height` | int | 页面高度 |\n| `is_scan` | bool | 是否扫描件 |\n| `page_angle` | int | 页面倾斜角度 |\n| `page_type` | string | 页面属性：text（正文）、contents（目录）、appendix（附录）、others（其他） |\n| `sheet_name` | string | Excel 的 sheet 名 |\n\n#### 版面元素（layouts[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `layout_id` | string | layout 元素唯一标志，格式 \"xxxxx-layout-{global_layout_index}\" |\n| `text` | string | layout 对应的文本内容。注：当 type 为 table/image 时该字段为空，需根据 type 和 layout_id 分别到 tables/images 字段里找到对应内容 |\n| `position` | list | 元素在页面中的位置 [x, y, w, h]，左上角和宽高 |\n| `type` | string | 版面元素类型（见下表） |\n| `sub_type` | string | 版面元素子类型（见下表） |\n| `parent` | string | 标题层级树中父节点的 layout_id，若为一级标题则 parent 为 \"root\" |\n| `children` | list | 标题层级树中子节点的 layout_id 列表 |\n\n**版面类型（type）**：\n\n| 类型 | 说明 |\n|------|------|\n| `para` | 段落 |\n| `table` | 表格 |\n| `image` | 文档中的插图 |\n| `head_tail` | 页面顶部（页眉/页脚） |\n| `contents` | 目录 |\n| `seal` | 印章 |\n| `title` | 标题 |\n| `formula` | 公式 |\n\n**子类型（sub_type）**：\n\n- **title 类**：`title_{n}`（n 级标题，如 title_2 代表二级标题）、`image_title`（图标题）、`table_title`（表标题）\n- **image 类**：`chart`（统计图表）、`figure`（普通插图）、`QR_code`（二维码）、`Bar_code`（条形码）\n\n#### 表格对象（tables[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `layout_id` | string | 与 layouts 中 type 为 table 的元素的 layout ID 对应 |\n| `markdown` | string | 表格内容的 Markdown 形式 |\n| `table_title_id` | list | 表格标题对应的 layout_id，默认为 null |\n| `position` | list | 边框数据 [x, y, w, h]（以页面坐标为原点），版式格式时有效 |\n| `cells` | list | 单元格的内版面信息，layout 类型为表格时有值 |\n| `matrix` | list | 二位数组，表示表格内布局位置信息，每个元素对应 cells 列表中元素的索引 |\n| `merge_table` | string | 跨页表格标记：begin（开始）、inner（中间，超过两页）、end（结束）；非跨页表格该字段为空 |\n\n#### 图片对象（images[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `layout_id` | string | 与 layouts 中 type 为 image 的元素的 layout ID 对应 |\n| `image_title_id` | list | 图片标题对应的 layout_id，默认为 null |\n| `position` | list | 边框数据 [x, y, w, h] |\n| `content_layouts` | list | 图片的内版面信息 |\n| `data_url` | string | 图片存储链接 |\n| `image_description` | string | 对统计图表进行内容解析和描述，输出结果为 JSON 字符串 |\n\n#### 分块对象（chunks[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `chunk_id` | string | 切片的 ID |\n| `content` | string | 切片的内容 |\n| `type` | string | 切片类型：text 或 table |\n| `meta.title` | list | chunk 所属的多级标题内容 |\n| `meta.position` | list | chunk 的位置，根据分块算法有可能 chunk 跨多个页 |\n| `meta.box` | list | chunk 的位置坐标 |\n| `meta.page_num` | int | chunk 内容所在页数 |\n\n## API 特性\n\n### 异步处理流程\n\n1. 调用提交请求接口 → 获取 `task_id`\n2. 通过 `task_id` 调用获取结果接口轮询\n\n### 轮询建议\n\n- 提交请求后 5~10 秒开始轮询\n- 轮询间隔：5 秒\n- 最大轮询时间：300 秒\n\n### QPS 限制\n\n- 提交请求接口：2 QPS\n- 获取结果接口：10 QPS\n\n## 文件限制\n\n| 限制项 | 说明 |\n|--------|------|\n| 文件大小（file_data） | ≤ 50MB，超过 50M 须使用 file_url |\n| 文件大小（file_url） | PDF ≤ 300MB，非 PDF ≤ 50MB |\n| URL 长度 | ≤ 1024 字节 |\n| 页数限制 | PDF ≤ 2000 页 |\n| 优先级 | file_data > file_url（同时存在时 file_url 字段失效） |\n\n## 错误处理\n\n常见错误码：\n\n| 错误码 | 说明 | 解决方案 |\n|--------|------|----------|\n| 110/111 | access_token 无效或过期 | 重新获取 access_token |\n| 216200 | 文件或 URL 为空 | 提供 file_data 或 file_url |\n| 216201 | 文件格式错误 | 检查文件格式是否支持 |\n| 216202 | 文件大小超限 | 缩减文件大小 |\n| 282000 | 内部错误 | 重试或联系技术支持 |\n| 282003 | 缺少必要参数 | 检查必填参数 |\n| 282007 | 任务不存在 | 检查 task_id 是否正确 |\n| 282018 | 服务繁忙 | 降低请求频率 |\n\n完整错误码参见 `references/error_codes.md`\n\n## 在线调试\n\n可在 [示例代码中心](https://console.bce.baidu.com/tools/#/api?product=AI&project=%E6%96%87%E5%AD%97%E8%AF%86%E5%88%AB&parent=%E6%99%BA%E8%83%BD%E6%96%87%E6%A1%A3%E5%88%86%E6%9E%90%E5%B9%B3%E5%8F%B0&api=rest/2.0/brain/online/v2/parser/task&method=post) 申请试该接口，可进行签名验证、查看在线调用的请求内容和返回结果、示例代码的自动生成。\n\n## 脚本\n\n- `scripts/baidu_doc_parser.py`：文档解析主程序，支持命令行快速调用\n\n## 参考文档\n\n- `references/parameters.md`：完整 API 参数与返回结构详解\n- `references/error_codes.md`：完整错误码参考\n- `references/apikey-fetch.md`：API Key 配置指南\n\n## 相关链接\n\n- [官方 API 文档](https://ai.baidu.com/ai-doc/OCR/llxst5nn0)\n- [百度云控制台](https://console.bce.baidu.com/ai/)\n- [智能文档分析平台](https://ai.baidu.com/solution/intelligent-document-analysis)\n\nFile v1.0.4:_meta.json\n\n{\n  \"ownerId\": \"kn75p1w9cr0c8ycct5wzrb5ken83e9nc\",\n  \"slug\": \"baidu-doc-pipeline-parser\",\n  \"version\": \"1.0.4\",\n  \"publishedAt\": 1778316580814\n}\n\nFile v1.0.4:references/apikey-fetch.md\n\n# 百度文档解析 API Key 配置指南\n\n## BAIDU_DOC_AI_API_KEY 和 BAIDU_DOC_AI_SECRET_KEY 未配置\n\n当环境变量 `BAIDU_DOC_AI_API_KEY` 和 `BAIDU_DOC_AI_SECRET_KEY` 未设置时，按照以下步骤操作：\n\n### 1. 获取 API Key 和 Secret Key\n\n访问：**https://ai.baidu.com/ai-doc/REFERENCE/Ck3dwjhhu#1-获取aksk**\n\n- 登录百度云账号\n- 创建应用或查看已有的 API Key 和 Secret Key\n- 复制你的 **API Key** 和 **Secret Key**\n\n### 2. 领取免费测试资源\n\n访问：**https://ai.baidu.com/ai-doc/OCR/dk3iqnq51**\n\n### 3. 配置环境变量\n\n#### 方式一：直接设置环境变量\n\n```bash\nexport BAIDU_DOC_AI_API_KEY=\"your_actual_api_key_here\"\nexport BAIDU_DOC_AI_SECRET_KEY=\"your_actual_secret_key_here\"\n```\n\n#### 方式二：通过配置文件\n\n编辑配置文件：`~/.claude/settings.json` 或项目 `.claude/settings.json`\n\n添加以下结构：\n\n```json\n{\n  \"skills\": {\n    \"entries\": {\n      \"baidu-doc-pipeline-parser\": {\n        \"env\": {\n          \"BAIDU_DOC_AI_API_KEY\": \"your_actual_api_key_here\",\n          \"BAIDU_DOC_AI_SECRET_KEY\": \"your_actual_secret_key_here\"\n        }\n      }\n    }\n  }\n}\n```\n\n将 `your_actual_api_key_here` 替换为你的实际 API Key，`your_actual_secret_key_here` 替换为你的实际 Secret Key。\n\n### 4. 验证配置\n\n```bash\n# 验证 access_token 是否可正常获取\ncurl -X POST 'https://aip.baidubce.com/oauth/2.0/token' \\\n  -d 'grant_type=client_credentials' \\\n  -d 'client_id={your_api_key}' \\\n  -d 'client_secret={your_secret_key}'\n```\n\n成功返回示例：\n\n```json\n{\n  \"access_token\": \"24.xxxxx.xxxxxx.xxxxxxx-xxxxxxx\",\n  \"expires_in\": 2592000\n}\n```\n\n`expires_in` 为 2592000 秒（30 天），到期后需重新获取。\n\n### 5. 测试\n\n```bash\npython3 scripts/baidu_doc_parser.py --file_url \"https://example.com/test.pdf\" --file_name \"test.pdf\"\npython3 scripts/baidu_doc_parser.py --file_data \"<文件的base64编码>\" --file_name \"test.pdf\"\n```\n\n## 常见问题\n\n- 确保环境变量已正确设置（可通过 `echo $BAIDU_DOC_AI_API_KEY` 验证）\n- 确认 API Key 有效且已开通百度智能文档分析平台服务\n- 检查百度云账户余额或免费额度\n- access_token 有效期 30 天，过期后会自动重新获取\n\n## 相关链接\n\n- [获取 AK/SK 文档](https://ai.baidu.com/ai-doc/REFERENCE/Ck3dwjhhu#1-获取aksk)\n- [领取免费资源](https://ai.baidu.com/ai-doc/OCR/dk3iqnq51)\n- [百度云控制台](https://console.bce.baidu.com/ai/)\n\nFile v1.0.4:references/error_codes.md\n\n# 百度文档解析 API 错误码参考\n\n## 通用错误\n\n### 认证相关错误\n\n| 错误码 | 错误信息 | 说明 | 解决方案 |\n|--------|---------|------|----------|\n| 1 | Unknown error | 未知错误 | 重试请求，持续出现请联系技术支持 |\n| 2 | Service temporarily unavailable | 服务暂不可用 | 重试请求，持续出现请联系技术支持 |\n| 3 | Unsupported openapi method | API 接口不存在 | 检查 URL 是否正确，去除非英文字符 |\n| 4 | Open api request limit reached | 集群超限额 | 重试请求，持续出现请联系技术支持 |\n| 6 | No permission to access data | 无 API 访问权限 | 在百度云控制台开通该 API 权限 |\n| 14 | IAM Certification failed | IAM 认证失败 | 检查签名生成方式或改用 AK/SK |\n| 17 | Open api daily request limit reached | 日配额超限 | 购买额度或等待次日重置 |\n| 18 | Open api qps request limit reached | QPS 超限 | 降低请求频率 |\n| 19 | Open api total request limit reached | 总量配额超限 | 购买额外配额 |\n| 100 | Invalid parameter | access_token 无效 | 重新获取 access_token |\n| 110 | Access token invalid or no longer valid | access_token 无效 | token 有效期 30 天，重新获取 |\n| 111 | Access token expired | access_token 过期 | token 有效期 30 天，重新获取 |\n\n### 文件相关错误\n\n| 错误码 | 错误信息 | 说明 | 解决方案 |\n|--------|---------|------|----------|\n| 216200 | empty file or fileurl | 文件或 URL 为空 | 提供 file_data 或 file_url |\n| 216201 | file format error | 文件格式不支持 | 使用支持的格式（PDF、Word、Excel 等） |\n| 216202 | file size error | 文件大小超限 | 缩减文件大小（file_data ≤ 50MB，file_url PDF ≤ 300MB） |\n\n### 任务处理错误\n\n| 错误码 | 错误信息 | 说明 | 解决方案 |\n|--------|---------|------|----------|\n| 282000 | internal error | 任务处理失败 | 重试或联系技术支持 |\n| 282001 | template not found | 合同类型未找到 | 检查合同类型名称 |\n| 282003 | missing parameters | 缺少必要参数 | 检查必填参数 |\n| 282005 | quota exceed error | 额度不足 | 申请增加配额 |\n| 282006 | check user auth error | 用户权限校验失败 | 验证用户权限 |\n| 282007 | task not exist, please check task id | 任务不存在 | 检查 task_id 是否正确 |\n| 282018 | Service busy | 服务繁忙 | 降低请求频率 |\n\n### URL 相关错误\n\n| 错误码 | 错误信息 | 说明 | 解决方案 |\n|--------|---------|------|----------|\n| 282111 | url format illegal | URL 格式不合法 | 检查 URL 格式 |\n| 282112 | url download timeout | URL 下载超时 | 检查 URL 是否可访问 |\n| 282113 | url response invalid | URL 响应无效 | 检查 URL 返回内容是否正确 |\n| 282114 | url size error | URL 长度超过 1024 字节 | 缩短 URL |\n\n### 参数错误\n\n| 错误码 | 错误信息 | 说明 | 解决方案 |\n|--------|---------|------|----------|\n| 283016 | parameters value error | 参数值无效 | 检查参数格式和取值 |\n\n## 错误响应格式\n\n```json\n{\n  \"log_id\": \"13665091038742503867108513247688\",\n  \"error_code\": \"282007\",\n  \"error_msg\": \"task not exist, please check task id\",\n  \"result\": \"null\"\n}\n```\n\n## 错误处理策略\n\n### 重试策略\n\n| 错误类型 | 错误码 | 建议处理方式 |\n|---------|--------|-------------|\n| 瞬时错误 | 1, 2, 4, 282000, 282018 | 指数退避重试 |\n| 认证错误 | 100, 110, 111 | 重新获取 access_token |\n| 配额错误 | 17, 18, 19, 282005 | 等待或购买额外配额 |\n| 参数错误 | 216200, 216201, 282003, 283016 | 修正参数后重试 |\n| URL 错误 | 282111, 282112, 282113, 282114 | 检查并修正 URL |\n\n### 指数退避重试示例\n\n```python\nimport time\n\ndef retry_with_backoff(func, max_retries=3):\n    for i in range(max_retries):\n        try:\n            return func()\n        except Exception as e:\n            if i == max_retries - 1:\n                raise\n            wait_time = 2 ** i  # 1s, 2s, 4s\n            time.sleep(wait_time)\n```\n\n### Token 刷新示例\n\n```python\ndef ensure_valid_token(client):\n    try:\n        client.query_task(\"test-task-id\")\n    except Exception as e:\n        if \"110\" in str(e) or \"111\" in str(e):\n            client.access_token = client._fetch_auth_credential()\n```\n\n## 获取帮助\n\n- [提交工单](https://ticket.bce.baidu.com/?_=1648086674827&fromai=1#/ticket/create~productId=96&questionId=1306&channel=2)\n- [权限诊断工具](https://console.bce.baidu.com/tools/#/aiInterfacePermissions)\n- [百度云控制台](https://console.bce.baidu.com/ai/)\n\nFile v1.0.4:references/parameters.md\n\n# 百度文档解析 API 参数详解\n\n## 接口概述\n\n百度文档解析 API 支持对 doc、pdf、图片、xlsx 等 18 种格式文档进行解析，输出文档的版面、表格、阅读顺序、标题层级、旋转角度等信息，支持中、英、日、韩、法等 20 余种语言类型，识别准确率可达 90% 以上。\n\n## API 接口地址\n\n### 提交请求接口\n\n```\nPOST https://aip.baidubce.com/rest/2.0/brain/online/v2/parser/task?access_token={access_token}\nContent-Type: application/x-www-form-urlencoded\n```\n\n### 获取结果接口\n\n```\nPOST https://aip.baidubce.com/rest/2.0/brain/online/v2/parser/task/query?access_token={access_token}\nContent-Type: application/x-www-form-urlencoded\n```\n\n## 提交请求参数\n\n### 文件参数（必选，二选一）\n\n| 参数 | 必选 | 类型 | 说明 |\n|------|------|------|------|\n| file_data | 和 file_url 二选一 | string | 文件的 base64 编码数据。版式文档：pdf, jpg, jpeg, png, bmp, tif, tiff, ofd, ppt, pptx；流式文档：doc, docx, txt, xls, xlsx, wps, html, mhtml。文档大小不超过 50M，其中 PDF 文档最大支持 2000 页。若文档大小超过 50M，须从 file_url 方式上传。优先级：file_data > file_url，当 file_data 字段存在时，file_url 字段失效 |\n| file_url | 和 file_data 二选一 | string | 文件数据 URL，URL 长度不超过 1024 字节，支持单个 URL 传入。PDF 文档大小不超过 300MB，非 PDF 文档大小不超过 50M，其中 PDF 文档最大支持 2000 页。优先级：file_data > file_url。**请注意关闭 URL 防盗链** |\n| file_name | 是 | string | 文件名，请保证文件名后缀正确，例如 \"1.pdf\" |\n\n### 核心功能参数\n\n| 参数 | 必选 | 类型 | 可选值范围 | 说明 |\n|------|------|------|----------|------|\n| recognize_formula | 否 | bool | True/False | 是否对版式类型文档进行公式识别 |\n| analysis_chart | 否 | bool | True/False | 是否对统计图表进行解析 |\n| angle_adjust | 否 | bool | True/False | 是否对图片进行角度矫正 |\n| parse_image_layout | 否 | bool | True/False | 是否返回文档中的图片位置信息 |\n\n### 语言与格式参数\n\n| 参数 | 必选 | 类型 | 默认值 | 说明 |\n|------|------|------|--------|------|\n| language_type | 否 | string | CHN_ENG | 识别语种类型 |\n| switch_digital_width | 否 | string | auto | 是否对数字进行全半角转换。auto：不转换；half：半角输出；full：全角输出 |\n| html_table_format | 否 | bool | True | 是否将识别出的表格转换为 HTML 格式返回 |\n\n### 支持语种列表\n\n| 代码 | 语言 | 代码 | 语言 |\n|------|------|------|------|\n| CHN_ENG | 中英文 | DAN | 丹麦语 |\n| JAP | 日语 | DUT | 荷兰语 |\n| KOR | 韩语 | MAL | 马来语 |\n| FRE | 法语 | SWE | 瑞典语 |\n| SPA | 西班牙语 | IND | 印尼语 |\n| POR | 葡萄牙语 | POL | 波兰语 |\n| GER | 德语 | ROM | 罗马尼亚语 |\n| ITA | 意大利语 | TUR | 土耳其语 |\n| RUS | 俄语 | GRE | 希腊语 |\n| HUN | 匈牙利语 | THA | 泰语 |\n| VIE | 越南语 | ARA | 阿拉伯语 |\n| HIN | 印地语 | - | - |\n\n### 文档分块参数\n\n| 参数 | 必选 | 类型 | 默认值 | 说明 |\n|------|------|------|--------|------|\n| return_doc_chunks | 否 | dict | - | 是否返回文档切分后的片段数据（按语义、字数、标点） |\n| + switch | 否 | bool | False | 是否进行文档内容切分 |\n| + split_type | 否 | str | chunk | 切分方式。chunk：按照 chunk_size 来切；mark：按照 separators 来切 |\n| + separators | 否 | list | ['。','；','！','？',';','!','?'] | 切分标点 |\n| + chunk_size | 否 | int | -1 | 切分块的大小，-1 表示按照语义自动切分，不限定块的大小 |\n\n## 获取结果请求参数\n\n| 参数 | 必选 | 类型 | 说明 |\n|------|------|------|------|\n| task_id | 是 | string | 发送提交请求时返回的 task_id |\n\n## 返回结构\n\n### 提交请求返回\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| log_id | uint64 | 唯一的 log id，用于问题定位 |\n| error_code | int | 错误码 |\n| error_msg | string | 错误描述信息 |\n| result | dict | 返回的结果列表 |\n| + task_id | string | 该请求生成的 task_id，后续使用该 task_id 获取结果 |\n\n成功返回示例：\n\n```json\n{\n  \"error_code\": 0,\n  \"error_msg\": \"\",\n  \"log_id\": \"10138598131137362685273585665433\",\n  \"result\": {\n    \"task_id\": \"task-3zy9Bg8CHt1M4pP0cX2q5bg28j268015\"\n  }\n}\n```\n\n### 获取结果返回\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| log_id | uint64 | 唯一的 log id，用于问题定位 |\n| error_code | int | 错误码 |\n| error_msg | string | 错误描述信息 |\n| result | dict | 返回的结果列表 |\n| + task_id | string | 任务 ID |\n| + status | string | 任务状态：pending（排队中）、processing（运行中）、success（成功）、failed（失败） |\n| + task_error | string | 解析报错信息，包含任务失败、额度不够 |\n| + markdown_url | string | 文档解析结果的 Markdown 格式链接，链接有效期 30 天 |\n| + parse_result_url | string | 文档解析结果的 BOS 链接，链接有效期 30 天 |\n\n成功返回示例：\n\n```json\n{\n  \"error_code\": 0,\n  \"error_msg\": \"\",\n  \"result\": {\n    \"task_id\": \"task-UnvGsgbYZp9pS3BZRHn11ifzjNvKzTgf\",\n    \"status\": \"success\",\n    \"task_error\": null,\n    \"duration\": 902.0,\n    \"parse_result_url\": \"https://xxxxxxxxxxxxxxxxxx\"\n  }\n}\n```\n\n### parse_result_url 返回的 JSON 结构\n\n通过 `parse_result_url` 下载解析结果的 JSON 文件，结构如下：\n\n#### 顶层结构\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| file_name | string | 文档名称 |\n| file_id | string | 文档 ID |\n| pages | list | 文件单页解析内容 |\n| chunks | list | 文件内容切分结果（return_doc_chunks 中 switch 为 True 时有值） |\n\n#### 页面对象（pages[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| page_id | string | 页码 ID |\n| page_num | int | 页码数 |\n| text | string | 当前页的所有纯文字内容 |\n| layouts | list | 页面内容版式分析的结果 |\n| tables | list | 页面表格解析结果 |\n| images | list | 页面中图片解析结果 |\n| meta | dict | 页元信息 |\n\n#### 页面元信息（meta）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| page_width | int | 页面宽度 |\n| page_height | int | 页面高度 |\n| is_scan | bool | 是否扫描件 |\n| page_angle | int | 页面倾斜角度 |\n| page_type | string | 页面属性：text（正文）、contents（目录）、appendix（附录）、others（其他） |\n| sheet_name | string | Excel 的 sheet 名 |\n\n#### 版面元素（layouts[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| layout_id | string | layout 元素唯一标志，以 \"xxxxx-layout-{global_layout_index}\" 形式，global_layout_index 为 layout 元素整个文档的全局索引 |\n| text | string | layout 对应的文本内容。注：当 type 为 table、image 时该字段为空，需根据 type 和 layout_id 分别到 tables、images 字段里找到对应的内容 |\n| position | list | layout 元素在页面中的位置，[x, y, w, h] box 框，左上角和宽高 |\n| type | string | 版面元素类型（见下表） |\n| sub_type | string | 版面元素子类型（见下表） |\n| parent | string | 标题层级树中父节点的 layout_id，若当前 layout 为一级标题，其 parent 为 \"root\"。在 table 和 image 的内版面信息中暂时都为空 |\n| children | list | 标题层级树中子节点的 layout_id。在 table 和 image 的内版面信息中暂时都为空 |\n\n**版面类型（type）取值**：\n\n| 类型 | 说明 |\n|------|------|\n| para | 段落 |\n| table | 表格 |\n| image | 文档中的插图 |\n| head_tail | 页面顶部 |\n| contents | 目录 |\n| seal | 印章 |\n| title | 标题 |\n| formula | 公式 |\n\n**子类型（sub_type）取值**：\n\n当 type 为 title 或 image 时，sub_type 有值：\n\n- **title 的 sub_type**：\n  - `title_{n}`：代表 n 级标题，比如 title_2 代表二级标题\n  - `image_title`：图标题\n  - `table_title`：表标题\n\n- **image 的 sub_type**：\n  - `chart`：统计图表\n  - `figure`：普通插图\n  - `QR_code`：二维码\n  - `Bar_code`：条形码\n\n#### 表格对象（tables[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| layout_id | string | layout ID，与 layouts 中的 type 为 table 的元素的 layout ID 对应 |\n| markdown | string | 表格内容的 Markdown 形式 |\n| table_title_id | list | 表格标题对应的 layout_id，默认为 null |\n| position | list | 边框数据 [x, y, w, h]（以页面坐标为原点），版式格式时有效 |\n| cells | list | 单元格的内版面信息，layout 类型为表格时有值 |\n| matrix | list | 二位数组，表示表格内布局位置信息，每个元素对应 cells 列表中元素的索引 |\n| merge_table | string | 跨页表格标记：\"begin\"（开始）、\"inner\"（中间，表格跨页超过两页）、\"end\"（结束）；非跨页表格该字段为空 |\n\n#### 图片对象（images[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| layout_id | string | layout ID，与 layouts 中 type 为 image 的元素的 layout ID 对应 |\n| image_title_id | list | 图片标题对应的 layout_id，默认为 null |\n| position | list | 边框数据 [x, y, w, h] |\n| content_layouts | list | 图片的内版面信息 |\n| data_url | string | 图片存储链接 |\n| image_description | string | 对统计图表进行内容解析和描述，输出结果为 JSON 字符串，可通过 json.loads 结构化为 JSON 格式 |\n\n#### 分块对象（chunks[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| chunk_id | string | 切片的 ID |\n| content | string | 切片的内容 |\n| type | string | 切片类型，为 text 或者 table |\n| meta | dict | chunk 元信息 |\n| + title | list | chunk 所属的多级标题内容 |\n| + position | list | chunk 的位置，根据分块算法有可能 chunk 跨多个页 |\n| + box | list | chunk 的位置坐标 |\n| + page_num | int | chunk 内容所在页数 |\n\n## 文件限制\n\n| 限制项 | 说明 |\n|--------|------|\n| 文件大小（file_data） | ≤ 50MB，超过 50M 须使用 file_url |\n| 文件大小（file_url） | PDF ≤ 300MB，非 PDF ≤ 50MB |\n| URL 长度 | ≤ 1024 字节 |\n| 页数限制 | PDF ≤ 2000 页 |\n| 优先级 | file_data > file_url（同时存在时 file_url 字段失效） |\n\n## QPS 限制\n\n- 提交请求接口：2 QPS\n- 获取结果接口：10 QPS\n\n## 轮询建议\n\n- 提交请求后 5~10 秒开始轮询\n- 轮询间隔：5 秒\n- 最大轮询时间：300 秒\n\n## 相关文档\n\n- [官方 API 文档](https://ai.baidu.com/ai-doc/OCR/llxst5nn0)\n- [错误码参考](error_codes.md)\n- [API Key 配置指南](apikey-fetch.md)\n\nArchive v1.0.3: 6 files, 16605 bytes\n\nFiles: _meta.json (144b), references/apikey-fetch.md (2455b), references/error_codes.md (4606b), references/parameters.md (10795b), scripts/baidu_doc_parser.py (11057b), SKILL.md (13091b)\n\nFile v1.0.3:SKILL.md\n\n---\nname: baidu-doc-pipeline-parser 百度文档解析\ndescription: 调用百度文档解析API解析文档。支持PDF、Word、Excel、PPT、图片等18+格式。提取文本、表格、版面分析、OCR识别及RAG文档分块。当用户需要解析文档、提取文本/表格、分析文档结构、处理扫描件时使用。触发词：文档解析、PDF解析、Word解析、表格提取、OCR、文档分析、提取文本、文档结构、扫描识别。\nlicense: MIT\n---\n\n# 百度文档解析 Skill\n\n基于百度智能文档分析平台 API，提供文档解析能力。\n\n## 功能概述\n\n- 支持对 doc、pdf、图片、xlsx 等 18 种格式文档进行解析\n- 输出文档的版面、表格、阅读顺序、标题层级、旋转角度等信息\n- 支持中、英、日、韩、法等 20 余种语言类型\n- 可返回 Markdown 格式内容，将非结构化数据转化为易于处理的结构化数据\n- 识别准确率可达 90% 以上\n- 文档分块（适用于 RAG 场景）\n\n## 适用场景\n\n当用户需要：\n- 解析 PDF、Word、Excel 等格式文档\n- 从文档中提取文本内容\n- 识别并提取表格数据\n- 分析文档结构（标题层级、章节、版面布局）\n- 对扫描件进行 OCR 文字识别\n- 将文档分块用于 RAG 应用\n\n## API 配置\n\n### 环境变量（必须）\n\n[百度智能文档分析平台 领取免费测试资源](https://ai.baidu.com/ai-doc/OCR/dk3iqnq51)\n\n使用前请设置以下环境变量：\n\n```bash\nexport BAIDU_DOC_AI_API_KEY=\"your_api_key\"\nexport BAIDU_DOC_AI_SECRET_KEY=\"your_secret_key\"\n```\n\n### 认证方式\n\n通过 API Key 和 Secret Key 获取 access_token，有效期 30 天。\n\n## 支持格式\n\n**版式文档**：pdf, jpg, jpeg, png, bmp, tif, tiff, ofd, ppt, pptx\n\n**流式文档**：doc, docx, txt, xls, xlsx, wps, html, mhtml\n\n## 支持语言\n\nCHN_ENG（中英文）、JAP（日语）、KOR（韩语）、FRE（法语）、SPA（西班牙语）、POR（葡萄牙语）、GER（德语）、ITA（意大利语）、RUS（俄语）、DAN（丹麦语）、DUT（荷兰语）、MAL（马来语）、SWE（瑞典语）、IND（印尼语）、POL（波兰语）、ROM（罗马尼亚语）、TUR（土耳其语）、GRE（希腊语）、HUN（匈牙利语）、THA（泰语）、VIE（越南语）、ARA（阿拉伯语）、HIN（印地语）\n\n## 使用方式\n\n```bash\npython3 scripts/baidu_doc_parser.py --file_data <文件的base64编码>\npython3 scripts/baidu_doc_parser.py --file_url <文件公网URL>\n```\n\n## API 接口\n\n文档解析 API 服务为异步接口，需要先调用**提交请求接口**获取 task_id，然后调用**获取结果接口**进行结果轮询。\n\n### 提交请求接口\n\n- **HTTP 方法**：POST\n- **请求 URL**：`https://aip.baidubce.com/rest/2.0/brain/online/v2/parser/task?access_token={token}`\n- **Content-Type**：`application/x-www-form-urlencoded`\n\n### 获取结果接口\n\n- **HTTP 方法**：POST\n- **请求 URL**：`https://aip.baidubce.com/rest/2.0/brain/online/v2/parser/task/query?access_token={token}`\n- **Content-Type**：`application/x-www-form-urlencoded`\n- **请求参数**：`task_id`（必填，提交请求时返回的 task_id）\n\n## 请求参数\n\n### 文件参数（必选，二选一）\n\n| 参数 | 必选 | 类型 | 说明 |\n|------|------|------|------|\n| `file_data` | 和 file_url 二选一 | string | 文件 Base64 编码数据。版式文档：pdf, jpg, jpeg, png, bmp, tif, tiff, ofd, ppt, pptx；流式文档：doc, docx, txt, xls, xlsx, wps, html, mhtml。文档大小不超过 50M，PDF 最大支持 2000 页。**若文档大小超过 50M，须从 file_url 方式上传**。优先级：file_data > file_url |\n| `file_url` | 和 file_data 二选一 | string | 文件数据 URL，长度不超过 1024 字节，支持单个 URL 传入。PDF 文档大小不超过 300MB，非 PDF 不超过 50M，PDF 最大支持 2000 页。**请注意关闭 URL 防盗链** |\n| `file_name` | 是 | string | 文件名，请保证文件名后缀正确，例如 \"1.pdf\" |\n\n### 核心功能参数\n\n| 参数 | 必选 | 类型 | 可选值范围 | 说明 |\n|------|------|------|----------|------|\n| `recognize_formula` | 否 | bool | True/False | 是否对版式类型文档进行公式识别 |\n| `analysis_chart` | 否 | bool | True/False | 是否对统计图表进行解析 |\n| `angle_adjust` | 否 | bool | True/False | 是否对图片进行角度矫正 |\n| `parse_image_layout` | 否 | bool | True/False | 是否返回文档中的图片位置信息 |\n\n### 语言与格式参数\n\n| 参数 | 必选 | 类型 | 说明 |\n|------|------|------|------|\n| `language_type` | 否 | string | 识别语种类型，默认为 CHN_ENG（中英文） |\n| `switch_digital_width` | 否 | string | 是否对数字进行全半角转换，默认为 auto。可选：auto（不转换）、half（半角输出）、full（全角输出） |\n| `html_table_format` | 否 | bool | 是否将识别出的表格转换为 HTML 格式返回，**default=True** |\n\n### 文档分块参数\n\n`return_doc_chunks` 为字典类型，用于返回文档切分后的片段数据（按语义、字数、标点）：\n\n| 参数 | 必选 | 类型 | 默认值 | 说明 |\n|------|------|------|--------|------|\n| `switch` | 否 | bool | False | 是否进行文档内容切分 |\n| `split_type` | 否 | str | chunk | 切分方式：chunk（按 chunk_size 来切）/ mark（按 separators 来切） |\n| `separators` | 否 | list | ['。','；','！','？',';','!','?'] | 切分标点 |\n| `chunk_size` | 否 | int | -1 | 切分块的大小，-1 表示按照语义自动切分，不限定块的大小 |\n\n## 返回结构\n\n### 提交请求返回\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `log_id` | uint64 | 唯一的 log id，用于问题定位 |\n| `error_code` | int | 错误码 |\n| `error_msg` | string | 错误描述信息 |\n| `result.task_id` | string | 该请求生成的 task_id，后续使用该 task_id 获取审查结果 |\n\n### 获取结果返回\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `log_id` | uint64 | 唯一的 log id |\n| `error_code` | int | 错误码 |\n| `error_msg` | string | 错误描述信息 |\n| `result.task_id` | string | 任务 ID |\n| `result.status` | string | 任务状态：pending（排队中）、processing（运行中）、success（成功）、failed（失败） |\n| `result.task_error` | string | 解析报错信息，包含任务失败、额度不够 |\n| `result.markdown_url` | string | 文档解析结果的 Markdown 格式链接，**链接有效期 30 天** |\n| `result.parse_result_url` | string | 文档解析结果的 BOS 链接（JSON），**链接有效期 30 天** |\n\n### 解析结果 JSON 结构（parse_result_url）\n\n#### 顶层结构\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `file_name` | string | 文档名称 |\n| `file_id` | string | 文档 ID |\n| `pages` | list | 文件单页解析内容 |\n| `chunks` | list | 文件内容切分结果（return_doc_chunks.switch=True 时有值） |\n\n#### 页面对象（pages[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `page_id` | string | 页码 ID |\n| `page_num` | int | 页码数 |\n| `text` | string | 当前页的所有纯文字内容 |\n| `layouts` | list | 页面内容版式分析的结果 |\n| `tables` | list | 页面表格解析结果 |\n| `images` | list | 页面中图片解析结果 |\n| `meta` | dict | 页元信息 |\n\n#### 页面元信息（meta）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `page_width` | int | 页面宽度 |\n| `page_height` | int | 页面高度 |\n| `is_scan` | bool | 是否扫描件 |\n| `page_angle` | int | 页面倾斜角度 |\n| `page_type` | string | 页面属性：text（正文）、contents（目录）、appendix（附录）、others（其他） |\n| `sheet_name` | string | Excel 的 sheet 名 |\n\n#### 版面元素（layouts[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `layout_id` | string | layout 元素唯一标志，格式 \"xxxxx-layout-{global_layout_index}\" |\n| `text` | string | layout 对应的文本内容。注：当 type 为 table/image 时该字段为空，需根据 type 和 layout_id 分别到 tables/images 字段里找到对应内容 |\n| `position` | list | 元素在页面中的位置 [x, y, w, h]，左上角和宽高 |\n| `type` | string | 版面元素类型（见下表） |\n| `sub_type` | string | 版面元素子类型（见下表） |\n| `parent` | string | 标题层级树中父节点的 layout_id，若为一级标题则 parent 为 \"root\" |\n| `children` | list | 标题层级树中子节点的 layout_id 列表 |\n\n**版面类型（type）**：\n\n| 类型 | 说明 |\n|------|------|\n| `para` | 段落 |\n| `table` | 表格 |\n| `image` | 文档中的插图 |\n| `head_tail` | 页面顶部（页眉/页脚） |\n| `contents` | 目录 |\n| `seal` | 印章 |\n| `title` | 标题 |\n| `formula` | 公式 |\n\n**子类型（sub_type）**：\n\n- **title 类**：`title_{n}`（n 级标题，如 title_2 代表二级标题）、`image_title`（图标题）、`table_title`（表标题）\n- **image 类**：`chart`（统计图表）、`figure`（普通插图）、`QR_code`（二维码）、`Bar_code`（条形码）\n\n#### 表格对象（tables[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `layout_id` | string | 与 layouts 中 type 为 table 的元素的 layout ID 对应 |\n| `markdown` | string | 表格内容的 Markdown 形式 |\n| `table_title_id` | list | 表格标题对应的 layout_id，默认为 null |\n| `position` | list | 边框数据 [x, y, w, h]（以页面坐标为原点），版式格式时有效 |\n| `cells` | list | 单元格的内版面信息，layout 类型为表格时有值 |\n| `matrix` | list | 二位数组，表示表格内布局位置信息，每个元素对应 cells 列表中元素的索引 |\n| `merge_table` | string | 跨页表格标记：begin（开始）、inner（中间，超过两页）、end（结束）；非跨页表格该字段为空 |\n\n#### 图片对象（images[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `layout_id` | string | 与 layouts 中 type 为 image 的元素的 layout ID 对应 |\n| `image_title_id` | list | 图片标题对应的 layout_id，默认为 null |\n| `position` | list | 边框数据 [x, y, w, h] |\n| `content_layouts` | list | 图片的内版面信息 |\n| `data_url` | string | 图片存储链接 |\n| `image_description` | string | 对统计图表进行内容解析和描述，输出结果为 JSON 字符串 |\n\n#### 分块对象（chunks[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `chunk_id` | string | 切片的 ID |\n| `content` | string | 切片的内容 |\n| `type` | string | 切片类型：text 或 table |\n| `meta.title` | list | chunk 所属的多级标题内容 |\n| `meta.position` | list | chunk 的位置，根据分块算法有可能 chunk 跨多个页 |\n| `meta.box` | list | chunk 的位置坐标 |\n| `meta.page_num` | int | chunk 内容所在页数 |\n\n## API 特性\n\n### 异步处理流程\n\n1. 调用提交请求接口 → 获取 `task_id`\n2. 通过 `task_id` 调用获取结果接口轮询\n\n### 轮询建议\n\n- 提交请求后 5~10 秒开始轮询\n- 轮询间隔：5 秒\n- 最大轮询时间：300 秒\n\n### QPS 限制\n\n- 提交请求接口：2 QPS\n- 获取结果接口：10 QPS\n\n## 文件限制\n\n| 限制项 | 说明 |\n|--------|------|\n| 文件大小（file_data） | ≤ 50MB，超过 50M 须使用 file_url |\n| 文件大小（file_url） | PDF ≤ 300MB，非 PDF ≤ 50MB |\n| URL 长度 | ≤ 1024 字节 |\n| 页数限制 | PDF ≤ 2000 页 |\n| 优先级 | file_data > file_url（同时存在时 file_url 字段失效） |\n\n## 错误处理\n\n常见错误码：\n\n| 错误码 | 说明 | 解决方案 |\n|--------|------|----------|\n| 110/111 | access_token 无效或过期 | 重新获取 access_token |\n| 216200 | 文件或 URL 为空 | 提供 file_data 或 file_url |\n| 216201 | 文件格式错误 | 检查文件格式是否支持 |\n| 216202 | 文件大小超限 | 缩减文件大小 |\n| 282000 | 内部错误 | 重试或联系技术支持 |\n| 282003 | 缺少必要参数 | 检查必填参数 |\n| 282007 | 任务不存在 | 检查 task_id 是否正确 |\n| 282018 | 服务繁忙 | 降低请求频率 |\n\n完整错误码参见 `references/error_codes.md`\n\n## 在线调试\n\n可在 [示例代码中心](https://console.bce.baidu.com/tools/#/api?product=AI&project=%E6%96%87%E5%AD%97%E8%AF%86%E5%88%AB&parent=%E6%99%BA%E8%83%BD%E6%96%87%E6%A1%A3%E5%88%86%E6%9E%90%E5%B9%B3%E5%8F%B0&api=rest/2.0/brain/online/v2/parser/task&method=post) 申请试该接口，可进行签名验证、查看在线调用的请求内容和返回结果、示例代码的自动生成。\n\n## 脚本\n\n- `scripts/baidu_doc_parser.py`：文档解析主程序，支持命令行快速调用\n\n## 参考文档\n\n- `references/parameters.md`：完整 API 参数与返回结构详解\n- `references/error_codes.md`：完整错误码参考\n- `references/apikey-fetch.md`：API Key 配置指南\n\n## 相关链接\n\n- [官方 API 文档](https://ai.baidu.com/ai-doc/OCR/llxst5nn0)\n- [百度云控制台](https://console.bce.baidu.com/ai/)\n- [智能文档分析平台](https://ai.baidu.com/solution/intelligent-document-analysis)\n\nFile v1.0.3:_meta.json\n\n{\n  \"ownerId\": \"kn75p1w9cr0c8ycct5wzrb5ken83e9nc\",\n  \"slug\": \"baidu-doc-pipeline-parser\",\n  \"version\": \"1.0.3\",\n  \"publishedAt\": 1778307517912\n}\n\nFile v1.0.3:references/apikey-fetch.md\n\n# 百度文档解析 API Key 配置指南\n\n## BAIDU_DOC_AI_API_KEY 和 BAIDU_DOC_AI_SECRET_KEY 未配置\n\n当环境变量 `BAIDU_DOC_AI_API_KEY` 和 `BAIDU_DOC_AI_SECRET_KEY` 未设置时，按照以下步骤操作：\n\n### 1. 获取 API Key 和 Secret Key\n\n访问：**https://ai.baidu.com/ai-doc/REFERENCE/Ck3dwjhhu#1-获取aksk**\n\n- 登录百度云账号\n- 创建应用或查看已有的 API Key 和 Secret Key\n- 复制你的 **API Key** 和 **Secret Key**\n\n### 2. 领取免费测试资源\n\n访问：**https://ai.baidu.com/ai-doc/OCR/dk3iqnq51**\n\n### 3. 配置环境变量\n\n#### 方式一：直接设置环境变量\n\n```bash\nexport BAIDU_DOC_AI_API_KEY=\"your_actual_api_key_here\"\nexport BAIDU_DOC_AI_SECRET_KEY=\"your_actual_secret_key_here\"\n```\n\n#### 方式二：通过配置文件\n\n编辑配置文件：`~/.claude/settings.json` 或项目 `.claude/settings.json`\n\n添加以下结构：\n\n```json\n{\n  \"skills\": {\n    \"entries\": {\n      \"baidu-doc-pipeline-parser\": {\n        \"env\": {\n          \"BAIDU_DOC_AI_API_KEY\": \"your_actual_api_key_here\",\n          \"BAIDU_DOC_AI_SECRET_KEY\": \"your_actual_secret_key_here\"\n        }\n      }\n    }\n  }\n}\n```\n\n将 `your_actual_api_key_here` 替换为你的实际 API Key，`your_actual_secret_key_here` 替换为你的实际 Secret Key。\n\n### 4. 验证配置\n\n```bash\n# 验证 access_token 是否可正常获取\ncurl -X POST 'https://aip.baidubce.com/oauth/2.0/token' \\\n  -d 'grant_type=client_credentials' \\\n  -d 'client_id={your_api_key}' \\\n  -d 'client_secret={your_secret_key}'\n```\n\n成功返回示例：\n\n```json\n{\n  \"access_token\": \"24.xxxxx.xxxxxx.xxxxxxx-xxxxxxx\",\n  \"expires_in\": 2592000\n}\n```\n\n`expires_in` 为 2592000 秒（30 天），到期后需重新获取。\n\n### 5. 测试\n\n```bash\npython3 scripts/baidu_doc_parser.py --file_url \"https://example.com/test.pdf\" --file_name \"test.pdf\"\npython3 scripts/baidu_doc_parser.py --file_data \"<文件的base64编码>\" --file_name \"test.pdf\"\n```\n\n## 常见问题\n\n- 确保环境变量已正确设置（可通过 `echo $BAIDU_DOC_AI_API_KEY` 验证）\n- 确认 API Key 有效且已开通百度智能文档分析平台服务\n- 检查百度云账户余额或免费额度\n- access_token 有效期 30 天，过期后会自动重新获取\n\n## 相关链接\n\n- [获取 AK/SK 文档](https://ai.baidu.com/ai-doc/REFERENCE/Ck3dwjhhu#1-获取aksk)\n- [领取免费资源](https://ai.baidu.com/ai-doc/OCR/dk3iqnq51)\n- [百度云控制台](https://console.bce.baidu.com/ai/)\n\nFile v1.0.3:references/error_codes.md\n\n# 百度文档解析 API 错误码参考\n\n## 通用错误\n\n### 认证相关错误\n\n| 错误码 | 错误信息 | 说明 | 解决方案 |\n|--------|---------|------|----------|\n| 1 | Unknown error | 未知错误 | 重试请求，持续出现请联系技术支持 |\n| 2 | Service temporarily unavailable | 服务暂不可用 | 重试请求，持续出现请联系技术支持 |\n| 3 | Unsupported openapi method | API 接口不存在 | 检查 URL 是否正确，去除非英文字符 |\n| 4 | Open api request limit reached | 集群超限额 | 重试请求，持续出现请联系技术支持 |\n| 6 | No permission to access data | 无 API 访问权限 | 在百度云控制台开通该 API 权限 |\n| 14 | IAM Certification failed | IAM 认证失败 | 检查签名生成方式或改用 AK/SK |\n| 17 | Open api daily request limit reached | 日配额超限 | 购买额度或等待次日重置 |\n| 18 | Open api qps request limit reached | QPS 超限 | 降低请求频率 |\n| 19 | Open api total request limit reached | 总量配额超限 | 购买额外配额 |\n| 100 | Invalid parameter | access_token 无效 | 重新获取 access_token |\n| 110 | Access token invalid or no longer valid | access_token 无效 | token 有效期 30 天，重新获取 |\n| 111 | Access token expired | access_token 过期 | token 有效期 30 天，重新获取 |\n\n### 文件相关错误\n\n| 错误码 | 错误信息 | 说明 | 解决方案 |\n|--------|---------|------|----------|\n| 216200 | empty file or fileurl | 文件或 URL 为空 | 提供 file_data 或 file_url |\n| 216201 | file format error | 文件格式不支持 | 使用支持的格式（PDF、Word、Excel 等） |\n| 216202 | file size error | 文件大小超限 | 缩减文件大小（file_data ≤ 50MB，file_url PDF ≤ 300MB） |\n\n### 任务处理错误\n\n| 错误码 | 错误信息 | 说明 | 解决方案 |\n|--------|---------|------|----------|\n| 282000 | internal error | 任务处理失败 | 重试或联系技术支持 |\n| 282001 | template not found | 合同类型未找到 | 检查合同类型名称 |\n| 282003 | missing parameters | 缺少必要参数 | 检查必填参数 |\n| 282005 | quota exceed error | 额度不足 | 申请增加配额 |\n| 282006 | check user auth error | 用户权限校验失败 | 验证用户权限 |\n| 282007 | task not exist, please check task id | 任务不存在 | 检查 task_id 是否正确 |\n| 282018 | Service busy | 服务繁忙 | 降低请求频率 |\n\n### URL 相关错误\n\n| 错误码 | 错误信息 | 说明 | 解决方案 |\n|--------|---------|------|----------|\n| 282111 | url format illegal | URL 格式不合法 | 检查 URL 格式 |\n| 282112 | url download timeout | URL 下载超时 | 检查 URL 是否可访问 |\n| 282113 | url response invalid | URL 响应无效 | 检查 URL 返回内容是否正确 |\n| 282114 | url size error | URL 长度超过 1024 字节 | 缩短 URL |\n\n### 参数错误\n\n| 错误码 | 错误信息 | 说明 | 解决方案 |\n|--------|---------|------|----------|\n| 283016 | parameters value error | 参数值无效 | 检查参数格式和取值 |\n\n## 错误响应格式\n\n```json\n{\n  \"log_id\": \"13665091038742503867108513247688\",\n  \"error_code\": \"282007\",\n  \"error_msg\": \"task not exist, please check task id\",\n  \"result\": \"null\"\n}\n```\n\n## 错误处理策略\n\n### 重试策略\n\n| 错误类型 | 错误码 | 建议处理方式 |\n|---------|--------|-------------|\n| 瞬时错误 | 1, 2, 4, 282000, 282018 | 指数退避重试 |\n| 认证错误 | 100, 110, 111 | 重新获取 access_token |\n| 配额错误 | 17, 18, 19, 282005 | 等待或购买额外配额 |\n| 参数错误 | 216200, 216201, 282003, 283016 | 修正参数后重试 |\n| URL 错误 | 282111, 282112, 282113, 282114 | 检查并修正 URL |\n\n### 指数退避重试示例\n\n```python\nimport time\n\ndef retry_with_backoff(func, max_retries=3):\n    for i in range(max_retries):\n        try:\n            return func()\n        except Exception as e:\n            if i == max_retries - 1:\n                raise\n            wait_time = 2 ** i  # 1s, 2s, 4s\n            time.sleep(wait_time)\n```\n\n### Token 刷新示例\n\n```python\ndef ensure_valid_token(client):\n    try:\n        client.query_task(\"test-task-id\")\n    except Exception as e:\n        if \"110\" in str(e) or \"111\" in str(e):\n            client.access_token = client._get_access_token()\n```\n\n## 获取帮助\n\n- [提交工单](https://ticket.bce.baidu.com/?_=1648086674827&fromai=1#/ticket/create~productId=96&questionId=1306&channel=2)\n- [权限诊断工具](https://console.bce.baidu.com/tools/#/aiInterfacePermissions)\n- [百度云控制台](https://console.bce.baidu.com/ai/)\n\nFile v1.0.3:references/parameters.md\n\n# 百度文档解析 API 参数详解\n\n## 接口概述\n\n百度文档解析 API 支持对 doc、pdf、图片、xlsx 等 18 种格式文档进行解析，输出文档的版面、表格、阅读顺序、标题层级、旋转角度等信息，支持中、英、日、韩、法等 20 余种语言类型，识别准确率可达 90% 以上。\n\n## API 接口地址\n\n### 提交请求接口\n\n```\nPOST https://aip.baidubce.com/rest/2.0/brain/online/v2/parser/task?access_token={access_token}\nContent-Type: application/x-www-form-urlencoded\n```\n\n### 获取结果接口\n\n```\nPOST https://aip.baidubce.com/rest/2.0/brain/online/v2/parser/task/query?access_token={access_token}\nContent-Type: application/x-www-form-urlencoded\n```\n\n## 提交请求参数\n\n### 文件参数（必选，二选一）\n\n| 参数 | 必选 | 类型 | 说明 |\n|------|------|------|------|\n| file_data | 和 file_url 二选一 | string | 文件的 base64 编码数据。版式文档：pdf, jpg, jpeg, png, bmp, tif, tiff, ofd, ppt, pptx；流式文档：doc, docx, txt, xls, xlsx, wps, html, mhtml。文档大小不超过 50M，其中 PDF 文档最大支持 2000 页。若文档大小超过 50M，须从 file_url 方式上传。优先级：file_data > file_url，当 file_data 字段存在时，file_url 字段失效 |\n| file_url | 和 file_data 二选一 | string | 文件数据 URL，URL 长度不超过 1024 字节，支持单个 URL 传入。PDF 文档大小不超过 300MB，非 PDF 文档大小不超过 50M，其中 PDF 文档最大支持 2000 页。优先级：file_data > file_url。**请注意关闭 URL 防盗链** |\n| file_name | 是 | string | 文件名，请保证文件名后缀正确，例如 \"1.pdf\" |\n\n### 核心功能参数\n\n| 参数 | 必选 | 类型 | 可选值范围 | 说明 |\n|------|------|------|----------|------|\n| recognize_formula | 否 | bool | True/False | 是否对版式类型文档进行公式识别 |\n| analysis_chart | 否 | bool | True/False | 是否对统计图表进行解析 |\n| angle_adjust | 否 | bool | True/False | 是否对图片进行角度矫正 |\n| parse_image_layout | 否 | bool | True/False | 是否返回文档中的图片位置信息 |\n\n### 语言与格式参数\n\n| 参数 | 必选 | 类型 | 默认值 | 说明 |\n|------|------|------|--------|------|\n| language_type | 否 | string | CHN_ENG | 识别语种类型 |\n| switch_digital_width | 否 | string | auto | 是否对数字进行全半角转换。auto：不转换；half：半角输出；full：全角输出 |\n| html_table_format | 否 | bool | True | 是否将识别出的表格转换为 HTML 格式返回 |\n\n### 支持语种列表\n\n| 代码 | 语言 | 代码 | 语言 |\n|------|------|------|------|\n| CHN_ENG | 中英文 | DAN | 丹麦语 |\n| JAP | 日语 | DUT | 荷兰语 |\n| KOR | 韩语 | MAL | 马来语 |\n| FRE | 法语 | SWE | 瑞典语 |\n| SPA | 西班牙语 | IND | 印尼语 |\n| POR | 葡萄牙语 | POL | 波兰语 |\n| GER | 德语 | ROM | 罗马尼亚语 |\n| ITA | 意大利语 | TUR | 土耳其语 |\n| RUS | 俄语 | GRE | 希腊语 |\n| HUN | 匈牙利语 | THA | 泰语 |\n| VIE | 越南语 | ARA | 阿拉伯语 |\n| HIN | 印地语 | - | - |\n\n### 文档分块参数\n\n| 参数 | 必选 | 类型 | 默认值 | 说明 |\n|------|------|------|--------|------|\n| return_doc_chunks | 否 | dict | - | 是否返回文档切分后的片段数据（按语义、字数、标点） |\n| + switch | 否 | bool | False | 是否进行文档内容切分 |\n| + split_type | 否 | str | chunk | 切分方式。chunk：按照 chunk_size 来切；mark：按照 separators 来切 |\n| + separators | 否 | list | ['。','；','！','？',';','!','?'] | 切分标点 |\n| + chunk_size | 否 | int | -1 | 切分块的大小，-1 表示按照语义自动切分，不限定块的大小 |\n\n## 获取结果请求参数\n\n| 参数 | 必选 | 类型 | 说明 |\n|------|------|------|------|\n| task_id | 是 | string | 发送提交请求时返回的 task_id |\n\n## 返回结构\n\n### 提交请求返回\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| log_id | uint64 | 唯一的 log id，用于问题定位 |\n| error_code | int | 错误码 |\n| error_msg | string | 错误描述信息 |\n| result | dict | 返回的结果列表 |\n| + task_id | string | 该请求生成的 task_id，后续使用该 task_id 获取结果 |\n\n成功返回示例：\n\n```json\n{\n  \"error_code\": 0,\n  \"error_msg\": \"\",\n  \"log_id\": \"10138598131137362685273585665433\",\n  \"result\": {\n    \"task_id\": \"task-3zy9Bg8CHt1M4pP0cX2q5bg28j268015\"\n  }\n}\n```\n\n### 获取结果返回\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| log_id | uint64 | 唯一的 log id，用于问题定位 |\n| error_code | int | 错误码 |\n| error_msg | string | 错误描述信息 |\n| result | dict | 返回的结果列表 |\n| + task_id | string | 任务 ID |\n| + status | string | 任务状态：pending（排队中）、processing（运行中）、success（成功）、failed（失败） |\n| + task_error | string | 解析报错信息，包含任务失败、额度不够 |\n| + markdown_url | string | 文档解析结果的 Markdown 格式链接，链接有效期 30 天 |\n| + parse_result_url | string | 文档解析结果的 BOS 链接，链接有效期 30 天 |\n\n成功返回示例：\n\n```json\n{\n  \"error_code\": 0,\n  \"error_msg\": \"\",\n  \"result\": {\n    \"task_id\": \"task-UnvGsgbYZp9pS3BZRHn11ifzjNvKzTgf\",\n    \"status\": \"success\",\n    \"task_error\": null,\n    \"duration\": 902.0,\n    \"parse_result_url\": \"https://xxxxxxxxxxxxxxxxxx\"\n  }\n}\n```\n\n### parse_result_url 返回的 JSON 结构\n\n通过 `parse_result_url` 下载解析结果的 JSON 文件，结构如下：\n\n#### 顶层结构\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| file_name | string | 文档名称 |\n| file_id | string | 文档 ID |\n| pages | list | 文件单页解析内容 |\n| chunks | list | 文件内容切分结果（return_doc_chunks 中 switch 为 True 时有值） |\n\n#### 页面对象（pages[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| page_id | string | 页码 ID |\n| page_num | int | 页码数 |\n| text | string | 当前页的所有纯文字内容 |\n| layouts | list | 页面内容版式分析的结果 |\n| tables | list | 页面表格解析结果 |\n| images | list | 页面中图片解析结果 |\n| meta | dict | 页元信息 |\n\n#### 页面元信息（meta）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| page_width | int | 页面宽度 |\n| page_height | int | 页面高度 |\n| is_scan | bool | 是否扫描件 |\n| page_angle | int | 页面倾斜角度 |\n| page_type | string | 页面属性：text（正文）、contents（目录）、appendix（附录）、others（其他） |\n| sheet_name | string | Excel 的 sheet 名 |\n\n#### 版面元素（layouts[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| layout_id | string | layout 元素唯一标志，以 \"xxxxx-layout-{global_layout_index}\" 形式，global_layout_index 为 layout 元素整个文档的全局索引 |\n| text | string | layout 对应的文本内容。注：当 type 为 table、image 时该字段为空，需根据 type 和 layout_id 分别到 tables、images 字段里找到对应的内容 |\n| position | list | layout 元素在页面中的位置，[x, y, w, h] box 框，左上角和宽高 |\n| type | string | 版面元素类型（见下表） |\n| sub_type | string | 版面元素子类型（见下表） |\n| parent | string | 标题层级树中父节点的 layout_id，若当前 layout 为一级标题，其 parent 为 \"root\"。在 table 和 image 的内版面信息中暂时都为空 |\n| children | list | 标题层级树中子节点的 layout_id。在 table 和 image 的内版面信息中暂时都为空 |\n\n**版面类型（type）取值**：\n\n| 类型 | 说明 |\n|------|------|\n| para | 段落 |\n| table | 表格 |\n| image | 文档中的插图 |\n| head_tail | 页面顶部 |\n| contents | 目录 |\n| seal | 印章 |\n| title | 标题 |\n| formula | 公式 |\n\n**子类型（sub_type）取值**：\n\n当 type 为 title 或 image 时，sub_type 有值：\n\n- **title 的 sub_type**：\n  - `title_{n}`：代表 n 级标题，比如 title_2 代表二级标题\n  - `image_title`：图标题\n  - `table_title`：表标题\n\n- **image 的 sub_type**：\n  - `chart`：统计图表\n  - `figure`：普通插图\n  - `QR_code`：二维码\n  - `Bar_code`：条形码\n\n#### 表格对象（tables[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| layout_id | string | layout ID，与 layouts 中的 type 为 table 的元素的 layout ID 对应 |\n| markdown | string | 表格内容的 Markdown 形式 |\n| table_title_id | list | 表格标题对应的 layout_id，默认为 null |\n| position | list | 边框数据 [x, y, w, h]（以页面坐标为原点），版式格式时有效 |\n| cells | list | 单元格的内版面信息，layout 类型为表格时有值 |\n| matrix | list | 二位数组，表示表格内布局位置信息，每个元素对应 cells 列表中元素的索引 |\n| merge_table | string | 跨页表格标记：\"begin\"（开始）、\"inner\"（中间，表格跨页超过两页）、\"end\"（结束）；非跨页表格该字段为空 |\n\n#### 图片对象（images[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| layout_id | string | layout ID，与 layouts 中 type 为 image 的元素的 layout ID 对应 |\n| image_title_id | list | 图片标题对应的 layout_id，默认为 null |\n| position | list | 边框数据 [x, y, w, h] |\n| content_layouts | list | 图片的内版面信息 |\n| data_url | string | 图片存储链接 |\n| image_description | string | 对统计图表进行内容解析和描述，输出结果为 JSON 字符串，可通过 json.loads 结构化为 JSON 格式 |\n\n#### 分块对象（chunks[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| chunk_id | string | 切片的 ID |\n| content | string | 切片的内容 |\n| type | string | 切片类型，为 text 或者 table |\n| meta | dict | chunk 元信息 |\n| + title | list | chunk 所属的多级标题内容 |\n| + position | list | chunk 的位置，根据分块算法有可能 chunk 跨多个页 |\n| + box | list | chunk 的位置坐标 |\n| + page_num | int | chunk 内容所在页数 |\n\n## 文件限制\n\n| 限制项 | 说明 |\n|--------|------|\n| 文件大小（file_data） | ≤ 50MB，超过 50M 须使用 file_url |\n| 文件大小（file_url） | PDF ≤ 300MB，非 PDF ≤ 50MB |\n| URL 长度 | ≤ 1024 字节 |\n| 页数限制 | PDF ≤ 2000 页 |\n| 优先级 | file_data > file_url（同时存在时 file_url 字段失效） |\n\n## QPS 限制\n\n- 提交请求接口：2 QPS\n- 获取结果接口：10 QPS\n\n## 轮询建议\n\n- 提交请求后 5~10 秒开始轮询\n- 轮询间隔：5 秒\n- 最大轮询时间：300 秒\n\n## 相关文档\n\n- [官方 API 文档](https://ai.baidu.com/ai-doc/OCR/llxst5nn0)\n- [错误码参考](error_codes.md)\n- [API Key 配置指南](apikey-fetch.md)\n\nArchive v1.0.2: 6 files, 16606 bytes\n\nFiles: _meta.json (144b), references/apikey-fetch.md (2455b), references/error_codes.md (4606b), references/parameters.md (10795b), scripts/baidu_doc_parser.py (11057b), SKILL.md (13091b)\n\nFile v1.0.2:SKILL.md\n\n---\nname: baidu-doc-pipeline-parser 百度文档解析\ndescription: 调用百度文档解析API解析文档。支持PDF、Word、Excel、PPT、图片等18+格式。提取文本、表格、版面分析、OCR识别及RAG文档分块。当用户需要解析文档、提取文本/表格、分析文档结构、处理扫描件时使用。触发词：文档解析、PDF解析、Word解析、表格提取、OCR、文档分析、提取文本、文档结构、扫描识别。\nlicense: MIT\n---\n\n# 百度文档解析 Skill\n\n基于百度智能文档分析平台 API，提供文档解析能力。\n\n## 功能概述\n\n- 支持对 doc、pdf、图片、xlsx 等 18 种格式文档进行解析\n- 输出文档的版面、表格、阅读顺序、标题层级、旋转角度等信息\n- 支持中、英、日、韩、法等 20 余种语言类型\n- 可返回 Markdown 格式内容，将非结构化数据转化为易于处理的结构化数据\n- 识别准确率可达 90% 以上\n- 文档分块（适用于 RAG 场景）\n\n## 适用场景\n\n当用户需要：\n- 解析 PDF、Word、Excel 等格式文档\n- 从文档中提取文本内容\n- 识别并提取表格数据\n- 分析文档结构（标题层级、章节、版面布局）\n- 对扫描件进行 OCR 文字识别\n- 将文档分块用于 RAG 应用\n\n## API 配置\n\n### 环境变量（必须）\n\n[百度智能文档分析平台 领取免费测试资源](https://ai.baidu.com/ai-doc/OCR/dk3iqnq51)\n\n使用前请设置以下环境变量：\n\n```bash\nexport BAIDU_DOC_AI_API_KEY=\"your_api_key\"\nexport BAIDU_DOC_AI_SECRET_KEY=\"your_secret_key\"\n```\n\n### 认证方式\n\n通过 API Key 和 Secret Key 获取 access_token，有效期 30 天。\n\n## 支持格式\n\n**版式文档**：pdf, jpg, jpeg, png, bmp, tif, tiff, ofd, ppt, pptx\n\n**流式文档**：doc, docx, txt, xls, xlsx, wps, html, mhtml\n\n## 支持语言\n\nCHN_ENG（中英文）、JAP（日语）、KOR（韩语）、FRE（法语）、SPA（西班牙语）、POR（葡萄牙语）、GER（德语）、ITA（意大利语）、RUS（俄语）、DAN（丹麦语）、DUT（荷兰语）、MAL（马来语）、SWE（瑞典语）、IND（印尼语）、POL（波兰语）、ROM（罗马尼亚语）、TUR（土耳其语）、GRE（希腊语）、HUN（匈牙利语）、THA（泰语）、VIE（越南语）、ARA（阿拉伯语）、HIN（印地语）\n\n## 使用方式\n\n```bash\npython3 scripts/baidu_doc_parser.py --file_data <文件的base64编码>\npython3 scripts/baidu_doc_parser.py --file_url <文件公网URL>\n```\n\n## API 接口\n\n文档解析 API 服务为异步接口，需要先调用**提交请求接口**获取 task_id，然后调用**获取结果接口**进行结果轮询。\n\n### 提交请求接口\n\n- **HTTP 方法**：POST\n- **请求 URL**：`https://aip.baidubce.com/rest/2.0/brain/online/v2/parser/task?access_token={token}`\n- **Content-Type**：`application/x-www-form-urlencoded`\n\n### 获取结果接口\n\n- **HTTP 方法**：POST\n- **请求 URL**：`https://aip.baidubce.com/rest/2.0/brain/online/v2/parser/task/query?access_token={token}`\n- **Content-Type**：`application/x-www-form-urlencoded`\n- **请求参数**：`task_id`（必填，提交请求时返回的 task_id）\n\n## 请求参数\n\n### 文件参数（必选，二选一）\n\n| 参数 | 必选 | 类型 | 说明 |\n|------|------|------|------|\n| `file_data` | 和 file_url 二选一 | string | 文件 Base64 编码数据。版式文档：pdf, jpg, jpeg, png, bmp, tif, tiff, ofd, ppt, pptx；流式文档：doc, docx, txt, xls, xlsx, wps, html, mhtml。文档大小不超过 50M，PDF 最大支持 2000 页。**若文档大小超过 50M，须从 file_url 方式上传**。优先级：file_data > file_url |\n| `file_url` | 和 file_data 二选一 | string | 文件数据 URL，长度不超过 1024 字节，支持单个 URL 传入。PDF 文档大小不超过 300MB，非 PDF 不超过 50M，PDF 最大支持 2000 页。**请注意关闭 URL 防盗链** |\n| `file_name` | 是 | string | 文件名，请保证文件名后缀正确，例如 \"1.pdf\" |\n\n### 核心功能参数\n\n| 参数 | 必选 | 类型 | 可选值范围 | 说明 |\n|------|------|------|----------|------|\n| `recognize_formula` | 否 | bool | True/False | 是否对版式类型文档进行公式识别 |\n| `analysis_chart` | 否 | bool | True/False | 是否对统计图表进行解析 |\n| `angle_adjust` | 否 | bool | True/False | 是否对图片进行角度矫正 |\n| `parse_image_layout` | 否 | bool | True/False | 是否返回文档中的图片位置信息 |\n\n### 语言与格式参数\n\n| 参数 | 必选 | 类型 | 说明 |\n|------|------|------|------|\n| `language_type` | 否 | string | 识别语种类型，默认为 CHN_ENG（中英文） |\n| `switch_digital_width` | 否 | string | 是否对数字进行全半角转换，默认为 auto。可选：auto（不转换）、half（半角输出）、full（全角输出） |\n| `html_table_format` | 否 | bool | 是否将识别出的表格转换为 HTML 格式返回，**default=True** |\n\n### 文档分块参数\n\n`return_doc_chunks` 为字典类型，用于返回文档切分后的片段数据（按语义、字数、标点）：\n\n| 参数 | 必选 | 类型 | 默认值 | 说明 |\n|------|------|------|--------|------|\n| `switch` | 否 | bool | False | 是否进行文档内容切分 |\n| `split_type` | 否 | str | chunk | 切分方式：chunk（按 chunk_size 来切）/ mark（按 separators 来切） |\n| `separators` | 否 | list | ['。','；','！','？',';','!','?'] | 切分标点 |\n| `chunk_size` | 否 | int | -1 | 切分块的大小，-1 表示按照语义自动切分，不限定块的大小 |\n\n## 返回结构\n\n### 提交请求返回\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `log_id` | uint64 | 唯一的 log id，用于问题定位 |\n| `error_code` | int | 错误码 |\n| `error_msg` | string | 错误描述信息 |\n| `result.task_id` | string | 该请求生成的 task_id，后续使用该 task_id 获取审查结果 |\n\n### 获取结果返回\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `log_id` | uint64 | 唯一的 log id |\n| `error_code` | int | 错误码 |\n| `error_msg` | string | 错误描述信息 |\n| `result.task_id` | string | 任务 ID |\n| `result.status` | string | 任务状态：pending（排队中）、processing（运行中）、success（成功）、failed（失败） |\n| `result.task_error` | string | 解析报错信息，包含任务失败、额度不够 |\n| `result.markdown_url` | string | 文档解析结果的 Markdown 格式链接，**链接有效期 30 天** |\n| `result.parse_result_url` | string | 文档解析结果的 BOS 链接（JSON），**链接有效期 30 天** |\n\n### 解析结果 JSON 结构（parse_result_url）\n\n#### 顶层结构\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `file_name` | string | 文档名称 |\n| `file_id` | string | 文档 ID |\n| `pages` | list | 文件单页解析内容 |\n| `chunks` | list | 文件内容切分结果（return_doc_chunks.switch=True 时有值） |\n\n#### 页面对象（pages[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `page_id` | string | 页码 ID |\n| `page_num` | int | 页码数 |\n| `text` | string | 当前页的所有纯文字内容 |\n| `layouts` | list | 页面内容版式分析的结果 |\n| `tables` | list | 页面表格解析结果 |\n| `images` | list | 页面中图片解析结果 |\n| `meta` | dict | 页元信息 |\n\n#### 页面元信息（meta）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `page_width` | int | 页面宽度 |\n| `page_height` | int | 页面高度 |\n| `is_scan` | bool | 是否扫描件 |\n| `page_angle` | int | 页面倾斜角度 |\n| `page_type` | string | 页面属性：text（正文）、contents（目录）、appendix（附录）、others（其他） |\n| `sheet_name` | string | Excel 的 sheet 名 |\n\n#### 版面元素（layouts[]）\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| `layout_id` | string | layout 元素唯一标志，格式 \"xxxxx-layout-{global_layout_index}\" |\n| `text` | string | layout 对应的文本内容。注：当 type 为 table/image 时该字段为空，需根据 type 和 layout_id 分别到 tables/images 字段里找到对应内容 |\n| `position` | list | 元素在页面中的位置 [x, y, w, h]，左上角和宽高 |\n| `type` | string | 版面元素类型（见下表） |\n| `sub_type` | string | 版面元素子类型（见下表） |\n| `parent` | string | 标题层级树中父节点的 layout_id，若为一级标题则 parent 为 \"root\" |\n| `children` | list | 标题层级树中子节点的 layout_id 列表 |","readmeExcerpt":"Skill: 百度文档解析pipeline-parser Owner: maglanyulan Summary: 调用百度文档解析API解析文档。支持PDF、Word、Excel、PPT、图片等18+格式。提取文本、表格、版面分析、OCR识别及RAG文档分块。当用户需要解析文档、提取文本/表格、分析文档结构、处理扫描件时使用。触发词：文档解析、PDF解析、Word解析、表格提取、OCR、文档分析、提取文本、文档结构、扫描识别。 Tags: latest:1.0.8 Version history: v1.0.8 | 2026-09-17T11:15:55.631Z | user - 移除 skill-card.md 文件。 - SKILL.md 文档中，调整了免费额度表：企业实名认证用户额度由 1000 页改为 200 页。 - 页面对象解析字段及部分类型补充、细化（如 page_num、text 字段描述、type/版面类型等","codeSnippets":[],"executableExamples":[{"language":"bash","snippet":"export BAIDU_DOC_AI_API_KEY=\"your_api_key\"\nexport BAIDU_DOC_AI_SECRET_KEY=\"your_secret_key\""},{"language":"bash","snippet":"python3 scripts/baidu_doc_parser.py --file_data <文件的base64编码>\npython3 scripts/baidu_doc_parser.py --file_url <文件公网URL>"},{"language":"bash","snippet":"export BAIDU_DOC_AI_API_KEY=\"your_actual_api_key_here\"\nexport BAIDU_DOC_AI_SECRET_KEY=\"your_actual_secret_key_here\""},{"language":"json","snippet":"{\n  \"skills\": {\n    \"entries\": {\n      \"baidu-doc-pipeline-parser\": {\n        \"env\": {\n          \"BAIDU_DOC_AI_API_KEY\": \"your_actual_api_key_here\",\n          \"BAIDU_DOC_AI_SECRET_KEY\": \"your_actual_secret_key_here\"\n        }\n      }\n    }\n  }\n}"},{"language":"bash","snippet":"curl -X POST 'https://aip.baidubce.com/oauth/2.0/token' \\\n  -d 'grant_type=client_credentials' \\\n  -d 'client_id={your_api_key}' \\\n  -d 'client_secret={your_secret_key}'"},{"language":"bash","snippet":"# 验证 access_token 是否可正常获取\ncurl -X POST 'https://aip.baidubce.com/oauth/2.0/token' \\\n  -d 'grant_type=client_credentials' \\\n  -d 'client_id={your_api_key}' \\\n  -d 'client_secret={your_secret_key}'"}],"parameters":null,"dependencies":[],"permissions":[],"extractedFiles":[{"path":"SKILL.md","content":"---\nname: baidu-doc-pipeline-parser 百度文档解析\ndescription: 调用百度文档解析API解析文档。支持PDF、Word、Excel、PPT、图片等18+格式。提取文本、表格、版面分析、OCR识别及RAG文档分块。当用户需要解析文档、提取文本/表格、分析文档结构、处理扫描件时使用。触发词：文档解析、PDF解析、Word解析、表格提取、OCR、文档分析、提取文本、文档结构、扫描识别。\nlicense: MIT\n---\n\n# 百度文档解析 Skill\n\n基于百度智能文档分析平台 API，提供文档解析能力。\n\n## 功能概述\n\n- 支持对 doc、pdf、图片、xlsx 等 18 种格式文档进行解析\n- 输出文档的版面、表格、阅读顺序、标题层级、旋转角度等信息\n- 支持中、英、日、韩、法等 20 余种语言类型\n- 可返回 Markdown 格式内容，将非结构化数据转化为易于处理的结构化数据\n- 识别准确率可达 90% 以上\n- 文档分块（适用于 RAG 场景）\n\n## 适用场景\n\n当用户需要：\n- 解析 PDF、Word、Excel 等格式文档\n- 从文档中提取文本内容\n- 识别并提取表格数据\n- 分析文档结构（标题层级、章节、版面布局）\n- 对扫描件进行 OCR 文字识别\n- 将文档分块用于 RAG 应用\n\n\n## 免费资源领取和计费说明\n[百度智能文档分析平台 领取免费测试资源](https://cloud.baidu.com/doc/OCR/s/fk3h7xu7h)\n\n\n[百度智能文档分析平台计费与购买方式](https://cloud.baidu.com/doc/OCR/s/Fls06fa15#%E6%96%87%E6%A1%A3%E8%A7%A3%E6%9E%90%EF%BC%88paddleocr-vl%EF%BC%89)\n\n\n| 用户类型 | 免费额度 |\n|---------|---------|\n| 个人实名认证用户 | **200 页** |\n| 企业实名认证用户 | **200 页** |\n\n## API 配置\n\n### 额度获取方式\n\n您可通过百度智能云平台获取[免费额度](https://cloud.baidu.com/doc/OCR/s/fk3h7xu7h)与[购买调用资源](https://cloud.baidu.com/doc/OCR/s/Fls06fa15#%E6%96%87%E6%A1%A3%E8%A7%A3%E6%9E%90%EF%BC%88paddleocr-vl%EF%BC%89)\n\n### 环境变量（必须）\n\n[百度智能文档分析平台 领取免费测试资源](https://ai.baidu.com/ai-doc/OCR/dk3iqnq51)\n\n使用前请设置以下环境变量：\n\n```bash\nexport BAIDU_DOC_AI_API_KEY=\"your_api_key\"\nexport BAIDU_DOC_AI_SECRET_KEY=\"your_secret_key\"\n```\n\n### 认证方式\n\n通过 API Key 和 Secret Key 获取 access_token，有效期 30 天。\n\n## 支持格式\n\n**版式文档**：pdf, jpg, jpeg, png, bmp, tif, tiff, ofd, ppt, pptx\n\n**流式文档**：doc, docx, txt, xls, xlsx, wps, html, mhtml\n\n## 支持语言\n\nCHN_ENG（中英文）、JAP（日语）、KOR（韩语）、FRE（法语）、SPA（西班牙语）、POR（葡萄牙语）、GER（德语）、ITA（意大利语）、RUS（俄语）、DAN（丹麦语）、DUT（荷兰语）、MAL（马来语）、SWE（瑞典语）、IND（印尼语）、POL（波兰语）、ROM（罗马尼亚语）、TUR（土耳其语）、GRE（希腊语）、HUN（匈牙利语）、THA（泰语）、VIE（越南语）、ARA（阿拉伯语）、HIN（印地语）\n\n## 使用方式\n\n```bash\npython3 scripts/baidu_doc_parser.py --file_data <文件的base64编码>\npython3 scripts/baidu_doc_parser.py --file_url <文件公网URL>\n```\n\n## API 接口\n\n文档解析 API 服务为异步接口，需要先调用**提交请求接口**获取 task_id，然后调用**获取结果接口**进行结果轮询。\n\n### 提交请求接口\n\n- **HTTP 方法**：POST\n- **请求 URL**：`https://aip.baidubce.com/rest/2.0/brain/online/v2/parser/task?access_token={token}`\n- **Content-Type**：`application/x-www-form-urlencoded`\n\n### 获取结果接口\n\n- **HTTP 方法**：POST\n- **请求 URL**：`https://aip.baidubce.com/rest/2.0/brain/online/v2/parser/task/query?access_token={token}`\n- **Content-Type**：`application/x-www-form-urlencoded`\n- **请求参数**：`task_id`（必填，提交请求时返回的 task_id）\n\n## 请求参数\n\n### 文件参数（必选，二选一）\n\n| 参数 | 必选 | 类型 | 说明 |\n|------|------|------|------|\n| `file_data` | 和 file_url 二选一 | string | 文件 Base64 编码数据。版式文档：pdf, jpg, jpeg, png, bmp, tif, tiff, ofd, ppt, pptx；流式文档：doc, docx, txt, xls, xlsx, wps, html, mhtml。文档大小不超过 50M，PDF 最大支持 2000 页。**若文档大小超过 50M，须从 file_url 方式上传**。优先级：file_data > file_url |\n| `file_url` | 和 file_data 二选一 | string | 文件数据 URL，长度不超过 1024 字节，支持单个 URL 传入。PDF 文档大小不超过 300MB，非 PDF 不超过 50M，PDF 最大支持 2000 页。**请注意关闭 URL 防盗链** |\n| `file_name` | 是 | string | 文件名，请保证文件名后缀正确，例如 \"1.pdf\" |\n\n### 核心功能参数\n\n| 参数 | 必选 | 类型 | 可选值范围 | 说明 |\n|------|------|------|----------|------|\n| `recognize_formula` | 否 | bool | Tru"},{"path":"_meta.json","content":"{\n  \"ownerId\": \"kn75p1w9cr0c8ycct5wzrb5ken83e9nc\",\n  \"slug\": \"baidu-doc-pipeline-parser\",\n  \"version\": \"1.0.8\",\n  \"publishedAt\": 1789643755631\n}"},{"path":"references/apikey-fetch.md","content":"# 百度文档解析 API Key 配置指南\n\n## BAIDU_DOC_AI_API_KEY 和 BAIDU_DOC_AI_SECRET_KEY 未配置\n\n当环境变量 `BAIDU_DOC_AI_API_KEY` 和 `BAIDU_DOC_AI_SECRET_KEY` 未设置时，按照以下步骤操作：\n\n### 1. 获取 API Key 和 Secret Key\n\n访问：**https://ai.baidu.com/ai-doc/REFERENCE/Ck3dwjhhu#1-获取aksk**\n\n- 登录百度云账号\n- 创建应用或查看已有的 API Key 和 Secret Key\n- 复制你的 **API Key** 和 **Secret Key**\n\n### 2. 领取免费测试资源\n\n访问：**https://ai.baidu.com/ai-doc/OCR/dk3iqnq51**\n\n### 3. 配置环境变量\n\n#### 方式一：直接设置环境变量\n\n```bash\nexport BAIDU_DOC_AI_API_KEY=\"your_actual_api_key_here\"\nexport BAIDU_DOC_AI_SECRET_KEY=\"your_actual_secret_key_here\"\n```\n\n#### 方式二：通过配置文件\n\n编辑配置文件：`~/.claude/settings.json` 或项目 `.claude/settings.json`\n\n添加以下结构：\n\n```json\n{\n  \"skills\": {\n    \"entries\": {\n      \"baidu-doc-pipeline-parser\": {\n        \"env\": {\n          \"BAIDU_DOC_AI_API_KEY\": \"your_actual_api_key_here\",\n          \"BAIDU_DOC_AI_SECRET_KEY\": \"your_actual_secret_key_here\"\n        }\n      }\n    }\n  }\n}\n```\n\n将 `your_actual_api_key_here` 替换为你的实际 API Key，`your_actual_secret_key_here` 替换为你的实际 Secret Key。\n\n### 4. 验证配置\n\n```bash\n# 验证 access_token 是否可正常获取\ncurl -X POST 'https://aip.baidubce.com/oauth/2.0/token' \\\n  -d 'grant_type=client_credentials' \\\n  -d 'client_id={your_api_key}' \\\n  -d 'client_secret={your_secret_key}'\n```\n\n成功返回示例：\n\n```json\n{\n  \"access_token\": \"24.xxxxx.xxxxxx.xxxxxxx-xxxxxxx\",\n  \"expires_in\": 2592000\n}\n```\n\n`expires_in` 为 2592000 秒（30 天），到期后需重新获取。\n\n### 5. 测试\n\n```bash\npython3 scripts/baidu_doc_parser.py --file_url \"https://example.com/test.pdf\" --file_name \"test.pdf\"\npython3 scripts/baidu_doc_parser.py --file_data \"<文件的base64编码>\" --file_name \"test.pdf\"\n```\n\n## 常见问题\n\n- 确保环境变量已正确设置（可通过 `echo $BAIDU_DOC_AI_API_KEY` 验证）\n- 确认 API Key 有效且已开通百度智能文档分析平台服务\n- 检查百度云账户余额或免费额度\n- access_token 有效期 30 天，过期后会自动重新获取\n\n## 相关链接\n\n- [获取 AK/SK 文档](https://ai.baidu.com/ai-doc/REFERENCE/Ck3dwjhhu#1-获取aksk)\n- [领取免费资源](https://ai.baidu.com/ai-doc/OCR/dk3iqnq51)\n- [百度云控制台](https://console.bce.baidu.com/ai/)"},{"path":"references/error_codes.md","content":"# 百度文档解析 API 错误码参考\n\n## 通用错误\n\n### 认证相关错误\n\n| 错误码 | 错误信息 | 说明 | 解决方案 |\n|--------|---------|------|----------|\n| 1 | Unknown error | 未知错误 | 重试请求，持续出现请联系技术支持 |\n| 2 | Service temporarily unavailable | 服务暂不可用 | 重试请求，持续出现请联系技术支持 |\n| 3 | Unsupported openapi method | API 接口不存在 | 检查 URL 是否正确，去除非英文字符 |\n| 4 | Open api request limit reached | 集群超限额 | 重试请求，持续出现请联系技术支持 |\n| 6 | No permission to access data | 无 API 访问权限 | 在百度云控制台开通该 API 权限 |\n| 14 | IAM Certification failed | IAM 认证失败 | 检查签名生成方式或改用 AK/SK |\n| 17 | Open api daily request limit reached | 日配额超限 | 购买额度或等待次日重置 |\n| 18 | Open api qps request limit reached | QPS 超限 | 降低请求频率 |\n| 19 | Open api total request limit reached | 总量配额超限 | 购买额外配额 |\n| 100 | Invalid parameter | access_token 无效 | 重新获取 access_token |\n| 110 | Access token invalid or no longer valid | access_token 无效 | token 有效期 30 天，重新获取 |\n| 111 | Access token expired | access_token 过期 | token 有效期 30 天，重新获取 |\n\n### 文件相关错误\n\n| 错误码 | 错误信息 | 说明 | 解决方案 |\n|--------|---------|------|----------|\n| 216200 | empty file or fileurl | 文件或 URL 为空 | 提供 file_data 或 file_url |\n| 216201 | file format error | 文件格式不支持 | 使用支持的格式（PDF、Word、Excel 等） |\n| 216202 | file size error | 文件大小超限 | 缩减文件大小（file_data ≤ 50MB，file_url PDF ≤ 300MB） |\n\n### 任务处理错误\n\n| 错误码 | 错误信息 | 说明 | 解决方案 |\n|--------|---------|------|----------|\n| 282000 | internal error | 任务处理失败 | 重试或联系技术支持 |\n| 282001 | template not found | 合同类型未找到 | 检查合同类型名称 |\n| 282003 | missing parameters | 缺少必要参数 | 检查必填参数 |\n| 282005 | quota exceed error | 额度不足 | 申请增加配额 |\n| 282006 | check user auth error | 用户权限校验失败 | 验证用户权限 |\n| 282007 | task not exist, please check task id | 任务不存在 | 检查 task_id 是否正确 |\n| 282018 | Service busy | 服务繁忙 | 降低请求频率 |\n\n### URL 相关错误\n\n| 错误码 | 错误信息 | 说明 | 解决方案 |\n|--------|---------|------|----------|\n| 282111 | url format illegal | URL 格式不合法 | 检查 URL 格式 |\n| 282112 | url download timeout | URL 下载超时 | 检查 URL 是否可访问 |\n| 282113 | url response invalid | URL 响应无效 | 检查 URL 返回内容是否正确 |\n| 282114 | url size error | URL 长度超过 1024 字节 | 缩短 URL |\n\n### 参数错误\n\n| 错误码 | 错误信息 | 说明 | 解决方案 |\n|--------|---------|------|----------|\n| 283016 | parameters value error | 参数值无效 | 检查参数格式和取值 |\n\n## 错误响应格式\n\n```json\n{\n  \"log_id\": \"13665091038742503867108513247688\",\n  \"error_code\": \"282007\",\n  \"error_msg\": \"task not exist, please check task id\",\n  \"result\": \"null\"\n}\n```\n\n## 错误处理策略\n\n### 重试策略\n\n| 错误类型 | 错误码 | 建议处理方式 |\n|---------|--------|-------------|\n| 瞬时错误 | 1, 2, 4, 282000, 282018 | 指数退避重试 |\n| 认证错误 | 100, 110, 111 | 重新获取 access_token |\n| 配额错误 | 17, 18, 19, 282005 | 等待或购买额外配额 |\n| 参数错误 | 216200, 216201, 282003, 283016 | 修正参数后重试 |\n| URL 错误 | 282111, 282112, 282113, 282114 | 检查并修正 URL |\n\n### 指数退避重试示例\n\n```python\nimport time\n\ndef retry_with_backoff(func, max_retries=3):\n    for i in range(max_retries):\n        try:\n            return func()\n        except Exception as e:\n            if i == max_retries - 1:\n                raise\n            wait_time = 2 ** i  # 1s, 2s, 4s\n            time.sleep(wait_time)\n```\n\n### Token 刷新示例\n\n```python\ndef ensure_valid_token(cli"},{"path":"references/parameters.md","content":"# 百度文档解析 API 参数详解\n\n## 接口概述\n\n百度文档解析 API 支持对 doc、pdf、图片、xlsx 等 18 种格式文档进行解析，输出文档的版面、表格、阅读顺序、标题层级、旋转角度等信息，支持中、英、日、韩、法等 20 余种语言类型，识别准确率可达 90% 以上。\n\n## API 接口地址\n\n### 提交请求接口\n\n```\nPOST https://aip.baidubce.com/rest/2.0/brain/online/v2/parser/task?access_token={access_token}\nContent-Type: application/x-www-form-urlencoded\n```\n\n### 获取结果接口\n\n```\nPOST https://aip.baidubce.com/rest/2.0/brain/online/v2/parser/task/query?access_token={access_token}\nContent-Type: application/x-www-form-urlencoded\n```\n\n## 提交请求参数\n\n### 文件参数（必选，二选一）\n\n| 参数 | 必选 | 类型 | 说明 |\n|------|------|------|------|\n| file_data | 和 file_url 二选一 | string | 文件的 base64 编码数据。版式文档：pdf, jpg, jpeg, png, bmp, tif, tiff, ofd, ppt, pptx；流式文档：doc, docx, txt, xls, xlsx, wps, html, mhtml。文档大小不超过 50M，其中 PDF 文档最大支持 2000 页。若文档大小超过 50M，须从 file_url 方式上传。优先级：file_data > file_url，当 file_data 字段存在时，file_url 字段失效 |\n| file_url | 和 file_data 二选一 | string | 文件数据 URL，URL 长度不超过 1024 字节，支持单个 URL 传入。PDF 文档大小不超过 300MB，非 PDF 文档大小不超过 50M，其中 PDF 文档最大支持 2000 页。优先级：file_data > file_url。**请注意关闭 URL 防盗链** |\n| file_name | 是 | string | 文件名，请保证文件名后缀正确，例如 \"1.pdf\" |\n\n### 核心功能参数\n\n| 参数 | 必选 | 类型 | 可选值范围 | 说明 |\n|------|------|------|----------|------|\n| recognize_formula | 否 | bool | True/False | 是否对版式类型文档进行公式识别 |\n| analysis_chart | 否 | bool | True/False | 是否对统计图表进行解析 |\n| angle_adjust | 否 | bool | True/False | 是否对图片进行角度矫正 |\n| parse_image_layout | 否 | bool | True/False | 是否返回文档中的图片位置信息 |\n\n### 语言与格式参数\n\n| 参数 | 必选 | 类型 | 默认值 | 说明 |\n|------|------|------|--------|------|\n| language_type | 否 | string | CHN_ENG | 识别语种类型 |\n| switch_digital_width | 否 | string | auto | 是否对数字进行全半角转换。auto：不转换；half：半角输出；full：全角输出 |\n| html_table_format | 否 | bool | True | 是否将识别出的表格转换为 HTML 格式返回 |\n\n### 支持语种列表\n\n| 代码 | 语言 | 代码 | 语言 |\n|------|------|------|------|\n| CHN_ENG | 中英文 | DAN | 丹麦语 |\n| JAP | 日语 | DUT | 荷兰语 |\n| KOR | 韩语 | MAL | 马来语 |\n| FRE | 法语 | SWE | 瑞典语 |\n| SPA | 西班牙语 | IND | 印尼语 |\n| POR | 葡萄牙语 | POL | 波兰语 |\n| GER | 德语 | ROM | 罗马尼亚语 |\n| ITA | 意大利语 | TUR | 土耳其语 |\n| RUS | 俄语 | GRE | 希腊语 |\n| HUN | 匈牙利语 | THA | 泰语 |\n| VIE | 越南语 | ARA | 阿拉伯语 |\n| HIN | 印地语 | - | - |\n\n### 文档分块参数\n\n| 参数 | 必选 | 类型 | 默认值 | 说明 |\n|------|------|------|--------|------|\n| return_doc_chunks | 否 | dict | - | 是否返回文档切分后的片段数据（按语义、字数、标点） |\n| + switch | 否 | bool | False | 是否进行文档内容切分 |\n| + split_type | 否 | str | chunk | 切分方式。chunk：按照 chunk_size 来切；mark：按照 separators 来切 |\n| + separators | 否 | list | ['。','；','！','？',';','!','?'] | 切分标点 |\n| + chunk_size | 否 | int | -1 | 切分块的大小，-1 表示按照语义自动切分，不限定块的大小 |\n\n## 获取结果请求参数\n\n| 参数 | 必选 | 类型 | 说明 |\n|------|------|------|------|\n| task_id | 是 | string | 发送提交请求时返回的 task_id |\n\n## 返回结构\n\n### 提交请求返回\n\n| 字段 | 类型 | 说明 |\n|------|------|------|\n| log_id | uint64 | 唯一的 log id，用于问题定位 |\n| error_code | int | 错误码 |\n| error_msg | string | 错误描述信息 |\n| result | dict | 返回的结果列表 |\n| + task_id | string | 该请求生成的 task_id，后续使用该 task_id 获取结果 |\n\n成功返回示例：\n\n```json\n{\n  \"error_code\": 0,\n  \"error_msg\": \"\",\n  \"log_id\": \"10138598131137362685273585665433\",\n  \"result\": {\n    \"task_id\": \"task-3zy9Bg8CHt1M4p"}],"languages":[],"docsSourceLabel":"CLAWHUB","editorialOverview":"调用百度文档解析API解析文档。支持PDF、Word、Excel、PPT、图片等18+格式。提取文本、表格、版面分析、OCR识别及RAG文档分块。当用户需要解析文档、提取文本/表格、分析文档结构、处理扫描件时使用。触发词：文档解析、PDF解析、Word解析、表格提取、OCR、文档分析、提取文本、文档结构、扫描识别。 Skill: 百度文档解析pipeline-parser Owner: maglanyulan Summary: 调用百度文档解析API解析文档。支持PDF、Word、Excel、PPT、图片等18+格式。提取文本、表格、版面分析、OCR识别及RAG文档分块。当用户需要解析文档、提取文本/表格、分析文档结构、处理扫描件时使用。触发词：文档解析、PDF解析、Word解析、表格提取、OCR、文档分析、提取文本、文档结构、扫描识别。 Tags: latest:1.0.8 Version history: v1.0.8 | 2026-09-17T11:15:55.631Z | user - 移除 skill-card.md 文件。 - SKILL.md 文档中，调整了免费额度表：企业实名认证用户额度由 1000 页改为 200 页。 - 页面对象解析字段及部分类型补充、细化（如 page_num、text 字段描述、type/版面类型等","editorialQuality":{"score":100,"threshold":65,"status":"ready","wordCount":1188,"uniquenessScore":48,"reasons":[]}},"media":{"evidence":{"source":"no-media","verified":false,"confidence":"low","updatedAt":"2026-10-10T12:10:35.808Z","emptyReason":"No screenshots, media assets, or demo links are available."},"primaryImageUrl":null,"mediaAssetCount":0,"assets":[],"demoUrl":null},"ownerResources":{"evidence":{"source":"unclaimed","verified":false,"confidence":"low","updatedAt":"2026-10-10T12:10:35.808Z","emptyReason":"This page has not been claimed by the agent owner."},"hasCustomPage":false,"customPageUpdatedAt":null,"customLinks":[],"structuredLinks":{"docsUrl":null,"demoUrl":null,"supportUrl":null,"pricingUrl":null,"statusUrl":null},"customPage":null},"relatedAgents":{"evidence":{"source":"protocol-neighbors","verified":false,"confidence":"medium","updatedAt":"2026-10-10T14:46:32.358Z","emptyReason":null},"items":[{"id":"8ebccd8e-3863-4187-8355-c3f14e1f9edf","entityType":"agent","canonicalPath":"/agent/iofficeai-aionui","slug":"iofficeai-aionui","name":"AionUi","description":"Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!","url":"https://github.com/iOfficeAI/AionUi","homepage":"https://www.aionui.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-10-09T19:11:12.944Z","createdAt":"2026-02-25T03:38:16.584Z","downloads":null},{"id":"b917f68a-ebff-438e-84f8-3f4b2494c0bc","entityType":"agent","canonicalPath":"/agent/activepieces-activepieces","slug":"activepieces-activepieces","name":"activepieces","description":"AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents","url":"https://github.com/activepieces/activepieces","homepage":"https://www.activepieces.com","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-15T02:22:12.426Z","createdAt":"2026-02-25T03:38:12.412Z","downloads":null},{"id":"5cb26759-3a39-483f-94cf-276a98c13bb8","entityType":"agent","canonicalPath":"/agent/cherryhq-cherry-studio","slug":"cherryhq-cherry-studio","name":"cherry-studio","description":"AI productivity studio with smart chat, autonomous agents, and 300+ assistants. Unified access to frontier LLMs","url":"https://github.com/CherryHQ/cherry-studio","homepage":"https://cherry-ai.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-11T14:38:40.986Z","createdAt":"2026-02-25T03:38:19.379Z","downloads":null},{"id":"6f6582d0-5d76-4f0f-b81d-86520247950b","entityType":"agent","canonicalPath":"/agent/copilotkit-copilotkit","slug":"copilotkit-copilotkit","name":"CopilotKit","description":"The Frontend for Agents & Generative UI. React + Angular","url":"https://github.com/CopilotKit/CopilotKit","homepage":"https://docs.copilotkit.ai","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-03-25T09:50:57.846Z","createdAt":"2026-02-25T03:39:14.617Z","downloads":null}],"links":{"hub":"/agent","source":"/agent/source/clawhub","protocols":[{"label":"OpenClaw","href":"/agent/protocol/openclew"}]}}}