{"id":"75ac7db8-a700-45ef-9941-a541aa1ec366","entityType":"agent","slug":"clawhub-370299455cx-web-ym-mediatoolkit","name":"YM-MediaToolkit(媒体处理工具集)","canonicalUrl":"https://www.xpersona.co/agent/clawhub-370299455cx-web-ym-mediatoolkit","canonicalPath":"/agent/clawhub-370299455cx-web-ym-mediatoolkit","generatedAt":"2026-10-10T14:44:18.920Z","source":"CLAWHUB","claimStatus":"UNCLAIMED","verificationTier":"NONE","summary":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T11:42:42.807Z","emptyReason":null},"description":"自然语言媒体助手 - 视频压缩、MP4/MOV 封面提取、音频转换、字幕识别","descriptionLabel":"Source description","evidenceSummary":"Capability contract not published. No trust telemetry is available yet. 1.5K downloads reported by the source. Last updated 10/10/2026.","installCommand":"clawhub skill install s179wpayjrbp1yrqzswm93bgsx84sbcs:ym-mediatoolkit","sourceUrl":"https://clawhub.ai/370299455cx-web/ym-mediatoolkit","homepage":"https://clawhub.ai/370299455cx-web/skills/ym-mediatoolkit","primaryLinks":[{"label":"View on ClawHub","url":"https://clawhub.ai/370299455cx-web/ym-mediatoolkit","kind":"source"},{"label":"Homepage","url":"https://clawhub.ai/370299455cx-web/skills/ym-mediatoolkit","kind":"homepage"}],"safetyScore":84,"overallRank":62,"popularityScore":63,"trustScore":null,"claimedByName":null,"isOwner":false,"seoDescription":"YM-MediaToolkit(媒体处理工具集) technical dossier on Xpersona with agent coverage, OPENCLEW support, and live trust metadata."},"coverage":{"evidence":{"source":"public-profile","verified":false,"confidence":"medium","updatedAt":"2026-10-10T11:42:42.807Z","emptyReason":null},"protocols":[{"protocol":"OPENCLEW","label":"OpenClaw","status":"self-declared","notes":"Declared in the public agent profile."}],"capabilities":[],"verifiedCount":0,"selfDeclaredCount":1,"capabilityMatrix":{"rows":[{"key":"OPENCLEW","type":"protocol","support":"unknown","confidenceSource":"profile","notes":"Listed on profile"}],"flattenedTokens":"protocol:OPENCLEW|unknown|profile"}},"adoption":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T11:42:42.807Z","emptyReason":null},"stars":null,"forks":null,"downloads":1457,"packageName":null,"latestVersion":"4.2.2","tractionLabel":"1.5K downloads"},"release":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T11:42:42.807Z","emptyReason":null},"lastUpdatedAt":"2026-10-10T11:42:42.807Z","lastCrawledAt":"2026-10-10T11:42:42.807Z","lastIndexedAt":null,"nextCrawlAt":"2026-10-11T11:42:42.807Z","lastVerifiedAt":null,"highlights":[{"version":"4.2.2","createdAt":"2026-06-25T08:57:44.696Z","changelog":"4.3.1 更新 - 修正 HTTP chat 短等待：job 在 wait_timeout_sec 内完成时返回 200 和最终结果，否则返回 202。 - /skill/jobs 接口的 params 仅允许为 JSON object，缺省/null 时自动为 {}，其他类型返回 invalid_params。 - async 仅支持 JSON boolean 或 \"auto\" 字符串，wait_timeout_sec 必须为 0–30 秒数字。 - 异步返回的 output_paths 去重并保持首次出现顺序。","fileCount":20,"zipByteSize":63166},{"version":"4.2.1","createdAt":"2026-06-22T02:46:14.001Z","changelog":"- 新增 chat 的 HTTP 异步闭环，支持 async=true/auto，将多媒体任务自动分流为异步 job。 - chat 支持 async=\"auto\"：音频、压缩、识别字幕、pipeline 等长任务自动异步，信息/封面保持同步。 - job 快照增加 created_by、intent、source、metadata 字段，便于对话上下文展示与管理。 - HTTP 服务启动自动清理过期 job，默认保留 7 天且不超过 200 条历史。 - 新增 .gitignore 文件用于忽略不需纳入版本管理的文件。","fileCount":19,"zipByteSize":61062},{"version":"4.2.0","createdAt":"2026-06-08T02:36:06.669Z","changelog":"# ym-mediatoolkit 4.2.0 更新日志 - 新增 HTTP 异步长任务接口 `/skill/jobs`，支持耗时 action（如压缩、ASR/OCR、字幕、pipeline）。 - 任务状态持久化到 `output/jobs/<job_id>/job.json`，可进行提交、轮询和列表查询。 - 服务重启后，已完成任务仍可查询，未完成的任务将自动标记为中断。 - 原有 `/skill/<action>` 同步接口和 CLI 用法保持不变。 - 文档补充异步任务用法说明。","fileCount":19,"zipByteSize":57729},{"version":"4.0.1","createdAt":"2026-05-26T04:16:15.212Z","changelog":"- Unified all action responses with stable protocol fields: `code`, `reply`, and `hint`. - Standardized error codes such as `missing_source`, `source_not_allowed`, `output_exists`, `parse_failed`, `missing_steps`, `unsupported_action`, `ffmpeg_failed`, `missing_asr_dependency`, and `missing_ocr_dependency`. - Maintained backward compatibility with existing action names, HTTP endpoints, and media processing logic. - No changes to core workflow; previous outputs and parameters remain compatible.","fileCount":17,"zipByteSize":45713},{"version":"4.0.0","createdAt":"2026-05-21T02:27:04.676Z","changelog":"ym-mediatoolkit 4.0.0 introduces natural language control and directory authorization: - Added a natural language `chat` action for intuitive media requests (e.g., extract audio, compress, get info). - Supports `media_roots` whitelist for local file access beyond the working directory. - Natural language requests now receive a user-friendly reply and structured `result` for further automation. - Enhanced maintenance process with new documentation in MAINTENANCE.md. - Added initial intent parsing logic and test cases for release behaviors.","fileCount":13,"zipByteSize":34412},{"version":"3.0.3","createdAt":"2026-05-18T06:47:34.572Z","changelog":"- 支持从当前工作目录内的本地视频文件作为输入源，输入字段兼容 video_url/url/source - 明确所有写文件接口支持 overwrite 控制（默认 true） - 默认输出目录说明：视频、音频、封面分别写入 output/videos、output/audio、output/thumbs - 输入示例及参数表同步支持本地文件与远程 URL","fileCount":9,"zipByteSize":21406},{"version":"3.0.2","createdAt":"2026-05-13T03:22:15.751Z","changelog":"Version 3.0.2 - 新增流式视频处理工具集支持 MP4/MOV 封面提取和批量任务 - 优化功能描述与参数说明，文档结构更加简明 - 明确仅支持 http/https URL，输出路径受限于当前目录 - 提供多种命令行与 HTTP 服务调用示例 - 精简安全细节与依赖说明，聚焦核心功能和使用方式","fileCount":9,"zipByteSize":19648},{"version":"3.0.1","createdAt":"2026-05-08T16:39:33.410Z","changelog":"- Version bump from 2.1.0 to 3.0.1 with no file changes detected. - No updates to features, code, documentation, or configuration. - All functionality and documentation remain unchanged from the previous version.","fileCount":9,"zipByteSize":21984}]},"execution":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No published capability contract is available yet."},"installCommand":"clawhub skill install s179wpayjrbp1yrqzswm93bgsx84sbcs:ym-mediatoolkit","setupComplexity":"low","setupSteps":["Install using `clawhub skill install s179wpayjrbp1yrqzswm93bgsx84sbcs:ym-mediatoolkit` in an isolated environment before connecting it to live workloads.","No published capability contract is available yet, so validate auth and request/response behavior manually.","Review the upstream CLAWHUB listing at https://clawhub.ai/370299455cx-web/ym-mediatoolkit before using production credentials."],"contract":{"contractStatus":"missing","authModes":[],"requires":[],"forbidden":[],"supportsMcp":false,"supportsA2a":false,"supportsStreaming":false,"inputSchemaRef":null,"outputSchemaRef":null,"dataRegion":null,"contractUpdatedAt":null,"sourceUpdatedAt":null,"freshnessSeconds":null},"invocationGuide":{"preferredApi":{"snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-370299455cx-web-ym-mediatoolkit/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-370299455cx-web-ym-mediatoolkit/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-370299455cx-web-ym-mediatoolkit/trust"},"curlExamples":["curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-370299455cx-web-ym-mediatoolkit/snapshot\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-370299455cx-web-ym-mediatoolkit/contract\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-370299455cx-web-ym-mediatoolkit/trust\""],"jsonRequestTemplate":{"query":"summarize this repo","constraints":{"maxLatencyMs":2000,"protocolPreference":["OPENCLEW"]}},"jsonResponseTemplate":{"ok":true,"result":{"summary":"...","confidence":0.9},"meta":{"source":"CLAWHUB","generatedAt":"2026-10-10T14:44:18.916Z"}},"retryPolicy":{"maxAttempts":3,"backoffMs":[500,1500,3500],"retryableConditions":["HTTP_429","HTTP_503","NETWORK_TIMEOUT"]}},"endpoints":{"dossierUrl":"https://www.xpersona.co/api/v1/agents/clawhub-370299455cx-web-ym-mediatoolkit/dossier","snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-370299455cx-web-ym-mediatoolkit/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-370299455cx-web-ym-mediatoolkit/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-370299455cx-web-ym-mediatoolkit/trust"}},"reliability":{"evidence":{"source":"runtime-metrics","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No trust, reliability, or runtime telemetry is available."},"trust":{"status":"unavailable","handshakeStatus":"UNKNOWN","verificationFreshnessHours":null,"reputationScore":null,"p95LatencyMs":null,"successRate30d":null,"fallbackRate":null,"attempts30d":null,"trustUpdatedAt":null,"trustConfidence":"unknown","sourceUpdatedAt":null,"freshnessSeconds":null},"decisionGuardrails":{"doNotUseIf":["Contract metadata is missing or unavailable for deterministic execution."],"safeUseWhen":[],"riskFlags":["missing_or_unavailable_contract","trust_data_unavailable","schema_references_missing"],"operationalConfidence":"low"},"executionMetrics":{"observedLatencyMsP50":null,"observedLatencyMsP95":null,"estimatedCostUsd":null,"uptime30d":null,"rateLimitRpm":null,"rateLimitBurst":null,"lastVerifiedAt":null,"verificationSource":null},"runtimeMetrics":{"successRate":null,"avgLatencyMs":null,"avgCostUsd":null,"hallucinationRate":null,"retryRate":null,"disputeRate":null,"p50Latency":null,"p95Latency":null,"lastUpdated":null}},"benchmarks":{"evidence":{"source":"no-benchmark-data","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No benchmark suites or observed failure patterns are available."},"suites":[],"failurePatterns":[]},"artifacts":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T11:42:42.807Z","emptyReason":null},"readme":"Skill: YM-MediaToolkit(媒体处理工具集)\n\nOwner: 370299455cx-web\n\nSummary: 自然语言媒体助手 - 视频压缩、MP4/MOV 封面提取、音频转换、字幕识别\n\nTags: latest:4.2.2\n\nVersion history:\n\nv4.2.2 | 2026-06-25T08:57:44.696Z | user\n\n4.3.1 更新\n\n- 修正 HTTP chat 短等待：job 在 wait_timeout_sec 内完成时返回 200 和最终结果，否则返回 202。\n- /skill/jobs 接口的 params 仅允许为 JSON object，缺省/null 时自动为 {}，其他类型返回 invalid_params。\n- async 仅支持 JSON boolean 或 \"auto\" 字符串，wait_timeout_sec 必须为 0–30 秒数字。\n- 异步返回的 output_paths 去重并保持首次出现顺序。\n\nv4.2.1 | 2026-06-22T02:46:14.001Z | user\n\n- 新增 chat 的 HTTP 异步闭环，支持 async=true/auto，将多媒体任务自动分流为异步 job。\n- chat 支持 async=\"auto\"：音频、压缩、识别字幕、pipeline 等长任务自动异步，信息/封面保持同步。\n- job 快照增加 created_by、intent、source、metadata 字段，便于对话上下文展示与管理。\n- HTTP 服务启动自动清理过期 job，默认保留 7 天且不超过 200 条历史。\n- 新增 .gitignore 文件用于忽略不需纳入版本管理的文件。\n\nv4.2.0 | 2026-06-08T02:36:06.669Z | user\n\n# ym-mediatoolkit 4.2.0 更新日志\n\n- 新增 HTTP 异步长任务接口 `/skill/jobs`，支持耗时 action（如压缩、ASR/OCR、字幕、pipeline）。\n- 任务状态持久化到 `output/jobs/<job_id>/job.json`，可进行提交、轮询和列表查询。\n- 服务重启后，已完成任务仍可查询，未完成的任务将自动标记为中断。\n- 原有 `/skill/<action>` 同步接口和 CLI 用法保持不变。\n- 文档补充异步任务用法说明。\n\nv4.0.1 | 2026-05-26T04:16:15.212Z | user\n\n- Unified all action responses with stable protocol fields: `code`, `reply`, and `hint`.\n- Standardized error codes such as `missing_source`, `source_not_allowed`, `output_exists`, `parse_failed`, `missing_steps`, `unsupported_action`, `ffmpeg_failed`, `missing_asr_dependency`, and `missing_ocr_dependency`.\n- Maintained backward compatibility with existing action names, HTTP endpoints, and media processing logic.\n- No changes to core workflow; previous outputs and parameters remain compatible.\n\nv4.0.0 | 2026-05-21T02:27:04.676Z | user\n\nym-mediatoolkit 4.0.0 introduces natural language control and directory authorization:\n\n- Added a natural language `chat` action for intuitive media requests (e.g., extract audio, compress, get info).\n- Supports `media_roots` whitelist for local file access beyond the working directory.\n- Natural language requests now receive a user-friendly reply and structured `result` for further automation.\n- Enhanced maintenance process with new documentation in MAINTENANCE.md.\n- Added initial intent parsing logic and test cases for release behaviors.\n\nv3.0.3 | 2026-05-18T06:47:34.572Z | user\n\n- 支持从当前工作目录内的本地视频文件作为输入源，输入字段兼容 video_url/url/source\n- 明确所有写文件接口支持 overwrite 控制（默认 true）\n- 默认输出目录说明：视频、音频、封面分别写入 output/videos、output/audio、output/thumbs\n- 输入示例及参数表同步支持本地文件与远程 URL\n\nv3.0.2 | 2026-05-13T03:22:15.751Z | user\n\nVersion 3.0.2\n\n- 新增流式视频处理工具集支持 MP4/MOV 封面提取和批量任务\n- 优化功能描述与参数说明，文档结构更加简明\n- 明确仅支持 http/https URL，输出路径受限于当前目录\n- 提供多种命令行与 HTTP 服务调用示例\n- 精简安全细节与依赖说明，聚焦核心功能和使用方式\n\nv3.0.1 | 2026-05-08T16:39:33.410Z | user\n\n- Version bump from 2.1.0 to 3.0.1 with no file changes detected.\n- No updates to features, code, documentation, or configuration.\n- All functionality and documentation remain unchanged from the previous version.\n\nv3.0.0 | 2026-04-29T16:43:23.365Z | user\n\n媒体处理工具集 - 压缩、封面提取、音频提取/格式转换,，无需下载完整视频\n\nv2.0.2 | 2026-04-29T16:25:47.051Z | user\n\n- Improved security: All user-supplied output paths (`output_path`, `save_path`, `output_dir`) are now strictly checked against path traversal, system reserved names, and enforced to remain within the working directory.\n- Enhanced SSRF protection: Now, if DNS resolution fails when validating video source URLs, requests are rejected (fail-close strategy), preventing prior accidental bypass on DNS errors.\n- Updated documentation: Security section now details new output path validation, clarifies DNS error handling, and notes HTTP service defaults to 127.0.0.1 (local only).\n- Minor notes added to configuration (`clawhub.note`) reminding users not to expose the HTTP service directly to the internet by default.\n\nv2.0.1 | 2026-04-29T16:08:36.042Z | user\n\n2.1.0 版本强调了安全防护增强，并补充了运行时部署指引：\n\n- 增强 validate_video_url()，多层校验防范 SSRF/LFI，包括协议、Unicode/Punycode 域名、IP段与 DNS 解析（含 IPv6）。\n- 文档显式新增“安全部署指南”：涵盖容器隔离、最小权限、磁盘配额、临时目录清理与前置认证建议。\n- 增加 validate_video_url() 检查细节和已知局限说明（如不防 DNS 重绑定、重定向限制）。\n- 测试和用法说明中涵盖内网 IPv4/v6 及恶意域名的安全验证例子。\n- 版本号从 2.0.0 升级到 2.1.0。\n\nv2.0.0 | 2026-04-29T15:54:25.152Z | user\n\nMajor update: Adds comprehensive video URL security validation and documentation.\n\n- Strict URL validation now blocks local files, internal IPs, loopbacks, and unsafe protocols to prevent LFI and SSRF attacks.\n- Validation is enforced before any video URL is processed by all features (compression, thumbnail, audio, info).\n- Only http:// and https:// video URLs are accepted.\n- New “security” tag added.\n- Documentation fully updated to detail security measures, endpoints, CLI usage, and architecture.\n\nv1.0.0 | 2026-04-27T08:12:40.724Z | user\n\n- Introduced stream-based video processing; no need to download entire files.\n- Added video compression with adjustable output sizes and quality preservation.\n- Implemented thumbnail extraction at any timecode or frame.\n- Enabled audio extraction to MP3, WAV, AAC, and M4A formats.\n- All features use efficient streaming to save time and disk space.\n\nArchive index:\n\nArchive v4.2.2: 20 files, 63166 bytes\n\nFiles: asr_engine.py (3005b), audio_extractor.py (9638b), caption_segmenter.py (8548b), frame_extractor.py (22227b), intent_parser.py (6678b), job_manager.py (11618b), MAINTENANCE.md (10045b), ocr_engine.py (3045b), output/jobs/c65f1520dca6436c87f1028354263718/job.json (826b), requirements.txt (157b), run.py (43830b), scripts/smoke_test.py (8721b), skill-card.md (2381b), skill.json (18110b), SKILL.md (16402b), subtitle_extractor.py (5490b), tests/test_release_behaviors.py (43065b), utils.py (8193b), video_compressor.py (6897b), _meta.json (134b)\n\nFile v4.2.2:SKILL.md\n\n---\nname: ym-mediatoolkit\nversion: 4.3.1\ndescription: 自然语言媒体助手 - 视频压缩、MP4/MOV 封面提取、音频转换、字幕识别\nauthor: your_name\ntags:\n  - video\n  - compression\n  - thumbnail\n  - audio\n  - streaming\n  - ffmpeg\ncategories:\n  - media\n  - utility\nclawhub:\n  entrypoint: python run.py\n  runtime: python3\n  http_port: 8080\n---\n\n# YM MediaToolkit\n\n自然语言媒体助手，支持从远程视频 URL、当前工作目录内的本地视频文件、或配置过 `media_roots` 的本地媒体目录直接处理：\n\n- 自然语言调用\n- 视频压缩\n- MP4/MOV 封面提取\n- 音频提取与转换：MP3 / WAV / AAC / M4A\n- OCR / ASR 字幕识别\n- emlet 字幕二次分句\n- 批量处理\n- JSON 驱动媒体流水线\n\n## 依赖\n\n```bash\npip install -r requirements.txt\n```\n\n系统需要安装：\n\n```bash\nffmpeg\nffprobe\n```\n\n字幕识别依赖已内置在 `requirements.txt`，包括 `faster-whisper`、`paddlepaddle`、`paddleocr`。\n\n## 维护\n\n维护、发布、测试和排障流程见 [MAINTENANCE.md](./MAINTENANCE.md)。\n\n## 命令行\n\n```bash\npython run.py -a <action> -i '<json>'\n```\n\n也可以从 JSON 文件读取参数：\n\n```bash\npython run.py -i params.json\n```\n\n## HTTP 服务\n\n```bash\npython run.py --serve\n```\n\n默认监听 `127.0.0.1:8080`。\n\n健康检查：\n\n```http\nGET /health\n```\n\n异步长任务：\n\n```http\nPOST /skill/jobs\nGET /skill/jobs/<job_id>\nGET /skill/jobs\n```\n\n## 功能\n\n输入源字段可使用 `video_url`、`url` 或 `source`，三者等价。远程输入仅支持 `http/https`，本地输入默认限制在当前工作目录内；需要访问绝对路径时，通过请求参数 `media_roots` 或环境变量 `YM_MEDIA_ROOTS` 配置允许的媒体根目录。\n\n## 返回协议\n\n从 `4.1.0` 开始，所有 action 都会返回稳定协议字段：\n\n| 字段 | 说明 |\n|------|------|\n| `status` | `success` / `partial` / `skipped` / `error` |\n| `code` | 稳定机器码，例如 `ok`、`missing_source`、`output_exists`、`parse_failed` |\n| `reply` | 适合聊天展示的简短中文回复 |\n| `hint` | 面向调用方或用户的下一步建议 |\n\n原有业务字段会继续保留，例如 `output_path`、`saved_path`、`outputPath`、`manifest_path`、`info`、`captions`、`result`。\n\n默认输出目录：\n\n| 类型 | 默认目录 |\n|------|----------|\n| 压缩视频 | `output/videos` |\n| 音频 | `output/audio` |\n| 封面 | `output/thumbs` |\n\n所有会写文件的接口都支持 `overwrite`，默认 `true`。设置为 `false` 时，如果输出文件已存在会直接返回错误。\n\n## 3.0.3 更新\n\n- 支持当前工作目录内的本地视频文件输入。\n- 统一 `video_url` / `url` / `source` 三种输入字段。\n- 增加默认输出目录：`output/videos`、`output/audio`、`output/thumbs`。\n- 增加 `overwrite` 覆盖策略，避免误覆盖已有文件。\n\n## 4.0.0 更新\n\n- 新增自然语言入口 `chat`，适合 Claw 直接转发用户聊天文本。\n- 新增 `media_roots` 白名单，支持处理授权目录内的绝对路径文件。\n- 自然语言命令支持提取音频、提取封面、压缩、查看信息、JSON 流水线。\n- `chat` 返回 `reply` 和结构化 `result`，同时兼顾聊天展示和自动化消费。\n\n## 4.0.1 更新\n\n- 新增 `subtitle` 推荐入口，支持 `asr` / `ocr` / `fusion` 模式。\n- 新增 `asr` 和 `ocr` 单独调试入口。\n- 字幕统一输出 SRT-like JSON：`captionTxt`、`startTimeUs`、`endTimeUs`、`source`、`confidence`。\n- `chat` 支持“识别字幕 / 提取字幕 / 转字幕 / 生成字幕”等自然语言命令。\n\n## 4.0.2 更新\n\n- 将 `faster-whisper`、`paddlepaddle`、`paddleocr` 纳入默认 `requirements.txt`。\n- 字幕识别从“可选依赖”调整为默认安装能力。\n- 运行时仍保留缺依赖 JSON error，方便定位未重新安装依赖的环境。\n\n## 4.1.0 更新\n\n- 所有 action 统一补齐 `code`、`reply`、`hint`，方便 Claw 和后续渠道适配。\n- 增加稳定错误码：`missing_source`、`source_not_allowed`、`output_exists`、`parse_failed`、`missing_steps`、`unsupported_action`、`ffmpeg_failed`、`missing_asr_dependency`、`missing_ocr_dependency`。\n- 保持旧字段兼容，不改变现有 action 名称、HTTP endpoint 和底层媒体处理逻辑。\n\n## 4.1.1 更新\n\n- 新增 `caption_segment`，用于 emlet 字幕二次分句。\n- 默认每句最多 `12` 个字符，按强标点、弱标点、连接词和长度切分。\n- 新增 `protected_terms` 和 `protected_terms_path`，用于保护品牌词、产品名、人名、术语不被拆开。\n- `pipeline` 支持 `subtitle -> caption_segment` 串联；分句步骤未传 `caption_path` 时会自动使用上一步字幕 JSON。\n\n## 4.2.0 更新\n\n- 新增 HTTP 异步长任务接口 `/skill/jobs`，适合压缩、ASR/OCR、字幕和 pipeline 等耗时 action。\n- 任务状态持久化到 `output/jobs/<job_id>/job.json`，支持提交、轮询和列表查询。\n- 现有 `/skill/<action>` 同步接口保持不变；CLI 仍保持同步执行。\n\n## 4.3.0 更新\n\n- `chat` 增加 HTTP 异步闭环：`async=true` 会把识别出的 action 提交为 job。\n- `async=\"auto\"` 会自动将 `audio`、`compress`、`asr`、`ocr`、`subtitle`、`caption_segment`、`batch`、`pipeline` 作为异步长任务执行；`info`、`audio_info`、`thumbnail` 保持同步返回。\n- job 快照新增 `created_by`、`intent`、`source`、`metadata`，方便 Claw 按聊天上下文展示结果。\n- HTTP 服务启动时会清理过旧终态 job，默认保留 7 天且最多 200 条。\n\n## 4.3.1 更新\n\n- 修正 HTTP chat 短等待行为：job 在 `wait_timeout_sec` 内完成时返回 `200` 和最终结果；仍在排队或运行时返回 `202`。\n- `/skill/jobs` 严格要求 `params` 为 JSON object；缺省或 `null` 使用 `{}`，数组、字符串等返回 `invalid_params`。\n- `async` 仅接受 JSON boolean 或 `\"auto\"`；`wait_timeout_sec` 必须是 `0-30` 秒数字。\n- 异步结果中的 `output_paths` 会按首次出现顺序去重。\n\n### 自然语言调用\n\nAction: `chat`\n\nClaw 推荐优先调用 `chat`。普通短命令可同步调用；压缩、字幕、pipeline 等长任务推荐 HTTP 调用时传入 `async:\"auto\"`。复杂、确定性要求高的多步骤流程继续使用 `pipeline`。\n\n```bash\npython run.py -a chat -i '{\"message\":\"将 \\\"sample.mp4\\\" 提取音频\"}'\npython run.py -a chat -i '{\"message\":\"给 \\\"sample.mp4\\\" 提取第 3 秒封面\"}'\npython run.py -a chat -i '{\"message\":\"压缩 \\\"sample.mp4\\\"\"}'\npython run.py -a chat -i '{\"message\":\"查看 \\\"sample.mp4\\\" 信息\"}'\npython run.py -a chat -i '{\"message\":\"识别 \\\"sample.mp4\\\" 的字幕\"}'\n```\n\nHTTP / Claw 长任务推荐：\n\n```bash\ncurl -X POST http://127.0.0.1:8080/skill/chat \\\n  -H 'Content-Type: application/json' \\\n  -d '{\"message\":\"识别 \\\"sample.mp4\\\" 的字幕\",\"async\":\"auto\"}'\n```\n\n返回会包含 `job_id`、`poll_url`、`job_path`，随后轮询 `poll_url` 获取 `reply`、`result` 和 `output_paths`。如需全部自然语言命令都进入 job，可传 `async:true`；如需短等待，可传 `wait_timeout_sec`。排队或运行中的异步响应返回 HTTP `202`，短等待内已完成的异步响应返回 HTTP `200`。\n\n处理绝对路径时需要配置媒体根目录：\n\n```bash\npython run.py -a chat -i '{\"message\":\"将 \\\"D:/AA.MP4\\\" 提取音频\",\"media_roots\":[\"D:/\"]}'\n```\n\n返回包含：\n\n| 字段 | 说明 |\n|------|------|\n| `reply` | 可直接展示给用户的聊天回复 |\n| `intent` | 识别出的意图 |\n| `action` | 实际调用的 action |\n| `params` | 传给底层 action 的参数 |\n| `result` | 底层 action 原始结果 |\n| `output_paths` | 本次生成的输出路径列表 |\n| `job_id` | 异步提交时的任务 id |\n| `poll_url` | 异步提交后的轮询地址 |\n\n### HTTP 异步长任务\n\nAction 可以继续同步调用，也可以通过 job API 异步执行。推荐对 `compress`、`asr`、`ocr`、`subtitle`、`pipeline` 等长任务使用异步接口。\n\n提交任务：\n\n```bash\ncurl -X POST http://127.0.0.1:8080/skill/jobs \\\n  -H 'Content-Type: application/json' \\\n  -d '{\"action\":\"pipeline\",\"params\":{\"source\":\"sample.mp4\",\"steps\":[{\"id\":\"metadata\",\"action\":\"info\",\"enabled\":true}]}}'\n```\n\n`params` 必须是 JSON object；不传或传 `null` 时按 `{}` 处理。\n\n返回：\n\n```json\n{\n  \"status\": \"queued\",\n  \"code\": \"ok\",\n  \"reply\": \"任务已提交：<job_id>\",\n  \"job_id\": \"<job_id>\",\n  \"job_path\": \"output/jobs/<job_id>/job.json\",\n  \"poll_url\": \"/skill/jobs/<job_id>\"\n}\n```\n\n轮询任务：\n\n```bash\ncurl http://127.0.0.1:8080/skill/jobs/<job_id>\n```\n\n任务状态包括：`queued`、`running`、`success`、`partial`、`skipped`、`error`。任务结果保存在 `result`，产物路径汇总在 `output_paths`。\n\n查询任务列表：\n\n```bash\ncurl 'http://127.0.0.1:8080/skill/jobs?status=success&limit=50'\n```\n\n任务文件存储在 `output/jobs/<job_id>/job.json`。服务重启后，已完成任务仍可查询；未完成的 `queued` / `running` 任务会标记为 `error`，`code=job_interrupted`。\n\n### 字幕识别\n\nAction: `subtitle`\n\n推荐使用 `subtitle`，默认 `mode=fusion`：ASR 负责主要时间轴和文本，OCR 做画面字幕校正。识别依赖随 `requirements.txt` 安装；如果环境未重新安装依赖，会返回 JSON error，不会抛未捕获异常。\n\n```bash\npython run.py -a subtitle -i '{\"source\":\"sample.mp4\",\"mode\":\"fusion\"}'\npython run.py -a subtitle -i '{\"source\":\"sample.mp4\",\"mode\":\"asr\",\"language\":\"zh\"}'\npython run.py -a subtitle -i '{\"source\":\"sample.mp4\",\"mode\":\"ocr\",\"sample_interval_sec\":1}'\n```\n\n参数：\n\n| 参数 | 类型 | 默认值 | 说明 |\n|------|------|--------|------|\n| `video_url` / `url` / `source` | string | 必填 | 远程 URL 或本地文件 |\n| `mode` | string | fusion | `asr` / `ocr` / `fusion` |\n| `language` | string | auto | ASR 语言，例：`zh` / `en` |\n| `model_size` | string | base | faster-whisper 模型规格 |\n| `sample_interval_sec` | number | 1.0 | OCR 抽帧间隔 |\n| `crop_bottom_ratio` | number | 0.35 | OCR 默认扫描画面下方比例 |\n| `output_path` | string | `output/subtitles/...` | 字幕 JSON 输出路径 |\n| `overwrite` | boolean | true | 是否覆盖已有文件 |\n\n字幕条目格式：\n\n```json\n{\n  \"captionTxt\": \"识别到的字幕文本\",\n  \"startTimeUs\": 1000000,\n  \"endTimeUs\": 2000000,\n  \"source\": \"asr\",\n  \"confidence\": 0.92\n}\n```\n\n底层调试入口：\n\n```bash\npython run.py -a asr -i '{\"source\":\"sample.mp4\",\"language\":\"zh\"}'\npython run.py -a ocr -i '{\"source\":\"sample.mp4\"}'\n```\n\n### 字幕二次分句\n\nAction: `caption_segment`\n\n`caption_segment` 是 emlet 字幕分句器：它不重新识别字幕，只处理已有 captions。默认 `max_chars=12`，输出仍是 SRT-like captions JSON。\n\n```bash\npython run.py -a caption_segment -i '{\n  \"caption_path\":\"output/subtitles/sample.captions.json\",\n  \"max_chars\":12,\n  \"protected_terms\":[\"苹果\",\"华为\",\"吉利\"]\n}'\n```\n\n也可以直接传 captions：\n\n```bash\npython run.py -a caption_segment -i '{\n  \"captions\":[\n    {\"captionTxt\":\"今天我们聊苹果华为和吉利的新产品\",\"startTimeUs\":0,\"endTimeUs\":3000000}\n  ],\n  \"max_chars\":12,\n  \"protected_terms\":\"苹果,华为,吉利\"\n}'\n```\n\n长期词库可以放在 JSON 文件中：\n\n```json\n{\n  \"brands\": [\"苹果\", \"华为\", \"吉利\"],\n  \"products\": [\"小米汽车\", \"Model Y\", \"ChatGPT\"]\n}\n```\n\n参数：\n\n| 参数 | 类型 | 默认值 | 说明 |\n|------|------|--------|------|\n| `caption_path` / `input_path` | string | - | 已有 captions JSON 文件 |\n| `captions` | array | - | 直接传入字幕条目 |\n| `max_chars` | integer | 12 | 单条字幕最大字符数 |\n| `protected_terms` | array/string | [] | 不拆开的品牌词、产品名、人名、术语 |\n| `protected_terms_path` | string | - | 当前工作目录内的保护词 JSON 文件 |\n| `auto_protect_ascii` | boolean | true | 自动保护英文、数字、型号、URL、路径 |\n| `output_path` | string | `output/subtitles/...` | 分句后字幕 JSON 输出路径 |\n| `overwrite` | boolean | true | 是否覆盖已有文件 |\n\n### 压缩视频\n\nAction: `compress`\n\n```bash\npython run.py -a compress -i '{\"source\":\"sample.mp4\",\"target_ratio\":0.1}'\n```\n\n参数：\n\n| 参数 | 类型 | 默认值 | 说明 |\n|------|------|--------|------|\n| `video_url` / `url` / `source` | string | 必填 | 远程 URL 或本地文件 |\n| `target_ratio` | number | 0.1 | 目标体积比例 |\n| `adaptive` | boolean | true | 是否自动尝试不同 CRF |\n| `crf` | integer | 24 | 非 adaptive 模式下使用 |\n| `preset` | string | veryfast | ffmpeg 编码预设 |\n| `output_path` | string | `output/videos/...` | 输出路径 |\n| `overwrite` | boolean | true | 是否覆盖已有文件 |\n\n### 提取封面\n\nAction: `thumbnail`\n\n当前支持 MP4/MOV 容器，主要适用于 H.264/H.265 视频轨道。\n\n```bash\npython run.py -a thumbnail -i '{\"source\":\"sample.mp4\",\"time_seconds\":5}'\n```\n\n参数：\n\n| 参数 | 类型 | 默认值 | 说明 |\n|------|------|--------|------|\n| `video_url` / `url` / `source` | string | 必填 | 远程 URL 或本地文件 |\n| `time_seconds` | number | 0 | 按时间点提取 |\n| `frame_number` | integer | - | 按帧号提取，优先于 `time_seconds` |\n| `save_path` | string | `output/thumbs/...` | 保存路径 |\n| `resize_width` | integer | - | 输出宽度 |\n| `quality` | integer | 85 | JPEG 质量 |\n| `overwrite` | boolean | true | 是否覆盖已有文件 |\n\n### 提取音频\n\nAction: `audio`\n\n```bash\npython run.py -a audio -i '{\"source\":\"sample.mp4\",\"format\":\"mp3\"}'\n```\n\n参数：\n\n| 参数 | 类型 | 默认值 | 说明 |\n|------|------|--------|------|\n| `video_url` / `url` / `source` | string | 必填 | 远程 URL 或本地文件 |\n| `format` | string | mp3 | mp3 / wav / aac / m4a |\n| `bitrate` | string | 128k | 音频比特率 |\n| `sample_rate` | integer | 44100 | 采样率 |\n| `channels` | integer | 2 | 声道数 |\n| `start_time` | number | - | 开始时间，秒 |\n| `duration` | number | - | 截取时长，秒 |\n| `output_path` | string | `output/audio/...` | 输出路径 |\n| `overwrite` | boolean | true | 是否覆盖已有文件 |\n\n### 批量音频\n\nAction: `audio_batch`\n\n```bash\npython run.py -a audio_batch -i '{\n  \"videos\":[\n    {\"source\":\"sample1.mp4\",\"name\":\"video1\"},\n    {\"url\":\"https://example.com/2.mp4\",\"name\":\"video2\"}\n  ],\n  \"output_dir\":\"output/audio\",\n  \"format\":\"mp3\"\n}'\n```\n\n### 批量处理\n\nAction: `batch`\n\n```bash\npython run.py -a batch -i '{\n  \"action\":\"thumbnail\",\n  \"videos\":[\n    {\"source\":\"sample1.mp4\",\"time_seconds\":5},\n    {\"url\":\"https://example.com/2.mp4\",\"time_seconds\":10}\n  ]\n}'\n```\n\n`action` 支持：`compress`、`thumbnail`、`audio`。\n\n### JSON 流水线\n\nAction: `pipeline`\n\n`steps` 是唯一流程控制入口，没有写进 `steps` 的动作不会执行。支持的 step action：`info`、`thumbnail`、`audio`、`compress`、`audio_info`、`asr`、`ocr`、`subtitle`、`caption_segment`。\n\n```bash\npython run.py -a pipeline -i '{\n  \"source\":\"sample.mp4\",\n  \"name\":\"sample\",\n  \"output_dir\":\"output/pipeline/sample\",\n  \"overwrite\":true,\n  \"steps\":[\n    {\"id\":\"metadata\",\"action\":\"info\",\"enabled\":true},\n    {\n      \"id\":\"cover\",\n      \"action\":\"thumbnail\",\n      \"enabled\":true,\n      \"params\":{\"time_seconds\":3,\"resize_width\":720}\n    },\n    {\n      \"id\":\"audio_mp3\",\n      \"action\":\"audio\",\n      \"enabled\":false,\n      \"params\":{\"format\":\"mp3\",\"bitrate\":\"128k\"}\n    }\n  ]\n}'\n```\n\n规则：\n\n| 参数 | 类型 | 默认值 | 说明 |\n|------|------|--------|------|\n| `video_url` / `url` / `source` | string | 必填 | 远程 URL 或本地文件 |\n| `name` | string | 输入文件名 | 流水线名称 |\n| `output_dir` | string | `output/pipeline/<name>` | manifest 和默认产物目录 |\n| `overwrite` | boolean | true | 是否覆盖已有产物 |\n| `steps` | array | 必填 | 按 JSON 顺序执行的步骤 |\n\n每个 step 需要 `id`、`action`、`enabled`。`enabled=false` 会记录为 `skipped`。每次执行都会生成 `manifest.json`。\n\n### 获取信息\n\n```bash\npython run.py -a info -i '{\"source\":\"sample.mp4\"}'\npython run.py -a audio_info -i '{\"source\":\"sample.mp4\"}'\n```\n\n## 注意\n\n- 远程输入仅支持 `http` / `https`。\n- 本地输入路径默认限制在当前工作目录内；通过 `media_roots` / `YM_MEDIA_ROOTS` 可授权额外媒体根目录。\n- 输出路径限制在当前工作目录内。\n- HTTP 服务默认只绑定本机地址。\n\nFile v4.2.2:_meta.json\n\n{\n  \"ownerId\": \"kn787sam7qk7fffsjc875yxp0984rpzh\",\n  \"slug\": \"ym-mediatoolkit\",\n  \"version\": \"4.2.2\",\n  \"publishedAt\": 1782377864696\n}\n\nFile v4.2.2:MAINTENANCE.md\n\n# YM MediaToolkit 维护手册\n\n本文档面向后续维护者，记录版本发布、测试验证、配置边界和排障流程。\n\n## 发布流程\n\n1. 同步版本号：\n   - `SKILL.md` front matter 的 `version`\n   - `skill.json` 的 `version`\n2. 更新用户文档：\n   - 新增 action 时同步更新 `SKILL.md` 和 `skill.json`\n   - 新增参数时同步补充默认值、用途和安全限制\n   - 变更行为时在 `SKILL.md` 的版本更新段落记录\n3. 运行验证：\n\n```bash\npython3 -B -m py_compile run.py utils.py intent_parser.py audio_extractor.py frame_extractor.py video_compressor.py asr_engine.py ocr_engine.py subtitle_extractor.py caption_segmenter.py job_manager.py tests/test_release_behaviors.py scripts/smoke_test.py\npython3 -m json.tool skill.json\npython3 -B -m unittest discover -s tests\npython3 scripts/smoke_test.py\n```\n\n4. 清理产物：\n\n```bash\nfind . -type d -name __pycache__ -prune -exec rm -r {} +\nfind . -name .DS_Store -delete\n```\n\n## 当前接口分层\n\n- `chat`：Claw 推荐入口，接收自然语言，解析后调用现有 action。\n- `pipeline`：确定性 JSON 流水线入口，适合多步骤、可重复的自动化流程。\n- `audio` / `thumbnail` / `compress` / `info` / `audio_info`：底层单步能力。\n- `subtitle` / `asr` / `ocr`：字幕识别能力，输出 SRT-like captions JSON。\n- `caption_segment`：emlet 字幕二次分句器，处理已有 captions，不重新识别媒体。\n- `batch` / `audio_batch`：批量处理入口。\n- `/skill/jobs`：HTTP 异步长任务入口，适合压缩、字幕识别、pipeline 等耗时 action。\n- `/skill/chat` + `async:\"auto\"`：Claw 推荐长任务入口，先解析自然语言，再自动提交 job。\n\n维护原则：新体验优先接到 `chat` 或 `pipeline`，媒体处理逻辑继续复用底层 action，避免重复实现 ffmpeg 调用。\n\n## Action 返回协议\n\n所有 action 必须返回 JSON 对象，并保留以下协议字段：\n\n- `status`：`success` / `partial` / `skipped` / `error`\n- `code`：稳定机器码，成功为 `ok`\n- `reply`：适合聊天展示的中文回复\n- `hint`：下一步建议或排障提示\n\n新增 handler 时只需要返回原始业务结果，`run.py` 的 action protocol wrapper 会补齐缺失字段。若底层模块已经返回 `code`，包装层会保留该值。\n\n常用错误码：\n\n| code | 场景 |\n|------|------|\n| `missing_source` | 缺少 `video_url` / `url` / `source` |\n| `source_not_allowed` | 本地路径不在当前工作目录或 `media_roots` 内 |\n| `output_exists` | 输出文件已存在且 `overwrite=false` |\n| `parse_failed` | `chat` 无法识别自然语言意图 |\n| `missing_steps` | `pipeline` 未传 `steps` |\n| `missing_captions` | `caption_segment` 未传 `captions` 或 `caption_path` |\n| `invalid_action` | 异步任务或 CLI 传入不支持的 action |\n| `invalid_params` | HTTP job 的 `params` 不是 JSON object |\n| `invalid_async_mode` | HTTP chat 的 `async` 不是 boolean 或 `\"auto\"` |\n| `invalid_wait_timeout` | HTTP chat 的 `wait_timeout_sec` 不是 `0-30` 秒数字 |\n| `invalid_job_id` | job id 格式不合法 |\n| `job_not_found` | job 文件不存在 |\n| `job_interrupted` | 服务重启或进程中断导致未完成 job 失效 |\n| `unsupported_action` | action 或 pipeline step 不支持 |\n| `invalid_step` | pipeline step 结构不合法 |\n| `ffmpeg_failed` | ffmpeg / ffprobe 缺失或执行失败 |\n| `missing_asr_dependency` | ASR 依赖缺失 |\n| `missing_ocr_dependency` | OCR 依赖缺失 |\n\n## HTTP 异步任务维护\n\n异步任务逻辑在 `job_manager.py`。HTTP 服务启动时会创建单 worker 队列，串行执行提交到 `/skill/jobs` 的 action。\n\n维护规则：\n\n- job 文件存储在 `output/jobs/<job_id>/job.json`。\n- job id 使用 32 位十六进制 UUID。\n- job 文件必须保留 `job_id`、`action`、`params`、`status`、`code`、`reply`、`hint`、`created_at`、`started_at`、`finished_at`、`result`、`output_paths`、`error`。\n- `chat` 提交的 job 会额外写入 `created_by=chat`、`intent`、`source`、`metadata.message`。\n- `/skill/jobs` 直接提交的 job 会写入 `created_by=jobs`。\n- 当前版本不支持取消任务，不记录进度百分比。\n- 服务重启后不恢复未完成任务；旧的 `queued` / `running` 会标记为 `error`，`code=job_interrupted`。\n- HTTP 服务启动时会调用 `cleanup_jobs(retention_days=7, max_jobs=200)`，只清理达到保留期或超量的终态 job，不清理 `queued` / `running`。\n- 新增 action 时，只要加入 `ACTIONS`，异步任务会自动支持。\n- `/skill/jobs` 的 `params` 必须是 JSON object；缺省或 `null` 会按 `{}` 处理，其他类型返回 `invalid_params` 且不创建 job。\n\n## HTTP chat 异步排障\n\n`/skill/chat` 支持 `async` 参数：\n\n- `false` 或不传：保持同步执行，兼容 CLI 和旧 HTTP 调用。\n- `true`：解析成功后总是提交 job，不立即执行底层 handler。\n- `\"auto\"`：只把 `audio`、`compress`、`asr`、`ocr`、`subtitle`、`caption_segment`、`batch`、`pipeline` 作为长任务提交；`info`、`audio_info`、`thumbnail` 继续同步。\n\n排查要点：\n\n- `async` 只接受 JSON boolean 或 `\"auto\"`；字符串 `\"true\"` / `\"false\"` 会返回 `invalid_async_mode`。\n- `wait_timeout_sec` 只接受 `0-30` 秒数字；非法值会返回 `invalid_wait_timeout`。\n- 异步任务排队或运行中返回 HTTP `202`；短等待内已完成时返回 HTTP `200` 和最终 job 结果。\n- 如果返回 `parse_failed`，说明自然语言没有解析出 action 或 source，不会创建 job。\n- 如果返回 `queued`，让调用方使用 `poll_url` 查询最终 `reply`、`result`、`output_paths`。\n- 如果任务完成后 `output_paths` 为空，检查底层 action 是否返回了 `output_path`、`saved_path`、`outputPath` 或 `manifest_path`。重复路径会按首次出现顺序去重。\n\n## media_roots 配置\n\n本地输入默认只允许当前工作目录内的文件。需要处理绝对路径时，必须配置媒体根目录白名单。\n\n请求级配置：\n\n```json\n{\n  \"message\": \"将 \\\"D:/AA.MP4\\\" 提取音频\",\n  \"media_roots\": [\"D:/\"]\n}\n```\n\n环境变量配置：\n\n```bash\nexport YM_MEDIA_ROOTS=\"/Users/me/Videos;/Volumes/Media\"\n```\n\n规则：\n\n- `media_roots` 优先级高于 `YM_MEDIA_ROOTS`\n- 未配置时只允许当前工作目录\n- 支持用逗号或分号分隔多个根目录\n- URL 仍只允许 `http` / `https`\n- 输出路径仍限制在当前工作目录内\n\n## 自然语言解析维护\n\n解析逻辑在 `intent_parser.py`。\n\n当前支持：\n\n- 提取音频：`将 \"sample.mp4\" 提取音频`\n- 提取封面：`给 \"sample.mp4\" 提取第 3 秒封面`\n- 压缩：`压缩 \"sample.mp4\"`\n- 查看信息：`查看 \"sample.mp4\" 信息`\n- JSON pipeline：消息本身是包含 `steps` 的 JSON\n- 字幕识别：`识别 \"sample.mp4\" 的字幕`\n\n新增自然语言规则时需要同时补：\n\n- `tests/test_release_behaviors.py` 的解析测试\n- `scripts/smoke_test.py` 的真实链路测试，若会产生文件\n- `SKILL.md` 的示例\n- `skill.json` 的 action schema 或 examples，若公开接口有变化\n\n## 验证策略\n\n单元测试覆盖轻量行为：\n\n- 输入字段兼容：`video_url` / `url` / `source`\n- 本地路径和 `media_roots` 权限\n- 默认输出目录和 `overwrite=false`\n- pipeline 顺序、跳过、失败继续、manifest\n- chat 解析、执行、失败不调用 handler\n\nSmoke test 覆盖真实 ffmpeg 链路：\n\n- 生成 1 秒本地测试视频\n- 跑 `info`、`thumbnail`、`audio`、`compress`、`batch`\n- 跑 `pipeline`\n- 跑 `chat` 的音频、封面、压缩、信息命令\n- 跑 `subtitle`；若环境未重新安装依赖，需返回 dependency error；安装完整依赖后验证 captions JSON\n\n## 字幕识别维护\n\n字幕输出统一使用 camelCase 和微秒整数：\n\n```json\n{\n  \"captionTxt\": \"字幕文本\",\n  \"startTimeUs\": 1000000,\n  \"endTimeUs\": 2000000,\n  \"source\": \"asr\",\n  \"confidence\": 0.92\n}\n```\n\n字幕识别依赖已进入默认 `requirements.txt`：\n\n```bash\npip install -r requirements.txt\n```\n\n维护规则：\n\n- `subtitle` 是推荐入口，`asr` / `ocr` 用于单独调试。\n- `mode=fusion` 默认 ASR 为主，OCR 只做高置信文本校正。\n- 环境未安装完整依赖时必须返回 JSON error，不能抛未捕获异常。\n- 新增字幕准确率策略时，优先补 `subtitle_extractor.py` 的纯函数测试。\n\n## emlet 字幕分句维护\n\n二次分句逻辑在 `caption_segmenter.py`，对已有 captions 做后处理，不调用 ASR/OCR。\n\n默认策略：\n\n- `max_chars=12`\n- 强标点优先切分：`。！？!?`\n- 弱标点其次：`，、；：,;:`\n- 再尝试连接词边界，最后按长度切分\n- `protected_terms` 和 `protected_terms_path` 内的词不允许被拆开\n- `auto_protect_ascii=true` 时自动保护英文、数字、型号、URL、路径\n\n维护规则：\n\n- 新增分句策略时先补纯函数测试，确保时间轴连续、不重叠、不倒退。\n- `caption_segment` 输出继续使用 `captionTxt`、`startTimeUs`、`endTimeUs`。\n- pipeline 中 `caption_segment` 如果没有显式传 `caption_path`，会自动使用前面步骤生成的 `.captions.json`。\n\n## 常见问题\n\n`本地输入路径超出允许的 media_roots`\n\n确认文件路径位于当前工作目录，或在请求中传入 `media_roots`。\n\n`输出路径超出工作目录`\n\n输出文件必须写入当前工作目录下，例如 `output/audio/demo.mp3`。\n\n`ffmpeg 错误，返回码: ...`\n\n先确认输入文件可播放，再用 smoke test 验证当前环境的 ffmpeg 是否可用。\n\n`DNS 解析失败` 或 `禁止访问私有/内网 IP`\n\n远程 URL 会做安全校验，不允许内网、回环、链路本地地址或无法验证的目标。\n\n`chat` 没识别出命令\n\n优先使用明确句式：`将 \"sample.mp4\" 提取音频`、`给 \"sample.mp4\" 提取第 3 秒封面`、`压缩 \"sample.mp4\"`。\n\n`缺少 ASR/OCR 依赖`\n\n重新运行 `pip install -r requirements.txt`，确认 `faster-whisper`、`paddlepaddle`、`paddleocr` 已安装。\n\nFile v4.2.2:skill-card.md\n\n## Description:\n\n自然语言媒体助手 - 视频压缩、MP4/MOV 封面提取、音频转换、字幕识别。\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[370299455cx-web](https://clawhub.ai/user/370299455cx-web)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers and external users use this skill to turn natural-language or JSON requests into media-processing tasks such as video compression, thumbnail extraction, audio conversion, subtitle recognition, batch processing, and asynchronous job polling.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: The unauthenticated local HTTP API can process local files and remote URLs.\n\nMitigation: Run it only in a trusted local environment and do not expose the HTTP server to a network.\n\nRisk: Changing the bind host to 0.0.0.0 can make the service reachable beyond localhost.\n\nMitigation: Use authentication, restricted CORS, and a controlled reverse proxy before any non-local deployment.\n\nRisk: Accepting media_roots from untrusted callers can expand local file access.\n\nMitigation: Do not accept media_roots from untrusted callers and restrict configured media roots to approved directories.\n\nRisk: Media processing and model dependencies can consume significant compute, memory, disk, and network resources.\n\nMitigation: Use pinned dependencies, trusted media sources, quota controls, and resource limits before production use.\n\n## Reference(s):\n\n- [ClawHub skill page](https://clawhub.ai/370299455cx-web/skills/ym-mediatoolkit)\n- [Publisher profile](https://clawhub.ai/user/370299455cx-web)\n\n## Skill Output:\n\n**Output Type(s):** [text, markdown, code, shell commands, configuration, guidance, files]\n\n**Output Format:** [Markdown guidance, shell commands, JSON responses, and generated media or caption files]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [Outputs can include stable status/code/reply/hint fields, media file paths, caption JSON, pipeline manifests, and asynchronous job metadata.]\n\n## Skill Version(s):\n\n4.2.2 (source: server release metadata)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nFile v4.2.2:output/jobs/c65f1520dca6436c87f1028354263718/job.json\n\n{\n  \"job_id\": \"c65f1520dca6436c87f1028354263718\",\n  \"action\": \"info\",\n  \"params\": {\n    \"source\": \"sample.mp4\",\n    \"name\": \"sample\"\n  },\n  \"created_by\": \"chat\",\n  \"intent\": \"info\",\n  \"source\": \"sample.mp4\",\n  \"metadata\": {\n    \"created_by\": \"chat\",\n    \"intent\": \"info\",\n    \"source\": \"sample.mp4\",\n    \"message\": \"查看 \\\"sample.mp4\\\" 信息\"\n  },\n  \"status\": \"success\",\n  \"code\": \"ok\",\n  \"reply\": \"ok\",\n  \"hint\": \"ok\",\n  \"created_at\": \"2026-06-25T02:22:13.021739+00:00\",\n  \"started_at\": \"2026-06-25T02:22:13.022226+00:00\",\n  \"finished_at\": \"2026-06-25T02:22:13.022565+00:00\",\n  \"result\": {\n    \"status\": \"success\",\n    \"code\": \"ok\",\n    \"reply\": \"ok\",\n    \"hint\": \"ok\",\n    \"output_path\": \"out/a.json\",\n    \"nested\": {\n      \"outputPath\": \"out/a.json\"\n    }\n  },\n  \"output_paths\": [\n    \"out/a.json\"\n  ],\n  \"error\": null\n}\n\nFile v4.2.2:skill.json\n\n{\n  \"name\": \"ym-mediatoolkit\",\n  \"version\": \"4.3.1\",\n  \"description\": \"自然语言媒体助手：视频压缩、MP4/MOV 封面提取、音频转换、字幕识别、JSON 流水线\",\n  \"author\": \"your_name\",\n  \"entrypoint\": \"python run.py --input {input_json}\",\n  \"http_port\": 8080,\n  \"http_bind\": \"127.0.0.1 (默认，仅本地；可通过 --host 0.0.0.0 改为公网，需反向代理认证)\",\n  \"external_binaries\": {\n    \"ffmpeg\": \"必需 - 视频压缩、音频提取、流式处理\",\n    \"ffprobe\": \"必需 - 获取视频/音频流信息\"\n  },\n  \"python_dependencies\": {\n    \"faster-whisper\": \"必需 - ASR 字幕识别\",\n    \"paddlepaddle\": \"必需 - PaddleOCR 推理运行时\",\n    \"paddleocr\": \"必需 - OCR 字幕识别\"\n  },\n  \"response_protocol\": {\n    \"description\": \"所有 action 都会返回稳定协议字段，并保留原有业务字段。\",\n    \"fields\": {\n      \"status\": \"success / partial / skipped / error\",\n      \"code\": \"稳定机器码，例如 ok、missing_source、output_exists、parse_failed\",\n      \"reply\": \"适合聊天展示的简短中文回复\",\n      \"hint\": \"面向调用方或用户的下一步建议\"\n    },\n    \"common_codes\": [\n      \"ok\",\n      \"missing_source\",\n      \"source_not_allowed\",\n      \"output_exists\",\n      \"parse_failed\",\n      \"missing_steps\",\n      \"missing_captions\",\n      \"invalid_action\",\n      \"invalid_params\",\n      \"invalid_json\",\n      \"invalid_async_mode\",\n      \"invalid_wait_timeout\",\n      \"invalid_job_id\",\n      \"job_not_found\",\n      \"job_interrupted\",\n      \"job_failed\",\n      \"unsupported_action\",\n      \"invalid_step\",\n      \"ffmpeg_failed\",\n      \"missing_asr_dependency\",\n      \"missing_ocr_dependency\"\n    ]\n  },\n  \"http_apis\": [\n    {\n      \"path\": \"/skill/jobs\",\n      \"method\": \"POST\",\n      \"description\": \"提交异步长任务，现有同步 /skill/<action> 接口保持不变\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"required\": [\"action\"],\n        \"properties\": {\n          \"action\": {\n            \"type\": \"string\",\n            \"description\": \"当前 actions 中支持的 action 名称\"\n          },\n          \"params\": {\n            \"type\": \"object\",\n            \"default\": {},\n            \"description\": \"传给 action 的参数；缺省或 null 按空对象处理，数组、字符串等会返回 invalid_params\"\n          }\n        }\n      },\n      \"response_schema\": {\n        \"type\": \"object\",\n        \"properties\": {\n          \"status\": {\"type\": \"string\", \"enum\": [\"queued\", \"error\"]},\n          \"code\": {\"type\": \"string\"},\n          \"reply\": {\"type\": \"string\"},\n          \"job_id\": {\"type\": \"string\"},\n          \"job_path\": {\"type\": \"string\"},\n          \"poll_url\": {\"type\": \"string\"},\n          \"created_by\": {\"type\": \"string\"},\n          \"intent\": {\"type\": \"string\"},\n          \"params\": {\"type\": \"object\"}\n        }\n      }\n    },\n    {\n      \"path\": \"/skill/jobs/<job_id>\",\n      \"method\": \"GET\",\n      \"description\": \"查询单个异步任务状态、结果和输出路径\"\n    },\n    {\n      \"path\": \"/skill/jobs\",\n      \"method\": \"GET\",\n      \"description\": \"查询异步任务列表，支持 status 和 limit query 参数\"\n    }\n  ],\n  \"actions\": [\n    {\n      \"name\": \"chat\",\n      \"description\": \"自然语言媒体助手入口，将用户聊天文本解析为媒体处理 action；HTTP 可用 async 自动提交长任务 job\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"required\": [\"message\"],\n        \"properties\": {\n          \"message\": {\n            \"type\": \"string\",\n            \"description\": \"自然语言命令，例如：将 \\\"D:/AA.MP4\\\" 提取音频\"\n          },\n          \"media_roots\": {\n            \"type\": [\"array\", \"string\"],\n            \"description\": \"允许访问的本地媒体根目录；未传时读取 YM_MEDIA_ROOTS，否则仅允许当前工作目录\"\n          },\n          \"overwrite\": {\"type\": \"boolean\", \"default\": true},\n          \"output_dir\": {\"type\": \"string\", \"description\": \"自然语言命令的默认输出目录\"},\n          \"async\": {\n            \"oneOf\": [\n              {\"type\": \"boolean\"},\n              {\"type\": \"string\", \"enum\": [\"auto\"]}\n            ],\n            \"default\": false,\n            \"description\": \"HTTP chat 异步控制；仅接受 JSON boolean 或字符串 auto，true 全部转 job，auto 仅长任务转 job\"\n          },\n          \"wait_timeout_sec\": {\n            \"type\": \"number\",\n            \"minimum\": 0,\n            \"maximum\": 30,\n            \"default\": 0,\n            \"description\": \"异步提交后最多短等待秒数；范围 0-30，任务很快完成时可直接返回 HTTP 200 最终结果\"\n          }\n        }\n      },\n      \"examples\": [\n        {\"message\": \"将 \\\"sample.mp4\\\" 提取音频\"},\n        {\"message\": \"给 \\\"sample.mp4\\\" 提取第 3 秒封面\"},\n        {\"message\": \"压缩 \\\"sample.mp4\\\"\"},\n        {\"message\": \"查看 \\\"sample.mp4\\\" 信息\"},\n        {\"message\": \"识别 \\\"sample.mp4\\\" 的字幕\", \"async\": \"auto\"}\n      ]\n    },\n    {\n      \"name\": \"compress\",\n      \"description\": \"流式压缩视频，保持清晰度，支持自适应 CRF\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"anyOf\": [\n          {\"required\": [\"video_url\"]},\n          {\"required\": [\"url\"]},\n          {\"required\": [\"source\"]}\n        ],\n        \"properties\": {\n          \"video_url\": {\"type\": \"string\", \"description\": \"视频 URL 或本地文件路径\"},\n          \"url\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"source\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"media_roots\": {\"type\": [\"array\", \"string\"], \"description\": \"允许访问的本地媒体根目录\"},\n          \"target_ratio\": {\"type\": \"number\", \"default\": 0.1},\n          \"adaptive\": {\"type\": \"boolean\", \"default\": true},\n          \"crf\": {\"type\": \"integer\", \"default\": 24},\n          \"preset\": {\"type\": \"string\", \"default\": \"veryfast\"},\n          \"output_path\": {\"type\": \"string\", \"description\": \"默认 output/videos\"},\n          \"overwrite\": {\"type\": \"boolean\", \"default\": true}\n        }\n      }\n    },\n    {\n      \"name\": \"thumbnail\",\n      \"description\": \"从 MP4/MOV 视频任意时间点或帧号提取封面，流式只下载必要部分\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"anyOf\": [\n          {\"required\": [\"video_url\"]},\n          {\"required\": [\"url\"]},\n          {\"required\": [\"source\"]}\n        ],\n        \"properties\": {\n          \"video_url\": {\"type\": \"string\", \"description\": \"视频 URL 或本地文件路径\"},\n          \"url\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"source\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"media_roots\": {\"type\": [\"array\", \"string\"], \"description\": \"允许访问的本地媒体根目录\"},\n          \"time_seconds\": {\"type\": \"number\"},\n          \"frame_number\": {\"type\": \"integer\"},\n          \"save_path\": {\"type\": \"string\"},\n          \"resize_width\": {\"type\": \"integer\"},\n          \"quality\": {\"type\": \"integer\", \"default\": 85},\n          \"overwrite\": {\"type\": \"boolean\", \"default\": true}\n        }\n      }\n    },\n    {\n      \"name\": \"audio\",\n      \"description\": \"流式提取音频，支持 MP3/WAV/AAC/M4A 格式\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"anyOf\": [\n          {\"required\": [\"video_url\"]},\n          {\"required\": [\"url\"]},\n          {\"required\": [\"source\"]}\n        ],\n        \"properties\": {\n          \"video_url\": {\"type\": \"string\", \"description\": \"视频 URL 或本地文件路径\"},\n          \"url\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"source\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"media_roots\": {\"type\": [\"array\", \"string\"], \"description\": \"允许访问的本地媒体根目录\"},\n          \"format\": {\"type\": \"string\", \"enum\": [\"mp3\", \"wav\", \"aac\", \"m4a\"], \"default\": \"mp3\"},\n          \"bitrate\": {\"type\": \"string\", \"default\": \"128k\", \"description\": \"比特率: 128k, 192k, 320k\"},\n          \"sample_rate\": {\"type\": \"integer\", \"default\": 44100, \"description\": \"采样率: 44100, 48000\"},\n          \"channels\": {\"type\": \"integer\", \"default\": 2, \"description\": \"声道: 1=单声道, 2=立体声\"},\n          \"start_time\": {\"type\": \"number\", \"description\": \"开始时间（秒）\"},\n          \"duration\": {\"type\": \"number\", \"description\": \"持续时间（秒）\"},\n          \"output_path\": {\"type\": \"string\", \"description\": \"默认 output/audio\"},\n          \"overwrite\": {\"type\": \"boolean\", \"default\": true}\n        }\n      }\n    },\n    {\n      \"name\": \"audio_batch\",\n      \"description\": \"批量提取多个视频的音频\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"required\": [\"videos\"],\n        \"properties\": {\n          \"videos\": {\"type\": \"array\", \"description\": \"视频列表，条目支持 video_url/url/source\"},\n          \"media_roots\": {\"type\": [\"array\", \"string\"], \"description\": \"允许访问的本地媒体根目录\"},\n          \"output_dir\": {\"type\": \"string\", \"default\": \"output/audio\"},\n          \"format\": {\"type\": \"string\", \"default\": \"mp3\"},\n          \"bitrate\": {\"type\": \"string\", \"default\": \"128k\"},\n          \"sample_rate\": {\"type\": \"integer\", \"default\": 44100},\n          \"overwrite\": {\"type\": \"boolean\", \"default\": true}\n        }\n      }\n    },\n    {\n      \"name\": \"audio_info\",\n      \"description\": \"获取视频的音频流信息\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"anyOf\": [\n          {\"required\": [\"video_url\"]},\n          {\"required\": [\"url\"]},\n          {\"required\": [\"source\"]}\n        ],\n        \"properties\": {\n          \"video_url\": {\"type\": \"string\", \"description\": \"视频 URL 或本地文件路径\"},\n          \"url\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"source\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"media_roots\": {\"type\": [\"array\", \"string\"], \"description\": \"允许访问的本地媒体根目录\"}\n        }\n      }\n    },\n    {\n      \"name\": \"asr\",\n      \"description\": \"从音频轨识别字幕，输出 SRT-like captions JSON\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"anyOf\": [\n          {\"required\": [\"video_url\"]},\n          {\"required\": [\"url\"]},\n          {\"required\": [\"source\"]}\n        ],\n        \"properties\": {\n          \"video_url\": {\"type\": \"string\", \"description\": \"视频 URL 或本地文件路径\"},\n          \"url\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"source\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"language\": {\"type\": \"string\", \"default\": \"auto\"},\n          \"model_size\": {\"type\": \"string\", \"default\": \"base\"},\n          \"output_path\": {\"type\": \"string\", \"description\": \"默认 output/subtitles/<name>.captions.json\"},\n          \"overwrite\": {\"type\": \"boolean\", \"default\": true},\n          \"media_roots\": {\"type\": [\"array\", \"string\"], \"description\": \"允许访问的本地媒体根目录\"}\n        }\n      }\n    },\n    {\n      \"name\": \"ocr\",\n      \"description\": \"从视频画面识别硬字幕或屏幕文字，输出 SRT-like captions JSON\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"anyOf\": [\n          {\"required\": [\"video_url\"]},\n          {\"required\": [\"url\"]},\n          {\"required\": [\"source\"]}\n        ],\n        \"properties\": {\n          \"video_url\": {\"type\": \"string\", \"description\": \"视频 URL 或本地文件路径\"},\n          \"url\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"source\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"sample_interval_sec\": {\"type\": \"number\", \"default\": 1.0},\n          \"crop_bottom_ratio\": {\"type\": \"number\", \"default\": 0.35},\n          \"output_path\": {\"type\": \"string\", \"description\": \"默认 output/subtitles/<name>.captions.json\"},\n          \"overwrite\": {\"type\": \"boolean\", \"default\": true},\n          \"media_roots\": {\"type\": [\"array\", \"string\"], \"description\": \"允许访问的本地媒体根目录\"}\n        }\n      }\n    },\n    {\n      \"name\": \"subtitle\",\n      \"description\": \"推荐字幕识别入口，支持 asr / ocr / fusion 模式，统一输出 captions JSON\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"anyOf\": [\n          {\"required\": [\"video_url\"]},\n          {\"required\": [\"url\"]},\n          {\"required\": [\"source\"]}\n        ],\n        \"properties\": {\n          \"video_url\": {\"type\": \"string\", \"description\": \"视频 URL 或本地文件路径\"},\n          \"url\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"source\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"mode\": {\"type\": \"string\", \"enum\": [\"asr\", \"ocr\", \"fusion\"], \"default\": \"fusion\"},\n          \"language\": {\"type\": \"string\", \"default\": \"auto\"},\n          \"model_size\": {\"type\": \"string\", \"default\": \"base\"},\n          \"sample_interval_sec\": {\"type\": \"number\", \"default\": 1.0},\n          \"crop_bottom_ratio\": {\"type\": \"number\", \"default\": 0.35},\n          \"output_path\": {\"type\": \"string\", \"description\": \"默认 output/subtitles/<name>.captions.json\"},\n          \"overwrite\": {\"type\": \"boolean\", \"default\": true},\n          \"media_roots\": {\"type\": [\"array\", \"string\"], \"description\": \"允许访问的本地媒体根目录\"}\n        }\n      },\n      \"examples\": [\n        {\"source\": \"sample.mp4\", \"mode\": \"fusion\"},\n        {\"source\": \"sample.mp4\", \"mode\": \"asr\", \"language\": \"zh\"}\n      ]\n    },\n    {\n      \"name\": \"caption_segment\",\n      \"description\": \"emlet 字幕二次分句器，对已有 captions JSON 按长度、标点和保护词重新切分\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"anyOf\": [\n          {\"required\": [\"caption_path\"]},\n          {\"required\": [\"input_path\"]},\n          {\"required\": [\"captions\"]}\n        ],\n        \"properties\": {\n          \"caption_path\": {\"type\": \"string\", \"description\": \"已有 captions JSON 文件路径\"},\n          \"input_path\": {\"type\": \"string\", \"description\": \"同 caption_path\"},\n          \"captions\": {\"type\": \"array\", \"description\": \"直接传入的字幕条目\"},\n          \"max_chars\": {\"type\": \"integer\", \"default\": 12, \"description\": \"单条字幕最大字符数\"},\n          \"protected_terms\": {\n            \"type\": [\"array\", \"string\"],\n            \"description\": \"不拆开的品牌词、产品名、人名、术语，例如 苹果,华为,吉利\"\n          },\n          \"protected_terms_path\": {\n            \"type\": \"string\",\n            \"description\": \"当前工作目录内的保护词 JSON 文件，可按 brands/products 等字段分组\"\n          },\n          \"auto_protect_ascii\": {\n            \"type\": \"boolean\",\n            \"default\": true,\n            \"description\": \"自动保护英文、数字、型号、URL、路径\"\n          },\n          \"output_path\": {\"type\": \"string\", \"description\": \"默认 output/subtitles/<name>.segmented.captions.json\"},\n          \"overwrite\": {\"type\": \"boolean\", \"default\": true}\n        }\n      },\n      \"examples\": [\n        {\n          \"caption_path\": \"output/subtitles/sample.captions.json\",\n          \"max_chars\": 12,\n          \"protected_terms\": [\"苹果\", \"华为\", \"吉利\"]\n        }\n      ]\n    },\n    {\n      \"name\": \"batch\",\n      \"description\": \"批量处理多个视频，支持 compress / thumbnail / audio\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"required\": [\"videos\"],\n        \"properties\": {\n          \"action\": {\"type\": \"string\", \"enum\": [\"compress\", \"thumbnail\", \"audio\"], \"default\": \"thumbnail\"},\n          \"videos\": {\"type\": \"array\", \"description\": \"按 action 传入对应参数对象列表，条目支持 video_url/url/source\"},\n          \"media_roots\": {\"type\": [\"array\", \"string\"], \"description\": \"允许访问的本地媒体根目录\"},\n          \"output_dir\": {\"type\": \"string\", \"description\": \"批量输出目录\"},\n          \"overwrite\": {\"type\": \"boolean\", \"default\": true}\n        }\n      }\n    },\n    {\n      \"name\": \"pipeline\",\n      \"description\": \"JSON 驱动媒体流水线，按 steps 顺序编排 info / thumbnail / audio / compress / audio_info / subtitle / caption_segment\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"anyOf\": [\n          {\"required\": [\"video_url\"]},\n          {\"required\": [\"url\"]},\n          {\"required\": [\"source\"]}\n        ],\n        \"required\": [\"steps\"],\n        \"properties\": {\n          \"video_url\": {\"type\": \"string\", \"description\": \"视频 URL 或本地文件路径\"},\n          \"url\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"source\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"media_roots\": {\"type\": [\"array\", \"string\"], \"description\": \"允许访问的本地媒体根目录\"},\n          \"name\": {\"type\": \"string\", \"description\": \"流水线名称\"},\n          \"output_dir\": {\"type\": \"string\", \"description\": \"默认 output/pipeline/<name>\"},\n          \"overwrite\": {\"type\": \"boolean\", \"default\": true},\n          \"steps\": {\n            \"type\": \"array\",\n            \"description\": \"步骤列表，每步包含 id/action/enabled/params\",\n            \"items\": {\n              \"type\": \"object\",\n              \"required\": [\"id\", \"action\", \"enabled\"],\n              \"properties\": {\n                \"id\": {\"type\": \"string\"},\n                \"action\": {\"type\": \"string\", \"enum\": [\"info\", \"thumbnail\", \"audio\", \"compress\", \"audio_info\", \"asr\", \"ocr\", \"subtitle\", \"caption_segment\"]},\n                \"enabled\": {\"type\": \"boolean\"},\n                \"params\": {\"type\": \"object\"}\n              }\n            }\n          }\n        }\n      }\n    },\n    {\n      \"name\": \"info\",\n      \"description\": \"获取完整视频信息（分辨率、时长、编码等）\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"anyOf\": [\n          {\"required\": [\"video_url\"]},\n          {\"required\": [\"url\"]},\n          {\"required\": [\"source\"]}\n        ],\n        \"properties\": {\n          \"video_url\": {\"type\": \"string\", \"description\": \"视频 URL 或本地文件路径\"},\n          \"url\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"source\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"media_roots\": {\"type\": [\"array\", \"string\"], \"description\": \"允许访问的本地媒体根目录\"}\n        }\n      }\n    }\n  ]\n}\n\nFile v4.2.2:requirements.txt\n\nrequests>=2.28.0\nopencv-python>=4.8.0\nnumpy>=1.24.0\naiohttp>=3.8.0\nflask>=2.3.0\nflask-cors>=4.0.0\nfaster-whisper>=1.0.0\npaddlepaddle>=2.6.0\npaddleocr>=2.7.0\n\nArchive v4.2.1: 19 files, 61062 bytes\n\nFiles: asr_engine.py (3005b), audio_extractor.py (9638b), caption_segmenter.py (8548b), frame_extractor.py (22227b), intent_parser.py (6678b), job_manager.py (11280b), MAINTENANCE.md (9258b), ocr_engine.py (3045b), requirements.txt (157b), run.py (42629b), scripts/smoke_test.py (8721b), skill-card.md (2387b), skill.json (17848b), SKILL.md (15750b), subtitle_extractor.py (5490b), tests/test_release_behaviors.py (36985b), utils.py (8193b), video_compressor.py (6897b), _meta.json (134b)\n\nFile v4.2.1:SKILL.md\n\n---\nname: ym-mediatoolkit\nversion: 4.3.0\ndescription: 自然语言媒体助手 - 视频压缩、MP4/MOV 封面提取、音频转换、字幕识别\nauthor: your_name\ntags:\n  - video\n  - compression\n  - thumbnail\n  - audio\n  - streaming\n  - ffmpeg\ncategories:\n  - media\n  - utility\nclawhub:\n  entrypoint: python run.py\n  runtime: python3\n  http_port: 8080\n---\n\n# YM MediaToolkit\n\n自然语言媒体助手，支持从远程视频 URL、当前工作目录内的本地视频文件、或配置过 `media_roots` 的本地媒体目录直接处理：\n\n- 自然语言调用\n- 视频压缩\n- MP4/MOV 封面提取\n- 音频提取与转换：MP3 / WAV / AAC / M4A\n- OCR / ASR 字幕识别\n- emlet 字幕二次分句\n- 批量处理\n- JSON 驱动媒体流水线\n\n## 依赖\n\n```bash\npip install -r requirements.txt\n```\n\n系统需要安装：\n\n```bash\nffmpeg\nffprobe\n```\n\n字幕识别依赖已内置在 `requirements.txt`，包括 `faster-whisper`、`paddlepaddle`、`paddleocr`。\n\n## 维护\n\n维护、发布、测试和排障流程见 [MAINTENANCE.md](./MAINTENANCE.md)。\n\n## 命令行\n\n```bash\npython run.py -a <action> -i '<json>'\n```\n\n也可以从 JSON 文件读取参数：\n\n```bash\npython run.py -i params.json\n```\n\n## HTTP 服务\n\n```bash\npython run.py --serve\n```\n\n默认监听 `127.0.0.1:8080`。\n\n健康检查：\n\n```http\nGET /health\n```\n\n异步长任务：\n\n```http\nPOST /skill/jobs\nGET /skill/jobs/<job_id>\nGET /skill/jobs\n```\n\n## 功能\n\n输入源字段可使用 `video_url`、`url` 或 `source`，三者等价。远程输入仅支持 `http/https`，本地输入默认限制在当前工作目录内；需要访问绝对路径时，通过请求参数 `media_roots` 或环境变量 `YM_MEDIA_ROOTS` 配置允许的媒体根目录。\n\n## 返回协议\n\n从 `4.1.0` 开始，所有 action 都会返回稳定协议字段：\n\n| 字段 | 说明 |\n|------|------|\n| `status` | `success` / `partial` / `skipped` / `error` |\n| `code` | 稳定机器码，例如 `ok`、`missing_source`、`output_exists`、`parse_failed` |\n| `reply` | 适合聊天展示的简短中文回复 |\n| `hint` | 面向调用方或用户的下一步建议 |\n\n原有业务字段会继续保留，例如 `output_path`、`saved_path`、`outputPath`、`manifest_path`、`info`、`captions`、`result`。\n\n默认输出目录：\n\n| 类型 | 默认目录 |\n|------|----------|\n| 压缩视频 | `output/videos` |\n| 音频 | `output/audio` |\n| 封面 | `output/thumbs` |\n\n所有会写文件的接口都支持 `overwrite`，默认 `true`。设置为 `false` 时，如果输出文件已存在会直接返回错误。\n\n## 3.0.3 更新\n\n- 支持当前工作目录内的本地视频文件输入。\n- 统一 `video_url` / `url` / `source` 三种输入字段。\n- 增加默认输出目录：`output/videos`、`output/audio`、`output/thumbs`。\n- 增加 `overwrite` 覆盖策略，避免误覆盖已有文件。\n\n## 4.0.0 更新\n\n- 新增自然语言入口 `chat`，适合 Claw 直接转发用户聊天文本。\n- 新增 `media_roots` 白名单，支持处理授权目录内的绝对路径文件。\n- 自然语言命令支持提取音频、提取封面、压缩、查看信息、JSON 流水线。\n- `chat` 返回 `reply` 和结构化 `result`，同时兼顾聊天展示和自动化消费。\n\n## 4.0.1 更新\n\n- 新增 `subtitle` 推荐入口，支持 `asr` / `ocr` / `fusion` 模式。\n- 新增 `asr` 和 `ocr` 单独调试入口。\n- 字幕统一输出 SRT-like JSON：`captionTxt`、`startTimeUs`、`endTimeUs`、`source`、`confidence`。\n- `chat` 支持“识别字幕 / 提取字幕 / 转字幕 / 生成字幕”等自然语言命令。\n\n## 4.0.2 更新\n\n- 将 `faster-whisper`、`paddlepaddle`、`paddleocr` 纳入默认 `requirements.txt`。\n- 字幕识别从“可选依赖”调整为默认安装能力。\n- 运行时仍保留缺依赖 JSON error，方便定位未重新安装依赖的环境。\n\n## 4.1.0 更新\n\n- 所有 action 统一补齐 `code`、`reply`、`hint`，方便 Claw 和后续渠道适配。\n- 增加稳定错误码：`missing_source`、`source_not_allowed`、`output_exists`、`parse_failed`、`missing_steps`、`unsupported_action`、`ffmpeg_failed`、`missing_asr_dependency`、`missing_ocr_dependency`。\n- 保持旧字段兼容，不改变现有 action 名称、HTTP endpoint 和底层媒体处理逻辑。\n\n## 4.1.1 更新\n\n- 新增 `caption_segment`，用于 emlet 字幕二次分句。\n- 默认每句最多 `12` 个字符，按强标点、弱标点、连接词和长度切分。\n- 新增 `protected_terms` 和 `protected_terms_path`，用于保护品牌词、产品名、人名、术语不被拆开。\n- `pipeline` 支持 `subtitle -> caption_segment` 串联；分句步骤未传 `caption_path` 时会自动使用上一步字幕 JSON。\n\n## 4.2.0 更新\n\n- 新增 HTTP 异步长任务接口 `/skill/jobs`，适合压缩、ASR/OCR、字幕和 pipeline 等耗时 action。\n- 任务状态持久化到 `output/jobs/<job_id>/job.json`，支持提交、轮询和列表查询。\n- 现有 `/skill/<action>` 同步接口保持不变；CLI 仍保持同步执行。\n\n## 4.3.0 更新\n\n- `chat` 增加 HTTP 异步闭环：`async=true` 会把识别出的 action 提交为 job。\n- `async=\"auto\"` 会自动将 `audio`、`compress`、`asr`、`ocr`、`subtitle`、`caption_segment`、`batch`、`pipeline` 作为异步长任务执行；`info`、`audio_info`、`thumbnail` 保持同步返回。\n- job 快照新增 `created_by`、`intent`、`source`、`metadata`，方便 Claw 按聊天上下文展示结果。\n- HTTP 服务启动时会清理过旧终态 job，默认保留 7 天且最多 200 条。\n\n### 自然语言调用\n\nAction: `chat`\n\nClaw 推荐优先调用 `chat`。普通短命令可同步调用；压缩、字幕、pipeline 等长任务推荐 HTTP 调用时传入 `async:\"auto\"`。复杂、确定性要求高的多步骤流程继续使用 `pipeline`。\n\n```bash\npython run.py -a chat -i '{\"message\":\"将 \\\"sample.mp4\\\" 提取音频\"}'\npython run.py -a chat -i '{\"message\":\"给 \\\"sample.mp4\\\" 提取第 3 秒封面\"}'\npython run.py -a chat -i '{\"message\":\"压缩 \\\"sample.mp4\\\"\"}'\npython run.py -a chat -i '{\"message\":\"查看 \\\"sample.mp4\\\" 信息\"}'\npython run.py -a chat -i '{\"message\":\"识别 \\\"sample.mp4\\\" 的字幕\"}'\n```\n\nHTTP / Claw 长任务推荐：\n\n```bash\ncurl -X POST http://127.0.0.1:8080/skill/chat \\\n  -H 'Content-Type: application/json' \\\n  -d '{\"message\":\"识别 \\\"sample.mp4\\\" 的字幕\",\"async\":\"auto\"}'\n```\n\n返回会包含 `job_id`、`poll_url`、`job_path`，随后轮询 `poll_url` 获取 `reply`、`result` 和 `output_paths`。如需全部自然语言命令都进入 job，可传 `async:true`；如需短等待，可传 `wait_timeout_sec`。\n\n处理绝对路径时需要配置媒体根目录：\n\n```bash\npython run.py -a chat -i '{\"message\":\"将 \\\"D:/AA.MP4\\\" 提取音频\",\"media_roots\":[\"D:/\"]}'\n```\n\n返回包含：\n\n| 字段 | 说明 |\n|------|------|\n| `reply` | 可直接展示给用户的聊天回复 |\n| `intent` | 识别出的意图 |\n| `action` | 实际调用的 action |\n| `params` | 传给底层 action 的参数 |\n| `result` | 底层 action 原始结果 |\n| `output_paths` | 本次生成的输出路径列表 |\n| `job_id` | 异步提交时的任务 id |\n| `poll_url` | 异步提交后的轮询地址 |\n\n### HTTP 异步长任务\n\nAction 可以继续同步调用，也可以通过 job API 异步执行。推荐对 `compress`、`asr`、`ocr`、`subtitle`、`pipeline` 等长任务使用异步接口。\n\n提交任务：\n\n```bash\ncurl -X POST http://127.0.0.1:8080/skill/jobs \\\n  -H 'Content-Type: application/json' \\\n  -d '{\"action\":\"pipeline\",\"params\":{\"source\":\"sample.mp4\",\"steps\":[{\"id\":\"metadata\",\"action\":\"info\",\"enabled\":true}]}}'\n```\n\n返回：\n\n```json\n{\n  \"status\": \"queued\",\n  \"code\": \"ok\",\n  \"reply\": \"任务已提交：<job_id>\",\n  \"job_id\": \"<job_id>\",\n  \"job_path\": \"output/jobs/<job_id>/job.json\",\n  \"poll_url\": \"/skill/jobs/<job_id>\"\n}\n```\n\n轮询任务：\n\n```bash\ncurl http://127.0.0.1:8080/skill/jobs/<job_id>\n```\n\n任务状态包括：`queued`、`running`、`success`、`partial`、`skipped`、`error`。任务结果保存在 `result`，产物路径汇总在 `output_paths`。\n\n查询任务列表：\n\n```bash\ncurl 'http://127.0.0.1:8080/skill/jobs?status=success&limit=50'\n```\n\n任务文件存储在 `output/jobs/<job_id>/job.json`。服务重启后，已完成任务仍可查询；未完成的 `queued` / `running` 任务会标记为 `error`，`code=job_interrupted`。\n\n### 字幕识别\n\nAction: `subtitle`\n\n推荐使用 `subtitle`，默认 `mode=fusion`：ASR 负责主要时间轴和文本，OCR 做画面字幕校正。识别依赖随 `requirements.txt` 安装；如果环境未重新安装依赖，会返回 JSON error，不会抛未捕获异常。\n\n```bash\npython run.py -a subtitle -i '{\"source\":\"sample.mp4\",\"mode\":\"fusion\"}'\npython run.py -a subtitle -i '{\"source\":\"sample.mp4\",\"mode\":\"asr\",\"language\":\"zh\"}'\npython run.py -a subtitle -i '{\"source\":\"sample.mp4\",\"mode\":\"ocr\",\"sample_interval_sec\":1}'\n```\n\n参数：\n\n| 参数 | 类型 | 默认值 | 说明 |\n|------|------|--------|------|\n| `video_url` / `url` / `source` | string | 必填 | 远程 URL 或本地文件 |\n| `mode` | string | fusion | `asr` / `ocr` / `fusion` |\n| `language` | string | auto | ASR 语言，例：`zh` / `en` |\n| `model_size` | string | base | faster-whisper 模型规格 |\n| `sample_interval_sec` | number | 1.0 | OCR 抽帧间隔 |\n| `crop_bottom_ratio` | number | 0.35 | OCR 默认扫描画面下方比例 |\n| `output_path` | string | `output/subtitles/...` | 字幕 JSON 输出路径 |\n| `overwrite` | boolean | true | 是否覆盖已有文件 |\n\n字幕条目格式：\n\n```json\n{\n  \"captionTxt\": \"识别到的字幕文本\",\n  \"startTimeUs\": 1000000,\n  \"endTimeUs\": 2000000,\n  \"source\": \"asr\",\n  \"confidence\": 0.92\n}\n```\n\n底层调试入口：\n\n```bash\npython run.py -a asr -i '{\"source\":\"sample.mp4\",\"language\":\"zh\"}'\npython run.py -a ocr -i '{\"source\":\"sample.mp4\"}'\n```\n\n### 字幕二次分句\n\nAction: `caption_segment`\n\n`caption_segment` 是 emlet 字幕分句器：它不重新识别字幕，只处理已有 captions。默认 `max_chars=12`，输出仍是 SRT-like captions JSON。\n\n```bash\npython run.py -a caption_segment -i '{\n  \"caption_path\":\"output/subtitles/sample.captions.json\",\n  \"max_chars\":12,\n  \"protected_terms\":[\"苹果\",\"华为\",\"吉利\"]\n}'\n```\n\n也可以直接传 captions：\n\n```bash\npython run.py -a caption_segment -i '{\n  \"captions\":[\n    {\"captionTxt\":\"今天我们聊苹果华为和吉利的新产品\",\"startTimeUs\":0,\"endTimeUs\":3000000}\n  ],\n  \"max_chars\":12,\n  \"protected_terms\":\"苹果,华为,吉利\"\n}'\n```\n\n长期词库可以放在 JSON 文件中：\n\n```json\n{\n  \"brands\": [\"苹果\", \"华为\", \"吉利\"],\n  \"products\": [\"小米汽车\", \"Model Y\", \"ChatGPT\"]\n}\n```\n\n参数：\n\n| 参数 | 类型 | 默认值 | 说明 |\n|------|------|--------|------|\n| `caption_path` / `input_path` | string | - | 已有 captions JSON 文件 |\n| `captions` | array | - | 直接传入字幕条目 |\n| `max_chars` | integer | 12 | 单条字幕最大字符数 |\n| `protected_terms` | array/string | [] | 不拆开的品牌词、产品名、人名、术语 |\n| `protected_terms_path` | string | - | 当前工作目录内的保护词 JSON 文件 |\n| `auto_protect_ascii` | boolean | true | 自动保护英文、数字、型号、URL、路径 |\n| `output_path` | string | `output/subtitles/...` | 分句后字幕 JSON 输出路径 |\n| `overwrite` | boolean | true | 是否覆盖已有文件 |\n\n### 压缩视频\n\nAction: `compress`\n\n```bash\npython run.py -a compress -i '{\"source\":\"sample.mp4\",\"target_ratio\":0.1}'\n```\n\n参数：\n\n| 参数 | 类型 | 默认值 | 说明 |\n|------|------|--------|------|\n| `video_url` / `url` / `source` | string | 必填 | 远程 URL 或本地文件 |\n| `target_ratio` | number | 0.1 | 目标体积比例 |\n| `adaptive` | boolean | true | 是否自动尝试不同 CRF |\n| `crf` | integer | 24 | 非 adaptive 模式下使用 |\n| `preset` | string | veryfast | ffmpeg 编码预设 |\n| `output_path` | string | `output/videos/...` | 输出路径 |\n| `overwrite` | boolean | true | 是否覆盖已有文件 |\n\n### 提取封面\n\nAction: `thumbnail`\n\n当前支持 MP4/MOV 容器，主要适用于 H.264/H.265 视频轨道。\n\n```bash\npython run.py -a thumbnail -i '{\"source\":\"sample.mp4\",\"time_seconds\":5}'\n```\n\n参数：\n\n| 参数 | 类型 | 默认值 | 说明 |\n|------|------|--------|------|\n| `video_url` / `url` / `source` | string | 必填 | 远程 URL 或本地文件 |\n| `time_seconds` | number | 0 | 按时间点提取 |\n| `frame_number` | integer | - | 按帧号提取，优先于 `time_seconds` |\n| `save_path` | string | `output/thumbs/...` | 保存路径 |\n| `resize_width` | integer | - | 输出宽度 |\n| `quality` | integer | 85 | JPEG 质量 |\n| `overwrite` | boolean | true | 是否覆盖已有文件 |\n\n### 提取音频\n\nAction: `audio`\n\n```bash\npython run.py -a audio -i '{\"source\":\"sample.mp4\",\"format\":\"mp3\"}'\n```\n\n参数：\n\n| 参数 | 类型 | 默认值 | 说明 |\n|------|------|--------|------|\n| `video_url` / `url` / `source` | string | 必填 | 远程 URL 或本地文件 |\n| `format` | string | mp3 | mp3 / wav / aac / m4a |\n| `bitrate` | string | 128k | 音频比特率 |\n| `sample_rate` | integer | 44100 | 采样率 |\n| `channels` | integer | 2 | 声道数 |\n| `start_time` | number | - | 开始时间，秒 |\n| `duration` | number | - | 截取时长，秒 |\n| `output_path` | string | `output/audio/...` | 输出路径 |\n| `overwrite` | boolean | true | 是否覆盖已有文件 |\n\n### 批量音频\n\nAction: `audio_batch`\n\n```bash\npython run.py -a audio_batch -i '{\n  \"videos\":[\n    {\"source\":\"sample1.mp4\",\"name\":\"video1\"},\n    {\"url\":\"https://example.com/2.mp4\",\"name\":\"video2\"}\n  ],\n  \"output_dir\":\"output/audio\",\n  \"format\":\"mp3\"\n}'\n```\n\n### 批量处理\n\nAction: `batch`\n\n```bash\npython run.py -a batch -i '{\n  \"action\":\"thumbnail\",\n  \"videos\":[\n    {\"source\":\"sample1.mp4\",\"time_seconds\":5},\n    {\"url\":\"https://example.com/2.mp4\",\"time_seconds\":10}\n  ]\n}'\n```\n\n`action` 支持：`compress`、`thumbnail`、`audio`。\n\n### JSON 流水线\n\nAction: `pipeline`\n\n`steps` 是唯一流程控制入口，没有写进 `steps` 的动作不会执行。支持的 step action：`info`、`thumbnail`、`audio`、`compress`、`audio_info`、`asr`、`ocr`、`subtitle`、`caption_segment`。\n\n```bash\npython run.py -a pipeline -i '{\n  \"source\":\"sample.mp4\",\n  \"name\":\"sample\",\n  \"output_dir\":\"output/pipeline/sample\",\n  \"overwrite\":true,\n  \"steps\":[\n    {\"id\":\"metadata\",\"action\":\"info\",\"enabled\":true},\n    {\n      \"id\":\"cover\",\n      \"action\":\"thumbnail\",\n      \"enabled\":true,\n      \"params\":{\"time_seconds\":3,\"resize_width\":720}\n    },\n    {\n      \"id\":\"audio_mp3\",\n      \"action\":\"audio\",\n      \"enabled\":false,\n      \"params\":{\"format\":\"mp3\",\"bitrate\":\"128k\"}\n    }\n  ]\n}'\n```\n\n规则：\n\n| 参数 | 类型 | 默认值 | 说明 |\n|------|------|--------|------|\n| `video_url` / `url` / `source` | string | 必填 | 远程 URL 或本地文件 |\n| `name` | string | 输入文件名 | 流水线名称 |\n| `output_dir` | string | `output/pipeline/<name>` | manifest 和默认产物目录 |\n| `overwrite` | boolean | true | 是否覆盖已有产物 |\n| `steps` | array | 必填 | 按 JSON 顺序执行的步骤 |\n\n每个 step 需要 `id`、`action`、`enabled`。`enabled=false` 会记录为 `skipped`。每次执行都会生成 `manifest.json`。\n\n### 获取信息\n\n```bash\npython run.py -a info -i '{\"source\":\"sample.mp4\"}'\npython run.py -a audio_info -i '{\"source\":\"sample.mp4\"}'\n```\n\n## 注意\n\n- 远程输入仅支持 `http` / `https`。\n- 本地输入路径默认限制在当前工作目录内；通过 `media_roots` / `YM_MEDIA_ROOTS` 可授权额外媒体根目录。\n- 输出路径限制在当前工作目录内。\n- HTTP 服务默认只绑定本机地址。\n\nFile v4.2.1:_meta.json\n\n{\n  \"ownerId\": \"kn787sam7qk7fffsjc875yxp0984rpzh\",\n  \"slug\": \"ym-mediatoolkit\",\n  \"version\": \"4.2.1\",\n  \"publishedAt\": 1782096374001\n}\n\nFile v4.2.1:MAINTENANCE.md\n\n# YM MediaToolkit 维护手册\n\n本文档面向后续维护者，记录版本发布、测试验证、配置边界和排障流程。\n\n## 发布流程\n\n1. 同步版本号：\n   - `SKILL.md` front matter 的 `version`\n   - `skill.json` 的 `version`\n2. 更新用户文档：\n   - 新增 action 时同步更新 `SKILL.md` 和 `skill.json`\n   - 新增参数时同步补充默认值、用途和安全限制\n   - 变更行为时在 `SKILL.md` 的版本更新段落记录\n3. 运行验证：\n\n```bash\npython3 -B -m py_compile run.py utils.py intent_parser.py audio_extractor.py frame_extractor.py video_compressor.py asr_engine.py ocr_engine.py subtitle_extractor.py caption_segmenter.py job_manager.py tests/test_release_behaviors.py scripts/smoke_test.py\npython3 -m json.tool skill.json\npython3 -B -m unittest discover -s tests\npython3 scripts/smoke_test.py\n```\n\n4. 清理产物：\n\n```bash\nfind . -type d -name __pycache__ -prune -exec rm -r {} +\nfind . -name .DS_Store -delete\n```\n\n## 当前接口分层\n\n- `chat`：Claw 推荐入口，接收自然语言，解析后调用现有 action。\n- `pipeline`：确定性 JSON 流水线入口，适合多步骤、可重复的自动化流程。\n- `audio` / `thumbnail` / `compress` / `info` / `audio_info`：底层单步能力。\n- `subtitle` / `asr` / `ocr`：字幕识别能力，输出 SRT-like captions JSON。\n- `caption_segment`：emlet 字幕二次分句器，处理已有 captions，不重新识别媒体。\n- `batch` / `audio_batch`：批量处理入口。\n- `/skill/jobs`：HTTP 异步长任务入口，适合压缩、字幕识别、pipeline 等耗时 action。\n- `/skill/chat` + `async:\"auto\"`：Claw 推荐长任务入口，先解析自然语言，再自动提交 job。\n\n维护原则：新体验优先接到 `chat` 或 `pipeline`，媒体处理逻辑继续复用底层 action，避免重复实现 ffmpeg 调用。\n\n## Action 返回协议\n\n所有 action 必须返回 JSON 对象，并保留以下协议字段：\n\n- `status`：`success` / `partial` / `skipped` / `error`\n- `code`：稳定机器码，成功为 `ok`\n- `reply`：适合聊天展示的中文回复\n- `hint`：下一步建议或排障提示\n\n新增 handler 时只需要返回原始业务结果，`run.py` 的 action protocol wrapper 会补齐缺失字段。若底层模块已经返回 `code`，包装层会保留该值。\n\n常用错误码：\n\n| code | 场景 |\n|------|------|\n| `missing_source` | 缺少 `video_url` / `url` / `source` |\n| `source_not_allowed` | 本地路径不在当前工作目录或 `media_roots` 内 |\n| `output_exists` | 输出文件已存在且 `overwrite=false` |\n| `parse_failed` | `chat` 无法识别自然语言意图 |\n| `missing_steps` | `pipeline` 未传 `steps` |\n| `missing_captions` | `caption_segment` 未传 `captions` 或 `caption_path` |\n| `invalid_action` | 异步任务或 CLI 传入不支持的 action |\n| `invalid_job_id` | job id 格式不合法 |\n| `job_not_found` | job 文件不存在 |\n| `job_interrupted` | 服务重启或进程中断导致未完成 job 失效 |\n| `unsupported_action` | action 或 pipeline step 不支持 |\n| `invalid_step` | pipeline step 结构不合法 |\n| `ffmpeg_failed` | ffmpeg / ffprobe 缺失或执行失败 |\n| `missing_asr_dependency` | ASR 依赖缺失 |\n| `missing_ocr_dependency` | OCR 依赖缺失 |\n\n## HTTP 异步任务维护\n\n异步任务逻辑在 `job_manager.py`。HTTP 服务启动时会创建单 worker 队列，串行执行提交到 `/skill/jobs` 的 action。\n\n维护规则：\n\n- job 文件存储在 `output/jobs/<job_id>/job.json`。\n- job id 使用 32 位十六进制 UUID。\n- job 文件必须保留 `job_id`、`action`、`params`、`status`、`code`、`reply`、`hint`、`created_at`、`started_at`、`finished_at`、`result`、`output_paths`、`error`。\n- `chat` 提交的 job 会额外写入 `created_by=chat`、`intent`、`source`、`metadata.message`。\n- `/skill/jobs` 直接提交的 job 会写入 `created_by=jobs`。\n- 当前版本不支持取消任务，不记录进度百分比。\n- 服务重启后不恢复未完成任务；旧的 `queued` / `running` 会标记为 `error`，`code=job_interrupted`。\n- HTTP 服务启动时会调用 `cleanup_jobs(retention_days=7, max_jobs=200)`，只清理过旧或超量的终态 job。\n- 新增 action 时，只要加入 `ACTIONS`，异步任务会自动支持。\n\n## HTTP chat 异步排障\n\n`/skill/chat` 支持 `async` 参数：\n\n- `false` 或不传：保持同步执行，兼容 CLI 和旧 HTTP 调用。\n- `true`：解析成功后总是提交 job，不立即执行底层 handler。\n- `\"auto\"`：只把 `audio`、`compress`、`asr`、`ocr`、`subtitle`、`caption_segment`、`batch`、`pipeline` 作为长任务提交；`info`、`audio_info`、`thumbnail` 继续同步。\n\n排查要点：\n\n- 如果返回 `parse_failed`，说明自然语言没有解析出 action 或 source，不会创建 job。\n- 如果返回 `queued`，让调用方使用 `poll_url` 查询最终 `reply`、`result`、`output_paths`。\n- 如果任务完成后 `output_paths` 为空，检查底层 action 是否返回了 `output_path`、`saved_path`、`outputPath` 或 `manifest_path`。\n\n## media_roots 配置\n\n本地输入默认只允许当前工作目录内的文件。需要处理绝对路径时，必须配置媒体根目录白名单。\n\n请求级配置：\n\n```json\n{\n  \"message\": \"将 \\\"D:/AA.MP4\\\" 提取音频\",\n  \"media_roots\": [\"D:/\"]\n}\n```\n\n环境变量配置：\n\n```bash\nexport YM_MEDIA_ROOTS=\"/Users/me/Videos;/Volumes/Media\"\n```\n\n规则：\n\n- `media_roots` 优先级高于 `YM_MEDIA_ROOTS`\n- 未配置时只允许当前工作目录\n- 支持用逗号或分号分隔多个根目录\n- URL 仍只允许 `http` / `https`\n- 输出路径仍限制在当前工作目录内\n\n## 自然语言解析维护\n\n解析逻辑在 `intent_parser.py`。\n\n当前支持：\n\n- 提取音频：`将 \"sample.mp4\" 提取音频`\n- 提取封面：`给 \"sample.mp4\" 提取第 3 秒封面`\n- 压缩：`压缩 \"sample.mp4\"`\n- 查看信息：`查看 \"sample.mp4\" 信息`\n- JSON pipeline：消息本身是包含 `steps` 的 JSON\n- 字幕识别：`识别 \"sample.mp4\" 的字幕`\n\n新增自然语言规则时需要同时补：\n\n- `tests/test_release_behaviors.py` 的解析测试\n- `scripts/smoke_test.py` 的真实链路测试，若会产生文件\n- `SKILL.md` 的示例\n- `skill.json` 的 action schema 或 examples，若公开接口有变化\n\n## 验证策略\n\n单元测试覆盖轻量行为：\n\n- 输入字段兼容：`video_url` / `url` / `source`\n- 本地路径和 `media_roots` 权限\n- 默认输出目录和 `overwrite=false`\n- pipeline 顺序、跳过、失败继续、manifest\n- chat 解析、执行、失败不调用 handler\n\nSmoke test 覆盖真实 ffmpeg 链路：\n\n- 生成 1 秒本地测试视频\n- 跑 `info`、`thumbnail`、`audio`、`compress`、`batch`\n- 跑 `pipeline`\n- 跑 `chat` 的音频、封面、压缩、信息命令\n- 跑 `subtitle`；若环境未重新安装依赖，需返回 dependency error；安装完整依赖后验证 captions JSON\n\n## 字幕识别维护\n\n字幕输出统一使用 camelCase 和微秒整数：\n\n```json\n{\n  \"captionTxt\": \"字幕文本\",\n  \"startTimeUs\": 1000000,\n  \"endTimeUs\": 2000000,\n  \"source\": \"asr\",\n  \"confidence\": 0.92\n}\n```\n\n字幕识别依赖已进入默认 `requirements.txt`：\n\n```bash\npip install -r requirements.txt\n```\n\n维护规则：\n\n- `subtitle` 是推荐入口，`asr` / `ocr` 用于单独调试。\n- `mode=fusion` 默认 ASR 为主，OCR 只做高置信文本校正。\n- 环境未安装完整依赖时必须返回 JSON error，不能抛未捕获异常。\n- 新增字幕准确率策略时，优先补 `subtitle_extractor.py` 的纯函数测试。\n\n## emlet 字幕分句维护\n\n二次分句逻辑在 `caption_segmenter.py`，对已有 captions 做后处理，不调用 ASR/OCR。\n\n默认策略：\n\n- `max_chars=12`\n- 强标点优先切分：`。！？!?`\n- 弱标点其次：`，、；：,;:`\n- 再尝试连接词边界，最后按长度切分\n- `protected_terms` 和 `protected_terms_path` 内的词不允许被拆开\n- `auto_protect_ascii=true` 时自动保护英文、数字、型号、URL、路径\n\n维护规则：\n\n- 新增分句策略时先补纯函数测试，确保时间轴连续、不重叠、不倒退。\n- `caption_segment` 输出继续使用 `captionTxt`、`startTimeUs`、`endTimeUs`。\n- pipeline 中 `caption_segment` 如果没有显式传 `caption_path`，会自动使用前面步骤生成的 `.captions.json`。\n\n## 常见问题\n\n`本地输入路径超出允许的 media_roots`\n\n确认文件路径位于当前工作目录，或在请求中传入 `media_roots`。\n\n`输出路径超出工作目录`\n\n输出文件必须写入当前工作目录下，例如 `output/audio/demo.mp3`。\n\n`ffmpeg 错误，返回码: ...`\n\n先确认输入文件可播放，再用 smoke test 验证当前环境的 ffmpeg 是否可用。\n\n`DNS 解析失败` 或 `禁止访问私有/内网 IP`\n\n远程 URL 会做安全校验，不允许内网、回环、链路本地地址或无法验证的目标。\n\n`chat` 没识别出命令\n\n优先使用明确句式：`将 \"sample.mp4\" 提取音频`、`给 \"sample.mp4\" 提取第 3 秒封面`、`压缩 \"sample.mp4\"`。\n\n`缺少 ASR/OCR 依赖`\n\n重新运行 `pip install -r requirements.txt`，确认 `faster-whisper`、`paddlepaddle`、`paddleocr` 已安装。\n\nFile v4.2.1:skill-card.md\n\n## Description: <br>\nYM-MediaToolkit is a natural-language media assistant for video compression, MP4/MOV thumbnail extraction, audio conversion, subtitle recognition, and JSON media pipelines. <br>\n\nThis skill is ready for commercial/non-commercial use. <br>\n\n## Publisher: <br>\n[370299455cx-web](https://clawhub.ai/user/370299455cx-web) <br>\n\n### License/Terms of Use: <br>\nMIT-0 <br>\n\n\n## Use Case: <br>\nDevelopers, operators, and content teams use this skill to process trusted media through chat, CLI, HTTP, or JSON pipeline flows for compression, cover extraction, audio extraction, subtitle recognition, caption segmentation, and long-running job tracking. <br>\n\n### Deployment Geography for Use: <br>\nGlobal <br>\n\n## Known Risks and Mitigations: <br>\nRisk: The skill fetches remote media URLs and processes local media with ffmpeg, ASR, and OCR. <br>\nMitigation: Use trusted URLs and authorized media inputs; keep media_roots narrow so local file access stays limited. <br>\nRisk: Generated output files can be overwritten by default. <br>\nMitigation: Set overwrite=false when preserving existing output files matters. <br>\nRisk: The HTTP service can be bound beyond localhost if configured that way. <br>\nMitigation: Keep the default localhost binding or place any wider binding behind access control. <br>\n\n\n## Reference(s): <br>\n- [ClawHub Skill Page](https://clawhub.ai/370299455cx-web/ym-mediatoolkit) <br>\n- [Publisher Profile](https://clawhub.ai/user/370299455cx-web) <br>\n- [Skill Documentation](artifact/SKILL.md) <br>\n- [Maintenance Guide](artifact/MAINTENANCE.md) <br>\n\n\n## Skill Output: <br>\n**Output Type(s):** [Text, JSON, Files] <br>\n**Output Format:** [JSON responses with status, code, reply, hint, result data, output paths, generated media files, caption JSON, manifests, and job records.] <br>\n**Output Parameters:** [1D] <br>\n**Other Properties Related to Output:** [Outputs may include workspace files under output/videos, output/audio, output/thumbs, output/subtitles, output/pipeline, and output/jobs.] <br>\n\n## Skill Version(s): <br>\n4.2.1 (source: server release metadata) <br>\n\n## Ethical Considerations: <br>\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment. <br>\n\nFile v4.2.1:skill.json\n\n{\n  \"name\": \"ym-mediatoolkit\",\n  \"version\": \"4.3.0\",\n  \"description\": \"自然语言媒体助手：视频压缩、MP4/MOV 封面提取、音频转换、字幕识别、JSON 流水线\",\n  \"author\": \"your_name\",\n  \"entrypoint\": \"python run.py --input {input_json}\",\n  \"http_port\": 8080,\n  \"http_bind\": \"127.0.0.1 (默认，仅本地；可通过 --host 0.0.0.0 改为公网，需反向代理认证)\",\n  \"external_binaries\": {\n    \"ffmpeg\": \"必需 - 视频压缩、音频提取、流式处理\",\n    \"ffprobe\": \"必需 - 获取视频/音频流信息\"\n  },\n  \"python_dependencies\": {\n    \"faster-whisper\": \"必需 - ASR 字幕识别\",\n    \"paddlepaddle\": \"必需 - PaddleOCR 推理运行时\",\n    \"paddleocr\": \"必需 - OCR 字幕识别\"\n  },\n  \"response_protocol\": {\n    \"description\": \"所有 action 都会返回稳定协议字段，并保留原有业务字段。\",\n    \"fields\": {\n      \"status\": \"success / partial / skipped / error\",\n      \"code\": \"稳定机器码，例如 ok、missing_source、output_exists、parse_failed\",\n      \"reply\": \"适合聊天展示的简短中文回复\",\n      \"hint\": \"面向调用方或用户的下一步建议\"\n    },\n    \"common_codes\": [\n      \"ok\",\n      \"missing_source\",\n      \"source_not_allowed\",\n      \"output_exists\",\n      \"parse_failed\",\n      \"missing_steps\",\n      \"missing_captions\",\n      \"invalid_action\",\n      \"invalid_params\",\n      \"invalid_json\",\n      \"invalid_job_id\",\n      \"job_not_found\",\n      \"job_interrupted\",\n      \"job_failed\",\n      \"unsupported_action\",\n      \"invalid_step\",\n      \"ffmpeg_failed\",\n      \"missing_asr_dependency\",\n      \"missing_ocr_dependency\"\n    ]\n  },\n  \"http_apis\": [\n    {\n      \"path\": \"/skill/jobs\",\n      \"method\": \"POST\",\n      \"description\": \"提交异步长任务，现有同步 /skill/<action> 接口保持不变\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"required\": [\"action\"],\n        \"properties\": {\n          \"action\": {\n            \"type\": \"string\",\n            \"description\": \"当前 actions 中支持的 action 名称\"\n          },\n          \"params\": {\n            \"type\": \"object\",\n            \"default\": {},\n            \"description\": \"传给 action 的参数\"\n          }\n        }\n      },\n      \"response_schema\": {\n        \"type\": \"object\",\n        \"properties\": {\n          \"status\": {\"type\": \"string\", \"enum\": [\"queued\", \"error\"]},\n          \"code\": {\"type\": \"string\"},\n          \"reply\": {\"type\": \"string\"},\n          \"job_id\": {\"type\": \"string\"},\n          \"job_path\": {\"type\": \"string\"},\n          \"poll_url\": {\"type\": \"string\"},\n          \"created_by\": {\"type\": \"string\"},\n          \"intent\": {\"type\": \"string\"},\n          \"params\": {\"type\": \"object\"}\n        }\n      }\n    },\n    {\n      \"path\": \"/skill/jobs/<job_id>\",\n      \"method\": \"GET\",\n      \"description\": \"查询单个异步任务状态、结果和输出路径\"\n    },\n    {\n      \"path\": \"/skill/jobs\",\n      \"method\": \"GET\",\n      \"description\": \"查询异步任务列表，支持 status 和 limit query 参数\"\n    }\n  ],\n  \"actions\": [\n    {\n      \"name\": \"chat\",\n      \"description\": \"自然语言媒体助手入口，将用户聊天文本解析为媒体处理 action；HTTP 可用 async 自动提交长任务 job\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"required\": [\"message\"],\n        \"properties\": {\n          \"message\": {\n            \"type\": \"string\",\n            \"description\": \"自然语言命令，例如：将 \\\"D:/AA.MP4\\\" 提取音频\"\n          },\n          \"media_roots\": {\n            \"type\": [\"array\", \"string\"],\n            \"description\": \"允许访问的本地媒体根目录；未传时读取 YM_MEDIA_ROOTS，否则仅允许当前工作目录\"\n          },\n          \"overwrite\": {\"type\": \"boolean\", \"default\": true},\n          \"output_dir\": {\"type\": \"string\", \"description\": \"自然语言命令的默认输出目录\"},\n          \"async\": {\n            \"oneOf\": [\n              {\"type\": \"boolean\"},\n              {\"type\": \"string\", \"enum\": [\"auto\"]}\n            ],\n            \"default\": false,\n            \"description\": \"HTTP chat 异步控制；true 全部转 job，auto 仅长任务转 job\"\n          },\n          \"wait_timeout_sec\": {\n            \"type\": \"number\",\n            \"default\": 0,\n            \"description\": \"异步提交后最多短等待秒数；任务很快完成时可直接返回最终结果\"\n          }\n        }\n      },\n      \"examples\": [\n        {\"message\": \"将 \\\"sample.mp4\\\" 提取音频\"},\n        {\"message\": \"给 \\\"sample.mp4\\\" 提取第 3 秒封面\"},\n        {\"message\": \"压缩 \\\"sample.mp4\\\"\"},\n        {\"message\": \"查看 \\\"sample.mp4\\\" 信息\"},\n        {\"message\": \"识别 \\\"sample.mp4\\\" 的字幕\", \"async\": \"auto\"}\n      ]\n    },\n    {\n      \"name\": \"compress\",\n      \"description\": \"流式压缩视频，保持清晰度，支持自适应 CRF\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"anyOf\": [\n          {\"required\": [\"video_url\"]},\n          {\"required\": [\"url\"]},\n          {\"required\": [\"source\"]}\n        ],\n        \"properties\": {\n          \"video_url\": {\"type\": \"string\", \"description\": \"视频 URL 或本地文件路径\"},\n          \"url\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"source\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"media_roots\": {\"type\": [\"array\", \"string\"], \"description\": \"允许访问的本地媒体根目录\"},\n          \"target_ratio\": {\"type\": \"number\", \"default\": 0.1},\n          \"adaptive\": {\"type\": \"boolean\", \"default\": true},\n          \"crf\": {\"type\": \"integer\", \"default\": 24},\n          \"preset\": {\"type\": \"string\", \"default\": \"veryfast\"},\n          \"output_path\": {\"type\": \"string\", \"description\": \"默认 output/videos\"},\n          \"overwrite\": {\"type\": \"boolean\", \"default\": true}\n        }\n      }\n    },\n    {\n      \"name\": \"thumbnail\",\n      \"description\": \"从 MP4/MOV 视频任意时间点或帧号提取封面，流式只下载必要部分\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"anyOf\": [\n          {\"required\": [\"video_url\"]},\n          {\"required\": [\"url\"]},\n          {\"required\": [\"source\"]}\n        ],\n        \"properties\": {\n          \"video_url\": {\"type\": \"string\", \"description\": \"视频 URL 或本地文件路径\"},\n          \"url\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"source\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"media_roots\": {\"type\": [\"array\", \"string\"], \"description\": \"允许访问的本地媒体根目录\"},\n          \"time_seconds\": {\"type\": \"number\"},\n          \"frame_number\": {\"type\": \"integer\"},\n          \"save_path\": {\"type\": \"string\"},\n          \"resize_width\": {\"type\": \"integer\"},\n          \"quality\": {\"type\": \"integer\", \"default\": 85},\n          \"overwrite\": {\"type\": \"boolean\", \"default\": true}\n        }\n      }\n    },\n    {\n      \"name\": \"audio\",\n      \"description\": \"流式提取音频，支持 MP3/WAV/AAC/M4A 格式\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"anyOf\": [\n          {\"required\": [\"video_url\"]},\n          {\"required\": [\"url\"]},\n          {\"required\": [\"source\"]}\n        ],\n        \"properties\": {\n          \"video_url\": {\"type\": \"string\", \"description\": \"视频 URL 或本地文件路径\"},\n          \"url\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"source\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"media_roots\": {\"type\": [\"array\", \"string\"], \"description\": \"允许访问的本地媒体根目录\"},\n          \"format\": {\"type\": \"string\", \"enum\": [\"mp3\", \"wav\", \"aac\", \"m4a\"], \"default\": \"mp3\"},\n          \"bitrate\": {\"type\": \"string\", \"default\": \"128k\", \"description\": \"比特率: 128k, 192k, 320k\"},\n          \"sample_rate\": {\"type\": \"integer\", \"default\": 44100, \"description\": \"采样率: 44100, 48000\"},\n          \"channels\": {\"type\": \"integer\", \"default\": 2, \"description\": \"声道: 1=单声道, 2=立体声\"},\n          \"start_time\": {\"type\": \"number\", \"description\": \"开始时间（秒）\"},\n          \"duration\": {\"type\": \"number\", \"description\": \"持续时间（秒）\"},\n          \"output_path\": {\"type\": \"string\", \"description\": \"默认 output/audio\"},\n          \"overwrite\": {\"type\": \"boolean\", \"default\": true}\n        }\n      }\n    },\n    {\n      \"name\": \"audio_batch\",\n      \"description\": \"批量提取多个视频的音频\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"required\": [\"videos\"],\n        \"properties\": {\n          \"videos\": {\"type\": \"array\", \"description\": \"视频列表，条目支持 video_url/url/source\"},\n          \"media_roots\": {\"type\": [\"array\", \"string\"], \"description\": \"允许访问的本地媒体根目录\"},\n          \"output_dir\": {\"type\": \"string\", \"default\": \"output/audio\"},\n          \"format\": {\"type\": \"string\", \"default\": \"mp3\"},\n          \"bitrate\": {\"type\": \"string\", \"default\": \"128k\"},\n          \"sample_rate\": {\"type\": \"integer\", \"default\": 44100},\n          \"overwrite\": {\"type\": \"boolean\", \"default\": true}\n        }\n      }\n    },\n    {\n      \"name\": \"audio_info\",\n      \"description\": \"获取视频的音频流信息\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"anyOf\": [\n          {\"required\": [\"video_url\"]},\n          {\"required\": [\"url\"]},\n          {\"required\": [\"source\"]}\n        ],\n        \"properties\": {\n          \"video_url\": {\"type\": \"string\", \"description\": \"视频 URL 或本地文件路径\"},\n          \"url\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"source\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"media_roots\": {\"type\": [\"array\", \"string\"], \"description\": \"允许访问的本地媒体根目录\"}\n        }\n      }\n    },\n    {\n      \"name\": \"asr\",\n      \"description\": \"从音频轨识别字幕，输出 SRT-like captions JSON\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"anyOf\": [\n          {\"required\": [\"video_url\"]},\n          {\"required\": [\"url\"]},\n          {\"required\": [\"source\"]}\n        ],\n        \"properties\": {\n          \"video_url\": {\"type\": \"string\", \"description\": \"视频 URL 或本地文件路径\"},\n          \"url\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"source\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"language\": {\"type\": \"string\", \"default\": \"auto\"},\n          \"model_size\": {\"type\": \"string\", \"default\": \"base\"},\n          \"output_path\": {\"type\": \"string\", \"description\": \"默认 output/subtitles/<name>.captions.json\"},\n          \"overwrite\": {\"type\": \"boolean\", \"default\": true},\n          \"media_roots\": {\"type\": [\"array\", \"string\"], \"description\": \"允许访问的本地媒体根目录\"}\n        }\n      }\n    },\n    {\n      \"name\": \"ocr\",\n      \"description\": \"从视频画面识别硬字幕或屏幕文字，输出 SRT-like captions JSON\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"anyOf\": [\n          {\"required\": [\"video_url\"]},\n          {\"required\": [\"url\"]},\n          {\"required\": [\"source\"]}\n        ],\n        \"properties\": {\n          \"video_url\": {\"type\": \"string\", \"description\": \"视频 URL 或本地文件路径\"},\n          \"url\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"source\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"sample_interval_sec\": {\"type\": \"number\", \"default\": 1.0},\n          \"crop_bottom_ratio\": {\"type\": \"number\", \"default\": 0.35},\n          \"output_path\": {\"type\": \"string\", \"description\": \"默认 output/subtitles/<name>.captions.json\"},\n          \"overwrite\": {\"type\": \"boolean\", \"default\": true},\n          \"media_roots\": {\"type\": [\"array\", \"string\"], \"description\": \"允许访问的本地媒体根目录\"}\n        }\n      }\n    },\n    {\n      \"name\": \"subtitle\",\n      \"description\": \"推荐字幕识别入口，支持 asr / ocr / fusion 模式，统一输出 captions JSON\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"anyOf\": [\n          {\"required\": [\"video_url\"]},\n          {\"required\": [\"url\"]},\n          {\"required\": [\"source\"]}\n        ],\n        \"properties\": {\n          \"video_url\": {\"type\": \"string\", \"description\": \"视频 URL 或本地文件路径\"},\n          \"url\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"source\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"mode\": {\"type\": \"string\", \"enum\": [\"asr\", \"ocr\", \"fusion\"], \"default\": \"fusion\"},\n          \"language\": {\"type\": \"string\", \"default\": \"auto\"},\n          \"model_size\": {\"type\": \"string\", \"default\": \"base\"},\n          \"sample_interval_sec\": {\"type\": \"number\", \"default\": 1.0},\n          \"crop_bottom_ratio\": {\"type\": \"number\", \"default\": 0.35},\n          \"output_path\": {\"type\": \"string\", \"description\": \"默认 output/subtitles/<name>.captions.json\"},\n          \"overwrite\": {\"type\": \"boolean\", \"default\": true},\n          \"media_roots\": {\"type\": [\"array\", \"string\"], \"description\": \"允许访问的本地媒体根目录\"}\n        }\n      },\n      \"examples\": [\n        {\"source\": \"sample.mp4\", \"mode\": \"fusion\"},\n        {\"source\": \"sample.mp4\", \"mode\": \"asr\", \"language\": \"zh\"}\n      ]\n    },\n    {\n      \"name\": \"caption_segment\",\n      \"description\": \"emlet 字幕二次分句器，对已有 captions JSON 按长度、标点和保护词重新切分\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"anyOf\": [\n          {\"required\": [\"caption_path\"]},\n          {\"required\": [\"input_path\"]},\n          {\"required\": [\"captions\"]}\n        ],\n        \"properties\": {\n          \"caption_path\": {\"type\": \"string\", \"description\": \"已有 captions JSON 文件路径\"},\n          \"input_path\": {\"type\": \"string\", \"description\": \"同 caption_path\"},\n          \"captions\": {\"type\": \"array\", \"description\": \"直接传入的字幕条目\"},\n          \"max_chars\": {\"type\": \"integer\", \"default\": 12, \"description\": \"单条字幕最大字符数\"},\n          \"protected_terms\": {\n            \"type\": [\"array\", \"string\"],\n            \"description\": \"不拆开的品牌词、产品名、人名、术语，例如 苹果,华为,吉利\"\n          },\n          \"protected_terms_path\": {\n            \"type\": \"string\",\n            \"description\": \"当前工作目录内的保护词 JSON 文件，可按 brands/products 等字段分组\"\n          },\n          \"auto_protect_ascii\": {\n            \"type\": \"boolean\",\n            \"default\": true,\n            \"description\": \"自动保护英文、数字、型号、URL、路径\"\n          },\n          \"output_path\": {\"type\": \"string\", \"description\": \"默认 output/subtitles/<name>.segmented.captions.json\"},\n          \"overwrite\": {\"type\": \"boolean\", \"default\": true}\n        }\n      },\n      \"examples\": [\n        {\n          \"caption_path\": \"output/subtitles/sample.captions.json\",\n          \"max_chars\": 12,\n          \"protected_terms\": [\"苹果\", \"华为\", \"吉利\"]\n        }\n      ]\n    },\n    {\n      \"name\": \"batch\",\n      \"description\": \"批量处理多个视频，支持 compress / thumbnail / audio\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"required\": [\"videos\"],\n        \"properties\": {\n          \"action\": {\"type\": \"string\", \"enum\": [\"compress\", \"thumbnail\", \"audio\"], \"default\": \"thumbnail\"},\n          \"videos\": {\"type\": \"array\", \"description\": \"按 action 传入对应参数对象列表，条目支持 video_url/url/source\"},\n          \"media_roots\": {\"type\": [\"array\", \"string\"], \"description\": \"允许访问的本地媒体根目录\"},\n          \"output_dir\": {\"type\": \"string\", \"description\": \"批量输出目录\"},\n          \"overwrite\": {\"type\": \"boolean\", \"default\": true}\n        }\n      }\n    },\n    {\n      \"name\": \"pipeline\",\n      \"description\": \"JSON 驱动媒体流水线，按 steps 顺序编排 info / thumbnail / audio / compress / audio_info / subtitle / caption_segment\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"anyOf\": [\n          {\"required\": [\"video_url\"]},\n          {\"required\": [\"url\"]},\n          {\"required\": [\"source\"]}\n        ],\n        \"required\": [\"steps\"],\n        \"properties\": {\n          \"video_url\": {\"type\": \"string\", \"description\": \"视频 URL 或本地文件路径\"},\n          \"url\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"source\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"media_roots\": {\"type\": [\"array\", \"string\"], \"description\": \"允许访问的本地媒体根目录\"},\n          \"name\": {\"type\": \"string\", \"description\": \"流水线名称\"},\n          \"output_dir\": {\"type\": \"string\", \"description\": \"默认 output/pipeline/<name>\"},\n          \"overwrite\": {\"type\": \"boolean\", \"default\": true},\n          \"steps\": {\n            \"type\": \"array\",\n            \"description\": \"步骤列表，每步包含 id/action/enabled/params\",\n            \"items\": {\n              \"type\": \"object\",\n              \"required\": [\"id\", \"action\", \"enabled\"],\n              \"properties\": {\n                \"id\": {\"type\": \"string\"},\n                \"action\": {\"type\": \"string\", \"enum\": [\"info\", \"thumbnail\", \"audio\", \"compress\", \"audio_info\", \"asr\", \"ocr\", \"subtitle\", \"caption_segment\"]},\n                \"enabled\": {\"type\": \"boolean\"},\n                \"params\": {\"type\": \"object\"}\n              }\n            }\n          }\n        }\n      }\n    },\n    {\n      \"name\": \"info\",\n      \"description\": \"获取完整视频信息（分辨率、时长、编码等）\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"anyOf\": [\n          {\"required\": [\"video_url\"]},\n          {\"required\": [\"url\"]},\n          {\"required\": [\"source\"]}\n        ],\n        \"properties\": {\n          \"video_url\": {\"type\": \"string\", \"description\": \"视频 URL 或本地文件路径\"},\n          \"url\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"source\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"media_roots\": {\"type\": [\"array\", \"string\"], \"description\": \"允许访问的本地媒体根目录\"}\n        }\n      }\n    }\n  ]\n}\n\nFile v4.2.1:requirements.txt\n\nrequests>=2.28.0\nopencv-python>=4.8.0\nnumpy>=1.24.0\naiohttp>=3.8.0\nflask>=2.3.0\nflask-cors>=4.0.0\nfaster-whisper>=1.0.0\npaddlepaddle>=2.6.0\npaddleocr>=2.7.0\n\nArchive v4.2.0: 19 files, 57729 bytes\n\nFiles: asr_engine.py (3005b), audio_extractor.py (9638b), caption_segmenter.py (8548b), frame_extractor.py (22227b), intent_parser.py (6678b), job_manager.py (8239b), MAINTENANCE.md (8052b), ocr_engine.py (3045b), requirements.txt (157b), run.py (38662b), scripts/smoke_test.py (8721b), skill-card.md (2806b), skill.json (17101b), SKILL.md (14576b), subtitle_extractor.py (5490b), tests/test_release_behaviors.py (29915b), utils.py (8193b), video_compressor.py (6897b), _meta.json (134b)\n\nFile v4.2.0:SKILL.md\n\n---\nname: ym-mediatoolkit\nversion: 4.2.0\ndescription: 自然语言媒体助手 - 视频压缩、MP4/MOV 封面提取、音频转换、字幕识别\nauthor: your_name\ntags:\n  - video\n  - compression\n  - thumbnail\n  - audio\n  - streaming\n  - ffmpeg\ncategories:\n  - media\n  - utility\nclawhub:\n  entrypoint: python run.py\n  runtime: python3\n  http_port: 8080\n---\n\n# YM MediaToolkit\n\n自然语言媒体助手，支持从远程视频 URL、当前工作目录内的本地视频文件、或配置过 `media_roots` 的本地媒体目录直接处理：\n\n- 自然语言调用\n- 视频压缩\n- MP4/MOV 封面提取\n- 音频提取与转换：MP3 / WAV / AAC / M4A\n- OCR / ASR 字幕识别\n- emlet 字幕二次分句\n- 批量处理\n- JSON 驱动媒体流水线\n\n## 依赖\n\n```bash\npip install -r requirements.txt\n```\n\n系统需要安装：\n\n```bash\nffmpeg\nffprobe\n```\n\n字幕识别依赖已内置在 `requirements.txt`，包括 `faster-whisper`、`paddlepaddle`、`paddleocr`。\n\n## 维护\n\n维护、发布、测试和排障流程见 [MAINTENANCE.md](./MAINTENANCE.md)。\n\n## 命令行\n\n```bash\npython run.py -a <action> -i '<json>'\n```\n\n也可以从 JSON 文件读取参数：\n\n```bash\npython run.py -i params.json\n```\n\n## HTTP 服务\n\n```bash\npython run.py --serve\n```\n\n默认监听 `127.0.0.1:8080`。\n\n健康检查：\n\n```http\nGET /health\n```\n\n异步长任务：\n\n```http\nPOST /skill/jobs\nGET /skill/jobs/<job_id>\nGET /skill/jobs\n```\n\n## 功能\n\n输入源字段可使用 `video_url`、`url` 或 `source`，三者等价。远程输入仅支持 `http/https`，本地输入默认限制在当前工作目录内；需要访问绝对路径时，通过请求参数 `media_roots` 或环境变量 `YM_MEDIA_ROOTS` 配置允许的媒体根目录。\n\n## 返回协议\n\n从 `4.1.0` 开始，所有 action 都会返回稳定协议字段：\n\n| 字段 | 说明 |\n|------|------|\n| `status` | `success` / `partial` / `skipped` / `error` |\n| `code` | 稳定机器码，例如 `ok`、`missing_source`、`output_exists`、`parse_failed` |\n| `reply` | 适合聊天展示的简短中文回复 |\n| `hint` | 面向调用方或用户的下一步建议 |\n\n原有业务字段会继续保留，例如 `output_path`、`saved_path`、`outputPath`、`manifest_path`、`info`、`captions`、`result`。\n\n默认输出目录：\n\n| 类型 | 默认目录 |\n|------|----------|\n| 压缩视频 | `output/videos` |\n| 音频 | `output/audio` |\n| 封面 | `output/thumbs` |\n\n所有会写文件的接口都支持 `overwrite`，默认 `true`。设置为 `false` 时，如果输出文件已存在会直接返回错误。\n\n## 3.0.3 更新\n\n- 支持当前工作目录内的本地视频文件输入。\n- 统一 `video_url` / `url` / `source` 三种输入字段。\n- 增加默认输出目录：`output/videos`、`output/audio`、`output/thumbs`。\n- 增加 `overwrite` 覆盖策略，避免误覆盖已有文件。\n\n## 4.0.0 更新\n\n- 新增自然语言入口 `chat`，适合 Claw 直接转发用户聊天文本。\n- 新增 `media_roots` 白名单，支持处理授权目录内的绝对路径文件。\n- 自然语言命令支持提取音频、提取封面、压缩、查看信息、JSON 流水线。\n- `chat` 返回 `reply` 和结构化 `result`，同时兼顾聊天展示和自动化消费。\n\n## 4.0.1 更新\n\n- 新增 `subtitle` 推荐入口，支持 `asr` / `ocr` / `fusion` 模式。\n- 新增 `asr` 和 `ocr` 单独调试入口。\n- 字幕统一输出 SRT-like JSON：`captionTxt`、`startTimeUs`、`endTimeUs`、`source`、`confidence`。\n- `chat` 支持“识别字幕 / 提取字幕 / 转字幕 / 生成字幕”等自然语言命令。\n\n## 4.0.2 更新\n\n- 将 `faster-whisper`、`paddlepaddle`、`paddleocr` 纳入默认 `requirements.txt`。\n- 字幕识别从“可选依赖”调整为默认安装能力。\n- 运行时仍保留缺依赖 JSON error，方便定位未重新安装依赖的环境。\n\n## 4.1.0 更新\n\n- 所有 action 统一补齐 `code`、`reply`、`hint`，方便 Claw 和后续渠道适配。\n- 增加稳定错误码：`missing_source`、`source_not_allowed`、`output_exists`、`parse_failed`、`missing_steps`、`unsupported_action`、`ffmpeg_failed`、`missing_asr_dependency`、`missing_ocr_dependency`。\n- 保持旧字段兼容，不改变现有 action 名称、HTTP endpoint 和底层媒体处理逻辑。\n\n## 4.1.1 更新\n\n- 新增 `caption_segment`，用于 emlet 字幕二次分句。\n- 默认每句最多 `12` 个字符，按强标点、弱标点、连接词和长度切分。\n- 新增 `protected_terms` 和 `protected_terms_path`，用于保护品牌词、产品名、人名、术语不被拆开。\n- `pipeline` 支持 `subtitle -> caption_segment` 串联；分句步骤未传 `caption_path` 时会自动使用上一步字幕 JSON。\n\n## 4.2.0 更新\n\n- 新增 HTTP 异步长任务接口 `/skill/jobs`，适合压缩、ASR/OCR、字幕和 pipeline 等耗时 action。\n- 任务状态持久化到 `output/jobs/<job_id>/job.json`，支持提交、轮询和列表查询。\n- 现有 `/skill/<action>` 同步接口保持不变；CLI 仍保持同步执行。\n\n### 自然语言调用\n\nAction: `chat`\n\nClaw 推荐优先调用 `chat`。复杂、确定性要求高的多步骤流程继续使用 `pipeline`。\n\n```bash\npython run.py -a chat -i '{\"message\":\"将 \\\"sample.mp4\\\" 提取音频\"}'\npython run.py -a chat -i '{\"message\":\"给 \\\"sample.mp4\\\" 提取第 3 秒封面\"}'\npython run.py -a chat -i '{\"message\":\"压缩 \\\"sample.mp4\\\"\"}'\npython run.py -a chat -i '{\"message\":\"查看 \\\"sample.mp4\\\" 信息\"}'\npython run.py -a chat -i '{\"message\":\"识别 \\\"sample.mp4\\\" 的字幕\"}'\n```\n\n处理绝对路径时需要配置媒体根目录：\n\n```bash\npython run.py -a chat -i '{\"message\":\"将 \\\"D:/AA.MP4\\\" 提取音频\",\"media_roots\":[\"D:/\"]}'\n```\n\n返回包含：\n\n| 字段 | 说明 |\n|------|------|\n| `reply` | 可直接展示给用户的聊天回复 |\n| `intent` | 识别出的意图 |\n| `action` | 实际调用的 action |\n| `params` | 传给底层 action 的参数 |\n| `result` | 底层 action 原始结果 |\n| `output_paths` | 本次生成的输出路径列表 |\n\n### HTTP 异步长任务\n\nAction 可以继续同步调用，也可以通过 job API 异步执行。推荐对 `compress`、`asr`、`ocr`、`subtitle`、`pipeline` 等长任务使用异步接口。\n\n提交任务：\n\n```bash\ncurl -X POST http://127.0.0.1:8080/skill/jobs \\\n  -H 'Content-Type: application/json' \\\n  -d '{\"action\":\"pipeline\",\"params\":{\"source\":\"sample.mp4\",\"steps\":[{\"id\":\"metadata\",\"action\":\"info\",\"enabled\":true}]}}'\n```\n\n返回：\n\n```json\n{\n  \"status\": \"queued\",\n  \"code\": \"ok\",\n  \"reply\": \"任务已提交：<job_id>\",\n  \"job_id\": \"<job_id>\",\n  \"job_path\": \"output/jobs/<job_id>/job.json\",\n  \"poll_url\": \"/skill/jobs/<job_id>\"\n}\n```\n\n轮询任务：\n\n```bash\ncurl http://127.0.0.1:8080/skill/jobs/<job_id>\n```\n\n任务状态包括：`queued`、`running`、`success`、`partial`、`skipped`、`error`。任务结果保存在 `result`，产物路径汇总在 `output_paths`。\n\n查询任务列表：\n\n```bash\ncurl 'http://127.0.0.1:8080/skill/jobs?status=success&limit=50'\n```\n\n任务文件存储在 `output/jobs/<job_id>/job.json`。服务重启后，已完成任务仍可查询；未完成的 `queued` / `running` 任务会标记为 `error`，`code=job_interrupted`。\n\n### 字幕识别\n\nAction: `subtitle`\n\n推荐使用 `subtitle`，默认 `mode=fusion`：ASR 负责主要时间轴和文本，OCR 做画面字幕校正。识别依赖随 `requirements.txt` 安装；如果环境未重新安装依赖，会返回 JSON error，不会抛未捕获异常。\n\n```bash\npython run.py -a subtitle -i '{\"source\":\"sample.mp4\",\"mode\":\"fusion\"}'\npython run.py -a subtitle -i '{\"source\":\"sample.mp4\",\"mode\":\"asr\",\"language\":\"zh\"}'\npython run.py -a subtitle -i '{\"source\":\"sample.mp4\",\"mode\":\"ocr\",\"sample_interval_sec\":1}'\n```\n\n参数：\n\n| 参数 | 类型 | 默认值 | 说明 |\n|------|------|--------|------|\n| `video_url` / `url` / `source` | string | 必填 | 远程 URL 或本地文件 |\n| `mode` | string | fusion | `asr` / `ocr` / `fusion` |\n| `language` | string | auto | ASR 语言，例：`zh` / `en` |\n| `model_size` | string | base | faster-whisper 模型规格 |\n| `sample_interval_sec` | number | 1.0 | OCR 抽帧间隔 |\n| `crop_bottom_ratio` | number | 0.35 | OCR 默认扫描画面下方比例 |\n| `output_path` | string | `output/subtitles/...` | 字幕 JSON 输出路径 |\n| `overwrite` | boolean | true | 是否覆盖已有文件 |\n\n字幕条目格式：\n\n```json\n{\n  \"captionTxt\": \"识别到的字幕文本\",\n  \"startTimeUs\": 1000000,\n  \"endTimeUs\": 2000000,\n  \"source\": \"asr\",\n  \"confidence\": 0.92\n}\n```\n\n底层调试入口：\n\n```bash\npython run.py -a asr -i '{\"source\":\"sample.mp4\",\"language\":\"zh\"}'\npython run.py -a ocr -i '{\"source\":\"sample.mp4\"}'\n```\n\n### 字幕二次分句\n\nAction: `caption_segment`\n\n`caption_segment` 是 emlet 字幕分句器：它不重新识别字幕，只处理已有 captions。默认 `max_chars=12`，输出仍是 SRT-like captions JSON。\n\n```bash\npython run.py -a caption_segment -i '{\n  \"caption_path\":\"output/subtitles/sample.captions.json\",\n  \"max_chars\":12,\n  \"protected_terms\":[\"苹果\",\"华为\",\"吉利\"]\n}'\n```\n\n也可以直接传 captions：\n\n```bash\npython run.py -a caption_segment -i '{\n  \"captions\":[\n    {\"captionTxt\":\"今天我们聊苹果华为和吉利的新产品\",\"startTimeUs\":0,\"endTimeUs\":3000000}\n  ],\n  \"max_chars\":12,\n  \"protected_terms\":\"苹果,华为,吉利\"\n}'\n```\n\n长期词库可以放在 JSON 文件中：\n\n```json\n{\n  \"brands\": [\"苹果\", \"华为\", \"吉利\"],\n  \"products\": [\"小米汽车\", \"Model Y\", \"ChatGPT\"]\n}\n```\n\n参数：\n\n| 参数 | 类型 | 默认值 | 说明 |\n|------|------|--------|------|\n| `caption_path` / `input_path` | string | - | 已有 captions JSON 文件 |\n| `captions` | array | - | 直接传入字幕条目 |\n| `max_chars` | integer | 12 | 单条字幕最大字符数 |\n| `protected_terms` | array/string | [] | 不拆开的品牌词、产品名、人名、术语 |\n| `protected_terms_path` | string | - | 当前工作目录内的保护词 JSON 文件 |\n| `auto_protect_ascii` | boolean | true | 自动保护英文、数字、型号、URL、路径 |\n| `output_path` | string | `output/subtitles/...` | 分句后字幕 JSON 输出路径 |\n| `overwrite` | boolean | true | 是否覆盖已有文件 |\n\n### 压缩视频\n\nAction: `compress`\n\n```bash\npython run.py -a compress -i '{\"source\":\"sample.mp4\",\"target_ratio\":0.1}'\n```\n\n参数：\n\n| 参数 | 类型 | 默认值 | 说明 |\n|------|------|--------|------|\n| `video_url` / `url` / `source` | string | 必填 | 远程 URL 或本地文件 |\n| `target_ratio` | number | 0.1 | 目标体积比例 |\n| `adaptive` | boolean | true | 是否自动尝试不同 CRF |\n| `crf` | integer | 24 | 非 adaptive 模式下使用 |\n| `preset` | string | veryfast | ffmpeg 编码预设 |\n| `output_path` | string | `output/videos/...` | 输出路径 |\n| `overwrite` | boolean | true | 是否覆盖已有文件 |\n\n### 提取封面\n\nAction: `thumbnail`\n\n当前支持 MP4/MOV 容器，主要适用于 H.264/H.265 视频轨道。\n\n```bash\npython run.py -a thumbnail -i '{\"source\":\"sample.mp4\",\"time_seconds\":5}'\n```\n\n参数：\n\n| 参数 | 类型 | 默认值 | 说明 |\n|------|------|--------|------|\n| `video_url` / `url` / `source` | string | 必填 | 远程 URL 或本地文件 |\n| `time_seconds` | number | 0 | 按时间点提取 |\n| `frame_number` | integer | - | 按帧号提取，优先于 `time_seconds` |\n| `save_path` | string | `output/thumbs/...` | 保存路径 |\n| `resize_width` | integer | - | 输出宽度 |\n| `quality` | integer | 85 | JPEG 质量 |\n| `overwrite` | boolean | true | 是否覆盖已有文件 |\n\n### 提取音频\n\nAction: `audio`\n\n```bash\npython run.py -a audio -i '{\"source\":\"sample.mp4\",\"format\":\"mp3\"}'\n```\n\n参数：\n\n| 参数 | 类型 | 默认值 | 说明 |\n|------|------|--------|------|\n| `video_url` / `url` / `source` | string | 必填 | 远程 URL 或本地文件 |\n| `format` | string | mp3 | mp3 / wav / aac / m4a |\n| `bitrate` | string | 128k | 音频比特率 |\n| `sample_rate` | integer | 44100 | 采样率 |\n| `channels` | integer | 2 | 声道数 |\n| `start_time` | number | - | 开始时间，秒 |\n| `duration` | number | - | 截取时长，秒 |\n| `output_path` | string | `output/audio/...` | 输出路径 |\n| `overwrite` | boolean | true | 是否覆盖已有文件 |\n\n### 批量音频\n\nAction: `audio_batch`\n\n```bash\npython run.py -a audio_batch -i '{\n  \"videos\":[\n    {\"source\":\"sample1.mp4\",\"name\":\"video1\"},\n    {\"url\":\"https://example.com/2.mp4\",\"name\":\"video2\"}\n  ],\n  \"output_dir\":\"output/audio\",\n  \"format\":\"mp3\"\n}'\n```\n\n### 批量处理\n\nAction: `batch`\n\n```bash\npython run.py -a batch -i '{\n  \"action\":\"thumbnail\",\n  \"videos\":[\n    {\"source\":\"sample1.mp4\",\"time_seconds\":5},\n    {\"url\":\"https://example.com/2.mp4\",\"time_seconds\":10}\n  ]\n}'\n```\n\n`action` 支持：`compress`、`thumbnail`、`audio`。\n\n### JSON 流水线\n\nAction: `pipeline`\n\n`steps` 是唯一流程控制入口，没有写进 `steps` 的动作不会执行。支持的 step action：`info`、`thumbnail`、`audio`、`compress`、`audio_info`、`asr`、`ocr`、`subtitle`、`caption_segment`。\n\n```bash\npython run.py -a pipeline -i '{\n  \"source\":\"sample.mp4\",\n  \"name\":\"sample\",\n  \"output_dir\":\"output/pipeline/sample\",\n  \"overwrite\":true,\n  \"steps\":[\n    {\"id\":\"metadata\",\"action\":\"info\",\"enabled\":true},\n    {\n      \"id\":\"cover\",\n      \"action\":\"thumbnail\",\n      \"enabled\":true,\n      \"params\":{\"time_seconds\":3,\"resize_width\":720}\n    },\n    {\n      \"id\":\"audio_mp3\",\n      \"action\":\"audio\",\n      \"enabled\":false,\n      \"params\":{\"format\":\"mp3\",\"bitrate\":\"128k\"}\n    }\n  ]\n}'\n```\n\n规则：\n\n| 参数 | 类型 | 默认值 | 说明 |\n|------|------|--------|------|\n| `video_url` / `url` / `source` | string | 必填 | 远程 URL 或本地文件 |\n| `name` | string | 输入文件名 | 流水线名称 |\n| `output_dir` | string | `output/pipeline/<name>` | manifest 和默认产物目录 |\n| `overwrite` | boolean | true | 是否覆盖已有产物 |\n| `steps` | array | 必填 | 按 JSON 顺序执行的步骤 |\n\n每个 step 需要 `id`、`action`、`enabled`。`enabled=false` 会记录为 `skipped`。每次执行都会生成 `manifest.json`。\n\n### 获取信息\n\n```bash\npython run.py -a info -i '{\"source\":\"sample.mp4\"}'\npython run.py -a audio_info -i '{\"source\":\"sample.mp4\"}'\n```\n\n## 注意\n\n- 远程输入仅支持 `http` / `https`。\n- 本地输入路径默认限制在当前工作目录内；通过 `media_roots` / `YM_MEDIA_ROOTS` 可授权额外媒体根目录。\n- 输出路径限制在当前工作目录内。\n- HTTP 服务默认只绑定本机地址。\n\nFile v4.2.0:_meta.json\n\n{\n  \"ownerId\": \"kn787sam7qk7fffsjc875yxp0984rpzh\",\n  \"slug\": \"ym-mediatoolkit\",\n  \"version\": \"4.2.0\",\n  \"publishedAt\": 1780886166669\n}\n\nFile v4.2.0:MAINTENANCE.md\n\n# YM MediaToolkit 维护手册\n\n本文档面向后续维护者，记录版本发布、测试验证、配置边界和排障流程。\n\n## 发布流程\n\n1. 同步版本号：\n   - `SKILL.md` front matter 的 `version`\n   - `skill.json` 的 `version`\n2. 更新用户文档：\n   - 新增 action 时同步更新 `SKILL.md` 和 `skill.json`\n   - 新增参数时同步补充默认值、用途和安全限制\n   - 变更行为时在 `SKILL.md` 的版本更新段落记录\n3. 运行验证：\n\n```bash\npython3 -B -m py_compile run.py utils.py intent_parser.py audio_extractor.py frame_extractor.py video_compressor.py asr_engine.py ocr_engine.py subtitle_extractor.py caption_segmenter.py job_manager.py tests/test_release_behaviors.py scripts/smoke_test.py\npython3 -m json.tool skill.json\npython3 -B -m unittest discover -s tests\npython3 scripts/smoke_test.py\n```\n\n4. 清理产物：\n\n```bash\nfind . -type d -name __pycache__ -prune -exec rm -r {} +\nfind . -name .DS_Store -delete\n```\n\n## 当前接口分层\n\n- `chat`：Claw 推荐入口，接收自然语言，解析后调用现有 action。\n- `pipeline`：确定性 JSON 流水线入口，适合多步骤、可重复的自动化流程。\n- `audio` / `thumbnail` / `compress` / `info` / `audio_info`：底层单步能力。\n- `subtitle` / `asr` / `ocr`：字幕识别能力，输出 SRT-like captions JSON。\n- `caption_segment`：emlet 字幕二次分句器，处理已有 captions，不重新识别媒体。\n- `batch` / `audio_batch`：批量处理入口。\n- `/skill/jobs`：HTTP 异步长任务入口，适合压缩、字幕识别、pipeline 等耗时 action。\n\n维护原则：新体验优先接到 `chat` 或 `pipeline`，媒体处理逻辑继续复用底层 action，避免重复实现 ffmpeg 调用。\n\n## Action 返回协议\n\n所有 action 必须返回 JSON 对象，并保留以下协议字段：\n\n- `status`：`success` / `partial` / `skipped` / `error`\n- `code`：稳定机器码，成功为 `ok`\n- `reply`：适合聊天展示的中文回复\n- `hint`：下一步建议或排障提示\n\n新增 handler 时只需要返回原始业务结果，`run.py` 的 action protocol wrapper 会补齐缺失字段。若底层模块已经返回 `code`，包装层会保留该值。\n\n常用错误码：\n\n| code | 场景 |\n|------|------|\n| `missing_source` | 缺少 `video_url` / `url` / `source` |\n| `source_not_allowed` | 本地路径不在当前工作目录或 `media_roots` 内 |\n| `output_exists` | 输出文件已存在且 `overwrite=false` |\n| `parse_failed` | `chat` 无法识别自然语言意图 |\n| `missing_steps` | `pipeline` 未传 `steps` |\n| `missing_captions` | `caption_segment` 未传 `captions` 或 `caption_path` |\n| `invalid_action` | 异步任务或 CLI 传入不支持的 action |\n| `invalid_job_id` | job id 格式不合法 |\n| `job_not_found` | job 文件不存在 |\n| `job_interrupted` | 服务重启或进程中断导致未完成 job 失效 |\n| `unsupported_action` | action 或 pipeline step 不支持 |\n| `invalid_step` | pipeline step 结构不合法 |\n| `ffmpeg_failed` | ffmpeg / ffprobe 缺失或执行失败 |\n| `missing_asr_dependency` | ASR 依赖缺失 |\n| `missing_ocr_dependency` | OCR 依赖缺失 |\n\n## HTTP 异步任务维护\n\n异步任务逻辑在 `job_manager.py`。HTTP 服务启动时会创建单 worker 队列，串行执行提交到 `/skill/jobs` 的 action。\n\n维护规则：\n\n- job 文件存储在 `output/jobs/<job_id>/job.json`。\n- job id 使用 32 位十六进制 UUID。\n- job 文件必须保留 `job_id`、`action`、`params`、`status`、`code`、`reply`、`created_at`、`started_at`、`finished_at`、`result`、`output_paths`、`error`。\n- 当前版本不支持取消任务，不记录进度百分比。\n- 服务重启后不恢复未完成任务；旧的 `queued` / `running` 会标记为 `error`，`code=job_interrupted`。\n- 新增 action 时，只要加入 `ACTIONS`，异步任务会自动支持。\n\n## media_roots 配置\n\n本地输入默认只允许当前工作目录内的文件。需要处理绝对路径时，必须配置媒体根目录白名单。\n\n请求级配置：\n\n```json\n{\n  \"message\": \"将 \\\"D:/AA.MP4\\\" 提取音频\",\n  \"media_roots\": [\"D:/\"]\n}\n```\n\n环境变量配置：\n\n```bash\nexport YM_MEDIA_ROOTS=\"/Users/me/Videos;/Volumes/Media\"\n```\n\n规则：\n\n- `media_roots` 优先级高于 `YM_MEDIA_ROOTS`\n- 未配置时只允许当前工作目录\n- 支持用逗号或分号分隔多个根目录\n- URL 仍只允许 `http` / `https`\n- 输出路径仍限制在当前工作目录内\n\n## 自然语言解析维护\n\n解析逻辑在 `intent_parser.py`。\n\n当前支持：\n\n- 提取音频：`将 \"sample.mp4\" 提取音频`\n- 提取封面：`给 \"sample.mp4\" 提取第 3 秒封面`\n- 压缩：`压缩 \"sample.mp4\"`\n- 查看信息：`查看 \"sample.mp4\" 信息`\n- JSON pipeline：消息本身是包含 `steps` 的 JSON\n- 字幕识别：`识别 \"sample.mp4\" 的字幕`\n\n新增自然语言规则时需要同时补：\n\n- `tests/test_release_behaviors.py` 的解析测试\n- `scripts/smoke_test.py` 的真实链路测试，若会产生文件\n- `SKILL.md` 的示例\n- `skill.json` 的 action schema 或 examples，若公开接口有变化\n\n## 验证策略\n\n单元测试覆盖轻量行为：\n\n- 输入字段兼容：`video_url` / `url` / `source`\n- 本地路径和 `media_roots` 权限\n- 默认输出目录和 `overwrite=false`\n- pipeline 顺序、跳过、失败继续、manifest\n- chat 解析、执行、失败不调用 handler\n\nSmoke test 覆盖真实 ffmpeg 链路：\n\n- 生成 1 秒本地测试视频\n- 跑 `info`、`thumbnail`、`audio`、`compress`、`batch`\n- 跑 `pipeline`\n- 跑 `chat` 的音频、封面、压缩、信息命令\n- 跑 `subtitle`；若环境未重新安装依赖，需返回 dependency error；安装完整依赖后验证 captions JSON\n\n## 字幕识别维护\n\n字幕输出统一使用 camelCase 和微秒整数：\n\n```json\n{\n  \"captionTxt\": \"字幕文本\",\n  \"startTimeUs\": 1000000,\n  \"endTimeUs\": 2000000,\n  \"source\": \"asr\",\n  \"confidence\": 0.92\n}\n```\n\n字幕识别依赖已进入默认 `requirements.txt`：\n\n```bash\npip install -r requirements.txt\n```\n\n维护规则：\n\n- `subtitle` 是推荐入口，`asr` / `ocr` 用于单独调试。\n- `mode=fusion` 默认 ASR 为主，OCR 只做高置信文本校正。\n- 环境未安装完整依赖时必须返回 JSON error，不能抛未捕获异常。\n- 新增字幕准确率策略时，优先补 `subtitle_extractor.py` 的纯函数测试。\n\n## emlet 字幕分句维护\n\n二次分句逻辑在 `caption_segmenter.py`，对已有 captions 做后处理，不调用 ASR/OCR。\n\n默认策略：\n\n- `max_chars=12`\n- 强标点优先切分：`。！？!?`\n- 弱标点其次：`，、；：,;:`\n- 再尝试连接词边界，最后按长度切分\n- `protected_terms` 和 `protected_terms_path` 内的词不允许被拆开\n- `auto_protect_ascii=true` 时自动保护英文、数字、型号、URL、路径\n\n维护规则：\n\n- 新增分句策略时先补纯函数测试，确保时间轴连续、不重叠、不倒退。\n- `caption_segment` 输出继续使用 `captionTxt`、`startTimeUs`、`endTimeUs`。\n- pipeline 中 `caption_segment` 如果没有显式传 `caption_path`，会自动使用前面步骤生成的 `.captions.json`。\n\n## 常见问题\n\n`本地输入路径超出允许的 media_roots`\n\n确认文件路径位于当前工作目录，或在请求中传入 `media_roots`。\n\n`输出路径超出工作目录`\n\n输出文件必须写入当前工作目录下，例如 `output/audio/demo.mp3`。\n\n`ffmpeg 错误，返回码: ...`\n\n先确认输入文件可播放，再用 smoke test 验证当前环境的 ffmpeg 是否可用。\n\n`DNS 解析失败` 或 `禁止访问私有/内网 IP`\n\n远程 URL 会做安全校验，不允许内网、回环、链路本地地址或无法验证的目标。\n\n`chat` 没识别出命令\n\n优先使用明确句式：`将 \"sample.mp4\" 提取音频`、`给 \"sample.mp4\" 提取第 3 秒封面`、`压缩 \"sample.mp4\"`。\n\n`缺少 ASR/OCR 依赖`\n\n重新运行 `pip install -r requirements.txt`，确认 `faster-whisper`、`paddlepaddle`、`paddleocr` 已安装。\n\nFile v4.2.0:skill-card.md\n\n## Description: <br>\nNatural-language media assistant for video compression, MP4/MOV thumbnail extraction, audio conversion, subtitle recognition, and JSON-driven media pipelines. <br>\n\nThis skill is ready for commercial/non-commercial use. <br>\n\n## Publisher: <br>\n[370299455cx-web](https://clawhub.ai/user/370299455cx-web) <br>\n\n### License/Terms of Use: <br>\nMIT-0 <br>\n\n\n## Use Case: <br>\nDevelopers and media operations teams use this skill to run local or HTTP-based media workflows for compression, thumbnail extraction, audio conversion, ASR/OCR subtitle generation, caption segmentation, batch processing, and ordered JSON pipelines. <br>\n\n### Deployment Geography for Use: <br>\nGlobal <br>\n\n## Known Risks and Mitigations: <br>\nRisk: Remote media fetching and local media access can reach unintended network locations or files if configured broadly. <br>\nMitigation: Install only where media processing and outbound HTTP/HTTPS fetching are acceptable, and restrict media_roots to intended directories. <br>\nRisk: The HTTP service could expose media-processing actions if bound beyond localhost without protection. <br>\nMitigation: Keep the service bound to localhost unless it is placed behind authentication and appropriate network controls. <br>\nRisk: Media outputs may overwrite existing workspace files when overwrite behavior is enabled. <br>\nMitigation: Set overwrite=false when existing files should be preserved and review output paths before running batch or pipeline jobs. <br>\nRisk: Persisted job records may retain output paths, statuses, and processing results. <br>\nMitigation: Store job outputs only in intended workspace locations and apply normal cleanup or access controls for generated artifacts. <br>\n\n\n## Reference(s): <br>\n- [ClawHub skill page](https://clawhub.ai/370299455cx-web/ym-mediatoolkit) <br>\n- [Publisher profile](https://clawhub.ai/user/370299455cx-web) <br>\n- [Skill documentation](artifact/SKILL.md) <br>\n- [Maintenance documentation](artifact/MAINTENANCE.md) <br>\n\n\n## Skill Output: <br>\n**Output Type(s):** [text, JSON, files, configuration guidance] <br>\n**Output Format:** [JSON responses with chat-ready text, stable status codes, output paths, and generated media or caption files on disk.] <br>\n**Output Parameters:** [1D] <br>\n**Other Properties Related to Output:** [Writes media outputs to configured output directories and can persist asynchronous job records under output/jobs.] <br>\n\n## Skill Version(s): <br>\n4.2.0 (source: frontmatter and server release evidence) <br>\n\n## Ethical Considerations: <br>\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment. <br>\n\nFile v4.2.0:skill.json\n\n{\n  \"name\": \"ym-mediatoolkit\",\n  \"version\": \"4.2.0\",\n  \"description\": \"自然语言媒体助手：视频压缩、MP4/MOV 封面提取、音频转换、字幕识别、JSON 流水线\",\n  \"author\": \"your_name\",\n  \"entrypoint\": \"python run.py --input {input_json}\",\n  \"http_port\": 8080,\n  \"http_bind\": \"127.0.0.1 (默认，仅本地；可通过 --host 0.0.0.0 改为公网，需反向代理认证)\",\n  \"external_binaries\": {\n    \"ffmpeg\": \"必需 - 视频压缩、音频提取、流式处理\",\n    \"ffprobe\": \"必需 - 获取视频/音频流信息\"\n  },\n  \"python_dependencies\": {\n    \"faster-whisper\": \"必需 - ASR 字幕识别\",\n    \"paddlepaddle\": \"必需 - PaddleOCR 推理运行时\",\n    \"paddleocr\": \"必需 - OCR 字幕识别\"\n  },\n  \"response_protocol\": {\n    \"description\": \"所有 action 都会返回稳定协议字段，并保留原有业务字段。\",\n    \"fields\": {\n      \"status\": \"success / partial / skipped / error\",\n      \"code\": \"稳定机器码，例如 ok、missing_source、output_exists、parse_failed\",\n      \"reply\": \"适合聊天展示的简短中文回复\",\n      \"hint\": \"面向调用方或用户的下一步建议\"\n    },\n    \"common_codes\": [\n      \"ok\",\n      \"missing_source\",\n      \"source_not_allowed\",\n      \"output_exists\",\n      \"parse_failed\",\n      \"missing_steps\",\n      \"missing_captions\",\n      \"invalid_action\",\n      \"invalid_params\",\n      \"invalid_json\",\n      \"invalid_job_id\",\n      \"job_not_found\",\n      \"job_interrupted\",\n      \"unsupported_action\",\n      \"invalid_step\",\n      \"ffmpeg_failed\",\n      \"missing_asr_dependency\",\n      \"missing_ocr_dependency\"\n    ]\n  },\n  \"http_apis\": [\n    {\n      \"path\": \"/skill/jobs\",\n      \"method\": \"POST\",\n      \"description\": \"提交异步长任务，现有同步 /skill/<action> 接口保持不变\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"required\": [\"action\"],\n        \"properties\": {\n          \"action\": {\n            \"type\": \"string\",\n            \"description\": \"当前 actions 中支持的 action 名称\"\n          },\n          \"params\": {\n            \"type\": \"object\",\n            \"default\": {},\n            \"description\": \"传给 action 的参数\"\n          }\n        }\n      },\n      \"response_schema\": {\n        \"type\": \"object\",\n        \"properties\": {\n          \"status\": {\"type\": \"string\", \"enum\": [\"queued\", \"error\"]},\n          \"code\": {\"type\": \"string\"},\n          \"reply\": {\"type\": \"string\"},\n          \"job_id\": {\"type\": \"string\"},\n          \"job_path\": {\"type\": \"string\"},\n          \"poll_url\": {\"type\": \"string\"}\n        }\n      }\n    },\n    {\n      \"path\": \"/skill/jobs/<job_id>\",\n      \"method\": \"GET\",\n      \"description\": \"查询单个异步任务状态、结果和输出路径\"\n    },\n    {\n      \"path\": \"/skill/jobs\",\n      \"method\": \"GET\",\n      \"description\": \"查询异步任务列表，支持 status 和 limit query 参数\"\n    }\n  ],\n  \"actions\": [\n    {\n      \"name\": \"chat\",\n      \"description\": \"自然语言媒体助手入口，将用户聊天文本解析为媒体处理 action 并返回聊天回复\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"required\": [\"message\"],\n        \"properties\": {\n          \"message\": {\n            \"type\": \"string\",\n            \"description\": \"自然语言命令，例如：将 \\\"D:/AA.MP4\\\" 提取音频\"\n          },\n          \"media_roots\": {\n            \"type\": [\"array\", \"string\"],\n            \"description\": \"允许访问的本地媒体根目录；未传时读取 YM_MEDIA_ROOTS，否则仅允许当前工作目录\"\n          },\n          \"overwrite\": {\"type\": \"boolean\", \"default\": true},\n          \"output_dir\": {\"type\": \"string\", \"description\": \"自然语言命令的默认输出目录\"}\n        }\n      },\n      \"examples\": [\n        {\"message\": \"将 \\\"sample.mp4\\\" 提取音频\"},\n        {\"message\": \"给 \\\"sample.mp4\\\" 提取第 3 秒封面\"},\n        {\"message\": \"压缩 \\\"sample.mp4\\\"\"},\n        {\"message\": \"查看 \\\"sample.mp4\\\" 信息\"}\n      ]\n    },\n    {\n      \"name\": \"compress\",\n      \"description\": \"流式压缩视频，保持清晰度，支持自适应 CRF\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"anyOf\": [\n          {\"required\": [\"video_url\"]},\n          {\"required\": [\"url\"]},\n          {\"required\": [\"source\"]}\n        ],\n        \"properties\": {\n          \"video_url\": {\"type\": \"string\", \"description\": \"视频 URL 或本地文件路径\"},\n          \"url\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"source\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"media_roots\": {\"type\": [\"array\", \"string\"], \"description\": \"允许访问的本地媒体根目录\"},\n          \"target_ratio\": {\"type\": \"number\", \"default\": 0.1},\n          \"adaptive\": {\"type\": \"boolean\", \"default\": true},\n          \"crf\": {\"type\": \"integer\", \"default\": 24},\n          \"preset\": {\"type\": \"string\", \"default\": \"veryfast\"},\n          \"output_path\": {\"type\": \"string\", \"description\": \"默认 output/videos\"},\n          \"overwrite\": {\"type\": \"boolean\", \"default\": true}\n        }\n      }\n    },\n    {\n      \"name\": \"thumbnail\",\n      \"description\": \"从 MP4/MOV 视频任意时间点或帧号提取封面，流式只下载必要部分\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"anyOf\": [\n          {\"required\": [\"video_url\"]},\n          {\"required\": [\"url\"]},\n          {\"required\": [\"source\"]}\n        ],\n        \"properties\": {\n          \"video_url\": {\"type\": \"string\", \"description\": \"视频 URL 或本地文件路径\"},\n          \"url\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"source\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"media_roots\": {\"type\": [\"array\", \"string\"], \"description\": \"允许访问的本地媒体根目录\"},\n          \"time_seconds\": {\"type\": \"number\"},\n          \"frame_number\": {\"type\": \"integer\"},\n          \"save_path\": {\"type\": \"string\"},\n          \"resize_width\": {\"type\": \"integer\"},\n          \"quality\": {\"type\": \"integer\", \"default\": 85},\n          \"overwrite\": {\"type\": \"boolean\", \"default\": true}\n        }\n      }\n    },\n    {\n      \"name\": \"audio\",\n      \"description\": \"流式提取音频，支持 MP3/WAV/AAC/M4A 格式\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"anyOf\": [\n          {\"required\": [\"video_url\"]},\n          {\"required\": [\"url\"]},\n          {\"required\": [\"source\"]}\n        ],\n        \"properties\": {\n          \"video_url\": {\"type\": \"string\", \"description\": \"视频 URL 或本地文件路径\"},\n          \"url\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"source\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"media_roots\": {\"type\": [\"array\", \"string\"], \"description\": \"允许访问的本地媒体根目录\"},\n          \"format\": {\"type\": \"string\", \"enum\": [\"mp3\", \"wav\", \"aac\", \"m4a\"], \"default\": \"mp3\"},\n          \"bitrate\": {\"type\": \"string\", \"default\": \"128k\", \"description\": \"比特率: 128k, 192k, 320k\"},\n          \"sample_rate\": {\"type\": \"integer\", \"default\": 44100, \"description\": \"采样率: 44100, 48000\"},\n          \"channels\": {\"type\": \"integer\", \"default\": 2, \"description\": \"声道: 1=单声道, 2=立体声\"},\n          \"start_time\": {\"type\": \"number\", \"description\": \"开始时间（秒）\"},\n          \"duration\": {\"type\": \"number\", \"description\": \"持续时间（秒）\"},\n          \"output_path\": {\"type\": \"string\", \"description\": \"默认 output/audio\"},\n          \"overwrite\": {\"type\": \"boolean\", \"default\": true}\n        }\n      }\n    },\n    {\n      \"name\": \"audio_batch\",\n      \"description\": \"批量提取多个视频的音频\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"required\": [\"videos\"],\n        \"properties\": {\n          \"videos\": {\"type\": \"array\", \"description\": \"视频列表，条目支持 video_url/url/source\"},\n          \"media_roots\": {\"type\": [\"array\", \"string\"], \"description\": \"允许访问的本地媒体根目录\"},\n          \"output_dir\": {\"type\": \"string\", \"default\": \"output/audio\"},\n          \"format\": {\"type\": \"string\", \"default\": \"mp3\"},\n          \"bitrate\": {\"type\": \"string\", \"default\": \"128k\"},\n          \"sample_rate\": {\"type\": \"integer\", \"default\": 44100},\n          \"overwrite\": {\"type\": \"boolean\", \"default\": true}\n        }\n      }\n    },\n    {\n      \"name\": \"audio_info\",\n      \"description\": \"获取视频的音频流信息\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"anyOf\": [\n          {\"required\": [\"video_url\"]},\n          {\"required\": [\"url\"]},\n          {\"required\": [\"source\"]}\n        ],\n        \"properties\": {\n          \"video_url\": {\"type\": \"string\", \"description\": \"视频 URL 或本地文件路径\"},\n          \"url\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"source\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"media_roots\": {\"type\": [\"array\", \"string\"], \"description\": \"允许访问的本地媒体根目录\"}\n        }\n      }\n    },\n    {\n      \"name\": \"asr\",\n      \"description\": \"从音频轨识别字幕，输出 SRT-like captions JSON\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"anyOf\": [\n          {\"required\": [\"video_url\"]},\n          {\"required\": [\"url\"]},\n          {\"required\": [\"source\"]}\n        ],\n        \"properties\": {\n          \"video_url\": {\"type\": \"string\", \"description\": \"视频 URL 或本地文件路径\"},\n          \"url\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"source\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"language\": {\"type\": \"string\", \"default\": \"auto\"},\n          \"model_size\": {\"type\": \"string\", \"default\": \"base\"},\n          \"output_path\": {\"type\": \"string\", \"description\": \"默认 output/subtitles/<name>.captions.json\"},\n          \"overwrite\": {\"type\": \"boolean\", \"default\": true},\n          \"media_roots\": {\"type\": [\"array\", \"string\"], \"description\": \"允许访问的本地媒体根目录\"}\n        }\n      }\n    },\n    {\n      \"name\": \"ocr\",\n      \"description\": \"从视频画面识别硬字幕或屏幕文字，输出 SRT-like captions JSON\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"anyOf\": [\n          {\"required\": [\"video_url\"]},\n          {\"required\": [\"url\"]},\n          {\"required\": [\"source\"]}\n        ],\n        \"properties\": {\n          \"video_url\": {\"type\": \"string\", \"description\": \"视频 URL 或本地文件路径\"},\n          \"url\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"source\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"sample_interval_sec\": {\"type\": \"number\", \"default\": 1.0},\n          \"crop_bottom_ratio\": {\"type\": \"number\", \"default\": 0.35},\n          \"output_path\": {\"type\": \"string\", \"description\": \"默认 output/subtitles/<name>.captions.json\"},\n          \"overwrite\": {\"type\": \"boolean\", \"default\": true},\n          \"media_roots\": {\"type\": [\"array\", \"string\"], \"description\": \"允许访问的本地媒体根目录\"}\n        }\n      }\n    },\n    {\n      \"name\": \"subtitle\",\n      \"description\": \"推荐字幕识别入口，支持 asr / ocr / fusion 模式，统一输出 captions JSON\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"anyOf\": [\n          {\"required\": [\"video_url\"]},\n          {\"required\": [\"url\"]},\n          {\"required\": [\"source\"]}\n        ],\n        \"properties\": {\n          \"video_url\": {\"type\": \"string\", \"description\": \"视频 URL 或本地文件路径\"},\n          \"url\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"source\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"mode\": {\"type\": \"string\", \"enum\": [\"asr\", \"ocr\", \"fusion\"], \"default\": \"fusion\"},\n          \"language\": {\"type\": \"string\", \"default\": \"auto\"},\n          \"model_size\": {\"type\": \"string\", \"default\": \"base\"},\n          \"sample_interval_sec\": {\"type\": \"number\", \"default\": 1.0},\n          \"crop_bottom_ratio\": {\"type\": \"number\", \"default\": 0.35},\n          \"output_path\": {\"type\": \"string\", \"description\": \"默认 output/subtitles/<name>.captions.json\"},\n          \"overwrite\": {\"type\": \"boolean\", \"default\": true},\n          \"media_roots\": {\"type\": [\"array\", \"string\"], \"description\": \"允许访问的本地媒体根目录\"}\n        }\n      },\n      \"examples\": [\n        {\"source\": \"sample.mp4\", \"mode\": \"fusion\"},\n        {\"source\": \"sample.mp4\", \"mode\": \"asr\", \"language\": \"zh\"}\n      ]\n    },\n    {\n      \"name\": \"caption_segment\",\n      \"description\": \"emlet 字幕二次分句器，对已有 captions JSON 按长度、标点和保护词重新切分\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"anyOf\": [\n          {\"required\": [\"caption_path\"]},\n          {\"required\": [\"input_path\"]},\n          {\"required\": [\"captions\"]}\n        ],\n        \"properties\": {\n          \"caption_path\": {\"type\": \"string\", \"description\": \"已有 captions JSON 文件路径\"},\n          \"input_path\": {\"type\": \"string\", \"description\": \"同 caption_path\"},\n          \"captions\": {\"type\": \"array\", \"description\": \"直接传入的字幕条目\"},\n          \"max_chars\": {\"type\": \"integer\", \"default\": 12, \"description\": \"单条字幕最大字符数\"},\n          \"protected_terms\": {\n            \"type\": [\"array\", \"string\"],\n            \"description\": \"不拆开的品牌词、产品名、人名、术语，例如 苹果,华为,吉利\"\n          },\n          \"protected_terms_path\": {\n            \"type\": \"string\",\n            \"description\": \"当前工作目录内的保护词 JSON 文件，可按 brands/products 等字段分组\"\n          },\n          \"auto_protect_ascii\": {\n            \"type\": \"boolean\",\n            \"default\": true,\n            \"description\": \"自动保护英文、数字、型号、URL、路径\"\n          },\n          \"output_path\": {\"type\": \"string\", \"description\": \"默认 output/subtitles/<name>.segmented.captions.json\"},\n          \"overwrite\": {\"type\": \"boolean\", \"default\": true}\n        }\n      },\n      \"examples\": [\n        {\n          \"caption_path\": \"output/subtitles/sample.captions.json\",\n          \"max_chars\": 12,\n          \"protected_terms\": [\"苹果\", \"华为\", \"吉利\"]\n        }\n      ]\n    },\n    {\n      \"name\": \"batch\",\n      \"description\": \"批量处理多个视频，支持 compress / thumbnail / audio\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"required\": [\"videos\"],\n        \"properties\": {\n          \"action\": {\"type\": \"string\", \"enum\": [\"compress\", \"thumbnail\", \"audio\"], \"default\": \"thumbnail\"},\n          \"videos\": {\"type\": \"array\", \"description\": \"按 action 传入对应参数对象列表，条目支持 video_url/url/source\"},\n          \"media_roots\": {\"type\": [\"array\", \"string\"], \"description\": \"允许访问的本地媒体根目录\"},\n          \"output_dir\": {\"type\": \"string\", \"description\": \"批量输出目录\"},\n          \"overwrite\": {\"type\": \"boolean\", \"default\": true}\n        }\n      }\n    },\n    {\n      \"name\": \"pipeline\",\n      \"description\": \"JSON 驱动媒体流水线，按 steps 顺序编排 info / thumbnail / audio / compress / audio_info / subtitle / caption_segment\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"anyOf\": [\n          {\"required\": [\"video_url\"]},\n          {\"required\": [\"url\"]},\n          {\"required\": [\"source\"]}\n        ],\n        \"required\": [\"steps\"],\n        \"properties\": {\n          \"video_url\": {\"type\": \"string\", \"description\": \"视频 URL 或本地文件路径\"},\n          \"url\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"source\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"media_roots\": {\"type\": [\"array\", \"string\"], \"description\": \"允许访问的本地媒体根目录\"},\n          \"name\": {\"type\": \"string\", \"description\": \"流水线名称\"},\n          \"output_dir\": {\"type\": \"string\", \"description\": \"默认 output/pipeline/<name>\"},\n          \"overwrite\": {\"type\": \"boolean\", \"default\": true},\n          \"steps\": {\n            \"type\": \"array\",\n            \"description\": \"步骤列表，每步包含 id/action/enabled/params\",\n            \"items\": {\n              \"type\": \"object\",\n              \"required\": [\"id\", \"action\", \"enabled\"],\n              \"properties\": {\n                \"id\": {\"type\": \"string\"},\n                \"action\": {\"type\": \"string\", \"enum\": [\"info\", \"thumbnail\", \"audio\", \"compress\", \"audio_info\", \"asr\", \"ocr\", \"subtitle\", \"caption_segment\"]},\n                \"enabled\": {\"type\": \"boolean\"},\n                \"params\": {\"type\": \"object\"}\n              }\n            }\n          }\n        }\n      }\n    },\n    {\n      \"name\": \"info\",\n      \"description\": \"获取完整视频信息（分辨率、时长、编码等）\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"anyOf\": [\n          {\"required\": [\"video_url\"]},\n          {\"required\": [\"url\"]},\n          {\"required\": [\"source\"]}\n        ],\n        \"properties\": {\n          \"video_url\": {\"type\": \"string\", \"description\": \"视频 URL 或本地文件路径\"},\n          \"url\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"source\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"media_roots\": {\"type\": [\"array\", \"string\"], \"description\": \"允许访问的本地媒体根目录\"}\n        }\n      }\n    }\n  ]\n}\n\nFile v4.2.0:requirements.txt\n\nrequests>=2.28.0\nopencv-python>=4.8.0\nnumpy>=1.24.0\naiohttp>=3.8.0\nflask>=2.3.0\nflask-cors>=4.0.0\nfaster-whisper>=1.0.0\npaddlepaddle>=2.6.0\npaddleocr>=2.7.0\n\nArchive v4.0.1: 17 files, 45713 bytes\n\nFiles: asr_engine.py (3005b), audio_extractor.py (9638b), frame_extractor.py (22227b), intent_parser.py (6678b), MAINTENANCE.md (5987b), ocr_engine.py (3045b), requirements.txt (157b), run.py (36241b), scripts/smoke_test.py (6216b), skill-card.md (2523b), skill.json (13903b), SKILL.md (10969b), subtitle_extractor.py (5490b), tests/test_release_behaviors.py (19501b), utils.py (8193b), video_compressor.py (6897b), _meta.json (134b)\n\nFile v4.0.1:SKILL.md\n\n---\nname: ym-mediatoolkit\nversion: 4.1.0\ndescription: 自然语言媒体助手 - 视频压缩、MP4/MOV 封面提取、音频转换、字幕识别\nauthor: your_name\ntags:\n  - video\n  - compression\n  - thumbnail\n  - audio\n  - streaming\n  - ffmpeg\ncategories:\n  - media\n  - utility\nclawhub:\n  entrypoint: python run.py\n  runtime: python3\n  http_port: 8080\n---\n\n# YM MediaToolkit\n\n自然语言媒体助手，支持从远程视频 URL、当前工作目录内的本地视频文件、或配置过 `media_roots` 的本地媒体目录直接处理：\n\n- 自然语言调用\n- 视频压缩\n- MP4/MOV 封面提取\n- 音频提取与转换：MP3 / WAV / AAC / M4A\n- OCR / ASR 字幕识别\n- 批量处理\n- JSON 驱动媒体流水线\n\n## 依赖\n\n```bash\npip install -r requirements.txt\n```\n\n系统需要安装：\n\n```bash\nffmpeg\nffprobe\n```\n\n字幕识别依赖已内置在 `requirements.txt`，包括 `faster-whisper`、`paddlepaddle`、`paddleocr`。\n\n## 维护\n\n维护、发布、测试和排障流程见 [MAINTENANCE.md](./MAINTENANCE.md)。\n\n## 命令行\n\n```bash\npython run.py -a <action> -i '<json>'\n```\n\n也可以从 JSON 文件读取参数：\n\n```bash\npython run.py -i params.json\n```\n\n## HTTP 服务\n\n```bash\npython run.py --serve\n```\n\n默认监听 `127.0.0.1:8080`。\n\n健康检查：\n\n```http\nGET /health\n```\n\n## 功能\n\n输入源字段可使用 `video_url`、`url` 或 `source`，三者等价。远程输入仅支持 `http/https`，本地输入默认限制在当前工作目录内；需要访问绝对路径时，通过请求参数 `media_roots` 或环境变量 `YM_MEDIA_ROOTS` 配置允许的媒体根目录。\n\n## 返回协议\n\n从 `4.1.0` 开始，所有 action 都会返回稳定协议字段：\n\n| 字段 | 说明 |\n|------|------|\n| `status` | `success` / `partial` / `skipped` / `error` |\n| `code` | 稳定机器码，例如 `ok`、`missing_source`、`output_exists`、`parse_failed` |\n| `reply` | 适合聊天展示的简短中文回复 |\n| `hint` | 面向调用方或用户的下一步建议 |\n\n原有业务字段会继续保留，例如 `output_path`、`saved_path`、`outputPath`、`manifest_path`、`info`、`captions`、`result`。\n\n默认输出目录：\n\n| 类型 | 默认目录 |\n|------|----------|\n| 压缩视频 | `output/videos` |\n| 音频 | `output/audio` |\n| 封面 | `output/thumbs` |\n\n所有会写文件的接口都支持 `overwrite`，默认 `true`。设置为 `false` 时，如果输出文件已存在会直接返回错误。\n\n## 3.0.3 更新\n\n- 支持当前工作目录内的本地视频文件输入。\n- 统一 `video_url` / `url` / `source` 三种输入字段。\n- 增加默认输出目录：`output/videos`、`output/audio`、`output/thumbs`。\n- 增加 `overwrite` 覆盖策略，避免误覆盖已有文件。\n\n## 4.0.0 更新\n\n- 新增自然语言入口 `chat`，适合 Claw 直接转发用户聊天文本。\n- 新增 `media_roots` 白名单，支持处理授权目录内的绝对路径文件。\n- 自然语言命令支持提取音频、提取封面、压缩、查看信息、JSON 流水线。\n- `chat` 返回 `reply` 和结构化 `result`，同时兼顾聊天展示和自动化消费。\n\n## 4.0.1 更新\n\n- 新增 `subtitle` 推荐入口，支持 `asr` / `ocr` / `fusion` 模式。\n- 新增 `asr` 和 `ocr` 单独调试入口。\n- 字幕统一输出 SRT-like JSON：`captionTxt`、`startTimeUs`、`endTimeUs`、`source`、`confidence`。\n- `chat` 支持“识别字幕 / 提取字幕 / 转字幕 / 生成字幕”等自然语言命令。\n\n## 4.0.2 更新\n\n- 将 `faster-whisper`、`paddlepaddle`、`paddleocr` 纳入默认 `requirements.txt`。\n- 字幕识别从“可选依赖”调整为默认安装能力。\n- 运行时仍保留缺依赖 JSON error，方便定位未重新安装依赖的环境。\n\n## 4.1.0 更新\n\n- 所有 action 统一补齐 `code`、`reply`、`hint`，方便 Claw 和后续渠道适配。\n- 增加稳定错误码：`missing_source`、`source_not_allowed`、`output_exists`、`parse_failed`、`missing_steps`、`unsupported_action`、`ffmpeg_failed`、`missing_asr_dependency`、`missing_ocr_dependency`。\n- 保持旧字段兼容，不改变现有 action 名称、HTTP endpoint 和底层媒体处理逻辑。\n\n### 自然语言调用\n\nAction: `chat`\n\nClaw 推荐优先调用 `chat`。复杂、确定性要求高的多步骤流程继续使用 `pipeline`。\n\n```bash\npython run.py -a chat -i '{\"message\":\"将 \\\"sample.mp4\\\" 提取音频\"}'\npython run.py -a chat -i '{\"message\":\"给 \\\"sample.mp4\\\" 提取第 3 秒封面\"}'\npython run.py -a chat -i '{\"message\":\"压缩 \\\"sample.mp4\\\"\"}'\npython run.py -a chat -i '{\"message\":\"查看 \\\"sample.mp4\\\" 信息\"}'\npython run.py -a chat -i '{\"message\":\"识别 \\\"sample.mp4\\\" 的字幕\"}'\n```\n\n处理绝对路径时需要配置媒体根目录：\n\n```bash\npython run.py -a chat -i '{\"message\":\"将 \\\"D:/AA.MP4\\\" 提取音频\",\"media_roots\":[\"D:/\"]}'\n```\n\n返回包含：\n\n| 字段 | 说明 |\n|------|------|\n| `reply` | 可直接展示给用户的聊天回复 |\n| `intent` | 识别出的意图 |\n| `action` | 实际调用的 action |\n| `params` | 传给底层 action 的参数 |\n| `result` | 底层 action 原始结果 |\n| `output_paths` | 本次生成的输出路径列表 |\n\n### 字幕识别\n\nAction: `subtitle`\n\n推荐使用 `subtitle`，默认 `mode=fusion`：ASR 负责主要时间轴和文本，OCR 做画面字幕校正。识别依赖随 `requirements.txt` 安装；如果环境未重新安装依赖，会返回 JSON error，不会抛未捕获异常。\n\n```bash\npython run.py -a subtitle -i '{\"source\":\"sample.mp4\",\"mode\":\"fusion\"}'\npython run.py -a subtitle -i '{\"source\":\"sample.mp4\",\"mode\":\"asr\",\"language\":\"zh\"}'\npython run.py -a subtitle -i '{\"source\":\"sample.mp4\",\"mode\":\"ocr\",\"sample_interval_sec\":1}'\n```\n\n参数：\n\n| 参数 | 类型 | 默认值 | 说明 |\n|------|------|--------|------|\n| `video_url` / `url` / `source` | string | 必填 | 远程 URL 或本地文件 |\n| `mode` | string | fusion | `asr` / `ocr` / `fusion` |\n| `language` | string | auto | ASR 语言，例：`zh` / `en` |\n| `model_size` | string | base | faster-whisper 模型规格 |\n| `sample_interval_sec` | number | 1.0 | OCR 抽帧间隔 |\n| `crop_bottom_ratio` | number | 0.35 | OCR 默认扫描画面下方比例 |\n| `output_path` | string | `output/subtitles/...` | 字幕 JSON 输出路径 |\n| `overwrite` | boolean | true | 是否覆盖已有文件 |\n\n字幕条目格式：\n\n```json\n{\n  \"captionTxt\": \"识别到的字幕文本\",\n  \"startTimeUs\": 1000000,\n  \"endTimeUs\": 2000000,\n  \"source\": \"asr\",\n  \"confidence\": 0.92\n}\n```\n\n底层调试入口：\n\n```bash\npython run.py -a asr -i '{\"source\":\"sample.mp4\",\"language\":\"zh\"}'\npython run.py -a ocr -i '{\"source\":\"sample.mp4\"}'\n```\n\n### 压缩视频\n\nAction: `compress`\n\n```bash\npython run.py -a compress -i '{\"source\":\"sample.mp4\",\"target_ratio\":0.1}'\n```\n\n参数：\n\n| 参数 | 类型 | 默认值 | 说明 |\n|------|------|--------|------|\n| `video_url` / `url` / `source` | string | 必填 | 远程 URL 或本地文件 |\n| `target_ratio` | number | 0.1 | 目标体积比例 |\n| `adaptive` | boolean | true | 是否自动尝试不同 CRF |\n| `crf` | integer | 24 | 非 adaptive 模式下使用 |\n| `preset` | string | veryfast | ffmpeg 编码预设 |\n| `output_path` | string | `output/videos/...` | 输出路径 |\n| `overwrite` | boolean | true | 是否覆盖已有文件 |\n\n### 提取封面\n\nAction: `thumbnail`\n\n当前支持 MP4/MOV 容器，主要适用于 H.264/H.265 视频轨道。\n\n```bash\npython run.py -a thumbnail -i '{\"source\":\"sample.mp4\",\"time_seconds\":5}'\n```\n\n参数：\n\n| 参数 | 类型 | 默认值 | 说明 |\n|------|------|--------|------|\n| `video_url` / `url` / `source` | string | 必填 | 远程 URL 或本地文件 |\n| `time_seconds` | number | 0 | 按时间点提取 |\n| `frame_number` | integer | - | 按帧号提取，优先于 `time_seconds` |\n| `save_path` | string | `output/thumbs/...` | 保存路径 |\n| `resize_width` | integer | - | 输出宽度 |\n| `quality` | integer | 85 | JPEG 质量 |\n| `overwrite` | boolean | true | 是否覆盖已有文件 |\n\n### 提取音频\n\nAction: `audio`\n\n```bash\npython run.py -a audio -i '{\"source\":\"sample.mp4\",\"format\":\"mp3\"}'\n```\n\n参数：\n\n| 参数 | 类型 | 默认值 | 说明 |\n|------|------|--------|------|\n| `video_url` / `url` / `source` | string | 必填 | 远程 URL 或本地文件 |\n| `format` | string | mp3 | mp3 / wav / aac / m4a |\n| `bitrate` | string | 128k | 音频比特率 |\n| `sample_rate` | integer | 44100 | 采样率 |\n| `channels` | integer | 2 | 声道数 |\n| `start_time` | number | - | 开始时间，秒 |\n| `duration` | number | - | 截取时长，秒 |\n| `output_path` | string | `output/audio/...` | 输出路径 |\n| `overwrite` | boolean | true | 是否覆盖已有文件 |\n\n### 批量音频\n\nAction: `audio_batch`\n\n```bash\npython run.py -a audio_batch -i '{\n  \"videos\":[\n    {\"source\":\"sample1.mp4\",\"name\":\"video1\"},\n    {\"url\":\"https://example.com/2.mp4\",\"name\":\"video2\"}\n  ],\n  \"output_dir\":\"output/audio\",\n  \"format\":\"mp3\"\n}'\n```\n\n### 批量处理\n\nAction: `batch`\n\n```bash\npython run.py -a batch -i '{\n  \"action\":\"thumbnail\",\n  \"videos\":[\n    {\"source\":\"sample1.mp4\",\"time_seconds\":5},\n    {\"url\":\"https://example.com/2.mp4\",\"time_seconds\":10}\n  ]\n}'\n```\n\n`action` 支持：`compress`、`thumbnail`、`audio`。\n\n### JSON 流水线\n\nAction: `pipeline`\n\n`steps` 是唯一流程控制入口，没有写进 `steps` 的动作不会执行。支持的 step action：`info`、`thumbnail`、`audio`、`compress`、`audio_info`、`asr`、`ocr`、`subtitle`。\n\n```bash\npython run.py -a pipeline -i '{\n  \"source\":\"sample.mp4\",\n  \"name\":\"sample\",\n  \"output_dir\":\"output/pipeline/sample\",\n  \"overwrite\":true,\n  \"steps\":[\n    {\"id\":\"metadata\",\"action\":\"info\",\"enabled\":true},\n    {\n      \"id\":\"cover\",\n      \"action\":\"thumbnail\",\n      \"enabled\":true,\n      \"params\":{\"time_seconds\":3,\"resize_width\":720}\n    },\n    {\n      \"id\":\"audio_mp3\",\n      \"action\":\"audio\",\n      \"enabled\":false,\n      \"params\":{\"format\":\"mp3\",\"bitrate\":\"128k\"}\n    }\n  ]\n}'\n```\n\n规则：\n\n| 参数 | 类型 | 默认值 | 说明 |\n|------|------|--------|------|\n| `video_url` / `url` / `source` | string | 必填 | 远程 URL 或本地文件 |\n| `name` | string | 输入文件名 | 流水线名称 |\n| `output_dir` | string | `output/pipeline/<name>` | manifest 和默认产物目录 |\n| `overwrite` | boolean | true | 是否覆盖已有产物 |\n| `steps` | array | 必填 | 按 JSON 顺序执行的步骤 |\n\n每个 step 需要 `id`、`action`、`enabled`。`enabled=false` 会记录为 `skipped`。每次执行都会生成 `manifest.json`。\n\n### 获取信息\n\n```bash\npython run.py -a info -i '{\"source\":\"sample.mp4\"}'\npython run.py -a audio_info -i '{\"source\":\"sample.mp4\"}'\n```\n\n## 注意\n\n- 远程输入仅支持 `http` / `https`。\n- 本地输入路径默认限制在当前工作目录内；通过 `media_roots` / `YM_MEDIA_ROOTS` 可授权额外媒体根目录。\n- 输出路径限制在当前工作目录内。\n- HTTP 服务默认只绑定本机地址。\n\nFile v4.0.1:_meta.json\n\n{\n  \"ownerId\": \"kn787sam7qk7fffsjc875yxp0984rpzh\",\n  \"slug\": \"ym-mediatoolkit\",\n  \"version\": \"4.0.1\",\n  \"publishedAt\": 1779768975212\n}\n\nFile v4.0.1:MAINTENANCE.md\n\n# YM MediaToolkit 维护手册\n\n本文档面向后续维护者，记录版本发布、测试验证、配置边界和排障流程。\n\n## 发布流程\n\n1. 同步版本号：\n   - `SKILL.md` front matter 的 `version`\n   - `skill.json` 的 `version`\n2. 更新用户文档：\n   - 新增 action 时同步更新 `SKILL.md` 和 `skill.json`\n   - 新增参数时同步补充默认值、用途和安全限制\n   - 变更行为时在 `SKILL.md` 的版本更新段落记录\n3. 运行验证：\n\n```bash\npython3 -B -m py_compile run.py utils.py intent_parser.py audio_extractor.py frame_extractor.py video_compressor.py asr_engine.py ocr_engine.py subtitle_extractor.py tests/test_release_behaviors.py scripts/smoke_test.py\npython3 -m json.tool skill.json\npython3 -B -m unittest discover -s tests\npython3 scripts/smoke_test.py\n```\n\n4. 清理产物：\n\n```bash\nfind . -type d -name __pycache__ -prune -exec rm -r {} +\nfind . -name .DS_Store -delete\n```\n\n## 当前接口分层\n\n- `chat`：Claw 推荐入口，接收自然语言，解析后调用现有 action。\n- `pipeline`：确定性 JSON 流水线入口，适合多步骤、可重复的自动化流程。\n- `audio` / `thumbnail` / `compress` / `info` / `audio_info`：底层单步能力。\n- `subtitle` / `asr` / `ocr`：字幕识别能力，输出 SRT-like captions JSON。\n- `batch` / `audio_batch`：批量处理入口。\n\n维护原则：新体验优先接到 `chat` 或 `pipeline`，媒体处理逻辑继续复用底层 action，避免重复实现 ffmpeg 调用。\n\n## Action 返回协议\n\n所有 action 必须返回 JSON 对象，并保留以下协议字段：\n\n- `status`：`success` / `partial` / `skipped` / `error`\n- `code`：稳定机器码，成功为 `ok`\n- `reply`：适合聊天展示的中文回复\n- `hint`：下一步建议或排障提示\n\n新增 handler 时只需要返回原始业务结果，`run.py` 的 action protocol wrapper 会补齐缺失字段。若底层模块已经返回 `code`，包装层会保留该值。\n\n常用错误码：\n\n| code | 场景 |\n|------|------|\n| `missing_source` | 缺少 `video_url` / `url` / `source` |\n| `source_not_allowed` | 本地路径不在当前工作目录或 `media_roots` 内 |\n| `output_exists` | 输出文件已存在且 `overwrite=false` |\n| `parse_failed` | `chat` 无法识别自然语言意图 |\n| `missing_steps` | `pipeline` 未传 `steps` |\n| `unsupported_action` | action 或 pipeline step 不支持 |\n| `invalid_step` | pipeline step 结构不合法 |\n| `ffmpeg_failed` | ffmpeg / ffprobe 缺失或执行失败 |\n| `missing_asr_dependency` | ASR 依赖缺失 |\n| `missing_ocr_dependency` | OCR 依赖缺失 |\n\n## media_roots 配置\n\n本地输入默认只允许当前工作目录内的文件。需要处理绝对路径时，必须配置媒体根目录白名单。\n\n请求级配置：\n\n```json\n{\n  \"message\": \"将 \\\"D:/AA.MP4\\\" 提取音频\",\n  \"media_roots\": [\"D:/\"]\n}\n```\n\n环境变量配置：\n\n```bash\nexport YM_MEDIA_ROOTS=\"/Users/me/Videos;/Volumes/Media\"\n```\n\n规则：\n\n- `media_roots` 优先级高于 `YM_MEDIA_ROOTS`\n- 未配置时只允许当前工作目录\n- 支持用逗号或分号分隔多个根目录\n- URL 仍只允许 `http` / `https`\n- 输出路径仍限制在当前工作目录内\n\n## 自然语言解析维护\n\n解析逻辑在 `intent_parser.py`。\n\n当前支持：\n\n- 提取音频：`将 \"sample.mp4\" 提取音频`\n- 提取封面：`给 \"sample.mp4\" 提取第 3 秒封面`\n- 压缩：`压缩 \"sample.mp4\"`\n- 查看信息：`查看 \"sample.mp4\" 信息`\n- JSON pipeline：消息本身是包含 `steps` 的 JSON\n- 字幕识别：`识别 \"sample.mp4\" 的字幕`\n\n新增自然语言规则时需要同时补：\n\n- `tests/test_release_behaviors.py` 的解析测试\n- `scripts/smoke_test.py` 的真实链路测试，若会产生文件\n- `SKILL.md` 的示例\n- `skill.json` 的 action schema 或 examples，若公开接口有变化\n\n## 验证策略\n\n单元测试覆盖轻量行为：\n\n- 输入字段兼容：`video_url` / `url` / `source`\n- 本地路径和 `media_roots` 权限\n- 默认输出目录和 `overwrite=false`\n- pipeline 顺序、跳过、失败继续、manifest\n- chat 解析、执行、失败不调用 handler\n\nSmoke test 覆盖真实 ffmpeg 链路：\n\n- 生成 1 秒本地测试视频\n- 跑 `info`、`thumbnail`、`audio`、`compress`、`batch`\n- 跑 `pipeline`\n- 跑 `chat` 的音频、封面、压缩、信息命令\n- 跑 `subtitle`；若环境未重新安装依赖，需返回 dependency error；安装完整依赖后验证 captions JSON\n\n## 字幕识别维护\n\n字幕输出统一使用 camelCase 和微秒整数：\n\n```json\n{\n  \"captionTxt\": \"字幕文本\",\n  \"startTimeUs\": 1000000,\n  \"endTimeUs\": 2000000,\n  \"source\": \"asr\",\n  \"confidence\": 0.92\n}\n```\n\n字幕识别依赖已进入默认 `requirements.txt`：\n\n```bash\npip install -r requirements.txt\n```\n\n维护规则：\n\n- `subtitle` 是推荐入口，`asr` / `ocr` 用于单独调试。\n- `mode=fusion` 默认 ASR 为主，OCR 只做高置信文本校正。\n- 环境未安装完整依赖时必须返回 JSON error，不能抛未捕获异常。\n- 新增字幕准确率策略时，优先补 `subtitle_extractor.py` 的纯函数测试。\n\n## 常见问题\n\n`本地输入路径超出允许的 media_roots`\n\n确认文件路径位于当前工作目录，或在请求中传入 `media_roots`。\n\n`输出路径超出工作目录`\n\n输出文件必须写入当前工作目录下，例如 `output/audio/demo.mp3`。\n\n`ffmpeg 错误，返回码: ...`\n\n先确认输入文件可播放，再用 smoke test 验证当前环境的 ffmpeg 是否可用。\n\n`DNS 解析失败` 或 `禁止访问私有/内网 IP`\n\n远程 URL 会做安全校验，不允许内网、回环、链路本地地址或无法验证的目标。\n\n`chat` 没识别出命令\n\n优先使用明确句式：`将 \"sample.mp4\" 提取音频`、`给 \"sample.mp4\" 提取第 3 秒封面`、`压缩 \"sample.mp4\"`。\n\n`缺少 ASR/OCR 依赖`\n\n重新运行 `pip install -r requirements.txt`，确认 `faster-whisper`、`paddlepaddle`、`paddleocr` 已安装。\n\nFile v4.0.1:skill-card.md\n\n## Description: <br>\nYM-MediaToolkit is a natural-language media assistant for video compression, MP4/MOV thumbnail extraction, audio conversion, subtitle recognition, and JSON-driven media pipelines. <br>\n\nThis skill is ready for commercial/non-commercial use. <br>\n\n## Publisher: <br>\n[370299455cx-web](https://clawhub.ai/user/370299455cx-web) <br>\n\n### License/Terms of Use: <br>\nMIT-0 <br>\n\n\n## Use Case: <br>\nDevelopers and media workflow users can use this skill to automate common local or public-URL media processing tasks through natural-language commands or structured JSON actions. It is useful for extracting audio, creating thumbnails, compressing video, identifying subtitles, and building repeatable multi-step processing pipelines. <br>\n\n### Deployment Geography for Use: <br>\nGlobal <br>\n\n## Known Risks and Mitigations: <br>\nRisk: The local HTTP service can expose media-processing actions if bound beyond localhost. <br>\nMitigation: Keep the service on 127.0.0.1 unless an authenticated reverse proxy or equivalent access control is in place. <br>\nRisk: The skill can read local media files from the current directory or configured media roots and write generated outputs. <br>\nMitigation: Use narrow media_roots values, review requested input and output paths, and keep output directories inside the intended workspace. <br>\nRisk: The skill fetches public URLs and runs media, ASR, and OCR processing on supplied content. <br>\nMitigation: Process only trusted media sources when possible and avoid using the local service while browsing untrusted sites. <br>\n\n\n## Reference(s): <br>\n- [ClawHub skill page](https://clawhub.ai/370299455cx-web/ym-mediatoolkit) <br>\n- [ClawHub publisher profile](https://clawhub.ai/user/370299455cx-web) <br>\n\n\n## Skill Output: <br>\n**Output Type(s):** [text, JSON, media files, configuration guidance] <br>\n**Output Format:** [JSON responses, chat-ready text replies, media output files, subtitles JSON, and pipeline manifests] <br>\n**Output Parameters:** [1D] <br>\n**Other Properties Related to Output:** [Writes generated assets under configured output paths and returns stable status, code, reply, and hint fields.] <br>\n\n## Skill Version(s): <br>\n4.0.1 (source: server release evidence) <br>\n\n## Ethical Considerations: <br>\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment. <br>\n\nFile v4.0.1:skill.json\n\n{\n  \"name\": \"ym-mediatoolkit\",\n  \"version\": \"4.1.0\",\n  \"description\": \"自然语言媒体助手：视频压缩、MP4/MOV 封面提取、音频转换、字幕识别、JSON 流水线\",\n  \"author\": \"your_name\",\n  \"entrypoint\": \"python run.py --input {input_json}\",\n  \"http_port\": 8080,\n  \"http_bind\": \"127.0.0.1 (默认，仅本地；可通过 --host 0.0.0.0 改为公网，需反向代理认证)\",\n  \"external_binaries\": {\n    \"ffmpeg\": \"必需 - 视频压缩、音频提取、流式处理\",\n    \"ffprobe\": \"必需 - 获取视频/音频流信息\"\n  },\n  \"python_dependencies\": {\n    \"faster-whisper\": \"必需 - ASR 字幕识别\",\n    \"paddlepaddle\": \"必需 - PaddleOCR 推理运行时\",\n    \"paddleocr\": \"必需 - OCR 字幕识别\"\n  },\n  \"response_protocol\": {\n    \"description\": \"所有 action 都会返回稳定协议字段，并保留原有业务字段。\",\n    \"fields\": {\n      \"status\": \"success / partial / skipped / error\",\n      \"code\": \"稳定机器码，例如 ok、missing_source、output_exists、parse_failed\",\n      \"reply\": \"适合聊天展示的简短中文回复\",\n      \"hint\": \"面向调用方或用户的下一步建议\"\n    },\n    \"common_codes\": [\n      \"ok\",\n      \"missing_source\",\n      \"source_not_allowed\",\n      \"output_exists\",\n      \"parse_failed\",\n      \"missing_steps\",\n      \"unsupported_action\",\n      \"invalid_step\",\n      \"ffmpeg_failed\",\n      \"missing_asr_dependency\",\n      \"missing_ocr_dependency\"\n    ]\n  },\n  \"actions\": [\n    {\n      \"name\": \"chat\",\n      \"description\": \"自然语言媒体助手入口，将用户聊天文本解析为媒体处理 action 并返回聊天回复\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"required\": [\"message\"],\n        \"properties\": {\n          \"message\": {\n            \"type\": \"string\",\n            \"description\": \"自然语言命令，例如：将 \\\"D:/AA.MP4\\\" 提取音频\"\n          },\n          \"media_roots\": {\n            \"type\": [\"array\", \"string\"],\n            \"description\": \"允许访问的本地媒体根目录；未传时读取 YM_MEDIA_ROOTS，否则仅允许当前工作目录\"\n          },\n          \"overwrite\": {\"type\": \"boolean\", \"default\": true},\n          \"output_dir\": {\"type\": \"string\", \"description\": \"自然语言命令的默认输出目录\"}\n        }\n      },\n      \"examples\": [\n        {\"message\": \"将 \\\"sample.mp4\\\" 提取音频\"},\n        {\"message\": \"给 \\\"sample.mp4\\\" 提取第 3 秒封面\"},\n        {\"message\": \"压缩 \\\"sample.mp4\\\"\"},\n        {\"message\": \"查看 \\\"sample.mp4\\\" 信息\"}\n      ]\n    },\n    {\n      \"name\": \"compress\",\n      \"description\": \"流式压缩视频，保持清晰度，支持自适应 CRF\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"anyOf\": [\n          {\"required\": [\"video_url\"]},\n          {\"required\": [\"url\"]},\n          {\"required\": [\"source\"]}\n        ],\n        \"properties\": {\n          \"video_url\": {\"type\": \"string\", \"description\": \"视频 URL 或本地文件路径\"},\n          \"url\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"source\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"media_roots\": {\"type\": [\"array\", \"string\"], \"description\": \"允许访问的本地媒体根目录\"},\n          \"target_ratio\": {\"type\": \"number\", \"default\": 0.1},\n          \"adaptive\": {\"type\": \"boolean\", \"default\": true},\n          \"crf\": {\"type\": \"integer\", \"default\": 24},\n          \"preset\": {\"type\": \"string\", \"default\": \"veryfast\"},\n          \"output_path\": {\"type\": \"string\", \"description\": \"默认 output/videos\"},\n          \"overwrite\": {\"type\": \"boolean\", \"default\": true}\n        }\n      }\n    },\n    {\n      \"name\": \"thumbnail\",\n      \"description\": \"从 MP4/MOV 视频任意时间点或帧号提取封面，流式只下载必要部分\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"anyOf\": [\n          {\"required\": [\"video_url\"]},\n          {\"required\": [\"url\"]},\n          {\"required\": [\"source\"]}\n        ],\n        \"properties\": {\n          \"video_url\": {\"type\": \"string\", \"description\": \"视频 URL 或本地文件路径\"},\n          \"url\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"source\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"media_roots\": {\"type\": [\"array\", \"string\"], \"description\": \"允许访问的本地媒体根目录\"},\n          \"time_seconds\": {\"type\": \"number\"},\n          \"frame_number\": {\"type\": \"integer\"},\n          \"save_path\": {\"type\": \"string\"},\n          \"resize_width\": {\"type\": \"integer\"},\n          \"quality\": {\"type\": \"integer\", \"default\": 85},\n          \"overwrite\": {\"type\": \"boolean\", \"default\": true}\n        }\n      }\n    },\n    {\n      \"name\": \"audio\",\n      \"description\": \"流式提取音频，支持 MP3/WAV/AAC/M4A 格式\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"anyOf\": [\n          {\"required\": [\"video_url\"]},\n          {\"required\": [\"url\"]},\n          {\"required\": [\"source\"]}\n        ],\n        \"properties\": {\n          \"video_url\": {\"type\": \"string\", \"description\": \"视频 URL 或本地文件路径\"},\n          \"url\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"source\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"media_roots\": {\"type\": [\"array\", \"string\"], \"description\": \"允许访问的本地媒体根目录\"},\n          \"format\": {\"type\": \"string\", \"enum\": [\"mp3\", \"wav\", \"aac\", \"m4a\"], \"default\": \"mp3\"},\n          \"bitrate\": {\"type\": \"string\", \"default\": \"128k\", \"description\": \"比特率: 128k, 192k, 320k\"},\n          \"sample_rate\": {\"type\": \"integer\", \"default\": 44100, \"description\": \"采样率: 44100, 48000\"},\n          \"channels\": {\"type\": \"integer\", \"default\": 2, \"description\": \"声道: 1=单声道, 2=立体声\"},\n          \"start_time\": {\"type\": \"number\", \"description\": \"开始时间（秒）\"},\n          \"duration\": {\"type\": \"number\", \"description\": \"持续时间（秒）\"},\n          \"output_path\": {\"type\": \"string\", \"description\": \"默认 output/audio\"},\n          \"overwrite\": {\"type\": \"boolean\", \"default\": true}\n        }\n      }\n    },\n    {\n      \"name\": \"audio_batch\",\n      \"description\": \"批量提取多个视频的音频\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"required\": [\"videos\"],\n        \"properties\": {\n          \"videos\": {\"type\": \"array\", \"description\": \"视频列表，条目支持 video_url/url/source\"},\n          \"media_roots\": {\"type\": [\"array\", \"string\"], \"description\": \"允许访问的本地媒体根目录\"},\n          \"output_dir\": {\"type\": \"string\", \"default\": \"output/audio\"},\n          \"format\": {\"type\": \"string\", \"default\": \"mp3\"},\n          \"bitrate\": {\"type\": \"string\", \"default\": \"128k\"},\n          \"sample_rate\": {\"type\": \"integer\", \"default\": 44100},\n          \"overwrite\": {\"type\": \"boolean\", \"default\": true}\n        }\n      }\n    },\n    {\n      \"name\": \"audio_info\",\n      \"description\": \"获取视频的音频流信息\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"anyOf\": [\n          {\"required\": [\"video_url\"]},\n          {\"required\": [\"url\"]},\n          {\"required\": [\"source\"]}\n        ],\n        \"properties\": {\n          \"video_url\": {\"type\": \"string\", \"description\": \"视频 URL 或本地文件路径\"},\n          \"url\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"source\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"media_roots\": {\"type\": [\"array\", \"string\"], \"description\": \"允许访问的本地媒体根目录\"}\n        }\n      }\n    },\n    {\n      \"name\": \"asr\",\n      \"description\": \"从音频轨识别字幕，输出 SRT-like captions JSON\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"anyOf\": [\n          {\"required\": [\"video_url\"]},\n          {\"required\": [\"url\"]},\n          {\"required\": [\"source\"]}\n        ],\n        \"properties\": {\n          \"video_url\": {\"type\": \"string\", \"description\": \"视频 URL 或本地文件路径\"},\n          \"url\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"source\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"language\": {\"type\": \"string\", \"default\": \"auto\"},\n          \"model_size\": {\"type\": \"string\", \"default\": \"base\"},\n          \"output_path\": {\"type\": \"string\", \"description\": \"默认 output/subtitles/<name>.captions.json\"},\n          \"overwrite\": {\"type\": \"boolean\", \"default\": true},\n          \"media_roots\": {\"type\": [\"array\", \"string\"], \"description\": \"允许访问的本地媒体根目录\"}\n        }\n      }\n    },\n    {\n      \"name\": \"ocr\",\n      \"description\": \"从视频画面识别硬字幕或屏幕文字，输出 SRT-like captions JSON\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"anyOf\": [\n          {\"required\": [\"video_url\"]},\n          {\"required\": [\"url\"]},\n          {\"required\": [\"source\"]}\n        ],\n        \"properties\": {\n          \"video_url\": {\"type\": \"string\", \"description\": \"视频 URL 或本地文件路径\"},\n          \"url\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"source\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"sample_interval_sec\": {\"type\": \"number\", \"default\": 1.0},\n          \"crop_bottom_ratio\": {\"type\": \"number\", \"default\": 0.35},\n          \"output_path\": {\"type\": \"string\", \"description\": \"默认 output/subtitles/<name>.captions.json\"},\n          \"overwrite\": {\"type\": \"boolean\", \"default\": true},\n          \"media_roots\": {\"type\": [\"array\", \"string\"], \"description\": \"允许访问的本地媒体根目录\"}\n        }\n      }\n    },\n    {\n      \"name\": \"subtitle\",\n      \"description\": \"推荐字幕识别入口，支持 asr / ocr / fusion 模式，统一输出 captions JSON\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"anyOf\": [\n          {\"required\": [\"video_url\"]},\n          {\"required\": [\"url\"]},\n          {\"required\": [\"source\"]}\n        ],\n        \"properties\": {\n          \"video_url\": {\"type\": \"string\", \"description\": \"视频 URL 或本地文件路径\"},\n          \"url\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"source\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"mode\": {\"type\": \"string\", \"enum\": [\"asr\", \"ocr\", \"fusion\"], \"default\": \"fusion\"},\n          \"language\": {\"type\": \"string\", \"default\": \"auto\"},\n          \"model_size\": {\"type\": \"string\", \"default\": \"base\"},\n          \"sample_interval_sec\": {\"type\": \"number\", \"default\": 1.0},\n          \"crop_bottom_ratio\": {\"type\": \"number\", \"default\": 0.35},\n          \"output_path\": {\"type\": \"string\", \"description\": \"默认 output/subtitles/<name>.captions.json\"},\n          \"overwrite\": {\"type\": \"boolean\", \"default\": true},\n          \"media_roots\": {\"type\": [\"array\", \"string\"], \"description\": \"允许访问的本地媒体根目录\"}\n        }\n      },\n      \"examples\": [\n        {\"source\": \"sample.mp4\", \"mode\": \"fusion\"},\n        {\"source\": \"sample.mp4\", \"mode\": \"asr\", \"language\": \"zh\"}\n      ]\n    },\n    {\n      \"name\": \"batch\",\n      \"description\": \"批量处理多个视频，支持 compress / thumbnail / audio\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"required\": [\"videos\"],\n        \"properties\": {\n          \"action\": {\"type\": \"string\", \"enum\": [\"compress\", \"thumbnail\", \"audio\"], \"default\": \"thumbnail\"},\n          \"videos\": {\"type\": \"array\", \"description\": \"按 action 传入对应参数对象列表，条目支持 video_url/url/source\"},\n          \"media_roots\": {\"type\": [\"array\", \"string\"], \"description\": \"允许访问的本地媒体根目录\"},\n          \"output_dir\": {\"type\": \"string\", \"description\": \"批量输出目录\"},\n          \"overwrite\": {\"type\": \"boolean\", \"default\": true}\n        }\n      }\n    },\n    {\n      \"name\": \"pipeline\",\n      \"description\": \"JSON 驱动媒体流水线，按 steps 顺序编排 info / thumbnail / audio / compress / audio_info / subtitle\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"anyOf\": [\n          {\"required\": [\"video_url\"]},\n          {\"required\": [\"url\"]},\n          {\"required\": [\"source\"]}\n        ],\n        \"required\": [\"steps\"],\n        \"properties\": {\n          \"video_url\": {\"type\": \"string\", \"description\": \"视频 URL 或本地文件路径\"},\n          \"url\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"source\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"media_roots\": {\"type\": [\"array\", \"string\"], \"description\": \"允许访问的本地媒体根目录\"},\n          \"name\": {\"type\": \"string\", \"description\": \"流水线名称\"},\n          \"output_dir\": {\"type\": \"string\", \"description\": \"默认 output/pipeline/<name>\"},\n          \"overwrite\": {\"type\": \"boolean\", \"default\": true},\n          \"steps\": {\n            \"type\": \"array\",\n            \"description\": \"步骤列表，每步包含 id/action/enabled/params\",\n            \"items\": {\n              \"type\": \"object\",\n              \"required\": [\"id\", \"action\", \"enabled\"],\n              \"properties\": {\n                \"id\": {\"type\": \"string\"},\n                \"action\": {\"type\": \"string\", \"enum\": [\"info\", \"thumbnail\", \"audio\", \"compress\", \"audio_info\", \"asr\", \"ocr\", \"subtitle\"]},\n                \"enabled\": {\"type\": \"boolean\"},\n                \"params\": {\"type\": \"object\"}\n              }\n            }\n          }\n        }\n      }\n    },\n    {\n      \"name\": \"info\",\n      \"description\": \"获取完整视频信息（分辨率、时长、编码等）\",\n      \"input_schema\": {\n        \"type\": \"object\",\n        \"anyOf\": [\n          {\"required\": [\"video_url\"]},\n          {\"required\": [\"url\"]},\n          {\"required\": [\"source\"]}\n        ],\n        \"properties\": {\n          \"video_url\": {\"type\": \"string\", \"description\": \"视频 URL 或本地文件路径\"},\n          \"url\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"source\": {\"type\": \"string\", \"description\": \"同 video_url\"},\n          \"media_roots\": {\"type\": [\"array\", \"string\"], \"description\": \"允许访问的本地媒体根目录\"}\n        }\n      }\n    }\n  ]\n}\n\nFile v4.0.1:requirements.txt\n\nrequests>=2.28.0\nopencv-python>=4.8.0\nnumpy>=1.24.0\naiohttp>=3.8.0\nflask>=2.3.0\nflask-cors>=4.0.0\nfaster-whisper>=1.0.0\npaddlepaddle>=2.6.0\npaddleocr>=2.7.0\n\nArchive v4.0.0: 13 files, 34412 bytes\n\nFiles: audio_extractor.py (9638b), frame_extractor.py (22227b), intent_parser.py (5987b), MAINTENANCE.md (3774b), requirements.txt (98b), run.py (28944b), scripts/smoke_test.py (5471b), skill.json (9557b), SKILL.md (7643b), tests/test_release_behaviors.py (14111b), utils.py (8193b), video_compressor.py (6897b), _meta.json (134b)\n\nFile v4.0.0:SKILL.md\n\n---\nname: ym-mediatoolkit\nversion: 4.0.0\ndescription: 自然语言媒体助手 - 视频压缩、MP4/MOV 封面提取、音频转换\nauthor: your_name\ntags:\n  - video\n  - compression\n  - thumbnail\n  - audio\n  - streaming\n  - ffmpeg\ncategories:\n  - media\n  - utility\nclawhub:\n  entrypoint: python run.py\n  runtime: python3\n  http_port: 8080\n---\n\n# YM MediaToolkit\n\n自然语言媒体助手，支持从远程视频 URL、当前工作目录内的本地视频文件、或配置过 `media_roots` 的本地媒体目录直接处理：\n\n- 自然语言调用\n- 视频压缩\n- MP4/MOV 封面提取\n- 音频提取与转换：MP3 / WAV / AAC / M4A\n- 批量处理\n- JSON 驱动媒体流水线\n\n## 依赖\n\n```bash\npip install -r requirements.txt\n```\n\n系统需要安装：\n\n```bash\nffmpeg\nffprobe\n```\n\n## 维护\n\n维护、发布、测试和排障流程见 [MAINTENANCE.md](./MAINTENANCE.md)。\n\n## 命令行\n\n```bash\npython run.py -a <action> -i '<json>'\n```\n\n也可以从 JSON 文件读取参数：\n\n```bash\npython run.py -i params.json\n```\n\n## HTTP 服务\n\n```bash\npython run.py --serve\n```\n\n默认监听 `127.0.0.1:8080`。\n\n健康检查：\n\n```http\nGET /health\n```\n\n## 功能\n\n输入源字段可使用 `video_url`、`url` 或 `source`，三者等价。远程输入仅支持 `http/https`，本地输入默认限制在当前工作目录内；需要访问绝对路径时，通过请求参数 `media_roots` 或环境变量 `YM_MEDIA_ROOTS` 配置允许的媒体根目录。\n\n默认输出目录：\n\n| 类型 | 默认目录 |\n|------|----------|\n| 压缩视频 | `output/videos` |\n| 音频 | `output/audio` |\n| 封面 | `output/thumbs` |\n\n所有会写文件的接口都支持 `overwrite`，默认 `true`。设置为 `false` 时，如果输出文件已存在会直接返回错误。\n\n## 3.0.3 更新\n\n- 支持当前工作目录内的本地视频文件输入。\n- 统一 `video_url` / `url` / `source` 三种输入字段。\n- 增加默认输出目录：`output/videos`、`output/audio`、`output/thumbs`。\n- 增加 `overwrite` 覆盖策略，避免误覆盖已有文件。\n\n## 4.0.0 更新\n\n- 新增自然语言入口 `chat`，适合 Claw 直接转发用户聊天文本。\n- 新增 `media_roots` 白名单，支持处理授权目录内的绝对路径文件。\n- 自然语言命令支持提取音频、提取封面、压缩、查看信息、JSON 流水线。\n- `chat` 返回 `reply` 和结构化 `result`，同时兼顾聊天展示和自动化消费。\n\n### 自然语言调用\n\nAction: `chat`\n\nClaw 推荐优先调用 `chat`。复杂、确定性要求高的多步骤流程继续使用 `pipeline`。\n\n```bash\npython run.py -a chat -i '{\"message\":\"将 \\\"sample.mp4\\\" 提取音频\"}'\npython run.py -a chat -i '{\"message\":\"给 \\\"sample.mp4\\\" 提取第 3 秒封面\"}'\npython run.py -a chat -i '{\"message\":\"压缩 \\\"sample.mp4\\\"\"}'\npython run.py -a chat -i '{\"message\":\"查看 \\\"sample.mp4\\\" 信息\"}'\n```\n\n处理绝对路径时需要配置媒体根目录：\n\n```bash\npython run.py -a chat -i '{\"message\":\"将 \\\"D:/AA.MP4\\\" 提取音频\",\"media_roots\":[\"D:/\"]}'\n```\n\n返回包含：\n\n| 字段 | 说明 |\n|------|------|\n| `reply` | 可直接展示给用户的聊天回复 |\n| `intent` | 识别出的意图 |\n| `action` | 实际调用的 action |\n| `params` | 传给底层 action 的参数 |\n| `result` | 底层 action 原始结果 |\n| `output_paths` | 本次生成的输出路径列表 |\n\n### 压缩视频\n\nAction: `compress`\n\n```bash\npython run.py -a compress -i '{\"source\":\"sample.mp4\",\"target_ratio\":0.1}'\n```\n\n参数：\n\n| 参数 | 类型 | 默认值 | 说明 |\n|------|------|--------|------|\n| `video_url` / `url` / `source` | string | 必填 | 远程 URL 或本地文件 |\n| `target_ratio` | number | 0.1 | 目标体积比例 |\n| `adaptive` | boolean | true | 是否自动尝试不同 CRF |\n| `crf` | integer | 24 | 非 adaptive 模式下使用 |\n| `preset` | string | veryfast | ffmpeg 编码预设 |\n| `output_path` | string | `output/videos/...` | 输出路径 |\n| `overwrite` | boolean | true | 是否覆盖已有文件 |\n\n### 提取封面\n\nAction: `thumbnail`\n\n当前支持 MP4/MOV 容器，主要适用于 H.264/H.265 视频轨道。\n\n```bash\npython run.py -a thumbnail -i '{\"source\":\"sample.mp4\",\"time_seconds\":5}'\n```\n\n参数：\n\n| 参数 | 类型 | 默认值 | 说明 |\n|------|------|--------|------|\n| `video_url` / `url` / `source` | string | 必填 | 远程 URL 或本地文件 |\n| `time_seconds` | number | 0 | 按时间点提取 |\n| `frame_number` | integer | - | 按帧号提取，优先于 `time_seconds` |\n| `save_path` | string | `output/thumbs/...` | 保存路径 |\n| `resize_width` | integer | - | 输出宽度 |\n| `quality` | integer | 85 | JPEG 质量 |\n| `overwrite` | boolean | true | 是否覆盖已有文件 |\n\n### 提取音频\n\nAction: `audio`\n\n```bash\npython run.py -a audio -i '{\"source\":\"sample.mp4\",\"format\":\"mp3\"}'\n```\n\n参数：\n\n| 参数 | 类型 | 默认值 | 说明 |\n|------|------|--------|------|\n| `video_url` / `url` / `source` | string | 必填 | 远程 URL 或本地文件 |\n| `format` | string | mp3 | mp3 / wav / aac / m4a |\n| `bitrate` | string | 128k | 音频比特率 |\n| `sample_rate` | integer | 44100 | 采样率 |\n| `channels` | integer | 2 | 声道数 |\n| `start_time` | number | - | 开始时间，秒 |\n| `duration` | number | - | 截取时长，秒 |\n| `output_path` | string | `output/audio/...` | 输出路径 |\n| `overwrite` | boolean | true | 是否覆盖已有文件 |\n\n### 批量音频\n\nAction: `audio_batch`\n\n```bash\npython run.py -a audio_batch -i '{\n  \"videos\":[\n    {\"source\":\"sample1.mp4\",\"name\":\"video1\"},\n    {\"url\":\"https://example.com/2.mp4\",\"name\":\"video2\"}\n  ],\n  \"output_dir\":\"output/audio\",\n  \"format\":\"mp3\"\n}'\n```\n\n### 批量处理\n\nAction: `batch`\n\n```bash\npython run.py -a batch -i '{\n  \"action\":\"thumbnail\",\n  \"videos\":[\n    {\"source\":\"sample1.mp4\",\"time_seconds\":5},\n    {\"url\":\"https://example.com/2.mp4\",\"time_seconds\":10}\n  ]\n}'\n```\n\n`action` 支持：`compress`、`thumbnail`、`audio`。\n\n### JSON 流水线\n\nAction: `pipeline`\n\n`steps` 是唯一流程控制入口，没有写进 `steps` 的动作不会执行。支持的 step action：`info`、`thumbnail`、`audio`、`comp\n\nArchive v3.0.3: 9 files, 21406 bytes\n\nFiles: audio_extractor.py (9489b), frame_extractor.py (22180b), requirements.txt (98b), run.py (17721b), skill.json (6547b), SKILL.md (4383b), utils.py (7175b), video_compressor.py (6766b), _meta.json (134b)\n\nArchive v3.0.2: 9 files, 19648 bytes\n\nFiles: audio_extractor.py (8884b), frame_extractor.py (20141b), requirements.txt (98b), run.py (13604b), skill.json (4936b), SKILL.md (3713b), utils.py (5326b), video_compressor.py (6097b), _meta.json (134b)\n\nArchive v3.0.1: 9 files, 21984 bytes\n\nFiles: audio_extractor.py (8694b), frame_extractor.py (19941b), requirements.txt (66b), run.py (12066b), skill.json (4987b), SKILL.md (11530b), utils.py (4684b), video_compressor.py (5952b), _meta.json (134b)\n\nArchive v3.0.0: 9 files, 21803 bytes\n\nFiles: audio_extractor.py (8708b), frame_extractor.py (19941b), requirements.txt (66b), run.py (12066b), skill.json (4435b), SKILL.md (11530b), utils.py (4664b), video_compressor.py (5501b), _meta.json (134b)\n\nArchive v2.0.2: 9 files, 21802 bytes\n\nFiles: audio_extractor.py (8708b), frame_extractor.py (19941b), requirements.txt (66b), run.py (12066b), skill.json (4435b), SKILL.md (11534b), utils.py (4664b), video_compressor.py (5501b), _meta.json (134b)","readmeExcerpt":"Skill: YM-MediaToolkit(媒体处理工具集) Owner: 370299455cx-web Summary: 自然语言媒体助手 - 视频压缩、MP4/MOV 封面提取、音频转换、字幕识别 Tags: latest:4.2.2 Version history: v4.2.2 | 2026-06-25T08:57:44.696Z | user 4.3.1 更新 - 修正 HTTP chat 短等待：job 在 wait_timeout_sec 内完成时返回 200 和最终结果，否则返回 202。 - /skill/jobs 接口的 params 仅允许为 JSON object，缺省/null 时自动为 {}，其他类型返回 invalid_params。 - async 仅支持 JSON boolean 或 \"auto\" 字符串，wait_timeout_sec 必须为 0–30 秒数字。 - 异步返回的 outp","codeSnippets":[],"executableExamples":[{"language":"bash","snippet":"pip install -r requirements.txt"},{"language":"bash","snippet":"ffmpeg\nffprobe"},{"language":"bash","snippet":"python run.py -a <action> -i '<json>'"},{"language":"bash","snippet":"python run.py -i params.json"},{"language":"bash","snippet":"python run.py --serve"},{"language":"http","snippet":"GET /health"}],"parameters":null,"dependencies":[],"permissions":[],"extractedFiles":[{"path":"SKILL.md","content":"---\nname: ym-mediatoolkit\nversion: 4.3.1\ndescription: 自然语言媒体助手 - 视频压缩、MP4/MOV 封面提取、音频转换、字幕识别\nauthor: your_name\ntags:\n  - video\n  - compression\n  - thumbnail\n  - audio\n  - streaming\n  - ffmpeg\ncategories:\n  - media\n  - utility\nclawhub:\n  entrypoint: python run.py\n  runtime: python3\n  http_port: 8080\n---\n\n# YM MediaToolkit\n\n自然语言媒体助手，支持从远程视频 URL、当前工作目录内的本地视频文件、或配置过 `media_roots` 的本地媒体目录直接处理：\n\n- 自然语言调用\n- 视频压缩\n- MP4/MOV 封面提取\n- 音频提取与转换：MP3 / WAV / AAC / M4A\n- OCR / ASR 字幕识别\n- emlet 字幕二次分句\n- 批量处理\n- JSON 驱动媒体流水线\n\n## 依赖\n\n```bash\npip install -r requirements.txt\n```\n\n系统需要安装：\n\n```bash\nffmpeg\nffprobe\n```\n\n字幕识别依赖已内置在 `requirements.txt`，包括 `faster-whisper`、`paddlepaddle`、`paddleocr`。\n\n## 维护\n\n维护、发布、测试和排障流程见 [MAINTENANCE.md](./MAINTENANCE.md)。\n\n## 命令行\n\n```bash\npython run.py -a <action> -i '<json>'\n```\n\n也可以从 JSON 文件读取参数：\n\n```bash\npython run.py -i params.json\n```\n\n## HTTP 服务\n\n```bash\npython run.py --serve\n```\n\n默认监听 `127.0.0.1:8080`。\n\n健康检查：\n\n```http\nGET /health\n```\n\n异步长任务：\n\n```http\nPOST /skill/jobs\nGET /skill/jobs/<job_id>\nGET /skill/jobs\n```\n\n## 功能\n\n输入源字段可使用 `video_url`、`url` 或 `source`，三者等价。远程输入仅支持 `http/https`，本地输入默认限制在当前工作目录内；需要访问绝对路径时，通过请求参数 `media_roots` 或环境变量 `YM_MEDIA_ROOTS` 配置允许的媒体根目录。\n\n## 返回协议\n\n从 `4.1.0` 开始，所有 action 都会返回稳定协议字段：\n\n| 字段 | 说明 |\n|------|------|\n| `status` | `success` / `partial` / `skipped` / `error` |\n| `code` | 稳定机器码，例如 `ok`、`missing_source`、`output_exists`、`parse_failed` |\n| `reply` | 适合聊天展示的简短中文回复 |\n| `hint` | 面向调用方或用户的下一步建议 |\n\n原有业务字段会继续保留，例如 `output_path`、`saved_path`、`outputPath`、`manifest_path`、`info`、`captions`、`result`。\n\n默认输出目录：\n\n| 类型 | 默认目录 |\n|------|----------|\n| 压缩视频 | `output/videos` |\n| 音频 | `output/audio` |\n| 封面 | `output/thumbs` |\n\n所有会写文件的接口都支持 `overwrite`，默认 `true`。设置为 `false` 时，如果输出文件已存在会直接返回错误。\n\n## 3.0.3 更新\n\n- 支持当前工作目录内的本地视频文件输入。\n- 统一 `video_url` / `url` / `source` 三种输入字段。\n- 增加默认输出目录：`output/videos`、`output/audio`、`output/thumbs`。\n- 增加 `overwrite` 覆盖策略，避免误覆盖已有文件。\n\n## 4.0.0 更新\n\n- 新增自然语言入口 `chat`，适合 Claw 直接转发用户聊天文本。\n- 新增 `media_roots` 白名单，支持处理授权目录内的绝对路径文件。\n- 自然语言命令支持提取音频、提取封面、压缩、查看信息、JSON 流水线。\n- `chat` 返回 `reply` 和结构化 `result`，同时兼顾聊天展示和自动化消费。\n\n## 4.0.1 更新\n\n- 新增 `subtitle` 推荐入口，支持 `asr` / `ocr` / `fusion` 模式。\n- 新增 `asr` 和 `ocr` 单独调试入口。\n- 字幕统一输出 SRT-like JSON：`captionTxt`、`startTimeUs`、`endTimeUs`、`source`、`confidence`。\n- `chat` 支持“识别字幕 / 提取字幕 / 转字幕 / 生成字幕”等自然语言命令。\n\n## 4.0.2 更新\n\n- 将 `faster-whisper`、`paddlepaddle`、`paddleocr` 纳入默认 `requirements.txt`。\n- 字幕识别从“可选依赖”调整为默认安装能力。\n- 运行时仍保留缺依赖 JSON error，方便定位未重新安装依赖的环境。\n\n## 4.1.0 更新\n\n- 所有 action 统一补齐 `code`、`reply`、`hint`，方便 Claw 和后续渠道适配。\n- 增加稳定错误码：`missing_source`、`source_not_allowed`、`output_exists`、`parse_failed`、`missing_steps`、`unsupported_action`、`ffmpeg_failed`、`missing_asr_dependency`、`missing_ocr_dependency`。\n- 保持旧字段兼容，不改变现有 action 名称、HTTP endpoint 和底层媒体处理逻辑。\n\n## 4.1.1 更新\n\n- 新增 `caption_segment`，用于 emlet 字幕二次分句。\n- 默认每句最多 `12` 个字符，按强标点、弱标点、连接词和长度切分。\n- 新增 `protected_terms` 和 `protected_terms_path`，用于保护品牌词、产品名、人名、术语不被拆开。\n- `pipeline` 支持 `subtitle -> caption_segment` 串联；分句步骤未传 `cap"},{"path":"_meta.json","content":"{\n  \"ownerId\": \"kn787sam7qk7fffsjc875yxp0984rpzh\",\n  \"slug\": \"ym-mediatoolkit\",\n  \"version\": \"4.2.2\",\n  \"publishedAt\": 1782377864696\n}"},{"path":"MAINTENANCE.md","content":"# YM MediaToolkit 维护手册\n\n本文档面向后续维护者，记录版本发布、测试验证、配置边界和排障流程。\n\n## 发布流程\n\n1. 同步版本号：\n   - `SKILL.md` front matter 的 `version`\n   - `skill.json` 的 `version`\n2. 更新用户文档：\n   - 新增 action 时同步更新 `SKILL.md` 和 `skill.json`\n   - 新增参数时同步补充默认值、用途和安全限制\n   - 变更行为时在 `SKILL.md` 的版本更新段落记录\n3. 运行验证：\n\n```bash\npython3 -B -m py_compile run.py utils.py intent_parser.py audio_extractor.py frame_extractor.py video_compressor.py asr_engine.py ocr_engine.py subtitle_extractor.py caption_segmenter.py job_manager.py tests/test_release_behaviors.py scripts/smoke_test.py\npython3 -m json.tool skill.json\npython3 -B -m unittest discover -s tests\npython3 scripts/smoke_test.py\n```\n\n4. 清理产物：\n\n```bash\nfind . -type d -name __pycache__ -prune -exec rm -r {} +\nfind . -name .DS_Store -delete\n```\n\n## 当前接口分层\n\n- `chat`：Claw 推荐入口，接收自然语言，解析后调用现有 action。\n- `pipeline`：确定性 JSON 流水线入口，适合多步骤、可重复的自动化流程。\n- `audio` / `thumbnail` / `compress` / `info` / `audio_info`：底层单步能力。\n- `subtitle` / `asr` / `ocr`：字幕识别能力，输出 SRT-like captions JSON。\n- `caption_segment`：emlet 字幕二次分句器，处理已有 captions，不重新识别媒体。\n- `batch` / `audio_batch`：批量处理入口。\n- `/skill/jobs`：HTTP 异步长任务入口，适合压缩、字幕识别、pipeline 等耗时 action。\n- `/skill/chat` + `async:\"auto\"`：Claw 推荐长任务入口，先解析自然语言，再自动提交 job。\n\n维护原则：新体验优先接到 `chat` 或 `pipeline`，媒体处理逻辑继续复用底层 action，避免重复实现 ffmpeg 调用。\n\n## Action 返回协议\n\n所有 action 必须返回 JSON 对象，并保留以下协议字段：\n\n- `status`：`success` / `partial` / `skipped` / `error`\n- `code`：稳定机器码，成功为 `ok`\n- `reply`：适合聊天展示的中文回复\n- `hint`：下一步建议或排障提示\n\n新增 handler 时只需要返回原始业务结果，`run.py` 的 action protocol wrapper 会补齐缺失字段。若底层模块已经返回 `code`，包装层会保留该值。\n\n常用错误码：\n\n| code | 场景 |\n|------|------|\n| `missing_source` | 缺少 `video_url` / `url` / `source` |\n| `source_not_allowed` | 本地路径不在当前工作目录或 `media_roots` 内 |\n| `output_exists` | 输出文件已存在且 `overwrite=false` |\n| `parse_failed` | `chat` 无法识别自然语言意图 |\n| `missing_steps` | `pipeline` 未传 `steps` |\n| `missing_captions` | `caption_segment` 未传 `captions` 或 `caption_path` |\n| `invalid_action` | 异步任务或 CLI 传入不支持的 action |\n| `invalid_params` | HTTP job 的 `params` 不是 JSON object |\n| `invalid_async_mode` | HTTP chat 的 `async` 不是 boolean 或 `\"auto\"` |\n| `invalid_wait_timeout` | HTTP chat 的 `wait_timeout_sec` 不是 `0-30` 秒数字 |\n| `invalid_job_id` | job id 格式不合法 |\n| `job_not_found` | job 文件不存在 |\n| `job_interrupted` | 服务重启或进程中断导致未完成 job 失效 |\n| `unsupported_action` | action 或 pipeline step 不支持 |\n| `invalid_step` | pipeline step 结构不合法 |\n| `ffmpeg_failed` | ffmpeg / ffprobe 缺失或执行失败 |\n| `missing_asr_dependency` | ASR 依赖缺失 |\n| `missing_ocr_dependency` | OCR 依赖缺失 |\n\n## HTTP 异步任务维护\n\n异步任务逻辑在 `job_manager.py`。HTTP 服务启动时会创建单 worker 队列，串行执行提交到 `/skill/jobs` 的 action。\n\n维护规则：\n\n- job 文件存储在 `output/jobs/<job_id>/job.json`。\n- job id 使用 32 位十六进制 UUID。\n- job 文件必须保留 `job_id`、`action`、`params`、`status`、`code`、`reply`、`hint`、`created_at`、`started_at`、`finished_at`、`result`、`output_paths`、`error`。\n- `chat` 提交的 job 会额外写入 `created_by=chat`、`intent`、`source`、`metadata.message`。\n- `/skill/jobs` 直接提交的 job 会写入 `created_by=jobs`。\n- 当前版本不支持取消任务，不记录进度百分比。\n- 服务重启后不恢复未完成任务；旧的 `queued` / `runnin"},{"path":"skill-card.md","content":"## Description:\n\n自然语言媒体助手 - 视频压缩、MP4/MOV 封面提取、音频转换、字幕识别。\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[370299455cx-web](https://clawhub.ai/user/370299455cx-web)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers and external users use this skill to turn natural-language or JSON requests into media-processing tasks such as video compression, thumbnail extraction, audio conversion, subtitle recognition, batch processing, and asynchronous job polling.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: The unauthenticated local HTTP API can process local files and remote URLs.\n\nMitigation: Run it only in a trusted local environment and do not expose the HTTP server to a network.\n\nRisk: Changing the bind host to 0.0.0.0 can make the service reachable beyond localhost.\n\nMitigation: Use authentication, restricted CORS, and a controlled reverse proxy before any non-local deployment.\n\nRisk: Accepting media_roots from untrusted callers can expand local file access.\n\nMitigation: Do not accept media_roots from untrusted callers and restrict configured media roots to approved directories.\n\nRisk: Media processing and model dependencies can consume significant compute, memory, disk, and network resources.\n\nMitigation: Use pinned dependencies, trusted media sources, quota controls, and resource limits before production use.\n\n## Reference(s):\n\n- [ClawHub skill page](https://clawhub.ai/370299455cx-web/skills/ym-mediatoolkit)\n- [Publisher profile](https://clawhub.ai/user/370299455cx-web)\n\n## Skill Output:\n\n**Output Type(s):** [text, markdown, code, shell commands, configuration, guidance, files]\n\n**Output Format:** [Markdown guidance, shell commands, JSON responses, and generated media or caption files]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [Outputs can include stable status/code/reply/hint fields, media file paths, caption JSON, pipeline manifests, and asynchronous job metadata.]\n\n## Skill Version(s):\n\n4.2.2 (source: server release metadata)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment."},{"path":"output/jobs/c65f1520dca6436c87f1028354263718/job.json","content":"{\n  \"job_id\": \"c65f1520dca6436c87f1028354263718\",\n  \"action\": \"info\",\n  \"params\": {\n    \"source\": \"sample.mp4\",\n    \"name\": \"sample\"\n  },\n  \"created_by\": \"chat\",\n  \"intent\": \"info\",\n  \"source\": \"sample.mp4\",\n  \"metadata\": {\n    \"created_by\": \"chat\",\n    \"intent\": \"info\",\n    \"source\": \"sample.mp4\",\n    \"message\": \"查看 \\\"sample.mp4\\\" 信息\"\n  },\n  \"status\": \"success\",\n  \"code\": \"ok\",\n  \"reply\": \"ok\",\n  \"hint\": \"ok\",\n  \"created_at\": \"2026-06-25T02:22:13.021739+00:00\",\n  \"started_at\": \"2026-06-25T02:22:13.022226+00:00\",\n  \"finished_at\": \"2026-06-25T02:22:13.022565+00:00\",\n  \"result\": {\n    \"status\": \"success\",\n    \"code\": \"ok\",\n    \"reply\": \"ok\",\n    \"hint\": \"ok\",\n    \"output_path\": \"out/a.json\",\n    \"nested\": {\n      \"outputPath\": \"out/a.json\"\n    }\n  },\n  \"output_paths\": [\n    \"out/a.json\"\n  ],\n  \"error\": null\n}"}],"languages":[],"docsSourceLabel":"CLAWHUB","editorialOverview":null,"editorialQuality":{"score":100,"threshold":65,"status":"thin","wordCount":1395,"uniquenessScore":43,"reasons":["uniqueness-below-45"]}},"media":{"evidence":{"source":"no-media","verified":false,"confidence":"low","updatedAt":"2026-10-10T11:42:42.807Z","emptyReason":"No screenshots, media assets, or demo links are available."},"primaryImageUrl":null,"mediaAssetCount":0,"assets":[],"demoUrl":null},"ownerResources":{"evidence":{"source":"unclaimed","verified":false,"confidence":"low","updatedAt":"2026-10-10T11:42:42.807Z","emptyReason":"This page has not been claimed by the agent owner."},"hasCustomPage":false,"customPageUpdatedAt":null,"customLinks":[],"structuredLinks":{"docsUrl":null,"demoUrl":null,"supportUrl":null,"pricingUrl":null,"statusUrl":null},"customPage":null},"relatedAgents":{"evidence":{"source":"protocol-neighbors","verified":false,"confidence":"medium","updatedAt":"2026-10-10T14:44:18.920Z","emptyReason":null},"items":[{"id":"8ebccd8e-3863-4187-8355-c3f14e1f9edf","entityType":"agent","canonicalPath":"/agent/iofficeai-aionui","slug":"iofficeai-aionui","name":"AionUi","description":"Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!","url":"https://github.com/iOfficeAI/AionUi","homepage":"https://www.aionui.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-10-09T19:11:12.944Z","createdAt":"2026-02-25T03:38:16.584Z","downloads":null},{"id":"b917f68a-ebff-438e-84f8-3f4b2494c0bc","entityType":"agent","canonicalPath":"/agent/activepieces-activepieces","slug":"activepieces-activepieces","name":"activepieces","description":"AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents","url":"https://github.com/activepieces/activepieces","homepage":"https://www.activepieces.com","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-15T02:22:12.426Z","createdAt":"2026-02-25T03:38:12.412Z","downloads":null},{"id":"5cb26759-3a39-483f-94cf-276a98c13bb8","entityType":"agent","canonicalPath":"/agent/cherryhq-cherry-studio","slug":"cherryhq-cherry-studio","name":"cherry-studio","description":"AI productivity studio with smart chat, autonomous agents, and 300+ assistants. Unified access to frontier LLMs","url":"https://github.com/CherryHQ/cherry-studio","homepage":"https://cherry-ai.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-11T14:38:40.986Z","createdAt":"2026-02-25T03:38:19.379Z","downloads":null},{"id":"6f6582d0-5d76-4f0f-b81d-86520247950b","entityType":"agent","canonicalPath":"/agent/copilotkit-copilotkit","slug":"copilotkit-copilotkit","name":"CopilotKit","description":"The Frontend for Agents & Generative UI. React + Angular","url":"https://github.com/CopilotKit/CopilotKit","homepage":"https://docs.copilotkit.ai","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-03-25T09:50:57.846Z","createdAt":"2026-02-25T03:39:14.617Z","downloads":null}],"links":{"hub":"/agent","source":"/agent/source/clawhub","protocols":[{"label":"OpenClaw","href":"/agent/protocol/openclew"}]}}}