agentCLAWHUBUnverified

Video Subtitle Extractor

Cross-platform video subtitle extraction using multi-engine ASR (speech-to-text). Downloads audio from video URLs via yt-dlp, transcribes with SenseVoice / w... Skill: Video Subtitle Extractor Owner: forhonourlx Summary: Cross-platform video subtitle extraction using multi-engine ASR (speech-to-text). Downloads audio from video URLs via yt-dlp, transcribes with SenseVoice / w... Tags: ai:1.0.0, asr:1.0.0, chinese:1.0.0, latest:2.0.0, major:2.0.0, subtitle:1.0.0, video:1.0.0, whisper:1.0.0 Version history: v2.0.0 | 2026-05-27T08:24:13.771Z | user Multi-engine ASR: SenseVoice

OpenClaw

Rank

62

Safety

84

Downloads

1.2k

Updated

Oct 10, 2026

Version

2.0.0

Source

CLAWHUB

About

What it does, and when to use it.

Capability contract not published. No trust telemetry is available yet. 1.2K downloads reported by the source. Last updated 10/10/2026.

Avoid when

  • Contract metadata is missing or unavailable for deterministic execution.

Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing

Public facts

Every fact links back to the source it came from.

Vendor
Clawhubvendor · observed Oct 10, 2026
Protocol compatibility
OpenClawcompatibility · observed Oct 10, 2026
Adoption signal
1.2K downloadsadoption · observed Oct 10, 2026
Latest release
2.0.0release · observed May 27, 2026
Handshake status
UNKNOWNsecurity

Install and run

Setup complexity: low.

clawhub skill install s171scmakmb14cmkc061v2rkvh875fc9:video-subtitle-extractor
  1. Setup complexity is classified as HIGH. You must provision dedicated cloud infrastructure or an isolated VM. Do not run this directly on your local workstation.
  2. Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.

Contract: missing

curl -s "https://www.xpersona.co/api/v1/agents/clawhub-forhonourlx-video-subtitle-extractor/snapshot"

Documentation

CLAWHUB

146,522 characters of source documentation, loaded on request.

Extracted files

5 files captured from the source.

SKILL.md

---
name: video-subtitle-extractor
description: |
  Cross-platform video subtitle extraction using multi-engine ASR (speech-to-text). Downloads audio from video URLs via yt-dlp, transcribes with SenseVoice / whisper.cpp / openai-whisper (default: SenseVoice Small for Chinese), and applies LLM-based text calibration for Chinese financial/technical content. Use when: (1) extracting subtitles from Bilibili, Xiaohongshu, YouTube, or any yt-dlp-supported platform, (2) the video has no built-in subtitles, (3) users say "下载字幕", "提取字幕", "语音转文字", "视频转文字", "字幕提取", "ASR转写", (4) needing to transcribe audio files to text, (5) working with Chinese-language video content requiring high-accuracy transcription. Automatically handles dependency installation (ffmpeg, yt-dlp, ASR backends) and model downloads.
---

# Video Subtitle Extractor 🎬→📝

Cross-platform multi-engine ASR subtitle extraction pipeline. Downloads audio from any yt-dlp-compatible video platform, transcribes with **SenseVoice / whisper.cpp / openai-whisper**, and applies LLM-based text calibration for Chinese content.

**Default engine**: SenseVoice Small (Alibaba FunASR) — ~1.5GB RAM, 234MB disk, 20× realtime speed, ~96% Chinese accuracy.

**Tested & verified** on Windows 11 with real Bilibili & Xiaohongshu videos.

## Quick Start

```bash
# One-command full pipeline (SenseVoice Small — default, blazing fast for Chinese)
python scripts/run.py <video_url> --output-dir ./output

# Use whisper.cpp GGML (even lighter, 2GB RAM)
python scripts/run.py <video_url> --backend whispercpp --model medium-q5_1

# Use openai-whisper (standard, 5GB RAM)
python scripts/run.py <video_url> --backend openai --model medium

# Download audio only
python scripts/download_audio.py <video_url> <output_dir>

# Download audio + video (keep both as middleware)
python scripts/download_audio.py <video_url> <output_dir> --save-video --video-quality 1080

# Transcribe existing audio with auto backend selection
python scripts/transcribe.py <audio_file> --backend auto --language zh

# Transcribe with specific backend
python scripts/transcribe.py <audio_file> --backend sensevoice --language zh
python scripts/transcribe.py <audio_file> --backend whispercpp --model medium-q5_1
python scripts/transcribe.py <audio_file> --backend openai --model medium
```

## When to Use This Skill

Use this skill when:
1. The video has **no built-in subtitles** (Bilibili, Xiaohongshu, YouTube, etc.)
2. You need **high-accuracy Chinese transcription** (~95% with medium model)
3. You want **multiple output formats** (TXT, SRT, VTT, JSON)
4. You need **LLM-assisted text calibration** for financial/technical terms
5. The user says: "下载字幕", "提取字幕", "语音转文字", "视频转文字", "字幕提取", "ASR转写"

## Workflow

### Step 0: Install Dependencies (once)

```bash
python scripts/install_deps.py
```

Auto-detects OS and installs: ffmpeg (winget/brew/apt), yt-dlp (pip), openai-whisper (pip). Handles Windows ffmpeg path

_meta.json

{
  "ownerId": "kn71y79mxxwmzynnpj6e284kch82gwd2",
  "slug": "video-subtitle-extractor",
  "version": "2.0.0",
  "publishedAt": 1779870253771
}

references/asr_models.md

# ASR Model Selection Guide

## Available Models

| Model | RAM | Disk | Speed | Quality | Best For |
|-------|-----|------|-------|---------|----------|
| `small` | ~2GB | 461MB | Fast (~1min/6min audio) | Good | Quick tests, low-resource systems |
| `medium` | ~5GB | 1.42GB | Medium (~3-5min) | High | **Recommended default** - best quality/speed ratio |
| `large-v3` | ~10GB | 2.88GB | Slow (~10-20min) | Best | Production quality, needs high RAM |
| `large-v3-turbo` | ~6GB | 1.6GB | Medium-fast | High | Good compromise, smaller than large-v3 |

## Language-Specific Notes

### Chinese (zh)
- `medium`: Good for general Chinese content. Some errors on homophones and financial terms.
- `large-v3`: Best Chinese accuracy, handles accents and domain terminology better.
- Common errors: 同音字混淆 (硬扛→硬钢), 金融术语 (抛压→抛押, 交筹→焦愁), K线术语 (十字星→14星)

### English (en)
- `small`: Sufficient for clear English speech.
- `medium`: Excellent accuracy for most content.

## Memory Constraints

On Windows, `large-v3` may be killed (SIGKILL) on systems with <16GB RAM due to FP32 fallback.
If killed, fall back to `medium` or use `larger-v3-turbo`.

## Future Model Compatibility

The `transcribe.py` script is designed for easy backend extension:
- `faster-whisper`: CTranslate2 backend, more memory efficient
- `whisper.cpp`: Native C++ implementation
- `mlx-whisper`: Apple Silicon optimized
- Cloud APIs: AssemblyAI, iFlytek, Whisper API

To add a new backend, implement a `transcribe_<backend>()` function in transcribe.py
following the same interface (audio_path, model_name, language, output_dir).

## Model Auto-Download

Models are downloaded automatically by openai-whisper on first use.
Cache location:
- Windows: `C:\Users\<user>\.cache\whisper\`
- macOS: `~/Library/Caches/whisper/`
- Linux: `~/.cache/whisper/`

references/calibration_guide.md

# Text Calibration Guide for Chinese ASR Output

Common transcription errors and their corrections. Apply these patterns when calibrating whisper output for Chinese financial/technical content.

## 1. Homophone Replacements (同音字混淆)

| Raw (Wrong) | Correct | Example |
|-------------|---------|---------|
| 硬钢 | 硬扛 | 硬钢→硬扛 |
| 抛押 | 抛压 | 消化抛押→消化抛压 |
| 模两个月 | 磨两个月 | 横盘调整模两个月→磨两个月 |
| 膜光短线 | 磨光短线 | 膜光短线→磨光短线 |
| 流通骨 | 流通股 | 流通骨的换手→流通股的换手 |
| 金接盘 | 新接盘 | 金接盘的成本→新接盘的成本 |
| 拉伸 | 拉升 | 拉伸成本→拉升成本 |
| 跟锋 | 跟风 | 跟锋买入→跟风买入 |
| 微转 | 微赚 | 微转就抛售→微赚就抛售 |
| 落带为安 | 落袋为安 | 落带为安→落袋为安 |
| 互盘 | 护盘 | 主力互盘明显→主力护盘明显 |
| 逼散互买 | 逼散户卖 | 洗盘本质是逼散互买→逼散户卖 |
| 仅 | 有 | 仅有资金拖住→有资金托住 |
| 军线 | 均线 | 关键军线→关键均线 |
| 快有动作 | 快有动作 | Already correct, but watch for 快会→快会 |

## 2. Financial Term Corrections

| Raw | Correct | Context |
|-----|---------|---------|
| 交筹 | 交筹 | 慢慢就焦愁→慢慢就交筹 |
| 再计 | 在即 | 拉升再计→拉升在即 |
| 没装 | 没仓 | 根本没装→根本没仓 |
| 利空 | 利空 | Already correct, verify |
| K线收14星 | K线收十字星 | 14→十 |
| 14星 | 十字星 | K线收十字星 |
| 洗崩 | 洗崩 | Already correct (跌太多) |
| 割肉 | 割肉 | Already correct |

## 3. Domain Term Patterns

Whisper often confuses financial jargon:
- 洗盘 (xǐ pán) vs 洗盘 (same pronunciation but context-dependent)
- 筹码 (chóu mǎ) - usually correct
- 建仓 (jiàn cāng) - usually correct
- 杠杆 (gàng gǎn) - usually correct
- 信托 (xìn tuō) - usually correct

## 4. Structural Cleanup

- Add proper punctuation (periods, commas) where ASR output lacks them
- Split long run-on sentences at natural topic breaks
- Format as flowing paragraphs, not timestamp-ordered fragments
- Add section headings for topic shifts: "洗盘核心目的", "三个核心指标", "三个信号", etc.
- Keep timestamps if user wants time-coded output (from .srt/.vtt)

## 5. Quality Indicators

After calibration, flag low-confidence sections:
- Unclear audio sections (background noise, overlapping speech)
- Rapid technical jargon sequences
- Sections where multiple interpretations are plausible

## 6. Multi-language Content

For bilingual content (Chinese + English):
- Preserve English terms: PE ratio, MA, MACD, KDJ, Bollinger Bands
- Mixed language phrases: "比如 PE 20倍", "MACD 金叉" are correct
- Don't translate technical terms to Chinese

## 7. Calibration Output Format

After applying corrections, present as:
- Clean prose with proper Chinese punctuation
- Optional: show what was changed vs raw output
- Optional: timestamp references from original SRT/VTT

## 8. AI / Tech Domain Corrections (AI 及科技领域)

For videos about AI, semiconductors, and tech investment topics, watch for:

### Chinese Company Names (whisper frequent errors)
| Raw (Wrong) | Correct | Context |
|-------------|---------|---------|
| 中繼續創 | 中际旭创 | A股光模块龙头 |
| 新益勝 | 新易盛 | A股光模块 |
| 天賦通信 | 天孚通信 | A股通信 |
| 元傑科技 | 源杰科技 | A股芯片 |
| 阿力 | 阿里 | 阿里巴巴 |
| Alley | 阿里 | Context: 腾讯、阿里、字节 |

### AI Product & Term Names
| Raw (Wrong) | Correct |
|-------------|---------|
| Deepseat | DeepSeek |
| ChadGPT | ChatGPT |
|

skill-card.md

## Description:

Cross-platform video subtitle extraction using multi-engine ASR: it downloads audio from video URLs with yt-dlp, transcribes with SenseVoice, whisper.cpp, or openai-whisper, and can calibrate Chinese financial and technical transcripts.

This skill is ready for commercial/non-commercial use.

## Publisher:

[forhonourlx](https://clawhub.ai/user/forhonourlx)

### License/Terms of Use:

MIT-0

## Use Case:

Developers, analysts, and content teams use this skill to extract subtitles or transcript files from videos that lack built-in captions, especially Chinese-language financial or technical videos. It guides agents through dependency setup, media download, ASR transcription, optional video retention, and transcript calibration.

### Deployment Geography for Use:

Global

## Known Risks and Mitigations:

Risk: The security review flags broad install, download, browser-cookie, and persistent storage behaviors.

Mitigation: Review the skill before use, run it only in a trusted environment, and approve dependency installation and media downloads deliberately.

Risk: Browser-cookie workflows can expose account session material to the local toolchain.

Mitigation: Avoid browser-cookie options unless the machine and toolchain are trusted; prefer videos that do not require cookies.

Risk: Downloaded media and generated transcripts can remain on disk after execution.

Mitigation: Use a dedicated output directory and delete media, transcripts, and metadata after they are no longer needed.

## Reference(s):

- [ASR Model Selection Guide](artifact/references/asr_models.md)
- [Text Calibration Guide for Chinese ASR Output](artifact/references/calibration_guide.md)
- [ClawHub Skill Page](https://clawhub.ai/forhonourlx/skills/video-subtitle-extractor)

## Skill Output:

**Output Type(s):** [text, markdown, code, shell commands, configuration, guidance]

**Output Format:** [Markdown guidance with shell commands and generated transcript files such as TXT, SRT, VTT, and JSON]

**Output Parameters:** [1D]

**Other Properties Related to Output:** [May create persistent media, transcript, subtitle, metadata, and calibrated transcript files in the selected output directory.]

## Skill Version(s):

2.0.0 (source: server release metadata)

## Ethical Considerations:

Users should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.
Github ReposUpdated 1d agoRank 70

AionUi

Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!

MCPOPENCLAW
Github ReposUpdated 6mo agoRank 70

activepieces

AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents

OPENCLAW
Github ReposUpdated 6mo agoRank 70

cherry-studio

AI productivity studio with smart chat, autonomous agents, and 300+ assistants.

MCPOPENCLAW
Github ReposUpdated 7mo agoRank 70

CopilotKit

The Frontend for Agents & Generative UI. React + Angular

OPENCLAW

Machine-readable data

The same record, as JSON, for agents and crawlers.

{
  "facts": [
    {
      "factKey": "vendor",
      "category": "vendor",
      "label": "Vendor",
      "value": "Clawhub",
      "href": "https://clawhub.ai/forhonourlx/skills/video-subtitle-extractor",
      "sourceUrl": "https://clawhub.ai/forhonourlx/skills/video-subtitle-extractor",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-10T22:42:44.197Z",
      "isPublic": true
    },
    {
      "factKey": "protocols",
      "category": "compatibility",
      "label": "Protocol compatibility",
      "value": "OpenClaw",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-forhonourlx-video-subtitle-extractor/contract",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-forhonourlx-video-subtitle-extractor/contract",
      "sourceType": "contract",
      "confidence": "medium",
      "observedAt": "2026-10-10T22:42:44.197Z",
      "isPublic": true
    },
    {
      "factKey": "traction",
      "category": "adoption",
      "label": "Adoption signal",
      "value": "1.2K downloads",
      "href": "https://clawhub.ai/forhonourlx/video-subtitle-extractor",
      "sourceUrl": "https://clawhub.ai/forhonourlx/video-subtitle-extractor",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-10T22:42:44.197Z",
      "isPublic": true
    },
    {
      "factKey": "latest_release",
      "category": "release",
      "label": "Latest release",
      "value": "2.0.0",
      "href": "https://clawhub.ai/forhonourlx/video-subtitle-extractor",
      "sourceUrl": "https://clawhub.ai/forhonourlx/video-subtitle-extractor",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-05-27T08:24:13.771Z",
      "isPublic": true
    },
    {
      "factKey": "handshake_status",
      "category": "security",
      "label": "Handshake status",
      "value": "UNKNOWN",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-forhonourlx-video-subtitle-extractor/trust",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-forhonourlx-video-subtitle-extractor/trust",
      "sourceType": "trust",
      "confidence": "medium",
      "observedAt": null,
      "isPublic": true
    }
  ],
  "events": [
    {
      "eventType": "release",
      "title": "Release 2.0.0",
      "description": "Multi-engine ASR: SenseVoice Small default + whisper.cpp GGML + openai-whisper. Auto-backend detection. ~1.5GB RAM, 5x faster, ~96% Chinese accuracy.",
      "href": "https://clawhub.ai/forhonourlx/video-subtitle-extractor",
      "sourceUrl": "https://clawhub.ai/forhonourlx/video-subtitle-extractor",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-05-27T08:24:13.771Z",
      "isPublic": true
    }
  ]
}

Record generated Oct 11, 2026.

Sponsored

Ads related to Video Subtitle Extractor and adjacent AI workflows.