video-add-captions
Add word-timed captions to an Open Recut program. Use this skill to map the canonical transcript through timeline.json, review a maintained style on source-backed pixels, render a local transparent HyperFrames PNG sequence, and register it as an overlay contribution for the shared delivery render.
Rank
62
Safety
84
Downloads
1.1k
Updated
Oct 11, 2026
Version
1.0.8
Source
CLAWHUB
About
What it does, and when to use it.
Capability contract not published. No trust telemetry is available yet. 1.1K downloads reported by the source. Last updated 10/11/2026.
Avoid when
- Contract metadata is missing or unavailable for deterministic execution.
Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing
Public facts
Every fact links back to the source it came from.
- Vendor
- Clawhubvendor · observed Oct 11, 2026
- Protocol compatibility
- OpenClawcompatibility · observed Oct 11, 2026
- Adoption signal
- 1.1K downloadsadoption · observed Oct 11, 2026
- Latest release
- 1.0.8release · observed Oct 7, 2026
- Handshake status
- UNKNOWNsecurity
Install and run
Setup complexity: low.
clawhub skill install s178pg569r7j3dkw6ekc9bn7258b0eme:video-add-captions- Install using `clawhub skill install s178pg569r7j3dkw6ekc9bn7258b0eme:video-add-captions` in an isolated environment before connecting it to live workloads.
- No published capability contract is available yet, so validate auth and request/response behavior manually.
- Review the upstream CLAWHUB listing at https://clawhub.ai/whitetowerai/video-add-captions before using production credentials.
Contract: missing
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-whitetowerai-video-add-captions/snapshot"
Documentation
CLAWHUB
160,000 characters of source documentation, loaded on request.
Extracted files
5 files captured from the source.
SKILL.md
--- name: video-add-captions description: > Add word-timed captions to an Open Recut program. Use this skill to map the canonical transcript through timeline.json, review a maintained style on source-backed pixels, render a local transparent HyperFrames PNG sequence, and register it as an overlay contribution for the shared delivery render. --- # Video Add Captions ## Dependencies `/video-understand` is a prerequisite. Run it first so captions use the validated word-level transcript and canonical timeline. Before starting, verify that it is installed. If it is not, warn the user that this prerequisite is missing and stop before processing media. Require `ffmpeg`/`ffprobe` on PATH, Python with `Pillow`, and Node.js >= 22 (for the `.mjs` scripts and `npx hyperframes`, fetched on demand). `hyperframes render`/`snapshot` drives a headless Chrome — it manages its own `chrome-headless-shell`, and falls back to a system Chrome (set `CHROME` to override) when the cached one is unusable. Check these before processing media. ## Scope This skill owns caption grouping, style selection, review, and the transparent caption track. It does not transcribe, cut, retime, grade, reframe, or choose the delivery audio policy. Run captions before content cards and motion graphics whenever either is selected. The approved caption layout establishes a reserved subtitle region that both later operations must keep clear. This relative order also applies when only one pair is active. Use `video-understand` first. If a cut exists, caption the active program timeline; do not treat source transcript seconds as program seconds. ## Protocol Inputs Required: - `work/project.json` - `work/understand/transcript.json`, with word-level source timestamps - `work/understand/media.json` - `work/timeline.json` - the source video named by `work/project.json` The timeline is the only source-to-program mapping. All ranges are half-open `[start_s, end_s)`. `scripts/build_captions.py` uses the shared `projectlib.map_transcript_to_timeline`; words in dropped source ranges disappear, retained words move to program time, and cues always break at clip boundaries. ## Durable Outputs ```text work/captions/ |-- captions-plan.json |-- caption-spatial-context.json (only for eligible B-roll composites) `-- caption-interaction.json work/cache/captions/ |-- preview-project/ |-- preview-snapshots/ |-- overlay-project/ `-- overlay-frames/frame_000001.png ... review/05-captions/ |-- captions.srt |-- captions-style-review-<UUID>.html |-- captions-style-review.html |-- captions-review.html |-- preview-early.png |-- preview-middle.png |-- preview-late.png |-- preview-no-caption.png |-- captions-evidence.json `-- captions-summary.md ``` Generated HTML, copied runtime files, extracted source frames, and overlay frames are cache. The plan, optional spatial context, SRT, decision receipt, and review evidence are durable. ## Caption Plan `work/captions/captions-plan.json` is the canonical c
_meta.json
{
"ownerId": "kn70hcphbcepjj43vwv7j0zbnh8b0jpm",
"slug": "video-add-captions",
"version": "1.0.8",
"publishedAt": 1791346398173
}scripts/caption-styles.json
{
"presets": {
"clean": {
"preset": "clean",
"font": {
"family": "system-ui, sans-serif",
"sizeRatio": 0.0416,
"weight": 800,
"color": "#F5F4ED",
"lineHeight": 1.32,
"letterSpacing": 0
},
"layout": {
"anchor": "bottom",
"align": "center",
"maxWidth": 0.92,
"paddingBottomRatio": 0.07
},
"background": {
"enabled": false,
"theme": "gray",
"shape": "rounded",
"color": "#000000",
"opacity": 0.48,
"radiusRatio": 0.018,
"paddingXRatio": 0.035,
"paddingYRatio": 0.018
},
"effects": {
"shadow": {
"strength": "none",
"color": "#000000",
"opacity": 0.55,
"offsetYRatio": 0.02,
"blurRatio": 0.04
},
"stroke": {
"strength": "none",
"color": "#000000",
"widthRatio": 0
}
},
"stroke": {
"enabled": false,
"theme": "black",
"color": "#000000",
"opacity": 0.85,
"widthRatio": 0
},
"wordHighlight": {
"enabled": true,
"mode": "textColor",
"activeColor": "#FF7A45",
"activeScale": 1.12,
"upcomingOpacity": 0.55,
"backgroundColor": "#FF7A45",
"backgroundOpacity": 0.24,
"backgroundRadiusRatio": 0.01
},
"animation": {
"type": "pop",
"popInFrames": 6,
"popOutFrames": 4,
"translateYPx": 22
}
},
"minimal": {
"preset": "minimal",
"font": {
"family": "system-ui, sans-serif",
"sizeRatio": 0.0352,
"weight": 650,
"color": "#FFFFFF",
"lineHeight": 1.342,
"letterSpacing": 0
},
"layout": {
"anchor": "bottom",
"align": "center",
"maxWidth": 0.86,
"paddingBottomRatio": 0.075
},
"background": {
"enabled": false,
"theme": "gray",
"shape": "rounded",
"color": "#000000",
"opacity": 0.32,
"radiusRatio": 0.014,
"paddingXRatio": 0.028,
"paddingYRatio": 0.014
},
"effects": {
"shadow": {
"strength": "none",
"color": "#000000",
"opacity": 0,
"offsetYRatio": 0,
"blurRatio": 0
},
"stroke": {
"strength": "none",
"color": "#000000",
"widthRatio": 0
}
},
"stroke": {
"enabled": false,
"theme": "black",
"color": "#000000",
"opacity": 0.85,
"widthRatio": 0
},
"wordHighlight": {
"enabled": false,
"mode": "none",
"activeColor": "#FFFFFF",
"activeScale": 1,
"upcomingOpacity": 1,
"backgroundColor": "#FFFFFF",
"backgroundOpacity": 0,
"backgroundRadiusRatio": 0.01
},
"animation": {
"reference/caption-feedback-mapping.md
# Caption Feedback Mapping This skill is agent-facing. In the canonical workflow, users choose a style and approve or revise source-backed previews by copying structured summaries from the bound HTML review pages. Standalone compatibility accepts an exact gallery combination ID or `skip`, and accepts `approve` only after source-backed preview evidence exists. Historical non-English aliases remain accepted silently but are not user instructions. The agent maps only recorded user feedback to optional JSON overrides accepted by `scripts/generate_caption_project.mjs`. ## Safe Edit Points - Use `scripts/caption-styles.json` as the source of official preset and theme names. - Record the exact gallery response with `scripts/caption_interaction.mjs select`. - Pass `--interaction-state` to the generator. The generator reads the selected preset, themes, and Karaoke value from that state and rejects conflicting flags. - Put only requested property overrides in a JSON file passed with `--overrides`. - Edit caption cue JSON only when correcting subtitle text or timing data. The user may skip gallery selection, which explicitly chooses `clean`. The source-backed preview confirmation cannot be skipped. Canonical full rendering requires the exact copied approval summary; standalone compatibility requires the exact public response `approve`. ## Official Style Vocabulary Official presets: - `clean` - `minimal` - `social-bold` - `pill` - `boxed` - `stroked` - `shorts` Background themes: - `gray` - `yellow` - `blue` - `pink` - `green` Stroke themes: - `black` - `yellow` - `blue` - `pink` - `green` Shorts highlight colors: - `yellow` -> `#F8F54F` - `green` -> `#21D32E` - `orange` -> `#F8BD6D` - `black` -> `#000000` - `blue` -> `#2563EB` - `pink` -> `#DB2777` Karaoke is an option, not a preset. Preview ids such as `pill-yellow`, `boxed-green`, `stroked-blue`, `shorts-yellow`, and `social-bold-karaoke` are preset/config combinations, not official presets. ## Feedback Rules Font and visual intensity: - "make the text bigger" -> increase `font.sizeRatio` - "make the text smaller" -> decrease `font.sizeRatio` - "cleaner" -> move toward `clean` or `minimal` - "more subtle" -> move toward `minimal` - "more eye-catching" -> move toward `social-bold` - "more like TikTok", "big short-video captions" -> move toward `social-bold` - "vertical shorts style", "Reels style", "YouTube Shorts style" -> use `shorts` Background: - "no background" -> `background.enabled = false` - "add a background" -> choose `pill` or `boxed` - "rounded background", "capsule background" -> `preset = "pill"` or `background.shape = "pill"` - "small rounded background", "rectangle background bar" -> `preset = "boxed"` or `background.shape = "rounded"` - "make the background more transparent" -> decrease `background.opacity` - "make the background stronger" -> increase `background.opacity` - "gray background" -> `background.theme = "gray"` - "yellow background" -> `background.theme = "y
reference/caption-rules.md
# Caption Rules and Data Shape
## How build_captions.py chunks words into cues
A new cue is closed when any of these conditions is met:
- the word ends a sentence: `.`, `?`, `!`, or ellipsis
- adding the next word would exceed the line budget: `--max-chars` x `--max-lines`
- adding the next word would exceed `--max-dur` seconds on screen
- there is a speech gap of at least `--gap` seconds before the next word
Each cue keeps per-word timings, so the renderer can highlight the word currently
being spoken. `lines[]` is the cue wrapped to `--max-chars` for display and for
the SRT file.
Useful defaults:
```text
--max-chars 42 --max-lines 2 --max-dur 6 --gap 0.6
```
For fast-cut vertical/social captions, try shorter cues:
```text
--max-chars 24 --max-dur 3
```
## captions.json schema
```jsonc
[
{
"index": 1,
"start": 1.0,
"end": 3.4,
"text": "Hey, it's Thariq from the Claude Code team.",
"lines": ["Hey, it's Thariq from the Claude", "Code team."],
"words": [
{ "word": "Hey,", "start": 1.0, "end": 1.2 }
]
}
]
```
`captions.srt` is the same content as portable SubRip. Use it as a sanity read
or to hand to a player or another tool.
## Renderer Contract
- Treat `captions.json` as renderer-neutral cue data.
- Preserve `start`, `end`, `text`, `lines[]`, and per-word timings when passing cues
to a renderer.
- Karaoke is a true/false option, not a preset. Use `karaoke: true` for per-word
highlight and `karaoke: false` for plain blocks.
- Keep style and compositing implementation outside this data-shaping contract.
## Expressive Planning
Standard is the default. A Standard canonical plan needs no `presentation` field,
and the legacy top-level cue array remains Standard-compatible. Expressive must be
explicitly requested with `--presentation-mode expressive`; it is a presentation
mode, not a preset.
`build_captions.py` creates only the base cues and a draft planning shell. It does
not infer placement. The Agent reads the transcript,
available understanding artifacts, timeline, generated cues, and necessary real
visual evidence, then fills the complete plan once for the whole program.
Agent layout rules:
1. Default to `bottom-standard`.
2. Keep ordinary explanatory sentences in a continuous bottom layout when possible.
3. Use `center-emphasis` for short keywords, numbers, or conclusions when emphasis is justified.
4. Do not mechanically alternate between bottom and center.
5. Do not change position in the middle of a sentence or cue.
6. Prefer merging adjacent ordinary cues into one stable layout beat.
7. Avoid repeated consecutive `center-emphasis` beats.
8. Every position change must have a semantic reason recorded in the beat rationale.
9. When uncertain, fall back to `bottom-standard`.
A layout beat covers one or more complete, contiguous cues. Beat IDs must be unique;
beats must follow time and cue order, must not overlap, and must not start or end
inside a cue. A completed Expressive plan covers eAionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
activepieces
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
CopilotKit
The Frontend for Agents & Generative UI. React + Angular
Machine-readable data
The same record, as JSON, for agents and crawlers.
{
"facts": [
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Clawhub",
"href": "https://clawhub.ai/whitetowerai/skills/video-add-captions",
"sourceUrl": "https://clawhub.ai/whitetowerai/skills/video-add-captions",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-11T09:51:19.684Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-whitetowerai-video-add-captions/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-whitetowerai-video-add-captions/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-11T09:51:19.684Z",
"isPublic": true
},
{
"factKey": "traction",
"category": "adoption",
"label": "Adoption signal",
"value": "1.1K downloads",
"href": "https://clawhub.ai/whitetowerai/video-add-captions",
"sourceUrl": "https://clawhub.ai/whitetowerai/video-add-captions",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-11T09:51:19.684Z",
"isPublic": true
},
{
"factKey": "latest_release",
"category": "release",
"label": "Latest release",
"value": "1.0.8",
"href": "https://clawhub.ai/whitetowerai/video-add-captions",
"sourceUrl": "https://clawhub.ai/whitetowerai/video-add-captions",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-10-07T04:13:18.173Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-whitetowerai-video-add-captions/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-whitetowerai-video-add-captions/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
],
"events": [
{
"eventType": "release",
"title": "Release 1.0.8",
"description": "video-add-captions 1.0.8 - Documentation updated in SKILL.md: clarified motion graphics ordering, improved several explanations, and corrected minor issues. - Removed outdated skill-card.md file. - No functional or interface changes; update focuses on documentation consistency and accuracy.",
"href": "https://clawhub.ai/whitetowerai/video-add-captions",
"sourceUrl": "https://clawhub.ai/whitetowerai/video-add-captions",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-10-07T04:13:18.173Z",
"isPublic": true
}
]
}Record generated Oct 11, 2026.
