embedded-captions
Add captions or subtitles to an existing single-subject talking-head video without editing the footage. Use for plain verbatim captions, cinematic captions embedded behind the subject, VFX captions, “炸/特效/酷炫字幕,” or a named identity from the 35-style catalog. Route by visual identity, not by backend engine. The quiet `anchor` rail is the default; embed every word only when the user explicitly wants a fully cinematic treatment. The workflow runs locally end to end, including transcription and subject matting; split multi-shot footage before applying it. Skill: embedded-captions Owner: heygen-com Summary: Add captions or subtitles to an existing single-subject talking-head video without editing the footage. Use for plain verbatim captions, cinematic captions embedded behind the subject, VFX captions, “炸/特效/酷炫字幕,” or a named identity from the 35-style catalog. Route by visual identity, not by backend engine. The quiet anchor rail is the default; embed every word only
Rank
62
Safety
84
Downloads
2.3k
Updated
Oct 9, 2026
Version
1.0.25
Source
CLAWHUB
About
What it does, and when to use it.
Capability contract not published. No trust telemetry is available yet. 2.3K downloads reported by the source. Last updated 10/9/2026.
Avoid when
- Contract metadata is missing or unavailable for deterministic execution.
Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing
Public facts
Every fact links back to the source it came from.
- Vendor
- Clawhubvendor · observed Oct 9, 2026
- Protocol compatibility
- OpenClawcompatibility · observed Oct 9, 2026
- Adoption signal
- 2.3K downloadsadoption · observed Oct 9, 2026
- Latest release
- 1.0.25release · observed Oct 4, 2026
- Handshake status
- UNKNOWNsecurity
Install and run
Setup complexity: low.
clawhub skill install s17fpgb0p797dzkbtbrxw5x1hh89qs64:embedded-captions- Setup complexity is classified as HIGH. You must provision dedicated cloud infrastructure or an isolated VM. Do not run this directly on your local workstation.
- Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.
Contract: missing
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-heygen-com-embedded-captions/snapshot"
Documentation
CLAWHUB
160,000 characters of source documentation, loaded on request.
Extracted files
5 files captured from the source.
SKILL.md
---
name: embedded-captions
description: >
Add captions or subtitles to an existing single-subject talking-head video without editing the
footage. Use for plain verbatim captions, cinematic captions embedded behind the subject, VFX
captions, “炸/特效/酷炫字幕,” or a named identity from the 35-style catalog. Route by visual
identity, not by backend engine. The quiet `anchor` rail is the default; embed every word only
when the user explicitly wants a fully cinematic treatment. The workflow runs locally end to end,
including transcription and subject matting; split multi-shot footage before applying it.
---
**Plugin installs:** Before setup or freshness commands, follow [plugin execution rules](../hyperframes/references/plugin-installation.md) when this skill is inside a HyperFrames plugin. Standalone installs keep the update instructions below.
> **First, keep this skill fresh — confirm with the user before running:** `npx hyperframes skills update embedded-captions`. A fast no-op when everything is current; otherwise it refreshes this skill plus the core domain skills it depends on before you rely on them.
# Embedded Captions
**One catalog, picked up front** ([CATALOG.md](CATALOG.md) — 35 identities; the engines behind it are backend detail). **Standard** (default) builds a clean verbatim **rail** (lower-third subtitle carrying most text) + an **embed** climax composited _into_ the scene behind the subject at the peak. **Cinematic** is pure embed — no rail, every caption composited behind the subject (hero typography, accumulation, occlusion as the effect). **Theme** is a complete themed constitution — body paradigm × hero setpiece × front fx × plate reaction, composed from registries ([themes/README.md](themes/README.md)): `ordnance` `terminal` `neonsign` `stardust` `stomp`. Most explainer / voiceover is **Standard**; **embed is the scarce, earned peak** — embedding every word is the common mistake; Theme is for VFX-grade asks ("炸", "特效", "像 AE 做的").
---
## Runtime prerequisites
Plugin installs use the bundled, manifest-pinned CLI for matting, transcription,
and rendering; no source checkout is required. The local preview and caption
measurement helpers also need Sharp, Puppeteer (with its Chromium browser), and
GSAP. Install these in the **caption project**, not inside the read-only plugin:
```bash
npm install --prefix <project> --save-dev --save-exact [email protected] [email protected] [email protected]
```
Keep the project's lockfile. If these dependencies already exist, use its locked
versions instead of overwriting them. Bash and FFmpeg/ffprobe must be on PATH.
Matting and transcription may download their own models on first use.
Rendering waits for the CLI to exit successfully before compositing. The old
`HF_TIMEOUT_S` shell watchdog is no longer used: a large partial file is not proof
that rendering finished. An explicit built-checkout argument or `HYPERFRAMES_ROOT`
selects the contributor CLI instead of the plugin pin. Cancel a stalledna/README.md
# DNA registry — pick a visual language, not a preset A **DNA** is a complete, art-directed visual language: typeface, palette logic, motion grammar, and hero orchestration. It **parameterizes per scene** instead of shipping a fixed look: the accent color is sampled from THIS scene, the contact shadow falls along THIS scene's light, embed text blur matches THIS scene's depth-of-field, and the hero's entrance amplitude follows how hard the word was actually spoken (RMS). This replaces the template grab-bag. Six deep languages × scene adaptation beats 54 shallow presets — every render is already fitted to its footage. ## Category lock (deliveries field, enforced by the compilers) Every classic DNA's **home is Cinematic (column)** — that is where all ten were built and validated. (Standard/rail mode was retired 2026-06-12; the verbatim-rail need is served by the `anchor` theme. The old rail combos are archived outside this repo and are not distributed with the skill.) ## The ten | DNA | Register | Scene fit | Voice | | --------------- | -------------- | ----------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **cream** | premium-warm | dark / mid warm scenes (band luma < 150) | Inter, warm cream, screen blend, glowing emergence hero. The poetic default. | | **ink** | premium | **bright scenes (band luma > 150)** | Inter, near-black, multiply blend — type reads as _printed on_ the wall. Fixes the bright-scene hole. | | **editorial** | editorial-luxe | introspective / fashion / poetic, mid-dark | Bodoni Moda, bone, _lowercase italic hero_ — magazine elegance over shout. | | **keynote** | tech-premium | product / launch / founder updates | Inter 800, opaque white, line-wipe reveals, hero wipes UP. Stillness = confidence. | | **documentary** | formal | interviews, serious subject matter | Inter, bone, **burn-in reveals**, no hero. Gravitas IS the style. | | **loud** | loud | hype / sport / music / social | Anton, scene-sampled accent hero, single-unit slam + caption-layer ripple; **body announc
modes/cinematic/README.md
# Cinematic mode (pure embed) — one engine, six DNAs > Cinematic mode compiles **[../../dna/](../../dna/README.md)** through > **[engine.html](engine.html)** (`make-composition.cjs`). The old per-template HTML > shells are retired — `cinematic-cream` maps to `dna: "cream"` automatically; the other > archived templates (memory-wall / champion / portrait-header, in [\_archive/](_archive/)) > remain as design references only. Use this mode for pure-embed asks (no rail): brand film, hype, social reel, showcase. The **DNA** locks the visual language (type, palette scheme, blend, motion grammar, hero three-act); **safe-zones v2** parameterizes it to the scene (sampled accent, light- direction contact shadow, depth-match blur); **the agent decides layout only** (planes, blocks, per-line typography within the DNA). ## Workflow 1. `bash scripts/prepare.sh <project>` → matte ∥ transcript ∥ envelope → safe-zones v2 2. Pick a DNA ([../../dna/README.md](../../dna/README.md)): bright hero band → `ink`, else by register (cream / editorial / keynote / documentary / loud). Recommend, let the user pick. 3. Author `<project>/cinematic.json` — `"dna": "<name>"` + thought-blocks (schema: `scripts/make-cinematic.cjs` header) 4. `node scripts/make-cinematic.cjs <project>` → plan.json → engine-compiled index.html 5. `node scripts/preview-frames.cjs <project>` → § Visual QA (failure checks + the 5 positive checks in [../../references/reference-bar.md](../../references/reference-bar.md)) 6. `bash scripts/render-and-composite.sh <project>` → gates → final.mp4 ## What the engine generates (never author these) - word timings from the transcript; accumulate-within-block / page-flip-between-blocks - the hero hand-off + **three-act orchestration** (dim → RMS-coupled per-letter entrance → breathe + glow), per the DNA's `hero` block - scene tokens: `--accent` (sampled), contact shadow, depth blur - reading order, re-slot from measured heights, hero size/collision post-pass ## What you DON'T do - Override `.cap` color / blend / shadow / filter / motion curves — that's the DNA. Scene fights the look → pick a different DNA (bright → `ink`), never recolor. - Hand-position the hero into a clean margin (it belongs ON the subject, ~30–55% occluded — safe-zones `heroBands.best`). - Add full-frame grades/textures over the footage (hard rule: the video ships untouched). ## Adding a DNA `dna/<name>.json` — copy one, change the voice (see [../../dna/README.md](../../dna/README.md) § Adding). The engine consumes it with no code change. A DNA must be a distinct voice with a reason to exist, not a recolor.
themes/README.md
# THEME mode — composed visual constitutions
Theme mode is the third compiler (`scripts/make-theme.cjs`), beside Standard and
Cinematic. It exists because "mode" was a bundle of orthogonal axes pretending
to be one switch. A theme DNA composes its identity from registries implemented
ONCE in the compiler — **paradigms are the unit of code; DNAs are the unit of
identity**. A new look is a JSON file; only a genuinely new paradigm/setpiece
(rare) touches the compiler.
```
theme DNA = body PARADIGM how the transcript surface lives
× body LAYER fg-alpha (rail.html channel) | bg-embed
× hero SETPIECE the climax choreography
× front FX flash / rings / sparks / scanband / crowdflash / paflash (fg, over subject)
× PLATE budget charge-dim (in-page) + punch/shake/grain (_postfx.sh)
× LINKAGES declarative theme interactions
```
Standard and Cinematic were, in retrospect, two fixed points of this space.
Standard is now RETIRED (2026-06-12): its rail×embed-climax point is served by
the `anchor` theme (rail paradigm × settle setpiece — the quiet default).
Cinematic remains a separate compiler; do not re-implement it as a theme yet.
## Unification roadmap (strangler fig — interface first, engines later)
The user-facing model is already unified (SKILL.md Step 0): one catalog of
LOOKS; classic looks pick a DELIVERY (rail | column), themed looks bind their
own. "Standard/Cinematic" are delivery/compiler names, not modes. Remaining
phases, each gated on need — never rewrite for tidiness alone:
- **Phase 2 — one authoring schema.** `lines`/`minors`/`hero` are already
~90% shared between standard.json and theme.json; a router that translates a
single `caption.json` into the engine-specific file removes the last
user-visible seam. cinematic.json's blocks/planes are the odd one — map the
common fields, pass engine-specific ones through.
- **Phase 3 — engine convergence.** Port a classic delivery into make-theme
ONLY when something forces it (e.g. a classic DNA wants a plate budget or a
setpiece). Acceptance bar: blind A/B on the cap_multi regression scenes vs
the old compiler — swap engines only when indistinguishable or better. Until
then the old compilers are the reference implementation of 8 rounds of
validated typography (lockup/orbit, multi-climax, ratio-lock, per-plane
legibility, occlusion adjudication) — that machinery is the moat, not debt.
## Body paradigms (registry)
| paradigm | surface _meta.json
{
"ownerId": "kn77d06grj6xqp3dqwkk4bavhn89pegt",
"slug": "embedded-captions",
"version": "1.0.25",
"publishedAt": 1791142135378
}AionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
activepieces
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
CopilotKit
The Frontend for Agents & Generative UI. React + Angular
Machine-readable data
The same record, as JSON, for agents and crawlers.
{
"facts": [
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Clawhub",
"href": "https://clawhub.ai/heygen-com/skills/embedded-captions",
"sourceUrl": "https://clawhub.ai/heygen-com/skills/embedded-captions",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-09T16:46:35.121Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-heygen-com-embedded-captions/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-heygen-com-embedded-captions/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-09T16:46:35.121Z",
"isPublic": true
},
{
"factKey": "traction",
"category": "adoption",
"label": "Adoption signal",
"value": "2.3K downloads",
"href": "https://clawhub.ai/heygen-com/embedded-captions",
"sourceUrl": "https://clawhub.ai/heygen-com/embedded-captions",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-09T16:46:35.121Z",
"isPublic": true
},
{
"factKey": "latest_release",
"category": "release",
"label": "Latest release",
"value": "1.0.25",
"href": "https://clawhub.ai/heygen-com/embedded-captions",
"sourceUrl": "https://clawhub.ai/heygen-com/embedded-captions",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-10-04T19:28:55.378Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-heygen-com-embedded-captions/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-heygen-com-embedded-captions/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
],
"events": [
{
"eventType": "release",
"title": "Release 1.0.25",
"description": "Synced from 0c3e244 (main)",
"href": "https://clawhub.ai/heygen-com/embedded-captions",
"sourceUrl": "https://clawhub.ai/heygen-com/embedded-captions",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-10-04T19:28:55.378Z",
"isPublic": true
}
]
}Record generated Oct 9, 2026.
