music-to-video
Turn a music track (an audio file, a video to pull audio from, or a track generated from a mood brief) into a beat-synced video — lyric video, slideshow, or kinetic promo. The music drives all pacing; any user-supplied images/videos are cut onto the same beat grid, and a complete video needs zero assets. Narrated pieces → the input-matched workflow (see /hyperframes). Unclear → /hyperframes.
Rank
62
Safety
84
Downloads
2.2k
Updated
Oct 9, 2026
Version
1.0.25
Source
CLAWHUB
About
What it does, and when to use it.
Capability contract not published. No trust telemetry is available yet. 2.2K downloads reported by the source. Last updated 10/9/2026.
Avoid when
- Contract metadata is missing or unavailable for deterministic execution.
Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing
Public facts
Every fact links back to the source it came from.
- Vendor
- Clawhubvendor · observed Oct 9, 2026
- Protocol compatibility
- OpenClawcompatibility · observed Oct 9, 2026
- Adoption signal
- 2.2K downloadsadoption · observed Oct 9, 2026
- Latest release
- 1.0.25release · observed Oct 7, 2026
- Handshake status
- UNKNOWNsecurity
Install and run
Setup complexity: low.
clawhub skill install s17fpgb0p797dzkbtbrxw5x1hh89qs64:music-to-video- Install using `clawhub skill install s17fpgb0p797dzkbtbrxw5x1hh89qs64:music-to-video` in an isolated environment before connecting it to live workloads.
- No published capability contract is available yet, so validate auth and request/response behavior manually.
- Review the upstream CLAWHUB listing at https://clawhub.ai/heygen-com/music-to-video before using production credentials.
Contract: missing
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-heygen-com-music-to-video/snapshot"
Documentation
CLAWHUB
160,000 characters of source documentation, loaded on request.
Extracted files
5 files captured from the source.
SKILL.md
--- name: music-to-video description: "Turn a music track (an audio file, a video to pull audio from, or a track generated from a mood brief) into a beat-synced video — lyric video, slideshow, or kinetic promo. The music drives all pacing; any user-supplied images/videos are cut onto the same beat grid, and a complete video needs zero assets. Narrated pieces → the input-matched workflow (see /hyperframes). Unclear → /hyperframes." --- **Plugin installs:** Before setup or freshness commands, follow [plugin execution rules](../hyperframes/references/plugin-installation.md) when this skill is inside a HyperFrames plugin. Standalone installs keep the update instructions below. > **First, keep this skill fresh — confirm with the user before running:** `npx hyperframes skills update music-to-video`. A fast no-op when everything is current; otherwise it refreshes this skill plus the core domain skills it depends on before you rely on them. # music-to-video — one music-grounded, beat-synced video workflow Use this skill to turn a **music track** into a beat-synced HyperFrames video. You analyze the track once, lay out the frames, fill in a per-frame plan, and build each frame as a composition. The input is a music track plus optional user images or videos — there is **no narration and no website capture**. Typography and templates are the floor (a complete video needs zero assets); any media the user supplies is cut in on the same beat grid. You are the **orchestrator**. Work in `videos/<project>/`. Run the steps in order and pass each **Gate** before moving on. Two steps need the user: **Step 3** (plan approval) and **Step 6** (render approval) — both are checkpoint gates per `../hyperframes/references/brief-contract.md` (read it before Step 0): in autonomous mode, post the Step 3 summary as a heads-up and proceed; Step 6 render approval is still asked, as the one kept question. Do every step yourself except **Step 4**, where you dispatch **one sub-agent per frame**. Keep design and motion rules out of this file — they live in `references/` and the `frame-worker` sub-agent. `SKILL_DIR` = this skill directory. `PROJECT_DIR` = `videos/<project-name>/`. Workflow: Step 0 setup → `hyperframes.json` + `assets/bgm.mp3`; Step 1 analyze → `audiomap.json`; Step 2 skeleton → `STORYBOARD.md` (frames, groups `TBD`); Step 3 plan → complete `STORYBOARD.md` + `frame.md`; Step 4 build → `compositions/frames/NN-*.html`; Step 5 assemble → `index.html`; Step 6 render → `renders/video.mp4`. ## Two ideas that shape everything - **One analyzer, and you trust it.** `analyze-beatgrid.py` is the only beat analyzer — never re-measure beats with another tool or by ear. Its energy / density / rolls / onsets / silences are always reliable. Its `bpm` and `beats_sec` are reliable **only when the music is genuinely rhythmic**; on calm music the grid is a metronome the tracker imposed, so pace by phrases and energy instead and never hard-cut to it. Deciding which case you're in
_meta.json
{
"ownerId": "kn77d06grj6xqp3dqwkk4bavhn89pegt",
"slug": "music-to-video",
"version": "1.0.25",
"publishedAt": 1791358267609
}references/frame-skeleton.md
# Frame skeleton (Step 2) — read the music, lay out the frames At Step 2 **you (the orchestrator)** read `audiomap.json` and write the **skeleton** of `STORYBOARD.md` directly: cut the track into **frames** (one frame = one composition file = one scene), and for each frame set its **span**, its **pacing** (does this stretch want hard beat-cuts, or calm phrase/energy flow?), its **mood**, and a one-line **feel** note. You **classify and lay out the spine only.** You do **not** pick templates, write copy, choose colors/fonts, or decide a frame's groups — those are Step 3 (the plan fills each frame in place). Leave every frame's `### Groups` as `TBD (Step 3)` and the frontmatter `style` blank. There is **no intermediate JSON** — the skeleton _is_ the start of `STORYBOARD.md`. Step 3 edits the same file. ## The trust boundary (read this first) `audiomap.json` is one analyzer's output. Some fields are robust on **any** music; some are reliable only when the music is **actually rhythmic**. This decides each frame's `pacing`: | Field | Trust | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `energy_phases[]` (level / energy / density / feel), `events[]` + `onset_rate`, `rolls[]` (and their **absence**), `silences[]`, `hard_stops[]`, `key_moments[]`, `phrases[]`, `audio.duration_sec` | **Always** — robust measurements | | `tempo.bpm`, `grid.beats_sec` / `downbeats_sec` **precision** | **Only when the music is rhythmic.** On calm / sparse material the beat grid is a metronome the tracker _imposes_ (often octave-doubled) — usually **more grid beats than real onsets**. Do **not** anchor cuts to it there. | - **Grid is reliable** when: rolls present, and/or dense phases, and/or high `onset_rate` with a steady grid. - **Grid is fictional** when: `rolls`≈0, mostly `sparse` phases, low `onset_rate` → pace by `phrases[]` + `energy_phases[]`, not beats. ## How to lay out f
references/montage.md
# Asset treatments — weaving user media onto the beat spine When the user supplies images/videos, a group can be an **asset treatment** instead of a typographic template/free-compose. Assets are an **additive ingredient on the same beat spine** — never a separate pipeline. Typography/templates stay the floor: if no asset fits a group, fall back to a template/free group (a complete video needs zero assets). The planner (Step 3) picks the treatment and the clips + anchors (WHAT); the frame-worker realizes it inside the frame file (HOW). **Obey the frame's `pacing`.** ## The three treatments ### `beat_cut` — one clip per anchor (only on a `beat_cut` frame) The asset-driven analogue of a per-onset typographic group: cut to a new clip on each anchor (the frame's beats/onsets from the audiomap). Each clip is a `class="clip"` element (`<img>` for a photo, **muted** `<video>` for a motion clip) placed at its anchor with `data-start`/`data-duration`/`data-track-index` per the core clip contract. (Muted on purpose: the music track drives the sound.) Between clips, crossfade the outgoing content to `opacity:0` ending **at** the next anchor. Cut on the **strong** anchors; land a hero clip on a `key_moment`/downbeat. ### `ken_burns` — slow push on one clip (fits a `phrase_flow` frame) For calm frames: one clip held over the span with a slow scale/translate push (e.g. scale 1.0→1.08 + a small drift) eased across the whole `span_sec` — paced by the frame, not by beats. No hard cuts. Crossfade in/out at the frame edges. This is the right asset treatment when the beat grid is unreliable (calm music). ### `bg_under_text` — clip dimmed behind a template/free group A full-bleed clip dimmed ~30–50% as the background of a group whose foreground is a template or free-compose typographic treatment. The text rides on the same anchors; the clip is the bed. Use when the user wants their footage present but the message must stay readable. ## Rules - **`pacing` decides the treatment**: `beat_cut` only on a `beat_cut` frame; on a `phrase_flow` frame use `ken_burns` or a slow crossfade — **never** per-onset hard cuts on the (unreliable) calm grid. - **Clips are muted; the root owns audio.** Mount each `<video class="clip">` **muted**, as a direct child of the frame root (never nested in another timed element, or the renderer freezes it). The BGM is the only audio in v1. - **Crossfades animate `opacity`/`autoAlpha`**, never `visibility`/`display` on a `.clip` (the framework owns clip visibility — that trips `gsap_animates_clip_element`). - **Backgrounds dim ~30–50%** so any foreground text stays legible. - Anchors are **track seconds from `audiomap.json`**; the worker subtracts the frame start to get frame-local time. - Local staged assets only (`assets/` via `stage-assets.mjs`); never remote URLs. ## Deferred hook (not v1) A clip that should play **its own sound** (interview cut, lyric clip) needs a sibling `<audio>` mounted at the **root** by the asse
references/motion-primitive-catalog.md
# Motion-primitive catalog — the free-compose menu The atomic layer: one anchor → one micro-move. When no template fits a group, free-compose by naming primitives from here. Scan **anchor** + **best span** + **what it does**, then pick the smallest set that carries the group. ## Timing & latency (applies to every primitive) - **Hard hits are 0ms.** Cuts, palette flips, content swaps, freezes are `tl.set(...)` with no duration — the percussion _is_ the motion. Easing a hit kills it. - **Lead the anchor.** A move that must _land_ on a beat (a wipe covering the frame, a count-up locking, two blocks colliding) starts **~40–190ms early** so it completes ON the anchor. Reactive entrances (something appearing _because_ of the hit) fire 0–45ms after. - **Eased entrances: 300–500ms** (scale punch, slides, camera pushes). **Macro builds: 800–2000ms** spanning a whole roll / silence. - **Per-bar caps:** one accumulating element per hit (not a burst); a camera move at most once per phrase, never per beat; a dense flip/strobe system runs ≤2–3s. - **Tension-builds lock.** A count-up / sequential build / morph must _resolve on_ a downbeat or hard_stop, never trail off mid-bar. - **Best span means active motion.** The catalog's span guidance is not a license to stretch one primitive over a whole frame. If a free-composed group runs longer than the listed span, add a hold / bed / next primitive, or split the frame into another group at the next musical anchor. ## Catalog | id | anchor | best span | what it does | | --------------------- | -------------------------------- | ----------------- | ------------------------------------------------------------------------- | | `hypercut-whip` | beat / hard_stop | 0.18-0.45s | fast whip-pan hard cut between frames | | `kinetic-letter-in` | downbeat / phrase | 0.4-1.2s | per-letter kinetic entrance | | `braam-punch` | drop / surge | 0.2-0.9s active | big impact: scale + weight slam | | `chromatic-split` | snare / glitch / surge | 0.1-0.6s | RGB channel split / glitch on a word | | `mask-reveal` | section_start / downbeat | 0.5-1.2s | clip-path mask wipe reveal | | `screen-shake` | drop / crash / kick | 0.1-0.5s | camera / screen shake jitter | | `binary-decrypt` | roll / build | 0.8-2.5s | scramble→decode text (binary → word) | | `dolly-zoom` | phrase / build | 1.2-2.5s | vertigo dolly-zoom (sc
AionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
activepieces
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
CopilotKit
The Frontend for Agents & Generative UI. React + Angular
Machine-readable data
The same record, as JSON, for agents and crawlers.
{
"facts": [
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Clawhub",
"href": "https://clawhub.ai/heygen-com/skills/music-to-video",
"sourceUrl": "https://clawhub.ai/heygen-com/skills/music-to-video",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-09T17:28:08.435Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-heygen-com-music-to-video/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-heygen-com-music-to-video/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-09T17:28:08.435Z",
"isPublic": true
},
{
"factKey": "traction",
"category": "adoption",
"label": "Adoption signal",
"value": "2.2K downloads",
"href": "https://clawhub.ai/heygen-com/music-to-video",
"sourceUrl": "https://clawhub.ai/heygen-com/music-to-video",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-09T17:28:08.435Z",
"isPublic": true
},
{
"factKey": "latest_release",
"category": "release",
"label": "Latest release",
"value": "1.0.25",
"href": "https://clawhub.ai/heygen-com/music-to-video",
"sourceUrl": "https://clawhub.ai/heygen-com/music-to-video",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-10-07T07:31:07.609Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-heygen-com-music-to-video/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-heygen-com-music-to-video/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
],
"events": [
{
"eventType": "release",
"title": "Release 1.0.25",
"description": "Synced from 4fcad1e (main)",
"href": "https://clawhub.ai/heygen-com/music-to-video",
"sourceUrl": "https://clawhub.ai/heygen-com/music-to-video",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-10-07T07:31:07.609Z",
"isPublic": true
}
]
}Record generated Oct 9, 2026.
