audio-prompting
Use when crafting TTS, music, or bed prompts for any generative audio model — director style, song structure, and post-production layering.
Rank
62
Safety
84
Downloads
1.0k
Updated
Oct 11, 2026
Version
1.0.14
Source
CLAWHUB
About
What it does, and when to use it.
Capability contract not published. No trust telemetry is available yet. 1K downloads reported by the source. Last updated 10/11/2026.
Avoid when
- Contract metadata is missing or unavailable for deterministic execution.
Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing
Public facts
Every fact links back to the source it came from.
- Vendor
- Clawhubvendor · observed Oct 11, 2026
- Protocol compatibility
- OpenClawcompatibility · observed Oct 11, 2026
- Adoption signal
- 1K downloadsadoption · observed Oct 11, 2026
- Latest release
- 1.0.14release · observed Sep 29, 2026
- Handshake status
- UNKNOWNsecurity
Install and run
Setup complexity: low.
clawhub skill install s173djpkkm1x4yfzg522h5vfp989mrn5:audio-prompting- Install using `clawhub skill install s173djpkkm1x4yfzg522h5vfp989mrn5:audio-prompting` in an isolated environment before connecting it to live workloads.
- No published capability contract is available yet, so validate auth and request/response behavior manually.
- Review the upstream CLAWHUB listing at https://clawhub.ai/pruna-ai/audio-prompting before using production credentials.
Contract: missing
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-pruna-ai-audio-prompting/snapshot"
Run-check
$0.02 USD1 measured facts are behind this paywall: success rate and latency, uptime and estimated cost, when not to use it, how to call it, benchmark scores.
Agents pay $0.02 in USDC. A card payment is $0.50, the smallest a card allows.
Documentation
CLAWHUB
146,307 characters of source documentation, loaded on request.
Extracted files
5 files captured from the source.
SKILL.md
--- name: audio-prompting description: Use when crafting TTS, music, or bed prompts for any generative audio model — director style, song structure, and post-production layering. license: MIT metadata: version: "1.0.14" package: pruna-skills --- # Audio prompting Vendor-neutral craft for **speech, music, and beds**. Works with Gemini TTS, ElevenLabs, Music 2.5, Stable Audio, Suno, and similar APIs. ## Install | Skill | Description | Install | | --- | --- | --- | | `audio-prompting` | Use when crafting TTS, music, or bed prompts for any generative audio model — director style, song structure, and post-production layering. | `npx skills add PrunaAI/pruna-skills@audio-prompting -y` | | `generation-diversity` | Use when writing any generative prompt — ritual seed, explicit structure, scenario axes, and quality gates before paid API calls. | `npx skills add PrunaAI/pruna-skills@generation-diversity -y` | ## When to use - Director-style TTS prompts and inline performance tags - Full songs with vocals vs instrumental beds - Choosing when to embed audio in a video model vs mix in post - Narration + bed layering pipelines ## Works with Gemini Flash TTS, ElevenLabs, Music 2.5, Stable Audio, Suno, Udio, and other audio models. Pair with `video-prompting` when uploading VO into a video model. ## When NOT to use Use a different skill instead: | Skill | Description | Install | | --- | --- | --- | | `music-2.5` | Use when someone wants an original AI song with vocals — sung lyrics, a style prompt track, or source audio for a music video. | `npx skills add PrunaAI/[email protected] -y` | | `stable-audio-2.5` | Use when someone wants light instrumental background music — an ambient bed under dialogue or underscore for reels and explainers. | `npx skills add PrunaAI/[email protected] -y` | | `gemini-3.1-flash-tts` | Use when someone needs spoken narration or voiceover — explainer tracks, documentary lines, or voice to pair with generated video. | `npx skills add PrunaAI/[email protected] -y` | | `video-prompting` | Use when crafting video or motion prompts for any generative model — dramaturgy, camera, physics-safe motion, frame anchors, and clip chaining. | `npx skills add PrunaAI/pruna-skills@video-prompting -y` | | `image-prompting` | Use when crafting still-image prompts for any generative model — composition, identity sheets, edits, try-on, and photoreal personas. | `npx skills add PrunaAI/pruna-skills@image-prompting -y` | | `video-editing` | Use when assembling or polishing already-rendered clips with ffmpeg — concat, crossfades, burned captions and subtitles, text/logo overlays, before/after sliders, background music beds, platform export — or when composing a multi-layer HTML combination video with Hyperframes. Not for AI video generation, prompt craft, or model-based video edits. | `npx skills add PrunaAI/pruna-skills@video-editing -y` | ## Guide habit In the **first reply**, name `` `audio-prompting` ``
_meta.json
{
"ownerId": "kn7cagwf7q3t0cxrgteb7xk0bh81j0eb",
"slug": "audio-prompting",
"version": "1.0.14",
"publishedAt": 1790695895155
}references/audio-post-production.md
# Audio post-production (Pruna + Replicate) How to choose and **layer** audio when building reels, multi-scene films, and launch videos. **Prompt craft (how to write):** [tts-style-prompting.md](./tts-style-prompting.md) · [music-and-bed-prompting.md](./music-and-bed-prompting.md). For in-video audio modes and talking-head VO, install and follow `video-prompting`. **Multi-scene narrated films:** scene anchor triple lives in `video-prompting` — pass TTS to **`p-video`** as `input.audio` with `image` + `last_frame_image`; do not post-mux unless re-render is impossible. Workflow: `narrated-multi-scene`. **Visual-only transitions (no VO):** scene anchor pair in `video-prompting` — `duration` instead of `audio`. Workflow: `visual-transition-reel`. ## Audio-led `p-video` (required when VO/narration exists) When narration, TTS, or a timed audio slice is available **before** video render: 1. Upload the audio file to Pruna (`POST /v1/files`) — see `pruna-api`. 2. Pass `urls.get` as **`input.audio`** on **`p-video`** (or **`p-video-avatar`** for human lip-sync). 3. **Omit `duration`** — clip length follows the audio (capped at **20s** on P-API); the model syncs motion to speech. 4. Set **`save_audio`: true** so the full line is embedded in the output clip. 5. **Probe TTS length** before render — per-scene lines should be **≤ ~19s** or the API truncates the tail even when `audio` is set. 6. **Concat** clips in order (narration already on each clip). Optional bed mixed **under** VO in post. **Never** generate silent `p-video` and ffmpeg-mux narration afterward unless re-render is impossible — post-mux **truncates** lines longer than the video slot (common with Gemini TTS). **Over 20s?** Shorten scene copy → tighten TTS pace in `style_prompt` → split into two scene rows (each with its own triple). See `narrated-multi-scene` duration gate. | Need | Approach | Skill | |------|----------|-------| | Lip-sync / duration locked to VO | Upload audio → `p-video` with `audio` | `p-video` | | Documentary / story narrator | Gemini Flash TTS → upload → video | `gemini-3.1-flash-tts` | | Light instrumental under dialogue | Stable Audio bed under VO | `stable-audio-2.5` | | Full song with sung vocals | Music 2.5 track | `music-2.5` | | Speaking on-camera character | Portrait + script / audio | `p-video-avatar` | **Env:** Pruna calls need `PRUNA_API_KEY`; Replicate audio tools need `REPLICATE_API_TOKEN`. Assembly steps need **`ffmpeg`** / **`ffprobe`**. Credentials: `pruna-api`. Shared ffmpeg recipes (concat, captions, bed mix, export): **`video-editing`**. ## Layering matrix | Stack | Primary audio | Secondary | Mix notes | |-------|---------------|-----------|-----------| | **Silent B-roll** | — | — | Concat video only | | **Native `p-video` sound** | Model output | — | Keep `save_audio` default; normalize in assembly if scenes differ | | **Narration only (fallback)** | Gemini TTS | — | Post-mux only when audio-led `p-video` is not suitable — prefer **Pipelin
references/music-and-bed-prompting.md
# Music and bed prompting Prompt craft for `music-2.5` (songs with vocals) and `stable-audio-2.5` (instrumental beds). Mix/stack: [audio-post-production.md](./audio-post-production.md). In-video sync: install `video-prompting`. ## Music 2.5 (full song) Stack: **genre + mood + vocal + tempo + instruments + production feel** (≤ ~2000 chars). Pair with a **lyrics** field — verse / chorus / bridge structure for vocal tracks. ```text Indie pop, uplifting, warm female vocal, 92 BPM, acoustic guitar and mellow synth pads, no harsh distortion ``` **Lyrics:** write singable lines per section (verse, chorus, optional bridge). Lock structure before the paid call; same lyrics + prompt still yield different arrangements — lock seeds only when the user asks. | Include | Avoid | |---------|-------| | Genre, BPM, vocal timbre, key instruments | Vague `epic cinematic masterpiece` | | Explicit `no harsh distortion` / energy caps when needed | Contradictions (`lo-fi quiet` + `stadium EDM drop`) | ## Stable Audio 2.5 (beds under VO) Instrumental, understated, mix-friendly: ```text Instrumental light electronic pop bed, soft groove and mellow synth pads, calm positive tech atmosphere, understated background music, no vocals, 94 BPM ``` Rules: - Always **`no vocals`** when under narration - Keep energy **below** dialogue — assembly mixes ~0.08–0.15 under VO - Tag style works well; keep prompts short ## Which tool? | Need | Tool | |------|------| | Sung song / music video source | Music 2.5 | | Quiet bed under TTS or avatar | Stable Audio 2.5 | | Diegetic SFX inside `p-video` | Native `save_audio` / prompt cues — not these models | ## Pre-send - [ ] Song vs bed chosen deliberately - [ ] Bed: no vocals + BPM + understated - [ ] Song: genre/mood/vocal/tempo + **lyrics** structure present - [ ] Duration matches scene or assembly plan ## Worked examples ### Full song (Music 2.5) — indie pop, remote-work theme User lock: warm female vocal, ~92 BPM, acoustic + mellow synth, **not** EDM drop. **Style prompt** (`prompt` field): ```text Indie pop, warm and hopeful, female vocal, 92 BPM, acoustic guitar and mellow synth pads, intimate bedroom-production feel, no harsh distortion, no stadium drop ``` **Lyrics** (`lyrics` field — verse / chorus / bridge): ```text [Verse 1] Coffee rings on the desk again Window light on a second screen Slack pings like a metronome Building something from my home [Chorus] We're still here, we're still on Pixels bridge what miles have drawn Heart in the work, voice in the song Remote but never alone [Verse 2] Cat walks across the keyboard line Deadline hums but the team's aligned Same sky, different time zones Same goal in our headphones [Bridge] When the Wi‑Fi stutters, we don't fold Call reconnects — the story holds [Chorus] We're still here, we're still on ... ``` Confirm lyrics + style before `POST`. For music-video cut points later → `whisperx` after the track exists. ### Instrumental bed (Stable Audio 2.5
references/tts-style-prompting.md
# TTS style prompting (Gemini 3.1 Flash TTS) Director-style `prompt` craft for `gemini-3.1-flash-tts`. Upload results to Pruna for Mode B in-video audio (install `video-prompting`). Layering: [audio-post-production.md](./audio-post-production.md). ## Align three channels | Channel | Role | |---------|------| | `text` | Spoken words (+ optional inline `[tags]`) | | `prompt` | Tone, pace, accent, character — max ~4k bytes | | `[tags]` in text | Momentary direction matching `prompt` | All three must point the **same** emotional direction. **Bracket clarity:** `[tags]` live only in this TTS `text` field. Still typography uses double-quoted `"[STRING]"` — see `image-prompting`. Native clip dialogue (`[subject] says "[LINE]"`) is Mode A in `video-prompting` — not this skill. ## Human narrator defaults ```text Warm storybook narrator, gentle pace, empathetic, no announcer voice. ``` ```text Natural documentary host, measured pacing, clear consonants, calm authority. ``` Avoid: radio-ad hype, “cinematic trailer voice”, reading the product brief into `prompt`. ## Duration gate for `p-video` Audio-led clips cap at **20s** (keep TTS ≤ **~19s**). If `ffprobe` is long: 1. Shorten `text` 2. Add pace to `prompt`: `brisk pace, ~2.3 words per second, no filler` 3. Split into two scene rows Never rely on post-mux over silent video. ## Avatar vs TTS | Path | Fields | |------|--------| | Narrator B-roll | Gemini TTS → `p-video` `input.audio` | | On-camera speaker | `p-video-avatar` `voice_script` + `voice_prompt` — **not** this TTS `prompt` | Do not paste VO into avatar `voice_prompt`. ## Pre-send - [ ] `prompt` / `text` / tags aligned - [ ] Length probed for `p-video` - [ ] Voice + language recorded in manifest for regen consistency ## Worked example — explainer narration (aligned channels) User lock: documentary explainer, **measured** pace, ~45s script, will feed `p-video` Mode B. **`prompt`** (director style): ```text Natural documentary host, measured pacing, clear consonants, calm authority, empathetic, no radio-ad hype ``` **`text`** (spoken words + inline tags): ```text [warm] Most teams treat diversity as a checkbox. [pause] But the ritual seed is what breaks repetition before you ever hit generate. [emphasis] Lock the brief first — then rotate free axes. [measured] Same subject, fresh camera and light, every panel. ``` **Alignment check:** tags match the calm documentary `prompt` — no `[shout]` hype against a gentle director line. **Duration gate:** 45s exceeds single **~19s** `p-video` audio-led cap → split into **three** scene rows (~15s each) or shorten copy; probe with `ffprobe` after TTS. Never plan one 45s embed clip. **Avatar redirect:** on-camera host speaking to lens → `p-video-avatar` `voice_script` + `voice_prompt` — do **not** paste this TTS `prompt` into avatar fields.
AionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
activepieces
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
CopilotKit
The Frontend for Agents & Generative UI. React + Angular
Machine-readable data
The same record, as JSON, for agents and crawlers.
{
"facts": [
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Clawhub",
"href": "https://clawhub.ai/pruna-ai/skills/audio-prompting",
"sourceUrl": "https://clawhub.ai/pruna-ai/skills/audio-prompting",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-11T17:08:17.830Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-pruna-ai-audio-prompting/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-pruna-ai-audio-prompting/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-11T17:08:17.830Z",
"isPublic": true
},
{
"factKey": "traction",
"category": "adoption",
"label": "Adoption signal",
"value": "1K downloads",
"href": "https://clawhub.ai/pruna-ai/audio-prompting",
"sourceUrl": "https://clawhub.ai/pruna-ai/audio-prompting",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-11T17:08:17.830Z",
"isPublic": true
},
{
"factKey": "latest_release",
"category": "release",
"label": "Latest release",
"value": "1.0.14",
"href": "https://clawhub.ai/pruna-ai/audio-prompting",
"sourceUrl": "https://clawhub.ai/pruna-ai/audio-prompting",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-09-29T15:31:35.155Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-pruna-ai-audio-prompting/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-pruna-ai-audio-prompting/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
],
"events": [
{
"eventType": "release",
"title": "Release 1.0.14",
"description": "- Updated version metadata to 1.0.14 in SKILL.md. - Removed the skill-card.md file. - No changes to user-facing instructions, features, or functionality.",
"href": "https://clawhub.ai/pruna-ai/audio-prompting",
"sourceUrl": "https://clawhub.ai/pruna-ai/audio-prompting",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-09-29T15:31:35.155Z",
"isPublic": true
}
]
}Record generated Oct 11, 2026.
