agentCLAWHUBUnverified

audio-prompting

Use when crafting TTS, music, or bed prompts for any generative audio model — director style, song structure, and post-production layering.

OpenClaw

Rank

62

Safety

84

Downloads

1.0k

Updated

Oct 11, 2026

Version

1.0.14

Source

CLAWHUB

About

What it does, and when to use it.

Capability contract not published. No trust telemetry is available yet. 1K downloads reported by the source. Last updated 10/11/2026.

Avoid when

  • Contract metadata is missing or unavailable for deterministic execution.

Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing

Public facts

Every fact links back to the source it came from.

Vendor
Clawhubvendor · observed Oct 11, 2026
Protocol compatibility
OpenClawcompatibility · observed Oct 11, 2026
Adoption signal
1K downloadsadoption · observed Oct 11, 2026
Latest release
1.0.14release · observed Sep 29, 2026
Handshake status
UNKNOWNsecurity

Install and run

Setup complexity: low.

clawhub skill install s173djpkkm1x4yfzg522h5vfp989mrn5:audio-prompting
  1. Install using `clawhub skill install s173djpkkm1x4yfzg522h5vfp989mrn5:audio-prompting` in an isolated environment before connecting it to live workloads.
  2. No published capability contract is available yet, so validate auth and request/response behavior manually.
  3. Review the upstream CLAWHUB listing at https://clawhub.ai/pruna-ai/audio-prompting before using production credentials.

Contract: missing

curl -s "https://www.xpersona.co/api/v1/agents/clawhub-pruna-ai-audio-prompting/snapshot"

Run-check

$0.02 USD

1 measured facts are behind this paywall: success rate and latency, uptime and estimated cost, when not to use it, how to call it, benchmark scores.

Agents pay $0.02 in USDC. A card payment is $0.50, the smallest a card allows.

Documentation

CLAWHUB

146,307 characters of source documentation, loaded on request.

Extracted files

5 files captured from the source.

SKILL.md

---
name: audio-prompting
description: Use when crafting TTS, music, or bed prompts for any generative audio model — director style, song structure, and post-production layering.
license: MIT
metadata:
  version: "1.0.14"
  package: pruna-skills
---

# Audio prompting

Vendor-neutral craft for **speech, music, and beds**. Works with Gemini TTS, ElevenLabs, Music 2.5, Stable Audio, Suno, and similar APIs.

## Install

| Skill | Description | Install |
| --- | --- | --- |
| `audio-prompting` | Use when crafting TTS, music, or bed prompts for any generative audio model — director style, song structure, and post-production layering. | `npx skills add PrunaAI/pruna-skills@audio-prompting -y` |
| `generation-diversity` | Use when writing any generative prompt — ritual seed, explicit structure, scenario axes, and quality gates before paid API calls. | `npx skills add PrunaAI/pruna-skills@generation-diversity -y` |

## When to use

- Director-style TTS prompts and inline performance tags
- Full songs with vocals vs instrumental beds
- Choosing when to embed audio in a video model vs mix in post
- Narration + bed layering pipelines

## Works with

Gemini Flash TTS, ElevenLabs, Music 2.5, Stable Audio, Suno, Udio, and other audio models. Pair with `video-prompting` when uploading VO into a video model.

## When NOT to use

Use a different skill instead:

| Skill | Description | Install |
| --- | --- | --- |
| `music-2.5` | Use when someone wants an original AI song with vocals — sung lyrics, a style prompt track, or source audio for a music video. | `npx skills add PrunaAI/[email protected] -y` |
| `stable-audio-2.5` | Use when someone wants light instrumental background music — an ambient bed under dialogue or underscore for reels and explainers. | `npx skills add PrunaAI/[email protected] -y` |
| `gemini-3.1-flash-tts` | Use when someone needs spoken narration or voiceover — explainer tracks, documentary lines, or voice to pair with generated video. | `npx skills add PrunaAI/[email protected] -y` |
| `video-prompting` | Use when crafting video or motion prompts for any generative model — dramaturgy, camera, physics-safe motion, frame anchors, and clip chaining. | `npx skills add PrunaAI/pruna-skills@video-prompting -y` |
| `image-prompting` | Use when crafting still-image prompts for any generative model — composition, identity sheets, edits, try-on, and photoreal personas. | `npx skills add PrunaAI/pruna-skills@image-prompting -y` |
| `video-editing` | Use when assembling or polishing already-rendered clips with ffmpeg — concat, crossfades, burned captions and subtitles, text/logo overlays, before/after sliders, background music beds, platform export — or when composing a multi-layer HTML combination video with Hyperframes. Not for AI video generation, prompt craft, or model-based video edits. | `npx skills add PrunaAI/pruna-skills@video-editing -y` |

## Guide habit

In the **first reply**, name `` `audio-prompting` `` 

_meta.json

{
  "ownerId": "kn7cagwf7q3t0cxrgteb7xk0bh81j0eb",
  "slug": "audio-prompting",
  "version": "1.0.14",
  "publishedAt": 1790695895155
}

references/audio-post-production.md

# Audio post-production (Pruna + Replicate)

How to choose and **layer** audio when building reels, multi-scene films, and launch videos.

**Prompt craft (how to write):** [tts-style-prompting.md](./tts-style-prompting.md) · [music-and-bed-prompting.md](./music-and-bed-prompting.md). For in-video audio modes and talking-head VO, install and follow `video-prompting`.

**Multi-scene narrated films:** scene anchor triple lives in `video-prompting` — pass TTS to **`p-video`** as `input.audio` with `image` + `last_frame_image`; do not post-mux unless re-render is impossible. Workflow: `narrated-multi-scene`.

**Visual-only transitions (no VO):** scene anchor pair in `video-prompting` — `duration` instead of `audio`. Workflow: `visual-transition-reel`.

## Audio-led `p-video` (required when VO/narration exists)

When narration, TTS, or a timed audio slice is available **before** video render:

1. Upload the audio file to Pruna (`POST /v1/files`) — see `pruna-api`.
2. Pass `urls.get` as **`input.audio`** on **`p-video`** (or **`p-video-avatar`** for human lip-sync).
3. **Omit `duration`** — clip length follows the audio (capped at **20s** on P-API); the model syncs motion to speech.
4. Set **`save_audio`: true** so the full line is embedded in the output clip.
5. **Probe TTS length** before render — per-scene lines should be **≤ ~19s** or the API truncates the tail even when `audio` is set.
6. **Concat** clips in order (narration already on each clip). Optional bed mixed **under** VO in post.

**Never** generate silent `p-video` and ffmpeg-mux narration afterward unless re-render is impossible — post-mux **truncates** lines longer than the video slot (common with Gemini TTS).

**Over 20s?** Shorten scene copy → tighten TTS pace in `style_prompt` → split into two scene rows (each with its own triple). See `narrated-multi-scene` duration gate.

| Need | Approach | Skill |
|------|----------|-------|
| Lip-sync / duration locked to VO | Upload audio → `p-video` with `audio` | `p-video` |
| Documentary / story narrator | Gemini Flash TTS → upload → video | `gemini-3.1-flash-tts` |
| Light instrumental under dialogue | Stable Audio bed under VO | `stable-audio-2.5` |
| Full song with sung vocals | Music 2.5 track | `music-2.5` |
| Speaking on-camera character | Portrait + script / audio | `p-video-avatar` |

**Env:** Pruna calls need `PRUNA_API_KEY`; Replicate audio tools need `REPLICATE_API_TOKEN`. Assembly steps need **`ffmpeg`** / **`ffprobe`**. Credentials: `pruna-api`. Shared ffmpeg recipes (concat, captions, bed mix, export): **`video-editing`**.

## Layering matrix

| Stack | Primary audio | Secondary | Mix notes |
|-------|---------------|-----------|-----------|
| **Silent B-roll** | — | — | Concat video only |
| **Native `p-video` sound** | Model output | — | Keep `save_audio` default; normalize in assembly if scenes differ |
| **Narration only (fallback)** | Gemini TTS | — | Post-mux only when audio-led `p-video` is not suitable — prefer **Pipelin

references/music-and-bed-prompting.md

# Music and bed prompting

Prompt craft for `music-2.5` (songs with vocals) and `stable-audio-2.5` (instrumental beds). Mix/stack: [audio-post-production.md](./audio-post-production.md). In-video sync: install `video-prompting`.

## Music 2.5 (full song)

Stack: **genre + mood + vocal + tempo + instruments + production feel** (≤ ~2000 chars). Pair with a **lyrics** field — verse / chorus / bridge structure for vocal tracks.

```text
Indie pop, uplifting, warm female vocal, 92 BPM, acoustic guitar and mellow synth pads, no harsh distortion
```

**Lyrics:** write singable lines per section (verse, chorus, optional bridge). Lock structure before the paid call; same lyrics + prompt still yield different arrangements — lock seeds only when the user asks.

| Include | Avoid |
|---------|-------|
| Genre, BPM, vocal timbre, key instruments | Vague `epic cinematic masterpiece` |
| Explicit `no harsh distortion` / energy caps when needed | Contradictions (`lo-fi quiet` + `stadium EDM drop`) |

## Stable Audio 2.5 (beds under VO)

Instrumental, understated, mix-friendly:

```text
Instrumental light electronic pop bed, soft groove and mellow synth pads, calm positive tech atmosphere, understated background music, no vocals, 94 BPM
```

Rules:

- Always **`no vocals`** when under narration  
- Keep energy **below** dialogue — assembly mixes ~0.08–0.15 under VO  
- Tag style works well; keep prompts short  

## Which tool?

| Need | Tool |
|------|------|
| Sung song / music video source | Music 2.5 |
| Quiet bed under TTS or avatar | Stable Audio 2.5 |
| Diegetic SFX inside `p-video` | Native `save_audio` / prompt cues — not these models |

## Pre-send

- [ ] Song vs bed chosen deliberately  
- [ ] Bed: no vocals + BPM + understated  
- [ ] Song: genre/mood/vocal/tempo + **lyrics** structure present  
- [ ] Duration matches scene or assembly plan

## Worked examples

### Full song (Music 2.5) — indie pop, remote-work theme

User lock: warm female vocal, ~92 BPM, acoustic + mellow synth, **not** EDM drop.

**Style prompt** (`prompt` field):

```text
Indie pop, warm and hopeful, female vocal, 92 BPM, acoustic guitar and mellow synth pads, intimate bedroom-production feel, no harsh distortion, no stadium drop
```

**Lyrics** (`lyrics` field — verse / chorus / bridge):

```text
[Verse 1]
Coffee rings on the desk again
Window light on a second screen
Slack pings like a metronome
Building something from my home

[Chorus]
We're still here, we're still on
Pixels bridge what miles have drawn
Heart in the work, voice in the song
Remote but never alone

[Verse 2]
Cat walks across the keyboard line
Deadline hums but the team's aligned
Same sky, different time zones
Same goal in our headphones

[Bridge]
When the Wi‑Fi stutters, we don't fold
Call reconnects — the story holds

[Chorus]
We're still here, we're still on
...
```

Confirm lyrics + style before `POST`. For music-video cut points later → `whisperx` after the track exists.

### Instrumental bed (Stable Audio 2.5

references/tts-style-prompting.md

# TTS style prompting (Gemini 3.1 Flash TTS)

Director-style `prompt` craft for `gemini-3.1-flash-tts`. Upload results to Pruna for Mode B in-video audio (install `video-prompting`). Layering: [audio-post-production.md](./audio-post-production.md).

## Align three channels

| Channel | Role |
|---------|------|
| `text` | Spoken words (+ optional inline `[tags]`) |
| `prompt` | Tone, pace, accent, character — max ~4k bytes |
| `[tags]` in text | Momentary direction matching `prompt` |

All three must point the **same** emotional direction.

**Bracket clarity:** `[tags]` live only in this TTS `text` field. Still typography uses double-quoted `"[STRING]"` — see `image-prompting`. Native clip dialogue (`[subject] says "[LINE]"`) is Mode A in `video-prompting` — not this skill.

## Human narrator defaults

```text
Warm storybook narrator, gentle pace, empathetic, no announcer voice.
```

```text
Natural documentary host, measured pacing, clear consonants, calm authority.
```

Avoid: radio-ad hype, “cinematic trailer voice”, reading the product brief into `prompt`.

## Duration gate for `p-video`

Audio-led clips cap at **20s** (keep TTS ≤ **~19s**). If `ffprobe` is long:

1. Shorten `text`  
2. Add pace to `prompt`: `brisk pace, ~2.3 words per second, no filler`  
3. Split into two scene rows  

Never rely on post-mux over silent video.

## Avatar vs TTS

| Path | Fields |
|------|--------|
| Narrator B-roll | Gemini TTS → `p-video` `input.audio` |
| On-camera speaker | `p-video-avatar` `voice_script` + `voice_prompt` — **not** this TTS `prompt` |

Do not paste VO into avatar `voice_prompt`.

## Pre-send

- [ ] `prompt` / `text` / tags aligned  
- [ ] Length probed for `p-video`  
- [ ] Voice + language recorded in manifest for regen consistency

## Worked example — explainer narration (aligned channels)

User lock: documentary explainer, **measured** pace, ~45s script, will feed `p-video` Mode B.

**`prompt`** (director style):

```text
Natural documentary host, measured pacing, clear consonants, calm authority, empathetic, no radio-ad hype
```

**`text`** (spoken words + inline tags):

```text
[warm] Most teams treat diversity as a checkbox.
[pause] But the ritual seed is what breaks repetition before you ever hit generate.
[emphasis] Lock the brief first — then rotate free axes.
[measured] Same subject, fresh camera and light, every panel.
```

**Alignment check:** tags match the calm documentary `prompt` — no `[shout]` hype against a gentle director line.

**Duration gate:** 45s exceeds single **~19s** `p-video` audio-led cap → split into **three** scene rows (~15s each) or shorten copy; probe with `ffprobe` after TTS. Never plan one 45s embed clip.

**Avatar redirect:** on-camera host speaking to lens → `p-video-avatar` `voice_script` + `voice_prompt` — do **not** paste this TTS `prompt` into avatar fields.
Github ReposUpdated 2d agoRank 70

AionUi

Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!

MCPOPENCLAW
Github ReposUpdated 6mo agoRank 70

activepieces

AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents

OPENCLAW
Github ReposUpdated 6mo agoRank 70

cherry-studio

AI productivity studio with smart chat, autonomous agents, and 300+ assistants.

MCPOPENCLAW
Github ReposUpdated 7mo agoRank 70

CopilotKit

The Frontend for Agents & Generative UI. React + Angular

OPENCLAW

Machine-readable data

The same record, as JSON, for agents and crawlers.

{
  "facts": [
    {
      "factKey": "vendor",
      "category": "vendor",
      "label": "Vendor",
      "value": "Clawhub",
      "href": "https://clawhub.ai/pruna-ai/skills/audio-prompting",
      "sourceUrl": "https://clawhub.ai/pruna-ai/skills/audio-prompting",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-11T17:08:17.830Z",
      "isPublic": true
    },
    {
      "factKey": "protocols",
      "category": "compatibility",
      "label": "Protocol compatibility",
      "value": "OpenClaw",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-pruna-ai-audio-prompting/contract",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-pruna-ai-audio-prompting/contract",
      "sourceType": "contract",
      "confidence": "medium",
      "observedAt": "2026-10-11T17:08:17.830Z",
      "isPublic": true
    },
    {
      "factKey": "traction",
      "category": "adoption",
      "label": "Adoption signal",
      "value": "1K downloads",
      "href": "https://clawhub.ai/pruna-ai/audio-prompting",
      "sourceUrl": "https://clawhub.ai/pruna-ai/audio-prompting",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-11T17:08:17.830Z",
      "isPublic": true
    },
    {
      "factKey": "latest_release",
      "category": "release",
      "label": "Latest release",
      "value": "1.0.14",
      "href": "https://clawhub.ai/pruna-ai/audio-prompting",
      "sourceUrl": "https://clawhub.ai/pruna-ai/audio-prompting",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-09-29T15:31:35.155Z",
      "isPublic": true
    },
    {
      "factKey": "handshake_status",
      "category": "security",
      "label": "Handshake status",
      "value": "UNKNOWN",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-pruna-ai-audio-prompting/trust",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-pruna-ai-audio-prompting/trust",
      "sourceType": "trust",
      "confidence": "medium",
      "observedAt": null,
      "isPublic": true
    }
  ],
  "events": [
    {
      "eventType": "release",
      "title": "Release 1.0.14",
      "description": "- Updated version metadata to 1.0.14 in SKILL.md. - Removed the skill-card.md file. - No changes to user-facing instructions, features, or functionality.",
      "href": "https://clawhub.ai/pruna-ai/audio-prompting",
      "sourceUrl": "https://clawhub.ai/pruna-ai/audio-prompting",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-09-29T15:31:35.155Z",
      "isPublic": true
    }
  ]
}

Record generated Oct 11, 2026.

Sponsored

Ads related to audio-prompting and adjacent AI workflows.