{"id":"6760df65-0255-4cba-bf3e-10ec56ad1f66","entityType":"agent","slug":"clawhub-pruna-ai-generation-diversity","name":"generation-diversity","canonicalUrl":"https://www.xpersona.co/agent/clawhub-pruna-ai-generation-diversity","canonicalPath":"/agent/clawhub-pruna-ai-generation-diversity","generatedAt":"2026-10-11T03:55:33.368Z","source":"CLAWHUB","claimStatus":"UNCLAIMED","verificationTier":"NONE","summary":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-11T00:37:51.974Z","emptyReason":null},"description":"Use when writing any generative prompt — ritual seed, explicit structure, scenario axes, and quality gates before paid API calls.","descriptionLabel":"Source description","evidenceSummary":"Capability contract not published. No trust telemetry is available yet. 1.2K downloads reported by the source. Last updated 10/11/2026.","installCommand":"clawhub skill install s173djpkkm1x4yfzg522h5vfp989mrn5:generation-diversity","sourceUrl":"https://clawhub.ai/pruna-ai/generation-diversity","homepage":"https://clawhub.ai/pruna-ai/skills/generation-diversity","primaryLinks":[{"label":"View on ClawHub","url":"https://clawhub.ai/pruna-ai/generation-diversity","kind":"source"},{"label":"Homepage","url":"https://clawhub.ai/pruna-ai/skills/generation-diversity","kind":"homepage"}],"safetyScore":84,"overallRank":62,"popularityScore":62,"trustScore":null,"claimedByName":null,"isOwner":false,"seoDescription":"generation-diversity technical dossier on Xpersona with agent coverage, OPENCLEW support, and live trust metadata."},"coverage":{"evidence":{"source":"public-profile","verified":false,"confidence":"medium","updatedAt":"2026-10-11T00:37:51.974Z","emptyReason":null},"protocols":[{"protocol":"OPENCLEW","label":"OpenClaw","status":"self-declared","notes":"Declared in the public agent profile."}],"capabilities":[],"verifiedCount":0,"selfDeclaredCount":1,"capabilityMatrix":{"rows":[{"key":"OPENCLEW","type":"protocol","support":"unknown","confidenceSource":"profile","notes":"Listed on profile"}],"flattenedTokens":"protocol:OPENCLEW|unknown|profile"}},"adoption":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-11T00:37:51.974Z","emptyReason":null},"stars":null,"forks":null,"downloads":1217,"packageName":null,"latestVersion":"1.0.14","tractionLabel":"1.2K downloads"},"release":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-11T00:37:51.900Z","emptyReason":null},"lastUpdatedAt":"2026-10-11T00:37:51.974Z","lastCrawledAt":"2026-10-11T00:37:51.900Z","lastIndexedAt":null,"nextCrawlAt":"2026-10-12T00:37:51.901Z","lastVerifiedAt":null,"highlights":[{"version":"1.0.14","createdAt":"2026-09-29T15:28:56.996Z","changelog":"generation-diversity v1.0.14 - Added explicit mention of `p-video-2-pro` `mode` (cost/speed/quality) in the clarification intake checklist. - Internal reference and documentation updates. - Removed the now-unnecessary skill-card.md file.","fileCount":9,"zipByteSize":35388},{"version":"1.0.13","createdAt":"2026-09-17T13:49:35.604Z","changelog":"- Add support and documentation for p-video-2-pro in skill and related skills. - Update video skill descriptions to clarify coverage, capabilities, and distinctions among p-video-2-pro, p-video-2, and p-video. - Expand clarification intake checklist to cover additional resolutions (480p/768p vs 720p/1080p). - Remove deprecated skill-card.md file.","fileCount":9,"zipByteSize":35423},{"version":"1.0.12","createdAt":"2026-09-10T13:46:57.119Z","changelog":"- Updated references to support new tools (`p-image-ideogram`, `p-video-2`) alongside existing ones. - Expanded related skills section to include new installation options for advanced image (`p-image-ideogram`) and video (`p-video-2`) generation. - Removed outdated file (`skill-card.md`). - Clarified documentation to better differentiate use cases for each supported tool.","fileCount":9,"zipByteSize":35148},{"version":"1.0.11","createdAt":"2026-09-03T14:03:55.298Z","changelog":"- Updated version to 1.0.11. - Documentation improvements to SKILL.md and reference guides. - Removed unnecessary file: skill-card.md. - No logic or workflow changes; content and usability updates only.","fileCount":9,"zipByteSize":34712},{"version":"1.0.10","createdAt":"2026-08-28T07:50:24.915Z","changelog":"- Bumped skill version to 1.0.10. - Removed the redundant skill-card.md file. - No substantive changes in documentation or guidance.","fileCount":9,"zipByteSize":34327},{"version":"1.0.9","createdAt":"2026-08-04T06:15:41.790Z","changelog":"generation-diversity 1.0.9 - Updated related skill descriptions for p-image to clarify it's for fastest, cheapest photo generation and not controlled photoreal/in-image text. - Removed the outdated skill-card.md file. - Minor clarifications and consistency updates to references and guides. - Version bump to 1.0.9 in metadata.","fileCount":9,"zipByteSize":34331},{"version":"1.0.8","createdAt":"2026-07-28T17:17:55.478Z","changelog":"- Adds a clarification-intake guide to prompt for missing job details, improving workflow intake. - Guide habit now directs agents to ask clarification questions when media source, brand, audio, structure, resolution, or approval are unclear. - Updates the generation workflow to include a clarification-intake step before prompt creation. - Removes redundant skill-card.md file.","fileCount":9,"zipByteSize":34167},{"version":"1.0.7","createdAt":"2026-07-23T12:32:05.188Z","changelog":"- Major update: Expanded skill to cover prompt diversity and quality assurance for all popular generative models, with new guides and stricter workflow checks. - Broadened scope to include image, video, and audio generation workflows—applies to both Pruna and third-party APIs. - Added three new references: quality checklists, still-image prompt flow, and workflow feedback gates. - Updated guidelines: new ritual seed process, scenario axes, and introduction of quality gates before proceeding to paid API steps. - Improved installation instructions and clarified required actions when red flags are present. - Removed legacy docs and outdated references for a streamlined, up-to-date playbook.","fileCount":8,"zipByteSize":30193}]},"execution":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No published capability contract is available yet."},"installCommand":"clawhub skill install s173djpkkm1x4yfzg522h5vfp989mrn5:generation-diversity","setupComplexity":"low","setupSteps":["Install using `clawhub skill install s173djpkkm1x4yfzg522h5vfp989mrn5:generation-diversity` in an isolated environment before connecting it to live workloads.","No published capability contract is available yet, so validate auth and request/response behavior manually.","Review the upstream CLAWHUB listing at https://clawhub.ai/pruna-ai/generation-diversity before using production credentials."],"contract":{"contractStatus":"missing","authModes":[],"requires":[],"forbidden":[],"supportsMcp":false,"supportsA2a":false,"supportsStreaming":false,"inputSchemaRef":null,"outputSchemaRef":null,"dataRegion":null,"contractUpdatedAt":null,"sourceUpdatedAt":null,"freshnessSeconds":null},"invocationGuide":{"preferredApi":{"snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-pruna-ai-generation-diversity/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-pruna-ai-generation-diversity/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-pruna-ai-generation-diversity/trust"},"curlExamples":["curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-pruna-ai-generation-diversity/snapshot\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-pruna-ai-generation-diversity/contract\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-pruna-ai-generation-diversity/trust\""],"jsonRequestTemplate":{"query":"summarize this repo","constraints":{"maxLatencyMs":2000,"protocolPreference":["OPENCLEW"]}},"jsonResponseTemplate":{"ok":true,"result":{"summary":"...","confidence":0.9},"meta":{"source":"CLAWHUB","generatedAt":"2026-10-11T03:55:33.363Z"}},"retryPolicy":{"maxAttempts":3,"backoffMs":[500,1500,3500],"retryableConditions":["HTTP_429","HTTP_503","NETWORK_TIMEOUT"]}},"endpoints":{"dossierUrl":"https://www.xpersona.co/api/v1/agents/clawhub-pruna-ai-generation-diversity/dossier","snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-pruna-ai-generation-diversity/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-pruna-ai-generation-diversity/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-pruna-ai-generation-diversity/trust"}},"reliability":{"evidence":{"source":"runtime-metrics","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No trust, reliability, or runtime telemetry is available."},"trust":{"status":"unavailable","handshakeStatus":"UNKNOWN","verificationFreshnessHours":null,"reputationScore":null,"p95LatencyMs":null,"successRate30d":null,"fallbackRate":null,"attempts30d":null,"trustUpdatedAt":null,"trustConfidence":"unknown","sourceUpdatedAt":null,"freshnessSeconds":null},"decisionGuardrails":{"doNotUseIf":["Contract metadata is missing or unavailable for deterministic execution."],"safeUseWhen":[],"riskFlags":["missing_or_unavailable_contract","trust_data_unavailable","schema_references_missing"],"operationalConfidence":"low"},"executionMetrics":{"observedLatencyMsP50":null,"observedLatencyMsP95":null,"estimatedCostUsd":null,"uptime30d":null,"rateLimitRpm":null,"rateLimitBurst":null,"lastVerifiedAt":null,"verificationSource":null},"runtimeMetrics":{"successRate":null,"avgLatencyMs":null,"avgCostUsd":null,"hallucinationRate":null,"retryRate":null,"disputeRate":null,"p50Latency":null,"p95Latency":null,"lastUpdated":null}},"benchmarks":{"evidence":{"source":"no-benchmark-data","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No benchmark suites or observed failure patterns are available."},"suites":[],"failurePatterns":[]},"artifacts":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-11T00:37:51.974Z","emptyReason":null},"readme":"Skill: generation-diversity\n\nOwner: pruna-ai\n\nSummary: Use when writing any generative prompt — ritual seed, explicit structure, scenario axes, and quality gates before paid API calls.\n\nTags: ai:1.0.14, generative:1.0.14, latest:1.0.14, pruna:1.0.14\n\nVersion history:\n\nv1.0.14 | 2026-09-29T15:28:56.996Z | auto\n\ngeneration-diversity v1.0.14\n\n- Added explicit mention of `p-video-2-pro` `mode` (cost/speed/quality) in the clarification intake checklist.\n- Internal reference and documentation updates.\n- Removed the now-unnecessary skill-card.md file.\n\nv1.0.13 | 2026-09-17T13:49:35.604Z | auto\n\n- Add support and documentation for p-video-2-pro in skill and related skills.\n- Update video skill descriptions to clarify coverage, capabilities, and distinctions among p-video-2-pro, p-video-2, and p-video.\n- Expand clarification intake checklist to cover additional resolutions (480p/768p vs 720p/1080p).\n- Remove deprecated skill-card.md file.\n\nv1.0.12 | 2026-09-10T13:46:57.119Z | auto\n\n- Updated references to support new tools (`p-image-ideogram`, `p-video-2`) alongside existing ones.\n- Expanded related skills section to include new installation options for advanced image (`p-image-ideogram`) and video (`p-video-2`) generation.\n- Removed outdated file (`skill-card.md`).\n- Clarified documentation to better differentiate use cases for each supported tool.\n\nv1.0.11 | 2026-09-03T14:03:55.298Z | auto\n\n- Updated version to 1.0.11.\n- Documentation improvements to SKILL.md and reference guides.\n- Removed unnecessary file: skill-card.md.\n- No logic or workflow changes; content and usability updates only.\n\nv1.0.10 | 2026-08-28T07:50:24.915Z | auto\n\n- Bumped skill version to 1.0.10.\n- Removed the redundant skill-card.md file.\n- No substantive changes in documentation or guidance.\n\nv1.0.9 | 2026-08-04T06:15:41.790Z | auto\n\ngeneration-diversity 1.0.9\n\n- Updated related skill descriptions for p-image to clarify it's for fastest, cheapest photo generation and not controlled photoreal/in-image text.\n- Removed the outdated skill-card.md file.\n- Minor clarifications and consistency updates to references and guides.\n- Version bump to 1.0.9 in metadata.\n\nv1.0.8 | 2026-07-28T17:17:55.478Z | auto\n\n- Adds a clarification-intake guide to prompt for missing job details, improving workflow intake.\n- Guide habit now directs agents to ask clarification questions when media source, brand, audio, structure, resolution, or approval are unclear.\n- Updates the generation workflow to include a clarification-intake step before prompt creation.\n- Removes redundant skill-card.md file.\n\nv1.0.7 | 2026-07-23T12:32:05.188Z | auto\n\n- Major update: Expanded skill to cover prompt diversity and quality assurance for all popular generative models, with new guides and stricter workflow checks.\n- Broadened scope to include image, video, and audio generation workflows—applies to both Pruna and third-party APIs.\n- Added three new references: quality checklists, still-image prompt flow, and workflow feedback gates.\n- Updated guidelines: new ritual seed process, scenario axes, and introduction of quality gates before proceeding to paid API steps.\n- Improved installation instructions and clarified required actions when red flags are present.\n- Removed legacy docs and outdated references for a streamlined, up-to-date playbook.\n\nv1.0.2 | 2026-07-16T13:23:03.930Z | auto\n\n- Bumped version to 1.0.2.\n- Documentation updates in SKILL.md and two reference files.\n- Removed skill-card.md file.\n\nv1.0.1 | 2026-07-14T15:44:46.216Z | auto\n\n- Bumped version metadata to 1.0.1 in SKILL.md.\n- Removed skill-card.md for simplification.\n- No changes to skill logic or usage instructions.\n\nv0.0.1 | 2026-07-14T15:00:48.809Z | auto\n\ngeneration-diversity 0.0.1\n\n- Initial release providing guidelines to increase output variety for Pruna models.\n- Introduces the \"random seed ritual\", explicit prompt construction, and scenario axis rotation.\n- No API integration; designed for manual application before generation tasks.\n- Includes quick flow steps and links to detailed checklists and related skills.\n\nArchive index:\n\nArchive v1.0.14: 9 files, 35388 bytes\n\nFiles: references/clarification-intake.md (7696b), references/generation-diversity.md (51115b), references/generation-quality-checklists.md (8656b), references/still-image-prompt-flow.md (6935b), references/workflow-feedback-gates.md (2141b), skill-card.md (2350b), skill.manifest.json (195b), SKILL.md (6216b), _meta.json (140b)\n\nFile v1.0.14:SKILL.md\n\n---\nname: generation-diversity\ndescription: Use when writing any generative prompt — ritual seed, explicit structure, scenario axes, and quality gates before paid API calls.\nlicense: MIT\nmetadata:\n  version: \"1.0.14\"\n  package: pruna-skills\n---\n\n# Generation diversity\n\nVendor-neutral playbook for **diverse, explicit prompts** and output QA. Apply before every generation on any model (Pruna, Flux, Midjourney, Runway, ElevenLabs, …).\n\n## Install\n\n| Skill | Description | Install |\n| --- | --- | --- |\n| `generation-diversity` | Use when writing any generative prompt — ritual seed, explicit structure, scenario axes, and quality gates before paid API calls. | `npx skills add PrunaAI/pruna-skills@generation-diversity -y` |\n\n## When to use\n\n- Starting a new image, video, or audio generation\n- Outputs feel repetitive or “AI sloppy”\n- Multi-example batches that need cast/setting/camera variety\n- Before advancing a multi-step workflow past a phase gate\n\n## Works with\n\nAny generative model. Pruna tools (`p-image-ideogram`, `p-image`, `p-video-2-pro`, `p-video-2`, `p-video`, …) and third-party APIs alike.\n\n## Guide habit\n\nIn the **first reply**, name `` `generation-diversity` `` in backticks. When the brief leaves media source, brand, audio, structure, resolution, or approval unclear, **[ask before spending](./references/clarification-intake.md)** — every tool and workflow defers here for shared intake topics. For still-image jobs (`p-image-ideogram`, `p-image`, `p-image-edit`), point agents at **[Still-image prompt flow](./references/still-image-prompt-flow.md)** — brief lock → ritual → axes → explicit prompt → fidelity check. **Mood boards:** new ritual per independent panel; user-locked brand hex / subject stays locked on every panel.\n\n## Before generating\n\n0. **[Clarification intake](./references/clarification-intake.md)** — generate vs existing assets, colors, narration/VO, music, captions, aspect/resolution (480p/768p vs 720p/1080p, canvas, MP), `p-video-2-pro` `mode` (cost/speed/quality), structure, approval (unless the user waived or already locked answers).\n1. **[Generation diversity](./references/generation-diversity.md)** — random seed ritual (SSoT), explicit prompt structure, rotate ≥2 scenario axes per session.\n2. **Still images (`p-image` family):** **[still-image-prompt-flow.md](./references/still-image-prompt-flow.md)** — generation flow, edit flow, mood-board rules, hero → edit handoff. Pair with `image-prompting` golden rules and edit craft.\n3. **[Quality checklists](./references/generation-quality-checklists.md)** — open outputs and judge pass/fail before the next paid step.\n4. **Workflows:** [workflow-feedback-gates.md](./references/workflow-feedback-gates.md) — pause at plan / stills / clips before paid video.\n\n## Red flags\n\nStop and ask (or show assets) before the next paid step if any of these are true:\n\n| Red flag | Required action |\n| --- | --- |\n| User says skip review / burn video credits / run everything now | Refuse same-turn plan+video; require **approve plan** (then stills/clips) unless they explicitly ask for automation |\n| Using `--yes-skip-*-gate` without the user requesting automation | Confirm explicitly before bypassing gates |\n| Outputs not opened / checklist not run | Open the file; run the matching quality checklist before the next `POST` |\n| Ritual seed skipped or copied from docs | Fresh random seed ritual first; do **not** pass the ritual string as API `seed` |\n\n## Related skills\n\nInstall related skills when the job needs them:\n\n| Skill | Description | Install |\n| --- | --- | --- |\n| `image-prompting` | Use when crafting still-image prompts for any generative model — composition, identity sheets, edits, try-on, and photoreal personas. | `npx skills add PrunaAI/pruna-skills@image-prompting -y` |\n| `video-prompting` | Use when crafting video or motion prompts for any generative model — dramaturgy, camera, physics-safe motion, frame anchors, and clip chaining. | `npx skills add PrunaAI/pruna-skills@video-prompting -y` |\n| `audio-prompting` | Use when crafting TTS, music, or bed prompts for any generative audio model — director style, song structure, and post-production layering. | `npx skills add PrunaAI/pruna-skills@audio-prompting -y` |\n| `pruna-api` | Use before any Pruna or Replicate HTTP call — credentials, upload/poll/download, parallel batches, and agent safety. | `npx skills add PrunaAI/pruna-skills@pruna-api -y` |\n| `p-image-ideogram` | Use when photo generation needs more control — photoreal results, text in the image, or structured JSON with hex colors and bounding boxes. Simpler photo generation, edits, and video use other skills in the suite. | `npx skills add PrunaAI/pruna-skills@p-image-ideogram -y` |\n| `p-image` | Use when someone explicitly wants the fastest, cheapest photo generation — mood boards, bulk panels, or quick iterations — not when controlled photoreal or in-image text is needed. | `npx skills add PrunaAI/pruna-skills@p-image -y` |\n| `p-image-edit` | Use when someone wants to edit an existing photo — change outfits or backgrounds, compose from reference images, or apply prompt-driven edits. | `npx skills add PrunaAI/pruna-skills@p-image-edit -y` |\n| `p-video-2-pro` | Use when someone wants a cinematic clip from text or start/end frames — product ads, documentary shots, or dialogue with generated audio. Not for 1080p, imported audio tracks, or talking-head-only hosts. | `npx skills add PrunaAI/pruna-skills@p-video-2-pro -y` |\n| `p-video-2` | Use when someone wants a polished short clip from text, images, or imported audio — 1080p B-roll, start/end frame animation, or a motion shot with a mixed track. Not for cinematic generated-audio clips or talking-head-only hosts. | `npx skills add PrunaAI/pruna-skills@p-video-2 -y` |\n| `p-video` | Use when someone wants a simple short clip from text or images — quick B-roll, drafts, or start/end frame animation. Not when the brief needs cinematic generation, highest quality, tight lip-sync, or imported audio at 1080p. | `npx skills add PrunaAI/pruna-skills@p-video -y` |\n\nOr install the full suite once: `npx skills add PrunaAI/pruna-skills@pruna -y`\n\nFile v1.0.14:_meta.json\n\n{\n  \"ownerId\": \"kn7cagwf7q3t0cxrgteb7xk0bh81j0eb\",\n  \"slug\": \"generation-diversity\",\n  \"version\": \"1.0.14\",\n  \"publishedAt\": 1790695736996\n}\n\nFile v1.0.14:references/clarification-intake.md\n\n# Clarification intake — ask before you spend\n\nUse when the user’s request could mean more than one deliverable, more than one media path, or more than one creative default. **Ask in the first reply** (or right after routing), bundle related questions, and **record answers** in the plan or manifest. Do not start paid `POST`s, bulk generation, or long renders until missing decisions are answered or the user explicitly waives them (“use your judgment”, “surprise me”, “just build it”).\n\nWorkflow skills with **`Intake: ask before generating`** tables are authoritative for that deliverable. This doc is the **shared topic list** every Pruna guide and tool should use when the brief is silent.\n\n## How to ask\n\n| Do | Don't |\n| --- | --- |\n| Offer **2–3 concrete options** plus “other” when the choice is structural (acts, route, layout) | Interrogate field-by-field when the user already gave a locked brief |\n| **Group** questions (media + audio in one message when both are open) | Same-turn **plan + paid video** without **approve plan** / gates |\n| Use structured choice UI when the host supports it (many independent decisions) | Invent brand colors, voice, or “generate vs existing” when cost or look changes |\n| Say what you **assume** if they waive, and proceed with a receipt in the summary | Treat inference as confirmation — restate inferred defaults separately |\n\n**Red flags** (must clarify or show plan): skip review, burn credits, automation flags without explicit opt-in → see `generation-diversity` **Red flags** and [workflow-feedback-gates.md](./workflow-feedback-gates.md).\n\n### Example bundles (adapt to the job)\n\nUse when several topics are open at once; trim if the user already locked some:\n\n1. **Deliverable shape:** “Should I **generate** new visuals/audio, **use files you already have**, or **mix** (e.g. your logo + generated B-roll)? Any **aspect ratio** and **resolution** target (480p/768p for cinematic `p-video-2-pro`, 720p vs 1080p for `p-video-2`, 9:16 vs 16:9)?”\n2. **Look and sound:** “**Brand palette** (named kit vs custom hex / reference image)? **Narration or VO** (none, TTS, lip-sync host, your upload)? **Music** (silent, bed under VO, full song)? **Captions** (none, burned after render, in-composition)?”\n3. **Structure and gates:** “Single clip or **multi-act** piece? If multi-act, I can propose **two orderings** — which direction? OK to pause for **approve plan** before paid video, or run end-to-end?”\n\nFor still-only jobs, add **aspect_ratio** and **megapixel / upscale target** when using `p-image-upscale` or print-sized exports.\n\n## Universal topics\n\nAsk when the brief does not already answer these:\n\n| Topic | Clarify |\n| --- | --- |\n| **Media source** | Generate new assets (image/video/audio API) vs **use existing** files/URLs vs **mix** (e.g. user logo + generated B-roll) |\n| **Brand / look** | Named brand kit vs ad-hoc palette (hex or reference), light vs dark, photoreal vs stylized |\n| **Narration / VO** | None · on-screen text only · **TTS/narration track** · **lip-sync host** · user-uploaded VO · music-only |\n| **Music / bed** | Silent · instrumental bed under VO · full song driving cuts · user-supplied track |\n| **Captions** | None · burned after render · embedded in composition · style (promo karaoke vs simple phrase) |\n| **Format** | Aspect ratio, duration target, destination (LinkedIn, TikTok, in-app, internal) |\n| **Resolution / canvas** | **Video:** 480p / 768p (`p-video-2-pro`) vs 720p vs 1080p vs 4K (and whether to upscale after gen). **`p-video-2-pro` recipe:** `mode` `cost` (same quality as `speed`, but cheaper and slower) vs `speed` (default, faster) vs `quality`. **HTML/composition:** export width×height (e.g. 1920×1080, 1080×1920). **Stills:** native aspect vs letterbox; target long edge or MP for upscale jobs |\n| **Frame rate** | 24 / 25 / 30 fps when the deliverable or platform cares (Reels often 30; cinematic often 24) |\n| **Structure** | Single clip vs multi-act / multi-scene; when vague, propose **two act orders** and let the user pick |\n| **Approval** | Phase gates (**approve plan** / stills / clips) vs one-shot automation |\n| **Locale / voice** | Language, accent, voice preset — never silently default from doc examples |\n| **Privacy / upload** | First remote upload in session → `pruna-api` agent-safety acknowledgment |\n\n## By skill type\n\n### Tools (`p-image-ideogram`, `p-image`, `p-video-2-pro`, `p-video-2`, `p-video`, TTS, beds, …)\n\nMinimum before first `POST`:\n\n- What to produce (subject, mandatory copy, refs)\n- **Generate vs upload** for each required input\n- Aspect / resolution / duration caps where the API cares (`aspect_ratio`, `resolution`, seconds, MP target)\n- Point to **generation-diversity** ritual + craft skill for prompt drafting\n\nTool-specific intake lives in each tool’s **Agent habit**; expand with rows from the universal table when silent.\n\n### Craft guides (`image-prompting`, `video-prompting`, `audio-prompting`)\n\nClarify **creative locks** before drafting prompts: identity continuity, camera/motion intent, embed-vs-post audio, lyrics vs instrumental.\n\n### `video-editing` / assembly\n\nConfirm **files on disk**, ffmpeg available, and whether missing pieces should be **generated** (redirect to tools) or **skipped**. For multi-act HTML combos, clarify structure, caption timing source, and bed/narration like the universal table.\n\n### Brand / visual identity\n\nWhen logos, colors, or on-image type might matter, ask before generating:\n\n- **Colors:** reference image, named swatches, or custom hex values?\n- **Logo / wordmark:** file the user will supply, none, or text-only labels in the frame?\n- **Look:** light vs dark, photoreal vs stylized (see universal **Brand / look** row)\n\nDo not invent palette or logo details when the user did not specify them. When they supplied official assets, use those files — do not redraw trademarks from a text prompt alone.\n\n### Workflows\n\nRun the workflow’s **Intake** table in full. Cross-check universal topics; add rows only when the workflow table is thinner than the job needs.\n\n### HyperFrames (optional companion)\n\nFresh creation: **`hyperframes`** runs the intent layer (`intent-interview.md`) → `BRIEF.md` — do not duplicate that interview in Pruna skills.\n\nWhen HyperFrames is used **without** a fresh interview (edit, resume, narrow fix, or Pruna-only assembly):\n\n| Topic | Ask if unclear |\n| --- | --- |\n| Media | Generate new visuals/audio vs use **existing** project or user assets |\n| Design | Existing `frame.md` / design spec vs new palette (offer preset vs custom hex) |\n| Narration | VO yes/no, script source, TTS vs upload vs none |\n| Music | Bed vs beat-driven vs none |\n| Captions | In-composition vs post-render burn (Pruna path: `video-editing`) |\n| **Canvas / export** | 1920×1080 vs 1080×1920 vs square; 720 vs 1080 render; fps if not default |\n\nSee companion **`hyperframes/references/clarification-before-build.md`** when installed (checklist + question bundles for edit/resume paths).\n\n## When you can skip\n\n- User supplied a **complete** manifest, `BRIEF.md`, or explicit “use these files only”\n- **Recipe / remembered defaults** adopted with confirmation (HyperFrames)\n- Single trivial op with one obvious input (e.g. upscale this file to 4 MP)\n- User already answered in the same thread — do not re-ask; restate in the plan\n\n## See also\n\n- [workflow-feedback-gates.md](./workflow-feedback-gates.md)\n- [still-image-prompt-flow.md](./still-image-prompt-flow.md) — brief lock for stills\n- `pruna-api` agent-safety reference\n- `video-editing` — **Structure and creativity** when act order is open\n\nFile v1.0.14:references/generation-diversity.md\n\n# Generation diversity (all models)\n\nOne policy for **every** generative output — images, video, try-on, avatars, replace, animate (Pruna or otherwise). Covers the **random seed ritual**, explicit prompt structure, scenario axis rotation, and **visual variety** ladders.\n\nUse the **full** checklist here for every generation.\n\n## Contents\n\n- [Random seed ritual](#random-seed-ritual-mandatory-before-every-generation)\n- [Three steps (every job)](#three-steps-every-job)\n- [Still-image prompt flow](./still-image-prompt-flow.md) — `p-image-ideogram` / `p-image` / `p-image-edit` agent pipeline (brief lock → ritual → POST)\n- [Explicit prompt structure](#explicit-prompt-structure-required)\n- [Text & typography by model](#text--typography-by-model)\n- [SSoT axis derivation](#ssot-axis-derivation-sum-mod)\n- [Scenario axes](#scenario-axes-rotate-across-outputs)\n- [Render categories](#render-categories)\n- [Crowded scenes](#crowded-scenes-p-image)\n- [Body type spread](#body-type-spread)\n- [Location-matched crowds](#location-matched-crowds)\n- [Group classes](#group-classes--courses)\n- [Framing & camera](#framing--camera)\n- [Scene spice](#scene-spice-when-it-fits)\n- [Photoreal anti-slop](#photoreal-anti-slop-neon--stylized-briefs)\n- [Aspect ratio](#aspect-ratio-multi-example-sets)\n- [By model](#by-model-minimum-diversity)\n- [When not to maximize diversity](#when-not-to-maximize-diversity)\n- [Visual variety](#visual-variety)\n- [Variety checklist](#variety-checklist-before-first-api-call)\n- [Prompt patterns (variety)](#prompt-patterns-variety)\n- [Anti-patterns](#anti-patterns)\n\n## Three steps (every job)\n\n1. **[Random seed ritual](#random-seed-ritual-mandatory-before-every-generation) (SSoT)** — **always first**, before the prompt. Generate a fresh random string, **state it in the turn**, derive axes via [sum-mod](#ssot-axis-derivation-sum-mod). **Do not** pass the ritual string to API `seed`. **One new ritual string per independent generation**; reuse only on same-brief slop retry.\n2. **Write an [explicit prompt](#explicit-prompt-structure-required)** — name specific people, animals, objects, actions, setting, and camera/light. Add text/typography only when the brief needs it — see [text rules by model](#text--typography-by-model).\n3. **Diversify the scenario row** — change at least **two axes** from the previous output in the same session (cast, setting, camera, **`render_category_tag`**, **aspect_ratio**, creatures, props, … — unless user asked for continuity).\n4. **Log** — `ritual_seed`, axes chosen, prediction id (manifest or turn text).\n\n\n\n## Random seed ritual (mandatory before every generation)\n\nThe random seed ritual is a lean [String Seed of Thought](https://pub.sakana.ai/ssot/) (DAG) protocol. **Every** Pruna generation — every prompt, every `POST /v1/predictions`, every scene row — starts here.\n\nThis prevents copy-pasting example strings (`k7Qm2xP9`, `482901`, …) and reduces accidental duplicate outputs across sessions.\n\n### The ritual (do this first)\n\nBefore writing prompts, curl, or runner JSON:\n\n1. **Generate a random string** in-agent (8–16 chars, mixed case + digits).\n2. **Log it** as `ritual_seed` in the manifest / internal plan. Do **not** require a user-visible *\"Ritual seed: …\"* line unless the user asks for transparency.\n3. **Derive prompt choices** from the string — sum char codes, mod N — pick axes from this doc (`aspect_ratio`, `camera_tag`, `render_category_tag`, …).\n4. **Write the prompt** using [explicit prompt structure](#explicit-prompt-structure-required) and derived axes.\n5. **Record** axes chosen and prediction id in the manifest alongside `ritual_seed`.\n\n**Do not pass the ritual string to API `seed`.** API runs without `seed` unless the user explicitly requests reproducibility (`api_seed`).\n\n**Never** proceed to `POST /v1/predictions` without completing steps 1–2 (unless the user supplied an explicit `api_seed` — see below).\n\n### Reuse rules\n\n| Situation | Action |\n|-----------|--------|\n| **New hero / independent still / mood-board panel** | Fresh ritual string |\n| **Same-brief slop retry** | Reuse same `ritual_seed`; note `retry_ritual_seed` in manifest |\n| **Same character arc** | Lock **hero plate URL** + cast descriptor; reuse `ritual_seed` only on same-brief regen |\n| **User says \"lock seed\" / provides integer** | Pass **their** number as `api_seed` → `input.seed`; skip new ritual for that chain |\n\nCharacter continuity = approved plate URL + cast descriptor — **not** the ritual string on the API.\n\n### Ritual anti-patterns\n\n| Wrong | Right |\n|-------|--------|\n| Copy example strings from SKILL.md | Fresh ritual string each independent generation |\n| Pass ritual string as API `seed` | Ritual is planning-only; `api_seed` only when user asks |\n| One ritual string for entire mood board | New ritual per independent **`p-image`** |\n| Skip ritual because API `seed` is optional | Ritual always; API omits `seed` by default |\n\n### Example (internal plan / optional user-visible)\n\nManifest: `\"ritual_seed\": \"k7Qm2xP9\"`. Derived: aspect_ratio 16:9, camera_tag fish-eye, render_category_tag cartoon_anime_fantasy.\nPrompt: Disco ball reflections on an otter DJ scratching vinyl at a packed 1970s roller rink, fish-eye lens, glitter confetti mid-air, funky energy.\n…then curl / runner **without** `\"seed\"` in `input`.\n\n### Manifest snippet\n\n```json\n{\n  \"ritual_seed_policy\": \"ssot_dag_before_every_generation\",\n  \"ritual_seed\": \"k7Qm2xP9\",\n  \"seed_log\": [\n    { \"phase\": \"hero_p_image\", \"ritual_seed\": \"k7Qm2xP9\", \"creature_tag\": \"otter_dj\", \"setting_tag\": \"1970s_roller_rink\", \"prompt_hash\": \"…\" },\n    { \"phase\": \"scene_2_avatar\", \"ritual_seed\": \"k7Qm2xP9\", \"scene_id\": 2 }\n  ]\n}\n```\n\n## Explicit prompt structure (required)\n\n**Vague prompts produce generic AI slop.** After the ritual and axis picks, every still prompt must be **specific and dynamic** — concrete nouns, frozen actions, named places. Prefer playground/creative briefs over marketing abstractions.\n\n**Name at least four of these per prompt (log tags in manifest):**\n\n| Clause | Log as | Agent must specify |\n|--------|--------|-------------------|\n| **People** | `cast_descriptor` | Named role + age band + expression (`fearless grandmother in floral apron`, not `woman`) |\n| **Animals / creatures** | `creature_tag` | Species + attitude (`otter DJ`, `luna moth knight`, `VIP anglerfish`) |\n| **Objects** | `prop_tag` | Concrete props (`vinyl record`, `chrome rocket sled`, `velvet rope`, `tiny boombox`) |\n| **Action** | `action_tag` | Frozen mid-motion verb (`scratching vinyl`, `lassoing runaway taco truck`, `cape mid-swing`) |\n| **Duration** | `duration_tag` | When timing matters (`1970s`, `8PM`, `45-minute spin class`, `Saturday-morning cartoon`) |\n| **Setting** | `setting_tag` | Named place + era + materials (`packed 1970s roller rink`, `abyss-depth jellyfish nightclub`, `Monument Valley dust storm`) |\n| **Text / typography** | `text_spec` | Only when brief needs readable type — exact strings + surface (see [by model](#text--typography-by-model)) |\n| **Camera + light** | `camera_tag`, `lighting_tag` | `fish-eye lens`, `tilt-shift macro`, `teal-magenta cinematic`, `golden hour sparkle` |\n| **Style** | `render_category_tag` | Medium (`cel-shaded anime`, `baroque oil painting`, `ink-wash storybook`, `photoreal documentary`) |\n\n**Template:**\n\n```text\n{people and/or creatures} {action} with/at {specific objects} in {named setting},\n{style or era cues}, {camera_tag}, {lighting_tag}\n```\n\n**Good examples (dynamic / specific):**\n\n```text\nDisco ball reflections on an otter DJ scratching vinyl at a packed 1970s roller rink,\nfish-eye lens, glitter confetti mid-air, funky energy\n```\n\n```text\nBioluminescent jellyfish nightclub at abyss depth, VIP anglerfish in sunglasses at velvet rope,\nteal-magenta cinematic lighting\n```\n\n```text\nCorgi cowboy lassoing a runaway taco truck through Monument Valley dust storm,\npulp western poster energy, dynamic diagonal composition\n```\n\n**Anti-pattern:** `cool cyberpunk portrait, neon vibes` — no subject, no action, no place. **Right:** name who, what they're doing, where, with which props.\n\n## Text & typography by model\n\n**Never use negation to suppress text** — `no text`, `without signs`, `no typography` often **invoke** the thing you are trying to avoid. Describe surfaces positively when you want blank walls (`plain unmarked walls`, `matte unprinted props`).\n\n| Model | Prompt upsampling | Typography in prompt |\n|-------|-------------------|----------------------|\n| **`p-image-ideogram`** | **`thinking: high`** + **`prompt_upsampling: true`** by default; **`false`** for JSON or locked text | **Controlled photo generation** — in-image text, hex/JSON/bbox, high-detail photoreal. Speed path: **`thinking: low`**, **`prompt_upsampling: false`**, nuanced explicit prompt. |\n| **`p-image`** | **No** effective prompt upsampling | **Simple, quick** photo generation from a short prompt. Avoid dense in-image text — route to **`p-image-ideogram`**. Collage triggers still apply: `interactive-explainer` (`flat lay`, `grid`, `collage`, …). |\n\n**`p-image` text hygiene:** prefer scenes without copy. If a screen appears: `monitor soft colorful blur glow only` — not legible UI unless the user explicitly asked for readable text (route to **`p-image-ideogram`** or simplify the brief).\n\n**Channel split (quotes vs tags):**\n\n- **`text_spec` (stills):** `\"[exact string]\"` + `[surface]` + `[placement]` → `image-prompting` §5\n- **Native clip dialogue:** `p-video-2` for quality (or `p-video` for a simpler clip) Mode A only (`[subject] says \"[LINE]\"` + mouth + gesture) → `video-prompting` — not `p-image`\n- **`[tags]`:** Gemini TTS `text` performance only — not still typography, not `p-video` motion prompt → `audio-prompting`\n\n**Collage triggers (photo generation models):** still avoid `flat lay`, `packshot`, `grid`, `collage`, `montage`, `contact sheet`, `split`, `before and after` — use `single frame`, `one camera angle` instead. Full table: `interactive-explainer`.\n\n## SSoT axis derivation (sum-mod)\n\nAfter stating `ritual_seed` (random string), derive prompt choices — sum Unicode/ASCII char codes, mod list length:\n\n```text\nRATIOS = [\"1:1\", \"16:9\", \"9:16\", \"4:3\", \"3:4\", \"3:2\", \"2:3\"]\naspect_ratio  ← RATIOS[ sum(codes(ritual_seed)) % 7 ]\ncamera_tag    ← camera_tags[ sum(codes(ritual_seed[0:4])) % len(camera_tags) ]\nrender_tag    ← render_tags[ sum(codes(ritual_seed[4:8])) % len(render_tags) ]\n```\n\n`camera_tags` and `render_tags` — see [framing & camera](#framing--camera) and [render categories](#render-categories). State derived picks in the turn (*\"Aspect ratio: 16:9, camera: over-shoulder\"*).\n\n**User `api_seed`:** when the user supplies an integer for reproducibility, pass it as `input.seed` — separate from the ritual string.\n\n## Scenario axes (rotate across outputs)\n\n| Axis | Vary with | Applies to |\n|------|-----------|------------|\n| **Cast** | age, ethnicity, gender, archetype, **hairstyle**, **body type** (rotate — see [below](#body-type-spread)), disability aids (wheelchair, cane), visible age band twice in prompt | all person/content gens |\n| **Medium** | `render_category_tag` — rotate across [render categories](#render-categories) | `p-image`, avatar stills |\n| **Setting** | unique `setting_tag` — specific room/street/venue/era, not repeat adjacent rows | stills + video plates |\n| **Camera** | `camera_tag` — rotate across [framing ladder](#framing--camera); never default MC facing lens | stills, `video_prompt` |\n| **Lighting** | `lighting_tag` — golden hour · neon · overcast · practical | stills, video mood |\n| **Motion** | unique `video_prompt` per clip | `p-video-2`, `p-video`, `p-video-avatar`, animate |\n| **Voice** | natural `voice_script`; one `voice` preset per character | avatar, TTS-led video |\n| **Seed** | new ritual string per **independent** job; reuse only on same-brief slop retry | all generation skills |\n| **Aspect ratio** | different `aspect_ratio` per independent still in a batch — see [below](#aspect-ratio-multi-example-sets) | `p-image`, `p-image-edit` |\n| **Crowd density** | layered background population + activity cues — see [below](#crowded-scenes-p-image) | `p-image` plates with busy worlds |\n\nFull style/camera/lighting ladders: [Visual variety](#visual-variety) below. Persona + try-on bar: `image-prompting`.\n\n## Render categories\n\nRotate **`render_category_tag`** (and log it) so diversity batches cover more than photoreal portraits or anime. Category families below mirror arena leaderboards — pick a **different tag per independent output**.\n\n**Random seed ritual still applies** to every generation in [step 1](#three-steps-every-job); categories describe *what* to vary, not *when* to pick `seed`.\n\n### Text-to-image — `p-image-ideogram` / `p-image`\n\nSources: [Arena text-to-image](https://arena.ai/leaderboard/text-to-image) · [AA text-to-image](https://artificialanalysis.ai/image/leaderboard/text-to-image)\n\n**Unified `render_category_tag`** (Arena bucket = tag — pick one per still):\n\n`product_branding_commercial` · `3d_imaging_modeling` · `cartoon_anime_fantasy` · `photoreal_cinematic` · `art` · `portraits` · `nature_environment` · `animals_creature` · `text_rendering`\n\n| Tag | Typical prompt lane |\n|-----|---------------------|\n| `product_branding_commercial` | single product on seamless studio, person + product in named setting, showroom (not `flat lay` / `packshot` words) |\n| `3d_imaging_modeling` | CG film still, clay/stop-motion, rounded 3D forms |\n| `cartoon_anime_fantasy` | cel anime, fantasy character, crowded stylized world |\n| `photoreal_cinematic` | documentary crowd scenes, film-scale wide, urban march |\n| `art` | oil, watercolor, gouache, charcoal, flat vector |\n| `portraits` | single-subject editorial or documentary portrait (crowd optional behind) |\n| `nature_environment` | landscape-wide; subject small in frame |\n| `animals_creature` | named species + handler; crowded market/park when it fits |\n| `text_rendering` | **user-requested only** — otherwise no readable text |\n\nLog `render_category_tag` in manifest. Combine with [crowded scenes](#crowded-scenes-p-image), [body type](#body-type-spread), and [scene spice](#scene-spice-when-it-fits) when the brief allows.\n\n### Image edit — `p-image-edit`\n\nSources: [Arena image edit](https://arena.ai/leaderboard/image-edit) · [AA image editing](https://artificialanalysis.ai/image/leaderboard/editing)\n\nArena modalities: `single_image_edit` · `multi_image_edit`\n\nEdit diversity tags: `background_swap` · `relight` · `wardrobe_on_plate` · `pose_or_angle_delta` · `multi_ref_composite` · `region_inpaint`\n\nVary **instruction** and **what changes** while identity URL stays fixed on character arcs.\n\n### Text-to-video — `p-video-2` (quality) / `p-video` (simpler)\n\nSources: [Arena text-to-video](https://arena.ai/leaderboard/text-to-video) · [AA text-to-video](https://artificialanalysis.ai/video/leaderboard/text-to-video)\n\nMotion/scene tags: `character_performance` · `landscape_broll` · `urban_street` · `product_demo` · `abstract_mood` · `crowd_scene` · `dialogue_beat`\n\nRotate `video_prompt` grammar, start plate world, and `camera_tag` per clip.\n\n### Image-to-video — `p-video-2` (quality) / `p-video` (simpler) (+ plate upload)\n\nSources: [Arena image-to-video](https://arena.ai/leaderboard/image-to-video) · [AA image-to-video](https://artificialanalysis.ai/video/leaderboard/image-to-video)\n\nPlate-driven tags: `animate_hero_still` · `camera_move_on_plate` · `environmental_parallax` · `avatar_lip_sync` · `hands_or_prop_motion`\n\nMatch motion to what the **still** already shows — do not contradict the plate.\n\n### Video edit — `p-video-replace` / `p-video-edit`\n\nSource: [Arena video edit](https://arena.ai/leaderboard/video-edit)\n\nEdit tags: `face_recast` · `wardrobe_swap` · `accessory_swap` · `background_replace` · `object_in_hand_swap` · `style_transfer_on_subject` · `attribute_recolor` · `object_remove` · `environment_restyle` · `relight` · `on_screen_text`\n\n**`p-video-replace`:** character swap from required refs. **`p-video-edit`:** one principal instruction change per run — lock the keep-list; do not invent a new scene.\n\nSame-gender / identity rules for talking-head beats still apply — see [Cast diversity](#cast-diversity) below.\n\n## Crowded scenes (`p-image`)\n\nWhen the brief asks for **busy**, **crowded**, or **lively** worlds — not a lone subject on a blank wall — stack density in the prompt:\n\n1. **Three depth layers** — sharp foreground subject · readable midground faces/hands/props · landmark bokeh (stage, temple, billboards, ferris wheel).\n2. **Named population count** — `hundreds of pedestrians`, `dozens of faces in midground`, `20+ tiny clay figures` (stylized sets need explicit counts; models under-deliver on vague \"busy\").\n3. **Activity verbs** — raised hands, umbrellas open, food steam, confetti, market haggling, commuters pressed shoulder-to-shoulder.\n4. **Shallow DOF + single subject** — `single subject one frame` keeps one identity readable while the crowd stays behind them.\n5. **Age & angle lock** — repeat age band twice (`woman in her late 50s, visibly fifty`) and use [framing & camera](#framing--camera) — models drift younger, center-frame, and front-facing without it.\n\n| Crowd family | Density cues |\n|--------------|--------------|\n| **Urban rush** | crosswalk stripes, wet reflections, umbrellas, billboard bokeh |\n| **Festival / parade** | confetti, raised hands, costume layers, smoke haze |\n| **Market / bazaar** | overflowing stalls, hanging goods, steam, price tags as color blobs |\n| **Transit crush** | strap hangers, door windows, blurred faces pressed together |\n| **Stylized miniature** | counted clay/figurine shoppers (`20+`), cramped aisle, stacked crates |\n| **Institutional / ER** | framed oil portraits on beige walls, triage number board, wall sanitizer, vending machine, scuffed linoleum, TV blur, mixed-age seated patients |\n| **Urban march / protest** | named city, local landmarks, multiracial crowd cues separate from hero — see [location-matched crowds](#location-matched-crowds) |\n| **Group fitness class** | class name + duration, mixed-gender riders, realistic warm studio light — see [group classes](#group-classes--courses) |\n\n**Anti-pattern:** one blurred smear behind a portrait — name **what** the crowd is doing and **where** layers sit. **Institutional** scenes (ER, airport, classroom) need `benches full`, `standing room only`, or `shoulder-to-shoulder` — otherwise models default to a quiet hallway. Name **set dressing** too: framed portraits on walls, triage number board, vending machine glow, scuffed linoleum — generic mint corridors read AI-empty.\n\n## Body type spread\n\nModels default to one “average fitness” body. In diversity batches, **name build on the hero and vary background bodies**:\n\n| Build tag | Prompt cue |\n|-----------|------------|\n| **Plus-size / curvy** | `plus-size`, `curvy build`, `full-figured` |\n| **Athletic / muscular** | `broad shoulders`, `muscular arms`, `athletic build` |\n| **Petite / slim** | `petite frame`, `slim build`, `narrow shoulders` |\n| **Tall / lanky** | `tall and lanky`, `6-foot frame`, `long limbs` |\n| **Stocky / heavyset** | `stocky build`, `heavyset`, `barrel chest` |\n| **Lean wiry** | `lean wiry frame`, `weathered thin face` |\n\n**Rule:** rotate build across independent panels in a session — not every hero “athletic build”. Background crowd should mix ages **and** silhouettes (`elderly thin woman`, `heavyset man`, `pregnant woman seated`, `toddler on lap`).\n\n## Location-matched crowds\n\nWhen the prompt names a **real city or country**, background faces must match that place’s **demographic mix** — not clone the hero’s ethnicity.\n\n| Wrong | Right |\n|-------|--------|\n| South Asian hero + only South Asian protesters in “New York” | Hero is one identity; crowd explicitly `multiracial NYC march — Black, Latino, white, East Asian protesters` |\n| “Dense city march” with no geography | Name city + 3–4 crowd ethnicity cues + local landmarks (yellow cabs, art deco towers, steam vent) |\n| Festival in Lagos with only Nordic faces | Match crowd to `setting_tag` region |\n\n**Prompt pattern:** lock hero cast in sentence 1; sentence 2 lists **four+ distinct background silhouettes** unrelated to hero ethnicity; sentence 3 names **local landmarks** so the plate cannot read as generic stock.\n\n**Applies to:** protests, airports, transit, street markets, sports crowds — any scene where “crowded” implies a real place.\n\n## Group classes & courses\n\nWhen the scene is a **class, workshop, or team activity**, name the **course type** and **who else is in the room** — models default to monochrome crowds (all men, all one age).\n\n| Specify | Example cues |\n|---------|----------------|\n| **Class type** | `45-minute evening spin class`, `beginner yoga flow`, `HIIT bootcamp circuit` |\n| **Room realism** | warm overhead track lights, mirror wall, rubber floor, water bottles, towels — **not** magenta-cyan neon strips unless brief is explicitly nightclub |\n| **Gender mix** | hero is one person; crowd `mixed-gender class — women with ponytails, men with beards, nonbinary cyclist` |\n| **Body + age mix** | plus-size rider, petite woman, athletic man, woman in her 50s — same as [body type spread](#body-type-spread) |\n\n**Lighting rule for fitness:** real boutique studios are **dim warm overhead** or **single spotlight on instructor** — avoid `split gel`, `neon LED strips`, `magenta-cyan` on photoreal gym plates; those read AI-fake.\n\n**Prompt pattern:** `Documentary fitness portrait` + class name + instructor on bike at front + `20+ mixed-gender cyclists` with 3–4 named background silhouettes + realistic room props.\n\n## Framing & camera\n\nModels default to **centered subject, eyes at camera**. In diversity batches, **rotate `camera_tag` and frame placement** every row — log both in manifest.\n\n**Gaze rule:** `glance off-lens`, `profile`, `back to camera`, `looking down at [prop]`, or `watching the crowd` — **not** `facing camera` or `looking at viewer` unless the user asked for a direct-address avatar plate.\n\n**Placement rule:** name where the subject sits in frame — `left third`, `right third`, `lower right corner`, `edge of frame`, `small in environmental wide` — **not** centered mugshot every time.\n\n| `camera_tag` | Prompt cue |\n|--------------|------------|\n| **Overhead / bird's eye** | `overhead aerial view`, `top-down`, `drone shot looking straight down` |\n| **High corner** | `high angle from corner`, `surveillance-style downward angle` |\n| **Worm's eye** | `ground-level worm's eye`, `camera on pavement` |\n| **Crane-down** | `slight high angle crane-down` |\n| **Over-shoulder** | `over-shoulder from behind`, `seen past someone's shoulder` |\n| **Profile / side** | `profile side angle`, `walking across frame` |\n| **From behind** | `back to camera`, `three-quarter from behind` |\n| **Dutch tilt** | `dutch tilt` — tension scenes only |\n| **Through crowd** | `subject visible through gap in crowd`, `foreground heads out of focus` |\n\n**Batch rule:** no two adjacent stills share the same `camera_tag` **and** placement corner (e.g. don't do `left third` twice in a row).\n\nAvatar / lip-sync exception: face must stay readable and mouth visible — use `slight angle from the side` or `three-quarter`, still **off-center** and **off-lens gaze** when not delivering VO to camera.\n\n## Scene spice (when it fits)\n\nDefault plates are person + crowd + place. Add **one or two specific attributes** when the setting naturally supports them — not random clutter on every row.\n\n| Spice type | When to add | Example |\n|------------|-------------|---------|\n| **Animals** | setting implies them | dog park → `golden retriever on leash`; harbor → `seagulls overhead`; rooftop → `pigeons on water tower`; parade → `police horse midground` |\n| **Held / worn props** | role or weather | `red umbrella tucked under arm`, `wire beekeeper smoker`, `chipped ceramic mug`, `sample strawberry basket` |\n| **Micro-detail** | one thumb-stopping oddity | `muddy paw prints on pavement`, `honey jar on crate`, `green parade beads on fence` |\n\nCamera and placement live in [framing & camera](#framing--camera) — not optional spice.\n\n**Rule:** pick **at most two** spice items per prompt. They must answer “what would a photographer notice here?” — not a checklist dump.\n\n**Skip spice when:** product hero, avatar MC talking head, try-on full-body (garment is the focus), or minimal studio brief.\n\n## Photoreal anti-slop (neon / stylized briefs)\n\nStylized settings still need **documentary skin discipline** or outputs go waxy:\n\n- Lead with `documentary portrait, natural skin pores, not CGI, not illustration` even for neon/cyberpunk worlds.\n- Prefer **worn real materials** — matte leather, faded denim, scratched CRT bezels, sticky carpet — over `holographic puffer`, `chrome armor`, `HUD`.\n- Name **gritty location cues** — basement arcade, wet alley, scuffed linoleum — not abstract `neon corridor`.\n- Background crowd faces need **imperfect texture**; blur is fine, plastic skin in midground is not.\n\n## Aspect ratio (multi-example sets)\n\nWhen generating **two or more** stills in one session (playground grid, demo batch, mood board), give each independent output a **different** `aspect_ratio` unless the user locked a format.\n\n**Allowed `p-image` values:** `1:1` · `16:9` · `9:16` · `4:3` · `3:4` · `3:2` · `2:3`\n\n**How to pick:** after the [random seed ritual](#random-seed-ritual-mandatory-before-every-generation), use [sum-mod](#ssot-axis-derivation-sum-mod) on `ritual_seed` — state it in the turn (*\"Aspect ratio: 16:9\"*). Do **not** default every example to `9:16` or `1:1`.\n\n| Ratio | Typical use |\n|-------|-------------|\n| `9:16` | vertical UGC, full-body fashion, avatar talking head |\n| `16:9` | environmental wide, cinematic landscape plate |\n| `3:4` | editorial portrait, try-on full-body |\n| `4:3` | classic portrait, product + person |\n| `1:1` | packshot grid, social tile |\n| `3:2` · `2:3` | magazine / poster crops |\n\nMatch prompt framing to ratio (e.g. `16:9 horizontal wide shot`, `9:16 vertical full body`). **`p-image-try-on`** inherits plate size when `preserve_input_size: true` — diversify person plates first.\n\n**Same character arc:** one ratio for the whole chain unless the user asks for reframes.\n\n## By model (minimum diversity)\n\n| Model | Besides ritual seed, always vary |\n|-------|-----------------------------------|\n| **`p-image-ideogram`** | same axes as `p-image` — photoreal / text / JSON control path |\n| **`p-image`** | cast/creature + objects + action + setting + camera + **`render_category_tag`** + **aspect_ratio**; [explicit structure](#explicit-prompt-structure-required); [text hygiene](#text--typography-by-model) (no upsampling) |\n| **`p-image-edit`** | edit tag + setting/angle delta; same identity URL |\n| **`p-image-try-on`** | person plate world + garment complexity; preserve scene |\n| **`p-image-upscale`** | N/A on prompt — diversify **source** stills |\n| **`p-video-2`** | motion/scene tag + `video_prompt`; differ start plates per scene (quality path) |\n| **`p-video`** | same axes — simpler / quicker clips |\n| **`p-video-avatar`** | `video_prompt` + still world per scene; lock voice per character |\n| **`p-video-animate`** | persona still style/setting per slider ref |\n| **`p-video-replace`** | video-edit tag + full cast spread on showcase reels |\n| **`p-video-edit`** | video-edit tag + one-change lock (keep camera / motion / unmentioned subjects) |\n\n## When **not** to maximize diversity\n\n- **Same character arc** — lock hero plate URL, one `voice`, cast descriptor; vary only setting/angle/motion per scene.\n- **User asked for continuity** — match their cast and approved plates.\n- **Draft → final** — same prompt; change only `draft: false`. Use `api_seed` only if user locked API reproducibility.\n\n## Anti-patterns\n\n| Wrong | Right |\n|-------|--------|\n| Copy doc example ritual strings | [Random seed ritual](#random-seed-ritual-mandatory-before-every-generation) — fresh string each time |\n| Pass ritual string as API `seed` | Ritual is SSoT planning only; `api_seed` when user requests |\n| White wall + MC CU on every demo | Rotate setting + camera + cast |\n| One `video_prompt` for whole reel | Unique motion per scene row |\n| New ritual string mid avatar chain on same brief | Reuse `ritual_seed` until recast or new independent output |\n| Same aspect ratio on every playground example | Rotate `1:1` · `16:9` · `9:16` · `4:3` · `3:4` · `3:2` · `2:3` per [aspect ratio rules](#aspect-ratio-multi-example-sets) |\n| Every hero same athletic body | Rotate [body type spread](#body-type-spread) |\n| Generic hospital hallway | Named ER set dressing + mixed body types in crowd |\n| `holographic` / `chrome` on photoreal cyber scenes | Worn leather, scratched cabinets, documentary skin cues |\n| Monoculture crowd in a named global city | [Location-matched crowds](#location-matched-crowds) — hero ≠ background ethnicity |\n| Magenta-cyan neon on photoreal gym | Warm overhead studio light, mirror wall, real spin bikes |\n| All-male or all-female group class | [Group classes](#group-classes--courses) — mixed-gender background cues |\n| Centered subject every frame | [Framing & camera](#framing--camera) — rotate `camera_tag` + placement |\n| Subject facing camera / at viewer | Off-lens gaze, profile, from behind, or watching crowd |\n| Random animals with no setting reason | Animals only when place implies them |\n| Every stylized panel is anime | Rotate [render categories](#render-categories) — use `cartoon_anime_fantasy` at most once per batch |\n| Vague `cool portrait, neon vibes` | [Explicit structure](#explicit-prompt-structure-required) — named subject, action, objects, setting |\n| `no text` / `without signage` in prompt | Negation invokes text — use [text rules by model](#text--typography-by-model) |\n| Dense typography on **`p-image`** | Drop copy or simplify the brief — `p-image` has no prompt upsampling |\n\n## Visual variety\n\nUse this whenever you plan **`p-image-ideogram`**, **`p-image`**, **`p-image-edit`**, **`p-video-2`**, **`p-video`**, **`p-video-avatar`**, **`p-video-animate`**, **`p-video-replace`**, or **`p-video-edit`** rows. Run the **Variety checklist** at the bottom before the first API call.\n\n### Goal\n\nShowcase and multi-scene work should feel **art-directed**, not like the same talking head in the same office repeated eight times. Deliberately vary:\n\n- **Cast** — gender, age band, ethnicity, persona archetype\n- **Setting** — background / environment (never repeat the same location + framing twice in a row)\n- **Camera** — angle, shot size, movement grammar\n- **Lighting** — time of day, key/fill mood, practical vs cinematic\n- **Visual style** — photoreal, pencil sketch, hand-drawn 2D, cel anime, flat vector, stop-motion clay, CG 3D film, cyberpunk, blockbuster film, editorial, etc.\n- **Render medium** — how the frame is made: `photoreal` · `pencil_sketch` · `hand_drawn_2d` · `cel_anime_2d` · `stop_motion_3d` · `cg_3d_film` (orthogonal to subject family)\n\n**Rule:** Within one **scene row**, lock a local **style bible** so references in that row match (e.g. all three anime refs share the same cel-shaded look). **Across scene rows**, push variety — alternate worlds, angles, and lighting.\n\n## Dynamic prompt stack (eye-catching)\n\nEvery still prompt should feel **art-directed and thumb-stopping**, not generic stock. Build in order:\n\n1. **Style + subject** — who they are + one statement wardrobe piece\n2. **World** — 2–3 concrete environment cues (city bokeh, mirror panels, miniature teal lamp, twin moons)\n3. **Lighting name** — in **`p-image` / reference stills**: bright environment (sunny window, cheerful daylight, golden afternoon). **Avoid** ring light, studio lighting, key/rim/gel light **wording** in still prompts — those belong in plan `lighting_tag` + `video_prompt`, not `p-image`.\n4. **Shot framing** — in still prompts: slight angle from the side, wide shot, slight high angle — **not** “facing camera”, “three-quarter”, or “3/4”. Record angle in plan `camera_tag`.\n5. **`swap_visual_bible`** (plan) — amplify contrast on persona-ladder refs\n\n**Anti-pattern:** flat “neutral wall, soft natural light” on **every row in a scene** — especially UGC/install beats with three grey-wall refs. **Fix:** distinct location family per ref (loft · rooftop · cafe · LED studio) + named gel rim + varied **`camera_tag`** (low angle, side angle, slight high angle).\n\nSee **Prompt patterns** below — flash without text/collage artifacts.\n\n## Creative attractiveness (beyond cast & medium)\n\nSubject diversity is necessary but not sufficient. Thumb-stopping frames also need **color**, **composition**, **texture**, and **motion** variety.\n\n### Color palette ladder\n\nAssign a **`palette_tag`** per scene or ref so sliders do not all read as teal-and-amber:\n\n| Palette | Wardrobe + light pairing |\n|---------|--------------------------|\n| **Warm punch** | coral wall + magenta-cyan LED + gold chain |\n| **Cool contrast** | cobalt hoodie + teal edge light |\n| **Split gel** | rose-gold key + cyan-magenta rim (editorial) |\n| **Monochrome pop** | charcoal + single vivid accent (lime crew, orange sculpture) |\n| **Earth luxe** | walnut desk + copper prop + tungsten accent |\n| **Neon editorial** | violet hair + cherry-blossom bokeh + magenta-teal ambient |\n\n**Rule:** one **dominant accent color** per ref at thumbnail scale — avoid muddy mid-tones everywhere.\n\n### Texture & material\n\nName fabrics and surfaces in prompts — models respond strongly to material words:\n\n- faux fur · holographic puffer · satin wrap · matte clay · crosshatching · glossy chrome armor · walnut grain · matte ceramic\n\nScene 3 already stacks texture beats (fur, holo, pearl/gold); reuse that pattern on wardrobe rows elsewhere.\n\n### Composition & depth\n\n- **Shallow DOF** + gel reflections (single-subject neon boutique) or city window bokeh (studio) — separates subject from background\n- **Foreground anchor** — mug, **closed hardcover notebook**, tumbler at chest gives replace sliders a readable swap target\n- **Single subject one frame** — always; negative space on one side reads cleaner in inset thumbnails\n\n### Age & profession spread\n\nNot every scene needs “tech founder early 30s.” Rotate:\n\n- Gen-Z UGC creator · mid-30s creative director · late-20s advocate · **40s+ expert/trainer** for one VO row\n- Archetypes beyond tech: stylist, chef, fitness creator, museum docent — when the narrative allows\n\n### Camera & motion (reel-level)\n\nCurrent plan anti-pattern to avoid: every scene `medium_cu_dolly_in`. Spread:\n\n| Scene role | Suggested `camera_tag` | `video_prompt` grammar |\n|------------|------------------------|-------------------------|\n| Hook ladder | medium_cu_dolly_in + quarter-orbit | dolly + orbit |\n| UGC install | low_angle_handheld | handheld sway + arc left + push |\n| UGC ref ladder | low angle · side angle · slight high angle | vary per ref within one scene row |\n| Editorial gate | medium_cu_handheld | slow arc right |\n| Desk props | medium_cu_slow_arc | arc + push |\n| CTA | medium_cu_crane_settle | dolly + crane-down |\n\nVary **gaze beats** in `video_prompt` (glance to prop, bookshelf, mirror) — not only straight-to-lens.\n\n### Slider pacing\n\nLong persona ladders (7–9 refs) need tighter **`slider_seconds`** (1.25–1.5) or trim refs — otherwise hook scene dominates reel runtime.\n\n### Quality gates before Phase B\n\n- Ref still readable at **256px wide** (identity + accent color)\n- Adjacent refs differ in **medium + palette + setting**, not just hair color\n- `instruction_prompt` colors/materials **match** reference prompt (lime crew ≠ forest green; copper ≠ silver)\n- Source `video_prompt` props **match** `still_edit` (no mug glance if no mug in plate)\n\n## Cast diversity\n\nPlan a **cast ledger** before generation. For **skills-library / showcase batches** (not single-spokesperson arcs):\n\n| Rule | Guidance |\n|------|----------|\n| **Source host** | **Different person per scene row** — `plate_mode: p-image` + unique `cast_descriptor`. Do not hero-edit one female presenter into every male/advocacy row. |\n| **Reference beats** | Prefer **full recasts** (different ethnicity, age, archetype per ref) over three wardrobe tweaks on one face when proving library range. |\n| **Gender** | Alternate **`persona_gender`** and matching Pruna **`voice`** (`Zephyr (Female)` / `Puck (Male)`) across scenes when lip-sync VO matters. Face-swap refs must stay **same gender** as the source subject on talking-head beats. |\n| **Ethnicity / region** | Name specific, respectful descriptors in prompts (South Asian, East Asian, Black, Latina, Middle Eastern, Nordic, Mediterranean, etc.) — spread representation across the reel, not one token face. |\n| **Age** | Mix early 20s creator energy, mid-30s founder, 40s+ expert — match wardrobe and setting to age. |\n| **Persona archetype** | UGC creator, corporate trainer, fantasy warrior, anime hero, clay character, cyberpunk netrunner, fairy-tale royal, **anthropomorphic otter/fox presenter**, documentary host, gym creator, stylist, etc. |\n| **Subject family** | Photoreal human · fictional character · anthropomorphic (humanoid) · stylized 3D · wardrobe-only · accessories-only · object prop |\n\n**Eye-catching persona ladder (replace hook / animate slider):** one source performance → 5–7 **wildly different** reference stills — e.g. photoreal UGC → premium anime → claymation → cyberpunk → epic film warrior → **anthropomorphic library host** → **fairy-tale 3D royal**. Each ref gets its **own environment, lighting, wardrobe, and subject type**.\n\n**Wardrobe & accessories:** dedicate whole slider steps to **outfit-only** (bolero, vest) and **accessory-only** (scarf, choker, hat, statement earrings) with per-reference `instruction_prompt` naming the slot — same talent, new look, lips unchanged.\n\n## Background & setting ladder\n\nNo two consecutive scene rows should share the **same location type + shot size**. Rotate through distinct worlds:\n\n| Setting family | Example backgrounds |\n|----------------|---------------------|\n| **Domestic / UGC** | bedroom ring light, **creative loft brick**, rooftop dusk, cozy cafe corner, moody LED studio — use **one per ref**, not grey wall ×3 |\n| **Commercial** | boutique, gym floor, outdoor cafe, rooftop at dusk |\n| **Institutional** | classroom whiteboard, news desk, museum gallery |\n| **Fantasy / sci-fi** | stone temple courtyard, alien canyon twin moons, neon arcade corridor, enchanted garden |\n| **Stylized miniature** | clay living room set, diorama street, stop-motion bookshelf nook |\n| **Urban / editorial** | cherry-blossom night street, brutalist plaza, subway platform bokeh |\n\nRecord **`setting_tag`** per scene in the plan (e.g. `\"neon_anime_alley\"`, `\"clay_living_room\"`, `\"temple_courtyard\"`) and verify no duplicate tags in adjacent rows.\n\n## Camera angle & movement ladder\n\nVary **shot size**, **angle**, and **movement** per scene. Never default every row to medium close-up + gentle dolly.\n\n| Angle / size | When to use |\n|--------------|-------------|\n| Extreme close-up (eyes / mouth) | Hook tension, lip-sync proof |\n| Medium close-up (chest-up) | Default VO rows — mouth visible |\n| Medium wide (waist-up) | Wardrobe beats, props in frame |\n| Low angle (heroic) | Game knight, blockbuster reveal |\n| High angle (vulnerable / editorial) | Documentary, stylized anime |\n| Over-shoulder turning in | Explainer, product demo |\n| Profile side angle | Stylized refs when motion allows |\n\n**Movement grammar** (prefix `video_prompt` with continuous camera):\n\n- gentle dolly push-in · slow arc left · subtle handheld sway · orbit quarter-left · crane-down settle · tracking follow (silent B-roll only)\n\n**Anti-pattern:** eight scenes, all `medium close-up, gentle dolly push-in`.\n\n## Lighting ladder\n\nName lighting in every **`p-image` prompt**, **`still_edit`**, and reference still:\n\n| Mood | Prompt cues |\n|------|-------------|\n| Soft overcast documentary | even skin, neutral shadows |\n| Golden hour warm | rim light, amber fill, long shadows |\n| Neon / cyberpunk | magenta-cyan edge light, wet reflections |\n| Anime film dramatic | strong key, colored bounce, neon bokeh |\n| Stop-motion practical | warm desk lamp, miniature set glow |\n| Blockbuster / game cinematic | motivated sun shafts, volumetric haze |\n| Clean educational | bright even key, soft classroom fill |\n| Low-key cinematic | single motivated source, deep background falloff |\n\nAlternate lighting families across scenes — not only \"soft natural light\" on every row.\n\n## Visual style ladder\n\nFor **showcase and multi-scene** work, plan at least **4 distinct visual styles** across the full piece. Pick from (mix and match):\n\n| Style tag | Prompt direction |\n|-----------|------------------|\n| **Photoreal UGC** | smartphone-adjacent, natural skin, real locations |\n| **Photoreal commercial** | crisp product labels, controlled studio or location |\n| **Pencil / charcoal sketch** | crosshatching on cream paper, art-studio daylight, mouth visible |\n| **Hand-drawn 2D animation** | ink outlines, watercolor wash, golden-age animation palette — **single frame** |\n| **Premium anime (2D cel)** | cel-shaded, film-grade compositing, stylized hair/eyes |\n| **Flat vector 2D** | bold shapes, limited palette, motion-graphics friendly |\n| **Disney / Pixar 3D (CG film)** | rounded forms, storybook warmth, enchanted environments |\n| **Claymation / stop-motion 3D** | visible clay texture, miniature sets, practical lighting |\n| **Cyberpunk** | chrome, neon arcade corridor, HUD-free (no readable UI text) |\n| **Blockbuster movie** | anamorphic cues, epic scale, costume drama |\n| **Editorial fashion** | bold wardrobe, shallow DOF, magazine angles |\n| **Documentary** | handheld honesty, available light |\n| **Meme / reaction** | dorm, gaming chair, exaggerated expression (reaction / meme beat pattern) |\n| **Anthropomorphic** | humanoid otter/fox/red panda presenter, expressive face, mouth visible, cozy set |\n\n**Rendering medium ladder (persona hooks):** aim for **5+ mediums** in one slider when showcasing range — e.g. photoreal → pencil sketch → 2D ink frame → cel anime → stop-motion clay → CG 3D royal. Tag optional `render_medium_tag` per ref in plans.\n\n**Animate rows:** generate **one persona still per style** on the same motion template — each still carries its own background, lighting, and rendering style while matching pose/framing to the template.\n\n**Replace rows:** stylized refs (anime, clay, cyberpunk, **anthropomorphic**, **fictional 3D**) work best on **persona-ladder** hooks or **character** beats; use dedicated rows for **wardrobe-only** and **accessory-only** swaps; keep object beats as single props in frame.\n\n## Plan fields (agents & JSON plans)\n\nAdd to scene plans and manifests:\n\n| Field | Purpose |\n|-------|---------|\n| `visual_style_tag` | e.g. `anime_cinematic`, `clay_stop_motion`, `pencil_sketch`, `hand_drawn_2d` |\n| `render_medium_tag` | optional: `photoreal` · `pencil_sketch` · `hand_drawn_2d` · `cel_anime_2d` · `stop_motion_3d` · `cg_3d_film` |\n| `palette_tag` | optional dominant accent pairing: `warm_punch` · `cool_contrast` · `split_gel` · `monochrome_pop` · `neon_editorial` |\n| `setting_tag` | unique environment label per row |\n| `camera_tag` | e.g. `low_angle_mc`, `extreme_cu`, `over_shoulder` |\n| `lighting_tag` | e.g. `golden_hour`, `neon_night`, `practical_clay` |\n| `persona_gender` | `female` / `male` — lock voice + face-swap gender |\n| `cast_descriptor` | one-line identity (ethnicity, age, archetype, **anthropomorphic otter host**, **fictional royal**) |\n| `subject_family` | optional: `photoreal_human` · `fictional_character` · `anthropomorphic` · `wardrobe` · `accessories` · `object` |\n\n**Style bible (project level):** one sentence for **technical** consistency (aspect ratio, photoreal skin when photoreal). **Do not** use the style bible to force every scene into the same look — use **`visual_style_tag`** per row for deliberate variety.\n\n**Never use prompt trigger words** that cause stray text or multi-panel collages — see blocked phrases in **Prompt patterns** below. Prefer positive single-frame wording only.\n\n## Prompt patterns (variety)\n\n### Photoreal recast (replace / avatar)\n\n```text\nDocumentary street portrait, woman mid-30s, South Asian, curly auburn hair, emerald coat,\nlow angle from below, city bokeh background, bright open-sky daylight,\nentire face visible including eyes and mouth, walking stride frozen mid-step, one person one frame.\n```\n\n### UGC install row (source + per-ref worlds)\n\n**Source plate:** creative loft, exposed brick, teal window bokeh, low angle handheld, amber key + magenta rim.\n\n**Ref A — rooftop recast:** low angle chest-up, cobalt hoodie, city lights bokeh, golden hour rim, **closed hardcover notebook** at chest.\n\n**Ref B — cafe recast:** side angle chest-up, orange hoodie, warm wood panels, teal edge light, **closed hardcover notebook** at chest.\n\n**Ref C — wardrobe:** slight high angle, lime crewneck, magenta-cyan LED studio wash, ring light on face, **closed hardcover notebook** at chest.\n\nNever reuse `neutral grey wall` on source and all three refs.\n\n### Pencil sketch persona\n\n```text\nStylized muted-tone woman presenter, soft grey tones, north-facing art studio with soft skylight,\nwide shot slight high angle, mouth open mid-word turning from profile, sole subject one frame.\n```\n\nAvoid in **`p-image` still prompts:** charcoal, pencil, paper, crosshatching, drawing, illustration, **cinematic portrait**, **greyscale cinematic**, **graphite portrait** — they trigger contact-sheet and split-screen collages. **`style_bible`** negations (e.g. no laptops) are fine at plan root — not in per-still positive prompts.\n\n### Hand-drawn 2D animation frame\n\n```text\nHand-painted cel illustration of woman presenter, fluid ink outlines and soft peach watercolor wash background,\nslight angle from the side, mouth visible mid-speech, warm golden afternoon atmosphere, single subject one frame.\n```\n\n### Anime persona (animate / replace ladder)\n\n```text\nPremium anime cinematic young woman hero, cel-shaded film look, violet hair, iridescent jacket,\ncherry-blossom rooftop at dusk with neon color bokeh, low heroic angle from the side,\nmouth visible mid-speech, bright clear evening atmosphere, single character one frame.\n```\n\n### Claymation persona\n\n```text\nStop-motion claymation character woman, visible clay texture, chunky knit scarf, round glasses,\nminiature handmade cozy living room set with tiny lamp and bookshelf, medium close-up,\nmouth sculpted for speech, warm practical stop-motion desk-lamp lighting.\n```\n\n### Disney / fairy-tale 3D\n\n```text\nClassic fairy-tale royal princess cinematic 3D render, elegant ball gown, delicate tiara,\nenchanted castle garden at golden hour with ivy arches and lantern bokeh, medium close-up,\nmouth visible, storybook blockbuster film lighting.\n```\n\n### Cyberpunk\n\n```text\nCyberpunk netrunner woman, chrome undercut, iridescent jacket, neon arcade corridor with magenta-cyan edge light,\nlow angle from below, mouth visible mid-speech, bright electric atmosphere, single subject one frame.\n```\n\n## Ecommerce try-on & photoreal personas\n\nPublic examples across **`p-image`**, **`p-image-try-on`**, and **`p-video-avatar`** should not share one “white studio + plain tee + medium dolly” template.\n\n| Rule | Guidance |\n|------|----------|\n| **Unified bar** | this document · `image-prompting`\n| **Person plate** | Photoreal **`p-image-ideogram`** editorial prompts → slop gate |\n| **Try-on** | Garment tiers + preservation — `image-prompting` |\n| **Avatar motion** | Unique **`video_prompt`** per clip; natural **`voice_script`** |\n| **Cast** | Diversity ledger — gender, age, ethnicity spread |\n| **Playground** | Pin try-on refs on [p-image-try-on](https://replicate.com/prunaai/p-image-try-on); match with diverse [p-video-avatar](https://replicate.com/prunaai/p-video-avatar) examples — @ShinyTaskForce |\n\n**Try-on → avatar handoff:** approved try-on still → optional upscale → **`p-video-avatar`**; lock **`seed`** from person-plate generation.\n\n## Workflow-specific notes\n\n| Workflow | Variety emphasis |\n|----------|-------------------|\n| `p-image-try-on` | Editorial plates + complex garment refs; preservation checklist; diversity across playground set |\n| `p-video-replace` | Scene 1 **persona ladder** + per-scene distinct `still_edit` backgrounds; optional **light bed** after concat |\n| `p-video-edit` | One-change lock + video-edit tag; keep camera/motion; optional refs mapped in `prompt` |\n| `p-video-animate` | 3–4 **style tags** per animate slider row |\n| `avatar-multi-scene` | Every avatar row: new `setting_tag` + `camera_tag` via `p-image-edit` |\n| WORKFLOW-RECIPES | Intake must capture variety plan before recipe execution |\n\n## Variety checklist (before first API call)\n\n- [ ] **Cast:** gender, age, and ethnicity spread across scenes — not one default face\n- [ ] **Settings:** no adjacent duplicate `setting_tag`; at least 5 distinct environments in an 8-scene project\n- [ ] **Palettes:** no three adjacent scenes share the same dominant accent; name gel/light pairs in prompts\n- [ ] **Textures:** wardrobe rows name fabric/material (fur, holo, satin, clay)\n- [ ] **Camera:** at least 3 different `camera_tag` values; no duplicate motion grammar on every `video_prompt`\n- [ ] **Motion props:** `video_prompt` glance targets match objects in `still_edit`\n- [ ] **UGC/install rows:** source + refs use **different** locations and cameras — not neutral grey wall on all stills\n- [ ] **Lighting:** at least 3 different `lighting_tag` values; stylized scenes name their light mood\n- [ ] **Styles:** at least 4 `visual_style_tag` values in a multi-scene piece (mix **photoreal**, **sketch/2D**, **stop-motion 3D**, **CG 3D**)\n- [ ] **Render mediums:** persona ladder includes 5+ distinct mediums when showcasing replace range\n- [ ] **Persona ladder:** hook or animate row includes 6+ visually distinct refs if showcasing range\n- [ ] **Local consistency:** ref\n\nFile v1.0.14:references/generation-quality-checklists.md\n\n# Generation quality checklist hub\n\nUse this as the shared quality gate across models and workflows.\nRun the **Core checklist** for every generation job, then run the model-specific checklist in the guide/workflow skill named below.\n\n## Who applies these checklists?\n\n**The coding agent** — by **opening the real output files** (images, video, or audio) and reviewing them with vision. These checklists are **not** automated test scripts. There is no separate scoring service: the agent reads each item and judges pass or fail from what it sees and hears.\n\nTypical flow:\n\n1. **Generate or download** the asset to a local path (`stills/`, `clips/`, etc.).\n2. **Inspect the file** — view the image, watch the video clip, or listen to narration when the checklist covers audio.\n3. Run the **Core checklist** (below), then the **model-specific checklist** for that job.\n4. **If something fails** — note which items failed, adjust prompt / settings / seed, and regenerate **only that asset** (do not advance to expensive video steps on a bad still).\n5. **If it passes** — show the user the file paths (and previews when helpful). In workflows, still follow [approval gates](#approval-gates-workflows): agent checklist review happens **before** you ask the user to approve stills or clips.\n\nThe user's **approve plan / approve stills / approve clips** gates are separate. Agent checklists catch obvious problems early so the user is not asked to sign off on broken outputs.\n\nMaintenance rule: keep tool/workflow mapping only in this file to avoid link drift. Cross-skill craft lives in the named skills — install them; do not hyperlink into their trees.\n\n## Match map (tool → checklist skill → workflows)\n\n| Tool/model | Guide / workflow | Checklist file (inside that skill) | Common workflows |\n|------------|------------------|------------------------------------|------------------|\n| `p-image-ideogram` | `image-prompting` | `p-image-quality-checklist.md` · persona: `realistic-persona-showcase.md` | `image-to-video`, `narrated-multi-scene`, `avatar-multi-scene` |\n| `p-image` | `image-prompting` | `p-image-quality-checklist.md` · persona: `realistic-persona-showcase.md` | same workflows on cheap / fast drafts |\n| `p-image-edit` | `image-prompting` | `p-image-edit-quality-checklist.md` | `avatar-single-scene`, `avatar-multi-scene` |\n| `p-image-upscale` | `image-prompting` | `p-image-upscale-quality-checklist.md` | `image-to-video`, `narrated-multi-scene` |\n| `p-image-try-on` | `image-prompting` | `p-image-try-on-quality-checklist.md` · persona: `realistic-persona-showcase.md` | `p-image-try-on` |\n| `p-video-2-pro` | `video-prompting` | `p-video-2-pro-quality-checklist.md` | `image-to-video` (visual-only), `visual-transition-reel` |\n| `p-video-2` | `video-prompting` | `p-video-2-quality-checklist.md` | `image-to-video` (audio-led), `narrated-multi-scene`, `interactive-explainer` |\n| `p-video` | `video-prompting` | `p-video-quality-checklist.md` | same workflows on simpler / quicker clips |\n| `p-video-avatar` | `video-prompting` | `p-video-avatar-quality-checklist.md` · persona: `realistic-persona-showcase.md` in `image-prompting` | `avatar-single-scene`, `avatar-multi-scene`, `interactive-explainer` |\n| `p-video-animate` | `video-prompting` | `p-video-animate-quality-checklist.md` | `avatar-multi-scene` |\n| `p-video-replace` | `video-prompting` | `p-video-replace-quality-checklist.md` | `p-video-replace`, `avatar-multi-scene` |\n| `p-video-edit` | `video-prompting` | `p-video-edit-quality-checklist.md` | `p-video-edit` |\n| `music-2.5` + MV assembly | `music-video` | `music-video-quality-checklist.md` | `music-video` |\n\n## Core checklist (all models)\n\n- **[Generation diversity](./generation-diversity.md)** — ritual seed, explicit prompts, visual variety; rotate scenario axes on **every** model (image, video, try-on, avatar, …).\n- Goal and acceptance criteria are explicit (what \"good\" looks like is written down).\n- Input assets are valid and licensed (URL/file reachable, rights cleared).\n- Prompt and settings match the intended output format (`aspect_ratio`, duration, resolution, style lock). **Video default:** `p-video-2-pro` → `768p` 24 fps; `p-video-2` / `p-video` → `720p`, `24` fps unless the brief asks for final `1080p` / `48`.\n- Output contains no accidental watermarks, UI overlays, or stray text unless requested.\n- Brand, legal, and safety constraints are satisfied before handoff.\n- Manifest/log captures model, input fields, prediction id, output URL, and **`ritual_seed`** for traceability.\n\n## Model-specific checklists\n\nInstall the guide/workflow, then open the checklist file inside it:\n\n| Tool | Skill | File |\n|------|-------|------|\n| `p-image-ideogram` | `image-prompting` | `p-image-quality-checklist.md` |\n| `p-image` | `image-prompting` | `p-image-quality-checklist.md` |\n| `p-image-edit` | `image-prompting` | `p-image-edit-quality-checklist.md` |\n| `p-image-upscale` | `image-prompting` | `p-image-upscale-quality-checklist.md` |\n| `p-image-try-on` | `image-prompting` | `p-image-try-on-quality-checklist.md` |\n| `p-video-2-pro` | `video-prompting` | `p-video-2-pro-quality-checklist.md` |\n| `p-video-2` | `video-prompting` | `p-video-2-quality-checklist.md` |\n| `p-video` | `video-prompting` | `p-video-quality-checklist.md` |\n| `p-video-avatar` | `video-prompting` | `p-video-avatar-quality-checklist.md` |\n| `p-video-animate` | `video-prompting` | `p-video-animate-quality-checklist.md` |\n| `p-video-replace` | `video-prompting` | `p-video-replace-quality-checklist.md` |\n| `p-video-edit` | `video-prompting` | `p-video-edit-quality-checklist.md` |\n| music video | `music-video` | `music-video-quality-checklist.md` |\n\n## Visual variety\n\nBefore **any** generation, run [generation-diversity.md](./generation-diversity.md) including the [Variety checklist](./generation-diversity.md#variety-checklist-before-first-api-call). Persona/playground bar: `realistic-persona-showcase.md` in `image-prompting`.\n\n## Approval gates (workflows)\n\nHuman-in-the-loop phases for multi-step workflows. **Video and replace jobs are expensive** — gate on approved stills before any `p-video-*` call. Final audio (bed mix, full-song mux) runs only after clip review.\n\n| Phase | Models | Cost | User interaction |\n|-------|--------|------|------------------|\n| **0 — Plan** | none | free | Present scene table, cast, scripts, `style_bible`; explicit **approve plan / go** |\n| **A — Stills** | `p-image-ideogram`, `p-image`, `p-image-edit` | low | Show hero + start/end plates; run checklists; **approve stills** |\n| **A2 — Audio prep** | Gemini TTS, Music 2.5, WhisperX align | low–medium | **Listen / read** narration or song; fix copy before video |\n| **B — Video** | `p-video-2-pro`, `p-video-2`, `p-video`, `p-video-avatar`, `p-video-animate`, `p-video-replace`, `p-video-edit` | **high** | Only after Phase A approval; **approve clips** before assembly |\n| **C — Assembly** | local ffmpeg concat / slider scripts | free | Review concat (embedded VO); compare MP4s before final mux |\n| **D — Final audio** | Stable Audio bed, bed mix, full-song mux | low | Only after Phase B clip approval |\n\n**Never** run Phase B in the same turn as Phase A without showing stills and waiting for approval. Parallelize within a phase, not across phases.\n\n### Red flags (pause before paid generation)\n\n| Red flag | Required action |\n|----------|-----------------|\n| Plan not presented or no **approve plan** | Phase 0 — scene table + sample prompts |\n| Stills not shown since last prompt edit | Phase A — paths in `stills/`; wait for **approve stills** |\n| Same turn: plan approval + video | Split turns; never batch |\n| **approve clips** missing before concat + bed | Phase C/D only after clip review |\n| Using `--yes-skip-*-gate` without user asking for automation | Confirm explicitly |\n| Regen prompts without deleting stills/clips | Delete targets or `--fresh` / `--regen-*` |\n| Missing `PRUNA_API_KEY` / `REPLICATE_API_TOKEN` | Stop; `pruna-api` |\n\n**When not to stall:** the user already replied **approve plan**, **approve stills**, or **approve clips** for the current phase — proceed with that phase only.\n\nRunner `--approve-*` flags and per-workflow commands: [workflow-feedback-gates.md](./workflow-feedback-gates.md).\n\n## Workflow note\n\nFor multi-scene projects, run these checks per scene and add a final continuity pass\n(style, character identity, voice, and pacing consistency across scenes).\n\n**Narrated cinematic B-roll:** validate scene anchor triple (`video-prompting`) inputs before `p-video-2` — start still, end still, uploaded narration URL per row.\n\nFile v1.0.14:references/still-image-prompt-flow.md\n\n# Still-image prompt flow (`p-image` family)\n\nAgent playbook for **photo generation** (`p-image`, `p-image-ideogram`) and **surgical edit** (`p-image-edit`). Full ritual, axes, and explicit structure live in [generation-diversity.md](./generation-diversity.md). Still craft (golden rules, change/keep, checklists) lives in `image-prompting` — this doc is the **pipeline glue** only.\n\n## When to use\n\n| Job | Tool | Start here |\n|-----|------|------------|\n| New photo generation | `p-image-ideogram` when photo generation needs control (text, JSON, hex/bbox, photoreal detail); else `p-image` | `p-image-ideogram` or `p-image` skill + [Generation flow](#generation-flow-p-image) |\n| New photo generation (cheap / fast / bulk) | `p-image` | [Generation flow](#generation-flow-p-image) |\n| Change existing photo | `p-image-edit` | [Edit flow](#edit-flow-p-image-edit) |\n| Mood board / multi-panel batch | `p-image-ideogram` (×N) unless user asked cheap/fast → `p-image` | [Batch / mood board](#batch--mood-board) |\n| Photo → edit → upscale → video | `p-image-ideogram` or `p-image` → `p-image-edit` → … | [Photo → edit handoff](#photo--edit-handoff) |\n\n**Not this path:** virtual try-on → `p-image-try-on`; sharpen only → `p-image-upscale`.\n\n## Brief lock vs free axes\n\n**Dynamic + faithful:** diversity fills what the brief left open — it never overrides user locks.\n\n| Lock first (user or approved plan) | Free to derive from ritual (when brief silent) |\n|-----------------------------------|------------------------------------------------|\n| Named subject, product, brand, species | Camera tag, lighting nuance, render category |\n| Must-keep props, outfit color, readable copy | Setting texture *when* setting not specified |\n| Continuity plate URL + cast descriptor | Aspect ratio (multi-example sets) |\n| Edit: **change** clause + every **keep** clause | Wording spice inside the change (materials, hex, weather detail) |\n\n**Fidelity check (before pay):** remove the user’s named subject/product/change from the prompt — if the job still “works,” the prompt is wrong. For edits, every stated **keep** must appear in the prompt string.\n\n## Generation flow (`p-image-ideogram` / `p-image`)\n\nPick the model first: **`p-image-ideogram`** when the still needs control (photoreal hero, readable text, JSON/hex/bbox); **`p-image`** for a simple cheap/fast draft.\n\nRun in order every time:\n\n1. **Lock the brief** — list user-required facts (subject, product, format, copy-on-surface if any). Ask if anything is missing.\n2. **Random seed ritual** — fresh string; state it in the turn when drafting; [sum-mod](./generation-diversity.md#ssot-axis-derivation-sum-mod) for free axes. **Do not** pass ritual string as API `seed`.\n3. **Derive axes** — rotate ≥2 free axes vs the previous still in session (`aspect_ratio`, `camera_tag`, `render_category_tag`, setting materials, …). See [by model (p-image)](./generation-diversity.md#by-model-minimum-diversity).\n4. **Draft explicit prompt** — name cast/creature, objects, frozen action, setting, camera/light, style tag. Follow [explicit prompt structure](./generation-diversity.md#explicit-prompt-structure-required) and `image-prompting` golden rules. **`p-image` has no upsampling** — concrete language is the whole craft. **`p-image-ideogram`** defaults to upsampling on (`thinking: high`) unless copy/JSON is locked.\n5. **Fidelity check** — brief locks still present; no copied SKILL curl examples.\n6. **Confirm** — show `prompt` + `aspect_ratio` before `POST` unless user locked wording.\n7. **POST + checklist** — `pruna-api` poll/download; run `image-prompting` **p-image quality checklist** before upscale/video.\n\n### Turn template (draft-only turn)\n\n```text\nTool: `p-image-ideogram` or `p-image`\nRitual seed: <fresh string>\nBrief locks: <user facts>\nFree axes: aspect_ratio <ratio>, camera <tag>, render <tag>, …\nDraft prompt: \"<explicit prompt>\"\nAspect ratio: <ratio>\nFidelity: ✓ subject/product preserved\n→ Confirm to POST (or edit wording)\n```\n\n### Typography\n\nAvoid dense readable type unless the user asked for copy on a surface. When they did, follow [text & typography by model](./generation-diversity.md#text--typography-by-model).\n\n## Edit flow (`p-image-edit`)\n\n1. **Lock change + keeps** — `Change [X]. Keep [Y, Z, …] identical.` from user brief.\n2. **Upload** — `POST /v1/files`; use `urls.get` in `input.images` (1–5 refs). Never invent file URLs.\n3. **Ritual seed** — fresh string before drafting edit wording (variety in *how* you phrase the change, not *what* changes).\n4. **Draft edit prompt** — surgical formula; multi-ref: `face from image 1, outfit from image 2`. See `image-prompting` **p-image-edit-prompting** reference.\n5. **Fidelity check** — change clause matches ask; every keep clause present; no mood-only rewrite.\n6. **Confirm** — `prompt`, refs, `aspect_ratio` (`match_input_image` when appropriate), **`turbo`** on/off (off for hard edits).\n7. **POST + checklist** — `image-prompting` **p-image-edit quality checklist**.\n\n### Turn template (edit draft)\n\n```text\nTool: `p-image-edit`\nRitual seed: <fresh string>\nChange: <user request>\nKeep: <face, pose, outfit, lighting, …>\nDraft prompt: \"Change … Keep … identical.\"\nRefs: image 1 = <role>\nTurbo: off (hard edit) | on (default)\n→ Confirm to upload + POST\n```\n\n## Batch / mood board\n\nIndependent panels (playground grid, demo batch, mood board):\n\n- **New ritual string per panel** — one ritual for the whole board is wrong.\n- **Different `aspect_ratio` per panel** when format not locked — derive from each panel’s ritual ([aspect ratio](./generation-diversity.md#aspect-ratio-multi-example-sets)).\n- **Same character arc** — opposite rule: one approved photo URL, lock cast; vary only setting/angle per scene ([when not to maximize diversity](./generation-diversity.md#when-not-to-maximize-diversity)).\n\n## Photo → edit handoff\n\n| Step | Action |\n|------|--------|\n| Photo approved | Save `urls.get` from `p-image` or `p-image-ideogram` output — lock as the edit source |\n| Surgical tweak | `p-image-edit` from that URL — never run photo generation again for “same person, new background” |\n| Video | Edit from **upscaled** photo when the pipeline requires; upscale again after edit before `p-video*` |\n\nRedirect to `p-image-ideogram` (or `p-image` for a cheap draft) only when the user wants a **new** subject or scene from scratch.\n\n## Anti-patterns\n\n| Wrong | Right |\n|-------|--------|\n| Copy otter/corgi curl examples from SKILL.md | Fresh ritual + brief-faithful prompt |\n| `cool product vibe, neon` | Named product, materials, action, setting, camera |\n| Regen new face for “change background” | `p-image-edit` on locked photo URL |\n| One ritual for 4 mood-board tiles | Ritual per independent still |\n| Edit prompt without keep clauses | Explicit `Keep … identical` for every user lock |\n\nFile v1.0.14:references/workflow-feedback-gates.md\n\n# Workflow feedback gates\n\nAgent-only approval phases for multi-step workflows. Universal QA: [generation-quality-checklists.md](./generation-quality-checklists.md#approval-gates-workflows).\n\n**Agents must pause** at each gate and **ask** when art direction is unclear. Never approve the plan and run paid video in the same turn. There are **no Python runners** — follow each workflow skill’s phase tables (curl + ffmpeg).\n\n## Phase table (all workflows)\n\n| Phase | What to show | Proceed when |\n|-------|--------------|--------------|\n| **0 — Plan** | Scene table, cast, scripts, `style_bible` | **approve plan** |\n| **A — Stills** | Generated images only | **approve stills** |\n| **A2 — Audio** (when used) | TTS / song / align preview | User listens / accepts |\n| **B — Video** | Paid clip outputs | **approve clips** |\n| **C — Assembly** | Concat / bed / final mux | User accepts deliverable |\n\n## Per-workflow gate order\n\n| Workflow | Install | Gates |\n|----------|---------|-------|\n| `image-to-video` | `@image-to-video` | plan → stills → TTS (if triple) → video → bed |\n| `narrated-multi-scene` | `@narrated-multi-scene` | plan → stills → TTS → video → bed |\n| `visual-transition-reel` | `@visual-transition-reel` | plan → stills → video → assemble ± bed |\n| `avatar-single-scene` | `@avatar-single-scene` | plan → still → avatar |\n| `avatar-multi-scene` | `@avatar-multi-scene` | plan → hero+stills → avatar/animate → assembly |\n| `interactive-explainer` | `@interactive-explainer` | plan → stills → TTS → video → assemble ± bed |\n| `music-video` | `@music-video` | plan/lyrics → song → align → stills → video → assemble |\n| `illustrated-story-reel` | `@illustrated-story-reel` | plan → stills → tts or music → [video if p-video] → assemble |\n\n## Agent rules\n\n1. Complete `generation-diversity` ritual before any prompt work.\n2. Install workflow tools via that skill’s **Prerequisites** (or `@pruna` once).\n3. Parallelize independent curl jobs within a phase; never skip stills approval before paid video.\n4. Recipe picker for humans: WORKFLOW-RECIPES.md.\n\nFile v1.0.14:skill-card.md\n\n## Description:\n\nGuides agents to write diverse, explicit generative-media prompts and check outputs before paid API calls.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[pruna-ai](https://clawhub.ai/user/pruna-ai)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nCreators and developers use this skill to plan varied image, video, and audio prompts, clarify creative requirements, and review generated assets before continuing paid workflows.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: Paid media generation may incur unwanted charges or produce outputs that miss the brief.\n\nMitigation: Keep approval gates enabled, review prompts before API calls, and inspect outputs before advancing to paid steps.\n\nRisk: Example gender or voice choices may override a user's explicit preferences.\n\nMitigation: Preserve the user's stated identity and voice requirements when selecting prompt variations.\n\nRisk: Installing related skills introduces publisher and installer supply-chain exposure.\n\nMitigation: Install only when the publisher and skill installer are trusted.\n\n## Reference(s):\n\n- [ClawHub skill release](https://clawhub.ai/pruna-ai/skills/generation-diversity)\n- [Clarification intake](references/clarification-intake.md)\n- [Generation diversity guidance](references/generation-diversity.md)\n- [Still-image prompt flow](references/still-image-prompt-flow.md)\n- [Generation quality checklists](references/generation-quality-checklists.md)\n- [Workflow feedback gates](references/workflow-feedback-gates.md)\n- [String Seed of Thought](https://pub.sakana.ai/ssot/)\n\n## Skill Output:\n\n**Output Type(s):** [Guidance, Text, Markdown, Shell commands]\n\n**Output Format:** [Markdown prompts, planning notes, and review checklists; optional API command drafts]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [Prompts preserve user-specified constraints and include varied scenario choices and approval checkpoints.]\n\n## Skill Version(s):\n\n1.0.14 (source: ClawHub release metadata and skill frontmatter)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nFile v1.0.14:skill.manifest.json\n\n{\n  \"references\": [\n    \"generation-diversity.md\",\n    \"still-image-prompt-flow.md\",\n    \"generation-quality-checklists.md\",\n    \"workflow-feedback-gates.md\",\n    \"clarification-intake.md\"\n  ]\n}\n\nArchive v1.0.13: 9 files, 35423 bytes\n\nFiles: references/clarification-intake.md (7561b), references/generation-diversity.md (51115b), references/generation-quality-checklists.md (8656b), references/still-image-prompt-flow.md (6935b), references/workflow-feedback-gates.md (2141b), skill-card.md (2772b), skill.manifest.json (195b), SKILL.md (6171b), _meta.json (140b)\n\nFile v1.0.13:SKILL.md\n\n---\nname: generation-diversity\ndescription: Use when writing any generative prompt — ritual seed, explicit structure, scenario axes, and quality gates before paid API calls.\nlicense: MIT\nmetadata:\n  version: \"1.0.13\"\n  package: pruna-skills\n---\n\n# Generation diversity\n\nVendor-neutral playbook for **diverse, explicit prompts** and output QA. Apply before every generation on any model (Pruna, Flux, Midjourney, Runway, ElevenLabs, …).\n\n## Install\n\n| Skill | Description | Install |\n| --- | --- | --- |\n| `generation-diversity` | Use when writing any generative prompt — ritual seed, explicit structure, scenario axes, and quality gates before paid API calls. | `npx skills add PrunaAI/pruna-skills@generation-diversity -y` |\n\n## When to use\n\n- Starting a new image, video, or audio generation\n- Outputs feel repetitive or “AI sloppy”\n- Multi-example batches that need cast/setting/camera variety\n- Before advancing a multi-step workflow past a phase gate\n\n## Works with\n\nAny generative model. Pruna tools (`p-image-ideogram`, `p-image`, `p-video-2-pro`, `p-video-2`, `p-video`, …) and third-party APIs alike.\n\n## Guide habit\n\nIn the **first reply**, name `` `generation-diversity` `` in backticks. When the brief leaves media source, brand, audio, structure, resolution, or approval unclear, **[ask before spending](./references/clarification-intake.md)** — every tool and workflow defers here for shared intake topics. For still-image jobs (`p-image-ideogram`, `p-image`, `p-image-edit`), point agents at **[Still-image prompt flow](./references/still-image-prompt-flow.md)** — brief lock → ritual → axes → explicit prompt → fidelity check. **Mood boards:** new ritual per independent panel; user-locked brand hex / subject stays locked on every panel.\n\n## Before generating\n\n0. **[Clarification intake](./references/clarification-intake.md)** — generate vs existing assets, colors, narration/VO, music, captions, aspect/resolution (480p/768p vs 720p/1080p, canvas, MP), structure, approval (unless the user waived or already locked answers).\n1. **[Generation diversity](./references/generation-diversity.md)** — random seed ritual (SSoT), explicit prompt structure, rotate ≥2 scenario axes per session.\n2. **Still images (`p-image` family):** **[still-image-prompt-flow.md](./references/still-image-prompt-flow.md)** — generation flow, edit flow, mood-board rules, hero → edit handoff. Pair with `image-prompting` golden rules and edit craft.\n3. **[Quality checklists](./references/generation-quality-checklists.md)** — open outputs and judge pass/fail before the next paid step.\n4. **Workflows:** [workflow-feedback-gates.md](./references/workflow-feedback-gates.md) — pause at plan / stills / clips before paid video.\n\n## Red flags\n\nStop and ask (or show assets) before the next paid step if any of these are true:\n\n| Red flag | Required action |\n| --- | --- |\n| User says skip review / burn video credits / run everything now | Refuse same-turn plan+video; require **approve plan** (then stills/clips) unless they explicitly ask for automation |\n| Using `--yes-skip-*-gate` without the user requesting automation | Confirm explicitly before bypassing gates |\n| Outputs not opened / checklist not run | Open the file; run the matching quality checklist before the next `POST` |\n| Ritual seed skipped or copied from docs | Fresh random seed ritual first; do **not** pass the ritual string as API `seed` |\n\n## Related skills\n\nInstall related skills when the job needs them:\n\n| Skill | Description | Install |\n| --- | --- | --- |\n| `image-prompting` | Use when crafting still-image prompts for any generative model — composition, identity sheets, edits, try-on, and photoreal personas. | `npx skills add PrunaAI/pruna-skills@image-prompting -y` |\n| `video-prompting` | Use when crafting video or motion prompts for any generative model — dramaturgy, camera, physics-safe motion, frame anchors, and clip chaining. | `npx skills add PrunaAI/pruna-skills@video-prompting -y` |\n| `audio-prompting` | Use when crafting TTS, music, or bed prompts for any generative audio model — director style, song structure, and post-production layering. | `npx skills add PrunaAI/pruna-skills@audio-prompting -y` |\n| `pruna-api` | Use before any Pruna or Replicate HTTP call — credentials, upload/poll/download, parallel batches, and agent safety. | `npx skills add PrunaAI/pruna-skills@pruna-api -y` |\n| `p-image-ideogram` | Use when photo generation needs more control — photoreal results, text in the image, or structured JSON with hex colors and bounding boxes. Simpler photo generation, edits, and video use other skills in the suite. | `npx skills add PrunaAI/pruna-skills@p-image-ideogram -y` |\n| `p-image` | Use when someone explicitly wants the fastest, cheapest photo generation — mood boards, bulk panels, or quick iterations — not when controlled photoreal or in-image text is needed. | `npx skills add PrunaAI/pruna-skills@p-image -y` |\n| `p-image-edit` | Use when someone wants to edit an existing photo — change outfits or backgrounds, compose from reference images, or apply prompt-driven edits. | `npx skills add PrunaAI/pruna-skills@p-image-edit -y` |\n| `p-video-2-pro` | Use when someone wants a cinematic clip from text or start/end frames — product ads, documentary shots, or dialogue with generated audio. Not for 1080p, imported audio tracks, or talking-head-only hosts. | `npx skills add PrunaAI/pruna-skills@p-video-2-pro -y` |\n| `p-video-2` | Use when someone wants a polished short clip from text, images, or imported audio — 1080p B-roll, start/end frame animation, or a motion shot with a mixed track. Not for cinematic generated-audio clips or talking-head-only hosts. | `npx skills add PrunaAI/pruna-skills@p-video-2 -y` |\n| `p-video` | Use when someone wants a simple short clip from text or images — quick B-roll, drafts, or start/end frame animation. Not when the brief needs cinematic generation, highest quality, tight lip-sync, or imported audio at 1080p. | `npx skills add PrunaAI/pruna-skills@p-video -y` |\n\nOr install the full suite once: `npx skills add PrunaAI/pruna-skills@pruna -y`\n\nFile v1.0.13:_meta.json\n\n{\n  \"ownerId\": \"kn7cagwf7q3t0cxrgteb7xk0bh81j0eb\",\n  \"slug\": \"generation-diversity\",\n  \"version\": \"1.0.13\",\n  \"publishedAt\": 1789652975604\n}\n\nFile v1.0.13:references/clarification-intake.md\n\n# Clarification intake — ask before you spend\n\nUse when the user’s request could mean more than one deliverable, more than one media path, or more than one creative default. **Ask in the first reply** (or right after routing), bundle related questions, and **record answers** in the plan or manifest. Do not start paid `POST`s, bulk generation, or long renders until missing decisions are answered or the user explicitly waives them (“use your judgment”, “surprise me”, “just build it”).\n\nWorkflow skills with **`Intake: ask before generating`** tables are authoritative for that deliverable. This doc is the **shared topic list** every Pruna guide and tool should use when the brief is silent.\n\n## How to ask\n\n| Do | Don't |\n| --- | --- |\n| Offer **2–3 concrete options** plus “other” when the choice is structural (acts, route, layout) | Interrogate field-by-field when the user already gave a locked brief |\n| **Group** questions (media + audio in one message when both are open) | Same-turn **plan + paid video** without **approve plan** / gates |\n| Use structured choice UI when the host supports it (many independent decisions) | Invent brand colors, voice, or “generate vs existing” when cost or look changes |\n| Say what you **assume** if they waive, and proceed with a receipt in the summary | Treat inference as confirmation — restate inferred defaults separately |\n\n**Red flags** (must clarify or show plan): skip review, burn credits, automation flags without explicit opt-in → see `generation-diversity` **Red flags** and [workflow-feedback-gates.md](./workflow-feedback-gates.md).\n\n### Example bundles (adapt to the job)\n\nUse when several topics are open at once; trim if the user already locked some:\n\n1. **Deliverable shape:** “Should I **generate** new visuals/audio, **use files you already have**, or **mix** (e.g. your logo + generated B-roll)? Any **aspect ratio** and **resolution** target (480p/768p for cinematic `p-video-2-pro`, 720p vs 1080p for `p-video-2`, 9:16 vs 16:9)?”\n2. **Look and sound:** “**Brand palette** (named kit vs custom hex / reference image)? **Narration or VO** (none, TTS, lip-sync host, your upload)? **Music** (silent, bed under VO, full song)? **Captions** (none, burned after render, in-composition)?”\n3. **Structure and gates:** “Single clip or **multi-act** piece? If multi-act, I can propose **two orderings** — which direction? OK to pause for **approve plan** before paid video, or run end-to-end?”\n\nFor still-only jobs, add **aspect_ratio** and **megapixel / upscale target** when using `p-image-upscale` or print-sized exports.\n\n## Universal topics\n\nAsk when the brief does not already answer these:\n\n| Topic | Clarify |\n| --- | --- |\n| **Media source** | Generate new assets (image/video/audio API) vs **use existing** files/URLs vs **mix** (e.g. user logo + generated B-roll) |\n| **Brand / look** | Named brand kit vs ad-hoc palette (hex or reference), light vs dark, photoreal vs stylized |\n| **Narration / VO** | None · on-screen text only · **TTS/narration track** · **lip-sync host** · user-uploaded VO · music-only |\n| **Music / bed** | Silent · instrumental bed under VO · full song driving cuts · user-supplied track |\n| **Captions** | None · burned after render · embedded in composition · style (promo karaoke vs simple phrase) |\n| **Format** | Aspect ratio, duration target, destination (LinkedIn, TikTok, in-app, internal) |\n| **Resolution / canvas** | **Video:** 480p / 768p (`p-video-2-pro`) vs 720p vs 1080p vs 4K (and whether to upscale after gen). **HTML/composition:** export width×height (e.g. 1920×1080, 1080×1920). **Stills:** native aspect vs letterbox; target long edge or MP for upscale jobs |\n| **Frame rate** | 24 / 25 / 30 fps when the deliverable or platform cares (Reels often 30; cinematic often 24) |\n| **Structure** | Single clip vs multi-act / multi-scene; when vague, propose **two act orders** and let the user pick |\n| **Approval** | Phase gates (**approve plan** / stills / clips) vs one-shot automation |\n| **Locale / voice** | Language, accent, voice preset — never silently default from doc examples |\n| **Privacy / upload** | First remote upload in session → `pruna-api` agent-safety acknowledgment |\n\n## By skill type\n\n### Tools (`p-image-ideogram`, `p-image`, `p-video-2-pro`, `p-video-2`, `p-video`, TTS, beds, …)\n\nMinimum before first `POST`:\n\n- What to produce (subject, mandatory copy, refs)\n- **Generate vs upload** for each required input\n- Aspect / resolution / duration caps where the API cares (`aspect_ratio`, `resolution`, seconds, MP target)\n- Point to **generation-diversity** ritual + craft skill for prompt drafting\n\nTool-specific intake lives in each tool’s **Agent habit**; expand with rows from the universal table when silent.\n\n### Craft guides (`image-prompting`, `video-prompting`, `audio-prompting`)\n\nClarify **creative locks** before drafting prompts: identity continuity, camera/motion intent, embed-vs-post audio, lyrics vs instrumental.\n\n### `video-editing` / assembly\n\nConfirm **files on disk**, ffmpeg available, and whether missing pieces should be **generated** (redirect to tools) or **skipped**. For multi-act HTML combos, clarify structure, caption timing source, and bed/narration like the universal table.\n\n### Brand / visual identity\n\nWhen logos, colors, or on-image type might matter, ask before generating:\n\n- **Colors:** reference image, named swatches, or custom hex values?\n- **Logo / wordmark:** file the user will supply, none, or text-only labels in the frame?\n- **Look:** light vs dark, photoreal vs stylized (see universal **Brand / look** row)\n\nDo not invent palette or logo details when the user did not specify them. When they supplied official assets, use those files — do not redraw trademarks from a text prompt alone.\n\n### Workflows\n\nRun the workflow’s **Intake** table in full. Cross-check universal topics; add rows only when the workflow table is thinner than the job needs.\n\n### HyperFrames (optional companion)\n\nFresh creation: **`hyperframes`** runs the intent layer (`intent-interview.md`) → `BRIEF.md` — do not duplicate that interview in Pruna skills.\n\nWhen HyperFrames is used **without** a fresh interview (edit, resume, narrow fix, or Pruna-only assembly):\n\n| Topic | Ask if unclear |\n| --- | --- |\n| Media | Generate new visuals/audio vs use **existing** project or user assets |\n| Design | Existing `frame.md` / design spec vs new palette (offer preset vs custom hex) |\n| Narration | VO yes/no, script source, TTS vs upload vs none |\n| Music | Bed vs beat-driven vs none |\n| Captions | In-composition vs post-render burn (Pruna path: `video-editing`) |\n| **Canvas / export** | 1920×1080 vs 1080×1920 vs square; 720 vs 1080 render; fps if not default |\n\nSee companion **`hyperframes/references/clarification-before-build.md`** when installed (checklist + question bundles for edit/resume paths).\n\n## When you can skip\n\n- User supplied a **complete** manifest, `BRIEF.md`, or explicit “use these files only”\n- **Recipe / remembered defaults** adopted with confirmation (HyperFrames)\n- Single trivial op with one obvious input (e.g. upscale this file to 4 MP)\n- User already answered in the same thread — do not re-ask; restate in the plan\n\n## See also\n\n- [workflow-feedback-gates.md](./workflow-feedback-gates.md)\n- [still-image-prompt-flow.md](./still-image-prompt-flow.md) — brief lock for stills\n- `pruna-api` agent-safety reference\n- `video-editing` — **Structure and creativity** when act order is open\n\nFile v1.0.13:references/generation-diversity.md\n\n# Generation diversity (all models)\n\nOne policy for **every** generative output — images, video, try-on, avatars, replace, animate (Pruna or otherwise). Covers the **random seed ritual**, explicit prompt structure, scenario axis rotation, and **visual variety** ladders.\n\nUse the **full** checklist here for every generation.\n\n## Contents\n\n- [Random seed ritual](#random-seed-ritual-mandatory-before-every-generation)\n- [Three steps (every job)](#three-steps-every-job)\n- [Still-image prompt flow](./still-image-prompt-flow.md) — `p-image-ideogram` / `p-image` / `p-image-edit` agent pipeline (brief lock → ritual → POST)\n- [Explicit prompt structure](#explicit-prompt-structure-required)\n- [Text & typography by model](#text--typography-by-model)\n- [SSoT axis derivation](#ssot-axis-derivation-sum-mod)\n- [Scenario axes](#scenario-axes-rotate-across-outputs)\n- [Render categories](#render-categories)\n- [Crowded scenes](#crowded-scenes-p-image)\n- [Body type spread](#body-type-spread)\n- [Location-matched crowds](#location-matched-crowds)\n- [Group classes](#group-classes--courses)\n- [Framing & camera](#framing--camera)\n- [Scene spice](#scene-spice-when-it-fits)\n- [Photoreal anti-slop](#photoreal-anti-slop-neon--stylized-briefs)\n- [Aspect ratio](#aspect-ratio-multi-example-sets)\n- [By model](#by-model-minimum-diversity)\n- [When not to maximize diversity](#when-not-to-maximize-diversity)\n- [Visual variety](#visual-variety)\n- [Variety checklist](#variety-checklist-before-first-api-call)\n- [Prompt patterns (variety)](#prompt-patterns-variety)\n- [Anti-patterns](#anti-patterns)\n\n## Three steps (every job)\n\n1. **[Random seed ritual](#random-seed-ritual-mandatory-before-every-generation) (SSoT)** — **always first**, before the prompt. Generate a fresh random string, **state it in the turn**, derive axes via [sum-mod](#ssot-axis-derivation-sum-mod). **Do not** pass the ritual string to API `seed`. **One new ritual string per independent generation**; reuse only on same-brief slop retry.\n2. **Write an [explicit prompt](#explicit-prompt-structure-required)** — name specific people, animals, objects, actions, setting, and camera/light. Add text/typography only when the brief needs it — see [text rules by model](#text--typography-by-model).\n3. **Diversify the scenario row** — change at least **two axes** from the previous output in the same session (cast, setting, camera, **`render_category_tag`**, **aspect_ratio**, creatures, props, … — unless user asked for continuity).\n4. **Log** — `ritual_seed`, axes chosen, prediction id (manifest or turn text).\n\n\n\n## Random seed ritual (mandatory before every generation)\n\nThe random seed ritual is a lean [String Seed of Thought](https://pub.sakana.ai/ssot/) (DAG) protocol. **Every** Pruna generation — every prompt, every `POST /v1/predictions`, every scene row — starts here.\n\nThis prevents copy-pasting example strings (`k7Qm2xP9`, `482901`, …) and reduces accidental duplicate outputs across sessions.\n\n### The ritual (do this first)\n\nBefore writing prompts, curl, or runner JSON:\n\n1. **Generate a random string** in-agent (8–16 chars, mixed case + digits).\n2. **Log it** as `ritual_seed` in the manifest / internal plan. Do **not** require a user-visible *\"Ritual seed: …\"* line unless the user asks for transparency.\n3. **Derive prompt choices** from the string — sum char codes, mod N — pick axes from this doc (`aspect_ratio`, `camera_tag`, `render_category_tag`, …).\n4. **Write the prompt** using [explicit prompt structure](#explicit-prompt-structure-required) and derived axes.\n5. **Record** axes chosen and prediction id in the manifest alongside `ritual_seed`.\n\n**Do not pass the ritual string to API `seed`.** API runs without `seed` unless the user explicitly requests reproducibility (`api_seed`).\n\n**Never** proceed to `POST /v1/predictions` without completing steps 1–2 (unless the user supplied an explicit `api_seed` — see below).\n\n### Reuse rules\n\n| Situation | Action |\n|-----------|--------|\n| **New hero / independent still / mood-board panel** | Fresh ritual string |\n| **Same-brief slop retry** | Reuse same `ritual_seed`; note `retry_ritual_seed` in manifest |\n| **Same character arc** | Lock **hero plate URL** + cast descriptor; reuse `ritual_seed` only on same-brief regen |\n| **User says \"lock seed\" / provides integer** | Pass **their** number as `api_seed` → `input.seed`; skip new ritual for that chain |\n\nCharacter continuity = approved plate URL + cast descriptor — **not** the ritual string on the API.\n\n### Ritual anti-patterns\n\n| Wrong | Right |\n|-------|--------|\n| Copy example strings from SKILL.md | Fresh ritual string each independent generation |\n| Pass ritual string as API `seed` | Ritual is planning-only; `api_seed` only when user asks |\n| One ritual string for entire mood board | New ritual per independent **`p-image`** |\n| Skip ritual because API `seed` is optional | Ritual always; API omits `seed` by default |\n\n### Example (internal plan / optional user-visible)\n\nManifest: `\"ritual_seed\": \"k7Qm2xP9\"`. Derived: aspect_ratio 16:9, camera_tag fish-eye, render_category_tag cartoon_anime_fantasy.\nPrompt: Disco ball reflections on an otter DJ scratching vinyl at a packed 1970s roller rink, fish-eye lens, glitter confetti mid-air, funky energy.\n…then curl / runner **without** `\"seed\"` in `input`.\n\n### Manifest snippet\n\n```json\n{\n  \"ritual_seed_policy\": \"ssot_dag_before_every_generation\",\n  \"ritual_seed\": \"k7Qm2xP9\",\n  \"seed_log\": [\n    { \"phase\": \"hero_p_image\", \"ritual_seed\": \"k7Qm2xP9\", \"creature_tag\": \"otter_dj\", \"setting_tag\": \"1970s_roller_rink\", \"prompt_hash\": \"…\" },\n    { \"phase\": \"scene_2_avatar\", \"ritual_seed\": \"k7Qm2xP9\", \"scene_id\": 2 }\n  ]\n}\n```\n\n## Explicit prompt structure (required)\n\n**Vague prompts produce generic AI slop.** After the ritual and axis picks, every still prompt must be **specific and dynamic** — concrete nouns, frozen actions, named places. Prefer playground/creative briefs over marketing abstractions.\n\n**Name at least four of these per prompt (log tags in manifest):**\n\n| Clause | Log as | Agent must specify |\n|--------|--------|-------------------|\n| **People** | `cast_descriptor` | Named role + age band + expression (`fearless grandmother in floral apron`, not `woman`) |\n| **Animals / creatures** | `creature_tag` | Species + attitude (`otter DJ`, `luna moth knight`, `VIP anglerfish`) |\n| **Objects** | `prop_tag` | Concrete props (`vinyl record`, `chrome rocket sled`, `velvet rope`, `tiny boombox`) |\n| **Action** | `action_tag` | Frozen mid-motion verb (`scratching vinyl`, `lassoing runaway taco truck`, `cape mid-swing`) |\n| **Duration** | `duration_tag` | When timing matters (`1970s`, `8PM`, `45-minute spin class`, `Saturday-morning cartoon`) |\n| **Setting** | `setting_tag` | Named place + era + materials (`packed 1970s roller rink`, `abyss-depth jellyfish nightclub`, `Monument Valley dust storm`) |\n| **Text / typography** | `text_spec` | Only when brief needs readable type — exact strings + surface (see [by model](#text--typography-by-model)) |\n| **Camera + light** | `camera_tag`, `lighting_tag` | `fish-eye lens`, `tilt-shift macro`, `teal-magenta cinematic`, `golden hour sparkle` |\n| **Style** | `render_category_tag` | Medium (`cel-shaded anime`, `baroque oil painting`, `ink-wash storybook`, `photoreal documentary`) |\n\n**Template:**\n\n```text\n{people and/or creatures} {action} with/at {specific objects} in {named setting},\n{style or era cues}, {camera_tag}, {lighting_tag}\n```\n\n**Good examples (dynamic / specific):**\n\n```text\nDisco ball reflections on an otter DJ scratching vinyl at a packed 1970s roller rink,\nfish-eye lens, glitter confetti mid-air, funky energy\n```\n\n```text\nBioluminescent jellyfish nightclub at abyss depth, VIP anglerfish in sunglasses at velvet rope,\nteal-magenta cinematic lighting\n```\n\n```text\nCorgi cowboy lassoing a runaway taco truck through Monument Valley dust storm,\npulp western poster energy, dynamic diagonal composition\n```\n\n**Anti-pattern:** `cool cyberpunk portrait, neon vibes` — no subject, no action, no place. **Right:** name who, what they're doing, where, with which props.\n\n## Text & typography by model\n\n**Never use negation to suppress text** — `no text`, `without signs`, `no typography` often **invoke** the thing you are trying to avoid. Describe surfaces positively when you want blank walls (`plain unmarked walls`, `matte unprinted props`).\n\n| Model | Prompt upsampling | Typography in prompt |\n|-------|-------------------|----------------------|\n| **`p-image-ideogram`** | **`thinking: high`** + **`prompt_upsampling: true`** by default; **`false`** for JSON or locked text | **Controlled photo generation** — in-image text, hex/JSON/bbox, high-detail photoreal. Speed path: **`thinking: low`**, **`prompt_upsampling: false`**, nuanced explicit prompt. |\n| **`p-image`** | **No** effective prompt upsampling | **Simple, quick** photo generation from a short prompt. Avoid dense in-image text — route to **`p-image-ideogram`**. Collage triggers still apply: `interactive-explainer` (`flat lay`, `grid`, `collage`, …). |\n\n**`p-image` text hygiene:** prefer scenes without copy. If a screen appears: `monitor soft colorful blur glow only` — not legible UI unless the user explicitly asked for readable text (route to **`p-image-ideogram`** or simplify the brief).\n\n**Channel split (quotes vs tags):**\n\n- **`text_spec` (stills):** `\"[exact string]\"` + `[surface]` + `[placement]` → `image-prompting` §5\n- **Native clip dialogue:** `p-video-2` for quality (or `p-video` for a simpler clip) Mode A only (`[subject] says \"[LINE]\"` + mouth + gesture) → `video-prompting` — not `p-image`\n- **`[tags]`:** Gemini TTS `text` performance only — not still typography, not `p-video` motion prompt → `audio-prompting`\n\n**Collage triggers (photo generation models):** still avoid `flat lay`, `packshot`, `grid`, `collage`, `montage`, `contact sheet`, `split`, `before and after` — use `single frame`, `one camera angle` instead. Full table: `interactive-explainer`.\n\n## SSoT axis derivation (sum-mod)\n\nAfter stating `ritual_seed` (random string), derive prompt choices — sum Unicode/ASCII char codes, mod list length:\n\n```text\nRATIOS = [\"1:1\", \"16:9\", \"9:16\", \"4:3\", \"3:4\", \"3:2\", \"2:3\"]\naspect_ratio  ← RATIOS[ sum(codes(ritual_seed)) % 7 ]\ncamera_tag    ← camera_tags[ sum(codes(ritual_seed[0:4])) % len(camera_tags) ]\nrender_tag    ← render_tags[ sum(codes(ritual_seed[4:8])) % len(render_tags) ]\n```\n\n`camera_tags` and `render_tags` — see [framing & camera](#framing--camera) and [render categories](#render-categories). State derived picks in the turn (*\"Aspect ratio: 16:9, camera: over-shoulder\"*).\n\n**User `api_seed`:** when the user supplies an integer for reproducibility, pass it as `input.seed` — separate from the ritual string.\n\n## Scenario axes (rotate across outputs)\n\n| Axis | Vary with | Applies to |\n|------|-----------|------------|\n| **Cast** | age, ethnicity, gender, archetype, **hairstyle**, **body type** (rotate — see [below](#body-type-spread)), disability aids (wheelchair, cane), visible age band twice in prompt | all person/content gens |\n| **Medium** | `render_category_tag` — rotate across [render categories](#render-categories) | `p-image`, avatar stills |\n| **Setting** | unique `setting_tag` — specific room/street/venue/era, not repeat adjacent rows | stills + video plates |\n| **Camera** | `camera_tag` — rotate across [framing ladder](#framing--camera); never default MC facing lens | stills, `video_prompt` |\n| **Lighting** | `lighting_tag` — golden hour · neon · overcast · practical | stills, video mood |\n| **Motion** | unique `video_prompt` per clip | `p-video-2`, `p-video`, `p-video-avatar`, animate |\n| **Voice** | natural `voice_script`; one `voice` preset per character | avatar, TTS-led video |\n| **Seed** | new ritual string per **independent** job; reuse only on same-brief slop retry | all generation skills |\n| **Aspect ratio** | different `aspect_ratio` per independent still in a batch — see [below](#aspect-ratio-multi-example-sets) | `p-image`, `p-image-edit` |\n| **Crowd density** | layered background population + activity cues — see [below](#crowded-scenes-p-image) | `p-image` plates with busy worlds |\n\nFull style/camera/lighting ladders: [Visual variety](#visual-variety) below. Persona + try-on bar: `image-prompting`.\n\n## Render categories\n\nRotate **`render_category_tag`** (and log it) so diversity batches cover more than photoreal portraits or anime. Category families below mirror arena leaderboards — pick a **different tag per independent output**.\n\n**Random seed ritual still applies** to every generation in [step 1](#three-steps-every-job); categories describe *what* to vary, not *when* to pick `seed`.\n\n### Text-to-image — `p-image-ideogram` / `p-image`\n\nSources: [Arena text-to-image](https://arena.ai/leaderboard/text-to-image) · [AA text-to-image](https://artificialanalysis.ai/image/leaderboard/text-to-image)\n\n**Unified `render_category_tag`** (Arena bucket = tag — pick one per still):\n\n`product_branding_commercial` · `3d_imaging_modeling` · `cartoon_anime_fantasy` · `photoreal_cinematic` · `art` · `portraits` · `nature_environment` · `animals_creature` · `text_rendering`\n\n| Tag | Typical prompt lane |\n|-----|---------------------|\n| `product_branding_commercial` | single product on seamless studio, person + product in named setting, showroom (not `flat lay` / `packshot` words) |\n| `3d_imaging_modeling` | CG film still, clay/stop-motion, rounded 3D forms |\n| `cartoon_anime_fantasy` | cel anime, fantasy character, crowded stylized world |\n| `photoreal_cinematic` | documentary crowd scenes, film-scale wide, urban march |\n| `art` | oil, watercolor, gouache, charcoal, flat vector |\n| `portraits` | single-subject editorial or documentary portrait (crowd optional behind) |\n| `nature_environment` | landscape-wide; subject small in frame |\n| `animals_creature` | named species + handler; crowded market/park when it fits |\n| `text_rendering` | **user-requested only** — otherwise no readable text |\n\nLog `render_category_tag` in manifest. Combine with [crowded scenes](#crowded-scenes-p-image), [body type](#body-type-spread), and [scene spice](#scene-spice-when-it-fits) when the brief allows.\n\n### Image edit — `p-image-edit`\n\nSources: [Arena image edit](https://arena.ai/leaderboard/image-edit) · [AA image editing](https://artificialanalysis.ai/image/leaderboard/editing)\n\nArena modalities: `single_image_edit` · `multi_image_edit`\n\nEdit diversity tags: `background_swap` · `relight` · `wardrobe_on_plate` · `pose_or_angle_delta` · `multi_ref_composite` · `region_inpaint`\n\nVary **instruction** and **what changes** while identity URL stays fixed on character arcs.\n\n### Text-to-video — `p-video-2` (quality) / `p-video` (simpler)\n\nSources: [Arena text-to-video](https://arena.ai/leaderboard/text-to-video) · [AA text-to-video](https://artificialanalysis.ai/video/leaderboard/text-to-video)\n\nMotion/scene tags: `character_performance` · `landscape_broll` · `urban_street` · `product_demo` · `abstract_mood` · `crowd_scene` · `dialogue_beat`\n\nRotate `video_prompt` grammar, start plate world, and `camera_tag` per clip.\n\n### Image-to-video — `p-video-2` (quality) / `p-video` (simpler) (+ plate upload)\n\nSources: [Arena image-to-video](https://arena.ai/leaderboard/image-to-video) · [AA image-to-video](https://artificialanalysis.ai/video/leaderboard/image-to-video)\n\nPlate-driven tags: `animate_hero_still` · `camera_move_on_plate` · `environmental_parallax` · `avatar_lip_sync` · `hands_or_prop_motion`\n\nMatch motion to what the **still** already shows — do not contradict the plate.\n\n### Video edit — `p-video-replace` / `p-video-edit`\n\nSource: [Arena video edit](https://arena.ai/leaderboard/video-edit)\n\nEdit tags: `face_recast` · `wardrobe_swap` · `accessory_swap` · `background_replace` · `object_in_hand_swap` · `style_transfer_on_subject` · `attribute_recolor` · `object_remove` · `environment_restyle` · `relight` · `on_screen_text`\n\n**`p-video-replace`:** character swap from required refs. **`p-video-edit`:** one principal instruction change per run — lock the keep-list; do not invent a new scene.\n\nSame-gender / identity rules for talking-head beats still apply — see [Cast diversity](#cast-diversity) below.\n\n## Crowded scenes (`p-image`)\n\nWhen the brief asks for **busy**, **crowded**, or **lively** worlds — not a lone subject on a blank wall — stack density in the prompt:\n\n1. **Three depth layers** — sharp foreground subject · readable midground faces/hands/props · landmark bokeh (stage, temple, billboards, ferris wheel).\n2. **Named population count** — `hundreds of pedestrians`, `dozens of faces in midground`, `20+ tiny clay figures` (stylized sets need explicit counts; models under-deliver on vague \"busy\").\n3. **Activity verbs** — raised hands, umbrellas open, food steam, confetti, market haggling, commuters pressed shoulder-to-shoulder.\n4. **Shallow DOF + single subject** — `single subject one frame` keeps one identity readable while the crowd stays behind them.\n5. **Age & angle lock** — repeat age band twice (`woman in her late 50s, visibly fifty`) and use [framing & camera](#framing--camera) — models drift younger, center-frame, and front-facing without it.\n\n| Crowd family | Density cues |\n|--------------|--------------|\n| **Urban rush** | crosswalk stripes, wet reflections, umbrellas, billboard bokeh |\n| **Festival / parade** | confetti, raised hands, costume layers, smoke haze |\n| **Market / bazaar** | overflowing stalls, hanging goods, steam, price tags as color blobs |\n| **Transit crush** | strap hangers, door windows, blurred faces pressed together |\n| **Stylized miniature** | counted clay/figurine shoppers (`20+`), cramped aisle, stacked crates |\n| **Institutional / ER** | framed oil portraits on beige walls, triage number board, wall sanitizer, vending machine, scuffed linoleum, TV blur, mixed-age seated patients |\n| **Urban march / protest** | named city, local landmarks, multiracial crowd cues separate from hero — see [location-matched crowds](#location-matched-crowds) |\n| **Group fitness class** | class name + duration, mixed-gender riders, realistic warm studio light — see [group classes](#group-classes--courses) |\n\n**Anti-pattern:** one blurred smear behind a portrait — name **what** the crowd is doing and **where** layers sit. **Institutional** scenes (ER, airport, classroom) need `benches full`, `standing room only`, or `shoulder-to-shoulder` — otherwise models default to a quiet hallway. Name **set dressing** too: framed portraits on walls, triage number board, vending machine glow, scuffed linoleum — generic mint corridors read AI-empty.\n\n## Body type spread\n\nModels default to one “average fitness” body. In diversity batches, **name build on the hero and vary background bodies**:\n\n| Build tag | Prompt cue |\n|-----------|------------|\n| **Plus-size / curvy** | `plus-size`, `curvy build`, `full-figured` |\n| **Athletic / muscular** | `broad shoulders`, `muscular arms`, `athletic build` |\n| **Petite / slim** | `petite frame`, `slim build`, `narrow shoulders` |\n| **Tall / lanky** | `tall and lanky`, `6-foot frame`, `long limbs` |\n| **Stocky / heavyset** | `stocky build`, `heavyset`, `barrel chest` |\n| **Lean wiry** | `lean wiry frame`, `weathered thin face` |\n\n**Rule:** rotate build across independent panels in a session — not every hero “athletic build”. Background crowd should mix ages **and** silhouettes (`elderly thin woman`, `heavyset man`, `pregnant woman seated`, `toddler on lap`).\n\n## Location-matched crowds\n\nWhen the prompt names a **real city or country**, background faces must match that place’s **demographic mix** — not clone the hero’s ethnicity.\n\n| Wrong | Right |\n|-------|--------|\n| South Asian hero + only South Asian protesters in “New York” | Hero is one identity; crowd explicitly `multiracial NYC march — Black, Latino, white, East Asian protesters` |\n| “Dense city march” with no geography | Name city + 3–4 crowd ethnicity cues + local landmarks (yellow cabs, art deco towers, steam vent) |\n| Festival in Lagos with only Nordic faces | Match crowd to `setting_tag` region |\n\n**Prompt pattern:** lock hero cast in sentence 1; sentence 2 lists **four+ distinct background silhouettes** unrelated to hero ethnicity; sentence 3 names **local landmarks** so the plate cannot read as generic stock.\n\n**Applies to:** protests, airports, transit, street markets, sports crowds — any scene where “crowded” implies a real place.\n\n## Group classes & courses\n\nWhen the scene is a **class, workshop, or team activity**, name the **course type** and **who else is in the room** — models default to monochrome crowds (all men, all one age).\n\n| Specify | Example cues |\n|---------|----------------|\n| **Class type** | `45-minute evening spin class`, `beginner yoga flow`, `HIIT bootcamp circuit` |\n| **Room realism** | warm overhead track lights, mirror wall, rubber floor, water bottles, towels — **not** magenta-cyan neon strips unless brief is explicitly nightclub |\n| **Gender mix** | hero is one person; crowd `mixed-gender class — women with ponytails, men with beards, nonbinary cyclist` |\n| **Body + age mix** | plus-size rider, petite woman, athletic man, woman in her 50s — same as [body type spread](#body-type-spread) |\n\n**Lighting rule for fitness:** real boutique studios are **dim warm overhead** or **single spotlight on instructor** — avoid `split gel`, `neon LED strips`, `magenta-cyan` on photoreal gym plates; those read AI-fake.\n\n**Prompt pattern:** `Documentary fitness portrait` + class name + instructor on bike at front + `20+ mixed-gender cyclists` with 3–4 named background silhouettes + realistic room props.\n\n## Framing & camera\n\nModels default to **centered subject, eyes at camera**. In diversity batches, **rotate `camera_tag` and frame placement** every row — log both in manifest.\n\n**Gaze rule:** `glance off-lens`, `profile`, `back to camera`, `looking down at [prop]`, or `watching the crowd` — **not** `facing camera` or `looking at viewer` unless the user asked for a direct-address avatar plate.\n\n**Placement rule:** name where the subject sits in frame — `left third`, `right third`, `lower right corner`, `edge of frame`, `small in environmental wide` — **not** centered mugshot every time.\n\n| `camera_tag` | Prompt cue |\n|--------------|------------|\n| **Overhead / bird's eye** | `overhead aerial view`, `top-down`, `drone shot looking straight down` |\n| **High corner** | `high angle from corner`, `surveillance-style downward angle` |\n| **Worm's eye** | `ground-level worm's eye`, `camera on pavement` |\n| **Crane-down** | `slight high angle crane-down` |\n| **Over-shoulder** | `over-shoulder from behind`, `seen past someone's shoulder` |\n| **Profile / side** | `profile side angle`, `walking across frame` |\n| **From behind** | `back to camera`, `three-quarter from behind` |\n| **Dutch tilt** | `dutch tilt` — tension scenes only |\n| **Through crowd** | `subject visible through gap in crowd`, `foreground heads out of focus` |\n\n**Batch rule:** no two adjacent stills share the same `camera_tag` **and** placement corner (e.g. don't do `left third` twice in a row).\n\nAvatar / lip-sync exception: face must stay readable and mouth visible — use `slight angle from the side` or `three-quarter`, still **off-center** and **off-lens gaze** when not delivering VO to camera.\n\n## Scene spice (when it fits)\n\nDefault plates are person + crowd + place. Add **one or two specific attributes** when the setting naturally supports them — not random clutter on every row.\n\n| Spice type | When to add | Example |\n|------------|-------------|---------|\n| **Animals** | setting implies them | dog park → `golden retriever on leash`; harbor → `seagulls overhead`; rooftop → `pigeons on water tower`; parade → `police horse midground` |\n| **Held / worn props** | role or weather | `red umbrella tucked under arm`, `wire beekeeper smoker`, `chipped ceramic mug`, `sample strawberry basket` |\n| **Micro-detail** | one thumb-stopping oddity | `muddy paw prints on pavement`, `honey jar on crate`, `green parade beads on fence` |\n\nCamera and placement live in [framing & camera](#framing--camera) — not optional spice.\n\n**Rule:** pick **at most two** spice items per prompt. They must answer “what would a photographer notice here?” — not a checklist dump.\n\n**Skip spice when:** product hero, avatar MC talking head, try-on full-body (garment is the focus), or minimal studio brief.\n\n## Photoreal anti-slop (neon / stylized briefs)\n\nStylized settings still need **documentary skin discipline** or outputs go waxy:\n\n- Lead with `documentary portrait, natural skin pores, not CGI, not illustration` even for neon/cyberpunk worlds.\n- Prefer **worn real materials** — matte leather, faded denim, scratched CRT bezels, sticky carpet — over `holographic puffer`, `chrome armor`, `HUD`.\n- Name **gritty location cues** — basement arcade, wet alley, scuffed linoleum — not abstract `neon corridor`.\n- Background crowd faces need **imperfect texture**; blur is fine, plastic skin in midground is not.\n\n## Aspect ratio (multi-example sets)\n\nWhen generating **two or more** stills in one session (playground grid, demo batch, mood board), give each independent output a **different** `aspect_ratio` unless the user locked a format.\n\n**Allowed `p-image` values:** `1:1` · `16:9` · `9:16` · `4:3` · `3:4` · `3:2` · `2:3`\n\n**How to pick:** after the [random seed ritual](#random-seed-ritual-mandatory-before-every-generation), use [sum-mod](#ssot-axis-derivation-sum-mod) on `ritual_seed` — state it in the turn (*\"Aspect ratio: 16:9\"*). Do **not** default every example to `9:16` or `1:1`.\n\n| Ratio | Typical use |\n|-------|-------------|\n| `9:16` | vertical UGC, full-body fashion, avatar talking head |\n| `16:9` | environmental wide, cinematic landscape plate |\n| `3:4` | editorial portrait, try-on full-body |\n| `4:3` | classic portrait, product + person |\n| `1:1` | packshot grid, social tile |\n| `3:2` · `2:3` | magazine / poster crops |\n\nMatch prompt framing to ratio (e.g. `16:9 horizontal wide shot`, `9:16 vertical full body`). **`p-image-try-on`** inherits plate size when `preserve_input_size: true` — diversify person plates first.\n\n**Same character arc:** one ratio for the whole chain unless the user asks for reframes.\n\n## By model (minimum diversity)\n\n| Model | Besides ritual seed, always vary |\n|-------|-----------------------------------|\n| **`p-image-ideogram`** | same axes as `p-image` — photoreal / text / JSON control path |\n| **`p-image`** | cast/creature + objects + action + setting + camera + **`render_category_tag`** + **aspect_ratio**; [explicit structure](#explicit-prompt-structure-required); [text hygiene](#text--typography-by-model) (no upsampling) |\n| **`p-image-edit`** | edit tag + setting/angle delta; same identity URL |\n| **`p-image-try-on`** | person plate world + garment complexity; preserve scene |\n| **`p-image-upscale`** | N/A on prompt — diversify **source** stills |\n| **`p-video-2`** | motion/scene tag + `video_prompt`; differ start plates per scene (quality path) |\n| **`p-video`** | same axes — simpler / quicker clips |\n| **`p-video-avatar`** | `video_prompt` + still world per scene; lock voice per character |\n| **`p-video-animate`** | persona still style/setting per slider ref |\n| **`p-video-replace`** | video-edit tag + full cast spread on showcase reels |\n| **`p-video-edit`** | video-edit tag + one-change lock (keep camera / motion / unmentioned subjects) |\n\n## When **not** to maximize diversity\n\n- **Same character arc** — lock hero plate URL, one `voice`, cast descriptor; vary only setting/angle/motion per scene.\n- **User asked for continuity** — match their cast and approved plates.\n- **Draft → final** — same prompt; change only `draft: false`. Use `api_seed` only if user locked API reproducibility.\n\n## Anti-patterns\n\n| Wrong | Right |\n|-------|--------|\n| Copy doc example ritual strings | [Random seed ritual](#random-seed-ritual-mandatory-before-every-generation) — fresh string each time |\n| Pass ritual string as API `seed` | Ritual is SSoT planning only; `api_seed` when user requests |\n| White wall + MC CU on every demo | Rotate setting + camera + cast |\n| One `video_prompt` for whole reel | Unique motion per scene row |\n| New ritual string mid avatar chain on same brief | Reuse `ritual_seed` until recast or new independent output |\n| Same aspect ratio on every playground example | Rotate `1:1` · `16:9` · `9:16` · `4:3` · `3:4` · `3:2` · `2:3` per [aspect ratio rules](#aspect-ratio-multi-example-sets) |\n| Every hero same athletic body | Rotate [body type spread](#body-type-spread) |\n| Generic hospital hallway | Named ER set dressing + mixed body types in crowd |\n| `holographic` / `chrome` on photoreal cyber scenes | Worn leather, scratched cabinets, documentary skin cues |\n| Monoculture crowd in a named global city | [Location-matched crowds](#location-matched-crowds) — hero ≠ background ethnicity |\n| Magenta-cyan neon on photoreal gym | Warm overhead studio light, mirror wall, real spin bikes |\n| All-male or all-female group class | [Group classes](#group-classes--courses) — mixed-gender background cues |\n| Centered subject every frame | [Framing & camera](#framing--camera) — rotate `camera_tag` + placement |\n| Subject facing camera / at viewer | Off-lens gaze, profile, from behind, or watching crowd |\n| Random animals with no setting reason | Animals only when place implies them |\n| Every stylized panel is anime | Rotate [render categories](#render-categories) — use `cartoon_anime_fantasy` at most once per batch |\n| Vague `cool portrait, neon vibes` | [Explicit structure](#explicit-prompt-structure-required) — named subject, action, objects, setting |\n| `no text` / `without signage` in prompt | Negation invokes text — use [text rules by model](#text--typography-by-model) |\n| Dense typography on **`p-image`** | Drop copy or simplify the brief — `p-image` has no prompt upsampling |\n\n## Visual variety\n\nUse this whenever you plan **`p-image-ideogram`**, **`p-image`**, **`p-image-edit`**, **`p-video-2`**, **`p-video`**, **`p-video-avatar`**, **`p-video-animate`**, **`p-video-replace`**, or **`p-video-edit`** rows. Run the **Variety checklist** at the bottom before the first API call.\n\n### Goal\n\nShowcase and multi-scene work should feel **art-directed**, not like the same talking head in the same office repeated eight times. Deliberately vary:\n\n- **Cast** — gender, age band, ethnicity, persona archetype\n- **Setting** — background / environment (never repeat the same location + framing twice in a row)\n- **Camera** — angle, shot size, movement grammar\n- **Lighting** — time of day, key/fill mood, practical vs cinematic\n- **Visual style** — photoreal, pencil sketch, hand-drawn 2D, cel anime, flat vector, stop-motion clay, CG 3D film, cyberpunk, blockbuster film, editorial, etc.\n- **Render medium** — how the frame is made: `photoreal` · `pencil_sketch` · `hand_drawn_2d` · `cel_anime_2d` · `stop_motion_3d` · `cg_3d_film` (orthogonal to subject family)\n\n**Rule:** Within one **scene row**, lock a local **style bible** so references in that row match (e.g. all three anime refs share the same cel-shaded look). **Across scene rows**, push variety — alternate worlds, angles, and lighting.\n\n## Dynamic prompt stack (eye-catching)\n\nEvery still prompt should feel **art-directed and thumb-stopping**, not generic stock. Build in order:\n\n1. **Style + subject** — who they are + one statement wardrobe piece\n2. **World** — 2–3 concrete environment cues (city bokeh, mirror panels, miniature teal lamp, twin moons)\n3. **Lighting name** — in **`p-image` / reference stills**: bright environment (sunny window, cheerful daylight, golden afternoon). **Avoid** ring light, studio lighting, key/rim/gel light **wording** in still prompts — those belong in plan `lighting_tag` + `video_prompt`, not `p-image`.\n4. **Shot framing** — in still prompts: slight angle from the side, wide shot, slight high angle — **not** “facing camera”, “three-quarter”, or “3/4”. Record angle in plan `camera_tag`.\n5. **`swap_visual_bible`** (plan) — amplify contrast on persona-ladder refs\n\n**Anti-pattern:** flat “neutral wall, soft natural light” on **every row in a scene** — especially UGC/install beats with three grey-wall refs. **Fix:** distinct location family per ref (loft · rooftop · cafe · LED studio) + named gel rim + varied **`camera_tag`** (low angle, side angle, slight high angle).\n\nSee **Prompt patterns** below — flash without text/collage artifacts.\n\n## Creative attractiveness (beyond cast & medium)\n\nSubject diversity is necessary but not sufficient. Thumb-stopping frames also need **color**, **composition**, **texture**, and **motion** variety.\n\n### Color palette ladder\n\nAssign a **`palette_tag`** per scene or ref so sliders do not all read as teal-and-amber:\n\n| Palette | Wardrobe + light pairing |\n|---------|--------------------------|\n| **Warm punch** | coral wall + magenta-cyan LED + gold chain |\n| **Cool contrast** | cobalt hoodie + teal edge light |\n| **Split gel** | rose-gold key + cyan-magenta rim (editorial) |\n| **Monochrome pop** | charcoal + single vivid accent (lime crew, orange sculpture) |\n| **Earth luxe** | walnut desk + copper prop + tungsten accent |\n| **Neon editorial** | violet hair + cherry-blossom bokeh + magenta-teal ambient |\n\n**Rule:** one **dominant accent color** per ref at thumbnail scale — avoid muddy mid-tones everywhere.\n\n### Texture & material\n\nName fabrics and surfaces in prompts — models respond strongly to material words:\n\n- faux fur · holographic puffer · satin wrap · matte clay · crosshatching · glossy chrome armor · walnut grain · matte ceramic\n\nScene 3 already stacks texture beats (fur, holo, pearl/gold); reuse that pattern on wardrobe rows elsewhere.\n\n### Composition & depth\n\n- **Shallow DOF** + gel reflections (single-subject neon boutique) or city window bokeh (studio) — separates subject from background\n- **Foreground anchor** — mug, **closed hardcover notebook**, tumbler at chest gives replace sliders a readable swap target\n- **Single subject one frame** — always; negative space on one side reads cleaner in inset thumbnails\n\n### Age & profession spread\n\nNot every scene needs “tech founder early 30s.” Rotate:\n\n- Gen-Z UGC creator · mid-30s creative director · late-20s advocate · **40s+ expert/trainer** for one VO row\n- Archetypes beyond tech: stylist, chef, fitness creator, museum docent — when the narrative allows\n\n### Camera & motion (reel-level)\n\nCurrent plan anti-pattern to avoid: every scene `medium_cu_dolly_in`. Spread:\n\n| Scene role | Suggested `camera_tag` | `video_prompt` grammar |\n|------------|------------------------|-------------------------|\n| Hook ladder | medium_cu_dolly_in + quarter-orbit | dolly + orbit |\n| UGC install | low_angle_handheld | handheld sway + arc left + push |\n| UGC ref ladder | low angle · side angle · slight high angle | vary per ref within one scene row |\n| Editorial gate | medium_cu_handheld | slow arc right |\n| Desk props | medium_cu_slow_arc | arc + push |\n| CTA | medium_cu_crane_settle | dolly + crane-down |\n\nVary **gaze beats** in `video_prompt` (glance to prop, bookshelf, mirror) — not only straight-to-lens.\n\n### Slider pacing\n\nLong persona ladders (7–9 refs) need tighter **`slider_seconds`** (1.25–1.5) or trim refs — otherwise hook scene dominates reel runtime.\n\n### Quality gates before Phase B\n\n- Ref still readable at **256px wide** (identity + accent color)\n- Adjacent refs differ in **medium + palette + setting**, not just hair color\n- `instruction_prompt` colors/materials **match** reference prompt (lime crew ≠ forest green; copper ≠ silver)\n- Source `video_prompt` props **match** `still_edit` (no mug glance if no mug in plate)\n\n## Cast diversity\n\nPlan a **cast ledger** before generation. For **skills-library / showcase batches** (not single-spokesperson arcs):\n\n| Rule | Guidance |\n|------|----------|\n| **Source host** | **Different person per scene row** — `plate_mode: p-image` + unique `cast_descriptor`. Do not hero-edit one female presenter into every male/advocacy row. |\n| **Reference beats** | Prefer **full recasts** (different ethnicity, age, archetype per ref) over three wardrobe tweaks on one face when proving library range. |\n| **Gender** | Alternate **`persona_gender`** and matching Pruna **`voice`** (`Zephyr (Female)` / `Puck (Male)`) across scenes when lip-sync VO matters. Face-swap refs must stay **same gender** as the source subject on talking-head beats. |\n| **Ethnicity / region** | Name specific, respectful descriptors in prompts (South Asian, East Asian, Black, Latina, Middle Eastern, Nordic, Mediterranean, etc.) — spread representation across the reel, not one token face. |\n| **Age** | Mix early 20s creator energy, mid-30s founder, 40s+ expert — match wardrobe and setting to age. |\n| **Persona archetype** | UGC creator, corporate trainer, fantasy warrior, anime hero, clay character, cyberpunk netrunner, fairy-tale royal, **anthropomorphic otter/fox presenter**, documentary host, gym creator, stylist, etc. |\n| **Subject family** | Photoreal human · fictional character · anthropomorphic (humanoid) · stylized 3D · wardrobe-only · accessories-only · object prop |\n\n**Eye-catching persona ladder (replace hook / animate slider):** one source performance → 5–7 **wildly different** reference stills — e.g. photoreal UGC → premium anime → claymation → cyberpunk → epic film warrior → **anthropomorphic library host** → **fairy-tale 3D royal**. Each ref gets its **own environment, lighting, wardrobe, and subject type**.\n\n**Wardrobe & accessories:** dedicate whole slider steps to **outfit-only** (bolero, vest) and **accessory-only** (scarf, choker, hat, statement earrings) with per-reference `instruction_prompt` naming the slot — same talent, new look, lips unchanged.\n\n## Background & setting ladder\n\nNo two consecutive scene rows should share the **same location type + shot size**. Rotate through distinct worlds:\n\n| Setting family | Example backgrounds |\n|----------------|---------------------|\n| **Domestic / UGC** | bedroom ring light, **creative loft brick**, rooftop dusk, cozy cafe corner, moody LED studio — use **one per ref**, not grey wall ×3 |\n| **Commercial** | boutique, gym floor, outdoor cafe, rooftop at dusk |\n| **Institutional** | classroom whiteboard, news desk, museum gallery |\n| **Fantasy / sci-fi** | stone temple courtyard, alien canyon twin moons, neon arcade corridor, enchanted garden |\n| **Stylized miniature** | clay living room set, diorama street, stop-motion bookshelf nook |\n| **Urban / editorial** | cherry-blossom night street, brutalist plaza, subway platform bokeh |\n\nRecord **`setting_tag`** per scene in the plan (e.g. `\"neon_anime_alley\"`, `\"clay_living_room\"`, `\"temple_courtyard\"`) and verify no duplicate tags in adjacent rows.\n\n## Camera angle & movement ladder\n\nVary **shot size**, **angle**, and **movement** per scene. Never default every row to medium close-up + gentle dolly.\n\n| Angle / size | When to use |\n|--------------|-------------|\n| Extreme close-up (eyes / mouth) | Hook tension, lip-sync proof |\n| Medium close-up (chest-up) | Default VO rows — mouth visible |\n| Medium wide (waist-up) | Wardrobe beats, props in frame |\n| Low angle (heroic) | Game knight, blockbuster reveal |\n| High angle (vulnerable / editorial) | Documentary, stylized anime |\n| Over-shoulder turning in | Explainer, product demo |\n| Profile side angle | Stylized refs when motion allows |\n\n**Movement grammar** (prefix `video_prompt` with continuous camera):\n\n- gentle dolly push-in · slow arc left · subtle handheld sway · orbit quarter-left · crane-down settle · tracking follow (silent B-roll only)\n\n**Anti-pattern:** eight scenes, all `medium close-up, gentle dolly push-in`.\n\n## Lighting ladder\n\nName lighting in every **`p-image` prompt**, **`still_edit`**, and reference still:\n\n| Mood | Prompt cues |\n|------|-------------|\n| Soft overcast documentary | even skin, neutral shadows |\n| Golden hour warm | rim light, amber fill, long shadows |\n| Neon / cyberpunk | magenta-cyan edge light, wet reflections |\n| Anime film dramatic | strong key, colored bounce, neon bokeh |\n| Stop-motion practical | warm desk lamp, miniature set glow |\n| Blockbuster / game cinematic | motivated sun shafts, volumetric haze |\n| Clean educational | bright even key, soft classroom fill |\n| Low-key cinematic | single motivated source, deep background falloff |\n\nAlternate lighting families across scenes — not only \"soft natural light\" on every row.\n\n## Visual style ladder\n\nFor **showcase and multi-scene** work, plan at least **4 distinct visual styles** across the full piece. Pick from (mix and match):\n\n| Style tag | Prompt direction |\n|-----------|------------------|\n| **Photoreal UGC** | smartphone-adjacent, natural skin, real locations |\n| **Photoreal commercial** | crisp product labels, controlled studio or location |\n| **Pencil / charcoal sketch** | crosshatching on cream paper, art-studio daylight, mouth visible |\n| **Hand-drawn 2D animation** | ink outlines, watercolor wash, golden-age animation palette — **single frame** |\n| **Premium anime (2D cel)** | cel-shaded, film-grade compositing, stylized hair/eyes |\n| **Flat vector 2D** | bold shapes, limited palette, motion-graphics friendly |\n| **Disney / Pixar 3D (CG film)** | rounded forms, storybook warmth, enchanted environments |\n| **Claymation / stop-motion 3D** | visible clay texture, miniature sets, practical lighting |\n| **Cyberpunk** | chrome, neon arcade corridor, HUD-free (no readable UI text) |\n| **Blockbuster movie** | anamorphic cues, epic scale, costume drama |\n| **Editorial fashion** | bold wardrobe, shallow DOF, magazine angles |\n| **Documentary** | handheld honesty, available light |\n| **Meme / reaction** | dorm, gaming chair, exaggerated expression (reaction / meme beat pattern) |\n| **Anthropomorphic** | humanoid otter/fox/red panda presenter, expressive face, mouth visible, cozy set |\n\n**Rendering medium ladder (persona hooks):** aim for **5+ mediums** in one slider when showcasing range — e.g. photoreal → pencil sketch → 2D ink frame → cel anime → stop-motion clay → CG 3D royal. Tag optional `render_medium_tag` per ref in plans.\n\n**Animate rows:** generate **one persona still per style** on the same motion template — each still carries its own \n\nArchive v1.0.12: 9 files, 35148 bytes\n\nFiles: references/clarification-intake.md (7454b), references/generation-diversity.md (51115b), references/generation-quality-checklists.md (8373b), references/still-image-prompt-flow.md (6935b), references/workflow-feedback-gates.md (2141b), skill-card.md (2491b), skill.manifest.json (195b), SKILL.md (5807b), _meta.json (140b)\n\nArchive v1.0.11: 9 files, 34712 bytes\n\nFiles: references/clarification-intake.md (7421b), references/generation-diversity.md (50711b), references/generation-quality-checklists.md (7894b), references/still-image-prompt-flow.md (6586b), references/workflow-feedback-gates.md (2141b), skill-card.md (2644b), skill.manifest.json (195b), SKILL.md (5151b), _meta.json (140b)\n\nArchive v1.0.10: 9 files, 34327 bytes\n\nFiles: references/clarification-intake.md (7421b), references/generation-diversity.md (50244b), references/generation-quality-checklists.md (7707b), references/still-image-prompt-flow.md (6586b), references/workflow-feedback-gates.md (2141b), skill-card.md (2117b), skill.manifest.json (195b), SKILL.md (5151b), _meta.json (140b)\n\nArchive v1.0.9: 9 files, 34331 bytes\n\nFiles: references/clarification-intake.md (7421b), references/generation-diversity.md (50244b), references/generation-quality-checklists.md (7707b), references/still-image-prompt-flow.md (6586b), references/workflow-feedback-gates.md (2141b), skill-card.md (2251b), skill.manifest.json (195b), SKILL.md (5150b), _meta.json (139b)\n\nArchive v1.0.8: 9 files, 34167 bytes\n\nFiles: references/clarification-intake.md (6978b), references/generation-diversity.md (49975b), references/generation-quality-checklists.md (7707b), references/still-image-prompt-flow.md (6204b), references/workflow-feedback-gates.md (2141b), skill-card.md (2763b), skill.manifest.json (195b), SKILL.md (5086b), _meta.json (139b)\n\nArchive v1.0.7: 8 files, 30193 bytes\n\nFiles: references/generation-diversity.md (49975b), references/generation-quality-checklists.md (7707b), references/still-image-prompt-flow.md (6204b), references/workflow-feedback-gates.md (2141b), skill-card.md (2137b), skill.manifest.json (164b), SKILL.md (4606b), _meta.json (139b)\n\nArchive v1.0.2: 6 files, 14051 bytes\n\nFiles: README-INSTALL.md (212b), references/generation-diversity.md (26090b), references/random-seed-ritual.md (4088b), skill.manifest.json (128b), SKILL.md (1305b), _meta.json (139b)\n\nArchive v1.0.1: 7 files, 15519 bytes\n\nFiles: README-INSTALL.md (212b), references/generation-diversity.md (26796b), references/random-seed-ritual.md (4025b), skill-card.md (2298b), skill.manifest.json (128b), SKILL.md (1305b), _meta.json (139b)","readmeExcerpt":"Skill: generation-diversity Owner: pruna-ai Summary: Use when writing any generative prompt — ritual seed, explicit structure, scenario axes, and quality gates before paid API calls. Tags: ai:1.0.14, generative:1.0.14, latest:1.0.14, pruna:1.0.14 Version history: v1.0.14 | 2026-09-29T15:28:56.996Z | auto generation-diversity v1.0.14 - Added explicit mention of p-video-2-pro mode (cost/speed/quality) in the clarificat","codeSnippets":[],"executableExamples":[{"language":"json","snippet":"{\n  \"ritual_seed_policy\": \"ssot_dag_before_every_generation\",\n  \"ritual_seed\": \"k7Qm2xP9\",\n  \"seed_log\": [\n    { \"phase\": \"hero_p_image\", \"ritual_seed\": \"k7Qm2xP9\", \"creature_tag\": \"otter_dj\", \"setting_tag\": \"1970s_roller_rink\", \"prompt_hash\": \"…\" },\n    { \"phase\": \"scene_2_avatar\", \"ritual_seed\": \"k7Qm2xP9\", \"scene_id\": 2 }\n  ]\n}"},{"language":"text","snippet":"{people and/or creatures} {action} with/at {specific objects} in {named setting},\n{style or era cues}, {camera_tag}, {lighting_tag}"},{"language":"text","snippet":"Disco ball reflections on an otter DJ scratching vinyl at a packed 1970s roller rink,\nfish-eye lens, glitter confetti mid-air, funky energy"},{"language":"text","snippet":"Bioluminescent jellyfish nightclub at abyss depth, VIP anglerfish in sunglasses at velvet rope,\nteal-magenta cinematic lighting"},{"language":"text","snippet":"Corgi cowboy lassoing a runaway taco truck through Monument Valley dust storm,\npulp western poster energy, dynamic diagonal composition"},{"language":"text","snippet":"RATIOS = [\"1:1\", \"16:9\", \"9:16\", \"4:3\", \"3:4\", \"3:2\", \"2:3\"]\naspect_ratio  ← RATIOS[ sum(codes(ritual_seed)) % 7 ]\ncamera_tag    ← camera_tags[ sum(codes(ritual_seed[0:4])) % len(camera_tags) ]\nrender_tag    ← render_tags[ sum(codes(ritual_seed[4:8])) % len(render_tags) ]"}],"parameters":null,"dependencies":[],"permissions":[],"extractedFiles":[{"path":"SKILL.md","content":"---\nname: generation-diversity\ndescription: Use when writing any generative prompt — ritual seed, explicit structure, scenario axes, and quality gates before paid API calls.\nlicense: MIT\nmetadata:\n  version: \"1.0.14\"\n  package: pruna-skills\n---\n\n# Generation diversity\n\nVendor-neutral playbook for **diverse, explicit prompts** and output QA. Apply before every generation on any model (Pruna, Flux, Midjourney, Runway, ElevenLabs, …).\n\n## Install\n\n| Skill | Description | Install |\n| --- | --- | --- |\n| `generation-diversity` | Use when writing any generative prompt — ritual seed, explicit structure, scenario axes, and quality gates before paid API calls. | `npx skills add PrunaAI/pruna-skills@generation-diversity -y` |\n\n## When to use\n\n- Starting a new image, video, or audio generation\n- Outputs feel repetitive or “AI sloppy”\n- Multi-example batches that need cast/setting/camera variety\n- Before advancing a multi-step workflow past a phase gate\n\n## Works with\n\nAny generative model. Pruna tools (`p-image-ideogram`, `p-image`, `p-video-2-pro`, `p-video-2`, `p-video`, …) and third-party APIs alike.\n\n## Guide habit\n\nIn the **first reply**, name `` `generation-diversity` `` in backticks. When the brief leaves media source, brand, audio, structure, resolution, or approval unclear, **[ask before spending](./references/clarification-intake.md)** — every tool and workflow defers here for shared intake topics. For still-image jobs (`p-image-ideogram`, `p-image`, `p-image-edit`), point agents at **[Still-image prompt flow](./references/still-image-prompt-flow.md)** — brief lock → ritual → axes → explicit prompt → fidelity check. **Mood boards:** new ritual per independent panel; user-locked brand hex / subject stays locked on every panel.\n\n## Before generating\n\n0. **[Clarification intake](./references/clarification-intake.md)** — generate vs existing assets, colors, narration/VO, music, captions, aspect/resolution (480p/768p vs 720p/1080p, canvas, MP), `p-video-2-pro` `mode` (cost/speed/quality), structure, approval (unless the user waived or already locked answers).\n1. **[Generation diversity](./references/generation-diversity.md)** — random seed ritual (SSoT), explicit prompt structure, rotate ≥2 scenario axes per session.\n2. **Still images (`p-image` family):** **[still-image-prompt-flow.md](./references/still-image-prompt-flow.md)** — generation flow, edit flow, mood-board rules, hero → edit handoff. Pair with `image-prompting` golden rules and edit craft.\n3. **[Quality checklists](./references/generation-quality-checklists.md)** — open outputs and judge pass/fail before the next paid step.\n4. **Workflows:** [workflow-feedback-gates.md](./references/workflow-feedback-gates.md) — pause at plan / stills / clips before paid video.\n\n## Red flags\n\nStop and ask (or show assets) before the next paid step if any of these are true:\n\n| Red flag | Required action |\n| --- | --- |\n| User says skip review / burn video credits / run everything now | Refuse same-turn plan"},{"path":"_meta.json","content":"{\n  \"ownerId\": \"kn7cagwf7q3t0cxrgteb7xk0bh81j0eb\",\n  \"slug\": \"generation-diversity\",\n  \"version\": \"1.0.14\",\n  \"publishedAt\": 1790695736996\n}"},{"path":"references/clarification-intake.md","content":"# Clarification intake — ask before you spend\n\nUse when the user’s request could mean more than one deliverable, more than one media path, or more than one creative default. **Ask in the first reply** (or right after routing), bundle related questions, and **record answers** in the plan or manifest. Do not start paid `POST`s, bulk generation, or long renders until missing decisions are answered or the user explicitly waives them (“use your judgment”, “surprise me”, “just build it”).\n\nWorkflow skills with **`Intake: ask before generating`** tables are authoritative for that deliverable. This doc is the **shared topic list** every Pruna guide and tool should use when the brief is silent.\n\n## How to ask\n\n| Do | Don't |\n| --- | --- |\n| Offer **2–3 concrete options** plus “other” when the choice is structural (acts, route, layout) | Interrogate field-by-field when the user already gave a locked brief |\n| **Group** questions (media + audio in one message when both are open) | Same-turn **plan + paid video** without **approve plan** / gates |\n| Use structured choice UI when the host supports it (many independent decisions) | Invent brand colors, voice, or “generate vs existing” when cost or look changes |\n| Say what you **assume** if they waive, and proceed with a receipt in the summary | Treat inference as confirmation — restate inferred defaults separately |\n\n**Red flags** (must clarify or show plan): skip review, burn credits, automation flags without explicit opt-in → see `generation-diversity` **Red flags** and [workflow-feedback-gates.md](./workflow-feedback-gates.md).\n\n### Example bundles (adapt to the job)\n\nUse when several topics are open at once; trim if the user already locked some:\n\n1. **Deliverable shape:** “Should I **generate** new visuals/audio, **use files you already have**, or **mix** (e.g. your logo + generated B-roll)? Any **aspect ratio** and **resolution** target (480p/768p for cinematic `p-video-2-pro`, 720p vs 1080p for `p-video-2`, 9:16 vs 16:9)?”\n2. **Look and sound:** “**Brand palette** (named kit vs custom hex / reference image)? **Narration or VO** (none, TTS, lip-sync host, your upload)? **Music** (silent, bed under VO, full song)? **Captions** (none, burned after render, in-composition)?”\n3. **Structure and gates:** “Single clip or **multi-act** piece? If multi-act, I can propose **two orderings** — which direction? OK to pause for **approve plan** before paid video, or run end-to-end?”\n\nFor still-only jobs, add **aspect_ratio** and **megapixel / upscale target** when using `p-image-upscale` or print-sized exports.\n\n## Universal topics\n\nAsk when the brief does not already answer these:\n\n| Topic | Clarify |\n| --- | --- |\n| **Media source** | Generate new assets (image/video/audio API) vs **use existing** files/URLs vs **mix** (e.g. user logo + generated B-roll) |\n| **Brand / look** | Named brand kit vs ad-hoc palette (hex or reference), light vs dark, photoreal vs stylized |\n| **Narration / VO** | None · on-screen text onl"},{"path":"references/generation-diversity.md","content":"# Generation diversity (all models)\n\nOne policy for **every** generative output — images, video, try-on, avatars, replace, animate (Pruna or otherwise). Covers the **random seed ritual**, explicit prompt structure, scenario axis rotation, and **visual variety** ladders.\n\nUse the **full** checklist here for every generation.\n\n## Contents\n\n- [Random seed ritual](#random-seed-ritual-mandatory-before-every-generation)\n- [Three steps (every job)](#three-steps-every-job)\n- [Still-image prompt flow](./still-image-prompt-flow.md) — `p-image-ideogram` / `p-image` / `p-image-edit` agent pipeline (brief lock → ritual → POST)\n- [Explicit prompt structure](#explicit-prompt-structure-required)\n- [Text & typography by model](#text--typography-by-model)\n- [SSoT axis derivation](#ssot-axis-derivation-sum-mod)\n- [Scenario axes](#scenario-axes-rotate-across-outputs)\n- [Render categories](#render-categories)\n- [Crowded scenes](#crowded-scenes-p-image)\n- [Body type spread](#body-type-spread)\n- [Location-matched crowds](#location-matched-crowds)\n- [Group classes](#group-classes--courses)\n- [Framing & camera](#framing--camera)\n- [Scene spice](#scene-spice-when-it-fits)\n- [Photoreal anti-slop](#photoreal-anti-slop-neon--stylized-briefs)\n- [Aspect ratio](#aspect-ratio-multi-example-sets)\n- [By model](#by-model-minimum-diversity)\n- [When not to maximize diversity](#when-not-to-maximize-diversity)\n- [Visual variety](#visual-variety)\n- [Variety checklist](#variety-checklist-before-first-api-call)\n- [Prompt patterns (variety)](#prompt-patterns-variety)\n- [Anti-patterns](#anti-patterns)\n\n## Three steps (every job)\n\n1. **[Random seed ritual](#random-seed-ritual-mandatory-before-every-generation) (SSoT)** — **always first**, before the prompt. Generate a fresh random string, **state it in the turn**, derive axes via [sum-mod](#ssot-axis-derivation-sum-mod). **Do not** pass the ritual string to API `seed`. **One new ritual string per independent generation**; reuse only on same-brief slop retry.\n2. **Write an [explicit prompt](#explicit-prompt-structure-required)** — name specific people, animals, objects, actions, setting, and camera/light. Add text/typography only when the brief needs it — see [text rules by model](#text--typography-by-model).\n3. **Diversify the scenario row** — change at least **two axes** from the previous output in the same session (cast, setting, camera, **`render_category_tag`**, **aspect_ratio**, creatures, props, … — unless user asked for continuity).\n4. **Log** — `ritual_seed`, axes chosen, prediction id (manifest or turn text).\n\n\n\n## Random seed ritual (mandatory before every generation)\n\nThe random seed ritual is a lean [String Seed of Thought](https://pub.sakana.ai/ssot/) (DAG) protocol. **Every** Pruna generation — every prompt, every `POST /v1/predictions`, every scene row — starts here.\n\nThis prevents copy-pasting example strings (`k7Qm2xP9`, `482901`, …) and reduces accidental duplicate outputs across sessions.\n\n### The ritual (do this first)\n\nB"},{"path":"references/generation-quality-checklists.md","content":"# Generation quality checklist hub\n\nUse this as the shared quality gate across models and workflows.\nRun the **Core checklist** for every generation job, then run the model-specific checklist in the guide/workflow skill named below.\n\n## Who applies these checklists?\n\n**The coding agent** — by **opening the real output files** (images, video, or audio) and reviewing them with vision. These checklists are **not** automated test scripts. There is no separate scoring service: the agent reads each item and judges pass or fail from what it sees and hears.\n\nTypical flow:\n\n1. **Generate or download** the asset to a local path (`stills/`, `clips/`, etc.).\n2. **Inspect the file** — view the image, watch the video clip, or listen to narration when the checklist covers audio.\n3. Run the **Core checklist** (below), then the **model-specific checklist** for that job.\n4. **If something fails** — note which items failed, adjust prompt / settings / seed, and regenerate **only that asset** (do not advance to expensive video steps on a bad still).\n5. **If it passes** — show the user the file paths (and previews when helpful). In workflows, still follow [approval gates](#approval-gates-workflows): agent checklist review happens **before** you ask the user to approve stills or clips.\n\nThe user's **approve plan / approve stills / approve clips** gates are separate. Agent checklists catch obvious problems early so the user is not asked to sign off on broken outputs.\n\nMaintenance rule: keep tool/workflow mapping only in this file to avoid link drift. Cross-skill craft lives in the named skills — install them; do not hyperlink into their trees.\n\n## Match map (tool → checklist skill → workflows)\n\n| Tool/model | Guide / workflow | Checklist file (inside that skill) | Common workflows |\n|------------|------------------|------------------------------------|------------------|\n| `p-image-ideogram` | `image-prompting` | `p-image-quality-checklist.md` · persona: `realistic-persona-showcase.md` | `image-to-video`, `narrated-multi-scene`, `avatar-multi-scene` |\n| `p-image` | `image-prompting` | `p-image-quality-checklist.md` · persona: `realistic-persona-showcase.md` | same workflows on cheap / fast drafts |\n| `p-image-edit` | `image-prompting` | `p-image-edit-quality-checklist.md` | `avatar-single-scene`, `avatar-multi-scene` |\n| `p-image-upscale` | `image-prompting` | `p-image-upscale-quality-checklist.md` | `image-to-video`, `narrated-multi-scene` |\n| `p-image-try-on` | `image-prompting` | `p-image-try-on-quality-checklist.md` · persona: `realistic-persona-showcase.md` | `p-image-try-on` |\n| `p-video-2-pro` | `video-prompting` | `p-video-2-pro-quality-checklist.md` | `image-to-video` (visual-only), `visual-transition-reel` |\n| `p-video-2` | `video-prompting` | `p-video-2-quality-checklist.md` | `image-to-video` (audio-led), `narrated-multi-scene`, `interactive-explainer` |\n| `p-video` | `video-prompting` | `p-video-quality-checklist.md` | same workflows on simpler / quicker cl"}],"languages":[],"docsSourceLabel":"CLAWHUB","editorialOverview":null,"editorialQuality":{"score":100,"threshold":65,"status":"thin","wordCount":2263,"uniquenessScore":41,"reasons":["uniqueness-below-45"]}},"media":{"evidence":{"source":"no-media","verified":false,"confidence":"low","updatedAt":"2026-10-11T00:37:51.974Z","emptyReason":"No screenshots, media assets, or demo links are available."},"primaryImageUrl":null,"mediaAssetCount":0,"assets":[],"demoUrl":null},"ownerResources":{"evidence":{"source":"unclaimed","verified":false,"confidence":"low","updatedAt":"2026-10-11T00:37:51.974Z","emptyReason":"This page has not been claimed by the agent owner."},"hasCustomPage":false,"customPageUpdatedAt":null,"customLinks":[],"structuredLinks":{"docsUrl":null,"demoUrl":null,"supportUrl":null,"pricingUrl":null,"statusUrl":null},"customPage":null},"relatedAgents":{"evidence":{"source":"protocol-neighbors","verified":false,"confidence":"medium","updatedAt":"2026-10-11T03:55:33.368Z","emptyReason":null},"items":[{"id":"8ebccd8e-3863-4187-8355-c3f14e1f9edf","entityType":"agent","canonicalPath":"/agent/iofficeai-aionui","slug":"iofficeai-aionui","name":"AionUi","description":"Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!","url":"https://github.com/iOfficeAI/AionUi","homepage":"https://www.aionui.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-10-09T19:11:12.944Z","createdAt":"2026-02-25T03:38:16.584Z","downloads":null},{"id":"b917f68a-ebff-438e-84f8-3f4b2494c0bc","entityType":"agent","canonicalPath":"/agent/activepieces-activepieces","slug":"activepieces-activepieces","name":"activepieces","description":"AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents","url":"https://github.com/activepieces/activepieces","homepage":"https://www.activepieces.com","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-15T02:22:12.426Z","createdAt":"2026-02-25T03:38:12.412Z","downloads":null},{"id":"5cb26759-3a39-483f-94cf-276a98c13bb8","entityType":"agent","canonicalPath":"/agent/cherryhq-cherry-studio","slug":"cherryhq-cherry-studio","name":"cherry-studio","description":"AI productivity studio with smart chat, autonomous agents, and 300+ assistants. Unified access to frontier LLMs","url":"https://github.com/CherryHQ/cherry-studio","homepage":"https://cherry-ai.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-11T14:38:40.986Z","createdAt":"2026-02-25T03:38:19.379Z","downloads":null},{"id":"6f6582d0-5d76-4f0f-b81d-86520247950b","entityType":"agent","canonicalPath":"/agent/copilotkit-copilotkit","slug":"copilotkit-copilotkit","name":"CopilotKit","description":"The Frontend for Agents & Generative UI. React + Angular","url":"https://github.com/CopilotKit/CopilotKit","homepage":"https://docs.copilotkit.ai","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-03-25T09:50:57.846Z","createdAt":"2026-02-25T03:39:14.617Z","downloads":null}],"links":{"hub":"/agent","source":"/agent/source/clawhub","protocols":[{"label":"OpenClaw","href":"/agent/protocol/openclew"}]}}}