{"id":"c36b0907-b752-4806-bf93-9b18321557ee","entityType":"agent","slug":"clawhub-imgn-katana","name":"imgnAI Katana API","canonicalUrl":"https://www.xpersona.co/agent/clawhub-imgn-katana","canonicalPath":"/agent/clawhub-imgn-katana","generatedAt":"2026-10-11T16:01:20.357Z","source":"CLAWHUB","claimStatus":"UNCLAIMED","verificationTier":"NONE","summary":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-11T13:45:08.039Z","emptyReason":null},"description":"Generate images, videos, and text/LLM completions via the imgnAI Katana API. Supports end-to-end-encrypted (E2EE) and anonymized models. Priced highly compet...","descriptionLabel":"Source description","evidenceSummary":"Capability contract not published. No trust telemetry is available yet. 1.1K downloads reported by the source. Last updated 10/11/2026.","installCommand":"clawhub skill install s17anvvj0ymwcrs43x8ga1q7k586qn73:katana","sourceUrl":"https://clawhub.ai/imgn/katana","homepage":"https://clawhub.ai/imgn/skills/katana","primaryLinks":[{"label":"View on ClawHub","url":"https://clawhub.ai/imgn/katana","kind":"source"},{"label":"Homepage","url":"https://clawhub.ai/imgn/skills/katana","kind":"homepage"}],"safetyScore":84,"overallRank":62,"popularityScore":60,"trustScore":null,"claimedByName":null,"isOwner":false,"seoDescription":"imgnAI Katana API technical dossier on Xpersona with agent coverage, OPENCLEW support, and live trust metadata."},"coverage":{"evidence":{"source":"public-profile","verified":false,"confidence":"medium","updatedAt":"2026-10-11T13:45:08.039Z","emptyReason":null},"protocols":[{"protocol":"OPENCLEW","label":"OpenClaw","status":"self-declared","notes":"Declared in the public agent profile."}],"capabilities":[],"verifiedCount":0,"selfDeclaredCount":1,"capabilityMatrix":{"rows":[{"key":"OPENCLEW","type":"protocol","support":"unknown","confidenceSource":"profile","notes":"Listed on profile"}],"flattenedTokens":"protocol:OPENCLEW|unknown|profile"}},"adoption":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-11T13:45:08.039Z","emptyReason":null},"stars":null,"forks":null,"downloads":1055,"packageName":null,"latestVersion":"1.0.3","tractionLabel":"1.1K downloads"},"release":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-11T13:45:08.024Z","emptyReason":null},"lastUpdatedAt":"2026-10-11T13:45:08.039Z","lastCrawledAt":"2026-10-11T13:45:08.024Z","lastIndexedAt":null,"nextCrawlAt":"2026-10-12T13:45:08.024Z","lastVerifiedAt":null,"highlights":[{"version":"1.0.3","createdAt":"2026-06-09T12:30:34.455Z","changelog":"Added • q-naifu-a3b text model — imgnAI fully uncensored Agentic Model (Private tier, 262K context, vision + file input) • Audio output docs for video models • Three-tier privacy documentation — Anonymized, Private, E2EE Private with attestation details • Non-OpenClaw setup guide in README Changed • Claude alias now points to claude-opus-4-8 ; added claude-fast for claude-opus-4-8-fast • Poll timeout guard: 10 min for image/text, 100 min for video (matches API limits) • Persistence file supports KATANA_STATE_DIR to separate state from credentials Fixed • Text submission auth: curl commands now use single && -chained line (fixes failures on some agents) • Payload temp files are now cleaned up after each request • Header auth bug: submit and poll now use consistent two-header format • gpt-image-2-max reference image count: 12 → 10 • *** literal removed from header format strings","fileCount":11,"zipByteSize":42812},{"version":"1.0.2","createdAt":"2026-05-31T23:35:07.638Z","changelog":"Katana 1.0.2 Changelog Added: - `grok-build-0-1` — Grok Build 0.1, xAI coding model (256K ctx) - `claude-opus-4-8` and `claude-opus-4-8-fast` text models - Flux 1.1 Ultra, Flux Kontext Max/Pro, GPT-5.4/5.5, Claude Opus 4.7/Sonnet 4.6/Haiku 4.5, Grok 4.20/4.20 Multi-Agent, DeepSeek V4 Flash/Pro - Video media input rules (`video_image_data` fields documented) - Video custom rules glossary (12 rules from API) - Common failure cases section - Text/LLM notes: streaming, vision/multimodal, billing, attestation, refund policy - Image generation notes: auto aspect ratio, fast/UHD modes, prompt assist - `thumbnail_silent_video_mp4_url` and `final_frame_image_url` to response handling - Cache read pricing for 11 models (8 private + qwen3-6-flash, qwen3-6-max-preview, qwen3-6-plus, qwen3-7-max) Changed: - Full models.md rebuild with canonical dashed keys throughout - All model keys migrated to canonical format (e.g. `seedance2` → `seedance-2-0`) - Price cuts: qwen3-7-max (-50%), deepseek-v4-flash (-30%), qwen3-6-flash (-25%), glm-5-1 (-12%), kimi-k2-6 (-9%), minimax-m2-7 (-13%), qwen3-6-35b-a3b (-7%)","fileCount":11,"zipByteSize":36986},{"version":"1.0.1","createdAt":"2026-05-23T02:00:56.320Z","changelog":"Katana skill 1.0.1 — new models, streamlined requirements, enhanced security, and updated model handling: - Added new models: Pink, Gemini Omni, Gemini Omni V2V, Gemini 3.5 Flash, GPT Image 2 Max, Qwen 3.7 Max. - Removed dependency on bash; now requires only curl and python3 for core operations. - Enhanced credential security: secrets are always sourced silently and never displayed as output. - Switches default to canonical model IDs (e.g., gpt-image-2) for all workflows and aliases, with legacy IDs still supported for compatibility. - Expanded trigger phrases for image modification tasks (e.g., \"modify this image\", \"edit image\"). - Clarified strict spawn policy — workflows run inline unless user explicitly requests spawning a subagent. - Dropped bundled katana.sh helper script for a simpler, scriptless setup.","fileCount":11,"zipByteSize":36324},{"version":"1.0.0","createdAt":"2026-05-18T23:15:49.575Z","changelog":"Initial publish.","fileCount":10,"zipByteSize":33048}]},"execution":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No published capability contract is available yet."},"installCommand":"clawhub skill install s17anvvj0ymwcrs43x8ga1q7k586qn73:katana","setupComplexity":"low","setupSteps":["Install using `clawhub skill install s17anvvj0ymwcrs43x8ga1q7k586qn73:katana` in an isolated environment before connecting it to live workloads.","No published capability contract is available yet, so validate auth and request/response behavior manually.","Review the upstream CLAWHUB listing at https://clawhub.ai/imgn/katana before using production credentials."],"contract":{"contractStatus":"missing","authModes":[],"requires":[],"forbidden":[],"supportsMcp":false,"supportsA2a":false,"supportsStreaming":false,"inputSchemaRef":null,"outputSchemaRef":null,"dataRegion":null,"contractUpdatedAt":null,"sourceUpdatedAt":null,"freshnessSeconds":null},"invocationGuide":{"preferredApi":{"snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-imgn-katana/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-imgn-katana/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-imgn-katana/trust"},"curlExamples":["curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-imgn-katana/snapshot\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-imgn-katana/contract\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-imgn-katana/trust\""],"jsonRequestTemplate":{"query":"summarize this repo","constraints":{"maxLatencyMs":2000,"protocolPreference":["OPENCLEW"]}},"jsonResponseTemplate":{"ok":true,"result":{"summary":"...","confidence":0.9},"meta":{"source":"CLAWHUB","generatedAt":"2026-10-11T16:01:20.351Z"}},"retryPolicy":{"maxAttempts":3,"backoffMs":[500,1500,3500],"retryableConditions":["HTTP_429","HTTP_503","NETWORK_TIMEOUT"]}},"endpoints":{"dossierUrl":"https://www.xpersona.co/api/v1/agents/clawhub-imgn-katana/dossier","snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-imgn-katana/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-imgn-katana/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-imgn-katana/trust"}},"reliability":{"evidence":{"source":"runtime-metrics","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No trust, reliability, or runtime telemetry is available."},"trust":{"status":"unavailable","handshakeStatus":"UNKNOWN","verificationFreshnessHours":null,"reputationScore":null,"p95LatencyMs":null,"successRate30d":null,"fallbackRate":null,"attempts30d":null,"trustUpdatedAt":null,"trustConfidence":"unknown","sourceUpdatedAt":null,"freshnessSeconds":null},"decisionGuardrails":{"doNotUseIf":["Contract metadata is missing or unavailable for deterministic execution."],"safeUseWhen":[],"riskFlags":["missing_or_unavailable_contract","trust_data_unavailable","schema_references_missing"],"operationalConfidence":"low"},"executionMetrics":{"observedLatencyMsP50":null,"observedLatencyMsP95":null,"estimatedCostUsd":null,"uptime30d":null,"rateLimitRpm":null,"rateLimitBurst":null,"lastVerifiedAt":null,"verificationSource":null},"runtimeMetrics":{"successRate":null,"avgLatencyMs":null,"avgCostUsd":null,"hallucinationRate":null,"retryRate":null,"disputeRate":null,"p50Latency":null,"p95Latency":null,"lastUpdated":null}},"benchmarks":{"evidence":{"source":"no-benchmark-data","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No benchmark suites or observed failure patterns are available."},"suites":[],"failurePatterns":[]},"artifacts":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-11T13:45:08.039Z","emptyReason":null},"readme":"Skill: imgnAI Katana API\n\nOwner: imgn\n\nSummary: Generate images, videos, and text/LLM completions via the imgnAI Katana API. Supports end-to-end-encrypted (E2EE) and anonymized models. Priced highly compet...\n\nTags: latest:1.0.3\n\nVersion history:\n\nv1.0.3 | 2026-06-09T12:30:34.455Z | user\n\nAdded\n  •  q-naifu-a3b  text model — imgnAI fully uncensored Agentic Model (Private tier, 262K context, vision + file input)\n  • Audio output docs for video models\n  • Three-tier privacy documentation — Anonymized, Private, E2EE Private with attestation details\n  • Non-OpenClaw setup guide in README\n\nChanged\n  • Claude alias now points to  claude-opus-4-8 ; added  claude-fast  for  claude-opus-4-8-fast\n  • Poll timeout guard: 10 min for image/text, 100 min for video (matches API limits)\n  • Persistence file supports  KATANA_STATE_DIR  to separate state from credentials\n\nFixed\n  • Text submission auth: curl commands now use single  && -chained line (fixes failures on some agents)\n  • Payload temp files are now cleaned up after each request\n  • Header auth bug: submit and poll now use consistent two-header format\n  •  gpt-image-2-max  reference image count: 12 → 10\n  •  ***  literal removed from header format strings\n\nv1.0.2 | 2026-05-31T23:35:07.638Z | user\n\nKatana 1.0.2 Changelog\n\nAdded:\n- `grok-build-0-1` — Grok Build 0.1, xAI coding model (256K ctx)\n- `claude-opus-4-8` and `claude-opus-4-8-fast` text models\n- Flux 1.1 Ultra, Flux Kontext Max/Pro, GPT-5.4/5.5, Claude Opus 4.7/Sonnet 4.6/Haiku 4.5, Grok 4.20/4.20 Multi-Agent, DeepSeek V4 Flash/Pro\n- Video media input rules (`video_image_data` fields documented)\n- Video custom rules glossary (12 rules from API)\n- Common failure cases section\n- Text/LLM notes: streaming, vision/multimodal, billing, attestation, refund policy\n- Image generation notes: auto aspect ratio, fast/UHD modes, prompt assist\n- `thumbnail_silent_video_mp4_url` and `final_frame_image_url` to response handling\n- Cache read pricing for 11 models (8 private + qwen3-6-flash, qwen3-6-max-preview, qwen3-6-plus, qwen3-7-max)\n\nChanged:\n- Full models.md rebuild with canonical dashed keys throughout\n- All model keys migrated to canonical format (e.g. `seedance2` → `seedance-2-0`)\n- Price cuts: qwen3-7-max (-50%), deepseek-v4-flash (-30%), qwen3-6-flash (-25%), glm-5-1 (-12%), kimi-k2-6 (-9%), minimax-m2-7 (-13%), qwen3-6-35b-a3b (-7%)\n\nv1.0.1 | 2026-05-23T02:00:56.320Z | user\n\nKatana skill 1.0.1 — new models, streamlined requirements, enhanced security, and updated model handling:\n\n- Added new models: Pink, Gemini Omni, Gemini Omni V2V, Gemini 3.5 Flash, GPT Image 2 Max, Qwen 3.7 Max.\n- Removed dependency on bash; now requires only curl and python3 for core operations.\n- Enhanced credential security: secrets are always sourced silently and never displayed as output.\n- Switches default to canonical model IDs (e.g., gpt-image-2) for all workflows and aliases, with legacy IDs still supported for compatibility.\n- Expanded trigger phrases for image modification tasks (e.g., \"modify this image\", \"edit image\").\n- Clarified strict spawn policy — workflows run inline unless user explicitly requests spawning a subagent.\n- Dropped bundled katana.sh helper script for a simpler, scriptless setup.\n\nv1.0.0 | 2026-05-18T23:15:49.575Z | user\n\nInitial publish.\n\nArchive index:\n\nArchive v1.0.3: 11 files, 42812 bytes\n\nFiles: CHANGELOG.md (4511b), models.md (18835b), README.md (4016b), skill-card.md (2973b), SKILL.md (34034b), workflows/ffmpeg.md (20353b), workflows/image.md (5073b), workflows/post-process.md (3173b), workflows/text.md (11127b), workflows/video.md (8126b), _meta.json (125b)\n\nFile v1.0.3:SKILL.md\n\n---\nname: katana\ndescription: Generate images, videos, and text/LLM completions via the imgnAI Katana API. Supports end-to-end-encrypted (E2EE) and anonymized models. Priced highly competitively, can be 40-70% cheaper than Venice AI and other platforms. Includes post-processing such as combining videos and images, cutting, slicing, splicing, transitions, drawing text, re-encoding, resizing and much more!\nversion: 1.0.3\nauthor: arfonzo (imgnAI)\nlicense: MIT-0\nmetadata: {\"openclaw\": {\"requires\": {\"bins\": [\"curl\", \"python3\"]}, \"homepage\": \"https://app.imgnai.com\"}}\n---\n\n# Katana Skill — imgnAI API\n\nGenerate images, videos, and text/LLM completions via the [imgnAI Katana API](https://app.imgnai.com/katana-api). Supports end-to-end-encrypted (E2EE) and anonymized models. Priced highly competitively: can be 40-70% cheaper than Venice AI and other platforms.\n\nIncludes post-processing such as combining videos and images, cutting, slicing, splicing, transitions, drawing text, re-encoding, resizing and much more!\n\nA complete workflow for content creation from start to finish, all from the comfort of your agent.\n\n\n## Triggers\n\n\"generate image of X\", \"create image\", \"make picture\", \"imgnai image\", \"generate video of X\", \"create video\", \"make video\", \"ask grok about X\", \"ask claude about X\", \"use gpt to X\", \"katana image\", \"katana video\", \"katana chat\", \"katana gpt\", \"katana claude\", \"list katana models\", \"modify this image\", \"edit this image\", \"change this image\", \"transform this image\", \"edit image\", \"modify image\"\n\n## Spawn Policy\n\n**NEVER spawn subagents for katana operations by default.** All katana workflows (image generation, video generation, text completions, post-processing) MUST be executed inline in the current session.\n\n**Exception:** Only spawn if the user **explicitly requests** spawning in their prompt (e.g. \"spawn a subagent to handle this\", \"run this as a background task\"). Do NOT spawn based on AGENTS.md spawn rules or default agent behavior — user intent is the only trigger for spawning with katana.\n\nLLM-specific triggers (gpt, claude, etc) also respond to \"katana \\<model\\>\" to avoid conflicts with direct integrations.\n\n## Configuration\n\n### Data Retention\nHistorical prompts and results are retained for a maximum of **72 hours** after generation. Prompt/result history can be switched off from the API page at https://app.imgnai.com/katana-api.\n\n**HTTPS-only:** Public API calls must use HTTPS. If an integration sees an `http://` Katana base URL, replace it with `https://` before making calls.\n\n## Model IDs\n\nThe Katana API uses `model_key` as the model identifier, not `public_model_name`. When building requests, always use the model_key value. See `{baseDir}/models.md` for the full mapping.\n\n**Dual-key system:** The API supports both **canonical keys** (e.g. `gpt-image-2`) and **legacy keys** (e.g. `gpt2image`). Both work identically. This skill now uses **canonical keys** as the default for all workflows and aliases. Legacy keys are documented in the \"Model ID\" column of `models.md` for backward-compatibility reference. You may use either format when constructing API requests.\n\n## Model Discovery\n\n**Endpoint:** `GET /v1/models`\n**Auth:** `Authorization: Bearer ${KATANA_API_KEY}:${KATANA_API_SECRET}`\n\nReturns available models. Text models are returned for authenticated requests.\nFor the complete model catalogue including image/video, see models.md.\n\n**Usage:** Generally not needed before requests — use models.md as reference.\n\n---\n\n## Payment Methods\n\nThe API supports two payment methods:\n- **API key + secret** (Bearer auth) — used by this skill, preferred\n- **x402 micropayment** — NOT used by this skill\n\nNote: x402 text requests must be non-streaming. This skill only uses API-key auth.\n\n### Text/LLM Notes\n- **Streaming:** `stream: true` supported with SSE for API-key billing. x402 text calls must be non-streaming.\n- **Vision/multimodal:** Send images via `image_url` with Base64 data URLs or HTTPS URLs in messages. Base64 inputs are converted to JPEG, capped at 4096px max side.\n- **Billing:** Pre-charge reserve (10 credit minimum), refund after actual usage. 0.1 credit minimum charge rounded up.\n- **Refunds:** If an image or video generation fails after it was charged, credits or x402 balance will be refunded within 5 minutes, pending no Terms of Service violation.\n- **Default max_tokens:** If omitted and model supports output caps, API defaults to 16000.\n- **Privacy tiers:** Three tiers exist — `Anonymized` (customer identity not sent, model operator may process content), `Private` (private in-house model path, no E2EE/hardware attestation), and `E2EE Private` (hardware-protected confidential-compute, attestation via `GET /v1/text/attestation?model={model}&nonce={64_hex_nonce}`).\n- **`generation_timed_out` error:** When the API server-side timeout fires, poll returns terminal `failed`/`partial_failure` with `responses[].error.code = generation_timed_out`, `error.retryable: true`, `error.details.timeout_seconds`. Retry by submitting a new request.\n\n---\n\n- **API Base URL:** `https://kat.imgnai.com`\n- **API Reference:** https://kat.imgnai.com/llms.txt\n- **Model catalogue:** `{baseDir}/models.md`\n- **Skill directory:** Resolve dynamically from this file's location as `{baseDir}`. Most agent frameworks resolve this automatically.\n\n## Credentials\n\n**Secrets file:** Store your API key and secret in a file (default: `~/.openclaw/secrets/katana.env`):\n```\nKATANA_API_KEY=your_key_here\nKATANA_API_SECRET=your_secret_here\n```\nCreate with `chmod 600`. Get your credentials from https://app.imgnai.com/katana-api.\n\n**Loading:** All curl examples in this skill use `.` (dot) source to load credentials into the shell environment:\n```bash\n. \"${KATANA_SECRETS_FILE:-$HOME/.openclaw/secrets/katana.env}\"\n```\nOverride the default path with the `KATANA_SECRETS_FILE` environment variable.\n\n### ⚠️ Credential Security (MANDATORY)\n\n**NEVER display secrets in tool output.** The `.` source command loads credentials into shell variables silently — no output is produced. This is the correct and secure approach.\n\n**Banned patterns:**\n- `cat ~/.openclaw/secrets/katana.env`\n- `KATANA_API_KEY=kat_live_... curl ...`\n- Any form of reading secrets into tool output\n\n**If credential loading fails:** Fix the secrets file path or contents. Do NOT bypass security by hardcoding values.\n\n---\n\n## Optional Dependencies\n\nThese are not required for core API usage but enable additional features:\n\n| Binary | Needed for | Install |\n|--------|------------|--------|\n| `jq` | JSON parsing for API responses | `apt install jq` / `brew install jq` |\n| `python3` | Payload building, JSON parsing fallback | Pre-installed on most systems |\n| `ffmpeg` | Video post-processing (trim, join, effects) | `apt install ffmpeg` / `brew install ffmpeg` |\n\n`jq` or `python3` is needed for JSON parsing. Post-processing requires `ffmpeg`.\n\n---\n\n## Terminology\n\nThis skill uses some agent-specific terms. Here's what they mean regardless of your agent framework:\n\n| Term | Meaning |\n|------|--------|\n| **exec call / shell invocation** | A single shell command execution. Some agents execute each line of a multi-line script as separate invocations — hence the `&&` chaining requirement to keep everything in one shell. |\n| **tool-result-loss** | A situation where a command was executed but its output never arrives back — the result shows as empty or a synthetic error message. The command likely ran successfully but the result was lost in transit. |\n| **compaction** | When an agent's context window fills up, older messages may be summarised or removed to make room. State stored only in conversation history (not in files) is at risk of being lost during compaction. |\n| **heartbeat** | A periodic check-in cycle where the agent re-evaluates its state (e.g., checking if a pending generation has completed). |\n| **session reset** | The conversation is restarted or reloaded, losing any in-memory state. File-based persistence survives this. |\n\n---\n\n## ⚠️ MANDATORY ROUTING — DO NOT SKIP\n\n**Before ANY generation or post-processing request, you MUST load the correct workflow file:**\n\n| Task | Load this file |\n|------|---------------|\n| Image generation | `{baseDir}/workflows/image.md` |\n| Video generation | `{baseDir}/workflows/video.md` |\n| Text/LLM generation | `{baseDir}/workflows/text.md` |\n| Post-processing (ffmpeg, combine, text overlay, etc) | `{baseDir}/workflows/post-process.md` |\n\n**NEVER attempt a generation without loading the workflow file first.**\n**NEVER guess parameters — the workflow file has the exact steps.**\n\n---\n\n## Cost Reporting (ALL Requests)\n\n**After every generation (text, image, video), send a separate follow-up message with a cost summary.** Include all relevant details from the response:\n\n```\n📊 Katana Summary\nModel: gemma-4-26b-a4b (Anonymized)\nRequest: bf11cf04-8747-480e-a7f7-7d6cb092c614\nTokens: 42 in / 176 out (text only)\nCost: 0.1 credits (~$0.001)\nPrivacy: Anonymized\nTime: ~3s\n```\n\nFor image/video, replace tokens with dimensions/duration as relevant. Always compute cost in USD using the current credit rate (see `{baseDir}/models.md`).\n\n---\n\n## Model Aliases (Quick Reference)\n\n### Text/LLM\n\n| User says | API model ID |\n|---|---|\n| grok | `grok-4-3` |\n| gpt / gpt-5 | `gpt-5-5` |\n| claude / claude-opus | `claude-opus-4-8` |\n| claude-fast | `claude-opus-4-8-fast` |\n| claude-sonnet | `claude-sonnet-4-6` |\n| claude-haiku | `claude-haiku-4-5` |\n| naifu / q-naifu | `q-naifu-a3b` |\n\n### Image\n\n| User says | API model ID |\n|---|---|\n| default / imgnai | `gen` |\n| anime | `ani` |\n| gpt-image | `gpt-image-2` |\n| nano | `nano-banana-2` |\n| flux | `flux-2-pro` |\n| pink | `pink-image` |\n\n### Video\n\n| User says | API model ID |\n|---|---|\n| default / seedance | `seedance-2-0-fast` |\n| seedance-hd | `seedance-2-0` |\n| ltx | `ltx-2-3` |\n| kling | `kling-3-0-kling30` |\n| veo | `veo3-1` |\n\nIf the user specifies an exact model ID, pass it through directly. Full alias tables in `{baseDir}/models.md`.\n\n---\n\n## Pre-Submission Confirmation (MANDATORY)\n\nBefore submitting ANY generation request, present a summary (model, cost in credits AND dollars, details, prompt) and **wait for user confirmation**. See each workflow file for details.\n\n**NO EXCEPTIONS:** There is no urgency override. \"just do it\", \"generate now\", /katana, or any other shortcut does NOT skip confirmation. ALWAYS present summary and wait for explicit approval before submitting.\n\n---\n\n## Error Protocol\n\n**ONE-ATTEMPT RULE: Every paid API call gets exactly ONE attempt per turn. If the tool result is lost, missing, or empty after a submission — STOP. Report to the user that the result was lost. Wait for user confirmation before retrying. NEVER retry a paid API call silently, even if the result seems to have vanished.**\n\n**STRICT — NO SILENT RETRIES.** Every error stops. Every retry needs approval. Tool-result-loss (result never arrives, empty, or vanishes) is a hard-stop condition equal to a visible error. See each workflow file for details.\n\n- ANY error or tool-result-loss → STOP, report to user (what happened, credits charged, total across attempts)\n- Tool-result-loss (result shows 'missing tool result' or similar synthetic error) → the API call likely already succeeded. STOP. Report to user. Do NOT retry the same request.\n**Terminal submission responses:** If the submission response itself is terminal (`status: \"failed\"`, `status: \"rejected\"`, or all response items rejected) — do NOT poll. Report the returned `responses[].error` or top-level error to the user immediately.\n\n- **Upstream errors are terminal.** If the API returns `upstream_error` (404, 500, etc), do NOT try a different model, do NOT retry with different parameters, do NOT submit to another endpoint. STOP and report the error to the user. You MAY suggest recommended next steps or options (e.g. \"model X returned 404 — want me to try model Y instead?\"), but ANY proposed plan requires explicit user approval before execution.\n- Propose fix → wait for explicit user approval\n- Banned: automatic retries, debug/test requests, parameter changes without telling user, lying about call counts, silent retries on lost results\n\n### Error Codes Quick Reference\n\n| Error Code | Context | Meaning | Retryable | Action |\n|---|---|---|---|---|\n| `generation_timed_out` | Poll response | Server-side timeout during generation | Yes | Submit a **new** request (same request ID won't work) |\n| `upstream_error` | Any | Provider/upstream API error | No | Report to user; may suggest alternative model if approved |\n| Auth errors | Submission (401/403) | Invalid or missing credentials | No | Check secrets file path and contents |\n| `status: \"rejected\"` | Submission response | Validation failure (bad params, content policy) | No | Fix parameters per model spec; rephrase if content blocked |\n| `status: \"failed\"` | Submission or poll | Generation failed after dispatch | No | Credits refunded within 5 min unless ToS violation |\n| Rate limit (429) | Any | Too many requests | Yes (after delay) | Wait per `Retry-After` header, then retry |\n\n---\n\n## Concurrency Guard\n\n**NEVER submit a new request while any previous request is still processing.** One request in flight at a time — no exceptions.\n\n- Before submitting, verify no pending/processing requests exist\n- If a previous request is still running (poll returns incomplete), either wait for it, ask the user to cancel, or ask the user to approve submitting a concurrent request\n- This applies across ALL endpoints: text, image, and video\n\n---\n\n## Immediate Status Updates\n\nAfter submitting async generations (image/video), deliver a confirmation to the user BEFORE starting the poll loop. Include the model, cost, and request_id.\n\n## Async Polling\n\nImage and video generations are asynchronous. After submitting, poll manually.\n\n**Poll command:**\n```bash\n. \"${KATANA_SECRETS_FILE:-$HOME/.openclaw/secrets/katana.env}\" && _H=$(mktemp) && chmod 600 \"$_H\" && printf 'X-API-Key: %s\\nX-API-Secret: %s\\n' \"$KATANA_API_KEY\" \"$KATANA_API_SECRET\" > \"$_H\" && curl -s \"https://kat.imgnai.com/v1/generation-requests/${REQUEST_ID}\" -H @\"$_H\" && rm -f \"$_H\"\n```\n\n**Raw response:** Pipe to `jq '.'`.\n\n**Formatted:** Pipe to:\n```bash\npython3 -c \"\nimport sys,json\nd=json.load(sys.stdin)\nr=d.get('responses',[])\nfor i in range(len(r)):\n    ri=r[i]; st=ri.get('status','?')\n    print(f'Status: {st}')\n    for a in ri.get('output_assets',[]):\n        print(f'URL: {a.get(\"original_data_url\",\"\")}')\n        print(f'Dims: {a.get(\"width\",\"?\")}x{a.get(\"height\",\"?\")}')\n        print(f'Expires: {a.get(\"expires_at\",\"\")}')\n    print(f'Credits: {ri.get(\"metadata\",{}).get(\"credits_spent\",\"?\")}')\n    if st=='failed':\n        e=ri.get('error',{}); print(f'Error: {e.get(\"message\",\"\")} retryable={e.get(\"retryable\",\"\")}')\n\"\n```\n\n**`wait` parameter:** `wait=true` is available for convenience (blocks until complete), but production integrations should prefer polling with `wait=false` (the default).\n\n**Polling pattern:** Extract `poll_after_seconds` from the submission response and use it as the initial polling interval. If the poll response includes a new `poll_after_seconds`, use that for the next interval. Fall back to polling every 30 seconds for the first 5 minutes, then every 60 seconds if `poll_after_seconds` is absent or null.\n\n**Agent responsibility:** The agent decides how to schedule polls (intervals, background tasks, etc). Do not use long-running background processes — use single polls at intervals.\n\n### ⚠️ Polling Pattern Constraints\n\n**Keep `.` source and `curl` in the same command chain.** Shell `sleep` or `process poll` between commands breaks the env var loading — env vars are lost.\n\n**Correct:** Single shell invocation containing the full chain (see poll command above).\n\n**Wrong:** Separating `.` + `sleep` + `curl` into different shell invocations.\n\n**If your agent cannot chain commands:** Use the agent-native polling mechanism (background execution, process polling, etc) with the full command as one unit.\n\n**Response handling for completed polls:**\n- Extract `original_data_url` for delivery (full-resolution)\n- Extract dimensions from `responses[].output_assets[].width/height` (NOT from submission response)\n- Extract credits from `responses[].metadata.credits_spent`\n- Extract expiry from `responses[].output_assets[].expires_at` — display in user's local timezone in delivery summary\n\n### ⚠️ Poll Timeout Guard (10-minute hard stop for image/text; 100-minute for video)\n\nAfter 10 minutes of cumulative polling for image/text (or 100 minutes for video), STOP polling and inform the user:\n\n\"⚠️ Poll timeout: generation has been processing for [10/100] minutes. The API poll endpoint still says 'processing' but this may be stale — generations that time out (600s image/text, 6000s video) or get blocked by content safety often don't update the poll status. Check the Katana dashboard at https://app.imgnai.com/katana-api for the real status. Should I keep polling, or consider this failed?\"\n\nWAIT for user response before continuing:\n- If user says \"keep polling\" → resume polling (same guard applies)\n- If user says \"stop\" or \"failed\" → stop and report the situation\n- Do NOT silently continue past the timeout mark\n\nTrack cumulative poll time via wall-clock: record submission timestamp after confirmation, check elapsed time before each poll cycle.\n\n**API timeout values:** Image/text = 600s (10 min). Video = 6000s (100 min). The 10-minute guard is appropriate for image/text but too aggressive for video — video jobs may legitimately run for up to 100 minutes. Use 100-minute guard for video requests.\n\nThis guard exists because the API poll endpoint has been observed returning \"processing\" even after:\n- The generation timed out (600s/6000s upstream timeout)\n- The generation was blocked by content safety policy\n- Credits were already refunded\n\nThe only reliable source of truth for stale generations is the Katana dashboard.\n\n### ⚠️ Generation Persistence (compaction-safe tracking)\n\nAfter submitting any async generation, IMMEDIATELY write the request metadata to a persistence file. Use the same `KATANA_SECRETS_FILE` env var pattern for the path, defaulting to the secrets directory:\n\n```python\nimport json, datetime, os\nbase = os.environ.get('KATANA_STATE_DIR', os.path.dirname(os.environ.get('KATANA_SECRETS_FILE', os.path.expanduser('~/.openclaw/secrets/katana.env'))))\npath = os.path.join(base, 'katana_pending.json')\nmeta = {\n    'request_id': 'REQUEST_ID',\n    'model': 'MODEL',\n    'credits': CREDITS,\n    'submitted': datetime.datetime.now().isoformat(),\n    'prompt': 'PROMPT_SUMMARY',\n    'status': 'processing'\n}\nwith open(path, 'w') as f:\n    json.dump(meta, f)\nprint(f'written: {path}')\n```\n\nThis file survives compaction. On recovery (after compaction, after tool result loss, or at heartbeat), use the same path derivation:\n\n```python\nimport json, os\nbase = os.environ.get('KATANA_STATE_DIR', os.path.dirname(os.environ.get('KATANA_SECRETS_FILE', os.path.expanduser('~/.openclaw/secrets/katana.env'))))\npath = os.path.join(base, 'katana_pending.json')\nif os.path.exists(path):\n    with open(path) as f:\n        meta = json.load(f)\n    print(f'request_id={meta[\"request_id\"]} status={meta[\"status\"]} model={meta[\"model\"]}')\n```\n\nRecovery steps:\n1. Check if persistence file exists AND status is \"processing\"\n2. Resume polling from the saved request_id\n3. If completed → deliver result, delete persistence file\n4. If still processing → check elapsed time against 10-minute guard\n5. If failed → report to user, delete persistence file\n\nDelete the persistence file ONLY when the generation reaches a terminal state (completed, failed, delivered to user). Never delete while still processing.\n\nThis prevents the pattern where: agent submits → compaction happens → agent forgets → user has to ask for status manually.\n\n## Response Handling\n\n### Dimensions (IMPORTANT)\n1. **Submission response** (`requests[].width/height`) — PREVIEW dimensions, NOT actual output size.\n2. **Completed poll response** (`responses[].output_assets[].width/height`) — ACTUAL output dimensions.\n\n**Always report dimensions from the completed poll response, never from the submission acknowledgement.**\n\n### URL Fields (IMPORTANT)\n- **`original_data_url`** — full-resolution original. **Always use this for delivery.**\n- **`url`** — may be a compressed/reduced version. Do NOT use for delivery.\n- **`thumbnail_image_url`** — small thumbnail only.\n- **`thumbnail_silent_video_mp4_url`** — silent lightweight MP4 preview for video galleries/hover previews. This is just the thumbnail preview — the full video output (via `original_data_url`) may include generated audio. NOT the full video.\n- **`final_frame_image_url`** — last frame still image for completed videos. Use as first-frame input for video continuation workflows. Blank string when unavailable.\n\n### CLIP Tag Metadata\n`responses[].output_assets[].metadata.tags` contains CLIP-derived tags with confidence scores (e.g. `{\"tag\": \"ceramic_mug\", \"confidence\": 0.94}`). Only available on in-house imgnAI models — external/provider-hosted models return no CLIP-tag metadata.\n\n### Model Normalization\nCompleted media responses may normalize `requests[].model` and `responses[].metadata.model` (e.g. legacy key → canonical key). Use `GET /v1/models` for canonical display names.\n\n### Item Timestamps\n- `responses[].started_at` — item-level processing start timestamp\n- `responses[].completed_at` — item-level processing end timestamp\n- `created_at` — top-level request submission timestamp\n- `updated_at` — top-level request last-modified timestamp\n- Useful for tracking actual generation time per item\n\n### Asset Type Fields\n- `responses[].output_assets[].kind` — asset type (e.g. `\"image\"`, `\"video\"`)\n- `responses[].output_assets[].mime_type` — MIME type (e.g. `\"image/png\"`, `\"video/mp4\"`)\n\n### ⚠️ Anti-Pattern Warning\nData is under `responses[].output_assets[]` — do NOT look for `results[].url`. That is NOT the Katana response shape.\n\n### ⚠️ `output` Object\nDo NOT send an `output` object for ordinary integrations. This is for internal/special use only.\n\n## Payload Submission\n\nBuild the JSON payload in a temp file (required for large payloads and to avoid secrets in process listings):\n\n```python\nimport json, tempfile\npayload = {\"requests\": [{\"type\": \"video\", \"model\": \"seedance-2-0-fast\", \"prompt\": \"<prompt>\", \"duration_seconds\": 5, \"aspect_ratio\": \"16:9\"}]}\nwith tempfile.NamedTemporaryFile(mode=\"w\", suffix=\".json\", delete=False) as f:\n    json.dump(payload, f)\n    tmpfile = f.name\nprint(tmpfile)\n```\n\n### Secure Header Pattern\n\nWrite auth headers to a temp file to keep secrets out of `/proc/*/cmdline`. Source credentials at the start of each command chain.\n\n**Image/Video requests** (X-API-Key + X-API-Secret):\n```bash\n. \"${KATANA_SECRETS_FILE:-$HOME/.openclaw/secrets/katana.env}\" && _H=$(mktemp) && chmod 600 \"$_H\" && printf 'Content-Type: application/json\\nX-API-Key: %s\\nX-API-Secret: %s\\n' \"$KATANA_API_KEY\" \"$KATANA_API_SECRET\" > \"$_H\" && curl -s -X POST \"https://kat.imgnai.com/v1/generation-requests?wait=false\" -H @\"$_H\" -d @\"$tmpfile\" && rm -f \"$_H\" && rm -f \"$tmpfile\"\n```\n\nThe `printf` format string writes **two separate header lines**: `X-API-Key` with the key value, and `X-API-Secret` with the secret value. Two `%s` format specifiers consume the two shell variable arguments. No extra literal text appears in the header values.\n\n**Text/LLM requests** (Bearer auth):\n```bash\n. \"${KATANA_SECRETS_FILE:-$HOME/.openclaw/secrets/katana.env}\" && _H=$(mktemp) && chmod 600 \"$_H\" && printf 'Content-Type: application/json\\nAuthorization: Bearer %s:%s\\n' \"$KATANA_API_KEY\" \"$KATANA_API_SECRET\" > \"$_H\" && curl -s -X POST \"https://kat.imgnai.com/v1/chat/completions\" -H @\"$_H\" -d @\"$tmpfile\" && rm -f \"$_H\" && rm -f \"$tmpfile\"\n```\n\nParse the JSON response. Extract `request_id`. Deliver confirmation to the user (model, cost, request_id).\n\n### Image Generation Notes\n- **`aspect_ratio: \"auto\"`** inspects the first image input and chooses closest supported ratio. Defaults to `1:1` if no image supplied.\n- **`is_fast`/`fast_mode`** request lower-cost half-resolution generation (imgnAI-hosted models only)\n- **`is_uhd`/`uhd_mode`** request UHD generation (imgnAI-hosted models only). Takes precedence over `is_fast`. On Pink Image, this is \"Enhanced Quality\" mode.\n- **`use_assistant`/`prompt_assist`** translate natural language to tag-style prompts (tag/booru models only)\n- **`output_format`** accepts `png`, `jpeg`, or `webp` (`jpg` is alias for `jpeg`)\n\n---\n\n## Credit Balance\n\n```bash\n. \"${KATANA_SECRETS_FILE:-$HOME/.openclaw/secrets/katana.env}\" && _H=$(mktemp) && chmod 600 \"$_H\" && printf 'Authorization: Bearer %s:%s\\n' \"$KATANA_API_KEY\" \"$KATANA_API_SECRET\" > \"$_H\" && curl -s \"https://kat.imgnai.com/v1/me/balance\" -H @\"$_H\" && rm -f \"$_H\"\n```\n\nCalls `GET /v1/me/balance`. The API returns `credits` as a decimal string. Converts to USD using current credit rate (see `{baseDir}/models.md`).\n\n---\n\n## Video Media Input Rules\n\nPut video media inputs in `video_image_data`:\n\n- **`first_frame_image_url`**: first/source frame image (HTTPS URL, data URL, or raw base64)\n- **`mid_frame_image_url`**: mid-frame image (only if model supports it)\n- **`last_frame_image_url`**: last/end frame image (only if model supports it)\n- **`reference_image_urls`**: array of reference images (only if model supports reference images). Obey `maximum_reference_images`\n- **`audio_input_urls`**: array of audio reference URLs (only if model supports audio input). Obey `maximum_reference_audio_files` and global cap of 4\n- **`video_list`**: array of video input clip objects. Each requires `url`. Optional `start`/`ends` second offsets when model has `video_offset_allowed`. Only for models with `supports_video_input: true`. Obey `maximum_reference_videos`.\n- **Audio output:** Most modern video models generate audio by default — the model interprets the prompt for contextual sound design (speech, music, effects, ambient). Legacy models with `audio_gen_model: false` produce silent output. See `{baseDir}/models.md` Audio Out column and `{baseDir}/workflows/video.md` Audio Output section for details.\n\nCompatibility aliases: top-level `image_url`, `input_image_url`, `input_image`, `input_image_b64` map to `first_frame_image_url`. Top-level `reference_image_urls` maps to `video_image_data.reference_image_urls`.\n\n**Rules:**\n- Do NOT send local filesystem paths, `file://` URLs, or `http://` media URLs — use HTTPS URLs or Base64 data URLs\n- Do NOT send `video_list` to models where `supports_video_input` is false/missing\n- Do NOT send audio references to models where `supports_audio_input` is false\n- Do NOT create an empty `video_image_data` object — omit missing fields entirely\n- Do NOT mix first/last frame with reference images unless the model's `custom_rules` allow it\n- Use only durations from `video_lengths_and_costs` and aspect ratios from `supported_aspects`\n- `aspect_ratio: \"auto\"` uses the first frame first, then first reference image; defaults to `1:1`\n\n## Video Custom Rules Glossary\n\nAlways inspect each video model's `custom_rules` before composing requests:\n\n| Rule | Description |\n|------|-------------|\n| `audio_15s_max` | Combined audio input limited to 15 seconds |\n| `audio_drives_duration` | Video duration follows audio duration |\n| `audio_ff_only` | Audio only with first-frame conditioning |\n| `audio_needs_reference_image` | Audio input requires at least one reference image |\n| `audio_or_fflf_exclusive` | Audio cannot combine with first/last frame |\n| `input_video_drives_length` | Input video clip drives output length |\n| `lf_needs_ff` | Last frame requires a first frame |\n| `reference_ff_only` | Reference images may combine with first-frame only |\n| `reference_is_voice_timbre` | Reference audio interpreted as voice timbre when images present |\n| `reference_no_ff_or_lf` | Reference images cannot combine with first/last frame |\n| `video_offset_allowed` | Model accepts `start`/`ends` second offsets in `video_list` |\n| `video_required` | Model requires at least one `video_list` object |\n\n## Common Failure Cases\n\n- **Unsupported aspect ratio:** Choose from model's `supported_aspects`\n- **Unsupported duration:** Choose from model's `video_lengths_and_costs`\n- **Corrupt base64 image:** Validate data decodes to actual image before submitting\n- **Local file path as media:** Convert to Base64 data URL first — API cannot fetch caller's filesystem\n- **Audio on unsupported model:** Only send `audio_input_urls` when `supports_audio_input: true`\n- **Video metadata on unsupported model:** Only send `video_list` when model has matching support flag\n- **Missing required video input:** Models with `video_required` must include `video_list`\n- **Video offsets on unsupported model:** Only send `start` and `ends` in `video_list` objects when the model has `video_offset_allowed` in `custom_rules`\n- **Last frame without first frame:** Prohibited on models with `lf_needs_ff`\n- **Reference images mixed with frames on incompatible model:** Check `custom_rules`, especially `reference_no_ff_or_lf`\n- **Too many reference images:** Clamp to `multi_image_inputs_allowed` (image) or `maximum_reference_images` (video)\n\n---\n\n## reference_assets (Typed Asset System)\n\n`reference_assets` is an alternative to `image_urls`/`video_image_data` for providing media inputs with explicit role labels. Each asset has a `kind` and either `url` or `base64_data`.\n\n### Image models\n\nAccepted image-like asset kinds:\n- `source_image` — primary source/input image\n- `image` — generic image input\n- `mask` — mask for inpainting/editing\n- `style_reference` — style transfer reference\n- `start_frame` — starting frame for animation\n\nExample:\n```json\n{\n  \"reference_assets\": [\n    {\"kind\": \"source_image\", \"url\": \"https://example.com/product.png\"},\n    {\"kind\": \"style_reference\", \"base64_data\": \"data:image/jpeg;base64,...\"}\n  ]\n}\n```\n\n### Video models\n\nImage kinds for video:\n- `style_reference`, `reference_image`, `image` — map to video reference images\n\nAudio kinds for video:\n- `audio`, `source_audio`, `reference_audio`, `audio_reference` — map to audio reference inputs\n\nExample:\n```json\n{\n  \"reference_assets\": [\n    {\"kind\": \"reference_image\", \"url\": \"https://example.com/person.png\"},\n    {\"kind\": \"audio\", \"url\": \"https://example.com/voice.mp3\"}\n  ]\n}\n```\n\n---\n\n## llms.txt Freshness\n\nThis skill was built from the Katana API llms.txt reference document.\n\n**Last synced:** 2026-06-08\n**llms.txt URL:** https://kat.imgnai.com/llms.txt\n**Stored checksum:** `09a695f3958a6d9f17d4139179e2323600292c929be5c494253ae7df9d1410b3`\n\n### Pre-generation check\n\nBefore submitting ANY generation request, check if the llms.txt checksum has been verified in the last 24 hours. If stale:\n\n1. Fetch: `curl -s https://kat.imgnai.com/llms.txt`\n2. Compute SHA256: `sha256sum` (Linux) or `shasum -a 256` (macOS)\n3. Compare to stored checksum\n4. If CHANGED → tell the user: \"The Katana API model list has been updated since this skill was last synced. This may include new models, pricing changes, or removed models. Would you like me to check for changes and update the skill?\"\n5. If user says YES → parse new llms.txt, update models.md, update checksum and date\n6. If user says NO → proceed with current models\n7. Update last-checked date regardless\n\n### llms.txt update process\n\nWhen llms.txt changes, compare old vs new **holistically**. Diff the full documents — do not limit the review to a predefined checklist. Document ALL changes found and update all affected skill files accordingly: `models.md`, `SKILL.md`, workflow files.\n\n**DO NOT auto-update without user confirmation.**\n\n**Explicit approval rule:** During the llms.txt update process, always summarise ALL changes found and ask the user for explicit permission before updating any skill files (models.md, SKILL.md, workflow files). Do not auto-update without confirmation.\n\n---\n\n## Delivery Patterns\n\nDeliver the generated media to the user via your agent's messaging/file capability. Include: model name, resolution/dimensions, credits, dollar cost, description, and the full-res URL (`original_data_url`).\n\n### ⚠️ URL Display (MANDATORY)\n\nALL image and video deliveries MUST include the **full download URL** (`original_data_url`) as clickable text in the delivery message — not just the inline media attachment.\n\nUsers need the URL to:\n- Download the full-resolution file\n- Share it externally\n- Archive it before expiry\n\n**Include ALL URLs returned** — `original_data_url`, `thumbnail_image_url`, `final_frame_image_url`, `thumbnail_silent_video_mp4_url` — any URL the API returns for the asset. Do not assume the user only wants one.\n\nExample:\n```\nMEDIA:https://k.imgnai.com/abc123.mp4\n\n🔗 Full-res: https://k.imgnai.com/abc123.mp4\n🖼️ Thumbnail: https://k.imgnai.com/def456.jpg\n🎞️ Silent preview: https://k.imgnai.com/ghi789.mp4\n⏰ Expires: Fri 16 May 2026, 14:00 BST\n```\n\n### ⚠️ Expiry Warning (MANDATORY)\n\nALL image and video generation summaries MUST include:\n1. The **expiry timestamp** extracted from `responses[].output_assets[].expires_at` in the completed poll response — convert to user's local timezone for display\n2. A clear warning that content must be downloaded before expiry if the user wishes to keep it\n\nExample format:\n```\n⏰ Expires: Fri 16 May 2026, 14:00 BST — download before expiry if you need it long-term.\n```\n\n**Do NOT calculate expiry manually.** The API provides `expires_at` in the poll response. Use it directly. The 72h retention window may change server-side; `expires_at` is always authoritative.\n\nFor text/LLM: return the model's response verbatim. Then send a separate follow-up message with a cost summary per the \"Cost Reporting\" section above. Text completions do not require an expiry warning (no media URL to expire).\n\n---\n\nFile v1.0.3:README.md\n\n# Katana Skill — imgnAI API\n\nGenerate images, videos, and text/LLM completions via the [imgnAI Katana API](https://app.imgnai.com/katana-api). Supports end-to-end-encrypted (E2EE) and anonymized models. Priced highly competitively: can be 40-70% cheaper than Venice AI and other platforms.\n\nIncludes post-processing such as combining videos and images, cutting, slicing, splicing, transitions, drawing text, re-encoding, resizing and much more!\n\nA complete workflow for content creation from start to finish, all from the comfort of your agent.\n\n\n## Features\n\n- **Images:** 40+ models including GPT Image 2, imgnAI Gen, FLUX.2, Seedream, WAN, Imagine Art\n- **Videos:** 15+ models including Seedance 2.0, Kling O3 4K, WAN 2.7, Veo 3.1, LTX\n- **Text/LLM:** 30+ models including Grok 4.3, GPT-5.5, Claude Opus 4.7, DeepSeek V4, Qwen 3.6\n- **Pre-submission confirmation** with cost estimates before generating\n- **Adaptive polling** that respects API-recommended intervals\n- **Error handling** with clear user-facing messages\n- **Reference image support** for editing workflows\n- **Agent-agnostic:** works with OpenClaw, Hermes, Claude, or standalone\n\n## Setup\n\n1. Get your API key from https://app.imgnai.com/katana-api\n2. Create the secrets file:\n   ```bash\n   mkdir -p ~/.openclaw/secrets\n   cat > ~/.openclaw/secrets/katana.env << 'EOF'\n   KATANA_API_KEY=your_key_here\n   KATANA_API_SECRET=your_secret_here\n   EOF\n   chmod 600 ~/.openclaw/secrets/katana.env\n   ```\n\n   **Non-OpenClaw users:** Set `KATANA_SECRETS_FILE` to your preferred location:\n   ```bash\n   mkdir -p ~/.config/katana\n   cat > ~/.config/katana/katana.env << 'EOF'\n   KATANA_API_KEY=your_key_here\n   KATANA_API_SECRET=your_secret_here\n   EOF\n   chmod 600 ~/.config/katana/katana.env\n   export KATANA_SECRETS_FILE=~/.config/katana/katana.env\n   ```\n\n3. Install the skill (see [Agent Integration](#agent-integration) or [Standalone Usage](#standalone-usage) below)\n\n## Agent Integration\n\nThis skill works with any agent framework. It provides a `SKILL.md` routing hub and workflow files that guide your agent through image, video, and text generation.\n\n### OpenClaw (example)\n\n1. Follow the [Setup](#setup) steps above\n2. Install the skill:\n   ```bash\n   openclaw skill install katana\n   ```\n   Or manually: clone/copy the `katana/` directory into your skills path (typically `~/.openclaw/skills/` or `~/workspace/skills/`).\n3. Ask your assistant to generate: \"Generate an image of a cat riding a skateboard\"\n\n### Other Agent Frameworks\n\nPoint your agent to `SKILL.md` as the entry point. The skill resolves `{baseDir}` dynamically from the file's location. Set `KATANA_SECRETS_FILE` if your secrets are stored elsewhere.\n\n## Prerequisites\n\n- **curl** — API requests\n- **jq** or **python3** — JSON parsing (jq preferred, python3 as fallback)\n- **ffmpeg** (optional) — Post-processing\n\n### ffmpeg optional capabilities\n\nIf you install ffmpeg for post-processing:\n- **Text overlays / drawtext:** requires ffmpeg built with `--enable-libfreetype`. Most package managers include this. Test with: `ffmpeg -filters 2>/dev/null | grep drawtext`\n- **H.264 encoding:** requires `libx264`. Nearly all ffmpeg builds include this.\n- **Audio encoding:** requires `libfdk_aac` or built-in AAC encoder. The built-in encoder (`aac`) is sufficient.\n\nOn Ubuntu/Debian: `sudo apt install ffmpeg` includes all of the above.\nOn macOS: `brew install ffmpeg` includes all of the above.\n\n## Files\n\n- `SKILL.md` — Routing hub (triggers, setup, mandatory routing table, delivery patterns)\n- `workflows/image.md` — Image generation workflow\n- `workflows/video.md` — Video generation workflow\n- `workflows/text.md` — Text/LLM/E2EE chat workflow\n- `workflows/post-process.md` — FFmpeg post-processing workflow\n- `workflows/ffmpeg.md` — FFmpeg command reference for post-processing\n- `models.md` — Complete model catalogue with pricing (single source of truth)\n- `README.md` — This file\n\n## License\n\nMIT-0 (MIT No Attribution)\n\n## Author\n\narfonzo (imgnAI)\n\nFile v1.0.3:_meta.json\n\n{\n  \"ownerId\": \"kn75xgwvf577we4f2ej81pg3hx86p4c7\",\n  \"slug\": \"katana\",\n  \"version\": \"1.0.3\",\n  \"publishedAt\": 1781008234455\n}\n\nFile v1.0.3:CHANGELOG.md\n\n# Changelog\n\n## 1.0.3\n\n### Added\n- `q-naifu-a3b` text model — imgnAI fully uncensored Agentic Model (Private tier, 262K context, vision + file input)\n- `naifu` / `q-naifu` model alias mapping to `q-naifu-a3b`\n- Text polling for long-running models — submit with `?wait=false`, poll via `GET /v1/generation-requests/{id}`\n- Audio output docs for video models — which models generate audio, how to strip it\n- Three-tier privacy documentation — Anonymized, Private, E2EE Private with attestation details\n- Error codes quick reference table\n- Agent-agnostic terminology glossary\n- Non-OpenClaw setup guide in README\n\n### Changed\n- Claude alias now points to `claude-opus-4-8`; added `claude-fast` for `claude-opus-4-8-fast`\n- Poll timeout guard: 10 min for image/text, 100 min for video (matches API limits)\n- Persistence file supports `KATANA_STATE_DIR` to separate state from credentials\n- Cache write column formatting fixed for models with no cache write pricing\n\n### Fixed\n- Text submission auth: curl commands now use single `&&`-chained line (fixes failures on some agents)\n- Payload temp files are now cleaned up after each request\n- Header auth bug: submit and poll now use consistent two-header format\n- `gpt-image-2-max` reference image count: 12 → 10\n- `***` literal removed from header format strings\n\n## 1.0.2\n\n### Added\n- `grok-build-0-1` — Grok Build 0.1, xAI coding model (256K ctx)\n- `claude-opus-4-8` and `claude-opus-4-8-fast` text models\n- Flux 1.1 Ultra, Flux Kontext Max/Pro, GPT-5.4/5.5, Claude Opus 4.7/Sonnet 4.6/Haiku 4.5, Grok 4.20/4.20 Multi-Agent, DeepSeek V4 Flash/Pro\n- Video media input rules (`video_image_data` fields documented)\n- Video custom rules glossary (12 rules from API)\n- Common failure cases section\n- Text/LLM notes: streaming, vision/multimodal, billing, attestation, refund policy\n- Image generation notes: auto aspect ratio, fast/UHD modes, prompt assist\n- `thumbnail_silent_video_mp4_url` and `final_frame_image_url` to response handling\n- Cache read pricing for 11 models (8 private + qwen3-6-flash, qwen3-6-max-preview, qwen3-6-plus, qwen3-7-max)\n\n### Changed\n- Full models.md rebuild with canonical dashed keys throughout\n- All model keys migrated to canonical format (e.g. `seedance2` → `seedance-2-0`)\n- Price cuts: qwen3-7-max (-50%), deepseek-v4-flash (-30%), qwen3-6-flash (-25%), glm-5-1 (-12%), kimi-k2-6 (-9%), minimax-m2-7 (-13%), qwen3-6-35b-a3b (-7%)\n\n## 1.0.1\n\n### Added\n- `pink-image` model — 1 credit, high-speed generalist\n- `qwen3-7-max` text model — flagship Qwen, 1M context\n- `gemini-3-5-flash` text model — near-Pro at Flash cost, multimodal\n- `gpt-image-2-max` image model — MAX variant, 28 cr, QHD output\n- `gemini-omni` video model — Google video, 4-10s, 5 ref images\n- `gemini-omni-v2v` video model — V2V with `video_list` input\n- Text alias `qwen-max`, `gemini-35-flash`; video aliases `gemini-omni`, `gemini-v2v`\n- V2V (video-to-video) workflow section in `workflows/video.md`\n- Custom rules: `video_required`, `video_offset_allowed`, `input_video_drives_length`\n- `video_image_data.video_list` parameter for V2V models\n- Text workflow note: API defaults to 16K max_tokens when omitted\n- Video continuation workflow (chain videos from last frame)\n- Streaming, E2EE, and audio input support documented\n- Image editing trigger phrases\n\n### Changed\n- Removed `katana.sh` — skill is now pure markdown with inline curl patterns\n- Migrated to canonical model keys (e.g. `gpt-image-2` instead of `gpt2image`)\n- DRY cleanup — pricing, aliases, and protocols in single source files\n- Uniform `.` source credential loading — single method for all agent platforms\n- Security: concurrency guard (one request in flight, concurrent needs approval)\n- Security: upstream errors are terminal with suggest/approve flow\n- Security: credential security rules — never display secrets in tool output\n- `python3` added to required bins\n- `image_urls` vs `reference_assets` warning — `reference_assets` silently ignored on image requests\n- ffmpeg `-c copy` compat warning — stream copy preserves source codec, may break playback\n- Pre-publish DRY: removed duplicate alias table from `workflows/text.md`, deduplicated pre-submission paragraphs\n- `seedance-lite` canonical key fix (was `seedance-lite-seedancelite`)\n\n### Removed\n- `bash` requirement (only `curl` needed now)\n- Duplicated alias tables from workflow files\n- `katana.sh` helper script (VirusTotal false positive)\n\n## 1.0.0\n\n### Added\n- Initial public release\n\nFile v1.0.3:models.md\n\n# Katana Model Catalogue\n\n> **Production API uses `model_key` as the model ID parameter, not `public_model_name`. All IDs below are model_key values.**\n\n> **Note:** This is a static snapshot synced from the live API reference at https://kat.imgnai.com/llms.txt. For the most current pricing and model availability, check the live endpoint.\n\nAuto-generated from https://kat.imgnai.com/llms.txt — Last synced: 2026-06-08\n\nExtracted from the imgnAI Katana API docs. Organised by type: text, image, video.\n\nReference price: $0.0052 per credit (Platinum Annual).\n\n## Text / LLM Models\n\nAll text models use `POST /v1/chat/completions` (OpenAI-compatible).\nAuth: `Authorization: Bearer <api_key>:<api_secret>`\nBilling: pre-charge reserve, then refund/charge difference from actual usage. Minimum 0.1 credits.\n\n| Model ID | Publisher | Context | Max Output | Input Types | Privacy | In (cr/$) | Out (cr/$) | Cache R / W | Legacy |\n|---|---|---|---|---|---|---|---|---|---|---|\n| `q-naifu-a3b` | imgnAI | 262144 | 262144 | text, image, file | Private | 38.5 / $0.20 | 240.4 / $1.25 | — | — |\n| `claude-opus-4-8` | Anthropic | 1000000 | 128000 | text, image, file | Anonymized | 1000.0 / $5.20 | 5000.0 / $26.00 | 100.0 cr ($0.52) / 1250.0 cr ($6.50) | — |\n| `claude-opus-4-8-fast` | Anthropic | 1000000 | 128000 | text, image, file | Anonymized | 2000.0 / $10.40 | 10000.0 / $52.00 | 200.0 cr ($1.04) / 2500.0 cr ($13.00) | — |\n| `qwen3-7-max` | Qwen | 1000000 | 65536 | text | Anonymized | 264.5 / $1.38 | 793.3 / $4.13 | 52.9 cr ($0.2750) / 330.6 cr ($1.72) | — |\n| `grok-build-0-1` | xAI | 256000 | not listed | text, image | Anonymized | 192.4 / $1.00 | 384.7 / $2.00 | 38.5 cr ($0.20) / — | — |\n| `gemini-3-5-flash` | Google | 1048576 | 65536 | text, image, video, file, audio | Anonymized | 317.4 / $1.65 | 1903.9 / $9.90 | 31.8 cr ($0.1650) / 17.7 cr ($0.0917) | — |\n| `grok-4-3` | xAI | 1000000 | not listed | text, image | Anonymized | 264.5 / $1.38 | 528.9 / $2.75 | 42.4 cr ($0.2200) / — | — |\n| `qwen3-6-35b-a3b` | Qwen | 262144 | 262140 | text, image, video | Anonymized | 29.7 / $0.1540 | 211.6 / $1.10 | — | — |\n| `qwen3-6-flash` | Qwen | 1000000 | 65536 | text, image, video | Anonymized | 39.7 / $0.2062 | 238.0 / $1.24 | 4.0 cr ($0.0206) / 49.6 cr ($0.2578) | — |\n| `qwen3-6-max-preview` | Qwen | 262144 | 65536 | text | Anonymized | 220.0 / $1.14 | 1320.0 / $6.86 | 22.0 cr ($0.1144) / 275.0 cr ($1.43) | — |\n| `deepseek-v4-flash` | DeepSeek | 1048576 | 384000 | text | Anonymized | 20.8 / $0.1081 | 41.6 / $0.2163 | 4.2 cr ($0.0217) / — | — |\n| `deepseek-v4-pro` | DeepSeek | 1048576 | 384000 | text | Anonymized | 92.1 / $0.4785 | 184.1 / $0.9570 | 0.8 cr ($0.0040) / — | — |\n| `gpt-5-5` | OpenAI | 1050000 | 128000 | file, image, text | Anonymized | 1057.7 / $5.50 | 6346.2 / $33.00 | 105.8 cr ($0.5500) / — | — |\n| `kimi-k2-6-private` | MoonshotAI | 262144 | 262144 | text, image | E2EE Private | 230.6 / $1.20 | 973.1 / $5.06 | 78.3 cr ($0.4070) | — |\n| `qwen3-coder-next-private` | Qwen | 262144 | 262144 | text | E2EE Private | 38.1 / $0.1980 | 253.9 / $1.32 | — | — |\n| `glm-5-1-private` | Z.ai | 202752 | 202752 | text | E2EE Private | 256.0 / $1.33 | 888.5 / $4.62 | 127.0 cr ($0.6600) | — |\n| `kimi-k2-6` | MoonshotAI | 262144 | 16384 | text, image | Anonymized | 144.7 / $0.7524 | 723.5 / $3.76 | 30.5 cr ($0.1584) / — | — |\n| `mimo-v2-flash-private` | Xiaomi | 262144 | 262144 | text | E2EE Private | 21.2 / $0.1100 | 63.5 / $0.3300 | — | — |\n| `claude-opus-4-7` | Anthropic | 1000000 | 128000 | text, image | Anonymized | 1057.7 / $5.50 | 5288.5 / $27.50 | 105.8 cr ($0.5500) / 1322.2 cr ($6.88) | — |\n| `glm-5-1` | Z.ai | 202752 | 65535 | text | Anonymized | 207.4 / $1.08 | 651.6 / $3.39 | 38.5 cr ($0.2002) / — | — |\n| `gemma-4-26b-a4b` | Google | 262144 | not listed | image, text, video | Anonymized | 12.7 / $0.0660 | 69.9 / $0.3630 | — | — |\n| `gemma-4-31b` | Google | 262144 | 16384 | image, text, video | Anonymized | 25.4 / $0.1320 | 78.3 / $0.4070 | — | — |\n| `qwen3-6-plus` | Qwen | 1000000 | 65536 | text, image, video | Anonymized | 68.8 / $0.3575 | 412.5 / $2.15 | 6.9 cr ($0.0358) / 86.0 cr ($0.4469) | — |\n| `grok-4-20` | xAI | 2000000 | not listed | text, image, file | Anonymized | 264.5 / $1.38 | 528.9 / $2.75 | 42.4 cr ($0.2200) / — | — |\n| `grok-4-20-multi-agent` | xAI | 2000000 | not listed | text, image, file | Anonymized | 423.1 / $2.20 | 1269.3 / $6.60 | 42.4 cr ($0.2200) / — | — |\n| `minimax-m2-7` | MiniMax | 196608 | 131072 | text | Anonymized | 55.0 / $0.2860 | 253.9 / $1.32 | — | — |\n| `gpt-5-4-mini` | OpenAI | 400000 | 128000 | file, image, text | Anonymized | 158.7 / $0.8250 | 952.0 / $4.95 | 15.9 cr ($0.0825) / — | — |\n| `glm-5-turbo` | Z.ai | 202752 | 131072 | text | Anonymized | 253.9 / $1.32 | 846.2 / $4.40 | 50.8 cr ($0.2640) / — | — |\n| `qwen3-5-27b-private` | Qwen | 262144 | 262144 | text, image, video | E2EE Private | 63.5 / $0.3300 | 507.7 / $2.64 | — | — |\n| `gpt-5-4` | OpenAI | 1050000 | 128000 | text, image, file | Anonymized | 528.9 / $2.75 | 3173.1 / $16.50 | 52.9 cr ($0.2750) / — | — |\n| `gemini-3-1-flash-lite-preview` | Google | 1048576 | 65536 | text, image, video, file, audio | Anonymized | 52.9 / $0.2750 | 317.4 / $1.65 | 5.3 cr ($0.0275) / 17.7 cr ($0.0917) | — |\n| `qwen3-5-397b-a17b-private` | Qwen | 262144 | 262144 | text, image, video | E2EE Private | 116.4 / $0.6050 | 740.4 / $3.85 | 47.6 cr ($0.2475) | — |\n| `minimax-m2-5-private` | MiniMax | 196608 | 196608 | text | E2EE Private | 42.4 / $0.2200 | 292.0 / $1.52 | 15.9 cr ($0.0825) | — |\n| `gemini-3-1-pro-preview` | Google | 1048576 | 65536 | audio, file, image, text, video | Anonymized | 423.1 / $2.20 | 2538.5 / $13.20 | 42.4 cr ($0.2200) / 79.4 cr ($0.4125) | — |\n| `claude-sonnet-4-6` | Anthropic | 1000000 | 128000 | text, image | Anonymized | 634.7 / $3.30 | 3173.1 / $16.50 | 63.5 cr ($0.3300) / 793.3 cr ($4.13) | — |\n| `glm-5` | Z.ai | 202752 | not listed | text | Anonymized | 127.0 / $0.6600 | 406.2 / $2.11 | 25.4 cr ($0.1320) / — | — |\n| `glm-5-private` | Z.ai | 202752 | 202752 | text | E2EE Private | 253.9 / $1.32 | 740.4 / $3.85 | 100.5 cr ($0.5225) | — |\n| `claude-opus-4-6` | Anthropic | 1000000 | 128000 | text, image | Anonymized | 1057.7 / $5.50 | 5288.5 / $27.50 | 105.8 cr ($0.5500) / 1322.2 cr ($6.88) | — |\n| `kimi-k2-5-private` | MoonshotAI | 262144 | 262144 | text, image | E2EE Private | 127.0 / $0.6600 | 634.7 / $3.30 | 46.6 cr ($0.2420) | — |\n| `gemini-3-flash-preview` | Google | 1048576 | 65536 | text, image, file, audio, video | Anonymized | 105.8 / $0.5500 | 634.7 / $3.30 | 10.6 cr ($0.0550) / 17.7 cr ($0.0917) | — |\n| `deepseek-v3-2-private` | DeepSeek | 163840 | 163840 | text | E2EE Private | 67.7 / $0.3520 | 101.6 / $0.5280 | 29.7 cr ($0.1540) | — |\n| `qwen3-coder-480b-a35b-private` | Qwen | 262000 | 262000 | text | E2EE Private | 423.1 / $2.20 | 423.1 / $2.20 | — | — |\n| `qwen3-vl-30b-a3b-instruct-private` | Qwen | 128000 | 128000 | text, image | E2EE Private | 42.4 / $0.2200 | 148.1 / $0.7700 | — | — |\n| `claude-haiku-4-5` | Anthropic | 200000 | 64000 | image, text | Anonymized | 211.6 / $1.10 | 1057.7 / $5.50 | 21.2 cr ($0.1100) / 264.5 cr ($1.38) | — |\n| `gemma-3-27b-private` | Google | 53920 | 53920 | text, image | E2EE Private | 23.3 / $0.1210 | 84.7 / $0.4400 | — | — |\n\n## Image Models\n\nAll image models use `POST /v1/generation-requests` with `type: \"image\"`.\nRequired fields: `model`, `prompt`, `aspect_ratio`.\n\n| Model ID | Cost (cr) | Aspects | Ref Images | Creator | Legacy | Notes |\n|---|---|---|---|---|---|---|\n| `pink-image` | 1 credit per image (~$0.0052) | 1:1, 16:9, 21:9, 3:4, 9:16, 5:2, 4:3, 5:4, 4:5, 9:21, 4:7 | 3 | imgnAI | `pinkimage` | ref support, edit/ref |\n| `gpt-image-2-max` | 28 | 1:1, 2:3, 3:2, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:9 | 10 | OpenAI | `gpt2maximage` | ref support, edit/ref |\n| `nano-banana-2` | 32 | 21:9, 16:9, 3:2, 4:3, 5:4, 1:1, 4:5, 3:4, 2:3, 9:16 | 14 | Google | `nanobanana2` | ref support, edit/ref |\n| `ani` | 2 | 1:1, 16:9, 21:9, 3:4, 9:16, 5:2, 4:3, 5:4, 4:5, 9:21, 4:7 | 0 | imgnAI | — | standard image generation |\n| `fur` | 2 | 1:1, 16:9, 21:9, 3:4, 9:16, 5:2, 4:3, 5:4, 4:5, 9:21, 4:7 | 0 | imgnAI | — | standard image generation |\n| `noob` | 2 | 1:1, 16:9, 21:9, 3:4, 9:16, 5:2, 4:3, 5:4, 4:5, 9:21, 4:7 | 0 | imgnAI | — | standard image generation |\n| `gen` | 3 | 1:1, 16:9, 21:9, 3:4, 9:16, 5:2, 4:3, 5:4, 4:5, 9:21, 4:7 | 0 | imgnAI | — | text prompts |\n| `gpt-image-2` | 14 | 1:1, 2:3, 3:2, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:9 | 16 | OpenAI | `gpt2image` | ref support, edit/ref |\n| `nano-banana-pro` | 28 | 1:1, 2:3, 3:2, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:9 | 8 | Google | `nanobananapro` | ref support, edit/ref |\n| `seedream-4-5` | 12 | 1:1, 16:9, 9:16, 4:3, 3:4, 4:5, 5:4, 4:7, 7:4, 5:2, 2:5 | 10 | ByteDance | `seedream45` | ref support, edit/ref, 4K |\n| `seedream-5-0-lite` | 7 | 1:1, 2:3, 3:2, 3:4, 4:3, 9:16, 16:9, 21:9 | 10 | ByteDance | `seedream5lite` | ref support, edit/ref |\n| `synth` | 2 | 1:1, 16:9, 21:9, 3:4, 9:16, 5:2, 4:3, 5:4, 4:5, 9:21, 4:7 | 0 | imgnAI | — | standard image generation |\n| `qwen-2-0` | 8 | 1:1, 2:3, 3:2, 3:4, 4:3, 9:16, 16:9, 21:9 | 3 | Alibaba | `qwen2image` | ref support, edit/ref |\n| `wan-2-7-image-pro` | 14 | 1:1, 3:4, 4:3, 1:8, 8:1, 9:16, 16:9, 21:9 | 9 | Alibaba | `wan27proimage` | ref support, edit/ref |\n| `aura` | 2 | 1:1, 16:9, 21:9, 3:4, 9:16, 5:2, 4:3, 5:4, 4:5, 9:21, 4:7 | 0 | imgnAI | — | standard image generation |\n| `flux-2-klein-9b` | 10 | 1:1, 16:9, 9:16, 4:3, 3:4, 4:5, 5:4, 4:7, 7:4, 5:2, 2:5 | 4 | Black Forest Labs | `fluxklein9b` | ref support, edit/ref |\n| `nano-banana` | 12 | 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 21:9 | 6 | Google | `nanobanana` | ref support, requires image input, edit/ref |\n| `pixel` | 2 | 1:1, 16:9, 21:9, 3:4, 9:16, 5:2, 4:3, 5:4, 4:5, 9:21, 4:7 | 0 | imgnAI | — | standard image generation |\n| `flux-2-flex` | 28 | 1:1, 2:3, 3:2, 3:4, 4:3, 9:16, 16:9 | 8 | Black Forest Labs | `flux2flex` | ref support, edit/ref |\n| `hyper-cgi` | 2 | 1:1, 16:9, 21:9, 3:4, 9:16, 5:2, 4:3, 5:4, 4:5, 9:21, 4:7 | 0 | imgnAI | `hypercgi` | standard image generation |\n| `imagine-art-1-5-pro` | 14 | 1:1, 16:9, 9:16, 4:3, 3:4, 3:1, 1:3, 3:2, 2:3 | 0 | Imagine Art | `imagineart15` | 4K |\n| `volt` | 2 | 1:1, 16:9, 21:9, 3:4, 9:16, 5:2, 4:3, 5:4, 4:5, 9:21, 4:7 | 0 | imgnAI | — | standard image generation |\n| `wan-2-7-image` | 6 | 1:1, 3:4, 4:3, 1:8, 8:1, 9:16, 16:9, 21:9 | 9 | Alibaba | `wan27image` | ref support, edit/ref |\n| `gpt-image-1-5` | 32 | 1:1 | 0 | OpenAI | `gptimage15` | text prompts |\n| `flux-2-pro` | 8 | 1:1, 2:3, 3:2, 3:4, 4:3, 9:16, 16:9 | 4 | Black Forest Labs | `flux2pro` | ref support, edit/ref |\n| `muse` | 3 | 1:1, 16:9, 21:9, 3:4, 9:16, 5:2, 4:3, 5:4, 4:5, 9:21, 4:7 | 0 | imgnAI | — | text prompts |\n| `gothic` | 2 | 1:1, 16:9, 21:9, 3:4, 9:16, 5:2, 4:3, 5:4, 4:5, 9:21, 4:7 | 0 | imgnAI | — | standard image generation |\n| `z-image-base` | 7 | 1:1, 16:9, 9:16, 4:3, 3:4, 4:5, 5:4, 4:7, 7:4, 5:2, 2:5 | 0 | Z.ai | `zimagebase` | text prompts |\n| `z-image-turbo` | 4 | 1:1, 16:9, 9:16, 4:3, 3:4, 4:5, 5:4, 4:7, 7:4, 5:2, 2:5 | 0 | Z.ai | `zimageturbo` | text prompts |\n| `flux-2-klein-4b` | 4 | 1:1, 16:9, 9:16, 4:3, 3:4, 4:5, 5:4, 4:7, 7:4, 5:2, 2:5 | 4 | Black Forest Labs | `fluxklein4b` | ref support, edit/ref |\n| `rend` | 2 | 1:1, 16:9, 21:9, 3:4, 9:16, 5:2, 4:3, 5:4, 4:5, 9:21, 4:7 | 0 | imgnAI | — | standard image generation |\n| `retro` | 2 | 1:1, 16:9, 21:9, 3:4, 9:16, 5:2, 4:3, 5:4, 4:5, 9:21, 4:7 | 0 | imgnAI | — | standard image generation |\n| `neo` | 2 | 1:1, 16:9, 21:9, 3:4, 9:16, 5:2, 4:3, 5:4, 4:5, 9:21, 4:7 | 0 | imgnAI | — | UHD |\n| `pony` | 2 | 1:1, 16:9, 21:9, 3:4, 9:16, 5:2, 4:3, 5:4, 4:5, 9:21, 4:7 | 0 | imgnAI | — | standard image generation |\n| `nai` | 2 | 1:1, 16:9, 21:9, 3:4, 9:16, 5:2, 4:3, 5:4, 4:5, 9:21, 4:7 | 0 | imgnAI | — | standard image generation |\n| `flux-1-1-ultra` | 18 | 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 4:5, 5:4, 21:9, 9:21, 2:1, 1:2 | 0 | Black Forest Labs | `flux1ultra` | text prompts |\n| `glitch` | 2 | 1:1, 16:9, 21:9, 3:4, 9:16, 5:2, 4:3, 5:4, 4:5, 9:21, 4:7 | 0 | imgnAI | — | standard image generation |\n| `qwen-image` | 8 | 1:1, 16:9, 9:16, 4:3, 3:4 | 0 | Alibaba Cloud | `qwenimage` | ⚠️ Legacy |\n| `seedream-4` | 12 | 1:1, 2:3, 3:2, 3:4, 4:3, 9:16, 16:9, 21:9 | 10 | ByteDance | `seedream4` | ⚠️ Legacy, ref support, edit/ref, 4K |\n| `wan-2-2-image` | 8 | 1:1, 16:9, 9:16, 4:3, 3:4, 21:9 | 0 | Alibaba Cloud | `wan22image` | ⚠️ Legacy |\n| `flux1-d` | 3 | 1:1, 16:9, 21:9, 3:4, 9:16, 5:2, 4:3, 5:4, 4:5, 9:21, 4:7 | 0 | Black Forest Labs | `flux1d` | text prompts |\n| `qwen-image-edit` | 12 | 1:1, 16:9, 9:16, 4:3, 3:4 | 0 | Alibaba Cloud | `qwenedit` | ⚠️ Legacy, requires image input, edit/ref |\n| `supra` | 2 | 1:1, 16:9, 21:9, 3:4, 9:16, 5:2, 4:3, 5:4, 4:5, 9:21, 4:7 | 0 | imgnAI | — | standard image generation |\n| `evo` | 2 | 1:1, 16:9, 21:9, 3:4, 9:16, 5:2, 4:3, 5:4, 4:5, 9:21, 4:7 | 0 | imgnAI | — | ⚠️ Legacy |\n| `toon` | 2 | 1:1, 16:9, 21:9, 3:4, 9:16, 5:2, 4:3, 5:4, 4:5, 9:21, 4:7 | 0 | imgnAI | — | standard image generation |\n| `wassie` | 2 | 1:1, 16:9, 21:9, 3:4, 9:16, 5:2, 4:3, 5:4, 4:5, 9:21, 4:7 | 0 | imgnAI | — | standard image generation |\n| `hyperx` | 2 | 1:1, 16:9, 21:9, 3:4, 9:16, 5:2, 4:3, 5:4, 4:5, 9:21, 4:7 | 0 | imgnAI | — | ⚠️ Legacy |\n| `flux-kontext-max` | 20 | 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 4:5, 5:4, 21:9, 9:21, 2:1, 1:2 | 0 | Black Forest Labs | `kontextmax` | ⚠️ Legacy, requires image input, edit/ref |\n| `furxl-classic` | 2 | 1:1, 16:9, 21:9, 3:4, 9:16, 5:2, 4:3, 5:4, 4:5, 9:21, 4:7 | 0 | imgnAI | `furxl` | ⚠️ Legacy |\n| `flux-kontext-pro` | 12 | 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 4:5, 5:4, 21:9, 9:21, 2:1, 1:2 | 0 | Black Forest Labs | `kontextpro` | ⚠️ Legacy, requires image input, edit/ref |\n| `supra-classic` | 2 | 1:1, 16:9, 21:9, 3:4, 9:16, 5:2, 4:3, 5:4, 4:5, 9:21, 4:7 | 0 | imgnAI | `supraclassic` | ⚠️ Legacy |\n\n## Video Models\n\nAll video models use `POST /v1/generation-requests` with `type: \"video\"`.\nRequired fields: `model`, `prompt`, `duration_seconds`, `aspect_ratio`.\n\n| Model ID | Duration Costs | Aspects | First Frame | Last Frame | Ref Images | Audio In | Audio Out | Video In | Custom Rules | Legacy |\n|---|---|---|---|---|---|---|---|---|---|---|\n| `seedance-2-0` | 5s: 200 cr; 10s: 375 cr; 15s: 550 cr | 16:9, 9:16, 1:1, 4:3, 3:4 | ✓ | ✓ | 7 | ✓ | ✓ | ✗ | Reference images cannot be combined with first-frame or last-frame inputs.; Audio input requires at least one referen... | `seedance2` |\n| `seedance-2-0-480p` | 5s: 120 cr; 10s: 230 cr; 15s: 340 cr | 16:9, 9:16, 1:1, 4:3, 3:4 | ✓ | ✓ | 7 | ✓ | ✓ | ✗ | Reference images cannot be combined with first-frame or last-frame inputs.; Audio input requires at least one referen... | `seedance2480p` |\n| `seedance-2-0-fast` | 5s: 120 cr; 10s: 230 cr; 15s: 340 cr | 16:9, 9:16, 1:1, 4:3, 3:4 | ✓ | ✓ | 7 | ✓ | ✓ | ✗ | Reference images cannot be combined with first-frame or last-frame inputs.; Audio input requires at least one referen... | `seedance2fast` |\n| `ltx-2-3` | 5s: 20 cr; 10s: 40 cr; 15s: 60 cr; 20s: 80 cr | 16:9, 9:16, 1:1 | ✓ | ✓ | 0 | ✓ | ✓ | ✗ | Audio input can only be used with first-frame conditioning.; Video duration follows the selected audio duration. | `ltx23` |\n| `gemini-omni` | 4s: 100 cr; 6s: 135 cr; 8s: 170 cr; 10s: 200 cr | 16:9, 9:16 | ✗ | ✗ | 5 | ✗ | ✓ | ✗ | none | `googlegeminiomnivideo` |\n| `gemini-omni-v2v` | max 10s input video: 500 cr | 16:9, 9:16 | ✗ | ✗ | 5 | ✗ | ✓ | ✓ | Input video duration drives the output length. The listed length is the maximum accepted input video reference length... | `googlegeminiomniv2vvideo` |\n| `happy-horse-1-0-1080p` | 5s: 300 cr; 10s: 600 cr; 15s: 900 cr | 16:9, 9:16, 1:1, 4:3, 3:4 | ✓ | ✗ | 9 | ✗ | ✓ | ✗ | Reference images cannot be combined with first-frame or last-frame inputs. | `happyhorse101080p` |\n| `happy-horse-1-0-720p` | 5s: 150 cr; 10s: 300 cr; 15s: 450 cr | 16:9, 9:16, 1:1, 4:3, 3:4 | ✓ | ✗ | 9 | ✗ | ✓ | ✗ | Reference images cannot be combined with first-frame or last-frame inputs. | `happyhorse10720p` |\n| `wan-2-7-1080p` | 5s: 130 cr; 10s: 260 cr | 16:9, 9:16, 1:1, 4:3, 3:4 | ✓ | ✓ | 5 | ✓ | ✓ | ✗ | Reference audio is interpreted as voice timbre when reference images are present.; Reference images cannot be combine... | `wan271080pvideo` |\n| `wan-2-7-720p` | 5s: 90 cr; 10s: 180 cr | 16:9, 9:16, 1:1, 4:3, 3:4 | ✓ | ✓ | 5 | ✓ | ✓ | ✗ | Reference audio is interpreted as voice timbre when reference images are present.; Reference images cannot be combine... | `wan27720pvideo` |\n| `kling-3-0-kling30pro` | 5s: 350 cr; 10s: 650 cr; 15s: 925 cr | 16:9, 9:16, 1:1 | ✓ | ✓ | 6 | ✗ | ✓ | ✗ | none | `kling30pro` |\n| `kling-o3-4k` | 5s: 450 cr; 10s: 900 cr; 15s: 1350 cr | 16:9, 9:16, 1:1 | ✓ | ✓ | 6 | ✗ | ✓ | ✗ | none | `kling304k` |\n| `kling-3-0-kling30` | 5s: 280 cr; 10s: 550 cr; 15s: 800 cr | 16:9, 9:16, 1:1 | ✓ | ✓ | 6 | ✗ | ✓ | ✗ | none | `kling30` |\n| `veo3-1` | 4s: 380 cr; 8s: 750 cr | 16:9, 9:16 | ✓ | ✗ | 0 | ✗ | ✓ | ✗ | none | `veo3` |\n| `veo3-1-fast` | 4s: 160 cr; 8s: 300 cr | 16:9, 9:16 | ✓ | ✗ | 0 | ✗ | ✓ | ✗ | none | `veo3fast` |\n| `veo3-1-lite` | 8s: 140 cr | 16:9, 9:16 | ✓ | ✗ | 0 | ✗ | ✓ | ✗ | none | `veo3lite` |\n| `seedance-pro` | 5s: 250 cr; 10s: 450 cr | 16:9, 9:16, 1:1, 4:3, 3:4 | ✓ | ✗ | 0 | ✗ | ✗ | ✗ | none | `seedancepro` ⚠️ |\n| `wan-2-5` | 5s: 220 cr; 10s: 400 cr | 16:9, 9:16, 1:1, 4:3, 3:4, 21:9, 9:21 | ✓ | ✗ | 0 | ✗ | ✓ | ✗ | none | `wan25` ⚠️ |\n| `hailuo-2-minimax` | 6s: 200 cr | 16:9, 9:16, 1:1, 4:3, 3:4, 21:9, 9:21 | ✓ | ✗ | 0 | ✗ | ✗ | ✗ | none | `hailuo2` ⚠️ |\n| `kling-2-1-kling21` | 5s: 160 cr; 10s: 300 cr | 16:9, 9:16, 1:1, 4:3, 3:4, 21:9, 9:21 | ✓ | ✗ | 0 | ✗ | ✗ | ✗ | none | `kling21` ⚠️ |\n| `kling-2-1-kling21loop` | 5s: 160 cr; 10s: 300 cr | 16:9, 9:16, 1:1, 4:3, 3:4, 21:9, 9:21 | ✓ | ✗ | 0 | ✗ | ✗ | ✗ | none | `kling21loop` ⚠️ |\n| `seedance-lite` | 5s: 120 cr; 10s: 200 cr | 16:9, 9:16, 1:1, 4:3, 3:4 | ✓ | ✗ | 0 | ✗ | ✗ | ✗ | none | `seedancelite` ⚠️ |\n| `seedance-lite-loop` | 5s: 120 cr; 10s: 200 cr | 16:9, 9:16, 1:1, 4:3, 3:4 | ✓ | ✗ | 0 | ✗ | ✗ | ✗ | none | `seedanceliteloop` ⚠️ |\n| `kling-2-0` | 5s: 350 cr; 10s: 650 cr | 16:9, 9:16, 1:1 | ✓ | ✗ | 0 | ✗ | ✗ | ✗ | none | `kling20` ⚠️ |\n| `kling-1-6` | 5s: 130 cr; 10s: 250 cr | 16:9, 9:16, 1:1 | ✓ | ✗ | 0 | ✗ | ✗ | ✗ | none | `kling16` ⚠️ |\n\nFile v1.0.3:skill-card.md\n\n## Description:\n\nGenerate images, videos, and text/LLM completions via the imgnAI Katana API, with optional media post-processing workflows.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[imgn](https://clawhub.ai/user/imgn)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers and external users use this skill to route agent requests to Katana for image generation, video generation, LLM completions, model lookup, cost-aware request confirmation, polling, and ffmpeg-based media post-processing.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: The skill uses Katana API credentials and can initiate paid external requests.\n\nMitigation: Use a dedicated low-balance API key, keep the secrets file private, and require explicit review of model, prompt, and cost before submission.\n\nRisk: Prompts, files, images, or videos may be sent to imgnAI's API, and local request tracking may persist request metadata.\n\nMitigation: Avoid sending sensitive data unless the selected privacy tier is appropriate, use trusted state and secrets paths, and disable prompt/result history from the Katana API page when needed.\n\nRisk: The skill can propose local file changes for model catalogue updates and media post-processing outputs.\n\nMitigation: Review proposed llms.txt-driven updates before approval and inspect generated ffmpeg or file-writing commands before execution.\n\nRisk: The skill instructs agents to avoid default subagent routing for Katana operations, which may conflict with repository isolation practices.\n\nMitigation: Use the skill only in workspaces where inline execution is acceptable, and be cautious in repositories that depend on AGENTS.md or subagent isolation rules.\n\n## Reference(s):\n\n- [ClawHub skill page](https://clawhub.ai/imgn/skills/katana)\n- [imgnAI homepage](https://app.imgnai.com)\n- [Katana API page](https://app.imgnai.com/katana-api)\n- [Katana API reference](https://kat.imgnai.com/llms.txt)\n- [Model catalogue](artifact/models.md)\n- [Image workflow](artifact/workflows/image.md)\n- [Video workflow](artifact/workflows/video.md)\n- [Text workflow](artifact/workflows/text.md)\n- [Post-processing workflow](artifact/workflows/post-process.md)\n\n## Skill Output:\n\n**Output Type(s):** [text, markdown, code, shell commands, configuration, guidance]\n\n**Output Format:** [Markdown guidance with inline JSON payloads and shell commands]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [May produce API request payloads, polling instructions, cost summaries, generated media URLs, text responses, and local ffmpeg commands.]\n\n## Skill Version(s):\n\n1.0.3 (source: frontmatter, release evidence, changelog)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nFile v1.0.3:workflows/ffmpeg.md\n\n# ffmpeg Post-Processing Reference\n\n**Load this file when the user asks to edit, join, trim, crop, add text, add effects, convert format, or otherwise post-process a generated video.**\n\n---\n\n## ⚠️ Delivering ffmpeg output\n\nffmpeg temp output may not be in your agent's allowed file paths. Copy the file to an appropriate location before sending.\n\n**Example output directory:**\n\n```bash\nSKILL_OUTBOUND_DIR=\"${SKILL_OUTBOUND_DIR:-${OPENCLAW_CONFIG_PATH:-$HOME/.openclaw}/media/outbound}\"\nmkdir -p \"$SKILL_OUTBOUND_DIR\"\ncp ${TMPDIR:-/tmp}/output.mp4 \"$SKILL_OUTBOUND_DIR/output.mp4\"\n```\n\nAdapt the output directory to your agent framework's file delivery requirements.\n\n---\n\n## Detection\n\n```bash\nwhich ffmpeg 2>/dev/null\n```\n\nIf not found, install for your platform:\n\n| Platform | Command |\n|---|---|\n| Linux (Debian/Ubuntu) | `sudo apt install ffmpeg` |\n| Linux (Fedora) | `sudo dnf install ffmpeg` |\n| macOS | `brew install ffmpeg` |\n| Windows | `winget install ffmpeg` or `choco install ffmpeg` |\n\nffmpeg CLI flags are identical across all platforms.\n\n---\n\n**Temp directory:** Examples use `$TMPDIR` for temp files (falls back to `/tmp` on POSIX, uses `%TEMP%` on Windows). Set `TMPDIR` if needed, or substitute your preferred temp directory.\n\n**Windows users:** Shell commands in this file use POSIX (bash) syntax. For PowerShell/CMD: replace `${TMPDIR:-/tmp}` with `$env:TEMP`, use backtick `` ` `` for line continuation, and use `\\` for path separators.\n\n## ⚠️ MANDATORY: Output Compatibility (NON-NEGOTIABLE)\n\n**ALL ffmpeg output videos MUST use these flags to ensure playback on Telegram, iOS, Android, and browsers:**\n\n```bash\n-c:v libx264 -profile:v high -level 4.0 -pix_fmt yuv420p -c:a aac -b:a 192k -ar 44100 -ac 2 -movflags +faststart\n```\n\n**Why:** Without these flags, ffmpeg may default to:\n- `High 4:4:4 Predictive` profile → unplayable on Telegram, iOS, most Android devices, and many browsers\n- Non-standard pixel formats → blank thumbnails, no playback\n- Missing `faststart` → video must fully download before playing\n\n**For crossfade/xfade operations specifically:** Normalise BOTH source videos to the above format BEFORE applying filters, then re-encode the output with the same flags. The xfade filter can produce non-standard pixel formats if fed unnormalised input.\n\n```bash\n# Step 1: Normalise source 1\nffmpeg -y -i source1.mp4 \\\n  -c:v libx264 -profile:v high -level 4.0 -pix_fmt yuv420p -crf 18 -preset medium \\\n  -c:a aac -b:a 192k -ar 44100 -ac 2 \\\n  -movflags +faststart ${TMPDIR:-/tmp}/src1_norm.mp4\n\n# Step 2: Normalise source 2\nffmpeg -y -i source2.mp4 \\\n  -c:v libx264 -profile:v high -level 4.0 -pix_fmt yuv420p -crf 18 -preset medium \\\n  -c:a aac -b:a 192k -ar 44100 -ac 2 \\\n  -movflags +faststart ${TMPDIR:-/tmp}/src2_norm.mp4\n\n# Step 3: Apply filters on normalised sources, re-encode output with same flags\nffmpeg -y -i ${TMPDIR:-/tmp}/src1_norm.mp4 -i ${TMPDIR:-/tmp}/src2_norm.mp4 \\\n  -filter_complex \"...\" \\\n  -c:v libx264 -profile:v high -level 4.0 -pix_fmt yuv420p -crf 18 -preset medium \\\n  -c:a aac -b:a 192k -ar 44100 -ac 2 \\\n  -movflags +faststart output.mp4\n```\n\n**Every command in this file that produces an output MP4 MUST include these flags.** No exceptions.\n\n---\n\n## Speed Guide\n\n| Symbol | Meaning |\n|--------|---------|\n| ⚡ | Stream copy (`-c copy`) — fast, no re-encode. **Note:** preserves source codec settings — if playback issues occur, use full re-encode with mandatory compatibility flags instead |\n| 🐢 | Re-encode required — slower, quality loss possible |\n| 🔧 | Re-encode recommended for best results |\n\n---\n\n## 1. Concatenation / Joining\n\n### Join with concat demuxer (⚡ for same codec/resolution)\n\nBest for videos with identical codecs and resolution.\n\n```bash\n# Create a file list\nprintf \"file 'video1.mp4'\\nfile 'video2.mp4'\\nfile 'video3.mp4'\\n\" > ${TMPDIR:-/tmp}/concat.txt\n\n# Concat without re-encoding (fastest — requires same codec, resolution, frame rate)\nffmpeg -f concat -safe 0 -i ${TMPDIR:-/tmp}/concat.txt -c copy output.mp4\n```\n\n**Gotcha:** All inputs must have identical codecs, resolution, and frame rate. If they differ, use the concat filter or normalise first.\n\n### Join with concat filter (🐢 — handles different properties)\n\n```bash\nffmpeg -i video1.mp4 -i video2.mp4 -filter_complex \\\n  \"[0:v]scale=1920:1080:force_original_aspect_ratio=decrease,pad=1920:1080:(ow-iw)/2:(oh-ih)/2,setsar=1,fps=30[v0]; \\\n   [1:v]scale=1920:1080:force_original_aspect_ratio=decrease,pad=1920:1080:(ow-iw)/2:(oh-ih)/2,setsar=1,fps=30[v1]; \\\n   [v0][0:a][v1][1:a]concat=n=2:v=1:a=1[outv][outa]\" \\\n  -map \"[outv]\" -map \"[outa]\" output.mp4\n```\n\n**Note:** Normalises all inputs to 1080p/30fps before joining.\n\n### Join with crossfade transition (🐢)\n\n```bash\n# 1-second crossfade between two videos\nffmpeg -i video1.mp4 -i video2.mp4 -filter_complex \\\n  \"[0:v]trim=0:5,setpts=PTS-STARTPTS[v0]; \\\n   [1:v]trim=0:5,setpts=PTS-STARTPTS[v1]; \\\n   [v0][v1]xfade=transition=fade:duration=1:offset=4[outv]; \\\n   [0:a]atrim=0:5,asetpts=PTS-STARTPTS[a0]; \\\n   [1:a]atrim=0:5,asetpts=PTS-STARTPTS[a1]; \\\n   [a0][a1]acrossfade=d=1[outa]\" \\\n  -map \"[outv]\" -map \"[outa]\" output.mp4\n```\n\n**Gotcha:** `offset` = duration of first video minus transition duration. Available transitions: `fade`, `wipeleft`, `wiperight`, `dissolve`, `slidedown`, `slideup`, `circleopen`, `circleclose`, etc.\n\n---\n\n## 2. Trimming / Cutting\n\n### Trim by timestamp without re-encoding (⚡)\n\n```bash\nffmpeg -i input.mp4 -ss 00:00:05 -to 00:00:15 -c copy output.mp4\n```\n\n**Gotcha:** Cuts at nearest keyframe — may not be frame-accurate. Put `-ss` before `-i` for faster seek (less accurate).\n\n### Trim with re-encoding for precise cuts (🐢)\n\n```bash\nffmpeg -i input.mp4 -ss 00:00:05.123 -to 00:00:15.456 -c:v libx264 -c:a aac output.mp4\n```\n\n**Note:** Frame-accurate but requires full re-encode.\n\n### Trim to duration (⚡ or 🐢)\n\n```bash\n# Fast (stream copy)\nffmpeg -i input.mp4 -t 10 -c copy output.mp4\n\n# Precise (re-encode)\nffmpeg -i input.mp4 -t 10 -c:v libx264 -c:a aac output.mp4\n```\n\n---\n\n## 3. Cropping & Resizing\n\n### Crop to specific dimensions (🐢)\n\n```bash\n# Crop 1920x1080 to 1080x1080 (centred)\nffmpeg -i input.mp4 -filter:v \"crop=1080:1080:(1920-1080)/2:0\" -c:a copy output.mp4\n```\n\n**Formula:** `crop=W:H:X:Y` where X,Y is top-left corner of crop region.\n\n### Crop to aspect ratio (🐢)\n\n```bash\n# Crop to 1:1 (square)\nffmpeg -i input.mp4 -filter:v \"crop=ih:ih:(iw-ih)/2:0\" -c:a copy output.mp4\n\n# Crop to 9:16 (vertical from landscape)\nffmpeg -i input.mp4 -filter:v \"crop=ih*9/16:ih:(iw-ih*9/16)/2:0\" -c:a copy output.mp4\n```\n\n### Resize / scale (🐢)\n\n```bash\n# Scale to 1080p (preserve aspect ratio)\nffmpeg -i input.mp4 -vf \"scale=1920:1080:force_original_aspect_ratio=decrease,pad=1920:1080:(ow-iw)/2:(oh-ih)/2\" -c:a copy output.mp4\n\n# Scale to exact size (may distort)\nffmpeg -i input.mp4 -vf \"scale=1280:720\" -c:a copy output.mp4\n\n# Scale by factor (half size)\nffmpeg -i input.mp4 -vf \"scale=iw/2:ih/2\" -c:a copy output.mp4\n```\n\n**Note:** Use even dimensions — ffmpeg requires widths/heights divisible by 2 for most codecs.\n\n### Pad / letterbox to different aspect ratio (🐢)\n\n```bash\n# Letterbox 16:9 to 4:3 with black bars\nffmpeg -i input.mp4 -vf \"pad=ih*4/3:ih:(ow-iw)/2:(oh-ih)/2:black\" -c:a copy output.mp4\n```\n\n---\n\n## 4. Effects\n\n### Fade in / fade out — video (🐢)\n\n```bash\n# 1-second fade in, 1-second fade out on a 10-second video\nffmpeg -i input.mp4 -filter:v \\\n  \"fade=t=in:st=0:d=1,fade=t=out:st=9:d=1\" -c:a copy output.mp4\n```\n\n### Fade in / fade out — audio (🐢)\n\n```bash\nffmpeg -i input.mp4 -filter:a \\\n  \"afade=t=in:st=0:d=1,afade=t=out:st=9:d=1\" -c:v copy output.mp4\n```\n\n### Combined video + audio fade (🐢)\n\n```bash\nffmpeg -i input.mp4 -filter_complex \\\n  \"[0:v]fade=t=in:st=0:d=1,fade=t=out:st=9:d=1[v]; \\\n   [0:a]afade=t=in:st=0:d=1,afade=t=out:st=9:d=1[a]\" \\\n  -map \"[v]\" -map \"[a]\" output.mp4\n```\n\n### Speed up / slow down (🔧)\n\n```bash\n# 2x speed (fast forward)\nffmpeg -i input.mp4 -filter_complex \"[0:v]setpts=0.5*PTS[v];[0:a]atempo=2.0[a]\" -map \"[v]\" -map \"[a]\" output.mp4\n\n# 0.5x speed (slow motion)\nffmpeg -i input.mp4 -filter_complex \"[0:v]setpts=2.0*PTS[v];[0:a]atempo=0.5[a]\" -map \"[v]\" -map \"[a]\" output.mp4\n```\n\n**Gotcha:** `atempo` range is 0.5–2.0. For >2x or <0.5x, chain multiple: `atempo=2.0,atempo=2.0` for 4x.\n\n### Reverse video (🐢)\n\n```bash\n# Reverse video and audio\nffmpeg -i input.mp4 -vf reverse -af areverse output.mp4\n```\n\n**Gotcha:** Entire file is loaded into memory. For long videos, trim first.\n\n### Black and white / grayscale (🐢)\n\n```bash\nffmpeg -i input.mp4 -vf \"hue=s=0\" -c:a copy output.mp4\n# Alternative:\nffmpeg -i input.mp4 -vf \"colorchannelmixer=.3:.6:.1:0:.3:.6:.1:0:.3:.6:.1\" -c:a copy output.mp4\n```\n\n### Brightness / contrast / saturation (🐢)\n\n```bash\n# Increase brightness +0.1, contrast 1.5x, saturation 1.2x\nffmpeg -i input.mp4 -vf \"eq=brightness=0.1:contrast=1.5:saturation=1.2\" -c:a copy output.mp4\n```\n\n### Blur (🐢)\n\n```bash\n# Full video Gaussian blur\nffmpeg -i input.mp4 -vf \"gblur=sigma=5\" -c:a copy output.mp4\n\n# Box blur\nffmpeg -i input.mp4 -vf \"boxblur=5:1\" -c:a copy output.mp4\n\n# Selective blur (region only) — blur top-right quarter\nffmpeg -i input.mp4 -filter_complex \\\n  \"[0:v]crop=iw/2:ih/2:iw/2:0,boxblur=10[blur]; \\\n   [0:v][blur]overlay=iw/2:0\" -c:a copy output.mp4\n```\n\n---\n\n## 5. Text / Overlays\n\n### Basic text overlay (🐢)\n\n```bash\nffmpeg -i input.mp4 -vf \\\n  \"drawtext=text='Hello World':fontcolor=white:fontsize=48:x=50:y=50\" \\\n  -c:a copy output.mp4\n```\n\n### Styled text with font, colour, background box (🐢)\n\nffmpeg uses a built-in default font. To use a custom font, provide `fontfile=` with a platform-specific path.\n\n| Platform | Common font path |\n|---|---|\n| Linux | `/usr/share/fonts/truetype/dejavu/DejaVuSans-Bold.ttf` |\n| macOS | `/System/Library/Fonts/Helvetica.ttc` |\n| Windows | `C:/Windows/Fonts/arial.ttf` |\n\n```bash\nffmpeg -i input.mp4 -vf \\\n  \"drawtext=text='Hello World':fontcolor=white:fontsize=64:box=1:boxcolor=black@0.5:boxborderw=10:x=(w-text_w)/2:y=(h-text_h)/2\" \\\n  -c:a copy output.mp4\n```\n\n**Parameters:**\n- `box=1:boxcolor=black@0.5` — semi-transparent background\n- `boxborderw=10` — padding around text\n- `x=(w-text_w)/2:y=(h-text_h)/2` — centre alignment\n\n### Moving / animated text (🐢)\n\n```bash\n# Scroll text from right to left\nffmpeg -i input.mp4 -vf \\\n  \"drawtext=text='Breaking News':fontcolor=white:fontsize=48:x=w-mod(w*t*100\\,w+text_w):y=h-60\" \\\n  -c:a copy output.mp4\n\n# Fade in text (appears at 2s, fully visible at 3s)\nffmpeg -i input.mp4 -vf \\\n  \"drawtext=text='Hello':fontcolor=white:fontsize=64:x=(w-text_w)/2:y=(h-text_h)/2:enable='between(t,2,10)':alpha=if(lt(t,3)\\,(t-2)/1\\,1)\" \\\n  -c:a copy output.mp4\n```\n\n### Timestamp / watermark overlay (🐢)\n\n```bash\n# Burn in timestamp\nffmpeg -i input.mp4 -vf \\\n  \"drawtext=text='%{localtime\\:%Y-%m-%d %H\\\\\\\\\\:%M\\\\\\\\\\:%S}':fontcolor=white:fontsize=24:x=10:y=10:box=1:boxcolor=black@0.5\" \\\n  -c:a copy output.mp4\n```\n\n### Image overlay — picture-in-picture / watermark logo (🐢)\n\n```bash\n# Logo watermark in bottom-right corner (10% of width, with 10px padding)\nffmpeg -i input.mp4 -i logo.png -filter_complex \\\n  \"[1:v]scale=iw*0.1:-1[logo];[0:v][logo]overlay=W-w-10:H-h-10\" \\\n  -c:a copy output.mp4\n```\n\n### Subtitle file overlay (🐢)\n\n```bash\n# Burn in SRT subtitles\nffmpeg -i input.mp4 -vf \"subtitles=subtitles.srt\" -c:a copy output.mp4\n\n# ASS subtitles (preserves styling)\nffmpeg -i input.mp4 -vf \"subtitles=subtitles.ass\" -c:a copy output.mp4\n```\n\n**Gotcha:** `subtitles` filter requires ffmpeg compiled with `--enable-libass`. On most systems this is included. If not, use `-c:s mov_text` to embed (soft subs) instead of burning in.\n\n---\n\n## 6. Audio\n\n### Add / replace audio track (⚡ or 🐢)\n\n```bash\n# Replace audio (re-encode video)\nffmpeg -i input_video.mp4 -i audio.mp3 -c:v libx264 -c:a aac -map 0:v:0 -map 1:a:0 -shortest output.mp4\n\n# Replace audio without re-encoding video (if codecs compatible)\nffmpeg -i input_video.mp4 -i audio.mp3 -c:v copy -c:a aac -map 0:v:0 -map 1:a:0 -shortest output.mp4\n```\n\n### Extract audio from video (⚡)\n\n```bash\nffmpeg -i input.mp4 -vn -c:a copy output.aac\n# Or convert to mp3\nffmpeg -i input.mp4 -vn -c:a libmp3lame -q:a 2 output.mp3\n```\n\n### Remove audio (⚡)\n\n```bash\nffmpeg -i input.mp4 -an -c:v copy output.mp4\n```\n\n### Adjust volume (🐢)\n\n```bash\n# Double volume\nffmpeg -i input.mp4 -filter:a \"volume=2.0\" -c:v copy output.mp4\n\n# Half volume\nffmpeg -i input.mp4 -filter:a \"volume=0.5\" -c:v copy output.mp4\n\n# Normalise audio (loudnorm — two-pass for best results)\nffmpeg -i input.mp4 -af \"loudnorm=I=-16:TP=-1.5:LRA=11\" -c:v copy output.mp4\n```\n\n### Mix multiple audio tracks (🐢)\n\n```bash\nffmpeg -i input.mp4 -i bg_music.mp3 -filter_complex \\\n  \"[0:a][1:a]amix=inputs=2:duration=longest:dropout_transition=2[a]\" \\\n  -map 0:v -map \"[a]\" -c:v copy output.mp4\n```\n\n**Gotcha:** Use `weights` to balance: `amix=inputs=2:weights=1 0.3` for 70% reduction on second track.\n\n### Fade audio in / out (🐢)\n\n```bash\n# 2s fade in, 3s fade out (on 30s video)\nffmpeg -i input.mp4 -af \"afade=t=in:st=0:d=2,afade=t=out:st=27:d=3\" -c:v copy output.mp4\n```\n\n### Audio delay / offset (🐢)\n\n```bash\n# Delay audio by 500ms\nffmpeg -i input.mp4 -filter_complex \"[0:a]adelay=500|500[a]\" -map 0:v -map \"[a]\" -c:v copy output.mp4\n```\n\n---\n\n## 7. Format Conversion\n\n### Common format conversions (🐢 unless same codec)\n\n```bash\n# MP4 to WebM (VP9 + Opus — best quality for web)\nffmpeg -i input.mp4 -c:v libvpx-vp9 -crf 30 -b:v 0 -c:a libopus output.webm\n\n# MP4 to MOV\nffmpeg -i input.mp4 -c:v libx264 -c:a aac output.mov\n\n# MOV to MP4 (stream copy if possible)\nffmpeg -i input.mov -c copy output.mp4\n\n# Any to MKV\nffmpeg -i input.avi -c:v libx264 -c:a aac output.mkv\n\n# AV1 encoding (slow but excellent compression)\nffmpeg -i input.mp4 -c:v libaom-av1 -crf 30 -c:a libopus output.mp4\n```\n\n### Convert to GIF (🐢)\n\n```bash\n# Optimised GIF with palette generation (best quality)\nffmpeg -i input.mp4 -filter_complex \\\n  \"[0:v]fps=15,scale=480:-1:flags=lanczos,palettegen[pal]; \\\n   [0:v]fps=15,scale=480:-1:flags=lanczos[x]; \\\n   [x][pal]paletteuse=dither=bayer:bayer_scale=5\" \\\n  output.gif\n```\n\n**Note:** Two-pass palette generation gives much better colour. Adjust `fps` and `scale` for file size.\n\n### Extract frames as images (⚡)\n\n```bash\n# Extract all frames as PNG\nffmpeg -i input.mp4 frame_%05d.png\n\n# Extract 1 frame per second\nffmpeg -i input.mp4 -vf fps=1 frame_%05d.png\n\n# Extract specific range\nffmpeg -i input.mp4 -vf \"select=between(n\\,0\\,100)\" -vsync vff frame_%05d.png\n```\n\n### Create video from image sequence (🐢)\n\n```bash\n# 30fps from numbered images\nffmpeg -framerate 30 -i frame_%05d.png -c:v libx264 -pix_fmt yuv420p output.mp4\n```\n\n### Extract single frame / thumbnail (⚡)\n\n```bash\n# Frame at 5 seconds\nffmpeg -i input.mp4 -ss 00:00:05 -frames:v 1 -q:v 2 thumbnail.jpg\n\n# Best quality thumbnail at specific time\nffmpeg -i input.mp4 -ss 00:00:05 -frames:v 1 -q:v 1 thumbnail.png\n```\n\n---\n\n## 8. Advanced\n\n### Picture-in-picture (PIP) (🐢)\n\n```bash\n# Small video overlaid on main video (bottom-right, 25% size)\nffmpeg -i main.mp4 -i pip.mp4 -filter_complex \\\n  \"[1:v]scale=iw*0.25:-1[pip];[0:v][pip]overlay=W-w-20:H-h-20\" \\\n  -c:a copy output.mp4\n```\n\n### Side-by-side / split screen (🐢)\n\n```bash\n# Horizontal split (two videos side by side)\nffmpeg -i left.mp4 -i right.mp4 -filter_complex \\\n  \"[0:v]scale=960:1080:force_original_aspect_ratio=decrease,pad=960:1080:(ow-iw)/2:(oh-ih)/2[left]; \\\n   [1:v]scale=960:1080:force_original_aspect_ratio=decrease,pad=960:1080:(ow-iw)/2:(oh-ih)/2[right]; \\\n   [left][right]hstack=inputs=2[v]\" \\\n  -map \"[v]\" -map 0:a output.mp4\n```\n\n### Split video into segments (⚡)\n\n```bash\n# Split into 30-second segments\nffmpeg -i input.mp4 -c copy -map 0 -segment_time 30 -f segment -reset_timestamps 1 segment_%03d.mp4\n```\n\n### Slideshow from images with transitions (🐢)\n\n```bash\n# Simple slideshow (5 seconds per image, crossfade)\nffmpeg -framerate 1/5 -i img%03d.jpg -c:v libx264 -vf \"fps=30,format=yuv420p\" output.mp4\n\n# With crossfade between images (requires explicit filter per image pair)\n# For N images, create a concat with xfade between each:\nffmpeg -loop 1 -t 5 -i img001.jpg -loop 1 -t 5 -i img002.jpg -loop 1 -t 5 -i img003.jpg \\\n  -filter_complex \\\n  \"[0:v]fps=30[v0];[1:v]fps=30[v1];[2:v]fps=30[v2]; \\\n   [v0][v1]xfade=transition=fade:duration=1:offset=4[x01]; \\\n   [x01][v2]xfade=transition=fade:duration=1:offset=8[outv]\" \\\n  -map \"[outv]\" -c:v libx264 output.mp4\n```\n\n### Add metadata (⚡)\n\n```bash\nffmpeg -i input.mp4 -metadata title=\"My Video\" -metadata artist=\"imgnAI\" -metadata comment=\"Generated video\" -c copy output.mp4\n```\n\n---\n\n## 9. Common Pipelines\n\n### Generate → trim → add text → fade in/out\n\n```bash\n# Step 1: Trim generated video\nffmpeg -i generated.mp4 -ss 1 -to 8 -c:v libx264 -profile:v high -level 4.0 -pix_fmt yuv420p -c:a aac -b:a 192k -ar 44100 -ac 2 -movflags +faststart trimmed.mp4\n\n# Step 2: Add text overlay and fades\nffmpeg -i trimmed.mp4 -filter_complex \\\n  \"[0:v]fade=t=in:st=0:d=1,fade=t=out:st=6:d=1,drawtext=text='Generated by imgnAI':fontcolor=white:fontsize=36:x=(w-text_w)/2:y=h-60[v]; \\\n   [0:a]afade=t=in:st=0:d=1,afade=t=out:st=6:d=1[a]\" \\\n  -map \"[v]\" -map \"[a]\" -c:v libx264 -profile:v high -level 4.0 -pix_fmt yuv420p -c:a aac -b:a 192k -ar 44100 -ac 2 -movflags +faststart final.mp4\n```\n\n### Join two generated videos with crossfade\n\n**Note:** For crossfade operations, normalise both source videos first (see MANDATORY section above).\n\n```bash\nffmpeg -i video1.mp4 -i video2.mp4 -filter_complex \\\n  \"[0:v]scale=1920:1080,setsar=1,fps=30[v0]; \\\n   [1:v]scale=1920:1080,setsar=1,fps=30[v1]; \\\n   [v0][v1]xfade=transition=fade:duration=1:offset=4[outv]; \\\n   [0:a]asetpts=PTS-STARTPTS[a0]; \\\n   [1:a]asetpts=PTS-STARTPTS[a1]; \\\n   [a0][a1]acrossfade=d=1[outa]\" \\\n  -map \"[outv]\" -map \"[outa]\" -c:v libx264 -profile:v high -level 4.0 -pix_fmt yuv420p -crf 18 -preset medium -c:a aac -b:a 192k -ar 44100 -ac 2 -movflags +faststart merged.mp4\n```\n\n### Add audio to a silent generated video\n\n```bash\nffmpeg -i generated_video.mp4 -i music.mp3 -c:v copy -c:a aac -map 0:v:0 -map 1:a:0 -shortest video_with_audio.mp4\n```\n\n### Full pipeline: generate → crop 9:16 → add text → convert to GIF\n\n```bash\n# Crop to vertical 9:16\nffmpeg -i generated.mp4 -vf \"crop=ih*9/16:ih:(iw-ih*9/16)/2:0\" -c:v libx264 -profile:v high -level 4.0 -pix_fmt yuv420p -c:a aac -b:a 192k -ar 44100 -ac 2 -movflags +faststart cropped.mp4\n\n# Add text\nffmpeg -i cropped.mp4 -vf \"drawtext=text='Made with imgnAI':fontcolor=white:fontsize=32:x=(w-text_w)/2:y=h-50\" -c:v libx264 -profile:v high -level 4.0 -pix_fmt yuv420p -c:a aac -b:a 192k -ar 44100 -ac 2 -movflags +faststart text_overlay.mp4\n\n# Convert to GIF\nffmpeg -i text_overlay.mp4 -filter_complex \\\n  \"[0:v]fps=12,scale=360:-1:flags=lanczos,palettegen[pal]; \\\n   [0:v]fps=12,scale=360:-1:flags=lanczos[x]; \\\n   [x][pal]paletteuse\" output.gif\n```\n\n---\n\n## 10. Hardware Acceleration\n\n### NVENC (NVIDIA)\n\n```bash\n# Encode with NVENC (much faster than software x264)\nffmpeg -i input.mp4 -c:v h264_nvenc -preset p4 -cq 23 -c:a aac output.mp4\n\n# HEVC (H.265) with NVENC\nffmpeg -i input.mp4 -c:v hevc_nvenc -preset p4 -cq 28 -c:a aac output.mp4\n```\n\n**Check availability:**\n```bash\nffmpeg -hide_banner -encoders | grep nvenc\n```\n\n### Quick Sync (Intel)\n\n```bash\nffmpeg -i input.mp4 -c:v h264_qsv -preset medium -c:a aac output.mp4\n```\n\n### VAAPI (Linux AMD/Intel)\n\n```bash\nffmpeg -i input.mp4 -c:v h264_vaapi -vaapi_device /dev/dri/renderD128 -c:a aac output.mp4\n```\n\n**Note:** Hardware encoders are faster but may produce slightly lower quality than `libx264` at equivalent bitrates. Use `-cq` (constant quality) mode for best quality/size tradeoff.\n\n---\n\n## 11. Quick Reference: Codec Flags\n\n| Codec | Encode flag | Typical use |\n|-------|------------|-------------|\n| H.264 | `-c:v libx264` | Maximum compatibility |\n| H.265/HEVC | `-c:v libx265` | Better compression, slower |\n| VP9 | `-c:v libvpx-vp9` | WebM / web |\n| AV1 | `-c:v libaom-av1` | Best compression, very slow |\n| AAC | `-c:a aac` | Standard audio for MP4 |\n| Opus | `-c:a libopus` | Best audio quality |\n| MP3 | `-c:a libmp3lame` | Universal audio |\n\n---\n\nFile v1.0.3:workflows/image.md\n\n# Image Generation Workflow\n\n**Load this file when the user requests image generation.**\n\n---\n\n## Model Selection\n\nCheck `{baseDir}/models.md` for full catalogue with pricing and capabilities.\n\nFor model aliases and selection, see `{baseDir}/models.md`. Top picks: `gpt-image-2` (gpt-image), `nano-banana-2` (nano), `gen` (default).\n\nIf the user specifies an exact model ID (e.g. `fluxklein9b`), pass it through directly.\n\n### ⚠️ Models Requiring Image Input\n\nThe following models **require an image input** (`reference_image`, `image_urls`, or `reference_assets` with a source image). They cannot generate from text-only prompts:\n- `nano-banana` — first-gen Nano Banana edit model\n- `qwen-image-edit` — Qwen Image Edit\n- `flux-kontext-max` — Kontext Max edit model\n- `flux-kontext-pro` — Kontext Pro edit model\n\n### Cost Reference\n\nSee `{baseDir}/models.md` for detailed pricing per model.\n\n---\n\nBefore submitting, present a pre-submission summary and wait for user confirmation (see `{baseDir}/SKILL.md` for the mandatory protocol). Include model, cost, resolution/aspect ratio, output format, exact prompt, and any reference images.\n\n---\n\n## Workflow\n\nFollow the general workflow defined in `{baseDir}/SKILL.md` (pre-submission → payload → submit → poll → deliver). Image-specific notes:\n\n- After parsing response: if `status` is `failed` or `rejected`, report error and stop immediately\n- Include resolution, output format, and any reference images in the pre-submission summary\n\n---\n\n## Image Editing / Reference Images\n\n**Accepted `image_url` formats:**\n- HTTPS URLs: `\"https://example.com/image.png\"`\n- Data URLs: `\"data:image/jpeg;base64,<BASE64>\"`\n- Raw base64 strings (for text/LLM vision models only)\n- **NOT accepted:** `file://` paths, local file paths — these cause `\"Invalid base64 data\"` errors\n\n**Local file workflow:** When the user sends a local image:\n1. Read the file and base64-encode it with python3\n2. Write the full JSON payload (with data URL) to a temp file using python3\n3. Submit using the secure header pattern from SKILL.md (\"Image/Video requests\")\n\nExample:\n```python\nimport base64, json, tempfile\nwith open(\"input.jpg\", \"rb\") as f:\n    b64 = base64.b64encode(f.read()).decode()\npayload = {\"requests\": [{\"type\": \"image\", \"model\": \"gpt-image-2\", \"prompt\": \"...\", \"image_urls\": [\"data:image/jpeg;base64,\" + b64], \"aspect_ratio\": \"1:1\", \"output_format\": \"png\"}]}\nwith tempfile.NamedTemporaryFile(mode=\"w\", suffix=\".json\", delete=False) as f:\n    json.dump(payload, f)\n    tmpfile = f.name\n```\n\n**Size limits:** For images >100KB, always use the temp file approach. Inline `-d` with large base64 exceeds shell argument limits.\n\n**Reference images count:** Check `models.md` for per-model reference image limits (e.g. `gpt-image-2` supports up to 16).\n\n### ⚠️ Image Input Field — MUST USE `image_urls`\n\nFor **image** type requests, always use `image_urls` (a flat list of URLs/data URLs). Do NOT use `reference_assets` — that field is for **video** requests only and will be silently ignored on image requests, causing the model to generate without seeing your reference images.\n\n```python\n# CORRECT — image_urls (flat list)\npayload = {\"requests\": [{\"type\": \"image\", \"model\": \"gpt-image-2\", \"prompt\": \"...\",\n    \"image_urls\": [\"data:image/jpeg;base64,\" + b64_1, \"data:image/jpeg;base64,\" + b64_2]}]}\n\n# WRONG — reference_assets (silently ignored on image requests)\npayload = {\"requests\": [{\"type\": \"image\", \"model\": \"gpt-image-2\", \"prompt\": \"...\",\n    \"reference_assets\": [{\"kind\": \"source_image\", \"image_url\": \"...\"}]}]}\n```\n\n### Image Input Compatibility Aliases\n\nThe API accepts multiple parameter names for image inputs:\n\n- `image_urls`, `input_images`, `input_image_urls` — list of URLs (aliases)\n- `image_url`, `input_image_url`, `input_image`, `input_image_b64` — single URL (aliases)\n\n`image_urls` is the preferred/canonical form.\n\n---\n\n## Image Parameters\n\n- `aspect_ratio`: `1:1`, `16:9`, `9:16`, `21:9`, etc. Default `1:1`. Use `auto` to match input image. Check `models.md` for per-model supported ratios.\n- `output_format`: `png`, `jpeg`, `webp`. Default `png`.\n- `is_fast`: `true` for cheaper half-resolution (imgnAI models only).\n- `is_uhd`: `true` for UHD (imgnAI models only, overrides `is_fast`).\n\n### Prompt Assist\n\nOn supported tag/booru-based image models only (e.g. `noob`, `ani`, `pony`, etc.), prompt assist fields (`use_assistant`, `prompt_assist`, or `use_prompt_assist`) let users write natural language that is automatically translated to tag-style prompts before dispatch.\n\n**Note:** Only works on imgnAI tag-based models — not external/provider-hosted models.\n\n---\n\n\n\n## Delivery\n\nFollow the delivery pattern defined in `{baseDir}/SKILL.md`. Deliver the generated image to the user with: model name, resolution, credits, dollar cost, description, and the full-res URL.\n\n---\n\n## Error Handling\n\nFollow the error handling protocol defined in `{baseDir}/SKILL.md`.\n\n---\n\n*Part of the Katana skill. See SKILL.md for routing, general configuration, and llms.txt freshness checks.*\n\nFile v1.0.3:workflows/post-process.md\n\n# Post-Processing Workflow\n\n**Load this file when the user asks to edit, join, trim, crop, add text, add effects, convert format, or otherwise post-process a generated image or video.**\n\nFor detailed ffmpeg commands and flags, see `{baseDir}/workflows/ffmpeg.md`.\n\n---\n\n## ⚠️ Audio Preservation (MANDATORY)\n\n**ALWAYS preserve audio tracks from source videos unless the user explicitly asks to remove or replace audio.**\n\nWhen source video has audio (check with `ffprobe`), ffmpeg commands MUST include audio mapping and encoding:\n- Map audio: `-map 0:a` (or appropriate stream index)\n- Encode audio: `-c:a aac -b:a 192k -ar 44100 -ac 2`\n- For crossfades: use `acrossfade` filter alongside video `xfade`\n- Never use `-an` (strips audio) unless explicitly requested\n\nMany video models (e.g. `seedance-2-0`) generate audio alongside video. Stripping it silently is a data loss bug.\n\n---\n\n## Detection\n\n```bash\nwhich ffmpeg 2>/dev/null\n```\n\nIf not found, tell the user: \"ffmpeg is not installed. Install it for enhanced post-processing features (joining, trimming, effects, text overlays, format conversion).\" See `{baseDir}/workflows/ffmpeg.md` for platform-specific installation instructions.\n\n---\n\n## ⚠️ Output Compatibility & Delivery\n\nApply **mandatory output compatibility flags** and **delivery steps** as documented in `{baseDir}/workflows/ffmpeg.md` (sections: \"MANDATORY: Output Compatibility\" and \"Delivering ffmpeg output\").\n\n---\n\n## Workflow\n\n1. **Identify the post-processing task** (trim, join, crop, text overlay, effects, format conversion, etc.)\n2. **Check ffmpeg availability**\n3. **Load `{baseDir}/workflows/ffmpeg.md`** for detailed commands for the specific operation\n4. **Apply mandatory output compatibility flags** from `{baseDir}/workflows/ffmpeg.md` to every output MP4\n5. **Copy output** per `{baseDir}/workflows/ffmpeg.md` delivery instructions\n6. **Deliver** the processed file to the user\n\n---\n\n## Common Operations\n\n| Task | Reference |\n|------|-----------|\n| Join/concatenate videos | workflows/ffmpeg.md §1 |\n| Trim/cut | workflows/ffmpeg.md §2 |\n| Crop/resize | workflows/ffmpeg.md §3 |\n| Effects (fade, speed, reverse, B&W, blur) | workflows/ffmpeg.md §4 |\n| Text overlays / watermarks | workflows/ffmpeg.md §5 |\n| Audio (add, replace, extract, adjust) | workflows/ffmpeg.md §6 |\n| Format conversion (MP4, WebM, GIF) | workflows/ffmpeg.md §7 |\n| Advanced (PIP, split screen, slideshow) | workflows/ffmpeg.md §8 |\n| Hardware acceleration | workflows/ffmpeg.md §10 |\n\n---\n\n## Media Allowlist Handling\n\nWhen delivering post-processed media:\n- Images: PNG, JPEG, WebP\n- Videos: MP4 (H.264 + AAC), GIF\n- Ensure output format is in the allowlist before attempting delivery\n\n---\n\n## Delivery\n\nFollow the delivery pattern defined in `{baseDir}/SKILL.md`. Deliver the post-processed file to the user with a description of the operation performed.\n\n---\n\n## Error Handling\n\n- If ffmpeg fails, report the exact error output to the user\n- Do not retry with different parameters without user approval\n- If a format is unsupported, suggest alternatives\n\n---\n\n*Part of the Katana skill. See SKILL.md for routing and general configuration.*\n\nFile v1.0.3:workflows/text.md\n\n# Text / LLM Chat Workflow\n\n**Load this file when the user requests text generation, LLM chat, or E2EE private model usage.**\n\n---\n\n## Model Selection\n\nCheck `{baseDir}/models.md` for full catalogue with pricing and capabilities.\n\n### Text Model Aliases\n\nSee `{baseDir}/SKILL.md` quick reference or `{baseDir}/models.md` for the full alias table.\n\nIf the user specifies an exact model ID (e.g. `grok-4-20-multi-agent`), pass it through directly.\n\n### Cost Reference\n\nSee `{baseDir}/models.md` for detailed pricing per model.\n\n### Cache Write Costs\n\nSee `{baseDir}/models.md` for the full cache cost column.\n\n---\n\n## Endpoint & Authentication\n\n- **Endpoint:** `POST /v1/chat/completions` (OpenAI-compatible)\n- **Auth:** `Authorization: Bearer <api_key>:<api_secret>` (different from image/video which use separate X-API-Key/X-API-Secret headers)\n\n---\n\n## Workflow\n\n### Submission\n\nBuild the JSON payload in a temp file:\n\n> **Default max_tokens:** If not supplied, the API defaults to 16,000 output tokens. Always set `max_tokens` explicitly to control cost.\n\n```python\nimport json, tempfile\npayload = {\"model\": \"grok-4-3\", \"messages\": [{\"role\": \"user\", \"content\": \"Hello\"}], \"max_tokens\": 16000}\nwith tempfile.NamedTemporaryFile(mode=\"w\", suffix=\".json\", delete=False) as f:\n    json.dump(payload, f)\n    tmpfile = f.name\nprint(tmpfile)\n```\n\nSubmit using the secure header pattern:\n```bash\n. \"${KATANA_SECRETS_FILE:-$HOME/.openclaw/secrets/katana.env}\" && _H=$(mktemp) && chmod 600 \"$_H\" && printf 'Content-Type: application/json\\nAuthorization: Bearer %s:%s\\n' \"$KATANA_API_KEY\" \"$KATANA_API_SECRET\" > \"$_H\" && curl -s -X POST 'https://kat.imgnai.com/v1/chat/completions' -H @\"$_H\" -d @\"$tmpfile\" && rm -f \"$_H\" && rm -f \"$tmpfile\"\n```\n\n---\n\n## Text Parameters\n\n- `model`: Model ID (see aliases above or use exact ID)\n- `messages`: Array of `{\"role\": \"user\"|\"assistant\"|\"system\", \"content\": \"...\"}` objects\n- `max_tokens`: Default `16000`. Check `models.md` for per-model max output limits. `max_completion_tokens` is accepted as an alias.\n- `temperature`, `top_p`: Standard OpenAI-compatible parameters\n- `top_k`: Integer (e.g. 40) — top-K sampling\n- `min_p`: Float (e.g. 0.05) — minimum probability threshold\n- `repeat_penalty`: Float (e.g. 1.1) — repetition penalty\n- `presence_penalty`: Float (e.g. 0.0) — presence penalty\n- `files`: Array of `{\"url\": \"https://...\"}` objects for document/file input (models with \"file\" input type only)\n- **Single `prompt` string:** For clients that send a single string `prompt` (instead of `messages` array), the API converts it into one user message when `messages` is omitted. This skill always uses the `messages` array format.\n\n---\n\n## Response Handling\n\nOpenAI-compatible format:\n```json\n{\n  \"choices\": [\n    {\n      \"message\": {\n        \"role\": \"assistant\",\n        \"content\": \"...\"\n      }\n    }\n  ],\n  \"usage\": { ... }\n}\n```\n\nExtract `choices[0].message.content` from response and return it.\n\n### reasoning_content\n\nWhen the model produces reasoning/thinking output:\n- **Non-streaming chat:** `choices[].message.reasoning_content`\n- **Polled text completions:** `responses[].output_assets[].reasoning_content`\n\nThis field may be absent for models that do not produce reasoning traces.\n\n### Text Response Billing Fields\n\nThe response includes `usage.imgnai` with:\n- `credits_reserved`: credits held before inference\n- `credits_charged`: actual final cost\n- `credits_refunded`: difference refunded if cost < reserve\n- `privacy_mode`: `Anonymized` or `E2EE Private`\n- `billing_source`: where the charge came from\n\n**Billing reserve mechanics:**\n- Normal minimum reserve is 10 credits\n- Large `max_tokens`/`max_completion_tokens` can reserve more\n- API pre-charges the reserve, then refunds unused after actual usage is known\n- Every call has a minimum 0.1 credit charge, rounded up to the nearest 0.1\n\n---\n\n## Streaming\n\nThe API supports SSE streaming for text/LLM calls with API-key billing:\n\n```json\n{\n  \"model\": \"grok-4-3\",\n  \"messages\": [{\"role\": \"user\", \"content\": \"Hello\"}],\n  \"max_tokens\": 16000,\n  \"stream\": true,\n  \"stream_options\": {\"include_usage\": true}\n}\n```\n\n**Response format:** Server-Sent Events (SSE) — chunks arrive as `data: {json}` lines, model name is rewritten to public name, usage included in final chunk, then `data: [DONE]`.\n\n**Restrictions:**\n- Streaming is only available with API-key billing\n- x402 text calls must be non-streaming (omit `stream` or send `stream: false`)\n\n**This skill defaults to non-streaming.** Set `stream: true` only when explicitly requested.\n\n**⚠️ Streaming cost caveat:** If a streaming provider does not return usage in the stream, imgnAI keeps the full reserve instead of refunding based on an unknown cost. Always set `max_tokens` or `max_completion_tokens` to keep the reserve predictable. This applies to API-key billing only.\n\n---\n\n## Multimodal / Vision\n\nVision-capable text models accept image inputs using the OpenAI-compatible content array format. Supported by models with \"image\" in their input types (check `models.md`).\n\n**Format:** Send `image_url` objects in the message content array:\n\n```json\n{\n  \"model\": \"grok-4-3\",\n  \"messages\": [\n    {\n      \"role\": \"user\",\n      \"content\": [\n        {\"type\": \"text\", \"text\": \"Describe this image.\"},\n        {\n          \"type\": \"image_url\",\n          \"image_url\": {\n            \"url\": \"data:image/png;base64,<BASE64_ENCODED_IMAGE_BYTES>\"\n          }\n        }\n      ]\n    }\n  ],\n  \"max_tokens\": 16000\n}\n```\n\n**Accepted image formats:** HTTPS URLs, full `data:image/...;base64,...` data URLs, or raw base64 strings in `image_url.url`. Do NOT send local file paths or `file://` URLs.\n\n**Image processing:** Base64 inputs are converted to JPEG, capped at 4096px max side, original aspect ratio preserved.\n\n**Pricing:** Image tokens are counted in input token pricing. No separate vision surcharge.\n\n### Audio Input\n\nSome text models support audio input (check `models.md` for models listing \"audio\" in input types):\n- `gemini-3-1-flash-lite-preview` — text, image, video, file, audio\n- `gemini-3-1-pro-preview` — text, image, video, file, audio\n- `gemini-3-flash-preview` — text, image, file, audio, video\n\n**Format:** Same OpenAI content array pattern with audio content type:\n\n```json\n{\n  \"role\": \"user\",\n  \"content\": [\n    {\"type\": \"text\", \"text\": \"Transcribe this audio.\"},\n    {\"type\": \"input_audio\", \"input_audio\": {\"data\": \"<base64_encoded_audio>\", \"format\": \"mp3\"}}\n  ]\n}\n```\n\n---\n\n## x402 Streaming Restriction\n\n**x402 text calls must be non-streaming** (omit `stream` or send `stream: false`). Streaming is only available with API-key billing.\n\n---\n\n## Text Polling (Long-Running Models)\n\nSome text models publish `recommended_polling` (check `GET /v1/models`). For these models, use async polling instead of synchronous completions:\n\n1. **Submit:** `POST /v1/chat/completions?wait=false` with the normal payload\n2. **Poll:** `GET /v1/generation-requests/{request_id}` (same endpoint as image/video polling)\n3. **Extract response:** Completed assistant text is in `responses[].output_assets[].content`\n4. **Token usage:** `responses[].output_assets[].metadata.usage`\n5. **Reasoning content:** `responses[].output_assets[].reasoning_content`\n6. **Timeout:** Use the model's `request_timeout_seconds` from `GET /v1/models` as the deadline. For `q-naifu-a3b`, allow at least 600 seconds.\n\n**Polling pattern:** Follow the same polling pattern as image/video (extract `poll_after_seconds`, use as interval). Apply the 10-minute guard for text polling (600s API timeout).\n\n---\n\n## Assistant Prefill\n\nSome models (e.g. `q-naifu-a3b`) support assistant response prefill — making the completion begin with specific text. To use:\n\n1. Add the desired prefix as the **final assistant message** in the messages array\n2. The prefill content **must not end with trailing whitespace**\n\nExample:\n```json\n{\n  \"model\": \"q-naifu-a3b\",\n  \"messages\": [\n    {\"role\": \"user\", \"content\": \"Write a haiku about APIs.\"},\n    {\"role\": \"assistant\", \"content\": \"Here is your haiku:\"}\n  ],\n  \"max_tokens\": 16000\n}\n```\n\n---\n\n## File Input\n\nModels with \"file\" in their input types accept file attachments via the `files` array parameter:\n\n```json\n{\n  \"files\": [{ \"url\": \"https://example.com/document.pdf\" }]\n}\n```\n\n### `q-naifu-a3b` Supported File Formats\n\n- **Image formats:** JPEG/JPG, PNG, WebP, GIF, AVIF, TIFF/TIF, BMP, HEIC/HEIF\n- **Document formats:** .pdf, .docx\n- **Code/text formats:** .txt, .md, .json, .yaml, .yml, .csv, .tsv, .xml, .html, .css, .js, .ts, .tsx, .jsx, .py, .rs, .go, .java, .kt, .c, .cpp, .h, .cs, .sh, .ps1, .toml, .ini, .cfg, .log, .sql\n- **Blocked:** video, audio, archives, executables, old Office formats\n- **Max size:** 40 MB per file\n- Unknown extensions allowed when content is detected as image or text\n\n---\n\n## Delivery\n\n**Model features:** Text models support varying capabilities including Tool calling, Structured output, Long context, Vision, Video input, Audio input, and File input. Check `models.md` for per-model feature flags.\n\n**Return the model's response VERBATIM.** Do not summarise, rephrase, paraphrase, or editorialise the LLM's output. Present it exactly as returned, unless the user explicitly asks you to summarise or reformat.\n\nFollow the delivery pattern defined in `{baseDir}/SKILL.md`.\n\n---\n\n## Private & E2EE Private Models\n\nThree privacy tiers exist for text models:\n\n1. **Anonymized:** Customer account identity is not sent with the inference request. The model operator may process prompt content.\n2. **Private:** The request is handled through a **private in-house model path**. No E2EE/hardware attestation — imgnAI can see the content but it stays in-house.\n3. **E2EE Private:** The request is handled through a **hardware-protected confidential-compute model path**. The user's prompts and model responses are encrypted end-to-end — imgnAI cannot read them.\n\n**Available private models** (check `models.md` for current list):\n- `q-naifu-a3b` (Private — in-house, no attestation)\n- `kimi-k2-6-private`, `qwen3-coder-next-private`, `glm-5-1-private` (E2EE Private)\n- `mimo-v2-flash-private`, `qwen3-5-27b-private`, `qwen3-5-397b-a17b-private` (E2EE Private)\n- `minimax-m2-5-private`, `glm-5-private`, `kimi-k2-5-private` (E2EE Private)\n- `deepseek-v3-2-private`, `qwen3-coder-480b-a35b-private` (E2EE Private)\n- `qwen3-vl-30b-a3b-instruct-private`, `gemma-3-27b-private` (E2EE Private)\n\nUsage is identical to regular models — just use the model ID. Privacy is handled transparently by the API.\n\n### Privacy Attestation (E2EE Private only)\n\n**Attestation only applies to E2EE Private models, not Private models.** Verify E2EE claims for E2EE Private models:\n\n```\nGET /v1/text/attestation?model=<private_model>&nonce=<64_hex_nonce>\n```\n\nReturns a privacy proof for the E2EE model. Uses Intel TDX and NVIDIA Confidential Computing where applicable. The `nonce` must be a 64-character hex string.\n\n---\n\n## Error Handling\n\nFollow the error handling protocol defined in `{baseDir}/SKILL.md`.\n\n---\n\n*Part of the Katana skill. See SKILL.md for routing, general configuration, and llms.txt freshness checks.*\n\nFile v1.0.3:workflows/video.md\n\n## Video Generation Workflow\n\n**Load this file when the user requests video generation.**\n\n---\n\n### ⚠️ Video Poll Timeout\n\nThe video API has a **6000-second (100-minute) server-side timeout**. Video jobs may legitimately run for extended periods. The standard 10-minute poll guard from SKILL.md is too aggressive for video — use a **100-minute guard** for video requests instead.\n\nTypical video returns: < 800 seconds. The 6000s deadline is a worst-case limit.\n\n---\n\n## Model Selection\n\nCheck `{baseDir}/models.md` for full catalogue with pricing, durations, and capabilities.\n\nFor model aliases and selection, see `{baseDir}/models.md`. Top picks: `seedance-2-0-fast` (default), `ltx-2-3` (ltx).\n\nIf the user specifies an exact model ID, pass it through directly.\n\n### Cost Warning\n\nVideo can be expensive. Always mention estimated cost before generating. See `{baseDir}/models.md` for full pricing.\n\n---\n\nBefore submitting, present a pre-submission summary and wait for user confirmation (see `{baseDir}/SKILL.md` for the mandatory protocol). Include model, cost, duration, aspect ratio, exact prompt, and any reference images.\n\n---\n\n## Workflow\n\nFollow the general workflow defined in `{baseDir}/SKILL.md` (pre-submission → payload → submit → poll → deliver). Video-specific notes:\n\n- Include duration, aspect ratio, and any reference images in the pre-submission summary\n- Video can be expensive — always highlight estimated cost\n- **Poll timeout:** Video API timeout is 6000s (100 min). Use 100-minute guard for video polling, not the 10-minute image/text guard.\n\n---\n\n## Video Parameters\n\n- `duration_seconds`: must match a value in the model's `video_lengths_and_costs`. Common: `5`, `10`, `15`.\n- `aspect_ratio`: `16:9`, `9:16`, `1:1`, `4:3`, `3:4`. Default `16:9`. Check `models.md` for per-model support.\n- `video_image_data.first_frame_image_url`: first frame image (URL or data URL).\n- `video_image_data.mid_frame_image_url`: mid-frame image (some models).\n- `video_image_data.last_frame_image_url`: last frame (some models).\n- `video_image_data.reference_image_urls`: reference images (some models).\n- `video_image_data.audio_input_urls`: audio references (some models).\n- `video_image_data.video_list`: array of video input clips for V2V models (each object requires `url`; optional `start`/`ends` second offsets with `video_offset_allowed` rule).\n- `reference_assets`: typed asset references (alternative to video_image_data). Image kinds (`style_reference`, `reference_image`, `image`) and audio kinds (`audio`, `source_audio`, `reference_audio`, `audio_reference`).\n\n### Top-Level Compatibility Aliases\n\nThe API accepts several top-level aliases that map to `video_image_data` fields. The `video_image_data` forms are preferred for new integrations:\n\n- Top-level `image_url`, `input_image_url`, `input_image`, `input_image_b64` → alias for `video_image_data.first_frame_image_url`\n- Top-level `reference_image_urls` → alias for `video_image_data.reference_image_urls`\n\n### Aspect Ratio: `auto`\n\n`aspect_ratio: \"auto\"` uses the first frame dimensions, then the first reference image. Defaults to `1:1` if neither exists.\n\n### Audio File Limits\n\nPer-model audio file limit via `maximum_reference_audio_files`. There is a global cap of 4 audio files. Check `models.md` for per-model limits.\n\n### Custom Rules (varies by model — check `models.md`)\n\nFull glossary of named custom rules:\n\n| Rule | Description |\n|------|-------------|\n| `audio_15s_max` | Combined audio limited to 15s |\n| `audio_drives_duration` | Video duration follows audio duration |\n| `audio_ff_only` | Audio only with first-frame conditioning |\n| `audio_needs_reference_image` | Audio requires reference image |\n| `audio_or_fflf_exclusive` | Audio cannot combine with frame inputs |\n| `lf_needs_ff` | Last frame requires first frame |\n| `reference_ff_only` | References may combine with first frame only |\n| `reference_is_voice_timbre` | Audio interpreted as voice timbre when refs present |\n| `reference_no_ff_or_lf` | References cannot combine with frame inputs |\n| `video_required` | Model requires video input via `video_image_data.video_list` |\n| `video_offset_allowed` | Accepts `start`/`ends` second offsets in `video_list` objects |\n| `input_video_drives_length` | Input video determines output length; listed duration is max input and fixed cost |\n\nSome models require image input (e.g. `seedance-pro`, `hailuo-2-minimax`, `kling-2-0`). Check `models.md` for per-model rules.\n\n### Video-to-Video (V2V)\n\nV2V models (e.g. `gemini-omni-v2v`) accept video input via `video_image_data.video_list`. Each object requires `url` (HTTPS video URL). Optional `start` and `ends` offsets in seconds (requires `video_offset_allowed` rule).\n\nThe input video clip drives the output length and billing tier (`input_video_drives_length`). The listed duration is the maximum accepted input and the fixed cost.\n\n---\n\n\n\n## Audio Output\n\nMost modern video models generate audio alongside the video by default. The model interprets the prompt content for sound design — it doesn't just add ambient noise, but generates contextually appropriate audio (speech, music, environmental sounds, effects) based on what's described in the prompt and depicted in the visual content.\n\n### Which models generate audio\n\n**Audio output enabled (✓):** `seedance-2-0`, `seedance-2-0-480p`, `seedance-2-0-fast`, `ltx-2-3`, `gemini-omni`, `gemini-omni-v2v`, `happy-horse-1-0-1080p`, `happy-horse-1-0-720p`, `wan-2-7-1080p`, `wan-2-7-720p`, `kling-3-0-kling30pro`, `kling-o3-4k`, `kling-3-0-kling30`, `veo3-1`, `veo3-1-fast`, `veo3-1-lite`, `wan-2-5`\n\n**Audio output silent (✗):** `seedance-pro`, `hailuo-2-minimax`, `kling-2-1-kling21`, `kling-2-1-kling21loop`, `seedance-lite`, `seedance-lite-loop`, `kling-2-0`, `kling-1-6`\n\nSee `{baseDir}/models.md` for the full table with the Audio Out column.\n\n### Audio input vs audio output\n\nAudio **input** (via `video_image_data.audio_input_urls`) and audio **output** are independent capabilities:\n- Some models accept audio references for voice/sound conditioning **and** generate audio output (e.g. `seedance-2-0`)\n- Some models generate audio output but don't accept audio references (e.g. `veo3-1`, `kling-3-0-kling30pro`)\n- Legacy models may do neither\n\nWhen a model supports both, the audio reference serves as a style/timbre guide — the model still generates new audio content interpreted from the prompt, not a copy of the reference.\n\n### Suppressing audio\n\nCurrently there is no API parameter to disable audio output on models that generate it. If silent video is required, post-process with ffmpeg to strip the audio track:\n\n```bash\nffmpeg -i input.mp4 -c:v copy -an silent_output.mp4\n```\n\n---\n\n## Delivery\n\nFollow the delivery pattern defined in `{baseDir}/SKILL.md`. Deliver the generated video to the user with: model name, duration, resolution, credits, dollar cost, description, and the full-res URL.\n\n**Thumbnail preview:** `responses[].output_assets[].thumbnail_silent_video_mp4_url` may be present for video outputs — a short/silent lightweight MP4 preview suited to galleries and hover previews. Returned as blank string when unavailable. This is NOT the full video — always use `original_data_url` for delivery.\n\n## Video Continuation\n\nTo chain video clips together:\n\n1. Extract `final_frame_image_url` from the completed poll response (blank string when unavailable)\n2. Use it as `video_image_data.first_frame_image_url` in the next video request\n3. Stitch the resulting clips with ffmpeg (see `{baseDir}/workflows/post-process.md`)\n\nThis enables multi-shot video generation by using the last frame of one clip as the starting frame of the next.\n\n---\n\n## Pre-Submission Confirmation (MANDATORY)\n\nIf the user asks to edit, join, trim, crop, add text, add effects, convert format, or otherwise post-process a generated video, load `{baseDir}/workflows/post-process.md`.\n\n---\n\n## Error Handling\n\nFollow the error handling protocol defined in `{baseDir}/SKILL.md`.\n\n---\n\n*Part of the Katana skill. See SKILL.md for routing, general configuration, and llms.txt freshness checks.*\n\nArchive v1.0.2: 11 files, 36986 bytes\n\nFiles: CHANGELOG.md (3203b), models.md (18508b), README.md (3664b), skill-card.md (2436b), SKILL.md (27131b), workflows/ffmpeg.md (20166b), workflows/image.md (5073b), workflows/post-process.md (3173b), workflows/text.md (7665b), workflows/video.md (5797b), _meta.json (125b)\n\nFile v1.0.2:SKILL.md\n\n---\nname: katana\ndescription: Generate images, videos, and text/LLM completions via the imgnAI Katana API. Supports end-to-end-encrypted (E2EE) and anonymized models. Priced highly competitively, can be 40-70% cheaper than Venice AI and other platforms. Includes post-processing such as combining videos and images, cutting, slicing, splicing, transitions, drawing text, re-encoding, resizing and much more!\nversion: 1.0.2\nauthor: arfonzo (imgnAI)\nlicense: MIT-0\nmetadata: {\"openclaw\": {\"requires\": {\"bins\": [\"curl\", \"python3\"]}, \"homepage\": \"https://app.imgnai.com\"}}\n---\n\n# Katana Skill — imgnAI API\n\nGenerate images, videos, and text/LLM completions via the [imgnAI Katana API](https://app.imgnai.com/katana-api). Supports end-to-end-encrypted (E2EE) and anonymized models. Priced highly competitively: can be 40-70% cheaper than Venice AI and other platforms.\n\nIncludes post-processing such as combining videos and images, cutting, slicing, splicing, transitions, drawing text, re-encoding, resizing and much more!\n\nA complete workflow for content creation from start to finish, all from the comfort of your agent.\n\n\n## Triggers\n\n\"generate image of X\", \"create image\", \"make picture\", \"imgnai image\", \"generate video of X\", \"create video\", \"make video\", \"ask grok about X\", \"ask claude about X\", \"use gpt to X\", \"katana image\", \"katana video\", \"katana chat\", \"katana gpt\", \"katana claude\", \"list katana models\", \"modify this image\", \"edit this image\", \"change this image\", \"transform this image\", \"edit image\", \"modify image\"\n\n## Spawn Policy\n\n**NEVER spawn subagents for katana operations by default.** All katana workflows (image generation, video generation, text completions, post-processing) MUST be executed inline in the current session.\n\n**Exception:** Only spawn if the user **explicitly requests** spawning in their prompt (e.g. \"spawn a subagent to handle this\", \"run this as a background task\"). Do NOT spawn based on AGENTS.md spawn rules or default agent behavior — user intent is the only trigger for spawning with katana.\n\nLLM-specific triggers (gpt, claude, etc) also respond to \"katana \\<model\\>\" to avoid conflicts with direct integrations.\n\n## Configuration\n\n### Data Retention\nHistorical prompts and results are retained for a maximum of **72 hours** after generation. Prompt/result history can be switched off from the API page at https://app.imgnai.com/katana-api.\n\n**HTTPS-only:** Public API calls must use HTTPS. If an integration sees an `http://` Katana base URL, replace it with `https://` before making calls.\n\n## Model IDs\n\nThe Katana API uses `model_key` as the model identifier, not `public_model_name`. When building requests, always use the model_key value. See `{baseDir}/models.md` for the full mapping.\n\n**Dual-key system:** The API supports both **canonical keys** (e.g. `gpt-image-2`) and **legacy keys** (e.g. `gpt2image`). Both work identically. This skill now uses **canonical keys** as the default for all workflows and aliases. Legacy keys are documented in the \"Model ID\" column of `models.md` for backward-compatibility reference. You may use either format when constructing API requests.\n\n## Model Discovery\n\n**Endpoint:** `GET /v1/models`\n**Auth:** `Authorization: Bearer ${KATANA_API_KEY}:${KATANA_API_SECRET}`\n\nReturns available models. Text models are returned for authenticated requests.\nFor the complete model catalogue including image/video, see models.md.\n\n**Usage:** Generally not needed before requests — use models.md as reference.\n\n---\n\n## Payment Methods\n\nThe API supports two payment methods:\n- **API key + secret** (Bearer auth) — used by this skill, preferred\n- **x402 micropayment** — NOT used by this skill\n\nNote: x402 text requests must be non-streaming. This skill only uses API-key auth.\n\n### Text/LLM Notes\n- **Streaming:** `stream: true` supported with SSE for API-key billing. x402 text calls must be non-streaming.\n- **Vision/multimodal:** Send images via `image_url` with Base64 data URLs or HTTPS URLs in messages. Base64 inputs are converted to JPEG, capped at 4096px max side.\n- **Billing:** Pre-charge reserve (10 credit minimum), refund after actual usage. 0.1 credit minimum charge rounded up.\n- **Refunds:** If an image or video generation fails after it was charged, credits or x402 balance will be refunded within 5 minutes, pending no Terms of Service violation.\n- **Default max_tokens:** If omitted and model supports output caps, API defaults to 16000.\n- **E2EE attestation:** `GET /v1/text/attestation?model={model}&nonce={64_hex_nonce}` for private/E2EE models.\n\n---\n\n- **API Base URL:** `https://kat.imgnai.com`\n- **API Reference:** https://kat.imgnai.com/llms.txt\n- **Model catalogue:** `{baseDir}/models.md`\n- **Skill directory:** Resolve dynamically from this file's location as `{baseDir}`. Most agent frameworks resolve this automatically.\n\n## Credentials\n\n**Secrets file:** Store your API key and secret in a file (default: `~/.openclaw/secrets/katana.env`):\n```\nKATANA_API_KEY=your_key_here\nKATANA_API_SECRET=your_secret_here\n```\nCreate with `chmod 600`. Get your credentials from https://app.imgnai.com/katana-api.\n\n**Loading:** All curl examples in this skill use `.` (dot) source to load credentials into the shell environment:\n```bash\n. \"${KATANA_SECRETS_FILE:-$HOME/.openclaw/secrets/katana.env}\"\n```\nOverride the default path with the `KATANA_SECRETS_FILE` environment variable.\n\n### ⚠️ Credential Security (MANDATORY)\n\n**NEVER display secrets in tool output.** The `.` source command loads credentials into shell variables silently — no output is produced. This is the correct and secure approach.\n\n**Banned patterns:**\n- `cat ~/.openclaw/secrets/katana.env`\n- `KATANA_API_KEY=kat_live_... curl ...`\n- Any form of reading secrets into tool output\n\n**If credential loading fails:** Fix the secrets file path or contents. Do NOT bypass security by hardcoding values.\n\n---\n\n## Optional Dependencies\n\nThese are not required for core API usage but enable additional features:\n\n| Binary | Needed for | Install |\n|--------|------------|--------|\n| `jq` | JSON parsing for API responses | `apt install jq` / `brew install jq` |\n| `python3` | Payload building, JSON parsing fallback | Pre-installed on most systems |\n| `ffmpeg` | Video post-processing (trim, join, effects) | `apt install ffmpeg` / `brew install ffmpeg` |\n\n`jq` or `python3` is needed for JSON parsing. Post-processing requires `ffmpeg`.\n\n---\n\n## ⚠️ MANDATORY ROUTING — DO NOT SKIP\n\n**Before ANY generation or post-processing request, you MUST load the correct workflow file:**\n\n| Task | Load this file |\n|------|---------------|\n| Image generation | `{baseDir}/workflows/image.md` |\n| Video generation | `{baseDir}/workflows/video.md` |\n| Text/LLM generation | `{baseDir}/workflows/text.md` |\n| Post-processing (ffmpeg, combine, text overlay, etc) | `{baseDir}/workflows/post-process.md` |\n\n**NEVER attempt a generation without loading the workflow file first.**\n**NEVER guess parameters — the workflow file has the exact steps.**\n\n---\n\n## Cost Reporting (ALL Requests)\n\n**After every generation (text, image, video), send a separate follow-up message with a cost summary.** Include all relevant details from the response:\n\n```\n📊 Katana Summary\nModel: gemma-4-26b-a4b (Anonymized)\nRequest: bf11cf04-8747-480e-a7f7-7d6cb092c614\nTokens: 42 in / 176 out (text only)\nCost: 0.1 credits (~$0.001)\nPrivacy: Anonymized\nTime: ~3s\n```\n\nFor image/video, replace tokens with dimensions/duration as relevant. Always compute cost in USD using the current credit rate (see `{baseDir}/models.md`).\n\n---\n\n## Model Aliases (Quick Reference)\n\n### Text/LLM\n\n| User says | API model ID |\n|---|---|\n| grok | `grok-4-3` |\n| gpt / gpt-5 | `gpt-5-5` |\n| claude / claude-opus | `claude-opus-4-7` |\n| claude-sonnet | `claude-sonnet-4-6` |\n| claude-haiku | `claude-haiku-4-5` |\n\n### Image\n\n| User says | API model ID |\n|---|---|\n| default / imgnai | `gen` |\n| anime | `ani` |\n| gpt-image | `gpt-image-2` |\n| nano | `nano-banana-2` |\n| flux | `flux-2-pro` |\n| pink | `pink-image` |\n\n### Video\n\n| User says | API model ID |\n|---|---|\n| default / seedance | `seedance-2-0-fast` |\n| seedance-hd | `seedance-2-0` |\n| ltx | `ltx-2-3` |\n| kling | `kling-3-0-kling30` |\n| veo | `veo3-1` |\n\nIf the user specifies an exact model ID, pass it through directly. Full alias tables in `{baseDir}/models.md`.\n\n---\n\n## Pre-Submission Confirmation (MANDATORY)\n\nBefore submitting ANY generation request, present a summary (model, cost in credits AND dollars, details, prompt) and **wait for user confirmation**. See each workflow file for details.\n\n**NO EXCEPTIONS:** There is no urgency override. \"just do it\", \"generate now\", /katana, or any other shortcut does NOT skip confirmation. ALWAYS present summary and wait for explicit approval before submitting.\n\n---\n\n## Error Protocol\n\n**ONE-ATTEMPT RULE: Every paid API call gets exactly ONE attempt per turn. If the tool result is lost, missing, or empty after a submission — STOP. Report to the user that the result was lost. Wait for user confirmation before retrying. NEVER retry a paid API call silently, even if the result seems to have vanished.**\n\n**STRICT — NO SILENT RETRIES.** Every error stops. Every retry needs approval. Tool-result-loss (result never arrives, empty, or vanishes) is a hard-stop condition equal to a visible error. See each workflow file for details.\n\n- ANY error or tool-result-loss → STOP, report to user (what happened, credits charged, total across attempts)\n- Tool-result-loss (result shows 'missing tool result' or similar synthetic error) → the API call likely already succeeded. STOP. Report to user. Do NOT retry the same request.\n**Terminal submission responses:** If the submission response itself is terminal (`status: \"failed\"`, `status: \"rejected\"`, or all response items rejected) — do NOT poll. Report the returned `responses[].error` or top-level error to the user immediately.\n\n- **Upstream errors are terminal.** If the API returns `upstream_error` (404, 500, etc), do NOT try a different model, do NOT retry with different parameters, do NOT submit to another endpoint. STOP and report the error to the user. You MAY suggest recommended next steps or options (e.g. \"model X returned 404 — want me to try model Y instead?\"), but ANY proposed plan requires explicit user approval before execution.\n- Propose fix → wait for explicit user approval\n- Banned: automatic retries, debug/test requests, parameter changes without telling user, lying about call counts, silent retries on lost results\n\n---\n\n## Concurrency Guard\n\n**NEVER submit a new request while any previous request is still processing.** One request in flight at a time — no exceptions.\n\n- Before submitting, verify no pending/processing requests exist\n- If a previous request is still running (poll returns incomplete), either wait for it, ask the user to cancel, or ask the user to approve submitting a concurrent request\n- This applies across ALL endpoints: text, image, and video\n\n---\n\n## Immediate Status Updates\n\nAfter submitting async generations (image/video), deliver a confirmation to the user BEFORE starting the poll loop. Include the model, cost, and request_id.\n\n## Async Polling\n\nImage and video generations are asynchronous. After submitting, poll manually.\n\n**Poll command:**\n```bash\n. \"${KATANA_SECRETS_FILE:-$HOME/.openclaw/secrets/katana.env}\" && _H=$(mktemp) && chmod 600 \"$_H\" && printf 'X-API-Key: %s\\nX-API-Secret: %s\\n' \"$KATANA_API_KEY\" \"$KATANA_API_SECRET\" > \"$_H\" && curl -s \"https://kat.imgnai.com/v1/generation-requests/${REQUEST_ID}\" -H @\"$_H\" && rm -f \"$_H\"\n```\n\n**Raw response:** Pipe to `jq '.'`.\n\n**Formatted:** Pipe to:\n```bash\npython3 -c \"import sys,json; d=json.load(sys.stdin); r=d.get('responses',[]); [print(f\\\"Status: {r[i].get('status','?')}\\\\nURL: {a.get('original_data_url','')}\\\\nDims: {a.get('width','?')}x{a.get('height','?')}\\\\nCredits: {r[i].get('metadata',{}).get('credits_spent','?')}\\\\nExpires: {a.get('expires_at','')}\\\") for i in range(len(r)) for a in r[i].get('output_assets',[])]\"\n```\n\n**`wait` parameter:** `wait=true` is available for convenience (blocks until complete), but production integrations should prefer polling with `wait=false` (the default).\n\n**Polling pattern:** Extract `poll_after_seconds` from the submission response and use it as the initial polling interval. If the poll response includes a new `poll_after_seconds`, use that for the next interval. Fall back to polling every 30 seconds for the first 5 minutes, then every 60 seconds if `poll_after_seconds` is absent or null.\n\n**Agent responsibility:** The agent decides how to schedule polls (intervals, background tasks, etc). Do not use long-running background processes — use single polls at intervals.\n\n### ⚠️ Polling Pattern Constraints\n\n**Keep `.` source and `curl` in the same command chain.** Shell `sleep` or `process poll` between commands breaks the env var loading — env vars are lost.\n\n**Correct:** Single exec call containing the full chain (see poll command above).\n\n**Wrong:** Separating `.` + `sleep` + `curl` into different exec calls.\n\n**If your agent cannot chain commands:** Use the agent-native polling mechanism (background exec, process poll, etc) with the full command as one unit.\n\n**Response handling for completed polls:**\n- Extract `original_data_url` for delivery (full-resolution)\n- Extract dimensions from `responses[].output_assets[].width/height` (NOT from submission response)\n- Extract credits from `responses[].metadata.credits_spent`\n- Extract expiry from `responses[].output_assets[].expires_at` — display in user's local timezone in delivery summary\n\n## Response Handling\n\n### Dimensions (IMPORTANT)\n1. **Submission response** (`requests[].width/height`) — PREVIEW dimensions, NOT actual output size.\n2. **Completed poll response** (`responses[].output_assets[].width/height`) — ACTUAL output dimensions.\n\n**Always report dimensions from the completed poll response, never from the submission acknowledgement.**\n\n### URL Fields (IMPORTANT)\n- **`original_data_url`** — full-resolution original. **Always use this for delivery.**\n- **`url`** — may be a compressed/reduced version. Do NOT use for delivery.\n- **`thumbnail_image_url`** — small thumbnail only.\n- **`thumbnail_silent_video_mp4_url`** — silent lightweight MP4 preview for video galleries/hover previews. NOT the full video.\n- **`final_frame_image_url`** — last frame still image for completed videos. Use as first-frame input for video continuation workflows. Blank string when unavailable.\n\n### CLIP Tag Metadata\n`responses[].output_assets[].metadata.tags` contains CLIP-derived tags with confidence scores (e.g. `{\"tag\": \"ceramic_mug\", \"confidence\": 0.94}`). Only available on in-house imgnAI models — external/provider-hosted models return no CLIP-tag metadata.\n\n### Model Normalization\nCompleted media responses may normalize `requests[].model` and `responses[].metadata.model` (e.g. legacy key → canonical key). Use `GET /v1/models` for canonical display names.\n\n### Item Timestamps\n- `responses[].started_at` — item-level processing start timestamp\n- `responses[].completed_at` — item-level processing end timestamp\n- `created_at` — top-level request submission timestamp\n- `updated_at` — top-level request last-modified timestamp\n- Useful for tracking actual generation time per item\n\n### Asset Type Fields\n- `responses[].output_assets[].kind` — asset type (e.g. `\"image\"`, `\"video\"`)\n- `responses[].output_assets[].mime_type` — MIME type (e.g. `\"image/png\"`, `\"video/mp4\"`)\n\n### ⚠️ Anti-Pattern Warning\nData is under `responses[].output_assets[]` — do NOT look for `results[].url`. That is NOT the Katana response shape.\n\n### ⚠️ `output` Object\nDo NOT send an `output` object for ordinary integrations. This is for internal/special use only.\n\n## Payload Submission\n\nBuild the JSON payload in a temp file (required for large payloads and to avoid secrets in process listings):\n\n```python\nimport json, tempfile\npayload = {\"requests\": [{\"type\": \"video\", \"model\": \"seedance-2-0-fast\", \"prompt\": \"<prompt>\", \"duration_seconds\": 5, \"aspect_ratio\": \"16:9\"}]}\nwith tempfile.NamedTemporaryFile(mode=\"w\", suffix=\".json\", delete=False) as f:\n    json.dump(payload, f)\n    tmpfile = f.name\nprint(tmpfile)\n```\n\n### Secure Header Pattern\n\nWrite auth headers to a temp file to keep secrets out of `/proc/*/cmdline`. Source credentials at the start of each command chain.\n\n**Image/Video requests** (X-API-Key + X-API-Secret):\n```bash\n. \"${KATANA_SECRETS_FILE:-$HOME/.openclaw/secrets/katana.env}\" && _H=$(mktemp) && chmod 600 \"$_H\" && printf 'Content-Type: application/json\\nX-API-Key: %s\\nX-API-Secret: %s\\n' \"$KATANA_API_KEY\" \"$KATANA_API_SECRET\" > \"$_H\" && curl -s -X POST \"https://kat.imgnai.com/v1/generation-requests?wait=false\" -H @\"$_H\" -d @\"$tmpfile\" && rm -f \"$_H\"\n```\n\n**Text/LLM requests** (Bearer auth):\n```bash\n. \"${KATANA_SECRETS_FILE:-$HOME/.openclaw/secrets/katana.env}\" && _H=$(mktemp) && chmod 600 \"$_H\" && printf 'Content-Type: application/json\\nAuthorization: Bearer %s:%s\\n' \"$KATANA_API_KEY\" \"$KATANA_API_SECRET\" > \"$_H\" && curl -s -X POST \"https://kat.imgnai.com/v1/chat/completions\" -H @\"$_H\" -d @\"$tmpfile\" && rm -f \"$_H\"\n```\n\nParse the JSON response. Extract `request_id`. Deliver confirmation to the user (model, cost, request_id).\n\n### Image Generation Notes\n- **`aspect_ratio: \"auto\"`** inspects the first image input and chooses closest supported ratio. Defaults to `1:1` if no image supplied.\n- **`is_fast`/`fast_mode`** request lower-cost half-resolution generation (imgnAI-hosted models only)\n- **`is_uhd`/`uhd_mode`** request UHD generation (imgnAI-hosted models only). Takes precedence over `is_fast`.\n- **`use_assistant`/`prompt_assist`** translate natural language to tag-style prompts (tag/booru models only)\n- **`output_format`** accepts `png`, `jpeg`, or `webp` (`jpg` is alias for `jpeg`)\n\n---\n\n## Credit Balance\n\n```bash\n. \"${KATANA_SECRETS_FILE:-$HOME/.openclaw/secrets/katana.env}\" && _H=$(mktemp) && chmod 600 \"$_H\" && printf 'Authorization: Bearer %s:%s\\n' \"$KATANA_API_KEY\" \"$KATANA_API_SECRET\" > \"$_H\" && curl -s \"https://kat.imgnai.com/v1/me/balance\" -H @\"$_H\" && rm -f \"$_H\"\n```\n\nCalls `GET /v1/me/balance`. The API returns `credits` as a decimal string. Converts to USD using current credit rate (see `{baseDir}/models.md`).\n\n---\n\n## Video Media Input Rules\n\nPut video media inputs in `video_image_data`:\n\n- **`first_frame_image_url`**: first/source frame image (HTTPS URL, data URL, or raw base64)\n- **`mid_frame_image_url`**: mid-frame image (only if model supports it)\n- **`last_frame_image_url`**: last/end frame image (only if model supports it)\n- **`reference_image_urls`**: array of reference images (only if model supports reference images). Obey `maximum_reference_images`\n- **`audio_input_urls`**: array of audio reference URLs (only if model supports audio input). Obey `maximum_reference_audio_files` and global cap of 4\n- **`video_list`**: array of video input clip objects. Each requires `url`. Optional `start`/`ends` second offsets when model has `video_offset_allowed`. Only for models with `supports_video_input: true`. Obey `maximum_reference_videos`.\n- **Audio output:** Most modern video models generate audio by default. Legacy models with `audio_gen_model: false` produce silent output. Check model metadata for audio capabilities.\n\nCompatibility aliases: top-level `image_url`, `input_image_url`, `input_image`, `input_image_b64` map to `first_frame_image_url`. Top-level `reference_image_urls` maps to `video_image_data.reference_image_urls`.\n\n**Rules:**\n- Do NOT send local filesystem paths, `file://` URLs, or `http://` media URLs — use HTTPS URLs or Base64 data URLs\n- Do NOT send `video_list` to models where `supports_video_input` is false/missing\n- Do NOT send audio references to models where `supports_audio_input` is false\n- Do NOT create an empty `video_image_data` object — omit missing fields entirely\n- Do NOT mix first/last frame with reference images unless the model's `custom_rules` allow it\n- Use only durations from `video_lengths_and_costs` and aspect ratios from `supported_aspects`\n- `aspect_ratio: \"auto\"` uses the first frame first, then first reference image; defaults to `1:1`\n\n## Video Custom Rules Glossary\n\nAlways inspect each video model's `custom_rules` before composing requests:\n\n| Rule | Description |\n|------|-------------|\n| `audio_15s_max` | Combined audio input limited to 15 seconds |\n| `audio_drives_duration` | Video duration follows audio duration |\n| `audio_ff_only` | Audio only with first-frame conditioning |\n| `audio_needs_reference_image` | Audio input requires at least one reference image |\n| `audio_or_fflf_exclusive` | Audio cannot combine with first/last frame |\n| `input_video_drives_length` | Input video clip drives output length |\n| `lf_needs_ff` | Last frame requires a first frame |\n| `reference_ff_only` | Reference images may combine with first-frame only |\n| `reference_is_voice_timbre` | Reference audio interpreted as voice timbre when images present |\n| `reference_no_ff_or_lf` | Reference images cannot combine with first/last frame |\n| `video_offset_allowed` | Model accepts `start`/`ends` second offsets in `video_list` |\n| `video_required` | Model requires at least one `video_list` object |\n\n## Common Failure Cases\n\n- **Unsupported aspect ratio:** Choose from model's `supported_aspects`\n- **Unsupported duration:** Choose from model's `video_lengths_and_costs`\n- **Corrupt base64 image:** Validate data decodes to actual image before submitting\n- **Local file path as media:** Convert to Base64 data URL first — API cannot fetch caller's filesystem\n- **Audio on unsupported model:** Only send `audio_input_urls` when `supports_audio_input: true`\n- **Video metadata on unsupported model:** Only send `video_list` when model has matching support flag\n- **Missing required video input:** Models with `video_required` must include `video_list`\n- **Video offsets on unsupported model:** Only send `start` and `ends` in `video_list` objects when the model has `video_offset_allowed` in `custom_rules`\n- **Last frame without first frame:** Prohibited on models with `lf_needs_ff`\n- **Reference images mixed with frames on incompatible model:** Check `custom_rules`, especially `reference_no_ff_or_lf`\n- **Too many reference images:** Clamp to `multi_image_inputs_allowed` (image) or `maximum_reference_images` (video)\n\n---\n\n## reference_assets (Typed Asset System)\n\n`reference_assets` is an alternative to `image_urls`/`video_image_data` for providing media inputs with explicit role labels. Each asset has a `kind` and either `url` or `base64_data`.\n\n### Image models\n\nAccepted image-like asset kinds:\n- `source_image` — primary source/input image\n- `image` — generic image input\n- `mask` — mask for inpainting/editing\n- `style_reference` — style transfer reference\n- `start_frame` — starting frame for animation\n\nExample:\n```json\n{\n  \"reference_assets\": [\n    {\"kind\": \"source_image\", \"url\": \"https://example.com/product.png\"},\n    {\"kind\": \"style_reference\", \"base64_data\": \"data:image/jpeg;base64,...\"}\n  ]\n}\n```\n\n### Video models\n\nImage kinds for video:\n- `style_reference`, `reference_image`, `image` — map to video reference images\n\nAudio kinds for video:\n- `audio`, `source_audio`, `reference_audio`, `audio_reference` — map to audio reference inputs\n\nExample:\n```json\n{\n  \"reference_assets\": [\n    {\"kind\": \"reference_image\", \"url\": \"https://example.com/person.png\"},\n    {\"kind\": \"audio\", \"url\": \"https://example.com/voice.mp3\"}\n  ]\n}\n```\n\n---\n\n## llms.txt Freshness\n\nThis skill was built from the Katana API llms.txt reference document.\n\n**Last synced:** 2026-05-30\n**llms.txt URL:** https://kat.imgnai.com/llms.txt\n**Stored checksum:** `326fb64f9512c447bd3acf376a2c5d4583ba9ec28b8d8ad76b457459432e906a`\n\n### Pre-generation check\n\nBefore submitting ANY generation request, check if the llms.txt checksum has been verified in the last 24 hours. If stale:\n\n1. Fetch: `curl -s https://kat.imgnai.com/llms.txt`\n2. Compute SHA256: `sha256sum` (Linux) or `shasum -a 256` (macOS)\n3. Compare to stored checksum\n4. If CHANGED → tell the user: \"The Katana API model list has been updated since this skill was last synced. This may include new models, pricing changes, or removed models. Would you like me to check for changes and update the skill?\"\n5. If user says YES → parse new llms.txt, update models.md, update checksum and date\n6. If user says NO → proceed with current models\n7. Update last-checked date regardless\n\n### llms.txt update process\n\nWhen llms.txt changes, compare old vs new **holistically**. Diff the full documents — do not limit the review to a predefined checklist. Document ALL changes found and update all affected skill files accordingly: `models.md`, `SKILL.md`, workflow files.\n\n**DO NOT auto-update without user confirmation.**\n\n**Explicit approval rule:** During the llms.txt update process, always summarise ALL changes found and ask the user for explicit permission before updating any skill files (models.md, SKILL.md, workflow files). Do not auto-update without confirmation.\n\n---\n\n## Delivery Patterns\n\nDeliver the generated media to the user via your agent's messaging/file capability. Include: model name, resolution/dimensions, credits, dollar cost, description, and the full-res URL (`original_data_url`).\n\n### ⚠️ URL Display (MANDATORY)\n\nALL image and video deliveries MUST include the **full download URL** (`original_data_url`) as clickable text in the delivery message — not just the inline media attachment.\n\nUsers need the URL to:\n- Download the full-resolution file\n- Share it externally\n- Archive it before expiry\n\n**Include ALL URLs returned** — `original_data_url`, `thumbnail_image_url`, `final_frame_image_url`, `thumbnail_silent_video_mp4_url` — any URL the API returns for the asset. Do not assume the user only wants one.\n\nExample:\n```\nMEDIA:https://k.imgnai.com/abc123.mp4\n\n🔗 Full-res: https://k.imgnai.com/abc123.mp4\n🖼️ Thumbnail: https://k.imgnai.com/def456.jpg\n🎞️ Silent preview: https://k.imgnai.com/ghi789.mp4\n⏰ Expires: Fri 16 May 2026, 14:00 BST\n```\n\n### ⚠️ Expiry Warning (MANDATORY)\n\nALL image and video generation summaries MUST include:\n1. The **expiry timestamp** extracted from `responses[].output_assets[].expires_at` in the completed poll response — convert to user's local timezone for display\n2. A clear warning that content must be downloaded before expiry if the user wishes to keep it\n\nExample format:\n```\n⏰ Expires: Fri 16 May 2026, 14:00 BST — download before expiry if you need it long-term.\n```\n\n**Do NOT calculate expiry manually.** The API provides `expires_at` in the poll response. Use it directly. The 72h retention window may change server-side; `expires_at` is always authoritative.\n\nFor text/LLM: return the model's response verbatim. Then send a separate follow-up message with a cost summary per the \"Cost Reporting\" section above. Text completions do not require an expiry warning (no media URL to expire).\n\n---\n\n*Last updated: 2026-05-30*\n\nFile v1.0.2:README.md\n\n# Katana Skill — imgnAI API\n\nGenerate images, videos, and text/LLM completions via the [imgnAI Katana API](https://app.imgnai.com/katana-api). Supports end-to-end-encrypted (E2EE) and anonymized models. Priced highly competitively: can be 40-70% cheaper than Venice AI and other platforms.\n\nIncludes post-processing such as combining videos and images, cutting, slicing, splicing, transitions, drawing text, re-encoding, resizing and much more!\n\nA complete workflow for content creation from start to finish, all from the comfort of your agent.\n\n\n## Features\n\n- **Images:** 40+ models including GPT Image 2, imgnAI Gen, FLUX.2, Seedream, WAN, Imagine Art\n- **Videos:** 15+ models including Seedance 2.0, Kling O3 4K, WAN 2.7, Veo 3.1, LTX\n- **Text/LLM:** 30+ models including Grok 4.3, GPT-5.5, Claude Opus 4.7, DeepSeek V4, Qwen 3.6\n- **Pre-submission confirmation** with cost estimates before generating\n- **Adaptive polling** that respects API-recommended intervals\n- **Error handling** with clear user-facing messages\n- **Reference image support** for editing workflows\n- **Agent-agnostic:** works with OpenClaw, Hermes, Claude, or standalone\n\n## Setup\n\n1. Get your API key from https://app.imgnai.com/katana-api\n2. Create the secrets file:\n   ```bash\n   mkdir -p ~/.openclaw/secrets\n   cat > ~/.openclaw/secrets/katana.env << 'EOF'\n   KATANA_API_KEY=your_key_here\n   KATANA_API_SECRET=your_secret_here\n   EOF\n   chmod 600 ~/.openclaw/secrets/katana.env\n   ```\n3. Install the skill (see [Agent Integration](#agent-integration) or [Standalone Usage](#standalone-usage) below)\n\n## Agent Integration\n\nThis skill works with any agent framework. It provides a `SKILL.md` routing hub and workflow files that guide your agent t\n\nArchive v1.0.1: 11 files, 36324 bytes\n\nFiles: CHANGELOG.md (2097b), models.md (21497b), README.md (3664b), skill-card.md (2958b), SKILL.md (21379b), workflows/ffmpeg.md (20166b), workflows/image.md (5073b), workflows/post-process.md (3173b), workflows/text.md (7665b), workflows/video.md (5797b), _meta.json (125b)\n\nArchive v1.0.0: 10 files, 33048 bytes\n\nFiles: katana.sh (13445b), models.md (20425b), README.md (4471b), SKILL.md (13889b), workflows/ffmpeg.md (19382b), workflows/image.md (4092b), workflows/post-process.md (3170b), workflows/text.md (5919b), workflows/video.md (3318b), _meta.json (125b)","readmeExcerpt":"Skill: imgnAI Katana API Owner: imgn Summary: Generate images, videos, and text/LLM completions via the imgnAI Katana API. Supports end-to-end-encrypted (E2EE) and anonymized models. Priced highly compet... Tags: latest:1.0.3 Version history: v1.0.3 | 2026-06-09T12:30:34.455Z | user Added • q-naifu-a3b text model — imgnAI fully uncensored Agentic Model (Private tier, 262K context, vision + file input) • Audio output ","codeSnippets":[],"executableExamples":[{"language":"text","snippet":"KATANA_API_KEY=your_key_here\nKATANA_API_SECRET=your_secret_here"},{"language":"bash","snippet":". \"${KATANA_SECRETS_FILE:-$HOME/.openclaw/secrets/katana.env}\""},{"language":"text","snippet":"📊 Katana Summary\nModel: gemma-4-26b-a4b (Anonymized)\nRequest: bf11cf04-8747-480e-a7f7-7d6cb092c614\nTokens: 42 in / 176 out (text only)\nCost: 0.1 credits (~$0.001)\nPrivacy: Anonymized\nTime: ~3s"},{"language":"bash","snippet":". \"${KATANA_SECRETS_FILE:-$HOME/.openclaw/secrets/katana.env}\" && _H=$(mktemp) && chmod 600 \"$_H\" && printf 'X-API-Key: %s\\nX-API-Secret: %s\\n' \"$KATANA_API_KEY\" \"$KATANA_API_SECRET\" > \"$_H\" && curl -s \"https://kat.imgnai.com/v1/generation-requests/${REQUEST_ID}\" -H @\"$_H\" && rm -f \"$_H\""},{"language":"bash","snippet":"python3 -c \"\nimport sys,json\nd=json.load(sys.stdin)\nr=d.get('responses',[])\nfor i in range(len(r)):\n    ri=r[i]; st=ri.get('status','?')\n    print(f'Status: {st}')\n    for a in ri.get('output_assets',[]):\n        print(f'URL: {a.get(\"original_data_url\",\"\")}')\n        print(f'Dims: {a.get(\"width\",\"?\")}x{a.get(\"height\",\"?\")}')\n        print(f'Expires: {a.get(\"expires_at\",\"\")}')\n    print(f'Credits: {ri.get(\"metadata\",{}).get(\"credits_spent\",\"?\")}')\n    if st=='failed':\n        e=ri.get('error',{}); print(f'Error: {e.get(\"message\",\"\")} retryable={e.get(\"retryable\",\"\")}')\n\""},{"language":"python","snippet":"import json, datetime, os\nbase = os.environ.get('KATANA_STATE_DIR', os.path.dirname(os.environ.get('KATANA_SECRETS_FILE', os.path.expanduser('~/.openclaw/secrets/katana.env'))))\npath = os.path.join(base, 'katana_pending.json')\nmeta = {\n    'request_id': 'REQUEST_ID',\n    'model': 'MODEL',\n    'credits': CREDITS,\n    'submitted': datetime.datetime.now().isoformat(),\n    'prompt': 'PROMPT_SUMMARY',\n    'status': 'processing'\n}\nwith open(path, 'w') as f:\n    json.dump(meta, f)\nprint(f'written: {path}')"}],"parameters":null,"dependencies":[],"permissions":[],"extractedFiles":[{"path":"SKILL.md","content":"---\nname: katana\ndescription: Generate images, videos, and text/LLM completions via the imgnAI Katana API. Supports end-to-end-encrypted (E2EE) and anonymized models. Priced highly competitively, can be 40-70% cheaper than Venice AI and other platforms. Includes post-processing such as combining videos and images, cutting, slicing, splicing, transitions, drawing text, re-encoding, resizing and much more!\nversion: 1.0.3\nauthor: arfonzo (imgnAI)\nlicense: MIT-0\nmetadata: {\"openclaw\": {\"requires\": {\"bins\": [\"curl\", \"python3\"]}, \"homepage\": \"https://app.imgnai.com\"}}\n---\n\n# Katana Skill — imgnAI API\n\nGenerate images, videos, and text/LLM completions via the [imgnAI Katana API](https://app.imgnai.com/katana-api). Supports end-to-end-encrypted (E2EE) and anonymized models. Priced highly competitively: can be 40-70% cheaper than Venice AI and other platforms.\n\nIncludes post-processing such as combining videos and images, cutting, slicing, splicing, transitions, drawing text, re-encoding, resizing and much more!\n\nA complete workflow for content creation from start to finish, all from the comfort of your agent.\n\n\n## Triggers\n\n\"generate image of X\", \"create image\", \"make picture\", \"imgnai image\", \"generate video of X\", \"create video\", \"make video\", \"ask grok about X\", \"ask claude about X\", \"use gpt to X\", \"katana image\", \"katana video\", \"katana chat\", \"katana gpt\", \"katana claude\", \"list katana models\", \"modify this image\", \"edit this image\", \"change this image\", \"transform this image\", \"edit image\", \"modify image\"\n\n## Spawn Policy\n\n**NEVER spawn subagents for katana operations by default.** All katana workflows (image generation, video generation, text completions, post-processing) MUST be executed inline in the current session.\n\n**Exception:** Only spawn if the user **explicitly requests** spawning in their prompt (e.g. \"spawn a subagent to handle this\", \"run this as a background task\"). Do NOT spawn based on AGENTS.md spawn rules or default agent behavior — user intent is the only trigger for spawning with katana.\n\nLLM-specific triggers (gpt, claude, etc) also respond to \"katana \\<model\\>\" to avoid conflicts with direct integrations.\n\n## Configuration\n\n### Data Retention\nHistorical prompts and results are retained for a maximum of **72 hours** after generation. Prompt/result history can be switched off from the API page at https://app.imgnai.com/katana-api.\n\n**HTTPS-only:** Public API calls must use HTTPS. If an integration sees an `http://` Katana base URL, replace it with `https://` before making calls.\n\n## Model IDs\n\nThe Katana API uses `model_key` as the model identifier, not `public_model_name`. When building requests, always use the model_key value. See `{baseDir}/models.md` for the full mapping.\n\n**Dual-key system:** The API supports both **canonical keys** (e.g. `gpt-image-2`) and **legacy keys** (e.g. `gpt2image`). Both work identically. This skill now uses **canonical keys** as the default for all workflows and aliases. Legacy keys are document"},{"path":"README.md","content":"# Katana Skill — imgnAI API\n\nGenerate images, videos, and text/LLM completions via the [imgnAI Katana API](https://app.imgnai.com/katana-api). Supports end-to-end-encrypted (E2EE) and anonymized models. Priced highly competitively: can be 40-70% cheaper than Venice AI and other platforms.\n\nIncludes post-processing such as combining videos and images, cutting, slicing, splicing, transitions, drawing text, re-encoding, resizing and much more!\n\nA complete workflow for content creation from start to finish, all from the comfort of your agent.\n\n\n## Features\n\n- **Images:** 40+ models including GPT Image 2, imgnAI Gen, FLUX.2, Seedream, WAN, Imagine Art\n- **Videos:** 15+ models including Seedance 2.0, Kling O3 4K, WAN 2.7, Veo 3.1, LTX\n- **Text/LLM:** 30+ models including Grok 4.3, GPT-5.5, Claude Opus 4.7, DeepSeek V4, Qwen 3.6\n- **Pre-submission confirmation** with cost estimates before generating\n- **Adaptive polling** that respects API-recommended intervals\n- **Error handling** with clear user-facing messages\n- **Reference image support** for editing workflows\n- **Agent-agnostic:** works with OpenClaw, Hermes, Claude, or standalone\n\n## Setup\n\n1. Get your API key from https://app.imgnai.com/katana-api\n2. Create the secrets file:\n   ```bash\n   mkdir -p ~/.openclaw/secrets\n   cat > ~/.openclaw/secrets/katana.env << 'EOF'\n   KATANA_API_KEY=your_key_here\n   KATANA_API_SECRET=your_secret_here\n   EOF\n   chmod 600 ~/.openclaw/secrets/katana.env\n   ```\n\n   **Non-OpenClaw users:** Set `KATANA_SECRETS_FILE` to your preferred location:\n   ```bash\n   mkdir -p ~/.config/katana\n   cat > ~/.config/katana/katana.env << 'EOF'\n   KATANA_API_KEY=your_key_here\n   KATANA_API_SECRET=your_secret_here\n   EOF\n   chmod 600 ~/.config/katana/katana.env\n   export KATANA_SECRETS_FILE=~/.config/katana/katana.env\n   ```\n\n3. Install the skill (see [Agent Integration](#agent-integration) or [Standalone Usage](#standalone-usage) below)\n\n## Agent Integration\n\nThis skill works with any agent framework. It provides a `SKILL.md` routing hub and workflow files that guide your agent through image, video, and text generation.\n\n### OpenClaw (example)\n\n1. Follow the [Setup](#setup) steps above\n2. Install the skill:\n   ```bash\n   openclaw skill install katana\n   ```\n   Or manually: clone/copy the `katana/` directory into your skills path (typically `~/.openclaw/skills/` or `~/workspace/skills/`).\n3. Ask your assistant to generate: \"Generate an image of a cat riding a skateboard\"\n\n### Other Agent Frameworks\n\nPoint your agent to `SKILL.md` as the entry point. The skill resolves `{baseDir}` dynamically from the file's location. Set `KATANA_SECRETS_FILE` if your secrets are stored elsewhere.\n\n## Prerequisites\n\n- **curl** — API requests\n- **jq** or **python3** — JSON parsing (jq preferred, python3 as fallback)\n- **ffmpeg** (optional) — Post-processing\n\n### ffmpeg optional capabilities\n\nIf you install ffmpeg for post-processing:\n- **Text overlays / drawtext:** requires ffmpeg built with `--enable-lib"},{"path":"_meta.json","content":"{\n  \"ownerId\": \"kn75xgwvf577we4f2ej81pg3hx86p4c7\",\n  \"slug\": \"katana\",\n  \"version\": \"1.0.3\",\n  \"publishedAt\": 1781008234455\n}"},{"path":"CHANGELOG.md","content":"# Changelog\n\n## 1.0.3\n\n### Added\n- `q-naifu-a3b` text model — imgnAI fully uncensored Agentic Model (Private tier, 262K context, vision + file input)\n- `naifu` / `q-naifu` model alias mapping to `q-naifu-a3b`\n- Text polling for long-running models — submit with `?wait=false`, poll via `GET /v1/generation-requests/{id}`\n- Audio output docs for video models — which models generate audio, how to strip it\n- Three-tier privacy documentation — Anonymized, Private, E2EE Private with attestation details\n- Error codes quick reference table\n- Agent-agnostic terminology glossary\n- Non-OpenClaw setup guide in README\n\n### Changed\n- Claude alias now points to `claude-opus-4-8`; added `claude-fast` for `claude-opus-4-8-fast`\n- Poll timeout guard: 10 min for image/text, 100 min for video (matches API limits)\n- Persistence file supports `KATANA_STATE_DIR` to separate state from credentials\n- Cache write column formatting fixed for models with no cache write pricing\n\n### Fixed\n- Text submission auth: curl commands now use single `&&`-chained line (fixes failures on some agents)\n- Payload temp files are now cleaned up after each request\n- Header auth bug: submit and poll now use consistent two-header format\n- `gpt-image-2-max` reference image count: 12 → 10\n- `***` literal removed from header format strings\n\n## 1.0.2\n\n### Added\n- `grok-build-0-1` — Grok Build 0.1, xAI coding model (256K ctx)\n- `claude-opus-4-8` and `claude-opus-4-8-fast` text models\n- Flux 1.1 Ultra, Flux Kontext Max/Pro, GPT-5.4/5.5, Claude Opus 4.7/Sonnet 4.6/Haiku 4.5, Grok 4.20/4.20 Multi-Agent, DeepSeek V4 Flash/Pro\n- Video media input rules (`video_image_data` fields documented)\n- Video custom rules glossary (12 rules from API)\n- Common failure cases section\n- Text/LLM notes: streaming, vision/multimodal, billing, attestation, refund policy\n- Image generation notes: auto aspect ratio, fast/UHD modes, prompt assist\n- `thumbnail_silent_video_mp4_url` and `final_frame_image_url` to response handling\n- Cache read pricing for 11 models (8 private + qwen3-6-flash, qwen3-6-max-preview, qwen3-6-plus, qwen3-7-max)\n\n### Changed\n- Full models.md rebuild with canonical dashed keys throughout\n- All model keys migrated to canonical format (e.g. `seedance2` → `seedance-2-0`)\n- Price cuts: qwen3-7-max (-50%), deepseek-v4-flash (-30%), qwen3-6-flash (-25%), glm-5-1 (-12%), kimi-k2-6 (-9%), minimax-m2-7 (-13%), qwen3-6-35b-a3b (-7%)\n\n## 1.0.1\n\n### Added\n- `pink-image` model — 1 credit, high-speed generalist\n- `qwen3-7-max` text model — flagship Qwen, 1M context\n- `gemini-3-5-flash` text model — near-Pro at Flash cost, multimodal\n- `gpt-image-2-max` image model — MAX variant, 28 cr, QHD output\n- `gemini-omni` video model — Google video, 4-10s, 5 ref images\n- `gemini-omni-v2v` video model — V2V with `video_list` input\n- Text alias `qwen-max`, `gemini-35-flash`; video aliases `gemini-omni`, `gemini-v2v`\n- V2V (video-to-video) workflow section in `workflows/video.md`\n- Custom rules: `video_required`, `video_offset"},{"path":"models.md","content":"# Katana Model Catalogue\n\n> **Production API uses `model_key` as the model ID parameter, not `public_model_name`. All IDs below are model_key values.**\n\n> **Note:** This is a static snapshot synced from the live API reference at https://kat.imgnai.com/llms.txt. For the most current pricing and model availability, check the live endpoint.\n\nAuto-generated from https://kat.imgnai.com/llms.txt — Last synced: 2026-06-08\n\nExtracted from the imgnAI Katana API docs. Organised by type: text, image, video.\n\nReference price: $0.0052 per credit (Platinum Annual).\n\n## Text / LLM Models\n\nAll text models use `POST /v1/chat/completions` (OpenAI-compatible).\nAuth: `Authorization: Bearer <api_key>:<api_secret>`\nBilling: pre-charge reserve, then refund/charge difference from actual usage. Minimum 0.1 credits.\n\n| Model ID | Publisher | Context | Max Output | Input Types | Privacy | In (cr/$) | Out (cr/$) | Cache R / W | Legacy |\n|---|---|---|---|---|---|---|---|---|---|---|\n| `q-naifu-a3b` | imgnAI | 262144 | 262144 | text, image, file | Private | 38.5 / $0.20 | 240.4 / $1.25 | — | — |\n| `claude-opus-4-8` | Anthropic | 1000000 | 128000 | text, image, file | Anonymized | 1000.0 / $5.20 | 5000.0 / $26.00 | 100.0 cr ($0.52) / 1250.0 cr ($6.50) | — |\n| `claude-opus-4-8-fast` | Anthropic | 1000000 | 128000 | text, image, file | Anonymized | 2000.0 / $10.40 | 10000.0 / $52.00 | 200.0 cr ($1.04) / 2500.0 cr ($13.00) | — |\n| `qwen3-7-max` | Qwen | 1000000 | 65536 | text | Anonymized | 264.5 / $1.38 | 793.3 / $4.13 | 52.9 cr ($0.2750) / 330.6 cr ($1.72) | — |\n| `grok-build-0-1` | xAI | 256000 | not listed | text, image | Anonymized | 192.4 / $1.00 | 384.7 / $2.00 | 38.5 cr ($0.20) / — | — |\n| `gemini-3-5-flash` | Google | 1048576 | 65536 | text, image, video, file, audio | Anonymized | 317.4 / $1.65 | 1903.9 / $9.90 | 31.8 cr ($0.1650) / 17.7 cr ($0.0917) | — |\n| `grok-4-3` | xAI | 1000000 | not listed | text, image | Anonymized | 264.5 / $1.38 | 528.9 / $2.75 | 42.4 cr ($0.2200) / — | — |\n| `qwen3-6-35b-a3b` | Qwen | 262144 | 262140 | text, image, video | Anonymized | 29.7 / $0.1540 | 211.6 / $1.10 | — | — |\n| `qwen3-6-flash` | Qwen | 1000000 | 65536 | text, image, video | Anonymized | 39.7 / $0.2062 | 238.0 / $1.24 | 4.0 cr ($0.0206) / 49.6 cr ($0.2578) | — |\n| `qwen3-6-max-preview` | Qwen | 262144 | 65536 | text | Anonymized | 220.0 / $1.14 | 1320.0 / $6.86 | 22.0 cr ($0.1144) / 275.0 cr ($1.43) | — |\n| `deepseek-v4-flash` | DeepSeek | 1048576 | 384000 | text | Anonymized | 20.8 / $0.1081 | 41.6 / $0.2163 | 4.2 cr ($0.0217) / — | — |\n| `deepseek-v4-pro` | DeepSeek | 1048576 | 384000 | text | Anonymized | 92.1 / $0.4785 | 184.1 / $0.9570 | 0.8 cr ($0.0040) / — | — |\n| `gpt-5-5` | OpenAI | 1050000 | 128000 | file, image, text | Anonymized | 1057.7 / $5.50 | 6346.2 / $33.00 | 105.8 cr ($0.5500) / — | — |\n| `kimi-k2-6-private` | MoonshotAI | 262144 | 262144 | text, image | E2EE Private | 230.6 / $1.20 | 973.1 / $5.06 | 78.3 cr ($0.4070) | — |\n| `qwen3-coder-next-private` | Qw"}],"languages":[],"docsSourceLabel":"CLAWHUB","editorialOverview":null,"editorialQuality":{"score":100,"threshold":65,"status":"thin","wordCount":2352,"uniquenessScore":38,"reasons":["uniqueness-below-45"]}},"media":{"evidence":{"source":"no-media","verified":false,"confidence":"low","updatedAt":"2026-10-11T13:45:08.039Z","emptyReason":"No screenshots, media assets, or demo links are available."},"primaryImageUrl":null,"mediaAssetCount":0,"assets":[],"demoUrl":null},"ownerResources":{"evidence":{"source":"unclaimed","verified":false,"confidence":"low","updatedAt":"2026-10-11T13:45:08.039Z","emptyReason":"This page has not been claimed by the agent owner."},"hasCustomPage":false,"customPageUpdatedAt":null,"customLinks":[],"structuredLinks":{"docsUrl":null,"demoUrl":null,"supportUrl":null,"pricingUrl":null,"statusUrl":null},"customPage":null},"relatedAgents":{"evidence":{"source":"protocol-neighbors","verified":false,"confidence":"medium","updatedAt":"2026-10-11T16:01:20.357Z","emptyReason":null},"items":[{"id":"8ebccd8e-3863-4187-8355-c3f14e1f9edf","entityType":"agent","canonicalPath":"/agent/iofficeai-aionui","slug":"iofficeai-aionui","name":"AionUi","description":"Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!","url":"https://github.com/iOfficeAI/AionUi","homepage":"https://www.aionui.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-10-09T19:11:12.944Z","createdAt":"2026-02-25T03:38:16.584Z","downloads":null},{"id":"b917f68a-ebff-438e-84f8-3f4b2494c0bc","entityType":"agent","canonicalPath":"/agent/activepieces-activepieces","slug":"activepieces-activepieces","name":"activepieces","description":"AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents","url":"https://github.com/activepieces/activepieces","homepage":"https://www.activepieces.com","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-15T02:22:12.426Z","createdAt":"2026-02-25T03:38:12.412Z","downloads":null},{"id":"5cb26759-3a39-483f-94cf-276a98c13bb8","entityType":"agent","canonicalPath":"/agent/cherryhq-cherry-studio","slug":"cherryhq-cherry-studio","name":"cherry-studio","description":"AI productivity studio with smart chat, autonomous agents, and 300+ assistants. Unified access to frontier LLMs","url":"https://github.com/CherryHQ/cherry-studio","homepage":"https://cherry-ai.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-11T14:38:40.986Z","createdAt":"2026-02-25T03:38:19.379Z","downloads":null},{"id":"6f6582d0-5d76-4f0f-b81d-86520247950b","entityType":"agent","canonicalPath":"/agent/copilotkit-copilotkit","slug":"copilotkit-copilotkit","name":"CopilotKit","description":"The Frontend for Agents & Generative UI. React + Angular","url":"https://github.com/CopilotKit/CopilotKit","homepage":"https://docs.copilotkit.ai","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-03-25T09:50:57.846Z","createdAt":"2026-02-25T03:39:14.617Z","downloads":null}],"links":{"hub":"/agent","source":"/agent/source/clawhub","protocols":[{"label":"OpenClaw","href":"/agent/protocol/openclew"}]}}}