{"id":"81cf9faa-c86e-452d-bac2-85c9dedff6ae","entityType":"agent","slug":"clawhub-jimliu-baoyu-imagine","name":"Baoyu Imagine","canonicalUrl":"https://www.xpersona.co/agent/clawhub-jimliu-baoyu-imagine","canonicalPath":"/agent/clawhub-jimliu-baoyu-imagine","generatedAt":"2026-10-09T22:38:20.864Z","source":"CLAWHUB","claimStatus":"UNCLAIMED","verificationTier":"NONE","summary":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-09T09:29:37.766Z","emptyReason":null},"description":"AI image generation with OpenAI GPT Image 2, Azure OpenAI, Google, OpenRouter, DashScope, Z.AI GLM-Image, MiniMax, Jimeng, Seedream and Replicate APIs. Suppo...","descriptionLabel":"Source description","evidenceSummary":"Capability contract not published. No trust telemetry is available yet. 3.2K downloads reported by the source. Last updated 10/9/2026.","installCommand":"clawhub skill install s17dfbhg0khk4fvtnx0stqg0yx83j25z:baoyu-imagine","sourceUrl":"https://clawhub.ai/jimliu/baoyu-imagine","homepage":"https://clawhub.ai/jimliu/skills/baoyu-imagine","primaryLinks":[{"label":"View on ClawHub","url":"https://clawhub.ai/jimliu/baoyu-imagine","kind":"source"},{"label":"Homepage","url":"https://clawhub.ai/jimliu/skills/baoyu-imagine","kind":"homepage"}],"safetyScore":84,"overallRank":62,"popularityScore":57,"trustScore":null,"claimedByName":null,"isOwner":false,"seoDescription":"Baoyu Imagine technical dossier on Xpersona with agent coverage, OPENCLEW support, and live trust metadata."},"coverage":{"evidence":{"source":"public-profile","verified":false,"confidence":"medium","updatedAt":"2026-10-09T09:29:37.766Z","emptyReason":null},"protocols":[{"protocol":"OPENCLEW","label":"OpenClaw","status":"self-declared","notes":"Declared in the public agent profile."}],"capabilities":[],"verifiedCount":0,"selfDeclaredCount":1,"capabilityMatrix":{"rows":[{"key":"OPENCLEW","type":"protocol","support":"unknown","confidenceSource":"profile","notes":"Listed on profile"}],"flattenedTokens":"protocol:OPENCLEW|unknown|profile"}},"adoption":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-09T09:29:37.766Z","emptyReason":null},"stars":null,"forks":null,"downloads":3154,"packageName":null,"latestVersion":"1.117.3","tractionLabel":"3.2K downloads"},"release":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-09T09:29:37.766Z","emptyReason":null},"lastUpdatedAt":"2026-10-09T09:29:37.766Z","lastCrawledAt":"2026-10-09T09:29:37.766Z","lastIndexedAt":null,"nextCrawlAt":"2026-10-10T09:29:37.766Z","lastVerifiedAt":null,"highlights":[{"version":"1.117.3","createdAt":"2026-05-24T22:13:53.081Z","changelog":"- Added documentation for reference-image identity preservation, including best practices and anti-patterns. - Updated usage examples with new guidance for prompts when preserving real subjects from reference images. - Added two new reference documents: `codex-image2-fallback.md` and `codex-oauth-vs-openai-api-key.md`. - Clarified recommendations for avoiding drift and over-stylization in identity-preserving generation workflows.","fileCount":38,"zipByteSize":96871},{"version":"1.117.2","createdAt":"2026-05-18T02:15:08.352Z","changelog":"## 1.117.2 - 2026-05-17 ### Documentation - `baoyu-cover-image`: ban programmatic text repair on generated bitmaps — disallow ImageMagick / Pillow / Canvas / SVG / HTML overlays to cover, rewrite, or replace title/subtitle text; regenerate from a corrected prompt or switch to a lower-text or no-title variant instead - `baoyu-article-illustrator`, `baoyu-comic`, `baoyu-image-cards`, `baoyu-xhs-images`, `baoyu-infographic`, `baoyu-slide-deck`: sync the same text-repair ban with skill-specific text categories (labels/captions, dialogue/sound effects, titles/body/tags, headings/data values, slide titles/bullets)","fileCount":35,"zipByteSize":91535},{"version":"1.115.4","createdAt":"2026-05-11T23:49:03.373Z","changelog":"### Documentation - Image generation backend selection: emphasize Codex `imagegen` as the priority runtime-native tool (invoke via the `Skill` tool with `skill: \"imagegen\"`) and forbid SVG/HTML/canvas substitution when no raster backend can be resolved — fall through to asking the user instead of silently emitting code-based art. Updated in `docs/image-generation-tools.md` and inlined into `baoyu-article-illustrator`, `baoyu-comic`, `baoyu-cover-image`, `baoyu-image-cards`, `baoyu-infographic`, `baoyu-slide-deck`, and `baoyu-xhs-images`.","fileCount":35,"zipByteSize":91536},{"version":"1.115.1","createdAt":"2026-05-10T03:12:19.033Z","changelog":"## 1.115.1 - 2026-05-10 ### Fixes - `baoyu-imagine`: change the default MiniMax image API endpoint to `https://api.minimaxi.com` to match the current official image generation documentation, while keeping `https://api.minimax.io` available through `MINIMAX_BASE_URL` overrides. - `baoyu-image-gen`: sync the deprecated image-generation entrypoint with the same MiniMax default endpoint and regression coverage.","fileCount":35,"zipByteSize":91535},{"version":"1.104.0","createdAt":"2026-04-21T20:15:53.791Z","changelog":"## baoyu-imagine 1.104.0 - Add `gpt-image-2` support for OpenAI image generation and edits. - Make `gpt-image-2` the default OpenAI model and update Azure deployment guidance. - Document official size and quality mapping, custom-size constraints, and 4K usage examples.","fileCount":35,"zipByteSize":87161},{"version":"1.103.1","createdAt":"2026-04-13T16:15:09.162Z","changelog":"## 1.103.1 - 2026-04-13 ### Fixes - `baoyu-markdown-to-html`: decode HTML entities and strip tags from article summary - `baoyu-post-to-weibo`: decode HTML entities and strip tags from article summary","fileCount":27,"zipByteSize":79275},{"version":"1.103.0","createdAt":"2026-04-13T01:20:08.024Z","changelog":"## 1.103.0 - 2026-04-12 ### Features - baoyu-diagram: add multi-diagram mode for article-wide diagram generation ### Fixes - baoyu-article-illustrator: prevent color names and hex codes from appearing as visible text in generated images - baoyu-cover-image: prevent color names and hex codes from appearing as visible text in generated images - baoyu-image-cards: prevent color names from appearing as visible text in generated images - baoyu-post-to-wechat: decode HTML entities and strip tags from article summary","fileCount":27,"zipByteSize":79276},{"version":"1.0.1","createdAt":"2026-03-26T01:11:22.604Z","changelog":"- Adds legacy compatibility: if `.baoyu-skills/baoyu-image-gen/EXTEND.md` exists and the new path does not, it will be renamed to `.baoyu-skills/baoyu-imagine/EXTEND.md` at runtime. - Both old and new EXTEND.md can coexist; if both exist, the new path is always used. - No user action required; migration is automatic for existing users upgrading from baoyu-image-gen.","fileCount":25,"zipByteSize":66326}]},"execution":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No published capability contract is available yet."},"installCommand":"clawhub skill install s17dfbhg0khk4fvtnx0stqg0yx83j25z:baoyu-imagine","setupComplexity":"low","setupSteps":["Install using `clawhub skill install s17dfbhg0khk4fvtnx0stqg0yx83j25z:baoyu-imagine` in an isolated environment before connecting it to live workloads.","No published capability contract is available yet, so validate auth and request/response behavior manually.","Review the upstream CLAWHUB listing at https://clawhub.ai/jimliu/baoyu-imagine before using production credentials."],"contract":{"contractStatus":"missing","authModes":[],"requires":[],"forbidden":[],"supportsMcp":false,"supportsA2a":false,"supportsStreaming":false,"inputSchemaRef":null,"outputSchemaRef":null,"dataRegion":null,"contractUpdatedAt":null,"sourceUpdatedAt":null,"freshnessSeconds":null},"invocationGuide":{"preferredApi":{"snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-jimliu-baoyu-imagine/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-jimliu-baoyu-imagine/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-jimliu-baoyu-imagine/trust"},"curlExamples":["curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-jimliu-baoyu-imagine/snapshot\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-jimliu-baoyu-imagine/contract\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-jimliu-baoyu-imagine/trust\""],"jsonRequestTemplate":{"query":"summarize this repo","constraints":{"maxLatencyMs":2000,"protocolPreference":["OPENCLEW"]}},"jsonResponseTemplate":{"ok":true,"result":{"summary":"...","confidence":0.9},"meta":{"source":"CLAWHUB","generatedAt":"2026-10-09T22:38:20.861Z"}},"retryPolicy":{"maxAttempts":3,"backoffMs":[500,1500,3500],"retryableConditions":["HTTP_429","HTTP_503","NETWORK_TIMEOUT"]}},"endpoints":{"dossierUrl":"https://www.xpersona.co/api/v1/agents/clawhub-jimliu-baoyu-imagine/dossier","snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-jimliu-baoyu-imagine/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-jimliu-baoyu-imagine/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-jimliu-baoyu-imagine/trust"}},"reliability":{"evidence":{"source":"runtime-metrics","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No trust, reliability, or runtime telemetry is available."},"trust":{"status":"unavailable","handshakeStatus":"UNKNOWN","verificationFreshnessHours":null,"reputationScore":null,"p95LatencyMs":null,"successRate30d":null,"fallbackRate":null,"attempts30d":null,"trustUpdatedAt":null,"trustConfidence":"unknown","sourceUpdatedAt":null,"freshnessSeconds":null},"decisionGuardrails":{"doNotUseIf":["Contract metadata is missing or unavailable for deterministic execution."],"safeUseWhen":[],"riskFlags":["missing_or_unavailable_contract","trust_data_unavailable","schema_references_missing"],"operationalConfidence":"low"},"executionMetrics":{"observedLatencyMsP50":null,"observedLatencyMsP95":null,"estimatedCostUsd":null,"uptime30d":null,"rateLimitRpm":null,"rateLimitBurst":null,"lastVerifiedAt":null,"verificationSource":null},"runtimeMetrics":{"successRate":null,"avgLatencyMs":null,"avgCostUsd":null,"hallucinationRate":null,"retryRate":null,"disputeRate":null,"p50Latency":null,"p95Latency":null,"lastUpdated":null}},"benchmarks":{"evidence":{"source":"no-benchmark-data","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No benchmark suites or observed failure patterns are available."},"suites":[],"failurePatterns":[]},"artifacts":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-09T09:29:37.766Z","emptyReason":null},"readme":"Skill: Baoyu Imagine\n\nOwner: jimliu\n\nSummary: AI image generation with OpenAI GPT Image 2, Azure OpenAI, Google, OpenRouter, DashScope, Z.AI GLM-Image, MiniMax, Jimeng, Seedream and Replicate APIs. Suppo...\n\nTags: latest:1.117.3\n\nVersion history:\n\nv1.117.3 | 2026-05-24T22:13:53.081Z | auto\n\n- Added documentation for reference-image identity preservation, including best practices and anti-patterns.\n- Updated usage examples with new guidance for prompts when preserving real subjects from reference images.\n- Added two new reference documents: `codex-image2-fallback.md` and `codex-oauth-vs-openai-api-key.md`.\n- Clarified recommendations for avoiding drift and over-stylization in identity-preserving generation workflows.\n\nv1.117.2 | 2026-05-18T02:15:08.352Z | user\n\n## 1.117.2 - 2026-05-17\n\n### Documentation\n- `baoyu-cover-image`: ban programmatic text repair on generated bitmaps — disallow ImageMagick / Pillow / Canvas / SVG / HTML overlays to cover, rewrite, or replace title/subtitle text; regenerate from a corrected prompt or switch to a lower-text or no-title variant instead\n- `baoyu-article-illustrator`, `baoyu-comic`, `baoyu-image-cards`, `baoyu-xhs-images`, `baoyu-infographic`, `baoyu-slide-deck`: sync the same text-repair ban with skill-specific text categories (labels/captions, dialogue/sound effects, titles/body/tags, headings/data values, slide titles/bullets)\n\nv1.115.4 | 2026-05-11T23:49:03.373Z | user\n\n### Documentation\n- Image generation backend selection: emphasize Codex `imagegen` as the priority runtime-native tool (invoke via the `Skill` tool with `skill: \"imagegen\"`) and forbid SVG/HTML/canvas substitution when no raster backend can be resolved — fall through to asking the user instead of silently emitting code-based art. Updated in `docs/image-generation-tools.md` and inlined into `baoyu-article-illustrator`, `baoyu-comic`, `baoyu-cover-image`, `baoyu-image-cards`, `baoyu-infographic`, `baoyu-slide-deck`, and `baoyu-xhs-images`.\n\nv1.115.1 | 2026-05-10T03:12:19.033Z | user\n\n## 1.115.1 - 2026-05-10\n\n### Fixes\n- `baoyu-imagine`: change the default MiniMax image API endpoint to `https://api.minimaxi.com` to match the current official image generation documentation, while keeping `https://api.minimax.io` available through `MINIMAX_BASE_URL` overrides.\n- `baoyu-image-gen`: sync the deprecated image-generation entrypoint with the same MiniMax default endpoint and regression coverage.\n\nv1.104.0 | 2026-04-21T20:15:53.791Z | user\n\n## baoyu-imagine 1.104.0\n\n- Add `gpt-image-2` support for OpenAI image generation and edits.\n- Make `gpt-image-2` the default OpenAI model and update Azure deployment guidance.\n- Document official size and quality mapping, custom-size constraints, and 4K usage examples.\n\nv1.103.1 | 2026-04-13T16:15:09.162Z | user\n\n## 1.103.1 - 2026-04-13\n\n### Fixes\n- `baoyu-markdown-to-html`: decode HTML entities and strip tags from article summary\n- `baoyu-post-to-weibo`: decode HTML entities and strip tags from article summary\n\nv1.103.0 | 2026-04-13T01:20:08.024Z | user\n\n## 1.103.0 - 2026-04-12\n\n### Features\n- baoyu-diagram: add multi-diagram mode for article-wide diagram generation\n\n### Fixes\n- baoyu-article-illustrator: prevent color names and hex codes from appearing as visible text in generated images\n- baoyu-cover-image: prevent color names and hex codes from appearing as visible text in generated images\n- baoyu-image-cards: prevent color names from appearing as visible text in generated images\n- baoyu-post-to-wechat: decode HTML entities and strip tags from article summary\n\nv1.0.1 | 2026-03-26T01:11:22.604Z | auto\n\n- Adds legacy compatibility: if `.baoyu-skills/baoyu-image-gen/EXTEND.md` exists and the new path does not, it will be renamed to `.baoyu-skills/baoyu-imagine/EXTEND.md` at runtime.\n- Both old and new EXTEND.md can coexist; if both exist, the new path is always used.\n- No user action required; migration is automatic for existing users upgrading from baoyu-image-gen.\n\nv1.0.0 | 2026-03-25T21:31:56.392Z | auto\n\nbaoyu-imagine 1.0.0\n\n- Initial release for AI image generation using multiple providers (OpenAI, Azure OpenAI, Google, OpenRouter, DashScope, MiniMax, Jimeng, Seedream, Replicate).\n- Supports text-to-image, reference images, aspect ratios, quality presets, and batch generation from prompt files.\n- Preferences and setup guided by EXTEND.md configuration file for key defaults (provider, model, quality, etc).\n- Includes single-image and batch (parallel worker) modes, with flexible CLI options and JSON batch input.\n- Comprehensive environment variable and provider-specific configuration support.\n\nArchive index:\n\nArchive v1.117.3: 38 files, 96871 bytes\n\nFiles: references/codex-image2-fallback.md (1842b), references/codex-oauth-vs-openai-api-key.md (1902b), references/config/first-time-setup.md (13173b), references/config/preferences-schema.md (4224b), references/providers/dashscope.md (4591b), references/providers/minimax.md (1226b), references/providers/openrouter.md (776b), references/providers/replicate.md (1810b), references/providers/zai.md (1095b), references/usage-examples.md (4760b), scripts/build-batch.test.ts (4569b), scripts/build-batch.ts (7454b), scripts/main.test.ts (17852b), scripts/main.ts (43411b), scripts/providers/azure.test.ts (5423b), scripts/providers/azure.ts (5820b), scripts/providers/dashscope.test.ts (11534b), scripts/providers/dashscope.ts (18456b), scripts/providers/google.test.ts (3417b), scripts/providers/google.ts (9873b), scripts/providers/jimeng.test.ts (2737b), scripts/providers/jimeng.ts (12845b), scripts/providers/minimax.test.ts (5254b), scripts/providers/minimax.ts (6331b), scripts/providers/openai.test.ts (5536b), scripts/providers/openai.ts (12671b), scripts/providers/openrouter.test.ts (5450b), scripts/providers/openrouter.ts (9903b), scripts/providers/replicate.test.ts (7190b), scripts/providers/replicate.ts (17816b), scripts/providers/seedream.test.ts (6544b), scripts/providers/seedream.ts (9855b), scripts/providers/zai.test.ts (5274b), scripts/providers/zai.ts (8946b), scripts/types.ts (2123b), skill-card.md (3243b), SKILL.md (18468b), _meta.json (134b)\n\nFile v1.117.3:SKILL.md\n\n---\nname: baoyu-imagine\ndescription: AI image generation with OpenAI GPT Image 2, Azure OpenAI, Google, OpenRouter, DashScope, Z.AI GLM-Image, MiniMax, Jimeng, Seedream and Replicate APIs. Supports text-to-image, reference images, aspect ratios, and batch generation from saved prompt files. Sequential by default; use batch parallel generation when the user already has multiple prompts or wants stable multi-image throughput. Use when user asks to generate, create, or draw images.\nversion: 1.117.3\nmetadata:\n  openclaw:\n    homepage: https://github.com/JimLiu/baoyu-skills#baoyu-imagine\n    requires:\n      anyBins:\n        - bun\n        - npx\n---\n\n# Image Generation (AI SDK)\n\nOfficial API-based image generation. Supports OpenAI GPT Image 2, Azure OpenAI, Google, OpenRouter, DashScope (阿里通义万象), Z.AI GLM-Image, MiniMax, Jimeng (即梦), Seedream (豆包) and Replicate.\n\n## User Input Tools\n\nWhen this skill prompts the user, follow this tool-selection rule (priority order):\n\n1. **Prefer built-in user-input tools** exposed by the current agent runtime — e.g., `AskUserQuestion`, `request_user_input`, `clarify`, `ask_user`, or any equivalent.\n2. **Fallback**: if no such tool exists, emit a numbered plain-text message and ask the user to reply with the chosen number/answer for each question.\n3. **Batching**: if the tool supports multiple questions per call, combine all applicable questions into a single call; if only single-question, ask them one at a time in priority order.\n\nConcrete `AskUserQuestion` references below are examples — substitute the local equivalent in other runtimes.\n\n## Script Directory\n\n`{baseDir}` = this SKILL.md's directory. Main script: `{baseDir}/scripts/main.ts`. Resolve `${BUN_X}`: prefer `bun`; else `npx -y bun`; else suggest `brew install oven-sh/bun/bun`.\n\n## Step 0: Load Preferences ⛔ BLOCKING\n\nThis step MUST complete before any image generation — generation is blocked until EXTEND.md exists.\n\nCheck these paths in order; first hit wins:\n\n| Path | Scope |\n|------|-------|\n| `.baoyu-skills/baoyu-imagine/EXTEND.md` | Project |\n| `${XDG_CONFIG_HOME:-$HOME/.config}/baoyu-skills/baoyu-imagine/EXTEND.md` | XDG |\n| `$HOME/.baoyu-skills/baoyu-imagine/EXTEND.md` | User home |\n\n- **Found** → load, parse, apply. If `default_model.[provider]` is null → ask model only.\n- **Not found** → run first-time setup (`references/config/first-time-setup.md`) using AskUserQuestion to collect provider + model + quality + save location. Save EXTEND.md, then continue. Do not generate images before this completes.\n\nLegacy compatibility: if `.baoyu-skills/baoyu-image-gen/EXTEND.md` exists and the new path doesn't, the runtime renames it to `baoyu-imagine`. If both exist, the runtime leaves them alone and uses the new path.\n\n**EXTEND.md keys**: default provider, default quality, default aspect ratio, default image size, OpenAI image API dialect, default models, batch worker cap, provider-specific batch limits. Schema: `references/config/preferences-schema.md`.\n\n## Usage\n\nMinimum working examples — see `references/usage-examples.md` for the full set including per-provider invocations and batch mode.\n\n### Identity-preserving reference prompts\n\nWhen the user wants a real person/character/object preserved from reference images, do **not** replace the reference with a long generic description. Prefer short, hard identity-preservation language:\n\n- \"Use the person/object in the reference image(s) as the same identity. Do not redesign it or create a similar-looking new subject.\"\n- \"Only change scene, clothing, pose, lighting, rendering style, and composition. Keep the face/proportions/hair/key accessories/overall identity from the references.\"\n- If using multiple references, state that they are the same subject and should jointly define identity.\n\nPitfall: long descriptions like \"young East Asian woman, oval face, clear eyes...\" can cause the model to synthesize a new person matching the description instead of preserving the referenced person.\n\n```bash\n# Basic\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A cat\" --image cat.png\n\n# With aspect ratio and high quality\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A landscape\" --image out.png --ar 16:9 --quality 2k\n\n# Prompt from files\n${BUN_X} {baseDir}/scripts/main.ts --promptfiles system.md content.md --image out.png\n\n# With reference image\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"Make blue\" --image out.png --ref source.png\n\n# Specific provider\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A cat\" --image out.png --provider dashscope --model qwen-image-2.0-pro\n\n# OpenAI GPT Image 2\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A cat\" --image out.png --provider openai --model gpt-image-2\n\n# Batch mode\n${BUN_X} {baseDir}/scripts/main.ts --batchfile batch.json --jobs 4\n```\n\n## Reference-Image Identity Preservation\n\nWhen the user wants a person/object preserved from reference images:\n\n- Prefer a small curated set of existing source references (usually 2–4) over many images; large multi-megabyte refs can destabilize streaming providers.\n- Make the prompt say the references are the same subject and the output must use that identity. Avoid long generic facial-feature descriptions that can cause the model to synthesize a new similar-looking person.\n- Do not use newly generated outputs as references unless the user explicitly asks; generated refs compound drift.\n- If results become too polished or influencer-like, reduce stylized refs and add explicit anti-beautification constraints (no face slimming, eye enlargement, heavy makeup, commercial travel shoot, over-smoothing).\n- If the subject should look younger/older, preserve the face and express age through clothing, posture, scene, and styling; do not ask the model to change facial identity.\n\n## Options\n\n| Option | Description |\n|--------|-------------|\n| `--prompt <text>`, `-p` | Prompt text |\n| `--promptfiles <files...>` | Read prompt from files (concatenated) |\n| `--image <path>` | Output image path (required in single-image mode) |\n| `--batchfile <path>` | JSON batch file for multi-image generation |\n| `--jobs <count>` | Worker count for batch mode (default: auto, max from config, built-in default 10) |\n| `--provider google\\|openai\\|azure\\|openrouter\\|dashscope\\|zai\\|minimax\\|jimeng\\|seedream\\|replicate` | Force provider (default: auto-detect) |\n| `--model <id>`, `-m` | Model ID — see provider references for defaults and allowed values |\n| `--ar <ratio>` | Aspect ratio (`16:9`, `1:1`, `4:3`, …) |\n| `--size <WxH>` | Explicit size (e.g., `1024x1024`; for `gpt-image-2`, width/height must be multiples of 16, max edge 3840px, ratio no wider than 3:1) |\n| `--quality normal\\|2k` | Quality preset (default: `2k`) |\n| `--imageSize 1K\\|2K\\|4K` | Image size for Google/OpenRouter (default: from quality) |\n| `--imageApiDialect openai-native\\|ratio-metadata` | OpenAI-compatible endpoint dialect — use `ratio-metadata` for gateways that expect aspect-ratio `size` plus `metadata.resolution` |\n| `--ref <files...>` | Reference images. Supported by Google multimodal, OpenAI GPT Image edits, Azure OpenAI edits (PNG/JPG only), OpenRouter multimodal models, Replicate supported families, MiniMax subject-reference, Seedream 5.0/4.5/4.0, DashScope `wan2.7-image-pro`/`wan2.7-image`. Not supported by Jimeng, Seedream 3.0, SeedEdit 3.0, or any DashScope model outside the `wan2.7-image*` family |\n| `--n <count>` | Number of images. Replicate requires `--n 1` (single-output save semantics) |\n| `--json` | JSON output |\n\n## Environment Variables\n\n| Variable | Description |\n|----------|-------------|\n| `OPENAI_API_KEY` | OpenAI API key |\n| `AZURE_OPENAI_API_KEY` | Azure OpenAI API key |\n| `OPENROUTER_API_KEY` | OpenRouter API key |\n| `GOOGLE_API_KEY` | Google API key |\n| `DASHSCOPE_API_KEY` | DashScope API key |\n| `ZAI_API_KEY` (alias `BIGMODEL_API_KEY`) | Z.AI API key |\n| `MINIMAX_API_KEY` | MiniMax API key |\n| `REPLICATE_API_TOKEN` | Replicate API token |\n| `JIMENG_ACCESS_KEY_ID`, `JIMENG_SECRET_ACCESS_KEY` | Jimeng (即梦) Volcengine credentials |\n| `ARK_API_KEY` | Seedream (豆包) Volcengine ARK API key |\n| `<PROVIDER>_IMAGE_MODEL` | Per-provider model override (`OPENAI_IMAGE_MODEL`, `GOOGLE_IMAGE_MODEL`, `DASHSCOPE_IMAGE_MODEL`, `ZAI_IMAGE_MODEL`/`BIGMODEL_IMAGE_MODEL`, `MINIMAX_IMAGE_MODEL`, `OPENROUTER_IMAGE_MODEL`, `REPLICATE_IMAGE_MODEL`, `JIMENG_IMAGE_MODEL`, `SEEDREAM_IMAGE_MODEL`) |\n| `AZURE_OPENAI_DEPLOYMENT` (alias `AZURE_OPENAI_IMAGE_MODEL`) | Azure default deployment |\n| `<PROVIDER>_BASE_URL` | Per-provider endpoint override |\n| `AZURE_API_VERSION` | Azure image API version (default `2025-04-01-preview`) |\n| `JIMENG_REGION` | Jimeng region (default `cn-north-1`) |\n| `OPENAI_IMAGE_API_DIALECT` | `openai-native` \\| `ratio-metadata` |\n| `OPENROUTER_HTTP_REFERER`, `OPENROUTER_TITLE` | Optional OpenRouter attribution |\n| `BAOYU_IMAGE_GEN_MAX_WORKERS` | Override batch worker cap |\n| `BAOYU_IMAGE_GEN_<PROVIDER>_CONCURRENCY` | Per-provider concurrency (e.g., `BAOYU_IMAGE_GEN_REPLICATE_CONCURRENCY`) |\n| `BAOYU_IMAGE_GEN_<PROVIDER>_START_INTERVAL_MS` | Per-provider start-gap |\n\n**Load priority**: CLI args > EXTEND.md > env vars > `<cwd>/.baoyu-skills/.env` > `~/.baoyu-skills/.env`\n\n### Codex/ChatGPT OAuth is not an OpenAI API key\n\n`--provider openai --model gpt-image-2` uses the standard OpenAI Images API (`/v1/images/generations` or `/v1/images/edits`) and requires `OPENAI_API_KEY`. A Codex or ChatGPT desktop login is a different entitlement and is not a drop-in replacement for `OPENAI_API_KEY`; do not paste a Codex OAuth token into `OPENAI_API_KEY` or only set `OPENAI_BASE_URL` to a Codex backend.\n\nIf the user wants to use their Codex subscription / GPT Image 2 entitlement without an OpenAI API key, route through a Codex-native backend instead of this skill's `openai` provider:\n\n- In Codex runtime: use the native `imagegen` skill/tool.\n- In non-Codex runtimes with `codex` CLI installed and logged in: use the repo-level `scripts/codex-imagegen.sh` wrapper when the calling skill supports it (for example `baoyu-cover-image`). Resolve it from the plugin/repo root and pass absolute prompt/output/reference paths.\n- In Hermes runtimes with a native `image_generate` tool: use that tool as a fallback, and state whether reference images were passed directly or reconstructed from extracted traits.\n\nDo not modify the existing `openai` provider to silently consume Codex OAuth. If first-class Codex OAuth support is added to `baoyu-imagine`, implement it as a distinct provider (for example `openai-codex`) with its own auth, route, request shape, docs, and tests. See `references/codex-oauth-vs-openai-api-key.md`.\n\n## Model Resolution\n\nPriority (highest → lowest) applies to every provider:\n\n1. CLI flag `--model <id>`\n2. EXTEND.md `default_model.[provider]`\n3. Env var `<PROVIDER>_IMAGE_MODEL`\n4. Built-in default\n\nFor OpenAI, the built-in default is `gpt-image-2`. `gpt-image-1.5`, `gpt-image-1`, and GPT Image snapshots remain selectable with `--model` or `OPENAI_IMAGE_MODEL`.\n\nFor Azure, `--model` / `default_model.azure` is the Azure deployment name. `AZURE_OPENAI_DEPLOYMENT` is the preferred env var; `AZURE_OPENAI_IMAGE_MODEL` is kept as a backward-compatible alias. If your Azure deployment is named after the underlying model, use `gpt-image-2`; otherwise use the exact custom deployment name.\n\nEXTEND.md overrides env vars: if EXTEND.md sets `default_model.google: \"gemini-3-pro-image-preview\"` and the env var sets `GOOGLE_IMAGE_MODEL=gemini-3.1-flash-image-preview`, EXTEND.md wins.\n\n**Display model info before each generation**:\n\n- `Using [provider] / [model]`\n- `Switch model: --model <id> | EXTEND.md default_model.[provider] | env <PROVIDER>_IMAGE_MODEL`\n\n## OpenAI-Compatible Gateway Dialects\n\n`provider=openai` means the auth and routing entrypoint is OpenAI-compatible. It does **not** guarantee the upstream image API uses OpenAI native semantics. When a gateway expects a different wire format, set `default_image_api_dialect` in EXTEND.md, `OPENAI_IMAGE_API_DIALECT`, or `--imageApiDialect`:\n\n- `openai-native`: pixel `size` (`1536x1024`) and native OpenAI quality fields\n- `ratio-metadata`: aspect-ratio `size` (`16:9`) plus `metadata.resolution` (`1K|2K|4K`) and `metadata.orientation`\n\nUse `openai-native` for the OpenAI native API or strict clones; try `ratio-metadata` for compatibility gateways in front of Gemini or similar models. Current limitation: `ratio-metadata` applies only to text-to-image; reference-image edits still need `openai-native` or a provider with first-class edit support.\n\n## Provider-Specific Guides\n\nEach provider has its own quirks (model families, size rules, ref support, limits). Read these when the user picks that provider or asks for non-default behavior:\n\n| Provider | Reference |\n|----------|-----------|\n| DashScope (Qwen-Image families, custom sizes) | `references/providers/dashscope.md` |\n| Z.AI (GLM-Image, cogview-4) | `references/providers/zai.md` |\n| MiniMax (image-01, subject-reference) | `references/providers/minimax.md` |\n| OpenRouter (multimodal models, `/chat/completions` flow) | `references/providers/openrouter.md` |\n| Replicate (nano-banana, Seedream, Wan) | `references/providers/replicate.md` |\n\n## Provider Selection\n\n1. `--ref` provided + no `--provider` → auto-select Google → OpenAI → Azure → OpenRouter → Replicate → Seedream → MiniMax (MiniMax's subject reference is more specialized toward character/portrait consistency)\n2. `--provider` specified → use it (if `--ref`, must be google/openai/azure/openrouter/replicate/seedream/minimax)\n3. Only one API key present → use that provider\n4. Multiple keys → default priority: Google → OpenAI → Azure → OpenRouter → DashScope → Z.AI → MiniMax → Replicate → Jimeng → Seedream\n\n## Quality Presets\n\n| Preset | Google imageSize | OpenAI size | OpenRouter size | Replicate resolution | Use case |\n|--------|------------------|-------------|-----------------|----------------------|----------|\n| `normal` | 1K | 1024px target | 1K | 1K | Quick previews |\n| `2k` (default) | 2K | 2048px target | 2K | 2K | Covers, illustrations, infographics |\n\nGoogle/OpenRouter `imageSize` can be overridden with `--imageSize 1K|2K|4K`.\n\nFor OpenAI native `gpt-image-2`, `normal` maps to `quality=medium` and a low-latency valid size near the requested aspect ratio; `2k` maps to `quality=high` and 2048px-class sizes such as `2048x2048`, `2048x1152`, or `1152x2048`. Use explicit `--size` for valid custom or 4K outputs, e.g. `3840x2160`.\n\n## Aspect Ratios\n\nSupported: `1:1`, `16:9`, `9:16`, `4:3`, `3:4`, `2.35:1`.\n\n- Google multimodal: `imageConfig.aspectRatio`\n- OpenAI: `gpt-image-2` uses the closest valid custom size for the requested ratio; older GPT Image and DALL·E models use their closest supported fixed size\n- OpenRouter: `imageGenerationOptions.aspect_ratio`; if only `--size <WxH>` is given, the ratio is inferred\n- Replicate: behavior is model-specific — `google/nano-banana*` uses `aspect_ratio`, `bytedance/seedream-*` uses documented Replicate ratios, Wan 2.7 maps `--ar` to a concrete `size`\n- MiniMax: official `aspect_ratio` values; if `--size <WxH>` is given without `--ar`, sends `width`/`height` for `image-01`\n\n## Generation Mode\n\n**Default**: sequential. **Batch parallel**: enabled automatically when `--batchfile` contains 2+ pending tasks.\n\n| Situation | Prefer | Why |\n|-----------|--------|-----|\n| One image, or 1-2 simple images | Sequential | Lower coordination overhead, easier debugging |\n| Multiple images with saved prompt files | Batch (`--batchfile`) | Reuses finalized prompts, applies shared throttling/retries, predictable throughput |\n| Each image still needs its own reasoning / prompt writing / style exploration | Subagents | Work is still exploratory, each needs independent analysis |\n| Input is `outline.md` + `prompts/` (e.g. from `baoyu-article-illustrator`) | Batch — use `scripts/build-batch.ts` to assemble the payload | The outline + prompt files already contain everything needed |\n\nRule of thumb: once prompt files are saved and the task is \"generate all of these\", prefer batch over subagents. Use subagents only when generation is coupled with per-image thinking or divergent creative exploration.\n\n**Parallel behavior**:\n\n- Default worker count is automatic, capped by config, built-in default 10\n- Provider-specific throttling applies only in batch mode; defaults are tuned for throughput while avoiding RPM bursts\n- Override with `--jobs <count>`\n- Each image retries up to 3 attempts\n- Final output includes success count, failure count, and per-image failure reasons\n\n## Error Handling\n\n- Missing API key → error with setup instructions\n- Generation failure → auto-retry up to 3 attempts per image\n- Invalid aspect ratio → warning, proceed with default\n- Reference images with unsupported provider/model → error with fix hint\n\n### Codex image2 fallback\n\nIf `--provider openai --model gpt-image-2` fails because `OPENAI_API_KEY` is missing but the current runtime has a native image-generation backend or the repo-level `codex-imagegen` wrapper is available, use that path rather than leaving the user waiting. Be explicit about whether the fallback is true reference-image generation or only a text-prompt reconstruction from extracted visual traits. See `references/codex-image2-fallback.md`.\n\n## References\n\n| File | Content |\n|------|---------|\n| `references/usage-examples.md` | Extended CLI examples across providers and batch mode |\n| `references/codex-oauth-vs-openai-api-key.md` | Why Codex/ChatGPT OAuth image2 entitlement is not usable through baoyu-imagine's standard OpenAI API-key provider |\n| `references/codex-image2-fallback.md` | Practical fallback behavior when OpenAI API credentials are absent but Codex/native image generation is available |\n| `references/providers/dashscope.md` | DashScope families, sizes, limits |\n| `references/providers/zai.md` | Z.AI GLM-image / cogview-4 |\n| `references/providers/minimax.md` | MiniMax image-01 + subject reference |\n| `references/providers/openrouter.md` | OpenRouter multimodal flow |\n| `references/providers/replicate.md` | Replicate supported families + guardrails |\n| `references/config/preferences-schema.md` | EXTEND.md schema |\n| `references/config/first-time-setup.md` | First-time setup flow |\n\n## Extension Support\n\nCustom configurations via EXTEND.md. See Step 0 for paths and schema.\n\nFile v1.117.3:_meta.json\n\n{\n  \"ownerId\": \"kn7csrrndw79hpke5d0gsnx93d82k67r\",\n  \"slug\": \"baoyu-imagine\",\n  \"version\": \"1.117.3\",\n  \"publishedAt\": 1779660833081\n}\n\nFile v1.117.3:references/codex-image2-fallback.md\n\n---\nname: codex-image2-fallback\ndescription: Fallback behavior when baoyu-imagine lacks OpenAI API credentials but Codex/native image generation is available\n---\n\n# Codex Image2 Fallback\n\nWhen using `baoyu-imagine` with `--provider openai --model gpt-image-2`, the CLI can fail with:\n\n```text\nOPENAI_API_KEY is required. Codex/ChatGPT desktop login does not automatically grant OpenAI Images API access to this script.\n```\n\nThis is expected. The `openai` provider uses the public OpenAI Images API and needs `OPENAI_API_KEY`. Codex / ChatGPT image2 entitlement is a separate runtime-native path.\n\n## Practical fallback pattern\n\n1. Try `baoyu-imagine` when provider credentials are available.\n2. If it fails only because `OPENAI_API_KEY` is missing, do not leave the user waiting.\n3. Prefer a Codex/native raster backend in this order:\n   - Codex runtime native `imagegen` skill/tool, if available.\n   - Repo-level `scripts/codex-imagegen.sh`, if `codex` CLI is installed/logged in and the calling skill supports the wrapper.\n   - Hermes native `image_generate`, if available.\n4. Be transparent about reference-image behavior:\n   - If the fallback backend accepts references, pass the reference images.\n   - If it does not, derive a concise identity-preserving prompt from the references and state that it is a text-description fallback, not strict reference-image editing.\n5. Return the generated media path or structured backend error promptly.\n\n## User-facing wording\n\nUse concise wording such as:\n\n> The OpenAI API path needs `OPENAI_API_KEY`; Codex login is a separate image2 backend. I used the available Codex/native image backend instead. Reference images were [passed directly / reconstructed from visual traits].\n\nAvoid implying that `baoyu-imagine --provider openai` can use Codex OAuth without a dedicated provider implementation.\n\nFile v1.117.3:references/codex-oauth-vs-openai-api-key.md\n\n# Codex OAuth vs OpenAI API key for baoyu-imagine\n\n`baoyu-imagine --provider openai` uses the standard OpenAI Images API and requires `OPENAI_API_KEY`. It calls OpenAI-compatible image endpoints such as `/images/generations` and `/images/edits`.\n\nCodex / ChatGPT login is different. Codex image generation is driven by Codex OAuth and the Codex runtime's `image_gen` capability, not by the public OpenAI Images API key path. A Codex OAuth token is not a drop-in replacement for `OPENAI_API_KEY`, and setting `OPENAI_BASE_URL` to a Codex backend will not make baoyu-imagine's existing `openai` provider work because the auth, route, and payload shape differ.\n\n## What to use instead\n\n- If running inside Codex and the native `imagegen` skill/tool is available, use it directly.\n- If running outside Codex but the `codex` CLI is installed and logged in, use the repo-level `scripts/codex-imagegen.sh` wrapper when the calling skill supports it. The wrapper invokes `codex exec` and the Codex `image_gen` tool; no `OPENAI_API_KEY` is required.\n- If running inside Hermes and a native `image_generate` tool is available, use that as a runtime-native fallback. Be explicit about whether reference images are passed directly or only reconstructed from extracted traits.\n- If the user wants `baoyu-imagine` itself to support Codex OAuth, add a distinct provider such as `openai-codex` rather than modifying the existing `openai` provider.\n\n## Reference-image prompting note\n\nWhen using actual reference images for identity preservation, avoid long generic descriptions of the subject. Long descriptions can cause the model to synthesize a new similar-looking person/object. Prefer direct wording:\n\n> Use the person/object in the reference image(s) as the same identity. Do not redesign it or create a similar-looking new subject. Only change scene, clothing, pose, lighting, rendering style, and composition.\n\nFile v1.117.3:references/config/first-time-setup.md\n\n---\nname: first-time-setup\ndescription: First-time setup and default model selection flow for baoyu-imagine\n---\n\n# First-Time Setup\n\n## Overview\n\nTriggered when:\n1. No EXTEND.md found → full setup (provider + model + preferences)\n2. EXTEND.md found but `default_model.[provider]` is null → model selection only\n\n## Setup Flow\n\n```\nNo EXTEND.md found          EXTEND.md found, model null\n        │                            │\n        ▼                            ▼\n┌─────────────────────┐    ┌──────────────────────┐\n│ AskUserQuestion     │    │ AskUserQuestion      │\n│ (full setup)        │    │ (model only)         │\n└─────────────────────┘    └──────────────────────┘\n        │                            │\n        ▼                            ▼\n┌─────────────────────┐    ┌──────────────────────┐\n│ Create EXTEND.md    │    │ Update EXTEND.md     │\n└─────────────────────┘    └──────────────────────┘\n        │                            │\n        ▼                            ▼\n    Continue                     Continue\n```\n\n## Flow 1: No EXTEND.md (Full Setup)\n\n**Language**: Use user's input language or saved language preference.\n\nUse AskUserQuestion with ALL questions in ONE call:\n\n### Question 1: Default Provider\n\n```yaml\nheader: \"Provider\"\nquestion: \"Default image generation provider?\"\noptions:\n  - label: \"Google (Recommended)\"\n    description: \"Gemini multimodal - high quality, reference images, flexible sizes\"\n  - label: \"OpenAI\"\n    description: \"GPT Image 2 - latest OpenAI image model, reference-image workflows\"\n  - label: \"Azure OpenAI\"\n    description: \"Azure-hosted GPT Image deployments with resource-specific routing\"\n  - label: \"OpenRouter\"\n    description: \"Router for Gemini/FLUX/OpenAI-compatible image models\"\n  - label: \"DashScope\"\n    description: \"Alibaba Cloud - Qwen-Image, strong Chinese/English text rendering\"\n  - label: \"Z.AI\"\n    description: \"GLM-image, strong poster and text-heavy image generation\"\n  - label: \"MiniMax\"\n    description: \"MiniMax image generation with subject-reference character workflows\"\n  - label: \"Replicate\"\n    description: \"Curated Replicate image families - nano-banana-2, Seedream, and Wan image models\"\n```\n\n### Question 2: Default Google Model\n\nOnly show if user selected Google or auto-detect (no explicit provider).\n\n```yaml\nheader: \"Google Model\"\nquestion: \"Default Google image generation model?\"\noptions:\n  - label: \"gemini-3-pro-image-preview (Recommended)\"\n    description: \"Highest quality, best for production use\"\n  - label: \"gemini-3.1-flash-image-preview\"\n    description: \"Fast generation, good quality, lower cost\"\n  - label: \"gemini-3-flash-preview\"\n    description: \"Fast generation, balanced quality and speed\"\n```\n\n### Question 2b: Default OpenRouter Model\n\nOnly show if user selected OpenRouter.\n\n```yaml\nheader: \"OpenRouter Model\"\nquestion: \"Default OpenRouter image generation model?\"\noptions:\n  - label: \"google/gemini-3.1-flash-image-preview (Recommended)\"\n    description: \"Best general-purpose OpenRouter image model with reference-image workflows\"\n  - label: \"google/gemini-2.5-flash-image-preview\"\n    description: \"Fast Gemini preview model on OpenRouter\"\n  - label: \"black-forest-labs/flux.2-pro\"\n    description: \"Strong text-to-image quality through OpenRouter\"\n```\n\n### Question 2c: Default Azure Deployment\n\nOnly show if user selected Azure OpenAI.\n\n```yaml\nheader: \"Azure Deploy\"\nquestion: \"Default Azure image deployment name?\"\noptions:\n  - label: \"gpt-image-2 (Recommended)\"\n    description: \"Use if your Azure deployment uses the GPT Image 2 model name\"\n  - label: \"gpt-image-1.5\"\n    description: \"Previous GPT Image deployment name\"\n  - label: \"gpt-image-1\"\n    description: \"Earlier GPT Image deployment name\"\n```\n\n### Question 2d: Default MiniMax Model\n\nOnly show if user selected MiniMax.\n\n```yaml\nheader: \"MiniMax Model\"\nquestion: \"Default MiniMax image generation model?\"\noptions:\n  - label: \"image-01 (Recommended)\"\n    description: \"Best default, supports aspect ratios and custom width/height\"\n  - label: \"image-01-live\"\n    description: \"Faster variant, use aspect ratio instead of custom size\"\n```\n\n### Question 2e: Default Z.AI Model\n\nOnly show if user selected Z.AI.\n\n```yaml\nheader: \"Z.AI Model\"\nquestion: \"Default Z.AI image generation model?\"\noptions:\n  - label: \"glm-image (Recommended)\"\n    description: \"Best default for posters, diagrams, and text-heavy images\"\n  - label: \"cogview-4-250304\"\n    description: \"Legacy Z.AI image model on the same endpoint\"\n```\n\n### Question 3: Default Quality\n\n```yaml\nheader: \"Quality\"\nquestion: \"Default image quality?\"\noptions:\n  - label: \"2k (Recommended)\"\n    description: \"2048px - covers, illustrations, infographics\"\n  - label: \"normal\"\n    description: \"1024px - quick previews, drafts\"\n```\n\n### Question 4: Save Location\n\n```yaml\nheader: \"Save\"\nquestion: \"Where to save preferences?\"\noptions:\n  - label: \"Project (Recommended)\"\n    description: \".baoyu-skills/ (this project only)\"\n  - label: \"User\"\n    description: \"~/.baoyu-skills/ (all projects)\"\n```\n\n### Save Locations\n\n| Choice | Path | Scope |\n|--------|------|-------|\n| Project | `.baoyu-skills/baoyu-imagine/EXTEND.md` | Current project |\n| User | `$HOME/.baoyu-skills/baoyu-imagine/EXTEND.md` | All projects |\n\n### EXTEND.md Template\n\n```yaml\n---\nversion: 1\ndefault_provider: [selected provider or null]\ndefault_quality: [selected quality]\ndefault_aspect_ratio: null\ndefault_image_size: null\ndefault_image_api_dialect: null\ndefault_model:\n  google: [selected google model or null]\n  openai: null\n  azure: [selected azure deployment or null]\n  openrouter: [selected openrouter model or null]\n  dashscope: null\n  zai: [selected Z.AI model or null]\n  minimax: [selected minimax model or null]\n  replicate: null\n---\n```\n\nIf the user selects `OpenAI` but says their endpoint is only OpenAI-compatible and fronts another image model family, save `default_image_api_dialect: ratio-metadata` when they explicitly confirm the gateway expects aspect-ratio `size` plus metadata-based resolution. Otherwise leave it `null` / `openai-native`.\n\n## Flow 2: EXTEND.md Exists, Model Null\n\nWhen EXTEND.md exists but `default_model.[current_provider]` is null, ask ONLY the model question for the current provider.\n\n### Google Model Selection\n\n```yaml\nheader: \"Google Model\"\nquestion: \"Choose a default Google image generation model?\"\noptions:\n  - label: \"gemini-3-pro-image-preview (Recommended)\"\n    description: \"Highest quality, best for production use\"\n  - label: \"gemini-3.1-flash-image-preview\"\n    description: \"Fast generation, good quality, lower cost\"\n  - label: \"gemini-3-flash-preview\"\n    description: \"Fast generation, balanced quality and speed\"\n```\n\n### OpenAI Model Selection\n\n```yaml\nheader: \"OpenAI Model\"\nquestion: \"Choose a default OpenAI image generation model?\"\noptions:\n  - label: \"gpt-image-2 (Recommended)\"\n    description: \"Latest GPT Image model, flexible sizes up to 4K, high-fidelity image inputs\"\n  - label: \"gpt-image-1.5\"\n    description: \"Previous GPT Image model\"\n  - label: \"gpt-image-1\"\n    description: \"Earlier GPT Image model\"\n```\n\n### Azure Deployment Selection\n\n```yaml\nheader: \"Azure Deploy\"\nquestion: \"Choose a default Azure image deployment name?\"\noptions:\n  - label: \"gpt-image-2 (Recommended)\"\n    description: \"Use when your Azure deployment name matches the GPT Image 2 model\"\n  - label: \"gpt-image-1.5\"\n    description: \"Use when your Azure deployment name matches the GPT Image 1.5 model\"\n  - label: \"gpt-image-1\"\n    description: \"Use when your Azure deployment name matches GPT-image-1\"\n```\n\nNotes for Azure setup:\n\n- In `baoyu-imagine`, Azure `--model` / `default_model.azure` should be the Azure deployment name, not just the underlying model family.\n- If the deployment name is custom, save that exact deployment name in `default_model.azure`.\n\n### OpenRouter Model Selection\n\n```yaml\nheader: \"OpenRouter Model\"\nquestion: \"Choose a default OpenRouter image generation model?\"\noptions:\n  - label: \"google/gemini-3.1-flash-image-preview (Recommended)\"\n    description: \"Recommended for image output and reference-image edits\"\n  - label: \"google/gemini-2.5-flash-image-preview\"\n    description: \"Fast preview-oriented image generation\"\n  - label: \"black-forest-labs/flux.2-pro\"\n    description: \"High-quality text-to-image through OpenRouter\"\n```\n\n### DashScope Model Selection\n\n```yaml\nheader: \"DashScope Model\"\nquestion: \"Choose a default DashScope image generation model?\"\noptions:\n  - label: \"qwen-image-2.0-pro (Recommended)\"\n    description: \"Best DashScope model for text rendering and custom sizes\"\n  - label: \"qwen-image-2.0\"\n    description: \"Faster 2.0 variant with flexible output size\"\n  - label: \"qwen-image-max\"\n    description: \"Legacy Qwen model with five fixed output sizes\"\n  - label: \"qwen-image-plus\"\n    description: \"Legacy Qwen model, same current capability as qwen-image\"\n  - label: \"wan2.7-image-pro\"\n    description: \"Wan 2.7 Pro — supports up to 4K text-to-image and reference-image editing\"\n  - label: \"wan2.7-image\"\n    description: \"Wan 2.7 base — faster generation, up to 2K, supports reference-image editing\"\n  - label: \"z-image-turbo\"\n    description: \"Legacy DashScope model for compatibility\"\n  - label: \"z-image-ultra\"\n    description: \"Legacy DashScope model, higher quality but slower\"\n```\n\nNotes for DashScope setup:\n\n- Prefer `qwen-image-2.0-pro` when the user needs custom `--size`, uncommon ratios like `21:9`, or strong Chinese/English text rendering.\n- `qwen-image-max` / `qwen-image-plus` / `qwen-image` only support five fixed sizes: `1664*928`, `1472*1104`, `1328*1328`, `1104*1472`, `928*1664`.\n- `wan2.7-image-pro` and `wan2.7-image` are the only DashScope models that accept `--ref`. Pick one of these when the user wants reference-image editing or multi-image fusion via DashScope.\n- In `baoyu-imagine`, `quality` is a compatibility preset. It is not a native DashScope parameter.\n\n### Z.AI Model Selection\n\n```yaml\nheader: \"Z.AI Model\"\nquestion: \"Choose a default Z.AI image generation model?\"\noptions:\n  - label: \"glm-image (Recommended)\"\n    description: \"Current flagship image model with better text rendering and poster layouts\"\n  - label: \"cogview-4-250304\"\n    description: \"Legacy model on the sync image endpoint\"\n```\n\nNotes for Z.AI setup:\n\n- Prefer `glm-image` for posters, diagrams, and Chinese/English text-heavy layouts.\n- In `baoyu-imagine`, Z.AI currently exposes text-to-image only; reference images are not wired for this provider.\n- The sync Z.AI image API returns a downloadable image URL, which the runtime saves locally after download.\n\n### Replicate Model Selection\n\n```yaml\nheader: \"Replicate Model\"\nquestion: \"Choose a default Replicate image generation model?\"\noptions:\n  - label: \"google/nano-banana-2 (Recommended)\"\n    description: \"Current default for general Replicate image generation in baoyu-imagine\"\n  - label: \"bytedance/seedream-4.5\"\n    description: \"Replicate Seedream 4.5 with validated local size/ref guardrails\"\n  - label: \"bytedance/seedream-5-lite\"\n    description: \"Replicate Seedream 5 Lite with validated local size/ref guardrails\"\n  - label: \"wan-video/wan-2.7-image-pro\"\n    description: \"Replicate Wan 2.7 Image Pro with 4K text-to-image support\"\n```\n\n### MiniMax Model Selection\n\n```yaml\nheader: \"MiniMax Model\"\nquestion: \"Choose a default MiniMax image generation model?\"\noptions:\n  - label: \"image-01 (Recommended)\"\n    description: \"Best general-purpose MiniMax image model with custom width/height support\"\n  - label: \"image-01-live\"\n    description: \"Lower-latency MiniMax image model using aspect ratios\"\n```\n\nNotes for MiniMax setup:\n\n- `image-01` is the safest default. It supports official `aspect_ratio` values and documented custom `width` / `height` output sizes.\n- `image-01-live` is useful when the user prefers faster generation and can work with aspect-ratio-based sizing.\n- MiniMax subject reference currently uses `subject_reference[].type = character`; docs recommend front-facing portrait references in JPG/JPEG/PNG under 10MB.\n\n### Update EXTEND.md\n\nAfter user selects a model:\n\n1. Read existing EXTEND.md\n2. If `default_model:` section exists → update the provider-specific key\n3. If `default_model:` section missing → add the full section:\n\n```yaml\ndefault_model:\n  google: [value or null]\n  openai: [value or null]\n  azure: [value or null]\n  openrouter: [value or null]\n  dashscope: [value or null]\n  zai: [value or null]\n  minimax: [value or null]\n  replicate: [value or null]\n```\n\nOnly set the selected provider's model; leave others as their current value or null.\n\n## After Setup\n\n1. Create directory if needed\n2. Write/update EXTEND.md with frontmatter\n3. Confirm: \"Preferences saved to [path]\"\n4. Continue with image generation\n\nFile v1.117.3:references/config/preferences-schema.md\n\n---\nname: preferences-schema\ndescription: EXTEND.md YAML schema for baoyu-imagine user preferences\n---\n\n# Preferences Schema\n\n## Full Schema\n\n```yaml\n---\nversion: 1\n\ndefault_provider: null      # google|openai|azure|openrouter|dashscope|zai|minimax|replicate|null (null = auto-detect)\n\ndefault_quality: null       # normal|2k|null (null = use default: 2k)\n\ndefault_aspect_ratio: null  # \"16:9\"|\"1:1\"|\"4:3\"|\"3:4\"|\"2.35:1\"|null\n\ndefault_image_size: null    # 1K|2K|4K|null (Google/OpenRouter, overrides quality)\n\ndefault_image_api_dialect: null  # openai-native|ratio-metadata|null (OpenAI-compatible gateways; null = use env/default)\n\ndefault_model:\n  google: null              # e.g., \"gemini-3-pro-image-preview\", \"gemini-3.1-flash-image-preview\"\n  openai: null              # e.g., \"gpt-image-2\", \"gpt-image-1.5\", \"gpt-image-1\"\n  azure: null               # Azure deployment name, e.g., \"gpt-image-2\" or \"image-prod\"\n  openrouter: null          # e.g., \"google/gemini-3.1-flash-image-preview\"\n  dashscope: null           # e.g., \"qwen-image-2.0-pro\"\n  zai: null                 # e.g., \"glm-image\"\n  minimax: null             # e.g., \"image-01\"\n  replicate: null           # e.g., \"google/nano-banana-2\"\n\nbatch:\n  max_workers: 10\n  provider_limits:\n    replicate:\n      concurrency: 5\n      start_interval_ms: 700\n    google:\n      concurrency: 3\n      start_interval_ms: 1100\n    openai:\n      concurrency: 3\n      start_interval_ms: 1100\n    azure:\n      concurrency: 3\n      start_interval_ms: 1100\n    openrouter:\n      concurrency: 3\n      start_interval_ms: 1100\n    dashscope:\n      concurrency: 3\n      start_interval_ms: 1100\n    zai:\n      concurrency: 3\n      start_interval_ms: 1100\n    minimax:\n      concurrency: 3\n      start_interval_ms: 1100\n---\n```\n\n## Field Reference\n\n| Field | Type | Default | Description |\n|-------|------|---------|-------------|\n| `version` | int | 1 | Schema version |\n| `default_provider` | string\\|null | null | Default provider (null = auto-detect) |\n| `default_quality` | string\\|null | null | Default quality (null = 2k) |\n| `default_aspect_ratio` | string\\|null | null | Default aspect ratio |\n| `default_image_size` | string\\|null | null | Google/OpenRouter image size (overrides quality) |\n| `default_image_api_dialect` | string\\|null | null | OpenAI-compatible image dialect (`openai-native` or `ratio-metadata`) |\n| `default_model.google` | string\\|null | null | Google default model |\n| `default_model.openai` | string\\|null | null | OpenAI default model |\n| `default_model.azure` | string\\|null | null | Azure default deployment name |\n| `default_model.openrouter` | string\\|null | null | OpenRouter default model |\n| `default_model.dashscope` | string\\|null | null | DashScope default model |\n| `default_model.zai` | string\\|null | null | Z.AI default model |\n| `default_model.minimax` | string\\|null | null | MiniMax default model |\n| `default_model.replicate` | string\\|null | null | Replicate default model |\n| `batch.max_workers` | int\\|null | 10 | Batch worker cap |\n| `batch.provider_limits.<provider>.concurrency` | int\\|null | provider default | Max simultaneous requests per provider |\n| `batch.provider_limits.<provider>.start_interval_ms` | int\\|null | provider default | Minimum gap between request starts per provider |\n\n## Examples\n\n**Minimal**:\n```yaml\n---\nversion: 1\ndefault_provider: google\ndefault_quality: 2k\ndefault_image_api_dialect: null\n---\n```\n\n**Full**:\n```yaml\n---\nversion: 1\ndefault_provider: google\ndefault_quality: 2k\ndefault_aspect_ratio: \"16:9\"\ndefault_image_size: 2K\ndefault_image_api_dialect: null\ndefault_model:\n  google: \"gemini-3-pro-image-preview\"\n  openai: \"gpt-image-2\"\n  azure: \"gpt-image-2\"\n  openrouter: \"google/gemini-3.1-flash-image-preview\"\n  dashscope: \"qwen-image-2.0-pro\"\n  zai: \"glm-image\"\n  minimax: \"image-01\"\n  replicate: \"google/nano-banana-2\"\nbatch:\n  max_workers: 10\n  provider_limits:\n    replicate:\n      concurrency: 5\n      start_interval_ms: 700\n    azure:\n      concurrency: 3\n      start_interval_ms: 1100\n    zai:\n      concurrency: 3\n      start_interval_ms: 1100\n    openrouter:\n      concurrency: 3\n      start_interval_ms: 1100\n    minimax:\n      concurrency: 3\n      start_interval_ms: 1100\n---\n```\n\nFile v1.117.3:references/providers/dashscope.md\n\n# DashScope (阿里通义万象)\n\nRead when the user picks `--provider dashscope`, sets `default_model.dashscope`, or asks for Qwen-Image behavior. The SKILL.md only names the default — this file covers model families, sizing rules, and limits.\n\n## Model Families\n\n**`qwen-image-2.0*`** — recommended modern family. Members: `qwen-image-2.0-pro`, `qwen-image-2.0-pro-2026-03-03`, `qwen-image-2.0`, `qwen-image-2.0-2026-03-03`.\n\n- Free-form `size` in `宽*高` format\n- Total pixels must be between `512*512` and `2048*2048`\n- Default ≈ `1024*1024`\n- Best choice for custom ratios (e.g. `21:9`) and text-heavy Chinese/English layouts\n\n**Fixed-size family** — `qwen-image-max`, `qwen-image-max-2025-12-30`, `qwen-image-plus`, `qwen-image-plus-2026-01-09`, `qwen-image`.\n\n- Only five sizes allowed: `1664*928`, `1472*1104`, `1328*1328`, `1104*1472`, `928*1664`\n- Default is `1664*928`\n- `qwen-image` currently has the same capability as `qwen-image-plus`\n\n**`wan2.7-image*`** — multimodal Wan 2.7 family. Members: `wan2.7-image-pro`, `wan2.7-image`.\n\n- Free-form `size` in `宽*高` format, plus aspect-ratio inference\n- `wan2.7-image-pro` text-to-image (no `--ref`): total pixels in `[768*768, 4096*4096]`, ratio in `[1:8, 8:1]`\n- `wan2.7-image-pro` with reference images and `wan2.7-image` (all scenarios): total pixels in `[768*768, 2048*2048]`, ratio in `[1:8, 8:1]`\n- Default: `1024*1024` (`--quality normal`) or `2048*2048` (`--quality 2k`); 4K requires explicit `--size`\n- Supports up to 9 reference images in `--ref` (image editing / multi-image fusion)\n- Reference images are sent inline as base64 (or passed through if the path is an `http(s)://` URL)\n- API does NOT use `prompt_extend`; the skill omits it for this family\n- The Wan 2.7 API defaults `n` to **4** in non-collage mode and bills per generated image. baoyu-imagine forces `n: 1` and rejects `--n > 1` to avoid silently paying for and discarding extra images.\n\n**Legacy** — `z-image-turbo`, `z-image-ultra`, `wanx-v1`. Only use when the user explicitly asks for legacy behavior.\n\n## Size Resolution\n\n- `--size` wins over `--ar`\n- For `qwen-image-2.0*`: prefer explicit `--size`; otherwise infer from `--ar` using the recommended table below\n- For `qwen-image-max/plus/image`: only use the five fixed sizes; if the requested ratio doesn't fit, switch to `qwen-image-2.0-pro`\n- For `wan2.7-image*`: explicit `--size` is validated against the per-mode pixel/ratio limits; otherwise the size is derived from `--ar` and `--quality` (`normal` ≈ 1K, `2k` ≈ 2K). To request 4K with `wan2.7-image-pro` text-to-image, pass `--size` explicitly (e.g. `4096*4096`, `3840*2160`)\n- `--quality` is a baoyu-imagine preset, not an official DashScope field. The mapping of `normal`/`2k` onto the `qwen-image-2.0*` and `wan2.7-image*` tables is an implementation choice, not an API guarantee\n\n### Recommended `qwen-image-2.0*` sizes\n\n| Ratio | `normal` | `2k` |\n|-------|----------|------|\n| `1:1` | `1024*1024` | `1536*1536` |\n| `2:3` | `768*1152` | `1024*1536` |\n| `3:2` | `1152*768` | `1536*1024` |\n| `3:4` | `960*1280` | `1080*1440` |\n| `4:3` | `1280*960` | `1440*1080` |\n| `9:16` | `720*1280` | `1080*1920` |\n| `16:9` | `1280*720` | `1920*1080` |\n| `21:9` | `1344*576` | `2048*872` |\n\n## Reference Images\n\n- Only `wan2.7-image-pro` and `wan2.7-image` accept `--ref`. Other DashScope models (qwen-image-2.0*, qwen-image-max/plus/image, legacy) reject `--ref` and the user is steered to a different provider/model.\n- Up to 9 reference images per request. Local files are inlined as base64 data URLs; `http(s)://` URLs are forwarded as-is.\n- Supplying any `--ref` automatically clamps the wan2.7-image-pro pixel ceiling from 4K to 2K (the API only supports 4K for pure text-to-image with no image input).\n\n## Not Exposed\n\nDashScope APIs also support `negative_prompt`, `prompt_extend`, `watermark`, `thinking_mode`, `seed`, `bbox_list`, `enable_sequential`, and `color_palette`. `baoyu-imagine` does not expose them as CLI flags today; the wan2.7 family relies on the API defaults (e.g. `thinking_mode=true`). The skill always sends `n=1` for wan2.7 — if you want grid/collage mode you currently need to call the API directly.\n\n## Official References\n\n- [Qwen-Image API](https://help.aliyun.com/zh/model-studio/qwen-image-api)\n- [Text-to-image guide](https://help.aliyun.com/zh/model-studio/text-to-image)\n- [Qwen-Image Edit API](https://help.aliyun.com/zh/model-studio/qwen-image-edit-api)\n- [Wan 2.7 image generation & editing API](https://help.aliyun.com/zh/model-studio/wan-image-generation-and-editing-api-reference)\n\nFile v1.117.3:references/providers/minimax.md\n\n# MiniMax\n\nRead when the user picks `--provider minimax` or sets `default_model.minimax`. Default model is `image-01`.\n\n## Models\n\n**`image-01`** (recommended default)\n\n- Supports text-to-image and subject-reference image generation\n- Supports official `aspect_ratio` values: `1:1`, `16:9`, `4:3`, `3:2`, `2:3`, `3:4`, `9:16`, `21:9`\n- Supports documented custom `width` / `height` via `--size <WxH>`\n- Both width and height must be in `[512, 2048]` and divisible by `8`\n\n**`image-01-live`** — lower-latency variant\n\n- Use `--ar` for sizing; MiniMax documents custom `width`/`height` only for `image-01`\n\n## Subject Reference\n\n- `--ref` files are sent as MiniMax `subject_reference`\n- `subject_reference[].type` is currently `character`\n- Official docs say `image_file` supports public URLs or Base64 Data URLs; baoyu-imagine sends local refs as Data URLs\n- Recommended refs: front-facing portraits, JPG/JPEG/PNG, under 10MB\n\n## Official References\n\n- [Image Generation Guide](https://platform.minimaxi.com/docs/guides/image-generation)\n- [Text-to-Image API](https://platform.minimaxi.com/docs/api-reference/image-generation-t2i)\n- [Image-to-Image API](https://platform.minimaxi.com/docs/api-reference/image-generation-i2i)\n\nFile v1.117.3:references/providers/openrouter.md\n\n# OpenRouter\n\nRead when the user picks `--provider openrouter`. Default model is `google/gemini-3.1-flash-image-preview`.\n\n## Common Models\n\nUse full OpenRouter model IDs:\n\n- `google/gemini-3.1-flash-image-preview` (recommended — supports image output and reference-image workflows)\n- `google/gemini-2.5-flash-image-preview`\n- `black-forest-labs/flux.2-pro`\n- Any other OpenRouter image-capable model ID\n\n## Behavior Notes\n\n- OpenRouter image generation uses `/chat/completions`, not the OpenAI `/images` endpoints\n- `--ref` requires a multimodal model that supports both image input and image output\n- `--imageSize` maps to `imageGenerationOptions.size`\n- `--size <WxH>` is converted to the nearest supported OpenRouter size, and the aspect ratio is inferred when possible\n\nFile v1.117.3:references/providers/replicate.md\n\n# Replicate\n\nRead when the user picks `--provider replicate`. Replicate support is intentionally scoped to model families baoyu-imagine can validate locally and save without dropping outputs.\n\n## Supported Families\n\n**`google/nano-banana*`** (default: `google/nano-banana-2`)\n\n- Supports prompt-only and reference-image generation\n- Uses Replicate `aspect_ratio`, `resolution`, and `output_format`\n- `--size <WxH>` is accepted only as a shorthand for a documented `aspect_ratio` plus `1K` / `2K`\n\n**`bytedance/seedream-4.5`**\n\n- Supports prompt-only and reference-image generation\n- Uses Replicate `size`, `aspect_ratio`, and `image_input`\n- Local validation blocks unsupported `1K` requests before the API call\n\n**`bytedance/seedream-5-lite`**\n\n- Supports prompt-only and reference-image generation\n- Uses Replicate `size`, `aspect_ratio`, and `image_input`\n- Local validation currently accepts `2K` / `3K` only\n\n**`wan-video/wan-2.7-image`**\n\n- Supports prompt-only and reference-image generation\n- Uses Replicate `size` and `images`\n- Max output is 2K\n\n**`wan-video/wan-2.7-image-pro`**\n\n- Supports prompt-only and reference-image generation\n- Uses Replicate `size` and `images`\n- 4K is allowed only for text-to-image; local validation blocks `4K + --ref`\n\n## Guardrails\n\n- Replicate currently supports only single-output save semantics in this tool — keep `--n 1`\n- If a model is outside the compatibility list above, baoyu-imagine treats it as prompt-only and rejects advanced local options instead of guessing a nano-banana-style schema\n\n## Examples\n\n```bash\n# Default model\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A cat\" --image out.png --provider replicate\n\n# Explicit model\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A cat\" --image out.png --provider replicate --model google/nano-banana\n```\n\nFile v1.117.3:references/providers/zai.md\n\n# Z.AI GLM-Image\n\nRead when the user picks `--provider zai` or sets `default_model.zai`. Default model is `glm-image`.\n\n## Models\n\n**`glm-image`** (recommended default)\n\n- Text-to-image only in baoyu-imagine (no `--ref` support yet)\n- Native `quality` options are `hd` and `standard`; this skill maps `2k → hd` and `normal → standard`\n- Recommended sizes: `1280x1280`, `1568x1056`, `1056x1568`, `1472x1088`, `1088x1472`, `1728x960`, `960x1728`\n- Custom `--size` requires width/height in `[1024, 2048]`, divisible by `32`, total pixels ≤ `2^22`\n\n**`cogview-4-250304`** (legacy family, same endpoint)\n\n- Custom `--size` requires width/height in `[512, 2048]`, divisible by `16`, total pixels ≤ `2^21`\n\n## Behavior Notes\n\n- The sync API returns a temporary URL; baoyu-imagine downloads it and writes locally\n- `--ref` is not supported for Z.AI in this skill yet\n- The sync API returns a single image, so `--n > 1` is rejected\n\n## Official References\n\n- [GLM-Image Guide](https://docs.z.ai/guides/image/glm-image)\n- [Generate Image API](https://docs.z.ai/api-reference/image/generate-image)\n\nFile v1.117.3:references/usage-examples.md\n\n# Usage Examples\n\nExtended CLI examples. SKILL.md shows the minimum set; read this file when the user asks about provider-specific invocation, batch generation, or less-common flags.\n\n## Core Patterns\n\n```bash\n# Basic text-to-image\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A cat\" --image cat.png\n\n# With aspect ratio\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A landscape\" --image out.png --ar 16:9\n\n# High quality\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A cat\" --image out.png --quality 2k\n\n# Prompt from files\n${BUN_X} {baseDir}/scripts/main.ts --promptfiles system.md content.md --image out.png\n\n# With reference images (any provider family that supports refs)\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"Make blue\" --image out.png --ref source.png\n```\n\n## Per-Provider\n\n```bash\n# OpenAI\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A cat\" --image out.png --provider openai --model gpt-image-2\n\n# Azure OpenAI (model = deployment name)\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A cat\" --image out.png --provider azure --model gpt-image-2\n\n# OpenAI GPT Image 2 custom 4K size\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A cinematic landscape\" --image out.png --provider openai --model gpt-image-2 --size 3840x2160\n\n# Google with explicit model\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"Make blue\" --image out.png --provider google --model gemini-3-pro-image-preview --ref source.png\n\n# OpenRouter (recommended default)\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A cat\" --image out.png --provider openrouter\n\n# OpenRouter with reference\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"Make blue\" --image out.png --provider openrouter --model google/gemini-3.1-flash-image-preview --ref source.png\n\n# DashScope (default model)\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"一只可爱的猫\" --image out.png --provider dashscope\n\n# DashScope Qwen-Image 2.0 Pro (custom size, Chinese text)\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"为咖啡品牌设计一张 21:9 横幅海报，包含清晰中文标题\" --image out.png --provider dashscope --model qwen-image-2.0-pro --size 2048x872\n\n# DashScope legacy fixed-size\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"一张电影感海报\" --image out.png --provider dashscope --model qwen-image-max --size 1664x928\n\n# DashScope Wan 2.7 Image Pro (4K text-to-image)\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"一间有着精致窗户的花店\" --image out.png --provider dashscope --model wan2.7-image-pro --size 4096x4096\n\n# DashScope Wan 2.7 Image with reference image (multi-image fusion)\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"把图2的涂鸦喷绘在图1的汽车上\" --image out.png --provider dashscope --model wan2.7-image-pro --ref car.webp paint.webp\n\n# Z.AI GLM-image\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"一张带清晰中文标题的科技海报\" --image out.png --provider zai\n\n# Z.AI with custom size\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A science illustration with labels\" --image out.png --provider zai --model glm-image --size 1472x1088\n\n# MiniMax\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A fashion editorial portrait\" --image out.jpg --provider minimax\n\n# MiniMax with subject reference (character/portrait consistency)\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A girl by the library window\" --image out.jpg --provider minimax --model image-01 --ref portrait.png --ar 16:9\n\n# Replicate (default: google/nano-banana-2)\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A cat\" --image out.png --provider replicate\n\n# Replicate Seedream 4.5\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A cinematic portrait\" --image out.png --provider replicate --model bytedance/seedream-4.5 --ar 3:2\n\n# Replicate Wan 2.7 Image Pro\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A concept frame\" --image out.png --provider replicate --model wan-video/wan-2.7-image-pro --size 2048x1152\n```\n\n## Batch Mode\n\n```bash\n# Batch from saved prompt files\n${BUN_X} {baseDir}/scripts/main.ts --batchfile batch.json\n\n# Batch with explicit worker count\n${BUN_X} {baseDir}/scripts/main.ts --batchfile batch.json --jobs 4 --json\n```\n\n### Batch File Format\n\n```json\n{\n  \"jobs\": 4,\n  \"tasks\": [\n    {\n      \"id\": \"hero\",\n      \"promptFiles\": [\"prompts/hero.md\"],\n      \"image\": \"out/hero.png\",\n      \"provider\": \"replicate\",\n      \"model\": \"google/nano-banana-2\",\n      \"ar\": \"16:9\",\n      \"quality\": \"2k\"\n    },\n    {\n      \"id\": \"diagram\",\n      \"promptFiles\": [\"prompts/diagram.md\"],\n      \"image\": \"out/diagram.png\",\n      \"ref\": [\"references/original.png\"]\n    }\n  ]\n}\n```\n\nPaths in `promptFiles`, `image`, and `ref` are resolved relative to the batch file's directory. `jobs` is optional (overridden by CLI `--jobs`). A top-level array without the `jobs` wrapper is also accepted.\n\nArchive v1.117.2: 35 files, 91535 bytes\n\nFiles: references/config/first-time-setup.md (13173b), references/config/preferences-schema.md (4224b), references/providers/dashscope.md (4591b), references/providers/minimax.md (1226b), references/providers/openrouter.md (776b), references/providers/replicate.md (1810b), references/providers/zai.md (1095b), references/usage-examples.md (4760b), scripts/build-batch.test.ts (4569b), scripts/build-batch.ts (7454b), scripts/main.test.ts (17852b), scripts/main.ts (43411b), scripts/providers/azure.test.ts (5423b), scripts/providers/azure.ts (5820b), scripts/providers/dashscope.test.ts (11534b), scripts/providers/dashscope.ts (18456b), scripts/providers/google.test.ts (3417b), scripts/providers/google.ts (9873b), scripts/providers/jimeng.test.ts (2737b), scripts/providers/jimeng.ts (12845b), scripts/providers/minimax.test.ts (5254b), scripts/providers/minimax.ts (6331b), scripts/providers/openai.test.ts (5536b), scripts/providers/openai.ts (12671b), scripts/providers/openrouter.test.ts (5450b), scripts/providers/openrouter.ts (9903b), scripts/providers/replicate.test.ts (7190b), scripts/providers/replicate.ts (17816b), scripts/providers/seedream.test.ts (6544b), scripts/providers/seedream.ts (9855b), scripts/providers/zai.test.ts (5274b), scripts/providers/zai.ts (8946b), scripts/types.ts (2123b), SKILL.md (14389b), _meta.json (134b)\n\nFile v1.117.2:SKILL.md\n\n---\nname: baoyu-imagine\ndescription: AI image generation with OpenAI GPT Image 2, Azure OpenAI, Google, OpenRouter, DashScope, Z.AI GLM-Image, MiniMax, Jimeng, Seedream and Replicate APIs. Supports text-to-image, reference images, aspect ratios, and batch generation from saved prompt files. Sequential by default; use batch parallel generation when the user already has multiple prompts or wants stable multi-image throughput. Use when user asks to generate, create, or draw images.\nversion: 1.58.0\nmetadata:\n  openclaw:\n    homepage: https://github.com/JimLiu/baoyu-skills#baoyu-imagine\n    requires:\n      anyBins:\n        - bun\n        - npx\n---\n\n# Image Generation (AI SDK)\n\nOfficial API-based image generation. Supports OpenAI GPT Image 2, Azure OpenAI, Google, OpenRouter, DashScope (阿里通义万象), Z.AI GLM-Image, MiniMax, Jimeng (即梦), Seedream (豆包) and Replicate.\n\n## User Input Tools\n\nWhen this skill prompts the user, follow this tool-selection rule (priority order):\n\n1. **Prefer built-in user-input tools** exposed by the current agent runtime — e.g., `AskUserQuestion`, `request_user_input`, `clarify`, `ask_user`, or any equivalent.\n2. **Fallback**: if no such tool exists, emit a numbered plain-text message and ask the user to reply with the chosen number/answer for each question.\n3. **Batching**: if the tool supports multiple questions per call, combine all applicable questions into a single call; if only single-question, ask them one at a time in priority order.\n\nConcrete `AskUserQuestion` references below are examples — substitute the local equivalent in other runtimes.\n\n## Script Directory\n\n`{baseDir}` = this SKILL.md's directory. Main script: `{baseDir}/scripts/main.ts`. Resolve `${BUN_X}`: prefer `bun`; else `npx -y bun`; else suggest `brew install oven-sh/bun/bun`.\n\n## Step 0: Load Preferences ⛔ BLOCKING\n\nThis step MUST complete before any image generation — generation is blocked until EXTEND.md exists.\n\nCheck these paths in order; first hit wins:\n\n| Path | Scope |\n|------|-------|\n| `.baoyu-skills/baoyu-imagine/EXTEND.md` | Project |\n| `${XDG_CONFIG_HOME:-$HOME/.config}/baoyu-skills/baoyu-imagine/EXTEND.md` | XDG |\n| `$HOME/.baoyu-skills/baoyu-imagine/EXTEND.md` | User home |\n\n- **Found** → load, parse, apply. If `default_model.[provider]` is null → ask model only.\n- **Not found** → run first-time setup (`references/config/first-time-setup.md`) using AskUserQuestion to collect provider + model + quality + save location. Save EXTEND.md, then continue. Do not generate images before this completes.\n\nLegacy compatibility: if `.baoyu-skills/baoyu-image-gen/EXTEND.md` exists and the new path doesn't, the runtime renames it to `baoyu-imagine`. If both exist, the runtime leaves them alone and uses the new path.\n\n**EXTEND.md keys**: default provider, default quality, default aspect ratio, default image size, OpenAI image API dialect, default models, batch worker cap, provider-specific batch limits. Schema: `references/config/preferences-schema.md`.\n\n## Usage\n\nMinimum working examples — see `references/usage-examples.md` for the full set including per-provider invocations and batch mode.\n\n```bash\n# Basic\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A cat\" --image cat.png\n\n# With aspect ratio and high quality\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A landscape\" --image out.png --ar 16:9 --quality 2k\n\n# Prompt from files\n${BUN_X} {baseDir}/scripts/main.ts --promptfiles system.md content.md --image out.png\n\n# With reference image\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"Make blue\" --image out.png --ref source.png\n\n# Specific provider\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A cat\" --image out.png --provider dashscope --model qwen-image-2.0-pro\n\n# OpenAI GPT Image 2\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A cat\" --image out.png --provider openai --model gpt-image-2\n\n# Batch mode\n${BUN_X} {baseDir}/scripts/main.ts --batchfile batch.json --jobs 4\n```\n\n## Options\n\n| Option | Description |\n|--------|-------------|\n| `--prompt <text>`, `-p` | Prompt text |\n| `--promptfiles <files...>` | Read prompt from files (concatenated) |\n| `--image <path>` | Output image path (required in single-image mode) |\n| `--batchfile <path>` | JSON batch file for multi-image generation |\n| `--jobs <count>` | Worker count for batch mode (default: auto, max from config, built-in default 10) |\n| `--provider google\\|openai\\|azure\\|openrouter\\|dashscope\\|zai\\|minimax\\|jimeng\\|seedream\\|replicate` | Force provider (default: auto-detect) |\n| `--model <id>`, `-m` | Model ID — see provider references for defaults and allowed values |\n| `--ar <ratio>` | Aspect ratio (`16:9`, `1:1`, `4:3`, …) |\n| `--size <WxH>` | Explicit size (e.g., `1024x1024`; for `gpt-image-2`, width/height must be multiples of 16, max edge 3840px, ratio no wider than 3:1) |\n| `--quality normal\\|2k` | Quality preset (default: `2k`) |\n| `--imageSize 1K\\|2K\\|4K` | Image size for Google/OpenRouter (default: from quality) |\n| `--imageApiDialect openai-native\\|ratio-metadata` | OpenAI-compatible endpoint dialect — use `ratio-metadata` for gateways that expect aspect-ratio `size` plus `metadata.resolution` |\n| `--ref <files...>` | Reference images. Supported by Google multimodal, OpenAI GPT Image edits, Azure OpenAI edits (PNG/JPG only), OpenRouter multimodal models, Replicate supported families, MiniMax subject-reference, Seedream 5.0/4.5/4.0, DashScope `wan2.7-image-pro`/`wan2.7-image`. Not supported by Jimeng, Seedream 3.0, SeedEdit 3.0, or any DashScope model outside the `wan2.7-image*` family |\n| `--n <count>` | Number of images. Replicate requires `--n 1` (single-output save semantics) |\n| `--json` | JSON output |\n\n## Environment Variables\n\n| Variable | Description |\n|----------|-------------|\n| `OPENAI_API_KEY` | OpenAI API key |\n| `AZURE_OPENAI_API_KEY` | Azure OpenAI API key |\n| `OPENROUTER_API_KEY` | OpenRouter API key |\n| `GOOGLE_API_KEY` | Google API key |\n| `DASHSCOPE_API_KEY` | DashScope API key |\n| `ZAI_API_KEY` (alias `BIGMODEL_API_KEY`) | Z.AI API key |\n| `MINIMAX_API_KEY` | MiniMax API key |\n| `REPLICATE_API_TOKEN` | Replicate API token |\n| `JIMENG_ACCESS_KEY_ID`, `JIMENG_SECRET_ACCESS_KEY` | Jimeng (即梦) Volcengine credentials |\n| `ARK_API_KEY` | Seedream (豆包) Volcengine ARK API key |\n| `<PROVIDER>_IMAGE_MODEL` | Per-provider model override (`OPENAI_IMAGE_MODEL`, `GOOGLE_IMAGE_MODEL`, `DASHSCOPE_IMAGE_MODEL`, `ZAI_IMAGE_MODEL`/`BIGMODEL_IMAGE_MODEL`, `MINIMAX_IMAGE_MODEL`, `OPENROUTER_IMAGE_MODEL`, `REPLICATE_IMAGE_MODEL`, `JIMENG_IMAGE_MODEL`, `SEEDREAM_IMAGE_MODEL`) |\n| `AZURE_OPENAI_DEPLOYMENT` (alias `AZURE_OPENAI_IMAGE_MODEL`) | Azure default deployment |\n| `<PROVIDER>_BASE_URL` | Per-provider endpoint override |\n| `AZURE_API_VERSION` | Azure image API version (default `2025-04-01-preview`) |\n| `JIMENG_REGION` | Jimeng region (default `cn-north-1`) |\n| `OPENAI_IMAGE_API_DIALECT` | `openai-native` \\| `ratio-metadata` |\n| `OPENROUTER_HTTP_REFERER`, `OPENROUTER_TITLE` | Optional OpenRouter attribution |\n| `BAOYU_IMAGE_GEN_MAX_WORKERS` | Override batch worker cap |\n| `BAOYU_IMAGE_GEN_<PROVIDER>_CONCURRENCY` | Per-provider concurrency (e.g., `BAOYU_IMAGE_GEN_REPLICATE_CONCURRENCY`) |\n| `BAOYU_IMAGE_GEN_<PROVIDER>_START_INTERVAL_MS` | Per-provider start-gap |\n\n**Load priority**: CLI args > EXTEND.md > env vars > `<cwd>/.baoyu-skills/.env` > `~/.baoyu-skills/.env`\n\n## Model Resolution\n\nPriority (highest → lowest) applies to every provider:\n\n1. CLI flag `--model <id>`\n2. EXTEND.md `default_model.[provider]`\n3. Env var `<PROVIDER>_IMAGE_MODEL`\n4. Built-in default\n\nFor OpenAI, the built-in default is `gpt-image-2`. `gpt-image-1.5`, `gpt-image-1`, and GPT Image snapshots remain selectable with `--model` or `OPENAI_IMAGE_MODEL`.\n\nFor Azure, `--model` / `default_model.azure` is the Azure deployment name. `AZURE_OPENAI_DEPLOYMENT` is the preferred env var; `AZURE_OPENAI_IMAGE_MODEL` is kept as a backward-compatible alias. If your Azure deployment is named after the underlying model, use `gpt-image-2`; otherwise use the exact custom deployment name.\n\nEXTEND.md overrides env vars: if EXTEND.md sets `default_model.google: \"gemini-3-pro-image-preview\"` and the env var sets `GOOGLE_IMAGE_MODEL=gemini-3.1-flash-image-preview`, EXTEND.md wins.\n\n**Display model info before each generation**:\n\n- `Using [provider] / [model]`\n- `Switch model: --model <id> | EXTEND.md default_model.[provider] | env <PROVIDER>_IMAGE_MODEL`\n\n## OpenAI-Compatible Gateway Dialects\n\n`provider=openai` means the auth and routing entrypoint is OpenAI-compatible. It does **not** guarantee the upstream image API uses OpenAI native semantics. When a gateway expects a different wire format, set `default_image_api_dialect` in EXTEND.md, `OPENAI_IMAGE_API_DIALECT`, or `--imageApiDialect`:\n\n- `openai-native`: pixel `size` (`1536x1024`) and native OpenAI quality fields\n- `ratio-metadata`: aspect-ratio `size` (`16:9`) plus `metadata.resolution` (`1K|2K|4K`) and `metadata.orientation`\n\nUse `openai-native` for the OpenAI native API or strict clones; try `ratio-metadata` for compatibility gateways in front of Gemini or similar models. Current limitation: `ratio-metadata` applies only to text-to-image; reference-image edits still need `openai-native` or a provider with first-class edit support.\n\n## Provider-Specific Guides\n\nEach provider has its own quirks (model families, size rules, ref support, limits). Read these when the user picks that provider or asks for non-default behavior:\n\n| Provider | Reference |\n|----------|-----------|\n| DashScope (Qwen-Image families, custom sizes) | `references/providers/dashscope.md` |\n| Z.AI (GLM-Image, cogview-4) | `references/providers/zai.md` |\n| MiniMax (image-01, subject-reference) | `references/providers/minimax.md` |\n| OpenRouter (multimodal models, `/chat/completions` flow) | `references/providers/openrouter.md` |\n| Replicate (nano-banana, Seedream, Wan) | `references/providers/replicate.md` |\n\n## Provider Selection\n\n1. `--ref` provided + no `--provider` → auto-select Google → OpenAI → Azure → OpenRouter → Replicate → Seedream → MiniMax (MiniMax's subject reference is more specialized toward character/portrait consistency)\n2. `--provider` specified → use it (if `--ref`, must be google/openai/azure/openrouter/replicate/seedream/minimax)\n3. Only one API key present → use that provider\n4. Multiple keys → default priority: Google → OpenAI → Azure → OpenRouter → DashScope → Z.AI → MiniMax → Replicate → Jimeng → Seedream\n\n## Quality Presets\n\n| Preset | Google imageSize | OpenAI size | OpenRouter size | Replicate resolution | Use case |\n|--------|------------------|-------------|-----------------|----------------------|----------|\n| `normal` | 1K | 1024px target | 1K | 1K | Quick previews |\n| `2k` (default) | 2K | 2048px target | 2K | 2K | Covers, illustrations, infographics |\n\nGoogle/OpenRouter `imageSize` can be overridden with `--imageSize 1K|2K|4K`.\n\nFor OpenAI native `gpt-image-2`, `normal` maps to `quality=medium` and a low-latency valid size near the requested aspect ratio; `2k` maps to `quality=high` and 2048px-class sizes such as `2048x2048`, `2048x1152`, or `1152x2048`. Use explicit `--size` for valid custom or 4K outputs, e.g. `3840x2160`.\n\n## Aspect Ratios\n\nSupported: `1:1`, `16:9`, `9:16`, `4:3`, `3:4`, `2.35:1`.\n\n- Google multimodal: `imageConfig.aspectRatio`\n- OpenAI: `gpt-image-2` uses the closest valid custom size for the requested ratio; older GPT Image and DALL·E models use their closest supported fixed size\n- OpenRouter: `imageGenerationOptions.aspect_ratio`; if only `--size <WxH>` is given, the ratio is inferred\n- Replicate: behavior is model-specific — `google/nano-banana*` uses `aspect_ratio`, `bytedance/seedream-*` uses documented Replicate ratios, Wan 2.7 maps `--ar` to a concrete `size`\n- MiniMax: official `aspect_ratio` values; if `--size <WxH>` is given without `--ar`, sends `width`/`height` for `image-01`\n\n## Generation Mode\n\n**Default**: sequential. **Batch parallel**: enabled automatically when `--batchfile` contains 2+ pending tasks.\n\n| Situation | Prefer | Why |\n|-----------|--------|-----|\n| One image, or 1-2 simple images | Sequential | Lower coordination overhead, easier debugging |\n| Multiple images with saved prompt files | Batch (`--batchfile`) | Reuses finalized prompts, applies shared throttling/retries, predictable throughput |\n| Each image still needs its own reasoning / prompt writing / style exploration | Subagents | Work is still exploratory, each needs independent analysis |\n| Input is `outline.md` + `prompts/` (e.g. from `baoyu-article-illustrator`) | Batch — use `scripts/build-batch.ts` to assemble the payload | The outline + prompt files already contain everything needed |\n\nRule of thumb: once prompt files are saved and the task is \"generate all of these\", prefer batch over subagents. Use subagents only when generation is coupled with per-image thinking or divergent creative exploration.\n\n**Parallel behavior**:\n\n- Default worker count is automatic, capped by config, built-in default 10\n- Provider-specific throttling applies only in batch mode; defaults are tuned for throughput while avoiding RPM bursts\n- Override with `--jobs <count>`\n- Each image retries up to 3 attempts\n- Final output includes success count, failure count, and per-image failure reasons\n\n## Error Handling\n\n- Missing API key → error with setup instructions\n- Generation failure → auto-retry up to 3 attempts per image\n- Invalid aspect ratio → warning, proceed with default\n- Reference images with unsupported provider/model → error with fix hint\n\n## References\n\n| File | Content |\n|------|---------|\n| `references/usage-examples.md` | Extended CLI examples across providers and batch mode |\n| `references/providers/dashscope.md` | DashScope families, sizes, limits |\n| `references/providers/zai.md` | Z.AI GLM-image / cogview-4 |\n| `references/providers/minimax.md` | MiniMax image-01 + subject reference |\n| `references/providers/openrouter.md` | OpenRouter multimodal flow |\n| `references/providers/replicate.md` | Replicate supported families + guardrails |\n| `references/config/preferences-schema.md` | EXTEND.md schema |\n| `references/config/first-time-setup.md` | First-time setup flow |\n\n## Extension Support\n\nCustom configurations via EXTEND.md. See Step 0 for paths and schema.\n\nFile v1.117.2:_meta.json\n\n{\n  \"ownerId\": \"kn7csrrndw79hpke5d0gsnx93d82k67r\",\n  \"slug\": \"baoyu-imagine\",\n  \"version\": \"1.117.2\",\n  \"publishedAt\": 1779070508352\n}\n\nFile v1.117.2:references/config/first-time-setup.md\n\n---\nname: first-time-setup\ndescription: First-time setup and default model selection flow for baoyu-imagine\n---\n\n# First-Time Setup\n\n## Overview\n\nTriggered when:\n1. No EXTEND.md found → full setup (provider + model + preferences)\n2. EXTEND.md found but `default_model.[provider]` is null → model selection only\n\n## Setup Flow\n\n```\nNo EXTEND.md found          EXTEND.md found, model null\n        │                            │\n        ▼                            ▼\n┌─────────────────────┐    ┌──────────────────────┐\n│ AskUserQuestion     │    │ AskUserQuestion      │\n│ (full setup)        │    │ (model only)         │\n└─────────────────────┘    └──────────────────────┘\n        │                            │\n        ▼                            ▼\n┌─────────────────────┐    ┌──────────────────────┐\n│ Create EXTEND.md    │    │ Update EXTEND.md     │\n└─────────────────────┘    └──────────────────────┘\n        │                            │\n        ▼                            ▼\n    Continue                     Continue\n```\n\n## Flow 1: No EXTEND.md (Full Setup)\n\n**Language**: Use user's input language or saved language preference.\n\nUse AskUserQuestion with ALL questions in ONE call:\n\n### Question 1: Default Provider\n\n```yaml\nheader: \"Provider\"\nquestion: \"Default image generation provider?\"\noptions:\n  - label: \"Google (Recommended)\"\n    description: \"Gemini multimodal - high quality, reference images, flexible sizes\"\n  - label: \"OpenAI\"\n    description: \"GPT Image 2 - latest OpenAI image model, reference-image workflows\"\n  - label: \"Azure OpenAI\"\n    description: \"Azure-hosted GPT Image deployments with resource-specific routing\"\n  - label: \"OpenRouter\"\n    description: \"Router for Gemini/FLUX/OpenAI-compatible image models\"\n  - label: \"DashScope\"\n    description: \"Alibaba Cloud - Qwen-Image, strong Chinese/English text rendering\"\n  - label: \"Z.AI\"\n    description: \"GLM-image, strong poster and text-heavy image generation\"\n  - label: \"MiniMax\"\n    description: \"MiniMax image generation with subject-reference character workflows\"\n  - label: \"Replicate\"\n    description: \"Curated Replicate image families - nano-banana-2, Seedream, and Wan image models\"\n```\n\n### Question 2: Default Google Model\n\nOnly show if user selected Google or auto-detect (no explicit provider).\n\n```yaml\nheader: \"Google Model\"\nquestion: \"Default Google image generation model?\"\noptions:\n  - label: \"gemini-3-pro-image-preview (Recommended)\"\n    description: \"Highest quality, best for production use\"\n  - label: \"gemini-3.1-flash-image-preview\"\n    description: \"Fast generation, good quality, lower cost\"\n  - label: \"gemini-3-flash-preview\"\n    description: \"Fast generation, balanced quality and speed\"\n```\n\n### Question 2b: Default OpenRouter Model\n\nOnly show if user selected OpenRouter.\n\n```yaml\nheader: \"OpenRouter Model\"\nquestion: \"Default OpenRouter image generation model?\"\noptions:\n  - label: \"google/gemini-3.1-flash-image-preview (Recommended)\"\n    description: \"Best general-purpose OpenRouter image model with reference-image workflows\"\n  - label: \"google/gemini-2.5-flash-image-preview\"\n    description: \"Fast Gemini preview model on OpenRouter\"\n  - label: \"black-forest-labs/flux.2-pro\"\n    description: \"Strong text-to-image quality through OpenRouter\"\n```\n\n### Question 2c: Default Azure Deployment\n\nOnly show if user selected Azure OpenAI.\n\n```yaml\nheader: \"Azure Deploy\"\nquestion: \"Default Azure image deployment name?\"\noptions:\n  - label: \"gpt-image-2 (Recommended)\"\n    description: \"Use if your Azure deployment uses the GPT Image 2 model name\"\n  - label: \"gpt-image-1.5\"\n    description: \"Previous GPT Image deployment name\"\n  - label: \"gpt-image-1\"\n    description: \"Earlier GPT Image deployment name\"\n```\n\n### Question 2d: Default MiniMax Model\n\nOnly show if user selected MiniMax.\n\n```yaml\nheader: \"MiniMax Model\"\nquestion: \"Default MiniMax image generation model?\"\noptions:\n  - label: \"image-01 (Recommended)\"\n    description: \"Best default, supports aspect ratios and custom width/height\"\n  - label: \"image-01-live\"\n    description: \"Faster variant, use aspect ratio instead of custom size\"\n```\n\n### Question 2e: Default Z.AI Model\n\nOnly show if user selected Z.AI.\n\n```yaml\nheader: \"Z.AI Model\"\nquestion: \"Default Z.AI image generation model?\"\noptions:\n  - label: \"glm-image (Recommended)\"\n    description: \"Best default for posters, diagrams, and text-heavy images\"\n  - label: \"cogview-4-250304\"\n    description: \"Legacy Z.AI image model on the same endpoint\"\n```\n\n### Question 3: Default Quality\n\n```yaml\nheader: \"Quality\"\nquestion: \"Default image quality?\"\noptions:\n  - label: \"2k (Recommended)\"\n    description: \"2048px - covers, illustrations, infographics\"\n  - label: \"normal\"\n    description: \"1024px - quick previews, drafts\"\n```\n\n### Question 4: Save Location\n\n```yaml\nheader: \"Save\"\nquestion: \"Where to save preferences?\"\noptions:\n  - label: \"Project (Recommended)\"\n    description: \".baoyu-skills/ (this project only)\"\n  - label: \"User\"\n    description: \"~/.baoyu-skills/ (all projects)\"\n```\n\n### Save Locations\n\n| Choice | Path | Scope |\n|--------|------|-------|\n| Project | `.baoyu-skills/baoyu-imagine/EXTEND.md` | Current project |\n| User | `$HOME/.baoyu-skills/baoyu-imagine/EXTEND.md` | All projects |\n\n### EXTEND.md Template\n\n```yaml\n---\nversion: 1\ndefault_provider: [selected provider or null]\ndefault_quality: [selected quality]\ndefault_aspect_ratio: null\ndefault_image_size: null\ndefault_image_api_dialect: null\ndefault_model:\n  google: [selected google model or null]\n  openai: null\n  azure: [selected azure deployment or null]\n  openrouter: [selected openrouter model or null]\n  dashscope: null\n  zai: [selected Z.AI model or null]\n  minimax: [selected minimax model or null]\n  replicate: null\n---\n```\n\nIf the user selects `OpenAI` but says their endpoint is only OpenAI-compatible and fronts another image model family, save `default_image_api_dialect: ratio-metadata` when they explicitly confirm the gateway expects aspect-ratio `size` plus metadata-based resolution. Otherwise leave it `null` / `openai-native`.\n\n## Flow 2: EXTEND.md Exists, Model Null\n\nWhen EXTEND.md exists but `default_model.[current_provider]` is null, ask ONLY the model question for the current provider.\n\n### Google Model Selection\n\n```yaml\nheader: \"Google Model\"\nquestion: \"Choose a default Google image generation model?\"\noptions:\n  - label: \"gemini-3-pro-image-preview (Recommended)\"\n    description: \"Highest quality, best for production use\"\n  - label: \"gemini-3.1-flash-image-preview\"\n    description: \"Fast generation, good quality, lower cost\"\n  - label: \"gemini-3-flash-preview\"\n    description: \"Fast generation, balanced quality and speed\"\n```\n\n### OpenAI Model Selection\n\n```yaml\nheader: \"OpenAI Model\"\nquestion: \"Choose a default OpenAI image generation model?\"\noptions:\n  - label: \"gpt-image-2 (Recommended)\"\n    description: \"Latest GPT Image model, flexible sizes up to 4K, high-fidelity image inputs\"\n  - label: \"gpt-image-1.5\"\n    description: \"Previous GPT Image model\"\n  - label: \"gpt-image-1\"\n    description: \"Earlier GPT Image model\"\n```\n\n### Azure Deployment Selection\n\n```yaml\nheader: \"Azure Deploy\"\nquestion: \"Choose a default Azure image deployment name?\"\noptions:\n  - label: \"gpt-image-2 (Recommended)\"\n    description: \"Use when your Azure deployment name matches the GPT Image 2 model\"\n  - label: \"gpt-image-1.5\"\n    description: \"Use when your Azure deployment name matches the GPT Image 1.5 model\"\n  - label: \"gpt-image-1\"\n    description: \"Use when your Azure deployment name matches GPT-image-1\"\n```\n\nNotes for Azure setup:\n\n- In `baoyu-imagine`, Azure `--model` / `default_model.azure` should be the Azure deployment name, not just the underlying model family.\n- If the deployment name is custom, save that exact deployment name in `default_model.azure`.\n\n### OpenRouter Model Selection\n\n```yaml\nheader: \"OpenRouter Model\"\nquestion: \"Choose a default OpenRouter image generation model?\"\noptions:\n  - label: \"google/gemini-3.1-flash-image-preview (Recommended)\"\n    description: \"Recommended for image output and reference-image edits\"\n  - label: \"google/gemini-2.5-flash-image-preview\"\n    description: \"Fast preview-oriented image generation\"\n  - label: \"black-forest-labs/flux.2-pro\"\n    description: \"High-quality text-to-image through OpenRouter\"\n```\n\n### DashScope Model Selection\n\n```yaml\nheader: \"DashScope Model\"\nquestion: \"Choose a default DashScope image generation model?\"\noptions:\n  - label: \"qwen-image-2.0-pro (Recommended)\"\n    description: \"Best DashScope model for text rendering and custom sizes\"\n  - label: \"qwen-image-2.0\"\n    description: \"Faster 2.0 variant with flexible output size\"\n  - label: \"qwen-image-max\"\n    description: \"Legacy Qwen model with five fixed output sizes\"\n  - label: \"qwen-image-plus\"\n    description: \"Legacy Qwen model, same current capability as qwen-image\"\n  - label: \"wan2.7-image-pro\"\n    description: \"Wan 2.7 Pro — supports up to 4K text-to-image and reference-image editing\"\n  - label: \"wan2.7-image\"\n    description: \"Wan 2.7 base — faster generation, up to 2K, supports reference-image editing\"\n  - label: \"z-image-turbo\"\n    description: \"Legacy DashScope model for compatibility\"\n  - label: \"z-image-ultra\"\n    description: \"Legacy DashScope model, higher quality but slower\"\n```\n\nNotes for DashScope setup:\n\n- Prefer `qwen-image-2.0-pro` when the user needs custom `--size`, uncommon ratios like `21:9`, or strong Chinese/English text rendering.\n- `qwen-image-max` / `qwen-image-plus` / `qwen-image` only support five fixed sizes: `1664*928`, `1472*1104`, `1328*1328`, `1104*1472`, `928*1664`.\n- `wan2.7-image-pro` and `wan2.7-image` are the only DashScope models that accept `--ref`. Pick one of these when the user wants reference-image editing or multi-image fusion via DashScope.\n- In `baoyu-imagine`, `quality` is a compatibility preset. It is not a native DashScope parameter.\n\n### Z.AI Model Selection\n\n```yaml\nheader: \"Z.AI Model\"\nquestion: \"Choose a default Z.AI image generation model?\"\noptions:\n  - label: \"glm-image (Recommended)\"\n    description: \"Current flagship image model with better text rendering and poster layouts\"\n  - label: \"cogview-4-250304\"\n    description: \"Legacy model on the sync image endpoint\"\n```\n\nNotes for Z.AI setup:\n\n- Prefer `glm-image` for posters, diagrams, and Chinese/English text-heavy layouts.\n- In `baoyu-imagine`, Z.AI currently exposes text-to-image only; reference images are not wired for this provider.\n- The sync Z.AI image API returns a downloadable image URL, which the runtime saves locally after download.\n\n### Replicate Model Selection\n\n```yaml\nheader: \"Replicate Model\"\nquestion: \"Choose a default Replicate image generation model?\"\noptions:\n  - label: \"google/nano-banana-2 (Recommended)\"\n    description: \"Current default for general Replicate image generation in baoyu-imagine\"\n  - label: \"bytedance/seedream-4.5\"\n    description: \"Replicate Seedream 4.5 with validated local size/ref guardrails\"\n  - label: \"bytedance/seedream-5-lite\"\n    description: \"Replicate Seedream 5 Lite with validated local size/ref guardrails\"\n  - label: \"wan-video/wan-2.7-image-pro\"\n    description: \"Replicate Wan 2.7 Image Pro with 4K text-to-image support\"\n```\n\n### MiniMax Model Selection\n\n```yaml\nheader: \"MiniMax Model\"\nquestion: \"Choose a default MiniMax image generation model?\"\noptions:\n  - label: \"image-01 (Recommended)\"\n    description: \"Best general-purpose MiniMax image model with custom width/height support\"\n  - label: \"image-01-live\"\n    description: \"Lower-latency MiniMax image model using aspect ratios\"\n```\n\nNotes for MiniMax setup:\n\n- `image-01` is the safest default. It supports official `aspect_ratio` values and documented custom `width` / `height` output sizes.\n- `image-01-live` is useful when the user prefers faster generation and can work with aspect-ratio-based sizing.\n- MiniMax subject reference currently uses `subject_reference[].type = character`; docs recommend front-facing portrait references in JPG/JPEG/PNG under 10MB.\n\n### Update EXTEND.md\n\nAfter user selects a model:\n\n1. Read existing EXTEND.md\n2. If `default_model:` section exists → update the provider-specific key\n3. If `default_model:` section missing → add the full section:\n\n```yaml\ndefault_model:\n  google: [value or null]\n  openai: [value or null]\n  azure: [value or null]\n  openrouter: [value or null]\n  dashscope: [value or null]\n  zai: [value or null]\n  minimax: [value or null]\n  replicate: [value or null]\n```\n\nOnly set the selected provider's model; leave others as their current value or null.\n\n## After Setup\n\n1. Create directory if needed\n2. Write/update EXTEND.md with frontmatter\n3. Confirm: \"Preferences saved to [path]\"\n4. Continue with image generation\n\nFile v1.117.2:references/config/preferences-schema.md\n\n---\nname: preferences-schema\ndescription: EXTEND.md YAML schema for baoyu-imagine user preferences\n---\n\n# Preferences Schema\n\n## Full Schema\n\n```yaml\n---\nversion: 1\n\ndefault_provider: null      # google|openai|azure|openrouter|dashscope|zai|minimax|replicate|null (null = auto-detect)\n\ndefault_quality: null       # normal|2k|null (null = use default: 2k)\n\ndefault_aspect_ratio: null  # \"16:9\"|\"1:1\"|\"4:3\"|\"3:4\"|\"2.35:1\"|null\n\ndefault_image_size: null    # 1K|2K|4K|null (Google/OpenRouter, overrides quality)\n\ndefault_image_api_dialect: null  # openai-native|ratio-metadata|null (OpenAI-compatible gateways; null = use env/default)\n\ndefault_model:\n  google: null              # e.g., \"gemini-3-pro-image-preview\", \"gemini-3.1-flash-image-preview\"\n  openai: null              # e.g., \"gpt-image-2\", \"gpt-image-1.5\", \"gpt-image-1\"\n  azure: null               # Azure deployment name, e.g., \"gpt-image-2\" or \"image-prod\"\n  openrouter: null          # e.g., \"google/gemini-3.1-flash-image-preview\"\n  dashscope: null           # e.g., \"qwen-image-2.0-pro\"\n  zai: null                 # e.g., \"glm-image\"\n  minimax: null             # e.g., \"image-01\"\n  replicate: null           # e.g., \"google/nano-banana-2\"\n\nbatch:\n  max_workers: 10\n  provider_limits:\n    replicate:\n      concurrency: 5\n      start_interval_ms: 700\n    google:\n      concurrency: 3\n      start_interval_ms: 1100\n    openai:\n      concurrency: 3\n      start_interval_ms: 1100\n    azure:\n      concurrency: 3\n      start_interval_ms: 1100\n    openrouter:\n      concurrency: 3\n      start_interval_ms: 1100\n    dashscope:\n      concurrency: 3\n      start_interval_ms: 1100\n    zai:\n      concurrency: 3\n      start_interval_ms: 1100\n    minimax:\n      concurrency: 3\n      start_interval_ms: 1100\n---\n```\n\n## Field Reference\n\n| Field | Type | Default | Description |\n|-------|------|---------|-------------|\n| `version` | int | 1 | Schema version |\n| `default_provider` | string\\|null | null | Default provider (null = auto-detect) |\n| `default_quality` | string\\|null | null | Default quality (null = 2k) |\n| `default_aspect_ratio` | string\\|null | null | Default aspect ratio |\n| `default_image_size` | string\\|null | null | Google/OpenRouter image size (overrides quality) |\n| `default_image_api_dialect` | string\\|null | null | OpenAI-compatible image dialect (`openai-native` or `ratio-metadata`) |\n| `default_model.google` | string\\|null | null | Google default model |\n| `default_model.openai` | string\\|null | null | OpenAI default model |\n| `default_model.azure` | string\\|null | null | Azure default deployment name |\n| `default_model.openrouter` | string\\|null | null | OpenRouter default model |\n| `default_model.dashscope` | string\\|null | null | DashScope default model |\n| `default_model.zai` | string\\|null | null | Z.AI default model |\n| `default_model.minimax` | string\\|null | null | MiniMax default model |\n| `default_model.replicate` | string\\|null | null | Replicate default model |\n| `batch.max_workers` | int\\|null | 10 | Batch worker cap |\n| `batch.provider_limits.<provider>.concurrency` | int\\|null | provider default | Max simultaneous requests per provider |\n| `batch.provider_limits.<provider>.start_interval_ms` | int\\|null | provider default | Minimum gap between request starts per provider |\n\n## Examples\n\n**Minimal**:\n```yaml\n---\nversion: 1\ndefault_provider: google\ndefault_quality: 2k\ndefault_image_api_dialect: null\n---\n```\n\n**Full**:\n```yaml\n---\nversion: 1\ndefault_provider: google\ndefault_quality: 2k\ndefault_aspect_ratio: \"16:9\"\ndefault_image_size: 2K\ndefault_image_api_dialect: null\ndefault_model:\n  google: \"gemini-3-pro-image-preview\"\n  openai: \"gpt-image-2\"\n  azure: \"gpt-image-2\"\n  openrouter: \"google/gemini-3.1-flash-image-preview\"\n  dashscope: \"qwen-image-2.0-pro\"\n  zai: \"glm-image\"\n  minimax: \"image-01\"\n  replicate: \"google/nano-banana-2\"\nbatch:\n  max_workers: 10\n  provider_limits:\n    replicate:\n      concurrency: 5\n      start_interval_ms: 700\n    azure:\n      concurrency: 3\n      start_interval_ms: 1100\n    zai:\n      concurrency: 3\n      start_interval_ms: 1100\n    openrouter:\n      concurrency: 3\n      start_interval_ms: 1100\n    minimax:\n      concurrency: 3\n      start_interval_ms: 1100\n---\n```\n\nFile v1.117.2:references/providers/dashscope.md\n\n# DashScope (阿里通义万象)\n\nRead when the user picks `--provider dashscope`, sets `default_model.dashscope`, or asks for Qwen-Image behavior. The SKILL.md only names the default — this file covers model families, sizing rules, and limits.\n\n## Model Families\n\n**`qwen-image-2.0*`** — recommended modern family. Members: `qwen-image-2.0-pro`, `qwen-image-2.0-pro-2026-03-03`, `qwen-image-2.0`, `qwen-image-2.0-2026-03-03`.\n\n- Free-form `size` in `宽*高` format\n- Total pixels must be between `512*512` and `2048*2048`\n- Default ≈ `1024*1024`\n- Best choice for custom ratios (e.g. `21:9`) and text-heavy Chinese/English layouts\n\n**Fixed-size family** — `qwen-image-max`, `qwen-image-max-2025-12-30`, `qwen-image-plus`, `qwen-image-plus-2026-01-09`, `qwen-image`.\n\n- Only five sizes allowed: `1664*928`, `1472*1104`, `1328*1328`, `1104*1472`, `928*1664`\n- Default is `1664*928`\n- `qwen-image` currently has the same capability as `qwen-image-plus`\n\n**`wan2.7-image*`** — multimodal Wan 2.7 family. Members: `wan2.7-image-pro`, `wan2.7-image`.\n\n- Free-form `size` in `宽*高` format, plus aspect-ratio inference\n- `wan2.7-image-pro` text-to-image (no `--ref`): total pixels in `[768*768, 4096*4096]`, ratio in `[1:8, 8:1]`\n- `wan2.7-image-pro` with reference images and `wan2.7-image` (all scenarios): total pixels in `[768*768, 2048*2048]`, ratio in `[1:8, 8:1]`\n- Default: `1024*1024` (`--quality normal`) or `2048*2048` (`--quality 2k`); 4K requires explicit `--size`\n- Supports up to 9 reference images in `--ref` (image editing / multi-image fusion)\n- Reference images are sent inline as base64 (or passed through if the path is an `http(s)://` URL)\n- API does NOT use `prompt_extend`; the skill omits it for this family\n- The Wan 2.7 API defaults `n` to **4** in non-collage mode and bills per generated image. baoyu-imagine forces `n: 1` and rejects `--n > 1` to avoid silently paying for and discarding extra images.\n\n**Legacy** — `z-image-turbo`, `z-image-ultra`, `wanx-v1`. Only use when the user explicitly asks for legacy behavior.\n\n## Size Resolution\n\n- `--size` wins over `--ar`\n- For `qwen-image-2.0*`: prefer explicit `--size`; otherwise infer from `--ar` using the recommended table below\n- For `qwen-image-max/plus/image`: only use the five fixed sizes; if the requested ratio doesn't fit, switch to `qwen-image-2.0-pro`\n- For `wan2.7-image*`: explicit `--size` is validated against the per-mode pixel/ratio limits; otherwise the size is derived from `--ar` and `--quality` (`normal` ≈ 1K, `2k` ≈ 2K). To request 4K with `wan2.7-image-pro` text-to-image, pass `--size` explicitly (e.g. `4096*4096`, `3840*2160`)\n- `--quality` is a baoyu-imagine preset, not an official DashScope field. The mapping of `normal`/`2k` onto the `qwen-image-2.0*` and `wan2.7-image*` tables is an implementation choice, not an API guarantee\n\n### Recommended `qwen-image-2.0*` sizes\n\n| Ratio | `normal` | `2k` |\n|-------|----------|------|\n| `1:1` | `1024*1024` | `1536*1536` |\n| `2:3` | `768*1152` | `1024*1536` |\n| `3:2` | `1152*768` | `1536*1024` |\n| `3:4` | `960*1280` | `1080*1440` |\n| `4:3` | `1280*960` | `1440*1080` |\n| `9:16` | `720*1280` | `1080*1920` |\n| `16:9` | `1280*720` | `1920*1080` |\n| `21:9` | `1344*576` | `2048*872` |\n\n## Reference Images\n\n- Only `wan2.7-image-pro` and `wan2.7-image` accept `--ref`. Other DashScope models (qwen-image-2.0*, qwen-image-max/plus/image, legacy) reject `--ref` and the user is steered to a different provider/model.\n- Up to 9 reference images per request. Local files are inlined as base64 data URLs; `http(s)://` URLs are forwarded as-is.\n- Supplying any `--ref` automatically clamps the wan2.7-image-pro pixel ceiling from 4K to 2K (the API only supports 4K for pure text-to-image with no image input).\n\n## Not Exposed\n\nDashScope APIs also support `negative_prompt`, `prompt_extend`, `watermark`, `thinking_mode`, `seed`, `bbox_list`, `enable_sequential`, and `color_palette`. `baoyu-imagine` does not expose them as CLI flags today; the wan2.7 family relies on the API defaults (e.g. `thinking_mode=true`). The skill always sends `n=1` for wan2.7 — if you want grid/collage mode you currently need to call the API directly.\n\n## Official References\n\n- [Qwen-Image API](https://help.aliyun.com/zh/model-studio/qwen-image-api)\n- [Text-to-image guide](https://help.aliyun.com/zh/model-studio/text-to-image)\n- [Qwen-Image Edit API](https://help.aliyun.com/zh/model-studio/qwen-image-edit-api)\n- [Wan 2.7 image generation & editing API](https://help.aliyun.com/zh/model-studio/wan-image-generation-and-editing-api-reference)\n\nFile v1.117.2:references/providers/minimax.md\n\n# MiniMax\n\nRead when the user picks `--provider minimax` or sets `default_model.minimax`. Default model is `image-01`.\n\n## Models\n\n**`image-01`** (recommended default)\n\n- Supports text-to-image and subject-reference image generation\n- Supports official `aspect_ratio` values: `1:1`, `16:9`, `4:3`, `3:2`, `2:3`, `3:4`, `9:16`, `21:9`\n- Supports documented custom `width` / `height` via `--size <WxH>`\n- Both width and height must be in `[512, 2048]` and divisible by `8`\n\n**`image-01-live`** — lower-latency variant\n\n- Use `--ar` for sizing; MiniMax documents custom `width`/`height` only for `image-01`\n\n## Subject Reference\n\n- `--ref` files are sent as MiniMax `subject_reference`\n- `subject_reference[].type` is currently `character`\n- Official docs say `image_file` supports public URLs or Base64 Data URLs; baoyu-imagine sends local refs as Data URLs\n- Recommended refs: front-facing portraits, JPG/JPEG/PNG, under 10MB\n\n## Official References\n\n- [Image Generation Guide](https://platform.minimaxi.com/docs/guides/image-generation)\n- [Text-to-Image API](https://platform.minimaxi.com/docs/api-reference/image-generation-t2i)\n- [Image-to-Image API](https://platform.minimaxi.com/docs/api-reference/image-generation-i2i)\n\nFile v1.117.2:references/providers/openrouter.md\n\n# OpenRouter\n\nRead when the user picks `--provider openrouter`. Default model is `google/gemini-3.1-flash-image-preview`.\n\n## Common Models\n\nUse full OpenRouter model IDs:\n\n- `google/gemini-3.1-flash-image-preview` (recommended — supports image output and reference-image workflows)\n- `google/gemini-2.5-flash-image-preview`\n- `black-forest-labs/flux.2-pro`\n- Any other OpenRouter image-capable model ID\n\n## Behavior Notes\n\n- OpenRouter image generation uses `/chat/completions`, not the OpenAI `/images` endpoints\n- `--ref` requires a multimodal model that supports both image input and image output\n- `--imageSize` maps to `imageGenerationOptions.size`\n- `--size <WxH>` is converted to the nearest supported OpenRouter size, and the aspect ratio is inferred when possible\n\nFile v1.117.2:references/providers/replicate.md\n\n# Replicate\n\nRead when the user picks `--provider replicate`. Replicate support is intentionally scoped to model families baoyu-imagine can validate locally and save without dropping outputs.\n\n## Supported Families\n\n**`google/nano-banana*`** (default: `google/nano-banana-2`)\n\n- Supports prompt-only and reference-image generation\n- Uses Replicate `aspect_ratio`, `resolution`, and `output_format`\n- `--size <WxH>` is accepted only as a shorthand for a documented `aspect_ratio` plus `1K` / `2K`\n\n**`bytedance/seedream-4.5`**\n\n- Supports prompt-only and reference-image generation\n- Uses Replicate `size`, `aspect_ratio`, and `image_input`\n- Local validation blocks unsupported `1K` requests before the API call\n\n**`bytedance/seedream-5-lite`**\n\n- Supports prompt-only and reference-image generation\n- Uses Replicate `size`, `aspect_ratio`, and `image_input`\n- Local validation currently accepts `2K` / `3K` only\n\n**`wan-video/wan-2.7-image`**\n\n- Supports prompt-only and reference-image generation\n- Uses Replicate `size` and `images`\n- Max output is 2K\n\n**`wan-video/wan-2.7-image-pro`**\n\n- Supports prompt-only and reference-image generation\n- Uses Replicate `size` and `images`\n- 4K is allowed only for text-to-image; local validation blocks `4K + --ref`\n\n## Guardrails\n\n- Replicate currently supports only single-output save semantics in this tool — keep `--n 1`\n- If a model is outside the compatibility list above, baoyu-imagine treats it as prompt-only and rejects advanced local options instead of guessing a nano-banana-style schema\n\n## Examples\n\n```bash\n# Default model\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A cat\" --image out.png --provider replicate\n\n# Explicit model\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A cat\" --image out.png --provider replicate --model google/nano-banana\n```\n\nFile v1.117.2:references/providers/zai.md\n\n# Z.AI GLM-Image\n\nRead when the user picks `--provider zai` or sets `default_model.zai`. Default model is `glm-image`.\n\n## Models\n\n**`glm-image`** (recommended default)\n\n- Text-to-image only in baoyu-imagine (no `--ref` support yet)\n- Native `quality` options are `hd` and `standard`; this skill maps `2k → hd` and `normal → standard`\n- Recommended sizes: `1280x1280`, `1568x1056`, `1056x1568`, `1472x1088`, `1088x1472`, `1728x960`, `960x1728`\n- Custom `--size` requires width/height in `[1024, 2048]`, divisible by `32`, total pixels ≤ `2^22`\n\n**`cogview-4-250304`** (legacy family, same endpoint)\n\n- Custom `--size` requires width/height in `[512, 2048]`, divisible by `16`, total pixels ≤ `2^21`\n\n## Behavior Notes\n\n- The sync API returns a temporary URL; baoyu-imagine downloads it and writes locally\n- `--ref` is not supported for Z.AI in this skill yet\n- The sync API returns a single image, so `--n > 1` is rejected\n\n## Official References\n\n- [GLM-Image Guide](https://docs.z.ai/guides/image/glm-image)\n- [Generate Image API](https://docs.z.ai/api-reference/image/generate-image)\n\nFile v1.117.2:references/usage-examples.md\n\n# Usage Examples\n\nExtended CLI examples. SKILL.md shows the minimum set; read this file when the user asks about provider-specific invocation, batch generation, or less-common flags.\n\n## Core Patterns\n\n```bash\n# Basic text-to-image\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A cat\" --image cat.png\n\n# With aspect ratio\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A landscape\" --image out.png --ar 16:9\n\n# High quality\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A cat\" --image out.png --quality 2k\n\n# Prompt from files\n${BUN_X} {baseDir}/scripts/main.ts --promptfiles system.md content.md --image out.png\n\n# With reference images (any provider family that supports refs)\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"Make blue\" --image out.png --ref source.png\n```\n\n## Per-Provider\n\n```bash\n# OpenAI\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A cat\" --image out.png --provider openai --model gpt-image-2\n\n# Azure OpenAI (model = deployment name)\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A cat\" --image out.png --provider azure --model gpt-image-2\n\n# OpenAI GPT Image 2 custom 4K size\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A cinematic landscape\" --image out.png --provider openai --model gpt-image-2 --size 3840x2160\n\n# Google with explicit model\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"Make blue\" --image out.png --provider google --model gemini-3-pro-image-preview --ref source.png\n\n# OpenRouter (recommended default)\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A cat\" --image out.png --provider openrouter\n\n# OpenRouter with reference\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"Make blue\" --image out.png --provider openrouter --model google/gemini-3.1-flash-image-preview --ref source.png\n\n# DashScope (default model)\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"一只可爱的猫\" --image out.png --provider dashscope\n\n# DashScope Qwen-Image 2.0 Pro (custom size, Chinese text)\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"为咖啡品牌设计一张 21:9 横幅海报，包含清晰中文标题\" --image out.png --provider dashscope --model qwen-image-2.0-pro --size 2048x872\n\n# DashScope legacy fixed-size\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"一张电影感海报\" --image out.png --provider dashscope --model qwen-image-max --size 1664x928\n\n# DashScope Wan 2.7 Image Pro (4K text-to-image)\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"一间有着精致窗户的花店\" --image out.png --provider dashscope --model wan2.7-image-pro --size 4096x4096\n\n# DashScope Wan 2.7 Image with reference image (multi-image fusion)\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"把图2的涂鸦喷绘在图1的汽车上\" --image out.png --provider dashscope --model wan2.7-image-pro --ref car.webp paint.webp\n\n# Z.AI GLM-image\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"一张带清晰中文标题的科技海报\" --image out.png --provider zai\n\n# Z.AI with custom size\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A science illustration with labels\" --image out.png --provider zai --model glm-image --size 1472x1088\n\n# MiniMax\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A fashion editorial portrait\" --image out.jpg --provider minimax\n\n# MiniMax with subject reference (character/portrait consistency)\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A girl by the library window\" --image out.jpg --provider minimax --model image-01 --ref portrait.png --ar 16:9\n\n# Replicate (default: google/nano-banana-2)\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A cat\" --image out.png --provider replicate\n\n# Replicate Seedream 4.5\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A cinematic portrait\" --image out.png --provider replicate --model bytedance/seedream-4.5 --ar 3:2\n\n# Replicate Wan 2.7 Image Pro\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A concept frame\" --image out.png --provider replicate --model wan-video/wan-2.7-image-pro --size 2048x1152\n```\n\n## Batch Mode\n\n```bash\n# Batch from saved prompt files\n${BUN_X} {baseDir}/scripts/main.ts --batchfile batch.json\n\n# Batch with explicit worker count\n${BUN_X} {baseDir}/scripts/main.ts --batchfile batch.json --jobs 4 --json\n```\n\n### Batch File Format\n\n```json\n{\n  \"jobs\": 4,\n  \"tasks\": [\n    {\n      \"id\": \"hero\",\n      \"promptFiles\": [\"prompts/hero.md\"],\n      \"image\": \"out/hero.png\",\n      \"provider\": \"replicate\",\n      \"model\": \"google/nano-banana-2\",\n      \"ar\": \"16:9\",\n      \"quality\": \"2k\"\n    },\n    {\n      \"id\": \"diagram\",\n      \"promptFiles\": [\"prompts/diagram.md\"],\n      \"image\": \"out/diagram.png\",\n      \"ref\": [\"references/original.png\"]\n    }\n  ]\n}\n```\n\nPaths in `promptFiles`, `image`, and `ref` are resolved relative to the batch file's directory. `jobs` is optional (overridden by CLI `--jobs`). A top-level array without the `jobs` wrapper is also accepted.\n\nArchive v1.115.4: 35 files, 91536 bytes\n\nFiles: references/config/first-time-setup.md (13173b), references/config/preferences-schema.md (4224b), references/providers/dashscope.md (4591b), references/providers/minimax.md (1226b), references/providers/openrouter.md (776b), references/providers/replicate.md (1810b), references/providers/zai.md (1095b), references/usage-examples.md (4760b), scripts/build-batch.test.ts (4569b), scripts/build-batch.ts (7454b), scripts/main.test.ts (17852b), scripts/main.ts (43411b), scripts/providers/azure.test.ts (5423b), scripts/providers/azure.ts (5820b), scripts/providers/dashscope.test.ts (11534b), scripts/providers/dashscope.ts (18456b), scripts/providers/google.test.ts (3417b), scripts/providers/google.ts (9873b), scripts/providers/jimeng.test.ts (2737b), scripts/providers/jimeng.ts (12845b), scripts/providers/minimax.test.ts (5254b), scripts/providers/minimax.ts (6331b), scripts/providers/openai.test.ts (5536b), scripts/providers/openai.ts (12671b), scripts/providers/openrouter.test.ts (5450b), scripts/providers/openrouter.ts (9903b), scripts/providers/replicate.test.ts (7190b), scripts/providers/replicate.ts (17816b), scripts/providers/seedream.test.ts (6544b), scripts/providers/seedream.ts (9855b), scripts/providers/zai.test.ts (5274b), scripts/providers/zai.ts (8946b), scripts/types.ts (2123b), SKILL.md (14389b), _meta.json (134b)\n\nFile v1.115.4:SKILL.md\n\n---\nname: baoyu-imagine\ndescription: AI image generation with OpenAI GPT Image 2, Azure OpenAI, Google, OpenRouter, DashScope, Z.AI GLM-Image, MiniMax, Jimeng, Seedream and Replicate APIs. Supports text-to-image, reference images, aspect ratios, and batch generation from saved prompt files. Sequential by default; use batch parallel generation when the user already has multiple prompts or wants stable multi-image throughput. Use when user asks to generate, create, or draw images.\nversion: 1.58.0\nmetadata:\n  openclaw:\n    homepage: https://github.com/JimLiu/baoyu-skills#baoyu-imagine\n    requires:\n      anyBins:\n        - bun\n        - npx\n---\n\n# Image Generation (AI SDK)\n\nOfficial API-based image generation. Supports OpenAI GPT Image 2, Azure OpenAI, Google, OpenRouter, DashScope (阿里通义万象), Z.AI GLM-Image, MiniMax, Jimeng (即梦), Seedream (豆包) and Replicate.\n\n## User Input Tools\n\nWhen this skill prompts the user, follow this tool-selection rule (priority order):\n\n1. **Prefer built-in user-input tools** exposed by the current agent runtime — e.g., `AskUserQuestion`, `request_user_input`, `clarify`, `ask_user`, or any equivalent.\n2. **Fallback**: if no such tool exists, emit a numbered plain-text message and ask the user to reply with the chosen number/answer for each question.\n3. **Batching**: if the tool supports multiple questions per call, combine all applicable questions into a single call; if only single-question, ask them one at a time in priority order.\n\nConcrete `AskUserQuestion` references below are examples — substitute the local equivalent in other runtimes.\n\n## Script Directory\n\n`{baseDir}` = this SKILL.md's directory. Main script: `{baseDir}/scripts/main.ts`. Resolve `${BUN_X}`: prefer `bun`; else `npx -y bun`; else suggest `brew install oven-sh/bun/bun`.\n\n## Step 0: Load Preferences ⛔ BLOCKING\n\nThis step MUST complete before any image generation — generation is blocked until EXTEND.md exists.\n\nCheck these paths in order; first hit wins:\n\n| Path | Scope |\n|------|-------|\n| `.baoyu-skills/baoyu-imagine/EXTEND.md` | Project |\n| `${XDG_CONFIG_HOME:-$HOME/.config}/baoyu-skills/baoyu-imagine/EXTEND.md` | XDG |\n| `$HOME/.baoyu-skills/baoyu-imagine/EXTEND.md` | User home |\n\n- **Found** → load, parse, apply. If `default_model.[provider]` is null → ask model only.\n- **Not found** → run first-time setup (`references/config/first-time-setup.md`) using AskUserQuestion to collect provider + model + quality + save location. Save EXTEND.md, then continue. Do not generate images before this completes.\n\nLegacy compatibility: if `.baoyu-skills/baoyu-image-gen/EXTEND.md` exists and the new path doesn't, the runtime renames it to `baoyu-imagine`. If both exist, the runtime leaves them alone and uses the new path.\n\n**EXTEND.md keys**: default provider, default quality, default aspect ratio, default image size, OpenAI image API dialect, default models, batch worker cap, provider-specific batch limits. Schema: `references/config/preferences-schema.md`.\n\n## Usage\n\nMinimum working examples — see `references/usage-examples.md` for the full set including per-provider invocations and batch mode.\n\n```bash\n# Basic\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A cat\" --image cat.png\n\n# With aspect ratio and high quality\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A landscape\" --image out.png --ar 16:9 --quality 2k\n\n# Prompt from files\n${BUN_X} {baseDir}/scripts/main.ts --promptfiles system.md content.md --image out.png\n\n# With reference image\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"Make blue\" --image out.png --ref source.png\n\n# Specific provider\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A cat\" --image out.png --provider dashscope --model qwen-image-2.0-pro\n\n# OpenAI GPT Image 2\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A cat\" --image out.png --provider openai --model gpt-image-2\n\n# Batch mode\n${BUN_X} {baseDir}/scripts/main.ts --batchfile batch.json --jobs 4\n```\n\n## Options\n\n| Option | Description |\n|--------|-------------|\n| `--prompt <text>`, `-p` | Prompt text |\n| `--promptfiles <files...>` | Read prompt from files (concatenated) |\n| `--image <path>` | Output image path (required in single-image mode) |\n| `--batchfile <path>` | JSON batch file for multi-image generation |\n| `--jobs <count>` | Worker count for batch mode (default: auto, max from config, built-in default 10) |\n| `--provider google\\|openai\\|azure\\|openrouter\\|dashscope\\|zai\\|minimax\\|jimeng\\|seedream\\|replicate` | Force provider (default: auto-detect) |\n| `--model <id>`, `-m` | Model ID — see provider references for defaults and allowed values |\n| `--ar <ratio>` | Aspect ratio (`16:9`, `1:1`, `4:3`, …) |\n| `--size <WxH>` | Explicit size (e.g., `1024x1024`; for `gpt-image-2`, width/height must be multiples of 16, max edge 3840px, ratio no wider than 3:1) |\n| `--quality normal\\|2k` | Quality preset (default: `2k`) |\n| `--imageSize 1K\\|2K\\|4K` | Image size for Google/OpenRouter (default: from quality) |\n| `--imageApiDialect openai-native\\|ratio-metadata` | OpenAI-compatible endpoint dialect — use `ratio-metadata` for gateways that expect aspect-ratio `size` plus `metadata.resolution` |\n| `--ref <files...>` | Reference images. Supported by Google multimodal, OpenAI GPT Image edits, Azure OpenAI edits (PNG/JPG only), OpenRouter multimodal models, Replicate supported families, MiniMax subject-reference, Seedream 5.0/4.5/4.0, DashScope `wan2.7-image-pro`/`wan2.7-image`. Not supported by Jimeng, Seedream 3.0, SeedEdit 3.0, or any DashScope model outside the `wan2.7-image*` family |\n| `--n <count>` | Number of images. Replicate requires `--n 1` (single-output save semantics) |\n| `--json` | JSON output |\n\n## Environment Variables\n\n| Variable | Description |\n|----------|-------------|\n| `OPENAI_API_KEY` | OpenAI API key |\n| `AZURE_OPENAI_API_KEY` | Azure OpenAI API key |\n| `OPENROUTER_API_KEY` | OpenRouter API key |\n| `GOOGLE_API_KEY` | Google API key |\n| `DASHSCOPE_API_KEY` | DashScope API key |\n| `ZAI_API_KEY` (alias `BIGMODEL_API_KEY`) | Z.AI API key |\n| `MINIMAX_API_KEY` | MiniMax API key |\n| `REPLICATE_API_TOKEN` | Replicate API token |\n| `JIMENG_ACCESS_KEY_ID`, `JIMENG_SECRET_ACCESS_KEY` | Jimeng (即梦) Volcengine credentials |\n| `ARK_API_KEY` | Seedream (豆包) Volcengine ARK API key |\n| `<PROVIDER>_IMAGE_MODEL` | Per-provider model override (`OPENAI_IMAGE_MODEL`, `GOOGLE_IMAGE_MODEL`, `DASHSCOPE_IMAGE_MODEL`, `ZAI_IMAGE_MODEL`/`BIGMODEL_IMAGE_MODEL`, `MINIMAX_IMAGE_MODEL`, `OPENROUTER_IMAGE_MODEL`, `REPLICATE_IMAGE_MODEL`, `JIMENG_IMAGE_MODEL`, `SEEDREAM_IMAGE_MODEL`) |\n| `AZURE_OPENAI_DEPLOYMENT` (alias `AZURE_OPENAI_IMAGE_MODEL`) | Azure default deployment |\n| `<PROVIDER>_BASE_URL` | Per-provider endpoint override |\n| `AZURE_API_VERSION` | Azure image API version (default `2025-04-01-preview`) |\n| `JIMENG_REGION` | Jimeng region (default `cn-north-1`) |\n| `OPENAI_IMAGE_API_DIALECT` | `openai-native` \\| `ratio-metadata` |\n| `OPENROUTER_HTTP_REFERER`, `OPENROUTER_TITLE` | Optional OpenRouter attribution |\n| `BAOYU_IMAGE_GEN_MAX_WORKERS` | Override batch worker cap |\n| `BAOYU_IMAGE_GEN_<PROVIDER>_CONCURRENCY` | Per-provider concurrency (e.g., `BAOYU_IMAGE_GEN_REPLICATE_CONCURRENCY`) |\n| `BAOYU_IMAGE_GEN_<PROVIDER>_START_INTERVAL_MS` | Per-provider start-gap |\n\n**Load priority**: CLI args > EXTEND.md > env vars > `<cwd>/.baoyu-skills/.env` > `~/.baoyu-skills/.env`\n\n## Model Resolution\n\nPriority (highest → lowest) applies to every provider:\n\n1. CLI flag `--model <id>`\n2. EXTEND.md `default_model.[provider]`\n3. Env var `<PROVIDER>_IMAGE_MODEL`\n4. Built-in default\n\nFor OpenAI, the built-in default is `gpt-image-2`. `gpt-image-1.5`, `gpt-image-1`, and GPT Image snapshots remain selectable with `--model` or `OPENAI_IMAGE_MODEL`.\n\nFor Azure, `--model` / `default_model.azure` is the Azure deployment name. `AZURE_OPENAI_DEPLOYMENT` is the preferred env var; `AZURE_OPENAI_IMAGE_MODEL` is kept as a backward-compatible alias. If your Azure deployment is named after the underlying model, use `gpt-image-2`; otherwise use the exact custom deployment name.\n\nEXTEND.md overrides env vars: if EXTEND.md sets `default_model.google: \"gemini-3-pro-image-preview\"` and the env var sets `GOOGLE_IMAGE_MODEL=gemini-3.1-flash-image-preview`, EXTEND.md wins.\n\n**Display model info before each generation**:\n\n- `Using [provider] / [model]`\n- `Switch model: --model <id> | EXTEND.md default_model.[provider] | env <PROVIDER>_IMAGE_MODEL`\n\n## OpenAI-Compatible Gateway Dialects\n\n`provider=openai` means the auth and routing entrypoint is OpenAI-compatible. It does **not** guarantee the upstream image API uses OpenAI native semantics. When a gateway expects a different wire format, set `default_image_api_dialect` in EXTEND.md, `OPENAI_IMAGE_API_DIALECT`, or `--imageApiDialect`:\n\n- `openai-native`: pixel `size` (`1536x1024`) and native OpenAI quality fields\n- `ratio-metadata`: aspect-ratio `size` (`16:9`) plus `metadata.resolution` (`1K|2K|4K`) and `metadata.orientation`\n\nUse `openai-native` for the OpenAI native API or strict clones; try `ratio-metadata` for compatibility gateways in front of Gemini or similar models. Current limitation: `ratio-metadata` applies only to text-to-image; reference-image edits still need `openai-native` or a provider with first-class edit support.\n\n## Provider-Specific Guides\n\nEach provider has its own quirks (model families, size rules, ref support, limits). Read these when the user picks that provider or asks for non-default behavior:\n\n| Provider | Reference |\n|----------|-----------|\n| DashScope (Qwen-Image families, custom sizes) | `references/providers/dashscope.md` |\n| Z.AI (GLM-Image, cogview-4) | `references/providers/zai.md` |\n| MiniMax (image-01, subject-reference) | `references/providers/minimax.md` |\n| OpenRouter (multimodal models, `/chat/completions` flow) | `references/providers/openrouter.md` |\n| Replicate (nano-banana, Seedream, Wan) | `references/providers/replicate.md` |\n\n## Provider Selection\n\n1. `--ref` provided + no `--provider` → auto-select Google → OpenAI → Azure → OpenRouter → Replicate → Seedream → MiniMax (MiniMax's subject reference is more specialized toward character/portrait consistency)\n2. `--provider` specified → use it (if `--ref`, must be google/openai/azure/openrouter/replicate/seedream/minimax)\n3. Only one API key present → use that provider\n4. Multiple keys → default priority: Google → OpenAI → Azure → OpenRouter → DashScope → Z.AI → MiniMax → Replicate → Jimeng → Seedream\n\n## Quality Presets\n\n| Preset | Google imageSize | OpenAI size | OpenRouter size | Replicate resolution | Use case |\n|--------|------------------|-------------|-----------------|----------------------|----------|\n| `normal` | 1K | 1024px target | 1K | 1K | Quick previews |\n| `2k` (default) | 2K | 2048px target | 2K | 2K | Covers, illustrations, infographics |\n\nGoogle/OpenRouter `imageSize` can be overridden with `--imageSize 1K|2K|4K`.\n\nFor OpenAI native `gpt-image-2`, `normal` maps to `quality=medium` and a low-latency valid size near the requested aspect ratio; `2k` maps to `quality=high` and 2048px-class sizes such as `2048x2048`, `2048x1152`, or `1152x2048`. Use explicit `--size` for valid custom or 4K outputs, e.g. `3840x2160`.\n\n## Aspect Ratios\n\nSupported: `1:1`, `16:9`, `9:16`, `4:3`, `3:4`, `2.35:1`.\n\n- Google multimodal: `imageConfig.aspectRatio`\n- OpenAI: `gpt-image-2` uses the closest valid custom size for the requested ratio; older GPT Image and DALL·E models use their closest supported fixed size\n- OpenRouter: `imageGenerationOptions.aspect_ratio`; if only `--size <WxH>` is given, the ratio is inferred\n- Replicate: behavior is model-specific — `google/nano-banana*` uses `aspect_ratio`, `bytedance/seedream-*` uses documented Replicate ratios, Wan 2.7 maps `--ar` to a concrete `size`\n- MiniMax: official `aspect_ratio` values; if `--size <WxH>` is given without `--ar`, sends `width`/`height` for `image-01`\n\n## Generation Mode\n\n**Default**: sequential. **Batch parallel**: enabled automatically when `--batchfile` contains 2+ pending tasks.\n\n| Situation | Prefer | Why |\n|-----------|--------|-----|\n| One image, or 1-2 simple images | Sequential | Lower coordination overhead, easier debugging |\n| Multiple images with saved prompt files | Batch (`--batchfile`) | Reuses finalized prompts, applies shared throttling/retries, predictable throughput |\n| Each image still needs its own reasoning / prompt writing / style exploration | Subagents | Work is still exploratory, each needs independent analysis |\n| Input is `outline.md` + `prompts/` (e.g. from `baoyu-article-illustrator`) | Batch — use `scripts/build-batch.ts` to assemble the payload | The outline + prompt files already contain everything needed |\n\nRule of thumb: once prompt files are saved and the task is \"generate all of these\", prefer batch over subagents. Use subagents only when generation is coupled with per-image thinking or divergent creative exploration.\n\n**Parallel behavior**:\n\n- Default worker count is automatic, capped by config, built-in default 10\n- Provider-specific throttling applies only in batch mode; defaults are tuned for throughput while avoiding RPM bursts\n- Override with `--jobs <count>`\n- Each image retries up to 3 attempts\n- Final output includes success count, failure count, and per-image failure reasons\n\n## Error Handling\n\n- Missing API key → error with setup instructions\n- Generation failure → auto-retry up to 3 attempts per image\n- Invalid aspect ratio → warning, proceed with default\n- Reference images with unsupported provider/model → error with fix hint\n\n## References\n\n| File | Content |\n|------|---------|\n| `references/usage-examples.md` | Extended CLI examples across providers and batch mode |\n| `references/providers/dashscope.md` | DashScope families, sizes, limits |\n| `references/providers/zai.md` | Z.AI GLM-image / cogview-4 |\n| `references/providers/minimax.md` | MiniMax image-01 + subject reference |\n| `references/providers/openrouter.md` | OpenRouter multimodal flow |\n| `references/providers/replicate.md` | Replicate supported families + guardrails |\n| `references/config/preferences-schema.md` | EXTEND.md schema |\n| `references/config/first-time-setup.md` | First-time setup flow |\n\n## Extension Support\n\nCustom configurations via EXTEND.md. See Step 0 for paths and schema.\n\nFile v1.115.4:_meta.json\n\n{\n  \"ownerId\": \"kn7csrrndw79hpke5d0gsnx93d82k67r\",\n  \"slug\": \"baoyu-imagine\",\n  \"version\": \"1.115.4\",\n  \"publishedAt\": 1778543343373\n}\n\nFile v1.115.4:references/config/first-time-setup.md\n\n---\nname: first-time-setup\ndescription: First-time setup and default model selection flow for baoyu-imagine\n---\n\n# First-Time Setup\n\n## Overview\n\nTriggered when:\n1. No EXTEND.md found → full setup (provider + model + preferences)\n2. EXTEND.md found but `default_model.[provider]` is null → model selection only\n\n## Setup Flow\n\n```\nNo EXTEND.md found          EXTEND.md found, model null\n        │                            │\n        ▼                            ▼\n┌─────────────────────┐    ┌──────────────────────┐\n│ AskUserQuestion     │    │ AskUserQuestion      │\n│ (full setup)        │    │ (model only)         │\n└─────────────────────┘    └──────────────────────┘\n        │                            │\n        ▼                            ▼\n┌─────────────────────┐    ┌──────────────────────┐\n│ Create EXTEND.md    │    │ Update EXTEND.md     │\n└─────────────────────┘    └──────────────────────┘\n        │                            │\n        ▼                            ▼\n    Continue                     Continue\n```\n\n## Flow 1: No EXTEND.md (Full Setup)\n\n**Language**: Use user's input language or saved language preference.\n\nUse AskUserQuestion with ALL questions in ONE call:\n\n### Question 1: Default Provider\n\n```yaml\nheader: \"Provider\"\nquestion: \"Default image generation provider?\"\noptions:\n  - label: \"Google (Recommended)\"\n    description: \"Gemini multimodal - high quality, reference images, flexible sizes\"\n  - label: \"OpenAI\"\n    description: \"GPT Image 2 - latest OpenAI image model, reference-image workflows\"\n  - label: \"Azure OpenAI\"\n    description: \"Azure-hosted GPT Image deployments with resource-specific routing\"\n  - label: \"OpenRouter\"\n    description: \"Router for Gemini/FLUX/OpenAI-compatible image models\"\n  - label: \"DashScope\"\n    description: \"Alibaba Cloud - Qwen-Image, strong Chinese/English text rendering\"\n  - label: \"Z.AI\"\n    description: \"GLM-image, strong poster and text-heavy image generation\"\n  - label: \"MiniMax\"\n    description: \"MiniMax image generation with subject-reference character workflows\"\n  - label: \"Replicate\"\n    description: \"Curated Replicate image families - nano-banana-2, Seedream, and Wan image models\"\n```\n\n### Question 2: Default Google Model\n\nOnly show if user selected Google or auto-detect (no explicit provider).\n\n```yaml\nheader: \"Google Model\"\nquestion: \"Default Google image generation model?\"\noptions:\n  - label: \"gemini-3-pro-image-preview (Recommended)\"\n    description: \"Highest quality, best for production use\"\n  - label: \"gemini-3.1-flash-image-preview\"\n    description: \"Fast generation, good quality, lower cost\"\n  - label: \"gemini-3-flash-preview\"\n    description: \"Fast generation, balanced quality and speed\"\n```\n\n### Question 2b: Default OpenRouter Model\n\nOnly show if user selected OpenRouter.\n\n```yaml\nheader: \"OpenRouter Model\"\nquestion: \"Default OpenRouter image generation model?\"\noptions:\n  - label: \"google/gemini-3.1-flash-image-preview (Recommended)\"\n    description: \"Best general-purpose OpenRouter image model with reference-image workflows\"\n  - label: \"google/gemini-2.5-flash-image-preview\"\n    description: \"Fast Gemini preview model on OpenRouter\"\n  - label: \"black-forest-labs/flux.2-pro\"\n    description: \"Strong text-to-image quality through OpenRouter\"\n```\n\n### Question 2c: Default Azure Deployment\n\nOnly show if user selected Azure OpenAI.\n\n```yaml\nheader: \"Azure Deploy\"\nquestion: \"Default Azure image deployment name?\"\noptions:\n  - label: \"gpt-image-2 (Recommended)\"\n    description: \"Use if your Azure deployment uses the GPT Image 2 model name\"\n  - label: \"gpt-image-1.5\"\n    description: \"Previous GPT Image deployment name\"\n  - label: \"gpt-image-1\"\n    description: \"Earlier GPT Image deployment name\"\n```\n\n### Question 2d: Default MiniMax Model\n\nOnly show if user selected MiniMax.\n\n```yaml\nheader: \"MiniMax Model\"\nquestion: \"Default MiniMax image generation model?\"\noptions:\n  - label: \"image-01 (Recommended)\"\n    description: \"Best default, supports aspect ratios and custom width/height\"\n  - label: \"image-01-live\"\n    description: \"Faster variant, use aspect ratio instead of custom size\"\n```\n\n### Question 2e: Default Z.AI Model\n\nOnly show if user selected Z.AI.\n\n```yaml\nheader: \"Z.AI Model\"\nquestion: \"Default Z.AI image generation model?\"\noptions:\n  - label: \"glm-image (Recommended)\"\n    description: \"Best default for posters, diagrams, and text-heavy images\"\n  - label: \"cogview-4-250304\"\n    description: \"Legacy Z.AI image model on the same endpoint\"\n```\n\n### Question 3: Default Quality\n\n```yaml\nheader: \"Quality\"\nquestion: \"Default image quality?\"\noptions:\n  - label: \"2k (Recommended)\"\n    description: \"2048px - covers, illustrations, infographics\"\n  - label: \"normal\"\n    description: \"1024px - quick previews, drafts\"\n```\n\n### Question 4: Save Location\n\n```yaml\nheader: \"Save\"\nquestion: \"Where to save preferences?\"\noptions:\n  - label: \"Project (Recommended)\"\n    description: \".baoyu-skills/ (this project only)\"\n  - label: \"User\"\n    description: \"~/.baoyu-skills/ (all projects)\"\n```\n\n### Save Locations\n\n| Choice | Path | Scope |\n|--------|------|-------|\n| Project | `.baoyu-skills/baoyu-imagine/EXTEND.md` | Current project |\n| User | `$HOME/.baoyu-skills/baoyu-imagine/EXTEND.md` | All projects |\n\n### EXTEND.md Template\n\n```yaml\n---\nversion: 1\ndefault_provider: [selected provider or null]\ndefault_quality: [selected quality]\ndefault_aspect_ratio: null\ndefault_image_size: null\ndefault_image_api_dialect: null\ndefault_model:\n  google: [selected google model or null]\n  openai: null\n  azure: [selected azure deployment or null]\n  openrouter: [selected openrouter model or null]\n  dashscope: null\n  zai: [selected Z.AI model or null]\n  minimax: [selected minimax model or null]\n  replicate: null\n---\n```\n\nIf the user selects `OpenAI` but says their endpoint is only OpenAI-compatible and fronts another image model family, save `default_image_api_dialect: ratio-metadata` when they explicitly confirm the gateway expects aspect-ratio `size` plus metadata-based resolution. Otherwise leave it `null` / `openai-native`.\n\n## Flow 2: EXTEND.md Exists, Model Null\n\nWhen EXTEND.md exists but `default_model.[current_provider]` is null, ask ONLY the model question for the current provider.\n\n### Google Model Selection\n\n```yaml\nheader: \"Google Model\"\nquestion: \"Choose a default Google image generation model?\"\noptions:\n  - label: \"gemini-3-pro-image-preview (Recommended)\"\n    description: \"Highest quality, best for production use\"\n  - label: \"gemini-3.1-flash-image-preview\"\n    description: \"Fast generation, good quality, lower cost\"\n  - label: \"gemini-3-flash-preview\"\n    description: \"Fast generation, balanced quality and speed\"\n```\n\n### OpenAI Model Selection\n\n```yaml\nheader: \"OpenAI Model\"\nquestion: \"Choose a default OpenAI image generation model?\"\noptions:\n  - label: \"gpt-image-2 (Recommended)\"\n    description: \"Latest GPT Image model, flexible sizes up to 4K, high-fidelity image inputs\"\n  - label: \"gpt-image-1.5\"\n    description: \"Previous GPT Image model\"\n  - label: \"gpt-image-1\"\n    description: \"Earlier GPT Image model\"\n```\n\n### Azure Deployment Selection\n\n```yaml\nheader: \"Azure Deploy\"\nquestion: \"Choose a default Azure image deployment name?\"\noptions:\n  - label: \"gpt-image-2 (Recommended)\"\n    description: \"Use when your Azure deployment name matches the GPT Image 2 model\"\n  - label: \"gpt-image-1.5\"\n    description: \"Use when your Azure deployment name matches the GPT Image 1.5 model\"\n  - label: \"gpt-image-1\"\n    description: \"Use when your Azure deployment name matches GPT-image-1\"\n```\n\nNotes for Azure setup:\n\n- In `baoyu-imagine`, Azure `--model` / `default_model.azure` should be the Azure deployment name, not just the underlying model family.\n- If the deployment name is custom, save that exact deployment name in `default_model.azure`.\n\n### OpenRouter Model Selection\n\n```yaml\nheader: \"OpenRouter Model\"\nquestion: \"Choose a default OpenRouter image generation model?\"\noptions:\n  - label: \"google/gemini-3.1-flash-image-preview (Recommended)\"\n    description: \"Recommended for image output and reference-image edits\"\n  - label: \"google/gemini-2.5-flash-image-preview\"\n    description: \"Fast preview-oriented image generation\"\n  - label: \"black-forest-labs/flux.2-pro\"\n    description: \"High-quality text-to-image through OpenRouter\"\n```\n\n### DashScope Model Selection\n\n```yaml\nheader: \"DashScope Model\"\nquestion: \"Choose a default DashScope image generation model?\"\noptions:\n  - label: \"qwen-image-2.0-pro (Recommended)\"\n    description: \"Best DashScope model for text rendering and custom sizes\"\n  - label: \"qwen-image-2.0\"\n    description: \"Faster 2.0 variant with flexible output size\"\n  - label: \"qwen-image-max\"\n    description: \"Legacy Qwen model with five fixed output sizes\"\n  - label: \"qwen-image-plus\"\n    description: \"Legacy Qwen model, same current capability as qwen-image\"\n  - label: \"wan2.7-image-pro\"\n    description: \"Wan 2.7 Pro — supports up to 4K text-to-image and reference-image editing\"\n  - label: \"wan2.7-image\"\n    description: \"Wan 2.7 base — faster generation, up to 2K, supports reference-image editing\"\n  - label: \"z-image-turbo\"\n    description: \"Legacy DashScope model for compatibility\"\n  - label: \"z-image-ultra\"\n    description: \"Legacy DashScope model, higher quality but slower\"\n```\n\nNotes for DashScope setup:\n\n- Prefer `qwen-image-2.0-pro` when the user needs custom `--size`, uncommon ratios like `21:9`, or strong Chinese/English text rendering.\n- `qwen-image-max` / `qwen-image-plus` / `qwen-image` only support five fixed sizes: `1664*928`, `1472*1104`, `1328*1328`, `1104*1472`, `928*1664`.\n- `wan2.7-image-pro` and `wan2.7-image` are the only DashScope models that accept `--ref`. Pick one of these when the user wants reference-image editing or multi-image fusion via DashScope.\n- In `baoyu-imagine`, `quality` is a compatibility preset. It is not a native DashScope parameter.\n\n### Z.AI Model Selection\n\n```yaml\nheader: \"Z.AI Model\"\nquestion: \"Choose a default Z.AI image generation model?\"\noptions:\n  - label: \"glm-image (Recommended)\"\n    description: \"Current flagship image model with better text rendering and poster layouts\"\n  - label: \"cogview-4-250304\"\n    description: \"Legacy model on the sync image endpoint\"\n```\n\nNotes for Z.AI setup:\n\n- Prefer `glm-image` for posters, diagrams, and Chinese/English text-heavy layouts.\n- In `baoyu-imagine`, Z.AI currently exposes text-to-image only; reference images are not wired for this provider.\n- The sync Z.AI image API returns a downloadable image URL, which the runtime saves locally after download.\n\n### Replicate Model Selection\n\n```yaml\nheader: \"Replicate Model\"\nquestion: \"Choose a default Replicate image generation model?\"\noptions:\n  - label: \"google/nano-banana-2 (Recommended)\"\n    description: \"Current default for general Replicate image generation in baoyu-imagine\"\n  - label: \"bytedance/seedream-4.5\"\n    description: \"Replicate Seedream 4.5 with validated local size/ref guardrails\"\n  - label: \"bytedance/seedream-5-lite\"\n    description: \"Replicate Seedream 5 Lite with validated local size/ref guardrails\"\n  - label: \"wan-video/wan-2.7-image-pro\"\n    description: \"Replicate Wan 2.7 Image Pro with 4K text-to-image support\"\n```\n\n### MiniMax Model Selection\n\n```yaml\nheader: \"MiniMax Model\"\nquestion: \"Choose a default MiniMax image generation model?\"\noptions:\n  - label: \"image-01 (Recommended)\"\n    description: \"Best general-purpose MiniMax image model with custom width/height support\"\n  - label: \"image-01-live\"\n    description: \"Lower-latency MiniMax image model using aspect ratios\"\n```\n\nNotes for MiniMax setup:\n\n- `image-01` is the safest default. It supports official `aspect_ratio` values and documented custom `width` / `height` output sizes.\n- `image-01-live` is useful when the user prefers faster generation and can work with aspect-ratio-based sizing.\n- MiniMax subject reference currently uses `subject_reference[].type = character`; docs recommend front-facing portrait references in JPG/JPEG/PNG under 10MB.\n\n### Update EXTEND.md\n\nAfter user selects a model:\n\n1. Read existing EXTEND.md\n2. If `default_model:` section exists → update the provider-specific key\n3. If `default_model:` section missing → add the full section:\n\n```yaml\ndefault_model:\n  google: [value or null]\n  openai: [value or null]\n  azure: [value or null]\n  openrouter: [value or null]\n  dashscope: [value or null]\n  zai: [value or null]\n  minimax: [value or null]\n  replicate: [value or null]\n```\n\nOnly set the selected provider's model; leave others as their current value or null.\n\n## After Setup\n\n1. Create directory if needed\n2. Write/update EXTEND.md with frontmatter\n3. Confirm: \"Preferences saved to [path]\"\n4. Continue with image generation\n\nFile v1.115.4:references/config/preferences-schema.md\n\n---\nname: preferences-schema\ndescription: EXTEND.md YAML schema for baoyu-imagine user preferences\n---\n\n# Preferences Schema\n\n## Full Schema\n\n```yaml\n---\nversion: 1\n\ndefault_provider: null      # google|openai|azure|openrouter|dashscope|zai|minimax|replicate|null (null = auto-detect)\n\ndefault_quality: null       # normal|2k|null (null = use default: 2k)\n\ndefault_aspect_ratio: null  # \"16:9\"|\"1:1\"|\"4:3\"|\"3:4\"|\"2.35:1\"|null\n\ndefault_image_size: null    # 1K|2K|4K|null (Google/OpenRouter, overrides quality)\n\ndefault_image_api_dialect: null  # openai-native|ratio-metadata|null (OpenAI-compatible gateways; null = use env/default)\n\ndefault_model:\n  google: null              # e.g., \"gemini-3-pro-image-preview\", \"gemini-3.1-flash-image-preview\"\n  openai: null              # e.g., \"gpt-image-2\", \"gpt-image-1.5\", \"gpt-image-1\"\n  azure: null               # Azure deployment name, e.g., \"gpt-image-2\" or \"image-prod\"\n  openrouter: null          # e.g., \"google/gemini-3.1-flash-image-preview\"\n  dashscope: null           # e.g., \"qwen-image-2.0-pro\"\n  zai: null                 # e.g., \"glm-image\"\n  minimax: null             # e.g., \"image-01\"\n  replicate: null           # e.g., \"google/nano-banana-2\"\n\nbatch:\n  max_workers: 10\n  provider_limits:\n    replicate:\n      concurrency: 5\n      start_interval_ms: 700\n    google:\n      concurrency: 3\n      start_interval_ms: 1100\n    openai:\n      concurrency: 3\n      start_interval_ms: 1100\n    azure:\n      concurrency: 3\n      start_interval_ms: 1100\n    openrouter:\n      concurrency: 3\n      start_interval_ms: 1100\n    dashscope:\n      concurrency: 3\n      start_interval_ms: 1100\n    zai:\n      concurrency: 3\n      start_interval_ms: 1100\n    minimax:\n      concurrency: 3\n      start_interval_ms: 1100\n---\n```\n\n## Field Reference\n\n| Field | Type | Default | Description |\n|-------|------|---------|-------------|\n| `version` | int | 1 | Schema version |\n| `default_provider` | string\\|null | null | Default provider (null = auto-detect) |\n| `default_quality` | string\\|null | null | Default quality (null = 2k) |\n| `default_aspect_ratio` | string\\|null | null | Default aspect ratio |\n| `default_image_size` | string\\|null | null | Google/OpenRouter image size (overrides quality) |\n| `default_image_api_dialect` | string\\|null | null | OpenAI-compatible image dialect (`openai-native` or `ratio-metadata`) |\n| `default_model.google` | string\\|null | null | Google default model |\n| `default_model.openai` | string\\|null | null | OpenAI default model |\n| `default_model.azure` | string\\|null | null | Azure default deployment name |\n| `default_model.openrouter` | string\\|null | null | OpenRouter default model |\n| `default_model.dashscope` | string\\|null | null | DashScope default model |\n| `default_model.zai` | string\\|null | null | Z.AI default model |\n| `default_model.minimax` | string\\|null | null | MiniMax default model |\n| `default_model.replicate` | string\\|null | null | Replicate default model |\n| `batch.max_workers` | int\\|null | 10 | Batch worker cap |\n| `batch.provider_limits.<provider>.concurrency` | int\\|null | provider default | Max simultaneous requests per provider |\n| `batch.provider_limits.<provider>.start_interval_ms` | int\\|null | provider default | Minimum gap between request starts per provider |\n\n## Examples\n\n**Minimal**:\n```yaml\n---\nversion: 1\ndefault_provider: google\ndefault_quality: 2k\ndefault_image_api_dialect: null\n---\n```\n\n**Full**:\n```yaml\n---\nversion: 1\ndefault_provider: google\ndefault_quality: 2k\ndefault_aspect_ratio: \"16:9\"\ndefault_image_size: 2K\ndefault_image_api_dialect: null\ndefault_model:\n  google: \"gemini-3-pro-image-preview\"\n  openai: \"gpt-image-2\"\n  azure: \"gpt-image-2\"\n  openrouter: \"google/gemini-3.1-flash-image-preview\"\n  dashscope: \"qwen-image-2.0-pro\"\n  zai: \"glm-image\"\n  minimax: \"image-01\"\n  replicate: \"google/nano-banana-2\"\nbatch:\n  max_workers: 10\n  provider_limits:\n    replicate:\n      concurrency: 5\n      start_interval_ms: 700\n    azure:\n      concurrency: 3\n      start_interval_ms: 1100\n    zai:\n      concurrency: 3\n      start_interval_ms: 1100\n    openrouter:\n      concurrency: 3\n      start_interval_ms: 1100\n    minimax:\n      concurrency: 3\n      start_interval_ms: 1100\n---\n```\n\nFile v1.115.4:references/providers/dashscope.md\n\n# DashScope (阿里通义万象)\n\nRead when the user picks `--provider dashscope`, sets `default_model.dashscope`, or asks for Qwen-Image behavior. The SKILL.md only names the default — this file covers model families, sizing rules, and limits.\n\n## Model Families\n\n**`qwen-image-2.0*`** — recommended modern family. Members: `qwen-image-2.0-pro`, `qwen-image-2.0-pro-2026-03-03`, `qwen-image-2.0`, `qwen-image-2.0-2026-03-03`.\n\n- Free-form `size` in `宽*高` format\n- Total pixels must be between `512*512` and `2048*2048`\n- Default ≈ `1024*1024`\n- Best choice for custom ratios (e.g. `21:9`) and text-heavy Chinese/English layouts\n\n**Fixed-size family** — `qwen-image-max`, `qwen-image-max-2025-12-30`, `qwen-image-plus`, `qwen-image-plus-2026-01-09`, `qwen-image`.\n\n- Only five sizes allowed: `1664*928`, `1472*1104`, `1328*1328`, `1104*1472`, `928*1664`\n- Default is `1664*928`\n- `qwen-image` currently has the same capability as `qwen-image-plus`\n\n**`wan2.7-image*`** — multimodal Wan 2.7 family. Members: `wan2.7-image-pro`, `wan2.7-image`.\n\n- Free-form `size` in `宽*高` format, plus aspect-ratio inference\n- `wan2.7-image-pro` text-to-image (no `--ref`): total pixels in `[768*768, 4096*4096]`, ratio in `[1:8, 8:1]`\n- `wan2.7-image-pro` with reference images and `wan2.7-image` (all scenarios): total pixels in `[768*768, 2048*2048]`, ratio in `[1:8, 8:1]`\n- Default: `1024*1024` (`--quality normal`) or `2048*2048` (`--quality 2k`); 4K requires explicit `--size`\n- Supports up to 9 reference images in `--ref` (image editing / multi-image fusion)\n- Reference images are sent inline as base64 (or passed through if the path is an `http(s)://` URL)\n- API does NOT use `prompt_extend`; the skill omits it for this family\n- The Wan 2.7 API defaults `n` to **4** in non-collage mode and bills per generated image. baoyu-imagine forces `n: 1` and rejects `--n > 1` to avoid silently paying for and discarding extra images.\n\n**Legacy** — `z-image-turbo`, `z-image-ultra`, `wanx-v1`. Only use when the user explicitly asks for legacy behavior.\n\n## Size Resolution\n\n- `--size` wins over `--ar`\n- For `qwen-image-2.0*`: prefer explicit `--size`; otherwise infer from `--ar` using the recommended table below\n- For `qwen-image-max/plus/image`: only use the five fixed sizes; if the requested ratio doesn't fit, switch to `qwen-image-2.0-pro`\n- For `wan2.7-image*`: explicit `--size` is validated against the per-mode pixel/ratio limits; otherwise the size is derived from `--ar` and `--quality` (`normal` ≈ 1K, `2k` ≈ 2K). To request 4K with `wan2.7-image-pro` text-to-image, pass `--size` explicitly (e.g. `4096*4096`, `3840*2160`)\n- `--quality` is a baoyu-imagine preset, not an official DashScope field. The mapping of `normal`/`2k` onto the `qwen-image-2.0*` and `wan2.7-image*` tables is an implementation choice, not an API guarantee\n\n### Recommended `qwen-image-2.0*` sizes\n\n| Ratio | `normal` | `2k` |\n|-------|----------|------|\n| `1:1` | `1024*1024` | `1536*1536` |\n| `2:3` | `768*1152` | `1024*1536` |\n| `3:2` | `1152*768` | `1536*1024` |\n| `3:4` | `960*1280` | `1080*1440` |\n| `4:3` | `1280*960` | `1440*1080` |\n| `9:16` | `720*1280` | `1080*1920` |\n| `16:9` | `1280*720` | `1920*1080` |\n| `21:9` | `1344*576` | `2048*872` |\n\n## Reference Images\n\n- Only `wan2.7-image-pro` and `wan2.7-image` accept `--ref`. Other DashScope models (qwen-image-2.0*, qwen-image-max/plus/image, legacy) reject `--ref` and the user is steered to a different provider/model.\n- Up to 9 reference images per request. Local files are inlined as base64 data URLs; `http(s)://` URLs are forwarded as-is.\n- Supplying any `--ref` automatically clamps the wan2.7-image-pro pixel ceiling from 4K to 2K (the API only supports 4K for pure text-to-image with no image input).\n\n## Not Exposed\n\nDashScope APIs also support `negative_prompt`, `prompt_extend`, `watermark`, `thinking_mode`, `seed`, `bbox_list`, `enable_sequential`, and `color_palette`. `baoyu-imagine` does not expose them as CLI flags today; the wan2.7 family relies on the API defaults (e.g. `thinking_mode=true`). The skill always sends `n=1` for wan2.7 — if you want grid/collage mode you currently need to call the API directly.\n\n## Official References\n\n- [Qwen-Image API](https://help.aliyun.com/zh/model-studio/qwen-image-api)\n- [Text-to-image guide](https://help.aliyun.com/zh/model-studio/text-to-image)\n- [Qwen-Image Edit API](https://help.aliyun.com/zh/model-studio/qwen-image-edit-api)\n- [Wan 2.7 image generation & editing API](https://help.aliyun.com/zh/model-studio/wan-image-generation-and-editing-api-reference)\n\nFile v1.115.4:references/providers/minimax.md\n\n# MiniMax\n\nRead when the user picks `--provider minimax` or sets `default_model.minimax`. Default model is `image-01`.\n\n## Models\n\n**`image-01`** (recommended default)\n\n- Supports text-to-image and subject-reference image generation\n- Supports official `aspect_ratio` values: `1:1`, `16:9`, `4:3`, `3:2`, `2:3`, `3:4`, `9:16`, `21:9`\n- Supports documented custom `width` / `height` via `--size <WxH>`\n- Both width and height must be in `[512, 2048]` and divisible by `8`\n\n**`image-01-live`** — lower-latency variant\n\n- Use `--ar` for sizing; MiniMax documents custom `width`/`height` only for `image-01`\n\n## Subject Reference\n\n- `--ref` files are sent as MiniMax `subject_reference`\n- `subject_reference[].type` is currently `character`\n- Official docs say `image_file` supports public URLs or Base64 Data URLs; baoyu-imagine sends local refs as Data URLs\n- Recommended refs: front-facing portraits, JPG/JPEG/PNG, under 10MB\n\n## Official References\n\n- [Image Generation Guide](https://platform.minimaxi.com/docs/guides/image-generation)\n- [Text-to-Image API](https://platform.minimaxi.com/docs/api-reference/image-generation-t2i)\n- [Image-to-Image API](https://platform.minimaxi.com/docs/api-reference/image-generation-i2i)\n\nFile v1.115.4:references/providers/openrouter.md\n\n# OpenRouter\n\nRead when the user picks `--provider openrouter`. Default model is `google/gemini-3.1-flash-image-preview`.\n\n## Common Models\n\nUse full OpenRouter model IDs:\n\n- `google/gemini-3.1-flash-image-preview` (recommended — supports image output and reference-image workflows)\n- `google/gemini-2.5-flash-image-preview`\n- `black-forest-labs/flux.2-pro`\n- Any other OpenRouter image-capable model ID\n\n## Behavior Notes\n\n- OpenRouter image generation uses `/chat/completions`, not the OpenAI `/images` endpoints\n- `--ref` requires a multimodal model that supports both image input and image output\n- `--imageSize` maps to `imageGenerationOptions.size`\n- `--size <WxH>` is converted to the nearest supported OpenRouter size, and the aspect ratio is inferred when possible\n\nFile v1.115.4:references/providers/replicate.md\n\n# Replicate\n\nRead when the user picks `--provider replicate`. Replicate support is intentionally scoped to model families baoyu-imagine can validate locally and save without dropping outputs.\n\n## Supported Families\n\n**`google/nano-banana*`** (default: `google/nano-banana-2`)\n\n- Supports prompt-only and reference-image generation\n- Uses Replicate `aspect_ratio`, `resolution`, and `output_format`\n- `--size <WxH>` is accepted only as a shorthand for a documented `aspect_ratio` plus `1K` / `2K`\n\n**`bytedance/seedream-4.5`**\n\n- Supports prompt-only and reference-image generation\n- Uses Replicate `size`, `aspect_ratio`, and `image_input`\n- Local validation blocks unsupported `1K` requests before the API call\n\n**`bytedance/seedream-5-lite`**\n\n- Supports prompt-only and reference-image generation\n- Uses Replicate `size`, `aspect_ratio`, and `image_input`\n- Local validation currently accepts `2K` / `3K` only\n\n**`wan-video/wan-2.7-image`**\n\n- Supports prompt-only and reference-image generation\n- Uses Replicate `size` and `images`\n- Max output is 2K\n\n**`wan-video/wan-2.7-image-pro`**\n\n- Supports prompt-only and reference-image generation\n- Uses Replicate `size` and `images`\n- 4K is allowed only for text-to-image; local validation blocks `4K + --ref`\n\n## Guardrails\n\n- Replicate currently supports only single-output save semantics in this tool — keep `--n 1`\n- If a model is outside the compatibility list above, baoyu-imagine treats it as prompt-only and rejects advanced local options instead of guessing a nano-banana-style schema\n\n## Examples\n\n```bash\n# Default model\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A cat\" --image out.png --provider replicate\n\n# Explicit model\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A cat\" --image out.png --provider replicate --model google/nano-banana\n```\n\nFile v1.115.4:references/providers/zai.md\n\n# Z.AI GLM-Image\n\nRead when the user picks `--provider zai` or sets `default_model.zai`. Default model is `glm-image`.\n\n## Models\n\n**`glm-image`** (recommended default)\n\n- Text-to-image only in baoyu-imagine (no `--ref` support yet)\n- Native `quality` options are `hd` and `standard`; this skill maps `2k → hd` and `normal → standard`\n- Recommended sizes: `1280x1280`, `1568x1056`, `1056x1568`, `1472x1088`, `1088x1472`, `1728x960`, `960x1728`\n- Custom `--size` requires width/height in `[1024, 2048]`, divisible by `32`, total pixels ≤ `2^22`\n\n**`cogview-4-250304`** (legacy family, same endpoint)\n\n- Custom `--size` requires width/height in `[512, 2048]`, divisible by `16`, total pixels ≤ `2^21`\n\n## Behavior Notes\n\n- The sync API returns a temporary URL; baoyu-imagine downloads it and writes locally\n- `--ref` is not supported for Z.AI in this skill yet\n- The sync API returns a single image, so `--n > 1` is rejected\n\n## Official References\n\n- [GLM-Image Guide](https://docs.z.ai/guides/image/glm-image)\n- [Generate Image API](https://docs.z.ai/api-reference/image/generate-image)\n\nFile v1.115.4:references/usage-examples.md\n\n# Usage Examples\n\nExtended CLI examples. SKILL.md shows the minimum set; read this file when the user asks about provider-specific invocation, batch generation, or less-common flags.\n\n## Core Patterns\n\n```bash\n# Basic text-to-image\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A cat\" --image cat.png\n\n# With aspect ratio\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A landscape\" --image out.png --ar 16:9\n\n# High quality\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A cat\" --image out.png --quality\n\nArchive v1.115.1: 35 files, 91535 bytes\n\nFiles: references/config/first-time-setup.md (13173b), references/config/preferences-schema.md (4224b), references/providers/dashscope.md (4591b), references/providers/minimax.md (1226b), references/providers/openrouter.md (776b), references/providers/replicate.md (1810b), references/providers/zai.md (1095b), references/usage-examples.md (4760b), scripts/build-batch.test.ts (4569b), scripts/build-batch.ts (7454b), scripts/main.test.ts (17852b), scripts/main.ts (43411b), scripts/providers/azure.test.ts (5423b), scripts/providers/azure.ts (5820b), scripts/providers/dashscope.test.ts (11534b), scripts/providers/dashscope.ts (18456b), scripts/providers/google.test.ts (3417b), scripts/providers/google.ts (9873b), scripts/providers/jimeng.test.ts (2737b), scripts/providers/jimeng.ts (12845b), scripts/providers/minimax.test.ts (5254b), scripts/providers/minimax.ts (6331b), scripts/providers/openai.test.ts (5536b), scripts/providers/openai.ts (12671b), scripts/providers/openrouter.test.ts (5450b), scripts/providers/openrouter.ts (9903b), scripts/providers/replicate.test.ts (7190b), scripts/providers/replicate.ts (17816b), scripts/providers/seedream.test.ts (6544b), scripts/providers/seedream.ts (9855b), scripts/providers/zai.test.ts (5274b), scripts/providers/zai.ts (8946b), scripts/types.ts (2123b), SKILL.md (14389b), _meta.json (134b)\n\nArchive v1.104.0: 35 files, 87161 bytes\n\nFiles: references/config/first-time-setup.md (12734b), references/config/preferences-schema.md (4224b), references/providers/dashscope.md (2382b), references/providers/minimax.md (1220b), references/providers/openrouter.md (776b), references/providers/replicate.md (1810b), references/providers/zai.md (1095b), references/usage-examples.md (4306b), scripts/build-batch.test.ts (4569b), scripts/build-batch.ts (7454b), scripts/main.test.ts (16085b), scripts/main.ts (42335b), scripts/providers/azure.test.ts (5423b), scripts/providers/azure.ts (5820b), scripts/providers/dashscope.test.ts (4359b), scripts/providers/dashscope.ts (13312b), scripts/providers/google.test.ts (3417b), scripts/providers/google.ts (9873b), scripts/providers/jimeng.test.ts (2737b), scripts/providers/jimeng.ts (12845b), scripts/providers/minimax.test.ts (5103b), scripts/providers/minimax.ts (6327b), scripts/providers/openai.test.ts (5536b), scripts/providers/openai.ts (12671b), scripts/providers/openrouter.test.ts (5450b), scripts/providers/openrouter.ts (9903b), scripts/providers/replicate.test.ts (7190b), scripts/providers/replicate.ts (17816b), scripts/providers/seedream.test.ts (6544b), scripts/providers/seedream.ts (9855b), scripts/providers/zai.test.ts (5274b), scripts/providers/zai.ts (8946b), scripts/types.ts (2123b), SKILL.md (14285b), _meta.json (134b)\n\nArchive v1.103.1: 27 files, 79275 bytes\n\nFiles: references/config/first-time-setup.md (12427b), references/config/preferences-schema.md (4215b), scripts/main.test.ts (16089b), scripts/main.ts (42339b), scripts/providers/azure.test.ts (5423b), scripts/providers/azure.ts (5822b), scripts/providers/dashscope.test.ts (4359b), scripts/providers/dashscope.ts (13312b), scripts/providers/google.test.ts (3417b), scripts/providers/google.ts (9873b), scripts/providers/jimeng.test.ts (2737b), scripts/providers/jimeng.ts (12845b), scripts/providers/minimax.test.ts (5103b), scripts/providers/minimax.ts (6327b), scripts/providers/openai.test.ts (4041b), scripts/providers/openai.ts (9586b), scripts/providers/openrouter.test.ts (5450b), scripts/providers/openrouter.ts (9903b), scripts/providers/replicate.test.ts (7190b), scripts/providers/replicate.ts (17816b), scripts/providers/seedream.test.ts (6544b), scripts/providers/seedream.ts (9855b), scripts/providers/zai.test.ts (5274b), scripts/providers/zai.ts (8946b), scripts/types.ts (2123b), SKILL.md (25745b), _meta.json (134b)\n\nArchive v1.103.0: 27 files, 79276 bytes\n\nFiles: references/config/first-time-setup.md (12427b), references/config/preferences-schema.md (4215b), scripts/main.test.ts (16089b), scripts/main.ts (42339b), scripts/providers/azure.test.ts (5423b), scripts/providers/azure.ts (5822b), scripts/providers/dashscope.test.ts (4359b), scripts/providers/dashscope.ts (13312b), scripts/providers/google.test.ts (3417b), scripts/providers/google.ts (9873b), scripts/providers/jimeng.test.ts (2737b), scripts/providers/jimeng.ts (12845b), scripts/providers/minimax.test.ts (5103b), scripts/providers/minimax.ts (6327b), scripts/providers/openai.test.ts (4041b), scripts/providers/openai.ts (9586b), scripts/providers/openrouter.test.ts (5450b), scripts/providers/openrouter.ts (9903b), scripts/providers/replicate.test.ts (7190b), scripts/providers/replicate.ts (17816b), scripts/providers/seedream.test.ts (6544b), scripts/providers/seedream.ts (9855b), scripts/providers/zai.test.ts (5274b), scripts/providers/zai.ts (8946b), scripts/types.ts (2123b), SKILL.md (25745b), _meta.json (134b)\n\nArchive v1.0.1: 25 files, 66326 bytes\n\nFiles: references/config/first-time-setup.md (10555b), references/config/preferences-schema.md (3648b), scripts/main.test.ts (13033b), scripts/main.ts (39118b), scripts/providers/azure.test.ts (5396b), scripts/providers/azure.ts (5822b), scripts/providers/dashscope.test.ts (4359b), scripts/providers/dashscope.ts (13312b), scripts/providers/google.test.ts (3390b), scripts/providers/google.ts (9873b), scripts/providers/jimeng.test.ts (2710b), scripts/providers/jimeng.ts (12845b), scripts/providers/minimax.test.ts (5076b), scripts/providers/minimax.ts (6327b), scripts/providers/openai.test.ts (1949b), scripts/providers/openai.ts (6233b), scripts/providers/openrouter.test.ts (5423b), scripts/providers/openrouter.ts (9903b), scripts/providers/replicate.test.ts (2430b), scripts/providers/replicate.ts (5987b), scripts/providers/seedream.test.ts (6517b), scripts/providers/seedream.ts (9855b), scripts/types.ts (1749b), SKILL.md (20524b), _meta.json (132b)\n\nArchive v1.0.0: 25 files, 65701 bytes\n\nFiles: references/config/first-time-setup.md (10555b), references/config/preferences-schema.md (3648b), scripts/main.test.ts (11138b), scripts/main.ts (38077b), scripts/providers/azure.test.ts (5396b), scripts/providers/azure.ts (5822b), scripts/providers/dashscope.test.ts (4359b), scripts/providers/dashscope.ts (13312b), scripts/providers/google.test.ts (3390b), scripts/providers/google.ts (9873b), scripts/providers/jimeng.test.ts (2710b), scripts/providers/jimeng.ts (12845b), scripts/providers/minimax.test.ts (5076b), scripts/providers/minimax.ts (6327b), scripts/providers/openai.test.ts (1949b), scripts/providers/openai.ts (6233b), scripts/providers/openrouter.test.ts (5423b), scripts/providers/openrouter.ts (9903b), scripts/providers/replicate.test.ts (2430b), scripts/providers/replicate.ts (5987b), scripts/providers/seedream.test.ts (6517b), scripts/providers/seedream.ts (9855b), scripts/types.ts (1749b), SKILL.md (20309b), _meta.json (132b)","readmeExcerpt":"Skill: Baoyu Imagine Owner: jimliu Summary: AI image generation with OpenAI GPT Image 2, Azure OpenAI, Google, OpenRouter, DashScope, Z.AI GLM-Image, MiniMax, Jimeng, Seedream and Replicate APIs. Suppo... Tags: latest:1.117.3 Version history: v1.117.3 | 2026-05-24T22:13:53.081Z | auto - Added documentation for reference-image identity preservation, including best practices and anti-patterns. - Updated usage examples ","codeSnippets":[],"executableExamples":[{"language":"bash","snippet":"# Basic\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A cat\" --image cat.png\n\n# With aspect ratio and high quality\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A landscape\" --image out.png --ar 16:9 --quality 2k\n\n# Prompt from files\n${BUN_X} {baseDir}/scripts/main.ts --promptfiles system.md content.md --image out.png\n\n# With reference image\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"Make blue\" --image out.png --ref source.png\n\n# Specific provider\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A cat\" --image out.png --provider dashscope --model qwen-image-2.0-pro\n\n# OpenAI GPT Image 2\n${BUN_X} {baseDir}/scripts/main.ts --prompt \"A cat\" --image out.png --provider openai --model gpt-image-2\n\n# Batch mode\n${BUN_X} {baseDir}/scripts/main.ts --batchfile batch.json --jobs 4"},{"language":"text","snippet":"OPENAI_API_KEY is required. Codex/ChatGPT desktop login does not automatically grant OpenAI Images API access to this script."},{"language":"text","snippet":"No EXTEND.md found          EXTEND.md found, model null\n        │                            │\n        ▼                            ▼\n┌─────────────────────┐    ┌──────────────────────┐\n│ AskUserQuestion     │    │ AskUserQuestion      │\n│ (full setup)        │    │ (model only)         │\n└─────────────────────┘    └──────────────────────┘\n        │                            │\n        ▼                            ▼\n┌─────────────────────┐    ┌──────────────────────┐\n│ Create EXTEND.md    │    │ Update EXTEND.md     │\n└─────────────────────┘    └──────────────────────┘\n        │                            │\n        ▼                            ▼\n    Continue                     Continue"},{"language":"yaml","snippet":"header: \"Provider\"\nquestion: \"Default image generation provider?\"\noptions:\n  - label: \"Google (Recommended)\"\n    description: \"Gemini multimodal - high quality, reference images, flexible sizes\"\n  - label: \"OpenAI\"\n    description: \"GPT Image 2 - latest OpenAI image model, reference-image workflows\"\n  - label: \"Azure OpenAI\"\n    description: \"Azure-hosted GPT Image deployments with resource-specific routing\"\n  - label: \"OpenRouter\"\n    description: \"Router for Gemini/FLUX/OpenAI-compatible image models\"\n  - label: \"DashScope\"\n    description: \"Alibaba Cloud - Qwen-Image, strong Chinese/English text rendering\"\n  - label: \"Z.AI\"\n    description: \"GLM-image, strong poster and text-heavy image generation\"\n  - label: \"MiniMax\"\n    description: \"MiniMax image generation with subject-reference character workflows\"\n  - label: \"Replicate\"\n    description: \"Curated Replicate image families - nano-banana-2, Seedream, and Wan image models\""},{"language":"yaml","snippet":"header: \"Google Model\"\nquestion: \"Default Google image generation model?\"\noptions:\n  - label: \"gemini-3-pro-image-preview (Recommended)\"\n    description: \"Highest quality, best for production use\"\n  - label: \"gemini-3.1-flash-image-preview\"\n    description: \"Fast generation, good quality, lower cost\"\n  - label: \"gemini-3-flash-preview\"\n    description: \"Fast generation, balanced quality and speed\""},{"language":"yaml","snippet":"header: \"OpenRouter Model\"\nquestion: \"Default OpenRouter image generation model?\"\noptions:\n  - label: \"google/gemini-3.1-flash-image-preview (Recommended)\"\n    description: \"Best general-purpose OpenRouter image model with reference-image workflows\"\n  - label: \"google/gemini-2.5-flash-image-preview\"\n    description: \"Fast Gemini preview model on OpenRouter\"\n  - label: \"black-forest-labs/flux.2-pro\"\n    description: \"Strong text-to-image quality through OpenRouter\""}],"parameters":null,"dependencies":[],"permissions":[],"extractedFiles":[{"path":"SKILL.md","content":"---\nname: baoyu-imagine\ndescription: AI image generation with OpenAI GPT Image 2, Azure OpenAI, Google, OpenRouter, DashScope, Z.AI GLM-Image, MiniMax, Jimeng, Seedream and Replicate APIs. Supports text-to-image, reference images, aspect ratios, and batch generation from saved prompt files. Sequential by default; use batch parallel generation when the user already has multiple prompts or wants stable multi-image throughput. Use when user asks to generate, create, or draw images.\nversion: 1.117.3\nmetadata:\n  openclaw:\n    homepage: https://github.com/JimLiu/baoyu-skills#baoyu-imagine\n    requires:\n      anyBins:\n        - bun\n        - npx\n---\n\n# Image Generation (AI SDK)\n\nOfficial API-based image generation. Supports OpenAI GPT Image 2, Azure OpenAI, Google, OpenRouter, DashScope (阿里通义万象), Z.AI GLM-Image, MiniMax, Jimeng (即梦), Seedream (豆包) and Replicate.\n\n## User Input Tools\n\nWhen this skill prompts the user, follow this tool-selection rule (priority order):\n\n1. **Prefer built-in user-input tools** exposed by the current agent runtime — e.g., `AskUserQuestion`, `request_user_input`, `clarify`, `ask_user`, or any equivalent.\n2. **Fallback**: if no such tool exists, emit a numbered plain-text message and ask the user to reply with the chosen number/answer for each question.\n3. **Batching**: if the tool supports multiple questions per call, combine all applicable questions into a single call; if only single-question, ask them one at a time in priority order.\n\nConcrete `AskUserQuestion` references below are examples — substitute the local equivalent in other runtimes.\n\n## Script Directory\n\n`{baseDir}` = this SKILL.md's directory. Main script: `{baseDir}/scripts/main.ts`. Resolve `${BUN_X}`: prefer `bun`; else `npx -y bun`; else suggest `brew install oven-sh/bun/bun`.\n\n## Step 0: Load Preferences ⛔ BLOCKING\n\nThis step MUST complete before any image generation — generation is blocked until EXTEND.md exists.\n\nCheck these paths in order; first hit wins:\n\n| Path | Scope |\n|------|-------|\n| `.baoyu-skills/baoyu-imagine/EXTEND.md` | Project |\n| `${XDG_CONFIG_HOME:-$HOME/.config}/baoyu-skills/baoyu-imagine/EXTEND.md` | XDG |\n| `$HOME/.baoyu-skills/baoyu-imagine/EXTEND.md` | User home |\n\n- **Found** → load, parse, apply. If `default_model.[provider]` is null → ask model only.\n- **Not found** → run first-time setup (`references/config/first-time-setup.md`) using AskUserQuestion to collect provider + model + quality + save location. Save EXTEND.md, then continue. Do not generate images before this completes.\n\nLegacy compatibility: if `.baoyu-skills/baoyu-image-gen/EXTEND.md` exists and the new path doesn't, the runtime renames it to `baoyu-imagine`. If both exist, the runtime leaves them alone and uses the new path.\n\n**EXTEND.md keys**: default provider, default quality, default aspect ratio, default image size, OpenAI image API dialect, default models, batch worker cap, provider-specific batch limits. Schema: `references/config/preferences-schema.md`.\n\n## Us"},{"path":"_meta.json","content":"{\n  \"ownerId\": \"kn7csrrndw79hpke5d0gsnx93d82k67r\",\n  \"slug\": \"baoyu-imagine\",\n  \"version\": \"1.117.3\",\n  \"publishedAt\": 1779660833081\n}"},{"path":"references/codex-image2-fallback.md","content":"---\nname: codex-image2-fallback\ndescription: Fallback behavior when baoyu-imagine lacks OpenAI API credentials but Codex/native image generation is available\n---\n\n# Codex Image2 Fallback\n\nWhen using `baoyu-imagine` with `--provider openai --model gpt-image-2`, the CLI can fail with:\n\n```text\nOPENAI_API_KEY is required. Codex/ChatGPT desktop login does not automatically grant OpenAI Images API access to this script.\n```\n\nThis is expected. The `openai` provider uses the public OpenAI Images API and needs `OPENAI_API_KEY`. Codex / ChatGPT image2 entitlement is a separate runtime-native path.\n\n## Practical fallback pattern\n\n1. Try `baoyu-imagine` when provider credentials are available.\n2. If it fails only because `OPENAI_API_KEY` is missing, do not leave the user waiting.\n3. Prefer a Codex/native raster backend in this order:\n   - Codex runtime native `imagegen` skill/tool, if available.\n   - Repo-level `scripts/codex-imagegen.sh`, if `codex` CLI is installed/logged in and the calling skill supports the wrapper.\n   - Hermes native `image_generate`, if available.\n4. Be transparent about reference-image behavior:\n   - If the fallback backend accepts references, pass the reference images.\n   - If it does not, derive a concise identity-preserving prompt from the references and state that it is a text-description fallback, not strict reference-image editing.\n5. Return the generated media path or structured backend error promptly.\n\n## User-facing wording\n\nUse concise wording such as:\n\n> The OpenAI API path needs `OPENAI_API_KEY`; Codex login is a separate image2 backend. I used the available Codex/native image backend instead. Reference images were [passed directly / reconstructed from visual traits].\n\nAvoid implying that `baoyu-imagine --provider openai` can use Codex OAuth without a dedicated provider implementation."},{"path":"references/codex-oauth-vs-openai-api-key.md","content":"# Codex OAuth vs OpenAI API key for baoyu-imagine\n\n`baoyu-imagine --provider openai` uses the standard OpenAI Images API and requires `OPENAI_API_KEY`. It calls OpenAI-compatible image endpoints such as `/images/generations` and `/images/edits`.\n\nCodex / ChatGPT login is different. Codex image generation is driven by Codex OAuth and the Codex runtime's `image_gen` capability, not by the public OpenAI Images API key path. A Codex OAuth token is not a drop-in replacement for `OPENAI_API_KEY`, and setting `OPENAI_BASE_URL` to a Codex backend will not make baoyu-imagine's existing `openai` provider work because the auth, route, and payload shape differ.\n\n## What to use instead\n\n- If running inside Codex and the native `imagegen` skill/tool is available, use it directly.\n- If running outside Codex but the `codex` CLI is installed and logged in, use the repo-level `scripts/codex-imagegen.sh` wrapper when the calling skill supports it. The wrapper invokes `codex exec` and the Codex `image_gen` tool; no `OPENAI_API_KEY` is required.\n- If running inside Hermes and a native `image_generate` tool is available, use that as a runtime-native fallback. Be explicit about whether reference images are passed directly or only reconstructed from extracted traits.\n- If the user wants `baoyu-imagine` itself to support Codex OAuth, add a distinct provider such as `openai-codex` rather than modifying the existing `openai` provider.\n\n## Reference-image prompting note\n\nWhen using actual reference images for identity preservation, avoid long generic descriptions of the subject. Long descriptions can cause the model to synthesize a new similar-looking person/object. Prefer direct wording:\n\n> Use the person/object in the reference image(s) as the same identity. Do not redesign it or create a similar-looking new subject. Only change scene, clothing, pose, lighting, rendering style, and composition."},{"path":"references/config/first-time-setup.md","content":"---\nname: first-time-setup\ndescription: First-time setup and default model selection flow for baoyu-imagine\n---\n\n# First-Time Setup\n\n## Overview\n\nTriggered when:\n1. No EXTEND.md found → full setup (provider + model + preferences)\n2. EXTEND.md found but `default_model.[provider]` is null → model selection only\n\n## Setup Flow\n\n```\nNo EXTEND.md found          EXTEND.md found, model null\n        │                            │\n        ▼                            ▼\n┌─────────────────────┐    ┌──────────────────────┐\n│ AskUserQuestion     │    │ AskUserQuestion      │\n│ (full setup)        │    │ (model only)         │\n└─────────────────────┘    └──────────────────────┘\n        │                            │\n        ▼                            ▼\n┌─────────────────────┐    ┌──────────────────────┐\n│ Create EXTEND.md    │    │ Update EXTEND.md     │\n└─────────────────────┘    └──────────────────────┘\n        │                            │\n        ▼                            ▼\n    Continue                     Continue\n```\n\n## Flow 1: No EXTEND.md (Full Setup)\n\n**Language**: Use user's input language or saved language preference.\n\nUse AskUserQuestion with ALL questions in ONE call:\n\n### Question 1: Default Provider\n\n```yaml\nheader: \"Provider\"\nquestion: \"Default image generation provider?\"\noptions:\n  - label: \"Google (Recommended)\"\n    description: \"Gemini multimodal - high quality, reference images, flexible sizes\"\n  - label: \"OpenAI\"\n    description: \"GPT Image 2 - latest OpenAI image model, reference-image workflows\"\n  - label: \"Azure OpenAI\"\n    description: \"Azure-hosted GPT Image deployments with resource-specific routing\"\n  - label: \"OpenRouter\"\n    description: \"Router for Gemini/FLUX/OpenAI-compatible image models\"\n  - label: \"DashScope\"\n    description: \"Alibaba Cloud - Qwen-Image, strong Chinese/English text rendering\"\n  - label: \"Z.AI\"\n    description: \"GLM-image, strong poster and text-heavy image generation\"\n  - label: \"MiniMax\"\n    description: \"MiniMax image generation with subject-reference character workflows\"\n  - label: \"Replicate\"\n    description: \"Curated Replicate image families - nano-banana-2, Seedream, and Wan image models\"\n```\n\n### Question 2: Default Google Model\n\nOnly show if user selected Google or auto-detect (no explicit provider).\n\n```yaml\nheader: \"Google Model\"\nquestion: \"Default Google image generation model?\"\noptions:\n  - label: \"gemini-3-pro-image-preview (Recommended)\"\n    description: \"Highest quality, best for production use\"\n  - label: \"gemini-3.1-flash-image-preview\"\n    description: \"Fast generation, good quality, lower cost\"\n  - label: \"gemini-3-flash-preview\"\n    description: \"Fast generation, balanced quality and speed\"\n```\n\n### Question 2b: Default OpenRouter Model\n\nOnly show if user selected OpenRouter.\n\n```yaml\nheader: \"OpenRouter Model\"\nquestion: \"Default OpenRouter image generation model?\"\noptions:\n  - label: \"google/gemini-3.1-flash-image-preview (Recommended)\"\n    description: \"Best general-purpose OpenR"}],"languages":[],"docsSourceLabel":"CLAWHUB","editorialOverview":null,"editorialQuality":{"score":100,"threshold":65,"status":"thin","wordCount":1964,"uniquenessScore":40,"reasons":["uniqueness-below-45"]}},"media":{"evidence":{"source":"no-media","verified":false,"confidence":"low","updatedAt":"2026-10-09T09:29:37.766Z","emptyReason":"No screenshots, media assets, or demo links are available."},"primaryImageUrl":null,"mediaAssetCount":0,"assets":[],"demoUrl":null},"ownerResources":{"evidence":{"source":"unclaimed","verified":false,"confidence":"low","updatedAt":"2026-10-09T09:29:37.766Z","emptyReason":"This page has not been claimed by the agent owner."},"hasCustomPage":false,"customPageUpdatedAt":null,"customLinks":[],"structuredLinks":{"docsUrl":null,"demoUrl":null,"supportUrl":null,"pricingUrl":null,"statusUrl":null},"customPage":null},"relatedAgents":{"evidence":{"source":"protocol-neighbors","verified":false,"confidence":"medium","updatedAt":"2026-10-09T22:38:20.864Z","emptyReason":null},"items":[{"id":"8ebccd8e-3863-4187-8355-c3f14e1f9edf","entityType":"agent","canonicalPath":"/agent/iofficeai-aionui","slug":"iofficeai-aionui","name":"AionUi","description":"Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!","url":"https://github.com/iOfficeAI/AionUi","homepage":"https://www.aionui.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-10-09T19:11:12.944Z","createdAt":"2026-02-25T03:38:16.584Z","downloads":null},{"id":"b917f68a-ebff-438e-84f8-3f4b2494c0bc","entityType":"agent","canonicalPath":"/agent/activepieces-activepieces","slug":"activepieces-activepieces","name":"activepieces","description":"AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents","url":"https://github.com/activepieces/activepieces","homepage":"https://www.activepieces.com","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-15T02:22:12.426Z","createdAt":"2026-02-25T03:38:12.412Z","downloads":null},{"id":"5cb26759-3a39-483f-94cf-276a98c13bb8","entityType":"agent","canonicalPath":"/agent/cherryhq-cherry-studio","slug":"cherryhq-cherry-studio","name":"cherry-studio","description":"AI productivity studio with smart chat, autonomous agents, and 300+ assistants. Unified access to frontier LLMs","url":"https://github.com/CherryHQ/cherry-studio","homepage":"https://cherry-ai.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-11T14:38:40.986Z","createdAt":"2026-02-25T03:38:19.379Z","downloads":null},{"id":"6f6582d0-5d76-4f0f-b81d-86520247950b","entityType":"agent","canonicalPath":"/agent/copilotkit-copilotkit","slug":"copilotkit-copilotkit","name":"CopilotKit","description":"The Frontend for Agents & Generative UI. React + Angular","url":"https://github.com/CopilotKit/CopilotKit","homepage":"https://docs.copilotkit.ai","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-03-25T09:50:57.846Z","createdAt":"2026-02-25T03:39:14.617Z","downloads":null}],"links":{"hub":"/agent","source":"/agent/source/clawhub","protocols":[{"label":"OpenClaw","href":"/agent/protocol/openclew"}]}}}