{"id":"42cc3f1d-023e-424b-b89b-fe346cb5148d","entityType":"agent","slug":"clawhub-sunshinejnjn-image-with-comfyui","name":"image-with-comfyui","canonicalUrl":"https://www.xpersona.co/agent/clawhub-sunshinejnjn-image-with-comfyui","canonicalPath":"/agent/clawhub-sunshinejnjn-image-with-comfyui","generatedAt":"2026-10-10T05:40:20.918Z","source":"CLAWHUB","claimStatus":"UNCLAIMED","verificationTier":"NONE","summary":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T00:30:13.191Z","emptyReason":null},"description":"Use to generate, edit, or animate images and videos through a user-defined ComfyUI server (COMFYUI_URL env var or comfyui_url in config.json; if no server is reachable the script exits with configuration instructions) — trigger with requests like \"make a picture of\", \"replace the background\", \"cut out / extract the subject of\", \"turn this into a video\", \"edit my photo\", or \"make a 3D mesh of this character\". Not for simple non-generation image tweaks. **Qwen-Image 2.1** is the default model for both text-to-image (T2I) and image editing/multi-image (I2I); it can also **cut out / extract the main subject or any image element into a transparent-background image**. Z-Image / SD3.5 (T2I), Qwen Image Edit (2511, single-image), Wan2.2 (I2V), and Hunyuan 3D v2.1 (I2M, image → 3D GLB mesh) are available on request.","descriptionLabel":"Source description","evidenceSummary":"Capability contract not published. No trust telemetry is available yet. 1.8K downloads reported by the source. Last updated 10/10/2026.","installCommand":"clawhub skill install s17ehnmpwm46zw2010e8g6j2z583gzv3:image-with-comfyui","sourceUrl":"https://clawhub.ai/sunshinejnjn/image-with-comfyui","homepage":"https://clawhub.ai/sunshinejnjn/skills/image-with-comfyui","primaryLinks":[{"label":"View on ClawHub","url":"https://clawhub.ai/sunshinejnjn/image-with-comfyui","kind":"source"},{"label":"Homepage","url":"https://clawhub.ai/sunshinejnjn/skills/image-with-comfyui","kind":"homepage"}],"safetyScore":84,"overallRank":62,"popularityScore":65,"trustScore":null,"claimedByName":null,"isOwner":false,"seoDescription":"image-with-comfyui technical dossier on Xpersona with agent coverage, OPENCLEW support, and live trust metadata."},"coverage":{"evidence":{"source":"public-profile","verified":false,"confidence":"medium","updatedAt":"2026-10-10T00:30:13.191Z","emptyReason":null},"protocols":[{"protocol":"OPENCLEW","label":"OpenClaw","status":"self-declared","notes":"Declared in the public agent profile."}],"capabilities":[],"verifiedCount":0,"selfDeclaredCount":1,"capabilityMatrix":{"rows":[{"key":"OPENCLEW","type":"protocol","support":"unknown","confidenceSource":"profile","notes":"Listed on profile"}],"flattenedTokens":"protocol:OPENCLEW|unknown|profile"}},"adoption":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T00:30:13.191Z","emptyReason":null},"stars":null,"forks":null,"downloads":1835,"packageName":null,"latestVersion":"3.6.0","tractionLabel":"1.8K downloads"},"release":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T00:30:13.191Z","emptyReason":null},"lastUpdatedAt":"2026-10-10T00:30:13.191Z","lastCrawledAt":"2026-10-10T00:30:13.191Z","lastIndexedAt":null,"nextCrawlAt":"2026-10-11T00:30:13.191Z","lastVerifiedAt":null,"highlights":[{"version":"3.6.0","createdAt":"2026-10-04T17:57:44.035Z","changelog":"Added documentation that Qwen-Image 2.1 can cut out / extract the subject or any image element into a transparent-background image.","fileCount":19,"zipByteSize":66793},{"version":"3.5.1","createdAt":"2026-10-03T17:42:26.849Z","changelog":"Fixed workspace directory detection so output files land on the correct standard path.","fileCount":19,"zipByteSize":66219},{"version":"3.5.0","createdAt":"2026-10-01T21:40:22.656Z","changelog":"Add reference to the sunshinejnjn/comfyui-auth custom node for Bearer API key auth; version bump to 3.5.0","fileCount":19,"zipByteSize":65887},{"version":"3.2.0","createdAt":"2026-09-30T07:14:04.925Z","changelog":"v3.2.0: i2m delivery scaling. The Hunyuan 3D v2.1 GLB comes out at a normalized unit scale (character ~2 units, no declared units), so the script now scales the GLB in place by a delivery factor (default x3000, new config mesh.hunyuan3d_v2.1.default_scale; --scale to override, --scale 1 to skip) BEFORE zipping it. New scale_glb_inplace() scales POSITION data + accessor min/max (and node translations) and rewrites a valid GLB; the result block prints the scaled bounding box. Unit-tested on real Hunyuan 3D output (raw positions verified x3000 exactly). Docs updated (SKILL.md, cli-reference, delivery, README).","fileCount":19,"zipByteSize":64429},{"version":"3.1.1","createdAt":"2026-09-30T06:41:55.459Z","changelog":"v3.1.1: i2m delivery rule (3) refined — a third-party online GLB viewer link IS allowed (the user views/rotates the mesh themselves; zip + character image remain the primary delivery). Still never render the mesh: no GLB->image/video conversion, no 3D-viewer screenshots. Supersedes the 3.1.0 wording ('no third-party viewer links').","fileCount":19,"zipByteSize":62455},{"version":"3.1.0","createdAt":"2026-09-30T06:25:29.554Z","changelog":"v3.1.0: i2m delivery rules. (1) The extracted character image from step 1 is attached to the user as an image (never just its path; with --extract none there is nothing new to send back). (2) The GLB is now automatically packaged by the script as i2m_<timestamp>.zip and sent as a zip attachment — the raw .glb is never sent alone. (3) The mesh is never rendered: no GLB->image/video conversion, no 3D-viewer screenshots, no third-party viewer links. New zip_mesh_files() in the script; docs updated (SKILL.md, cli-reference, delivery, README).","fileCount":19,"zipByteSize":62399},{"version":"3.0.0","createdAt":"2026-09-29T21:43:42.968Z","changelog":"v3.0.0: new i2m mode — image → 3D mesh via Hunyuan 3D v2.1 (GLB). Auto-extracts the main character first (i2i transparent cutout by default, or --extract rembg / none), then runs the 3D workflow. Adds workflows/hunyuan3d-v2.1_i2m_api.json (user graph, topology kept 1:1), mesh config block, GLB download + sanitization, and docs (SKILL/README/cli-reference/attribution/MODEL_URL). Verified end-to-end on live server: RGBA cutout + valid glTF v2 GLB (174k verts). Published under the beta tag — 'latest' intentionally unchanged (2.5.0).","fileCount":19,"zipByteSize":60953},{"version":"2.5.0","createdAt":"2026-09-29T16:53:36.421Z","changelog":"v2.5.0: PNG/transparency routing rule — if the user asks for a PNG image/file, or an RGBA image with transparency/alpha, the agent must pass --format png and deliver the PNG file as-is (server PNG bytes pass through unchanged; never flatten alpha onto white via JPG). Rule documented in SKILL.md output-format section + cli-reference.md Output Format with the trigger-phrase list. Verified: fmt=png saves byte-identical RGBA PNG; fmt=jpg still composites to white.","fileCount":18,"zipByteSize":54554}]},"execution":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No published capability contract is available yet."},"installCommand":"clawhub skill install s17ehnmpwm46zw2010e8g6j2z583gzv3:image-with-comfyui","setupComplexity":"low","setupSteps":["Install using `clawhub skill install s17ehnmpwm46zw2010e8g6j2z583gzv3:image-with-comfyui` in an isolated environment before connecting it to live workloads.","No published capability contract is available yet, so validate auth and request/response behavior manually.","Review the upstream CLAWHUB listing at https://clawhub.ai/sunshinejnjn/image-with-comfyui before using production credentials."],"contract":{"contractStatus":"missing","authModes":[],"requires":[],"forbidden":[],"supportsMcp":false,"supportsA2a":false,"supportsStreaming":false,"inputSchemaRef":null,"outputSchemaRef":null,"dataRegion":null,"contractUpdatedAt":null,"sourceUpdatedAt":null,"freshnessSeconds":null},"invocationGuide":{"preferredApi":{"snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-sunshinejnjn-image-with-comfyui/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-sunshinejnjn-image-with-comfyui/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-sunshinejnjn-image-with-comfyui/trust"},"curlExamples":["curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-sunshinejnjn-image-with-comfyui/snapshot\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-sunshinejnjn-image-with-comfyui/contract\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-sunshinejnjn-image-with-comfyui/trust\""],"jsonRequestTemplate":{"query":"summarize this repo","constraints":{"maxLatencyMs":2000,"protocolPreference":["OPENCLEW"]}},"jsonResponseTemplate":{"ok":true,"result":{"summary":"...","confidence":0.9},"meta":{"source":"CLAWHUB","generatedAt":"2026-10-10T05:40:20.914Z"}},"retryPolicy":{"maxAttempts":3,"backoffMs":[500,1500,3500],"retryableConditions":["HTTP_429","HTTP_503","NETWORK_TIMEOUT"]}},"endpoints":{"dossierUrl":"https://www.xpersona.co/api/v1/agents/clawhub-sunshinejnjn-image-with-comfyui/dossier","snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-sunshinejnjn-image-with-comfyui/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-sunshinejnjn-image-with-comfyui/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-sunshinejnjn-image-with-comfyui/trust"}},"reliability":{"evidence":{"source":"runtime-metrics","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No trust, reliability, or runtime telemetry is available."},"trust":{"status":"unavailable","handshakeStatus":"UNKNOWN","verificationFreshnessHours":null,"reputationScore":null,"p95LatencyMs":null,"successRate30d":null,"fallbackRate":null,"attempts30d":null,"trustUpdatedAt":null,"trustConfidence":"unknown","sourceUpdatedAt":null,"freshnessSeconds":null},"decisionGuardrails":{"doNotUseIf":["Contract metadata is missing or unavailable for deterministic execution."],"safeUseWhen":[],"riskFlags":["missing_or_unavailable_contract","trust_data_unavailable","schema_references_missing"],"operationalConfidence":"low"},"executionMetrics":{"observedLatencyMsP50":null,"observedLatencyMsP95":null,"estimatedCostUsd":null,"uptime30d":null,"rateLimitRpm":null,"rateLimitBurst":null,"lastVerifiedAt":null,"verificationSource":null},"runtimeMetrics":{"successRate":null,"avgLatencyMs":null,"avgCostUsd":null,"hallucinationRate":null,"retryRate":null,"disputeRate":null,"p50Latency":null,"p95Latency":null,"lastUpdated":null}},"benchmarks":{"evidence":{"source":"no-benchmark-data","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No benchmark suites or observed failure patterns are available."},"suites":[],"failurePatterns":[]},"artifacts":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T00:30:13.191Z","emptyReason":null},"readme":"Skill: image-with-comfyui\n\nOwner: sunshinejnjn\n\nSummary: Use to generate, edit, or animate images and videos through a user-defined ComfyUI server (COMFYUI_URL env var or comfyui_url in config.json; if no server is reachable the script exits with configuration instructions) — trigger with requests like \"make a picture of\", \"replace the background\", \"cut out / extract the subject of\", \"turn this into a video\", \"edit my photo\", or \"make a 3D mesh of this character\". Not for simple non-generation image tweaks. **Qwen-Image 2.1** is the default model for both text-to-image (T2I) and image editing/multi-image (I2I); it can also **cut out / extract the main subject or any image element into a transparent-background image**. Z-Image / SD3.5 (T2I), Qwen Image Edit (2511, single-image), Wan2.2 (I2V), and Hunyuan 3D v2.1 (I2M, image → 3D GLB mesh) are available on request.\n\nTags: beta:3.0.0, latest:3.6.0\n\nVersion history:\n\nv3.6.0 | 2026-10-04T17:57:44.035Z | user\n\nAdded documentation that Qwen-Image 2.1 can cut out / extract the subject or any image element into a transparent-background image.\n\nv3.5.1 | 2026-10-03T17:42:26.849Z | user\n\nFixed workspace directory detection so output files land on the correct standard path.\n\nv3.5.0 | 2026-10-01T21:40:22.656Z | user\n\nAdd reference to the sunshinejnjn/comfyui-auth custom node for Bearer API key auth; version bump to 3.5.0\n\nv3.2.0 | 2026-09-30T07:14:04.925Z | user\n\nv3.2.0: i2m delivery scaling. The Hunyuan 3D v2.1 GLB comes out at a normalized unit scale (character ~2 units, no declared units), so the script now scales the GLB in place by a delivery factor (default x3000, new config mesh.hunyuan3d_v2.1.default_scale; --scale to override, --scale 1 to skip) BEFORE zipping it. New scale_glb_inplace() scales POSITION data + accessor min/max (and node translations) and rewrites a valid GLB; the result block prints the scaled bounding box. Unit-tested on real Hunyuan 3D output (raw positions verified x3000 exactly). Docs updated (SKILL.md, cli-reference, delivery, README).\n\nv3.1.1 | 2026-09-30T06:41:55.459Z | user\n\nv3.1.1: i2m delivery rule (3) refined — a third-party online GLB viewer link IS allowed (the user views/rotates the mesh themselves; zip + character image remain the primary delivery). Still never render the mesh: no GLB->image/video conversion, no 3D-viewer screenshots. Supersedes the 3.1.0 wording ('no third-party viewer links').\n\nv3.1.0 | 2026-09-30T06:25:29.554Z | user\n\nv3.1.0: i2m delivery rules. (1) The extracted character image from step 1 is attached to the user as an image (never just its path; with --extract none there is nothing new to send back). (2) The GLB is now automatically packaged by the script as i2m_<timestamp>.zip and sent as a zip attachment — the raw .glb is never sent alone. (3) The mesh is never rendered: no GLB->image/video conversion, no 3D-viewer screenshots, no third-party viewer links. New zip_mesh_files() in the script; docs updated (SKILL.md, cli-reference, delivery, README).\n\nv3.0.0 | 2026-09-29T21:43:42.968Z | user\n\nv3.0.0: new i2m mode — image → 3D mesh via Hunyuan 3D v2.1 (GLB). Auto-extracts the main character first (i2i transparent cutout by default, or --extract rembg / none), then runs the 3D workflow. Adds workflows/hunyuan3d-v2.1_i2m_api.json (user graph, topology kept 1:1), mesh config block, GLB download + sanitization, and docs (SKILL/README/cli-reference/attribution/MODEL_URL). Verified end-to-end on live server: RGBA cutout + valid glTF v2 GLB (174k verts). Published under the beta tag — 'latest' intentionally unchanged (2.5.0).\n\nv2.5.0 | 2026-09-29T16:53:36.421Z | user\n\nv2.5.0: PNG/transparency routing rule — if the user asks for a PNG image/file, or an RGBA image with transparency/alpha, the agent must pass --format png and deliver the PNG file as-is (server PNG bytes pass through unchanged; never flatten alpha onto white via JPG). Rule documented in SKILL.md output-format section + cli-reference.md Output Format with the trigger-phrase list. Verified: fmt=png saves byte-identical RGBA PNG; fmt=jpg still composites to white.\n\nv2.4.1 | 2026-09-29T16:19:24.225Z | user\n\nv2.4.1: (1) wording — endpoint is now described as user-defined (COMFYUI_URL env var > config.json), loopback-first phrasing removed from SKILL.md/README/CLI help; intro now names Qwen-Image 2.1 as the default T2I+I2I model and lists Qwen Image Edit 2511 among on-request options; (2) t2i/i2i/wan2.2 pre-check the endpoint (GET /history, 5s) and on failure exit with step-by-step configuration instructions pointing at README.md setup docs; (3) error-handling.md documents the pre-check.\n\nv2.4.0 | 2026-09-29T15:27:35.448Z | user\n\nv2.4.0: qwen21 I2I --aspect now works via the workflow's latent switch (node 14): explicit --aspect routes the canvas through the ResolutionSelector (switch OFF) so a portrait reference CAN become a 16:9 landscape; without --aspect the canvas follows the first reference image's own size (no silent 1:1). Reference images are always uploaded at native size, never stretched.\n\nv2.3.0 | 2026-09-27T21:06:52.869Z | user\n\nv2.3.0: (1) unified output filename scheme — qwen21 t2i='qt2i', qwen21 i2i='qi2i', qwen_imageedit='qei2i', z-image='zt2i', sd35='st2i' + timestamp + counter; (2) trailing underscore stripped from saved filenames (client-side); (3) Qwen-Image 2.1 now exposes --negative for both T2I and I2I (existing negative_prompt input wired through CLI); docs updated\n\nv2.2.1 | 2026-09-27T20:17:22.262Z | user\n\nv2.2.1: default --aspect is now 1:1 (1024x1024) for t2i/i2i; all Chinese debug/log messages translated to English; default JPG quality lowered to 89 (config image.jpg_quality); workflow ResolutionSelector default set to 1:1 (Square); docs updated\n\nv2.2.0 | 2026-09-27T20:00:53.421Z | user\n\nv2.2.0: runtime fallback chain documented (t2i: qwen21->z-image->sd35->any available checkpoint; i2i: qwen21->qwen_imageedit; config problems abort without degrading; --no-fallback flag); JPG output: --format jpg/--jpg-quality with client-side Pillow conversion (alpha composited to white, PNG fallback); config image.output_format defaults to jpg; fixed stale model names in cli-reference (z-image, qwen_imageedit); version bump 2.1.2->2.2.0\n\nv2.1.2 | 2026-09-27T14:56:39.931Z | user\n\nSFW-only model list: removed all UC/NSFW sources + server private IP from MODEL_URL.txt; added filename-matching caveat for SFW weights; removed all keep_models_resident descriptions from SKILL.md; version bump 2.1.1 -> 2.1.2.\n\nv2.1.1 | 2026-09-27T14:19:14.339Z | user\n\nimage-with-comfyui 2.1.1 — Single universal model Qwen-Image 2.1 for T2I and I2I (dual-mode switch, no dual-weight set). Multi-image reference input up to 16 refs via --image (slot order = <image1>..<image16>). Added MODEL_URL.txt with verified SFW/HuggingFace download sources for every model. Documented ComfyUI workflow provenance + modifications (replaced GGUF load with standard safetensors via UNETLoader, added UnloadAllModels VRAM-cleanup) in references/attribution.md. Fix SD3.5 double-'.safetensors' suffix bug. Wan2.2 switched to 4-step Lightspeed weights. keep_models_resident default = false (VRAM cleared each run). SKILL.md Execution Discipline — Fast Path (no re-reading docs/config, single pre-check, minimal post-run verification) + mandatory --prompt flag guidance (100% avoidable failure). Full docs in English with bilingual prompt keywords preserved; added COMFYUI_KEEP_MODELS_RESIDENT env override; license/description_zh added to frontmatter.\n\nv1.7.1 | 2026-09-20T12:23:12.585Z | user\n\nv1.7.1: ClawHub security audit remediation. (HIGH) Bundled config.json default endpoint changed from a private-network LAN address to loopback-only http://127.0.0.1:8188; non-loopback COMFYUI_URL now prints an explicit data-flow disclosure warning (prompts + full source images for I2I/I2V) on every run; script docstring documents exact data flow. (MEDIUM) README custom-node install instructions now pin reviewed commits (Impact-Pack 429d0159a, WAS Nodes 9934caa92, ComfyUI-Manager 946ef8fe6), use git checkout --detach, recommend isolated venvs / non-root installs; corrected two dead repo URLs (ComfyUI-WAS-Nodes renamed to was-node-suite-comfyui; ComfyUI-Manager migrated to Comfy-Org). Docs: fixed argparse/docstring 'T2I/I2I' contradiction (now T2I/I2I/I2V), removed non-existent 'sd35' subcommand and 'test_all_workflows.py' references, added Data & Privacy section, English-prompt requirement softened to a recommendation, image-first mode scoped to same user/session, outbound staging cleanup rule added, added COMFYUI_TIMEOUT/COMFYUI_OUTPUT_DIR env support in the script. Added LICENSE (BSD-3-Clause) and VERSION.\n\nv1.7.0 | 2026-08-28T17:25:39.871Z | user\n\nAdd trigger-phrase-rich description for better skill routing (e.g. 'make a picture of', 'replace the background', 'turn this into a video', 'edit my photo'). Description now surfaces the four modes — Z-Image/SD3.5 T2I, Qwen I2I edits, Wan2.2 I2V — plus image-first mode, so the agent can match user intent faster.\n\nv1.5.0 | 2026-06-27T07:23:15.231Z | user\n\nMade env vars optional: COMFYUI_URL, COMFYUI_TIMEOUT, COMFYUI_POLL_INTERVAL, COMFYUI_OUTPUT_DIR, OPENCLAW_WORKSPACE no longer required in requires\n\nv1.4.9 | 2026-04-26T10:28:13.753Z | user\n\nMinor update\n\nv1.4.8 | 2026-04-26T10:05:41.950Z | user\n\nFull English translation of Chinese content; removed internal lesson source notes; updated to 1.4.8\n\nv1.4.7 | 2026-04-22T16:50:52.230Z | user\n\nFixed print_result_video: local_paths was str list but code called .stat() expecting Path objects; added Path() conversion for reliable video output printing.\n\nv1.4.6 | 2026-04-22T16:37:05.130Z | user\n\nAdded rule: deliver generated media in the user's original request session/thread, not in a separate topic.\n\nv1.4.5 | 2026-04-22T16:00:36.974Z | user\n\n### Security\n\n- Added sanitize_filename() to prevent path traversal attacks: filenames from the ComfyUI API are now validated against output_dir before being written. Rejects empty names, .., path separators, and any filename that resolves outside the output directory. Falls back to timestamped name if unsafe.\n- Applied sanitization across all output paths: get_output_images(), get_output_videos(), save_b64_image(), save_b64_video().,\n\nv1.4.4 | 2026-04-22T15:00:58.936Z | user\n\n### Added\n\n- Declared required environment variables in SKILL.md metadata (COMFYUI_URL, COMFYUI_TIMEOUT, COMFYUI_POLL_INTERVAL, COMFYUI_OUTPUT_DIR, OPENCLAW_WORKSPACE) to resolve registry warning about undocumented env vars\n\n### Fixed\n\n- Fixed typo in ImpactKSamplerAdvancedBasicPipe node name in NODE_PACKAGE_MAP\n\nv1.4.3 | 2026-04-22T13:01:22.225Z | user\n\n### Changed\n\n- Removed hardcoded git clone install commands from NODE_PACKAGE_MAP. The script no longer includes auto-executable install commands since ComfyUI may be on a remote server not reachable from this agent. Install commands now displayed as manual instructions in error output.\n\n### Added\n\n- OPENCLAW_WORKSPACE environment variable documented in configuration table\n\n### Removed\n\n- test_all_workflows.py (removed in 1.4.1)\n\nv1.4.2 | 2026-04-22T11:21:10.368Z | user\n\n### Changed\n\n- Updated default ComfyUI URL in documentation from remote server to localhost:8188 to reflect standard local installation\n\nv1.4.1 | 2026-04-22T11:17:14.967Z | user\n\n### Removed\n\n- Removed test_all_workflows.py script that contained hardcoded system-specific paths (user media paths, localhost defaults)\n\nv1.4.0 | 2026-04-22T10:53:01.084Z | user\n\n### Added\n\n- **wait_for_completion error detection**: Detects ComfyUI status_str: error and returns immediately instead of polling until timeout. Fixes 'No images found' false negatives.\n\n- **print_result output**: Shows filename, size, path on successful generation. Previously empty function leaving users uncertain.\n\n- **Image-First Conversational Pattern**: Auto-routes image+text within 2 min to Qwen I2I or Wan2.2 I2V.\n\n### Fixed\n\n- I2I no longer silently fails on ComfyUI backend errors. Error details now visible.\n\nv1.3.0 | 2026-04-22T08:24:47.223Z | user\n\n1.3.0: output dir auto-creation, I2V auto resolution detection (no more portrait-for-landscape), extended WAN2_2_RES with 16:9/1:1/4:3/3:2 categories\n\nv1.2.0 | 2026-04-22T06:04:17.877Z | user\n\nFull English translation, system output localization, config URL updates, README docs, UnloadAllModels bypass, error handling, t2i/i2i/i2v workflows\n\nArchive index:\n\nArchive v3.6.0: 19 files, 66793 bytes\n\nFiles: config.json (2438b), image_with_comfyui.py (110332b), MODEL_URL.txt (7015b), README.md (17687b), references/attribution.md (2476b), references/cli-reference.md (7159b), references/delivery.md (2435b), references/error-handling.md (4382b), references/image-first.md (2231b), references/prompts.md (6357b), skill-card.md (1998b), SKILL.md (13051b), workflows/hunyuan3d-v2.1_i2m_api.json (3286b), workflows/qwen_image_2.1_api.json (3702b), workflows/qwen_image-edit_api.json (4241b), workflows/sd3.5-med_t2i_api.json (2063b), workflows/wan2.2_i2v_api.json (8022b), workflows/z-image_t2i_api.json (2674b), _meta.json (137b)\n\nFile v3.6.0:SKILL.md\n\n---\nname: image-with-comfyui\ndescription: 'Use to generate, edit, or animate images and videos through a user-defined ComfyUI server (COMFYUI_URL env var or comfyui_url in config.json; if no server is reachable the script exits with configuration instructions) — trigger with requests like \"make a picture of\", \"replace the background\", \"cut out / extract the subject of\", \"turn this into a video\", \"edit my photo\", or \"make a 3D mesh of this character\". Not for simple non-generation image tweaks. **Qwen-Image 2.1** is the default model for both text-to-image (T2I) and image editing/multi-image (I2I); it can also **cut out / extract the main subject or any image element into a transparent-background image**. Z-Image / SD3.5 (T2I), Qwen Image Edit (2511, single-image), Wan2.2 (I2V), and Hunyuan 3D v2.1 (I2M, image → 3D GLB mesh) are available on request.'\nversion: \"3.5.0\"\nlicense: BSD-3-Clause\ncompatibility: 'ComfyUI ≥ 0.37.0; requires python3; the endpoint is user-defined (COMFYUI_URL or config.json) — point it at a server you control and trust.'\nallowed-tools: Read, Write, Edit, Exec\nmetadata:\n  openclaw:\n    emoji: \"🎨\"\n    requires:\n      anyBins: [\"python3\"]\n    config:\n      path: \"config.json\"\n---\n\n# Image with ComfyUI\n\nCall a user-defined ComfyUI server to generate or edit images and videos, or turn an image into a 3D mesh. **Qwen-Image 2.1** is the default model for both T2I and I2I; Z-Image / SD3.5 (T2I), Qwen Image Edit 2511 (single-image edit), Wan2.2 (I2V), and Hunyuan 3D v2.1 (I2M, image → 3D GLB) are available on request. If no ComfyUI server is reachable, the script exits with configuration instructions (see [README.md](README.md)).\n\n- **T2I** (Text → Image) → **Qwen-Image 2.1** (default) or Z-Image / SD3.5 Medium\n- **I2I** (Image → Image / Edit / Multi-image, up to 16 refs) → **Qwen-Image 2.1** (default) or Qwen Image Edit (2511, single image)\n- **I2V** (Image → Video) → Wan2.2 model\n- **Cutout** (extract the main subject or any image element) → **Qwen-Image 2.1** — redraws the target on a transparent background, producing a clean, alpha-only image; see [Cutout below](#cutout-extract-the-subject).\n- **I2M** (Image → 3D Mesh) → **Hunyuan 3D v2.1** — extracts the main character first (transparent PNG), generates a GLB, scales it ×3000, and zips it for delivery\n\n`config.json` at the skill root holds every default; any value can be overridden by an env var (see the Data & Privacy table there).\n\n## Cutout — Extract the Subject\n\n> **Note:** cutout is a genuine image-generation task, not a simple non-generation tweak — use it when the user wants to isolate the subject or any image element from its surroundings.\n\n**Qwen-Image 2.1** can **cut out / extract** the main subject (or a specified element) from an image and return it on a transparent background. The output is a clean PNG with alpha only — it re-draws the target rather than doing a hard pixel mask, so it works well for organic shapes (people, animals, products, characters) and can target a specific described element (e.g. \"extract just the red car\", \"isolate the vase\") via the prompt.\n\n- Use the **`i2i`** subcommand with `--image <path>` and a prompt that names the subject to extract. Keep the prompt focused on *what to keep*, not what to remove.\n- The result is transparent (alpha-only). Deliver it as-is — never convert to JPG (which has no alpha channel, so it would flatten transparency onto white).\n- This is the same redraw-to-transparent step that I2M runs first (see the i2m rules), but here it is used on its own to produce a standalone cutout image.\n\n## When to Use\n\n- Generate images from text → **t2i**\n- Edit one image, or compose 2–16 reference images → **i2i** (`--image` order = slot order = `<image1>`, `<image2>` …)\n- Turn an image into a short video → **wan2.2**\n- Turn the character in an image into a **3D mesh (GLB)** → **i2m** (Hunyuan 3D v2.1; the main character is extracted first, see below)\n- The user sends an image, then within 2 minutes asks to edit or animate it → **image-first mode**, see [references/image-first.md](references/image-first.md)\n\n## Data & Privacy (read before running)\n\n- **What is transmitted:** prompts and the workflow JSON go to the configured endpoint for **every** mode; **I2I, I2V and I2M additionally upload the full source image**.\n- **Endpoint** is user-defined (`COMFYUI_URL` env var > `comfyui_url` in `config.json`; bundled default `http://127.0.0.1:8188`). A non-local host prints a data-flow warning on every run; if no server answers, the script exits with configuration instructions instead of a raw timeout.\n- **Written to disk:** generated media is saved under `output_dir` (`config.json` or `COMFYUI_OUTPUT_DIR`). Don't submit sensitive content if the endpoint or disk is a concern.\n- **Secrets:** never put credentials in the URL or logs.\n\n## Workflow Files\n\n| Mode | Workflow | File |\n|---|---|---|\n| T2I (**Qwen-Image 2.1, default**) | Qwen-Image 2.1 dual-mode | `workflows/qwen_image_2.1_api.json` |\n| I2I (**Qwen-Image 2.1, default**, incl. multi-image) | same file, switch ON | `workflows/qwen_image_2.1_api.json` |\n| T2I (Z-Image) | Z-Image | `workflows/z-image_t2i_api.json` |\n| T2I (SD3.5) | SD3.5 Medium | `workflows/sd3.5-med_t2i_api.json` |\n| I2I (legacy) | Qwen Image Edit (2511) | `workflows/qwen_image-edit_api.json` |\n| I2V | Wan2.2 Image-to-Video | `workflows/wan2.2_i2v_api.json` |\n| I2M (Image → 3D Mesh) | Hunyuan 3D v2.1 | `workflows/hunyuan3d-v2.1_i2m_api.json` |\n\nEach workflow is an **adapted** version of a real ComfyUI graph. Original source + exactly what was changed → [references/attribution.md](references/attribution.md).\n\n## CLI Usage\n\n> **⚠️ The prompt is MANDATORY and must be passed via `--prompt` — never as a positional word.**\n>\n> Every subcommand (`t2i`, `i2i`, `wan2.2`) reads the prompt from `--prompt \"...\"`. Passing the prompt as a bare positional word fails immediately with exit code 2:\n> `python3 image_with_comfyui.py t2i \"some description\"` → `t2i: error: the following arguments are required: --prompt`\n>\n> This is a **100%-avoidable failure loop** — the command never reaches ComfyUI, so the model can't help. Always write the prompt through `--prompt` on the **first** attempt.\n>\n> `i2m` takes no prompt — the character itself is the condition (`--image` is its only required flag).\n\nFull commands and the timeout table → [references/cli-reference.md](references/cli-reference.md)\n\n## Prompt Formatting\n\nPer-model guidance (formulas, rules, example prompts) lives in [references/prompts.md](references/prompts.md). The English prose is there; example prompts keep their bilingual keyword pairs where the original was bilingual (e.g. `blurry 模糊`).\n\n- **Qwen-Image 2.1 (T2I + I2I default):** natural-language prompts. **CFG locked to 1.0** (distilled) — never pass `--cfg`. Optional `--negative` supported in both T2I and I2I (keep it short); I2I prompts stay concise, in the user's language.\n- **SD3.5 Medium:** natural language + optional `--negative`; default **1:1 (1024×1024)**, 20 steps, CFG 4.01.\n- **Z-Image / Qwen-Image 2.1:** `--aspect` is **optional** — T2I omits it and the output defaults to **1:1 (1024×1024)**; I2I omits it and the output canvas matches the **first reference image's size** (workflow latent switch ON). Pass `--aspect 16:9`, `9:16`, etc. to force a shape — in I2I this flips the latent switch to route the canvas through the ResolutionSelector (reference stays at native size, never stretched).\n- **Z-Image:** natural language — **no negative prompts**.\n- **Qwen Image Edit (2511):** concise, positive-only, user's exact words.\n- **Wan2.2 (I2V):** motion/action prompts; default 81 frames, 4 steps, CFG 4.5.\n- **i2m (Hunyuan 3D v2.1, no prompt):** the only input is the character image. Default extraction is **i2i** (Qwen-Image 2.1 redraws the main character centered on a transparent background, saved as PNG — this same redraw-to-transparent cutout is what makes the rest of I2M work; the same Qwen-Image 2.1 I2I call can produce a standalone cutout without any 3D follow-up). Use `--extract rembg` for a non-generative u2net human cutout, or `--extract none` to feed the image straight in (only if it's already a clean cutout). Defaults: resolution 4096, 30 steps, CFG 5.0, delivery scale **×3000** — the raw Hunyuan 3D v2.1 mesh is at a normalized unit scale (character ≈ 2 units, no declared units), so the GLB is scaled in place before zipping (`--scale 1` keeps it unscaled).\n\n## Error Handling\n\nThe script auto-detects three categories of server-side problems and reports them with install/source guidance, and `t2i`/`i2i` run through a **runtime fallback chain** when the default model fails. Details → [references/error-handling.md](references/error-handling.md):\n\n1. **Missing nodes** — reports which node, its package, and a pinned GitHub install hint (the script never runs `git clone` itself).\n2. **Missing models** — substitutes a compatible model (e.g. SD3.5 medium variants → SD3.5 large).\n3. **Non-critical utility nodes** (e.g. `UnloadAllModels`) — bypassed automatically without interrupting generation.\n4. **Runtime fallback chain** — `t2i`/`i2i` auto-descend through alternate models (and, for t2i, any available checkpoint) on runtime failures; config problems (missing files/nodes) abort instead of degrading. Disable with `--no-fallback`.\n\n## Output Delivery\n\n**Send the generated file — do not just describe it.** Deliver in the **user's original session**, never raw paths/URLs. Per-channel prefixes and the staging/cleanup routine → [references/delivery.md](references/delivery.md)\n\n**Output format:** generated images default to **JPG** (`config.json` → `image.output_format`). **Routing rule: if the user asks for a PNG image/file, or an RGBA image with transparency/alpha, always pass `--format png` and deliver the PNG file as-is — never convert to JPG** (JPEG has no alpha channel; the conversion would flatten transparency onto white). PNG output passes the server's PNG bytes through unchanged, so transparency is preserved byte-for-byte. Otherwise `--format jpg --jpg-quality N` is available per run; JPG conversion happens client-side (ComfyUI always saves PNG server-side). **i2m delivery rules (3D output):** (1) attach the **extracted character image** from step 1 as an **image** — never just its path (with `--extract none` the step-1 image *is* the user's own input, so there's nothing to send back); (2) send the mesh **packaged as a zip** — the script first **scales the GLB in place by the delivery factor** (default ×3000; the raw mesh is at a normalized unit scale, character ≈ 2 units, no declared units), then zips it (`i2m_<timestamp>.zip`); attach the **zip**, never the raw `.glb`; (3) **never render the mesh** — no GLB → image/video conversion and no 3D-viewer screenshots (a third-party viewer link is fine to share). Format conversion never applies to 3D files.\n\n## Execution Discipline — Fast Path\n\nThe **measured GPU sampling time is ~12–20s**. When a task takes minutes, the delay is **not** inference — it's either a model cold-reload or agent-side overhead. Kill the waste:\n\n**i2m exception:** 3D generation is legitimately long (cold checkpoint load + 30 sampling steps + voxel decode ≈ minutes on a 4090). Run it once in the background (`exec` with a generous `timeoutSeconds`), report the queue state, and when it completes deliver the character image + the GLB zip per the i2m delivery rules above — don't spin tool calls around it.\n\n1. **Never re-read on repeat.** Don't re-read `SKILL.md`, `config.json`, or the script for a known config / repeat / same-model job. Defaults are baked in; `--aspect` / `--steps` / `--cfg` are in the links above.\n2. **At most one pre-check merge.** Gather any needed info in a single batched command, or skip it entirely for trivial repeat jobs.\n3. **Run immediately after one check.** Don't inspect, second-guess, or add tool calls between knowing the prompt and running.\n4. **Verify only the minimum.** After the run, check only what the user asked (size / seed / file exists).\n5. **Golden rule:** sampling is ~12–20s. Anything longer = reload or overhead — never pile on tool calls before the image exists.\n6. **Prompt-flag trap:** pass `--prompt \"...\"` on the first attempt; a positional word wastes a whole turn.\n\n## References (index)\n\n- [references/image-first.md](references/image-first.md) — image-then-edit detection rules + multi-image reference composition\n- [references/attribution.md](references/attribution.md) — original source + modifications per workflow\n- [references/cli-reference.md](references/cli-reference.md) — full commands + timeout reference\n- [references/prompts.md](references/prompts.md) — per-model prompt guidance\n- [references/error-handling.md](references/error-handling.md) — auto-handled error categories\n- [references/delivery.md](references/delivery.md) — per-channel delivery + staging\n\nFile v3.6.0:README.md\n\n# Image with ComfyUI\n\nCall a user-defined ComfyUI server (local or remote) for **text-to-image**, **image-to-image/edit**, **image-to-video**, and **image-to-3D-mesh** generation. **Qwen-Image 2.1** is the default model for both T2I and image editing (including multi-image reference composition); Z-Image, SD3.5 Medium, Qwen Image Edit (2511), Wan2.2, and Hunyuan 3D v2.1 (3D mesh) are available as options.\n\n> **New here?** Qwen-Image 2.1 needs a one-time ComfyUI setup — see [Qwen-Image 2.1 ComfyUI Setup](#qwen-image-21-comfyui-setup).\n\n## Quick Start\n\n1. **Ensure a ComfyUI server is running** (locally, or on a host you can reach).\n2. **Set the URL** (via environment variable or `config.json`). The endpoint is user-defined — the bundled default is `http://127.0.0.1:8188`; point `COMFYUI_URL` at any ComfyUI server you control — a data-flow warning is printed on every run for non-local hosts. If no server is reachable, the script exits with configuration instructions:\n   ```bash\n   export COMFYUI_URL=http://comfyui.host:api-port\n   ```\n3. **Test a workflow** (see [Testing](#testing) below).\n\n## Configuration\n\nRead `config.json` at the skill root. All values can be overridden by environment variables:\n\n| Env Variable | Overrides | Default |\n|---|---|---|\n| `COMFYUI_URL` | `comfyui_url` | `http://localhost:8188` |\n| `COMFYUI_TIMEOUT` | `timeout_seconds` | `120` |\n| `COMFYUI_POLL_INTERVAL` | `poll_interval_seconds` | `1` |\n| `COMFYUI_KEEP_MODELS_RESIDENT` | `keep_models_resident` | `false` |\n| `COMFYUI_OUTPUT_DIR` | `output_dir` | `media/comfyui` |\n| `COMFYUI_APIKEY` | `comfyui_api_key` | _(none → no auth header)_ |\n\n**Priority:** Environment variables > `config.json`.\n\n> Note: The config.json key `comfyui_api_key` (string) is optional. When empty/absent, **no** `Authorization` header is added and behavior is unchanged. When set, every ComfyUI API request carries `Authorization: Bearer <key>` (env var wins over the config value).\n\n### 🔒 ComfyUI API Key (Bearer Auth)\n\nIf your ComfyUI instance requires an API key, set it via either of these (env var takes precedence):\n\n```bash\nexport COMFYUI_APIKEY=sk-your-key\n```\n\nor, in `config.json`:\n\n```json\n\"comfyui_api_key\": \"sk-your-key\"\n```\n\nWhen present, the skill adds `Authorization: Bearer <key>` to all ComfyUI API calls (`/prompt`, `/history`, `/object_info`, `/upload/image`, `/view`, `/upload/*`). Do not commit your key to version control — prefer the env var, and never paste it into the chat or logs.\n\n> **⚠️ This skill only *sends* the key — it doesn't authenticate your ComfyUI itself.** For the Bearer key to be honored, the ComfyUI server must run an authentication custom node that recognizes that key. The reference implementation is [`sunshinejnjn/comfyui-auth`](https://github.com/sunshinejnjn/comfyui-auth) (a fork of the original [`ivellioscolin/comfyui-auth`](https://github.com/ivellioscolin/comfyui-auth)). To set it up:\n> 1. Clone it into your ComfyUI: `git clone https://github.com/sunshinejnjn/comfyui-auth.git ComfyUI/custom_nodes/comfyui-auth`\n> 2. Edit its `apikeys.conf` — add a `label = <your-key>` line (the key is read **as written**, plain or as configured), then restart ComfyUI (the plugin reloads its config per request, but a re-clone/restart makes the new code load).\n> 3. Point this skill at that server via `COMFYUI_URL` and set `COMFYUI_APIKEY` to that same key.\n>\n> Without the plugin installed on the server, the key is silently ignored and the API responds as if no auth is required. The `test` command (`python3 image_with_comfyui.py test`) will tell you whether the endpoint answers before you try generating.\n\n### Data & Privacy\n\n- **Transmitted to the endpoint**: prompts and workflow JSON (all modes); **full source images** for I2I and I2V.\n- **Endpoint** is user-defined (`COMFYUI_URL` env var > `comfyui_url` in `config.json`; bundled default `http://127.0.0.1:8188`). A non-local endpoint prints a warning on every run, and an unreachable server exits with configuration instructions (this section + [Qwen-Image 2.1 ComfyUI Setup](#qwen-image-21-comfyui-setup)). Use HTTPS for remote servers and never put credentials in the URL or logs.\n- **Written to disk**: generated media is saved under `output_dir` (config) or `COMFYUI_OUTPUT_DIR` in your workspace; delivery staging copies live in `~/.openclaw/media/outbound/` and are removed after a successful send. If your content is sensitive, review the endpoint and output location before running.\n\n## Models & Custom Nodes\n\nThis skill uses the **existing** ComfyUI installation — it does not ship models or nodes. Before using, ensure your ComfyUI has:\n\n### Required Models\n\nEvery model below has a **verified download source** in [`MODEL_URL.txt`](MODEL_URL.txt) — open that file for the exact HuggingFace repo, folder, and filename. Our server currently runs a **UC/NSFW-merged set** (Qwen-Image 2.1 = uncensored, Wan2.2 = NSFW creator merges); `MODEL_URL.txt` lists **SFW/official equivalents** for a clean setup.\n\n| Model | Expected Location | Source |\n|---|---|---|\n| Z-Image (SVD XT) | `ComfyUI/models/checkpoints/` | [`MODEL_URL.txt`](MODEL_URL.txt) |\n| SD3.5 Medium | `ComfyUI/models/checkpoints/` | [`MODEL_URL.txt`](MODEL_URL.txt) |\n| Qwen Image Edit Plus | `ComfyUI/models/checkpoints/` | [`MODEL_URL.txt`](MODEL_URL.txt) |\n| **Qwen-Image 2.1** (default T2I + edit) | see [Qwen-Image 2.1 ComfyUI Setup](#qwen-image-21-comfyui-setup) | [`MODEL_URL.txt`](MODEL_URL.txt) |\n| Wan2.2 1.3B | `ComfyUI/models/checkpoints/` | [`MODEL_URL.txt`](MODEL_URL.txt) |\n| Wan2.2 VAEMODEL | `ComfyUI/models/vae/` | [`MODEL_URL.txt`](MODEL_URL.txt) |\n| Wan2.2 Audio Model | `ComfyUI/models/audio_encoders/` | [`MODEL_URL.txt`](MODEL_URL.txt) |\n| T5XXL & FluxFill | `ComfyUI/models/clip/` | [`MODEL_URL.txt`](MODEL_URL.txt) |\n| **Hunyuan 3D v2.1** (I2M) | `ComfyUI/models/checkpoints/` | [`MODEL_URL.txt`](MODEL_URL.txt) |\n\n### Required Custom Nodes\n\nInstall **pinned at reviewed commits** (as of 2026-09-20). Check out the exact commit rather than a mutable default branch, and install dependencies in an isolated venv under a non-privileged user — never `pip install` as root:\n\n| Node | Package (repo) | Pinned commit |\n|---|---|---|\n| Impact Pack | [ltdrdata/ComfyUI-Impact-Pack](https://github.com/ltdrdata/ComfyUI-Impact-Pack) | `429d0159a` |\n| WAS Nodes | [WASasquatch/was-node-suite-comfyui](https://github.com/WASasquatch/was-node-suite-comfyui) *(repo renamed from `ComfyUI-WAS-Nodes`)* | `9934caa92` |\n| ComfyUI-Manager | [Comfy-Org/ComfyUI-Manager](https://github.com/Comfy-Org/ComfyUI-Manager) *(migrated from `comfyanonymous/`)* | `946ef8fe6` |\n| rgthree | [rgthree/rgthree-comfy](https://github.com/rgthree/rgthree-comfy) | — (required by the Qwen-Image 2.1 workflow's LoRA pass-through node) |\n\n```bash\ncd ComfyUI/custom_nodes\ngit clone https://github.com/ltdrdata/ComfyUI-Impact-Pack.git\ncd ComfyUI-Impact-Pack && git checkout --detach 429d0159ad429e64d2b3916e6e7be9c22d025c3c\npython3 -m venv .venv && .venv/bin/pip install -r requirements.txt\n```\n\nRepeat the same pin-and-checkout pattern for the other two repositories (use their pinned commits above). Review dependency lists before installing, and re-check upstream commits before upgrading a pin.\n\nWhen a missing node or model is detected, the script reports the package name and a manual install hint — it never executes `git clone` itself (ComfyUI may run on another host).\n\n## Qwen-Image 2.1 ComfyUI Setup\n\n**Qwen-Image 2.1** (model id `qwen21`) is the **default** model for both `t2i` and `i2i`. One workflow covers text-to-image, single-image edit, and multi-image reference composition. It is a distilled model — **CFG must stay at 1.0** (the script enforces this; raising it burns the image).\n\n### Prerequisites\n- **ComfyUI ≥ 0.37.0** (frontend 1.53.6). The workflow uses core nodes that only exist in recent builds:\n  - `TextEncodeQwenImage21` (`comfy_extras.nodes_qwen`)\n  - `ResolutionSelector` (`comfy_extras.nodes_resolution`)\n  - `ComfySwitchNode` / `If/Else Switch` (`comfy_extras.nodes_logic`)\n  - `PrimitiveBoolean` (`comfy_extras.nodes_primitive`)\n  \n  If any show up as a red *missing node*, **update ComfyUI** — they are core, not a custom pack. (`If/Else Switch` is experimental and may be hidden from node search until you enable experimental nodes; loading the workflow works either way.)\n\n### Model files (3)\nPlace each in the folder shown, under your ComfyUI install. These are the exact filenames the workflow expects:\n\n| File | Destination folder |\n|---|---|\n| `QWEN_Image/qwen-image-2.1-UC-int8_convrot.safetensors` | `ComfyUI/models/diffusion_models/` (keep the `QWEN_Image/` subfolder; older builds may use `models/unet/`) |\n| `qwen3vl_8b_int8_convrot.safetensors` | `ComfyUI/models/text_encoders/` (Qwen3-VL 8B int8; loaded with `type: qwen_image`) |\n| `qwen_image_2.1_vae_bf16.safetensors` | `ComfyUI/models/vae/` |\n\n> Unlike the original foprc workflow (which used a GGUF Q8 model + ComfyUI-GGUF), this workflow loads the diffusion model via **`UNETLoader`** with `weight_dtype: fp8_e4m3fn_fast`. That **removes the ComfyUI-GGUF dependency** — the only custom node needed is **rgthree**.\n\nIf you use a different quantization/filename, update `unet_name` in `workflows/qwen_image_2.1_api.json` (node `17`) to match what's on disk.\n\n### Custom nodes\n- **rgthree-comfy** (for the `Power Lora Loader` node) — install via ComfyUI Manager or `git clone https://github.com/rgthree/rgthree-comfy` into `custom_nodes/`, then restart.\n- All other nodes are core ComfyUI.\n- *Optional:* the LoRA stack is off by default and passes the model/CLIP straight through. If you don't have rgthree, you can instead delete node `10` and rewire `17 → 6` (model) and `2 → 5` (clip).\n\n### Verifying your setup\n```bash\npython3 image_with_comfyui.py test          # server reachable?\npython3 image_with_comfyui.py t2i --prompt \"A red fox in snow\" --aspect 1:1\npython3 image_with_comfyui.py i2i --prompt \"Change the sky to sunset\" --image /path/in.jpg\n```\nIf a loader dropdown is blank in the web UI, the file isn't where ComfyUI is looking — check the folder above. If a node reports *missing*, update ComfyUI (see Prerequisites).\n\n### How the dual-mode switch works\nNode `12` (`PrimitiveBoolean`) is the master switch:\n- **OFF** (`false`) → pure **T2I**: the sampler starts from an empty latent sized by `ResolutionSelector` (aspect ratio + megapixels).\n- **ON** (`true`) → **edit / multi-image**: reference image(s) feed `TextEncodeQwenImage21`, whose output latent is sampled instead. The script wires 1 image to `image_1`, and extra images to `image_2 … image_N` (up to 16).\n\n### Multi-image prompt format\nWith 2+ reference images, the model sees them in **slot order** (`image_1`, `image_2`, …). Reference each by index in the prompt so it knows which is which:\n```\nPut the garment from <image2> onto the person in <image1>, keep <image1>'s pose and lighting.\nCombine both photos into one scene: subject of <image1> on the left, subject of <image2> on the right.\n```\nKeep the instruction concise and in the user's language; the skill passes it through verbatim. `--image` accepts multiple paths (order = slot order = `<image1>`, `<image2>` …).\n\n## Testing\n\nUse the built-in test script:\n\n```bash\n# Health check\npython3 image_with_comfyui.py test\n\n# Test each workflow individually\npython3 image_with_comfyui.py t2i --prompt \"A cat on a windowsill\"\n\n# Run the health check before each batch of jobs\npython3 image_with_comfyui.py test\n```\n\n### Output filename scheme\n\nGenerated files use a short mode prefix + timestamp (`YYYYMMDD_HHMMSS`) + ComfyUI's counter:\n\n| Mode | Prefix | Example (JPG) |\n|---|---|---|\n| Qwen-Image 2.1 T2I | `qt2i` | `qt2i_20260928_050328_00001.jpg` |\n| Qwen-Image 2.1 I2I | `qi2i` | `qi2i_20260928_050328_00001.jpg` |\n| Qwen Image Edit (2511) | `qei2i` | `qei2i_20260928_050328_00001.jpg` |\n| Z-Image T2I | `zt2i` | `zt2i_20260928_050328_00001.jpg` |\n| SD3.5 Medium T2I | `st2i` | `st2i_20260928_050328_00001.jpg` |\n| Character cutout (i2m step 1) | `mcut` | `mcut_20260928_050328_00001.png` |\n| Hunyuan 3D v2.1 i2m | `i2m` | `i2m_20260928_050328_00001.glb` |\n| i2m delivery package | `i2m` | `i2m_20260928_050328_00001.zip` (the ×3000-scaled GLB inside) |\n\nThe trailing underscore ComfyUI appends is stripped, so names never end in `_`. The i2m GLB is **scaled ×3000 in place, then packaged as `<same-stem>.zip`** — the `.glb` stays on disk, but what gets sent to the user is the **zip** (delivery rules: SKILL.md → Output Delivery; the character image is also sent, as an image, and the mesh is never rendered).\n\n### Manual Workflow Testing\n\nYou can also test workflows directly in the ComfyUI web UI. Copy any workflow file from `workflows/` into ComfyUI, load it, and run:\n\n- **Qwen-Image 2.1 (T2I + edit + multi-image, default)**: `workflows/qwen_image_2.1_api.json`\n- **Z-Image T2I**: `workflows/z-image_t2i_api.json`\n- **SD3.5 Medium T2I**: `workflows/sd3.5-med_t2i_api.json`\n- **Qwen Image Edit (2511)**: `workflows/qwen_image-edit_api.json`\n- **Wan2.2 I2V**: `workflows/wan2.2_i2v_api.json`\n- **Hunyuan 3D v2.1 I2M (image → 3D mesh)**: `workflows/hunyuan3d-v2.1_i2m_api.json`\n\n## Important Notes\n\n### Local Skill — No Automatic Installation\n\nThis skill **does not install anything automatically**. It connects to your existing ComfyUI instance via API calls; the workflows, models, and custom nodes must already be configured in your ComfyUI setup. If a node is missing, the skill only *reports* the package and a pinned manual install hint (it never runs `git clone` itself, since ComfyUI may be on another host). The manual commands in [Required Custom Nodes](#required-custom-nodes) are optional setup steps for your own ComfyUI — review and pin them before running.\n\n### VRAM Release vs. Speed Tradeoff\n\nBy default this skill uses the `UnloadAllModels` node (and `UnloadModel` in the Qwen edit branch) to free VRAM after each generation, preventing VRAM exhaustion during repeated tasks. The unload nodes are stripped only when `keep_models_resident` is explicitly set to `true` (config or `COMFYUI_KEEP_MODELS_RESIDENT`); otherwise they run and release VRAM on every run. The unload nodes are bypassed automatically if a given ComfyUI instance doesn't have them — the system detects and skips them transparently.\n\n> **The slow part of a “5-minute” run is almost never inference.** It's model reload *or* agent-side overhead before the run, not the sampling. See the skill's **Execution Discipline — Fast Path** section for how the run is expected to be invoked.\n\n### Workflow Modifications\n\nAll workflow files in the `workflows/` directory are **adapted versions** of original ComfyUI workflows. They are repackaged into self-contained HTTP-API graphs — the original source is **not** kept inside the JSON. The original source for each workflow plus exactly what was changed is documented in [references/attribution.md](references/attribution.md). Known modifications include:\n- API-ready output node placement + timestamped output filenames\n- Parameter exposure (aspect, steps, CFG, seed, neg, etc.) for external control\n- Prompt auto-routing (e.g. Qwen Edit routes to the positive prompt node)\n- A standard **safetensors** model load replacing the original GGUF load (Qwen-Image 2.1), dropping the ComfyUI-GGUF dependency\n- VRAM-cleanup nodes (`UnloadAllModels`, `UnloadModel`) added, auto-bypassed when absent\n- The SD3.5 double-`.safetensors` suffix bug fixed; Wan2.2 pointed at 4-step Lightspeed weights\n\n### Error Handling\n\nThe system handles three categories of errors automatically:\n\n1. **Missing nodes** — Reports which node is missing, provides package name, install command, and GitHub URL.\n2. **Missing models** — Attempts to substitute a compatible model (e.g., SD3.5 medium variants → SD3.5 large).\n3. **Non-critical utility nodes** — Automatically bypasses nodes like `UnloadAllModels` without interrupting generation.\n\n## CLI Reference\n\n```bash\n# Generate image from text (default: Qwen-Image 2.1)\npython3 image_with_comfyui.py t2i --prompt \"Your description\" --aspect 1:1\n\n# Generate with an explicit model\npython3 image_with_comfyui.py t2i --model z-image --prompt \"Your description\"\npython3 image_with_comfyui.py t2i --model sd35 --prompt \"Your description\" --negative \"blurry\"\n\n# Edit an image (default: Qwen-Image 2.1)\npython3 image_with_comfyui.py i2i --prompt \"Replace background\" --image /path/to/image.png\n\n# Multi-image reference (2+ images; order = <image1>, <image2>, ...)\npython3 image_with_comfyui.py i2i \\\n  --prompt \"Put the shirt from <image2> on the person in <image1>\" \\\n  --image /path/person.jpg /path/shirt.jpg\n\n# Generate video from image\npython3 image_with_comfyui.py wan2.2 --prompt \"Camera pans left\" --image /path/to/image.png\n\n# Turn the character in an image into a 3D mesh (GLB)\n# Default: extract the main character first (i2i transparent cutout), then generate\npython3 image_with_comfyui.py i2m --image /path/to/character.jpg\n# --extract rembg (u2net cutout, no redraw) | --extract none (image used as-is)\n# Delivery: character image (as an image) + i2m_*.zip (as an attachment);\n# the GLB is scaled x3000 in place before zipping (--scale 1 to skip);\n# the mesh is never rendered (no GLB -> image/video conversion; viewer links OK)\n\n# Test ComfyUI connection\npython3 image_with_comfyui.py test\n```\n\n## Timeout Reference\n\n| Mode | Timeout |\n|---|---|\n| T2I (Qwen-Image 2.1) | 300s |\n| T2I (Z-Image) | 100s |\n| T2I (SD3.5) | 100s |\n| I2I (Qwen-Image 2.1, incl. multi-image) | 600s |\n| I2I (Qwen 2511) | 600s |\n| I2V (Wan2.2) | 1000s |\n| I2M (Hunyuan 3D v2.1) | 900s |\n\nFile v3.6.0:_meta.json\n\n{\n  \"ownerId\": \"kn7dn131pvgfaveyzzzcrc5xf182yb84\",\n  \"slug\": \"image-with-comfyui\",\n  \"version\": \"3.6.0\",\n  \"publishedAt\": 1791136664035\n}\n\nFile v3.6.0:references/attribution.md\n\n# Source Attribution (image-with-comfyui)\n\nEvery workflow in `workflows/` is an **adapted** version of a real ComfyUI\nworkflow. Original source + the specific modifications applied by this skill are\ndocumented here. The workflow JSONs carry only ComfyUI's standard per-node\nlabels (`_meta` `title`); they do **not** embed provenance or attribution\nmetadata, so the original source is not duplicated inside them — it lives here\ninstead, and each file stays runnable.\n\n| Workflow | Original source | Key modifications applied |\n|---|---|---|\n| `qwen_image_2.1_api.json` | [foprc/qwen-image-2.1-comfyui-workflow](https://github.com/foprc/qwen-image-2.1-comfyui-workflow) (ships GGUF Q8_0 + ComfyUI-GGUF) | Replaced GGUF model with standard **safetensors** via `UNETLoader` (`weight_dtype: fp8_e4m3fn_fast`) — removes the ComfyUI-GGUF dependency; added `UnloadModel` (node 18) + `UnloadAllModels` (node 19) VRAM-cleanup; dual-mode reference switch (nodes 12/13/14); `ResolutionSelector` (node 9); API/`timestamp` output |\n| `qwen_image-edit_api.json` | Qwen-Image-Edit (Comfy-Org) default edit graph | Prompt auto-routing to the positive `TextEncodeQwenImageEditPlus` node (115:111) by the script |\n| `sd3.5-med_t2i_api.json` | Comfy-Org / sd3-medium default graph | Fixed the `.safetensors.safetensors` double-suffix `ckpt_name` bug; `sd3_medium` → `sd3.5_large.safetensors` fallback; `UnloadAllModels` VRAM-cleanup |\n| `z-image_t2i_api.json` | Comfy-Org / z-image default graph | `UnloadAllModels` VRAM-cleanup |\n| `wan2.2_i2v_api.json` | Comfy-Org / Wan2.2 I2V default graph | Pointed the diffusion model at **4-step Lightspeed** weights; LoRA stack nodes (109/110); `UnloadAllModels` VRAM-cleanup |\n| `hunyuan3d-v2.1_i2m_api.json` | Tencent Hunyuan3D-2.1 ComfyUI image→mesh reference graph (user-supplied, [tencent/Hunyuan3D-2.1](https://huggingface.co/tencent/Hunyuan3D-2.1)) | Topology kept 1:1 (including the user's VRAM-unload sink nodes 19/20/21/22); placeholder `image` set to `character.png` (the script always rewrites it to the uploaded filename); default `filename_prefix` kept (`mesh/hy21`) — the script stamps it per run to `mesh/i2m_<timestamp>`; no model/parameter changes — resolution 4096 / 30 steps / CFG 5.0 are the graph's own defaults and stay the skill defaults\n\nAll six are **adapted** (as opposed to) the original; they add prompt routing,\nresolution selection, dual-mode switches, and VRAM management on top of the\noriginal ComfyUI graphs.\n\nFile v3.6.0:references/cli-reference.md\n\n# CLI Usage Reference (image-with-comfyui)\n\nFull command examples for every subcommand. The SKILL.md procedure only lists the\ncore rule; open this file for the exact flags.\n\n> **⚠️ The prompt is MANDATORY and must be passed via `--prompt` — never as a positional argument.**\n>\n> Every subcommand that takes a prompt (`t2i`, `i2i`, `wan2.2`) reads it from `--prompt \"...\"` (`i2m` has no prompt — see below). Passing the prompt as a bare positional word fails immediately with exit code 2:\n> `python3 image_with_comfyui.py t2i \"some description\"` → `t2i: error: the following arguments are required: --prompt`\n>\n> This is a **100%-avoidable failure loop** — the model can't help you because the command never reaches ComfyUI. Always write the prompt through the `--prompt` flag on the first attempt.\n\n### T2I (Text → Image)\n\n```bash\n# Qwen-Image 2.1 (default model; CFG locked to 1.0 — no --cfg)\n# --aspect is OPTIONAL: omit it and the output is 1:1 (1024x1024)\npython3 image_with_comfyui.py t2i \\\n  --prompt \"Your detailed image description\" \\\n  --steps 25\n\n# With a negative prompt (Qwen-Image 2.1 supports it)\npython3 image_with_comfyui.py t2i \\\n  --prompt \"A detailed image description\" \\\n  --negative \"text, watermark, blurry\" \\\n  --steps 25\n\n# Request a specific aspect ratio explicitly (16:9, 9:16, 4:3, ...)\npython3 image_with_comfyui.py t2i \\\n  --prompt \"Your detailed image description\" \\\n  --aspect 16:9 \\\n  --steps 25\n\n\n# Z-Image\npython3 image_with_comfyui.py t2i \\\n  --model z-image \\\n  --prompt \"Your detailed image description\" \\\n  --aspect 1:1 \\\n  --steps 9\n\n# SD3.5 Medium (defaults to 1:1)\npython3 image_with_comfyui.py t2i \\\n  --model sd35 \\\n  --prompt \"A beautiful sunset over mountains\" \\\n  --negative \"text, watermark, blurry\" \\\n  --steps 20 \\\n  --cfg 5.5\n\n# Run ONLY the requested model (skip the runtime fallback chain)\npython3 image_with_comfyui.py t2i \\\n  --prompt \"...\" --model z-image --no-fallback\n\n# JPG output (converted client-side; default from config.json image.output_format)\npython3 image_with_comfyui.py t2i \\\n  --prompt \"...\" --format jpg --jpg-quality 90\n```\n\n### I2I (Edit Image)\n\n```bash\n# Qwen-Image 2.1 (default I2I)\npython3 image_with_comfyui.py i2i \\\n  --prompt \"Change background to a beach\" \\\n  --image /path/to/source.jpg \\\n  --steps 25\n\n# Explicit aspect ratio: routes the I2I latent canvas through the\n# ResolutionSelector (workflow latent switch OFF) instead of matching the\n# first reference image's size, so a portrait photo CAN become 16:9 landscape.\n# The reference is uploaded at native size, never stretched. Without --aspect,\n# the output follows the first reference image's own size (no 1:1 forcing).\npython3 image_with_comfyui.py i2i \\\n  --prompt \"Extend the scene into a wide landscape\" \\\n  --image /path/to/portrait.jpg \\\n  --aspect 16:9\n\n# Multi-image reference (up to 16): --image order = slot order\npython3 image_with_comfyui.py i2i \\\n  --prompt \"put the shirt from <image2> on the person in <image1>\" \\\n  --image /path/to/base.jpg /path/to/reference.jpg\n\n# Legacy single-image edit\npython3 image_with_comfyui.py i2i \\\n  --model qwen_imageedit \\\n  --prompt \"put on a red jacket\" \\\n  --image /path/to/source.jpg \\\n  --steps 4\n\n# Qwen-Image 2.1 edit with a negative prompt\npython3 image_with_comfyui.py i2i \\\n  --prompt \"replace the background with a night city\" \\\n  --negative \"text, watermark\" \\\n  --image /path/to/source.jpg\n\n# Run ONLY the requested model (skip the runtime fallback chain)\npython3 image_with_comfyui.py i2i \\\n  --prompt \"...\" --image /path/to/source.jpg --no-fallback\n\n# JPG output (see t2i example)\npython3 image_with_comfyui.py i2i \\\n  --prompt \"...\" --image /path/to/source.jpg --format jpg\n```\n\n### I2V (Image → Video)\n\n```bash\npython3 image_with_comfyui.py wan2.2 \\\n  --prompt \"the person walks forward and smiles\" \\\n  --image /path/to/source.jpg \\\n  --length 81 --steps 4\n```\n\n\n### I2M (Image → 3D Mesh, Hunyuan 3D v2.1 → GLB)\n\n```bash\n# Default: extract the main character first (Qwen-Image 2.1 i2i redraw on a\n# transparent background, PNG), then generate the mesh.\npython3 image_with_comfyui.py i2m --image /path/to/character.jpg\n\n# Non-generative extraction: u2net human cutout (pixel-precise, no redraw)\npython3 image_with_comfyui.py i2m --image /path/to/character.jpg --extract rembg\n\n# Image is already a clean cutout — skip extraction entirely\npython3 image_with_comfyui.py i2m --image /path/to/cutout.png --extract none\n\n# Higher-detail / more-iterative generation\npython3 image_with_comfyui.py i2m --image /path/to/character.jpg \\\n  --resolution 8192 --steps 35 --cfg 6.0 --seed 42\n\n# Custom delivery scale (default: x3000; 1 = keep the normalized units)\npython3 image_with_comfyui.py i2m --image /path/to/character.jpg --scale 1\n```\n\n- Output: a **GLB file** under `output_dir/meshes/` (prefix `i2m_<timestamp>`),\n  **scaled in place by the delivery factor** (default ×3000 — the raw Hunyuan\n  3D v2.1 mesh is at a normalized unit scale, character ≈ 2 units, no declared\n  units; `--scale 1` keeps it unscaled) and **automatically packaged as\n  `<same-stem>.zip`** — the result block prints both paths plus the delivery hints.\n- **Delivery (i2m rules):** (1) send the extracted **character image** (the\n  step-1 PNG; the script prints it as `CHARACTER IMAGE`) **as an image** — with\n  `--extract none` there is no new image to send back. (2) send the **zip** (the\n  scaled GLB) as an attachment, never the raw `.glb`. (3) **Never render the\n  mesh** — no\n  GLB → image/video conversion, no 3D-viewer screenshots (a third-party\n  viewer *link* is fine).\n- Requires `hunyuan_3d_v2.1.safetensors` in the server's `models/checkpoints/`\n  and ComfyUI **≥ 0.37.0** (all Hunyuan 3D v2.1 nodes are core — no custom\n  packages).\n- This mode is the slow one by design: expect minutes, not ~20 s (see SKILL.md\n  \"i2m exception\"). Run it in the background and deliver on completion.\n\n### Health Check\n\n```bash\npython3 image_with_comfyui.py test\n```\n\n### Output Format (t2i / i2i)\n\nComfyUI's `SaveImage` node always writes PNG server-side. `--format jpg` converts the\n**downloaded** image client-side (Pillow) to JPEG — alpha channels are composited\nonto white, and conversion failures fall back to PNG automatically. `--format png`\npasses the server's PNG bytes through **unchanged** (no conversion), so RGBA\ntransparency survives intact. Defaults come from `config.json` → `image.output_format`\n(default `\"jpg\"`) and `image.jpg_quality` (default 89). Videos (wan2.2) are always\nMP4 and are not affected.\n\n**Routing rule:** the user's wording decides the format —\n\n- \"PNG image/file\", \"PNG format\", \"keep the alpha channel\", or **\"RGBA image with transparency\"** → always `--format png`, and **deliver the PNG file directly without any JPG/JPEG conversion** (never flatten transparency onto a background).\n- No format mentioned → use the config default (JPG).\n\n---\n\n## Timeout Reference\n\n| Model | Timeout |\n|---|---|\n| T2I (Qwen-Image 2.1) | 600s |\n| T2I (Z-Image) | 100s |\n| SD3.5 Medium | 100s |\n| I2I (Qwen 2.1) | 600s |\n| I2I (Qwen Image Edit 2511) | 600s |\n| I2V (Wan2.2) | 1000s |\n| I2M (Hunyuan 3D v2.1) | 900s |\n\nFile v3.6.0:references/delivery.md\n\n# Delivery Reference (image-with-comfyui)\n\nThe SKILL.md core says \"send the media, in the user's session, no paths\". This\nfile holds the per-channel delivery formats and the staging steps. Open it when\nyou need the exact prefix for a channel.\n\n## The one rule\n\nAfter generating, **send the file**, don't describe it. A text line like \"video\nattachment\" is NOT a send.\n\n## Per-channel format\n\n| Channel | Format | Example |\n|---|---|---|\n| WhatsApp | `MEDIA:` + absolute path | `MEDIA:/home/user/.openclaw/media/outbound/angel_video.mp4` |\n| Telegram | `MEDIA:` or `filePath:` | varies by implementation |\n| Discord | Direct attachment | varies by implementation |\n\n## Staging steps (WhatsApp)\n\n1. Copy the output to `~/.openclaw/media/outbound/` first.\n2. Send via `MEDIA:/home/user/.openclaw/media/outbound/<filename>` (use absolute file path).\n3. The `MEDIA:` line must be the **sole content** of the WhatsApp message — no\n   `[[reply_to_current]]`, no text before it — otherwise WhatsApp splits the file\n   and text into two messages. For a caption use `MEDIA:./file.ext caption=...`.\n4. **Clean up:** delete the staged file after a successful send so sensitive\n   generated media doesn't linger across sessions.\n\n## i2m (3D mesh) delivery\n\ni2m produces **two deliverables**, and both must go to the user:\n\n1. **The extracted character image** (step-1 PNG) — send it **as an image**\n   (an image message / attachment), never just its path. With `--extract none`\n   the step-1 image *is* the user's own input, so there is nothing new to send.\n2. **The mesh, as a zip** — the script first **scales the GLB in place by the\n   delivery factor** (default ×3000 — the raw Hunyuan 3D v2.1 mesh is at a\n   normalized unit scale, character ≈ 2 units, no declared units; `--scale 1`\n   keeps it unscaled), then packages it (`i2m_<timestamp>.glb` →\n   `i2m_<timestamp>.zip`). Send the **zip** as an attachment; never send the\n   raw `.glb` on its own.\n\n**Never render the mesh.** Do not convert the GLB to an image or video, or take\nscreenshots of a 3D viewer. A **third-party online GLB viewer link is allowed** —\nthe user views/rotates the mesh themselves; the zip + character image remain the\nprimary delivery.\n\n## Always\n\n- Deliver in the **user's original request session/thread** — not a separate\n  topic/group/thread unless explicitly told.\n- Never send raw local file paths, ComfyUI URLs, or host info unless asked.\n\nFile v3.6.0:references/error-handling.md\n\n# Error Handling Reference (image-with-comfyui)\n\nThe script auto-detects two categories of server-side problems (missing nodes and\nmissing models) and reports them with install/source guidance. Open this for the\nfull behavior; the SKILL.md procedure only tells you \"if you see ⚠️ report it\".\n\n## 1. Missing Node Detection\n\nWhen a workflow references a custom node that isn't installed, the system detects it and reports:\n- **Which node** is missing (class type name)\n- **Which package** provides it\n- **GitHub URL** for manual download\n- **Manual install instruction** (the script does NOT execute git clone; ComfyUI may be on a remote server)\n\nExample: If `ImpactKSamplerBasicPipe` is missing:\n```\n⚠️ Missing node: `ImpactKSamplerBasicPipe`\n📦 Package: ComfyUI-Impact-Pack\n🔗 GitHub: https://github.com/ltdrdata/ComfyUI-Impact-Pack\nℹ️ Install manually: see README \"Required Custom Nodes\" (pinned commit + venv; the script never runs git clone itself)\n```\n\n## 2. Missing Model Substitution\n\nWhen a workflow references a model file that doesn't exist, the system attempts to find a compatible substitute:\n\n| Requested Model | Substitute |\n|---|---|\n| `sd3.5_medium` variants | `sd3.5_large.safetensors` |\n| WAN High → Low or vice versa | Swap between variants |\n| Other unknown models | No substitution (error returned) |\n\nExample: If `my_custom_sd3_medium_v2.safetensors` is missing:\n```\n⚠️ Model missing: `my_custom_sd3_medium_v2.safetensors`\n🔄 Substituted: `sd3.5_large.safetensors`\n📦 Loader: CheckpointLoaderSimple.ckpt_name\n```\n\nAfter substitution, the workflow is retried automatically with the substitute model.\n\n## 3. Missing Utility Node Bypass (UnloadAllModels)\n\nWhen the workflow references `UnloadAllModels` (a memory cleanup node) which isn't available, the system **automatically bypasses it** by rerouting the signal path:\n- Removes the missing `UnloadAllModels` node\n- Redirects the upstream processing node directly to the downstream output node\n- Generation continues without interruption\n- User receives a warning about the bypass\n\nExample:\n```\n⚠️ Workflow missing node: `UnloadAllModels` (memory cleanup, non-critical)\n🔄 Auto-bypassed — generation continues\n```\n\n## 4. Runtime Fallback Chain (t2i / i2i)\n\n`t2i` and `i2i` commands run through a fallback engine (`_run_with_fallback`) instead of the single-model path. A run proceeds through ordered **tiers**:\n\n| Kind | Tier order |\n|---|---|\n| t2i | `qwen21` → `z-image` → `sd35` → (any remaining server checkpoint, standard topology) |\n| i2i | `qwen21` → `qwen_imageedit` |\n\n- The **requested model** (via `--model` or the config default) is always the first tier; the remaining tiers follow in the order above.\n- A tier fails on: send errors, server timeout (> `timeout_seconds`), `status_str == error`, or no output images.\n- **Config problems do NOT degrade** — missing model files or missing nodes abort the chain with `config problem (missing model file / node) — no fallback` (install/fix the model instead of silently switching).\n- **Server unreachable** mid-run is treated as an environment problem — every tier on the same server would fail identically, so the chain aborts.\n- **Server unreachable before a run:** `t2i` / `i2i` / `wan2.2` pre-check the user-defined endpoint (GET `/history`, 5 s timeout) and, if nothing answers, exit with configuration instructions (start a ComfyUI server, set `COMFYUI_URL` or `comfyui_url` in `config.json`, verify with `test`) instead of failing mid-run with a raw timeout.\n- **i2i never degrades blind:** a plain T2I model cannot honor the source image, so when the i2i tiers are exhausted the run exits with the available-model list for manual choice (no fallback to a text-only generation).\n- When a lower tier succeeds, the output includes a note: `🔄 used <tier> — fell back because <failed tiers> failed this run.`\n- `qwen_imageedit` skips the tier when the request has multiple input images (single-image model) and the chain continues.\n- **`--no-fallback`** (t2i / i2i): run only the requested model — original single-model behavior, no runtime descent.\n\nExample (qwen21 times out, z-image succeeds):\n\n```\n🎨 trying qwen21 ...\n⚠️ tier qwen21 — failed: server timeout (>600s)\n🎨 trying z-image ...\n✅ 1 image(s) saved ...\n🔄 used z-image — fell back because qwen21 failed this run.\n```\n\nFile v3.6.0:references/image-first.md\n\n# Image-First Mode Reference (image-with-comfyui)\n\nThe SKILL.md core says \"image-first: route remembered image + new text to\n`i2i`/`wan2.2`\". This file holds the full detection logic and worked examples.\nOpen it when a user sends an image and then a short edit request.\n\n## Detection\n\n1. The user sends **only an image** (no other text that turn).\n2. Within **2 minutes**, the user sends a **text** message that looks like an\n   edit or video request. The agent matches whichever language the user wrote —\n   CN or EN keywords, e.g. `修一下/fix`, `换背景/change background`,\n   `加特效/add an effect`, `变动画/turn into a video`, `改颜色/change color`.\n3. The text intent is **I2I** (edit the image) or **I2V** (animate the image).\n\n## Action\n\n- Route the **remembered latest image** + the **new text** to\n  `image_with_comfyui.py i2i` or `wan2.2`.\n- `--image <path>` = the latest image received. `--prompt` = the new text.\n- If unsure whether it's I2I or I2V, **default to I2I** unless the text clearly\n  says video/animation.\n- Do **NOT** ask the user for the image again — you already have it from the\n  previous turn.\n- **Scope:** only the immediately preceding image from the **same user, same\n  session**, within the 2-minute window. Never reuse images from other users,\n  other sessions, or older than 2 minutes.\n\n## Context tracking\n\n- Store the latest image media path (or URL) when no text is received.\n- Clear the stored image after it's used, or after 2 minutes of no new text.\n\n## Examples\n\n- `[image: a photo of a dog]` → (wait)\n- `[text: change the background to a beach]` → `i2i --image <path> --prompt \"change the background to a beach\"`\n- `[image: a cat sitting on a chair]` → (wait)\n- `[text: make it stand up and walk]` → `wan2.2 --image <path> --prompt \"the cat stands up and walks\"`\n\n## Multi-image reference (up to 16)\n\nWith 2+ reference images the model sees them in **slot order**\n(`image_1`, `image_2`, …). Reference each by index in the prompt:\n\n```\nPut the garment from <image2> onto the person in <image1>, keep <image1>'s pose and lighting.\n```\n\n`--image` accepts multiple paths in slot order: `--image /base.jpg /ref.jpg`\n= `<image1>` `/base.jpg`, `<image2>` `/ref.jpg`.\n\nFile v3.6.0:references/prompts.md\n\n# Prompt Formatting Reference (image-with-comfyui)\n\nPer-model guidance for writing prompts. This is reference material — the SKILL.md\nprocedure just says \"format your prompt; see below\". Prompt keywords are kept\nbilingual where they originally appeared bilingual (e.g. `blurry 模糊`).\n\n## Z-Image (T2I)\n\nZ-Image works best with **structured natural language prompts**, not keyword spam.\n\n**6-part formula:**\n```\nSubject + Scene + Composition + Lighting + Style + Constraints\n```\n\n**Rules:**\n- ✅ Use natural language sentences (not comma-separated tags)\n- ✅ Be specific about subject, camera, lighting, style\n- ❌ **NO negative prompts** — Z-Image Turbo ignores them completely\n- ❌ No weighted tags like `(word:1.2)`\n\n**Example:**\n```\nA young woman with long wavy blonde hair sits at a wooden café table,\nsteam rising from a ceramic cup. Shot from a 3/4 angle, close-up framing.\nSoft morning light filters through sheer curtains, casting warm golden tones.\nCinematic photography, shallow depth of field, Kodak Portra 400 aesthetic.\nNo text, no logos, photorealistic skin texture.\n```\n\n**Aspect ratios:** `1:1`, `4:3`, `3:4`, `16:9`, `9:16`, `3:2`, `2:3`\nShould be ommited (defaults to 1:1) unless specified by user or as an extracted requirement from user's requests.\nIn I2I situation, it should match the original image which the user asks to modify.\n\n---\n\n## SD3.5 Medium (T2I)\n\nSD3.5 Medium uses **natural language prompts** with optional negative prompts.\n\n**Prompt formula:**\n```\n[Composition/Angle] + [Subject] + [Scene/Environment] + [Lighting/Color] + [Style/Texture] + [Details]\n```\n\n**Rules:**\n- ✅ **Complete natural language sentences** — describe telling a human what to see\n- ✅ Subject first (model prioritizes early text)\n- ✅ Be specific about colors, materials, mood, atmosphere\n- ✅ Mixed CN/EN is fine (Chinese works better for Chinese scenes)\n- ✅ Use `--negative` for elements to exclude\n- ✅ Default 1:1 (1024×1024), use `--aspect` to change\n- ✅ Default 20 steps, CFG 4.01 (higher = stronger control)\n- ✅ Seed defaults to random; specify `--seed` for reproducibility\n- ❌ No comma-separated keyword spam (`beautiful, amazing, 4k`)\n- ❌ No weighted tags `(word:1.2)` — SD3.5 doesn't recognize them\n\n**Parameter recommendations:**\n- **CFG**: 4-7 (4.01 = softer, 5-7 = stronger control)\n- **Steps**: 20-25 (below 20 may lack detail)\n- **Negative prompt**: Highly effective in SD3.5\n\n**Common negative prompt words (mixed CN/EN — either or both work):**\n```\nblurry 模糊, low quality 低质量, pixelated 像素化, grainy 颗粒感,\noverexposed 过曝, underexposed 欠曝, flat lighting 平光,\ntext 文字, watermark 水印, logo 商标, signature 署名, caption 字幕,\npoorly drawn face 画坏的脸, deformed 变形, mutated 畸变, disfigured 毁容, extra limbs 多余肢体,\ncartoonish 卡通风(写实时用)\n```\n\n**Mixed CN/EN example — both work; for Chinese scenes the prompt may be written in Chinese:**\n```\n上海魔都春日花海 — 黄浦江畔，大片郁金香、樱花、油菜花盛开，繁花似锦；\n春日和煦阳光，远景陆家嘴三件套天际线；湿润的滨江步道倒映花影；\n低饱和胶片色调，文艺清新，广角视野。\n```\n\n**English example:**\n```\nCinematic photography, wide-angle shot of a bustling Tokyo street at night,\nneon signs reflecting on wet pavement, people with transparent umbrellas,\nmoody atmospheric lighting, deep blues and vibrant reds, street photography,\nshallow depth of field with bokeh background\n```\n\n---\n\n## Qwen-Image 2.1 (T2I default & I2I default)\n\n- T2I: same 6-part natural-language formula as Z-Image. **Optional `--negative`** (kept short, e.g. `text, watermark, blurry`). **CFG locked to 1.0** (distilled) — never pass `--cfg`.\n- I2I: concise prompts, user's original language; **optional `--negative`** allowed (concise, positive-only edits still work best). Multi-image: up to 16 refs, `--image` order = slot order, referenced as `<image1>`, `<image2>`, ...; with the reference switch ON the output follows the first reference's aspect ratio.\n\n## Qwen Image Edit (I2I) — Concise Prompts\n\nI2I prompts must be **concise and direct**. Keep the user's original language.\n\n**Rules:**\n- ✅ **Positive prompt only** — no negative prompts\n- ✅ Use user's exact words (don't translate or expand)\n- ✅ Concise (either language): 换红外套/put on a red jacket, 把背景换成蓝天白云/change the background to a blue sky with clouds, 把女孩换成男孩/turn the girl into a boy\n- ❌ Don't translate between languages\n- ❌ Don't over-explain or add details\n\n**Prompt routing fix (2026-04-22):**\n- The Qwen workflow has TWO `TextEncodeQwenImageEditPlus` nodes:\n  - `115:110` — empty negative prompt node\n  - `115:111` — positive prompt node (contains default text like \"the girl\")\n- Script must route prompts to **node 111 (positive)**, not node 110\n- The `prepare_i2i_workflow()` function auto-detects by scanning for existing default text\n\n---\n\n## Wan2.2 I2V (Image → Video)\n\nWan2.2 generates short videos (~5 seconds) from a static image + motion description.\n\n**Rules:**\n- ✅ Prompt describes **actions/movement** (not scene description)\n- ✅ English motion descriptions tend to give the best results; the user's own language is acceptable — do not force a translation the user didn't ask for\n- ✅ Focus on \"who does what\" and \"how the camera moves\"\n- ❌ Don't describe static scene elements in motion prompt\n\n- Default: 81 frames (~5s @ 16fps), 4 steps, CFG 4.5\n- Base resolution: **560×720** (3:4, fast and OK quality)\n- Auto-detect input image aspect ratio and select reference resolution:\n\n### Resolution Reference\n\n| Aspect | Fast & OK | User Fav | WAN 2.2 Native |\n|--------|-----------|----------|----------------|\n| 3:4 | 560×720 | 720×912 | 848×1088 |\n| 2:3 | 528×768 | 656×960 | 784×1136 |\n| 9:16 | 480×848 | 608×1072 | 720×1264 |\n\nOther available resolutions:\n- **3:4**: 416×544, 672×864, 784×1008\n- **2:3**: 384×576, 624×912, 736×1072\n- **9:16**: 368×624, 576×1008, 672×1184\n\n**Examples:**\n```\nprompt: \"the cat walks forward and looks at the camera, tail wagging\"\nprompt: \"the girl smiles and turns her head, wind blowing her hair\"\nprompt: \"the person stands in a busy street, camera pans left and slowly zooms in, cars driving, red flag fluttering\"\n```\n\nFile v3.6.0:skill-card.md\n\n## Description:\n\nGuides an agent in generating and editing images, animating images into video, and creating 3D meshes through a user-configured ComfyUI server.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[sunshinejnjn](https://clawhub.ai/user/sunshinejnjn)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers and other ComfyUI users can ask an agent to create or edit images, extract subjects onto transparent backgrounds, turn images into videos, or package character meshes for delivery.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: Prompts, workflow data, source images, and any configured API key are sent to the selected ComfyUI server.\n\nMitigation: Use a server you control or trust, prefer localhost or trusted HTTPS, and avoid sensitive inputs unless its handling is acceptable.\n\nRisk: Generated media may be stored in local output or delivery locations.\n\nMitigation: Check the configured output location before processing sensitive media.\n\n## Reference(s):\n\n- [Publisher profile](https://clawhub.ai/user/sunshinejnjn)\n- [ComfyUI skill setup and usage](artifact/README.md)\n- [Model download sources](artifact/MODEL_URL.txt)\n- [Workflow source attribution](artifact/references/attribution.md)\n\n## Skill Output:\n\n**Output Type(s):** [Guidance, Shell commands, Generated media files]\n\n**Output Format:** [Text guidance and attached JPG or PNG images, video files, or ZIP archives containing GLB meshes]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [Transparent image cutouts require PNG; 3D meshes are delivered as ZIP archives.]\n\n## Skill Version(s):\n\n3.6.0 (source: ClawHub release metadata; bundled frontmatter says 3.5.0)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nFile v3.6.0:config.json\n\n{\n  \"comfyui_url\": \"http://127.0.0.1:8188\",\n  \"comfyui_api_key\": \"\",\n  \"poll_interval_seconds\": 1,\n  \"keep_models_resident\": false,\n  \"output_dir\": \"media/comfyui\",\n  \"image\": {\n    \"output_format\": \"jpg\",\n    \"jpg_quality\": 89,\n    \"t2i\": {\n      \"z-image\": {\n        \"workflow_path\": \"workflows/z-image_t2i_api.json\",\n        \"timeout_seconds\": 100,\n        \"default_width\": 1280,\n        \"default_height\": 720,\n        \"default_steps\": 9,\n        \"output_subdir\": \"images\"\n      },\n      \"sd35\": {\n        \"workflow_path\": \"workflows/sd3.5-med_t2i_api.json\",\n        \"timeout_seconds\": 100,\n        \"default_width\": 1024,\n        \"default_height\": 1024,\n        \"default_steps\": 20,\n        \"default_cfg\": 4.01,\n        \"output_subdir\": \"images\"\n      },\n      \"qwen21\": {\n        \"workflow_path\": \"workflows/qwen_image_2.1_api.json\",\n        \"timeout_seconds\": 300,\n        \"default_aspect\": \"1:1\",\n        \"default_steps\": 25,\n        \"default_cfg\": 1.0,\n        \"output_subdir\": \"images\"\n      }\n    },\n    \"i2i\": {\n      \"qwen_imageedit\": {\n        \"workflow_path\": \"workflows/qwen_image-edit_api.json\",\n        \"timeout_seconds\": 600,\n        \"default_steps\": 4,\n        \"output_subdir\": \"images\"\n      },\n      \"qwen21\": {\n        \"workflow_path\": \"workflows/qwen_image_2.1_api.json\",\n        \"timeout_seconds\": 600,\n        \"default_aspect\": \"1:1\",\n        \"default_steps\": 25,\n        \"default_cfg\": 1.0,\n        \"output_subdir\": \"images\"\n      }\n    }\n  },\n  \"mesh\": {\n    \"hunyuan3d_v2.1\": {\n      \"workflow_path\": \"workflows/hunyuan3d-v2.1_i2m_api.json\",\n      \"timeout_seconds\": 900,\n      \"default_resolution\": 4096,\n      \"default_steps\": 30,\n      \"default_cfg\": 5.0,\n      \"default_scale\": 3000,\n      \"output_subdir\": \"meshes\",\n      \"extract\": {\n        \"enabled\": true,\n        \"workflow\": \"i2i\",\n        \"model\": \"qwen21\",\n        \"aspect\": \"1:1\",\n        \"prompt_suffix\": \"isolate the main character, full body, centered, plain transparent background, nothing else\"\n      }\n    }\n  },\n  \"video\": {\n    \"wan2.2_i2v\": {\n      \"workflow_path\": \"workflows/wan2.2_i2v_api.json\",\n      \"timeout_seconds\": 1000,\n      \"default_width\": 528,\n      \"default_height\": 768,\n      \"default_length\": 81,\n      \"default_steps\": 4,\n      \"default_cfg\": 4.5,\n      \"frame_rate\": 16,\n      \"output_subdir\": \"video\",\n      \"keep_interpolated_only\": false\n    }\n  },\n  \"default_model\": {\n    \"t2i\": \"qwen21\",\n    \"i2i\": \"qwen21\"\n  }\n}\n\nFile v3.6.0:workflows/hunyuan3d-v2.1_i2m_api.json\n\n{\n  \"1\": {\n    \"inputs\": {\n      \"ckpt_name\": \"hunyuan_3d_v2.1.safetensors\"\n    },\n    \"class_type\": \"ImageOnlyCheckpointLoader\",\n    \"_meta\": {\n      \"title\": \"Load Checkpoint Image Only (Hunyuan 3D v2.1)\"\n    }\n  },\n  \"2\": {\n    \"inputs\": {\n      \"image\": \"character.png\"\n    },\n    \"class_type\": \"LoadImage\",\n    \"_meta\": {\n      \"title\": \"Load Image\"\n    }\n  },\n  \"3\": {\n    \"inputs\": {\n      \"shift\": 1,\n      \"sampling\": \"flow\",\n      \"model\": [\n        \"1\",\n        0\n      ]\n    },\n    \"class_type\": \"ModelSamplingAuraFlow\",\n    \"_meta\": {\n      \"title\": \"ModelSamplingAuraFlow\"\n    }\n  },\n  \"4\": {\n    \"inputs\": {\n      \"resolution\": 4096,\n      \"batch_size\": 1\n    },\n    \"class_type\": \"EmptyLatentHunyuan3Dv2\",\n    \"_meta\": {\n      \"title\": \"Empty Hunyuan 3D v2 Latent\"\n    }\n  },\n  \"6\": {\n    \"inputs\": {\n      \"clip_vision_output\": [\n        \"22\",\n        0\n      ]\n    },\n    \"class_type\": \"Hunyuan3Dv2Conditioning\",\n    \"_meta\": {\n      \"title\": \"Hunyuan3Dv2Conditioning\"\n    }\n  },\n  \"7\": {\n    \"inputs\": {\n      \"seed\": 384326697438209,\n      \"steps\": 30,\n      \"cfg\": 5,\n      \"sampler_name\": \"euler\",\n      \"scheduler\": \"normal\",\n      \"denoise\": 1,\n      \"model\": [\n        \"3\",\n        0\n      ],\n      \"positive\": [\n        \"6\",\n        0\n      ],\n      \"negative\": [\n        \"6\",\n        1\n      ],\n      \"latent_image\": [\n        \"4\",\n        0\n      ]\n    },\n    \"class_type\": \"KSampler\",\n    \"_meta\": {\n      \"title\": \"KSampler\"\n    }\n  },\n  \"8\": {\n    \"inputs\": {\n      \"num_chunks\": 8000,\n      \"octree_resolution\": 256,\n      \"samples\": [\n        \"19\",\n        0\n      ],\n      \"vae\": [\n        \"1\",\n        2\n      ]\n    },\n    \"class_type\": \"VAEDecodeHunyuan3D\",\n    \"_meta\": {\n      \"title\": \"Hunyuan 3D VAE Decode\"\n    }\n  },\n  \"9\": {\n    \"inputs\": {\n      \"algorithm\": \"surface net\",\n      \"threshold\": 0.6,\n      \"voxel\": [\n        \"21\",\n        0\n      ]\n    },\n    \"class_type\": \"VoxelToMesh\",\n    \"_meta\": {\n      \"title\": \"Voxel to Mesh\"\n    }\n  },\n  \"10\": {\n    \"inputs\": {\n      \"filename_prefix\": \"mesh/hy21\",\n      \"mesh\": [\n        \"20\",\n        0\n      ]\n    },\n    \"class_type\": \"SaveGLB\",\n    \"_meta\": {\n      \"title\": \"Save 3D Model\"\n    }\n  },\n  \"13\": {\n    \"inputs\": {\n      \"crop\": \"center\",\n      \"clip_vision\": [\n        \"1\",\n        1\n      ],\n      \"image\": [\n        \"2\",\n        0\n      ]\n    },\n    \"class_type\": \"CLIPVisionEncode\",\n    \"_meta\": {\n      \"title\": \"CLIP Vision Encode\"\n    }\n  },\n  \"19\": {\n    \"inputs\": {\n      \"value\": [\n        \"7\",\n        0\n      ],\n      \"model\": [\n        \"1\",\n        0\n      ]\n    },\n    \"class_type\": \"UnloadModel\",\n    \"_meta\": {\n      \"title\": \"UnloadModel\"\n    }\n  },\n  \"20\": {\n    \"inputs\": {\n      \"value\": [\n        \"9\",\n        0\n      ]\n    },\n    \"class_type\": \"UnloadAllModels\",\n    \"_meta\": {\n      \"title\": \"UnloadAllModels\"\n    }\n  },\n  \"21\": {\n    \"inputs\": {\n      \"value\": [\n        \"8\",\n        0\n      ],\n      \"model\": [\n        \"1\",\n        2\n      ]\n    },\n    \"class_type\": \"UnloadModel\",\n    \"_meta\": {\n      \"title\": \"UnloadModel\"\n    }\n  },\n  \"22\": {\n    \"inputs\": {\n      \"value\": [\n        \"13\",\n        0\n      ],\n      \"model\": [\n        \"1\",\n        1\n      ]\n    },\n    \"class_type\": \"UnloadModel\",\n    \"_meta\": {\n      \"title\": \"UnloadModel\"\n    }\n  }\n}\n\nArchive v3.5.1: 19 files, 66219 bytes\n\nFiles: config.json (2438b), image_with_comfyui.py (110332b), MODEL_URL.txt (7015b), README.md (17687b), references/attribution.md (2476b), references/cli-reference.md (7159b), references/delivery.md (2435b), references/error-handling.md (4382b), references/image-first.md (2231b), references/prompts.md (6357b), skill-card.md (2039b), SKILL.md (11518b), workflows/hunyuan3d-v2.1_i2m_api.json (3286b), workflows/qwen_image_2.1_api.json (3702b), workflows/qwen_image-edit_api.json (4241b), workflows/sd3.5-med_t2i_api.json (2063b), workflows/wan2.2_i2v_api.json (8022b), workflows/z-image_t2i_api.json (2674b), _meta.json (137b)\n\nFile v3.5.1:SKILL.md\n\n---\nname: image-with-comfyui\ndescription: 'Use to generate, edit, or animate images and videos through a user-defined ComfyUI server (COMFYUI_URL env var or comfyui_url in config.json; if no server is reachable the script exits with configuration instructions) — trigger with requests like \"make a picture of\", \"replace the background\", \"turn this into a video\", \"edit my photo\", or \"make a 3D mesh of this character\". Not for simple non-generation image tweaks. **Qwen-Image 2.1** is the default model for both text-to-image (T2I) and image editing/multi-image (I2I); Z-Image / SD3.5 (T2I), Qwen Image Edit (2511, single-image), Wan2.2 (I2V), and Hunyuan 3D v2.1 (I2M, image → 3D GLB mesh) are available on request.'\nversion: \"3.5.0\"\nlicense: BSD-3-Clause\ncompatibility: 'ComfyUI ≥ 0.37.0; requires python3; the endpoint is user-defined (COMFYUI_URL or config.json) — point it at a server you control and trust.'\nallowed-tools: Read, Write, Edit, Exec\nmetadata:\n  openclaw:\n    emoji: \"🎨\"\n    requires:\n      anyBins: [\"python3\"]\n    config:\n      path: \"config.json\"\n---\n\n# Image with ComfyUI\n\nCall a user-defined ComfyUI server to generate or edit images and videos, or turn an image into a 3D mesh. **Qwen-Image 2.1** is the default model for both T2I and I2I; Z-Image / SD3.5 (T2I), Qwen Image Edit 2511 (single-image edit), Wan2.2 (I2V), and Hunyuan 3D v2.1 (I2M, image → 3D GLB) are available on request. If no ComfyUI server is reachable, the script exits with configuration instructions (see [README.md](README.md)).\n\n- **T2I** (Text → Image) → **Qwen-Image 2.1** (default) or Z-Image / SD3.5 Medium\n- **I2I** (Image → Image / Edit / Multi-image, up to 16 refs) → **Qwen-Image 2.1** (default) or Qwen Image Edit (2511, single image)\n- **I2V** (Image → Video) → Wan2.2 model\n- **I2M** (Image → 3D Mesh) → **Hunyuan 3D v2.1** — extracts the main character first (transparent PNG), generates a GLB, scales it ×3000, and zips it for delivery\n\n`config.json` at the skill root holds every default; any value can be overridden by an env var (see the Data & Privacy table there).\n\n## When to Use\n\n- Generate images from text → **t2i**\n- Edit one image, or compose 2–16 reference images → **i2i** (`--image` order = slot order = `<image1>`, `<image2>` …)\n- Turn an image into a short video → **wan2.2**\n- Turn the character in an image into a **3D mesh (GLB)** → **i2m** (Hunyuan 3D v2.1; the main character is extracted first, see below)\n- The user sends an image, then within 2 minutes asks to edit or animate it → **image-first mode**, see [references/image-first.md](references/image-first.md)\n\n## Data & Privacy (read before running)\n\n- **What is transmitted:** prompts and the workflow JSON go to the configured endpoint for **every** mode; **I2I, I2V and I2M additionally upload the full source image**.\n- **Endpoint** is user-defined (`COMFYUI_URL` env var > `comfyui_url` in `config.json`; bundled default `http://127.0.0.1:8188`). A non-local host prints a data-flow warning on every run; if no server answers, the script exits with configuration instructions instead of a raw timeout.\n- **Written to disk:** generated media is saved under `output_dir` (`config.json` or `COMFYUI_OUTPUT_DIR`). Don't submit sensitive content if the endpoint or disk is a concern.\n- **Secrets:** never put credentials in the URL or logs.\n\n## Workflow Files\n\n| Mode | Workflow | File |\n|---|---|---|\n| T2I (**Qwen-Image 2.1, default**) | Qwen-Image 2.1 dual-mode | `workflows/qwen_image_2.1_api.json` |\n| I2I (**Qwen-Image 2.1, default**, incl. multi-image) | same file, switch ON | `workflows/qwen_image_2.1_api.json` |\n| T2I (Z-Image) | Z-Image | `workflows/z-image_t2i_api.json` |\n| T2I (SD3.5) | SD3.5 Medium | `workflows/sd3.5-med_t2i_api.json` |\n| I2I (legacy) | Qwen Image Edit (2511) | `workflows/qwen_image-edit_api.json` |\n| I2V | Wan2.2 Image-to-Video | `workflows/wan2.2_i2v_api.json` |\n| I2M (Image → 3D Mesh) | Hunyuan 3D v2.1 | `workflows/hunyuan3d-v2.1_i2m_api.json` |\n\nEach workflow is an **adapted** version of a real ComfyUI graph. Original source + exactly what was changed → [references/attribution.md](references/attribution.md).\n\n## CLI Usage\n\n> **⚠️ The prompt is MANDATORY and must be passed via `--prompt` — never as a positional word.**\n>\n> Every subcommand (`t2i`, `i2i`, `wan2.2`) reads the prompt from `--prompt \"...\"`. Passing the prompt as a bare positional word fails immediately with exit code 2:\n> `python3 image_with_comfyui.py t2i \"some description\"` → `t2i: error: the following arguments are required: --prompt`\n>\n> This is a **100%-avoidable failure loop** — the command never reaches ComfyUI, so the model can't help. Always write the prompt through `--prompt` on the **first** attempt.\n>\n> `i2m` takes no prompt — the character itself is the condition (`--image` is its only required flag).\n\nFull commands and the timeout table → [references/cli-reference.md](references/cli-reference.md)\n\n## Prompt Formatting\n\nPer-model guidance (formulas, rules, example prompts) lives in [references/prompts.md](references/prompts.md). The English prose is there; example prompts keep their bilingual keyword pairs where the original was bilingual (e.g. `blurry 模糊`).\n\n- **Qwen-Image 2.1 (T2I + I2I default):** natural-language prompts. **CFG locked to 1.0** (distilled) — never pass `--cfg`. Optional `--negative` supported in both T2I and I2I (keep it short); I2I prompts stay concise, in the user's language.\n- **SD3.5 Medium:** natural language + optional `--negative`; default **1:1 (1024×1024)**, 20 steps, CFG 4.01.\n- **Z-Image / Qwen-Image 2.1:** `--aspect` is **optional** — T2I omits it and the output defaults to **1:1 (1024×1024)**; I2I omits it and the output canvas matches the **first reference image's size** (workflow latent switch ON). Pass `--aspect 16:9`, `9:16`, etc. to force a shape — in I2I this flips the latent switch to route the canvas through the ResolutionSelector (reference stays at native size, never stretched).\n- **Z-Image:** natural language — **no negative prompts**.\n- **Qwen Image Edit (2511):** concise, positive-only, user's exact words.\n- **Wan2.2 (I2V):** motion/action prompts; default 81 frames, 4 steps, CFG 4.5.\n- **i2m (Hunyuan 3D v2.1, no prompt):** the only input is the character image. Default extraction is **i2i** (Qwen-Image 2.1 redraws the main character centered on a transparent background, saved as PNG — alpha is preserved end-to-end: the Hunyuan 3D CLIP-vision conditioner drops the alpha channel itself, and the output is a GLB, never flattened). Use `--extract rembg` for a non-generative u2net human cutout, or `--extract none` to feed the image straight in (only if it's already a clean cutout). Defaults: resolution 4096, 30 steps, CFG 5.0, delivery scale **×3000** — the raw Hunyuan 3D v2.1 mesh is at a normalized unit scale (character ≈ 2 units, no declared units), so the GLB is scaled in place before zipping (`--scale 1` keeps it unscaled).\n\n## Error Handling\n\nThe script auto-detects three categories of server-side problems and reports them with install/source guidance, and `t2i`/`i2i` run through a **runtime fallback chain** when the default model fails. Details → [references/error-handling.md](references/error-handling.md):\n\n1. **Missing nodes** — reports which node, its package, and a pinned GitHub install hint (the script never runs `git clone` itself).\n2. **Missing models** — substitutes a compatible model (e.g. SD3.5 medium variants → SD3.5 large).\n3. **Non-critical utility nodes** (e.g. `UnloadAllModels`) — bypassed automatically without interrupting generation.\n4. **Runtime fallback chain** — `t2i`/`i2i` auto-descend through alternate models (and, for t2i, any available checkpoint) on runtime failures; config problems (missing files/nodes) abort instead of degrading. Disable with `--no-fallback`.\n\n## Output Delivery\n\n**Send the generated file — do not just describe it.** Deliver in the **user's original session**, never raw paths/URLs. Per-channel prefixes and the staging/cleanup routine → [references/delivery.md](references/delivery.md)\n\n**Output format:** generated images default to **JPG** (`config.json` → `image.output_format`). **Routing rule: if the user asks for a PNG image/file, or an RGBA image with transparency/alpha, always pass `--format png` and deliver the PNG file as-is — never convert to JPG** (JPEG has no alpha channel; the conversion would flatten transparency onto white). PNG output passes the server's PNG bytes through unchanged, so transparency is preserved byte-for-byte. Otherwise `--format jpg --jpg-quality N` is available per run; JPG conversion happens client-side (ComfyUI always saves PNG server-side). **i2m delivery rules (3D output):** (1) attach the **extracted character image** from step 1 as an **image** — never just its path (with `--extract none` the step-1 image *is* the user's own input, so there's nothing to send back); (2) send the mesh **packaged as a zip** — the script first **scales the GLB in place by the delivery factor** (default ×3000; the raw mesh is at a normalized unit scale, character ≈ 2 units, no declared units), then zips it (`i2m_<timestamp>.zip`); attach the **zip**, never the raw `.glb`; (3) **never render the mesh** — no GLB → image/video conversion and no 3D-viewer screenshots (a third-party viewer link is fine to share). Format conversion never applies to 3D files.\n\n## Execution Discipline — Fast Path\n\nThe **measured GPU sampling time is ~12–20s**. When a task takes minutes, the delay is **not** inference — it's either a model cold-reload or agent-side overhead. Kill the waste:\n\n**i2m exception:** 3D generation is legitimately long (cold checkpoint load + 30 sampling steps + voxel decode ≈ minutes on a 4090). Run it once in the background (`exec` with a generous `timeoutSeconds`), report the queue state, and when it completes deliver the character image + the GLB zip per the i2m delivery rules above — don't spin tool calls around it.\n\n1. **Never re-read on repeat.** Don't re-read `SKILL.md`, `config.json`, or the script for a known config / repeat / same-model job. Defaults are baked in; `--aspect` / `--steps` / `--cfg` are in the links above.\n2. **At most one pre-check merge.** Gather any needed info in a single batched command, or skip it entirely for trivial repeat jobs.\n3. **Run immediately after one check.** Don't inspect, second-guess, or add tool calls between knowing the prompt and running.\n4. **Verify only the minimum.** After the run, check only what the user asked (size / seed / file exists).\n5. **Golden rule:** sampling is ~12–20s. Anything longer = reload or overhead — never pile on tool calls before the image exists.\n6. **Prompt-flag trap:** pass `--prompt \"...\"` on the first attempt; a positional word wastes a whole turn.\n\n## References (index)\n\n- [references/image-first.md](references/image-first.md) — image-then-edit detection rules + multi-image reference composition\n- [references/attribution.md](references/attribution.md) — original source + modifications per workflow\n- [references/cli-reference.md](references/cli-reference.md) — full commands + timeout reference\n- [references/prompts.md](references/prompts.md) — per-model prompt guidance\n- [references/error-handling.md](references/error-handling.md) — auto-handled error categories\n- [references/delivery.md](references/delivery.md) — per-channel delivery + staging\n\nFile v3.5.1:README.md\n\n# Image with ComfyUI\n\nCall a user-defined ComfyUI server (local or remote) for **text-to-image**, **image-to-image/edit**, **image-to-video**, and **image-to-3D-mesh** generation. **Qwen-Image 2.1** is the default model for both T2I and image editing (including multi-image reference composition); Z-Image, SD3.5 Medium, Qwen Image Edit (2511), Wan2.2, and Hunyuan 3D v2.1 (3D mesh) are available as options.\n\n> **New here?** Qwen-Image 2.1 needs a one-time ComfyUI setup — see [Qwen-Image 2.1 ComfyUI Setup](#qwen-image-21-comfyui-setup).\n\n## Quick Start\n\n1. **Ensure a ComfyUI server is running** (locally, or on a host you can reach).\n2. **Set the URL** (via environment variable or `config.json`). The endpoint is user-defined — the bundled default is `http://127.0.0.1:8188`; point `COMFYUI_URL` at any ComfyUI server you control — a data-flow warning is printed on every run for non-local hosts. If no server is reachable, the script exits with configuration instructions:\n   ```bash\n   export COMFYUI_URL=http://comfyui.host:api-port\n   ```\n3. **Test a workflow** (see [Testing](#testing) below).\n\n## Configuration\n\nRead `config.json` at the skill root. All values can be overridden by environment variables:\n\n| Env Variable | Overrides | Default |\n|---|---|---|\n| `COMFYUI_URL` | `comfyui_url` | `http://localhost:8188` |\n| `COMFYUI_TIMEOUT` | `timeout_seconds` | `120` |\n| `COMFYUI_POLL_INTERVAL` | `poll_interval_seconds` | `1` |\n| `COMFYUI_KEEP_MODELS_RESIDENT` | `keep_models_resident` | `false` |\n| `COMFYUI_OUTPUT_DIR` | `output_dir` | `media/comfyui` |\n| `COMFYUI_APIKEY` | `comfyui_api_key` | _(none → no auth header)_ |\n\n**Priority:** Environment variables > `config.json`.\n\n> Note: The config.json key `comfyui_api_key` (string) is optional. When empty/absent, **no** `Authorization` header is added and behavior is unchanged. When set, every ComfyUI API request carries `Authorization: Bearer <key>` (env var wins over the config value).\n\n### 🔒 ComfyUI API Key (Bearer Auth)\n\nIf your ComfyUI instance requires an API key, set it via either of these (env var takes precedence):\n\n```bash\nexport COMFYUI_APIKEY=sk-your-key\n```\n\nor, in `config.json`:\n\n```json\n\"comfyui_api_key\": \"sk-your-key\"\n```\n\nWhen present, the skill adds `Authorization: Bearer <key>` to all ComfyUI API calls (`/prompt`, `/history`, `/object_info`, `/upload/image`, `/view`, `/upload/*`). Do not commit your key to version control — prefer the env var, and never paste it into the chat or logs.\n\n> **⚠️ This skill only *sends* the key — it doesn't authenticate your ComfyUI itself.** For the Bearer key to be honored, the ComfyUI server must run an authentication custom node that recognizes that key. The reference implementation is [`sunshinejnjn/comfyui-auth`](https://github.com/sunshinejnjn/comfyui-auth) (a fork of the original [`ivellioscolin/comfyui-auth`](https://github.com/ivellioscolin/comfyui-auth)). To set it up:\n> 1. Clone it into your ComfyUI: `git clone https://github.com/sunshinejnjn/comfyui-auth.git ComfyUI/custom_nodes/comfyui-auth`\n> 2. Edit its `apikeys.conf` — add a `label = <your-key>` line (the key is read **as written**, plain or as configured), then restart ComfyUI (the plugin reloads its config per request, but a re-clone/restart makes the new code load).\n> 3. Point this skill at that server via `COMFYUI_URL` and set `COMFYUI_APIKEY` to that same key.\n>\n> Without the plugin installed on the server, the key is silently ignored and the API responds as if no auth is required. The `test` command (`python3 image_with_comfyui.py test`) will tell you whether the endpoint answers before you try generating.\n\n### Data & Privacy\n\n- **Transmitted to the endpoint**: prompts and workflow JSON (all modes); **full source images** for I2I and I2V.\n- **Endpoint** is user-defined (`COMFYUI_URL` env var > `comfyui_url` in `config.json`; bundled default `http://127.0.0.1:8188`). A non-local endpoint prints a warning on every run, and an unreachable server exits with configuration instructions (this section + [Qwen-Image 2.1 ComfyUI Setup](#qwen-image-21-comfyui-setup)). Use HTTPS for remote servers and never put credentials in the URL or logs.\n- **Written to disk**: generated media is saved under `output_dir` (config) or `COMFYUI_OUTPUT_DIR` in your workspace; delivery staging copies live in `~/.openclaw/media/outbound/` and are removed after a successful send. If your content is sensitive, review the endpoint and output location before running.\n\n## Models & Custom Nodes\n\nThis skill uses the **existing** ComfyUI installation — it does not ship models or nodes. Before using, ensure your ComfyUI has:\n\n### Required Models\n\nEvery model below has a **verified download source** in [`MODEL_URL.txt`](MODEL_URL.txt) — open that file for the exact HuggingFace repo, folder, and filename. Our server currently runs a **UC/NSFW-merged set** (Qwen-Image 2.1 = uncensored, Wan2.2 = NSFW creator merges); `MODEL_URL.txt` lists **SFW/official equivalents** for a clean setup.\n\n| Model | Expected Location | Source |\n|---|---|---|\n| Z-Image (SVD XT) | `ComfyUI/models/checkpoints/` | [`MODEL_URL.txt`](MODEL_URL.txt) |\n| SD3.5 Medium | `ComfyUI/models/checkpoints/` | [`MODEL_URL.txt`](MODEL_URL.txt) |\n| Qwen Image Edit Plus | `ComfyUI/models/checkpoints/` | [`MODEL_URL.txt`](MODEL_URL.txt) |\n| **Qwen-Image 2.1** (default T2I + edit) | see [Qwen-Image 2.1 ComfyUI Setup](#qwen-image-21-comfyui-setup) | [`MODEL_URL.txt`](MODEL_URL.txt) |\n| Wan2.2 1.3B | `ComfyUI/models/checkpoints/` | [`MODEL_URL.txt`](MODEL_URL.txt) |\n| Wan2.2 VAEMODEL | `ComfyUI/models/vae/` | [`MODEL_URL.txt`](MODEL_URL.txt) |\n| Wan2.2 Audio Model | `ComfyUI/models/audio_encoders/` | [`MODEL_URL.txt`](MODEL_URL.txt) |\n| T5XXL & FluxFill | `ComfyUI/models/clip/` | [`MODEL_URL.txt`](MODEL_URL.txt) |\n| **Hunyuan 3D v2.1** (I2M) | `ComfyUI/models/checkpoints/` | [`MODEL_URL.txt`](MODEL_URL.txt) |\n\n### Required Custom Nodes\n\nInstall **pinned at reviewed commits** (as of 2026-09-20). Check out the exact commit rather than a mutable default branch, and install dependencies in an isolated venv under a non-privileged user — never `pip install` as root:\n\n| Node | Package (repo) | Pinned commit |\n|---|---|---|\n| Impact Pack | [ltdrdata/ComfyUI-Impact-Pack](https://github.com/ltdrdata/ComfyUI-Impact-Pack) | `429d0159a` |\n| WAS Nodes | [WASasquatch/was-node-suite-comfyui](https://github.com/WASasquatch/was-node-suite-comfyui) *(repo renamed from `ComfyUI-WAS-Nodes`)* | `9934caa92` |\n| ComfyUI-Manager | [Comfy-Org/ComfyUI-Manager](https://github.com/Comfy-Org/ComfyUI-Manager) *(migrated from `comfyanonymous/`)* | `946ef8fe6` |\n| rgthree | [rgthree/rgthree-comfy](https://github.com/rgthree/rgthree-comfy) | — (required by the Qwen-Image 2.1 workflow's LoRA pass-through node) |\n\n```bash\ncd ComfyUI/custom_nodes\ngit clone https://github.com/ltdrdata/ComfyUI-Impact-Pack.git\ncd ComfyUI-Impact-Pack && git checkout --detach 429d0159ad429e64d2b3916e6e7be9c22d025c3c\npython3 -m venv .venv && .venv/bin/pip install -r requirements.txt\n```\n\nRepeat the same pin-and-checkout pattern for the other two repositories (use their pinned commits above). Review dependency lists before installing, and re-check upstream commits before upgrading a pin.\n\nWhen a missing node or model is detected, the script reports the package name and a manual install hint — it never executes `git clone` itself (ComfyUI may run on another host).\n\n## Qwen-Image 2.1 ComfyUI Setup\n\n**Qwen-Image 2.1** (model id `qwen21`) is the **default** model for both `t2i` and `i2i`. One workflow covers text-to-image, single-image edit, and multi-image reference composition. It is a distilled model — **CFG must stay at 1.0** (the script enforces this; raising it burns the image).\n\n### Prerequisites\n- **ComfyUI ≥ 0.37.0** (frontend 1.53.6). The workflow uses core nodes that only exist in recent builds:\n  - `TextEncodeQwenImage21` (`comfy_extras.nodes_qwen`)\n  - `ResolutionSelector` (`comfy_extras.nodes_resolution`)\n  - `ComfySwitchNode` / `If/Else Switch` (`comfy_extras.nodes_logic`)\n  - `PrimitiveBoolean` (`comfy_extras.nodes_primitive`)\n  \n  If any show up as a red *missing node*, **update ComfyUI** — they are core, not a custom pack. (`If/Else Switch` is experimental and may be hidden from node search until you enable experimental nodes; loading the workflow works either way.)\n\n### Model files (3)\nPlace each in the folder shown, under your ComfyUI install. These are the exact filenames the workflow expects:\n\n| File | Destination folder |\n|---|---|\n| `QWEN_Image/qwen-image-2.1-UC-int8_convrot.safetensors` | `ComfyUI/models/diffusion_models/` (keep the `QWEN_Image/` subfolder; older builds may use `models/unet/`) |\n| `qwen3vl_8b_int8_convrot.safetensors` | `ComfyUI/models/text_encoders/` (Qwen3-VL 8B int8; loaded with `type: qwen_image`) |\n| `qwen_image_2.1_vae_bf16.safetensors` | `ComfyUI/models/vae/` |\n\n> Unlike the original foprc workflow (which used a GGUF Q8 model + ComfyUI-GGUF), this workflow loads the diffusion model via **`UNETLoader`** with `weight_dtype: fp8_e4m3fn_fast`. That **removes the ComfyUI-GGUF dependency** — the only custom node needed is **rgthree**.\n\nIf you use a different quantization/filename, update `unet_name` in `workflows/qwen_image_2.1_api.json` (node `17`) to match what's on disk.\n\n### Custom nodes\n- **rgthree-comfy** (for the `Power Lora Loader` node) — install via ComfyUI Manager or `git clone https://github.com/rgthree/rgthree-comfy` into `custom_nodes/`, then restart.\n- All other nodes are core ComfyUI.\n- *Optional:* the LoRA stack is off by default and passes the model/CLIP straight through. If you don't have rgthree, you can instead delete node `10` and rewire `17 → 6` (model) and `2 → 5` (clip).\n\n### Verifying your setup\n```bash\npython3 image_with_comfyui.py test          # server reachable?\npython3 image_with_comfyui.py t2i --prompt \"A red fox in snow\" --aspect 1:1\npython3 image_with_comfyui.py i2i --prompt \"Change the sky to sunset\" --image /path/in.jpg\n```\nIf a loader dropdown is blank in the web UI, the file isn't where ComfyUI is looking — check the folder above. If a node reports *missing*, update ComfyUI (see Prerequisites).\n\n### How the dual-mode switch works\nNode `12` (`PrimitiveBoolean`) is the master switch:\n- **OFF** (`false`) → pure **T2I**: the sampler starts from an empty latent sized by `ResolutionSelector` (aspect ratio + megapixels).\n- **ON** (`true`) → **edit / multi-image**: reference image(s) feed `TextEncodeQwenImage21`, whose output latent is sampled instead. The script wires 1 image to `image_1`, and extra images to `image_2 … image_N` (up to 16).\n\n### Multi-image prompt format\nWith 2+ reference images, the model sees them in **slot order** (`image_1`, `image_2`, …). Reference each by index in the prompt so it knows which is which:\n```\nPut the garment from <image2> onto the person in <image1>, keep <image1>'s pose and lighting.\nCombine both photos into one scene: subject of <image1> on the left, subject of <image2> on the right.\n```\nKeep the instruction concise and in the user's language; the skill passes it through verbatim. `--image` accepts multiple paths (order = slot order = `<image1>`, `<image2>` …).\n\n## Testing\n\nUse the built-in test script:\n\n```bash\n# Health check\npython3 image_with_comfyui.py test\n\n# Test each workflow individually\npython3 image_with_comfyui.py t2i --prompt \"A cat on a windowsill\"\n\n# Run the health check before each batch of jobs\npython3 image_with_comfyui.py test\n```\n\n### Output filename scheme\n\nGenerated files use a short mode prefix + timestamp (`YYYYMMDD_HHMMSS`) + ComfyUI's counter:\n\n| Mode | Prefix | Example (JPG) |\n|---|---|---|\n| Qwen-Image 2.1 T2I | `qt2i` | `qt2i_20260928_050328_00001.jpg` |\n| Qwen-Image 2.1 I2I | `qi2i` | `qi2i_20260928_050328_00001.jpg` |\n| Qwen Image Edit (2511) | `qei2i` | `qei2i_20260928_050328_00001.jpg` |\n| Z-Image T2I | `zt2i` | `zt2i_20260928_050328_00001.jpg` |\n| SD3.5 Medium T2I | `st2i` | `st2i_20260928_050328_00001.jpg` |\n| Character cutout (i2m step 1) | `mcut` | `mcut_20260928_050328_00001.png` |\n| Hunyuan 3D v2.1 i2m | `i2m` | `i2m_20260928_050328_00001.glb` |\n| i2m delivery package | `i2m` | `i2m_20260928_050328_00001.zip` (the ×3000-scaled GLB inside) |\n\nThe trailing underscore ComfyUI appends is stripped, so names never end in `_`. The i2m GLB is **scaled ×3000 in place, then packaged as `<same-stem>.zip`** — the `.glb` stays on disk, but what gets sent to the user is the **zip** (delivery rules: SKILL.md → Output Delivery; the character image is also sent, as an image, and the mesh is never rendered).\n\n### Manual Workflow Testing\n\nYou can also test workflows directly in the ComfyUI web UI. Copy any workflow file from `workflows/` into ComfyUI, load it, and run:\n\n- **Qwen-Image 2.1 (T2I + edit + multi-image, default)**: `workflows/qwen_image_2.1_api.json`\n- **Z-Image T2I**: `workflows/z-image_t2i_api.json`\n- **SD3.5 Medium T2I**: `workflows/sd3.5-med_t2i_api.json`\n- **Qwen Image Edit (2511)**: `workflows/qwen_image-edit_api.json`\n- **Wan2.2 I2V**: `workflows/wan2.2_i2v_api.json`\n- **Hunyuan 3D v2.1 I2M (image → 3D mesh)**: `workflows/hunyuan3d-v2.1_i2m_api.json`\n\n## Important Notes\n\n### Local Skill — No Automatic Installation\n\nThis skill **does not install anything automatically**. It connects to your existing ComfyUI instance via API calls; the workflows, models, and custom nodes must already be configured in your ComfyUI setup. If a node is missing, the skill only *reports* the package and a pinned manual install hint (it never runs `git clone` itself, since ComfyUI may be on another host). The manual commands in [Required Custom Nodes](#required-custom-nodes) are optional setup steps for your own ComfyUI — review and pin them before running.\n\n### VRAM Release vs. Speed Tradeoff\n\nBy default this skill uses the `UnloadAllModels` node (and `UnloadModel` in the Qwen edit branch) to free VRAM after each generation, preventing VRAM exhaustion during repeated tasks. The unload nodes are stripped only when `keep_models_resident` is explicitly set to `true` (config or `COMFYUI_KEEP_MODELS_RESIDENT`); otherwise they run and release VRAM on every run. The unload nodes are bypassed automatically if a given ComfyUI instance doesn't have them — the system detects and skips them transparently.\n\n> **The slow part of a “5-minute” run is almost never inference.** It's model reload *or* agent-side overhead before the run, not the sampling. See the skill's **Execution Discipline — Fast Path** section for how the run is expected to be invoked.\n\n### Workflow Modifications\n\nAll workflow files in the `workflows/` directory are **adapted versions** of original ComfyUI workflows. They are repackaged into self-contained HTTP-API graphs — the original source is **not** kept inside the JSON. The original source for each workflow plus exactly what was changed is documented in [references/attribution.md](references/attribution.md). Known modifications include:\n- API-ready output node placement + timestamped output filenames\n- Parameter exposure (aspect, steps, CFG, seed, neg, etc.) for external control\n- Prompt auto-routing (e.g. Qwen Edit routes to the positive prompt node)\n- A standard **safetensors** model load replacing the original GGUF load (Qwen-Image 2.1), dropping the ComfyUI-GGUF dependency\n- VRAM-cleanup nodes (`UnloadAllModels`, `UnloadModel`) added, auto-bypassed when absent\n- The SD3.5 double-`.safetensors` suffix bug fixed; Wan2.2 pointed at 4-step Lightspeed weights\n\n### Error Handling\n\nThe system handles three categories of errors automatically:\n\n1. **Missing nodes** — Reports which node is missing, provides package name, install command, and GitHub URL.\n2. **Missing models** — Attempts to substitute a compatible model (e.g., SD3.5 medium variants → SD3.5 large).\n3. **Non-critical utility nodes** — Automatically bypasses nodes like `UnloadAllModels` without interrupting generation.\n\n## CLI Reference\n\n```bash\n# Generate image from text (default: Qwen-Image 2.1)\npython3 image_with_comfyui.py t2i --prompt \"Your description\" --aspect 1:1\n\n# Generate with an explicit model\npython3 image_with_comfyui.py t2i --model z-image --prompt \"Your description\"\npython3 image_with_comfyui.py t2i --model sd35 --prompt \"Your description\" --negative \"blurry\"\n\n# Edit an image (default: Qwen-Image 2.1)\npython3 image_with_comfyui.py i2i --prompt \"Replace background\" --image /path/to/image.png\n\n# Multi-image reference (2+ images; order = <image1>, <image2>, ...)\npython3 image_with_comfyui.py i2i \\\n  --prompt \"Put the shirt from <image2> on the person in <image1>\" \\\n  --image /path/person.jpg /path/shirt.jpg\n\n# Generate video from image\npython3 image_with_comfyui.py wan2.2 --prompt \"Camera pans left\" --image /path/to/image.png\n\n# Turn the character in an image into a 3D mesh (GLB)\n# Default: extract the main character first (i2i transparent cutout), then generate\npython3 image_with_comfyui.py i2m --image /path/to/character.jpg\n# --extract rembg (u2net cutout, no redraw) | --extract none (image used as-is)\n# Delivery: character image (as an image) + i2m_*.zip (as an attachment);\n# the GLB is scaled x3000 in place before zipping (--scale 1 to skip);\n# the mesh is never rendered (no GLB -> image/video conversion; viewer links OK)\n\n# Test ComfyUI connection\npython3 image_with_comfyui.py test\n```\n\n## Timeout Reference\n\n| Mode | Timeout |\n|---|---|\n| T2I (Qwen-Image 2.1) | 300s |\n| T2I (Z-Image) | 100s |\n| T2I (SD3.5) | 100s |\n| I2I (Qwen-Image 2.1, incl. multi-image) | 600s |\n| I2I (Qwen 2511) | 600s |\n| I2V (Wan2.2) | 1000s |\n| I2M (Hunyuan 3D v2.1) | 900s |\n\nFile v3.5.1:_meta.json\n\n{\n  \"ownerId\": \"kn7dn131pvgfaveyzzzcrc5xf182yb84\",\n  \"slug\": \"image-with-comfyui\",\n  \"version\": \"3.5.1\",\n  \"publishedAt\": 1791049346849\n}\n\nFile v3.5.1:references/attribution.md\n\n# Source Attribution (image-with-comfyui)\n\nEvery workflow in `workflows/` is an **adapted** version of a real ComfyUI\nworkflow. Original source + the specific modifications applied by this skill are\ndocumented here. The workflow JSONs carry only ComfyUI's standard per-node\nlabels (`_meta` `title`); they do **not** embed provenance or attribution\nmetadata, so the original source is not duplicated inside them — it lives here\ninstead, and each file stays runnable.\n\n| Workflow | Original source | Key modifications applied |\n|---|---|---|\n| `qwen_image_2.1_api.json` | [foprc/qwen-image-2.1-comfyui-workflow](https://github.com/foprc/qwen-image-2.1-comfyui-workflow) (ships GGUF Q8_0 + ComfyUI-GGUF) | Replaced GGUF model with standard **safetensors** via `UNETLoader` (`weight_dtype: fp8_e4m3fn_fast`) — removes the ComfyUI-GGUF dependency; added `UnloadModel` (node 18) + `UnloadAllModels` (node 19) VRAM-cleanup; dual-mode reference switch (nodes 12/13/14); `ResolutionSelector` (node 9); API/`timestamp` output |\n| `qwen_image-edit_api.json` | Qwen-Image-Edit (Comfy-Org) default edit graph | Prompt auto-routing to the positive `TextEncodeQwenImageEditPlus` node (115:111) by the script |\n| `sd3.5-med_t2i_api.json` | Comfy-Org / sd3-medium default graph | Fixed the `.safetensors.safetensors` double-suffix `ckpt_name` bug; `sd3_medium` → `sd3.5_large.safetensors` fallback; `UnloadAllModels` VRAM-cleanup |\n| `z-image_t2i_api.json` | Comfy-Org / z-image default graph | `UnloadAllModels` VRAM-cleanup |\n| `wan2.2_i2v_api.json` | Comfy-Org / Wan2.2 I2V default graph | Pointed the diffusion model at **4-step Lightspeed** weights; LoRA stack nodes (109/110); `UnloadAllModels` VRAM-cleanup |\n| `hunyuan3d-v2.1_i2m_api.json` | Tencent Hunyuan3D-2.1 ComfyUI image→mesh reference graph (user-supplied, [tencent/Hunyuan3D-2.1](https://huggingface.co/tencent/Hunyuan3D-2.1)) | Topology kept 1:1 (including the user's VRAM-unload sink nodes 19/20/21/22); placeholder `image` set to `character.png` (the script always rewrites it to the uploaded filename); default `filename_prefix` kept (`mesh/hy21`) — the script stamps it per run to `mesh/i2m_<timestamp>`; no model/parameter changes — resolution 4096 / 30 steps / CFG 5.0 are the graph's own defaults and stay the skill defaults\n\nAll six are **adapted** (as opposed to) the original; they add prompt routing,\nresolution selection, dual-mode switches, and VRAM management on top of the\noriginal ComfyUI graphs.\n\nFile v3.5.1:references/cli-reference.md\n\n# CLI Usage Reference (image-with-comfyui)\n\nFull command examples for every subcommand. The SKILL.md procedure only lists the\ncore rule; open this file for the exact flags.\n\n> **⚠️ The prompt is MANDATORY and must be passed via `--prompt` — never as a positional argument.**\n>\n> Every subcommand that takes a prompt (`t2i`, `i2i`, `wan2.2`) reads it from `--prompt \"...\"` (`i2m` has no prompt — see below). Passing the prompt as a bare positional word fails immediately with exit code 2:\n> `python3 image_with_comfyui.py t2i \"some description\"` → `t2i: error: the following arguments are required: --prompt`\n>\n> This is a **100%-avoidable failure loop** — the model can't help you because the command never reaches ComfyUI. Always write the prompt through the `--prompt` flag on the first attempt.\n\n### T2I (Text → Image)\n\n```bash\n# Qwen-Image 2.1 (default model; CFG locked to 1.0 — no --cfg)\n# --aspect is OPTIONAL: omit it and the output is 1:1 (1024x1024)\npython3 image_with_comfyui.py t2i \\\n  --prompt \"Your detailed image description\" \\\n  --steps 25\n\n# With a negative prompt (Qwen-Image 2.1 supports it)\npython3 image_with_comfyui.py t2i \\\n  --prompt \"A detailed image description\" \\\n  --negative \"text, watermark, blurry\" \\\n  --steps 25\n\n# Request a specific aspect ratio explicitly (16:9, 9:16, 4:3, ...)\npython3 image_with_comfyui.py t2i \\\n  --prompt \"Your detailed image description\" \\\n  --aspect 16:9 \\\n  --steps 25\n\n\n# Z-Image\npython3 image_with_comfyui.py t2i \\\n  --model z-image \\\n  --prompt \"Your detailed image description\" \\\n  --aspect 1:1 \\\n  --steps 9\n\n# SD3.5 Medium (defaults to 1:1)\npython3 image_with_comfyui.py t2i \\\n  --model sd35 \\\n  --prompt \"A beautiful sunset over mountains\" \\\n  --negative \"text, watermark, blurry\" \\\n  --steps 20 \\\n  --cfg 5.5\n\n# Run ONLY the requested model (skip the runtime fallback chain)\npython3 image_with_comfyui.py t2i \\\n  --prompt \"...\" --model z-image --no-fallback\n\n# JPG output (converted client-side; default from config.json image.output_format)\npython3 image_with_comfyui.py t2i \\\n  --prompt \"...\" --format jpg --jpg-quality 90\n```\n\n### I2I (Edit Image)\n\n```bash\n# Qwen-Image 2.1 (default I2I)\npython3 image_with_comfyui.py i2i \\\n  --prompt \"Change background to a beach\" \\\n  --image /path/to/source.jpg \\\n  --steps 25\n\n# Explicit aspect ratio: routes the I2I latent canvas through the\n# ResolutionSelector (workflow latent switch OFF) instead of matching the\n# first reference image's size, so a portrait photo CAN become 16:9 landscape.\n# The reference is uploaded at native size, never stretched. Without --aspect,\n# the output follows the first reference image's own size (no 1:1 forcing).\npython3 image_with_comfyui.py i2i \\\n  --prompt \"Extend the scene into a wide landscape\" \\\n  --image /path/to/portrait.jpg \\\n  --aspect 16:9\n\n# Multi-image reference (up to 16): --image order = slot order\npython3 image_with_comfyui.py i2i \\\n  --prompt \"put the shirt from <image2> on the person in <image1>\" \\\n  --image /path/to/base.jpg /path/to/reference.jpg\n\n# Legacy single-image edit\npython3 image_with_comfyui.py i2i \\\n  --model qwen_imageedit \\\n  --prompt \"put on a red jacket\" \\\n  --image /path/to/source.jpg \\\n  --steps 4\n\n# Qwen-Image 2.1 edit with a negative prompt\npython3 image_with_comfyui.py i2i \\\n  --prompt \"replace the background with a night city\" \\\n  --negative \"text, watermark\" \\\n  --image /path/to/source.jpg\n\n# Run ONLY the requested model (skip the runtime fallback chain)\npython3 image_with_comfyui.py i2i \\\n  --prompt \"...\" --image /path/to/source.jpg --no-fallback\n\n# JPG output (see t2i example)\npython3 image_with_comfyui.py i2i \\\n  --prompt \"...\" --image /path/to/source.jpg --format jpg\n```\n\n### I2V (Image → Video)\n\n```bash\npython3 image_with_comfyui.py wan2.2 \\\n  --prompt \"the person walks forward and smiles\" \\\n  --image /path/to/source.jpg \\\n  --length 81 --steps 4\n```\n\n\n### I2M (Image → 3D Mesh, Hunyuan 3D v2.1 → GLB)\n\n```bash\n# Default: extract the main character first (Qwen-Image 2.1 i2i redraw on a\n# transparent background, PNG), then generate the mesh.\npython3 image_with_comfyui.py i2m --image /path/to/character.jpg\n\n# Non-generative extraction: u2net human cutout (pixel-precise, no redraw)\npython3 image_with_comfyui.py i2m --image /path/to/character.jpg --extract rembg\n\n# Image is already a clean cutout — skip extraction entirely\npython3 image_with_comfyui.py i2m --image /path/to/cutout.png --extract none\n\n# Higher-detail / more-iterative generation\npython3 image_with_comfyui.py i2m --image /path/to/character.jpg \\\n  --resolution 8192 --steps 35 --cfg 6.0 --seed 42\n\n# Custom delivery scale (default: x3000; 1 = keep the normalized units)\npython3 image_with_comfyui.py i2m --image /path/to/character.jpg --scale 1\n```\n\n- Output: a **GLB file** under `output_dir/meshes/` (prefix `i2m_<timestamp>`),\n  **scaled in place by the delivery factor** (default ×3000 — the raw Hunyuan\n  3D v2.1 mesh is at a normalized unit scale, character ≈ 2 units, no declared\n  units; `--scale 1` keeps it unscaled) and **automatically packaged as\n  `<same-stem>.zip`** — the result block prints both paths plus the delivery hints.\n- **Delivery (i2m rules):** (1) send the extracted **character image** (the\n  step-1 PNG; the script prints it as `CHARACTER IMAGE`) **as an image** — with\n  `--extract none` there is no new image to send back. (2) send the **zip** (the\n  scaled GLB) as an attachment, never the raw `.glb`. (3) **Never render the\n  mesh** — no\n  GLB → image/video conversion, no 3D-viewer screenshots (a third-party\n  viewer *link* is fine).\n- Requires `hunyuan_3d_v2.1.safetensors` in the server's `models/checkpoints/`\n  and ComfyUI **≥ 0.37.0** (all Hunyuan 3D v2.1 nodes are core — no custom\n  packages).\n- This mode is the slow one by design: expect minutes, not ~20 s (see SKILL.md\n  \"i2m exception\"). Run it in the background and deliver on completion.\n\n### Health Check\n\n```bash\npython3 image_with_comfyui.py test\n```\n\n### Output Format (t2i / i2i)\n\nComfyUI's `SaveImage` node always writes PNG server-side. `--format jpg` converts the\n**downloaded** image client-side (Pillow) to JPEG — alpha channels are composited\nonto white, and conversion failures fall back to PNG automatically. `--format png`\npasses the server's PNG bytes through **unchanged** (no conversion), so RGBA\ntransparency survives intact. Defaults come from `config.json` → `image.output_format`\n(default `\"jpg\"`) and `image.jpg_quality` (default 89). Videos (wan2.2) are always\nMP4 and are not affected.\n\n**Routing rule:** the user's wording decides the format —\n\n- \"PNG image/file\", \"PNG format\", \"keep the alpha channel\", or **\"RGBA image with transparency\"** → always `--format png`, and **deliver the PNG file directly without any JPG/JPEG conversion** (never flatten transparency onto a background).\n- No format mentioned → use the config default (JPG).\n\n---\n\n## Timeout Reference\n\n| Model | Timeout |\n|---|---|\n| T2I (Qwen-Image 2.1) | 600s |\n| T2I (Z-Image) | 100s |\n| SD3.5 Medium | 100s |\n| I2I (Qwen 2.1) | 600s |\n| I2I (Qwen Image Edit 2511) | 600s |\n| I2V (Wan2.2) | 1000s |\n| I2M (Hunyuan 3D v2.1) | 900s |\n\nFile v3.5.1:references/delivery.md\n\n# Delivery Reference (image-with-comfyui)\n\nThe SKILL.md core says \"send the media, in the user's session, no paths\". This\nfile holds the per-channel delivery formats and the staging steps. Open it when\nyou need the exact prefix for a channel.\n\n## The one rule\n\nAfter generating, **send the file**, don't describe it. A text line like \"video\nattachment\" is NOT a send.\n\n## Per-channel format\n\n| Channel | Format | Example |\n|---|---|---|\n| WhatsApp | `MEDIA:` + absolute path | `MEDIA:/home/user/.openclaw/media/outbound/angel_video.mp4` |\n| Telegram | `MEDIA:` or `filePath:` | varies by implementation |\n| Discord | Direct attachment | varies by implementation |\n\n## Staging steps (WhatsApp)\n\n1. Copy the output to `~/.openclaw/media/outbound/` first.\n2. Send via `MEDIA:/home/user/.openclaw/media/outbound/<filename>` (use absolute file path).\n3. The `MEDIA:` line must be the **sole content** of the WhatsApp message — no\n   `[[reply_to_current]]`, no text before it — otherwise WhatsApp splits the file\n   and text into two messages. For a caption use `MEDIA:./file.ext caption=...`.\n4. **Clean up:** delete the staged file after a successful send so sensitive\n   generated media doesn't linger across sessions.\n\n## i2m (3D mesh) delivery\n\ni2m produces **two deliverables**, and both must go to the user:\n\n1. **The extracted character image** (step-1 PNG) — send it **as an image**\n   (an image message / attachment), never just its path. With `--extract none`\n   the step-1 image *is* the user's own input, so there is nothing new to send.\n2. **The mesh, as a zip** — the script first **scales the GLB in place by the\n   delivery factor** (default ×3000 — the raw Hunyuan 3D v2.1 mesh is at a\n   normalized unit scale, character ≈ 2 units, no declared units; `--scale 1`\n   keeps it unscaled), then packages it (`i2m_<timestamp>.glb` →\n   `i2m_<timestamp>.zip`). Send the **zip** as an attachment; never send the\n   raw `.glb` on its own.\n\n**Never render the mesh.** Do not convert the GLB to an image or video, or take\nscreenshots of a 3D viewer. A **third-party online GLB viewer link is allowed** —\nthe user views/rotates the mesh themselves; the zip + character image remain the\nprimary delivery.\n\n## Always\n\n- Deliver in the **user's original request session/thread** — not a separate\n  topic/group/thread unless explicitly told.\n- Never send raw local file paths, ComfyUI URLs, or host info unless asked.\n\nFile v3.5.1:references/error-handling.md\n\n# Error Handling Reference (image-with-comfyui)\n\nThe script auto-detects two categories of server-side problems (missing nodes and\nmissing models) and reports them with install/source guidance. Open this for the\nfull behavior; the SKILL.md procedure only tells you \"if you see ⚠️ report it\".\n\n## 1. Missing Node Detection\n\nWhen a workflow references a custom node that isn't installed, the system detects it and reports:\n- **Which node** is missing (class type name)\n- **Which package** provides it\n- **GitHub URL** for manual download\n- **Manual install instruction** (the script does NOT execute git clone; ComfyUI may be on a remote server)\n\nExample: If `ImpactKSamplerBasicPipe` is missing:\n```\n⚠️ Missing node: `ImpactKSamplerBasicPipe`\n📦 Package: ComfyUI-Impact-Pack\n🔗 GitHub: https://github.com/ltdrdata/ComfyUI-Impact-Pack\nℹ️ Install manually: see README \"Required Custom Nodes\" (pinned commit + venv; the script never runs git clone itself)\n```\n\n## 2. Missing Model Substitution\n\nWhen a workflow references a model file that doesn't exist, the system attempts to find a compatible substitute:\n\n| Requested Model | Substitute |\n|---|---|\n| `sd3.5_medium` variants | `sd3.5_large.safetensors` |\n| WAN High → Low or vice versa | Swap between variants |\n| Other unknown models | No substitution (error returned) |\n\nExample: If `my_custom_sd3_medium_v2.safetensors` is missing:\n```\n⚠️ Model missing: `my_custom_sd3_medium_v2.safetensors`\n🔄 Substituted: `sd3.5_large.safetensors`\n📦 Loader: CheckpointLoaderSimple.ckpt_name\n```\n\nAfter substitution, the workflow is retried automatically with the substitute model.\n\n## 3. Missing Utility Node Bypass (UnloadAllModels)\n\nWhen the workflow references `UnloadAllModels` (a memory cleanup node) which isn't available, the system **automatically bypasses it** by rerouting the signal path:\n- Removes the missing `UnloadAllModels` node\n- Redirects the upstream processing node directly to the downstream output node\n- Generation continues without interruption\n- User receives a warning about the bypass\n\nExample:\n```\n⚠️ Workflow missing node: `UnloadAllModels` (memory cleanup, non-critical)\n🔄 Auto-bypassed — generation continues\n```\n\n## 4. Runtime Fallback Chain (t2i / i2i)\n\n`t2i` and `i2i` commands run through a fallback engine (`_run_with_fallback`) instead of the single-model path. A run proceeds through ordered **tiers**:\n\n| Kind | Tier order |\n|---|---|\n| t2i | `qwen21` → `z-image` → `sd35` → (any remaining server checkpoint, standard topology) |\n| i2i | `qwen21` → `qwen_imageedit` |\n\n- The **requested model** (via `--model` or the config default) is always the first tier; the remaining tiers follow in the order above.\n- A tier fails on: send errors, server timeout (> `timeout_seconds`), `status_str == error`, or no output images.\n- **Config problems do NOT degrade** — missing model files or missing nodes abort the chain with `config problem (missing model file / node) — no fallback` (install/fix the model instead of silently switching).\n- **Server unreachable** mid-run is treated as an environment problem — every tier on the same server would fail identically, so the chain aborts.\n- **Server unreachable before a run:** `t2i` / `i2i` / `wan2.2` pre-check the user-defined endpoint (GET `/history`, 5 s timeout) and, if nothing answers, exit with configuration instructions (start a ComfyUI server, set `COMFYUI_URL` or `comfyui_url` in `config.json`, verify with `test`) instead of failing mid-run with a raw timeout.\n- **i2i never degrades blind:** a plain T2I model cannot honor the source image, so when the i2i tiers are exhausted the run exits with the available-model list for manual choice (no fallback to a text-only generation).\n- When a lower tier succeeds, the output includes a note: `🔄 used <tier> — fell back because <failed tiers> failed this run.`\n- `qwen_imageedit` skips the tier when the request has multiple input images (single-image model) and the chain continues.\n- **`--no-fallback`** (t2i / i2i): run only the requested model — original single-model behavior, no runtime descent.\n\nExample (qwen21 times out, z-image succeeds):\n\n```\n🎨 trying qwen21 ...\n⚠️ tier qwen21 — failed: server timeout (>600s)\n🎨 trying z-image ...\n✅ 1 image(s) saved ...\n🔄 used z-image — fell back because qwen21 failed this run.\n```\n\nFile v3.5.1:references/image-first.md\n\n# Image-First Mode Reference (image-with-comfyui)\n\nThe SKILL.md core says \"image-first: route remembered image + new text to\n`i2i`/`wan2.2`\". This file holds the full detection logic and worked examples.\nOpen it when a user sends an image and then a short edit request.\n\n## Detection\n\n1. The user sends **only an image** (no other text that turn).\n2. Within **2 minutes**, the user sends a **text** message that looks like an\n   edit or video request. The agent matches whichever language the user wrote —\n   CN or EN keywords, e.g. `修一下/fix`, `换背景/change background`,\n   `加特效/add an effect`, `变动画/turn into a video`, `改颜色/change color`.\n3. The text intent is **I2I** (edit the image) or **I2V** (animate the image).\n\n## Action\n\n- Route the **remembered latest image** + the **new text** to\n  `image_with_comfyui.py i2i` or `wan2.2`.\n- `--image <path>` = the latest image received. `--prompt` = the new text.\n- If unsure whether it's I2I or I2V, **default to I2I** unless the text clearly\n  says video/animation.\n- Do **NOT** ask the user for the image again — you already have it from the\n  previous turn.\n- **Scope:** only the immediately preceding image from the **same user, same\n  session**, within the 2-minute window. Never reuse images from other users,\n  other sessions, or older than 2 minutes.\n\n## Context tracking\n\n- Store the latest image media path (or URL) when no text is received.\n- Clear the stored image after it's used, or after 2 minutes of no new text.\n\n## Examples\n\n- `[image: a photo of a dog]` → (wait)\n- `[text: change the background to a beach]` → `i2i --image <path> --prompt \"change the background to a beach\"`\n- `[image: a cat sitting on a chair]` → (wait)\n- `[text: make it stand up and walk]` → `wan2.2 --image <path> --prompt \"the cat stands up and walks\"`\n\n## Multi-image reference (up to 16)\n\nWith 2+ reference images the model sees them in **slot order**\n(`image_1`, `image_2`, …). Reference each by index in the prompt:\n\n```\nPut the garment from <image2> onto the person in <image1>, keep <image1>'s pose and lighting.\n```\n\n`--image` accepts multiple paths in slot order: `--image /base.jpg /ref.jpg`\n= `<image1>` `/base.jpg`, `<image2>` `/ref.jpg`.\n\nFile v3.5.1:references/prompts.md\n\n# Prompt Formatting Reference (image-with-comfyui)\n\nPer-model guidance for writing prompts. This is reference material — the SKILL.md\nprocedure just says \"format your prompt; see below\". Prompt keywords are kept\nbilingual where they originally appeared bilingual (e.g. `blurry 模糊`).\n\n## Z-Image (T2I)\n\nZ-Image works best with **structured natural language prompts**, not keyword spam.\n\n**6-part formula:**\n```\nSubject + Scene + Composition + Lighting + Style + Constraints\n```\n\n**Rules:**\n- ✅ Use natural language sentences (not comma-separated tags)\n- ✅ Be specific about subject, camera, lighting, style\n- ❌ **NO negative prompts** — Z-Image Turbo ignores them completely\n- ❌ No weighted tags like `(word:1.2)`\n\n**Example:**\n```\nA young woman with long wavy blonde hair sits at a wooden café table,\nsteam rising from a ceramic cup. Shot from a 3/4 angle, close-up framing.\nSoft morning light filters through sheer curtains, casting warm golden tones.\nCinematic photography, shallow depth of field, Kodak Portra 400 aesthetic.\nNo text, no logos, photorealistic skin texture.\n```\n\n**Aspect ratios:** `1:1`, `4:3`, `3:4`, `16:9`, `9:16`, `3:2`, `2:3`\nShould be ommited (defaults to 1:1) unless specified by user or as an extracted requirement from user's requests.\nIn I2I situation, it should match the original image which the user asks to modify.\n\n---\n\n## SD3.5 Medium (T2I)\n\nSD3.5 Medium uses **natural language prompts** with optional negative prompts.\n\n**Prompt formula:**\n```\n[Composition/Angle] + [Subject] + [Scene/Environment] + [Lighting/Color] + [Style/Texture] + [Details]\n```\n\n**Rules:**\n- ✅ **Complete natural language sentences** — describe telling a human what to see\n- ✅ Subject first (model prioritizes early text)\n- ✅ Be specific about colors, materials, mood, atmosphere\n- ✅ Mixed CN/EN is fine (Chinese works better for Chinese scenes)\n- ✅ Use `--negative` for elements to exclude\n- ✅ Default 1:1 (1024×1024), use `--aspect` to change\n- ✅ Default 20 steps, CFG 4.01 (higher = stronger control)\n- ✅ Seed defaults to random; specify `--seed` for reproducibility\n- ❌ No comma-separated keyword spam (`beautiful, amazing, 4k`)\n- ❌ No weighted tags `(word:1.2)` — SD3.5 doesn't recognize them\n\n**Parameter recommendations:**\n- **CFG**: 4-7 (4.01 = softer, 5-7 = stronger control)\n- **Steps**: 20-25 (below 20 may lack detail)\n- **Negative prompt**: Highly effective in SD3.5\n\n**Common negative prompt words (mixed CN/EN — either or both work):**\n```\nblurry 模糊, low quality 低质量, pixelated 像素化, grainy 颗粒感,\noverexposed 过曝, underexposed 欠曝, flat lighting 平光,\ntext 文字, watermark 水印, logo 商标, signature 署名, caption 字幕,\npoorly drawn face 画坏的脸, deformed 变形, mutated 畸变, disfigured 毁容, extra limbs 多余肢体,\ncartoonish 卡通风(写实时用)\n```\n\n**Mixed CN/EN example — both work; for Chinese scenes the prompt may be written in Chinese:**\n```\n上海魔都春日花海 — 黄浦江畔，大片郁金香、樱花、油菜花盛开，繁花似锦；\n春日和煦阳光，远景陆家嘴三件套天际线；湿润的滨江步道倒映花影；\n低饱和胶片色调，文艺清新，广角视野。\n```\n\n**English example:**\n```\nCinematic photography, wide-angle shot of a bustling Tokyo street at night,\nneon signs reflecting on wet pavement, people with transparent umbrellas,\nmoody atmospheric lighting, deep blues and vibrant reds, street photography,\nshallow depth of field with bokeh background\n```\n\n---\n\n## Qwen-Image 2.1 (T2I default & I2I default)\n\n- T2I: same 6-part natural-language formula as Z-Image. **Optional `--negative`** (kept short, e.g. `text, watermark, blurry`). **CFG locked to 1.0** (distilled) — never pass `--cfg`.\n- I2I: concise prompts, user's original language; **optional `--negative`** allowed (concise, positive-only edits still work best). Multi-image: up to 16 refs, `--image` order = slot order, referenced as `<image1>`, `<image2>`, ...; with the reference switch ON the output follows the first reference's aspect ratio.\n\n## Qwen Image Edit (I2I) — Concise Prompts\n\nI2I prompts must be **concise and direct**. Keep the user's original language.\n\n**Rules:**\n- ✅ **Positive prompt only** — no negative prompts\n- ✅ Use user's exact words (don't translate or expand)\n- ✅ Concise (either language): 换红外套/put on a red jacket, 把背景换成蓝天白云/change the background to a blue sky with clouds, 把女孩换成男孩/turn the girl into a boy\n- ❌ Don't translate between languages\n- ❌ Don't over-explain or add details\n\n**Prompt routing fix (2026-04-22):**\n- The Qwen workflow has TWO `TextEncodeQwenImageEditPlus` nodes:\n  - `115:110` — empty negative prompt node\n  - `115:111` — positive prompt node (contains default text like \"the girl\")\n- Script must route prompts to **node 111 (positive)**, not node 110\n- The `prepare_i2i_workflow()` function auto-detects by scanning for existing default text\n\n---\n\n## Wan2.2 I2V (Image → Video)\n\nWan2.2 generates short videos (~5 seconds) from a static image + motion description.\n\n**Rules:**\n- ✅ Prompt describes **actions/movement** (not scene description)\n- ✅ English motion descriptions tend to give the best results; the user's own language is acceptable — do not force a translation the user didn't ask for\n- ✅ Focus on \"who does what\" and \"how the camera moves\"\n- ❌ Don't describe static scene elements in motion prompt\n\n- Default: 81 frames (~5s @ 16fps), 4 steps, CFG 4.5\n- Base resolution: **560×720** (3:4, fast and OK quality)\n- Auto-detect input image aspect ratio and select reference resolution:\n\n### Resolution Reference\n\n| Aspect | Fast & OK | User Fav | WAN 2.2 Native |\n|--------|-----------|----------|----------------|\n| 3:4 | 560×720 | 720×912 | 848×1088 |\n| 2:3 | 528×768 | 656×960 | 784×1136 |\n| 9:16 | 480×848 | 608×1072 | 720×1264 |\n\nOther available resolutions:\n- **3:4**: 416×544, 672×864, 784×1008\n- **2:3**: 384×576, 624×912, 736×1072\n- **9:16**: 368×624, 576×1008, 672×1184\n\n**Examples:**\n```\nprompt: \"the cat walks forward and looks at the camera, tail wagging\"\nprompt: \"the girl smiles and turns her head, wind blowing her hair\"\nprompt: \"the person stands in a busy street, camera pans left and slowly zooms in, cars driving, red flag fluttering\"\n```\n\nFile v3.5.1:skill-card.md\n\n## Description:\n\nGenerate and edit images, animate images into videos, and turn character images into 3D meshes using a user-configured ComfyUI server.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[sunshinejnjn](https://clawhub.ai/user/sunshinejnjn)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nCreators and developers use this skill to request generated or edited images, short videos, and 3D character meshes from a ComfyUI server they configure.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: Prompts and selected source images are sent to the configured ComfyUI endpoint.\n\nMitigation: Keep the default localhost endpoint or use a server you control; avoid sending sensitive images to an untrusted endpoint.\n\nRisk: Generated media is saved to a local output directory.\n\nMitigation: Use a trusted output directory and review access to saved media.\n\nRisk: Manual installation of ComfyUI models or nodes can introduce unreviewed code or files.\n\nMitigation: Review model and node sources before installing; prefer environment variables for API keys.\n\n## Reference(s):\n\n- [ClawHub skill release](https://clawhub.ai/sunshinejnjn/skills/image-with-comfyui)\n- [Skill README](artifact/README.md)\n- [CLI usage reference](artifact/references/cli-reference.md)\n- [Qwen-Image 2.1 model files](https://huggingface.co/Comfy-Org/Qwen-Image-2.1/tree/main?recursive=true)\n\n## Skill Output:\n\n**Output Type(s):** [Images, Video, 3D mesh files]\n\n**Output Format:** [JPG or PNG images, video files, and ZIP-packaged GLB meshes]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [Generated files are saved locally for delivery to the user.]\n\n## Skill Version(s):\n\n3.5.1 (source: ClawHub release metadata)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nFile v3.5.1:config.json\n\n{\n  \"comfyui_url\": \"http://127.0.0.1:8188\",\n  \"comfyui_api_key\": \"\",\n  \"poll_interval_seconds\": 1,\n  \"keep_models_resident\": false,\n  \"output_dir\": \"media/comfyui\",\n  \"image\": {\n    \"output_format\": \"jpg\",\n    \"jpg_quality\": 89,\n    \"t2i\": {\n      \"z-image\": {\n        \"workflow_path\": \"workflows/z-image_t2i_api.json\",\n        \"timeout_seconds\": 100,\n        \"default_width\": 1280,\n        \"default_height\": 720,\n        \"default_steps\": 9,\n        \"output_subdir\": \"images\"\n      },\n      \"sd35\": {\n        \"workflow_path\": \"workflows/sd3.5-med_t2i_api.json\",\n        \"timeout_seconds\": 100,\n        \"default_width\": 1024,\n        \"default_height\": 1024,\n        \"default_steps\": 20,\n        \"default_cfg\": 4.01,\n        \"output_subdir\": \"images\"\n      },\n      \"qwen21\": {\n        \"workflow_path\": \"workflows/qwen_image_2.1_api.json\",\n        \"timeout_seconds\": 300,\n        \"default_aspect\": \"1:1\",\n        \"default_steps\": 25,\n        \"default_cfg\": 1.0,\n        \"output_subdir\": \"images\"\n      }\n    },\n    \"i2i\": {\n      \"qwen_imageedit\": {\n        \"workflow_path\": \"workflows/qwen_image-edit_api.json\",\n        \"timeout_seconds\": 600,\n        \"default_steps\": 4,\n        \"output_subdir\": \"images\"\n      },\n      \"qwen21\": {\n        \"workflow_path\": \"workflows/qwen_image_2.1_api.json\",\n        \"timeout_seconds\": 600,\n        \"default_aspect\": \"1:1\",\n        \"default_steps\": 25,\n        \"default_cfg\": 1.0,\n        \"output_subdir\": \"images\"\n      }\n    }\n  },\n  \"mesh\": {\n    \"hunyuan3d_v2.1\": {\n      \"workflow_path\": \"workflows/hunyuan3d-v2.1_i2m_api.json\",\n      \"timeout_seconds\": 900,\n      \"default_resolution\": 4096,\n      \"default_steps\": 30,\n      \"default_cfg\": 5.0,\n      \"default_scale\": 3000,\n      \"output_subdir\": \"meshes\",\n      \"extract\": {\n        \"enabled\": true,\n        \"workflow\": \"i2i\",\n        \"model\": \"qwen21\",\n        \"aspect\": \"1:1\",\n        \"prompt_suffix\": \"isolate the main character, full body, centered, plain transparent background, nothing else\"\n      }\n    }\n  },\n  \"video\": {\n    \"wan2.2_i2v\": {\n      \"workflow_path\": \"workflows/wan2.2_i2v_api.json\",\n      \"timeout_seconds\": 1000,\n      \"default_width\": 528,\n      \"default_height\": 768,\n      \"default_length\": 81,\n      \"default_steps\": 4,\n      \"default_cfg\": 4.5,\n      \"frame_rate\": 16,\n      \"output_subdir\": \"video\",\n      \"keep_interpolated_only\": false\n    }\n  },\n  \"default_model\": {\n    \"t2i\": \"qwen21\",\n    \"i2i\": \"qwen21\"\n  }\n}\n\nFile v3.5.1:workflows/hunyuan3d-v2.1_i2m_api.json\n\n{\n  \"1\": {\n    \"inputs\": {\n      \"ckpt_name\": \"hunyuan_3d_v2.1.safetensors\"\n    },\n    \"class_type\": \"ImageOnlyCheckpointLoader\",\n    \"_meta\": {\n      \"title\": \"Load Checkpoint Image Only (Hunyuan 3D v2.1)\"\n    }\n  },\n  \"2\": {\n    \"inputs\": {\n      \"image\": \"character.png\"\n    },\n    \"class_type\": \"LoadImage\",\n    \"_meta\": {\n      \"title\": \"Load Image\"\n    }\n  },\n  \"3\": {\n    \"inputs\": {\n      \"shift\": 1,\n      \"sampling\": \"flow\",\n      \"model\": [\n        \"1\",\n        0\n      ]\n    },\n    \"class_type\": \"ModelSamplingAuraFlow\",\n    \"_meta\": {\n      \"title\": \"ModelSamplingAuraFlow\"\n    }\n  },\n  \"4\": {\n    \"inputs\": {\n      \"resolution\": 4096,\n      \"batch_size\": 1\n    },\n    \"class_type\": \"EmptyLatentHunyuan3Dv2\",\n    \"_meta\": {\n      \"title\": \"Empty Hunyuan 3D v2 Latent\"\n    }\n  },\n  \"6\": {\n    \"inputs\": {\n      \"clip_vision_output\": [\n        \"22\",\n        0\n      ]\n    },\n    \"class_type\": \"Hunyuan3Dv2Conditioning\",\n    \"_meta\": {\n      \"title\": \"Hunyuan3Dv2Conditioning\"\n    }\n  },\n  \"7\": {\n    \"inputs\": {\n      \"seed\": 384326697438209,\n      \"steps\": 30,\n      \"cfg\": 5,\n      \"sampler_name\": \"euler\",\n      \"scheduler\": \"normal\",\n      \"denoise\": 1,\n      \"model\": [\n        \"3\",\n        0\n      ],\n      \"positive\": [\n        \"6\",\n        0\n      ],\n      \"negative\": [\n        \"6\",\n        1\n      ],\n      \"latent_image\": [\n        \"4\",\n        0\n      ]\n    },\n    \"class_type\": \"KSampler\",\n    \"_meta\": {\n      \"title\": \"KSampler\"\n    }\n  },\n  \"8\": {\n    \"inputs\": {\n      \"num_chunks\": 8000,\n      \"octree_resolution\": 256,\n      \"samples\": [\n        \"19\",\n        0\n      ],\n      \"vae\": [\n        \"1\",\n        2\n      ]\n    },\n    \"class_type\": \"VAEDecodeHunyuan3D\",\n    \"_meta\": {\n      \"title\": \"Hunyuan 3D VAE Decode\"\n    }\n  },\n  \"9\": {\n    \"inputs\": {\n      \"algorithm\": \"surface net\",\n      \"threshold\": 0.6,\n      \"voxel\": [\n        \"21\",\n        0\n      ]\n    },\n    \"class_type\": \"VoxelToMesh\",\n    \"_meta\": {\n      \"title\": \"Voxel to Mesh\"\n    }\n  },\n  \"10\": {\n    \"inputs\": {\n      \"filename_prefix\": \"mesh/hy21\",\n      \"mesh\": [\n        \"20\",\n        0\n      ]\n    },\n    \"class_type\": \"SaveGLB\",\n    \"_meta\": {\n      \"title\": \"Save 3D Model\"\n    }\n  },\n  \"13\": {\n    \"inputs\": {\n      \"crop\": \"center\",\n      \"clip_vision\": [\n        \"1\",\n        1\n      ],\n      \"image\": [\n        \"2\",\n        0\n      ]\n    },\n    \"class_type\": \"CLIPVisionEncode\",\n    \"_meta\": {\n      \"title\": \"CLIP Vision Encode\"\n    }\n  },\n  \"19\": {\n    \"inputs\": {\n      \"value\": [\n        \"7\",\n        0\n      ],\n      \"model\": [\n        \"1\",\n        0\n      ]\n    },\n    \"class_type\": \"UnloadModel\",\n    \"_meta\": {\n      \"title\": \"UnloadModel\"\n    }\n  },\n  \"20\": {\n    \"inputs\": {\n      \"value\": [\n        \"9\",\n        0\n      ]\n    },\n    \"class_type\": \"UnloadAllModels\",\n    \"_meta\": {\n      \"title\": \"UnloadAllModels\"\n    }\n  },\n  \"21\": {\n    \"inputs\": {\n      \"value\": [\n        \"8\",\n        0\n      ],\n      \"model\": [\n        \"1\",\n        2\n      ]\n    },\n    \"class_type\": \"UnloadModel\",\n    \"_meta\": {\n      \"title\": \"UnloadModel\"\n    }\n  },\n  \"22\": {\n    \"inputs\": {\n      \"value\": [\n        \"13\",\n        0\n      ],\n      \"model\": [\n        \"1\",\n        1\n      ]\n    },\n    \"class_type\": \"UnloadModel\",\n    \"_meta\": {\n      \"title\": \"UnloadModel\"\n    }\n  }\n}\n\nArchive v3.5.0: 19 files, 65887 bytes\n\nFiles: config.json (2438b), image_with_comfyui.py (108658b), MODEL_URL.txt (7015b), README.md (17687b), references/attribution.md (2476b), references/cli-reference.md (7159b), references/delivery.md (2435b), references/error-handling.md (4382b), references/image-first.md (2231b), references/prompts.md (6357b), skill-card.md (2206b), SKILL.md (11518b), workflows/hunyuan3d-v2.1_i2m_api.json (3286b), workflows/qwen_image_2.1_api.json (3702b), workflows/qwen_image-edit_api.json (4241b), workflows/sd3.5-med_t2i_api.json (2063b), workflows/wan2.2_i2v_api.json (8022b), workflows/z-image_t2i_api.json (2674b), _meta.json (137b)\n\nFile v3.5.0:SKILL.md\n\n---\nname: image-with-comfyui\ndescription: 'Use to generate, edit, or animate images and videos through a user-defined ComfyUI server (COMFYUI_URL env var or comfyui_url in config.json; if no server is reachable the script exits with configuration instructions) — trigger with requests like \"make a picture of\", \"replace the background\", \"turn this into a video\", \"edit my photo\", or \"make a 3D mesh of this character\". Not for simple non-generation image tweaks. **Qwen-Image 2.1** is the default model for both text-to-image (T2I) and image editing/multi-image (I2I); Z-Image / SD3.5 (T2I), Qwen Image Edit (2511, single-image), Wan2.2 (I2V), and Hunyuan 3D v2.1 (I2M, image → 3D GLB mesh) are available on request.'\nversion: \"3.5.0\"\nlicense: BSD-3-Clause\ncompatibility: 'ComfyUI ≥ 0.37.0; requires python3; the endpoint is user-defined (COMFYUI_URL or config.json) — point it at a server you control and trust.'\nallowed-tools: Read, Write, Edit, Exec\nmetadata:\n  openclaw:\n    emoji: \"🎨\"\n    requires:\n      anyBins: [\"python3\"]\n    config:\n      path: \"config.json\"\n---\n\n# Image with ComfyUI\n\nCall a user-defined ComfyUI server to generate or edit images and videos, or turn an image into a 3D mesh. **Qwen-Image 2.1** is the default model for both T2I and I2I; Z-Image / SD3.5 (T2I), Qwen Image Edit 2511 (single-image edit), Wan2.2 (I2V), and Hunyuan 3D v2.1 (I2M, image → 3D GLB) are available on request. If no ComfyUI server is reachable, the script exits with configuration instructions (see [README.md](README.md)).\n\n- **T2I** (Text → Image) → **Qwen-Image 2.1** (default) or Z-Image / SD3.5 Medium\n- **I2I** (Image → Image / Edit / Multi-image, up to 16 refs) → **Qwen-Image 2.1** (default) or Qwen Image Edit (2511, single image)\n- **I2V** (Image → Video) → Wan2.2 model\n- **I2M** (Image → 3D Mesh) → **Hunyuan 3D v2.1** — extracts the main character first (transparent PNG), generates a GLB, scales it ×3000, and zips it for delivery\n\n`config.json` at the skill root holds every default; any value can be overridden by an env var (see the Data & Privacy table there).\n\n## When to Use\n\n- Generate images from text → **t2i**\n- Edit one image, or compose 2–16 reference images → **i2i** (`--image` order = slot order = `<image1>`, `<image2>` …)\n- Turn an image into a short video → **wan2.2**\n- Turn the character in an image into a **3D mesh (GLB)** → **i2m** (Hunyuan 3D v2.1; the main character is extracted first, see below)\n- The user sends an image, then within 2 minutes asks to edit or animate it → **image-first mode**, see [references/image-first.md](references/image-first.md)\n\n## Data & Privacy (read before running)\n\n- **What is transmitted:** prompts and the workflow JSON go to the configured endpoint for **every** mode; **I2I, I2V and I2M additionally upload the full source image**.\n- **Endpoint** is user-defined (`COMFYUI_URL` env var > `comfyui_url` in `config.json`; bundled default `http://127.0.0.1:8188`). A non-local host prints a data-flow warning on every run; if no server answers, the script exits with configuration instructions instead of a raw timeout.\n- **Written to disk:** generated media is saved under `output_dir` (`config.json` or `COMFYUI_OUTPUT_DIR`). Don't submit sensitive content if the endpoint or disk is a concern.\n- **Secrets:** never put credentials in the URL or logs.\n\n## Workflow Files\n\n| Mode | Workflow | File |\n|---|---|---|\n| T2I (**Qwen-Image 2.1, default**) | Qwen-Image 2.1 dual-mode | `workflows/qwen_image_2.1_api.json` |\n| I2I (**Qwen-Image 2.1, default**, incl. multi-image) | same file, switch ON | `workflows/qwen_image_2.1_api.json` |\n| T2I (Z-Image) | Z-Image | `workflows/z-image_t2i_api.json` |\n| T2I (SD3.5) | SD3.5 Medium | `workflows/sd3.5-med_t2i_api.json` |\n| I2I (legacy) | Qwen Image Edit (2511) | `workflows/qwen_image-edit_api.json` |\n| I2V | Wan2.2 Image-to-Video | `workflows/wan2.2_i2v_api.json` |\n| I2M (Image → 3D Mesh) | Hunyuan 3D v2.1 | `workflows/hunyuan3d-v2.1_i2m_api.json` |\n\nEach workflow is an **adapted** version of a real ComfyUI graph. Original source + exactly what was changed → [references/attribution.md](references/attribution.md).\n\n## CLI Usage\n\n> **⚠️ The prompt is MANDATORY and must be passed via `--prompt` — never as a positional word.**\n>\n> Every subcommand (`t2i`, `i2i`, `wan2.2`) reads the prompt from `--prompt \"...\"`. Passing the prompt as a bare positional word fails immediately with exit code 2:\n> `python3 image_with_comfyui.py t2i \"some description\"` → `t2i: error: the following arguments are required: --prompt`\n>\n> This is a **100%-avoidable failure loop** — the command never reaches ComfyUI, so the model can't help. Always write the prompt through `--prompt` on the **first** attempt.\n>\n> `i2m` takes no prompt — the character itself is the condition (`--image` is its only required flag).\n\nFull commands and the timeout table → [references/cli-reference.md](references/cli-reference.md)\n\n## Prompt Formatting\n\nPer-model guidance (formulas, rules, example prompts) lives in [references/prompts.md](references/prompts.md). The English prose is there; example prompts keep their bilingual keyword pairs where the original was bilingual (e.g. `blurry 模糊`).\n\n- **Qwen-Image 2.1 (T2I + I2I default):** natural-language prompts. **CFG locked to 1.0** (distilled) — never pass `--cfg`. Optional `--negative` supported in both T2I and I2I (keep it short); I2I prompts stay concise, in the user's language.\n- **SD3.5 Medium:** natural language + optional `--negative`; default **1:1 (1024×1024)**, 20 steps, CFG 4.01.\n- **Z-Image / Qwen-Image 2.1:** `--aspect` is **optional** — T2I omits it and the output defaults to **1:1 (1024×1024)**; I2I omits it and the output canvas matches the **first reference image's size** (workflow latent switch ON). Pass `--aspect 16:9`, `9:16`, etc. to force a shape — in I2I this flips the latent switch to route the canvas through the ResolutionSelector (reference stays at native size, never stretched).\n- **Z-Image:** natural language — **no negative prompts**.\n- **Qwen Image Edit (2511):** concise, positive-only, user's exact words.\n- **Wan2.2 (I2V):** motion/action prompts; default 81 frames, 4 steps, CFG 4.5.\n- **i2m (Hunyuan 3D v2.1, no prompt):** the only input is the character image. Default extraction is **i2i** (Qwen-Image 2.1 redraws the main character centered on a transparent background, saved as PNG — alpha is preserved end-to-end: the Hunyuan 3D CLIP-vision conditioner drops the alpha channel itself, and the output is a GLB, never flattened). Use `--extract rembg` for a non-generative u2net human cutout, or `--extract none` to feed the image straight in (only if it's already a clean cutout). Defaults: resolution 4096, 30 steps, CFG 5.0, delivery scale **×3000** — the raw Hunyuan 3D v2.1 mesh is at a normalized unit scale (character ≈ 2 units, no declared units), so the GLB is scaled in place before zipping (`--scale 1` keeps it unscaled).\n\n## Error Handling\n\nThe script auto-detects three categories of server-side problems and reports them with install/source guidance, and `t2i`/`i2i` run through a **runtime fallback chain** when the default model fails. Details → [references/error-handling.md](references/error-handling.md):\n\n1. **Missing nodes** — reports which node, its package, and a pinned GitHub install hint (the script never runs `git clone` itself).\n2. **Missing models** — substitutes a compatible model (e.g. SD3.5 medium variants → SD3.5 large).\n3. **Non-critical utility nodes** (e.g. `UnloadAllModels`) — bypassed automatically without interrupting generation.\n4. **Runtime fallback chain** — `t2i`/`i2i` auto-descend through alternate models (and, for t2i, any available checkpoint) on runtime failures; config problems (missing files/nodes) abort instead of degrading. Disable with `--no-fallback`.\n\n## Output Delivery\n\n**Send the generated file — do not just describe it.** Deliver in the **user's original session**, never raw paths/URLs. Per-channel prefixes and the staging/cleanup routine → [references/delivery.md](references/delivery.md)\n\n**Output format:** generated images default to **JPG** (`config.json` → `image.output_format`). **Routing rule: if the user asks for a PNG image/file, or an RGBA image with transparency/alpha, always pass `--format png` and deliver the PNG file as-is — never convert to JPG** (JPEG has no alpha channel; the conversion would flatten transparency onto white). PNG output passes the server's PNG bytes through unchanged, so transparency is preserved byte-for-byte. Otherwise `--format jpg --jpg-quality N` is available per run; JPG conversion happens client-side (ComfyUI always saves PNG server-side). **i2m delivery rules (3D output):** (1) attach the **extracted character image** from step 1 as an **image** — never just its path (with `--extract none` the step-1 image *is* the user's own input, so there's nothing to send back); (2) send the mesh **packaged as a zip** — the script first **scales the GLB in place by the delivery factor** (default ×3000; the raw mesh is at a normalized unit scale, character ≈ 2 units, no declared units), then zips it (`i2m_<timestamp>.zip`); attach the **zip**, never the raw `.glb`; (3) **never render the mesh** — no GLB → image/video conversion and no 3D-viewer screenshots (a third-party viewer link is fine to share). Format conversion never applies to 3D files.\n\n## Execution Discipline — Fast Path\n\nThe **measured GPU sampling time is ~12–20s**. When a task takes minutes, the delay is **not** inference — it's either a model cold-reload or agent-side overhead. Kill the waste:\n\n**i2m exception:** 3D generation is legitimately long (cold checkpoint load + 30 sampling steps + voxel decode ≈ minutes on a 4090). Run it once in the background (`exec` with a generous `timeoutSeconds`), report the queue state, and when it completes deliver the character image + the GLB zip per the i2m delivery rules above — don't spin tool calls around it.\n\n1. **Never re-read on repeat.** Don't re-read `SKILL.md`, `config.json`, or the script for a known config / repeat / same-model job. Defaults are baked in; `--aspect` / `--steps` / `--cfg` are in the links above.\n2. **At most one pre-check merge.** Gather any needed info in a single batched command, or skip it entirely for trivial repeat jobs.\n3. **Run immediately after one check.** Don't inspect, second-guess, or add tool calls between knowing the prompt and running.\n4. **Verify only the minimum.** After the run, check only what the user asked (size / seed / file exists).\n5. **Golden rule:** sampling is ~12–20s. Anything longer = reload or overhead — never pile on tool calls before the image exists.\n6. **Prompt-flag trap:** pass `--prompt \"...\"` on the first attempt; a positional word wastes a whole turn.\n\n## References (index)\n\n- [references/image-first.md](references/image-first.md) — image-then-edit detection rules + multi-image reference composition\n- [references/attribution.md](references/attribution.md) — original source + modifications per workflow\n- [references/cli-reference.md](references/cli-reference.md) — full commands + timeout reference\n- [references/prompts.md](references/prompts.md) — per-model prompt guidance\n- [references/error-handling.md](references/error-handling.md) — auto-handled error categories\n- [references/delivery.md](references/delivery.md) — per-channel delivery + staging\n\nFile v3.5.0:README.md\n\n# Image with ComfyUI\n\nCall a user-defined ComfyUI server (local or remote) for **text-to-image**, **image-to-image/edit**, **image-to-video**, and **image-to-3D-mesh** generation. **Qwen-Image 2.1** is the default model for both T2I and image editing (including multi-image reference composition); Z-Image, SD3.5 Medium, Qwen Image Edit (2511), Wan2.2, and Hunyuan 3D v2.1 (3D mesh) are available as options.\n\n> **New here?** Qwen-Image 2.1 needs a one-time ComfyUI setup — see [Qwen-Image 2.1 ComfyUI Setup](#qwen-image-21-comfyui-setup).\n\n## Quick Start\n\n1. **Ensure a ComfyUI server is running** (locally, or on a host you can reach).\n2. **Set the URL** (via environment variable or `config.json`). The endpoint is user-defined — the bundled default is `http://127.0.0.1:8188`; point `COMFYUI_URL` at any ComfyUI server you control — a data-flow warning is printed on every run for non-local hosts. If no server is reachable, the script exits with configuration instructions:\n   ```bash\n   export COMFYUI_URL=http://comfyui.host:api-port\n   ```\n3. **Test a workflow** (see [Testing](#testing) below).\n\n## Configuration\n\nRead `config.json` at the skill root. All values can be overridden by environment variables:\n\n| Env Variable | Overrides | Default |\n|---|---|---|\n| `COMFYUI_URL` | `comfyui_url` | `http://localhost:8188` |\n| `COMFYUI_TIMEOUT` | `timeout_seconds` | `120` |\n| `COMFYUI_POLL_INTERVAL` | `poll_interval_seconds` | `1` |\n| `COMFYUI_KEEP_MODELS_RESIDENT` | `keep_models_resident` | `false` |\n| `COMFYUI_OUTPUT_DIR` | `output_dir` | `media/comfyui` |\n| `COMFYUI_APIKEY` | `comfyui_api_key` | _(none → no auth header)_ |\n\n**Priority:** Environment variables > `config.json`.\n\n> Note: The config.json key `comfyui_api_key` (string) is optional. When empty/absent, **no** `Authorization` header is added and behavior is unchanged. When set, every ComfyUI API request carries `Authorization: Bearer <key>` (env var wins over the config value).\n\n### 🔒 ComfyUI API Key (Bearer Auth)\n\nIf your ComfyUI instance requires an API key, set it via either of these (env var takes precedence):\n\n```bash\nexport COMFYUI_APIKEY=sk-your-key\n```\n\nor, in `config.json`:\n\n```json\n\"comfyui_api_key\": \"sk-your-key\"\n```\n\nWhen present, the skill adds `Authorization: Bearer <key>` to all ComfyUI API calls (`/prompt`, `/history`, `/object_info`, `/upload/image`, `/view`, `/upload/*`). Do not commit your key to version control — prefer the env var, and never paste it into the chat or logs.\n\n> **⚠️ This skill only *sends* the key — it doesn't authenticate your ComfyUI itself.** For the Bearer key to be honored, the ComfyUI server must run an authentication custom node that recognizes that key. The reference implementation is [`sunshinejnjn/comfyui-auth`](https://github.com/sunshinejnjn/comfyui-auth) (a fork of the original [`ivellioscolin/comfyui-auth`](https://github.com/ivellioscolin/comfyui-auth)). To set it up:\n> 1. Clone it into your ComfyUI: `git clone https://github.com/sunshinejnjn/comfyui-auth.git ComfyUI/custom_nodes/comfyui-auth`\n> 2. Edit its `apikeys.conf` — add a `label = <your-key>` line (the key is read **as written**, plain or as configured), then restart ComfyUI (the plugin reloads its config per request, but a re-clone/restart makes the new code load).\n> 3. Point this skill at that server via `COMFYUI_URL` and set `COMFYUI_APIKEY` to that same key.\n>\n> Without the plugin installed on the server, the key is silently ignored and the API responds as if no auth is required. The `test` command (`python3 image_with_comfyui.py test`) will tell you whether the endpoint answers before you try generating.\n\n### Data & Privacy\n\n- **Transmitted to the endpoint**: prompts and workflow JSON (all modes); **full source images** for I2I and I2V.\n- **Endpoint** is user-defined (`COMFYUI_URL` env var > `comfyui_url` in `config.json`; bundled default `http://127.0.0.1:8188`). A non-local endpoint prints a warning on every run, and an unreachable server exits with configuration instructions (this section + [Qwen-Image 2.1 ComfyUI Setup](#qwen-image-21-comfyui-setup)). Use HTTPS for remote servers and never put credentials in the URL or logs.\n- **Written to disk**: generated media is saved under `output_dir` (config) or `COMFYUI_OUTPUT_DIR` in your workspace; delivery staging copies live in `~/.openclaw/media/outbound/` and are removed after a successful send. If your content is sensitive, review the endpoint and output location before running.\n\n## Models & Custom Nodes\n\nThis skill uses th\n\nArchive v3.2.0: 19 files, 64429 bytes\n\nFiles: config.json (2413b), image_with_comfyui.py (107267b), MODEL_URL.txt (7015b), README.md (15664b), references/attribution.md (2476b), references/cli-reference.md (7159b), references/delivery.md (2435b), references/error-handling.md (4382b), references/image-first.md (2231b), references/prompts.md (6357b), skill-card.md (1929b), SKILL.md (11518b), workflows/hunyuan3d-v2.1_i2m_api.json (3286b), workflows/qwen_image_2.1_api.json (3702b), workflows/qwen_image-edit_api.json (4241b), workflows/sd3.5-med_t2i_api.json (2063b), workflows/wan2.2_i2v_api.json (8022b), workflows/z-image_t2i_api.json (2674b), _meta.json (137b)\n\nArchive v3.1.1: 19 files, 62455 bytes\n\nFiles: config.json (2384b), image_with_comfyui.py (102782b), MODEL_URL.txt (7015b), README.md (15564b), references/attribution.md (2476b), references/cli-reference.md (6788b), references/delivery.md (2229b), references/error-handling.md (4382b), references/image-first.md (2231b), references/prompts.md (6357b), skill-card.md (2070b), SKILL.md (11121b), workflows/hunyuan3d-v2.1_i2m_api.json (3286b), workflows/qwen_image_2.1_api.json (3702b), workflows/qwen_image-edit_api.json (4241b), workflows/sd3.5-med_t2i_api.json (2063b), workflows/wan2.2_i2v_api.json (8022b), workflows/z-image_t2i_api.json (2674b), _meta.json (137b)\n\nArchive v3.1.0: 19 files, 62399 bytes\n\nFiles: config.json (2384b), image_with_comfyui.py (102782b), MODEL_URL.txt (7015b), README.md (15547b), references/attribution.md (2476b), references/cli-reference.md (6757b), references/delivery.md (2166b), references/error-handling.md (4382b), references/image-first.md (2231b), references/prompts.md (6357b), skill-card.md (2137b), SKILL.md (11102b), workflows/hunyuan3d-v2.1_i2m_api.json (3286b), workflows/qwen_image_2.1_api.json (3702b), workflows/qwen_image-edit_api.json (4241b), workflows/sd3.5-med_t2i_api.json (2063b), workflows/wan2.2_i2v_api.json (8022b), workflows/z-image_t2i_api.json (2674b), _meta.json (137b)\n\nArchive v3.0.0: 19 files, 60953 bytes\n\nFiles: config.json (2384b), image_with_comfyui.py (101790b), MODEL_URL.txt (7015b), README.md (15059b), references/attribution.md (2476b), references/cli-reference.md (6299b), references/delivery.md (1432b), references/error-handling.md (4382b), references/image-first.md (2231b), references/prompts.md (6357b), skill-card.md (1830b), SKILL.md (10591b), workflows/hunyuan3d-v2.1_i2m_api.json (3286b), workflows/qwen_image_2.1_api.json (3702b), workflows/qwen_image-edit_api.json (4241b), workflows/sd3.5-med_t2i_api.json (2063b), workflows/wan2.2_i2v_api.json (8022b), workflows/z-image_t2i_api.json (2674b), _meta.json (137b)\n\nArchive v2.5.0: 18 files, 54554 bytes\n\nFiles: config.json (1873b), image_with_comfyui.py (86217b), MODEL_URL.txt (6079b), README.md (14466b), references/attribution.md (1887b), references/cli-reference.md (5049b), references/delivery.md (1432b), references/error-handling.md (4382b), references/image-first.md (2231b), references/prompts.md (6357b), skill-card.md (2130b), SKILL.md (9007b), workflows/qwen_image_2.1_api.json (3702b), workflows/qwen_image-edit_api.json (4241b), workflows/sd3.5-med_t2i_api.json (2063b), workflows/wan2.2_i2v_api.json (8022b), workflows/z-image_t2i_api.json (2674b), _meta.json (137b)\n\nArchive v2.4.1: 18 files, 54185 bytes\n\nFiles: config.json (1873b), image_with_comfyui.py (86217b), MODEL_URL.txt (6079b), README.md (14466b), references/attribution.md (1887b), references/cli-reference.md (4566b), references/delivery.md (1432b), references/error-handling.md (4382b), references/image-first.md (2231b), references/prompts.md (6357b), skill-card.md (2166b), SKILL.md (8648b), workflows/qwen_image_2.1_api.json (3702b), workflows/qwen_image-edit_api.json (4241b), workflows/sd3.5-med_t2i_api.json (2063b), workflows/wan2.2_i2v_api.json (8022b), workflows/z-image_t2i_api.json (2674b), _meta.json (137b)\n\nArchive v2.4.0: 18 files, 53160 bytes\n\nFiles: config.json (1873b), image_with_comfyui.py (84497b), MODEL_URL.txt (6079b), README.md (14147b), references/attribution.md (1887b), references/cli-reference.md (4566b), references/delivery.md (1432b), references/error-handling.md (4045b), references/image-first.md (2231b), references/prompts.md (6357b), skill-card.md (1852b), SKILL.md (8286b), workflows/qwen_image_2.1_api.json (3702b), workflows/qwen_image-edit_api.json (4241b), workflows/sd3.5-med_t2i_api.json (2063b), workflows/wan2.2_i2v_api.json (8022b), workflows/z-image_t2i_api.json (2674b), _meta.json (137b)","readmeExcerpt":"Skill: image-with-comfyui Owner: sunshinejnjn Summary: Use to generate, edit, or animate images and videos through a user-defined ComfyUI server (COMFYUI_URL env var or comfyui_url in config.json; if no server is reachable the script exits with configuration instructions) — trigger with requests like \"make a picture of\", \"replace the background\", \"cut out / extract the subject of\", \"turn this into a video\", \"edit my ","codeSnippets":[],"executableExamples":[{"language":"bash","snippet":"export COMFYUI_URL=http://comfyui.host:api-port"},{"language":"bash","snippet":"export COMFYUI_APIKEY=sk-your-key"},{"language":"json","snippet":"\"comfyui_api_key\": \"sk-your-key\""},{"language":"bash","snippet":"cd ComfyUI/custom_nodes\ngit clone https://github.com/ltdrdata/ComfyUI-Impact-Pack.git\ncd ComfyUI-Impact-Pack && git checkout --detach 429d0159ad429e64d2b3916e6e7be9c22d025c3c\npython3 -m venv .venv && .venv/bin/pip install -r requirements.txt"},{"language":"bash","snippet":"python3 image_with_comfyui.py test          # server reachable?\npython3 image_with_comfyui.py t2i --prompt \"A red fox in snow\" --aspect 1:1\npython3 image_with_comfyui.py i2i --prompt \"Change the sky to sunset\" --image /path/in.jpg"},{"language":"text","snippet":"Put the garment from <image2> onto the person in <image1>, keep <image1>'s pose and lighting.\nCombine both photos into one scene: subject of <image1> on the left, subject of <image2> on the right."}],"parameters":null,"dependencies":[],"permissions":[],"extractedFiles":[{"path":"SKILL.md","content":"---\nname: image-with-comfyui\ndescription: 'Use to generate, edit, or animate images and videos through a user-defined ComfyUI server (COMFYUI_URL env var or comfyui_url in config.json; if no server is reachable the script exits with configuration instructions) — trigger with requests like \"make a picture of\", \"replace the background\", \"cut out / extract the subject of\", \"turn this into a video\", \"edit my photo\", or \"make a 3D mesh of this character\". Not for simple non-generation image tweaks. **Qwen-Image 2.1** is the default model for both text-to-image (T2I) and image editing/multi-image (I2I); it can also **cut out / extract the main subject or any image element into a transparent-background image**. Z-Image / SD3.5 (T2I), Qwen Image Edit (2511, single-image), Wan2.2 (I2V), and Hunyuan 3D v2.1 (I2M, image → 3D GLB mesh) are available on request.'\nversion: \"3.5.0\"\nlicense: BSD-3-Clause\ncompatibility: 'ComfyUI ≥ 0.37.0; requires python3; the endpoint is user-defined (COMFYUI_URL or config.json) — point it at a server you control and trust.'\nallowed-tools: Read, Write, Edit, Exec\nmetadata:\n  openclaw:\n    emoji: \"🎨\"\n    requires:\n      anyBins: [\"python3\"]\n    config:\n      path: \"config.json\"\n---\n\n# Image with ComfyUI\n\nCall a user-defined ComfyUI server to generate or edit images and videos, or turn an image into a 3D mesh. **Qwen-Image 2.1** is the default model for both T2I and I2I; Z-Image / SD3.5 (T2I), Qwen Image Edit 2511 (single-image edit), Wan2.2 (I2V), and Hunyuan 3D v2.1 (I2M, image → 3D GLB) are available on request. If no ComfyUI server is reachable, the script exits with configuration instructions (see [README.md](README.md)).\n\n- **T2I** (Text → Image) → **Qwen-Image 2.1** (default) or Z-Image / SD3.5 Medium\n- **I2I** (Image → Image / Edit / Multi-image, up to 16 refs) → **Qwen-Image 2.1** (default) or Qwen Image Edit (2511, single image)\n- **I2V** (Image → Video) → Wan2.2 model\n- **Cutout** (extract the main subject or any image element) → **Qwen-Image 2.1** — redraws the target on a transparent background, producing a clean, alpha-only image; see [Cutout below](#cutout-extract-the-subject).\n- **I2M** (Image → 3D Mesh) → **Hunyuan 3D v2.1** — extracts the main character first (transparent PNG), generates a GLB, scales it ×3000, and zips it for delivery\n\n`config.json` at the skill root holds every default; any value can be overridden by an env var (see the Data & Privacy table there).\n\n## Cutout — Extract the Subject\n\n> **Note:** cutout is a genuine image-generation task, not a simple non-generation tweak — use it when the user wants to isolate the subject or any image element from its surroundings.\n\n**Qwen-Image 2.1** can **cut out / extract** the main subject (or a specified element) from an image and return it on a transparent background. The output is a clean PNG with alpha only — it re-draws the target rather than doing a hard pixel mask, so it works well for organic shapes (people, animals, products, characters) and can tar"},{"path":"README.md","content":"# Image with ComfyUI\n\nCall a user-defined ComfyUI server (local or remote) for **text-to-image**, **image-to-image/edit**, **image-to-video**, and **image-to-3D-mesh** generation. **Qwen-Image 2.1** is the default model for both T2I and image editing (including multi-image reference composition); Z-Image, SD3.5 Medium, Qwen Image Edit (2511), Wan2.2, and Hunyuan 3D v2.1 (3D mesh) are available as options.\n\n> **New here?** Qwen-Image 2.1 needs a one-time ComfyUI setup — see [Qwen-Image 2.1 ComfyUI Setup](#qwen-image-21-comfyui-setup).\n\n## Quick Start\n\n1. **Ensure a ComfyUI server is running** (locally, or on a host you can reach).\n2. **Set the URL** (via environment variable or `config.json`). The endpoint is user-defined — the bundled default is `http://127.0.0.1:8188`; point `COMFYUI_URL` at any ComfyUI server you control — a data-flow warning is printed on every run for non-local hosts. If no server is reachable, the script exits with configuration instructions:\n   ```bash\n   export COMFYUI_URL=http://comfyui.host:api-port\n   ```\n3. **Test a workflow** (see [Testing](#testing) below).\n\n## Configuration\n\nRead `config.json` at the skill root. All values can be overridden by environment variables:\n\n| Env Variable | Overrides | Default |\n|---|---|---|\n| `COMFYUI_URL` | `comfyui_url` | `http://localhost:8188` |\n| `COMFYUI_TIMEOUT` | `timeout_seconds` | `120` |\n| `COMFYUI_POLL_INTERVAL` | `poll_interval_seconds` | `1` |\n| `COMFYUI_KEEP_MODELS_RESIDENT` | `keep_models_resident` | `false` |\n| `COMFYUI_OUTPUT_DIR` | `output_dir` | `media/comfyui` |\n| `COMFYUI_APIKEY` | `comfyui_api_key` | _(none → no auth header)_ |\n\n**Priority:** Environment variables > `config.json`.\n\n> Note: The config.json key `comfyui_api_key` (string) is optional. When empty/absent, **no** `Authorization` header is added and behavior is unchanged. When set, every ComfyUI API request carries `Authorization: Bearer <key>` (env var wins over the config value).\n\n### 🔒 ComfyUI API Key (Bearer Auth)\n\nIf your ComfyUI instance requires an API key, set it via either of these (env var takes precedence):\n\n```bash\nexport COMFYUI_APIKEY=sk-your-key\n```\n\nor, in `config.json`:\n\n```json\n\"comfyui_api_key\": \"sk-your-key\"\n```\n\nWhen present, the skill adds `Authorization: Bearer <key>` to all ComfyUI API calls (`/prompt`, `/history`, `/object_info`, `/upload/image`, `/view`, `/upload/*`). Do not commit your key to version control — prefer the env var, and never paste it into the chat or logs.\n\n> **⚠️ This skill only *sends* the key — it doesn't authenticate your ComfyUI itself.** For the Bearer key to be honored, the ComfyUI server must run an authentication custom node that recognizes that key. The reference implementation is [`sunshinejnjn/comfyui-auth`](https://github.com/sunshinejnjn/comfyui-auth) (a fork of the original [`ivellioscolin/comfyui-auth`](https://github.com/ivellioscolin/comfyui-auth)). To set it up:\n> 1. Clone it into your ComfyUI: `git clone https://github.com/sunshinejnjn/comfyui"},{"path":"_meta.json","content":"{\n  \"ownerId\": \"kn7dn131pvgfaveyzzzcrc5xf182yb84\",\n  \"slug\": \"image-with-comfyui\",\n  \"version\": \"3.6.0\",\n  \"publishedAt\": 1791136664035\n}"},{"path":"references/attribution.md","content":"# Source Attribution (image-with-comfyui)\n\nEvery workflow in `workflows/` is an **adapted** version of a real ComfyUI\nworkflow. Original source + the specific modifications applied by this skill are\ndocumented here. The workflow JSONs carry only ComfyUI's standard per-node\nlabels (`_meta` `title`); they do **not** embed provenance or attribution\nmetadata, so the original source is not duplicated inside them — it lives here\ninstead, and each file stays runnable.\n\n| Workflow | Original source | Key modifications applied |\n|---|---|---|\n| `qwen_image_2.1_api.json` | [foprc/qwen-image-2.1-comfyui-workflow](https://github.com/foprc/qwen-image-2.1-comfyui-workflow) (ships GGUF Q8_0 + ComfyUI-GGUF) | Replaced GGUF model with standard **safetensors** via `UNETLoader` (`weight_dtype: fp8_e4m3fn_fast`) — removes the ComfyUI-GGUF dependency; added `UnloadModel` (node 18) + `UnloadAllModels` (node 19) VRAM-cleanup; dual-mode reference switch (nodes 12/13/14); `ResolutionSelector` (node 9); API/`timestamp` output |\n| `qwen_image-edit_api.json` | Qwen-Image-Edit (Comfy-Org) default edit graph | Prompt auto-routing to the positive `TextEncodeQwenImageEditPlus` node (115:111) by the script |\n| `sd3.5-med_t2i_api.json` | Comfy-Org / sd3-medium default graph | Fixed the `.safetensors.safetensors` double-suffix `ckpt_name` bug; `sd3_medium` → `sd3.5_large.safetensors` fallback; `UnloadAllModels` VRAM-cleanup |\n| `z-image_t2i_api.json` | Comfy-Org / z-image default graph | `UnloadAllModels` VRAM-cleanup |\n| `wan2.2_i2v_api.json` | Comfy-Org / Wan2.2 I2V default graph | Pointed the diffusion model at **4-step Lightspeed** weights; LoRA stack nodes (109/110); `UnloadAllModels` VRAM-cleanup |\n| `hunyuan3d-v2.1_i2m_api.json` | Tencent Hunyuan3D-2.1 ComfyUI image→mesh reference graph (user-supplied, [tencent/Hunyuan3D-2.1](https://huggingface.co/tencent/Hunyuan3D-2.1)) | Topology kept 1:1 (including the user's VRAM-unload sink nodes 19/20/21/22); placeholder `image` set to `character.png` (the script always rewrites it to the uploaded filename); default `filename_prefix` kept (`mesh/hy21`) — the script stamps it per run to `mesh/i2m_<timestamp>`; no model/parameter changes — resolution 4096 / 30 steps / CFG 5.0 are the graph's own defaults and stay the skill defaults\n\nAll six are **adapted** (as opposed to) the original; they add prompt routing,\nresolution selection, dual-mode switches, and VRAM management on top of the\noriginal ComfyUI graphs."},{"path":"references/cli-reference.md","content":"# CLI Usage Reference (image-with-comfyui)\n\nFull command examples for every subcommand. The SKILL.md procedure only lists the\ncore rule; open this file for the exact flags.\n\n> **⚠️ The prompt is MANDATORY and must be passed via `--prompt` — never as a positional argument.**\n>\n> Every subcommand that takes a prompt (`t2i`, `i2i`, `wan2.2`) reads it from `--prompt \"...\"` (`i2m` has no prompt — see below). Passing the prompt as a bare positional word fails immediately with exit code 2:\n> `python3 image_with_comfyui.py t2i \"some description\"` → `t2i: error: the following arguments are required: --prompt`\n>\n> This is a **100%-avoidable failure loop** — the model can't help you because the command never reaches ComfyUI. Always write the prompt through the `--prompt` flag on the first attempt.\n\n### T2I (Text → Image)\n\n```bash\n# Qwen-Image 2.1 (default model; CFG locked to 1.0 — no --cfg)\n# --aspect is OPTIONAL: omit it and the output is 1:1 (1024x1024)\npython3 image_with_comfyui.py t2i \\\n  --prompt \"Your detailed image description\" \\\n  --steps 25\n\n# With a negative prompt (Qwen-Image 2.1 supports it)\npython3 image_with_comfyui.py t2i \\\n  --prompt \"A detailed image description\" \\\n  --negative \"text, watermark, blurry\" \\\n  --steps 25\n\n# Request a specific aspect ratio explicitly (16:9, 9:16, 4:3, ...)\npython3 image_with_comfyui.py t2i \\\n  --prompt \"Your detailed image description\" \\\n  --aspect 16:9 \\\n  --steps 25\n\n\n# Z-Image\npython3 image_with_comfyui.py t2i \\\n  --model z-image \\\n  --prompt \"Your detailed image description\" \\\n  --aspect 1:1 \\\n  --steps 9\n\n# SD3.5 Medium (defaults to 1:1)\npython3 image_with_comfyui.py t2i \\\n  --model sd35 \\\n  --prompt \"A beautiful sunset over mountains\" \\\n  --negative \"text, watermark, blurry\" \\\n  --steps 20 \\\n  --cfg 5.5\n\n# Run ONLY the requested model (skip the runtime fallback chain)\npython3 image_with_comfyui.py t2i \\\n  --prompt \"...\" --model z-image --no-fallback\n\n# JPG output (converted client-side; default from config.json image.output_format)\npython3 image_with_comfyui.py t2i \\\n  --prompt \"...\" --format jpg --jpg-quality 90\n```\n\n### I2I (Edit Image)\n\n```bash\n# Qwen-Image 2.1 (default I2I)\npython3 image_with_comfyui.py i2i \\\n  --prompt \"Change background to a beach\" \\\n  --image /path/to/source.jpg \\\n  --steps 25\n\n# Explicit aspect ratio: routes the I2I latent canvas through the\n# ResolutionSelector (workflow latent switch OFF) instead of matching the\n# first reference image's size, so a portrait photo CAN become 16:9 landscape.\n# The reference is uploaded at native size, never stretched. Without --aspect,\n# the output follows the first reference image's own size (no 1:1 forcing).\npython3 image_with_comfyui.py i2i \\\n  --prompt \"Extend the scene into a wide landscape\" \\\n  --image /path/to/portrait.jpg \\\n  --aspect 16:9\n\n# Multi-image reference (up to 16): --image order = slot order\npython3 image_with_comfyui.py i2i \\\n  --prompt \"put the shirt from <image2> on the person in <image1>\" \\\n  --image /path/to/base.jpg /path"}],"languages":[],"docsSourceLabel":"CLAWHUB","editorialOverview":null,"editorialQuality":{"score":100,"threshold":65,"status":"thin","wordCount":2381,"uniquenessScore":37,"reasons":["uniqueness-below-45"]}},"media":{"evidence":{"source":"no-media","verified":false,"confidence":"low","updatedAt":"2026-10-10T00:30:13.191Z","emptyReason":"No screenshots, media assets, or demo links are available."},"primaryImageUrl":null,"mediaAssetCount":0,"assets":[],"demoUrl":null},"ownerResources":{"evidence":{"source":"unclaimed","verified":false,"confidence":"low","updatedAt":"2026-10-10T00:30:13.191Z","emptyReason":"This page has not been claimed by the agent owner."},"hasCustomPage":false,"customPageUpdatedAt":null,"customLinks":[],"structuredLinks":{"docsUrl":null,"demoUrl":null,"supportUrl":null,"pricingUrl":null,"statusUrl":null},"customPage":null},"relatedAgents":{"evidence":{"source":"protocol-neighbors","verified":false,"confidence":"medium","updatedAt":"2026-10-10T05:40:20.918Z","emptyReason":null},"items":[{"id":"8ebccd8e-3863-4187-8355-c3f14e1f9edf","entityType":"agent","canonicalPath":"/agent/iofficeai-aionui","slug":"iofficeai-aionui","name":"AionUi","description":"Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!","url":"https://github.com/iOfficeAI/AionUi","homepage":"https://www.aionui.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-10-09T19:11:12.944Z","createdAt":"2026-02-25T03:38:16.584Z","downloads":null},{"id":"b917f68a-ebff-438e-84f8-3f4b2494c0bc","entityType":"agent","canonicalPath":"/agent/activepieces-activepieces","slug":"activepieces-activepieces","name":"activepieces","description":"AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents","url":"https://github.com/activepieces/activepieces","homepage":"https://www.activepieces.com","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-15T02:22:12.426Z","createdAt":"2026-02-25T03:38:12.412Z","downloads":null},{"id":"5cb26759-3a39-483f-94cf-276a98c13bb8","entityType":"agent","canonicalPath":"/agent/cherryhq-cherry-studio","slug":"cherryhq-cherry-studio","name":"cherry-studio","description":"AI productivity studio with smart chat, autonomous agents, and 300+ assistants. Unified access to frontier LLMs","url":"https://github.com/CherryHQ/cherry-studio","homepage":"https://cherry-ai.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-11T14:38:40.986Z","createdAt":"2026-02-25T03:38:19.379Z","downloads":null},{"id":"6f6582d0-5d76-4f0f-b81d-86520247950b","entityType":"agent","canonicalPath":"/agent/copilotkit-copilotkit","slug":"copilotkit-copilotkit","name":"CopilotKit","description":"The Frontend for Agents & Generative UI. React + Angular","url":"https://github.com/CopilotKit/CopilotKit","homepage":"https://docs.copilotkit.ai","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-03-25T09:50:57.846Z","createdAt":"2026-02-25T03:39:14.617Z","downloads":null}],"links":{"hub":"/agent","source":"/agent/source/clawhub","protocols":[{"label":"OpenClaw","href":"/agent/protocol/openclew"}]}}}