{"id":"72f95e9c-552a-4f7b-8583-733d2324aca1","entityType":"agent","slug":"clawhub-jimliu-baoyu-url-to-markdown","name":"Baoyu Url To Markdown","canonicalUrl":"https://www.xpersona.co/agent/clawhub-jimliu-baoyu-url-to-markdown","canonicalPath":"/agent/clawhub-jimliu-baoyu-url-to-markdown","generatedAt":"2026-10-10T03:54:22.805Z","source":"CLAWHUB","claimStatus":"UNCLAIMED","verificationTier":"NONE","summary":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-09T05:42:36.985Z","emptyReason":null},"description":"Fetch any URL and convert to markdown using baoyu-fetch CLI (Chrome CDP with site-specific adapters). Built-in adapters for X/Twitter, YouTube transcripts, H...","descriptionLabel":"Source description","evidenceSummary":"Capability contract not published. No trust telemetry is available yet. 4.3K downloads reported by the source. Last updated 10/9/2026.","installCommand":"clawhub skill install s17dfbhg0khk4fvtnx0stqg0yx83j25z:baoyu-url-to-markdown","sourceUrl":"https://clawhub.ai/jimliu/baoyu-url-to-markdown","homepage":"https://clawhub.ai/jimliu/skills/baoyu-url-to-markdown","primaryLinks":[{"label":"View on ClawHub","url":"https://clawhub.ai/jimliu/baoyu-url-to-markdown","kind":"source"},{"label":"Homepage","url":"https://clawhub.ai/jimliu/skills/baoyu-url-to-markdown","kind":"homepage"}],"safetyScore":84,"overallRank":62,"popularityScore":62,"trustScore":null,"claimedByName":null,"isOwner":false,"seoDescription":"Baoyu Url To Markdown technical dossier on Xpersona with agent coverage, OPENCLEW support, and live trust metadata."},"coverage":{"evidence":{"source":"public-profile","verified":false,"confidence":"medium","updatedAt":"2026-10-09T05:42:36.985Z","emptyReason":null},"protocols":[{"protocol":"OPENCLEW","label":"OpenClaw","status":"self-declared","notes":"Declared in the public agent profile."}],"capabilities":[],"verifiedCount":0,"selfDeclaredCount":1,"capabilityMatrix":{"rows":[{"key":"OPENCLEW","type":"protocol","support":"unknown","confidenceSource":"profile","notes":"Listed on profile"}],"flattenedTokens":"protocol:OPENCLEW|unknown|profile"}},"adoption":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-09T05:42:36.985Z","emptyReason":null},"stars":null,"forks":null,"downloads":4268,"packageName":null,"latestVersion":"1.117.2","tractionLabel":"4.3K downloads"},"release":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-09T05:42:36.984Z","emptyReason":null},"lastUpdatedAt":"2026-10-09T05:42:36.985Z","lastCrawledAt":"2026-10-09T05:42:36.984Z","lastIndexedAt":null,"nextCrawlAt":"2026-10-10T05:42:36.985Z","lastVerifiedAt":null,"highlights":[{"version":"1.117.2","createdAt":"2026-05-18T02:16:39.623Z","changelog":"## 1.117.2 - 2026-05-17 ### Documentation - `baoyu-cover-image`: ban programmatic text repair on generated bitmaps — disallow ImageMagick / Pillow / Canvas / SVG / HTML overlays to cover, rewrite, or replace title/subtitle text; regenerate from a corrected prompt or switch to a lower-text or no-title variant instead - `baoyu-article-illustrator`, `baoyu-comic`, `baoyu-image-cards`, `baoyu-xhs-images`, `baoyu-infographic`, `baoyu-slide-deck`: sync the same text-repair ban with skill-specific text categories (labels/captions, dialogue/sound effects, titles/body/tags, headings/data values, slide titles/bullets)","fileCount":50,"zipByteSize":96542},{"version":"1.115.4","createdAt":"2026-05-11T23:50:17.162Z","changelog":"### Documentation - Image generation backend selection: emphasize Codex `imagegen` as the priority runtime-native tool (invoke via the `Skill` tool with `skill: \"imagegen\"`) and forbid SVG/HTML/canvas substitution when no raster backend can be resolved — fall through to asking the user instead of silently emitting code-based art. Updated in `docs/image-generation-tools.md` and inlined into `baoyu-article-illustrator`, `baoyu-comic`, `baoyu-cover-image`, `baoyu-image-cards`, `baoyu-infographic`, `baoyu-slide-deck`, and `baoyu-xhs-images`.","fileCount":49,"zipByteSize":94970},{"version":"1.109.0","createdAt":"2026-04-21T18:40:22.242Z","changelog":"## 1.109.0 - 2026-04-21 ### Features - `baoyu-url-to-markdown`: vendor the `baoyu-fetch` runtime into `scripts/lib` and run it through a local `scripts/baoyu-fetch` CLI so published skill installs are self-contained ### Fixes - `baoyu-fetch`: extract playable X/Twitter video MP4 variants for single posts and X Articles, choosing the highest-bitrate MP4 and rendering article videos as `[video](...)` - `sync-clawhub`: publish from the shared release file list so extensionless CLI entrypoints, `bun.lock`, and vendored `scripts/lib` files are uploaded ### Maintenance - Upgrade `defuddle` to 0.17.0 and `jsdom` to 29.0.2; override `@xmldom/xmldom` to 0.8.13 to keep the Defuddle dependency chain vulnerability-free","fileCount":49,"zipByteSize":94970},{"version":"1.103.1","createdAt":"2026-04-13T16:18:15.080Z","changelog":"## 1.103.1 - 2026-04-13 ### Fixes - `baoyu-markdown-to-html`: decode HTML entities and strip tags from article summary - `baoyu-post-to-weibo`: decode HTML entities and strip tags from article summary","fileCount":49,"zipByteSize":110029},{"version":"1.103.0","createdAt":"2026-04-13T01:22:59.026Z","changelog":"## 1.103.0 - 2026-04-12 ### Features - baoyu-diagram: add multi-diagram mode for article-wide diagram generation ### Fixes - baoyu-article-illustrator: prevent color names and hex codes from appearing as visible text in generated images - baoyu-cover-image: prevent color names and hex codes from appearing as visible text in generated images - baoyu-image-cards: prevent color names from appearing as visible text in generated images - baoyu-post-to-wechat: decode HTML entities and strip tags from article summary","fileCount":49,"zipByteSize":110029},{"version":"1.82.2","createdAt":"2026-04-09T16:30:06.118Z","changelog":"**Major update: Migrated to a new vendored `baoyu-fetch` CLI with site-specific adapters, streamlining codebase and usage.** - CLI rewritten to use `baoyu-fetch` (now under `scripts/vendor/baoyu-fetch/`), supporting adapters for X/Twitter, YouTube, Hacker News, and generic content. - Legacy scripts and internal parser/converter infrastructure removed and replaced with the new adapter-based system. - Usage and agent instructions updated for the new CLI entrypoint; old `scripts/main.ts` and friends are gone. - EXTEND.md preference logic and setup retained with minor documentation clarifications. - Features such as login/CAPTCHA handling, media download, profile persistence, and debug output now standardized and improved via the new CLI.","fileCount":48,"zipByteSize":86106},{"version":"1.82.1","createdAt":"2026-03-25T21:34:21.172Z","changelog":"- Removed bun.lock file from scripts directory. - No changes to core functionality or user-facing features. - Maintenance update to streamline repository and dependency management.","fileCount":29,"zipByteSize":56003},{"version":"1.82.0","createdAt":"2026-03-25T03:41:05.635Z","changelog":"## 1.82.0 - 2026-03-24 ### Features - `baoyu-url-to-markdown`: add browser fallback strategy — headless first, automatic retry in visible Chrome on technical failure; new `--browser auto|headless|headed` flag with `--headless`/`--headed` shortcuts - `baoyu-url-to-markdown`: add content cleaner module for HTML preprocessing before extraction (remove ads, base64 images, scripts, styles) - `baoyu-url-to-markdown`: support base64 data URI images in media localizer alongside remote URLs - `baoyu-url-to-markdown`: capture final URL from browser to track redirects for output path generation - `baoyu-url-to-markdown`: add agent quality gate documentation for post-capture content validation ### Dependencies - `baoyu-url-to-markdown`: upgrade defuddle ^0.12.0 → ^0.14.0 ### Tests - `baoyu-url-to-markdown`: add unit tests for content-cleaner, html-to-markdown, legacy-converter, media-localizer","fileCount":30,"zipByteSize":64789}]},"execution":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No published capability contract is available yet."},"installCommand":"clawhub skill install s17dfbhg0khk4fvtnx0stqg0yx83j25z:baoyu-url-to-markdown","setupComplexity":"low","setupSteps":["Install using `clawhub skill install s17dfbhg0khk4fvtnx0stqg0yx83j25z:baoyu-url-to-markdown` in an isolated environment before connecting it to live workloads.","No published capability contract is available yet, so validate auth and request/response behavior manually.","Review the upstream CLAWHUB listing at https://clawhub.ai/jimliu/baoyu-url-to-markdown before using production credentials."],"contract":{"contractStatus":"missing","authModes":[],"requires":[],"forbidden":[],"supportsMcp":false,"supportsA2a":false,"supportsStreaming":false,"inputSchemaRef":null,"outputSchemaRef":null,"dataRegion":null,"contractUpdatedAt":null,"sourceUpdatedAt":null,"freshnessSeconds":null},"invocationGuide":{"preferredApi":{"snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-jimliu-baoyu-url-to-markdown/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-jimliu-baoyu-url-to-markdown/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-jimliu-baoyu-url-to-markdown/trust"},"curlExamples":["curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-jimliu-baoyu-url-to-markdown/snapshot\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-jimliu-baoyu-url-to-markdown/contract\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-jimliu-baoyu-url-to-markdown/trust\""],"jsonRequestTemplate":{"query":"summarize this repo","constraints":{"maxLatencyMs":2000,"protocolPreference":["OPENCLEW"]}},"jsonResponseTemplate":{"ok":true,"result":{"summary":"...","confidence":0.9},"meta":{"source":"CLAWHUB","generatedAt":"2026-10-10T03:54:22.802Z"}},"retryPolicy":{"maxAttempts":3,"backoffMs":[500,1500,3500],"retryableConditions":["HTTP_429","HTTP_503","NETWORK_TIMEOUT"]}},"endpoints":{"dossierUrl":"https://www.xpersona.co/api/v1/agents/clawhub-jimliu-baoyu-url-to-markdown/dossier","snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-jimliu-baoyu-url-to-markdown/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-jimliu-baoyu-url-to-markdown/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-jimliu-baoyu-url-to-markdown/trust"}},"reliability":{"evidence":{"source":"runtime-metrics","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No trust, reliability, or runtime telemetry is available."},"trust":{"status":"unavailable","handshakeStatus":"UNKNOWN","verificationFreshnessHours":null,"reputationScore":null,"p95LatencyMs":null,"successRate30d":null,"fallbackRate":null,"attempts30d":null,"trustUpdatedAt":null,"trustConfidence":"unknown","sourceUpdatedAt":null,"freshnessSeconds":null},"decisionGuardrails":{"doNotUseIf":["Contract metadata is missing or unavailable for deterministic execution."],"safeUseWhen":[],"riskFlags":["missing_or_unavailable_contract","trust_data_unavailable","schema_references_missing"],"operationalConfidence":"low"},"executionMetrics":{"observedLatencyMsP50":null,"observedLatencyMsP95":null,"estimatedCostUsd":null,"uptime30d":null,"rateLimitRpm":null,"rateLimitBurst":null,"lastVerifiedAt":null,"verificationSource":null},"runtimeMetrics":{"successRate":null,"avgLatencyMs":null,"avgCostUsd":null,"hallucinationRate":null,"retryRate":null,"disputeRate":null,"p50Latency":null,"p95Latency":null,"lastUpdated":null}},"benchmarks":{"evidence":{"source":"no-benchmark-data","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No benchmark suites or observed failure patterns are available."},"suites":[],"failurePatterns":[]},"artifacts":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-09T05:42:36.985Z","emptyReason":null},"readme":"Skill: Baoyu Url To Markdown\n\nOwner: jimliu\n\nSummary: Fetch any URL and convert to markdown using baoyu-fetch CLI (Chrome CDP with site-specific adapters). Built-in adapters for X/Twitter, YouTube transcripts, H...\n\nTags: latest:1.117.2\n\nVersion history:\n\nv1.117.2 | 2026-05-18T02:16:39.623Z | user\n\n## 1.117.2 - 2026-05-17\n\n### Documentation\n- `baoyu-cover-image`: ban programmatic text repair on generated bitmaps — disallow ImageMagick / Pillow / Canvas / SVG / HTML overlays to cover, rewrite, or replace title/subtitle text; regenerate from a corrected prompt or switch to a lower-text or no-title variant instead\n- `baoyu-article-illustrator`, `baoyu-comic`, `baoyu-image-cards`, `baoyu-xhs-images`, `baoyu-infographic`, `baoyu-slide-deck`: sync the same text-repair ban with skill-specific text categories (labels/captions, dialogue/sound effects, titles/body/tags, headings/data values, slide titles/bullets)\n\nv1.115.4 | 2026-05-11T23:50:17.162Z | user\n\n### Documentation\n- Image generation backend selection: emphasize Codex `imagegen` as the priority runtime-native tool (invoke via the `Skill` tool with `skill: \"imagegen\"`) and forbid SVG/HTML/canvas substitution when no raster backend can be resolved — fall through to asking the user instead of silently emitting code-based art. Updated in `docs/image-generation-tools.md` and inlined into `baoyu-article-illustrator`, `baoyu-comic`, `baoyu-cover-image`, `baoyu-image-cards`, `baoyu-infographic`, `baoyu-slide-deck`, and `baoyu-xhs-images`.\n\nv1.109.0 | 2026-04-21T18:40:22.242Z | user\n\n## 1.109.0 - 2026-04-21\n\n### Features\n- `baoyu-url-to-markdown`: vendor the `baoyu-fetch` runtime into `scripts/lib` and run it through a local `scripts/baoyu-fetch` CLI so published skill installs are self-contained\n\n### Fixes\n- `baoyu-fetch`: extract playable X/Twitter video MP4 variants for single posts and X Articles, choosing the highest-bitrate MP4 and rendering article videos as `[video](...)`\n- `sync-clawhub`: publish from the shared release file list so extensionless CLI entrypoints, `bun.lock`, and vendored `scripts/lib` files are uploaded\n\n### Maintenance\n- Upgrade `defuddle` to 0.17.0 and `jsdom` to 29.0.2; override `@xmldom/xmldom` to 0.8.13 to keep the Defuddle dependency chain vulnerability-free\n\nv1.103.1 | 2026-04-13T16:18:15.080Z | user\n\n## 1.103.1 - 2026-04-13\n\n### Fixes\n- `baoyu-markdown-to-html`: decode HTML entities and strip tags from article summary\n- `baoyu-post-to-weibo`: decode HTML entities and strip tags from article summary\n\nv1.103.0 | 2026-04-13T01:22:59.026Z | user\n\n## 1.103.0 - 2026-04-12\n\n### Features\n- baoyu-diagram: add multi-diagram mode for article-wide diagram generation\n\n### Fixes\n- baoyu-article-illustrator: prevent color names and hex codes from appearing as visible text in generated images\n- baoyu-cover-image: prevent color names and hex codes from appearing as visible text in generated images\n- baoyu-image-cards: prevent color names from appearing as visible text in generated images\n- baoyu-post-to-wechat: decode HTML entities and strip tags from article summary\n\nv1.82.2 | 2026-04-09T16:30:06.118Z | auto\n\n**Major update: Migrated to a new vendored `baoyu-fetch` CLI with site-specific adapters, streamlining codebase and usage.**\n\n- CLI rewritten to use `baoyu-fetch` (now under `scripts/vendor/baoyu-fetch/`), supporting adapters for X/Twitter, YouTube, Hacker News, and generic content.\n- Legacy scripts and internal parser/converter infrastructure removed and replaced with the new adapter-based system.\n- Usage and agent instructions updated for the new CLI entrypoint; old `scripts/main.ts` and friends are gone.\n- EXTEND.md preference logic and setup retained with minor documentation clarifications.\n- Features such as login/CAPTCHA handling, media download, profile persistence, and debug output now standardized and improved via the new CLI.\n\nv1.82.1 | 2026-03-25T21:34:21.172Z | auto\n\n- Removed bun.lock file from scripts directory.\n- No changes to core functionality or user-facing features.\n- Maintenance update to streamline repository and dependency management.\n\nv1.82.0 | 2026-03-25T03:41:05.635Z | user\n\n## 1.82.0 - 2026-03-24\n\n### Features\n- `baoyu-url-to-markdown`: add browser fallback strategy — headless first, automatic retry in visible Chrome on technical failure; new `--browser auto|headless|headed` flag with `--headless`/`--headed` shortcuts\n- `baoyu-url-to-markdown`: add content cleaner module for HTML preprocessing before extraction (remove ads, base64 images, scripts, styles)\n- `baoyu-url-to-markdown`: support base64 data URI images in media localizer alongside remote URLs\n- `baoyu-url-to-markdown`: capture final URL from browser to track redirects for output path generation\n- `baoyu-url-to-markdown`: add agent quality gate documentation for post-capture content validation\n\n### Dependencies\n- `baoyu-url-to-markdown`: upgrade defuddle ^0.12.0 → ^0.14.0\n\n### Tests\n- `baoyu-url-to-markdown`: add unit tests for content-cleaner, html-to-markdown, legacy-converter, media-localizer\n\nv1.76.1 | 2026-03-22T04:25:12.559Z | auto\n\n- Internal code update in vendor/baoyu-chrome-cdp (index.ts); user-visible behavior and documentation unchanged.\n- No changes to features, usage, or setup.\n\nv1.69.1 | 2026-03-16T18:02:09.141Z | user\n\n## 1.69.1 - 2026-03-16\n\n### Fixes\n- `baoyu-chrome-cdp`: tighten chrome auto-connect logic to reduce false positives\n\nv1.63.0 | 2026-03-13T05:25:02.608Z | user\n\n### Features\n- Add hosted `defuddle.md` API fallback when local browser capture fails\n- Extract YouTube transcript/caption text into markdown output\n- Materialize shadow DOM content for better web-component page conversion\n- Include language hint in markdown front matter when available\n\n### Refactor\n- Split monolithic converter into defuddle, legacy, and shared modules\n\nv1.60.0 | 2026-03-12T03:45:58.920Z | user\n\n## 1.60.0 - 2026-03-11\n\n### Features\n- `baoyu-url-to-markdown`: support reusing existing Chrome CDP instances and fix port detection order\n\n### Fixes\n- `baoyu-post-to-x`: add missing `fs` import in x-article\n\n### Refactor\n- Unify all CDP skills to use shared `baoyu-chrome-cdp` package with vendored copies\n- Simplify CLAUDE.md, move detailed documentation to `docs/` directory\n- Publish skills directly from synced vendor, removing separate artifact preparation step\n\nv1.0.0 | 2026-03-09T17:20:12.404Z | auto\n\n- First public release of baoyu-url-to-markdown: fetch any URL, render with Chrome CDP, and convert to markdown.\n- Saves both the clean markdown and a full HTML snapshot for each conversion.\n- Supports automatic fallback from Defuddle markdown conversion to legacy parser when needed.\n- Two modes available: capture on page load, or wait for user signal (for login-required pages).\n- First-time setup requires user input for media handling, output directory, and preference storage; setup enforces explicit user choice.\n- Media (images/videos) download and output options can be controlled via CLI arguments or saved preferences.\n\nArchive index:\n\nArchive v1.117.2: 50 files, 96542 bytes\n\nFiles: references/adapters.md (3231b), references/config/first-time-setup.md (2459b), references/quality-gate.md (2490b), scripts/baoyu-fetch (122b), scripts/bun.lock (35037b), scripts/lib/adapters/generic/index.ts (2451b), scripts/lib/adapters/hn/index.ts (10454b), scripts/lib/adapters/index.ts (834b), scripts/lib/adapters/types.ts (2100b), scripts/lib/adapters/x/article.ts (12677b), scripts/lib/adapters/x/index.ts (4048b), scripts/lib/adapters/x/login.ts (2258b), scripts/lib/adapters/x/match.ts (297b), scripts/lib/adapters/x/payloads.ts (1681b), scripts/lib/adapters/x/session.ts (1404b), scripts/lib/adapters/x/shared.ts (12980b), scripts/lib/adapters/x/single.ts (2583b), scripts/lib/adapters/x/thread-loader.ts (9078b), scripts/lib/adapters/x/thread.ts (9196b), scripts/lib/adapters/x/types.ts (600b), scripts/lib/adapters/youtube/index.ts (898b), scripts/lib/adapters/youtube/transcript.ts (13301b), scripts/lib/adapters/youtube/utils.ts (6523b), scripts/lib/browser/cdp-client.ts (6879b), scripts/lib/browser/chrome-launcher.ts (5883b), scripts/lib/browser/cookie-sidecar.ts (2941b), scripts/lib/browser/interaction-gates.ts (4095b), scripts/lib/browser/network-journal.ts (7153b), scripts/lib/browser/page-snapshot.ts (3429b), scripts/lib/browser/profile.ts (5670b), scripts/lib/browser/session.ts (4642b), scripts/lib/cli.ts (7066b), scripts/lib/commands/convert.ts (17456b), scripts/lib/extract/document.ts (846b), scripts/lib/extract/html-cleaner.ts (11016b), scripts/lib/extract/html-extractor.ts (2243b), scripts/lib/extract/html-to-markdown.ts (21344b), scripts/lib/extract/markdown-renderer.ts (4661b), scripts/lib/media/default-downloader.ts (5215b), scripts/lib/media/markdown-media.ts (13186b), scripts/lib/media/media-utils.ts (6222b), scripts/lib/media/types.ts (710b), scripts/lib/types/defuddle-node.d.ts (605b), scripts/lib/types/shims.d.ts (422b), scripts/lib/utils/logger.ts (645b), scripts/lib/utils/url.ts (300b), scripts/package.json (557b), skill-card.md (2870b), SKILL.md (8520b), _meta.json (142b)\n\nFile v1.117.2:SKILL.md\n\n---\nname: baoyu-url-to-markdown\ndescription: Fetch any URL and convert to markdown using baoyu-fetch CLI (Chrome CDP with site-specific adapters). Built-in adapters for X/Twitter, YouTube transcripts, Hacker News threads, and generic pages via Defuddle. Handles login/CAPTCHA via interaction wait modes. Use when user wants to save a webpage as markdown.\nversion: 1.61.0\nmetadata:\n  openclaw:\n    homepage: https://github.com/JimLiu/baoyu-skills#baoyu-url-to-markdown\n    requires:\n      anyBins:\n        - bun\n---\n\n# URL to Markdown\n\nFetches any URL via `baoyu-fetch` CLI (Chrome CDP + site-specific adapters) and converts it to clean markdown.\n\n## User Input Tools\n\nWhen this skill prompts the user, follow this tool-selection rule (priority order):\n\n1. **Prefer built-in user-input tools** exposed by the current agent runtime — e.g., `AskUserQuestion`, `request_user_input`, `clarify`, `ask_user`, or any equivalent.\n2. **Fallback**: if no such tool exists, emit a numbered plain-text message and ask the user to reply with the chosen number/answer for each question.\n3. **Batching**: if the tool supports multiple questions per call, combine all applicable questions into a single call; if only single-question, ask them one at a time in priority order.\n\nConcrete `AskUserQuestion` references below are examples — substitute the local equivalent in other runtimes.\n\n## CLI Setup\n\n**Important**: The CLI source is vendored in `{baseDir}/scripts/lib`. `scripts/package.json` installs only third-party runtime dependencies.\n\n**Agent Execution Instructions**:\n1. Determine this SKILL.md file's directory path as `{baseDir}`\n2. Resolve `${BUN}` runtime: if `bun` installed → `bun`; else suggest installing Bun\n3. If `{baseDir}/scripts/node_modules` does not exist, run `${BUN} install --cwd {baseDir}/scripts`\n4. `${READER}` = `{baseDir}/scripts/baoyu-fetch`\n5. Replace all `${READER}` in this document with the resolved value\n\n## Preferences (EXTEND.md)\n\nCheck EXTEND.md in priority order — the first one found wins:\n\n| Priority | Path | Scope |\n|----------|------|-------|\n| 1 | `.baoyu-skills/baoyu-url-to-markdown/EXTEND.md` | Project |\n| 2 | `${XDG_CONFIG_HOME:-$HOME/.config}/baoyu-skills/baoyu-url-to-markdown/EXTEND.md` | XDG |\n| 3 | `$HOME/.baoyu-skills/baoyu-url-to-markdown/EXTEND.md` | User home |\n\n| Result | Action |\n|--------|--------|\n| Found | Read, parse, apply settings |\n| Not found | **MUST** run first-time setup (see below) — do NOT silently create defaults |\n\n**EXTEND.md supports**: download media by default, default output directory.\n\n### First-Time Setup ⛔ BLOCKING\n\nWhen EXTEND.md is not found, you **MUST** use `AskUserQuestion` to gather preferences before creating EXTEND.md. **NEVER** create EXTEND.md with silent defaults. Generation is BLOCKED until setup completes. Batch all three questions into a single call:\n\n- **Q1 — Media** (header \"Media\"): \"How to handle images and videos in pages?\"\n  - \"Ask each time (Recommended)\" — Prompt after each save\n  - \"Always download\" — Download to local `imgs/` and `videos/`\n  - \"Never download\" — Keep remote URLs\n- **Q2 — Output** (header \"Output\"): \"Default output directory?\"\n  - \"url-to-markdown (Recommended)\" — Save to `./url-to-markdown/{domain}/{slug}.md`\n  - User may pick \"Other\" and type a custom path\n- **Q3 — Save** (header \"Save\"): \"Where to save preferences?\"\n  - \"User (Recommended)\" — `~/.baoyu-skills/` (all projects)\n  - \"Project\" — `.baoyu-skills/` (this project only)\n\nAfter answers, write EXTEND.md, confirm \"Preferences saved to [path]\", then continue.\n\nFull template: [references/config/first-time-setup.md](references/config/first-time-setup.md).\n\n### Supported Keys\n\n| Key | Default | Values | Description |\n|-----|---------|--------|-------------|\n| `download_media` | `ask` | `ask` / `1` / `0` | `ask` = prompt each time, `1` = always, `0` = never |\n| `default_output_dir` | empty | path or empty | Default output directory (empty = `./url-to-markdown/`) |\n\n**EXTEND.md → CLI mapping**:\n\n| EXTEND.md key | CLI argument | Notes |\n|---------------|-------------|-------|\n| `download_media: 1` | `--download-media` | Requires `--output` to be set |\n| `default_output_dir: ./posts/` | Agent constructs `--output ./posts/{domain}/{slug}.md` | Agent generates path, not a direct flag |\n\n**Value priority**: CLI arguments → EXTEND.md → skill defaults.\n\n## Usage\n\n```bash\n# Default: headless capture, markdown to stdout\n${READER} <url>\n\n# Save to file\n${READER} <url> --output article.md\n\n# Save with media download\n${READER} <url> --output article.md --download-media\n\n# Wait for interaction (login/CAPTCHA) — auto-detect and continue\n${READER} <url> --wait-for interaction --output article.md\n\n# Wait for interaction — manual control (Enter to continue)\n${READER} <url> --wait-for force --output article.md\n\n# JSON output\n${READER} <url> --format json --output article.json\n\n# Force specific adapter\n${READER} <url> --adapter youtube --output transcript.md\n```\n\n## Options\n\n| Option | Description |\n|--------|-------------|\n| `<url>` | URL to fetch |\n| `--output <path>` | Output file path (default: stdout) |\n| `--format <type>` | Output format: `markdown` (default) or `json` |\n| `--json` | Shorthand for `--format json` |\n| `--adapter <name>` | Force adapter: `x`, `youtube`, `hn`, or `generic` (default: auto-detect) |\n| `--headless` | Force headless Chrome (no visible window) |\n| `--wait-for <mode>` | Interaction wait mode: `none` (default), `interaction`, or `force` |\n| `--wait-for-interaction` | Alias for `--wait-for interaction` |\n| `--wait-for-login` | Alias for `--wait-for interaction` |\n| `--timeout <ms>` | Page load timeout (default: 30000) |\n| `--interaction-timeout <ms>` | Login/CAPTCHA wait timeout (default: 600000 = 10 min) |\n| `--interaction-poll-interval <ms>` | Poll interval for interaction checks (default: 1500) |\n| `--download-media` | Download images/videos to local `imgs/` and `videos/`, rewrite markdown links. Requires `--output` |\n| `--media-dir <dir>` | Base directory for downloaded media (default: same as `--output` directory) |\n| `--cdp-url <url>` | Reuse existing Chrome DevTools Protocol endpoint |\n| `--browser-path <path>` | Custom Chrome/Chromium binary path |\n| `--chrome-profile-dir <path>` | Chrome user data directory (default: `BAOYU_CHROME_PROFILE_DIR` env or `./baoyu-skills/chrome-profile`) |\n| `--debug-dir <dir>` | Write debug artifacts (document.json, markdown.md, page.html, network.json) |\n\n## Agent Quality Gate\n\n**CRITICAL**: treat default headless capture as provisional. Some sites render differently in headless mode and can silently return low-quality content without failing the CLI.\n\nAfter every headless run, inspect the saved markdown. See [references/quality-gate.md](references/quality-gate.md) for the full checklist, recovery workflow, and capture-mode table. Read it whenever a run looks suspicious or the user asks about login/CAPTCHA handling.\n\n## Output Path Generation\n\nThe agent must construct the output file path — `baoyu-fetch` does not auto-generate paths.\n\n**Algorithm**:\n1. Determine base directory from EXTEND.md `default_output_dir` or default `./url-to-markdown/`\n2. Extract domain from URL (e.g., `example.com`)\n3. Generate slug from URL path or page title (kebab-case, 2-6 words)\n4. Construct: `{base_dir}/{domain}/{slug}/{slug}.md` — each URL gets its own directory so media files stay isolated\n5. Conflict resolution: append timestamp `{slug}-YYYYMMDD-HHMMSS/{slug}-YYYYMMDD-HHMMSS.md`\n\nPass the constructed path to `--output`. Media files (`--download-media`) are saved into subdirectories next to the markdown file, keeping each URL's assets self-contained.\n\n## Adapters & Media\n\nSee [references/adapters.md](references/adapters.md) for the adapter catalog (X, YouTube, Hacker News, generic), per-adapter notes, the media download flow (`ask` / always / never), and the JSON output schema. Read it before answering adapter-specific questions or handling media prompts.\n\n## Environment Variables\n\n| Variable | Description |\n|----------|-------------|\n| `BAOYU_CHROME_PROFILE_DIR` | Chrome user data directory (can also use `--chrome-profile-dir`) |\n\n**Troubleshooting**: Chrome not found → use `--browser-path`. Timeout → increase `--timeout`. Login/CAPTCHA → `--wait-for interaction`. Debug → `--debug-dir` to inspect captured HTML and network logs.\n\n## Extension Support\n\nCustom configurations via EXTEND.md. See **Preferences** section above for paths and supported keys.\n\nFile v1.117.2:_meta.json\n\n{\n  \"ownerId\": \"kn7csrrndw79hpke5d0gsnx93d82k67r\",\n  \"slug\": \"baoyu-url-to-markdown\",\n  \"version\": \"1.117.2\",\n  \"publishedAt\": 1779070599623\n}\n\nFile v1.117.2:references/adapters.md\n\n# Adapters & Media\n\nRead when choosing an adapter, handling media, or answering adapter-specific questions.\n\n## Built-in Adapters\n\n| Adapter | URLs | Key Features |\n|---------|------|-------------|\n| `x` | x.com, twitter.com | Tweets, threads, X Articles, media, login detection |\n| `youtube` | youtube.com, youtu.be | Transcript/captions, chapters, cover image, metadata |\n| `hn` | news.ycombinator.com | Threaded comments, story metadata, nested replies |\n| `generic` | Any URL (fallback) | Defuddle extraction, Readability fallback, auto-scroll, network idle detection |\n\nAdapter is auto-selected based on URL. Override with `--adapter <name>`.\n\n### YouTube\n\n- Extracts transcripts/captions when available\n- Transcript format: `[MM:SS] Text segment` with chapter headings\n- Availability depends on YouTube exposing a caption track; videos with captions disabled or restricted playback may produce description-only output\n- Use `--wait-for force` if the page needs time to finish loading player metadata\n\n### X/Twitter\n\n- Extracts single tweets, threads, and X Articles\n- Auto-detects login state; if logged out and content requires auth, JSON output shows `\"status\": \"needs_interaction\"`\n- Use `--wait-for interaction` for login-protected content\n\n### Hacker News\n\n- Parses threaded comments with proper nesting and reply hierarchy\n- Includes story metadata (title, URL, author, score, comment count)\n- Shows comment deletion/dead status\n\n## Media Download Workflow\n\nDriven by `download_media` in EXTEND.md:\n\n| Setting | Behavior |\n|---------|----------|\n| `1` (always) | Run CLI with `--download-media --output <path>` |\n| `0` (never) | Run CLI with `--output <path>` (no media download) |\n| `ask` (default) | Follow the ask-each-time flow below |\n\n### Ask-Each-Time Flow\n\n1. Run the CLI **without** `--download-media` with `--output <path>` → markdown saved\n2. Check the saved markdown for remote media URLs (`https://` in image/video links)\n3. **If no remote media found** → done, no prompt needed\n4. **If remote media found** → ask via `AskUserQuestion`:\n   - header: \"Media\", question: \"Download N images/videos to local files?\"\n   - \"Yes\" — Download to local directories\n   - \"No\" — Keep remote URLs\n5. If the user confirms → run the CLI **again** with `--download-media --output <same-path>` (overwrites markdown with localized links)\n\n### Media Layout\n\nWhen `--download-media` is enabled:\n\n- Images → `imgs/` next to the output file (or `--media-dir`)\n- Videos → `videos/` next to the output file (or `--media-dir`)\n- Markdown media links are rewritten to local relative paths\n\n## Output Format\n\nMarkdown to stdout (or file with `--output`).\n\nJSON output (`--format json`) returns structured data:\n\n- `adapter` — which adapter handled the URL\n- `status` — `\"ok\"` or `\"needs_interaction\"`\n- `login` — login state detection (`logged_in`, `logged_out`, `unknown`)\n- `interaction` — interaction gate details (kind, provider, prompt)\n- `document` — structured content (url, title, author, publishedAt, content blocks, metadata)\n- `media` — collected media assets with url, kind, role\n- `markdown` — converted markdown text\n- `downloads` — media download results (when `--download-media` used)\n\nFile v1.117.2:references/config/first-time-setup.md\n\n---\nname: first-time-setup\ndescription: First-time setup flow for baoyu-url-to-markdown preferences\n---\n\n# First-Time Setup\n\n## Overview\n\nWhen no EXTEND.md is found, guide user through preference setup.\n\n**BLOCKING OPERATION**: This setup MUST complete before ANY other workflow steps. Do NOT:\n- Start converting URLs\n- Ask about URLs or output paths\n- Proceed to any conversion\n\nONLY ask the questions in this setup flow, save EXTEND.md, then continue.\n\n## Setup Flow\n\n```\nNo EXTEND.md found\n        |\n        v\n+---------------------+\n| AskUserQuestion     |\n| (all questions)     |\n+---------------------+\n        |\n        v\n+---------------------+\n| Create EXTEND.md    |\n+---------------------+\n        |\n        v\n    Continue conversion\n```\n\n## Questions\n\n**Language**: Use user's input language or saved language preference.\n\nUse AskUserQuestion with ALL questions in ONE call:\n\n### Question 1: Download Media\n\n```yaml\nheader: \"Media\"\nquestion: \"How to handle images and videos in pages?\"\noptions:\n  - label: \"Ask each time (Recommended)\"\n    description: \"After saving markdown, ask whether to download media\"\n  - label: \"Always download\"\n    description: \"Always download media to local imgs/ and videos/ directories\"\n  - label: \"Never download\"\n    description: \"Keep original remote URLs in markdown\"\n```\n\n### Question 2: Default Output Directory\n\n```yaml\nheader: \"Output\"\nquestion: \"Default output directory?\"\noptions:\n  - label: \"url-to-markdown (Recommended)\"\n    description: \"Save to ./url-to-markdown/{domain}/{slug}.md\"\n```\n\nNote: User will likely choose \"Other\" to type a custom path.\n\n### Question 3: Save Location\n\n```yaml\nheader: \"Save\"\nquestion: \"Where to save preferences?\"\noptions:\n  - label: \"User (Recommended)\"\n    description: \"~/.baoyu-skills/ (all projects)\"\n  - label: \"Project\"\n    description: \".baoyu-skills/ (this project only)\"\n```\n\n## Save Locations\n\n| Choice | Path | Scope |\n|--------|------|-------|\n| User | `~/.baoyu-skills/baoyu-url-to-markdown/EXTEND.md` | All projects |\n| Project | `.baoyu-skills/baoyu-url-to-markdown/EXTEND.md` | Current project |\n\n## After Setup\n\n1. Create directory if needed\n2. Write EXTEND.md\n3. Confirm: \"Preferences saved to [path]\"\n4. Continue with conversion using saved preferences\n\n## EXTEND.md Template\n\n```md\ndownload_media: [ask/1/0]\ndefault_output_dir: [path or empty]\n```\n\n## Modifying Preferences Later\n\nUsers can edit EXTEND.md directly or delete it to trigger setup again.\n\nFile v1.117.2:references/quality-gate.md\n\n# Quality Gate & Recovery\n\nHeadless Chrome can silently return low-quality content — layout shells, login walls, or framework payloads — without the CLI returning a non-zero exit code. Read this after every headless run so you can catch and recover from those cases.\n\n## Checks the Agent Must Run\n\n1. Confirm the markdown title matches the target page, not a generic site shell\n2. Confirm the body contains the expected article/page content, not just navigation, footer, or a generic error\n3. Watch for obvious failure signs:\n   - `Application error`\n   - `This page could not be found`\n   - Login, signup, subscribe, or verification shells\n   - Extremely short markdown for a page that should be long-form\n   - Raw framework payloads or mostly boilerplate content\n4. Do NOT accept a run as successful just because the CLI exited `0`\n\n**Tip**: run with `--format json` to get structured signals including `status`, `login.state`, and `interaction`. `\"status\": \"needs_interaction\"` means the page requires manual interaction.\n\n## Recovery Workflow\n\n1. Start headless (default) unless there is already a clear reason to use interaction mode\n2. Review markdown quality immediately after the run\n3. If the content is low quality or indicates login/CAPTCHA:\n   - `--wait-for interaction` for auto-detected gates (login, CAPTCHA, Cloudflare)\n   - `--wait-for force` when the page needs manual browsing, scroll loading, or complex interaction\n4. If `--wait-for` is used, tell the user exactly what to do:\n   - Login required → sign in in the browser\n   - CAPTCHA visible → solve it\n   - Slow loading → wait until content is visible\n   - `--wait-for force` → press Enter when ready\n5. If JSON output shows `\"status\": \"needs_interaction\"`, switch to `--wait-for interaction` automatically\n\n## Capture Modes\n\n| Mode | Behavior | Use When |\n|------|----------|----------|\n| Default | Headless Chrome, auto-extract on network idle | Public pages, static content |\n| `--headless` | Explicit headless (same as default) | Clarify intent |\n| `--wait-for interaction` | Opens visible Chrome, auto-detects login/CAPTCHA gates, waits for them to clear, then continues | Login-required, CAPTCHA-protected |\n| `--wait-for force` | Opens visible Chrome, auto-detects OR accepts Enter keypress to continue | Complex flows, lazy loading, paywalls |\n\n**Interaction gate auto-detection**: Cloudflare Turnstile / \"just a moment\" pages, Google reCAPTCHA, hCaptcha, custom challenge / verification screens.\n\nFile v1.117.2:scripts/package.json\n\n{\n  \"name\": \"baoyu-url-to-markdown-scripts\",\n  \"private\": true,\n  \"type\": \"module\",\n  \"bin\": {\n    \"baoyu-fetch\": \"./baoyu-fetch\"\n  },\n  \"scripts\": {\n    \"reader\": \"bun ./lib/cli.ts\"\n  },\n  \"dependencies\": {\n    \"@mozilla/readability\": \"^0.6.0\",\n    \"chrome-launcher\": \"^1.2.1\",\n    \"defuddle\": \"^0.17.0\",\n    \"jsdom\": \"^29.0.2\",\n    \"remark-gfm\": \"^4.0.1\",\n    \"remark-parse\": \"^11.0.0\",\n    \"turndown\": \"^7.2.0\",\n    \"turndown-plugin-gfm\": \"^1.0.2\",\n    \"unified\": \"^11.0.5\",\n    \"ws\": \"^8.18.3\"\n  },\n  \"overrides\": {\n    \"@xmldom/xmldom\": \"0.8.13\"\n  }\n}\n\nFile v1.117.2:skill-card.md\n\n## Description:\n\nFetch any URL and convert to markdown using baoyu-fetch CLI with Chrome CDP and site-specific adapters for X/Twitter, YouTube transcripts, Hacker News threads, and generic pages.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[jimliu](https://clawhub.ai/user/jimliu)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers and agents use this skill to fetch web pages, convert them into Markdown or JSON, and optionally save page media beside the generated document. It is intended for webpage capture workflows that may require site-specific handling or interactive login/CAPTCHA resolution.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: Authenticated browser sessions may expose private page content or account state during capture.\n\nMitigation: Use the skill only on URLs approved for browser automation, and avoid private, internal, invitation-only, password-reset, or token-bearing URLs unless the workflow has been reviewed.\n\nRisk: Generic Markdown extraction can send target URLs to defuddle.md as a remote fallback.\n\nMitigation: Review remote fallback behavior before use on confidential pages, and disable or modify the fallback path when sensitive URLs must remain local.\n\nRisk: X/Twitter session cookies can be persisted in a plaintext sidecar file.\n\nMitigation: Treat the Chrome profile directory and x-session-cookies.json as sensitive, restrict access to them, and delete saved session files when they are no longer needed.\n\nRisk: Debug output can save page HTML, generated Markdown, document metadata, and network contents.\n\nMitigation: Avoid --debug-dir on confidential pages, or store debug artifacts only in approved protected locations and remove them after review.\n\n## Reference(s):\n\n- [ClawHub skill page](https://clawhub.ai/jimliu/skills/baoyu-url-to-markdown)\n- [Project homepage](https://github.com/JimLiu/baoyu-skills#baoyu-url-to-markdown)\n- [Adapters & Media](references/adapters.md)\n- [Quality Gate & Recovery](references/quality-gate.md)\n- [First-Time Setup](references/config/first-time-setup.md)\n\n## Skill Output:\n\n**Output Type(s):** [Markdown, JSON, Files, Shell commands, Configuration]\n\n**Output Format:** [Markdown or JSON output, optionally written to files with downloaded media assets in adjacent directories]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [Requires Bun and a Chrome or Chromium browser; output quality should be checked after headless captures.]\n\n## Skill Version(s):\n\n1.117.2 (source: ClawHub release metadata, changelog dated 2026-05-17)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nArchive v1.115.4: 49 files, 94970 bytes\n\nFiles: references/adapters.md (3231b), references/config/first-time-setup.md (2459b), references/quality-gate.md (2490b), scripts/baoyu-fetch (122b), scripts/bun.lock (35037b), scripts/lib/adapters/generic/index.ts (2451b), scripts/lib/adapters/hn/index.ts (10454b), scripts/lib/adapters/index.ts (834b), scripts/lib/adapters/types.ts (2100b), scripts/lib/adapters/x/article.ts (12677b), scripts/lib/adapters/x/index.ts (4048b), scripts/lib/adapters/x/login.ts (2258b), scripts/lib/adapters/x/match.ts (297b), scripts/lib/adapters/x/payloads.ts (1681b), scripts/lib/adapters/x/session.ts (1404b), scripts/lib/adapters/x/shared.ts (12980b), scripts/lib/adapters/x/single.ts (2583b), scripts/lib/adapters/x/thread-loader.ts (9078b), scripts/lib/adapters/x/thread.ts (9196b), scripts/lib/adapters/x/types.ts (600b), scripts/lib/adapters/youtube/index.ts (898b), scripts/lib/adapters/youtube/transcript.ts (13301b), scripts/lib/adapters/youtube/utils.ts (6523b), scripts/lib/browser/cdp-client.ts (6879b), scripts/lib/browser/chrome-launcher.ts (5883b), scripts/lib/browser/cookie-sidecar.ts (2941b), scripts/lib/browser/interaction-gates.ts (4095b), scripts/lib/browser/network-journal.ts (7153b), scripts/lib/browser/page-snapshot.ts (3429b), scripts/lib/browser/profile.ts (5670b), scripts/lib/browser/session.ts (4642b), scripts/lib/cli.ts (7066b), scripts/lib/commands/convert.ts (17456b), scripts/lib/extract/document.ts (846b), scripts/lib/extract/html-cleaner.ts (11016b), scripts/lib/extract/html-extractor.ts (2243b), scripts/lib/extract/html-to-markdown.ts (21344b), scripts/lib/extract/markdown-renderer.ts (4661b), scripts/lib/media/default-downloader.ts (5215b), scripts/lib/media/markdown-media.ts (13186b), scripts/lib/media/media-utils.ts (6222b), scripts/lib/media/types.ts (710b), scripts/lib/types/defuddle-node.d.ts (605b), scripts/lib/types/shims.d.ts (422b), scripts/lib/utils/logger.ts (645b), scripts/lib/utils/url.ts (300b), scripts/package.json (557b), SKILL.md (8520b), _meta.json (142b)\n\nFile v1.115.4:SKILL.md\n\n---\nname: baoyu-url-to-markdown\ndescription: Fetch any URL and convert to markdown using baoyu-fetch CLI (Chrome CDP with site-specific adapters). Built-in adapters for X/Twitter, YouTube transcripts, Hacker News threads, and generic pages via Defuddle. Handles login/CAPTCHA via interaction wait modes. Use when user wants to save a webpage as markdown.\nversion: 1.61.0\nmetadata:\n  openclaw:\n    homepage: https://github.com/JimLiu/baoyu-skills#baoyu-url-to-markdown\n    requires:\n      anyBins:\n        - bun\n---\n\n# URL to Markdown\n\nFetches any URL via `baoyu-fetch` CLI (Chrome CDP + site-specific adapters) and converts it to clean markdown.\n\n## User Input Tools\n\nWhen this skill prompts the user, follow this tool-selection rule (priority order):\n\n1. **Prefer built-in user-input tools** exposed by the current agent runtime — e.g., `AskUserQuestion`, `request_user_input`, `clarify`, `ask_user`, or any equivalent.\n2. **Fallback**: if no such tool exists, emit a numbered plain-text message and ask the user to reply with the chosen number/answer for each question.\n3. **Batching**: if the tool supports multiple questions per call, combine all applicable questions into a single call; if only single-question, ask them one at a time in priority order.\n\nConcrete `AskUserQuestion` references below are examples — substitute the local equivalent in other runtimes.\n\n## CLI Setup\n\n**Important**: The CLI source is vendored in `{baseDir}/scripts/lib`. `scripts/package.json` installs only third-party runtime dependencies.\n\n**Agent Execution Instructions**:\n1. Determine this SKILL.md file's directory path as `{baseDir}`\n2. Resolve `${BUN}` runtime: if `bun` installed → `bun`; else suggest installing Bun\n3. If `{baseDir}/scripts/node_modules` does not exist, run `${BUN} install --cwd {baseDir}/scripts`\n4. `${READER}` = `{baseDir}/scripts/baoyu-fetch`\n5. Replace all `${READER}` in this document with the resolved value\n\n## Preferences (EXTEND.md)\n\nCheck EXTEND.md in priority order — the first one found wins:\n\n| Priority | Path | Scope |\n|----------|------|-------|\n| 1 | `.baoyu-skills/baoyu-url-to-markdown/EXTEND.md` | Project |\n| 2 | `${XDG_CONFIG_HOME:-$HOME/.config}/baoyu-skills/baoyu-url-to-markdown/EXTEND.md` | XDG |\n| 3 | `$HOME/.baoyu-skills/baoyu-url-to-markdown/EXTEND.md` | User home |\n\n| Result | Action |\n|--------|--------|\n| Found | Read, parse, apply settings |\n| Not found | **MUST** run first-time setup (see below) — do NOT silently create defaults |\n\n**EXTEND.md supports**: download media by default, default output directory.\n\n### First-Time Setup ⛔ BLOCKING\n\nWhen EXTEND.md is not found, you **MUST** use `AskUserQuestion` to gather preferences before creating EXTEND.md. **NEVER** create EXTEND.md with silent defaults. Generation is BLOCKED until setup completes. Batch all three questions into a single call:\n\n- **Q1 — Media** (header \"Media\"): \"How to handle images and videos in pages?\"\n  - \"Ask each time (Recommended)\" — Prompt after each save\n  - \"Always download\" — Download to local `imgs/` and `videos/`\n  - \"Never download\" — Keep remote URLs\n- **Q2 — Output** (header \"Output\"): \"Default output directory?\"\n  - \"url-to-markdown (Recommended)\" — Save to `./url-to-markdown/{domain}/{slug}.md`\n  - User may pick \"Other\" and type a custom path\n- **Q3 — Save** (header \"Save\"): \"Where to save preferences?\"\n  - \"User (Recommended)\" — `~/.baoyu-skills/` (all projects)\n  - \"Project\" — `.baoyu-skills/` (this project only)\n\nAfter answers, write EXTEND.md, confirm \"Preferences saved to [path]\", then continue.\n\nFull template: [references/config/first-time-setup.md](references/config/first-time-setup.md).\n\n### Supported Keys\n\n| Key | Default | Values | Description |\n|-----|---------|--------|-------------|\n| `download_media` | `ask` | `ask` / `1` / `0` | `ask` = prompt each time, `1` = always, `0` = never |\n| `default_output_dir` | empty | path or empty | Default output directory (empty = `./url-to-markdown/`) |\n\n**EXTEND.md → CLI mapping**:\n\n| EXTEND.md key | CLI argument | Notes |\n|---------------|-------------|-------|\n| `download_media: 1` | `--download-media` | Requires `--output` to be set |\n| `default_output_dir: ./posts/` | Agent constructs `--output ./posts/{domain}/{slug}.md` | Agent generates path, not a direct flag |\n\n**Value priority**: CLI arguments → EXTEND.md → skill defaults.\n\n## Usage\n\n```bash\n# Default: headless capture, markdown to stdout\n${READER} <url>\n\n# Save to file\n${READER} <url> --output article.md\n\n# Save with media download\n${READER} <url> --output article.md --download-media\n\n# Wait for interaction (login/CAPTCHA) — auto-detect and continue\n${READER} <url> --wait-for interaction --output article.md\n\n# Wait for interaction — manual control (Enter to continue)\n${READER} <url> --wait-for force --output article.md\n\n# JSON output\n${READER} <url> --format json --output article.json\n\n# Force specific adapter\n${READER} <url> --adapter youtube --output transcript.md\n```\n\n## Options\n\n| Option | Description |\n|--------|-------------|\n| `<url>` | URL to fetch |\n| `--output <path>` | Output file path (default: stdout) |\n| `--format <type>` | Output format: `markdown` (default) or `json` |\n| `--json` | Shorthand for `--format json` |\n| `--adapter <name>` | Force adapter: `x`, `youtube`, `hn`, or `generic` (default: auto-detect) |\n| `--headless` | Force headless Chrome (no visible window) |\n| `--wait-for <mode>` | Interaction wait mode: `none` (default), `interaction`, or `force` |\n| `--wait-for-interaction` | Alias for `--wait-for interaction` |\n| `--wait-for-login` | Alias for `--wait-for interaction` |\n| `--timeout <ms>` | Page load timeout (default: 30000) |\n| `--interaction-timeout <ms>` | Login/CAPTCHA wait timeout (default: 600000 = 10 min) |\n| `--interaction-poll-interval <ms>` | Poll interval for interaction checks (default: 1500) |\n| `--download-media` | Download images/videos to local `imgs/` and `videos/`, rewrite markdown links. Requires `--output` |\n| `--media-dir <dir>` | Base directory for downloaded media (default: same as `--output` directory) |\n| `--cdp-url <url>` | Reuse existing Chrome DevTools Protocol endpoint |\n| `--browser-path <path>` | Custom Chrome/Chromium binary path |\n| `--chrome-profile-dir <path>` | Chrome user data directory (default: `BAOYU_CHROME_PROFILE_DIR` env or `./baoyu-skills/chrome-profile`) |\n| `--debug-dir <dir>` | Write debug artifacts (document.json, markdown.md, page.html, network.json) |\n\n## Agent Quality Gate\n\n**CRITICAL**: treat default headless capture as provisional. Some sites render differently in headless mode and can silently return low-quality content without failing the CLI.\n\nAfter every headless run, inspect the saved markdown. See [references/quality-gate.md](references/quality-gate.md) for the full checklist, recovery workflow, and capture-mode table. Read it whenever a run looks suspicious or the user asks about login/CAPTCHA handling.\n\n## Output Path Generation\n\nThe agent must construct the output file path — `baoyu-fetch` does not auto-generate paths.\n\n**Algorithm**:\n1. Determine base directory from EXTEND.md `default_output_dir` or default `./url-to-markdown/`\n2. Extract domain from URL (e.g., `example.com`)\n3. Generate slug from URL path or page title (kebab-case, 2-6 words)\n4. Construct: `{base_dir}/{domain}/{slug}/{slug}.md` — each URL gets its own directory so media files stay isolated\n5. Conflict resolution: append timestamp `{slug}-YYYYMMDD-HHMMSS/{slug}-YYYYMMDD-HHMMSS.md`\n\nPass the constructed path to `--output`. Media files (`--download-media`) are saved into subdirectories next to the markdown file, keeping each URL's assets self-contained.\n\n## Adapters & Media\n\nSee [references/adapters.md](references/adapters.md) for the adapter catalog (X, YouTube, Hacker News, generic), per-adapter notes, the media download flow (`ask` / always / never), and the JSON output schema. Read it before answering adapter-specific questions or handling media prompts.\n\n## Environment Variables\n\n| Variable | Description |\n|----------|-------------|\n| `BAOYU_CHROME_PROFILE_DIR` | Chrome user data directory (can also use `--chrome-profile-dir`) |\n\n**Troubleshooting**: Chrome not found → use `--browser-path`. Timeout → increase `--timeout`. Login/CAPTCHA → `--wait-for interaction`. Debug → `--debug-dir` to inspect captured HTML and network logs.\n\n## Extension Support\n\nCustom configurations via EXTEND.md. See **Preferences** section above for paths and supported keys.\n\nFile v1.115.4:_meta.json\n\n{\n  \"ownerId\": \"kn7csrrndw79hpke5d0gsnx93d82k67r\",\n  \"slug\": \"baoyu-url-to-markdown\",\n  \"version\": \"1.115.4\",\n  \"publishedAt\": 1778543417162\n}\n\nFile v1.115.4:references/adapters.md\n\n# Adapters & Media\n\nRead when choosing an adapter, handling media, or answering adapter-specific questions.\n\n## Built-in Adapters\n\n| Adapter | URLs | Key Features |\n|---------|------|-------------|\n| `x` | x.com, twitter.com | Tweets, threads, X Articles, media, login detection |\n| `youtube` | youtube.com, youtu.be | Transcript/captions, chapters, cover image, metadata |\n| `hn` | news.ycombinator.com | Threaded comments, story metadata, nested replies |\n| `generic` | Any URL (fallback) | Defuddle extraction, Readability fallback, auto-scroll, network idle detection |\n\nAdapter is auto-selected based on URL. Override with `--adapter <name>`.\n\n### YouTube\n\n- Extracts transcripts/captions when available\n- Transcript format: `[MM:SS] Text segment` with chapter headings\n- Availability depends on YouTube exposing a caption track; videos with captions disabled or restricted playback may produce description-only output\n- Use `--wait-for force` if the page needs time to finish loading player metadata\n\n### X/Twitter\n\n- Extracts single tweets, threads, and X Articles\n- Auto-detects login state; if logged out and content requires auth, JSON output shows `\"status\": \"needs_interaction\"`\n- Use `--wait-for interaction` for login-protected content\n\n### Hacker News\n\n- Parses threaded comments with proper nesting and reply hierarchy\n- Includes story metadata (title, URL, author, score, comment count)\n- Shows comment deletion/dead status\n\n## Media Download Workflow\n\nDriven by `download_media` in EXTEND.md:\n\n| Setting | Behavior |\n|---------|----------|\n| `1` (always) | Run CLI with `--download-media --output <path>` |\n| `0` (never) | Run CLI with `--output <path>` (no media download) |\n| `ask` (default) | Follow the ask-each-time flow below |\n\n### Ask-Each-Time Flow\n\n1. Run the CLI **without** `--download-media` with `--output <path>` → markdown saved\n2. Check the saved markdown for remote media URLs (`https://` in image/video links)\n3. **If no remote media found** → done, no prompt needed\n4. **If remote media found** → ask via `AskUserQuestion`:\n   - header: \"Media\", question: \"Download N images/videos to local files?\"\n   - \"Yes\" — Download to local directories\n   - \"No\" — Keep remote URLs\n5. If the user confirms → run the CLI **again** with `--download-media --output <same-path>` (overwrites markdown with localized links)\n\n### Media Layout\n\nWhen `--download-media` is enabled:\n\n- Images → `imgs/` next to the output file (or `--media-dir`)\n- Videos → `videos/` next to the output file (or `--media-dir`)\n- Markdown media links are rewritten to local relative paths\n\n## Output Format\n\nMarkdown to stdout (or file with `--output`).\n\nJSON output (`--format json`) returns structured data:\n\n- `adapter` — which adapter handled the URL\n- `status` — `\"ok\"` or `\"needs_interaction\"`\n- `login` — login state detection (`logged_in`, `logged_out`, `unknown`)\n- `interaction` — interaction gate details (kind, provider, prompt)\n- `document` — structured content (url, title, author, publishedAt, content blocks, metadata)\n- `media` — collected media assets with url, kind, role\n- `markdown` — converted markdown text\n- `downloads` — media download results (when `--download-media` used)\n\nFile v1.115.4:references/config/first-time-setup.md\n\n---\nname: first-time-setup\ndescription: First-time setup flow for baoyu-url-to-markdown preferences\n---\n\n# First-Time Setup\n\n## Overview\n\nWhen no EXTEND.md is found, guide user through preference setup.\n\n**BLOCKING OPERATION**: This setup MUST complete before ANY other workflow steps. Do NOT:\n- Start converting URLs\n- Ask about URLs or output paths\n- Proceed to any conversion\n\nONLY ask the questions in this setup flow, save EXTEND.md, then continue.\n\n## Setup Flow\n\n```\nNo EXTEND.md found\n        |\n        v\n+---------------------+\n| AskUserQuestion     |\n| (all questions)     |\n+---------------------+\n        |\n        v\n+---------------------+\n| Create EXTEND.md    |\n+---------------------+\n        |\n        v\n    Continue conversion\n```\n\n## Questions\n\n**Language**: Use user's input language or saved language preference.\n\nUse AskUserQuestion with ALL questions in ONE call:\n\n### Question 1: Download Media\n\n```yaml\nheader: \"Media\"\nquestion: \"How to handle images and videos in pages?\"\noptions:\n  - label: \"Ask each time (Recommended)\"\n    description: \"After saving markdown, ask whether to download media\"\n  - label: \"Always download\"\n    description: \"Always download media to local imgs/ and videos/ directories\"\n  - label: \"Never download\"\n    description: \"Keep original remote URLs in markdown\"\n```\n\n### Question 2: Default Output Directory\n\n```yaml\nheader: \"Output\"\nquestion: \"Default output directory?\"\noptions:\n  - label: \"url-to-markdown (Recommended)\"\n    description: \"Save to ./url-to-markdown/{domain}/{slug}.md\"\n```\n\nNote: User will likely choose \"Other\" to type a custom path.\n\n### Question 3: Save Location\n\n```yaml\nheader: \"Save\"\nquestion: \"Where to save preferences?\"\noptions:\n  - label: \"User (Recommended)\"\n    description: \"~/.baoyu-skills/ (all projects)\"\n  - label: \"Project\"\n    description: \".baoyu-skills/ (this project only)\"\n```\n\n## Save Locations\n\n| Choice | Path | Scope |\n|--------|------|-------|\n| User | `~/.baoyu-skills/baoyu-url-to-markdown/EXTEND.md` | All projects |\n| Project | `.baoyu-skills/baoyu-url-to-markdown/EXTEND.md` | Current project |\n\n## After Setup\n\n1. Create directory if needed\n2. Write EXTEND.md\n3. Confirm: \"Preferences saved to [path]\"\n4. Continue with conversion using saved preferences\n\n## EXTEND.md Template\n\n```md\ndownload_media: [ask/1/0]\ndefault_output_dir: [path or empty]\n```\n\n## Modifying Preferences Later\n\nUsers can edit EXTEND.md directly or delete it to trigger setup again.\n\nFile v1.115.4:references/quality-gate.md\n\n# Quality Gate & Recovery\n\nHeadless Chrome can silently return low-quality content — layout shells, login walls, or framework payloads — without the CLI returning a non-zero exit code. Read this after every headless run so you can catch and recover from those cases.\n\n## Checks the Agent Must Run\n\n1. Confirm the markdown title matches the target page, not a generic site shell\n2. Confirm the body contains the expected article/page content, not just navigation, footer, or a generic error\n3. Watch for obvious failure signs:\n   - `Application error`\n   - `This page could not be found`\n   - Login, signup, subscribe, or verification shells\n   - Extremely short markdown for a page that should be long-form\n   - Raw framework payloads or mostly boilerplate content\n4. Do NOT accept a run as successful just because the CLI exited `0`\n\n**Tip**: run with `--format json` to get structured signals including `status`, `login.state`, and `interaction`. `\"status\": \"needs_interaction\"` means the page requires manual interaction.\n\n## Recovery Workflow\n\n1. Start headless (default) unless there is already a clear reason to use interaction mode\n2. Review markdown quality immediately after the run\n3. If the content is low quality or indicates login/CAPTCHA:\n   - `--wait-for interaction` for auto-detected gates (login, CAPTCHA, Cloudflare)\n   - `--wait-for force` when the page needs manual browsing, scroll loading, or complex interaction\n4. If `--wait-for` is used, tell the user exactly what to do:\n   - Login required → sign in in the browser\n   - CAPTCHA visible → solve it\n   - Slow loading → wait until content is visible\n   - `--wait-for force` → press Enter when ready\n5. If JSON output shows `\"status\": \"needs_interaction\"`, switch to `--wait-for interaction` automatically\n\n## Capture Modes\n\n| Mode | Behavior | Use When |\n|------|----------|----------|\n| Default | Headless Chrome, auto-extract on network idle | Public pages, static content |\n| `--headless` | Explicit headless (same as default) | Clarify intent |\n| `--wait-for interaction` | Opens visible Chrome, auto-detects login/CAPTCHA gates, waits for them to clear, then continues | Login-required, CAPTCHA-protected |\n| `--wait-for force` | Opens visible Chrome, auto-detects OR accepts Enter keypress to continue | Complex flows, lazy loading, paywalls |\n\n**Interaction gate auto-detection**: Cloudflare Turnstile / \"just a moment\" pages, Google reCAPTCHA, hCaptcha, custom challenge / verification screens.\n\nFile v1.115.4:scripts/package.json\n\n{\n  \"name\": \"baoyu-url-to-markdown-scripts\",\n  \"private\": true,\n  \"type\": \"module\",\n  \"bin\": {\n    \"baoyu-fetch\": \"./baoyu-fetch\"\n  },\n  \"scripts\": {\n    \"reader\": \"bun ./lib/cli.ts\"\n  },\n  \"dependencies\": {\n    \"@mozilla/readability\": \"^0.6.0\",\n    \"chrome-launcher\": \"^1.2.1\",\n    \"defuddle\": \"^0.17.0\",\n    \"jsdom\": \"^29.0.2\",\n    \"remark-gfm\": \"^4.0.1\",\n    \"remark-parse\": \"^11.0.0\",\n    \"turndown\": \"^7.2.0\",\n    \"turndown-plugin-gfm\": \"^1.0.2\",\n    \"unified\": \"^11.0.5\",\n    \"ws\": \"^8.18.3\"\n  },\n  \"overrides\": {\n    \"@xmldom/xmldom\": \"0.8.13\"\n  }\n}\n\nArchive v1.109.0: 49 files, 94970 bytes\n\nFiles: references/adapters.md (3231b), references/config/first-time-setup.md (2459b), references/quality-gate.md (2490b), scripts/baoyu-fetch (122b), scripts/bun.lock (35037b), scripts/lib/adapters/generic/index.ts (2451b), scripts/lib/adapters/hn/index.ts (10454b), scripts/lib/adapters/index.ts (834b), scripts/lib/adapters/types.ts (2100b), scripts/lib/adapters/x/article.ts (12677b), scripts/lib/adapters/x/index.ts (4048b), scripts/lib/adapters/x/login.ts (2258b), scripts/lib/adapters/x/match.ts (297b), scripts/lib/adapters/x/payloads.ts (1681b), scripts/lib/adapters/x/session.ts (1404b), scripts/lib/adapters/x/shared.ts (12980b), scripts/lib/adapters/x/single.ts (2583b), scripts/lib/adapters/x/thread-loader.ts (9078b), scripts/lib/adapters/x/thread.ts (9196b), scripts/lib/adapters/x/types.ts (600b), scripts/lib/adapters/youtube/index.ts (898b), scripts/lib/adapters/youtube/transcript.ts (13301b), scripts/lib/adapters/youtube/utils.ts (6523b), scripts/lib/browser/cdp-client.ts (6879b), scripts/lib/browser/chrome-launcher.ts (5883b), scripts/lib/browser/cookie-sidecar.ts (2941b), scripts/lib/browser/interaction-gates.ts (4095b), scripts/lib/browser/network-journal.ts (7153b), scripts/lib/browser/page-snapshot.ts (3429b), scripts/lib/browser/profile.ts (5670b), scripts/lib/browser/session.ts (4642b), scripts/lib/cli.ts (7066b), scripts/lib/commands/convert.ts (17456b), scripts/lib/extract/document.ts (846b), scripts/lib/extract/html-cleaner.ts (11016b), scripts/lib/extract/html-extractor.ts (2243b), scripts/lib/extract/html-to-markdown.ts (21344b), scripts/lib/extract/markdown-renderer.ts (4661b), scripts/lib/media/default-downloader.ts (5215b), scripts/lib/media/markdown-media.ts (13186b), scripts/lib/media/media-utils.ts (6222b), scripts/lib/media/types.ts (710b), scripts/lib/types/defuddle-node.d.ts (605b), scripts/lib/types/shims.d.ts (422b), scripts/lib/utils/logger.ts (645b), scripts/lib/utils/url.ts (300b), scripts/package.json (557b), SKILL.md (8520b), _meta.json (142b)\n\nFile v1.109.0:SKILL.md\n\n---\nname: baoyu-url-to-markdown\ndescription: Fetch any URL and convert to markdown using baoyu-fetch CLI (Chrome CDP with site-specific adapters). Built-in adapters for X/Twitter, YouTube transcripts, Hacker News threads, and generic pages via Defuddle. Handles login/CAPTCHA via interaction wait modes. Use when user wants to save a webpage as markdown.\nversion: 1.61.0\nmetadata:\n  openclaw:\n    homepage: https://github.com/JimLiu/baoyu-skills#baoyu-url-to-markdown\n    requires:\n      anyBins:\n        - bun\n---\n\n# URL to Markdown\n\nFetches any URL via `baoyu-fetch` CLI (Chrome CDP + site-specific adapters) and converts it to clean markdown.\n\n## User Input Tools\n\nWhen this skill prompts the user, follow this tool-selection rule (priority order):\n\n1. **Prefer built-in user-input tools** exposed by the current agent runtime — e.g., `AskUserQuestion`, `request_user_input`, `clarify`, `ask_user`, or any equivalent.\n2. **Fallback**: if no such tool exists, emit a numbered plain-text message and ask the user to reply with the chosen number/answer for each question.\n3. **Batching**: if the tool supports multiple questions per call, combine all applicable questions into a single call; if only single-question, ask them one at a time in priority order.\n\nConcrete `AskUserQuestion` references below are examples — substitute the local equivalent in other runtimes.\n\n## CLI Setup\n\n**Important**: The CLI source is vendored in `{baseDir}/scripts/lib`. `scripts/package.json` installs only third-party runtime dependencies.\n\n**Agent Execution Instructions**:\n1. Determine this SKILL.md file's directory path as `{baseDir}`\n2. Resolve `${BUN}` runtime: if `bun` installed → `bun`; else suggest installing Bun\n3. If `{baseDir}/scripts/node_modules` does not exist, run `${BUN} install --cwd {baseDir}/scripts`\n4. `${READER}` = `{baseDir}/scripts/baoyu-fetch`\n5. Replace all `${READER}` in this document with the resolved value\n\n## Preferences (EXTEND.md)\n\nCheck EXTEND.md in priority order — the first one found wins:\n\n| Priority | Path | Scope |\n|----------|------|-------|\n| 1 | `.baoyu-skills/baoyu-url-to-markdown/EXTEND.md` | Project |\n| 2 | `${XDG_CONFIG_HOME:-$HOME/.config}/baoyu-skills/baoyu-url-to-markdown/EXTEND.md` | XDG |\n| 3 | `$HOME/.baoyu-skills/baoyu-url-to-markdown/EXTEND.md` | User home |\n\n| Result | Action |\n|--------|--------|\n| Found | Read, parse, apply settings |\n| Not found | **MUST** run first-time setup (see below) — do NOT silently create defaults |\n\n**EXTEND.md supports**: download media by default, default output directory.\n\n### First-Time Setup ⛔ BLOCKING\n\nWhen EXTEND.md is not found, you **MUST** use `AskUserQuestion` to gather preferences before creating EXTEND.md. **NEVER** create EXTEND.md with silent defaults. Generation is BLOCKED until setup completes. Batch all three questions into a single call:\n\n- **Q1 — Media** (header \"Media\"): \"How to handle images and videos in pages?\"\n  - \"Ask each time (Recommended)\" — Prompt after each save\n  - \"Always download\" — Download to local `imgs/` and `videos/`\n  - \"Never download\" — Keep remote URLs\n- **Q2 — Output** (header \"Output\"): \"Default output directory?\"\n  - \"url-to-markdown (Recommended)\" — Save to `./url-to-markdown/{domain}/{slug}.md`\n  - User may pick \"Other\" and type a custom path\n- **Q3 — Save** (header \"Save\"): \"Where to save preferences?\"\n  - \"User (Recommended)\" — `~/.baoyu-skills/` (all projects)\n  - \"Project\" — `.baoyu-skills/` (this project only)\n\nAfter answers, write EXTEND.md, confirm \"Preferences saved to [path]\", then continue.\n\nFull template: [references/config/first-time-setup.md](references/config/first-time-setup.md).\n\n### Supported Keys\n\n| Key | Default | Values | Description |\n|-----|---------|--------|-------------|\n| `download_media` | `ask` | `ask` / `1` / `0` | `ask` = prompt each time, `1` = always, `0` = never |\n| `default_output_dir` | empty | path or empty | Default output directory (empty = `./url-to-markdown/`) |\n\n**EXTEND.md → CLI mapping**:\n\n| EXTEND.md key | CLI argument | Notes |\n|---------------|-------------|-------|\n| `download_media: 1` | `--download-media` | Requires `--output` to be set |\n| `default_output_dir: ./posts/` | Agent constructs `--output ./posts/{domain}/{slug}.md` | Agent generates path, not a direct flag |\n\n**Value priority**: CLI arguments → EXTEND.md → skill defaults.\n\n## Usage\n\n```bash\n# Default: headless capture, markdown to stdout\n${READER} <url>\n\n# Save to file\n${READER} <url> --output article.md\n\n# Save with media download\n${READER} <url> --output article.md --download-media\n\n# Wait for interaction (login/CAPTCHA) — auto-detect and continue\n${READER} <url> --wait-for interaction --output article.md\n\n# Wait for interaction — manual control (Enter to continue)\n${READER} <url> --wait-for force --output article.md\n\n# JSON output\n${READER} <url> --format json --output article.json\n\n# Force specific adapter\n${READER} <url> --adapter youtube --output transcript.md\n```\n\n## Options\n\n| Option | Description |\n|--------|-------------|\n| `<url>` | URL to fetch |\n| `--output <path>` | Output file path (default: stdout) |\n| `--format <type>` | Output format: `markdown` (default) or `json` |\n| `--json` | Shorthand for `--format json` |\n| `--adapter <name>` | Force adapter: `x`, `youtube`, `hn`, or `generic` (default: auto-detect) |\n| `--headless` | Force headless Chrome (no visible window) |\n| `--wait-for <mode>` | Interaction wait mode: `none` (default), `interaction`, or `force` |\n| `--wait-for-interaction` | Alias for `--wait-for interaction` |\n| `--wait-for-login` | Alias for `--wait-for interaction` |\n| `--timeout <ms>` | Page load timeout (default: 30000) |\n| `--interaction-timeout <ms>` | Login/CAPTCHA wait timeout (default: 600000 = 10 min) |\n| `--interaction-poll-interval <ms>` | Poll interval for interaction checks (default: 1500) |\n| `--download-media` | Download images/videos to local `imgs/` and `videos/`, rewrite markdown links. Requires `--output` |\n| `--media-dir <dir>` | Base directory for downloaded media (default: same as `--output` directory) |\n| `--cdp-url <url>` | Reuse existing Chrome DevTools Protocol endpoint |\n| `--browser-path <path>` | Custom Chrome/Chromium binary path |\n| `--chrome-profile-dir <path>` | Chrome user data directory (default: `BAOYU_CHROME_PROFILE_DIR` env or `./baoyu-skills/chrome-profile`) |\n| `--debug-dir <dir>` | Write debug artifacts (document.json, markdown.md, page.html, network.json) |\n\n## Agent Quality Gate\n\n**CRITICAL**: treat default headless capture as provisional. Some sites render differently in headless mode and can silently return low-quality content without failing the CLI.\n\nAfter every headless run, inspect the saved markdown. See [references/quality-gate.md](references/quality-gate.md) for the full checklist, recovery workflow, and capture-mode table. Read it whenever a run looks suspicious or the user asks about login/CAPTCHA handling.\n\n## Output Path Generation\n\nThe agent must construct the output file path — `baoyu-fetch` does not auto-generate paths.\n\n**Algorithm**:\n1. Determine base directory from EXTEND.md `default_output_dir` or default `./url-to-markdown/`\n2. Extract domain from URL (e.g., `example.com`)\n3. Generate slug from URL path or page title (kebab-case, 2-6 words)\n4. Construct: `{base_dir}/{domain}/{slug}/{slug}.md` — each URL gets its own directory so media files stay isolated\n5. Conflict resolution: append timestamp `{slug}-YYYYMMDD-HHMMSS/{slug}-YYYYMMDD-HHMMSS.md`\n\nPass the constructed path to `--output`. Media files (`--download-media`) are saved into subdirectories next to the markdown file, keeping each URL's assets self-contained.\n\n## Adapters & Media\n\nSee [references/adapters.md](references/adapters.md) for the adapter catalog (X, YouTube, Hacker News, generic), per-adapter notes, the media download flow (`ask` / always / never), and the JSON output schema. Read it before answering adapter-specific questions or handling media prompts.\n\n## Environment Variables\n\n| Variable | Description |\n|----------|-------------|\n| `BAOYU_CHROME_PROFILE_DIR` | Chrome user data directory (can also use `--chrome-profile-dir`) |\n\n**Troubleshooting**: Chrome not found → use `--browser-path`. Timeout → increase `--timeout`. Login/CAPTCHA → `--wait-for interaction`. Debug → `--debug-dir` to inspect captured HTML and network logs.\n\n## Extension Support\n\nCustom configurations via EXTEND.md. See **Preferences** section above for paths and supported keys.\n\nFile v1.109.0:_meta.json\n\n{\n  \"ownerId\": \"kn7csrrndw79hpke5d0gsnx93d82k67r\",\n  \"slug\": \"baoyu-url-to-markdown\",\n  \"version\": \"1.109.0\",\n  \"publishedAt\": 1776796822242\n}\n\nFile v1.109.0:references/adapters.md\n\n# Adapters & Media\n\nRead when choosing an adapter, handling media, or answering adapter-specific questions.\n\n## Built-in Adapters\n\n| Adapter | URLs | Key Features |\n|---------|------|-------------|\n| `x` | x.com, twitter.com | Tweets, threads, X Articles, media, login detection |\n| `youtube` | youtube.com, youtu.be | Transcript/captions, chapters, cover image, metadata |\n| `hn` | news.ycombinator.com | Threaded comments, story metadata, nested replies |\n| `generic` | Any URL (fallback) | Defuddle extraction, Readability fallback, auto-scroll, network idle detection |\n\nAdapter is auto-selected based on URL. Override with `--adapter <name>`.\n\n### YouTube\n\n- Extracts transcripts/captions when available\n- Transcript format: `[MM:SS] Text segment` with chapter headings\n- Availability depends on YouTube exposing a caption track; videos with captions disabled or restricted playback may produce description-only output\n- Use `--wait-for force` if the page needs time to finish loading player metadata\n\n### X/Twitter\n\n- Extracts single tweets, threads, and X Articles\n- Auto-detects login state; if logged out and content requires auth, JSON output shows `\"status\": \"needs_interaction\"`\n- Use `--wait-for interaction` for login-protected content\n\n### Hacker News\n\n- Parses threaded comments with proper nesting and reply hierarchy\n- Includes story metadata (title, URL, author, score, comment count)\n- Shows comment deletion/dead status\n\n## Media Download Workflow\n\nDriven by `download_media` in EXTEND.md:\n\n| Setting | Behavior |\n|---------|----------|\n| `1` (always) | Run CLI with `--download-media --output <path>` |\n| `0` (never) | Run CLI with `--output <path>` (no media download) |\n| `ask` (default) | Follow the ask-each-time flow below |\n\n### Ask-Each-Time Flow\n\n1. Run the CLI **without** `--download-media` with `--output <path>` → markdown saved\n2. Check the saved markdown for remote media URLs (`https://` in image/video links)\n3. **If no remote media found** → done, no prompt needed\n4. **If remote media found** → ask via `AskUserQuestion`:\n   - header: \"Media\", question: \"Download N images/videos to local files?\"\n   - \"Yes\" — Download to local directories\n   - \"No\" — Keep remote URLs\n5. If the user confirms → run the CLI **again** with `--download-media --output <same-path>` (overwrites markdown with localized links)\n\n### Media Layout\n\nWhen `--download-media` is enabled:\n\n- Images → `imgs/` next to the output file (or `--media-dir`)\n- Videos → `videos/` next to the output file (or `--media-dir`)\n- Markdown media links are rewritten to local relative paths\n\n## Output Format\n\nMarkdown to stdout (or file with `--output`).\n\nJSON output (`--format json`) returns structured data:\n\n- `adapter` — which adapter handled the URL\n- `status` — `\"ok\"` or `\"needs_interaction\"`\n- `login` — login state detection (`logged_in`, `logged_out`, `unknown`)\n- `interaction` — interaction gate details (kind, provider, prompt)\n- `document` — structured content (url, title, author, publishedAt, content blocks, metadata)\n- `media` — collected media assets with url, kind, role\n- `markdown` — converted markdown text\n- `downloads` — media download results (when `--download-media` used)\n\nFile v1.109.0:references/config/first-time-setup.md\n\n---\nname: first-time-setup\ndescription: First-time setup flow for baoyu-url-to-markdown preferences\n---\n\n# First-Time Setup\n\n## Overview\n\nWhen no EXTEND.md is found, guide user through preference setup.\n\n**BLOCKING OPERATION**: This setup MUST complete before ANY other workflow steps. Do NOT:\n- Start converting URLs\n- Ask about URLs or output paths\n- Proceed to any conversion\n\nONLY ask the questions in this setup flow, save EXTEND.md, then continue.\n\n## Setup Flow\n\n```\nNo EXTEND.md found\n        |\n        v\n+---------------------+\n| AskUserQuestion     |\n| (all questions)     |\n+---------------------+\n        |\n        v\n+---------------------+\n| Create EXTEND.md    |\n+---------------------+\n        |\n        v\n    Continue conversion\n```\n\n## Questions\n\n**Language**: Use user's input language or saved language preference.\n\nUse AskUserQuestion with ALL questions in ONE call:\n\n### Question 1: Download Media\n\n```yaml\nheader: \"Media\"\nquestion: \"How to handle images and videos in pages?\"\noptions:\n  - label: \"Ask each time (Recommended)\"\n    description: \"After saving markdown, ask whether to download media\"\n  - label: \"Always download\"\n    description: \"Always download media to local imgs/ and videos/ directories\"\n  - label: \"Never download\"\n    description: \"Keep original remote URLs in markdown\"\n```\n\n### Question 2: Default Output Directory\n\n```yaml\nheader: \"Output\"\nquestion: \"Default output directory?\"\noptions:\n  - label: \"url-to-markdown (Recommended)\"\n    description: \"Save to ./url-to-markdown/{domain}/{slug}.md\"\n```\n\nNote: User will likely choose \"Other\" to type a custom path.\n\n### Question 3: Save Location\n\n```yaml\nheader: \"Save\"\nquestion: \"Where to save preferences?\"\noptions:\n  - label: \"User (Recommended)\"\n    description: \"~/.baoyu-skills/ (all projects)\"\n  - label: \"Project\"\n    description: \".baoyu-skills/ (this project only)\"\n```\n\n## Save Locations\n\n| Choice | Path | Scope |\n|--------|------|-------|\n| User | `~/.baoyu-skills/baoyu-url-to-markdown/EXTEND.md` | All projects |\n| Project | `.baoyu-skills/baoyu-url-to-markdown/EXTEND.md` | Current project |\n\n## After Setup\n\n1. Create directory if needed\n2. Write EXTEND.md\n3. Confirm: \"Preferences saved to [path]\"\n4. Continue with conversion using saved preferences\n\n## EXTEND.md Template\n\n```md\ndownload_media: [ask/1/0]\ndefault_output_dir: [path or empty]\n```\n\n## Modifying Preferences Later\n\nUsers can edit EXTEND.md directly or delete it to trigger setup again.\n\nFile v1.109.0:references/quality-gate.md\n\n# Quality Gate & Recovery\n\nHeadless Chrome can silently return low-quality content — layout shells, login walls, or framework payloads — without the CLI returning a non-zero exit code. Read this after every headless run so you can catch and recover from those cases.\n\n## Checks the Agent Must Run\n\n1. Confirm the markdown title matches the target page, not a generic site shell\n2. Confirm the body contains the expected article/page content, not just navigation, footer, or a generic error\n3. Watch for obvious failure signs:\n   - `Application error`\n   - `This page could not be found`\n   - Login, signup, subscribe, or verification shells\n   - Extremely short markdown for a page that should be long-form\n   - Raw framework payloads or mostly boilerplate content\n4. Do NOT accept a run as successful just because the CLI exited `0`\n\n**Tip**: run with `--format json` to get structured signals including `status`, `login.state`, and `interaction`. `\"status\": \"needs_interaction\"` means the page requires manual interaction.\n\n## Recovery Workflow\n\n1. Start headless (default) unless there is already a clear reason to use interaction mode\n2. Review markdown quality immediately after the run\n3. If the content is low quality or indicates login/CAPTCHA:\n   - `--wait-for interaction` for auto-detected gates (login, CAPTCHA, Cloudflare)\n   - `--wait-for force` when the page needs manual browsing, scroll loading, or complex interaction\n4. If `--wait-for` is used, tell the user exactly what to do:\n   - Login required → sign in in the browser\n   - CAPTCHA visible → solve it\n   - Slow loading → wait until content is visible\n   - `--wait-for force` → press Enter when ready\n5. If JSON output shows `\"status\": \"needs_interaction\"`, switch to `--wait-for interaction` automatically\n\n## Capture Modes\n\n| Mode | Behavior | Use When |\n|------|----------|----------|\n| Default | Headless Chrome, auto-extract on network idle | Public pages, static content |\n| `--headless` | Explicit headless (same as default) | Clarify intent |\n| `--wait-for interaction` | Opens visible Chrome, auto-detects login/CAPTCHA gates, waits for them to clear, then continues | Login-required, CAPTCHA-protected |\n| `--wait-for force` | Opens visible Chrome, auto-detects OR accepts Enter keypress to continue | Complex flows, lazy loading, paywalls |\n\n**Interaction gate auto-detection**: Cloudflare Turnstile / \"just a moment\" pages, Google reCAPTCHA, hCaptcha, custom challenge / verification screens.\n\nFile v1.109.0:scripts/package.json\n\n{\n  \"name\": \"baoyu-url-to-markdown-scripts\",\n  \"private\": true,\n  \"type\": \"module\",\n  \"bin\": {\n    \"baoyu-fetch\": \"./baoyu-fetch\"\n  },\n  \"scripts\": {\n    \"reader\": \"bun ./lib/cli.ts\"\n  },\n  \"dependencies\": {\n    \"@mozilla/readability\": \"^0.6.0\",\n    \"chrome-launcher\": \"^1.2.1\",\n    \"defuddle\": \"^0.17.0\",\n    \"jsdom\": \"^29.0.2\",\n    \"remark-gfm\": \"^4.0.1\",\n    \"remark-parse\": \"^11.0.0\",\n    \"turndown\": \"^7.2.0\",\n    \"turndown-plugin-gfm\": \"^1.0.2\",\n    \"unified\": \"^11.0.5\",\n    \"ws\": \"^8.18.3\"\n  },\n  \"overrides\": {\n    \"@xmldom/xmldom\": \"0.8.13\"\n  }\n}\n\nArchive v1.103.1: 49 files, 110029 bytes\n\nFiles: references/config/first-time-setup.md (2459b), scripts/bun.lock (56924b), scripts/package.json (157b), scripts/vendor/baoyu-fetch/package.json (1624b), scripts/vendor/baoyu-fetch/README.md (4477b), scripts/vendor/baoyu-fetch/README.zh-CN.md (4326b), scripts/vendor/baoyu-fetch/src/adapters/generic/index.ts (2451b), scripts/vendor/baoyu-fetch/src/adapters/hn/index.ts (10454b), scripts/vendor/baoyu-fetch/src/adapters/index.ts (834b), scripts/vendor/baoyu-fetch/src/adapters/types.ts (2100b), scripts/vendor/baoyu-fetch/src/adapters/x/article.ts (12122b), scripts/vendor/baoyu-fetch/src/adapters/x/index.ts (4048b), scripts/vendor/baoyu-fetch/src/adapters/x/login.ts (2258b), scripts/vendor/baoyu-fetch/src/adapters/x/match.ts (297b), scripts/vendor/baoyu-fetch/src/adapters/x/payloads.ts (1681b), scripts/vendor/baoyu-fetch/src/adapters/x/session.ts (1404b), scripts/vendor/baoyu-fetch/src/adapters/x/shared.ts (11727b), scripts/vendor/baoyu-fetch/src/adapters/x/single.ts (2583b), scripts/vendor/baoyu-fetch/src/adapters/x/thread-loader.ts (9078b), scripts/vendor/baoyu-fetch/src/adapters/x/thread.ts (9196b), scripts/vendor/baoyu-fetch/src/adapters/x/types.ts (600b), scripts/vendor/baoyu-fetch/src/adapters/youtube/index.ts (898b), scripts/vendor/baoyu-fetch/src/adapters/youtube/transcript.ts (13301b), scripts/vendor/baoyu-fetch/src/adapters/youtube/utils.ts (6523b), scripts/vendor/baoyu-fetch/src/browser/cdp-client.ts (6879b), scripts/vendor/baoyu-fetch/src/browser/chrome-launcher.ts (5883b), scripts/vendor/baoyu-fetch/src/browser/cookie-sidecar.ts (2941b), scripts/vendor/baoyu-fetch/src/browser/interaction-gates.ts (4095b), scripts/vendor/baoyu-fetch/src/browser/network-journal.ts (7153b), scripts/vendor/baoyu-fetch/src/browser/page-snapshot.ts (3429b), scripts/vendor/baoyu-fetch/src/browser/profile.ts (5670b), scripts/vendor/baoyu-fetch/src/browser/session.ts (4642b), scripts/vendor/baoyu-fetch/src/cli.ts (7066b), scripts/vendor/baoyu-fetch/src/commands/convert.ts (17456b), scripts/vendor/baoyu-fetch/src/extract/document.ts (846b), scripts/vendor/baoyu-fetch/src/extract/html-cleaner.ts (11016b), scripts/vendor/baoyu-fetch/src/extract/html-extractor.ts (2243b), scripts/vendor/baoyu-fetch/src/extract/html-to-markdown.ts (21344b), scripts/vendor/baoyu-fetch/src/extract/markdown-renderer.ts (4661b), scripts/vendor/baoyu-fetch/src/media/default-downloader.ts (5215b), scripts/vendor/baoyu-fetch/src/media/markdown-media.ts (13186b), scripts/vendor/baoyu-fetch/src/media/media-utils.ts (6222b), scripts/vendor/baoyu-fetch/src/media/types.ts (710b), scripts/vendor/baoyu-fetch/src/types/defuddle-node.d.ts (605b), scripts/vendor/baoyu-fetch/src/types/shims.d.ts (422b), scripts/vendor/baoyu-fetch/src/utils/logger.ts (645b), scripts/vendor/baoyu-fetch/src/utils/url.ts (300b), SKILL.md (15269b), _meta.json (142b)\n\nFile v1.103.1:SKILL.md\n\n---\nname: baoyu-url-to-markdown\ndescription: Fetch any URL and convert to markdown using baoyu-fetch CLI (Chrome CDP with site-specific adapters). Built-in adapters for X/Twitter, YouTube transcripts, Hacker News threads, and generic pages via Defuddle. Handles login/CAPTCHA via interaction wait modes. Use when user wants to save a webpage as markdown.\nversion: 1.60.0\nmetadata:\n  openclaw:\n    homepage: https://github.com/JimLiu/baoyu-skills#baoyu-url-to-markdown\n    requires:\n      anyBins:\n        - bun\n        - npx\n---\n\n# URL to Markdown\n\nFetches any URL via `baoyu-fetch` CLI (Chrome CDP + site-specific adapters) and converts it to clean markdown.\n\n## CLI Setup\n\n**Important**: The CLI source is vendored in the `scripts/vendor/baoyu-fetch/` subdirectory of this skill.\n\n**Agent Execution Instructions**:\n1. Determine this SKILL.md file's directory path as `{baseDir}`\n2. CLI entry point = `{baseDir}/scripts/vendor/baoyu-fetch/src/cli.ts`\n3. Resolve `${BUN_X}` runtime: if `bun` installed → `bun`; if `npx` available → `npx -y bun`; else suggest installing bun\n4. `${READER}` = `${BUN_X} {baseDir}/scripts/vendor/baoyu-fetch/src/cli.ts`\n5. Replace all `${READER}` in this document with the resolved value\n\n## Preferences (EXTEND.md)\n\nCheck EXTEND.md existence (priority order):\n\n```bash\n# macOS, Linux, WSL, Git Bash\ntest -f .baoyu-skills/baoyu-url-to-markdown/EXTEND.md && echo \"project\"\ntest -f \"${XDG_CONFIG_HOME:-$HOME/.config}/baoyu-skills/baoyu-url-to-markdown/EXTEND.md\" && echo \"xdg\"\ntest -f \"$HOME/.baoyu-skills/baoyu-url-to-markdown/EXTEND.md\" && echo \"user\"\n```\n\n```powershell\n# PowerShell (Windows)\nif (Test-Path .baoyu-skills/baoyu-url-to-markdown/EXTEND.md) { \"project\" }\n$xdg = if ($env:XDG_CONFIG_HOME) { $env:XDG_CONFIG_HOME } else { \"$HOME/.config\" }\nif (Test-Path \"$xdg/baoyu-skills/baoyu-url-to-markdown/EXTEND.md\") { \"xdg\" }\nif (Test-Path \"$HOME/.baoyu-skills/baoyu-url-to-markdown/EXTEND.md\") { \"user\" }\n```\n\n| Path | Location |\n|------|----------|\n| `.baoyu-skills/baoyu-url-to-markdown/EXTEND.md` | Project directory |\n| `$HOME/.baoyu-skills/baoyu-url-to-markdown/EXTEND.md` | User home |\n\n| Result | Action |\n|--------|--------|\n| Found | Read, parse, apply settings |\n| Not found | **MUST** run first-time setup (see below) — do NOT silently create defaults |\n\n**EXTEND.md Supports**: Download media by default | Default output directory\n\n### First-Time Setup (BLOCKING)\n\n**CRITICAL**: When EXTEND.md is not found, you **MUST use `AskUserQuestion`** to ask the user for their preferences before creating EXTEND.md. **NEVER** create EXTEND.md with defaults without asking. This is a **BLOCKING** operation — do NOT proceed with any conversion until setup is complete.\n\nUse `AskUserQuestion` with ALL questions in ONE call:\n\n**Question 1** — header: \"Media\", question: \"How to handle images and videos in pages?\"\n- \"Ask each time (Recommended)\" — After saving markdown, ask whether to download media\n- \"Always download\" — Always download media to local imgs/ and videos/ directories\n- \"Never download\" — Keep original remote URLs in markdown\n\n**Question 2** — header: \"Output\", question: \"Default output directory?\"\n- \"url-to-markdown (Recommended)\" — Save to ./url-to-markdown/{domain}/{slug}.md\n- (User may choose \"Other\" to type a custom path)\n\n**Question 3** — header: \"Save\", question: \"Where to save preferences?\"\n- \"User (Recommended)\" — ~/.baoyu-skills/ (all projects)\n- \"Project\" — .baoyu-skills/ (this project only)\n\nAfter user answers, create EXTEND.md at the chosen location, confirm \"Preferences saved to [path]\", then continue.\n\nFull reference: [references/config/first-time-setup.md](references/config/first-time-setup.md)\n\n### Supported Keys\n\n| Key | Default | Values | Description |\n|-----|---------|--------|-------------|\n| `download_media` | `ask` | `ask` / `1` / `0` | `ask` = prompt each time, `1` = always download, `0` = never |\n| `default_output_dir` | empty | path or empty | Default output directory (empty = `./url-to-markdown/`) |\n\n**EXTEND.md → CLI mapping**:\n| EXTEND.md key | CLI argument | Notes |\n|---------------|-------------|-------|\n| `download_media: 1` | `--download-media` | Requires `--output` to be set |\n| `default_output_dir: ./posts/` | Agent constructs `--output ./posts/{domain}/{slug}.md` | Agent generates path, not a direct CLI flag |\n\n**Value priority**:\n1. CLI arguments (`--download-media`, `--output`)\n2. EXTEND.md\n3. Skill defaults\n\n## Features\n\n- Chrome CDP for full JavaScript rendering via `baoyu-fetch` CLI\n- Site-specific adapters: X/Twitter, YouTube, Hacker News, generic (Defuddle)\n- Automatic adapter selection based on URL, or force with `--adapter`\n- Interaction gate detection: Cloudflare, reCAPTCHA, hCAPTCHA, custom challenges\n- Two capture modes: headless (default) or interactive with wait-for-interaction\n- Clean markdown output with YAML front matter\n- Structured JSON output available via `--format json`\n- X/Twitter: extracts tweets, threads, and X Articles with media\n- YouTube: transcript/caption extraction, chapters, cover images\n- Hacker News: threaded comment parsing with proper nesting\n- Generic: Defuddle extraction with Readability fallback\n- Download images and videos to local directories\n- Chrome profile persistence for authenticated sessions\n- Debug artifact output for troubleshooting\n\n## Usage\n\n```bash\n# Default: headless capture, markdown to stdout\n${READER} <url>\n\n# Save to file\n${READER} <url> --output article.md\n\n# Save with media download\n${READER} <url> --output article.md --download-media\n\n# Headless mode (explicit)\n${READER} <url> --headless --output article.md\n\n# Wait for interaction (login/CAPTCHA) — auto-detect and continue\n${READER} <url> --wait-for interaction --output article.md\n\n# Wait for interaction — manual control (Enter to continue)\n${READER} <url> --wait-for force --output article.md\n\n# JSON output\n${READER} <url> --format json --output article.json\n\n# Force specific adapter\n${READER} <url> --adapter youtube --output transcript.md\n\n# Connect to existing Chrome\n${READER} <url> --cdp-url http://localhost:9222 --output article.md\n\n# Debug artifacts\n${READER} <url> --output article.md --debug-dir ./debug/\n```\n\n## Options\n\n| Option | Description |\n|--------|-------------|\n| `<url>` | URL to fetch |\n| `--output <path>` | Output file path (default: stdout) |\n| `--format <type>` | Output format: `markdown` (default) or `json` |\n| `--json` | Shorthand for `--format json` |\n| `--adapter <name>` | Force adapter: `x`, `youtube`, `hn`, or `generic` (default: auto-detect) |\n| `--headless` | Force headless Chrome (no visible window) |\n| `--wait-for <mode>` | Interaction wait mode: `none` (default), `interaction`, or `force` |\n| `--wait-for-interaction` | Alias for `--wait-for interaction` |\n| `--wait-for-login` | Alias for `--wait-for interaction` |\n| `--timeout <ms>` | Page load timeout (default: 30000) |\n| `--interaction-timeout <ms>` | Login/CAPTCHA wait timeout (default: 600000 = 10 min) |\n| `--interaction-poll-interval <ms>` | Poll interval for interaction checks (default: 1500) |\n| `--download-media` | Download images/videos to local `imgs/` and `videos/`, rewrite markdown links. Requires `--output` |\n| `--media-dir <dir>` | Base directory for downloaded media (default: same as `--output` directory) |\n| `--cdp-url <url>` | Reuse existing Chrome DevTools Protocol endpoint |\n| `--browser-path <path>` | Custom Chrome/Chromium binary path |\n| `--chrome-profile-dir <path>` | Chrome user data directory (default: `BAOYU_CHROME_PROFILE_DIR` env or `./baoyu-skills/chrome-profile`) |\n| `--debug-dir <dir>` | Write debug artifacts (document.json, markdown.md, page.html, network.json) |\n\n## Capture Modes\n\n| Mode | Behavior | Use When |\n|------|----------|----------|\n| Default | Headless Chrome, auto-extract on network idle | Public pages, static content |\n| `--headless` | Explicit headless (same as default) | Clarify intent |\n| `--wait-for interaction` | Opens visible Chrome, auto-detects login/CAPTCHA gates, waits for them to clear, then continues | Login-required, CAPTCHA-protected |\n| `--wait-for force` | Opens visible Chrome, auto-detects OR accepts Enter keypress to continue | Complex flows, lazy loading, paywalls |\n\n**Interaction gate auto-detection**:\n- Cloudflare Turnstile / \"just a moment\" pages\n- Google reCAPTCHA\n- hCaptcha\n- Custom challenge / verification screens\n\n**Wait-for-interaction workflow**:\n1. Run with `--wait-for interaction` → Chrome opens visibly\n2. CLI auto-detects login/CAPTCHA gates\n3. User completes login or solves CAPTCHA in the browser\n4. CLI auto-detects gate cleared → captures page\n5. If `--wait-for force` is used, user can also press Enter to trigger capture manually\n\n## Agent Quality Gate\n\n**CRITICAL**: The agent must treat default headless capture as provisional. Some sites render differently in headless mode and can silently return low-quality content without causing the CLI to fail.\n\nAfter every headless run, the agent **MUST** inspect the saved markdown output.\n\n### Quality checks the agent must perform\n\n1. Confirm the markdown title matches the target page, not a generic site shell\n2. Confirm the body contains the expected article or page content, not just navigation, footer, or a generic error\n3. Watch for obvious failure signs:\n   - `Application error`\n   - `This page could not be found`\n   - Login, signup, subscribe, or verification shells\n   - Extremely short markdown for a page that should be long-form\n   - Raw framework payloads or mostly boilerplate content\n4. If the result is low quality, incomplete, or clearly wrong, do **not** accept the run as successful just because the CLI exited with code 0\n\n**Tip**: Use `--format json` to get structured output including `status`, `login.state`, and `interaction` fields for programmatic quality assessment. A `\"status\": \"needs_interaction\"` response means the page requires manual interaction.\n\n### Recovery workflow the agent must follow\n\n1. First run headless (default) unless there is already a clear reason to use interaction mode\n2. Review markdown quality immediately after the run\n3. If the content is low quality or indicates login/CAPTCHA:\n   - `--wait-for interaction` for auto-detected gates (login, CAPTCHA, Cloudflare)\n   - `--wait-for force` when the page needs manual browsing, scroll loading, or complex interaction\n4. If `--wait-for` is used, tell the user exactly what to do:\n   - If login is required, ask them to sign in in the browser\n   - If CAPTCHA appears, ask them to solve it\n   - If the page needs time to load, ask them to wait until content is visible\n   - For `--wait-for force`: tell them to press Enter when ready\n5. If JSON output shows `\"status\": \"needs_interaction\"`, switch to `--wait-for interaction` automatically\n\n## Output Path Generation\n\nThe agent must construct the output file path since `baoyu-fetch` does not auto-generate paths.\n\n**Algorithm**:\n1. Determine base directory from EXTEND.md `default_output_dir` or default `./url-to-markdown/`\n2. Extract domain from URL (e.g., `example.com`)\n3. Generate slug from URL path or page title (kebab-case, 2-6 words)\n4. Construct: `{base_dir}/{domain}/{slug}/{slug}.md` — each URL gets its own directory so media files stay isolated\n5. Conflict resolution: append timestamp `{slug}-YYYYMMDD-HHMMSS/{slug}-YYYYMMDD-HHMMSS.md`\n\nPass the constructed path to `--output`. Media files (`--download-media`) are saved into subdirectories next to the markdown file, keeping each URL's assets self-contained.\n\n## Output Format\n\nMarkdown output to stdout (or file with `--output`) as clean markdown text.\n\nJSON output (`--format json`) returns structured data including:\n- `adapter` — which adapter handled the URL\n- `status` — `\"ok\"` or `\"needs_interaction\"`\n- `login` — login state detection (`logged_in`, `logged_out`, `unknown`)\n- `interaction` — interaction gate details (kind, provider, prompt)\n- `document` — structured content (url, title, author, publishedAt, content blocks, metadata)\n- `media` — collected media assets with url, kind, role\n- `markdown` — converted markdown text\n- `downloads` — media download results (when `--download-media` used)\n\nWhen `--download-media` is enabled:\n- Images are saved to `imgs/` next to the output file (or in `--media-dir`)\n- Videos are saved to `videos/` next to the output file (or in `--media-dir`)\n- Markdown media links are rewritten to local relative paths\n\n## Built-in Adapters\n\n| Adapter | URLs | Key Features |\n|---------|------|-------------|\n| `x` | x.com, twitter.com | Tweets, threads, X Articles, media, login detection |\n| `youtube` | youtube.com, youtu.be | Transcript/captions, chapters, cover image, metadata |\n| `hn` | news.ycombinator.com | Threaded comments, story metadata, nested replies |\n| `generic` | Any URL (fallback) | Defuddle extraction, Readability fallback, auto-scroll, network idle detection |\n\nAdapter is auto-selected based on URL. Use `--adapter <name>` to override.\n\n## Media Download Workflow\n\nBased on `download_media` setting in EXTEND.md:\n\n| Setting | Behavior |\n|---------|----------|\n| `1` (always) | Run CLI with `--download-media --output <path>` |\n| `0` (never) | Run CLI with `--output <path>` (no media download) |\n| `ask` (default) | Follow the ask-each-time flow below |\n\n### Ask-Each-Time Flow\n\n1. Run CLI **without** `--download-media` with `--output <path>` → markdown saved\n2. Check saved markdown for remote media URLs (`https://` in image/video links)\n3. **If no remote media found** → done, no prompt needed\n4. **If remote media found** → use `AskUserQuestion`:\n   - header: \"Media\", question: \"Download N images/videos to local files?\"\n   - \"Yes\" — Download to local directories\n   - \"No\" — Keep remote URLs\n5. If user confirms → run CLI **again** with `--download-media --output <same-path>` (overwrites markdown with localized links)\n\n## Environment Variables\n\n| Variable | Description |\n|----------|-------------|\n| `BAOYU_CHROME_PROFILE_DIR` | Chrome user data directory (can also use `--chrome-profile-dir`) |\n\n**Troubleshooting**: Chrome not found → use `--browser-path`. Timeout → increase `--timeout`. Login/CAPTCHA pages → use `--wait-for interaction`. Debug → use `--debug-dir` to inspect captured HTML and network logs.\n\n### YouTube Notes\n\n- YouTube adapter extracts transcripts/captions automatically when available\n- Transcript format: `[MM:SS] Text segment` with chapter headings\n- Transcript availability depends on YouTube exposing a caption track. Videos with captions disabled or restricted playback may produce description-only output\n- Use `--wait-for force` if the page needs time to finish loading player metadata\n\n### X/Twitter Notes\n\n- Extracts single tweets, threads, and X Articles\n- Auto-detects login state; if logged out and content requires auth, JSON output will show `\"status\": \"needs_interaction\"`\n- Use `--wait-for interaction` for login-protected content\n\n### Hacker News Notes\n\n- Parses threaded comments with proper nesting and reply hierarchy\n- Includes story metadata (title, URL, author, score, comment count)\n- Shows comment deletion/dead status\n\n## Extension Support\n\nCustom configurations via EXTEND.md. See **Preferences** section for paths and supported options.\n\nFile v1.103.1:scripts/vendor/baoyu-fetch/README.md\n\n# baoyu-fetch\n\nEnglish | [简体中文](./README.zh-CN.md) | [Changelog](./CHANGELOG.md) | [中文更新日志](./CHANGELOG.zh-CN.md)\n\n`baoyu-fetch` is a Bun CLI built on Chrome CDP. Give it a URL and it returns\nhigh-quality `markdown` or `json`. When a site adapter matches, it prefers API\nresponses or structured page data; otherwise it falls back to generic HTML\nextraction.\n\n## Features\n\n- Capture rendered page content through Chrome CDP\n- Observe network requests and responses, and fetch bodies when needed\n- Adapter registry that auto-selects a handler from the URL\n- Built-in adapters for `x`, `youtube`, and `hn`\n- Generic fallback: Defuddle first, then Readability + HTML-to-Markdown; when `--format markdown` is requested, it can also fall back to `defuddle.md`\n- Print `markdown` / `json` to stdout or save with `--output`\n- Optionally download extracted images or videos and rewrite Markdown links\n- Optional wait modes for login and verification flows\n- Chrome profile defaults to `baoyu-skills/chrome-profile`\n\n## Installation\n\n```bash\nbun install\n```\n\nFor package usage, the quickest option is:\n\n```bash\nbunx baoyu-fetch https://example.com\n```\n\nYou can also install it globally:\n\n```bash\nnpm install -g baoyu-fetch\n```\n\nThe npm package ships TypeScript source entrypoints instead of a prebuilt\n`dist`, so Bun is required at runtime.\n\n## Usage\n\n```bash\nbun run src/cli.ts https://example.com\nbunx baoyu-fetch https://example.com\nbaoyu-fetch https://example.com\nbaoyu-fetch https://example.com --format markdown --output article.md\nbaoyu-fetch https://example.com --format markdown --output article.md --download-media\nbaoyu-fetch https://x.com/jack/status/20 --format json --output article.json\nbaoyu-fetch https://x.com/jack/status/20 --json\nbaoyu-fetch https://x.com/jack/status/20 --wait-for interaction\nbaoyu-fetch https://x.com/jack/status/20 --wait-for force\nbaoyu-fetch https://x.com/jack/status/20 --chrome-profile-dir ~/Library/Application\\\\ Support/baoyu-skills/chrome-profile\n```\n\n## Options\n\n```bash\nbaoyu-fetch <url> [options]\n\nOptions:\n  --output <file>       Save output to file\n  --format <type>       Output format: markdown | json\n  --json                Alias for --format json\n  --adapter <name>      Force an adapter (for example x / hn / generic)\n  --download-media      Download adapter-reported media into ./imgs and ./videos, then rewrite markdown links\n  --media-dir <dir>     Base directory for downloaded media. Defaults to the output directory\n  --debug-dir <dir>     Write debug artifacts (html, document.json, network.json)\n  --cdp-url <url>       Reuse an existing Chrome DevTools endpoint\n  --browser-path <path> Explicit Chrome binary path\n  --chrome-profile-dir <path>\n                        Chrome user data dir. Defaults to BAOYU_CHROME_PROFILE_DIR\n                        or baoyu-skills/chrome-profile\n  --headless            Launch a temporary headless Chrome if needed\n  --wait-for <mode>     Wait mode: interaction | force\n  --wait-for-interaction\n                        Alias for --wait-for interaction\n  --wait-for-login      Alias for --wait-for interaction\n  --interaction-timeout <ms>\n                        Manual interaction timeout. Default: 600000\n  --interaction-poll-interval <ms>\n                        Poll interval while waiting. Default: 1500\n  --login-timeout <ms>  Alias for --interaction-timeout\n  --login-poll-interval <ms>\n                        Alias for --interaction-poll-interval\n  --timeout <ms>        Page load timeout. Default: 30000\n  --help                Show help\n```\n\n## How It Works\n\n1. The CLI parses the target URL and options.\n2. It opens or connects to a Chrome CDP session and creates a controlled tab.\n3. `NetworkJournal` records requests and responses.\n4. The adapter registry resolves a site-specific adapter when possible.\n5. The adapter returns a structured `ExtractedDocument`.\n6. If nothing matches, generic HTML extraction runs instead.\n7. The result is rendered as Markdown, or returned as JSON with both\n   `document` and `markdown`.\n\n## Development\n\n```bash\nbun run check\nbun run test\nbun run build\n```\n\n## Release\n\nWhen you make a user-visible change, add a changeset first:\n\n```bash\nbunx changeset\n```\n\nAfter the generated `.changeset/*.md` file lands on `main`, GitHub Actions will\nopen or update the release PR. Merging that release PR publishes the package to\nnpm.\n\nThe publish flow does not build `dist`; it publishes `src/*.ts` for Bun\nexecution directly.\n\nFile v1.103.1:_meta.json\n\n{\n  \"ownerId\": \"kn7csrrndw79hpke5d0gsnx93d82k67r\",\n  \"slug\": \"baoyu-url-to-markdown\",\n  \"version\": \"1.103.1\",\n  \"publishedAt\": 1776097095080\n}\n\nFile v1.103.1:references/config/first-time-setup.md\n\n---\nname: first-time-setup\ndescription: First-time setup flow for baoyu-url-to-markdown preferences\n---\n\n# First-Time Setup\n\n## Overview\n\nWhen no EXTEND.md is found, guide user through preference setup.\n\n**BLOCKING OPERATION**: This setup MUST complete before ANY other workflow steps. Do NOT:\n- Start converting URLs\n- Ask about URLs or output paths\n- Proceed to any conversion\n\nONLY ask the questions in this setup flow, save EXTEND.md, then continue.\n\n## Setup Flow\n\n```\nNo EXTEND.md found\n        |\n        v\n+---------------------+\n| AskUserQuestion     |\n| (all questions)     |\n+---------------------+\n        |\n        v\n+---------------------+\n| Create EXTEND.md    |\n+---------------------+\n        |\n        v\n    Continue conversion\n```\n\n## Questions\n\n**Language**: Use user's input language or saved language preference.\n\nUse AskUserQuestion with ALL questions in ONE call:\n\n### Question 1: Download Media\n\n```yaml\nheader: \"Media\"\nquestion: \"How to handle images and videos in pages?\"\noptions:\n  - label: \"Ask each time (Recommended)\"\n    description: \"After saving markdown, ask whether to download media\"\n  - label: \"Always download\"\n    description: \"Always download media to local imgs/ and videos/ directories\"\n  - label: \"Never download\"\n    description: \"Keep original remote URLs in markdown\"\n```\n\n### Question 2: Default Output Directory\n\n```yaml\nheader: \"Output\"\nquestion: \"Default output directory?\"\noptions:\n  - label: \"url-to-markdown (Recommended)\"\n    description: \"Save to ./url-to-markdown/{domain}/{slug}.md\"\n```\n\nNote: User will likely choose \"Other\" to type a custom path.\n\n### Question 3: Save Location\n\n```yaml\nheader: \"Save\"\nquestion: \"Where to save preferences?\"\noptions:\n  - label: \"User (Recommended)\"\n    description: \"~/.baoyu-skills/ (all projects)\"\n  - label: \"Project\"\n    description: \".baoyu-skills/ (this project only)\"\n```\n\n## Save Locations\n\n| Choice | Path | Scope |\n|--------|------|-------|\n| User | `~/.baoyu-skills/baoyu-url-to-markdown/EXTEND.md` | All projects |\n| Project | `.baoyu-skills/baoyu-url-to-markdown/EXTEND.md` | Current project |\n\n## After Setup\n\n1. Create directory if needed\n2. Write EXTEND.md\n3. Confirm: \"Preferences saved to [path]\"\n4. Continue with conversion using saved preferences\n\n## EXTEND.md Template\n\n```md\ndownload_media: [ask/1/0]\ndefault_output_dir: [path or empty]\n```\n\n## Modifying Preferences Later\n\nUsers can edit EXTEND.md directly or delete it to trigger setup again.\n\nFile v1.103.1:scripts/package.json\n\n{\n  \"name\": \"baoyu-url-to-markdown-scripts\",\n  \"private\": true,\n  \"type\": \"module\",\n  \"dependencies\": {\n    \"baoyu-fetch\": \"file:./vendor/baoyu-fetch\"\n  }\n}\n\nFile v1.103.1:scripts/vendor/baoyu-fetch/package.json\n\n{\n  \"name\": \"baoyu-fetch\",\n  \"version\": \"0.1.1\",\n  \"description\": \"Read URLs into high-quality Markdown or JSON with Chrome CDP and site adapters.\",\n  \"type\": \"module\",\n  \"bin\": {\n    \"baoyu-fetch\": \"./src/cli.ts\"\n  },\n  \"files\": [\n    \"README.zh-CN.md\",\n    \"src/adapters\",\n    \"src/browser\",\n    \"src/cli.ts\",\n    \"src/commands\",\n    \"src/extract\",\n    \"src/media\",\n    \"src/types\",\n    \"src/utils\",\n    \"README.md\"\n  ],\n  \"repository\": {\n    \"type\": \"git\",\n    \"url\": \"git+https://github.com/JimLiu/baoyu-skills.git\",\n    \"directory\": \"packages/baoyu-fetch\"\n  },\n  \"bugs\": {\n    \"url\": \"https://github.com/JimLiu/baoyu-skills/issues\"\n  },\n  \"homepage\": \"https://github.com/JimLiu/baoyu-skills/tree/main/packages/baoyu-fetch#readme\",\n  \"publishConfig\": {\n    \"access\": \"public\"\n  },\n  \"scripts\": {\n    \"build\": \"rm -rf dist && bun build ./src/cli.ts --target bun --outfile ./dist/cli.js && chmod +x ./dist/cli.js\",\n    \"check\": \"tsc --noEmit\",\n    \"dev\": \"bun run ./src/cli.ts\",\n    \"release\": \"changeset publish\",\n    \"test\": \"bun test\",\n    \"version-packages\": \"changeset version\"\n  },\n  \"engines\": {\n    \"bun\": \">=1.2.0\"\n  },\n  \"dependencies\": {\n    \"@mozilla/readability\": \"^0.6.0\",\n    \"chrome-launcher\": \"^1.2.1\",\n    \"defuddle\": \"^0.14.0\",\n    \"jsdom\": \"^26.0.0\",\n    \"remark-gfm\": \"^4.0.1\",\n    \"remark-parse\": \"^11.0.0\",\n    \"turndown\": \"^7.2.0\",\n    \"turndown-plugin-gfm\": \"^1.0.2\",\n    \"unified\": \"^11.0.5\",\n    \"ws\": \"^8.18.3\"\n  },\n  \"devDependencies\": {\n    \"@changesets/cli\": \"^2.30.0\",\n    \"@types/bun\": \"^1.2.23\",\n    \"@types/jsdom\": \"^21.1.7\",\n    \"@types/ws\": \"^8.18.1\",\n    \"typescript\": \"^5.9.2\"\n  }\n}\n\nFile v1.103.1:scripts/vendor/baoyu-fetch/README.zh-CN.md\n\n# baoyu-fetch\n\n[English](./README.md) | 简体中文 | [更新日志](./CHANGELOG.zh-CN.md) | [English Changelog](./CHANGELOG.md)\n\n`baoyu-fetch` 是一个基于 Chrome CDP 的 Bun CLI。输入 URL，它会输出高质量\n`markdown` 或 `json`；命中站点 adapter 时优先消费 API 返回或页面内结构化\n数据，未命中时回退到通用 HTML 提取。\n\n## 当前能力\n\n- 通过 Chrome CDP 抓取渲染后的页面内容\n- 监听网络请求与响应，按需拉取响应体\n- adapter registry，支持按 URL 自动命中站点处理器\n- 内置 `x`、`youtube`、`hn` adapters\n- 通用 fallback：Defuddle 优先，Readability + HTML to Markdown 回退；`--format markdown` 时会再尝试 `defuddle.md` 兜底\n- `stdout` 或 `--output` 输出 `markdown` / `json`\n- 可选下载提取出的图片/视频并重写 Markdown 链接\n- 提供登录/验证场景下的交互等待模式\n- Chrome profile 默认对齐 `baoyu-skills/chrome-profile`\n\n## 安装\n\n```bash\nbun install\n```\n\n作为包使用时，推荐直接这样运行：\n\n```bash\nbunx baoyu-fetch https://example.com\n```\n\n也可以全局安装：\n\n```bash\nnpm install -g baoyu-fetch\n```\n\nnpm 包发布的是 TypeScript 源码入口，不包含预编译的 `dist`，所以运行时需要\nBun。\n\n## 用法\n\n```bash\nbun run src/cli.ts https://example.com\nbunx baoyu-fetch https://example.com\nbaoyu-fetch https://example.com\nbaoyu-fetch https://example.com --format markdown --output article.md\nbaoyu-fetch https://example.com --format markdown --output article.md --download-media\nbaoyu-fetch https://x.com/jack/status/20 --format json --output article.json\nbaoyu-fetch https://x.com/jack/status/20 --json\nbaoyu-fetch https://x.com/jack/status/20 --wait-for interaction\nbaoyu-fetch https://x.com/jack/status/20 --wait-for force\nbaoyu-fetch https://x.com/jack/status/20 --chrome-profile-dir ~/Library/Application\\\\ Support/baoyu-skills/chrome-profile\n```\n\n## 主要参数\n\n```bash\nbaoyu-fetch <url> [options]\n\nOptions:\n  --output <file>       保存输出内容到文件\n  --format <type>       输出格式：markdown | json\n  --json                `--format json` 的兼容别名\n  --adapter <name>      强制使用指定 adapter（如 x / hn / generic）\n  --download-media      下载 adapter 返回的媒体到 ./imgs 和 ./videos，并重写 markdown 链接\n  --media-dir <dir>     指定媒体下载根目录；默认使用输出文件所在目录\n  --debug-dir <dir>     导出调试信息（html、document.json、network.json）\n  --cdp-url <url>       连接现有 Chrome 调试地址\n  --browser-path <path> 指定 Chrome 可执行文件\n  --chrome-profile-dir <path>\n                        指定 Chrome profile 目录。默认使用 BAOYU_CHROME_PROFILE_DIR，\n                        否则回退到 baoyu-skills/chrome-profile\n  --headless            启动临时 headless Chrome（未连现有实例时）\n  --wait-for <mode>     等待模式：interaction | force\n  --wait-for-interaction\n                        `--wait-for interaction` 的别名\n  --wait-for-login      `--wait-for interaction` 的别名\n  --interaction-timeout <ms>\n                        手动交互等待超时，默认 600000\n  --interaction-poll-interval <ms>\n                        等待期间的轮询间隔，默认 1500\n  --login-timeout <ms>  `--interaction-timeout` 的别名\n  --login-poll-interval <ms>\n                        `--interaction-poll-interval` 的别名\n  --timeout <ms>        页面加载超时，默认 30000\n  --help                显示帮助\n```\n\n## 设计\n\n核心链路：\n\n1. CLI 解析 URL 和选项\n2. 建立 CDP 会话并创建受控 tab\n3. 启动 `NetworkJournal` 收集所有请求/响应\n4. 由 adapter registry 匹配站点 adapter\n5. adapter 返回结构化 `ExtractedDocument`\n6. 没命中则走通用 HTML 提取\n7. 按请求输出 Markdown，或输出包含 `document` 和 `markdown` 的 JSON\n\n## 开发\n\n```bash\nbun run check\nbun run test\nbun run build\n```\n\n## 发版\n\n新增用户可见改动后，先添加一个 changeset：\n\n```bash\nbunx changeset\n```\n\n把生成的 `.changeset/*.md` 一起合并到 `main` 后，GitHub Actions 会自动创建或\n更新 release PR；合并 release PR 之后，会自动发布到 npm。\n\n发布流程不会编译 `dist`，而是直接把 `src/*.ts` 发布出去供 Bun 执行。\n\nArchive v1.103.0: 49 files, 110029 bytes\n\nFiles: references/config/first-time-setup.md (2459b), scripts/bun.lock (56924b), scripts/package.json (157b), scripts/vendor/baoyu-fetch/package.json (1624b), scripts/vendor/baoyu-fetch/README.md (4477b), scripts/vendor/baoyu-fetch/README.zh-CN.md (4326b), scripts/vendor/baoyu-fetch/src/adapters/generic/index.ts (2451b), scripts/vendor/baoyu-fetch/src/adapters/hn/index.ts (10454b), scripts/vendor/baoyu-fetch/src/adapters/index.ts (834b), scripts/vendor/baoyu-fetch/src/adapters/types.ts (2100b), scripts/vendor/baoyu-fetch/src/adapters/x/article.ts (12122b), scripts/vendor/baoyu-fetch/src/adapters/x/index.ts (4048b), scripts/vendor/baoyu-fetch/src/adapters/x/login.ts (2258b), scripts/vendor/baoyu-fetch/src/adapters/x/match.ts (297b), scripts/vendor/baoyu-fetch/src/adapters/x/payloads.ts (1681b), scripts/vendor/baoyu-fetch/src/adapters/x/session.ts (1404b), scripts/vendor/baoyu-fetch/src/adapters/x/shared.ts (11727b), scripts/vendor/baoyu-fetch/src/adapters/x/single.ts (2583b), scripts/vendor/baoyu-fetch/src/adapters/x/thread-loader.ts (9078b), scripts/vendor/baoyu-fetch/src/adapters/x/thread.ts (9196b), scripts/vendor/baoyu-fetch/src/adapters/x/types.ts (600b), scripts/vendor/baoyu-fetch/src/adapters/youtube/index.ts (898b), scripts/vendor/baoyu-fetch/src/adapters/youtube/transcript.ts (13301b), scripts/vendor/baoyu-fetch/src/adapters/youtube/utils.ts (6523b), scripts/vendor/baoyu-fetch/src/browser/cdp-client.ts (6879b), scripts/vendor/baoyu-fetch/src/browser/chrome-launcher.ts (5883b), scripts/vendor/baoyu-fetch/src/browser/cookie-sidecar.ts (2941b), scripts/vendor/baoyu-fetch/src/browser/interaction-gates.ts (4095b), scripts/vendor/baoyu-fetch/src/browser/network-journal.ts (7153b), scripts/vendor/baoyu-fetch/src/browser/page-snapshot.ts (3429b), scripts/vendor/baoyu-fetch/src/browser/profile.ts (5670b), scripts/vendor/baoyu-fetch/src/browser/session.ts (4642b), scripts/vendor/baoyu-fetch/src/cli.ts (7066b), scripts/vendor/baoyu-fetch/src/commands/convert.ts (17456b), scripts/vendor/baoyu-fetch/src/extract/document.ts (846b), scripts/vendor/baoyu-fetch/src/extract/html-cleaner.ts (11016b), scripts/vendor/baoyu-fetch/src/extract/html-extractor.ts (2243b), scripts/vendor/baoyu-fetch/src/extract/html-to-markdown.ts (21344b), scripts/vendor/baoyu-fetch/src/extract/markdown-renderer.ts (4661b), scripts/vendor/baoyu-fetch/src/media/default-downloader.ts (5215b), scripts/vendor/baoyu-fetch/src/media/markdown-media.ts (13186b), scripts/vendor/baoyu-fetch/src/media/media-utils.ts (6222b), scripts/vendor/baoyu-fetch/src/media/types.ts (710b), scripts/vendor/baoyu-fetch/src/types/defuddle-node.d.ts (605b), scripts/vendor/baoyu-fetch/src/types/shims.d.ts (422b), scripts/vendor/baoyu-fetch/src/utils/logger.ts (645b), scripts/vendor/baoyu-fetch/src/utils/url.ts (300b), SKILL.md (15269b), _meta.json (142b)\n\nFile v1.103.0:SKILL.md\n\n---\nname: baoyu-url-to-markdown\ndescription: Fetch any URL and convert to markdown using baoyu-fetch CLI (Chrome CDP with site-specific adapters). Built-in adapters for X/Twitter, YouTube transcripts, Hacker News threads, and generic pages via Defuddle. Handles login/CAPTCHA via interaction wait modes. Use when user wants to save a webpage as markdown.\nversion: 1.60.0\nmetadata:\n  openclaw:\n    homepage: https://github.com/JimLiu/baoyu-skills#baoyu-url-to-markdown\n    requires:\n      anyBins:\n        - bun\n        - npx\n---\n\n# URL to Markdown\n\nFetches any URL via `baoyu-fetch` CLI (Chrome CDP + site-specific adapters) and converts it to clean markdown.\n\n## CLI Setup\n\n**Important**: The CLI source is vendored in the `scripts/vendor/baoyu-fetch/` subdirectory of this skill.\n\n**Agent Execution Instructions**:\n1. Determine this SKILL.md file's directory path as `{baseDir}`\n2. CLI entry point = `{baseDir}/scripts/vendor/baoyu-fetch/src/cli.ts`\n3. Resolve `${BUN_X}` runtime: if `bun` installed → `bun`; if `npx` available → `npx -y bun`; else suggest installing bun\n4. `${READER}` = `${BUN_X} {baseDir}/scripts/vendor/baoyu-fetch/src/cli.ts`\n5. Replace all `${READER}` in this document with the resolved value\n\n## Preferences (EXTEND.md)\n\nCheck EXTEND.md existence (priority order):\n\n```bash\n# macOS, Linux, WSL, Git Bash\ntest -f .baoyu-skills/baoyu-url-to-markdown/EXTEND.md && echo \"project\"\ntest -f \"${XDG_CONFIG_HOME:-$HOME/.config}/baoyu-skills/baoyu-url-to-markdown/EXTEND.md\" && echo \"xdg\"\ntest -f \"$HOME/.baoyu-skills/baoyu-url-to-markdown/EXTEND.md\" && echo \"user\"\n```\n\n```powershell\n# PowerShell (Windows)\nif (Test-Path .baoyu-skills/baoyu-url-to-markdown/EXTEND.md) { \"project\" }\n$xdg = if ($env:XDG_CONFIG_HOME) { $env:XDG_CONFIG_HOME } else { \"$HOME/.config\" }\nif (Test-Path \"$xdg/baoyu-skills/baoyu-url-to-markdown/EXTEND.md\") { \"xdg\" }\nif (Test-Path \"$HOME/.baoyu-skills/baoyu-url-to-markdown/EXTEND.md\") { \"user\" }\n```\n\n| Path | Location |\n|------|----------|\n| `.baoyu-skills/baoyu-url-to-markdown/EXTEND.md` | Project directory |\n| `$HOME/.baoyu-skills/baoyu-url-to-markdown/EXTEND.md` | User home |\n\n| Result | Action |\n|--------|--------|\n| Found | Read, parse, apply settings |\n| Not found | **MUST** run first-time setup (see below) — do NOT silently create defaults |\n\n**EXTEND.md Supports**: Download media by default | Default output directory\n\n### First-Time Setup (BLOCKING)\n\n**CRITICAL**: When EXTEND.md is not found, you **MUST use `AskUserQuestion`** to ask the user for their preferences before creating EXTEND.md. **NEVER** create EXTEND.md with defaults without asking. This is a **BLOCKING** operation — do NOT proceed with any conversion until setup is complete.\n\nUse `AskUserQuestion` with ALL questions in ONE call:\n\n**Question 1** — header: \"Media\", question: \"How to handle images and videos in pages?\"\n- \"Ask each time (Recommended)\" — After saving markdown, ask whether to download media\n- \"Always download\" — Always download media to local imgs/ and videos/ directories\n- \"Never download\" — Keep original remote URLs in markdown\n\n**Question 2** — header: \"Output\", question: \"Default output directory?\"\n- \"url-to-markdown (Recommended)\" — Save to ./url-to-markdown/{domain}/{slug}.md\n- (User may choose \"Other\" to type a custom path)\n\n**Question 3** — header: \"Save\", question: \"Where to save preferences?\"\n- \"User (Recommended)\" — ~/.baoyu-skills/ (all projects)\n- \"Project\" — .baoyu-skills/ (this project only)\n\nAfter user answers, create EXTEND.md at the chosen location, confirm \"Preferences saved to [path]\", then continue.\n\nFull reference: [references/config/first-time-setup.md](references/config/first-time-setup.md)\n\n### Supported Keys\n\n| Key | Default | Values | Description |\n|-----|---------|--------|-------------|\n| `download_media` | `ask` | `ask` / `1` / `0` | `ask` = prompt each time, `1` = always download, `0` = never |\n| `default_output_dir` | empty | path or empty | Default output directory (empty = `./url-to-markdown/`) |\n\n**EXTEND.md → CLI mapping**:\n| EXTEND.md key | CLI argument | Notes |\n|---------------|-------------|-------|\n| `download_media: 1` | `--download-media` | Requires `--output` to be set |\n| `default_output_dir: ./posts/` | Agent constructs `--output ./posts/{domain}/{slug}.md` | Agent generates path, not a direct CLI flag |\n\n**Value priority**:\n1. CLI arguments (`--download-media`, `--output`)\n2. EXTEND.md\n3. Skill defaults\n\n## Features\n\n- Chrome CDP for full JavaScript rendering via `baoyu-fetch` CLI\n- Site-specific adapters: X/Twitter, YouTube, Hacker News, generic (Defuddle)\n- Automatic adapter selection based on URL, or force with `--adapter`\n- Interaction gate detection: Cloudflare, reCAPTCHA, hCAPTCHA, custom challenges\n- Two capture modes: headless (default) or interactive with wait-for-interaction\n- Clean markdown output with YAML front matter\n- Structured JSON output available via `--format json`\n- X/Twitter: extracts tweets, threads, and X Articles with media\n- YouTube: transcript/caption extraction, chapters, cover images\n- Hacker News: threaded comment parsing with proper nesting\n- Generic: Defuddle extraction with Readability fallback\n- Download images and videos to local directories\n- Chrome profile persistence for authenticated sessions\n- Debug artifact output for troubleshooting\n\n## Usage\n\n```bash\n# Default: headless capture, markdown to stdout\n${READER} <url>\n\n# Save to file\n${READER} <url> --output article.md\n\n# Save with media download\n${READER} <url> --output article.md --download-media\n\n# Headless mode (explicit)\n${READER} <url> --headless --output article.md\n\n# Wait for interaction (login/CAPTCHA) — auto-detect and continue\n${READER} <url> --wait-for interaction --output article.md\n\n# Wait for interaction — manual control (Enter to continue)\n${READER} <url> --wait-for force --output article.md\n\n# JSON output\n${READER} <url> --format json --output article.json\n\n# Force specific adapter\n${READER} <url> --adapter youtube --output transcript.md\n\n# Connect to existing Chrome\n${READER} <url> --cdp-url http://localhost:9222 --output article.md\n\n# Debug artifacts\n${READER} <url> --output article.md --debug-dir ./debug/\n```\n\n## Options\n\n| Option | Description |\n|--------|-------------|\n| `<url>` | URL to fetch |\n| `--output <path>` | Output file path (default: stdout) |\n| `--format <type>` | Output format: `markdown` (default) or `json` |\n| `--json` | Shorthand for `--format json` |\n| `--adapter <name>` | Force adapter: `x`, `youtube`, `hn`, or `generic` (default: auto-detect) |\n| `--headless` | Force headless Chrome (no visible window) |\n| `--wait-for <mode>` | Interaction wait mode: `none` (default), `interaction`, or `force` |\n| `--wait-for-interaction` | Alias for `--wait-for interaction` |\n| `--wait-for-login` | Alias for `--wait-for interaction` |\n| `--timeout <ms>` | Page load timeout (default: 30000) |\n| `--interaction-timeout <ms>` | Login/CAPTCHA wait timeout (default: 600000 = 10 min) |\n| `--interaction-poll-interval <ms>` | Poll interval for interaction checks (default: 1500) |\n| `--download-media` | Download images/videos to local `imgs/` and `videos/`, rewrite markdown links. Requires `--output` |\n| `--media-dir <dir>` | Base directory for downloaded media (default: same as `--output` directory) |\n| `--cdp-url <url>` | Reuse existing Chrome DevTools Protocol endpoint |\n| `--browser-path <path>` | Custom Chrome/Chromium binary path |\n| `--chrome-profile-dir <path>` | Chrome user data directory (default: `BAOYU_CHROME_PROFILE_DIR` env or `./baoyu-skills/chrome-profile`) |\n| `--debug-dir <dir>` | Write debug artifacts (document.json, markdown.md, page.html, network.json) |\n\n## Capture Modes\n\n| Mode | Behavior | Use When |\n|------|----------|----------|\n| Default | Headless Chrome, auto-extract on network idle | Public pages, static content |\n| `--headless` | Explicit headless (same as default) | Clarify intent |\n| `--wait-for interaction` | Opens visible Chrome, auto-detects login/CAPTCHA gates, waits for them to clear, then continues | Login-required, CAPTCHA-protected |\n| `--wait-for force` | Opens visible Chrome, auto-detects OR accepts Enter keypress to continue | Complex flows, lazy loading, paywalls |\n\n**Interaction gate auto-detection**:\n- Cloudflare Turnstile / \"just a moment\" pages\n- Google reCAPTCHA\n- hCaptcha\n- Custom challenge / verification screens\n\n**Wait-for-interaction workflow**:\n1. Run with `--wait-for interaction` → Chrome opens visibly\n2. CLI auto-detects login/CAPTCHA gates\n3. User completes login or solves CAPTCHA in the browser\n4. CLI auto-detects gate cleared → captures page\n5. If `--wait-for force` is used, user can also press Enter to trigger capture manually\n\n## Agent Quality Gate\n\n**CRITICAL**: The agent must treat default headless capture as provisional. Some sites render differently in headless mode and can silently return low-quality content without causing the CLI to fail.\n\nAfter every headless run, the agent **MUST** inspect the saved markdown output.\n\n### Quality checks the agent must perform\n\n1. Confirm the markdown title matches the target page, not a generic site shell\n2. Confirm the body contains the expected article or page content, not just navigation, footer, or a generic error\n3. Watch for obvious failure signs:\n   - `Application error`\n   - `This page could not be found`\n   - Login, signup, subscribe, or verification shells\n   - Extremely short markdown for a page that should be long-form\n   - Raw framework payloads or mostly boilerplate content\n4. If the result is low quality, incomplete, or clearly wrong, do **not** accept the run as successful just because the CLI exited with code 0\n\n**Tip**: Use `--format json` to get structured output including `status`, `login.state`, and `interaction` fields for programmatic quality assessment. A `\"status\": \"needs_interaction\"` response means the page requires manual interaction.\n\n### Recovery workflow the agent must follow\n\n1. First run headless (default) unless there is already a clear reason to use interaction mode\n2. Review markdown quality immediately after the run\n3. If the content is low quality or indicates login/CAPTCHA:\n   - `--wait-for interaction` for auto-detected gates (login, CAPTCHA, Cloudflare)\n   - `--wait-for force` when the page needs manual browsing, scroll loading, or complex interaction\n4. If `--wait-for` is used, tell the user exactly what to do:\n   - If login is required, ask them to sign in in the browser\n   - If CAPTCHA appears, ask them to solve it\n   - If the page needs time to load, ask them to wait until content is visible\n   - For `--wait-for force`: tell them to press Enter when ready\n5. If JSON output shows `\"status\": \"needs_interaction\"`, switch to `--wait-for interaction` automatically\n\n## Output Path Generation\n\nThe agent must construct the output file path since `baoyu-fetch` does not auto-generate paths.\n\n**Algorithm**:\n1. Determine base directory from EXTEND.md `default_output_dir` or default `./url-to-markdown/`\n2. Extract domain from URL (e.g., `example.com`)\n3. Generate slug from URL path or page title (kebab-case, 2-6 words)\n4. Construct: `{base_dir}/{domain}/{slug}/{slug}.md` — each URL gets its own directory so media files stay isolated\n5. Conflict resolution: append timestamp `{slug}-YYYYMMDD-HHMMSS/{slug}-YYYYMMDD-HHMMSS.md`\n\nPass the constructed path to `--output`. Media files (`--download-media`) are saved into subdirectories next to the markdown file, keeping each URL's assets self-contained.\n\n## Output Format\n\nMarkdown output to stdout (or file with `--output`) as clean markdown text.\n\nJSON output (`--format json`) returns structured data including:\n- `adapter` — which adapter handled the URL\n- `status` — `\"ok\"` or `\"needs_interaction\"`\n- `login` — login state detection (`logged_in`, `logged_out`, `unknown`)\n- `interaction` — interaction gate details (kind, provider, prompt)\n- `document` — structured content (url, title, author, publishedAt, content blocks, metadata)\n- `media` — collected media assets with url, kind, role\n- `markdown` — converted markdown text\n- `downloads` — media download results (when `--download-media` used)\n\nWhen `--download-media` is enabled:\n- Images are saved to `imgs/` next to the output file (or in `--media-dir`)\n- Videos are saved to `videos/` next to the output file (or in `--media-dir`)\n- Markdown media links are rewritten to local relative paths\n\n## Built-in Adapters\n\n| Adapter | URLs | Key Features |\n|---------|------|-------------|\n| `x` | x.com, twitter.com | Tweets, threads, X Articles, media, login detection |\n| `youtube` | youtube.com, youtu.be | Transcript/captions, chapters, cover image, metadata |\n| `hn` | news.ycombinator.com | Threaded comments, story metadata, nested replies |\n| `generic` | Any URL (fallback) | Defuddle extraction, Readability fallback, auto-scroll, network idle detection |\n\nAdapter is auto-selected based on URL. Use `--adapter <name>` to override.\n\n## Media Download Workflow\n\nBased on `download_media` setting in EXTEND.md:\n\n| Setting | Behavior |\n|---------|----------|\n| `1` (always) | Run CLI with `--download-media --output <path>` |\n| `0` (never) | Run CLI with `--output <path>` (no media download) |\n| `ask` (default) | Follow the ask-each-time flow below |\n\n### Ask-Each-Time Flow\n\n1. Run CLI **without** `--download-media` with `--output <path>` → markdown saved\n2. Check saved markdown for remote media URLs (`https://` in image/video links)\n3. **If no remote media found** → done, no prompt needed\n4. **If remote media found** → use `AskUserQuestion`:\n   - header: \"Media\", question: \"Download N images/videos to local files?\"\n   - \"Yes\" — Download to local directories\n   - \"No\" — Keep remote URLs\n5. If user confirms → run CLI **again** with `--download-media --output <same-path>` (overwrites markdown with localized links)\n\n## Environment Variables\n\n| Variable | Description |\n|----------|-------------|\n| `BAOYU_CHROME_PROFILE_DIR` | Chrome user data directory (can also use `--chrome-profile-dir`) |\n\n**Troubleshooting**: Chrome not found → use `--browser-path`. Timeout → increase `--timeout`. Login/CAPTCHA pages → use `--wait-for interaction`. Debug → use `--debug-dir` to inspect captured HTML and network logs.\n\n### YouTube Notes\n\n- YouTube adapter extracts transcripts/captions automatically when available\n- Transcript format: `[MM:SS] Text segment` with chapter headings\n- Transcript availability depends on YouTube exposing a caption track. Videos with captions disabled or restricted playback may produce description-only output\n- Use `--wait-for force` if the page needs time to finish loading player metadata\n\n### X/Twitter Notes\n\n- Extracts single tweets, threads, and X Articles\n- Auto-detects login state; if logged out and content requires auth, JSON output will show `\"status\": \"needs_interaction\"`\n- Use `--wait-for interaction` for login-protected content\n\n### Hacker News Notes\n\n- Parses threaded comments with proper nesting and reply hierarchy\n- Includes story metadata (title, URL, author, score, comment count)\n- Shows comment deletion/dead status\n\n## Extension Support\n\nCustom configurations via EXTEND.md. See **Preferences** section for paths and supported options.\n\nFile v1.103.0:scripts/vendor/baoyu-fetch/README.md\n\n# baoyu-fetch\n\nEnglish | [简体中文](./README.zh-CN.md) | [Changelog](./CHANGELOG.md) | [中文更新日志](./CHANGELOG.zh-CN.md)\n\n`baoyu-fetch` is a Bun CLI built on Chrome CDP. Give it a URL and it returns\nhigh-quality `markdown` or `json`. When a site adapter matches, it prefers API\nresponses or structured page data; otherwise it falls back to generic HTML\nextraction.\n\n## Features\n\n- Capture rendered page content through Chrome CDP\n- Observe network requests and responses, and fetch bodies when needed\n- Adapter registry that auto-selects a handler from the URL\n- Built-in adapters for `x`, `youtube`, and `hn`\n- Generic fallback: Defuddle first, then Readability + HTML-to-Markdown; when `--format markdown` is requested, it can also fall back to `defuddle.md`\n- Print `markdown` / `json` to stdout or save with `--output`\n- Optionally download extracted images or videos and rewrite Markdown links\n- Optional wait modes for login and verification flows\n- Chrome profile defaults to `baoyu-skills/chrome-profile`\n\n## Installation\n\n```bash\nbun install\n```\n\nFor package usage, the quickest option is:\n\n```bash\nbunx baoyu-fetch https://example.com\n```\n\nYou can also install it globally:\n\n```bash\nnpm install -g baoyu-fetch\n```\n\nThe npm package ships TypeScript source entrypoints instead of a prebuilt\n`dist`, so Bun is required at runtime.\n\n## Usage\n\n```bash\nbun run src/cli.ts https://example.com\nbunx baoyu-fetch https://example.com\nbaoyu-fetch https://example.com\nbaoyu-fetch https://example.com --format markdown --output article.md\nbaoyu-fetch https://example.com --format markdown --output article.md --download-media\nbaoyu-fetch https://x.com/jack/status/20 --format json --output article.json\nbaoyu-fetch https://x.com/jack/status/20 --json\nbaoyu-fetch https://x.com/jack/status/20 --wait-for interaction\nbaoyu-fetch https://x.com/jack/status/20 --wait-for force\nbaoyu-fetch https://x.com/jack/status/20 --chrome-profile-dir ~/Library/Application\\\\ Support/baoyu-skills/chrome-profile\n```\n\n## Options\n\n```bash\nbaoyu-fetch <url> [options]\n\nOptions:\n  --output <file>       Save output to file\n  --format <type>       Output format: markdown | json\n  --json                Alias for --format json\n  --adapter <name>      Force an adapter (for example x / hn / generic)\n  --download-media      Download adapter-reported media into ./imgs and ./videos, then rewrite markdown links\n  --media-dir <dir>     Base directory for downloaded media. Defaults to the output directory\n  --debug-dir <dir>     Write debug artifacts (html, document.json, network.json)\n  --cdp-url <url>       Reuse an existing Chrome DevTools endpoint\n  --browser-path <path> Explicit Chrome binary path\n  --chrome-profile-dir <path>\n                        Chrome user data dir. Defaults to BAOYU_CHROME_PROFILE_DIR\n                        or baoyu-skills/chrome-profile\n  --headless            Launch a temporary headless Chrome if needed\n  --wait-for <mode>     Wait mode: interaction | force\n  --wait-for-interaction\n                        Alias for --wait-for interaction\n  --wait-for-login      Alias for --wait-for interaction\n  --interaction-timeout <ms>\n                        Manual interaction timeout. Default: 600000\n  --interaction-poll-interval <ms>\n                        Poll interval while waiting. Default: 1500\n  --login-timeout <ms>  Alias for --interaction-timeout\n  --login-poll-interval <ms>\n                        Alias for --interaction-poll-interval\n  --timeout <ms>        Page load timeout. Default: 30000\n  --help                Show help\n```\n\n## How It Works\n\n1. The CLI parses the target URL and options.\n2. It opens or connects to a Chrome CDP session and creates a controlled tab.\n3. `NetworkJournal` records requests and responses.\n4. The adapter registry resolves a site-specific adapter when possible.\n5. The adapter returns a structured `ExtractedDocument`.\n6. If nothing matches, generic HTML extraction runs instead.\n7. The result is rendered as Markdown, or returned as JSON with both\n   `document` and `markdown`.\n\n## Development\n\n```bash\nbun run check\nbun run test\nbun run build\n```\n\n## Release\n\nWhen you make a user-visible change, add a changeset first:\n\n```bash\nbunx changeset\n```\n\nAfter the generated `.changeset/*.md` file lands on `main`, GitHub Actions will\nopen or update the release PR. Merging that release PR publishes the package to\nnpm.\n\nThe publish flow does not build `dist`; it publishes `src/*.ts` for Bun\nexecution directly.\n\nFile v1.103.0:_meta.json\n\n{\n  \"ownerId\": \"kn7csrrndw79hpke5d0gsnx93d82k67r\",\n  \"slug\": \"baoyu-url-to-markdown\",\n  \"version\": \"1.103.0\",\n  \"publishedAt\": 1776043379026\n}\n\nFile v1.103.0:references/config/first-time-setup.md\n\n---\nname: first-time-setup\ndescription: First-time setup flow for baoyu-url-to-markdown preferences\n---\n\n# First-Time Setup\n\n## Overview\n\nWhen no EXTEND.md is found, guide user through preference setup.\n\n**BLOCKING OPERATION**: This setup MUST complete before ANY other workflow steps. Do NOT:\n- Start converting URLs\n- Ask about URLs or output paths\n- Proceed to any conversion\n\nONLY ask the questions in this setup flow, save EXTEND.md, then continue.\n\n## Setup Flow\n\n```\nNo EXTEND.md found\n        |\n        v\n+---------------------+\n| AskUserQuestion     |\n| (all questions)     |\n+---------------------+\n        |\n        v\n+---------------------+\n| Create EXTEND.md    |\n+---------------------+\n        |\n        v\n    Continue conversion\n```\n\n## Questions\n\n**Language**: Use user's input language or saved language preference.\n\nUse AskUserQuestion with ALL questions in ONE call:\n\n### Question 1: Download Media\n\n```yaml\nheader: \"Media\"\nquestion: \"How to handle images and videos in pages?\"\noptions:\n  - label: \"Ask each time (Recommended)\"\n    description: \"After saving markdown, ask whether to download media\"\n  - label: \"Always download\"\n    description: \"Always download media to local imgs/ and videos/ directories\"\n  - label: \"Never download\"\n    description: \"Keep original remote URLs in markdown\"\n```\n\n### Question 2: Default Output Directory\n\n```yaml\nheader: \"Output\"\nquestion: \"Default output directory?\"\noptions:\n  - label: \"url-to-markdown (Recommended)\"\n    description: \"Save to ./url-to-markdown/{domain}/{slug}.md\"\n```\n\nNote: User will likely choose \"Other\" to type a custom path.\n\n### Question 3: Save Location\n\n```yaml\nheader: \"Save\"\nquestion: \"Where to save preferences?\"\noptions:\n  - label: \"User (Recommended)\"\n    description: \"~/.baoyu-skills/ (all projects)\"\n  - label: \"Project\"\n    description: \".baoyu-skills/ (this project only)\"\n```\n\n## Save Locations\n\n| Choice | Path | Scope |\n|--------|------|-------|\n| User | `~/.baoyu-skills/baoyu-url-to-markdown/EXTEND.md` | All projects |\n| Project | `.baoyu-skills/baoyu-url-to-markdown/EXTEND.md` | Current project |\n\n## After Setup\n\n1. Create directory if needed\n2. Write EXTEND.md\n3. Confirm: \"Preferences saved to [path]\"\n4. Continue with conversion using saved preferences\n\n## EXTEND.md Template\n\n```md\ndownload_media: [ask/1/0]\ndefault_output_dir: [path or empty]\n```\n\n## Modifying Preferences Later\n\nUsers can edit EXTEND.md directly or delete it to trigger setup again.\n\nFile v1.103.0:scripts/package.json\n\n{\n  \"name\": \"baoyu-url-to-markdown-scripts\",\n  \"private\": true,\n  \"type\": \"module\",\n  \"dependencies\": {\n    \"baoyu-fetch\": \"file:./vendor/baoyu-fetch\"\n  }\n}\n\nFile v1.103.0:scripts/vendor/baoyu-fetch/package.json\n\n{\n  \"name\": \"baoyu-fetch\",\n  \"version\": \"0.1.1\",\n  \"description\": \"Read URLs into high-quality Markdown or JSON with Chrome CDP and site adapters.\",\n  \"type\": \"module\",\n  \"bin\": {\n    \"baoyu-fetch\": \"./src/cli.ts\"\n  },\n  \"files\": [\n    \"README.zh-CN.md\",\n    \"src/adapters\",\n    \"src/browser\",\n    \"src/cli.ts\",\n    \"src/commands\",\n    \"src/extract\",\n    \"src/media\",\n    \"src/types\",\n    \"src/utils\",\n    \"README.md\"\n  ],\n  \"repository\": {\n    \"type\": \"git\",\n    \"url\": \"git+https://github.com/JimLiu/baoyu-skills.git\",\n    \"directory\": \"packages/baoyu-fetch\"\n  },\n  \"bugs\": {\n    \"url\": \"https://github.com/JimLiu/baoyu-skills/issues\"\n  },\n  \"homepage\": \"https://github.com/JimLiu/baoyu-skills/tree/main/packages/baoyu-fetch#readme\",\n  \"publishConfig\": {\n    \"access\": \"public\"\n  },\n  \"scripts\": {\n    \"build\": \"rm -rf dist && bun build ./src/cli.ts --target bun --outfile ./dist/cli.js && chmod +x ./dist/cli.js\",\n    \"check\": \"tsc --noEmit\",\n    \"dev\": \"bun run ./src/cli.ts\",\n    \"release\": \"changeset publish\",\n    \"test\": \"bun test\",\n    \"version-packages\": \"changeset version\"\n  },\n  \"engines\": {\n    \"bun\": \">=1.2.0\"\n  },\n  \"dependencies\": {\n    \"@mozilla/readability\": \"^0.6.0\",\n    \"chrome-launcher\": \"^1.2.1\",\n    \"defuddle\": \"^0.14.0\",\n    \"jsdom\": \"^26.0.0\",\n    \"remark-gfm\": \"^4.0.1\",\n    \"remark-parse\": \"^11.0.0\",\n    \"turndown\": \"^7.2.0\",\n    \"turndown-plugin-gfm\": \"^1.0.2\",\n    \"unified\": \"^11.0.5\",\n    \"ws\": \"^8.18.3\"\n  },\n  \"devDependencies\": {\n    \"@changesets/cli\": \"^2.30.0\",\n    \"@types/bun\": \"^1.2.23\",\n    \"@types/jsdom\": \"^21.1.7\",\n    \"@types/ws\": \"^8.18.1\",\n    \"typescript\": \"^5.9.2\"\n  }\n}\n\nFile v1.103.0:scripts/vendor/baoyu-fetch/README.zh-CN.md\n\n# baoyu-fetch\n\n[English](./README.md) | 简体中文 | [更新日志](./CHANGELOG.zh-CN.md) | [English Changelog](./CHANGELOG.md)\n\n`baoyu-fetch` 是一个基于 Chrome CDP 的 Bun CLI。输入 URL，它会输出高质量\n`markdown` 或 `json`；命中站点 adapter 时优先消费 API 返回或页面内结构化\n数据，未命中时回退到通用 HTML 提取。\n\n## 当前能力\n\n- 通过 Chrome CDP 抓取渲染后的页面内容\n- 监听网络请求与响应，按需拉取响应体\n- adapter registry，支持按 URL 自动命中站点处理器\n- 内置 `x`、`youtube`、`hn` adapters\n- 通用 fallback：Defuddle 优先，Readability + HTML to Markdown 回退；`--format markdown` 时会再尝试 `defuddle.md` 兜底\n- `stdout` 或 `--output` 输出 `markdown` / `json`\n- 可选下载提取出的图片/视频并重写 Markdown 链接\n- 提供登录/验证场景下的交互等待模式\n- Chrome profile 默认对齐 `baoyu-skills/chrome-profile`\n\n## 安装\n\n```bash\nbun install\n```\n\n作为包使用时，推荐直接这样运行：\n\n```bash\nbunx baoyu-fetch https://example.com\n```\n\n也可以全局安装：\n\n```bash\nnpm install -g baoyu-fetch\n```\n\nnpm 包发布的是 TypeScript 源码入口，不包含预编译的 `dist`，所以运行时需要\nBun。\n\n## 用法\n\n```bash\nbun run src/cli.ts https://example.com\nbunx baoyu-fetch https://example.com\nbaoyu-fetch https://example.com\nbaoyu-fetch https://example.com --format markdown --output article.md\nbaoyu-fetch https://example.com --format markdown --output article.md --download-media\nbaoyu-fetch https://x.com/jack/status/20 --format json --output article.json\nbaoyu-fetch https://x.com/jack/status/20 --json\nbaoyu-fetch https://x.com/jack/status/20 --wait-for interaction\nbaoyu-fetch https://x.com/jack/status/20 --wait-for force\nbaoyu-fetch https://x.com/jack/status/20 --chrome-profile-dir ~/Library/Application\\\\ Support/baoyu-skills/chrome-profile\n```\n\n## 主要参数\n\n```bash\nbaoyu-fetch <url> [options]\n\nOptions:\n  --output <file>       保存输出内容到文件\n  --format <type>       输出格式：markdown | json\n  --json                `--format json` 的兼容别名\n  --adapter <name>      强制使用指定 adapter（如 x / hn / generic）\n  --download-media      下载 adapter 返回的媒体到 ./imgs 和 ./videos，并重写 markdown 链接\n  --media-dir <dir>     指定媒体下载根目录；默认使用输出文件所在目录\n  --debug-dir <dir>     导出调试信息（html、document.json、network.json）\n  --cdp-url <url>       连接现有 Chrome 调试地址\n  --browser-path <path> 指定 Chrome 可执行文件\n  --chrome-profile-dir <path>\n                        指定 Chrome profile 目录。默认使用 BAOYU_CHROME_PROFILE_DIR，\n                        否则回退到 baoyu-skills/chrome-profile\n  --headless            启动临时 headless Chrome（未连现有实例时）\n  --wait-for <mode>     等待模式：interaction | force\n  --wait-for-interaction\n                        `--wait-for interaction` 的别名\n  --wait-for-login      `--wait-for interaction` 的别名\n  --interaction-timeout <ms>\n                        手动交互等待超时，默认 600000\n  --interaction-poll-interval <ms>\n                        等待期间的轮询间隔，默认 1500\n  --login-timeout <ms>  `--interaction-timeout` 的别名\n  --login-poll-interval <ms>\n                        `--interaction-poll-interval` 的别名\n  --timeout <ms>        页面加载超时，默认 30000\n  --help                显示帮助\n```\n\n## 设计\n\n核心链路：\n\n1. CLI 解析 URL 和选项\n2. 建立 CDP 会话并创建受控 tab\n3. 启动 `NetworkJournal` 收集所有请求/响应\n4. 由 adapter registry 匹配站点 adapter\n5. adapter 返回结构化 `ExtractedDocument`\n6. 没命中则走通用 HTML 提取\n7. 按请求输出 Markdown，或输出包含 `document` 和 `markdown` 的 JSON\n\n## 开发\n\n```bash\nbun run check\nbun run test\nbun run build\n```\n\n## 发版\n\n新增用户可见改动后，先添加一个 changeset：\n\n```bash\nbunx changeset\n```\n\n把生成的 `.changeset/*.md` 一起合并到 `main` 后，GitHub Actions 会自动创建或\n更新 release PR；合并 release PR 之后，会自动发布到 npm。\n\n发布流程不会编译 `dist`，而是直接把 `src/*.ts` 发布出去供 Bun 执行。\n\nArchive v1.82.2: 48 files, 86106 bytes\n\nFiles: references/config/first-time-setup.md (2459b), scripts/package.json (157b), scripts/vendor/baoyu-fetch/package.json (1624b), scripts/vendor/baoyu-fetch/README.md (4477b), scripts/vendor/baoyu-fetch/README.zh-CN.md (4326b), scripts/vendor/baoyu-fetch/src/adapters/generic/index.ts (2451b), scripts/vendor/baoyu-fetch/src/adapters/hn/index.ts (10454b), scripts/vendor/baoyu-fetch/src/adapters/index.ts (834b), scripts/vendor/baoyu-fetch/src/adapters/types.ts (2100b), scripts/vendor/baoyu-fetch/src/adapters/x/article.ts (12122b), scripts/vendor/baoyu-fetch/src/adapters/x/index.ts (4048b), scripts/vendor/baoyu-fetch/src/adapters/x/login.ts (2258b), scripts/vendor/baoyu-fetch/src/adapters/x/match.ts (297b), scripts/vendor/baoyu-fetch/src/adapters/x/payloads.ts (1681b), scripts/vendor/baoyu-fetch/src/adapters/x/session.ts (1404b), scripts/vendor/baoyu-fetch/src/adapters/x/shared.ts (11727b), scripts/vendor/baoyu-fetch/src/adapters/x/single.ts (2583b), scripts/vendor/baoyu-fetch/src/adapters/x/thread-loader.ts (9078b), scripts/vendor/baoyu-fetch/src/adapters/x/thread.ts (9196b), scripts/vendor/baoyu-fetch/src/adapters/x/types.ts (600b), scripts/vendor/baoyu-fetch/src/adapters/youtube/index.ts (898b), scripts/vendor/baoyu-fetch/src/adapters/youtube/transcript.ts (13301b), scripts/vendor/baoyu-fetch/src/adapters/youtube/utils.ts (6523b), scripts/vendor/baoyu-fetch/src/browser/cdp-client.ts (6879b), scripts/vendor/baoyu-fetch/src/browser/chrome-launcher.ts (5883b), scripts/vendor/baoyu-fetch/src/browser/cookie-sidecar.ts (2941b), scripts/vendor/baoyu-fetch/src/browser/interaction-gates.ts (4095b), scripts/vendor/baoyu-fetch/src/browser/network-journal.ts (7153b), scripts/vendor/baoyu-fetch/src/browser/page-snapshot.ts (3429b), scripts/vendor/baoyu-fetch/src/browser/profile.ts (5670b), scripts/vendor/baoyu-fetch/src/browser/session.ts (4642b), scripts/vendor/baoyu-fetch/src/cli.ts (7066b), scripts/vendor/baoyu-fetch/src/commands/convert.ts (17456b), scripts/vendor/baoyu-fetch/src/extract/document.ts (846b), scripts/vendor/baoyu-fetch/src/extract/html-cleaner.ts (11016b), scripts/vendor/baoyu-fetch/src/extract/html-extractor.ts (2243b), scripts/vendor/baoyu-fetch/src/extract/html-to-markdown.ts (21344b), scripts/vendor/baoyu-fetch/src/extract/markdown-renderer.ts (4661b), scripts/vendor/baoyu-fetch/src/media/default-downloader.ts (5215b), scripts/vendor/baoyu-fetch/src/media/markdown-media.ts (13186b), scripts/vendor/baoyu-fetch/src/media/media-utils.ts (6222b), scripts/vendor/baoyu-fetch/src/media/types.ts (710b), scripts/vendor/baoyu-fetch/src/types/defuddle-node.d.ts (605b), scripts/vendor/baoyu-fetch/src/types/shims.d.ts (422b), scripts/vendor/baoyu-fetch/src/utils/logger.ts (645b), scripts/vendor/baoyu-fetch/src/utils/url.ts (300b), SKILL.md (15269b), _meta.json (141b)\n\nFile v1.82.2:SKILL.md\n\n---\nname: baoyu-url-to-markdown\ndescription: Fetch any URL and convert to markdown using baoyu-fetch CLI (Chrome CDP with site-specific adapters). Built-in adapters for X/Twitter, YouTube transcripts, Hacker News threads, and generic pages via Defuddle. Handles login/CAPTCHA via interaction wait modes. Use when user wants to save a webpage as markdown.\nversion: 1.60.0\nmetadata:\n  openclaw:\n    homepage: https://github.com/JimLiu/baoyu-skills#baoyu-url-to-markdown\n    requires:\n      anyBins:\n        - bun\n        - npx\n---\n\n# URL to Markdown\n\nFetches any URL via `baoyu-fetch` CLI (Chrome CDP + site-specific adapters) and converts it to clean markdown.\n\n## CLI Setup\n\n**Important**: The CLI source is vendored in the `scripts/vendor/baoyu-fetch/` subdirectory of this skill.\n\n**Agent Execution Instructions**:\n1. Determine this SKILL.md file's directory path as `{baseDir}`\n2. CLI entry point = `{baseDir}/scripts/vendor/baoyu-fetch/src/cli.ts`\n3. Resolve `${BUN_X}` runtime: if `bun` installed → `bun`; if `npx` available → `npx -y bun`; else suggest installing bun\n4. `${READER}` = `${BUN_X} {baseDir}/scripts/vendor/baoyu-fetch/src/cli.ts`\n5. Replace all `${READER}` in this document with the resolved value\n\n## Preferences (EXTEND.md)\n\nCheck EXTEND.md existence (priority order):\n\n```bash\n# macOS, Linux, WSL, Git Bash\ntest -f .baoyu-skills/baoyu-url-to-markdown/EXTEND.md && echo \"project\"\ntest -f \"${XDG_CONFIG_HOME:-$HOME/.config}/baoyu-skills/baoyu-url-to-markdown/EXTEND.md\" && echo \"xdg\"\ntest -f \"$HOME/.baoyu-skills/baoyu-url-to-markdown/EXTEND.md\" && echo \"user\"\n```\n\n```powershell\n# PowerShell (Windows)\nif (Test-Path .baoyu-skills/baoyu-url-to-markdown/EXTEND.md) { \"project\" }\n$xdg = if ($env:XDG_CONFIG_HOME) { $env:XDG_CONFIG_HOME } else { \"$HOME/.config\" }\nif (Test-Path \"$xdg/baoyu-skills/baoyu-url-to-markdown/EXTEND.md\") { \"xdg\" }\nif (Test-Path \"$HOME/.baoyu-skills/baoyu-url-to-markdown/EXTEND.md\") { \"user\" }\n```\n\n| Path | Location |\n|------|----------|\n| `.baoyu-skills/baoyu-url-to-markdown/EXTEND.md` | Project directory |\n| `$HOME/.baoyu-skills/baoyu-url-to-markdown/EXTEND.md` | User home |\n\n| Result | Action |\n|--------|--------|\n| Found | Read, parse, apply settings |\n| Not found | **MUST** run first-time setup (see below) — do NOT silently create defaults |\n\n**EXTEND.md Supports**: Download media by default | Default output directory\n\n### First-Time Setup (BLOCKING)\n\n**CRITICAL**: When EXTEND.md is not found, you **MUST use `AskUserQuestion`** to ask the user for their preferences before creating EXTEND.md. **NEVER** create EXTEND.md with defaults without asking. This is a **BLOCKING** operation — do NOT proceed with any conversion until setup is complete.\n\nUse `AskUserQuestion` with ALL questions in ONE call:\n\n**Question 1** — header: \"Media\", question: \"How to handle images and videos in pages?\"\n- \"Ask each time (Recommended)\" — After saving markdown, ask whether to download media\n- \"Always download\" — Always download media to local imgs/ and videos/ directories\n- \"Never download\" — Keep original remote URLs in markdown\n\n**Question 2** — header: \"Output\", question: \"Default output directory?\"\n- \"url-to-markdown (Recommended)\" — Save to ./url-to-markdown/{domain}/{slug}.md\n- (User may choose \"Other\" to type a custom path)\n\n**Question 3** — header: \"Save\", question: \"Where to save preferences?\"\n- \"User (Recommended)\" — ~/.baoyu-skills/ (all projects)\n- \"Project\" — .baoyu-skills/ (this project only)\n\nAfter user answers, create EXTEND.md at the chosen location, confirm \"Preferences saved to [path]\", then continue.\n\nFull reference: [references/config/first-time-setup.md](references/config/first-time-setup.md)\n\n### Supported Keys\n\n| Key | Default | Values | Description |\n|-----|---------|--------|-------------|\n| `download_media` | `ask` | `ask` / `1` / `0` | `ask` = prompt each time, `1` = always download, `0` = never |\n| `default_output_dir` | empty | path or empty | Default output directory (empty = `./url-to-markdown/`) |\n\n**EXTEND.md → CLI mapping**:\n| EXTEND.md key | CLI argument | Notes |\n|---------------|-------------|-------|\n| `download_media: 1` | `--download-media` | Requires `--output` to be set |\n| `default_output_dir: ./posts/` | Agent constructs `--output ./posts/{domain}/{slug}.md` | Agent generates path, not a direct CLI flag |\n\n**Value priority**:\n1. CLI arguments (`--download-media`, `--output`)\n2. EXTEND.md\n3. Skill defaults\n\n## Features\n\n- Chrome CDP for full JavaScript rendering via `baoyu-fetch` CLI\n- Site-specific adapters: X/Twitter, YouTube, Hacker News, generic (Defuddle)\n- Automatic adapter selection based on URL, or force with `--adapter`\n- Interaction gate detection: Cloudflare, reCAPTCHA, hCAPTCHA, custom challenges\n- Two capture modes: headless (default) or interactive with wait-for-interaction\n- Clean markdown output with YAML front matter\n- Structured JSON output available via `--format json`\n- X/Twitter: extracts tweets, threads, and X Articles with media\n- YouTube: transcript/caption extraction, chapters, cover images\n- Hacker News: threaded comment parsing with proper nesting\n- Generic: Defuddle extraction with Readability fallback\n- Download images and videos to local directories\n- Chrome profile persistence for authenticated sessions\n- Debug artifact output for troubleshooting\n\n## Usage\n\n```bash\n# Default: headless capture, markdown to stdout\n${READER} <url>\n\n# Save to file\n${READER} <url> --output article.md\n\n# Save with media download\n${READER} <url> --output article.md --download-media\n\n# Headless mode (explicit)\n${READER} <url> --headless --output article.md\n\n# Wait for interaction (login/CAPTCHA) — auto-detect and continue\n${READER} <url> --wait-for interaction --output article.md\n\n# Wait for interaction — manual control (Enter to continue)\n${READER} <url> --wait-for force --output article.md\n\n# JSON output\n${READER} <url> --format json --output article.json\n\n# Force specific adapter\n${READER} <url> --adapter youtube --output transcript.md\n\n# Connect to existing Chrome\n${READER} <url> --cdp-url http://localhost:9222 --output article.md\n\n# Debug artifacts\n${READER} <url> --output article.md --debug-dir ./debug/\n```\n\n## Options\n\n| Option | Description |\n|--------|-------------|\n| `<url>` | URL to fetch |\n| `--output <path>` | Output file path (default: stdout) |\n| `--format <type>` | Output format: `markdown` (default) or `json` |\n| `--json` | Shorthand for `--format json` |\n| `--adapter <name>` | Force adapter: `x`, `youtube`, `hn`, or `generic` (default: auto-detect) |\n| `--headless` | Force headless Chrome (no visible window) |\n| `--wait-for <mode>` | Interaction wait mode: `none` (default), `interaction`, or `force` |\n| `--wait-for-interaction` | Alias for `--wait-for interaction` |\n| `--wait-for-login` | Alias for `--wait-for interaction` |\n| `--timeout <ms>` | Page load timeout (default: 30000) |\n| `--interaction-timeout <ms>` | Login/CAPTCHA wait timeout (default: 600000 = 10 min) |\n| `--interaction-poll-interval <ms>` | Poll interval for interaction checks (default: 1500) |\n| `--download-media` | Download images/videos to local `imgs/` and `videos/`, rewrite markdown links. Requires `--output` |\n| `--media-dir <dir>` | Base directory for downloaded media (default: same as `--output` directory) |\n| `--cdp-url <url>` | Reuse existing Chrome DevTools Protocol endpoint |\n| `--browser-path <path>` | Custom Chrome/Chromium binary path |\n| `--chrome-profile-dir <path>` | Chrome user data directory (default: `BAOYU_CHROME_PROFILE_DIR` env or `./baoyu-skills/chrome-profile`) |\n| `--debug-dir <dir>` | Write debug artifacts (document.json, markdown.md, page.html, network.json) |\n\n## Capture Modes\n\n| Mode | Behavior | Use When |\n|------|----------|----------|\n| Default | Headless Chrome, auto-extract on network idle | Public pages, static content |\n| `--headless` | Explicit headless (same as default) | Clarify intent |\n| `--wait-for interaction` | Opens visible Chrome, auto-detects login/CAPTCHA gates, waits for them to clear, then continues | Login-required, CAPTCHA-protected |\n| `--wait-for force` | Opens visible Chrome, auto-detects OR accepts Enter keypress to continue | Complex flows, lazy loading, paywalls |\n\n**Interaction gate auto-detection**:\n- Cloudflare Turnstile / \"just a moment\" pages\n- Google reCAPTCHA\n- hCaptcha\n- Custom challenge / verification screens\n\n**Wait-for-interaction workflow**:\n1. Run with `--wait-for interaction` → Chrome opens visibly\n2. CLI auto-detects login/CAPTCHA gates\n3. User completes login or solves CAPTCHA in the browser\n4. CLI auto-detects gate cleared → captures page\n5. If `--wait-for force` is used, user can also press Enter to trigger capture manually\n\n## Agent Quality Gate\n\n**CRITICAL**: The agent must treat default headless capture as provisional. Some sites render differently in headless mode and can silently return low-quality content without causing the CLI to fail.\n\nAfter every headless run, the agent **MUST** inspect the saved markdown output.\n\n### Quality checks the agent must perform\n\n1. Confirm the markdown title matches the target page, not a generic site shell\n2. Confirm the body contains the expected article or page content, not just navigation, footer, or a generic error\n3. Watch for obvious failure signs:\n   - `Application error`\n   - `This page could not be found`\n   - Login, signup, subscribe, or verification shells\n   - Extremely short markdown for a page that should be long-form\n   - Raw framework payloads or mostly boilerplate content\n4. If the result is low quality, incomplete, or clearly wrong, do **not** accept the run as successful just because the CLI exited with code 0\n\n**Tip**: Use `--format json` to get structured output including `status`, `login.state`, and `interaction` fields for programmatic quality assessment. A `\"status\": \"needs_interaction\"` response means the page requires manual interaction.\n\n### Recovery workflow the agent must follow\n\n1. First run headless (default) unless there is already a clear reason to use interaction mode\n2. Review markdown quality immediately after the run\n3. If the content is low quality or indicates login/CAPTCHA:\n   - `--wait-for interaction` for auto-detected gates (login, CAPTCHA, Cloudflare)\n   - `--wait-for force` when the page needs manual browsing, scroll loading, or complex interaction\n4. If `--wait-for` is used, tell the user exactly what to do:\n   - If login is required, ask them to sign in in the browser\n   - If CAPTCHA appears, ask them to solve it\n   - If the page needs time to load, ask them to wait until content is visible\n   - For `--wait-for force`: tell them to press Enter when ready\n5. If JSON output shows `\"status\": \"needs_interaction\"`, switch to `--wait-for interaction` automatically\n\n## Output Path Generation\n\nThe agent must construct the output file path since `baoyu-fetch` does not auto-generate paths.\n\n**Algorithm**:\n1. Determine base directory from EXTEND.md `default_output_dir` or default `./url-to-markdown/`\n2. Extract domain from URL (e.g., `example.com`)\n3. Generate slug from URL path or page title (kebab-case, 2-6 words)\n4. Construct: `{base_dir}/{domain}/{slug}/{slug}.md` — each URL gets its own directory so media files stay isolated\n5. Conflict resolution: append timestamp `{slug}-YYYYMMDD-HHMMSS/{slug}-YYYYMMDD-HHMMSS.md`\n\nPass the constructed path to `--output`. Media files (`--download-media`) are saved into subdirectories next to the markdown file, keeping each URL's assets self-contained.\n\n## Output Format\n\nMarkdown output to stdout (or file with `--output`) as clean markdown text.\n\nJSON output (`--format json`) returns structured data including:\n- `adapter` — which adapter handled the URL\n- `status` — `\"ok\"` or `\"needs_interaction\"`\n- `login` — login state detection (`logged_in`, `logged_out`, `unknown`)\n- `interaction` — interaction gate details (kind, provider, prompt)\n- `document` — structured content (url, title, author, publishedAt, content blocks, metadata)\n- `media` — collected media assets with url, kind, role\n- `markdown` — converted markdown text\n- `downloads` — media download results (when `--download-media` used)\n\nWhen `--download-media` is enabled:\n- Images are saved to `imgs/` next to the output file (or in `--media-dir`)\n- Videos are saved to `videos/` next to the output file (or in `--media-dir`)\n- Markdown media links are rewritten to local relative paths\n\n## Built-in Adapters\n\n| Adapter | URLs | Key Features |\n|---------|------|-------------|\n| `x` | x.com, twitter.com | Tweets, threads, X Articles, media, login detection |\n| `youtube` | youtube.com, youtu.be | Transcript/captions, chapters, cover image, metadata |\n| `hn` | news.ycombinator.com | Threaded comments, story metadata, nested replies |\n| `generic` | Any URL (fallback) | Defuddle extraction, Readability fallback, auto-scroll, network idle detection |\n\nAdapter is auto-selected based on URL. Use `--adapter <name>` to override.\n\n## Media Download Workflow\n\nBased on `download_media` setting in EXTEND.md:\n\n| Setting | Behavior |\n|---------|----------|\n| `1` (always) | Run CLI with `--download-media --output <path>` |\n| `0` (never) | Run CLI with `--output <path>` (no media download) |\n| `ask` (default) | Follow the ask-each-time flow below |\n\n### Ask-Each-Time Flow\n\n1. Run CLI **without** `--download-media` with `--output <path>` → markdown saved\n2. Check saved markdown for remote media URLs (`https://` in image/video links)\n3. **If no remote media found** → done, no prompt needed\n4. **If remote media found** → use `AskUserQuestion`:\n   - header: \"Media\", question: \"Download N images/videos to local files?\"\n   - \"Yes\" — Download to local directories\n   - \"No\" — Keep remote URLs\n5. If user confirms → run CLI **again** with `--download-media --output <same-path>` (overwrites markdown with localized links)\n\n## Environment Variables\n\n| Variable | Description |\n|----------|-------------|\n| `BAOYU_CHROME_PROFILE_DIR` | Chrome user data directory (can also use `--chrome-profile-dir`) |\n\n**Troubleshooting**: Chrome not found → use `--browser-path`. Timeout → increase `--timeout`. Login/CAPTCHA pages → use `--wait-for interaction`. Debug → use `--debug-dir` to inspect captured HTML and network logs.\n\n### YouTube Notes\n\n- YouTube adapter extracts transcripts/captions automatically when available\n- Transcript format: `[MM:SS] Text segment` with chapter headings\n- Transcript availability depends on YouTube exposing a caption track. Videos with captions disabled or restricted playback may produce description-only output\n- Use `--wait-for force` if the page needs time to finish loading player metadata\n\n### X/Twitter Notes\n\n- Extracts single tweets, threads, and X Articles\n- Auto-detects login state; if logged out and content requires auth, JSON output will show `\"status\": \"needs_interaction\"`\n- Use `--wait-for interaction` for login-protected content\n\n### Hacker News Notes\n\n- Parses threaded comments with proper nesting and reply hierarchy\n- Includes story metadata (title, URL, author, score, comment count)\n- Shows comment deletion/dead status\n\n## Extension Support\n\nCustom configurations via EXTEND.md. See **Preferences** section for paths and supported options.\n\nFile v1.82.2:scripts/vendor/baoyu-fetch/README.md\n\n# baoyu-fetch\n\nEnglish | [简体中文](./README.zh-CN.md) | [Changelog](./CHANGELOG.md) | [中文更新日志](./CHANGELOG.zh-CN.md)\n\n`baoyu-fetch` is a Bun CLI built on Chrome CDP. Give it a URL and it returns\nhigh-quality `markdown` or `json`. When a site adapter matches, it prefers API\nresponses or structured page data; otherwise it falls back to generic HTML\nextraction.\n\n## Features\n\n- Capture rendered page content through Chrome CDP\n- Observe network requests and responses, and fetch bodies when needed\n- Adapter registry that auto-selects a handler from the URL\n- Built-in adapters for `x`, `youtube`, and `hn`\n- Generic fallback: Defuddle first, then Readability + HTML-to-Markdown; when `--format markdown` is requested, it can also fall back to `defuddle.md`\n- Print `markdown` / `json` to stdout or save with `--output`\n- Optionally download extracted images or videos and rewrite Markdown links\n- Optional wait modes for login and verification flows\n- Chrome profile defaults to `baoyu-skills/chrome-profile`\n\n## Installation\n\n```bash\nbun install\n```\n\nFor package usage, the quickest option is:\n\n```bash\nbunx baoyu-fetch https://example.com\n```\n\nYou can also install it globally:\n\n```bash\nnpm install -g baoyu-fetch\n```\n\nThe npm package ships TypeScript source entrypoints instead of a prebuilt\n`dist`, so Bun is required at runtime.\n\n## Usage\n\n```bash\nbun run src/cli.ts https://example.com\nbunx baoyu-fetch https://example.com\nbaoyu-fetch https://example.com\nbaoyu-fetch https://example.com --format markdown --output article.md\nbaoyu-fetch https://example.com --format markdown --output article.md --download-media\nbaoyu-fetch https://x.com/jack/status/20 --format json --output article.json\nbaoyu-fetch https://x.com/jack/status/20 --json\nbaoyu-fetch https://x.com/jack/status/20 --wait-for interaction\nbaoyu-fetch https://x.com/jack/status/20 --wait-for force\nbaoyu-fetch https://x.com/jack/status/20 --chrome-profile-dir ~/Library/Application\\\\ Support/baoyu-skills/chrome-profile\n```\n\n## Options\n\n```bash\nbaoyu-fetch <url> [options]\n\nOptions:\n  --output <file>       Save output to file\n  --format <type>       Output format: markdown | json\n  --json                Alias for --format json\n  --adapter <name>      Force an adapter (for example x / hn / generic)\n  --download-media      Download adapter-reported media into ./imgs and ./videos, then rewrite markdown links\n  --media-dir <dir>     Base directory for downloaded media. Defaults to the output directory\n  --debug-dir <dir>     Write debug artifacts (html, document.json, network.json)\n  --cdp-url <url>       Reuse an existing Chrome DevTools endpoint\n  --browser-path <path> Explicit Chrome binary path\n  --chrome-profile-dir <path>\n                        Chrome user data dir. Defaults to BAOYU_CHROME_PROFILE_DIR\n                        or baoyu-skills/chrome-profile\n  --headless            Launch a temporary headless Chrome if needed\n  --wait-for <mode>     Wait mode: interaction | force\n  --wait-for-interaction\n                        Alias for --wait-for interaction\n  --wait-for-login      Alias for --wait-for interaction\n  --interaction-timeout <ms>\n                        Manual interaction timeout. Default: 600000\n  --interaction-poll-interval <ms>\n                        Poll interval while waiting. Default: 1500\n  --login-timeout <ms>  Alias for --interaction-timeout\n  --login-poll-interval <ms>\n                        Alias for --interaction-poll-interval\n  --timeout <ms>        Page load timeout. Default: 30000\n  --help                Show help\n```\n\n## How It Works\n\n1. The CLI parses the target URL and options.\n2. It opens or connects to a Chrome CDP session and creates a controlled tab.\n3. `NetworkJournal` records requests and responses.\n4. The adapter registry resolves a site-specific adapter when possible.\n5. The adapter returns a structured `ExtractedDocument`.\n6. If nothing matches, generic HTML extraction runs instead.\n7. The result is rendered as Markdown, or returned as JSON with both\n   `document` and `markdown`.\n\n## Development\n\n```bash\nbun run check\nbun run test\nbun run build\n```\n\n## Release\n\nWhen you make a user-visible change, add a changeset first:\n\n```bash\nbunx changeset\n```\n\nAfter the generated `.changeset/*.md` file lands on `main`, GitHub Actions will\nopen or update the release PR. Merging that release PR publishes the package to\nnpm.\n\nThe publish flow does not build `dist`; it publishes `src/*.ts` for Bun\nexecution directly.\n\nFile v1.82.2:_meta.json\n\n{\n  \"ownerId\": \"kn7csrrndw79hpke5d0gsnx93d82k67r\",\n  \"slug\": \"baoyu-url-to-markdown\",\n  \"version\": \"1.82.2\",\n  \"publishedAt\": 1775752206118\n}\n\nFile v1.82.2:references/config/first-time-setup.md\n\n---\nname: first-time-setup\ndescription: First-time setup flow for baoyu-url-to-markdown preferences\n---\n\n# First-Time Setup\n\n## Overview\n\nWhen no EXTEND.md is found, guide user through preference setup.\n\n**BLOCKING OPERATION**: This setup MUST complete before ANY other workflow steps. Do NOT:\n- Start converting URLs\n- Ask about URLs or output paths\n- Proceed to any conversion\n\nONLY ask the questions in this setup flow, save EXTEND.md, then continue.\n\n## Setup Flow\n\n```\nNo EXTEND.md found\n        |\n        v\n+---------------------+\n| AskUserQuestion     |\n| (all questions)     |\n+---------------------+\n        |\n        v\n+---------------------+\n| Create EXTEND.md    |\n+---------------------+\n        |\n        v\n    Continue conversion\n```\n\n## Questions\n\n**Language**: Use user's input language or saved language preference.\n\nUse AskUserQuestion with ALL questions in ONE call:\n\n### Question 1: Download Media\n\n```yaml\nheader: \"Media\"\nquestion: \"How to handle images and videos in pages?\"\noptions:\n  - label: \"Ask each time (Recommended)\"\n    description: \"After saving markdown, ask whether to download media\"\n  - label: \"Always download\"\n    description: \"Always download media to local imgs/ and videos/ directories\"\n  - label: \"Never download\"\n    description: \"Keep original remote URLs in markdown\"\n```\n\n### Question 2: Default Output Directory\n\n```yaml\nheader: \"Output\"\nquestion: \"Default output directory?\"\noptions:\n  - label: \"url-to-markdown (Recommended)\"\n    description: \"Save to ./url-to-markdown/{domain}/{slug}.md\"\n```\n\nNote: User will likely choose \"Other\" to type a custom path.\n\n### Question 3: Save Location\n\n```yaml\nheader: \"Save\"\nquestion: \"Where to save preferences?\"\noptions:\n  - label: \"User (Recommended)\"\n    description: \"~/.baoyu-skills/ (all projects)\"\n  - label: \"Project\"\n    description: \".baoyu-skills/ (this project only)\"\n```\n\n## Save Locations\n\n| Choice | Path | Scope |\n|--------|------|-------|\n| User | `~/.baoyu-skills/baoyu-url-to-markdown/EXTEND.md` | All projects |\n| Project | `.baoyu-skills/baoyu-url-to-markdown/EXTEND.md` | Current project |\n\n## After Setup\n\n1. Create directory if needed\n2. Write EXTEND.md\n3. Confirm: \"Preferences saved to [path]\"\n4. Continue with conversion using saved preferences\n\n## EXTEND.md Template\n\n```md\ndownload_media: [ask/1/0]\ndefault_output_dir: [path or empty]\n```\n\n## Modifying Preferences Later\n\nUsers can edit EXTEND.md directly or delete it to trigger setup again.\n\nFile v1.82.2:scripts/package.json\n\n{\n  \"name\": \"baoyu-url-to-markdown-scripts\",\n  \"private\": true,\n  \"type\": \"module\",\n  \"dependencies\": {\n    \"baoyu-fetch\": \"file:./vendor/baoyu-fetch\"\n  }\n}\n\nFile v1.82.2:scripts/vendor/baoyu-fetch/package.json\n\n{\n  \"name\": \"baoyu-fetch\",\n  \"version\": \"0.1.1\",\n  \"description\": \"Read URLs into high-quality Markdown or JSON with Chrome CDP and site adapters.\",\n  \"type\": \"module\",\n  \"bin\": {\n    \"baoyu-fetch\": \"./src/cli.ts\"\n  },\n  \"files\": [\n    \"README.zh-CN.md\",\n    \"src/adapters\",\n    \"src/browser\",\n    \"src/cli.ts\",\n    \"src/commands\",\n    \"src/extract\",\n    \"src/media\",\n    \"src/types\",\n    \"src/utils\",\n    \"README.md\"\n  ],\n  \"repository\": {\n    \"type\": \"git\",\n    \"url\": \"git+https://github.com/JimLiu/baoyu-skills.git\",\n    \"directory\": \"packages/baoyu-fetch\"\n  },\n  \"bugs\": {\n    \"url\": \"https://github.com/JimLiu/baoyu-skills/issues\"\n  },\n  \"homepage\": \"https://github.com/JimLiu/baoyu-skills/tree/main/packages/baoyu-fetch#readme\",\n  \"publishConfig\": {\n    \"access\": \"public\"\n  },\n  \"scripts\": {\n    \"build\": \"rm -rf dist && bun build ./src/cli.ts --target bun --outfile ./dist/cli.js && chmod +x ./dist/cli.js\",\n    \"check\": \"tsc --noEmit\",\n    \"dev\": \"bun run ./src/cli.ts\",\n    \"release\": \"changeset publish\",\n    \"test\": \"bun test\",\n    \"version-packages\": \"changeset version\"\n  },\n  \"engines\": {\n    \"bun\": \">=1.2.0\"\n  },\n  \"dependencies\": {\n    \"@mozilla/readability\": \"^0.6.0\",\n    \"chrome-launcher\": \"^1.2.1\",\n    \"defuddle\": \"^0.14.0\",\n    \"jsdom\": \"^26.0.0\",\n    \"remark-gfm\": \"^4.0.1\",\n    \"remark-parse\": \"^11.0.0\",\n    \"turndown\": \"^7.2.0\",\n    \"turndown-plugin-gfm\": \"^1.0.2\",\n    \"unified\": \"^11.0.5\",\n    \"ws\": \"^8.18.3\"\n  },\n  \"devDependencies\": {\n    \"@changesets/cli\": \"^2.30.0\",\n    \"@types/bun\": \"^1.2.23\",\n    \"@types/jsdom\": \"^21.1.7\",\n    \"@types/ws\": \"^8.18.1\",\n    \"typescript\": \"^5.9.2\"\n  }\n}\n\nFile v1.82.2:scripts/vendor/baoyu-fetch/README.zh-CN.md\n\n# baoyu-fetch\n\n[English](./README.md) | 简体中文 | [更新日志](./CHANGELOG.zh-CN.md) | [English Changelog](./CHANGELOG.md)\n\n`baoyu-fetch` 是一个基于 Chrome CDP 的 Bun CLI。输入 URL，它会输出高质量\n`markdown` 或 `json`；命中站点 adapter 时优先消费 API 返回或页面内结构化\n数据，未命中时回退到通用 HTML 提取。\n\n## 当前能力\n\n- 通过 Chrome CDP 抓取渲染后的页面内容\n- 监听网络请求与响应，按需拉取响应体\n- adapter registry，支持按 URL 自动命中站点处理器\n- 内置 `x`、`youtube`、`hn` adapters\n- 通用 fallback：Defuddle 优先，Readability + HTML to Markdown 回退；`--format markdown` 时会再尝试 `defuddle.md` 兜底\n- `stdout` 或 `--output` 输出 `markdown` / `json`\n- 可选下载提取出的图片/视频并重写 Markdown 链接\n- 提供登录/验证场景下的交互等待模式\n- Chrome profile 默认对齐 `baoyu-skills/chrome-profile`\n\n## 安装\n\n```bash\nbun install\n```\n\n作为包使用时，推荐直接这样运行：\n\n```bash\nbunx baoyu-fetch https://example.com\n```\n\n也可以全局安装：\n\n```bash\nnpm install -g baoyu-fetch\n```\n\nnpm 包发布的是 TypeScript 源码入口，不包含预编译的 `dist`，所以运行时需要\nBun。\n\n## 用法\n\n```bash\nbun run src/cli.ts https://example.com\nbunx baoyu-fetch https://example.com\nbaoyu-fetch https://example.com\nbaoyu-fetch https://example.com --format markdown --output article.md\nbaoyu-fetch https://example.com --format markdown --output article.md --download-media\nbaoyu-fetch https://x.com/jack/status/20 --format json --output article.json\nbaoyu-fetch https://x.com/jack/status/20 --json\nbaoyu-fetch https://x.com/jack/status/20 --wait-for interaction\nbaoyu-fetch https://x.com/jack/status/20 --wait-for force\nbaoyu-fetch https://x.com/jack/status/20 --chrome-profile-dir ~/Library/Application\\\\ Support/baoyu-skills/chrome-profile\n```\n\n## 主要参数\n\n```bash\nbaoyu-fetch <url> [options]\n\nOptions:\n  --output <file>       保存输出内容到文件\n  --format <type>       输出格式：markdown | json\n  --json                `--format json` 的兼容别名\n  --adapter <name>      强制使用指定 adapter（如 x / hn / generic）\n  --download-media      下载 adapter 返回的媒体到 ./imgs 和 ./videos，并重写 markdown 链接\n  --media-dir <dir>     指定媒体下载根目录；默认使用输出文件所在目录\n  --debug-dir <dir>     导出调试信息（html、document.json、network.json）\n  --cdp-url <url>       连接现有 Chrome 调试地址\n  --browser-path <path> 指定 Chrome 可执行文件\n  --chrome-profile-dir <path>\n                        指定 Chrome profile 目录。默认使用 BAOYU_CHROME_PROFILE_DIR，\n                        否则回退到 baoyu-skills/chrome-profile\n  --headless            启动临时 headless Chrome（未连现有实例时）\n  --wait-for <mode>     等待模式：interaction | force\n  --wait-for-interaction\n                        `--wait-for interaction` 的别名\n  --wait-for-login      `--wait-for interaction` 的别名\n  --interaction-timeout <ms>\n                        手动交互等待超时，默认 600000\n  --interaction-poll-interval <ms>\n                        等待期间的轮询间隔，默认 1500\n  --login-timeout <ms>  `--interaction-timeout` 的别名\n  --login-poll-interval <ms>\n                        `--interaction-poll-interval` 的别名\n  --timeout <ms>        页面加载超时，默认 30000\n  --help                显示帮助\n```\n\n## 设计\n\n核心链路：\n\n1. CLI 解析 URL 和选项\n2. 建立 CDP 会话并创建受控 tab\n3. 启动 `NetworkJournal` 收集所有请求/响应\n4. 由 adapter registry 匹配站点 adapter\n5. adapter 返回结构化 `ExtractedDocument`\n6. 没命中则走通用 HTML 提取\n7. 按请求输出 Markdown，或输出包含 `document` 和 `markdown` 的 JSON\n\n## 开发\n\n```bash\nbun run check\nbun run test\nbun run build\n```\n\n## 发版\n\n新增用户可见改动后，先添加一个 changeset：\n\n```bash\nbunx changeset\n```\n\n把生成的 `.changeset...","readmeExcerpt":"Skill: Baoyu Url To Markdown Owner: jimliu Summary: Fetch any URL and convert to markdown using baoyu-fetch CLI (Chrome CDP with site-specific adapters). Built-in adapters for X/Twitter, YouTube transcripts, H... Tags: latest:1.117.2 Version history: v1.117.2 | 2026-05-18T02:16:39.623Z | user 1.117.2 - 2026-05-17 Documentation - baoyu-cover-image: ban programmatic text repair on generated bitmaps — disallow ImageMagi","codeSnippets":[],"executableExamples":[{"language":"bash","snippet":"# Default: headless capture, markdown to stdout\n${READER} <url>\n\n# Save to file\n${READER} <url> --output article.md\n\n# Save with media download\n${READER} <url> --output article.md --download-media\n\n# Wait for interaction (login/CAPTCHA) — auto-detect and continue\n${READER} <url> --wait-for interaction --output article.md\n\n# Wait for interaction — manual control (Enter to continue)\n${READER} <url> --wait-for force --output article.md\n\n# JSON output\n${READER} <url> --format json --output article.json\n\n# Force specific adapter\n${READER} <url> --adapter youtube --output transcript.md"},{"language":"text","snippet":"No EXTEND.md found\n        |\n        v\n+---------------------+\n| AskUserQuestion     |\n| (all questions)     |\n+---------------------+\n        |\n        v\n+---------------------+\n| Create EXTEND.md    |\n+---------------------+\n        |\n        v\n    Continue conversion"},{"language":"yaml","snippet":"header: \"Media\"\nquestion: \"How to handle images and videos in pages?\"\noptions:\n  - label: \"Ask each time (Recommended)\"\n    description: \"After saving markdown, ask whether to download media\"\n  - label: \"Always download\"\n    description: \"Always download media to local imgs/ and videos/ directories\"\n  - label: \"Never download\"\n    description: \"Keep original remote URLs in markdown\""},{"language":"yaml","snippet":"header: \"Output\"\nquestion: \"Default output directory?\"\noptions:\n  - label: \"url-to-markdown (Recommended)\"\n    description: \"Save to ./url-to-markdown/{domain}/{slug}.md\""},{"language":"yaml","snippet":"header: \"Save\"\nquestion: \"Where to save preferences?\"\noptions:\n  - label: \"User (Recommended)\"\n    description: \"~/.baoyu-skills/ (all projects)\"\n  - label: \"Project\"\n    description: \".baoyu-skills/ (this project only)\""},{"language":"md","snippet":"download_media: [ask/1/0]\ndefault_output_dir: [path or empty]"}],"parameters":null,"dependencies":[],"permissions":[],"extractedFiles":[{"path":"SKILL.md","content":"---\nname: baoyu-url-to-markdown\ndescription: Fetch any URL and convert to markdown using baoyu-fetch CLI (Chrome CDP with site-specific adapters). Built-in adapters for X/Twitter, YouTube transcripts, Hacker News threads, and generic pages via Defuddle. Handles login/CAPTCHA via interaction wait modes. Use when user wants to save a webpage as markdown.\nversion: 1.61.0\nmetadata:\n  openclaw:\n    homepage: https://github.com/JimLiu/baoyu-skills#baoyu-url-to-markdown\n    requires:\n      anyBins:\n        - bun\n---\n\n# URL to Markdown\n\nFetches any URL via `baoyu-fetch` CLI (Chrome CDP + site-specific adapters) and converts it to clean markdown.\n\n## User Input Tools\n\nWhen this skill prompts the user, follow this tool-selection rule (priority order):\n\n1. **Prefer built-in user-input tools** exposed by the current agent runtime — e.g., `AskUserQuestion`, `request_user_input`, `clarify`, `ask_user`, or any equivalent.\n2. **Fallback**: if no such tool exists, emit a numbered plain-text message and ask the user to reply with the chosen number/answer for each question.\n3. **Batching**: if the tool supports multiple questions per call, combine all applicable questions into a single call; if only single-question, ask them one at a time in priority order.\n\nConcrete `AskUserQuestion` references below are examples — substitute the local equivalent in other runtimes.\n\n## CLI Setup\n\n**Important**: The CLI source is vendored in `{baseDir}/scripts/lib`. `scripts/package.json` installs only third-party runtime dependencies.\n\n**Agent Execution Instructions**:\n1. Determine this SKILL.md file's directory path as `{baseDir}`\n2. Resolve `${BUN}` runtime: if `bun` installed → `bun`; else suggest installing Bun\n3. If `{baseDir}/scripts/node_modules` does not exist, run `${BUN} install --cwd {baseDir}/scripts`\n4. `${READER}` = `{baseDir}/scripts/baoyu-fetch`\n5. Replace all `${READER}` in this document with the resolved value\n\n## Preferences (EXTEND.md)\n\nCheck EXTEND.md in priority order — the first one found wins:\n\n| Priority | Path | Scope |\n|----------|------|-------|\n| 1 | `.baoyu-skills/baoyu-url-to-markdown/EXTEND.md` | Project |\n| 2 | `${XDG_CONFIG_HOME:-$HOME/.config}/baoyu-skills/baoyu-url-to-markdown/EXTEND.md` | XDG |\n| 3 | `$HOME/.baoyu-skills/baoyu-url-to-markdown/EXTEND.md` | User home |\n\n| Result | Action |\n|--------|--------|\n| Found | Read, parse, apply settings |\n| Not found | **MUST** run first-time setup (see below) — do NOT silently create defaults |\n\n**EXTEND.md supports**: download media by default, default output directory.\n\n### First-Time Setup ⛔ BLOCKING\n\nWhen EXTEND.md is not found, you **MUST** use `AskUserQuestion` to gather preferences before creating EXTEND.md. **NEVER** create EXTEND.md with silent defaults. Generation is BLOCKED until setup completes. Batch all three questions into a single call:\n\n- **Q1 — Media** (header \"Media\"): \"How to handle images and videos in pages?\"\n  - \"Ask each time (Recommended)\" — Prompt after each save\n  - \"Always d"},{"path":"_meta.json","content":"{\n  \"ownerId\": \"kn7csrrndw79hpke5d0gsnx93d82k67r\",\n  \"slug\": \"baoyu-url-to-markdown\",\n  \"version\": \"1.117.2\",\n  \"publishedAt\": 1779070599623\n}"},{"path":"references/adapters.md","content":"# Adapters & Media\n\nRead when choosing an adapter, handling media, or answering adapter-specific questions.\n\n## Built-in Adapters\n\n| Adapter | URLs | Key Features |\n|---------|------|-------------|\n| `x` | x.com, twitter.com | Tweets, threads, X Articles, media, login detection |\n| `youtube` | youtube.com, youtu.be | Transcript/captions, chapters, cover image, metadata |\n| `hn` | news.ycombinator.com | Threaded comments, story metadata, nested replies |\n| `generic` | Any URL (fallback) | Defuddle extraction, Readability fallback, auto-scroll, network idle detection |\n\nAdapter is auto-selected based on URL. Override with `--adapter <name>`.\n\n### YouTube\n\n- Extracts transcripts/captions when available\n- Transcript format: `[MM:SS] Text segment` with chapter headings\n- Availability depends on YouTube exposing a caption track; videos with captions disabled or restricted playback may produce description-only output\n- Use `--wait-for force` if the page needs time to finish loading player metadata\n\n### X/Twitter\n\n- Extracts single tweets, threads, and X Articles\n- Auto-detects login state; if logged out and content requires auth, JSON output shows `\"status\": \"needs_interaction\"`\n- Use `--wait-for interaction` for login-protected content\n\n### Hacker News\n\n- Parses threaded comments with proper nesting and reply hierarchy\n- Includes story metadata (title, URL, author, score, comment count)\n- Shows comment deletion/dead status\n\n## Media Download Workflow\n\nDriven by `download_media` in EXTEND.md:\n\n| Setting | Behavior |\n|---------|----------|\n| `1` (always) | Run CLI with `--download-media --output <path>` |\n| `0` (never) | Run CLI with `--output <path>` (no media download) |\n| `ask` (default) | Follow the ask-each-time flow below |\n\n### Ask-Each-Time Flow\n\n1. Run the CLI **without** `--download-media` with `--output <path>` → markdown saved\n2. Check the saved markdown for remote media URLs (`https://` in image/video links)\n3. **If no remote media found** → done, no prompt needed\n4. **If remote media found** → ask via `AskUserQuestion`:\n   - header: \"Media\", question: \"Download N images/videos to local files?\"\n   - \"Yes\" — Download to local directories\n   - \"No\" — Keep remote URLs\n5. If the user confirms → run the CLI **again** with `--download-media --output <same-path>` (overwrites markdown with localized links)\n\n### Media Layout\n\nWhen `--download-media` is enabled:\n\n- Images → `imgs/` next to the output file (or `--media-dir`)\n- Videos → `videos/` next to the output file (or `--media-dir`)\n- Markdown media links are rewritten to local relative paths\n\n## Output Format\n\nMarkdown to stdout (or file with `--output`).\n\nJSON output (`--format json`) returns structured data:\n\n- `adapter` — which adapter handled the URL\n- `status` — `\"ok\"` or `\"needs_interaction\"`\n- `login` — login state detection (`logged_in`, `logged_out`, `unknown`)\n- `interaction` — interaction gate details (kind, provider, prompt)\n- `document` — structured content (url, title, author, publi"},{"path":"references/config/first-time-setup.md","content":"---\nname: first-time-setup\ndescription: First-time setup flow for baoyu-url-to-markdown preferences\n---\n\n# First-Time Setup\n\n## Overview\n\nWhen no EXTEND.md is found, guide user through preference setup.\n\n**BLOCKING OPERATION**: This setup MUST complete before ANY other workflow steps. Do NOT:\n- Start converting URLs\n- Ask about URLs or output paths\n- Proceed to any conversion\n\nONLY ask the questions in this setup flow, save EXTEND.md, then continue.\n\n## Setup Flow\n\n```\nNo EXTEND.md found\n        |\n        v\n+---------------------+\n| AskUserQuestion     |\n| (all questions)     |\n+---------------------+\n        |\n        v\n+---------------------+\n| Create EXTEND.md    |\n+---------------------+\n        |\n        v\n    Continue conversion\n```\n\n## Questions\n\n**Language**: Use user's input language or saved language preference.\n\nUse AskUserQuestion with ALL questions in ONE call:\n\n### Question 1: Download Media\n\n```yaml\nheader: \"Media\"\nquestion: \"How to handle images and videos in pages?\"\noptions:\n  - label: \"Ask each time (Recommended)\"\n    description: \"After saving markdown, ask whether to download media\"\n  - label: \"Always download\"\n    description: \"Always download media to local imgs/ and videos/ directories\"\n  - label: \"Never download\"\n    description: \"Keep original remote URLs in markdown\"\n```\n\n### Question 2: Default Output Directory\n\n```yaml\nheader: \"Output\"\nquestion: \"Default output directory?\"\noptions:\n  - label: \"url-to-markdown (Recommended)\"\n    description: \"Save to ./url-to-markdown/{domain}/{slug}.md\"\n```\n\nNote: User will likely choose \"Other\" to type a custom path.\n\n### Question 3: Save Location\n\n```yaml\nheader: \"Save\"\nquestion: \"Where to save preferences?\"\noptions:\n  - label: \"User (Recommended)\"\n    description: \"~/.baoyu-skills/ (all projects)\"\n  - label: \"Project\"\n    description: \".baoyu-skills/ (this project only)\"\n```\n\n## Save Locations\n\n| Choice | Path | Scope |\n|--------|------|-------|\n| User | `~/.baoyu-skills/baoyu-url-to-markdown/EXTEND.md` | All projects |\n| Project | `.baoyu-skills/baoyu-url-to-markdown/EXTEND.md` | Current project |\n\n## After Setup\n\n1. Create directory if needed\n2. Write EXTEND.md\n3. Confirm: \"Preferences saved to [path]\"\n4. Continue with conversion using saved preferences\n\n## EXTEND.md Template\n\n```md\ndownload_media: [ask/1/0]\ndefault_output_dir: [path or empty]\n```\n\n## Modifying Preferences Later\n\nUsers can edit EXTEND.md directly or delete it to trigger setup again."},{"path":"references/quality-gate.md","content":"# Quality Gate & Recovery\n\nHeadless Chrome can silently return low-quality content — layout shells, login walls, or framework payloads — without the CLI returning a non-zero exit code. Read this after every headless run so you can catch and recover from those cases.\n\n## Checks the Agent Must Run\n\n1. Confirm the markdown title matches the target page, not a generic site shell\n2. Confirm the body contains the expected article/page content, not just navigation, footer, or a generic error\n3. Watch for obvious failure signs:\n   - `Application error`\n   - `This page could not be found`\n   - Login, signup, subscribe, or verification shells\n   - Extremely short markdown for a page that should be long-form\n   - Raw framework payloads or mostly boilerplate content\n4. Do NOT accept a run as successful just because the CLI exited `0`\n\n**Tip**: run with `--format json` to get structured signals including `status`, `login.state`, and `interaction`. `\"status\": \"needs_interaction\"` means the page requires manual interaction.\n\n## Recovery Workflow\n\n1. Start headless (default) unless there is already a clear reason to use interaction mode\n2. Review markdown quality immediately after the run\n3. If the content is low quality or indicates login/CAPTCHA:\n   - `--wait-for interaction` for auto-detected gates (login, CAPTCHA, Cloudflare)\n   - `--wait-for force` when the page needs manual browsing, scroll loading, or complex interaction\n4. If `--wait-for` is used, tell the user exactly what to do:\n   - Login required → sign in in the browser\n   - CAPTCHA visible → solve it\n   - Slow loading → wait until content is visible\n   - `--wait-for force` → press Enter when ready\n5. If JSON output shows `\"status\": \"needs_interaction\"`, switch to `--wait-for interaction` automatically\n\n## Capture Modes\n\n| Mode | Behavior | Use When |\n|------|----------|----------|\n| Default | Headless Chrome, auto-extract on network idle | Public pages, static content |\n| `--headless` | Explicit headless (same as default) | Clarify intent |\n| `--wait-for interaction` | Opens visible Chrome, auto-detects login/CAPTCHA gates, waits for them to clear, then continues | Login-required, CAPTCHA-protected |\n| `--wait-for force` | Opens visible Chrome, auto-detects OR accepts Enter keypress to continue | Complex flows, lazy loading, paywalls |\n\n**Interaction gate auto-detection**: Cloudflare Turnstile / \"just a moment\" pages, Google reCAPTCHA, hCaptcha, custom challenge / verification screens."}],"languages":[],"docsSourceLabel":"CLAWHUB","editorialOverview":null,"editorialQuality":{"score":100,"threshold":65,"status":"thin","wordCount":2375,"uniquenessScore":40,"reasons":["uniqueness-below-45"]}},"media":{"evidence":{"source":"no-media","verified":false,"confidence":"low","updatedAt":"2026-10-09T05:42:36.985Z","emptyReason":"No screenshots, media assets, or demo links are available."},"primaryImageUrl":null,"mediaAssetCount":0,"assets":[],"demoUrl":null},"ownerResources":{"evidence":{"source":"unclaimed","verified":false,"confidence":"low","updatedAt":"2026-10-09T05:42:36.985Z","emptyReason":"This page has not been claimed by the agent owner."},"hasCustomPage":false,"customPageUpdatedAt":null,"customLinks":[],"structuredLinks":{"docsUrl":null,"demoUrl":null,"supportUrl":null,"pricingUrl":null,"statusUrl":null},"customPage":null},"relatedAgents":{"evidence":{"source":"protocol-neighbors","verified":false,"confidence":"medium","updatedAt":"2026-10-10T03:54:22.805Z","emptyReason":null},"items":[{"id":"8ebccd8e-3863-4187-8355-c3f14e1f9edf","entityType":"agent","canonicalPath":"/agent/iofficeai-aionui","slug":"iofficeai-aionui","name":"AionUi","description":"Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!","url":"https://github.com/iOfficeAI/AionUi","homepage":"https://www.aionui.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-10-09T19:11:12.944Z","createdAt":"2026-02-25T03:38:16.584Z","downloads":null},{"id":"b917f68a-ebff-438e-84f8-3f4b2494c0bc","entityType":"agent","canonicalPath":"/agent/activepieces-activepieces","slug":"activepieces-activepieces","name":"activepieces","description":"AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents","url":"https://github.com/activepieces/activepieces","homepage":"https://www.activepieces.com","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-15T02:22:12.426Z","createdAt":"2026-02-25T03:38:12.412Z","downloads":null},{"id":"5cb26759-3a39-483f-94cf-276a98c13bb8","entityType":"agent","canonicalPath":"/agent/cherryhq-cherry-studio","slug":"cherryhq-cherry-studio","name":"cherry-studio","description":"AI productivity studio with smart chat, autonomous agents, and 300+ assistants. Unified access to frontier LLMs","url":"https://github.com/CherryHQ/cherry-studio","homepage":"https://cherry-ai.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-11T14:38:40.986Z","createdAt":"2026-02-25T03:38:19.379Z","downloads":null},{"id":"6f6582d0-5d76-4f0f-b81d-86520247950b","entityType":"agent","canonicalPath":"/agent/copilotkit-copilotkit","slug":"copilotkit-copilotkit","name":"CopilotKit","description":"The Frontend for Agents & Generative UI. React + Angular","url":"https://github.com/CopilotKit/CopilotKit","homepage":"https://docs.copilotkit.ai","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-03-25T09:50:57.846Z","createdAt":"2026-02-25T03:39:14.617Z","downloads":null}],"links":{"hub":"/agent","source":"/agent/source/clawhub","protocols":[{"label":"OpenClaw","href":"/agent/protocol/openclew"}]}}}