Baoyu Url To Markdown
Fetch any URL and convert to markdown using baoyu-fetch CLI (Chrome CDP with site-specific adapters). Built-in adapters for X/Twitter, YouTube transcripts, H...
Rank
62
Safety
84
Downloads
4.3k
Updated
Oct 9, 2026
Version
1.117.2
Source
CLAWHUB
About
What it does, and when to use it.
Capability contract not published. No trust telemetry is available yet. 4.3K downloads reported by the source. Last updated 10/9/2026.
Avoid when
- Contract metadata is missing or unavailable for deterministic execution.
Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing
Public facts
Every fact links back to the source it came from.
- Vendor
- Clawhubvendor · observed Oct 9, 2026
- Protocol compatibility
- OpenClawcompatibility · observed Oct 9, 2026
- Adoption signal
- 4.3K downloadsadoption · observed Oct 9, 2026
- Latest release
- 1.117.2release · observed May 18, 2026
- Handshake status
- UNKNOWNsecurity
Install and run
Setup complexity: low.
clawhub skill install s17dfbhg0khk4fvtnx0stqg0yx83j25z:baoyu-url-to-markdown- Install using `clawhub skill install s17dfbhg0khk4fvtnx0stqg0yx83j25z:baoyu-url-to-markdown` in an isolated environment before connecting it to live workloads.
- No published capability contract is available yet, so validate auth and request/response behavior manually.
- Review the upstream CLAWHUB listing at https://clawhub.ai/jimliu/baoyu-url-to-markdown before using production credentials.
Contract: missing
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-jimliu-baoyu-url-to-markdown/snapshot"
Documentation
CLAWHUB
160,000 characters of source documentation, loaded on request.
Extracted files
5 files captured from the source.
SKILL.md
---
name: baoyu-url-to-markdown
description: Fetch any URL and convert to markdown using baoyu-fetch CLI (Chrome CDP with site-specific adapters). Built-in adapters for X/Twitter, YouTube transcripts, Hacker News threads, and generic pages via Defuddle. Handles login/CAPTCHA via interaction wait modes. Use when user wants to save a webpage as markdown.
version: 1.61.0
metadata:
openclaw:
homepage: https://github.com/JimLiu/baoyu-skills#baoyu-url-to-markdown
requires:
anyBins:
- bun
---
# URL to Markdown
Fetches any URL via `baoyu-fetch` CLI (Chrome CDP + site-specific adapters) and converts it to clean markdown.
## User Input Tools
When this skill prompts the user, follow this tool-selection rule (priority order):
1. **Prefer built-in user-input tools** exposed by the current agent runtime — e.g., `AskUserQuestion`, `request_user_input`, `clarify`, `ask_user`, or any equivalent.
2. **Fallback**: if no such tool exists, emit a numbered plain-text message and ask the user to reply with the chosen number/answer for each question.
3. **Batching**: if the tool supports multiple questions per call, combine all applicable questions into a single call; if only single-question, ask them one at a time in priority order.
Concrete `AskUserQuestion` references below are examples — substitute the local equivalent in other runtimes.
## CLI Setup
**Important**: The CLI source is vendored in `{baseDir}/scripts/lib`. `scripts/package.json` installs only third-party runtime dependencies.
**Agent Execution Instructions**:
1. Determine this SKILL.md file's directory path as `{baseDir}`
2. Resolve `${BUN}` runtime: if `bun` installed → `bun`; else suggest installing Bun
3. If `{baseDir}/scripts/node_modules` does not exist, run `${BUN} install --cwd {baseDir}/scripts`
4. `${READER}` = `{baseDir}/scripts/baoyu-fetch`
5. Replace all `${READER}` in this document with the resolved value
## Preferences (EXTEND.md)
Check EXTEND.md in priority order — the first one found wins:
| Priority | Path | Scope |
|----------|------|-------|
| 1 | `.baoyu-skills/baoyu-url-to-markdown/EXTEND.md` | Project |
| 2 | `${XDG_CONFIG_HOME:-$HOME/.config}/baoyu-skills/baoyu-url-to-markdown/EXTEND.md` | XDG |
| 3 | `$HOME/.baoyu-skills/baoyu-url-to-markdown/EXTEND.md` | User home |
| Result | Action |
|--------|--------|
| Found | Read, parse, apply settings |
| Not found | **MUST** run first-time setup (see below) — do NOT silently create defaults |
**EXTEND.md supports**: download media by default, default output directory.
### First-Time Setup ⛔ BLOCKING
When EXTEND.md is not found, you **MUST** use `AskUserQuestion` to gather preferences before creating EXTEND.md. **NEVER** create EXTEND.md with silent defaults. Generation is BLOCKED until setup completes. Batch all three questions into a single call:
- **Q1 — Media** (header "Media"): "How to handle images and videos in pages?"
- "Ask each time (Recommended)" — Prompt after each save
- "Always d_meta.json
{
"ownerId": "kn7csrrndw79hpke5d0gsnx93d82k67r",
"slug": "baoyu-url-to-markdown",
"version": "1.117.2",
"publishedAt": 1779070599623
}references/adapters.md
# Adapters & Media Read when choosing an adapter, handling media, or answering adapter-specific questions. ## Built-in Adapters | Adapter | URLs | Key Features | |---------|------|-------------| | `x` | x.com, twitter.com | Tweets, threads, X Articles, media, login detection | | `youtube` | youtube.com, youtu.be | Transcript/captions, chapters, cover image, metadata | | `hn` | news.ycombinator.com | Threaded comments, story metadata, nested replies | | `generic` | Any URL (fallback) | Defuddle extraction, Readability fallback, auto-scroll, network idle detection | Adapter is auto-selected based on URL. Override with `--adapter <name>`. ### YouTube - Extracts transcripts/captions when available - Transcript format: `[MM:SS] Text segment` with chapter headings - Availability depends on YouTube exposing a caption track; videos with captions disabled or restricted playback may produce description-only output - Use `--wait-for force` if the page needs time to finish loading player metadata ### X/Twitter - Extracts single tweets, threads, and X Articles - Auto-detects login state; if logged out and content requires auth, JSON output shows `"status": "needs_interaction"` - Use `--wait-for interaction` for login-protected content ### Hacker News - Parses threaded comments with proper nesting and reply hierarchy - Includes story metadata (title, URL, author, score, comment count) - Shows comment deletion/dead status ## Media Download Workflow Driven by `download_media` in EXTEND.md: | Setting | Behavior | |---------|----------| | `1` (always) | Run CLI with `--download-media --output <path>` | | `0` (never) | Run CLI with `--output <path>` (no media download) | | `ask` (default) | Follow the ask-each-time flow below | ### Ask-Each-Time Flow 1. Run the CLI **without** `--download-media` with `--output <path>` → markdown saved 2. Check the saved markdown for remote media URLs (`https://` in image/video links) 3. **If no remote media found** → done, no prompt needed 4. **If remote media found** → ask via `AskUserQuestion`: - header: "Media", question: "Download N images/videos to local files?" - "Yes" — Download to local directories - "No" — Keep remote URLs 5. If the user confirms → run the CLI **again** with `--download-media --output <same-path>` (overwrites markdown with localized links) ### Media Layout When `--download-media` is enabled: - Images → `imgs/` next to the output file (or `--media-dir`) - Videos → `videos/` next to the output file (or `--media-dir`) - Markdown media links are rewritten to local relative paths ## Output Format Markdown to stdout (or file with `--output`). JSON output (`--format json`) returns structured data: - `adapter` — which adapter handled the URL - `status` — `"ok"` or `"needs_interaction"` - `login` — login state detection (`logged_in`, `logged_out`, `unknown`) - `interaction` — interaction gate details (kind, provider, prompt) - `document` — structured content (url, title, author, publi
references/config/first-time-setup.md
---
name: first-time-setup
description: First-time setup flow for baoyu-url-to-markdown preferences
---
# First-Time Setup
## Overview
When no EXTEND.md is found, guide user through preference setup.
**BLOCKING OPERATION**: This setup MUST complete before ANY other workflow steps. Do NOT:
- Start converting URLs
- Ask about URLs or output paths
- Proceed to any conversion
ONLY ask the questions in this setup flow, save EXTEND.md, then continue.
## Setup Flow
```
No EXTEND.md found
|
v
+---------------------+
| AskUserQuestion |
| (all questions) |
+---------------------+
|
v
+---------------------+
| Create EXTEND.md |
+---------------------+
|
v
Continue conversion
```
## Questions
**Language**: Use user's input language or saved language preference.
Use AskUserQuestion with ALL questions in ONE call:
### Question 1: Download Media
```yaml
header: "Media"
question: "How to handle images and videos in pages?"
options:
- label: "Ask each time (Recommended)"
description: "After saving markdown, ask whether to download media"
- label: "Always download"
description: "Always download media to local imgs/ and videos/ directories"
- label: "Never download"
description: "Keep original remote URLs in markdown"
```
### Question 2: Default Output Directory
```yaml
header: "Output"
question: "Default output directory?"
options:
- label: "url-to-markdown (Recommended)"
description: "Save to ./url-to-markdown/{domain}/{slug}.md"
```
Note: User will likely choose "Other" to type a custom path.
### Question 3: Save Location
```yaml
header: "Save"
question: "Where to save preferences?"
options:
- label: "User (Recommended)"
description: "~/.baoyu-skills/ (all projects)"
- label: "Project"
description: ".baoyu-skills/ (this project only)"
```
## Save Locations
| Choice | Path | Scope |
|--------|------|-------|
| User | `~/.baoyu-skills/baoyu-url-to-markdown/EXTEND.md` | All projects |
| Project | `.baoyu-skills/baoyu-url-to-markdown/EXTEND.md` | Current project |
## After Setup
1. Create directory if needed
2. Write EXTEND.md
3. Confirm: "Preferences saved to [path]"
4. Continue with conversion using saved preferences
## EXTEND.md Template
```md
download_media: [ask/1/0]
default_output_dir: [path or empty]
```
## Modifying Preferences Later
Users can edit EXTEND.md directly or delete it to trigger setup again.references/quality-gate.md
# Quality Gate & Recovery Headless Chrome can silently return low-quality content — layout shells, login walls, or framework payloads — without the CLI returning a non-zero exit code. Read this after every headless run so you can catch and recover from those cases. ## Checks the Agent Must Run 1. Confirm the markdown title matches the target page, not a generic site shell 2. Confirm the body contains the expected article/page content, not just navigation, footer, or a generic error 3. Watch for obvious failure signs: - `Application error` - `This page could not be found` - Login, signup, subscribe, or verification shells - Extremely short markdown for a page that should be long-form - Raw framework payloads or mostly boilerplate content 4. Do NOT accept a run as successful just because the CLI exited `0` **Tip**: run with `--format json` to get structured signals including `status`, `login.state`, and `interaction`. `"status": "needs_interaction"` means the page requires manual interaction. ## Recovery Workflow 1. Start headless (default) unless there is already a clear reason to use interaction mode 2. Review markdown quality immediately after the run 3. If the content is low quality or indicates login/CAPTCHA: - `--wait-for interaction` for auto-detected gates (login, CAPTCHA, Cloudflare) - `--wait-for force` when the page needs manual browsing, scroll loading, or complex interaction 4. If `--wait-for` is used, tell the user exactly what to do: - Login required → sign in in the browser - CAPTCHA visible → solve it - Slow loading → wait until content is visible - `--wait-for force` → press Enter when ready 5. If JSON output shows `"status": "needs_interaction"`, switch to `--wait-for interaction` automatically ## Capture Modes | Mode | Behavior | Use When | |------|----------|----------| | Default | Headless Chrome, auto-extract on network idle | Public pages, static content | | `--headless` | Explicit headless (same as default) | Clarify intent | | `--wait-for interaction` | Opens visible Chrome, auto-detects login/CAPTCHA gates, waits for them to clear, then continues | Login-required, CAPTCHA-protected | | `--wait-for force` | Opens visible Chrome, auto-detects OR accepts Enter keypress to continue | Complex flows, lazy loading, paywalls | **Interaction gate auto-detection**: Cloudflare Turnstile / "just a moment" pages, Google reCAPTCHA, hCaptcha, custom challenge / verification screens.
AionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
activepieces
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
CopilotKit
The Frontend for Agents & Generative UI. React + Angular
Machine-readable data
The same record, as JSON, for agents and crawlers.
{
"facts": [
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Clawhub",
"href": "https://clawhub.ai/jimliu/skills/baoyu-url-to-markdown",
"sourceUrl": "https://clawhub.ai/jimliu/skills/baoyu-url-to-markdown",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-09T05:42:36.985Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-jimliu-baoyu-url-to-markdown/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-jimliu-baoyu-url-to-markdown/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-09T05:42:36.985Z",
"isPublic": true
},
{
"factKey": "traction",
"category": "adoption",
"label": "Adoption signal",
"value": "4.3K downloads",
"href": "https://clawhub.ai/jimliu/baoyu-url-to-markdown",
"sourceUrl": "https://clawhub.ai/jimliu/baoyu-url-to-markdown",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-09T05:42:36.985Z",
"isPublic": true
},
{
"factKey": "latest_release",
"category": "release",
"label": "Latest release",
"value": "1.117.2",
"href": "https://clawhub.ai/jimliu/baoyu-url-to-markdown",
"sourceUrl": "https://clawhub.ai/jimliu/baoyu-url-to-markdown",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-05-18T02:16:39.623Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-jimliu-baoyu-url-to-markdown/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-jimliu-baoyu-url-to-markdown/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
],
"events": [
{
"eventType": "release",
"title": "Release 1.117.2",
"description": "## 1.117.2 - 2026-05-17 ### Documentation - `baoyu-cover-image`: ban programmatic text repair on generated bitmaps — disallow ImageMagick / Pillow / Canvas / SVG / HTML overlays to cover, rewrite, or replace title/subtitle text; regenerate from a corrected prompt or switch to a lower-text or no-title variant instead - `baoyu-article-illustrator`, `baoyu-comic`, `baoyu-image-cards`, `baoyu-xhs-images`, `baoyu-infographic`, `baoyu-slide-deck`: sync the same text-repair ban with skill-specific text categories (labels/captions, dialogue/sound effects, titles/body/tags, headings/data values, slide titles/bullets)",
"href": "https://clawhub.ai/jimliu/baoyu-url-to-markdown",
"sourceUrl": "https://clawhub.ai/jimliu/baoyu-url-to-markdown",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-05-18T02:16:39.623Z",
"isPublic": true
}
]
}Record generated Oct 9, 2026.
