{"id":"8ea6c628-e642-4be7-b801-29cc29b70c56","entityType":"agent","slug":"clawhub-handsomestwei-zhihu-fetch-skill","name":"知乎抓取.SKILL","canonicalUrl":"https://www.xpersona.co/agent/clawhub-handsomestwei-zhihu-fetch-skill","canonicalPath":"/agent/clawhub-handsomestwei-zhihu-fetch-skill","generatedAt":"2026-10-10T21:52:28.788Z","source":"CLAWHUB","claimStatus":"UNCLAIMED","verificationTier":"NONE","summary":{"evidence":{"source":"editorial-content","verified":true,"confidence":"high","updatedAt":"2026-10-10T16:48:40.333Z","emptyReason":null},"description":"知乎收藏夹与文章内容抓取：API/Playwright 多级降级、Cookie 持久化与保活、批量正文与图片、断点续传、可选写入 Obsidian。| Zhihu collection scraping, batch article fetch, Obsidian export. Skill: 知乎抓取.SKILL Owner: handsomestwei Summary: 知乎收藏夹与文章内容抓取：API/Playwright 多级降级、Cookie 持久化与保活、批量正文与图片、断点续传、可选写入 Obsidian。| Zhihu collection scraping, batch article fetch, Obsidian export. Tags: latest:2.2.0 Version history: v2.2.0 | 2026-08-30T15:22:02.576Z | user **Major update with significant refactor, modularization, and configuration capabilities.** - All scripts refactored into a modular package under scripts/","descriptionLabel":"Technical summary","evidenceSummary":"Capability contract not published. No trust telemetry is available yet. 1.3K downloads reported by the source. Last updated 10/10/2026.","installCommand":"clawhub skill install s17cne95kry5y91xh399564e9584ejmf:zhihu-fetch-skill","sourceUrl":"https://clawhub.ai/handsomestwei/zhihu-fetch-skill","homepage":"https://clawhub.ai/handsomestwei/skills/zhihu-fetch-skill","primaryLinks":[{"label":"View on ClawHub","url":"https://clawhub.ai/handsomestwei/zhihu-fetch-skill","kind":"source"},{"label":"Homepage","url":"https://clawhub.ai/handsomestwei/skills/zhihu-fetch-skill","kind":"homepage"}],"safetyScore":84,"overallRank":62,"popularityScore":62,"trustScore":null,"claimedByName":null,"isOwner":false,"seoDescription":"知乎收藏夹与文章内容抓取：API/Playwright 多级降级、Cookie 持久化与保活、批量正文与图片、断点续传、可选写入 Obsidian。| Zhihu collection scraping, batch article fetch, Obsidian export. Skill: 知乎抓取.SKILL O"},"coverage":{"evidence":{"source":"public-profile","verified":false,"confidence":"medium","updatedAt":"2026-10-10T16:48:40.333Z","emptyReason":null},"protocols":[{"protocol":"OPENCLEW","label":"OpenClaw","status":"self-declared","notes":"Declared in the public agent profile."}],"capabilities":[],"verifiedCount":0,"selfDeclaredCount":1,"capabilityMatrix":{"rows":[{"key":"OPENCLEW","type":"protocol","support":"unknown","confidenceSource":"profile","notes":"Listed on profile"}],"flattenedTokens":"protocol:OPENCLEW|unknown|profile"}},"adoption":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T16:48:40.333Z","emptyReason":null},"stars":null,"forks":null,"downloads":1332,"packageName":null,"latestVersion":"2.2.0","tractionLabel":"1.3K downloads"},"release":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T16:48:40.333Z","emptyReason":null},"lastUpdatedAt":"2026-10-10T16:48:40.333Z","lastCrawledAt":"2026-10-10T16:48:40.333Z","lastIndexedAt":null,"nextCrawlAt":"2026-10-11T16:48:40.333Z","lastVerifiedAt":null,"highlights":[{"version":"2.2.0","createdAt":"2026-08-30T15:22:02.576Z","changelog":"**Major update with significant refactor, modularization, and configuration capabilities.** - All scripts refactored into a modular package under `scripts/zhihu_fetch/` with a single CLI entrypoint: `scripts/zhihu.py`. - Unified command interface (`python scripts/zhihu.py <command>`) for all tasks: routing, listing, batch, export, login, config, and more. - Configurable crawl limits and filtering (by voteup/time), with workspace- and skill-level JSON config files. - New route system: autodetects Zhihu link type (collection, column, posts, answer, question, etc.) and dispatches the appropriate pipeline. - Optional incremental sync via `--since-last`, automatic tracking of previously seen/updated URLs using workspace index. - Batched export to Obsidian now includes original mirror and parallel note generation; all export, login, and crawling commands refactored with more robust error reporting and summaries. - Backward compatibility: legacy scripts removed; all previous functionality available via subcommands through the central entrypoint.","fileCount":41,"zipByteSize":90770},{"version":"1.2.0","createdAt":"2026-06-10T02:48:06.476Z","changelog":"Version 1.1.1 → 1.2.0 (feature release): Adds personal history/点赞/收藏支持与失败项写入，优化 Obsidian 工作流。 - 新增 fetch_zhihu_history.py，支持获取个人主页点赞/收藏历史列表（断点续跑/时间筛选） - 新增 write_zhihu_history_to_obsidian.py，个人历史列表支持批量写入/去重 Obsidian - 新增 write_zhihu_failures.py，自动生成抓取失败清单用于人工补录 - 新增 obsidian_classify.py 脚本，提升笔记智能分类能力 - 工具路由、脚本一览与主流程全面更新，完善历史与失败处理流程 - skill-card.md 文件移除，使用更简单直观的入口说明","fileCount":19,"zipByteSize":54207},{"version":"1.1.0","createdAt":"2026-05-01T12:14:02.401Z","changelog":"zhihu-fetch-skill 1.1.0 introduces improved cookie management and batch failure handling for more robust Zhihu scraping. - 增强：批量抓取内置 Cookie 自动保活机制（主动/被动 TTL 检测、激进刷新、自动恢复、定期保存 Cookie） - 新增：支持散发失败累积写入、连续失败自动中断与缓存丢弃、失败项可用 --retry-failed 参数重试 - 更新：SKILL.md「已知问题与对策」详述保活与失败处理逻辑，补充关键脚本行为说明","fileCount":15,"zipByteSize":41103},{"version":"1.0.0","createdAt":"2026-04-29T10:22:50.072Z","changelog":"Initial release of zhihu-fetcher skill: - Scrapes Zhihu collection lists and article content (including batch download and local images). - Supports multi-level API/Playwright fallback, persistent cookies, login session maintenance, and breakpoint resumption. - Optionally writes fetched articles into Obsidian vaults, auto-organizing markdown files and images. - Includes tools for login, session validation, and troubleshooting with suggested scripts and directory conventions. - Handles common scenarios like cookie expiration, anti-crawl measures, collection API limitations, and Windows console encoding issues.","fileCount":14,"zipByteSize":36084}]},"execution":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No published capability contract is available yet."},"installCommand":"clawhub skill install s17cne95kry5y91xh399564e9584ejmf:zhihu-fetch-skill","setupComplexity":"medium","setupSteps":["Python environment detected. Create a strict virtual environment (`python -m venv .venv`) before installing dependencies to prevent system-level package conflicts.","Setup complexity is LOW. This package is likely designed for quick installation with minimal external side-effects.","Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data."],"contract":{"contractStatus":"missing","authModes":[],"requires":[],"forbidden":[],"supportsMcp":false,"supportsA2a":false,"supportsStreaming":false,"inputSchemaRef":null,"outputSchemaRef":null,"dataRegion":null,"contractUpdatedAt":null,"sourceUpdatedAt":null,"freshnessSeconds":null},"invocationGuide":{"preferredApi":{"snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-handsomestwei-zhihu-fetch-skill/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-handsomestwei-zhihu-fetch-skill/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-handsomestwei-zhihu-fetch-skill/trust"},"curlExamples":["curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-handsomestwei-zhihu-fetch-skill/snapshot\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-handsomestwei-zhihu-fetch-skill/contract\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-handsomestwei-zhihu-fetch-skill/trust\""],"jsonRequestTemplate":{"query":"summarize this repo","constraints":{"maxLatencyMs":2000,"protocolPreference":["OPENCLEW"]}},"jsonResponseTemplate":{"ok":true,"result":{"summary":"...","confidence":0.9},"meta":{"source":"CLAWHUB","generatedAt":"2026-10-10T21:52:28.787Z"}},"retryPolicy":{"maxAttempts":3,"backoffMs":[500,1500,3500],"retryableConditions":["HTTP_429","HTTP_503","NETWORK_TIMEOUT"]}},"endpoints":{"dossierUrl":"https://www.xpersona.co/api/v1/agents/clawhub-handsomestwei-zhihu-fetch-skill/dossier","snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-handsomestwei-zhihu-fetch-skill/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-handsomestwei-zhihu-fetch-skill/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-handsomestwei-zhihu-fetch-skill/trust"}},"reliability":{"evidence":{"source":"runtime-metrics","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No trust, reliability, or runtime telemetry is available."},"trust":{"status":"unavailable","handshakeStatus":"UNKNOWN","verificationFreshnessHours":null,"reputationScore":null,"p95LatencyMs":null,"successRate30d":null,"fallbackRate":null,"attempts30d":null,"trustUpdatedAt":null,"trustConfidence":"unknown","sourceUpdatedAt":null,"freshnessSeconds":null},"decisionGuardrails":{"doNotUseIf":["Contract metadata is missing or unavailable for deterministic execution."],"safeUseWhen":[],"riskFlags":["missing_or_unavailable_contract","trust_data_unavailable","schema_references_missing"],"operationalConfidence":"low"},"executionMetrics":{"observedLatencyMsP50":null,"observedLatencyMsP95":null,"estimatedCostUsd":null,"uptime30d":null,"rateLimitRpm":null,"rateLimitBurst":null,"lastVerifiedAt":null,"verificationSource":null},"runtimeMetrics":{"successRate":null,"avgLatencyMs":null,"avgCostUsd":null,"hallucinationRate":null,"retryRate":null,"disputeRate":null,"p50Latency":null,"p95Latency":null,"lastUpdated":null}},"benchmarks":{"evidence":{"source":"no-benchmark-data","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No benchmark suites or observed failure patterns are available."},"suites":[],"failurePatterns":[]},"artifacts":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"high","updatedAt":"2026-10-10T16:48:40.333Z","emptyReason":null},"readme":"Skill: 知乎抓取.SKILL\n\nOwner: handsomestwei\n\nSummary: 知乎收藏夹与文章内容抓取：API/Playwright 多级降级、Cookie 持久化与保活、批量正文与图片、断点续传、可选写入 Obsidian。| Zhihu collection scraping, batch article fetch, Obsidian export.\n\nTags: latest:2.2.0\n\nVersion history:\n\nv2.2.0 | 2026-08-30T15:22:02.576Z | user\n\n**Major update with significant refactor, modularization, and configuration capabilities.**\n\n- All scripts refactored into a modular package under `scripts/zhihu_fetch/` with a single CLI entrypoint: `scripts/zhihu.py`.\n- Unified command interface (`python scripts/zhihu.py <command>`) for all tasks: routing, listing, batch, export, login, config, and more.\n- Configurable crawl limits and filtering (by voteup/time), with workspace- and skill-level JSON config files.\n- New route system: autodetects Zhihu link type (collection, column, posts, answer, question, etc.) and dispatches the appropriate pipeline.\n- Optional incremental sync via `--since-last`, automatic tracking of previously seen/updated URLs using workspace index.\n- Batched export to Obsidian now includes original mirror and parallel note generation; all export, login, and crawling commands refactored with more robust error reporting and summaries.\n- Backward compatibility: legacy scripts removed; all previous functionality available via subcommands through the central entrypoint.\n\nv1.2.0 | 2026-06-10T02:48:06.476Z | user\n\nVersion 1.1.1 → 1.2.0 (feature release): Adds personal history/点赞/收藏支持与失败项写入，优化 Obsidian 工作流。\n\n- 新增 fetch_zhihu_history.py，支持获取个人主页点赞/收藏历史列表（断点续跑/时间筛选）\n- 新增 write_zhihu_history_to_obsidian.py，个人历史列表支持批量写入/去重 Obsidian\n- 新增 write_zhihu_failures.py，自动生成抓取失败清单用于人工补录\n- 新增 obsidian_classify.py 脚本，提升笔记智能分类能力\n- 工具路由、脚本一览与主流程全面更新，完善历史与失败处理流程\n- skill-card.md 文件移除，使用更简单直观的入口说明\n\nv1.1.0 | 2026-05-01T12:14:02.401Z | user\n\nzhihu-fetch-skill 1.1.0 introduces improved cookie management and batch failure handling for more robust Zhihu scraping.\n\n- 增强：批量抓取内置 Cookie 自动保活机制（主动/被动 TTL 检测、激进刷新、自动恢复、定期保存 Cookie）\n- 新增：支持散发失败累积写入、连续失败自动中断与缓存丢弃、失败项可用 --retry-failed 参数重试\n- 更新：SKILL.md「已知问题与对策」详述保活与失败处理逻辑，补充关键脚本行为说明\n\nv1.0.0 | 2026-04-29T10:22:50.072Z | user\n\nInitial release of zhihu-fetcher skill:\n\n- Scrapes Zhihu collection lists and article content (including batch download and local images).\n- Supports multi-level API/Playwright fallback, persistent cookies, login session maintenance, and breakpoint resumption.\n- Optionally writes fetched articles into Obsidian vaults, auto-organizing markdown files and images.\n- Includes tools for login, session validation, and troubleshooting with suggested scripts and directory conventions.\n- Handles common scenarios like cookie expiration, anti-crawl measures, collection API limitations, and Windows console encoding issues.\n\nArchive index:\n\nArchive v2.2.0: 41 files, 90770 bytes\n\nFiles: README.md (7352b), scripts/requirements.txt (87b), scripts/zhihu_fetch/__init__.py (33b), scripts/zhihu_fetch/__main__.py (377b), scripts/zhihu_fetch/auth/__init__.py (0b), scripts/zhihu_fetch/auth/login_save.py (2749b), scripts/zhihu_fetch/auth/login.py (3781b), scripts/zhihu_fetch/auth/relogin.py (2551b), scripts/zhihu_fetch/body/__init__.py (0b), scripts/zhihu_fetch/body/api.py (6145b), scripts/zhihu_fetch/body/interactive.py (5334b), scripts/zhihu_fetch/body/stealth.py (6371b), scripts/zhihu_fetch/core/__init__.py (0b), scripts/zhihu_fetch/core/filters.py (6493b), scripts/zhihu_fetch/core/limits.py (6158b), scripts/zhihu_fetch/core/paths.py (1282b), scripts/zhihu_fetch/core/seen.py (6251b), scripts/zhihu_fetch/core/summary.py (2534b), scripts/zhihu_fetch/core/times.py (1508b), scripts/zhihu_fetch/core/url.py (3714b), scripts/zhihu_fetch/export/__init__.py (0b), scripts/zhihu_fetch/export/classify.py (4376b), scripts/zhihu_fetch/export/failures.py (2616b), scripts/zhihu_fetch/export/history.py (8344b), scripts/zhihu_fetch/export/notes.py (7767b), scripts/zhihu_fetch/export/obsidian.py (11213b), scripts/zhihu_fetch/fetch/__init__.py (0b), scripts/zhihu_fetch/fetch/batch.py (39343b), scripts/zhihu_fetch/fetch/collection.py (20826b), scripts/zhihu_fetch/fetch/columns.py (13980b), scripts/zhihu_fetch/fetch/follow.py (2362b), scripts/zhihu_fetch/fetch/history.py (16400b), scripts/zhihu_fetch/fetch/posts.py (9900b), scripts/zhihu_fetch/fetch/question.py (6216b), scripts/zhihu_fetch/fetch/route.py (2702b), scripts/zhihu_fetch/fetch/single.py (5099b), scripts/zhihu.py (3354b), skill-card.md (2616b), SKILL.md (22985b), zhihu_fetch_config.json (441b), _meta.json (136b)\n\nFile v2.2.0:SKILL.md\n\n---\nname: zhihu-fetcher\ndescription: \"知乎收藏夹与文章内容抓取：API/Playwright 多级降级、Cookie 持久化与保活、批量正文与图片、断点续传、可选写入 Obsidian。| Zhihu collection scraping, batch article fetch, Obsidian export.\"\nversion: \"2.2.0\"\nuser-invocable: true\nargument-hint: \"[知乎链接：收藏夹/专栏/文章/回答/问题页/个人页；或输出目录、Vault 路径]\"\nallowed-tools: Read, Write, Edit, Grep, Glob, Bash, WebFetch\n---\n\n# 知乎数据抓取\n\n从知乎获取**收藏夹文章列表**与**正文 Markdown**（含图片本地化），支持写入 **Obsidian** 知识库。命令与路径约定见下文；可视化说明见仓库根目录 [`README.md`](README.md)。\n\n---\n\n## 环境与约定\n\n- **语言**：默认与用户语种一致。\n- **技能根目录**：本仓库根目录（含 `SKILL.md` 与 `scripts/`）。下文命令均从该目录执行，写作 `python scripts/...`。\n- **工作区目录**：Cookie、浏览器用户数据、默认文章输出等放在工作区（已 gitignore，勿提交）。\n  - 环境变量 **`ZHIHU_WORKSPACE`** 优先；\n  - 未设置时默认为技能根目录下的 **`zhihu-fetch-workspace/`**。\n- **依赖**：在 **`scripts/`** 下执行 **`pip install -r requirements.txt`**，并 **`playwright install chromium`**。\n- **命令入口**：根目录只留 [`scripts/zhihu.py`](scripts/zhihu.py)；业务代码在 [`scripts/zhihu_fetch/`](scripts/zhihu_fetch/) 分模块。统一写成 `python scripts/zhihu.py <命令>`。\n\n### 抓取上限（配置优先，对话可固化）\n\n不带条数时**不会全量爬**。读取顺序：当次命令行 → 工作区配置 → 技能根配置 → 代码默认值。\n\n| 文件 / 命令 | 作用 |\n|------|------|\n| [`zhihu_fetch_config.json`](zhihu_fetch_config.json) | **技能级**默认上限；用户说「以后默认…」时改这个并保存 |\n| `{workspace}/zhihu_fetch_config.json` | **本机覆盖**（在 gitignore 的工作区内） |\n| `python scripts/zhihu.py limits` | 查看当前生效值 |\n| `python scripts/zhihu.py limits --set collection.max_items=10` | 写入技能根配置（默认 `--where skill`） |\n| `--where workspace` | 只改本机覆盖 |\n| `--all` 或配置 `\"unlimited\": true` | 取消上限 |\n\n用户说「这次多抓一点」→ 命令行 `--max-items` / `--all`。用户说「以后默认每夹 10 篇」→ **改配置并固化**，不要只改当次命令。\n\n默认：收藏夹最多 10 个、每夹 20 篇；专栏最多 5 个、每栏 20 篇；个人文章/回答各 20 篇；问题页回答 20 条；历史/批量各 20 篇。\n\n**增量**：列表脚本支持 **`--since-last`**，对照工作区 `zhihu_url_index.json`（含 `content_updated`）、已有 `zhihu_*.json`、`_progress.json` 与 Markdown frontmatter 的 `url:`。未更新的已见 URL 跳过；列表里的更新时间**新于索引**则标 `refresh` 再抓，`batch` **不会**因 `_progress.json` 的 `completed` 跳过这些条。个人「文章」默认还会按 URL 排除已在专栏 JSON 里的篇目（两者重叠，有更新仍会刷新）。\n\n**列表过滤**（条数上限之外）：`--min-voteup N`、`--days N`、`--since ISO`。配置 `filter.min_voteup` / `filter.since_days`（`0` = 不过滤）。作用在收藏夹 / 专栏 / 文章 / 回答 / 问题页 / 跟读包。`max_items` 只计通过过滤且为 new/refresh 的条目。\n\n**登录态正文**：工作区有 Cookie 时，API / 页面 / 批量图片下载都会自动带上，降低专栏文章 403。未登录先跑 `python scripts/zhihu.py login` / `relogin`。\n\n**每次 run 摘要**：列表与 batch 结束会打印并写入 `{workspace}/zhihu_run_summary.json`（成功 / 跳过空项 / 跳过已抓 / 失败 / 403 / 需登录）。Agent 回复用户时读这份摘要；失败项仍可用 `python scripts/zhihu.py failures` 写入 Vault。\n\n### 登录与可选页面验证\n\n- **`python scripts/zhihu.py login`**：打开浏览器等待登录，默认以检测到 **`z_c0`** 为成功条件即可结束（不要求额外跳转）。\n- **可选二次校验**：若用户希望登录后再确认「某一内需登录页」是否可访问（如某收藏夹页、专栏后台、关注动态等），属**可选项**，不设则不执行：\n  - **环境变量** **`ZHIHU_VERIFY_URL`**：值为完整 **`http://` 或 `https://`** URL；\n  - **或**命令行第一个参数传入同一完整 URL：`python scripts/zhihu.py login \"https://www.zhihu.com/...\"`。\n  - 脚本会访问该 URL，若正文仍出现知乎通用提示「请登录后查看」，则提示可能未登录完成；否则认为当前会话可访问该页。**不限定于收藏夹**，任意知乎链接均可（只要登录态相关）。\n- **`python scripts/zhihu.py relogin`**：Cookie 失效、需重新登录并写回 **`zhihu_cookies.json`** 时使用（会打开浏览器）。\n\n---\n\n## 触发条件\n\n在用户使用以下任一方式时启用本技能：\n\n- 明确提及：知乎、Zhihu、专栏、收藏夹、文章抓取、批量下载、Cookie、验证码、Obsidian、知识库同步等\n- 粘贴 **zhihu.com** / **zhuanlan.zhihu.com** 链接并希望获取正文或列表\n- 需要 **断点续传**、**图片落盘**、**反爬 / Stealth** 相关协助\n\n---\n\n## 工具与脚本路由\n\n按任务选用能力；具体工具名以当前 Agent 环境为准。\n\n### 统一入口（先识别链接，再调现有脚本）\n\n用户丢什么链接就走哪条流水线。**优先** `python scripts/zhihu.py route <URL> [原命令参数]`（`fetch` 遇到非单篇链接也会转过来）。受 `zhihu_fetch_config.json` 上限约束。\n\n| 用户丢的链接 | kind | 调用 |\n|--------------|------|------|\n| `/p/` 或 `zhuanlan.zhihu.com/p/{id}` | article | `zhihu.py fetch` |\n| `/question/{qid}/answer/{aid}` | answer | `zhihu.py fetch`（回答 API，带 Cookie） |\n| `/collection/{id}` | collection | `zhihu.py collection` |\n| `/people/{slug}/collections` | collections | `zhihu.py collection`；`--collection 名称` 只爬一夹 |\n| `/people/{slug}/columns` 或 `/column/{id}` | columns / column | `zhihu.py columns`；`--column 名称` 只爬一栏 |\n| `/people/{slug}/posts` | posts | `zhihu.py posts --kind articles` |\n| `/people/{slug}/answers` | answers | `zhihu.py posts --kind answers` |\n| `/question/{id}`（无 `/answer/`） | question | `zhihu.py question`：默认排序回答列表 → JSON，再 `batch`；上限 `question.max_answers` |\n| 裸 `/people/{slug}` | people | **跟读包**：专栏 + 文章 + 回答（默认 `--since-last`，可用 `--no-since-last` 关掉）。单栏目仍用 `--posts` / `--answers` / `--columns` / `--collections` / `--history` |\n\n### 常见任务与建议方式\n\n| 任务 | 建议方式 |\n|------|----------|\n| 任意知乎链接（不知道类型） | **`Bash`** → `python scripts/zhihu.py route <URL>` |\n| 获取收藏夹 JSON 列表 | **`Bash`** → `python scripts/zhihu.py collection <收藏夹URL或ID>`；个人页加 `--collection CS` 只爬该夹；`--since-last` 只补新 |\n| 获取用户专栏列表与文章 | **`Bash`** → `python scripts/zhihu.py columns <people URL 或 /columns>`；`--column 名称` 只爬指定专栏；`--since-last` 只补新；层级 JSON + 每栏 `zhihu_column_{id}.json` 可交给 batch |\n| 获取个人页文章 / 回答 | **`Bash`** → `python scripts/zhihu.py posts <people URL> --kind articles\\|answers\\|both`；文章默认排除已在专栏里的 URL；`--since-last` 只补新 |\n| 裸个人页跟读 | **`Bash`** → `python scripts/zhihu.py route <people URL>` 或 `follow`；默认专栏+文章+回答+增量 |\n| 问题页回答列表 | **`Bash`** → `python scripts/zhihu.py route <question URL>` 或 `question`；再 `batch` |\n| 获取个人主页点赞/收藏历史 | **`Bash`** → `python scripts/zhihu.py history <people URL 或 slug> <起始时间ISO> <输出.json> [--until <结束时间ISO>]`；按活动时间保留 `interaction_*` 元数据，支持断点续跑 |\n| 批量抓取正文与图片 | **`Bash`** → `python scripts/zhihu.py batch <列表.json> [输出目录] [图片目录]`；默认输出目录见「路径约定」；结束写 `zhihu_run_summary.json` |\n| 写入 Obsidian 原文镜像 | **`Bash`** → `python scripts/zhihu.py obsidian <文章目录> [Vault路径]`；写入 **`{Vault}/知乎收藏/{分类}/`**（会删工作区源 md）；Vault：命令行优先，否则 **`OBSIDIAN_VAULT`** |\n| 从镜像生成笔记 | **`Bash`** → `python scripts/zhihu.py notes [Vault路径]`；扫描「知乎收藏」，写入并列根目录 **`{Vault}/知乎笔记/`**，**不改、不删镜像**；已有笔记默认跳过，`--force` 覆盖 |\n| 写入个人历史到 Obsidian | **`Bash`** → `python scripts/zhihu.py history-obsidian <文章目录> <Vault路径> [.]`；默认写入 `{Vault}/知乎收藏/{分类}/`，按 URL 去重更新 |\n| 写入失败项清单 | **`Bash`** → `python scripts/zhihu.py failures <Vault路径> <标签>:<progress.json> ...`；生成 `{Vault}/知乎收藏/抓取失败.md` |\n| Cookie 失效需人工登录 | **`Bash`** → `python scripts/zhihu.py relogin`（会打开浏览器窗口） |\n| 首次登录辅助（可选验证页） | **`Bash`** → `python scripts/zhihu.py login`；可选 **`ZHIHU_VERIFY_URL`** 或首个参数传入完整 http(s) 链接，见「登录与可选页面验证」 |\n| 单篇快速验证 | **`Bash`** → `python scripts/zhihu.py fetch`（文章/回答；其它 URL 转路由）或 `api` / `stealth` / `interactive` |\n| 查看 / 固化抓取上限 | **`Bash`** → `python scripts/zhihu.py limits`；改默认用 `--set key=value`（见「抓取上限」） |\n\n---\n\n## 模块一览\n\n根入口：`python scripts/zhihu.py <命令>`。实现按目录划分：\n\n```\nscripts/\n  zhihu.py                 # 唯一 CLI\n  requirements.txt\n  zhihu_fetch/\n    core/                  # paths, limits, url, seen, times, filters, summary\n    fetch/                 # route, collection, columns, posts, follow, question, history, batch, single\n    body/                  # api, stealth, interactive\n    auth/                  # login, relogin, login_save\n    export/                # obsidian, notes, history, failures, classify\n```\n\n| 命令 | 模块 | 用途 |\n|------|------|------|\n| `route` | `fetch/route.py` | 识别链接后转调对应流水线 |\n| `collection` | `fetch/collection.py` | 收藏夹列表；`--collection`、`--since-last` |\n| `columns` | `fetch/columns.py` | 用户专栏；`--column`、`--since-last` |\n| `posts` | `fetch/posts.py` | 个人文章 / 回答；与专栏 URL 去重 |\n| `follow` | `fetch/follow.py` | 裸主页跟读包（专栏+文章+回答） |\n| `question` | `fetch/question.py` | 问题页默认排序回答列表 |\n| `history` | `fetch/history.py` | 点赞/收藏动态 |\n| `batch` | `fetch/batch.py` | 批量正文、图片、断点续传、摘要 |\n| `fetch` | `fetch/single.py` | 单篇文章/回答 |\n| `api` / `stealth` / `interactive` | `body/` | 单篇调试 |\n| `login` / `relogin` / `login-save` | `auth/` | 登录与 Cookie |\n| `obsidian` / `notes` / `history-obsidian` / `failures` | `export/` | 原文镜像 / 并列笔记 / 历史 / 失败清单 |\n| `limits` | `core/limits.py` | 抓取上限配置 |\n\n---\n\n## 主流程（推荐执行顺序）\n\n1. **安装依赖**：`scripts/requirements.txt` + Chromium。\n2. **用户给出链接**：**`python scripts/zhihu.py route <URL>`**（不要凭印象选模块）。\n3. 列表得到 JSON 后 **`python scripts/zhihu.py batch`** → **`zhihu_articles_*/`**（含 **`_progress.json`**、**`images/`**、编号 **`*.md`**）。\n4. （可选）**`python scripts/zhihu.py obsidian`** → 原文镜像到 **`{Vault}/知乎收藏/{分类}/`**。\n5. （可选）**`python scripts/zhihu.py notes`** → 从「知乎收藏」生成并列的 **`{Vault}/知乎笔记/`**（不改镜像）。\n6. 回复用户前读 **`zhihu_run_summary.json`**。\n\n中断批量任务时：**重新运行同一条** `python scripts/zhihu.py batch` 命令即可续跑（已完成 URL 记录在 `_progress.json`）。\n\n### 个人历史流程（点赞 / 收藏）\n\n适用于个人主页动态中的 **赞同了回答 / 赞同了文章 / 收藏了回答 / 收藏了文章**。时间采用 ISO 格式，**建议显式带时区**（如 `+08:00`）；若省略时区，默认按 **Asia/Shanghai** 解释，可用环境变量 **`ZHIHU_TIMEZONE`** 或 **`TZ`** 覆盖。\n\n```bash\n# 1. 收集活动列表（起始时间含，结束时间不含）\npython scripts/zhihu.py history \\\n  https://www.zhihu.com/people/<slug> \\\n  2026-01-01T00:00:00+08:00 \\\n  /path/to/runtime/zhihu_history_2026-01-01_to_2026-04-05.json \\\n  --until 2026-04-05T00:00:00+08:00\n\n# 2. 抓取正文与图片；失败默认自动重试 3 次\npython scripts/zhihu.py batch \\\n  /path/to/runtime/zhihu_history_2026-01-01_to_2026-04-05.json \\\n  /path/to/runtime/zhihu_articles_history_2026-01-01_to_2026-04-05\n\n# 3. 写入 Obsidian 的知乎收藏根目录分类文件夹\npython scripts/zhihu.py history-obsidian \\\n  /path/to/runtime/zhihu_articles_history_2026-01-01_to_2026-04-05 \\\n  /path/to/ObsidianVault \\\n  .\n```\n\n历史笔记会保留：\n\n```yaml\ninteraction_action: \"赞同了回答\"\ninteraction_time: 2026-03-20T10:17:57.235000+00:00\ninteraction_date: 2026-03-20\ntags: [zhihu, 编程与开发, 赞同了回答]\n```\n\n历史列表中断时重新运行同一条命令即可续跑；加 `--fresh` 可忽略现有 checkpoint 重建。写入 Obsidian 时会扫描已有笔记的 `url` 并按 URL 更新，避免重复导入。\n\n### 用户专栏流程（他的专栏）\n\n适用于 `https://www.zhihu.com/people/<slug>/columns`：先列出专栏，再按栏抓文章。多专栏是列表，栏下文章是层级；可用 **`--column 专栏名`** 只爬其中一个。不写条数时走配置 `column.max_columns` / `column.items_per_column`。\n\n```bash\n# 列出并抓取（受配置默认上限）\npython scripts/zhihu.py columns https://www.zhihu.com/people/<slug>/columns\n\n# 只爬指定专栏名，每栏 2 篇\npython scripts/zhihu.py columns https://www.zhihu.com/people/<slug>/columns --column 远东轶事 --per-column 2\n\n# 仅列专栏、不抓文章\npython scripts/zhihu.py columns https://www.zhihu.com/people/<slug>/columns --list-only\n\n# 正文与图片\npython scripts/zhihu.py batch zhihu-fetch-workspace/zhihu_column_<id>.json\n```\n\n输出：`zhihu_columns_{slug}.json`（层级）+ `zhihu_column_{id}.json`（单栏，可交给 batch）。加 `--since-last` 只补尚未抓过的文章。\n\n### 个人页文章 / 回答\n\n专栏只是作者产出的一部分。跟读某个作者时，裸主页默认跑跟读包；也可以只用 `/posts` 与 `/answers`。文章与专栏按 URL 去重后再抓。默认上限 `people.max_articles` / `people.max_answers`。\n\n```bash\npython scripts/zhihu.py route https://www.zhihu.com/people/<slug>\npython scripts/zhihu.py follow https://www.zhihu.com/people/<slug> --no-since-last\npython scripts/zhihu.py route https://www.zhihu.com/people/<slug>/posts\npython scripts/zhihu.py route https://www.zhihu.com/people/<slug>/answers\npython scripts/zhihu.py posts https://www.zhihu.com/people/<slug> --kind both --since-last --min-voteup 50 --days 30\npython scripts/zhihu.py batch zhihu-fetch-workspace/zhihu_posts_<slug>.json\n```\n\n输出：`zhihu_posts_{slug}.json`、`zhihu_answers_{slug}.json`（可交给 batch）。\n\n### 问题页回答\n\n`https://www.zhihu.com/question/{id}`（不含 `/answer/`）拉默认排序回答列表，上限 `question.max_answers`。\n\n```bash\npython scripts/zhihu.py route https://www.zhihu.com/question/<id> --max-items 2\npython scripts/zhihu.py batch zhihu-fetch-workspace/zhihu_question_<id>.json\n```\n\n---\n\n## 路径与输出约定\n\n### 批量抓取命令格式\n\n```bash\npython scripts/zhihu.py batch <列表文件> [输出目录] [图片目录]\n```\n\n| 参数 | 说明 |\n|------|------|\n| **列表文件** | `zhihu.py collection` 产出的 JSON |\n| **输出目录** | 可选；省略时默认为 **`{workspace}/zhihu_articles_{collectionId}/`**（`collectionId` 由列表文件名推导） |\n| **图片目录** | 可选；省略时默认为 **`{输出目录}/images/`** |\n\n### 目录结构示例\n\n```\nzhihu_articles_{collectionId}/\n├── _progress.json          # 断点续传\n├── images/                 # 默认图片目录\n│   └── ...\n├── 0001_文章标题.md\n└── ...\n```\n\n### 单篇文章格式要点\n\n- YAML frontmatter：`title`、`author`、`source`、`url`、`voteup`、`images` 等\n- 正文为 Markdown；图片引用指向本地 **`images/`** 下文件名（或脚本生成的相对路径）\n\n示例结构：\n\n```markdown\n---\ntitle: \"文章标题\"\nauthor: \"作者\"\nsource: zhihu\nurl: \"https://...\"\nvoteup: 123\nimages: 5\n---\n\n# 文章标题\n\n> 作者: xxx | 原文: [知乎链接](https://...)\n\n正文...\n```\n\n### 持久化文件（默认 workspace）\n\n未设置 **`ZHIHU_WORKSPACE`** 时，`{workspace}` 为技能根目录下的 **`zhihu-fetch-workspace/`**。\n\n| 用途 | 路径 |\n|------|------|\n| Cookie | `{workspace}/zhihu_cookies.json` |\n| URL 增量索引 | `{workspace}/zhihu_url_index.json` |\n| 当次 run 摘要 | `{workspace}/zhihu_run_summary.json` |\n| Playwright 用户数据 | `{workspace}/chrome_user_data/` |\n| 默认文章目录 | `{workspace}/zhihu_articles_{collectionId}/` |\n| 默认图片目录 | `{文章输出目录}/images/` |\n\n---\n\n## Obsidian 写入要点\n\n- **原文镜像**：`python scripts/zhihu.py obsidian` 写入 **`{Vault}/知乎收藏/{分类}/{标题}.md`**，行为与旧版相同（分类、图片、删除工作区源 md）。**不要**把镜像改成笔记、不要覆盖已有镜像来当笔记。\n- **笔记板块**：`python scripts/zhihu.py notes [Vault]` 扫描「知乎收藏」，在 Vault **并列根目录**写入 **`{Vault}/知乎笔记/{分类}/`**，并维护 `知乎笔记/作者/`、`知乎笔记/问题/` 索引。不删除、不改写镜像。已有笔记默认跳过；Agent 可再润色摘要。\n- **Vault**：① **命令行参数**；② 环境变量 **`OBSIDIAN_VAULT`**；③ 常见目录扫描。\n- **分类**：优先对齐已有 **`知乎收藏/`** 子目录；否则按内容关键词；无法归类则 **「未分类」**。\n\n---\n\n## 已知问题与对策\n\n| # | 现象 / 原因 | 处理 |\n|---|-------------|------|\n| 1 | **Cookie 失效**：标题「安全验证」、`/account/unhuman` | **自动恢复**：脚本内置 3 次重试（激进保活：访问文章页+模拟阅读）；仍失败则 **`python scripts/zhihu.py relogin`** |\n| 2 | **收藏夹 API 分页**：带 `include` 时列表可能被截断 | **`zhihu.py collection`** 已内置 API ↔ DOM 切换；必要时减少 `include` 或走浏览器分页 |\n| 3 | **反爬**：Headless 被识别 | Stealth、UA、间隔；必要时 **`zhihu.py interactive`** |\n| 4 | **API 正文不完整**：`include` 只给摘要 | 批量与单篇流程中已优先 **页面 DOM** 拉全文 |\n| 5 | **图片下载失败** | 正文仍保留原 URL；排查网络、Referer、过期链接 |\n| 6 | **Windows 控制台 GBK** | 脚本已 **`sys.stdout.reconfigure(encoding='utf-8')`** |\n| 7 | **批量中断** | 直接再次运行 **`python scripts/zhihu.py batch`**，依赖 **`_progress.json`** |\n| 8 | **失败项累积** | 散发失败自动记录到 `_progress.json`（含 url/reason/title/timestamp）；连续失败 ≥5 次中断并丢弃缓存；用 **`--retry-failed`** 参数可重试 |\n\n### Cookie 保活机制\n\n脚本内置多层 Cookie 保活策略：\n\n1. **主动 TTL 检测**（每篇文章）：解析 z_c0 的 `expires` 字段，剩余 < 30 分钟时自动触发激进刷新\n2. **常规保活**（每 5-8 篇）：访问知乎列表页 + 模拟滚动\n3. **激进保活**（每 ~20 篇）：访问实际文章页 + 模拟阅读（停留 2-5 秒 + 滚动）\n4. **被动检测**：每次访问文章时检查是否被重定向到 `/account/unhuman` 或 `/signin`\n5. **自动恢复**：检测到失效时，自动尝试 3 次激进保活恢复\n6. **Cookie 备份**：每次保活后自动从浏览器提取最新 Cookie 保存到文件（扩展格式含 expires）\n7. **安全退出**：脚本结束前保存最新 Cookie + 当前进度\n\n### 失败处理策略\n\n脚本采用**两级失败处理**，区分「文章本身问题」和「环境问题」：\n\n| 场景 | 行为 | 说明 |\n|------|------|------|\n| 散发失败（中间有成功） | 记录到 `_progress.json` 的 `failed` 字段 | 视为文章本身问题（已删除/不可访问），后续跳过 |\n| 连续失败 ≥ 5 次 | 中断抓取，**丢弃**缓存的失败记录 | 视为环境问题（Cookie/网络），下次重试仍可跑 |\n\n**工作原理：**\n- 失败先缓存在内存中，不立即写入进度文件\n- 下一条成功时，将缓存的失败记录批量写入进度文件（确认是文章问题）\n- 连续失败达到阈值（5 次）时，中断抓取，丢弃缓存（保留重试机会）\n\n**相关常量：**\n- `CONSECUTIVE_FAIL_THRESHOLD = 5`：连续失败阈值\n- `CONSECUTIVE_FAIL_INTERRUPT = True`：是否在连续失败时中断\n\n**重试模式：**\n```bash\npython scripts/zhihu.py batch <列表文件> [输出目录] [图片目录] --retry-failed\n```\n此模式会清空 `failed` 列表，只重试之前记录为失败的文章。\n\n---\n\n## 故障排查流程\n\n```\n正文全空？\n  → Cookie（含 z_c0）→ 是否跳转验证页 → python scripts/zhihu.py relogin\n\n图片失败？\n  → URL/网络/Referer → Markdown 中仍可保留链接\n\n批量中途停止？\n  → 确认 _progress.json → 原命令重跑\n```\n\n---\n\n## Agent 自用工作流检查清单\n\n```\n□ 已确认 scripts 依赖与 playwright chromium 可用；必要时提示用户设置 ZHIHU_WORKSPACE\n□ 用户丢了链接：先 python scripts/zhihu.py route，不要猜错模块；裸个人页默认跟读包（专栏+文章+回答）；单栏目用 --posts / --answers / --columns / --collections / --history\n□ 收藏夹：zhihu.py collection；指定夹名用 --collection；增量 --since-last；赞数/时间过滤 --min-voteup / --days / --since；再 batch\n□ 专栏 / 文章 / 回答 / 问题页：columns、posts、question；文章按 URL 去重；内容更新会 refresh；默认受配置上限，全量才 --all\n□ 正文入口有 Cookie 就带上；专栏 403 优先登录而非换抓取器\n□ 批量输出路径：知悉默认 {workspace}/zhihu_articles_* 与 images/ 子目录；第三个参数仅在自定义图片目录时需要\n□ 回复用户前读 zhihu_run_summary.json（成功/跳过/403/需登录）；失败清单可用 python scripts/zhihu.py failures\n□ Obsidian 原文：`zhihu.py obsidian` → `{Vault}/知乎收藏/`；笔记：`zhihu.py notes` → `{Vault}/知乎笔记/`（并列，不改镜像）\n□ 遇验证页或全文为空：优先 Cookie/重登录，而非重复盲目加大并发\n□ 用户仅需单篇或调试：选用 python scripts/zhihu.py fetch（文章/回答），避免不必要批量\n```\n\nFile v2.2.0:README.md\n\n<div align=\"center\">\n\n# 知乎抓取.skill\n\n> 从知乎**收藏夹列表**到**批量正文与图片**，再到 **Obsidian 自动分类入库**：API / Playwright 多级降级、Cookie 持久化与保活、断点续传。\n\n[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](LICENSE)\n[![Python 3.10+](https://img.shields.io/badge/Python-3.10%2B-blue.svg)](https://www.python.org/)\n[![Playwright](https://img.shields.io/badge/Playwright-Chromium-45ba4b.svg)](https://playwright.dev/)\n[![AgentSkills](https://img.shields.io/badge/AgentSkills-Standard-green)](https://agentskills.io)\n\n<br>\n\n收藏夹里上千篇文章想**归档成 Markdown**？<br>\n需要**配图本地化**、中断后能**接着抓**？<br>\n希望落库到 Obsidian，并**按主题自动分类**？<br>\nCookie 经常失效，想要**持久化上下文 + 保活**？\n\n**本 Skill 按 AgentSkills 约定编排全流程，入口见根目录 [`SKILL.md`](SKILL.md)，脚本集中在 `scripts/`。**\n\n[功能特性](#功能特性) · [运行效果](#运行效果) · [安装](#安装) · [使用](#使用) · [项目结构](#项目结构) · [参考文档](#参考文档)\n\n</div>\n\n---\n\n## 功能特性\n\n| 能力 | 说明 |\n|------|------|\n| 收藏夹列表 | `zhihu.py collection`：优先 API，失败降级 Playwright DOM；`--collection 名称`、`--since-last` |\n| 用户专栏 | `zhihu.py columns`：`--column 名称`、`--since-last`，层级 JSON 可交给 batch |\n| 个人文章 / 回答 | `zhihu.py posts`：与专栏按 URL 去重；`--since-last` 只补新；内容更新会 refresh |\n| 跟读包 | `zhihu.py follow` / 裸主页 `route`：专栏 + 文章 + 回答 |\n| 问题页 | `zhihu.py question`：默认排序回答列表 |\n| 统一入口 | `zhihu.py route`：识别 `/collection/` `/columns` `/posts` `/answers` `/question/` `/p/` 回答链接 个人主页 |\n| 个人历史列表 | `zhihu.py history`：个人主页点赞/收藏动态，支持时间范围、断点续跑、互动时间元数据 |\n| 批量抓取 | `zhihu.py batch`：正文 Markdown、图片默认写入 `{输出目录}/images/`、`_progress.json` 断点续传、失败自动重试、API 回退 |\n| Cookie | 持久化浏览器上下文 + 定时保活；失效时用 `zhihu.py relogin` 手动登录 |\n| 单篇 / 调试 | `zhihu.py fetch` / `api` / `stealth` / `interactive` |\n| Obsidian | `zhihu.py obsidian`：原文镜像到 `{Vault}/知乎收藏/`；`zhihu.py notes`：笔记到并列的 `{Vault}/知乎笔记/` |\n\n**依赖**：见 [`scripts/requirements.txt`](scripts/requirements.txt)，并需 `playwright install chromium`。\n\n---\n\n## 运行效果\n\n<table width=\"100%\" border=\"1\" cellpadding=\"12\" cellspacing=\"0\">\n<tr>\n<th width=\"50%\" align=\"center\">批量抓取<br><sub>Agent 对话中的进度、剩余篇数与 Cookie 保活（OpenClaw 示例）</sub></th>\n<th width=\"50%\" align=\"center\">写入 Obsidian<br><sub>「知乎收藏」主题分类与关系图谱</sub></th>\n</tr>\n<tr>\n<td width=\"50%\" valign=\"top\" align=\"center\">\n<img src=\"docs/openclaw-run.jpg\" alt=\"Agent 对话：批量抓取进度与 Cookie 保活\" width=\"100%\" />\n</td>\n<td width=\"50%\" valign=\"top\" align=\"center\">\n<img src=\"docs/obs.jpg\" alt=\"Obsidian：知乎收藏分类与关系图谱\" width=\"100%\" />\n</td>\n</tr>\n</table>\n\n---\n\n## 安装\n\n### 加载技能\n\n将本仓库放到 Agent 宿主约定的 skills 路径（与 [`SKILL.md`](SKILL.md) 同级为 skill 根目录），重启后在技能列表中确认已加载。路径因宿主而异，例如 Claude Code、Cursor、OpenClaw 等。\n\n```bash\n# 示例：克隆到项目的 skills 目录（按宿主调整目标路径）\ngit clone https://github.com/handsomestWei/zhihu-fetch-skill.git\n```\n\n### 依赖\n\n```bash\ncd scripts\npip install -r requirements.txt\nplaywright install chromium\n```\n\n仓库根目录运行测试（访问真实知乎，仅最近少量条目；账号见 `tests/live_profile.py`）：\n\n```bash\npython -m pytest\n```\n\n抓取上限集中在根目录 [`zhihu_fetch_config.json`](zhihu_fetch_config.json)，运行时优先读配置；对话里改默认用 `python scripts/zhihu.py limits --set key=value`。详情见 [`SKILL.md`](SKILL.md)。\n\n---\n\n## 使用\n\n在 Agent 中用自然语言描述即可，例如：知乎文章、收藏夹、批量抓取、写入 Obsidian、Cookie 失效。\n\n典型三步（默认工作区为技能根下 `zhihu-fetch-workspace/`，可用环境变量 **`ZHIHU_WORKSPACE`** 覆盖，详见 [`SKILL.md`](SKILL.md)）：\n\n```bash\n# 1. 收藏夹 → JSON 列表\npython scripts/zhihu.py collection <收藏夹URL或ID>\n\n# 2. 批量抓取正文与图片\npython scripts/zhihu.py batch <列表.json>\n\n# 3. 原文镜像写入 Obsidian「知乎收藏」（可选 Vault 路径）\npython scripts/zhihu.py obsidian <文章目录> [Vault路径]\n\n# 4. 从「知乎收藏」生成并列的「知乎笔记」（不改镜像）\npython scripts/zhihu.py notes [Vault路径]\n```\n\n用户专栏（`/people/<slug>/columns`，支持 `--column 名称`、`--since-last`）：\n\n```bash\npython scripts/zhihu.py route https://www.zhihu.com/people/<slug>/columns --column 远东轶事 --per-column 2\npython scripts/zhihu.py batch zhihu-fetch-workspace/zhihu_column_<id>.json\n```\n\n个人文章 / 回答（与专栏按 URL 去重）：\n\n```bash\npython scripts/zhihu.py route https://www.zhihu.com/people/<slug>/posts\npython scripts/zhihu.py route https://www.zhihu.com/people/<slug>/answers --since-last\npython scripts/zhihu.py route https://www.zhihu.com/people/<slug>\npython scripts/zhihu.py route https://www.zhihu.com/question/<id> --max-items 2\n```\n\n收藏夹按名称筛选：\n\n```bash\npython scripts/zhihu.py collection https://www.zhihu.com/people/<slug> --per-collection 20 --collection CS --since-last\n```\n\n个人历史（点赞 / 收藏）示例：\n\n```bash\n# 1. 个人动态 → JSON 列表（起始时间含，结束时间不含）\npython scripts/zhihu.py history \\\n  https://www.zhihu.com/people/<slug> \\\n  2026-01-01T00:00:00+08:00 \\\n  runtime/zhihu_history_2026-01-01_to_2026-04-05.json \\\n  --until 2026-04-05T00:00:00+08:00\n\n# 2. 批量抓取正文与图片（失败默认自动重试 3 次）\npython scripts/zhihu.py batch \\\n  runtime/zhihu_history_2026-01-01_to_2026-04-05.json \\\n  runtime/zhihu_articles_history_2026-01-01_to_2026-04-05\n\n# 3. 写入 Obsidian 的「知乎收藏/{分类}/」根分类文件夹，按 url 去重更新\npython scripts/zhihu.py history-obsidian \\\n  runtime/zhihu_articles_history_2026-01-01_to_2026-04-05 \\\n  /path/to/ObsidianVault \\\n  .\n```\n\nCookie 异常时：\n\n```bash\npython scripts/zhihu.py relogin\n```\n\n---\n\n## 项目结构\n\n本仓库遵循 [AgentSkills](https://agentskills.io)，根目录即一个 skill：\n\n```\nzhihu-fetch-skill/\n├── SKILL.md\n├── README.md\n├── docs/\n├── tests/                 # core / fetch / live 与脚本模块对应\n├── scripts/\n│   ├── zhihu.py           # 唯一 CLI\n│   ├── requirements.txt\n│   └── zhihu_fetch/       # core, fetch, body, auth, export\n└── zhihu-fetch-workspace/\n```\n\n默认路径与命令以 [`SKILL.md`](SKILL.md) 为准。\n\n---\n\n## 参考文档\n\n- [技能入口与完整命令说明](SKILL.md)（依赖、脚本表、故障排查）\n- [脚本依赖清单](scripts/requirements.txt)\n\n---\n\n<div align=\"center\">\n\nMIT License © [handsomestWei](https://github.com/handsomestWei/)\n\n</div>\n\nFile v2.2.0:_meta.json\n\n{\n  \"ownerId\": \"kn71nkfhpcw7dp6pkyqrj43dgd84eg1k\",\n  \"slug\": \"zhihu-fetch-skill\",\n  \"version\": \"2.2.0\",\n  \"publishedAt\": 1788103322576\n}\n\nFile v2.2.0:scripts/requirements.txt\n\nrequests>=2.28.0\nbeautifulsoup4>=4.12.0\nplaywright>=1.40.0\npytest>=8.0.0\npytest>=8.0.0\n\nFile v2.2.0:skill-card.md\n\n## Description:\n\nFetches Zhihu collection lists, articles, answers, questions, user posts, and history into local JSON and Markdown, with Playwright fallback, cookie persistence, image localization, resumable batch runs, and optional Obsidian export.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[handsomestwei](https://clawhub.ai/user/handsomestwei)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers and agent users use this skill to collect Zhihu content, resume interrupted crawls, save article bodies and images as Markdown, and optionally organize mirrored content and notes inside an Obsidian vault.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: The skill stores Zhihu session cookies and can reuse them while fetching pages and images.\n\nMitigation: Use a dedicated workspace, restrict access to zhihu_cookies.json, and delete or rotate cookies when the run is complete.\n\nRisk: Untrusted batch JSON or fetched article content may drive authenticated requests or introduce unsafe Markdown content.\n\nMitigation: Run only trusted batch lists, review fetched content before importing it into a knowledge base, and avoid sensitive logged-in sessions for untrusted inputs.\n\nRisk: Obsidian export commands write into the selected vault and may update or delete source mirror files during import workflows.\n\nMitigation: Back up the vault, test with a small batch first, and verify the target vault path before running export commands.\n\n## Reference(s):\n\n- [ClawHub skill page](https://clawhub.ai/handsomestwei/skills/zhihu-fetch-skill)\n- [Publisher profile](https://clawhub.ai/user/handsomestwei)\n- [Python](https://www.python.org/)\n- [Playwright](https://playwright.dev/)\n- [AgentSkills](https://agentskills.io)\n\n## Skill Output:\n\n**Output Type(s):** [text, markdown, shell commands, configuration, guidance]\n\n**Output Format:** [Markdown guidance with inline shell commands; generated artifacts include JSON lists, Markdown articles, local image files, progress files, run summaries, and Obsidian notes.]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [Uses local workspace state for cookies, browser data, progress checkpoints, crawl limits, URL indexes, and run summaries.]\n\n## Skill Version(s):\n\n2.2.0 (source: server release and skill frontmatter)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nFile v2.2.0:zhihu_fetch_config.json\n\n{\n  \"unlimited\": false,\n  \"collection\": {\n    \"max_collections\": 10,\n    \"items_per_collection\": 20,\n    \"max_items\": 20\n  },\n  \"column\": {\n    \"max_columns\": 5,\n    \"items_per_column\": 20\n  },\n  \"history\": {\n    \"max_items\": 20\n  },\n  \"batch\": {\n    \"max_items\": 20\n  },\n  \"people\": {\n    \"max_articles\": 20,\n    \"max_answers\": 20\n  },\n  \"question\": {\n    \"max_answers\": 20\n  },\n  \"filter\": {\n    \"min_voteup\": 0,\n    \"since_days\": 0\n  }\n}\n\nArchive v1.2.0: 19 files, 54207 bytes\n\nFiles: README.md (6023b), scripts/fetch_zhihu_api.py (5256b), scripts/fetch_zhihu_batch.py (37280b), scripts/fetch_zhihu_collection.py (10018b), scripts/fetch_zhihu_history.py (15512b), scripts/fetch_zhihu_interactive.py (5435b), scripts/fetch_zhihu_stealth.py (6371b), scripts/fetch_zhihu.py (3548b), scripts/obsidian_classify.py (4167b), scripts/requirements.txt (59b), scripts/write_to_obsidian.py (11528b), scripts/write_zhihu_failures.py (2709b), scripts/write_zhihu_history_to_obsidian.py (8601b), scripts/zhihu_login_save.py (2780b), scripts/zhihu_login.py (4070b), scripts/zhihu_relogin.py (2962b), skill-card.md (2633b), SKILL.md (14586b), _meta.json (136b)\n\nFile v1.2.0:SKILL.md\n\n---\r\nname: zhihu-fetcher\r\ndescription: \"知乎收藏夹与文章内容抓取：API/Playwright 多级降级、Cookie 持久化与保活、批量正文与图片、断点续传、可选写入 Obsidian。| Zhihu collection scraping, batch article fetch, Obsidian export.\"\r\nversion: \"1.2.0\"\r\nuser-invocable: true\r\nargument-hint: \"[可选：收藏夹 URL 或 ID、单篇链接、输出目录、Vault 路径]\"\r\nallowed-tools: Read, Write, Edit, Grep, Glob, Bash, WebFetch\r\n---\r\n\r\n# 知乎数据抓取\r\n\r\n从知乎获取**收藏夹文章列表**与**正文 Markdown**（含图片本地化），支持写入 **Obsidian** 知识库。命令与路径约定见下文；可视化说明见仓库根目录 [`README.md`](README.md)。\r\n\r\n---\r\n\r\n## 环境与约定\r\n\r\n- **语言**：默认与用户语种一致。\r\n- **技能根目录**：下文 `${CLAUDE_SKILL_DIR}` 表示本 skill 仓库根目录（部分宿主 UI 中写作 **`{baseDir}`**，含义相同）。脚本均在 **`scripts/`** 下。\r\n- **工作区目录**：脚本默认将 Cookie、浏览器用户数据、默认文章输出等放在 **`OPENCLAW_WORKSPACE`** 环境变量指定的目录；未设置时为 **`~/.openclaw/workspace/`**。\r\n- **依赖**：在 **`scripts/`** 下执行 **`pip install -r requirements.txt`**，并 **`playwright install chromium`**。\r\n\r\n### 登录与可选页面验证\r\n\r\n- **`zhihu_login.py`**：打开浏览器等待登录，默认以检测到 **`z_c0`** 为成功条件即可结束（不要求额外跳转）。\r\n- **可选二次校验**：若用户希望登录后再确认「某一内需登录页」是否可访问（如某收藏夹页、专栏后台、关注动态等），属**可选项**，不设则不执行：\r\n  - **环境变量** **`ZHIHU_VERIFY_URL`**：值为完整 **`http://` 或 `https://`** URL；\r\n  - **或**命令行第一个参数传入同一完整 URL：`python \"${CLAUDE_SKILL_DIR}/scripts/zhihu_login.py\" \"https://www.zhihu.com/...\"`。\r\n  - 脚本会访问该 URL，若正文仍出现知乎通用提示「请登录后查看」，则提示可能未登录完成；否则认为当前会话可访问该页。**不限定于收藏夹**，任意知乎链接均可（只要登录态相关）。\r\n- **`zhihu_relogin.py`**：Cookie 失效、需重新登录并写回 **`zhihu_cookies.json`** 时使用（会打开浏览器）。\r\n\r\n---\r\n\r\n## 触发条件\r\n\r\n在用户使用以下任一方式时启用本技能：\r\n\r\n- 明确提及：知乎、Zhihu、专栏、收藏夹、文章抓取、批量下载、Cookie、验证码、Obsidian、知识库同步等\r\n- 粘贴 **zhihu.com** / **zhuanlan.zhihu.com** 链接并希望获取正文或列表\r\n- 需要 **断点续传**、**图片落盘**、**反爬 / Stealth** 相关协助\r\n\r\n---\r\n\r\n## 工具与脚本路由\r\n\r\n按任务选用能力；具体工具名以当前 Agent 环境为准。\r\n\r\n### 常见任务与建议方式\r\n\r\n| 任务 | 建议方式 |\r\n|------|----------|\r\n| 获取收藏夹 JSON 列表 | **`Bash`** → `python \"${CLAUDE_SKILL_DIR}/scripts/fetch_zhihu_collection.py\" <收藏夹URL或ID>`；优先 API，失败降级 Playwright DOM |\r\n| 获取个人主页点赞/收藏历史 | **`Bash`** → `python \"${CLAUDE_SKILL_DIR}/scripts/fetch_zhihu_history.py\" <people URL 或 slug> <起始时间ISO> <输出.json> [--until <结束时间ISO>]`；按活动时间保留 `interaction_*` 元数据，支持断点续跑 |\r\n| 批量抓取正文与图片 | **`Bash`** → `python \"${CLAUDE_SKILL_DIR}/scripts/fetch_zhihu_batch.py\" <列表.json> [输出目录] [图片目录]`；默认输出目录见「路径约定」 |\r\n| 写入 Obsidian Vault | **`Bash`** → `python \"${CLAUDE_SKILL_DIR}/scripts/write_to_obsidian.py\" <文章目录> [Vault路径]`；Vault：命令行优先，否则环境变量 **`OBSIDIAN_VAULT`**；会先找 `<文章目录>/images`，否则兼容同级 **`zhihu_images`** |\r\n| 写入个人历史到 Obsidian | **`Bash`** → `python \"${CLAUDE_SKILL_DIR}/scripts/write_zhihu_history_to_obsidian.py\" <文章目录> <Vault路径> [.]`；默认写入 `{Vault}/知乎收藏/{分类}/`，按 URL 去重更新 |\r\n| 写入失败项清单 | **`Bash`** → `python \"${CLAUDE_SKILL_DIR}/scripts/write_zhihu_failures.py\" <Vault路径> <标签>:<progress.json> ...`；生成 `{Vault}/知乎收藏/抓取失败.md` |\r\n| Cookie 失效需人工登录 | **`Bash`** → `python \"${CLAUDE_SKILL_DIR}/scripts/zhihu_relogin.py\"`（会打开浏览器窗口） |\r\n| 首次登录辅助（可选验证页） | **`Bash`** → `zhihu_login.py`；可选 **`ZHIHU_VERIFY_URL`** 或首个参数传入完整 http(s) 链接，见「登录与可选页面验证」 |\r\n| 单篇快速验证 | **`Bash`** → `fetch_zhihu_api.py` / `fetch_zhihu_stealth.py` / `fetch_zhihu_interactive.py` / 汇总 **`fetch_zhihu.py`**（自动多策略），按场景选用 |\r\n| 读本地已抓取 Markdown、排查 `_progress.json` | **`Read`** / **`Grep`** |\r\n\r\n---\r\n\r\n## 脚本一览\r\n\r\n| 脚本 | 用途 | 典型场景 |\r\n|------|------|----------|\r\n| `fetch_zhihu_collection.py` | **收藏夹列表**，智能版 | 输出 `zhihu_collection_{id}.json` |\r\n| `fetch_zhihu_history.py` | **个人历史列表** | 点赞/收藏动态，支持 `--until`、`--fresh`、断点续跑 |\r\n| `fetch_zhihu_batch.py` | **批量抓取**，推荐 | 大量文章、图片、`images/`、`_progress.json`，失败自动重试，DOM 不足时 API 回退 |\r\n| `fetch_zhihu.py` | 自动降级抓取 | 单篇、多策略串联 |\r\n| `fetch_zhihu_api.py` | API 直连 | 快速测试 |\r\n| `fetch_zhihu_stealth.py` | Playwright 隐身 | 绕过常见自动化检测 |\r\n| `fetch_zhihu_interactive.py` | 交互式浏览器 | 登录页、验证码 |\r\n| `write_to_obsidian.py` | 写入 Obsidian | 自动检测 Vault、智能分类、`知乎收藏/` |\r\n| `write_zhihu_history_to_obsidian.py` | 写入个人历史到 Obsidian | 智能分类、互动元数据、按 URL 去重 |\r\n| `write_zhihu_failures.py` | 写入失败项清单 | 生成 `抓取失败.md` 方便人工重试 |\r\n| `zhihu_relogin.py` | 重新登录 | Cookie 不可用 |\r\n| `zhihu_login.py` | 登录辅助 | 检测 `z_c0`；可选访问 **`ZHIHU_VERIFY_URL`** / 命令行 URL 做页面级验证 |\r\n| `zhihu_login_save.py` | 登录并保存 | 按需配合 Cookie 流程 |\r\n\r\n---\r\n\r\n## 主流程（推荐执行顺序）\r\n\r\n1. **安装依赖**：`scripts/requirements.txt` + Chromium。\r\n2. **`fetch_zhihu_collection.py`** → 得到收藏夹 **JSON 列表**。\r\n3. **`fetch_zhihu_batch.py`** → 生成 **`zhihu_articles_{collectionId}/`**（含 **`_progress.json`**、**`images/`**、编号 **`*.md`**）。\r\n4. （可选）**`write_to_obsidian.py`** → 同步到 **`{Vault}/知乎收藏/{分类}/`**。\r\n\r\n中断批量任务时：**重新运行同一条** `fetch_zhihu_batch.py` 命令即可续跑（已完成 URL 记录在 `_progress.json`）。\r\n\r\n### 个人历史流程（点赞 / 收藏）\r\n\r\n适用于个人主页动态中的 **赞同了回答 / 赞同了文章 / 收藏了回答 / 收藏了文章**。时间采用 ISO 格式，**建议显式带时区**（如 `+08:00`）；若省略时区，默认按 **Asia/Shanghai** 解释，可用环境变量 **`ZHIHU_TIMEZONE`** 或 **`TZ`** 覆盖。\r\n\r\n```bash\r\n# 1. 收集活动列表（起始时间含，结束时间不含）\r\npython scripts/fetch_zhihu_history.py \\\r\n  https://www.zhihu.com/people/<slug> \\\r\n  2026-01-01T00:00:00+08:00 \\\r\n  /path/to/runtime/zhihu_history_2026-01-01_to_2026-04-05.json \\\r\n  --until 2026-04-05T00:00:00+08:00\r\n\r\n# 2. 抓取正文与图片；失败默认自动重试 3 次\r\npython scripts/fetch_zhihu_batch.py \\\r\n  /path/to/runtime/zhihu_history_2026-01-01_to_2026-04-05.json \\\r\n  /path/to/runtime/zhihu_articles_history_2026-01-01_to_2026-04-05\r\n\r\n# 3. 写入 Obsidian 的知乎收藏根目录分类文件夹\r\npython scripts/write_zhihu_history_to_obsidian.py \\\r\n  /path/to/runtime/zhihu_articles_history_2026-01-01_to_2026-04-05 \\\r\n  /path/to/ObsidianVault \\\r\n  .\r\n```\r\n\r\n历史笔记会保留：\r\n\r\n```yaml\r\ninteraction_action: \"赞同了回答\"\r\ninteraction_time: 2026-03-20T10:17:57.235000+00:00\r\ninteraction_date: 2026-03-20\r\ntags: [zhihu, 编程与开发, 赞同了回答]\r\n```\r\n\r\n历史列表中断时重新运行同一条命令即可续跑；加 `--fresh` 可忽略现有 checkpoint 重建。写入 Obsidian 时会扫描已有笔记的 `url` 并按 URL 更新，避免重复导入。\r\n\r\n---\r\n\r\n## 路径与输出约定\r\n\r\n### 批量抓取命令格式\r\n\r\n```bash\r\npython fetch_zhihu_batch.py <列表文件> [输出目录] [图片目录]\r\n```\r\n\r\n| 参数 | 说明 |\r\n|------|------|\r\n| **列表文件** | `fetch_zhihu_collection.py` 产出的 JSON |\r\n| **输出目录** | 可选；省略时默认为 **`{workspace}/zhihu_articles_{collectionId}/`**（`collectionId` 由列表文件名推导） |\r\n| **图片目录** | 可选；省略时默认为 **`{输出目录}/images/`** |\r\n\r\n### 目录结构示例\r\n\r\n```\r\nzhihu_articles_{collectionId}/\r\n├── _progress.json          # 断点续传\r\n├── images/                 # 默认图片目录\r\n│   └── ...\r\n├── 0001_文章标题.md\r\n└── ...\r\n```\r\n\r\n### 单篇文章格式要点\r\n\r\n- YAML frontmatter：`title`、`author`、`source`、`url`、`voteup`、`images` 等\r\n- 正文为 Markdown；图片引用指向本地 **`images/`** 下文件名（或脚本生成的相对路径）\r\n\r\n示例结构：\r\n\r\n```markdown\r\n---\r\ntitle: \"文章标题\"\r\nauthor: \"作者\"\r\nsource: zhihu\r\nurl: \"https://...\"\r\nvoteup: 123\r\nimages: 5\r\n---\r\n\r\n# 文章标题\r\n\r\n> 作者: xxx | 原文: [知乎链接](https://...)\r\n\r\n正文...\r\n```\r\n\r\n### 持久化文件（默认 workspace）\r\n\r\n| 用途 | 路径 |\r\n|------|------|\r\n| Cookie | `{workspace}/zhihu_cookies.json` |\r\n| Playwright 用户数据 | `{workspace}/chrome_user_data/` |\r\n| 默认文章目录 | `{workspace}/zhihu_articles_{collectionId}/` |\r\n| 默认图片目录 | `{文章输出目录}/images/` |\r\n\r\n---\r\n\r\n## Obsidian 写入要点\r\n\r\n- **Vault**：① **命令行第二个参数**（优先让用户直接写出 Vault 根路径）；② 未传时使用环境变量 **`OBSIDIAN_VAULT`**（单个路径）；③ 仍无时脚本按常见目录扫描，多个命中时再交互选择。\r\n- **分类**：优先对齐已有 **`知乎收藏/`** 子目录；否则按内容关键词；无法归类则 **「未分类」**。\r\n- **落盘**：**`{Vault}/知乎收藏/{分类}/{文章标题}.md`**；图片同步规则见 **`write_to_obsidian.py`**（目标侧常有集中 **`images`** 目录）。\r\n\r\n---\r\n\r\n## 已知问题与对策\r\n\r\n| # | 现象 / 原因 | 处理 |\r\n|---|-------------|------|\r\n| 1 | **Cookie 失效**：标题「安全验证」、`/account/unhuman` | **自动恢复**：脚本内置 3 次重试（激进保活：访问文章页+模拟阅读）；仍失败则 **`zhihu_relogin.py`** |\r\n| 2 | **收藏夹 API 分页**：带 `include` 时列表可能被截断 | **`fetch_zhihu_collection.py`** 已内置 API ↔ DOM 切换；必要时减少 `include` 或走浏览器分页 |\r\n| 3 | **反爬**：Headless 被识别 | Stealth、UA、间隔；必要时 **`fetch_zhihu_interactive.py`** |\r\n| 4 | **API 正文不完整**：`include` 只给摘要 | 批量与单篇流程中已优先 **页面 DOM** 拉全文 |\r\n| 5 | **图片下载失败** | 正文仍保留原 URL；排查网络、Referer、过期链接 |\r\n| 6 | **Windows 控制台 GBK** | 脚本已 **`sys.stdout.reconfigure(encoding='utf-8')`** |\r\n| 7 | **批量中断** | 直接再次运行 **`fetch_zhihu_batch.py`**，依赖 **`_progress.json`** |\r\n| 8 | **失败项累积** | 散发失败自动记录到 `_progress.json`（含 url/reason/title/timestamp）；连续失败 ≥5 次中断并丢弃缓存；用 **`--retry-failed`** 参数可重试 |\r\n\r\n### Cookie 保活机制\r\n\r\n脚本内置多层 Cookie 保活策略：\r\n\r\n1. **主动 TTL 检测**（每篇文章）：解析 z_c0 的 `expires` 字段，剩余 < 30 分钟时自动触发激进刷新\r\n2. **常规保活**（每 5-8 篇）：访问知乎列表页 + 模拟滚动\r\n3. **激进保活**（每 ~20 篇）：访问实际文章页 + 模拟阅读（停留 2-5 秒 + 滚动）\r\n4. **被动检测**：每次访问文章时检查是否被重定向到 `/account/unhuman` 或 `/signin`\r\n5. **自动恢复**：检测到失效时，自动尝试 3 次激进保活恢复\r\n6. **Cookie 备份**：每次保活后自动从浏览器提取最新 Cookie 保存到文件（扩展格式含 expires）\r\n7. **安全退出**：脚本结束前保存最新 Cookie + 当前进度\r\n\r\n### 失败处理策略\r\n\r\n脚本采用**两级失败处理**，区分「文章本身问题」和「环境问题」：\r\n\r\n| 场景 | 行为 | 说明 |\r\n|------|------|------|\r\n| 散发失败（中间有成功） | 记录到 `_progress.json` 的 `failed` 字段 | 视为文章本身问题（已删除/不可访问），后续跳过 |\r\n| 连续失败 ≥ 5 次 | 中断抓取，**丢弃**缓存的失败记录 | 视为环境问题（Cookie/网络），下次重试仍可跑 |\r\n\r\n**工作原理：**\r\n- 失败先缓存在内存中，不立即写入进度文件\r\n- 下一条成功时，将缓存的失败记录批量写入进度文件（确认是文章问题）\r\n- 连续失败达到阈值（5 次）时，中断抓取，丢弃缓存（保留重试机会）\r\n\r\n**相关常量：**\r\n- `CONSECUTIVE_FAIL_THRESHOLD = 5`：连续失败阈值\r\n- `CONSECUTIVE_FAIL_INTERRUPT = True`：是否在连续失败时中断\r\n\r\n**重试模式：**\r\n```bash\r\npython fetch_zhihu_batch.py <列表文件> [输出目录] [图片目录] --retry-failed\r\n```\r\n此模式会清空 `failed` 列表，只重试之前记录为失败的文章。\r\n\r\n---\r\n\r\n## 故障排查流程\r\n\r\n```\r\n正文全空？\r\n  → Cookie（含 z_c0）→ 是否跳转验证页 → zhihu_relogin.py\r\n\r\n图片失败？\r\n  → URL/网络/Referer → Markdown 中仍可保留链接\r\n\r\n批量中途停止？\r\n  → 确认 _progress.json → 原命令重跑\r\n```\r\n\r\n---\r\n\r\n## Agent 自用工作流检查清单\r\n\r\n```\r\n□ 已确认 scripts 依赖与 playwright chromium 可用；必要时提示用户设置 OPENCLAW_WORKSPACE\r\n□ 收藏夹任务：已运行 fetch_zhihu_collection.py 并得到合法 JSON，再执行 fetch_zhihu_batch.py\r\n□ 批量输出路径：知悉默认 {workspace}/zhihu_articles_* 与 images/ 子目录；第三个参数仅在自定义图片目录时需要\r\n□ Obsidian：`write_to_obsidian.py` 的文章目录含 *.md 与 images/；Vault 优先命令行路径或 **`OBSIDIAN_VAULT`**\r\n□ 遇验证页或全文为空：优先 Cookie/重登录，而非重复盲目加大并发\r\n□ 用户仅需单篇或调试：选用 fetch_zhihu_api / stealth / interactive / fetch_zhihu，避免不必要批量\r\n```\n\nFile v1.2.0:README.md\n\n<div align=\"center\">\r\n\r\n# 知乎抓取.skill\r\n\r\n> 从知乎**收藏夹列表**到**批量正文与图片**，再到 **Obsidian 自动分类入库**：API / Playwright 多级降级、Cookie 持久化与保活、断点续传。\r\n\r\n[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](LICENSE)\r\n[![Python 3.10+](https://img.shields.io/badge/Python-3.10%2B-blue.svg)](https://www.python.org/)\r\n[![Playwright](https://img.shields.io/badge/Playwright-Chromium-45ba4b.svg)](https://playwright.dev/)\r\n[![AgentSkills](https://img.shields.io/badge/AgentSkills-Standard-green)](https://agentskills.io)\r\n\r\n<br>\r\n\r\n收藏夹里上千篇文章想**归档成 Markdown**？<br>\r\n需要**配图本地化**、中断后能**接着抓**？<br>\r\n希望落库到 Obsidian，并**按主题自动分类**？<br>\r\nCookie 经常失效，想要**持久化上下文 + 保活**？\r\n\r\n**本 Skill 按 AgentSkills 约定编排全流程，入口见根目录 [`SKILL.md`](SKILL.md)，脚本集中在 `scripts/`。**\r\n\r\n[功能特性](#功能特性) · [安装](#安装) · [使用](#使用) · [项目结构](#项目结构) · [运行效果](#运行效果) · [参考文档](#参考文档)\r\n\r\n</div>\r\n\r\n---\r\n\r\n## 功能特性\r\n\r\n| 能力 | 说明 |\r\n|------|------|\r\n| 收藏夹列表 | `fetch_zhihu_collection.py` 优先 API，失败降级 Playwright DOM；输出 JSON 列表 |\r\n| 个人历史列表 | `fetch_zhihu_history.py`：个人主页点赞/收藏动态，支持时间范围、断点续跑、互动时间元数据 |\r\n| 批量抓取 | `fetch_zhihu_batch.py`：正文 Markdown、图片默认写入 `{输出目录}/images/`、`_progress.json` 断点续传、失败自动重试、API 回退 |\r\n| Cookie | 持久化浏览器上下文 + 定时保活；失效时用 `zhihu_relogin.py` 手动登录 |\r\n| 单篇 / 调试 | `fetch_zhihu.py`、`fetch_zhihu_api.py`、`fetch_zhihu_stealth.py`、`fetch_zhihu_interactive.py` 等多路径 |\r\n| Obsidian | `write_to_obsidian.py`：Vault 检测、按内容与已有「知乎收藏」结构智能分类、同步图片；`write_zhihu_history_to_obsidian.py` 支持历史项 URL 去重导入 |\r\n\r\n**依赖**：见 [`scripts/requirements.txt`](scripts/requirements.txt)，并需 `playwright install chromium`。\r\n\r\n---\r\n\r\n## 安装\r\n\r\n### Claude Code / Cursor\r\n\r\n将本仓库放到宿主约定的 skills 路径（与 [`SKILL.md`](SKILL.md) 同级为 skill 根目录），重启后在规则或技能列表中确认已加载。\r\n\r\n```bash\r\n# 示例：克隆到项目的 skills 目录\r\nmkdir -p .cursor/skills\r\ngit clone https://github.com/handsomestWei/zhihu-fetch-skill.git .cursor/skills/zhihu-fetch-skill\r\n```\r\n\r\n### 依赖\r\n\r\n```bash\r\ncd scripts\r\npip install -r requirements.txt\r\nplaywright install chromium\r\n```\r\n\r\n---\r\n\r\n## 使用\r\n\r\n在 Agent 中用自然语言描述即可，例如：知乎文章、收藏夹、批量抓取、写入 Obsidian、Cookie 失效。\r\n\r\n典型三步（路径请按本机 `{workspace}` 调整，详见 [`SKILL.md`](SKILL.md)）：\r\n\r\n```bash\r\n# 1. 收藏夹 → JSON 列表\r\npython scripts/fetch_zhihu_collection.py <收藏夹URL或ID>\r\n\r\n# 2. 批量抓取正文与图片\r\npython scripts/fetch_zhihu_batch.py <列表.json>\r\n\r\n# 3. 写入 Obsidian Vault（可选 Vault 路径）\r\npython scripts/write_to_obsidian.py <文章目录> [Vault路径]\r\n```\r\n\r\n个人历史（点赞 / 收藏）示例：\r\n\r\n```bash\r\n# 1. 个人动态 → JSON 列表（起始时间含，结束时间不含）\r\npython scripts/fetch_zhihu_history.py \\\r\n  https://www.zhihu.com/people/<slug> \\\r\n  2026-01-01T00:00:00+08:00 \\\r\n  runtime/zhihu_history_2026-01-01_to_2026-04-05.json \\\r\n  --until 2026-04-05T00:00:00+08:00\r\n\r\n# 2. 批量抓取正文与图片（失败默认自动重试 3 次）\r\npython scripts/fetch_zhihu_batch.py \\\r\n  runtime/zhihu_history_2026-01-01_to_2026-04-05.json \\\r\n  runtime/zhihu_articles_history_2026-01-01_to_2026-04-05\r\n\r\n# 3. 写入 Obsidian 的「知乎收藏/{分类}/」根分类文件夹，按 url 去重更新\r\npython scripts/write_zhihu_history_to_obsidian.py \\\r\n  runtime/zhihu_articles_history_2026-01-01_to_2026-04-05 \\\r\n  /path/to/ObsidianVault \\\r\n  .\r\n```\r\n\r\nCookie 异常时：\r\n\r\n```bash\r\npython scripts/zhihu_relogin.py\r\n```\r\n\r\n---\r\n\r\n## 项目结构\r\n\r\n本仓库遵循 [AgentSkills](https://agentskills.io)，根目录即一个 skill：\r\n\r\n```\r\nzhihu-fetch-skill/\r\n├── SKILL.md                 # 技能入口：触发条件、命令与路径约定\r\n├── README.md                # 本说明\r\n├── LICENSE\r\n├── .gitignore\r\n├── docs/                    # 文档配图（运行效果截图）\r\n│   ├── openclaw-run.jpg\r\n│   └── obs.jpg\r\n└── scripts/\r\n    ├── requirements.txt\r\n    ├── fetch_zhihu_collection.py\r\n    ├── fetch_zhihu_history.py\r\n    ├── fetch_zhihu_batch.py\r\n    ├── fetch_zhihu.py\r\n    ├── fetch_zhihu_api.py\r\n    ├── fetch_zhihu_stealth.py\r\n    ├── fetch_zhihu_interactive.py\r\n    ├── obsidian_classify.py\r\n    ├── write_to_obsidian.py\r\n    ├── write_zhihu_history_to_obsidian.py\r\n    ├── write_zhihu_failures.py\r\n    ├── zhihu_login.py\r\n    ├── zhihu_login_save.py\r\n    └── zhihu_relogin.py\r\n```\r\n\r\n默认文章与图片目录等行为以 [`SKILL.md`](SKILL.md)「批量抓取详解」「文件路径」为准。\r\n\r\n---\r\n\r\n## 运行效果\r\n\r\n**在 OpenClaw 对话中执行批量抓取**（工具输出中可见进度、剩余篇数、图片数量与 Cookie 保活提示）\r\n\r\n![OpenClaw 聊天：批量抓取进度与 Cookie 保活](./docs/openclaw-run.jpg)\r\n\r\n**写入 Obsidian 后的 Vault 结构**（「知乎收藏」下主题分类与关系图谱）\r\n\r\n![Obsidian：知乎收藏分类与关系图谱](./docs/obs.jpg)\r\n\r\n---\r\n\r\n## 参考文档\r\n\r\n- [技能入口与完整命令说明](SKILL.md)（依赖、脚本表、故障排查）\r\n- [脚本依赖清单](scripts/requirements.txt)\r\n\r\n---\r\n\r\n<div align=\"center\">\r\n\r\nMIT License © [handsomestWei](https://github.com/handsomestWei/)\r\n\r\n</div>\n\nFile v1.2.0:_meta.json\n\n{\n  \"ownerId\": \"kn71nkfhpcw7dp6pkyqrj43dgd84eg1k\",\n  \"slug\": \"zhihu-fetch-skill\",\n  \"version\": \"1.2.0\",\n  \"publishedAt\": 1781059686476\n}\n\nFile v1.2.0:scripts/requirements.txt\n\nrequests>=2.28.0\nbeautifulsoup4>=4.12.0\nplaywright>=1.40.0\n\nFile v1.2.0:skill-card.md\n\n## Description: <br>\nFetches Zhihu collection lists, articles, images, and profile activity into Markdown, with resumable batch jobs and optional Obsidian import. <br>\n\nThis skill is ready for commercial/non-commercial use. <br>\n\n## Publisher: <br>\n[handsomestwei](https://clawhub.ai/user/handsomestwei) <br>\n\n### License/Terms of Use: <br>\nMIT-0 <br>\n\n\n## Use Case: <br>\nExternal users and developers use this skill to guide an agent through authenticated Zhihu collection, article, and profile-activity capture workflows, then save the results as Markdown, images, JSON progress files, or Obsidian notes. <br>\n\n### Deployment Geography for Use: <br>\nGlobal <br>\n\n## Known Risks and Mitigations: <br>\nRisk: Authenticated Zhihu automation and stealth browsing may trigger verification or create account-policy risk. <br>\nMitigation: Use only accounts and content you are authorized to access, preferably with a low-privilege Zhihu account, and stop if Zhihu presents verification or access warnings. <br>\nRisk: The skill stores Zhihu session cookies and browser user data locally. <br>\nMitigation: Run it in a dedicated workspace, restrict access to `zhihu_cookies.json` and `chrome_user_data`, and do not share or commit those files. <br>\nRisk: Obsidian import behavior can move or delete local export files. <br>\nMitigation: Back up the vault and review the generated Markdown and image directories before importing or running write-to-Obsidian scripts. <br>\n\n\n## Reference(s): <br>\n- [ClawHub Skill Page](https://clawhub.ai/handsomestwei/zhihu-fetch-skill) <br>\n- [Skill Entry and Command Reference](SKILL.md) <br>\n- [Script Dependencies](scripts/requirements.txt) <br>\n- [Playwright](https://playwright.dev/) <br>\n- [AgentSkills](https://agentskills.io) <br>\n\n\n## Skill Output: <br>\n**Output Type(s):** [text, markdown, code, shell commands, configuration, guidance] <br>\n**Output Format:** [Markdown guidance with shell command examples; scripts produce JSON lists, Markdown article files, local images, progress files, and Obsidian notes.] <br>\n**Output Parameters:** [1D] <br>\n**Other Properties Related to Output:** [May persist Zhihu cookies and browser user data in the configured workspace; batch jobs use progress files for resume and retry behavior.] <br>\n\n## Skill Version(s): <br>\n1.2.0 (source: server release and SKILL.md frontmatter) <br>\n\n## Ethical Considerations: <br>\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment. <br>\n\nArchive v1.1.0: 15 files, 41103 bytes\n\nFiles: README.md (4650b), scripts/fetch_zhihu_api.py (5256b), scripts/fetch_zhihu_batch.py (30379b), scripts/fetch_zhihu_collection.py (10018b), scripts/fetch_zhihu_interactive.py (5435b), scripts/fetch_zhihu_stealth.py (6371b), scripts/fetch_zhihu.py (3548b), scripts/requirements.txt (59b), scripts/write_to_obsidian.py (15602b), scripts/zhihu_login_save.py (2780b), scripts/zhihu_login.py (4070b), scripts/zhihu_relogin.py (2962b), skill-card.md (2308b), SKILL.md (11725b), _meta.json (136b)\n\nFile v1.1.0:SKILL.md\n\n---\nname: zhihu-fetcher\ndescription: \"知乎收藏夹与文章内容抓取：API/Playwright 多级降级、Cookie 持久化与保活、批量正文与图片、断点续传、可选写入 Obsidian。| Zhihu collection scraping, batch article fetch, Obsidian export.\"\nversion: \"1.1.0\"\nuser-invocable: true\nargument-hint: \"[可选：收藏夹 URL 或 ID、单篇链接、输出目录、Vault 路径]\"\nallowed-tools: Read, Write, Edit, Grep, Glob, Bash, WebFetch\n---\n\n# 知乎数据抓取\n\n从知乎获取**收藏夹文章列表**与**正文 Markdown**（含图片本地化），支持写入 **Obsidian** 知识库。命令与路径约定见下文；可视化说明见仓库根目录 [`README.md`](README.md)。\n\n---\n\n## 环境与约定\n\n- **语言**：默认与用户语种一致。\n- **技能根目录**：下文 `${CLAUDE_SKILL_DIR}` 表示本 skill 仓库根目录（部分宿主 UI 中写作 **`{baseDir}`**，含义相同）。脚本均在 **`scripts/`** 下。\n- **工作区目录**：脚本默认将 Cookie、浏览器用户数据、默认文章输出等放在 **`OPENCLAW_WORKSPACE`** 环境变量指定的目录；未设置时为 **`~/.openclaw/workspace/`**。\n- **依赖**：在 **`scripts/`** 下执行 **`pip install -r requirements.txt`**，并 **`playwright install chromium`**。\n\n### 登录与可选页面验证\n\n- **`zhihu_login.py`**：打开浏览器等待登录，默认以检测到 **`z_c0`** 为成功条件即可结束（不要求额外跳转）。\n- **可选二次校验**：若用户希望登录后再确认「某一内需登录页」是否可访问（如某收藏夹页、专栏后台、关注动态等），属**可选项**，不设则不执行：\n  - **环境变量** **`ZHIHU_VERIFY_URL`**：值为完整 **`http://` 或 `https://`** URL；\n  - **或**命令行第一个参数传入同一完整 URL：`python \"${CLAUDE_SKILL_DIR}/scripts/zhihu_login.py\" \"https://www.zhihu.com/...\"`。\n  - 脚本会访问该 URL，若正文仍出现知乎通用提示「请登录后查看」，则提示可能未登录完成；否则认为当前会话可访问该页。**不限定于收藏夹**，任意知乎链接均可（只要登录态相关）。\n- **`zhihu_relogin.py`**：Cookie 失效、需重新登录并写回 **`zhihu_cookies.json`** 时使用（会打开浏览器）。\n\n---\n\n## 触发条件\n\n在用户使用以下任一方式时启用本技能：\n\n- 明确提及：知乎、Zhihu、专栏、收藏夹、文章抓取、批量下载、Cookie、验证码、Obsidian、知识库同步等\n- 粘贴 **zhihu.com** / **zhuanlan.zhihu.com** 链接并希望获取正文或列表\n- 需要 **断点续传**、**图片落盘**、**反爬 / Stealth** 相关协助\n\n---\n\n## 工具与脚本路由\n\n按任务选用能力；具体工具名以当前 Agent 环境为准。\n\n### 常见任务与建议方式\n\n| 任务 | 建议方式 |\n|------|----------|\n| 获取收藏夹 JSON 列表 | **`Bash`** → `python \"${CLAUDE_SKILL_DIR}/scripts/fetch_zhihu_collection.py\" <收藏夹URL或ID>`；优先 API，失败降级 Playwright DOM |\n| 批量抓取正文与图片 | **`Bash`** → `python \"${CLAUDE_SKILL_DIR}/scripts/fetch_zhihu_batch.py\" <列表.json> [输出目录] [图片目录]`；默认输出目录见「路径约定」 |\n| 写入 Obsidian Vault | **`Bash`** → `python \"${CLAUDE_SKILL_DIR}/scripts/write_to_obsidian.py\" <文章目录> [Vault路径]`；Vault：命令行优先，否则环境变量 **`OBSIDIAN_VAULT`**；会先找 `<文章目录>/images`，否则兼容同级 **`zhihu_images`** |\n| Cookie 失效需人工登录 | **`Bash`** → `python \"${CLAUDE_SKILL_DIR}/scripts/zhihu_relogin.py\"`（会打开浏览器窗口） |\n| 首次登录辅助（可选验证页） | **`Bash`** → `zhihu_login.py`；可选 **`ZHIHU_VERIFY_URL`** 或首个参数传入完整 http(s) 链接，见「登录与可选页面验证」 |\n| 单篇快速验证 | **`Bash`** → `fetch_zhihu_api.py` / `fetch_zhihu_stealth.py` / `fetch_zhihu_interactive.py` / 汇总 **`fetch_zhihu.py`**（自动多策略），按场景选用 |\n| 读本地已抓取 Markdown、排查 `_progress.json` | **`Read`** / **`Grep`** |\n\n---\n\n## 脚本一览\n\n| 脚本 | 用途 | 典型场景 |\n|------|------|----------|\n| `fetch_zhihu_collection.py` | **收藏夹列表**，智能版 | 输出 `zhihu_collection_{id}.json` |\n| `fetch_zhihu_batch.py` | **批量抓取**，推荐 | 大量文章、图片、`images/`、`_progress.json` |\n| `fetch_zhihu.py` | 自动降级抓取 | 单篇、多策略串联 |\n| `fetch_zhihu_api.py` | API 直连 | 快速测试 |\n| `fetch_zhihu_stealth.py` | Playwright 隐身 | 绕过常见自动化检测 |\n| `fetch_zhihu_interactive.py` | 交互式浏览器 | 登录页、验证码 |\n| `write_to_obsidian.py` | 写入 Obsidian | 自动检测 Vault、智能分类、`知乎收藏/` |\n| `zhihu_relogin.py` | 重新登录 | Cookie 不可用 |\n| `zhihu_login.py` | 登录辅助 | 检测 `z_c0`；可选访问 **`ZHIHU_VERIFY_URL`** / 命令行 URL 做页面级验证 |\n| `zhihu_login_save.py` | 登录并保存 | 按需配合 Cookie 流程 |\n\n---\n\n## 主流程（推荐执行顺序）\n\n1. **安装依赖**：`scripts/requirements.txt` + Chromium。\n2. **`fetch_zhihu_collection.py`** → 得到收藏夹 **JSON 列表**。\n3. **`fetch_zhihu_batch.py`** → 生成 **`zhihu_articles_{collectionId}/`**（含 **`_progress.json`**、**`images/`**、编号 **`*.md`**）。\n4. （可选）**`write_to_obsidian.py`** → 同步到 **`{Vault}/知乎收藏/{分类}/`**。\n\n中断批量任务时：**重新运行同一条** `fetch_zhihu_batch.py` 命令即可续跑（已完成 URL 记录在 `_progress.json`）。\n\n---\n\n## 路径与输出约定\n\n### 批量抓取命令格式\n\n```bash\npython fetch_zhihu_batch.py <列表文件> [输出目录] [图片目录]\n```\n\n| 参数 | 说明 |\n|------|------|\n| **列表文件** | `fetch_zhihu_collection.py` 产出的 JSON |\n| **输出目录** | 可选；省略时默认为 **`{workspace}/zhihu_articles_{collectionId}/`**（`collectionId` 由列表文件名推导） |\n| **图片目录** | 可选；省略时默认为 **`{输出目录}/images/`** |\n\n### 目录结构示例\n\n```\nzhihu_articles_{collectionId}/\n├── _progress.json          # 断点续传\n├── images/                 # 默认图片目录\n│   └── ...\n├── 0001_文章标题.md\n└── ...\n```\n\n### 单篇文章格式要点\n\n- YAML frontmatter：`title`、`author`、`source`、`url`、`voteup`、`images` 等\n- 正文为 Markdown；图片引用指向本地 **`images/`** 下文件名（或脚本生成的相对路径）\n\n示例结构：\n\n```markdown\n---\ntitle: \"文章标题\"\nauthor: \"作者\"\nsource: zhihu\nurl: \"https://...\"\nvoteup: 123\nimages: 5\n---\n\n# 文章标题\n\n> 作者: xxx | 原文: [知乎链接](https://...)\n\n正文...\n```\n\n### 持久化文件（默认 workspace）\n\n| 用途 | 路径 |\n|------|------|\n| Cookie | `{workspace}/zhihu_cookies.json` |\n| Playwright 用户数据 | `{workspace}/chrome_user_data/` |\n| 默认文章目录 | `{workspace}/zhihu_articles_{collectionId}/` |\n| 默认图片目录 | `{文章输出目录}/images/` |\n\n---\n\n## Obsidian 写入要点\n\n- **Vault**：① **命令行第二个参数**（优先让用户直接写出 Vault 根路径）；② 未传时使用环境变量 **`OBSIDIAN_VAULT`**（单个路径）；③ 仍无时脚本按常见目录扫描，多个命中时再交互选择。\n- **分类**：优先对齐已有 **`知乎收藏/`** 子目录；否则按内容关键词；无法归类则 **「未分类」**。\n- **落盘**：**`{Vault}/知乎收藏/{分类}/{文章标题}.md`**；图片同步规则见 **`write_to_obsidian.py`**（目标侧常有集中 **`images`** 目录）。\n\n---\n\n## 已知问题与对策\n\n| # | 现象 / 原因 | 处理 |\n|---|-------------|------|\n| 1 | **Cookie 失效**：标题「安全验证」、`/account/unhuman` | **自动恢复**：脚本内置 3 次重试（激进保活：访问文章页+模拟阅读）；仍失败则 **`zhihu_relogin.py`** |\n| 2 | **收藏夹 API 分页**：带 `include` 时列表可能被截断 | **`fetch_zhihu_collection.py`** 已内置 API ↔ DOM 切换；必要时减少 `include` 或走浏览器分页 |\n| 3 | **反爬**：Headless 被识别 | Stealth、UA、间隔；必要时 **`fetch_zhihu_interactive.py`** |\n| 4 | **API 正文不完整**：`include` 只给摘要 | 批量与单篇流程中已优先 **页面 DOM** 拉全文 |\n| 5 | **图片下载失败** | 正文仍保留原 URL；排查网络、Referer、过期链接 |\n| 6 | **Windows 控制台 GBK** | 脚本已 **`sys.stdout.reconfigure(encoding='utf-8')`** |\n| 7 | **批量中断** | 直接再次运行 **`fetch_zhihu_batch.py`**，依赖 **`_progress.json`** |\n| 8 | **失败项累积** | 散发失败自动记录到 `_progress.json`（含 url/reason/title/timestamp）；连续失败 ≥5 次中断并丢弃缓存；用 **`--retry-failed`** 参数可重试 |\n\n### Cookie 保活机制\n\n脚本内置多层 Cookie 保活策略：\n\n1. **主动 TTL 检测**（每篇文章）：解析 z_c0 的 `expires` 字段，剩余 < 30 分钟时自动触发激进刷新\n2. **常规保活**（每 5-8 篇）：访问知乎列表页 + 模拟滚动\n3. **激进保活**（每 ~20 篇）：访问实际文章页 + 模拟阅读（停留 2-5 秒 + 滚动）\n4. **被动检测**：每次访问文章时检查是否被重定向到 `/account/unhuman` 或 `/signin`\n5. **自动恢复**：检测到失效时，自动尝试 3 次激进保活恢复\n6. **Cookie 备份**：每次保活后自动从浏览器提取最新 Cookie 保存到文件（扩展格式含 expires）\n7. **安全退出**：脚本结束前保存最新 Cookie + 当前进度\n\n### 失败处理策略\n\n脚本采用**两级失败处理**，区分「文章本身问题」和「环境问题」：\n\n| 场景 | 行为 | 说明 |\n|------|------|------|\n| 散发失败（中间有成功） | 记录到 `_progress.json` 的 `failed` 字段 | 视为文章本身问题（已删除/不可访问），后续跳过 |\n| 连续失败 ≥ 5 次 | 中断抓取，**丢弃**缓存的失败记录 | 视为环境问题（Cookie/网络），下次重试仍可跑 |\n\n**工作原理：**\n- 失败先缓存在内存中，不立即写入进度文件\n- 下一条成功时，将缓存的失败记录批量写入进度文件（确认是文章问题）\n- 连续失败达到阈值（5 次）时，中断抓取，丢弃缓存（保留重试机会）\n\n**相关常量：**\n- `CONSECUTIVE_FAIL_THRESHOLD = 5`：连续失败阈值\n- `CONSECUTIVE_FAIL_INTERRUPT = True`：是否在连续失败时中断\n\n**重试模式：**\n```bash\npython fetch_zhihu_batch.py <列表文件> [输出目录] [图片目录] --retry-failed\n```\n此模式会清空 `failed` 列表，只重试之前记录为失败的文章。\n\n---\n\n## 故障排查流程\n\n```\n正文全空？\n  → Cookie（含 z_c0）→ 是否跳转验证页 → zhihu_relogin.py\n\n图片失败？\n  → URL/网络/Referer → Markdown 中仍可保留链接\n\n批量中途停止？\n  → 确认 _progress.json → 原命令重跑\n```\n\n---\n\n## Agent 自用工作流检查清单\n\n```\n□ 已确认 scripts 依赖与 playwright chromium 可用；必要时提示用户设置 OPENCLAW_WORKSPACE\n□ 收藏夹任务：已运行 fetch_zhihu_collection.py 并得到合法 JSON，再执行 fetch_zhihu_batch.py\n□ 批量输出路径：知悉默认 {workspace}/zhihu_articles_* 与 images/ 子目录；第三个参数仅在自定义图片目录时需要\n□ Obsidian：`write_to_obsidian.py` 的文章目录含 *.md 与 images/；Vault 优先命令行路径或 **`OBSIDIAN_VAULT`**\n□ 遇验证页或全文为空：优先 Cookie/重登录，而非重复盲目加大并发\n□ 用户仅需单篇或调试：选用 fetch_zhihu_api / stealth / interactive / fetch_zhihu，避免不必要批量\n```\n\nFile v1.1.0:README.md\n\n<div align=\"center\">\n\n# 知乎抓取.skill\n\n> 从知乎**收藏夹列表**到**批量正文与图片**，再到 **Obsidian 自动分类入库**：API / Playwright 多级降级、Cookie 持久化与保活、断点续传。\n\n[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](LICENSE)\n[![Python 3.10+](https://img.shields.io/badge/Python-3.10%2B-blue.svg)](https://www.python.org/)\n[![Playwright](https://img.shields.io/badge/Playwright-Chromium-45ba4b.svg)](https://playwright.dev/)\n[![AgentSkills](https://img.shields.io/badge/AgentSkills-Standard-green)](https://agentskills.io)\n\n<br>\n\n收藏夹里上千篇文章想**归档成 Markdown**？<br>\n需要**配图本地化**、中断后能**接着抓**？<br>\n希望落库到 Obsidian，并**按主题自动分类**？<br>\nCookie 经常失效，想要**持久化上下文 + 保活**？\n\n**本 Skill 按 AgentSkills 约定编排全流程，入口见根目录 [`SKILL.md`](SKILL.md)，脚本集中在 `scripts/`。**\n\n[功能特性](#功能特性) · [安装](#安装) · [使用](#使用) · [项目结构](#项目结构) · [运行效果](#运行效果) · [参考文档](#参考文档)\n\n</div>\n\n---\n\n## 功能特性\n\n| 能力 | 说明 |\n|------|------|\n| 收藏夹列表 | `fetch_zhihu_collection.py` 优先 API，失败降级 Playwright DOM；输出 JSON 列表 |\n| 批量抓取 | `fetch_zhihu_batch.py`：正文 Markdown、图片默认写入 `{输出目录}/images/`、`_progress.json` 断点续传 |\n| Cookie | 持久化浏览器上下文 + 定时保活；失效时用 `zhihu_relogin.py` 手动登录 |\n| 单篇 / 调试 | `fetch_zhihu.py`、`fetch_zhihu_api.py`、`fetch_zhihu_stealth.py`、`fetch_zhihu_interactive.py` 等多路径 |\n| Obsidian | `write_to_obsidian.py`：Vault 检测、按内容与已有「知乎收藏」结构智能分类、同步图片 |\n\n**依赖**：见 [`scripts/requirements.txt`](scripts/requirements.txt)，并需 `playwright install chromium`。\n\n---\n\n## 安装\n\n### Claude Code / Cursor\n\n将本仓库放到宿主约定的 skills 路径（与 [`SKILL.md`](SKILL.md) 同级为 skill 根目录），重启后在规则或技能列表中确认已加载。\n\n```bash\n# 示例：克隆到项目的 skills 目录\nmkdir -p .cursor/skills\ngit clone https://github.com/handsomestWei/zhihu-fetch-skill.git .cursor/skills/zhihu-fetch-skill\n```\n\n### 依赖\n\n```bash\ncd scripts\npip install -r requirements.txt\nplaywright install chromium\n```\n\n---\n\n## 使用\n\n在 Agent 中用自然语言描述即可，例如：知乎文章、收藏夹、批量抓取、写入 Obsidian、Cookie 失效。\n\n典型三步（路径请按本机 `{workspace}` 调整，详见 [`SKILL.md`](SKILL.md)）：\n\n```bash\n# 1. 收藏夹 → JSON 列表\npython scripts/fetch_zhihu_collection.py <收藏夹URL或ID>\n\n# 2. 批量抓取正文与图片\npython scripts/fetch_zhihu_batch.py <列表.json>\n\n# 3. 写入 Obsidian Vault（可选 Vault 路径）\npython scripts/write_to_obsidian.py <文章目录> [Vault路径]\n```\n\nCookie 异常时：\n\n```bash\npython scripts/zhihu_relogin.py\n```\n\n---\n\n## 项目结构\n\n本仓库遵循 [AgentSkills](https://agentskills.io)，根目录即一个 skill：\n\n```\nzhihu-fetch-skill/\n├── SKILL.md                 # 技能入口：触发条件、命令与路径约定\n├── README.md                # 本说明\n├── LICENSE\n├── .gitignore\n├── docs/                    # 文档配图（运行效果截图）\n│   ├── openclaw-run.jpg\n│   └── obs.jpg\n└── scripts/\n    ├── requirements.txt\n    ├── fetch_zhihu_collection.py\n    ├── fetch_zhihu_batch.py\n    ├── fetch_zhihu.py\n    ├── fetch_zhihu_api.py\n    ├── fetch_zhihu_stealth.py\n    ├── fetch_zhihu_interactive.py\n    ├── write_to_obsidian.py\n    ├── zhihu_login.py\n    ├── zhihu_login_save.py\n    └── zhihu_relogin.py\n```\n\n默认文章与图片目录等行为以 [`SKILL.md`](SKILL.md)「批量抓取详解」「文件路径」为准。\n\n---\n\n## 运行效果\n\n**在 OpenClaw 对话中执行批量抓取**（工具输出中可见进度、剩余篇数、图片数量与 Cookie 保活提示）\n\n![OpenClaw 聊天：批量抓取进度与 Cookie 保活](./docs/openclaw-run.jpg)\n\n**写入 Obsidian 后的 Vault 结构**（「知乎收藏」下主题分类与关系图谱）\n\n![Obsidian：知乎收藏分类与关系图谱](./docs/obs.jpg)\n\n---\n\n## 参考文档\n\n- [技能入口与完整命令说明](SKILL.md)（依赖、脚本表、故障排查）\n- [脚本依赖清单](scripts/requirements.txt)\n\n---\n\n<div align=\"center\">\n\nMIT License © [handsomestWei](https://github.com/handsomestWei/)\n\n</div>\n\nFile v1.1.0:_meta.json\n\n{\n  \"ownerId\": \"kn71nkfhpcw7dp6pkyqrj43dgd84eg1k\",\n  \"slug\": \"zhihu-fetch-skill\",\n  \"version\": \"1.1.0\",\n  \"publishedAt\": 1777637642401\n}\n\nFile v1.1.0:scripts/requirements.txt\n\nrequests>=2.28.0\nbeautifulsoup4>=4.12.0\nplaywright>=1.40.0\n\nFile v1.1.0:skill-card.md\n\n## Description: <br>\nFetches Zhihu collection lists and article bodies into Markdown with image downloads, resume support, cookie-assisted browser automation, and optional Obsidian export. <br>\n\nThis skill is ready for commercial/non-commercial use. <br>\n\n## Publisher: <br>\n[handsomestwei](https://clawhub.ai/user/handsomestwei) <br>\n\n### License/Terms of Use: <br>\nMIT-0 <br>\n\n\n## Use Case: <br>\nExternal users and developers use this skill to archive Zhihu articles or collections as Markdown, download related images, resume interrupted batch runs, and optionally organize the result in an Obsidian vault. <br>\n\n### Deployment Geography for Use: <br>\nGlobal <br>\n\n## Known Risks and Mitigations: <br>\nRisk: The skill uses logged-in browser automation with stealth behavior and stores Zhihu session cookies. <br>\nMitigation: Run it only in a dedicated workspace, use a disposable or low-risk Zhihu session where possible, and review where cookies are stored before execution. <br>\nRisk: The export and Obsidian workflows can move images or delete source Markdown files. <br>\nMitigation: Back up export folders and any Obsidian vault before import, and test the workflow on a copy before using it with important notes. <br>\n\n\n## Reference(s): <br>\n- [ClawHub Skill Page](https://clawhub.ai/handsomestwei/zhihu-fetch-skill) <br>\n- [Playwright Documentation](https://playwright.dev/) <br>\n- [Python](https://www.python.org/) <br>\n- [AgentSkills](https://agentskills.io) <br>\n\n\n## Skill Output: <br>\n**Output Type(s):** [Shell commands, Guidance, JSON, Markdown, Files, Configuration] <br>\n**Output Format:** [Markdown articles, JSON collection lists, local image files, progress metadata, and optional Obsidian vault files] <br>\n**Output Parameters:** [1D] <br>\n**Other Properties Related to Output:** [May persist cookies, browser profile data, progress records, downloaded images, and generated article files in the configured workspace.] <br>\n\n## Skill Version(s): <br>\n1.1.0 (source: server evidence and skill frontmatter) <br>\n\n## Ethical Considerations: <br>\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment. <br>\n\nArchive v1.0.0: 14 files, 36084 bytes\n\nFiles: README.md (4650b), scripts/fetch_zhihu_api.py (5256b), scripts/fetch_zhihu_batch.py (18820b), scripts/fetch_zhihu_collection.py (10018b), scripts/fetch_zhihu_interactive.py (5435b), scripts/fetch_zhihu_stealth.py (6371b), scripts/fetch_zhihu.py (3548b), scripts/requirements.txt (59b), scripts/write_to_obsidian.py (15602b), scripts/zhihu_login_save.py (2780b), scripts/zhihu_login.py (4070b), scripts/zhihu_relogin.py (2962b), SKILL.md (9635b), _meta.json (136b)\n\nFile v1.0.0:SKILL.md\n\n---\nname: zhihu-fetcher\ndescription: \"知乎收藏夹与文章内容抓取：API/Playwright 多级降级、Cookie 持久化与保活、批量正文与图片、断点续传、可选写入 Obsidian。| Zhihu collection scraping, batch article fetch, Obsidian export.\"\nversion: \"1.0.0\"\nuser-invocable: true\nargument-hint: \"[可选：收藏夹 URL 或 ID、单篇链接、输出目录、Vault 路径]\"\nallowed-tools: Read, Write, Edit, Grep, Glob, Bash, WebFetch\n---\n\n# 知乎数据抓取\n\n从知乎获取**收藏夹文章列表**与**正文 Markdown**（含图片本地化），支持写入 **Obsidian** 知识库。命令与路径约定见下文；可视化说明见仓库根目录 [`README.md`](README.md)。\n\n---\n\n## 环境与约定\n\n- **语言**：默认与用户语种一致。\n- **技能根目录**：下文 `${CLAUDE_SKILL_DIR}` 表示本 skill 仓库根目录（部分宿主 UI 中写作 **`{baseDir}`**，含义相同）。脚本均在 **`scripts/`** 下。\n- **工作区目录**：脚本默认将 Cookie、浏览器用户数据、默认文章输出等放在 **`OPENCLAW_WORKSPACE`** 环境变量指定的目录；未设置时为 **`~/.openclaw/workspace/`**。\n- **依赖**：在 **`scripts/`** 下执行 **`pip install -r requirements.txt`**，并 **`playwright install chromium`**。\n\n### 登录与可选页面验证\n\n- **`zhihu_login.py`**：打开浏览器等待登录，默认以检测到 **`z_c0`** 为成功条件即可结束（不要求额外跳转）。\n- **可选二次校验**：若用户希望登录后再确认「某一内需登录页」是否可访问（如某收藏夹页、专栏后台、关注动态等），属**可选项**，不设则不执行：\n  - **环境变量** **`ZHIHU_VERIFY_URL`**：值为完整 **`http://` 或 `https://`** URL；\n  - **或**命令行第一个参数传入同一完整 URL：`python \"${CLAUDE_SKILL_DIR}/scripts/zhihu_login.py\" \"https://www.zhihu.com/...\"`。\n  - 脚本会访问该 URL，若正文仍出现知乎通用提示「请登录后查看」，则提示可能未登录完成；否则认为当前会话可访问该页。**不限定于收藏夹**，任意知乎链接均可（只要登录态相关）。\n- **`zhihu_relogin.py`**：Cookie 失效、需重新登录并写回 **`zhihu_cookies.json`** 时使用（会打开浏览器）。\n\n---\n\n## 触发条件\n\n在用户使用以下任一方式时启用本技能：\n\n- 明确提及：知乎、Zhihu、专栏、收藏夹、文章抓取、批量下载、Cookie、验证码、Obsidian、知识库同步等\n- 粘贴 **zhihu.com** / **zhuanlan.zhihu.com** 链接并希望获取正文或列表\n- 需要 **断点续传**、**图片落盘**、**反爬 / Stealth** 相关协助\n\n---\n\n## 工具与脚本路由\n\n按任务选用能力；具体工具名以当前 Agent 环境为准。\n\n### 常见任务与建议方式\n\n| 任务 | 建议方式 |\n|------|----------|\n| 获取收藏夹 JSON 列表 | **`Bash`** → `python \"${CLAUDE_SKILL_DIR}/scripts/fetch_zhihu_collection.py\" <收藏夹URL或ID>`；优先 API，失败降级 Playwright DOM |\n| 批量抓取正文与图片 | **`Bash`** → `python \"${CLAUDE_SKILL_DIR}/scripts/fetch_zhihu_batch.py\" <列表.json> [输出目录] [图片目录]`；默认输出目录见「路径约定」 |\n| 写入 Obsidian Vault | **`Bash`** → `python \"${CLAUDE_SKILL_DIR}/scripts/write_to_obsidian.py\" <文章目录> [Vault路径]`；Vault：命令行优先，否则环境变量 **`OBSIDIAN_VAULT`**；会先找 `<文章目录>/images`，否则兼容同级 **`zhihu_images`** |\n| Cookie 失效需人工登录 | **`Bash`** → `python \"${CLAUDE_SKILL_DIR}/scripts/zhihu_relogin.py\"`（会打开浏览器窗口） |\n| 首次登录辅助（可选验证页） | **`Bash`** → `zhihu_login.py`；可选 **`ZHIHU_VERIFY_URL`** 或首个参数传入完整 http(s) 链接，见「登录与可选页面验证」 |\n| 单篇快速验证 | **`Bash`** → `fetch_zhihu_api.py` / `fetch_zhihu_stealth.py` / `fetch_zhihu_interactive.py` / 汇总 **`fetch_zhihu.py`**（自动多策略），按场景选用 |\n| 读本地已抓取 Markdown、排查 `_progress.json` | **`Read`** / **`Grep`** |\n\n---\n\n## 脚本一览\n\n| 脚本 | 用途 | 典型场景 |\n|------|------|----------|\n| `fetch_zhihu_collection.py` | **收藏夹列表**，智能版 | 输出 `zhihu_collection_{id}.json` |\n| `fetch_zhihu_batch.py` | **批量抓取**，推荐 | 大量文章、图片、`images/`、`_progress.json` |\n| `fetch_zhihu.py` | 自动降级抓取 | 单篇、多策略串联 |\n| `fetch_zhihu_api.py` | API 直连 | 快速测试 |\n| `fetch_zhihu_stealth.py` | Playwright 隐身 | 绕过常见自动化检测 |\n| `fetch_zhihu_interactive.py` | 交互式浏览器 | 登录页、验证码 |\n| `write_to_obsidian.py` | 写入 Obsidian | 自动检测 Vault、智能分类、`知乎收藏/` |\n| `zhihu_relogin.py` | 重新登录 | Cookie 不可用 |\n| `zhihu_login.py` | 登录辅助 | 检测 `z_c0`；可选访问 **`ZHIHU_VERIFY_URL`** / 命令行 URL 做页面级验证 |\n| `zhihu_login_save.py` | 登录并保存 | 按需配合 Cookie 流程 |\n\n---\n\n## 主流程（推荐执行顺序）\n\n1. **安装依赖**：`scripts/requirements.txt` + Chromium。\n2. **`fetch_zhihu_collection.py`** → 得到收藏夹 **JSON 列表**。\n3. **`fetch_zhihu_batch.py`** → 生成 **`zhihu_articles_{collectionId}/`**（含 **`_progress.json`**、**`images/`**、编号 **`*.md`**）。\n4. （可选）**`write_to_obsidian.py`** → 同步到 **`{Vault}/知乎收藏/{分类}/`**。\n\n中断批量任务时：**重新运行同一条** `fetch_zhihu_batch.py` 命令即可续跑（已完成 URL 记录在 `_progress.json`）。\n\n---\n\n## 路径与输出约定\n\n### 批量抓取命令格式\n\n```bash\npython fetch_zhihu_batch.py <列表文件> [输出目录] [图片目录]\n```\n\n| 参数 | 说明 |\n|------|------|\n| **列表文件** | `fetch_zhihu_collection.py` 产出的 JSON |\n| **输出目录** | 可选；省略时默认为 **`{workspace}/zhihu_articles_{collectionId}/`**（`collectionId` 由列表文件名推导） |\n| **图片目录** | 可选；省略时默认为 **`{输出目录}/images/`** |\n\n### 目录结构示例\n\n```\nzhihu_articles_{collectionId}/\n├── _progress.json          # 断点续传\n├── images/                 # 默认图片目录\n│   └── ...\n├── 0001_文章标题.md\n└── ...\n```\n\n### 单篇文章格式要点\n\n- YAML frontmatter：`title`、`author`、`source`、`url`、`voteup`、`images` 等\n- 正文为 Markdown；图片引用指向本地 **`images/`** 下文件名（或脚本生成的相对路径）\n\n示例结构：\n\n```markdown\n---\ntitle: \"文章标题\"\nauthor: \"作者\"\nsource: zhihu\nurl: \"https://...\"\nvoteup: 123\nimages: 5\n---\n\n# 文章标题\n\n> 作者: xxx | 原文: [知乎链接](https://...)\n\n正文...\n```\n\n### 持久化文件（默认 workspace）\n\n| 用途 | 路径 |\n|------|------|\n| Cookie | `{workspace}/zhihu_cookies.json` |\n| Playwright 用户数据 | `{workspace}/chrome_user_data/` |\n| 默认文章目录 | `{workspace}/zhihu_articles_{collectionId}/` |\n| 默认图片目录 | `{文章输出目录}/images/` |\n\n---\n\n## Obsidian 写入要点\n\n- **Vault**：① **命令行第二个参数**（优先让用户直接写出 Vault 根路径）；② 未传时使用环境变量 **`OBSIDIAN_VAULT`**（单个路径）；③ 仍无时脚本按常见目录扫描，多个命中时再交互选择。\n- **分类**：优先对齐已有 **`知乎收藏/`** 子目录；否则按内容关键词；无法归类则 **「未分类」**。\n- **落盘**：**`{Vault}/知乎收藏/{分类}/{文章标题}.md`**；图片同步规则见 **`write_to_obsidian.py`**（目标侧常有集中 **`images`** 目录）。\n\n---\n\n## 已知问题与对策\n\n| # | 现象 / 原因 | 处理 |\n|---|-------------|------|\n| 1 | **Cookie 失效**：标题「安全验证」、`/account/unhuman` | 持久化上下文 + 内置保活；仍失败则 **`zhihu_relogin.py`** |\n| 2 | **收藏夹 API 分页**：带 `include` 时列表可能被截断 | **`fetch_zhihu_collection.py`** 已内置 API ↔ DOM 切换；必要时减少 `include` 或走浏览器分页 |\n| 3 | **反爬**：Headless 被识别 | Stealth、UA、间隔；必要时 **`fetch_zhihu_interactive.py`** |\n| 4 | **API 正文不完整**：`include` 只给摘要 | 批量与单篇流程中已优先 **页面 DOM** 拉全文 |\n| 5 | **图片下载失败** | 正文仍保留原 URL；排查网络、Referer、过期链接 |\n| 6 | **Windows 控制台 GBK** | 脚本已 **`sys.stdout.reconfigure(encoding='utf-8')`** |\n| 7 | **批量中断** | 直接再次运行 **`fetch_zhihu_batch.py`**，依赖 **`_progress.json`** |\n\n---\n\n## 故障排查流程（摘要）\n\n```\n正文全空？\n  → Cookie（含 z_c0）→ 是否跳转验证页 → zhihu_relogin.py\n\n图片失败？\n  → URL/网络/Referer → Markdown 中仍可保留链接\n\n批量中途停止？\n  → 确认 _progress.json → 原命令重跑\n```\n\n---\n\n## Agent 自用工作流检查清单\n\n```\n□ 已确认 scripts 依赖与 playwright chromium 可用；必要时提示用户设置 OPENCLAW_WORKSPACE\n□ 收藏夹任务：已运行 fetch_zhihu_collection.py 并得到合法 JSON，再执行 fetch_zhihu_batch.py\n□ 批量输出路径：知悉默认 {workspace}/zhihu_articles_* 与 images/ 子目录；第三个参数仅在自定义图片目录时需要\n□ Obsidian：`write_to_obsidian.py` 的文章目录含 *.md 与 images/；Vault 优先命令行路径或 **`OBSIDIAN_VAULT`**\n□ 遇验证页或全文为空：优先 Cookie/重登录，而非重复盲目加大并发\n□ 用户仅需单篇或调试：选用 fetch_zhihu_api / stealth / interactive / fetch_zhihu，避免不必要批量\n```\n\nFile v1.0.0:README.md\n\n<div align=\"center\">\n\n# 知乎抓取.skill\n\n> 从知乎**收藏夹列表**到**批量正文与图片**，再到 **Obsidian 自动分类入库**：API / Playwright 多级降级、Cookie 持久化与保活、断点续传。\n\n[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](LICENSE)\n[![Python 3.10+](https://img.shields.io/badge/Python-3.10%2B-blue.svg)](https://www.python.org/)\n[![Playwright](https://img.shields.io/badge/Playwright-Chromium-45ba4b.svg)](https://playwright.dev/)\n[![AgentSkills](https://img.shields.io/badge/AgentSkills-Standard-green)](https://agentskills.io)\n\n<br>\n\n收藏夹里上千篇文章想**归档成 Markdown**？<br>\n需要**配图本地化**、中断后能**接着抓**？<br>\n希望落库到 Obsidian，并**按主题自动分类**？<br>\nCookie 经常失效，想要**持久化上下文 + 保活**？\n\n**本 Skill 按 AgentSkills 约定编排全流程，入口见根目录 [`SKILL.md`](SKILL.md)，脚本集中在 `scripts/`。**\n\n[功能特性](#功能特性) · [安装](#安装) · [使用](#使用) · [项目结构](#项目结构) · [运行效果](#运行效果) · [参考文档](#参考文档)\n\n</div>\n\n---\n\n## 功能特性\n\n| 能力 | 说明 |\n|------|------|\n| 收藏夹列表 | `fetch_zhihu_collection.py` 优先 API，失败降级 Playwright DOM；输出 JSON 列表 |\n| 批量抓取 | `fetch_zhihu_batch.py`：正文 Markdown、图片默认写入 `{输出目录}/images/`、`_progress.json` 断点续传 |\n| Cookie | 持久化浏览器上下文 + 定时保活；失效时用 `zhihu_relogin.py` 手动登录 |\n| 单篇 / 调试 | `fetch_zhihu.py`、`fetch_zhihu_api.py`、`fetch_zhihu_stealth.py`、`fetch_zhihu_interactive.py` 等多路径 |\n| Obsidian | `write_to_obsidian.py`：Vault 检测、按内容与已有「知乎收藏」结构智能分类、同步图片 |\n\n**依赖**：见 [`scripts/requirements.txt`](scripts/requirements.txt)，并需 `playwright install chromium`。\n\n---\n\n## 安装\n\n### Claude Code / Cursor\n\n将本仓库放到宿主约定的 skills 路径（与 [`SKILL.md`](SKILL.md) 同级为 skill 根目录），重启后在规则或技能列表中确认已加载。\n\n```bash\n# 示例：克隆到项目的 skills 目录\nmkdir -p .cursor/skills\ngit clone https://github.com/handsomestWei/zhihu-fetch-skill.git .cursor/skills/zhihu-fetch-skill\n```\n\n### 依赖\n\n```bash\ncd scripts\npip install -r requirements.txt\nplaywright install chromium\n```\n\n---\n\n## 使用\n\n在 Agent 中用自然语言描述即可，例如：知乎文章、收藏夹、批量抓取、写入 Obsidian、Cookie 失效。\n\n典型三步（路径请按本机 `{workspace}` 调整，详见 [`SKILL.md`](SKILL.md)）：\n\n```bash\n# 1. 收藏夹 → JSON 列表\npython scripts/fetch_zhihu_collection.py <收藏夹URL或ID>\n\n# 2. 批量抓取正文与图片\npython scripts/fetch_zhihu_batch.py <列表.json>\n\n# 3. 写入 Obsidian Vault（可选 Vault 路径）\npython scripts/write_to_obsidian.py <文章目录> [Vault路径]\n```\n\nCookie 异常时：\n\n```bash\npython scripts/zhihu_relogin.py\n```\n\n---\n\n## 项目结构\n\n本仓库遵循 [AgentSkills](https://agentskills.io)，根目录即一个 skill：\n\n```\nzhihu-fetch-skill/\n├── SKILL.md                 # 技能入口：触发条件、命令与路径约定\n├── README.md                # 本说明\n├── LICENSE\n├── .gitignore\n├── docs/                    # 文档配图（运行效果截图）\n│   ├── openclaw-run.jpg\n│   └── obs.jpg\n└── scripts/\n    ├── requirements.txt\n    ├── fetch_zhihu_collection.py\n    ├── fetch_zhihu_batch.py\n    ├── fetch_zhihu.py\n    ├── fetch_zhihu_api.py\n    ├── fetch_zhihu_stealth.py\n    ├── fetch_zhihu_interactive.py\n    ├── write_to_obsidian.py\n    ├── zhihu_login.py\n    ├── zhihu_login_save.py\n    └── zhihu_relogin.py\n```\n\n默认文章与图片目录等行为以 [`SKILL.md`](SKILL.md)「批量抓取详解」「文件路径」为准。\n\n---\n\n## 运行效果\n\n**在 OpenClaw 对话中执行批量抓取**（工具输出中可见进度、剩余篇数、图片数量与 Cookie 保活提示）\n\n![OpenClaw 聊天：批量抓取进度与 Cookie 保活](./docs/openclaw-run.jpg)\n\n**写入 Obsidian 后的 Vault 结构**（「知乎收藏」下主题分类与关系图谱）\n\n![Obsidian：知乎收藏分类与关系图谱](./docs/obs.jpg)\n\n---\n\n## 参考文档\n\n- [技能入口与完整命令说明](SKILL.md)（依赖、脚本表、故障排查）\n- [脚本依赖清单](scripts/requirements.txt)\n\n---\n\n<div align=\"center\">\n\nMIT License © [handsomestWei](https://github.com/handsomestWei/)\n\n</div>\n\nFile v1.0.0:_meta.json\n\n{\n  \"ownerId\": \"kn71nkfhpcw7dp6pkyqrj43dgd84eg1k\",\n  \"slug\": \"zhihu-fetch-skill\",\n  \"version\": \"1.0.0\",\n  \"publishedAt\": 1777458170072\n}\n\nFile v1.0.0:scripts/requirements.txt\n\nrequests>=2.28.0\nbeautifulsoup4>=4.12.0\nplaywright>=1.40.0","readmeExcerpt":"Skill: 知乎抓取.SKILL Owner: handsomestwei Summary: 知乎收藏夹与文章内容抓取：API/Playwright 多级降级、Cookie 持久化与保活、批量正文与图片、断点续传、可选写入 Obsidian。| Zhihu collection scraping, batch article fetch, Obsidian export. Tags: latest:2.2.0 Version history: v2.2.0 | 2026-08-30T15:22:02.576Z | user **Major update with significant refactor, modularization, and configuration capabilities.** - All scripts refactored into a modular package under scripts/","codeSnippets":[],"executableExamples":[{"language":"text","snippet":"scripts/\n  zhihu.py                 # 唯一 CLI\n  requirements.txt\n  zhihu_fetch/\n    core/                  # paths, limits, url, seen, times, filters, summary\n    fetch/                 # route, collection, columns, posts, follow, question, history, batch, single\n    body/                  # api, stealth, interactive\n    auth/                  # login, relogin, login_save\n    export/                # obsidian, notes, history, failures, classify"},{"language":"bash","snippet":"# 1. 收集活动列表（起始时间含，结束时间不含）\npython scripts/zhihu.py history \\\n  https://www.zhihu.com/people/<slug> \\\n  2026-01-01T00:00:00+08:00 \\\n  /path/to/runtime/zhihu_history_2026-01-01_to_2026-04-05.json \\\n  --until 2026-04-05T00:00:00+08:00\n\n# 2. 抓取正文与图片；失败默认自动重试 3 次\npython scripts/zhihu.py batch \\\n  /path/to/runtime/zhihu_history_2026-01-01_to_2026-04-05.json \\\n  /path/to/runtime/zhihu_articles_history_2026-01-01_to_2026-04-05\n\n# 3. 写入 Obsidian 的知乎收藏根目录分类文件夹\npython scripts/zhihu.py history-obsidian \\\n  /path/to/runtime/zhihu_articles_history_2026-01-01_to_2026-04-05 \\\n  /path/to/ObsidianVault \\\n  ."},{"language":"yaml","snippet":"interaction_action: \"赞同了回答\"\ninteraction_time: 2026-03-20T10:17:57.235000+00:00\ninteraction_date: 2026-03-20\ntags: [zhihu, 编程与开发, 赞同了回答]"},{"language":"bash","snippet":"# 列出并抓取（受配置默认上限）\npython scripts/zhihu.py columns https://www.zhihu.com/people/<slug>/columns\n\n# 只爬指定专栏名，每栏 2 篇\npython scripts/zhihu.py columns https://www.zhihu.com/people/<slug>/columns --column 远东轶事 --per-column 2\n\n# 仅列专栏、不抓文章\npython scripts/zhihu.py columns https://www.zhihu.com/people/<slug>/columns --list-only\n\n# 正文与图片\npython scripts/zhihu.py batch zhihu-fetch-workspace/zhihu_column_<id>.json"},{"language":"bash","snippet":"python scripts/zhihu.py route https://www.zhihu.com/people/<slug>\npython scripts/zhihu.py follow https://www.zhihu.com/people/<slug> --no-since-last\npython scripts/zhihu.py route https://www.zhihu.com/people/<slug>/posts\npython scripts/zhihu.py route https://www.zhihu.com/people/<slug>/answers\npython scripts/zhihu.py posts https://www.zhihu.com/people/<slug> --kind both --since-last --min-voteup 50 --days 30\npython scripts/zhihu.py batch zhihu-fetch-workspace/zhihu_posts_<slug>.json"},{"language":"bash","snippet":"python scripts/zhihu.py route https://www.zhihu.com/question/<id> --max-items 2\npython scripts/zhihu.py batch zhihu-fetch-workspace/zhihu_question_<id>.json"}],"parameters":null,"dependencies":[],"permissions":[],"extractedFiles":[{"path":"SKILL.md","content":"---\nname: zhihu-fetcher\ndescription: \"知乎收藏夹与文章内容抓取：API/Playwright 多级降级、Cookie 持久化与保活、批量正文与图片、断点续传、可选写入 Obsidian。| Zhihu collection scraping, batch article fetch, Obsidian export.\"\nversion: \"2.2.0\"\nuser-invocable: true\nargument-hint: \"[知乎链接：收藏夹/专栏/文章/回答/问题页/个人页；或输出目录、Vault 路径]\"\nallowed-tools: Read, Write, Edit, Grep, Glob, Bash, WebFetch\n---\n\n# 知乎数据抓取\n\n从知乎获取**收藏夹文章列表**与**正文 Markdown**（含图片本地化），支持写入 **Obsidian** 知识库。命令与路径约定见下文；可视化说明见仓库根目录 [`README.md`](README.md)。\n\n---\n\n## 环境与约定\n\n- **语言**：默认与用户语种一致。\n- **技能根目录**：本仓库根目录（含 `SKILL.md` 与 `scripts/`）。下文命令均从该目录执行，写作 `python scripts/...`。\n- **工作区目录**：Cookie、浏览器用户数据、默认文章输出等放在工作区（已 gitignore，勿提交）。\n  - 环境变量 **`ZHIHU_WORKSPACE`** 优先；\n  - 未设置时默认为技能根目录下的 **`zhihu-fetch-workspace/`**。\n- **依赖**：在 **`scripts/`** 下执行 **`pip install -r requirements.txt`**，并 **`playwright install chromium`**。\n- **命令入口**：根目录只留 [`scripts/zhihu.py`](scripts/zhihu.py)；业务代码在 [`scripts/zhihu_fetch/`](scripts/zhihu_fetch/) 分模块。统一写成 `python scripts/zhihu.py <命令>`。\n\n### 抓取上限（配置优先，对话可固化）\n\n不带条数时**不会全量爬**。读取顺序：当次命令行 → 工作区配置 → 技能根配置 → 代码默认值。\n\n| 文件 / 命令 | 作用 |\n|------|------|\n| [`zhihu_fetch_config.json`](zhihu_fetch_config.json) | **技能级**默认上限；用户说「以后默认…」时改这个并保存 |\n| `{workspace}/zhihu_fetch_config.json` | **本机覆盖**（在 gitignore 的工作区内） |\n| `python scripts/zhihu.py limits` | 查看当前生效值 |\n| `python scripts/zhihu.py limits --set collection.max_items=10` | 写入技能根配置（默认 `--where skill`） |\n| `--where workspace` | 只改本机覆盖 |\n| `--all` 或配置 `\"unlimited\": true` | 取消上限 |\n\n用户说「这次多抓一点」→ 命令行 `--max-items` / `--all`。用户说「以后默认每夹 10 篇」→ **改配置并固化**，不要只改当次命令。\n\n默认：收藏夹最多 10 个、每夹 20 篇；专栏最多 5 个、每栏 20 篇；个人文章/回答各 20 篇；问题页回答 20 条；历史/批量各 20 篇。\n\n**增量**：列表脚本支持 **`--since-last`**，对照工作区 `zhihu_url_index.json`（含 `content_updated`）、已有 `zhihu_*.json`、`_progress.json` 与 Markdown frontmatter 的 `url:`。未更新的已见 URL 跳过；列表里的更新时间**新于索引**则标 `refresh` 再抓，`batch` **不会**因 `_progress.json` 的 `completed` 跳过这些条。个人「文章」默认还会按 URL 排除已在专栏 JSON 里的篇目（两者重叠，有更新仍会刷新）。\n\n**列表过滤**（条数上限之外）：`--min-voteup N`、`--days N`、`--since ISO`。配置 `filter.min_voteup` / `filter.since_days`（`0` = 不过滤）。作用在收藏夹 / 专栏 / 文章 / 回答 / 问题页 / 跟读包。`max_items` 只计通过过滤且为 new/refresh 的条目。\n\n**登录态正文**：工作区有 Cookie 时，API / 页面 / 批量图片下载都会自动带上，降低专栏文章 403。未登录先跑 `python scripts/zhihu.py login` / `relogin`。\n\n**每次 run 摘要**：列表与 batch 结束会打印并写入 `{workspace}/zhihu_run_summary.json`（成功 / 跳过空项 / 跳过已抓 / 失败 / 403 / 需登录）。Agent 回复用户时读这份摘要；失败项仍可用 `python scripts/zhihu.py failures` 写入 Vault。\n\n### 登录与可选页面验证\n\n- **`python scripts/zhihu.py login`**：打开浏览器等待登录，默认以检测到 **`z_c0`** 为成功条件即可结束（不要求额外跳转）。\n- **可选二次校验**：若用户希望登录后再确认「某一内需登录页」是否可访问（如某收藏夹页、专栏后台、关注动态等），属**可选项**，不设则不执行：\n  - **环境变量** **`ZHIHU_VERIFY_URL`**：值为完整 **`http://` 或 `https://`** URL；\n  - **或**命令行第一个参数传入同一完整 URL：`python scripts/zhihu.py login \"https://www.zhihu.com/...\"`。\n  - 脚本会访问该 URL，若正文仍出现知乎通用提示「请登录后查看」，则提示可能未登录完成；否则认为当前会话可访问该页。**不限定于收藏夹**，任意知乎链接均可（只要登录态相关）。\n- **`python scripts/zhihu.py relogin`**：Cookie 失效、需重新登录并写回 **`zhihu_cookies.json`** 时使用（会打开浏览器）。\n\n---\n\n## 触发条件\n\n在用户使用以下任一方式时启用本技能：\n\n- 明确提及：知乎、Zhihu、专栏、收藏夹、文章抓取、批量下载、Co"},{"path":"README.md","content":"<div align=\"center\">\n\n# 知乎抓取.skill\n\n> 从知乎**收藏夹列表**到**批量正文与图片**，再到 **Obsidian 自动分类入库**：API / Playwright 多级降级、Cookie 持久化与保活、断点续传。\n\n[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](LICENSE)\n[![Python 3.10+](https://img.shields.io/badge/Python-3.10%2B-blue.svg)](https://www.python.org/)\n[![Playwright](https://img.shields.io/badge/Playwright-Chromium-45ba4b.svg)](https://playwright.dev/)\n[![AgentSkills](https://img.shields.io/badge/AgentSkills-Standard-green)](https://agentskills.io)\n\n<br>\n\n收藏夹里上千篇文章想**归档成 Markdown**？<br>\n需要**配图本地化**、中断后能**接着抓**？<br>\n希望落库到 Obsidian，并**按主题自动分类**？<br>\nCookie 经常失效，想要**持久化上下文 + 保活**？\n\n**本 Skill 按 AgentSkills 约定编排全流程，入口见根目录 [`SKILL.md`](SKILL.md)，脚本集中在 `scripts/`。**\n\n[功能特性](#功能特性) · [运行效果](#运行效果) · [安装](#安装) · [使用](#使用) · [项目结构](#项目结构) · [参考文档](#参考文档)\n\n</div>\n\n---\n\n## 功能特性\n\n| 能力 | 说明 |\n|------|------|\n| 收藏夹列表 | `zhihu.py collection`：优先 API，失败降级 Playwright DOM；`--collection 名称`、`--since-last` |\n| 用户专栏 | `zhihu.py columns`：`--column 名称`、`--since-last`，层级 JSON 可交给 batch |\n| 个人文章 / 回答 | `zhihu.py posts`：与专栏按 URL 去重；`--since-last` 只补新；内容更新会 refresh |\n| 跟读包 | `zhihu.py follow` / 裸主页 `route`：专栏 + 文章 + 回答 |\n| 问题页 | `zhihu.py question`：默认排序回答列表 |\n| 统一入口 | `zhihu.py route`：识别 `/collection/` `/columns` `/posts` `/answers` `/question/` `/p/` 回答链接 个人主页 |\n| 个人历史列表 | `zhihu.py history`：个人主页点赞/收藏动态，支持时间范围、断点续跑、互动时间元数据 |\n| 批量抓取 | `zhihu.py batch`：正文 Markdown、图片默认写入 `{输出目录}/images/`、`_progress.json` 断点续传、失败自动重试、API 回退 |\n| Cookie | 持久化浏览器上下文 + 定时保活；失效时用 `zhihu.py relogin` 手动登录 |\n| 单篇 / 调试 | `zhihu.py fetch` / `api` / `stealth` / `interactive` |\n| Obsidian | `zhihu.py obsidian`：原文镜像到 `{Vault}/知乎收藏/`；`zhihu.py notes`：笔记到并列的 `{Vault}/知乎笔记/` |\n\n**依赖**：见 [`scripts/requirements.txt`](scripts/requirements.txt)，并需 `playwright install chromium`。\n\n---\n\n## 运行效果\n\n<table width=\"100%\" border=\"1\" cellpadding=\"12\" cellspacing=\"0\">\n<tr>\n<th width=\"50%\" align=\"center\">批量抓取<br><sub>Agent 对话中的进度、剩余篇数与 Cookie 保活（OpenClaw 示例）</sub></th>\n<th width=\"50%\" align=\"center\">写入 Obsidian<br><sub>「知乎收藏」主题分类与关系图谱</sub></th>\n</tr>\n<tr>\n<td width=\"50%\" valign=\"top\" align=\"center\">\n<img src=\"docs/openclaw-run.jpg\" alt=\"Agent 对话：批量抓取进度与 Cookie 保活\" width=\"100%\" />\n</td>\n<td width=\"50%\" valign=\"top\" align=\"center\">\n<img src=\"docs/obs.jpg\" alt=\"Obsidian：知乎收藏分类与关系图谱\" width=\"100%\" />\n</td>\n</tr>\n</table>\n\n---\n\n## 安装\n\n### 加载技能\n\n将本仓库放到 Agent 宿主约定的 skills 路径（与 [`SKILL.md`](SKILL.md) 同级为 skill 根目录），重启后在技能列表中确认已加载。路径因宿主而异，例如 Claude Code、Cursor、OpenClaw 等。\n\n```bash\n# 示例：克隆到项目的 skills 目录（按宿主调整目标路径）\ngit clone https://github.com/handsomestWei/zhihu-fetch-skill.git\n```\n\n### 依赖\n\n```bash\ncd scripts\npip install -r requirements.txt\nplaywright install chromium\n```\n\n仓库根目录运行测试（访问真实知乎，仅最近少量条目；账号见 `tests/live_profile.py`）：\n\n```bash\npython -m pytest\n```\n\n抓取上限集中在根目录 [`zhihu_fetch_config.json`](zhihu_fetch_config.json)，运行时优先读配置；对话里改默认用 `python scripts/zhihu.py limits --set key=value`。详情见 [`SKILL.md`](SKILL.md)。\n\n---\n\n## 使用\n\n在 Agent 中用自然语言描述即可，例如：知乎文章、收藏夹、批量抓取、写入 Obsidian、Cookie 失效。\n"},{"path":"_meta.json","content":"{\n  \"ownerId\": \"kn71nkfhpcw7dp6pkyqrj43dgd84eg1k\",\n  \"slug\": \"zhihu-fetch-skill\",\n  \"version\": \"2.2.0\",\n  \"publishedAt\": 1788103322576\n}"},{"path":"scripts/requirements.txt","content":"requests>=2.28.0\nbeautifulsoup4>=4.12.0\nplaywright>=1.40.0\npytest>=8.0.0\npytest>=8.0.0"},{"path":"skill-card.md","content":"## Description:\n\nFetches Zhihu collection lists, articles, answers, questions, user posts, and history into local JSON and Markdown, with Playwright fallback, cookie persistence, image localization, resumable batch runs, and optional Obsidian export.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[handsomestwei](https://clawhub.ai/user/handsomestwei)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers and agent users use this skill to collect Zhihu content, resume interrupted crawls, save article bodies and images as Markdown, and optionally organize mirrored content and notes inside an Obsidian vault.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: The skill stores Zhihu session cookies and can reuse them while fetching pages and images.\n\nMitigation: Use a dedicated workspace, restrict access to zhihu_cookies.json, and delete or rotate cookies when the run is complete.\n\nRisk: Untrusted batch JSON or fetched article content may drive authenticated requests or introduce unsafe Markdown content.\n\nMitigation: Run only trusted batch lists, review fetched content before importing it into a knowledge base, and avoid sensitive logged-in sessions for untrusted inputs.\n\nRisk: Obsidian export commands write into the selected vault and may update or delete source mirror files during import workflows.\n\nMitigation: Back up the vault, test with a small batch first, and verify the target vault path before running export commands.\n\n## Reference(s):\n\n- [ClawHub skill page](https://clawhub.ai/handsomestwei/skills/zhihu-fetch-skill)\n- [Publisher profile](https://clawhub.ai/user/handsomestwei)\n- [Python](https://www.python.org/)\n- [Playwright](https://playwright.dev/)\n- [AgentSkills](https://agentskills.io)\n\n## Skill Output:\n\n**Output Type(s):** [text, markdown, shell commands, configuration, guidance]\n\n**Output Format:** [Markdown guidance with inline shell commands; generated artifacts include JSON lists, Markdown articles, local image files, progress files, run summaries, and Obsidian notes.]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [Uses local workspace state for cookies, browser data, progress checkpoints, crawl limits, URL indexes, and run summaries.]\n\n## Skill Version(s):\n\n2.2.0 (source: server release and skill frontmatter)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment."}],"languages":[],"docsSourceLabel":"CLAWHUB","editorialOverview":"知乎收藏夹与文章内容抓取：API/Playwright 多级降级、Cookie 持久化与保活、批量正文与图片、断点续传、可选写入 Obsidian。| Zhihu collection scraping, batch article fetch, Obsidian export. Skill: 知乎抓取.SKILL Owner: handsomestwei Summary: 知乎收藏夹与文章内容抓取：API/Playwright 多级降级、Cookie 持久化与保活、批量正文与图片、断点续传、可选写入 Obsidian。| Zhihu collection scraping, batch article fetch, Obsidian export. Tags: latest:2.2.0 Version history: v2.2.0 | 2026-08-30T15:22:02.576Z | user **Major update with significant refactor, modularization, and configuration capabilities.** - All scripts refactored into a modular package under scripts/","editorialQuality":{"score":100,"threshold":65,"status":"ready","wordCount":1369,"uniquenessScore":48,"reasons":[]}},"media":{"evidence":{"source":"no-media","verified":false,"confidence":"low","updatedAt":"2026-10-10T16:48:40.333Z","emptyReason":"No screenshots, media assets, or demo links are available."},"primaryImageUrl":null,"mediaAssetCount":0,"assets":[],"demoUrl":null},"ownerResources":{"evidence":{"source":"unclaimed","verified":false,"confidence":"low","updatedAt":"2026-10-10T16:48:40.333Z","emptyReason":"This page has not been claimed by the agent owner."},"hasCustomPage":false,"customPageUpdatedAt":null,"customLinks":[],"structuredLinks":{"docsUrl":null,"demoUrl":null,"supportUrl":null,"pricingUrl":null,"statusUrl":null},"customPage":null},"relatedAgents":{"evidence":{"source":"protocol-neighbors","verified":false,"confidence":"medium","updatedAt":"2026-10-10T21:52:28.788Z","emptyReason":null},"items":[{"id":"8ebccd8e-3863-4187-8355-c3f14e1f9edf","entityType":"agent","canonicalPath":"/agent/iofficeai-aionui","slug":"iofficeai-aionui","name":"AionUi","description":"Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!","url":"https://github.com/iOfficeAI/AionUi","homepage":"https://www.aionui.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-10-09T19:11:12.944Z","createdAt":"2026-02-25T03:38:16.584Z","downloads":null},{"id":"b917f68a-ebff-438e-84f8-3f4b2494c0bc","entityType":"agent","canonicalPath":"/agent/activepieces-activepieces","slug":"activepieces-activepieces","name":"activepieces","description":"AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents","url":"https://github.com/activepieces/activepieces","homepage":"https://www.activepieces.com","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-15T02:22:12.426Z","createdAt":"2026-02-25T03:38:12.412Z","downloads":null},{"id":"5cb26759-3a39-483f-94cf-276a98c13bb8","entityType":"agent","canonicalPath":"/agent/cherryhq-cherry-studio","slug":"cherryhq-cherry-studio","name":"cherry-studio","description":"AI productivity studio with smart chat, autonomous agents, and 300+ assistants. Unified access to frontier LLMs","url":"https://github.com/CherryHQ/cherry-studio","homepage":"https://cherry-ai.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-11T14:38:40.986Z","createdAt":"2026-02-25T03:38:19.379Z","downloads":null},{"id":"6f6582d0-5d76-4f0f-b81d-86520247950b","entityType":"agent","canonicalPath":"/agent/copilotkit-copilotkit","slug":"copilotkit-copilotkit","name":"CopilotKit","description":"The Frontend for Agents & Generative UI. React + Angular","url":"https://github.com/CopilotKit/CopilotKit","homepage":"https://docs.copilotkit.ai","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-03-25T09:50:57.846Z","createdAt":"2026-02-25T03:39:14.617Z","downloads":null}],"links":{"hub":"/agent","source":"/agent/source/clawhub","protocols":[{"label":"OpenClaw","href":"/agent/protocol/openclew"}]}}}