agentCLAWHUBUnverified

知乎抓取.SKILL

知乎收藏夹与文章内容抓取:API/Playwright 多级降级、Cookie 持久化与保活、批量正文与图片、断点续传、可选写入 Obsidian。| Zhihu collection scraping, batch article fetch, Obsidian export. Skill: 知乎抓取.SKILL Owner: handsomestwei Summary: 知乎收藏夹与文章内容抓取:API/Playwright 多级降级、Cookie 持久化与保活、批量正文与图片、断点续传、可选写入 Obsidian。| Zhihu collection scraping, batch article fetch, Obsidian export. Tags: latest:2.2.0 Version history: v2.2.0 | 2026-08-30T15:22:02.576Z | user **Major update with significant refactor, modularization, and configuration capabilities.** - All scripts refactored into a modular package under scripts/

OpenClaw

Rank

62

Safety

84

Downloads

1.3k

Updated

Oct 10, 2026

Version

2.2.0

Source

CLAWHUB

About

What it does, and when to use it.

Capability contract not published. No trust telemetry is available yet. 1.3K downloads reported by the source. Last updated 10/10/2026.

Avoid when

  • Contract metadata is missing or unavailable for deterministic execution.

Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing

Public facts

Every fact links back to the source it came from.

Vendor
Clawhubvendor · observed Oct 10, 2026
Protocol compatibility
OpenClawcompatibility · observed Oct 10, 2026
Adoption signal
1.3K downloadsadoption · observed Oct 10, 2026
Latest release
2.2.0release · observed Aug 30, 2026
Handshake status
UNKNOWNsecurity

Install and run

Setup complexity: medium.

clawhub skill install s17cne95kry5y91xh399564e9584ejmf:zhihu-fetch-skill
  1. Python environment detected. Create a strict virtual environment (`python -m venv .venv`) before installing dependencies to prevent system-level package conflicts.
  2. Setup complexity is LOW. This package is likely designed for quick installation with minimal external side-effects.
  3. Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.

Contract: missing

curl -s "https://www.xpersona.co/api/v1/agents/clawhub-handsomestwei-zhihu-fetch-skill/snapshot"

Documentation

CLAWHUB

69,869 characters of source documentation, loaded on request.

Extracted files

5 files captured from the source.

SKILL.md

---
name: zhihu-fetcher
description: "知乎收藏夹与文章内容抓取:API/Playwright 多级降级、Cookie 持久化与保活、批量正文与图片、断点续传、可选写入 Obsidian。| Zhihu collection scraping, batch article fetch, Obsidian export."
version: "2.2.0"
user-invocable: true
argument-hint: "[知乎链接:收藏夹/专栏/文章/回答/问题页/个人页;或输出目录、Vault 路径]"
allowed-tools: Read, Write, Edit, Grep, Glob, Bash, WebFetch
---

# 知乎数据抓取

从知乎获取**收藏夹文章列表**与**正文 Markdown**(含图片本地化),支持写入 **Obsidian** 知识库。命令与路径约定见下文;可视化说明见仓库根目录 [`README.md`](README.md)。

---

## 环境与约定

- **语言**:默认与用户语种一致。
- **技能根目录**:本仓库根目录(含 `SKILL.md` 与 `scripts/`)。下文命令均从该目录执行,写作 `python scripts/...`。
- **工作区目录**:Cookie、浏览器用户数据、默认文章输出等放在工作区(已 gitignore,勿提交)。
  - 环境变量 **`ZHIHU_WORKSPACE`** 优先;
  - 未设置时默认为技能根目录下的 **`zhihu-fetch-workspace/`**。
- **依赖**:在 **`scripts/`** 下执行 **`pip install -r requirements.txt`**,并 **`playwright install chromium`**。
- **命令入口**:根目录只留 [`scripts/zhihu.py`](scripts/zhihu.py);业务代码在 [`scripts/zhihu_fetch/`](scripts/zhihu_fetch/) 分模块。统一写成 `python scripts/zhihu.py <命令>`。

### 抓取上限(配置优先,对话可固化)

不带条数时**不会全量爬**。读取顺序:当次命令行 → 工作区配置 → 技能根配置 → 代码默认值。

| 文件 / 命令 | 作用 |
|------|------|
| [`zhihu_fetch_config.json`](zhihu_fetch_config.json) | **技能级**默认上限;用户说「以后默认…」时改这个并保存 |
| `{workspace}/zhihu_fetch_config.json` | **本机覆盖**(在 gitignore 的工作区内) |
| `python scripts/zhihu.py limits` | 查看当前生效值 |
| `python scripts/zhihu.py limits --set collection.max_items=10` | 写入技能根配置(默认 `--where skill`) |
| `--where workspace` | 只改本机覆盖 |
| `--all` 或配置 `"unlimited": true` | 取消上限 |

用户说「这次多抓一点」→ 命令行 `--max-items` / `--all`。用户说「以后默认每夹 10 篇」→ **改配置并固化**,不要只改当次命令。

默认:收藏夹最多 10 个、每夹 20 篇;专栏最多 5 个、每栏 20 篇;个人文章/回答各 20 篇;问题页回答 20 条;历史/批量各 20 篇。

**增量**:列表脚本支持 **`--since-last`**,对照工作区 `zhihu_url_index.json`(含 `content_updated`)、已有 `zhihu_*.json`、`_progress.json` 与 Markdown frontmatter 的 `url:`。未更新的已见 URL 跳过;列表里的更新时间**新于索引**则标 `refresh` 再抓,`batch` **不会**因 `_progress.json` 的 `completed` 跳过这些条。个人「文章」默认还会按 URL 排除已在专栏 JSON 里的篇目(两者重叠,有更新仍会刷新)。

**列表过滤**(条数上限之外):`--min-voteup N`、`--days N`、`--since ISO`。配置 `filter.min_voteup` / `filter.since_days`(`0` = 不过滤)。作用在收藏夹 / 专栏 / 文章 / 回答 / 问题页 / 跟读包。`max_items` 只计通过过滤且为 new/refresh 的条目。

**登录态正文**:工作区有 Cookie 时,API / 页面 / 批量图片下载都会自动带上,降低专栏文章 403。未登录先跑 `python scripts/zhihu.py login` / `relogin`。

**每次 run 摘要**:列表与 batch 结束会打印并写入 `{workspace}/zhihu_run_summary.json`(成功 / 跳过空项 / 跳过已抓 / 失败 / 403 / 需登录)。Agent 回复用户时读这份摘要;失败项仍可用 `python scripts/zhihu.py failures` 写入 Vault。

### 登录与可选页面验证

- **`python scripts/zhihu.py login`**:打开浏览器等待登录,默认以检测到 **`z_c0`** 为成功条件即可结束(不要求额外跳转)。
- **可选二次校验**:若用户希望登录后再确认「某一内需登录页」是否可访问(如某收藏夹页、专栏后台、关注动态等),属**可选项**,不设则不执行:
  - **环境变量** **`ZHIHU_VERIFY_URL`**:值为完整 **`http://` 或 `https://`** URL;
  - **或**命令行第一个参数传入同一完整 URL:`python scripts/zhihu.py login "https://www.zhihu.com/..."`。
  - 脚本会访问该 URL,若正文仍出现知乎通用提示「请登录后查看」,则提示可能未登录完成;否则认为当前会话可访问该页。**不限定于收藏夹**,任意知乎链接均可(只要登录态相关)。
- **`python scripts/zhihu.py relogin`**:Cookie 失效、需重新登录并写回 **`zhihu_cookies.json`** 时使用(会打开浏览器)。

---

## 触发条件

在用户使用以下任一方式时启用本技能:

- 明确提及:知乎、Zhihu、专栏、收藏夹、文章抓取、批量下载、Co

README.md

<div align="center">

# 知乎抓取.skill

> 从知乎**收藏夹列表**到**批量正文与图片**,再到 **Obsidian 自动分类入库**:API / Playwright 多级降级、Cookie 持久化与保活、断点续传。

[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](LICENSE)
[![Python 3.10+](https://img.shields.io/badge/Python-3.10%2B-blue.svg)](https://www.python.org/)
[![Playwright](https://img.shields.io/badge/Playwright-Chromium-45ba4b.svg)](https://playwright.dev/)
[![AgentSkills](https://img.shields.io/badge/AgentSkills-Standard-green)](https://agentskills.io)

<br>

收藏夹里上千篇文章想**归档成 Markdown**?<br>
需要**配图本地化**、中断后能**接着抓**?<br>
希望落库到 Obsidian,并**按主题自动分类**?<br>
Cookie 经常失效,想要**持久化上下文 + 保活**?

**本 Skill 按 AgentSkills 约定编排全流程,入口见根目录 [`SKILL.md`](SKILL.md),脚本集中在 `scripts/`。**

[功能特性](#功能特性) · [运行效果](#运行效果) · [安装](#安装) · [使用](#使用) · [项目结构](#项目结构) · [参考文档](#参考文档)

</div>

---

## 功能特性

| 能力 | 说明 |
|------|------|
| 收藏夹列表 | `zhihu.py collection`:优先 API,失败降级 Playwright DOM;`--collection 名称`、`--since-last` |
| 用户专栏 | `zhihu.py columns`:`--column 名称`、`--since-last`,层级 JSON 可交给 batch |
| 个人文章 / 回答 | `zhihu.py posts`:与专栏按 URL 去重;`--since-last` 只补新;内容更新会 refresh |
| 跟读包 | `zhihu.py follow` / 裸主页 `route`:专栏 + 文章 + 回答 |
| 问题页 | `zhihu.py question`:默认排序回答列表 |
| 统一入口 | `zhihu.py route`:识别 `/collection/` `/columns` `/posts` `/answers` `/question/` `/p/` 回答链接 个人主页 |
| 个人历史列表 | `zhihu.py history`:个人主页点赞/收藏动态,支持时间范围、断点续跑、互动时间元数据 |
| 批量抓取 | `zhihu.py batch`:正文 Markdown、图片默认写入 `{输出目录}/images/`、`_progress.json` 断点续传、失败自动重试、API 回退 |
| Cookie | 持久化浏览器上下文 + 定时保活;失效时用 `zhihu.py relogin` 手动登录 |
| 单篇 / 调试 | `zhihu.py fetch` / `api` / `stealth` / `interactive` |
| Obsidian | `zhihu.py obsidian`:原文镜像到 `{Vault}/知乎收藏/`;`zhihu.py notes`:笔记到并列的 `{Vault}/知乎笔记/` |

**依赖**:见 [`scripts/requirements.txt`](scripts/requirements.txt),并需 `playwright install chromium`。

---

## 运行效果

<table width="100%" border="1" cellpadding="12" cellspacing="0">
<tr>
<th width="50%" align="center">批量抓取<br><sub>Agent 对话中的进度、剩余篇数与 Cookie 保活(OpenClaw 示例)</sub></th>
<th width="50%" align="center">写入 Obsidian<br><sub>「知乎收藏」主题分类与关系图谱</sub></th>
</tr>
<tr>
<td width="50%" valign="top" align="center">
<img src="docs/openclaw-run.jpg" alt="Agent 对话:批量抓取进度与 Cookie 保活" width="100%" />
</td>
<td width="50%" valign="top" align="center">
<img src="docs/obs.jpg" alt="Obsidian:知乎收藏分类与关系图谱" width="100%" />
</td>
</tr>
</table>

---

## 安装

### 加载技能

将本仓库放到 Agent 宿主约定的 skills 路径(与 [`SKILL.md`](SKILL.md) 同级为 skill 根目录),重启后在技能列表中确认已加载。路径因宿主而异,例如 Claude Code、Cursor、OpenClaw 等。

```bash
# 示例:克隆到项目的 skills 目录(按宿主调整目标路径)
git clone https://github.com/handsomestWei/zhihu-fetch-skill.git
```

### 依赖

```bash
cd scripts
pip install -r requirements.txt
playwright install chromium
```

仓库根目录运行测试(访问真实知乎,仅最近少量条目;账号见 `tests/live_profile.py`):

```bash
python -m pytest
```

抓取上限集中在根目录 [`zhihu_fetch_config.json`](zhihu_fetch_config.json),运行时优先读配置;对话里改默认用 `python scripts/zhihu.py limits --set key=value`。详情见 [`SKILL.md`](SKILL.md)。

---

## 使用

在 Agent 中用自然语言描述即可,例如:知乎文章、收藏夹、批量抓取、写入 Obsidian、Cookie 失效。

_meta.json

{
  "ownerId": "kn71nkfhpcw7dp6pkyqrj43dgd84eg1k",
  "slug": "zhihu-fetch-skill",
  "version": "2.2.0",
  "publishedAt": 1788103322576
}

scripts/requirements.txt

requests>=2.28.0
beautifulsoup4>=4.12.0
playwright>=1.40.0
pytest>=8.0.0
pytest>=8.0.0

skill-card.md

## Description:

Fetches Zhihu collection lists, articles, answers, questions, user posts, and history into local JSON and Markdown, with Playwright fallback, cookie persistence, image localization, resumable batch runs, and optional Obsidian export.

This skill is ready for commercial/non-commercial use.

## Publisher:

[handsomestwei](https://clawhub.ai/user/handsomestwei)

### License/Terms of Use:

MIT-0

## Use Case:

Developers and agent users use this skill to collect Zhihu content, resume interrupted crawls, save article bodies and images as Markdown, and optionally organize mirrored content and notes inside an Obsidian vault.

### Deployment Geography for Use:

Global

## Known Risks and Mitigations:

Risk: The skill stores Zhihu session cookies and can reuse them while fetching pages and images.

Mitigation: Use a dedicated workspace, restrict access to zhihu_cookies.json, and delete or rotate cookies when the run is complete.

Risk: Untrusted batch JSON or fetched article content may drive authenticated requests or introduce unsafe Markdown content.

Mitigation: Run only trusted batch lists, review fetched content before importing it into a knowledge base, and avoid sensitive logged-in sessions for untrusted inputs.

Risk: Obsidian export commands write into the selected vault and may update or delete source mirror files during import workflows.

Mitigation: Back up the vault, test with a small batch first, and verify the target vault path before running export commands.

## Reference(s):

- [ClawHub skill page](https://clawhub.ai/handsomestwei/skills/zhihu-fetch-skill)
- [Publisher profile](https://clawhub.ai/user/handsomestwei)
- [Python](https://www.python.org/)
- [Playwright](https://playwright.dev/)
- [AgentSkills](https://agentskills.io)

## Skill Output:

**Output Type(s):** [text, markdown, shell commands, configuration, guidance]

**Output Format:** [Markdown guidance with inline shell commands; generated artifacts include JSON lists, Markdown articles, local image files, progress files, run summaries, and Obsidian notes.]

**Output Parameters:** [1D]

**Other Properties Related to Output:** [Uses local workspace state for cookies, browser data, progress checkpoints, crawl limits, URL indexes, and run summaries.]

## Skill Version(s):

2.2.0 (source: server release and skill frontmatter)

## Ethical Considerations:

Users should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.
Github ReposUpdated 1d agoRank 70

AionUi

Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!

MCPOPENCLAW
Github ReposUpdated 6mo agoRank 70

activepieces

AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents

OPENCLAW
Github ReposUpdated 6mo agoRank 70

cherry-studio

AI productivity studio with smart chat, autonomous agents, and 300+ assistants.

MCPOPENCLAW
Github ReposUpdated 7mo agoRank 70

CopilotKit

The Frontend for Agents & Generative UI. React + Angular

OPENCLAW

Machine-readable data

The same record, as JSON, for agents and crawlers.

{
  "facts": [
    {
      "factKey": "vendor",
      "category": "vendor",
      "label": "Vendor",
      "value": "Clawhub",
      "href": "https://clawhub.ai/handsomestwei/skills/zhihu-fetch-skill",
      "sourceUrl": "https://clawhub.ai/handsomestwei/skills/zhihu-fetch-skill",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-10T16:48:40.333Z",
      "isPublic": true
    },
    {
      "factKey": "protocols",
      "category": "compatibility",
      "label": "Protocol compatibility",
      "value": "OpenClaw",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-handsomestwei-zhihu-fetch-skill/contract",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-handsomestwei-zhihu-fetch-skill/contract",
      "sourceType": "contract",
      "confidence": "medium",
      "observedAt": "2026-10-10T16:48:40.333Z",
      "isPublic": true
    },
    {
      "factKey": "traction",
      "category": "adoption",
      "label": "Adoption signal",
      "value": "1.3K downloads",
      "href": "https://clawhub.ai/handsomestwei/zhihu-fetch-skill",
      "sourceUrl": "https://clawhub.ai/handsomestwei/zhihu-fetch-skill",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-10T16:48:40.333Z",
      "isPublic": true
    },
    {
      "factKey": "latest_release",
      "category": "release",
      "label": "Latest release",
      "value": "2.2.0",
      "href": "https://clawhub.ai/handsomestwei/zhihu-fetch-skill",
      "sourceUrl": "https://clawhub.ai/handsomestwei/zhihu-fetch-skill",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-08-30T15:22:02.576Z",
      "isPublic": true
    },
    {
      "factKey": "handshake_status",
      "category": "security",
      "label": "Handshake status",
      "value": "UNKNOWN",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-handsomestwei-zhihu-fetch-skill/trust",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-handsomestwei-zhihu-fetch-skill/trust",
      "sourceType": "trust",
      "confidence": "medium",
      "observedAt": null,
      "isPublic": true
    }
  ],
  "events": [
    {
      "eventType": "release",
      "title": "Release 2.2.0",
      "description": "**Major update with significant refactor, modularization, and configuration capabilities.** - All scripts refactored into a modular package under `scripts/zhihu_fetch/` with a single CLI entrypoint: `scripts/zhihu.py`. - Unified command interface (`python scripts/zhihu.py <command>`) for all tasks: routing, listing, batch, export, login, config, and more. - Configurable crawl limits and filtering (by voteup/time), with workspace- and skill-level JSON config files. - New route system: autodetects Zhihu link type (collection, column, posts, answer, question, etc.) and dispatches the appropriate pipeline. - Optional incremental sync via `--since-last`, automatic tracking of previously seen/updated URLs using workspace index. - Batched export to Obsidian now includes original mirror and parallel note generation; all export, login, and crawling commands refactored with more robust error reporting and summaries. - Backward compatibility: legacy scripts removed; all previous functionality available via subcommands through the central entrypoint.",
      "href": "https://clawhub.ai/handsomestwei/zhihu-fetch-skill",
      "sourceUrl": "https://clawhub.ai/handsomestwei/zhihu-fetch-skill",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-08-30T15:22:02.576Z",
      "isPublic": true
    }
  ]
}

Record generated Oct 10, 2026.

Sponsored

Ads related to 知乎抓取.SKILL and adjacent AI workflows.