agentCLAWHUBUnverified

Article Fetcher(文章抓取+Notion/Obsidian知识库存档)

抓取微信公众号、小红书、豆瓣、知乎文章,自动上传 OSS 图片,LLM 智能提取关键词,一键存档到 Obsidian 本地知识库(可选 Notion)

OpenClaw

Rank

62

Safety

84

Downloads

1.6k

Updated

Oct 10, 2026

Version

1.3.6

Source

CLAWHUB

About

What it does, and when to use it.

Capability contract not published. No trust telemetry is available yet. 1.6K downloads reported by the source. Last updated 10/10/2026.

Avoid when

  • Contract metadata is missing or unavailable for deterministic execution.

Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing

Public facts

Every fact links back to the source it came from.

Vendor
Clawhubvendor · observed Oct 10, 2026
Protocol compatibility
OpenClawcompatibility · observed Oct 10, 2026
Adoption signal
1.6K downloadsadoption · observed Oct 10, 2026
Latest release
1.3.6release · observed Aug 8, 2026
Handshake status
UNKNOWNsecurity

Install and run

Setup complexity: low.

clawhub skill install s17145mdmb9rthykrxn51j1yrx849jek:article-fetcher
  1. Install using `clawhub skill install s17145mdmb9rthykrxn51j1yrx849jek:article-fetcher` in an isolated environment before connecting it to live workloads.
  2. No published capability contract is available yet, so validate auth and request/response behavior manually.
  3. Review the upstream CLAWHUB listing at https://clawhub.ai/ajayhao/article-fetcher before using production credentials.

Contract: missing

curl -s "https://www.xpersona.co/api/v1/agents/clawhub-ajayhao-article-fetcher/snapshot"

Documentation

CLAWHUB

153,960 characters of source documentation, loaded on request.

Extracted files

5 files captured from the source.

SKILL.md

---
name: article-fetcher
description: "抓取微信公众号、小红书、豆瓣、知乎文章,自动上传 OSS 图片,LLM 智能提取关键词,一键存档到 Obsidian 本地知识库(可选 Notion)"
homepage: https://github.com/AjayHao/article-fetcher
metadata:
  { "hermes": { "emoji": "📰", "version": "1.3.6", "requires": { "bins": ["python3"], "env": ["ALIYUN_OSS_AK", "ALIYUN_OSS_SK", "ALIYUN_OSS_BUCKET_ID", "ALIYUN_OSS_ENDPOINT", "NOTION_API_KEY", "LLM_API_KEY"] }, "primaryEnv": "OBSIDIAN_VAULT_PATH", "permissions": ["env:read", "net:outbound", "fs:write"], "allowedEnv": ["ALIYUN_OSS_AK", "ALIYUN_OSS_SK", "ALIYUN_OSS_BUCKET_ID", "ALIYUN_OSS_ENDPOINT", "OBSIDIAN_VAULT_PATH", "NOTION_API_KEY", "NOTION_ARTICLE_DATABASE_ID", "LLM_API_KEY", "LLM_BASE_URL", "LLM_MODEL", "WECHAT_COOKIES_FILE", "ZHIHU_COOKIES_FILE"], "securityNote": "OSS 凭证用于图片上传存储(图片外发至阿里云 OSS 为必需);Notion 和 LLM 为可选集成,不配置则跳过对应功能;配置 LLM 时文章文本(前 12000 字符)会发送至你配置的 LLM API 端点", "install": [{ "id": "pip", "kind": "pip", "packages": "requests oss2 python-dotenv beautifulsoup4 lxml notion-client markdownify pyyaml", "label": "Install Python dependencies" }, { "id": "playwright", "kind": "shell", "command": "playwright install chromium", "label": "Install Playwright Chromium browser" }] } }
---

# Article Fetcher v1.3.6

抓取微信公众号、小红书、豆瓣、知乎文章,自动上传 OSS 图床,LLM 智能关键词提取,默认存档到 Obsidian 本地知识库(可选 Notion 双写)。

## 存档模式

四种场景按需灵活切换,不存档时仅终端输出抓取结果:

| 场景 | Obsidian | Notion | 行为 |
|------|:---:|:---:|------|
| 🏠 **本地优先** | ✅ | ❌ | 存档到 Obsidian 本地知识库 |
| ☁️ **纯云端** | ❌ | ✅ | 存档到 Notion 数据库 |
| 🏠+☁️ **双写** | ✅ | ✅ | Obsidian + Notion 双存档 |
| 🔍 **预览** | ❌ | ❌ | 仅终端输出,不存档 |

## 快速开始

### 1. 安装依赖

```bash
pip install -r requirements.txt
```

### 2. 配置环境变量

skill 通过 `$AGENT_HOME/.env` 加载配置(自动兼容 Hermes / OpenClaw 等 agent)。设置方式:

```bash
# 当前 agent 为 Hermes,只需设置一次:
export AGENT_HOME=$HERMES_HOME
```

环境变量清单(写入 `$AGENT_HOME/.env` 或系统环境变量):

```bash
# ========== 必需:OSS 图床 ==========
ALIYUN_OSS_AK=your_ak
ALIYUN_OSS_SK=your_sk
ALIYUN_OSS_BUCKET_ID=your_bucket
ALIYUN_OSS_ENDPOINT=oss-cn-shanghai.aliyuncs.com

# ========== 推荐:Obsidian 本地存档 ==========
# Windows 示例:
OBSIDIAN_VAULT_PATH=D:\GitRepo\AjayObsidianVault
# macOS 示例:
# OBSIDIAN_VAULT_PATH=/Users/yourname/ObsidianVault
# Linux / 云服务器示例:
# OBSIDIAN_VAULT_PATH=/home/yourname/ObsidianVault

# ========== 可选:Notion 云端存档 ==========
NOTION_API_KEY=secret_xxx
NOTION_ARTICLE_DATABASE_ID=database_id

# ========== 可选:LLM 关键词提取(OpenAI 兼容接口,与 video-summarizer 共用配置)==========
LLM_API_KEY=***
LLM_BASE_URL=https://api.deepseek.com
LLM_MODEL=deepseek-v4-pro

# ========== 可选:Cookies(反爬,Netscape 格式)==========
WECHAT_COOKIES_FILE=~/.cookies/wechat_cookies.txt
ZHIHU_COOKIES_FILE=~/.cookies/zhihu_cookies.txt
```

### 3. 使用

```bash
cd <skill-dir>
python3 main.py "文章 URL" [标签1] [标签2]
```

**支持平台**:微信公众号 (`mp.weixin.qq.com`)、小红书 (`xiaohongshu.com` / `xhslink.com`)、豆瓣 (`douban.com`)、知乎 (`zhihu.com`)

## 离线输入(HTML / MHTML)

除 URL 外,支持直接喂入 HTML 文本或 `.mhtml` 文件(源自 wechat-article-capture 

README.md

# Article Fetcher — Hermes Skill

抓取微信公众号、小红书、豆瓣、知乎等平台文章,自动处理图片上传至阿里云 OSS,
LLM 智能提取关键词(本地词频降级),默认存档到 Obsidian 本地知识库(可选 Notion 双写)。

**版本**: 1.3.6 | **许可**: MIT | **作者**: Ajay Hao

---

## 🎯 核心能力

- **多平台支持**: 微信公众号、小红书、豆瓣、知乎
- **智能识别**: 自动识别文章来源平台
- **内容提取**: 标题、作者、发布时间、正文(HTML)、图片
- **图床集成**: 阿里云 OSS 自动上传,按 `article-001.jpg` 格式命名
- **智能关键词**: LLM 优先理解文章核心内容,本地词频分析降级兜底
- **灵活存档**: 默认 Obsidian 本地知识库(推荐),可选 Notion 双写,都不配则仅终端输出

## 🏗️ 存档模式

四种场景按需灵活切换:

| 场景 | Obsidian | Notion | 行为 |
|------|:---:|:---:|------|
| 🏠 **本地优先** | ✅ | ❌ | 存档到 Obsidian 本地知识库 |
| ☁️ **纯云端** | ❌ | ✅ | 存档到 Notion 数据库 |
| 🏠+☁️ **双写** | ✅ | ✅ | Obsidian + Notion 双存档 |
| 🔍 **预览** | ❌ | ❌ | 仅终端输出,不存档 |

## 🏗️ 处理流程

```
URL → 平台识别 (detector/) → 内容抓取 (fetchers/) → 图片上传 OSS (processors/)
   → 关键词提取 (LLM 优先 → 词频降级) → 字数统计 → Obsidian / Notion 存档 (archiver/)
```

## 📂 模块结构

```
article-fetcher/
├── main.py                      # 主入口 + 抓取器注册表 + 存档调度
├── config.py                    # 配置管理(环境变量)
├── detector/
│   └── platform_detector.py     # URL → 平台识别
├── fetchers/
│   ├── base_fetcher.py          # 基础抓取器(HTTP + Cookies)
│   ├── wechat_fetcher.py        # 微信公众号
│   ├── xhs_fetcher.py           # 小红书
│   ├── douban_fetcher.py        # 豆瓣
│   └── zhihu_fetcher.py         # 知乎
├── processors/
│   └── image_processor.py       # 图片上传 OSS
├── archiver/
│   ├── obsidian_archiver.py     # 🆕 HTML → Obsidian Markdown
│   └── notion_archiver.py       # HTML → Notion 结构化块(可选)
└── utils/
    ├── http_client.py           # HTTP 客户端(Session 复用)
    ├── logger.py                # 日志配置
    ├── word_counter.py          # 字数统计
    └── tag_extractor.py         # 关键词提取(LLM + 词频降级)
```

## ⚙️ 配置

### 环境变量

skill 读取 `$AGENT_HOME/.env`(通用,兼容 Hermes / OpenClaw 等任意 agent),回退 `$HERMES_HOME/.env`,最终回退当前目录 `.env`。

```bash
# 当前 agent 为 Hermes,设置一次即可:
export AGENT_HOME=$HERMES_HOME
```

```bash
# ========== 必需:OSS 图床 ==========
ALIYUN_OSS_AK=your_access_key_id
ALIYUN_OSS_SK=your_access_key_secret
ALIYUN_OSS_BUCKET_ID=your_bucket_name
ALIYUN_OSS_ENDPOINT=oss-cn-shanghai.aliyuncs.com

# ========== 推荐:Obsidian 本地存档 ==========
# Windows 示例:
OBSIDIAN_VAULT_PATH=D:\GitRepo\AjayObsidianVault
# macOS 示例:
# OBSIDIAN_VAULT_PATH=/Users/yourname/ObsidianVault
# Linux / 云服务器示例:
# OBSIDIAN_VAULT_PATH=/home/yourname/ObsidianVault

# ========== 可选:Notion 云端存档(双写时配置)==========
NOTION_API_KEY=secret_xxx
NOTION_ARTICLE_DATABASE_ID=xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx

# ========== 可选:LLM 关键词提取(OpenAI 兼容接口,与 video-summarizer 共用配置)==========
LLM_API_KEY=sk-xxx
LLM_BASE_URL=https://api.deepseek.com
LLM_MODEL=deepseek-v4-pro

# ========== 可选:Cookies(反爬)==========
ZHIHU_COOKIES_FILE=~/.cookies/zhihu_cookies.txt
WECHAT_COOKIES_FILE=~/.cookies/wechat_cookies.txt
```

### Obsidian 存档格式

文章存入 `{OBSIDIAN_VAULT_PATH}/1-输入-收件箱/文章收藏/`。

- **文件命名**: `{YYYY-MM-DD}_{platform}_{title}.md`
- **Frontmatter**: YAML 格式,包含 `title` / `sourc

_meta.json

{
  "ownerId": "kn73qb7xbge03wr1vsyg56snts837gqq",
  "slug": "article-fetcher",
  "version": "1.3.6",
  "publishedAt": 1786194974670
}

CHANGELOG.md

# Changelog

## v1.3.6 (2026-08-08)

### 🔒 安全修复(基于 ClawHub SkillSpector 审计)

- **P0 — lxml XXE 漏洞(High)**:`requirements.txt` 中 `lxml==6.0.4` → `lxml==6.1.0`,修复 CVE-2026-41066(`iterparse` / `ETCompatXMLParser` 默认配置 XXE)。经核查代码未使用 lxml 解析不可信 XML(MHTML 走 `email` 库、HTML 走 `BeautifulSoup('html.parser')`),升级即消除风险面。
- **P1 — 校正误导性「零网络外发」声明(Medium, Intent-Code Divergence)**:删除 README / SKILL 中「Obsidian 数据完全本地 / 零网络外发 / 不经过网络」等绝对化表述,统一为「Obsidian `.md` 本地落盘 + 文章图片上传阿里云 OSS(必需)」的准确描述。
- **P2 — LLM 外发声明补强(Medium, External Transmission)**:在 `utils/tag_extractor.py` 的 LLM 调用处补充安全注释(可选触发 / endpoint 来自 `LLM_BASE_URL` 配置 / 仅外发前 12000 字符);README 安全说明新增 LLM 文本外发条目,SKILL `securityNote` 同步声明。
- **P3 — 删除「绕过安全拦截」笔记**:移除 `references/tirith-blocking.md`(审计报告指出的 bypass-prior-blocking 笔记),并确认无其它文件引用。

## v1.3.5 (2026-08-08)

### ✨ 离线输入(HTML / MHTML)

- **新增 `fetchers/offline_parser.py`**:将 HTML 文本 / `.mhtml` 文件解析为标准 `article_data`,复用既有后半段管线(OSS 上传 → 替换 URL → 打标 → 字数 → Obsidian/Notion 归档)
- **`main.py` 抽出 `process_and_archive(article_data, platform, url, tags)`**:URL 抓取与离线输入共用同一归档流程,行为零变化
- **新增离线入口**:`archive_from_html()` / `archive_from_mhtml()`
- **CLI 扩展(`sys.argv` 手动解析,不使用 argparse,遵循安全合规约定)**:`--html`(支持 `-` 读 stdin)/`--mhtml`/`--platform`(默认 wechat)/`--url`(可选)
- **三陷阱修复(源自 wechat-article-capture 技能沉淀)**:
  - 微信懒加载:`data-src`→`src` 并 `del data-src`,根治 markdownify 读占位符
  - MHTML 编码:`get_payload(decode=True)` 取 bytes 后按 `get_content_charset()` 或 `utf-8` 解码,避免乱码
  - `data:image/svg+xml` 占位符过滤,图片列表归一化(去 query、去重)确保 URL 替换可命中

### 🔧 本地适配

- **Vault 去序号**:存档目录从 `1-收件箱/` 改为 `收件箱/`(知衍库 2026-07 架构调整)
- **移除 `distilled: false`**:蒸馏状态判定已改为文件存在性模型,不再依赖 frontmatter 标记
- **裁撤 `wechat-article-capture` 技能**:v1.3.5 的 MHTML/HTML 离线模式已覆盖其全部功能,无需单独维护

### 🐛 已知问题

- ~~**Notion 转换**:`### | title` 格式标题会被 Notion API 误判为表格分隔符~~ **已修复 (v1.3.5-local)**:根因为 `_parse_inline_html` 将内联 `<img>` 当作 image block 塞入 `rich_text`,导致 `text` 字段缺失。修复为内联图片转文本链接 `[alt](url)`。

---

## v1.3.4 (2026-07-19)

### 🔒 Tirith `requires.env` 补全

- **`NOTION_API_KEY` / `LLM_API_KEY` 加入 `requires.env`**:代码实际读取的可选凭证必须在元数据声明,消除 CRITICAL exfiltration 误报

---

## v1.3.3 (2026-07-06)

### 🔒 Tirith 安全扫描兼容

- **元数据 `allowedEnv` 声明**:显式列出所有读取的环境变量及用途,避免 `NOTION_API_KEY`/`LLM_API_KEY` 被误判为外泄
- **元数据 `securityNote`**:说明凭证用途,OSS 必需,Notion/LLM 可选
- **移除 `pip install` 文本**:代码日志和文档中的 `pip install` 语句替换为 `requirements.txt` 引用,消除供应链误报

---

## v1.3.2 (2026-07-06)

### 🐛 修复

- **文件名补日期前缀**:`_build_filename` 恢复 `{YYYY-MM-DD}_{title}.md` 格式,Obsidian 内按时间线排序
- **图片 Referer 空优先 + 403 回退**:新增 `_download_image()`,先空 Referer 请求,403 时自动回退平台 Referer

---

## v1.3.1 (2026-07-06)

### 🔒 安全审计修复

- **依赖升级**:`lxml` 6.0.2→6.0.4 (CVE-2026-41066 XXE)、`urllib3` 2.0.7→2.7.0 (多项 CVE)
- **权限声明**:SKILL.md metadata 新增 `permissions: [env:read, net:outbound, fs:write]`
- **隐私声明修正**:移除"数据完全本地"误导性表述,明确 LLM 外发范围和图片下载行为
- **安全章节补充**:新增 LLM 外发范围说明、图片下载风险提示

---

## v1.3.0 (2026-07-06)

skill-card.md

## Description:

Fetches articles from WeChat, Xiaohongshu, Douban, and Zhihu, uploads article images to Aliyun OSS, extracts keywords with optional LLM support, and archives the result to Obsidian Markdown with optional Notion sync.

This skill is ready for commercial/non-commercial use.

## Publisher:

[ajayhao](https://clawhub.ai/user/ajayhao)

### License/Terms of Use:

MIT-0

## Use Case:

Developers and knowledge workers use this skill to turn supported social and article URLs, HTML, or MHTML captures into archived knowledge-base entries. It is suited for workflows that need Markdown notes, optional Notion records, OSS-hosted images, and generated tags.

### Deployment Geography for Use:

Global

## Known Risks and Mitigations:

Risk: Untrusted HTML or MHTML can contain image URLs that the runtime fetches and uploads to OSS before image URL validation is added.

Mitigation: Process only trusted article captures and URLs, and avoid untrusted HTML/MHTML inputs until image URL validation is available.

Risk: Article images are uploaded to a configured Aliyun OSS bucket, so image content leaves the local environment.

Mitigation: Use a dedicated least-privilege OSS bucket and credentials limited to the needed object operations.

Risk: When LLM keyword extraction is configured, article text is sent to the configured LLM endpoint and the endpoint is not forced to HTTPS by the skill.

Mitigation: Configure LLM_BASE_URL with HTTPS only, or leave LLM settings unset to use local keyword extraction.

Risk: The release has a suspicious security verdict and should be reviewed before deployment.

Mitigation: Prefer the pinned requirements.txt installation path and review the configured cloud accounts and input sources before execution.

## Reference(s):

- [ClawHub Skill Page](https://clawhub.ai/ajayhao/skills/article-fetcher)
- [Project Homepage](https://github.com/AjayHao/article-fetcher)

## Skill Output:

**Output Type(s):** [text, markdown, code, shell commands, configuration]

**Output Format:** [Markdown files with YAML frontmatter, optional Notion pages, terminal status text, and Python dictionary results when used as a module]

**Output Parameters:** [1D]

**Other Properties Related to Output:** [Requires environment configuration for Aliyun OSS; Obsidian, Notion, cookies, and LLM keyword extraction are optional configuration-dependent outputs.]

## Skill Version(s):

1.3.6 (source: server release evidence, metadata, and changelog)

## Ethical Considerations:

Users should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.
Github ReposUpdated 15h agoRank 70

AionUi

Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!

MCPOPENCLAW
Github ReposUpdated 6mo agoRank 70

activepieces

AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents

OPENCLAW
Github ReposUpdated 6mo agoRank 70

cherry-studio

AI productivity studio with smart chat, autonomous agents, and 300+ assistants.

MCPOPENCLAW
Github ReposUpdated 7mo agoRank 70

CopilotKit

The Frontend for Agents & Generative UI. React + Angular

OPENCLAW

Machine-readable data

The same record, as JSON, for agents and crawlers.

{
  "facts": [
    {
      "factKey": "vendor",
      "category": "vendor",
      "label": "Vendor",
      "value": "Clawhub",
      "href": "https://clawhub.ai/ajayhao/skills/article-fetcher",
      "sourceUrl": "https://clawhub.ai/ajayhao/skills/article-fetcher",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-10T06:35:15.362Z",
      "isPublic": true
    },
    {
      "factKey": "protocols",
      "category": "compatibility",
      "label": "Protocol compatibility",
      "value": "OpenClaw",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-ajayhao-article-fetcher/contract",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-ajayhao-article-fetcher/contract",
      "sourceType": "contract",
      "confidence": "medium",
      "observedAt": "2026-10-10T06:35:15.362Z",
      "isPublic": true
    },
    {
      "factKey": "traction",
      "category": "adoption",
      "label": "Adoption signal",
      "value": "1.6K downloads",
      "href": "https://clawhub.ai/ajayhao/article-fetcher",
      "sourceUrl": "https://clawhub.ai/ajayhao/article-fetcher",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-10T06:35:15.362Z",
      "isPublic": true
    },
    {
      "factKey": "latest_release",
      "category": "release",
      "label": "Latest release",
      "value": "1.3.6",
      "href": "https://clawhub.ai/ajayhao/article-fetcher",
      "sourceUrl": "https://clawhub.ai/ajayhao/article-fetcher",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-08-08T13:16:14.670Z",
      "isPublic": true
    },
    {
      "factKey": "handshake_status",
      "category": "security",
      "label": "Handshake status",
      "value": "UNKNOWN",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-ajayhao-article-fetcher/trust",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-ajayhao-article-fetcher/trust",
      "sourceType": "trust",
      "confidence": "medium",
      "observedAt": null,
      "isPublic": true
    }
  ],
  "events": [
    {
      "eventType": "release",
      "title": "Release 1.3.6",
      "description": "## v1.3.6 — 安全修复(基于 ClawHub SkillSpector 审计) - P0(High): lxml 6.0.4→6.1.0 修复 CVE-2026-41066 XXE - P1(Medium): 校正 README/SKILL 误导性「零网络外发/完全本地」声明 - P2(Medium): 补强 LLM 文本外发声明(可选/来自配置/截断 12000 字符) - P3: 删除 references/tirith-blocking.md 绕过拦截笔记",
      "href": "https://clawhub.ai/ajayhao/article-fetcher",
      "sourceUrl": "https://clawhub.ai/ajayhao/article-fetcher",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-08-08T13:16:14.670Z",
      "isPublic": true
    }
  ]
}

Record generated Oct 10, 2026.

Sponsored

Ads related to Article Fetcher(文章抓取+Notion/Obsidian知识库存档) and adjacent AI workflows.