每日新闻搜索与智能摘要
Automatically scrape, process, and generate daily news digests from Chinese news sources. Covers industry dynamics, policy updates, economy, tech, energy, an... Skill: 每日新闻搜索与智能摘要 Owner: zigu-creator Summary: Automatically scrape, process, and generate daily news digests from Chinese news sources. Covers industry dynamics, policy updates, economy, tech, energy, an... Tags: latest:1.0.12 Version history: v1.0.12 | 2026-06-09T02:40:26.589Z | user v1.0.12 (2026-06-09) — Filtering Rules Expansion & Sync Fix Added 3 new exclusion categories to rules_config.py: cultural events/pro
Rank
62
Safety
84
Downloads
1.1k
Updated
Oct 11, 2026
Version
1.0.12
Source
CLAWHUB
About
What it does, and when to use it.
Capability contract not published. No trust telemetry is available yet. 1.1K downloads reported by the source. Last updated 10/11/2026.
Avoid when
- Contract metadata is missing or unavailable for deterministic execution.
Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing
Public facts
Every fact links back to the source it came from.
- Vendor
- Clawhubvendor · observed Oct 11, 2026
- Protocol compatibility
- OpenClawcompatibility · observed Oct 11, 2026
- Adoption signal
- 1.1K downloadsadoption · observed Oct 11, 2026
- Latest release
- 1.0.12release · observed Jun 9, 2026
- Handshake status
- UNKNOWNsecurity
Install and run
Setup complexity: low.
clawhub skill install s174gcx5nz7epm7zrz1qwamqt184w949:news-digest-v1- Setup complexity is classified as HIGH. You must provision dedicated cloud infrastructure or an isolated VM. Do not run this directly on your local workstation.
- Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.
Contract: missing
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-zigu-creator-news-digest-v1/snapshot"
Documentation
CLAWHUB
85,748 characters of source documentation, loaded on request.
Extracted files
3 files captured from the source.
SKILL.md
---
name: news-digest
description: "Automatically scrape, process, and generate daily news digests from Chinese news sources. Covers industry dynamics, policy updates, economy, tech, energy, and pricing information. Use when: user asks for daily news summary, news digest, 每日新闻摘要, 新闻汇总, 新闻摘要, or wants to set up automated news monitoring from Chinese news websites. Outputs formatted summaries with source attribution and original links."
version: 1.0.12
---
# News Digest - 每日新闻摘要
Automated pipeline for Chinese news aggregation and digest generation.
## Quick Start (3 步搞定)
```bash
# 第 1 步:安装依赖
pip install requests beautifulsoup4
# 第 2 步:一键初始化(建表 + 插入示例网站 + 关键词)
python scripts/news_digest_v2/init_db.py
# 第 3 步:运行摘要
python scripts/news_digest_v2/run_all_stages.py
```
或者一条命令全部搞定:
```bash
python scripts/news_digest_v2/quick_start.py
```
Output: `.news-digest-out.md` (workspace) + `新闻摘要_YYYYMMDD_HHMMSS.txt` (desktop)
## Architecture
```
Stage 1: Fetch → Scrape websites → Filter → Save to SQLite DB
Stage 2: Process → Deduplicate (≥90% similarity) → Tag keywords
Stage 2.5: LLM → Batch LLM summarization (optional, requires API key)
Stage 3: Output → Read LLM summaries (fallback to rule summaries) → Save to files
```
## Setup
### Prerequisites
- Python 3.8+ with: `requests`, `beautifulsoup4`
- SQLite (built-in)
### Initialize Database
Run the init script to create tables and seed with sample data:
```bash
python scripts/news_digest_v2/init_db.py
```
This creates:
- Database tables (articles, monitor_websites, system_keywords, digest_output)
- 10 sample news websites (People.cn, Xinhua, 36Kr, etc.)
- 18 sample keywords (产业, 政策, 经济, 科技, etc.)
Default database path: `news.db` (in the skill directory).
Override with environment variable: `NEWS_DIGEST_DB=/your/path/news.db`
### Customizing Your Sources
After initialization, add or remove websites and keywords via SQL:
```sql
-- Add a website
INSERT INTO monitor_websites (name, url, selector, category, priority)
VALUES ('示例网站', 'https://example.com', 'a', '财经', 1);
-- Add a keyword
INSERT INTO system_keywords (keyword, category, weight)
VALUES ('新能源', 'core', 5);
```
### Core Database Tables
| Table | Purpose |
|-------|---------|
| `articles` | Scraped news articles (title, content, URL, date, keywords, duplicate flag) |
| `monitor_websites` | Monitored websites (name, URL, CSS selector, category, enabled) |
| `system_keywords` | Keywords for relevance scoring (core vs auxiliary, with weight) |
| `digest_output` | LLM-generated summaries (optional) |
## Usage
### Full Pipeline
```bash
python scripts/news_digest_v2/run_all_stages.py
```
Takes ~13 minutes (network + LLM bound).
### One-Command Quick Start
```bash
python scripts/news_digest_v2/quick_start.py
```
Runs init + fetch + process + output in one shot.
### Cron Job Example
```yaml
schedule: "0 20 * * *" # Daily 20:00
payload:
run: python scripts/news_digest_v2/run_all_stages.py
then: read .news-d_meta.json
{
"ownerId": "kn7983et5m2qgha7y8gvb8dgsn84w87m",
"slug": "news-digest-v1",
"version": "1.0.12",
"publishedAt": 1780972826589
}skill-card.md
## Description: Automatically scrapes Chinese news sources, filters and deduplicates articles, and generates daily summaries with source attribution and original links. This skill is ready for commercial/non-commercial use. ## Publisher: [zigu-creator](https://clawhub.ai/user/zigu-creator) ### License/Terms of Use: MIT-0 ## Use Case: Developers and operators use this skill to set up a local Chinese news monitoring workflow that fetches news, filters low-relevance content, optionally summarizes with an LLM, and produces a daily digest for review or distribution. ### Deployment Geography for Use: Global ## Known Risks and Mitigations: Risk: The skill can make broad network requests to configured news sites and custom article links. Mitigation: Run it only with reviewed monitor_websites entries and trusted article sources before using the full pipeline. Risk: The LLM stage can reuse an OpenClaw API key from local configuration when dedicated NEWS_DIGEST_LLM credentials are not set. Mitigation: Set dedicated, limited NEWS_DIGEST_LLM_API_KEY and NEWS_DIGEST_LLM_BASE_URL values, or remove the OpenClaw config fallback before execution. Risk: The skill writes a local SQLite database plus digest files in the workspace and Desktop. Mitigation: Run it in a workspace where these writes are expected, and review generated digest content before sharing it. ## Reference(s): - [ClawHub skill page](https://clawhub.ai/zigu-creator/skills/news-digest-v1) - [Publisher profile](https://clawhub.ai/user/zigu-creator) ## Skill Output: **Output Type(s):** [text, markdown, shell commands, configuration, guidance] **Output Format:** [Markdown and plain text digest files with source titles, summaries, dates, and original links; setup guidance includes Markdown with bash and SQL snippets.] **Output Parameters:** [1D] **Other Properties Related to Output:** [Writes .news-digest-out.md in the workspace and a timestamped Chinese text digest on Desktop; optional LLM summarization uses configured API credentials.] ## Skill Version(s): 1.0.12 (source: SKILL.md frontmatter and ClawHub release evidence, released 2026-06-09) ## Ethical Considerations: Users should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.
AionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
activepieces
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
CopilotKit
The Frontend for Agents & Generative UI. React + Angular
Machine-readable data
The same record, as JSON, for agents and crawlers.
{
"facts": [
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Clawhub",
"href": "https://clawhub.ai/zigu-creator/skills/news-digest-v1",
"sourceUrl": "https://clawhub.ai/zigu-creator/skills/news-digest-v1",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-11T06:42:50.372Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-zigu-creator-news-digest-v1/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-zigu-creator-news-digest-v1/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-11T06:42:50.372Z",
"isPublic": true
},
{
"factKey": "traction",
"category": "adoption",
"label": "Adoption signal",
"value": "1.1K downloads",
"href": "https://clawhub.ai/zigu-creator/news-digest-v1",
"sourceUrl": "https://clawhub.ai/zigu-creator/news-digest-v1",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-11T06:42:50.372Z",
"isPublic": true
},
{
"factKey": "latest_release",
"category": "release",
"label": "Latest release",
"value": "1.0.12",
"href": "https://clawhub.ai/zigu-creator/news-digest-v1",
"sourceUrl": "https://clawhub.ai/zigu-creator/news-digest-v1",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-06-09T02:40:26.589Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-zigu-creator-news-digest-v1/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-zigu-creator-news-digest-v1/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
],
"events": [
{
"eventType": "release",
"title": "Release 1.0.12",
"description": "v1.0.12 (2026-06-09) — Filtering Rules Expansion & Sync Fix Added 3 new exclusion categories to rules_config.py: cultural events/propaganda sessions (诵读会, 宣讲会, scientist spirit readings), credit-knowledge Q&A/encyclopedia content, and news briefings/morning digest formats (e.g. \"8点1氪\") — 8 new keywords total Fixed sync_to_skill.py: restored cross_day_dedup.py to the sync file list (was accidentally removed by cleanup logic) Manually cleaned residual duplicate entries in digest_output table v1.0.11 (2026-06-08) — People's Daily Primary Site & Cross-Day Dedup Improvement Added People's Daily primary site (paper.people.com.cn/rmrb/) as a new monitored source (id=47, priority=1) Enhanced cross_day_dedup.py with Hard Rule 4: mutual-title-inclusion detection (catches same-event reports from different outlets with completely different titles) Tuned Jaccard weight from 0.5 → 0.6 for better title-similarity scoring Restored header format: \"Sources: X | Articles: Y items\" in formatter.py Simplified SQL queries and optimized logging in stage2_5_llm_summary.py v1.0.10 (2026-06-05) — Header Count Fix & 5 New Filter Categories Fixed formatter.py article count in header: changed from pre-filter count (filtered_news length) to actual output count via placeholder + post-replacement Added title truncation protection: titles ending with commas/enumeration commas now get ellipsis appended to avoid misleading DB-truncated titles Added 5 new exclusion categories: Party-building/historical commemoration articles (17 keywords), solar-term/astronomical science popularization, official dismissal/prosecution notices (12 keywords, also added to corporate_scandal) v1.0.8 (2026-06-05) — Cross-Day Dedup & GBK Decoding New cross_day_dedup.py: compares today's candidates against the last 3 days of historical digests using a weighted scoring model — Title Jaccard(0.5) + Number match(0.25) + Content-word overlap(0.25) Hard rule: articles sharing no significant numbers → similarity forced to 0 (auto-passes recurring reports like PMI/CPI without whitelist) Three-tier verdict: ≥0.75 block / 0.60–0.75 warn-keep / <0.60 pass Title rewrite rule 9: exhibition/attendance titles → event subject focus; IPO/financing/earnings → retain company name Anti-hallucination rule 10: summaries must derive exclusively from source text, no external knowledge injection stage1_fetch.py: source count changed from hardcoded 42 to dynamic len(WEBSITES) GBK decoding: decode_response() merged into main fetch flow, supports all known GBK-encoded sources v1.0.7 (2026-06-01) — GBK Encoding Fix & Content-Type Blacklist fetcher.py: new decode_response() function forces GBK decoding for known GBK sources (People's Daily Overseas Edition), fixing Cyrillic mojibake at the root stage2_5_llm_summary.py: mojibake title detection — prompts LLM to generate accurate title from article body when garbled text is detected formatter.py: fallback filter to discard garbled titles New TITLE_EXCLUDE_KEYWORDS: excludes opinion pieces (评论丨, 时评, 社评), investigative reports (深度观察, 记者观察), HR/recruitment (招聘, 面试, 人事任免), obituaries (讣告), exclusive interviews (专访) New URL_EXCLUDE_PATTERNS: blocks People's Daily opinion channel and newspaper commentary pages Timeout recommendation: 900s → 1200s (实测 total runtime ~1150s for 50-article LLM batch + fetch)",
"href": "https://clawhub.ai/zigu-creator/news-digest-v1",
"sourceUrl": "https://clawhub.ai/zigu-creator/news-digest-v1",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-06-09T02:40:26.589Z",
"isPublic": true
}
]
}Record generated Oct 11, 2026.
