{"id":"d98cb929-3026-49d1-a2a8-632af49d4dbe","entityType":"agent","slug":"clawhub-zigu-creator-news-digest-v1","name":"每日新闻搜索与智能摘要","canonicalUrl":"https://www.xpersona.co/agent/clawhub-zigu-creator-news-digest-v1","canonicalPath":"/agent/clawhub-zigu-creator-news-digest-v1","generatedAt":"2026-10-11T10:51:53.537Z","source":"CLAWHUB","claimStatus":"UNCLAIMED","verificationTier":"NONE","summary":{"evidence":{"source":"editorial-content","verified":true,"confidence":"high","updatedAt":"2026-10-11T06:42:50.372Z","emptyReason":null},"description":"Automatically scrape, process, and generate daily news digests from Chinese news sources. Covers industry dynamics, policy updates, economy, tech, energy, an... Skill: 每日新闻搜索与智能摘要 Owner: zigu-creator Summary: Automatically scrape, process, and generate daily news digests from Chinese news sources. Covers industry dynamics, policy updates, economy, tech, energy, an... Tags: latest:1.0.12 Version history: v1.0.12 | 2026-06-09T02:40:26.589Z | user v1.0.12 (2026-06-09) — Filtering Rules Expansion & Sync Fix Added 3 new exclusion categories to rules_config.py: cultural events/pro","descriptionLabel":"Technical summary","evidenceSummary":"Capability contract not published. No trust telemetry is available yet. 1.1K downloads reported by the source. Last updated 10/11/2026.","installCommand":"clawhub skill install s174gcx5nz7epm7zrz1qwamqt184w949:news-digest-v1","sourceUrl":"https://clawhub.ai/zigu-creator/news-digest-v1","homepage":"https://clawhub.ai/zigu-creator/skills/news-digest-v1","primaryLinks":[{"label":"View on ClawHub","url":"https://clawhub.ai/zigu-creator/news-digest-v1","kind":"source"},{"label":"Homepage","url":"https://clawhub.ai/zigu-creator/skills/news-digest-v1","kind":"homepage"}],"safetyScore":84,"overallRank":62,"popularityScore":61,"trustScore":null,"claimedByName":null,"isOwner":false,"seoDescription":"Automatically scrape, process, and generate daily news digests from Chinese news sources. Covers industry dynamics, policy updates, economy, tech, energy, an..."},"coverage":{"evidence":{"source":"public-profile","verified":false,"confidence":"medium","updatedAt":"2026-10-11T06:42:50.372Z","emptyReason":null},"protocols":[{"protocol":"OPENCLEW","label":"OpenClaw","status":"self-declared","notes":"Declared in the public agent profile."}],"capabilities":[],"verifiedCount":0,"selfDeclaredCount":1,"capabilityMatrix":{"rows":[{"key":"OPENCLEW","type":"protocol","support":"unknown","confidenceSource":"profile","notes":"Listed on profile"}],"flattenedTokens":"protocol:OPENCLEW|unknown|profile"}},"adoption":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-11T06:42:50.372Z","emptyReason":null},"stars":null,"forks":null,"downloads":1131,"packageName":null,"latestVersion":"1.0.12","tractionLabel":"1.1K downloads"},"release":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-11T06:42:50.314Z","emptyReason":null},"lastUpdatedAt":"2026-10-11T06:42:50.372Z","lastCrawledAt":"2026-10-11T06:42:50.314Z","lastIndexedAt":null,"nextCrawlAt":"2026-10-12T06:42:50.314Z","lastVerifiedAt":null,"highlights":[{"version":"1.0.12","createdAt":"2026-06-09T02:40:26.589Z","changelog":"v1.0.12 (2026-06-09) — Filtering Rules Expansion & Sync Fix Added 3 new exclusion categories to rules_config.py: cultural events/propaganda sessions (诵读会, 宣讲会, scientist spirit readings), credit-knowledge Q&A/encyclopedia content, and news briefings/morning digest formats (e.g. \"8点1氪\") — 8 new keywords total Fixed sync_to_skill.py: restored cross_day_dedup.py to the sync file list (was accidentally removed by cleanup logic) Manually cleaned residual duplicate entries in digest_output table v1.0.11 (2026-06-08) — People's Daily Primary Site & Cross-Day Dedup Improvement Added People's Daily primary site (paper.people.com.cn/rmrb/) as a new monitored source (id=47, priority=1) Enhanced cross_day_dedup.py with Hard Rule 4: mutual-title-inclusion detection (catches same-event reports from different outlets with completely different titles) Tuned Jaccard weight from 0.5 → 0.6 for better title-similarity scoring Restored header format: \"Sources: X | Articles: Y items\" in formatter.py Simplified SQL queries and optimized logging in stage2_5_llm_summary.py v1.0.10 (2026-06-05) — Header Count Fix & 5 New Filter Categories Fixed formatter.py article count in header: changed from pre-filter count (filtered_news length) to actual output count via placeholder + post-replacement Added title truncation protection: titles ending with commas/enumeration commas now get ellipsis appended to avoid misleading DB-truncated titles Added 5 new exclusion categories: Party-building/historical commemoration articles (17 keywords), solar-term/astronomical science popularization, official dismissal/prosecution notices (12 keywords, also added to corporate_scandal) v1.0.8 (2026-06-05) — Cross-Day Dedup & GBK Decoding New cross_day_dedup.py: compares today's candidates against the last 3 days of historical digests using a weighted scoring model — Title Jaccard(0.5) + Number match(0.25) + Content-word overlap(0.25) Hard rule: articles sharing no significant numbers → similarity forced to 0 (auto-passes recurring reports like PMI/CPI without whitelist) Three-tier verdict: ≥0.75 block / 0.60–0.75 warn-keep / <0.60 pass Title rewrite rule 9: exhibition/attendance titles → event subject focus; IPO/financing/earnings → retain company name Anti-hallucination rule 10: summaries must derive exclusively from source text, no external knowledge injection stage1_fetch.py: source count changed from hardcoded 42 to dynamic len(WEBSITES) GBK decoding: decode_response() merged into main fetch flow, supports all known GBK-encoded sources v1.0.7 (2026-06-01) — GBK Encoding Fix & Content-Type Blacklist fetcher.py: new decode_response() function forces GBK decoding for known GBK sources (People's Daily Overseas Edition), fixing Cyrillic mojibake at the root stage2_5_llm_summary.py: mojibake title detection — prompts LLM to generate accurate title from article body when garbled text is detected formatter.py: fallback filter to discard garbled titles New TITLE_EXCLUDE_KEYWORDS: excludes opinion pieces (评论丨, 时评, 社评), investigative reports (深度观察, 记者观察), HR/recruitment (招聘, 面试, 人事任免), obituaries (讣告), exclusive interviews (专访) New URL_EXCLUDE_PATTERNS: blocks People's Daily opinion channel and newspaper commentary pages Timeout recommendation: 900s → 1200s (实测 total runtime ~1150s for 50-article LLM batch + fetch)","fileCount":19,"zipByteSize":61343},{"version":"1.0.6","createdAt":"2026-06-05T10:34:29.678Z","changelog":"**v1.0.6 Changelog** - Added cross-day deduplication support with `scripts/news_digest_v2/cross_day_dedup.py` to detect and handle duplicate news across multiple days. - Removed the file `skill-card.md` from the project. - No user-facing feature changes beyond improved deduplication and project cleanup.","fileCount":19,"zipByteSize":59788},{"version":"1.0.5","createdAt":"2026-06-02T14:39:05.375Z","changelog":"news-digest v1.0.5 New Headline rewriting rules (LLM stage): When an original headline is led by a specific company/enterprise as the subject and describes participation in an exhibition, event, or activity (e.g., \"XX Company debuts at XX Expo\"), the headline will be automatically rewritten to center on the event, activity, or industry trend, with the company name retained in the summary body. Headline rewriting exemptions: Headlines will remain unchanged in scenarios where the company itself is the core subject, including IPO, listing, financing, acquisition, merger, financial report release, major contract signing, and core product launch. Fixed Authoritative priority strategy for LLM article selection: The guaranteed minimum 2-item strategy is also applied to batch summarized selections in Phase 2.5 LLM, ensuring at least 2 articles are selected from each authoritative source. config.py: Switched to use the NEWS_DIGEST_DB environment variable to eliminate hard-coded paths. Cleaned up Removed redundant documents under skills/news-digest/ (CURRENT_SETUP.md, README.md, README_STAGES.md, TASK_GUIDE.md, main.py, update_db_schema.py), reducing over 1,300 lines of duplicate files. Changed Added 12 system keywords to rule files. MAX_OUTPUT_COUNT changed from 35 to 50. Added 3 new monitoring sources. Other minor improvements.","fileCount":18,"zipByteSize":53818},{"version":"1.0.3","createdAt":"2026-05-22T08:59:38.486Z","changelog":"### v1.0.3 (2026-05-22) - Added an independent parser for CNR.cn to separately process <strong> tags as headlines, resolving the issue of content filtering caused by mixed headlines and body text. - Implemented a guaranteed authoritative source strategy for both Phase 2.5 (LLM) and Phase 3 outputs, with a minimum of 2 news items selected from each authoritative source. - Unified sub-channels of Xinhuanet.com under the single source label \"Xinhuanet\". - Enhanced social news filtering by adding new filtering keywords related to traffic violations and penalties, including fines, administrative detention, traffic police, etc. - Optimized web crawling for websites using the GB2312 encoding (e.g., CNR.cn) to improve the success rate of encoding detection.","fileCount":18,"zipByteSize":48022},{"version":"1.0.2","createdAt":"2026-05-08T10:33:59.172Z","changelog":"## v1.0.2 (2026-05-08) - LLM summary limit increased to 300 characters (was 200), with more comprehensive content - New requirement: full names of referenced policies, plans, and documents must be preserved - Prompt optimization: emphasize retaining key data and core facts","fileCount":17,"zipByteSize":42091},{"version":"1.0.1","createdAt":"2026-05-06T13:46:26.266Z","changelog":"**v1.0.1 (2026-05-06)** - Added `init_db.py` script for one-click database initialization (creates tables and seeds sample data) - Added `quick_start.py` script for one-command full pipeline setup and execution - Updated documentation for a simplified 3-step installation and quick start process - Included FAQ section and improved onboarding guidance - Now ships with 10 sample Chinese news sources and 18 sample keywords by default","fileCount":17,"zipByteSize":42217},{"version":"1.0.0","createdAt":"2026-05-05T06:47:45.258Z","changelog":"## v1.0.0 - 初始发布 自动抓取自定义来源的新闻网站，生成每日新闻摘要。覆盖指定领域。 ### 安装后配置（必须完成） **1. 安装依赖** ```bash pip install requests beautifulsoup4 2. 初始化数据库 bash Copy # 创建数据库和表结构 python -c \" import sqlite3, os db = sqlite3.connect('news.db') db.execute('''CREATE TABLE IF NOT EXISTS articles ( id INTEGER PRIMARY KEY AUTOINCREMENT, title TEXT NOT NULL, source TEXT NOT NULL, publish_date TEXT NOT NULL, summary TEXT, content TEXT, url TEXT UNIQUE NOT NULL, keywords TEXT, is_duplicate INTEGER DEFAULT 0, similarity_score REAL DEFAULT 0, created_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP )''') db.execute('''CREATE TABLE IF NOT EXISTS monitor_websites ( id INTEGER PRIMARY KEY AUTOINCREMENT, name TEXT UNIQUE NOT NULL, url TEXT NOT NULL, selector TEXT DEFAULT 'a', category TEXT, priority INTEGER DEFAULT 3, enabled INTEGER DEFAULT 1 )''') db.execute('''CREATE TABLE IF NOT EXISTS system_keywords ( id INTEGER PRIMARY KEY AUTOINCREMENT, keyword TEXT UNIQUE NOT NULL, category TEXT, weight INTEGER DEFAULT 1, enabled INTEGER DEFAULT 1 )''') db.execute('CREATE INDEX IF NOT EXISTS idx_publish_date ON articles(publish_date)') db.execute('CREATE INDEX IF NOT EXISTS idx_keywords ON articles(keywords)') db.commit() db.close() print('OK: news.db created') \" 3. 添加监测网站（示例） bash Copy python -c \" import sqlite3 db = sqlite3.connect('news.db') sites = [ ('中国经济网', 'http://www.ce.cn/', 'a', '财经', 1), # 添加你想监控的网站... ] for name, url, sel, cat, pri in sites: db.execute('INSERT OR IGNORE INTO monitor_websites (name, url, selector, category, priority) VALUES (?,?,?,?,?)', (name, url, sel, cat, pri)) db.commit() db.close() print(f'OK: added {len(sites)} websites') \" 4. 添加关键词（示例） bash Copy python -c \" import sqlite3 db = sqlite3.connect('news.db') kws = [ ('市场', 'auxiliary', 2), ('企业', 'auxiliary', 2), ] for kw, cat, w in kws: db.execute('INSERT OR IGNORE INTO system_keywords (keyword, category, weight) VALUES (?,?,?)', (kw, cat, w)) db.commit() db.close() print(f'OK: added {len(kws)} keywords') \" 5. 运行 bash Copy python scripts/news_digest_v2/run_all_stages.py 可选配置 环境变量 默认值 说明 NEWS_DIGEST_DB news.db 数据库路径 NEWS_DIGEST_LLM_API_KEY (空) LLM API 密钥（可选，启用智能总结） NEWS_DIGEST_LLM_BASE_URL (空) LLM API 地址 NEWS_DIGEST_LLM_MODEL qwen-plus LLM 模型名 功能特性 • 🕷️ 网站智能抓取（支持站点独立解析器） • 🧹 6 大类无效内容过滤（娱乐/社会/科普/教程等） • 🔄 相似度 ≥90% 自动去重 • ✍️ 智能摘要提取（非简单截断，基于信息密度评分） • 🤖 可选 LLM 批量总结（生成更专业的摘要） • 📄 紧凑格式输出（来源+标题+摘要+链接）","fileCount":15,"zipByteSize":38224}]},"execution":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No published capability contract is available yet."},"installCommand":"clawhub skill install s174gcx5nz7epm7zrz1qwamqt184w949:news-digest-v1","setupComplexity":"low","setupSteps":["Setup complexity is classified as HIGH. You must provision dedicated cloud infrastructure or an isolated VM. Do not run this directly on your local workstation.","Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data."],"contract":{"contractStatus":"missing","authModes":[],"requires":[],"forbidden":[],"supportsMcp":false,"supportsA2a":false,"supportsStreaming":false,"inputSchemaRef":null,"outputSchemaRef":null,"dataRegion":null,"contractUpdatedAt":null,"sourceUpdatedAt":null,"freshnessSeconds":null},"invocationGuide":{"preferredApi":{"snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zigu-creator-news-digest-v1/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zigu-creator-news-digest-v1/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zigu-creator-news-digest-v1/trust"},"curlExamples":["curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-zigu-creator-news-digest-v1/snapshot\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-zigu-creator-news-digest-v1/contract\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-zigu-creator-news-digest-v1/trust\""],"jsonRequestTemplate":{"query":"summarize this repo","constraints":{"maxLatencyMs":2000,"protocolPreference":["OPENCLEW"]}},"jsonResponseTemplate":{"ok":true,"result":{"summary":"...","confidence":0.9},"meta":{"source":"CLAWHUB","generatedAt":"2026-10-11T10:51:53.537Z"}},"retryPolicy":{"maxAttempts":3,"backoffMs":[500,1500,3500],"retryableConditions":["HTTP_429","HTTP_503","NETWORK_TIMEOUT"]}},"endpoints":{"dossierUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zigu-creator-news-digest-v1/dossier","snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zigu-creator-news-digest-v1/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zigu-creator-news-digest-v1/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zigu-creator-news-digest-v1/trust"}},"reliability":{"evidence":{"source":"runtime-metrics","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No trust, reliability, or runtime telemetry is available."},"trust":{"status":"unavailable","handshakeStatus":"UNKNOWN","verificationFreshnessHours":null,"reputationScore":null,"p95LatencyMs":null,"successRate30d":null,"fallbackRate":null,"attempts30d":null,"trustUpdatedAt":null,"trustConfidence":"unknown","sourceUpdatedAt":null,"freshnessSeconds":null},"decisionGuardrails":{"doNotUseIf":["Contract metadata is missing or unavailable for deterministic execution."],"safeUseWhen":[],"riskFlags":["missing_or_unavailable_contract","trust_data_unavailable","schema_references_missing"],"operationalConfidence":"low"},"executionMetrics":{"observedLatencyMsP50":null,"observedLatencyMsP95":null,"estimatedCostUsd":null,"uptime30d":null,"rateLimitRpm":null,"rateLimitBurst":null,"lastVerifiedAt":null,"verificationSource":null},"runtimeMetrics":{"successRate":null,"avgLatencyMs":null,"avgCostUsd":null,"hallucinationRate":null,"retryRate":null,"disputeRate":null,"p50Latency":null,"p95Latency":null,"lastUpdated":null}},"benchmarks":{"evidence":{"source":"no-benchmark-data","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No benchmark suites or observed failure patterns are available."},"suites":[],"failurePatterns":[]},"artifacts":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"high","updatedAt":"2026-10-11T06:42:50.372Z","emptyReason":null},"readme":"Skill: 每日新闻搜索与智能摘要\n\nOwner: zigu-creator\n\nSummary: Automatically scrape, process, and generate daily news digests from Chinese news sources. Covers industry dynamics, policy updates, economy, tech, energy, an...\n\nTags: latest:1.0.12\n\nVersion history:\n\nv1.0.12 | 2026-06-09T02:40:26.589Z | user\n\nv1.0.12 (2026-06-09) — Filtering Rules Expansion & Sync Fix\n\nAdded 3 new exclusion categories to rules_config.py: cultural events/propaganda sessions (诵读会, 宣讲会, scientist spirit readings), credit-knowledge Q&A/encyclopedia content, and news briefings/morning digest formats (e.g. \"8点1氪\") — 8 new keywords total\nFixed sync_to_skill.py: restored cross_day_dedup.py to the sync file list (was accidentally removed by cleanup logic)\nManually cleaned residual duplicate entries in digest_output table\nv1.0.11 (2026-06-08) — People's Daily Primary Site & Cross-Day Dedup Improvement\n\nAdded People's Daily primary site (paper.people.com.cn/rmrb/) as a new monitored source (id=47, priority=1)\nEnhanced cross_day_dedup.py with Hard Rule 4: mutual-title-inclusion detection (catches same-event reports from different outlets with completely different titles)\nTuned Jaccard weight from 0.5 → 0.6 for better title-similarity scoring\nRestored header format: \"Sources: X | Articles: Y items\" in formatter.py\nSimplified SQL queries and optimized logging in stage2_5_llm_summary.py\nv1.0.10 (2026-06-05) — Header Count Fix & 5 New Filter Categories\n\nFixed formatter.py article count in header: changed from pre-filter count (filtered_news length) to actual output count via placeholder + post-replacement\nAdded title truncation protection: titles ending with commas/enumeration commas now get ellipsis appended to avoid misleading DB-truncated titles\nAdded 5 new exclusion categories: Party-building/historical commemoration articles (17 keywords), solar-term/astronomical science popularization, official dismissal/prosecution notices (12 keywords, also added to corporate_scandal)\nv1.0.8 (2026-06-05) — Cross-Day Dedup & GBK Decoding\n\nNew cross_day_dedup.py: compares today's candidates against the last 3 days of historical digests using a weighted scoring model — Title Jaccard(0.5) + Number match(0.25) + Content-word overlap(0.25)\nHard rule: articles sharing no significant numbers → similarity forced to 0 (auto-passes recurring reports like PMI/CPI without whitelist)\nThree-tier verdict: ≥0.75 block / 0.60–0.75 warn-keep / <0.60 pass\nTitle rewrite rule 9: exhibition/attendance titles → event subject focus; IPO/financing/earnings → retain company name\nAnti-hallucination rule 10: summaries must derive exclusively from source text, no external knowledge injection\nstage1_fetch.py: source count changed from hardcoded 42 to dynamic len(WEBSITES)\nGBK decoding: decode_response() merged into main fetch flow, supports all known GBK-encoded sources\nv1.0.7 (2026-06-01) — GBK Encoding Fix & Content-Type Blacklist\n\nfetcher.py: new decode_response() function forces GBK decoding for known GBK sources (People's Daily Overseas Edition), fixing Cyrillic mojibake at the root\nstage2_5_llm_summary.py: mojibake title detection — prompts LLM to generate accurate title from article body when garbled text is detected\nformatter.py: fallback filter to discard garbled titles\nNew TITLE_EXCLUDE_KEYWORDS: excludes opinion pieces (评论丨, 时评, 社评), investigative reports (深度观察, 记者观察), HR/recruitment (招聘, 面试, 人事任免), obituaries (讣告), exclusive interviews (专访)\nNew URL_EXCLUDE_PATTERNS: blocks People's Daily opinion channel and newspaper commentary pages\nTimeout recommendation: 900s → 1200s (实测 total runtime ~1150s for 50-article LLM batch + fetch)\n\nv1.0.6 | 2026-06-05T10:34:29.678Z | user\n\n**v1.0.6 Changelog**\n\n- Added cross-day deduplication support with `scripts/news_digest_v2/cross_day_dedup.py` to detect and handle duplicate news across multiple days.\n- Removed the file `skill-card.md` from the project.\n- No user-facing feature changes beyond improved deduplication and project cleanup.\n\nv1.0.5 | 2026-06-02T14:39:05.375Z | user\n\nnews-digest v1.0.5\n\nNew\nHeadline rewriting rules (LLM stage): When an original headline is led by a specific company/enterprise as the subject and describes participation in an exhibition, event, or activity (e.g., \"XX Company debuts at XX Expo\"), the headline will be automatically rewritten to center on the event, activity, or industry trend, with the company name retained in the summary body.\nHeadline rewriting exemptions: Headlines will remain unchanged in scenarios where the company itself is the core subject, including IPO, listing, financing, acquisition, merger, financial report release, major contract signing, and core product launch.\nFixed\nAuthoritative priority strategy for LLM article selection: The guaranteed minimum 2-item strategy is also applied to batch summarized selections in Phase 2.5 LLM, ensuring at least 2 articles are selected from each authoritative source.\nconfig.py: Switched to use the NEWS_DIGEST_DB environment variable to eliminate hard-coded paths.\nCleaned up\nRemoved redundant documents under skills/news-digest/ (CURRENT_SETUP.md, README.md, README_STAGES.md, TASK_GUIDE.md, main.py, update_db_schema.py), reducing over 1,300 lines of duplicate files.\nChanged\nAdded 12 system keywords to rule files.\nMAX_OUTPUT_COUNT changed from 35 to 50.\nAdded 3 new monitoring sources.\nOther minor improvements.\n\nv1.0.3 | 2026-05-22T08:59:38.486Z | user\n\n### v1.0.3 (2026-05-22)\n- Added an independent parser for CNR.cn to separately process <strong> tags as headlines, resolving the issue of content filtering caused by mixed headlines and body text.\n- Implemented a guaranteed authoritative source strategy for both Phase 2.5 (LLM) and Phase 3 outputs, with a minimum of 2 news items selected from each authoritative source.\n- Unified sub-channels of Xinhuanet.com under the single source label \"Xinhuanet\".\n- Enhanced social news filtering by adding new filtering keywords related to traffic violations and penalties, including fines, administrative detention, traffic police, etc.\n- Optimized web crawling for websites using the GB2312 encoding (e.g., CNR.cn) to improve the success rate of encoding detection.\n\nv1.0.2 | 2026-05-08T10:33:59.172Z | user\n\n## v1.0.2 (2026-05-08)\n- LLM summary limit increased to 300 characters (was 200), with more comprehensive content\n- New requirement: full names of referenced policies, plans, and documents must be preserved\n- Prompt optimization: emphasize retaining key data and core facts\n\nv1.0.1 | 2026-05-06T13:46:26.266Z | user\n\n**v1.0.1 (2026-05-06)**\n- Added `init_db.py` script for one-click database initialization (creates tables and seeds sample data)\n- Added `quick_start.py` script for one-command full pipeline setup and execution\n- Updated documentation for a simplified 3-step installation and quick start process\n- Included FAQ section and improved onboarding guidance\n- Now ships with 10 sample Chinese news sources and 18 sample keywords by default\n\nv1.0.0 | 2026-05-05T06:47:45.258Z | user\n\n## v1.0.0 - 初始发布\n自动抓取自定义来源的新闻网站，生成每日新闻摘要。覆盖指定领域。\n### 安装后配置（必须完成）\n\n**1. 安装依赖**\n```bash\npip install requests beautifulsoup4\n2. 初始化数据库\nbash\nCopy\n# 创建数据库和表结构\npython -c \"\nimport sqlite3, os\ndb = sqlite3.connect('news.db')\ndb.execute('''CREATE TABLE IF NOT EXISTS articles (\n    id INTEGER PRIMARY KEY AUTOINCREMENT,\n    title TEXT NOT NULL, source TEXT NOT NULL,\n    publish_date TEXT NOT NULL, summary TEXT,\n    content TEXT, url TEXT UNIQUE NOT NULL,\n    keywords TEXT, is_duplicate INTEGER DEFAULT 0,\n    similarity_score REAL DEFAULT 0,\n    created_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP\n)''')\ndb.execute('''CREATE TABLE IF NOT EXISTS monitor_websites (\n    id INTEGER PRIMARY KEY AUTOINCREMENT,\n    name TEXT UNIQUE NOT NULL, url TEXT NOT NULL,\n    selector TEXT DEFAULT 'a', category TEXT,\n    priority INTEGER DEFAULT 3, enabled INTEGER DEFAULT 1\n)''')\ndb.execute('''CREATE TABLE IF NOT EXISTS system_keywords (\n    id INTEGER PRIMARY KEY AUTOINCREMENT,\n    keyword TEXT UNIQUE NOT NULL,\n    category TEXT, weight INTEGER DEFAULT 1,\n    enabled INTEGER DEFAULT 1\n)''')\ndb.execute('CREATE INDEX IF NOT EXISTS idx_publish_date ON articles(publish_date)')\ndb.execute('CREATE INDEX IF NOT EXISTS idx_keywords ON articles(keywords)')\ndb.commit()\ndb.close()\nprint('OK: news.db created')\n\"\n3. 添加监测网站（示例）\nbash\nCopy\npython -c \"\nimport sqlite3\ndb = sqlite3.connect('news.db')\nsites = [\n    ('中国经济网', 'http://www.ce.cn/', 'a', '财经', 1),\n    # 添加你想监控的网站...\n]\nfor name, url, sel, cat, pri in sites:\n    db.execute('INSERT OR IGNORE INTO monitor_websites (name, url, selector, category, priority) VALUES (?,?,?,?,?)',\n               (name, url, sel, cat, pri))\ndb.commit()\ndb.close()\nprint(f'OK: added {len(sites)} websites')\n\"\n4. 添加关键词（示例）\nbash\nCopy\npython -c \"\nimport sqlite3\ndb = sqlite3.connect('news.db')\nkws = [\n    ('市场', 'auxiliary', 2), ('企业', 'auxiliary', 2),\n]\nfor kw, cat, w in kws:\n    db.execute('INSERT OR IGNORE INTO system_keywords (keyword, category, weight) VALUES (?,?,?)',\n               (kw, cat, w))\ndb.commit()\ndb.close()\nprint(f'OK: added {len(kws)} keywords')\n\"\n5. 运行\nbash\nCopy\npython scripts/news_digest_v2/run_all_stages.py\n可选配置\n环境变量\n默认值\n说明\nNEWS_DIGEST_DB\nnews.db\n数据库路径\nNEWS_DIGEST_LLM_API_KEY\n(空)\nLLM API 密钥（可选，启用智能总结）\nNEWS_DIGEST_LLM_BASE_URL\n(空)\nLLM API 地址\nNEWS_DIGEST_LLM_MODEL\nqwen-plus\nLLM 模型名\n功能特性\n• 🕷️ 网站智能抓取（支持站点独立解析器）\n• 🧹 6 大类无效内容过滤（娱乐/社会/科普/教程等）\n• 🔄 相似度 ≥90% 自动去重\n• ✍️ 智能摘要提取（非简单截断，基于信息密度评分）\n• 🤖 可选 LLM 批量总结（生成更专业的摘要）\n• 📄 紧凑格式输出（来源+标题+摘要+链接）\n\nArchive index:\n\nArchive v1.0.12: 19 files, 61343 bytes\n\nFiles: scripts/news_digest_v2/__init__.py (310b), scripts/news_digest_v2/config.py (6769b), scripts/news_digest_v2/cross_day_dedup.py (9285b), scripts/news_digest_v2/database.py (13334b), scripts/news_digest_v2/fetcher.py (33837b), scripts/news_digest_v2/filters.py (3258b), scripts/news_digest_v2/formatter.py (6815b), scripts/news_digest_v2/init_db.py (6849b), scripts/news_digest_v2/quick_start.py (1966b), scripts/news_digest_v2/rules_config.py (25517b), scripts/news_digest_v2/run_all_stages.py (3685b), scripts/news_digest_v2/similarity.py (2348b), scripts/news_digest_v2/stage1_fetch.py (1808b), scripts/news_digest_v2/stage2_5_llm_summary.py (14888b), scripts/news_digest_v2/stage2_process.py (2431b), scripts/news_digest_v2/stage3_output.py (7967b), skill-card.md (2417b), SKILL.md (15145b), _meta.json (134b)\n\nFile v1.0.12:SKILL.md\n\n---\nname: news-digest\ndescription: \"Automatically scrape, process, and generate daily news digests from Chinese news sources. Covers industry dynamics, policy updates, economy, tech, energy, and pricing information. Use when: user asks for daily news summary, news digest, 每日新闻摘要, 新闻汇总, 新闻摘要, or wants to set up automated news monitoring from Chinese news websites. Outputs formatted summaries with source attribution and original links.\"\nversion: 1.0.12\n---\n\n# News Digest - 每日新闻摘要\n\nAutomated pipeline for Chinese news aggregation and digest generation.\n\n## Quick Start (3 步搞定)\n\n```bash\n# 第 1 步：安装依赖\npip install requests beautifulsoup4\n\n# 第 2 步：一键初始化（建表 + 插入示例网站 + 关键词）\npython scripts/news_digest_v2/init_db.py\n\n# 第 3 步：运行摘要\npython scripts/news_digest_v2/run_all_stages.py\n```\n\n或者一条命令全部搞定：\n```bash\npython scripts/news_digest_v2/quick_start.py\n```\n\nOutput: `.news-digest-out.md` (workspace) + `新闻摘要_YYYYMMDD_HHMMSS.txt` (desktop)\n\n## Architecture\n\n```\nStage 1:   Fetch     →  Scrape websites → Filter → Save to SQLite DB\nStage 2:   Process   →  Deduplicate (≥90% similarity) → Tag keywords\nStage 2.5: LLM       →  Batch LLM summarization (optional, requires API key)\nStage 3:   Output    →  Read LLM summaries (fallback to rule summaries) → Save to files\n```\n\n## Setup\n\n### Prerequisites\n\n- Python 3.8+ with: `requests`, `beautifulsoup4`\n- SQLite (built-in)\n\n### Initialize Database\n\nRun the init script to create tables and seed with sample data:\n\n```bash\npython scripts/news_digest_v2/init_db.py\n```\n\nThis creates:\n- Database tables (articles, monitor_websites, system_keywords, digest_output)\n- 10 sample news websites (People.cn, Xinhua, 36Kr, etc.)\n- 18 sample keywords (产业, 政策, 经济, 科技, etc.)\n\nDefault database path: `news.db` (in the skill directory).\nOverride with environment variable: `NEWS_DIGEST_DB=/your/path/news.db`\n\n### Customizing Your Sources\n\nAfter initialization, add or remove websites and keywords via SQL:\n\n```sql\n-- Add a website\nINSERT INTO monitor_websites (name, url, selector, category, priority)\nVALUES ('示例网站', 'https://example.com', 'a', '财经', 1);\n\n-- Add a keyword\nINSERT INTO system_keywords (keyword, category, weight)\nVALUES ('新能源', 'core', 5);\n```\n\n### Core Database Tables\n\n| Table | Purpose |\n|-------|---------|\n| `articles` | Scraped news articles (title, content, URL, date, keywords, duplicate flag) |\n| `monitor_websites` | Monitored websites (name, URL, CSS selector, category, enabled) |\n| `system_keywords` | Keywords for relevance scoring (core vs auxiliary, with weight) |\n| `digest_output` | LLM-generated summaries (optional) |\n\n## Usage\n\n### Full Pipeline\n\n```bash\npython scripts/news_digest_v2/run_all_stages.py\n```\n\nTakes ~13 minutes (network + LLM bound).\n\n### One-Command Quick Start\n\n```bash\npython scripts/news_digest_v2/quick_start.py\n```\n\nRuns init + fetch + process + output in one shot.\n\n### Cron Job Example\n\n```yaml\nschedule: \"0 20 * * *\"  # Daily 20:00\npayload:\n  run: python scripts/news_digest_v2/run_all_stages.py\n  then: read .news-digest-out.md and send to messaging\ntimeout: 900  # 15 minutes\n```\n\n## Output Format\n\n```\n【来源：标题】\n摘要内容（智能选段，300字以内，包含关键数据和核心事实）\n发布时间：YYYY-MM-DD\n原文链接：http://...\n```\n\n### 摘要质量保证\n\n**不完整句子自动过滤**：\n- 摘要末尾以逗号、顿号、分号、冒号等结尾 → 回退截断到上一个句号\n- 全文没有句号（整段残缺）→ 直接丢弃，不输出\n- 截断时信息损失超过 40% → 整段放弃，宁缺毋滥\n\n**教程/指南类内容全部过滤**：\n- 标题或内容包含\"教程\"、\"指南\"、\"攻略\"、\"手把手\"、\"从零开始\"等 → 自动排除\n- 详见 `rules_config.py` 中 `social` 分类的教程关键词列表\n\n## Key Features\n\n### Smart Summary Extraction (fetcher.py → extract_brief_summary)\n\nNot simple truncation. Each paragraph is scored by:\n- **Position**: Lead paragraph +10, top-3 +5 (inverted pyramid journalism)\n- **Data density**: Numbers × 2.5\n- **Signal words**: 印发/发布/宣布/决定/完成/启动 (+2 each)\n- **Entity density**: Organizations, locations (+1 each)\n- **Completeness**: Full sentence ending +3\n\nThen filtered: removes image captions, journalist bylines, ads, subtitles, boilerplate.\n\n### 摘要后处理 (formatter.py → clean_summary)\n\n- 电头/记者署名清理（预编译正则，支持新华社、中新网、财联社等）\n- 不完整句子过滤：以逗号/顿号/分号结尾 → 回退到上一个句号\n- 全文无句号 → 丢弃（不输出残缺内容）\n\n### Filtering Rules (rules_config.py)\n\nExcluded topics: entertainment, social news, violence, crime cases, health/wellness, education, automotive consumer news, science popularization (科普类), animal/archaeology news.\n\n**教程类（全部过滤）**：教程、指南、攻略、入门、自学、从零开始、手把手、保姆级教程、怎么做、如何使用、操作步骤、图文教程、视频教程、科研绘图、PS教程、Illustrator 等。\n\n**企业宣传稿/软文（全部过滤）**：产能突破、全线投产、技术溢出、供应链底气、跨界营销、负面舆情、品鉴官、品牌定位等。\n\n**教育/社会活动/颁奖（全部过滤）**：十佳、颁奖仪式、表彰、职校生、职业院校、评选、杰出代表、工匠精神等。\n\nInvalid keywords: clickbait patterns, advertising, webpage navigation elements.\n\n### Deduplication (similarity.py)\n\n- Jaccard 2-gram similarity\n- Threshold: ≥90% → mark as duplicate\n- Only one version appears in output\n\n### Date Filtering\n\n- Normal: within 3 days\n- Holidays: within 7 days\n- No date → discard\n- Old URLs (year > 1 year ago) → skip\n\n## Configuration\n\n### Environment Variables\n\n| Variable | Default | Description |\n|----------|---------|-------------|\n| `NEWS_DIGEST_DB` | `news.db` | SQLite database path |\n| `NEWS_DIGEST_LLM_API_KEY` | (empty) | LLM API key for Stage 2.5 summarization |\n| `NEWS_DIGEST_LLM_BASE_URL` | (empty) | LLM API base URL |\n| `NEWS_DIGEST_LLM_MODEL` | `qwen3.6-plus` | LLM model name |\n\nIf LLM env vars are not set, Stage 2.5 is silently skipped and rule-based summaries are used instead.\n\n## Files\n\n```\nnews-digest/\n├── SKILL.md\n└── scripts/\n    └── news_digest_v2/\n        ├── __init__.py\n        ├── config.py               # DB path, websites, keywords, holidays, LLM config\n        ├── database.py             # SQLite operations\n        ├── fetcher.py              # Web scraping + smart summary extraction\n        ├── filters.py              # Content filtering logic\n        ├── formatter.py            # Output formatting + incomplete sentence handling\n        ├── init_db.py              # One-click database initialization (NEW in v1.0.1)\n        ├── quick_start.py          # One-command full pipeline (NEW in v1.0.1)\n        ├── rules_config.py         # Exclusion rules, keywords, dateline patterns\n        ├── similarity.py           # Jaccard deduplication\n        ├── stage1_fetch.py         # Stage 1 entry (fetch)\n        ├── stage2_process.py       # Stage 2 entry (dedup + keywords)\n        ├── stage2_5_llm_summary.py # Stage 2.5 (LLM batch summarization)\n        ├── stage3_output.py        # Stage 3 entry (read + format + save)\n        └── run_all_stages.py       # Full pipeline entry\n```\n\n## FAQ\n\n**Q: 安装后跑不起来？**\nA: 确保先运行了 `init_db.py` 初始化数据库。没有数据库和示例数据，后续步骤会失败。\n\n**Q: pip install 失败？**\nA: 尝试 `pip install --upgrade pip` 后再安装。如果网络问题，使用 `pip install -i https://pypi.tuna.tsinghua.edu.cn/simple requests beautifulsoup4`。\n\n**Q: 某些网站抓取失败？**\nA: 正常现象。部分网站有反爬或 SSL 问题，脚本会继续处理其他网站。不影响最终输出。\n\n**Q: 输出是空的？**\nA: 检查数据库中是否有数据。运行 `python scripts/news_digest_v2/init_db.py` 重新初始化。\n\n**Q: 如何自定义监测网站？**\nA: 通过 SQL 插入 `monitor_websites` 表，字段：name, url, selector, category, priority。\n\n**Q: 数据库会越来越大吗？**\nA: 约 30-50 条/天。建议定期清理旧数据，或删除 `news.db` 后重新初始化。\n\n## Performance Notes\n\n- ~5 minutes for full scrape (network I/O bound)\n- Some sites may fail (SSL issues, 521 errors, 404s) — pipeline continues\n- Recommended cron timeout: 900 seconds\n- **数据库是增量追加的**，不会被清空。新新闻按 URL 去重插入（`INSERT OR IGNORE`），旧新闻保留。\n- 重复新闻标记 `is_duplicate = 1`，不删除。\n\n## Changelog\n\n### v1.0.12 (2026-06-09)\n- **新增 6 类过滤规则**（`rules_config.py`）：\n  - 文化活动/精神宣讲：诵读会、宣讲会、科学家精神、阅读角、推广计划、弘扬.*精神等\n  - 信用知识科普/问答类：信用知识、百问百答、信用体系、信用记录等\n  - 新闻简报/早报类：8点1氪、点早闻、早晚报、每日速递、资讯速览等\n- **去重修复**：清理 digest_output 中 is_duplicate=1 的残留条目\n- **sync_to_skill.py 修复**：新增 cross_day_dedup.py 到同步列表\n\n### v1.0.11 (2026-06-08)\n- **新增人民日报主站监控**：`parse_rmrbhwb` 路由增加 `paper.people.com.cn/rmrb/` 支持（数据库新增 id=47, priority=1）\n- **跨天去重修复**：`cross_day_dedup.py` 新增硬规则4（标题互相包含检测）+ Jaccard权重 0.5→0.6\n- **header 恢复**：formatter.py 恢复'来源网站: X | 收录新闻: Y 条'格式\n- **stage2_5 优化**：SQL查询简化 + 日志优化\n\n### v1.0.10 (2026-06-05)\n- **Header 条数修复**：`formatter.py` 中\"收录新闻\"统计从预过滤数改为实际输出数\n  - 旧逻辑：统计 `filtered_news` 长度（排除重复后），但未减去内容类型黑名单过滤的条目\n  - 新逻辑：用占位符+后置替换，确保 header 数字与实际输出条数一致\n- **标题截断保护**：标题以逗号/顿号结尾时自动补省略号，避免数据库截断误导\n- **新增过滤规则 5 类**：\n  - 党建历史/纪念性：伟大征程、永放光芒、精神永存、丰碑、铸魂、初心等 17 词\n  - 节气科普：节气、农忙、天文专家、正本清源、二十四节气等\n  - 官员落马/公诉：提起公诉、受贿案、指定管辖、监察调查等 12 词（同时加入 `corporate_scandal` 分类）\n\n### v1.0.8 (2026-06-05)\n- **跨天去重**：新增 `cross_day_dedup.py`，对比最近 3 天历史摘要自动拦截跨天重复新闻\n  - 核心算法：标题 Jaccard(0.5) + 数字匹配(0.25) + 内容词重叠(0.25)\n  - 硬规则：显著数字不共享 → 直接判 0（自动放行 PMI/CPI 等周期性新闻，无需白名单）\n  - 三档判定：≥0.75 拦截 / 0.60~0.75 警告保留 / <0.60 正常通过\n- **标题改写规则 9**：参展/出席类标题改为事件主体，IPO/融资/财报等保留公司名\n- **禁止编造规则 10**：摘要中所有信息必须来源于原文，不得自行补充外部知识\n- **来源数显示修复**：`stage1_fetch.py` 从硬编码 42 改为动态读取 `len(WEBSITES)`\n- **GBK 解码增强**：`decode_response()` 已合并入主流程，支持所有已知 GBK 来源\n\n### v1.0.7 (2026-06-01)\n- **标题乱码修复**：\n  - `fetcher.py` 新增 `decode_response()` 函数，对已知 GBK 编码来源（人民日报海外版）强制使用 GBK 解码，从根源修复 Cyrillic 乱码\n  - `stage2_5_llm_summary.py` 新增乱码标题检测，发现乱码时提示 LLM 从正文生成准确标题\n  - `formatter.py` 新增乱码标题兜底过滤\n- **内容类型黑名单**：\n  - 新增 `TITLE_EXCLUDE_KEYWORDS`（评论丨/时评/社评/深度观察/记者观察/招聘/面试/递补/人事任免/讣告/专访等）\n  - 新增 `URL_EXCLUDE_PATTERNS`（人民网评论频道等）\n  - `formatter.py` 输出时自动跳过非硬新闻类型\n- **推荐超时**：900s→1200s（实测 LLM 总结 50 条 + 抓取总耗时 ~1150s）\n\n### v1.0.6 (2026-05-29)\n- **新闻源增至 46 个**（45 启用，1 个\"新华每日电讯\"禁用）\n- **LLM 模型修复**：默认模型从 `qwen-plus`（不存在，400 错误）改为 `qwen3.6-plus`\n- **LLM_BATCH_SIZE**：35→50，与 MAX_OUTPUT_COUNT 一致\n- **来源数显示修复**：输出头部\"来源网站\"从动态统计改为固定 46\n- **新增过滤规则**：`corporate_pr`（企业宣传稿/软文）+ `education_social`（教育/颁奖/评选）\n- **推荐超时**：600s→900s（因 LLM 总结 50 条耗时增加）\n- **飞书队列清理**：不再自动写入飞书\n\n### v1.0.5 (2026-05-28)\n- **新增新闻源**：安徽日报、人民日报海外版、新华每日电讯（42→45个启用源）\n- **MAX_OUTPUT_COUNT**：35→50条\n- **新增关键词**：十五五、标准、纲要、公报、全覆盖、创新药、芯片、测评、公共服务（12个）\n- **fetcher.py 兼容修复**：4处相对导入改为 try/except fallback，支持直接运行和包导入\n- **中国工信网超时**：10秒→30秒\n- **数据库**：新闻源从 SQLite 加载，关键词从数据库读取\n\n### v1.0.4 (2026-05-25)\n- **标题清理**：自动去除标题首尾的多余符号（如中点 `·`、空格）\n- **输出排序优化**：权威来源（人民网/新华网等）按级别升序排列，同级别按时间倒序\n- **过滤规则增强**：新增艺术展览/书画捐赠过滤词（避免非产业类文化新闻干扰）\n- **权威选文修复**：修正文章数 <35 时未排序的 Bug\n\n### v1.0.3 (2026-05-22)\n- **央广网独立解析器** (`parse_cnr`): 央广网页面标题和正文在同一 `<a>` 标签内，新增独立解析器只取 `<strong>` 作为标题，避免标题+正文混一起导致标题过长被过滤\n- **权威来源优先选文**: 阶段 2.5（LLM 批量总结）和阶段 3（输出）都应用权威来源保底策略，每个权威来源（人民网、新华网、央广网、经济日报、科技日报、科学网、中国科技网、科创版日报、中国经济网）至少入选 2 条，避免被中国经济网和中宏网等高产源淹没\n- **新华网子频道归并**: 新华能源、新华科创、新华时政、新华汽车等子频道统一归并到\"新华网\"来源\n- **社会新闻过滤增强**: 新增交通违法/行政处罚类社会新闻过滤词（罚款、行拘、拘留、交警、变造号牌等）\n- **编码检测优化**: 央广网等 GB2312 编码网站从 HTML `<meta>` 标签检测编码，提高抓取成功率\n\n### v1.0.1 (2026-05-06)\n- Added `init_db.py` for one-click database initialization with sample data\n- Added `quick_start.py` for one-command full pipeline\n- Simplified SKILL.md installation guide to 3 steps\n- Added FAQ section\n- Updated example websites to 10 mainstream Chinese news sources\n\n### v1.0.0 (2026-05-05)\n- Initial release\n\nFile v1.0.12:_meta.json\n\n{\n  \"ownerId\": \"kn7983et5m2qgha7y8gvb8dgsn84w87m\",\n  \"slug\": \"news-digest-v1\",\n  \"version\": \"1.0.12\",\n  \"publishedAt\": 1780972826589\n}\n\nFile v1.0.12:skill-card.md\n\n## Description:\n\nAutomatically scrapes Chinese news sources, filters and deduplicates articles, and generates daily summaries with source attribution and original links.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[zigu-creator](https://clawhub.ai/user/zigu-creator)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers and operators use this skill to set up a local Chinese news monitoring workflow that fetches news, filters low-relevance content, optionally summarizes with an LLM, and produces a daily digest for review or distribution.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: The skill can make broad network requests to configured news sites and custom article links.\n\nMitigation: Run it only with reviewed monitor_websites entries and trusted article sources before using the full pipeline.\n\nRisk: The LLM stage can reuse an OpenClaw API key from local configuration when dedicated NEWS_DIGEST_LLM credentials are not set.\n\nMitigation: Set dedicated, limited NEWS_DIGEST_LLM_API_KEY and NEWS_DIGEST_LLM_BASE_URL values, or remove the OpenClaw config fallback before execution.\n\nRisk: The skill writes a local SQLite database plus digest files in the workspace and Desktop.\n\nMitigation: Run it in a workspace where these writes are expected, and review generated digest content before sharing it.\n\n## Reference(s):\n\n- [ClawHub skill page](https://clawhub.ai/zigu-creator/skills/news-digest-v1)\n- [Publisher profile](https://clawhub.ai/user/zigu-creator)\n\n## Skill Output:\n\n**Output Type(s):** [text, markdown, shell commands, configuration, guidance]\n\n**Output Format:** [Markdown and plain text digest files with source titles, summaries, dates, and original links; setup guidance includes Markdown with bash and SQL snippets.]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [Writes .news-digest-out.md in the workspace and a timestamped Chinese text digest on Desktop; optional LLM summarization uses configured API credentials.]\n\n## Skill Version(s):\n\n1.0.12 (source: SKILL.md frontmatter and ClawHub release evidence, released 2026-06-09)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nArchive v1.0.6: 19 files, 59788 bytes\n\nFiles: scripts/news_digest_v2/__init__.py (296b), scripts/news_digest_v2/config.py (6769b), scripts/news_digest_v2/cross_day_dedup.py (7583b), scripts/news_digest_v2/database.py (13334b), scripts/news_digest_v2/fetcher.py (33547b), scripts/news_digest_v2/filters.py (3258b), scripts/news_digest_v2/formatter.py (6710b), scripts/news_digest_v2/init_db.py (6849b), scripts/news_digest_v2/quick_start.py (1966b), scripts/news_digest_v2/rules_config.py (24979b), scripts/news_digest_v2/run_all_stages.py (3685b), scripts/news_digest_v2/similarity.py (2348b), scripts/news_digest_v2/stage1_fetch.py (1808b), scripts/news_digest_v2/stage2_5_llm_summary.py (14948b), scripts/news_digest_v2/stage2_process.py (2431b), scripts/news_digest_v2/stage3_output.py (7967b), skill-card.md (2359b), SKILL.md (14169b), _meta.json (133b)\n\nFile v1.0.6:SKILL.md\n\n---\nname: news-digest\ndescription: \"Automatically scrape, process, and generate daily news digests from Chinese news sources. Covers industry dynamics, policy updates, economy, tech, energy, and pricing information. Use when: user asks for daily news summary, news digest, 每日新闻摘要, 新闻汇总, 新闻摘要, or wants to set up automated news monitoring from Chinese news websites. Outputs formatted summaries with source attribution and original links.\"\nversion: 1.0.9\n---\n\n# News Digest - 每日新闻摘要\n\nAutomated pipeline for Chinese news aggregation and digest generation.\n\n## Quick Start (3 步搞定)\n\n```bash\n# 第 1 步：安装依赖\npip install requests beautifulsoup4\n\n# 第 2 步：一键初始化（建表 + 插入示例网站 + 关键词）\npython scripts/news_digest_v2/init_db.py\n\n# 第 3 步：运行摘要\npython scripts/news_digest_v2/run_all_stages.py\n```\n\n或者一条命令全部搞定：\n```bash\npython scripts/news_digest_v2/quick_start.py\n```\n\nOutput: `.news-digest-out.md` (workspace) + `新闻摘要_YYYYMMDD_HHMMSS.txt` (desktop)\n\n## Architecture\n\n```\nStage 1:   Fetch     →  Scrape websites → Filter → Save to SQLite DB\nStage 2:   Process   →  Deduplicate (≥90% similarity) → Tag keywords\nStage 2.5: LLM       →  Batch LLM summarization (optional, requires API key)\nStage 3:   Output    →  Read LLM summaries (fallback to rule summaries) → Save to files\n```\n\n## Setup\n\n### Prerequisites\n\n- Python 3.8+ with: `requests`, `beautifulsoup4`\n- SQLite (built-in)\n\n### Initialize Database\n\nRun the init script to create tables and seed with sample data:\n\n```bash\npython scripts/news_digest_v2/init_db.py\n```\n\nThis creates:\n- Database tables (articles, monitor_websites, system_keywords, digest_output)\n- 10 sample news websites (People.cn, Xinhua, 36Kr, etc.)\n- 18 sample keywords (产业, 政策, 经济, 科技, etc.)\n\nDefault database path: `news.db` (in the skill directory).\nOverride with environment variable: `NEWS_DIGEST_DB=/your/path/news.db`\n\n### Customizing Your Sources\n\nAfter initialization, add or remove websites and keywords via SQL:\n\n```sql\n-- Add a website\nINSERT INTO monitor_websites (name, url, selector, category, priority)\nVALUES ('示例网站', 'https://example.com', 'a', '财经', 1);\n\n-- Add a keyword\nINSERT INTO system_keywords (keyword, category, weight)\nVALUES ('新能源', 'core', 5);\n```\n\n### Core Database Tables\n\n| Table | Purpose |\n|-------|---------|\n| `articles` | Scraped news articles (title, content, URL, date, keywords, duplicate flag) |\n| `monitor_websites` | Monitored websites (name, URL, CSS selector, category, enabled) |\n| `system_keywords` | Keywords for relevance scoring (core vs auxiliary, with weight) |\n| `digest_output` | LLM-generated summaries (optional) |\n\n## Usage\n\n### Full Pipeline\n\n```bash\npython scripts/news_digest_v2/run_all_stages.py\n```\n\nTakes ~13 minutes (network + LLM bound).\n\n### One-Command Quick Start\n\n```bash\npython scripts/news_digest_v2/quick_start.py\n```\n\nRuns init + fetch + process + output in one shot.\n\n### Cron Job Example\n\n```yaml\nschedule: \"0 20 * * *\"  # Daily 20:00\npayload:\n  run: python scripts/news_digest_v2/run_all_stages.py\n  then: read .news-digest-out.md and send to messaging\ntimeout: 900  # 15 minutes\n```\n\n## Output Format\n\n```\n【来源：标题】\n摘要内容（智能选段，300字以内，包含关键数据和核心事实）\n发布时间：YYYY-MM-DD\n原文链接：http://...\n```\n\n### 摘要质量保证\n\n**不完整句子自动过滤**：\n- 摘要末尾以逗号、顿号、分号、冒号等结尾 → 回退截断到上一个句号\n- 全文没有句号（整段残缺）→ 直接丢弃，不输出\n- 截断时信息损失超过 40% → 整段放弃，宁缺毋滥\n\n**教程/指南类内容全部过滤**：\n- 标题或内容包含\"教程\"、\"指南\"、\"攻略\"、\"手把手\"、\"从零开始\"等 → 自动排除\n- 详见 `rules_config.py` 中 `social` 分类的教程关键词列表\n\n## Key Features\n\n### Smart Summary Extraction (fetcher.py → extract_brief_summary)\n\nNot simple truncation. Each paragraph is scored by:\n- **Position**: Lead paragraph +10, top-3 +5 (inverted pyramid journalism)\n- **Data density**: Numbers × 2.5\n- **Signal words**: 印发/发布/宣布/决定/完成/启动 (+2 each)\n- **Entity density**: Organizations, locations (+1 each)\n- **Completeness**: Full sentence ending +3\n\nThen filtered: removes image captions, journalist bylines, ads, subtitles, boilerplate.\n\n### 摘要后处理 (formatter.py → clean_summary)\n\n- 电头/记者署名清理（预编译正则，支持新华社、中新网、财联社等）\n- 不完整句子过滤：以逗号/顿号/分号结尾 → 回退到上一个句号\n- 全文无句号 → 丢弃（不输出残缺内容）\n\n### Filtering Rules (rules_config.py)\n\nExcluded topics: entertainment, social news, violence, crime cases, health/wellness, education, automotive consumer news, science popularization (科普类), animal/archaeology news.\n\n**教程类（全部过滤）**：教程、指南、攻略、入门、自学、从零开始、手把手、保姆级教程、怎么做、如何使用、操作步骤、图文教程、视频教程、科研绘图、PS教程、Illustrator 等。\n\n**企业宣传稿/软文（全部过滤）**：产能突破、全线投产、技术溢出、供应链底气、跨界营销、负面舆情、品鉴官、品牌定位等。\n\n**教育/社会活动/颁奖（全部过滤）**：十佳、颁奖仪式、表彰、职校生、职业院校、评选、杰出代表、工匠精神等。\n\nInvalid keywords: clickbait patterns, advertising, webpage navigation elements.\n\n### Deduplication (similarity.py)\n\n- Jaccard 2-gram similarity\n- Threshold: ≥90% → mark as duplicate\n- Only one version appears in output\n\n### Date Filtering\n\n- Normal: within 3 days\n- Holidays: within 7 days\n- No date → discard\n- Old URLs (year > 1 year ago) → skip\n\n## Configuration\n\n### Environment Variables\n\n| Variable | Default | Description |\n|----------|---------|-------------|\n| `NEWS_DIGEST_DB` | `news.db` | SQLite database path |\n| `NEWS_DIGEST_LLM_API_KEY` | (empty) | LLM API key for Stage 2.5 summarization |\n| `NEWS_DIGEST_LLM_BASE_URL` | (empty) | LLM API base URL |\n| `NEWS_DIGEST_LLM_MODEL` | `qwen3.6-plus` | LLM model name |\n\nIf LLM env vars are not set, Stage 2.5 is silently skipped and rule-based summaries are used instead.\n\n## Files\n\n```\nnews-digest/\n├── SKILL.md\n└── scripts/\n    └── news_digest_v2/\n        ├── __init__.py\n        ├── config.py               # DB path, websites, keywords, holidays, LLM config\n        ├── database.py             # SQLite operations\n        ├── fetcher.py              # Web scraping + smart summary extraction\n        ├── filters.py              # Content filtering logic\n        ├── formatter.py            # Output formatting + incomplete sentence handling\n        ├── init_db.py              # One-click database initialization (NEW in v1.0.1)\n        ├── quick_start.py          # One-command full pipeline (NEW in v1.0.1)\n        ├── rules_config.py         # Exclusion rules, keywords, dateline patterns\n        ├── similarity.py           # Jaccard deduplication\n        ├── stage1_fetch.py         # Stage 1 entry (fetch)\n        ├── stage2_process.py       # Stage 2 entry (dedup + keywords)\n        ├── stage2_5_llm_summary.py # Stage 2.5 (LLM batch summarization)\n        ├── stage3_output.py        # Stage 3 entry (read + format + save)\n        └── run_all_stages.py       # Full pipeline entry\n```\n\n## FAQ\n\n**Q: 安装后跑不起来？**\nA: 确保先运行了 `init_db.py` 初始化数据库。没有数据库和示例数据，后续步骤会失败。\n\n**Q: pip install 失败？**\nA: 尝试 `pip install --upgrade pip` 后再安装。如果网络问题，使用 `pip install -i https://pypi.tuna.tsinghua.edu.cn/simple requests beautifulsoup4`。\n\n**Q: 某些网站抓取失败？**\nA: 正常现象。部分网站有反爬或 SSL 问题，脚本会继续处理其他网站。不影响最终输出。\n\n**Q: 输出是空的？**\nA: 检查数据库中是否有数据。运行 `python scripts/news_digest_v2/init_db.py` 重新初始化。\n\n**Q: 如何自定义监测网站？**\nA: 通过 SQL 插入 `monitor_websites` 表，字段：name, url, selector, category, priority。\n\n**Q: 数据库会越来越大吗？**\nA: 约 30-50 条/天。建议定期清理旧数据，或删除 `news.db` 后重新初始化。\n\n## Performance Notes\n\n- ~5 minutes for full scrape (network I/O bound)\n- Some sites may fail (SSL issues, 521 errors, 404s) — pipeline continues\n- Recommended cron timeout: 900 seconds\n- **数据库是增量追加的**，不会被清空。新新闻按 URL 去重插入（`INSERT OR IGNORE`），旧新闻保留。\n- 重复新闻标记 `is_duplicate = 1`，不删除。\n\n## Changelog\n\n### v1.0.9 (2026-06-05)\n- **Header 条数修复**：`formatter.py` 中\"收录新闻\"统计从预过滤数改为实际输出数\n  - 旧逻辑：统计 `filtered_news` 长度（排除重复后），但未减去内容类型黑名单过滤的条目\n  - 新逻辑：用占位符+后置替换，确保 header 数字与实际输出条数一致\n- **标题截断保护**：标题以逗号/顿号结尾时自动补省略号，避免数据库截断误导\n- **新增过滤规则 5 类**：\n  - 党建历史/纪念性：伟大征程、永放光芒、精神永存、丰碑、铸魂、初心等 17 词\n  - 节气科普：节气、农忙、天文专家、正本清源、二十四节气等\n  - 官员落马/公诉：提起公诉、受贿案、指定管辖、监察调查等 12 词（同时加入 `corporate_scandal` 分类）\n\n### v1.0.8 (2026-06-05)\n- **跨天去重**：新增 `cross_day_dedup.py`，对比最近 3 天历史摘要自动拦截跨天重复新闻\n  - 核心算法：标题 Jaccard(0.5) + 数字匹配(0.25) + 内容词重叠(0.25)\n  - 硬规则：显著数字不共享 → 直接判 0（自动放行 PMI/CPI 等周期性新闻，无需白名单）\n  - 三档判定：≥0.75 拦截 / 0.60~0.75 警告保留 / <0.60 正常通过\n- **标题改写规则 9**：参展/出席类标题改为事件主体，IPO/融资/财报等保留公司名\n- **禁止编造规则 10**：摘要中所有信息必须来源于原文，不得自行补充外部知识\n- **来源数显示修复**：`stage1_fetch.py` 从硬编码 42 改为动态读取 `len(WEBSITES)`\n- **GBK 解码增强**：`decode_response()` 已合并入主流程，支持所有已知 GBK 来源\n\n### v1.0.7 (2026-06-01)\n- **标题乱码修复**：\n  - `fetcher.py` 新增 `decode_response()` 函数，对已知 GBK 编码来源（人民日报海外版）强制使用 GBK 解码，从根源修复 Cyrillic 乱码\n  - `stage2_5_llm_summary.py` 新增乱码标题检测，发现乱码时提示 LLM 从正文生成准确标题\n  - `formatter.py` 新增乱码标题兜底过滤\n- **内容类型黑名单**：\n  - 新增 `TITLE_EXCLUDE_KEYWORDS`（评论丨/时评/社评/深度观察/记者观察/招聘/面试/递补/人事任免/讣告/专访等）\n  - 新增 `URL_EXCLUDE_PATTERNS`（人民网评论频道等）\n  - `formatter.py` 输出时自动跳过非硬新闻类型\n- **推荐超时**：900s→1200s（实测 LLM 总结 50 条 + 抓取总耗时 ~1150s）\n\n### v1.0.6 (2026-05-29)\n- **新闻源增至 46 个**（45 启用，1 个\"新华每日电讯\"禁用）\n- **LLM 模型修复**：默认模型从 `qwen-plus`（不存在，400 错误）改为 `qwen3.6-plus`\n- **LLM_BATCH_SIZE**：35→50，与 MAX_OUTPUT_COUNT 一致\n- **来源数显示修复**：输出头部\"来源网站\"从动态统计改为固定 46\n- **新增过滤规则**：`corporate_pr`（企业宣传稿/软文）+ `education_social`（教育/颁奖/评选）\n- **推荐超时**：600s→900s（因 LLM 总结 50 条耗时增加）\n- **飞书队列清理**：不再自动写入飞书\n\n### v1.0.5 (2026-05-28)\n- **新增新闻源**：安徽日报、人民日报海外版、新华每日电讯（42→45个启用源）\n- **MAX_OUTPUT_COUNT**：35→50条\n- **新增关键词**：十五五、标准、纲要、公报、全覆盖、创新药、芯片、测评、公共服务（12个）\n- **fetcher.py 兼容修复**：4处相对导入改为 try/except fallback，支持直接运行和包导入\n- **中国工信网超时**：10秒→30秒\n- **数据库**：新闻源从 SQLite 加载，关键词从数据库读取\n\n### v1.0.4 (2026-05-25)\n- **标题清理**：自动去除标题首尾的多余符号（如中点 `·`、空格）\n- **输出排序优化**：权威来源（人民网/新华网等）按级别升序排列，同级别按时间倒序\n- **过滤规则增强**：新增艺术展览/书画捐赠过滤词（避免非产业类文化新闻干扰）\n- **权威选文修复**：修正文章数 <35 时未排序的 Bug\n\n### v1.0.3 (2026-05-22)\n- **央广网独立解析器** (`parse_cnr`): 央广网页面标题和正文在同一 `<a>` 标签内，新增独立解析器只取 `<strong>` 作为标题，避免标题+正文混一起导致标题过长被过滤\n- **权威来源优先选文**: 阶段 2.5（LLM 批量总结）和阶段 3（输出）都应用权威来源保底策略，每个权威来源（人民网、新华网、央广网、经济日报、科技日报、科学网、中国科技网、科创版日报、中国经济网）至少入选 2 条，避免被中国经济网和中宏网等高产源淹没\n- **新华网子频道归并**: 新华能源、新华科创、新华时政、新华汽车等子频道统一归并到\"新华网\"来源\n- **社会新闻过滤增强**: 新增交通违法/行政处罚类社会新闻过滤词（罚款、行拘、拘留、交警、变造号牌等）\n- **编码检测优化**: 央广网等 GB2312 编码网站从 HTML `<meta>` 标签检测编码，提高抓取成功率\n\n### v1.0.1 (2026-05-06)\n- Added `init_db.py` for one-click database initialization with sample data\n- Added `quick_start.py` for one-command full pipeline\n- Simplified SKILL.md installation guide to 3 steps\n- Added FAQ section\n- Updated example websites to 10 mainstream Chinese news sources\n\n### v1.0.0 (2026-05-05)\n- Initial release\n\nFile v1.0.6:_meta.json\n\n{\n  \"ownerId\": \"kn7983et5m2qgha7y8gvb8dgsn84w87m\",\n  \"slug\": \"news-digest-v1\",\n  \"version\": \"1.0.6\",\n  \"publishedAt\": 1780655669678\n}\n\nFile v1.0.6:skill-card.md\n\n## Description: <br>\nAutomatically scrapes Chinese news sources, filters and deduplicates articles, and generates daily news digests with source attribution and original links. <br>\n\nThis skill is ready for commercial/non-commercial use. <br>\n\n## Publisher: <br>\n[zigu-creator](https://clawhub.ai/user/zigu-creator) <br>\n\n### License/Terms of Use: <br>\nMIT-0 <br>\n\n\n## Use Case: <br>\nDevelopers and operators use this skill to run an automated daily monitoring pipeline for Chinese news websites, producing concise digests for industry dynamics, policy updates, economy, technology, energy, and pricing information. <br>\n\n### Deployment Geography for Use: <br>\nGlobal <br>\n\n## Known Risks and Mitigations: <br>\nRisk: The skill may use local OpenClaw or LLM API credentials and send scraped article text and metadata to the configured LLM endpoint. <br>\nMitigation: Use a trusted LLM endpoint, prefer dedicated low-privilege credentials, and run only on news sources whose content is appropriate to share with that service. <br>\nRisk: Some website fetches may use weakened TLS verification or encounter unreliable third-party news sites. <br>\nMitigation: Restrict monitored sources to trusted sites and review generated summaries and links before redistribution. <br>\nRisk: The pipeline writes generated files to both the workspace and Desktop. <br>\nMitigation: Run it in an intended workspace and review output paths before scheduling unattended runs. <br>\n\n\n## Reference(s): <br>\n- [ClawHub skill page](https://clawhub.ai/zigu-creator/news-digest-v1) <br>\n\n\n## Skill Output: <br>\n**Output Type(s):** [text, markdown, shell commands, configuration, guidance] <br>\n**Output Format:** [Markdown and plain-text digest files with source labels, summaries, publish dates, and original article links.] <br>\n**Output Parameters:** [1D] <br>\n**Other Properties Related to Output:** [Writes .news-digest-out.md in the workspace and a timestamped news digest text file on the Desktop when the full pipeline is run.] <br>\n\n## Skill Version(s): <br>\n1.0.6 (source: server release evidence) <br>\n\n## Ethical Considerations: <br>\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment. <br>\n\nArchive v1.0.5: 18 files, 53818 bytes\n\nFiles: scripts/news_digest_v2/__init__.py (296b), scripts/news_digest_v2/config.py (6337b), scripts/news_digest_v2/database.py (13281b), scripts/news_digest_v2/fetcher.py (33472b), scripts/news_digest_v2/filters.py (3258b), scripts/news_digest_v2/formatter.py (5913b), scripts/news_digest_v2/init_db.py (6849b), scripts/news_digest_v2/quick_start.py (1966b), scripts/news_digest_v2/rules_config.py (23868b), scripts/news_digest_v2/run_all_stages.py (3685b), scripts/news_digest_v2/similarity.py (2348b), scripts/news_digest_v2/stage1_fetch.py (1761b), scripts/news_digest_v2/stage2_5_llm_summary.py (12329b), scripts/news_digest_v2/stage2_process.py (2431b), scripts/news_digest_v2/stage3_output.py (7967b), skill-card.md (2648b), SKILL.md (12529b), _meta.json (133b)\n\nFile v1.0.5:SKILL.md\n\n---\nname: news-digest\ndescription: \"Automatically scrape, process, and generate daily news digests from Chinese news sources. Covers industry dynamics, policy updates, economy, tech, energy, and pricing information. Use when: user asks for daily news summary, news digest, 每日新闻摘要, 新闻汇总, 新闻摘要, or wants to set up automated news monitoring from Chinese news websites. Outputs formatted summaries with source attribution and original links.\"\nversion: 1.0.7\n---\n\n# News Digest - 每日新闻摘要\n\nAutomated pipeline for Chinese news aggregation and digest generation.\n\n## Quick Start (3 步搞定)\n\n```bash\n# 第 1 步：安装依赖\npip install requests beautifulsoup4\n\n# 第 2 步：一键初始化（建表 + 插入示例网站 + 关键词）\npython scripts/news_digest_v2/init_db.py\n\n# 第 3 步：运行摘要\npython scripts/news_digest_v2/run_all_stages.py\n```\n\n或者一条命令全部搞定：\n```bash\npython scripts/news_digest_v2/quick_start.py\n```\n\nOutput: `.news-digest-out.md` (workspace) + `新闻摘要_YYYYMMDD_HHMMSS.txt` (desktop)\n\n## Architecture\n\n```\nStage 1:   Fetch     →  Scrape websites → Filter → Save to SQLite DB\nStage 2:   Process   →  Deduplicate (≥90% similarity) → Tag keywords\nStage 2.5: LLM       →  Batch LLM summarization (optional, requires API key)\nStage 3:   Output    →  Read LLM summaries (fallback to rule summaries) → Save to files\n```\n\n## Setup\n\n### Prerequisites\n\n- Python 3.8+ with: `requests`, `beautifulsoup4`\n- SQLite (built-in)\n\n### Initialize Database\n\nRun the init script to create tables and seed with sample data:\n\n```bash\npython scripts/news_digest_v2/init_db.py\n```\n\nThis creates:\n- Database tables (articles, monitor_websites, system_keywords, digest_output)\n- 10 sample news websites (People.cn, Xinhua, 36Kr, etc.)\n- 18 sample keywords (产业, 政策, 经济, 科技, etc.)\n\nDefault database path: `news.db` (in the skill directory).\nOverride with environment variable: `NEWS_DIGEST_DB=/your/path/news.db`\n\n### Customizing Your Sources\n\nAfter initialization, add or remove websites and keywords via SQL:\n\n```sql\n-- Add a website\nINSERT INTO monitor_websites (name, url, selector, category, priority)\nVALUES ('示例网站', 'https://example.com', 'a', '财经', 1);\n\n-- Add a keyword\nINSERT INTO system_keywords (keyword, category, weight)\nVALUES ('新能源', 'core', 5);\n```\n\n### Core Database Tables\n\n| Table | Purpose |\n|-------|---------|\n| `articles` | Scraped news articles (title, content, URL, date, keywords, duplicate flag) |\n| `monitor_websites` | Monitored websites (name, URL, CSS selector, category, enabled) |\n| `system_keywords` | Keywords for relevance scoring (core vs auxiliary, with weight) |\n| `digest_output` | LLM-generated summaries (optional) |\n\n## Usage\n\n### Full Pipeline\n\n```bash\npython scripts/news_digest_v2/run_all_stages.py\n```\n\nTakes ~13 minutes (network + LLM bound).\n\n### One-Command Quick Start\n\n```bash\npython scripts/news_digest_v2/quick_start.py\n```\n\nRuns init + fetch + process + output in one shot.\n\n### Cron Job Example\n\n```yaml\nschedule: \"0 20 * * *\"  # Daily 20:00\npayload:\n  run: python scripts/news_digest_v2/run_all_stages.py\n  then: read .news-digest-out.md and send to messaging\ntimeout: 900  # 15 minutes\n```\n\n## Output Format\n\n```\n【来源：标题】\n摘要内容（智能选段，300字以内，包含关键数据和核心事实）\n发布时间：YYYY-MM-DD\n原文链接：http://...\n```\n\n### 摘要质量保证\n\n**不完整句子自动过滤**：\n- 摘要末尾以逗号、顿号、分号、冒号等结尾 → 回退截断到上一个句号\n- 全文没有句号（整段残缺）→ 直接丢弃，不输出\n- 截断时信息损失超过 40% → 整段放弃，宁缺毋滥\n\n**教程/指南类内容全部过滤**：\n- 标题或内容包含\"教程\"、\"指南\"、\"攻略\"、\"手把手\"、\"从零开始\"等 → 自动排除\n- 详见 `rules_config.py` 中 `social` 分类的教程关键词列表\n\n## Key Features\n\n### Smart Summary Extraction (fetcher.py → extract_brief_summary)\n\nNot simple truncation. Each paragraph is scored by:\n- **Position**: Lead paragraph +10, top-3 +5 (inverted pyramid journalism)\n- **Data density**: Numbers × 2.5\n- **Signal words**: 印发/发布/宣布/决定/完成/启动 (+2 each)\n- **Entity density**: Organizations, locations (+1 each)\n- **Completeness**: Full sentence ending +3\n\nThen filtered: removes image captions, journalist bylines, ads, subtitles, boilerplate.\n\n### 摘要后处理 (formatter.py → clean_summary)\n\n- 电头/记者署名清理（预编译正则，支持新华社、中新网、财联社等）\n- 不完整句子过滤：以逗号/顿号/分号结尾 → 回退到上一个句号\n- 全文无句号 → 丢弃（不输出残缺内容）\n\n### Filtering Rules (rules_config.py)\n\nExcluded topics: entertainment, social news, violence, crime cases, health/wellness, education, automotive consumer news, science popularization (科普类), animal/archaeology news.\n\n**教程类（全部过滤）**：教程、指南、攻略、入门、自学、从零开始、手把手、保姆级教程、怎么做、如何使用、操作步骤、图文教程、视频教程、科研绘图、PS教程、Illustrator 等。\n\n**企业宣传稿/软文（全部过滤）**：产能突破、全线投产、技术溢出、供应链底气、跨界营销、负面舆情、品鉴官、品牌定位等。\n\n**教育/社会活动/颁奖（全部过滤）**：十佳、颁奖仪式、表彰、职校生、职业院校、评选、杰出代表、工匠精神等。\n\nInvalid keywords: clickbait patterns, advertising, webpage navigation elements.\n\n### Deduplication (similarity.py)\n\n- Jaccard 2-gram similarity\n- Threshold: ≥90% → mark as duplicate\n- Only one version appears in output\n\n### Date Filtering\n\n- Normal: within 3 days\n- Holidays: within 7 days\n- No date → discard\n- Old URLs (year > 1 year ago) → skip\n\n## Configuration\n\n### Environment Variables\n\n| Variable | Default | Description |\n|----------|---------|-------------|\n| `NEWS_DIGEST_DB` | `news.db` | SQLite database path |\n| `NEWS_DIGEST_LLM_API_KEY` | (empty) | LLM API key for Stage 2.5 summarization |\n| `NEWS_DIGEST_LLM_BASE_URL` | (empty) | LLM API base URL |\n| `NEWS_DIGEST_LLM_MODEL` | `qwen3.6-plus` | LLM model name |\n\nIf LLM env vars are not set, Stage 2.5 is silently skipped and rule-based summaries are used instead.\n\n## Files\n\n```\nnews-digest/\n├── SKILL.md\n└── scripts/\n    └── news_digest_v2/\n        ├── __init__.py\n        ├── config.py               # DB path, websites, keywords, holidays, LLM config\n        ├── database.py             # SQLite operations\n        ├── fetcher.py              # Web scraping + smart summary extraction\n        ├── filters.py              # Content filtering logic\n        ├── formatter.py            # Output formatting + incomplete sentence handling\n        ├── init_db.py              # One-click database initialization (NEW in v1.0.1)\n        ├── quick_start.py          # One-command full pipeline (NEW in v1.0.1)\n        ├── rules_config.py         # Exclusion rules, keywords, dateline patterns\n        ├── similarity.py           # Jaccard deduplication\n        ├── stage1_fetch.py         # Stage 1 entry (fetch)\n        ├── stage2_process.py       # Stage 2 entry (dedup + keywords)\n        ├── stage2_5_llm_summary.py # Stage 2.5 (LLM batch summarization)\n        ├── stage3_output.py        # Stage 3 entry (read + format + save)\n        └── run_all_stages.py       # Full pipeline entry\n```\n\n## FAQ\n\n**Q: 安装后跑不起来？**\nA: 确保先运行了 `init_db.py` 初始化数据库。没有数据库和示例数据，后续步骤会失败。\n\n**Q: pip install 失败？**\nA: 尝试 `pip install --upgrade pip` 后再安装。如果网络问题，使用 `pip install -i https://pypi.tuna.tsinghua.edu.cn/simple requests beautifulsoup4`。\n\n**Q: 某些网站抓取失败？**\nA: 正常现象。部分网站有反爬或 SSL 问题，脚本会继续处理其他网站。不影响最终输出。\n\n**Q: 输出是空的？**\nA: 检查数据库中是否有数据。运行 `python scripts/news_digest_v2/init_db.py` 重新初始化。\n\n**Q: 如何自定义监测网站？**\nA: 通过 SQL 插入 `monitor_websites` 表，字段：name, url, selector, category, priority。\n\n**Q: 数据库会越来越大吗？**\nA: 约 30-50 条/天。建议定期清理旧数据，或删除 `news.db` 后重新初始化。\n\n## Performance Notes\n\n- ~5 minutes for full scrape (network I/O bound)\n- Some sites may fail (SSL issues, 521 errors, 404s) — pipeline continues\n- Recommended cron timeout: 900 seconds\n- **数据库是增量追加的**，不会被清空。新新闻按 URL 去重插入（`INSERT OR IGNORE`），旧新闻保留。\n- 重复新闻标记 `is_duplicate = 1`，不删除。\n\n## Changelog\n\n### v1.0.7 (2026-06-01)\n- **标题乱码修复**：\n  - `fetcher.py` 新增 `decode_response()` 函数，对已知 GBK 编码来源（人民日报海外版）强制使用 GBK 解码，从根源修复 Cyrillic 乱码\n  - `stage2_5_llm_summary.py` 新增乱码标题检测，发现乱码时提示 LLM 从正文生成准确标题\n  - `formatter.py` 新增乱码标题兜底过滤\n- **内容类型黑名单**：\n  - 新增 `TITLE_EXCLUDE_KEYWORDS`（评论丨/时评/社评/深度观察/记者观察/招聘/面试/递补/人事任免/讣告/专访等）\n  - 新增 `URL_EXCLUDE_PATTERNS`（人民网评论频道等）\n  - `formatter.py` 输出时自动跳过非硬新闻类型\n- **推荐超时**：900s→1200s（实测 LLM 总结 50 条 + 抓取总耗时 ~1150s）\n\n### v1.0.6 (2026-05-29)\n- **新闻源增至 46 个**（45 启用，1 个\"新华每日电讯\"禁用）\n- **LLM 模型修复**：默认模型从 `qwen-plus`（不存在，400 错误）改为 `qwen3.6-plus`\n- **LLM_BATCH_SIZE**：35→50，与 MAX_OUTPUT_COUNT 一致\n- **来源数显示修复**：输出头部\"来源网站\"从动态统计改为固定 46\n- **新增过滤规则**：`corporate_pr`（企业宣传稿/软文）+ `education_social`（教育/颁奖/评选）\n- **推荐超时**：600s→900s（因 LLM 总结 50 条耗时增加）\n- **飞书队列清理**：不再自动写入飞书\n\n### v1.0.5 (2026-05-28)\n- **新增新闻源**：安徽日报、人民日报海外版、新华每日电讯（42→45个启用源）\n- **MAX_OUTPUT_COUNT**：35→50条\n- **新增关键词**：十五五、标准、纲要、公报、全覆盖、创新药、芯片、测评、公共服务（12个）\n- **fetcher.py 兼容修复**：4处相对导入改为 try/except fallback，支持直接运行和包导入\n- **中国工信网超时**：10秒→30秒\n- **数据库**：新闻源从 SQLite 加载，关键词从数据库读取\n\n### v1.0.4 (2026-05-25)\n- **标题清理**：自动去除标题首尾的多余符号（如中点 `·`、空格）\n- **输出排序优化**：权威来源（人民网/新华网等）按级别升序排列，同级别按时间倒序\n- **过滤规则增强**：新增艺术展览/书画捐赠过滤词（避免非产业类文化新闻干扰）\n- **权威选文修复**：修正文章数 <35 时未排序的 Bug\n\n### v1.0.3 (2026-05-22)\n- **央广网独立解析器** (`parse_cnr`): 央广网页面标题和正文在同一 `<a>` 标签内，新增独立解析器只取 `<strong>` 作为标题，避免标题+正文混一起导致标题过长被过滤\n- **权威来源优先选文**: 阶段 2.5（LLM 批量总结）和阶段 3（输出）都应用权威来源保底策略，每个权威来源（人民网、新华网、央广网、经济日报、科技日报、科学网、中国科技网、科创版日报、中国经济网）至少入选 2 条，避免被中国经济网和中宏网等高产源淹没\n- **新华网子频道归并**: 新华能源、新华科创、新华时政、新华汽车等子频道统一归并到\"新华网\"来源\n- **社会新闻过滤增强**: 新增交通违法/行政处罚类社会新闻过滤词（罚款、行拘、拘留、交警、变造号牌等）\n- **编码检测优化**: 央广网等 GB2312 编码网站从 HTML `<meta>` 标签检测编码，提高抓取成功率\n\n### v1.0.1 (2026-05-06)\n- Added `init_db.py` for one-click database initialization with sample data\n- Added `quick_start.py` for one-command full pipeline\n- Simplified SKILL.md installation guide to 3 steps\n- Added FAQ section\n- Updated example websites to 10 mainstream Chinese news sources\n\n### v1.0.0 (2026-05-05)\n- Initial release\n\nFile v1.0.5:_meta.json\n\n{\n  \"ownerId\": \"kn7983et5m2qgha7y8gvb8dgsn84w87m\",\n  \"slug\": \"news-digest-v1\",\n  \"version\": \"1.0.5\",\n  \"publishedAt\": 1780411145375\n}\n\nFile v1.0.5:skill-card.md\n\n## Description: <br>\nAutomatically scrapes Chinese news sources, filters and deduplicates articles, and generates attributed daily news digests with optional LLM summaries. <br>\n\nThis skill is ready for commercial/non-commercial use. <br>\n\n## Publisher: <br>\n[zigu-creator](https://clawhub.ai/user/zigu-creator) <br>\n\n### License/Terms of Use: <br>\nMIT-0 <br>\n\n\n## Use Case: <br>\nDevelopers and news-monitoring users use this skill to collect Chinese news from configured sources, rank and filter articles, optionally summarize them with an LLM, and produce daily digest files with source attribution. <br>\n\n### Deployment Geography for Use: <br>\nGlobal <br>\n\n## Known Risks and Mitigations: <br>\nRisk: The skill scrapes many external news sites and some fetch paths disable TLS certificate verification. <br>\nMitigation: Run only against trusted configured sources, review fetched content before relying on it, and avoid sources that require disabled TLS verification. <br>\nRisk: Article text may be sent to a configured LLM endpoint for batch summarization. <br>\nMitigation: Set explicit NEWS_DIGEST_LLM_API_KEY, NEWS_DIGEST_LLM_BASE_URL, and NEWS_DIGEST_LLM_MODEL values only for an approved provider, or leave LLM configuration unset to use rule-based summaries. <br>\nRisk: The skill may reuse credentials from an OpenClaw configuration fallback when explicit LLM settings are not set. <br>\nMitigation: Prefer explicit NEWS_DIGEST_LLM_* environment variables and remove or review the fallback configuration before execution. <br>\nRisk: The pipeline stores scraped article data in a local SQLite database and writes digest files to the workspace and Desktop. <br>\nMitigation: Set NEWS_DIGEST_DB to an intended local path and review generated files before sharing or syncing them. <br>\n\n\n## Reference(s): <br>\n- [ClawHub skill page](https://clawhub.ai/zigu-creator/news-digest-v1) <br>\n\n\n## Skill Output: <br>\n**Output Type(s):** [Markdown, Text, Shell commands, Configuration, Files] <br>\n**Output Format:** [Markdown and plain text digest files with source names, summaries, dates, and original links] <br>\n**Output Parameters:** [1D] <br>\n**Other Properties Related to Output:** [Writes a workspace digest file, may write a Desktop text digest, and maintains a local SQLite news database.] <br>\n\n## Skill Version(s): <br>\n1.0.5 (source: server release metadata) <br>\n\n## Ethical Considerations: <br>\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment. <br>\n\nArchive v1.0.3: 18 files, 48022 bytes\n\nFiles: scripts/news_digest_v2/__init__.py (296b), scripts/news_digest_v2/config.py (6161b), scripts/news_digest_v2/database.py (12845b), scripts/news_digest_v2/fetcher.py (29295b), scripts/news_digest_v2/filters.py (3258b), scripts/news_digest_v2/formatter.py (5519b), scripts/news_digest_v2/init_db.py (6849b), scripts/news_digest_v2/quick_start.py (1966b), scripts/news_digest_v2/rules_config.py (18095b), scripts/news_digest_v2/run_all_stages.py (3685b), scripts/news_digest_v2/similarity.py (2348b), scripts/news_digest_v2/stage1_fetch.py (1761b), scripts/news_digest_v2/stage2_5_llm_summary.py (11317b), scripts/news_digest_v2/stage2_process.py (2431b), scripts/news_digest_v2/stage3_output.py (6248b), skill-card.md (2925b), SKILL.md (9904b), _meta.json (133b)\n\nFile v1.0.3:SKILL.md\n\n---\nname: news-digest\ndescription: \"Automatically scrape, process, and generate daily news digests from Chinese news sources. Covers industry dynamics, policy updates, economy, tech, energy, and pricing information. Use when: user asks for daily news summary, news digest, 每日新闻摘要, 新闻汇总, 新闻摘要, or wants to set up automated news monitoring from Chinese news websites. Outputs formatted summaries with source attribution and original links.\"\nversion: 1.0.3\n---\n\n# News Digest - 每日新闻摘要\n\nAutomated pipeline for Chinese news aggregation and digest generation.\n\n## Quick Start (3 步搞定)\n\n```bash\n# 第 1 步：安装依赖\npip install requests beautifulsoup4\n\n# 第 2 步：一键初始化（建表 + 插入示例网站 + 关键词）\npython scripts/news_digest_v2/init_db.py\n\n# 第 3 步：运行摘要\npython scripts/news_digest_v2/run_all_stages.py\n```\n\n或者一条命令全部搞定：\n```bash\npython scripts/news_digest_v2/quick_start.py\n```\n\nOutput: `.news-digest-out.md` (workspace) + `新闻摘要_YYYYMMDD_HHMMSS.txt` (desktop)\n\n## Architecture\n\n```\nStage 1:   Fetch     →  Scrape websites → Filter → Save to SQLite DB\nStage 2:   Process   →  Deduplicate (≥90% similarity) → Tag keywords\nStage 2.5: LLM       →  Batch LLM summarization (optional, requires API key)\nStage 3:   Output    →  Read LLM summaries (fallback to rule summaries) → Save to files\n```\n\n## Setup\n\n### Prerequisites\n\n- Python 3.8+ with: `requests`, `beautifulsoup4`\n- SQLite (built-in)\n\n### Initialize Database\n\nRun the init script to create tables and seed with sample data:\n\n```bash\npython scripts/news_digest_v2/init_db.py\n```\n\nThis creates:\n- Database tables (articles, monitor_websites, system_keywords, digest_output)\n- 10 sample news websites (People.cn, Xinhua, 36Kr, etc.)\n- 18 sample keywords (产业, 政策, 经济, 科技, etc.)\n\nDefault database path: `news.db` (in the skill directory).\nOverride with environment variable: `NEWS_DIGEST_DB=/your/path/news.db`\n\n### Customizing Your Sources\n\nAfter initialization, add or remove websites and keywords via SQL:\n\n```sql\n-- Add a website\nINSERT INTO monitor_websites (name, url, selector, category, priority)\nVALUES ('示例网站', 'https://example.com', 'a', '财经', 1);\n\n-- Add a keyword\nINSERT INTO system_keywords (keyword, category, weight)\nVALUES ('新能源', 'core', 5);\n```\n\n### Core Database Tables\n\n| Table | Purpose |\n|-------|---------|\n| `articles` | Scraped news articles (title, content, URL, date, keywords, duplicate flag) |\n| `monitor_websites` | Monitored websites (name, URL, CSS selector, category, enabled) |\n| `system_keywords` | Keywords for relevance scoring (core vs auxiliary, with weight) |\n| `digest_output` | LLM-generated summaries (optional) |\n\n## Usage\n\n### Full Pipeline\n\n```bash\npython scripts/news_digest_v2/run_all_stages.py\n```\n\nTakes ~5 minutes (network-bound).\n\n### One-Command Quick Start\n\n```bash\npython scripts/news_digest_v2/quick_start.py\n```\n\nRuns init + fetch + process + output in one shot.\n\n### Cron Job Example\n\n```yaml\nschedule: \"0 20 * * *\"  # Daily 20:00\npayload:\n  run: python scripts/news_digest_v2/run_all_stages.py\n  then: read .news-digest-out.md and send to messaging\ntimeout: 600  # 10 minutes\n```\n\n## Output Format\n\n```\n【来源：标题】\n摘要内容（智能选段，300字以内，包含关键数据和核心事实）\n发布时间：YYYY-MM-DD\n原文链接：http://...\n```\n\n### 摘要质量保证\n\n**不完整句子自动过滤**：\n- 摘要末尾以逗号、顿号、分号、冒号等结尾 → 回退截断到上一个句号\n- 全文没有句号（整段残缺）→ 直接丢弃，不输出\n- 截断时信息损失超过 40% → 整段放弃，宁缺毋滥\n\n**教程/指南类内容全部过滤**：\n- 标题或内容包含\"教程\"、\"指南\"、\"攻略\"、\"手把手\"、\"从零开始\"等 → 自动排除\n- 详见 `rules_config.py` 中 `social` 分类的教程关键词列表\n\n## Key Features\n\n### Smart Summary Extraction (fetcher.py → extract_brief_summary)\n\nNot simple truncation. Each paragraph is scored by:\n- **Position**: Lead paragraph +10, top-3 +5 (inverted pyramid journalism)\n- **Data density**: Numbers × 2.5\n- **Signal words**: 印发/发布/宣布/决定/完成/启动 (+2 each)\n- **Entity density**: Organizations, locations (+1 each)\n- **Completeness**: Full sentence ending +3\n\nThen filtered: removes image captions, journalist bylines, ads, subtitles, boilerplate.\n\n### 摘要后处理 (formatter.py → clean_summary)\n\n- 电头/记者署名清理（预编译正则，支持新华社、中新网、财联社等）\n- 不完整句子过滤：以逗号/顿号/分号结尾 → 回退到上一个句号\n- 全文无句号 → 丢弃（不输出残缺内容）\n\n### Filtering Rules (rules_config.py)\n\nExcluded topics: entertainment, social news, violence, crime cases, health/wellness, education, automotive consumer news, science popularization (科普类), animal/archaeology news.\n\n**教程类（全部过滤）**：教程、指南、攻略、入门、自学、从零开始、手把手、保姆级教程、怎么做、如何使用、操作步骤、图文教程、视频教程、科研绘图、PS教程、Illustrator 等。\n\nInvalid keywords: clickbait patterns, advertising, webpage navigation elements.\n\n### Deduplication (similarity.py)\n\n- Jaccard 2-gram similarity\n- Threshold: ≥90% → mark as duplicate\n- Only one version appears in output\n\n### Date Filtering\n\n- Normal: within 3 days\n- Holidays: within 7 days\n- No date → discard\n- Old URLs (year > 1 year ago) → skip\n\n## Configuration\n\n### Environment Variables\n\n| Variable | Default | Description |\n|----------|---------|-------------|\n| `NEWS_DIGEST_DB` | `news.db` | SQLite database path |\n| `NEWS_DIGEST_LLM_API_KEY` | (empty) | LLM API key for Stage 2.5 summarization |\n| `NEWS_DIGEST_LLM_BASE_URL` | (empty) | LLM API base URL |\n| `NEWS_DIGEST_LLM_MODEL` | `qwen-plus` | LLM model name |\n\nIf LLM env vars are not set, Stage 2.5 is silently skipped and rule-based summaries are used instead.\n\n## Files\n\n```\nnews-digest/\n├── SKILL.md\n└── scripts/\n    └── news_digest_v2/\n        ├── __init__.py\n        ├── config.py               # DB path, websites, keywords, holidays, LLM config\n        ├── database.py             # SQLite operations\n        ├── fetcher.py              # Web scraping + smart summary extraction\n        ├── filters.py              # Content filtering logic\n        ├── formatter.py            # Output formatting + incomplete sentence handling\n        ├── init_db.py              # One-click database initialization (NEW in v1.0.1)\n        ├── quick_start.py          # One-command full pipeline (NEW in v1.0.1)\n        ├── rules_config.py         # Exclusion rules, keywords, dateline patterns\n        ├── similarity.py           # Jaccard deduplication\n        ├── stage1_fetch.py         # Stage 1 entry (fetch)\n        ├── stage2_process.py       # Stage 2 entry (dedup + keywords)\n        ├── stage2_5_llm_summary.py # Stage 2.5 (LLM batch summarization)\n        ├── stage3_output.py        # Stage 3 entry (read + format + save)\n        └── run_all_stages.py       # Full pipeline entry\n```\n\n## FAQ\n\n**Q: 安装后跑不起来？**\nA: 确保先运行了 `init_db.py` 初始化数据库。没有数据库和示例数据，后续步骤会失败。\n\n**Q: pip install 失败？**\nA: 尝试 `pip install --upgrade pip` 后再安装。如果网络问题，使用 `pip install -i https://pypi.tuna.tsinghua.edu.cn/simple requests beautifulsoup4`。\n\n**Q: 某些网站抓取失败？**\nA: 正常现象。部分网站有反爬或 SSL 问题，脚本会继续处理其他网站。不影响最终输出。\n\n**Q: 输出是空的？**\nA: 检查数据库中是否有数据。运行 `python scripts/news_digest_v2/init_db.py` 重新初始化。\n\n**Q: 如何自定义监测网站？**\nA: 通过 SQL 插入 `monitor_websites` 表，字段：name, url, selector, category, priority。\n\n**Q: 数据库会越来越大吗？**\nA: 约 30-50 条/天。建议定期清理旧数据，或删除 `news.db` 后重新初始化。\n\n## Performance Notes\n\n- ~5 minutes for full scrape (network I/O bound)\n- Some sites may fail (SSL issues, 521 errors, 404s) — pipeline continues\n- Recommended cron timeout: 600 seconds\n- **数据库是增量追加的**，不会被清空。新新闻按 URL 去重插入（`INSERT OR IGNORE`），旧新闻保留。\n- 重复新闻标记 `is_duplicate = 1`，不删除。\n\n## Changelog\n\n### v1.0.3 (2026-05-22)\n- **央广网独立解析器** (`parse_cnr`): 央广网页面标题和正文在同一 `<a>` 标签内，新增独立解析器只取 `<strong>` 作为标题，避免标题+正文混一起导致标题过长被过滤\n- **权威来源优先选文**: 阶段 2.5（LLM 批量总结）和阶段 3（输出）都应用权威来源保底策略，每个权威来源（人民网、新华网、央广网、经济日报、科技日报、科学网、中国科技网、科创版日报、中国经济网）至少入选 2 条，避免被中国经济网和中宏网等高产源淹没\n- **新华网子频道归并**: 新华能源、新华科创、新华时政、新华汽车等子频道统一归并到\"新华网\"来源\n- **社会新闻过滤增强**: 新增交通违法/行政处罚类社会新闻过滤词（罚款、行拘、拘留、交警、变造号牌等）\n- **编码检测优化**: 央广网等 GB2312 编码网站从 HTML `<meta>` 标签检测编码，提高抓取成功率\n\n### v1.0.1 (2026-05-06)\n- Added `init_db.py` for one-click database initialization with sample data\n- Added `quick_start.py` for one-command full pipeline\n- Simplified SKILL.md installation guide to 3 steps\n- Added FAQ section\n- Updated example websites to 10 mainstream Chinese news sources\n\n### v1.0.0 (2026-05-05)\n- Initial release\n\nFile v1.0.3:_meta.json\n\n{\n  \"ownerId\": \"kn7983et5m2qgha7y8gvb8dgsn84w87m\",\n  \"slug\": \"news-digest-v1\",\n  \"version\": \"1.0.3\",\n  \"publishedAt\": 1779440378486\n}\n\nFile v1.0.3:skill-card.md\n\n## Description: <br>\nAutomatically scrape, process, and generate daily news digests from Chinese news sources, covering industry dynamics, policy updates, economy, technology, energy, and pricing information with source attribution and original links. <br>\n\nThis skill is ready for commercial/non-commercial use. <br>\n\n## Publisher: <br>\n[zigu-creator](https://clawhub.ai/user/zigu-creator) <br>\n\n### License/Terms of Use: <br>\nMIT-0 <br>\n\n\n## Use Case: <br>\nEmployees, external users, and developers use this skill to set up automated monitoring of Chinese news websites and generate concise daily digests with source links. It supports daily news summary, news digest, 每日新闻摘要, 新闻汇总, and 新闻摘要 workflows. <br>\n\n### Deployment Geography for Use: <br>\nGlobal <br>\n\n## Known Risks and Mitigations: <br>\nRisk: The skill scrapes external news websites and may encounter site failures, encoding issues, or untrusted remote content. <br>\nMitigation: Review monitored sources before running, keep selectors and source lists limited to expected news sites, and inspect generated summaries before forwarding them. <br>\nRisk: Optional LLM summarization may send news titles and snippets to the configured model endpoint using local OpenClaw credentials. <br>\nMitigation: Run without the LLM stage unless that data sharing is acceptable, or configure a trusted endpoint and credential scope before enabling summarization. <br>\nRisk: The security review notes an unsafe TLS verify=False fallback and hard-coded local database paths. <br>\nMitigation: Remove or disable the TLS fallback and set an explicit database path appropriate for the deployment environment before scheduled or unattended use. <br>\nRisk: The skill writes digest files to the workspace and Desktop. <br>\nMitigation: Run it in a controlled workspace and confirm output locations before automation so generated files do not expose sensitive monitoring interests. <br>\n\n\n## Reference(s): <br>\n- [ClawHub skill page](https://clawhub.ai/zigu-creator/news-digest-v1) <br>\n\n\n## Skill Output: <br>\n**Output Type(s):** [Text, Markdown, Shell commands, Configuration] <br>\n**Output Format:** [Markdown digest plus timestamped text file with source attribution and original links] <br>\n**Output Parameters:** [1D] <br>\n**Other Properties Related to Output:** [Writes .news-digest-out.md in the workspace and a timestamped Chinese news digest text file on the Desktop; optional LLM summarization writes summaries into a local SQLite digest table before output.] <br>\n\n## Skill Version(s): <br>\n1.0.3 (source: frontmatter and server release evidence, released 2026-05-22) <br>\n\n## Ethical Considerations: <br>\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment. <br>\n\nArchive v1.0.2: 17 files, 42091 bytes\n\nFiles: scripts/news_digest_v2/__init__.py (296b), scripts/news_digest_v2/config.py (6060b), scripts/news_digest_v2/database.py (11878b), scripts/news_digest_v2/fetcher.py (25098b), scripts/news_digest_v2/filters.py (3258b), scripts/news_digest_v2/formatter.py (5519b), scripts/news_digest_v2/init_db.py (6849b), scripts/news_digest_v2/quick_start.py (1966b), scripts/news_digest_v2/rules_config.py (17126b), scripts/news_digest_v2/run_all_stages.py (3685b), scripts/news_digest_v2/similarity.py (2348b), scripts/news_digest_v2/stage1_fetch.py (1761b), scripts/news_digest_v2/stage2_5_llm_summary.py (7968b), scripts/news_digest_v2/stage2_process.py (2431b), scripts/news_digest_v2/stage3_output.py (4057b), SKILL.md (8915b), _meta.json (133b)\n\nFile v1.0.2:SKILL.md\n\n---\nname: news-digest\ndescription: \"Automatically scrape, process, and generate daily news digests from Chinese news sources. Covers industry dynamics, policy updates, economy, tech, energy, and pricing information. Use when: user asks for daily news summary, news digest, 每日新闻摘要, 新闻汇总, 新闻摘要, or wants to set up automated news monitoring from Chinese news websites. Outputs formatted summaries with source attribution and original links.\"\nversion: 1.0.1\n---\n\n# News Digest - 每日新闻摘要\n\nAutomated pipeline for Chinese news aggregation and digest generation.\n\n## Quick Start (3 步搞定)\n\n```bash\n# 第 1 步：安装依赖\npip install requests beautifulsoup4\n\n# 第 2 步：一键初始化（建表 + 插入示例网站 + 关键词）\npython scripts/news_digest_v2/init_db.py\n\n# 第 3 步：运行摘要\npython scripts/news_digest_v2/run_all_stages.py\n```\n\n或者一条命令全部搞定：\n```bash\npython scripts/news_digest_v2/quick_start.py\n```\n\nOutput: `.news-digest-out.md` (workspace) + `新闻摘要_YYYYMMDD_HHMMSS.txt` (desktop)\n\n## Architecture\n\n```\nStage 1:   Fetch     →  Scrape websites → Filter → Save to SQLite DB\nStage 2:   Process   →  Deduplicate (≥90% similarity) → Tag keywords\nStage 2.5: LLM       →  Batch LLM summarization (optional, requires API key)\nStage 3:   Output    →  Read LLM summaries (fallback to rule summaries) → Save to files\n```\n\n## Setup\n\n### Prerequisites\n\n- Python 3.8+ with: `requests`, `beautifulsoup4`\n- SQLite (built-in)\n\n### Initialize Database\n\nRun the init script to create tables and seed with sample data:\n\n```bash\npython scripts/news_digest_v2/init_db.py\n```\n\nThis creates:\n- Database tables (articles, monitor_websites, system_keywords, digest_output)\n- 10 sample news websites (People.cn, Xinhua, 36Kr, etc.)\n- 18 sample keywords (产业, 政策, 经济, 科技, etc.)\n\nDefault database path: `news.db` (in the skill directory).\nOverride with environment variable: `NEWS_DIGEST_DB=/your/path/news.db`\n\n### Customizing Your Sources\n\nAfter initialization, add or remove websites and keywords via SQL:\n\n```sql\n-- Add a website\nINSERT INTO monitor_websites (name, url, selector, category, priority)\nVALUES ('示例网站', 'https://example.com', 'a', '财经', 1);\n\n-- Add a keyword\nINSERT INTO system_keywords (keyword, category, weight)\nVALUES ('新能源', 'core', 5);\n```\n\n### Core Database Tables\n\n| Table | Purpose |\n|-------|---------|\n| `articles` | Scraped news articles (title, content, URL, date, keywords, duplicate flag) |\n| `monitor_websites` | Monitored websites (name, URL, CSS selector, category, enabled) |\n| `system_keywords` | Keywords for relevance scoring (core vs auxiliary, with weight) |\n| `digest_output` | LLM-generated summaries (optional) |\n\n## Usage\n\n### Full Pipeline\n\n```bash\npython scripts/news_digest_v2/run_all_stages.py\n```\n\nTakes ~5 minutes (network-bound).\n\n### One-Command Quick Start\n\n```bash\npython scripts/news_digest_v2/quick_start.py\n```\n\nRuns init + fetch + process + output in one shot.\n\n### Cron Job Example\n\n```yaml\nschedule: \"0 20 * * *\"  # Daily 20:00\npayload:\n  run: python scripts/news_digest_v2/run_all_stages.py\n  then: read .news-digest-out.md and send to messaging\ntimeout: 600  # 10 minutes\n```\n\n## Output Format\n\n```\n【来源：标题】\n摘要内容（智能选段，300字以内，包含关键数据和核心事实）\n发布时间：YYYY-MM-DD\n原文链接：http://...\n```\n\n### 摘要质量保证\n\n**不完整句子自动过滤**：\n- 摘要末尾以逗号、顿号、分号、冒号等结尾 → 回退截断到上一个句号\n- 全文没有句号（整段残缺）→ 直接丢弃，不输出\n- 截断时信息损失超过 40% → 整段放弃，宁缺毋滥\n\n**教程/指南类内容全部过滤**：\n- 标题或内容包含\"教程\"、\"指南\"、\"攻略\"、\"手把手\"、\"从零开始\"等 → 自动排除\n- 详见 `rules_config.py` 中 `social` 分类的教程关键词列表\n\n## Key Features\n\n### Smart Summary Extraction (fetcher.py → extract_brief_summary)\n\nNot simple truncation. Each paragraph is scored by:\n- **Position**: Lead paragraph +10, top-3 +5 (inverted pyramid journalism)\n- **Data density**: Numbers × 2.5\n- **Signal words**: 印发/发布/宣布/决定/完成/启动 (+2 each)\n- **Entity density**: Organizations, locations (+1 each)\n- **Completeness**: Full sentence ending +3\n\nThen filtered: removes image captions, journalist bylines, ads, subtitles, boilerplate.\n\n### 摘要后处理 (formatter.py → clean_summary)\n\n- 电头/记者署名清理（预编译正则，支持新华社、中新网、财联社等）\n- 不完整句子过滤：以逗号/顿号/分号结尾 → 回退到上一个句号\n- 全文无句号 → 丢弃（不输出残缺内容）\n\n### Filtering Rules (rules_config.py)\n\nExcluded topics: entertainment, social news, violence, crime cases, health/wellness, education, automotive consumer news, science popularization (科普类), animal/archaeology news.\n\n**教程类（全部过滤）**：教程、指南、攻略、入门、自学、从零开始、手把手、保姆级教程、怎么做、如何使用、操作步骤、图文教程、视频教程、科研绘图、PS教程、Illustrator 等。\n\nInvalid keywords: clickbait patterns, advertising, webpage navigation elements.\n\n### Deduplication (similarity.py)\n\n- Jaccard 2-gram similarity\n- Threshold: ≥90% → mark as duplicate\n- Only one version appears in output\n\n### Date Filtering\n\n- Normal: within 3 days\n- Holidays: within 7 days\n- No date → discard\n- Old URLs (year > 1 year ago) → skip\n\n## Configuration\n\n### Environment Variables\n\n| Variable | Default | Description |\n|----------|---------|-------------|\n| `NEWS_DIGEST_DB` | `news.db` | SQLite database path |\n| `NEWS_DIGEST_LLM_API_KEY` | (empty) | LLM API key for Stage 2.5 summarization |\n| `NEWS_DIGEST_LLM_BASE_URL` | (empty) | LLM API base URL |\n| `NEWS_DIGEST_LLM_MODEL` | `qwen-plus` | LLM model name |\n\nIf LLM env vars are not set, Stage 2.5 is silently skipped and rule-based summaries are used instead.\n\n## Files\n\n```\nnews-digest/\n├── SKILL.md\n└── scripts/\n    └── news_digest_v2/\n        ├── __init__.py\n        ├── config.py               # DB path, websites, keywords, holidays, LLM config\n        ├── database.py             # SQLite operations\n        ├── fetcher.py              # Web scraping + smart summary extraction\n        ├── filters.py              # Content filtering logic\n        ├── formatter.py            # Output formatting + incomplete sentence handling\n        ├── init_db.py              # One-click database initialization (NEW in v1.0.1)\n        ├── quick_start.py          # One-command full pipeline (NEW in v1.0.1)\n        ├── rules_config.py         # Exclusion rules, keywords, dateline patterns\n        ├── similarity.py           # Jaccard deduplication\n        ├── stage1_fetch.py         # Stage 1 entry (fetch)\n        ├── stage2_process.py       # Stage 2 entry (dedup + keywords)\n        ├── stage2_5_llm_summary.py # Stage 2.5 (LLM batch summarization)\n        ├── stage3_output.py        # Stage 3 entry (read + format + save)\n        └── run_all_stages.py       # Full pipeline entry\n```\n\n## FAQ\n\n**Q: 安装后跑不起来？**\nA: 确保先运行了 `init_db.py` 初始化数据库。没有数据库和示例数据，后续步骤会失败。\n\n**Q: pip install 失败？**\nA: 尝试 `pip install --upgrade pip` 后再安装。如果网络问题，使用 `pip install -i https://pypi.tuna.tsinghua.edu.cn/simple requests beautifulsoup4`。\n\n**Q: 某些网站抓取失败？**\nA: 正常现象。部分网站有反爬或 SSL 问题，脚本会继续处理其他网站。不影响最终输出。\n\n**Q: 输出是空的？**\nA: 检查数据库中是否有数据。运行 `python scripts/news_digest_v2/init_db.py` 重新初始化。\n\n**Q: 如何自定义监测网站？**\nA: 通过 SQL 插入 `monitor_websites` 表，字段：name, url, selector, category, priority。\n\n**Q: 数据库会越来越大吗？**\nA: 约 30-50 条/天。建议定期清理旧数据，或删除 `news.db` 后重新初始化。\n\n## Performance Notes\n\n- ~5 minutes for full scrape (network I/O bound)\n- Some sites may fail (SSL issues, 521 errors, 404s) — pipeline continues\n- Recommended cron timeout: 600 seconds\n- **数据库是增量追加的**，不会被清空。新新闻按 URL 去重插入（`INSERT OR IGNORE`），旧新闻保留。\n- 重复新闻标记 `is_duplicate = 1`，不删除。\n\n## Changelog\n\n### v1.0.1 (2026-05-06)\n- Added `init_db.py` for one-click database initialization with sample data\n- Added `quick_start.py` for one-command full pipeline\n- Simplified SKILL.md installation guide to 3 steps\n- Added FAQ section\n- Updated example websites to 10 mainstream Chinese news sources\n\n### v1.0.0 (2026-05-05)\n- Initial release\n\nFile v1.0.2:_meta.json\n\n{\n  \"ownerId\": \"kn7983et5m2qgha7y8gvb8dgsn84w87m\",\n  \"slug\": \"news-digest-v1\",\n  \"version\": \"1.0.2\",\n  \"publishedAt\": 1778236439172\n}\n\nArchive v1.0.1: 17 files, 42217 bytes\n\nFiles: scripts/news_digest_v2/__init__.py (296b), scripts/news_digest_v2/config.py (6060b), scripts/news_digest_v2/database.py (11878b), scripts/news_digest_v2/fetcher.py (25098b), scripts/news_digest_v2/filters.py (3258b), scripts/news_digest_v2/formatter.py (5519b), scripts/news_digest_v2/init_db.py (6849b), scripts/news_digest_v2/quick_start.py (1966b), scripts/news_digest_v2/rules_config.py (17126b), scripts/news_digest_v2/run_all_stages.py (3685b), scripts/news_digest_v2/similarity.py (2348b), scripts/news_digest_v2/stage1_fetch.py (1761b), scripts/news_digest_v2/stage2_5_llm_summary.py (8773b), scripts/news_digest_v2/stage2_process.py (2431b), scripts/news_digest_v2/stage3_output.py (4057b), SKILL.md (8915b), _meta.json (133b)\n\nFile v1.0.1:SKILL.md\n\n---\nname: news-digest\ndescription: \"Automatically scrape, process, and generate daily news digests from Chinese news sources. Covers industry dynamics, policy updates, economy, tech, energy, and pricing information. Use when: user asks for daily news summary, news digest, 每日新闻摘要, 新闻汇总, 新闻摘要, or wants to set up automated news monitoring from Chinese news websites. Outputs formatted summaries with source attribution and original links.\"\nversion: 1.0.1\n---\n\n# News Digest - 每日新闻摘要\n\nAutomated pipeline for Chinese news aggregation and digest generation.\n\n## Quick Start (3 步搞定)\n\n```bash\n# 第 1 步：安装依赖\npip install requests beautifulsoup4\n\n# 第 2 步：一键初始化（建表 + 插入示例网站 + 关键词）\npython scripts/news_digest_v2/init_db.py\n\n# 第 3 步：运行摘要\npython scripts/news_digest_v2/run_all_stages.py\n```\n\n或者一条命令全部搞定：\n```bash\npython scripts/news_digest_v2/quick_start.py\n```\n\nOutput: `.news-digest-out.md` (workspace) + `新闻摘要_YYYYMMDD_HHMMSS.txt` (desktop)\n\n## Architecture\n\n```\nStage 1:   Fetch     →  Scrape websites → Filter → Save to SQLite DB\nStage 2:   Process   →  Deduplicate (≥90% similarity) → Tag keywords\nStage 2.5: LLM       →  Batch LLM summarization (optional, requires API key)\nStage 3:   Output    →  Read LLM summaries (fallback to rule summaries) → Save to files\n```\n\n## Setup\n\n### Prerequisites\n\n- Python 3.8+ with: `requests`, `beautifulsoup4`\n- SQLite (built-in)\n\n### Initialize Database\n\nRun the init script to create tables and seed with sample data:\n\n```bash\npython scripts/news_digest_v2/init_db.py\n```\n\nThis creates:\n- Database tables (articles, monitor_websites, system_keywords, digest_output)\n- 10 sample news websites (People.cn, Xinhua, 36Kr, etc.)\n- 18 sample keywords (产业, 政策, 经济, 科技, etc.)\n\nDefault database path: `news.db` (in the skill directory).\nOverride with environment variable: `NEWS_DIGEST_DB=/your/path/news.db`\n\n### Customizing Your Sources\n\nAfter initialization, add or remove websites and keywords via SQL:\n\n```sql\n-- Add a website\nINSERT INTO monitor_websites (name, url, selector, category, priority)\nVALUES ('示例网站', 'https://example.com', 'a', '财经', 1);\n\n-- Add a keyword\nINSERT INTO system_keywords (keyword, category, weight)\nVALUES ('新能源', 'core', 5);\n```\n\n### Core Database Tables\n\n| Table | Purpose |\n|-------|---------|\n| `articles` | Scraped news articles (title, content, URL, date, keywords, duplicate flag) |\n| `monitor_websites` | Monitored websites (name, URL, CSS selector, category, enabled) |\n| `system_keywords` | Keywords for relevance scoring (core vs auxiliary, with weight) |\n| `digest_output` | LLM-generated summaries (optional) |\n\n## Usage\n\n### Full Pipeline\n\n```bash\npython scripts/news_digest_v2/run_all_stages.py\n```\n\nTakes ~5 minutes (network-bound).\n\n### One-Command Quick Start\n\n```bash\npython scripts/news_digest_v2/quick_start.py\n```\n\nRuns init + fetch + process + output in one shot.\n\n### Cron Job Example\n\n```yaml\nschedule: \"0 20 * * *\"  # Daily 20:00\npayload:\n  run: python scripts/news_digest_v2/run_all_stages.py\n  then: read .news-digest-out.md and send to messaging\ntimeout: 600  # 10 minutes\n```\n\n## Output Format\n\n```\n【来源：标题】\n摘要内容（智能选段，300字以内，包含关键数据和核心事实）\n发布时间：YYYY-MM-DD\n原文链接：http://...\n```\n\n### 摘要质量保证\n\n**不完整句子自动过滤**：\n- 摘要末尾以逗号、顿号、分号、冒号等结尾 → 回退截断到上一个句号\n- 全文没有句号（整段残缺）→ 直接丢弃，不输出\n- 截断时信息损失超过 40% → 整段放弃，宁缺毋滥\n\n**教程/指南类内容全部过滤**：\n- 标题或内容包含\"教程\"、\"指南\"、\"攻略\"、\"手把手\"、\"从零开始\"等 → 自动排除\n- 详见 `rules_config.py` 中 `social` 分类的教程关键词列表\n\n## Key Features\n\n### Smart Summary Extraction (fetcher.py → extract_brief_summary)\n\nNot simple truncation. Each paragraph is scored by:\n- **Position**: Lead paragraph +10, top-3 +5 (inverted pyramid journalism)\n- **Data density**: Numbers × 2.5\n- **Signal words**: 印发/发布/宣布/决定/完成/启动 (+2 each)\n- **Entity density**: Organizations, locations (+1 each)\n- **Completeness**: Full sentence ending +3\n\nThen filtered: removes image captions, journalist bylines, ads, subtitles, boilerplate.\n\n### 摘要后处理 (formatter.py → clean_summary)\n\n- 电头/记者署名清理（预编译正则，支持新华社、中新网、财联社等）\n- 不完整句子过滤：以逗号/顿号/分号结尾 → 回退到上一个句号\n- 全文无句号 → 丢弃（不输出残缺内容）\n\n### Filtering Rules (rules_config.py)\n\nExcluded topics: entertainment, social news, violence, crime cases, health/wellness, education, automotive consumer news, science popularization (科普类), animal/archaeology news.\n\n**教程类（全部过滤）**：教程、指南、攻略、入门、自学、从零开始、手把手、保姆级教程、怎么做、如何使用、操作步骤、图文教程、视频教程、科研绘图、PS教程、Illustrator 等。\n\nInvalid keywords: clickbait patterns, advertising, webpage navigation elements.\n\n### Deduplication (similarity.py)\n\n- Jaccard 2-gram similarity\n- Threshold: ≥90% → mark as duplicate\n- Only one version appears in output\n\n### Date Filtering\n\n- Normal: within 3 days\n- Holidays: within 7 days\n- No date → discard\n- Old URLs (year > 1 year ago) → skip\n\n## Configuration\n\n### Environment Variables\n\n| Variable | Default | Description |\n|----------|---------|-------------|\n| `NEWS_DIGEST_DB` | `news.db` | SQLite database path |\n| `NEWS_DIGEST_LLM_API_KEY` | (empty) | LLM API key for Stage 2.5 summarization |\n| `NEWS_DIGEST_LLM_BASE_URL` | (empty) | LLM API base URL |\n| `NEWS_DIGEST_LLM_MODEL` | `qwen-plus` | LLM model name |\n\nIf LLM env vars are not set, Stage 2.5 is silently skipped and rule-based summaries are used instead.\n\n## Files\n\n```\nnews-digest/\n├── SKILL.md\n└── scripts/\n    └── news_digest_v2/\n        ├── __init__.py\n        ├── config.py               # DB path, websites, keywords, holidays, LLM config\n        ├── database.py             # SQLite operations\n        ├── fetcher.py              # Web scraping + smart summary extraction\n        ├── filters.py              # Content filtering logic\n        ├── formatter.py            # Output formatting + incomplete sentence handling\n        ├── init_db.py              # One-click database initialization (NEW in v1.0.1)\n        ├── quick_start.py          # One-command full pipeline (NEW in v1.0.1)\n        ├── rules_config.py         # Exclusion rules, keywords, dateline patterns\n        ├── similarity.py           # Jaccard deduplication\n        ├── stage1_fetch.py         # Stage 1 entry (fetch)\n        ├── stage2_process.py       # Stage 2 entry (dedup + keywords)\n        ├── stage2_5_llm_summary.py # Stage 2.5 (LLM batch summarization)\n        ├── stage3_output.py        # Stage 3 entry (read + format + save)\n        └── run_all_stages.py       # Full pipeline entry\n```\n\n## FAQ\n\n**Q: 安装后跑不起来？**\nA: 确保先运行了 `init_db.py` 初始化数据库。没有数据库和示例数据，后续步骤会失败。\n\n**Q: pip install 失败？**\nA: 尝试 `pip install --upgrade pip` 后再安装。如果网络问题，使用 `pip install -i https://pypi.tuna.tsinghua.edu.cn/simple requests beautifulsoup4`。\n\n**Q: 某些网站抓取失败？**\nA: 正常现象。部分网站有反爬或 SSL 问题，脚本会继续处理其他网站。不影响最终输出。\n\n**Q: 输出是空的？**\nA: 检查数据库中是否有数据。运行 `python scripts/news_digest_v2/init_db.py` 重新初始化。\n\n**Q: 如何自定义监测网站？**\nA: 通过 SQL 插入 `monitor_websites` 表，字段：name, url, selector, category, priority。\n\n**Q: 数据库会越来越大吗？**\nA: 约 30-50 条/天。建议定期清理旧数据，或删除 `news.db` 后重新初始化。\n\n## Performance Notes\n\n- ~5 minutes for full scrape (network I/O bound)\n- Some sites may fail (SSL issues, 521 errors, 404s) — pipeline continues\n- Recommended cron timeout: 600 seconds\n- **数据库是增量追加的**，不会被清空。新新闻按 URL 去重插入（`INSERT OR IGNORE`），旧新闻保留。\n- 重复新闻标记 `is_duplicate = 1`，不删除。\n\n## Changelog\n\n### v1.0.1 (2026-05-06)\n- Added `init_db.py` for one-click database initialization with sample data\n- Added `quick_start.py` for one-command full pipeline\n- Simplified SKILL.md installation guide to 3 steps\n- Added FAQ section\n- Updated example websites to 10 mainstream Chinese news sources\n\n### v1.0.0 (2026-05-05)\n- Initial release\n\nFile v1.0.1:_meta.json\n\n{\n  \"ownerId\": \"kn7983et5m2qgha7y8gvb8dgsn84w87m\",\n  \"slug\": \"news-digest-v1\",\n  \"version\": \"1.0.1\",\n  \"publishedAt\": 1778075186266\n}\n\nArchive v1.0.0: 15 files, 38224 bytes\n\nFiles: scripts/news_digest_v2/__init__.py (296b), scripts/news_digest_v2/config.py (6060b), scripts/news_digest_v2/database.py (11878b), scripts/news_digest_v2/fetcher.py (25098b), scripts/news_digest_v2/filters.py (3258b), scripts/news_digest_v2/formatter.py (5519b), scripts/news_digest_v2/rules_config.py (17126b), scripts/news_digest_v2/run_all_stages.py (3685b), scripts/news_digest_v2/similarity.py (2348b), scripts/news_digest_v2/stage1_fetch.py (1761b), scripts/news_digest_v2/stage2_5_llm_summary.py (8773b), scripts/news_digest_v2/stage2_process.py (2431b), scripts/news_digest_v2/stage3_output.py (4057b), SKILL.md (7517b), _meta.json (133b)\n\nFile v1.0.0:SKILL.md\n\n---\nname: news-digest\ndescription: \"Automatically scrape, process, and generate daily news digests from 42 Chinese news sources. Covers industry dynamics, policy updates, economy, tech, energy, and pricing information. Use when: user asks for daily news summary, news digest, 每日新闻摘要, 新闻汇总, 新闻摘要, or wants to set up automated news monitoring from Chinese news websites. Outputs formatted summaries with source attribution and original links.\"\n---\n\n# News Digest - 每日新闻摘要\n\nAutomated 3-stage pipeline for Chinese news aggregation and digest generation.\n\n## Quick Start\n\n```bash\npython scripts/news_digest_v2/run_all_stages.py\n```\n\nOutput: `.news-digest-out.md` (workspace) + `新闻摘要_YYYYMMDD_HHMMSS.txt` (desktop)\n\n## Architecture\n\n```\nStage 1:   Fetch     →  Scrape 42 websites → Filter → Save to SQLite DB\nStage 2:   Process   →  Deduplicate (≥90% similarity) → Tag keywords\nStage 2.5: LLM       →  Batch LLM summarization (optional, requires API key)\nStage 3:   Output    →  Read LLM summaries (fallback to rule summaries) → Save to files\n```\n\n## Setup\n\n### Prerequisites\n\n- Python 3.8+ with: `requests`, `beautifulsoup4`\n- SQLite (built-in)\n\n### Initialize Database\n\nThe database stores articles and configuration. Default path: `news_digest_v2/news.db` (relative to scripts directory).\n\nOverride with environment variable: `NEWS_DIGEST_DB=/your/path/news.db`\n\nThen seed the database with monitored websites and system keywords using SQL insertion into `monitor_websites` and `system_keywords` tables.\n\n### Core Database Tables\n\n| Table | Purpose |\n|-------|---------|\n| `articles` | Scraped news articles (title, content, URL, date, keywords, duplicate flag) |\n| `monitor_websites` | 42 monitored websites (name, URL, CSS selector, category, enabled) |\n| `system_keywords` | Keywords for relevance scoring (core vs auxiliary, with weight) |\n\n## Usage\n\n### Full Pipeline\n\n```bash\npython scripts/news_digest_v2/run_all_stages.py\n```\n\nTakes ~5 minutes (network-bound, 42 websites).\n\n### Entry Point (PowerShell wrapper)\n\nFor OpenClaw or automated integration, create a wrapper script that:\n1. Runs the pipeline\n2. Reads the output file\n3. Sends to your preferred messaging platform\n\n### Cron Job Example\n\n```yaml\nschedule: \"0 20 * * *\"  # Daily 20:00\npayload:\n  run: python scripts/news_digest_v2/run_all_stages.py\n  then: read .news-digest-out.md and send to messaging\ntimeout: 600  # 10 minutes\n```\n\n## Output Format\n\n```\n【来源：标题】\n摘要内容（智能选段，300字以内，包含关键数据和核心事实）\n发布时间：YYYY-MM-DD\n原文链接：http://...\n```\n\n### 摘要质量保证\n\n**不完整句子自动过滤**：\n- 摘要末尾以逗号、顿号、分号、冒号等结尾 → 回退截断到上一个句号\n- 全文没有句号（整段残缺）→ 直接丢弃，不输出\n- 截断时信息损失超过 40% → 整段放弃，宁缺毋滥\n\n**教程/指南类内容全部过滤**：\n- 标题或内容包含\"教程\"、\"指南\"、\"攻略\"、\"手把手\"、\"从零开始\"等 → 自动排除\n- 科研绘图/PS教程/Illustrator教程 → 自动排除\n- 详见 `rules_config.py` 中 `social` 分类的教程关键词列表\n\n## Key Features\n\n### Smart Summary Extraction (fetcher.py → extract_brief_summary)\n\nNot simple truncation. Each paragraph is scored by:\n- **Position**: Lead paragraph +10, top-3 +5 (inverted pyramid journalism)\n- **Data density**: Numbers × 1.5\n- **Signal words**: 印发/发布/宣布/决定/完成/启动 (+2 each)\n- **Entity density**: Organizations, locations (+1 each)\n- **Completeness**: Full sentence ending +3\n\nThen filtered: removes image captions, journalist bylines, ads, subtitles, boilerplate.\n\n**截断保护**：截断时信息损失 >40% → 整段放弃。\n\n### 摘要后处理 (formatter.py → clean_summary)\n\n- 电头/记者署名清理（预编译正则，支持新华社、中新网、财联社等）\n- 不完整句子过滤：以逗号/顿号/分号结尾 → 回退到上一个句号\n- 全文无句号 → 丢弃（不输出残缺内容）\n\n### Filtering Rules (rules_config.py)\n\nExcluded topics: entertainment, social news, violence, crime cases, health/wellness, education, automotive consumer news, science popularization (科普类), animal/archaeology news.\n\n**教程类（全部过滤）**：教程、指南、攻略、入门、自学、从零开始、手把手、保姆级教程、怎么做、如何使用、操作步骤、图文教程、视频教程、科研绘图、PS教程、Illustrator、AI教程、钢笔工具、高斯模糊、路径查找器等。\n\nInvalid keywords: clickbait patterns, advertising, webpage navigation elements.\n\nSee `scripts/news_digest_v2/rules_config.py` for full lists.\n\n### Deduplication (similarity.py)\n\n- Jaccard similarity on keyword sets\n- Threshold: ≥90% → mark as duplicate\n- Only one version appears in output\n\n### Date Filtering\n\n- Normal: within 3 days\n- Holidays: within 7 days\n- No date → discard\n- Old URLs (year > 1 year ago) → skip\n\n## Configuration\n\n### Environment Variables\n\n| Variable | Default | Description |\n|----------|---------|-------------|\n| `NEWS_DIGEST_DB` | `news_digest_v2/news.db` | SQLite database path |\n| `NEWS_DIGEST_LLM_API_KEY` | (empty) | LLM API key for Stage 2.5 summarization |\n| `NEWS_DIGEST_LLM_BASE_URL` | (empty) | LLM API base URL |\n| `NEWS_DIGEST_LLM_MODEL` | `qwen-plus` | LLM model name |\n\nIf LLM env vars are not set, Stage 2.5 is silently skipped and rule-based summaries are used instead.\n\n### Add/Remove Websites\n\nEdit `monitor_websites` table:\n\n```sql\nINSERT INTO monitor_websites (name, url, selector, category, enabled)\nVALUES ('示例网站', 'https://example.com', 'a', '财经', 1);\n```\n\n### Customize Keywords\n\nEdit `system_keywords` table:\n\n```sql\nINSERT INTO system_keywords (keyword, category, weight, enabled)\nVALUES ('新能源', 'core', 5, 1);\n```\n\n### Adjust Output\n\nIn `config.py`:\n- `MAX_OUTPUT_COUNT = 35` (max articles per digest)\n- `SIMILARITY_THRESHOLD = 0.90`\n\n## Files\n\n```\nnews-digest/\n├── SKILL.md\n└── scripts/\n    └── news_digest_v2/\n        ├── __init__.py\n        ├── config.py              # DB path, websites, keywords, holidays, LLM config\n        ├── database.py            # SQLite operations\n        ├── fetcher.py             # Web scraping + smart summary extraction\n        ├── filters.py             # Content filtering logic\n        ├── formatter.py           # Output formatting + incomplete sentence handling\n        ├── rules_config.py        # Exclusion rules, keywords, dateline patterns\n        ├── similarity.py          # Jaccard deduplication\n        ├── stage1_fetch.py        # Stage 1 entry (fetch)\n        ├── stage2_process.py      # Stage 2 entry (dedup + keywords)\n        ├── stage2_5_llm_summary.py # Stage 2.5 (LLM batch summarization)\n        ├── stage3_output.py       # Stage 3 entry (read + format + save)\n        └── run_all_stages.py      # Full pipeline entry\n```\n\n## Performance Notes\n\n- ~5 minutes for full 42-website scrape (network I/O bound)\n- Some sites may fail (SSL issues, 521 errors, 404s) — pipeline continues\n- Recommended cron timeout: 600 seconds\n- **数据库是增量追加的**，不会被清空。新新闻按 URL 去重插入（`INSERT OR IGNORE`），旧新闻保留。\n- 重复新闻标记 `is_duplicate = 1`，不删除。\n- 数据库增长约 30-50 条/天，建议定期清理（可选）。\n\nFile v1.0.0:_meta.json\n\n{\n  \"ownerId\": \"kn7983et5m2qgha7y8gvb8dgsn84w87m\",\n  \"slug\": \"news-digest-v1\",\n  \"version\": \"1.0.0\",\n  \"publishedAt\": 1777963665258\n}","readmeExcerpt":"Skill: 每日新闻搜索与智能摘要 Owner: zigu-creator Summary: Automatically scrape, process, and generate daily news digests from Chinese news sources. Covers industry dynamics, policy updates, economy, tech, energy, an... Tags: latest:1.0.12 Version history: v1.0.12 | 2026-06-09T02:40:26.589Z | user v1.0.12 (2026-06-09) — Filtering Rules Expansion & Sync Fix Added 3 new exclusion categories to rules_config.py: cultural events/pro","codeSnippets":[],"executableExamples":[{"language":"bash","snippet":"pip install requests beautifulsoup4\n2. 初始化数据库\nbash\nCopy\n# 创建数据库和表结构\npython -c \"\nimport sqlite3, os\ndb = sqlite3.connect('news.db')\ndb.execute('''CREATE TABLE IF NOT EXISTS articles (\n    id INTEGER PRIMARY KEY AUTOINCREMENT,\n    title TEXT NOT NULL, source TEXT NOT NULL,\n    publish_date TEXT NOT NULL, summary TEXT,\n    content TEXT, url TEXT UNIQUE NOT NULL,\n    keywords TEXT, is_duplicate INTEGER DEFAULT 0,\n    similarity_score REAL DEFAULT 0,\n    created_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP\n)''')\ndb.execute('''CREATE TABLE IF NOT EXISTS monitor_websites (\n    id INTEGER PRIMARY KEY AUTOINCREMENT,\n    name TEXT UNIQUE NOT NULL, url TEXT NOT NULL,\n    selector TEXT DEFAULT 'a', category TEXT,\n    priority INTEGER DEFAULT 3, enabled INTEGER DEFAULT 1\n)''')\ndb.execute('''CREATE TABLE IF NOT EXISTS system_keywords (\n    id INTEGER PRIMARY KEY AUTOINCREMENT,\n    keyword TEXT UNIQUE NOT NULL,\n    category TEXT, weight INTEGER DEFAULT 1,\n    enabled INTEGER DEFAULT 1\n)''')\ndb.execute('CREATE INDEX IF NOT EXISTS idx_publish_date ON articles(publish_date)')\ndb.execute('CREATE INDEX IF NOT EXISTS idx_keywords ON articles(keywords)')\ndb.commit()\ndb.close()\nprint('OK: news.db created')\n\"\n3. 添加监测网站（示例）\nbash\nCopy\npython -c \"\nimport sqlite3\ndb = sqlite3.connect('news.db')\nsites = [\n    ('中国经济网', 'http://www.ce.cn/', 'a', '财经', 1),\n    # 添加你想监控的网站...\n]\nfor name, url, sel, cat, pri in sites:\n    db.execute('INSERT OR IGNORE INTO monitor_websites (name, url, selector, category, priority) VALUES (?,?,?,?,?)',\n               (name, url, sel, cat, pri))\ndb.commit()\ndb.close()\nprint(f'OK: added {len(sites)} websites')\n\"\n4. 添加关键词（示例）\nbash\nCopy\npython -c \"\nimport sqlite3\ndb = sqlite3.connect('news.db')\nkws = [\n    ('市场', 'auxiliary', 2), ('企业', 'auxiliary', 2),\n]\nfor kw, cat, w in kws:\n    db.execute('INSERT OR IGNORE INTO system_keywords (keyword, category, weight) VALUES (?,?,?)',\n               (kw, cat, w))\ndb.commit()\ndb.close()\nprint(f'OK: added {len(kws)} keywords')\n\"\n5. 运行\nbash"},{"language":"text","snippet":"或者一条命令全部搞定："},{"language":"text","snippet":"Output: `.news-digest-out.md` (workspace) + `新闻摘要_YYYYMMDD_HHMMSS.txt` (desktop)\n\n## Architecture"},{"language":"text","snippet":"## Setup\n\n### Prerequisites\n\n- Python 3.8+ with: `requests`, `beautifulsoup4`\n- SQLite (built-in)\n\n### Initialize Database\n\nRun the init script to create tables and seed with sample data:"},{"language":"text","snippet":"This creates:\n- Database tables (articles, monitor_websites, system_keywords, digest_output)\n- 10 sample news websites (People.cn, Xinhua, 36Kr, etc.)\n- 18 sample keywords (产业, 政策, 经济, 科技, etc.)\n\nDefault database path: `news.db` (in the skill directory).\nOverride with environment variable: `NEWS_DIGEST_DB=/your/path/news.db`\n\n### Customizing Your Sources\n\nAfter initialization, add or remove websites and keywords via SQL:"},{"language":"text","snippet":"### Core Database Tables\n\n| Table | Purpose |\n|-------|---------|\n| `articles` | Scraped news articles (title, content, URL, date, keywords, duplicate flag) |\n| `monitor_websites` | Monitored websites (name, URL, CSS selector, category, enabled) |\n| `system_keywords` | Keywords for relevance scoring (core vs auxiliary, with weight) |\n| `digest_output` | LLM-generated summaries (optional) |\n\n## Usage\n\n### Full Pipeline"}],"parameters":null,"dependencies":[],"permissions":[],"extractedFiles":[{"path":"SKILL.md","content":"---\nname: news-digest\ndescription: \"Automatically scrape, process, and generate daily news digests from Chinese news sources. Covers industry dynamics, policy updates, economy, tech, energy, and pricing information. Use when: user asks for daily news summary, news digest, 每日新闻摘要, 新闻汇总, 新闻摘要, or wants to set up automated news monitoring from Chinese news websites. Outputs formatted summaries with source attribution and original links.\"\nversion: 1.0.12\n---\n\n# News Digest - 每日新闻摘要\n\nAutomated pipeline for Chinese news aggregation and digest generation.\n\n## Quick Start (3 步搞定)\n\n```bash\n# 第 1 步：安装依赖\npip install requests beautifulsoup4\n\n# 第 2 步：一键初始化（建表 + 插入示例网站 + 关键词）\npython scripts/news_digest_v2/init_db.py\n\n# 第 3 步：运行摘要\npython scripts/news_digest_v2/run_all_stages.py\n```\n\n或者一条命令全部搞定：\n```bash\npython scripts/news_digest_v2/quick_start.py\n```\n\nOutput: `.news-digest-out.md` (workspace) + `新闻摘要_YYYYMMDD_HHMMSS.txt` (desktop)\n\n## Architecture\n\n```\nStage 1:   Fetch     →  Scrape websites → Filter → Save to SQLite DB\nStage 2:   Process   →  Deduplicate (≥90% similarity) → Tag keywords\nStage 2.5: LLM       →  Batch LLM summarization (optional, requires API key)\nStage 3:   Output    →  Read LLM summaries (fallback to rule summaries) → Save to files\n```\n\n## Setup\n\n### Prerequisites\n\n- Python 3.8+ with: `requests`, `beautifulsoup4`\n- SQLite (built-in)\n\n### Initialize Database\n\nRun the init script to create tables and seed with sample data:\n\n```bash\npython scripts/news_digest_v2/init_db.py\n```\n\nThis creates:\n- Database tables (articles, monitor_websites, system_keywords, digest_output)\n- 10 sample news websites (People.cn, Xinhua, 36Kr, etc.)\n- 18 sample keywords (产业, 政策, 经济, 科技, etc.)\n\nDefault database path: `news.db` (in the skill directory).\nOverride with environment variable: `NEWS_DIGEST_DB=/your/path/news.db`\n\n### Customizing Your Sources\n\nAfter initialization, add or remove websites and keywords via SQL:\n\n```sql\n-- Add a website\nINSERT INTO monitor_websites (name, url, selector, category, priority)\nVALUES ('示例网站', 'https://example.com', 'a', '财经', 1);\n\n-- Add a keyword\nINSERT INTO system_keywords (keyword, category, weight)\nVALUES ('新能源', 'core', 5);\n```\n\n### Core Database Tables\n\n| Table | Purpose |\n|-------|---------|\n| `articles` | Scraped news articles (title, content, URL, date, keywords, duplicate flag) |\n| `monitor_websites` | Monitored websites (name, URL, CSS selector, category, enabled) |\n| `system_keywords` | Keywords for relevance scoring (core vs auxiliary, with weight) |\n| `digest_output` | LLM-generated summaries (optional) |\n\n## Usage\n\n### Full Pipeline\n\n```bash\npython scripts/news_digest_v2/run_all_stages.py\n```\n\nTakes ~13 minutes (network + LLM bound).\n\n### One-Command Quick Start\n\n```bash\npython scripts/news_digest_v2/quick_start.py\n```\n\nRuns init + fetch + process + output in one shot.\n\n### Cron Job Example\n\n```yaml\nschedule: \"0 20 * * *\"  # Daily 20:00\npayload:\n  run: python scripts/news_digest_v2/run_all_stages.py\n  then: read .news-d"},{"path":"_meta.json","content":"{\n  \"ownerId\": \"kn7983et5m2qgha7y8gvb8dgsn84w87m\",\n  \"slug\": \"news-digest-v1\",\n  \"version\": \"1.0.12\",\n  \"publishedAt\": 1780972826589\n}"},{"path":"skill-card.md","content":"## Description:\n\nAutomatically scrapes Chinese news sources, filters and deduplicates articles, and generates daily summaries with source attribution and original links.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[zigu-creator](https://clawhub.ai/user/zigu-creator)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers and operators use this skill to set up a local Chinese news monitoring workflow that fetches news, filters low-relevance content, optionally summarizes with an LLM, and produces a daily digest for review or distribution.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: The skill can make broad network requests to configured news sites and custom article links.\n\nMitigation: Run it only with reviewed monitor_websites entries and trusted article sources before using the full pipeline.\n\nRisk: The LLM stage can reuse an OpenClaw API key from local configuration when dedicated NEWS_DIGEST_LLM credentials are not set.\n\nMitigation: Set dedicated, limited NEWS_DIGEST_LLM_API_KEY and NEWS_DIGEST_LLM_BASE_URL values, or remove the OpenClaw config fallback before execution.\n\nRisk: The skill writes a local SQLite database plus digest files in the workspace and Desktop.\n\nMitigation: Run it in a workspace where these writes are expected, and review generated digest content before sharing it.\n\n## Reference(s):\n\n- [ClawHub skill page](https://clawhub.ai/zigu-creator/skills/news-digest-v1)\n- [Publisher profile](https://clawhub.ai/user/zigu-creator)\n\n## Skill Output:\n\n**Output Type(s):** [text, markdown, shell commands, configuration, guidance]\n\n**Output Format:** [Markdown and plain text digest files with source titles, summaries, dates, and original links; setup guidance includes Markdown with bash and SQL snippets.]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [Writes .news-digest-out.md in the workspace and a timestamped Chinese text digest on Desktop; optional LLM summarization uses configured API credentials.]\n\n## Skill Version(s):\n\n1.0.12 (source: SKILL.md frontmatter and ClawHub release evidence, released 2026-06-09)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment."}],"languages":[],"docsSourceLabel":"CLAWHUB","editorialOverview":"Automatically scrape, process, and generate daily news digests from Chinese news sources. Covers industry dynamics, policy updates, economy, tech, energy, an... Skill: 每日新闻搜索与智能摘要 Owner: zigu-creator Summary: Automatically scrape, process, and generate daily news digests from Chinese news sources. Covers industry dynamics, policy updates, economy, tech, energy, an... Tags: latest:1.0.12 Version history: v1.0.12 | 2026-06-09T02:40:26.589Z | user v1.0.12 (2026-06-09) — Filtering Rules Expansion & Sync Fix Added 3 new exclusion categories to rules_config.py: cultural events/pro","editorialQuality":{"score":100,"threshold":65,"status":"ready","wordCount":899,"uniquenessScore":52,"reasons":[]}},"media":{"evidence":{"source":"no-media","verified":false,"confidence":"low","updatedAt":"2026-10-11T06:42:50.372Z","emptyReason":"No screenshots, media assets, or demo links are available."},"primaryImageUrl":null,"mediaAssetCount":0,"assets":[],"demoUrl":null},"ownerResources":{"evidence":{"source":"unclaimed","verified":false,"confidence":"low","updatedAt":"2026-10-11T06:42:50.372Z","emptyReason":"This page has not been claimed by the agent owner."},"hasCustomPage":false,"customPageUpdatedAt":null,"customLinks":[],"structuredLinks":{"docsUrl":null,"demoUrl":null,"supportUrl":null,"pricingUrl":null,"statusUrl":null},"customPage":null},"relatedAgents":{"evidence":{"source":"protocol-neighbors","verified":false,"confidence":"medium","updatedAt":"2026-10-11T10:51:53.537Z","emptyReason":null},"items":[{"id":"8ebccd8e-3863-4187-8355-c3f14e1f9edf","entityType":"agent","canonicalPath":"/agent/iofficeai-aionui","slug":"iofficeai-aionui","name":"AionUi","description":"Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!","url":"https://github.com/iOfficeAI/AionUi","homepage":"https://www.aionui.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-10-09T19:11:12.944Z","createdAt":"2026-02-25T03:38:16.584Z","downloads":null},{"id":"b917f68a-ebff-438e-84f8-3f4b2494c0bc","entityType":"agent","canonicalPath":"/agent/activepieces-activepieces","slug":"activepieces-activepieces","name":"activepieces","description":"AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents","url":"https://github.com/activepieces/activepieces","homepage":"https://www.activepieces.com","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-15T02:22:12.426Z","createdAt":"2026-02-25T03:38:12.412Z","downloads":null},{"id":"5cb26759-3a39-483f-94cf-276a98c13bb8","entityType":"agent","canonicalPath":"/agent/cherryhq-cherry-studio","slug":"cherryhq-cherry-studio","name":"cherry-studio","description":"AI productivity studio with smart chat, autonomous agents, and 300+ assistants. Unified access to frontier LLMs","url":"https://github.com/CherryHQ/cherry-studio","homepage":"https://cherry-ai.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-11T14:38:40.986Z","createdAt":"2026-02-25T03:38:19.379Z","downloads":null},{"id":"6f6582d0-5d76-4f0f-b81d-86520247950b","entityType":"agent","canonicalPath":"/agent/copilotkit-copilotkit","slug":"copilotkit-copilotkit","name":"CopilotKit","description":"The Frontend for Agents & Generative UI. React + Angular","url":"https://github.com/CopilotKit/CopilotKit","homepage":"https://docs.copilotkit.ai","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-03-25T09:50:57.846Z","createdAt":"2026-02-25T03:39:14.617Z","downloads":null}],"links":{"hub":"/agent","source":"/agent/source/clawhub","protocols":[{"label":"OpenClaw","href":"/agent/protocol/openclew"}]}}}