agentCLAWHUBUnverified

boss-resume-crawler

从 Boss 直聘批量爬取职位详情(含 security_id、职位描述),支持 PUA 薪资解码和增量去重。 Skill: boss-resume-crawler Owner: iichaner Summary: 从 Boss 直聘批量爬取职位详情(含 security_id、职位描述),支持 PUA 薪资解码和增量去重。 Tags: latest:0.1.0 Version history: v0.1.0 | 2026-08-18T10:45:58.143Z | auto Initial release of boss-resume-crawler. - Supports batch crawling of Boss直聘 job details, including security_id and job description - Implements PUA salary decoding and incremental deduplication - Provides strict dependency checks (Pyth

OpenClaw

Rank

62

Safety

84

Downloads

1.4k

Updated

Oct 10, 2026

Version

0.1.0

Source

CLAWHUB

About

What it does, and when to use it.

Capability contract not published. No trust telemetry is available yet. 1.4K downloads reported by the source. Last updated 10/10/2026.

Avoid when

  • Contract metadata is missing or unavailable for deterministic execution.

Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing

Public facts

Every fact links back to the source it came from.

Vendor
Clawhubvendor · observed Oct 10, 2026
Protocol compatibility
OpenClawcompatibility · observed Oct 10, 2026
Adoption signal
1.4K downloadsadoption · observed Oct 10, 2026
Latest release
0.1.0release · observed Aug 18, 2026
Handshake status
UNKNOWNsecurity

Install and run

Setup complexity: low.

clawhub skill install s17evhyc3z1a82bvpm10c418p9849pg7:boss-resume-crawler
  1. Setup complexity is LOW. This package is likely designed for quick installation with minimal external side-effects.
  2. Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.

Contract: missing

curl -s "https://www.xpersona.co/api/v1/agents/clawhub-iichaner-boss-resume-crawler/snapshot"

Documentation

CLAWHUB

26,297 characters of source documentation, loaded on request.

Extracted files

5 files captured from the source.

SKILL.md

---
name: boss-resume-crawler
description: "从 Boss 直聘批量爬取职位详情(含 security_id、职位描述),支持 PUA 薪资解码和增量去重。"
metadata:
  {
    "openclaw":
      {
        "emoji": "🕷️",
        "requires": { "bins": ["curl", "python3"] },
      },
  }
read_when:
  - 用户要求爬取 Boss 直聘职位
  - 用户提到"爬取"、"抓取"、"JD 数据"、"Boss 直聘"
  - 用户要求批量获取职位详情或 security_id
allowed-tools: Bash,Read,Write,exec
---

# Boss直聘职位爬取

## 快速验证(环境 OK 后立即跑通)

```bash
# 一条命令验证:爬取 5 条职位,输出到 /tmp
python3 scripts/boss_extract_cdp.py --max-scroll 3 --max-jobs 5 --output /tmp/boss_test
```

验证通过 → 正式爬取。失败 → 检下方依赖和 CDP 连接。

---

## 性能基准(实测)

| 指标 | 数值 |
|------|------|
| 单页滚动加载 | ~2 秒/次 |
| 详情页提取 | ~22 秒/条(含 20 秒等待 + 随机波动) |
| 100 条职位总耗时 | ~40 分钟 |
| 500 条职位总耗时 | ~3 小时 |

> 详情页耗时主要由等待时间决定(20 秒/条),这是 Boss 直聘客户端渲染的硬限制。

---

## 首次使用:依赖检查(必须)

**在执行任何爬取操作之前,必须先检查以下依赖是否就绪。缺失时提示用户安装。**

### 检查脚本

```bash
# 1. Python3
python3 --version 2>/dev/null || echo "❌ 未安装 Python3 → https://www.python.org/downloads/"

# 2. websocket-client(Python 库)
python3 -c "import websocket" 2>/dev/null || echo "❌ 缺少 websocket-client → pip3 install websocket-client"

# 3. CDP 连接(CloakBrowser 是否启动)
curl -s http://localhost:9222/json >/dev/null 2>&1 || echo "❌ CDP 未连接,请先启动 CloakBrowser(见下方说明)"
```

### CloakBrowser 启动方法

CloakBrowser 是一个反检测 Chromium 浏览器,用于绕过 Boss 直聘的自动化检测。

**安装:**
```bash
# 安装依赖
npm install cloakbrowser playwright-core

# 下载 Chromium(需要代理)
export https_proxy=http://127.0.0.1:7890  # 根据你的代理配置
curl -L --max-time 600 -o /tmp/cloakbrowser-darwin-x64.tar.gz <下载链接>
tar -xzf /tmp/cloakbrowser-darwin-x64.tar.gz -C ~/.cache/cloakbrowser/

# 设置环境变量
export CLOAKBROWSER_BINARY_PATH=~/.cache/cloakbrowser/Chromium.app/Contents/MacOS/Chromium
```

> CloakBrowser 项目地址:https://github.com/nickspaargaren/cloakbrowser

**启动(有头模式,必须):**
```bash
open ~/.cache/cloakbrowser/Chromium.app --args \
  --remote-debugging-port=9222 \
  "--remote-allow-origins=*" \
  --user-data-dir=<你的浏览器数据目录> \
  "<Boss直聘列表页URL>"
```

> ⚠️ Boss 直聘会检测 headless 模式,**必须使用有头模式**(能看到浏览器窗口)。

### 依赖就绪标志

所有 ✅ 后方可执行爬取:
- [ ] Python3 可用
- [ ] websocket-client 已安装
- [ ] CDP 连接正常(`curl -s http://localhost:9222/json` 返回页面列表)
- [ ] 用户已登录 Boss 直聘(见下方登录检查)

---

## 输入要求

- **必须提供** Boss 直聘列表页 URL(含 `zhipin.com/web/geek/jobs`)
- 未提供 URL 时必须主动询问,不要自行构造

## 登录状态检查(必须在 Phase 1 之前执行)

打开列表页后,**首先检查登录状态**,未登录则暂停等待人类操作:

```bash
snapshot=$(agent-browser --cdp 9222 snapshot -i --timeout 8000 2>/dev/null)
if echo "$snapshot" | grep -qE "登录/注册|立即登录|登录"; then
  echo "⚠️ 未登录状态,请手动扫码登录"
  echo "登录完成后告知我,我再继续"
fi
```

**判断逻辑:**
- ❌ 出现「登录/注册」「立即登录」「我要找工作」等链接 → 未登录
- ✅ 出现用户名或用户头像链接 → 已登录

**未登录时的处理:**
1. 暂停所有爬取工作
2. 提示用户:「页面显示未登录,请在浏览器中扫码登录,完成后告知我」
3. 等待用户明确说「已登录」或「继续」后,再执行后续 Phase

---

## 执行流程

### Phase 1:列表页滚动加载
使用 CDP `Input.dispatchMouseEvent mouseWheel` 模拟真实鼠标滚轮(agent-browser scroll 无效)。最多 100 次滚动,随机等待 1.5-3.5 秒,连续 3 次数量不变则停止。启动浏览器需添加 `"--remote-allow-origins=*"` 参数。
详见 [references/sop.md](references/sop.md) SOP-1。

### Phase 2:职位列表提取
通过 CDP 执行 JavaScript 从 `.job-card-wrap` 提取职位基础信息,薪资需 PUA 解码

README.md

# Boss直聘职位爬取 Skill

> 🕷️ 从 Boss 直聘批量提取职位详情(含 security_id、职位描述),支持 PUA 薪资解码和增量去重。

一个 [OpenClaw](https://github.com/openclaw/openclaw) Skill,帮助 AI Agent 或用户自动化爬取 Boss 直聘的职位数据。

---

## 功能特性

- 🔄 **智能滚动加载** — 使用 CDP 模拟真实鼠标滚轮,绕过 Boss 直聘反爬检测
- 🔐 **PUA 薪资解码** — 自动将 Boss 直聘的特殊 Unicode 字符还原为真实薪资数字
- 📋 **详情页深度提取** — 逐条打开详情页,提取 security_id 和完整职位描述
- 💾 **即时写入存储** — 每提取 1 条立即写入 CSV,中断不丢数据
- ✅ **质量报告** — 每次爬取后输出字段完整率和错误统计
- 🛡️ **两层容错** — 失败自动重试(指数退避),仍失败则跳过并记录错误日志
- 🕶️ **反爬优化** — 独立连接、随机等待、tab 即关,降低被检测风险

## 前置依赖

| 依赖 | 必需 | 说明 |
|------|------|------|
| Python 3.6+ | ✅ | 运行脚本 |
| websocket-client | ✅ | Python CDP 通信库 |
| CloakBrowser | ✅ | 反检测 Chromium 浏览器 |

## 安装

### 1. 克隆仓库

```bash
git clone https://github.com/iichaner/boss-resume-crawler.git
cd boss-resume-crawler
```

### 2. 安装 Python 依赖

```bash
pip3 install websocket-client
```

### 3. 安装 CloakBrowser

```bash
npm install cloakbrowser playwright-core

# 下载 Chromium(需要代理访问 GitHub)
export https_proxy=http://127.0.0.1:7890
curl -L --max-time 600 -o /tmp/cloakbrowser-darwin-x64.tar.gz <下载链接>
tar -xzf /tmp/cloakbrowser-darwin-x64.tar.gz -C ~/.cache/cloakbrowser/
```

> 📖 [CloakBrowser 文档](https://github.com/nickspaargaren/cloakbrowser)

## 快速开始

```bash
# 1. 启动 CloakBrowser(有头模式,必须)
open ~/.cache/cloakbrowser/Chromium.app --args \
  --remote-debugging-port=9222 \
  "--remote-allow-origins=*" \
  --user-data-dir=/tmp/chrome-cdp-profile \
  "https://www.zhipin.com/web/geek/jobs?query=总经理助理&city=101010100"

# 2. 在浏览器中扫码登录 Boss 直聘

# 3. 快速验证(爬取 5 条)
python3 scripts/boss_extract_cdp.py --max-scroll 3 --max-jobs 5 --output /tmp/boss_test

# 4. 正式爬取
python3 scripts/boss_extract_cdp.py --output ~/Desktop/jobs --max-scroll 50
```

## 脚本说明

| 脚本 | 用途 | 推荐度 |
|------|------|--------|
| `scripts/boss_extract_cdp.py` | **默认脚本**:纯 CDP 模式,反爬优化 | ⭐⭐⭐ 推荐 |
| `scripts/boss_extract_pure.py` | 旧版:纯 CDP 模式(无反爬优化) | ⭐ 备用 |
| `scripts/boss_extract_final.py` | agent-browser 模式 | ⭐⭐ 可选 |

### 默认脚本参数

```bash
python3 scripts/boss_extract_cdp.py [OPTIONS]

--output DIR       输出目录(默认当前目录)
--max-scroll N     最大滚动次数(默认 100)
--max-jobs N       最大爬取条数(默认全部)
--base-wait N      详情页基础等待秒数(默认 20)
```

### 示例

```bash
# 爬取全部职位,输出到桌面
python3 scripts/boss_extract_cdp.py --output ~/Desktop/jobs

# 只爬 20 条,滚动 30 次
python3 scripts/boss_extract_cdp.py --max-jobs 20 --max-scroll 30

# 网络较慢时增加等待
python3 scripts/boss_extract_cdp.py --base-wait 25
```

## 反爬设计

| 措施 | 实现 | 说明 |
|------|------|------|
| 有头模式 | CloakBrowser | headless 会被拦截 |
| 反检测浏览器 | CloakBrowser | 绕过自动化检测 |
| 真实滚轮模拟 | CDP `Input.dispatchMouseEvent` | 普通 scroll 不触发加载 |
| 随机滚动间隔 | 1.5-3.5 秒随机 | 模拟人类阅读节奏 |
| 随机详情页等待 | 20 秒 + 0-3 秒波动 | 降低请求规律性 |
| 独立 WebSocket 连接 | 每次操作新建,用完即关 | 防长连接被检测 |
| Tab 即关 | 提取后立即关闭 | 防 tab 堆积 |
| 指数退避重试 | +5 秒/次,最多 3 次 | 避免频繁重试触发风控 |

## 输出格式

### CSV 字段

| 字段名 | 说明 | 示例 |
|--------|------|------|
| 职位名称 | 职位标题 | 总经理助理 |
| 薪资 | PUA 解码后的真实薪资 | 15-20K·13薪 |
| 经验要求 | 工作经验 | 5-10年 |
| 学历要求 | 最低学历 | 本科 |
| 公司名称 | 招聘公司 | 某某科技有限公司 |
| 城市 | 工作城市 | 上海 |
| 区域 |

_meta.json

{
  "ownerId": "kn76dgfmrcmcfw91f0m92tj4rd84805k",
  "slug": "boss-resume-crawler",
  "version": "0.1.0",
  "publishedAt": 1787049958143
}

references/data-spec.md

# 数据字段规格

## 必要字段(缺失即失败)

| 字段名 | 类型 | 校验标准 | 提取来源 |
|--------|------|---------|---------|
| job_id | Text | 非空,>=20 字符 | 列表页 href 正则 `/job_detail/(.+?)\.html` |
| security_id | Text | 非空,>=30 字符 | 详情页 script 标签 |
| 薪资 | Text | 包含 "K" | 列表页 `.job-salary`,需 PUA 解码 |
| 职位描述 | Text | 非空,>=100 字符 | 详情页 body.innerText |
| 公司名称 | Text | 非空,>=2 字符 | 列表页 `.boss-name` link text |

## 普通字段

| 字段名 | 提取来源 |
|--------|---------|
| 职位名称 | 列表页 `.job-name` link text |
| 经验要求 | 列表页 `.tag-list li`,匹配 `\d+-\d+年` |
| 学历要求 | 列表页 `.tag-list li`,匹配 `本科\|大专\|硕士\|博士\|学历不限` |
| 城市 | 列表页 `.company-location`,按 `·` 拆分取第一段 |
| 区域 | 列表页 `.company-location`,按 `·` 拆分取剩余 |
| 招聘者 | 详情页(如有) |
| 招聘者职位 | 详情页(如有) |
| 创建日期 | 爬取时间,格式 `YYYY-MM-DD HH:MM` |

## CSV 字段顺序

```
职位名称,薪资,经验要求,学历要求,公司名称,城市,区域,job_id,security_id,职位描述,创建日期
```

## PUA 薪资解码

Boss 直聘使用 PUA Unicode 字符隐藏真实薪资数字:

```
0xe031 → 0    0xe036 → 5
0xe032 → 1    0xe037 → 6
0xe033 → 2    0xe038 → 7
0xe034 → 3    0xe039 → 8
0xe035 → 4    0xe03a → 9
```

Python 解码函数(已内嵌于脚本):

```python
PUA_MAP = {
    0xe031: '0', 0xe032: '1', 0xe033: '2', 0xe034: '3', 0xe035: '4',
    0xe036: '5', 0xe037: '6', 0xe038: '7', 0xe039: '8', 0xe03a: '9'
}
def decode_pua(text):
    if not text: return text
    return ''.join(PUA_MAP.get(ord(c), c) for c in text)
```

## 页面选择器(当前有效)

| 元素 | 选择器 | 备注 |
|------|--------|------|
| 职位名称 | `.job-name` | ✅ |
| 薪资 | `.job-salary` | ✅ 需 PUA 解码 |
| 公司名称 | `.boss-name` | ✅ |
| 地区 | `.company-location` | ✅ |
| 经验/学历 | `.tag-list li` | ✅ |

**已失效选择器**:`.salary`, `.job-title`, `.company-name a`, `.area`

references/error-handling.md

# 错误处理

## 两层 Fallback 机制

### 第一层:重试

- **触发**:页面加载慢、网络波动、反爬延迟
- **策略**:增加 50% 等待时间,重试当前步骤
- **上限**:同一位置最多重试 3 次,超过则升级到第二层

### 第二层:修复并继续

- **触发**:一层重试仍失败
- **策略**:
  1. 跳过当前错误职位
  2. 记录错误到 `temp/error_log.csv`
  3. 继续处理下一个职位

### 错误日志格式

```csv
title,job_id,error,timestamp
```

---

## 当前需注意的陷阱

- **详情页等待必须 20 秒以上**:不足会导致职位描述全部为空(0 字节)
- **滚动必须模拟人类节奏**:快速连续滚动会导致列表只加载部分(如 50 条只加载 15 条)
- **CSV 必须用 `'a'` 追加模式**:`'w'` 模式会覆盖历史数据

---

## 已修复问题归档

| # | 问题 | 修复方案 | 状态 |
|---|------|---------|------|
| 0 | CSV 覆盖历史数据 | 改用 `'a'` 追加 + job_id 去重 | ✅ |
| 1 | 列表页上限 49 条 | 连续 5 次滚动不变则停止 | ✅ |
| 2 | 职位描述提取不完整 | 从“职位描述”标题开始,多结束标记 | ✅ |
| 3 | job_id 提取失败 | 改用 `.+?\.html` 正则 | ✅ |
| 4 | 公司名称解析错误 | 直接从 link 元素获取 | ✅ |
| 5 | security_id 提取为空 | 从详情页 script 标签正则提取 | ✅ |
| 6 | 职位描述全部为空 | 等待增至 20 秒 + readyState 检测 | ✅ |
| 7 | 快速滚动加载不全 | 人类滚动策略:随机等待 + 停留 | ✅ |
| 8 | WebSocket 长连接超时 | 每次 CDP 操作新建独立连接,用完即关 | ✅ |
| 9 | 进程中断数据全丢 | 每提取 1 条立即追加写入 CSV | ✅ |
| 10 | 详情页 tab 堆积 | 提取后立即 close_tab | ✅ |
| 11 | 反爬检测(固定间隔) | 所有等待时间加随机波动 | ✅ |
Github ReposUpdated 21h agoRank 70

AionUi

Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!

MCPOPENCLAW
Github ReposUpdated 6mo agoRank 70

activepieces

AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents

OPENCLAW
Github ReposUpdated 6mo agoRank 70

cherry-studio

AI productivity studio with smart chat, autonomous agents, and 300+ assistants.

MCPOPENCLAW
Github ReposUpdated 7mo agoRank 70

CopilotKit

The Frontend for Agents & Generative UI. React + Angular

OPENCLAW

Machine-readable data

The same record, as JSON, for agents and crawlers.

{
  "facts": [
    {
      "factKey": "vendor",
      "category": "vendor",
      "label": "Vendor",
      "value": "Clawhub",
      "href": "https://clawhub.ai/iichaner/skills/boss-resume-crawler",
      "sourceUrl": "https://clawhub.ai/iichaner/skills/boss-resume-crawler",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-10T14:22:16.495Z",
      "isPublic": true
    },
    {
      "factKey": "protocols",
      "category": "compatibility",
      "label": "Protocol compatibility",
      "value": "OpenClaw",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-iichaner-boss-resume-crawler/contract",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-iichaner-boss-resume-crawler/contract",
      "sourceType": "contract",
      "confidence": "medium",
      "observedAt": "2026-10-10T14:22:16.495Z",
      "isPublic": true
    },
    {
      "factKey": "traction",
      "category": "adoption",
      "label": "Adoption signal",
      "value": "1.4K downloads",
      "href": "https://clawhub.ai/iichaner/boss-resume-crawler",
      "sourceUrl": "https://clawhub.ai/iichaner/boss-resume-crawler",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-10T14:22:16.495Z",
      "isPublic": true
    },
    {
      "factKey": "latest_release",
      "category": "release",
      "label": "Latest release",
      "value": "0.1.0",
      "href": "https://clawhub.ai/iichaner/boss-resume-crawler",
      "sourceUrl": "https://clawhub.ai/iichaner/boss-resume-crawler",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-08-18T10:45:58.143Z",
      "isPublic": true
    },
    {
      "factKey": "handshake_status",
      "category": "security",
      "label": "Handshake status",
      "value": "UNKNOWN",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-iichaner-boss-resume-crawler/trust",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-iichaner-boss-resume-crawler/trust",
      "sourceType": "trust",
      "confidence": "medium",
      "observedAt": null,
      "isPublic": true
    }
  ],
  "events": [
    {
      "eventType": "release",
      "title": "Release 0.1.0",
      "description": "Initial release of boss-resume-crawler. - Supports batch crawling of Boss直聘 job details, including security_id and job description - Implements PUA salary decoding and incremental deduplication - Provides strict dependency checks (Python3, websocket-client, CDP connection via CloakBrowser) - Enforces manual login verification before crawling begins - Features resilient crawling logic: randomized delays, per-job CSV append, error logging and automatic retries - Includes detailed usage instructions, performance benchmarks, error handling, and anti-scraping precautions",
      "href": "https://clawhub.ai/iichaner/boss-resume-crawler",
      "sourceUrl": "https://clawhub.ai/iichaner/boss-resume-crawler",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-08-18T10:45:58.143Z",
      "isPublic": true
    }
  ]
}

Record generated Oct 10, 2026.

Sponsored

Ads related to boss-resume-crawler and adjacent AI workflows.