amazon-scraper
Containerized Amazon.com scraper (Docker + Playwright) for BSR/new-releases/movers rankings, keyword search results, and product detail pages. Requires a paid ISP/residential proxy. Use when the request is specifically about Amazon product or marketplace data: 亚马逊/Amazon, ASIN, BSR, Best Sellers, 畅销榜, 新品榜, 飙升榜, 选品, 竞品分析, 类目分析, listing 分析, 月销量 (bought in past month), 评分分布, 评论分析, Amazon 关键词搜索结果, Amazon 产品详情. Do NOT use for general-purpose web scraping just because the user said 爬取/抓取/采集/scrape/crawl. Every run spends metered proxy bandwidth, and the generic mode exists only as a fallback for pages related to an Amazon task. For unrelated sites prefer web_search/web_fetch, or ask first.
Rank
62
Safety
84
Downloads
4.6k
Updated
Oct 9, 2026
Version
4.0.1
Source
CLAWHUB
About
What it does, and when to use it.
Capability contract not published. No trust telemetry is available yet. 4.6K downloads reported by the source. Last updated 10/9/2026.
Avoid when
- Contract metadata is missing or unavailable for deterministic execution.
Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing
Public facts
Every fact links back to the source it came from.
- Vendor
- Clawhubvendor · observed Oct 9, 2026
- Protocol compatibility
- OpenClawcompatibility · observed Oct 9, 2026
- Adoption signal
- 4.6K downloadsadoption · observed Oct 9, 2026
- Latest release
- 4.0.1release · observed Oct 2, 2026
- Handshake status
- UNKNOWNsecurity
Install and run
Setup complexity: low.
clawhub skill install s171dacsa9hkrd446pra6ma8sd83q3hc:amazon-scraper- Install using `clawhub skill install s171dacsa9hkrd446pra6ma8sd83q3hc:amazon-scraper` in an isolated environment before connecting it to live workloads.
- No published capability contract is available yet, so validate auth and request/response behavior manually.
- Review the upstream CLAWHUB listing at https://clawhub.ai/jiafar/amazon-scraper before using production credentials.
Contract: missing
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-jiafar-amazon-scraper/snapshot"
Documentation
CLAWHUB
144,108 characters of source documentation, loaded on request.
Extracted files
5 files captured from the source.
SKILL.md
---
name: amazon-scraper
description: >
Containerized Amazon.com scraper (Docker + Playwright) for BSR/new-releases/movers rankings,
keyword search results, and product detail pages. Requires a paid ISP/residential proxy.
Use when the request is specifically about Amazon product or marketplace data:
亚马逊/Amazon, ASIN, BSR, Best Sellers, 畅销榜, 新品榜, 飙升榜,
选品, 竞品分析, 类目分析, listing 分析, 月销量 (bought in past month),
评分分布, 评论分析, Amazon 关键词搜索结果, Amazon 产品详情.
Do NOT use for general-purpose web scraping just because the user said
爬取/抓取/采集/scrape/crawl. Every run spends metered proxy bandwidth, and the
generic mode exists only as a fallback for pages related to an Amazon task.
For unrelated sites prefer web_search/web_fetch, or ask first.
metadata:
openclaw:
requires:
bins:
- docker
---
# Amazon Scraper
Docker 容器化爬虫,Playwright Chromium。不使用 stealth 插件。支持亚马逊榜单/搜索/详情及通用动态页。
## 第 0 步(开跑前必须先做,不过就停)
先确认出口,再碰亚马逊:
```bash
export AMAZON_PROXIES="http://USER:PASS@HOST:PORT" # 从密码管理器取,不要贴进对话
curl -s -x "$AMAZON_PROXIES" http://api.ipify.org
```
返回的纯文本必须等于你配置的代理 host。不是这个 IP、超时、407、403,都停,不要开爬,也不要重建后直接跑。
- **凭证不进仓库、不进镜像、不进命令回显。** 用 `-e AMAZON_PROXIES`(值从环境取,不要写在命令行里)或挂载 `config/proxies.json`。该文件已在 `.gitignore` / `.dockerignore` 里。
- 这一步只证明代理层通。裸 curl 打亚马逊拿到 500/202,不算第 0 步失败。
- 第 0 步过了,才允许 `docker run amazon-scraper`。
## 系统要求
- **Docker Engine 20.10+**(必须已安装并运行)
- **磁盘空间**:~2GB(镜像 + Playwright 浏览器二进制文件)
- **内存**:建议 2GB+(Playwright 运行时需要)
## 快速开始
首次使用:在 skill 目录下执行一键构建脚本:
```bash
bash scripts/setup.sh
```
脚本会自动完成:构建 `amazon-scraper` 镜像 + 创建 `~/scrapes` 输出目录。
## 模式选择规则
### 1. Amazon模式 (`amazon_handler.js`)
**自动触发条件:** URL包含 `amazon.com`,或用户提到亚马逊/Amazon/ASIN/BSR/选品/竞品/畅销榜/类目分析等关键词
根据URL自动识别页面类型:
| URL特征 | 页面类型 | 可获取字段 |
|---|---|---|
| `/gp/bestsellers/` | 畅销榜 | rank, title, asin, price, rating, reviews, image, url |
| `/zg/new-releases/` | 新品榜 | 同上 |
| `/zg/movers-and-shakers/` | 飙升榜 | 同上 |
| `/s?k=` 或 `/s/` | 搜索结果 | title, asin, price, rating, reviews, image, url, **boughtPastMonth**, sponsored |
| `/dp/` 或 `/gp/product/` | 产品详情 | title, asin, price, rating, reviews, brand, bsr, **boughtPastMonth**, **seller**, dateFirstAvailable, category, bullets, details |
**⚠️ 重要规则:**
- **Best Sellers 页面没有月销量(boughtPastMonth)数据** — 亚马逊不在榜单页显示此信息
- **要获取月销量,必须用搜索页(`/s?k=关键词`)或产品详情页(`/dp/ASIN`)**
- 如果用户同时需要排名+月销量,建议:先爬 Best Sellers 拿排名,再用搜索页补月销
- **BSR URL 必须使用 `/gp/bestsellers/`**,`/zgbs/` 会返回 Page Not Found
- **BSR 单 URL 只能拿 60 个产品**(2 页限制)。拿 Top 100 用搜索页 `--pages 5` 或多子类目 BSR 合并。详见 `references/bsr-top100-strategy.md`
- **评论(rating/title/body/date/helpful/verified)不在这张表里**:需要登录态 + CDP,走 `scripts/scrape_reviews.py`(`/portal/customer-reviews/ASIN` 瀑布流)。先读 `references/reviews-strategy.md` 顶部的账号/合规代价说明
- 视觉化选品 fallback:把 `image` URL 喂给 `vision_analyze`,让它"看图识物"。参考 `references/product-form-fusion.md`
```bash
# 畅销榜(有排名,无月销,最多 60 个)
docker run --rm amazon-scraper node assets/amazon_handler.js "https://www.amazon.com/gp/bestsellers/_meta.json
{
"ownerId": "kn7ewmmms6dthpk632rzrc05v981ah4f",
"slug": "amazon-scraper",
"version": "4.0.1",
"publishedAt": 1790926130735
}references/batch-parallel-scraping.md
# Batch Parallel Scraping Pattern
**当前不适用。** 活代理只有 1 条(出口 IP 和端口见你本地的 `config/proxies.json`,不要写进文档)。代码把并发卡成 `min(请求值, 任务数, proxies.length)`,所以 `--concurrency 5` 仍是 1。`-e AMAZON_PROXIES` 在 `proxies.json` 非空时不生效。下面的 5 容器 × Oxylabs 8001–8005 是旧方案,现在执行会让 5 个浏览器打同一个 IP。
大批量在只有 1 条 IP 时:一个容器,`--asins`,`--output` 必须配 `-v 宿主机目录:/data`,超过约 10 分钟改后台。先用 `http://api.ipify.org` 确认出口 IP 等于你配置的代理 host,再跑亚马逊。裸 curl 拿到 500/202 不算爬虫失败,以 handler 的 JSON 为准。
旧方案(要有 5 个不同出口才能用,且必须挂载覆盖 `/app/config/proxies.json`,不能靠环境变量):
Safe high-throughput detail page scraping using multiple Docker containers with dedicated proxy ports.
## Problem
Single container `--asins "100_asins" --concurrency 5` is slow (~20 min for 100 ASINs). But naive parallelism (5 containers × concurrency 5 = 25 total concurrent requests) causes ISP proxy soft-blocks — 76% of ASINs return all-null data.
## Safe Pattern: 5 Containers × 2-3 Concurrency
```bash
PROXY_USER="user-XXX"
PROXY_PASS="XXX"
# Split 100 ASINs into 5 batches of 20
# Each batch → its own Docker container with a dedicated proxy port
for i in 0 1 2 3 4; do
PORT=$((8001 + i))
BATCH="<20 ASINs comma-separated>"
docker run --rm -v ~/scrapes:/data \
-e AMAZON_PROXIES="http://${PROXY_USER}:${PROXY_PASS}@isp.oxylabs.io:${PORT}" \
amazon-scraper node assets/amazon_handler.js \
--asins "$BATCH" --concurrency 2 \
--output "batch_${i}.json" \
> "batch_${i}.log" 2>&1 &
done
wait
```
## Key Parameters
| Parameter | Safe Value | Risk Value | Notes |
|---|---|---|---|
| Containers | 5 | >8 | One per proxy port (8001-8005) |
| Concurrency per container | 2-3 | ≥5 | Each proxy port handles 2-3 concurrent connections |
| Total concurrent | 10-15 | ≥25 | >15 triggers Amazon soft-blocks on ISP proxies |
| ASINs per batch | 20 | >30 | Bigger batches = longer single-container runtime |
## Performance
| Mode | 100 ASINs | Success Rate |
|---|---|---|
| Single container, concurrency 5 | ~20 min | ~90% |
| 5 containers × concurrency 3 | ~5 min | ~25% (too aggressive) |
| 5 containers × concurrency 2 | ~5-7 min | ~25% (still aggressive on retry) |
| 5 containers × concurrency 2, then retry failed with concurrency 1 | ~10 min | ~50% cumulative |
## Retry Strategy for Failed ASINs
After the first pass, collect ASINs that returned all-null (proxy soft-blocked), then retry with lower concurrency:
```python
# Collect failed ASINs
for p in all_products:
if not (p.get('title') or p.get('brand')):
failed_asins.append(p['asin'])
```
Retry with `--concurrency 1` or `--concurrency 2` on the same 5-container pattern. Typical improvement: +10-15% success on retry.
## Merging Results
After all batches + retries complete, merge by ASIN — prefer rows with actual data (title/brand not null):
```python
detail_map = {}
for prefix in ['batch_', 'retry_']:
for i in range(5):
d = json.load(open(f'{prefix}{i}.json'))
for p in d['products']:
if p.get('title') or p.get('brand'): # has real data
references/bsr-top100-strategy.md
# Amazon BSR Top 100 抓取策略
## 核心限制(必读)
**`/gp/bestsellers/` URL 只能拿到 2 页 = 60 个产品。** Amazon 官方限制。`?pg=3` 会返回 "Page Not Found"。
这意味着:
- 拿 Top 30 ✅ 直接 `/gp/bestsellers/{category}`
- 拿 Top 50/60 ✅ `/gp/bestsellers/{category} --pages 2`
- **拿 Top 100 ❌ 单个 BSR URL 不行**
## 三种 Top 100 策略
### 策略 A:搜索页替代(最快、推荐)⭐
```bash
docker run --rm -v /tmp/top100:/data amazon-scraper \
node assets/amazon_handler.js \
"https://www.amazon.com/s?k=cable+management" \
--pages 5 \
--output /data/cm.json
```
- 一次调用,~3 分钟
- 5 页 = 100+ 个产品(去重后约 60-80 个独立 ASIN)
- **数据是搜索算法排序的,不是 BSR 严格排名**(混了广告位)
- 对"市场分析"够用
**适合**:快速拿数据做品类分析、价格带分析、视觉调研。
### 策略 B:多子类目 BSR 合并
BSR 父类目下钻到 5 个子类目,每个拿 Top 30 = 150 个产品:
```python
# 主类目 electronics 没有子节点
# 子类目节点 ID(在 URL 里能看到)
subcats = [
"electronics/172541", # Audio & Video
"electronics/281407", # Computers & Accessories
"electronics/2407745011", # Wearable Technology
"electronics/13896617011", # Computer & Accessories
"electronics/3024167031", # Cell Phones
]
# 拼 URL: https://www.amazon.com/gp/bestsellers/{subcat}
```
每个子 BSR 限 2 页 = 30 个。**5 × 30 = 150,去重 ~120 个**。
**适合**:要做严格的"畅销榜"分析(不被广告位污染)。
### 策略 C:单 session 串行(最稳但最慢)
把 BSR Top 30 拿到 ASIN,逐个爬详情,3 个代理并发:
```bash
# /tmp/top100.sh
while read asin; do
docker run --rm amazon-scraper node assets/amazon_handler.js \
"https://www.amazon.com/dp/${asin}" \
--output "/tmp/details/${asin}.json" 2>/dev/null
sleep 3
done < /tmp/asins.txt
```
**耗时**:30 个 ASIN × 15-25 秒 = 7-12 分钟(单容器,1 代理)
**适合**:要拿详情做品牌/BSR/详情页分析。
## 提速方案
| 方案 | 速度 | 限制 |
|---|---|---|
| 单容器 `--pages 5` | 1x | skill 本身 |
| 多 Docker 容器并发 | Nx(N=容器数) | 需 N 个不同代理 |
| 改 handler.js 加并发 | 内部可控 | 需重新 build 镜像 |
**多容器并发的现实约束**:你的代理数 = 最大并发数。
- 3 个 DDC IP → 最多 3 个并发 → 3x 提速
- 5 个 ISP 端口 → 最多 5 个并发 → 5x 提速
skill 自带轮询:handler 内部已支持 1 个容器内多代理轮询 + 故障切换,无需额外配置。
## 速度与限流
- 单个 session:~30 秒/页(含 15 秒冷启动 + 3-5 秒 waitFor)
- Amazon 限流:30-60 请求/小时/同一 IP 是安全线,超了会触发 503
- 跑 Top 100 一次消耗约 5-10 个"请求单位"
## 输出文件路径
⚠️ **坑**:`--output` 是**容器内路径**。要保存到主机必须挂载卷:
```bash
# 错:文件在容器里
docker run --rm amazon-scraper node .../amazon_handler.js "URL" --output /tmp/x.json
# 对:挂载 /data
docker run --rm -v /tmp/results:/data amazon-scraper node .../amazon_handler.js \
"URL" --output /data/x.json
# 文件在主机的 /tmp/results/x.json
```
代码里 `--output` 的路径会拼到 `/data/` 前缀下。references/cdp-fallback-strategy.md
# CDP Fallback 爬取策略
## 问题场景
批量爬取 Amazon 详情页时(`--asins` 模式),Docker 代理方案在以下情况会被 Amazon 软拦截:
- 总并发 ≥ 15(5 容器 × concurrency 3)
- 单代理端口并发 ≥ 3
- 短时间内同 IP 大量请求
软拦截表现:`status: SUCCESS` 但 `title/brand/seller/bsr` 全部 null,页面未渲染。
## 解决方案:Chrome CDP 直连
**核心思路**:绕过 Docker 代理,用 VPS 本地 Chrome 浏览器(CDP 协议)逐个串行爬取,配合 2-4 秒随机延迟。
**实测结果(2026-07-04)**:
- Docker 代理方案:75 个失败 ASIN,成功率 0%(全部被拦)
- CDP 直连方案:75 个失败 ASIN,**成功率 100%**(0 失败)
- 耗时:75 个 ASIN × ~5 秒/个 = 约 6 分钟
## 前置条件
1. Chrome 已启动并监听 CDP:
```bash
export DISPLAY=:99
Xvfb :99 -screen 0 1920x1080x24 &>/dev/null &
/opt/google/chrome/chrome --disable-gpu --no-first-run \
--no-default-browser-check \
--remote-debugging-port=9222 \
--remote-debugging-address=127.0.0.1 \
--user-data-dir="$HOME/.cache/amazon-scraper-chrome" \
"https://www.amazon.com"
```
> ### ⚠️ 不要加 `--remote-allow-origins=*`
>
> CDP **没有任何认证机制**:谁能连上 9222,谁就完全控制这个浏览器 —— 读 cookie、以登录用户身份下单、改收货地址、导出 session。本方案的前提恰恰是这个 Chrome 带着真实 Amazon 登录态,所以它的调试口就等于账号凭证。
>
> - `--remote-allow-origins=*` 关掉了 WebSocket 的 Origin 校验,于是**你在这个浏览器里打开的任意网页**都能连上 9222 接管它。只在确实遇到 Origin 报错时,针对具体来源写白名单,不要用 `*`。
> - 必须显式 `--remote-debugging-address=127.0.0.1`。绑到 `0.0.0.0` 的 VPS 等于开了一个公网无密码浏览器后门;即使绑回环,也要确认没有端口转发或 docker 规则把它暴露出去(`ss -lntp | grep 9222` 自查)。
> - 不要用 `--no-sandbox`。VPS 上常以 root 跑 Chrome,关掉沙箱意味着一个渲染器漏洞就能拿到 root。需要在容器里跑就改用非 root 用户 + `--user-ns` 之类的方案。
> - profile 不要放 `/tmp`(全局可写,其他本地用户可读你的 cookie)。放 `$HOME/.cache/...` 并保持 `chmod 700`。
> - 用完把这个 Chrome 关掉,别长期挂着一个带登录态的调试口。
本目录的两个 CDP 脚本都会**自己新开一个标签页**并在结束时关掉,不会劫持你正在用的标签页。
2. Python 依赖:
```bash
pip install websocket-client
```
## 使用方法
```bash
# 方式1:直接传 ASIN 列表
python3 ~/.openclaw/skills/amazon-scraper/scripts/cdp_fallback_scrape.py \
--asins "B07XXX,B08YYY,B09ZZZ" \
--output cdp_results.json
# 方式2:从 JSON 文件读 ASIN 列表
python3 ~/.openclaw/skills/amazon-scraper/scripts/cdp_fallback_scrape.py \
--asin-file failed_asins.json \
--output cdp_results.json
```
## 完整工作流:Docker 批量 + CDP 补漏
```bash
# Step 1: Docker 批量爬取(快但有失败)
# 每个容器一个独立出口端口。没有 N 个不同出口就不要开 N 个容器。
# 凭证从环境变量取,不要写进命令行(会进 shell history 和 ps 输出)。
read -rsp 'proxy user: ' PROXY_USER; echo
read -rsp 'proxy pass: ' PROXY_PASS; echo
for i in 0 1 2 3 4; do
PORT=$((8001 + i))
AMAZON_PROXIES="http://${PROXY_USER}:${PROXY_PASS}@isp.oxylabs.io:${PORT}" \
docker run --rm -v ~/scrapes:/data \
-e AMAZON_PROXIES \
amazon-scraper node assets/amazon_handler.js \
--asins "$BATCH_$i" --concurrency 2 \
--output "batch_${i}.json" &
done
wait
# Step 2: 找出失败的 ASIN(title/brand 全 null)
python3 -c "
import json
failed = []
for i in range(5):
d = json.load(open(f'batch_{i}.json'))
for p in d.get('products', []):
if not (p.get('title') or p.get('brand')):
failed.append(p['asin'])
json.dump(failed, open('failed_asins.json', 'w'))
print(f'{len(failed)} failed ASINs')
"
# Step 3: CDP 补漏(串行,100% 成功率)
python3 ~/.openclaw/skills/amazon-scraper/scripts/cdp_fallback_scrape.py \
--asin-file failed_asins.json \
--outpactivepieces
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
AionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
CopilotKit
The Frontend for Agents & Generative UI. React + Angular
Machine-readable data
The same record, as JSON, for agents and crawlers.
{
"facts": [
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Clawhub",
"href": "https://clawhub.ai/jiafar/skills/amazon-scraper",
"sourceUrl": "https://clawhub.ai/jiafar/skills/amazon-scraper",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-09T05:12:46.315Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-jiafar-amazon-scraper/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-jiafar-amazon-scraper/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-09T05:12:46.315Z",
"isPublic": true
},
{
"factKey": "traction",
"category": "adoption",
"label": "Adoption signal",
"value": "4.6K downloads",
"href": "https://clawhub.ai/jiafar/amazon-scraper",
"sourceUrl": "https://clawhub.ai/jiafar/amazon-scraper",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-09T05:12:46.315Z",
"isPublic": true
},
{
"factKey": "latest_release",
"category": "release",
"label": "Latest release",
"value": "4.0.1",
"href": "https://clawhub.ai/jiafar/amazon-scraper",
"sourceUrl": "https://clawhub.ai/jiafar/amazon-scraper",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-10-02T07:28:50.735Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-jiafar-amazon-scraper/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-jiafar-amazon-scraper/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
],
"events": [
{
"eventType": "release",
"title": "Release 4.0.1",
"description": "OpenClaw skill format. Docker is required via SKILL.md metadata. Credentials stay out of the bundle. Scheduled monitors use openclaw automations.",
"href": "https://clawhub.ai/jiafar/amazon-scraper",
"sourceUrl": "https://clawhub.ai/jiafar/amazon-scraper",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-10-02T07:28:50.735Z",
"isPublic": true
}
]
}Record generated Oct 9, 2026.
