scrapling
高性能自适应 Python 网页抓取框架,内置反爬虫绕过(Cloudflare Turnstile)、智能元素重定位、完整爬虫框架和 MCP 服务器,适合 AI 辅助数据提取和大规模爬取任务
Rank
62
Safety
84
Downloads
1.0k
Updated
Oct 11, 2026
Version
0.1.0
Source
CLAWHUB
About
What it does, and when to use it.
Capability contract not published. No trust telemetry is available yet. 1K downloads reported by the source. Last updated 10/11/2026.
Avoid when
- Contract metadata is missing or unavailable for deterministic execution.
Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing
Public facts
Every fact links back to the source it came from.
- Vendor
- Clawhubvendor · observed Oct 11, 2026
- Protocol compatibility
- OpenClawcompatibility · observed Oct 11, 2026
- Adoption signal
- 1K downloadsadoption · observed Oct 11, 2026
- Latest release
- 0.1.0release · observed Apr 22, 2026
- Handshake status
- UNKNOWNsecurity
Install and run
Setup complexity: low.
clawhub skill install s17dbd0j0kazk5ay31x9kv1z3n84f04v:cn-scrapling- Install using `clawhub skill install s17dbd0j0kazk5ay31x9kv1z3n84f04v:cn-scrapling` in an isolated environment before connecting it to live workloads.
- No published capability contract is available yet, so validate auth and request/response behavior manually.
- Review the upstream CLAWHUB listing at https://clawhub.ai/cn-big-cabbage/cn-scrapling before using production credentials.
Contract: missing
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-cn-big-cabbage-cn-scrapling/snapshot"
Run-check
$0.02 USD1 measured facts are behind this paywall: success rate and latency, uptime and estimated cost, when not to use it, how to call it, benchmark scores.
Agents pay $0.02 in USDC. A card payment is $0.50, the smallest a card allows.
Documentation
CLAWHUB
26,305 characters of source documentation, loaded on request.
Extracted files
5 files captured from the source.
SKILL.md
---
name: scrapling
description: 高性能自适应 Python 网页抓取框架,内置反爬虫绕过(Cloudflare Turnstile)、智能元素重定位、完整爬虫框架和 MCP 服务器,适合 AI 辅助数据提取和大规模爬取任务
version: 0.1.0
metadata:
openclaw_requires: ">=1.0.0"
emoji: 🕷️
homepage: https://scrapling.readthedocs.io
---
# Scrapling — 自适应网页抓取框架
Scrapling 是 Google Chrome DevTools 生态之外最强大的 Python 网页抓取框架之一,能够处理从单次 HTTP 请求到大规模并发爬取的所有场景。它的自适应解析引擎在网页改版后自动重新定位元素,内置 Cloudflare Turnstile 绕过能力,Spider 框架支持暂停/恢复,并提供 MCP 服务器让 AI 直接辅助数据提取,从源头减少 Token 消耗。
## 核心使用场景
- **反爬虫网站抓取**:`StealthyFetcher` 内置 Cloudflare Turnstile 绕过,支持 TLS 指纹伪装和浏览器自动化
- **自适应数据采集**:网页改版后,`auto_save=True` 保存元素快照,`adaptive=True` 自动重新定位变化元素
- **大规模并发爬取**:Spider 框架支持多 Session、代理轮换、暂停恢复,像 Scrapy 一样定义爬虫
- **AI 辅助提取**:内置 MCP 服务器,Claude/Cursor 等 AI 工具可直接调用 Scrapling 提取目标内容
- **动态页面处理**:`DynamicFetcher` 基于 Playwright,支持完整浏览器自动化和网络空闲等待
## AI 辅助使用流程
1. **安装依赖** — AI 执行 `pip install scrapling` 并按需安装浏览器驱动
2. **选择 Fetcher** — AI 根据目标网站类型推荐 `Fetcher`/`StealthyFetcher`/`DynamicFetcher`
3. **编写抓取逻辑** — AI 生成 CSS/XPath 选择器代码,配置 `auto_save` 实现自适应
4. **调试与优化** — AI 分析响应结果,调整选择器或切换 Fetcher 策略
5. **扩展为 Spider** — AI 将单页抓取扩展为完整 Spider 类,配置并发和代理
6. **MCP 模式** — 启动 Scrapling MCP Server,让 AI 直接操控浏览器提取数据
## 关键章节导航
- [安装指南](guides/01-installation.md) — pip 安装、浏览器驱动、Docker 镜像
- [快速开始](guides/02-quickstart.md) — Fetcher 选型、CSS/XPath 选择器、自适应抓取
- [高级用法](guides/03-advanced-usage.md) — Spider 框架、代理轮换、MCP 服务器、CLI 工具
- [故障排查](troubleshooting.md) — 反爬虫、浏览器驱动、超时、代理问题
## AI 助手能力
使用本技能时,AI 可以:
- ✅ 安装 Scrapling 并配置浏览器驱动(`scrapling install playwright` / `scrapling install camoufox`)
- ✅ 根据目标网站自动选择最合适的 Fetcher 类
- ✅ 编写 CSS/XPath 选择器提取目标数据
- ✅ 配置 `auto_save=True` 和 `adaptive=True` 实现自适应抓取
- ✅ 构建完整的 Spider 类实现并发爬取,配置暂停/恢复
- ✅ 设置代理轮换和防 DNS 泄露(DoH 模式)
- ✅ 启动和配置 Scrapling MCP 服务器
- ✅ 使用 CLI 工具快速测试 URL 抓取效果
## 核心功能
- ✅ **三种 Fetcher** — `Fetcher`(快速 HTTP)、`StealthyFetcher`(反爬绕过)、`DynamicFetcher`(浏览器自动化)
- ✅ **自适应解析** — 网页改版后自动重定位元素,降低维护成本
- ✅ **Cloudflare 绕过** — 内置 Turnstile/Interstitial 解决方案,免额外服务
- ✅ **Spider 框架** — Scrapy 风格 API,支持并发、多 Session、暂停恢复
- ✅ **流式输出** — `spider.stream()` 实时推送抓取结果,适合大规模任务
- ✅ **MCP 服务器** — AI 工具直接调用 Scrapling 提取数据,减少 Token 消耗
- ✅ **代理轮换** — 内置 `ProxyRotator`,支持循环或自定义策略
- ✅ **会话管理** — `FetcherSession`/`StealthySession`/`DynamicSession` 跨请求保持状态
- ✅ **开发模式** — 首次运行缓存响应,后续离线回放,快速迭代解析逻辑
- ✅ **CLI 工具** — 无需写代码直接从终端抓取页面
- ✅ **IPython Shell** — 交互式调试,内置 curl 转换工具
- ✅ **Docker 镜像** — 预置所有浏览器的生产就绪镜像
## 快速示例
```python
from scrapling.fetchers import Fetcher, StealthyFetcher, DynamicFetcher
# 普通 HTTP 抓取(最快)
page = Fetcher.get('https://quotes.toscrape.com/')
quotes = page.css('.quote .text::text').getall()
# 隐身模式绕过 Cloudflare
page = StealthyFetcher.fetch('https://protected-site.com', headless=True)
data = page.css('.content::text').get()
# 自适应抓取(网站改版后自动重定位)
page = Fetcher.get('https://example.com/products')
products = page.css('.product', auto_save=True) # 首次保存元素快照
# 网站改版后:
products = page.css('.product', adaptive=True) # 自动重新定位
```
```bash
# CLI 快速测试(无需写代码)
scrapling fetc_meta.json
{
"ownerId": "kn7c5ry95ff7x22mt3fx6j0q3d81p8t3",
"slug": "cn-scrapling",
"version": "0.1.0",
"publishedAt": 1776818626652
}guides/01-installation.md
# 安装指南
## 适用场景
- 安装 Scrapling 核心库并配置浏览器驱动
- 在 Docker 环境中使用预置镜像
- 为不同 Fetcher 安装对应依赖
---
## 基础安装
> **AI 可自动执行**
```bash
pip install scrapling
```
验证安装:
```bash
python -c "import scrapling; print(scrapling.__version__)"
```
---
## 按需安装浏览器驱动
Scrapling 有三种 Fetcher,各需不同依赖:
### Fetcher(纯 HTTP,无需额外驱动)
```bash
pip install scrapling
# 无需额外安装,开箱即用
```
### StealthyFetcher(隐身模式,需要 Camoufox)
```bash
pip install scrapling
scrapling install camoufox # 安装修改版 Firefox 驱动
```
或手动安装:
```bash
pip install camoufox[geoip]
python -m camoufox fetch
```
### DynamicFetcher(完整浏览器自动化,需要 Playwright)
```bash
pip install scrapling
scrapling install playwright # 安装 Playwright 和 Chromium
```
或手动安装:
```bash
pip install playwright
playwright install chromium
```
---
## 一次性安装全部依赖
```bash
pip install scrapling
scrapling install all
```
---
## Docker 安装(推荐生产环境)
使用官方预置镜像(含所有浏览器驱动):
```bash
# 拉取最新镜像
docker pull d4vinci/scrapling:latest
# 运行容器
docker run -it d4vinci/scrapling:latest python3
# 在容器内直接使用
docker run --rm d4vinci/scrapling:latest python3 -c "
from scrapling.fetchers import StealthyFetcher
page = StealthyFetcher.fetch('https://example.com', headless=True)
print(page.css('title::text').get())
"
```
---
## 虚拟环境(推荐)
```bash
python -m venv scrapling-env
source scrapling-env/bin/activate # Windows: scrapling-env\Scripts\activate
pip install scrapling
scrapling install playwright
```
---
## MCP 服务器安装(AI 集成)
Scrapling 内置 MCP 服务器,让 Claude/Cursor 等 AI 直接调用:
### Claude Code
```bash
claude mcp add scrapling --scope user npx -y scrapling-mcp
```
或手动安装后配置:
```json
{
"mcpServers": {
"scrapling": {
"command": "python",
"args": ["-m", "scrapling.mcp"]
}
}
}
```
### 直接启动 MCP 服务器
```bash
scrapling mcp
```
---
## 验证安装
```python
# 验证核心安装
from scrapling.fetchers import Fetcher
page = Fetcher.get('https://httpbin.org/get')
print(page.status) # 期望:200
# 验证 Playwright(DynamicFetcher)
from scrapling.fetchers import DynamicFetcher
page = DynamicFetcher.fetch('https://example.com', headless=True)
print(page.css('title::text').get())
# 验证 Camoufox(StealthyFetcher)
from scrapling.fetchers import StealthyFetcher
page = StealthyFetcher.fetch('https://example.com', headless=True)
print(page.css('title::text').get())
```
---
## 完成确认检查清单
- [ ] `pip install scrapling` 执行成功
- [ ] `python -c "import scrapling"` 无报错
- [ ] 按需安装了 Playwright 或 Camoufox(视场景而定)
- [ ] `Fetcher.get('https://httpbin.org/get').status == 200` 验证通过
---
## 下一步
- [快速开始](02-quickstart.md) — Fetcher 选型指南、CSS/XPath 选择器、自适应抓取guides/02-quickstart.md
# 快速开始
## 适用场景
- 从静态或动态页面提取数据
- 处理需要绕过反爬保护的网站
- 使用 CSS/XPath 选择器精准提取内容
- 实现网页改版后的自适应抓取
---
## 选择正确的 Fetcher
| 场景 | Fetcher | 速度 |
|------|---------|------|
| 普通 HTTP 请求 | `Fetcher` | 最快 |
| 需要 TLS 指纹伪装 | `Fetcher(impersonate='chrome')` | 快 |
| Cloudflare / 反爬 | `StealthyFetcher` | 中 |
| 需要 JS 渲染 | `DynamicFetcher` | 慢 |
---
## 基础 HTTP 抓取
```python
from scrapling.fetchers import Fetcher
# 单次请求
page = Fetcher.get('https://quotes.toscrape.com/')
# 提取数据
quotes = page.css('.quote .text::text').getall()
authors = page.css('.quote .author::text').getall()
print(quotes[:3])
# XPath 方式
titles = page.xpath('//span[@class="text"]/text()').getall()
```
---
## Session 复用(跨请求保持状态)
```python
from scrapling.fetchers import FetcherSession
with FetcherSession(impersonate='chrome') as session:
# 使用最新版 Chrome TLS 指纹
page1 = session.get('https://example.com/', stealthy_headers=True)
page2 = session.get('https://example.com/products') # 复用 cookie
products = page2.css('.product h2::text').getall()
```
---
## 绕过 Cloudflare 保护
```python
from scrapling.fetchers import StealthyFetcher, StealthySession
# 单次请求(每次打开/关闭浏览器)
page = StealthyFetcher.fetch(
'https://protected-site.com',
headless=True,
solve_cloudflare=True # 自动处理 Cloudflare Turnstile
)
data = page.css('.content').get()
# Session 模式(保持浏览器,效率更高)
with StealthySession(headless=True) as session:
page = session.fetch('https://protected-site.com', google_search=False)
links = page.css('a.product-link::attr(href)').getall()
```
---
## 动态页面(JS 渲染)
```python
from scrapling.fetchers import DynamicFetcher, DynamicSession
# 等待网络空闲后抓取(确保异步数据加载完成)
page = DynamicFetcher.fetch(
'https://spa-app.com',
headless=True,
network_idle=True
)
items = page.css('.item-list .item::text').getall()
# Session 模式(连续操作多个页面)
with DynamicSession(headless=True) as session:
page = session.fetch('https://example.com/login')
# 可以执行页面交互(通过 Playwright)
page2 = session.fetch('https://example.com/dashboard')
data = page2.css('.metric::text').getall()
```
---
## CSS 和 XPath 选择器
```python
page = Fetcher.get('https://quotes.toscrape.com/')
# CSS 选择器
title = page.css('h1::text').get() # 第一个元素
all_texts = page.css('.quote .text::text').getall() # 全部
# XPath 选择器
author = page.xpath('//small[@class="author"]/text()').get()
hrefs = page.xpath('//a/@href').getall()
# 属性提取
link = page.css('a::attr(href)').get()
src = page.css('img::attr(src)').get()
# 正则提取
price = page.css('.price::text').re_first(r'\$[\d.]+')
numbers = page.css('p::text').re(r'\d+')
```
---
## 自适应抓取(网页改版后自动重定位)
这是 Scrapling 最独特的功能:
```python
from scrapling.fetchers import Fetcher
# 第一次抓取:保存元素快照(存入本地数据库)
page = Fetcher.get('https://example.com/products')
products = page.css('.product-card', auto_save=True)
print(f"找到 {len(products)} 个商品")
for p in products:
print(p.css('h2::text').get(), p.css('.price::text').get())
# 网站改版后(.product-card 变成了 .item-wrapper)
# 传入 adaptguides/03-advanced-usage.md
# 高级用法
## Spider 框架(大规模并发爬取)
Spider 框架提供 Scrapy 风格的 API,适合需要跨多页面爬取的任务:
```python
from scrapling.spiders import Spider, Request, Response
class QuotesSpider(Spider):
name = "quotes"
start_urls = ["https://quotes.toscrape.com/"]
concurrent_requests = 10 # 最大并发数
download_delay = 0.5 # 请求间隔(秒)
robots_txt_obey = True # 遵守 robots.txt
async def parse(self, response: Response):
for quote in response.css('.quote'):
yield {
"text": quote.css('.text::text').get(),
"author": quote.css('.author::text').get(),
"tags": quote.css('.tag::text').getall(),
}
# 翻页
next_page = response.css('li.next a::attr(href)').get()
if next_page:
yield Request(response.urljoin(next_page))
# 运行爬虫
result = QuotesSpider().start()
result.items.to_json('quotes.json') # 导出 JSON
result.items.to_jsonl('quotes.jsonl') # 导出 JSONL
```
---
## 暂停与恢复爬虫
```python
from scrapling.spiders import Spider
class LargeSpider(Spider):
name = "large"
start_urls = ["https://large-site.com/"]
# 开启检查点持久化
checkpoint = True
async def parse(self, response):
# ...爬取逻辑...
pass
# 启动时按 Ctrl+C 优雅停止
result = LargeSpider().start()
# 下次运行时自动从上次停止的地方继续
result = LargeSpider().start() # 无需额外配置
```
---
## 流式输出(实时处理)
```python
from scrapling.spiders import Spider
import asyncio
class StreamingSpider(Spider):
name = "stream"
start_urls = ["https://news-site.com/"]
async def parse(self, response):
for article in response.css('article'):
yield {"title": article.css('h2::text').get()}
async def main():
spider = StreamingSpider()
async for item in spider.stream():
# 实时处理每个抓取到的 item
print(item)
# 实时写入数据库、发送到队列等
asyncio.run(main())
```
---
## 多 Session 路由(混用 HTTP 和浏览器)
```python
from scrapling.spiders import Spider, Request, Response
class HybridSpider(Spider):
name = "hybrid"
start_urls = ["https://example.com/catalog"]
async def parse(self, response: Response):
for link in response.css('.product-link::attr(href)').getall():
# 普通页面用 HTTP,详情页面用浏览器
if '/protected/' in link:
yield Request(link, session_id='stealthy') # 使用 StealthyFetcher
else:
yield Request(link) # 使用默认 HTTP
async def parse_product(self, response: Response):
yield {"title": response.css('h1::text').get()}
```
---
## 代理轮换
```python
from scrapling.fetchers import Fetcher, FetcherSession, StealthyFetcher
from scrapling.proxy_rotator import ProxyRotator
# 设置代理列表
proxies = [
"http://user:[email protected]:8080",
"http://user:[email protected]:8080",
"socks5://user:[email protected]:1080",
]
# 循环轮换
rotator = ProxyRotator(proxies, mode='cyclic')
with FetcherSession(proxy=rotator) as session:
page = session.get(AionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
activepieces
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
CopilotKit
The Frontend for Agents & Generative UI. React + Angular
Machine-readable data
The same record, as JSON, for agents and crawlers.
{
"facts": [
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Clawhub",
"href": "https://clawhub.ai/cn-big-cabbage/skills/cn-scrapling",
"sourceUrl": "https://clawhub.ai/cn-big-cabbage/skills/cn-scrapling",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-11T18:08:57.473Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-cn-big-cabbage-cn-scrapling/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-cn-big-cabbage-cn-scrapling/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-11T18:08:57.473Z",
"isPublic": true
},
{
"factKey": "traction",
"category": "adoption",
"label": "Adoption signal",
"value": "1K downloads",
"href": "https://clawhub.ai/cn-big-cabbage/cn-scrapling",
"sourceUrl": "https://clawhub.ai/cn-big-cabbage/cn-scrapling",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-11T18:08:57.473Z",
"isPublic": true
},
{
"factKey": "latest_release",
"category": "release",
"label": "Latest release",
"value": "0.1.0",
"href": "https://clawhub.ai/cn-big-cabbage/cn-scrapling",
"sourceUrl": "https://clawhub.ai/cn-big-cabbage/cn-scrapling",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-04-22T00:43:46.652Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-cn-big-cabbage-cn-scrapling/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-cn-big-cabbage-cn-scrapling/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
],
"events": [
{
"eventType": "release",
"title": "Release 0.1.0",
"description": "Initial release of Scrapling, a high-performance adaptive Python web scraping framework. - Supports automated anti-bot bypass (Cloudflare Turnstile), adaptive element re-location, and complete spider framework. - Includes MCP server for AI-assisted data extraction to reduce token consumption. - Offers three Fetcher classes (Fetcher, StealthyFetcher, DynamicFetcher) for HTTP, stealth/anti-bot, and browser automation tasks. - Features proxy rotation, session management, pause/resume for spiders, and CLI utilities. - Provides comprehensive documentation and integration guides for rapid use and extension.",
"href": "https://clawhub.ai/cn-big-cabbage/cn-scrapling",
"sourceUrl": "https://clawhub.ai/cn-big-cabbage/cn-scrapling",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-04-22T00:43:46.652Z",
"isPublic": true
}
]
}Record generated Oct 11, 2026.
