agentCLAWHUBUnverified

Scrapling - Stealth Web Scraper

Web scraping using Scrapling — a Python framework with anti-bot bypass (Cloudflare Turnstile, fingerprint spoofing), adaptive element tracking, stealth headl... Skill: Scrapling - Stealth Web Scraper Owner: jeminay Summary: Web scraping using Scrapling — a Python framework with anti-bot bypass (Cloudflare Turnstile, fingerprint spoofing), adaptive element tracking, stealth headl... Tags: latest:1.0.3 Version history: v1.0.3 | 2026-02-25T07:47:52.031Z | user Added license (MIT) and metadata.source/pypi fields to frontmatter so registry shows verified provenance instead of 'So

OpenClaw

Rank

62

Safety

84

Downloads

2.9k

Updated

Oct 9, 2026

Version

1.0.3

Source

CLAWHUB

About

What it does, and when to use it.

Capability contract not published. No trust telemetry is available yet. 2.9K downloads reported by the source. Last updated 10/9/2026.

Avoid when

  • Contract metadata is missing or unavailable for deterministic execution.

Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing

Public facts

Every fact links back to the source it came from.

Vendor
Clawhubvendor · observed Oct 9, 2026
Protocol compatibility
OpenClawcompatibility · observed Oct 9, 2026
Adoption signal
2.9K downloadsadoption · observed Oct 9, 2026
Latest release
1.0.3release · observed Feb 25, 2026
Handshake status
UNKNOWNsecurity

Install and run

Setup complexity: low.

clawhub skill install s170dx4nw2adr8scv090bb28dd84pbe2:scrapling-fetcher
  1. Setup complexity is LOW. This package is likely designed for quick installation with minimal external side-effects.
  2. Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.

Contract: missing

curl -s "https://www.xpersona.co/api/v1/agents/clawhub-jeminay-scrapling-fetcher/snapshot"

Documentation

CLAWHUB

28,255 characters of source documentation, loaded on request.

Extracted files

4 files captured from the source.

SKILL.md

---
name: scrapling
description: "Web scraping using Scrapling — a Python framework with anti-bot bypass (Cloudflare Turnstile, fingerprint spoofing), adaptive element tracking, stealth headless browser, and full CSS/XPath extraction. Use when web_fetch fails (Cloudflare, JS-rendered pages), or when extracting structured data from websites (prices, articles, lists). Supports HTTP, stealth, and full browser modes. Source: github.com/D4Vinci/Scrapling (PyPI: scrapling). Only use on sites you have permission to scrape."
license: MIT
metadata:
  source: https://github.com/D4Vinci/Scrapling
  pypi: https://pypi.org/project/scrapling/
---

# Scrapling Skill

**Source:** https://github.com/D4Vinci/Scrapling (open source, MIT-like license)
**PyPI:** `scrapling` — install before first use (see below)

> ⚠️ Only scrape sites you have permission to access. Respect `robots.txt` and Terms of Service. Do not use stealth modes to bypass paywalls or access restricted content without authorization.

## Installation (one-time, confirm with user before running)

```bash
pip install scrapling[all]
patchright install chromium  # required for stealth/dynamic modes
```

- `scrapling[all]` installs `patchright` (a stealth fork of Playwright, bundled as a PyPI package — not a typo), `curl_cffi`, MCP server deps, and IPython shell.
- `patchright install chromium` downloads Chromium (~100 MB) via patchright's own installer (same mechanism as `playwright install chromium`).
- Confirm with user before running — installs ~200 MB of dependencies and browser binaries.

## Script

`scripts/scrape.py` — CLI wrapper for all three fetcher modes.

```bash
# Basic fetch (text output)
python3 ~/skills/scrapling/scripts/scrape.py <url> -q

# CSS selector extraction
python3 ~/skills/scrapling/scripts/scrape.py <url> --selector ".class" -q

# Stealth mode (Cloudflare bypass) — only on sites you're authorized to access
python3 ~/skills/scrapling/scripts/scrape.py <url> --mode stealth -q

# JSON output
python3 ~/skills/scrapling/scripts/scrape.py <url> --selector "h2" --json -q
```

## Fetcher Modes

- **http** (default) — Fast HTTP with browser TLS fingerprint spoofing. Most sites.
- **stealth** — Headless Chrome with anti-detect. For Cloudflare/anti-bot.
- **dynamic** — Full Playwright browser. For heavy JS SPAs.

## When to Use Each Mode

- `web_fetch` returns 403/429/Cloudflare challenge → use `--mode stealth`
- Page content requires JS execution → use `--mode dynamic`
- Regular site, just need text/data → use `--mode http` (default)

## Python Inline Usage

For custom logic beyond the CLI, write inline Python. See `references/patterns.md` for:
- Adaptive scraping (`auto_save` / `adaptive` — saves element fingerprints locally)
- Session/cookie handling
- Async usage
- XPath, find_similar, attribute extraction

## Notes

- **MCP server** (`scrapling mcp`): starts a local network service for AI-native scraping. Only start if explicitly needed and trusted — it exposes a local HTTP server.

_meta.json

{
  "ownerId": "kn73589k2rv646tm107cp6d90s81g7xy",
  "slug": "scrapling-fetcher",
  "version": "1.0.3",
  "publishedAt": 1772005672031
}

references/patterns.md

# Scrapling Patterns Reference

## Fetcher Selection Guide

| Scenario | Fetcher | Notes |
|---|---|---|
| Regular sites, APIs | `Fetcher` | Fastest, HTTP-only |
| Cloudflare, anti-bot | `StealthyFetcher` | Headless Chrome, fingerprint spoofing |
| Heavy JS rendering | `DynamicFetcher` | Full Playwright browser |
| Async pipeline | `AsyncFetcher` | Async equivalent of Fetcher |

## Python Quick Patterns

### Basic HTTP fetch
```python
from scrapling.fetchers import Fetcher
page = Fetcher.get('https://example.com')
print(page.status)  # 200
text = page.get_all_text(ignore_tags=('script', 'style'))
```

### Stealth fetch (bypass Cloudflare)
```python
from scrapling.fetchers import StealthyFetcher
page = StealthyFetcher.fetch('https://protected-site.com', headless=True, network_idle=True)
```

### Dynamic fetch (JS-rendered content)
```python
from scrapling.fetchers import DynamicFetcher
page = DynamicFetcher.fetch('https://spa-site.com', headless=True, network_idle=True)
```

### CSS selector extraction
```python
titles = page.css('h2.title')
for t in titles:
    print(t.text)

# Get attribute
links = page.css('a.product-link')
for a in links:
    print(a.attrib['href'])
```

### XPath extraction
```python
items = page.xpath('//div[@class="item"]/span/text()')
```

### Adaptive scraping (survives site redesigns)
```python
# First run: auto_save=True saves element fingerprints
products = page.css('.product-card', auto_save=True)
# Later runs: adaptive=True finds them even if CSS changed
products = page.css('.product-card', adaptive=True)
```

### Find similar elements
```python
first = page.css('.price')[0]
all_prices = first.find_similar()
```

### Session with cookies
```python
from scrapling.fetchers import FetcherSession
session = FetcherSession()
session.get('https://example.com/login', data={'user': 'x', 'pass': 'y'})
page = session.get('https://example.com/dashboard')
```

### Async usage
```python
import asyncio
from scrapling.fetchers import AsyncFetcher

async def scrape():
    page = await AsyncFetcher.get('https://example.com')
    return page.css('h1')[0].text

asyncio.run(scrape())
```

## CLI Usage

```bash
# Simple text extraction
python3 scrape.py https://example.com

# CSS selector extraction
python3 scrape.py https://example.com --selector "h2.title"

# Extract attribute value
python3 scrape.py https://example.com --selector "a.product" --attr href

# Stealth mode for protected sites
python3 scrape.py https://cloudflare-site.com --mode stealth

# JSON output
python3 scrape.py https://example.com --selector ".price" --json

# Quiet mode (no INFO logs)
python3 scrape.py https://example.com -q
```

## MCP Server Setup

> ⚠️ The MCP server starts a local HTTP service. Only use in trusted environments.

```bash
scrapling mcp
# or
python3 -m scrapling.mcp
```

Add to OpenClaw MCP config (mcporter) to get scraping as a native tool. Confirm with user before starting.

skill-card.md

## Description:

Provides a Scrapling-based web scraping workflow for fetching pages, extracting structured data with CSS or XPath selectors, and using HTTP, stealth, or dynamic browser modes on sites the user is authorized to scrape.

This skill is ready for commercial/non-commercial use.

## Publisher:

[jeminay](https://clawhub.ai/user/jeminay)

### License/Terms of Use:

MIT

## Use Case:

Developers and automation engineers use this skill to retrieve web page text or structured fields when ordinary fetch tools fail because of JavaScript rendering, rate limits, or anti-bot challenges. It is intended for authorized scraping workflows and selector-based extraction, not for bypassing restricted access.

### Deployment Geography for Use:

Global

## Known Risks and Mitigations:

Risk: The skill can send requests through HTTP, stealth, and browser-based modes to external sites.

Mitigation: Use it only for sites the user is authorized to scrape, respect applicable terms and robots.txt, and avoid passing credentials or sensitive URLs unless the destination is trusted.

Risk: Installing full Scrapling dependencies and browser binaries expands the local runtime surface.

Mitigation: Install and run the skill in a virtual environment or disposable container without administrator privileges.

Risk: The optional MCP server starts a local HTTP service.

Mitigation: Start the server only when explicitly needed and only in a trusted local environment.

Risk: Adaptive scraping can persist element fingerprints or local state in the working directory.

Mitigation: Run adaptive workflows in a controlled directory and remove stored state when it is no longer needed.

## Reference(s):

- [Scrapling Patterns Reference](references/patterns.md)
- [Scrapling on PyPI](https://pypi.org/project/scrapling/)
- [ClawHub Skill Page](https://clawhub.ai/jeminay/skills/scrapling-fetcher)

## Skill Output:

**Output Type(s):** [Text, Markdown, Code, Shell commands, Configuration, Guidance]

**Output Format:** [Markdown guidance with inline shell commands, Python examples, plain text extraction, and optional JSON arrays]

**Output Parameters:** [1D]

**Other Properties Related to Output:** [Can emit selector matches as text lines or JSON; stealth and dynamic modes may use headless Chromium.]

## Skill Version(s):

1.0.3 (source: server release metadata)

## Ethical Considerations:

Users should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.
Github ReposUpdated 2h agoRank 70

AionUi

Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!

MCPOPENCLAW
Github ReposUpdated 6mo agoRank 70

activepieces

AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents

OPENCLAW
Github ReposUpdated 6mo agoRank 70

cherry-studio

AI productivity studio with smart chat, autonomous agents, and 300+ assistants.

MCPOPENCLAW
Github ReposUpdated 7mo agoRank 70

CopilotKit

The Frontend for Agents & Generative UI. React + Angular

OPENCLAW

Machine-readable data

The same record, as JSON, for agents and crawlers.

{
  "facts": [
    {
      "factKey": "vendor",
      "category": "vendor",
      "label": "Vendor",
      "value": "Clawhub",
      "href": "https://clawhub.ai/jeminay/skills/scrapling-fetcher",
      "sourceUrl": "https://clawhub.ai/jeminay/skills/scrapling-fetcher",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-09T10:38:41.911Z",
      "isPublic": true
    },
    {
      "factKey": "protocols",
      "category": "compatibility",
      "label": "Protocol compatibility",
      "value": "OpenClaw",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-jeminay-scrapling-fetcher/contract",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-jeminay-scrapling-fetcher/contract",
      "sourceType": "contract",
      "confidence": "medium",
      "observedAt": "2026-10-09T10:38:41.911Z",
      "isPublic": true
    },
    {
      "factKey": "traction",
      "category": "adoption",
      "label": "Adoption signal",
      "value": "2.9K downloads",
      "href": "https://clawhub.ai/jeminay/scrapling-fetcher",
      "sourceUrl": "https://clawhub.ai/jeminay/scrapling-fetcher",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-09T10:38:41.911Z",
      "isPublic": true
    },
    {
      "factKey": "latest_release",
      "category": "release",
      "label": "Latest release",
      "value": "1.0.3",
      "href": "https://clawhub.ai/jeminay/scrapling-fetcher",
      "sourceUrl": "https://clawhub.ai/jeminay/scrapling-fetcher",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-02-25T07:47:52.031Z",
      "isPublic": true
    },
    {
      "factKey": "handshake_status",
      "category": "security",
      "label": "Handshake status",
      "value": "UNKNOWN",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-jeminay-scrapling-fetcher/trust",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-jeminay-scrapling-fetcher/trust",
      "sourceType": "trust",
      "confidence": "medium",
      "observedAt": null,
      "isPublic": true
    }
  ],
  "events": [
    {
      "eventType": "release",
      "title": "Release 1.0.3",
      "description": "Added license (MIT) and metadata.source/pypi fields to frontmatter so registry shows verified provenance instead of 'Source: unknown'.",
      "href": "https://clawhub.ai/jeminay/scrapling-fetcher",
      "sourceUrl": "https://clawhub.ai/jeminay/scrapling-fetcher",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-02-25T07:47:52.031Z",
      "isPublic": true
    }
  ]
}

Record generated Oct 9, 2026.

Sponsored

Ads related to Scrapling - Stealth Web Scraper and adjacent AI workflows.