agentCLAWHUBUnverified

102 Playwright Scraper Skill

Playwright-based web scraping OpenClaw Skill with anti-bot protection. Successfully tested on complex sites like Discuss.com.hk. Skill: 102 Playwright Scraper Skill Owner: smallkeyboy Summary: Playwright-based web scraping OpenClaw Skill with anti-bot protection. Successfully tested on complex sites like Discuss.com.hk. Tags: latest:1.0.0 Version history: v1.0.0 | 2026-04-16T07:09:40.650Z | auto **Playwright Scraper Skill v1.2.0:** - Added comprehensive usage guide and script descriptions, including detailed guidance for various website anti-b

OpenClaw

Rank

62

Safety

84

Downloads

1.0k

Updated

Oct 11, 2026

Version

1.0.0

Source

CLAWHUB

About

What it does, and when to use it.

Capability contract not published. No trust telemetry is available yet. 1K downloads reported by the source. Last updated 10/11/2026.

Avoid when

  • Contract metadata is missing or unavailable for deterministic execution.

Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing

Public facts

Every fact links back to the source it came from.

Vendor
Clawhubvendor · observed Oct 11, 2026
Protocol compatibility
OpenClawcompatibility · observed Oct 11, 2026
Adoption signal
1K downloadsadoption · observed Oct 11, 2026
Latest release
1.0.0release · observed Apr 16, 2026
Handshake status
UNKNOWNsecurity

Install and run

Setup complexity: low.

clawhub skill install s170nr6nxfsrk9n96c9r21ydcx83j5pf:102-playwright-scraper-skill
  1. Setup complexity is LOW. This package is likely designed for quick installation with minimal external side-effects.
  2. Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.

Contract: missing

curl -s "https://www.xpersona.co/api/v1/agents/clawhub-smallkeyboy-102-playwright-scraper-skill/snapshot"

Run-check

$0.02 USD

1 measured facts are behind this paywall: success rate and latency, uptime and estimated cost, when not to use it, how to call it, benchmark scores.

Agents pay $0.02 in USDC. A card payment is $0.50, the smallest a card allows.

Documentation

CLAWHUB

30,783 characters of source documentation, loaded on request.

Extracted files

5 files captured from the source.

SKILL.md

---
name: playwright-scraper-skill
description: Playwright-based web scraping OpenClaw Skill with anti-bot protection. Successfully tested on complex sites like Discuss.com.hk.
version: 1.2.0
author: Simon Chan
---

# Playwright Scraper Skill

A Playwright-based web scraping OpenClaw Skill with anti-bot protection. Choose the best approach based on the target website's anti-bot level.

---

## 🎯 Use Case Matrix

| Target Website | Anti-Bot Level | Recommended Method | Script |
|---------------|----------------|-------------------|--------|
| **Regular Sites** | Low | web_fetch tool | N/A (built-in) |
| **Dynamic Sites** | Medium | Playwright Simple | `scripts/playwright-simple.js` |
| **Cloudflare Protected** | High | **Playwright Stealth** ⭐ | `scripts/playwright-stealth.js` |
| **YouTube** | Special | deep-scraper | Install separately |
| **Reddit** | Special | reddit-scraper | Install separately |

---

## 📦 Installation

```bash
cd playwright-scraper-skill
npm install
npx playwright install chromium
```

---

## 🚀 Quick Start

### 1️⃣ Simple Sites (No Anti-Bot)

Use OpenClaw's built-in `web_fetch` tool:

```bash
# Invoke directly in OpenClaw
Hey, fetch me the content from https://example.com
```

---

### 2️⃣ Dynamic Sites (Requires JavaScript)

Use **Playwright Simple**:

```bash
node scripts/playwright-simple.js "https://example.com"
```

**Example output:**
```json
{
  "url": "https://example.com",
  "title": "Example Domain",
  "content": "...",
  "elapsedSeconds": "3.45"
}
```

---

### 3️⃣ Anti-Bot Protected Sites (Cloudflare etc.)

Use **Playwright Stealth**:

```bash
node scripts/playwright-stealth.js "https://m.discuss.com.hk/#hot"
```

**Features:**
- Hide automation markers (`navigator.webdriver = false`)
- Realistic User-Agent (iPhone, Android)
- Random delays to mimic human behavior
- Screenshot and HTML saving support

---

### 4️⃣ YouTube Video Transcripts

Use **deep-scraper** (install separately):

```bash
# Install deep-scraper skill
npx clawhub install deep-scraper

# Use it
cd skills/deep-scraper
node assets/youtube_handler.js "https://www.youtube.com/watch?v=VIDEO_ID"
```

---

## 📖 Script Descriptions

### `scripts/playwright-simple.js`
- **Use Case:** Regular dynamic websites
- **Speed:** Fast (3-5 seconds)
- **Anti-Bot:** None
- **Output:** JSON (title, content, URL)

### `scripts/playwright-stealth.js` ⭐
- **Use Case:** Sites with Cloudflare or anti-bot protection
- **Speed:** Medium (5-20 seconds)
- **Anti-Bot:** Medium-High (hides automation, realistic UA)
- **Output:** JSON + Screenshot + HTML file
- **Verified:** 100% success on Discuss.com.hk

---

## 🎓 Best Practices

### 1. Try web_fetch First
If the site doesn't have dynamic loading, use OpenClaw's `web_fetch` tool—it's fastest.

### 2. Need JavaScript? Use Playwright Simple
If you need to wait for JavaScript rendering, use `playwright-simple.js`.

### 3. Getting Blocked? Use Stealth
If you encounter 403 or Cloudflare challenges, use `playwright-stealth.j

examples/README.md

# Usage Examples

## Basic Usage

### 1. Quick Scrape (Example.com)

```bash
node scripts/playwright-simple.js https://example.com
```

**Output:**
```json
{
  "title": "Example Domain",
  "url": "https://example.com/",
  "content": "Example Domain\n\nThis domain is for use...",
  "metaDescription": "",
  "elapsedSeconds": "3.42"
}
```

---

### 2. Anti-Bot Protected Site (Discuss.com.hk)

```bash
node scripts/playwright-stealth.js "https://m.discuss.com.hk/#hot"
```

**Output:**
```json
{
  "title": "香港討論區 discuss.com.hk",
  "url": "https://m.discuss.com.hk/#hot",
  "htmlLength": 186345,
  "contentPreview": "...",
  "cloudflare": false,
  "screenshot": "./screenshot-1770467444364.png",
  "data": {
    "links": [
      {
        "text": "區議員周潔瑩疑消防通道違泊 道歉稱急於搬貨",
        "href": "https://m.discuss.com.hk/index.php?action=thread&tid=32148378..."
      }
    ]
  },
  "elapsedSeconds": "19.59"
}
```

---

## Advanced Usage

### 3. Custom Wait Time

```bash
WAIT_TIME=15000 node scripts/playwright-stealth.js <URL>
```

### 4. Show Browser (Debug Mode)

```bash
HEADLESS=false node scripts/playwright-stealth.js <URL>
```

### 5. Save Screenshot and HTML

```bash
SCREENSHOT_PATH=/tmp/my-page.png \
SAVE_HTML=true \
node scripts/playwright-stealth.js <URL>
```

### 6. Custom User-Agent

```bash
USER_AGENT="Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36" \
node scripts/playwright-stealth.js <URL>
```

---

## Integration Examples

### Using in Shell Scripts

```bash
#!/bin/bash
# Run from playwright-scraper-skill directory

URL="https://example.com"
OUTPUT_FILE="result.json"

echo "🕷️  Starting scrape: $URL"

node scripts/playwright-stealth.js "$URL" > "$OUTPUT_FILE"

if [ $? -eq 0 ]; then
  echo "✅ Success! Results saved to: $OUTPUT_FILE"
else
  echo "❌ Failed"
  exit 1
fi
```

### Batch Scraping Multiple URLs

```bash
#!/bin/bash

URLS=(
  "https://example.com"
  "https://example.org"
  "https://example.net"
)

for url in "${URLS[@]}"; do
  echo "Scraping: $url"
  node scripts/playwright-stealth.js "$url" > "output_$(date +%s).json"
  sleep 5  # Avoid IP blocking
done
```

---

## Calling from Node.js

```javascript
const { spawn } = require('child_process');

function scrape(url) {
  return new Promise((resolve, reject) => {
    const proc = spawn('node', [
      'scripts/playwright-stealth.js',
      url
    ]);
    
    let output = '';
    
    proc.stdout.on('data', (data) => {
      output += data.toString();
    });
    
    proc.on('close', (code) => {
      if (code === 0) {
        try {
          // Extract JSON (last line)
          const lines = output.trim().split('\n');
          const json = JSON.parse(lines[lines.length - 1]);
          resolve(json);
        } catch (e) {
          reject(e);
        }
      } else {
        reject(new Error(`Exit code: ${code}`));
      }
    });
  });
}

// Usage
(async () => {
  const result = await scrape('https://example.com');
  console.log(result.title);
})();
```

---

## Common Scen

README.md

# Playwright Scraper Skill 🕷️

[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)
[![Node.js](https://img.shields.io/badge/Node.js-18+-green.svg)](https://nodejs.org/)
[![Playwright](https://img.shields.io/badge/Playwright-1.40+-blue.svg)](https://playwright.dev/)

**[中文文檔](README_ZH.md)** | English

A Playwright-based web scraping OpenClaw Skill with anti-bot protection. Successfully tested on complex websites like Discuss.com.hk.

> 📦 **Installation:** See [INSTALL.md](INSTALL.md)  
> 📚 **Full Documentation:** See [SKILL.md](SKILL.md)  
> 💡 **Examples:** See [examples/README.md](examples/README.md)

---

## ✨ Features

- ✅ **Pure Playwright** — Modern, powerful, easy to use
- ✅ **Anti-Bot Protection** — Hides automation, realistic UA
- ✅ **Verified** — 100% success on Discuss.com.hk
- ✅ **Simple to Use** — One-line commands
- ✅ **Customizable** — Environment variable support

---

## 🚀 Quick Start

### Installation

```bash
npm install
npx playwright install chromium
```

### Usage

```bash
# Quick scraping
node scripts/playwright-simple.js https://example.com

# Stealth mode (recommended)
node scripts/playwright-stealth.js "https://m.discuss.com.hk/#hot"
```

---

## 📖 Two Modes

| Mode | Use Case | Speed | Anti-Bot |
|------|----------|-------|----------|
| **Simple** | Regular dynamic sites | Fast (3-5s) | None |
| **Stealth** ⭐ | Sites with anti-bot | Medium (5-20s) | Medium-High |

### Simple Mode

For sites without anti-bot protection:

```bash
node scripts/playwright-simple.js <URL>
```

### Stealth Mode (Recommended)

For sites with Cloudflare or anti-bot protection:

```bash
node scripts/playwright-stealth.js <URL>
```

**Anti-Bot Techniques:**
- Hide `navigator.webdriver`
- Realistic User-Agent (iPhone)
- Human-like behavior simulation
- Screenshot and HTML saving support

---

## 🎯 Customization

All scripts support environment variables:

```bash
# Show browser
HEADLESS=false node scripts/playwright-stealth.js <URL>

# Custom wait time (milliseconds)
WAIT_TIME=10000 node scripts/playwright-stealth.js <URL>

# Save screenshot
SCREENSHOT_PATH=/tmp/page.png node scripts/playwright-stealth.js <URL>

# Save HTML
SAVE_HTML=true node scripts/playwright-stealth.js <URL>

# Custom User-Agent
USER_AGENT="Mozilla/5.0 ..." node scripts/playwright-stealth.js <URL>
```

---

## 📊 Test Results

| Website | Result | Time |
|---------|--------|------|
| **Discuss.com.hk** | ✅ 200 OK | 5-20s |
| **Example.com** | ✅ 200 OK | 3-5s |
| **Cloudflare Protected** | ✅ Mostly successful | 10-30s |

---

## 📁 File Structure

```
playwright-scraper-skill/
├── scripts/
│   ├── playwright-simple.js       # Simple mode
│   └── playwright-stealth.js      # Stealth mode ⭐
├── examples/
│   ├── discuss-hk.sh              # Discuss.com.hk example
│   └── README.md                  # More examples
├── SKILL.md                       # Full documentation
├── INSTALL.md                     # Installation g

_meta.json

{
  "ownerId": "kn7ez4h284sw1q2f7zvmge91wx83jdbt",
  "slug": "102-playwright-scraper-skill",
  "version": "1.0.0",
  "publishedAt": 1776323380650
}

CHANGELOG.md

# Changelog

## [1.2.0] - 2026-02-07

### 🔄 Major Changes

- **Project Renamed** — `web-scraper` → `playwright-scraper-skill`
- Updated all documentation and links
- Updated GitHub repo name
- **Bilingual Documentation** — All docs now in English (with Chinese README available)

---

## [1.1.0] - 2026-02-07

### ✅ Added

- **LICENSE** — MIT License
- **CONTRIBUTING.md** — Contribution guidelines
- **examples/README.md** — Detailed usage examples
- **test.sh** — Automated test script
- **README.md** — Redesigned with badges

### 🔧 Improvements

- Clearer file structure
- More detailed documentation
- More practical examples

---

## [1.0.0] - 2026-02-07

### ✅ Initial Release

**Tools Created:**
- ✅ `playwright-simple.js` — Fast simple scraper
- ✅ `playwright-stealth.js` — Anti-bot protected version (primary) ⭐

**Test Results:**
- ✅ Discuss.com.hk success (200 OK, 19.6s)
- ✅ Example.com success (3.4s)
- ✅ Auto fallback to deep-scraper's Playwright

**Documentation:**
- ✅ SKILL.md (full documentation)
- ✅ README.md (quick reference)
- ✅ Example scripts (discuss-hk.sh)
- ✅ package.json

**Key Findings:**
1. **Playwright Stealth is the best solution** (100% success on Discuss.com.hk)
2. **Don't use Crawlee** (easily detected)
3. **Chaser (Rust) doesn't work currently** (blocked by Cloudflare)
4. **Hiding `navigator.webdriver` is key**

---

## Future Plans

- [ ] Add proxy IP rotation
- [ ] CAPTCHA handling integration
- [ ] Cookie management (maintain login state)
- [ ] Batch scraping (parallel processing)
- [ ] Integration with OpenClaw browser tool
Github ReposUpdated 2d agoRank 70

AionUi

Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!

MCPOPENCLAW
Github ReposUpdated 6mo agoRank 70

activepieces

AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents

OPENCLAW
Github ReposUpdated 6mo agoRank 70

cherry-studio

AI productivity studio with smart chat, autonomous agents, and 300+ assistants.

MCPOPENCLAW
Github ReposUpdated 7mo agoRank 70

CopilotKit

The Frontend for Agents & Generative UI. React + Angular

OPENCLAW

Machine-readable data

The same record, as JSON, for agents and crawlers.

{
  "facts": [
    {
      "factKey": "vendor",
      "category": "vendor",
      "label": "Vendor",
      "value": "Clawhub",
      "href": "https://clawhub.ai/smallkeyboy/skills/102-playwright-scraper-skill",
      "sourceUrl": "https://clawhub.ai/smallkeyboy/skills/102-playwright-scraper-skill",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-11T14:54:26.941Z",
      "isPublic": true
    },
    {
      "factKey": "protocols",
      "category": "compatibility",
      "label": "Protocol compatibility",
      "value": "OpenClaw",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-smallkeyboy-102-playwright-scraper-skill/contract",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-smallkeyboy-102-playwright-scraper-skill/contract",
      "sourceType": "contract",
      "confidence": "medium",
      "observedAt": "2026-10-11T14:54:26.941Z",
      "isPublic": true
    },
    {
      "factKey": "traction",
      "category": "adoption",
      "label": "Adoption signal",
      "value": "1K downloads",
      "href": "https://clawhub.ai/smallkeyboy/102-playwright-scraper-skill",
      "sourceUrl": "https://clawhub.ai/smallkeyboy/102-playwright-scraper-skill",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-11T14:54:26.941Z",
      "isPublic": true
    },
    {
      "factKey": "latest_release",
      "category": "release",
      "label": "Latest release",
      "value": "1.0.0",
      "href": "https://clawhub.ai/smallkeyboy/102-playwright-scraper-skill",
      "sourceUrl": "https://clawhub.ai/smallkeyboy/102-playwright-scraper-skill",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-04-16T07:09:40.650Z",
      "isPublic": true
    },
    {
      "factKey": "handshake_status",
      "category": "security",
      "label": "Handshake status",
      "value": "UNKNOWN",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-smallkeyboy-102-playwright-scraper-skill/trust",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-smallkeyboy-102-playwright-scraper-skill/trust",
      "sourceType": "trust",
      "confidence": "medium",
      "observedAt": null,
      "isPublic": true
    }
  ],
  "events": [
    {
      "eventType": "release",
      "title": "Release 1.0.0",
      "description": "**Playwright Scraper Skill v1.2.0:** - Added comprehensive usage guide and script descriptions, including detailed guidance for various website anti-bot levels. - Introduced a use case matrix and performance comparison chart for choosing the best scraping method. - Documented anti-bot protection strategies employed in the provided scripts. - Included troubleshooting steps, environmental variable customization, and best practices for higher scraping success. - Outlined future improvement plans and provided helpful external references.",
      "href": "https://clawhub.ai/smallkeyboy/102-playwright-scraper-skill",
      "sourceUrl": "https://clawhub.ai/smallkeyboy/102-playwright-scraper-skill",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-04-16T07:09:40.650Z",
      "isPublic": true
    }
  ]
}

Record generated Oct 11, 2026.

Sponsored

Ads related to 102 Playwright Scraper Skill and adjacent AI workflows.