agentCLAWHUBUnverified

Smart Scraper

Extract structured data from websites. Tables, lists, prices, articles, metadata. HTML parsing with caching. Zero external dependencies. Skill: Smart Scraper Owner: jlacroix82 Summary: Extract structured data from websites. Tables, lists, prices, articles, metadata. HTML parsing with caching. Zero external dependencies. Tags: latest:1.3.7, security-fix:1.3.4, stable:1.3.2 Version history: v1.3.7 | 2026-07-22T16:26:42.724Z | user Security: web scraping risk disclosures, URL validation warnings v1.3.4 | 2026-07-20T01:56:30.329Z | user SSRF IPv6 bypass f

OpenClaw

Rank

62

Safety

84

Downloads

2.1k

Updated

Oct 9, 2026

Version

1.3.7

Source

CLAWHUB

About

What it does, and when to use it.

Capability contract not published. No trust telemetry is available yet. 2.1K downloads reported by the source. Last updated 10/9/2026.

Avoid when

  • Contract metadata is missing or unavailable for deterministic execution.

Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing

Public facts

Every fact links back to the source it came from.

Vendor
Clawhubvendor · observed Oct 9, 2026
Protocol compatibility
OpenClawcompatibility · observed Oct 9, 2026
Adoption signal
2.1K downloadsadoption · observed Oct 9, 2026
Latest release
1.3.7release · observed Jul 22, 2026
Handshake status
UNKNOWNsecurity

Install and run

Setup complexity: low.

clawhub skill install s175p518b8g47fx6r9zyvs95ks876t4t:smart-scraper-web
  1. Setup complexity is classified as HIGH. You must provision dedicated cloud infrastructure or an isolated VM. Do not run this directly on your local workstation.
  2. Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.

Contract: missing

curl -s "https://www.xpersona.co/api/v1/agents/clawhub-jlacroix82-smart-scraper-web/snapshot"

Documentation

CLAWHUB

151,468 characters of source documentation, loaded on request.

Extracted files

5 files captured from the source.

SKILL.md

# Smart Scraper

Extract structured data from websites with integrated security protections. Supports tables, lists, prices, articles, metadata extraction, plus change monitoring with structured diffs.

**Security-first design**:
- SSRF protection: blocks private IPs, cloud metadata endpoints, dangerous URL schemes
- Cache is **disabled by default** (opt-in via `--cache` flag)
- Redirect validation prevents SSRF bypass attacks
- Rate-limited requests (100ms minimum interval)
- No dynamic code execution (no `eval`, no `execSync`, no arbitrary `require`)
- Bounded regex: all HTML pattern matching has content length limits

**Cache behavior**:
- Disk caching is **opt-in** — use `--cache` flag to enable
- Default: every request fetches fresh data
- Cache TTL: 5 minutes, max 50 entries, max 10MB total
- Cache location: `memory/scraper-cache/cache.json` (in workspace)
- Privacy warning displayed when caching is activated
- When cache is disabled, no persistent data is written to disk during `--extract`

## ⚠️ Important Warnings

### Watch Mode Persistence
`--watch <url>` writes data to disk regardless of cache setting:
- Baseline snapshots are saved per-watched-URL
- Diff results are stored on each poll interval
- These files persist until manually deleted

### HTTP Connections
This tool allows both `http://` and `https://` URLs. HTTP connections transmit data in cleartext over the network.
- Prefer HTTPS for sensitive scraping targets
- HTTP responses may be intercepted or modified in transit
- SSRF protection applies to both protocols

**Usage modes**:
- `--extract <url>` — extract structured data (tables, lists, prices, articles, or all)
- `--parse <html>` — parse raw HTML and display structure
- `--watch <url>` — monitor for content changes with baseline comparison
- `--status` — show cache statistics

**Programmatic API**: All functions exported as `module.exports` for testability.

skills/smart-scraper-web/SKILL.md

---
name: web-data-extractor
description: Extract structured data from websites. Tables, lists, prices, articles, metadata. Zero external dependencies.
---

# Web Data Extractor 🕷️

> ⚠️ **Security Note** — This skill **sends user-provided URLs over the network**. Do not use with sensitive, authenticated, internal, or attacker-controlled URLs until redirect targets are revalidated.
>
> **Privacy Notice** — Caching is **enabled by default** for performance. A visible warning is shown before first cache write. Use `--no-cache` to disable local persistence of scraped content to `memory/scraper-cache/cache.json`. Each scrape writes **title, headings, paragraphs, links, tables, lists, prices, images, and metadata** to `cache.json`.

**Stop copying data by hand. Start extracting it automatically.**

## The Problem

Web content is everywhere but inaccessible to agents. `web_fetch` gets raw HTML, but you need structure — tables, prices, lists, article text — to make it useful.

Web Data Extractor turns raw HTML into structured data with one command.

## Quick Start

### Extract everything from a page

```bash
node skills/smart-scraper/smart-scraper.js --extract https://example.com
```

Returns title, headings, paragraphs, links, tables, lists, prices, images, and metadata.

### Extract without caching (privacy mode)

```bash
node skills/smart-scraper/smart-scraper.js --extract --no-cache https://example.com
```

Disables local cache persistence — scraped content is not written to disk.

### Extract with caching enabled (default)

```bash
node skills/smart-scraper/smart-scraper.js --extract https://example.com
```

Caching is enabled by default for performance. A visible warning is shown before first cache write.

### Extract tables only

```bash
node skills/smart-scraper/smart-scraper.js --extract --table https://example.com/pricing
```

### Extract lists only

```bash
node skills/smart-scraper/smart-scraper.js --extract --list https://example.com/blog
```

### Extract prices

```bash
node skills/smart-scraper/smart-scraper.js --extract --price https://example.com/products
```

### Extract article content

```bash
node skills/smart-scraper/smart-scraper.js --extract --article https://example.com/blog/post
```

### Parse raw HTML

```bash
node skills/smart-scraper/smart-scraper.js --parse "<html>...</html>"
```

### Status overview

```bash
node skills/smart-scraper/smart-scraper.js --status
```

## Features

### HTML Parsing

- Title extraction
- Heading hierarchy (h1-h6)
- Paragraph extraction (filters short fragments)
- Link extraction with text
- Image extraction with alt text
- Metadata/meta tag extraction

### Table Extraction

- Full table structure with rows and cells
- Handles th and td elements
- Strips nested HTML from cells

### List Extraction

- Both ordered and unordered lists
- List item text extraction
- Preserves list structure

### Price Detection

- Matches USD ($), EUR (€), GBP (£), JPY (¥) formats
- Handles comma-separated thousands

README.md

# Smart Scraper — Structured Web Data Extraction

Extract structured data from websites with zero external dependencies. Built-in SSRF protection, rate limiting, caching, and change monitoring.

## Features

- **Extraction modes**: tables, lists, prices, articles, metadata, or everything
- **HTML parsing**: title, headings (h1–h6), paragraphs, links, images, tables, lists, prices, meta tags
- **Change monitoring**: watch URLs for content changes with diff output
- **Security**: SSRF blocklist, URL validation, redirect limits, rate limiting
- **Caching**: optional disk cache with TTL, size limits, and eviction
- **Zero dependencies**: uses only Node.js built-in modules

## Usage

```bash
# Extract structured data from a URL
node smart-scraper.js --extract https://example.com
node smart-scraper.js --extract --all https://example.com

# Extract specific content types
node smart-scraper.js --extract --table https://example.com
node smart-scraper.js --extract --list https://example.com
node smart-scraper.js --extract --price https://example.com
node smart-scraper.js --extract --article https://example.com

# Parse raw HTML
node smart-scraper.js --parse "<html><title>Hello</title></html>"

# Change monitoring
node smart-scraper.js --watch https://example.com           # First run: capture baseline
node smart-scraper.js --watch https://example.com            # Second run: compare
node smart-scraper.js --watch https://example.com --interval 300  # Poll every 5 min
node smart-scraper.js --watch https://example.com --alert-on-change  # CI mode

# Cache control
node smart-scraper.js --extract https://example.com --cache  # Enable disk caching
node smart-scraper.js --extract https://example.com --no-cache  # Disable caching

# Status
node smart-scraper.js --status
```

## API (for programmatic use)

```javascript
const SS = require('./smart-scraper.js');

// URL validation
const result = SS.validateUrl('https://example.com');
console.log(result.valid); // true

// Parse HTML
const data = SS.parseHtml('<html>...</html>');
console.log(data.title, data.headings, data.paragraphs);

// Extract from URL (async)
const extracted = await SS.extractFromUrl('https://example.com');

// Diff two snapshots
const changes = SS.diffSnapshots(oldData, newData);

// Watch mode (async)
const exitCode = await SS.watchMode(url, interval, alertOnChange, diffOnly);
```

## Security

- **SSRF protection**: blocks private IPs, localhost, cloud metadata endpoints
- **Blocked schemes**: `file:`, `gopher:`, `data:`, `javascript:`, `ftp:`
- **Redirect validation**: re-validates redirect targets to prevent SSRF bypass
- **Rate limiting**: 100ms minimum delay between requests
- **Cache opt-in**: disk caching is disabled by default; requires `--cache` flag
- **No dynamic evaluation**: no `eval()`, no `execSync()`, no `require()` of user input
- **Bounded regex**: content length limits on all HTML regex operations

## Testing

```bash
node test/run-tests.js
```

Runs 36 tests covering URL va

skills/smart-scraper-web/_meta.json

{
  "ownerId": "kn7b6eyf5vc7khg5fr63pjm8xd82qvw5",
  "slug": "smart-scraper-web",
  "version": "1.2.1",
  "publishedAt": 1780786140000
}

_meta.json

{
  "ownerId": "kn7b6eyf5vc7khg5fr63pjm8xd82qvw5",
  "slug": "smart-scraper-web",
  "version": "1.3.7",
  "publishedAt": 1784737602724
}
Github ReposUpdated 4h agoRank 70

AionUi

Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!

MCPOPENCLAW
Github ReposUpdated 6mo agoRank 70

activepieces

AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents

OPENCLAW
Github ReposUpdated 6mo agoRank 70

cherry-studio

AI productivity studio with smart chat, autonomous agents, and 300+ assistants.

MCPOPENCLAW
Github ReposUpdated 7mo agoRank 70

CopilotKit

The Frontend for Agents & Generative UI. React + Angular

OPENCLAW

Machine-readable data

The same record, as JSON, for agents and crawlers.

{
  "facts": [
    {
      "factKey": "vendor",
      "category": "vendor",
      "label": "Vendor",
      "value": "Clawhub",
      "href": "https://clawhub.ai/jlacroix82/skills/smart-scraper-web",
      "sourceUrl": "https://clawhub.ai/jlacroix82/skills/smart-scraper-web",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-09T19:29:59.254Z",
      "isPublic": true
    },
    {
      "factKey": "protocols",
      "category": "compatibility",
      "label": "Protocol compatibility",
      "value": "OpenClaw",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-jlacroix82-smart-scraper-web/contract",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-jlacroix82-smart-scraper-web/contract",
      "sourceType": "contract",
      "confidence": "medium",
      "observedAt": "2026-10-09T19:29:59.254Z",
      "isPublic": true
    },
    {
      "factKey": "traction",
      "category": "adoption",
      "label": "Adoption signal",
      "value": "2.1K downloads",
      "href": "https://clawhub.ai/jlacroix82/smart-scraper-web",
      "sourceUrl": "https://clawhub.ai/jlacroix82/smart-scraper-web",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-09T19:29:59.254Z",
      "isPublic": true
    },
    {
      "factKey": "latest_release",
      "category": "release",
      "label": "Latest release",
      "value": "1.3.7",
      "href": "https://clawhub.ai/jlacroix82/smart-scraper-web",
      "sourceUrl": "https://clawhub.ai/jlacroix82/smart-scraper-web",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-07-22T16:26:42.724Z",
      "isPublic": true
    },
    {
      "factKey": "handshake_status",
      "category": "security",
      "label": "Handshake status",
      "value": "UNKNOWN",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-jlacroix82-smart-scraper-web/trust",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-jlacroix82-smart-scraper-web/trust",
      "sourceType": "trust",
      "confidence": "medium",
      "observedAt": null,
      "isPublic": true
    }
  ],
  "events": [
    {
      "eventType": "release",
      "title": "Release 1.3.7",
      "description": "Security: web scraping risk disclosures, URL validation warnings",
      "href": "https://clawhub.ai/jlacroix82/smart-scraper-web",
      "sourceUrl": "https://clawhub.ai/jlacroix82/smart-scraper-web",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-07-22T16:26:42.724Z",
      "isPublic": true
    }
  ]
}

Record generated Oct 9, 2026.

Sponsored

Ads related to Smart Scraper and adjacent AI workflows.