Claim this agent
agentCLAWHUBUnverified

Crawl

Crawl any website and save pages as local markdown files. Use when you need to download documentation, knowledge bases, or web content for offline access or analysis. No code required - just provide a URL. Skill: Crawl Owner: barneyjm Summary: Crawl any website and save pages as local markdown files. Use when you need to download documentation, knowledge bases, or web content for offline access or analysis. No code required - just provide a URL. Tags: latest:0.1.0 Version history: v0.1.0 | 2026-02-03T15:26:18.843Z | auto Initial release of the Crawl skill—extract content from websites as markdown for offline access or

OpenClaw

Rank

62

Safety

84

Downloads

2.2k

Updated

Oct 9, 2026

Version

0.1.0

Source

CLAWHUB

About

What it does, and when to use it.

Capability contract not published. No trust telemetry is available yet. 2.2K downloads reported by the source. Last updated 10/9/2026.

Avoid when

  • Contract metadata is missing or unavailable for deterministic execution.

Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing

Public facts

Every fact links back to the source it came from.

Vendor
Clawhubvendor · observed Oct 9, 2026
Protocol compatibility
OpenClawcompatibility · observed Oct 9, 2026
Adoption signal
2.2K downloadsadoption · observed Oct 9, 2026
Latest release
0.1.0release · observed Feb 3, 2026
Handshake status
UNKNOWNsecurity

Install and run

Setup complexity: low.

clawhub skill install s175g00ed7n2ej0r13jbm8m3k18842ww:crawl
  1. Setup complexity is LOW. This package is likely designed for quick installation with minimal external side-effects.
  2. Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.

Contract: missing

curl -s "https://www.xpersona.co/api/v1/agents/clawhub-barneyjm-crawl/snapshot"

Documentation

CLAWHUB

7,812 characters of source documentation, loaded on request.

Extracted files

2 files captured from the source.

SKILL.md

---
name: crawl
description: "Crawl any website and save pages as local markdown files. Use when you need to download documentation, knowledge bases, or web content for offline access or analysis. No code required - just provide a URL."
---

# Crawl Skill

Crawl websites to extract content from multiple pages. Ideal for documentation, knowledge bases, and site-wide content extraction.

## Prerequisites

**Tavily API Key Required** - Get your key at https://tavily.com

Add to `~/.claude/settings.json`:
```json
{
  "env": {
    "TAVILY_API_KEY": "tvly-your-api-key-here"
  }
}
```

## Quick Start

### Using the Script

```bash
./scripts/crawl.sh '<json>' [output_dir]
```

**Examples:**
```bash
# Basic crawl
./scripts/crawl.sh '{"url": "https://docs.example.com"}'

# Deeper crawl with limits
./scripts/crawl.sh '{"url": "https://docs.example.com", "max_depth": 2, "limit": 50}'

# Save to files
./scripts/crawl.sh '{"url": "https://docs.example.com", "max_depth": 2}' ./docs

# Focused crawl with path filters
./scripts/crawl.sh '{"url": "https://example.com", "max_depth": 2, "select_paths": ["/docs/.*", "/api/.*"], "exclude_paths": ["/blog/.*"]}'

# With semantic instructions (for agentic use)
./scripts/crawl.sh '{"url": "https://docs.example.com", "instructions": "Find API documentation", "chunks_per_source": 3}'
```

When `output_dir` is provided, each crawled page is saved as a separate markdown file.

### Basic Crawl

```bash
curl --request POST \
  --url https://api.tavily.com/crawl \
  --header "Authorization: Bearer $TAVILY_API_KEY" \
  --header 'Content-Type: application/json' \
  --data '{
    "url": "https://docs.example.com",
    "max_depth": 1,
    "limit": 20
  }'
```

### Focused Crawl with Instructions

```bash
curl --request POST \
  --url https://api.tavily.com/crawl \
  --header "Authorization: Bearer $TAVILY_API_KEY" \
  --header 'Content-Type: application/json' \
  --data '{
    "url": "https://docs.example.com",
    "max_depth": 2,
    "instructions": "Find API documentation and code examples",
    "chunks_per_source": 3,
    "select_paths": ["/docs/.*", "/api/.*"]
  }'
```

## API Reference

### Endpoint

```
POST https://api.tavily.com/crawl
```

### Headers

| Header | Value |
|--------|-------|
| `Authorization` | `Bearer <TAVILY_API_KEY>` |
| `Content-Type` | `application/json` |

### Request Body

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `url` | string | Required | Root URL to begin crawling |
| `max_depth` | integer | 1 | Levels deep to crawl (1-5) |
| `max_breadth` | integer | 20 | Links per page |
| `limit` | integer | 50 | Total pages cap |
| `instructions` | string | null | Natural language guidance for focus |
| `chunks_per_source` | integer | 3 | Chunks per page (1-5, requires instructions) |
| `extract_depth` | string | `"basic"` | `basic` or `advanced` |
| `format` | string | `"markdown"` | `markdown` or `text` |
| `select_paths` | array | null | Regex patterns to include |
| 

_meta.json

{
  "ownerId": "kn7e3j3w1x5et0yppacy4m90y18084ba",
  "slug": "crawl",
  "version": "0.1.0",
  "publishedAt": 1770132378843
}
Github ReposUpdated 1h agoRank 70

AionUi

Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!

MCPOPENCLAW
Github ReposUpdated 6mo agoRank 70

activepieces

AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents

OPENCLAW
Github ReposUpdated 6mo agoRank 70

cherry-studio

AI productivity studio with smart chat, autonomous agents, and 300+ assistants.

MCPOPENCLAW
Github ReposUpdated 7mo agoRank 70

CopilotKit

The Frontend for Agents & Generative UI. React + Angular

OPENCLAW

Machine-readable data

The same record, as JSON, for agents and crawlers.

{
  "facts": [
    {
      "factKey": "vendor",
      "category": "vendor",
      "label": "Vendor",
      "value": "Clawhub",
      "href": "https://clawhub.ai/barneyjm/skills/crawl",
      "sourceUrl": "https://clawhub.ai/barneyjm/skills/crawl",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-09T17:29:54.797Z",
      "isPublic": true
    },
    {
      "factKey": "protocols",
      "category": "compatibility",
      "label": "Protocol compatibility",
      "value": "OpenClaw",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-barneyjm-crawl/contract",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-barneyjm-crawl/contract",
      "sourceType": "contract",
      "confidence": "medium",
      "observedAt": "2026-10-09T17:29:54.797Z",
      "isPublic": true
    },
    {
      "factKey": "traction",
      "category": "adoption",
      "label": "Adoption signal",
      "value": "2.2K downloads",
      "href": "https://clawhub.ai/barneyjm/crawl",
      "sourceUrl": "https://clawhub.ai/barneyjm/crawl",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-09T17:29:54.797Z",
      "isPublic": true
    },
    {
      "factKey": "latest_release",
      "category": "release",
      "label": "Latest release",
      "value": "0.1.0",
      "href": "https://clawhub.ai/barneyjm/crawl",
      "sourceUrl": "https://clawhub.ai/barneyjm/crawl",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-02-03T15:26:18.843Z",
      "isPublic": true
    },
    {
      "factKey": "handshake_status",
      "category": "security",
      "label": "Handshake status",
      "value": "UNKNOWN",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-barneyjm-crawl/trust",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-barneyjm-crawl/trust",
      "sourceType": "trust",
      "confidence": "medium",
      "observedAt": null,
      "isPublic": true
    }
  ],
  "events": [
    {
      "eventType": "release",
      "title": "Release 0.1.0",
      "description": "Initial release of the Crawl skill—extract content from websites as markdown for offline access or analysis. - Crawl any website using a simple script or direct curl calls to the Tavily API. - Save multiple pages locally as markdown files with flexible filtering options (by path, depth, limit). - Supports focused crawling using natural language instructions and path selection. - Offers specialized agentic mode for LLM context ingestion by returning relevant content chunks. - Includes clear API examples, guidance, and best practices for both data collection and agentic use.",
      "href": "https://clawhub.ai/barneyjm/crawl",
      "sourceUrl": "https://clawhub.ai/barneyjm/crawl",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-02-03T15:26:18.843Z",
      "isPublic": true
    }
  ]
}

Record generated Oct 9, 2026.

Sponsored

Ads related to Crawl and adjacent AI workflows.