Crawl4AI Web Crawler
Use Crawl4AI for web scraping and content extraction. Use when users need to scrape web content, extract structured data, convert web pages to Markdown, perf... Skill: Crawl4AI Web Crawler Owner: openlark Summary: Use Crawl4AI for web scraping and content extraction. Use when users need to scrape web content, extract structured data, convert web pages to Markdown, perf... Tags: latest:1.0.1 Version history: v1.0.1 | 2026-05-12T01:15:47.164Z | user Version 1.0.1 - Renamed and rebranded the skill to "crawl4ai-web-crawler" with updated purpose and trigger words. - Replaced the
Rank
62
Safety
84
Downloads
1.3k
Updated
Oct 10, 2026
Version
1.0.1
Source
CLAWHUB
About
What it does, and when to use it.
Capability contract not published. No trust telemetry is available yet. 1.3K downloads reported by the source. Last updated 10/10/2026.
Avoid when
- Contract metadata is missing or unavailable for deterministic execution.
Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing
Public facts
Every fact links back to the source it came from.
- Vendor
- Clawhubvendor · observed Oct 10, 2026
- Protocol compatibility
- OpenClawcompatibility · observed Oct 10, 2026
- Adoption signal
- 1.3K downloadsadoption · observed Oct 10, 2026
- Latest release
- 1.0.1release · observed May 12, 2026
- Handshake status
- UNKNOWNsecurity
Install and run
Setup complexity: low.
clawhub skill install s1727wv2g20pc729snzcm4nf8183hy72:crawl4ai-web-crawler- Setup complexity is LOW. This package is likely designed for quick installation with minimal external side-effects.
- Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.
Contract: missing
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-openlark-crawl4ai-web-crawler/snapshot"
Documentation
CLAWHUB
48,114 characters of source documentation, loaded on request.
Extracted files
4 files captured from the source.
SKILL.md
---
name: crawl4ai-web-crawler
description: Use Crawl4AI for web scraping and content extraction. Use when users need to scrape web content, extract structured data, convert web pages to Markdown, perform batch crawling, or use AI-driven web data collection.
---
# Crawl4AI Web Crawler
[Crawl4AI](https://github.com/unclecode/crawl4ai) is an open-source, LLM-friendly web crawler on GitHub that converts web pages into clean Markdown or structured JSON, ideal for RAG, AI Agents, and data pipelines.
For detailed API parameters, see [references/api-reference.md](references/api-reference.md).
## Trigger Words
"scrape," "crawl," "crawl," "extract webpage," "convert webpage to markdown," "structured extraction," etc.
## Installation
```bash
pip install -U crawl4ai
crawl4ai-setup # Automatically installs the Playwright browser
crawl4ai-doctor # Verifies the installation
```
If the browser installation fails, run manually:
```bash
python -m playwright install --with-deps chromium
```
## Core Architecture
Three core classes:
| Class | Purpose |
|-------|---------|
| `AsyncWebCrawler` | Main async crawler class, manages the browser lifecycle |
| `BrowserConfig` | Browser settings (headless, UA, proxy, viewport, etc.) |
| `CrawlerRunConfig` | Per-crawl settings (cache, extraction strategy, JS, screenshots, etc.) |
## Basic Usage
### Simplest Crawl
```python
import asyncio
from crawl4ai import AsyncWebCrawler
async def main():
async with AsyncWebCrawler() as crawler:
result = await crawler.arun(url="https://example.com")
print(result.markdown) # LLM-ready Markdown
asyncio.run(main())
```
### Crawl with Configuration
```python
from crawl4ai import AsyncWebCrawler, BrowserConfig, CrawlerRunConfig, CacheMode
browser_cfg = BrowserConfig(headless=True, verbose=True)
run_cfg = CrawlerRunConfig(
cache_mode=CacheMode.BYPASS, # BYPASS=no cache, ENABLED=enable, WRITE_ONLY, READ_ONLY
css_selector="main.article", # Only extract the specified area
word_count_threshold=10, # Filter out short text blocks
screenshot=True, # Take a screenshot
)
async with AsyncWebCrawler(config=browser_cfg) as crawler:
result = await crawler.arun(url="https://example.com", config=run_cfg)
print(result.markdown)
if result.screenshot:
print(f"Screenshot: {len(result.screenshot)} bytes base64")
```
### Command Line Tool
```bash
# Basic crawl
crwl https://example.com -o markdown
# Deep crawl (BFS, up to 10 pages)
crwl https://docs.crawl4ai.com --deep-crawl bfs --max-pages 10
# LLM extraction
crwl https://example.com/products -q "Extract all product prices"
```
## Markdown Generation
### Using Content Filters
Raw Markdown is generated by default. Use `DefaultMarkdownGenerator` + content filters to get cleaner output:
```python
from crawl4ai.content_filter_strategy import PruningContentFilter, BM25ContentFilter
from crawl4ai.markdown_generation_strategy import DefaultM_meta.json
{
"ownerId": "kn75qrtb885pznwsjwsh18dvf1813bv0",
"slug": "crawl4ai-web-crawler",
"version": "1.0.1",
"publishedAt": 1778548547164
}references/api-reference.md
# Crawl4AI API Reference
Detailed class and parameter reference. Read the main SKILL.md first before using this reference.
## BrowserConfig
```python
BrowserConfig(
browser_type: str = "chromium", # "chromium" | "firefox" | "webkit"
headless: bool = True,
viewport_width: int = 1080,
viewport_height: int = 600,
user_agent: Optional[str] = None, # Custom User-Agent
proxy: Optional[str] = None, # Proxy URL
proxy_config: Optional[dict] = None, # Advanced proxy configuration
use_managed_browser: bool = False, # Use an existing browser (anti-detection)
user_data_dir: Optional[str] = None, # Browser profile path
channel: Optional[str] = None, # Playwright channel
ignore_https_errors: bool = True,
java_script_enabled: bool = True,
cookies: list = [], # Pre-set cookies
headers: dict = {}, # Additional HTTP headers
accept_downloads: bool = False,
downloads_path: Optional[str] = None,
storage_state: Optional[str] = None, # Authentication state file path
text_mode: bool = False, # Disable image loading (faster)
light_mode: bool = False, # Lightweight mode
verbose: bool = True,
extra_args: Optional[list] = None, # Additional browser launch arguments
cdp_url: Optional[str] = None, # Connect to remote Chrome DevTools
)
```
## CrawlerRunConfig
```python
CrawlerRunConfig(
# Cache
cache_mode: CacheMode = CacheMode.BYPASS, # BYPASS | ENABLED | WRITE_ONLY | READ_ONLY
# Content Extraction
css_selector: Optional[str] = None, # Only extract matching CSS regions
word_count_threshold: int = 10, # Filter short text (discard if below this value)
excluded_tags: list = [], # Excluded HTML tags
excluded_selector: Optional[str] = None, # Excluded CSS selector
keep_data_attributes: bool = False, # Preserve data-* attributes
remove_forms: bool = False,
remove_overlay_elements: bool = False, # Remove pop-ups/masks
# Markdown Generation
markdown_generator: Optional[DefaultMarkdownGenerator] = None,
# Extraction Strategy
extraction_strategy: Optional = None, # JsonCssExtractionStrategy | LLMExtractionStrategy
# Dynamic Pages
js_code: Optional[list] = None, # List of JS code snippets to execute
js_only: bool = False, # Use JS only (do not load HTML)
wait_for: Optional[str] = None, # CSS selector to wait for
wait_for_images: bool = False, # Wait for images to load
delay_before_return_html: float = 0.0, # Additional wait in seconds before returning
page_timeout: int = 60000, # Page load timeout (ms)
# Media
screenshot: bool = False,
screenshot_wait_for: Optional[float] = None, # Wait skill-card.md
## Description: Use Crawl4AI for web scraping and content extraction, including structured data extraction, Markdown conversion, batch crawling, and AI-driven web data collection. This skill is ready for commercial/non-commercial use. ## Publisher: [openlark](https://clawhub.ai/user/openlark) ### License/Terms of Use: MIT-0 ## Use Case: Developers and engineers use this skill to configure and run Crawl4AI for web scraping, content-to-Markdown conversion, structured JSON extraction, batch crawling, deep crawling, and optional LLM-assisted extraction. ### Deployment Geography for Use: Global ## Known Risks and Mitigations: Risk: Crawler and browser tooling can access broad network targets, persistent browser profiles, downloads, and cached page content. Mitigation: Use an isolated environment, avoid private or authenticated sites unless explicitly intended, avoid normal browser profiles, and keep downloads and caches in disposable directories. Risk: Unpinned package and Docker image installation can change behavior across releases. Mitigation: Pin the Crawl4AI package version and Docker image digest before use. Risk: Remote LLM extraction can disclose crawled page content to the configured provider. Mitigation: Use remote LLM extraction only for content approved for that provider, or use a local model when page content should remain local. ## Reference(s): - [Crawl4AI API Reference](references/api-reference.md) - [Crawl4AI GitHub Repository](https://github.com/unclecode/crawl4ai) - [Crawl4AI Documentation](https://docs.crawl4ai.com) ## Skill Output: **Output Type(s):** [text, markdown, code, shell commands, configuration, guidance] **Output Format:** [Markdown guidance with Python and bash code blocks] **Output Parameters:** [1D] **Other Properties Related to Output:** [May include Crawl4AI installation commands, crawler configuration, API usage examples, and safety scoping guidance.] ## Skill Version(s): 1.0.1 (source: server release evidence) ## Ethical Considerations: Users should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.
AionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
activepieces
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
CopilotKit
The Frontend for Agents & Generative UI. React + Angular
Machine-readable data
The same record, as JSON, for agents and crawlers.
{
"facts": [
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Clawhub",
"href": "https://clawhub.ai/openlark/skills/crawl4ai-web-crawler",
"sourceUrl": "https://clawhub.ai/openlark/skills/crawl4ai-web-crawler",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-10T18:23:08.768Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-openlark-crawl4ai-web-crawler/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-openlark-crawl4ai-web-crawler/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-10T18:23:08.768Z",
"isPublic": true
},
{
"factKey": "traction",
"category": "adoption",
"label": "Adoption signal",
"value": "1.3K downloads",
"href": "https://clawhub.ai/openlark/crawl4ai-web-crawler",
"sourceUrl": "https://clawhub.ai/openlark/crawl4ai-web-crawler",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-10T18:23:08.768Z",
"isPublic": true
},
{
"factKey": "latest_release",
"category": "release",
"label": "Latest release",
"value": "1.0.1",
"href": "https://clawhub.ai/openlark/crawl4ai-web-crawler",
"sourceUrl": "https://clawhub.ai/openlark/crawl4ai-web-crawler",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-05-12T01:15:47.164Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-openlark-crawl4ai-web-crawler/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-openlark-crawl4ai-web-crawler/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
],
"events": [
{
"eventType": "release",
"title": "Release 1.0.1",
"description": "Version 1.0.1 - Renamed and rebranded the skill to \"crawl4ai-web-crawler\" with updated purpose and trigger words. - Replaced the prior RAGFlow documentation with comprehensive Crawl4AI usage guide, including installation, core classes, and real-world examples. - Added API reference pointer: references/api-reference.md. - Removed prior references: architecture.md, cli-reference.md, and deployment.md, focusing documentation on Crawl4AI features and usage.",
"href": "https://clawhub.ai/openlark/crawl4ai-web-crawler",
"sourceUrl": "https://clawhub.ai/openlark/crawl4ai-web-crawler",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-05-12T01:15:47.164Z",
"isPublic": true
}
]
}Record generated Oct 10, 2026.
