agentCLAWHUBUnverified

crawlora-datasets

Queries Crawlora's pre-built hosted datasets — Airbnb markets, App Store/Google Play apps, GitHub/Instagram/X users, job postings, US housing markets, Google Maps businesses, Goodreads, PitchBook, Steam, TrustMRR, Product Hunt, SEC companies, tech-stack, and more — via search/facets/item/nearby endpoints, returning clean JSON without live-crawling each platform. Use when the user wants bulk or aggregate analysis, to search a pre-indexed corpus, to facet/filter a large population, or to look up one record by its dataset id, instead of scraping pages one at a time. Skill: crawlora-datasets Owner: crawlora-org Summary: Queries Crawlora's pre-built hosted datasets — Airbnb markets, App Store/Google Play apps, GitHub/Instagram/X users, job postings, US housing markets, Google Maps businesses, Goodreads, PitchBook, Steam, TrustMRR, Product Hunt, SEC companies, tech-stack, and more — via search/facets/item/nearby endpoints, returning clean JSON without live-crawling each platform. U

OpenClaw

Rank

62

Safety

84

Downloads

1.2k

Updated

Oct 11, 2026

Version

1.0.20

Source

CLAWHUB

About

What it does, and when to use it.

Capability contract not published. No trust telemetry is available yet. 1.2K downloads reported by the source. Last updated 10/11/2026.

Avoid when

  • Contract metadata is missing or unavailable for deterministic execution.

Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing

Public facts

Every fact links back to the source it came from.

Vendor
Clawhubvendor · observed Oct 11, 2026
Protocol compatibility
OpenClawcompatibility · observed Oct 11, 2026
Adoption signal
1.2K downloadsadoption · observed Oct 11, 2026
Latest release
1.0.20release · observed Oct 5, 2026
Handshake status
UNKNOWNsecurity

Install and run

Setup complexity: low.

clawhub skill install s17d53nb8nd03gyyfdy32rgde58e574f:crawlora-datasets
  1. Setup complexity is classified as HIGH. You must provision dedicated cloud infrastructure or an isolated VM. Do not run this directly on your local workstation.
  2. Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.

Contract: missing

curl -s "https://www.xpersona.co/api/v1/agents/clawhub-crawlora-org-crawlora-datasets/snapshot"

Documentation

CLAWHUB

145,452 characters of source documentation, loaded on request.

Extracted files

4 files captured from the source.

SKILL.md

---
name: crawlora-datasets
description: Queries Crawlora's pre-built hosted datasets — Airbnb markets, App Store/Google Play apps, GitHub/Instagram/X users, job postings, US housing markets, Google Maps businesses, Goodreads, PitchBook, Steam, TrustMRR, Product Hunt, SEC companies, tech-stack, and more — via search/facets/item/nearby endpoints, returning clean JSON without live-crawling each platform. Use when the user wants bulk or aggregate analysis, to search a pre-indexed corpus, to facet/filter a large population, or to look up one record by its dataset id, instead of scraping pages one at a time.
---

# Crawlora hosted datasets

Query Crawlora's own **pre-crawled, pre-indexed datasets** — search, facet, and
fetch-by-id over corpora Crawlora already built and refreshes on a schedule.
This is different from the other skills in this repo: those hit a live
per-platform endpoint (one request, one page); this skill hits a **search
index** over millions of already-collected records, so it's the right tool
for population-level questions ("how many", "top N by X", "everything
matching Y") rather than one-off lookups.

## When to use this skill

- "How many / what share of X match Y?" — facet/aggregate questions.
- "Find all X with property Y" (e.g. jobs paying > $150k, apps with 4.5+
  rating, GitHub users near a city, houses in a metro).
- "Give me the full list of Z" instead of one record — bulk/list research.
- Any of: Airbnb markets, app-store apps/reviews/charts, GitHub/Instagram/X
  users, job postings + which companies are hiring, US housing markets
  (Redfin-sourced), Google Maps businesses, Goodreads authors/books, Apple
  Podcasts shows, Chrome Web Store extensions, PitchBook companies/funds/
  investors/advisors/LPs, PlayStation games, Product Hunt makers/products/
  trends, Reddit trending, SEC companies + institutional positions, Steam
  games/prices/playercounts/reviews/news/achievements/charts, TrustMRR
  startups, journalists, Numbeo cost-of-living cities/countries, website
  tech-stack.
- Prefer the platform-specific skill instead when the job is "look up this
  one profile/listing right now" (e.g. `youtube-research`, `movie-tv-research`)
  — datasets are refreshed periodically, not real-time.

## Setup (one-time)

- Get a free Crawlora API key (2,000 credits/mo, no card) at [https://crawlora.net](https://crawlora.net?utm_source=github&utm_medium=referral&utm_campaign=crawlora-skills).
- Set `CRAWLORA_API_KEY` in the environment before running the helper.
- The helper reads `CRAWLORA_API_KEY` from the environment and sends requests to `https://api.crawlora.net/api/v1`. Missing/invalid key → `401`.

## How it works

Every dataset follows the same shape under `/datasets/<dataset-id>/...`:

1. **Discover** — `GET /datasets` lists every available dataset id and its
   capabilities (search / facets / item / nearby).
2. **Search** — `GET /datasets/<id>/search` full-text + filtered search;
   paginate with `page`/`size` (see `reference/en

_meta.json

{
  "ownerId": "kn70shhkf6qpfwgfrbgtep2wkd8c6b4t",
  "slug": "crawlora-datasets",
  "version": "1.0.20",
  "publishedAt": 1791162745933
}

reference/endpoints.md

# crawlora-datasets — endpoint reference

> Generated from `scripts/tools.json` by `scripts/generate.mjs` — do not edit by hand.

Endpoints this skill uses, grouped by platform. Call them via `scripts/crawlora.sh` (see SKILL.md).

All paths are relative to the API base `https://api.crawlora.net/api/v1` and require the header `x-api-key: $CRAWLORA_API_KEY`. Path params like `{id}` are substituted into the URL; `GET` params go in the query string; `POST` params go in a JSON body.

**130 endpoints across 1 platform group(s).**

## Datasets (130)

### `datasets_airbnb_facets`

- **HTTP:** `GET /datasets/airbnb-markets/facets`
- **What:** Facet the Airbnb markets dataset. Returns suppressed distribution counts over the Airbnb markets dataset, honoring the same filters as search. Facet enum: `country`, `market`, `currency`, `superhost`, `guest_favorite`, `rating_band`, `review_band`, `admin1` (top subdivision), `locality` (settlement), `room_type` (`entire_place`/`private_room`/`hotel`/`shared_room`), `property_type` (Airbnb's canonical listing type from the detail page), `amenities` (each amenity with the count of listings offering it). The `admin1`, `locality`, `room_type`, `property_type` and `amenities` facets stay empty until their enrichment coverage is high enough to be reliable. group_by enum: `country`, `market`, `admin1`, `locality`, `room_type`, `property_type`.
- **Params:** `active_since` (string, optional) — Freshness filter, an ISO-8601 date (YYYY-MM-DD); `country` (string, optional) — Exact ISO-3166-1 alpha-2 country filter, e.g. FR; `facet` (string, **required**) — Facet enum: country, market, currency, superhost, guest_favorite, rating_band, review_band, admin1, locality, room_type, property_type, amenities; `group_by` (string, optional) — Aggregate cell dimension enum: country, market, admin1, locality, room_type, property_type. Defaults to country; `guest_favorite` (boolean, optional) — Count only Guest Favorite listings (an observed lower bound; the badge under-counts); `market` (string, optional) — Exact metro-market filter, max 128 characters; `min_listings` (integer, optional) — Minimum listings per bucket; raises the small-cell suppression floor; `min_rating` (number, optional) — Minimum listing rating, from 0 through 5; `min_review_count` (integer, optional) — Minimum listing review count, 0 or greater; `superhost` (boolean, optional) — Count only Superhost listings

### `datasets_airbnb_item`

- **HTTP:** `GET /datasets/airbnb-markets/items/{country}`
- **What:** Get an Airbnb market from the dataset. Returns one country's full aggregate Airbnb market profile from dataset id enum value `airbnb-markets` — headline supply, Superhost share, Guest Favorite share (`guest_favorite_pct`, an observed lower bound), `avg_person_capacity` (average guests a listing sleeps over the detail-page-enriched sample), ratings, its top metros, bounding box, per-currency nightly-price percentiles, and a USD-normalized `price_usd` percentile block 

skill-card.md

## Description:

Queries Crawlora's hosted public datasets for searches, aggregate breakdowns, nearby records, and individual records, returning JSON without crawling each source live.

This skill is ready for commercial/non-commercial use.

## Publisher:

[crawlora-org](https://clawhub.ai/user/crawlora-org)

### License/Terms of Use:

MIT-0

## Use Case:

Developers and researchers use this skill to search and analyze Crawlora's pre-indexed public datasets, compare populations through facets, and retrieve records by dataset ID.

### Deployment Geography for Use:

Global

## Known Risks and Mitigations:

Risk: Queries and the API key are sent to Crawlora's hosted API.

Mitigation: Use a dedicated Crawlora key and keep secrets out of query text.

Risk: Public-contact and business records may require careful handling.

Mitigation: Use returned data in accordance with applicable policy and law.

## Reference(s):

- [Crawlora dataset endpoint reference](reference/endpoints.md)
- [Crawlora](https://crawlora.net)

## Skill Output:

**Output Type(s):** [JSON, Text]

**Output Format:** [JSON API responses and concise text summaries]

**Output Parameters:** [1D]

**Other Properties Related to Output:** [Paginated results; hosted datasets are refreshed periodically rather than in real time.]

## Skill Version(s):

1.0.20 (source: server-resolved release metadata)

## Ethical Considerations:

Users should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.
Github ReposUpdated 1d agoRank 70

AionUi

Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!

MCPOPENCLAW
Github ReposUpdated 6mo agoRank 70

activepieces

AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents

OPENCLAW
Github ReposUpdated 6mo agoRank 70

cherry-studio

AI productivity studio with smart chat, autonomous agents, and 300+ assistants.

MCPOPENCLAW
Github ReposUpdated 7mo agoRank 70

CopilotKit

The Frontend for Agents & Generative UI. React + Angular

OPENCLAW

Machine-readable data

The same record, as JSON, for agents and crawlers.

{
  "facts": [
    {
      "factKey": "vendor",
      "category": "vendor",
      "label": "Vendor",
      "value": "Clawhub",
      "href": "https://clawhub.ai/crawlora-org/skills/crawlora-datasets",
      "sourceUrl": "https://clawhub.ai/crawlora-org/skills/crawlora-datasets",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-11T04:01:55.945Z",
      "isPublic": true
    },
    {
      "factKey": "protocols",
      "category": "compatibility",
      "label": "Protocol compatibility",
      "value": "OpenClaw",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-crawlora-org-crawlora-datasets/contract",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-crawlora-org-crawlora-datasets/contract",
      "sourceType": "contract",
      "confidence": "medium",
      "observedAt": "2026-10-11T04:01:55.945Z",
      "isPublic": true
    },
    {
      "factKey": "traction",
      "category": "adoption",
      "label": "Adoption signal",
      "value": "1.2K downloads",
      "href": "https://clawhub.ai/crawlora-org/crawlora-datasets",
      "sourceUrl": "https://clawhub.ai/crawlora-org/crawlora-datasets",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-11T04:01:55.945Z",
      "isPublic": true
    },
    {
      "factKey": "latest_release",
      "category": "release",
      "label": "Latest release",
      "value": "1.0.20",
      "href": "https://clawhub.ai/crawlora-org/crawlora-datasets",
      "sourceUrl": "https://clawhub.ai/crawlora-org/crawlora-datasets",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-10-05T01:12:25.933Z",
      "isPublic": true
    },
    {
      "factKey": "handshake_status",
      "category": "security",
      "label": "Handshake status",
      "value": "UNKNOWN",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-crawlora-org-crawlora-datasets/trust",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-crawlora-org-crawlora-datasets/trust",
      "sourceType": "trust",
      "confidence": "medium",
      "observedAt": null,
      "isPublic": true
    }
  ],
  "events": [
    {
      "eventType": "release",
      "title": "Release 1.0.20",
      "description": "Sync skill instructions, references, and helper from GitHub 83bb98ef1362f25cecb5ddd4bc1ea0e564555d97",
      "href": "https://clawhub.ai/crawlora-org/crawlora-datasets",
      "sourceUrl": "https://clawhub.ai/crawlora-org/crawlora-datasets",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-10-05T01:12:25.933Z",
      "isPublic": true
    }
  ]
}

Record generated Oct 11, 2026.

Sponsored

Ads related to crawlora-datasets and adjacent AI workflows.