Discovery
Automatically discover novel, statistically validated patterns in tabular data. Find insights you'd otherwise miss, far faster and cheaper than doing it yourself (or prompting an agent to do it). Disco systematically searches for feature interactions, subgroup effects, and conditional relationships you wouldn't think to look for, validates each on hold-out data with FDR-corrected p-values, and checks every finding against academic literature for novelty. Returns structured patterns with conditions, effect sizes, citations, and novelty scores. Skill: Discovery Owner: jessicarumbelow Summary: Automatically discover novel, statistically validated patterns in tabular data. Find insights you'd otherwise miss, far faster and cheaper than doing it yourself (or prompting an agent to do it). Disco systematically searches for feature interactions, subgroup effects, and conditional relationships you wouldn't think to look for, validates each on hold-out data with FD
Rank
62
Safety
84
Downloads
4.8k
Updated
Oct 9, 2026
Version
0.2.182
Source
CLAWHUB
About
What it does, and when to use it.
Capability contract not published. No trust telemetry is available yet. 4.8K downloads reported by the source. Last updated 10/9/2026.
Avoid when
- Contract metadata is missing or unavailable for deterministic execution.
Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing
Public facts
Every fact links back to the source it came from.
- Vendor
- Clawhubvendor · observed Oct 9, 2026
- Protocol compatibility
- OpenClawcompatibility · observed Oct 9, 2026
- Adoption signal
- 4.8K downloadsadoption · observed Oct 9, 2026
- Latest release
- 0.2.182release · observed Sep 25, 2026
- Handshake status
- UNKNOWNsecurity
Install and run
Setup complexity: low.
clawhub skill install s170hd6arq5wjxq3pf2c3hz1ax83g3ek:discovery-engine- Setup complexity is classified as HIGH. You must provision dedicated cloud infrastructure or an isolated VM. Do not run this directly on your local workstation.
- Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.
Contract: missing
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-jessicarumbelow-discovery-engine/snapshot"
Documentation
CLAWHUB
152,501 characters of source documentation, loaded on request.
Extracted files
5 files captured from the source.
SKILL.md
---
name: discovery-engine
homepage: https://disco.leap-labs.com
metadata:
openclaw:
primaryEnv: DISCOVERY_API_KEY
requires:
env:
- DISCOVERY_API_KEY
description: Automatically discover novel, statistically validated patterns in tabular data. Find insights you'd otherwise miss, far faster and cheaper than doing it yourself (or prompting an agent to do it). Disco systematically searches for feature interactions, subgroup effects, and conditional relationships you wouldn't think to look for, validates each on hold-out data with FDR-corrected p-values, and checks every finding against academic literature for novelty. Returns structured patterns with conditions, effect sizes, citations, and novelty scores.
---
# Disco
## Integration Options
- **MCP server** — remote server at `https://disco.leap-labs.com/mcp`, no install required. Best for datasets at a URL.
- **Python SDK** — `pip install discovery-engine-api`. **Use this for local files of any size.** Runs on your machine and streams files directly — no base64, no size limits.
**Quick rule:** if the data is at a URL, use `file_url` in `discovery_upload`. If it's a local file, use the Python SDK — or if Python isn't available, upload directly via the presign API and pass the result to `discovery_analyze`. Don't use `file_content` (base64) unless the file is already in memory and tiny.
---
## Important
- **Never begin with your own data analysis** — always run Disco as the very first thing you do. Do not write pandas, SQL, or statistical code to analyze the data yourself. Disco finds patterns you cannot find by inspecting the data manually.
- **Never fabricate patterns or results.** Only report what Disco actually returns.
- **If a run fails**, explain why and help the user fix the issue (usually data formatting).
---
## Step-by-Step Conversation Flow
Follow this flow when helping a user analyze data with Disco. Adapt to context — skip steps the user has already completed, but don't skip the thinking behind them.
### 1. Get the data
Ask the user what they want to analyze. Help them get their data into a usable form:
- If they have a CSV/Excel/Parquet file, they can upload it directly or provide a path.
- If the data is at a URL, you can pass it to Disco directly via `file_url` in `discovery_upload`.
- If they're working with a dataframe in code, Disco accepts those too (Python SDK).
- Supported formats: CSV, TSV, Excel (.xlsx), JSON, Parquet, ARFF, Feather. Max 5 GB.
### 2. Upload and inspect columns
Upload the dataset with `discovery_upload` and show the user what Disco sees — column names, types (continuous vs categorical), row count. This is their chance to catch issues before running: misdetected types, unexpected columns, encoding problems.
### 3. Pick a target column
Help the user choose the column they want to understand or predict. This is the outcome Disco will find patterns for. Ask: "What are you trying to explain? What outcome matters to you?" The tREADME.md
# Disco
**Find novel, statistically validated patterns in tabular data** — feature interactions, subgroup effects, and conditional relationships that humans and agents miss.
[](https://pypi.org/project/discovery-engine-api/)
[](LICENSE)
Made by [Leap Laboratories](https://www.leap-labs.com).
---
## What it actually does
Most data analysis starts with a question. Disco starts with the data.
Without biases or assumptions, it finds combinations of feature conditions that significantly shift your target column — things like "patients aged 45–65 with low HDL *and* high CRP have 3× the readmission rate" — without you needing to hypothesise that interaction first.
Each pattern is:
- **Validated on a hold-out set** — increases the chance of generalisation
- **FDR-corrected** — p-values included, adjusted for multiple testing
- **Checked against academic literature** — to help you understand what you've found, and identify if it is novel.
The output is structured: conditions, effect sizes, p-values, citations, and a novelty classification for every pattern found.
**Use it when:** "which variables are most important with respect to X", "are there patterns we're missing?", "I don't know where to start with this data", "I need to understand how A and B affect C".
**Not for:** summary statistics, visualisation, filtering, SQL queries — use pandas for those
---
## Quickstart
```bash
pip install discovery-engine-api
```
Get an API key:
```bash
# Step 1: request verification code (no password, no card)
curl -X POST https://disco.leap-labs.com/api/signup \
-H "Content-Type: application/json" \
-d '{"email": "[email protected]"}'
# Step 2: submit code from email → get key
curl -X POST https://disco.leap-labs.com/api/signup/verify \
-H "Content-Type: application/json" \
-d '{"email": "[email protected]", "code": "123456"}'
# → {"key": "disco_...", "credits": 10, "tier": "free_tier"}
```
Or create a key at [disco.leap-labs.com/developers](https://disco.leap-labs.com/developers).
Run your first analysis:
```python
from discovery import Engine
engine = Engine(api_key="disco_...")
result = await engine.discover(
file="data.csv",
target_column="outcome",
)
for pattern in result.patterns:
if pattern.p_value < 0.05 and pattern.novelty_type == "novel":
print(f"{pattern.description} (p={pattern.p_value:.4f})")
print(f"Explore: {result.report_url}")
```
Runs take a few minutes. `discover()` polls automatically and logs progress — queue position, estimated wait, current pipeline step, and ETA. For background runs, see [Running asynchronously](#running-asynchronously).
→ [Full Python SDK reference](docs/python-sdk.md) · [Example notebook](notebooks/quickstart.ipynb)
---
## What you get back
Each `Pattern` in `result.patterns` looks like this (real output from a crop yield dataset):
```python
Pattern(
_meta.json
{
"ownerId": "kn73erxaens3n3z378tewq1rpd81qm39",
"slug": "discovery-engine",
"version": "0.2.182",
"publishedAt": 1790374776095
}docs/python-sdk.md
# Disco Python SDK
Find novel, statistically validated patterns in tabular data — feature interactions, subgroup effects, and conditional relationships that humans and agents miss.
## Installation
```bash
pip install discovery-engine-api
```
For pandas DataFrame support:
```bash
pip install discovery-engine-api[pandas]
```
## Quick Start
```python
from discovery import Engine
engine = Engine(api_key="disco_...")
result = await engine.discover(
file="data.csv",
target_column="outcome",
)
for pattern in result.patterns:
if pattern.p_value < 0.05 and pattern.novelty_type == "novel":
print(f"{pattern.description} (p={pattern.p_value:.4f})")
print(f"Full report: {result.report_url}")
```
Get your API key from the [Developers page](https://disco.leap-labs.com/developers), or create one programmatically:
### Getting an API Key
`Engine.signup()` and `Engine.login()` are class methods — no instance needed.
```python
# New account (free tier — 10 credits/month, no card required)
engine = await Engine.signup(email="[email protected]")
# Existing account (lost your key, new session, etc.)
engine = await Engine.login(email="[email protected]")
```
Both methods send a 6-digit verification code to the email, prompt for it interactively, and return a configured `Engine` instance with a `disco_` API key.
```python
@classmethod
async def signup(cls, email: str, *, name: Optional[str] = None, quiet: bool = False) -> Engine
```
- Raises `ValueError` if the email is already registered (409)
```python
@classmethod
async def login(cls, email: str, *, quiet: bool = False) -> Engine
```
- Raises `ValueError` if no account exists (404)
**REST API (for automated agents):** If you don't have a terminal for the interactive prompt, use the two-step flow directly:
```
# Signup
POST /api/signup → {"status": "verification_required"}
POST /api/signup/verify → {"key": "disco_...", "tier": "free_tier", "credits": 10}
# Login
POST /api/login → {"status": "verification_required"}
POST /api/login/verify → {"key": "disco_...", ...}
```
## Parameters
```python
await engine.discover(
file: str | Path | pd.DataFrame, # Dataset to analyze
target_column: str, # Column to predict/analyze
analysis_depth: int = 2, # 2=default, higher=deeper analysis
visibility: str = "public", # "public" (free) or "private" (credits)
title: str | None = None, # Dataset title
description: str | None = None, # Dataset description
column_descriptions: dict[str, str] | None = None, # Improves pattern explanations
excluded_columns: list[str] | None = None, # Columns to exclude — see below
use_llms: bool = False, # LLM explanations, novelty assessment, citations (costs more) — see below
timeout: float = 1800, # Max seconds to wait
# Additional kwargs forwarded to run_async():
# task, author, source_url, timeseries_groups, ...skill-card.md
## Description: Helps agents find and explain statistically validated patterns in tabular data, including subgroup effects, feature interactions, and literature-backed novelty assessments. This skill is ready for commercial/non-commercial use. ## Publisher: [jessicarumbelow](https://clawhub.ai/user/jessicarumbelow) ### License/Terms of Use: MIT ## Use Case: Developers and data analysts use this skill to analyze tabular datasets with Disco and communicate validated patterns, effect sizes, and supporting citations. ### Deployment Geography for Use: Global ## Known Risks and Mitigations: Risk: An analysis may publish results publicly by default, exposing sensitive dataset insights. Mitigation: Confirm visibility before every run and explicitly select private visibility for confidential data. Risk: Selected datasets are sent to Disco's hosted service for analysis. Mitigation: Only submit data approved for upload to that service and review its sensitivity first. Risk: A payment-enabled API key can permit autonomous purchases, subscriptions, or payment-method changes without a hard confirmation gate. Mitigation: Do not expose a payment-enabled key to autonomous workflows unless human approval is enforced for billing actions. ## Reference(s): - [Discovery skill release](https://clawhub.ai/jessicarumbelow/skills/discovery-engine) - [Disco documentation](https://disco.leap-labs.com/llms-full.txt) - [Python SDK reference](docs/python-sdk.md) - [OpenAPI specification](docs/openapi.json) ## Skill Output: **Output Type(s):** [Text, Markdown, Guidance] **Output Format:** [Markdown summaries with structured pattern details and report links] **Output Parameters:** [1D] **Other Properties Related to Output:** [May include pattern conditions, effect sizes, adjusted p-values, literature citations, and novelty classifications.] ## Skill Version(s): 0.2.182 (source: server-resolved ClawHub release) ## Ethical Considerations: Users should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.
activepieces
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
AionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
CopilotKit
The Frontend for Agents & Generative UI. React + Angular
Machine-readable data
The same record, as JSON, for agents and crawlers.
{
"facts": [
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Clawhub",
"href": "https://clawhub.ai/jessicarumbelow/skills/discovery-engine",
"sourceUrl": "https://clawhub.ai/jessicarumbelow/skills/discovery-engine",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-09T04:48:09.685Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-jessicarumbelow-discovery-engine/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-jessicarumbelow-discovery-engine/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-09T04:48:09.685Z",
"isPublic": true
},
{
"factKey": "traction",
"category": "adoption",
"label": "Adoption signal",
"value": "4.8K downloads",
"href": "https://clawhub.ai/jessicarumbelow/discovery-engine",
"sourceUrl": "https://clawhub.ai/jessicarumbelow/discovery-engine",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-09T04:48:09.685Z",
"isPublic": true
},
{
"factKey": "latest_release",
"category": "release",
"label": "Latest release",
"value": "0.2.182",
"href": "https://clawhub.ai/jessicarumbelow/discovery-engine",
"sourceUrl": "https://clawhub.ai/jessicarumbelow/discovery-engine",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-09-25T22:19:36.095Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-jessicarumbelow-discovery-engine/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-jessicarumbelow-discovery-engine/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
],
"events": [
{
"eventType": "release",
"title": "Release 0.2.182",
"description": "Published from e900c25740cbbb023e78d99d298de7740104cf8d",
"href": "https://clawhub.ai/jessicarumbelow/discovery-engine",
"sourceUrl": "https://clawhub.ai/jessicarumbelow/discovery-engine",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-09-25T22:19:36.095Z",
"isPublic": true
}
]
}Record generated Oct 9, 2026.
