agentCLAWHUBUnverified

data-scientist

You are a data scientist with expertise in statistical analysis, machine learning, data visualization, and experimental design. Use when: statistical analysi... Skill: data-scientist Owner: mtsatryan Summary: You are a data scientist with expertise in statistical analysis, machine learning, data visualization, and experimental design. Use when: statistical analysi... Tags: latest:1.0.0 Version history: v1.0.0 | 2026-04-30T12:42:45.733Z | user Initial release — part of 188 AI agent skills collection by MTNT Solutions Archive index: Archive v1.0.0: 4 files, 9247 bytes Files: r

OpenClaw

Rank

62

Safety

84

Downloads

1.1k

Updated

Oct 11, 2026

Version

1.0.0

Source

CLAWHUB

About

What it does, and when to use it.

Capability contract not published. No trust telemetry is available yet. 1.1K downloads reported by the source. Last updated 10/11/2026.

Avoid when

  • Contract metadata is missing or unavailable for deterministic execution.

Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing

Public facts

Every fact links back to the source it came from.

Vendor
Clawhubvendor · observed Oct 11, 2026
Protocol compatibility
OpenClawcompatibility · observed Oct 11, 2026
Adoption signal
1.1K downloadsadoption · observed Oct 11, 2026
Latest release
1.0.0release · observed Apr 30, 2026
Handshake status
UNKNOWNsecurity

Install and run

Setup complexity: low.

clawhub skill install s17bvyvkfhp17ybx0q3ak5dcsn85nqpv:ah-data-scientist
  1. Setup complexity is LOW. This package is likely designed for quick installation with minimal external side-effects.
  2. Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.

Contract: missing

curl -s "https://www.xpersona.co/api/v1/agents/clawhub-mtsatryan-ah-data-scientist/snapshot"

Documentation

CLAWHUB

29,706 characters of source documentation, loaded on request.

Extracted files

4 files captured from the source.

SKILL.md

---
name: data-scientist
description: 'You are a data scientist with expertise in statistical analysis, machine learning, data visualization, and experimental design. Use when: statistical analysis and hypothesis testing, machine learning model development and evaluation, data visualization and storytelling, experimental design and a/b testing, feature engineering and selection.'
---

# Data Scientist

You are a data scientist with expertise in statistical analysis, machine learning, data visualization, and experimental design.

## Core Expertise
- Statistical analysis and hypothesis testing
- Machine learning model development and evaluation
- Data visualization and storytelling
- Experimental design and A/B testing
- Feature engineering and selection
- Time series analysis and forecasting
- Deep learning and neural networks
- Causal inference and econometrics

## Technical Skills
- **Languages**: Python, R, SQL, Scala, Julia
- **ML Libraries**: scikit-learn, XGBoost, LightGBM, CatBoost
- **Deep Learning**: TensorFlow, PyTorch, Keras, JAX
- **Data Manipulation**: pandas, numpy, polars, dplyr
- **Visualization**: matplotlib, seaborn, plotly, ggplot2, Tableau
- **Big Data**: Spark, Dask, Ray, Databricks
- **Cloud Platforms**: AWS SageMaker, Google AI Platform, Azure ML

## Statistical Analysis Framework
> 📎 **Code example 1** (python) — see [references/examples.md](references/examples.md)

## Machine Learning Pipeline
> 📎 **Code example 2** (python) — see [references/examples.md](references/examples.md)

## Time Series Analysis
> 📎 **Code example 3** (python) — see [references/examples.md](references/examples.md)

## A/B Testing Framework
> 📎 **Code example 4** (python) — see [references/examples.md](references/examples.md)

## Data Visualization Suite
> 📎 **Code example 5** (python) — see [references/examples.md](references/examples.md)

## Best Practices
1. **Data Quality**: Always validate and clean data before analysis
2. **Reproducibility**: Use random seeds and version control for experiments
3. **Cross-Validation**: Use proper validation techniques to avoid overfitting
4. **Feature Engineering**: Invest time in creating meaningful features
5. **Model Interpretability**: Use SHAP, LIME for model explanation
6. **Statistical Significance**: Don't confuse statistical and practical significance
7. **Documentation**: Document assumptions, methodologies, and findings

## Experimental Design
- Design experiments with proper controls and randomization
- Calculate required sample sizes before data collection
- Account for multiple testing corrections
- Use appropriate statistical tests for your data type
- Consider confounding variables and bias sources
- Plan for missing data and outlier handling

## Approach
- Start with exploratory data analysis and data quality assessment
- Define clear hypotheses and success metrics
- Choose appropriate statistical methods and models
- Validate results using multiple approaches
- Communicate findings with 

_meta.json

{
  "ownerId": "kn7fhxm2kjnxpxwkk5x3h3xj1985nhw1",
  "slug": "ah-data-scientist",
  "version": "1.0.0",
  "publishedAt": 1777552965733
}

references/examples.md

# Data Scientist — Code Examples

## Example 1

```python
import pandas as pd
import numpy as np
import scipy.stats as stats
from scipy.stats import ttest_ind, chi2_contingency, mannwhitneyu
import matplotlib.pyplot as plt
import seaborn as sns
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import train_test_split
from sklearn.metrics import classification_report, confusion_matrix

class StatisticalAnalyzer:
    def __init__(self, data):
        self.data = data
        self.results = {}
    
    def descriptive_statistics(self, columns=None):
        """Generate comprehensive descriptive statistics"""
        if columns is None:
            columns = self.data.select_dtypes(include=[np.number]).columns
        
        stats_summary = {}
        for col in columns:
            stats_summary[col] = {
                'count': self.data[col].count(),
                'mean': self.data[col].mean(),
                'median': self.data[col].median(),
                'std': self.data[col].std(),
                'min': self.data[col].min(),
                'max': self.data[col].max(),
                'q25': self.data[col].quantile(0.25),
                'q75': self.data[col].quantile(0.75),
                'skewness': stats.skew(self.data[col].dropna()),
                'kurtosis': stats.kurtosis(self.data[col].dropna())
            }
        
        return pd.DataFrame(stats_summary).T
    
    def hypothesis_testing(self, group_col, target_col, test_type='auto'):
        """Perform appropriate hypothesis tests"""
        groups = self.data[group_col].unique()
        
        if len(groups) != 2:
            raise ValueError("Currently supports only two-group comparisons")
        
        group1 = self.data[self.data[group_col] == groups[0]][target_col].dropna()
        group2 = self.data[self.data[group_col] == groups[1]][target_col].dropna()
        
        # Normality tests
        _, p_norm1 = stats.shapiro(group1.sample(min(5000, len(group1))))
        _, p_norm2 = stats.shapiro(group2.sample(min(5000, len(group2))))
        
        # Equal variance test
        _, p_var = stats.levene(group1, group2)
        
        results = {
            'group1_size': len(group1),
            'group2_size': len(group2),
            'group1_mean': group1.mean(),
            'group2_mean': group2.mean(),
            'normality_p1': p_norm1,
            'normality_p2': p_norm2,
            'equal_variance_p': p_var
        }
        
        # Choose appropriate test
        if test_type == 'auto':
            if p_norm1 > 0.05 and p_norm2 > 0.05:
                # Both normal, use t-test
                if p_var > 0.05:
                    # Equal variances
                    stat, p_value = ttest_ind(group1, group2)
                    test_used = "Independent t-test (equal variances)"
                else:
                    # Unequal variances
                    stat, p_value = ttest_ind(group1, group2, equal_var=False)

skill-card.md

## Description:

You are a data scientist with expertise in statistical analysis, machine learning, data visualization, and experimental design.

This skill is ready for commercial/non-commercial use.

## Publisher:

[mtsatryan](https://clawhub.ai/user/mtsatryan)

### License/Terms of Use:

MIT-0

## Use Case:

Developers, analysts, and data science teams use this skill to plan and produce statistical analyses, machine learning workflows, visualizations, experimental designs, and reproducible data science deliverables.

### Deployment Geography for Use:

Global

## Known Risks and Mitigations:

Risk: Generated analysis code may process local or sensitive datasets and may rely on third-party Python libraries in the user's environment.

Mitigation: Review generated analysis code before running it, especially on sensitive data, and confirm dependencies and data handling practices fit the deployment environment.

Risk: Statistical or machine learning guidance can be misapplied if assumptions, validation choices, or data quality issues are not checked.

Mitigation: Validate datasets, methods, assumptions, and model results before using outputs for decisions.

## Reference(s):

- [Data Scientist Code Examples](references/examples.md)

## Skill Output:

**Output Type(s):** [text, markdown, code, guidance]

**Output Format:** [Markdown with code examples and analysis guidance]

**Output Parameters:** [1D]

**Other Properties Related to Output:** [May include reproducible analysis code, statistical interpretations, visualizations, notebook-style explanations, assumptions, limitations, and recommendations.]

## Skill Version(s):

1.0.0 (source: server release evidence)

## Ethical Considerations:

Users should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.
Github ReposUpdated 1d agoRank 70

AionUi

Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!

MCPOPENCLAW
Github ReposUpdated 6mo agoRank 70

activepieces

AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents

OPENCLAW
Github ReposUpdated 6mo agoRank 70

cherry-studio

AI productivity studio with smart chat, autonomous agents, and 300+ assistants.

MCPOPENCLAW
Github ReposUpdated 7mo agoRank 70

CopilotKit

The Frontend for Agents & Generative UI. React + Angular

OPENCLAW

Machine-readable data

The same record, as JSON, for agents and crawlers.

{
  "facts": [
    {
      "factKey": "vendor",
      "category": "vendor",
      "label": "Vendor",
      "value": "Clawhub",
      "href": "https://clawhub.ai/mtsatryan/skills/ah-data-scientist",
      "sourceUrl": "https://clawhub.ai/mtsatryan/skills/ah-data-scientist",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-11T05:26:59.369Z",
      "isPublic": true
    },
    {
      "factKey": "protocols",
      "category": "compatibility",
      "label": "Protocol compatibility",
      "value": "OpenClaw",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-mtsatryan-ah-data-scientist/contract",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-mtsatryan-ah-data-scientist/contract",
      "sourceType": "contract",
      "confidence": "medium",
      "observedAt": "2026-10-11T05:26:59.369Z",
      "isPublic": true
    },
    {
      "factKey": "traction",
      "category": "adoption",
      "label": "Adoption signal",
      "value": "1.1K downloads",
      "href": "https://clawhub.ai/mtsatryan/ah-data-scientist",
      "sourceUrl": "https://clawhub.ai/mtsatryan/ah-data-scientist",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-11T05:26:59.369Z",
      "isPublic": true
    },
    {
      "factKey": "latest_release",
      "category": "release",
      "label": "Latest release",
      "value": "1.0.0",
      "href": "https://clawhub.ai/mtsatryan/ah-data-scientist",
      "sourceUrl": "https://clawhub.ai/mtsatryan/ah-data-scientist",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-04-30T12:42:45.733Z",
      "isPublic": true
    },
    {
      "factKey": "handshake_status",
      "category": "security",
      "label": "Handshake status",
      "value": "UNKNOWN",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-mtsatryan-ah-data-scientist/trust",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-mtsatryan-ah-data-scientist/trust",
      "sourceType": "trust",
      "confidence": "medium",
      "observedAt": null,
      "isPublic": true
    }
  ],
  "events": [
    {
      "eventType": "release",
      "title": "Release 1.0.0",
      "description": "Initial release — part of 188 AI agent skills collection by MTNT Solutions",
      "href": "https://clawhub.ai/mtsatryan/ah-data-scientist",
      "sourceUrl": "https://clawhub.ai/mtsatryan/ah-data-scientist",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-04-30T12:42:45.733Z",
      "isPublic": true
    }
  ]
}

Record generated Oct 11, 2026.

Sponsored

Ads related to data-scientist and adjacent AI workflows.