Crawler Summary

skill-studio answer-first brief

Comprehensive skill evaluation and debugging framework for testing agent skills. Use when users need to (1) evaluate a skill's recall rate (how often it triggers correctly), (2) test a skill's accuracy against expected outputs, (3) analyze skill performance with various prompts, or (4) generate improvement recommendations for existing skills. Requires test cases with standard answers. --- name: skill-studio description: "Comprehensive skill evaluation and debugging framework for testing agent skills. Use when users need to (1) evaluate a skill's recall rate (how often it triggers correctly), (2) test a skill's accuracy against expected outputs, (3) analyze skill performance with various prompts, or (4) generate improvement recommendations for existing skills. Requires test cases with standard answ Capability contract not published. No trust telemetry is available yet. 1 GitHub stars reported by the source. Last updated 4/15/2026.

Freshness

Last checked 4/15/2026

Best For

skill-studio is best for general automation workflows where OpenClaw compatibility matters.

Not Ideal For

Contract metadata is missing or unavailable for deterministic execution.

Evidence Sources Checked

editorial-content, GITHUB OPENCLEW, runtime-metrics, public facts pack

Claim this agent
Agent DossierGitHubSafety: 94/100

skill-studio

Comprehensive skill evaluation and debugging framework for testing agent skills. Use when users need to (1) evaluate a skill's recall rate (how often it triggers correctly), (2) test a skill's accuracy against expected outputs, (3) analyze skill performance with various prompts, or (4) generate improvement recommendations for existing skills. Requires test cases with standard answers. --- name: skill-studio description: "Comprehensive skill evaluation and debugging framework for testing agent skills. Use when users need to (1) evaluate a skill's recall rate (how often it triggers correctly), (2) test a skill's accuracy against expected outputs, (3) analyze skill performance with various prompts, or (4) generate improvement recommendations for existing skills. Requires test cases with standard answ

OpenClawself-declared

Public facts

5

Change events

1

Artifacts

0

Freshness

Apr 15, 2026

Verifiededitorial-contentNo verified compatibility signals1 GitHub stars

Capability contract not published. No trust telemetry is available yet. 1 GitHub stars reported by the source. Last updated 4/15/2026.

1 GitHub starsTrust evidence available

Trust score

Unknown

Compatibility

OpenClaw

Freshness

Apr 15, 2026

Vendor

Lzw12w

Artifacts

0

Benchmarks

0

Last release

Unpublished

Executive Summary

Key links, install path, and a quick operational read before the deeper crawl record.

Verifiededitorial-content

Summary

Capability contract not published. No trust telemetry is available yet. 1 GitHub stars reported by the source. Last updated 4/15/2026.

Setup snapshot

git clone https://github.com/lzw12w/agent-skills-studio.git
  1. 1

    Setup complexity is LOW. This package is likely designed for quick installation with minimal external side-effects.

  2. 2

    Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.

Evidence Ledger

Everything public we have scraped or crawled about this agent, grouped by evidence type with provenance.

Verifiededitorial-content
Vendor (1)

Vendor

Lzw12w

profilemedium
Observed Apr 15, 2026Source linkProvenance
Compatibility (1)

Protocol compatibility

OpenClaw

contractmedium
Observed Apr 15, 2026Source linkProvenance
Adoption (1)

Adoption signal

1 GitHub stars

profilemedium
Observed Apr 15, 2026Source linkProvenance
Security (1)

Handshake status

UNKNOWN

trustmedium
Observed unknownSource linkProvenance
Integration (1)

Crawlable docs

6 indexed pages on the official domain

search_documentmedium
Observed Apr 15, 2026Source linkProvenance

Release & Crawl Timeline

Merged public release, docs, artifact, benchmark, pricing, and trust refresh events.

Self-declaredagent-index

Artifacts Archive

Extracted files, examples, snippets, parameters, dependencies, permissions, and artifact metadata.

Self-declaredGITHUB OPENCLEW

Extracted files

0

Examples

3

Snippets

0

Languages

typescript

Parameters

Executable Examples

bash

python3 scripts/test_recall.py \
  --skill-name <skill-name> \
  --test-cases <path-to-test-cases.json> \
  --variations 5 \
  --output outputs/recall_results.json

bash

python3 scripts/test_accuracy.py \
  --skill-name <skill-name> \
  --test-cases <path-to-test-cases.json> \
  --standard-answers <path-to-standard-answers.json> \
  --workspace <path-to-workspace-directory> \
  --artifacts-dir test_artifacts \
  --output outputs/accuracy_results.json

bash

python3 scripts/analyze_results.py \
  --recall-results outputs/recall_results.json \
  --accuracy-results outputs/accuracy_results.json \
  --skill-path <path-to-skill-folder> \
  --output outputs/recommendations.md

Docs & README

Full documentation captured from public sources, including the complete README when available.

Self-declaredGITHUB OPENCLEW

Docs source

GITHUB OPENCLEW

Editorial quality

ready

Comprehensive skill evaluation and debugging framework for testing agent skills. Use when users need to (1) evaluate a skill's recall rate (how often it triggers correctly), (2) test a skill's accuracy against expected outputs, (3) analyze skill performance with various prompts, or (4) generate improvement recommendations for existing skills. Requires test cases with standard answers. --- name: skill-studio description: "Comprehensive skill evaluation and debugging framework for testing agent skills. Use when users need to (1) evaluate a skill's recall rate (how often it triggers correctly), (2) test a skill's accuracy against expected outputs, (3) analyze skill performance with various prompts, or (4) generate improvement recommendations for existing skills. Requires test cases with standard answ

Full README

name: skill-studio description: "Comprehensive skill evaluation and debugging framework for testing agent skills. Use when users need to (1) evaluate a skill's recall rate (how often it triggers correctly), (2) test a skill's accuracy against expected outputs, (3) analyze skill performance with various prompts, or (4) generate improvement recommendations for existing skills. Requires test cases with standard answers."

Skill Studio

A systematic framework for evaluating and debugging agent skills through automated testing of recall rates and output accuracy.

Overview

Skill Studio provides tools and workflows to:

  • Test if a skill triggers correctly (recall rate)
  • Verify skill outputs match expected results (accuracy rate)
  • Evaluate code outputs with specialized metrics (syntax, structure, functionality, quality)
  • Generate actionable improvement recommendations
  • Track performance across iterations

Core Workflow

Step 1: Collect Test Requirements

Gather from the user:

  • Target skill name: The skill being evaluated (e.g., "pptx", "docx", "custom-skill")
  • Test cases file: Path to test prompts (or help create one)
  • Standard answers file: Path to expected outputs (or help create one)
  • Test parameters: Number of prompt variations, testing approach

Step 2: Prepare Test Cases

If test cases don't exist, help create test_cases.json using the format in references/test_format.md.

Key elements:

  • Diverse prompts that should trigger the skill
  • Negative cases (prompts that should NOT trigger)
  • Category labels for analysis
  • Expected behaviors

Step 3: Run Recall Testing

Execute recall rate evaluation:

python3 scripts/test_recall.py \
  --skill-name <skill-name> \
  --test-cases <path-to-test-cases.json> \
  --variations 5 \
  --output outputs/recall_results.json

What it does:

  • Generates prompt variations for each test case
  • Tests if the skill triggers with each variation
  • Calculates recall rates overall and by category
  • Identifies false positives and false negatives

Key metrics:

  • Overall recall rate: % of prompts that correctly triggered
  • Per-category recall: Performance by test case type
  • False positive rate: Incorrect triggering

Step 4: Run Accuracy Testing

Execute output accuracy evaluation:

python3 scripts/test_accuracy.py \
  --skill-name <skill-name> \
  --test-cases <path-to-test-cases.json> \
  --standard-answers <path-to-standard-answers.json> \
  --workspace <path-to-workspace-directory> \
  --artifacts-dir test_artifacts \
  --output outputs/accuracy_results.json

What it does:

  • Runs the skill with test prompts
  • Captures file changes (for coding agent skills)
  • Compares outputs to standard answers
  • Evaluates both exact matches and semantic similarity
  • Saves before/after states and diffs
  • Categorizes errors by type

Key parameters:

  • --workspace: Optional. Path to workspace directory for capturing code changes
  • --artifacts-dir: Optional. Directory to save test artifacts (default: test_artifacts)

Key metrics:

  • Exact match accuracy: % of perfect matches
  • Semantic similarity: Average similarity score (0-1)
  • Error categories: Common failure patterns

Artifacts saved (when --workspace is provided):

For each test case, the following artifacts are saved in test_artifacts/<test_id>/:

  • metadata.json: Test case information
  • before_state.json: File states before skill execution
  • after_state.json: File states after skill execution
  • changes.json: Summary of added/modified/deleted files
  • diffs/<file>.diff: Unified diffs for each modified file
  • full_output.json: Complete skill response

This allows you to:

  • Review exact code changes made by the skill
  • Compare actual changes to expected changes
  • Debug why tests failed
  • Track how the skill evolves over iterations

For code outputs:

  • Syntax validity: Does code parse correctly?
  • Structure match: Has expected functions/classes/imports?
  • Functionality: Passes test cases?
  • Code quality: Comments, docstrings, error handling, type hints

Step 5: Analyze and Generate Recommendations

Run the analysis to get improvement suggestions:

python3 scripts/analyze_results.py \
  --recall-results outputs/recall_results.json \
  --accuracy-results outputs/accuracy_results.json \
  --skill-path <path-to-skill-folder> \
  --output outputs/recommendations.md

Output includes:

  • Performance summary with visual indicators
  • Specific issues identified (low-performing categories)
  • Concrete recommendations for SKILL.md improvements
  • Suggested trigger phrases for description
  • Priority-ranked action items

Step 6: Implement Improvements

Based on recommendations:

  1. Review the generated recommendations document
  2. Update the skill's SKILL.md frontmatter description
  3. Enhance instructions in SKILL.md body
  4. Add clarifying examples or decision trees
  5. Re-run tests to validate improvements

Iterate until target performance is achieved.

Interpreting Results

Recall Rate Thresholds

  • >90%: Excellent - skill triggers reliably
  • 70-90%: Good - minor description improvements helpful
  • 50-70%: Fair - description needs refinement
  • <50%: Poor - major triggering issues, likely unclear description

Accuracy Rate Thresholds

  • >85%: Excellent - outputs are high quality
  • 70-85%: Good - minor instruction improvements helpful
  • 55-70%: Fair - instructions need clarification
  • <55%: Poor - major workflow or instruction issues

Common Issues and Solutions

Low Recall Rate:

  • Description is too vague or generic
  • Missing key trigger keywords
  • Overlaps with other skills' descriptions
  • Solution: Add specific file types, actions, or scenarios to description

High False Positive Rate:

  • Description is too broad
  • Includes common generic terms
  • Solution: Be more specific about exact use cases

Low Accuracy Rate:

  • Instructions are unclear or incomplete
  • Missing critical steps in workflow
  • Examples don't match user expectations
  • Solution: Add decision trees, more examples, clearer step-by-step guidance

Variable Performance by Category:

  • Skill handles some use cases well but not others
  • Solution: Add specific guidance for underperforming categories

Best Practices

Test Case Design

  • Include 20-30 diverse test cases minimum
  • Cover all major use cases the skill should handle
  • Add negative cases (should NOT trigger) at ~20% ratio
  • Use realistic, varied phrasing
  • Label test cases by category for granular analysis

Standard Answers

  • Define clear quality criteria
  • Use semantic similarity for flexible matching when appropriate
  • Include both structure and content expectations
  • Specify file formats and key elements

Iterative Testing

  • Test after each modification to track progress
  • Focus on lowest-performing categories first
  • Keep a log of changes and their impact
  • Aim for incremental improvements (5-10% gains per iteration)

Skill Description Optimization

  • Include specific file extensions (e.g., ".docx", ".pptx")
  • List concrete actions (e.g., "creating", "editing", "analyzing")
  • Mention key scenarios (e.g., "when user uploads", "when user requests")
  • Use numbered lists for multiple trigger conditions
  • Avoid generic terms that apply to many skills

Resources

scripts/

  • test_recall.py: Measures skill triggering consistency
  • test_accuracy.py: Evaluates output quality
  • analyze_results.py: Generates improvement recommendations
  • generate_variations.py: Creates prompt variations for testing

references/

  • test_format.md: Detailed test case format specification
  • answer_format.md: Standard answer format specification
  • code_evaluation.md: Code output evaluation format and best practices
  • metrics_explained.md: Deep dive into evaluation metrics

Contract & API

Machine endpoints, protocol fit, contract coverage, invocation examples, and guardrails for agent-to-agent use.

MissingGITHUB OPENCLEW

Contract coverage

Status

missing

Auth

None

Streaming

No

Data region

Unspecified

Protocol support

OpenClaw: self-declared

Requires: none

Forbidden: none

Guardrails

Operational confidence: low

No positive guardrails captured.
Invocation examples
curl -s "https://www.xpersona.co/api/v1/agents/lzw12w-agent-skills-studio/snapshot"
curl -s "https://www.xpersona.co/api/v1/agents/lzw12w-agent-skills-studio/contract"
curl -s "https://www.xpersona.co/api/v1/agents/lzw12w-agent-skills-studio/trust"

Reliability & Benchmarks

Trust and runtime signals, benchmark suites, failure patterns, and practical risk constraints.

Missingruntime-metrics

Trust signals

Handshake

UNKNOWN

Confidence

unknown

Attempts 30d

unknown

Fallback rate

unknown

Runtime metrics

Observed P50

unknown

Observed P95

unknown

Rate limit

unknown

Estimated cost

unknown

Do not use if

Contract metadata is missing or unavailable for deterministic execution.
No benchmark suites or observed failure patterns are available.

Media & Demo

Every public screenshot, visual asset, demo link, and owner-provided destination tied to this agent.

Missingno-media
No screenshots, media assets, or demo links are available.

Related Agents

Neighboring agents from the same protocol and source ecosystem for comparison and shortlist building.

Self-declaredprotocol-neighbors
Github ReposUpdated 6mo agoRank 70

activepieces

AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents

OPENCLAW
Github ReposUpdated 6mo agoRank 70

cherry-studio

AI productivity studio with smart chat, autonomous agents, and 300+ assistants.

MCPOPENCLAW
Github ReposUpdated 6mo agoRank 70

AionUi

Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!

MCPOPENCLAW
Github ReposUpdated 7mo agoRank 70

CopilotKit

The Frontend for Agents & Generative UI. React + Angular

OPENCLAW
Machine Appendix

Contract JSON

{
  "contractStatus": "missing",
  "authModes": [],
  "requires": [],
  "forbidden": [],
  "supportsMcp": false,
  "supportsA2a": false,
  "supportsStreaming": false,
  "inputSchemaRef": null,
  "outputSchemaRef": null,
  "dataRegion": null,
  "contractUpdatedAt": null,
  "sourceUpdatedAt": null,
  "freshnessSeconds": null
}

Invocation Guide

{
  "preferredApi": {
    "snapshotUrl": "https://www.xpersona.co/api/v1/agents/lzw12w-agent-skills-studio/snapshot",
    "contractUrl": "https://www.xpersona.co/api/v1/agents/lzw12w-agent-skills-studio/contract",
    "trustUrl": "https://www.xpersona.co/api/v1/agents/lzw12w-agent-skills-studio/trust"
  },
  "curlExamples": [
    "curl -s \"https://www.xpersona.co/api/v1/agents/lzw12w-agent-skills-studio/snapshot\"",
    "curl -s \"https://www.xpersona.co/api/v1/agents/lzw12w-agent-skills-studio/contract\"",
    "curl -s \"https://www.xpersona.co/api/v1/agents/lzw12w-agent-skills-studio/trust\""
  ],
  "jsonRequestTemplate": {
    "query": "summarize this repo",
    "constraints": {
      "maxLatencyMs": 2000,
      "protocolPreference": [
        "OPENCLEW"
      ]
    }
  },
  "jsonResponseTemplate": {
    "ok": true,
    "result": {
      "summary": "...",
      "confidence": 0.9
    },
    "meta": {
      "source": "GITHUB_OPENCLEW",
      "generatedAt": "2026-10-09T12:17:18.311Z"
    }
  },
  "retryPolicy": {
    "maxAttempts": 3,
    "backoffMs": [
      500,
      1500,
      3500
    ],
    "retryableConditions": [
      "HTTP_429",
      "HTTP_503",
      "NETWORK_TIMEOUT"
    ]
  }
}

Trust JSON

{
  "status": "unavailable",
  "handshakeStatus": "UNKNOWN",
  "verificationFreshnessHours": null,
  "reputationScore": null,
  "p95LatencyMs": null,
  "successRate30d": null,
  "fallbackRate": null,
  "attempts30d": null,
  "trustUpdatedAt": null,
  "trustConfidence": "unknown",
  "sourceUpdatedAt": null,
  "freshnessSeconds": null
}

Capability Matrix

{
  "rows": [
    {
      "key": "OPENCLEW",
      "type": "protocol",
      "support": "unknown",
      "confidenceSource": "profile",
      "notes": "Listed on profile"
    }
  ],
  "flattenedTokens": "protocol:OPENCLEW|unknown|profile"
}

Facts JSON

[
  {
    "factKey": "docs_crawl",
    "label": "Crawlable docs",
    "value": "6 indexed pages on the official domain",
    "category": "integration",
    "href": "https://github.com/login?return_to=https%3A%2F%2Fgithub.com%2Fopenclaw%2Fskills%2Ftree%2Fmain%2Fskills%2Fasleep123%2Fcaldav-calendar",
    "sourceUrl": "https://github.com/login?return_to=https%3A%2F%2Fgithub.com%2Fopenclaw%2Fskills%2Ftree%2Fmain%2Fskills%2Fasleep123%2Fcaldav-calendar",
    "sourceType": "search_document",
    "confidence": "medium",
    "observedAt": "2026-04-15T05:03:46.393Z",
    "isPublic": true,
    "metadata": {}
  },
  {
    "factKey": "vendor",
    "label": "Vendor",
    "value": "Lzw12w",
    "category": "vendor",
    "href": "https://github.com/lzw12w/agent-skills-studio",
    "sourceUrl": "https://github.com/lzw12w/agent-skills-studio",
    "sourceType": "profile",
    "confidence": "medium",
    "observedAt": "2026-04-15T03:13:23.576Z",
    "isPublic": true,
    "metadata": {}
  },
  {
    "factKey": "protocols",
    "label": "Protocol compatibility",
    "value": "OpenClaw",
    "category": "compatibility",
    "href": "https://www.xpersona.co/api/v1/agents/lzw12w-agent-skills-studio/contract",
    "sourceUrl": "https://www.xpersona.co/api/v1/agents/lzw12w-agent-skills-studio/contract",
    "sourceType": "contract",
    "confidence": "medium",
    "observedAt": "2026-04-15T03:13:23.576Z",
    "isPublic": true,
    "metadata": {}
  },
  {
    "factKey": "traction",
    "label": "Adoption signal",
    "value": "1 GitHub stars",
    "category": "adoption",
    "href": "https://github.com/lzw12w/agent-skills-studio",
    "sourceUrl": "https://github.com/lzw12w/agent-skills-studio",
    "sourceType": "profile",
    "confidence": "medium",
    "observedAt": "2026-04-15T03:13:23.576Z",
    "isPublic": true,
    "metadata": {}
  },
  {
    "factKey": "handshake_status",
    "label": "Handshake status",
    "value": "UNKNOWN",
    "category": "security",
    "href": "https://www.xpersona.co/api/v1/agents/lzw12w-agent-skills-studio/trust",
    "sourceUrl": "https://www.xpersona.co/api/v1/agents/lzw12w-agent-skills-studio/trust",
    "sourceType": "trust",
    "confidence": "medium",
    "observedAt": null,
    "isPublic": true,
    "metadata": {}
  }
]

Change Events JSON

[
  {
    "eventType": "docs_update",
    "title": "Docs refreshed: Sign in to GitHub · GitHub",
    "description": "Fresh crawlable documentation was indexed for the official domain.",
    "href": "https://github.com/login?return_to=https%3A%2F%2Fgithub.com%2Fopenclaw%2Fskills%2Ftree%2Fmain%2Fskills%2Fasleep123%2Fcaldav-calendar",
    "sourceUrl": "https://github.com/login?return_to=https%3A%2F%2Fgithub.com%2Fopenclaw%2Fskills%2Ftree%2Fmain%2Fskills%2Fasleep123%2Fcaldav-calendar",
    "sourceType": "search_document",
    "confidence": "medium",
    "observedAt": "2026-04-15T05:03:46.393Z",
    "isPublic": true,
    "metadata": {}
  }
]

Sponsored

Ads related to skill-studio and adjacent AI workflows.