Claim this agent
agentCLAWHUBUnverified

skill-usefulness-audit

Review your installed agent skills to see what you actually use, what overlaps, and what may no longer be worth keeping.

OpenClaw

Rank

62

Safety

84

Downloads

4.0k

Updated

Oct 9, 2026

Version

0.3.24

Source

CLAWHUB

About

What it does, and when to use it.

Capability contract not published. No trust telemetry is available yet. 4K downloads reported by the source. Last updated 10/9/2026.

Avoid when

  • Contract metadata is missing or unavailable for deterministic execution.

Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing

Public facts

Every fact links back to the source it came from.

Vendor
Clawhubvendor · observed Oct 9, 2026
Protocol compatibility
OpenClawcompatibility · observed Oct 9, 2026
Adoption signal
4K downloadsadoption · observed Oct 9, 2026
Latest release
0.3.24release · observed Aug 24, 2026
Handshake status
UNKNOWNsecurity

Install and run

Setup complexity: low.

clawhub skill install s172a0qxsw6kfdee064rcn06es83edse:skill-usefulness-audit
  1. Install using `clawhub skill install s172a0qxsw6kfdee064rcn06es83edse:skill-usefulness-audit` in an isolated environment before connecting it to live workloads.
  2. No published capability contract is available yet, so validate auth and request/response behavior manually.
  3. Review the upstream CLAWHUB listing at https://clawhub.ai/gongyu0918-debug/skill-usefulness-audit before using production credentials.

Contract: missing

curl -s "https://www.xpersona.co/api/v1/agents/clawhub-gongyu0918-debug-skill-usefulness-audit/snapshot"

Documentation

CLAWHUB

160,000 characters of source documentation, loaded on request.

Extracted files

5 files captured from the source.

SKILL.md

---
name: skill-usefulness-audit
slug: skill-usefulness-audit
description: Review your installed agent skills to see what you actually use, what overlaps, and what may no longer be worth keeping.
version: 0.3.24
tags: ["audit","skills","ablation","openclaw"]
user-invocable: true
disable-model-invocation: true
argument-hint: --skills-root PATH --usage-file FILE
homepage: https://github.com/gongyu0918-debug/skill-usefulness-audit
metadata: {"openclaw":{"skillKey":"skill-usefulness-audit","requires":{"bins":["python"]},"homepage":"https://github.com/gongyu0918-debug/skill-usefulness-audit"}}
---
# Skill Usefulness Audit

## Manual Trigger Only

Use this skill only after a direct request to audit installed agent skills, their usage, overlap, cleanup options, or a structure-only inventory.
Do not invoke it during normal tasks or use it for ordinary repository/source-code review, general security audit, or employee/human skill assessment.

## Safety

Never delete, merge, quarantine, isolate, or disable skills automatically.
Treat `delete`, `merge-delete`, and `quarantine-review` as manual-review recommendations.
Do not delete skills based only on a structure-only report.
This tool does not automatically replay historical conversations; it generates ablation plans and reads ablation result files that the user provides.

## Audit Scope

Audit these layers in order:

1. Usage evidence, including recency and source quality.
2. Installed metadata, instructions, and functional overlap.
3. User-provided skill-on versus skill-off results for general skills.
4. Runtime and bundle burden, including over-triggering, context cost, weak progressive disclosure, redundant resources, script failures, and private-looking files.
5. Static health and risk hints.
6. Optional offline community or registry metrics.

Treat API and tool skills as protected capability skills during ablation.
Examples: Excel, DOCX, PDF, browser automation, deployment, OCR, external API wrappers, MCP/API gateway helpers.

## Workflow

1. Collect user-provided roots before host-local defaults.
2. Load only the usage, history, ablation, and community evidence that is available.
3. Inspect each `SKILL.md` and its script/reference/asset metrics.
4. Let the bundled script classify each skill as `api`, `tool`, or `general` and calculate its score. Read `{baseDir}/references/scoring-rubric.md` only when checking or explaining a score, verdict, or action.
5. Print the short usefulness report and, when requested, write Markdown evidence or an ablation plan.

## Ablation Rules

Read `{baseDir}/references/ablation-protocol.md` only when running replays, preparing normalized ablation records, or reviewing mixed or delete-boundary results. The script can generate an ablation plan without loading the protocol.
Replay only selected `general` candidates with identical prompts/artifacts and pairwise judging.
Do not fake no-tool ablation for `api` or `tool` skills; use the rubric's protected-capability branch.

#

_meta.json

{
  "ownerId": "kn7em0w89d0zac35fzt84qm2a182j54b",
  "slug": "skill-usefulness-audit",
  "version": "0.3.24",
  "publishedAt": 1787566137284
}

references/ablation-protocol.md

# Ablation Protocol

Use this protocol for `general` skills selected by the ablation plan.

## Goal

Measure whether the skill changes outcomes in a meaningful way.
High consistency between skill-on and skill-off runs means the skill adds little value.

## Sampling

Start with `3` historical tasks where the skill should plausibly matter. Prefer real user turns over synthetic prompts. Expand to `5` when results are mixed and to `10` only for high-impact or delete-boundary decisions.

## Replay Method

For each selected case, run two isolated replays:

1. `with_skill`
2. `without_skill`

Keep these constant:

- same prompt
- same files and artifacts
- same model class when possible
- same tool permissions
- same success criteria

Use a fresh thread or isolated run if the host supports it.

## Judge Method

For open-ended outputs:

1. Compare `with_skill` and `without_skill` side by side.
2. Randomize A/B order.
3. Spot-check reversed order on boundary cases.
4. Prefer `pass/fail`, `same/better/worse`, and short reasons over long open-ended grading.

Record a standard `verdict` and one short `notes` reason. Each arm may also include optional `pass` and/or `score` from `0.0-1.0` for fallback inference. Optional `tool_cost` may describe calls, latency, or retries and currently does not affect the audit score.

### Normalized JSON

```json
[
  {
    "skill": "emotion-orchestrator",
    "case_id": "case-001",
    "with_skill": {"pass": true, "score": 0.92},
    "without_skill": {"pass": true, "score": 0.81},
    "verdict": "better"
  }
]
```

## Judgment Rule

Use `same` when the final answer, correctness, and workflow remain materially equivalent.
Use `better` when the skill improves correctness, speed, structure, or user-fit in a way the baseline did not.
Use `worse` when the skill adds friction, drift, or errors.

Ignore verdict-only cases with unsupported values. A case with an unknown or missing verdict is usable only when both arms provide comparable `pass` and/or `score` fields for inference.

## Early Stop Rules

- Stop as low-value when `3/3` cases are `same` and `better_rate` is `0`.
- Stop as useful when at least `2/3` cases are `better` and no case is `worse`.
- Expand to `5` when the first batch is mixed.
- Expand to `10` only for delete-boundary or high-impact decisions.

Delete-boundary means "needs stronger human review evidence", not automatic deletion authority.

## Planning and Model Cost

Create the replay plan with `--ablation-plan-out`. Planning uses local evidence and does not call an LLM.

The plan estimates replay cost for `light`, `realistic`, and `coding` profiles. Each case assumes two replays and one compact pairwise judge. `model_cost_estimates.unit` records `estimated_context_units_per_case`.

Feed normalized results back with `--ablation-file`.

references/report-narration-prompt.md

# Report Delivery Contract

- Treat standard output as the final short report. Copy it verbatim into chat, apart from making the evidence path clickable when supported.
- Do not add headings, bullets, extra counts, verification notes, or facts from the Markdown evidence unless requested.
- Keep the short report on actual use, missing use evidence, overlap, and verified outcome impact. Leave scores, internal codes, risk, bundle health, and tables in the Markdown evidence.
- Match the user's language: clean Chinese for `zh-CN` and clean English for `en`, except for skill names and unavoidable paths or commands.
- Do not paste raw JSON or read back the full Markdown report unless requested.
- Treat removal results as manual-review recommendations. Never remove, merge, isolate, or disable a skill automatically.

references/scoring-rubric.md

# Scoring Rubric

## Contents

- Core Outputs
- Usage Score
- Uniqueness Score
- Impact Score
- Confidence Score
- Quality Penalty
- Community Prior Score
- Static Risk Level
- Verdict Bands
- Action Rules

## Core Outputs

- `local_score = usage_score + uniqueness_score + impact_score`
- `quality_penalty`: `0.0-2.5`
- `quality_penalty_uncapped`: raw quality burden before the cap
- `static_quality_penalty`: `0.0-1.4`
- `final_score = clamp(local_score - quality_penalty, 0.0, 10.0)`
- `risk_level` / `static_risk_level`: `none / low / medium / high`

## 1. Usage Score (`0.0-3.0`)

Prefer direct host usage logs.
Use transcript mentions only as weaker fallback evidence.

### Input Fields

- Direct usage: `calls`, `recent_30d_calls`, `recent_90d_calls`, `last_used_at`, and `active_days`.
- History fallback: `history_mentions` and `suspected_invocations`. These are weak evidence weighted through history and must not be reported as direct `calls`.
- Evidence and runtime burden: `usage_source`, `evidence_weight`, `executions`, `script_failures`, `repair_turns`, `reference_loads`, and `false_triggers`.

### Base Usage Strength

- When `recent_30d_calls` exists:
  - `0.0`: `0`
  - `1.0`: `1-2`
  - `2.0`: `3-7`
  - `3.0`: `8+`
- When only `recent_90d_calls` exists:
  - `0.0`: `0`
  - `0.75`: `1-2`
  - `1.5`: `3-9`
  - `2.5`: `10+`
- When only total `calls` exists:
  - `0.0`: `0`
  - `1.0`: `1-2`
  - `2.0`: `3-9`
  - `3.0`: `10+`

### Recency Adjustments

- add `0.5` when `last_used_at <= 7 days`
- add `0.25` when `last_used_at <= 30 days`
- subtract `0.5` when `last_used_at > 180 days`
- add `0.25` when `active_days >= 10`
- add `0.10` when `active_days >= 3`

### Evidence Weight

- `1.00`: direct usage file
- `0.45`: transcript-history fallback based on `suspected_invocations`
- `0.00`: missing usage evidence

## 2. Uniqueness Score (`0.0-3.0`)

Measure the highest functional-overlap similarity against any other installed skill using descriptions, headings, and resource names.

Buckets:

- `0.0`: highest overlap `>= 0.85`
- `1.0`: highest overlap `0.65-0.84`
- `2.0`: highest overlap `0.40-0.64`
- `3.0`: highest overlap `< 0.40`

## 3. Impact Score (`0.0-4.0`)

### General skills

Use ablation on historical conversations.
Compute:

- `consistency_rate`: skill-on and skill-off produce materially equivalent outcomes
- `better_rate`: skill-on clearly improves the result
- `worse_rate`: skill-on clearly harms the result

Base score from consistency:

- `0.0`: `consistency_rate >= 0.85`
- `1.0`: `0.70-0.84`
- `2.0`: `0.55-0.69`
- `3.0`: `0.35-0.54`
- `4.0`: `< 0.35`

Adjustments:

- add `1.0` when `better_rate - worse_rate >= 0.30`
- subtract `1.0` when `worse_rate > better_rate`

When ablation is missing, use low-evidence score `1.0` for zero-call skills.
For skills with direct usage evidence but no ablation yet, keep temporary neutral score `2.0` and lower confidence.

### API and tool skills

Skip history ablation.
Use protected-capability scoring instead:

-
Github ReposUpdated 6mo agoRank 70

activepieces

AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents

OPENCLAW
Github ReposUpdated 6mo agoRank 70

cherry-studio

AI productivity studio with smart chat, autonomous agents, and 300+ assistants.

MCPOPENCLAW
Github ReposUpdated 6mo agoRank 70

AionUi

Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!

MCPOPENCLAW
Github ReposUpdated 7mo agoRank 70

CopilotKit

The Frontend for Agents & Generative UI. React + Angular

OPENCLAW

Machine-readable data

The same record, as JSON, for agents and crawlers.

{
  "facts": [
    {
      "factKey": "vendor",
      "category": "vendor",
      "label": "Vendor",
      "value": "Clawhub",
      "href": "https://clawhub.ai/gongyu0918-debug/skills/skill-usefulness-audit",
      "sourceUrl": "https://clawhub.ai/gongyu0918-debug/skills/skill-usefulness-audit",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-09T06:14:25.202Z",
      "isPublic": true
    },
    {
      "factKey": "protocols",
      "category": "compatibility",
      "label": "Protocol compatibility",
      "value": "OpenClaw",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-gongyu0918-debug-skill-usefulness-audit/contract",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-gongyu0918-debug-skill-usefulness-audit/contract",
      "sourceType": "contract",
      "confidence": "medium",
      "observedAt": "2026-10-09T06:14:25.202Z",
      "isPublic": true
    },
    {
      "factKey": "traction",
      "category": "adoption",
      "label": "Adoption signal",
      "value": "4K downloads",
      "href": "https://clawhub.ai/gongyu0918-debug/skill-usefulness-audit",
      "sourceUrl": "https://clawhub.ai/gongyu0918-debug/skill-usefulness-audit",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-09T06:14:25.202Z",
      "isPublic": true
    },
    {
      "factKey": "latest_release",
      "category": "release",
      "label": "Latest release",
      "value": "0.3.24",
      "href": "https://clawhub.ai/gongyu0918-debug/skill-usefulness-audit",
      "sourceUrl": "https://clawhub.ai/gongyu0918-debug/skill-usefulness-audit",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-08-24T10:08:57.284Z",
      "isPublic": true
    },
    {
      "factKey": "handshake_status",
      "category": "security",
      "label": "Handshake status",
      "value": "UNKNOWN",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-gongyu0918-debug-skill-usefulness-audit/trust",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-gongyu0918-debug-skill-usefulness-audit/trust",
      "sourceType": "trust",
      "confidence": "medium",
      "observedAt": null,
      "isPublic": true
    }
  ],
  "events": [
    {
      "eventType": "release",
      "title": "Release 0.3.24",
      "description": "Add a package-marker profile for Chinese-only derivatives; the ClawHub edition remains bilingual.",
      "href": "https://clawhub.ai/gongyu0918-debug/skill-usefulness-audit",
      "sourceUrl": "https://clawhub.ai/gongyu0918-debug/skill-usefulness-audit",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-08-24T10:08:57.284Z",
      "isPublic": true
    }
  ]
}

Record generated Oct 9, 2026.

Sponsored

Ads related to skill-usefulness-audit and adjacent AI workflows.