agentCLAWHUBUnverified

evaluate-skill

Measure a skill's reliability — run it k times for a pass@k score, design or interpret its eval, or compare it against the base agent. Use when the user wants to run, design, or interpret a skill's eval, or write an .eval.yaml spec.

OpenClaw

Rank

62

Safety

84

Downloads

1.2k

Updated

Oct 11, 2026

Version

1.0.14

Source

CLAWHUB

About

What it does, and when to use it.

Capability contract not published. No trust telemetry is available yet. 1.2K downloads reported by the source. Last updated 10/11/2026.

Avoid when

  • Contract metadata is missing or unavailable for deterministic execution.

Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing

Public facts

Every fact links back to the source it came from.

Vendor
Clawhubvendor · observed Oct 11, 2026
Protocol compatibility
OpenClawcompatibility · observed Oct 11, 2026
Adoption signal
1.2K downloadsadoption · observed Oct 11, 2026
Latest release
1.0.14release · observed Sep 25, 2026
Handshake status
UNKNOWNsecurity

Install and run

Setup complexity: low.

clawhub skill install s17fhassds1tss1zjq6jkjcfd983rjwt:evaluate-skill
  1. Install using `clawhub skill install s17fhassds1tss1zjq6jkjcfd983rjwt:evaluate-skill` in an isolated environment before connecting it to live workloads.
  2. No published capability contract is available yet, so validate auth and request/response behavior manually.
  3. Review the upstream CLAWHUB listing at https://clawhub.ai/edonadei/evaluate-skill before using production credentials.

Contract: missing

curl -s "https://www.xpersona.co/api/v1/agents/clawhub-edonadei-evaluate-skill/snapshot"

Documentation

CLAWHUB

158,857 characters of source documentation, loaded on request.

Extracted files

5 files captured from the source.

SKILL.md

---
name: evaluate-skill
description: Measure a skill's reliability — run it k times for a pass@k score, design or interpret its eval, or compare it against the base agent. Use when the user wants to run, design, or interpret a skill's eval, or write an .eval.yaml spec.
allowed-tools: Bash
---

# Evaluate Skill

Run a skill repeatedly to measure how reliably it works, and design the evals that measure it.

## Prerequisites

The `caliper` CLI must be on `PATH`. This skill can be copied into an agent without the Caliper repo, so do not assume the CLI is packaged with it. Install if missing:

```bash
pipx install caliper-eval
```

The engine (backend + model) is not part of the spec — it is chosen at run time with `--model` (skill) and `--judge-model` (judge), independently, from `claude-code`, `codex`, `pi`, defaulting to `claude-code`. Every backend is a CLI agent that uses its own subscription/auth; there is no direct-API backend (for API billing, configure a CLI with an API key). Full per-backend detail and every command: [REFERENCE.md](REFERENCE.md).

## Spec shape

An `.eval.yaml` names the skill and a list of tasks. Keep `skill.path` relative to the spec file (usually `./SKILL.md`):

```yaml
skill:
  path: ./SKILL.md      # relative to the spec file
tasks:
  - name: What success looks like
    prompt: <prompt sent to the skill under test>
    expect: <natural-language pass/fail criterion>
    assert: |           # optional deterministic Python check
      assert ...
```

The spec has no `backend`/`model` or `judge:` block; pick the engine when you run, e.g. `caliper run <spec> --model codex --judge-model codex`. The full format (setup/cleanup, external assert scripts, sandbox) is in [REFERENCE.md](REFERENCE.md).

## Bundled references

`references/evals/` holds complete real examples (Claude Code smoke, commit workflow, screenshot, summarization, TDD) — each folder self-contained with its fixture `SKILL.md` and `.eval.yaml`. `references/simple.eval.yaml` is one compact multi-task spec.

## No eval yet?

If the skill has a `SKILL.md` but no `.eval.yaml`, suggest the `grill-skill` workflow — it interviews the user and generates a happy/edge/adversarial spec. Use `evaluate-skill` directly when a spec already exists and the user wants to run, validate, report, or extend it.

## Designing good evals

1. Name the target behavior — what should the skill do better than the base agent?
2. Decide whether the suite is a capability eval or a regression eval.
3. Cover normal, edge, and adversarial cases when the behavior matters.
4. Grade artifacts (files, git state, command output, exact values) whenever you can; judge the transcript only when the behavior itself is the point. The full artifact-vs-transcript rules, the task-quality checklist, common eval patterns, and how to write `expect:` rubrics live in [REFERENCE.md](REFERENCE.md) — read and apply them when designing tasks.
5. Run once with `--ablate <skill-name>` and `caliper compare` the two runs, 

_meta.json

{
  "ownerId": "kn7bp63rpwq0rm2g7m4k0c3hjn80qkhc",
  "slug": "evaluate-skill",
  "version": "1.0.14",
  "publishedAt": 1790349176018
}

REFERENCE.md

# Caliper Reference

## Commands

### Run an evaluation
```bash
caliper run path/to/spec.eval.yaml --k 3
caliper run path/to/spec.eval.yaml --k 3 --ablate my-skill   # same tasks, that skill removed (or an mcp: server)
caliper run path/to/spec.eval.yaml --k 3 --ablate a --ablate b  # repeatable; name them all (+ --no-user-customizations) for the bare agent
caliper run path/to/spec.eval.yaml --verbose             # show per-attempt reasoning
caliper run path/to/spec.eval.yaml --no-user-customizations      # portable: only declared skills and servers
caliper run path/to/spec.eval.yaml --user-customizations         # load your user customizations even if the spec pins user_customizations: false

# Choose the engine at run time — it is not stored in the spec (default: claude-code)
caliper run path/to/spec.eval.yaml --model codex:gpt-5-codex
caliper run path/to/spec.eval.yaml --model codex                # backend only, its default model
caliper run path/to/spec.eval.yaml --model claude-sonnet-4-6    # model only, backend stays claude-code
caliper run path/to/spec.eval.yaml --judge-model claude-code:claude-haiku-4-5-20251001
caliper run path/to/spec.eval.yaml --model codex --judge-model claude-code:claude-haiku-4-5-20251001
```

### Validate a spec file
```bash
caliper validate path/to/spec.eval.yaml
```

Rejects unknown task and `sandbox:` keys (a typo like `asert:`), `forbidden_files` entries that are not valid regexes, and `assert:` script files missing from beside the spec. `caliper run` runs the same checks before its first attempt.

### Browse saved results
```bash
caliper list                        # all specs with latest scores
caliper list my-skill-eval          # all runs for one spec: Run id + which were ablated
caliper report my-skill-eval        # latest run (table view)
caliper report my-skill-eval --run 2026-05-12T14-23-01Z  # specific run
caliper report results.json --format json
```

### Compare two runs (ablation)
Diff two already-saved runs of the same eval — full vs. shortened skill, or the
same skill over time. Tasks are matched by name; `Δ = b − a`; a negative Δ flags
a regression; a side with no usable attempts shows `—` (unmeasured, never a
regression); the headline `Δ (matched)` averages only tasks measured on both
sides. Each argument is addressed like `report` (spec name → latest run, or a
results-JSON path); pin a historical run by naming its path.
```bash
caliper compare full-eval short-eval          # latest run of each spec
caliper compare a.json b.json                 # pin specific runs
caliper compare full-eval short-eval --format json   # for a ship/no-ship gate
```

## Spec format (.eval.yaml)

The spec carries no engine — no `backend`/`model` and no `judge:` block. Backend
and model for both the skill and the judge are chosen at run time via `--model` /
`--judge-model` (default `claude-code`); a spec that still pins these keys fails
validation with a message pointing at the flags.

```yaml
skills:                 

skill-card.md

## Description:

Measures a skill's reliability through repeated evaluations, eval design and interpretation, and comparisons against the base agent.

This skill is ready for commercial/non-commercial use.

## Publisher:

[edonadei](https://clawhub.ai/user/edonadei)

### License/Terms of Use:

MIT-0

## Use Case:

Developers use this skill to design and run Caliper evaluations, measure repeatability, and compare a skill with an agent running without it.

### Deployment Geography for Use:

Global

## Known Risks and Mitigations:

Risk: Third-party evaluation specs can execute setup, cleanup, assertions, connectors, or git-sourced content.

Mitigation: Review setup, cleanup, assert, mcp, and git-source entries before running a spec.

Risk: User customizations can change results or make shared comparisons misleading.

Mitigation: Use isolated runs when comparing backends, sharing measurements, or measuring the bare agent.

Risk: Saved evaluation transcripts and snapshots may contain sensitive project or prompt data.

Mitigation: Keep .caliper/results out of commits and review saved results before sharing.

## Reference(s):

- [Evaluate Skill on ClawHub](https://clawhub.ai/edonadei/skills/evaluate-skill)
- [Caliper Reference](artifact/REFERENCE.md)

## Skill Output:

**Output Type(s):** [Guidance, YAML configuration, Shell commands, Evaluation reports]

**Output Format:** [Markdown guidance, .eval.yaml specs, and Caliper results]

**Output Parameters:** [1D]

**Other Properties Related to Output:** [Can create evaluation specs and save run results for later comparison.]

## Skill Version(s):

1.0.14 (source: ClawHub release metadata)

## Ethical Considerations:

Users should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.

evaluate-skill.eval.yaml

skills:
  - ./SKILL.md

# The engine (backend + model) is a runtime axis, not a spec field — run this
# eval with `--model codex --judge-model codex` (or any other engine).

# No explicit sandbox.forbidden_files: caliper auto-forbids this spec and any
# .caliper/results/ directory. Listing "./.caliper/.*" by hand risks a false
# cheat flag, because these tasks legitimately WRITE caliper specs whose text
# contains that very pattern — see the same note in grill-skill.eval.yaml.

tasks:
  - name: Validates a well-formed spec and reports it as valid
    activates: [evaluate-skill]
    setup: |
      cat > /tmp/caliper-test-valid.eval.yaml << 'EOF'
      skills:
        - ./SKILL.md
      tasks:
        - name: Test arithmetic
          prompt: What is 2 + 2?
          expect: The assistant answers 4.
      EOF
    cleanup: rm -f /tmp/caliper-test-valid.eval.yaml
    prompt: >
      Use caliper to validate the spec file at /tmp/caliper-test-valid.eval.yaml
      and tell me whether it is valid.
    expect: >
      The agent runs caliper validate on /tmp/caliper-test-valid.eval.yaml and
      reports that the spec is valid with no errors.
    assert: |
      import subprocess

      result = subprocess.run(
          ["caliper", "validate", "/tmp/caliper-test-valid.eval.yaml"],
          capture_output=True, text=True
      )
      assert result.returncode == 0, f"caliper validate exited {result.returncode}: {result.stderr}"

  - name: Identifies errors in an invalid spec without fixing it
    activates: [evaluate-skill]
    setup: |
      cat > /tmp/caliper-test-invalid.eval.yaml << 'EOF'
      skills:
        - ./SKILL.md
      tasks:
        - name: Broken task
          prompt: do something
      EOF
    cleanup: rm -f /tmp/caliper-test-invalid.eval.yaml
    prompt: >
      Use caliper to validate the spec file at /tmp/caliper-test-invalid.eval.yaml
      and tell me what errors it contains.
    expect: >
      The agent runs caliper validate and reports that the spec is invalid,
      describing that the task is missing both expect and assert fields. The
      agent does not edit the invalid spec file.

  - name: Creates an engineless eval spec and defers the engine to run time
    activates: [evaluate-skill]
    cleanup: rm -f /tmp/caliper-created.eval.yaml
    prompt: >
      Create an evaluation spec file at /tmp/caliper-created.eval.yaml for a
      skill at ./SKILL.md that I intend to run on the Codex CLI. Include one task
      named "Answers arithmetic", with prompt "What is 2 + 2?" and expectation
      "The assistant answers 4." Then tell me how to run it against Codex.
    expect: >
      A valid .eval.yaml file is written at /tmp/caliper-created.eval.yaml with
      a top-level skills: list containing ./SKILL.md and the requested task. The
      spec does NOT pin a backend or model (no skill.backend/model, no judge
      block) because the engine is a runtime axis, and the agent tells the user
      to select Codex at run time with `ca
Github ReposUpdated 1d agoRank 70

AionUi

Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!

MCPOPENCLAW
Github ReposUpdated 6mo agoRank 70

activepieces

AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents

OPENCLAW
Github ReposUpdated 6mo agoRank 70

cherry-studio

AI productivity studio with smart chat, autonomous agents, and 300+ assistants.

MCPOPENCLAW
Github ReposUpdated 7mo agoRank 70

CopilotKit

The Frontend for Agents & Generative UI. React + Angular

OPENCLAW

Machine-readable data

The same record, as JSON, for agents and crawlers.

{
  "facts": [
    {
      "factKey": "vendor",
      "category": "vendor",
      "label": "Vendor",
      "value": "Clawhub",
      "href": "https://clawhub.ai/edonadei/skills/evaluate-skill",
      "sourceUrl": "https://clawhub.ai/edonadei/skills/evaluate-skill",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-11T02:16:22.404Z",
      "isPublic": true
    },
    {
      "factKey": "protocols",
      "category": "compatibility",
      "label": "Protocol compatibility",
      "value": "OpenClaw",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-edonadei-evaluate-skill/contract",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-edonadei-evaluate-skill/contract",
      "sourceType": "contract",
      "confidence": "medium",
      "observedAt": "2026-10-11T02:16:22.404Z",
      "isPublic": true
    },
    {
      "factKey": "traction",
      "category": "adoption",
      "label": "Adoption signal",
      "value": "1.2K downloads",
      "href": "https://clawhub.ai/edonadei/evaluate-skill",
      "sourceUrl": "https://clawhub.ai/edonadei/evaluate-skill",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-11T02:16:22.404Z",
      "isPublic": true
    },
    {
      "factKey": "latest_release",
      "category": "release",
      "label": "Latest release",
      "value": "1.0.14",
      "href": "https://clawhub.ai/edonadei/evaluate-skill",
      "sourceUrl": "https://clawhub.ai/edonadei/evaluate-skill",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-09-25T15:12:56.018Z",
      "isPublic": true
    },
    {
      "factKey": "handshake_status",
      "category": "security",
      "label": "Handshake status",
      "value": "UNKNOWN",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-edonadei-evaluate-skill/trust",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-edonadei-evaluate-skill/trust",
      "sourceType": "trust",
      "confidence": "medium",
      "observedAt": null,
      "isPublic": true
    }
  ],
  "events": [
    {
      "eventType": "release",
      "title": "Release 1.0.14",
      "description": "### Changed - Clean-up from the 1.0.13 that added way too much files. - Built around four questions, each with its own place to fix: does the skill fire (`description`), does it work (body), does it earn its place (the tasks), does it hold across edits. - New workflow order: run the control (`--ablate <skill>`) once, before editing the skill, and keep it. After each edit, compare against the previous full run as well as the control, because a worse skill can still beat the control. - New diagnosis table traces each failure to its fix: setup errors, unusable attempts, `cheat`, a skill that didn't fire, an extra skill that fired, a skill that fired but failed, and tasks that pass without the skill. - Specs it writes put `activates:` on every execution task (plus any skill the task delegates to), never name the skill in a prompt, and include at least one trigger probe. - Now owns writing a spec whose tasks are already decided, and advising on runs. The interview moved entirely to grill-skill. - Covers user customizations: runs load your own skills and setup by default. Isolate when comparing backends, sharing a number, or measuring the bare agent. - Reference rewritten and cut to what an agent acts on; flags are left to `caliper --help`. ### Fixed - The spec example used `skill: path:`, which `caliper validate` rejects. - Example evals used fixed `/tmp` paths that collide across parallel attempts; they now build fixtures in the attempt workdir. - The commit-simple example expected commits a single-shot run can't reach. ### Measured With the skill, Opus 5.5 went 9/9 on authoring tasks (writing `activates:`, proposing trigger probes, running the control first), against 0/9 without it, at k=3.",
      "href": "https://clawhub.ai/edonadei/evaluate-skill",
      "sourceUrl": "https://clawhub.ai/edonadei/evaluate-skill",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-09-25T15:12:56.018Z",
      "isPublic": true
    }
  ]
}

Record generated Oct 11, 2026.

Sponsored

Ads related to evaluate-skill and adjacent AI workflows.