evaluate-skill
Measure a skill's reliability — run it k times for a pass@k score, design or interpret its eval, or compare it against the base agent. Use when the user wants to run, design, or interpret a skill's eval, or write an .eval.yaml spec.
Rank
62
Safety
84
Downloads
1.2k
Updated
Oct 11, 2026
Version
1.0.14
Source
CLAWHUB
About
What it does, and when to use it.
Capability contract not published. No trust telemetry is available yet. 1.2K downloads reported by the source. Last updated 10/11/2026.
Avoid when
- Contract metadata is missing or unavailable for deterministic execution.
Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing
Public facts
Every fact links back to the source it came from.
- Vendor
- Clawhubvendor · observed Oct 11, 2026
- Protocol compatibility
- OpenClawcompatibility · observed Oct 11, 2026
- Adoption signal
- 1.2K downloadsadoption · observed Oct 11, 2026
- Latest release
- 1.0.14release · observed Sep 25, 2026
- Handshake status
- UNKNOWNsecurity
Install and run
Setup complexity: low.
clawhub skill install s17fhassds1tss1zjq6jkjcfd983rjwt:evaluate-skill- Install using `clawhub skill install s17fhassds1tss1zjq6jkjcfd983rjwt:evaluate-skill` in an isolated environment before connecting it to live workloads.
- No published capability contract is available yet, so validate auth and request/response behavior manually.
- Review the upstream CLAWHUB listing at https://clawhub.ai/edonadei/evaluate-skill before using production credentials.
Contract: missing
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-edonadei-evaluate-skill/snapshot"
Documentation
CLAWHUB
158,857 characters of source documentation, loaded on request.
Extracted files
5 files captured from the source.
SKILL.md
---
name: evaluate-skill
description: Measure a skill's reliability — run it k times for a pass@k score, design or interpret its eval, or compare it against the base agent. Use when the user wants to run, design, or interpret a skill's eval, or write an .eval.yaml spec.
allowed-tools: Bash
---
# Evaluate Skill
Run a skill repeatedly to measure how reliably it works, and design the evals that measure it.
## Prerequisites
The `caliper` CLI must be on `PATH`. This skill can be copied into an agent without the Caliper repo, so do not assume the CLI is packaged with it. Install if missing:
```bash
pipx install caliper-eval
```
The engine (backend + model) is not part of the spec — it is chosen at run time with `--model` (skill) and `--judge-model` (judge), independently, from `claude-code`, `codex`, `pi`, defaulting to `claude-code`. Every backend is a CLI agent that uses its own subscription/auth; there is no direct-API backend (for API billing, configure a CLI with an API key). Full per-backend detail and every command: [REFERENCE.md](REFERENCE.md).
## Spec shape
An `.eval.yaml` names the skill and a list of tasks. Keep `skill.path` relative to the spec file (usually `./SKILL.md`):
```yaml
skill:
path: ./SKILL.md # relative to the spec file
tasks:
- name: What success looks like
prompt: <prompt sent to the skill under test>
expect: <natural-language pass/fail criterion>
assert: | # optional deterministic Python check
assert ...
```
The spec has no `backend`/`model` or `judge:` block; pick the engine when you run, e.g. `caliper run <spec> --model codex --judge-model codex`. The full format (setup/cleanup, external assert scripts, sandbox) is in [REFERENCE.md](REFERENCE.md).
## Bundled references
`references/evals/` holds complete real examples (Claude Code smoke, commit workflow, screenshot, summarization, TDD) — each folder self-contained with its fixture `SKILL.md` and `.eval.yaml`. `references/simple.eval.yaml` is one compact multi-task spec.
## No eval yet?
If the skill has a `SKILL.md` but no `.eval.yaml`, suggest the `grill-skill` workflow — it interviews the user and generates a happy/edge/adversarial spec. Use `evaluate-skill` directly when a spec already exists and the user wants to run, validate, report, or extend it.
## Designing good evals
1. Name the target behavior — what should the skill do better than the base agent?
2. Decide whether the suite is a capability eval or a regression eval.
3. Cover normal, edge, and adversarial cases when the behavior matters.
4. Grade artifacts (files, git state, command output, exact values) whenever you can; judge the transcript only when the behavior itself is the point. The full artifact-vs-transcript rules, the task-quality checklist, common eval patterns, and how to write `expect:` rubrics live in [REFERENCE.md](REFERENCE.md) — read and apply them when designing tasks.
5. Run once with `--ablate <skill-name>` and `caliper compare` the two runs, _meta.json
{
"ownerId": "kn7bp63rpwq0rm2g7m4k0c3hjn80qkhc",
"slug": "evaluate-skill",
"version": "1.0.14",
"publishedAt": 1790349176018
}REFERENCE.md
# Caliper Reference ## Commands ### Run an evaluation ```bash caliper run path/to/spec.eval.yaml --k 3 caliper run path/to/spec.eval.yaml --k 3 --ablate my-skill # same tasks, that skill removed (or an mcp: server) caliper run path/to/spec.eval.yaml --k 3 --ablate a --ablate b # repeatable; name them all (+ --no-user-customizations) for the bare agent caliper run path/to/spec.eval.yaml --verbose # show per-attempt reasoning caliper run path/to/spec.eval.yaml --no-user-customizations # portable: only declared skills and servers caliper run path/to/spec.eval.yaml --user-customizations # load your user customizations even if the spec pins user_customizations: false # Choose the engine at run time — it is not stored in the spec (default: claude-code) caliper run path/to/spec.eval.yaml --model codex:gpt-5-codex caliper run path/to/spec.eval.yaml --model codex # backend only, its default model caliper run path/to/spec.eval.yaml --model claude-sonnet-4-6 # model only, backend stays claude-code caliper run path/to/spec.eval.yaml --judge-model claude-code:claude-haiku-4-5-20251001 caliper run path/to/spec.eval.yaml --model codex --judge-model claude-code:claude-haiku-4-5-20251001 ``` ### Validate a spec file ```bash caliper validate path/to/spec.eval.yaml ``` Rejects unknown task and `sandbox:` keys (a typo like `asert:`), `forbidden_files` entries that are not valid regexes, and `assert:` script files missing from beside the spec. `caliper run` runs the same checks before its first attempt. ### Browse saved results ```bash caliper list # all specs with latest scores caliper list my-skill-eval # all runs for one spec: Run id + which were ablated caliper report my-skill-eval # latest run (table view) caliper report my-skill-eval --run 2026-05-12T14-23-01Z # specific run caliper report results.json --format json ``` ### Compare two runs (ablation) Diff two already-saved runs of the same eval — full vs. shortened skill, or the same skill over time. Tasks are matched by name; `Δ = b − a`; a negative Δ flags a regression; a side with no usable attempts shows `—` (unmeasured, never a regression); the headline `Δ (matched)` averages only tasks measured on both sides. Each argument is addressed like `report` (spec name → latest run, or a results-JSON path); pin a historical run by naming its path. ```bash caliper compare full-eval short-eval # latest run of each spec caliper compare a.json b.json # pin specific runs caliper compare full-eval short-eval --format json # for a ship/no-ship gate ``` ## Spec format (.eval.yaml) The spec carries no engine — no `backend`/`model` and no `judge:` block. Backend and model for both the skill and the judge are chosen at run time via `--model` / `--judge-model` (default `claude-code`); a spec that still pins these keys fails validation with a message pointing at the flags. ```yaml skills:
skill-card.md
## Description: Measures a skill's reliability through repeated evaluations, eval design and interpretation, and comparisons against the base agent. This skill is ready for commercial/non-commercial use. ## Publisher: [edonadei](https://clawhub.ai/user/edonadei) ### License/Terms of Use: MIT-0 ## Use Case: Developers use this skill to design and run Caliper evaluations, measure repeatability, and compare a skill with an agent running without it. ### Deployment Geography for Use: Global ## Known Risks and Mitigations: Risk: Third-party evaluation specs can execute setup, cleanup, assertions, connectors, or git-sourced content. Mitigation: Review setup, cleanup, assert, mcp, and git-source entries before running a spec. Risk: User customizations can change results or make shared comparisons misleading. Mitigation: Use isolated runs when comparing backends, sharing measurements, or measuring the bare agent. Risk: Saved evaluation transcripts and snapshots may contain sensitive project or prompt data. Mitigation: Keep .caliper/results out of commits and review saved results before sharing. ## Reference(s): - [Evaluate Skill on ClawHub](https://clawhub.ai/edonadei/skills/evaluate-skill) - [Caliper Reference](artifact/REFERENCE.md) ## Skill Output: **Output Type(s):** [Guidance, YAML configuration, Shell commands, Evaluation reports] **Output Format:** [Markdown guidance, .eval.yaml specs, and Caliper results] **Output Parameters:** [1D] **Other Properties Related to Output:** [Can create evaluation specs and save run results for later comparison.] ## Skill Version(s): 1.0.14 (source: ClawHub release metadata) ## Ethical Considerations: Users should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.
evaluate-skill.eval.yaml
skills:
- ./SKILL.md
# The engine (backend + model) is a runtime axis, not a spec field — run this
# eval with `--model codex --judge-model codex` (or any other engine).
# No explicit sandbox.forbidden_files: caliper auto-forbids this spec and any
# .caliper/results/ directory. Listing "./.caliper/.*" by hand risks a false
# cheat flag, because these tasks legitimately WRITE caliper specs whose text
# contains that very pattern — see the same note in grill-skill.eval.yaml.
tasks:
- name: Validates a well-formed spec and reports it as valid
activates: [evaluate-skill]
setup: |
cat > /tmp/caliper-test-valid.eval.yaml << 'EOF'
skills:
- ./SKILL.md
tasks:
- name: Test arithmetic
prompt: What is 2 + 2?
expect: The assistant answers 4.
EOF
cleanup: rm -f /tmp/caliper-test-valid.eval.yaml
prompt: >
Use caliper to validate the spec file at /tmp/caliper-test-valid.eval.yaml
and tell me whether it is valid.
expect: >
The agent runs caliper validate on /tmp/caliper-test-valid.eval.yaml and
reports that the spec is valid with no errors.
assert: |
import subprocess
result = subprocess.run(
["caliper", "validate", "/tmp/caliper-test-valid.eval.yaml"],
capture_output=True, text=True
)
assert result.returncode == 0, f"caliper validate exited {result.returncode}: {result.stderr}"
- name: Identifies errors in an invalid spec without fixing it
activates: [evaluate-skill]
setup: |
cat > /tmp/caliper-test-invalid.eval.yaml << 'EOF'
skills:
- ./SKILL.md
tasks:
- name: Broken task
prompt: do something
EOF
cleanup: rm -f /tmp/caliper-test-invalid.eval.yaml
prompt: >
Use caliper to validate the spec file at /tmp/caliper-test-invalid.eval.yaml
and tell me what errors it contains.
expect: >
The agent runs caliper validate and reports that the spec is invalid,
describing that the task is missing both expect and assert fields. The
agent does not edit the invalid spec file.
- name: Creates an engineless eval spec and defers the engine to run time
activates: [evaluate-skill]
cleanup: rm -f /tmp/caliper-created.eval.yaml
prompt: >
Create an evaluation spec file at /tmp/caliper-created.eval.yaml for a
skill at ./SKILL.md that I intend to run on the Codex CLI. Include one task
named "Answers arithmetic", with prompt "What is 2 + 2?" and expectation
"The assistant answers 4." Then tell me how to run it against Codex.
expect: >
A valid .eval.yaml file is written at /tmp/caliper-created.eval.yaml with
a top-level skills: list containing ./SKILL.md and the requested task. The
spec does NOT pin a backend or model (no skill.backend/model, no judge
block) because the engine is a runtime axis, and the agent tells the user
to select Codex at run time with `caAionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
activepieces
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
CopilotKit
The Frontend for Agents & Generative UI. React + Angular
Machine-readable data
The same record, as JSON, for agents and crawlers.
{
"facts": [
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Clawhub",
"href": "https://clawhub.ai/edonadei/skills/evaluate-skill",
"sourceUrl": "https://clawhub.ai/edonadei/skills/evaluate-skill",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-11T02:16:22.404Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-edonadei-evaluate-skill/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-edonadei-evaluate-skill/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-11T02:16:22.404Z",
"isPublic": true
},
{
"factKey": "traction",
"category": "adoption",
"label": "Adoption signal",
"value": "1.2K downloads",
"href": "https://clawhub.ai/edonadei/evaluate-skill",
"sourceUrl": "https://clawhub.ai/edonadei/evaluate-skill",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-11T02:16:22.404Z",
"isPublic": true
},
{
"factKey": "latest_release",
"category": "release",
"label": "Latest release",
"value": "1.0.14",
"href": "https://clawhub.ai/edonadei/evaluate-skill",
"sourceUrl": "https://clawhub.ai/edonadei/evaluate-skill",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-09-25T15:12:56.018Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-edonadei-evaluate-skill/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-edonadei-evaluate-skill/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
],
"events": [
{
"eventType": "release",
"title": "Release 1.0.14",
"description": "### Changed - Clean-up from the 1.0.13 that added way too much files. - Built around four questions, each with its own place to fix: does the skill fire (`description`), does it work (body), does it earn its place (the tasks), does it hold across edits. - New workflow order: run the control (`--ablate <skill>`) once, before editing the skill, and keep it. After each edit, compare against the previous full run as well as the control, because a worse skill can still beat the control. - New diagnosis table traces each failure to its fix: setup errors, unusable attempts, `cheat`, a skill that didn't fire, an extra skill that fired, a skill that fired but failed, and tasks that pass without the skill. - Specs it writes put `activates:` on every execution task (plus any skill the task delegates to), never name the skill in a prompt, and include at least one trigger probe. - Now owns writing a spec whose tasks are already decided, and advising on runs. The interview moved entirely to grill-skill. - Covers user customizations: runs load your own skills and setup by default. Isolate when comparing backends, sharing a number, or measuring the bare agent. - Reference rewritten and cut to what an agent acts on; flags are left to `caliper --help`. ### Fixed - The spec example used `skill: path:`, which `caliper validate` rejects. - Example evals used fixed `/tmp` paths that collide across parallel attempts; they now build fixtures in the attempt workdir. - The commit-simple example expected commits a single-shot run can't reach. ### Measured With the skill, Opus 5.5 went 9/9 on authoring tasks (writing `activates:`, proposing trigger probes, running the control first), against 0/9 without it, at k=3.",
"href": "https://clawhub.ai/edonadei/evaluate-skill",
"sourceUrl": "https://clawhub.ai/edonadei/evaluate-skill",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-09-25T15:12:56.018Z",
"isPublic": true
}
]
}Record generated Oct 11, 2026.
