agentCLAWHUBUnverified

grill-skill

Build and harden a skill with evals — interview to design its eval tasks, then run, measure, and iterate. Use when the user wants to create or improve a skill's eval, or run the create → test → improve loop for a skill.

OpenClaw

Rank

62

Safety

84

Downloads

1.2k

Updated

Oct 11, 2026

Version

1.0.11

Source

CLAWHUB

About

What it does, and when to use it.

Capability contract not published. No trust telemetry is available yet. 1.2K downloads reported by the source. Last updated 10/11/2026.

Avoid when

  • Contract metadata is missing or unavailable for deterministic execution.

Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing

Public facts

Every fact links back to the source it came from.

Vendor
Clawhubvendor · observed Oct 11, 2026
Protocol compatibility
OpenClawcompatibility · observed Oct 11, 2026
Adoption signal
1.2K downloadsadoption · observed Oct 11, 2026
Latest release
1.0.11release · observed Sep 25, 2026
Handshake status
UNKNOWNsecurity

Install and run

Setup complexity: low.

clawhub skill install s17fhassds1tss1zjq6jkjcfd983rjwt:grill-skill
  1. Install using `clawhub skill install s17fhassds1tss1zjq6jkjcfd983rjwt:grill-skill` in an isolated environment before connecting it to live workloads.
  2. No published capability contract is available yet, so validate auth and request/response behavior manually.
  3. Review the upstream CLAWHUB listing at https://clawhub.ai/edonadei/grill-skill before using production credentials.

Contract: missing

curl -s "https://www.xpersona.co/api/v1/agents/clawhub-edonadei-grill-skill/snapshot"

Documentation

CLAWHUB

147,506 characters of source documentation, loaded on request.

Extracted files

5 files captured from the source.

SKILL.md

---
name: grill-skill
description: Build and harden a skill with evals — interview to design its eval tasks, then run, measure, and iterate. Use when the user wants to create or improve a skill's eval, or run the create → test → improve loop for a skill.
allowed-tools: Bash, Read, Write, Edit
---

# Grill Skill

Interview the user to design a skill's eval, then loop run → measure → improve until it ships. Requires `caliper` (`pipx install caliper-eval` if missing). Commands, spec skeleton, and expect/assert guidance: [REFERENCE.md](REFERENCE.md).

## Entry point

`/grill-skill [path]` — optional path to a `SKILL.md`.

- **Path given** — use it.
- **No path** — look for `SKILL.md` in the cwd; if found, confirm before proceeding, else ask where it is.

## Phase 1 — Understand

Read the `SKILL.md`. Summarize what it does, when it triggers, and what a successful run looks like. Ask the user to confirm your reading. **Wait for confirmation before continuing.**

## Phase 2 — Detect eval mode

Look for `*.eval.yaml` beside the `SKILL.md` (try `<dir-name>.eval.yaml` first).

- **None** → New eval. **Found** → Gap-fill.

Interview one question at a time and wait for each answer. Never invent the user's answers or write the spec before interviewing.

### New eval — three tasks

Elicit three tasks, one question at a time:

1. **Happy path** — the most common successful use. What did the agent do, and what would confirm it worked?
2. **Edge case** — a tricky-but-valid input that might trip the raw agent.
3. **Adversarial** — what the skill should refuse or avoid.

Turn each answer into a task: a realistic `prompt`, an observable `expect`, and an `assert` when the outcome is checkable (see [REFERENCE.md](REFERENCE.md)). Show the proposed YAML and confirm before writing.

Write the spec beside `SKILL.md`, named `<dir-name>.eval.yaml`, with `skill.path: ./SKILL.md` and `claude-code` as the default backend for both `skill` and `judge` unless the SKILL.md targets another.

### Gap-fill

Read the existing spec and report its tasks. **Ask what behaviors are missing or under-tested before proposing or writing anything** — even if the user only asked you to inspect it, report first, then ask. Sharpen each gap into a task, show it, and confirm before writing it in.

## Whose setup is measured

Runs load the user's own customizations by default (user skills, plugins, rules, settings and connectors; see REFERENCE.md for backend exceptions), which answers "does my skill work in *my* agent?". **Isolate** (`--no-user-customizations`, or `user_customizations: false` in the spec) when comparing backends or models, when the number leaves this machine (shared, published, compared with someone else's run), or when measuring the bare agent: each setup is different, so otherwise part of the delta is the setups. `--ablate` of the user's own skill needs no isolation, since both runs load the same setup.

**Always tell the user which mode ran** and what it loaded, from the report he

_meta.json

{
  "ownerId": "kn7bp63rpwq0rm2g7m4k0c3hjn80qkhc",
  "slug": "grill-skill",
  "version": "1.0.11",
  "publishedAt": 1790347950116
}

REFERENCE.md

# Grill Skill Reference

## Caliper commands used by this skill

```bash
# Check spec is valid before running (unknown task or sandbox keys, bad forbidden_files
# regexes, missing assert: files)
caliper validate path/to/spec.eval.yaml

# First run — fast, catches spec errors
caliper run path/to/spec.eval.yaml --k 1

# Reliability run — after iterating on the skill
caliper run path/to/spec.eval.yaml --k 3

# Ablated run — before committing, proves the skill makes a difference.
# Run once and keep it: it cannot move when the skill's text changes.
# A declared mcp: server can be ablated the same way; qualify as skill:/mcp:
# if both declare the name.
caliper run path/to/spec.eval.yaml --k 3 --ablate my-skill
# Then diff it against the full run. A bare spec name resolves to that spec's
# LATEST run, so address the older side by its saved results path.
caliper compare .caliper/results/<spec>/<ablated-run>.json <spec>

# Choose the engine at run time — it is not stored in the spec (default: claude-code)
caliper run path/to/spec.eval.yaml --model codex:gpt-5-codex
caliper run path/to/spec.eval.yaml --model codex
caliper run path/to/spec.eval.yaml --judge-model claude-code:claude-haiku-4-5-20251001

# Runs load your user customizations (skills, plugins, rules, settings and connectors) by default; isolate for a
# portable score (or pin user_customizations: false in the spec)
caliper run path/to/spec.eval.yaml --no-user-customizations

# Browse past results
caliper list
caliper report path/to/spec.eval.yaml

# Compare two saved runs of the same eval (ablation: full vs. shortened, or over time)
caliper compare full-eval short-eval           # spec name -> latest run, or a results-JSON path
caliper compare a.json b.json --format json     # per-task Δ, regression flags, for scripting
```

`caliper compare <A> <B>` diffs two already-saved runs task by task: tasks are
matched by name, `Δ = b − a`, a negative Δ flags a regression (any-below), and a
side with no usable attempts shows `—` (unmeasured, never a regression) so
infra/judge noise can't fake a loss. Under the success-rate headline, `compare` also
shows **token and wall-clock deltas** (green = cheaper) — the "same quality, 40%
fewer tokens" signal an ablation looks for. These are secondary: a token/time
change is **never** a regression (only the score is), and dollar cost is not tracked
(tokens are the volume signal). Each attempt in the report also shows its tokens
next to its duration under `--verbose`.

`compare` also reports **skill drift** — a member of the neighbourhood whose
*text* changed between the two runs, read from the per-file hashes in each run's
snapshots. It is graded by provenance, not role: a drifted **git source** warns,
because the spec claimed where those bytes came from and the delta you are
reading is confounded; a drifted **path source** is shown without alarm, because
nothing was promised about a working file and that edit is usually the thing the
run exists to measure.

```
 ⚠ 

skill-card.md

## Description:

Interviews skill authors to design eval tasks, then uses Caliper results to test and improve their skills.

This skill is ready for commercial/non-commercial use.

## Publisher:

[edonadei](https://clawhub.ai/user/edonadei)

### License/Terms of Use:

MIT-0

## Use Case:

Skill authors and developers use this skill to design evaluation tasks, run Caliper tests, and identify changes that improve skill reliability.

### Deployment Geography for Use:

Global

## Known Risks and Mitigations:

Risk: The skill can read and edit selected skill files and write evaluation specifications.

Mitigation: Use it in a trusted project and review generated evaluation YAML before confirming writes.

Risk: Caliper tests may execute task setup or cleanup commands and load local agent customizations.

Mitigation: Review test commands and use isolated Caliper mode for portable or shared results.

## Reference(s):

- [Grill Skill on ClawHub](https://clawhub.ai/edonadei/skills/grill-skill)
- [Grill Skill Reference](artifact/REFERENCE.md)

## Skill Output:

**Output Type(s):** [Guidance, Configuration, Shell commands]

**Output Format:** [Markdown guidance and YAML evaluation specifications]

**Output Parameters:** [1D]

**Other Properties Related to Output:** [Includes test results and comparisons when Caliper is run.]

## Skill Version(s):

1.0.11 (source: ClawHub release metadata)

## Ethical Considerations:

Users should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.

grill-skill.eval.yaml

skills:
  - ./SKILL.md

# The engine (backend + model) is a runtime axis, not a spec field — this eval
# runs on whatever `--model` / `--judge-model` select (default claude-code).

# No explicit sandbox.forbidden_files: caliper already auto-forbids the eval
# spec and any .caliper/results/ directory. Listing "./.caliper/.*" here caused a
# false cheat flag, because these grill-skill tasks legitimately WRITE caliper
# specs whose text contains that very pattern.

tasks:
  # Task 1 — First-turn interview discipline (open prompt, no eval present).
  # Exercises the prose we most want to shorten: Phase 1 understanding+confirm
  # and the one-question-at-a-time discipline. Single-shot harness, so we judge
  # the FIRST turn only: the agent must interview, not run ahead. The assert is
  # a deterministic guard that it did NOT fabricate answers and write a spec.
  - name: Interviews before generating — asks, then stops (does not run ahead)
    activates: [grill-skill]
    setup: |
      rm -rf /tmp/grill-fresh
      mkdir -p /tmp/grill-fresh
      cat > /tmp/grill-fresh/SKILL.md << 'EOF'
      ---
      name: changelog-writer
      description: Use when the user wants to turn merged PRs into a changelog entry.
      allowed-tools: Bash, Read, Write
      ---

      # Changelog Writer

      Read the merged PRs since the last tag and write a grouped changelog
      entry (Features / Fixes / Chore) to CHANGELOG.md.
      EOF
    cleanup: rm -rf /tmp/grill-fresh
    prompt: >
      I want to create a caliper eval for my skill at
      /tmp/grill-fresh/SKILL.md. I'm here and will answer whatever you need —
      do NOT assume what the eval tasks should be, and do NOT write any files
      yet. What do you need to know from me to get started?
    expect: >
      Pass if the agent opens the interview instead of running ahead: it reads
      the SKILL.md, gives some understanding of the skill, and asks the user for
      input before generating anything — then STOPS to wait. The number of
      questions does not matter. Fail if the agent skips the interview: it
      invents the user's answers, generates the eval tasks itself without
      asking, or writes any .eval.yaml file in this turn.
    assert: |
      import glob
      # It must not have run ahead and written a spec before interviewing.
      specs = glob.glob("/tmp/grill-fresh/*.eval.yaml")
      assert not specs, f"Agent wrote a spec without interviewing: {specs}"

  # Task 2 — Gap-fill detection + no silent overwrite (open prompt, eval present).
  # Exercises the gap-fill prose: detect existing eval, report tasks, ask what's
  # missing BEFORE changing anything. Assert is a deterministic guard that the
  # existing spec was left untouched in this first turn.
  - name: Detects an existing eval and asks before touching it
    activates: [grill-skill]
    setup: |
      rm -rf /tmp/grill-existing
      mkdir -p /tmp/grill-existing
      cat > /tmp/grill-existing/SKILL.md << 'EOF'
      ---
      name
Github ReposUpdated 1d agoRank 70

AionUi

Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!

MCPOPENCLAW
Github ReposUpdated 6mo agoRank 70

activepieces

AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents

OPENCLAW
Github ReposUpdated 6mo agoRank 70

cherry-studio

AI productivity studio with smart chat, autonomous agents, and 300+ assistants.

MCPOPENCLAW
Github ReposUpdated 7mo agoRank 70

CopilotKit

The Frontend for Agents & Generative UI. React + Angular

OPENCLAW

Machine-readable data

The same record, as JSON, for agents and crawlers.

{
  "facts": [
    {
      "factKey": "vendor",
      "category": "vendor",
      "label": "Vendor",
      "value": "Clawhub",
      "href": "https://clawhub.ai/edonadei/skills/grill-skill",
      "sourceUrl": "https://clawhub.ai/edonadei/skills/grill-skill",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-11T04:54:05.694Z",
      "isPublic": true
    },
    {
      "factKey": "protocols",
      "category": "compatibility",
      "label": "Protocol compatibility",
      "value": "OpenClaw",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-edonadei-grill-skill/contract",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-edonadei-grill-skill/contract",
      "sourceType": "contract",
      "confidence": "medium",
      "observedAt": "2026-10-11T04:54:05.694Z",
      "isPublic": true
    },
    {
      "factKey": "traction",
      "category": "adoption",
      "label": "Adoption signal",
      "value": "1.2K downloads",
      "href": "https://clawhub.ai/edonadei/grill-skill",
      "sourceUrl": "https://clawhub.ai/edonadei/grill-skill",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-11T04:54:05.694Z",
      "isPublic": true
    },
    {
      "factKey": "latest_release",
      "category": "release",
      "label": "Latest release",
      "value": "1.0.11",
      "href": "https://clawhub.ai/edonadei/grill-skill",
      "sourceUrl": "https://clawhub.ai/edonadei/grill-skill",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-09-25T14:52:30.116Z",
      "isPublic": true
    },
    {
      "factKey": "handshake_status",
      "category": "security",
      "label": "Handshake status",
      "value": "UNKNOWN",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-edonadei-grill-skill/trust",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-edonadei-grill-skill/trust",
      "sourceType": "trust",
      "confidence": "medium",
      "observedAt": null,
      "isPublic": true
    }
  ],
  "events": [
    {
      "eventType": "release",
      "title": "Release 1.0.11",
      "description": "### Changed - The interview now covers triggering: it asks which skills yours could be confused with, adds a neighbour probe for each, and always proposes a silence probe. - Generated tasks put `activates:` on every execution task and never name the skill in the prompt, so a run can tell a `description` failure from a body failure. - New Phase 4: run the control (`--ablate <skill>`) before editing the skill. A task that passes without the skill gets sharpened before any iteration. - New Phase 5: each failure is traced to its fix (`description`, body, task, or setup). After each edit, compare against the previous full run and against the kept control. - Warns that the harness is single-shot: for a skill that asks before acting, tasks judge the first turn. - Now focused on the interview (helping you decide what to test). Writing a spec whose tasks you've already decided moved to evaluate-skill. - Reference rewritten around authoring: spec skeleton with probes, task-writing rules, attempt workdir, user customizations. ### Fixed - It told the agent to write a backend into the spec, which `caliper validate` rejects.",
      "href": "https://clawhub.ai/edonadei/grill-skill",
      "sourceUrl": "https://clawhub.ai/edonadei/grill-skill",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-09-25T14:52:30.116Z",
      "isPublic": true
    }
  ]
}

Record generated Oct 11, 2026.

Sponsored

Ads related to grill-skill and adjacent AI workflows.