grill-skill
Build and harden a skill with evals — interview to design its eval tasks, then run, measure, and iterate. Use when the user wants to create or improve a skill's eval, or run the create → test → improve loop for a skill.
Rank
62
Safety
84
Downloads
1.2k
Updated
Oct 11, 2026
Version
1.0.11
Source
CLAWHUB
About
What it does, and when to use it.
Capability contract not published. No trust telemetry is available yet. 1.2K downloads reported by the source. Last updated 10/11/2026.
Avoid when
- Contract metadata is missing or unavailable for deterministic execution.
Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing
Public facts
Every fact links back to the source it came from.
- Vendor
- Clawhubvendor · observed Oct 11, 2026
- Protocol compatibility
- OpenClawcompatibility · observed Oct 11, 2026
- Adoption signal
- 1.2K downloadsadoption · observed Oct 11, 2026
- Latest release
- 1.0.11release · observed Sep 25, 2026
- Handshake status
- UNKNOWNsecurity
Install and run
Setup complexity: low.
clawhub skill install s17fhassds1tss1zjq6jkjcfd983rjwt:grill-skill- Install using `clawhub skill install s17fhassds1tss1zjq6jkjcfd983rjwt:grill-skill` in an isolated environment before connecting it to live workloads.
- No published capability contract is available yet, so validate auth and request/response behavior manually.
- Review the upstream CLAWHUB listing at https://clawhub.ai/edonadei/grill-skill before using production credentials.
Contract: missing
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-edonadei-grill-skill/snapshot"
Documentation
CLAWHUB
147,506 characters of source documentation, loaded on request.
Extracted files
5 files captured from the source.
SKILL.md
--- name: grill-skill description: Build and harden a skill with evals — interview to design its eval tasks, then run, measure, and iterate. Use when the user wants to create or improve a skill's eval, or run the create → test → improve loop for a skill. allowed-tools: Bash, Read, Write, Edit --- # Grill Skill Interview the user to design a skill's eval, then loop run → measure → improve until it ships. Requires `caliper` (`pipx install caliper-eval` if missing). Commands, spec skeleton, and expect/assert guidance: [REFERENCE.md](REFERENCE.md). ## Entry point `/grill-skill [path]` — optional path to a `SKILL.md`. - **Path given** — use it. - **No path** — look for `SKILL.md` in the cwd; if found, confirm before proceeding, else ask where it is. ## Phase 1 — Understand Read the `SKILL.md`. Summarize what it does, when it triggers, and what a successful run looks like. Ask the user to confirm your reading. **Wait for confirmation before continuing.** ## Phase 2 — Detect eval mode Look for `*.eval.yaml` beside the `SKILL.md` (try `<dir-name>.eval.yaml` first). - **None** → New eval. **Found** → Gap-fill. Interview one question at a time and wait for each answer. Never invent the user's answers or write the spec before interviewing. ### New eval — three tasks Elicit three tasks, one question at a time: 1. **Happy path** — the most common successful use. What did the agent do, and what would confirm it worked? 2. **Edge case** — a tricky-but-valid input that might trip the raw agent. 3. **Adversarial** — what the skill should refuse or avoid. Turn each answer into a task: a realistic `prompt`, an observable `expect`, and an `assert` when the outcome is checkable (see [REFERENCE.md](REFERENCE.md)). Show the proposed YAML and confirm before writing. Write the spec beside `SKILL.md`, named `<dir-name>.eval.yaml`, with `skill.path: ./SKILL.md` and `claude-code` as the default backend for both `skill` and `judge` unless the SKILL.md targets another. ### Gap-fill Read the existing spec and report its tasks. **Ask what behaviors are missing or under-tested before proposing or writing anything** — even if the user only asked you to inspect it, report first, then ask. Sharpen each gap into a task, show it, and confirm before writing it in. ## Whose setup is measured Runs load the user's own customizations by default (user skills, plugins, rules, settings and connectors; see REFERENCE.md for backend exceptions), which answers "does my skill work in *my* agent?". **Isolate** (`--no-user-customizations`, or `user_customizations: false` in the spec) when comparing backends or models, when the number leaves this machine (shared, published, compared with someone else's run), or when measuring the bare agent: each setup is different, so otherwise part of the delta is the setups. `--ablate` of the user's own skill needs no isolation, since both runs load the same setup. **Always tell the user which mode ran** and what it loaded, from the report he
_meta.json
{
"ownerId": "kn7bp63rpwq0rm2g7m4k0c3hjn80qkhc",
"slug": "grill-skill",
"version": "1.0.11",
"publishedAt": 1790347950116
}REFERENCE.md
# Grill Skill Reference ## Caliper commands used by this skill ```bash # Check spec is valid before running (unknown task or sandbox keys, bad forbidden_files # regexes, missing assert: files) caliper validate path/to/spec.eval.yaml # First run — fast, catches spec errors caliper run path/to/spec.eval.yaml --k 1 # Reliability run — after iterating on the skill caliper run path/to/spec.eval.yaml --k 3 # Ablated run — before committing, proves the skill makes a difference. # Run once and keep it: it cannot move when the skill's text changes. # A declared mcp: server can be ablated the same way; qualify as skill:/mcp: # if both declare the name. caliper run path/to/spec.eval.yaml --k 3 --ablate my-skill # Then diff it against the full run. A bare spec name resolves to that spec's # LATEST run, so address the older side by its saved results path. caliper compare .caliper/results/<spec>/<ablated-run>.json <spec> # Choose the engine at run time — it is not stored in the spec (default: claude-code) caliper run path/to/spec.eval.yaml --model codex:gpt-5-codex caliper run path/to/spec.eval.yaml --model codex caliper run path/to/spec.eval.yaml --judge-model claude-code:claude-haiku-4-5-20251001 # Runs load your user customizations (skills, plugins, rules, settings and connectors) by default; isolate for a # portable score (or pin user_customizations: false in the spec) caliper run path/to/spec.eval.yaml --no-user-customizations # Browse past results caliper list caliper report path/to/spec.eval.yaml # Compare two saved runs of the same eval (ablation: full vs. shortened, or over time) caliper compare full-eval short-eval # spec name -> latest run, or a results-JSON path caliper compare a.json b.json --format json # per-task Δ, regression flags, for scripting ``` `caliper compare <A> <B>` diffs two already-saved runs task by task: tasks are matched by name, `Δ = b − a`, a negative Δ flags a regression (any-below), and a side with no usable attempts shows `—` (unmeasured, never a regression) so infra/judge noise can't fake a loss. Under the success-rate headline, `compare` also shows **token and wall-clock deltas** (green = cheaper) — the "same quality, 40% fewer tokens" signal an ablation looks for. These are secondary: a token/time change is **never** a regression (only the score is), and dollar cost is not tracked (tokens are the volume signal). Each attempt in the report also shows its tokens next to its duration under `--verbose`. `compare` also reports **skill drift** — a member of the neighbourhood whose *text* changed between the two runs, read from the per-file hashes in each run's snapshots. It is graded by provenance, not role: a drifted **git source** warns, because the spec claimed where those bytes came from and the delta you are reading is confounded; a drifted **path source** is shown without alarm, because nothing was promised about a working file and that edit is usually the thing the run exists to measure. ``` ⚠
skill-card.md
## Description: Interviews skill authors to design eval tasks, then uses Caliper results to test and improve their skills. This skill is ready for commercial/non-commercial use. ## Publisher: [edonadei](https://clawhub.ai/user/edonadei) ### License/Terms of Use: MIT-0 ## Use Case: Skill authors and developers use this skill to design evaluation tasks, run Caliper tests, and identify changes that improve skill reliability. ### Deployment Geography for Use: Global ## Known Risks and Mitigations: Risk: The skill can read and edit selected skill files and write evaluation specifications. Mitigation: Use it in a trusted project and review generated evaluation YAML before confirming writes. Risk: Caliper tests may execute task setup or cleanup commands and load local agent customizations. Mitigation: Review test commands and use isolated Caliper mode for portable or shared results. ## Reference(s): - [Grill Skill on ClawHub](https://clawhub.ai/edonadei/skills/grill-skill) - [Grill Skill Reference](artifact/REFERENCE.md) ## Skill Output: **Output Type(s):** [Guidance, Configuration, Shell commands] **Output Format:** [Markdown guidance and YAML evaluation specifications] **Output Parameters:** [1D] **Other Properties Related to Output:** [Includes test results and comparisons when Caliper is run.] ## Skill Version(s): 1.0.11 (source: ClawHub release metadata) ## Ethical Considerations: Users should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.
grill-skill.eval.yaml
skills:
- ./SKILL.md
# The engine (backend + model) is a runtime axis, not a spec field — this eval
# runs on whatever `--model` / `--judge-model` select (default claude-code).
# No explicit sandbox.forbidden_files: caliper already auto-forbids the eval
# spec and any .caliper/results/ directory. Listing "./.caliper/.*" here caused a
# false cheat flag, because these grill-skill tasks legitimately WRITE caliper
# specs whose text contains that very pattern.
tasks:
# Task 1 — First-turn interview discipline (open prompt, no eval present).
# Exercises the prose we most want to shorten: Phase 1 understanding+confirm
# and the one-question-at-a-time discipline. Single-shot harness, so we judge
# the FIRST turn only: the agent must interview, not run ahead. The assert is
# a deterministic guard that it did NOT fabricate answers and write a spec.
- name: Interviews before generating — asks, then stops (does not run ahead)
activates: [grill-skill]
setup: |
rm -rf /tmp/grill-fresh
mkdir -p /tmp/grill-fresh
cat > /tmp/grill-fresh/SKILL.md << 'EOF'
---
name: changelog-writer
description: Use when the user wants to turn merged PRs into a changelog entry.
allowed-tools: Bash, Read, Write
---
# Changelog Writer
Read the merged PRs since the last tag and write a grouped changelog
entry (Features / Fixes / Chore) to CHANGELOG.md.
EOF
cleanup: rm -rf /tmp/grill-fresh
prompt: >
I want to create a caliper eval for my skill at
/tmp/grill-fresh/SKILL.md. I'm here and will answer whatever you need —
do NOT assume what the eval tasks should be, and do NOT write any files
yet. What do you need to know from me to get started?
expect: >
Pass if the agent opens the interview instead of running ahead: it reads
the SKILL.md, gives some understanding of the skill, and asks the user for
input before generating anything — then STOPS to wait. The number of
questions does not matter. Fail if the agent skips the interview: it
invents the user's answers, generates the eval tasks itself without
asking, or writes any .eval.yaml file in this turn.
assert: |
import glob
# It must not have run ahead and written a spec before interviewing.
specs = glob.glob("/tmp/grill-fresh/*.eval.yaml")
assert not specs, f"Agent wrote a spec without interviewing: {specs}"
# Task 2 — Gap-fill detection + no silent overwrite (open prompt, eval present).
# Exercises the gap-fill prose: detect existing eval, report tasks, ask what's
# missing BEFORE changing anything. Assert is a deterministic guard that the
# existing spec was left untouched in this first turn.
- name: Detects an existing eval and asks before touching it
activates: [grill-skill]
setup: |
rm -rf /tmp/grill-existing
mkdir -p /tmp/grill-existing
cat > /tmp/grill-existing/SKILL.md << 'EOF'
---
nameAionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
activepieces
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
CopilotKit
The Frontend for Agents & Generative UI. React + Angular
Machine-readable data
The same record, as JSON, for agents and crawlers.
{
"facts": [
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Clawhub",
"href": "https://clawhub.ai/edonadei/skills/grill-skill",
"sourceUrl": "https://clawhub.ai/edonadei/skills/grill-skill",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-11T04:54:05.694Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-edonadei-grill-skill/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-edonadei-grill-skill/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-11T04:54:05.694Z",
"isPublic": true
},
{
"factKey": "traction",
"category": "adoption",
"label": "Adoption signal",
"value": "1.2K downloads",
"href": "https://clawhub.ai/edonadei/grill-skill",
"sourceUrl": "https://clawhub.ai/edonadei/grill-skill",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-11T04:54:05.694Z",
"isPublic": true
},
{
"factKey": "latest_release",
"category": "release",
"label": "Latest release",
"value": "1.0.11",
"href": "https://clawhub.ai/edonadei/grill-skill",
"sourceUrl": "https://clawhub.ai/edonadei/grill-skill",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-09-25T14:52:30.116Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-edonadei-grill-skill/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-edonadei-grill-skill/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
],
"events": [
{
"eventType": "release",
"title": "Release 1.0.11",
"description": "### Changed - The interview now covers triggering: it asks which skills yours could be confused with, adds a neighbour probe for each, and always proposes a silence probe. - Generated tasks put `activates:` on every execution task and never name the skill in the prompt, so a run can tell a `description` failure from a body failure. - New Phase 4: run the control (`--ablate <skill>`) before editing the skill. A task that passes without the skill gets sharpened before any iteration. - New Phase 5: each failure is traced to its fix (`description`, body, task, or setup). After each edit, compare against the previous full run and against the kept control. - Warns that the harness is single-shot: for a skill that asks before acting, tasks judge the first turn. - Now focused on the interview (helping you decide what to test). Writing a spec whose tasks you've already decided moved to evaluate-skill. - Reference rewritten around authoring: spec skeleton with probes, task-writing rules, attempt workdir, user customizations. ### Fixed - It told the agent to write a backend into the spec, which `caliper validate` rejects.",
"href": "https://clawhub.ai/edonadei/grill-skill",
"sourceUrl": "https://clawhub.ai/edonadei/grill-skill",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-09-25T14:52:30.116Z",
"isPublic": true
}
]
}Record generated Oct 11, 2026.
