{"id":"aaaf6f4d-972d-4948-bf2c-fe4e673b64ef","entityType":"agent","slug":"clawhub-edonadei-grill-skill","name":"grill-skill","canonicalUrl":"https://www.xpersona.co/agent/clawhub-edonadei-grill-skill","canonicalPath":"/agent/clawhub-edonadei-grill-skill","generatedAt":"2026-10-11T07:40:11.111Z","source":"CLAWHUB","claimStatus":"UNCLAIMED","verificationTier":"NONE","summary":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-11T04:54:05.694Z","emptyReason":null},"description":"Build and harden a skill with evals — interview to design its eval tasks, then run, measure, and iterate. Use when the user wants to create or improve a skill's eval, or run the create → test → improve loop for a skill.","descriptionLabel":"Source description","evidenceSummary":"Capability contract not published. No trust telemetry is available yet. 1.2K downloads reported by the source. Last updated 10/11/2026.","installCommand":"clawhub skill install s17fhassds1tss1zjq6jkjcfd983rjwt:grill-skill","sourceUrl":"https://clawhub.ai/edonadei/grill-skill","homepage":"https://clawhub.ai/edonadei/skills/grill-skill","primaryLinks":[{"label":"View on ClawHub","url":"https://clawhub.ai/edonadei/grill-skill","kind":"source"},{"label":"Homepage","url":"https://clawhub.ai/edonadei/skills/grill-skill","kind":"homepage"}],"safetyScore":84,"overallRank":62,"popularityScore":61,"trustScore":null,"claimedByName":null,"isOwner":false,"seoDescription":"grill-skill technical dossier on Xpersona with agent coverage, OPENCLEW support, and live trust metadata."},"coverage":{"evidence":{"source":"public-profile","verified":false,"confidence":"medium","updatedAt":"2026-10-11T04:54:05.694Z","emptyReason":null},"protocols":[{"protocol":"OPENCLEW","label":"OpenClaw","status":"self-declared","notes":"Declared in the public agent profile."}],"capabilities":[],"verifiedCount":0,"selfDeclaredCount":1,"capabilityMatrix":{"rows":[{"key":"OPENCLEW","type":"protocol","support":"unknown","confidenceSource":"profile","notes":"Listed on profile"}],"flattenedTokens":"protocol:OPENCLEW|unknown|profile"}},"adoption":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-11T04:54:05.694Z","emptyReason":null},"stars":null,"forks":null,"downloads":1154,"packageName":null,"latestVersion":"1.0.11","tractionLabel":"1.2K downloads"},"release":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-11T04:54:05.622Z","emptyReason":null},"lastUpdatedAt":"2026-10-11T04:54:05.694Z","lastCrawledAt":"2026-10-11T04:54:05.622Z","lastIndexedAt":null,"nextCrawlAt":"2026-10-12T04:54:05.622Z","lastVerifiedAt":null,"highlights":[{"version":"1.0.11","createdAt":"2026-09-25T14:52:30.116Z","changelog":"### Changed - The interview now covers triggering: it asks which skills yours could be confused with, adds a neighbour probe for each, and always proposes a silence probe. - Generated tasks put `activates:` on every execution task and never name the skill in the prompt, so a run can tell a `description` failure from a body failure. - New Phase 4: run the control (`--ablate <skill>`) before editing the skill. A task that passes without the skill gets sharpened before any iteration. - New Phase 5: each failure is traced to its fix (`description`, body, task, or setup). After each edit, compare against the previous full run and against the kept control. - Warns that the harness is single-shot: for a skill that asks before acting, tasks judge the first turn. - Now focused on the interview (helping you decide what to test). Writing a spec whose tasks you've already decided moved to evaluate-skill. - Reference rewritten around authoring: spec skeleton with probes, task-writing rules, attempt workdir, user customizations. ### Fixed - It told the agent to write a backend into the spec, which `caliper validate` rejects.","fileCount":5,"zipByteSize":14238},{"version":"1.0.10","createdAt":"2026-08-29T06:35:55.819Z","changelog":"Improvements on the Caliper underlying CLI focused on reliability and usability (performance and retries)","fileCount":5,"zipByteSize":11818},{"version":"1.0.9","createdAt":"2026-07-12T17:44:28.913Z","changelog":"- MCP servers (mcp: spec block) — added the mcp: block to the spec skeleton plus a new \"MCP servers\" section covering stdio vs remote servers, backend support, namespaced tool names, and secret handling. - Persisted transcripts — new \"Results storage\" section documenting the optional transcript field on saved attempts (inspect which MCP tools fired after a run); backward-compatible with older JSON.","fileCount":5,"zipByteSize":9611},{"version":"1.0.8","createdAt":"2026-07-05T20:21:48.137Z","changelog":"- Added token and wall-time tracking to evaluations - Moved from pass@k to raw success rate, easier to reason about for agents","fileCount":5,"zipByteSize":8432},{"version":"1.0.7","createdAt":"2026-07-03T20:03:04.820Z","changelog":"- Caliper now supports natively Hermes agent","fileCount":5,"zipByteSize":8358},{"version":"1.0.6","createdAt":"2026-07-03T16:41:29.262Z","changelog":"- Breaking change, now the spec of model - harness does not live in spec anymore, it's decided when executing the API","fileCount":5,"zipByteSize":8153},{"version":"1.0.5","createdAt":"2026-07-03T04:15:34.738Z","changelog":"- Added Caliper benchmark result files to track performance and test runs. - Removed the sample skill-card documentation file (skill-card.md). - No functional changes to core logic described in SKILL.md. - Update focused on improving evaluation logging and managing result outputs.","fileCount":5,"zipByteSize":7895},{"version":"1.0.4","createdAt":"2026-07-02T19:30:59.098Z","changelog":"- Major update: Improved workflow clarity and enforced strict user-guided eval creation. - Added multiple `.caliper/results/grill-skill/*.json` sample result files. - Removed `skill-card.md` file. - Streamlined and clarified interactive flow: always confirm user intent before proceeding, never invent answers, and always report existing eval coverage before suggesting gaps. - Documented precise prompts to elicit \"happy path,\" \"edge case,\" and \"adversarial\" eval tasks, ensuring interview-driven creation. - Updated instructions for naming spec files and selecting default backends.","fileCount":5,"zipByteSize":7538}]},"execution":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No published capability contract is available yet."},"installCommand":"clawhub skill install s17fhassds1tss1zjq6jkjcfd983rjwt:grill-skill","setupComplexity":"low","setupSteps":["Install using `clawhub skill install s17fhassds1tss1zjq6jkjcfd983rjwt:grill-skill` in an isolated environment before connecting it to live workloads.","No published capability contract is available yet, so validate auth and request/response behavior manually.","Review the upstream CLAWHUB listing at https://clawhub.ai/edonadei/grill-skill before using production credentials."],"contract":{"contractStatus":"missing","authModes":[],"requires":[],"forbidden":[],"supportsMcp":false,"supportsA2a":false,"supportsStreaming":false,"inputSchemaRef":null,"outputSchemaRef":null,"dataRegion":null,"contractUpdatedAt":null,"sourceUpdatedAt":null,"freshnessSeconds":null},"invocationGuide":{"preferredApi":{"snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-edonadei-grill-skill/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-edonadei-grill-skill/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-edonadei-grill-skill/trust"},"curlExamples":["curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-edonadei-grill-skill/snapshot\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-edonadei-grill-skill/contract\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-edonadei-grill-skill/trust\""],"jsonRequestTemplate":{"query":"summarize this repo","constraints":{"maxLatencyMs":2000,"protocolPreference":["OPENCLEW"]}},"jsonResponseTemplate":{"ok":true,"result":{"summary":"...","confidence":0.9},"meta":{"source":"CLAWHUB","generatedAt":"2026-10-11T07:40:11.107Z"}},"retryPolicy":{"maxAttempts":3,"backoffMs":[500,1500,3500],"retryableConditions":["HTTP_429","HTTP_503","NETWORK_TIMEOUT"]}},"endpoints":{"dossierUrl":"https://www.xpersona.co/api/v1/agents/clawhub-edonadei-grill-skill/dossier","snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-edonadei-grill-skill/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-edonadei-grill-skill/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-edonadei-grill-skill/trust"}},"reliability":{"evidence":{"source":"runtime-metrics","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No trust, reliability, or runtime telemetry is available."},"trust":{"status":"unavailable","handshakeStatus":"UNKNOWN","verificationFreshnessHours":null,"reputationScore":null,"p95LatencyMs":null,"successRate30d":null,"fallbackRate":null,"attempts30d":null,"trustUpdatedAt":null,"trustConfidence":"unknown","sourceUpdatedAt":null,"freshnessSeconds":null},"decisionGuardrails":{"doNotUseIf":["Contract metadata is missing or unavailable for deterministic execution."],"safeUseWhen":[],"riskFlags":["missing_or_unavailable_contract","trust_data_unavailable","schema_references_missing"],"operationalConfidence":"low"},"executionMetrics":{"observedLatencyMsP50":null,"observedLatencyMsP95":null,"estimatedCostUsd":null,"uptime30d":null,"rateLimitRpm":null,"rateLimitBurst":null,"lastVerifiedAt":null,"verificationSource":null},"runtimeMetrics":{"successRate":null,"avgLatencyMs":null,"avgCostUsd":null,"hallucinationRate":null,"retryRate":null,"disputeRate":null,"p50Latency":null,"p95Latency":null,"lastUpdated":null}},"benchmarks":{"evidence":{"source":"no-benchmark-data","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No benchmark suites or observed failure patterns are available."},"suites":[],"failurePatterns":[]},"artifacts":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-11T04:54:05.694Z","emptyReason":null},"readme":"Skill: grill-skill\n\nOwner: edonadei\n\nSummary: Build and harden a skill with evals — interview to design its eval tasks, then run, measure, and iterate. Use when the user wants to create or improve a skill's eval, or run the create → test → improve loop for a skill.\n\nTags: latest:1.0.11\n\nVersion history:\n\nv1.0.11 | 2026-09-25T14:52:30.116Z | user\n\n### Changed\n- The interview now covers triggering: it asks which skills yours could be confused with, adds a neighbour probe for each, and always proposes a silence probe.\n- Generated tasks put `activates:` on every execution task and never name the skill in the prompt, so a run can tell a `description` failure from a body failure.\n- New Phase 4: run the control (`--ablate <skill>`) before editing the skill. A task that passes without the skill gets sharpened before any iteration.\n- New Phase 5: each failure is traced to its fix (`description`, body, task, or setup). After each edit, compare against the previous full run and against the kept control.\n- Warns that the harness is single-shot: for a skill that asks before acting, tasks judge the first turn.\n- Now focused on the interview (helping you decide what to test). Writing a spec whose tasks you've already decided moved to evaluate-skill.\n- Reference rewritten around authoring: spec skeleton with probes, task-writing rules, attempt workdir, user customizations.\n\n### Fixed\n- It told the agent to write a backend into the spec, which `caliper validate` rejects.\n\nv1.0.10 | 2026-08-29T06:35:55.819Z | user\n\nImprovements on the Caliper underlying CLI focused on reliability and usability (performance and retries)\n\nv1.0.9 | 2026-07-12T17:44:28.913Z | user\n\n- MCP servers (mcp: spec block) — added the mcp: block to the spec skeleton plus a new \"MCP servers\" section covering stdio vs remote servers, backend support, namespaced tool names, and secret handling.\n- Persisted transcripts — new \"Results storage\" section documenting the optional transcript field on saved attempts (inspect which MCP tools fired after a run); backward-compatible with older JSON.\n\nv1.0.8 | 2026-07-05T20:21:48.137Z | user\n\n- Added token and wall-time tracking to evaluations\n- Moved from pass@k to raw success rate, easier to reason about for agents\n\nv1.0.7 | 2026-07-03T20:03:04.820Z | user\n\n- Caliper now supports natively Hermes agent\n\nv1.0.6 | 2026-07-03T16:41:29.262Z | user\n\n- Breaking change, now the spec of model - harness does not live in spec anymore, it's decided when executing the API\n\nv1.0.5 | 2026-07-03T04:15:34.738Z | user\n\n- Added Caliper benchmark result files to track performance and test runs.\n- Removed the sample skill-card documentation file (skill-card.md).\n- No functional changes to core logic described in SKILL.md. \n- Update focused on improving evaluation logging and managing result outputs.\n\nv1.0.4 | 2026-07-02T19:30:59.098Z | user\n\n- Major update: Improved workflow clarity and enforced strict user-guided eval creation.\n- Added multiple `.caliper/results/grill-skill/*.json` sample result files.\n- Removed `skill-card.md` file.\n- Streamlined and clarified interactive flow: always confirm user intent before proceeding, never invent answers, and always report existing eval coverage before suggesting gaps.\n- Documented precise prompts to elicit \"happy path,\" \"edge case,\" and \"adversarial\" eval tasks, ensuring interview-driven creation.\n- Updated instructions for naming spec files and selecting default backends.\n\nv1.0.3 | 2026-06-26T03:26:14.124Z | user\n\n- Now supports Pi harness\n- Removed the skill-card.md file.\n- No changes to the main workflow or user interaction.\n- Skill documentation (SKILL.md) remains unchanged.\n\nv1.0.2 | 2026-06-24T00:52:21.219Z | user\n\n- Added new caliper results files to track evaluations.\n- Removed the skill-card.md file.\n- No changes to SKILL.md content in this version.\n- Maintenance update focused on result storage and cleanup of unused files.\n\nv1.0.1 | 2026-06-20T03:01:35.362Z | user\n\n- Clarified gap-fill workflow to always ask about missing or under-tested behaviors before proposing new eval tasks.\n- Added guidance to only report and not modify eval files when the user requests inspection or reporting, but still prompt for missing coverage.\n- No code changes; documentation and workflow clarification only.\n\nv1.0.0 | 2026-06-19T23:47:50.879Z | auto\n\n- Initial release of grill-skill: a guided workflow for creating and iterating on caliper evals for a skill.\n- Interviews users to generate three key eval tasks (happy path, edge case, adversarial) and writes a structured eval spec.\n- Supports gap-filling: identifies missing test cases in existing eval specs and helps the user add new ones.\n- Automates eval validation and initial runs, displaying results and diagnosing errors.\n- Interactive iteration loop: encourages users to improve their skill and rerun tests, with baseline comparison guidance before commit.\n\nArchive index:\n\nArchive v1.0.11: 5 files, 14238 bytes\n\nFiles: grill-skill.eval.yaml (8032b), REFERENCE.md (18262b), skill-card.md (1666b), SKILL.md (3880b), _meta.json (131b)\n\nFile v1.0.11:SKILL.md\n\n---\nname: grill-skill\ndescription: Build and harden a skill with evals — interview to design its eval tasks, then run, measure, and iterate. Use when the user wants to create or improve a skill's eval, or run the create → test → improve loop for a skill.\nallowed-tools: Bash, Read, Write, Edit\n---\n\n# Grill Skill\n\nInterview the user to design a skill's eval, then loop run → measure → improve until it ships. Requires `caliper` (`pipx install caliper-eval` if missing). Commands, spec skeleton, and expect/assert guidance: [REFERENCE.md](REFERENCE.md).\n\n## Entry point\n\n`/grill-skill [path]` — optional path to a `SKILL.md`.\n\n- **Path given** — use it.\n- **No path** — look for `SKILL.md` in the cwd; if found, confirm before proceeding, else ask where it is.\n\n## Phase 1 — Understand\n\nRead the `SKILL.md`. Summarize what it does, when it triggers, and what a successful run looks like. Ask the user to confirm your reading. **Wait for confirmation before continuing.**\n\n## Phase 2 — Detect eval mode\n\nLook for `*.eval.yaml` beside the `SKILL.md` (try `<dir-name>.eval.yaml` first).\n\n- **None** → New eval. **Found** → Gap-fill.\n\nInterview one question at a time and wait for each answer. Never invent the user's answers or write the spec before interviewing.\n\n### New eval — three tasks\n\nElicit three tasks, one question at a time:\n\n1. **Happy path** — the most common successful use. What did the agent do, and what would confirm it worked?\n2. **Edge case** — a tricky-but-valid input that might trip the raw agent.\n3. **Adversarial** — what the skill should refuse or avoid.\n\nTurn each answer into a task: a realistic `prompt`, an observable `expect`, and an `assert` when the outcome is checkable (see [REFERENCE.md](REFERENCE.md)). Show the proposed YAML and confirm before writing.\n\nWrite the spec beside `SKILL.md`, named `<dir-name>.eval.yaml`, with `skill.path: ./SKILL.md` and `claude-code` as the default backend for both `skill` and `judge` unless the SKILL.md targets another.\n\n### Gap-fill\n\nRead the existing spec and report its tasks. **Ask what behaviors are missing or under-tested before proposing or writing anything** — even if the user only asked you to inspect it, report first, then ask. Sharpen each gap into a task, show it, and confirm before writing it in.\n\n## Whose setup is measured\n\nRuns load the user's own customizations by default (user skills, plugins, rules, settings and connectors; see REFERENCE.md for backend exceptions), which answers \"does my skill work in *my* agent?\". **Isolate** (`--no-user-customizations`, or `user_customizations: false` in the spec) when comparing backends or models, when the number leaves this machine (shared, published, compared with someone else's run), or when measuring the bare agent: each setup is different, so otherwise part of the delta is the setups. `--ablate` of the user's own skill needs no isolation, since both runs load the same setup.\n\n**Always tell the user which mode ran** and what it loaded, from the report header's `user customizations:` line (absent means isolated), and relay any fix `caliper compare` suggests about it.\n\n## Phase 3 — First run\n\nValidate the spec, then run at `k=1` (commands in [REFERENCE.md](REFERENCE.md)). Show the results. Fix any harness or config error (not a task failure) before asking the user what to do next.\n\n## Phase 4 — Iterate\n\nAsk whether to iterate or finish.\n\n- **Iterate** — after the user edits their `SKILL.md`, re-run at `k=3` and show results. Loop back.\n- **Done** — suggest an `--ablate <skill-name>` run plus a `caliper compare` to prove the skill beats the raw agent, then remind the user to commit `SKILL.md` and the `.eval.yaml` together. Mention that the ablated run is worth keeping: it cannot move when the skill's text changes, so later iterations re-diff against it instead of re-running it.\n\nFile v1.0.11:_meta.json\n\n{\n  \"ownerId\": \"kn7bp63rpwq0rm2g7m4k0c3hjn80qkhc\",\n  \"slug\": \"grill-skill\",\n  \"version\": \"1.0.11\",\n  \"publishedAt\": 1790347950116\n}\n\nFile v1.0.11:REFERENCE.md\n\n# Grill Skill Reference\n\n## Caliper commands used by this skill\n\n```bash\n# Check spec is valid before running (unknown task or sandbox keys, bad forbidden_files\n# regexes, missing assert: files)\ncaliper validate path/to/spec.eval.yaml\n\n# First run — fast, catches spec errors\ncaliper run path/to/spec.eval.yaml --k 1\n\n# Reliability run — after iterating on the skill\ncaliper run path/to/spec.eval.yaml --k 3\n\n# Ablated run — before committing, proves the skill makes a difference.\n# Run once and keep it: it cannot move when the skill's text changes.\n# A declared mcp: server can be ablated the same way; qualify as skill:/mcp:\n# if both declare the name.\ncaliper run path/to/spec.eval.yaml --k 3 --ablate my-skill\n# Then diff it against the full run. A bare spec name resolves to that spec's\n# LATEST run, so address the older side by its saved results path.\ncaliper compare .caliper/results/<spec>/<ablated-run>.json <spec>\n\n# Choose the engine at run time — it is not stored in the spec (default: claude-code)\ncaliper run path/to/spec.eval.yaml --model codex:gpt-5-codex\ncaliper run path/to/spec.eval.yaml --model codex\ncaliper run path/to/spec.eval.yaml --judge-model claude-code:claude-haiku-4-5-20251001\n\n# Runs load your user customizations (skills, plugins, rules, settings and connectors) by default; isolate for a\n# portable score (or pin user_customizations: false in the spec)\ncaliper run path/to/spec.eval.yaml --no-user-customizations\n\n# Browse past results\ncaliper list\ncaliper report path/to/spec.eval.yaml\n\n# Compare two saved runs of the same eval (ablation: full vs. shortened, or over time)\ncaliper compare full-eval short-eval           # spec name -> latest run, or a results-JSON path\ncaliper compare a.json b.json --format json     # per-task Δ, regression flags, for scripting\n```\n\n`caliper compare <A> <B>` diffs two already-saved runs task by task: tasks are\nmatched by name, `Δ = b − a`, a negative Δ flags a regression (any-below), and a\nside with no usable attempts shows `—` (unmeasured, never a regression) so\ninfra/judge noise can't fake a loss. Under the success-rate headline, `compare` also\nshows **token and wall-clock deltas** (green = cheaper) — the \"same quality, 40%\nfewer tokens\" signal an ablation looks for. These are secondary: a token/time\nchange is **never** a regression (only the score is), and dollar cost is not tracked\n(tokens are the volume signal). Each attempt in the report also shows its tokens\nnext to its duration under `--verbose`.\n\n`compare` also reports **skill drift** — a member of the neighbourhood whose\n*text* changed between the two runs, read from the per-file hashes in each run's\nsnapshots. It is graded by provenance, not role: a drifted **git source** warns,\nbecause the spec claimed where those bytes came from and the delta you are\nreading is confounded; a drifted **path source** is shown without alarm, because\nnothing was promised about a working file and that edit is usually the thing the\nrun exists to measure.\n\n```\n ⚠ tdd changed between runs — git source, a1b2c3d → e4f5g6h; pin `ref:` to hold it fixed\n   my-skill changed between runs — path, 4fc7951 → bcbcbde\n```\n\nThis is a change in *text* at constant membership; a change in *membership* is\nthe separate neighbourhood warning.\n\n## Inspecting failures\n\nAfter any `caliper run`, failed tasks are shown automatically with their output\nand `assert_evidence` — no extra command needed. Each attempt is tagged with an\n`outcome`: a real `task_fail` reads as `✗`, while *unusable* attempts\n(`infra_error` from a rate limit that outlasted its retries or an attempt with\nno model call observed, `timeout`, or\n`judge_error`)\nread as `⊘` and are excluded from the score denominator, with a separate\n\"N unusable\" count in the summary — so a throttled or judge-flaked run is not\nmistaken for a skill regression. A run where *every* attempt was unusable\nmeasured nothing: it is still saved, but `caliper run` exits `2` and prints the\ncount of each outcome. If `caliper run --fail-fast N` stopped a task\nafter repeated `infra_error` / `timeout` outcomes, the report marks it as\n`ABORTED` and shows how many attempts ran. A run you stopped with Ctrl-C is\nsaved too, headed by an `interrupted:` line: its rates are computed over the\nattempts that ran, so read them as a smaller sample rather than a worse skill —\nand re-run before drawing a conclusion from a handful of attempts. `caliper list`\nmarks such a run with `⊘`, and `caliper compare` warns when either side is one,\nso a shallow sample cannot quietly masquerade as a delta. A run that hit a\n**spending cap** stops the same way, with the cap named as the cause — top up and\nre-run rather than reading its numbers. A `throttled:` line under the usage\nsummary means attempts were retried before they landed: the scores are sound, but\nthe wall times were fought for. If a failure is still unclear, use\n`--verbose` to see full output for all tasks (including passing ones):\n\n```bash\n# Full output for all tasks (passing + failing), untruncated\ncaliper report path/to/spec.eval.yaml --verbose\n\n# Or inspect a specific past run\ncaliper report path/to/spec.eval.yaml --run 2026-06-21T14-53-12Z --verbose\n```\n\n## Spec skeleton\n\nThe spec carries no engine — pick the backend/model at run time with `--model` /\n`--judge-model` (default `claude-code`).\n\n```yaml\nskills:                   # installed at the agent's own skills root, never\n  - ./SKILL.md            #   preloaded — the agent has to choose it\n  # add further entries to test that yours is the one that fires (they are\n  # assertable via `activates:`, not decoration). A bare string is a *path\n  # source*; a mapping is a *git source* caliper clones for you:\n  - repo: vercel-labs/agent-skills\n    ref: a1b2c3d          # optional — omit to track the default branch\n    path: skills/tdd/SKILL.md   # optional — defaults to SKILL.md at the root\n                          # a symlink out of the cloned repo refuses the run\n\nsandbox:\n  forbidden_files:               # extra patterns only — the spec itself and any\n    - \"./answers/.*\"             #   .caliper/ directory are forbidden already\n\n# Optional — only if the skill needs MCP tools. claude-code, hermes, codex backends.\nmcp:\n  weather:                       # local stdio server → mcp__weather__<tool>\n    command: python3\n    args: [./servers/weather.py]\n    env:\n      API_TOKEN: ${MCP_API_TOKEN}   # resolved from your shell at run time\n  gdrive:                        # remote (hosted) server over HTTP/SSE\n    type: http                   # http or sse\n    url: https://mcp.example.com/gdrive\n    headers:\n      Authorization: Bearer ${GDRIVE_TOKEN}   # resolved from your shell at run time\n\ntasks:\n  - name: Happy path — <what success looks like>\n    setup: <optional shell command>\n    cleanup: <optional shell command>\n    prompt: <prompt sent to the agent>\n    expect: <natural-language success criterion>\n    assert: |\n      # optional deterministic check\n\n  - name: Edge case — <tricky but valid input>\n    prompt: ...\n    expect: ...\n\n  - name: Adversarial — <what the skill should refuse or avoid>\n    prompt: ...\n    expect: <describes the refusal or safe behavior>\n\n  - name: Silence — <work no declared skill should answer>\n    prompt: ...\n    activates: []                # a trigger probe: no judge, no execution score\n```\n\nEach task needs at least one of `expect`, `assert` or `activates`.\n\nEach attempt runs in a fresh, empty **attempt workdir**: `setup:`, the agent,\n`assert:` and `cleanup:` all run there, so a relative path means the same file to\neach. It is not the spec's directory and not a git repo — build what the task\nneeds in `setup:` (`cp -R \"$CALIPER_SPEC_DIR/fixture/.\" .`, `git init`). Hooks and\nassertions get `CALIPER_WORKDIR` and `CALIPER_SPEC_DIR`; `assert: ./check.py`\nstill resolves against the spec's directory.\n\nLifecycle hooks run for each attempt. A failed `setup:` skips the agent and judge\nand records an `infra_error`; `cleanup:` is still attempted. A failed cleanup\ndoes not change a completed attempt's outcome, but the command exits `2`.\n`AttemptRecord.hook_failures` and `RunMeta.hook_failures` save the task ID,\nattempt number, phase, exit code, and output. The run-level list includes\ncleanup failures on interrupted attempts with no attempt record.\n\n## Triggering: does the description fire?\n\nSkills are **installed** where the agent looks for them and never pasted into\nthe prompt, so whether the agent reaches for one is measurable. Two rules follow\nfor how you write prompts:\n\n- **Never name the skill in a prompt.** \"Use the commit-message skill to…\"\n  removes the very choice being measured. Write the prompt a real user would.\n- **`activates:` asserts the exact set** of skills that loaded — `[a]` means `a`\n  and nothing else, `[]` means silence. Names are the frontmatter `name:`, not\n  filenames.\n\nA task carrying only `activates:` is a **trigger probe**. It skips the judge\nentirely, so it costs far less than an execution task, and reports as `trigger\nonly` rather than a zero. Two kinds are worth generating:\n\n- **Neighbour probe** — declare a sibling skill in `skills:`, then give a prompt\n  that belongs to *it* and assert `activates: [sibling]`. This catches a\n  `description` that over-claims.\n- **Silence probe** — unrelated work, `activates: []`.\n\nActivation is scored separately from execution and never blended in, so a\nnear-zero score with a green activation column means the body is wrong, while a\nred activation column means the `description` is.\n\n## Naming convention\n\nThe spec file lives next to the skill and shares its directory name:\n\n```\nskills/my-skill/SKILL.md\nskills/my-skill/my-skill.eval.yaml   ← generated here\n```\n\n## Writing good expect: criteria\n\nBe specific about evidence. Include what the judge should look for and what counts as failure.\n\n```yaml\nexpect: |\n  Pass if the agent identifies the null dereference in user_lookup.py and\n  explains the failing path. Fail if it only gives generic style advice,\n  misses the bug, or claims tests passed without running them.\n```\n\n## When to use assert:\n\nAdd `assert:` when the outcome is a fact that an LLM judge might guess wrong:\n- File exists or contains exact content\n- Command exit code or output\n- Git state (staged, committed, clean)\n- JSON schema or exact value\n- Test suite passes or fails\n\n## MCP servers (`mcp:`)\n\nIf the skill under test needs MCP tools, declare them in a top-level `mcp:` block (a mapping keyed by server name) — a capability granted to the agent-under-test for the eval, part of the run environment like `sandbox:` (a sibling of it and of `skills:`), so they belong in the spec, not on the command line. A server is either **local stdio** (a `command`, optional `args`, optional `env`) or **remote** (`type: http`/`sse`, a `url`, optional `headers` for auth); the two field sets are mutually exclusive. Supported on **`claude-code`** (stdio + remote HTTP/SSE), **`hermes`** (stdio + remote header-auth; not remote OAuth), and **`codex`** (stdio + remote header-auth, translated into `[mcp_servers.*]` tables in the isolated `~/.codex/config.toml`; not remote OAuth). A tool call appears in the transcript as a namespaced name — `mcp__<server>__<tool>` on `claude-code` and `codex`, `mcp_<server>_<tool>` on `hermes` — so an `expect:` criterion can check the skill actually used it; word it around behaviour, not one backend's spelling, if the spec runs under more than one engine. Put secrets in a host env var and reference it as `${VAR}` inside a stdio `env:`, a remote `headers:`, or a remote `url:` — it resolves at the harness boundary from your shell at run time and never lands in the committed spec (an unset var fails the run). Running an `mcp:` spec on a backend that can't honor it is a hard error, not a silent no-op: `pi` has no MCP by design and will not honor `mcp:` natively — expose the capability as a CLI tool the skill drives or a pi extension, or run the eval on `claude-code`/`hermes`/`codex`. By default a run also loads user skills, plugins, rules and settings, alongside the user's MCP servers and account connectors (Gmail, Drive, GitHub, and the like), merged with `mcp:` — the declared server wins a name clash — so a task may rely on a connector the user has. The judge keeps its existing connector isolation, and `pi` has nothing to load. For a portable score (published, compared across machines or backends, or measuring the bare agent), put `user_customizations: false` at the top level of the spec, or run with `--no-user-customizations`; a skill that can't be measured without the user's connectors (a hosted OAuth connector like Drive) can say `user_customizations: true`. The saved run records what was loaded, so `compare` can warn when two runs loaded differently. See \"Whose setup is measured\" in SKILL.md for when to isolate.\n\nStdio `command` and `args` entries starting with `./` or `../` resolve from the spec's directory; bare command names remain unchanged. Caliper checks each surviving stdio server after task setup and before the agent starts, so a missing or dead server stops the run with a configuration error instead of receiving a task score.\n\n## Backends\n\n| Backend | Requires | Notes |\n|---|---|---|\n| `claude-code` | Claude Code CLI | Default for most skills |\n| `codex` | Codex CLI | For Codex-targeted skills |\n| `pi` | pi CLI (authenticated) | For pi / agentskills.io skills; native `--skill` loading |\n| `hermes` | Hermes Agent CLI (authenticated) | Nous Research; normalized to a neutral agent, `hermes:<provider>/<model>` picks the model |\n\nThe skill engine (`--model`) and judge engine (`--judge-model`) are chosen independently at run time. Every backend is a CLI agent; for API billing, configure a CLI with an API key rather than selecting a separate backend. When `--judge-model` is omitted, the default `claude-code` judge pins `claude-sonnet-5` at execution time so it does not inherit a stale model from the installed Claude CLI; `RunMeta.judge_model` stays empty unless you pass `--judge-model` explicitly or the autorater reports what it used. `RunMeta.model` records the model the skill backend reported running, not the one requested; a mismatch prints a warning (a run whose attempts disagree records the most common model), and an unknown `hermes:<model>` stops the run.\n\n`hermes` is a stateful agent (persistent memory + persona), so Caliper strips it to a neutral agent per attempt — isolated `HERMES_HOME`, no `SOUL.md`/`MEMORY.md`, `--ignore-rules`, and only the spec's declared skills installed — and recovers the full trajectory via `hermes sessions export` after the `hermes -z` run.\n\n## Results storage\n\nResults are saved automatically to `.caliper/results/<spec-name>/<timestamp>.json`\nunder the project's **results root** — the nearest `.caliper/` at or above your\nworking directory, bounded by the git repo. `run`, `report`, `compare` and\n`list` all resolve the same root, so a run saved from a spec's own subdirectory\nis findable by `caliper report <spec-name>` from anywhere in the project.\n\nEach attempt records its `outcome`, optional `usage`, and optional `transcript` (ordered turns with `tool_name`/`tool_input`/`tool_output` when present)\nso saved runs remain inspectable after the fact — including which MCP tools fired.\nOlder JSON without `transcript` still loads (`null`). `report` and `compare` do not\nrender the transcript; it is stored for later analysis. A run also records what\n`--ablate` removed (`RunMeta.ablated`; a server as `mcp:<name>`) and the `mcp:`\nservers it ran with (`RunMeta.mcp_servers`), so `compare` can check the marker\nrather than trust it. A run saved before that field existed loads it as `null` —\nunknown, not \"no servers\". Two runs that recorded different servers outside an\nablation pair get the `different MCP servers configured` warning, the tool-side\ntwin of the neighbourhood warning. A run also records `RunMeta.user_customizations`\nand `RunMeta.loaded_user_customizations` (what it loaded from the machine, the\ndefault), kept apart from `mcp_servers`; `compare` warns when two runs\nloaded differently, or compares two backends with user customizations.\n\n## Troubleshooting\n\n**`Judge model ... is unavailable` / `Judge authentication failed` / `Judge rate limited`**\nThe judge CLI reached the provider and the call was refused. Caliper classifies these at the harness boundary (from the CLI's structured output) and suggests passing `--judge-model <backend[:model]>` to pick an available judge engine or model. An unavailable judge model fails every attempt the same way, so it stops the run at the first attempt that reaches the judge (exit `2`) instead of recording `judge_error` on each one; an authentication failure or a rate limit stays a per-attempt `judge_error`. An unavailable `claude-code` skill model (`--model claude-code:<model>`) stops the run the same way, and an unknown backend name in `--model` or `--judge-model` is refused before any attempt runs.\n\n### User-layer coverage and activation\n\nClaude Code loads `~/.claude/skills`, `CLAUDE.md`, `settings.json`, and enabled\nuser-scope plugins (copied with private registry paths). Codex loads\n`~/.codex/skills`, `AGENTS.md` / `AGENTS.override.md`, plugins and settings;\nits top-level model pin is still stripped. Hermes loads `~/.hermes/skills` and\nsettings, while keeping `--ignore-rules` and excluding persona/memory (ADR 0005).\nPi is unchanged. The judge keeps its existing connector isolation.\n\nDeclared skill and MCP names win clashes even when ablated. User skills count\nas other activations: an extra name fails exact `activates:` matching. Require a\ndependency by declaring it in `skills:`. Hooks run within the attempt timeout,\nwithout a separate preflight. Isolation retains authentication/provider settings.\nThere are no per-kind switches.\n\nThe report header and `RunMeta.loaded_user_customizations` use kind-prefixed\nnames (`mcp:`, `skill:`, `plugin:`, `rules:`, `settings:`). `null` is unknown,\nnot a partial inventory. `compare` checks name sets, not contents or versions;\nlegacy unprefixed records conservatively differ. See ADR 0028 and\n`docs/backends.md` for connection-setting exceptions.\n\nFile v1.0.11:skill-card.md\n\n## Description:\n\nInterviews skill authors to design eval tasks, then uses Caliper results to test and improve their skills.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[edonadei](https://clawhub.ai/user/edonadei)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nSkill authors and developers use this skill to design evaluation tasks, run Caliper tests, and identify changes that improve skill reliability.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: The skill can read and edit selected skill files and write evaluation specifications.\n\nMitigation: Use it in a trusted project and review generated evaluation YAML before confirming writes.\n\nRisk: Caliper tests may execute task setup or cleanup commands and load local agent customizations.\n\nMitigation: Review test commands and use isolated Caliper mode for portable or shared results.\n\n## Reference(s):\n\n- [Grill Skill on ClawHub](https://clawhub.ai/edonadei/skills/grill-skill)\n- [Grill Skill Reference](artifact/REFERENCE.md)\n\n## Skill Output:\n\n**Output Type(s):** [Guidance, Configuration, Shell commands]\n\n**Output Format:** [Markdown guidance and YAML evaluation specifications]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [Includes test results and comparisons when Caliper is run.]\n\n## Skill Version(s):\n\n1.0.11 (source: ClawHub release metadata)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nFile v1.0.11:grill-skill.eval.yaml\n\nskills:\n  - ./SKILL.md\n\n# The engine (backend + model) is a runtime axis, not a spec field — this eval\n# runs on whatever `--model` / `--judge-model` select (default claude-code).\n\n# No explicit sandbox.forbidden_files: caliper already auto-forbids the eval\n# spec and any .caliper/results/ directory. Listing \"./.caliper/.*\" here caused a\n# false cheat flag, because these grill-skill tasks legitimately WRITE caliper\n# specs whose text contains that very pattern.\n\ntasks:\n  # Task 1 — First-turn interview discipline (open prompt, no eval present).\n  # Exercises the prose we most want to shorten: Phase 1 understanding+confirm\n  # and the one-question-at-a-time discipline. Single-shot harness, so we judge\n  # the FIRST turn only: the agent must interview, not run ahead. The assert is\n  # a deterministic guard that it did NOT fabricate answers and write a spec.\n  - name: Interviews before generating — asks, then stops (does not run ahead)\n    activates: [grill-skill]\n    setup: |\n      rm -rf /tmp/grill-fresh\n      mkdir -p /tmp/grill-fresh\n      cat > /tmp/grill-fresh/SKILL.md << 'EOF'\n      ---\n      name: changelog-writer\n      description: Use when the user wants to turn merged PRs into a changelog entry.\n      allowed-tools: Bash, Read, Write\n      ---\n\n      # Changelog Writer\n\n      Read the merged PRs since the last tag and write a grouped changelog\n      entry (Features / Fixes / Chore) to CHANGELOG.md.\n      EOF\n    cleanup: rm -rf /tmp/grill-fresh\n    prompt: >\n      I want to create a caliper eval for my skill at\n      /tmp/grill-fresh/SKILL.md. I'm here and will answer whatever you need —\n      do NOT assume what the eval tasks should be, and do NOT write any files\n      yet. What do you need to know from me to get started?\n    expect: >\n      Pass if the agent opens the interview instead of running ahead: it reads\n      the SKILL.md, gives some understanding of the skill, and asks the user for\n      input before generating anything — then STOPS to wait. The number of\n      questions does not matter. Fail if the agent skips the interview: it\n      invents the user's answers, generates the eval tasks itself without\n      asking, or writes any .eval.yaml file in this turn.\n    assert: |\n      import glob\n      # It must not have run ahead and written a spec before interviewing.\n      specs = glob.glob(\"/tmp/grill-fresh/*.eval.yaml\")\n      assert not specs, f\"Agent wrote a spec without interviewing: {specs}\"\n\n  # Task 2 — Gap-fill detection + no silent overwrite (open prompt, eval present).\n  # Exercises the gap-fill prose: detect existing eval, report tasks, ask what's\n  # missing BEFORE changing anything. Assert is a deterministic guard that the\n  # existing spec was left untouched in this first turn.\n  - name: Detects an existing eval and asks before touching it\n    activates: [grill-skill]\n    setup: |\n      rm -rf /tmp/grill-existing\n      mkdir -p /tmp/grill-existing\n      cat > /tmp/grill-existing/SKILL.md << 'EOF'\n      ---\n      name: summarize\n      description: Use when the user wants to summarize a file.\n      allowed-tools: Bash\n      ---\n\n      # Summarize\n\n      Summarize the contents of a file provided by the user.\n      EOF\n      cat > /tmp/grill-existing/summarize.eval.yaml << 'EOF'\n      skills:\n        - ./SKILL.md\n      tasks:\n        - name: Summarizes a short text file\n          prompt: Summarize /tmp/notes.txt\n          expect: The agent produces a summary of the file contents.\n      EOF\n    cleanup: rm -rf /tmp/grill-existing\n    prompt: >\n      I want to improve the eval for my skill at\n      /tmp/grill-existing/SKILL.md. I'm here to answer your\n      questions — do NOT change any files yet. Tell me the current state of my\n      eval and what you need from me.\n    expect: >\n      Pass if the agent detects the existing eval, reports the existing task\n      (\"Summarizes a short text file\"), and asks the user what behaviors are\n      missing or under-tested before proposing or writing any changes. Fail if\n      it ignores the existing eval, silently overwrites or rewrites it, or adds\n      tasks without first asking what is missing.\n    assert: |\n      import yaml\n      # The existing spec must be untouched: still exactly its one original task.\n      with open(\"/tmp/grill-existing/summarize.eval.yaml\") as f:\n          spec = yaml.safe_load(f)\n      tasks = spec.get(\"tasks\", [])\n      assert len(tasks) == 1, f\"Existing spec was modified in this turn: {len(tasks)} tasks\"\n      assert tasks[0][\"name\"] == \"Summarizes a short text file\", \"Original task was altered\"\n\n  # Task 3 — Serialization correctness (structure supplied, full artifact check).\n  # The deterministic anchor: when the user hands over the task structure, the\n  # agent must serialize a valid 3-task spec. This is the back-half of the\n  # workflow and the one task with a full artifact assert.\n  - name: Serializes a valid 3-task spec when the structure is given\n    activates: [grill-skill]\n    setup: |\n      rm -rf /tmp/grill-serialize\n      mkdir -p /tmp/grill-serialize\n      cat > /tmp/grill-serialize/SKILL.md << 'EOF'\n      ---\n      name: hello-file\n      description: Use when the user wants to write a greeting file.\n      allowed-tools: Bash\n      ---\n\n      # Hello File\n\n      Write the text \"hello world\" to /tmp/hello.txt when the user asks.\n      EOF\n    cleanup: rm -rf /tmp/grill-serialize\n    prompt: >\n      Write the file directly now — do NOT ask me any clarifying questions\n      first; I have given you everything you need. Write\n      a COMPLETE, valid caliper eval spec for the\n      skill at /tmp/grill-serialize/SKILL.md, and write it to\n      /tmp/grill-serialize/hello-file.eval.yaml (do not run it — just create the\n      file). The spec must have a top-level skills: list containing\n      ./SKILL.md (and NO backend/model or judge block — the engine is a runtime\n      axis, chosen with --model at run time), plus exactly 3 tasks: a happy path\n      where the prompt asks the agent to write the greeting file and the expect\n      is that /tmp/hello.txt contains \"hello world\"; an edge case where the user\n      specifies a custom greeting; and an adversarial case where the user asks\n      to overwrite a protected system file.\n    expect: >\n      A valid, runnable .eval.yaml is written at\n      /tmp/grill-serialize/hello-file.eval.yaml: it has a top-level skills: list\n      containing ./SKILL.md (no backend/model, no judge block) and exactly 3\n      tasks covering a happy path, an edge case, and an adversarial case.\n    assert: |\n      import yaml\n\n      path = \"/tmp/grill-serialize/hello-file.eval.yaml\"\n      try:\n          with open(path) as f:\n              spec = yaml.safe_load(f)\n      except FileNotFoundError:\n          assert False, \"Spec file was not created\"\n      except Exception as e:\n          assert False, f\"spec is not valid YAML: {e}\"\n      assert isinstance(spec, dict), \"spec is not a mapping\"\n      assert \"skill\" not in spec, \\\n          \"`skill:` was replaced by the `skills:` neighbourhood\"\n      skills = spec.get(\"skills\")\n      assert isinstance(skills, list) and skills, \\\n          f\"spec needs a top-level skills: list, got {skills!r}\"\n      # Entries are bare paths; tolerate the mapping form in case the agent\n      # writes one, so the assert tests the schema, not the agent's YAML style.\n      def pa(e): return e if isinstance(e, str) else (e or {}).get(\"path\")\n      assert pa(skills[0]) in (\"./SKILL.md\", \"SKILL.md\"), \\\n          f\"skills[0] should be ./SKILL.md, got {pa(skills[0])!r}\"\n      assert \"judge\" not in spec, \"engine is a runtime axis: no judge block belongs in the spec\"\n      tasks = spec.get(\"tasks\", [])\n      assert len(tasks) == 3, f\"Expected 3 tasks, got {len(tasks)}\"\n      for task in tasks:\n          assert \"name\" in task, \"Each task needs a name\"\n          assert \"prompt\" in task, \"Each task needs a prompt\"\n          assert \"expect\" in task or \"assert\" in task, \\\n              f\"Task '{task.get('name')}' needs expect or assert\"\n\nArchive v1.0.10: 5 files, 11818 bytes\n\nFiles: grill-skill.eval.yaml (8027b), REFERENCE.md (12484b), skill-card.md (1966b), SKILL.md (3049b), _meta.json (131b)\n\nFile v1.0.10:SKILL.md\n\n---\nname: grill-skill\ndescription: Build and harden a skill with evals — interview to design its eval tasks, then run, measure, and iterate. Use when the user wants to create or improve a skill's eval, or run the create → test → improve loop for a skill.\nallowed-tools: Bash, Read, Write, Edit\n---\n\n# Grill Skill\n\nInterview the user to design a skill's eval, then loop run → measure → improve until it ships. Requires `caliper` (`pipx install caliper-eval` if missing). Commands, spec skeleton, and expect/assert guidance: [REFERENCE.md](REFERENCE.md).\n\n## Entry point\n\n`/grill-skill [path]` — optional path to a `SKILL.md`.\n\n- **Path given** — use it.\n- **No path** — look for `SKILL.md` in the cwd; if found, confirm before proceeding, else ask where it is.\n\n## Phase 1 — Understand\n\nRead the `SKILL.md`. Summarize what it does, when it triggers, and what a successful run looks like. Ask the user to confirm your reading. **Wait for confirmation before continuing.**\n\n## Phase 2 — Detect eval mode\n\nLook for `*.eval.yaml` beside the `SKILL.md` (try `<dir-name>.eval.yaml` first).\n\n- **None** → New eval. **Found** → Gap-fill.\n\nInterview one question at a time and wait for each answer. Never invent the user's answers or write the spec before interviewing.\n\n### New eval — three tasks\n\nElicit three tasks, one question at a time:\n\n1. **Happy path** — the most common successful use. What did the agent do, and what would confirm it worked?\n2. **Edge case** — a tricky-but-valid input that might trip the raw agent.\n3. **Adversarial** — what the skill should refuse or avoid.\n\nTurn each answer into a task: a realistic `prompt`, an observable `expect`, and an `assert` when the outcome is checkable (see [REFERENCE.md](REFERENCE.md)). Show the proposed YAML and confirm before writing.\n\nWrite the spec beside `SKILL.md`, named `<dir-name>.eval.yaml`, with `skill.path: ./SKILL.md` and `claude-code` as the default backend for both `skill` and `judge` unless the SKILL.md targets another.\n\n### Gap-fill\n\nRead the existing spec and report its tasks. **Ask what behaviors are missing or under-tested before proposing or writing anything** — even if the user only asked you to inspect it, report first, then ask. Sharpen each gap into a task, show it, and confirm before writing it in.\n\n## Phase 3 — First run\n\nValidate the spec, then run at `k=1` (commands in [REFERENCE.md](REFERENCE.md)). Show the results. Fix any harness or config error (not a task failure) before asking the user what to do next.\n\n## Phase 4 — Iterate\n\nAsk whether to iterate or finish.\n\n- **Iterate** — after the user edits their `SKILL.md`, re-run at `k=3` and show results. Loop back.\n- **Done** — suggest an `--ablate <skill-name>` run plus a `caliper compare` to prove the skill beats the raw agent, then remind the user to commit `SKILL.md` and the `.eval.yaml` together. Mention that the ablated run is worth keeping: it cannot move when the skill's text changes, so later iterations re-diff against it instead of re-running it.\n\nFile v1.0.10:_meta.json\n\n{\n  \"ownerId\": \"kn7bp63rpwq0rm2g7m4k0c3hjn80qkhc\",\n  \"slug\": \"grill-skill\",\n  \"version\": \"1.0.10\",\n  \"publishedAt\": 1787985355819\n}\n\nFile v1.0.10:REFERENCE.md\n\n# Grill Skill Reference\n\n## Caliper commands used by this skill\n\n```bash\n# Check spec is valid before running\ncaliper validate path/to/spec.eval.yaml\n\n# First run — fast, catches spec errors\ncaliper run path/to/spec.eval.yaml --k 1\n\n# Reliability run — after iterating on the skill\ncaliper run path/to/spec.eval.yaml --k 3\n\n# Ablated run — before committing, proves the skill makes a difference.\n# Run once and keep it: it cannot move when the skill's text changes.\ncaliper run path/to/spec.eval.yaml --k 3 --ablate my-skill\n# Then diff it against the full run. A bare spec name resolves to that spec's\n# LATEST run, so address the older side by its saved results path.\ncaliper compare .caliper/results/<spec>/<ablated-run>.json <spec>\n\n# Choose the engine at run time — it is not stored in the spec (default: claude-code)\ncaliper run path/to/spec.eval.yaml --model codex:gpt-5-codex\ncaliper run path/to/spec.eval.yaml --model codex\ncaliper run path/to/spec.eval.yaml --judge-model claude-code:claude-haiku-4-5-20251001\n\n# Browse past results\ncaliper list\ncaliper report path/to/spec.eval.yaml\n\n# Compare two saved runs of the same eval (ablation: full vs. shortened, or over time)\ncaliper compare full-eval short-eval           # spec name -> latest run, or a results-JSON path\ncaliper compare a.json b.json --format json     # per-task Δ, regression flags, for scripting\n```\n\n`caliper compare <A> <B>` diffs two already-saved runs task by task: tasks are\nmatched by name, `Δ = b − a`, a negative Δ flags a regression (any-below), and a\nside with no usable attempts shows `—` (unmeasured, never a regression) so\ninfra/judge noise can't fake a loss. Under the success-rate headline, `compare` also\nshows **token and wall-clock deltas** (green = cheaper) — the \"same quality, 40%\nfewer tokens\" signal an ablation looks for. These are secondary: a token/time\nchange is **never** a regression (only the score is), and dollar cost is not tracked\n(tokens are the volume signal). Each attempt in the report also shows its tokens\nnext to its duration under `--verbose`.\n\n`compare` also reports **skill drift** — a member of the neighbourhood whose\n*text* changed between the two runs, read from the per-file hashes in each run's\nsnapshots. It is graded by provenance, not role: a drifted **git source** warns,\nbecause the spec claimed where those bytes came from and the delta you are\nreading is confounded; a drifted **path source** is shown without alarm, because\nnothing was promised about a working file and that edit is usually the thing the\nrun exists to measure.\n\n```\n ⚠ tdd changed between runs — git source, a1b2c3d → e4f5g6h; pin `ref:` to hold it fixed\n   my-skill changed between runs — path, 4fc7951 → bcbcbde\n```\n\nThis is a change in *text* at constant membership; a change in *membership* is\nthe separate neighbourhood warning.\n\n## Inspecting failures\n\nAfter any `caliper run`, failed tasks are shown automatically with their output\nand `assert_evidence` — no extra command needed. Each attempt is tagged with an\n`outcome`: a real `task_fail` reads as `✗`, while *unusable* attempts\n(`infra_error` from a rate limit that outlasted its retries, `timeout`, or\n`judge_error`)\nread as `⊘` and are excluded from the score denominator, with a separate\n\"N unusable\" count in the summary — so a throttled or judge-flaked run is not\nmistaken for a skill regression. If `caliper run --fail-fast N` stopped a task\nafter repeated `infra_error` / `timeout` outcomes, the report marks it as\n`ABORTED` and shows how many attempts ran. A run you stopped with Ctrl-C is\nsaved too, headed by an `interrupted:` line: its rates are computed over the\nattempts that ran, so read them as a smaller sample rather than a worse skill —\nand re-run before drawing a conclusion from a handful of attempts. `caliper list`\nmarks such a run with `⊘`, and `caliper compare` warns when either side is one,\nso a shallow sample cannot quietly masquerade as a delta. A run that hit a\n**spending cap** stops the same way, with the cap named as the cause — top up and\nre-run rather than reading its numbers. A `throttled:` line under the usage\nsummary means attempts were retried before they landed: the scores are sound, but\nthe wall times were fought for. If a failure is still unclear, use\n`--verbose` to see full output for all tasks (including passing ones):\n\n```bash\n# Full output for all tasks (passing + failing), untruncated\ncaliper report path/to/spec.eval.yaml --verbose\n\n# Or inspect a specific past run\ncaliper report path/to/spec.eval.yaml --run 2026-06-21T14-53-12Z --verbose\n```\n\n## Spec skeleton\n\nThe spec carries no engine — pick the backend/model at run time with `--model` /\n`--judge-model` (default `claude-code`).\n\n```yaml\nskills:                   # installed at the agent's own skills root, never\n  - ./SKILL.md            #   preloaded — the agent has to choose it\n  # add further entries to test that yours is the one that fires (they are\n  # assertable via `activates:`, not decoration). A bare string is a *path\n  # source*; a mapping is a *git source* caliper clones for you:\n  - repo: vercel-labs/agent-skills\n    ref: a1b2c3d          # optional — omit to track the default branch\n    path: skills/tdd/SKILL.md   # optional — defaults to SKILL.md at the root\n\nsandbox:\n  forbidden_files:\n    - \".*\\\\.eval\\\\.yaml$\"\n    - \"./.caliper/.*\"\n\n# Optional — only if the skill needs MCP tools. claude-code, hermes, codex backends.\nmcp:\n  weather:                       # local stdio server → mcp__weather__<tool>\n    command: python3\n    args: [./servers/weather.py]\n    env:\n      API_TOKEN: ${MCP_API_TOKEN}   # resolved from your shell at run time\n  gdrive:                        # remote (hosted) server over HTTP/SSE\n    type: http                   # http or sse\n    url: https://mcp.example.com/gdrive\n    headers:\n      Authorization: Bearer ${GDRIVE_TOKEN}   # resolved from your shell at run time\n\ntasks:\n  - name: Happy path — <what success looks like>\n    setup: <optional shell command>\n    cleanup: <optional shell command>\n    prompt: <prompt sent to the agent>\n    expect: <natural-language success criterion>\n    assert: |\n      # optional deterministic check\n\n  - name: Edge case — <tricky but valid input>\n    prompt: ...\n    expect: ...\n\n  - name: Adversarial — <what the skill should refuse or avoid>\n    prompt: ...\n    expect: <describes the refusal or safe behavior>\n\n  - name: Silence — <work no declared skill should answer>\n    prompt: ...\n    activates: []                # a trigger probe: no judge, no execution score\n```\n\nEach task needs at least one of `expect`, `assert` or `activates`.\n\n## Triggering: does the description fire?\n\nSkills are **installed** where the agent looks for them and never pasted into\nthe prompt, so whether the agent reaches for one is measurable. Two rules follow\nfor how you write prompts:\n\n- **Never name the skill in a prompt.** \"Use the commit-message skill to…\"\n  removes the very choice being measured. Write the prompt a real user would.\n- **`activates:` asserts the exact set** of skills that loaded — `[a]` means `a`\n  and nothing else, `[]` means silence. Names are the frontmatter `name:`, not\n  filenames.\n\nA task carrying only `activates:` is a **trigger probe**. It skips the judge\nentirely, so it costs far less than an execution task, and reports as `trigger\nonly` rather than a zero. Two kinds are worth generating:\n\n- **Neighbour probe** — declare a sibling skill in `skills:`, then give a prompt\n  that belongs to *it* and assert `activates: [sibling]`. This catches a\n  `description` that over-claims.\n- **Silence probe** — unrelated work, `activates: []`.\n\nActivation is scored separately from execution and never blended in, so a\nnear-zero score with a green activation column means the body is wrong, while a\nred activation column means the `description` is.\n\n## Naming convention\n\nThe spec file lives next to the skill and shares its directory name:\n\n```\nskills/my-skill/SKILL.md\nskills/my-skill/my-skill.eval.yaml   ← generated here\n```\n\n## Writing good expect: criteria\n\nBe specific about evidence. Include what the judge should look for and what counts as failure.\n\n```yaml\nexpect: |\n  Pass if the agent identifies the null dereference in user_lookup.py and\n  explains the failing path. Fail if it only gives generic style advice,\n  misses the bug, or claims tests passed without running them.\n```\n\n## When to use assert:\n\nAdd `assert:` when the outcome is a fact that an LLM judge might guess wrong:\n- File exists or contains exact content\n- Command exit code or output\n- Git state (staged, committed, clean)\n- JSON schema or exact value\n- Test suite passes or fails\n\n## MCP servers (`mcp:`)\n\nIf the skill under test needs MCP tools, declare them in a top-level `mcp:` block (a mapping keyed by server name) — a capability granted to the agent-under-test for the eval, part of the run environment like `sandbox:` (a sibling of it and of `skills:`), so they belong in the spec, not on the command line. A server is either **local stdio** (a `command`, optional `args`, optional `env`) or **remote** (`type: http`/`sse`, a `url`, optional `headers` for auth); the two field sets are mutually exclusive. Supported on **`claude-code`** (stdio + remote HTTP/SSE), **`hermes`** (stdio + remote header-auth; not remote OAuth), and **`codex`** (stdio + remote header-auth, translated into `[mcp_servers.*]` tables in the isolated `~/.codex/config.toml`; not remote OAuth). A tool call appears in the transcript as a namespaced name — `mcp__<server>__<tool>` on `claude-code` and `codex`, `mcp_<server>_<tool>` on `hermes` — so an `expect:` criterion can check the skill actually used it; word it around behaviour, not one backend's spelling, if the spec runs under more than one engine. Put secrets in a host env var and reference it as `${VAR}` inside a stdio `env:`, a remote `headers:`, or a remote `url:` — it resolves at the harness boundary from your shell at run time and never lands in the committed spec (an unset var fails the run). Running an `mcp:` spec on a backend that can't honor it is a hard error, not a silent no-op: `pi` has no MCP by design and will not honor `mcp:` natively — expose the capability as a CLI tool the skill drives or a pi extension, or run the eval on `claude-code`/`hermes`/`codex`.\n\n## Backends\n\n| Backend | Requires | Notes |\n|---|---|---|\n| `claude-code` | Claude Code CLI | Default for most skills |\n| `codex` | Codex CLI | For Codex-targeted skills |\n| `pi` | pi CLI (authenticated) | For pi / agentskills.io skills; native `--skill` loading |\n| `hermes` | Hermes Agent CLI (authenticated) | Nous Research; normalized to a neutral agent, `hermes:<provider>/<model>` picks the model |\n\nThe skill engine (`--model`) and judge engine (`--judge-model`) are chosen independently at run time. Every backend is a CLI agent; for API billing, configure a CLI with an API key rather than selecting a separate backend. When `--judge-model` is omitted, the default `claude-code` judge pins `claude-sonnet-5` at execution time so it does not inherit a stale model from the installed Claude CLI; `RunMeta.judge_model` stays empty unless you pass `--judge-model` explicitly or the autorater reports what it used.\n\n`hermes` is a stateful agent (persistent memory + persona), so Caliper strips it to a neutral agent per attempt — isolated `HERMES_HOME`, no `SOUL.md`/`MEMORY.md`, `--ignore-rules`, and only the spec's declared skills installed — and recovers the full trajectory via `hermes sessions export` after the `hermes -z` run.\n\n## Results storage\n\nResults are saved automatically to `.caliper/results/<spec-name>/<timestamp>.json`\nalongside the spec file. Each attempt records its `outcome`, optional `usage`, and\noptional `transcript` (ordered turns with `tool_name`/`tool_input`/`tool_output` when present)\nso saved runs remain inspectable after the fact — including which MCP tools fired.\nOlder JSON without `transcript` still loads (`null`). `report` and `compare` do not\nrender the transcript; it is stored for later analysis.\n\n## Troubleshooting\n\n**`Judge model ... is unavailable` / `Judge authentication failed` / `Judge rate limited`**\nThe judge CLI reached the provider and the call was refused. Caliper classifies these at the harness boundary (from the CLI's structured output) and suggests passing `--judge-model <backend[:model]>` to pick an available judge engine or model.\n\nFile v1.0.10:skill-card.md\n\n## Description:\n\nBuild and harden a skill with evals by interviewing to design eval tasks, then running, measuring, and iterating on them.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[edonadei](https://clawhub.ai/user/edonadei)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers and engineers use this skill to create or improve Caliper evals for agent skills, including new eval design, gap-filling existing evals, validation runs, and iteration based on results.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: The skill can write or edit eval YAML beside a target skill.\n\nMitigation: Review proposed YAML before allowing edits and keep the target skill and eval spec under version control.\n\nRisk: Caliper runs may invoke configured agent and judge CLIs.\n\nMitigation: Use a pinned, reviewed Caliper version and inspect commands before running them in sensitive workspaces.\n\nRisk: Eval runs can create temporary fixtures, saved results, and transcripts.\n\nMitigation: Run evals in a workspace where temporary files and persisted transcripts are acceptable.\n\n## Reference(s):\n\n- [Grill Skill Reference](artifact/REFERENCE.md)\n- [ClawHub Skill Page](https://clawhub.ai/edonadei/skills/grill-skill)\n\n## Skill Output:\n\n**Output Type(s):** [text, markdown, code, shell commands, configuration, guidance]\n\n**Output Format:** [Markdown with YAML and shell command snippets]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [May write or edit adjacent eval YAML after confirmation and summarize Caliper validation, run, report, and compare results.]\n\n## Skill Version(s):\n\n1.0.10 (source: server release evidence)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nFile v1.0.10:grill-skill.eval.yaml\n\nskills:\n  - ./SKILL.md\n\n# The engine (backend + model) is a runtime axis, not a spec field — this eval\n# runs on whatever `--model` / `--judge-model` select (default claude-code).\n\n# No explicit sandbox.forbidden_files: caliper already auto-forbids the eval\n# spec and .caliper/ by absolute path. Listing \"./.caliper/.*\" here caused a\n# false cheat flag, because these grill-skill tasks legitimately WRITE caliper\n# specs whose text contains that very pattern.\n\ntasks:\n  # Task 1 — First-turn interview discipline (open prompt, no eval present).\n  # Exercises the prose we most want to shorten: Phase 1 understanding+confirm\n  # and the one-question-at-a-time discipline. Single-shot harness, so we judge\n  # the FIRST turn only: the agent must interview, not run ahead. The assert is\n  # a deterministic guard that it did NOT fabricate answers and write a spec.\n  - name: Interviews before generating — asks, then stops (does not run ahead)\n    activates: [grill-skill]\n    setup: |\n      rm -rf /tmp/grill-fresh\n      mkdir -p /tmp/grill-fresh\n      cat > /tmp/grill-fresh/SKILL.md << 'EOF'\n      ---\n      name: changelog-writer\n      description: Use when the user wants to turn merged PRs into a changelog entry.\n      allowed-tools: Bash, Read, Write\n      ---\n\n      # Changelog Writer\n\n      Read the merged PRs since the last tag and write a grouped changelog\n      entry (Features / Fixes / Chore) to CHANGELOG.md.\n      EOF\n    cleanup: rm -rf /tmp/grill-fresh\n    prompt: >\n      I want to create a caliper eval for my skill at\n      /tmp/grill-fresh/SKILL.md. I'm here and will answer whatever you need —\n      do NOT assume what the eval tasks should be, and do NOT write any files\n      yet. What do you need to know from me to get started?\n    expect: >\n      Pass if the agent opens the interview instead of running ahead: it reads\n      the SKILL.md, gives some understanding of the skill, and asks the user for\n      input before generating anything — then STOPS to wait. The number of\n      questions does not matter. Fail if the agent skips the interview: it\n      invents the user's answers, generates the eval tasks itself without\n      asking, or writes any .eval.yaml file in this turn.\n    assert: |\n      import glob\n      # It must not have run ahead and written a spec before interviewing.\n      specs = glob.glob(\"/tmp/grill-fresh/*.eval.yaml\")\n      assert not specs, f\"Agent wrote a spec without interviewing: {specs}\"\n\n  # Task 2 — Gap-fill detection + no silent overwrite (open prompt, eval present).\n  # Exercises the gap-fill prose: detect existing eval, report tasks, ask what's\n  # missing BEFORE changing anything. Assert is a deterministic guard that the\n  # existing spec was left untouched in this first turn.\n  - name: Detects an existing eval and asks before touching it\n    activates: [grill-skill]\n    setup: |\n      rm -rf /tmp/grill-existing\n      mkdir -p /tmp/grill-existing\n      cat > /tmp/grill-existing/SKILL.md << 'EOF'\n      ---\n      name: summarize\n      description: Use when the user wants to summarize a file.\n      allowed-tools: Bash\n      ---\n\n      # Summarize\n\n      Summarize the contents of a file provided by the user.\n      EOF\n      cat > /tmp/grill-existing/summarize.eval.yaml << 'EOF'\n      skills:\n        - ./SKILL.md\n      tasks:\n        - name: Summarizes a short text file\n          prompt: Summarize /tmp/notes.txt\n          expect: The agent produces a summary of the file contents.\n      EOF\n    cleanup: rm -rf /tmp/grill-existing\n    prompt: >\n      I want to improve the eval for my skill at\n      /tmp/grill-existing/SKILL.md. I'm here to answer your\n      questions — do NOT change any files yet. Tell me the current state of my\n      eval and what you need from me.\n    expect: >\n      Pass if the agent detects the existing eval, reports the existing task\n      (\"Summarizes a short text file\"), and asks the user what behaviors are\n      missing or under-tested before proposing or writing any changes. Fail if\n      it ignores the existing eval, silently overwrites or rewrites it, or adds\n      tasks without first asking what is missing.\n    assert: |\n      import yaml\n      # The existing spec must be untouched: still exactly its one original task.\n      with open(\"/tmp/grill-existing/summarize.eval.yaml\") as f:\n          spec = yaml.safe_load(f)\n      tasks = spec.get(\"tasks\", [])\n      assert len(tasks) == 1, f\"Existing spec was modified in this turn: {len(tasks)} tasks\"\n      assert tasks[0][\"name\"] == \"Summarizes a short text file\", \"Original task was altered\"\n\n  # Task 3 — Serialization correctness (structure supplied, full artifact check).\n  # The deterministic anchor: when the user hands over the task structure, the\n  # agent must serialize a valid 3-task spec. This is the back-half of the\n  # workflow and the one task with a full artifact assert.\n  - name: Serializes a valid 3-task spec when the structure is given\n    activates: [grill-skill]\n    setup: |\n      rm -rf /tmp/grill-serialize\n      mkdir -p /tmp/grill-serialize\n      cat > /tmp/grill-serialize/SKILL.md << 'EOF'\n      ---\n      name: hello-file\n      description: Use when the user wants to write a greeting file.\n      allowed-tools: Bash\n      ---\n\n      # Hello File\n\n      Write the text \"hello world\" to /tmp/hello.txt when the user asks.\n      EOF\n    cleanup: rm -rf /tmp/grill-serialize\n    prompt: >\n      Write the file directly now — do NOT ask me any clarifying questions\n      first; I have given you everything you need. Write\n      a COMPLETE, valid caliper eval spec for the\n      skill at /tmp/grill-serialize/SKILL.md, and write it to\n      /tmp/grill-serialize/hello-file.eval.yaml (do not run it — just create the\n      file). The spec must have a top-level skills: list containing\n      ./SKILL.md (and NO backend/model or judge block — the engine is a runtime\n      axis, chosen with --model at run time), plus exactly 3 tasks: a happy path\n      where the prompt asks the agent to write the greeting file and the expect\n      is that /tmp/hello.txt contains \"hello world\"; an edge case where the user\n      specifies a custom greeting; and an adversarial case where the user asks\n      to overwrite a protected system file.\n    expect: >\n      A valid, runnable .eval.yaml is written at\n      /tmp/grill-serialize/hello-file.eval.yaml: it has a top-level skills: list\n      containing ./SKILL.md (no backend/model, no judge block) and exactly 3\n      tasks covering a happy path, an edge case, and an adversarial case.\n    assert: |\n      import yaml\n\n      path = \"/tmp/grill-serialize/hello-file.eval.yaml\"\n      try:\n          with open(path) as f:\n              spec = yaml.safe_load(f)\n      except FileNotFoundError:\n          assert False, \"Spec file was not created\"\n      except Exception as e:\n          assert False, f\"spec is not valid YAML: {e}\"\n      assert isinstance(spec, dict), \"spec is not a mapping\"\n      assert \"skill\" not in spec, \\\n          \"`skill:` was replaced by the `skills:` neighbourhood\"\n      skills = spec.get(\"skills\")\n      assert isinstance(skills, list) and skills, \\\n          f\"spec needs a top-level skills: list, got {skills!r}\"\n      # Entries are bare paths; tolerate the mapping form in case the agent\n      # writes one, so the assert tests the schema, not the agent's YAML style.\n      def pa(e): return e if isinstance(e, str) else (e or {}).get(\"path\")\n      assert pa(skills[0]) in (\"./SKILL.md\", \"SKILL.md\"), \\\n          f\"skills[0] should be ./SKILL.md, got {pa(skills[0])!r}\"\n      assert \"judge\" not in spec, \"engine is a runtime axis: no judge block belongs in the spec\"\n      tasks = spec.get(\"tasks\", [])\n      assert len(tasks) == 3, f\"Expected 3 tasks, got {len(tasks)}\"\n      for task in tasks:\n          assert \"name\" in task, \"Each task needs a name\"\n          assert \"prompt\" in task, \"Each task needs a prompt\"\n          assert \"expect\" in task or \"assert\" in task, \\\n              f\"Task '{task.get('name')}' needs expect or assert\"\n\nArchive v1.0.9: 5 files, 9611 bytes\n\nFiles: grill-skill.eval.yaml (7809b), REFERENCE.md (8078b), skill-card.md (1787b), SKILL.md (2854b), _meta.json (130b)\n\nFile v1.0.9:SKILL.md\n\n---\nname: grill-skill\ndescription: Build and harden a skill with evals — interview to design its eval tasks, then run, measure, and iterate. Use when the user wants to create or improve a skill's eval, or run the create → test → improve loop for a skill.\nallowed-tools: Bash, Read, Write, Edit\n---\n\n# Grill Skill\n\nInterview the user to design a skill's eval, then loop run → measure → improve until it ships. Requires `caliper` (`pipx install caliper-eval` if missing). Commands, spec skeleton, and expect/assert guidance: [REFERENCE.md](REFERENCE.md).\n\n## Entry point\n\n`/grill-skill [path]` — optional path to a `SKILL.md`.\n\n- **Path given** — use it.\n- **No path** — look for `SKILL.md` in the cwd; if found, confirm before proceeding, else ask where it is.\n\n## Phase 1 — Understand\n\nRead the `SKILL.md`. Summarize what it does, when it triggers, and what a successful run looks like. Ask the user to confirm your reading. **Wait for confirmation before continuing.**\n\n## Phase 2 — Detect eval mode\n\nLook for `*.eval.yaml` beside the `SKILL.md` (try `<dir-name>.eval.yaml` first).\n\n- **None** → New eval. **Found** → Gap-fill.\n\nInterview one question at a time and wait for each answer. Never invent the user's answers or write the spec before interviewing.\n\n### New eval — three tasks\n\nElicit three tasks, one question at a time:\n\n1. **Happy path** — the most common successful use. What did the agent do, and what would confirm it worked?\n2. **Edge case** — a tricky-but-valid input that might trip the raw agent.\n3. **Adversarial** — what the skill should refuse or avoid.\n\nTurn each answer into a task: a realistic `prompt`, an observable `expect`, and an `assert` when the outcome is checkable (see [REFERENCE.md](REFERENCE.md)). Show the proposed YAML and confirm before writing.\n\nWrite the spec beside `SKILL.md`, named `<dir-name>.eval.yaml`, with `skill.path: ./SKILL.md` and `claude-code` as the default backend for both `skill` and `judge` unless the SKILL.md targets another.\n\n### Gap-fill\n\nRead the existing spec and report its tasks. **Ask what behaviors are missing or under-tested before proposing or writing anything** — even if the user only asked you to inspect it, report first, then ask. Sharpen each gap into a task, show it, and confirm before writing it in.\n\n## Phase 3 — First run\n\nValidate the spec, then run at `k=1` (commands in [REFERENCE.md](REFERENCE.md)). Show the results. Fix any harness or config error (not a task failure) before asking the user what to do next.\n\n## Phase 4 — Iterate\n\nAsk whether to iterate or finish.\n\n- **Iterate** — after the user edits their `SKILL.md`, re-run at `k=3` and show results. Loop back.\n- **Done** — suggest a `--baseline` run to prove the skill beats the raw agent, then remind the user to commit `SKILL.md` and the `.eval.yaml` together.\n\nFile v1.0.9:_meta.json\n\n{\n  \"ownerId\": \"kn7bp63rpwq0rm2g7m4k0c3hjn80qkhc\",\n  \"slug\": \"grill-skill\",\n  \"version\": \"1.0.9\",\n  \"publishedAt\": 1783878268913\n}\n\nFile v1.0.9:REFERENCE.md\n\n# Grill Skill Reference\n\n## Caliper commands used by this skill\n\n```bash\n# Check spec is valid before running\ncaliper validate path/to/spec.eval.yaml\n\n# First run — fast, catches spec errors\ncaliper run path/to/spec.eval.yaml --k 1\n\n# Reliability run — after iterating on the skill\ncaliper run path/to/spec.eval.yaml --k 3\n\n# Baseline run — before committing, proves the skill makes a difference\ncaliper run path/to/spec.eval.yaml --k 3 --baseline\n\n# Choose the engine at run time — it is not stored in the spec (default: claude-code)\ncaliper run path/to/spec.eval.yaml --model codex:gpt-5-codex\ncaliper run path/to/spec.eval.yaml --model codex\ncaliper run path/to/spec.eval.yaml --judge-model claude-code:claude-haiku-4-5-20251001\n\n# Browse past results\ncaliper list\ncaliper report path/to/spec.eval.yaml\n\n# Compare two saved runs of the same eval (ablation: full vs. shortened, or over time)\ncaliper compare full-eval short-eval           # spec name -> latest run, or a results-JSON path\ncaliper compare a.json b.json --format json     # per-task Δ, regression flags, for scripting\n```\n\n`caliper compare <A> <B>` diffs two already-saved runs task by task: tasks are\nmatched by name, `Δ = b − a`, a negative Δ flags a regression (any-below), and a\nside with no usable attempts shows `—` (unmeasured, never a regression) so\ninfra/judge noise can't fake a loss. Under the success-rate headline, `compare` also\nshows **token and wall-clock deltas** (green = cheaper) — the \"same quality, 40%\nfewer tokens\" signal an ablation looks for. These are secondary: a token/time\nchange is **never** a regression (only the score is), and dollar cost is not tracked\n(tokens are the volume signal). Each attempt in the report also shows its tokens\nnext to its duration under `--verbose`.\n\n## Inspecting failures\n\nAfter any `caliper run`, failed tasks are shown automatically with their output\nand `assert_evidence` — no extra command needed. Each attempt is tagged with an\n`outcome`: a real `task_fail` reads as `✗`, while *unusable* attempts\n(`infra_error` from a rate-limit / spending-cap, `timeout`, or `judge_error`)\nread as `⊘` and are excluded from the score denominator, with a separate\n\"N unusable\" count in the summary — so a throttled or judge-flaked run is not\nmistaken for a skill regression. If `caliper run --fail-fast N` stopped a task\nafter repeated `infra_error` / `timeout` outcomes, the report marks it as\n`ABORTED` and shows how many attempts ran. If a failure is still unclear, use\n`--verbose` to see full output for all tasks (including passing ones):\n\n```bash\n# Full output for all tasks (passing + failing), untruncated\ncaliper report path/to/spec.eval.yaml --verbose\n\n# Or inspect a specific past run\ncaliper report path/to/spec.eval.yaml --run 2026-06-21T14-53-12Z --verbose\n```\n\n## Spec skeleton\n\nThe spec carries no engine — pick the backend/model at run time with `--model` /\n`--judge-model` (default `claude-code`).\n\n```yaml\nskill:\n  path: ./SKILL.md\n\nsandbox:\n  forbidden_files:\n    - \".*\\\\.eval\\\\.yaml$\"\n    - \"./.caliper/.*\"\n\n# Optional — only if the skill needs MCP tools. claude-code, hermes, codex backends.\nmcp:\n  weather:                       # local stdio server → mcp__weather__<tool>\n    command: python3\n    args: [./servers/weather.py]\n    env:\n      API_TOKEN: ${MCP_API_TOKEN}   # resolved from your shell at run time\n  gdrive:                        # remote (hosted) server over HTTP/SSE\n    type: http                   # http or sse\n    url: https://mcp.example.com/gdrive\n    headers:\n      Authorization: Bearer ${GDRIVE_TOKEN}   # resolved from your shell at run time\n\ntasks:\n  - name: Happy path — <what success looks like>\n    setup: <optional shell command>\n    cleanup: <optional shell command>\n    prompt: <prompt sent to the agent>\n    expect: <natural-language success criterion>\n    assert: |\n      # optional deterministic check\n\n  - name: Edge case — <tricky but valid input>\n    prompt: ...\n    expect: ...\n\n  - name: Adversarial — <what the skill should refuse or avoid>\n    prompt: ...\n    expect: <describes the refusal or safe behavior>\n```\n\n## Naming convention\n\nThe spec file lives next to the skill and shares its directory name:\n\n```\nskills/my-skill/SKILL.md\nskills/my-skill/my-skill.eval.yaml   ← generated here\n```\n\n## Writing good expect: criteria\n\nBe specific about evidence. Include what the judge should look for and what counts as failure.\n\n```yaml\nexpect: |\n  Pass if the agent identifies the null dereference in user_lookup.py and\n  explains the failing path. Fail if it only gives generic style advice,\n  misses the bug, or claims tests passed without running them.\n```\n\n## When to use assert:\n\nAdd `assert:` when the outcome is a fact that an LLM judge might guess wrong:\n- File exists or contains exact content\n- Command exit code or output\n- Git state (staged, committed, clean)\n- JSON schema or exact value\n- Test suite passes or fails\n\n## MCP servers (`mcp:`)\n\nIf the skill under test needs MCP tools, declare them in a top-level `mcp:` block (a mapping keyed by server name) — a capability granted to the agent-under-test for the eval, part of the run environment like `sandbox:` (a sibling of it, not nested under `skill:`), so they belong in the spec, not on the command line. A server is either **local stdio** (a `command`, optional `args`, optional `env`) or **remote** (`type: http`/`sse`, a `url`, optional `headers` for auth); the two field sets are mutually exclusive. Supported on **`claude-code`** (stdio + remote HTTP/SSE), **`hermes`** (stdio + remote header-auth; not remote OAuth), and **`codex`** (stdio + remote header-auth, translated into `[mcp_servers.*]` tables in the isolated `~/.codex/config.toml`; not remote OAuth). A tool call appears in the transcript as a namespaced name — `mcp__<server>__<tool>` on `claude-code` and `codex`, `mcp_<server>_<tool>` on `hermes` — so an `expect:` criterion can check the skill actually used it; word it around behaviour, not one backend's spelling, if the spec runs under more than one engine. Put secrets in a host env var and reference it as `${VAR}` inside a stdio `env:`, a remote `headers:`, or a remote `url:` — it resolves at the harness boundary from your shell at run time and never lands in the committed spec (an unset var fails the run). Running an `mcp:` spec on a backend that can't honor it is a hard error, not a silent no-op: `pi` has no MCP by design and will not honor `mcp:` natively — expose the capability as a CLI tool the skill drives or a pi extension, or run the eval on `claude-code`/`hermes`/`codex`.\n\n## Backends\n\n| Backend | Requires | Notes |\n|---|---|---|\n| `claude-code` | Claude Code CLI | Default for most skills |\n| `codex` | Codex CLI | For Codex-targeted skills |\n| `pi` | pi CLI (authenticated) | For pi / agentskills.io skills; native `--skill` loading |\n| `hermes` | Hermes Agent CLI (authenticated) | Nous Research; normalized to a neutral agent, `hermes:<provider>/<model>` picks the model |\n\nThe skill engine (`--model`) and judge engine (`--judge-model`) are chosen independently at run time. Every backend is a CLI agent; for API billing, configure a CLI with an API key rather than selecting a separate backend.\n\n`hermes` is a stateful agent (persistent memory + persona), so Caliper strips it to a neutral agent per attempt — isolated `HERMES_HOME`, no `SOUL.md`/`MEMORY.md`, `--ignore-rules`, skill-under-test staged as the only local skill — and recovers the full trajectory via `hermes sessions export` after the `hermes -z` run.\n\n## Results storage\n\nResults are saved automatically to `.caliper/results/<spec-name>/<timestamp>.json`\nalongside the spec file. Each attempt records its `outcome`, optional `usage`, and\noptional `transcript` (ordered turns with `tool_name`/`tool_input`/`tool_output` when present)\nso saved runs remain inspectable after the fact — including which MCP tools fired.\nOlder JSON without `transcript` still loads (`null`). `report` and `compare` do not\nrender the transcript; it is stored for later analysis.\n\nFile v1.0.9:skill-card.md\n\n## Description: <br>\nGrill Skill helps developers design, run, measure, and iterate skill evals with Caliper. <br>\n\nThis skill is ready for commercial/non-commercial use. <br>\n\n## Publisher: <br>\n[edonadei](https://clawhub.ai/user/edonadei) <br>\n\n### License/Terms of Use: <br>\nMIT-0 <br>\n\n\n## Use Case: <br>\nDevelopers and skill authors use this skill to create or improve Caliper evals for agent skills, then run quick and reliability checks before release. <br>\n\n### Deployment Geography for Use: <br>\nGlobal <br>\n\n## Known Risks and Mitigations: <br>\nRisk: Generated eval setup, cleanup, MCP, backend, or secret configuration could cause unintended file changes or external service use when run. <br>\nMitigation: Review generated eval YAML and Caliper commands before execution, especially sections that write files, configure MCP servers, select backends, or reference secrets. <br>\n\n\n## Reference(s): <br>\n- [Grill Skill Reference](artifact/REFERENCE.md) <br>\n- [ClawHub skill page](https://clawhub.ai/edonadei/skills/grill-skill) <br>\n\n\n## Skill Output: <br>\n**Output Type(s):** [text, markdown, code, shell commands, configuration, guidance] <br>\n**Output Format:** [Markdown guidance with YAML snippets and shell commands] <br>\n**Output Parameters:** [1D] <br>\n**Other Properties Related to Output:** [May create or edit Caliper eval YAML files and recommend Caliper validation, run, report, compare, and baseline commands.] <br>\n\n## Skill Version(s): <br>\n1.0.9 (source: server release metadata) <br>\n\n## Ethical Considerations: <br>\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment. <br>\n\nFile v1.0.9:grill-skill.eval.yaml\n\nskill:\n  path: ./SKILL.md\n\n# The engine (backend + model) is a runtime axis, not a spec field — this eval\n# runs on whatever `--model` / `--judge-model` select (default claude-code).\n\n# No explicit sandbox.forbidden_files: caliper already auto-forbids the eval\n# spec and .caliper/ by absolute path. Listing \"./.caliper/.*\" here caused a\n# false cheat flag, because these grill-skill tasks legitimately WRITE caliper\n# specs whose text contains that very pattern.\n\ntasks:\n  # Task 1 — First-turn interview discipline (open prompt, no eval present).\n  # Exercises the prose we most want to shorten: Phase 1 understanding+confirm\n  # and the one-question-at-a-time discipline. Single-shot harness, so we judge\n  # the FIRST turn only: the agent must interview, not run ahead. The assert is\n  # a deterministic guard that it did NOT fabricate answers and write a spec.\n  - name: Interviews before generating — asks, then stops (does not run ahead)\n    setup: |\n      rm -rf /tmp/grill-fresh\n      mkdir -p /tmp/grill-fresh\n      cat > /tmp/grill-fresh/SKILL.md << 'EOF'\n      ---\n      name: changelog-writer\n      description: Use when the user wants to turn merged PRs into a changelog entry.\n      allowed-tools: Bash, Read, Write\n      ---\n\n      # Changelog Writer\n\n      Read the merged PRs since the last tag and write a grouped changelog\n      entry (Features / Fixes / Chore) to CHANGELOG.md.\n      EOF\n    cleanup: rm -rf /tmp/grill-fresh\n    prompt: >\n      I want to create a caliper eval for my skill at\n      /tmp/grill-fresh/SKILL.md. I'm here and will answer whatever you need —\n      do NOT assume what the eval tasks should be, and do NOT write any files\n      yet. What do you need to know from me to get started?\n    expect: >\n      Pass if the agent opens the interview instead of running ahead: it reads\n      the SKILL.md, gives some understanding of the skill, and asks the user for\n      input before generating anything — then STOPS to wait. The number of\n      questions does not matter. Fail if the agent skips the interview: it\n      invents the user's answers, generates the eval tasks itself without\n      asking, or writes any .eval.yaml file in this turn.\n    assert: |\n      import glob\n      # It must not have run ahead and written a spec before interviewing.\n      specs = glob.glob(\"/tmp/grill-fresh/*.eval.yaml\")\n      assert not specs, f\"Agent wrote a spec without interviewing: {specs}\"\n\n  # Task 2 — Gap-fill detection + no silent overwrite (open prompt, eval present).\n  # Exercises the gap-fill prose: detect existing eval, report tasks, ask what's\n  # missing BEFORE changing anything. Assert is a deterministic guard that the\n  # existing spec was left untouched in this first turn.\n  - name: Detects an existing eval and asks before touching it\n    setup: |\n      rm -rf /tmp/grill-existing\n      mkdir -p /tmp/grill-existing\n      cat > /tmp/grill-existing/SKILL.md << 'EOF'\n      ---\n      name: summarize\n      description: Use when the user wants to summarize a file.\n      allowed-tools: Bash\n      ---\n\n      # Summarize\n\n      Summarize the contents of a file provided by the user.\n      EOF\n      cat > /tmp/grill-existing/summarize.eval.yaml << 'EOF'\n      skill:\n        path: ./SKILL.md\n      tasks:\n        - name: Summarizes a short text file\n          prompt: Summarize /tmp/notes.txt\n          expect: The agent produces a summary of the file contents.\n      EOF\n    cleanup: rm -rf /tmp/grill-existing\n    prompt: >\n      I want to improve the eval for my skill at\n      /tmp/grill-existing/SKILL.md using grill-skill. I'm here to answer your\n      questions — do NOT change any files yet. Tell me the current state of my\n      eval and what you need from me.\n    expect: >\n      Pass if the agent detects the existing eval, reports the existing task\n      (\"Summarizes a short text file\"), and asks the user what behaviors are\n      missing or under-tested before proposing or writing any changes. Fail if\n      it ignores the existing eval, silently overwrites or rewrites it, or adds\n      tasks without first asking what is missing.\n    assert: |\n      import yaml\n      # The existing spec must be untouched: still exactly its one original task.\n      with open(\"/tmp/grill-existing/summarize.eval.yaml\") as f:\n          spec = yaml.safe_load(f)\n      tasks = spec.get(\"tasks\", [])\n      assert len(tasks) == 1, f\"Existing spec was modified in this turn: {len(tasks)} tasks\"\n      assert tasks[0][\"name\"] == \"Summarizes a short text file\", \"Original task was altered\"\n\n  # Task 3 — Serialization correctness (structure supplied, full artifact check).\n  # The deterministic anchor: when the user hands over the task structure, the\n  # agent must serialize a valid 3-task spec. This is the back-half of the\n  # workflow and the one task with a full artifact assert.\n  - name: Serializes a valid 3-task spec when the structure is given\n    setup: |\n      rm -rf /tmp/grill-serialize\n      mkdir -p /tmp/grill-serialize\n      cat > /tmp/grill-serialize/SKILL.md << 'EOF'\n      ---\n      name: hello-file\n      description: Use when the user wants to write a greeting file.\n      allowed-tools: Bash\n      ---\n\n      # Hello File\n\n      Write the text \"hello world\" to /tmp/hello.txt when the user asks.\n      EOF\n    cleanup: rm -rf /tmp/grill-serialize\n    prompt: >\n      Write the file directly now — do NOT ask me any clarifying questions\n      first; I have given you everything you need. Use the grill-skill to write\n      a COMPLETE, valid caliper eval spec for the\n      skill at /tmp/grill-serialize/SKILL.md, and write it to\n      /tmp/grill-serialize/hello-file.eval.yaml (do not run it — just create the\n      file). The spec must have a top-level skill block with skill.path\n      ./SKILL.md (and NO backend/model or judge block — the engine is a runtime\n      axis, chosen with --model at run time), plus exactly 3 tasks: a happy path\n      where the prompt asks the agent to write the greeting file and the expect\n      is that /tmp/hello.txt contains \"hello world\"; an edge case where the user\n      specifies a custom greeting; and an adversarial case where the user asks\n      to overwrite a protected system file.\n    expect: >\n      A valid, runnable .eval.yaml is written at\n      /tmp/grill-serialize/hello-file.eval.yaml: it has a top-level skill block\n      (path ./SKILL.md, no backend/model, no judge block) and exactly 3 tasks\n      covering a happy path, an edge case, and an adversarial case.\n    assert: |\n      import yaml\n      def pa(n): return n if isinstance(n, str) else (n or {}).get(\"path\")\n\n      path = \"/tmp/grill-serialize/hello-file.eval.yaml\"\n      try:\n          with open(path) as f:\n              spec = yaml.safe_load(f)\n      except FileNotFoundError:\n          assert False, \"Spec file was not created\"\n      except Exception as e:\n          assert False, f\"spec is not valid YAML: {e}\"\n      assert isinstance(spec, dict), \"spec is not a mapping\"\n      assert pa(spec.get(\"skill\")) in (\"./SKILL.md\", \"SKILL.md\"), \\\n          f\"skill.path should be ./SKILL.md, got {pa(spec.get('skill'))!r}\"\n      skill = spec.get(\"skill\") or {}\n      if isinstance(skill, dict):\n          assert \"backend\" not in skill and \"model\" not in skill, \\\n              \"engine is a runtime axis: skill must not pin backend/model\"\n      assert \"judge\" not in spec, \"engine is a runtime axis: no judge block belongs in the spec\"\n      tasks = spec.get(\"tasks\", [])\n      assert len(tasks) == 3, f\"Expected 3 tasks, got {len(tasks)}\"\n      for task in tasks:\n          assert \"name\" in task, \"Each task needs a name\"\n          assert \"prompt\" in task, \"Each task needs a prompt\"\n          assert \"expect\" in task or \"assert\" in task, \\\n              f\"Task '{task.get('name')}' needs expect or assert\"\n\nArchive v1.0.8: 5 files, 8432 bytes\n\nFiles: grill-skill.eval.yaml (7809b), REFERENCE.md (5337b), skill-card.md (1773b), SKILL.md (2854b), _meta.json (130b)\n\nFile v1.0.8:SKILL.md\n\n---\nname: grill-skill\ndescription: Build and harden a skill with evals — interview to design its eval tasks, then run, measure, and iterate. Use when the user wants to create or improve a skill's eval, or run the create → test → improve loop for a skill.\nallowed-tools: Bash, Read, Write, Edit\n---\n\n# Grill Skill\n\nInterview the user to design a skill's eval, then loop run → measure → improve until it ships. Requires `caliper` (`pipx install caliper-eval` if missing). Commands, spec skeleton, and expect/assert guidance: [REFERENCE.md](REFERENCE.md).\n\n## Entry point\n\n`/grill-skill [path]` — optional path to a `SKILL.md`.\n\n- **Path given** — use it.\n- **No path** — look for `SKILL.md` in the cwd; if found, confirm before proceeding, else ask where it is.\n\n## Phase 1 — Understand\n\nRead the `SKILL.md`. Summarize what it does, when it triggers, and what a successful run looks like. Ask the user to confirm your reading. **Wait for confirmation before continuing.**\n\n## Phase 2 — Detect eval mode\n\nLook for `*.eval.yaml` beside the `SKILL.md` (try `<dir-name>.eval.yaml` first).\n\n- **None** → New eval. **Found** → Gap-fill.\n\nInterview one question at a time and wait for each answer. Never invent the user's answers or write the spec before interviewing.\n\n### New eval — three tasks\n\nElicit three tasks, one question at a time:\n\n1. **Happy path** — the most common successful use. What did the agent do, and what would confirm it worked?\n2. **Edge case** — a tricky-but-valid input that might trip the raw agent.\n3. **Adversarial** — what the skill should refuse or avoid.\n\nTurn each answer into a task: a realistic `prompt`, an observable `expect`, and an `assert` when the outcome is checkable (see [REFERENCE.md](REFERENCE.md)). Show the proposed YAML and confirm before writing.\n\nWrite the spec beside `SKILL.md`, named `<dir-name>.eval.yaml`, with `skill.path: ./SKILL.md` and `claude-code` as the default backend for both `skill` and `judge` unless the SKILL.md targets another.\n\n### Gap-fill\n\nRead the existing spec and report its tasks. **Ask what behaviors are missing or under-tested before proposing or writing anything** — even if the user only asked you to inspect it, report first, then ask. Sharpen each gap into a task, show it, and confirm before writing it in.\n\n## Phase 3 — First run\n\nValidate the spec, then run at `k=1` (commands in [REFERENCE.md](REFERENCE.md)). Show the results. Fix any harness or config error (not a task failure) before asking the user what to do next.\n\n## Phase 4 — Iterate\n\nAsk whether to iterate or finish.\n\n- **Iterate** — after the user edits their `SKILL.md`, re-run at `k=3` and show results. Loop back.\n- **Done** — suggest a `--baseline` run to prove the skill beats the raw agent, then remind the user to commit `SKILL.md` and the `.eval.yaml` together.\n\nFile v1.0.8:_meta.json\n\n{\n  \"ownerId\": \"kn7bp63rpwq0rm2g7m4k0c3hjn80qkhc\",\n  \"slug\": \"grill-skill\",\n  \"version\": \"1.0.8\",\n  \"publishedAt\": 1783282908137\n}\n\nFile v1.0.8:REFERENCE.md\n\n# Grill Skill Reference\n\n## Caliper commands used by this skill\n\n```bash\n# Check spec is valid before running\ncaliper validate path/to/spec.eval.yaml\n\n# First run — fast, catches spec errors\ncaliper run path/to/spec.eval.yaml --k 1\n\n# Reliability run — after iterating on the skill\ncaliper run path/to/spec.eval.yaml --k 3\n\n# Baseline run — before committing, proves the skill makes a difference\ncaliper run path/to/spec.eval.yaml --k 3 --baseline\n\n# Choose the engine at run time — it is not stored in the spec (default: claude-code)\ncaliper run path/to/spec.eval.yaml --model codex:gpt-5-codex\ncaliper run path/to/spec.eval.yaml --model codex\ncaliper run path/to/spec.eval.yaml --judge-model claude-code:claude-haiku-4-5-20251001\n\n# Browse past results\ncaliper list\ncaliper report path/to/spec.eval.yaml\n\n# Compare two saved runs of the same eval (ablation: full vs. shortened, or over time)\ncaliper compare full-eval short-eval           # spec name -> latest run, or a results-JSON path\ncaliper compare a.json b.json --format json     # per-task Δ, regression flags, for scripting\n```\n\n`caliper compare <A> <B>` diffs two already-saved runs task by task: tasks are\nmatched by name, `Δ = b − a`, a negative Δ flags a regression (any-below), and a\nside with no usable attempts shows `—` (unmeasured, never a regression) so\ninfra/judge noise can't fake a loss. Under the success-rate headline, `compare` also\nshows **token and wall-clock deltas** (green = cheaper) — the \"same quality, 40%\nfewer tokens\" signal an ablation looks for. These are secondary: a token/time\nchange is **never** a regression (only the score is), and dollar cost is not tracked\n(tokens are the volume signal). Each attempt in the report also shows its tokens\nnext to its duration under `--verbose`.\n\n## Inspecting failures\n\nAfter any `caliper run`, failed tasks are shown automatically with their output\nand `assert_evidence` — no extra command needed. Each attempt is tagged with an\n`outcome`: a real `task_fail` reads as `✗`, while *unusable* attempts\n(`infra_error` from a rate-limit / spending-cap, `timeout`, or `judge_error`)\nread as `⊘` and are excluded from the score denominator, with a separate\n\"N unusable\" count in the summary — so a throttled or judge-flaked run is not\nmistaken for a skill regression. If `caliper run --fail-fast N` stopped a task\nafter repeated `infra_error` / `timeout` outcomes, the report marks it as\n`ABORTED` and shows how many attempts ran. If a failure is still unclear, use\n`--verbose` to see full output for all tasks (including passing ones):\n\n```bash\n# Full output for all tasks (passing + failing), untruncated\ncaliper report path/to/spec.eval.yaml --verbose\n\n# Or inspect a specific past run\ncaliper report path/to/spec.eval.yaml --run 2026-06-21T14-53-12Z --verbose\n```\n\n## Spec skeleton\n\nThe spec carries no engine — pick the backend/model at run time with `--model` /\n`--judge-model` (default `claude-code`).\n\n```yaml\nskill:\n  path: ./SKILL.md\n\nsandbox:\n  forbidden_files:\n    - \".*\\\\.eval\\\\.yaml$\"\n    - \"./.caliper/.*\"\n\ntasks:\n  - name: Happy path — <what success looks like>\n    setup: <optional shell command>\n    cleanup: <optional shell command>\n    prompt: <prompt sent to the agent>\n    expect: <natural-language success criterion>\n    assert: |\n      # optional deterministic check\n\n  - name: Edge case — <tricky but valid input>\n    prompt: ...\n    expect: ...\n\n  - name: Adversarial — <what the skill should refuse or avoid>\n    prompt: ...\n    expect: <describes the refusal or safe behavior>\n```\n\n## Naming convention\n\nThe spec file lives next to the skill and shares its directory name:\n\n```\nskills/my-skill/SKILL.md\nskills/my-skill/my-skill.eval.yaml   ← generated here\n```\n\n## Writing good expect: criteria\n\nBe specific about evidence. Include what the judge should look for and what counts as failure.\n\n```yaml\nexpect: |\n  Pass if the agent identifies the null dereference in user_lookup.py and\n  explains the failing path. Fail if it only gives generic style advice,\n  misses the bug, or claims tests passed without running them.\n```\n\n## When to use assert:\n\nAdd `assert:` when the outcome is a fact that an LLM judge might guess wrong:\n- File exists or contains exact content\n- Command exit code or output\n- Git state (staged, committed, clean)\n- JSON schema or exact value\n- Test suite passes or fails\n\n## Backends\n\n| Backend | Requires | Notes |\n|---|---|---|\n| `claude-code` | Claude Code CLI | Default for most skills |\n| `codex` | Codex CLI | For Codex-targeted skills |\n| `pi` | pi CLI (authenticated) | For pi / agentskills.io skills; native `--skill` loading |\n| `hermes` | Hermes Agent CLI (authenticated) | Nous Research; normalized to a neutral agent, `hermes:<provider>/<model>` picks the model |\n\nThe skill engine (`--model`) and judge engine (`--judge-model`) are chosen independently at run time. Every backend is a CLI agent; for API billing, configure a CLI with an API key rather than selecting a separate backend.\n\n`hermes` is a stateful agent (persistent memory + persona), so Caliper strips it to a neutral agent per attempt — isolated `HERMES_HOME`, no `SOUL.md`/`MEMORY.md`, `--ignore-rules`, skill-under-test staged as the only local skill — and recovers the full trajectory via `hermes sessions export` after the `hermes -z` run.\n\nFile v1.0.8:skill-card.md\n\n## Description: <br>\nBuild and harden a skill with evals by interviewing to design eval tasks, then running, measuring, and iterating on the skill. <br>\n\nThis skill is ready for commercial/non-commercial use. <br>\n\n## Publisher: <br>\n[edonadei](https://clawhub.ai/user/edonadei) <br>\n\n### License/Terms of Use: <br>\nMIT-0 <br>\n\n\n## Use Case: <br>\nDevelopers and skill authors use Grill Skill to create or improve Caliper evals for an agent skill, including new eval design, gap filling for existing evals, first-run validation, and iteration toward a stronger skill. <br>\n\n### Deployment Geography for Use: <br>\nGlobal <br>\n\n## Known Risks and Mitigations: <br>\nRisk: The skill can write evaluation specs beside a skill and run Caliper or agent CLI commands. <br>\nMitigation: Review proposed eval content before confirming file writes or command runs. <br>\n\n\n## Reference(s): <br>\n- [Grill Skill on ClawHub](https://clawhub.ai/edonadei/skills/grill-skill) <br>\n- [Grill Skill Reference](artifact/REFERENCE.md) <br>\n\n\n## Skill Output: <br>\n**Output Type(s):** [text, markdown, code, shell commands, configuration, guidance] <br>\n**Output Format:** [Markdown with YAML snippets, Python assertions, and shell commands] <br>\n**Output Parameters:** [1D] <br>\n**Other Properties Related to Output:** [Can write .eval.yaml files beside a target skill and may run Caliper or agent CLI commands after user confirmation.] <br>\n\n## Skill Version(s): <br>\n1.0.8 (source: server evidence release.version) <br>\n\n## Ethical Considerations: <br>\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment. <br>\n\nFile v1.0.8:grill-skill.eval.yaml\n\nskill:\n  path: ./SKILL.md\n\n# The engine (backend + model) is a runtime axis, not a spec field — this eval\n# runs on whatever `--model` / `--judge-model` select (default claude-code).\n\n# No explicit sandbox.forbidden_files: caliper already auto-forbids the eval\n# spec and .caliper/ by absolute path. Listing \"./.caliper/.*\" here caused a\n# false cheat flag, because these grill-skill tasks legitimately WRITE caliper\n# specs whose text contains that very pattern.\n\ntasks:\n  # Task 1 — First-turn interview discipline (open prompt, no eval present).\n  # Exercises the prose we most want to shorten: Phase 1 understanding+confirm\n  # and the one-question-at-a-time discipline. Single-shot harness, so we judge\n  # the FIRST turn only: the agent must interview, not run ahead. The assert is\n  # a deterministic guard that it did NOT fabricate answers and write a spec.\n  - name: Interviews before generating — asks, then stops (does not run ahead)\n    setup: |\n      rm -rf /tmp/grill-fresh\n      mkdir -p /tmp/grill-fresh\n      cat > /tmp/grill-fresh/SKILL.md << 'EOF'\n      ---\n      name: changelog-writer\n      description: Use when the user wants to turn merged PRs into a changelog entry.\n      allowed-tools: Bash, Read, Write\n      ---\n\n      # Changelog Writer\n\n      Read the merged PRs since the last tag and write a grouped changelog\n      entry (Features / Fixes / Chore) to CHANGELOG.md.\n      EOF\n    cleanup: rm -rf /tmp/grill-fresh\n    prompt: >\n      I want to create a caliper eval for my skill at\n      /tmp/grill-fresh/SKILL.md. I'm here and will answer whatever you need —\n      do NOT assume what the eval tasks should be, and do NOT write any files\n      yet. What do you need to know from me to get started?\n    expect: >\n      Pass if the agent opens the interview instead of running ahead: it reads\n      the SKILL.md, gives some understanding of the skill, and asks the user for\n      input before generating anything — then STOPS to wait. The number of\n      questions does not matter. Fail if the agent skips the interview: it\n      invents the user's answers, generates the eval tasks itself without\n      asking, or writes any .eval.yaml file in this turn.\n    assert: |\n      import glob\n      # It must not have run ahead and written a spec before interviewing.\n      specs = glob.glob(\"/tmp/grill-fresh/*.eval.yaml\")\n      assert not specs, f\"Agent wrote a spec without interviewing: {specs}\"\n\n  # Task 2 — Gap-fill detection + no silent overwrite (open prompt, eval present).\n  # Exercises the gap-fill prose: detect existing eval, report tasks, ask what's\n  # missing BEFORE changing anything. Assert is a deterministic guard that the\n  # existing spec was left untouched in this first turn.\n  - name: Detects an existing eval and asks before touching it\n    setup: |\n      rm -rf /tmp/grill-existing\n      mkdir -p /tmp/grill-existing\n      cat > /tmp/grill-existing/SKILL.md << 'EOF'\n      ---\n      name: summarize\n      description: Use when the user wants to summarize a file.\n      allowed-tools: Bash\n      ---\n\n      # Summarize\n\n      Summarize the contents of a file provided by the user.\n      EOF\n      cat > /tmp/grill-existing/summarize.eval.yaml << 'EOF'\n      skill:\n        path: ./SKILL.md\n      tasks:\n        - name: Summarizes a short text file\n          prompt: Summarize /tmp/notes.txt\n          expect: The agent produces a summary of the file contents.\n      EOF\n    cleanup: rm -rf /tmp/grill-existing\n    prompt: >\n      I want to improve the eval for my skill at\n      /tmp/grill-existing/SKILL.md using grill-skill. I'm here to answer your\n      questions — do NOT change any files yet. Tell me the current state of my\n      eval and what you need from me.\n    expect: >\n      Pass if the agent detects the existing eval, reports the existing task\n      (\"Summarizes a short text file\"), and asks the user what behaviors are\n      missing or under-tested before proposing or writing any changes. Fail if\n      it ignores the existing eval, silently overwrites or rewrites it, or adds\n      tasks without first asking what is missing.\n    assert: |\n      import yaml\n      # The existing spec must be untouched: still exactly its one original task.\n      with open(\"/tmp/grill-existing/summarize.eval.yaml\") as f:\n          spec = yaml.safe_load(f)\n      tasks = spec.get(\"tasks\", [])\n      assert len(tasks) == 1, f\"Existing spec was modified in this turn: {len(tasks)} tasks\"\n      assert tasks[0][\"name\"] == \"Summarizes a short text file\", \"Original task was altered\"\n\n  # Task 3 — Serialization correctness (structure supplied, full artifact check).\n  # The deterministic anchor: when the user hands over the task structure, the\n  # agent must serialize a valid 3-task spec. This is the back-half of the\n  # workflow and the one task with a full artifact assert.\n  - name: Serializes a valid 3-task spec when the structure is given\n    setup: |\n      rm -rf /tmp/grill-serialize\n      mkdir -p /tmp/grill-serialize\n      cat > /tmp/grill-serialize/SKILL.md << 'EOF'\n      ---\n      name: hello-file\n      description: Use when the user wants to write a greeting file.\n      allowed-tools: Bash\n      ---\n\n      # Hello File\n\n      Write the text \"hello world\" to /tmp/hello.txt when the user asks.\n      EOF\n    cleanup: rm -rf /tmp/grill-serialize\n    prompt: >\n      Write the file directly now — do NOT ask me any clarifying questions\n      first; I have given you everything you need. Use the grill-skill to write\n      a COMPLETE, valid caliper eval spec for the\n      skill at /tmp/grill-serialize/SKILL.md, and write it to\n      /tmp/grill-serialize/hello-file.eval.yaml (do not run it — just create the\n      file). The spec must have a top-level skill block with skill.path\n      ./SKILL.md (and NO backend/model or judge block — the engine is a runtime\n      axis, chosen with --model at run time), plus exactly 3 tasks: a happy path\n      where the prompt asks the agent to write the greeting file and the expect\n      is that /tmp/hello.txt contains \"hello world\"; an edge case where the user\n      specifies a custom greeting; and an adversarial case where the user asks\n      to overwrite a protected system file.\n    expect: >\n      A valid, runnable .eval.yaml is written at\n      /tmp/grill-serialize/hello-file.eval.yaml: it has a top-level skill block\n      (path ./SKILL.md, no backend/model, no judge block) and exactly 3 tasks\n      covering a happy path, an edge case, and an adversarial case.\n    assert: |\n      import yaml\n      def pa(n): return n if isinstance(n, str) else (n or {}).get(\"path\")\n\n      path = \"/tmp/grill-serialize/hello-file.eval.yaml\"\n      try:\n          with open(path) as f:\n              spec = yaml.safe_load(f)\n      except FileNotFoundError:\n          assert False, \"Spec file was not created\"\n      except Exception as e:\n          assert False, f\"spec is not valid YAML: {e}\"\n      assert isinstance(spec, dict), \"spec is not a mapping\"\n      assert pa(spec.get(\"skill\")) in (\"./SKILL.md\", \"SKILL.md\"), \\\n          f\"skill.path should be ./SKILL.md, got {pa(spec.get('skill'))!r}\"\n      skill = spec.get(\"skill\") or {}\n      if isinstance(skill, dict):\n          assert \"backend\" not in skill and \"model\" not in skill, \\\n              \"engine is a runtime axis: skill must not pin backend/model\"\n      assert \"judge\" not in spec, \"engine is a runtime axis: no judge block belongs in the spec\"\n      tasks = spec.get(\"tasks\", [])\n      assert len(tasks) == 3, f\"Expected 3 tasks, got {len(tasks)}\"\n      for task in tasks:\n          assert \"name\" in task, \"Each task needs a name\"\n          assert \"prompt\" in task, \"Each task needs a prompt\"\n          assert \"expect\" in task or \"assert\" in task, \\\n              f\"Task '{task.get('name')}' needs expect or assert\"\n\nArchive v1.0.7: 5 files, 8358 bytes\n\nFiles: grill-skill.eval.yaml (7809b), REFERENCE.md (4923b), skill-card.md (2092b), SKILL.md (2854b), _meta.json (130b)\n\nFile v1.0.7:SKILL.md\n\n---\nname: grill-skill\ndescription: Build and harden a skill with evals — interview to design its eval tasks, then run, measure, and iterate. Use when the user wants to create or improve a skill's eval, or run the create → test → improve loop for a skill.\nallowed-tools: Bash, Read, Write, Edit\n---\n\n# Grill Skill\n\nInterview the user to design a skill's eval, then loop run → measure → improve until it ships. Requires `caliper` (`pipx install caliper-eval` if missing). Commands, spec skeleton, and expect/assert guidance: [REFERENCE.md](REFERENCE.md).\n\n## Entry point\n\n`/grill-skill [path]` — optional path to a `SKILL.md`.\n\n- **Path given** — use it.\n- **No path** — look for `SKILL.md` in the cwd; if found, confirm before proceeding, else ask where it is.\n\n## Phase 1 — Understand\n\nRead the `SKILL.md`. Summarize what it does, when it triggers, and what a successful run looks like. Ask the user to confirm your reading. **Wait for confirmation before continuing.**\n\n## Phase 2 — Detect eval mode\n\nLook for `*.eval.yaml` beside the `SKILL.md` (try `<dir-name>.eval.yaml` first).\n\n- **None** → New eval. **Found** → Gap-fill.\n\nInterview one question at a time and wait for each answer. Never invent the user's answers or write the spec before interviewing.\n\n### New eval — three tasks\n\nElicit three tasks, one question at a time:\n\n1. **Happy path** — the most common successful use. What did the agent do, and what would confirm it worked?\n2. **Edge case** — a tricky-but-valid input that might trip the raw agent.\n3. **Adversarial** — what the skill should refuse or avoid.\n\nTurn each answer into a task: a realistic `prompt`, an observable `expect`, and an `assert` when the outcome is checkable (see [REFERENCE.md](REFERENCE.md)). Show the proposed YAML and confirm before writing.\n\nWrite the spec beside `SKILL.md`, named `<dir-name>.eval.yaml`, with `skill.path: ./SKILL.md` and `claude-code` as the default backend for both `skill` and `judge` unless the SKILL.md targets another.\n\n### Gap-fill\n\nRead the existing spec and report its tasks. **Ask what behaviors are missing or under-tested before proposing or writing anything** — even if the user only asked you to inspect it, report first, then ask. Sharpen each gap into a task, show it, and confirm before writing it in.\n\n## Phase 3 — First run\n\nValidate the spec, then run at `k=1` (commands in [REFERENCE.md](REFERENCE.md)). Show the results. Fix any harness or config error (not a task failure) before asking the user what to do next.\n\n## Phase 4 — Iterate\n\nAsk whether to iterate or finish.\n\n- **Iterate** — after the user edits their `SKILL.md`, re-run at `k=3` and show results. Loop back.\n- **Done** — suggest a `--baseline` run to prove the skill beats the raw agent, then remind the user to commit `SKILL.md` and the `.eval.yaml` together.\n\nFile v1.0.7:_meta.json\n\n{\n  \"ownerId\": \"kn7bp63rpwq0rm2g7m4k0c3hjn80qkhc\",\n  \"slug\": \"grill-skill\",\n  \"version\": \"1.0.7\",\n  \"publishedAt\": 1783108984820\n}\n\nFile v1.0.7:REFERENCE.md\n\n# Grill Skill Reference\n\n## Caliper commands used by this skill\n\n```bash\n# Check spec is valid before running\ncaliper validate path/to/spec.eval.yaml\n\n# First run — fast, catches spec errors\ncaliper run path/to/spec.eval.yaml --k 1\n\n# Reliability run — after iterating on the skill\ncaliper run path/to/spec.eval.yaml --k 3\n\n# Baseline run — before committing, proves the skill makes a difference\ncaliper run path/to/spec.eval.yaml --k 3 --baseline\n\n# Choose the engine at run time — it is not stored in the spec (default: claude-code)\ncaliper run path/to/spec.eval.yaml --model codex:gpt-5-codex\ncaliper run path/to/spec.eval.yaml --model codex\ncaliper run path/to/spec.eval.yaml --judge-model claude-code:claude-haiku-4-5-20251001\n\n# Browse past results\ncaliper list\ncaliper report path/to/spec.eval.yaml\n\n# Compare two saved runs of the same eval (ablation: full vs. shortened, or over time)\ncaliper compare full-eval short-eval           # spec name -> latest run, or a results-JSON path\ncaliper compare a.json b.json --format json     # per-task Δ, regression flags, for scripting\n```\n\n`caliper compare <A> <B>` diffs two already-saved runs task by task: tasks are\nmatched by name, `Δ = b − a`, a negative Δ flags a regression (any-below), and a\nside with no usable attempts shows `—` (unmeasured, never a regression) so\ninfra/judge noise can't fake a loss.\n\n## Inspecting failures\n\nAfter any `caliper run`, failed tasks are shown automatically with their output\nand `assert_evidence` — no extra command needed. Each attempt is tagged with an\n`outcome`: a real `task_fail` reads as `✗`, while *unusable* attempts\n(`infra_error` from a rate-limit / spending-cap, `timeout`, or `judge_error`)\nread as `⊘` and are excluded from the pass@k denominator, with a separate\n\"N unusable\" count in the summary — so a throttled or judge-flaked run is not\nmistaken for a skill regression. If `caliper run --fail-fast N` stopped a task\nafter repeated `infra_error` / `timeout` outcomes, the report marks it as\n`ABORTED` and shows how many attempts ran. If a failure is still unclear, use\n`--verbose` to see full output for all tasks (including passing ones):\n\n```bash\n# Full output for all tasks (passing + failing), untruncated\ncaliper report path/to/spec.eval.yaml --verbose\n\n# Or inspect a specific past run\ncaliper report path/to/spec.eval.yaml --run 2026-06-21T14-53-12Z --verbose\n```\n\n## Spec skeleton\n\nThe spec carries no engine — pick the backend/model at run time with `--model` /\n`--judge-model` (default `claude-code`).\n\n```yaml\nskill:\n  path: ./SKILL.md\n\nsandbox:\n  forbidden_files:\n    - \".*\\\\.eval\\\\.yaml$\"\n    - \"./.caliper/.*\"\n\ntasks:\n  - name: Happy path — <what success looks like>\n    setup: <optional shell command>\n    cleanup: <optional shell command>\n    prompt: <prompt sent to the agent>\n    expect: <natural-language success criterion>\n    assert: |\n      # optional deterministic check\n\n  - name: Edge case — <tricky but valid input>\n    prompt: ...\n    expect: ...\n\n  - name: Adversarial — <what the skill should refuse or avoid>\n    prompt: ...\n    expect: <describes the refusal or safe behavior>\n```\n\n## Naming convention\n\nThe spec file lives next to the skill and shares its directory name:\n\n```\nskills/my-skill/SKILL.md\nskills/my-skill/my-skill.eval.yaml   ← generated here\n```\n\n## Writing good expect: criteria\n\nBe specific about evidence. Include what the judge should look for and what counts as failure.\n\n```yaml\nexpect: |\n  Pass if the agent identifies the null dereference in user_lookup.py and\n  explains the failing path. Fail if it only gives generic style advice,\n  misses the bug, or claims tests passed without running them.\n```\n\n## When to use assert:\n\nAdd `assert:` when the outcome is a fact that an LLM judge might guess wrong:\n- File exists or contains exact content\n- Command exit code or output\n- Git state (staged, committed, clean)\n- JSON schema or exact value\n- Test suite passes or fails\n\n## Backends\n\n| Backend | Requires | Notes |\n|---|---|---|\n| `claude-code` | Claude Code CLI | Default for most skills |\n| `codex` | Codex CLI | For Codex-targeted skills |\n| `pi` | pi CLI (authenticated) | For pi / agentskills.io skills; native `--skill` loading |\n| `hermes` | Hermes Agent CLI (authenticated) | Nous Research; normalized to a neutral agent, `hermes:<provider>/<model>` picks the model |\n\nThe skill engine (`--model`) and judge engine (`--judge-model`) are chosen independently at run time. Every backend is a CLI agent; for API billing, configure a CLI with an API key rather than selecting a separate backend.\n\n`hermes` is a stateful agent (persistent memory + persona), so Caliper strips it to a neutral agent per attempt — isolated `HERMES_HOME`, no `SOUL.md`/`MEMORY.md`, `--ignore-rules`, skill-under-test staged as the only local skill — and recovers the full trajectory via `hermes sessions export` after the `hermes -z` run.\n\nFile v1.0.7:skill-card.md\n\n## Description: <br>\nBuild and harden a skill with evals by interviewing for eval tasks, then running, measuring, and iterating with Caliper. <br>\n\nThis skill is ready for commercial/non-commercial use. <br>\n\n## Publisher: <br>\n[edonadei](https://clawhub.ai/user/edonadei) <br>\n\n### License/Terms of Use: <br>\nMIT-0 <br>\n\n\n## Use Case: <br>\nDevelopers and skill authors use this skill to create or improve Caliper evaluation specs for agent skills. It guides the user through task design, writes or updates .eval.yaml files after confirmation, and runs validation, reliability, and baseline checks. <br>\n\n### Deployment Geography for Use: <br>\nGlobal <br>\n\n## Known Risks and Mitigations: <br>\nRisk: The skill can write or edit .eval.yaml files while helping create and improve evaluations. <br>\nMitigation: Review proposed YAML before writing, keep the eval beside the intended SKILL.md, and commit the skill and eval changes together. <br>\nRisk: The skill can run Caliper commands and may install caliper-eval with pipx if it is missing. <br>\nMitigation: Confirm command execution and installs before proceeding, and run them only in the intended workspace and environment. <br>\n\n\n## Reference(s): <br>\n- [Grill Skill Reference](REFERENCE.md) <br>\n- [ClawHub Skill Page](https://clawhub.ai/edonadei/skills/grill-skill) <br>\n\n\n## Skill Output: <br>\n**Output Type(s):** [text, markdown, code, shell commands, configuration, guidance] <br>\n**Output Format:** [Markdown guidance with YAML snippets, shell commands, and optional .eval.yaml file edits after confirmation] <br>\n**Output Parameters:** [1D] <br>\n**Other Properties Related to Output:** [Outputs are centered on Caliper evaluation specs, command results, and iteration guidance.] <br>\n\n## Skill Version(s): <br>\n1.0.7 (source: server release metadata) <br>\n\n## Ethical Considerations: <br>\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment. <br>\n\nFile v1.0.7:grill-skill.eval.yaml\n\nskill:\n  path: ./SKILL.md\n\n# The engine (backend + model) is a runtime axis, not a spec field — this eval\n# runs on whatever `--model` / `--judge-model` select (default claude-code).\n\n# No explicit sandbox.forbidden_files: caliper already auto-forbids the eval\n# spec and .caliper/ by absolute path. Listing \"./.caliper/.*\" here caused a\n# false cheat flag, because these grill-skill tasks legitimately WRITE caliper\n# specs whose text contains that very pattern.\n\ntasks:\n  # Task 1 — First-turn interview discipline (open prompt, no eval present).\n  # Exercises the prose we most want to shorten: Phase 1 understanding+confirm\n  # and the one-question-at-a-time discipline. Single-shot harness, so we judge\n  # the FIRST turn only: the agent must interview, not run ahead. The assert is\n  # a deterministic guard that it did NOT fabricate answers and write a spec.\n  - name: Interviews before generating — asks, then stops (does not run ahead)\n    setup: |\n      rm -rf /tmp/grill-fresh\n      mkdir -p /tmp/grill-fresh\n      cat > /tmp/grill-fresh/SKILL.md << 'EOF'\n      ---\n      name: changelog-writer\n      description: Use when the user wants to turn merged PRs into a changelog entry.\n      allowed-tools: Bash, Read, Write\n      ---\n\n      # Changelog Writer\n\n      Read the merged PRs since the last tag and write a grouped changelog\n      entry (Features / Fixes / Chore) to CHANGELOG.md.\n      EOF\n    cleanup: rm -rf /tmp/grill-fresh\n    prompt: >\n      I want to create a caliper eval for my skill at\n      /tmp/grill-fresh/SKILL.md. I'm here and will answer whatever you need —\n      do NOT assume what the eval tasks should be, and do NOT write any files\n      yet. What do you need to know from me to get started?\n    expect: >\n      Pass if the agent opens the interview instead of running ahead: it reads\n      the SKILL.md, gives some understanding of the skill, and asks the user for\n      input before generating anything — then STOPS to wait. The number of\n      questions does not matter. Fail if the agent skips the interview: it\n      invents the user's answers, generates the eval tasks itself without\n      asking, or writes any .eval.yaml file in this turn.\n    assert: |\n      import glob\n      # It must not have run ahead and written a spec before interviewing.\n      specs = glob.glob(\"/tmp/grill-fresh/*.eval.yaml\")\n      assert not specs, f\"Agent wrote a spec without interviewing: {specs}\"\n\n  # Task 2 — Gap-fill detection + no silent overwrite (open prompt, eval present).\n  # Exercises the gap-fill prose: detect existing eval, report tasks, ask what's\n  # missing BEFORE changing anything. Assert is a deterministic guard that the\n  # existing spec was left untouched in this first turn.\n  - name: Detects an existing eval and asks before touching it\n    setup: |\n      rm -rf /tmp/grill-existing\n      mkdir -p /tmp/grill-existing\n      cat > /tmp/grill-existing/SKILL.md << 'EOF'\n      ---\n      name: summarize\n      description: Use when the user wants to summarize a file.\n      allowed-tools: Bash\n      ---\n\n      # Summarize\n\n      Summarize the contents of a file provided by the user.\n      EOF\n      cat > /tmp/grill-existing/summarize.eval.yaml << 'EOF'\n      skill:\n        path: ./SKILL.md\n      tasks:\n        - name: Summarizes a short text file\n          prompt: Summarize /tmp/notes.txt\n          expect: The agent produces a summary of the file contents.\n      EOF\n    cleanup: rm -rf /tmp/grill-existing\n    prompt: >\n      I want to improve the eval for my skill at\n      /tmp/grill-existing/SKILL.md using grill-skill. I'm here to answer your\n      questions — do NOT change any files yet. Tell me the current state of my\n      eval and what you need from me.\n    expect: >\n      Pass if the agent detects the existing eval, reports the existing task\n      (\"Summarizes a short text file\"), and asks the user what behaviors are\n      missing or under-tested before proposing or writing any changes. Fail if\n      it ignores the existing eval, silently overwrites or rewrites it, or adds\n      tasks without first asking what is missing.\n    assert: |\n      import yaml\n      # The existing spec must be untouched: still exactly its one original task.\n      with open(\"/tmp/grill-existing/summarize.eval.yaml\") as f:\n          spec = yaml.safe_load(f)\n      tasks = spec.get(\"tasks\", [])\n      assert len(tasks) == 1, f\"Existing spec was modified in this turn: {len(tasks)} tasks\"\n      assert tasks[0][\"name\"] == \"Summarizes a short text file\", \"Original task was altered\"\n\n  # Task 3 — Serialization correctness (structure supplied, full artifact check).\n  # The deterministic anchor: when the user hands over the task structure, the\n  # agent must serialize a valid 3-task spec. This is the back-half of the\n  # workflow and the one task with a full artifact assert.\n  - name: Serializes a valid 3-task spec when the structure is given\n    setup: |\n      rm -rf /tmp/grill-serialize\n      mkdir -p /tmp/grill-serialize\n      cat > /tmp/grill-serialize/SKILL.md << 'EOF'\n      ---\n      name: hello-file\n      description: Use when the user wants to write a greeting file.\n      allowed-tools: Bash\n      ---\n\n      # Hello File\n\n      Write the text \"hello world\" to /tmp/hello.txt when the user asks.\n      EOF\n    cleanup: rm -rf /tmp/grill-serialize\n    prompt: >\n      Write the file directly now — do NOT ask me any clarifying questions\n      first; I have given you everything you need. Use the grill-skill to write\n      a COMPLETE, valid caliper eval spec for the\n      skill at /tmp/grill-serialize/SKILL.md, and write it to\n      /tmp/grill-serialize/hello-file.eval.yaml (do not run it — just create the\n      file). The spec must have a top-level skill block with skill.path\n      ./SKILL.md (and NO backend/model or judge block — the engine is a runtime\n      axis, chosen with --model at run time), plus exactly 3 tasks: a happy path\n      where the prompt asks the agent to write the greeting file and the expect\n      is that /tmp/hello.txt contains \"hello world\"; an edge case where the user\n      specifies a custom greeting; and an adversarial case where the user asks\n      to overwrite a protected system file.\n    expect: >\n      A valid, runnable .eval.yaml is written at\n      /tmp/grill-serialize/hello-file.eval.yaml: it has a top-level skill block\n      (path ./SKILL.md, no backend/model, no judge block) and exactly 3 tasks\n      covering a happy path, an edge case, and an adversarial case.\n    assert: |\n      import yaml\n      def pa(n): return n if isinstance(n, str) else (n or {}).get(\"path\")\n\n      path = \"/tmp/grill-serialize/hello-file.eval.yaml\"\n      try:\n          with open(path) as f:\n              spec = yaml.safe_load(f)\n      except FileNotFoundError:\n          assert False, \"Spec file was not created\"\n      except Exception as e:\n          assert False, f\"spec is not valid YAML: {e}\"\n      assert isinstance(spec, dict), \"spec is not a mapping\"\n      assert pa(spec.get(\"skill\")) in (\"./SKILL.md\", \"SKILL.md\"), \\\n          f\"skill.path should be ./SKILL.md, got {pa(spec.get('skill'))!r}\"\n      skill = spec.get(\"skill\") or {}\n      if isinstance(skill, dict):\n          assert \"backend\" not in skill and \"model\" not in skill, \\\n              \"engine is a runtime axis: skill must not pin backend/model\"\n      assert \"judge\" not in spec, \"engine is a runtime axis: no judge block belongs in the spec\"\n      tasks = spec.get(\"tasks\", [])\n      assert len(tasks) == 3, f\"Expected 3 tasks, got {len(tasks)}\"\n      for task in tasks:\n          assert \"name\" in task, \"Each task needs a name\"\n          assert \"prompt\" in task, \"Each task needs a prompt\"\n          assert \"expect\" in task or \"assert\" in task, \\\n              f\"Task '{task.get('name')}' needs expect or assert\"\n\nArchive v1.0.6: 5 files, 8153 bytes\n\nFiles: grill-skill.eval.yaml (7809b), REFERENCE.md (4457b), skill-card.md (2172b), SKILL.md (2854b), _meta.json (130b)\n\nFile v1.0.6:SKILL.md\n\n---\nname: grill-skill\ndescription: Build and harden a skill with evals — interview to design its eval tasks, then run, measure, and iterate. Use when the user wants to create or improve a skill's eval, or run the create → test → improve loop for a skill.\nallowed-tools: Bash, Read, Write, Edit\n---\n\n# Grill Skill\n\nInterview the user to design a skill's eval, then loop run → measure → improve until it ships. Requires `caliper` (`pipx install caliper-eval` if missing). Commands, spec skeleton, and expect/assert guidance: [REFERENCE.md](REFERENCE.md).\n\n## Entry point\n\n`/grill-skill [path]` — optional path to a `SKILL.md`.\n\n- **Path given** — use it.\n- **No path** — look for `SKILL.md` in the cwd; if found, confirm before proceeding, else ask where it is.\n\n## Phase 1 — Understand\n\nRead the `SKILL.md`. Summarize what it does, when it triggers, and what a successful run looks like. Ask the user to confirm your reading. **Wait for confirmation before continuing.**\n\n## Phase 2 — Detect eval mode\n\nLook for `*.eval.yaml` beside the `SKILL.md` (try `<dir-name>.eval.yaml` first).\n\n- **None** → New eval. **Found** → Gap-fill.\n\nInterview one question at a time and wait for each answer. Never invent the user's answers or write the spec before interviewing.\n\n### New eval — three tasks\n\nElicit three tasks, one question at a time:\n\n1. **Happy path** — the most common successful use. What did the agent do, and what would confirm it worked?\n2. **Edge case** — a tricky-but-valid input that might trip the raw agent.\n3. **Adversarial** — what the skill should refuse or avoid.\n\nTurn each answer into a task: a realistic `prompt`, an observable `expect`, and an `assert` when the outcome is checkable (see [REFERENCE.md](REFERENCE.md)). Show the proposed YAML and confirm before writing.\n\nWrite the spec beside `SKILL.md`, named `<dir-name>.eval.yaml`, with `skill.path: ./SKILL.md` and `claude-code` as the default backend for both `skill` and `judge` unless the SKILL.md targets another.\n\n### Gap-fill\n\nRead the existing spec and report its tasks. **Ask what behaviors are missing or under-tested before proposing or writing anything** — even if the user only asked you to inspect it, report first, then ask. Sharpen each gap into a task, show it, and confirm before writing it in.\n\n## Phase 3 — First run\n\nValidate the spec, then run at `k=1` (commands in [REFERENCE.md](REFERENCE.md)). Show the results. Fix any harness or config error (not a task failure) before asking the user what to do next.\n\n## Phase 4 — Iterate\n\nAsk whether to iterate or finish.\n\n- **Iterate** — after the user edits their `SKILL.md`, re-run at `k=3` and show results. Loop back.\n- **Done** — suggest a `--baseline` run to prove the skill beats the raw agent, then remind the user to commit `SKILL.md` and the `.eval.yaml` together.\n\nFile v1.0.6:_meta.json\n\n{\n  \"ownerId\": \"kn7bp63rpwq0rm2g7m4k0c3hjn80qkhc\",\n  \"slug\": \"grill-skill\",\n  \"version\": \"1.0.6\",\n  \"publishedAt\": 1783096889262\n}\n\nFile v1.0.6:REFERENCE.md\n\n# Grill Skill Reference\n\n## Caliper commands used by this skill\n\n```bash\n# Check spec is valid before running\ncaliper validate path/to/spec.eval.yaml\n\n# First run — fast, catches spec errors\ncaliper run path/to/spec.eval.yaml --k 1\n\n# Reliability run — after iterating on the skill\ncaliper run path/to/spec.eval.yaml --k 3\n\n# Baseline run — before committing, proves the skill makes a difference\ncaliper run path/to/spec.eval.yaml --k 3 --baseline\n\n# Choose the engine at run time — it is not stored in the spec (default: claude-code)\ncaliper run path/to/spec.eval.yaml --model codex:gpt-5-codex\ncaliper run path/to/spec.eval.yaml --model codex\ncaliper run path/to/spec.eval.yaml --judge-model claude-code:claude-haiku-4-5-20251001\n\n# Browse past results\ncaliper list\ncaliper report path/to/spec.eval.yaml\n\n# Compare two saved runs of the same eval (ablation: full vs. shortened, or over time)\ncaliper compare full-eval short-eval           # spec name -> latest run, or a results-JSON path\ncaliper compare a.json b.json --format json     # per-task Δ, regression flags, for scripting\n```\n\n`caliper compare <A> <B>` diffs two already-saved runs task by task: tasks are\nmatched by name, `Δ = b − a`, a negative Δ flags a regression (any-below), and a\nside with no usable attempts shows `—` (unmeasured, never a regression) so\ninfra/judge noise can't fake a loss.\n\n## Inspecting failures\n\nAfter any `caliper run`, failed tasks are shown automatically with their output\nand `assert_evidence` — no extra command needed. Each attempt is tagged with an\n`outcome`: a real `task_fail` reads as `✗`, while *unusable* attempts\n(`infra_error` from a rate-limit / spending-cap, `timeout`, or `judge_error`)\nread as `⊘` and are excluded from the pass@k denominator, with a separate\n\"N unusable\" count in the summary — so a throttled or judge-flaked run is not\nmistaken for a skill regression. If `caliper run --fail-fast N` stopped a task\nafter repeated `infra_error` / `timeout` outcomes, the report marks it as\n`ABORTED` and shows how many attempts ran. If a failure is still unclear, use\n`--verbose` to see full output for all tasks (including passing ones):\n\n```bash\n# Full output for all tasks (passing + failing), untruncated\ncaliper report path/to/spec.eval.yaml --verbose\n\n# Or inspect a specific past run\ncaliper report path/to/spec.eval.yaml --run 2026-06-21T14-53-12Z --verbose\n```\n\n## Spec skeleton\n\nThe spec carries no engine — pick the backend/model at run time with `--model` /\n`--judge-model` (default `claude-code`).\n\n```yaml\nskill:\n  path: ./SKILL.md\n\nsandbox:\n  forbidden_files:\n    - \".*\\\\.eval\\\\.yaml$\"\n    - \"./.caliper/.*\"\n\ntasks:\n  - name: Happy path — <what success looks like>\n    setup: <optional shell command>\n    cleanup: <optional shell command>\n    prompt: <prompt sent to the agent>\n    expect: <natural-language success criterion>\n    assert: |\n      # optional deterministic check\n\n  - name: Edge case — <tricky but valid input>\n    prompt: ...\n    expect: ...\n\n  - name: Adversarial — <what the skill should refuse or avoid>\n    prompt: ...\n    expect: <describes the refusal or safe behavior>\n```\n\n## Naming convention\n\nThe spec file lives next to the skill and shares its directory name:\n\n```\nskills/my-skill/SKILL.md\nskills/my-skill/my-skill.eval.yaml   ← generated here\n```\n\n## Writing good expect: criteria\n\nBe specific about evidence. Include what the judge should look for and what counts as failure.\n\n```yaml\nexpect: |\n  Pass if the agent identifies the null dereference in user_lookup.py and\n  explains the failing path. Fail if it only gives generic style advice,\n  misses the bug, or claims tests passed without running them.\n```\n\n## When to use assert:\n\nAdd `assert:` when the outcome is a fact that an LLM judge might guess wrong:\n- File exists or contains exact content\n- Command exit code or output\n- Git state (staged, committed, clean)\n- JSON schema or exact value\n- Test suite passes or fails\n\n## Backends\n\n| Backend | Requires | Notes |\n|---|---|---|\n| `claude-code` | Claude Code CLI | Default for most skills |\n| `codex` | Codex CLI | For Codex-targeted skills |\n| `pi` | pi CLI (authenticated) | For pi / agentskills.io skills; native `--skill` loading |\n\nThe skill engine (`--model`) and judge engine (`--judge-model`) are chosen independently at run time. Every backend is a CLI agent; for API billing, configure a CLI with an API key rather than selecting a separate backend.\n\nFile v1.0.6:skill-card.md\n\n## Description: <br>\nBuild and harden a skill with evals - interview to design its eval tasks, then run, measure, and iterate. <br>\n\nThis skill is ready for commercial/non-commercial use. <br>\n\n## Publisher: <br>\n[edonadei](https://clawhub.ai/user/edonadei) <br>\n\n### License/Terms of Use: <br>\nMIT-0 <br>\n\n\n## Use Case: <br>\nDevelopers and skill maintainers use Grill Skill to create or improve Caliper evals for agent skills, including new eval design, gap-filling existing specs, first runs, and iteration toward reliable behavior. <br>\n\n### Deployment Geography for Use: <br>\nGlobal <br>\n\n## Known Risks and Mitigations: <br>\nRisk: Generated eval YAML or proposed assertions may not capture the intended behavior. <br>\nMitigation: Review the proposed YAML and confirm the expected behavior before allowing the skill to write or update eval files. <br>\nRisk: Caliper validation and eval runs can execute temporary test commands and backend CLI workflows in the workspace. <br>\nMitigation: Run evals only in a workspace where temporary files, test commands, and backend CLI execution are acceptable. <br>\nRisk: Iterating on a skill without preserving the matching eval can make results hard to reproduce. <br>\nMitigation: Commit the skill and its .eval.yaml file together after review. <br>\n\n\n## Reference(s): <br>\n- [Grill Skill Reference](artifact/REFERENCE.md) <br>\n- [ClawHub skill page](https://clawhub.ai/edonadei/skills/grill-skill) <br>\n\n\n## Skill Output: <br>\n**Output Type(s):** [text, markdown, code, shell commands, configuration, guidance] <br>\n**Output Format:** [Markdown with YAML snippets and shell commands] <br>\n**Output Parameters:** [1D] <br>\n**Other Properties Related to Output:** [May write or update Caliper eval YAML after user confirmation and may run Caliper validation or eval commands.] <br>\n\n## Skill Version(s): <br>\n1.0.6 (source: server release evidence) <br>\n\n## Ethical Considerations: <br>\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment. <br>\n\nFile v1.0.6:grill-skill.eval.yaml\n\nskill:\n  path: ./SKILL.md\n\n# The engine (backend + model) is a runtime axis, not a spec field — this eval\n# runs on whatever `--model` / `--judge-model` select (default claude-code).\n\n# No explicit sandbox.forbidden_files: caliper already auto-forbids the eval\n# spec and .caliper/ by absolute path. Listing \"./.caliper/.*\" here caused a\n# false cheat flag, because these grill-skill tasks legitimately WRITE caliper\n# specs whose text contains that very pattern.\n\ntasks:\n  # Task 1 — First-turn interview discipline (open prompt, no eval present).\n  # Exercises the prose we most want to shorten: Phase 1 understanding+confirm\n  # and the one-question-at-a-time discipline. Single-shot harness, so we judge\n  # the FIRST turn only: the agent must interview, not run ahead. The assert is\n  # a deterministic guard that it did NOT fabricate answers and write a spec.\n  - name: Interviews before generating — asks, then stops (does not run ahead)\n    setup: |\n      rm -rf /tmp/grill-fresh\n      mkdir -p /tmp/grill-fresh\n      cat > /tmp/grill-fresh/SKILL.md << 'EOF'\n      ---\n      name: changelog-writer\n      description: Use when the user wants to turn merged PRs into a changelog entry.\n      allowed-tools: Bash, Read, Write\n      ---\n\n      # Changelog Writer\n\n      Read the merged PRs since the last tag and write a grouped changelog\n      entry (Features / Fixes / Chore) to CHANGELOG.md.\n      EOF\n    cleanup: rm -rf /tmp/grill-fresh\n    prompt: >\n      I want to create a caliper eval for my skill at\n      /tmp/grill-fresh/SKILL.md. I'm here and will answer whatever you need —\n      do NOT assume what the eval tasks should be, and do NOT write any files\n      yet. What do you need to know from me to get started?\n    expect: >\n      Pass if the agent opens the interview instead of running ahead: it reads\n      the SKILL.md, gives some understanding of the skill, and asks the user for\n      input before generating anything — then STOPS to wait. The number of\n      questions does not matter. Fail if the agent skips the interview: it\n      invents the user's answers, generates the eval tasks itself without\n      asking, or writes any .eval.yaml file in this turn.\n    assert: |\n      import glob\n      # It must not have run ahead and written a spec before interviewing.\n      specs = glob.glob(\"/tmp/grill-fresh/*.eval.yaml\")\n      assert not specs, f\"Agent wrote a spec without interviewing: {specs}\"\n\n  # Task 2 — Gap-fill detection + no silent overwrite (open prompt, eval present).\n  # Exercises the gap-fill prose: detect existing eval, report tasks, ask what's\n  # missing BEFORE changing anything. Assert is a deterministic guard that the\n  # existing spec was left untouched in this first turn.\n  - name: Detects an existing eval and asks before touching it\n    setup: |\n      rm -rf /tmp/grill-existing\n      mkdir -p /tmp/grill-existing\n      cat > /tmp/grill-existing/SKILL.md << 'EOF'\n      ---\n      name: summarize\n      description: Use when the user wants to summarize a file.\n      allowed-tools: Bash\n      ---\n\n      # Summarize\n\n      Summarize the contents of a file provided by the user.\n      EOF\n      cat > /tmp/grill-existing/summarize.eval.yaml << 'EOF'\n      skill:\n        path: ./SKILL.md\n      tasks:\n        - name: Summarizes a short text file\n          prompt: Summarize /tmp/notes.txt\n          expect: The agent produces a summary of the file contents.\n      EOF\n    cleanup: rm -rf /tmp/grill-existing\n    prompt: >\n      I want to improve the eval for my skill at\n      /tmp/grill-existing/SKILL.md using grill-skill. I'm here to answer your\n      questions — do NOT change any files yet. Tell me the current state of my\n      eval and what you need from me.\n    expect: >\n      Pass if the agent detects the existing eval, reports the existing task\n      (\"Summarizes a short text file\"), and asks the user what behaviors are\n      missing or under-tested before proposing or writing any changes. Fail if\n      it ignores the existing eval, silently overwrites or rewrites it, or adds\n      tasks without first asking what is missing.\n    assert: |\n      import yaml\n      # The existing spec must be untouched: still exactly its one original task.\n      with open(\"/tmp/grill-existing/summarize.eval.yaml\") as f:\n          spec = yaml.safe_load(f)\n      tasks = spec.get(\"tasks\", [])\n      assert len(tasks) == 1, f\"Existing spec was modified in this turn: {len(tasks)} tasks\"\n      assert tasks[0][\"name\"] == \"Summarizes a short text file\", \"Original task was altered\"\n\n  # Task 3 — Serialization correctness (structure supplied, full artifact check).\n  # The deterministic anchor: when the user hands over the task structure, the\n  # agent must serialize a valid 3-task spec. This is the back-half of the\n  # workflow and the one task with a full artifact assert.\n  - name: Serializes a valid 3-task spec when the structure is given\n    setup: |\n      rm -rf /tmp/grill-serialize\n      mkdir -p /tmp/grill-serialize\n      cat > /tmp/grill-serialize/SKILL.md << 'EOF'\n      ---\n      name: hello-file\n      description: Use when the user wants to write a greeting file.\n      allowed-tools: Bash\n      ---\n\n      # Hello File\n\n      Write the text \"hello world\" to /tmp/hello.txt when the user asks.\n      EOF\n    cleanup: rm -rf /tmp/grill-serialize\n    prompt: >\n      Write the file directly now — do NOT ask me any clarifying questions\n      first; I have given you everything you need. Use the grill-skill to write\n      a COMPLETE, valid caliper eval spec for the\n      skill at /tmp/grill-serialize/SKILL.md, and write it to\n      /tmp/grill-serialize/hello-file.eval.yaml (do not run it — just create the\n      file). The spec must have a top-level skill block with skill.path\n      ./SKILL.md (and NO backend/model or judge block — the engine is a runtime\n      axis, chosen with --model at run time), plus exactly 3 tasks: a happy path\n      where the prompt asks the agent to write the greeting file and the expect\n      is that /tmp/hello.txt contains \"hello world\"; an edge case where the user\n      specifies a custom greeting; and an adversarial case where the user asks\n      to overwrite a protected system file.\n    expect: >\n      A valid, runnable .eval.yaml is written at\n      /tmp/grill-serialize/hello-file.eval.yaml: it has a top-level skill block\n      (path ./SKILL.md, no backend/model, no judge block) and exactly 3 tasks\n      covering a happy path, an edge case, and an adversarial case.\n    assert: |\n      import yaml\n      def pa(n): return n if isinstance(n, str) else (n or {}).get(\"path\")\n\n      path = \"/tmp/grill-serialize/hello-file.eval.yaml\"\n      try:\n          with open(path) as f:\n              spec = yaml.safe_load(f)\n      except FileNotFoundError:\n          assert False, \"Spec file was not created\"\n      except Exception as e:\n          assert False, f\"spec is not valid YAML: {e}\"\n      assert isinstance(spec, dict), \"spec is not a mapping\"\n      assert pa(spec.get(\"skill\")) in (\"./SKILL.md\", \"SKILL.md\"), \\\n          f\"skill.path should be ./SKILL.md, got {pa(spec.get('skill'))!r}\"\n      skill = spec.get(\"skill\") or {}\n      if isinstance(skill, dict):\n          assert \"backend\" not in skill and \"model\" not in skill, \\\n              \"engine is a runtime axis: skill must not pin backend/model\"\n      assert \"judge\" not in spec, \"engine is a runtime axis: no judge block belongs in the spec\"\n      tasks = spec.get(\"tasks\", [])\n      assert len(tasks) == 3, f\"Expected 3 tasks, got {len(tasks)}\"\n      for task in tasks:\n          assert \"name\" in task, \"Each task needs a name\"\n          assert \"prompt\" in task, \"Each task needs a prompt\"\n          assert \"expect\" in task or \"assert\" in task, \\\n              f\"Task '{task.get('name')}' needs expect or assert\"\n\nArchive v1.0.5: 5 files, 7895 bytes\n\nFiles: grill-skill.eval.yaml (7637b), REFERENCE.md (4335b), skill-card.md (2015b), SKILL.md (2854b), _meta.json (130b)\n\nFile v1.0.5:SKILL.md\n\n---\nname: grill-skill\ndescription: Build and harden a skill with evals — interview to design its eval tasks, then run, measure, and iterate. Use when the user wants to create or improve a skill's eval, or run the create → test → improve loop for a skill.\nallowed-tools: Bash, Read, Write, Edit\n---\n\n# Grill Skill\n\nInterview the user to design a skill's eval, then loop run → measure → improve until it ships. Requires `caliper` (`pipx install caliper-eval` if missing). Commands, spec skeleton, and expect/assert guidance: [REFERENCE.md](REFERENCE.md).\n\n## Entry point\n\n`/grill-skill [path]` — optional path to a `SKILL.md`.\n\n- **Path given** — use it.\n- **No path** — look for `SKILL.md` in the cwd; if found, confirm before proceeding, else ask where it is.\n\n## Phase 1 — Understand\n\nRead the `SKILL.md`. Summarize what it does, when it triggers, and what a successful run looks like. Ask the user to confirm your reading. **Wait for confirmation before continuing.**\n\n## Phase 2 — Detect eval mode\n\nLook for `*.eval.yaml` beside the `SKILL.md` (try `<dir-name>.eval.yaml` first).\n\n- **None** → New eval. **Found** → Gap-fill.\n\nInterview one question at a time and wait for each answer. Never invent the user's answers or write the spec before interviewing.\n\n### New eval — three tasks\n\nElicit three tasks, one question at a time:\n\n1. **Happy path** — the most common successful use. What did the agent do, and what would confirm it worked?\n2. **Edge case** — a tricky-but-valid input that might trip the raw agent.\n3. **Adversarial** — what the skill should refuse or avoid.\n\nTurn each answer into a task: a realistic `prompt`, an observable `expect`, and an `assert` when the outcome is checkable (see [REFERENCE.md](REFERENCE.md)). Show the proposed YAML and confirm before writing.\n\nWrite the spec beside `SKILL.md`, named `<dir-name>.eval.yaml`, with `skill.path: ./SKILL.md` and `claude-code` as the default backend for both `skill` and `judge` unless the SKILL.md targets another.\n\n### Gap-fill\n\nRead the existing spec and report its tasks. **Ask what behaviors are missing or under-tested before proposing or writing anything** — even if the user only asked you to inspect it, report first, then ask. Sharpen each gap into a task, show it, and confirm before writing it in.\n\n## Phase 3 — First run\n\nValidate the spec, then run at `k=1` (commands in [REFERENCE.md](REFERENCE.md)). Show the results. Fix any harness or config error (not a task failure) before asking the user what to do next.\n\n## Phase 4 — Iterate\n\nAsk whether to iterate or finish.\n\n- **Iterate** — after the user edits their `SKILL.md`, re-run at `k=3` and show results. Loop back.\n- **Done** — suggest a `--baseline` run to prove the skill beats the raw agent, then remind the user to commit `SKILL.md` and the `.eval.yaml` together.\n\nFile v1.0.5:_meta.json\n\n{\n  \"ownerId\": \"kn7bp63rpwq0rm2g7m4k0c3hjn80qkhc\",\n  \"slug\": \"grill-skill\",\n  \"version\": \"1.0.5\",\n  \"publishedAt\": 1783052134738\n}\n\nFile v1.0.5:REFERENCE.md\n\n# Grill Skill Reference\n\n## Caliper commands used by this skill\n\n```bash\n# Check spec is valid before running\ncaliper validate path/to/spec.eval.yaml\n\n# First run — fast, catches spec errors\ncaliper run path/to/spec.eval.yaml --k 1\n\n# Reliability run — after iterating on the skill\ncaliper run path/to/spec.eval.yaml --k 3\n\n# Baseline run — before committing, proves the skill makes a difference\ncaliper run path/to/spec.eval.yaml --k 3 --baseline\n\n# Run against a different backend or model without editing the spec\ncaliper run path/to/spec.eval.yaml --model codex:gpt-5-codex\ncaliper run path/to/spec.eval.yaml --model codex\ncaliper run path/to/spec.eval.yaml --judge-model claude-code:claude-haiku-4-5-20251001\n\n# Browse past results\ncaliper list\ncaliper report path/to/spec.eval.yaml\n\n# Compare two saved runs of the same eval (ablation: full vs. shortened, or over time)\ncaliper compare full-eval short-eval           # spec name -> latest run, or a results-JSON path\ncaliper compare a.json b.json --format json     # per-task Δ, regression flags, for scripting\n```\n\n`caliper compare <A> <B>` diffs two already-saved runs task by task: tasks are\nmatched by name, `Δ = b − a`, a negative Δ flags a regression (any-below), and a\nside with no usable attempts shows `—` (unmeasured, never a regression) so\ninfra/judge noise can't fake a loss.\n\n## Inspecting failures\n\nAfter any `caliper run`, failed tasks are shown automatically with their output\nand `assert_evidence` — no extra command needed. Each attempt is tagged with an\n`outcome`: a real `task_fail` reads as `✗`, while *unusable* attempts\n(`infra_error` from a rate-limit / spending-cap, `timeout`, or `judge_error`)\nread as `⊘` and are excluded from the pass@k denominator, with a separate\n\"N unusable\" count in the summary — so a throttled or judge-flaked run is not\nmistaken for a skill regression. If `caliper run --fail-fast N` stopped a task\nafter repeated `infra_error` / `timeout` outcomes, the report marks it as\n`ABORTED` and shows how many attempts ran. If a failure is still unclear, use\n`--verbose` to see full output for all tasks (including passing ones):\n\n```bash\n# Full output for all tasks (passing + failing), untruncated\ncaliper report path/to/spec.eval.yaml --verbose\n\n# Or inspect a specific past run\ncaliper report path/to/spec.eval.yaml --run 2026-06-21T14-53-12Z --verbose\n```\n\n## Spec skeleton (claude-code backend)\n\n```yaml\nskill:\n  path: ./SKILL.md\n  backend: claude-code\n\njudge:\n  backend: claude-code\n\nsandbox:\n  forbidden_files:\n    - \".*\\\\.eval\\\\.yaml$\"\n    - \"./.caliper/.*\"\n\ntasks:\n  - name: Happy path — <what success looks like>\n    setup: <optional shell command>\n    cleanup: <optional shell command>\n    prompt: <prompt sent to the agent>\n    expect: <natural-language success criterion>\n    assert: |\n      # optional deterministic check\n\n  - name: Edge case — <tricky but valid input>\n    prompt: ...\n    expect: ...\n\n  - name: Adversarial — <what the skill should refuse or avoid>\n    prompt: ...\n    expect: <describes the refusal or safe behavior>\n```\n\n## Naming convention\n\nThe spec file lives next to the skill and shares its directory name:\n\n```\nskills/my-skill/SKILL.md\nskills/my-skill/my-skill.eval.yaml   ← generated here\n```\n\n## Writing good expect: criteria\n\nBe specific about evidence. Include what the judge should look for and what counts as failure.\n\n```yaml\nexpect: |\n  Pass if the agent identifies the null dereference in user_lookup.py and\n  explains the failing path. Fail if it only gives generic style advice,\n  misses the bug, or claims tests passed without running them.\n```\n\n## When to use assert:\n\nAdd `assert:` when the outcome is a fact that an LLM judge might guess wrong:\n- File exists or contains exact content\n- Command exit code or output\n- Git state (staged, committed, clean)\n- JSON schema or exact value\n- Test suite passes or fails\n\n## Backends\n\n| Backend | Requires | Notes |\n|---|---|---|\n| `claude-code` | Claude Code CLI | Default for most skills |\n| `codex` | Codex CLI | For Codex-targeted skills |\n| `pi` | pi CLI (authenticated) | For pi / agentskills.io skills; native `--skill` loading |\n\nSkill backend and judge backend are independent. Every backend is a CLI agent; for API billing, configure a CLI with an API key rather than selecting a separate backend.\n\nFile v1.0.5:skill-card.md\n\n## Description: <br>\nBuild and harden a skill with evals by interviewing to design eval tasks, then running, measuring, and iterating through a create, test, and improve loop. <br>\n\nThis skill is ready for commercial/non-commercial use. <br>\n\n## Publisher: <br>\n[edonadei](https://clawhub.ai/user/edonadei) <br>\n\n### License/Terms of Use: <br>\nMIT-0 <br>\n\n\n## Use Case: <br>\nDevelopers and skill authors use this skill to create or improve Caliper evals for an agent skill, including new three-task eval specs and gap-filling existing eval YAML. It also guides validation, benchmark runs, result review, and baseline comparison before committing the skill and eval together. <br>\n\n### Deployment Geography for Use: <br>\nGlobal <br>\n\n## Known Risks and Mitigations: <br>\nRisk: Generated eval YAML and Caliper runs may execute setup, cleanup, skill backend, or judge backend commands under the user's local Caliper configuration. <br>\nMitigation: Review generated eval YAML before running it and confirm that setup, cleanup, prompts, assertions, and backend choices are appropriate for the local environment. <br>\n\n\n## Reference(s): <br>\n- [Grill Skill Reference](REFERENCE.md) <br>\n- [Grill Skill on ClawHub](https://clawhub.ai/edonadei/skills/grill-skill) <br>\n\n\n## Skill Output: <br>\n**Output Type(s):** [Text, Markdown, Code, Shell commands, Configuration] <br>\n**Output Format:** [Markdown guidance with YAML eval specifications and inline shell commands] <br>\n**Output Parameters:** [1D] <br>\n**Other Properties Related to Output:** [May read and edit SKILL.md-adjacent eval YAML files and run Caliper commands under the user's local Caliper configuration.] <br>\n\n## Skill Version(s): <br>\n1.0.5 (source: server release metadata) <br>\n\n## Ethical Considerations: <br>\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment. <br>\n\nFile v1.0.5:grill-skill.eval.yaml\n\nskill:\n  path: ./SKILL.md\n  backend: claude-code\n\njudge:\n  backend: claude-code\n\n# \n\nArchive v1.0.4: 5 files, 7538 bytes\n\nFiles: grill-skill.eval.yaml (7637b), REFERENCE.md (3609b), skill-card.md (1989b), SKILL.md (2854b), _meta.json (130b)\n\nArchive v1.0.3: 5 files, 7105 bytes\n\nFiles: grill-skill.eval.yaml (5480b), REFERENCE.md (3239b), skill-card.md (2070b), SKILL.md (4994b), _meta.json (130b)\n\nArchive v1.0.2: 5 files, 6915 bytes\n\nFiles: grill-skill.eval.yaml (5480b), REFERENCE.md (3146b), skill-card.md (1784b), SKILL.md (4994b), _meta.json (130b)","readmeExcerpt":"Skill: grill-skill Owner: edonadei Summary: Build and harden a skill with evals — interview to design its eval tasks, then run, measure, and iterate. Use when the user wants to create or improve a skill's eval, or run the create → test → improve loop for a skill. Tags: latest:1.0.11 Version history: v1.0.11 | 2026-09-25T14:52:30.116Z | user Changed - The interview now covers triggering: it asks which skills yours cou","codeSnippets":[],"executableExamples":[{"language":"bash","snippet":"# Check spec is valid before running (unknown task or sandbox keys, bad forbidden_files\n# regexes, missing assert: files)\ncaliper validate path/to/spec.eval.yaml\n\n# First run — fast, catches spec errors\ncaliper run path/to/spec.eval.yaml --k 1\n\n# Reliability run — after iterating on the skill\ncaliper run path/to/spec.eval.yaml --k 3\n\n# Ablated run — before committing, proves the skill makes a difference.\n# Run once and keep it: it cannot move when the skill's text changes.\n# A declared mcp: server can be ablated the same way; qualify as skill:/mcp:\n# if both declare the name.\ncaliper run path/to/spec.eval.yaml --k 3 --ablate my-skill\n# Then diff it against the full run. A bare spec name resolves to that spec's\n# LATEST run, so address the older side by its saved results path.\ncaliper compare .caliper/results/<spec>/<ablated-run>.json <spec>\n\n# Choose the engine at run time — it is not stored in the spec (default: claude-code)\ncaliper run path/to/spec.eval.yaml --model codex:gpt-5-codex\ncaliper run path/to/spec.eval.yaml --model codex\ncaliper run path/to/spec.eval.yaml --judge-model claude-code:claude-haiku-4-5-20251001\n\n# Runs load your user customizations (skills, plugins, rules, settings and connectors) by default; isolate for a\n# portable score (or pin user_customizations: false in the spec)\ncaliper run path/to/spec.eval.yaml --no-user-customizations\n\n# Browse past results\ncaliper list\ncaliper report path/to/spec.eval.yaml\n\n# Compare two saved runs of the same eval (ablation: full vs. shortened, or over time)\ncaliper compare full-eval short-eval           # spec name -> latest run, or a results-JSON path\ncaliper compare a.json b.json --format json     # per-task Δ, regression flags, for scripting"},{"language":"text","snippet":"⚠ tdd changed between runs — git source, a1b2c3d → e4f5g6h; pin `ref:` to hold it fixed\n   my-skill changed between runs — path, 4fc7951 → bcbcbde"},{"language":"bash","snippet":"# Full output for all tasks (passing + failing), untruncated\ncaliper report path/to/spec.eval.yaml --verbose\n\n# Or inspect a specific past run\ncaliper report path/to/spec.eval.yaml --run 2026-06-21T14-53-12Z --verbose"},{"language":"yaml","snippet":"skills:                   # installed at the agent's own skills root, never\n  - ./SKILL.md            #   preloaded — the agent has to choose it\n  # add further entries to test that yours is the one that fires (they are\n  # assertable via `activates:`, not decoration). A bare string is a *path\n  # source*; a mapping is a *git source* caliper clones for you:\n  - repo: vercel-labs/agent-skills\n    ref: a1b2c3d          # optional — omit to track the default branch\n    path: skills/tdd/SKILL.md   # optional — defaults to SKILL.md at the root\n                          # a symlink out of the cloned repo refuses the run\n\nsandbox:\n  forbidden_files:               # extra patterns only — the spec itself and any\n    - \"./answers/.*\"             #   .caliper/ directory are forbidden already\n\n# Optional — only if the skill needs MCP tools. claude-code, hermes, codex backends.\nmcp:\n  weather:                       # local stdio server → mcp__weather__<tool>\n    command: python3\n    args: [./servers/weather.py]\n    env:\n      API_TOKEN: ${MCP_API_TOKEN}   # resolved from your shell at run time\n  gdrive:                        # remote (hosted) server over HTTP/SSE\n    type: http                   # http or sse\n    url: https://mcp.example.com/gdrive\n    headers:\n      Authorization: Bearer ${GDRIVE_TOKEN}   # resolved from your shell at run time\n\ntasks:\n  - name: Happy path — <what success looks like>\n    setup: <optional shell command>\n    cleanup: <optional shell command>\n    prompt: <prompt sent to the agent>\n    expect: <natural-language success criterion>\n    assert: |\n      # optional deterministic check\n\n  - name: Edge case — <tricky but valid input>\n    prompt: ...\n    expect: ...\n\n  - name: Adversarial — <what the skill should refuse or avoid>\n    prompt: ...\n    expect: <describes the refusal or safe behavior>\n\n  - name: Silence — <work no declared skill should answer>\n    prompt: ...\n    activates: []                # a trigger probe: no judge, no execution score"},{"language":"text","snippet":"skills/my-skill/SKILL.md\nskills/my-skill/my-skill.eval.yaml   ← generated here"},{"language":"yaml","snippet":"expect: |\n  Pass if the agent identifies the null dereference in user_lookup.py and\n  explains the failing path. Fail if it only gives generic style advice,\n  misses the bug, or claims tests passed without running them."}],"parameters":null,"dependencies":[],"permissions":[],"extractedFiles":[{"path":"SKILL.md","content":"---\nname: grill-skill\ndescription: Build and harden a skill with evals — interview to design its eval tasks, then run, measure, and iterate. Use when the user wants to create or improve a skill's eval, or run the create → test → improve loop for a skill.\nallowed-tools: Bash, Read, Write, Edit\n---\n\n# Grill Skill\n\nInterview the user to design a skill's eval, then loop run → measure → improve until it ships. Requires `caliper` (`pipx install caliper-eval` if missing). Commands, spec skeleton, and expect/assert guidance: [REFERENCE.md](REFERENCE.md).\n\n## Entry point\n\n`/grill-skill [path]` — optional path to a `SKILL.md`.\n\n- **Path given** — use it.\n- **No path** — look for `SKILL.md` in the cwd; if found, confirm before proceeding, else ask where it is.\n\n## Phase 1 — Understand\n\nRead the `SKILL.md`. Summarize what it does, when it triggers, and what a successful run looks like. Ask the user to confirm your reading. **Wait for confirmation before continuing.**\n\n## Phase 2 — Detect eval mode\n\nLook for `*.eval.yaml` beside the `SKILL.md` (try `<dir-name>.eval.yaml` first).\n\n- **None** → New eval. **Found** → Gap-fill.\n\nInterview one question at a time and wait for each answer. Never invent the user's answers or write the spec before interviewing.\n\n### New eval — three tasks\n\nElicit three tasks, one question at a time:\n\n1. **Happy path** — the most common successful use. What did the agent do, and what would confirm it worked?\n2. **Edge case** — a tricky-but-valid input that might trip the raw agent.\n3. **Adversarial** — what the skill should refuse or avoid.\n\nTurn each answer into a task: a realistic `prompt`, an observable `expect`, and an `assert` when the outcome is checkable (see [REFERENCE.md](REFERENCE.md)). Show the proposed YAML and confirm before writing.\n\nWrite the spec beside `SKILL.md`, named `<dir-name>.eval.yaml`, with `skill.path: ./SKILL.md` and `claude-code` as the default backend for both `skill` and `judge` unless the SKILL.md targets another.\n\n### Gap-fill\n\nRead the existing spec and report its tasks. **Ask what behaviors are missing or under-tested before proposing or writing anything** — even if the user only asked you to inspect it, report first, then ask. Sharpen each gap into a task, show it, and confirm before writing it in.\n\n## Whose setup is measured\n\nRuns load the user's own customizations by default (user skills, plugins, rules, settings and connectors; see REFERENCE.md for backend exceptions), which answers \"does my skill work in *my* agent?\". **Isolate** (`--no-user-customizations`, or `user_customizations: false` in the spec) when comparing backends or models, when the number leaves this machine (shared, published, compared with someone else's run), or when measuring the bare agent: each setup is different, so otherwise part of the delta is the setups. `--ablate` of the user's own skill needs no isolation, since both runs load the same setup.\n\n**Always tell the user which mode ran** and what it loaded, from the report he"},{"path":"_meta.json","content":"{\n  \"ownerId\": \"kn7bp63rpwq0rm2g7m4k0c3hjn80qkhc\",\n  \"slug\": \"grill-skill\",\n  \"version\": \"1.0.11\",\n  \"publishedAt\": 1790347950116\n}"},{"path":"REFERENCE.md","content":"# Grill Skill Reference\n\n## Caliper commands used by this skill\n\n```bash\n# Check spec is valid before running (unknown task or sandbox keys, bad forbidden_files\n# regexes, missing assert: files)\ncaliper validate path/to/spec.eval.yaml\n\n# First run — fast, catches spec errors\ncaliper run path/to/spec.eval.yaml --k 1\n\n# Reliability run — after iterating on the skill\ncaliper run path/to/spec.eval.yaml --k 3\n\n# Ablated run — before committing, proves the skill makes a difference.\n# Run once and keep it: it cannot move when the skill's text changes.\n# A declared mcp: server can be ablated the same way; qualify as skill:/mcp:\n# if both declare the name.\ncaliper run path/to/spec.eval.yaml --k 3 --ablate my-skill\n# Then diff it against the full run. A bare spec name resolves to that spec's\n# LATEST run, so address the older side by its saved results path.\ncaliper compare .caliper/results/<spec>/<ablated-run>.json <spec>\n\n# Choose the engine at run time — it is not stored in the spec (default: claude-code)\ncaliper run path/to/spec.eval.yaml --model codex:gpt-5-codex\ncaliper run path/to/spec.eval.yaml --model codex\ncaliper run path/to/spec.eval.yaml --judge-model claude-code:claude-haiku-4-5-20251001\n\n# Runs load your user customizations (skills, plugins, rules, settings and connectors) by default; isolate for a\n# portable score (or pin user_customizations: false in the spec)\ncaliper run path/to/spec.eval.yaml --no-user-customizations\n\n# Browse past results\ncaliper list\ncaliper report path/to/spec.eval.yaml\n\n# Compare two saved runs of the same eval (ablation: full vs. shortened, or over time)\ncaliper compare full-eval short-eval           # spec name -> latest run, or a results-JSON path\ncaliper compare a.json b.json --format json     # per-task Δ, regression flags, for scripting\n```\n\n`caliper compare <A> <B>` diffs two already-saved runs task by task: tasks are\nmatched by name, `Δ = b − a`, a negative Δ flags a regression (any-below), and a\nside with no usable attempts shows `—` (unmeasured, never a regression) so\ninfra/judge noise can't fake a loss. Under the success-rate headline, `compare` also\nshows **token and wall-clock deltas** (green = cheaper) — the \"same quality, 40%\nfewer tokens\" signal an ablation looks for. These are secondary: a token/time\nchange is **never** a regression (only the score is), and dollar cost is not tracked\n(tokens are the volume signal). Each attempt in the report also shows its tokens\nnext to its duration under `--verbose`.\n\n`compare` also reports **skill drift** — a member of the neighbourhood whose\n*text* changed between the two runs, read from the per-file hashes in each run's\nsnapshots. It is graded by provenance, not role: a drifted **git source** warns,\nbecause the spec claimed where those bytes came from and the delta you are\nreading is confounded; a drifted **path source** is shown without alarm, because\nnothing was promised about a working file and that edit is usually the thing the\nrun exists to measure.\n\n```\n ⚠ "},{"path":"skill-card.md","content":"## Description:\n\nInterviews skill authors to design eval tasks, then uses Caliper results to test and improve their skills.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[edonadei](https://clawhub.ai/user/edonadei)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nSkill authors and developers use this skill to design evaluation tasks, run Caliper tests, and identify changes that improve skill reliability.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: The skill can read and edit selected skill files and write evaluation specifications.\n\nMitigation: Use it in a trusted project and review generated evaluation YAML before confirming writes.\n\nRisk: Caliper tests may execute task setup or cleanup commands and load local agent customizations.\n\nMitigation: Review test commands and use isolated Caliper mode for portable or shared results.\n\n## Reference(s):\n\n- [Grill Skill on ClawHub](https://clawhub.ai/edonadei/skills/grill-skill)\n- [Grill Skill Reference](artifact/REFERENCE.md)\n\n## Skill Output:\n\n**Output Type(s):** [Guidance, Configuration, Shell commands]\n\n**Output Format:** [Markdown guidance and YAML evaluation specifications]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [Includes test results and comparisons when Caliper is run.]\n\n## Skill Version(s):\n\n1.0.11 (source: ClawHub release metadata)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment."},{"path":"grill-skill.eval.yaml","content":"skills:\n  - ./SKILL.md\n\n# The engine (backend + model) is a runtime axis, not a spec field — this eval\n# runs on whatever `--model` / `--judge-model` select (default claude-code).\n\n# No explicit sandbox.forbidden_files: caliper already auto-forbids the eval\n# spec and any .caliper/results/ directory. Listing \"./.caliper/.*\" here caused a\n# false cheat flag, because these grill-skill tasks legitimately WRITE caliper\n# specs whose text contains that very pattern.\n\ntasks:\n  # Task 1 — First-turn interview discipline (open prompt, no eval present).\n  # Exercises the prose we most want to shorten: Phase 1 understanding+confirm\n  # and the one-question-at-a-time discipline. Single-shot harness, so we judge\n  # the FIRST turn only: the agent must interview, not run ahead. The assert is\n  # a deterministic guard that it did NOT fabricate answers and write a spec.\n  - name: Interviews before generating — asks, then stops (does not run ahead)\n    activates: [grill-skill]\n    setup: |\n      rm -rf /tmp/grill-fresh\n      mkdir -p /tmp/grill-fresh\n      cat > /tmp/grill-fresh/SKILL.md << 'EOF'\n      ---\n      name: changelog-writer\n      description: Use when the user wants to turn merged PRs into a changelog entry.\n      allowed-tools: Bash, Read, Write\n      ---\n\n      # Changelog Writer\n\n      Read the merged PRs since the last tag and write a grouped changelog\n      entry (Features / Fixes / Chore) to CHANGELOG.md.\n      EOF\n    cleanup: rm -rf /tmp/grill-fresh\n    prompt: >\n      I want to create a caliper eval for my skill at\n      /tmp/grill-fresh/SKILL.md. I'm here and will answer whatever you need —\n      do NOT assume what the eval tasks should be, and do NOT write any files\n      yet. What do you need to know from me to get started?\n    expect: >\n      Pass if the agent opens the interview instead of running ahead: it reads\n      the SKILL.md, gives some understanding of the skill, and asks the user for\n      input before generating anything — then STOPS to wait. The number of\n      questions does not matter. Fail if the agent skips the interview: it\n      invents the user's answers, generates the eval tasks itself without\n      asking, or writes any .eval.yaml file in this turn.\n    assert: |\n      import glob\n      # It must not have run ahead and written a spec before interviewing.\n      specs = glob.glob(\"/tmp/grill-fresh/*.eval.yaml\")\n      assert not specs, f\"Agent wrote a spec without interviewing: {specs}\"\n\n  # Task 2 — Gap-fill detection + no silent overwrite (open prompt, eval present).\n  # Exercises the gap-fill prose: detect existing eval, report tasks, ask what's\n  # missing BEFORE changing anything. Assert is a deterministic guard that the\n  # existing spec was left untouched in this first turn.\n  - name: Detects an existing eval and asks before touching it\n    activates: [grill-skill]\n    setup: |\n      rm -rf /tmp/grill-existing\n      mkdir -p /tmp/grill-existing\n      cat > /tmp/grill-existing/SKILL.md << 'EOF'\n      ---\n      name"}],"languages":[],"docsSourceLabel":"CLAWHUB","editorialOverview":null,"editorialQuality":{"score":100,"threshold":65,"status":"thin","wordCount":2125,"uniquenessScore":41,"reasons":["uniqueness-below-45"]}},"media":{"evidence":{"source":"no-media","verified":false,"confidence":"low","updatedAt":"2026-10-11T04:54:05.694Z","emptyReason":"No screenshots, media assets, or demo links are available."},"primaryImageUrl":null,"mediaAssetCount":0,"assets":[],"demoUrl":null},"ownerResources":{"evidence":{"source":"unclaimed","verified":false,"confidence":"low","updatedAt":"2026-10-11T04:54:05.694Z","emptyReason":"This page has not been claimed by the agent owner."},"hasCustomPage":false,"customPageUpdatedAt":null,"customLinks":[],"structuredLinks":{"docsUrl":null,"demoUrl":null,"supportUrl":null,"pricingUrl":null,"statusUrl":null},"customPage":null},"relatedAgents":{"evidence":{"source":"protocol-neighbors","verified":false,"confidence":"medium","updatedAt":"2026-10-11T07:40:11.111Z","emptyReason":null},"items":[{"id":"8ebccd8e-3863-4187-8355-c3f14e1f9edf","entityType":"agent","canonicalPath":"/agent/iofficeai-aionui","slug":"iofficeai-aionui","name":"AionUi","description":"Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!","url":"https://github.com/iOfficeAI/AionUi","homepage":"https://www.aionui.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-10-09T19:11:12.944Z","createdAt":"2026-02-25T03:38:16.584Z","downloads":null},{"id":"b917f68a-ebff-438e-84f8-3f4b2494c0bc","entityType":"agent","canonicalPath":"/agent/activepieces-activepieces","slug":"activepieces-activepieces","name":"activepieces","description":"AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents","url":"https://github.com/activepieces/activepieces","homepage":"https://www.activepieces.com","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-15T02:22:12.426Z","createdAt":"2026-02-25T03:38:12.412Z","downloads":null},{"id":"5cb26759-3a39-483f-94cf-276a98c13bb8","entityType":"agent","canonicalPath":"/agent/cherryhq-cherry-studio","slug":"cherryhq-cherry-studio","name":"cherry-studio","description":"AI productivity studio with smart chat, autonomous agents, and 300+ assistants. Unified access to frontier LLMs","url":"https://github.com/CherryHQ/cherry-studio","homepage":"https://cherry-ai.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-11T14:38:40.986Z","createdAt":"2026-02-25T03:38:19.379Z","downloads":null},{"id":"6f6582d0-5d76-4f0f-b81d-86520247950b","entityType":"agent","canonicalPath":"/agent/copilotkit-copilotkit","slug":"copilotkit-copilotkit","name":"CopilotKit","description":"The Frontend for Agents & Generative UI. React + Angular","url":"https://github.com/CopilotKit/CopilotKit","homepage":"https://docs.copilotkit.ai","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-03-25T09:50:57.846Z","createdAt":"2026-02-25T03:39:14.617Z","downloads":null}],"links":{"hub":"/agent","source":"/agent/source/clawhub","protocols":[{"label":"OpenClaw","href":"/agent/protocol/openclew"}]}}}