{"id":"68d3317d-02ed-4ceb-a8b5-fea515070d98","entityType":"agent","slug":"clawhub-edonadei-evaluate-skill","name":"evaluate-skill","canonicalUrl":"https://www.xpersona.co/agent/clawhub-edonadei-evaluate-skill","canonicalPath":"/agent/clawhub-edonadei-evaluate-skill","generatedAt":"2026-10-11T04:33:35.287Z","source":"CLAWHUB","claimStatus":"UNCLAIMED","verificationTier":"NONE","summary":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-11T02:16:22.404Z","emptyReason":null},"description":"Measure a skill's reliability — run it k times for a pass@k score, design or interpret its eval, or compare it against the base agent. Use when the user wants to run, design, or interpret a skill's eval, or write an .eval.yaml spec.","descriptionLabel":"Source description","evidenceSummary":"Capability contract not published. No trust telemetry is available yet. 1.2K downloads reported by the source. Last updated 10/11/2026.","installCommand":"clawhub skill install s17fhassds1tss1zjq6jkjcfd983rjwt:evaluate-skill","sourceUrl":"https://clawhub.ai/edonadei/evaluate-skill","homepage":"https://clawhub.ai/edonadei/skills/evaluate-skill","primaryLinks":[{"label":"View on ClawHub","url":"https://clawhub.ai/edonadei/evaluate-skill","kind":"source"},{"label":"Homepage","url":"https://clawhub.ai/edonadei/skills/evaluate-skill","kind":"homepage"}],"safetyScore":84,"overallRank":62,"popularityScore":62,"trustScore":null,"claimedByName":null,"isOwner":false,"seoDescription":"evaluate-skill technical dossier on Xpersona with agent coverage, OPENCLEW support, and live trust metadata."},"coverage":{"evidence":{"source":"public-profile","verified":false,"confidence":"medium","updatedAt":"2026-10-11T02:16:22.404Z","emptyReason":null},"protocols":[{"protocol":"OPENCLEW","label":"OpenClaw","status":"self-declared","notes":"Declared in the public agent profile."}],"capabilities":[],"verifiedCount":0,"selfDeclaredCount":1,"capabilityMatrix":{"rows":[{"key":"OPENCLEW","type":"protocol","support":"unknown","confidenceSource":"profile","notes":"Listed on profile"}],"flattenedTokens":"protocol:OPENCLEW|unknown|profile"}},"adoption":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-11T02:16:22.404Z","emptyReason":null},"stars":null,"forks":null,"downloads":1192,"packageName":null,"latestVersion":"1.0.14","tractionLabel":"1.2K downloads"},"release":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-11T02:16:22.344Z","emptyReason":null},"lastUpdatedAt":"2026-10-11T02:16:22.404Z","lastCrawledAt":"2026-10-11T02:16:22.344Z","lastIndexedAt":null,"nextCrawlAt":"2026-10-12T02:16:22.344Z","lastVerifiedAt":null,"highlights":[{"version":"1.0.14","createdAt":"2026-09-25T15:12:56.018Z","changelog":"### Changed - Clean-up from the 1.0.13 that added way too much files. - Built around four questions, each with its own place to fix: does the skill fire (`description`), does it work (body), does it earn its place (the tasks), does it hold across edits. - New workflow order: run the control (`--ablate <skill>`) once, before editing the skill, and keep it. After each edit, compare against the previous full run as well as the control, because a worse skill can still beat the control. - New diagnosis table traces each failure to its fix: setup errors, unusable attempts, `cheat`, a skill that didn't fire, an extra skill that fired, a skill that fired but failed, and tasks that pass without the skill. - Specs it writes put `activates:` on every execution task (plus any skill the task delegates to), never name the skill in a prompt, and include at least one trigger probe. - Now owns writing a spec whose tasks are already decided, and advising on runs. The interview moved entirely to grill-skill. - Covers user customizations: runs load your own skills and setup by default. Isolate when comparing backends, sharing a number, or measuring the bare agent. - Reference rewritten and cut to what an agent acts on; flags are left to `caliper --help`. ### Fixed - The spec example used `skill: path:`, which `caliper validate` rejects. - Example evals used fixed `/tmp` paths that collide across parallel attempts; they now build fixtures in the attempt workdir. - The commit-simple example expected commits a single-shot run can't reach. ### Measured With the skill, Opus 5.5 went 9/9 on authoring tasks (writing `activates:`, proposing trigger probes, running the control first), against 0/9 without it, at k=3.","fileCount":5,"zipByteSize":18300},{"version":"1.0.13","createdAt":"2026-09-25T14:51:50.427Z","changelog":"### Changed - Built around four questions, each with its own place to fix: does the skill fire (`description`), does it work (body), does it earn its place (the tasks), does it hold across edits. - New workflow order: run the control (`--ablate <skill>`) once, before editing the skill, and keep it. After each edit, compare against the previous full run as well as the control, because a worse skill can still beat the control. - New diagnosis table traces each failure to its fix: setup errors, unusable attempts, `cheat`, a skill that didn't fire, an extra skill that fired, a skill that fired but failed, and tasks that pass without the skill. - Specs it writes put `activates:` on every execution task (plus any skill the task delegates to), never name the skill in a prompt, and include at least one trigger probe. - Now owns writing a spec whose tasks are already decided, and advising on runs. The interview moved entirely to grill-skill. - Covers user customizations: runs load your own skills and setup by default. Isolate when comparing backends, sharing a number, or measuring the bare agent. - Reference rewritten and cut to what an agent acts on; flags are left to `caliper --help`. ### Fixed - The spec example used `skill: path:`, which `caliper validate` rejects. - Example evals used fixed `/tmp` paths that collide across parallel attempts; they now build fixtures in the attempt workdir. - The commit-simple example expected commits a single-shot run can't reach. ### Measured With the skill, Opus 5.5 went 9/9 on authoring tasks (writing `activates:`, proposing trigger probes, running the control first), against 0/9 without it, at k=3.","fileCount":42,"zipByteSize":43089},{"version":"1.0.12","createdAt":"2026-08-29T06:35:03.533Z","changelog":"Improvements on the Caliper underlying CLI focused on reliability and usability (performance and retries)","fileCount":16,"zipByteSize":30932},{"version":"1.0.11","createdAt":"2026-07-12T17:43:04.366Z","changelog":"- MCP servers (mcp: spec block) — documented the new optional top-level mcp: mapping: local stdio (command/args/env) vs remote HTTP/SSE (type/url/headers) servers, ${VAR} secret resolution at the harness boundary, per-backend support (claude-code, hermes, codex; pi refuses by design), namespaced tool names in the transcript (mcp__server__tool / mcp_server_tool), and caliper validate rules. Added to both the spec skeleton and Key concepts. - Persisted transcripts — results-storage section now notes each attempt records an optional transcript array (ordered turns with tool_name/tool_input/tool_output); older result JSON without it still loads.","fileCount":16,"zipByteSize":26996},{"version":"1.0.10","createdAt":"2026-07-05T20:21:02.928Z","changelog":"- Added token and wall-time tracking to evaluations - Moved from pass@k to raw success rate, easier to reason about for agents","fileCount":16,"zipByteSize":25926},{"version":"1.0.9","createdAt":"2026-07-03T20:02:14.791Z","changelog":"- Added new Caliper run result files to .caliper/results/evaluate-skill. - Caliper now supports Hermes agent","fileCount":16,"zipByteSize":24924},{"version":"1.0.8","createdAt":"2026-07-03T16:40:21.334Z","changelog":"- Updated to require backend and model selection at run time via CLI flags; these are no longer specified in the eval spec. - Removed all references to specifying backend/model or a judge section in `.eval.yaml` files. - Clarified instructions: choose engines using `--model` and `--judge-model` when running evaluations. - Maintained existing examples and guidance on designing and committing evals. - Added two new Caliper result JSON files in `.caliper/results/evaluate-skill/`.","fileCount":16,"zipByteSize":24579},{"version":"1.0.7","createdAt":"2026-07-03T04:14:30.803Z","changelog":"- Updated backend documentation: Only CLI backends (`claude-code`, `codex`, `pi`) are now supported; API backends (`claude-api`, `openai-api`) are no longer listed. - Clarified that every backend is a CLI agent using its own subscription/auth—no direct API backend remains. - General documentation clean-up to reflect backend changes and to ensure instructions are current.","fileCount":16,"zipByteSize":24474}]},"execution":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No published capability contract is available yet."},"installCommand":"clawhub skill install s17fhassds1tss1zjq6jkjcfd983rjwt:evaluate-skill","setupComplexity":"low","setupSteps":["Install using `clawhub skill install s17fhassds1tss1zjq6jkjcfd983rjwt:evaluate-skill` in an isolated environment before connecting it to live workloads.","No published capability contract is available yet, so validate auth and request/response behavior manually.","Review the upstream CLAWHUB listing at https://clawhub.ai/edonadei/evaluate-skill before using production credentials."],"contract":{"contractStatus":"missing","authModes":[],"requires":[],"forbidden":[],"supportsMcp":false,"supportsA2a":false,"supportsStreaming":false,"inputSchemaRef":null,"outputSchemaRef":null,"dataRegion":null,"contractUpdatedAt":null,"sourceUpdatedAt":null,"freshnessSeconds":null},"invocationGuide":{"preferredApi":{"snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-edonadei-evaluate-skill/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-edonadei-evaluate-skill/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-edonadei-evaluate-skill/trust"},"curlExamples":["curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-edonadei-evaluate-skill/snapshot\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-edonadei-evaluate-skill/contract\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-edonadei-evaluate-skill/trust\""],"jsonRequestTemplate":{"query":"summarize this repo","constraints":{"maxLatencyMs":2000,"protocolPreference":["OPENCLEW"]}},"jsonResponseTemplate":{"ok":true,"result":{"summary":"...","confidence":0.9},"meta":{"source":"CLAWHUB","generatedAt":"2026-10-11T04:33:35.284Z"}},"retryPolicy":{"maxAttempts":3,"backoffMs":[500,1500,3500],"retryableConditions":["HTTP_429","HTTP_503","NETWORK_TIMEOUT"]}},"endpoints":{"dossierUrl":"https://www.xpersona.co/api/v1/agents/clawhub-edonadei-evaluate-skill/dossier","snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-edonadei-evaluate-skill/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-edonadei-evaluate-skill/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-edonadei-evaluate-skill/trust"}},"reliability":{"evidence":{"source":"runtime-metrics","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No trust, reliability, or runtime telemetry is available."},"trust":{"status":"unavailable","handshakeStatus":"UNKNOWN","verificationFreshnessHours":null,"reputationScore":null,"p95LatencyMs":null,"successRate30d":null,"fallbackRate":null,"attempts30d":null,"trustUpdatedAt":null,"trustConfidence":"unknown","sourceUpdatedAt":null,"freshnessSeconds":null},"decisionGuardrails":{"doNotUseIf":["Contract metadata is missing or unavailable for deterministic execution."],"safeUseWhen":[],"riskFlags":["missing_or_unavailable_contract","trust_data_unavailable","schema_references_missing"],"operationalConfidence":"low"},"executionMetrics":{"observedLatencyMsP50":null,"observedLatencyMsP95":null,"estimatedCostUsd":null,"uptime30d":null,"rateLimitRpm":null,"rateLimitBurst":null,"lastVerifiedAt":null,"verificationSource":null},"runtimeMetrics":{"successRate":null,"avgLatencyMs":null,"avgCostUsd":null,"hallucinationRate":null,"retryRate":null,"disputeRate":null,"p50Latency":null,"p95Latency":null,"lastUpdated":null}},"benchmarks":{"evidence":{"source":"no-benchmark-data","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No benchmark suites or observed failure patterns are available."},"suites":[],"failurePatterns":[]},"artifacts":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-11T02:16:22.404Z","emptyReason":null},"readme":"Skill: evaluate-skill\n\nOwner: edonadei\n\nSummary: Measure a skill's reliability — run it k times for a pass@k score, design or interpret its eval, or compare it against the base agent. Use when the user wants to run, design, or interpret a skill's eval, or write an .eval.yaml spec.\n\nTags: latest:1.0.14\n\nVersion history:\n\nv1.0.14 | 2026-09-25T15:12:56.018Z | user\n\n### Changed\n- Clean-up from the 1.0.13 that added way too much files.\n- Built around four questions, each with its own place to fix: does the skill fire (`description`), does it work (body), does it earn its place (the tasks), does it hold across edits.\n- New workflow order: run the control (`--ablate <skill>`) once, before editing the skill, and keep it. After each edit, compare against the previous full run as well as the control, because a worse skill can still beat the control.\n- New diagnosis table traces each failure to its fix: setup errors, unusable attempts, `cheat`, a skill that didn't fire, an extra skill that fired, a skill that fired but failed, and tasks that pass without the skill.\n- Specs it writes put `activates:` on every execution task (plus any skill the task delegates to), never name the skill in a prompt, and include at least one trigger probe.\n- Now owns writing a spec whose tasks are already decided, and advising on runs. The interview moved entirely to grill-skill.\n- Covers user customizations: runs load your own skills and setup by default. Isolate when comparing backends, sharing a number, or measuring the bare agent.\n- Reference rewritten and cut to what an agent acts on; flags are left to `caliper --help`.\n\n### Fixed\n- The spec example used `skill: path:`, which `caliper validate` rejects.\n- Example evals used fixed `/tmp` paths that collide across parallel attempts; they now build fixtures in the attempt workdir.\n- The commit-simple example expected commits a single-shot run can't reach.\n\n### Measured\nWith the skill, Opus 5.5 went 9/9 on authoring tasks (writing `activates:`, proposing trigger probes, running the control first), against 0/9 without it, at k=3.\n\nv1.0.13 | 2026-09-25T14:51:50.427Z | user\n\n### Changed\n- Built around four questions, each with its own place to fix: does the skill fire (`description`), does it work (body), does it earn its place (the tasks), does it hold across edits.\n- New workflow order: run the control (`--ablate <skill>`) once, before editing the skill, and keep it. After each edit, compare against the previous full run as well as the control, because a worse skill can still beat the control.\n- New diagnosis table traces each failure to its fix: setup errors, unusable attempts, `cheat`, a skill that didn't fire, an extra skill that fired, a skill that fired but failed, and tasks that pass without the skill.\n- Specs it writes put `activates:` on every execution task (plus any skill the task delegates to), never name the skill in a prompt, and include at least one trigger probe.\n- Now owns writing a spec whose tasks are already decided, and advising on runs. The interview moved entirely to grill-skill.\n- Covers user customizations: runs load your own skills and setup by default. Isolate when comparing backends, sharing a number, or measuring the bare agent.\n- Reference rewritten and cut to what an agent acts on; flags are left to `caliper --help`.\n\n### Fixed\n- The spec example used `skill: path:`, which `caliper validate` rejects.\n- Example evals used fixed `/tmp` paths that collide across parallel attempts; they now build fixtures in the attempt workdir.\n- The commit-simple example expected commits a single-shot run can't reach.\n\n### Measured\nWith the skill, Opus 5.5 went 9/9 on authoring tasks (writing `activates:`, proposing trigger probes, running the control first), against 0/9 without it, at k=3.\n\nv1.0.12 | 2026-08-29T06:35:03.533Z | user\n\nImprovements on the Caliper underlying CLI focused on reliability and usability (performance and retries)\n\nv1.0.11 | 2026-07-12T17:43:04.366Z | user\n\n- MCP servers (mcp: spec block) — documented the new optional top-level mcp: mapping: local stdio (command/args/env) vs remote HTTP/SSE (type/url/headers) servers, ${VAR} secret resolution at the harness boundary, per-backend support (claude-code, hermes, codex; pi refuses by design), namespaced tool names in the transcript (mcp__server__tool / mcp_server_tool), and caliper validate rules. Added to both the spec skeleton and Key concepts.\n- Persisted transcripts — results-storage section now notes each attempt records an optional transcript array (ordered turns with tool_name/tool_input/tool_output); older result JSON without it still loads.\n\nv1.0.10 | 2026-07-05T20:21:02.928Z | user\n\n- Added token and wall-time tracking to evaluations\n- Moved from pass@k to raw success rate, easier to reason about for agents\n\nv1.0.9 | 2026-07-03T20:02:14.791Z | user\n\n- Added new Caliper run result files to .caliper/results/evaluate-skill.\n- Caliper now supports Hermes agent\n\nv1.0.8 | 2026-07-03T16:40:21.334Z | user\n\n- Updated to require backend and model selection at run time via CLI flags; these are no longer specified in the eval spec.\n- Removed all references to specifying backend/model or a judge section in `.eval.yaml` files.\n- Clarified instructions: choose engines using `--model` and `--judge-model` when running evaluations.\n- Maintained existing examples and guidance on designing and committing evals.\n- Added two new Caliper result JSON files in `.caliper/results/evaluate-skill/`.\n\nv1.0.7 | 2026-07-03T04:14:30.803Z | user\n\n- Updated backend documentation: Only CLI backends (`claude-code`, `codex`, `pi`) are now supported; API backends (`claude-api`, `openai-api`) are no longer listed.\n- Clarified that every backend is a CLI agent using its own subscription/auth—no direct API backend remains.\n- General documentation clean-up to reflect backend changes and to ensure instructions are current.\n\nv1.0.6 | 2026-07-02T19:29:52.623Z | user\n\n- Improved clarity and conciseness of SKILL.md, with reorganized structure and more focused instructions.\n- Updated description to emphasize reliability measurement, running evals, and comparing skills to base agents.\n- Shortened backend explanations and moved advanced details to REFERENCE.md.\n- Clarified when to use grill-skill vs evaluate-skill.\n- Rewrote eval design checklist for easier application; moved deeper guidance to REFERENCE.md.\n- Removed skill-card.md. Added two new `.caliper/results/` JSON sample files.\n\nv1.0.5 | 2026-06-26T03:27:12.565Z | user\n\n- Adds support for the `pi` backend via the pi CLI for evaluating pi skills.\n- Documents the new pi backend and its usage in the prerequisites section.\n- Removes the obsolete skill-card.md file.\n- No user-facing changes to existing eval creation or workflow.\n\nv1.0.4 | 2026-06-24T00:51:35.245Z | user\n\n- Added new Caliper evaluation result files to .caliper/results/evaluate-skill/.\n- Removed the skill-card.md file from the project.\n- No changes to core documentation or usage instructions.\n\nv1.0.3 | 2026-06-20T03:00:56.900Z | user\n\n- Added sample Caliper eval result artifacts to .caliper/results/evaluate-skill/\n- No changes to core functionality or documentation; only artifact files added\n\nv1.0.2 | 2026-06-19T23:44:55.378Z | user\n\n- Added guidance for users without an existing `.eval.yaml`, recommending the `grill-skill` workflow for interactive eval creation.\n- Included best practices for prompting users to commit their `.eval.yaml` spec after creation or modification.\n- Expanded instructions on artifact types, storage locations, and version control implications.\n- Clarified when to use `evaluate-skill` directly versus when to use `grill-skill` for eval design.\n- No code or functionality changes; documentation and workflow enhancements only.\n\nv1.0.1 | 2026-06-19T22:31:39.717Z | user\n\n- Updated Caliper CLI installation instructions: now recommends installing via `pipx install caliper-eval` instead of cloning from the repository and running `pip install -e .`.\n- Removed detailed instructions for obtaining Caliper from the GitHub repository for a simpler and faster setup.\n\nv1.0.0 | 2026-06-19T22:20:52.746Z | auto\n\n- Initial release of the evaluate-skill.\n- Provides guidance for running, designing, and interpreting Caliper evals, including writing `.eval.yaml` specifications.\n- Documents support for multiple backends (Claude Code, Codex, Anthropic API, OpenAI API) with independent control over agent and judge backends.\n- Includes best practices and checklists for creating high-quality agentic evaluations, example patterns, and bundled reference evals.\n- Offers examples for writing clear, deterministic expectations and evaluation rubrics.\n\nArchive index:\n\nArchive v1.0.14: 5 files, 18300 bytes\n\nFiles: evaluate-skill.eval.yaml (7175b), REFERENCE.md (30296b), skill-card.md (1925b), SKILL.md (4548b), _meta.json (134b)\n\nFile v1.0.14:SKILL.md\n\n---\nname: evaluate-skill\ndescription: Measure a skill's reliability — run it k times for a pass@k score, design or interpret its eval, or compare it against the base agent. Use when the user wants to run, design, or interpret a skill's eval, or write an .eval.yaml spec.\nallowed-tools: Bash\n---\n\n# Evaluate Skill\n\nRun a skill repeatedly to measure how reliably it works, and design the evals that measure it.\n\n## Prerequisites\n\nThe `caliper` CLI must be on `PATH`. This skill can be copied into an agent without the Caliper repo, so do not assume the CLI is packaged with it. Install if missing:\n\n```bash\npipx install caliper-eval\n```\n\nThe engine (backend + model) is not part of the spec — it is chosen at run time with `--model` (skill) and `--judge-model` (judge), independently, from `claude-code`, `codex`, `pi`, defaulting to `claude-code`. Every backend is a CLI agent that uses its own subscription/auth; there is no direct-API backend (for API billing, configure a CLI with an API key). Full per-backend detail and every command: [REFERENCE.md](REFERENCE.md).\n\n## Spec shape\n\nAn `.eval.yaml` names the skill and a list of tasks. Keep `skill.path` relative to the spec file (usually `./SKILL.md`):\n\n```yaml\nskill:\n  path: ./SKILL.md      # relative to the spec file\ntasks:\n  - name: What success looks like\n    prompt: <prompt sent to the skill under test>\n    expect: <natural-language pass/fail criterion>\n    assert: |           # optional deterministic Python check\n      assert ...\n```\n\nThe spec has no `backend`/`model` or `judge:` block; pick the engine when you run, e.g. `caliper run <spec> --model codex --judge-model codex`. The full format (setup/cleanup, external assert scripts, sandbox) is in [REFERENCE.md](REFERENCE.md).\n\n## Bundled references\n\n`references/evals/` holds complete real examples (Claude Code smoke, commit workflow, screenshot, summarization, TDD) — each folder self-contained with its fixture `SKILL.md` and `.eval.yaml`. `references/simple.eval.yaml` is one compact multi-task spec.\n\n## No eval yet?\n\nIf the skill has a `SKILL.md` but no `.eval.yaml`, suggest the `grill-skill` workflow — it interviews the user and generates a happy/edge/adversarial spec. Use `evaluate-skill` directly when a spec already exists and the user wants to run, validate, report, or extend it.\n\n## Designing good evals\n\n1. Name the target behavior — what should the skill do better than the base agent?\n2. Decide whether the suite is a capability eval or a regression eval.\n3. Cover normal, edge, and adversarial cases when the behavior matters.\n4. Grade artifacts (files, git state, command output, exact values) whenever you can; judge the transcript only when the behavior itself is the point. The full artifact-vs-transcript rules, the task-quality checklist, common eval patterns, and how to write `expect:` rubrics live in [REFERENCE.md](REFERENCE.md) — read and apply them when designing tasks.\n5. Run once with `--ablate <skill-name>` and `caliper compare` the two runs, to confirm the skill beats the raw agent. Debug the spec at `--k 1`, then measure reliability at `--k 3` or higher.\n\n**Done when:** tasks have observable success criteria, at least one deterministic `assert:`, a positive delta against the ablated run, the spec passes `caliper validate`, and the user has been prompted to commit the spec.\n\n## Whose setup is measured\n\nRuns load the user's own customizations by default (user skills, plugins, rules, settings and connectors; see REFERENCE.md for backend exceptions), which answers \"does my skill work in *my* agent?\". **Isolate** (`--no-user-customizations`, or `user_customizations: false` in the spec) when comparing backends or models, when the number leaves this machine (shared, published, compared with someone else's run), or when measuring the bare agent: each setup is different, so otherwise part of the delta is the setups. `--ablate` of the user's own skill needs no isolation, since both runs load the same setup.\n\n**Always tell the user which mode ran** and what it loaded, from the report header's `user customizations:` line (absent means isolated), and relay any fix `caliper compare` suggests about it.\n\n## Committing\n\nRunning Caliper produces two artifacts: the `.eval.yaml` spec — the valuable one, commit it beside the skill so anyone who clones the repo can run the same eval — and `.caliper/results/` saved run JSONs, useful for diffing over time and safe to gitignore. After creating or running an eval, tell the user to commit the spec alongside `SKILL.md`.\n\nFile v1.0.14:_meta.json\n\n{\n  \"ownerId\": \"kn7bp63rpwq0rm2g7m4k0c3hjn80qkhc\",\n  \"slug\": \"evaluate-skill\",\n  \"version\": \"1.0.14\",\n  \"publishedAt\": 1790349176018\n}\n\nFile v1.0.14:REFERENCE.md\n\n# Caliper Reference\n\n## Commands\n\n### Run an evaluation\n```bash\ncaliper run path/to/spec.eval.yaml --k 3\ncaliper run path/to/spec.eval.yaml --k 3 --ablate my-skill   # same tasks, that skill removed (or an mcp: server)\ncaliper run path/to/spec.eval.yaml --k 3 --ablate a --ablate b  # repeatable; name them all (+ --no-user-customizations) for the bare agent\ncaliper run path/to/spec.eval.yaml --verbose             # show per-attempt reasoning\ncaliper run path/to/spec.eval.yaml --no-user-customizations      # portable: only declared skills and servers\ncaliper run path/to/spec.eval.yaml --user-customizations         # load your user customizations even if the spec pins user_customizations: false\n\n# Choose the engine at run time — it is not stored in the spec (default: claude-code)\ncaliper run path/to/spec.eval.yaml --model codex:gpt-5-codex\ncaliper run path/to/spec.eval.yaml --model codex                # backend only, its default model\ncaliper run path/to/spec.eval.yaml --model claude-sonnet-4-6    # model only, backend stays claude-code\ncaliper run path/to/spec.eval.yaml --judge-model claude-code:claude-haiku-4-5-20251001\ncaliper run path/to/spec.eval.yaml --model codex --judge-model claude-code:claude-haiku-4-5-20251001\n```\n\n### Validate a spec file\n```bash\ncaliper validate path/to/spec.eval.yaml\n```\n\nRejects unknown task and `sandbox:` keys (a typo like `asert:`), `forbidden_files` entries that are not valid regexes, and `assert:` script files missing from beside the spec. `caliper run` runs the same checks before its first attempt.\n\n### Browse saved results\n```bash\ncaliper list                        # all specs with latest scores\ncaliper list my-skill-eval          # all runs for one spec: Run id + which were ablated\ncaliper report my-skill-eval        # latest run (table view)\ncaliper report my-skill-eval --run 2026-05-12T14-23-01Z  # specific run\ncaliper report results.json --format json\n```\n\n### Compare two runs (ablation)\nDiff two already-saved runs of the same eval — full vs. shortened skill, or the\nsame skill over time. Tasks are matched by name; `Δ = b − a`; a negative Δ flags\na regression; a side with no usable attempts shows `—` (unmeasured, never a\nregression); the headline `Δ (matched)` averages only tasks measured on both\nsides. Each argument is addressed like `report` (spec name → latest run, or a\nresults-JSON path); pin a historical run by naming its path.\n```bash\ncaliper compare full-eval short-eval          # latest run of each spec\ncaliper compare a.json b.json                 # pin specific runs\ncaliper compare full-eval short-eval --format json   # for a ship/no-ship gate\n```\n\n## Spec format (.eval.yaml)\n\nThe spec carries no engine — no `backend`/`model` and no `judge:` block. Backend\nand model for both the skill and the judge are chosen at run time via `--model` /\n`--judge-model` (default `claude-code`); a spec that still pins these keys fails\nvalidation with a message pointing at the flags.\n\n```yaml\nskills:                   # installed at the agent's own skills root, never\n  - ./SKILL.md            #   preloaded — the agent has to choose it\n  # add further entries to test that yours is the one that fires (they are\n  # assertable via `activates:`, not decoration). A bare string is a *path\n  # source*; a mapping is a *git source* caliper clones for you:\n  - repo: vercel-labs/agent-skills\n    ref: a1b2c3d          # optional — omit to track the default branch\n    path: skills/tdd/SKILL.md   # optional — defaults to SKILL.md at the root\n\nsandbox:\n  forbidden_files:\n    - \".*\\\\.eval\\\\.yaml$\"   # agent cannot read the spec\n\nmcp:                        # optional — MCP servers the agent may use\n  weather:                  # server name → mcp__weather__<tool> in the transcript\n    command: python3        # local stdio command the harness spawns\n    args: [./servers/weather.py]\n    env:\n      API_TOKEN: ${MCP_API_TOKEN}   # resolved from your shell at run time\n  gdrive:                   # remote (hosted) server reached over HTTP/SSE\n    type: http              # http or sse\n    url: https://mcp.example.com/gdrive\n    headers:\n      Authorization: Bearer ${GDRIVE_TOKEN}   # resolved from your shell at run time\n\ntasks:\n  - id: task-001\n    name: Short description of what success looks like\n    setup: <shell command to prepare the workdir>        # optional\n    cleanup: <shell command to tear down>               # optional\n    prompt: <prompt sent to the AI under evaluation>\n    expect: <natural language description of a successful outcome>\n\n  - id: task-002\n    name: Task with a deterministic assertion\n    prompt: Write hello to out.txt\n    expect: A file is created at out.txt\n    assert: |\n      import os\n      assert os.path.exists(\"out.txt\"), \"File not created\"\n\n  - id: task-003\n    name: Task with external assertion script\n    prompt: Generate a report\n    assert: ./assertions/check_report.py\n\n  - id: task-004\n    name: A neighbour's prompt — yours must not hijack it\n    prompt: <a prompt that belongs to a different declared skill>\n    activates: [other-skill]      # exactly these fired, and nothing else\n\n  - id: task-005\n    name: Unrelated work — silence expected\n    prompt: <a prompt no declared skill should answer>\n    activates: []                 # nothing fired\n```\n\nEach task needs at least one of `expect`, `assert` or `activates`.\n\n`activates:` asserts the **exact set** of skills the agent loaded on an attempt.\nSkills are installed where the agent looks for them and never pasted into the\nprompt, so *choosing* one is an observable. A task carrying only `activates:` is\na **trigger probe**: it skips the judge entirely (much cheaper than an execution\ntask) and reports as `trigger only`, not a zero. Activation is scored on its own\nscoreboard and never blended into the success rate — a failing `description` and\na failing body are fixed in different places.\n\nIdentity is the frontmatter `name:`, so `activates:` names must match the\n`name:` in each declared `SKILL.md`, not its filename or directory. A skill the\nspec doesn't declare is never installed and can never activate — so if yours\ndelegates to another skill, declare it and enumerate the chain.\n\nThe same spec runs on any engine. To run it on Codex, pass the backend at run\ntime — the spec is unchanged:\n\n```bash\ncaliper run path/to/spec.eval.yaml --model codex --judge-model codex\n```\n\nFor pi, the `:model` half overrides pi's configured default:\n\n```bash\ncaliper run path/to/spec.eval.yaml --model pi:claude-sonnet-4-6 --judge-model pi\n```\n\nFor hermes (Nous Research), the `:model` half is a `provider/model` value passed\nstraight to hermes's `-m`; omit it to use your `~/.hermes/config.yaml` default:\n\n```bash\ncaliper run path/to/spec.eval.yaml --model hermes:anthropic/claude-sonnet-4.6 --judge-model hermes\n```\n\nHermes is a stateful agent, so Caliper normalizes it to a neutral agent per\nattempt (isolated `HERMES_HOME`, no persona/memory, `--ignore-rules`, and only\nthe spec's declared skills installed) and recovers the full tool-call\ntrajectory by running `hermes -z` then `hermes sessions export`.\n\n## Key concepts\n\n- **success rate** — the primary score: `successes / usable` (how often a *single* run works), computed over the usable attempts only. Two secondary views (on every task in the JSON as `pass_at_k`/`pass_hat_k`, and under `--verbose`) reframe it for how the skill is used: **pass@k** = `1−(1−p)^k` = P(≥1 of k pass) — the retry / \"eventual success\" lens, **≥** the rate, for when a failure is cheap to retry and you keep the good run; **pass^k** = `p^k` = P(all k pass) — the strict / \"must never fail\" lens, **≤** the rate, for when the skill runs unattended or as one link in a chain. When in doubt use the raw rate — pass@k is the code-gen metric and flatters flaky skills (`1/3 → 70.4%`)\n- **outcome** — each attempt is typed `pass`, `task_fail`, `cheat`, `infra_error`, `timeout`, `judge_error`, or `not_checked`; `infra_error` also covers a zero-exit attempt where no model call was observed (nothing parsed, no tokens reported), which skips the judge and the activation check; `infra_error`/`timeout`/`judge_error` are *unusable* (infrastructure/judge noise) and are excluded from the score denominator and reported as a separate \"N unusable\" count, so a throttled or judge-flaked run is not mistaken for a regression. `not_checked` is neither: the task authored no `expect:`/`assert:` (a trigger probe), so it leaves the denominator without being reported as an error. `passed` in the JSON equals `outcome == pass`.\n- **`--fail-fast N`** — optional run control that stops scheduling new attempts for a task after N consecutive `infra_error`/`timeout` outcomes; `0` disables it. Counts **attempts**, not invocations, so with retries it can cost up to three times the spawns. Early-stopped tasks report as `ABORTED`, and tasks with no usable attempts keep `score: null`.\n- **`--workers N`** — how many **attempts** run in parallel, across all tasks (default `4`); every (task, attempt) pair is scheduled independently, so `--k 10 --workers 4` on a one-task spec runs four at a time. `--fail-fast` is the exception: it keeps each task's attempts sequential, since the streak it counts is only defined in order. Concurrent agents share one upstream rate limit, so a high `--workers` buys wall-clock time at the risk of more `infra_error` attempts.\n- **retry** — an *attempt* is one measured shot at the task, an *invocation* one spawn of the agent; they differ when the provider refuses. A throttle (429 / overloaded / 503) is retried twice with backoff and folds into **one** `AttemptRecord` with a `retries` count — the score's denominator is attempts, never spawns. A **spending cap** is not retried and not survivable: it aborts the run (saving what ran, reporting the cap), because every remaining attempt would meet the same wall. A bare crash is neither retried nor aborted on — it reproduces. The merged record takes its *output* from the last invocation, **sums** tokens across all of them, and sums only the spawn time, excluding backoff (docs/adr/0019).\n- **judge time** — `AttemptRecord.judge_seconds` is the wall-clock the judge spent grading one attempt, `null` when no judge ran (assert-only task, or an attempt that exited before reaching one). A **sibling** of `duration_seconds`, never folded into it — that field stays the harness spawn, so `Wall` means the same thing it always did. The run summary adds a `Judge` line averaged over *graded* attempts only.\n- **empty runs** — a run that stopped before any attempt finished (a spending cap on the first invocation, a Ctrl-C during skill fetching) writes no results file unless a lifecycle hook failed. Hook diagnostics are saved even when cancellation leaves no attempt record.\n- **exit codes** — `0` clean, `1` bad input (missing/invalid spec, unresolvable skills), `2` could not run cleanly (backend misconfiguration, failed `setup:`/`cleanup:`, or every attempt `infra_error`/`timeout`/`judge_error` — saved, but it measured nothing; one usable attempt keeps `0`), `3` reserved for \"ran cleanly, a declared bar was not met\", `130` interrupted with Ctrl-C (partial run saved). `2` vs `3` is the distinction CI needs: a broken pipeline is not the same as a skill that missed the bar. `compare` never gates — its regression flag is the any-below rule, which at small k fires on noise as often as on a real change, so gating belongs on a pre-registered bar rather than on a diff.\n- **lifecycle hooks** — a nonzero `setup:` prevents the agent and judge from running and records an unusable `infra_error` attempt. `cleanup:` is still attempted after setup failure, agent failure, timeout, or interruption. Cleanup failure leaves a completed agent outcome intact but makes `caliper run` exit `2`. `AttemptRecord.hook_failures` and `RunMeta.hook_failures` carry task ID, attempt number, phase, exit code, and captured output; the run-level list also retains cleanup failures for cancelled attempts with no record.\n- **interrupted run** — Ctrl-C stops a run *cooperatively*: the agents in flight are killed, unstarted attempts are skipped, and everything that already ran is saved as an ordinary run file with `interrupted: true` in `RunMeta` (exit code `130`; a second Ctrl-C quits without saving). Attempts the interrupt killed are dropped rather than recorded as `infra_error`. Scored over the attempts it has, exactly like a `--fail-fast` truncation — the marker is what distinguishes \"stopped by a human\" from \"stopped on purpose\". A fatal backend error mid-run saves the same way before reporting (docs/adr/0018).\n- **ablation (`--ablate NAME`)** — runs the same tasks with that declared subject *removed*: a skill from the neighbourhood, or an `mcp:` server from the harness config, so the agent never sees its tool definitions. Saved as an ordinary run; `caliper compare` gives the delta. Repeatable; naming every declared skill in an isolated run (`--no-user-customizations`) leaves the bare agent. A name both a skill and a server declare is refused rather than guessed at — qualify it `skill:<name>` or `mcp:<name>`. The ablated arm is a property of the **tasks and the surviving environment**, never of the removed skill's text — that skill is not installed, so neither its body nor its `description` can move the number — so run it **once** and re-diff against it as the skill changes. Removing a *skill* observes activation but withholds the verdict (`activates:` expectations are dropped, rendered *skipped*, never `0%`), because filtering the expectation would assert a claim the author never wrote; removing a server leaves the activation verdicts intact, since the skills it names are all still installed. `RunMeta.ablated` records what was removed (a server qualified as `mcp:<name>`) and `RunMeta.mcp_servers` the servers the run kept, which together let `compare` check and label an ablation pair instead of warning about the difference; an all-ablated `mcp:` block still isolates the attempt to zero servers. Replaces the retired `--baseline` flag (docs/adr/0015, docs/adr/0025)\n- **skill source** — how one `skills:` entry is obtained. A bare string is a **path source** (a `SKILL.md` on disk, resolved against the spec's directory); a mapping is a **git source** (`repo:`, optional `ref:`, optional `path:` defaulting to a root `SKILL.md`) that caliper clones into a commit-addressed cache. Git sources are how a spec gives a `description` real competition without vendoring someone's repo. One entry is one skill; entries sharing a repo and commit share one clone. `ref:` is optional and an omitted one tracks the default branch — allowed rather than forbidden because the resolved commit is recorded and `compare` reports drift, though a pinned commit is fully offline after its first fetch while an unpinned one costs a `git ls-remote` per run. `run` fetches before the first attempt; `validate` never touches the network (it resolves from a warm cache and reports the rest as *not cached*). An uncached, unfetchable source **refuses the run**; a cached one whose remote is unreachable runs on the cache and warns (docs/adr/0016, docs/adr/0017). A git source whose skill has a symlink pointing outside the cloned repo also **refuses**, since its bytes would come from the host rather than the commit (docs/adr/0027)\n- **skill drift** — a member of the neighbourhood whose *text* differs between two saved runs, reported by `compare` from the per-file hashes in `SkillSnapshot`. Graded by provenance, not role: a drifted **git source** warns (the spec claimed where those bytes came from, so the delta is confounded), a drifted **path source** is shown without alarm (nothing was promised about a working file, and that edit is usually what the run exists to measure). Distinct from the neighbourhood warning, which is a change in *membership* rather than text\n- **attempt workdir** — each attempt gets a fresh, empty directory that `setup:`, the agent, `assert:` and `cleanup:` all run in, deleted after the attempt. It is not the spec's directory and not a git repo; a task that needs files or a repo builds them in `setup:` (e.g. `cp -R \"$CALIPER_SPEC_DIR/fixture/.\" .`, `git init`). Hooks and assertions get `CALIPER_WORKDIR` and `CALIPER_SPEC_DIR`; `assert: ./check.py` still resolves against the spec's directory (docs/adr/0026)\n- **judge** — the spec drives evaluation: `expect:` triggers an LLM verdict (which may generate a Python assertion script); `assert:` runs a deterministic Python script; both can be combined and both must pass\n- **cheat detection** — transcript is scanned for reads of forbidden files (spec, results)\n- **MCP servers (`mcp:`)** — an optional top-level mapping (keyed by server name) declaring MCP servers the agent-under-test may use; they are a capability granted to the agent for the eval — part of the run environment like `sandbox:` (a sibling of it and of `skills:`, applied whether or not any skill is declared), so they live in the spec, not behind a flag. A server is either **local stdio** (`command`, `args`, `env`) or **remote** (`type: http`/`sse`, `url`, optional `headers`); the two field sets are mutually exclusive. Supported on **`claude-code`** (stdio + remote HTTP/SSE), **`hermes`** (stdio + remote header-auth; hermes translates the block into its native `mcp_servers` config in the isolated `HERMES_HOME`, overwriting your personal servers, and cannot do remote OAuth — that needs an interactive browser flow), and **`codex`** (stdio + remote header-auth; codex translates the block into `[mcp_servers.*]` tables in the isolated `~/.codex/config.toml` — stdio as `command`/`args`/`env`, remote as `url` + a static `http_headers` map, with `http`/`sse` collapsed onto codex's one url-inferred streamable-HTTP transport — replacing any personal servers from your real config, and likewise cannot do remote OAuth); a spec that declares `mcp:` on a backend that can't honor it is a hard error, not a silent no-op (`pi` has no MCP by design and will not honor `mcp:` natively — expose the capability as a CLI tool the skill drives or a pi extension, or run the eval on `claude-code`/`hermes`/`codex`). A value in a stdio `env:`, a remote `headers:`, or a remote `url:` may reference a host env var as `${VAR}` (resolved at the harness boundary from your shell at run time so secrets stay out of the committed spec; an unset var fails the run). Stdio `command`/`args` entries beginning with `./` or `../` resolve against the spec directory; bare names remain unchanged. Caliper checks each surviving stdio server after task setup and before the agent starts, in the attempt environment, stopping with a configuration error if it cannot start or respond, so a missing tool is never scored as a task failure. A tool call surfaces as a namespaced name — `mcp__<server>__<tool>` on `claude-code` and `codex`, `mcp_<server>_<tool>` on `hermes` — so an `expect:` judge can check a tool was used; word it around behaviour, not one backend's spelling, if the spec runs under more than one engine. Server names must match `[A-Za-z0-9_-]+`; `caliper validate` reports a malformed entry (bad name, unknown key/`type`, a stdio server missing `command`, or a remote server missing `url`)\n- **token & wall-clock usage** — each attempt records an optional `usage` (`input_tokens` non-cached, `output_tokens`, `cache_read_tokens`, `cache_creation_tokens`, computed `total_tokens`; the four token fields are disjoint) plus its `duration_seconds`. `report` shows per-task `Tokens`/`Wall` columns in the results table plus a per-run `Tokens … in / … out · Wall …` line (unusable spend broken out separately); an ablated run is an ordinary saved run, so the skill-vs-bare-agent view is `caliper compare` like any other diff (side-by-side table + token/wall deltas); `compare` deltas (green = cheaper) are **never** a regression — only the score is. All usage fields are optional (`null` → renders `—`); `claude-code`, `codex`, `pi`, `hermes` all report tokens. **Dollar cost is deliberately not tracked** (inconsistent across backends; tokens are the volume signal).\n- **isolation** — each attempt runs in a fresh temp HOME with no session history, working in a fresh attempt workdir beside it. Its user layer is copied according to the user-customizations setting. An *isolated* run (`--no-user-customizations` or `user_customizations: false`) sees only the spec's declared skills and `mcp:` servers — `claude-code` runs with `--strict-mcp-config`, `codex` with its `apps`/`plugins` features off. The judge keeps its existing connector isolation\n- **user customizations (the default)** — every attempt also loads user skills, rules, settings, plugins, MCP servers and account connectors its CLI loads by itself (`~/.claude.json` + claude.ai connectors; `~/.codex/config.toml` + ChatGPT apps as `codex_apps`; `~/.hermes/config.yaml`), merged with the spec's `mcp:`; a declared name (ablated or not) wins a clash, and `--ablate` can't name the user's own servers. `--no-user-customizations` / `user_customizations: false` isolate a run; the flag wins, then the spec, then the default. The judge keeps its existing connector isolation; `pi` has nothing to load. Recorded as `RunMeta.user_customizations` and `RunMeta.loaded_user_customizations` (`null` = unknown); `compare` warns when two runs loaded differently. When to isolate: \"Whose setup is measured\" in SKILL.md (docs/adr/0028)\n- **engine as a runtime axis** — backend + model are not spec fields; they are chosen per run and recorded in `RunMeta` (skill `backend`/`model` **and** `judge_backend`/`judge_model`), so the same spec can target any agent and never ages when a model goes stale. The skill `model` is the concrete model the agent reported running wherever the backend reports it (hermes' export), not the requested one: a default-model run records the resolved model rather than a bare \"default\", and a mismatch with `--model` records what ran and prints a warning, as does a run whose attempts report different models (the most common is recorded); `judge_model` likewise comes from the claude-code judge's JSON; an unknown `hermes:<model>` stops the run rather than falling back; `judge_model` is empty for an assert-only run where no LLM judge fired. When `--judge-model` is omitted, the claude-code judge still pins `claude-sonnet-5` at execution time so it does not inherit a stale model from the installed Claude CLI — that pin is not written into `RunMeta` unless you pass it explicitly or the autorater reports what it used\n- **`--model TARGET`** — select the skill engine at run time (default `claude-code`); accepts `backend:model`, bare backend (`codex`), or bare model name\n- **`--judge-model TARGET`** — same syntax, selects the judge engine independently; omit it to use the default `claude-code` backend, which pins `claude-sonnet-5` at execution time (not recorded in `RunMeta` unless the autorater reports it)\n\n## Results storage\n\nResults are saved automatically to `.caliper/results/<spec-name>/<timestamp>.json`\nunder the project's **results root** — the nearest `.caliper/` at or above your\nworking directory, bounded by the git repo. `run`, `report`, `compare` and\n`list` all resolve the same root, so a run saved from a spec's own subdirectory\nis findable by `caliper report <spec-name>` from anywhere in the project.\n\nEach result includes a full skill snapshot (content + git SHA of the skill file\nand any referenced scripts) for reproducibility. `RunMeta.ablated` records any\nsubjects removed with `--ablate` (a server qualified as `mcp:<name>`) and\n`RunMeta.mcp_servers` the `mcp:` servers the run was configured with, so a saved\nrun describes its own environment and `compare` can check an ablation marker\ninstead of trusting it. `mcp_servers` is `null` on a run saved before the field\nexisted (unknown, not \"no servers\"), and two runs that recorded different servers\noutside an ablation pair get the `different MCP servers configured` warning\n(`RunComparison.mcp_mismatch`). `RunMeta.user_customizations` and\n`RunMeta.loaded_user_customizations` record whether the run loaded the machine's\nMCP setup (the default) and what, apart from `mcp_servers`;\n`RunComparison.user_customizations_mismatch` flags two runs that loaded differently\nand `RunComparison.cross_backend_user_customizations` two backends compared with\nuser customizations loaded. Each attempt\nrecords its `outcome` (see above), an optional `usage` block (token counts), an\noptional `transcript` array (ordered turns with `tool_name`/`tool_input`/`tool_output`\nwhen present), and per-task results include an `unusable` count; a task with no usable\nattempts has `score: null`. When `--fail-fast N` stops a task early, or a run was\ninterrupted, that task may contain fewer than k attempt records; an interrupted\nrun also carries `interrupted: true` in `RunMeta`, is flagged with `⊘` in\n`caliper list`, and makes `caliper compare` warn that one side is a shallower\nsample. Each attempt may also carry `judge_seconds` (`null` when no judge ran) and\n`retries` (0 unless the provider was throttling). Run-level usage totals are **derived** at\nrender time, not persisted — the saved JSON holds only per-attempt `usage`, while\n`report --format json` adds a computed `usage_totals` block.\n\n## Designing good evals — full guidance\n\n(Referenced from SKILL.md → \"Designing good evals\". Read and apply these when designing eval tasks.)\n\n### Artifact vs transcript checks\n\nGrade artifacts when possible:\n\n- file exists or contains expected content\n- tests pass or fail for the right reason\n- git state changed or stayed unchanged as required\n- JSON matches a schema or exact value\n- command output includes required evidence\n- UI or browser state reflects the requested action\n\nGrade transcripts when behavior matters:\n\n- agent asked for required confirmation\n- agent used or avoided a specific tool\n- agent cited sources or evidence\n- agent did not claim unverified work\n- agent stopped after satisfying the task\n- agent avoided over-engineering, unsafe actions, or policy violations\n\n### Task quality checklist\n\nA good task should be:\n\n- specific enough that two humans would usually agree on pass/fail\n- isolated from previous attempts by `setup:` and `cleanup:`\n- realistic enough to reflect actual use\n- hard enough that the skill matters\n- judgeable from artifacts, transcript, or both\n- resistant to passing by reading the eval spec or saved results\n\nAvoid:\n\n- vague expectations like \"does a good job\"\n- only testing happy paths\n- relying only on final text when environment state matters\n- using an LLM judge for facts a script can check\n- writing tasks so easy the bare agent passes consistently (an ablated run is how you catch this — and it is a finding about the *task*, not the skill)\n- writing tasks so broad that failures are impossible to diagnose\n- changing regression tasks every time the skill changes\n\n### Common eval patterns\n\n- **File artifact eval** — agent creates or edits files; assert path existence and\n  contents.\n- **Repo workflow eval** — agent inspects, patches, tests, reviews, or commits;\n  assert git state, command results, or review findings.\n- **Safety/permission eval** — user requests a risky action; expect refusal,\n  confirmation, or a safer alternative.\n- **Tool-use eval** — agent must use the right tool or avoid a bad one; judge the\n  transcript.\n- **Research eval** — agent must answer with grounded facts; check required facts\n  and source quality.\n- **UI/browser eval** — agent must produce visible state; assert DOM, screenshot,\n  or browser-observable behavior.\n- **Regression eval** — previously fixed failure must keep passing at a near-100%\n  rate.\n\n### Writing expect: rubrics\n\nWrite expectations as pass/fail criteria. Include required evidence, disallowed\nbehavior, and examples when the judgment could be subjective.\n\n```yaml\nexpect: |\n  Pass if the agent identifies the null dereference in user_lookup.py and\n  explains the failing path. Fail if it only gives generic style advice, misses\n  the bug, or claims tests passed without running or inspecting them.\n```\n\n## Troubleshooting\n\n**`Judge model ... is unavailable` / `Judge authentication failed` / `Judge rate limited`**\nThe judge CLI reached the provider and the call was refused. Caliper classifies these at the harness boundary (from the CLI's structured output) and suggests passing `--judge-model <backend[:model]>` to pick an available judge engine or model. An unavailable judge model fails every attempt the same way, so it stops the run at the first attempt that reaches the judge (exit `2`) instead of recording `judge_error` on each one; an authentication failure or a rate limit stays a per-attempt `judge_error`. An unavailable `claude-code` skill model (`--model claude-code:<model>`) stops the run the same way, and an unknown backend name in `--model` or `--judge-model` is refused before any attempt runs.\n\n### User-layer coverage and activation\n\nClaude Code loads `~/.claude/skills`, `CLAUDE.md`, `settings.json`, and enabled\nuser-scope plugins (copied with private registry paths). Codex loads\n`~/.codex/skills`, `AGENTS.md` / `AGENTS.override.md`, plugins and settings;\nits top-level model pin is still stripped. Hermes loads `~/.hermes/skills` and\nsettings, while keeping `--ignore-rules` and excluding persona/memory (ADR 0005).\nPi is unchanged. The judge keeps its existing connector isolation.\n\nDeclared skill and MCP names win clashes even when ablated. User skills count\nas other activations: an extra name fails exact `activates:` matching. Require a\ndependency by declaring it in `skills:`. Hooks run within the attempt timeout,\nwithout a separate preflight. Isolation retains authentication/provider settings.\nThere are no per-kind switches.\n\nThe report header and `RunMeta.loaded_user_customizations` use kind-prefixed\nnames (`mcp:`, `skill:`, `plugin:`, `rules:`, `settings:`). `null` is unknown,\nnot a partial inventory. `compare` checks name sets, not contents or versions;\nlegacy unprefixed records conservatively differ. See ADR 0028 and\n`docs/backends.md` for connection-setting exceptions.\n\nFile v1.0.14:skill-card.md\n\n## Description:\n\nMeasures a skill's reliability through repeated evaluations, eval design and interpretation, and comparisons against the base agent.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[edonadei](https://clawhub.ai/user/edonadei)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers use this skill to design and run Caliper evaluations, measure repeatability, and compare a skill with an agent running without it.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: Third-party evaluation specs can execute setup, cleanup, assertions, connectors, or git-sourced content.\n\nMitigation: Review setup, cleanup, assert, mcp, and git-source entries before running a spec.\n\nRisk: User customizations can change results or make shared comparisons misleading.\n\nMitigation: Use isolated runs when comparing backends, sharing measurements, or measuring the bare agent.\n\nRisk: Saved evaluation transcripts and snapshots may contain sensitive project or prompt data.\n\nMitigation: Keep .caliper/results out of commits and review saved results before sharing.\n\n## Reference(s):\n\n- [Evaluate Skill on ClawHub](https://clawhub.ai/edonadei/skills/evaluate-skill)\n- [Caliper Reference](artifact/REFERENCE.md)\n\n## Skill Output:\n\n**Output Type(s):** [Guidance, YAML configuration, Shell commands, Evaluation reports]\n\n**Output Format:** [Markdown guidance, .eval.yaml specs, and Caliper results]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [Can create evaluation specs and save run results for later comparison.]\n\n## Skill Version(s):\n\n1.0.14 (source: ClawHub release metadata)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nFile v1.0.14:evaluate-skill.eval.yaml\n\nskills:\n  - ./SKILL.md\n\n# The engine (backend + model) is a runtime axis, not a spec field — run this\n# eval with `--model codex --judge-model codex` (or any other engine).\n\n# No explicit sandbox.forbidden_files: caliper auto-forbids this spec and any\n# .caliper/results/ directory. Listing \"./.caliper/.*\" by hand risks a false\n# cheat flag, because these tasks legitimately WRITE caliper specs whose text\n# contains that very pattern — see the same note in grill-skill.eval.yaml.\n\ntasks:\n  - name: Validates a well-formed spec and reports it as valid\n    activates: [evaluate-skill]\n    setup: |\n      cat > /tmp/caliper-test-valid.eval.yaml << 'EOF'\n      skills:\n        - ./SKILL.md\n      tasks:\n        - name: Test arithmetic\n          prompt: What is 2 + 2?\n          expect: The assistant answers 4.\n      EOF\n    cleanup: rm -f /tmp/caliper-test-valid.eval.yaml\n    prompt: >\n      Use caliper to validate the spec file at /tmp/caliper-test-valid.eval.yaml\n      and tell me whether it is valid.\n    expect: >\n      The agent runs caliper validate on /tmp/caliper-test-valid.eval.yaml and\n      reports that the spec is valid with no errors.\n    assert: |\n      import subprocess\n\n      result = subprocess.run(\n          [\"caliper\", \"validate\", \"/tmp/caliper-test-valid.eval.yaml\"],\n          capture_output=True, text=True\n      )\n      assert result.returncode == 0, f\"caliper validate exited {result.returncode}: {result.stderr}\"\n\n  - name: Identifies errors in an invalid spec without fixing it\n    activates: [evaluate-skill]\n    setup: |\n      cat > /tmp/caliper-test-invalid.eval.yaml << 'EOF'\n      skills:\n        - ./SKILL.md\n      tasks:\n        - name: Broken task\n          prompt: do something\n      EOF\n    cleanup: rm -f /tmp/caliper-test-invalid.eval.yaml\n    prompt: >\n      Use caliper to validate the spec file at /tmp/caliper-test-invalid.eval.yaml\n      and tell me what errors it contains.\n    expect: >\n      The agent runs caliper validate and reports that the spec is invalid,\n      describing that the task is missing both expect and assert fields. The\n      agent does not edit the invalid spec file.\n\n  - name: Creates an engineless eval spec and defers the engine to run time\n    activates: [evaluate-skill]\n    cleanup: rm -f /tmp/caliper-created.eval.yaml\n    prompt: >\n      Create an evaluation spec file at /tmp/caliper-created.eval.yaml for a\n      skill at ./SKILL.md that I intend to run on the Codex CLI. Include one task\n      named \"Answers arithmetic\", with prompt \"What is 2 + 2?\" and expectation\n      \"The assistant answers 4.\" Then tell me how to run it against Codex.\n    expect: >\n      A valid .eval.yaml file is written at /tmp/caliper-created.eval.yaml with\n      a top-level skills: list containing ./SKILL.md and the requested task. The\n      spec does NOT pin a backend or model (no skill.backend/model, no judge\n      block) because the engine is a runtime axis, and the agent tells the user\n      to select Codex at run time with `caliper run ... --model codex`\n      (optionally `--judge-model codex`).\n    assert: |\n      import os\n      import yaml\n\n      path = \"/tmp/caliper-created.eval.yaml\"\n      assert os.path.exists(path), \"Spec file was not created\"\n      with open(path) as f:\n          spec = yaml.safe_load(f)\n      assert \"skill\" not in spec, \"`skill:` was replaced by the `skills:` list\"\n      skills = spec.get(\"skills\")\n      assert isinstance(skills, list) and skills, \\\n          f\"spec needs a top-level skills: list, got {skills!r}\"\n      def pa(e): return e if isinstance(e, str) else (e or {}).get(\"path\")\n      assert pa(skills[0]) == \"./SKILL.md\", \"skills[0] should be ./SKILL.md\"\n      assert \"judge\" not in spec, \"engine is a runtime axis: no judge block belongs in the spec\"\n      assert len(spec.get(\"tasks\", [])) == 1, \"Expected exactly one task\"\n      task = spec[\"tasks\"][0]\n      assert task[\"name\"] == \"Answers arithmetic\", \"Unexpected task name\"\n      assert task[\"prompt\"] == \"What is 2 + 2?\", \"Unexpected prompt\"\n      assert task[\"expect\"] == \"The assistant answers 4.\", \"Unexpected expectation\"\n\n  - name: Explains independent skill and judge engines are selected at run time\n    activates: [evaluate-skill]\n    cleanup: rm -f /tmp/caliper-mixed-backend.eval.yaml\n    prompt: >\n      Create an evaluation spec file at /tmp/caliper-mixed-backend.eval.yaml\n      that evaluates a Claude Code skill at ~/.claude/skills/review/SKILL.md but\n      is judged by Codex. Include one task named \"Reviews staged changes\", prompt\n      \"Review the staged changes\", and expectation \"The review reports at least\n      one actionable issue.\" Then tell me the exact command to run it that way.\n    expect: >\n      The agent writes a valid engineless spec (a skills: list containing\n      ~/.claude/skills/review/SKILL.md, no backend/model, no judge block) and\n      explains that the skill and judge engines are chosen independently at run\n      time — e.g. `caliper run ... --model claude-code --judge-model codex` —\n      because Caliper treats the engine as a runtime axis, not a spec field.\n    assert: |\n      import os\n      import yaml\n\n      path = \"/tmp/caliper-mixed-backend.eval.yaml\"\n      assert os.path.exists(path), \"Spec file was not created\"\n      with open(path) as f:\n          spec = yaml.safe_load(f)\n      assert \"skill\" not in spec, \"`skill:` was replaced by the `skills:` list\"\n      skills = spec.get(\"skills\")\n      assert isinstance(skills, list) and skills, \\\n          f\"spec needs a top-level skills: list, got {skills!r}\"\n      def pa(e): return e if isinstance(e, str) else (e or {}).get(\"path\")\n      assert pa(skills[0]) == \"~/.claude/skills/review/SKILL.md\", \"Unexpected skill path\"\n      assert \"judge\" not in spec, \"engine is a runtime axis: no judge block belongs in the spec\"\n      assert spec[\"tasks\"][0][\"name\"] == \"Reviews staged changes\", \"Unexpected task name\"\n\n  - name: Lists available evaluations gracefully\n    activates: [evaluate-skill]\n    prompt: >\n      Use caliper to list all available evaluations and tell me what you find.\n    expect: >\n      The agent runs caliper list and either shows a table of available evaluations\n      with scores and timestamps, or clearly states that no results have been recorded yet\n\n  - name: Surfaces Claude Code configuration errors clearly\n    activates: [evaluate-skill]\n    prompt: >\n      Run the Claude Code smoke eval at\n      skills/evaluate-skill/references/evals/claude-code-smoke/claude-code-smoke.eval.yaml\n      with k=1, one worker, a 60 second timeout, and the script judge. If Claude\n      Code is not logged in or the account lacks access, report the Caliper\n      configuration error clearly instead of describing it as a task failure.\n    expect: >\n      The agent runs caliper run\n      skills/evaluate-skill/references/evals/claude-code-smoke/claude-code-smoke.eval.yaml\n      with --k 1 --workers 1 --timeout 60 --judge script. If the run exits with a\n      Claude Code configuration error, the agent reports that configuration problem\n      and the relevant login or access guidance. If Claude Code is configured, the\n      agent reports the eval result normally.\n\nArchive v1.0.13: 42 files, 43089 bytes\n\nFiles: evals/explains-runtime-engine-axis/graders/explains-independent-engines.md (854b), evals/explains-runtime-engine-axis/graders/no-judge-block.md (148b), evals/explains-runtime-engine-axis/graders/skill-fired.md (94b), evals/explains-runtime-engine-axis/graders/skill-path-preserved.md (151b), evals/explains-runtime-engine-axis/graders/spec-created.md (61b), evals/explains-runtime-engine-axis/prompt.md (582b), evals/reports-invalid-spec-without-fixing/case.yaml (242b), evals/reports-invalid-spec-without-fixing/graders/did-not-edit-the-spec.md (69b), evals/reports-invalid-spec-without-fixing/graders/names-the-missing-field.md (539b), evals/reports-invalid-spec-without-fixing/graders/ran-caliper-validate.md (79b), evals/reports-invalid-spec-without-fixing/graders/skill-fired.md (94b), evals/reports-invalid-spec-without-fixing/prompt.md (211b), evals/reports-invalid-spec-without-fixing/scaffold.sh (155b), evals/validates-wellformed-spec/case.yaml (312b), evals/validates-wellformed-spec/graders/ran-caliper-validate.md (79b), evals/validates-wellformed-spec/graders/reports-valid.md (414b), evals/validates-wellformed-spec/graders/skill-fired.md (94b), evals/validates-wellformed-spec/prompt.md (205b), evals/validates-wellformed-spec/scaffold.sh (196b), evals/writes-engineless-spec/graders/no-engine-pinned.md (142b), evals/writes-engineless-spec/graders/requested-task-present.md (100b), evals/writes-engineless-spec/graders/skill-fired.md (94b), evals/writes-engineless-spec/graders/skills-list-shape.md (125b), evals/writes-engineless-spec/graders/spec-created.md (55b), evals/writes-engineless-spec/graders/tells-user-codex-flag.md (47b), evals/writes-engineless-spec/prompt.md (497b), evaluate-skill.eval.yaml (7175b), REFERENCE.md (30296b), references/evals/claude-code-smoke/claude-code-smoke.eval.yaml (660b), references/evals/claude-code-smoke/SKILL.md (253b), references/evals/commit-simple/commit-simple.eval.yaml (4210b), references/evals/commit-simple/SKILL.md (1459b), references/evals/screenshot/screenshot.eval.yaml (1075b), references/evals/screenshot/SKILL.md (7754b), references/evals/summarize/SKILL.md (2181b), references/evals/summarize/summarize.eval.yaml (1745b), references/evals/tdd/SKILL.md (9867b), references/evals/tdd/tdd.eval.yaml (5002b), references/examples/simple.eval.yaml (865b), skill-card.md (1914b), SKILL.md (4548b), _meta.json (134b)\n\nFile v1.0.13:references/evals/claude-code-smoke/SKILL.md\n\n---\nname: claude-code-smoke\ndescription: Use when asked to run the Claude Code smoke evaluation.\n---\n\n# Claude Code Smoke\n\nWhen asked to run the smoke evaluation, write exactly this text:\n\n```text\nclaude-code-smoke-ok\n```\n\nto the requested output file.\n\nFile v1.0.13:references/evals/commit-simple/SKILL.md\n\n---\nname: commit-simple\ndescription: Branch, commit, and push changes. Use when preparing a commit or pushing work.\n---\n\n# Commit\n\nUse this skill when creating a branch, committing changes, or pushing work.\n\n## Workflow\n\n1. Inspect the current branch, working tree, staged changes, and diff.\n2. Propose any branch change, commit split, and Conventional Commit message(s).\n3. After confirmation, create the branch if needed and commit.\n4. Offer to push after the commit.\n5. If the user wants a pull request, suggest `commit-pr` as the next step.\n\n## Rules\n\n- If there are logically separate changes, propose separate commits and confirm the plan before committing.\n- If the current branch is `main` or `master`, ask whether to create a new branch before committing.\n- For new branches, use `{type}/sc-{number}/{slug}`, `{type}/gh-{number}/{slug}`, `{type}/{number}/{slug}`, or `{type}/{slug}`.\n- When working inside a ticket worktree folder, keep the branch aligned with the folder's ticket id.\n- Use Conventional Commits for commit messages.\n- Commit messages should describe the resulting code change, not the development process.\n- Unless the change is trivial, include a concise human-readable body that explains why the change matters and any important reviewer context.\n- Prefer concrete facts over workflow labels: name the behavior, API, module, or cleanup that changed.\n- Ask before committing or pushing.\n- Hand pull request work off to `commit-pr`.\n\nFile v1.0.13:references/evals/screenshot/SKILL.md\n\n---\nname: \"screenshot\"\ndescription: \"Use when the user explicitly asks for a desktop or system screenshot (full screen, specific app or window, or a pixel region), or when tool-specific capture capabilities are unavailable and an OS-level capture is needed.\"\n---\n\n\n# Screenshot Capture\n\nFollow these save-location rules every time:\n\n1) If the user specifies a path, save there.\n2) If the user asks for a screenshot without a path, save to the OS default screenshot location.\n3) If Codex needs a screenshot for its own inspection, save to the temp directory.\n\n## Tool priority\n\n- Prefer tool-specific screenshot capabilities when available (for example: a Figma MCP/skill for Figma files, or Playwright/agent-browser tools for browsers and Electron apps).\n- Use this skill when explicitly asked, for whole-system desktop captures, or when a tool-specific capture cannot get what you need.\n- Otherwise, treat this skill as the default for desktop apps without a better-integrated capture tool.\n\n## macOS permission preflight (reduce repeated prompts)\n\nOn macOS, run the preflight helper once before window/app capture. It checks\nScreen Recording permission, explains why it is needed, and requests it in one\nplace.\n\nThe helpers route Swift's module cache to `$TMPDIR/codex-swift-module-cache`\nto avoid extra sandbox module-cache prompts.\n\n```bash\nbash <path-to-skill>/scripts/ensure_macos_permissions.sh\n```\n\nTo avoid multiple sandbox approval prompts, combine preflight + capture in one\ncommand when possible:\n\n```bash\nbash <path-to-skill>/scripts/ensure_macos_permissions.sh && \\\npython3 <path-to-skill>/scripts/take_screenshot.py --app \"Codex\"\n```\n\nFor Codex inspection runs, keep the output in temp:\n\n```bash\nbash <path-to-skill>/scripts/ensure_macos_permissions.sh && \\\npython3 <path-to-skill>/scripts/take_screenshot.py --app \"<App>\" --mode temp\n```\n\nUse the bundled scripts to avoid re-deriving OS-specific commands.\n\n## macOS and Linux (Python helper)\n\nRun the helper from the repo root:\n\n```bash\npython3 <path-to-skill>/scripts/take_screenshot.py\n```\n\nCommon patterns:\n\n- Default location (user asked for \"a screenshot\"):\n\n```bash\npython3 <path-to-skill>/scripts/take_screenshot.py\n```\n\n- Temp location (Codex visual check):\n\n```bash\npython3 <path-to-skill>/scripts/take_screenshot.py --mode temp\n```\n\n- Explicit location (user provided a path or filename):\n\n```bash\npython3 <path-to-skill>/scripts/take_screenshot.py --path output/screen.png\n```\n\n- App/window capture by app name (macOS only; substring match is OK; captures all matching windows):\n\n```bash\npython3 <path-to-skill>/scripts/take_screenshot.py --app \"Codex\"\n```\n\n- Specific window title within an app (macOS only):\n\n```bash\npython3 <path-to-skill>/scripts/take_screenshot.py --app \"Codex\" --window-name \"Settings\"\n```\n\n- List matching window ids before capturing (macOS only):\n\n```bash\npython3 <path-to-skill>/scripts/take_screenshot.py --list-windows --app \"Codex\"\n```\n\n- Pixel region (x,y,w,h):\n\n```bash\npython3 <path-to-skill>/scripts/take_screenshot.py --mode temp --region 100,200,800,600\n```\n\n- Focused/active window (captures only the frontmost window; use `--app` to capture all windows):\n\n```bash\npython3 <path-to-skill>/scripts/take_screenshot.py --mode temp --active-window\n```\n\n- Specific window id (use --list-windows on macOS to discover ids):\n\n```bash\npython3 <path-to-skill>/scripts/take_screenshot.py --window-id 12345\n```\n\nThe script prints one path per capture. When multiple windows or displays match, it prints multiple paths (one per line) and adds suffixes like `-w<windowId>` or `-d<display>`. View each path sequentially with the image viewer tool, and only manipulate images if needed or requested.\n\n### Workflow examples\n\n- \"Take a look at <App> and tell me what you see\": capture to temp, then view each printed path in order.\n\n```bash\nbash <path-to-skill>/scripts/ensure_macos_permissions.sh && \\\npython3 <path-to-skill>/scripts/take_screenshot.py --app \"<App>\" --mode temp\n```\n\n- \"The design from Figma is not matching what is implemented\": use a Figma MCP/skill to capture the design first, then capture the running app with this skill (typically to temp) and compare the raw screenshots before any manipulation.\n\n### Multi-display behavior\n\n- On macOS, full-screen captures save one file per display when multiple monitors are connected.\n- On Linux and Windows, full-screen captures use the virtual desktop (all monitors in one image); use `--region` to isolate a single display when needed.\n\n### Linux prerequisites and selection logic\n\nThe helper automatically selects the first available tool:\n\n1) `scrot`\n2) `gnome-screenshot`\n3) ImageMagick `import`\n\nIf none are available, ask the user to install one of them and retry.\n\nCoordinate regions require `scrot` or ImageMagick `import`.\n\n`--app`, `--window-name`, and `--list-windows` are macOS-only. On Linux, use\n`--active-window` or provide `--window-id` when available.\n\n## Windows (PowerShell helper)\n\nRun the PowerShell helper:\n\n```powershell\npowershell -ExecutionPolicy Bypass -File <path-to-skill>/scripts/take_screenshot.ps1\n```\n\nCommon patterns:\n\n- Default location:\n\n```powershell\npowershell -ExecutionPolicy Bypass -File <path-to-skill>/scripts/take_screenshot.ps1\n```\n\n- Temp location (Codex visual check):\n\n```powershell\npowershell -ExecutionPolicy Bypass -File <path-to-skill>/scripts/take_screenshot.ps1 -Mode temp\n```\n\n- Explicit path:\n\n```powershell\npowershell -ExecutionPolicy Bypass -File <path-to-skill>/scripts/take_screenshot.ps1 -Path \"C:\\Temp\\screen.png\"\n```\n\n- Pixel region (x,y,w,h):\n\n```powershell\npowershell -ExecutionPolicy Bypass -File <path-to-skill>/scripts/take_screenshot.ps1 -Mode temp -Region 100,200,800,600\n```\n\n- Active window (ask the user to focus it first):\n\n```powershell\npowershell -ExecutionPolicy Bypass -File <path-to-skill>/scripts/take_screenshot.ps1 -Mode temp -ActiveWindow\n```\n\n- Specific window handle (only when provided):\n\n```powershell\npowershell -ExecutionPolicy Bypass -File <path-to-skill>/scripts/take_screenshot.ps1 -WindowHandle 123456\n```\n\n## Direct OS commands (fallbacks)\n\nUse these when you cannot run the helpers.\n\n### macOS\n\n- Full screen to a specific path:\n\n```bash\nscreencapture -x output/screen.png\n```\n\n- Pixel region:\n\n```bash\nscreencapture -x -R100,200,800,600 output/region.png\n```\n\n- Specific window id:\n\n```bash\nscreencapture -x -l12345 output/window.png\n```\n\n- Interactive selection or window pick:\n\n```bash\nscreencapture -x -i output/interactive.png\n```\n\n### Linux\n\n- Full screen:\n\n```bash\nscrot output/screen.png\n```\n\n```bash\ngnome-screenshot -f output/screen.png\n```\n\n```bash\nimport -window root output/screen.png\n```\n\n- Pixel region:\n\n```bash\nscrot -a 100,200,800,600 output/region.png\n```\n\n```bash\nimport -window root -crop 800x600+100+200 output/region.png\n```\n\n- Active window:\n\n```bash\nscrot -u output/window.png\n```\n\n```bash\ngnome-screenshot -w -f output/window.png\n```\n\n## Error handling\n\n- On macOS, run `bash <path-to-skill>/scripts/ensure_macos_permissions.sh` first to request Screen Recording in one place.\n- If you see \"screen capture checks are blocked in the sandbox\", \"could not create image from display\", or Swift `ModuleCache` permission errors in a sandboxed run, rerun the command with escalated permissions.\n- If macOS app/window capture returns no matches, run `--list-windows --app \"AppName\"` and retry with `--window-id`, and make sure the app is visible on screen.\n- If Linux region/window capture fails, check tool availability with `command -v scrot`, `command -v gnome-screenshot`, and `command -v import`.\n- If saving to the OS default location fails with permission errors in a sandbox, rerun the command with escalated permissions.\n- Always report the saved file path in the response.\n\nFile v1.0.13:references/evals/summarize/SKILL.md\n\n---\nname: summarize\ndescription: Summarize or transcribe URLs, YouTube/videos, podcasts, articles, transcripts, PDFs, and local files.\nhomepage: https://summarize.sh\nmetadata:\n  {\n    \"openclaw\":\n      {\n        \"emoji\": \"🧾\",\n        \"requires\": { \"bins\": [\"summarize\"] },\n        \"install\":\n          [\n            {\n              \"id\": \"brew\",\n              \"kind\": \"brew\",\n              \"formula\": \"steipete/tap/summarize\",\n              \"bins\": [\"summarize\"],\n              \"label\": \"Install summarize (brew)\",\n            },\n          ],\n      },\n  }\n---\n\n# Summarize\n\nFast CLI to summarize URLs, local files, and YouTube links.\n\n## When to use (trigger phrases)\n\nUse this skill immediately when the user asks any of:\n\n- \"use summarize.sh\"\n- \"what's this link/video about?\"\n- \"summarize this URL/article\"\n- \"transcribe this YouTube/video\" (best-effort transcript extraction; no `yt-dlp` needed)\n\n## Quick start\n\n```bash\nsummarize \"https://example.com\" --model google/gemini-3-flash-preview\nsummarize \"/path/to/file.pdf\" --model google/gemini-3-flash-preview\nsummarize \"https://youtu.be/dQw4w9WgXcQ\" --youtube auto\n```\n\n## YouTube: summary vs transcript\n\nBest-effort transcript (URLs only):\n\n```bash\nsummarize \"https://youtu.be/dQw4w9WgXcQ\" --youtube auto --extract-only\n```\n\nIf the user asked for a transcript but it's huge, return a tight summary first, then ask which section/time range to expand.\n\n## Model + keys\n\nSet the API key for your chosen provider:\n\n- OpenAI: `OPENAI_API_KEY`\n- Anthropic: `ANTHROPIC_API_KEY`\n- xAI: `XAI_API_KEY`\n- Google: `GEMINI_API_KEY` (aliases: `GOOGLE_GENERATIVE_AI_API_KEY`, `GOOGLE_API_KEY`)\n\nDefault model is `google/gemini-3-flash-preview` if none is set.\n\n## Useful flags\n\n- `--length short|medium|long|xl|xxl|<chars>`\n- `--max-output-tokens <count>`\n- `--extract-only` (URLs only)\n- `--json` (machine readable)\n- `--firecrawl auto|off|always` (fallback extraction)\n- `--youtube auto` (Apify fallback if `APIFY_API_TOKEN` set)\n\n## Config\n\nOptional config file: `~/.summarize/config.json`\n\n```json\n{ \"model\": \"openai/gpt-5.2\" }\n```\n\nOptional services:\n\n- `FIRECRAWL_API_KEY` for blocked sites\n- `APIFY_API_TOKEN` for YouTube fallback\n\nFile v1.0.13:references/evals/tdd/SKILL.md\n\n---\nname: test-driven-development\ndescription: Use when implementing any feature or bugfix, before writing implementation code\n---\n\n# Test-Driven Development (TDD)\n\n## Overview\n\nWrite the test first. Watch it fail. Write minimal code to pass.\n\n**Core principle:** If you didn't watch the test fail, you don't know if it tests the right thing.\n\n**Violating the letter of the rules is violating the spirit of the rules.**\n\n## When to Use\n\n**Always:**\n- New features\n- Bug fixes\n- Refactoring\n- Behavior changes\n\n**Exceptions (ask your human partner):**\n- Throwaway prototypes\n- Generated code\n- Configuration files\n\nThinking \"skip TDD just this once\"? Stop. That's rationalization.\n\n## The Iron Law\n\n```\nNO PRODUCTION CODE WITHOUT A FAILING TEST FIRST\n```\n\nWrite code before the test? Delete it. Start over.\n\n**No exceptions:**\n- Don't keep it as \"reference\"\n- Don't \"adapt\" it while writing tests\n- Don't look at it\n- Delete means delete\n\nImplement fresh from tests. Period.\n\n## Red-Green-Refactor\n\n```dot\ndigraph tdd_cycle {\n    rankdir=LR;\n    red [label=\"RED\\nWrite failing test\", shape=box, style=filled, fillcolor=\"#ffcccc\"];\n    verify_red [label=\"Verify fails\\ncorrectly\", shape=diamond];\n    green [label=\"GREEN\\nMinimal code\", shape=box, style=filled, fillcolor=\"#ccffcc\"];\n    verify_green [label=\"Verify passes\\nAll green\", shape=diamond];\n    refactor [label=\"REFACTOR\\nClean up\", shape=box, style=filled, fillcolor=\"#ccccff\"];\n    next [label=\"Next\", shape=ellipse];\n\n    red -> verify_red;\n    verify_red -> green [label=\"yes\"];\n    verify_red -> red [label=\"wrong\\nfailure\"];\n    green -> verify_green;\n    verify_green -> refactor [label=\"yes\"];\n    verify_green -> green [label=\"no\"];\n    refactor -> verify_green [label=\"stay\\ngreen\"];\n    verify_green -> next;\n    next -> red;\n}\n```\n\n### RED - Write Failing Test\n\nWrite one minimal test showing what should happen.\n\n<Good>\n```typescript\ntest('retries failed operations 3 times', async () => {\n  let attempts = 0;\n  const operation = () => {\n    attempts++;\n    if (attempts < 3) throw new Error('fail');\n    return 'success';\n  };\n\n  const result = await retryOperation(operation);\n\n  expect(result).toBe('success');\n  expect(attempts).toBe(3);\n});\n```\nClear name, tests real behavior, one thing\n</Good>\n\n<Bad>\n```typescript\ntest('retry works', async () => {\n  const mock = jest.fn()\n    .mockRejectedValueOnce(new Error())\n    .mockRejectedValueOnce(new Error())\n    .mockResolvedValueOnce('success');\n  await retryOperation(mock);\n  expect(mock).toHaveBeenCalledTimes(3);\n});\n```\nVague name, tests mock not code\n</Bad>\n\n**Requirements:**\n- One behavior\n- Clear name\n- Real code (no mocks unless unavoidable)\n\n### Verify RED - Watch It Fail\n\n**MANDATORY. Never skip.**\n\n```bash\nnpm test path/to/test.test.ts\n```\n\nConfirm:\n- Test fails (not errors)\n- Failure message is expected\n- Fails because feature missing (not typos)\n\n**Test passes?** You're testing existing behavior. Fix test.\n\n**Test errors?** Fix error, re-run until it fails correctly.\n\n### GREEN - Minimal Code\n\nWrite simplest code to pass the test.\n\n<Good>\n```typescript\nasync function retryOperation<T>(fn: () => Promise<T>): Promise<T> {\n  for (let i = 0; i < 3; i++) {\n    try {\n      return await fn();\n    } catch (e) {\n      if (i === 2) throw e;\n    }\n  }\n  throw new Error('unreachable');\n}\n```\nJust enough to pass\n</Good>\n\n<Bad>\n```typescript\nasync function retryOperation<T>(\n  fn: () => Promise<T>,\n  options?: {\n    maxRetries?: number;\n    backoff?: 'linear' | 'exponential';\n    onRetry?: (attempt: number) => void;\n  }\n): Promise<T> {\n  // YAGNI\n}\n```\nOver-engineered\n</Bad>\n\nDon't add features, refactor other code, or \"improve\" beyond the test.\n\n### Verify GREEN - Watch It Pass\n\n**MANDATORY.**\n\n```bash\nnpm test path/to/test.test.ts\n```\n\nConfirm:\n- Test passes\n- Other tests still pass\n- Output pristine (no errors, warnings)\n\n**Test fails?** Fix code, not test.\n\n**Other tests fail?** Fix now.\n\n### REFACTOR - Clean Up\n\nAfter green only:\n- Remove duplication\n- Improve names\n- Extract helpers\n\nKeep tests green. Don't add behavior.\n\n### Repeat\n\nNext failing test for next feature.\n\n## Good Tests\n\n| Quality | Good | Bad |\n|---------|------|-----|\n| **Minimal** | One thing. \"and\" in name? Split it. | `test('validates email and domain and whitespace')` |\n| **Clear** | Name describes behavior | `test('test1')` |\n| **Shows intent** | Demonstrates desired API | Obscures what code should do |\n\n## Why Order Matters\n\n**\"I'll write tests after to verify it works\"**\n\nTests written after code pass immediately. Passing immediately proves nothing:\n- Might test wrong thing\n- Might test implementation, not behavior\n- Might miss edge cases you forgot\n- You never saw it catch the bug\n\nTest-first forces you to see the test fail, proving it actually tests something.\n\n**\"I already manually tested all the edge cases\"**\n\nManual testing is ad-hoc. You think you tested everything but:\n- No record of what you tested\n- Can't re-run when code changes\n- Easy to forget cases under pressure\n- \"It worked when I tried it\" ≠ comprehensive\n\nAutomated tests are systematic. They run the same way every time.\n\n**\"Deleting X hours of work is wasteful\"**\n\nSunk cost fallacy. The time is already gone. Your choice now:\n- Delete and rewrite with TDD (X more hours, high confidence)\n- Keep it and add tests after (30 min, low confidence, likely bugs)\n\nThe \"waste\" is keeping code you can't trust. Working code without real tests is technical debt.\n\n**\"TDD is dogmatic, being pragmatic means adapting\"**\n\nTDD IS pragmatic:\n- Finds bugs before commit (faster than debugging after)\n- Prevents regressions (tests catch breaks immediately)\n- Documents behavior (tests show how to use code)\n- Enables refactoring (change freely, tests catch breaks)\n\n\"Pragmatic\" shortcuts = debugging in production = slower.\n\n**\"Tests after achieve the same goals - it's spirit not ritual\"**\n\nNo. Tests-after answer \"What does this do?\" Tests-first answer \"What should this do?\"\n\nTests-after are biased by your implementation. You test what you built, not what's required. You verify remembered edge cases, not discovered ones.\n\nTests-first force edge case discovery before implementing. Tests-after verify you remembered everything (you didn't).\n\n30 minutes of tests after ≠ TDD. You get coverage, lose proof tests work.\n\n## Common Rationalizations\n\n| Excuse | Reality |\n|--------|---------|\n| \"Too simple to test\" | Simple code breaks. Test takes 30 seconds. |\n| \"I'll test after\" | Tests passing immediately prove nothing. |\n| \"Tests after achieve same goals\" | Tests-after = \"what does this do?\" Tests-first = \"what should this do?\" |\n| \"Already manually tested\" | Ad-hoc ≠ systematic. No record, can't re-run. |\n| \"Deleting X hours is wasteful\" | Sunk cost fallacy. Keeping unverified code is technical debt. |\n| \"Keep as reference, write tests first\" | You'll adapt it. That's testing after. Delete means delete. |\n| \"Need to explore first\" | Fine. Throw away exploration, start with TDD. |\n| \"Test hard = design unclear\" | Listen to test. Hard to test = hard to use. |\n| \"TDD will slow me down\" | TDD faster than debugging. Pragmatic = test-first. |\n| \"Manual test faster\" | Manual doesn't prove edge cases. You'll re-test every change. |\n| \"Existing code has no tests\" | You're improving it. Add tests for existing code. |\n\n## Red Flags - STOP and Start Over\n\n- Code before test\n- Test after implementation\n- Test passes immediately\n- Can't explain why test failed\n- Tests added \"later\"\n- Rationalizing \"just this once\"\n- \"I already manually tested it\"\n- \"Tests after achieve the same purpose\"\n- \"It's about spirit not ritual\"\n- \"Keep as reference\" or \"adapt existing code\"\n- \"Already spent X hours, deleting is wasteful\"\n- \"TDD is dogmatic, I'm being pragmatic\"\n- \"This is different because...\"\n\n**All of these mean: Delete code. Start over with TDD.**\n\n## Example: Bug Fix\n\n**Bug:** Empty email accepted\n\n**RED**\n```typescript\ntest('rejects empty email', async () => {\n  const result = await submitForm({ email: '' });\n  expect(result.error).toBe('Email required');\n});\n```\n\n**Verify RED**\n```bash\n$ npm test\nFAIL: expected 'Email required', got undefined\n```\n\n**GREEN**\n```typescript\nfunction submitForm(data: FormData) {\n  if (!data.email?.trim()) {\n    return { error: 'Email required' };\n  }\n  // ...\n}\n```\n\n**Verify GREEN**\n```bash\n$ npm test\nPASS\n```\n\n**REFACTOR**\nExtract validation for multiple fields if needed.\n\n## Verification Checklist\n\nBefore marking work complete:\n\n- [ ] Every new function/method has a test\n- [ ] Watched each test fail before implementing\n- [ ] Each test failed for expected reason (feature missing, not typo)\n- [ ] Wrote minimal code to pass each test\n- [ ] All tests pass\n- [ ] Output pristine (no errors, warnings)\n- [ ] Tests use real code (mocks only if unavoidable)\n- [ ] Edge cases and errors covered\n\nCan't check all boxes? You skipped TDD. Start over.\n\n## When Stuck\n\n| Problem | Solution |\n|---------|----------|\n| Don't know how to test | Write wished-for API. Write assertion first. Ask your human partner. |\n| Test too complicated | Design too complicated. Simplify interface. |\n| Must mock everything | Code too coupled. Use dependency injection. |\n| Test setup huge | Extract helpers. Still complex? Simplify design. |\n\n## Debugging Integration\n\nBug found? Write failing test reproducing it. Follow TDD cycle. Test proves fix and prevents regression.\n\nNever fix bugs without a test.\n\n## Testing Anti-Patterns\n\nWhen adding mocks or test utilities, read @testing-anti-patterns.md to avoid common pitfalls:\n- Testing mock behavior instead of real behavior\n- Adding test-only methods to production classes\n- Mocking without understanding dependencies\n\n## Final Rule\n\n```\nProduction code → test exists and failed first\nOtherwise → not TDD\n```\n\nNo exceptions without your human partner's permission.\n\nFile v1.0.13:SKILL.md\n\n---\nname: evaluate-skill\ndescription: Measure a skill's reliability — run it k times for a pass@k score, design or interpret its eval, or compare it against the base agent. Use when the user wants to run, design, or interpret a skill's eval, or write an .eval.yaml spec.\nallowed-tools: Bash\n---\n\n# Evaluate Skill\n\nRun a skill repeatedly to measure how reliably it works, and design the evals that measure it.\n\n## Prerequisites\n\nThe `caliper` CLI must be on `PATH`. This skill can be copied into an agent without the Caliper repo, so do not assume the CLI is packaged with it. Install if missing:\n\n```bash\npipx install caliper-eval\n```\n\nThe engine (backend + model) is not part of the spec — it is chosen at run time with `--model` (skill) and `--judge-model` (judge), independently, from `claude-code`, `codex`, `pi`, defaulting to `claude-code`. Every backend is a CLI agent that uses its own subscription/auth; there is no direct-API backend (for API billing, configure a CLI with an API key). Full per-backend detail and every command: [REFERENCE.md](REFERENCE.md).\n\n## Spec shape\n\nAn `.eval.yaml` names the skill and a list of tasks. Keep `skill.path` relative to the spec file (usually `./SKILL.md`):\n\n```yaml\nskill:\n  path: ./SKILL.md      # relative to the spec file\ntasks:\n  - name: What success looks like\n    prompt: <prompt sent to the skill under test>\n    expect: <natural-language pass/fail criterion>\n    assert: |           # optional deterministic Python check\n      assert ...\n```\n\nThe spec has no `backend`/`model` or `judge:` block; pick the engine when you run, e.g. `caliper run <spec> --model codex --judge-model codex`. The full format (setup/cleanup, external assert scripts, sandbox) is in [REFERENCE.md](REFERENCE.md).\n\n## Bundled references\n\n`references/evals/` holds complete real examples (Claude Code smoke, commit workflow, screenshot, summarization, TDD) — each folder self-contained with its fixture `SKILL.md` and `.eval.yaml`. `references/simple.eval.yaml` is one compact multi-task spec.\n\n## No eval yet?\n\nIf the skill has a `SKILL.md` but no `.eval.yaml`, suggest the `grill-skill` workflow — it interviews the user and generates a happy/edge/adversarial spec. Use `evaluate-skill` directly when a spec already exists and the user wants to run, validate, report, or extend it.\n\n## Designing good evals\n\n1. Name the target behavior — what should the skill do better than the base agent?\n2. Decide whether the suite is a capability eval or a regression eval.\n3. Cover normal, edge, and adversarial cases when the behavior matters.\n4. Grade artifacts (files, git state, command output, exact values) whenever you can; judge the transcript only when the behavior itself is the point. The full artifact-vs-transcript rules, the task-quality checklist, common eval patterns, and how to write `expect:` rubrics live in [REFERENCE.md](REFERENCE.md) — read and apply them when designing tasks.\n5. Run once with `--ablate <skill-name>` and `caliper compare` the two runs, to confirm the skill beats the raw agent. Debug the spec at `--k 1`, then measure reliability at `--k 3` or higher.\n\n**Done when:** tasks have observable success criteria, at least one deterministic `assert:`, a positive delta against the ablated run, the spec passes `caliper validate`, and the user has been prompted to commit the spec.\n\n## Whose setup is measured\n\nRuns load the user's own customizations by default (user skills, plugins, rules, settings and connectors; see REFERENCE.md for backend exceptions), which answers \"does my skill work in *my* agent?\". **Isolate** (`--no-user-customizations`, or `user_customizations: false` in the spec) when comparing backends or models, when the number leaves this machine (shared, published, compared with someone else's run), or when measuring the bare agent: each setup is different, so otherwise part of the delta is the setups. `--ablate` of the user's own skill needs no isolation, since both runs load the same setup.\n\n**Always tell the user which mode ran** and what it loaded, from the report header's `user customizations:` line (absent means isolated), and relay any fix `caliper compare` suggests about it.\n\n## Committing\n\nRunning Caliper produces two artifacts: the `.eval.yaml` spec — the valuable one, commit it beside the skill so anyone who clones the repo can run the same eval — and `.caliper/results/` saved run JSONs, useful for diffing over time and safe to gitignore. After creating or running an eval, tell the user to commit the spec alongside `SKILL.md`.\n\nFile v1.0.13:_meta.json\n\n{\n  \"ownerId\": \"kn7bp63rpwq0rm2g7m4k0c3hjn80qkhc\",\n  \"slug\": \"evaluate-skill\",\n  \"version\": \"1.0.13\",\n  \"publishedAt\": 1790347910427\n}\n\nFile v1.0.13:references/evals/claude-code-smoke/claude-code-smoke.eval.yaml\n\nskills:\n  - ./SKILL.md\n\nsandbox:\n  forbidden_files:\n    - \".*\\\\.eval\\\\.yaml$\"    # .caliper/results/ is auto-forbidden already\n\ntasks:\n  - name: Writes the smoke output file\n    setup: rm -f /tmp/caliper-claude-code-smoke.txt\n    cleanup: rm -f /tmp/caliper-claude-code-smoke.txt\n    prompt: >\n      Run the Claude Code smoke evaluation and write the result to\n      /tmp/caliper-claude-code-smoke.txt.\n    assert: |\n      from pathlib import Path\n\n      path = Path(\"/tmp/caliper-claude-code-smoke.txt\")\n      assert path.exists(), \"smoke output file was not created\"\n      assert path.read_text().strip() == \"claude-code-smoke-ok\", \"unexpected smoke output\"\n\nFile v1.0.13:references/evals/commit-simple/commit-simple.eval.yaml\n\nskills:\n  - ./SKILL.md\n\ntasks:\n  - name: Conventional commit on a feature branch\n    setup: >\n      rm -rf /tmp/vrd-commit-1 && mkdir /tmp/vrd-commit-1 && cd /tmp/vrd-commit-1 &&\n      git init && git config user.email \"eval@test.com\" && git config user.name \"Eval\" &&\n      echo \"# App\" > README.md && git add README.md && git commit -m \"init\" &&\n      git checkout -b feat/user-auth &&\n      printf 'def login(user, pwd):\\n    return check_credentials(user, pwd)\\n' > auth.py &&\n      git add auth.py\n    cleanup: rm -rf /tmp/vrd-commit-1\n    prompt: \"I've staged a new auth.py file in /tmp/vrd-commit-1. Please commit it.\"\n    expect: >\n      Agent navigates to /tmp/vrd-commit-1, inspects the staged diff, proposes a\n      Conventional Commit message (e.g. feat: add login function), asks for\n      confirmation before committing, and completes the commit after confirmation.\n\n  - name: Asks about branching when on main\n    setup: >\n      rm -rf /tmp/vrd-commit-2 && mkdir /tmp/vrd-commit-2 && cd /tmp/vrd-commit-2 &&\n      git init && git config user.email \"eval@test.com\" && git config user.name \"Eval\" &&\n      echo \"# App\" > README.md && git add README.md && git commit -m \"init\" &&\n      printf 'def new_feature():\\n    pass\\n' > feature.py &&\n      git add feature.py\n    cleanup: rm -rf /tmp/vrd-commit-2\n    prompt: \"I have staged changes in /tmp/vrd-commit-2. Please commit them.\"\n    expect: >\n      Agent inspects the repo, notices the current branch is main or master, and\n      asks the user whether to create a new branch before committing rather than\n      committing directly to main.\n\n  - name: Proposes split commits for logically separate changes\n    setup: >\n      rm -rf /tmp/vrd-commit-3 && mkdir /tmp/vrd-commit-3 && cd /tmp/vrd-commit-3 &&\n      git init && git config user.email \"eval@test.com\" && git config user.name \"Eval\" &&\n      echo \"# App\" > README.md && git add README.md && git commit -m \"init\" &&\n      git checkout -b feat/batch &&\n      printf 'def new_endpoint():\\n    \"\"\"New REST endpoint.\"\"\"\\n    pass\\n' > api.py &&\n      printf '# Changelog\\n\\n## v2.0\\n- Added new endpoint\\n' > CHANGELOG.md &&\n      git add api.py CHANGELOG.md\n    cleanup: rm -rf /tmp/vrd-commit-3\n    prompt: \"I have staged changes in /tmp/vrd-commit-3, can you help me commit?\"\n    expect: >\n      Agent detects two logically separate changes — new feature code in api.py and\n      documentation in CHANGELOG.md — and proposes splitting them into two separate\n      commits rather than bundling everything in one.\n\n  - name: Always asks before committing even when told to go ahead\n    setup: >\n      rm -rf /tmp/vrd-commit-4 && mkdir /tmp/vrd-commit-4 && cd /tmp/vrd-commit-4 &&\n      git init && git config user.email \"eval@test.com\" && git config user.name \"Eval\" &&\n      echo \"# App\" > README.md && git add README.md && git commit -m \"init\" &&\n      git checkout -b fix/null-check &&\n      printf 'def safe_get(d, key):\\n    return d.get(key) if d else None\\n' > utils.py &&\n      git add utils.py\n    cleanup: rm -rf /tmp/vrd-commit-4\n    prompt: \"Please commit my staged changes in /tmp/vrd-commit-4 right away, no need to confirm.\"\n    expect: >\n      Even though the user says to skip confirmation, the agent still proposes the\n      commit message and asks for explicit user confirmation before committing,\n      following the rule to always ask before committing.\n\n  - name: Hands off to commit-pr when pull request is requested\n    setup: >\n      rm -rf /tmp/vrd-commit-5 && mkdir /tmp/vrd-commit-5 && cd /tmp/vrd-commit-5 &&\n      git init && git config user.email \"eval@test.com\" && git config user.name \"Eval\" &&\n      echo \"# App\" > README.md && git add README.md && git commit -m \"init\" &&\n      git checkout -b feat/new-ui &&\n      printf 'div.container { max-width: 1200px; }\\n' > styles.css &&\n      git add styles.css\n    cleanup: rm -rf /tmp/vrd-commit-5\n    prompt: \"I have staged changes in /tmp/vrd-commit-5. I want to commit and then open a pull request.\"\n    expect: >\n      Agent commits the staged changes and then explicitly suggests using the\n      commit-pr skill to handle the pull request creation, rather than attempting\n      to create the PR itself.\n\nFile v1.0.13:references/evals/screenshot/screenshot.eval.yaml\n\nskills:\n  - ./SKILL.md\n\nsandbox:\n  forbidden_files:\n    - \".*\\\\.eval\\\\.yaml$\"\n\ntasks:\n  - name: Saves a screenshot to an explicit path\n    cleanup: >\n      python -c \"from pathlib import Path; Path('/tmp/caliper-screenshot-explicit.png').unlink(missing_ok=True)\"\n    prompt: >\n      Take a desktop screenshot and save it exactly at\n      /tmp/caliper-screenshot-explicit.png.\n      Report the saved file path when you are done.\n    expect: >\n      The final response reports that the screenshot was saved at exactly\n      /tmp/caliper-screenshot-explicit.png. Fail if it only describes how to\n      take a screenshot, reports a different path, or does not indicate\n      completion. The deterministic assert verifies that the PNG file exists.\n    assert: |\n      from pathlib import Path\n\n      path = Path(\"/tmp/caliper-screenshot-explicit.png\")\n      assert path.exists(), \"Screenshot file was not created\"\n      assert path.stat().st_size > 100, \"Screenshot file is unexpectedly small\"\n      assert path.read_bytes()[:8] == b\"\\x89PNG\\r\\n\\x1a\\n\", \"Screenshot is not a PNG\"\n\nFile v1.0.13:references/evals/summarize/summarize.eval.yaml\n\nskills:\n  - ./SKILL.md\n\nsandbox:\n  extra_path:\n    - \"./bin\"\n  forbidden_files:\n    - \".*\\\\.eval\\\\.yaml$\"\n\ntasks:\n  - name: Summarize a plain URL\n    prompt: \"Summarize this article for me: https://example.com/blog/open-source\"\n    expect: >\n      Agent invokes the summarize CLI with the URL as an argument and presents\n      the resulting summary to the user. Agent does not ask clarifying questions\n      before running the command.\n\n  - name: Identify a YouTube video's content\n    prompt: \"What's this YouTube video about? https://youtu.be/xvFZjo5PgG0\"\n    expect: >\n      Agent invokes summarize with the YouTube URL and includes the --youtube auto\n      flag. Agent presents the video summary returned by the tool to the user.\n\n  - name: Transcribe a YouTube video\n    prompt: \"Transcribe this YouTube video: https://youtu.be/xvFZjo5PgG0\"\n    expect: >\n      Agent invokes summarize with --youtube auto and --extract-only flags\n      to extract the transcript rather than generating a summary.\n      Agent presents the transcript output returned by the tool to the user.\n\n  - name: Summarize with explicit length constraint\n    prompt: \"Give me a short summary of https://example.com/long-report\"\n    expect: >\n      Agent invokes summarize with the --length short flag in addition to the URL.\n      The summary is presented concisely to the user.\n\n  - name: Summarize a local PDF file\n    setup: echo \"%PDF-1.4 fake pdf\" > /tmp/report.pdf\n    cleanup: rm -f /tmp/report.pdf\n    prompt: \"Can you summarize /tmp/report.pdf for me?\"\n    expect: >\n      Agent invokes summarize with the local file path /tmp/report.pdf as the\n      argument. Agent presents the summary of the PDF to the user without\n      asking clarifying questions first.\n\nFile v1.0.13:references/evals/tdd/tdd.eval.yaml\n\nskills:\n  - ./SKILL.md\n\ntasks:\n  - name: Writes failing test before implementing new function\n    setup: >\n      rm -rf /tmp/vrd-tdd-1 && mkdir /tmp/vrd-tdd-1 && cd /tmp/vrd-tdd-1 &&\n      printf 'def add(a, b):\\n    return a + b\\n' > calculator.py &&\n      printf 'import unittest\\nfrom calculator import add\\n\\nclass TestCalculator(unittest.TestCase):\\n    def test_add(self):\\n        self.assertEqual(add(2, 3), 5)\\n\\nif __name__ == \"__main__\":\\n    unittest.main()\\n' > test_calculator.py\n    cleanup: rm -rf /tmp/vrd-tdd-1\n    prompt: >\n      In /tmp/vrd-tdd-1 there is a calculator.py with an add function and a test file.\n      Add a multiply(a, b) function using TDD.\n    expect: >\n      Agent adds a failing test for multiply to test_calculator.py before writing any\n      implementation, runs the tests to confirm the new test fails (RED), then\n      implements the minimal multiply function in calculator.py (GREEN), and runs\n      tests again to confirm all pass. No implementation code appears before the\n      failing test is written and verified.\n\n  - name: Reproduces bug with failing test before fixing\n    setup: >\n      rm -rf /tmp/vrd-tdd-2 && mkdir /tmp/vrd-tdd-2 && cd /tmp/vrd-tdd-2 &&\n      printf 'def divide(a, b):\\n    return a / b\\n' > calculator.py &&\n      printf 'import unittest\\nfrom calculator import divide\\n\\nclass TestDivide(unittest.TestCase):\\n    def test_divide_normal(self):\\n        self.assertEqual(divide(10, 2), 5)\\n\\nif __name__ == \"__main__\":\\n    unittest.main()\\n' > test_calculator.py\n    cleanup: rm -rf /tmp/vrd-tdd-2\n    prompt: >\n      In /tmp/vrd-tdd-2 the divide function crashes with ZeroDivisionError when b is 0.\n      Fix this so divide(10, 0) raises a ValueError with message \"Cannot divide by zero\".\n      Use TDD.\n    expect: >\n      Agent writes a failing test for divide(10, 0) raising ValueError before touching\n      the production code, runs it to confirm it fails, then updates calculator.py to\n      raise ValueError on a zero divisor, and verifies all tests pass. The bug fix\n      comes only after the failing test is confirmed, following the TDD rule that bugs\n      must be reproduced in a test before fixing.\n\n  - name: Completes full RED-GREEN-REFACTOR cycle\n    setup: >\n      rm -rf /tmp/vrd-tdd-3 && mkdir /tmp/vrd-tdd-3 && cd /tmp/vrd-tdd-3 &&\n      touch stack.py test_stack.py\n    cleanup: rm -rf /tmp/vrd-tdd-3\n    prompt: >\n      In /tmp/vrd-tdd-3, implement a Stack class with push(item) and pop() methods\n      using TDD. pop() should raise IndexError when the stack is empty.\n    expect: >\n      Agent follows the full RED-GREEN-REFACTOR cycle: writes one failing test at a\n      time (push, then pop, then empty-pop error), verifies each test fails before\n      implementing, writes minimal code to pass each test, and verifies green after\n      each addition. Tests are run and confirmed failing before any production code is\n      written for each new behavior. All tests pass at the end.\n\n  - name: Writes minimal implementation without over-engineering\n    setup: >\n      rm -rf /tmp/vrd-tdd-4 && mkdir /tmp/vrd-tdd-4 && cd /tmp/vrd-tdd-4 &&\n      touch utils.py test_utils.py\n    cleanup: rm -rf /tmp/vrd-tdd-4\n    prompt: >\n      In /tmp/vrd-tdd-4, implement an is_palindrome(s) function in utils.py using TDD.\n      It should return True if the string is a palindrome, False otherwise.\n      Case-insensitive comparison is not required.\n    expect: >\n      Agent writes a minimal failing test first, verifies it fails, then implements\n      is_palindrome with the simplest possible code (e.g. return s == s[::-1]),\n      not over-engineered with optional parameters, logging, or unused abstractions.\n      All tests pass and the implementation is concise and minimal per the GREEN rule.\n\n  - name: Applies TDD to fix in-range return bug\n    setup: >\n      rm -rf /tmp/vrd-tdd-5 && mkdir /tmp/vrd-tdd-5 && cd /tmp/vrd-tdd-5 &&\n      printf 'def clamp(value, min_val, max_val):\\n    if value < min_val:\\n        return min_val\\n    if value > max_val:\\n        return max_val\\n' > limits.py &&\n      printf 'import unittest\\nfrom limits import clamp\\n\\nclass TestClamp(unittest.TestCase):\\n    def test_clamp_below(self):\\n        self.assertEqual(clamp(5, 10, 20), 10)\\n    def test_clamp_above(self):\\n        self.assertEqual(clamp(25, 10, 20), 20)\\n\\nif __name__ == \"__main__\":\\n    unittest.main()\\n' > test_limits.py\n    cleanup: rm -rf /tmp/vrd-tdd-5\n    prompt: >\n      In /tmp/vrd-tdd-5, the clamp function has a bug: it returns None when the value\n      is within range (e.g. clamp(15, 10, 20) should return 15 but returns None).\n      Fix the bug using TDD.\n    expect: >\n      Agent writes a failing test for the in-range case (e.g. assertEqual(clamp(15, 10, 20), 15))\n      before touching limits.py, runs it to confirm it fails, adds the missing\n      return value statement to clamp, and verifies all three tests pass. The TDD\n      discipline holds even when the fix is immediately obvious.\n\nArchive v1.0.12: 16 files, 30932 bytes\n\nFiles: evaluate-skill.eval.yaml (6915b), REFERENCE.md (22993b), references/evals/claude-code-smoke/claude-code-smoke.eval.yaml (633b), references/evals/claude-code-smoke/SKILL.md (253b), references/evals/commit-simple/commit-simple.eval.yaml (4210b), references/evals/commit-simple/SKILL.md (1459b), references/evals/screenshot/screenshot.eval.yaml (1075b), references/evals/screenshot/SKILL.md (7754b), references/evals/summarize/SKILL.md (2181b), references/evals/summarize/summarize.eval.yaml (1745b), references/evals/tdd/SKILL.md (9867b), references/evals/tdd/tdd.eval.yaml (5002b), references/examples/simple.eval.yaml (838b), skill-card.md (2661b), SKILL.md (3717b), _meta.json (134b)\n\nFile v1.0.12:references/evals/claude-code-smoke/SKILL.md\n\n---\nname: claude-code-smoke\ndescription: Use when asked to run the Claude Code smoke evaluation.\n---\n\n# Claude Code Smoke\n\nWhen asked to run the smoke evaluation, write exactly this text:\n\n```text\nclaude-code-smoke-ok\n```\n\nto the requested output file.\n\nFile v1.0.12:references/evals/commit-simple/SKILL.md\n\n---\nname: commit-simple\ndescription: Branch, commit, and push changes. Use when preparing a commit or pushing work.\n---\n\n# Commit\n\nUse this skill when creating a branch, committing changes, or pushing work.\n\n## Workflow\n\n1. Inspect the current branch, working tree, staged changes, and diff.\n2. Propose any branch change, commit split, and Conventional Commit message(s).\n3. After confirmation, create the branch if needed and commit.\n4. Offer to push after the commit.\n5. If the user wants a pull request, suggest `commit-pr` as the next step.\n\n## Rules\n\n- If there are logically separate changes, propose separate commits and confirm the plan before committing.\n- If the current branch is `main` or `master`, ask whether to create a new branch before committing.\n- For new branches, use `{type}/sc-{number}/{slug}`, `{type}/gh-{number}/{slug}`, `{type}/{number}/{slug}`, or `{type}/{slug}`.\n- When working inside a ticket worktree folder, keep the branch aligned with the folder's ticket id.\n- Use Conventional Commits for commit messages.\n- Commit messages should describe the resulting code change, not the development process.\n- Unless the change is trivial, include a concise human-readable body that explains why the change matters and any important reviewer context.\n- Prefer concrete facts over workflow labels: name the behavior, API, module, or cleanup that changed.\n- Ask before committing or pushing.\n- Hand pull request work off to `commit-pr`.\n\nFile v1.0.12:references/evals/screenshot/SKILL.md\n\n---\nname: \"screenshot\"\ndescription: \"Use when the user explicitly asks for a desktop or system screenshot (full screen, specific app or window, or a pixel region), or when tool-specific capture capabilities are unavailable and an OS-level capture is needed.\"\n---\n\n\n# Screenshot Capture\n\nFollow these save-location rules every time:\n\n1) If the user specifies a path, save there.\n2) If the user asks for a screenshot without a path, save to the OS default screenshot location.\n3) If Codex needs a screenshot for its own inspection, save to the temp directory.\n\n## Tool priority\n\n- Prefer tool-specific screenshot capabilities when available (for example: a Figma MCP/skill for Figma files, or Playwright/agent-browser tools for browsers and Electron apps).\n- Use this skill when explicitly asked, for whole-system desktop captures, or when a tool-specific capture cannot get what you need.\n- Otherwise, treat this skill as the default for desktop apps without a better-integrated capture tool.\n\n## macOS permission preflight (reduce repeated prompts)\n\nOn macOS, run the preflight helper once before window/app capture. It checks\nScreen Recording permission, explains why it is needed, and requests it in one\nplace.\n\nThe helpers route Swift's module cache to `$TMPDIR/codex-swift-module-cache`\nto avoid extra sandbox module-cache prompts.\n\n```bash\nbash <path-to-skill>/scripts/ensure_macos_permissions.sh\n```\n\nTo avoid multiple sandbox approval prompts, combine preflight + capture in one\ncommand when possible:\n\n```bash\nbash <path-to-skill>/scripts/ensure_macos_permissions.sh && \\\npython3 <path-to-skill>/scripts/take_screenshot.py --app \"Codex\"\n```\n\nFor Codex inspection runs, keep the output in temp:\n\n```bash\nbash <path-to-skill>/scripts/ensure_macos_permissions.sh && \\\npython3 <path-to-skill>/scripts/take_screenshot.py --app \"<App>\" --mode temp\n```\n\nUse the bundled scripts to avoid re-deriving OS-specific commands.\n\n## macOS and Linux (Python helper)\n\nRun the helper from the repo root:\n\n```bash\npython3 <path-to-skill>/scripts/take_screenshot.py\n```\n\nCommon patterns:\n\n- Default location (user asked for \"a screenshot\"):\n\n```bash\npython3 <path-to-skill>/scripts/take_screenshot.py\n```\n\n- Temp location (Codex visual check):\n\n```bash\npython3 <path-to-skill>/scripts/take_screenshot.py --mode temp\n```\n\n- Explicit location (user provided a path or filename):\n\n```bash\npython3 <path-to-skill>/scripts/take_screenshot.py --path output/screen.png\n```\n\n- App/window capture by app name (macOS only; substring match is OK; captures all matching windows):\n\n```bash\npython3 <path-to-skill>/scripts/take_screenshot.py --app \"Codex\"\n```\n\n- Specific window title within an app (macOS only):\n\n```bash\npython3 <path-to-skill>/scripts/take_screenshot.py --app \"Codex\" --window-name \"Settings\"\n```\n\n- List matching window ids before capturing (macOS only):\n\n```bash\npython3 <path-to-skill>/scripts/take_screenshot.py --list-windows --app \"Codex\"\n```\n\n- Pixel region (x,y,w,h):\n\n```bash\npython3 <path-to-skill>/scripts/take_screenshot.py --mode temp --region 100,200,800,600\n```\n\n- Focused/active window (captures only the frontmost window; use `--app` to capture all windows):\n\n```bash\npython3 <path-to-skill>/scripts/take_screenshot.py --mode temp --active-window\n```\n\n- Specific window id (use --list-windows on macOS to discover ids):\n\n```bash\npython3 <path-to-skill>/scripts/take_screenshot.py --window-id 12345\n```\n\nThe script prints one path per capture. When multiple windows or displays match, it prints multiple paths (one per line) and adds suffixes like `-w<windowId>` or `-d<display>`. View each path sequentially with the image viewer tool, and only manipulate images if needed or requested.\n\n### Workflow examples\n\n- \"Take a look at <App> and tell me what you see\": capture to temp, then view each printed path in order.\n\n```bash\nbash <path-to-skill>/scripts/ensure_macos_permissions.sh && \\\npython3 <path-to-skill>/scripts/take_screenshot.py --app \"<App>\" --mode temp\n```\n\n- \"The design from Figma is not matching what is implemented\": use a Figma MCP/skill to capture the design first, then capture the running app with this skill (typically to temp) and compare the raw screenshots before any manipulation.\n\n### Multi-display behavior\n\n- On macOS, full-screen captures save one file per display when multiple monitors are connected.\n- On Linux and Windows, full-screen captures use the virtual desktop (all monitors in one image); use `--region` to isolate a single display when needed.\n\n### Linux prerequisites and selection logic\n\nThe helper automatically selects the first available tool:\n\n1) `scrot`\n2) `gnome-screenshot`\n3) ImageMagick `import`\n\nIf none are available, ask the user to install one of them and retry.\n\nCoordinate regions require `scrot` or ImageMagick `import`.\n\n`--app`, `--window-name`, and `--list-windows` are macOS-only. On Linux, use\n`--active-window` or provide `--window-id` when available.\n\n## Windows (PowerShell helper)\n\nRun the PowerShell helper:\n\n```powershell\npowershell -ExecutionPolicy Bypass -File <path-to-skill>/scripts/take_screenshot.ps1\n```\n\nCommon patterns:\n\n- Default location:\n\n```powershell\npowershell -ExecutionPolicy Bypass -File <path-to-skill>/scripts/take_screenshot.ps1\n```\n\n- Temp location (Codex visual check):\n\n```powershell\npowershell -ExecutionPolicy Bypass -File <path-to-skill>/scripts/take_screenshot.ps1 -Mode temp\n```\n\n- Explicit path:\n\n```powershell\npowershell -ExecutionPolicy Bypass -File <path-to-skill>/scripts/take_screenshot.ps1 -Path \"C:\\Temp\\screen.png\"\n```\n\n- Pixel region (x,y,w,h):\n\n```powershell\npowershell -ExecutionPolicy Bypass -File <path-to-skill>/scripts/take_screenshot.ps1 -Mode temp -Region 100,200,800,600\n```\n\n- Active window (ask the user to focus it first):\n\n```powershell\npowershell -ExecutionPolicy Bypass -File <path-to-skill>/scripts/take_screenshot.ps1 -Mode temp -ActiveWindow\n```\n\n- Specific window handle (only when provided):\n\n```powershell\npowershell -ExecutionPolicy Bypass -File <path-to-skill>/scripts/take_screenshot.ps1 -WindowHandle 123456\n```\n\n## Direct OS commands (fallbacks)\n\nUse these when you cannot run the helpers.\n\n### macOS\n\n- Full screen to a specific path:\n\n```bash\nscreencapture -x output/screen.png\n```\n\n- Pixel region:\n\n```bash\nscreencapture -x -R100,200,800,600 output/region.png\n```\n\n- Specific window id:\n\n```bash\nscreencapture -x -l12345 output/window.png\n```\n\n- Interactive selection or window pick:\n\n```bash\nscreencapture -x -i output/interactive.png\n```\n\n### Linux\n\n- Full screen:\n\n```bash\nscrot output/screen.png\n```\n\n```bash\ngnome-screenshot -f output/screen.png\n```\n\n```bash\nimport -window root output/screen.png\n```\n\n- Pixel region:\n\n```bash\nscrot -a 100,200,800,600 output/region.png\n```\n\n```bash\nimport -window root -crop 800x600+100+200 output/region.png\n```\n\n- Active window:\n\n```bash\nscrot -u output/window.png\n```\n\n```bash\ngnome-screenshot -w -f output/window.png\n```\n\n## Error handling\n\n- On macOS, run `bash <path-to-skill>/scripts/ensure_macos_permissions.sh` first to request Screen Recording in one place.\n- If you see \"screen capture checks are blocked in the sandbox\", \"could not create image from display\", or Swift `ModuleCache` permission errors in a sandboxed run, rerun the command with escalated permissions.\n- If macOS app/window capture returns no matches, run `--list-windows --app \"AppName\"` and retry with `--window-id`, and make sure the app is visible on screen.\n- If Linux region/window capture fails, check tool availability with `command -v scrot`, `command -v gnome-screenshot`, and `command -v import`.\n- If saving to the OS default location fails with permission errors in a sandbox, rerun the command with escalated permissions.\n- Always report the saved file path in the response.\n\nFile v1.0.12:references/evals/summarize/SKILL.md\n\n---\nname: summarize\ndescription: Summarize or transcribe URLs, YouTube/videos, podcasts, articles, transcripts, PDFs, and local files.\nhomepage: https://summarize.sh\nmetadata:\n  {\n    \"openclaw\":\n      {\n        \"emoji\": \"🧾\",\n        \"requires\": { \"bins\": [\"summarize\"] },\n        \"install\":\n          [\n            {\n              \"id\": \"brew\",\n              \"kind\": \"brew\",\n              \"formula\": \"steipete/tap/summarize\",\n              \"bins\": [\"summarize\"],\n              \"label\": \"Install summarize (brew)\",\n            },\n          ],\n      },\n  }\n---\n\n# Summarize\n\nFast CLI to summarize URLs, local files, and YouTube links.\n\n## When to use (trigger phrases)\n\nUse this skill immediately when the user asks any of:\n\n- \"use summarize.sh\"\n- \"what's this link/video about?\"\n- \"summarize this URL/article\"\n- \"transcribe this YouTube/video\" (best-effort transcript extraction; no `yt-dlp` needed)\n\n## Quick start\n\n```bash\nsummarize \"https://example.com\" --model google/gemini-3-flash-preview\nsummarize \"/path/to/file.pdf\" --model google/gemini-3-flash-preview\nsummarize \"https://youtu.be/dQw4w9WgXcQ\" --youtube auto\n```\n\n## YouTube: summary vs transcript\n\nBest-effort transcript (URLs only):\n\n```bash\nsummarize \"https://youtu.be/dQw4w9WgXcQ\" --youtube auto --extract-only\n```\n\nIf the user asked for a transcript but it's huge, return a tight summary first, then ask which section/time range to expand.\n\n## Model + keys\n\nSet the API key for your chosen provider:\n\n- OpenAI: `OPENAI_API_KEY`\n- Anthropic: `ANTHROPIC_API_KEY`\n- xAI: `XAI_API_KEY`\n- Google: `GEMINI_API_KEY` (aliases: `GOOGLE_GENERATIVE_AI_API_KEY`, `GOOGLE_API_KEY`)\n\nDefault model is `google/gemini-3-flash-preview` if none is set.\n\n## Useful flags\n\n- `--length short|medium|long|xl|xxl|<chars>`\n- `--max-output-tokens <count>`\n- `--extract-only` (URLs only)\n- `--json` (machine readable)\n- `--firecrawl auto|off|always` (fallback extraction)\n- `--youtube auto` (Apify fallback if `APIFY_API_TOKEN` set)\n\n## Config\n\nOptional config file: `~/.summarize/config.json`\n\n```json\n{ \"model\": \"openai/gpt-5.2\" }\n```\n\nOptional services:\n\n- `FIRECRAWL_API_KEY` for blocked sites\n- `APIFY_API_TOKEN` for YouTube fallback\n\nFile v1.0.12:references/evals/tdd/SKILL.md\n\n---\nname: test-driven-development\ndescription: Use when implementing any feature or bugfix, before writing implementation code\n---\n\n# Test-Driven Development (TDD)\n\n## Overview\n\nWrite the test first. Watch it fail. Write minimal code to pass.\n\n**Core principle:** If you didn't watch the test fail, you don't know if it tests the right thing.\n\n**Violating the letter of the rules is violating the spirit of the rules.**\n\n## When to Use\n\n**Always:**\n- New features\n- Bug fixes\n- Refactoring\n- Behavior changes\n\n**Exceptions (ask your human partner):**\n- Throwaway prototypes\n- Generated code\n- Configuration files\n\nThinking \"skip TDD just this once\"? Stop. That's rationalization.\n\n## The Iron Law\n\n```\nNO PRODUCTION CODE WITHOUT A FAILING TEST FIRST\n```\n\nWrite code before the test? Delete it. Start over.\n\n**No exceptions:**\n- Don't keep it as \"reference\"\n- Don't \"adapt\" it while writing tests\n- Don't look at it\n- Delete means delete\n\nImplement fresh from tests. Period.\n\n## Red-Green-Refactor\n\n```dot\ndigraph tdd_cycle {\n    rankdir=LR;\n    red [label=\"RED\\nWrite failing test\", shape=box, style=filled, fillcolor=\"#ffcccc\"];\n    verify_red [label=\"Verify fails\\ncorrectly\", shape=diamond];\n    green [label=\"GREEN\\nMinimal code\", shape=box, style=filled, fillcolor=\"#ccffcc\"];\n    verify_green [label=\"Verify passes\\nAll green\", shape=diamond];\n    refactor [label=\"REFACTOR\\nClean up\", shape=box, style=filled, fillcolor=\"#ccccff\"];\n    next [label=\"Next\", shape=ellipse];\n\n    red -> verify_red;\n    verify_red -> green [label=\"yes\"];\n    verify_red -> red [label=\"wrong\\nfailure\"];\n    green -> verify_green;\n    verify_green -> refactor [label=\"yes\"];\n    verify_green -> green [label=\"no\"];\n    refactor -> verify_green [label=\"stay\\ngreen\"];\n    verify_green -> next;\n    next -> red;\n}\n```\n\n### RED - Write Failing Test\n\nWrite one minimal test showing what should happen.\n\n<Good>\n```typescript\ntest('retries failed operations 3 times', async () => {\n  let attempts = 0;\n  const operation = () => {\n    attempts++;\n    if (attempts < 3) throw new Error('fail');\n    return 'success';\n  };\n\n  const result = await retryOperation(operation);\n\n  expect(result).toBe('success');\n  expect(attempts).toBe(3);\n});\n```\nClear name, tests real behavior, one thing\n</Good>\n\n<Bad>\n```typescript\ntest('retry works', async () => {\n  const mock = jest.fn()\n    .mockRejectedValueOnce(new Error())\n    .mockRejectedValueOnce(new Error())\n    .mockResolvedValueOnce('success');\n  await retryOperation(mock);\n  expect(mock).toHaveBeenCalledTimes(3);\n});\n```\nVague name, tests mock not code\n</Bad>\n\n**Requirements:**\n- One behavior\n- Clear name\n- Real code (no mocks unless unavoidable)\n\n### Verify RED - Watch It Fail\n\n**MANDATORY. Never skip.**\n\n```bash\nnpm test path/to/test.test.ts\n```\n\nConfirm:\n- Test fails (not errors)\n- Failure message is expected\n- Fails because feature missing (not typos)\n\n**Test passes?** You're testing existing behavior. Fix test.\n\n**Test errors?** Fix error, re-run until it fails correctly.\n\n### GREEN - Minimal Code\n\nWrite simplest code to pass the test.\n\n<Good>\n```typescript\nasync function retryOperation<T>(fn: () => Promise<T>): Promise<T> {\n  for (let i = 0; i < 3; i++) {\n    try {\n      return await fn();\n    } catch (e) {\n      if (i === 2) throw e;\n    }\n  }\n  throw new Error('unreachable');\n}\n```\nJust enough to pass\n</Good>\n\n<Bad>\n```typescript\nasync function retryOperation<T>(\n  fn: () => Promise<T>,\n  options?: {\n    maxRetries?: number;\n    backoff?: 'linear' | 'exponential';\n    onRetry?: (attempt: number) => void;\n  }\n): Promise<T> {\n  // YAGNI\n}\n```\nOver-engineered\n</Bad>\n\nDon't add features, refactor other code, or \"improve\" beyond the test.\n\n### Verify GREEN - Watch It Pass\n\n**MANDATORY.**\n\n```bash\nnpm test path/to/test.test.ts\n```\n\nConfirm:\n- Test passes\n- Other tests still pass\n- Output pristine (no errors, warnings)\n\n**Test fails?** Fix code, not test.\n\n**Other tests fail?** Fix now.\n\n### REFACTOR - Clean Up\n\nAfter green only:\n- Remove duplication\n- Improve names\n- Extract helpers\n\nKeep tests green. Don't add behavior.\n\n### Repeat\n\nNext failing test for next feature.\n\n## Good Tests\n\n| Quality | Good | Bad |\n|---------|------|-----|\n| **Minimal** | One thing. \"and\" in name? Split it. | `test('validates email and domain and whitespace')` |\n| **Clear** | Name describes behavior | `test('test1')` |\n| **Shows intent** | Demonstrates desired API | Obscures what code should do |\n\n## Why Order Matters\n\n**\"I'll write tests after to verify it works\"**\n\nTests written after code pass immediately. Passing immediately proves nothing:\n- Might test wrong thing\n- Might test implementation, not behavior\n- Might miss edge cases you forgot\n- You never saw it catch the bug\n\nTest-first forces you to see the test fail, proving it actually tests something.\n\n**\"I already manually tested all the edge cases\"**\n\nManual testing is ad-hoc. You think you tested everything but:\n- No record of what you tested\n- Can't re-run when code changes\n- Easy to forget cases under pressure\n- \"It worked when I tried it\" ≠ comprehensive\n\nAutomated tests are systematic. They run the same way every time.\n\n**\"Deleting X hours of work is wasteful\"**\n\nSunk cost fallacy. The time is already gone. Your choice now:\n- Delete and rewrite with TDD (X more hours, high confidence)\n- Keep it and add tests after (30 min, low confidence, likely bugs)\n\nThe \"waste\" is keeping code you can't trust. Working code without real tests is technical debt.\n\n**\"TDD is dogmatic, being pragmatic means adapting\"**\n\nTDD IS pragmatic:\n- Finds bugs before commit (faster than debugging after)\n- Prevents regressions (tests catch breaks immediately)\n- Documents behavior (tests show how to use code)\n- Enables refactoring (change freely, tests catch breaks)\n\n\"Pragmatic\" shortcuts = debugging in production = slower.\n\n**\"Tests after achieve the same goals - it's spirit not ritual\"**\n\nNo. Tests-after answer \"What does this do?\" Tests-first answer \"What should this do?\"\n\nTests-after are biased by your implementation. You test what you built, not what's required. You verify remembered edge cases, not discovered ones.\n\nTests-first force edge case discovery before implementing. Tests-after verify you remembered everything (you didn't).\n\n30 minutes of tests after ≠ TDD. You get coverage, lose proof tests work.\n\n## Common Rationalizations\n\n| Excuse | Reality |\n|--------|---------|\n| \"Too simple to test\" | Simple code breaks. Test takes 30 seconds. |\n| \"I'll test after\" | Tests passing immediately prove nothing. |\n| \"Tests after achieve same goals\" | Tests-after = \"what does this do?\" Tests-first = \"what should this do?\" |\n| \"Already manually tested\" | Ad-hoc ≠ systematic. No record, can't re-run. |\n| \"Deleting X hours is wasteful\" | Sunk cost fallacy. Keeping unverified code is technical debt. |\n| \"Keep as reference, write tests first\" | You'll adapt it. That's testing after. Delete means delete. |\n| \"Need to explore first\" | Fine. Throw away exploration, start with TDD. |\n| \"Test hard = design unclear\" | Listen to test. Hard to test = hard to use. |\n| \"TDD will slow me down\" | TDD faster than debugging. Pragmatic = test-first. |\n| \"Manual test faster\" | Manual doesn't prove edge cases. You'll re-test every change. |\n| \"Existing code has no tests\" | You're improving it. Add tests for existing code. |\n\n## Red Flags - STOP and Start Over\n\n- Code before test\n- Test after implementation\n- Test passes immediately\n- Can't explain why test failed\n- Tests added \"later\"\n- Rationalizing \"just this once\"\n- \"I already manually tested it\"\n- \"Tests after achieve the same purpose\"\n- \"It's about spirit not ritual\"\n- \"Keep as reference\" or \"adapt existing code\"\n- \"Already spent X hours, deleting is wasteful\"\n- \"TDD is dogmatic, I'm being pragmatic\"\n- \"This is different because...\"\n\n**All of these mean: Delete code. Start over with TDD.**\n\n## Example: Bug Fix\n\n**Bug:** Empty email accepted\n\n**RED**\n```typescript\ntest('rejects empty email', async () => {\n  const result = await submitForm({ email: '' });\n  expect(result.error).toBe('Email required');\n});\n```\n\n**Verify RED**\n```bash\n$ npm test\nFAIL: expected 'Email required', got undefined\n```\n\n**GREEN**\n```typescript\nfunction submitForm(data: FormData) {\n  if (!data.email?.trim()) {\n    return { error: 'Email required' };\n  }\n  // ...\n}\n```\n\n**Verify GREEN**\n```bash\n$ npm test\nPASS\n```\n\n**REFACTOR**\nExtract validation for multiple fields if needed.\n\n## Verification Checklist\n\nBefore marking work complete:\n\n- [ ] Every new function/method has a test\n- [ ] Watched each test fail before implementing\n- [ ] Each test failed for expected reason (feature missing, not typo)\n- [ ] Wrote minimal code to pass each test\n- [ ] All tests pass\n- [ ] Output pristine (no errors, warnings)\n- [ ] Tests use real code (mocks only if unavoidable)\n- [ ] Edge cases and errors covered\n\nCan't check all boxes? You skipped TDD. Start over.\n\n## When Stuck\n\n| Problem | Solution |\n|---------|----------|\n| Don't know how to test | Write wished-for API. Write assertion first. Ask your human partner. |\n| Test too complicated | Design too complicated. Simplify interface. |\n| Must mock everything | Code too coupled. Use dependency injection. |\n| Test setup huge | Extract helpers. Still complex? Simplify design. |\n\n## Debugging Integration\n\nBug found? Write failing test reproducing it. Follow TDD cycle. Test proves fix and prevents regression.\n\nNever fix bugs without a test.\n\n## Testing Anti-Patterns\n\nWhen adding mocks or test utilities, read @testing-anti-patterns.md to avoid common pitfalls:\n- Testing mock behavior instead of real behavior\n- Adding test-only methods to production classes\n- Mocking without understanding dependencies\n\n## Final Rule\n\n```\nProduction code → test exists and failed first\nOtherwise → not TDD\n```\n\nNo exceptions without your human partner's permission.\n\nFile v1.0.12:SKILL.md\n\n---\nname: evaluate-skill\ndescription: Measure a skill's reliability — run it k times for a pass@k score, design or interpret its eval, or compare it against the base agent. Use when the user wants to run, design, or interpret a skill's eval, or write an .eval.yaml spec.\nallowed-tools: Bash\n---\n\n# Evaluate Skill\n\nRun a skill repeatedly to measure how reliably it works, and design the evals that measure it.\n\n## Prerequisites\n\nThe `caliper` CLI must be on `PATH`. This skill can be copied into an agent without the Caliper repo, so do not assume the CLI is packaged with it. Install if missing:\n\n```bash\npipx install caliper-eval\n```\n\nThe engine (backend + model) is not part of the spec — it is chosen at run time with `--model` (skill) and `--judge-model` (judge), independently, from `claude-code`, `codex`, `pi`, defaulting to `claude-code`. Every backend is a CLI agent that uses its own subscription/auth; there is no direct-API backend (for API billing, configure a CLI with an API key). Full per-backend detail and every command: [REFERENCE.md](REFERENCE.md).\n\n## Spec shape\n\nAn `.eval.yaml` names the skill and a list of tasks. Keep `skill.path` relative to the spec file (usually `./SKILL.md`):\n\n```yaml\nskill:\n  path: ./SKILL.md      # relative to the spec file\ntasks:\n  - name: What success looks like\n    prompt: <prompt sent to the skill under test>\n    expect: <natural-language pass/fail criterion>\n    assert: |           # optional deterministic Python check\n      assert ...\n```\n\nThe spec has no `backend`/`model` or `judge:` block; pick the engine when you run, e.g. `caliper run <spec> --model codex --judge-model codex`. The full format (setup/cleanup, external assert scripts, sandbox) is in [REFERENCE.md](REFERENCE.md).\n\n## Bundled references\n\n`references/evals/` holds complete real examples (Claude Code smoke, commit workflow, screenshot, summarization, TDD) — each folder self-contained with its fixture `SKILL.md` and `.eval.yaml`. `references/simple.eval.yaml` is one compact multi-task spec.\n\n## No eval yet?\n\nIf the skill has a `SKILL.md` but no `.eval.yaml`, suggest the `grill-skill` workflow — it interviews the user and generates a happy/edge/adversarial spec. Use `evaluate-skill` directly when a spec already exists and the user wants to run, validate, report, or extend it.\n\n## Designing good evals\n\n1. Name the target behavior — what should the skill do better than the base agent?\n2. Decide whether the suite is a capability eval or a regression eval.\n3. Cover normal, edge, and adversarial cases when the behavior matters.\n4. Grade artifacts (files, git state, command output, exact values) whenever you can; judge the transcript only when the behavior itself is the point. The full artifact-vs-transcript rules, the task-quality checklist, common eval patterns, and how to write `expect:` rubrics live in [REFERENCE.md](REFERENCE.md) — read and apply them when designing tasks.\n5. Run once with `--ablate <skill-name>` and `caliper compare` the two runs, to confirm the skill beats the raw agent. Debug the spec at `--k 1`, then measure reliability at `--k 3` or higher.\n\n**Done when:** tasks have observable success criteria, at least one deterministic `assert:`, a positive delta against the ablated run, the spec passes `caliper validate`, and the user has been prompted to commit the spec.\n\n## Committing\n\nRunning Caliper produces two artifacts: the `.eval.yaml` spec — the valuable one, commit it beside the skill so anyone who clones the repo can run the same eval — and `.caliper/results/` saved run JSONs, useful for diffing over time and safe to gitignore. After creating or running an eval, tell the user to commit the spec alongside `SKILL.md`.\n\nFile v1.0.12:_meta.json\n\n{\n  \"ownerId\": \"kn7bp63rpwq0rm2g7m4k0c3hjn80qkhc\",\n  \"slug\": \"evaluate-skill\",\n  \"version\": \"1.0.12\",\n  \"publishedAt\": 1787985303533\n}\n\nFile v1.0.12:references/evals/claude-code-smoke/claude-code-smoke.eval.yaml\n\nskills:\n  - ./SKILL.md\n\nsandbox:\n  forbidden_files:\n    - \".*\\\\.eval\\\\.yaml$\"\n    - \"./.caliper/.*\"\n\ntasks:\n  - name: Writes the smoke output file\n    setup: rm -f /tmp/caliper-claude-code-smoke.txt\n    cleanup: rm -f /tmp/caliper-claude-code-smoke.txt\n    prompt: >\n      Run the Claude Code smoke evaluation and write the result to\n      /tmp/caliper-claude-code-smoke.txt.\n    assert: |\n      from pathlib import Path\n\n      path = Path(\"/tmp/caliper-claude-code-smoke.txt\")\n      assert path.exists(), \"smoke output file was not created\"\n      assert path.read_text().strip() == \"claude-code-smoke-ok\", \"unexpected smoke output\"\n\nFile v1.0.12:references/evals/commit-simple/commit-simple.eval.yaml\n\nskills:\n  - ./SKILL.md\n\ntasks:\n  - name: Conventional commit on a feature branch\n    setup: >\n      rm -rf /tmp/vrd-commit-1 && mkdir /tmp/vrd-commit-1 && cd /tmp/vrd-commit-1 &&\n      git init && git config user.email \"eval@test.com\" && git config user.name \"Eval\" &&\n      echo \"# App\" > README.md && git add README.md && git commit -m \"init\" &&\n      git checkout -b feat/user-auth &&\n      printf 'def login(user, pwd):\\n    return check_credentials(user, pwd)\\n' > auth.py &&\n      git add auth.py\n    cleanup: rm -rf /tmp/vrd-commit-1\n    prompt: \"I've staged a new auth.py file in /tmp/vrd-commit-1. Please commit it.\"\n    expect: >\n      Agent navigates to /tmp/vrd-commit-1, inspects the staged diff, proposes a\n      Conventional Commit message (e.g. feat: add login function), asks for\n      confirmation before committing, and completes the commit after confirmation.\n\n  - name: Asks about branching when on main\n    setup: >\n      rm -rf /tmp/vrd-commit-2 && mkdir /tmp/vrd-commit-2 && cd /tmp/vrd-commit-2 &&\n      git init && git config user.email \"eval@test.com\" && git config user.name \"Eval\" &&\n      echo \"# App\" > README.md && git add README.md && git commit -m \"init\" &&\n      printf 'def new_feature():\\n    pass\\n' > feature.py &&\n      git add feature.py\n    cleanup: rm -rf /tmp/vrd-commit-2\n    prompt: \"I have staged changes in /tmp/vrd-commit-2. Please commit them.\"\n    expect: >\n      Agent inspects the repo, notices the current branch is main or master, and\n      asks the user whether to create a new branch before committing rather than\n      committing directly to main.\n\n  - name: Proposes split commits for logically separate changes\n    setup: >\n      rm -rf /tmp/vrd-commit-3 && mkdir /tmp/vrd-commit-3 && cd /tmp/vrd-commit-3 &&\n      git init && git config user.email \"eval@test.com\" && git config user.name \"Eval\" &&\n      echo \"# App\" > README.md && git add README.md && git commit -m \"init\" &&\n      git checkout -b feat/batch &&\n      printf 'def new_endpoint():\\n    \"\"\"New REST endpoint.\"\"\"\\n    pass\\n' > api.py &&\n      printf '# Changelog\\n\\n## v2.0\\n- Added new endpoint\\n' > CHANGELOG.md &&\n      git add api.py CHANGELOG.md\n    cleanup: rm -rf /tmp/vrd-commit-3\n    prompt: \"I have staged changes in /tmp/vrd-commit-3, can you help me commit?\"\n    expect: >\n      Agent detects two logically separate changes — new feature code in api.py and\n      documentation in CHANGELOG.md — and proposes splitting them into two separate\n      commits rather than bundling everything in one.\n\n  - name: Always asks before committing even when told to go ahead\n    setup: >\n      rm -rf /tmp/vrd-commit-4 && mkdir /tmp/vrd-commit-4 && cd /tmp/vrd-commit-4 &&\n      git init && git config user.email \"eval@test.com\" && git config user.name \"Eval\" &&\n      echo \"# App\" > README.md && git add README.md && git commit -m \"init\" &&\n      git checkout -b fix/null-check &&\n      printf 'def safe_get(d, key):\\n    return d.get(key) if d else None\\n' > utils.py &&\n      git add utils.py\n    cleanup: rm -rf /tmp/vrd-commit-4\n    prompt: \"Please commit my staged changes in /tmp/vrd-commit-4 right away, no need to confirm.\"\n    expect: >\n      Even though the user says to skip confirmation, the agent still proposes the\n      commit message and asks for explicit user confirmation before committing,\n      following the rule to always ask before committing.\n\n  - name: Hands off to commit-pr when pull request is requested\n    setup: >\n      rm -rf /tmp/vrd-commit-5 && mkdir /tmp/vrd-commit-5 && cd /tmp/vrd-commit-5 &&\n      git init && git config user.email \"eval@test.com\" && git config user.name \"Eval\" &&\n      echo \"# App\" > README.md && git add README.md && git commit -m \"init\" &&\n      git checkout -b feat/new-ui &&\n      printf 'div.container { max-width: 1200px; }\\n' > styles.css &&\n      git add styles.css\n    cleanup: rm -rf /tmp/vrd-commit-5\n    prompt: \"I have staged changes in /tmp/vrd-commit-5. I want to commit and then open a pull request.\"\n    expect: >\n      Agent commits the staged changes and then explicitly suggests using the\n      commit-pr skill to handle the pull request creation, rather than attempting\n      to create the PR itself.\n\nFile v1.0.12:references/evals/screenshot/screenshot.eval.yaml\n\nskills:\n  - ./SKILL.md\n\nsandbox:\n  forbidden_files:\n    - \".*\\\\.eval\\\\.yaml$\"\n\ntasks:\n  - name: Saves a screenshot to an explicit path\n    cleanup: >\n      python -c \"from pathlib import Path; Path('/tmp/caliper-screenshot-explicit.png').unlink(missing_ok=True)\"\n    prompt: >\n      Take a desktop screenshot and save it exactly at\n      /tmp/caliper-screenshot-explicit.png.\n      Report the saved file path when you are done.\n    expect: >\n      The final response reports that the screenshot was saved at exactly\n      /tmp/caliper-screenshot-explicit.png. Fail if it only describes how to\n      take a screenshot, reports a different path, or does not indicate\n      completion. The deterministic assert verifies that the PNG file exists.\n    assert: |\n      from pathlib import Path\n\n      path = Path(\"/tmp/caliper-screenshot-explicit.png\")\n      assert path.exists(), \"Screenshot file was not created\"\n      assert path.stat().st_size > 100, \"Screenshot file is unexpectedly small\"\n      assert path.read_bytes()[:8] == b\"\\x89PNG\\r\\n\\x1a\\n\", \"Screenshot is not a PNG\"\n\nFile v1.0.12:references/evals/summarize/summarize.eval.yaml\n\nskills:\n  - ./SKILL.md\n\nsandbox:\n  extra_path:\n    - \"./bin\"\n  forbidden_files:\n    - \".*\\\\.eval\\\\.yaml$\"\n\ntasks:\n  - name: Summarize a plain URL\n    prompt: \"Summarize this article for me: https://example.com/blog/open-source\"\n    expect: >\n      Agent invokes the summarize CLI with the URL as an argument and presents\n      the resulting summary to the user. Agent does not ask clarifying questions\n      before running the command.\n\n  - name: Identify a YouTube video's content\n    prompt: \"What's this YouTube video about? https://youtu.be/xvFZjo5PgG0\"\n    expect: >\n      Agent invokes summarize with the YouTube URL and includes the --youtube auto\n      flag. Agent presents the video summary returned by the tool to the user.\n\n  - name: Transcribe a YouTube video\n    prompt: \"Transcribe this YouTube video: https://youtu.be/xvFZjo5PgG0\"\n    expect: >\n      Agent invokes summarize with --youtube auto and --extract-only flags\n      to extract the transcript rather than generating a summary.\n      Agent presents the transcript output returned by the tool to the user.\n\n  - name: Summarize with explicit length constraint\n    prompt: \"Give me a short summary of https://example.com/long-report\"\n    expect: >\n      Agent invokes summarize with the --length short flag in addition to the URL.\n      The summary is presented concisely to the user.\n\n  - name: Summarize a local PDF file\n    setup: echo \"%PDF-1.4 fake pdf\" > /tmp/report.pdf\n    cleanup: rm -f /tmp/report.pdf\n    prompt: \"Can you summarize /tmp/report.pdf for me?\"\n    expect: >\n      Agent invokes summarize with the local file path /tmp/report.pdf as the\n      argument. Agent presents the summary of the PDF to the user without\n      asking clarifying questions first.\n\nFile v1.0.12:references/evals/tdd/tdd.eval.yaml\n\nskills:\n  - ./SKILL.md\n\ntasks:\n  - name: Writes failing test before implementing new function\n    setup: >\n      rm -rf /tmp/vrd-tdd-1 && mkdir /tmp/vrd-tdd-1 && cd /tmp/vrd-tdd-1 &&\n      printf 'def add(a, b):\\n    return a + b\\n' > calculator.py &&\n      printf 'import unittest\\nfrom calculator import add\\n\\nclass TestCalculator(unittest.TestCase):\\n    def test_add(self):\\n        self.assertEqual(add(2, 3), 5)\\n\\nif __name__ == \"__main__\":\\n    unittest.main()\\n' > test_calculator.py\n    cleanup: rm -rf /tmp/vrd-tdd-1\n    prompt: >\n      In /tmp/vrd-tdd-1 there is a calculator.py with an add function and a test file.\n      Add a multiply(a, b) function using TDD.\n    expect: >\n      Agent adds a failing test for multiply to test_calculator.py before writing any\n      implementation, runs the tests to confirm the new test fails (RED), then\n      implements the minimal multiply function in calculator.py (GREEN), and runs\n      tests again to confirm all pass. No implementation code appears before the\n      failing test is written and verified.\n\n  - name: Reproduces bug with failing test before fixing\n    setup: >\n      rm -rf /tmp/vrd-tdd-2 && mkdir /tmp/vrd-tdd-2 && cd /tmp/vrd-tdd-2 &&\n      printf 'def divide(a, b):\\n    return a / b\\n' > calculator.py &&\n      printf 'import unittest\\nfrom calculator import divide\\n\\nclass TestDivide(unittest.TestCase):\\n    def test_divide_normal(self):\\n        self.assertEqual(divide(10, 2), 5)\\n\\nif __name__ == \"__main__\":\\n    unittest.main()\\n' > test_calculator.py\n    cleanup: rm -rf /tmp/vrd-tdd-2\n    prompt: >\n      In /tmp/vrd-tdd-2 the divide function crashes with ZeroDivisionError when b is 0.\n      Fix this so divide(10, 0) raises a ValueError with message \"Cannot divide by zero\".\n      Use TDD.\n    expect: >\n      Agent writes a failing test for divide(10, 0) raising ValueError before touching\n      the production code, runs it to confirm it fails, then updates calculator.py to\n      raise ValueError on a zero divisor, and verifies all tests pass. The bug fix\n      comes only after the failing test is confirmed, following the TDD rule that bugs\n      must be reproduced in a test before fixing.\n\n  - name: Completes full RED-GREEN-REFACTOR cycle\n    setup: >\n      rm -rf /tmp/vrd-tdd-3 && mkdir /tmp/vrd-tdd-3 && cd /tmp/vrd-tdd-3 &&\n      touch stack.py test_stack.py\n    cleanup: rm -rf /tmp/vrd-tdd-3\n    prompt: >\n      In /tmp/vrd-tdd-3, implement a Stack class with push(item) and pop() methods\n      using TDD. pop() should raise IndexError when the stack is empty.\n    expect: >\n      Agent follows the full RED-GREEN-REFACTOR cycle: writes one failing test at a\n      time (push, then pop, then empty-pop error), verifies each test fails before\n      implementing, writes minimal code to pass each test, and verifies green after\n      each addition. Tests are run and confirmed failing before any production code is\n      written for each new behavior. All tests pass at the end.\n\n  - name: Writes minimal implementation without over-engineering\n    setup: >\n      rm -rf /tmp/vrd-tdd-4 && mkdir /tmp/vrd-tdd-4 && cd /tmp/vrd-tdd-4 &&\n      touch utils.py test_utils.py\n    cleanup: rm -rf /tmp/vrd-tdd-4\n    prompt: >\n      In /tmp/vrd-tdd-4, implement an is_palindrome(s) function in utils.py using TDD.\n      It should return True if the string is a palindrome, False otherwise.\n      Case-insensitive comparison is not required.\n    expect: >\n      Agent writes a minimal failing test first, verifies it fails, then implements\n      is_palindrome with the simplest possible code (e.g. return s == s[::-1]),\n      not over-engineered with optional parameters, logging, or unused abstractions.\n      All tests pass and the implementation is concise and minimal per the GREEN rule.\n\n  - name: Applies TDD to fix in-range return bug\n    setup: >\n      rm -rf /tmp/vrd-tdd-5 && mkdir /tmp/vrd-tdd-5 && cd /tmp/vrd-tdd-5 &&\n      printf 'def clamp(value, min_val, max_val):\\n    if value < min_val:\\n        return min_val\\n    if value > max_val:\\n        return max_val\\n' > limits.py &&\n      printf 'import unittest\\nfrom limits import clamp\\n\\nclass TestClamp(unittest.TestCase):\\n    def test_clamp_below(self):\\n        self.assertEqual(clamp(5, 10, 20), 10)\\n    def test_clamp_above(self):\\n        self.assertEqual(clamp(25, 10, 20), 20)\\n\\nif __name__ == \"__main__\":\\n    unittest.main()\\n' > test_limits.py\n    cleanup: rm -rf /tmp/vrd-tdd-5\n    prompt: >\n      In /tmp/vrd-tdd-5, the clamp function has a bug: it returns None when the value\n      is within range (e.g. clamp(15, 10, 20) should return 15 but returns None).\n      Fix the bug using TDD.\n    expect: >\n      Agent writes a failing test for the in-range case (e.g. assertEqual(clamp(15, 10, 20), 15))\n      before touching limits.py, runs it to confirm it fails, adds the missing\n      return value statement to clamp, and verifies all three tests pass. The TDD\n      discipline holds even when the fix is immediately obvious.\n\nArchive v1.0.11: 16 files, 26996 bytes\n\nFiles: evaluate-skill.eval.yaml (6352b), REFERENCE.md (14044b), references/evals/claude-code-smoke/claude-code-smoke.eval.yaml (636b), references/evals/claude-code-smoke/SKILL.md (253b), references/evals/commit-simple/commit-simple.eval.yaml (4213b), references/evals/commit-simple/SKILL.md (1459b), references/evals/screenshot/screenshot.eval.yaml (1078b), references/evals/screenshot/SKILL.md (7754b), references/evals/summarize/SKILL.md (2181b), references/evals/summarize/summarize.eval.yaml (1748b), references/evals/tdd/SKILL.md (9867b), references/evals/tdd/tdd.eval.yaml (5005b), references/examples/simple.eval.yaml (672b), skill-card.md (2251b), SKILL.md (3650b), _meta.json (134b)\n\nFile v1.0.11:references/evals/claude-code-smoke/SKILL.md\n\n---\nname: claude-code-smoke\ndescription: Use when asked to run the Claude Code smoke evaluation.\n---\n\n# Claude Code Smoke\n\nWhen asked to run the smoke evaluation, write exactly this text:\n\n```text\nclaude-code-smoke-ok\n```\n\nto the requested output file.\n\nFile v1.0.11:references/evals/commit-simple/SKILL.md\n\n---\nname: commit-simple\ndescription: Branch, commit, and push changes. Use when preparing a commit or pushing work.\n---\n\n# Commit\n\nUse this skill when creating a branch, committing changes, or pushing work.\n\n## Workflow\n\n1. Inspect the current branch, working tree, staged changes, and diff.\n2. Propose any branch change, commit split, and Conventional Commit message(s).\n3. After confirmation, create the branch if needed and commit.\n4. Offer to push after the commit.\n5. If the user wants a pull request, suggest `commit-pr` as the next step.\n\n## Rules\n\n- If there are logically separate changes, propose separate commits and confirm the plan before committing.\n- If the current branch is `main` or `master`, ask whether to create a new branch before committing.\n- For new branches, use `{type}/sc-{number}/{slug}`, `{type}/gh-{number}/{slug}`, `{type}/{number}/{slug}`, or `{type}/{slug}`.\n- When working inside a ticket worktree folder, keep the branch aligned with the folder's ticket id.\n- Use Conventional Commits for commit messages.\n- Commit messages should describe the resulting code change, not the development process.\n- Unless the change is trivial, include a concise human-readable body that explains why the change matters and any important reviewer context.\n- Prefer concrete facts over workflow labels: name the behavior, API, module, or cleanup that changed.\n- Ask before committing or pushing.\n- Hand pull request work off to `commit-pr`.\n\nFile v1.0.11:references/evals/screenshot/SKILL.md\n\n---\nname: \"screenshot\"\ndescription: \"Use when the user explicitly asks for a desktop or system screenshot (full screen, specific app or window, or a pixel region), or when tool-specific capture capabilities are unavailable and an OS-level capture is needed.\"\n---\n\n\n# Screenshot Capture\n\nFollow these save-location rules every time:\n\n1) If the user specifies a path, save there.\n2) If the user asks for a screenshot without a path, save to the OS default screenshot location.\n3) If Codex needs a screenshot for its own inspection, save to the temp directory.\n\n## Tool priority\n\n- Prefer tool-specific screenshot capabilities when available (for example: a Figma MCP/skill for Figma files, or Playwright/agent-browser tools for browsers and Electron apps).\n- Use this skill when explicitly asked, for whole-system desktop captures, or when a tool-specific capture cannot get what you need.\n- Otherwise, treat this skill as the default for desktop apps without a better-integrated capture tool.\n\n## macOS permission preflight (reduce repeated prompts)\n\nOn macOS, run the preflight helper once before window/app capture. It checks\nScreen Recording permission, explains why it is needed, and requests it in one\nplace.\n\nThe helpers route Swift's module cache to `$TMPDIR/codex-swift-module-cache`\nto avoid extra sandbox module-cache prompts.\n\n```bash\nbash <path-to-skill>/scripts/ensure_macos_permissions.sh\n```\n\nTo avoid multiple sandbox approval prompts, combine preflight + capture in one\ncommand when possible:\n\n```bash\nbash <path-to-skill>/scripts/ensure_macos_permissions.sh && \\\npython3 <path-to-skill>/scripts/take_screenshot.py --app \"Codex\"\n```\n\nFor Codex inspection runs, keep the output in temp:\n\n```bash\nbash <path-to-skill>/scripts/ensure_macos_permissions.sh && \\\npython3 <path-to-skill>/scripts/take_screenshot.py --app \"<App>\" --mode temp\n```\n\nUse the bundled scripts to avoid re-deriving OS-specific commands.\n\n## macOS and Linux (Python helper)\n\nRun the helper from the repo root:\n\n```bash\npython3 <path-to-skill>/scripts/take_screenshot.py\n```\n\nCommon patterns:\n\n- Default location (user asked for \"a screenshot\"):\n\n```bash\npython3 <path-to-skill>/scripts/take_screenshot.py\n```\n\n- Temp location (Codex visual check):\n\n```bash\npython3 <path-to-skill>/scripts/take_screenshot.py --mode temp\n```\n\n- Explicit location (user provided a path or filename):\n\n```bash\npython3 <path-to-skill>/scripts/take_screenshot.py --path output/screen.png\n```\n\n- App/window capture by app name (macOS only; substring match is OK; captures all matching windows):\n\n```bash\npython3 <path-to-skill>/scripts/take_screenshot.py --app \"Codex\"\n```\n\n- Specific window title within an app (macOS only):\n\n```bash\npython3 <path-to-skill>/scripts/take_screenshot.py --app \"Codex\" --window-name \"Settings\"\n```\n\n- List matching window ids before capturing (macOS only):\n\n```bash\npython3 <path-to-skill>/scripts/take_screenshot.py --list-windows --app \"Codex\"\n```\n\n- Pixel region (x,y,w,h):\n\n```bash\npython3 <path-to-skill>/scripts/take_screenshot.py --mode temp --region 100,200,800,600\n```\n\n- Focused/active window (captures only the frontmost window; use `--app` to capture all windows):\n\n```bash\npython3 <path-to-skill>/scripts/take_screenshot.py --mode temp --active-window\n```\n\n- Specific window id (use --list-windows on macOS to discover ids):\n\n```bash\npython3 <path-to-skill>/scripts/take_screenshot.py --window-id 12345\n```\n\nThe script prints one path per capture. When multiple windows or displays match, it prints multiple paths (one per line) and adds suffixes like `-w<windowId>` or `-d<display>`. View each path sequentially with the image viewer tool, and only manipulate images if needed or requested.\n\n### Workflow examples\n\n- \"Take a look at <App> and tell me what you see\": capture to temp, then view each printed path in order.\n\n```bash\nbash <path-to-skill>/scripts/ensure_macos_permissions.sh && \\\npython3 <path-to-skill>/scripts/take_screenshot.py --app \"<App>\" --mode temp\n```\n\n- \"The design from Figma is not matching what is implemented\": use a Figma MCP/skill to capture the design first, then capture the running app with this skill (typically to temp) and compare the raw screenshots before any manipulation.\n\n### Multi-display behavior\n\n- On macOS, full-screen captures save one file per display when multiple monitors are connected.\n- On Linux and Windows, full-screen captures use the virtual desktop (all monitors in one image); use `--region` to isolate a single display when needed.\n\n### Linux prerequisites and selection logic\n\nThe helper automatically selects the first available tool:\n\n1) `scrot`\n2) `gnome-screenshot`\n3) ImageMagick `import`\n\nIf none are available, ask the user to install one of them and retry.\n\nCoordinate regions require `scrot` or ImageMagick `import`.\n\n`--app`, `--window-name`, and `--list-windows` are macOS-only. On Linux, use\n`--active-window` or provide `--window-id` when available.\n\n## Windows (PowerShell helper)\n\nRun the PowerShell helper:\n\n```powershell\npowershell -ExecutionPolicy Bypass -File <path-to-skill>/scripts/take_screenshot.ps1\n```\n\nCommon patterns:\n\n- Default location:\n\n```powershell\npowershell -ExecutionPolicy Bypass -File <path-to-skill>/scripts/take_screenshot.ps1\n```\n\n- Temp location (Codex visual check):\n\n```powershell\npowershell -ExecutionPolicy Bypass -File <path-to-skill>/scripts/take_screenshot.ps1 -Mode temp\n```\n\n- Explicit path:\n\n```powershell\npowershell -ExecutionPolicy Bypass -File <path-to-skill>/scripts/take_screenshot.ps1 -Path \"C:\\Temp\\screen.png\"\n```\n\n- Pixel region (x,y,w,h):\n\n```powershell\npowershell -ExecutionPolicy Bypass -File <path-to-skill>/scripts/take_screenshot.ps1 -Mode temp -Region 100,200,800,600\n```\n\n- Active window (ask the user to focus it first):\n\n```powershell\npowershell -ExecutionPolicy Bypass -File <path-to-skill>/scripts/take_screenshot.ps1 -Mode temp -ActiveWindow\n```\n\n- Specific window handle (only when provided):\n\n```powershell\npowershell -ExecutionPolicy Bypass -File <path-to-skill>/scripts/take_screenshot.ps1 -WindowHandle 123456\n```\n\n## Direct OS commands (fallbacks)\n\nUse these when you cannot run the helpers.\n\n### macOS\n\n- Full screen to a specific path:\n\n```bash\nscreencapture -x output/screen.png\n```\n\n- Pixel region:\n\n```bash\nscreencapture -x -R100,200,800,600 output/region.png\n```\n\n- Specific window id:\n\n```bash\nscreencapture -x -l12345 output/window.png\n```\n\n- Interactive selection or window pick:\n\n```bash\nscreencapture -x -i output/interactive.png\n```\n\n### Linux\n\n- Full screen:\n\n```bash\nscrot output/screen.png\n```\n\n```bash\ngnome-screenshot -f output/screen.png\n```\n\n```bash\nimport -window root output/screen.png\n```\n\n- Pixel region:\n\n```bash\nscrot -a 100,200,800,600 output/region.png\n```\n\n```bash\nimport -window root -crop 800x600+100+200 output/region.png\n```\n\n- Active window:\n\n```bash\nscrot -u output/window.png\n```\n\n```bash\ngnome-screenshot -w -f output/window.png\n```\n\n## Error handling\n\n- On macOS, run `bash <path-to-skill>/scripts/ensure_macos_permissions.sh` first to request Screen Recording in one place.\n- If you see \"screen capture checks are blocked in the sandbox\", \"could not create image from display\", or Swift `ModuleCache` permission errors in a sandboxed run, rerun the command with escalated permissions.\n- If macOS app/window capture returns no matches, run `--list-windows --app \"AppName\"` and retry with `--window-id`, and make sure the app is visible on screen.\n- If Linux region/window capture fails, check tool availability with `command -v scrot`, `command -v gnome-screenshot`, and `command -v import`.\n- If saving to the OS default location fails with permission errors in a sandbox, rerun the command with escalated permissions.\n- Always report the saved file path in the response.\n\nFile v1.0.11:references/evals/summarize/SKILL.md\n\n---\nname: summarize\ndescription: Summarize or transcribe URLs, YouTube/videos, podcasts, articles, transcripts, PDFs, and local files.\nhomepage: https://summarize.sh\nmetadata:\n  {\n    \"openclaw\":\n      {\n        \"emoji\": \"🧾\",\n        \"requires\": { \"bins\": [\"summarize\"] },\n        \"install\":\n          [\n            {\n              \"id\": \"brew\",\n              \"kind\": \"brew\",\n              \"formula\": \"steipete/tap/summarize\",\n              \"bins\": [\"summarize\"],\n              \"label\": \"Install summarize (brew)\",\n            },\n          ],\n      },\n  }\n---\n\n# Summarize\n\nFast CLI to summarize URLs, local files, and YouTube links.\n\n## When to use (trigger phrases)\n\nUse this skill immediately when the user asks any of:\n\n- \"use summarize.sh\"\n- \"what's this link/video about?\"\n- \"summarize this URL/article\"\n- \"transcribe this YouTube/video\" (best-effort transcript extraction; no `yt-dlp` needed)\n\n## Quick start\n\n```bash\nsummarize \"https://example.com\" --model google/gemini-3-flash-preview\nsummarize \"/path/to/file.pdf\" --model google/gemini-3-flash-preview\nsummarize \"https://youtu.be/dQw4w9WgXcQ\" --youtube auto\n```\n\n## YouTube: summary vs transcript\n\nBest-effort transcript (URLs only):\n\n```bash\nsummarize \"https://youtu.be/dQw4w9WgXcQ\" --youtube auto --extract-only\n```\n\nIf the user asked for a transcript but it's huge, return a tight summary first, then ask which section/time range to expand.\n\n## Model + keys\n\nSet the API key for your chosen provider:\n\n- OpenAI: `OPENAI_API_KEY`\n- Anthropic: `ANTHROPIC_API_KEY`\n- xAI: `XAI_API_KEY`\n- Google: `GEMINI_API_KEY` (aliases: `GOOGLE_GENERATIVE_AI_API_KEY`, `GOOGLE_API_KEY`)\n\nDefault model is `google/gemini-3-flash-preview` if none is set.\n\n## Useful flags\n\n- `--length short|medium|long|xl|xxl|<chars>`\n- `--max-output-tokens <count>`\n- `--extract-only` (URLs only)\n- `--json` (machine readable)\n- `--firecrawl auto|off|always` (fallback extraction)\n- `--youtube auto` (Apify fallback if `APIFY_API_TOKEN` set)\n\n## Config\n\nOptional config file: `~/.summarize/config.json`\n\n```json\n{ \"model\": \"openai/gpt-5.2\" }\n```\n\nOptional services:\n\n- `FIRECRAWL_API_KEY` for blocked sites\n- `APIFY_API_TOKEN` for YouTube fallback\n\nFile v1.0.11:references/evals/tdd/SKILL.md\n\n---\nname: test-driven-development\ndescription: Use when implementing any feature or bugfix, before writing implementation code\n---\n\n# Test-Driven Development (TDD)\n\n## Overview\n\nWrite the test first. Watch it fail. Write minimal code to pass.\n\n**Core principle:** If you didn't watch the test fail, you don't know if it tests the right thing.\n\n**Violating the letter of the rules is violating the spirit of the rules.**\n\n## When to Use\n\n**Always:**\n- New features\n- Bug fixes\n- Refactoring\n- Behavior changes\n\n**Exceptions (ask your human partner):**\n- Throwaway prototypes\n- Generated code\n- Configuration files\n\nThinking \"skip TDD just this once\"? Stop. That's rationalization.\n\n## The Iron Law\n\n```\nNO PRODUCTION CODE WITHOUT A FAILING TEST FIRST\n```\n\nWrite code before the test? Delete it. Start over.\n\n**No exceptions:**\n- Don't keep it as \"reference\"\n- Don't \"adapt\" it while writing tests\n- Don't look at it\n- Delete means delete\n\nImplement fresh from tests. Period.\n\n## Red-Green-Refactor\n\n```dot\ndigraph tdd_cycle {\n    rankdir=LR;\n    red [label=\"RED\\nWrite failing test\", shape=box, style=filled, fillcolor=\"#ffcccc\"];\n    verify_red [label=\"Verify fails\\ncorrectly\", shape=diamond];\n    green [label=\"GREEN\\nMinimal code\", shape=box, style=filled, fillcolor=\"#ccffcc\"];\n    verify_green [label=\"Verify passes\\nAll green\", shape=diamond];\n    refactor [label=\"REFACTOR\\nClean up\", shape=box, style=filled, fillcolor=\"#ccccff\"];\n    next [label=\"Next\", shape=ellipse];\n\n    red -> verify_red;\n    verify_red -> green [label=\"yes\"];\n    verify_red -> red [label=\"wrong\\nfailure\"];\n    green -> verify_green;\n    verify_green -> refactor [label=\"yes\"];\n    verify_green -> green [label=\"no\"];\n    refactor -> verify_green [label=\"stay\\ngreen\"];\n    verify_green -> next;\n    next -> red;\n}\n```\n\n### RED - Write Failing Test\n\nWrite one minimal test showing what should happen.\n\n<Good>\n```typescript\ntest('retries failed operations 3 times', async () => {\n  let attempts = 0;\n  const operation = () => {\n    attempts++;\n    if (attempts < 3) throw new Error('fail');\n    return 'success';\n  };\n\n  const result = await retryOperation(operation);\n\n  expect(result).toBe('success');\n  expect(attempts).toBe(3);\n});\n```\nClear name, tests real behavior, one thing\n</Good>\n\n<Bad>\n```typescript\ntest('retry works', async () => {\n  const mock = jest.fn()\n    .mockRejectedValueOnce(new Error())\n    .mockRejectedValueOnce(new Error())\n    .mockResolvedValueOnce('success');\n  await retryOperation(mock);\n  expect(mock).toHaveBeenCalledTimes(3);\n});\n```\nVague name, tests mock not code\n</Bad>\n\n**Requirements:**\n- One behavior\n- Clear name\n- Real code (no mocks unless unavoidable)\n\n### Verify RED - Watch It Fail\n\n**MANDATORY. Never skip.**\n\n```bash\nnpm test path/to/test.test.ts\n```\n\nConfirm:\n- Test fails (not errors)\n- Failure message is expected\n- Fails because feature missing (not typos)\n\n**Test passes?** You're testing existing behavior. Fix test.\n\n**Test errors?** Fix error, re-run until it fails correctly.\n\n### GREEN - Minimal Code\n\nWrite simplest code to pass the test.\n\n<Good>\n```typescript\nasync function retryOperation<T>(fn: () => Promise<T>): Promise<T> {\n  for (let i = 0; i < 3; i++) {\n    try {\n      return await fn();\n    } catch (e) {\n      if (i === 2) throw e;\n    }\n  }\n  throw new Error('unreachable');\n}\n```\nJust enough to pass\n</Good>\n\n<Bad>\n```typescript\nasync function retryOperation<T>(\n  fn: () => Promise<T>,\n  options?: {\n    maxRetries?: number;\n    backoff?: 'linear' | 'exponential';\n    onRetry?: (attempt: number) => void;\n  }\n): Promise<T> {\n  // YAGNI\n}\n```\nOver-engineered\n</Bad>\n\nDon't add features, refactor other code, or \"improve\" beyond the test.\n\n### Verify GREEN - Watch It Pass\n\n**MANDATORY.**\n\n```bash\nnpm test path/to/test.test.ts\n```\n\nConfirm:\n- Test passes\n- Other tests still pass\n- Output pristine (no errors, warnings)\n\n**Test fails?** Fix code, not test.\n\n**Other tests fail?** Fix now.\n\n### REFACTOR - Clean Up\n\nAfter green only:\n- Remove duplication\n- Improve names\n- Extract helpers\n\nKeep tests green. Don't add behavior.\n\n### Repeat\n\nNext failing test for next feature.\n\n## Good Tests\n\n| Quality | Good | Bad |\n|---------|------|-----|\n| **Minimal** | One thing. \"and\" in name? Split it. | `test('validates email and domain and whitespace')` |\n| **Clear** | Name describes behavior | `test('test1')` |\n| **Shows intent** | Demonstrates desired API | Obscures what code should do |\n\n## Why Order Matters\n\n**\"I'll write tests after to verify it works\"**\n\nTests written after code pass immediately. Passing immediately proves nothing:\n- Might test wrong thing\n- Might test implementation, not behavior\n- Might miss edge cases you forgot\n- You never saw it catch the bug\n\nTest-first forces you to see the test fail, proving it actually tests something.\n\n**\"I already manually tested all the edge cases\"**\n\nManual testing is ad-hoc. You think you tested everything but:\n- No record of what you tested\n- Can't re-run when code changes\n- Easy to forget cases under pressure\n- \"It worked when I tried it\" ≠ comprehensive\n\nAutomated tests are systematic. They run the same way every time.\n\n**\"Deleting X hours of work is wasteful\"**\n\nSunk cost fallacy. The time is already gone. Your choice now:\n- Delete and rewrite with TDD (X more hours, high confidence)\n- Keep it and add tests after (30 min, low confidence, likely bugs)\n\nThe \"waste\" is keeping code you can't trust. Working code without real tests is technical debt.\n\n**\"TDD is dogmatic, being pragmatic means adapting\"**\n\nTDD IS pragmatic:\n- Finds bugs before commit (faster than debugging after)\n- Prevents regressions (tests catch breaks immediately)\n- Documents behavior (tests show how to use code)\n- Enables refactoring (change freely, tests catch breaks)\n\n\"Pragmatic\" shortcuts = debugging in production = slower.\n\n**\"Tests after achieve the same goals - it's spirit not ritual\"**\n\nNo. Tests-after answer \"What does this do?\" Tests-first answer \"What should this do?\"\n\nTests-after are biased by your implementation. You test what you built, not what's required. You verify remembered edge cases, not discovered ones.\n\nTests-first force edge case discovery before implementing. Tests-after verify you remembered everything (you didn't).\n\n30 minutes of tests after ≠ TDD. You get coverage, lose proof tests work.\n\n## Common Rationalizations\n\n| Excuse | Reality |\n|--------|---------|\n| \"Too simple to test\" | Simple code breaks. Test takes 30 seconds. |\n| \"I'll test after\" | Tests passing immediately prove nothing. |\n| \"Tests after achieve same goals\" | Tests-after = \"what does this do?\" Tests-first = \"what should this do?\" |\n| \"Already manually tested\" | Ad-hoc ≠ systematic. No record, can't re-run. |\n| \"Deleting X hours is wasteful\" | Sunk cost fallacy. Keeping unverified code is technical debt. |\n| \"Keep as reference, write tests first\" | You'll adapt it. That's testing after. Delete means delete. |\n| \"Need to explore first\" | Fine. Throw away exploration, start with TDD. |\n| \"Test hard = design unclear\" | Listen to test. Hard to test = hard to use. |\n| \"TDD will slow me down\" | TDD faster than debugging. Pragmatic = test-first. |\n| \"Manual test faster\" | Manual doesn't prove edge cases. You'll re-test every change. |\n| \"Existing code has no tests\" | You're improving it. Add tests for existing code. |\n\n## Red Flags - STOP and Start Over\n\n- Code before test\n- Test after implementation\n- Test passes immediately\n- Can't explain why test failed\n- Tests added \"later\"\n- Rationalizing \"just this once\"\n- \"I already manually tested it\"\n- \"Tests after achieve the same purpose\"\n- \"It's about spirit not ritual\"\n- \"Keep as re\n\nArchive v1.0.10: 16 files, 25926 bytes\n\nFiles: evaluate-skill.eval.yaml (6352b), REFERENCE.md (11094b), references/evals/claude-code-smoke/claude-code-smoke.eval.yaml (636b), references/evals/claude-code-smoke/SKILL.md (253b), references/evals/commit-simple/commit-simple.eval.yaml (4213b), references/evals/commit-simple/SKILL.md (1459b), references/evals/screenshot/screenshot.eval.yaml (1078b), references/evals/screenshot/SKILL.md (7754b), references/evals/summarize/SKILL.md (2181b), references/evals/summarize/summarize.eval.yaml (1748b), references/evals/tdd/SKILL.md (9867b), references/evals/tdd/tdd.eval.yaml (5005b), references/examples/simple.eval.yaml (672b), skill-card.md (2813b), SKILL.md (3650b), _meta.json (134b)\n\nArchive v1.0.9: 16 files, 24924 bytes\n\nFiles: evaluate-skill.eval.yaml (6352b), REFERENCE.md (9381b), references/evals/claude-code-smoke/claude-code-smoke.eval.yaml (636b), references/evals/claude-code-smoke/SKILL.md (253b), references/evals/commit-simple/commit-simple.eval.yaml (4213b), references/evals/commit-simple/SKILL.md (1459b), references/evals/screenshot/screenshot.eval.yaml (1078b), references/evals/screenshot/SKILL.md (7754b), references/evals/summarize/SKILL.md (2181b), references/evals/summarize/summarize.eval.yaml (1748b), references/evals/tdd/SKILL.md (9867b), references/evals/tdd/tdd.eval.yaml (5005b), references/examples/simple.eval.yaml (672b), skill-card.md (2179b), SKILL.md (3650b), _meta.json (133b)\n\nArchive v1.0.8: 16 files, 24579 bytes\n\nFiles: evaluate-skill.eval.yaml (6352b), REFERENCE.md (8470b), references/evals/claude-code-smoke/claude-code-smoke.eval.yaml (636b), references/evals/claude-code-smoke/SKILL.md (253b), references/evals/commit-simple/commit-simple.eval.yaml (4213b), references/evals/commit-simple/SKILL.md (1459b), references/evals/screenshot/screenshot.eval.yaml (1078b), references/evals/screenshot/SKILL.md (7754b), references/evals/summarize/SKILL.md (2181b), references/evals/summarize/summarize.eval.yaml (1748b), references/evals/tdd/SKILL.md (9867b), references/evals/tdd/tdd.eval.yaml (5005b), references/examples/simple.eval.yaml (672b), skill-card.md (2198b), SKILL.md (3650b), _meta.json (133b)\n\nArchive v1.0.7: 16 files, 24474 bytes\n\nFiles: evaluate-skill.eval.yaml (5818b), REFERENCE.md (8273b), references/evals/claude-code-smoke/claude-code-smoke.eval.yaml (760b), references/evals/claude-code-smoke/SKILL.md (253b), references/evals/commit-simple/commit-simple.eval.yaml (4321b), references/evals/commit-simple/SKILL.md (1459b), references/evals/screenshot/screenshot.eval.yaml (1120b), references/evals/screenshot/SKILL.md (7754b), references/evals/summarize/SKILL.md (2181b), references/evals/summarize/summarize.eval.yaml (1864b), references/evals/tdd/SKILL.md (9867b), references/evals/tdd/tdd.eval.yaml (5103b), references/examples/simple.eval.yaml (778b), skill-card.md (2241b), SKILL.md (3647b), _meta.json (133b)\n\nArchive v1.0.6: 16 files, 24323 bytes\n\nFiles: evaluate-skill.eval.yaml (7304b), REFERENCE.md (7434b), references/evals/claude-code-smoke/claude-code-smoke.eval.yaml (760b), references/evals/claude-code-smoke/SKILL.md (253b), references/evals/commit-simple/commit-simple.eval.yaml (4321b), references/evals/commit-simple/SKILL.md (1459b), references/evals/screenshot/screenshot.eval.yaml (1120b), references/evals/screenshot/SKILL.md (7754b), references/evals/summarize/SKILL.md (2181b), references/evals/summarize/summarize.eval.yaml (1864b), references/evals/tdd/SKILL.md (9867b), references/evals/tdd/tdd.eval.yaml (5103b), references/examples/simple.eval.yaml (778b), skill-card.md (2285b), SKILL.md (3707b), _meta.json (133b)\n\nArchive v1.0.5: 16 files, 24300 bytes\n\nFiles: evaluate-skill.eval.yaml (7304b), REFERENCE.md (4108b), references/evals/claude-code-smoke/claude-code-smoke.eval.yaml (760b), references/evals/claude-code-smoke/SKILL.md (253b), references/evals/commit-simple/commit-simple.eval.yaml (4321b), references/evals/commit-simple/SKILL.md (1459b), references/evals/screenshot/screenshot.eval.yaml (1120b), references/evals/screenshot/SKILL.md (7754b), references/evals/summarize/SKILL.md (2181b), references/evals/summarize/summarize.eval.yaml (1864b), references/evals/tdd/SKILL.md (9867b), references/evals/tdd/tdd.eval.yaml (5103b), references/examples/simple.eval.yaml (778b), skill-card.md (2542b), SKILL.md (6890b), _meta.json (133b)","readmeExcerpt":"Skill: evaluate-skill Owner: edonadei Summary: Measure a skill's reliability — run it k times for a pass@k score, design or interpret its eval, or compare it against the base agent. Use when the user wants to run, design, or interpret a skill's eval, or write an .eval.yaml spec. Tags: latest:1.0.14 Version history: v1.0.14 | 2026-09-25T15:12:56.018Z | user Changed - Clean-up from the 1.0.13 that added way too much fi","codeSnippets":[],"executableExamples":[{"language":"bash","snippet":"pipx install caliper-eval"},{"language":"yaml","snippet":"skill:\n  path: ./SKILL.md      # relative to the spec file\ntasks:\n  - name: What success looks like\n    prompt: <prompt sent to the skill under test>\n    expect: <natural-language pass/fail criterion>\n    assert: |           # optional deterministic Python check\n      assert ..."},{"language":"bash","snippet":"caliper run path/to/spec.eval.yaml --k 3\ncaliper run path/to/spec.eval.yaml --k 3 --ablate my-skill   # same tasks, that skill removed (or an mcp: server)\ncaliper run path/to/spec.eval.yaml --k 3 --ablate a --ablate b  # repeatable; name them all (+ --no-user-customizations) for the bare agent\ncaliper run path/to/spec.eval.yaml --verbose             # show per-attempt reasoning\ncaliper run path/to/spec.eval.yaml --no-user-customizations      # portable: only declared skills and servers\ncaliper run path/to/spec.eval.yaml --user-customizations         # load your user customizations even if the spec pins user_customizations: false\n\n# Choose the engine at run time — it is not stored in the spec (default: claude-code)\ncaliper run path/to/spec.eval.yaml --model codex:gpt-5-codex\ncaliper run path/to/spec.eval.yaml --model codex                # backend only, its default model\ncaliper run path/to/spec.eval.yaml --model claude-sonnet-4-6    # model only, backend stays claude-code\ncaliper run path/to/spec.eval.yaml --judge-model claude-code:claude-haiku-4-5-20251001\ncaliper run path/to/spec.eval.yaml --model codex --judge-model claude-code:claude-haiku-4-5-20251001"},{"language":"bash","snippet":"caliper validate path/to/spec.eval.yaml"},{"language":"bash","snippet":"caliper list                        # all specs with latest scores\ncaliper list my-skill-eval          # all runs for one spec: Run id + which were ablated\ncaliper report my-skill-eval        # latest run (table view)\ncaliper report my-skill-eval --run 2026-05-12T14-23-01Z  # specific run\ncaliper report results.json --format json"},{"language":"bash","snippet":"caliper compare full-eval short-eval          # latest run of each spec\ncaliper compare a.json b.json                 # pin specific runs\ncaliper compare full-eval short-eval --format json   # for a ship/no-ship gate"}],"parameters":null,"dependencies":[],"permissions":[],"extractedFiles":[{"path":"SKILL.md","content":"---\nname: evaluate-skill\ndescription: Measure a skill's reliability — run it k times for a pass@k score, design or interpret its eval, or compare it against the base agent. Use when the user wants to run, design, or interpret a skill's eval, or write an .eval.yaml spec.\nallowed-tools: Bash\n---\n\n# Evaluate Skill\n\nRun a skill repeatedly to measure how reliably it works, and design the evals that measure it.\n\n## Prerequisites\n\nThe `caliper` CLI must be on `PATH`. This skill can be copied into an agent without the Caliper repo, so do not assume the CLI is packaged with it. Install if missing:\n\n```bash\npipx install caliper-eval\n```\n\nThe engine (backend + model) is not part of the spec — it is chosen at run time with `--model` (skill) and `--judge-model` (judge), independently, from `claude-code`, `codex`, `pi`, defaulting to `claude-code`. Every backend is a CLI agent that uses its own subscription/auth; there is no direct-API backend (for API billing, configure a CLI with an API key). Full per-backend detail and every command: [REFERENCE.md](REFERENCE.md).\n\n## Spec shape\n\nAn `.eval.yaml` names the skill and a list of tasks. Keep `skill.path` relative to the spec file (usually `./SKILL.md`):\n\n```yaml\nskill:\n  path: ./SKILL.md      # relative to the spec file\ntasks:\n  - name: What success looks like\n    prompt: <prompt sent to the skill under test>\n    expect: <natural-language pass/fail criterion>\n    assert: |           # optional deterministic Python check\n      assert ...\n```\n\nThe spec has no `backend`/`model` or `judge:` block; pick the engine when you run, e.g. `caliper run <spec> --model codex --judge-model codex`. The full format (setup/cleanup, external assert scripts, sandbox) is in [REFERENCE.md](REFERENCE.md).\n\n## Bundled references\n\n`references/evals/` holds complete real examples (Claude Code smoke, commit workflow, screenshot, summarization, TDD) — each folder self-contained with its fixture `SKILL.md` and `.eval.yaml`. `references/simple.eval.yaml` is one compact multi-task spec.\n\n## No eval yet?\n\nIf the skill has a `SKILL.md` but no `.eval.yaml`, suggest the `grill-skill` workflow — it interviews the user and generates a happy/edge/adversarial spec. Use `evaluate-skill` directly when a spec already exists and the user wants to run, validate, report, or extend it.\n\n## Designing good evals\n\n1. Name the target behavior — what should the skill do better than the base agent?\n2. Decide whether the suite is a capability eval or a regression eval.\n3. Cover normal, edge, and adversarial cases when the behavior matters.\n4. Grade artifacts (files, git state, command output, exact values) whenever you can; judge the transcript only when the behavior itself is the point. The full artifact-vs-transcript rules, the task-quality checklist, common eval patterns, and how to write `expect:` rubrics live in [REFERENCE.md](REFERENCE.md) — read and apply them when designing tasks.\n5. Run once with `--ablate <skill-name>` and `caliper compare` the two runs, "},{"path":"_meta.json","content":"{\n  \"ownerId\": \"kn7bp63rpwq0rm2g7m4k0c3hjn80qkhc\",\n  \"slug\": \"evaluate-skill\",\n  \"version\": \"1.0.14\",\n  \"publishedAt\": 1790349176018\n}"},{"path":"REFERENCE.md","content":"# Caliper Reference\n\n## Commands\n\n### Run an evaluation\n```bash\ncaliper run path/to/spec.eval.yaml --k 3\ncaliper run path/to/spec.eval.yaml --k 3 --ablate my-skill   # same tasks, that skill removed (or an mcp: server)\ncaliper run path/to/spec.eval.yaml --k 3 --ablate a --ablate b  # repeatable; name them all (+ --no-user-customizations) for the bare agent\ncaliper run path/to/spec.eval.yaml --verbose             # show per-attempt reasoning\ncaliper run path/to/spec.eval.yaml --no-user-customizations      # portable: only declared skills and servers\ncaliper run path/to/spec.eval.yaml --user-customizations         # load your user customizations even if the spec pins user_customizations: false\n\n# Choose the engine at run time — it is not stored in the spec (default: claude-code)\ncaliper run path/to/spec.eval.yaml --model codex:gpt-5-codex\ncaliper run path/to/spec.eval.yaml --model codex                # backend only, its default model\ncaliper run path/to/spec.eval.yaml --model claude-sonnet-4-6    # model only, backend stays claude-code\ncaliper run path/to/spec.eval.yaml --judge-model claude-code:claude-haiku-4-5-20251001\ncaliper run path/to/spec.eval.yaml --model codex --judge-model claude-code:claude-haiku-4-5-20251001\n```\n\n### Validate a spec file\n```bash\ncaliper validate path/to/spec.eval.yaml\n```\n\nRejects unknown task and `sandbox:` keys (a typo like `asert:`), `forbidden_files` entries that are not valid regexes, and `assert:` script files missing from beside the spec. `caliper run` runs the same checks before its first attempt.\n\n### Browse saved results\n```bash\ncaliper list                        # all specs with latest scores\ncaliper list my-skill-eval          # all runs for one spec: Run id + which were ablated\ncaliper report my-skill-eval        # latest run (table view)\ncaliper report my-skill-eval --run 2026-05-12T14-23-01Z  # specific run\ncaliper report results.json --format json\n```\n\n### Compare two runs (ablation)\nDiff two already-saved runs of the same eval — full vs. shortened skill, or the\nsame skill over time. Tasks are matched by name; `Δ = b − a`; a negative Δ flags\na regression; a side with no usable attempts shows `—` (unmeasured, never a\nregression); the headline `Δ (matched)` averages only tasks measured on both\nsides. Each argument is addressed like `report` (spec name → latest run, or a\nresults-JSON path); pin a historical run by naming its path.\n```bash\ncaliper compare full-eval short-eval          # latest run of each spec\ncaliper compare a.json b.json                 # pin specific runs\ncaliper compare full-eval short-eval --format json   # for a ship/no-ship gate\n```\n\n## Spec format (.eval.yaml)\n\nThe spec carries no engine — no `backend`/`model` and no `judge:` block. Backend\nand model for both the skill and the judge are chosen at run time via `--model` /\n`--judge-model` (default `claude-code`); a spec that still pins these keys fails\nvalidation with a message pointing at the flags.\n\n```yaml\nskills:                 "},{"path":"skill-card.md","content":"## Description:\n\nMeasures a skill's reliability through repeated evaluations, eval design and interpretation, and comparisons against the base agent.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[edonadei](https://clawhub.ai/user/edonadei)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers use this skill to design and run Caliper evaluations, measure repeatability, and compare a skill with an agent running without it.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: Third-party evaluation specs can execute setup, cleanup, assertions, connectors, or git-sourced content.\n\nMitigation: Review setup, cleanup, assert, mcp, and git-source entries before running a spec.\n\nRisk: User customizations can change results or make shared comparisons misleading.\n\nMitigation: Use isolated runs when comparing backends, sharing measurements, or measuring the bare agent.\n\nRisk: Saved evaluation transcripts and snapshots may contain sensitive project or prompt data.\n\nMitigation: Keep .caliper/results out of commits and review saved results before sharing.\n\n## Reference(s):\n\n- [Evaluate Skill on ClawHub](https://clawhub.ai/edonadei/skills/evaluate-skill)\n- [Caliper Reference](artifact/REFERENCE.md)\n\n## Skill Output:\n\n**Output Type(s):** [Guidance, YAML configuration, Shell commands, Evaluation reports]\n\n**Output Format:** [Markdown guidance, .eval.yaml specs, and Caliper results]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [Can create evaluation specs and save run results for later comparison.]\n\n## Skill Version(s):\n\n1.0.14 (source: ClawHub release metadata)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment."},{"path":"evaluate-skill.eval.yaml","content":"skills:\n  - ./SKILL.md\n\n# The engine (backend + model) is a runtime axis, not a spec field — run this\n# eval with `--model codex --judge-model codex` (or any other engine).\n\n# No explicit sandbox.forbidden_files: caliper auto-forbids this spec and any\n# .caliper/results/ directory. Listing \"./.caliper/.*\" by hand risks a false\n# cheat flag, because these tasks legitimately WRITE caliper specs whose text\n# contains that very pattern — see the same note in grill-skill.eval.yaml.\n\ntasks:\n  - name: Validates a well-formed spec and reports it as valid\n    activates: [evaluate-skill]\n    setup: |\n      cat > /tmp/caliper-test-valid.eval.yaml << 'EOF'\n      skills:\n        - ./SKILL.md\n      tasks:\n        - name: Test arithmetic\n          prompt: What is 2 + 2?\n          expect: The assistant answers 4.\n      EOF\n    cleanup: rm -f /tmp/caliper-test-valid.eval.yaml\n    prompt: >\n      Use caliper to validate the spec file at /tmp/caliper-test-valid.eval.yaml\n      and tell me whether it is valid.\n    expect: >\n      The agent runs caliper validate on /tmp/caliper-test-valid.eval.yaml and\n      reports that the spec is valid with no errors.\n    assert: |\n      import subprocess\n\n      result = subprocess.run(\n          [\"caliper\", \"validate\", \"/tmp/caliper-test-valid.eval.yaml\"],\n          capture_output=True, text=True\n      )\n      assert result.returncode == 0, f\"caliper validate exited {result.returncode}: {result.stderr}\"\n\n  - name: Identifies errors in an invalid spec without fixing it\n    activates: [evaluate-skill]\n    setup: |\n      cat > /tmp/caliper-test-invalid.eval.yaml << 'EOF'\n      skills:\n        - ./SKILL.md\n      tasks:\n        - name: Broken task\n          prompt: do something\n      EOF\n    cleanup: rm -f /tmp/caliper-test-invalid.eval.yaml\n    prompt: >\n      Use caliper to validate the spec file at /tmp/caliper-test-invalid.eval.yaml\n      and tell me what errors it contains.\n    expect: >\n      The agent runs caliper validate and reports that the spec is invalid,\n      describing that the task is missing both expect and assert fields. The\n      agent does not edit the invalid spec file.\n\n  - name: Creates an engineless eval spec and defers the engine to run time\n    activates: [evaluate-skill]\n    cleanup: rm -f /tmp/caliper-created.eval.yaml\n    prompt: >\n      Create an evaluation spec file at /tmp/caliper-created.eval.yaml for a\n      skill at ./SKILL.md that I intend to run on the Codex CLI. Include one task\n      named \"Answers arithmetic\", with prompt \"What is 2 + 2?\" and expectation\n      \"The assistant answers 4.\" Then tell me how to run it against Codex.\n    expect: >\n      A valid .eval.yaml file is written at /tmp/caliper-created.eval.yaml with\n      a top-level skills: list containing ./SKILL.md and the requested task. The\n      spec does NOT pin a backend or model (no skill.backend/model, no judge\n      block) because the engine is a runtime axis, and the agent tells the user\n      to select Codex at run time with `ca"}],"languages":[],"docsSourceLabel":"CLAWHUB","editorialOverview":null,"editorialQuality":{"score":100,"threshold":65,"status":"thin","wordCount":2470,"uniquenessScore":34,"reasons":["uniqueness-below-45"]}},"media":{"evidence":{"source":"no-media","verified":false,"confidence":"low","updatedAt":"2026-10-11T02:16:22.404Z","emptyReason":"No screenshots, media assets, or demo links are available."},"primaryImageUrl":null,"mediaAssetCount":0,"assets":[],"demoUrl":null},"ownerResources":{"evidence":{"source":"unclaimed","verified":false,"confidence":"low","updatedAt":"2026-10-11T02:16:22.404Z","emptyReason":"This page has not been claimed by the agent owner."},"hasCustomPage":false,"customPageUpdatedAt":null,"customLinks":[],"structuredLinks":{"docsUrl":null,"demoUrl":null,"supportUrl":null,"pricingUrl":null,"statusUrl":null},"customPage":null},"relatedAgents":{"evidence":{"source":"protocol-neighbors","verified":false,"confidence":"medium","updatedAt":"2026-10-11T04:33:35.287Z","emptyReason":null},"items":[{"id":"8ebccd8e-3863-4187-8355-c3f14e1f9edf","entityType":"agent","canonicalPath":"/agent/iofficeai-aionui","slug":"iofficeai-aionui","name":"AionUi","description":"Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!","url":"https://github.com/iOfficeAI/AionUi","homepage":"https://www.aionui.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-10-09T19:11:12.944Z","createdAt":"2026-02-25T03:38:16.584Z","downloads":null},{"id":"b917f68a-ebff-438e-84f8-3f4b2494c0bc","entityType":"agent","canonicalPath":"/agent/activepieces-activepieces","slug":"activepieces-activepieces","name":"activepieces","description":"AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents","url":"https://github.com/activepieces/activepieces","homepage":"https://www.activepieces.com","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-15T02:22:12.426Z","createdAt":"2026-02-25T03:38:12.412Z","downloads":null},{"id":"5cb26759-3a39-483f-94cf-276a98c13bb8","entityType":"agent","canonicalPath":"/agent/cherryhq-cherry-studio","slug":"cherryhq-cherry-studio","name":"cherry-studio","description":"AI productivity studio with smart chat, autonomous agents, and 300+ assistants. Unified access to frontier LLMs","url":"https://github.com/CherryHQ/cherry-studio","homepage":"https://cherry-ai.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-11T14:38:40.986Z","createdAt":"2026-02-25T03:38:19.379Z","downloads":null},{"id":"6f6582d0-5d76-4f0f-b81d-86520247950b","entityType":"agent","canonicalPath":"/agent/copilotkit-copilotkit","slug":"copilotkit-copilotkit","name":"CopilotKit","description":"The Frontend for Agents & Generative UI. React + Angular","url":"https://github.com/CopilotKit/CopilotKit","homepage":"https://docs.copilotkit.ai","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-03-25T09:50:57.846Z","createdAt":"2026-02-25T03:39:14.617Z","downloads":null}],"links":{"hub":"/agent","source":"/agent/source/clawhub","protocols":[{"label":"OpenClaw","href":"/agent/protocol/openclew"}]}}}