{"slug":"clawhub-edonadei-evaluate-skill","facts":[{"factKey":"vendor","category":"vendor","label":"Vendor","value":"Clawhub","href":"https://clawhub.ai/edonadei/skills/evaluate-skill","sourceUrl":"https://clawhub.ai/edonadei/skills/evaluate-skill","sourceType":"profile","confidence":"medium","observedAt":"2026-10-11T02:16:22.404Z","isPublic":true},{"factKey":"protocols","category":"compatibility","label":"Protocol compatibility","value":"OpenClaw","href":"https://www.xpersona.co/api/v1/agents/clawhub-edonadei-evaluate-skill/contract","sourceUrl":"https://www.xpersona.co/api/v1/agents/clawhub-edonadei-evaluate-skill/contract","sourceType":"contract","confidence":"medium","observedAt":"2026-10-11T02:16:22.404Z","isPublic":true},{"factKey":"traction","category":"adoption","label":"Adoption signal","value":"1.2K downloads","href":"https://clawhub.ai/edonadei/evaluate-skill","sourceUrl":"https://clawhub.ai/edonadei/evaluate-skill","sourceType":"profile","confidence":"medium","observedAt":"2026-10-11T02:16:22.404Z","isPublic":true},{"factKey":"latest_release","category":"release","label":"Latest release","value":"1.0.14","href":"https://clawhub.ai/edonadei/evaluate-skill","sourceUrl":"https://clawhub.ai/edonadei/evaluate-skill","sourceType":"release","confidence":"medium","observedAt":"2026-09-25T15:12:56.018Z","isPublic":true},{"factKey":"handshake_status","category":"security","label":"Handshake status","value":"UNKNOWN","href":"https://www.xpersona.co/api/v1/agents/clawhub-edonadei-evaluate-skill/trust","sourceUrl":"https://www.xpersona.co/api/v1/agents/clawhub-edonadei-evaluate-skill/trust","sourceType":"trust","confidence":"medium","observedAt":null,"isPublic":true}],"changeEvents":[{"eventType":"release","title":"Release 1.0.14","description":"### Changed - Clean-up from the 1.0.13 that added way too much files. - Built around four questions, each with its own place to fix: does the skill fire (`description`), does it work (body), does it earn its place (the tasks), does it hold across edits. - New workflow order: run the control (`--ablate <skill>`) once, before editing the skill, and keep it. After each edit, compare against the previous full run as well as the control, because a worse skill can still beat the control. - New diagnosis table traces each failure to its fix: setup errors, unusable attempts, `cheat`, a skill that didn't fire, an extra skill that fired, a skill that fired but failed, and tasks that pass without the skill. - Specs it writes put `activates:` on every execution task (plus any skill the task delegates to), never name the skill in a prompt, and include at least one trigger probe. - Now owns writing a spec whose tasks are already decided, and advising on runs. The interview moved entirely to grill-skill. - Covers user customizations: runs load your own skills and setup by default. Isolate when comparing backends, sharing a number, or measuring the bare agent. - Reference rewritten and cut to what an agent acts on; flags are left to `caliper --help`. ### Fixed - The spec example used `skill: path:`, which `caliper validate` rejects. - Example evals used fixed `/tmp` paths that collide across parallel attempts; they now build fixtures in the attempt workdir. - The commit-simple example expected commits a single-shot run can't reach. ### Measured With the skill, Opus 5.5 went 9/9 on authoring tasks (writing `activates:`, proposing trigger probes, running the control first), against 0/9 without it, at k=3.","href":"https://clawhub.ai/edonadei/evaluate-skill","sourceUrl":"https://clawhub.ai/edonadei/evaluate-skill","sourceType":"release","confidence":"medium","observedAt":"2026-09-25T15:12:56.018Z","isPublic":true}]}