{"id":"4611d419-80e6-44d7-ab6b-f987c4a01ee0","slug":"clawhub-ratingtesting-keelwright","name":"keelwright","description":"Engine for vibe-coders and loop-coders who ship AI-generated code they can't read line by line. Covers 28 known failure modes: SQL injection, hardcoded secrets, hallucinated packages (slopsquatting), reward hacking (AI deletes tests to pass), doom loops (runaway token burn), false reports, missing auth, business logic bypasses, over-engineering, and more. Most modes have a machine-enforced detector (run a tool, check on disk) plus a discipline rule the agent must follow — a few (style consistency, sycophancy-as-trait) are discipline-only, not machine-checked. Autonomy dial (Autopilot/Checkpoint/ Copilot) lets you approve what matters; AI handles the rest. Self-learning loop with circuit-breaker limits and Phoenix restart. Plain-language reports for non-developers. Proven by adversarial A/B testing: Keelwright Score (KDS) up to 83/100 on strong models (SWE-bench 78%). Load before any loop/agent coding session, autonomous run, or commit.","canonicalUrl":"https://www.xpersona.co/agent/clawhub-ratingtesting-keelwright","sourceUrl":"https://clawhub.ai/ratingtesting/keelwright","homepage":"https://clawhub.ai/ratingtesting/skills/keelwright","source":"CLAWHUB","vendor":{"slug":"clawhub","label":"Clawhub","url":"https://clawhub.ai/ratingtesting/skills/keelwright"},"protocols":["OPENCLEW"],"capabilities":[],"trustScore":null,"trustConfidence":"unknown","artifactCount":0,"benchmarkCount":0,"lastRelease":"1.11.0","freshnessAt":"2026-10-09T23:24:37.811Z","freshnessLabel":"Oct 9, 2026","securityReviewed":true,"openapiReady":false,"stats":[{"label":"Trust score","value":"Unknown"},{"label":"Compatibility","value":"OpenClaw"},{"label":"Freshness","value":"Oct 9, 2026"},{"label":"Vendor","value":"Clawhub"},{"label":"Artifacts","value":"0"},{"label":"Benchmarks","value":"0"},{"label":"Last release","value":"1.11.0"}],"factsPreview":[{"factKey":"vendor","category":"vendor","label":"Vendor","value":"Clawhub","href":"https://clawhub.ai/ratingtesting/skills/keelwright","sourceUrl":"https://clawhub.ai/ratingtesting/skills/keelwright","sourceType":"profile","confidence":"medium","observedAt":"2026-10-09T23:24:37.823Z","isPublic":true},{"factKey":"protocols","category":"compatibility","label":"Protocol compatibility","value":"OpenClaw","href":"https://www.xpersona.co/api/v1/agents/clawhub-ratingtesting-keelwright/contract","sourceUrl":"https://www.xpersona.co/api/v1/agents/clawhub-ratingtesting-keelwright/contract","sourceType":"contract","confidence":"medium","observedAt":"2026-10-09T23:24:37.823Z","isPublic":true},{"factKey":"traction","category":"adoption","label":"Adoption signal","value":"1.9K downloads","href":"https://clawhub.ai/ratingtesting/keelwright","sourceUrl":"https://clawhub.ai/ratingtesting/keelwright","sourceType":"profile","confidence":"medium","observedAt":"2026-10-09T23:24:37.823Z","isPublic":true},{"factKey":"latest_release","category":"release","label":"Latest release","value":"1.11.0","href":"https://clawhub.ai/ratingtesting/keelwright","sourceUrl":"https://clawhub.ai/ratingtesting/keelwright","sourceType":"release","confidence":"medium","observedAt":"2026-09-01T13:36:11.668Z","isPublic":true},{"factKey":"handshake_status","category":"security","label":"Handshake status","value":"UNKNOWN","href":"https://www.xpersona.co/api/v1/agents/clawhub-ratingtesting-keelwright/trust","sourceUrl":"https://www.xpersona.co/api/v1/agents/clawhub-ratingtesting-keelwright/trust","sourceType":"trust","confidence":"medium","observedAt":null,"isPublic":true}],"highlights":["1.9K downloads","Trust evidence available"],"agentCard":{"name":"keelwright","description":"Engine for vibe-coders and loop-coders who ship AI-generated code they can't read line by line. Covers 28 known failure modes: SQL injection, hardcoded secrets, hallucinated packages (slopsquatting), reward hacking (AI deletes tests to pass), doom loops (runaway token burn), false reports, missing auth, business logic bypasses, over-engineering, and more. Most modes have a machine-enforced detector (run a tool, check on disk) plus a discipline rule the agent must follow — a few (style consistency, sycophancy-as-trait) are discipline-only, not machine-checked. Autonomy dial (Autopilot/Checkpoint/ Copilot) lets you approve what matters; AI handles the rest. Self-learning loop with circuit-breaker limits and Phoenix restart. Plain-language reports for non-developers. Proven by adversarial A/B testing: Keelwright Score (KDS) up to 83/100 on strong models (SWE-bench 78%). Load before any loop/agent coding session, autonomous run, or commit.","source":"CLAWHUB","sourceId":"clawhub:s173w241ctx00fqb5chj1845bd8a21ec:keelwright","homepage":"https://clawhub.ai/ratingtesting/skills/keelwright","repository":"https://clawhub.ai/ratingtesting/keelwright","documentation":"https://www.xpersona.co/agent/clawhub-ratingtesting-keelwright","protocols":["OPENCLEW"],"examples":[{"kind":"example","language":"text","snippet":"File v1.11.0:qa-results/README.md\n\n# QA Results — Adversarial Test Runs\r\n\r\nkeelwright is battle-tested with adversarial A/B testing (control vs treatment, fact-checked on\r\ndisk, never self-report). This folder holds **machine-verified results** so every claim is backed\r\nby artifacts, not marketing.\r\n\r\n## Keelwright Score (KDS)\r\n\r\n**KDS = ER × DR / 100** — one number (0–100) that tells you how well a model understands and\r\napplies the skill's checks.\r\n\r\n- **ER** (Execution Rate): can the model run an A/B test at all? `valid_tests / total_tests × 100`\r\n- **DR** (Discrimination Rate): does the skill change the model's behavior? `DISCRIMINATES / valid_tests × 100`\r\n\r\n| KDS | What it means |\r\n|-----|---------------|\r\n| **0** | Model can't run A/B tests (below threshold) |\r\n| **1–10** | Weak / medium — skill adds some checks |\r\n| **10–30** | Medium-strong — skill adds meaningful checks |\r\n| **30–50** | Strong — skill adds security & quality gates |\r\n| **50+** | Frontier — skill deeply understood and applied |\r\n\r\n**KDS is not a general intelligence benchmark.** It measures \"how much does keelwright improve\r\nthis model's outcomes\" — a dimension no SWE-bench or GPQA captures.\r\n\r\n## Scoreboard\r\n\r\n| Model | Tier | SWE-Bench | Tests | DISC | DR | **KDS** |\r\n|-------|------|-----------|-------|------|----|---------|\r\n| poolside/laguna-s-2.1:free | STRONG | ML 78.5%, Pro 59.4% | 18 | 15 | 83% | **83** |\r\n| stepfun/step-3.7-flash:free | MEDIUM | Pro ~56% | 6 | 4 | 67% | **67** |\r\n| nvidia/nemotron-3-ultra-550b:free | STRONG | ML 67.7% | 5 | 2 | 40% | **40** |\r\n| deepseek-v4-flash-free | STRONG | Verified ~79% | 14 | 4 | 29% | **29** |\r\n| kimi-k3:free | STRONG | Terminal-Bench 88.3, ProgramBench 77.8 | 12 | 3 | 25% | **25** |\r\n| inclusionai/ling-3.0-flash:free | UNKNOWN | SWE-bench/GPQA not published | 18 | 4 | 29% | **22** |\r\n| mimo-v2.5-free | MEDIUM | Verified 78.9%, Pro 57.2% | 11 | 2 | 22% | **18** |\r\n| claude-opus-4-8 | STRONG | frontier | 6 | 1 | 17% | **17** |\r\n| claude-opu"},{"kind":"example","language":"text","snippet":"File v1.11.0:README.md\n\n# keelwright\r\n\r\n**Layered skill (index + on-demand references) for safe AI coding.**\r\nCatches SQL injection, hardcoded secrets, hallucinated packages, reward hacking,\r\ndoom loops, and 23 other failure modes — with **machine-enforced gates** (not prompt\r\nsuggestions) and **plain-language reports** for non-developers.\r\n\r\n[![security](https://github.com/ratingtesting/keelwright/actions/workflows/security.yml/badge.svg)](https://github.com/ratingtesting/keelwright/actions/workflows/security.yml)\r\n[![license](https://img.shields.io/badge/license-MIT--0-blue.svg)](LICENSE)\r\n[![kds](https://img.shields.io/badge/KDS-83%2F100-brightgreen.svg)](#keelwright-score-kds)\r\n\r\n---\r\n\r\n## What's new in v1.10.0\r\n\r\n**Layered skill (ADR-001).** `SKILL.md` is now a thin **index** (~3K tokens, 84% smaller).\r\nHeavy content lives in `references/*.md` and loads on demand. Public registries\r\n(skills.sh / ClawHub / askill.sh) display the **assembled full document** built by\r\n`scripts/build_skill.py`. Saves ~14K tokens per session start across Hermes, Cursor,\r\nCodex, Cline, and OpenClaw.\r\n\r\nSee [`docs/ADR-001-layered-skill.md`](docs/ADR-001-layered-skill.md) for the decision\r\nand `SKILL.md §Architecture` for runtime usage.\r\n\r\n---\r\n\r\n## The problem\r\n\r\nYou use AI to write code. You're not a developer — you're a founder, a builder, a\r\nproduct person. The AI writes fast. You ship fast. And somewhere in that code:\r\n\r\n- A password is hardcoded in plain text\r\n- A database query is wide open to SQL injection\r\n- A package name is one letter off from a real one — and it's malware\r\n- The AI deleted a test to make the build go green\r\n- A loop ran for 6 hours and burned $80 in tokens before you noticed\r\n- The AI \"fixed\" a bug by removing the check that caught it\r\n\r\nNone of this shows up in a code review you can do. Because you can't read the code.\r\n\r\n**keelwright fixes this.** It wraps your AI agent with machine-enforced checks that\r\ncatch these problems automatically — before they sh"}]}}