{"id":"4611d419-80e6-44d7-ab6b-f987c4a01ee0","slug":"clawhub-ratingtesting-keelwright","name":"keelwright","description":"Engine for vibe-coders and loop-coders who ship AI-generated code they can't read line by line. Covers 28 known failure modes: SQL injection, hardcoded secrets, hallucinated packages (slopsquatting), reward hacking (AI deletes tests to pass), doom loops (runaway token burn), false reports, missing auth, business logic bypasses, over-engineering, and more. Most modes have a machine-enforced detector (run a tool, check on disk) plus a discipline rule the agent must follow — a few (style consistency, sycophancy-as-trait) are discipline-only, not machine-checked. Autonomy dial (Autopilot/Checkpoint/ Copilot) lets you approve what matters; AI handles the rest. Self-learning loop with circuit-breaker limits and Phoenix restart. Plain-language reports for non-developers. Proven by adversarial A/B testing: Keelwright Score (KDS) up to 83/100 on strong models (SWE-bench 78%). Load before any loop/agent coding session, autonomous run, or commit.","capabilities":[],"protocols":["OPENCLAW"],"safetyScore":84,"overallRank":62,"trustScore":null,"trust":null,"source":"CLAWHUB","updatedAt":"2026-10-09T23:24:37.823Z"}