agentCLAWHUBUnverified

keelwright

Engine for vibe-coders and loop-coders who ship AI-generated code they can't read line by line. Covers 28 known failure modes: SQL injection, hardcoded secrets, hallucinated packages (slopsquatting), reward hacking (AI deletes tests to pass), doom loops (runaway token burn), false reports, missing auth, business logic bypasses, over-engineering, and more. Most modes have a machine-enforced detector (run a tool, check on disk) plus a discipline rule the agent must follow — a few (style consistency, sycophancy-as-trait) are discipline-only, not machine-checked. Autonomy dial (Autopilot/Checkpoint/ Copilot) lets you approve what matters; AI handles the rest. Self-learning loop with circuit-breaker limits and Phoenix restart. Plain-language reports for non-developers. Proven by adversarial A/B testing: Keelwright Score (KDS) up to 83/100 on strong models (SWE-bench 78%). Load before any loop/agent coding session, autonomous run, or commit. Skill: keelwright Owner: ratingtesting Summary: Engine for vibe-coders and loop-coders who ship AI-generated code they can't read line by line. Covers 28 known failure modes: SQL injection, hardcoded secrets, hallucinated packages (slopsquatting), reward hacking (AI deletes tests to pass), doom loops (runaway token burn), false reports, missing auth, business logic bypasses, over-engineering, and more. Most modes hav

OpenClaw

Rank

62

Safety

84

Downloads

1.9k

Updated

Oct 9, 2026

Version

1.11.0

Source

CLAWHUB

About

What it does, and when to use it.

Capability contract not published. No trust telemetry is available yet. 1.9K downloads reported by the source. Last updated 10/9/2026.

Avoid when

  • Contract metadata is missing or unavailable for deterministic execution.

Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing

Public facts

Every fact links back to the source it came from.

Vendor
Clawhubvendor · observed Oct 9, 2026
Protocol compatibility
OpenClawcompatibility · observed Oct 9, 2026
Adoption signal
1.9K downloadsadoption · observed Oct 9, 2026
Latest release
1.11.0release · observed Sep 1, 2026
Handshake status
UNKNOWNsecurity

Install and run

Setup complexity: low.

clawhub skill install s173w241ctx00fqb5chj1845bd8a21ec:keelwright
  1. Setup complexity is classified as HIGH. You must provision dedicated cloud infrastructure or an isolated VM. Do not run this directly on your local workstation.
  2. Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.

Contract: missing

curl -s "https://www.xpersona.co/api/v1/agents/clawhub-ratingtesting-keelwright/snapshot"

Documentation

CLAWHUB

160,000 characters of source documentation, loaded on request.

Extracted files

5 files captured from the source.

SKILL.md

---
name: keelwright
slug: keelwright
description: >-
  Engine for vibe-coders and loop-coders who ship AI-generated code they can't read line
  by line. Covers 28 known failure modes: SQL injection, hardcoded secrets, hallucinated
  packages (slopsquatting), reward hacking (AI deletes tests to pass), doom loops (runaway
  token burn), false reports, missing auth, business logic bypasses, over-engineering, and
  more. Most modes have a machine-enforced detector (run a tool, check on disk) plus a
  discipline rule the agent must follow — a few (style consistency, sycophancy-as-trait)
  are discipline-only, not machine-checked. Autonomy dial (Autopilot/Checkpoint/
  Copilot) lets you approve what matters; AI handles the rest. Self-learning loop with
  circuit-breaker limits and Phoenix restart. Plain-language reports for non-developers.
  Proven by adversarial A/B testing: Keelwright Score (KDS) up to 83/100 on strong models
  (SWE-bench 78%). Load before any loop/agent coding session, autonomous run, or commit.
version: 1.11.0
license: MIT-0
author: ratingtesting (https://github.com/ratingtesting)
platforms: [windows, linux, macos]
triggers:
  - vibe-code session starting
  - loop-code / autonomous agent run
  - unattended swarm / overnight job
  - commit touching auth/payments/data
  - agent asks "should I run this?"
metadata:
  runtime-agnostic: true
  self-contained: true
permissions:
  filesystem:
    - read
    - write
  shell:
    - run_scripts
    - run_tests
  network:
    - web_lookup
    - github_release_check
  install:
    - require_explicit_opt_in
---

# keelwright — an engine for vibe/loop coding

**One skill that combines four things a non-programmer needs to ship AI-generated code
safely and autonomously:** an autonomous loop, machine-enforced safety gates, an autonomy
dial, and self-learning. **Thin index** — heavy content lives in `references/*.md`,
load on demand. Saves ~14K tokens per session start vs a monolithic SKILL.md.

## ⚠️ Safety & consent (read first)

Keelwright is an **operational** skill. When loaded by an agent it can:

- Read and write files in your project (including `git add` / `git commit` during work).
- Invoke shell commands, run scripts, and execute local Python (verification recipes).
- Perform network checks (self-update, web guard) and, if you enable it, install optional tooling.

Loading the skill alone is **read-only context** until you answer the bootstrap question
or give explicit instruction. Every gate produces on-disk evidence, not a self-report.

---

## 🛡️ Critical rules (must hold even without reading references)

**These are duplicated here so they survive any context trim. Do not skip.**

- **R1 OWASP / R2 secrets / R3 business logic** = blockers EVEN in Autopilot. Never proceed past them without explicit human OK.
- **R4 80% problem (tech debt)**: agent delivers 80% of feature, silently skips critical 20% (tests, error handli

examples/README.md

# Examples — toy apps to try keelwright on

Three minimal projects to see keelwright's gates fire. Each is a deliberately small
loop-coding target; run keelwright alongside your agent and watch the gates.

## 1. `toy-flask-api/` — a 1-file web API
- **Task:** "build a /login endpoint that checks a hardcoded user".
- **What keelwright catches:** R2 (hardcoded password), R1 (SQL string concat if you use a DB).
- **Try:** `cd toy-flask-api && python app.py` then `curl localhost:5000/login`.

## 2. `toy-cli/` — a command-line tool
- **Task:** "a CLI that renames files by a pattern".
- **What keelwright catches:** R8 slopsquatting if the agent suggests a fake package;
  R3 business-logic review if the rename is destructive.
- **Try:** `cd toy-cli && python main.py --help`.

## 3. `toy-loop/` — an autonomous loop
- **Task:** "loop: fetch a number, double it, write to file, repeat 10x".
- **What keelwright catches:** circuit-breaker (doom-loop guard), R12 preflight.
- **Try:** `cd toy-loop && python loop.py` — watch breaker.py cap iterations.

## 30-second try (no install of keelwright internals needed)
1. Load the skill by name (`keelwright`) in your agent before coding.
2. Paste any toy task above into your agent.
3. Read the gate report at session end: `Keelwright this session: <N> gates passed,
   <M> traps avoided, <K> attacks blocked.`

No agent? Run the demo directly:
```bash
python scripts/validate_run.py --self-test   # exercises GATE 1-8 on a built-in sample
```

qa-results/README.md

# QA Results — Adversarial Test Runs

keelwright is battle-tested with adversarial A/B testing (control vs treatment, fact-checked on
disk, never self-report). This folder holds **machine-verified results** so every claim is backed
by artifacts, not marketing.

## Keelwright Score (KDS)

**KDS = ER × DR / 100** — one number (0–100) that tells you how well a model understands and
applies the skill's checks.

- **ER** (Execution Rate): can the model run an A/B test at all? `valid_tests / total_tests × 100`
- **DR** (Discrimination Rate): does the skill change the model's behavior? `DISCRIMINATES / valid_tests × 100`

| KDS | What it means |
|-----|---------------|
| **0** | Model can't run A/B tests (below threshold) |
| **1–10** | Weak / medium — skill adds some checks |
| **10–30** | Medium-strong — skill adds meaningful checks |
| **30–50** | Strong — skill adds security & quality gates |
| **50+** | Frontier — skill deeply understood and applied |

**KDS is not a general intelligence benchmark.** It measures "how much does keelwright improve
this model's outcomes" — a dimension no SWE-bench or GPQA captures.

## Scoreboard

| Model | Tier | SWE-Bench | Tests | DISC | DR | **KDS** |
|-------|------|-----------|-------|------|----|---------|
| poolside/laguna-s-2.1:free | STRONG | ML 78.5%, Pro 59.4% | 18 | 15 | 83% | **83** |
| stepfun/step-3.7-flash:free | MEDIUM | Pro ~56% | 6 | 4 | 67% | **67** |
| nvidia/nemotron-3-ultra-550b:free | STRONG | ML 67.7% | 5 | 2 | 40% | **40** |
| deepseek-v4-flash-free | STRONG | Verified ~79% | 14 | 4 | 29% | **29** |
| kimi-k3:free | STRONG | Terminal-Bench 88.3, ProgramBench 77.8 | 12 | 3 | 25% | **25** |
| inclusionai/ling-3.0-flash:free | UNKNOWN | SWE-bench/GPQA not published | 18 | 4 | 29% | **22** |
| mimo-v2.5-free | MEDIUM | Verified 78.9%, Pro 57.2% | 11 | 2 | 22% | **18** |
| claude-opus-4-8 | STRONG | frontier | 6 | 1 | 17% | **17** |
| claude-opus-5 | STRONG | Verified 96.0% | 15 | 2 | 18% | **13** |
| tencent/hy3:free | STRONG | ML 75.8%, Verified 78% | 43 | 3 | 7% | **7** |
| cohere/north-mini-code:free | WEAK | Agentic 3.1 | — | — | — | **0** |
| nvidia/nemotron-nano-9b-v2:free | WEAK | — | — | — | — | **0** |
| nvidia/nemotron-3-super-120b-a12b:free | STRONG | Verified 60.47% | 2* | 2* | 100%* | **PARTIAL** |

*\* `nvidia/nemotron-3-super-120b-a12b:free` — PARTIAL run: only sectors 1.1–1.2 completed
(2/18 tests) due to tool-call limit. Both showed DISCRIMINATES (code quality + task fidelity),
but KDS is not computed until ≥ a meaningful fraction of the battery runs. Re-run pending.

**Key findings:**
- **Laguna S 2.1** (KDS 83): strong model + skill adds 83% more checks. Best result recorded.
- **Step 3.7** (KDS 67): medium model gets MORE value from skill than some strong models.
  The skill compensates for gaps the model can't fill alone.
- **Weak models** (KDS 0): can't execute A/B tests — fabricate results instead. The skill
  can't help 

README.md

# keelwright

**Layered skill (index + on-demand references) for safe AI coding.**
Catches SQL injection, hardcoded secrets, hallucinated packages, reward hacking,
doom loops, and 23 other failure modes — with **machine-enforced gates** (not prompt
suggestions) and **plain-language reports** for non-developers.

[![security](https://github.com/ratingtesting/keelwright/actions/workflows/security.yml/badge.svg)](https://github.com/ratingtesting/keelwright/actions/workflows/security.yml)
[![license](https://img.shields.io/badge/license-MIT--0-blue.svg)](LICENSE)
[![kds](https://img.shields.io/badge/KDS-83%2F100-brightgreen.svg)](#keelwright-score-kds)

---

## What's new in v1.10.0

**Layered skill (ADR-001).** `SKILL.md` is now a thin **index** (~3K tokens, 84% smaller).
Heavy content lives in `references/*.md` and loads on demand. Public registries
(skills.sh / ClawHub / askill.sh) display the **assembled full document** built by
`scripts/build_skill.py`. Saves ~14K tokens per session start across Hermes, Cursor,
Codex, Cline, and OpenClaw.

See [`docs/ADR-001-layered-skill.md`](docs/ADR-001-layered-skill.md) for the decision
and `SKILL.md §Architecture` for runtime usage.

---

## The problem

You use AI to write code. You're not a developer — you're a founder, a builder, a
product person. The AI writes fast. You ship fast. And somewhere in that code:

- A password is hardcoded in plain text
- A database query is wide open to SQL injection
- A package name is one letter off from a real one — and it's malware
- The AI deleted a test to make the build go green
- A loop ran for 6 hours and burned $80 in tokens before you noticed
- The AI "fixed" a bug by removing the check that caught it

None of this shows up in a code review you can do. Because you can't read the code.

**keelwright fixes this.** It wraps your AI agent with machine-enforced checks that
catch these problems automatically — before they ship, before they cost you money,
before they become a security incident.

---

## What it does

![Architecture](assets/architecture.png)

**1. Machine-enforced security gates (R1–R12)**
28 known failure modes, checked automatically on every iteration. Every gate produces
on-disk evidence — not a self-report. Full implementation → `references/security-gates.md`.

**2. Autonomy dial**
Three modes you control: `Autopilot` (runs unattended, escalates on blockers),
`Checkpoint` (pauses at phase boundaries), `Copilot` (proposes, you approve every step).
Auth, payments, and production deploys always come to you.

**3. Circuit-breaker**
Stops runaway loops: 50 iterations max, 5 no-progress cap, 2-hour wall-clock, 3× same-error
repeat. Enforced by `scripts/breaker.py` (file-backed counters). Full philosophy →
`references/circuit-breaker.md`.

**4. Plain-language reporting**
Every gate outcome, every blocker, every decision point is explained in plain English —
what happened, why it matters to y

_meta.json

{
  "ownerId": "kn7ffn8e60z6nasp2f7gdbah0s8a2pxy",
  "slug": "keelwright",
  "version": "1.11.0",
  "publishedAt": 1788269771668
}
Github ReposUpdated 12h agoRank 70

AionUi

Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!

MCPOPENCLAW
Github ReposUpdated 6mo agoRank 70

activepieces

AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents

OPENCLAW
Github ReposUpdated 6mo agoRank 70

cherry-studio

AI productivity studio with smart chat, autonomous agents, and 300+ assistants.

MCPOPENCLAW
Github ReposUpdated 7mo agoRank 70

CopilotKit

The Frontend for Agents & Generative UI. React + Angular

OPENCLAW

Machine-readable data

The same record, as JSON, for agents and crawlers.

{
  "facts": [
    {
      "factKey": "vendor",
      "category": "vendor",
      "label": "Vendor",
      "value": "Clawhub",
      "href": "https://clawhub.ai/ratingtesting/skills/keelwright",
      "sourceUrl": "https://clawhub.ai/ratingtesting/skills/keelwright",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-09T23:24:37.823Z",
      "isPublic": true
    },
    {
      "factKey": "protocols",
      "category": "compatibility",
      "label": "Protocol compatibility",
      "value": "OpenClaw",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-ratingtesting-keelwright/contract",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-ratingtesting-keelwright/contract",
      "sourceType": "contract",
      "confidence": "medium",
      "observedAt": "2026-10-09T23:24:37.823Z",
      "isPublic": true
    },
    {
      "factKey": "traction",
      "category": "adoption",
      "label": "Adoption signal",
      "value": "1.9K downloads",
      "href": "https://clawhub.ai/ratingtesting/keelwright",
      "sourceUrl": "https://clawhub.ai/ratingtesting/keelwright",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-09T23:24:37.823Z",
      "isPublic": true
    },
    {
      "factKey": "latest_release",
      "category": "release",
      "label": "Latest release",
      "value": "1.11.0",
      "href": "https://clawhub.ai/ratingtesting/keelwright",
      "sourceUrl": "https://clawhub.ai/ratingtesting/keelwright",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-09-01T13:36:11.668Z",
      "isPublic": true
    },
    {
      "factKey": "handshake_status",
      "category": "security",
      "label": "Handshake status",
      "value": "UNKNOWN",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-ratingtesting-keelwright/trust",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-ratingtesting-keelwright/trust",
      "sourceType": "trust",
      "confidence": "medium",
      "observedAt": null,
      "isPublic": true
    }
  ],
  "events": [
    {
      "eventType": "release",
      "title": "Release 1.11.0",
      "description": "v1.11.0: SkillSpector response bundle (permissions, viral opt-in, verify_web_guard tighten, strategy move, QA gating)",
      "href": "https://clawhub.ai/ratingtesting/keelwright",
      "sourceUrl": "https://clawhub.ai/ratingtesting/keelwright",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-09-01T13:36:11.668Z",
      "isPublic": true
    }
  ]
}

Record generated Oct 10, 2026.

Sponsored

Ads related to keelwright and adjacent AI workflows.