OpenClaw Self-Improvement
A reusable operator-guided workflow improvement skill for OpenClaw and ClawLite that turns repeated failures into logged learnings, binary eval loops, SOPs,... Skill: OpenClaw Self-Improvement Owner: x-rayluan Summary: A reusable operator-guided workflow improvement skill for OpenClaw and ClawLite that turns repeated failures into logged learnings, binary eval loops, SOPs,... Tags: latest:0.2.11 Version history: v0.2.11 | 2026-05-01T08:13:57.502Z | user Add scorecard repair loop with recovery ticket generation. v0.2.10 | 2026-05-01T08:07:07.960Z | user Add harness observabi
Rank
62
Safety
84
Downloads
3.4k
Updated
Oct 9, 2026
Version
0.2.11
Source
CLAWHUB
About
What it does, and when to use it.
Capability contract not published. No trust telemetry is available yet. 3.4K downloads reported by the source. Last updated 10/9/2026.
Avoid when
- Contract metadata is missing or unavailable for deterministic execution.
Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing
Public facts
Every fact links back to the source it came from.
- Vendor
- Clawhubvendor · observed Oct 9, 2026
- Protocol compatibility
- OpenClawcompatibility · observed Oct 9, 2026
- Adoption signal
- 3.4K downloadsadoption · observed Oct 9, 2026
- Latest release
- 0.2.11release · observed May 1, 2026
- Handshake status
- UNKNOWNsecurity
Install and run
Setup complexity: low.
clawhub skill install s17c0n0ce78mre7mrytgzehcwn83hkp1:openclaw-self-improvement- Setup complexity is classified as HIGH. You must provision dedicated cloud infrastructure or an isolated VM. Do not run this directly on your local workstation.
- Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.
Contract: missing
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-x-rayluan-openclaw-self-improvement/snapshot"
Documentation
CLAWHUB
150,598 characters of source documentation, loaded on request.
Extracted files
5 files captured from the source.
SKILL.md
---
name: openclaw-self-improvement
description: A reusable operator-guided workflow improvement skill for OpenClaw and ClawLite that turns repeated failures into logged learnings, binary eval loops, SOPs, checklists, and proof-based operational improvements.
metadata:
{
"openclaw":
{
"requires": { "bins": ["node"] },
"writes": [".learnings/", "memory/harness-backlog-latest.md", "mission-control/data/delivery-receipts/agent-scorecard-YYYY-MM-DD.md", "AGENTS.md", "TOOLS.md", "SOUL.md"],
"env": ["WORKSPACE", "OBSIDIAN_LEARNINGS_DIR"],
"network": false,
"notes": "Local-file workflow only. Promotion writes should be reviewed and can be previewed with --dry-run."
}
}
---
# OpenClaw / ClawLite Self-Improvement
Use this skill to turn mistakes, corrections, blockers, and better approaches into durable operating knowledge.
## What problem this solves
AI ops often repeat the same failures because mistakes stay in chat history instead of becoming system rules. This skill creates a lightweight improvement loop:
- log failures and learnings
- separate errors from feature requests
- run small eval-driven experiments on repeated failures
- classify harness/runtime failures instead of blaming vague “model issues”
- generate daily agent scorecards from real evidence chains
- promote important patterns into AGENTS.md / TOOLS.md / SOUL.md
- write operator notes into Obsidian vault
- support stricter acceptance via Karen / Mission Control
## When to use
Use this skill when the user asks:
- "make the agent improve itself"
- "capture learnings"
- "log mistakes so we do not repeat them"
- "record blockers / corrections / feature gaps"
- "build a self-improving OpenClaw workflow"
- "operationalize lessons learned"
- "test whether this new rule actually helps"
- "run an eval loop on this workflow/skill/SOP"
- "should we keep this new guardrail or discard it"
- "why did the agents fail today"
- "why is daily marketing not closing automatically"
- "classify OpenClaw harness failures"
- "generate agent delivery scorecard"
## Files this skill uses
- `.learnings/LEARNINGS.md`
- `.learnings/ERRORS.md`
- `.learnings/FEATURE_REQUESTS.md`
- `.learnings/EXPERIMENTS.md`
- `memory/harness-backlog-latest.md`
- `mission-control/data/delivery-receipts/agent-scorecard-YYYY-MM-DD.md`
- Optional export under `.learnings/exports/obsidian/` by default, or `OBSIDIAN_LEARNINGS_DIR` if explicitly configured
## Safety boundaries
- Local-file workflow only, no network I/O
- Promotion can append to `AGENTS.md`, `TOOLS.md`, or `SOUL.md`
- Always review promotion targets first, or run `scripts/promote-learning.mjs ... --dry-run`
- `OBSIDIAN_LEARNINGS_DIR` should only point at a path you intend to modify
## Command examples
```bash
node {baseDir}/scripts/log-learning.mjs learning "Summary" "Details" "Suggested action"
node {baseDir}/scripts/log-learning.mjs error "Summary" "Error details" "Suggested fix"
node {baseDir}/scripts/log-leREADME.md
# OpenClaw Self-Improvement **OpenClaw Self-Improvement** is a reusable agent skill for turning repeated AI-agent mistakes into durable operational improvements, measurable guardrails, and inspectable workflow upgrades. ## Landing-page summary Most AI agents do not really improve. They repeat mistakes, hide partial failures behind optimistic language, and leave lessons trapped in chat history. OpenClaw Self-Improvement gives you a practical operating loop for **self-improving AI agents**: - capture repeated failures - test one guardrail at a time - verify whether it actually reduces failure - promote proven fixes into SOPs, checklists, policies, and workflow rules If you want **AI agents that get more reliable over time**, **multi-agent workflows that stop repeating the same mistakes**, or **proof-based QA for agent operations**, this skill is built for that exact use case. ## Multilingual summary **中文:** 这是一个面向 AI agent 自我改进的实战型 skill,用来减少重复犯错、建立 guardrails、验证修复是否真的有效,并把经验沉淀成 SOP、检查清单和可复用规则。 **日本語:** これは自己改善する AI エージェント向けの実践的な skill です。繰り返し発生する失敗を減らし、ガードレールを検証し、改善を SOP・チェックリスト・再利用可能な運用ルールへ昇格させます。 **한국어:** 이 스킬은 스스로 개선하는 AI 에이전트를 위한 운영형 skill입니다. 반복 실수를 줄이고, 가드레일이 실제로 효과가 있는지 검증하며, 개선 사항을 SOP·체크리스트·재사용 가능한 규칙으로 승격합니다. **Español:** Esta skill está diseñada para agentes de IA que deben mejorar con el tiempo. Ayuda a reducir errores repetidos, validar guardrails operativos y convertir mejoras en SOP, checklists y reglas reutilizables. If you want a practical way to deploy OpenClaw with cheaper tokens, BYOK flexibility, and operator control, see **[ClawLite](https://clawlite.ai)**. ## TL;DR If you are looking for a practical system for **self-improving AI agents**, **AI workflow optimization**, **multi-agent failure prevention**, **binary eval loops**, or **agent operations QA**, this skill is designed for that exact job. It helps OpenClaw / ClawLite operators and agent teams: - log recurring failures - separate one-off errors from reusable lessons - run lightweight **binary eval loops** on new guardrails - classify changes as **keep**, **partial_keep**, or **discard** - promote proven fixes into SOPs, checklists, workflow rules, and operating policy If you care about reducing fake-complete states, tightening QA truth, improving deploy closeout, and making agent learning inspectable, this skill is built for that job. --- ## Why this skill exists Many AI systems say they "learn," but most only store lessons in chat history or loose notes. That is not enough. Operationally, repeated failures tend to come back in the same forms: - delivery gets described as complete before proof exists - receipts are missing or too thin - back-end fixes never reach the operator-facing surface - code-ready states get confused with production-ready states - teams add new rules without checking whether those rules actually reduce failure OpenClaw Self-Improvement gives you a lightweight operating loop for fixing that. --- ## What problem it solves T
_meta.json
{
"ownerId": "kn7d88952ey3hbm158x7ejqs1d81zmyq",
"slug": "openclaw-self-improvement",
"version": "0.2.11",
"publishedAt": 1777623237502
}references/decision-rules.md
# Decision Rules for Self-Improvement Use this reference to decide whether a new issue should become: - a simple learning entry - an experiment with binary evals - a promoted operating rule --- ## Option 1 — Log only Use **log only** when: - the issue happened once - root cause is still unclear - there is not enough evidence yet to turn it into a rule - the lesson is useful, but not broadly reusable yet Typical output: - `learning` - `error` - `feature` Examples: - one-off API outage - first-time tool glitch with unclear cause - user preference that does not affect system-wide ops --- ## Option 2 — Run an experiment Use an **experiment** when: - the same failure happened 2+ times - a new guardrail/SOP/checklist/schema change is being proposed - you can define 3-5 binary evals - you want evidence that the change helped before promoting it broadly Typical output: - `experiment` - baseline + mutation + binary evals + keep/discard decision Examples: - Mission Control summaries repeatedly missing links/details - deploy closeout repeatedly confusing code-ready with live - repeated missing receipts or incomplete proof bundles - repeated front-end / operator-surface mismatch after backend fixes --- ## Option 3 — Promote immediately Use **promote immediately** when: - the rule is already obviously correct and low-risk - the issue is severe enough that waiting would be irresponsible - the required change is a principle or ownership rule, not an uncertain optimization - operator review already confirms the new rule should become standard Typical output: - promotion into `AGENTS.md`, `TOOLS.md`, `SOUL.md`, or `docs/ops/*.md` Examples: - deployment owner must be explicit - code-ready is not the same as live - missing receipt cannot be treated as delivered - summary without proof links is not operator-complete --- ## Promote after experiment Use **experiment first, then promote** when: - the rule sounds plausible but may add friction - you are not sure whether the added checklist/schema field actually reduces errors - the change could create process overhead without improving truth quality Examples: - adding new summary schema fields - adding new receipt requirements - adding extra verification steps to handoff or QA lanes --- ## Anti-patterns Do **not** run an experiment when: - there is no clear repeated failure - the evals would be vague or subjective - the issue is really just missing execution, not missing learning - the fix requires immediate owner action, not more analysis Do **not** promote immediately when: - the rule is still based on one anecdote - the change is likely to create bureaucracy without proof of benefit - the actual failure surface is still unclear --- ## Quick decision tree 1. Did this happen only once? - Yes → log only - No → continue 2. Is the new rule obviously necessary and low-risk? - Yes → promote immediately - No → continue 3. Can you define 3-5 binary evals for the proposed change? - Yes → run an exp
references/eval-loop.md
# Eval Loop for Self-Improvement Use this reference when a repeated failure should become a tested operational improvement instead of only a logged lesson. ## Goal Do not only ask "what did we learn?" Also ask: - what is the current baseline? - what exact guardrail or rule changed? - how will we measure whether it helped? - should we keep or discard the change? ## Use this loop for - repeated Mission Control wording failures - missing receipts / missing proof chains - deploy closeout failures - stale operator-facing surfaces - repeated handoff mistakes between agents - recurring SOP/checklist changes ## 1. Define the target State one concrete thing you want to improve. Examples: - Hunter summary should always include concrete links and details - ClawLite deploy closeout should never stop at code-ready status - Mission Control front-end should render source links from structured fields ## 2. Write 3-5 binary evals Each eval must be yes/no. Examples for summary quality: - Does the summary include at least one artifact path or URL? - Does the summary include evidence links when external proof matters? - Does the summary include a detail block describing what actually changed? - Does the summary include the next handoff or recovery action? - Does the operator-facing surface actually render these fields? Examples for deploy closeout: - Is the deployed commit hash recorded? - Is a deployment ref/URL recorded? - Was the production page or sitemap actually verified? - Was a structured receipt written? - Is the final state classified with the correct deploy-state vocabulary? ## 3. Capture baseline Before changing the rule/SOP/skill/checklist: - record the current failure pattern - record which evals currently fail - treat this as the baseline state ## 4. Change only one thing Good changes: - one wording rule - one new checklist item - one schema field - one render mapping - one validation step Bad changes: - rewriting everything at once - adding five new rules at once - changing wording and schema and code together unless absolutely required ## 5. Re-check and classify After the single change: - run the same evals again - note which checks improved - decide: - KEEP - DISCARD - PARTIAL_KEEP ## 6. Promotion rule Only promote broadly reusable changes after they pass the eval loop or after operator review confirms the change materially reduced the failure. ## Suggested experiment entry format ```md ## [EXP-YYYYMMDD-XXX] experiment **Logged**: ISO-8601 timestamp **Priority**: medium | high | critical **Status**: baseline | testing | keep | discard | partial_keep **Area**: workflow | tools | product | growth | security | infra | ops ### Target What repeated problem is being improved ### Baseline What was failing before the change ### Mutation The single change introduced ### Binary Evals - [ ] Eval 1 - [ ] Eval 2 - [ ] Eval 3 ### Result What improved / did not improve ### Keep or Discard keep | discard | partial_keep ### Meta
activepieces
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
AionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
CopilotKit
The Frontend for Agents & Generative UI. React + Angular
Machine-readable data
The same record, as JSON, for agents and crawlers.
{
"facts": [
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Clawhub",
"href": "https://clawhub.ai/x-rayluan/skills/openclaw-self-improvement",
"sourceUrl": "https://clawhub.ai/x-rayluan/skills/openclaw-self-improvement",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-09T08:11:38.169Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-x-rayluan-openclaw-self-improvement/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-x-rayluan-openclaw-self-improvement/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-09T08:11:38.169Z",
"isPublic": true
},
{
"factKey": "traction",
"category": "adoption",
"label": "Adoption signal",
"value": "3.4K downloads",
"href": "https://clawhub.ai/x-rayluan/openclaw-self-improvement",
"sourceUrl": "https://clawhub.ai/x-rayluan/openclaw-self-improvement",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-09T08:11:38.169Z",
"isPublic": true
},
{
"factKey": "latest_release",
"category": "release",
"label": "Latest release",
"value": "0.2.11",
"href": "https://clawhub.ai/x-rayluan/openclaw-self-improvement",
"sourceUrl": "https://clawhub.ai/x-rayluan/openclaw-self-improvement",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-05-01T08:13:57.502Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-x-rayluan-openclaw-self-improvement/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-x-rayluan-openclaw-self-improvement/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
],
"events": [
{
"eventType": "release",
"title": "Release 0.2.11",
"description": "Add scorecard repair loop with recovery ticket generation.",
"href": "https://clawhub.ai/x-rayluan/openclaw-self-improvement",
"sourceUrl": "https://clawhub.ai/x-rayluan/openclaw-self-improvement",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-05-01T08:13:57.502Z",
"isPublic": true
}
]
}Record generated Oct 9, 2026.
