agentCLAWHUBUnverified

Goodhart's Law

Activate when: our KPI is going up but the real outcome isn't improving; people seem to be gaming the metric; we're about to tie bonuses or promotions to a n... Skill: Goodhart's Law Owner: deciqai Summary: Activate when: our KPI is going up but the real outcome isn't improving; people seem to be gaming the metric; we're about to tie bonuses or promotions to a n... Tags: latest:1.0.5 Version history: v1.0.5 | 2026-07-16T18:01:32.110Z | user Description tail link + agents machine-readable metadata line (deciqai.com/s/goodharts-law.json) v1.0.4 | 2026-07-09T11:17:54.034Z | use

OpenClaw

Rank

62

Safety

84

Downloads

1.2k

Updated

Oct 11, 2026

Version

1.0.5

Source

CLAWHUB

About

What it does, and when to use it.

Capability contract not published. No trust telemetry is available yet. 1.2K downloads reported by the source. Last updated 10/11/2026.

Avoid when

  • Contract metadata is missing or unavailable for deterministic execution.

Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing

Public facts

Every fact links back to the source it came from.

Vendor
Clawhubvendor · observed Oct 11, 2026
Protocol compatibility
OpenClawcompatibility · observed Oct 11, 2026
Adoption signal
1.2K downloadsadoption · observed Oct 11, 2026
Latest release
1.0.5release · observed Jul 16, 2026
Handshake status
UNKNOWNsecurity

Install and run

Setup complexity: low.

clawhub skill install s17a4mqcnk515kvaca5ze55d0x88pfpx:goodharts-law
  1. Setup complexity is classified as HIGH. You must provision dedicated cloud infrastructure or an isolated VM. Do not run this directly on your local workstation.
  2. Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.

Contract: missing

curl -s "https://www.xpersona.co/api/v1/agents/clawhub-deciqai-goodharts-law/snapshot"

Documentation

CLAWHUB

143,441 characters of source documentation, loaded on request.

Extracted files

5 files captured from the source.

SKILL.md

---
name: goodharts-law
description: "Activate when: our KPI is going up but the real outcome isn't improving; people seem to be gaming the metric; we're about to tie bonuses or promotions to a number; an algorithm is producing results nobody intended; a test or audit system is being designed.
  Do NOT activate when: the metric IS the goal with no proxy gap; measurement is purely descriptive with zero stakes attached. More: deciqai.com/c/goodharts-law"
---

# Goodhart's Law

## Overview

**Goodhart's Law:** when a metric controls behavior, people optimize the metric rather than the underlying goal. Formulated by economist Charles Goodhart (1975) on UK monetary policy; sharpened by Marilyn Strathern (1997): *"When a measure becomes a target, it ceases to be a good measure."* Four failure mechanisms (Manheim & Garrabrant 2018): **Regressional**, **Extremal**, **Causal**, **Adversarial**. Countermeasure is always multi-metric + audit + rotation.

Composes with `feedback-loops`, `principal-agent`, `okr-goal-setting`, `survivorship-bias`.

## When to Use

- A KPI is being introduced or its weight is increasing in performance evaluation
- A metric is "improving" without corresponding improvement in the underlying goal
- People are visibly optimizing for a number rather than the work it was meant to track
- Algorithmic optimization is producing outcomes the designers didn't intend
- Resource allocation is driven by a single composite score or ranking
- An AI model, benchmark, or engagement metric is being optimized (or used to justify AI capex / adoption / AI-native competition) and the score is rising faster than real capability or user value

**Not when:** metric and goal are identical; stakes too low for gaming; metric is purely descriptive with no reward/punishment; question is which metric to use, not whether the measurement-reward system is sound.

## Coaching Novices (Adaptive Front Door)

- **Engine mode:** user has a concrete metric or system → run The Process directly.
- **Coach mode:** user is unfamiliar or has no concrete case → guide step by step.

In Coach mode, respond one step at a time. Each [WAIT] is a hard stop — output only that step's question, then stop.

1. One-line: before relying on a metric to control behavior, predict how people will game it — choose the system that survives that prediction.
2. Check fit: if the metric is purely descriptive (no reward attached), Goodhart's law doesn't apply yet.
3. Elicit the specific metric and the underlying goal: what's being measured? What's the actual outcome you care about?
> **[WAIT — do not advance until user responds]**
4. One question at a time: proxy gap? How would a clever agent game this? Which Goodhart category? What countermeasure fits?
> **[WAIT — do not advance until user responds]**
5. Close: name the gaming-resistant design (multi-metric, audit, rotation, paired-constraint) + monitoring schedule.
> **[WAIT — do not advance until user responds]**

## The Process

**Step 1 — S

_meta.json

{
  "ownerId": "kn754b8sk22s8c6gjxt02bftbn88q7ye",
  "slug": "goodharts-law",
  "version": "1.0.5",
  "publishedAt": 1784224892110
}

references/sources.md

# Sources — goodharts-law

> *Primary sources for the [goodharts-law](../SKILL.md) skill.*

- Goodhart, C. A. E. (1975). "Problems of monetary management: The U.K. experience." Papers in Monetary Economics, Reserve Bank of Australia. Reprinted in *Monetary Theory and Practice: The U.K. Experience* (1984), Macmillan. The original.
- Strathern, M. (1997). "'Improving ratings': Audit in the British university system." *European Review*, 5(3), 305-321. The modern aphoristic formulation.
- Campbell, D. T. (1979). "Assessing the impact of planned social change." *Evaluation and Program Planning*, 2(1), 67-90. Independent formulation ("Campbell's Law").
- Lucas, R. E. (1976). "Econometric policy evaluation: A critique." *Carnegie-Rochester Conference Series on Public Policy*, 1(1), 19-46. The adjacent macro-econometric "Lucas critique."
- Manheim, D., & Garrabrant, S. (2018). "Categorizing variants of Goodhart's Law." *arXiv:1803.04585*. The four-mechanism taxonomy.
- Muller, J. Z. (2018). *The Tyranny of Metrics.* Princeton University Press. ISBN 978-0691174952. Comprehensive case-study survey.
- Doerr, J. (2018). *Measure What Matters: How Google, Bono, and the Gates Foundation Rock the World with OKRs.* Portfolio. ISBN 978-0525536222.
- Wells Fargo / Office of the Comptroller of the Currency (2016, 2018). Various enforcement actions and consent orders documenting the cross-selling case.
- Zhou, K., et al. (2023). "Don't Make Your LLM an Evaluation Benchmark Cheater." *arXiv:2311.01964*. Documents benchmark data contamination and how public evaluation scores decay as a capability signal once test data leaks into training — a direct 2020s AI instance of Goodhart's law.
- Chiang, W.-L., et al. (2024). "Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference." *arXiv:2403.04132*. Motivates live, blind, human-preference evaluation as a harder-to-game complement to static leaderboards; illustrates both the countermeasure and its own residual gaming risks.

examples/ai-benchmark-and-engagement-gaming-2023-2026.md

# Method in Action: AI Benchmarks and Engagement Metrics as Targets (2023–2026)

> *Example for the [goodharts-law](../SKILL.md) skill.*

By the mid-2020s, Goodhart's law had become one of the most-cited frames inside the AI industry itself — because two of its own core metrics visibly decayed under optimization pressure. First, **public benchmark scores** (MMLU, GSM8K, HumanEval, and a proliferation of leaderboards) came to dominate model marketing, funding narratives, and internal go/no-go decisions — and, predictably, models began scoring well without a matching gain in real-world capability. Second, **consumer-app engagement metrics** (watch time, session length, daily active use) continued their long slide from "signal of user value" to "target that no longer measures it." This walks both cases through the skill's own six-step Process.

---

## Step 1 — State metric and goal

**Case A — AI benchmarks.**
- **Metric being targeted:** score on a fixed public benchmark (e.g., a multiple-choice knowledge test like MMLU, a grade-school math set like GSM8K, or a coding pass-rate like HumanEval).
- **Underlying goal:** general, transferable model capability — does the model actually reason, code, and generalize on tasks users bring that were *not* in the test set?
- **Current proxy–goal correlation:** initially high on a genuinely held-out test, and a legitimate research signal. It degrades as the benchmark ages, becomes a marketing target, and (critically) as the test's questions leak into training data.
- **Who is measured:** frontier and open-weight model developers; the score is read by press, investors, enterprise buyers, and internal leadership deciding what to ship.
- **Stakes:** very high — leaderboard position drives valuations, capex-justification narratives, and enterprise procurement in an intensely competitive, AI-native market.

**Case B — engagement metrics.**
- **Metric being targeted:** engagement (watch time, session length, DAU/MAU, scroll depth) feeding a recommendation or ranking system.
- **Underlying goal:** users getting durable value — time well spent, learning, connection, satisfaction they'd endorse on reflection.
- **Correlation:** engagement is a real proxy for value at low intensity, but the two diverge sharply as the system optimizes hard against engagement.
- **Who is measured:** the ranking model and the teams whose OKRs it feeds.
- **Stakes:** very high — engagement drives ad revenue and growth targets.

## Step 2 — Predict the gaming (≥3 vectors each)

**Case A — benchmarks:**
1. **Train on the test (contamination).** Benchmark questions and answers, published openly on the web, get scraped into pretraining or fine-tuning corpora — accidentally or deliberately. The model then "knows" the answers rather than deriving them. Contamination of popular benchmarks was widely documented and discussed across the research community by 2023–2024.
2. **Overfit the format.** Tune specifically to the benchmark's answer style, pr

examples/goodhart-1975-m3-and-strathern-1997-rae.md

# Method in Action: Goodhart 1975 (M3) and Strathern 1997 (RAE)

> *Example for the [goodharts-law](../SKILL.md) skill.*

The empirical foundation has two key moments. The first is **Charles Goodhart's 1975 critique of UK monetary policy**, originally a conference paper for the Reserve Bank of Australia, later expanded in his 1984 book *Monetary Theory and Practice*.

The context: in the mid-1970s, the Bank of England (and many other central banks) had identified a robust historical correlation between growth in broad money supply (the M3 aggregate) and subsequent inflation. The natural policy implication was straightforward: target M3 growth at a level consistent with low inflation, and inflation would be controlled.

Goodhart was skeptical. His objection was structural, not empirical:

> "It is not the case that there is some fixed and stable relationship between an observed monetary aggregate, such as M3, and other variables in the economic system. Rather, the relationships we observe statistically are the equilibrium outcomes of behavior by banks, depositors, and borrowers, each of whom is responding to a complex set of incentives. When the regulator targets M3 — and especially when the targeting carries policy weight that will affect interest rates and reserve requirements — these economic actors will reorganize their behavior to operate around the regulation. Liquid funds will be reclassified into categories that fall outside the M3 definition. The pre-targeting M3-inflation correlation will not survive the targeting. Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes."
>
> — Goodhart (1975), as reprinted in Goodhart (1984), pp. 96-98.

Goodhart's prediction was empirical and falsifiable. It came true within a few years: as the Bank of England's M3 targets bit, UK banks began creating money-substitutes that escaped the M3 definition (most notably, the rise of the eurodollar market and certain types of negotiable certificates of deposit). The M3-inflation correlation collapsed; the Bank quietly abandoned M3 targeting by the mid-1980s.

The principle's second formative moment came two decades later, in **Marilyn Strathern's 1997 ethnographic study of British universities** during the rise of the Research Assessment Exercise (RAE). Strathern, an anthropologist at Cambridge, observed the RAE's effect on academic behavior:

> "The Research Assessment Exercise sets out to evaluate the quality of research conducted in British universities. The exercise was introduced with the goal of identifying excellent research and directing funding toward it. The instrument: publication count, citation analysis, peer-rated quality. The exercise's effect on the academy has been precisely the inversion of its stated goal. Academics have shifted from publishing fewer, more substantial works to publishing more, smaller, shallower ones. They have shifted topic selection toward areas where rapid publication
Github ReposUpdated 1d agoRank 70

AionUi

Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!

MCPOPENCLAW
Github ReposUpdated 6mo agoRank 70

activepieces

AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents

OPENCLAW
Github ReposUpdated 6mo agoRank 70

cherry-studio

AI productivity studio with smart chat, autonomous agents, and 300+ assistants.

MCPOPENCLAW
Github ReposUpdated 7mo agoRank 70

CopilotKit

The Frontend for Agents & Generative UI. React + Angular

OPENCLAW

Machine-readable data

The same record, as JSON, for agents and crawlers.

{
  "facts": [
    {
      "factKey": "vendor",
      "category": "vendor",
      "label": "Vendor",
      "value": "Clawhub",
      "href": "https://clawhub.ai/deciqai/skills/goodharts-law",
      "sourceUrl": "https://clawhub.ai/deciqai/skills/goodharts-law",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-11T00:08:25.475Z",
      "isPublic": true
    },
    {
      "factKey": "protocols",
      "category": "compatibility",
      "label": "Protocol compatibility",
      "value": "OpenClaw",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-deciqai-goodharts-law/contract",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-deciqai-goodharts-law/contract",
      "sourceType": "contract",
      "confidence": "medium",
      "observedAt": "2026-10-11T00:08:25.475Z",
      "isPublic": true
    },
    {
      "factKey": "traction",
      "category": "adoption",
      "label": "Adoption signal",
      "value": "1.2K downloads",
      "href": "https://clawhub.ai/deciqai/goodharts-law",
      "sourceUrl": "https://clawhub.ai/deciqai/goodharts-law",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-11T00:08:25.475Z",
      "isPublic": true
    },
    {
      "factKey": "latest_release",
      "category": "release",
      "label": "Latest release",
      "value": "1.0.5",
      "href": "https://clawhub.ai/deciqai/goodharts-law",
      "sourceUrl": "https://clawhub.ai/deciqai/goodharts-law",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-07-16T18:01:32.110Z",
      "isPublic": true
    },
    {
      "factKey": "handshake_status",
      "category": "security",
      "label": "Handshake status",
      "value": "UNKNOWN",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-deciqai-goodharts-law/trust",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-deciqai-goodharts-law/trust",
      "sourceType": "trust",
      "confidence": "medium",
      "observedAt": null,
      "isPublic": true
    }
  ],
  "events": [
    {
      "eventType": "release",
      "title": "Release 1.0.5",
      "description": "Description tail link + agents machine-readable metadata line (deciqai.com/s/goodharts-law.json)",
      "href": "https://clawhub.ai/deciqai/goodharts-law",
      "sourceUrl": "https://clawhub.ai/deciqai/goodharts-law",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-07-16T18:01:32.110Z",
      "isPublic": true
    }
  ]
}

Record generated Oct 11, 2026.

Sponsored

Ads related to Goodhart's Law and adjacent AI workflows.