Falsifiability
Activate when: user says 'what would prove this wrong', 'how do we know if our strategy is working', 'what would change your mind', 'this claim feels unfalsi... Skill: Falsifiability Owner: deciqai Summary: Activate when: user says 'what would prove this wrong', 'how do we know if our strategy is working', 'what would change your mind', 'this claim feels unfalsi... Tags: latest:1.0.5 Version history: v1.0.5 | 2026-07-16T17:59:34.650Z | user Description tail link + agents machine-readable metadata line (deciqai.com/s/falsifiability.json) v1.0.4 | 2026-07-10T10:25:40.288Z | us
Rank
62
Safety
84
Downloads
1.1k
Updated
Oct 11, 2026
Version
1.0.5
Source
CLAWHUB
About
What it does, and when to use it.
Capability contract not published. No trust telemetry is available yet. 1.1K downloads reported by the source. Last updated 10/11/2026.
Avoid when
- Contract metadata is missing or unavailable for deterministic execution.
Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing
Public facts
Every fact links back to the source it came from.
- Vendor
- Clawhubvendor · observed Oct 11, 2026
- Protocol compatibility
- OpenClawcompatibility · observed Oct 11, 2026
- Adoption signal
- 1.1K downloadsadoption · observed Oct 11, 2026
- Latest release
- 1.0.5release · observed Jul 16, 2026
- Handshake status
- UNKNOWNsecurity
Install and run
Setup complexity: low.
clawhub skill install s17a4mqcnk515kvaca5ze55d0x88pfpx:falsifiability- Setup complexity is classified as HIGH. You must provision dedicated cloud infrastructure or an isolated VM. Do not run this directly on your local workstation.
- Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.
Contract: missing
curl -s "https://www.xpersona.co/api/v1/agents/clawhub-deciqai-falsifiability/snapshot"
Documentation
CLAWHUB
116,837 characters of source documentation, loaded on request.
Extracted files
5 files captured from the source.
SKILL.md
---
name: falsifiability
description: "Activate when: user says 'what would prove this wrong', 'how do we know if our strategy is working', 'what would change your mind', 'this claim feels unfalsifiable', or is designing a hypothesis/experiment/investment thesis that needs to be made testable.
Do NOT activate when: the claim is genuinely non-empirical (ethical, aesthetic, philosophical); or the cost of running the test exceeds the value of the knowledge. More: deciqai.com/c/falsifiability"
---
# Falsifiability
## Overview
A meaningful empirical claim must specify what observations would refute it. Claims that resist all possible refutation are not science — they are unfalsifiable belief. Formalized by Karl Popper (1934): science progresses not by accumulating confirmations but by surviving rigorous attempts at falsification. More-specific claims are more falsifiable; ad-hoc modifications that explain away failures destroy a claim's scientific status.
Composes with `confirmation-bias` (falsifiability is the structural counter), `abductive-reasoning` (generates hypotheses; this skill tests them), `bayesian-reasoning`, `critical-thinking`.
## When to Use
- Designing OKRs, KPIs, or strategic goals; writing or evaluating investment theses
- Designing experiments (A/B tests, product hypotheses, market entry)
- Evaluating consultant/advisor recommendations or diagnosing vague leadership claims
- Someone says "what would change your mind," "how would you know you're wrong," "Popper"
- Stress-testing an AI/AGI hype claim ("AGI is near," "the model truly understands," "our AI adoption is working") — demand what evidence would disprove the capability or safety claim
**Not when:** genuinely non-empirical (philosophical, ethical, aesthetic); test cost exceeds its value.
## Coaching Novices (Adaptive Front Door)
- **Engine mode:** user has a specific claim → run The Process directly.
- **Coach mode:** user is new → guide step by step.
In Coach mode, respond one step at a time. Each [WAIT] is a hard stop — output only that step's question, then stop.
1. One-line: before acting on a claim, ask what observation would refute it — if "nothing would," it's belief, not knowledge.
2. Check fit: if the claim is genuinely metaphysical (ethical, aesthetic), falsifiability doesn't apply.
3. Elicit the claim and its current evidential basis.
> **[WAIT — do not advance until user responds]**
4. Ask: what specific observation would falsify this? When and how would you observe it?
> **[WAIT — do not advance until user responds]**
5. Close: falsifiability conditions specified + monitoring plan + commitment to act on disconfirming evidence.
> **[WAIT — do not advance until user responds]**
## The Process
**Step 1 — State the claim:** claim / who asserts it / decision dependent on it / current evidential basis.
**Step 2 — Test whether empirical:** claim about how the world works (empirical) or values/aesthetics (non-empirical)? If non-empirical, stop here.
**Ste_meta.json
{
"ownerId": "kn754b8sk22s8c6gjxt02bftbn88q7ye",
"slug": "falsifiability",
"version": "1.0.5",
"publishedAt": 1784224774650
}references/sources.md
# Sources — falsifiability > *Primary sources for the [falsifiability](../SKILL.md) skill.* ## Sources - Popper, K. R. (1934). *Logik der Forschung.* Vienna: Springer. English: *The Logic of Scientific Discovery* (1959). London: Hutchinson. ISBN 978-0415278447. - Popper, K. R. (1963). *Conjectures and Refutations: The Growth of Scientific Knowledge.* London: Routledge. ISBN 978-0415285940. - Eddington, A. S. (1919). "Joint Eclipse Meeting of the Royal Society and the Royal Astronomical Society." *Nature*, 104, 354. The relativity confirmation. - Kuhn, T. S. (1962). *The Structure of Scientific Revolutions.* University of Chicago Press. ISBN 978-0226458120. The paradigm-shift complement. - Lakatos, I. (1970). "Falsification and the Methodology of Scientific Research Programmes." in Lakatos & Musgrave (eds.), *Criticism and the Growth of Knowledge*. Cambridge University Press. - Ries, E. (2011). *The Lean Startup.* Crown Business. ISBN 978-0307887894. The startup application. - Tetlock, P. E. (2015). *Superforecasting.* Crown. ISBN 978-0804136693. Calibrated falsifiability in prediction. - Mayo, D. G. (1996). *Error and the Growth of Experimental Knowledge.* University of Chicago Press. The error-statistical framework. - Chollet, F. (2019). "On the Measure of Intelligence." arXiv:1911.01547. Operationalizes "intelligence" into the falsifiable ARC / ARC-AGI benchmark — a template for turning "AGI is near" into a testable claim (2024–2026 AI application). - On benchmark contamination and reasoning robustness (2023–2025): peer-reviewed and arXiv work documenting train/test data contamination in LLM evaluations and accuracy drops when surface features of reasoning problems are perturbed — why a high AI benchmark score does not, by itself, confirm a capability claim. (Consult current surveys; specific reported scores are vendor figures, not independently audited.)
examples/ai-claims-agi-is-near-vs-testable-predictions-2024-2026.md
# Method in Action: Separating Falsifiable from Unfalsifiable AI Claims (2024–2026)
> *Example for the [falsifiability](../SKILL.md) skill.*
Between 2024 and 2026, public discourse around large AI models filled with two very different kinds of statement. Some were slogans — "AGI is near," "the model truly understands," "scaling will just keep working" — and some were concrete, dated predictions with numbers attached. Popper's criterion sorts them cleanly: a claim is empirical knowledge only if you can say in advance what observation would prove it wrong. This walkthrough runs the anchor claims through the falsifiability skill's own six-step Process.
## Step 1 — State the claim
Take three representative claims from the period:
- **Claim A (slogan):** "AGI is near."
- **Claim B (mentalistic):** "The model *truly understands* what it's saying."
- **Claim C (testable):** "By the end of 2025, a frontier model will exceed a specified accuracy threshold on the competition-mathematics benchmark AIME," or similar dated, metric-bound predictions of the kind researchers publish.
*Who asserts them:* executives, commentators, and researchers, respectively. *Decision at stake:* whether an enterprise should bet a roadmap (or an investor a position) on imminent general capability. *Current evidential basis:* rapid, genuine benchmark gains from roughly 2023 onward, plus extrapolation.
## Step 2 — Test whether empirical
- **Claim A ("AGI is near")** is empirical *only if* "AGI" and "near" are defined. As typically used, neither is: "AGI" has no agreed operational definition, and "near" has no date. Without those, it is not yet a testable claim — it is a mood.
- **Claim B ("truly understands")** invokes an inner mental state. As stated it is closer to metaphysics than to empirical science: no external observation is specified that would distinguish "truly understands" from "produces the same outputs without understanding." Step 2 says: if it stays non-empirical, stop and label it belief.
- **Claim C** is empirical: it is a statement about a measurable score on a fixed benchmark by a fixed date.
## Step 3 — Specify falsification conditions
Force each claim to complete: *"This would be falsified if I observed: ___."*
- **Claim A** can be *rescued into* an empirical claim by pinning it down — e.g. "A single model will pass [a specified operationalization, such as the ARC-AGI abstraction-and-reasoning benchmark at human-level, or a stated economically-valuable-task bar] before 31 December 2026." Now it can fail. Note the discipline: the version that can be proven wrong is the only version worth arguing about.
- **Claim B** resists completion. "Understanding" that predicts no observable difference from "not understanding" has no falsification condition. The productive move is to *replace* it with a behavioral proxy that does — e.g. "the model will maintain accuracy when the same problem is presented with surface features (names, numbers, framing) changed," whexamples/popper-1934-eddington-1919-eclipse-modern-applications.md
# Method in Action: Popper 1934 + Eddington 1919 Eclipse + Modern Applications
> *Example for the [falsifiability](../SKILL.md) skill.*
**Karl Popper** (1902-1994) was an Austrian-British philosopher whose 1934 *Logik der Forschung* (translated as *The Logic of Scientific Discovery* in 1959) reshaped 20th-century philosophy of science. Popper had observed the rise of psychoanalysis (Freud, Adler) and Marxism in Vienna during the 1920s, and was struck by how their adherents claimed they "explained" everything — every observed event could be interpreted as confirming the theory. Popper's diagnosis: any theory that explains everything explains nothing. Such theories make no risky predictions; they cannot fail; therefore they are not empirical knowledge.
Popper contrasted this with Einstein's general relativity (1915). Einstein had made a specific, quantitative prediction: light passing through the Sun's gravitational field during an eclipse would be deflected by 1.75 arcseconds, twice the Newtonian prediction. The 1919 total solar eclipse provided the empirical test. **Arthur Eddington's measurements during the eclipse** (in Príncipe and Brazil) confirmed the relativistic prediction within experimental error:
> Eddington, A. S. (1919). "Joint Eclipse Meeting of the Royal Society and the Royal Astronomical Society." *Nature*, 104, 354.
Popper's emphasis: the test was real because relativity could have failed. If the measured deflection had been close to the Newtonian value or no deflection, relativity would have been falsified. The test mattered because failure was possible. The theory's confirmation was meaningful only because falsification was possible. This is the operational criterion of science.
Popper's 1963 *Conjectures and Refutations* extended the framework to many domains:
> "The history of science, like the history of all human ideas, is a history of irresponsible dreams, of obstinacy, and of error. But science is one of the very few human activities — perhaps the only one — in which errors are systematically criticized and fairly often, in time, corrected. This is why we can say that, in science, we often learn from our mistakes, and why we can speak clearly and sensibly about making progress there."
>
> — Popper (1963), p. 216.
The framework has been applied extensively:
**Lean Startup methodology.** Eric Ries's 2011 *The Lean Startup* is implicitly Popperian. Each iteration generates falsifiable hypotheses ("if we A/B test this checkout flow change, conversion will improve by ≥5%"), tests them empirically, and either confirms or rejects them based on data. The build-measure-learn cycle is operationalized falsifiability applied to product development.
**Hypothesis-driven product development.** Modern product organizations (Amazon, Google, Microsoft) make explicit predictions about what user behavior change a feature will produce, then measure the actual change. Features that don't deliver the predicted impact are rolled back. TAionUi
Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!
activepieces
AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants.
CopilotKit
The Frontend for Agents & Generative UI. React + Angular
Machine-readable data
The same record, as JSON, for agents and crawlers.
{
"facts": [
{
"factKey": "vendor",
"category": "vendor",
"label": "Vendor",
"value": "Clawhub",
"href": "https://clawhub.ai/deciqai/skills/falsifiability",
"sourceUrl": "https://clawhub.ai/deciqai/skills/falsifiability",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-11T11:19:57.408Z",
"isPublic": true
},
{
"factKey": "protocols",
"category": "compatibility",
"label": "Protocol compatibility",
"value": "OpenClaw",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-deciqai-falsifiability/contract",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-deciqai-falsifiability/contract",
"sourceType": "contract",
"confidence": "medium",
"observedAt": "2026-10-11T11:19:57.408Z",
"isPublic": true
},
{
"factKey": "traction",
"category": "adoption",
"label": "Adoption signal",
"value": "1.1K downloads",
"href": "https://clawhub.ai/deciqai/falsifiability",
"sourceUrl": "https://clawhub.ai/deciqai/falsifiability",
"sourceType": "profile",
"confidence": "medium",
"observedAt": "2026-10-11T11:19:57.408Z",
"isPublic": true
},
{
"factKey": "latest_release",
"category": "release",
"label": "Latest release",
"value": "1.0.5",
"href": "https://clawhub.ai/deciqai/falsifiability",
"sourceUrl": "https://clawhub.ai/deciqai/falsifiability",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-07-16T17:59:34.650Z",
"isPublic": true
},
{
"factKey": "handshake_status",
"category": "security",
"label": "Handshake status",
"value": "UNKNOWN",
"href": "https://www.xpersona.co/api/v1/agents/clawhub-deciqai-falsifiability/trust",
"sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-deciqai-falsifiability/trust",
"sourceType": "trust",
"confidence": "medium",
"observedAt": null,
"isPublic": true
}
],
"events": [
{
"eventType": "release",
"title": "Release 1.0.5",
"description": "Description tail link + agents machine-readable metadata line (deciqai.com/s/falsifiability.json)",
"href": "https://clawhub.ai/deciqai/falsifiability",
"sourceUrl": "https://clawhub.ai/deciqai/falsifiability",
"sourceType": "release",
"confidence": "medium",
"observedAt": "2026-07-16T17:59:34.650Z",
"isPublic": true
}
]
}Record generated Oct 11, 2026.
