Claim this agent
agentCLAWHUBUnverified

ia-writing-tests

Generic test writing discipline: test quality, real assertions, anti-patterns, and rationalization resistance. Use when writing tests, adding test coverage, or fixing failing tests for any language or framework. Complements language-specific skills. Skill: ia-writing-tests Owner: iliaal Summary: Generic test writing discipline: test quality, real assertions, anti-patterns, and rationalization resistance. Use when writing tests, adding test coverage, or fixing failing tests for any language or framework. Complements language-specific skills. Tags: latest:5.0.1 Version history: v5.0.1 | 2026-10-03T17:08:52.717Z | user v5.0.1 v5.0.0 | 2026-09-26T23:25:09.836Z | use

OpenClaw

Rank

62

Safety

84

Downloads

2.2k

Updated

Oct 9, 2026

Version

5.0.1

Source

CLAWHUB

About

What it does, and when to use it.

Capability contract not published. No trust telemetry is available yet. 2.2K downloads reported by the source. Last updated 10/9/2026.

Avoid when

  • Contract metadata is missing or unavailable for deterministic execution.

Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing

Public facts

Every fact links back to the source it came from.

Vendor
Clawhubvendor · observed Oct 9, 2026
Protocol compatibility
OpenClawcompatibility · observed Oct 9, 2026
Adoption signal
2.2K downloadsadoption · observed Oct 9, 2026
Latest release
5.0.1release · observed Oct 3, 2026
Handshake status
UNKNOWNsecurity

Install and run

Setup complexity: low.

clawhub skill install s17bcar8wq0xhegs0ny6f57ypd8484bw:compound-eng-writing-tests
  1. Setup complexity is classified as HIGH. You must provision dedicated cloud infrastructure or an isolated VM. Do not run this directly on your local workstation.
  2. Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.

Contract: missing

curl -s "https://www.xpersona.co/api/v1/agents/clawhub-iliaal-compound-eng-writing-tests/snapshot"

Documentation

CLAWHUB

146,490 characters of source documentation, loaded on request.

Extracted files

5 files captured from the source.

SKILL.md

---
name: ia-writing-tests
class: discipline
description: >-
  Generic test writing discipline: test quality, real assertions, anti-patterns,
  and rationalization resistance. Use when writing tests, adding test coverage,
  or fixing failing tests for any language or framework. Complements
  language-specific skills.
---

# Writing tests

Produce tests that prove the requested behavior and fail when that behavior breaks. Follow user scope and the repository's actual contracts; this skill does not authorize implementation, external actions, or changes to acceptance criteria.

## Procedure

1. Discover the repository's runner, pinned wrapper, configuration, neighboring tests, and CI command before writing tests. Learn both focused and full-suite commands. Use project tooling rather than a global default; keep tracked scripts portable and apply personal wrappers only to the outer invocation.
2. Derive cases from user requirements, implemented behavior, and claims intended for the handoff. Map each acceptance criterion to a discriminating case that a plausible wrong implementation fails. Name tests by observable behavior; keep one behavior per test and make fixtures adversarial on the axis under test.
3. For bug fixes, write the reproducer, observe the intended failure, apply the smallest fix, and observe green. For new features, follow the selected test posture in vertical slices: default to tests-after (one small implementation slice, then its test); with an explicit test-first choice, write one test, then the minimal implementation that passes it. Repeat by slice; writing all tests first and all implementation afterward (horizontal slicing) is an anti-pattern that produces tests for imagined behavior. Refactor after green. Never alter a specification, assertion, fixture, snapshot, or expected output merely to make the implementation pass.
4. Prefer real internal objects, real temporary files, and real test databases. Mock only external boundaries, at the last owned adapter; framework-maintained fakes are appropriate where the framework recommends them. Check the real contract and side effects before mocking.
5. Assert consumer-visible outcomes rather than private structure, mock calls, or framework behavior. Include relevant boundary, invalid-input, concurrency, and failure cases. If error handling catches, logs, substitutes, or rolls back, assert its observable result as well as error visibility.
6. Prove absence/isolation assertions with a run-unique forbidden violation and observe that specific assertion fail. Confirm the control mutation actually landed and its build completed before evaluating it. Remove the control, restore the artifact, and verify green.
7. Run the narrow checks during edits and the applicable complete checks before handoff. Inspect passed/executed counts; an empty or skipped suite is not positive evidence.

## Select the detail needed

- When choosing cases, fixture shape, mock boundaries, or unit/integration/E2E balance, 

_meta.json

{
  "ownerId": "kn715jrbbh71q9zncr0bqdkr8n848q1a",
  "slug": "compound-eng-writing-tests",
  "version": "5.0.1",
  "publishedAt": 1791047332717
}

references/anti-patterns-extended.md

# Anti-Patterns: Extended Notes

Detail offloaded from the SKILL.md Anti-Patterns section. The symptom/fix one-liners live inline; this file holds the mechanics, root-cause narratives, and fix ladders.

## Persistent test infrastructure state contamination

**Root cause:** persistent test infrastructure (a long-running `docker compose up`, a shared local database, a volume left between iterations) accumulates state across test runs. The current run's data sits on top of the previous run's data; assertions counting rows or jobs see the sum. The numbers look like a code bug ("the loop runs N times instead of once"), but they are clean integer multiples of the expected value, and the same test passes in CI on a fresh container.

**Fix ladder**, in order of preference:

1. **Ephemeral containers per test session** (`testcontainers`, `pytest-postgresql`, or `docker compose run --rm <service>` for one-shot runs): slowest to start, strongest isolation. Default for CI.
2. **Fixture-driven `TRUNCATE` / `DROP DATABASE`** in a session-scoped or per-test fixture: fast, but requires careful coverage of every stateful table.
3. **Volume teardown between iterations** (`docker compose down -v` before each run) when running locally: manual but reliable.

Never rely on tests "cleaning up after themselves." If a previous run errored mid-test, the cleanup didn't run, and the next run inherits the partial state.

## Synchronous adapters hide timing-dependent races

**Wire-latency mechanics:** a test fires two or more parallel requests through a mock/adapter that resolves synchronously (a promise that settles in the same microtask, an in-memory fake with zero latency) and asserts a coalescing/dedup/single-flight guard held. It passes, but only because every call observed the shared in-flight state before any reset ran. Under real wire latency the staggered arrivals miss the window, and the guard spawns N operations instead of one.

Same-tick microtask concurrency is not a proxy for production burst behavior. For dedup/coalescing logic, inject controllable latency (fake timers, deferred resolution staggered across ticks) so a later arrival lands after the reset, and assert the guard holds for arrival-staggered bursts.

## Constructing the object-under-test below the layer that transforms it

**Extended rationale:** when a fix guards or transforms a field in an upstream layer (a parser, normalizer, `from_api_response` constructor, serializer) and the test builds the object directly via the leaf constructor (`Model(field=x)`, `new T(...)`, the raw initializer), the test injects the already-correct value. The upstream strip/transform never runs, the guard never fires, and the test is green while production is still broken. The test cannot fail for the exact bug it was written to catch.

Enter through the same entry point production uses. If a test must construct the leaf form directly for other reasons, it is not covering the transform; add a separate test that feeds the 

references/false-pass-oracle-traps.md

# False-Pass Oracle Traps

Assertion oracles that can report success without observing failure, moved from the SKILL.md Anti-Patterns section.

## Piping a command into `grep -q` to assert on its output

**Symptom:** the assertion reports "absent" for a string plainly present in a manual run. Under `set -o pipefail` the pipeline's status is the *writer's*: `grep -q` exits on first match and closes the pipe, so the producer dies of SIGPIPE and the pipeline fails **because the assertion matched**. Independently, a command that legitimately exits non-zero (a refusal path, a status code that is part of the contract) fails the pipeline regardless of the match. Either way the false negative reads as a behavioral finding and sends you into the production code.

**Fix:** capture, then match. `out=$(cmd 2>&1)` on one line, `grep -q 'needle' <<<"$out"` on the next; never put the command and the matcher in one pipeline. Two adjacent shapes to avoid in assertions: `grep -c` prints `0` *and* exits 1, so `grep -c p f || printf 0` emits `0\n0`; and `cmd && x || y` is not if-then-else: `y` also runs when `x` fails.

## A comparison oracle that fails open

**Symptom:** the harness compares two producers with `diff -q <(producer_a | filter) <(producer_b | filter)`. `diff` observes the streams, not whether either producer succeeded: two failed producers yield two empty streams and compare equal. Comparing only added lines has the same hole: two deletions of *different* content both produce an empty `^+` stream, and identical file and line counts do not mean identical content.

**Fix:** capture each producer to a temporary file and check its exit status before comparing. Compare both added and removed hunk bodies from a zero-context diff, not summaries or diffstats. Classify a producer failure, a binary or metadata-only patch, or a comparison I/O error as *undecidable*, never as *identical*.

## A green suite over a feature the test environment disables

**Symptom:** The code path that would fail is behind a config or environment flag that defaults off, and the test environment sets no override. Every test exercising the affected object passes, including tests written for it, and the failure appears on the first write in an environment where the flag is on. Grepping the repository reinforces the wrong conclusion, because the enabled value lives in deployment configuration (a task definition, a parameter store), not in the codebase; the only value in the tree is the `false` default.

**Fix:** Before reading a pass as coverage, check whether the flag gating the consumer that would fail is on under test. Re-run one existing test with the flag forced on, alongside a test touching only unaffected objects as a control, so a failure is attributable to the flag rather than to the environment change. When the disabling guard keys on the test runner itself (an environment variable the runner sets), no in-suite test can observe the behavior at all. Verify that surface out of b

references/generated-corpus-techniques.md

# Generated-Corpus Techniques

Cases enumerated by a generator rather than typed by hand, offloaded from the SKILL.md Writing Good Tests section. Reach for these where the behavior under test emerges from a whole table, or where the change is mechanical and the existing suite is the wrong instrument.

## A hand-picked corpus for emergent-precedence behavior

**Symptom:** the behavior emerges from a whole table (phrase lists, route priorities, rule sets), but the cases are typed by hand, so they test the author's model of the table rather than the table. Every case the author did not think of is a precedence interaction nobody has seen.

**Fix:** build the corpus as the cross-product of the axes the table enumerates, drive it through the real matcher against both the old and the new table, and diff the outcomes. Grouping the analysis on the hand-picked axis re-imposes the same blind spot the corpus was built to remove. Sample a real corpus wherever one exists, and keep the generated diff as the review artifact rather than a prose summary of it.

## Proving a mechanical refactor behavior-preserving with a generated transcript

**Symptom:** a mechanical refactor is declared safe because the suite is green. The suite covers what someone thought to test, and a mechanical refactor can move anything else, including the surfaces nobody wrote a case for.

**Fix:** enumerate the public surface by reflection, call each entry with a per-type pool of edge values varying one parameter at a time, and print one deterministic line per call: return value, warning, exception. Run that against both revisions and diff the transcripts. Rebuild fixtures before every call, so a mutated fixture does not read as a behavior change. Require every surviving difference to map to an intended change, and treat an unexplained difference as the finding rather than as transcript noise.

## Numerical tolerance in differential comparisons

When a field's contract permits floating-point rounding differences, declare the affected field, units, absolute or relative bounds, and justification before comparing outputs. Derive the tolerance from the contract and numerical method. Keep identifiers, integer counts, text, and exact decimal contracts exact. Apply tolerance only to the declared fields; do not mask whole records or widen bounds after seeing an unexpected difference. Confirm that a plausible wrong numerical result outside the permitted bounds fails the comparison.
Github ReposUpdated 1h agoRank 70

AionUi

Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!

MCPOPENCLAW
Github ReposUpdated 6mo agoRank 70

activepieces

AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents

OPENCLAW
Github ReposUpdated 6mo agoRank 70

cherry-studio

AI productivity studio with smart chat, autonomous agents, and 300+ assistants.

MCPOPENCLAW
Github ReposUpdated 7mo agoRank 70

CopilotKit

The Frontend for Agents & Generative UI. React + Angular

OPENCLAW

Machine-readable data

The same record, as JSON, for agents and crawlers.

{
  "facts": [
    {
      "factKey": "vendor",
      "category": "vendor",
      "label": "Vendor",
      "value": "Clawhub",
      "href": "https://clawhub.ai/iliaal/skills/compound-eng-writing-tests",
      "sourceUrl": "https://clawhub.ai/iliaal/skills/compound-eng-writing-tests",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-09T17:15:34.126Z",
      "isPublic": true
    },
    {
      "factKey": "protocols",
      "category": "compatibility",
      "label": "Protocol compatibility",
      "value": "OpenClaw",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-iliaal-compound-eng-writing-tests/contract",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-iliaal-compound-eng-writing-tests/contract",
      "sourceType": "contract",
      "confidence": "medium",
      "observedAt": "2026-10-09T17:15:34.126Z",
      "isPublic": true
    },
    {
      "factKey": "traction",
      "category": "adoption",
      "label": "Adoption signal",
      "value": "2.2K downloads",
      "href": "https://clawhub.ai/iliaal/compound-eng-writing-tests",
      "sourceUrl": "https://clawhub.ai/iliaal/compound-eng-writing-tests",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-09T17:15:34.126Z",
      "isPublic": true
    },
    {
      "factKey": "latest_release",
      "category": "release",
      "label": "Latest release",
      "value": "5.0.1",
      "href": "https://clawhub.ai/iliaal/compound-eng-writing-tests",
      "sourceUrl": "https://clawhub.ai/iliaal/compound-eng-writing-tests",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-10-03T17:08:52.717Z",
      "isPublic": true
    },
    {
      "factKey": "handshake_status",
      "category": "security",
      "label": "Handshake status",
      "value": "UNKNOWN",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-iliaal-compound-eng-writing-tests/trust",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-iliaal-compound-eng-writing-tests/trust",
      "sourceType": "trust",
      "confidence": "medium",
      "observedAt": null,
      "isPublic": true
    }
  ],
  "events": [
    {
      "eventType": "release",
      "title": "Release 5.0.1",
      "description": "v5.0.1",
      "href": "https://clawhub.ai/iliaal/compound-eng-writing-tests",
      "sourceUrl": "https://clawhub.ai/iliaal/compound-eng-writing-tests",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-10-03T17:08:52.717Z",
      "isPublic": true
    }
  ]
}

Record generated Oct 9, 2026.

Sponsored

Ads related to ia-writing-tests and adjacent AI workflows.