agentCLAWHUBUnverified

failure-forensics

Use when an agent task fails or produces unexpected results. Performs structured post-mortem root cause analysis: categorizes the failure, traces the exact failure point through tool-call logs, reconstructs the decision chain, generates a post-mortem report, and saves lessons to prevent recurrence. Skill: failure-forensics Owner: voronindenis5 Summary: Use when an agent task fails or produces unexpected results. Performs structured post-mortem root cause analysis: categorizes the failure, traces the exact failure point through tool-call logs, reconstructs the decision chain, generates a post-mortem report, and saves lessons to prevent recurrence. Tags: latest:0.1.1 Version history: v0.1.1 | 2026-08-11T11:58:54.

OpenClaw

Rank

62

Safety

84

Downloads

2.5k

Updated

Oct 9, 2026

Version

0.1.1

Source

CLAWHUB

About

What it does, and when to use it.

Capability contract not published. No trust telemetry is available yet. 2.5K downloads reported by the source. Last updated 10/9/2026.

Avoid when

  • Contract metadata is missing or unavailable for deterministic execution.

Risk flags: missing_or_unavailable_contract, trust_data_unavailable, schema_references_missing

Public facts

Every fact links back to the source it came from.

Vendor
Clawhubvendor · observed Oct 9, 2026
Protocol compatibility
OpenClawcompatibility · observed Oct 9, 2026
Adoption signal
2.5K downloadsadoption · observed Oct 9, 2026
Latest release
0.1.1release · observed Aug 11, 2026
Handshake status
UNKNOWNsecurity

Install and run

Setup complexity: low.

clawhub skill install s17b6amkd3wzqgg640v03a9r1n83gxs1:failure-forensics
  1. Setup complexity is LOW. This package is likely designed for quick installation with minimal external side-effects.
  2. Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data.

Contract: missing

curl -s "https://www.xpersona.co/api/v1/agents/clawhub-voronindenis5-failure-forensics/snapshot"

Documentation

CLAWHUB

69,821 characters of source documentation, loaded on request.

Extracted files

5 files captured from the source.

SKILL.md

---
name: failure-forensics
description: "Use when an agent task fails or produces unexpected results. Performs structured post-mortem root cause analysis: categorizes the failure, traces the exact failure point through tool-call logs, reconstructs the decision chain, generates a post-mortem report, and saves lessons to prevent recurrence."
version: 1.0.0
author: Denis Voronin
license: MIT
metadata:
  hermes:
    tags: [debugging, post-mortem, forensics, root-cause-analysis, failure-analysis, agent-reliability]
    related_skills: [systematic-debugging, debugging-hermes-tui-commands]
---

# Failure Forensics

## Overview

When an agent task fails, the default response is to retry — hoping for a different outcome. **Failure Forensics** rejects that reflex. Instead, the agent performs structured root cause analysis *before* retrying, treating every failure as evidence to be collected, categorized, and learned from.

The workflow has four phases:

1. **Triage** — Categorize the failure using the taxonomy in [`references/failure-taxonomy.md`](references/failure-taxonomy.md).
2. **Timeline Reconstruction** — Parse tool-call logs and agent decision points to build a chronological failure timeline. The script [`scripts/failure_forensics.py`](scripts/failure_forensics.py) automates this from JSON or JSONL log formats.
3. **Causal Chain Analysis** — Trace the chain of decisions, assumptions, and actions that led from the task kickoff to the failure point. Identify the *root cause*, not just the proximate symptom.
4. **Post-Mortem Report** — Generate a structured report from the template in [`references/post-mortem-template.md`](references/post-mortem-template.md) and persist it so future sessions can learn.

This skill turns a single failure into a permanent, reusable lesson.

## When to Use

- **An agent task failed** and retrying without understanding *why* is risky.
- **A failure recurs** across attempts — you suspect a systemic cause, not bad luck.
- **You need an artifact** documenting what went wrong for a team review or audit.
- **A complex multi-step task** partially completed then broke — you need to understand which step is safe to resume from.
- **You want to improve agent reliability** by building a corpus of past failure patterns.

### Don't use for:

- **Trivial failures with obvious fixes** (typo in a command, missing flag). Fix and move on.
- **Live debugging** of an actively failing process — use `systematic-debugging` for that. Run forensics *after* the process is dead or the task is abandoned.
- **Human performance reviews.** This skill analyzes agent + tool behavior, not people.

## The Forensics Workflow

### Phase 1: Triage — Categorize the Failure

Read the full taxonomy in [`references/failure-taxonomy.md`](references/failure-taxonomy.md). At a high level, every failure falls into one of six categories:

| Category | Signature | First Question |
|---|---|---|
| **Network** | Connection refused, timeout, DNS, TLS, 5xx HTTP | "Is the

README.md

# Failure Forensics

> A structured post-mortem analysis skill for AI agents. When a task fails, don't just retry — investigate.

[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](LICENSE)

## What It Does

**Failure Forensics** is a skill for AI agent frameworks (Hermes Agent, OpenClaw, and compatible). When an agent task fails, instead of blindly retrying, the agent performs a four-phase structured root cause analysis:

1. **Triage** — Categorize the failure (network, permissions, logic, environment, dependency, resource)
2. **Timeline Reconstruction** — Parse tool-call logs to build a chronological failure timeline
3. **Causal Chain Analysis** — Trace the decision chain backward to the root cause
4. **Post-Mortem Report** — Generate a structured report and save lessons learned

## Why

Retrying a failed task without understanding why it failed is a gamble. You might:
- Hit the same failure again (wasted effort)
- Mask the real cause with incidental changes (harder to debug later)
- Miss a systemic issue that will recur in different forms

Failure Forensics turns each failure into a reusable lesson, building institutional memory that makes the agent more reliable over time.

## Repository Structure

```
failure-forensics/
├── SKILL.md                          # Main skill definition (YAML frontmatter + workflow)
├── README.md                         # This file
├── LICENSE                           # MIT
├── references/
│   ├── failure-taxonomy.md           # Six-category failure taxonomy with signatures
│   └── post-mortem-template.md       # Fill-in-the-blanks report template
└── scripts/
    └── failure_forensics.py          # Log parser, categorizer, report generator
```

## Quick Start

### As a Hermes Agent Skill

Copy or symlink this directory to your skills folder:

```bash
cp -r failure-forensics/ ~/.hermes/skills/
```

The skill auto-loads. When a task fails, the agent will follow the forensics workflow described in `SKILL.md`.

### Standalone (Script Only)

The Python script works independently — no agent required:

```bash
# Analyze a JSONL log of tool calls
python3 scripts/failure_forensics.py analyze --log session.jsonl --output timeline.md

# Categorize an error message
python3 scripts/failure_forensics.py categorize --error "ConnectionRefusedError: Connection refused"

# Generate a pre-filled post-mortem report
python3 scripts/failure_forensics.py report --log session.jsonl --title "Deploy failure"
```

### Log Format

The analyzer reads JSON or JSONL files where each entry represents a tool call:

```json
{
  "timestamp": "2024-01-15T10:23:45Z",
  "tool": "terminal",
  "args": {"command": "npm install"},
  "result": {"success": false, "error": "EACCES: permission denied"},
  "duration_ms": 1200
}
```

See `scripts/sample_log.jsonl` for a working example.

## Failure Taxonomy (Summary)

| Category | Signature | Example |
|---|---|---|
| **Network** | Connection refused, timeout, DNS, TLS | `curl: (7) Failed 

_meta.json

{
  "ownerId": "kn75wwn4x6djaf28jbykeamazd81gtdp",
  "slug": "failure-forensics",
  "version": "0.1.1",
  "publishedAt": 1786449534013
}

references/failure-taxonomy.md

# Failure Taxonomy

A reference for categorizing failures during Phase 1 (Triage) of the forensics workflow.

## How to Use This Taxonomy

1. Read the error message / failure signature.
2. Match it against the patterns below.
3. If multiple categories match, pick the **most specific** one. A `ModuleNotFoundError` is a *dependency* failure, not an *environment* failure, even though both relate to the system.
4. If no category fits cleanly, record it as **uncategorized** and note the novel pattern. The taxonomy grows by accretion.

---

## 1. Network Failures

**Core question:** Is the endpoint reachable *right now*, and from this environment?

### Signatures

| Pattern | Meaning |
|---|---|
| `Connection refused`, `ConnectionRefusedError` | Port open but nothing listening / rejected |
| `Connection timed out`, `ETIMEDOUT` | Packet dropped, firewall, or host unreachable |
| `Name or service not known`, `NXDOMAIN` | DNS resolution failure |
| `SSL: CERTIFICATE_VERIFY_FAILED` | TLS cert expired, self-signed, or MITM |
| `HTTP 502 Bad Gateway`, `503 Service Unavailable`, `504 Gateway Timeout` | Server-side failure |
| `curl: (7) Failed to connect`, `curl: (28) Connection timed out` | CLI-level network failure |
| `ECONNRESET`, `Connection reset by peer` | Remote end dropped the connection |

### Diagnostic Questions

- Does `curl -v <url>` or `nc -zv <host> <port>` work from the same environment?
- Is this an internal vs. external endpoint? (internal may need VPN/peering)
- Is there a proxy or corporate firewall in play?
- Did this work before? What changed? (network config, DNS, certs)

### Common Root Causes

- Service is down or not started
- Wrong port number (e.g., `:443` vs `:80`)
- DNS misconfiguration or stale cache
- Expired TLS certificate
- Firewall / security group blocking the port
- IPv6 vs IPv4 resolution mismatch

---

## 2. Permissions Failures

**Core question:** Does the credential/token/user have the needed scope for this action?

### Signatures

| Pattern | Meaning |
|---|---|
| `401 Unauthorized`, `HTTP 401` | No credentials, or credentials rejected |
| `403 Forbidden`, `HTTP 403` | Credentials valid, but lack permission |
| `Permission denied`, `EACCES`, `PermissionError` | Filesystem permission denied |
| `Access denied`, `UnauthorizedAccess` | Cloud API / IAM denial |
| `insufficient privileges`, `requires elevated permissions` | OS-level privilege denial |
| `invalid token`, `token expired`, `invalid_grant` | Auth token problem |

### Diagnostic Questions

- What user/service account is the agent running as?
- What scopes/roles does the token have? (check the token's claims, not assumptions)
- Is this a filesystem permission issue (check `ls -la`, `id`, `getfacl`)?
- Is this an API/IAM issue (check the service's permission model)?
- Did the token expire? Check issuance and expiry timestamps.

### Common Root Causes

- Token expired and wasn't refreshed
- Token has correct identity but wrong scope/role
- File owned by a differ

references/post-mortem-template.md

# Post-Mortem Report Template

Fill in every section. If a section doesn't apply, write "N/A — [reason]" rather than deleting it. A blank section is information; a missing section is ambiguity.

---

# Post-Mortem: [TITLE]

**Date:** YYYY-MM-DD
**Author:** [agent name / human name]
**Task:** [one-line description of what the agent was trying to do]
**Status:** [Failed / Partially completed / Recovered after intervention]

## Summary

[One paragraph, plain language. Describe what happened, not just that it failed. A reader who wasn't present should understand the failure and its impact from this paragraph alone. Aim for 3-5 sentences.]

**Failure category:** [network / permissions / logic / environment / dependency / resource / uncategorized]

## Timeline

Reconstruct the sequence of events leading to the failure. Use timestamps from logs where available. Mark the failure point explicitly with **[FAILURE]**.

| Time (UTC) | Event | Outcome |
|---|---|---|
| 10:23:01 | Agent received task: "Deploy service to staging" | Task started |
| 10:23:15 | Agent ran `git pull origin main` | Success |
| 10:23:45 | Agent ran `npm install` | **[FAILURE]** — EACCES: permission denied |
| 10:24:02 | Agent retried `npm install` with `sudo` | Different error: EACCES on different path |
| 10:24:30 | Agent abandoned task | Task failed |

If using the forensics script, paste the generated timeline here.

## Impact

- **What was affected:** [services, data, users, downstream tasks]
- **Severity:** [low / medium / high / critical]
- **Duration of impact:** [how long the system was in a bad state, if applicable]
- **Data loss:** [yes/no — if yes, what and how much]
- **Recovery actions taken:** [what was done to restore service, if anything]

## Root Cause

[The terminal link of the causal chain. State this plainly and specifically. This should be a single, clear sentence that explains *why* the failure happened at the deepest level you could trace.]

**Example (bad):** "npm install failed."
**Example (good):** "The agent ran as a non-root user in a container where the global npm directory (`/usr/lib/node_modules`) was owned by root with no write permission for others, and `npm install` without `--prefix` defaults to global installation."

## Causal Chain

Trace backward from the failure point. Each entry should answer "why did the previous step happen/ matter?"

1. **[FAILURE]** `npm install` returned EACCES on `/usr/lib/node_modules`
2. **Because:** npm attempted a global install (no `--prefix` or local `package.json`)
3. **Because:** The agent assumed the install target was local, but the working directory had no `package.json`
4. **Because:** The agent didn't verify the working directory contents before running the install
5. **Because:** The task description referenced a project at a path the agent assumed existed without checking ← **ROOT CAUSE**

**Root cause:** The agent operated on an unverified assumption about the filesystem state (project path) and cascaded i
Github ReposUpdated 3h agoRank 70

AionUi

Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!

MCPOPENCLAW
Github ReposUpdated 6mo agoRank 70

activepieces

AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents

OPENCLAW
Github ReposUpdated 6mo agoRank 70

cherry-studio

AI productivity studio with smart chat, autonomous agents, and 300+ assistants.

MCPOPENCLAW
Github ReposUpdated 7mo agoRank 70

CopilotKit

The Frontend for Agents & Generative UI. React + Angular

OPENCLAW

Machine-readable data

The same record, as JSON, for agents and crawlers.

{
  "facts": [
    {
      "factKey": "vendor",
      "category": "vendor",
      "label": "Vendor",
      "value": "Clawhub",
      "href": "https://clawhub.ai/voronindenis5/skills/failure-forensics",
      "sourceUrl": "https://clawhub.ai/voronindenis5/skills/failure-forensics",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-09T14:15:58.668Z",
      "isPublic": true
    },
    {
      "factKey": "protocols",
      "category": "compatibility",
      "label": "Protocol compatibility",
      "value": "OpenClaw",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-voronindenis5-failure-forensics/contract",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-voronindenis5-failure-forensics/contract",
      "sourceType": "contract",
      "confidence": "medium",
      "observedAt": "2026-10-09T14:15:58.668Z",
      "isPublic": true
    },
    {
      "factKey": "traction",
      "category": "adoption",
      "label": "Adoption signal",
      "value": "2.5K downloads",
      "href": "https://clawhub.ai/voronindenis5/failure-forensics",
      "sourceUrl": "https://clawhub.ai/voronindenis5/failure-forensics",
      "sourceType": "profile",
      "confidence": "medium",
      "observedAt": "2026-10-09T14:15:58.668Z",
      "isPublic": true
    },
    {
      "factKey": "latest_release",
      "category": "release",
      "label": "Latest release",
      "value": "0.1.1",
      "href": "https://clawhub.ai/voronindenis5/failure-forensics",
      "sourceUrl": "https://clawhub.ai/voronindenis5/failure-forensics",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-08-11T11:58:54.013Z",
      "isPublic": true
    },
    {
      "factKey": "handshake_status",
      "category": "security",
      "label": "Handshake status",
      "value": "UNKNOWN",
      "href": "https://www.xpersona.co/api/v1/agents/clawhub-voronindenis5-failure-forensics/trust",
      "sourceUrl": "https://www.xpersona.co/api/v1/agents/clawhub-voronindenis5-failure-forensics/trust",
      "sourceType": "trust",
      "confidence": "medium",
      "observedAt": null,
      "isPublic": true
    }
  ],
  "events": [
    {
      "eventType": "release",
      "title": "Release 0.1.1",
      "description": "- Removed the file skill-card.md to clean up redundant or outdated documentation. - No changes to core logic or usage. - All main documentation and workflow details now remain in SKILL.md.",
      "href": "https://clawhub.ai/voronindenis5/failure-forensics",
      "sourceUrl": "https://clawhub.ai/voronindenis5/failure-forensics",
      "sourceType": "release",
      "confidence": "medium",
      "observedAt": "2026-08-11T11:58:54.013Z",
      "isPublic": true
    }
  ]
}

Record generated Oct 9, 2026.

Sponsored

Ads related to failure-forensics and adjacent AI workflows.