{"id":"658311ac-d555-4970-ae54-9b7ec7ce4d21","entityType":"agent","slug":"clawhub-zw008-vmware-debug","name":"vmware-debug","canonicalUrl":"https://www.xpersona.co/agent/clawhub-zw008-vmware-debug","canonicalPath":"/agent/clawhub-zw008-vmware-debug","generatedAt":"2026-10-10T10:42:31.798Z","source":"CLAWHUB","claimStatus":"UNCLAIMED","verificationTier":"NONE","summary":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T06:37:53.811Z","emptyReason":null},"description":"Use this skill whenever the user is troubleshooting a VMware/vSphere problem — a reported error, an exception, a log dump, a slow or failed VM, a host that went sideways — and needs help locating the root cause. It is the diagnostic brain of the VMware family: it drives a systematic investigation, pulls the right signals from the other skills, correlates events into one timeline, ranks root-cause hypotheses, and tells you what to check next even when you don't know where to start. Always use this skill for \"diagnose this VMware issue\", \"why is my VM slow\", \"troubleshoot this vSphere error\", \"what does this log mean\", \"help me figure out what broke\" when the context is explicitly VMware/vSphere/ESXi/NSX. It never touches vSphere: its only writes are to a local case ledger. Do NOT use it to execute fixes — single fixes go to vmware-aiops, multi-step gated remediation goes to vmware-pilot. Do NOT use it for routine inventory or health checks with no problem to solve — use vmware-monitor.","descriptionLabel":"Source description","evidenceSummary":"Capability contract not published. No trust telemetry is available yet. 1.6K downloads reported by the source. Last updated 10/10/2026.","installCommand":"clawhub skill install s171xgnmqse0nqvgqvqnaq5f9183kyre:vmware-debug","sourceUrl":"https://clawhub.ai/zw008/vmware-debug","homepage":"https://clawhub.ai/zw008/skills/vmware-debug","primaryLinks":[{"label":"View on ClawHub","url":"https://clawhub.ai/zw008/vmware-debug","kind":"source"},{"label":"Homepage","url":"https://clawhub.ai/zw008/skills/vmware-debug","kind":"homepage"}],"safetyScore":84,"overallRank":62,"popularityScore":64,"trustScore":null,"claimedByName":null,"isOwner":false,"seoDescription":"vmware-debug technical dossier on Xpersona with agent coverage, OPENCLEW support, and live trust metadata."},"coverage":{"evidence":{"source":"public-profile","verified":false,"confidence":"medium","updatedAt":"2026-10-10T06:37:53.811Z","emptyReason":null},"protocols":[{"protocol":"OPENCLEW","label":"OpenClaw","status":"self-declared","notes":"Declared in the public agent profile."}],"capabilities":[],"verifiedCount":0,"selfDeclaredCount":1,"capabilityMatrix":{"rows":[{"key":"OPENCLEW","type":"protocol","support":"unknown","confidenceSource":"profile","notes":"Listed on profile"}],"flattenedTokens":"protocol:OPENCLEW|unknown|profile"}},"adoption":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T06:37:53.811Z","emptyReason":null},"stars":null,"forks":null,"downloads":1614,"packageName":null,"latestVersion":"1.13.1","tractionLabel":"1.6K downloads"},"release":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T06:37:53.810Z","emptyReason":null},"lastUpdatedAt":"2026-10-10T06:37:53.811Z","lastCrawledAt":"2026-10-10T06:37:53.810Z","lastIndexedAt":null,"nextCrawlAt":"2026-10-11T06:37:53.810Z","lastVerifiedAt":null,"highlights":[{"version":"1.13.1","createdAt":"2026-09-16T05:19:39.714Z","changelog":"Exclusion is decided per hypothesis; an empty ledger rules nothing out","fileCount":9,"zipByteSize":23538},{"version":"1.13.0","createdAt":"2026-09-15T14:41:24.361Z","changelog":"Exclusion is per hypothesis: the case is Excluded only when every hypothesis is ruled out; case_grade returns hypotheses.","fileCount":9,"zipByteSize":23394},{"version":"1.12.2","createdAt":"2026-09-15T08:55:11.605Z","changelog":"Every MCP tool call is audited, failures included; timeline results not copied into audit rows; documents what debug inherits from vmware-policy. Requires vmware-policy>=1.16.0.","fileCount":9,"zipByteSize":23355},{"version":"1.12.1","createdAt":"2026-09-15T06:02:04.348Z","changelog":"CLI reads are audited under their MCP tool names; every CLI command declares what it reaches (needs vmware-policy 1.15.0)","fileCount":9,"zipByteSize":21709},{"version":"1.12.0","createdAt":"2026-09-15T03:13:41.339Z","changelog":"real alert titles get a category; case_readiness recognises vmware-aiops","fileCount":9,"zipByteSize":21784},{"version":"1.11.4","createdAt":"2026-09-12T00:20:36.379Z","changelog":"Concurrent writers stop losing case entries: ledger writes take a per-case lock and land atomically, so two agents on one shared OPS_HOME keep every entry. Every machine sharing a cases folder needs this version. The manifests now advertise all 14 tools.","fileCount":9,"zipByteSize":21230},{"version":"1.11.3","createdAt":"2026-08-31T00:37:34.927Z","changelog":"fix: run the suite on a non-UTF-8 machine, and stop one skill answering for another","fileCount":9,"zipByteSize":20278},{"version":"1.11.2","createdAt":"2026-08-30T15:20:48.967Z","changelog":"Second-round fixes from the 2026-08-30 VCF 9.1 re-test; vmware-policy floor raised to 1.11.0 (the engine no longer fails open when rules.yaml cannot be read).","fileCount":9,"zipByteSize":20074}]},"execution":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No published capability contract is available yet."},"installCommand":"clawhub skill install s171xgnmqse0nqvgqvqnaq5f9183kyre:vmware-debug","setupComplexity":"low","setupSteps":["Install using `clawhub skill install s171xgnmqse0nqvgqvqnaq5f9183kyre:vmware-debug` in an isolated environment before connecting it to live workloads.","No published capability contract is available yet, so validate auth and request/response behavior manually.","Review the upstream CLAWHUB listing at https://clawhub.ai/zw008/vmware-debug before using production credentials."],"contract":{"contractStatus":"missing","authModes":[],"requires":[],"forbidden":[],"supportsMcp":false,"supportsA2a":false,"supportsStreaming":false,"inputSchemaRef":null,"outputSchemaRef":null,"dataRegion":null,"contractUpdatedAt":null,"sourceUpdatedAt":null,"freshnessSeconds":null},"invocationGuide":{"preferredApi":{"snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-vmware-debug/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-vmware-debug/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-vmware-debug/trust"},"curlExamples":["curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-zw008-vmware-debug/snapshot\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-zw008-vmware-debug/contract\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-zw008-vmware-debug/trust\""],"jsonRequestTemplate":{"query":"summarize this repo","constraints":{"maxLatencyMs":2000,"protocolPreference":["OPENCLEW"]}},"jsonResponseTemplate":{"ok":true,"result":{"summary":"...","confidence":0.9},"meta":{"source":"CLAWHUB","generatedAt":"2026-10-10T10:42:31.793Z"}},"retryPolicy":{"maxAttempts":3,"backoffMs":[500,1500,3500],"retryableConditions":["HTTP_429","HTTP_503","NETWORK_TIMEOUT"]}},"endpoints":{"dossierUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-vmware-debug/dossier","snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-vmware-debug/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-vmware-debug/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-vmware-debug/trust"}},"reliability":{"evidence":{"source":"runtime-metrics","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No trust, reliability, or runtime telemetry is available."},"trust":{"status":"unavailable","handshakeStatus":"UNKNOWN","verificationFreshnessHours":null,"reputationScore":null,"p95LatencyMs":null,"successRate30d":null,"fallbackRate":null,"attempts30d":null,"trustUpdatedAt":null,"trustConfidence":"unknown","sourceUpdatedAt":null,"freshnessSeconds":null},"decisionGuardrails":{"doNotUseIf":["Contract metadata is missing or unavailable for deterministic execution."],"safeUseWhen":[],"riskFlags":["missing_or_unavailable_contract","trust_data_unavailable","schema_references_missing"],"operationalConfidence":"low"},"executionMetrics":{"observedLatencyMsP50":null,"observedLatencyMsP95":null,"estimatedCostUsd":null,"uptime30d":null,"rateLimitRpm":null,"rateLimitBurst":null,"lastVerifiedAt":null,"verificationSource":null},"runtimeMetrics":{"successRate":null,"avgLatencyMs":null,"avgCostUsd":null,"hallucinationRate":null,"retryRate":null,"disputeRate":null,"p50Latency":null,"p95Latency":null,"lastUpdated":null}},"benchmarks":{"evidence":{"source":"no-benchmark-data","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No benchmark suites or observed failure patterns are available."},"suites":[],"failurePatterns":[]},"artifacts":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T06:37:53.811Z","emptyReason":null},"readme":"Skill: vmware-debug\n\nOwner: zw008\n\nSummary: Use this skill whenever the user is troubleshooting a VMware/vSphere problem — a reported error, an exception, a log dump, a slow or failed VM, a host that went sideways — and needs help locating the root cause. It is the diagnostic brain of the VMware family: it drives a systematic investigation, pulls the right signals from the other skills, correlates events into one timeline, ranks root-cause hypotheses, and tells you what to check next even when you don't know where to start. Always use this skill for \"diagnose this VMware issue\", \"why is my VM slow\", \"troubleshoot this vSphere error\", \"what does this log mean\", \"help me figure out what broke\" when the context is explicitly VMware/vSphere/ESXi/NSX. It never touches vSphere: its only writes are to a local case ledger. Do NOT use it to execute fixes — single fixes go to vmware-aiops, multi-step gated remediation goes to vmware-pilot. Do NOT use it for routine inventory or health checks with no problem to solve — use vmware-monitor.\n\nTags: latest:1.13.1\n\nVersion history:\n\nv1.13.1 | 2026-09-16T05:19:39.714Z | user\n\nExclusion is decided per hypothesis; an empty ledger rules nothing out\n\nv1.13.0 | 2026-09-15T14:41:24.361Z | user\n\nExclusion is per hypothesis: the case is Excluded only when every hypothesis is ruled out; case_grade returns hypotheses.\n\nv1.12.2 | 2026-09-15T08:55:11.605Z | user\n\nEvery MCP tool call is audited, failures included; timeline results not copied into audit rows; documents what debug inherits from vmware-policy. Requires vmware-policy>=1.16.0.\n\nv1.12.1 | 2026-09-15T06:02:04.348Z | user\n\nCLI reads are audited under their MCP tool names; every CLI command declares what it reaches (needs vmware-policy 1.15.0)\n\nv1.12.0 | 2026-09-15T03:13:41.339Z | user\n\nreal alert titles get a category; case_readiness recognises vmware-aiops\n\nv1.11.4 | 2026-09-12T00:20:36.379Z | user\n\nConcurrent writers stop losing case entries: ledger writes take a per-case lock and land atomically, so two agents on one shared OPS_HOME keep every entry. Every machine sharing a cases folder needs this version. The manifests now advertise all 14 tools.\n\nv1.11.3 | 2026-08-31T00:37:34.927Z | user\n\nfix: run the suite on a non-UTF-8 machine, and stop one skill answering for another\n\nv1.11.2 | 2026-08-30T15:20:48.967Z | user\n\nSecond-round fixes from the 2026-08-30 VCF 9.1 re-test; vmware-policy floor raised to 1.11.0 (the engine no longer fails open when rules.yaml cannot be read).\n\nv1.11.1 | 2026-08-30T09:36:10.546Z | user\n\nParameter descriptions now reach the MCP JSON schema (0% -> 100% coverage); additionalProperties closed; vmware-policy floor raised to 1.10.0.\n\nv1.11.0 | 2026-08-30T07:46:56.744Z | user\n\nHost lifecycle events are classified instead of falling through; spike binning chooses by event density, not window length; symptom routing ranks by phrase length; unclassified share is reported instead of absorbed.\n\nv1.10.0 | 2026-08-30T03:22:40.306Z | user\n\nThe knowledge layer is implemented, not just described. case_knowledge lists the six accepted formats (Markdown front-matter, YAML, JSON, JSONL, CSV/TSV, text+sidecar) and names what needs conversion. An entry is decisive only if its own applies_to was checked against the case scope and passed — a constraint the scope cannot answer, or one the checker cannot evaluate, is never a pass.\n\nv1.9.0 | 2026-08-30T01:52:54.818Z | user\n\nThe investigation layer: 11 new tools implementing the eight-step evidence loop. case_grade takes no grade — the level is recomputed from the ledger. Recording a gap is free; only a gap that could overturn a hypothesis holds a case back. On a stock install the ceiling is Probable, and every tool says so.\n\nv1.8.9 | 2026-08-28T02:54:16.737Z | user\n\nFixes the server's self-reported version and the advertised tool count; adds a Claude Code plugin manifest.\n\nv1.8.8 | 2026-08-01T03:10:42.126Z | user\n\nMoved to vmware-skills GitHub org; MCP Registry namespace → io.github.vmware-skills. Links updated.\n\nv1.8.7 | 2026-07-21T11:42:26.725Z | user\n\nRemove read-only switch and approval tiers; read/write authz delegated to RBAC. Plus accumulated fixes since 1.8.5.\n\nv1.8.5 | 2026-07-20T13:08:03.555Z | user\n\nA failure that is returned is now audited as a failure, and certificate/URL detail no longer reaches the agent. Both fixes v1.8.4 announced were incomplete.\n\nv1.8.4 | 2026-07-20T08:28:28.453Z | user\n\nTeaching error messages, domain exceptions no longer redacted on the way to the agent, and tool descriptions that state when to use each tool and what to call next.\n\nv1.8.3 | 2026-07-20T03:54:45.624Z | user\n\nFamily version alignment; the credential env-var override lands in the skills that have per-target credentials\n\nv1.8.2 | 2026-07-19T18:14:34.707Z | user\n\nMCP server moved into the package namespace — fixes two skills in one environment silently overwriting each other's server; agent-guardrails.md for local/small models now ships in every skill\n\nv1.8.1 | 2026-07-19T11:38:00.468Z | user\n\nRead-only mode now documented on every surface that teaches it (SKILL.md, setup-guide, capabilities)\n\nv1.8.0 | 2026-07-19T09:52:01.626Z | user\n\nRead-only mode verified at start-up (2/2 tools read), declared environments; RELEASE_NOTES no longer claims a config.yaml switch this package lacks\n\nv1.6.1 | 2026-06-24T00:01:41.127Z | user\n\nv1.6.1 initial release\n\nArchive index:\n\nArchive v1.13.1: 9 files, 23538 bytes\n\nFiles: references/agent-guardrails.md (9251b), references/capabilities.md (3160b), references/cli-reference.md (1585b), references/event-envelope.md (3622b), references/routing.md (2958b), references/setup-guide.md (4533b), skill-card.md (2841b), SKILL.md (18813b), _meta.json (132b)\n\nFile v1.13.1:SKILL.md\n\n---\nname: vmware-debug\ndescription: >\n  Use this skill whenever the user is troubleshooting a VMware/vSphere problem —\n  a reported error, an exception, a log dump, a slow or failed VM, a host that\n  went sideways — and needs help locating the root cause. It is the diagnostic\n  brain of the VMware family: it drives a systematic investigation, pulls the\n  right signals from the other skills, correlates events into one timeline,\n  ranks root-cause hypotheses, and tells you what to check next even when you\n  don't know where to start. Always use this skill for \"diagnose this VMware\n  issue\", \"why is my VM slow\", \"troubleshoot this vSphere error\", \"what does\n  this log mean\", \"help me figure out what broke\" when the context is explicitly\n  VMware/vSphere/ESXi/NSX. It never touches vSphere: its only writes are to a\n  local case ledger. Do NOT\n  use it to execute fixes — single fixes go to vmware-aiops, multi-step gated\n  remediation goes to vmware-pilot. Do NOT use it for routine inventory or\n  health checks with no problem to solve — use vmware-monitor.\ninstaller:\n  kind: uv\n  package: vmware-debug\nallowed-tools:\n  - Bash\nmetadata: {\"openclaw\":{\"requires\":{\"anyBins\":[\"vmware-debug\",\"uvx\"]},\"optional\":{\"env\":[\"VMWARE_AUDIT_APPROVED_BY\",\"VMWARE_AUDIT_RATIONALE\"],\"bins\":[\"vmware-policy\"]},\"homepage\":\"https://github.com/vmware-skills/VMware-Debug\",\"os\":[\"macos\",\"linux\"]}}\n---\n\n# VMware Debug\n\n> **Disclaimer**: Community-maintained open-source project, **not affiliated with,\n> endorsed by, or sponsored by VMware, Inc. or Broadcom Inc.** \"VMware\" and \"vSphere\"\n> are trademarks of Broadcom. Source is publicly auditable under the MIT license.\n\nThe diagnostic brain of the VMware skill family. You bring the symptom; this skill\nruns the investigation and points at the root cause. It **reads and reasons** — it\nnever writes to vSphere; its only writes go to its own local case ledger.\nCompanion skills do the data collection and the fixing.\n\n## What This Skill Does\n\n| Category | What | Read or Write |\n|---|---|---|\n| Incident correlation | Merge events from many sources into one timeline, detect spikes | Read |\n| Root-cause ranking | Score symptom clusters, surface the most likely cause first | Read |\n| Next-check ideas | Suggest exactly what to look at next (which skill/tool) when you're stuck | Read |\n| Remediation routing | Hand the fix to vmware-aiops (single) or vmware-pilot (gated, multi-step) | Read (routes only) |\n| Investigation ledger | Open a case; record evidence, gaps and hypotheses; grade and close it | Write (local ledger only) |\n\n**No network access of its own, and no write to any VMware system.** It correlates\ndata the agent has already gathered with the other skills' read tools; its seven\nwrite tools touch only the local case ledger.\n\n## Quick Install\n\n```bash\nuv tool install vmware-debug==1.13.1\nvmware-debug categories          # see what it can diagnose\n```\n\n## When to Use This Skill\n\nUse it when there is a **problem to solve**: an error message, a stack of logs, an\nalarm storm, \"my VM won't power on\", \"storage feels slow\", \"the host disconnected\".\n\n- Need raw inventory/health with no incident? → **vmware-monitor**\n- Need to actually run a fix? → **vmware-aiops** (single op) or **vmware-pilot** (gated workflow)\n- Need metrics/anomalies? → **vmware-aria**; centralized logs? → **vmware-log-insight**\n\n**Do NOT use when** there is nothing wrong (routine listing → monitor), or when the\nuser wants the fix executed (→ aiops/pilot). This skill stops at the diagnosis and\na recommended plan.\n\n## Related Skills — Skill Routing\n\n| Symptom touches | Pull signals from | Then |\n|---|---|---|\n| Storage / datastore / vSAN | vmware-storage, vmware-log-insight | rank → route fix to aiops/pilot |\n| Network / firewall / vMotion | vmware-nsx, vmware-nsx-security | run traceflow, check DFW |\n| CPU / memory contention | vmware-aria (metrics/anomalies) | rightsizing via pilot |\n| HA / DRS / cluster | vmware-monitor, vmware-aiops | cluster remediation via pilot |\n| Power / clone / snapshot | vmware-aiops, vmware-monitor | task status, then fix via aiops |\n| Auth / cert / login | check creds & cert; (security) | fix config/.env |\n\n## Common Workflows\n\n### 1. \"Here's a pile of logs / alarms — what broke?\"\n1. Collect events with the data-source skills (e.g. `vmware-monitor event_list --vm web01 --since 1h`, `vmware-log-insight log_search ...`, `vmware-aria alert_query ...`).\n2. Pass them all to **`incident_timeline`** (envelope below). Read the top hypothesis + `next_checks`.\n3. Follow `next_checks` to pull more targeted data; re-run `incident_timeline` to confirm.\n4. **Failure branch — no events come back:** the affected target may be unreachable. Run the source skill's `doctor`/health first; a 503/timeout is a *signal* (platform not ready), not a dead end.\n5. Produce a diagnosis + recommended fix. Route execution to aiops/pilot. **Do not fix here.**\n\n### 2. \"I don't even know what to check\"\n1. Run **`list_symptom_categories`** (or `vmware-debug categories`) to see the catalogue.\n2. Describe the symptom; map it to a category; the `suggested_check` tells you which skill/tool to run first.\n3. Collect → `incident_timeline` → narrow. Loop until one hypothesis dominates.\n\n### 3. Hand off the fix (advisor → executor, like vmware-harden)\n1. Debug emits a structured diagnosis + a proposed remediation (steps).\n2. **Single, low-risk fix** → call the matching **vmware-aiops** tool (it has its own double-confirm).\n3. **Multi-step / needs approval / cross-skill** → submit the plan to **vmware-pilot**, which owns the state machine, approval gate, rollback, and audit.\n4. **Failure branch — fix is ambiguous or risky:** stop and present the hypotheses to the user; never guess-execute.\n\n## Usage Mode\n\n- **MCP** (in an agent): the agent calls the other skills' read tools, then `incident_timeline` to correlate. This is the primary mode — that's where the cross-skill \"联动\" happens.\n- **CLI** (humans): `vmware-debug triage --events events.json` correlates a JSON array you collected yourself.\n\n## MCP Tools (14 — 7 read, 7 write)\n\n**Correlation** — stateless, for a single look:\n\n| Tool | What |\n|---|---|\n| `incident_timeline` | [READ] Correlate pre-fetched events → timeline + spikes + ranked hypotheses + next-check ideas |\n| `list_symptom_categories` | [READ] List recognised symptom categories + what to check for each |\n\n**Investigation ledger** — for an incident you will reason about over time:\n\n| Tool | What |\n|---|---|\n| `case_open` | [WRITE] Define the event; returns a case id and the grade this environment can reach |\n| `case_readiness` | [READ] What grade this environment can reach, per symptom category, **before** you start |\n| `case_knowledge` | [READ] Which knowledge formats are accepted, what is mounted, and which entries apply to a case |\n| `case_plan` | [READ] What to fetch next — skill, tool and purpose per step; recomputed from the case's current state |\n| `case_list` | [READ] Cases, newest first |\n| `case_get` | [READ] One case: scope, ledger sizes, grade history |\n| `case_hypotheses` | [WRITE] Register a candidate explanation, or read the ledger of what supports and refutes each |\n| `case_submit_evidence` | [WRITE] Record one retrieved fact, with its source, query and time basis |\n| `case_record_gap` | [WRITE] Record what could **not** be retrieved, and how to close it |\n| `case_timeline` | [WRITE] Correlate everything the case has collected into one timeline |\n| `case_grade` | [WRITE] Recompute the conclusion grade from the ledger and record it |\n| `case_close` | [WRITE] Record the final grade, archive, and name what was left open |\n\nThe seven writes go to `$OPS_HOME` (default `~/.vmware/cases/`) and nowhere else.\n\n**List envelope** (output of `list_symptom_categories`): `{items, returned, limit, total, truncated, hint}` — read the rows from `items`. `truncated` is always `false` here, which is the point: it states that the catalogue is complete instead of leaving you to infer it.\n\n**Event envelope** (input to `incident_timeline`): `{ts, source, severity, entity, text, fields}`.\nSee `references/event-envelope.md`. The agent normalises each source's events into this\nshape; debug stays source-agnostic and has no dependency on the other packages.\nKeep each event's `event_type` in `fields` — the classifier matches it as well as\nthe message, and on a modern `EventEx` it is the only thing that says what the\nevent was.\n\n## The Investigation Ledger\n\nCorrelating events answers \"what happened together\". A case answers \"what do we\nbelieve, on what evidence, and what is still missing\" — and keeps answering it\nacross sessions and across people.\n\nThe case folder is the deliverable, not an implementation detail. Everything in\nit but the index is plain text, so a customer can take the folder away and audit\nhow a conclusion was reached with none of this installed:\n\n```\n~/.vmware/cases/<case-id>/\n├── scope.json    what is being investigated, and how that was decided\n├── evidence/     one file per fact: source skill, exact query, time basis\n├── gaps.json     what could NOT be obtained, what it blocks, how to close it\n├── conclusion.md the grade, appended — including every time it went down\n└── timeline.md · hypotheses.md · plan.jsonl · case.json\n```\n\nHypotheses get ids (H1, H2, …), and those ids are what `case_record_gap(blocks=…)`\nand `case_submit_evidence(falsifies=…)` refer to. **An id that was never\nregistered is refused, not ignored** — a dangling reference blocks nothing and\nfalsifies nothing, which quietly reports a stronger case than you have.\n\n**You cannot state a conclusion level.** `case_grade` has no parameter for one;\nthe grade is recomputed from the ledger on every call. To change it, change the\nledger — submit the missing evidence, or record the gap that is blocking it.\n\n- **Candidate** — a hypothesis exists\n- **Probable** — ≥2 *independent* sources agree (two calls to one skill are one\n  source) and nothing outstanding could overturn it\n- **Confirmed** — that, plus a decisive item: a direct hardware diagnostic, a\n  version-checked knowledge-base entry, or a vendor SR; and no gap left open\n- **Excluded** — an observation that actually rules it out. \"We looked and found\n  nothing\" is a gap, not an exclusion. Exclusion is per hypothesis: two\n  independent sources must falsify *that* hypothesis, and the case is Excluded\n  only when every registered hypothesis is ruled out. A hypothesis contradicted\n  by too few sources holds the case at Candidate. Otherwise `case_grade`\n  grades what remains, counting only evidence that falsifies nothing, and lists\n  each hypothesis as `open` or `excluded` in `hypotheses`\n\n`case_plan` is not a checklist: submit evidence and the next plan is shorter,\nlose a source and it routes around it. Its `unavailable` half is the important\none — a source this install cannot reach is listed there with how to supply it,\nso the gap is visible now rather than when the conclusion refuses to firm up.\n\n`case_readiness` answers this per symptom category rather than as one number —\n\"storage reaches Probable, hardware reaches Candidate\" can be acted on;\n\"readiness 78%\" cannot. Two classes served by the *same* skill count as one\nsource, so it agrees with what `case_grade` will actually award.\n\n### Mounting a knowledge library\n\n`case_knowledge` answers \"what can I add\" without anyone reading this file.\nEntries go under `$OPS_HOME/knowledge/{kb,runbook,sr,cases}/`:\n\n| Format | How metadata travels |\n|---|---|\n| `.md` `.markdown` | YAML front-matter between `---` fences, body below — **preferred** |\n| `.yaml` `.yml` | the whole file is one entry |\n| `.json` | one entry per file |\n| `.jsonl` | one JSON object per line — what ticketing systems export |\n| `.csv` `.tsv` | one entry per row; use dotted columns (`driver.version`) for nested constraints |\n| `.txt` `.log` | a sibling `<name>.yaml` carries the metadata |\n\nPDF, DOCX, PPTX and HTML must be converted to Markdown first — `case_knowledge`\nnames them rather than ignoring them.\n\n**Every entry needs an `applies_to` block to be decisive:**\n\n```yaml\n---\nid: KB-2026-0417\napplies_to:\n  product: vsphere\n  build: \">=8.0.3, <9.0\"\n  driver: {name: nvme_pcie, version: \">=1.2.4\"}\n  firmware: {vendor: dell, version: \">=52.26\"}\n---\n```\n\n`product`, `build`, `driver` and `firmware` are the constraints the checker\nevaluates. **Any other key leaves the entry non-decisive** — including a typo —\nbecause an unverified constraint is not a satisfied one, and `case_knowledge`\nnames the key it could not check.\n\nKnowledge evidence must say **which** entry it is\n(`case_submit_evidence(source_skill=\"knowledge-kb\", knowledge_entry_id=\"KB-…\")`)\n— \"some applicable entry is mounted somewhere\" is a different claim from \"this\none applies\", and only the second can carry a conclusion.\n\nMatching is **by version applicability, never by similarity** — an entry written\nfor the wrong build reads exactly like the right one, and similarity is the only\nthing that would let it through. An entry with no `applies_to` can support a\nhypothesis but can never make a case Confirmed. A constraint the case scope\ncannot answer is not a match either: silence is not a pass.\n\n> **On a stock install the ceiling is Probable.** Confirmed needs a decisive\n> source, and there is neither a hardware-diagnostic channel (no Redfish/BMC, no\n> SMART/NVMe) nor a knowledge library mounted — `~/.vmware/knowledge/` ships\n> empty. Every tool that can reach the ceiling says so in its output rather than\n> letting you wonder why a well-supported case never goes higher. Mount a\n> knowledge library and the ceiling rises on its own.\n\nRecording a gap is meant to be free: a missing confirmation caps the grade, it\ndoes not demote it. Only a gap that could *overturn* the hypothesis holds a case\nat Candidate.\n\n## No Network, By Design\n\nvmware-debug connects to nothing and holds no credentials. The calling agent\nfetches with the other skills' read tools and submits the results here. Its only\nwrites are to the local case ledger. Running with local or small models? See\n[`references/agent-guardrails.md`](references/agent-guardrails.md).\n\n## CLI Quick Reference\n\n```bash\nvmware-debug categories                        # what can it diagnose\nvmware-debug triage --events events.json       # correlate a collected event set\ncat events.json | vmware-debug triage          # or via stdin\nvmware-debug mcp                                # start stdio MCP server (proxy-safe)\n```\n\n## Troubleshooting\n\n- **`incident_timeline` raises \"event[N] could not be normalised\"** — event N is missing a timestamp or has an unparseable one. Every event needs `ts` (ISO-8601, epoch seconds, or millis).\n- **Most events come back \"uncategorized\"** — read `classification` in the result: it says how many, what share, and quotes the texts that matched nothing. Do **not** widen the window; a wider window adds baseline, not signal. The spikes still tell you *when*, which is answerable without a category. If the samples name a subsystem the taxonomy does not know, add a signature (see `references/routing.md`).\n- **No spikes detected on an obvious burst** — check `binning` for the resolution you were given. The width is chosen from event density so that bins average ≥4 events; a burst shorter than one bin can still be flattened by it. Pass `bin_seconds` to narrow. Below three bins there is no baseline at all and nothing is reported.\n- **`case_timeline` says zero events after you submitted results** — check `payload_events` in each `case_submit_evidence` reply. Events are read from a bare list, or from `items`/`events`/`rows`; a *summary* of a tool's result carries none. `note` names which items carried nothing and what keys they held instead.\n- **`case_readiness` says a skill you have is not installed** — both spellings are accepted (`monitor` and `vmware-monitor`). A name it did not recognise comes back in `unrecognised_skills` rather than being read as missing.\n- **It won't execute the fix** — by design. Route to vmware-aiops or vmware-pilot.\n\n## Audit & Safety\n\nNo network, nothing executed. The seven [WRITE] tools write only to the local case\nledger under `$OPS_HOME`; nothing here touches a remote VMware estate. Evidence, gaps,\nhypotheses and grade history are only ever added to — `timeline.md` is regenerated from\nthe evidence and `case.json` holds the current grade and state. Remediation is always routed to\naiops/pilot, where the double-confirm / approval gates live.\n\n**What the audit row holds.** Every debug tool call — a failed one included — writes one\nrow to `~/.vmware/audit.db` through `@vmware_tool`, as do the `categories` and `triage`\nCLI commands under their MCP tools' names. The row records the tool, time, OS user,\ndetected agent, arguments, status (`ok`, `error`, `denied`, `budget_exceeded`,\n`interrupted`) and a result. Two arguments are stored as `***`: evidence `payload` and\n`incident_timeline` `events`. Two results are not stored at all — `incident_timeline` and\n`case_timeline` quote event text (`sample_text`, `unmatched_samples`, `rejected`), which\ncan carry anything the source logged, so their row says\n`[redacted: return value declared sensitive]`. Every other result (ids, counts, grades,\nyour own summaries and statements) is stored after credential scrubbing. The data itself\nis in the case folder, or with you.\n\n**What debug inherits from vmware-policy**, since its tools pass through `guard()`:\n- **Runaway guard** — the 26th call to one tool with identical arguments within 120 s is\n  refused (`budget_exceeded`; `VMWARE_RUNAWAY_MAX`, `VMWARE_RUNAWAY_WINDOW_SEC`). It\n  compares a digest of the arguments the tool received (vmware-policy ≥1.16.0), so\n  calls with different `events` are not identical.\n- **A malformed `~/.vmware/rules.yaml` fails closed** — every tool is refused with a reason\n  naming the file; `categories`/`triage` exit 1 with a traceback ending in that message.\n- **A `deny` rule without `operations`** matches every tool, debug's included.\n- **Environment-scoped rules never match** — debug registers no resolver (no targets).\n- **Maintenance windows** gate only `high`/`critical` risk; every debug tool is `low`.\n\nSee `references/setup-guide.md`.\n\n**Case data is sensitive and is kept until you delete it.** Each case lives in\n`$OPS_HOME/cases/<case-id>/` (default `~/.vmware/cases/`), created owner-only (`0700`).\nIt holds whatever evidence was submitted — host names, addresses, log and event text.\nNothing is deleted automatically: `case_close` records the grade and marks the case\nclosed but keeps its files. Remove a case by deleting its folder.\n\n## License\n\nMIT.\n\nFile v1.13.1:_meta.json\n\n{\n  \"ownerId\": \"kn7b067awq2s97bn3d7p5qfhw5827pxc\",\n  \"slug\": \"vmware-debug\",\n  \"version\": \"1.13.1\",\n  \"publishedAt\": 1789535979714\n}\n\nFile v1.13.1:references/agent-guardrails.md\n\n# Operating vmware-debug with a local / small model\n\nClaude-class models drive this skill without special instruction. Smaller and\nlocally-hosted models — Llama 3.3 70B, Qwen, Mistral, and similar, served\nthrough Goose, Ollama, or OpenShift AI — need explicit operating rules to call\ntools reliably.\n\nThis page exists because an operator wrote those rules by hand first. The\nguardrails below are adapted, with thanks, from the working configuration\n[@juanpf-ha](https://github.com/juanpf-ha) developed while running\nvmware-monitor and vmware-aria against a production vSphere estate with Llama\n3.3 70B FP8 on an on-prem H100\n([VMware-AIops#31](https://github.com/vmware-skills/VMware-AIops/issues/31)). The\ncross-skill rules are identical across this family; the parts below marked\nvmware-debug are specific to this skill.\n\nvmware-debug exposes 14 MCP tools: 7 reads and 7 writes. It connects to nothing\nand holds no credentials — the calling agent gathers events from the other\nskills, normalises them, and hands them over. The seven writes go to the local\ninvestigation ledger under `~/.vmware/cases/`, never to a VMware system. That\nmakes it the safest skill in the family to point a small model at — and the one\nmost exposed to the model's reasoning, because its output *is* an\ninterpretation.\n\nFor a small model the ledger is more than bookkeeping: `case_grade` computes the\nconclusion level from recorded evidence rather than accepting one, so a model\nthat would happily narrate \"root cause confirmed\" cannot record that unless the\nevidence for it is actually in the folder.\n\n> **Disclaimer**: This is a community-maintained open-source project and is\n> **not affiliated with, endorsed by, or sponsored by VMware, Inc. or Broadcom\n> Inc.** \"VMware\" and \"vSphere\" are trademarks of Broadcom.\n\n---\n\n## First: the rules you no longer need to write\n\nSeveral guardrails from the original configuration are now enforced by the\nskill itself. Prompt instructions are advisory — a model can ignore them.\nThese are structural, so it cannot.\n\n| Guardrail you would otherwise prompt for | Now enforced by |\n|---|---|\n| \"Work read-only and never modify anything\" | **The tool surface itself.** No tool reaches a VMware system. The seven [WRITE] tools write only to this skill's own local case ledger, so there is nothing in vSphere to withhold and nothing to switch off. |\n| \"Diagnose only — never apply the fix you propose\" | **Structural.** This skill has no tool that changes anything outside its own ledger, and it holds no connection to vCenter, NSX or anything else. Remediation is routed to vmware-aiops or vmware-pilot by the calling agent. |\n| \"Do not fabricate a timeline — build it from the events I gave you\" | **`incident_timeline` correlates only its input.** It is source-agnostic and has no way to fetch anything, so the timeline cannot contain an event the agent did not supply. |\n| \"Tell me when the symptom is outside what you can recognise\" | **`list_symptom_categories`** states the catalogue, and unmatched symptoms come back as `uncategorized` rather than being forced into the nearest signature. |\n| \"Use explicit limits for queries that may return large amounts of data\" | **The list envelope.** `list_symptom_categories` returns `{items, returned, limit, total, truncated, hint}` with `truncated` always `false` — which is the point: it states that the catalogue is complete instead of leaving you to infer it. |\n| \"Log everything you looked at\" | **`~/.vmware/audit.db` and the case ledger.** Every debug tool call, a failed one included, writes one audit row (tool, time, user, agent, arguments, status and result — evidence `payload` and `events` are stored as `***`, and the results of `incident_timeline` and `case_timeline`, which quote event text, are not stored at all). Evidence submitted to a case is also recorded in the ledger with its source skill, tool, query and fetch time — so pass the real fetch time, never a placeholder. |\n\n---\n\n## The system prompt\n\nEverything below still benefits from being stated explicitly. Copy this into\nyour agent's instruction block.\n\n```text\n## Tool use\n\n- Gather real events with the companion skills' read tools before calling\n  incident_timeline. Never answer from memory or assumption, and never\n  hand-write events to represent what you believe happened.\n- Never describe a tool call, and never output a JSON example, instead of\n  executing the tool. If you intend to call a tool, call it.\n- If a tool fails, report the actual error text. Do not complete the answer\n  with assumptions about what the result would have been.\n- Use explicit limits and time windows when gathering events. Do not request\n  unlimited results unless the user asks for them.\n\n## Skill routing\n\n- vmware-debug: correlate pre-fetched events into a timeline, detect spikes,\n  rank root-cause hypotheses, suggest next checks.\n- Gather the events from: vmware-monitor (alarms, events, host logs),\n  vmware-log-insight (centralised logs), vmware-aria (metrics, anomalies),\n  vmware-nsx / vmware-nsx-security (network and firewall), vmware-storage.\n- Route the fix to vmware-aiops, or to vmware-pilot when it needs approval\n  gating. This skill executes nothing.\n\n## Building the event set\n\n- Every event needs ts (ISO-8601, epoch seconds, or millis), and the envelope\n  shape {ts, source, severity, entity, text, fields}. An event that cannot be\n  normalised is rejected by index — fix that event, do not drop the batch.\n- Carry each event's original source and severity through unchanged. Do not\n  re-grade a severity to make a story cohere.\n- Pull from more than one source before concluding. A single source's view of\n  an incident is not a correlation.\n- Widen the window before concluding \"no spike\": spike detection needs at least\n  three time bins, so a short window or a large bin_seconds can hide a burst.\n\n## Data fidelity\n\n- Never invent events, timestamps, entities, or relationships between them. If\n  an event was not in the input, it does not exist for this answer.\n- Preserve the exact severity and source values. Do not translate, normalise,\n  or prettify them.\n- Report the timeline in the order the tool returned it.\n- If a requested field was not returned, show it as \"not available\".\n- When a response is long, report every item it contains.\n\n## Analysis discipline\n\n- Separate observed data from interpretation. State which is which. In this\n  skill that separation is the deliverable.\n- Report the ranked hypotheses the tool returned, with its ranking. Do not\n  promote your own preferred explanation above them, and do not present the\n  top hypothesis as a diagnosis.\n- Report next_checks as checks still to be run, not as findings.\n- Do not claim a root cause is confirmed. The tool ranks; it does not conclude.\n- Avoid generic recommendations that are not directly supported by the results.\n```\n\n---\n\n## Known failure modes on small models\n\nObserved with Llama 3.3 70B FP8 (Goose, on-prem H100), and useful as a\nchecklist when evaluating any local model against these skills:\n\n| Symptom | Mitigation |\n|---|---|\n| Describes a tool call, or emits a JSON example, instead of executing it | The \"never describe a tool call\" rule above. This skill is unusually exposed to it: its input *is* JSON, so a model that has just been shown an event envelope will sometimes emit a fabricated one instead of gathering real events. Check that the events came from tool calls. |\n| Long tool responses: omits items, or reports \"no data returned\" when data was present | Ask for explicit limits and narrow windows when gathering. `list_symptom_categories` states `truncated: false`, so a \"no categories\" claim is checkable against `returned`. |\n| Adds generic recommendations unsupported by results | The \"analysis discipline\" rules. Root-cause output attracts invented advice more than any other shape of result in this family. |\n| Drops requested fields or reorders results | State the required fields and ordering in the request itself. A reordered timeline is a different incident. |\n| Multi-tool workflows take 30–50s end to end | Unavoidable in part — this skill's whole premise is a fan-out gather. Narrow each source's window before widening, and prefer the companion skills' aggregate tools when gathering. |\n| Presents the top-ranked hypothesis as the confirmed cause | The \"the tool ranks, it does not conclude\" rule. |\n| Re-grades an event's severity so the narrative fits | The \"carry severity through unchanged\" rule. |\n| Concludes from one source because the others were slow to query | The \"pull from more than one source\" rule. |\n| Reports \"no spike\" from a window too short to have a baseline | Three bins minimum. Widen the window or shrink `bin_seconds` before concluding. |\n| Offers to apply the fix it proposed | It cannot. Route to vmware-aiops or vmware-pilot. |\n\n## Reporting results\n\nLocal-model compatibility is an explicit design constraint for this family, and\nthe evidence base is small. If you evaluate a model against this skill —\nQwen, Mistral, Granite, or anything else — a report of what worked and what did\nnot is genuinely useful:\n[github.com/vmware-skills/VMware-Debug/issues](https://github.com/vmware-skills/VMware-Debug/issues).\n\nFile v1.13.1:references/capabilities.md\n\n# vmware-debug Capabilities\n\nOffline incident correlation. No network, no credentials, no writes to any VMware\nsystem. This table covers the two stateless correlation tools; the twelve `case_*`\ninvestigation-ledger tools (seven of which write, to the local ledger only) are\nlisted in `SKILL.md`, and their response sizes are not yet measured here.\n\n| Tool | What it returns | Typical response tokens |\n|---|---|---|\n| `incident_timeline` | `{event_count, window, spikes:[{start,end,count,zscore}], hypotheses:[{category, score, summary, evidence_count, first_seen, last_seen, sample_text, suggested_check}], next_checks:[...]}` | 300–2000 (scales with hypotheses) |\n| `list_symptom_categories` | `{items: [{category, example_keywords, suggested_check}], returned, limit, total, truncated, hint}` | ~400 |\n\n`list_symptom_categories` returns the family list envelope — read the rows from\n`items`. It has no `limit` parameter, which is exactly why the envelope matters:\n`truncated: false` states that this is every category there is, rather than\nleaving a model to guess whether it is holding page one. The catalogue is a\nfixed in-process constant, so `total` is a real count and `limit` is `null`.\n\n## Correlation engine\n\n- **Timeline**: events normalised to the unified envelope, sorted, and time-binned\n  (auto bin width ≈ span/30, or caller-specified).\n- **Spike detection**: z-score over bin counts (≥3 bins required for a baseline;\n  flat series yields no false spikes).\n- **Hypothesis ranking**: events clustered by symptom category (keyword match on\n  text + entity), scored by summed severity weight, tie-broken by recency.\n  Uncategorised events are kept visible, not dropped.\n- **Next-check routing**: each category carries a concrete \"which skill/tool to run\n  next\" suggestion — the value when the user doesn't know what to check.\n\n## Symptom categories\n\n`storage`, `network`, `compute`, `ha_drs`, `host_lifecycle`, `power_lifecycle`,\n`auth`, `platform`, `hardware`, `licensing`, `data_collection`.\nSee `references/routing.md` for keyword signatures and the skill each routes to.\n\n`hardware`, `licensing` and `data_collection` were added after real alert titles\n(\"Host TPM attestation alarm\", \"License will soon expire\", \"Objects are not\nreceiving data from adapter instance\") matched no category at all. A roll-up\nsuch as \"Group population health is degraded\" is deliberately left\nuncategorized: it names no subsystem, and the cause is in one of its members.\n\n`host_lifecycle` is a host changing its own availability state — maintenance\nmode, shutdown, reboot, standby, connection loss, sync failure.\n`power_lifecycle` is the VM-level equivalent. They are separate because they are\nseparate investigations: the second is a task question for vmware-aiops, the\nfirst is a cluster, DPM, vLCM or drift question.\n\n## Design properties\n\n- **Zero cross-skill runtime deps** — correlation is pure functions over plain\n  dicts; the agent fans out to other skills' read tools (踩坑 #21/#32).\n- **JSON-serialisable output** — suitable for direct MCP responses.\n- **Immutable** — inputs are never mutated; every function returns new values.\n\nFile v1.13.1:references/cli-reference.md\n\n# vmware-debug CLI Reference\n\nAll commands are offline (no network, no credentials). All but `mcp` are\nread-only; `mcp` starts the MCP server, whose seven `case_*` write tools record\ninto the local case ledger.\n\n## triage — correlate a set of collected events\n\n```bash\nvmware-debug triage [OPTIONS]\n  -e, --events PATH     JSON file of event envelopes (reads stdin if omitted)\n      --bin-seconds N   Time-bin width (auto if omitted)\n      --top-n N         Max hypotheses to return   [default: 5]\n```\n\nInput is a JSON array of event envelopes (see `references/event-envelope.md`):\n\n```bash\ncat events.json | vmware-debug triage\nvmware-debug triage --events events.json --top-n 3\n```\n\nOutput (JSON): `{event_count, window, spikes, hypotheses, next_checks}`.\n\n## categories — list recognised symptom categories\n\n```bash\nvmware-debug categories\n```\n\nPrints each category, sample keywords, and the suggested next check (which\nskill/tool to run). Use when you don't know what to look at.\n\n## version / mcp\n\n```bash\nvmware-debug version    # installed version\nvmware-debug mcp        # start the stdio MCP server (no network at startup)\n```\n\n## How the agent uses it\n\nIn an agent, the cross-skill correlation happens at the agent layer:\n\n1. Fetch events with the data-source skills (vmware-monitor `event_list`,\n   vmware-log-insight `log_search`/`log_aggregate`, vmware-aria alerts/anomaly,\n   vmware-nsx).\n2. Normalise each into the event envelope.\n3. Call the `incident_timeline` MCP tool to correlate and rank.\n4. Follow `next_checks`; route any fix to vmware-aiops / vmware-pilot.\n\nFile v1.13.1:references/event-envelope.md\n\n# The Unified Event Envelope\n\nThis is the contract between `vmware-debug` and every data-source skill. The\norchestrating agent fetches events with each skill's own read tools, normalises\neach into this shape, and passes the list to `incident_timeline`. Debug has **no\nruntime dependency** on the other packages (no version lockstep, no heavy install).\n\n## Shape\n\n```json\n{\n  \"ts\":       \"2026-06-23T10:15:30Z\",\n  \"source\":   \"monitor\",\n  \"severity\": \"error\",\n  \"entity\":   \"vm-web01\",\n  \"text\":     \"Device naa.600... performance has deteriorated\",\n  \"fields\":   { \"event_type\": \"esx.problem.scsi.device.io.latency.high\",\n                \"host\": \"esxi-03\", \"datastore\": \"ds1\" }\n}\n```\n\n| Field | Type | Notes |\n|---|---|---|\n| `ts` | string \\| number | ISO-8601, epoch **seconds**, or epoch **millis** (auto-detected). Required. |\n| `source` | string | `monitor` \\| `aria` \\| `loginsight` \\| `nsx` \\| `nsx-security` \\| `storage` \\| ... The catalogue in `rules/evidence_sources.yaml` spells the same skills `vmware-monitor`, `vmware-aria`, `vmware-log-insight`. Both spellings are understood wherever a skill is named — `case_readiness(available_skills=...)`, and the grader's count of independent sources, which treats two spellings of one skill as one source. |\n| `severity` | string | Free text; normalised to `critical`/`error`/`warning`/`info`/`unknown`. |\n| `entity` | string | The object the event is about (VM/host/datastore). May be empty. |\n| `text` | string | Human-readable message. Matched by the symptom classifier. |\n| `fields` | object | Any source-specific extras; preserved, never dropped. `event_type` / `eventTypeId` is matched by the classifier too — see below. |\n\n### `fields.event_type` is worth passing\n\nThe classifier matches the identifier as well as the prose, with camelCase and\ndots split so a two-word keyword can reach it (`HostShutdownEvent` → `host\nshutdown`, `esx.problem.cpu.ready` → `cpu ready`).\n\nFor a classic vCenter event this changes almost nothing — `VmMigratedEvent` and\nits message say the same words. It matters for `EventEx`, which is how modern\nvSphere emits most events: the message is generic boilerplate (\"Issue detected on\nesx01\") and the whole identity is in `eventTypeId`. Drop that field and a storage\nincident arrives as an unreadable event; keep it and it classifies as storage.\n\n`vmware-monitor`'s `get_events` returns it as `event_type`; passing the row\nthrough unchanged is enough.\n\nThe normaliser is tolerant of common field-name variants (e.g. `timestamp`,\n`createTime`, `startTimeUTC` for `ts`; `criticality`, `level` for `severity`;\n`resourceName`, `vm_name`, `fullFormattedMessage` for entity/text), so most\nsources map with little or no adaptation.\n\n## Mapping cheatsheet per source\n\n| Source tool (example) | ts | severity | entity | text |\n|---|---|---|---|---|\n| vmware-monitor `event_list` | `createdTime` | `severity` | `vm`/`host` | `fullFormattedMessage` |\n| vmware-aria `alert_query` | `startTimeUTC` | `criticality` | `resourceName` | `alertDefinitionName` |\n| vmware-aria `anomaly` | `timestamp` | (derive) | `resourceName` | stat + value |\n| vmware-log-insight `log_search` | `timestamp` | `severity`/derive | `hostname` | `text` |\n| vmware-nsx (firewall/traceflow) | `time` | (derive) | src/dst | rule/verdict |\n\n## Why this design\n\n- **Decoupling** — debug never imports monitor/aria/log-insight (CLAUDE.md 踩坑 #21/#32).\n- **Testability** — correlation is pure functions over `Event`; unit tests feed synthetic events.\n- **Transparency** — the cross-skill \"联动\" happens at the agent layer, visibly, not hidden inside debug.\n\nFile v1.13.1:references/routing.md\n\n# Symptom → Signal → Skill Routing\n\nThe catalogue debug uses to turn \"something is wrong\" into concrete next steps.\nKeep this in sync with `_CATEGORY_SIGNATURES` in `vmware_debug/ops/timeline.py`\n(a regression test asserts the two match).\n\n| Category | Keyword signatures (sample) | Pull signals from | Fix via |\n|---|---|---|---|\n| **storage** | datastore, scsi, latency, vsan, apd, pdl, no space, vmfs, iscsi | vmware-storage, vmware-log-insight (vmkernel scsi/apd) | aiops / pilot |\n| **network** | vmotion, uplink, link down, mtu, firewall, dfw, segment, tier-0, bgp | vmware-nsx, vmware-nsx-security (traceflow, DFW hits) | pilot |\n| **compute** | cpu ready, memory, balloon, swap, contention, numa | vmware-aria (metrics + anomalies) | pilot (rightsizing) |\n| **ha_drs** | ha, high availability, drs, failover, admission control, isolation | vmware-monitor, vmware-aiops (cluster) | pilot |\n| **host_lifecycle** | maintenance mode, shut down of, host reboot, standby mode, lost connection to, cannot synchronize, not responding | vmware-monitor (host/cluster state), vmware-harden (drift — a host that left service on cue was told to), vmware-log-insight (vpxd/hostd) | pilot |\n| **power_lifecycle** | power on/off, failed to start, boot, vmx, ovf, clone, snapshot | vmware-aiops (snapshot tree), vmware-monitor | aiops |\n| **auth** | login, authentication, denied, 401, 403, token, certificate, tls, password | config/.env, target cert + time sync | config fix |\n| **platform** | vpxd, hostd, service restart, crash, 503, not responding, disconnected | vmware-monitor (connection/service), vmware-log-insight (vpxd/hostd) | pilot |\n| **hardware** | tpm, attestation, ipmi, sensor, bmc | vmware-monitor (host sensors and the CIM Server that supplies them, host logs for ipmi/cim), the server's BMC | site / vendor |\n| **licensing** | license | vmware-monitor (license_status — which asset holds which key); an Aria alert on \"Unlicensed Group\" is Aria's own license | license portal |\n| **data_collection** | adapter instance, not receiving data, collector group, cloud proxy | vmware-aria (adapters, collector groups, node health, objects that stopped reporting), vmware-monitor (is the source vCenter reachable) | aria admin |\n\n## Remediation handoff (advisor → executor)\n\nDebug never executes. It mirrors the `vmware-harden → vmware-pilot` pattern:\n\n- **Single, low-risk fix** → call the matching **vmware-aiops** tool (own double-confirm).\n- **Multi-step / approval / cross-skill** → submit the proposed plan to **vmware-pilot**:\n  state machine + approval gate + rollback + audit all live there.\n\n## Adding a new symptom category\n\n1. Add a `(name, keywords, suggested_check)` tuple to `_CATEGORY_SIGNATURES`.\n2. Add the matching row to the table above.\n3. Add a focused unit test in `tests/test_timeline.py` (a sample text → expected category).\n4. Add a playbook under `references/playbooks/` if the investigation is non-obvious.\n\nFile v1.13.1:references/setup-guide.md\n\n# vmware-debug Setup Guide\n\nvmware-debug has **no configuration, no credentials, and no network access** — it\nis a pure, offline correlation engine. There is no `config.yaml` and no `.env`.\n\n## Install\n\n```bash\nuv tool install vmware-debug==1.13.1\nvmware-debug categories      # verify it runs\n```\n\n## MCP client configuration\n\n```json\n{\n  \"command\": \"uvx\",\n  \"args\": [\"--from\", \"vmware-debug==1.13.1\", \"vmware-debug-mcp\"]\n}\n```\n\nIf installed with `uv tool install`, prefer the entry point `vmware-debug mcp`\n(no PyPI resolution at startup — robust behind corporate TLS proxies, 踩坑 #25).\n\nFor full cross-skill diagnosis, also install the data-source skills it correlates\n(vmware-monitor, vmware-log-insight, vmware-aria, vmware-nsx) and the executors it\nroutes fixes to (vmware-aiops, vmware-pilot).\n\n## Security\n\n> **Disclaimer**: Community-maintained open-source project, **not affiliated with,\n> endorsed by, or sponsored by VMware, Inc. or Broadcom Inc.**\n\n1. **Source Code** — https://github.com/vmware-skills/VMware-Debug (MIT).\n2. **Credentials** — none. debug holds no secrets and connects to nothing.\n3. **Network** — none. The correlation tools are local pure functions over\n   event data the agent supplies; the case tools read and write local files.\n4. **Writes** — only to the local investigation ledger under `$OPS_HOME`\n   (the seven [WRITE] `case_*` tools). Evidence, gaps, hypotheses and grade\n   history are only ever added to; `timeline.md` is regenerated from the evidence and\n   `case.json` holds the current grade and state. Nothing is written to any\n   VMware system: debug only diagnoses and recommends; remediation is routed\n   to vmware-aiops / vmware-pilot, where confirmation/approval/audit live.\n   Case data is sensitive: each case folder (`$OPS_HOME/cases/<case-id>/`, default\n   `~/.vmware/cases/`) is created owner-only (`0700`) and holds the submitted\n   evidence — host names, addresses, log and event text. There is no automatic\n   retention limit: `case_close` marks a case closed and keeps its files; delete\n   the folder to remove a case.\n5. **No cross-skill coupling** — events arrive as plain dicts (the event\n   envelope); debug imports no other skill package at runtime.\n6. **Environment scoping** — policy rules can scope by environment, and skills\n   that connect to a VMware estate may declare `environment:` (`production` /\n   `staging` / `lab`) per target in their own `config.yaml` as an optional label\n   an environment-scoped `deny` rule can match on. debug has no config and no\n   connection to declare one about, so it registers no environment resolver —\n   it has no basis to answer for any target — and an environment-scoped rule\n   never matches its tools. Every other policy rule does apply: each MCP tool\n   and the `categories` / `triage` CLI commands pass through vmware-policy's\n   `guard()`. Verified behaviour:\n   - A `deny` rule naming a debug tool, or with no `operations` key at all,\n     refuses it (status `denied`, with the rule's reason).\n   - A malformed `~/.vmware/rules.yaml` fails closed: every tool is refused with\n     a reason naming the file and the way out (fix it, or\n     `VMWARE_POLICY_DISABLED=1`). The `categories` and `triage` commands exit 1\n     with a Python traceback whose last lines are that message; `version` and\n     `mcp` are unaffected.\n   - Runaway guard: within one MCP server process, the 26th call to the same\n     tool with identical arguments inside 120 s is refused (`budget_exceeded`).\n     Tune with `VMWARE_RUNAWAY_MAX` / `VMWARE_RUNAWAY_WINDOW_SEC`. The guard\n     compares a SHA-256 digest of the arguments the tool actually received\n     (vmware-policy ≥1.16.0), not the redacted copy in the audit row — so calls\n     with different `events` or `payload` are never counted as identical.\n   - A `maintenance_window` gates only `high`/`critical` risk operations; every\n     debug tool is `low`, so a window never blocks one.\n7. **Audit row contents** — one row per call in `~/.vmware/audit.db`: tool,\n   time, OS user, detected agent, arguments, status and result. `payload` and\n   `events` are stored as `***`. The results of `incident_timeline` and\n   `case_timeline` are not stored (`[redacted: return value declared\n   sensitive]`) because they quote event text; other results are stored after\n   credential scrubbing. The status is still truthful — a call that returns\n   `{\"error\": …}` is `error`.\n8. **Static analysis** — `uvx bandit -r vmware_debug/` (release bar:\n   0 Medium+).\n\nFile v1.13.1:skill-card.md\n\n## Description:\n\nVMware Debug helps agents troubleshoot VMware, vSphere, ESXi, and NSX incidents by correlating supplied events, ranking root-cause hypotheses, suggesting next checks, and recording investigation evidence in a local case ledger.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[zw008](https://clawhub.ai/user/zw008)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers and operations engineers use this skill to diagnose active VMware or vSphere problems, correlate logs and events from companion skills, and decide what data or remediation path to pursue next. It is intended for troubleshooting incidents, not for executing fixes against VMware systems.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: Auth troubleshooting may lead an agent to inspect credential-bearing configuration or .env files.\n\nMitigation: Use narrowly scoped, masked authentication checks and avoid submitting raw secrets as evidence.\n\nRisk: Case folders may retain sensitive host names, addresses, log lines, event text, and investigation notes.\n\nMitigation: Store case data in owner-only locations and periodically delete closed case folders when retention is no longer needed.\n\nRisk: Runtime package resolution can introduce deployment uncertainty in controlled environments.\n\nMitigation: Prefer a verified local install over runtime package resolution where package integrity and repeatability matter.\n\nRisk: Ranked hypotheses can be mistaken for confirmed root cause if an agent overstates the tool output.\n\nMitigation: Treat hypotheses as diagnostic leads, require corroborating evidence from independent sources, and route any fix to the appropriate gated remediation skill.\n\n## Reference(s):\n\n- [ClawHub skill page](https://clawhub.ai/zw008/skills/vmware-debug)\n- [Project homepage](https://github.com/vmware-skills/VMware-Debug)\n- [Agent guardrails](references/agent-guardrails.md)\n- [Capabilities](references/capabilities.md)\n- [CLI reference](references/cli-reference.md)\n- [Event envelope](references/event-envelope.md)\n- [Routing](references/routing.md)\n- [Setup guide](references/setup-guide.md)\n\n## Skill Output:\n\n**Output Type(s):** [text, markdown, JSON, shell commands, configuration, guidance]\n\n**Output Format:** [Markdown guidance with JSON tool and CLI outputs]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [May create local case ledger files under the operator's configured VMware operations directory.]\n\n## Skill Version(s):\n\n1.13.1 (source: server release metadata)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nArchive v1.13.0: 9 files, 23394 bytes\n\nFiles: references/agent-guardrails.md (9251b), references/capabilities.md (3160b), references/cli-reference.md (1585b), references/event-envelope.md (3622b), references/routing.md (2958b), references/setup-guide.md (4533b), skill-card.md (2653b), SKILL.md (18675b), _meta.json (132b)\n\nFile v1.13.0:SKILL.md\n\n---\nname: vmware-debug\ndescription: >\n  Use this skill whenever the user is troubleshooting a VMware/vSphere problem —\n  a reported error, an exception, a log dump, a slow or failed VM, a host that\n  went sideways — and needs help locating the root cause. It is the diagnostic\n  brain of the VMware family: it drives a systematic investigation, pulls the\n  right signals from the other skills, correlates events into one timeline,\n  ranks root-cause hypotheses, and tells you what to check next even when you\n  don't know where to start. Always use this skill for \"diagnose this VMware\n  issue\", \"why is my VM slow\", \"troubleshoot this vSphere error\", \"what does\n  this log mean\", \"help me figure out what broke\" when the context is explicitly\n  VMware/vSphere/ESXi/NSX. It never touches vSphere: its only writes are to a\n  local case ledger. Do NOT\n  use it to execute fixes — single fixes go to vmware-aiops, multi-step gated\n  remediation goes to vmware-pilot. Do NOT use it for routine inventory or\n  health checks with no problem to solve — use vmware-monitor.\ninstaller:\n  kind: uv\n  package: vmware-debug\nallowed-tools:\n  - Bash\nmetadata: {\"openclaw\":{\"requires\":{\"anyBins\":[\"vmware-debug\",\"uvx\"]},\"optional\":{\"env\":[\"VMWARE_AUDIT_APPROVED_BY\",\"VMWARE_AUDIT_RATIONALE\"],\"bins\":[\"vmware-policy\"]},\"homepage\":\"https://github.com/vmware-skills/VMware-Debug\",\"os\":[\"macos\",\"linux\"]}}\n---\n\n# VMware Debug\n\n> **Disclaimer**: Community-maintained open-source project, **not affiliated with,\n> endorsed by, or sponsored by VMware, Inc. or Broadcom Inc.** \"VMware\" and \"vSphere\"\n> are trademarks of Broadcom. Source is publicly auditable under the MIT license.\n\nThe diagnostic brain of the VMware skill family. You bring the symptom; this skill\nruns the investigation and points at the root cause. It **reads and reasons** — it\nnever writes to vSphere; its only writes go to its own local case ledger.\nCompanion skills do the data collection and the fixing.\n\n## What This Skill Does\n\n| Category | What | Read or Write |\n|---|---|---|\n| Incident correlation | Merge events from many sources into one timeline, detect spikes | Read |\n| Root-cause ranking | Score symptom clusters, surface the most likely cause first | Read |\n| Next-check ideas | Suggest exactly what to look at next (which skill/tool) when you're stuck | Read |\n| Remediation routing | Hand the fix to vmware-aiops (single) or vmware-pilot (gated, multi-step) | Read (routes only) |\n| Investigation ledger | Open a case; record evidence, gaps and hypotheses; grade and close it | Write (local ledger only) |\n\n**No network access of its own, and no write to any VMware system.** It correlates\ndata the agent has already gathered with the other skills' read tools; its seven\nwrite tools touch only the local case ledger.\n\n## Quick Install\n\n```bash\nuv tool install vmware-debug==1.13.0\nvmware-debug categories          # see what it can diagnose\n```\n\n## When to Use This Skill\n\nUse it when there is a **problem to solve**: an error message, a stack of logs, an\nalarm storm, \"my VM won't power on\", \"storage feels slow\", \"the host disconnected\".\n\n- Need raw inventory/health with no incident? → **vmware-monitor**\n- Need to actually run a fix? → **vmware-aiops** (single op) or **vmware-pilot** (gated workflow)\n- Need metrics/anomalies? → **vmware-aria**; centralized logs? → **vmware-log-insight**\n\n**Do NOT use when** there is nothing wrong (routine listing → monitor), or when the\nuser wants the fix executed (→ aiops/pilot). This skill stops at the diagnosis and\na recommended plan.\n\n## Related Skills — Skill Routing\n\n| Symptom touches | Pull signals from | Then |\n|---|---|---|\n| Storage / datastore / vSAN | vmware-storage, vmware-log-insight | rank → route fix to aiops/pilot |\n| Network / firewall / vMotion | vmware-nsx, vmware-nsx-security | run traceflow, check DFW |\n| CPU / memory contention | vmware-aria (metrics/anomalies) | rightsizing via pilot |\n| HA / DRS / cluster | vmware-monitor, vmware-aiops | cluster remediation via pilot |\n| Power / clone / snapshot | vmware-aiops, vmware-monitor | task status, then fix via aiops |\n| Auth / cert / login | check creds & cert; (security) | fix config/.env |\n\n## Common Workflows\n\n### 1. \"Here's a pile of logs / alarms — what broke?\"\n1. Collect events with the data-source skills (e.g. `vmware-monitor event_list --vm web01 --since 1h`, `vmware-log-insight log_search ...`, `vmware-aria alert_query ...`).\n2. Pass them all to **`incident_timeline`** (envelope below). Read the top hypothesis + `next_checks`.\n3. Follow `next_checks` to pull more targeted data; re-run `incident_timeline` to confirm.\n4. **Failure branch — no events come back:** the affected target may be unreachable. Run the source skill's `doctor`/health first; a 503/timeout is a *signal* (platform not ready), not a dead end.\n5. Produce a diagnosis + recommended fix. Route execution to aiops/pilot. **Do not fix here.**\n\n### 2. \"I don't even know what to check\"\n1. Run **`list_symptom_categories`** (or `vmware-debug categories`) to see the catalogue.\n2. Describe the symptom; map it to a category; the `suggested_check` tells you which skill/tool to run first.\n3. Collect → `incident_timeline` → narrow. Loop until one hypothesis dominates.\n\n### 3. Hand off the fix (advisor → executor, like vmware-harden)\n1. Debug emits a structured diagnosis + a proposed remediation (steps).\n2. **Single, low-risk fix** → call the matching **vmware-aiops** tool (it has its own double-confirm).\n3. **Multi-step / needs approval / cross-skill** → submit the plan to **vmware-pilot**, which owns the state machine, approval gate, rollback, and audit.\n4. **Failure branch — fix is ambiguous or risky:** stop and present the hypotheses to the user; never guess-execute.\n\n## Usage Mode\n\n- **MCP** (in an agent): the agent calls the other skills' read tools, then `incident_timeline` to correlate. This is the primary mode — that's where the cross-skill \"联动\" happens.\n- **CLI** (humans): `vmware-debug triage --events events.json` correlates a JSON array you collected yourself.\n\n## MCP Tools (14 — 7 read, 7 write)\n\n**Correlation** — stateless, for a single look:\n\n| Tool | What |\n|---|---|\n| `incident_timeline` | [READ] Correlate pre-fetched events → timeline + spikes + ranked hypotheses + next-check ideas |\n| `list_symptom_categories` | [READ] List recognised symptom categories + what to check for each |\n\n**Investigation ledger** — for an incident you will reason about over time:\n\n| Tool | What |\n|---|---|\n| `case_open` | [WRITE] Define the event; returns a case id and the grade this environment can reach |\n| `case_readiness` | [READ] What grade this environment can reach, per symptom category, **before** you start |\n| `case_knowledge` | [READ] Which knowledge formats are accepted, what is mounted, and which entries apply to a case |\n| `case_plan` | [READ] What to fetch next — skill, tool and purpose per step; recomputed from the case's current state |\n| `case_list` | [READ] Cases, newest first |\n| `case_get` | [READ] One case: scope, ledger sizes, grade history |\n| `case_hypotheses` | [WRITE] Register a candidate explanation, or read the ledger of what supports and refutes each |\n| `case_submit_evidence` | [WRITE] Record one retrieved fact, with its source, query and time basis |\n| `case_record_gap` | [WRITE] Record what could **not** be retrieved, and how to close it |\n| `case_timeline` | [WRITE] Correlate everything the case has collected into one timeline |\n| `case_grade` | [WRITE] Recompute the conclusion grade from the ledger and record it |\n| `case_close` | [WRITE] Record the final grade, archive, and name what was left open |\n\nThe seven writes go to `$OPS_HOME` (default `~/.vmware/cases/`) and nowhere else.\n\n**List envelope** (output of `list_symptom_categories`): `{items, returned, limit, total, truncated, hint}` — read the rows from `items`. `truncated` is always `false` here, which is the point: it states that the catalogue is complete instead of leaving you to infer it.\n\n**Event envelope** (input to `incident_timeline`): `{ts, source, severity, entity, text, fields}`.\nSee `references/event-envelope.md`. The agent normalises each source's events into this\nshape; debug stays source-agnostic and has no dependency on the other packages.\nKeep each event's `event_type` in `fields` — the classifier matches it as well as\nthe message, and on a modern `EventEx` it is the only thing that says what the\nevent was.\n\n## The Investigation Ledger\n\nCorrelating events answers \"what happened together\". A case answers \"what do we\nbelieve, on what evidence, and what is still missing\" — and keeps answering it\nacross sessions and across people.\n\nThe case folder is the deliverable, not an implementation detail. Everything in\nit but the index is plain text, so a customer can take the folder away and audit\nhow a conclusion was reached with none of this installed:\n\n```\n~/.vmware/cases/<case-id>/\n├── scope.json    what is being investigated, and how that was decided\n├── evidence/     one file per fact: source skill, exact query, time basis\n├── gaps.json     what could NOT be obtained, what it blocks, how to close it\n├── conclusion.md the grade, appended — including every time it went down\n└── timeline.md · hypotheses.md · plan.jsonl · case.json\n```\n\nHypotheses get ids (H1, H2, …), and those ids are what `case_record_gap(blocks=…)`\nand `case_submit_evidence(falsifies=…)` refer to. **An id that was never\nregistered is refused, not ignored** — a dangling reference blocks nothing and\nfalsifies nothing, which quietly reports a stronger case than you have.\n\n**You cannot state a conclusion level.** `case_grade` has no parameter for one;\nthe grade is recomputed from the ledger on every call. To change it, change the\nledger — submit the missing evidence, or record the gap that is blocking it.\n\n- **Candidate** — a hypothesis exists\n- **Probable** — ≥2 *independent* sources agree (two calls to one skill are one\n  source) and nothing outstanding could overturn it\n- **Confirmed** — that, plus a decisive item: a direct hardware diagnostic, a\n  version-checked knowledge-base entry, or a vendor SR; and no gap left open\n- **Excluded** — an observation that actually rules it out. \"We looked and found\n  nothing\" is a gap, not an exclusion. Exclusion is per hypothesis: the case is\n  Excluded only when every registered hypothesis is ruled out. Otherwise\n  `case_grade` grades what remains, counting only evidence that rules nothing\n  out, and lists each hypothesis as `open` or `excluded` in `hypotheses`\n\n`case_plan` is not a checklist: submit evidence and the next plan is shorter,\nlose a source and it routes around it. Its `unavailable` half is the important\none — a source this install cannot reach is listed there with how to supply it,\nso the gap is visible now rather than when the conclusion refuses to firm up.\n\n`case_readiness` answers this per symptom category rather than as one number —\n\"storage reaches Probable, hardware reaches Candidate\" can be acted on;\n\"readiness 78%\" cannot. Two classes served by the *same* skill count as one\nsource, so it agrees with what `case_grade` will actually award.\n\n### Mounting a knowledge library\n\n`case_knowledge` answers \"what can I add\" without anyone reading this file.\nEntries go under `$OPS_HOME/knowledge/{kb,runbook,sr,cases}/`:\n\n| Format | How metadata travels |\n|---|---|\n| `.md` `.markdown` | YAML front-matter between `---` fences, body below — **preferred** |\n| `.yaml` `.yml` | the whole file is one entry |\n| `.json` | one entry per file |\n| `.jsonl` | one JSON object per line — what ticketing systems export |\n| `.csv` `.tsv` | one entry per row; use dotted columns (`driver.version`) for nested constraints |\n| `.txt` `.log` | a sibling `<name>.yaml` carries the metadata |\n\nPDF, DOCX, PPTX and HTML must be converted to Markdown first — `case_knowledge`\nnames them rather than ignoring them.\n\n**Every entry needs an `applies_to` block to be decisive:**\n\n```yaml\n---\nid: KB-2026-0417\napplies_to:\n  product: vsphere\n  build: \">=8.0.3, <9.0\"\n  driver: {name: nvme_pcie, version: \">=1.2.4\"}\n  firmware: {vendor: dell, version: \">=52.26\"}\n---\n```\n\n`product`, `build`, `driver` and `firmware` are the constraints the checker\nevaluates. **Any other key leaves the entry non-decisive** — including a typo —\nbecause an unverified constraint is not a satisfied one, and `case_knowledge`\nnames the key it could not check.\n\nKnowledge evidence must say **which** entry it is\n(`case_submit_evidence(source_skill=\"knowledge-kb\", knowledge_entry_id=\"KB-…\")`)\n— \"some applicable entry is mounted somewhere\" is a different claim from \"this\none applies\", and only the second can carry a conclusion.\n\nMatching is **by version applicability, never by similarity** — an entry written\nfor the wrong build reads exactly like the right one, and similarity is the only\nthing that would let it through. An entry with no `applies_to` can support a\nhypothesis but can never make a case Confirmed. A constraint the case scope\ncannot answer is not a match either: silence is not a pass.\n\n> **On a stock install the ceiling is Probable.** Confirmed needs a decisive\n> source, and there is neither a hardware-diagnostic channel (no Redfish/BMC, no\n> SMART/NVMe) nor a knowledge library mounted — `~/.vmware/knowledge/` ships\n> empty. Every tool that can reach the ceiling says so in its output rather than\n> letting you wonder why a well-supported case never goes higher. Mount a\n> knowledge library and the ceiling rises on its own.\n\nRecording a gap is meant to be free: a missing confirmation caps the grade, it\ndoes not demote it. Only a gap that could *overturn* the hypothesis holds a case\nat Candidate.\n\n## No Network, By Design\n\nvmware-debug connects to nothing and holds no credentials. The calling agent\nfetches with the other skills' read tools and submits the results here. Its only\nwrites are to the local case ledger. Running with local or small models? See\n[`references/agent-guardrails.md`](references/agent-guardrails.md).\n\n## CLI Quick Reference\n\n```bash\nvmware-debug categories                        # what can it diagnose\nvmware-debug triage --events events.json       # correlate a collected event set\ncat events.json | vmware-debug triage          # or via stdin\nvmware-debug mcp                                # start stdio MCP server (proxy-safe)\n```\n\n## Troubleshooting\n\n- **`incident_timeline` raises \"event[N] could not be normalised\"** — event N is missing a timestamp or has an unparseable one. Every event needs `ts` (ISO-8601, epoch seconds, or millis).\n- **Most events come back \"uncategorized\"** — read `classification` in the result: it says how many, what share, and quotes the texts that matched nothing. Do **not** widen the window; a wider window adds baseline, not signal. The spikes still tell you *when*, which is answerable without a category. If the samples name a subsystem the taxonomy does not know, add a signature (see `references/routing.md`).\n- **No spikes detected on an obvious burst** — check `binning` for the resolution you were given. The width is chosen from event density so that bins average ≥4 events; a burst shorter than one bin can still be flattened by it. Pass `bin_seconds` to narrow. Below three bins there is no baseline at all and nothing is reported.\n- **`case_timeline` says zero events after you submitted results** — check `payload_events` in each `case_submit_evidence` reply. Events are read from a bare list, or from `items`/`events`/`rows`; a *summary* of a tool's result carries none. `note` names which items carried nothing and what keys they held instead.\n- **`case_readiness` says a skill you have is not installed** — both spellings are accepted (`monitor` and `vmware-monitor`). A name it did not recognise comes back in `unrecognised_skills` rather than being read as missing.\n- **It won't execute the fix** — by design. Route to vmware-aiops or vmware-pilot.\n\n## Audit & Safety\n\nNo network, nothing executed. The seven [WRITE] tools write only to the local case\nledger under `$OPS_HOME`; nothing here touches a remote VMware estate. Evidence, gaps,\nhypotheses and grade history are only ever added to — `timeline.md` is regenerated from\nthe evidence and `case.json` holds the current grade and state. Remediation is always routed to\naiops/pilot, where the double-confirm / approval gates live.\n\n**What the audit row holds.** Every debug tool call — a failed one included — writes one\nrow to `~/.vmware/audit.db` through `@vmware_tool`, as do the `categories` and `triage`\nCLI commands under their MCP tools' names. The row records the tool, time, OS user,\ndetected agent, arguments, status (`ok`, `error`, `denied`, `budget_exceeded`,\n`interrupted`) and a result. Two arguments are stored as `***`: evidence `payload` and\n`incident_timeline` `events`. Two results are not stored at all — `incident_timeline` and\n`case_timeline` quote event text (`sample_text`, `unmatched_samples`, `rejected`), which\ncan carry anything the source logged, so their row says\n`[redacted: return value declared sensitive]`. Every other result (ids, counts, grades,\nyour own summaries and statements) is stored after credential scrubbing. The data itself\nis in the case folder, or with you.\n\n**What debug inherits from vmware-policy**, since its tools pass through `guard()`:\n- **Runaway guard** — the 26th call to one tool with identical arguments within 120 s is\n  refused (`budget_exceeded`; `VMWARE_RUNAWAY_MAX`, `VMWARE_RUNAWAY_WINDOW_SEC`). It\n  compares a digest of the arguments the tool received (vmware-policy ≥1.16.0), so\n  calls with different `events` are not identical.\n- **A malformed `~/.vmware/rules.yaml` fails closed** — every tool is refused with a reason\n  naming the file; `categories`/`triage` exit 1 with a traceback ending in that message.\n- **A `deny` rule without `operations`** matches every tool, debug's included.\n- **Environment-scoped rules never match** — debug registers no resolver (no targets).\n- **Maintenance windows** gate only `high`/`critical` risk; every debug tool is `low`.\n\nSee `references/setup-guide.md`.\n\n**Case data is sensitive and is kept until you delete it.** Each case lives in\n`$OPS_HOME/cases/<case-id>/` (default `~/.vmware/cases/`), created owner-only (`0700`).\nIt holds whatever evidence was submitted — host names, addresses, log and event text.\nNothing is deleted automatically: `case_close` records the grade and marks the case\nclosed but keeps its files. Remove a case by deleting its folder.\n\n## License\n\nMIT.\n\nFile v1.13.0:_meta.json\n\n{\n  \"ownerId\": \"kn7b067awq2s97bn3d7p5qfhw5827pxc\",\n  \"slug\": \"vmware-debug\",\n  \"version\": \"1.13.0\",\n  \"publishedAt\": 1789483284361\n}\n\nFile v1.13.0:references/agent-guardrails.md\n\n# Operating vmware-debug with a local / small model\n\nClaude-class models drive this skill without special instruction. Smaller and\nlocally-hosted models — Llama 3.3 70B, Qwen, Mistral, and similar, served\nthrough Goose, Ollama, or OpenShift AI — need explicit operating rules to call\ntools reliably.\n\nThis page exists because an operator wrote those rules by hand first. The\nguardrails below are adapted, with thanks, from the working configuration\n[@juanpf-ha](https://github.com/juanpf-ha) developed while running\nvmware-monitor and vmware-aria against a production vSphere estate with Llama\n3.3 70B FP8 on an on-prem H100\n([VMware-AIops#31](https://github.com/vmware-skills/VMware-AIops/issues/31)). The\ncross-skill rules are identical across this family; the parts below marked\nvmware-debug are specific to this skill.\n\nvmware-debug exposes 14 MCP tools: 7 reads and 7 writes. It connects to nothing\nand holds no credentials — the calling agent gathers events from the other\nskills, normalises them, and hands them over. The seven writes go to the local\ninvestigation ledger under `~/.vmware/cases/`, never to a VMware system. That\nmakes it the safest skill in the family to point a small model at — and the one\nmost exposed to the model's reasoning, because its output *is* an\ninterpretation.\n\nFor a small model the ledger is more than bookkeeping: `case_grade` computes the\nconclusion level from recorded evidence rather than accepting one, so a model\nthat would happily narrate \"root cause confirmed\" cannot record that unless the\nevidence for it is actually in the folder.\n\n> **Disclaimer**: This is a community-maintained open-source project and is\n> **not affiliated with, endorsed by, or sponsored by VMware, Inc. or Broadcom\n> Inc.** \"VMware\" and \"vSphere\" are trademarks of Broadcom.\n\n---\n\n## First: the rules you no longer need to write\n\nSeveral guardrails from the original configuration are now enforced by the\nskill itself. Prompt instructions are advisory — a model can ignore them.\nThese are structural, so it cannot.\n\n| Guardrail you would otherwise prompt for | Now enforced by |\n|---|---|\n| \"Work read-only and never modify anything\" | **The tool surface itself.** No tool reaches a VMware system. The seven [WRITE] tools write only to this skill's own local case ledger, so there is nothing in vSphere to withhold and nothing to switch off. |\n| \"Diagnose only — never apply the fix you propose\" | **Structural.** This skill has no tool that changes anything outside its own ledger, and it holds no connection to vCenter, NSX or anything else. Remediation is routed to vmware-aiops or vmware-pilot by the calling agent. |\n| \"Do not fabricate a timeline — build it from the events I gave you\" | **`incident_timeline` correlates only its input.** It is source-agnostic and has no way to fetch anything, so the timeline cannot contain an event the agent did not supply. |\n| \"Tell me when the symptom is outside what you can recognise\" | **`list_symptom_categories`** states the catalogue, and unmatched symptoms come back as `uncategorized` rather than being forced into the nearest signature. |\n| \"Use explicit limits for queries that may return large amounts of data\" | **The list envelope.** `list_symptom_categories` returns `{items, returned, limit, total, truncated, hint}` with `truncated` always `false` — which is the point: it states that the catalogue is complete instead of leaving you to infer it. |\n| \"Log everything you looked at\" | **`~/.vmware/audit.db` and the case ledger.** Every debug tool call, a failed one included, writes one audit row (tool, time, user, agent, arguments, status and result — evidence `payload` and `events` are stored as `***`, and the results of `incident_timeline` and `case_timeline`, which quote event text, are not stored at all). Evidence submitted to a case is also recorded in the ledger with its source skill, tool, query and fetch time — so pass the real fetch time, never a placeholder. |\n\n---\n\n## The system prompt\n\nEverything below still benefits from being stated explicitly. Copy this into\nyour agent's instruction block.\n\n```text\n## Tool use\n\n- Gather real events with the companion skills' read tools before calling\n  incident_timeline. Never answer from memory or assumption, and never\n  hand-write events to represent what you believe happened.\n- Never describe a tool call, and never output a JSON example, instead of\n  executing the tool. If you intend to call a tool, call it.\n- If a tool fails, report the actual error text. Do not complete the answer\n  with assumptions about what the result would have been.\n- Use explicit limits and time windows when gathering events. Do not request\n  unlimited results unless the user asks for them.\n\n## Skill routing\n\n- vmware-debug: correlate pre-fetched events into a timeline, detect spikes,\n  rank root-cause hypotheses, suggest next checks.\n- Gather the events from: vmware-monitor (alarms, events, host logs),\n  vmware-log-insight (centralised logs), vmware-aria (metrics, anomalies),\n  vmware-nsx / vmware-nsx-security (network and firewall), vmware-storage.\n- Route the fix to vmware-aiops, or to vmware-pilot when it needs approval\n  gating. This skill executes nothing.\n\n## Building the event set\n\n- Every event needs ts (ISO-8601, epoch seconds, or millis), and the envelope\n  shape {ts, source, severity, entity, text, fields}. An event that cannot be\n  normalised is rejected by index — fix that event, do not drop the batch.\n- Carry each event's original source and severity through unchanged. Do not\n  re-grade a severity to make a story cohere.\n- Pull from more than one source before concluding. A single source's view of\n  an incident is not a correlation.\n- Widen the window before concluding \"no spike\": spike detection needs at least\n  three time bins, so a short window or a large bin_seconds can hide a burst.\n\n## Data fidelity\n\n- Never invent events, timestamps, entities, or relationships between them. If\n  an event was not in the input, it does not exist for this answer.\n- Preserve the exact severity and source values. Do not translate, normalise,\n  or prettify them.\n- Report the timeline in the order the tool returned it.\n- If a requested field was not returned, show it as \"not available\".\n- When a response is long, report every item it contains.\n\n## Analysis discipline\n\n- Separate observed data from interpretation. State which is which. In this\n  skill that separation is the deliverable.\n- Report the ranked hypotheses the tool returned, with its ranking. Do not\n  promote your own preferred explanation above them, and do not present the\n  top hypothesis as a diagnosis.\n- Report next_checks as checks still to be run, not as findings.\n- Do not claim a root cause is confirmed. The tool ranks; it does not conclude.\n- Avoid generic recommendations that are not directly supported by the results.\n```\n\n---\n\n## Known failure modes on small models\n\nObserved with Llama 3.3 70B FP8 (Goose, on-prem H100), and useful as a\nchecklist when evaluating any local model against these skills:\n\n| Symptom | Mitigation |\n|---|---|\n| Describes a tool call, or emits a JSON example, instead of executing it | The \"never describe a tool call\" rule above. This skill is unusually exposed to it: its input *is* JSON, so a model that has just been shown an event envelope will sometimes emit a fabricated one instead of gathering real events. Check that the events came from tool calls. |\n| Long tool responses: omits items, or reports \"no data returned\" when data was present | Ask for explicit limits and narrow windows when gathering. `list_symptom_categories` states `truncated: false`, so a \"no categories\" claim is checkable against `returned`. |\n| Adds generic recommendations unsupported by results | The \"analysis discipline\" rules. Root-cause output attracts invented advice more than any other shape of result in this family. |\n| Drops requested fields or reorders results | State the required fields and ordering in the request itself. A reordered timeline is a different incident. |\n| Multi-tool workflows take 30–50s end to end | Unavoidable in part — this skill's whole premise is a fan-out gather. Narrow each source's window before widening, and prefer the companion skills' aggregate tools when gathering. |\n| Presents the top-ranked hypothesis as the confirmed cause | The \"the tool ranks, it does not conclude\" rule. |\n| Re-grades an event's severity so the narrative fits | The \"carry severity through unchanged\" rule. |\n| Concludes from one source because the others were slow to query | The \"pull from more than one source\" rule. |\n| Reports \"no spike\" from a window too short to have a baseline | Three bins minimum. Widen the window or shrink `bin_seconds` before concluding. |\n| Offers to apply the fix it proposed | It cannot. Route to vmware-aiops or vmware-pilot. |\n\n## Reporting results\n\nLocal-model compatibility is an explicit design constraint for this family, and\nthe evidence base is small. If you evaluate a model against this skill —\nQwen, Mistral, Granite, or anything else — a report of what worked and what did\nnot is genuinely useful:\n[github.com/vmware-skills/VMware-Debug/issues](https://github.com/vmware-skills/VMware-Debug/issues).\n\nFile v1.13.0:references/capabilities.md\n\n# vmware-debug Capabilities\n\nOffline incident correlation. No network, no credentials, no writes to any VMware\nsystem. This table covers the two stateless correlation tools; the twelve `case_*`\ninvestigation-ledger tools (seven of which write, to the local ledger only) are\nlisted in `SKILL.md`, and their response sizes are not yet measured here.\n\n| Tool | What it returns | Typical response tokens |\n|---|---|---|\n| `incident_timeline` | `{event_count, window, spikes:[{start,end,count,zscore}], hypotheses:[{category, score, summary, evidence_count, first_seen, last_seen, sample_text, suggested_check}], next_checks:[...]}` | 300–2000 (scales with hypotheses) |\n| `list_symptom_categories` | `{items: [{category, example_keywords, suggested_check}], returned, limit, total, truncated, hint}` | ~400 |\n\n`list_symptom_categories` returns the family list envelope — read the rows from\n`items`. It has no `limit` parameter, which is exactly why the envelope matters:\n`truncated: false` states that this is every category there is, rather than\nleaving a model to guess whether it is holding page one. The catalogue is a\nfixed in-process constant, so `total` is a real count and `limit` is `null`.\n\n## Correlation engine\n\n- **Timeline**: events normalised to the unified envelope, sorted, and time-binned\n  (auto bin width ≈ span/30, or caller-specified).\n- **Spike detection**: z-score over bin counts (≥3 bins required for a baseline;\n  flat series yields no false spikes).\n- **Hypothesis ranking**: events clustered by symptom category (keyword match on\n  text + entity), scored by summed severity weight, tie-broken by recency.\n  Uncategorised events are kept visible, not dropped.\n- **Next-check routing**: each category carries a concrete \"which skill/tool to run\n  next\" suggestion — the value when the user doesn't know what to check.\n\n## Symptom categories\n\n`storage`, `network`, `compute`, `ha_drs`, `host_lifecycle`, `power_lifecycle`,\n`auth`, `platform`, `hardware`, `licensing`, `data_collection`.\nSee `references/routing.md` for keyword signatures and the skill each routes to.\n\n`hardware`, `licensing` and `data_collection` were added after real alert titles\n(\"Host TPM attestation alarm\", \"License will soon expire\", \"Objects are not\nreceiving data from adapter instance\") matched no category at all. A roll-up\nsuch as \"Group population health is degraded\" is deliberately left\nuncategorized: it names no subsystem, and the cause is in one of its members.\n\n`host_lifecycle` is a host changing its own availability state — maintenance\nmode, shutdown, reboot, standby, connection loss, sync failure.\n`power_lifecycle` is the VM-level equivalent. They are separate because they are\nseparate investigations: the second is a task question for vmware-aiops, the\nfirst is a cluster, DPM, vLCM or drift question.\n\n## Design properties\n\n- **Zero cross-skill runtime deps** — correlation is pure functions over plain\n  dicts; the agent fans out to other skills' read tools (踩坑 #21/#32).\n- **JSON-serialisable output** — suitable for direct MCP responses.\n- **Immutable** — inputs are never mutated; every function returns new values.\n\nFile v1.13.0:references/cli-reference.md\n\n# vmware-debug CLI Reference\n\nAll commands are offline (no network, no credentials). All but `mcp` are\nread-only; `mcp` starts the MCP server, whose seven `case_*` write tools record\ninto the local case ledger.\n\n## triage — correlate a set of collected events\n\n```bash\nvmware-debug triage [OPTIONS]\n  -e, --events PATH     JSON file of event envelopes (reads stdin if omitted)\n      --bin-seconds N   Time-bin width (auto if omitted)\n      --top-n N         Max hypotheses to return   [default: 5]\n```\n\nInput is a JSON array of event envelopes (see `references/event-envelope.md`):\n\n```bash\ncat events.json | vmware-debug triage\nvmware-debug triage --events events.json --top-n 3\n```\n\nOutput (JSON): `{event_count, window, spikes, hypotheses, next_checks}`.\n\n## categories — list recognised symptom categories\n\n```bash\nvmware-debug categories\n```\n\nPrints each category, sample keywords, and the suggested next check (which\nskill/tool to run). Use when you don't know what to look at.\n\n## version / mcp\n\n```bash\nvmware-debug version    # installed version\nvmware-debug mcp        # start the stdio MCP server (no network at startup)\n```\n\n## How the agent uses it\n\nIn an agent, the cross-skill correlation happens at the agent layer:\n\n1. Fetch events with the data-source skills (vmware-monitor `event_list`,\n   vmware-log-insight `log_search`/`log_aggregate`, vmware-aria alerts/anomaly,\n   vmware-nsx).\n2. Normalise each into the event envelope.\n3. Call the `incident_timeline` MCP tool to correlate and rank.\n4. Follow `next_checks`; route any fix to vmware-aiops / vmware-pilot.\n\nFile v1.13.0:references/event-envelope.md\n\n# The Unified Event Envelope\n\nThis is the contract between `vmware-debug` and every data-source skill. The\norchestrating agent fetches events with each skill's own read tools, normalises\neach into this shape, and passes the list to `incident_timeline`. Debug has **no\nruntime dependency** on the other packages (no version lockstep, no heavy install).\n\n## Shape\n\n```json\n{\n  \"ts\":       \"2026-06-23T10:15:30Z\",\n  \"source\":   \"monitor\",\n  \"severity\": \"error\",\n  \"entity\":   \"vm-web01\",\n  \"text\":     \"Device naa.600... performance has deteriorated\",\n  \"fields\":   { \"event_type\": \"esx.problem.scsi.device.io.latency.high\",\n                \"host\": \"esxi-03\", \"datastore\": \"ds1\" }\n}\n```\n\n| Field | Type | Notes |\n|---|---|---|\n| `ts` | string \\| number | ISO-8601, epoch **seconds**, or epoch **millis** (auto-detected). Required. |\n| `source` | string | `monitor` \\| `aria` \\| `loginsight` \\| `nsx` \\| `nsx-security` \\| `storage` \\| ... The catalogue in `rules/evidence_sources.yaml` spells the same skills `vmware-monitor`, `vmware-aria`, `vmware-log-insight`. Both spellings are understood wherever a skill is named — `case_readiness(available_skills=...)`, and the grader's count of independent sources, which treats two spellings of one skill as one source. |\n| `severity` | string | Free text; normalised to `critical`/`error`/`warning`/`info`/`unknown`. |\n| `entity` | string | The object the event is about (VM/host/datastore). May be empty. |\n| `text` | string | Human-readable message. Matched by the symptom classifier. |\n| `fields` | object | Any source-specific extras; preserved, never dropped. `event_type` / `eventTypeId` is matched by the classifier too — see below. |\n\n### `fields.event_type` is worth passing\n\nThe classifier matches the identifier as well as the prose, with camelCase and\ndots split so a two-word keyword can reach it (`HostShutdownEvent` → `host\nshutdown`, `esx.problem.cpu.ready` → `cpu ready`).\n\nFor a classic vCenter event this changes almost nothing — `VmMigratedEvent` and\nits message say the same words. It matters for `EventEx`, which is how modern\nvSphere emits most events: the message is generic boilerplate (\"Issue detected on\nesx01\") and the whole identity is in `eventTypeId`. Drop that field and a storage\nincident arrives as an unreadable event; keep it and it classifies as storage.\n\n`vmware-monitor`'s `get_events` returns it as `event_type`; passing the row\nthrough unchanged is enough.\n\nThe normaliser is tolerant of common field-name variants (e.g. `timestamp`,\n`createTime`, `startTimeUTC` for `ts`; `criticality`, `level` for `severity`;\n`resourceName`, `vm_name`, `fullFormattedMessage` for entity/text), so most\nsources map with little or no adaptation.\n\n## Mapping cheatsheet per source\n\n| Source tool (example) | ts | severity | entity | text |\n|---|---|---|---|---|\n| vmware-monitor `event_list` | `createdTime` | `severity` | `vm`/`host` | `fullFormattedMessage` |\n| vmware-aria `alert_query` | `startTimeUTC` | `criticality` | `resourceName` | `alertDefinitionName` |\n| vmware-aria `anomaly` | `timestamp` | (derive) | `resourceName` | stat + value |\n| vmware-log-insight `log_search` | `timestamp` | `severity`/derive | `hostname` | `text` |\n| vmware-nsx (firewall/traceflow) | `time` | (derive) | src/dst | rule/verdict |\n\n## Why this design\n\n- **Decoupling** — debug never imports monitor/aria/log-insight (CLAUDE.md 踩坑 #21/#32).\n- **Testability** — correlation is pure functions over `Event`; unit tests feed synthetic events.\n- **Transparency** — the cross-skill \"联动\" happens at the agent layer, visibly, not hidden inside debug.\n\nFile v1.13.0:references/routing.md\n\n# Symptom → Signal → Skill Routing\n\nThe catalogue debug uses to turn \"something is wrong\" into concrete next steps.\nKeep this in sync with `_CATEGORY_SIGNATURES` in `vmware_debug/ops/timeline.py`\n(a regression test asserts the two match).\n\n| Category | Keyword signatures (sample) | Pull signals from | Fix via |\n|---|---|---|---|\n| **storage** | datastore, scsi, latency, vsan, apd, pdl, no space, vmfs, iscsi | vmware-storage, vmware-log-insight (vmkernel scsi/apd) | aiops / pilot |\n| **network** | vmotion, uplink, link down, mtu, firewall, dfw, segment, tier-0, bgp | vmware-nsx, vmware-nsx-security (traceflow, DFW hits) | pilot |\n| **compute** | cpu ready, memory, balloon, swap, contention, numa | vmware-aria (metrics + anomalies) | pilot (rightsizing) |\n| **ha_drs** | ha, high availability, drs, failover, admission control, isolation | vmware-monitor, vmware-aiops (cluster) | pilot |\n| **host_lifecycle** | maintenance mode, shut down of, host reboot, standby mode, lost connection to, cannot synchronize, not responding | vmware-monitor (host/cluster state), vmware-harden (drift — a host that left service on cue was told to), vmware-log-insight (vpxd/hostd) | pilot |\n| **power_lifecycle** | power on/off, failed to start, boot, vmx, ovf, clone, snapshot | vmware-aiops (snapshot tree), vmware-monitor | aiops |\n| **auth** | login, authentication, denied, 401, 403, token, certificate, tls, password | config/.env, target cert + time sync | config fix |\n| **platform** | vpxd, hostd, service restart, crash, 503, not responding, disconnected | vmware-monitor (connection/service), vmware-log-insight (vpxd/hostd) | pilot |\n| **hardware** | tpm, attestation, ipmi, sensor, bmc | vmware-monitor (host sensors and the CIM Server that supplies them, host logs for ipmi/cim), the server's BMC | site / vendor |\n| **licensing** | license | vmware-monitor (license_status — which asset holds which key); an Aria alert on \"Unlicensed Group\" is Aria's own license | license portal |\n| **data_collection** | adapter instance, not receiving data, collector group, cloud proxy | vmware-aria (adapters, collector groups, node health, objects that stopped reporting), vmware-monitor (is the source vCenter reachable) | aria admin |\n\n## Remediation handoff (advisor → executor)\n\nDebug never executes. It mirrors the `vmware-harden → vmware-pilot` pattern:\n\n- **Single, low-risk fix** → call the matching **vmware-aiops** tool (own double-confirm).\n- **Multi-step / approval / cross-skill** → submit the proposed plan to **vmware-pilot**:\n  state machine + approval gate + rollback + audit all live there.\n\n## Adding a new symptom category\n\n1. Add a `(name, keywords, suggested_check)` tuple to `_CATEGORY_SIGNATURES`.\n2. Add the matching row to the table above.\n3. Add a focused unit test in `tests/test_timeline.py` (a sample text → expected category).\n4. Add a playbook under `references/playbooks/` if the investigation is non-obvious.\n\nFile v1.13.0:references/setup-guide.md\n\n# vmware-debug Setup Guide\n\nvmware-debug has **no configuration, no credentials, and no network access** — it\nis a pure, offline correlation engine. There is no `config.yaml` and no `.env`.\n\n## Install\n\n```bash\nuv tool install vmware-debug==1.13.0\nvmware-debug categories      # verify it runs\n```\n\n## MCP client configuration\n\n```json\n{\n  \"command\": \"uvx\",\n  \"args\": [\"--from\", \"vmware-debug==1.13.0\", \"vmware-debug-mcp\"]\n}\n```\n\nIf installed with `uv tool install`, prefer the entry point `vmware-debug mcp`\n(no PyPI resolution at startup — robust behind corporate TLS proxies, 踩坑 #25).\n\nFor full cross-skill diagnosis, also install the data-source skills it correlates\n(vmware-monitor, vmware-log-insight, vmware-aria, vmware-nsx) and the executors it\nroutes fixes to (vmware-aiops, vmware-pilot).\n\n## Security\n\n> **Disclaimer**: Community-maintained open-source project, **not affiliated with,\n> endorsed by, or sponsored by VMware, Inc. or Broadcom Inc.**\n\n1. **Source Code** — https://github.com/vmware-skills/VMware-Debug (MIT).\n2. **Credentials** — none. debug holds no secrets and connects to nothing.\n3. **Network** — none. The correlation tools are local pure functions over\n   event data the agent supplies; the case tools read and write local files.\n4. **Writes** — only to the local investigation ledger under `$OPS_HOME`\n   (the seven [WRITE] `case_*` tools). Evidence, gaps, hypotheses and grade\n   history are only ever added to; `timeline.md` is regenerated from the evidence and\n   `case.json` holds the current grade and state. Nothing is written to any\n   VMware system: debug only diagnoses and recommends; remediation is routed\n   to vmware-aiops / vmware-pilot, where confirmation/approval/audit live.\n   Case data is sensitive: each case folder (`$OPS_HOME/cases/<case-id>/`, default\n   `~/.vmware/cases/`) is created owner-only (`0700`) and holds the submitted\n   evidence — host names, addresses, log and event text. There is no automatic\n   retention limit: `case_close` marks a case closed and keeps its files; delete\n   the folder to remove a case.\n5. **No cross-skill coupling** — events arrive as plain dicts (the event\n   envelope); debug imports no other skill package at runtime.\n6. **Environment scoping** — policy rules can scope by environment, and skills\n   that connect to a VMware estate may declare `environment:` (`production` /\n   `staging` / `lab`) per target in their own `config.yaml` as an optional label\n   an environment-scoped `deny` rule can match on. debug has no config and no\n   connection to declare one about, so it registers no environment resolver —\n   it has no basis to answer for any target — and an environment-scoped rule\n   never matches its tools. Every other policy rule does apply: each MCP tool\n   and the `categories` / `triage` CLI commands pass through vmware-policy's\n   `guard()`. Verified behaviour:\n   - A `deny` rule naming a debug tool, or with no `operations` key at all,\n     refuses it (status `denied`, with the rule's reason).\n   - A malformed `~/.vmware/rules.yaml` fails closed: every tool is refused with\n     a reason naming the file and the way out (fix it, or\n     `VMWARE_POLICY_DISABLED=1`). The `categories` and `triage` commands exit 1\n     with a Python traceback whose last lines are that message; `version` and\n     `mcp` are unaffected.\n   - Runaway guard: within one MCP server process, the 26th call to the same\n     tool with identical arguments inside 120 s is refused (`budget_exceeded`).\n     Tune with `VMWARE_RUNAWAY_MAX` / `VMWARE_RUNAWAY_WINDOW_SEC`. The guard\n     compares a SHA-256 digest of the arguments the tool actually received\n     (vmware-policy ≥1.16.0), not the redacted copy in the audit row — so calls\n     with different `events` or `payload` are never counted as identical.\n   - A `maintenance_window` gates only `high`/`critical` risk operations; every\n     debug tool is `low`, so a window never blocks one.\n7. **Audit row contents** — one row per call in `~/.vmware/audit.db`: tool,\n   time, OS user, detected agent, arguments, status and result. `payload` and\n   `events` are stored as `***`. The results of `incident_timeline` and\n   `case_timeline` are not stored (`[redacted: return value declared\n   sensitive]`) because they quote event text; other results are stored after\n   credential scrubbing. The status is still truthful — a call that returns\n   `{\"error\": …}` is `error`.\n8. **Static analysis** — `uvx bandit -r vmware_debug/` (release bar:\n   0 Medium+).\n\nFile v1.13.0:skill-card.md\n\n## Description:\n\nvmware-debug helps agents troubleshoot VMware, vSphere, ESXi, and NSX incidents by correlating supplied events, ranking root-cause hypotheses, suggesting next checks, and recording local investigation evidence.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[zw008](https://clawhub.ai/user/zw008)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers, operators, and support engineers use this skill to investigate VMware incidents from collected logs, alerts, and events, then decide what evidence to gather next. It is diagnostic only: remediation is routed to companion execution skills rather than performed by this skill.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: Auth troubleshooting can lead an agent toward credential-bearing .env files or other secret material.\n\nMitigation: Do not provide raw .env contents; share only redacted variable names, presence checks, file permissions, and certificate metadata.\n\nRisk: Local case and audit files can contain host names, addresses, log text, event text, and other incident evidence, and they are retained until deleted.\n\nMitigation: Treat the case directory and audit database as sensitive operational records, restrict filesystem access, and delete retained case folders when they are no longer needed.\n\n## Reference(s):\n\n- [ClawHub skill page](https://clawhub.ai/zw008/skills/vmware-debug)\n- [Project homepage](https://github.com/vmware-skills/VMware-Debug)\n- [vmware-debug Setup Guide](references/setup-guide.md)\n- [vmware-debug Capabilities](references/capabilities.md)\n- [The Unified Event Envelope](references/event-envelope.md)\n- [Operating vmware-debug with a local / small model](references/agent-guardrails.md)\n- [Symptom to Signal to Skill Routing](references/routing.md)\n- [vmware-debug CLI Reference](references/cli-reference.md)\n\n## Skill Output:\n\n**Output Type(s):** [text, markdown, JSON, shell commands, guidance]\n\n**Output Format:** [Markdown guidance with JSON tool inputs, JSON diagnostic outputs, and shell command examples]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [May create local case ledger files under the configured operations home when case tools are used; diagnostic responses can include sensitive submitted event text.]\n\n## Skill Version(s):\n\n1.13.0 (source: server release evidence)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nArchive v1.12.2: 9 files, 23355 bytes\n\nFiles: references/agent-guardrails.md (9251b), references/capabilities.md (3160b), references/cli-reference.md (1585b), references/event-envelope.md (3622b), references/routing.md (2958b), references/setup-guide.md (4533b), skill-card.md (2808b), SKILL.md (18409b), _meta.json (132b)\n\nFile v1.12.2:SKILL.md\n\n---\nname: vmware-debug\ndescription: >\n  Use this skill whenever the user is troubleshooting a VMware/vSphere problem —\n  a reported error, an exception, a log dump, a slow or failed VM, a host that\n  went sideways — and needs help locating the root cause. It is the diagnostic\n  brain of the VMware family: it drives a systematic investigation, pulls the\n  right signals from the other skills, correlates events into one timeline,\n  ranks root-cause hypotheses, and tells you what to check next even when you\n  don't know where to start. Always use this skill for \"diagnose this VMware\n  issue\", \"why is my VM slow\", \"troubleshoot this vSphere error\", \"what does\n  this log mean\", \"help me figure out what broke\" when the context is explicitly\n  VMware/vSphere/ESXi/NSX. It never touches vSphere: its only writes are to a\n  local case ledger. Do NOT\n  use it to execute fixes — single fixes go to vmware-aiops, multi-step gated\n  remediation goes to vmware-pilot. Do NOT use it for routine inventory or\n  health checks with no problem to solve — use vmware-monitor.\ninstaller:\n  kind: uv\n  package: vmware-debug\nallowed-tools:\n  - Bash\nmetadata: {\"openclaw\":{\"requires\":{\"anyBins\":[\"vmware-debug\",\"uvx\"]},\"optional\":{\"env\":[\"VMWARE_AUDIT_APPROVED_BY\",\"VMWARE_AUDIT_RATIONALE\"],\"bins\":[\"vmware-policy\"]},\"homepage\":\"https://github.com/vmware-skills/VMware-Debug\",\"os\":[\"macos\",\"linux\"]}}\n---\n\n# VMware Debug\n\n> **Disclaimer**: Community-maintained open-source project, **not affiliated with,\n> endorsed by, or sponsored by VMware, Inc. or Broadcom Inc.** \"VMware\" and \"vSphere\"\n> are trademarks of Broadcom. Source is publicly auditable under the MIT license.\n\nThe diagnostic brain of the VMware skill family. You bring the symptom; this skill\nruns the investigation and points at the root cause. It **reads and reasons** — it\nnever writes to vSphere; its only writes go to its own local case ledger.\nCompanion skills do the data collection and the fixing.\n\n## What This Skill Does\n\n| Category | What | Read or Write |\n|---|---|---|\n| Incident correlation | Merge events from many sources into one timeline, detect spikes | Read |\n| Root-cause ranking | Score symptom clusters, surface the most likely cause first | Read |\n| Next-check ideas | Suggest exactly what to look at next (which skill/tool) when you're stuck | Read |\n| Remediation routing | Hand the fix to vmware-aiops (single) or vmware-pilot (gated, multi-step) | Read (routes only) |\n| Investigation ledger | Open a case; record evidence, gaps and hypotheses; grade and close it | Write (local ledger only) |\n\n**No network access of its own, and no write to any VMware system.** It correlates\ndata the agent has already gathered with the other skills' read tools; its seven\nwrite tools touch only the local case ledger.\n\n## Quick Install\n\n```bash\nuv tool install vmware-debug==1.12.2\nvmware-debug categories          # see what it can diagnose\n```\n\n## When to Use This Skill\n\nUse it when there is a **problem to solve**: an error message, a stack of logs, an\nalarm storm, \"my VM won't power on\", \"storage feels slow\", \"the host disconnected\".\n\n- Need raw inventory/health with no incident? → **vmware-monitor**\n- Need to actually run a fix? → **vmware-aiops** (single op) or **vmware-pilot** (gated workflow)\n- Need metrics/anomalies? → **vmware-aria**; centralized logs? → **vmware-log-insight**\n\n**Do NOT use when** there is nothing wrong (routine listing → monitor), or when the\nuser wants the fix executed (→ aiops/pilot). This skill stops at the diagnosis and\na recommended plan.\n\n## Related Skills — Skill Routing\n\n| Symptom touches | Pull signals from | Then |\n|---|---|---|\n| Storage / datastore / vSAN | vmware-storage, vmware-log-insight | rank → route fix to aiops/pilot |\n| Network / firewall / vMotion | vmware-nsx, vmware-nsx-security | run traceflow, check DFW |\n| CPU / memory contention | vmware-aria (metrics/anomalies) | rightsizing via pilot |\n| HA / DRS / cluster | vmware-monitor, vmware-aiops | cluster remediation via pilot |\n| Power / clone / snapshot | vmware-aiops, vmware-monitor | task status, then fix via aiops |\n| Auth / cert / login | check creds & cert; (security) | fix config/.env |\n\n## Common Workflows\n\n### 1. \"Here's a pile of logs / alarms — what broke?\"\n1. Collect events with the data-source skills (e.g. `vmware-monitor event_list --vm web01 --since 1h`, `vmware-log-insight log_search ...`, `vmware-aria alert_query ...`).\n2. Pass them all to **`incident_timeline`** (envelope below). Read the top hypothesis + `next_checks`.\n3. Follow `next_checks` to pull more targeted data; re-run `incident_timeline` to confirm.\n4. **Failure branch — no events come back:** the affected target may be unreachable. Run the source skill's `doctor`/health first; a 503/timeout is a *signal* (platform not ready), not a dead end.\n5. Produce a diagnosis + recommended fix. Route execution to aiops/pilot. **Do not fix here.**\n\n### 2. \"I don't even know what to check\"\n1. Run **`list_symptom_categories`** (or `vmware-debug categories`) to see the catalogue.\n2. Describe the symptom; map it to a category; the `suggested_check` tells you which skill/tool to run first.\n3. Collect → `incident_timeline` → narrow. Loop until one hypothesis dominates.\n\n### 3. Hand off the fix (advisor → executor, like vmware-harden)\n1. Debug emits a structured diagnosis + a proposed remediation (steps).\n2. **Single, low-risk fix** → call the matching **vmware-aiops** tool (it has its own double-confirm).\n3. **Multi-step / needs approval / cross-skill** → submit the plan to **vmware-pilot**, which owns the state machine, approval gate, rollback, and audit.\n4. **Failure branch — fix is ambiguous or risky:** stop and present the hypotheses to the user; never guess-execute.\n\n## Usage Mode\n\n- **MCP** (in an agent): the agent calls the other skills' read tools, then `incident_timeline` to correlate. This is the primary mode — that's where the cross-skill \"联动\" happens.\n- **CLI** (humans): `vmware-debug triage --events events.json` correlates a JSON array you collected yourself.\n\n## MCP Tools (14 — 7 read, 7 write)\n\n**Correlation** — stateless, for a single look:\n\n| Tool | What |\n|---|---|\n| `incident_timeline` | [READ] Correlate pre-fetched events → timeline + spikes + ranked hypotheses + next-check ideas |\n| `list_symptom_categories` | [READ] List recognised symptom categories + what to check for each |\n\n**Investigation ledger** — for an incident you will reason about over time:\n\n| Tool | What |\n|---|---|\n| `case_open` | [WRITE] Define the event; returns a case id and the grade this environment can reach |\n| `case_readiness` | [READ] What grade this environment can reach, per symptom category, **before** you start |\n| `case_knowledge` | [READ] Which knowledge formats are accepted, what is mounted, and which entries apply to a case |\n| `case_plan` | [READ] What to fetch next — skill, tool and purpose per step; recomputed from the case's current state |\n| `case_list` | [READ] Cases, newest first |\n| `case_get` | [READ] One case: scope, ledger sizes, grade history |\n| `case_hypotheses` | [WRITE] Register a candidate explanation, or read the ledger of what supports and refutes each |\n| `case_submit_evidence` | [WRITE] Record one retrieved fact, with its source, query and time basis |\n| `case_record_gap` | [WRITE] Record what could **not** be retrieved, and how to close it |\n| `case_timeline` | [WRITE] Correlate everything the case has collected into one timeline |\n| `case_grade` | [WRITE] Recompute the conclusion grade from the ledger and record it |\n| `case_close` | [WRITE] Record the final grade, archive, and name what was left open |\n\nThe seven writes go to `$OPS_HOME` (default `~/.vmware/cases/`) and nowhere else.\n\n**List envelope** (output of `list_symptom_categories`): `{items, returned, limit, total, truncated, hint}` — read the rows from `items`. `truncated` is always `false` here, which is the point: it states that the catalogue is complete instead of leaving you to infer it.\n\n**Event envelope** (input to `incident_timeline`): `{ts, source, severity, entity, text, fields}`.\nSee `references/event-envelope.md`. The agent normalises each source's events into this\nshape; debug stays source-agnostic and has no dependency on the other packages.\nKeep each event's `event_type` in `fields` — the classifier matches it as well as\nthe message, and on a modern `EventEx` it is the only thing that says what the\nevent was.\n\n## The Investigation Ledger\n\nCorrelating events answers \"what happened together\". A case answers \"what do we\nbelieve, on what evidence, and what is still missing\" — and keeps answering it\nacross sessions and across people.\n\nThe case folder is the deliverable, not an implementation detail. Everything in\nit but the index is plain text, so a customer can take the folder away and audit\nhow a conclusion was reached with none of this installed:\n\n```\n~/.vmware/cases/<case-id>/\n├── scope.json    what is being investigated, and how that was decided\n├── evidence/     one file per fact: source skill, exact query, time basis\n├── gaps.json     what could NOT be obtained, what it blocks, how to close it\n├── conclusion.md the grade, appended — including every time it went down\n└── timeline.md · hypotheses.md · plan.jsonl · case.json\n```\n\nHypotheses get ids (H1, H2, …), and those ids are what `case_record_gap(blocks=…)`\nand `case_submit_evidence(falsifies=…)` refer to. **An id that was never\nregistered is refused, not ignored** — a dangling reference blocks nothing and\nfalsifies nothing, which quietly reports a stronger case than you have.\n\n**You cannot state a conclusion level.** `case_grade` has no parameter for one;\nthe grade is recomputed from the ledger on every call. To change it, change the\nledger — submit the missing evidence, or record the gap that is blocking it.\n\n- **Candidate** — a hypothesis exists\n- **Probable** — ≥2 *independent* sources agree (two calls to one skill are one\n  source) and nothing outstanding could overturn it\n- **Confirmed** — that, plus a decisive item: a direct hardware diagnostic, a\n  version-checked knowledge-base entry, or a vendor SR; and no gap left open\n- **Excluded** — an observation that actually rules it out. \"We looked and found\n  nothing\" is a gap, not an exclusion\n\n`case_plan` is not a checklist: submit evidence and the next plan is shorter,\nlose a source and it routes around it. Its `unavailable` half is the important\none — a source this install cannot reach is listed there with how to supply it,\nso the gap is visible now rather than when the conclusion refuses to firm up.\n\n`case_readiness` answers this per symptom category rather than as one number —\n\"storage reaches Probable, hardware reaches Candidate\" can be acted on;\n\"readiness 78%\" cannot. Two classes served by the *same* skill count as one\nsource, so it agrees with what `case_grade` will actually award.\n\n### Mounting a knowledge library\n\n`case_knowledge` answers \"what can I add\" without anyone reading this file.\nEntries go under `$OPS_HOME/knowledge/{kb,runbook,sr,cases}/`:\n\n| Format | How metadata travels |\n|---|---|\n| `.md` `.markdown` | YAML front-matter between `---` fences, body below — **preferred** |\n| `.yaml` `.yml` | the whole file is one entry |\n| `.json` | one entry per file |\n| `.jsonl` | one JSON object per line — what ticketing systems export |\n| `.csv` `.tsv` | one entry per row; use dotted columns (`driver.version`) for nested constraints |\n| `.txt` `.log` | a sibling `<name>.yaml` carries the metadata |\n\nPDF, DOCX, PPTX and HTML must be converted to Markdown first — `case_knowledge`\nnames them rather than ignoring them.\n\n**Every entry needs an `applies_to` block to be decisive:**\n\n```yaml\n---\nid: KB-2026-0417\napplies_to:\n  product: vsphere\n  build: \">=8.0.3, <9.0\"\n  driver: {name: nvme_pcie, version: \">=1.2.4\"}\n  firmware: {vendor: dell, version: \">=52.26\"}\n---\n```\n\n`product`, `build`, `driver` and `firmware` are the constraints the checker\nevaluates. **Any other key leaves the entry non-decisive** — including a typo —\nbecause an unverified constraint is not a satisfied one, and `case_knowledge`\nnames the key it could not check.\n\nKnowledge evidence must say **which** entry it is\n(`case_submit_evidence(source_skill=\"knowledge-kb\", knowledge_entry_id=\"KB-…\")`)\n— \"some applicable entry is mounted somewhere\" is a different claim from \"this\none applies\", and only the second can carry a conclusion.\n\nMatching is **by version applicability, never by similarity** — an entry written\nfor the wrong build reads exactly like the right one, and similarity is the only\nthing that would let it through. An entry with no `applies_to` can support a\nhypothesis but can never make a case Confirmed. A constraint the case scope\ncannot answer is not a match either: silence is not a pass.\n\n> **On a stock install the ceiling is Probable.** Confirmed needs a decisive\n> source, and there is neither a hardware-diagnostic channel (no Redfish/BMC, no\n> SMART/NVMe) nor a knowledge library mounted — `~/.vmware/knowledge/` ships\n> empty. Every tool that can reach the ceiling says so in its output rather than\n> letting you wonder why a well-supported case never goes higher. Mount a\n> knowledge library and the ceiling rises on its own.\n\nRecording a gap is meant to be free: a missing confirmation caps the grade, it\ndoes not demote it. Only a gap that could *overturn* the hypothesis holds a case\nat Candidate.\n\n## No Network, By Design\n\nvmware-debug connects to nothing and holds no credentials. The calling agent\nfetches with the other skills' read tools and submits the results here. Its only\nwrites are to the local case ledger. Running with local or small models? See\n[`references/agent-guardrails.md`](references/agent-guardrails.md).\n\n## CLI Quick Reference\n\n```bash\nvmware-debug categories                        # what can it diagnose\nvmware-debug triage --events events.json       # correlate a collected event set\ncat events.json | vmware-debug triage          # or via stdin\nvmware-debug mcp                                # start stdio MCP server (proxy-safe)\n```\n\n## Troubleshooting\n\n- **`incident_timeline` raises \"event[N] could not be normalised\"** — event N is missing a timestamp or has an unparseable one. Every event needs `ts` (ISO-8601, epoch seconds, or millis).\n- **Most events come back \"uncategorized\"** — read `classification` in the result: it says how many, what share, and quotes the texts that matched nothing. Do **not** widen the window; a wider window adds baseline, not signal. The spikes still tell you *when*, which is answerable without a category. If the samples name a subsystem the taxonomy does not know, add a signature (see `references/routing.md`).\n- **No spikes detected on an obvious burst** — check `binning` for the resolution you were given. The width is chosen from event density so that bins average ≥4 events; a burst shorter than one bin can still be flattened by it. Pass `bin_seconds` to narrow. Below three bins there is no baseline at all and nothing is reported.\n- **`case_timeline` says zero events after you submitted results** — check `payload_events` in each `case_submit_evidence` reply. Events are read from a bare list, or from `items`/`events`/`rows`; a *summary* of a tool's result carries none. `note` names which items carried nothing and what keys they held instead.\n- **`case_readiness` says a skill you have is not installed** — both spellings are accepted (`monitor` and `vmware-monitor`). A name it did not recognise comes back in `unrecognised_skills` rather than being read as missing.\n- **It won't execute the fix** — by design. Route to vmware-aiops or vmware-pilot.\n\n## Audit & Safety\n\nNo network, nothing executed. The seven [WRITE] tools write only to the local case\nledger under `$OPS_HOME`; nothing here touches a remote VMware estate. Evidence, gaps,\nhypotheses and grade history are only ever added to — `timeline.md` is regenerated from\nthe evidence and `case.json` holds the current grade and state. Remediation is always routed to\naiops/pilot, where the double-confirm / approval gates live.\n\n**What the audit row holds.** Every debug tool call — a failed one included — writes one\nrow to `~/.vmware/audit.db` through `@vmware_tool`, as do the `categories` and `triage`\nCLI commands under their MCP tools' names. The row records the tool, time, OS user,\ndetected agent, arguments, status (`ok`, `error`, `denied`, `budget_exceeded`,\n`interrupted`) and a result. Two arguments are stored as `***`: evidence `payload` and\n`incident_timeline` `events`. Two results are not stored at all — `incident_timeline` and\n`case_timeline` quote event text (`sample_text`, `unmatched_samples`, `rejected`), which\ncan carry anything the source logged, so their row says\n`[redacted: return value declared sensitive]`. Every other result (ids, counts, grades,\nyour own summaries and statements) is stored after credential scrubbing. The data itself\nis in the case folder, or with you.\n\n**What debug inherits from vmware-policy**, since its tools pass through `guard()`:\n- **Runaway guard** — the 26th call to one tool with identical arguments within 120 s is\n  refused (`budget_exceeded`; `VMWARE_RUNAWAY_MAX`, `VMWARE_RUNAWAY_WINDOW_SEC`). It\n  compares a digest of the arguments the tool received (vmware-policy ≥1.16.0), so\n  calls with different `events` are not identical.\n- **A malformed `~/.vmware/rules.yaml` fails closed** — every tool is refused with a reason\n  naming the file; `categories`/`triage` exit 1 with a traceback ending in that message.\n- **A `deny` rule without `operations`** matches every tool, debug's included.\n- **Environment-scoped rules never match** — debug registers no resolver (no targets).\n- **Maintenance windows** gate only `high`/`critical` risk; every debug tool is `low`.\n\nSee `references/setup-guide.md`.\n\n**Case data is sensitive and is kept until you delete it.** Each case lives in\n`$OPS_HOME/cases/<case-id>/` (default `~/.vmware/cases/`), created owner-only (`0700`).\nIt holds whatever evidence was submitted — host names, addresses, log and event text.\nNothing is deleted automatically: `case_close` records the grade and marks the case\nclosed but keeps its files. Remove a case by deleting its folder.\n\n## License\n\nMIT.\n\nFile v1.12.2:_meta.json\n\n{\n  \"ownerId\": \"kn7b067awq2s97bn3d7p5qfhw5827pxc\",\n  \"slug\": \"vmware-debug\",\n  \"version\": \"1.12.2\",\n  \"publishedAt\": 1789462511605\n}\n\nFile v1.12.2:references/agent-guardrails.md\n\n# Operating vmware-debug with a local / small model\n\nClaude-class models drive this skill without special instruction. Smaller and\nlocally-hosted models — Llama 3.3 70B, Qwen, Mistral, and similar, served\nthrough Goose, Ollama, or OpenShift AI — need explicit operating rules to call\ntools reliably.\n\nThis page exists because an operator wrote those rules by hand first. The\nguardrails below are adapted, with thanks, from the working configuration\n[@juanpf-ha](https://github.com/juanpf-ha) developed while running\nvmware-monitor and vmware-aria against a production vSphere estate with Llama\n3.3 70B FP8 on an on-prem H100\n([VMware-AIops#31](https://github.com/vmware-skills/VMware-AIops/issues/31)). The\ncross-skill rules are identical across this family; the parts below marked\nvmware-debug are specific to this skill.\n\nvmware-debug exposes 14 MCP tools: 7 reads and 7 writes. It connects to nothing\nand holds no credentials — the calling agent gathers events from the other\nskills, normalises them, and hands them over. The seven writes go to the local\ninvestigation ledger under `~/.vmware/cases/`, never to a VMware system. That\nmakes it the safest skill in the family to point a small model at — and the one\nmost exposed to the model's reasoning, because its output *is* an\ninterpretation.\n\nFor a small model the ledger is more than bookkeeping: `case_grade` computes the\nconclusion level from recorded evidence rather than accepting one, so a model\nthat would happily narrate \"root cause confirmed\" cannot record that unless the\nevidence for it is actually in the folder.\n\n> **Disclaimer**: This is a community-maintained open-source project and is\n> **not affiliated with, endorsed by, or sponsored by VMware, Inc. or Broadcom\n> Inc.** \"VMware\" and \"vSphere\" are trademarks of Broadcom.\n\n---\n\n## First: the rules you no longer need to write\n\nSeveral guardrails from the original configuration are now enforced by the\nskill itself. Prompt instructions are advisory — a model can ignore them.\nThese are structural, so it cannot.\n\n| Guardrail you would otherwise prompt for | Now enforced by |\n|---|---|\n| \"Work read-only and never modify anything\" | **The tool surface itself.** No tool reaches a VMware system. The seven [WRITE] tools write only to this skill's own local case ledger, so there is nothing in vSphere to withhold and nothing to switch off. |\n| \"Diagnose only — never apply the fix you propose\" | **Structural.** This skill has no tool that changes anything outside its own ledger, and it holds no connection to vCenter, NSX or anything else. Remediation is routed to vmware-aiops or vmware-pilot by the calling agent. |\n| \"Do not fabricate a timeline — build it from the events I gave you\" | **`incident_timeline` correlates only its input.** It is source-agnostic and has no way to fetch anything, so the timeline cannot contain an event the agent did not supply. |\n| \"Tell me when the symptom is outside what you can recognise\" | **`list_symptom_categories`** states the catalogue, and unmatched symptoms come back as `uncategorized` rather than being forced into the nearest signature. |\n| \"Use explicit limits for queries that may return large amounts of data\" | **The list envelope.** `list_symptom_categories` returns `{items, returned, limit, total, truncated, hint}` with `truncated` always `false` — which is the point: it states that the catalogue is complete instead of leaving you to infer it. |\n| \"Log everything you looked at\" | **`~/.vmware/audit.db` and the case ledger.** Every debug tool call, a failed one included, writes one audit row (tool, time, user, agent, arguments, status and result — evidence `payload` and `events` are stored as `***`, and the results of `incident_timeline` and `case_timeline`, which quote event text, are not stored at all). Evidence submitted to a case is also recorded in the ledger with its source skill, tool, query and fetch time — so pass the real fetch time, never a placeholder. |\n\n---\n\n## The system prompt\n\nEverything below still benefits from being stated explicitly. Copy this into\nyour agent's instruction block.\n\n```text\n## Tool use\n\n- Gather real events with the companion skills' read tools before calling\n  incident_timeline. Never answer from memory or assumption, and never\n  hand-write events to represent what you believe happened.\n- Never describe a tool call, and never output a JSON example, instead of\n  executing the tool. If you intend to call a tool, call it.\n- If a tool fails, report the actual error text. Do not complete the answer\n  with assumptions about what the result would have been.\n- Use explicit limits and time windows when gathering events. Do not request\n  unlimited results unless the user asks for them.\n\n## Skill routing\n\n- vmware-debug: correlate pre-fetched events into a timeline, detect spikes,\n  rank root-cause hypotheses, suggest next checks.\n- Gather the events from: vmware-monitor (alarms, events, host logs),\n  vmware-log-insight (centralised logs), vmware-aria (metrics, anomalies),\n  vmware-nsx / vmware-nsx-security (network and firewall), vmware-storage.\n- Route the fix to vmware-aiops, or to vmware-pilot when it needs approval\n  gating. This skill executes nothing.\n\n## Building the event set\n\n- Every event needs ts (ISO-8601, epoch seconds, or millis), and the envelope\n  shape {ts, source, severity, entity, text, fields}. An event that cannot be\n  normalised is rejected by index — fix that event, do not drop the batch.\n- Carry each event's original source and severity through unchanged. Do not\n  re-grade a severity to make a story cohere.\n- Pull from more than one source before concluding. A single source's view of\n  an incident is not a correlation.\n- Widen the window before concluding \"no spike\": spike detection needs at least\n  three time bins, so a short window or a large bin_seconds can hide a burst.\n\n## Data fidelity\n\n- Never invent events, timestamps, entities, or relationships between them. If\n  an event was not in the input, it does not exist for this answer.\n- Preserve the exact severity and source values. Do not translate, normalise,\n  or prettify them.\n- Report the timeline in the order the tool returned it.\n- If a requested field was not returned, show it as \"not available\".\n- When a response is long, report every item it contains.\n\n## Analysis discipline\n\n- Separate observed data from interpretation. State which is which. In this\n  skill that separation is the deliverable.\n- Report the ranked hypotheses the tool returned, with its ranking. Do not\n  promote your own preferred explanation above them, and do not present the\n  top hypothesis as a diagnosis.\n- Report next_checks as checks still to be run, not as findings.\n- Do not claim a root cause is confirmed. The tool ranks; it does not conclude.\n- Avoid generic recommendations that are not directly supported by the results.\n```\n\n---\n\n## Known failure modes on small models\n\nObserved with Llama 3.3 70B FP8 (Goose, on-prem H100), and useful as a\nchecklist when evaluating any local model against these skills:\n\n| Symptom | Mitigation |\n|---|---|\n| Describes a tool call, or emits a JSON example, instead of executing it | The \"never describe a tool call\" rule above. This skill is unusually exposed to it: its input *is* JSON, so a model that has just been shown an event envelope will sometimes emit a fabricated one instead of gathering real events. Check that the events came from tool calls. |\n| Long tool responses: omits items, or reports \"no data returned\" when data was present | Ask for explicit limits and narrow windows when gathering. `list_symptom_categories` states `truncated: false`, so a \"no categories\" claim is checkable against `returned`. |\n| Adds generic recommendations unsupported by results | The \"analysis discipline\" rules. Root-cause output attracts invented advice more than any other shape of result in this family. |\n| Drops requested fields or reorders results | State the required fields and ordering in the request itself. A reordered timeline is a different incident. |\n| Multi-tool workflows take 30–50s end to end | Unavoidable in part — this skill's whole premise is a fan-out gather. Narrow each source's window before widening, and prefer the companion skills' aggregate tools when gathering. |\n| Presents the top-ranked hypothesis as the confirmed cause | The \"the tool ranks, it does not conclude\" rule. |\n| Re-grades an event's severity so the narrative fits | The \"carry severity through unchanged\" rule. |\n| Concludes from one source because the others were slow to query | The \"pull from more than one source\" rule. |\n| Reports \"no spike\" from a window too short to have a baseline | Three bins minimum. Widen the window or shrink `bin_seconds` before concluding. |\n| Offers to apply the fix it proposed | It cannot. Route to vmware-aiops or vmware-pilot. |\n\n## Reporting results\n\nLocal-model compatibility is an explicit design constraint for this family, and\nthe evidence base is small. If you evaluate a model against this skill —\nQwen, Mistral, Granite, or anything else — a report of what worked and what did\nnot is genuinely useful:\n[github.com/vmware-skills/VMware-Debug/issues](https://github.com/vmware-skills/VMware-Debug/issues).\n\nFile v1.12.2:references/capabilities.md\n\n# vmware-debug Capabilities\n\nOffline incident correlation. No network, no credentials, no writes to any VMware\nsystem. This table covers the two stateless correlation tools; the twelve `case_*`\ninvestigation-ledger tools (seven of which write, to the local ledger only) are\nlisted in `SKILL.md`, and their response sizes are not yet measured here.\n\n| Tool | What it returns | Typical response tokens |\n|---|---|---|\n| `incident_timeline` | `{event_count, window, spikes:[{start,end,count,zscore}], hypotheses:[{category, score, summary, evidence_count, first_seen, last_seen, sample_text, suggested_check}], next_checks:[...]}` | 300–2000 (scales with hypotheses) |\n| `list_symptom_categories` | `{items: [{category, example_keywords, suggested_check}], returned, limit, total, truncated, hint}` | ~400 |\n\n`list_symptom_categories` returns the family list envelope — read the rows from\n`items`. It has no `limit` parameter, which is exactly why the envelope matters:\n`truncated: false` states that this is every category there is, rather than\nleaving a model to guess whether it is holding page one. The catalogue is a\nfixed in-process constant, so `total` is a real count and `limit` is `null`.\n\n## Correlation engine\n\n- **Timeline**: events normalised to the unified envelope, sorted, and time-binned\n  (auto bin width ≈ span/30, or caller-specified).\n- **Spike detection**: z-score over bin counts (≥3 bins required for a baseline;\n  flat series yields no false spikes).\n- **Hypothesis ranking**: events clustered by symptom category (keyword match on\n  text + entity), scored by summed severity weight, tie-broken by recency.\n  Uncategorised events are kept visible, not dropped.\n- **Next-check routing**: each category carries a concrete \"which skill/tool to run\n  next\" suggestion — the value when the user doesn't know what to check.\n\n## Symptom categories\n\n`storage`, `network`, `compute`, `ha_drs`, `host_lifecycle`, `power_lifecycle`,\n`auth`, `platform`, `hardware`, `licensing`, `data_collection`.\nSee `references/routing.md` for keyword signatures and the skill each routes to.\n\n`hardware`, `licensing` and `data_collection` were added after real alert titles\n(\"Host TPM attestation alarm\", \"License will soon expire\", \"Objects are not\nreceiving data from adapter instance\") matched no category at all. A roll-up\nsuch as \"Group population health is degraded\" is deliberately left\nuncategorized: it names no subsystem, and the cause is in one of its members.\n\n`host_lifecycle` is a host changing its own availability state — maintenance\nmode, shutdown, reboot, standby, connection loss, sync failure.\n`power_lifecycle` is the VM-level equivalent. They are separate because they are\nseparate investigations: the second is a task question for vmware-aiops, the\nfirst is a cluster, DPM, vLCM or drift question.\n\n## Design properties\n\n- **Zero cross-skill runtime deps** — correlation is pure functions over plain\n  dicts; the agent fans out to other skills' read tools (踩坑 #21/#32).\n- **JSON-serialisable output** — suitable for direct MCP responses.\n- **Immutable** — inputs are never mutated; every function returns new values.\n\nFile v1.12.2:references/cli-reference.md\n\n# vmware-debug CLI Reference\n\nAll commands are offline (no network, no credentials). All but `mcp` are\nread-only; `mcp` starts the MCP server, whose seven `case_*` write tools record\ninto the local case ledger.\n\n## triage — correlate a set of collected events\n\n```bash\nvmware-debug triage [OPTIONS]\n  -e, --events PATH     JSON file of event envelopes (reads stdin if omitted)\n      --bin-seconds N   Time-bin width (auto if omitted)\n      --top-n N         Max hypotheses to return   [default: 5]\n```\n\nInput is a JSON array of event envelopes (see `references/event-envelope.md`):\n\n```bash\ncat events.json | vmware-debug triage\nvmware-debug triage --events events.json --top-n 3\n```\n\nOutput (JSON): `{event_count, window, spikes, hypotheses, next_checks}`.\n\n## categories — list recognised symptom categories\n\n```bash\nvmware-debug categories\n```\n\nPrints each category, sample keywords, and the suggested next check (which\nskill/tool to run). Use when you don't know what to look at.\n\n## version / mcp\n\n```bash\nvmware-debug version    # installed version\nvmware-debug mcp        # start the stdio MCP server (no network at startup)\n```\n\n## How the agent uses it\n\nIn an agent, the cross-skill correlation happens at the agent layer:\n\n1. Fetch events with the data-source skills (vmware-monitor `event_list`,\n   vmware-log-insight `log_search`/`log_aggregate`, vmware-aria alerts/anomaly,\n   vmware-nsx).\n2. Normalise each into the event envelope.\n3. Call the `incident_timeline` MCP tool to correlate and rank.\n4. Follow `next_checks`; route any fix to vmware-aiops / vmware-pilot.\n\nFile v1.12.2:references/event-envelope.md\n\n# The Unified Event Envelope\n\nThis is the contract between `vmware-debug` and every data-source skill. The\norchestrating agent fetches events with each skill's own read tools, normalises\neach into this shape, and passes the list to `incident_timeline`. Debug has **no\nruntime dependency** on the other packages (no version lockstep, no heavy install).\n\n## Shape\n\n```json\n{\n  \"ts\":       \"2026-06-23T10:15:30Z\",\n  \"source\":   \"monitor\",\n  \"severity\": \"error\",\n  \"entity\":   \"vm-web01\",\n  \"text\":     \"Device naa.600... performance has deteriorated\",\n  \"fields\":   { \"event_type\": \"esx.problem.scsi.device.io.latency.high\",\n                \"host\": \"esxi-03\", \"datastore\": \"ds1\" }\n}\n```\n\n| Field | Type | Notes |\n|---|---|---|\n| `ts` | string \\| number | ISO-8601, epoch **seconds**, or epoch **millis** (auto-detected). Required. |\n| `source` | string | `monitor` \\| `aria` \\| `loginsight` \\| `nsx` \\| `nsx-security` \\| `storage` \\| ... The catalogue in `rules/evidence_sources.yaml` spells the same skills `vmware-monitor`, `vmware-aria`, `vmware-log-insight`. Both spellings are understood wherever a skill is named — `case_readiness(available_skills=...)`, and the grader's count of independent sources, which treats two spellings of one skill as one source. |\n| `severity` | string | Free text; normalised to `critical`/`error`/`warning`/`info`/`unknown`. |\n| `entity` | string | The object the event is about (VM/host/datastore). May be empty. |\n| `text` | string | Human-readable message. Matched by the symptom classifier. |\n| `fields` | object | Any source-specific extras; preserved, never dropped. `event_type` / `eventTypeId` is matched by the classifier too — see below. |\n\n### `fields.event_type` is worth passing\n\nThe classifier matches the identifier as well as the prose, with camelCase and\ndots split so a two-word keyword can reach it (`HostShutdownEvent` → `host\nshutdown`, `esx.problem.cpu.ready` → `cpu ready`).\n\nFor a classic vCenter event this changes almost nothing — `VmMigratedEvent` and\nits message say the same words. It matters for `EventEx`, which is how modern\nvSphere emits most events: the message is generic boilerplate (\"Issue detected on\nesx01\") and the whole identity is in `eventTypeId`. Drop that field and a storage\nincident arrives as an unreadable event; keep it and it classifies as storage.\n\n`vmware-monitor`'s `get_events` returns it as `event_type`; passing the row\nthrough unchanged is enough.\n\nThe normaliser is tolerant of common field-name variants (e.g. `timestamp`,\n`createTime`, `startTimeUTC` for `ts`; `criticality`, `level` for `severity`;\n`resourceName`, `vm_name`, `fullFormattedMessage` for entity/text), so most\nsources map with little or no adaptation.\n\n## Mapping cheatsheet per source\n\n| Source tool (example) | ts | severity | entity | text |\n|---|---|---|---|---|\n| vmware-monitor `event_list` | `createdTime` | `severity` | `vm`/`host` | `fullFormattedMessage` |\n| vmware-aria `alert_query` | `startTimeUTC` | `criticality` | `resourceName` | `alertDefinitionName` |\n| vmware-aria `anomaly` | `timestamp` | (derive) | `resourceName` | stat + value |\n| vmware-log-insight `log_search` | `timestamp` | `severity`/derive | `hostname` | `text` |\n| vmware-nsx (firewall/traceflow) | `time` | (derive) | src/dst | rule/verdict |\n\n## Why this design\n\n- **Decoupling** — debug never imports monitor/aria/log-insight (CLAUDE.md 踩坑 #21/#32).\n- **Testability** — correlation is pure functions over `Event`; unit tests feed synthetic events.\n- **Transparency** — the cross-skill \"联动\" happens at the agent layer, visibly, not hidden inside debug.\n\nFile v1.12.2:references/routing.md\n\n# Symptom → Signal → Skill Routing\n\nThe catalogue debug uses to turn \"something is wrong\" into concrete next steps.\nKeep this in sync with `_CATEGORY_SIGNATURES` in `vmware_debug/ops/timeline.py`\n(a regression test asserts the two match).\n\n| Category | Keyword signatures (sample) | Pull signals from | Fix via |\n|---|---|---|---|\n| **storage** | datastore, scsi, latency, vsan, apd, pdl, no space, vmfs, iscsi | vmware-storage, vmware-log-insight (vmkernel scsi/apd) | aiops / pilot |\n| **network** | vmotion, uplink, link down, mtu, firewall, dfw, segment, tier-0, bgp | vmware-nsx, vmware-nsx-security (traceflow, DFW hits) | pilot |\n| **compute** | cpu ready, memory, balloon, swap, contention, numa | vmware-aria (metrics + anomalies) | pilot (rightsizing) |\n| **ha_drs** | ha, high availability, drs, failover, admission control, isolation | vmware-monitor, vmware-aiops (cluster) | pilot |\n| **host_lifecycle** | maintenance mode, shut down of, host reboot, standby mode, lost connection to, cannot synchronize, not responding | vmware-monitor (host/cluster state), vmware-harden (drift — a host that left service on cue was told to), vmware-log-insight (vpxd/hostd) | pilot |\n| **power_lifecycle** | power on/off, failed to start, boot, vmx, ovf, clone, snapshot | vmware-aiops (snapshot tree), vmware-monitor | aiops |\n| **auth** | login, authentication, denied, 401, 403, token, certificate, tls, password | config/.env, target cert + time sync | config fix |\n| **platform** | vpxd, hostd, service restart, crash, 503, not responding, disconnected | vmware-monitor (connection/service), vmware-log-insight (vpxd/hostd) | pilot |\n| **hardware** | tpm, attestation, ipmi, sensor, bmc | vmware-monitor (host sensors and the CIM Server that supplies them, host logs for ipmi/cim), the server's BMC | site / vendor |\n| **licensing** | license | vmware-monitor (license_status — which asset holds which key); an Aria alert on \"Unlicensed Group\" is Aria's own license | license portal |\n| **data_collection** | adapter instance, not receiving data, collector group, cloud proxy | vmware-aria (adapters, collector groups, node health, objects that stopped reporting), vmware-monitor (is the source vCenter reachable) | aria admin |\n\n## Remediation handoff (advisor → executor)\n\nDebug never executes. It mirrors the `vmware-harden → vmware-pilot` pattern:\n\n- **Single, low-risk fix** → call the matching **vmware-aiops** tool (own double-confirm).\n- **Multi-step / approval / cross-skill** → submit the proposed plan to **vmware-pilot**:\n  state machine + approval gate + rollback + audit all live there.\n\n## Adding a new symptom category\n\n1. Add a `(name, keywords, suggested_check)` tuple to `_CATEGORY_SIGNATURES`.\n2. Add the matching row to the table above.\n3. Add a focused unit test in `tests/test_timeline.py` (a sample text → expected category).\n4. Add a playbook under `references/playbooks/` if the investigation is non-obvious.\n\nFile v1.12.2:references/setup-guide.md\n\n# vmware-debug Setup Guide\n\nvmware-debug has **no configuration, no credentials, and no network access** — it\nis a pure, offline correlation engine. There is no `config.yaml` and no `.env`.\n\n## Install\n\n```bash\nuv tool install vmware-debug==1.12.2\nvmware-debug categories      # verify it runs\n```\n\n## MCP client configuration\n\n```json\n{\n  \"command\": \"uvx\",\n  \"args\": [\"--from\", \"vmware-debug==1.12.2\", \"vmware-debug-mcp\"]\n}\n```\n\nIf installed with `uv tool install`, prefer the entry point `vmware-debug mcp`\n(no PyPI resolution at startup — robust behind corporate TLS proxies, 踩坑 #25).\n\nFor full cross-skill diagnosis, also install the data-source skills it correlates\n(vmware-monitor, vmware-log-insight, vmware-aria, vmware-nsx) and the executors it\nroutes fixes to (vmware-aiops, vmware-pilot).\n\n## Security\n\n> **Disclaimer**: Community-maintained open-source project, **not affiliated with,\n> endorsed by, or sponsored by VMware, Inc. or Broadcom Inc.**\n\n1. **Source Code** — https://github.com/vmware-skills/VMware-Debug (MIT).\n2. **Credentials** — none. debug holds no secrets and connects to nothing.\n3. **Network** — none. The correlation tools are local pure functions over\n   event data the agent supplies; the case tools read and write local files.\n4. **Writes** — only to the local investigation ledger under `$OPS_HOME`\n   (the seven [WRITE] `case_*` tools). Evidence, gaps, hypotheses and grade\n   history are only ever added to; `timeline.md` is regenerated from the evidence and\n   `case.json` holds the current grade and state. Nothing is written to any\n   VMware system: debug only diagnoses and recommends; remediation is routed\n   to vmware-aiops / vmware-pilot, where confirmation/approval/audit live.\n   Case data is sensitive: each case folder (`$OPS_HOME/cases/<case-id>/`, default\n   `~/.vmware/cases/`) is created owner-only (`0700`) and holds the submitted\n   evidence — host names, addresses, log and event text. There is no automatic\n   retention limit: `case_close` marks a case closed and keeps its files; delete\n   the folder to remove a case.\n5. **No cross-skill coupling** — events arrive as plain dicts (the event\n   envelope); debug imports no other skill package at runtime.\n6. **Environment scoping** — policy rules can scope by environment, and skills\n   that connect to a VMware estate may declare `environment:` (`production` /\n   `staging` / `lab`) per target in their own `config.yaml` as an optional label\n   an environment-scoped `deny` rule can match on. debug has no config and no\n   connection to declare one about, so it registers no environment resolver —\n   it has no basis to answer for any target — and an environment-scoped rule\n   never matches its tools. Every other policy rule does apply: each MCP tool\n   and the `categories` / `triage` CLI commands pass through vmware-policy's\n   `guard()`. Verified behaviour:\n   - A `deny` rule naming a debug tool, or with no `operations` key at all,\n     refuses it (status `denied`, with the rule's reason).\n   - A malformed `~/.vmware/rules.yaml` fails closed: every tool is refused with\n     a reason naming the file and the way out (fix it, or\n     `VMWARE_POLICY_DISABLED=1`). The `categories` and `triage` commands exit 1\n     with a Python traceback whose last lines are that message; `version` and\n     `mcp` are unaffected.\n   - Runaway guard: within one MCP server process, the 26th call to the same\n     tool with identical arguments inside 120 s is refused (`budget_exceeded`).\n     Tune with `VMWARE_RUNAWAY_MAX` / `VMWARE_RUNAWAY_WINDOW_SEC`. The guard\n     compares a SHA-256 digest of the arguments the tool actually received\n     (vmware-policy ≥1.16.0), not the redacted copy in the audit row — so calls\n     with different `events` or `payload` are never counted as identical.\n   - A `maintenance_window` gates only `high`/`critical` risk operations; every\n     debug tool is `low`, so a window never blocks one.\n7. **Audit row contents** — one row per call in `~/.vmware/audit.db`: tool,\n   time, OS user, detected agent, arguments, status and result. `payload` and\n   `events` are stored as `***`. The results of `incident_timeline` and\n   `case_timeline` are not stored (`[redacted: return value declared\n   sensitive]`) because they quote event text; other results are stored after\n   credential scrubbing. The status is still truthful — a call that returns\n   `{\"error\": …}` is `error`.\n8. **Static analysis** — `uvx bandit -r vmware_debug/` (release bar:\n   0 Medium+).\n\nFile v1.12.2:skill-card.md\n\n## Description:\n\nUse this skill when troubleshooting VMware, vSphere, ESXi, or NSX incidents: it correlates gathered events into timelines, ranks root-cause hypotheses, suggests next checks, and records diagnostic case evidence without writing to VMware systems.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[zw008](https://clawhub.ai/user/zw008)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers, operators, and support engineers use this skill to diagnose VMware incidents from pre-collected events, logs, alerts, and case evidence. It helps decide what to inspect next and routes remediation to other VMware skills instead of executing fixes itself.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: The skill can direct agents toward VMware credentials, configuration files, sensitive logs, and locally persisted case data.\n\nMitigation: Review before installing in environments where those materials may be reachable, use a dedicated account, limit filesystem access, and do not allow secret values to be read unless explicitly necessary.\n\nRisk: Submitted case evidence and audit records may contain host names, addresses, log text, and diagnostic conclusions that remain on disk.\n\nMitigation: Protect or periodically delete ~/.vmware/cases and ~/.vmware/audit.db according to the environment's retention and access-control requirements.\n\nRisk: Runtime package resolution can introduce supply-chain uncertainty for a diagnostic tool used around sensitive infrastructure.\n\nMitigation: Prefer a verified local install over runtime uvx resolution when possible.\n\n## Reference(s):\n\n- [VMware Debug homepage](https://github.com/vmware-skills/VMware-Debug)\n- [agent-guardrails.md](references/agent-guardrails.md)\n- [capabilities.md](references/capabilities.md)\n- [cli-reference.md](references/cli-reference.md)\n- [event-envelope.md](references/event-envelope.md)\n- [routing.md](references/routing.md)\n- [setup-guide.md](references/setup-guide.md)\n- [VMware-AIops issue 31](https://github.com/vmware-skills/VMware-AIops/issues/31)\n\n## Skill Output:\n\n**Output Type(s):** [analysis, markdown, JSON, shell commands, guidance]\n\n**Output Format:** [Markdown guidance with JSON tool outputs and shell command examples]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [Diagnostic outputs may include timelines, ranked hypotheses, next checks, case ledger summaries, and local audit references.]\n\n## Skill Version(s):\n\n1.12.2 (source: server release evidence)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nArchive v1.12.1: 9 files, 21709 bytes\n\nFiles: references/agent-guardrails.md (8989b), references/capabilities.md (3160b), references/cli-reference.md (1585b), references/event-envelope.md (3622b), references/routing.md (2958b), references/setup-guide.md (2888b), skill-card.md (2666b), SKILL.md (16899b), _meta.json (132b)\n\nFile v1.12.1:SKILL.md\n\n---\nname: vmware-debug\ndescription: >\n  Use this skill whenever the user is troubleshooting a VMware/vSphere problem —\n  a reported error, an exception, a log dump, a slow or failed VM, a host that\n  went sideways — and needs help locating the root cause. It is the diagnostic\n  brain of the VMware family: it drives a systematic investigation, pulls the\n  right signals from the other skills, correlates events into one timeline,\n  ranks root-cause hypotheses, and tells you what to check next even when you\n  don't know where to start. Always use this skill for \"diagnose this VMware\n  issue\", \"why is my VM slow\", \"troubleshoot this vSphere error\", \"what does\n  this log mean\", \"help me figure out what broke\" when the context is explicitly\n  VMware/vSphere/ESXi/NSX. It never touches vSphere: its only writes are to a\n  local case ledger. Do NOT\n  use it to execute fixes — single fixes go to vmware-aiops, multi-step gated\n  remediation goes to vmware-pilot. Do NOT use it for routine inventory or\n  health checks with no pr\n\nArchive v1.12.0: 9 files, 21784 bytes\n\nFiles: references/agent-guardrails.md (8989b), references/capabilities.md (3160b), references/cli-reference.md (1585b), references/event-envelope.md (3622b), references/routing.md (2958b), references/setup-guide.md (2888b), skill-card.md (2832b), SKILL.md (16899b), _meta.json (132b)\n\nArchive v1.11.4: 9 files, 21230 bytes\n\nFiles: references/agent-guardrails.md (8989b), references/capabilities.md (2736b), references/cli-reference.md (1585b), references/event-envelope.md (3622b), references/routing.md (2364b), references/setup-guide.md (2888b), skill-card.md (2683b), SKILL.md (16899b), _meta.json (132b)\n\nArchive v1.11.3: 9 files, 20278 bytes\n\nFiles: references/agent-guardrails.md (8772b), references/capabilities.md (2497b), references/cli-reference.md (1473b), references/event-envelope.md (3622b), references/routing.md (2364b), references/setup-guide.md (2057b), skill-card.md (2831b), SKILL.md (15969b), _meta.json (132b)\n\nArchive v1.11.2: 9 files, 20074 bytes\n\nFiles: references/agent-guardrails.md (8772b), references/capabilities.md (2497b), references/cli-reference.md (1473b), references/event-envelope.md (3622b), references/routing.md (2364b), references/setup-guide.md (2057b), skill-card.md (2361b), SKILL.md (15969b), _meta.json (132b)\n\nArchive v1.11.1: 9 files, 19649 bytes\n\nFiles: references/agent-guardrails.md (8772b), references/capabilities.md (2497b), references/cli-reference.md (1473b), references/event-envelope.md (2726b), references/routing.md (2364b), references/setup-guide.md (2057b), skill-card.md (2659b), SKILL.md (15795b), _meta.json (132b)\n\nArchive v1.11.0: 9 files, 19675 bytes\n\nFiles: references/agent-guardrails.md (8772b), references/capabilities.md (2497b), references/cli-reference.md (1473b), references/event-envelope.md (2726b), references/routing.md (2364b), references/setup-guide.md (2057b), skill-card.md (2633b), SKILL.md (15795b), _meta.json (132b)","readmeExcerpt":"Skill: vmware-debug Owner: zw008 Summary: Use this skill whenever the user is troubleshooting a VMware/vSphere problem — a reported error, an exception, a log dump, a slow or failed VM, a host that went sideways — and needs help locating the root cause. It is the diagnostic brain of the VMware family: it drives a systematic investigation, pulls the right signals from the other skills, correlates events into one timel","codeSnippets":[],"executableExamples":[{"language":"bash","snippet":"uv tool install vmware-debug==1.13.1\nvmware-debug categories          # see what it can diagnose"},{"language":"text","snippet":"~/.vmware/cases/<case-id>/\n├── scope.json    what is being investigated, and how that was decided\n├── evidence/     one file per fact: source skill, exact query, time basis\n├── gaps.json     what could NOT be obtained, what it blocks, how to close it\n├── conclusion.md the grade, appended — including every time it went down\n└── timeline.md · hypotheses.md · plan.jsonl · case.json"},{"language":"yaml","snippet":"---\nid: KB-2026-0417\napplies_to:\n  product: vsphere\n  build: \">=8.0.3, <9.0\"\n  driver: {name: nvme_pcie, version: \">=1.2.4\"}\n  firmware: {vendor: dell, version: \">=52.26\"}\n---"},{"language":"bash","snippet":"vmware-debug categories                        # what can it diagnose\nvmware-debug triage --events events.json       # correlate a collected event set\ncat events.json | vmware-debug triage          # or via stdin\nvmware-debug mcp                                # start stdio MCP server (proxy-safe)"},{"language":"text","snippet":"## Tool use\n\n- Gather real events with the companion skills' read tools before calling\n  incident_timeline. Never answer from memory or assumption, and never\n  hand-write events to represent what you believe happened.\n- Never describe a tool call, and never output a JSON example, instead of\n  executing the tool. If you intend to call a tool, call it.\n- If a tool fails, report the actual error text. Do not complete the answer\n  with assumptions about what the result would have been.\n- Use explicit limits and time windows when gathering events. Do not request\n  unlimited results unless the user asks for them.\n\n## Skill routing\n\n- vmware-debug: correlate pre-fetched events into a timeline, detect spikes,\n  rank root-cause hypotheses, suggest next checks.\n- Gather the events from: vmware-monitor (alarms, events, host logs),\n  vmware-log-insight (centralised logs), vmware-aria (metrics, anomalies),\n  vmware-nsx / vmware-nsx-security (network and firewall), vmware-storage.\n- Route the fix to vmware-aiops, or to vmware-pilot when it needs approval\n  gating. This skill executes nothing.\n\n## Building the event set\n\n- Every event needs ts (ISO-8601, epoch seconds, or millis), and the envelope\n  shape {ts, source, severity, entity, text, fields}. An event that cannot be\n  normalised is rejected by index — fix that event, do not drop the batch.\n- Carry each event's original source and severity through unchanged. Do not\n  re-grade a severity to make a story cohere.\n- Pull from more than one source before concluding. A single source's view of\n  an incident is not a correlation.\n- Widen the window before concluding \"no spike\": spike detection needs at least\n  three time bins, so a short window or a large bin_seconds can hide a burst.\n\n## Data fidelity\n\n- Never invent events, timestamps, entities, or relationships between them. If\n  an event was not in the input, it does not exist for this answer.\n- Preserve the exact severity and source values. Do not translate, normalise,\n  or pr"},{"language":"bash","snippet":"vmware-debug triage [OPTIONS]\n  -e, --events PATH     JSON file of event envelopes (reads stdin if omitted)\n      --bin-seconds N   Time-bin width (auto if omitted)\n      --top-n N         Max hypotheses to return   [default: 5]"}],"parameters":null,"dependencies":[],"permissions":[],"extractedFiles":[{"path":"SKILL.md","content":"---\nname: vmware-debug\ndescription: >\n  Use this skill whenever the user is troubleshooting a VMware/vSphere problem —\n  a reported error, an exception, a log dump, a slow or failed VM, a host that\n  went sideways — and needs help locating the root cause. It is the diagnostic\n  brain of the VMware family: it drives a systematic investigation, pulls the\n  right signals from the other skills, correlates events into one timeline,\n  ranks root-cause hypotheses, and tells you what to check next even when you\n  don't know where to start. Always use this skill for \"diagnose this VMware\n  issue\", \"why is my VM slow\", \"troubleshoot this vSphere error\", \"what does\n  this log mean\", \"help me figure out what broke\" when the context is explicitly\n  VMware/vSphere/ESXi/NSX. It never touches vSphere: its only writes are to a\n  local case ledger. Do NOT\n  use it to execute fixes — single fixes go to vmware-aiops, multi-step gated\n  remediation goes to vmware-pilot. Do NOT use it for routine inventory or\n  health checks with no problem to solve — use vmware-monitor.\ninstaller:\n  kind: uv\n  package: vmware-debug\nallowed-tools:\n  - Bash\nmetadata: {\"openclaw\":{\"requires\":{\"anyBins\":[\"vmware-debug\",\"uvx\"]},\"optional\":{\"env\":[\"VMWARE_AUDIT_APPROVED_BY\",\"VMWARE_AUDIT_RATIONALE\"],\"bins\":[\"vmware-policy\"]},\"homepage\":\"https://github.com/vmware-skills/VMware-Debug\",\"os\":[\"macos\",\"linux\"]}}\n---\n\n# VMware Debug\n\n> **Disclaimer**: Community-maintained open-source project, **not affiliated with,\n> endorsed by, or sponsored by VMware, Inc. or Broadcom Inc.** \"VMware\" and \"vSphere\"\n> are trademarks of Broadcom. Source is publicly auditable under the MIT license.\n\nThe diagnostic brain of the VMware skill family. You bring the symptom; this skill\nruns the investigation and points at the root cause. It **reads and reasons** — it\nnever writes to vSphere; its only writes go to its own local case ledger.\nCompanion skills do the data collection and the fixing.\n\n## What This Skill Does\n\n| Category | What | Read or Write |\n|---|---|---|\n| Incident correlation | Merge events from many sources into one timeline, detect spikes | Read |\n| Root-cause ranking | Score symptom clusters, surface the most likely cause first | Read |\n| Next-check ideas | Suggest exactly what to look at next (which skill/tool) when you're stuck | Read |\n| Remediation routing | Hand the fix to vmware-aiops (single) or vmware-pilot (gated, multi-step) | Read (routes only) |\n| Investigation ledger | Open a case; record evidence, gaps and hypotheses; grade and close it | Write (local ledger only) |\n\n**No network access of its own, and no write to any VMware system.** It correlates\ndata the agent has already gathered with the other skills' read tools; its seven\nwrite tools touch only the local case ledger.\n\n## Quick Install\n\n```bash\nuv tool install vmware-debug==1.13.1\nvmware-debug categories          # see what it can diagnose\n```\n\n## When to Use This Skill\n\nUse it when there is a **problem to solve**: an error message"},{"path":"_meta.json","content":"{\n  \"ownerId\": \"kn7b067awq2s97bn3d7p5qfhw5827pxc\",\n  \"slug\": \"vmware-debug\",\n  \"version\": \"1.13.1\",\n  \"publishedAt\": 1789535979714\n}"},{"path":"references/agent-guardrails.md","content":"# Operating vmware-debug with a local / small model\n\nClaude-class models drive this skill without special instruction. Smaller and\nlocally-hosted models — Llama 3.3 70B, Qwen, Mistral, and similar, served\nthrough Goose, Ollama, or OpenShift AI — need explicit operating rules to call\ntools reliably.\n\nThis page exists because an operator wrote those rules by hand first. The\nguardrails below are adapted, with thanks, from the working configuration\n[@juanpf-ha](https://github.com/juanpf-ha) developed while running\nvmware-monitor and vmware-aria against a production vSphere estate with Llama\n3.3 70B FP8 on an on-prem H100\n([VMware-AIops#31](https://github.com/vmware-skills/VMware-AIops/issues/31)). The\ncross-skill rules are identical across this family; the parts below marked\nvmware-debug are specific to this skill.\n\nvmware-debug exposes 14 MCP tools: 7 reads and 7 writes. It connects to nothing\nand holds no credentials — the calling agent gathers events from the other\nskills, normalises them, and hands them over. The seven writes go to the local\ninvestigation ledger under `~/.vmware/cases/`, never to a VMware system. That\nmakes it the safest skill in the family to point a small model at — and the one\nmost exposed to the model's reasoning, because its output *is* an\ninterpretation.\n\nFor a small model the ledger is more than bookkeeping: `case_grade` computes the\nconclusion level from recorded evidence rather than accepting one, so a model\nthat would happily narrate \"root cause confirmed\" cannot record that unless the\nevidence for it is actually in the folder.\n\n> **Disclaimer**: This is a community-maintained open-source project and is\n> **not affiliated with, endorsed by, or sponsored by VMware, Inc. or Broadcom\n> Inc.** \"VMware\" and \"vSphere\" are trademarks of Broadcom.\n\n---\n\n## First: the rules you no longer need to write\n\nSeveral guardrails from the original configuration are now enforced by the\nskill itself. Prompt instructions are advisory — a model can ignore them.\nThese are structural, so it cannot.\n\n| Guardrail you would otherwise prompt for | Now enforced by |\n|---|---|\n| \"Work read-only and never modify anything\" | **The tool surface itself.** No tool reaches a VMware system. The seven [WRITE] tools write only to this skill's own local case ledger, so there is nothing in vSphere to withhold and nothing to switch off. |\n| \"Diagnose only — never apply the fix you propose\" | **Structural.** This skill has no tool that changes anything outside its own ledger, and it holds no connection to vCenter, NSX or anything else. Remediation is routed to vmware-aiops or vmware-pilot by the calling agent. |\n| \"Do not fabricate a timeline — build it from the events I gave you\" | **`incident_timeline` correlates only its input.** It is source-agnostic and has no way to fetch anything, so the timeline cannot contain an event the agent did not supply. |\n| \"Tell me when the symptom is outside what you can recognise\" | **`list_symptom_categories`** states the cat"},{"path":"references/capabilities.md","content":"# vmware-debug Capabilities\n\nOffline incident correlation. No network, no credentials, no writes to any VMware\nsystem. This table covers the two stateless correlation tools; the twelve `case_*`\ninvestigation-ledger tools (seven of which write, to the local ledger only) are\nlisted in `SKILL.md`, and their response sizes are not yet measured here.\n\n| Tool | What it returns | Typical response tokens |\n|---|---|---|\n| `incident_timeline` | `{event_count, window, spikes:[{start,end,count,zscore}], hypotheses:[{category, score, summary, evidence_count, first_seen, last_seen, sample_text, suggested_check}], next_checks:[...]}` | 300–2000 (scales with hypotheses) |\n| `list_symptom_categories` | `{items: [{category, example_keywords, suggested_check}], returned, limit, total, truncated, hint}` | ~400 |\n\n`list_symptom_categories` returns the family list envelope — read the rows from\n`items`. It has no `limit` parameter, which is exactly why the envelope matters:\n`truncated: false` states that this is every category there is, rather than\nleaving a model to guess whether it is holding page one. The catalogue is a\nfixed in-process constant, so `total` is a real count and `limit` is `null`.\n\n## Correlation engine\n\n- **Timeline**: events normalised to the unified envelope, sorted, and time-binned\n  (auto bin width ≈ span/30, or caller-specified).\n- **Spike detection**: z-score over bin counts (≥3 bins required for a baseline;\n  flat series yields no false spikes).\n- **Hypothesis ranking**: events clustered by symptom category (keyword match on\n  text + entity), scored by summed severity weight, tie-broken by recency.\n  Uncategorised events are kept visible, not dropped.\n- **Next-check routing**: each category carries a concrete \"which skill/tool to run\n  next\" suggestion — the value when the user doesn't know what to check.\n\n## Symptom categories\n\n`storage`, `network`, `compute`, `ha_drs`, `host_lifecycle`, `power_lifecycle`,\n`auth`, `platform`, `hardware`, `licensing`, `data_collection`.\nSee `references/routing.md` for keyword signatures and the skill each routes to.\n\n`hardware`, `licensing` and `data_collection` were added after real alert titles\n(\"Host TPM attestation alarm\", \"License will soon expire\", \"Objects are not\nreceiving data from adapter instance\") matched no category at all. A roll-up\nsuch as \"Group population health is degraded\" is deliberately left\nuncategorized: it names no subsystem, and the cause is in one of its members.\n\n`host_lifecycle` is a host changing its own availability state — maintenance\nmode, shutdown, reboot, standby, connection loss, sync failure.\n`power_lifecycle` is the VM-level equivalent. They are separate because they are\nseparate investigations: the second is a task question for vmware-aiops, the\nfirst is a cluster, DPM, vLCM or drift question.\n\n## Design properties\n\n- **Zero cross-skill runtime deps** — correlation is pure functions over plain\n  dicts; the agent fans out to other skills' read tools (踩坑 #21/#32).\n- **JSON-"},{"path":"references/cli-reference.md","content":"# vmware-debug CLI Reference\n\nAll commands are offline (no network, no credentials). All but `mcp` are\nread-only; `mcp` starts the MCP server, whose seven `case_*` write tools record\ninto the local case ledger.\n\n## triage — correlate a set of collected events\n\n```bash\nvmware-debug triage [OPTIONS]\n  -e, --events PATH     JSON file of event envelopes (reads stdin if omitted)\n      --bin-seconds N   Time-bin width (auto if omitted)\n      --top-n N         Max hypotheses to return   [default: 5]\n```\n\nInput is a JSON array of event envelopes (see `references/event-envelope.md`):\n\n```bash\ncat events.json | vmware-debug triage\nvmware-debug triage --events events.json --top-n 3\n```\n\nOutput (JSON): `{event_count, window, spikes, hypotheses, next_checks}`.\n\n## categories — list recognised symptom categories\n\n```bash\nvmware-debug categories\n```\n\nPrints each category, sample keywords, and the suggested next check (which\nskill/tool to run). Use when you don't know what to look at.\n\n## version / mcp\n\n```bash\nvmware-debug version    # installed version\nvmware-debug mcp        # start the stdio MCP server (no network at startup)\n```\n\n## How the agent uses it\n\nIn an agent, the cross-skill correlation happens at the agent layer:\n\n1. Fetch events with the data-source skills (vmware-monitor `event_list`,\n   vmware-log-insight `log_search`/`log_aggregate`, vmware-aria alerts/anomaly,\n   vmware-nsx).\n2. Normalise each into the event envelope.\n3. Call the `incident_timeline` MCP tool to correlate and rank.\n4. Follow `next_checks`; route any fix to vmware-aiops / vmware-pilot."}],"languages":[],"docsSourceLabel":"CLAWHUB","editorialOverview":null,"editorialQuality":{"score":100,"threshold":65,"status":"thin","wordCount":2202,"uniquenessScore":41,"reasons":["uniqueness-below-45"]}},"media":{"evidence":{"source":"no-media","verified":false,"confidence":"low","updatedAt":"2026-10-10T06:37:53.811Z","emptyReason":"No screenshots, media assets, or demo links are available."},"primaryImageUrl":null,"mediaAssetCount":0,"assets":[],"demoUrl":null},"ownerResources":{"evidence":{"source":"unclaimed","verified":false,"confidence":"low","updatedAt":"2026-10-10T06:37:53.811Z","emptyReason":"This page has not been claimed by the agent owner."},"hasCustomPage":false,"customPageUpdatedAt":null,"customLinks":[],"structuredLinks":{"docsUrl":null,"demoUrl":null,"supportUrl":null,"pricingUrl":null,"statusUrl":null},"customPage":null},"relatedAgents":{"evidence":{"source":"protocol-neighbors","verified":false,"confidence":"medium","updatedAt":"2026-10-10T10:42:31.798Z","emptyReason":null},"items":[{"id":"8ebccd8e-3863-4187-8355-c3f14e1f9edf","entityType":"agent","canonicalPath":"/agent/iofficeai-aionui","slug":"iofficeai-aionui","name":"AionUi","description":"Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!","url":"https://github.com/iOfficeAI/AionUi","homepage":"https://www.aionui.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-10-09T19:11:12.944Z","createdAt":"2026-02-25T03:38:16.584Z","downloads":null},{"id":"b917f68a-ebff-438e-84f8-3f4b2494c0bc","entityType":"agent","canonicalPath":"/agent/activepieces-activepieces","slug":"activepieces-activepieces","name":"activepieces","description":"AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents","url":"https://github.com/activepieces/activepieces","homepage":"https://www.activepieces.com","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-15T02:22:12.426Z","createdAt":"2026-02-25T03:38:12.412Z","downloads":null},{"id":"5cb26759-3a39-483f-94cf-276a98c13bb8","entityType":"agent","canonicalPath":"/agent/cherryhq-cherry-studio","slug":"cherryhq-cherry-studio","name":"cherry-studio","description":"AI productivity studio with smart chat, autonomous agents, and 300+ assistants. Unified access to frontier LLMs","url":"https://github.com/CherryHQ/cherry-studio","homepage":"https://cherry-ai.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-11T14:38:40.986Z","createdAt":"2026-02-25T03:38:19.379Z","downloads":null},{"id":"6f6582d0-5d76-4f0f-b81d-86520247950b","entityType":"agent","canonicalPath":"/agent/copilotkit-copilotkit","slug":"copilotkit-copilotkit","name":"CopilotKit","description":"The Frontend for Agents & Generative UI. React + Angular","url":"https://github.com/CopilotKit/CopilotKit","homepage":"https://docs.copilotkit.ai","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-03-25T09:50:57.846Z","createdAt":"2026-02-25T03:39:14.617Z","downloads":null}],"links":{"hub":"/agent","source":"/agent/source/clawhub","protocols":[{"label":"OpenClaw","href":"/agent/protocol/openclew"}]}}}