{"id":"fc583fa6-384c-4695-91a8-477e298498f1","entityType":"agent","slug":"clawhub-zw008-vmware-monitor","name":"vmware-monitor","canonicalUrl":"https://www.xpersona.co/agent/clawhub-zw008-vmware-monitor","canonicalPath":"/agent/clawhub-zw008-vmware-monitor","generatedAt":"2026-10-10T05:25:33.687Z","source":"CLAWHUB","claimStatus":"UNCLAIMED","verificationTier":"NONE","summary":{"evidence":{"source":"editorial-content","verified":true,"confidence":"high","updatedAt":"2026-10-09T03:41:48.278Z","emptyReason":null},"description":"Use this skill for safe, read-only queries of VMware infrastructure — no destructive operations exist in the codebase; a test enforces it. Directly handles: a one-glance cross-cluster health summary, object-centered VM/host/datastore investigation drill-downs (correlating surrounding infrastructure + recent events), a cross-vCenter \"what needs attention now?\" rollup, list VMs/hosts/datastores/clusters, active alarms, recent events, VM details. Always use vmware-monitor when the user asks to \"list VMs\", \"check vSphere alarms\", \"show host status\", \"is anything on fire\", \"what needs attention now\", \"what is happening around this VM/host/datastore\", \"investigate this VM\" — or needs read-only VMware info before making changes. Do NOT use for any write operations — this skill is read-only and has no code path that creates, modifies, or deletes a vSphere resource. For VM modifications use vmware-aiops, for networking use vmware-nsx, for metrics/capacity use vmware-aria. For load balancing/AVI/AKO use vmware-avi. Skill: vmware-monitor Owner: zw008 Summary: Use this skill for safe, read-only queries of VMware infrastructure — no destructive operations exist in the codebase; a test enforces it. Directly handles: a one-glance cross-cluster health summary, object-centered VM/host/datastore investigation drill-downs (correlating surrounding infrastructure + recent events), a cross-vCenter \"what needs attention now?\" rollup, list V","descriptionLabel":"Technical summary","evidenceSummary":"Capability contract not published. No trust telemetry is available yet. 5.9K downloads reported by the source. Last updated 10/9/2026.","installCommand":"clawhub skill install s171xgnmqse0nqvgqvqnaq5f9183kyre:vmware-monitor","sourceUrl":"https://clawhub.ai/zw008/vmware-monitor","homepage":"https://clawhub.ai/zw008/skills/vmware-monitor","primaryLinks":[{"label":"View on ClawHub","url":"https://clawhub.ai/zw008/vmware-monitor","kind":"source"},{"label":"Homepage","url":"https://clawhub.ai/zw008/skills/vmware-monitor","kind":"homepage"}],"safetyScore":84,"overallRank":62,"popularityScore":58,"trustScore":null,"claimedByName":null,"isOwner":false,"seoDescription":"Use this skill for safe, read-only queries of VMware infrastructure — no destructive operations exist in the codebase; a test enforces it. Directly handles: a o"},"coverage":{"evidence":{"source":"public-profile","verified":false,"confidence":"medium","updatedAt":"2026-10-09T03:41:48.278Z","emptyReason":null},"protocols":[{"protocol":"OPENCLEW","label":"OpenClaw","status":"self-declared","notes":"Declared in the public agent profile."}],"capabilities":[],"verifiedCount":0,"selfDeclaredCount":1,"capabilityMatrix":{"rows":[{"key":"OPENCLEW","type":"protocol","support":"unknown","confidenceSource":"profile","notes":"Listed on profile"}],"flattenedTokens":"protocol:OPENCLEW|unknown|profile"}},"adoption":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-09T03:41:48.278Z","emptyReason":null},"stars":null,"forks":null,"downloads":5866,"packageName":null,"latestVersion":"1.16.0","tractionLabel":"5.9K downloads"},"release":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-09T03:41:48.278Z","emptyReason":null},"lastUpdatedAt":"2026-10-09T03:41:48.278Z","lastCrawledAt":"2026-10-09T03:41:48.278Z","lastIndexedAt":null,"nextCrawlAt":"2026-10-10T03:41:48.278Z","lastVerifiedAt":null,"highlights":[{"version":"1.16.0","createdAt":"2026-09-20T14:51:28.294Z","changelog":"MCP instructions now name the configured targets and how to choose one; a config that cannot be read says so instead of falling silent.","fileCount":10,"zipByteSize":44534},{"version":"1.15.1","createdAt":"2026-09-16T05:18:43.515Z","changelog":"A stop signal always ends the server (bounded logout); active_tasks prints its unreadable-task note beside the table","fileCount":10,"zipByteSize":44620},{"version":"1.15.0","createdAt":"2026-09-15T14:36:54.728Z","changelog":"Results name the target that answered (default_target, target field, instructions list targets); active_tasks counts unreadable tasks; Stopping the MCP server now logs out its vCenter session.","fileCount":10,"zipByteSize":44532},{"version":"1.14.0","createdAt":"2026-09-15T08:53:58.214Z","changelog":"Stale alarms re-checked (license alarm cleared by evidence; unknown keeps severity), appliance alarms labelled, attention counts overlapping ESXi/vCenter hosts once, capacity views explained, perf vms balloon/swap, deployment-size without traceback.","fileCount":10,"zipByteSize":44180},{"version":"1.13.1","createdAt":"2026-09-15T05:59:05.408Z","changelog":"CLI reads are audited under their MCP tool names; every CLI command declares what it reaches (needs vmware-policy 1.15.0)","fileCount":10,"zipByteSize":42080},{"version":"1.13.0","createdAt":"2026-09-15T03:13:05.996Z","changelog":"health sensors say why a host has none; NTP source mismatch; over-committed datastores in the summary","fileCount":10,"zipByteSize":41818},{"version":"1.12.0","createdAt":"2026-09-14T15:12:20.640Z","changelog":"Stale-alarm verdicts, newest-first event reads with windows and routine folding, real event names ranked, grouped host logs, sessions and license assignments.","fileCount":10,"zipByteSize":41576},{"version":"1.11.3","createdAt":"2026-09-12T00:11:10.763Z","changelog":"host_log_scan reads host logs for the first time (0 findings before, 131 live after); unreadable logs are reported instead of looking clean; the daemon reports each line once and only critical host-log lines page the webhook; standalone hosts are no longer counted as clusters.","fileCount":10,"zipByteSize":40270}]},"execution":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No published capability contract is available yet."},"installCommand":"clawhub skill install s171xgnmqse0nqvgqvqnaq5f9183kyre:vmware-monitor","setupComplexity":"low","setupSteps":["Setup complexity is classified as HIGH. You must provision dedicated cloud infrastructure or an isolated VM. Do not run this directly on your local workstation.","Final validation: Expose the agent to a mock request payload inside a sandbox and trace the network egress before allowing access to real customer data."],"contract":{"contractStatus":"missing","authModes":[],"requires":[],"forbidden":[],"supportsMcp":false,"supportsA2a":false,"supportsStreaming":false,"inputSchemaRef":null,"outputSchemaRef":null,"dataRegion":null,"contractUpdatedAt":null,"sourceUpdatedAt":null,"freshnessSeconds":null},"invocationGuide":{"preferredApi":{"snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-vmware-monitor/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-vmware-monitor/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-vmware-monitor/trust"},"curlExamples":["curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-zw008-vmware-monitor/snapshot\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-zw008-vmware-monitor/contract\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-zw008-vmware-monitor/trust\""],"jsonRequestTemplate":{"query":"summarize this repo","constraints":{"maxLatencyMs":2000,"protocolPreference":["OPENCLEW"]}},"jsonResponseTemplate":{"ok":true,"result":{"summary":"...","confidence":0.9},"meta":{"source":"CLAWHUB","generatedAt":"2026-10-10T05:25:33.683Z"}},"retryPolicy":{"maxAttempts":3,"backoffMs":[500,1500,3500],"retryableConditions":["HTTP_429","HTTP_503","NETWORK_TIMEOUT"]}},"endpoints":{"dossierUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-vmware-monitor/dossier","snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-vmware-monitor/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-vmware-monitor/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-vmware-monitor/trust"}},"reliability":{"evidence":{"source":"runtime-metrics","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No trust, reliability, or runtime telemetry is available."},"trust":{"status":"unavailable","handshakeStatus":"UNKNOWN","verificationFreshnessHours":null,"reputationScore":null,"p95LatencyMs":null,"successRate30d":null,"fallbackRate":null,"attempts30d":null,"trustUpdatedAt":null,"trustConfidence":"unknown","sourceUpdatedAt":null,"freshnessSeconds":null},"decisionGuardrails":{"doNotUseIf":["Contract metadata is missing or unavailable for deterministic execution."],"safeUseWhen":[],"riskFlags":["missing_or_unavailable_contract","trust_data_unavailable","schema_references_missing"],"operationalConfidence":"low"},"executionMetrics":{"observedLatencyMsP50":null,"observedLatencyMsP95":null,"estimatedCostUsd":null,"uptime30d":null,"rateLimitRpm":null,"rateLimitBurst":null,"lastVerifiedAt":null,"verificationSource":null},"runtimeMetrics":{"successRate":null,"avgLatencyMs":null,"avgCostUsd":null,"hallucinationRate":null,"retryRate":null,"disputeRate":null,"p50Latency":null,"p95Latency":null,"lastUpdated":null}},"benchmarks":{"evidence":{"source":"no-benchmark-data","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No benchmark suites or observed failure patterns are available."},"suites":[],"failurePatterns":[]},"artifacts":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"high","updatedAt":"2026-10-09T03:41:48.278Z","emptyReason":null},"readme":"Skill: vmware-monitor\n\nOwner: zw008\n\nSummary: Use this skill for safe, read-only queries of VMware infrastructure — no destructive operations exist in the codebase; a test enforces it. Directly handles: a one-glance cross-cluster health summary, object-centered VM/host/datastore investigation drill-downs (correlating surrounding infrastructure + recent events), a cross-vCenter \"what needs attention now?\" rollup, list VMs/hosts/datastores/clusters, active alarms, recent events, VM details. Always use vmware-monitor when the user asks to \"list VMs\", \"check vSphere alarms\", \"show host status\", \"is anything on fire\", \"what needs attention now\", \"what is happening around this VM/host/datastore\", \"investigate this VM\" — or needs read-only VMware info before making changes. Do NOT use for any write operations — this skill is read-only and has no code path that creates, modifies, or deletes a vSphere resource. For VM modifications use vmware-aiops, for networking use vmware-nsx, for metrics/capacity use vmware-aria. For load balancing/AVI/AKO use vmware-avi.\n\nTags: esxi:1.16.0, latest:1.16.0, monitoring:1.16.0, read-only:1.16.0, safety:1.16.0, vcenter:1.16.0, vmware:1.16.0\n\nVersion history:\n\nv1.16.0 | 2026-09-20T14:51:28.294Z | user\n\nMCP instructions now name the configured targets and how to choose one; a config that cannot be read says so instead of falling silent.\n\nv1.15.1 | 2026-09-16T05:18:43.515Z | user\n\nA stop signal always ends the server (bounded logout); active_tasks prints its unreadable-task note beside the table\n\nv1.15.0 | 2026-09-15T14:36:54.728Z | user\n\nResults name the target that answered (default_target, target field, instructions list targets); active_tasks counts unreadable tasks; Stopping the MCP server now logs out its vCenter session.\n\nv1.14.0 | 2026-09-15T08:53:58.214Z | user\n\nStale alarms re-checked (license alarm cleared by evidence; unknown keeps severity), appliance alarms labelled, attention counts overlapping ESXi/vCenter hosts once, capacity views explained, perf vms balloon/swap, deployment-size without traceback.\n\nv1.13.1 | 2026-09-15T05:59:05.408Z | user\n\nCLI reads are audited under their MCP tool names; every CLI command declares what it reaches (needs vmware-policy 1.15.0)\n\nv1.13.0 | 2026-09-15T03:13:05.996Z | user\n\nhealth sensors say why a host has none; NTP source mismatch; over-committed datastores in the summary\n\nv1.12.0 | 2026-09-14T15:12:20.640Z | user\n\nStale-alarm verdicts, newest-first event reads with windows and routine folding, real event names ranked, grouped host logs, sessions and license assignments.\n\nv1.11.3 | 2026-09-12T00:11:10.763Z | user\n\nhost_log_scan reads host logs for the first time (0 findings before, 131 live after); unreadable logs are reported instead of looking clean; the daemon reports each line once and only critical host-log lines page the webhook; standalone hosts are no longer counted as clusters.\n\nv1.11.2 | 2026-09-05T01:05:54.463Z | user\n\na dropped connection no longer keeps itself alive\n\nv1.11.1 | 2026-09-03T08:51:15.114Z | user\n\nthe window this does not cover, said out loud\n\nv1.11.0 | 2026-09-03T02:03:04.895Z | user\n\nan open backup is not a failed one\n\nv1.10.0 | 2026-08-31T07:21:17.130Z | user\n\n_is_cluster used a leaf type name pyVmomi never reports, so every bundle lost its cluster context; event catalogue recovers 1887 extended ids; bundles no longer crash on standalone ESXi; vLCM stops counting unscannable hosts as non-compliant\n\nv1.9.2 | 2026-08-31T01:17:31.474Z | user\n\nFirst live run of the issue #26 backup-window tool found the coverage note blaming vpxd.task.maxAge, a name that raises vim.fault.InvalidName; the option is task.maxAge and the note now reads its real value.\n\nv1.9.1 | 2026-08-31T00:31:50.949Z | user\n\nfix: run the suite on a non-UTF-8 machine, and stop one skill answering for another\n\nv1.9.0 | 2026-08-30T15:52:42.018Z | user\n\nSee RELEASE_NOTES.md for v1.9.0\n\nv1.8.15 | 2026-08-30T15:17:30.283Z | user\n\nSecond-round fixes from the 2026-08-30 VCF 9.1 re-test; vmware-policy floor raised to 1.11.0 (the engine no longer fails open when rules.yaml cannot be read).\n\nv1.8.14 | 2026-08-30T09:33:52.676Z | user\n\nRead-only enforcement replaced a 33-string grep with an AST allowlist derived from pyVmomi's own metadata. Parameter descriptions now reach the MCP JSON schema (0% -> 100% coverage); additionalProperties closed; vmware-policy floor raised to 1.10.0.\n\nv1.8.13 | 2026-08-30T07:37:54.752Z | user\n\nUnreachable hosts no longer reported as misconfigured; event severity from vCenter's own catalogue; doctor authenticates every target; one config path across CLI/doctor/MCP; server.json starts the MCP server.\n\nv1.8.12 | 2026-08-29T15:18:20.031Z | user\n\nStandalone ESXi reported 0 VMs and 0% usage; vCenter-level alarms were invisible to the health summary; the documented SSL config key was one the code never read. Found on real hardware.\n\nv1.8.11 | 2026-08-28T02:55:19.187Z | user\n\nFixes the server's self-reported version and the advertised tool count; adds a Claude Code plugin manifest.\n\nv1.8.10 | 2026-08-06T09:43:56.096Z | user\n\nvSphere 9.1 memory tiering + vLCM patch compliance + deployment size — 4 read tools (27→31). + review hardening (CLI 503-tolerance, compliance count). Beta caveats in RELEASE_NOTES.\n\nv1.8.9 | 2026-08-01T03:11:25.082Z | user\n\nMoved to vmware-skills GitHub org; MCP Registry namespace → io.github.vmware-skills. Links updated.\n\nv1.8.8 | 2026-07-21T15:46:21.465Z | user\n\nCLI writes now route through the shared guard()+audit_call() core via @guarded, exactly like the MCP tools (HLD I-1/I-8). Requires vmware-policy>=1.8.8.\n\nv1.8.7 | 2026-07-21T11:37:40.201Z | user\n\nRemove read-only switch and approval tiers; read/write authz delegated to RBAC. Plus accumulated fixes since 1.8.5.\n\nv1.8.5 | 2026-07-20T13:02:20.347Z | user\n\nA failure that is returned is now audited as a failure, and certificate/URL detail no longer reaches the agent. Both fixes v1.8.4 announced were incomplete.\n\nv1.8.4 | 2026-07-20T08:24:17.782Z | user\n\nTeaching error messages, domain exceptions no longer redacted on the way to the agent, and tool descriptions that state when to use each tool and what to call next.\n\nv1.8.3 | 2026-07-20T03:38:23.080Z | user\n\nPer-target username can now come from an env var, resolved per access like the password; documented credential variables corrected against what each repo's code actually reads\n\nv1.8.2 | 2026-07-19T17:58:11.380Z | user\n\nMCP server moved into the package namespace — fixes two skills in one environment silently overwriting each other's server; agent-guardrails.md for local/small models now ships in every skill\n\nv1.8.1 | 2026-07-19T11:20:39.041Z | user\n\nRead-only mode now documented on every surface that teaches it (SKILL.md, setup-guide, capabilities) and reported by doctor\n\nv1.8.0 | 2026-07-19T09:37:30.067Z | user\n\nRead-only mode (structural gate, 27/27 tools verified read at start-up), list-result envelope, declared environments; tool-count and tool-name doc corrections\n\nv1.7.7 | 2026-07-17T06:55:58.741Z | user\n\nSession-probe eviction fix (dead cached sessions were never evicted; None currentSession now treated as dead) + lockfile mcp 1.28.1 clearing three GHSA HIGH advisories.\n\nv1.7.6 | 2026-07-14T08:55:16.937Z | user\n\nObject investigation bundles (vm/host/datastore) + cross-vCenter attention: correlated drill-down with event timeline + offline HTML. MCP 23→27.\n\nv1.7.5 | 2026-07-13T07:16:45.385Z | user\n\ncluster_health_summary tool (22→23): one-glance cross-cluster triage with ranked top-N issues + offline timestamped HTML snapshots (summary --html); editable display template\n\nv1.7.4 | 2026-07-13T04:52:00.945Z | user\n\nBatch host-check boundary reads (NTP/services/certificate) via new _collect_objects helper — per-host serviceInfo/certificateInfo lazy reads collapse to one RetrievePropertiesEx (issue #31 tail).\n\nv1.7.3 | 2026-07-03T00:48:22.140Z | user\n\nAdd host_log_scan MCP tool (issue #31 request); 21->22 tools\n\nv1.7.2 | 2026-07-02T14:30:54.913Z | user\n\nBatch all monitoring read paths via shared PropertyCollector helper\n\nv1.7.1 | 2026-07-02T10:52:35.952Z | user\n\nLarge-inventory scale fix (issue #31): PropertyCollector batching replaces per-object lazy SOAP round-trips.\n\nv1.7.0 | 2026-06-27T01:00:47.728Z | user\n\n10 new read-only observability tools (MCP 11->21) + guided init wizard\n\nv1.6.1 | 2026-06-24T00:00:40.518Z | user\n\nv1.6.1 .env password b64 obfuscation\n\nv1.6.0 | 2026-06-22T09:17:05.834Z | user\n\nv1.6.0 trust architecture: undo tokens + governance harness (budget/audit/risk-tiers)\n\nv1.5.39 | 2026-06-22T00:42:10.446Z | user\n\nv1.5.39: AIops snapshot-delete async + honest timeout (token-burn fix), Storage browse timeout fix; others version-aligned\n\nv1.5.38 | 2026-06-12T06:59:21.937Z | user\n\nrelease alignment\n\nv1.5.37 | 2026-06-12T01:58:15.292Z | user\n\nbacklog: wire up advertised health tools\n\nv1.5.36 | 2026-06-11T23:21:43.994Z | user\n\nMCP error-shape fix, resilient daemon, no false \"all clear\"\n\nv1.5.35 | 2026-06-10T00:44:58.980Z | user\n\nSecurity hardening: safe error handling, TLS/path/permission fixes\n\nv1.5.32 | 2026-06-08T02:47:14.191Z | user\n\nv1.5.32: sensor health fix + vm_list_snapshots MCP tool added (8 tools)\n\nv1.5.30 | 2026-06-07T13:22:47.850Z | user\n\nv1.5.30: family version alignment, no functional changes\n\nv1.5.29 | 2026-05-29T02:19:46.655Z | user\n\nNEW: --folder-filter CLI flag (CLI/MCP parity, 踩坑 #34); 7 new unit tests; doc cleanup\n\nv1.5.28 | 2026-05-20T10:00:03.734Z | user\n\nFix subclass() arg 1 must be a class in goose/old-mcp environments. v1.5.25-1.5.27 only addressed PEP 604 X|None -> Optional[X] but kept 'from __future__ import annotations'; under mcp 1.10-1.13 FastMCP's issubclass() on string annotations crashed server load. This release removes the future import. CLAUDE.md pitfall #33 updated.\n\nv1.5.27 | 2026-05-20T06:57:03.714Z | user\n\nLoosen Python requirement to >= 3.10 (was >=3.11). v1.5.25/26 PEP 604 fix already enables 3.10 at runtime; this release lifts pip download/install block.\n\nArchive index:\n\nArchive v1.16.0: 10 files, 44534 bytes\n\nFiles: evals/evals.json (1251b), references/agent-guardrails.md (6782b), references/capabilities.md (20921b), references/cli-reference.md (12083b), references/health-summary-template.md (9123b), references/investigation-protocol.md (6758b), references/setup-guide.md (14729b), skill-card.md (2899b), SKILL.md (25005b), _meta.json (134b)\n\nFile v1.16.0:SKILL.md\n\n---\nname: vmware-monitor\ndescription: >\n  Use this skill for safe, read-only queries of VMware infrastructure — no destructive operations exist in the codebase; a test enforces it.\n  Directly handles: a one-glance cross-cluster health summary, object-centered VM/host/datastore investigation drill-downs (correlating surrounding infrastructure + recent events), a cross-vCenter \"what needs attention now?\" rollup, list VMs/hosts/datastores/clusters, active alarms, recent events, VM details.\n  Always use vmware-monitor when the user asks to \"list VMs\", \"check vSphere alarms\", \"show host status\", \"is anything on fire\", \"what needs attention now\", \"what is happening around this VM/host/datastore\", \"investigate this VM\" — or needs read-only VMware info before making changes.\n  Do NOT use for any write operations — this skill is read-only and has no code path that creates, modifies, or deletes a vSphere resource.\n  For VM modifications use vmware-aiops, for networking use vmware-nsx, for metrics/capacity use vmware-aria. For load balancing/AVI/AKO use vmware-avi.\ninstaller:\n  kind: uv\n  package: vmware-monitor\nallowed-tools:\n  - Bash\nmetadata: {\"openclaw\":{\"requires\":{\"anyBins\":[\"vmware-monitor\",\"uvx\"]},\"optional\":{\"env\":[\"VMWARE_MONITOR_CONFIG\",\"VMWARE_TARGET_PASSWORD\",\"VMWARE_<TARGET>_USERNAME\",\"SLACK_WEBHOOK_URL\",\"DISCORD_WEBHOOK_URL\",\"VMWARE_AUDIT_APPROVED_BY\"],\"bins\":[\"vmware-policy\"]},\"homepage\":\"https://github.com/vmware-skills/VMware-Monitor\",\"emoji\":\"📊\",\"os\":[\"macos\",\"linux\"]}}\ncompatibility: >\n  vmware-policy auto-installed as Python dependency (provides @vmware_tool decorator and audit logging). MCP tool calls and remote CLI commands audited to ~/.vmware/audit.db; CLI queries also to ~/.vmware-monitor/audit.log.\n  Credentials: Each vCenter/ESXi target requires a per-target password env var in ~/.vmware-monitor/.env following the pattern VMWARE_<TARGET_NAME_UPPER>_PASSWORD (e.g., target \"vcenter-prod\" → VMWARE_VCENTER_PROD_PASSWORD). SLACK_WEBHOOK_URL and DISCORD_WEBHOOK_URL are optional — disabled by default, user-configured only, used solely by the opt-in daemon scanner. Daemon: the background scanner (vmware-monitor daemon start) is user-initiated only, never auto-started. Webhook payloads carry issue counts plus every critical issue and every alarm/event warning (host-log warnings and info rows are not sent): entity name and the sanitized, truncated alarm, vCenter event, or ESXi log text, or a connection error — which can include host names, IPs, and user names. No credentials from the skill's config are sent.\n---\n\n# VMware Monitor (Read-Only)\n\n> **Disclaimer**: This is a community-maintained open-source project and is **not affiliated with, endorsed by, or sponsored by VMware, Inc. or Broadcom Inc.** \"VMware\" and \"vSphere\" are trademarks of Broadcom. Source code is publicly auditable at [github.com/vmware-skills/VMware-Monitor](https://github.com/vmware-skills/VMware-Monitor) under the MIT license.\n\nRead-only VMware vCenter/ESXi monitoring — 32 MCP tools, zero destructive code.\n\n> **Read-only toward vSphere**: no code path changes vCenter/ESXi state — no power, create, delete, snapshot, or reconfigure call exists. On vCenter it opens only its login session and short-lived query handles it releases. Gate: [`tests/eval/regression/test_read_only_enforcement.py`](https://github.com/vmware-skills/VMware-Monitor/blob/main/tests/eval/regression/test_read_only_enforcement.py) (source repo, not this bundle) requires every vSphere method called to be on a reviewed allowlist, checked against pyVmomi's type metadata. It checks source, not runtime, and no CI runs it. Independent of this code: use a dedicated account with vCenter's Read-Only role.\n> **Local writes** (this machine only): `init` writes `~/.vmware-monitor/config.yaml` and `.env` (0600; plaintext passwords are rewritten as `b64:` on load); audit logs `~/.vmware/audit.db` (MCP and CLI) and `~/.vmware-monitor/audit.log` (CLI); `--html` snapshots in `~/vmware-health/`; after `daemon start` only, `scan.log`, `daemon.pid` and opt-in webhook posts.\n> **Companion skills**: [vmware-aiops](https://github.com/vmware-skills/VMware-AIops) (VM lifecycle), [vmware-storage](https://github.com/vmware-skills/VMware-Storage) (iSCSI/vSAN), [vmware-vks](https://github.com/vmware-skills/VMware-VKS) (Tanzu Kubernetes), [vmware-nsx](https://github.com/vmware-skills/VMware-NSX) (NSX networking), [vmware-nsx-security](https://github.com/vmware-skills/VMware-NSX-Security) (DFW/firewall), [vmware-aria](https://github.com/vmware-skills/VMware-Aria) (metrics/alerts/capacity), [vmware-avi](https://github.com/vmware-skills/VMware-AVI) (AVI/ALB/AKO), [vmware-harden](https://github.com/vmware-skills/VMware-Harden) (compliance baselines).\n> | [vmware-pilot](../vmware-pilot/SKILL.md) (workflow orchestration) | [vmware-policy](../vmware-policy/SKILL.md) (audit/policy)\n\n## What This Skill Does\n\n| Category | Capabilities |\n|----------|-------------|\n| **Cluster Triage** | One-glance `cluster_health_summary` — cross-cluster Problems/Capacity/Health rollup with an opinionated status; customizable view |\n| **Object Investigation** | \"What is happening around this VM / host / datastore?\" — one correlated drill-down bundle per object, plus `cross_vcenter_attention` — one ranked \"what needs attention now?\" list across every configured vCenter |\n| **Inventory** | List VMs, ESXi hosts, datastores, clusters, networks |\n| **Health** | Active alarms, recent events (filter by severity/time), hardware sensors, host services |\n| **Performance** | Real-time host & VM CPU/memory/disk/network utilisation (PerfManager) |\n| **Capacity** | Datastore thin-provisioning over-commit, resource-pool reservation/usage |\n| **Infra Health** | ESXi certificate expiry, license usage/expiry, NTP configuration health |\n| **Snapshots** | Inventory-wide snapshot aging & sprawl (flag old snapshots) |\n| **Activity** | In-flight tasks, active login sessions |\n| **VM Details** | CPU, memory, disks, NICs, snapshots, guest OS, IP |\n| **Scanning** | Scheduled alarm/log scanning with Slack/Discord webhooks |\n| **vSphere 9.1** | Host memory tiering (DRAM/NVMe uplift), vLCM cluster patch compliance & last-apply result, vCenter deployment size |\n\n## Quick Install\n\n```bash\nuv tool install vmware-monitor==1.16.0\nvmware-monitor doctor\n```\n\n## When to Use This Skill\n\n- List or search VMs, hosts, datastores, clusters\n- Check active alarms or recent events\n- Get detailed info about a specific VM\n- Set up scheduled monitoring with webhook alerts\n- Any read-only VMware query where safety is paramount\n\n### Alarm/Event Output: `suggested_actions` Field\n\n`get_alarms` and `get_events` results include a `suggested_actions` list. Each\nitem is a ready-to-use hint naming the correct companion skill and tool call\n(e.g. `\"vmware-aiops: acknowledge_vcenter_alarm(entity_name=..., alarm_name=...)\"`),\nso agents — especially smaller local models — can follow them directly without\nreasoning about skill routing. Example payload: `references/capabilities.md`.\n\n**Use companion skills for**:\n- Power on/off, deploy, clone, migrate --> `vmware-aiops`\n- iSCSI, vSAN, datastore management --> `vmware-storage`\n- Tanzu Kubernetes clusters --> `vmware-vks`\n- Load balancing, AVI/ALB, AKO, Ingress --> `vmware-avi`\n\n## Related Skills — Skill Routing\n\n| User Intent | Recommended Skill |\n|-------------|------------------|\n| Read-only vSphere monitoring | **vmware-monitor** ← this skill |\n| Storage: iSCSI, vSAN, datastores | **vmware-storage** |\n| VM lifecycle, deployment, guest ops | **vmware-aiops** |\n| Tanzu Kubernetes (vSphere 8.x+) | **vmware-vks** |\n| NSX networking: segments, gateways, NAT | **vmware-nsx** |\n| NSX security: DFW rules, security groups | **vmware-nsx-security** |\n| Aria Ops: metrics, alerts, capacity planning | **vmware-aria** |\n| Multi-step workflows with approval | **vmware-pilot** |\n| Compliance baselines (CIS / 等保 / PCI-DSS), drift detection, LLM remediation advisor | **vmware-harden** (`uv tool install vmware-harden`) |\n| Load balancer, AVI, ALB, AKO, Ingress | **vmware-avi** (`uv tool install vmware-avi`) |\n| Audit log query | **vmware-policy** (`vmware-audit` CLI) |\n\n## Common Workflows\n\n> **Diagnostic investigations**: Before running any \"why is X failing / down / abnormal\" workflow, follow [`references/investigation-protocol.md`](references/investigation-protocol.md). It enforces the four root-cause completeness criteria (falsifiability / sufficiency / necessity / mechanism) and the up-to-three-rounds deepening loop. Since vmware-monitor is read-only, it serves as the data source — actuation belongs to companion skills like vmware-aiops.\n\n### Cluster Health Check (\"is anything on fire?\" / \"what's wrong right now?\")\n\n**Judgment**: this is the 5-second triage glance, not an Aria replacement. One call rolls every cluster's hosts, VM power, live CPU/memory and alarms up, flattens the individual anomalies into a ranked **top-N focus list** (`top_issues`), and gives each cluster an opinionated `status`. On a big fleet, lead with the focus list — scanning per-cluster rows is too slow.\n\n1. One glance --> `vmware-monitor summary` (MCP: `cluster_health_summary`). Read `top_issues` first (worst first, each with a drill-down `next step`); the per-cluster table is context. `issues_total` shows how many anomalies existed before the top-N cap\n2. Tighten or widen the focus --> `--top 5` for the 5 most urgent, `--top 20` for more, `--top 0` to hide the list and just see the table\n3. Drill into what the list points at --> e.g. a `host_down` row → `inventory hosts`; an `alarm` row → `get_alarms`; a `capacity` row → `perf hosts` / `capacity datastores`; scope with `--cluster prod-a`\n4. Reshape the view on request --> the output ends with a friendly hint; the operator can say \"add datastore free space\", \"drop the DRS column\", \"only show clusters needing attention\", or \"save this as an HTML page\". Default layout, columns, and thresholds live in [`references/health-summary-template.md`](references/health-summary-template.md) and are meant to be edited\n5. Save an offline snapshot --> `vmware-monitor summary --html` writes a self-contained HTML file (no external assets, nothing uploaded) to `~/vmware-health/cluster-health-<vc>-<timestamp>.html`; `--html-path <file>` for an explicit path. The timestamped filename means a folder of them becomes a browsable point-in-time history. It is a snapshot, not a live page — re-run to refresh\n6. **On a very large fleet** --> add `--no-vms` to skip the VM rollup pass when you only need host/alarm/capacity signals\n7. **`totals.clusters` is 0** --> not empty: un-clustered hosts (standalone ESXi, or a cluster-less vCenter) form the `(standalone hosts)` row, still in `totals` and `top_issues`\n\n### Daily Health Check\n\n**Judgment**: alarms tell you what vCenter has decided is wrong, events tell you what happened. They diverge — an event burst with no alarms often signals a metric threshold miscalibration, not \"everything is fine.\" Read both.\n\n1. Check alarms --> `vmware-monitor health alarms --target prod-vcenter` — focus on Red severity AND alarms older than 1 hour (transient ones self-clear)\n2. Review recent events --> `vmware-monitor health events --hours 24 --severity warning` — look for repeated events from the same entity (a single event is noise; 50 events in an hour is a pattern)\n3. List hosts --> `vmware-monitor inventory hosts` — flag hosts disconnected, in maintenance mode unexpectedly, or memory > 90%\n4. **If connection fails** --> run `vmware-monitor doctor` to diagnose config/network issues\n\n### Object-Centered Investigation (\"what is happening around this VM / host / datastore?\")\n\n**Judgment**: this is the drill-down the operator wants after triage points at a problem — one call *correlates* the object with its surrounding infrastructure and recent history, so you explain the aggregated result in operational language instead of stitching five tools yourself. The tool aggregates; you never dump raw inventory into the conversation.\n\nOffer the levels progressively — do **not** ask for details the environment already fixes:\n1. **Start at the top** --> `cross_vcenter_attention` (CLI: `vmware-monitor attention`) for \"what needs attention now?\" across every vCenter. If only one vCenter is configured, skip straight to its `cluster_health_summary` — no need to ask which target\n2. **Offer to drill into an object** the top-issues list points at. Ask *which level* only when it is genuinely ambiguous:\n   - a VM --> `vm_investigation_bundle` (CLI: `vmware-monitor investigate vm <name>`) → VM state, recent events, snapshots, alarms & recent changes, the host it runs on, the cluster context, the datastores backing it, performance signals, and a correlated event timeline\n   - a host --> `host_investigation_bundle` (CLI: `investigate host <name>`) → host state, cluster context, the VMs it runs, mounted datastores, alarms, performance, correlated timeline\n   - a datastore --> `datastore_investigation_bundle` (CLI: `investigate datastore <name>`) → capacity/free, mounting hosts, VMs it backs, alarms, correlated timeline\n3. **Widen or narrow** on request --> `--hours 72` for a longer event window; the bundle ends with a hint listing what is adjustable\n4. **Make it tangible** --> add `--html` to any `investigate`/`attention` command for a self-contained offline snapshot (drill-down sections collapse/expand natively, no JS, nothing uploaded) written to `~/vmware-health/`\n5. **If the object name is unknown** --> the tool returns a *teaching* error naming exactly how to list the objects (`list_virtual_machines` / `list_esxi_hosts` / `list_all_datastores`); get the exact name and retry\n6. **If a vCenter is unreachable** (attention only) --> it is listed under `unreachable` with a reason and the rest still aggregate — surface the gap, don't fail the whole view\n\n### Performance Triage (\"the cluster feels slow\")\n**Judgment**: inventory shows *configured* capacity (cores, GB); it cannot tell you what is actually hot. Use the real-time perf tools, then narrow.\n1. Rank hosts --> `vmware-monitor perf hosts` — the busiest host floats to the top (sorted by CPU%)\n2. Rank VMs on the suspect --> `vmware-monitor perf vms --limit 25` — find the noisy neighbour\n3. Check for hidden storage pressure --> `vmware-monitor capacity datastores` — over-commit % > 100 means a thin datastore can fill mid-run even with \"free\" space showing\n4. Rule out snapshot drag --> `vmware-monitor snapshots aging --only-old` — old snapshots silently degrade I/O\n5. **If perf tools return empty** --> the host/VM may be disconnected or powered off (no real-time provider); confirm with `inventory hosts` / `inventory vms`\n\n### Scheduled-Outage Pre-flight (certs, licenses, time)\n1. Cert expiry --> `vmware-monitor infra certs --warn-days 60` — an expired ESXi cert drops host management\n2. License headroom --> `vmware-monitor infra licenses` — catch over-allocation before it disables features\n3. Time sync --> `vmware-monitor infra ntp` — `healthy: no` breaks SSO/Kerberos/log correlation (note: live offset is not exposed by the SOAP API, only config health)\n\n### Set Up Continuous Monitoring\n1. Configure webhook in `~/.vmware-monitor/config.yaml`\n2. Start daemon --> `vmware-monitor daemon start`\n3. Daemon scans every 15 min, sends alerts to Slack/Discord\n\n## Usage Mode\n\n| Scenario | Recommended | Why |\n|----------|:-----------:|-----|\n| Local/small models (Ollama, Qwen) | **CLI** | ~2K tokens vs ~8K for MCP |\n| Cloud models (Claude, GPT-4o) | Either | MCP gives structured JSON I/O |\n| Automated pipelines | **MCP** | Type-safe parameters, structured output |\n\n## MCP Tools (32 — all read-only)\n\n| Tool | Description |\n|------|------------|\n| `list_virtual_machines` | List VMs with filtering (power state, sort, limit, `folder_filter`); each VM includes `folder_path` |\n| `list_esxi_hosts` | ESXi hosts with CPU, memory, version, uptime |\n| `list_all_datastores` | Datastores with capacity, free space, type |\n| `list_all_clusters` | Clusters with host count, DRS/HA status |\n| `cluster_health_summary` | One-glance triage across all clusters — ranked `top_issues` focus list + per-cluster rollup with opinionated `status`. Params: `cluster_filter`, `include_vms`, `top_n`. Render per `references/health-summary-template.md` |\n| `vm_investigation_bundle` | \"What is happening around this VM?\" — correlated drill-down with a merged, newest-first **event timeline** (contents listed in the workflow above). Params: `vm_name`, `hours`. Aggregated in the tool; explain, don't dump raw |\n| `host_investigation_bundle` | Same correlated drill-down around an ESXi host. Params: `host_name`, `hours` |\n| `datastore_investigation_bundle` | Same correlated drill-down around a datastore. Params: `datastore_name`, `hours` |\n| `cross_vcenter_attention` | \"What needs attention now?\" across **every** configured vCenter — one globally-ranked `top_issues` list (each tagged with its `vcenter`) + per-target rollup; unreachable targets degrade gracefully. Params: `cluster_filter`, `top_n` |\n| `list_all_networks` | Networks with attached VM count and accessibility |\n| `get_alarms` | Active alarms: `suggested_actions`, acknowledger, `condition_now` (`cleared` = stale) |\n| `get_events` | Events by severity and time window (`start`/`end`) |\n| `get_host_sensors` | Hardware sensor status (temperature/voltage/fan) per host with green/yellow/red health |\n| `get_host_services` | Host service status (running state and startup policy), optionally filtered by host |\n| `vm_info` | Detailed VM info (CPU, memory, disks, NICs, snapshots) |\n| `vm_list_snapshots` | Snapshot list for one VM with nesting hierarchy (read-only) |\n| `host_performance` | **Real-time** host CPU/mem/disk/net utilisation (PerfManager); busiest first |\n| `vm_performance` | **Real-time** VM CPU/mem/balloon/swap/disk/net utilisation (top 25 by default); powered-on only |\n| `snapshot_aging` | Inventory-wide snapshot sweep with age + sprawl; flags snapshots older than N days |\n| `vm_backup_snapshot_history` | Backup windows for one VM from snapshot task history; a lower bound, not job duration |\n| `certificate_status` | Per-host ESXi management certificate expiry (days until expiry, expiring flag) |\n| `license_status` | Licenses and per-asset `assignments` |\n| `ntp_status` | Per-host NTP config health (servers + ntpd state); live offset not in SOAP API |\n| `datastore_capacity` | Datastore over-commit (provisioned vs capacity); thin-provisioning risk |\n| `resource_pool_usage` | Resource-pool CPU/memory reservation, limit, and current usage |\n| `active_tasks` | In-flight (and recently completed) vCenter tasks with progress/errors |\n| `active_sessions` | Who is logged in, and from which client |\n| `host_log_scan` | ESXi host log trouble lines, grouped by pattern (CLI: `scan logs`) |\n| `host_memory_tiering` | **vSphere 9.1** — per-host memory tiering (DRAM/NVMe tiers) + NVMe uplift ratio (pyVmomi `hardware.memoryTierInfo`, needs ESXi 8.0U3+). Params: `host_name`, `limit`. Returns the list envelope |\n| `cluster_patch_compliance` | **vSphere 9.1** — vLCM software (patch) compliance for one cluster over vSphere Automation REST. Param: `cluster` (MoID, e.g. `domain-c123`). `available:false` = vCenter answered 503 (likely mid-patch), not an error; `non_compliant_hosts` is `null` when the host-status field is unknown, never a false 0 |\n| `cluster_last_apply_result` | **vSphere 9.1** — result of the last vLCM remediation (apply) on one cluster (REST). Param: `cluster` (MoID). Reports outcome only; never runs a remediation |\n| `vcenter_deployment_size` | **vSphere 9.1** — vCenter appliance deployment size class (REST, NEW in 9.1). `available:false` on 503 |\n\n> **vSphere 9.1 field-parse honesty**: the 3 REST tools' *endpoints* are spec-verified, but their JSON field names have NOT yet been replayed against a live 9.1 vCenter — every field is read defensively and each result self-labels via its `note` (`endpoint verified; field parse best-effort pending live 9.1 vCenter`). `host_memory_tiering` requires vCenter/ESXi 8.0U3+; older targets raise a teaching error naming the missing property.\n\nNo tool modifies, creates, or deletes any vCenter/ESXi resource.\nPerformance/capacity readings are point-in-time samples — this skill retains no\nhistory, so it never reports a fabricated \"trend\" or runway date.\n\n### List result shape\n\nThe 21 row-listing tools above (including `host_memory_tiering`) return the family list envelope\n`{items, returned, limit, total, truncated, hint}`, not a bare array. Read\n`truncated` before summarising: `true` means more rows exist — never call\n`items` the whole picture; `false` means complete, so empty `items` means\n\"checked, found none\" — for `host_log_scan`, only if `logs_unavailable`\n(logs it could not read) is empty too. A `null` `total`\n(`get_events`, `host_log_scan`) is deliberate. Aggregate tools return\npurpose-built objects — see `references/capabilities.md`.\n\n## Read-Only by Design\n\nAll 32 tools are vSphere reads (local writes: see top). Running\nwith local or small models? See\n[`references/agent-guardrails.md`](references/agent-guardrails.md).\n\n## CLI Quick Reference\n\n```bash\nvmware-monitor summary [--top 10] [--cluster <substr>] [--html] [--target <t>]\nvmware-monitor inventory vms|hosts|datastores|clusters|networks [--target <t>]\nvmware-monitor health alarms|events|sensors|services [--target <t>]\nvmware-monitor perf hosts|vms [--target <t>]\nvmware-monitor capacity datastores|pools [--target <t>]\nvmware-monitor infra certs|licenses|ntp [--target <t>]\nvmware-monitor snapshots aging [--only-old] [--target <t>]\nvmware-monitor vm info <vm-name> [--target <t>]\nvmware-monitor memory tiering [--host <esxi>] [--target <t>]        # vSphere 9.1\nvmware-monitor patch compliance|last-apply <cluster-moid> [--target <t>]   # vSphere 9.1 (vLCM)\nvmware-monitor deployment-size [--target <t>]                      # vSphere 9.1\nvmware-monitor scan now | scan logs [--host <name>] | daemon start|stop|status | doctor [--skip-auth]\n```\n\n> Full CLI reference (all flags + activity/tasks/sessions): see `references/cli-reference.md`\n\n## Troubleshooting\n\n### Alarms returns empty but vCenter shows alarms\nThe `get_alarms` tool queries triggered alarms at the root folder level. Some alarms are entity-specific — try checking events instead: `get_events --hours 1 --severity info`.\n\n### \"Connection refused\" error\n1. Run `vmware-monitor doctor` to diagnose\n2. Verify target hostname/IP and port (443) in config.yaml\n3. For self-signed certs: set `verify_ssl: false`\n\n### Events returns too many results\nUse severity filter: `--severity warning` (default) filters out info-level events. Use `--hours 4` to narrow time range.\n\n### VM info shows \"guest_os: unknown\"\nVMware Tools not installed or not running in the guest. Install/start VMware Tools for guest OS detection, IP address, and guest family info.\n\n### Doctor passes but commands fail with timeout\nvCenter may be under heavy load. Try targeting a specific ESXi host directly instead of vCenter, or increase connection timeout in config.yaml.\n\n### Should I set `environment:` on a read-only skill?\nYou can — add `environment: production` (or `staging`, `lab`, your own label)\nto each target in `~/.vmware-monitor/config.yaml`. It's an optional label; this\nskill has zero write tools, so nothing it exposes is ever gated by it — reads\nare never gated. It matters for the write skills (`vmware-aiops`,\n`vmware-storage`, `vmware-nsx`) pointed at the same vCenter: an\nenvironment-scoped `deny` rule in `~/.vmware/rules.yaml` can match on the label\nto block their writes (e.g. freeze `production`). A target with no label is\nsimply not matched by such a rule. Config example: `references/setup-guide.md`.\n\n## Setup\n\n```bash\nuv tool install vmware-monitor==1.16.0\nvmware-monitor init      # guided: prompts for host/user/password, writes config + .env (chmod 600), then verifies\n```\n\n`init` stores the password grep-safe (obfuscated `b64:`, never plaintext) and\nlocks `.env` to 0600. Prefer it over hand-editing; manual steps:\n`references/setup-guide.md`.\n\n> Full setup guide, security details, and AI platform compatibility: see `references/setup-guide.md`\n\n## Audit & Safety\n\nMCP calls (`@vmware_tool`) and remote CLI commands (`@audited`) are audited via vmware-policy; CLI queries also append to `~/.vmware-monitor/audit.log`:\n- Every MCP call and remote CLI command logged to `~/.vmware/audit.db` (SQLite)\n- Policy rules enforced via `~/.vmware/rules.yaml` (deny rules, maintenance windows, risk levels)\n- Risk classification: each tool tagged as low/medium/high/critical\n- View recent operations: `vmware-audit log --last 20`\n- View denied operations: `vmware-audit log --status denied`\n\n## License\n\nMIT — [github.com/vmware-skills/VMware-Monitor](https://github.com/vmware-skills/VMware-Monitor)\n\nFile v1.16.0:_meta.json\n\n{\n  \"ownerId\": \"kn7b067awq2s97bn3d7p5qfhw5827pxc\",\n  \"slug\": \"vmware-monitor\",\n  \"version\": \"1.16.0\",\n  \"publishedAt\": 1789915888294\n}\n\nFile v1.16.0:references/agent-guardrails.md\n\n# Operating the VMware skills with a local / small model\n\nClaude-class models drive these skills without special instruction. Smaller and\nlocally-hosted models — Llama 3.3 70B, Qwen, Mistral, and similar, served\nthrough Goose, Ollama, or OpenShift AI — need explicit operating rules to call\ntools reliably.\n\nThis page exists because an operator wrote those rules by hand first. The\nguardrails below are adapted, with thanks, from the working configuration\n[@juanpf-ha](https://github.com/juanpf-ha) developed while running\nvmware-monitor and vmware-aria against a production vSphere estate with Llama\n3.3 70B FP8 on an on-prem H100\n([VMware-AIops#31](https://github.com/vmware-skills/VMware-AIops/issues/31)).\n\n> **Disclaimer**: This is a community-maintained open-source project and is\n> **not affiliated with, endorsed by, or sponsored by VMware, Inc. or Broadcom\n> Inc.** \"VMware\" and \"vSphere\" are trademarks of Broadcom.\n\n---\n\n## First: the rules you no longer need to write\n\nSeveral guardrails from the original configuration are now enforced by the\nskills themselves. Prompt instructions are advisory — a model can ignore them.\nThese are structural, so it cannot.\n\n| Guardrail you would otherwise prompt for | Now enforced by |\n|---|---|\n| \"Work read-only and never modify anything\" | **Zero write tools.** This skill has no vSphere create, modify, delete, or power operation in its registry at all — `list_tools()` only ever offers reads, so the model cannot call what does not exist. |\n| \"First resolve the affected resource_id through vmware-aria before querying vmware-monitor\" | **`investigate_alert`** does the whole alert → resource → confirmed-name sequence in one call. |\n| \"Do not confuse the alert ID with the affected resource ID\" | `investigate_alert` returns a `correlation` block with both UUIDs explicitly labelled. |\n| \"Only correlate Aria and vCenter data after the resource name and type have been confirmed\" | `investigate_alert` returns `correlation.confirmed`, and withholds its `next_step` handoff until name and kind are known. |\n| \"Use explicit limits for queries that may return large amounts of data\" | **The list envelope.** Every list tool returns `{items, returned, limit, total, truncated, hint}`, so the model reads truncation instead of guessing at it. Most tools are unlimited unless you pass `limit` (`vm_performance` defaults to 25); the envelope states which case you got. |\n| \"If a requested field was not returned by any tool, show it as not available\" | Tools return explicit `null` for unresolved fields rather than omitting the key. |\n\n---\n\n## The system prompt\n\nEverything below still benefits from being stated explicitly. Copy this into\nyour agent's instruction block.\n\n```text\n## Tool use\n\n- Always call an MCP tool before answering any question about the current\n  VMware environment. Never answer from memory or assumption.\n- Never describe a tool call, and never output a JSON example, instead of\n  executing the tool. If you intend to call a tool, call it.\n- If a tool fails, report the actual error text. Do not complete the answer\n  with assumptions about what the result would have been.\n- Use explicit limits on queries that may return large amounts of data. Do not\n  request unlimited results unless the user asks for them.\n\n## Skill routing\n\n- vmware-monitor: vCenter inventory, ESXi hosts, clusters, datastores, VMs,\n  snapshots, alarms, events, performance.\n- vmware-aria: Aria Operations health, resources, metrics, alerts, capacity.\n- vmware-aiops: VM lifecycle (power, clone, snapshot, migrate, delete).\n- vmware-nsx / vmware-nsx-security: networking and firewall.\n- vmware-storage: datastores, iSCSI, vSAN.\n\n## Data fidelity\n\n- Never invent infrastructure objects, metrics, alarms, events, or\n  relationships. If a tool did not return it, it does not exist for this answer.\n- Preserve the exact criticality, status, impact, and control-state values the\n  tools return. Do not translate, normalise, or prettify enum values.\n- If a requested field was not returned, show it as \"not available\". Do not\n  infer it from other fields.\n- Preserve the original order and the full set of fields when the user asks\n  for specific ones.\n- When a response is long, report every item it contains. If a result is\n  truncated, the tool says so explicitly — report the truncation rather than\n  describing the visible subset as the whole.\n\n## Analysis discipline\n\n- Separate observed data from interpretation. State which is which.\n- Do not claim a security, performance, storage, or capacity problem unless\n  the tool output contains explicit supporting evidence.\n- Avoid generic recommendations that are not directly supported by the results.\n\n## Correlating Aria alerts with vCenter\n\n- Use investigate_alert to go from an alert to its affected resource. It\n  resolves the resource and confirms its name and kind in one call.\n- Do not pass an alert UUID where a resource UUID is expected. investigate_alert\n  labels both in its correlation block.\n- Only query vCenter for a resource once correlation.confirmed is true. The\n  next_step block names the exact tool and argument to use.\n```\n\n---\n\n## Known failure modes on small models\n\nObserved with Llama 3.3 70B FP8 (Goose, on-prem H100), and useful as a\nchecklist when evaluating any local model against these skills:\n\n| Symptom | Mitigation |\n|---|---|\n| Describes a tool call, or emits a JSON example, instead of executing it | The \"never describe a tool call\" rule above. Also check your harness is not echoing tool schemas into context — models imitate the nearest format they see. |\n| Long tool responses: omits items, or reports \"no data returned\" when data was present | Ask for explicit limits so responses stay small. Check the envelope's `truncated` / `returned` / `total` fields rather than trusting the model's summary — every list tool states them, so a \"no data\" claim is checkable against `returned`. |\n| Adds generic recommendations unsupported by results | The \"analysis discipline\" rules. |\n| Drops requested fields or reorders results | State the required fields and ordering in the request itself, not only in the system prompt. |\n| Multi-tool workflows take 30–50s end to end | Prefer the aggregate tools — `investigate_alert`, `cluster_health_summary`, `vm_investigation_bundle` — which collapse a 3-4 call sequence into one round trip. |\n\n## Reporting results\n\nLocal-model compatibility is an explicit design constraint for this family, and\nthe evidence base is small. If you evaluate a model against these skills —\nQwen, Mistral, Granite, or anything else — a report of what worked and what did\nnot is genuinely useful:\n[github.com/vmware-skills/VMware-Monitor/issues](https://github.com/vmware-skills/VMware-Monitor/issues).\n\nFile v1.16.0:references/capabilities.md\n\n# Capabilities (Read-Only)\n\nDetailed feature tables for `vmware-monitor`.\n\n## List result envelope\n\nThe 20 list-returning MCP tools — `list_virtual_machines`, `list_esxi_hosts`,\n`list_all_datastores`,\n`list_all_clusters`, `list_all_networks`, `get_alarms`, `get_events`,\n`get_host_sensors`, `get_host_services`, `host_log_scan`, `active_tasks`,\n`active_sessions`, `datastore_capacity`, `resource_pool_usage`,\n`certificate_status`, `license_status`, `ntp_status`, `host_performance`,\n`vm_performance`, `vm_list_snapshots` — return the family list envelope rather\nthan a bare array:\n\n```json\n{\"items\": [...], \"returned\": 50, \"limit\": 50, \"total\": 213,\n \"truncated\": true, \"hint\": \"Showing 50 of 213. Raise limit or narrow the query...\"}\n```\n\n| Key | Meaning |\n|-----|---------|\n| `items` | The rows, in the tool's documented order |\n| `returned` | `len(items)` — check this before claiming \"no data\" |\n| `limit` | The limit that produced this page; `null` when unlimited |\n| `total` | Real collection size, or `null` when the backing API does not report one |\n| `truncated` | `true` = more rows exist behind this page; `false` = this is complete |\n| `hint` | What to do about truncation; `null` when complete |\n\nTwo tools report `total: null` on purpose: `get_events` (events are read newest\nfirst and the read stops at 5000 — `read_truncated: true` and `read_note` say\nwhen it did and how far back it got) and `host_log_scan` (only the last N lines\nper log are read). The investigation bundles add `timeline_note` when their\ntimeline is not every event in the window: how many of how many are shown, and\nwhich scopes' reads stopped at 5000. Everywhere else the total is a real count taken before the\nlimit was applied, which is what lets a full page be recognised as complete\ninstead of flagged as possibly-truncated.\n\n`host_log_scan` adds one field to the envelope: `logs_unavailable`, one row per\nhost/log it could **not** read, with the reason (for example, the account lacks\n`Global.Diagnostics`, which `BrowseDiagnosticLog` requires). Its `items` holds\nonly the lines that matched a trouble pattern in the logs it *did* read, so an\nempty `items` with `truncated: false` means \"checked, found none\" only when\n`logs_unavailable` is empty as well. With unread logs, the honest answer is\n\"nothing matched in the logs that could be read\" — name the ones that could not.\n\nBy default `host_log_scan` groups repeated lines by pattern, per log: each item has `count`,\n`hosts`, `first_seen` / `last_seen` (the log's own timestamps), `severity`, `log_level`, `pattern`\nand one `sample`; `lines_matched` is the ungrouped count and `group=false` returns one row per\nline. Severity follows the level ESXi wrote on the line (`Cr`/`Al`/`Em` critical, `Er`/`Wa`\nwarning, `In`/`No`/`Db` info); a line containing \"critical\", \"panic\" or \"corrupt\" is critical\nwhatever its level.\n\nThe envelope adds ~30 tokens to a response. It exists because a bare list gave\nsmaller models nothing to distinguish a complete answer from page one, and they\nsometimes resolved that ambiguity as \"no data was returned\"\n(VMware-AIops issue #31).\n\n`list_virtual_machines` adds one extra key to the envelope, `mode` (`\"full\"` or\n`\"compact\"`), and reuses `hint` for the compact-mode note; its `total` is the\ncount after `power_state` / `folder_filter` are applied and before `limit`.\n\nTools with purpose-built return objects — `vm_info`, `snapshot_aging`,\n`cluster_health_summary`, `cross_vcenter_attention`, and the three\n`*_investigation_bundle` tools — are unaffected.\n\n## Automation Level Reference\n\nEach operation is classified by autonomy level per the Enterprise Harness Engineering framework. **vmware-monitor is L1/L2 only by design** — no vSphere write operations exist in the codebase, gated by an allowlist test in the source repository.\n\n| Level | Meaning | Agent autonomy | Examples in this skill |\n|:-:|---|---|---|\n| **L1** | Read-only, raw data | Always auto-run | `list_virtual_machines`, `list_esxi_hosts`, `get_alarms`, `get_events`, `list_all_datastores`, `list_all_clusters`, `host_performance` |\n| **L2** | Read + analysis / recommendation | Always auto-run | `cluster_health_summary`, `cross_vcenter_attention`, `snapshot_aging`, the three `*_investigation_bundle` tools, scheduled scan reports, log pattern matching (error/fail/critical/panic/timeout), alarm correlation, daemon-driven webhook digests |\n| **L3** | Single write — user must approve | *N/A* | — *(use [vmware-aiops](https://github.com/vmware-skills/VMware-AIops) for write operations)* |\n| **L4** | Multi-step plan / apply workflow | *N/A* | — *(use [vmware-pilot](https://github.com/vmware-skills/VMware-Pilot) for orchestration)* |\n| **L5** | Auto-remediation from learned pattern | *N/A* | — *(remediation is out of scope by design)* |\n\n**Notes**:\n- No tool changes vCenter/ESXi state, so agents can call them without confirmation — gated by [`tests/eval/regression/test_read_only_enforcement.py`](https://github.com/vmware-skills/VMware-Monitor/blob/main/tests/eval/regression/test_read_only_enforcement.py) (source repository; a check on the code as written, run by the test suite — there is no CI). Results still carry sensitive inventory, event, log, and session data: scope the account as in `setup-guide.md` → Least Privilege.\n- Local files the skill writes (config, `.env`, audit logs, HTML snapshots, daemon state) are listed in `setup-guide.md` → What \"read-only\" covers.\n\n## 0. Cluster Health Summary (triage)\n\nCLI `summary`, MCP `cluster_health_summary`. One aggregated read for a fast\ncross-cluster \"is anything on fire?\" glance — the operator's first look, not an\nAria Operations replacement.\n\n| Aspect | Detail |\n|--------|--------|\n| Passes | 4 batched `RetrievePropertiesEx` calls (clusters, hosts, VMs, datastores) — never one per object (issue #31 class) |\n| Focus list | `top_issues`: individual anomalies (disconnected hosts, triggered alarms, capacity/HA, datastores thin-provisioned past `DATASTORE_OVERCOMMIT_WARN_PCT=100` — `scope: datastore`) flattened + ranked worst-first, capped at `top_n`; `issues_total` reports pre-cap count. Alarm names and definitions resolved in one batched call (no N+1) |\n| Alarm issues | Carry `condition_now` (`holds` / `cleared` / `unknown`, as `get_alarms`) and `acknowledged_days` (null when not acknowledged). Only `cleared` ranks after the other issues; `unknown` keeps its severity rank — not re-checked, possibly still live. `object` is `vCenter appliance <address>` for alarms vCenter raises about itself |\n| Next step | `drilldown` names MCP tools; the CLI table and HTML snapshot name CLI commands |\n| Rollup | Per cluster: hosts connected/total, VM power, live CPU/mem %, HA/DRS, alarm counts (cluster + host) |\n| Status | Opinionated `ok` / `warn` / `critical` + plain-language `attention` reasons; sorted worst-first |\n| Thresholds | `CPU_MEM_WARN_PCT=85`, `CPU_MEM_CRIT_PCT=95` (named constants in `ops/cluster_summary.py`); disconnected host or critical alarm forces `critical` |\n| Customizable | Columns/thresholds/layout in [`health-summary-template.md`](health-summary-template.md); response carries a `customization_hint` |\n\n### `cluster_health_summary` — input parameters\n\n| Parameter | Type | Default | Behavior |\n|-----------|------|---------|----------|\n| `target` | str (optional) | default target | Named vCenter/ESXi target |\n| `cluster_filter` | str (optional) | None (all) | Case-insensitive substring; suppresses standalone-hosts bucket |\n| `include_vms` | bool | True | Roll up VM power counts; False skips the VM pass (faster on huge fleets) |\n| `top_n` | int | 10 | Cap the `top_issues` focus list; `issues_total` keeps the pre-cap count; 0 hides the list |\n\n**`totals.clusters` counts real clusters only.** Hosts that belong to no cluster\n— a standalone ESXi target, or standalone hosts under a vCenter — are rolled into\na `(standalone hosts)` row, and alarms raised above any cluster or host get a\n`(vCenter-level)` row when there are any; neither row is counted as a cluster. Their hosts, VMs and alarms still count in the\nother `totals` fields and feed `top_issues`. So a vCenter with only standalone\nhosts reports `clusters: 0` next to a non-zero `hosts_total` and a populated\nstandalone row: read that row, not an alternative tool. `cluster_filter` hides\nthe standalone row, so a filtered `clusters: 0` means the filter matched nothing.\n\n**Typical response tokens**: ~120–400 (one compact row per cluster + totals);\nscales with cluster count, not VM count. This is the aggregation-in-the-tool\npattern — the model never sees raw inventory.\n\n### `cross_vcenter_attention` — overlapping targets\n\nCLI `attention`. Targets can overlap: an ESXi host configured as its own target\nmay also be managed by a configured vCenter. Identity is read, never taken from\nnames:\n\n| Field | Meaning |\n|-------|---------|\n| `totals.vcenters` / `esxi_targets` / `unidentified_targets` | Targets by endpoint type (`about.apiType`); a type that cannot be read is `unidentified`, never counted as a vCenter |\n| `totals.hosts_total` / `hosts_connected` | A host reached through a vCenter and as its own ESXi target counts once (`summary.hardware.uuid`). Not applied under `cluster_filter` |\n| `targets[].endpoint` / `shared_hosts` | `vCenter` / `ESXi host` / `unknown`, and how many of the target's hosts another target already counted |\n| `top_issues[].also_seen_via` | On a datastore issue: the same volume (`summary.url`) seen through an ESXi target that the vCenter manages — `[{vcenter, detail}]`, that view's own figure |\n\nOnly a real overlap is merged: an ESXi target whose host a single vCenter target\nreports. A hardware UUID seen on two hosts of one target, on hosts of two vCenter\ntargets, or a known SMBIOS placeholder (`03000200-0400-0500-0006-000700080009`,\nall zeros, all f) is not an identity. Two vCenters sharing an NFS datastore keep\nboth issues, each with its own figure.\n\n## 1. Inventory\n\n| Feature | vCenter | ESXi | Details |\n|---------|:-------:|:----:|---------|\n| Cluster health summary | Y | N | Cross-cluster triage rollup with opinionated status — CLI `summary`, MCP `cluster_health_summary` |\n| List VMs | Y | Y | Name, power state, CPU, memory, guest OS, IP, `folder_path` (vCenter inventory folder, e.g. `/Datacenters/Production/Web Tier`) |\n| List Hosts | Y | Self only | CPU cores, memory, ESXi version, VM count, uptime |\n| List Datastores | Y | Y | Capacity, free/used, type (VMFS/NFS), usage % |\n| List Clusters | Y | N | Host count, DRS/HA status |\n| List Networks | Y | Y | Network name, associated VM count, accessibility — CLI `inventory networks`, MCP `list_all_networks` |\n\n### `list_virtual_machines` — input parameters\n\n| Parameter | Type | Default | Behavior |\n|-----------|------|---------|----------|\n| `target` | str (optional) | default target | Named vCenter/ESXi target from `config.yaml` |\n| `limit` | int (optional) | None (all) | Max VMs to return |\n| `sort_by` | str | `name` | `name` \\| `cpu` \\| `memory_mb` \\| `power_state` \\| `folder_path` |\n| `power_state` | str (optional) | None | `poweredOn` \\| `poweredOff` \\| `suspended` |\n| `fields` | list[str] (optional) | auto | Subset of: `name`, `power_state`, `cpu`, `memory_mb`, `guest_os`, `ip_address`, `host`, `uuid`, `tools_status`, `folder_path` |\n| `folder_filter` | str (optional) | None | Case-insensitive substring match against `folder_path` (CLI `--folder-filter`, MCP `folder_filter`). Example: `folder_filter=\"Production\"` returns VMs anywhere under any folder whose path contains \"production\" (including nested subfolders like `/Datacenters/Production/Web Tier`). |\n\n### `list_virtual_machines` — response fields\n\nEach VM dict in the `vms` array contains:\n\n| Field | Description |\n|-------|-------------|\n| `name` | VM name |\n| `power_state` | `poweredOn` / `poweredOff` / `suspended` |\n| `cpu` | vCPU count |\n| `memory_mb` | RAM in MB |\n| `guest_os` | Guest OS full name (full mode only) |\n| `ip_address` | Guest IP from VMware Tools (full mode only) |\n| `host` | ESXi host name (full mode only) |\n| `uuid` | VM UUID (full mode only) |\n| `tools_status` | VMware Tools running status (full mode only) |\n| `folder_path` | vCenter inventory folder path, e.g. `/Datacenters/Production/Web Tier`. Returned in both compact and full modes. |\n\n## 2. Health & Monitoring\n\n| Feature | vCenter | ESXi | Details |\n|---------|:-------:|:----:|---------|\n| Active Alarms | Y | Y | Severity, alarm name, entity, timestamp, who acknowledged it and when, and `condition_now` (holds / cleared / unknown) |\n| Event/Log Query | Y | Y | Filter by time range, severity; 50+ event types |\n| Hardware Sensors | Y | Y | Per-sensor `type` (temperature/voltage/fan...), reading, unit, and health `status` (green/yellow/red); connected hosts with no sensors in `hosts_without_sensors` with their CIM Server (`sfcbd-watchdog`) state and a `sensors_note` — CLI `health sensors`, MCP `get_host_sensors` |\n| Host Services | Y | Y | hostd, vpxa running/stopped status — CLI `health services`, MCP `get_host_services` |\n\n### Stale alarms — `condition_now`\n\nvCenter keeps an alarm triggered until something resets it. `get_alarms` re-evaluates the\n**state** parts of each alarm's definition against the object's current properties:\n`holds` (live), `cleared` (still shown, but the condition is false now — reset it after\nconfirming), or `unknown` (event- or metric-based, or a property could not be read; never\nguessed). `stale_alarms` counts the cleared ones and `stale_note` explains. On a lab vCenter\n8.0.3, \"Host connection and power state\" was red for eleven days on a connected host and now\nreads `cleared` with the note `runtime.connectionState is connected, not notResponding`.\n\nOne event-raised alarm is decided anyway, because the skill reads the evidence: an\nexpired-vCenter-license alarm (`com.vmware.license.VcLicenseExpiredEvent`) is checked\nagainst the license this vCenter itself is assigned (matched by `about.instanceUuid`, one\n`QueryAssignedLicenses` call, only when such an alarm is present). A readable, unexpired\nexpiry gives `cleared`; an expired one gives `holds`; unreadable assignments, no row for\nthis vCenter, or evaluation mode stay `unknown`.\n\n`object_label` names alarms about the vCenter appliance itself that sit on the inventory\nroot (\"Datacenters\") — e.g. `vCenter appliance 192.168.60.16` — recognised from the\ndefinition's event type (`vim.event.ResourceExhaustionStatusChangedEvent`,\n`com.vmware.vc.system.RootPasswordExpiredEvent`) and its `_sourcehost_` comparison. It is\nnull otherwise. `entity_name` is unchanged, because vmware-aiops resolves it.\n\n## 2a. Performance, capacity and vSphere 9.1 reads — result fields\n\n| Tool (CLI) | Field | Meaning |\n|------------|-------|---------|\n| `vm_performance` (`perf vms`) | `mem_ballooned_mb` | `mem.vmmemctl.average` — memory reclaimed by the balloon driver; above 0 means host memory pressure or a limit |\n| `vm_performance` | `mem_swapped_mb` | `mem.swapped.average` — guest memory the host swapped to disk |\n| `vm_performance` | `mem_usage_pct` | `mem.usage.average` — *active* guest memory over configured, not consumed; can read low on a VM under pressure (CLI column \"Active mem %\") |\n| `vm_performance` | `counters` (envelope) | The vSphere counter and meaning behind every row field |\n| `datastore_capacity` (`capacity datastores`) | `view` / `view_note` (envelope) | `vCenter` / `ESXi host` / `unknown`: whose figure this is. Each endpoint sums provisioned space over the VMs *it* has registered, so a vCenter and a directly-reached ESXi host can differ for one datastore |\n| `datastore_capacity` | `vm_count` | VMs the endpoint has registered on the datastore — the ones `provisioned_gb` covers |\n| `datastore_capacity` | `vms_not_connected` | Those VMs the endpoint reports as orphaned / inaccessible / disconnected, as `name (state)`; null when the state could not be read |\n| `vcenter_deployment_size` (`deployment-size`) | `available: false`, `reason`, `requires` | The appliance reported a version below 9.1 — an answer, not a failure (audited `ok`). If the version cannot be read, the call still returns an error that explains the 9.1 floor without claiming a build. An ESXi target gets an error saying these REST reads need a vCenter target |\n\n### Alarm/Event `suggested_actions` example\n\n`get_alarms` and `get_events` results include a `suggested_actions` list — each\nitem is a ready-to-use hint pointing to the correct companion skill and tool:\n\n```json\n{\n  \"alarm_name\": \"VM CPU Ready High\",\n  \"entity_name\": \"prod-db-01\",\n  \"suggested_actions\": [\n    \"vmware-aiops: acknowledge_vcenter_alarm(entity_name='prod-db-01', alarm_name='VM CPU Ready High')\",\n    \"vmware-aiops: reset_vcenter_alarm(entity_name='prod-db-01', alarm_name='VM CPU Ready High')\"\n  ]\n}\n```\n\n### Monitored Event Types\n\n| Category | Events |\n|----------|--------|\n| VM Failures | `VmFailedToPowerOnEvent`, `VmDiskFailedEvent`, `VmFailoverFailed` |\n| Host Issues | `HostConnectionLostEvent`, `HostShutdownEvent`, `HostIpChangedEvent` |\n| Storage | `DatastoreCapacityIncreasedEvent`, SCSI high latency |\n| HA/DRS | `DasHostFailedEvent`, `DrsVmMigratedEvent`, `DrsSoftRuleViolationEvent` |\n| Auth | `UserLoginSessionEvent`, `BadUsernameSessionEvent` |\n\n## 3. VM Info & Snapshot List\n\n| Feature | Details |\n|---------|---------|\n| VM Info | Name, power state, guest OS, CPU, memory, IP, VMware Tools, disks, NICs, `folder_path` |\n| Snapshot List | List existing snapshots with name and creation time (no create/revert/delete) — CLI `vm snapshot-list`, MCP tool `vm_list_snapshots` |\n| Backup Window | How long backups held a snapshot open on a VM, from vCenter task history — CLI `snapshots backup-window`, MCP tool `vm_backup_snapshot_history`. A lower bound on the backup job, never its official duration |\n\n## 4. Scheduled Scanning & Notifications\n\n| Feature | Details |\n|---------|---------|\n| Daemon | APScheduler-based, configurable interval (default 15 min) |\n| Multi-target Scan | Sequentially scan all configured vCenter/ESXi targets |\n| Scan Content | Each cycle: triggered alarms, vCenter events from the last `lookback_hours`, and new lines in the ESXi host logs `hostd`, `vmkernel`, `vpxa` |\n| Host Logs | Read incrementally: each line is reported once per daemon run (a restart re-reads each log's last 500 lines once). A rotated log, or more than 500 new lines between cycles, adds an `info` row. Needs `Global.Diagnostics` (not in vCenter's Read-Only role); an unreadable log becomes an `info` row with the reason |\n| Log Analysis | Host-log lines matching error, fail, critical, panic, lost access, cannot, timeout, refused, corrupt — critical/panic/corrupt lines are `critical`, the rest `warning` |\n| Webhook | Slack, Discord, or any HTTP endpoint. Every critical issue and every alarm/event warning; host-log warnings stay in `scan.log`; `info` rows are never sent |\n| Cycle Summary | One line per cycle: findings (and how many were sent), unreadable host logs, logs with unscanned lines, failed passes; `Scan INCOMPLETE` if any pass failed or a target could not be reached |\n\n## Safety Features\n\n| Feature | Details |\n|---------|---------|\n| Code-Level Isolation | Independent repository — zero destructive functions in codebase, gated by an AST allowlist over every vSphere call |\n| Audit Trail | MCP tool calls and every CLI command that reaches vCenter logged to `~/.vmware/audit.db` (SQLite WAL, via vmware-policy); CLI queries also append to `~/.vmware-monitor/audit.log` (JSON Lines) |\n| Password Protection | `.env` file loading with permission check (warn if not 600) |\n| SSL Self-signed Support | `verify_ssl: false` — **only** for ESXi hosts with self-signed certificates in isolated lab/home environments. Production environments should use CA-signed certificates with full TLS verification enabled. |\n\n## FORBIDDEN Operations — DO NOT EXIST IN CODEBASE\n\nThese operations are **not** performed by this skill — no code path here does them, and the allowlist gate fails if one is added:\n\n- `vm power-on/off`, `vm reset`, `vm suspend`\n- `vm create/delete/reconfigure`\n- `vm snapshot-create/revert/delete`\n- `vm clone/migrate`\n\nDirect users to **VMware-AIops** (`uv tool install vmware-aiops`) for these.\n\n## Version Compatibility\n\n| vSphere / VCF Version | Support | Notes |\n|----------------|---------|-------|\n| VCF 9.1 / vSphere 9.1 | ✅ Full | Released 2026-05-12. pyVmomi `<10.0` resolves and connects via SOAP. |\n| VCF 9.0 / vSphere 9.0 | ✅ Full | pyVmomi 8.0.3+ connects against vSphere 9 SOAP API. |\n| 8.0 / 8.0U1-U3 | Full | pyVmomi 8.0.3+ |\n| 7.0 / 7.0U1-U3 | Full | All read-only APIs supported |\n| 6.7 | Compatible | Backward-compatible, tested |\n| 6.5 | Compatible | Backward-compatible, tested |\n\n> pyVmomi auto-negotiates the API version during SOAP handshake — no manual configuration needed.\n\nFile v1.16.0:references/cli-reference.md\n\n# CLI Reference\n\nFull command reference for `vmware-monitor`.\n\n> **Output shape**: the CLI renders tables, so command output is unchanged. The\n> equivalent MCP tools wrap their rows in the list envelope\n> (`{items, returned, limit, total, truncated, hint}`) — see\n> [`capabilities.md`](capabilities.md#list-result-envelope).\n\n## Diagnostics\n\n```bash\nvmware-monitor doctor [--skip-auth]\n```\n\nChecks config file, connectivity, authentication, and pyVmomi version. Use `--skip-auth` to test config parsing without connecting.\n\n## MCP Config Generator\n\n```bash\nvmware-monitor mcp-config generate --agent <goose|cursor|claude-code|continue|vscode-copilot|localcowork|mcp-agent>\nvmware-monitor mcp-config list\n```\n\nGenerates MCP configuration files for supported agents. `list` shows all available agent templates.\n\n## Cluster Health Summary\n\n```bash\nvmware-monitor summary [--top <n>] [--cluster <substring>] [--no-vms] [--html] [--html-path <file>] [--target <name>]\n```\n\n- One aggregated read across all clusters. Leads with a ranked **top-N issues** focus list (the individual anomalies — disconnected hosts, triggered alarms, capacity/HA — worst first, each with a drill-down hint), then a per-cluster table: hosts connected/total, VM power rollup, live CPU/memory %, HA/DRS, alarm counts, and an opinionated `status` (ok/warn/critical) with `attention` reasons.\n- `--top`: size of the focus list (default 10; `--top 5` tighter, `--top 0` hides it and shows only the table). The header shows \"Top N (of TOTAL)\" when truncated.\n- `--cluster`: case-insensitive substring; show only matching clusters (also suppresses the standalone-hosts bucket).\n- `--no-vms`: skip the VM rollup pass (faster on very large fleets when only host/alarm/capacity signals are needed).\n- `--html`: write a self-contained, offline HTML snapshot (no external CSS/JS/fonts — nothing leaves the machine) to `~/vmware-health/cluster-health-<vc>-<YYYYMMDD-HHMMSS>.html`. The timestamped filename turns a folder of snapshots into a browsable point-in-time history. It is a snapshot, not a live page — re-run to refresh.\n- `--html-path <file>`: write the HTML snapshot to an explicit path instead of the auto-timestamped default (implies `--html`).\n- The rendered view is customizable — columns, thresholds, and layout live in [`health-summary-template.md`](health-summary-template.md). MCP tool: `cluster_health_summary`.\n\n## Object Investigation (drill-down)\n\n```bash\nvmware-monitor investigate vm <name>        [--hours <n>] [--html] [--html-path <file>] [--target <name>]\nvmware-monitor investigate host <name>      [--hours <n>] [--html] [--html-path <file>] [--target <name>]\nvmware-monitor investigate datastore <name> [--hours <n>] [--html] [--html-path <file>] [--target <name>]\n```\n\n- \"What is happening around this object?\" — one call *correlates* the object with its surrounding infrastructure and recent history, so the model explains an aggregated result instead of stitching several tools. Read-only; all cross-object reads are batched (cheap on large fleets).\n- **vm**: VM state, the host it runs on, its cluster context, backing datastores, snapshots, alarms (across VM/host/cluster/datastore), live performance, and a merged newest-first **event timeline** correlating the VM/host/cluster/datastore.\n- **host**: host state (connection, CPU/memory, ESXi version, uptime), cluster context, a rollup of the VMs it runs, mounted datastores, alarms, performance, correlated timeline.\n- **datastore**: capacity/free/accessibility, the hosts mounting it, a rollup of the VMs it backs, alarms, correlated timeline (per-datastore latency is a separate perf report, not included).\n- `--hours`: event-timeline look-back window (default 24).\n- `--html` / `--html-path`: self-contained offline snapshot to `~/vmware-health/investigate-<kind>-<object>-<timestamp>.html`; drill-down sections collapse/expand natively (no JS), nothing uploaded.\n- Unknown object names return a *teaching* error naming how to list objects. MCP tools: `vm_investigation_bundle`, `host_investigation_bundle`, `datastore_investigation_bundle`.\n\n## Cross-vCenter Attention\n\n```bash\nvmware-monitor attention [--top <n>] [--cluster <substring>] [--html] [--html-path <file>]\n```\n\n- \"What needs attention now?\" across **every** configured vCenter — one globally-ranked `top_issues` list (each tagged with its `vcenter`) plus a per-target table. Use it instead of running `summary` once per target and merging by hand.\n- Degrades gracefully: a target that can't be reached (or errors mid-summary) is listed as `unreachable` with a reason, and the rest still aggregate — one dead vCenter never sinks the view.\n- `--top`: size of the merged focus list (default 10). `--cluster`: case-insensitive substring applied to every target.\n- `--html` / `--html-path`: offline snapshot to `~/vmware-health/attention-<timestamp>.html`. MCP tool: `cross_vcenter_attention`.\n\n## Inventory\n\n```bash\nvmware-monitor inventory vms [--target <name>] [--limit <n>] [--sort-by name|cpu|memory_mb|power_state|folder_path] [--power-state poweredOn|poweredOff] [--folder-filter <pattern>]\nvmware-monitor inventory hosts [--target <name>]\nvmware-monitor inventory datastores [--target <name>]\nvmware-monitor inventory clusters [--target <name>]\nvmware-monitor inventory networks [--target <name>]\n```\n\n- `--target`: Named target from `config.yaml` (default: `default_target` in `config.yaml`, else the first target)\n- `--limit`: Max VMs to return (default: unlimited)\n- `--sort-by`: Sort field for VM listing — `name` | `cpu` | `memory_mb` | `power_state` | `folder_path`\n- `--power-state`: Filter VMs by power state\n- `--folder-filter`: Case-insensitive substring match against `folder_path` (e.g. `--folder-filter \"Production\"` returns VMs anywhere under a folder containing \"production\", including nested subfolders like `/Datacenters/Production/Web Tier`).\n\n**Output fields** (CLI and MCP): each VM entry includes `folder_path` — the vCenter inventory folder path (e.g. `/Datacenters/Production/Web Tier`). Present in both compact and full modes.\n\n## Health\n\n```bash\nvmware-monitor health alarms [--target <name>]\nvmware-monitor health events [--hours 24] [--severity warning] [--start <iso>] [--end <iso>] [--include-routine] [--target <name>]\nvmware-monitor health sensors [--target <name>]\nvmware-monitor health services [--host <esxi-name>] [--target <name>]\n```\n\n- `--hours`: Time range for event query (default: 24)\n- `--severity`: Minimum severity filter — `info`, `warning`, `error`, `critical` (default: `warning`)\n- `--start` / `--end`: Query a specific window instead of the last `--hours` — ISO 8601, a time with no zone is UTC (e.g. `--start 2026-09-03T12:00Z --end 2026-09-03T16:00Z`). `--start` alone runs to now.\n- `--include-routine`: List routine login/logout events too. By default they are folded and counted, because one local agent's logins can outnumber everything else\n- `--host` (services only): Filter service status to a single host by exact name (default: all hosts)\n- `sensors`: a connected host that reports no sensors is named under the table with its CIM Server (`sfcbd-watchdog`) state — ESXi reads sensors through it, so a stopped CIM Server means none are reported; a running one usually means no BMC/IPMI device. MCP `get_host_sensors` returns the same as `hosts_without_sensors` and `sensors_note`\n\n## VM Info (Read-Only)\n\n```bash\nvmware-monitor vm info <vm-name> [--target <name>]\nvmware-monitor vm snapshot-list <vm-name> [--target <name>]\n```\n\nReturns detailed VM information: CPU, memory, disks, NICs, guest OS, IP, VMware Tools status, and snapshots.\n\n`snapshot-list` lists existing snapshots with name and creation time. No create, revert, or delete operations exist. The same data is exposed via the MCP tool `vm_list_snapshots`.\n\n```bash\nvmware-monitor snapshots backup-window <vm> [--days 30] [--cycles]\n```\n\nInfers backup windows for one VM from its snapshot task history: image-level\nbackup products snapshot the VM, copy the frozen disks, then delete the\nsnapshot, and vCenter records both ends. Reports four hour statistics\n(backup active / snapshot present / total window / removal itself) plus\ncreations that never got a matching removal.\n\nTwo lines in the output change what the numbers mean, and both are printed\nrather than left to the reader: a note when vCenter could not be read at all\n(which is not \"no backups\"), and a note when part of the requested window has\nalready aged out of task history (`task.maxAge`, 30 days by default — the\noption is not called `vpxd.task.maxAge`; that name raises `vim.fault.InvalidName`).\nDuplicate VM names are refused rather than resolved to one. The same data is\nexposed via the MCP tool `vm_backup_snapshot_history`.\n\n## vSphere 9.1 (read-only)\n\n```bash\nvmware-monitor memory tiering [--host <esxi>] [--limit <n>] [--target <name>]\nvmware-monitor patch compliance <cluster-moid> [--target <name>]\nvmware-monitor patch last-apply <cluster-moid> [--target <name>]\nvmware-monitor deployment-size [--target <name>]\n```\n\n- `memory tiering`: Per-host memory tiering (DRAM/NVMe tiers) and NVMe uplift ratio via pyVmomi `hardware.memoryTierInfo`. Requires ESXi/vCenter **8.0U3+**; on older targets the command fails with a teaching error naming the missing property. Same data as the MCP tool `host_memory_tiering`.\n- `patch compliance`: vLCM software (patch) compliance for one cluster over vSphere Automation REST. `<cluster-moid>` is the cluster **MoID** (e.g. `domain-c123`) — get it from `inventory clusters`, not the display name. When vCenter is mid-patch it answers 503 and the command prints a \"state unavailable, retry shortly\" line instead of crashing; `non-compliant` shows `?` when the host-status field is unknown (never a false 0). MCP tool: `cluster_patch_compliance`.\n- `patch last-apply`: Result (status + end time) of the last vLCM remediation on one cluster. `<cluster-moid>` as above. Reports the outcome only; it never runs a remediation. MCP tool: `cluster_last_apply_result`.\n- `deployment-size`: vCenter appliance deployment size class over REST (NEW in vSphere 9.1). MCP tool: `vcenter_deployment_size`.\n\n> These four commands resolve config **without opening a SOAP session** and only ever GET, so they still degrade gracefully when vCenter is mid-patch (answering 503). The three REST commands' endpoints are spec-verified, but their JSON field names are read defensively pending a live 9.1 vCenter replay — each result carries a `note` that self-labels the parse as best-effort.\n\n## Scanning & Daemon\n\n```bash\nvmware-monitor scan now [--target <name>]\nvmware-monitor scan logs [--host <name>] [--lines 500] [--raw] [--target <name>]\nvmware-monitor daemon start\nvmware-monitor daemon stop\nvmware-monitor daemon status\n```\n\n- `scan now`: Run a one-time scan of alarms and events. It does not read ESXi host logs — use `scan logs` (MCP `host_log_scan`); the daemon's host-log pass reads them too\n- `scan logs`: Read the last `--lines` lines of hostd/vmkernel/vpxa on each host (or `--host`) and list the lines matching a trouble pattern, grouped by pattern with a count, the hosts and the log time span; severity follows the level ESXi wrote on the line. `--raw` prints one row per line\n- `daemon start`: Start APScheduler-based background scanner (default: every 15 min)\n- `daemon stop`: Stop the background scanner\n- `daemon status`: Check if the daemon is running\n\n## Setup\n\n```bash\nmkdir -p ~/.vmware-monitor\ncp config.example.yaml ~/.vmware-monitor/config.yaml\ncp .env.example ~/.vmware-monitor/.env\nchmod 600 ~/.vmware-monitor/.env\n```\n\nCopies the `config.yaml` and `.env` templates into `~/.vmware-monitor/`; then edit them with your target details.\n\n## Architecture\n\n```\nUser (Natural Language)\n  |\nAI Tool (Claude Code / Aider / Gemini / Codex / Cursor / Trae / Kimi)\n  |\n  +-- CLI mode (default): vmware-monitor CLI --> pyVmomi --> vSphere API\n  |\n  +-- MCP mode (optional): MCP Server (stdio) --> pyVmomi --> vSphere API\n  |\nvCenter Server --> ESXi Clusters --> VMs\n    or\nESXi Standalone --> VMs\n```\n\nFile v1.16.0:references/health-summary-template.md\n\n# Cluster Health Summary — Display Template\n\nThis is the **editable template** for the `cluster_health_summary` view (CLI:\n`vmware-monitor summary`, MCP tool: `cluster_health_summary`). It is intentionally\nopinionated and intentionally *not* a widget builder — the goal is the 5-second\n\"is anything on fire?\" glance, not a replacement for Aria Operations dashboards.\n\nEverything below is meant to be changed. Add columns, drop columns, retune\nthresholds, regroup, or switch the output to an HTML page. The tool returns clean\nstructured data; **how it is displayed is entirely up to this template and to what\nthe operator asks for in plain language.**\n\n> **Not affiliated with, endorsed by, or sponsored by VMware, Inc. or Broadcom Inc.**\n> \"VMware\" and \"vSphere\" are trademarks of Broadcom.\n\n---\n\n## Default layout\n\nThe rollup returns `{ totals, top_issues, issues_total, clusters, snapshot,\ncustomization_hint }` and renders in two parts: a **top-N focus list** (the\nheadline) followed by the **per-cluster table** (the context).\n\n### Part 1 — Top-N issues (the focus list)\n\nFor a large fleet, scanning every cluster row is too slow, so the individual\nanomalies are flattened into one ranked list — the top N things wrong right now,\nworst first. This is the headline; lead with it.\n\nThe terminal shows four columns so the text that matters stays readable at 80\ncolumns:\n\n| Column               | Field(s)                  | Meaning |\n|----------------------|---------------------------|---------|\n| #                    | (position)                | Rank, 1 = most urgent |\n| Severity             | `severity`                | `CRITICAL` / `WARNING` |\n| Object               | `object`, then `cluster`  | The host, cluster, datastore or vCenter appliance the problem is on, with the cluster it rolls up to on the line below |\n| Problem · next step  | `detail`, then a CLI hint | Plain-language description (alarm name, \"memory at 96%\", \"host notResponding\"), with the CLI command to run next on the line below |\n\n`object` names the vCenter appliance (e.g. `vCenter appliance 192.168.60.16`)\nfor alarms vCenter raises about itself — memory exhaustion, root password\nexpiry — recognised from the alarm definition's event type, not its name.\n\nAlarm issues also carry `condition_now` (`holds` / `cleared` / `unknown`) and\n`acknowledged_days` (null when not acknowledged). The `detail` says when a\ncondition no longer holds, or when an alarm acknowledged 7+ days ago has a\ncondition that is not re-checked.\n\n`drilldown` in the data names MCP tools (it is read by a model); the terminal\nand the HTML snapshot name CLI commands instead.\n\n`top_n` (CLI `--top`, default 10) caps the list; `issues_total` reports the\npre-cap count so truncation is visible (e.g. \"Top 5 issues (of 8)\"). `--top 0`\nhides the list and shows only the table. Ranking: `cleared` alarms (the\ncondition is read to be false now) after every other issue; then by severity,\nthen kind (`host_down` → `alarm` → `capacity` → `config`), then hottest capacity\nfirst. An `unknown` alarm keeps its severity rank however long ago it was\nacknowledged — an acknowledgement is not evidence the condition went away.\nTune the order in `staleness_rank` (`ops/alarm_triage.py`) and `_KIND_RANK` /\n`_rank_issues` (`ops/cluster_summary.py`).\n\n### Part 2 — Per-cluster table\n\nThe rollup's `clusters` array renders one row per cluster, worst status first:\n\n| Column       | Field                         | Meaning |\n|--------------|-------------------------------|---------|\n| Status       | `status`                      | Opinionated verdict: `CRITICAL` / `WARN` / `OK` |\n| Cluster      | `name`                        | Cluster name (`(standalone hosts)` bucket for un-clustered hosts) |\n| Hosts        | `hosts_connected`/`hosts_total` | Connected vs total ESXi hosts |\n| VMs on       | `vms_on`/`vms_total`          | Powered-on vs total VMs (omitted when `include_vms=false`) |\n| CPU%         | `cpu_used_pct`                | Live cluster CPU utilisation |\n| Mem%         | `mem_used_pct`                | Live cluster memory utilisation |\n| HA           | `ha_enabled`                  | vSphere HA on/off; `null` (shown `n/a`) on the `(standalone hosts)` and `(vCenter-level)` rows |\n| DRS          | `drs_enabled`                 | DRS on/off; `null` (shown `n/a`) on the same two rows |\n| Alarms C/W   | `alarms.critical`/`.warning`  | Triggered alarm counts (cluster + its hosts) |\n| Attention    | `attention[]`                 | Plain-language reasons the status is not OK |\n\nHeader line shows cross-cluster totals and the single worst status. The last line\nis always the **friendly customization hint** (`customization_hint`) — keep it.\n\n### The three signals (why these columns)\n\nAn operator's first look is always the same three questions. The default columns\nmap to exactly these and nothing more:\n\n- **Problems** → `Status`, `Alarms C/W`, `Attention`, disconnected hosts\n- **Capacity** → `CPU%`, `Mem%`\n- **Health**  → `HA`, `DRS`, hosts connected\n\nIf a proposed new column does not answer one of those three, it probably belongs\nin a drill-down tool (`get_alarms`, `datastore_capacity`, `host_performance`,\n`vm_info`), not here.\n\n---\n\n## Opinionated thresholds\n\nDefined in `vmware_monitor/ops/cluster_summary.py` as named constants — edit there\nto retune:\n\n| Constant             | Default | Drives |\n|----------------------|:-------:|--------|\n| `CPU_MEM_WARN_PCT`   | 85      | CPU or memory at/above → contributes `warn` |\n| `CPU_MEM_CRIT_PCT`   | 95      | CPU or memory at/above → `critical` |\n| `DATASTORE_OVERCOMMIT_WARN_PCT` (in `ops/capacity.py`) | 100 | Datastore provisioned above this % of capacity → a `warning` capacity issue in `top_issues` (`scope: datastore`); the same line `capacity datastores` colours red |\n\nStatus is also forced to `critical` by **any disconnected host** or **any critical\nalarm**, and to at least `warn` by warning alarms or **HA disabled on a multi-host\ncluster**. The HA rule applies only to real clusters — HA is a cluster setting, so\nthe `(standalone hosts)` row never gets it. Change `_rollup_status()` to add or relax rules.\n\n---\n\n## Adding a column — worked example (datastore free space)\n\nDatastore headroom is the most common request. It is deliberately left out of the\ndefault view so this template can show the full add-a-column path end to end:\n\n1. **Data** — in `get_cluster_health_summary`, add a fourth batched pass over\n   `vim.Datastore` collecting `[\"name\", \"summary.freeSpace\", \"summary.capacity\",\n   \"host\"]`. Map each datastore to clusters via its mounted hosts\n   (`host[].key` → `host_to_cluster`), and record the minimum free % per cluster\n   on the accumulator (e.g. `rec[\"ds_min_free_pct\"]`).\n2. **Status** — in `_rollup_status`, flag `warn` when `ds_min_free_pct` drops\n   below a new `DS_FREE_WARN_PCT` constant, appending a reason like\n   `\"datastore <name> at 8% free\"`.\n3. **Render** — add a `Datastore free%` column in `cluster_summary_cmd` (CLI) and\n   note the field in the table above. The MCP tool needs no change; the new field\n   flows through automatically.\n\nRemoving a column is the reverse: drop the render column and, if you no longer\nwant it to affect status, remove its rule from `_rollup_status`.\n\n---\n\n## Output modes\n\nThe **same structured data** renders three ways — pick per request, no code change:\n\n- **In chat (default):** a Markdown table of the `clusters` rows, ending with the\n  customization hint line. Best for a quick answer.\n- **CLI terminal:** `vmware-monitor summary` — colourised Rich table.\n- **Offline HTML snapshot:** `vmware-monitor summary --html` (or `--html-path\n  <file>`) renders `ops/health_html.py` → a **self-contained** file (all CSS\n  inlined, no external CSS/JS/fonts, theme-aware) written to\n  `~/vmware-health/cluster-health-<vc>-<YYYYMMDD-HHMMSS>.html`. It is offline by\n  design — internal host/cluster names never leave the machine (a cloud-hosted\n  artifact would upload them). The timestamped filename makes a folder of\n  snapshots a browsable point-in-time history. It is a **snapshot, not a live\n  page** — re-run to refresh. Every vSphere-sourced value is HTML-escaped at\n  render time. To restyle it, edit the `_CSS` block and the row/card builders in\n  `ops/health_html.py`; it reads the same `top_issues` + `clusters` fields, so no\n  raw inventory is ever embedded.\n\n---\n\n## The friendly closing line\n\nEvery rendered summary ends with one inviting line so the operator always knows\nthe view bends to them. The tool returns it as `customization_hint`; echo it\nverbatim, or adapt the phrasing while keeping the spirit:\n\n> Want a different view? Just say so — e.g. \"add datastore free space\", \"drop the\n> DRS column\", \"only show clusters that need attention\", \"add per-VM CPU ready\",\n> or \"render this as an HTML page\". Columns, thresholds, and grouping are all\n> adjustable.\n\nThis line is the whole point of the template being editable: the operator does not\nneed to know the field names — they describe what they want in words and the\nassistant reshapes the table.\n\nFile v1.16.0:references/investigation-protocol.md\n\n# Investigation Protocol — Causal Chain Root Cause Analysis\n\nA protocol for AI agents performing diagnostic investigations on VMware infrastructure (alarms, performance regressions, availability incidents). Adopted from Enterprise Harness Engineering, drawing on 5 Whys, Google SRE, ITIL, and NASA Fault Tree Analysis.\n\n## When to Apply\n\nUse this protocol whenever the user asks:\n\n- \"Why is X slow / failing / down?\"\n- \"What caused this alarm / alert / incident?\"\n- \"Investigate / diagnose / debug …\"\n- Any open-ended question that requires identifying a root cause rather than just reading state.\n\nDo NOT apply for:\n\n- Simple state lookups (\"is the VM on?\", \"list datastores\")\n- Operational requests (\"clone this VM\", \"create this rule\")\n- Configuration questions (\"what's the default for X?\")\n\n## The Four Criteria for Root Cause Completeness\n\nA diagnostic conclusion is **incomplete** unless ALL four criteria are satisfied. The agent must self-check against each one before outputting a report.\n\n### 1. Falsifiability (可证伪性)\n\nThe root cause must be independently measurable and verifiable. If you cannot test it, it is a hypothesis, not a root cause.\n\n- ✅ \"Datastore latency exceeded 50ms because IOPS hit the SAN cap of 10,000\" — directly testable via `get_metrics datastore.iops`\n- ❌ \"Network was congested\" — too vague to verify\n\n### 2. Sufficiency (充分性)\n\nRemoving the root cause must make the symptom disappear. If the symptom persists after the supposed fix, the cause was wrong or partial.\n\n- ✅ \"Deleting the orphaned snapshot freed 200 GB and the alarm cleared within 60 seconds\"\n- ❌ \"Restarted the VM and the issue went away\" — correlation, not causation\n\n### 3. Necessity (必要性)\n\nThe symptom must occur whenever the root cause is present. If the same condition exists elsewhere without the symptom, you have not found the true root cause.\n\n- ✅ \"Every cluster with 80%+ memory overcommit shows the same vMotion stall\"\n- ❌ \"Only this one VM has the issue\" — without explaining why this VM specifically\n\n### 4. Mechanism (机制性)\n\nYou must explain the propagation chain: root cause → propagation → amplification → impact. A single point claim with no mechanism is a guess.\n\n- ✅ \"Snapshot delta files filled the datastore (root) → VM I/O blocked on write (propagation) → guest filesystem went read-only (amplification) → application timeout (impact)\"\n- ❌ \"The datastore was full\" — describes a state, not a chain\n\n## Investigation Workflow — Up to Three Depth Rounds\n\n### Round 1 — Initial Hypothesis\n\n1. Gather symptoms via L1/L2 read tools (alarms, metrics, events, logs)\n2. Form an initial causal chain hypothesis\n3. Apply the four criteria\n\nIf all four pass → output report.\nIf any criterion fails → proceed to Round 2 with that criterion as the focus.\n\n### Round 2 — Targeted Deepening\n\n1. Identify which criterion failed\n2. Gather additional evidence aimed specifically at that criterion (e.g. failed Necessity → compare against unaffected peers; failed Mechanism → trace next propagation step)\n3. Refine the causal chain\n4. Re-apply the four criteria\n\nIf all four pass → output report.\nIf any still fails → proceed to Round 3.\n\n### Round 3 — Final Deepen or Escalate\n\n1. If a deeper cause is reachable, gather final evidence and finalize the chain\n2. If evidence is unavailable, system-bounded, or beyond the agent's tool surface, **escalate to a human** and explicitly label the conclusion as `⚠️ INCOMPLETE — <criterion> unsatisfied`\n3. **Never** silently output a partial conclusion as if it were complete\n\n## Output Format\n\nEvery investigation report must structure findings exactly as:\n\n```\n🔴 [ROOT CAUSE]   <falsifiable, mechanism-explained statement>\n  → [PROPAGATION] <how the root cause spread to neighboring systems>\n    → [AMPLIFICATION] <what made the impact worse, if applicable>\n      → [IMPACT]    <observable user / business / SLA effect>\n\n✅ Falsifiability:  <evidence — metric name, log query, command output>\n✅ Sufficiency:     <evidence or stated counterfactual>\n✅ Necessity:       <evidence or peer comparison>\n✅ Mechanism:       <see propagation chain above>\n```\n\nIf any criterion is unmet, mark it `⚠️ INCOMPLETE — <reason>` and state explicitly what additional evidence would be required to satisfy it.\n\n## Anti-Patterns\n\n| ❌ Pattern | Why it fails |\n|---|---|\n| \"thanos-cn unreachable\" alone | Describes symptom; does not answer **why** unreachable |\n| \"Datastore full\" alone | No propagation, no impact chain |\n| \"Try restarting it\" | Skips diagnosis entirely |\n| \"Probably the network\" | Not falsifiable |\n| Stopping at the first plausible cause | Skips Necessity check |\n| Silent partial conclusion | Hides incompleteness from the user |\n\n## Worked Examples\n\n### Bad — Incomplete Diagnosis\n\n> \"VM is slow because the host is busy.\"\n\nMissing:\n- **Falsifiability**: which metric, what threshold?\n- **Necessity**: why this VM only?\n- **Mechanism**: how does host load translate into VM slowness?\n\n### Good — Complete Diagnosis\n\n> 🔴 [ROOT] Host `esx-03` CPU ready time exceeds 15% (validated via `get_metrics host.cpu.ready`)\n>   → [PROPAGATION] vCPU contention from 4-VM reservation collision in resource pool `prod-rp`\n>     → [AMPLIFICATION] DRS is in manual mode, so VMs are not rebalanced\n>       → [IMPACT] Application p99 latency doubled from 200 ms to 400 ms\n>\n> ✅ Falsifiability: `host.cpu.ready` metric directly observable; threshold defined in vSphere docs\n> ✅ Sufficiency: vMotion `vm-A` off `esx-03` reduced ready time to 3% and p99 latency back to 200 ms\n> ✅ Necessity: only VMs in `prod-rp` with active reservations are affected; identical workloads in `staging-rp` are healthy\n> ✅ Mechanism: cpu.ready = vCPU waiting for pCPU → guest perceives as CPU starvation → app threadpool exhaustion → tail latency\n\n## Related Skills\n\nA complete investigation often chains across skills:\n\n- **vmware-monitor** (this skill): inventory, alarms, events — read-only data source\n- [vmware-aria](https://github.com/vmware-skills/VMware-Aria): metrics, alerts, anomaly detection — primary L1/L2 data source for time-series analysis\n- [vmware-aiops](https://github.com/vmware-skills/VMware-AIops): VM/host state, deployment history; can also remediate at L3+ once the investigation is complete and approved\n- [vmware-pilot](https://github.com/vmware-skills/VMware-Pilot): orchestrate the investigation itself as a multi-step Dispatcher → Subagent workflow\n\nThe agent should treat investigation as **read-heavy first**: gather across skills, reason centrally, only invoke L3+ write tools after the four criteria are satisfied AND the user has approved a remediation plan.\n\nFile v1.16.0:references/setup-guide.md\n\n# Setup Guide\n\n## Install\n\nAll install methods fetch from the same source: [github.com/vmware-skills/VMware-Monitor](https://github.com/vmware-skills/VMware-Monitor) (MIT licensed). We recommend reviewing the source code before installing.\n\n```bash\n# Via PyPI (recommended for version pinning)\nuv tool install vmware-monitor==1.16.0\n\n# Via Skills.sh (fetches from GitHub)\nnpx skills add vmware-skills/VMware-Monitor#v1.16.0\n\n# Via ClawHub (fetches from ClawHub registry snapshot of GitHub)\nclawhub install @zw008/vmware-monitor --version 1.16.0\n```\n\n### Claude Code\n\n`npx skills add` and `clawhub install` both place the skill in Claude Code's skills\ndirectory. To install it manually from a clone:\n\n```bash\nmkdir -p ~/.claude/skills/vmware-monitor\ncp -r skills/vmware-monitor/. ~/.claude/skills/vmware-monitor/\n```\n\nFor tool access (not just skill context), register the MCP server:\n\n```bash\nclaude mcp add vmware-monitor -- vmware-monitor mcp\n```\n\n## What Gets Installed\n\nThe `vmware-monitor` package installs a Python CLI binary — read-only toward vSphere — and its dependencies (pyVmomi, Click, Rich, APScheduler, python-dotenv). No background services, daemons, or system-level changes are made during installation. The scheduled scanner (`daemon start`) only runs when explicitly started by the user.\n\n### What \"read-only\" covers, and what it writes locally\n\n\"Read-only\" is a claim about vCenter/ESXi: no code path changes their state. On\nvCenter the skill opens only its own login session (SOAP, plus a REST session for\nthe three vSphere 9.1 REST tools) and short-lived query handles — container views\nand a task-history collector — that it releases after use.\n\nIt does write to the local machine — by default under your home directory:\n\n| Path | Written by | When |\n|---|---|---|\n| `~/.vmware-monitor/config.yaml`, `~/.vmware-monitor/.env` | `vmware-monitor init` | Once, when you run it (`.env` is set to 0600) |\n| `~/.vmware-monitor/.env` (rewritten in place) | Every run | Only if it holds a plaintext `*_PASSWORD`: the value is rewritten as `b64:` (see below) |\n| `~/.vmware/audit.db` (SQLite; `OPS_HOME` moves it) | vmware-policy | Every MCP tool call, and every CLI command that reaches vCenter |\n| `~/.vmware-monitor/audit.log` (JSON Lines) | CLI | Every CLI query command |\n| `~/vmware-health/*.html` | `summary` / `investigate` / `attention` | Only with `--html` (or `--html-path <file>`) |\n| `~/.vmware-monitor/daemon.pid`, `~/.vmware-monitor/scan.log` (`notify.log_file` moves the log) | Scanner daemon | Only after `daemon start` |\n| The file you name | `mcp-config generate --output <file>` | Only with `--output`; otherwise it prints |\n\nOutbound traffic beyond vCenter/ESXi: the daemon posts to your Slack/Discord\nwebhook, only if you configured one and a scan finds an issue it sends — any\ncritical issue, or an alarm/event warning (host-log warnings and `info` rows\nare never sent).\n\n## Configuration\n\n```bash\n# 1. Install\nuv tool install vmware-monitor==1.16.0\n\n# 2. Verify\nvmware-monitor --version\n\n# 3. Configure\nmkdir -p ~/.vmware-monitor\ncp config.example.yaml ~/.vmware-monitor/config.yaml\ncp .env.example ~/.vmware-monitor/.env\nchmod 600 ~/.vmware-monitor/.env\n# Edit ~/.vmware-monitor/config.yaml and .env with your target details\n```\n\n### Declare `environment:` on each target\n\n```yaml\ntargets:\n  - name: prod-vcenter\n    host: vcenter-prod.example.com\n    environment: production   # production | staging | lab | <your own label>\n```\n\n`environment` is an optional label. This skill has zero write tools, so nothing\nit exposes is ever gated by it — reads are never gated under any setting. It\nmatters for the write skills (`vmware-aiops`, `vmware-storage`, `vmware-nsx`)\npointed at the same vCenter: an environment-scoped `deny` rule in\n`~/.vmware/rules.yaml` can match on the label to block their writes (e.g. freeze\n`production`). A target with no label is simply not matched by such a rule.\n\n### Choose the default target\n\n```yaml\ndefault_target: prod-vcenter   # optional; without it the first entry under targets: is used\n```\n\nA tool uses the default only when the request names no target. The MCP server\nlists every configured target (name, type, host) in its instructions and asks the\nagent to pick the one the question is about: a vCenter for the environment,\nclusters or several hosts; the managing vCenter for a named ESXi host, since\nvCenter holds that host's alarms, events and tasks; a standalone `esxi` target\nonly when asked about that host directly. When the request does not say which and\nthe answers would differ, the agent is told to ask. A result from a tool that takes\n`target` carries `target: {name, type}` naming the target that answered, so the\nanswer can say where it came from; a request for a target name that is not\nconfigured returns an error listing the configured names instead. A\n`default_target` that names no configured target is an error, not a fallback —\nfalling back to the first entry is how a standalone ESXi host once answered a\nquestion about vCenter.\n\n## Development Install\n\n```bash\ngit clone --branch v1.16.0 https://github.com/vmware-skills/VMware-Monitor.git\ncd VMware-Monitor\nuv venv && source .venv/bin/activate\nuv pip install -e .\n```\n\n## Usage Mode\n\nChoose the best mode based on your AI tool:\n\n| Platform | Recommended Mode | Why |\n|----------|-----------------|-----|\n| Claude Code, Cursor | **MCP** | Structured tool calls, no interactive confirmation needed, seamless experience |\n| Aider, Codex, Gemini CLI, Continue | **CLI** | Lightweight, low context overhead, universal compatibility |\n| Ollama + local models | **CLI** | Minimal context usage, works with any model size |\n\n### Calling Priority\n\n- **MCP-native tools** (Claude Code, Cursor): MCP first, CLI fallback\n- **All other tools**: CLI first (MCP not needed)\n\n> **Tip**: If your AI tool supports MCP, check whether `vmware-monitor` MCP server is loaded (`/mcp` in Claude Code). If not, configure it first — MCP provides the best hands-free experience.\n\n## MCP Mode Configuration\n\nFor Claude Code / Cursor users who prefer structured tool calls, add to `~/.claude/settings.json`:\n\n```json\n{\n  \"mcpServers\": {\n    \"vmware-monitor\": {\n      \"command\": \"vmware-monitor\",\n      \"args\": [\"mcp\"],\n      \"env\": {\n        \"VMWARE_MONITOR_CONFIG\": \"~/.vmware-monitor/config.yaml\"\n      }\n    }\n  }\n}\n```\n\n> v1.5.15+ recommends the single-command form `vmware-monitor mcp`. Pre-1.5.15 used\n> `uvx --from vmware-monitor vmware-monitor-mcp`, which still works but re-resolves from <!-- install-pin: historical -->\n> PyPI on each launch and breaks behind corporate TLS proxies. The legacy\n> `vmware-monitor-mcp` entry point is also kept for backward compatibility.\n\nMCP exposes 32 read-only tools:\n\n| Group | Tools |\n|---|---|\n| Inventory | `list_virtual_machines`, `list_esxi_hosts`, `list_all_datastores`, `list_all_clusters`, `list_all_networks`, `resource_pool_usage` |\n| Health & triage | `get_alarms`, `get_events`, `cluster_health_summary`, `cross_vcenter_attention` |\n| VM detail | `vm_info`, `vm_list_snapshots`, `vm_performance`, `snapshot_aging` |\n| Host detail | `host_performance`, `host_log_scan`, `get_host_sensors`, `get_host_services`, `ntp_status` |\n| Platform state | `certificate_status`, `license_status`, `active_sessions`, `active_tasks`, `datastore_capacity` |\n| Investigation bundles | `vm_investigation_bundle`, `host_investigation_bundle`, `datastore_investigation_bundle` |\n| Patching & vSphere 9 | `cluster_patch_compliance`, `cluster_last_apply_result`, `host_memory_tiering`, `vcenter_deployment_size` |\n| Backups | `vm_backup_snapshot_history` |\n\nAll accept an optional `target` parameter except `cross_vcenter_attention`, which\nsweeps every configured target by design.\n\nList-style tools take `limit` (default 50) to keep large inventories from flooding\ncontext; `list_virtual_machines` additionally supports `sort_by`, `power_state`,\n`fields`, and `folder_filter`.\n\n### Password obfuscation at rest\n\nOn first load, any plaintext `*_PASSWORD` value in `.env` is automatically\nrewritten to a grep-safe `b64:<encoded>` form and decoded transparently at\nruntime, so a casual `grep` of the file no longer reveals the password. Values\nare read and written through python-dotenv's own parser, so the stored secret\nnever drifts from what you configured (quotes, inline comments, and trailing\nwhitespace are handled correctly).\n\n> **This is obfuscation, not encryption.** Anyone who can read the file can\n> still decode it. For real secrecy at rest, do not store the password in `.env`\n> at all — inject it from a secret manager (HashiCorp Vault, CyberArk, AWS\n> Secrets Manager, or a Kubernetes Secret) into the `*_PASSWORD` environment\n> variable at process start. The code reads the env var either way.\n\n## Security\n\n> **Disclaimer**: This is a community-maintained open-source project and is **not affiliated with, endorsed by, or sponsored by VMware, Inc. or Broadcom Inc.** \"VMware\" and \"vSphere\" are trademarks of Broadcom.\n\n- **Read-Only toward vSphere**: This is an independent repository with zero destructive code paths. No power off, delete, create, reconfigure, or migrate functions exist in the codebase. Enforced by [`tests/eval/regression/test_read_only_enforcement.py`](https://github.com/vmware-skills/VMware-Monitor/blob/main/tests/eval/regression/test_read_only_enforcement.py) — in the source repository, not in this skill bundle — which allowlists every vSphere method the package calls against pyVmomi's own type metadata. It is a gate on source code, not a runtime block, and this repo has no CI — for defence that does not depend on it, use a least-privilege account (next item). Local files the skill writes are listed under [What \"read-only\" covers](#what-read-only-covers-and-what-it-writes-locally).\n- **Least Privilege**: Give the skill a dedicated vCenter service account holding the built-in **Read-Only** role (propagated from the inventory root), not an administrator. That role has no privilege to change anything, so the read-only claim no longer rests on this code. Two reads need more than it grants: `active_sessions` reads the session list, which vCenter gates on `Sessions.TerminateSession` (a privilege that also allows ending sessions — without it the tool returns a one-row explanation instead); and ESXi host-log reads (`host_log_scan` and the daemon's host-log pass) call `BrowseDiagnosticLog`, gated on `Global.Diagnostics`. Grant those only if you need those reads.\n- **Source Code**: Fully open source at [github.com/vmware-skills/VMware-Monitor](https://github.com/vmware-skills/VMware-Monitor) (MIT). The `uv` installer fetches the `vmware-monitor` package from PyPI, which is built from this GitHub repository. We recommend reviewing the source code and commit history before deploying in production.\n- **TLS Verification**: Enabled by default. Setting `verify_ssl: false` is solely for ESXi hosts using self-signed certificates in isolated lab/home environments. In production, always use CA-signed certificates with full TLS verification.\n- **Credentials & Config**: This skill requires the following secrets, all stored in `~/.vmware-monitor/.env` (`chmod 600`, loaded via `python-dotenv`):\n  - `VMWARE_<TARGET>_PASSWORD` — per-target password where `<TARGET>` is the uppercased target name from `config.yaml` (hyphens become underscores). Example: target named `vcenter-prod` uses `VMWARE_VCENTER_PROD_PASSWORD`.\n  - (Optional) Webhook URLs for Slack/Discord notifications\n\n  The config file `~/.vmware-monitor/config.yaml` stores only target hostnames, ports, and usernames — it does **not** contain passwords or tokens. The env var `VMWARE_MONITOR_CONFIG` points to this YAML file.\n- **Webhook Data Scope**: Webhook notifications are **disabled by default**. When enabled, the daemon posts to **user-configured URLs only** (Slack, Discord, or any HTTP endpoint you control); no data is sent to any other service. Each payload carries critical/warning counts plus every critical issue and every alarm/event warning from that scan — host-log warnings go to `scan.log` only, and `info` rows (unreadable or partly-read host logs) are never sent. Each issue carries the entity name and one of: the alarm name, vCenter event message (sanitized, ≤500 chars), ESXi log line matching critical/panic/corrupt (sanitized, ≤200 chars), or the error text for a target the daemon could not connect to. Event, log, and error text can contain host names, IP addresses, and user names — treat the webhook destination as receiving operational data. No credentials from the skill's config or `.env` are included.\n- **Prompt Injection Protection**: All vSphere-sourced content (event messages, host logs) is truncated, stripped of control characters, and wrapped in boundary markers (`[VSPHERE_EVENT]`/`[VSPHERE_HOST_LOG]`) before output to prevent prompt injection when consumed by LLM agents.\n\n## Supported AI Platforms\n\n| Platform | Status | Config File |\n|----------|--------|-------------|\n| Claude Code | Native Skill | `skills/vmware-monitor/SKILL.md` |\n| Gemini CLI | Context file + MCP | `skills/vmware-monitor/SKILL.md` |\n| OpenAI Codex CLI | Skill + AGENTS.md | `skills/vmware-monitor/SKILL.md` |\n| Aider | Conventions | `skills/vmware-monitor/SKILL.md` |\n| Continue CLI | Rules | `skills/vmware-monitor/SKILL.md` |\n| Trae IDE | Rules | `skills/vmware-monitor/SKILL.md` |\n| Kimi Code CLI | Skill | `skills/vmware-monitor/SKILL.md` |\n| MCP Server | MCP Protocol | `vmware_monitor/mcp_server/` |\n| Python CLI | Standalone | N/A |\n\n### MCP Server — Local Agent Compatibility\n\nThe MCP server works with any MCP-compatible agent via stdio transport. All 32 tools are **read-only** toward vSphere. Config templates in `examples/mcp-configs/`:\n\n| Agent | Local Models | Config Template |\n|-------|:----------:|-----------------|\n| Goose (Block) | Ollama, LM Studio | `goose.json` |\n| LocalCowork (Liquid AI) | Fully offline | `localcowork.json` |\n| mcp-agent (LastMile AI) | Ollama, vLLM | `mcp-agent.yaml` |\n| VS Code Copilot | — | `vscode-copilot.json` |\n| Cursor | — | `cursor.json` |\n| Continue | Ollama | `continue.yaml` |\n| Claude Code | — | `claude-code.json` |\n\n```bash\n# Example: Aider + Ollama (fully local, no cloud API)\naider --conventions skills/vmware-monitor/SKILL.md --model ollama/qwen2.5-coder:32b\n```\n\n## First Interaction: Environment Selection\n\nWhen the user starts a conversation, **always ask first**:\n\n1. **Which environment** do they want to monitor? (vCenter Server or standalone ESXi host)\n2. **Which target** from their config? (e.g., `prod-vcenter`, `lab-esxi`)\n3. If no config exists yet, guide them through creating `~/.vmware-monitor/config.yaml`\n\nFile v1.16.0:skill-card.md\n\n## Description:\n\nvmware-monitor helps agents perform safe, read-only VMware vCenter and ESXi monitoring across inventory, alarms, events, performance, capacity, snapshots, and object investigations.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[zw008](https://clawhub.ai/user/zw008)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nInfrastructure operators, administrators, and support engineers use this skill to ask an agent for read-only VMware vCenter and ESXi health, inventory, alarm, event, performance, capacity, snapshot, and investigation summaries before taking action in other tools.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: The skill handles VMware infrastructure credentials and operational data.\n\nMitigation: Use a dedicated VMware account with the built-in Read-Only role and inject passwords from a secret manager instead of storing them in .env.\n\nRisk: Webhook alerts can disclose operational details such as hostnames, IP addresses, usernames, event text, and log snippets.\n\nMitigation: Enable Slack or Discord webhooks only for destinations the operator controls, and treat those destinations as receiving operational data.\n\nRisk: Additional VMware privileges for session listing or host log access can exceed the baseline read-only role.\n\nMitigation: Grant privileges such as session listing or host log access only when those specific reads are needed.\n\n## Reference(s):\n\n- [ClawHub skill page](https://clawhub.ai/zw008/skills/vmware-monitor)\n- [Project homepage](https://github.com/vmware-skills/VMware-Monitor)\n- [Read-only enforcement test](https://github.com/vmware-skills/VMware-Monitor/blob/main/tests/eval/regression/test_read_only_enforcement.py)\n- [Capabilities](references/capabilities.md)\n- [CLI Reference](references/cli-reference.md)\n- [Setup Guide](references/setup-guide.md)\n- [Agent Guardrails](references/agent-guardrails.md)\n- [Investigation Protocol](references/investigation-protocol.md)\n- [Cluster Health Summary Display Template](references/health-summary-template.md)\n\n## Skill Output:\n\n**Output Type(s):** [text, markdown, shell commands, configuration, guidance]\n\n**Output Format:** [Markdown and text responses with CLI commands, structured tool-result summaries, and optional local HTML snapshot guidance.]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [The skill is read-only toward vSphere; some outputs may summarize sensitive operational data such as hostnames, IP addresses, usernames, events, alarms, and log snippets.]\n\n## Skill Version(s):\n\n1.16.0 (source: server release evidence)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nFile v1.16.0:evals/evals.json\n\n{\n  \"skill_name\": \"vmware-monitor\",\n  \"evals\": [\n    {\n      \"id\": 1,\n      \"prompt\": \"Show me all VMs that are powered off and how much memory each has\",\n      \"expected_output\": \"Filtered VM list with power state and memory\",\n      \"files\": [],\n      \"expectations\": [\n        \"Uses list_virtual_machines with power_state filter\",\n        \"Results include memory_mb for each VM\",\n        \"Does NOT attempt any write operations\"\n      ]\n    },\n    {\n      \"id\": 2,\n      \"prompt\": \"Are there any critical alarms? If so, what should I do about them?\",\n      \"expected_output\": \"Alarm list with remediation suggestions\",\n      \"files\": [],\n      \"expectations\": [\n        \"Uses get_alarms to retrieve active alarms\",\n        \"Includes suggested_actions in the response\",\n        \"Routes user to vmware-aiops for remediation if needed\"\n      ]\n    },\n    {\n      \"id\": 3,\n      \"prompt\": \"Give me a health report of host esxi-02 including hardware status\",\n      \"expected_output\": \"Host details with CPU, memory, version, uptime\",\n      \"files\": [],\n      \"expectations\": [\n        \"Uses list_esxi_hosts or vm_info for host details\",\n        \"Presents information in a readable format\",\n        \"Does NOT attempt any modifications\"\n      ]\n    }\n  ]\n}\n\nArchive v1.15.1: 10 files, 44620 bytes\n\nFiles: evals/evals.json (1251b), references/agent-guardrails.md (6782b), references/capabilities.md (20921b), references/cli-reference.md (12083b), references/health-summary-template.md (9123b), references/investigation-protocol.md (6758b), references/setup-guide.md (14729b), skill-card.md (2992b), SKILL.md (25005b), _meta.json (134b)\n\nFile v1.15.1:SKILL.md\n\n---\nname: vmware-monitor\ndescription: >\n  Use this skill for safe, read-only queries of VMware infrastructure — no destructive operations exist in the codebase; a test enforces it.\n  Directly handles: a one-glance cross-cluster health summary, object-centered VM/host/datastore investigation drill-downs (correlating surrounding infrastructure + recent events), a cross-vCenter \"what needs attention now?\" rollup, list VMs/hosts/datastores/clusters, active alarms, recent events, VM details.\n  Always use vmware-monitor when the user asks to \"list VMs\", \"check vSphere alarms\", \"show host status\", \"is anything on fire\", \"what needs attention now\", \"what is happening around this VM/host/datastore\", \"investigate this VM\" — or needs read-only VMware info before making changes.\n  Do NOT use for any write operations — this skill is read-only and has no code path that creates, modifies, or deletes a vSphere resource.\n  For VM modifications use vmware-aiops, for networking use vmware-nsx, for metrics/capacity use vmware-aria. For load balancing/AVI/AKO use vmware-avi.\ninstaller:\n  kind: uv\n  package: vmware-monitor\nallowed-tools:\n  - Bash\nmetadata: {\"openclaw\":{\"requires\":{\"anyBins\":[\"vmware-monitor\",\"uvx\"]},\"optional\":{\"env\":[\"VMWARE_MONITOR_CONFIG\",\"VMWARE_TARGET_PASSWORD\",\"VMWARE_<TARGET>_USERNAME\",\"SLACK_WEBHOOK_URL\",\"DISCORD_WEBHOOK_URL\",\"VMWARE_AUDIT_APPROVED_BY\"],\"bins\":[\"vmware-policy\"]},\"homepage\":\"https://github.com/vmware-skills/VMware-Monitor\",\"emoji\":\"📊\",\"os\":[\"macos\",\"linux\"]}}\ncompatibility: >\n  vmware-policy auto-installed as Python dependency (provides @vmware_tool decorator and audit logging). MCP tool calls and remote CLI commands audited to ~/.vmware/audit.db; CLI queries also to ~/.vmware-monitor/audit.log.\n  Credentials: Each vCenter/ESXi target requires a per-target password env var in ~/.vmware-monitor/.env following the pattern VMWARE_<TARGET_NAME_UPPER>_PASSWORD (e.g., target \"vcenter-prod\" → VMWARE_VCENTER_PROD_PASSWORD). SLACK_WEBHOOK_URL and DISCORD_WEBHOOK_URL are optional — disabled by default, user-configured only, used solely by the opt-in daemon scanner. Daemon: the background scanner (vmware-monitor daemon start) is user-initiated only, never auto-started. Webhook payloads carry issue counts plus every critical issue and every alarm/event warning (host-log warnings and info rows are not sent): entity name and the sanitized, truncated alarm, vCenter event, or ESXi log text, or a connection error — which can include host names, IPs, and user names. No credentials from the skill's config are sent.\n---\n\n# VMware Monitor (Read-Only)\n\n> **Disclaimer**: This is a community-maintained open-source project and is **not affiliated with, endorsed by, or sponsored by VMware, Inc. or Broadcom Inc.** \"VMware\" and \"vSphere\" are trademarks of Broadcom. Source code is publicly auditable at [github.com/vmware-skills/VMware-Monitor](https://github.com/vmware-skills/VMware-Monitor) under the MIT license.\n\nRead-only VMware vCenter/ESXi monitoring — 32 MCP tools, zero destructive code.\n\n> **Read-only toward vSphere**: no code path changes vCenter/ESXi state — no power, create, delete, snapshot, or reconfigure call exists. On vCenter it opens only its login session and short-lived query handles it releases. Gate: [`tests/eval/regression/test_read_only_enforcement.py`](https://github.com/vmware-skills/VMware-Monitor/blob/main/tests/eval/regression/test_read_only_enforcement.py) (source repo, not this bundle) requires every vSphere method called to be on a reviewed allowlist, checked against pyVmomi's type metadata. It checks source, not runtime, and no CI runs it. Independent of this code: use a dedicated account with vCenter's Read-Only role.\n> **Local writes** (this machine only): `init` writes `~/.vmware-monitor/config.yaml` and `.env` (0600; plaintext passwords are rewritten as `b64:` on load); audit logs `~/.vmware/audit.db` (MCP and CLI) and `~/.vmware-monitor/audit.log` (CLI); `--html` snapshots in `~/vmware-health/`; after `daemon start` only, `scan.log`, `daemon.pid` and opt-in webhook posts.\n> **Companion skills**: [vmware-aiops](https://github.com/vmware-skills/VMware-AIops) (VM lifecycle), [vmware-storage](https://github.com/vmware-skills/VMware-Storage) (iSCSI/vSAN), [vmware-vks](https://github.com/vmware-skills/VMware-VKS) (Tanzu Kubernetes), [vmware-nsx](https://github.com/vmware-skills/VMware-NSX) (NSX networking), [vmware-nsx-security](https://github.com/vmware-skills/VMware-NSX-Security) (DFW/firewall), [vmware-aria](https://github.com/vmware-skills/VMware-Aria) (metrics/alerts/capacity), [vmware-avi](https://github.com/vmware-skills/VMware-AVI) (AVI/ALB/AKO), [vmware-harden](https://github.com/vmware-skills/VMware-Harden) (compliance baselines).\n> | [vmware-pilot](../vmware-pilot/SKILL.md) (workflow orchestration) | [vmware-policy](../vmware-policy/SKILL.md) (audit/policy)\n\n## What This Skill Does\n\n| Category | Capabilities |\n|----------|-------------|\n| **Cluster Triage** | One-glance `cluster_health_summary` — cross-cluster Problems/Capacity/Health rollup with an opinionated status; customizable view |\n| **Object Investigation** | \"What is happening around this VM / host / datastore?\" — one correlated drill-down bundle per object, plus `cross_vcenter_attention` — one ranked \"what needs attention now?\" list across every configured vCenter |\n| **Inventory** | List VMs, ESXi hosts, datastores, clusters, networks |\n| **Health** | Active alarms, recent events (filter by severity/time), hardware sensors, host services |\n| **Performance** | Real-time host & VM CPU/memory/disk/network utilisation (PerfManager) |\n| **Capacity** | Datastore thin-provisioning over-commit, resource-pool reservation/usage |\n| **Infra Health** | ESXi certificate expiry, license usage/expiry, NTP configuration health |\n| **Snapshots** | Inventory-wide snapshot aging & sprawl (flag old snapshots) |\n| **Activity** | In-flight tasks, active login sessions |\n| **VM Details** | CPU, memory, disks, NICs, snapshots, guest OS, IP |\n| **Scanning** | Scheduled alarm/log scanning with Slack/Discord webhooks |\n| **vSphere 9.1** | Host memory tiering (DRAM/NVMe uplift), vLCM cluster patch compliance & last-apply result, vCenter deployment size |\n\n## Quick Install\n\n```bash\nuv tool install vmware-monitor==1.15.1\nvmware-monitor doctor\n```\n\n## When to Use This Skill\n\n- List or search VMs, hosts, datastores, clusters\n- Check active alarms or recent events\n- Get detailed info about a specific VM\n- Set up scheduled monitoring with webhook alerts\n- Any read-only VMware query where safety is paramount\n\n### Alarm/Event Output: `suggested_actions` Field\n\n`get_alarms` and `get_events` results include a `suggested_actions` list. Each\nitem is a ready-to-use hint naming the correct companion skill and tool call\n(e.g. `\"vmware-aiops: acknowledge_vcenter_alarm(entity_name=..., alarm_name=...)\"`),\nso agents — especially smaller local models — can follow them directly without\nreasoning about skill routing. Example payload: `references/capabilities.md`.\n\n**Use companion skills for**:\n- Power on/off, deploy, clone, migrate --> `vmware-aiops`\n- iSCSI, vSAN, datastore management --> `vmware-storage`\n- Tanzu Kubernetes clusters --> `vmware-vks`\n- Load balancing, AVI/ALB, AKO, Ingress --> `vmware-avi`\n\n## Related Skills — Skill Routing\n\n| User Intent | Recommended Skill |\n|-------------|------------------|\n| Read-only vSphere monitoring | **vmware-monitor** ← this skill |\n| Storage: iSCSI, vSAN, datastores | **vmware-storage** |\n| VM lifecycle, deployment, guest ops | **vmware-aiops** |\n| Tanzu Kubernetes (vSphere 8.x+) | **vmware-vks** |\n| NSX networking: segments, gateways, NAT | **vmware-nsx** |\n| NSX security: DFW rules, security groups | **vmware-nsx-security** |\n| Aria Ops: metrics, alerts, capacity planning | **vmware-aria** |\n| Multi-step workflows with approval | **vmware-pilot** |\n| Compliance baselines (CIS / 等保 / PCI-DSS), drift detection, LLM remediation advisor | **vmware-harden** (`uv tool install vmware-harden`) |\n| Load balancer, AVI, ALB, AKO, Ingress | **vmware-avi** (`uv tool install vmware-avi`) |\n| Audit log query | **vmware-policy** (`vmware-audit` CLI) |\n\n## Common Workflows\n\n> **Diagnostic investigations**: Before running any \"why is X failing / down / abnormal\" workflow, follow [`references/investigation-protocol.md`](references/investigation-protocol.md). It enforces the four root-cause completeness criteria (falsifiability / sufficiency / necessity / mechanism) and the up-to-three-rounds deepening loop. Since vmware-monitor is read-only, it serves as the data source — actuation belongs to companion skills like vmware-aiops.\n\n### Cluster Health Check (\"is anything on fire?\" / \"what's wrong right now?\")\n\n**Judgment**: this is the 5-second triage glance, not an Aria replacement. One call rolls every cluster's hosts, VM power, live CPU/memory and alarms up, flattens the individual anomalies into a ranked **top-N focus list** (`top_issues`), and gives each cluster an opinionated `status`. On a big fleet, lead with the focus list — scanning per-cluster rows is too slow.\n\n1. One glance --> `vmware-monitor summary` (MCP: `cluster_health_summary`). Read `top_issues` first (worst first, each with a drill-down `next step`); the per-cluster table is context. `issues_total` shows how many anomalies existed before the top-N cap\n2. Tighten or widen the focus --> `--top 5` for the 5 most urgent, `--top 20` for more, `--top 0` to hide the list and just see the table\n3. Drill into what the list points at --> e.g. a `host_down` row → `inventory hosts`; an `alarm` row → `get_alarms`; a `capacity` row → `perf hosts` / `capacity datastores`; scope with `--cluster prod-a`\n4. Reshape the view on request --> the output ends with a friendly hint; the operator can say \"add datastore free space\", \"drop the DRS column\", \"only show clusters needing attention\", or \"save this as an HTML page\". Default layout, columns, and thresholds live in [`references/health-summary-template.md`](references/health-summary-template.md) and are meant to be edited\n5. Save an offline snapshot --> `vmware-monitor summary --html` writes a self-contained HTML file (no external assets, nothing uploaded) to `~/vmware-health/cluster-health-<vc>-<timestamp>.html`; `--html-path <file>` for an explicit path. The timestamped filename means a folder of them becomes a browsable point-in-time history. It is a snapshot, not a live page — re-run to refresh\n6. **On a very large fleet** --> add `--no-vms` to skip the VM rollup pass when you only need host/alarm/capacity signals\n7. **`totals.clusters` is 0** --> not empty: un-clustered hosts (standalone ESXi, or a cluster-less vCenter) form the `(standalone hosts)` row, still in `totals` and `top_issues`\n\n### Daily Health Check\n\n**Judgment**: alarms tell you what vCenter has decided is wrong, events tell you what happened. They diverge — an event burst with no alarms often signals a metric threshold miscalibration, not \"everything is fine.\" Read both.\n\n1. Check alarms --> `vmware-monitor health alarms --target prod-vcenter` — focus on Red severity AND alarms older than 1 hour (transient ones self-clear)\n2. Review recent events --> `vmware-monitor health events --hours 24 --severity warning` — look for repeated events from the same entity (a single event is noise; 50 events in an hour is a pattern)\n3. List hosts --> `vmware-monitor inventory hosts` — flag hosts disconnected, in maintenance mode unexpectedly, or memory > 90%\n4. **If connection fails** --> run `vmware-monitor doctor` to diagnose config/network issues\n\n### Object-Centered Investigation (\"what is happening around this VM / host / datastore?\")\n\n**Judgment**: this is the drill-down the operator wants after triage points at a problem — one call *correlates* the object with its surrounding infrastructure and recent history, so you explain the aggregated result in operational language instead of stitching five tools yourself. The tool aggregates; you never dump raw inventory into the conversation.\n\nOffer the levels progressively — do **not** ask for details the environment already fixes:\n1. **Start at the top** --> `cross_vcenter_attention` (CLI: `vmware-monitor attention`) for \"what needs attention now?\" across every vCenter. If only one vCenter is configured, skip straight to its `cluster_health_summary` — no need to ask which target\n2. **Offer to drill into an object** the top-issues list points at. Ask *which level* only when it is genuinely ambiguous:\n   - a VM --> `vm_investigation_bundle` (CLI: `vmware-monitor investigate vm <name>`) → VM state, recent events, snapshots, alarms & recent changes, the host it runs on, the cluster context, the datastores backing it, performance signals, and a correlated event timeline\n   - a host --> `host_investigation_bundle` (CLI: `investigate host <name>`) → host state, cluster context, the VMs it runs, mounted datastores, alarms, performance, correlated timeline\n   - a datastore --> `datastore_investigation_bundle` (CLI: `investigate datastore <name>`) → capacity/free, mounting hosts, VMs it backs, alarms, correlated timeline\n3. **Widen or narrow** on request --> `--hours 72` for a longer event window; the bundle ends with a hint listing what is adjustable\n4. **Make it tangible** --> add `--html` to any `investigate`/`attention` command for a self-contained offline snapshot (drill-down sections collapse/expand natively, no JS, nothing uploaded) written to `~/vmware-health/`\n5. **If the object name is unknown** --> the tool returns a *teaching* error naming exactly how to list the objects (`list_virtual_machines` / `list_esxi_hosts` / `list_all_datastores`); get the exact name and retry\n6. **If a vCenter is unreachable** (attention only) --> it is listed under `unreachable` with a reason and the rest still aggregate — surface the gap, don't fail the whole view\n\n### Performance Triage (\"the cluster feels slow\")\n**Judgment**: inventory shows *configured* capacity (cores, GB); it cannot tell you what is actually hot. Use the real-time perf tools, then narrow.\n1. Rank hosts --> `vmware-monitor perf hosts` — the busiest host floats to the top (sorted by CPU%)\n2. Rank VMs on the suspect --> `vmware-monitor perf vms --limit 25` — find the noisy neighbour\n3. Check for hidden storage pressure --> `vmware-monitor capacity datastores` — over-commit % > 100 means a thin datastore can fill mid-run even with \"free\" space showing\n4. Rule out snapshot drag --> `vmware-monitor snapshots aging --only-old` — old snapshots silently degrade I/O\n5. **If perf tools return empty** --> the host/VM may be disconnected or powered off (no real-time provider); confirm with `inventory hosts` / `inventory vms`\n\n### Scheduled-Outage Pre-flight (certs, licenses, time)\n1. Cert expiry --> `vmware-monitor infra certs --warn-days 60` — an expired ESXi cert drops host management\n2. License headroom --> `vmware-monitor infra licenses` — catch over-allocation before it disables features\n3. Time sync --> `vmware-monitor infra ntp` — `healthy: no` breaks SSO/Kerberos/log correlation (note: live offset is not exposed by the SOAP API, only config health)\n\n### Set Up Continuous Monitoring\n1. Configure webhook in `~/.vmware-monitor/config.yaml`\n2. Start daemon --> `vmware-monitor daemon start`\n3. Daemon scans every 15 min, sends alerts to Slack/Discord\n\n## Usage Mode\n\n| Scenario | Recommended | Why |\n|----------|:-----------:|-----|\n| Local/small models (Ollama, Qwen) | **CLI** | ~2K tokens vs ~8K for MCP |\n| Cloud models (Claude, GPT-4o) | Either | MCP gives structured JSON I/O |\n| Automated pipelines | **MCP** | Type-safe parameters, structured output |\n\n## MCP Tools (32 — all read-only)\n\n| Tool | Description |\n|------|------------|\n| `list_virtual_machines` | List VMs with filtering (power state, sort, limit, `folder_filter`); each VM includes `folder_path` |\n| `list_esxi_hosts` | ESXi hosts with CPU, memory, version, uptime |\n| `list_all_datastores` | Datastores with capacity, free space, type |\n| `list_all_clusters` | Clusters with host count, DRS/HA status |\n| `cluster_health_summary` | One-glance triage across all clusters — ranked `top_issues` focus list + per-cluster rollup with opinionated `status`. Params: `cluster_filter`, `include_vms`, `top_n`. Render per `references/health-summary-template.md` |\n| `vm_investigation_bundle` | \"What is happening around this VM?\" — correlated drill-down with a merged, newest-first **event timeline** (contents listed in the workflow above). Params: `vm_name`, `hours`. Aggregated in the tool; explain, don't dump raw |\n| `host_investigation_bundle` | Same correlated drill-down around an ESXi host. Params: `host_name`, `hours` |\n| `datastore_investigation_bundle` | Same correlated drill-down around a datastore. Params: `datastore_name`, `hours` |\n| `cross_vcenter_attention` | \"What needs attention now?\" across **every** configured vCenter — one globally-ranked `top_issues` list (each tagged with its `vcenter`) + per-target rollup; unreachable targets degrade gracefully. Params: `cluster_filter`, `top_n` |\n| `list_all_networks` | Networks with attached VM count and accessibility |\n| `get_alarms` | Active alarms: `suggested_actions`, acknowledger, `condition_now` (`cleared` = stale) |\n| `get_events` | Events by severity and time window (`start`/`end`) |\n| `get_host_sensors` | Hardware sensor status (temperature/voltage/fan) per host with green/yellow/red health |\n| `get_host_services` | Host service status (running state and startup policy), optionally filtered by host |\n| `vm_info` | Detailed VM info (CPU, memory, disks, NICs, snapshots) |\n| `vm_list_snapshots` | Snapshot list for one VM with nesting hierarchy (read-only) |\n| `host_performance` | **Real-time** host CPU/mem/disk/net utilisation (PerfManager); busiest first |\n| `vm_performance` | **Real-time** VM CPU/mem/balloon/swap/disk/net utilisation (top 25 by default); powered-on only |\n| `snapshot_aging` | Inventory-wide snapshot sweep with age + sprawl; flags snapshots older than N days |\n| `vm_backup_snapshot_history` | Backup windows for one VM from snapshot task history; a lower bound, not job duration |\n| `certificate_status` | Per-host ESXi management certificate expiry (days until expiry, expiring flag) |\n| `license_status` | Licenses and per-asset `assignments` |\n| `ntp_status` | Per-host NTP config health (servers + ntpd state); live offset not in SOAP API |\n| `datastore_capacity` | Datastore over-commit (provisioned vs capacity); thin-provisioning risk |\n| `resource_pool_usage` | Resource-pool CPU/memory reservation, limit, and current usage |\n| `active_tasks` | In-flight (and recently completed) vCenter tasks with progress/errors |\n| `active_sessions` | Who is logged in, and from which client |\n| `host_log_scan` | ESXi host log trouble lines, grouped by pattern (CLI: `scan logs`) |\n| `host_memory_tiering` | **vSphere 9.1** — per-host memory tiering (DRAM/NVMe tiers) + NVMe uplift ratio (pyVmomi `hardware.memoryTierInfo`, needs ESXi 8.0U3+). Params: `host_name`, `limit`. Returns the list envelope |\n| `cluster_patch_compliance` | **vSphere 9.1** — vLCM software (patch) compliance for one cluster over vSphere Automation REST. Param: `cluster` (MoID, e.g. `domain-c123`). `available:false` = vCenter answered 503 (likely mid-patch), not an error; `non_compliant_hosts` is `null` when the host-status field is unknown, never a false 0 |\n| `cluster_last_apply_result` | **vSphere 9.1** — result of the last vLCM remediation (apply) on one cluster (REST). Param: `cluster` (MoID). Reports outcome only; never runs a remediation |\n| `vcenter_deployment_size` | **vSphere 9.1** — vCenter appliance deployment size class (REST, NEW in 9.1). `available:false` on 503 |\n\n> **vSphere 9.1 field-parse honesty**: the 3 REST tools' *endpoints* are spec-verified, but their JSON field names have NOT yet been replayed against a live 9.1 vCenter — every field is read defensively and each result self-labels via its `note` (`endpoint verified; field parse best-effort pending live 9.1 vCenter`). `host_memory_tiering` requires vCenter/ESXi 8.0U3+; older targets raise a teaching error naming the missing property.\n\nNo tool modifies, creates, or deletes any vCenter/ESXi resource.\nPerformance/capacity readings are point-in-time samples — this skill retains no\nhistory, so it never reports a fabricated \"trend\" or runway date.\n\n### List result shape\n\nThe 21 row-listing tools above (including `host_memory_tiering`) return the family list envelope\n`{items, returned, limit, total, truncated, hint}`, not a bare array. Read\n`truncated` before summarising: `true` means more rows exist — never call\n`items` the whole picture; `false` means complete, so empty `items` means\n\"checked, found none\" — for `host_log_scan`, only if `logs_unavailable`\n(logs it could not read) is empty too. A `null` `total`\n(`get_events`, `host_log_scan`) is deliberate. Aggregate tools return\npurpose-built objects — see `references/capabilities.md`.\n\n## Read-Only by Design\n\nAll 32 tools are vSphere reads (local writes: see top). Running\nwith local or small models? See\n[`references/agent-guardrails.md`](references/agent-guardrails.md).\n\n## CLI Quick Reference\n\n```bash\nvmware-monitor summary [--top 10] [--cluster <substr>] [--html] [--target <t>]\nvmware-monitor inventory vms|hosts|datastores|clusters|networks [--target <t>]\nvmware-monitor health alarms|events|sensors|services [--target <t>]\nvmware-monitor perf hosts|vms [--target <t>]\nvmware-monitor capacity datastores|pools [--target <t>]\nvmware-monitor infra certs|licenses|ntp [--target <t>]\nvmware-monitor snapshots aging [--only-old] [--target <t>]\nvmware-monitor vm info <vm-name> [--target <t>]\nvmware-monitor memory tiering [--host <esxi>] [--target <t>]        # vSphere 9.1\nvmware-monitor patch compliance|last-apply <cluster-moid> [--target <t>]   # vSphere 9.1 (vLCM)\nvmware-monitor deployment-size [--target <t>]                      # vSphere 9.1\nvmware-monitor scan now | scan logs [--host <name>] | daemon start|stop|status | doctor [--skip-auth]\n```\n\n> Full CLI reference (all flags + activity/tasks/sessions): see `references/cli-reference.md`\n\n## Troubleshooting\n\n### Alarms returns empty but vCenter shows alarms\nThe `get_alarms` tool queries triggered alarms at the root folder level. Some alarms are entity-specific — try checking events instead: `get_events --hours 1 --severity info`.\n\n### \"Connection refused\" error\n1. Run `vmware-monitor doctor` to diagnose\n2. Verify target hostname/IP and port (443) in config.yaml\n3. For self-signed certs: set `verify_ssl: false`\n\n### Events returns too many results\nUse severity filter: `--severity warning` (default) filters out info-level events. Use `--hours 4` to narrow time range.\n\n### VM info shows \"guest_os: unknown\"\nVMware Tools not installed or not running in the guest. Install/start VMware Tools for guest OS detection, IP address, and guest family info.\n\n### Doctor passes but commands fail with timeout\nvCenter may be under heavy load. Try targeting a specific ESXi host directly instead of vCenter, or increase connection timeout in config.yaml.\n\n### Should I set `environment:` on a read-only skill?\nYou can — add `environment: production` (or `staging`, `lab`, your own label)\nto each target in `~/.vmware-monitor/config.yaml`. It's an optional label; this\nskill has zero write tools, so nothing it exposes is ever gated by it — reads\nare never gated. It matters for the write skills (`vmware-aiops`,\n`vmware-storage`, `vmware-nsx`) pointed at the same vCenter: an\nenvironment-scoped `deny` rule in `~/.vmware/rules.yaml` can match on the label\nto block their writes (e.g. freeze `production`). A target with no label is\nsimply not matched by such a rule. Config example: `references/setup-guide.md`.\n\n## Setup\n\n```bash\nuv tool install vmware-monitor==1.15.1\nvmware-monitor init      # guided: prompts for host/user/password, writes config + .env (chmod 600), then verifies\n```\n\n`init` stores the password grep-safe (obfuscated `b64:`, never plaintext) and\nlocks `.env` to 0600. Prefer it over hand-editing; manual steps:\n`references/setup-guide.md`.\n\n> Full setup guide, security details, and AI platform compatibility: see `references/setup-guide.md`\n\n## Audit & Safety\n\nMCP calls (`@vmware_tool`) and remote CLI commands (`@audited`) are audited via vmware-policy; CLI queries also append to `~/.vmware-monitor/audit.log`:\n- Every MCP call and remote CLI command logged to `~/.vmware/audit.db` (SQLite)\n- Policy rules enforced via `~/.vmware/rules.yaml` (deny rules, maintenance windows, risk levels)\n- Risk classification: each tool tagged as low/medium/high/critical\n- View recent operations: `vmware-audit log --last 20`\n- View denied operations: `vmware-audit log --status denied`\n\n## License\n\nMIT — [github.com/vmware-skills/VMware-Monitor](https://github.com/vmware-skills/VMware-Monitor)\n\nFile v1.15.1:_meta.json\n\n{\n  \"ownerId\": \"kn7b067awq2s97bn3d7p5qfhw5827pxc\",\n  \"slug\": \"vmware-monitor\",\n  \"version\": \"1.15.1\",\n  \"publishedAt\": 1789535923515\n}\n\nFile v1.15.1:references/agent-guardrails.md\n\n# Operating the VMware skills with a local / small model\n\nClaude-class models drive these skills without special instruction. Smaller and\nlocally-hosted models — Llama 3.3 70B, Qwen, Mistral, and similar, served\nthrough Goose, Ollama, or OpenShift AI — need explicit operating rules to call\ntools reliably.\n\nThis page exists because an operator wrote those rules by hand first. The\nguardrails below are adapted, with thanks, from the working configuration\n[@juanpf-ha](https://github.com/juanpf-ha) developed while running\nvmware-monitor and vmware-aria against a production vSphere estate with Llama\n3.3 70B FP8 on an on-prem H100\n([VMware-AIops#31](https://github.com/vmware-skills/VMware-AIops/issues/31)).\n\n> **Disclaimer**: This is a community-maintained open-source project and is\n> **not affiliated with, endorsed by, or sponsored by VMware, Inc. or Broadcom\n> Inc.** \"VMware\" and \"vSphere\" are trademarks of Broadcom.\n\n---\n\n## First: the rules you no longer need to write\n\nSeveral guardrails from the original configuration are now enforced by the\nskills themselves. Prompt instructions are advisory — a model can ignore them.\nThese are structural, so it cannot.\n\n| Guardrail you would otherwise prompt for | Now enforced by |\n|---|---|\n| \"Work read-only and never modify anything\" | **Zero write tools.** This skill has no vSphere create, modify, delete, or power operation in its registry at all — `list_tools()` only ever offers reads, so the model cannot call what does not exist. |\n| \"First resolve the affected resource_id through vmware-aria before querying vmware-monitor\" | **`investigate_alert`** does the whole alert → resource → confirmed-name sequence in one call. |\n| \"Do not confuse the alert ID with the affected resource ID\" | `investigate_alert` returns a `correlation` block with both UUIDs explicitly labelled. |\n| \"Only correlate Aria and vCenter data after the resource name and type have been confirmed\" | `investigate_alert` returns `correlation.confirmed`, and withholds its `next_step` handoff until name and kind are known. |\n| \"Use explicit limits for queries that may return large amounts of data\" | **The list envelope.** Every list tool returns `{items, returned, limit, total, truncated, hint}`, so the model reads truncation instead of guessing at it. Most tools are unlimited unless you pass `limit` (`vm_performance` defaults to 25); the envelope states which case you got. |\n| \"If a requested field was not returned by any tool, show it as not available\" | Tools return explicit `null` for unresolved fields rather than omitting the key. |\n\n---\n\n## The system prompt\n\nEverything below still benefits from being stated explicitly. Copy this into\nyour agent's instruction block.\n\n```text\n## Tool use\n\n- Always call an MCP tool before answering any question about the current\n  VMware environment. Never answer from memory or assumption.\n- Never describe a tool call, and never output a JSON example, instead of\n  executing the tool. If you intend to call a tool, call it.\n- If a tool fails, report the actual error text. Do not complete the answer\n  with assumptions about what the result would have been.\n- Use explicit limits on queries that may return large amounts of data. Do not\n  request unlimited results unless the user asks for them.\n\n## Skill routing\n\n- vmware-monitor: vCenter inventory, ESXi hosts, clusters, datastores, VMs,\n  snapshots, alarms, events, performance.\n- vmware-aria: Aria Operations health, resources, metrics, alerts, capacity.\n- vmware-aiops: VM lifecycle (power, clone, snapshot, migrate, delete).\n- vmware-nsx / vmware-nsx-security: networking and firewall.\n- vmware-storage: datastores, iSCSI, vSAN.\n\n## Data fidelity\n\n- Never invent infrastructure objects, metrics, alarms, events, or\n  relationships. If a tool did not return it, it does not exist for this answer.\n- Preserve the exact criticality, status, impact, and control-state values the\n  tools return. Do not translate, normalise, or prettify enum values.\n- If a requested field was not returned, show it as \"not available\". Do not\n  infer it from other fields.\n- Preserve the original order and the full set of fields when the user asks\n  for specific ones.\n- When a response is long, report every item it contains. If a result is\n  truncated, the tool says so explicitly — report the truncation rather than\n  describing the visible subset as the whole.\n\n## Analysis discipline\n\n- Separate observed data from interpretation. State which is which.\n- Do not claim a security, performance, storage, or capacity problem unless\n  the tool output contains explicit supporting evidence.\n- Avoid generic recommendations that are not directly supported by the results.\n\n## Correlating Aria alerts with vCenter\n\n- Use investigate_alert to go from an alert to its affected resource. It\n  resolves the resource and confirms its name and kind in one call.\n- Do not pass an alert UUID where a resource UUID is expected. investigate_alert\n  labels both in its correlation block.\n- Only query vCenter for a resource once correlation.confirmed is true. The\n  next_step block names the exact tool and argument to use.\n```\n\n---\n\n## Known failure modes on small models\n\nObserved with Llama 3.3 70B FP8 (Goose, on-prem H100), and useful as a\nchecklist when evaluating any local model against these skills:\n\n| Symptom | Mitigation |\n|---|---|\n| Describes a tool call, or emits a JSON example, instead of executing it | The \"never describe a tool call\" rule above. Also check your harness is not echoing tool schemas into context — models imitate the nearest format they see. |\n| Long tool responses: omits items, or reports \"no data returned\" when data was present | Ask for explicit limits so responses stay small. Check the envelope's `truncated` / `returned` / `total` fields rather than trusting the model's summary — every list tool states them, so a \"no data\" claim is checkable against `returned`. |\n| Adds generic recommendations unsupported by results | The \"analysis discipline\" rules. |\n| Drops requested fields or reorders results | State the required fields and ordering in the request itself, not only in the system prompt. |\n| Multi-tool workflows take 30–50s end to end | Prefer the aggregate tools — `investigate_alert`, `cluster_health_summary`, `vm_investigation_bundle` — which collapse a 3-4 call sequence into one round trip. |\n\n## Reporting results\n\nLocal-model compatibility is an explicit design constraint for this family, and\nthe evidence base is small. If you evaluate a model against these skills —\nQwen, Mistral, Granite, or anything else — a report of what worked and what did\nnot is genuinely useful:\n[github.com/vmware-skills/VMware-Monitor/issues](https://github.com/vmware-skills/VMware-Monitor/issues).\n\nFile v1.15.1:references/capabilities.md\n\n# Capabilities (Read-Only)\n\nDetailed feature tables for `vmware-monitor`.\n\n## List result envelope\n\nThe 20 list-returning MCP tools — `list_virtual_machines`, `list_esxi_hosts`,\n`list_all_datastores`,\n`list_all_clusters`, `list_all_networks`, `get_alarms`, `get_events`,\n`get_host_sensors`, `get_host_services`, `host_log_scan`, `active_tasks`,\n`active_sessions`, `datastore_capacity`, `resource_pool_usage`,\n`certificate_status`, `license_status`, `ntp_status`, `host_performance`,\n`vm_performance`, `vm_list_snapshots` — return the family list envelope rather\nthan a bare array:\n\n```json\n{\"items\": [...], \"returned\": 50, \"limit\": 50, \"total\": 213,\n \"truncated\": true, \"hint\": \"Showing 50 of 213. Raise limit or narrow the query...\"}\n```\n\n| Key | Meaning |\n|-----|---------|\n| `items` | The rows, in the tool's documented order |\n| `returned` | `len(items)` — check this before claiming \"no data\" |\n| `limit` | The limit that produced this page; `null` when unlimited |\n| `total` | Real collection size, or `null` when the backing API does not report one |\n| `truncated` | `true` = more rows exist behind this page; `false` = this is complete |\n| `hint` | What to do about truncation; `null` when complete |\n\nTwo tools report `total: null` on purpose: `get_events` (events are read newest\nfirst and the read stops at 5000 — `read_truncated: true` and `read_note` say\nwhen it did and how far back it got) and `host_log_scan` (only the last N lines\nper log are read). The investigation bundles add `timeline_note` when their\ntimeline is not every event in the window: how many of how many are shown, and\nwhich scopes' reads stopped at 5000. Everywhere else the total is a real count taken before the\nlimit was applied, which is what lets a full page be recognised as complete\ninstead of flagged as possibly-truncated.\n\n`host_log_scan` adds one field to the envelope: `logs_unavailable`, one row per\nhost/log it could **not** read, with the reason (for example, the account lacks\n`Global.Diagnostics`, which `BrowseDiagnosticLog` requires). Its `items` holds\nonly the lines that matched a trouble pattern in the logs it *did* read, so an\nempty `items` with `truncated: false` means \"checked, found none\" only when\n`logs_unavailable` is empty as well. With unread logs, the honest answer is\n\"nothing matched in the logs that could be read\" — name the ones that could not.\n\nBy default `host_log_scan` groups repeated lines by pattern, per log: each item has `count`,\n`hosts`, `first_seen` / `last_seen` (the log's own timestamps), `severity`, `log_level`, `pattern`\nand one `sample`; `lines_matched` is the ungrouped count and `group=false` returns one row per\nline. Severity follows the level ESXi wrote on the line (`Cr`/`Al`/`Em` critical, `Er`/`Wa`\nwarning, `In`/`No`/`Db` info); a line containing \"critical\", \"panic\" or \"corrupt\" is critical\nwhatever its level.\n\nThe envelope adds ~30 tokens to a response. It exists because a bare list gave\nsmaller models nothing to distinguish a complete answer from page one, and they\nsometimes resolved that ambiguity as \"no data was returned\"\n(VMware-AIops issue #31).\n\n`list_virtual_machines` adds one extra key to the envelope, `mode` (`\"full\"` or\n`\"compact\"`), and reuses `hint` for the compact-mode note; its `total` is the\ncount after `power_state` / `folder_filter` are applied and before `limit`.\n\nTools with purpose-built return objects — `vm_info`, `snapshot_aging`,\n`cluster_health_summary`, `cross_vcenter_attention`, and the three\n`*_investigation_bundle` tools — are unaffected.\n\n## Automation Level Reference\n\nEach operation is classified by autonomy level per the Enterprise Harness Engineering framework. **vmware-monitor is L1/L2 only by design** — no vSphere write operations exist in the codebase, gated by an allowlist test in the source repository.\n\n| Level | Meaning | Agent autonomy | Examples in this skill |\n|:-:|---|---|---|\n| **L1** | Read-only, raw data | Always auto-run | `list_virtual_machines`, `list_esxi_hosts`, `get_alarms`, `get_events`, `list_all_datastores`, `list_all_clusters`, `host_performance` |\n| **L2** | Read + analysis / recommendation | Always auto-run | `cluster_health_summary`, `cross_vcenter_attention`, `snapshot_aging`, the three `*_investigation_bundle` tools, scheduled scan reports, log pattern matching (error/fail/critical/panic/timeout), alarm correlation, daemon-driven webhook digests |\n| **L3** | Single write — user must approve | *N/A* | — *(use [vmware-aiops](https://github.com/vmware-skills/VMware-AIops) for write operations)* |\n| **L4** | Multi-step plan / apply workflow | *N/A* | — *(use [vmware-pilot](https://github.com/vmware-skills/VMware-Pilot) for orchestration)* |\n| **L5** | Auto-remediation from learned pattern | *N/A* | — *(remediation is out of scope by design)* |\n\n**Notes**:\n- No tool changes vCenter/ESXi state, so agents can call them without confirmation — gated by [`tests/eval/regression/test_read_only_enforcement.py`](https://github.com/vmware-skills/VMware-Monitor/blob/main/tests/eval/regression/test_read_only_enforcement.py) (source repository; a check on the code as written, run by the test suite — there is no CI). Results still carry sensitive inventory, event, log, and session data: scope the account as in `setup-guide.md` → Least Privilege.\n- Local files the skill writes (config, `.env`, audit logs, HTML snapshots, daemon state) are listed in `setup-guide.md` → What \"read-only\" covers.\n\n## 0. Cluster Health Summary (triage)\n\nCLI `summary`, MCP `cluster_health_summary`. One aggregated read for a fast\ncross-cluster \"is anything on fire?\" glance — the operator's first look, not an\nAria Operations replacement.\n\n| Aspect | Detail |\n|--------|--------|\n| Passes | 4 batched `RetrievePropertiesEx` calls (clusters, hosts, VMs, datastores) — never one per object (issue #31 class) |\n| Focus list | `top_issues`: individual anomalies (disconnected hosts, triggered alarms, capacity/HA, datastores thin-provisioned past `DATASTORE_OVERCOMMIT_WARN_PCT=100` — `scope: datastore`) flattened + ranked worst-first, capped at `top_n`; `issues_total` reports pre-cap count. Alarm names and definitions resolved in one batched call (no N+1) |\n| Alarm issues | Carry `condition_now` (`holds` / `cleared` / `unknown`, as `get_alarms`) and `acknowledged_days` (null when not acknowledged). Only `cleared` ranks after the other issues; `unknown` keeps its severity rank — not re-checked, possibly still live. `object` is `vCenter appliance <address>` for alarms vCenter raises about itself |\n| Next step | `drilldown` names MCP tools; the CLI table and HTML snapshot name CLI commands |\n| Rollup | Per cluster: hosts connected/total, VM power, live CPU/mem %, HA/DRS, alarm counts (cluster + host) |\n| Status | Opinionated `ok` / `warn` / `critical` + plain-language `attention` reasons; sorted worst-first |\n| Thresholds | `CPU_MEM_WARN_PCT=85`, `CPU_MEM_CRIT_PCT=95` (named constants in `ops/cluster_summary.py`); disconnected host or critical alarm forces `critical` |\n| Customizable | Columns/thresholds/layout in [`health-summary-template.md`](health-summary-template.md); response carries a `customization_hint` |\n\n### `cluster_health_summary` — input parameters\n\n| Parameter | Type | Default | Behavior |\n|-----------|------|---------|----------|\n| `target` | str (optional) | default target | Named vCenter/ESXi target |\n| `cluster_filter` | str (optional) | None (all) | Case-insensitive substring; suppresses standalone-hosts bucket |\n| `include_vms` | bool | True | Roll up VM power counts; False skips the VM pass (faster on huge fleets) |\n| `top_n` | int | 10 | Cap the `top_issues` focus list; `issues_total` keeps the pre-cap count; 0 hides the list |\n\n**`totals.clusters` counts real clusters only.** Hosts that belong to no cluster\n— a standalone ESXi target, or standalone hosts under a vCenter — are rolled into\na `(standalone hosts)` row, and alarms raised above any cluster or host get a\n`(vCenter-level)` row when there are any; neither row is counted as a cluster. Their hosts, VMs and alarms still count in the\nother `totals` fields and feed `top_issues`. So a vCenter with only standalone\nhosts reports `clusters: 0` next to a non-zero `hosts_total` and a populated\nstandalone row: read that row, not an alternative tool. `cluster_filter` hides\nthe standalone row, so a filtered `clusters: 0` means the filter matched nothing.\n\n**Typical response tokens**: ~120–400 (one compact row per cluster + totals);\nscales with cluster count, not VM count. This is the aggregation-in-the-tool\npattern — the model never sees raw inventory.\n\n### `cross_vcenter_attention` — overlapping targets\n\nCLI `attention`. Targets can overlap: an ESXi host configured as its own target\nmay also be managed by a configured vCenter. Identity is read, never taken from\nnames:\n\n| Field | Meaning |\n|-------|---------|\n| `totals.vcenters` / `esxi_targets` / `unidentified_targets` | Targets by endpoint type (`about.apiType`); a type that cannot be read is `unidentified`, never counted as a vCenter |\n| `totals.hosts_total` / `hosts_connected` | A host reached through a vCente\n\nArchive v1.15.0: 10 files, 44532 bytes\n\nFiles: evals/evals.json (1251b), references/agent-guardrails.md (6782b), references/capabilities.md (20921b), references/cli-reference.md (12039b), references/health-summary-template.md (9123b), references/investigation-protocol.md (6758b), references/setup-guide.md (14729b), skill-card.md (2828b), SKILL.md (25005b), _meta.json (134b)\n\nArchive v1.14.0: 10 files, 44180 bytes\n\nFiles: evals/evals.json (1251b), references/agent-guardrails.md (6782b), references/capabilities.md (20921b), references/cli-reference.md (12039b), references/health-summary-template.md (9123b), references/investigation-protocol.md (6758b), references/setup-guide.md (13627b), skill-card.md (2951b), SKILL.md (25005b), _meta.json (134b)\n\nArchive v1.13.1: 10 files, 42080 bytes\n\nFiles: evals/evals.json (1251b), references/agent-guardrails.md (6782b), references/capabilities.md (16556b), references/cli-reference.md (12039b), references/health-summary-template.md (7971b), references/investigation-protocol.md (6758b), references/setup-guide.md (13627b), skill-card.md (3086b), SKILL.md (24992b), _meta.json (134b)\n\nArchive v1.13.0: 10 files, 41818 bytes\n\nFiles: evals/evals.json (1251b), references/agent-guardrails.md (6782b), references/capabilities.md (16502b), references/cli-reference.md (12039b), references/health-summary-template.md (7971b), references/investigation-protocol.md (6758b), references/setup-guide.md (13583b), skill-card.md (2667b), SKILL.md (24902b), _meta.json (134b)\n\nArchive v1.12.0: 10 files, 41576 bytes\n\nFiles: evals/evals.json (1251b), references/agent-guardrails.md (6782b), references/capabilities.md (16269b), references/cli-reference.md (11697b), references/health-summary-template.md (7740b), references/investigation-protocol.md (6758b), references/setup-guide.md (13583b), skill-card.md (2925b), SKILL.md (24902b), _meta.json (134b)\n\nArchive v1.11.3: 10 files, 40270 bytes\n\nFiles: evals/evals.json (1251b), references/agent-guardrails.md (6782b), references/capabilities.md (14765b), references/cli-reference.md (10904b), references/health-summary-template.md (7502b), references/investigation-protocol.md (6758b), references/setup-guide.md (13583b), skill-card.md (2535b), SKILL.md (25022b), _meta.json (134b)\n\nArchive v1.11.2: 10 files, 37670 bytes\n\nFiles: evals/evals.json (1251b), references/agent-guardrails.md (6774b), references/capabilities.md (12202b), references/cli-reference.md (10798b), references/health-summary-template.md (7502b), references/investigation-protocol.md (6758b), references/setup-guide.md (10208b), skill-card.md (2735b), SKILL.md (24627b), _meta.json (134b)\n\nArchive v1.11.1: 10 files, 37758 bytes\n\nFiles: evals/evals.json (1251b), references/agent-guardrails.md (6774b), references/capabilities.md (12202b), references/cli-reference.md (10798b), references/health-summary-template.md (7502b), references/investigation-protocol.md (6758b), references/setup-guide.md (10208b), skill-card.md (2971b), SKILL.md (24627b), _meta.json (134b)","readmeExcerpt":"Skill: vmware-monitor Owner: zw008 Summary: Use this skill for safe, read-only queries of VMware infrastructure — no destructive operations exist in the codebase; a test enforces it. Directly handles: a one-glance cross-cluster health summary, object-centered VM/host/datastore investigation drill-downs (correlating surrounding infrastructure + recent events), a cross-vCenter \"what needs attention now?\" rollup, list V","codeSnippets":[],"executableExamples":[{"language":"bash","snippet":"uv tool install vmware-monitor==1.16.0\nvmware-monitor doctor"},{"language":"bash","snippet":"vmware-monitor summary [--top 10] [--cluster <substr>] [--html] [--target <t>]\nvmware-monitor inventory vms|hosts|datastores|clusters|networks [--target <t>]\nvmware-monitor health alarms|events|sensors|services [--target <t>]\nvmware-monitor perf hosts|vms [--target <t>]\nvmware-monitor capacity datastores|pools [--target <t>]\nvmware-monitor infra certs|licenses|ntp [--target <t>]\nvmware-monitor snapshots aging [--only-old] [--target <t>]\nvmware-monitor vm info <vm-name> [--target <t>]\nvmware-monitor memory tiering [--host <esxi>] [--target <t>]        # vSphere 9.1\nvmware-monitor patch compliance|last-apply <cluster-moid> [--target <t>]   # vSphere 9.1 (vLCM)\nvmware-monitor deployment-size [--target <t>]                      # vSphere 9.1\nvmware-monitor scan now | scan logs [--host <name>] | daemon start|stop|status | doctor [--skip-auth]"},{"language":"bash","snippet":"uv tool install vmware-monitor==1.16.0\nvmware-monitor init      # guided: prompts for host/user/password, writes config + .env (chmod 600), then verifies"},{"language":"text","snippet":"## Tool use\n\n- Always call an MCP tool before answering any question about the current\n  VMware environment. Never answer from memory or assumption.\n- Never describe a tool call, and never output a JSON example, instead of\n  executing the tool. If you intend to call a tool, call it.\n- If a tool fails, report the actual error text. Do not complete the answer\n  with assumptions about what the result would have been.\n- Use explicit limits on queries that may return large amounts of data. Do not\n  request unlimited results unless the user asks for them.\n\n## Skill routing\n\n- vmware-monitor: vCenter inventory, ESXi hosts, clusters, datastores, VMs,\n  snapshots, alarms, events, performance.\n- vmware-aria: Aria Operations health, resources, metrics, alerts, capacity.\n- vmware-aiops: VM lifecycle (power, clone, snapshot, migrate, delete).\n- vmware-nsx / vmware-nsx-security: networking and firewall.\n- vmware-storage: datastores, iSCSI, vSAN.\n\n## Data fidelity\n\n- Never invent infrastructure objects, metrics, alarms, events, or\n  relationships. If a tool did not return it, it does not exist for this answer.\n- Preserve the exact criticality, status, impact, and control-state values the\n  tools return. Do not translate, normalise, or prettify enum values.\n- If a requested field was not returned, show it as \"not available\". Do not\n  infer it from other fields.\n- Preserve the original order and the full set of fields when the user asks\n  for specific ones.\n- When a response is long, report every item it contains. If a result is\n  truncated, the tool says so explicitly — report the truncation rather than\n  describing the visible subset as the whole.\n\n## Analysis discipline\n\n- Separate observed data from interpretation. State which is which.\n- Do not claim a security, performance, storage, or capacity problem unless\n  the tool output contains explicit supporting evidence.\n- Avoid generic recommendations that are not directly supported by the results.\n\n## Correlating Aria alerts with "},{"language":"json","snippet":"{\"items\": [...], \"returned\": 50, \"limit\": 50, \"total\": 213,\n \"truncated\": true, \"hint\": \"Showing 50 of 213. Raise limit or narrow the query...\"}"},{"language":"json","snippet":"{\n  \"alarm_name\": \"VM CPU Ready High\",\n  \"entity_name\": \"prod-db-01\",\n  \"suggested_actions\": [\n    \"vmware-aiops: acknowledge_vcenter_alarm(entity_name='prod-db-01', alarm_name='VM CPU Ready High')\",\n    \"vmware-aiops: reset_vcenter_alarm(entity_name='prod-db-01', alarm_name='VM CPU Ready High')\"\n  ]\n}"}],"parameters":null,"dependencies":[],"permissions":[],"extractedFiles":[{"path":"SKILL.md","content":"---\nname: vmware-monitor\ndescription: >\n  Use this skill for safe, read-only queries of VMware infrastructure — no destructive operations exist in the codebase; a test enforces it.\n  Directly handles: a one-glance cross-cluster health summary, object-centered VM/host/datastore investigation drill-downs (correlating surrounding infrastructure + recent events), a cross-vCenter \"what needs attention now?\" rollup, list VMs/hosts/datastores/clusters, active alarms, recent events, VM details.\n  Always use vmware-monitor when the user asks to \"list VMs\", \"check vSphere alarms\", \"show host status\", \"is anything on fire\", \"what needs attention now\", \"what is happening around this VM/host/datastore\", \"investigate this VM\" — or needs read-only VMware info before making changes.\n  Do NOT use for any write operations — this skill is read-only and has no code path that creates, modifies, or deletes a vSphere resource.\n  For VM modifications use vmware-aiops, for networking use vmware-nsx, for metrics/capacity use vmware-aria. For load balancing/AVI/AKO use vmware-avi.\ninstaller:\n  kind: uv\n  package: vmware-monitor\nallowed-tools:\n  - Bash\nmetadata: {\"openclaw\":{\"requires\":{\"anyBins\":[\"vmware-monitor\",\"uvx\"]},\"optional\":{\"env\":[\"VMWARE_MONITOR_CONFIG\",\"VMWARE_TARGET_PASSWORD\",\"VMWARE_<TARGET>_USERNAME\",\"SLACK_WEBHOOK_URL\",\"DISCORD_WEBHOOK_URL\",\"VMWARE_AUDIT_APPROVED_BY\"],\"bins\":[\"vmware-policy\"]},\"homepage\":\"https://github.com/vmware-skills/VMware-Monitor\",\"emoji\":\"📊\",\"os\":[\"macos\",\"linux\"]}}\ncompatibility: >\n  vmware-policy auto-installed as Python dependency (provides @vmware_tool decorator and audit logging). MCP tool calls and remote CLI commands audited to ~/.vmware/audit.db; CLI queries also to ~/.vmware-monitor/audit.log.\n  Credentials: Each vCenter/ESXi target requires a per-target password env var in ~/.vmware-monitor/.env following the pattern VMWARE_<TARGET_NAME_UPPER>_PASSWORD (e.g., target \"vcenter-prod\" → VMWARE_VCENTER_PROD_PASSWORD). SLACK_WEBHOOK_URL and DISCORD_WEBHOOK_URL are optional — disabled by default, user-configured only, used solely by the opt-in daemon scanner. Daemon: the background scanner (vmware-monitor daemon start) is user-initiated only, never auto-started. Webhook payloads carry issue counts plus every critical issue and every alarm/event warning (host-log warnings and info rows are not sent): entity name and the sanitized, truncated alarm, vCenter event, or ESXi log text, or a connection error — which can include host names, IPs, and user names. No credentials from the skill's config are sent.\n---\n\n# VMware Monitor (Read-Only)\n\n> **Disclaimer**: This is a community-maintained open-source project and is **not affiliated with, endorsed by, or sponsored by VMware, Inc. or Broadcom Inc.** \"VMware\" and \"vSphere\" are trademarks of Broadcom. Source code is publicly auditable at [github.com/vmware-skills/VMware-Monitor](https://github.com/vmware-skills/VMware-Monitor) under the MIT license.\n\nRead-only VMware vCenter/ESXi monitoring"},{"path":"_meta.json","content":"{\n  \"ownerId\": \"kn7b067awq2s97bn3d7p5qfhw5827pxc\",\n  \"slug\": \"vmware-monitor\",\n  \"version\": \"1.16.0\",\n  \"publishedAt\": 1789915888294\n}"},{"path":"references/agent-guardrails.md","content":"# Operating the VMware skills with a local / small model\n\nClaude-class models drive these skills without special instruction. Smaller and\nlocally-hosted models — Llama 3.3 70B, Qwen, Mistral, and similar, served\nthrough Goose, Ollama, or OpenShift AI — need explicit operating rules to call\ntools reliably.\n\nThis page exists because an operator wrote those rules by hand first. The\nguardrails below are adapted, with thanks, from the working configuration\n[@juanpf-ha](https://github.com/juanpf-ha) developed while running\nvmware-monitor and vmware-aria against a production vSphere estate with Llama\n3.3 70B FP8 on an on-prem H100\n([VMware-AIops#31](https://github.com/vmware-skills/VMware-AIops/issues/31)).\n\n> **Disclaimer**: This is a community-maintained open-source project and is\n> **not affiliated with, endorsed by, or sponsored by VMware, Inc. or Broadcom\n> Inc.** \"VMware\" and \"vSphere\" are trademarks of Broadcom.\n\n---\n\n## First: the rules you no longer need to write\n\nSeveral guardrails from the original configuration are now enforced by the\nskills themselves. Prompt instructions are advisory — a model can ignore them.\nThese are structural, so it cannot.\n\n| Guardrail you would otherwise prompt for | Now enforced by |\n|---|---|\n| \"Work read-only and never modify anything\" | **Zero write tools.** This skill has no vSphere create, modify, delete, or power operation in its registry at all — `list_tools()` only ever offers reads, so the model cannot call what does not exist. |\n| \"First resolve the affected resource_id through vmware-aria before querying vmware-monitor\" | **`investigate_alert`** does the whole alert → resource → confirmed-name sequence in one call. |\n| \"Do not confuse the alert ID with the affected resource ID\" | `investigate_alert` returns a `correlation` block with both UUIDs explicitly labelled. |\n| \"Only correlate Aria and vCenter data after the resource name and type have been confirmed\" | `investigate_alert` returns `correlation.confirmed`, and withholds its `next_step` handoff until name and kind are known. |\n| \"Use explicit limits for queries that may return large amounts of data\" | **The list envelope.** Every list tool returns `{items, returned, limit, total, truncated, hint}`, so the model reads truncation instead of guessing at it. Most tools are unlimited unless you pass `limit` (`vm_performance` defaults to 25); the envelope states which case you got. |\n| \"If a requested field was not returned by any tool, show it as not available\" | Tools return explicit `null` for unresolved fields rather than omitting the key. |\n\n---\n\n## The system prompt\n\nEverything below still benefits from being stated explicitly. Copy this into\nyour agent's instruction block.\n\n```text\n## Tool use\n\n- Always call an MCP tool before answering any question about the current\n  VMware environment. Never answer from memory or assumption.\n- Never describe a tool call, and never output a JSON example, instead of\n  executing the tool. If you intend to call a t"},{"path":"references/capabilities.md","content":"# Capabilities (Read-Only)\n\nDetailed feature tables for `vmware-monitor`.\n\n## List result envelope\n\nThe 20 list-returning MCP tools — `list_virtual_machines`, `list_esxi_hosts`,\n`list_all_datastores`,\n`list_all_clusters`, `list_all_networks`, `get_alarms`, `get_events`,\n`get_host_sensors`, `get_host_services`, `host_log_scan`, `active_tasks`,\n`active_sessions`, `datastore_capacity`, `resource_pool_usage`,\n`certificate_status`, `license_status`, `ntp_status`, `host_performance`,\n`vm_performance`, `vm_list_snapshots` — return the family list envelope rather\nthan a bare array:\n\n```json\n{\"items\": [...], \"returned\": 50, \"limit\": 50, \"total\": 213,\n \"truncated\": true, \"hint\": \"Showing 50 of 213. Raise limit or narrow the query...\"}\n```\n\n| Key | Meaning |\n|-----|---------|\n| `items` | The rows, in the tool's documented order |\n| `returned` | `len(items)` — check this before claiming \"no data\" |\n| `limit` | The limit that produced this page; `null` when unlimited |\n| `total` | Real collection size, or `null` when the backing API does not report one |\n| `truncated` | `true` = more rows exist behind this page; `false` = this is complete |\n| `hint` | What to do about truncation; `null` when complete |\n\nTwo tools report `total: null` on purpose: `get_events` (events are read newest\nfirst and the read stops at 5000 — `read_truncated: true` and `read_note` say\nwhen it did and how far back it got) and `host_log_scan` (only the last N lines\nper log are read). The investigation bundles add `timeline_note` when their\ntimeline is not every event in the window: how many of how many are shown, and\nwhich scopes' reads stopped at 5000. Everywhere else the total is a real count taken before the\nlimit was applied, which is what lets a full page be recognised as complete\ninstead of flagged as possibly-truncated.\n\n`host_log_scan` adds one field to the envelope: `logs_unavailable`, one row per\nhost/log it could **not** read, with the reason (for example, the account lacks\n`Global.Diagnostics`, which `BrowseDiagnosticLog` requires). Its `items` holds\nonly the lines that matched a trouble pattern in the logs it *did* read, so an\nempty `items` with `truncated: false` means \"checked, found none\" only when\n`logs_unavailable` is empty as well. With unread logs, the honest answer is\n\"nothing matched in the logs that could be read\" — name the ones that could not.\n\nBy default `host_log_scan` groups repeated lines by pattern, per log: each item has `count`,\n`hosts`, `first_seen` / `last_seen` (the log's own timestamps), `severity`, `log_level`, `pattern`\nand one `sample`; `lines_matched` is the ungrouped count and `group=false` returns one row per\nline. Severity follows the level ESXi wrote on the line (`Cr`/`Al`/`Em` critical, `Er`/`Wa`\nwarning, `In`/`No`/`Db` info); a line containing \"critical\", \"panic\" or \"corrupt\" is critical\nwhatever its level.\n\nThe envelope adds ~30 tokens to a response. It exists because a bare list gave\nsmaller models nothing to distinguish a complete answer f"},{"path":"references/cli-reference.md","content":"# CLI Reference\n\nFull command reference for `vmware-monitor`.\n\n> **Output shape**: the CLI renders tables, so command output is unchanged. The\n> equivalent MCP tools wrap their rows in the list envelope\n> (`{items, returned, limit, total, truncated, hint}`) — see\n> [`capabilities.md`](capabilities.md#list-result-envelope).\n\n## Diagnostics\n\n```bash\nvmware-monitor doctor [--skip-auth]\n```\n\nChecks config file, connectivity, authentication, and pyVmomi version. Use `--skip-auth` to test config parsing without connecting.\n\n## MCP Config Generator\n\n```bash\nvmware-monitor mcp-config generate --agent <goose|cursor|claude-code|continue|vscode-copilot|localcowork|mcp-agent>\nvmware-monitor mcp-config list\n```\n\nGenerates MCP configuration files for supported agents. `list` shows all available agent templates.\n\n## Cluster Health Summary\n\n```bash\nvmware-monitor summary [--top <n>] [--cluster <substring>] [--no-vms] [--html] [--html-path <file>] [--target <name>]\n```\n\n- One aggregated read across all clusters. Leads with a ranked **top-N issues** focus list (the individual anomalies — disconnected hosts, triggered alarms, capacity/HA — worst first, each with a drill-down hint), then a per-cluster table: hosts connected/total, VM power rollup, live CPU/memory %, HA/DRS, alarm counts, and an opinionated `status` (ok/warn/critical) with `attention` reasons.\n- `--top`: size of the focus list (default 10; `--top 5` tighter, `--top 0` hides it and shows only the table). The header shows \"Top N (of TOTAL)\" when truncated.\n- `--cluster`: case-insensitive substring; show only matching clusters (also suppresses the standalone-hosts bucket).\n- `--no-vms`: skip the VM rollup pass (faster on very large fleets when only host/alarm/capacity signals are needed).\n- `--html`: write a self-contained, offline HTML snapshot (no external CSS/JS/fonts — nothing leaves the machine) to `~/vmware-health/cluster-health-<vc>-<YYYYMMDD-HHMMSS>.html`. The timestamped filename turns a folder of snapshots into a browsable point-in-time history. It is a snapshot, not a live page — re-run to refresh.\n- `--html-path <file>`: write the HTML snapshot to an explicit path instead of the auto-timestamped default (implies `--html`).\n- The rendered view is customizable — columns, thresholds, and layout live in [`health-summary-template.md`](health-summary-template.md). MCP tool: `cluster_health_summary`.\n\n## Object Investigation (drill-down)\n\n```bash\nvmware-monitor investigate vm <name>        [--hours <n>] [--html] [--html-path <file>] [--target <name>]\nvmware-monitor investigate host <name>      [--hours <n>] [--html] [--html-path <file>] [--target <name>]\nvmware-monitor investigate datastore <name> [--hours <n>] [--html] [--html-path <file>] [--target <name>]\n```\n\n- \"What is happening around this object?\" — one call *correlates* the object with its surrounding infrastructure and recent history, so the model explains an aggregated result instead of stitching several tools. Read-only; all cross-object"}],"languages":[],"docsSourceLabel":"CLAWHUB","editorialOverview":"Use this skill for safe, read-only queries of VMware infrastructure — no destructive operations exist in the codebase; a test enforces it. Directly handles: a one-glance cross-cluster health summary, object-centered VM/host/datastore investigation drill-downs (correlating surrounding infrastructure + recent events), a cross-vCenter \"what needs attention now?\" rollup, list VMs/hosts/datastores/clusters, active alarms, recent events, VM details. Always use vmware-monitor when the user asks to \"list VMs\", \"check vSphere alarms\", \"show host status\", \"is anything on fire\", \"what needs attention now\", \"what is happening around this VM/host/datastore\", \"investigate this VM\" — or needs read-only VMware info before making changes. Do NOT use for any write operations — this skill is read-only and has no code path that creates, modifies, or deletes a vSphere resource. For VM modifications use vmware-aiops, for networking use vmware-nsx, for metrics/capacity use vmware-aria. For load balancing/AVI/AKO use vmware-avi. Skill: vmware-monitor Owner: zw008 Summary: Use this skill for safe, read-only queries of VMware infrastructure — no destructive operations exist in the codebase; a test enforces it. Directly handles: a one-glance cross-cluster health summary, object-centered VM/host/datastore investigation drill-downs (correlating surrounding infrastructure + recent events), a cross-vCenter \"what needs attention now?\" rollup, list V","editorialQuality":{"score":100,"threshold":65,"status":"ready","wordCount":1693,"uniquenessScore":45,"reasons":[]}},"media":{"evidence":{"source":"no-media","verified":false,"confidence":"low","updatedAt":"2026-10-09T03:41:48.278Z","emptyReason":"No screenshots, media assets, or demo links are available."},"primaryImageUrl":null,"mediaAssetCount":0,"assets":[],"demoUrl":null},"ownerResources":{"evidence":{"source":"unclaimed","verified":false,"confidence":"low","updatedAt":"2026-10-09T03:41:48.278Z","emptyReason":"This page has not been claimed by the agent owner."},"hasCustomPage":false,"customPageUpdatedAt":null,"customLinks":[],"structuredLinks":{"docsUrl":null,"demoUrl":null,"supportUrl":null,"pricingUrl":null,"statusUrl":null},"customPage":null},"relatedAgents":{"evidence":{"source":"protocol-neighbors","verified":false,"confidence":"medium","updatedAt":"2026-10-10T05:25:33.687Z","emptyReason":null},"items":[{"id":"8ebccd8e-3863-4187-8355-c3f14e1f9edf","entityType":"agent","canonicalPath":"/agent/iofficeai-aionui","slug":"iofficeai-aionui","name":"AionUi","description":"Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!","url":"https://github.com/iOfficeAI/AionUi","homepage":"https://www.aionui.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-10-09T19:11:12.944Z","createdAt":"2026-02-25T03:38:16.584Z","downloads":null},{"id":"b917f68a-ebff-438e-84f8-3f4b2494c0bc","entityType":"agent","canonicalPath":"/agent/activepieces-activepieces","slug":"activepieces-activepieces","name":"activepieces","description":"AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents","url":"https://github.com/activepieces/activepieces","homepage":"https://www.activepieces.com","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-15T02:22:12.426Z","createdAt":"2026-02-25T03:38:12.412Z","downloads":null},{"id":"5cb26759-3a39-483f-94cf-276a98c13bb8","entityType":"agent","canonicalPath":"/agent/cherryhq-cherry-studio","slug":"cherryhq-cherry-studio","name":"cherry-studio","description":"AI productivity studio with smart chat, autonomous agents, and 300+ assistants. Unified access to frontier LLMs","url":"https://github.com/CherryHQ/cherry-studio","homepage":"https://cherry-ai.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-11T14:38:40.986Z","createdAt":"2026-02-25T03:38:19.379Z","downloads":null},{"id":"6f6582d0-5d76-4f0f-b81d-86520247950b","entityType":"agent","canonicalPath":"/agent/copilotkit-copilotkit","slug":"copilotkit-copilotkit","name":"CopilotKit","description":"The Frontend for Agents & Generative UI. React + Angular","url":"https://github.com/CopilotKit/CopilotKit","homepage":"https://docs.copilotkit.ai","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-03-25T09:50:57.846Z","createdAt":"2026-02-25T03:39:14.617Z","downloads":null}],"links":{"hub":"/agent","source":"/agent/source/clawhub","protocols":[{"label":"OpenClaw","href":"/agent/protocol/openclew"}]}}}