{"id":"038d5ae1-50f4-4f81-9473-ab8cce6687a9","entityType":"agent","slug":"clawhub-zw008-vmware-aiops","name":"vmware-aiops","canonicalUrl":"https://www.xpersona.co/agent/clawhub-zw008-vmware-aiops","canonicalPath":"/agent/clawhub-zw008-vmware-aiops","generatedAt":"2026-10-09T13:58:52.055Z","source":"CLAWHUB","claimStatus":"UNCLAIMED","verificationTier":"NONE","summary":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-09T04:26:49.944Z","emptyReason":null},"description":"Use this skill whenever the user needs to manage VMs in VMware/vSphere/ESXi — it's the entry point for all VM operations. Directly handles: power on/off, clone, snapshot, migrate, deploy from OVA or templates, run commands inside VMs, batch operations, cluster management, vCenter alarm acknowledgment, a one-glance cluster-health triage (\"is anything on fire?\"), and VM/host/datastore investigation drill-downs. Always use this skill for any \"power on\", \"clone\", \"deploy\", \"migrate\", \"batch\", \"guest exec\", \"alarm\", or VM lifecycle task, and for triage like \"is anything on fire\" / \"what needs attention now\" / \"investigate this VM\", when the context is explicitly VMware, vSphere, or ESXi. Do NOT use for general read-only queries (inventory/events/VM details — use vmware-monitor), NSX networking (use vmware-nsx), storage/iSCSI/vSAN (use vmware-storage), or Kubernetes cluster lifecycle (use vmware-vks). For multi-step workflows use vmware-pilot. For load balancing/AVI/AKO use vmware-avi.","descriptionLabel":"Source description","evidenceSummary":"Capability contract not published. No trust telemetry is available yet. 5K downloads reported by the source. Last updated 10/9/2026.","installCommand":"clawhub skill install s171xgnmqse0nqvgqvqnaq5f9183kyre:vmware-aiops","sourceUrl":"https://clawhub.ai/zw008/vmware-aiops","homepage":"https://clawhub.ai/zw008/skills/vmware-aiops","primaryLinks":[{"label":"View on ClawHub","url":"https://clawhub.ai/zw008/vmware-aiops","kind":"source"},{"label":"Homepage","url":"https://clawhub.ai/zw008/skills/vmware-aiops","kind":"homepage"}],"safetyScore":84,"overallRank":62,"popularityScore":59,"trustScore":null,"claimedByName":null,"isOwner":false,"seoDescription":"vmware-aiops technical dossier on Xpersona with agent coverage, OPENCLEW support, and live trust metadata."},"coverage":{"evidence":{"source":"public-profile","verified":false,"confidence":"medium","updatedAt":"2026-10-09T04:26:49.944Z","emptyReason":null},"protocols":[{"protocol":"OPENCLEW","label":"OpenClaw","status":"self-declared","notes":"Declared in the public agent profile."}],"capabilities":[],"verifiedCount":0,"selfDeclaredCount":1,"capabilityMatrix":{"rows":[{"key":"OPENCLEW","type":"protocol","support":"unknown","confidenceSource":"profile","notes":"Listed on profile"}],"flattenedTokens":"protocol:OPENCLEW|unknown|profile"}},"adoption":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-09T04:26:49.944Z","emptyReason":null},"stars":null,"forks":null,"downloads":5040,"packageName":null,"latestVersion":"1.12.0","tractionLabel":"5K downloads"},"release":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-09T04:26:49.943Z","emptyReason":null},"lastUpdatedAt":"2026-10-09T04:26:49.944Z","lastCrawledAt":"2026-10-09T04:26:49.943Z","lastIndexedAt":null,"nextCrawlAt":"2026-10-10T04:26:49.943Z","lastVerifiedAt":null,"highlights":[{"version":"1.12.0","createdAt":"2026-09-20T14:51:38.030Z","changelog":"MCP instructions now name the configured targets and how to choose one; a config that cannot be read says so instead of falling silent.","fileCount":9,"zipByteSize":42251},{"version":"1.11.0","createdAt":"2026-09-19T03:54:57.901Z","changelog":"Destructive MCP tools preview by default (confirm=False) and state their blast radius; confirm=True refuses on blockers or unreadable measurements. Requires vmware-policy>=1.17.0.","fileCount":9,"zipByteSize":42068},{"version":"1.10.0","createdAt":"2026-09-19T00:04:20.931Z","changelog":"vm_delete previews its blast radius and deletes only when the preview is acknowledged; duplicated VM names are refused; guest tools' risk level raised.","fileCount":9,"zipByteSize":39687},{"version":"1.9.7","createdAt":"2026-09-16T05:18:55.146Z","changelog":"Structured results name the target that answered; server instructions list the configured targets; bounded logout on stop","fileCount":9,"zipByteSize":39204},{"version":"1.9.5","createdAt":"2026-09-15T14:37:56.262Z","changelog":"Stopping the MCP server now logs out its vCenter session.","fileCount":9,"zipByteSize":39328},{"version":"1.9.4","createdAt":"2026-09-15T08:56:31.980Z","changelog":"Scanner daemon audited: TTL deletes pass guard() as vm_delete and write one row per outcome change; scan cycles and webhook sends audited; interrupts recorded; daemon start @audited.","fileCount":9,"zipByteSize":39469},{"version":"1.9.3","createdAt":"2026-09-15T05:59:52.104Z","changelog":"CLI reads are audited under their MCP tool names; every CLI command declares what it reaches (needs vmware-policy 1.15.0)","fileCount":9,"zipByteSize":39422},{"version":"1.9.2","createdAt":"2026-09-15T03:13:52.505Z","changelog":"health summary names over-committed datastores (requires vmware-monitor 1.13.0)","fileCount":9,"zipByteSize":39280}]},"execution":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No published capability contract is available yet."},"installCommand":"clawhub skill install s171xgnmqse0nqvgqvqnaq5f9183kyre:vmware-aiops","setupComplexity":"low","setupSteps":["Install using `clawhub skill install s171xgnmqse0nqvgqvqnaq5f9183kyre:vmware-aiops` in an isolated environment before connecting it to live workloads.","No published capability contract is available yet, so validate auth and request/response behavior manually.","Review the upstream CLAWHUB listing at https://clawhub.ai/zw008/vmware-aiops before using production credentials."],"contract":{"contractStatus":"missing","authModes":[],"requires":[],"forbidden":[],"supportsMcp":false,"supportsA2a":false,"supportsStreaming":false,"inputSchemaRef":null,"outputSchemaRef":null,"dataRegion":null,"contractUpdatedAt":null,"sourceUpdatedAt":null,"freshnessSeconds":null},"invocationGuide":{"preferredApi":{"snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-vmware-aiops/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-vmware-aiops/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-vmware-aiops/trust"},"curlExamples":["curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-zw008-vmware-aiops/snapshot\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-zw008-vmware-aiops/contract\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-zw008-vmware-aiops/trust\""],"jsonRequestTemplate":{"query":"summarize this repo","constraints":{"maxLatencyMs":2000,"protocolPreference":["OPENCLEW"]}},"jsonResponseTemplate":{"ok":true,"result":{"summary":"...","confidence":0.9},"meta":{"source":"CLAWHUB","generatedAt":"2026-10-09T13:58:52.051Z"}},"retryPolicy":{"maxAttempts":3,"backoffMs":[500,1500,3500],"retryableConditions":["HTTP_429","HTTP_503","NETWORK_TIMEOUT"]}},"endpoints":{"dossierUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-vmware-aiops/dossier","snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-vmware-aiops/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-vmware-aiops/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-vmware-aiops/trust"}},"reliability":{"evidence":{"source":"runtime-metrics","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No trust, reliability, or runtime telemetry is available."},"trust":{"status":"unavailable","handshakeStatus":"UNKNOWN","verificationFreshnessHours":null,"reputationScore":null,"p95LatencyMs":null,"successRate30d":null,"fallbackRate":null,"attempts30d":null,"trustUpdatedAt":null,"trustConfidence":"unknown","sourceUpdatedAt":null,"freshnessSeconds":null},"decisionGuardrails":{"doNotUseIf":["Contract metadata is missing or unavailable for deterministic execution."],"safeUseWhen":[],"riskFlags":["missing_or_unavailable_contract","trust_data_unavailable","schema_references_missing"],"operationalConfidence":"low"},"executionMetrics":{"observedLatencyMsP50":null,"observedLatencyMsP95":null,"estimatedCostUsd":null,"uptime30d":null,"rateLimitRpm":null,"rateLimitBurst":null,"lastVerifiedAt":null,"verificationSource":null},"runtimeMetrics":{"successRate":null,"avgLatencyMs":null,"avgCostUsd":null,"hallucinationRate":null,"retryRate":null,"disputeRate":null,"p50Latency":null,"p95Latency":null,"lastUpdated":null}},"benchmarks":{"evidence":{"source":"no-benchmark-data","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No benchmark suites or observed failure patterns are available."},"suites":[],"failurePatterns":[]},"artifacts":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-09T04:26:49.944Z","emptyReason":null},"readme":"Skill: vmware-aiops\n\nOwner: zw008\n\nSummary: Use this skill whenever the user needs to manage VMs in VMware/vSphere/ESXi — it's the entry point for all VM operations. Directly handles: power on/off, clone, snapshot, migrate, deploy from OVA or templates, run commands inside VMs, batch operations, cluster management, vCenter alarm acknowledgment, a one-glance cluster-health triage (\"is anything on fire?\"), and VM/host/datastore investigation drill-downs. Always use this skill for any \"power on\", \"clone\", \"deploy\", \"migrate\", \"batch\", \"guest exec\", \"alarm\", or VM lifecycle task, and for triage like \"is anything on fire\" / \"what needs attention now\" / \"investigate this VM\", when the context is explicitly VMware, vSphere, or ESXi. Do NOT use for general read-only queries (inventory/events/VM details — use vmware-monitor), NSX networking (use vmware-nsx), storage/iSCSI/vSAN (use vmware-storage), or Kubernetes cluster lifecycle (use vmware-vks). For multi-step workflows use vmware-pilot. For load balancing/AVI/AKO use vmware-avi.\n\nTags: aiops:1.12.0, esxi:1.12.0, latest:1.12.0, monitoring:1.12.0, operations:1.12.0, vcenter:1.12.0, vmware:1.12.0\n\nVersion history:\n\nv1.12.0 | 2026-09-20T14:51:38.030Z | user\n\nMCP instructions now name the configured targets and how to choose one; a config that cannot be read says so instead of falling silent.\n\nv1.11.0 | 2026-09-19T03:54:57.901Z | user\n\nDestructive MCP tools preview by default (confirm=False) and state their blast radius; confirm=True refuses on blockers or unreadable measurements. Requires vmware-policy>=1.17.0.\n\nv1.10.0 | 2026-09-19T00:04:20.931Z | user\n\nvm_delete previews its blast radius and deletes only when the preview is acknowledged; duplicated VM names are refused; guest tools' risk level raised.\n\nv1.9.7 | 2026-09-16T05:18:55.146Z | user\n\nStructured results name the target that answered; server instructions list the configured targets; bounded logout on stop\n\nv1.9.5 | 2026-09-15T14:37:56.262Z | user\n\nStopping the MCP server now logs out its vCenter session.\n\nv1.9.4 | 2026-09-15T08:56:31.980Z | user\n\nScanner daemon audited: TTL deletes pass guard() as vm_delete and write one row per outcome change; scan cycles and webhook sends audited; interrupts recorded; daemon start @audited.\n\nv1.9.3 | 2026-09-15T05:59:52.104Z | user\n\nCLI reads are audited under their MCP tool names; every CLI command declares what it reaches (needs vmware-policy 1.15.0)\n\nv1.9.2 | 2026-09-15T03:13:52.505Z | user\n\nhealth summary names over-committed datastores (requires vmware-monitor 1.13.0)\n\nv1.9.1 | 2026-09-14T15:13:51.763Z | user\n\nEvent sweep and daemon scan read newest-first and rank real pyVmomi event names; requires vmware-monitor 1.12.0.\n\nv1.9.0 | 2026-09-12T00:12:12.135Z | user\n\nBREAKING: guest operations require --user/username (they defaulted to root). CLI writes are authorised and audited under their MCP tool names, so one deny rule covers both surfaces. The daemon's host-log pass reads logs for the first time (0 before, 228 live after) and only critical host-log lines page the webhook.\n\nv1.8.22 | 2026-09-05T01:06:07.948Z | user\n\na dropped connection no longer keeps itself alive\n\nv1.8.21 | 2026-09-03T02:03:17.370Z | user\n\nsnapshot-delete gains the flag its docs already promised\n\nv1.8.20 | 2026-09-02T14:42:53.651Z | user\n\n`vm_cancel_ttl` is not destructive — it prevents a deletion\n\nv1.8.19 | 2026-08-31T07:23:46.417Z | user\n\none answer per .env, on every platform\n\nv1.8.18 | 2026-08-31T00:32:47.781Z | user\n\nfix: run the suite on a non-UTF-8 machine, and stop one skill answering for another\n\nv1.8.17 | 2026-08-30T15:52:53.276Z | user\n\nSee RELEASE_NOTES.md for v1.8.17\n\nv1.8.16 | 2026-08-30T15:17:46.966Z | user\n\nSecond-round fixes from the 2026-08-30 VCF 9.1 re-test; vmware-policy floor raised to 1.11.0 (the engine no longer fails open when rules.yaml cannot be read).\n\nv1.8.15 | 2026-08-30T09:34:03.190Z | user\n\nThree lying MCP annotations corrected; errors now marked isError on the wire. Parameter descriptions now reach the MCP JSON schema (0% -> 100% coverage); additionalProperties closed; vmware-policy floor raised to 1.10.0.\n\nv1.8.14 | 2026-08-30T07:41:48.912Z | user\n\nUnreachable hosts no longer vanish from list_host_vmks while the envelope claims completeness; three sibling write tools' bare AttributeError replaced with teaching errors; doctor authenticates every target; one config path across CLI/doctor/MCP.\n\nv1.8.13 | 2026-08-29T15:18:31.860Z | user\n\nThree list tools returned a shape outside the family envelope, so agents reading 'items' got nothing; offset paging claimed more data at the end of a list; the documented SSL config key was one the code never read.\n\nv1.8.12 | 2026-08-28T02:53:16.829Z | user\n\nFixes the server's self-reported version and the advertised tool count; adds a Claude Code plugin manifest.\n\nv1.8.11 | 2026-08-01T03:10:00.889Z | user\n\nMoved to vmware-skills GitHub org; MCP Registry namespace → io.github.vmware-skills. Links updated.\n\nv1.8.10 | 2026-07-25T07:22:48.572Z | user\n\nDRS VM-VM rule suite (list/create/delete/enable-disable) + CLI parity (PR #37), set_vmk_service host-service tagging (PR #36), UTF-8 I/O pin (PR #38). 55→60 tools. Community contributions by @wright-bench.\n\nv1.8.9 | 2026-07-23T01:22:45.554Z | user\n\nNetwork authoring: dvSwitch portgroups + host VMkernel adapters + DF-bit MTU-path ping (PR #35). 49 to 55 tools (17 read / 38 write). Writes preview/confirm gated; remove_host_vmk fail-closed.\n\nv1.8.8 | 2026-07-21T15:40:42.640Z | user\n\nCLI writes now route through the shared guard()+audit_call() core via @guarded, exactly like the MCP tools (HLD I-1/I-8). Requires vmware-policy>=1.8.8.\n\nv1.8.7 | 2026-07-21T11:38:38.501Z | user\n\nRemove read-only switch and approval tiers; read/write authz delegated to RBAC. Plus accumulated fixes since 1.8.5.\n\nv1.8.5 | 2026-07-20T13:03:02.826Z | user\n\nA failure that is returned is now audited as a failure, and certificate/URL detail no longer reaches the agent. Both fixes v1.8.4 announced were incomplete.\n\nv1.8.4 | 2026-07-20T08:24:48.887Z | user\n\nTeaching error messages, domain exceptions no longer redacted on the way to the agent, and tool descriptions that state when to use each tool and what to call next.\n\nv1.8.3 | 2026-07-20T03:40:02.284Z | user\n\nPer-target username can now come from an env var, resolved per access like the password; documented credential variables corrected against what each repo's code actually reads\n\nv1.8.2 | 2026-07-19T17:59:41.529Z | user\n\nMCP server moved into the package namespace — fixes two skills in one environment silently overwriting each other's server; agent-guardrails.md for local/small models now ships in every skill\n\nv1.8.1 | 2026-07-19T11:22:07.521Z | user\n\nRead-only mode now documented on every surface that teaches it (SKILL.md, setup-guide, capabilities) and reported by doctor\n\nv1.8.0 | 2026-07-19T09:39:47.726Z | user\n\nRead-only mode (36 write-effecting tools withheld), list-result envelope, declared environments; vm_list_plans no longer deletes plans; tool-count and capabilities corrections\n\nv1.7.7 | 2026-07-17T06:56:08.904Z | user\n\nSession-probe eviction fix (dead cached sessions were never evicted; None currentSession now treated as dead) + lockfile mcp 1.28.1 clearing three GHSA HIGH advisories.\n\nv1.7.6 | 2026-07-14T08:55:22.335Z | user\n\nObject investigation bundles + cross-vCenter attention re-exposed from the AIops entry point (delegates to vmware-monitor). MCP 45→49.\n\nv1.7.5 | 2026-07-13T07:16:51.955Z | user\n\ncluster_health_summary tool (44→45) + summary CLI delegating to vmware-monitor — family triage from the AIops entry point; new dep vmware-monitor>=1.7.5\n\nv1.7.4 | 2026-07-13T04:51:55.095Z | user\n\nFamily version alignment to 1.7.4 (substantive change this cycle is in vmware-monitor: host-check boundary read batching).\n\nv1.7.3 | 2026-07-03T00:48:27.532Z | user\n\nFamily version alignment (v1.7.3)\n\nv1.7.2 | 2026-07-02T14:30:48.850Z | user\n\nBatch alarm & health reads via PropertyCollector (issue #31 follow-up)\n\nv1.7.1 | 2026-07-02T10:52:31.304Z | user\n\nLarge-inventory scale fix (issue #31): PropertyCollector batching replaces per-object lazy SOAP round-trips.\n\nv1.7.0 | 2026-06-27T01:01:09.957Z | user\n\nguided init wizard, doctor fix, auth/TLS error teaching\n\nv1.6.1 | 2026-06-24T00:00:35.064Z | user\n\nv1.6.1 .env password b64 obfuscation\n\nv1.6.0 | 2026-06-22T09:16:58.899Z | user\n\nv1.6.0 trust architecture: undo tokens + governance harness (budget/audit/risk-tiers)\n\nv1.5.39 | 2026-06-22T00:42:00.889Z | user\n\nv1.5.39: AIops snapshot-delete async + honest timeout (token-burn fix), Storage browse timeout fix; others version-aligned\n\nv1.5.38 | 2026-06-12T06:58:56.934Z | user\n\nbacklog finish: MCP create/reconfigure, server split\n\nv1.5.37 | 2026-06-12T01:57:54.126Z | user\n\nbacklog: OVA deploy robustness, multi-DC, snapshot/TTL safety\n\nv1.5.36 | 2026-06-11T23:21:23.749Z | user\n\ncode-quality fix pack: teaching errors reach agents, TTL safety, CLI error translation\n\nv1.5.35 | 2026-06-10T00:44:30.017Z | user\n\nSecurity hardening: safe error handling, TLS/path/permission fixes\n\nv1.5.32 | 2026-06-08T02:47:08.548Z | user\n\nv1.5.32: invented pyVmomi methods fixed + alarm/sensor/migrate corrections\n\nv1.5.30 | 2026-06-07T13:22:34.534Z | user\n\nv1.5.30: Glama TDQS tool description quality rewrite\n\nv1.5.29 | 2026-05-29T02:19:27.956Z | user\n\nDoc sync for v1.5.26 tools: 41 MCP tools documented (8 read / 33 write); 7 VM lifecycle tools added to SKILL.md/capabilities.md/cli-reference.md\n\nArchive index:\n\nArchive v1.12.0: 9 files, 42251 bytes\n\nFiles: evals/evals.json (2055b), references/agent-guardrails.md (9940b), references/capabilities.md (34420b), references/cli-reference.md (4671b), references/investigation-protocol.md (6771b), references/setup-guide.md (11082b), skill-card.md (3258b), SKILL.md (26823b), _meta.json (132b)\n\nFile v1.12.0:SKILL.md\n\n---\nname: vmware-aiops\ndescription: >\n  Use this skill whenever the user needs to manage VMs in VMware/vSphere/ESXi — it's the entry point for all VM operations.\n  Directly handles: power on/off, clone, snapshot, migrate, deploy from OVA or templates, run commands inside VMs, batch operations, cluster management, vCenter alarm acknowledgment, a one-glance cluster-health triage (\"is anything on fire?\"), and VM/host/datastore investigation drill-downs.\n  Always use this skill for any \"power on\", \"clone\", \"deploy\", \"migrate\", \"batch\", \"guest exec\", \"alarm\", or VM lifecycle task, and for triage like \"is anything on fire\" / \"what needs attention now\" / \"investigate this VM\", when the context is explicitly VMware, vSphere, or ESXi.\n  Do NOT use for general read-only queries (inventory/events/VM details — use vmware-monitor), NSX networking (use vmware-nsx), storage/iSCSI/vSAN (use vmware-storage), or Kubernetes cluster lifecycle (use vmware-vks).\n  For multi-step workflows use vmware-pilot. For load balancing/AVI/AKO use vmware-avi.\ninstaller:\n  kind: uv\n  package: vmware-aiops\nargument-hint: \"[vm-name or describe your task]\"\nallowed-tools:\n  - Bash\nmetadata: {\"openclaw\":{\"requires\":{\"anyBins\":[\"vmware-aiops\",\"uvx\"]},\"optional\":{\"env\":[\"VMWARE_AIOPS_CONFIG\",\"VMWARE_TARGET_PASSWORD\",\"VMWARE_<TARGET>_USERNAME\",\"SLACK_WEBHOOK_URL\",\"DISCORD_WEBHOOK_URL\",\"VMWARE_AUDIT_APPROVED_BY\"],\"bins\":[\"vmware-policy\"]},\"homepage\":\"https://github.com/vmware-skills/VMware-AIops\",\"emoji\":\"🖥️\",\"os\":[\"macos\",\"linux\"]}}\ncompatibility: >\n  vmware-policy auto-installed as Python dependency (provides @vmware_tool decorator and audit logging). All write operations audited to ~/.vmware/audit.db.\n  Credentials: Each vCenter/ESXi target requires a per-target password env var in ~/.vmware-aiops/.env following the pattern VMWARE_<TARGET_NAME_UPPER>_PASSWORD. Passwords are never logged or echoed.\n  Destructive operations: All write tools require explicit parameters and pass through the @vmware_tool decorator (policy check + audit + sanitize). Every MCP write tool annotated destructive (22 of 43: power-off, delete, migrate, snapshot revert/delete, guest exec/upload/provision, cluster delete/remove-host, TTL, Clean Slate, plan apply/rollback, host-network and DRS) takes one confirm argument whose default returns a no-write blast-radius preview (HLD section 7); confirm=True is refused, and audited as a failure, on a blocker or an unreadable measurement, and vm_delete also requires the preview's acknowledgement echoed back and still matching; the other 21 write tools (create, clone, deploy, power-on, reconfigure) act on the first call. The enforcement boundary is the RBAC of the vCenter/ESXi account the server connects with, so run it under a dedicated least-privilege service account (a read-only role makes it read-only). Optional deny rules in ~/.vmware/rules.yaml are checked before every MCP call and remote CLI command; the shipped baseline denies nothing. CLI destructive commands additionally require double confirmation and most CLI writes support --dry-run; neither applies to MCP calls.\n  Guest operations: vm_name and command are required; no implicit or background execution. The command is unbounded and runs with the guest credentials supplied — the username is required on MCP and CLI alike (no root default) and over MCP the password is a tool argument the agent sees (redacted from the audit row). The guest account is a second authorization boundary that a read-only vCenter role does not limit; pass a least-privilege guest account. vm_guest_upload reads any local file the server process can read.\n  Webhooks: Disabled by default. When enabled, the daemon posts to user-configured URLs only: issue counts plus every critical issue and every alarm/event warning (host-log warnings and info rows are not sent), each with its entity name and the sanitized alarm, event, or ESXi log text, or a connection error — which can include host names, IPs, and user names. No credentials from the skill's config are sent. Reading host logs needs the Global.Diagnostics privilege; an unreadable log is recorded, not skipped.\n  TLS verification is on by default (verify_ssl: true); set verify_ssl: false only for self-signed certs in isolated lab environments.\n  Transitive dependencies: Only vmware-policy (audit/policy). No post-install scripts or background services.\n---\n\n# VMware AIops\n\n> **Disclaimer**: This is a community-maintained open-source project and is **not affiliated with, endorsed by, or sponsored by VMware, Inc. or Broadcom Inc.** \"VMware\" and \"vSphere\" are trademarks of Broadcom. Source code is publicly auditable at [github.com/vmware-skills/VMware-AIops](https://github.com/vmware-skills/VMware-AIops) under the MIT license.\n\nVMware family entry point — AI-powered VM lifecycle, deployment, and alarm management — 60 MCP tools.\n\n> **Start here**: install vmware-aiops first, then add modules as needed.\n> Run `vmware-aiops hub status` to see which family members are installed.\n> **Family**: [vmware-monitor](https://github.com/vmware-skills/VMware-Monitor) (inventory/health), [vmware-storage](https://github.com/vmware-skills/VMware-Storage) (iSCSI/vSAN), [vmware-vks](https://github.com/vmware-skills/VMware-VKS) (Tanzu Kubernetes), [vmware-nsx](https://github.com/vmware-skills/VMware-NSX) (NSX networking), [vmware-nsx-security](https://github.com/vmware-skills/VMware-NSX-Security) (DFW/firewall), [vmware-aria](https://github.com/vmware-skills/VMware-Aria) (metrics/alerts/capacity), [vmware-avi](https://github.com/vmware-skills/VMware-AVI) (AVI/ALB/AKO), [vmware-harden](https://github.com/vmware-skills/VMware-Harden) (compliance baselines).\n> | [vmware-pilot](../vmware-pilot/SKILL.md) (workflow orchestration) | [vmware-policy](../vmware-policy/SKILL.md) (audit/policy)\n\n## What This Skill Does\n\n| Category | Tools | Count |\n|----------|-------|:-----:|\n| **VM Lifecycle** | power on/off, create, reconfigure, clone, migrate, delete, snapshot CRUD, TTL auto-delete, clean slate | 16 |\n| **Deployment** | OVA, template, linked clone, batch clone/deploy | 8 |\n| **Guest Ops** | exec commands, upload/download files, provision | 5 |\n| **Plan/Apply** | multi-step planning with rollback | 4 |\n| **Cluster** | create, delete, HA/DRS config, add/remove hosts, DRS VM-VM rules (list/create/delete/enable-disable) | 10 |\n| **Datastore** | browse files, scan for images | 2 |\n| **Network** | dvSwitch portgroup list/create, host VMkernel list/add/remove/tag-service, DF-bit MTU-path ping | 7 |\n| **Alarm Management** | list alarms, acknowledge, reset | 3 |\n| **Triage & Investigation** (read-only, delegates to vmware-monitor) | one-glance cluster health summary, object-centered VM/host/datastore drill-down bundles, cross-vCenter \"what needs attention now?\" | 5 |\n\n## Audit & Safety\n\nRead before connecting an agent. Per-tool inventory: `references/capabilities.md`.\n\n- **MCP gates default to a no-write preview.** 22 write tools — every destructive one — return their blast radius unless `confirm=True`, which is refused on a blocker or unreadable measurement; `vm_delete` also needs `acknowledge_blast_radius` echoing the preview. The other 21 (create, clone, deploy, power-on, reconfigure) act on the first call.\n- **The enforcement boundary is vCenter/ESXi RBAC**: an agent can do whatever the configured account can. Use a dedicated, least-privilege service account scoped to what the agent may change (a read-only role makes the skill read-only). Store its password in `~/.vmware-aiops/.env` (0600) or a secret manager (`VMWARE_<TARGET>_PASSWORD`).\n- **CLI only**: destructive commands require double confirmation; most CLI writes take `--dry-run`. Neither applies to MCP.\n- **Policy**: deny rules and a maintenance window in `~/.vmware/rules.yaml` are checked before every MCP and remote CLI call (e.g. deny writes to `environment: production` targets). The shipped baseline denies nothing. An in-process guardrail, not a substitute for RBAC.\n- **Audit**: every MCP call is recorded in `~/.vmware/audit.db`, credentials redacted (`vmware-audit log --last 20`). Best-effort: a failed audit write warns, never blocks.\n- **Guest ops** run any command or file write the guest account allows — a read-only vCenter role does not limit this. `username` is required (no default account) and over MCP the password is a tool argument the agent sees; pass a minimal guest account. `vm_guest_upload` reads any local file the server can read.\n\n## Quick Install\n\n```bash\nuv tool install vmware-aiops==1.12.0\nvmware-aiops doctor\nvmware-aiops hub status   # see which family members are installed\n```\n\n## VMware Family — Install What You Need\n\nvmware-aiops is the entry point. Add modules for additional capabilities:\n\n| Module | Install | Adds |\n|--------|---------|------|\n| **vmware-monitor** | `uv tool install vmware-monitor` | Read-only inventory, alarms, events |\n| **vmware-storage** | `uv tool install vmware-storage` | iSCSI, vSAN, datastore management |\n| **vmware-vks** | `uv tool install vmware-vks` | Tanzu Kubernetes (vSphere 8.x+) |\n| **vmware-nsx** | `uv tool install vmware-nsx-mgmt` | NSX networking: segments, gateways, NAT |\n| **vmware-nsx-security** | `uv tool install vmware-nsx-security` | DFW microsegmentation, security groups |\n| **vmware-aria** | `uv tool install vmware-aria` | Aria Ops metrics, alerts, capacity |\n| **vmware-avi** | `uv tool install vmware-avi` | AVI load balancer, ALB, AKO, Ingress |\n\n> Each module stays independent — small tool count keeps local models (Ollama, Qwen) accurate.\n\n## When to Use This Skill\n\n- Power on/off, create, delete, snapshot, clone, or migrate VMs\n- Deploy VMs from OVA, templates, linked clones, or batch specs\n- Run commands or transfer files inside a VM (Guest Operations)\n- Create/configure clusters (HA/DRS)\n- Browse datastores for deployable images\n- Plan and execute multi-step operations with rollback\n- List, acknowledge, and clear vCenter triggered alarms (clear matches by entity type + status — see MCP Tools section)\n\n**Use companion skills for**:\n- Inventory, health, alarms, VM info → `vmware-monitor`\n- iSCSI, vSAN, datastore management → `vmware-storage`\n- Tanzu Kubernetes (Supervisor, Namespace, TKC) → `vmware-vks`\n- Load balancing, AVI/ALB, AKO, Ingress → `vmware-avi`\n\n## Related Skills — Skill Routing\n\n| User Intent | Recommended Skill |\n|-------------|------------------|\n| Read-only monitoring, zero risk | **vmware-monitor** (`uv tool install vmware-monitor`) |\n| Storage: iSCSI, vSAN, datastores | **vmware-storage** (`uv tool install vmware-storage`) |\n| VM lifecycle, deployment, guest ops | **vmware-aiops** ← this skill |\n| Tanzu Kubernetes (vSphere 8.x+) | **vmware-vks** (`uv tool install vmware-vks`) |\n| NSX networking: segments, gateways, NAT | **vmware-nsx** (`uv tool install vmware-nsx-mgmt`) |\n| NSX security: DFW rules, security groups | **vmware-nsx-security** (`uv tool install vmware-nsx-security`) |\n| Aria Ops: metrics, alerts, capacity | **vmware-aria** (`uv tool install vmware-aria`) |\n| Multi-step workflows with approval | **vmware-pilot** |\n| Compliance baselines (CIS / 等保 / PCI-DSS), drift detection, LLM remediation advisor | **vmware-harden** (`uv tool install vmware-harden`) |\n| Load balancer, AVI, ALB, AKO, Ingress | **vmware-avi** (`uv tool install vmware-avi`) |\n| Audit log query | **vmware-policy** (`vmware-audit` CLI) |\n\n## Common Workflows\n\n> **Diagnostic investigations**: Before remediating any \"why is X slow / failing / down\" issue, follow [`references/investigation-protocol.md`](references/investigation-protocol.md). It enforces the four root-cause completeness criteria (falsifiability / sufficiency / necessity / mechanism) and the up-to-three-rounds deepening loop. Only invoke L3+ write tools after the four criteria are satisfied AND the user has approved a remediation plan.\n\n### Cluster Health Triage (\"what's wrong right now?\")\n\nStart here when the ask is \"is anything on fire?\" before diving into a specific VM. This is a read-only rollup delegated to vmware-monitor, exposed here so triage-then-act stays in one conversation.\n\n1. One glance --> `cluster_health_summary` (MCP) or `vmware-aiops summary` (CLI). Read `top_issues` first — ranked anomalies (disconnected hosts, red/yellow alarms, capacity pressure), each with a drill-down hint; the per-cluster table is context\n2. Act on what it surfaces --> a `host_down` row → investigate the host; an alarm → `acknowledge_vcenter_alarm` / `reset_vcenter_alarm`; a hot VM implicated → `vm_migrate` or `vm_reconfigure` (after the investigation protocol)\n3. Save/share a snapshot --> `vmware-aiops summary --html` writes an offline, timestamped HTML file (identical to `vmware-monitor summary --html` — same shared renderer)\n4. **If vmware-monitor is not installed** --> this command/tool is unavailable (AIops delegates to it); install `vmware-monitor`, or use the deeper per-object read tools in that skill\n\n### Object-Centered Investigation → Act (drill-down before you change anything)\n\n**Judgment**: after triage points at a problem object, drill in with one correlated read *before* actuating — the bundle aggregates the object with its surrounding infrastructure and recent history so you (and the operator) see the full picture, not a guess. AIops is the conversational entry point, so triage → investigate → act stays in one conversation.\n\n1. Estate-wide --> `cross_vcenter_attention` (\"what needs attention now?\" across every vCenter). One vCenter configured → skip to `cluster_health_summary`\n2. Drill into the flagged object (offer the level; skip the question when unambiguous):\n   - a VM --> `vm_investigation_bundle` → state, host it runs on, cluster context, backing datastores, snapshots, alarms & recent changes, performance signals, correlated event timeline\n   - a host --> `host_investigation_bundle`; a datastore --> `datastore_investigation_bundle`\n3. **Then** act on the evidence --> e.g. bundle shows a wedged VM on a hot host → `vm_migrate`; a full datastore → `vm_delete_snapshot` on the sprawl the bundle surfaced. Follow the investigation protocol before any destructive action\n4. Widen the window with `hours=72`; render an offline snapshot with `--html` (drill-down sections collapse natively, nothing uploaded)\n5. **If the object name is unknown** --> the bundle returns a teaching error naming how to list objects; get the exact name and retry. **If vmware-monitor is not installed** --> these delegated tools are unavailable\n\n### Deploy a Lab Environment\n\n**Pre-flight (judgment, not blind sequence)**:\n- Free space: target datastore must have ≥ OVA size × 2 (delta files + thin-provision overhead). If multiple datastores qualify, prefer one with lowest current IOPS pressure (cross-check `vmware-aria` if available).\n- Name hygiene: prefix with date or owner (`lab-2026-04-30-alice`) so the TTL cleanup audit trail is meaningful.\n- TTL: always set. 480 min for a single test session, 7200 min for a week-long sandbox. **Never deploy a \"lab\" VM without a TTL** — that is how datastores fill up at 3 AM.\n- Snapshot timing: take the baseline **after** provisioning succeeds, not before — a pre-provision snapshot is just an empty checkpoint.\n\n**Steps**:\n1. `vmware-aiops datastore browse <ds> --pattern \"*.ova\"` → confirm image present and size\n2. `vmware-aiops deploy ova <path> --name <date>-<owner>-<purpose> --datastore <ds>`\n3. `vmware-aiops vm guest-exec <name> --cmd /usr/bin/python3 --args \"setup.py\" --user admin` → if exit ≠ 0, **stop**, do not snapshot a half-provisioned VM\n4. `vmware-aiops vm snapshot-create <name> --name baseline` (only if multi-iteration testing; skip for one-shot)\n5. `vmware-aiops vm set-ttl <name> --minutes 480`\n\n### Batch Clone for Testing\n\n**Pre-flight**:\n- Source VM state: powered-off is safest. If powered-on, VMware Tools must be running and quiesce-capable, else clones may have inconsistent disk state.\n- Capacity math: `free_space ≥ source.size × count × 1.2` (full clone) or `≥ count × 2 GB` (linked clone, delta-only).\n- Decision rule: **count > 10 → use linked clones** (`deploy linked-clone`); seconds vs minutes per clone, ~100× less storage. Tradeoff: linked clones depend on source snapshot — deleting the snapshot breaks all children.\n- Network exhaustion: each clone gets a unique MAC from the vSphere pool; if you batch > 200, verify pool capacity in advance.\n- TTL: every clone must have one. Use the plan's metadata to track ownership.\n\n**Steps**:\n1. `vm_create_plan` with clone + reconfigure + set-ttl steps grouped per VM (atomic per clone)\n2. Review the plan with the user — surface count, datastore, irreversible warnings\n3. `vm_apply_plan` — stops on first failure (intentional, do not auto-resume)\n4. On failure: `vm_rollback_plan` → reverses completed clones; manually verify rollback before retrying\n\n### Migrate VM to Another Host\n\n**Pre-flight (ALL must pass before issuing migrate)**:\n- CPU compatibility: target host CPU family must match source, OR cluster must be in EVC mode. Live migration across mismatched CPUs **fails mid-flight** and may leave the VM stunned.\n- Network parity: every portgroup the VM uses must exist on the target host's vSwitch with the same VLAN. Missing portgroup → vNICs disconnected post-migration.\n- Storage visibility: target host must see all of the VM's datastores; otherwise this is a Storage vMotion, not a host migration — different (slower) operation.\n- Affinity rules: if the VM is pinned to source by a DRS host-affinity rule, migration silently violates intent. Check `cluster info` first.\n- Hardware passthrough: VMs with PCI passthrough (GPU, USB) **cannot live-migrate** — schedule a cold migration window.\n\n**Steps**:\n1. Verify VM state and current host via `vmware-monitor vm info <name>`\n2. Verify target host: same cluster, EVC compatible, has required networks/datastores\n3. `vmware-aiops vm migrate <name> --to-host <target>` — wait for task completion, do not assume success on return\n4. Post-check: `vm info` confirms new host AND power state unchanged AND vNICs connected\n\n## Usage Mode\n\n| Scenario | Recommended | Why |\n|----------|:-----------:|-----|\n| Local/small models (Ollama, Qwen) | **CLI** | ~2K tokens vs ~8K for MCP |\n| Cloud models (Claude, GPT-4o) | Either | MCP gives structured JSON I/O |\n| Automated pipelines | **MCP** | Type-safe parameters, structured output |\n\n## MCP Tools (60 — 17 read, 43 write)\n\n| Category | Tools | R/W |\n|----------|-------|:---:|\n| VM Lifecycle (16) | `vm_list_ttl`, `vm_list_snapshots`, `vm_task_status` | Read |\n| | `vm_power_on`, `vm_power_off`, `vm_create`, `vm_reconfigure`, `vm_clone`, `vm_migrate`, `vm_delete`, `vm_create_snapshot`, `vm_revert_snapshot`, `vm_delete_snapshot`, `vm_set_ttl`, `vm_cancel_ttl`, `vm_clean_slate` | Write |\n| Deployment (8) | `deploy_vm_from_ova`, `deploy_vm_from_template`, `deploy_linked_clone`, `attach_iso_to_vm`, `convert_vm_to_template`, `batch_clone_vms`, `batch_linked_clone_vms`, `batch_deploy_from_spec` | Write |\n| Guest Ops (5) | `vm_guest_exec`, `vm_guest_exec_output`, `vm_guest_upload`, `vm_guest_download`, `vm_guest_provision` | Write |\n| Plan/Apply (4) | `vm_list_plans` | Read |\n| | `vm_create_plan`, `vm_apply_plan`, `vm_rollback_plan` | Write |\n| Datastore (2) | `browse_datastore`, `scan_datastore_images` | Read |\n| Network (7) | `list_dvs_portgroups`, `list_host_vmks`, `vmk_ping` | Read |\n| | `create_dvs_portgroup`, `add_host_vmk`, `remove_host_vmk`, `set_vmk_service` | Write |\n| Cluster (10) | `cluster_info`, `list_drs_rules` | Read |\n| | `cluster_create`, `cluster_delete`, `cluster_add_host`, `cluster_remove_host`, `cluster_configure`, `set_drs_rule_enabled`, `create_drs_rule`, `delete_drs_rule` | Write |\n| Alarm Management (3) | `list_vcenter_alarms` | Read |\n| | `acknowledge_vcenter_alarm`, `reset_vcenter_alarm` | Write |\n| Cluster Triage (1) | `cluster_health_summary` (delegates to vmware-monitor) | Read |\n| Object Investigation (4) | `vm_investigation_bundle`, `host_investigation_bundle`, `datastore_investigation_bundle`, `cross_vcenter_attention` (all delegate to vmware-monitor) | Read |\n\n**List envelope**: the read list tools — `browse_datastore`, `list_vcenter_alarms`, `vm_list_plans`, `vm_list_snapshots`, `vm_list_ttl` — return `{items, returned, limit, total, truncated, hint}` rather than a bare array. Read the rows from `items` and check `truncated` before concluding a listing is complete; empty `items` with `truncated: false` means checked-and-none, not a failure. The write `batch_*` tools keep their bare list (complete by construction). Rationale, `total` semantics, error shape: `references/capabilities.md`.\n\n**Read/write split**: 17 tools are read-only (per `[READ]` docstring marker), 43 modify state — gating in [Audit & Safety](#audit--safety). `vm_set_ttl` schedules an unattended auto-delete.\n\n**Network write gating**: all four network writes are preview/confirm-gated — `confirm=False` (default) returns the exact spec without writing. `remove_host_vmk` is **fail-closed**: it refuses when the vmk is selected for a host service (management/vMotion/vSAN), lives on a non-default netstack (NSX TEPs, dedicated vMotion stacks), carries a default gateway route, or when any of that cannot be verified — pass `force_unprotected=True` to override the non-absolute protections. The host's only management-enabled vmk is never removable (no override). `set_vmk_service` is **fail-closed** too: it refuses both directions when the host's service map is unreadable, and refuses (no override) to untag `management` from the host's only management-enabled vmk — the call rides the interface it would untag.\n\n**DRS rule gating**: `set_drs_rule_enabled`, `create_drs_rule`, `delete_drs_rule` are preview/confirm-gated and idempotent (matching state returns a no-write noop). `create_drs_rule` handles VM-VM affinity/anti-affinity only (≥2 distinct VMs, all cluster members); VM-Host rules are read via `list_drs_rules` but managed in the vSphere UI. `delete_drs_rule` **refuses non-VM-VM rules** (they can carry licensing/compliance placement constraints) and records the full rule definition in both preview and result so a mistaken delete can be recreated from the audit trail.\n\n**Alarm reset blast radius**: vSphere has no per-alarm clear API. `reset_vcenter_alarm` uses `AlarmManager.ClearTriggeredAlarms`, which clears **all** triggered alarms matching the named alarm's entity type (host/VM/all) and current status (red/yellow) — not just the one named. The response's `scope` field states exactly what was cleared. The named alarm is looked up first, so a typo fails fast without clearing anything.\n\n## CLI Quick Reference\n\n```bash\n# VM operations\nvmware-aiops vm power-on <name> [--target <t>]\nvmware-aiops vm power-off <name> [--force]\nvmware-aiops vm create <name> --cpu 4 --memory 8192 --disk 100\nvmware-aiops vm delete <name>\nvmware-aiops vm clone <name> --new-name <new> [--to-host <host>] [--to-datastore <ds>] [--power-on]\nvmware-aiops vm migrate <name> --to-host <host> [--to-datastore <ds>]\nvmware-aiops vm snapshot-create <name> --name <snap> [--description <text>] [--memory]\nvmware-aiops vm snapshot-list <name>\nvmware-aiops vm snapshot-revert <name> --name <snap>\nvmware-aiops vm snapshot-delete <name> --name <snap> [--remove-children] [--no-wait]\nvmware-aiops vm task-status <task-id>                      # poll an async (--no-wait) operation by id\nvmware-aiops vm set-ttl <name> --minutes 480 [--dry-run]   # double confirm; daemon auto-deletes VM on expiry\n\n# Guest operations (requires VMware Tools)\nvmware-aiops vm guest-exec <name> --cmd <script-path> --args \"<args>\" --user <username>\nvmware-aiops vm guest-upload <name> --local ./script.sh --guest /tmp/script.sh --user <username>\n\n# Deploy\nvmware-aiops deploy ova <path> --name <vm> --datastore <ds>\nvmware-aiops deploy linked-clone --source <vm> --snapshot <snap> --name <new>\n\n# Cluster\nvmware-aiops cluster create <name> --ha --drs\nvmware-aiops cluster info <name>\nvmware-aiops cluster drs-rules <name>                                     # list DRS rules\nvmware-aiops cluster drs-rule-set <name> --rule <r> --enable|--disable [--dry-run]\nvmware-aiops cluster drs-rule-create <name> --rule <r> --type antiAffinity --vm <vm1> --vm <vm2> [--disabled] [--dry-run]\nvmware-aiops cluster drs-rule-delete <name> --rule <r> [--dry-run]        # VM-VM only; double confirm\n\n# Datastore\nvmware-aiops datastore browse <ds> --pattern \"*.ova\"\n\n# Alarm management\nvmware-aiops alarm list [--target <t>]\nvmware-aiops alarm acknowledge <entity_name> <alarm_name> [--target <t>]\nvmware-aiops alarm reset <entity_name> <alarm_name> [--target <t>]   # double confirm; see blast radius above\n\n# Family\nvmware-aiops hub status        # show installed family members + install commands\n```\n\n> Full CLI reference: see `references/cli-reference.md`\n\n## Troubleshooting\n\n### \"VM not found\" error\nVM names are case-sensitive in vSphere. Use exact name from `vmware-monitor inventory vms`.\n\n### Guest exec returns empty output\nUse `vm_guest_exec_output` instead of `vm_guest_exec` — it auto-captures stdout/stderr. Basic `vm_guest_exec` only returns exit code.\n\n### Deploy OVA times out\nLarge OVA files (>10GB) may exceed the default 120s timeout. The upload happens via HTTP NFC lease — ensure network between the machine running vmware-aiops and ESXi is stable.\n\n### Snapshot delete is slow / \"still running after Ns\"\nDeleting an old or large snapshot consolidates its delta disk into the parent — the slowest write\noperation, often several minutes. `vm snapshot-delete` waits up to 30 min by default; if it still\nreturns a \"still running, NOT failed\" message with a task id, the delete did **not** fail — poll it with\n`vm task-status <task-id>`. Do not re-issue the delete or hand-roll polling. For very large snapshots,\nprefer `vm snapshot-delete <name> --name <snap> --no-wait` to get the task id immediately and poll.\n\n### Plan apply fails mid-way\nRun `vmware-aiops plan list` to see failed plan status. Ask user if they want to rollback with `vm_rollback_plan`. Irreversible steps (delete_vm) are skipped during rollback.\n\n### Connection refused / SSL error\n1. Verify target is reachable: `vmware-aiops doctor`\n2. For self-signed certs: set `verify_ssl: false` in config.yaml (lab environments only)\n\n## Setup\n\n```bash\nuv tool install vmware-aiops==1.12.0\nmkdir -p ~/.vmware-aiops\nvmware-aiops init  # generates config.yaml and .env templates\nchmod 600 ~/.vmware-aiops/.env\n```\n\n> Full setup guide, security details, and AI platform compatibility: see `references/setup-guide.md`\n\n## License\n\nMIT — [github.com/vmware-skills/VMware-AIops](https://github.com/vmware-skills/VMware-AIops)\n\nFile v1.12.0:_meta.json\n\n{\n  \"ownerId\": \"kn7b067awq2s97bn3d7p5qfhw5827pxc\",\n  \"slug\": \"vmware-aiops\",\n  \"version\": \"1.12.0\",\n  \"publishedAt\": 1789915898030\n}\n\nFile v1.12.0:references/agent-guardrails.md\n\n# Operating vmware-aiops with a local / small model\n\nClaude-class models drive this skill without special instruction. Smaller and\nlocally-hosted models — Llama 3.3 70B, Qwen, Mistral, and similar, served\nthrough Goose, Ollama, or OpenShift AI — need explicit operating rules to call\ntools reliably.\n\nThis page exists because an operator wrote those rules by hand first. The\nguardrails below are adapted, with thanks, from the working configuration\n[@juanpf-ha](https://github.com/juanpf-ha) developed while running\nvmware-monitor and vmware-aria against a production vSphere estate with Llama\n3.3 70B FP8 on an on-prem H100\n([VMware-AIops#31](https://github.com/vmware-skills/VMware-AIops/issues/31)). The\ncross-skill rules are identical across this family; the parts below marked\nvmware-aiops are specific to this skill.\n\nvmware-aiops carries the family's largest write surface — 43 of its 60 MCP\ntools change state, including `vm_delete`, cluster deletion, host VMkernel\nremoval and guest command execution. Of every skill here, this is the one where a model's discipline\nshould not be the only thing standing between a prompt and a destroyed VM — and\nover MCP, apart from optional deny rules, the only enforcement is the RBAC of\nthe vCenter/ESXi account the server connects with. Run it under a dedicated,\nleast-privilege service account scoped to what the agent may change.\n\n> **Disclaimer**: This is a community-maintained open-source project and is\n> **not affiliated with, endorsed by, or sponsored by VMware, Inc. or Broadcom\n> Inc.** \"VMware\" and \"vSphere\" are trademarks of Broadcom.\n\n---\n\n## First: the rules you no longer need to write\n\nSeveral guardrails from the original configuration are now enforced by the\nskill itself. Prompt instructions are advisory — a model can ignore them.\nThese are structural, so it cannot.\n\n| Guardrail you would otherwise prompt for | Now enforced by |\n|---|---|\n| \"Use explicit limits for queries that may return large amounts of data\" | **The list envelope.** `browse_datastore`, `list_vcenter_alarms`, `vm_list_plans`, `vm_list_snapshots` and `vm_list_ttl` return `{items, returned, limit, total, truncated, hint}`, so the model reads truncation instead of guessing at it. |\n| \"If a listing came back empty, say so rather than claiming the call failed\" | Same envelope. Empty `items` with `truncated: false` means checked-and-none — a stated result, not a silence the model has to interpret. |\n| \"Log every state change you make\" | **The `@vmware_tool` decorator.** Every write is recorded to `~/.vmware/audit.db` before the model sees the result, and policy rules are evaluated ahead of execution. Neither depends on the model cooperating. |\n| \"Block state-changing writes against a production target\" | **Policy.** An opt-in environment-scoped `deny` rule in `~/.vmware/rules.yaml` matches a target's `environment:` label and refuses matching writes before execution. |\n\n---\n\n## The system prompt\n\nEverything below still benefits from being stated explicitly. Copy this into\nyour agent's instruction block.\n\n```text\n## Tool use\n\n- Always call an MCP tool before answering any question about the current\n  VMware environment. Never answer from memory or assumption.\n- Never describe a tool call, and never output a JSON example, instead of\n  executing the tool. If you intend to call a tool, call it.\n- If a tool fails, report the actual error text. Do not complete the answer\n  with assumptions about what the result would have been.\n- Use explicit limits on queries that may return large amounts of data. Do not\n  request unlimited results unless the user asks for them.\n- Before any write, restate the exact object you are about to change and wait\n  for the user to confirm it. VM names are case-sensitive and near-duplicates\n  are common.\n\n## Skill routing\n\n- vmware-aiops: VM lifecycle (power, create, clone, migrate, delete,\n  snapshots), OVA/template deployment, guest operations, clusters, plan/apply.\n- vmware-monitor: read-only vCenter inventory, hosts, datastores, alarms,\n  events, performance. Prefer it for any question that only reads.\n- vmware-storage: iSCSI, vSAN, datastore capacity.\n- vmware-vks: Supervisor, namespaces, Tanzu Kubernetes clusters.\n- vmware-nsx / vmware-nsx-security: networking and firewall.\n- vmware-aria: Aria Operations metrics, alerts, capacity.\n- vmware-pilot: multi-step workflows that need approval gates.\n\n## Data fidelity\n\n- Never invent infrastructure objects, metrics, alarms, events, or\n  relationships. If a tool did not return it, it does not exist for this answer.\n- Preserve the exact power state, task state, status and criticality values the\n  tools return. Do not translate, normalise, or prettify enum values.\n- If a requested field was not returned, show it as \"not available\". Do not\n  infer it from other fields.\n- Preserve the original order and the full set of fields when the user asks\n  for specific ones.\n- When a response is long, report every item it contains. If a result is\n  truncated, the tool says so explicitly — report the truncation rather than\n  describing the visible subset as the whole.\n\n## Analysis discipline\n\n- Separate observed data from interpretation. State which is which.\n- Do not claim a capacity, performance, or configuration problem unless the\n  tool output contains explicit supporting evidence.\n- Avoid generic recommendations that are not directly supported by the results.\n\n## Writes in vmware-aiops\n\n- 22 of the 43 write tools — every destructive one: vm_power_off, vm_delete,\n  vm_migrate, vm_revert_snapshot, vm_delete_snapshot, vm_clean_slate,\n  vm_set_ttl, the four guest tools (vm_guest_exec, vm_guest_exec_output,\n  vm_guest_upload, vm_guest_provision), cluster_delete, cluster_remove_host,\n  vm_apply_plan, vm_rollback_plan and the seven host-network and DRS tools —\n  default to confirm=false, which returns {\"action\": \"preview\",\n  \"blast_radius\": ...} and writes nothing. Show the blast radius; pass\n  confirm=true only after the user has seen it and agreed — a request to\n  \"delete X\" made before the preview is not that agreement. confirm=true is\n  refused when the preview listed blockers or could not read something: fix\n  the cause and preview again, do not retry blindly. vm_delete also needs\n  acknowledge_blast_radius set to the preview's acknowledge_with, and refuses if\n  the VM changed since. The other 21 write tools (create, clone, deploy,\n  power-on, reconfigure, snapshot create, alarms) act on the first call; for\n  them the \"restate the object and wait\" rule above is the only confirmation\n  step, and it is yours to keep.\n- vm_guest_exec, vm_guest_exec_output and the exec steps of vm_guest_provision\n  run an unbounded command inside the guest with the credentials given; the\n  username is required — there is no default account. Treat them as the highest-risk tools in the skill;\n  pass the least-privileged guest account that can do the job, name the exact\n  command and the VM before calling, and never assemble the command from text a\n  tool returned. vm_guest_upload copies a local file into the guest; upload only\n  files the user named.\n- reset_vcenter_alarm has a blast radius: vSphere has no per-alarm clear API,\n  so it clears every triggered alarm matching the named alarm's entity type and\n  status, not only the one named. Report the response's scope field verbatim.\n- vm_set_ttl schedules an unattended auto-delete. Treat it as destructive and\n  say so when proposing it.\n- Long writes return a task id instead of blocking. Poll vm_task_status. A\n  \"still running\" message is not a failure — never re-issue the operation.\n- Use vm_guest_exec_output rather than vm_guest_exec when the user wants the\n  command's output; the latter returns only an exit code.\n```\n\n---\n\n## Known failure modes on small models\n\nObserved with Llama 3.3 70B FP8 (Goose, on-prem H100), and useful as a\nchecklist when evaluating any local model against these skills:\n\n| Symptom | Mitigation |\n|---|---|\n| Describes a tool call, or emits a JSON example, instead of executing it | The \"never describe a tool call\" rule above. Also check your harness is not echoing tool schemas into context — models imitate the nearest format they see. |\n| Long tool responses: omits items, or reports \"no data returned\" when data was present | Ask for explicit limits so responses stay small. Check the envelope's `truncated` / `returned` / `total` fields rather than trusting the model's summary — a \"no data\" claim is checkable against `returned`. |\n| Adds generic recommendations unsupported by results | The \"analysis discipline\" rules. |\n| Drops requested fields or reorders results | State the required fields and ordering in the request itself, not only in the system prompt. |\n| Multi-tool workflows take 30–50s end to end | Prefer the aggregate tools — `cluster_health_summary`, `vm_investigation_bundle`, `host_investigation_bundle`, `datastore_investigation_bundle`, `cross_vcenter_attention` — which collapse a 3-4 call sequence into one round trip. |\n| Picks a write tool for a question that only reads | Route read questions to vmware-monitor. A model that can see 43 write tools will sometimes reach for one to \"check\" something. |\n| Treats a long-running task's \"still running\" reply as a failure and re-issues the write | The `vm_task_status` rule above. A re-issued clone or delete is the worst outcome in this skill. |\n| Assumes an alarm reset cleared only the alarm it named | Report `scope` from the response. The clear is entity-type-wide by design. |\n\n## Reporting results\n\nLocal-model compatibility is an explicit design constraint for this family, and\nthe evidence base is small. If you evaluate a model against this skill —\nQwen, Mistral, Granite, or anything else — a report of what worked and what did\nnot is genuinely useful:\n[github.com/vmware-skills/VMware-AIops/issues](https://github.com/vmware-skills/VMware-AIops/issues).\n\nFile v1.12.0:references/capabilities.md\n\n# Capabilities Reference\n\n## Automation Level Reference\n\nEach operation is classified by autonomy level per the Enterprise Harness Engineering framework. This tells AI agents how much human gating each tool needs:\n\n| Level | Meaning | Agent autonomy | Examples in this skill |\n|:-:|---|---|---|\n| **L1** | Read-only, raw data | Always auto-run | `cluster_info`, `browse_datastore`, `scan_datastore_images`, `list_vcenter_alarms`, `vm_list_snapshots`, `vm_list_ttl`, `vm_task_status` |\n| **L2** | Read + analysis / recommendation | Always auto-run | `cluster_health_summary`, `cross_vcenter_attention`, `vm_investigation_bundle`, `host_investigation_bundle`, `datastore_investigation_bundle`; scheduled scan reports, alarm/event correlation, log pattern analysis |\n| **L3** | Single write | The level is a statement about blast radius, not about an enforced gate. On the CLI the destructive ones double-confirm (`vm power-on` and `vm snapshot-create` do not); over MCP every destructive one (`vm_power_off`, `vm_delete`, `vm_migrate`, …) returns a no-write preview of its blast radius unless called with `confirm=True`, while `vm_power_on`, `vm_create_snapshot` and `vm_clone` act on the first call — see [What gates a write](#what-gates-a-write) | `vm_power_on`, `vm_power_off`, `vm_delete`, `vm_create_snapshot`, `vm_clone`, `vm_migrate` |\n| **L4** | Multi-step plan / apply workflow | Plan generation auto. Review with the user before applying — an agent convention, not enforced: `vm_apply_plan` takes only a plan id | `vm_create_plan` → `vm_apply_plan` → `vm_rollback_plan`, batch-clone, batch-deploy YAML |\n| **L5** | Auto-remediation from learned pattern | Pattern library only; requires `risk:low` + `reversible:true` + `repeatable:true` + signed approval | *(roadmap — not implemented; candidates: snapshot consolidation, orphaned VM cleanup)* |\n\n**Notes**:\n- L1/L2 tools are read-only and safe for agents to call unprompted.\n- **The levels describe risk, not enforcement.** The MCP previews stop an agent acting blind, and `confirm=True` is refused on a blocker, but nothing in this skill stops an agent that passes `confirm=True` from calling an L3 or L4 tool the account may call. What decides whether the write lands is the vCenter account — see [What gates a write](#what-gates-a-write).\n- **List envelope**: the read list tools (`browse_datastore`, `list_vcenter_alarms`, `vm_list_plans`, `vm_list_snapshots`, `vm_list_ttl`) return `{items, returned, limit, total, truncated, hint}` instead of a bare array, so an agent can tell a complete answer from a first page rather than inferring it (issue #31). All five enumerate their collection in full before any limit is applied, so `total` is always the real count; only `list_vcenter_alarms` takes a `limit` and can therefore report `truncated: true`. The write `batch_*` tools deliberately keep a bare list — each row is a per-item result of work already done, complete by construction. Errors from these read tools are `{error, hint}` (a dict, not a one-element list).\n- L3+ tools always pass through the `@vmware_tool` decorator: connection check → policy check (opt-in `deny` rules only; nothing is denied by default) → audit log. Confirmation is not in that chain; where a tool has one, it is the tool's own `confirm` argument (see [What gates a write](#what-gates-a-write)).\n- Multi-party approval, where it is genuinely required, is [vmware-pilot](https://github.com/vmware-skills/VMware-Pilot)'s job — it has the state machine and a real human approval step. See it also for cross-skill L4 orchestration and the Dispatcher/Subagent pattern.\n\n## What gates a write\n\n<!-- gate-inventory: the counts and names below are checked against the live\n     tool registry by tests/eval/regression/test_documented_gates_match_the_registry.py.\n     Adding a write tool makes them wrong and turns that test red. -->\n\n| Surface | Confirmation | Preview |\n|---|---|---|\n| **CLI** | Two interactive `typer.confirm` prompts before every irreversible or guest-writing command. The set is derived from the MCP `destructiveHint` annotations, so a new such command fails the test suite until it has them, and a declined prompt is audited as `rejected`. *Honest limitation:* an agent with a shell satisfies both prompts by piping `yes` into the command. This defends the mistyped command, not a determined caller. | `--dry-run` on every write command except `deploy iso`, `deploy mark-template`, `vm cancel-ttl` and `vm guest-download` |\n| **MCP** | **One argument, `confirm`, defaulting to a no-write preview** (family security HLD §7, revised 2026-09-16). Decision **D-2** (2026-07-21) cut a confirmation handshake as a speed-bump that does not authorize anything; it is superseded. The handshake still is not authorization — the account is (below) — but a preview is what stops an agent acting on a guess, and a refusal is what stops it acting on a guess that has gone stale. The read-only switch removed in v1.8.7 stays removed. Every write tool annotated `destructiveHint: true` takes it: a bare call measures with reads only and returns `{\"action\": \"preview\", \"blast_radius\": {...}}` without changing anything; `confirm=True` re-measures and is refused — a teaching error, audited as a failure — when the measurement found a blocker or could not read a field it needs. `vm_delete` additionally requires echoing the preview's `acknowledge_with` back in `acknowledge_blast_radius`, and is refused if the VM changed since, is powered on or suspended. | 22 of the 43 write tools default to a no-write preview (below) |\n\n**What actually protects the estate over MCP is the vCenter/ESXi service account.**\nWrites the account may not perform are refused by vCenter itself, whatever the\nagent intends, on every surface, with no way around it from inside the skill. To\nrun the server under a dedicated, least-privilege service account scoped to\nwhat the agent may change; to run it read-only, give it a read-only vCenter role\nand point the skill's `.env` at that account — one decision, enforced where it\nis made. What happened is then recoverable from `~/.vmware/audit.db`, which\nrecords every MCP call (credentials redacted) before the caller sees a result;\nthe write is best-effort, so a failing audit store warns on stderr rather than\nblocking the call. The skill's previews and refusals stop an agent acting blind;\nthey do not stop one determined to delete a VM the account is allowed to delete.\n\nThe one skill-side control that can refuse a call is optional: `deny` rules and a maintenance window in\n`~/.vmware/rules.yaml` (vmware-policy) are evaluated before every MCP call and every CLI command that reaches vCenter, and\ncan refuse operations — for example, writes to targets whose `config.yaml` entry\ndeclares `environment: production`. The shipped baseline denies nothing, and an\nunreadable rules file fails closed. It runs in the same process as the tools\n(`VMWARE_POLICY_DISABLED=1` switches it off), so it is a guardrail for\nwell-meaning agents, not a replacement for RBAC.\n\n- **Write tools: 43** — every tool whose description starts `[WRITE]` and whose `readOnlyHint` is `false`.\n- **Confirm-gated: 22** — `add_host_vmk`, `cluster_delete`, `cluster_remove_host`, `create_drs_rule`, `create_dvs_portgroup`, `delete_drs_rule`, `remove_host_vmk`, `set_drs_rule_enabled`, `set_vmk_service`, `vm_apply_plan`, `vm_clean_slate`, `vm_delete`, `vm_delete_snapshot`, `vm_guest_exec`, `vm_guest_exec_output`, `vm_guest_provision`, `vm_guest_upload`, `vm_migrate`, `vm_power_off`, `vm_revert_snapshot`, `vm_rollback_plan`, `vm_set_ttl`\n  <br>Each takes a `confirm` argument that defaults to false, in which case it measures with reads only and returns `blast_radius` without writing. With `confirm=True` it re-measures and refuses on a blocker or an unreadable field instead of guessing. What each one measures and refuses:\n  - **Guest** (`vm_guest_exec`, `vm_guest_exec_output`, `vm_guest_upload`, `vm_guest_provision`): the VM, its OS family, the guest account, the command or files, each with its full length (`command_length`, `arguments_length`) and `truncated`. Refused when the VM is not powered on, VMware Tools is not running, a local upload source is missing or unreadable, a provision step is malformed, a `service` step targets a Windows guest (it runs `systemctl`), or a command, argument string or path is longer than the preview shows (1000 characters for commands) — put a long command in a script and upload it.\n  - **VM** (`vm_power_off`, `vm_migrate`, `vm_revert_snapshot`, `vm_delete_snapshot`): power state, host, Tools status, the snapshot and what depends on it. Refused on a graceful shutdown without running Tools or of a suspended VM (`force=True` is a hard power-off — preview it), a target host that is missing, disconnected, in maintenance mode or cannot reach the VM's storage, and a snapshot that is not found or whose name is duplicated.\n  - **Cluster** (`cluster_delete`, `cluster_remove_host`): member hosts, VMs, datastores. Refused while hosts or VMs are still in the cluster, or when the host is not in maintenance mode or still runs powered-on VMs.\n  - **Clean Slate / TTL** (`vm_clean_slate`, `vm_set_ttl`): the snapshot and what the revert discards; for a TTL, the VM that will be deleted unattended and when. Clean Slate is refused when the baseline snapshot is missing or its name is ambiguous — before anything is powered off — and reports an error, not `reverted`, if the revert did not happen. The TTL entry records the VM's target and instance UUID; at expiry the daemon deletes only a VM that still has that UUID (another VM under the same name is left alone, the entry dropped and the refusal audited). A TTL is refused on every `vm_delete` blocker except the power state (the daemon powers off first).\n  - **Snapshot names** (`vm_revert_snapshot`, `vm_delete_snapshot`, `vm_clean_slate`): a name matches a snapshot's raw name or the name `vm_list_snapshots` shows (control/invisible characters stripped). Two snapshots matching the same name — duplicates, or names that differ only by stripped characters — are refused by the gate and by the executors, which no longer take the first.\n  - **Plans** (`vm_apply_plan`, `vm_rollback_plan`): every step with its full parameters (passwords and other secret-looking keys redacted), which are destructive and which cannot be rolled back. Each destructive step is measured as its own tool measures it, and that tool's `blockers` / `unmeasured` become the plan's, prefixed `Step N (action):`. A step on an object an earlier step creates or changes (power off, then delete; a VM a step creates) is shown `check: \"deferred\"` and measured immediately before it runs; every destructive step is re-checked then, and a failed check stops the plan with a teaching error in that step's result. Refused when `target` does not match the plan's (no target is a target of its own), a step would be refused by its tool, or a plan step with action delete_vm carries no `acknowledge_blast_radius`. Every action the executor can run is classified: measured by its tool's gate (including `migrate`, measured as `vm_migrate`), refused, or non-destructive with a stated reason; anything else is refused rather than run unchecked. `iscsi_enable`, `iscsi_add_target`, `iscsi_remove_target` and `storage_rescan` are gated in vmware-storage and cannot be measured here, so a plan containing one is refused at preview, apply and rollback — use `storage_iscsi_*` / `storage_rescan` instead. Rollback steps get the same checks: each destructive rollback step (power off, delete snapshot, cluster delete, host removal) is measured in the `vm_rollback_plan` preview and again just before it runs; a refused check stops the rollback with a teaching error for that step, and the plan stays `failed` so rollback can be run again. A step that creates a VM records its instance UUID; rollback deletes the VM only if it still has that UUID, and refuses that rollback step otherwise. Plan responses never echo step passwords; secret keys are matched as whole name tokens (`password`, `passwd`, `pwd`, `secret`, `token`, `api_key`, `auth`, `credential(s)`, `private_key`).\n  - **Refusal messages** lead with the first blocker (problem and remedy) and count the rest (\"…and 2 more blockers\"), and stay within 480 characters; the preview lists every blocker.\n  - **Network / DRS** (`add_host_vmk`, `remove_host_vmk`, `set_vmk_service`, `create_dvs_portgroup`, `create_drs_rule`, `delete_drs_rule`, `set_drs_rule_enabled`): what would be created, changed or removed, with `blockers` and `unmeasured`; `remove_host_vmk` and `set_vmk_service` fail closed on an unreadable service map. Their refusals (a protected or the only management vmk, a VM-Host DRS rule) are reported as blockers in the preview and still raise on `confirm=True`.\n  <br>Only `vm_delete` also requires `acknowledge_blast_radius` to echo the preview's `acknowledge_with` (instance UUID, disk count, snapshot count), refused if those no longer match when re-measured. Undo tokens are filed only when a call actually changed something, never for a preview.\n- **Ungated: 21** — none is annotated destructive, and each acts on the first call: create/clone/deploy (`vm_create`, `vm_clone`, `deploy_vm_from_ova`, the `batch_*` tools, …), `vm_power_on`, `vm_create_snapshot`, `vm_reconfigure`, `cluster_create`, `cluster_configure`, `cluster_add_host`, `attach_iso_to_vm`, `convert_vm_to_template`, `vm_guest_download`, `vm_cancel_ttl`, `vm_create_plan` and the two alarm tools. \"Not destructive\" is not \"harmless\": `vm_reconfigure` and `cluster_configure` change live configuration without a preview.\n\n### `vm_guest_exec` deserves naming\n\n`vm_guest_exec` runs a caller-supplied command inside the guest OS through\nVMware Tools, with the credentials passed to it — its `username` parameter\nis required — there is no default account. It is the widest blast radius in the\nskill. Over MCP it now previews by default — the VM, the guest account and the\ncommand as it would run — as do `vm_guest_exec_output`, `vm_guest_upload` and\n`vm_guest_provision`. But the preview describes the command; it does not judge\nit. Nothing in this skill bounds what the command may be once `confirm=True`:\n`rm -rf /` is a well-formed argument, and the same is true of the `exec` steps\ninside `vm_guest_provision`.\n\nTwo things follow. First, the guest credentials are a second, separate\nauthorization boundary — a read-only vCenter role does not constrain what these\ntools do *inside* a VM, because that is decided by the guest account. Give the\nskill a guest account with the privileges the work actually needs, and pass it —\n`username` has no default, so the account is always a deliberate choice. Over MCP the guest\npassword is an ordinary tool argument, so the agent — and its transcript — sees\nit; only the audit row redacts it. The skill stores no guest credentials of its\nown. Relatedly, `vm_guest_upload` and the `upload` steps of `vm_guest_provision`\nread any local file the server process can read and copy it into the guest.\nSecond, until\n2026-08-30 all four tools that push content into a guest were annotated\n`destructiveHint: false` — the field a client consults before deciding whether\nto ask its user. They now declare `true`. These tools exist only in this repo,\nso that correction is complete here; whether the comparable high-blast-radius\ntools in the other thirteen skills carry honest annotations is a family-wide\nquestion this repo cannot settle on its own.\n\n<!-- /gate-inventory -->\n\n## Triage & Object Investigation (read-only)\n\nFive opinionated read-only reports that **aggregate and correlate server-side** and\nreturn high-signal results — never raw inventory. They exist so the agent can decide\n*where to look* before actuating anything. All five delegate to the\n[vmware-monitor](https://github.com/vmware-skills/VMware-Monitor) library using AIops' own\nvCenter connection, so **`vmware-monitor` must be installed**; without it these tools\nare unavailable. All are point-in-time (no trending). Each has a `--html` CLI form\nthat writes a self-contained, timestamped offline snapshot (no external references,\ndrill-downs collapse via native `<details>`, zero JavaScript).\n\n| Operation | CLI | MCP Tool | vCenter | ESXi |\n|-----------|-----|----------|:-------:|:----:|\n| Cluster health summary | `summary` | `cluster_health_summary` | ✅ | ❌ |\n| Cross-vCenter attention | `attention` | `cross_vcenter_attention` | ✅ | ❌ |\n| VM investigation bundle | `investigate vm <name>` | `vm_investigation_bundle` | ✅ | ✅ |\n| Host investigation bundle | `investigate host <name>` | `host_investigation_bundle` | ✅ | ✅ |\n| Datastore investigation bundle | `investigate datastore <name>` | `datastore_investigation_bundle` | ✅ | ✅ |\n\n### `cluster_health_summary` — \"is anything on fire?\"\n\nThe first look. Rolls up hosts, VM power state, live CPU/memory pressure and triggered\nalarms per cluster, assigns an opinionated `ok` / `warn` / `critical` status, and\nflattens individual anomalies into a ranked `top_issues` focus list (worst first, each\ncarrying a drill-down hint). Returns `{totals, top_issues, issues_total, clusters,\nsnapshot, customization_hint}` — lead with `top_issues`, show `clusters` as context.\n\n| Parameter | Type | Default | Behavior |\n|-----------|------|---------|----------|\n| `target` | str (optional) | default target | Named vCenter/ESXi target from `config.yaml` |\n| `cluster_filter` | str (optional) | None (all) | Case-insensitive substring; suppresses standalone-hosts bucket |\n| `include_vms` | bool | True | Roll up VM power counts; False skips the VM pass (faster on huge fleets) |\n| `top_n` | int | 10 | Cap the `top_issues` focus list; `issues_total` keeps the pre-cap count; 0 hides the list |\n\n**Typical response tokens**: ~120–400 (one compact row per cluster + totals); scales\nwith cluster count, not VM count. Aggregation happens in the tool — the model never\nsees raw inventory.\n\n### `cross_vcenter_attention` — \"where do I look first, anywhere in the estate?\"\n\nMerges every configured target's cluster-health summary into a single globally ranked\n`top_issues` list (each item tagged with its `vcenter`) plus a per-target rollup.\nDegrades gracefully: an unreachable target is listed under `unreachable` and the rest\nstill aggregate. Use it before `cluster_health_summary` when more than one vCenter is\nconfigured; with a single target, go straight to `cluster_health_summary`.\n\n| Parameter | Type | Default | Behavior |\n|-----------|------|---------|----------|\n| `cluster_filter` | str (optional) | None (all) | Case-insensitive cluster substring applied to every target |\n| `top_n` | int | 10 | Cap the merged `top_issues` focus list |\n\n**Typical response tokens**: ~200–600 (ranked issue list + one row per target); scales\nwith target count, not inventory size.\n\n### `*_investigation_bundle` — one correlated drill-down per object\n\nUse **after** triage points at a specific object. Each bundle collects and *correlates*\nthe object with its surrounding infrastructure and recent history in one batched call,\nso the agent does not stitch together separate info/alarm/snapshot/performance/event\nreads. All three accept `hours` (event-timeline look-back, default 24) and an optional\n`target`. An unknown object name returns a teaching error naming how to list objects.\n\n| Tool | Required arg | Correlates |\n|------|--------------|------------|\n| `vm_investigation_bundle` | `vm_name` | VM state, the host it runs on, cluster context, backing datastores, snapshots, triggered alarms, live performance, merged event timeline (VM + host + cluster + datastores, newest first) |\n| `host_investigation_bundle` | `host_name` | Connection state, CPU/memory, ESXi version, uptime, cluster context, rollup of VMs it runs, datastores it mounts, alarms across host/cluster/datastore, live performance, merged event timeline |\n| `datastore_investigation_bundle` | `datastore_name` | Capacity/free space/accessibility, hosts that mount it, rollup of VMs it backs, alarms across datastore/host, merged event timeline. (Per-datastore latency is a separate perf report, not included.) |\n\n**Typical response tokens**: ~400–1200 per bundle (correlated summary + capped event\ntimeline); grows with the `hours` window, not with fleet size. Explain the result in\noperational language — do not dump it raw.\n\n## VM Lifecycle\n\n| Operation | Command | Confirmation | vCenter | ESXi |\n|-----------|---------|:------------:|:-------:|:----:|\n| Power On | `vm power-on <name>` | — | ✅ | ✅ |\n| Graceful Shutdown | `vm power-off <name>` | Double | ✅ | ✅ |\n| Force Power Off | `vm power-off <name> --force` | Double | ✅ | ✅ |\n| Reset | plan action `reset` via `vm_create_plan` (MCP; no CLI command) | — | ✅ | ✅ |\n| Suspend | plan action `suspend` via `vm_create_plan` (MCP; no CLI command) | — | ✅ | ✅ |\n| VM Info | `vmware-monitor vm info <name>` (companion skill) | — | ✅ | ✅ |\n| Create VM | `vm create <name> --cpu --memory --disk` | — | ✅ | ✅ |\n| Delete VM | `vm delete <name>` | Double | ✅ | ✅ |\n| Reconfigure | `vm reconfigure <name> --cpu --memory` | Double | ✅ | ✅ |\n| Create Snapshot | `vm snapshot-create <name> --name <snap> [--description <text>] [--memory]` | — | ✅ | ✅ |\n| List Snapshots | `vm snapshot-list <name>` | — | ✅ | ✅ |\n| Revert Snapshot | `vm snapshot-revert <name> --name <snap>` | Double | ✅ | ✅ |\n| Delete Snapshot | `vm snapshot-delete <name> --name <snap> [--remove-children]` | Double | ✅ | ✅ |\n| Poll Async Task | `vm task-status <task-id>` | — | ✅ | ✅ |\n| Clone VM | `vm clone <name> --new-name <new> [--to-host <host>] [--to-datastore <ds>]` | Double | ✅ | ✅ |\n| vMotion | `vm migrate <name> --to-host <host> [--to-datastore <ds>]` | Double | ✅ | ❌ |\n| Set TTL | `vm set-ttl <name> --minutes <n>` | Double | ✅ | ✅ |\n| Cancel TTL | `vm cancel-ttl <name>` | — | ✅ | ✅ |\n| List TTLs | `vm list-ttl` | — | ✅ | ✅ |\n| Clean Slate | `vm clean-slate <name> [--snapshot baseline]` | Double | ✅ | ✅ |\n| Guest Exec | `vm guest-exec <name> --cmd /bin/bash --args \"-c 'whoami'\" --user <account>` | Double | ✅ | ✅ |\n| Guest Upload | `vm guest-upload <name> --local f.sh --guest /tmp/f.sh --user <account>` | Double | ✅ | ✅ |\n| Guest Download | `vm guest-download <name> --guest /var/log/syslog --local ./syslog --user <account>` | — | ✅ | ✅ |\n\n> Guest Operations require VMware Tools running inside the guest OS. `--user` is required on all three commands (and `username` on the MCP guest tools): there is no default account, so no call runs as root without choosing root.\n\n> `vm task-status` / `vm_task_status` polls a vSphere task id returned by an async\n> write (today: `vm_delete_snapshot`) instead of re-running the operation. Returns\n> state (`queued` / `running` / `success` / `error` / `gone`), progress percent, and\n> the entity name. `gone` means vCenter already garbage-collected a completed task —\n> re-list the resource to confirm the final state. A failed task carries its fault\n> under `task_error`, not `error` — the poll succeeded, the task did not.\n> **Typical response tokens**: ~40–80 (single status record).\n\n## Plan → Apply (Multi-step Operations)\n\nFor complex operations involving 2+ steps or 2+ VMs, use the plan/apply workflow:\n\n| Step | MCP Tool / CLI | Description |\n|------|---------------|-------------|\n| 1. Create Plan | `vm_create_plan` | Validates actions, checks targets in vSphere, generates plan with rollback info |\n| 2. Review | — | AI shows plan to user: steps, affected VMs, irreversible warnings |\n| 3. Apply | `vm_apply_plan` | Executes sequentially; stops on failure |\n| 4. Rollback (if failed) | `vm_rollback_plan` | Asks user, then reverses executed steps (skips irreversible) |\n\nPlans are stored in `~/.vmware-aiops/plans/`, deleted on success, auto-cleaned after 24h.\n\n## VM Deployment & Provisioning\n\n| Operation | Command | Speed | vCenter | ESXi |\n|-----------|---------|:-----:|:-------:|:----:|\n| Deploy from OVA | `deploy ova <path> --name <vm>` | Minutes | ✅ | ✅ |\n| Deploy from Template | `deploy template <tmpl> --name <vm>` | Minutes | ✅ | ✅ |\n| Linked Clone | `deploy linked-clone --source <vm> --snapshot <snap> --name <new>` | Seconds | ✅ | ✅ |\n| Attach ISO | `deploy iso <vm> --iso \"[ds] path/to.iso\"` | Instant | ✅ | ✅ |\n| Convert to Template | `deploy mark-template <vm>` | Instant | ✅ | ✅ |\n| Batch Clone | `deploy batch-clone --source <vm> --count <n>` | Minutes | ✅ | ✅ |\n| Batch Deploy (YAML) | `deploy batch spec.yaml` | Auto | ✅ | ✅ |\n\n### Guest Operations Notes\n\n`vm_guest_exec_output` — execute a shell command and **capture stdout/stderr** automatically. OS auto-detected (Linux/Windows) via `vm.guest.guestFamily`. No manual redirection needed.\n\n`vm_guest_provision` — run an ordered sequence of exec/upload/service steps in one call. Stops on first failure. Typical use: SSH key injection → package install → service start.\n\n## Datastore Browser\n\n| Feature | vCenter | ESXi | Details |\n|---------|:-------:|:----:|---------|\n| Browse Files | ✅ | ✅ | List files/folders in any datastore path |\n| Scan Images | ✅ | ✅ | Discover ISO, OVA, OVF, VMDK across all datastores |\n\n> For datastore management, iSCSI, and vSAN, use [vmware-storage](https://github.com/vmware-skills/VMware-Storage). For Tanzu Kubernetes, use [vmware-vks](https://github.com/vmware-skills/VMware-VKS).\n\n## Network (dvSwitch portgroups + host VMkernel)\n\nMCP-only (no CLI subcommand). Seven tools for distributed-switch portgroup and host VMkernel authoring, plus an MTU-path diagnostic. Writes are preview/confirm gated; `remove_host_vmk` and `set_vmk_service` are fail-closed. For NSX overlay segments/gateways/NAT, use [vmware-nsx](https://github.com/vmware-skills/VMware-NSX) — this surface is the underlay (VLAN-backed DVS portgroups, host kernel interfaces).\n\n| Tool | R/W | Risk | Operation |\n|------|:---:|:----:|-----------|\n| `list_dvs_portgroups` | R | low | Distributed portgroups: binding type, VLAN (id/trunk/pvlan), port count, uplink flag. Scope with `dvs_name`. |\n| `create_dvs_portgroup` | W | medium | VLAN-tagged portgroup on a dvSwitch; `earlyBinding` or `ephemeral` (ephemeral attaches with vCenter down — self-hosted-VCSA use case). `confirm=False` previews. |\n| `list_host_vmks` | R | low | VMkernel adapters per host: IP/netmask/dhcp, MTU, MAC, portgroup, netstack, selected services. `services` is `null` (not `[]`) when a host's service map can't be read. A host vCenter could not reach appears as a row with `reachable: false` and null facts rather than being dropped; check `hosts_unreachable` before treating the list as a complete estate inventory. |\n| `add_host_vmk` | W | medium | Static-IP vmk on a DVS portgroup — no gateway, no services (throwaway test-vmk shape). `confirm=False` previews; returns the assigned device (`vmk2`). |\n| `remove_host_vmk` | W | high | Fail-closed removal. Refuses on service selection / non-default netstack / default route / unverifiable state; `force_unprotected=True` overrides all but the only-management-vmk absolute. |\n| `set_vmk_service` | W | medium | Tag/untag a host service (nicType: vmotion, management, vsan, vSphereProvisioning, …) on an existing vmk — completes `add_host_vmk` (adapters are created serviceless). Idempotent; `confirm=False` previews. Fail-closed on unreadable service map; refuses (no override) to untag `management` from the only management-enabled vmk. |\n| `vmk_ping` | R | medium | DF-bit-capable ping sourced from a vmk via esxcli-over-API (no SSH). `df=True size=1572` proves a ≥1600 overlay floor; `size=8972` proves full jumbo. Oversized DF'd packets report `fault` structurally, not as an error. |\n\n> **Typical response tokens**: `list_*` ~60–400 (one compact row per portgroup/vmk, paginated at 200/100); `create`/`add`/`remove` ~40–120 (preview or result record); `vmk_ping` ~80–200 (request + per-summary stats or the esxcli fault text).\n\n## Cluster Management\n\n| Operation | Command | Confirmation | vCenter | ESXi |\n|-----------|---------|:------------:|:-------:|:----:|\n| Cluster Info | `cluster info <name>` | — | ✅ | ❌ |\n| Create Cluster | `cluster create <name> [--ha] [--drs]` | — | ✅ | ❌ |\n| Delete Cluster | `cluster delete <name>` | Double | ✅ | ❌ |\n| Add Host | `cluster add-host <cluster> --host <host>` | Double | ✅ | ❌ |\n| Remove Host | `cluster remove-host <cluster> --host <host>` | Double | ✅ | ❌ |\n| Configure HA/DRS | `cluster configure <name> [--ha/--no-ha] [--drs/--no-drs]` | Double | ✅ | ❌ |\n| List DRS Rules | `cluster drs-rules <name>` | — | ✅ | ❌ |\n| Enable/Disable DRS Rule | `cluster drs-rule-set <name> --rule <r> --enable\\|--disable` | Double | ✅ | ❌ |\n| Create DRS Rule | `cluster drs-rule-create <name> --rule <r> --type affinity\\|antiAffinity --vm <v1> --vm <v2>` | Double | ✅ | ❌ |\n| Delete DRS Rule | `cluster drs-rule-delete <name> --rule <r>` | Double | ✅ | ❌ |\n\n> `remove-host` requires the host to be in **maintenance mode** first; the host is moved out of the cluster into the datacenter's host folder as a standalone host (`Folder.MoveIntoFolder_Task`).\n>\n> **DRS rules**: `drs-rule-create` handles VM-VM affinity/anti-affinity only (≥2 distinct VMs, all cluster members); `drs-rule-delete` refuses VM-Host and other rule types (they can carry licensing/compliance placement constraints — manage those in the vSphere UI) and records the full definition for recreate. All three writes are idempotent (matching state = no-write noop) and support `--dry-run`.\n\n## Alarm Management\n\n| Operation | Command | Confirmation | vCenter | ESXi |\n|-----------|---------|:------------:|:-------:|:----:|\n| List Triggered Alarms | `alarm list [--target <t>]` | — | ✅ | ❌ |\n| Acknowledge Alarm | `alarm acknowledge <entity> <alarm>` | — | ✅ | ❌ |\n| Clear (Reset) Alarms | `alarm reset <entity> <alarm>` | Double | ✅ | ❌ |\n\n> **Blast radius**: vSphere has no per-alarm clear API. `alarm reset` / `reset_vcenter_alarm` uses `AlarmManager.ClearTriggeredAlarms`, which clears **all** triggered alarms matching the named alarm's entity type (host/VM/all) and current status (red/yellow) — not just the named one. The named alarm is looked up first (typos fail fast), and the result's `scope` field reports exactly what was cleared. Cleared alarms re-trigger automatically if their underlying condition persists.\n\n## Scheduled Scanning & Notifications\n\n| Feature | Details |\n|---------|---------|\n| Daemon | APScheduler-based, configurable interval (default 15 min) |\n| Multi-target Scan | Sequentially scan all configured vCenter/ESXi targets |\n| Scan Content | Each cycle: triggered alarms, vCenter events from the last `lookback_hours`, and new lines in the ESXi host logs `hostd`, `vmkernel`, `vpxa` (`scan now` reads alarms and events only) |\n| Host Logs | Read incrementally: each line is reported once per daemon run (a restart re-reads each log's last 500 lines once). A rotated log, or more than 500 new lines between cycles, adds an `info` row. Needs `Global.Diagnostics` (not in vCenter's Read-Only role); an unreadable log becomes an `info` row with the reason |\n| Log Analysis | Host-log lines matching error, fail, critical, panic, lost access, cannot, timeout, refused, corrupt — critical/panic/corrupt lines are `critical`, the rest `warning` |\n| Webhook | Slack, Discord, or any HTTP endpoint. Every critical issue and every alarm/event warning; host-log warnings stay in `scan.log`; `info` rows are never sent. Issue text (entity names, event/log text, connection errors) can include host names, IPs and user names; no credentials |\n| Cycle Summary | One line per cycle: findings (and how many were sent), unreadable host logs, logs with unscanned lines, failed passes; `Scan INCOMPLETE` if any pass failed or a target could not be reached |\n\n## Safety Features\n\n| Feature | Details |\n|---------|---------|\n| Plan → Confirm → Execute → Log | CLI workflow: show current state, confirm changes, execute, audit log |\n| Double Confirmation (**CLI only**) | CLI destructive and deploy commands (`vm` power-off, delete, reconfigure, snapshot-revert/delete, clone, migrate, set-ttl, clean-slate, guest-exec, guest-upload; `deploy` ova, template, linked-clone, batch, batch-clone, mark-template; `cluster` delete, add-host, remove-host, configure, drs-rule-set/create/delete; `alarm reset`) require 2 sequential prompts and take no bypass flag. **The MCP tools have no confirmation step at all** — see [What gates a write](#what-gates-a-write) |\n| Rejection Logging | Declined CLI confirmations are recorded in the audit trail for security review |\n| Audit Trail | Every MCP call and every CLI command that reaches vCenter logged to `~/.vmware/audit.db` (SQLite WAL, via vmware-policy; parameters, result, status, caller — credentials redacted). Most CLI writes also append to `~/.vmware-aiops/audit.log`, with before/after state where the command captures it (power, delete, reconfigure, snapshot-revert, clone, migrate, clean-slate, cluster delete/configure, DRS rule delete) |\n| Input Validation | VM name length/format, CPU (1-128), memory (128-1048576 MB), disk (1-65536 GB) validated before execution |\n| Password Protection | `.env` file loading, never in command line or shell history; file permission check at startup |\n| SSL Self-signed Support | `verify_ssl: false` — **only** for ESXi hosts with self-signed certificates in isolated lab/home environments. Production environments should use CA-signed certificates with full TLS verification enabled. |\n| Task Waiting | All async operations wait for completion and report result |\n| State Validation | Pre-operation checks (VM exists, power state correct) |\n\n## Version Compatibility\n\n| vSphere Version | Support | Notes |\n|----------------|---------|-------|\n| 8.0 / 8.0U1-U3 | ✅ Full | `CreateSnapshot_Task` deprecated → use `CreateSnapshotEx_Task` |\n| 7.0 / 7.0U1-U3 | ✅ Full | All APIs supported |\n| 6.7 | ✅ Compatible | Backward-compatible, tested |\n| 6.5 | ✅ Compatible | Backward-compatible, tested |\n\n> pyVmomi auto-negotiates the API version during SOAP handshake — no manual configuration needed.\n\nFile v1.12.0:references/cli-reference.md\n\n# CLI Reference\n\nDestructive and deploy commands ask for two confirmations and most write\ncommands take `--dry-run` (not `deploy iso`, `deploy mark-template`,\n`vm cancel-ttl`, `vm guest-download`). These are CLI-only: over MCP the 22 destructive write tools take `confirm` and preview by\ndefault, the other 21 write tools act immediately, and the enforcement boundary there is the RBAC of the vCenter/ESXi account — see\n`capabilities.md` → \"What gates a write\".\n\n```bash\n# Diagnostics\nvmware-aiops doctor [--skip-auth]   # --skip-auth only skips doctor's own vSphere login check; no other command has it\n\n# MCP Config Generator\nvmware-aiops mcp-config generate --agent <goose|cursor|claude-code|continue|vscode-copilot|localcowork|mcp-agent>\nvmware-aiops mcp-config list\n\n# VM Operations\nvmware-aiops vm power-on <vm-name>\nvmware-aiops vm power-off <vm-name> [--force]\nvmware-aiops vm create <name> [--cpu <n>] [--memory <mb>] [--disk <gb>]\nvmware-aiops vm delete <vm-name>\nvmware-aiops vm reconfigure <vm-name> [--cpu <n>] [--memory <mb>]\nvmware-aiops vm snapshot-create <vm-name> --name <snap-name> [--description <text>] [--memory]\nvmware-aiops vm snapshot-list <vm-name>\nvmware-aiops vm snapshot-revert <vm-name> --name <snap-name>\nvmware-aiops vm snapshot-delete <vm-name> --name <snap-name> [--remove-children]\nvmware-aiops vm clone <vm-name> --new-name <name> [--to-host <host>] [--to-datastore <ds>] [--power-on]\nvmware-aiops vm migrate <vm-name> --to-host <host> [--to-datastore <ds>]\nvmware-aiops vm set-ttl <vm-name> --minutes <n>\nvmware-aiops vm cancel-ttl <vm-name>\nvmware-aiops vm list-ttl\nvmware-aiops vm clean-slate <vm-name> [--snapshot baseline]\n\n# Guest Operations (requires VMware Tools)\nvmware-aiops vm guest-exec <vm-name> --cmd /bin/bash --args \"-c 'ls -la /tmp'\" --user root\nvmware-aiops vm guest-upload <vm-name> --local ./script.sh --guest /tmp/script.sh --user root\nvmware-aiops vm guest-download <vm-name> --guest /var/log/syslog --local ./syslog.txt --user root\n\n# Plan → Apply (multi-step operations)\nvmware-aiops plan list\n\n# Deploy\nvmware-aiops deploy ova <path> --name <vm-name> [--datastore <ds>] [--network <net>]\nvmware-aiops deploy template <template-name> --name <vm-name> [--datastore <ds>]\nvmware-aiops deploy linked-clone --source <vm> --snapshot <snap> --name <new-name>\nvmware-aiops deploy iso <vm-name> --iso \"[datastore] path/file.iso\"\nvmware-aiops deploy mark-template <vm-name>\nvmware-aiops deploy batch-clone --source <vm> --count <n> [--prefix <prefix>]\nvmware-aiops deploy batch <spec.yaml>\n\n# Cluster\nvmware-aiops cluster info <name>\nvmware-aiops cluster create <name> [--ha] [--drs] [--drs-behavior fullyAutomated|partiallyAutomated|manual] [--datacenter <dc>]\nvmware-aiops cluster delete <name>\nvmware-aiops cluster add-host <cluster> --host <hostname>\nvmware-aiops cluster remove-host <cluster> --host <hostname>   # host must be in maintenance mode; moved to datacenter host folder as standalone\nvmware-aiops cluster configure <name> [--ha/--no-ha] [--drs/--no-drs] [--drs-behavior <behavior>]\nvmware-aiops cluster drs-rules <name>                                                # list VM-VM + VM-Host DRS rules\nvmware-aiops cluster drs-rule-set <name> --rule <name> --enable|--disable [--dry-run] # idempotent; double confirm\nvmware-aiops cluster drs-rule-create <name> --rule <name> --type affinity|antiAffinity --vm <vm1> --vm <vm2> [--disabled] [--dry-run]\nvmware-aiops cluster drs-rule-delete <name> --rule <name> [--dry-run]                 # VM-VM only; double confirm\n\n# Alarm Management\nvmware-aiops alarm list [--target <name>]\nvmware-aiops alarm acknowledge <entity_name> <alarm_name> [--target <name>]\nvmware-aiops alarm reset <entity_name> <alarm_name> [--target <name>]\n# NOTE: 'alarm reset' clears ALL triggered alarms matching the named alarm's\n# entity type (host/VM/all) and current status (red/yellow) — vSphere has no\n# per-alarm clear API. The CLI double confirmation applies; output reports the\n# scope. The reset_vcenter_alarm MCP tool has no confirmation — it clears on the\n# first call.\n\n# Datastore\nvmware-aiops datastore browse <ds-name> [--path <subdir>]\nvmware-aiops datastore scan-images [--target <name>]\n\n# Scanning & Daemon\nvmware-aiops scan now [--target <name>]   # alarms + events only; host logs are read by the daemon\nvmware-aiops daemon start\nvmware-aiops daemon stop\nvmware-aiops daemon status\n\n# Moved to companion skills:\n# vmware-monitor inventory vms/hosts/datastores/clusters, health alarms/events, vm info\n# vmware-storage iscsi-enable/status/add-target/remove-target, rescan, vsan health/capacity\n# vmware-vks list-namespaces, create-tkc, scale-tkc, etc.\n```\n\nFile v1.12.0:references/investigation-protocol.md\n\n# Investigation Protocol — Causal Chain Root Cause Analysis\n\nA protocol for AI agents performing diagnostic investigations on VMware infrastructure (alarms, performance regressions, availability incidents). Adopted from Enterprise Harness Engineering, drawing on 5 Whys, Google SRE, ITIL, and NASA Fault Tree Analysis.\n\n## When to Apply\n\nUse this protocol whenever the user asks:\n\n- \"Why is X slow / failing / down?\"\n- \"What caused this alarm / alert / incident?\"\n- \"Investigate / diagnose / debug …\"\n- Any open-ended question that requires identifying a root cause rather than just reading state.\n\nDo NOT apply for:\n\n- Simple state lookups (\"is the VM on?\", \"list datastores\")\n- Operational requests (\"clone this VM\", \"create this rule\")\n- Configuration questions (\"what's the default for X?\")\n\n## The Four Criteria for Root Cause Completeness\n\nA diagnostic conclusion is **incomplete** unless ALL four criteria are satisfied. The agent must self-check against each one before outputting a report.\n\n### 1. Falsifiability (可证伪性)\n\nThe root cause must be independently measurable and verifiable. If you cannot test it, it is a hypothesis, not a root cause.\n\n- ✅ \"Datastore latency exceeded 50ms because IOPS hit the SAN cap of 10,000\" — directly testable via `get_metrics datastore.iops`\n- ❌ \"Network was congested\" — too vague to verify\n\n### 2. Sufficiency (充分性)\n\nRemoving the root cause must make the symptom disappear. If the symptom persists after the supposed fix, the cause was wrong or partial.\n\n- ✅ \"Deleting the orphaned snapshot freed 200 GB and the alarm cleared within 60 seconds\"\n- ❌ \"Restarted the VM and the issue went away\" — correlation, not causation\n\n### 3. Necessity (必要性)\n\nThe symptom must occur whenever the root cause is present. If the same condition exists elsewhere without the symptom, you have not found the true root cause.\n\n- ✅ \"Every cluster with 80%+ memory overcommit shows the same vMotion stall\"\n- ❌ \"Only this one VM has the issue\" — without explaining why this VM specifically\n\n### 4. Mechanism (机制性)\n\nYou must explain the propagation chain: root cause → propagation → amplification → impact. A single point claim with no mechanism is a guess.\n\n- ✅ \"Snapshot delta files filled the datastore (root) → VM I/O blocked on write (propagation) → guest filesystem went read-only (amplification) → application timeout (impact)\"\n- ❌ \"The datastore was full\" — describes a state, not a chain\n\n## Investigation Workflow — Up to Three Depth Rounds\n\n### Round 1 — Initial Hypothesis\n\n1. Gather symptoms via L1/L2 read tools (alarms, metrics, events, logs)\n2. Form an initial causal chain hypothesis\n3. Apply the four criteria\n\nIf all four pass → output report.\nIf any criterion fails → proceed to Round 2 with that criterion as the focus.\n\n### Round 2 — Targeted Deepening\n\n1. Identify which criterion failed\n2. Gather additional evidence aimed specifically at that criterion (e.g. failed Necessity → compare against unaffected peers; failed Mechanism → trace next propagation step)\n3. Refine the causal chain\n4. Re-apply the four criteria\n\nIf all four pass → output report.\nIf any still fails → proceed to Round 3.\n\n### Round 3 — Final Deepen or Escalate\n\n1. If a deeper cause is reachable, gather final evidence and finalize the chain\n2. If evidence is unavailable, system-bounded, or beyond the agent's tool surface, **escalate to a human** and explicitly label the conclusion as `⚠️ INCOMPLETE — <criterion> unsatisfied`\n3. **Never** silently output a partial conclusion as if it were complete\n\n## Output Format\n\nEvery investigation report must structure findings exactly as:\n\n```\n🔴 [ROOT CAUSE]   <falsifiable, mechanism-explained statement>\n  → [PROPAGATION] <how the root cause spread to neighboring systems>\n    → [AMPLIFICATION] <what made the impact worse, if applicable>\n      → [IMPACT]    <observable user / business / SLA effect>\n\n✅ Falsifiability:  <evidence — metric name, log query, command output>\n✅ Sufficiency:     <evidence or stated counterfactual>\n✅ Necessity:       <evidence or peer comparison>\n✅ Mechanism:       <see propagation chain above>\n```\n\nIf any criterion is unmet, mark it `⚠️ INCOMPLETE — <reason>` and state explicitly what additional evidence would be required to satisfy it.\n\n## Anti-Patterns\n\n| ❌ Pattern | Why it fails |\n|---|---|\n| \"thanos-cn unreachable\" alone | Describes symptom; does not answer **why** unreachable |\n| \"Datastore full\" alone | No propagation, no impact chain |\n| \"Try restarting it\" | Skips diagnosis entirely |\n| \"Probably the network\" | Not falsifiable |\n| Stopping at the first plausible cause | Skips Necessity check |\n| Silent partial conclusion | Hides incompleteness from the user |\n\n## Worked Examples\n\n### Bad — Incomplete Diagnosis\n\n> \"VM is slow because the host is busy.\"\n\nMissing:\n- **Falsifiability**: which metric, what threshold?\n- **Necessity**: why this VM only?\n- **Mechanism**: how does host load translate into VM slowness?\n\n### Good — Complete Diagnosis\n\n> 🔴 [ROOT] Host `esx-03` CPU ready time exceeds 15% (validated via `get_metrics host.cpu.ready`)\n>   → [PROPAGATION] vCPU contention from 4-VM reservation collision in resource pool `prod-rp`\n>     → [AMPLIFICATION] DRS is in manual mode, so VMs are not rebalanced\n>       → [IMPACT] Application p99 latency doubled from 200 ms to 400 ms\n>\n> ✅ Falsifiability: `host.cpu.ready` metric directly observable; threshold defined in vSphere docs\n> ✅ Sufficiency: vMotion `vm-A` off `esx-03` reduced ready time to 3% and p99 latency back to 200 ms\n> ✅ Necessity: only VMs in `prod-rp` with active reservations are affected; identical workloads in `staging-rp` are healthy\n> ✅ Mechanism: cpu.ready = vCPU waiting for pCPU → guest perceives as CPU starvation → app threadpool exhaustion → tail latency\n\n## Related Skills\n\nA complete investigation often chains across skills:\n\n- **vmware-aiops** (this skill): VM/host state, deployment history; can also remediate at L3+ once the investigation is complete and approved\n- [vmware-aria](https://github.com/vmware-skills/VMware-Aria): metrics, alerts, anomaly detection — primary L1/L2 data source for time-series analysis\n- [vmware-monitor](https://github.com/vmware-skills/VMware-Monitor): inventory, alarms, events — code-level read-only data source\n- [vmware-pilot](https://github.com/vmware-skills/VMware-Pilot): orchestrate the investigation itself as a multi-step Dispatcher → Subagent workflow\n\nThe agent should treat investigation as **read-heavy first**: gather across skills, reason centrally, only invoke L3+ write tools after the four criteria are satisfied AND the user has approved a remediation plan.\n\nFile v1.12.0:references/setup-guide.md\n\n# Setup Guide\n\n## Installation\n\nAll install methods fetch from the same source: [github.com/vmware-skills/VMware-AIops](https://github.com/vmware-skills/VMware-AIops) (MIT licensed). We recommend reviewing the source code before installing.\n\n```bash\n# Via PyPI (recommended for version pinning)\nuv tool install vmware-aiops==1.12.0\n\n# Via Skills.sh (fetches from GitHub)\nnpx skills add vmware-skills/VMware-AIops#v1.12.0\n\n# Via ClawHub (fetches from ClawHub registry snapshot of GitHub)\nclawhub install @zw008/vmware-aiops --version 1.12.0\n```\n\n### Claude Code\n\n`npx skills add` and `clawhub install` both place the skill in Claude Code's skills\ndirectory. To install it manually from a clone:\n\n```bash\nmkdir -p ~/.claude/skills/vmware-aiops\ncp -r skills/vmware-aiops/. ~/.claude/skills/vmware-aiops/\n```\n\nFor tool access (not just skill context), register the MCP server:\n\n```bash\nclaude mcp add vmware-aiops -- vmware-aiops mcp\n```\n\n## Configuration\n\n```bash\n# 1. Install from PyPI (source: github.com/vmware-skills/VMware-AIops)\nuv tool install vmware-aiops==1.12.0\n\n# 2. Verify installation source\nvmware-aiops --version  # confirms installed version\n\n# 3. Configure\nmkdir -p ~/.vmware-aiops\nvmware-aiops init  # generates config.yaml and .env templates\nchmod 600 ~/.vmware-aiops/.env\n# Edit ~/.vmware-aiops/config.yaml and .env with your target details\n```\n\n### Declare `environment:` on each target\n\n```yaml\ntargets:\n  - name: prod-vcenter\n    host: vcenter-prod.example.com\n    environment: production   # production | staging | lab | <your own label>\n```\n\n`environment:` is an optional free-form label. Policy scopes its rules by this\nvalue, so an environment-scoped `deny` rule in `~/.vmware/rules.yaml` can match\non it — for example, to freeze state-changing writes on `production`. A target\nwith no label is simply not matched by such a rule. A rule refuses only what its\n`operations` / `min_risk_level` filters match, so scope it to write operations if\nreads should keep working. Rules are evaluated before every MCP tool call and\nevery CLI command that reaches vCenter, and `operations` are MCP tool names — CLI commands are\nauthorised under the same names, so one rule covers both. The shipped baseline\ndenies nothing.\n\n## What Gets Installed\n\nThe `vmware-aiops` package installs a Python CLI binary and its dependencies (pyVmomi, Click, Rich, APScheduler, python-dotenv). No background services, daemons, or system-level changes are made during installation. The scheduled scanner (`daemon start`) only runs when explicitly started by the user.\n\n## Development Install\n\n```bash\ngit clone --branch v1.12.0 https://github.com/vmware-skills/VMware-AIops.git\ncd VMware-AIops\nuv venv && source .venv/bin/activate\n# --no-sources: pyproject's [tool.uv.sources] points vmware-monitor at a sibling\n# checkout (../VMware-Monitor) for family development. A fresh clone has none, so\n# plain `uv pip install -e .` fails with \"Distribution not found\"; this takes\n# vmware-monitor (>=1.11.3) from PyPI instead.\nuv pip install --no-sources -e .\n```\n\nTo develop against an unreleased vmware-monitor, clone\n[VMware-Monitor](https://github.com/vmware-skills/VMware-Monitor) next to this\ncheckout (so `../VMware-Monitor` exists) and drop `--no-sources`. Plain `pip`\nignores `[tool.uv.sources]` and needs neither.\n\n### Password obfuscation at rest\n\nOn first load, any plaintext `*_PASSWORD` value in `.env` is automatically\nrewritten to a grep-safe `b64:<encoded>` form and decoded transparently at\nruntime, so a casual `grep` of the file no longer reveals the password. Values\nare read and written through python-dotenv's own parser, so the stored secret\nnever drifts from what you configured (quotes, inline comments, and trailing\nwhitespace are handled correctly).\n\n> **This is obfuscation, not encryption.** Anyone who can read the file can\n> still decode it. For real secrecy at rest, do not store the password in `.env`\n> at all — inject it from a secret manager (HashiCorp Vault, CyberArk, AWS\n> Secrets Manager, or a Kubernetes Secret) into the `*_PASSWORD` environment\n> variable at process start. The code reads the env var either way.\n\n## Security\n\n> **Disclaimer**: This is a community-maintained open-source project and is **not affiliated with, endorsed by, or sponsored by VMware, Inc. or Broadcom Inc.** \"VMware\" and \"vSphere\" are trademarks of Broadcom.\n\n- **Source Code**: Fully open source at [github.com/vmware-skills/VMware-AIops](https://github.com/vmware-skills/VMware-AIops) (MIT). The `uv` installer fetches the `vmware-aiops` package from PyPI, which is built from this GitHub repository. We recommend reviewing the source code and commit history before deploying in production.\n- **TLS Verification**: Enabled by default. Setting `verify_ssl: false` is solely for ESXi hosts using self-signed certificates in isolated lab/home environments. In production, always use CA-signed certificates with full TLS verification.\n- **Credentials & Config**: This skill requires the following secrets, all stored in `~/.vmware-aiops/.env` (`chmod 600`, loaded via `python-dotenv`):\n  - `VMWARE_<TARGET>_PASSWORD` — per-target password where `<TARGET>` is the uppercased target name from `config.yaml` (hyphens become underscores). Example: target named `vcenter-prod` uses `VMWARE_VCENTER_PROD_PASSWORD`.\n  - (Optional) Webhook URLs for Slack/Discord notifications\n\n  The config file `~/.vmware-aiops/config.yaml` stores only target hostnames, ports, and usernames — it does **not** contain passwords or tokens. The env var `VMWARE_AIOPS_CONFIG` points to this YAML file.\n- **Webhook Data Scope**: Webhook notifications are **disabled by default**. When enabled, the daemon posts to **user-configured URLs only** (Slack, Discord, or any HTTP endpoint you control); no data is sent to any other service. Each payload carries critical/warning counts plus every critical issue and every alarm/event warning from that scan — host-log warnings go to `scan.log` only, and `info` rows (unreadable or partly-read host logs) are never sent. Each issue carries the entity name and one of: the alarm name, vCenter event message (sanitized, ≤500 chars), ESXi log line matching critical/panic/corrupt (sanitized, ≤200 chars), or the error text for a target the daemon could not connect to. Event, log, and error text can contain host names, IP addresses, and user names — treat the webhook destination as receiving operational data. No credentials from the skill's config or `.env` are included.\n- **Daemon host-log reads**: the scanner daemon reads the ESXi `hostd`, `vmkernel` and `vpxa` logs, which needs the `Global.Diagnostics` privilege — vCenter's built-in Read-Only role does not include it. Without it each log is recorded in `scan.log` as an `info` row with the reason instead of being scanned; grant it only if you want host-log scanning.\n- **Prompt Injection Protection**: All vSphere-sourced content (event messages, host logs) is truncated, stripped of control characters, and wrapped in boundary markers (`[VSPHERE_EVENT]`/`[VSPHERE_HOST_LOG]`) before output to prevent prompt injection when consumed by LLM agents.\n- **Least Privilege**: 22 of the 43 MCP write tools — every destructive one — return a no-write blast-radius preview unless called with `confirm=True`, which is refused on a blocker or an unreadable measurement; `vm_delete` also requires the preview's acknowledgement echoed back. The other 21 (create, clone, deploy, power-on, reconfigure) act on the first call. A preview is not authorization. The enforcement boundary is the RBAC of the vCenter/ESXi account in `.env`, so use a dedicated service account scoped to what the agent may change. For monitoring-only use cases, prefer the read-only [VMware-Monitor](https://github.com/vmware-skills/VMware-Monitor) skill which has zero destructive code paths. The CLI's double confirmation and `--dry-run` do not apply to MCP calls.\n- **Guest Credentials**: Guest operations run with whatever guest account is passed to them — the `username` is required (there is no default account) and over MCP the password is a tool argument the agent sees (the audit row redacts it). A read-only vCenter role does not limit what they do inside a VM. Pass a least-privilege guest account; avoid root unless the task needs it. `vm_guest_upload` reads any local file the server process can read.\n- **Policy & Audit**: Optional `deny` rules in `~/.vmware/rules.yaml` refuse matching operations before every MCP call and every CLI command that reaches vCenter (see `environment:` above); they run in-process and are a guardrail, not a substitute for RBAC. Every such call is recorded in `~/.vmware/audit.db` with credentials redacted (best-effort: an audit write failure warns and does not block).\n\nTo run the agent read-only, give it a read-only vCenter/ESXi service account (RBAC) — enforced at the platform.\n\n## Supported AI Platforms\n\n| Platform | Status | Config File |\n|----------|--------|-------------|\n| Claude Code | ✅ Native Skill | `skills/vmware-aiops/SKILL.md` |\n| Gemini CLI | ✅ Context file + MCP | `skills/vmware-aiops/SKILL.md` |\n| OpenAI Codex CLI | ✅ Skill + AGENTS.md | `skills/vmware-aiops/SKILL.md` |\n| Aider | ✅ Conventions | `skills/vmware-aiops/SKILL.md` |\n| Continue CLI | ✅ Rules | `skills/vmware-aiops/SKILL.md` |\n| Trae IDE | ✅ Rules | `skills/vmware-aiops/SKILL.md` |\n| Kimi Code CLI | ✅ Skill | `skills/vmware-aiops/SKILL.md` |\n| MCP Server | ✅ MCP Protocol | `vmware_aiops/mcp_server/` |\n| Python CLI | ✅ Standalone | N/A |\n\n## MCP Server — Local Agent Compatibility\n\nThe MCP server works with any MCP-compatible agent via stdio transport. Config templates in `examples/mcp-configs/`:\n\n| Agent | Local Models | Config Template |\n|-------|:----------:|-----------------|\n| Goose (Block) | ✅ Ollama, LM Studio | `goose.json` |\n| LocalCowork (Liquid AI) | ✅ Fully offline | `localcowork.json` |\n| mcp-agent (LastMile AI) | ✅ Ollama, vLLM | `mcp-agent.yaml` |\n| VS Code Copilot | — | `vscode-copilot.json` |\n| Cursor | — | `cursor.json` |\n| Continue | ✅ Ollama | `continue.yaml` |\n| Claude Code | — | `claude-code.json` |\n\n```bash\n# Example: Aider + Ollama (fully local, no cloud API)\naider --conventions skills/vmware-aiops/SKILL.md --model ollama/qwen2.5-coder:32b\n```\n\n## MCP Mode (Optional)\n\nFor Claude Code / Cursor users who prefer structured tool calls, add to `~/.claude/settings.json`:\n\n```json\n{\n  \"mcpServers\": {\n    \"vmware-aiops\": {\n      \"command\": \"vmware-aiops\",\n      \"args\": [\"mcp\"],\n      \"env\": {\n        \"VMWARE_AIOPS_CONFIG\": \"~/.vmware-aiops/config.yaml\"\n      }\n    }\n  }\n}\n```\n\n> v1.5.15+ recommends the single-command form `vmware-aiops mcp`. Pre-1.5.15 used\n> `uvx --from vmware-aiops vmware-aiops-mcp`, which still works but re-resolves from <!-- install-pin: historical -->\n> PyPI on each launch and breaks behind corporate TLS proxies. The legacy\n> `vmware-aiops-mcp` entry point is also kept for backward compatibility.\n\nMCP exposes 60 tools across 10 categories. All accept optional `target` parameter.\n\nFile v1.12.0:skill-card.md\n\n## Description:\n\nUse this skill to manage VMware, vSphere, and ESXi VM lifecycle operations, deployments, guest operations, clusters, alarms, and investigation workflows.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[zw008](https://clawhub.ai/user/zw008)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers, operators, and infrastructure engineers use this skill to operate VMware environments from an agent, including VM lifecycle work, deployment, cluster changes, guest commands, alarm handling, and triage before remediation.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: The skill can perform infrastructure-changing VMware operations, and some write actions can run without a built-in confirmation step.\n\nMitigation: Install with a dedicated least-privilege VMware service account, add deny rules before production use, and require operator review for writes such as reconfigure, clone, deploy, cluster changes, alarm reset, and plan apply.\n\nRisk: Guest execution and upload operations can run commands or copy files with the supplied guest credentials.\n\nMitigation: Use least-privilege guest accounts, avoid root or administrator credentials unless necessary, and review the exact command, file path, and target VM before approving execution.\n\nRisk: Destructive or broad-scope actions such as power-off, delete, migration, snapshot changes, TTL deletion, and alarm reset can affect availability or clear more than one object.\n\nMitigation: Inspect preview or blast-radius output when available, confirm target names and scopes, and re-check final state after execution.\n\nRisk: Webhook notifications and diagnostics may expose operational metadata such as host names, IP addresses, user names, alarm text, event messages, or connection errors.\n\nMitigation: Enable webhooks only to trusted destinations and treat notifications as operational data.\n\nRisk: Disabling TLS verification can expose connections to interception outside isolated lab environments.\n\nMitigation: Keep TLS verification enabled in production and use CA-signed certificates for vCenter or ESXi endpoints.\n\n## Reference(s):\n\n- [ClawHub Skill Page](https://clawhub.ai/zw008/skills/vmware-aiops)\n- [Project Homepage](https://github.com/vmware-skills/VMware-AIops)\n- [Capabilities Reference](references/capabilities.md)\n- [Setup Guide](references/setup-guide.md)\n- [CLI Reference](references/cli-reference.md)\n- [Agent Guardrails](references/agent-guardrails.md)\n- [Investigation Protocol](references/investigation-protocol.md)\n\n## Skill Output:\n\n**Output Type(s):** [text, markdown, shell commands, configuration, guidance]\n\n**Output Format:** [Markdown guidance with inline shell commands and configuration examples]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [May direct an agent to call VMware CLI or MCP tools that read or change infrastructure state.]\n\n## Skill Version(s):\n\n1.12.0 (source: server release evidence)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nFile v1.12.0:evals/evals.json\n\n{\n  \"skill_name\": \"vmware-aiops\",\n  \"evals\": [\n    {\n      \"id\": 1,\n      \"prompt\": \"Deploy a new VM from the ubuntu-22.04.ova file on datastore1, name it web-test-01, power it on, and set it to auto-delete in 8 hours\",\n      \"expected_output\": \"VM deployed from OVA, powered on, TTL set to 8 hours\",\n      \"files\": [],\n      \"expectations\": [\n        \"Uses deploy_vm_from_ova tool with correct OVA path and VM name\",\n        \"Powers on the VM after deployment\",\n        \"Sets TTL to 480 minutes using vm_set_ttl\"\n      ]\n    },\n    {\n      \"id\": 2,\n      \"prompt\": \"Clone prod-db-01 five times for load testing, name them load-test-01 through 05, give each 4 CPUs and 8GB RAM\",\n      \"expected_output\": \"5 clones created with specified resources\",\n      \"files\": [],\n      \"expectations\": [\n        \"Uses batch_clone_vms or creates 5 individual clones\",\n        \"Each clone has cpu=4 and memory_mb=8192\",\n        \"VM names follow the load-test-01 through load-test-05 pattern\"\n      ]\n    },\n    {\n      \"id\": 3,\n      \"prompt\": \"There's a critical alarm on esxi-host-03, acknowledge it and then check if any VMs on that host need attention\",\n      \"expected_output\": \"Alarm acknowledged, VM status checked\",\n      \"files\": [],\n      \"expectations\": [\n        \"Uses list_vcenter_alarms to find the alarm details\",\n        \"Uses acknowledge_vcenter_alarm with correct entity and alarm name\",\n        \"Routes to vmware-monitor for VM health check or uses get_alarms\"\n      ]\n    },\n    {\n      \"id\": 4,\n      \"prompt\": \"Create a plan to safely restart the database cluster: power off db-replica first, then db-primary, wait, power on db-primary, then db-replica\",\n      \"expected_output\": \"Execution plan created with correct order and rollback\",\n      \"files\": [],\n      \"expectations\": [\n        \"Uses vm_create_plan with sequential power_off and power_on operations\",\n        \"Order is correct: replica off → primary off → primary on → replica on\",\n        \"Shows plan summary and waits for user approval before vm_apply_plan\"\n      ]\n    }\n  ]\n}\n\nArchive v1.11.0: 9 files, 42068 bytes\n\nFiles: evals/evals.json (2055b), references/agent-guardrails.md (9940b), references/capabilities.md (34420b), references/cli-reference.md (4671b), references/investigation-protocol.md (6771b), references/setup-guide.md (11082b), skill-card.md (2833b), SKILL.md (26823b), _meta.json (132b)\n\nFile v1.11.0:SKILL.md\n\n---\nname: vmware-aiops\ndescription: >\n  Use this skill whenever the user needs to manage VMs in VMware/vSphere/ESXi — it's the entry point for all VM operations.\n  Directly handles: power on/off, clone, snapshot, migrate, deploy from OVA or templates, run commands inside VMs, batch operations, cluster management, vCenter alarm acknowledgment, a one-glance cluster-health triage (\"is anything on fire?\"), and VM/host/datastore investigation drill-downs.\n  Always use this skill for any \"power on\", \"clone\", \"deploy\", \"migrate\", \"batch\", \"guest exec\", \"alarm\", or VM lifecycle task, and for triage like \"is anything on fire\" / \"what needs attention now\" / \"investigate this VM\", when the context is explicitly VMware, vSphere, or ESXi.\n  Do NOT use for general read-only queries (inventory/events/VM details — use vmware-monitor), NSX networking (use vmware-nsx), storage/iSCSI/vSAN (use vmware-storage), or Kubernetes cluster lifecycle (use vmware-vks).\n  For multi-step workflows use vmware-pilot. For load balancing/AVI/AKO use vmware-avi.\ninstaller:\n  kind: uv\n  package: vmware-aiops\nargument-hint: \"[vm-name or describe your task]\"\nallowed-tools:\n  - Bash\nmetadata: {\"openclaw\":{\"requires\":{\"anyBins\":[\"vmware-aiops\",\"uvx\"]},\"optional\":{\"env\":[\"VMWARE_AIOPS_CONFIG\",\"VMWARE_TARGET_PASSWORD\",\"VMWARE_<TARGET>_USERNAME\",\"SLACK_WEBHOOK_URL\",\"DISCORD_WEBHOOK_URL\",\"VMWARE_AUDIT_APPROVED_BY\"],\"bins\":[\"vmware-policy\"]},\"homepage\":\"https://github.com/vmware-skills/VMware-AIops\",\"emoji\":\"🖥️\",\"os\":[\"macos\",\"linux\"]}}\ncompatibility: >\n  vmware-policy auto-installed as Python dependency (provides @vmware_tool decorator and audit logging). All write operations audited to ~/.vmware/audit.db.\n  Credentials: Each vCenter/ESXi target requires a per-target password env var in ~/.vmware-aiops/.env following the pattern VMWARE_<TARGET_NAME_UPPER>_PASSWORD. Passwords are never logged or echoed.\n  Destructive operations: All write tools require explicit parameters and pass through the @vmware_tool decorator (policy check + audit + sanitize). Every MCP write tool annotated destructive (22 of 43: power-off, delete, migrate, snapshot revert/delete, guest exec/upload/provision, cluster delete/remove-host, TTL, Clean Slate, plan apply/rollback, host-network and DRS) takes one confirm argument whose default returns a no-write blast-radius preview (HLD section 7); confirm=True is refused, and audited as a failure, on a blocker or an unreadable measurement, and vm_delete also requires the preview's acknowledgement echoed back and still matching; the other 21 write tools (create, clone, deploy, power-on, reconfigure) act on the first call. The enforcement boundary is the RBAC of the vCenter/ESXi account the server connects with, so run it under a dedicated least-privilege service account (a read-only role makes it read-only). Optional deny rules in ~/.vmware/rules.yaml are checked before every MCP call and remote CLI command; the shipped baseline denies nothing. CLI destructive commands additionally require double confirmation and most CLI writes support --dry-run; neither applies to MCP calls.\n  Guest operations: vm_name and command are required; no implicit or background execution. The command is unbounded and runs with the guest credentials supplied — the username is required on MCP and CLI alike (no root default) and over MCP the password is a tool argument the agent sees (redacted from the audit row). The guest account is a second authorization boundary that a read-only vCenter role does not limit; pass a least-privilege guest account. vm_guest_upload reads any local file the server process can read.\n  Webhooks: Disabled by default. When enabled, the daemon posts to user-configured URLs only: issue counts plus every critical issue and every alarm/event warning (host-log warnings and info rows are not sent), each with its entity name and the sanitized alarm, event, or ESXi log text, or a connection error — which can include host names, IPs, and user names. No credentials from the skill's config are sent. Reading host logs needs the Global.Diagnostics privilege; an unreadable log is recorded, not skipped.\n  TLS verification is on by default (verify_ssl: true); set verify_ssl: false only for self-signed certs in isolated lab environments.\n  Transitive dependencies: Only vmware-policy (audit/policy). No post-install scripts or background services.\n---\n\n# VMware AIops\n\n> **Disclaimer**: This is a community-maintained open-source project and is **not affiliated with, endorsed by, or sponsored by VMware, Inc. or Broadcom Inc.** \"VMware\" and \"vSphere\" are trademarks of Broadcom. Source code is publicly auditable at [github.com/vmware-skills/VMware-AIops](https://github.com/vmware-skills/VMware-AIops) under the MIT license.\n\nVMware family entry point — AI-powered VM lifecycle, deployment, and alarm management — 60 MCP tools.\n\n> **Start here**: install vmware-aiops first, then add modules as needed.\n> Run `vmware-aiops hub status` to see which family members are installed.\n> **Family**: [vmware-monitor](https://github.com/vmware-skills/VMware-Monitor) (inventory/health), [vmware-storage](https://github.com/vmware-skills/VMware-Storage) (iSCSI/vSAN), [vmware-vks](https://github.com/vmware-skills/VMware-VKS) (Tanzu Kubernetes), [vmware-nsx](https://github.com/vmware-skills/VMware-NSX) (NSX networking), [vmware-nsx-security](https://github.com/vmware-skills/VMware-NSX-Security) (DFW/firewall), [vmware-aria](https://github.com/vmware-skills/VMware-Aria) (metrics/alerts/capacity), [vmware-avi](https://github.com/vmware-skills/VMware-AVI) (AVI/ALB/AKO), [vmware-harden](https://github.com/vmware-skills/VMware-Harden) (compliance baselines).\n> | [vmware-pilot](../vmware-pilot/SKILL.md) (workflow orchestration) | [vmware-policy](../vmware-policy/SKILL.md) (audit/policy)\n\n## What This Skill Does\n\n| Category | Tools | Count |\n|----------|-------|:-----:|\n| **VM Lifecycle** | power on/off, create, reconfigure, clone, migrate, delete, snapshot CRUD, TTL auto-delete, clean slate | 16 |\n| **Deployment** | OVA, template, linked clone, batch clone/deploy | 8 |\n| **Guest Ops** | exec commands, upload/download files, provision | 5 |\n| **Plan/Apply** | multi-step planning with rollback | 4 |\n| **Cluster** | create, delete, HA/DRS config, add/remove hosts, DRS VM-VM rules (list/create/delete/enable-disable) | 10 |\n| **Datastore** | browse files, scan for images | 2 |\n| **Network** | dvSwitch portgroup list/create, host VMkernel list/add/remove/tag-service, DF-bit MTU-path ping | 7 |\n| **Alarm Management** | list alarms, acknowledge, reset | 3 |\n| **Triage & Investigation** (read-only, delegates to vmware-monitor) | one-glance cluster health summary, object-centered VM/host/datastore drill-down bundles, cross-vCenter \"what needs attention now?\" | 5 |\n\n## Audit & Safety\n\nRead before connecting an agent. Per-tool inventory: `references/capabilities.md`.\n\n- **MCP gates default to a no-write preview.** 22 write tools — every destructive one — return their blast radius unless `confirm=True`, which is refused on a blocker or unreadable measurement; `vm_delete` also needs `acknowledge_blast_radius` echoing the preview. The other 21 (create, clone, deploy, power-on, reconfigure) act on the first call.\n- **The enforcement boundary is vCenter/ESXi RBAC**: an agent can do whatever the configured account can. Use a dedicated, least-privilege service account scoped to what the agent may change (a read-only role makes the skill read-only). Store its password in `~/.vmware-aiops/.env` (0600) or a secret manager (`VMWARE_<TARGET>_PASSWORD`).\n- **CLI only**: destructive commands require double confirmation; most CLI writes take `--dry-run`. Neither applies to MCP.\n- **Policy**: deny rules and a maintenance window in `~/.vmware/rules.yaml` are checked before every MCP and remote CLI call (e.g. deny writes to `environment: production` targets). The shipped baseline denies nothing. An in-process guardrail, not a substitute for RBAC.\n- **Audit**: every MCP call is recorded in `~/.vmware/audit.db`, credentials redacted (`vmware-audit log --last 20`). Best-effort: a failed audit write warns, never blocks.\n- **Guest ops** run any command or file write the guest account allows — a read-only vCenter role does not limit this. `username` is required (no default account) and over MCP the password is a tool argument the agent sees; pass a minimal guest account. `vm_guest_upload` reads any local file the server can read.\n\n## Quick Install\n\n```bash\nuv tool install vmware-aiops==1.11.0\nvmware-aiops doctor\nvmware-aiops hub status   # see which family members are installed\n```\n\n## VMware Family — Install What You Need\n\nvmware-aiops is the entry point. Add modules for additional capabilities:\n\n| Module | Install | Adds |\n|--------|---------|------|\n| **vmware-monitor** | `uv tool install vmware-monitor` | Read-only inventory, alarms, events |\n| **vmware-storage** | `uv tool install vmware-storage` | iSCSI, vSAN, datastore management |\n| **vmware-vks** | `uv tool install vmware-vks` | Tanzu Kubernetes (vSphere 8.x+) |\n| **vmware-nsx** | `uv tool install vmware-nsx-mgmt` | NSX networking: segments, gateways, NAT |\n| **vmware-nsx-security** | `uv tool install vmware-nsx-security` | DFW microsegmentation, security groups |\n| **vmware-aria** | `uv tool install vmware-aria` | Aria Ops metrics, alerts, capacity |\n| **vmware-avi** | `uv tool install vmware-avi` | AVI load balancer, ALB, AKO, Ingress |\n\n> Each module stays independent — small tool count keeps local models (Ollama, Qwen) accurate.\n\n## When to Use This Skill\n\n- Power on/off, create, delete, snapshot, clone, or migrate VMs\n- Deploy VMs from OVA, templates, linked clones, or batch specs\n- Run commands or transfer files inside a VM (Guest Operations)\n- Create/configure clusters (HA/DRS)\n- Browse datastores for deployable images\n- Plan and execute multi-step operations with rollback\n- List, acknowledge, and clear vCenter triggered alarms (clear matches by entity type + status — see MCP Tools section)\n\n**Use companion skills for**:\n- Inventory, health, alarms, VM info → `vmware-monitor`\n- iSCSI, vSAN, datastore management → `vmware-storage`\n- Tanzu Kubernetes (Supervisor, Namespace, TKC) → `vmware-vks`\n- Load balancing, AVI/ALB, AKO, Ingress → `vmware-avi`\n\n## Related Skills — Skill Routing\n\n| User Intent | Recommended Skill |\n|-------------|------------------|\n| Read-only monitoring, zero risk | **vmware-monitor** (`uv tool install vmware-monitor`) |\n| Storage: iSCSI, vSAN, datastores | **vmware-storage** (`uv tool install vmware-storage`) |\n| VM lifecycle, deployment, guest ops | **vmware-aiops** ← this skill |\n| Tanzu Kubernetes (vSphere 8.x+) | **vmware-vks** (`uv tool install vmware-vks`) |\n| NSX networking: segments, gateways, NAT | **vmware-nsx** (`uv tool install vmware-nsx-mgmt`) |\n| NSX security: DFW rules, security groups | **vmware-nsx-security** (`uv tool install vmware-nsx-security`) |\n| Aria Ops: metrics, alerts, capacity | **vmware-aria** (`uv tool install vmware-aria`) |\n| Multi-step workflows with approval | **vmware-pilot** |\n| Compliance baselines (CIS / 等保 / PCI-DSS), drift detection, LLM remediation advisor | **vmware-harden** (`uv tool install vmware-harden`) |\n| Load balancer, AVI, ALB, AKO, Ingress | **vmware-avi** (`uv tool install vmware-avi`) |\n| Audit log query | **vmware-policy** (`vmware-audit` CLI) |\n\n## Common Workflows\n\n> **Diagnostic investigations**: Before remediating any \"why is X slow / failing / down\" issue, follow [`references/investigation-protocol.md`](references/investigation-protocol.md). It enforces the four root-cause completeness criteria (falsifiability / sufficiency / necessity / mechanism) and the up-to-three-rounds deepening loop. Only invoke L3+ write tools after the four criteria are satisfied AND the user has approved a remediation plan.\n\n### Cluster Health Triage (\"what's wrong right now?\")\n\nStart here when the ask is \"is anything on fire?\" before diving into a specific VM. This is a read-only rollup delegated to vmware-monitor, exposed here so triage-then-act stays in one conversation.\n\n1. One glance --> `cluster_health_summary` (MCP) or `vmware-aiops summary` (CLI). Read `top_issues` first — ranked anomalies (disconnected hosts, red/yellow alarms, capacity pressure), each with a drill-down hint; the per-cluster table is context\n2. Act on what it surfaces --> a `host_down` row → investigate the host; an alarm → `acknowledge_vcenter_alarm` / `reset_vcenter_alarm`; a hot VM implicated → `vm_migrate` or `vm_reconfigure` (after the investigation protocol)\n3. Save/share a snapshot --> `vmware-aiops summary --html` writes an offline, timestamped HTML file (identical to `vmware-monitor summary --html` — same shared renderer)\n4. **If vmware-monitor is not installed** --> this command/tool is unavailable (AIops delegates to it); install `vmware-monitor`, or use the deeper per-object read tools in that skill\n\n### Object-Centered Investigation → Act (drill-down before you change anything)\n\n**Judgment**: after triage points at a problem object, drill in with one correlated read *before* actuating — the bundle aggregates the object with its surrounding infrastructure and recent history so you (and the operator) see the full picture, not a guess. AIops is the conversational entry point, so triage → investigate → act stays in one conversation.\n\n1. Estate-wide --> `cross_vcenter_attention` (\"what needs attention now?\" across every vCenter). One vCenter configured → skip to `cluster_health_summary`\n2. Drill into the flagged object (offer the level; skip the question when unambiguous):\n   - a VM --> `vm_investigation_bundle` → state, host it runs on, cluster context, backing datastores, snapshots, alarms & recent changes, performance signals, correlated event timeline\n   - a host --> `host_investigation_bundle`; a datastore --> `datastore_investigation_bundle`\n3. **Then** act on the evidence --> e.g. bundle shows a wedged VM on a hot host → `vm_migrate`; a full datastore → `vm_delete_snapshot` on the sprawl the bundle surfaced. Follow the investigation protocol before any destructive action\n4. Widen the window with `hours=72`; render an offline snapshot with `--html` (drill-down sections collapse natively, nothing uploaded)\n5. **If the object name is unknown** --> the bundle returns a teaching error naming how to list objects; get the exact name and retry. **If vmware-monitor is not installed** --> these delegated tools are unavailable\n\n### Deploy a Lab Environment\n\n**Pre-flight (judgment, not blind sequence)**:\n- Free space: target datastore must have ≥ OVA size × 2 (delta files + thin-provision overhead). If multiple datastores qualify, prefer one with lowest current IOPS pressure (cross-check `vmware-aria` if available).\n- Name hygiene: prefix with date or owner (`lab-2026-04-30-alice`) so the TTL cleanup audit trail is meaningful.\n- TTL: always set. 480 min for a single test session, 7200 min for a week-long sandbox. **Never deploy a \"lab\" VM without a TTL** — that is how datastores fill up at 3 AM.\n- Snapshot timing: take the baseline **after** provisioning succeeds, not before — a pre-provision snapshot is just an empty checkpoint.\n\n**Steps**:\n1. `vmware-aiops datastore browse <ds> --pattern \"*.ova\"` → confirm image present and size\n2. `vmware-aiops deploy ova <path> --name <date>-<owner>-<purpose> --datastore <ds>`\n3. `vmware-aiops vm guest-exec <name> --cmd /usr/bin/python3 --args \"setup.py\" --user admin` → if exit ≠ 0, **stop**, do not snapshot a half-provisioned VM\n4. `vmware-aiops vm snapshot-create <name> --name baseline` (only if multi-iteration testing; skip for one-shot)\n5. `vmware-aiops vm set-ttl <name> --minutes 480`\n\n### Batch Clone for Testing\n\n**Pre-flight**:\n- Source VM state: powered-off is safest. If powered-on, VMware Tools must be running and quiesce-capable, else clones may have inconsistent disk state.\n- Capacity math: `free_space ≥ source.size × count × 1.2` (full clone) or `≥ count × 2 GB` (linked clone, delta-only).\n- Decision rule: **count > 10 → use linked clones** (`deploy linked-clone`); seconds vs minutes per clone, ~100× less storage. Tradeoff: linked clones depend on source snapshot — deleting the snapshot breaks all children.\n- Network exhaustion: each clone gets a unique MAC from the vSphere pool; if you batch > 200, verify pool capacity in advance.\n- TTL: every clone must have one. Use the plan's metadata to track ownership.\n\n**Steps**:\n1. `vm_create_plan` with clone + reconfigure + set-ttl steps grouped per VM (atomic per clone)\n2. Review the plan with the user — surface count, datastore, irreversible warnings\n3. `vm_apply_plan` — stops on first failure (intentional, do not auto-resume)\n4. On failure: `vm_rollback_plan` → reverses completed clones; manually verify rollback before retrying\n\n### Migrate VM to Another Host\n\n**Pre-flight (ALL must pass before issuing migrate)**:\n- CPU compatibility: target host CPU family must match source, OR cluster must be in EVC mode. Live migration across mismatched CPUs **fails mid-flight** and may leave the VM stunned.\n- Network parity: every portgroup the VM uses must exist on the target host's vSwitch with the same VLAN. Missing portgroup → vNICs disconnected post-migration.\n- Storage visibility: target host must see all of the VM's datastores; otherwise this is a Storage vMotion, not a host migration — different (slower) operation.\n- Affinity rules: if the VM is pinned to source by a DRS host-affinity rule, migration silently violates intent. Check `cluster info` first.\n- Hardware passthrough: VMs with PCI passthrough (GPU, USB) **cannot live-migrate** — schedule a cold migration window.\n\n**Steps**:\n1. Verify VM state and current host via `vmware-monitor vm info <name>`\n2. Verify target host: same cluster, EVC compatible, has required networks/datastores\n3. `vmware-aiops vm migrate <name> --to-host <target>` — wait for task completion, do not assume success on return\n4. Post-check: `vm info` confirms new host AND power state unchanged AND vNICs connected\n\n## Usage Mode\n\n| Scenario | Recommended | Why |\n|----------|:-----------:|-----|\n| Local/small models (Ollama, Qwen) | **CLI** | ~2K tokens vs ~8K for MCP |\n| Cloud models (Claude, GPT-4o) | Either | MCP gives structured JSON I/O |\n| Automated pipelines | **MCP** | Type-safe parameters, structured output |\n\n## MCP Tools (60 — 17 read, 43 write)\n\n| Category | Tools | R/W |\n|----------|-------|:---:|\n| VM Lifecycle (16) | `vm_list_ttl`, `vm_list_snapshots`, `vm_task_status` | Read |\n| | `vm_power_on`, `vm_power_off`, `vm_create`, `vm_reconfigure`, `vm_clone`, `vm_migrate`, `vm_delete`, `vm_create_snapshot`, `vm_revert_snapshot`, `vm_delete_snapshot`, `vm_set_ttl`, `vm_cancel_ttl`, `vm_clean_slate` | Write |\n| Deployment (8) | `deploy_vm_from_ova`, `deploy_vm_from_template`, `deploy_linked_clone`, `attach_iso_to_vm`, `convert_vm_to_template`, `batch_clone_vms`, `batch_linked_clone_vms`, `batch_deploy_from_spec` | Write |\n| Guest Ops (5) | `vm_guest_exec`, `vm_guest_exec_output`, `vm_guest_upload`, `vm_guest_download`, `vm_guest_provision` | Write |\n| Plan/Apply (4) | `vm_list_plans` | Read |\n| | `vm_create_plan`, `vm_apply_plan`, `vm_rollback_plan` | Write |\n| Datastore (2) | `browse_datastore`, `scan_datastore_images` | Read |\n| Network (7) | `list_dvs_portgroups`, `list_host_vmks`, `vmk_ping` | Read |\n| | `create_dvs_portgroup`, `add_host_vmk`, `remove_host_vmk`, `set_vmk_service` | Write |\n| Cluster (10) | `cluster_info`, `list_drs_rules` | Read |\n| | `cluster_create`, `cluster_delete`, `cluster_add_host`, `cluster_remove_host`, `cluster_configure`, `set_drs_rule_enabled`, `create_drs_rule`, `delete_drs_rule` | Write |\n| Alarm Management (3) | `list_vcenter_alarms` | Read |\n| | `acknowledge_vcenter_alarm`, `reset_vcenter_alarm` | Write |\n| Cluster Triage (1) | `cluster_health_summary` (delegates to vmware-monitor) | Read |\n| Object Investigation (4) | `vm_investigation_bundle`, `host_investigation_bundle`, `datastore_investigation_bundle`, `cross_vcenter_attention` (all delegate to vmware-monitor) | Read |\n\n**List envelope**: the read list tools — `browse_datastore`, `list_vcenter_alarms`, `vm_list_plans`, `vm_list_snapshots`, `vm_list_ttl` — return `{items, returned, limit, total, truncated, hint}` rather than a bare array. Read the rows from `items` and check `truncated` before concluding a listing is complete; empty `items` with `truncated: false` means checked-and-none, not a failure. The write `batch_*` tools keep their bare list (complete by construction). Rationale, `total` semantics, error shape: `references/capabilities.md`.\n\n**Read/write split**: 17 tools are read-only (per `[READ]` docstring marker), 43 modify state — gating in [Audit & Safety](#audit--safety). `vm_set_ttl` schedules an unattended auto-delete.\n\n**Network write gating**: all four network writes are preview/confirm-gated — `confirm=False` (default) returns the exact spec without writing. `remove_host_vmk` is **fail-closed**: it refuses when the vmk is selected for a host service (management/vMotion/vSAN), lives on a non-default netstack (NSX TEPs, dedicated vMotion stacks), carries a default gateway route, or when any of that cannot be verified — pass `force_unprotected=True` to override the non-absolute protections. The host's only management-enabled vmk is never removable (no override). `set_vmk_service` is **fail-closed** too: it refuses both directions when the host's service map is unreadable, and refuses (no override) to untag `management` from the host's only management-enabled vmk — the call rides the interface it would untag.\n\n**DRS rule gating**: `set_drs_rule_enabled`, `create_drs_rule`, `delete_drs_rule` are preview/confirm-gated and idempotent (matching state returns a no-write noop). `create_drs_rule` handles VM-VM affinity/anti-affinity only (≥2 distinct VMs, all cluster members); VM-Host rules are read via `list_drs_rules` but managed in the vSphere UI. `delete_drs_rule` **refuses non-VM-VM rules** (they can carry licensing/compliance placement constraints) and records the full rule definition in both preview and result so a mistaken delete can be recreated from the audit trail.\n\n**Alarm reset blast radius**: vSphere has no per-alarm clear API. `reset_vcenter_alarm` uses `AlarmManager.ClearTriggeredAlarms`, which clears **all** triggered alarms matching the named alarm's entity type (host/VM/all) and current status (red/yellow) — not just the one named. The response's `scope` field states exactly what was cleared. The named alarm is looked up first, so a typo fails fast without clearing anything.\n\n## CLI Quick Reference\n\n```bash\n# VM operations\nvmware-aiops vm power-on <name> [--target <t>]\nvmware-aiops vm power-off <name> [--force]\nvmware-aiops vm create <name> --cpu 4 --memory 8192 --disk 100\nvmware-aiops vm delete <name>\nvmware-aiops vm clone <name> --new-name <new> [--to-host <host>] [--to-datastore <ds>] [--power-on]\nvmware-aiops vm migrate <name> --to-host <host> [--to-datastore <ds>]\nvmware-aiops vm snapshot-create <name> --name <snap> [--description <text>] [--memory]\nvmware-aiops vm snapshot-list <name>\nvmware-aiops vm snapshot-revert <name> --name <snap>\nvmware-aiops vm snapshot-delete <name> --name <snap> [--remove-children] [--no-wait]\nvmware-aiops vm task-status <task-id>                      # poll an async (--no-wait) operation by id\nvmware-aiops vm set-ttl <name> --minutes 480 [--dry-run]   # double confirm; daemon auto-deletes VM on expiry\n\n# Guest operations (requires VMware Tools)\nvmware-aiops vm guest-exec <name> --cmd <script-path> --args \"<args>\" --user <username>\nvmware-aiops vm guest-upload <name> --local ./script.sh --guest /tmp/script.sh --user <username>\n\n# Deploy\nvmware-aiops deploy ova <path> --name <vm> --datastore <ds>\nvmware-aiops deploy linked-clone --source <vm> --snapshot <snap> --name <new>\n\n# Cluster\nvmware-aiops cluster create <name> --ha --drs\nvmware-aiops cluster info <name>\nvmware-aiops cluster drs-rules <name>                                     # list DRS rules\nvmware-aiops cluster drs-rule-set <name> --rule <r> --enable|--disable [--dry-run]\nvmware-aiops cluster drs-rule-create <name> --rule <r> --type antiAffinity --vm <vm1> --vm <vm2> [--disabled] [--dry-run]\nvmware-aiops cluster drs-rule-delete <name> --rule <r> [--dry-run]        # VM-VM only; double confirm\n\n# Datastore\nvmware-aiops datastore browse <ds> --pattern \"*.ova\"\n\n# Alarm management\nvmware-aiops alarm list [--target <t>]\nvmware-aiops alarm acknowledge <entity_name> <alarm_name> [--target <t>]\nvmware-aiops alarm reset <entity_name> <alarm_name> [--target <t>]   # double confirm; see blast radius above\n\n# Family\nvmware-aiops hub status        # show installed family members + install commands\n```\n\n> Full CLI reference: see `references/cli-reference.md`\n\n## Troubleshooting\n\n### \"VM not found\" error\nVM names are case-sensitive in vSphere. Use exact name from `vmware-monitor inventory vms`.\n\n### Guest exec returns empty output\nUse `vm_guest_exec_output` instead of `vm_guest_exec` — it auto-captures stdout/stderr. Basic `vm_guest_exec` only returns exit code.\n\n### Deploy OVA times out\nLarge OVA files (>10GB) may exceed the default 120s timeout. The upload happens via HTTP NFC lease — ensure network between the machine running vmware-aiops and ESXi is stable.\n\n### Snapshot delete is slow / \"still running after Ns\"\nDeleting an old or large snapshot consolidates its delta disk into the parent — the slowest write\noperation, often several minutes. `vm snapshot-delete` waits up to 30 min by default; if it still\nreturns a \"still running, NOT failed\" message with a task id, the delete did **not** fail — poll it with\n`vm task-status <task-id>`. Do not re-issue the delete or hand-roll polling. For very large snapshots,\nprefer `vm snapshot-delete <name> --name <snap> --no-wait` to get the task id immediately and poll.\n\n### Plan apply fails mid-way\nRun `vmware-aiops plan list` to see failed plan status. Ask user if they want to rollback with `vm_rollback_plan`. Irreversible steps (delete_vm) are skipped during rollback.\n\n### Connection refused / SSL error\n1. Verify target is reachable: `vmware-aiops doctor`\n2. For self-signed certs: set `verify_ssl: false` in config.yaml (lab environments only)\n\n## Setup\n\n```bash\nuv tool install vmware-aiops==1.11.0\nmkdir -p ~/.vmware-aiops\nvmware-aiops init  # generates config.yaml and .env templates\nchmod 600 ~/.vmware-aiops/.env\n```\n\n> Full setup guide, security details, and AI platform compatibility: see `references/setup-guide.md`\n\n## License\n\nMIT — [github.com/vmware-skills/VMware-AIops](https://github.com/vmware-skills/VMware-AIops)\n\nFile v1.11.0:_meta.json\n\n{\n  \"ownerId\": \"kn7b067awq2s97bn3d7p5qfhw5827pxc\",\n  \"slug\": \"vmware-aiops\",\n  \"version\": \"1.11.0\",\n  \"publishedAt\": 1789790097901\n}\n\nFile v1.11.0:references/agent-guardrails.md\n\n# Operating vmware-aiops with a local / small model\n\nClaude-class models drive this skill without special instruction. Smaller and\nlocally-hosted models — Llama 3.3 70B, Qwen, Mistral, and similar, served\nthrough Goose, Ollama, or OpenShift AI — need explicit operating rules to call\ntools reliably.\n\nThis page exists because an operator wrote those rules by hand first. The\nguardrails below are adapted, with thanks, from the working configuration\n[@juanpf-ha](https://github.com/juanpf-ha) developed while running\nvmware-monitor and vmware-aria against a production vSphere estate with Llama\n3.3 70B FP8 on an on-prem H100\n([VMware-AIops#31](https://github.com/vmware-skills/VMware-AIops/issues/31)). The\ncross-skill rules are identical across this family; the parts below marked\nvmware-aiops are specific to this skill.\n\nvmware-aiops carries the family's largest write surface — 43 of its 60 MCP\ntools change state, including `vm_delete`, cluster deletion, host VMkernel\nremoval and guest command execution. Of every skill here, this is the one where a model's discipline\nshould not be the only thing standing between a prompt and a destroyed VM — and\nover MCP, apart from optional deny rules, the only enforcement is the RBAC of\nthe vCenter/ESXi account the server connects with. Run it under a dedicated,\nleast-privilege service account scoped to what the agent may change.\n\n> **Disclaimer**: This is a community-maintained open-source project and is\n> **not affiliated with, endorsed by, or sponsored by VMware, Inc. or Broadcom\n> Inc.** \"VMware\" and \"vSphere\" are trademarks of Broadcom.\n\n---\n\n## First: the rules you no longer need to write\n\nSeveral guardrails from the original configuration are now enforced by the\nskill itself. Prompt instructions are advisory — a model can ignore them.\nThese are structural, so it cannot.\n\n| Guardrail you would otherwise prompt for | Now enforced by |\n|---|---|\n| \"Use explicit limits for queries that may return large amounts of data\" | **The list envelope.** `browse_datastore`, `list_vcenter_alarms`, `vm_list_plans`, `vm_list_snapshots` and `vm_list_ttl` return `{items, returned, limit, total, truncated, hint}`, so the model reads truncation instead of guessing at it. |\n| \"If a listing came back empty, say so rather than claiming the call failed\" | Same envelope. Empty `items` with `truncated: false` means checked-and-none — a stated result, not a silence the model has to interpret. |\n| \"Log every state change you make\" | **The `@vmware_tool` decorator.** Every write is recorded to `~/.vmware/audit.db` before the model sees the result, and policy rules are evaluated ahead of execution. Neither depends on the model cooperating. |\n| \"Block state-changing writes against a production target\" | **Policy.** An opt-in environment-scoped `deny` rule in `~/.vmware/rules.yaml` matches a target's `environment:` label and refuses matching writes before execution. |\n\n---\n\n## The system prompt\n\nEverything below still benefits from being stated explicitly. Copy this into\nyour agent's instruction block.\n\n```text\n## Tool use\n\n- Always call an MCP tool before answering any question about the current\n  VMware environment. Never answer from memory or assumption.\n- Never describe a tool call, and never output a JSON example, instead of\n  executing the tool. If you intend to call a tool, call it.\n- If a tool fails, report the actual error text. Do not complete the answer\n  with assumptions about what the result would have been.\n- Use explicit limits on queries that may return large amounts of data. Do not\n  request unlimited results unless the user asks for them.\n- Before any write, restate the exact object you are about to change and wait\n  for the user to confirm it. VM names are case-sensitive and near-duplicates\n  are common.\n\n## Skill routing\n\n- vmware-aiops: VM lifecycle (power, create, clone, migrate, delete,\n  snapshots), OVA/template deployment, guest operations, clusters, plan/apply.\n- vmware-monitor: read-only vCenter inventory, hosts, datastores, alarms,\n  events, performance. Prefer it for any question that only reads.\n- vmware-storage: iSCSI, vSAN, datastore capacity.\n- vmware-vks: Supervisor, namespaces, Tanzu Kubernetes clusters.\n- vmware-nsx / vmware-nsx-security: networking and firewall.\n- vmware-aria: Aria Operations metrics, alerts, capacity.\n- vmware-pilot: multi-step workflows that need approval gates.\n\n## Data fidelity\n\n- Never invent infrastructure objects, metrics, alarms, events, or\n  relationships. If a tool did not return it, it does not exist for this answer.\n- Preserve the exact power state, task state, status and criticality values the\n  tools return. Do not translate, normalise, or prettify enum values.\n- If a requested field was not returned, show it as \"not available\". Do not\n  infer it from other fields.\n- Preserve the original order and the full set of fields when the user asks\n  for specific ones.\n- When a response is long, report every item it contains. If a result is\n  truncated, the tool says so explicitly — report the truncation rather than\n  describing the visible subset as the whole.\n\n## Analysis discipline\n\n- Separate observed data from interpretation. State which is which.\n- Do not claim a capacity, performance, or configuration problem unless the\n  tool output contains explicit supporting evidence.\n- Avoid generic recommendations that are not directly supported by the results.\n\n## Writes in vmware-aiops\n\n- 22 of the 43 write tools — every destructive one: vm_power_off, vm_delete,\n  vm_migrate, vm_revert_snapshot, vm_delete_snapshot, vm_clean_slate,\n  vm_set_ttl, the four guest tools (vm_guest_exec, vm_guest_exec_output,\n  vm_guest_upload, vm_guest_provision), cluster_delete, cluster_remove_host,\n  vm_apply_plan, vm_rollback_plan and the seven host-network and DRS tools —\n  default to confirm=false, which returns {\"action\": \"preview\",\n  \"blast_radius\": ...} and writes nothing. Show the blast radius; pass\n  confirm=true only after the user has seen it and agreed — a request to\n  \"delete X\" made before the preview is not that agreement. confirm=true is\n  refused when the preview listed blockers or could not read something: fix\n  the cause and preview again, do not retry blindly. vm_delete also needs\n  acknowledge_blast_radius set to the preview's acknowledge_with, and refuses if\n  the VM changed since. The other 21 write tools (create, clone, deploy,\n  power-on, reconfigure, snapshot create, alarms) act on the first call; for\n  them the \"restate the object and wait\" rule above is the only confirmation\n  step, and it is yours to keep.\n- vm_guest_exec, vm_guest_exec_output and the exec steps of vm_guest_provision\n  run an unbounded command inside the guest with the credentials given; the\n  username is required — there is no default account. Treat them as the highest-risk tools in the skill;\n  pass the least-privileged guest account that can do the job, name the exact\n  command and the VM before calling, and never assemble the command from text a\n  tool returned. vm_guest_upload copies a local file into the guest; upload only\n  files the user named.\n- reset_vcenter_alarm has a blast radius: vSphere has no per-alarm clear API,\n  so it clears every triggered alarm matching the named alarm's entity type and\n  status, not only the one named. Report the response's scope field verbatim.\n- vm_set_ttl schedules an unattended auto-delete. Treat it as destructive and\n  say so when proposing it.\n- Long writes return a task id instead of blocking. Poll vm_task_status. A\n  \"still running\" message is not a failure — never re-issue the operation.\n- Use vm_guest_exec_output rather than vm_guest_exec when the user wants the\n  command's output; the latter returns only an exit code.\n```\n\n---\n\n## Known failure modes on small models\n\nObserved with Llama 3.3 70B FP8 (Goose, on-prem H100), and useful as a\nchecklist when evaluating any local model against these skills:\n\n| Symptom | Mitigation |\n|---|---|\n| Describes a tool call, or emits a JSON example, instead of executing it | The \"never describe a tool call\" rule above. Also check your harness is not echoing tool schemas into context — models imitate the nearest format they see. |\n| Long tool responses: omits items, or reports \"no data returned\" when data was present | Ask for explicit limits so responses stay small. Check the envelope's `truncated` / `returned` / `total` fields rather than trusting the model's summary — a \"no data\" claim is checkable against `returned`. |\n| Adds generic recommendations unsupported by results | The \"analysis discipline\" rules. |\n| Drops requested fields or reorders results | State the required fields and ordering in the request itself, not only in the system prompt. |\n| Multi-tool workflows take 30–50s end to end | Prefer the aggregate tools — `cluster_health_summary`, `vm_investigation_bundle`, `host_investigation_bundle`, `datastore_investigation_bundle`, `cross_vcenter_attention` — which collapse a 3-4 call sequence into one round trip. |\n| Picks a write tool for a question that only reads | Route read questions to vmware-monitor. A model that can see 43 write tools will sometimes reach for one to \"check\" something. |\n| Treats a long-running task's \"still running\" reply as a failure and re-issues the write | The `vm_task_status` rule above. A re-issued clone or delete is the worst outcome in this skill. |\n| Assumes an alarm reset cleared only the alarm it named | Report `scope` from the response. The clear is entity-type-wide by design. |\n\n## Reporting results\n\nLocal-model compatibility is an explicit design constraint for this family, and\nthe evidence base is small. If you evaluate a model against this skill —\nQwen, Mistral, Granite, or anything else — a report of what worked and what did\nnot is genuinely useful:\n[github.com/vmware-skills/VMware-AIops/issues](https://github.com/vmware-skills/VMware-AIops/issues).\n\nFile v1.11.0:references/capabilities.md\n\n# Capabilities Reference\n\n## Automation Level Reference\n\nEach operation is classified by autonomy level per the Enterprise Harness Engineering framework. This tells AI agents how much human gating each tool needs:\n\n| Level | Meaning | Agent autonomy | Examples in this skill |\n|:-:|---|---|---|\n| **L1** | Read-only, raw data | Always auto-run | `cluster_info`, `browse_datastore`, `scan_datastore_images`, `list_vcenter_alarms`, `vm_list_snapshots`, `vm_list_ttl`, `vm_task_status` |\n| **L2** | Read + analysis / recommendation | Always auto-run | `cluster_health_summary`, `cross_vcenter_attention`, `vm_investigation_bundle`, `host_investigation_bundle`, `datastore_investigation_bundle`; scheduled scan reports, alarm/event correlation, log pattern analysis |\n| **L3** | Single write | The level is a statement about blast radius, not about an enforced gate. On the CLI the destructive ones double-confirm (`vm power-on` and `vm snapshot-create` do not); over MCP every destructive one (`vm_power_off`, `vm_delete`, `vm_migrate`, …) returns a no-write preview of its blast radius unless called with `confirm=True`, while `vm_power_on`, `vm_create_snapshot` and `vm_clone` act on the first call — see [What gates a write](#what-gates-a-write) | `vm_power_on`, `vm_power_off`, `vm_delete`, `vm_create_snapshot`, `vm_clone`, `vm_migrate` |\n| **L4** | Multi-step plan / apply workflow | Plan generation auto. Review with the user before applying — an agent convention, not enforced: `vm_apply_plan` takes only a plan id | `vm_create_plan` → `vm_apply_plan` → `vm_rollback_plan`, batch-clone, batch-deploy YAML |\n| **L5** | Auto-remediation from learned pattern | Pattern library only; requires `risk:low` + `reversible:true` + `repeatable:true` + signed approval | *(roadmap — not implemented; candidates: snapshot consolidation, orphaned VM cleanup)* |\n\n**Notes**:\n- L1/L2 tools are read-only and safe for agents to call unprompted.\n- **The levels describe risk, not enforcement.** The MCP previews stop an agent acting blind, and `confirm=True` is refused on a blocker, but nothing in this skill stops an agent that passes `confirm=True` from calling an L3 or L4 tool the account may call. What decides whether the write lands is the vCenter account — see [What gates a write](#what-gates-a-write).\n- **List envelope**: the read list tools (`browse_datastore`, `list_vcenter_alarms`, `vm_list_plans`, `vm_list_snapshots`, `vm_list_ttl`) return `{items, returned, limit, total, truncated, hint}` instead of a bare array, so an agent can tell a complete answer from a first page rather than inferring it (issue #31). All five enumerate their collection in full before any limit is applied, so `total` is always the real count; only `list_vcenter_alarms` takes a `limit` and can therefore report `truncated: true`. The write `batch_*` tools deliberately keep a bare list — each row is a per-item result of work already done, complete by construction. Errors from these read tools are `{error, hint}` (a dict, not a one-element list).\n- L3+ tools always pass through the `@vmware_tool` decorator: connection check → policy check (opt-in `deny` rules only; nothing is denied by default) → audit log. Confirmation is not in that chain; where a tool has one, it is the tool's own `confirm` argument (see [What gates a write](#what-gates-a-write)).\n- Multi-party approval, where it is genuinely required, is [vmware-pilot](https://github.com/vmware-skills/VMware-Pilot)'s job — it has the state machine and a real human approval step. See it also for cross-skill L4 orchestration and the Dispatcher/Subagent pattern.\n\n## What gates a write\n\n<!-- gate-inventory: the counts and names below are checked against the live\n     tool registry by tests/eval/regression/test_documented_gates_match_the_registry.py.\n     Adding a write tool makes them wrong and turns that test red. -->\n\n| Surface | Confirmation | Preview |\n|---|---|---|\n| **CLI** | Two interactive `typer.confirm` prompts before every irreversible or guest-writing command. The set is derived from the MCP `destructiveHint` annotations, so a new such command fails the test suite until it has them, and a declined prompt is audited as `rejected`. *Honest limitation:* an agent with a shell satisfies both prompts by piping `yes` into the command. This defends the mistyped command, not a determined caller. | `--dry-run` on every write command except `deploy iso`, `deploy mark-template`, `vm cancel-ttl` and `vm guest-download` |\n| **MCP** | **One argument, `confirm`, defaulting to a no-write preview** (family security HLD §7, revised 2026-09-16). Decision **D-2** (2026-07-21) cut a confirmation handshake as a speed-bump that does not authorize anything; it is superseded. The handshake still is not authorization — the account is (below) — but a preview is what stops an agent acting on a guess, and a refusal is what stops it acting on a guess that has gone stale. The read-only \n\nArchive v1.10.0: 9 files, 39687 bytes\n\nFiles: evals/evals.json (2055b), references/agent-guardrails.md (9554b), references/capabilities.md (28093b), references/cli-reference.md (4650b), references/investigation-protocol.md (6771b), references/setup-guide.md (10920b), skill-card.md (2710b), SKILL.md (26631b), _meta.json (132b)\n\nArchive v1.9.7: 9 files, 39204 bytes\n\nFiles: evals/evals.json (2055b), references/agent-guardrails.md (9378b), references/capabilities.md (27424b), references/cli-reference.md (4587b), references/investigation-protocol.md (6771b), references/setup-guide.md (10881b), skill-card.md (2686b), SKILL.md (26434b), _meta.json (131b)\n\nArchive v1.9.5: 9 files, 39328 bytes\n\nFiles: evals/evals.json (2055b), references/agent-guardrails.md (9378b), references/capabilities.md (27424b), references/cli-reference.md (4587b), references/investigation-protocol.md (6771b), references/setup-guide.md (10881b), skill-card.md (2900b), SKILL.md (26434b), _meta.json (131b)\n\nArchive v1.9.4: 9 files, 39469 bytes\n\nFiles: evals/evals.json (2055b), references/agent-guardrails.md (9378b), references/capabilities.md (27424b), references/cli-reference.md (4587b), references/investigation-protocol.md (6771b), references/setup-guide.md (10881b), skill-card.md (3296b), SKILL.md (26434b), _meta.json (131b)\n\nArchive v1.9.3: 9 files, 39422 bytes\n\nFiles: evals/evals.json (2055b), references/agent-guardrails.md (9378b), references/capabilities.md (27424b), references/cli-reference.md (4587b), references/investigation-protocol.md (6771b), references/setup-guide.md (10881b), skill-card.md (3097b), SKILL.md (26434b), _meta.json (131b)\n\nArchive v1.9.2: 9 files, 39280 bytes\n\nFiles: evals/evals.json (2055b), references/agent-guardrails.md (9378b), references/capabilities.md (27394b), references/cli-reference.md (4587b), references/investigation-protocol.md (6771b), references/setup-guide.md (10850b), skill-card.md (2828b), SKILL.md (26438b), _meta.json (131b)\n\nArchive v1.9.1: 9 files, 39165 bytes\n\nFiles: evals/evals.json (2055b), references/agent-guardrails.md (9378b), references/capabilities.md (27394b), references/cli-reference.md (4587b), references/investigation-protocol.md (6771b), references/setup-guide.md (10850b), skill-card.md (2577b), SKILL.md (26443b), _meta.json (131b)\n\nArchive v1.9.0: 9 files, 39397 bytes\n\nFiles: evals/evals.json (2055b), references/agent-guardrails.md (9378b), references/capabilities.md (27394b), references/cli-reference.md (4587b), references/investigation-protocol.md (6771b), references/setup-guide.md (10850b), skill-card.md (3033b), SKILL.md (26443b), _meta.json (131b)","readmeExcerpt":"Skill: vmware-aiops Owner: zw008 Summary: Use this skill whenever the user needs to manage VMs in VMware/vSphere/ESXi — it's the entry point for all VM operations. Directly handles: power on/off, clone, snapshot, migrate, deploy from OVA or templates, run commands inside VMs, batch operations, cluster management, vCenter alarm acknowledgment, a one-glance cluster-health triage (\"is anything on fire?\"), and VM/host/da","codeSnippets":[],"executableExamples":[{"language":"bash","snippet":"uv tool install vmware-aiops==1.12.0\nvmware-aiops doctor\nvmware-aiops hub status   # see which family members are installed"},{"language":"bash","snippet":"# VM operations\nvmware-aiops vm power-on <name> [--target <t>]\nvmware-aiops vm power-off <name> [--force]\nvmware-aiops vm create <name> --cpu 4 --memory 8192 --disk 100\nvmware-aiops vm delete <name>\nvmware-aiops vm clone <name> --new-name <new> [--to-host <host>] [--to-datastore <ds>] [--power-on]\nvmware-aiops vm migrate <name> --to-host <host> [--to-datastore <ds>]\nvmware-aiops vm snapshot-create <name> --name <snap> [--description <text>] [--memory]\nvmware-aiops vm snapshot-list <name>\nvmware-aiops vm snapshot-revert <name> --name <snap>\nvmware-aiops vm snapshot-delete <name> --name <snap> [--remove-children] [--no-wait]\nvmware-aiops vm task-status <task-id>                      # poll an async (--no-wait) operation by id\nvmware-aiops vm set-ttl <name> --minutes 480 [--dry-run]   # double confirm; daemon auto-deletes VM on expiry\n\n# Guest operations (requires VMware Tools)\nvmware-aiops vm guest-exec <name> --cmd <script-path> --args \"<args>\" --user <username>\nvmware-aiops vm guest-upload <name> --local ./script.sh --guest /tmp/script.sh --user <username>\n\n# Deploy\nvmware-aiops deploy ova <path> --name <vm> --datastore <ds>\nvmware-aiops deploy linked-clone --source <vm> --snapshot <snap> --name <new>\n\n# Cluster\nvmware-aiops cluster create <name> --ha --drs\nvmware-aiops cluster info <name>\nvmware-aiops cluster drs-rules <name>                                     # list DRS rules\nvmware-aiops cluster drs-rule-set <name> --rule <r> --enable|--disable [--dry-run]\nvmware-aiops cluster drs-rule-create <name> --rule <r> --type antiAffinity --vm <vm1> --vm <vm2> [--disabled] [--dry-run]\nvmware-aiops cluster drs-rule-delete <name> --rule <r> [--dry-run]        # VM-VM only; double confirm\n\n# Datastore\nvmware-aiops datastore browse <ds> --pattern \"*.ova\"\n\n# Alarm management\nvmware-aiops alarm list [--target <t>]\nvmware-aiops alarm acknowledge <entity_name> <alarm_name> [--target <t>]\nvmware-aiops alarm reset <entity_name> <alarm_name> [--target <t>]   # double confirm; see b"},{"language":"bash","snippet":"uv tool install vmware-aiops==1.12.0\nmkdir -p ~/.vmware-aiops\nvmware-aiops init  # generates config.yaml and .env templates\nchmod 600 ~/.vmware-aiops/.env"},{"language":"text","snippet":"## Tool use\n\n- Always call an MCP tool before answering any question about the current\n  VMware environment. Never answer from memory or assumption.\n- Never describe a tool call, and never output a JSON example, instead of\n  executing the tool. If you intend to call a tool, call it.\n- If a tool fails, report the actual error text. Do not complete the answer\n  with assumptions about what the result would have been.\n- Use explicit limits on queries that may return large amounts of data. Do not\n  request unlimited results unless the user asks for them.\n- Before any write, restate the exact object you are about to change and wait\n  for the user to confirm it. VM names are case-sensitive and near-duplicates\n  are common.\n\n## Skill routing\n\n- vmware-aiops: VM lifecycle (power, create, clone, migrate, delete,\n  snapshots), OVA/template deployment, guest operations, clusters, plan/apply.\n- vmware-monitor: read-only vCenter inventory, hosts, datastores, alarms,\n  events, performance. Prefer it for any question that only reads.\n- vmware-storage: iSCSI, vSAN, datastore capacity.\n- vmware-vks: Supervisor, namespaces, Tanzu Kubernetes clusters.\n- vmware-nsx / vmware-nsx-security: networking and firewall.\n- vmware-aria: Aria Operations metrics, alerts, capacity.\n- vmware-pilot: multi-step workflows that need approval gates.\n\n## Data fidelity\n\n- Never invent infrastructure objects, metrics, alarms, events, or\n  relationships. If a tool did not return it, it does not exist for this answer.\n- Preserve the exact power state, task state, status and criticality values the\n  tools return. Do not translate, normalise, or prettify enum values.\n- If a requested field was not returned, show it as \"not available\". Do not\n  infer it from other fields.\n- Preserve the original order and the full set of fields when the user asks\n  for specific ones.\n- When a response is long, report every item it contains. If a result is\n  truncated, the tool says so explicitly — report the truncation rather tha"},{"language":"bash","snippet":"# Diagnostics\nvmware-aiops doctor [--skip-auth]   # --skip-auth only skips doctor's own vSphere login check; no other command has it\n\n# MCP Config Generator\nvmware-aiops mcp-config generate --agent <goose|cursor|claude-code|continue|vscode-copilot|localcowork|mcp-agent>\nvmware-aiops mcp-config list\n\n# VM Operations\nvmware-aiops vm power-on <vm-name>\nvmware-aiops vm power-off <vm-name> [--force]\nvmware-aiops vm create <name> [--cpu <n>] [--memory <mb>] [--disk <gb>]\nvmware-aiops vm delete <vm-name>\nvmware-aiops vm reconfigure <vm-name> [--cpu <n>] [--memory <mb>]\nvmware-aiops vm snapshot-create <vm-name> --name <snap-name> [--description <text>] [--memory]\nvmware-aiops vm snapshot-list <vm-name>\nvmware-aiops vm snapshot-revert <vm-name> --name <snap-name>\nvmware-aiops vm snapshot-delete <vm-name> --name <snap-name> [--remove-children]\nvmware-aiops vm clone <vm-name> --new-name <name> [--to-host <host>] [--to-datastore <ds>] [--power-on]\nvmware-aiops vm migrate <vm-name> --to-host <host> [--to-datastore <ds>]\nvmware-aiops vm set-ttl <vm-name> --minutes <n>\nvmware-aiops vm cancel-ttl <vm-name>\nvmware-aiops vm list-ttl\nvmware-aiops vm clean-slate <vm-name> [--snapshot baseline]\n\n# Guest Operations (requires VMware Tools)\nvmware-aiops vm guest-exec <vm-name> --cmd /bin/bash --args \"-c 'ls -la /tmp'\" --user root\nvmware-aiops vm guest-upload <vm-name> --local ./script.sh --guest /tmp/script.sh --user root\nvmware-aiops vm guest-download <vm-name> --guest /var/log/syslog --local ./syslog.txt --user root\n\n# Plan → Apply (multi-step operations)\nvmware-aiops plan list\n\n# Deploy\nvmware-aiops deploy ova <path> --name <vm-name> [--datastore <ds>] [--network <net>]\nvmware-aiops deploy template <template-name> --name <vm-name> [--datastore <ds>]\nvmware-aiops deploy linked-clone --source <vm> --snapshot <snap> --name <new-name>\nvmware-aiops deploy iso <vm-name> --iso \"[datastore] path/file.iso\"\nvmware-aiops deploy mark-template <vm-name>\nvmware-aiops deploy batch-clone --source <vm> "},{"language":"text","snippet":"🔴 [ROOT CAUSE]   <falsifiable, mechanism-explained statement>\n  → [PROPAGATION] <how the root cause spread to neighboring systems>\n    → [AMPLIFICATION] <what made the impact worse, if applicable>\n      → [IMPACT]    <observable user / business / SLA effect>\n\n✅ Falsifiability:  <evidence — metric name, log query, command output>\n✅ Sufficiency:     <evidence or stated counterfactual>\n✅ Necessity:       <evidence or peer comparison>\n✅ Mechanism:       <see propagation chain above>"}],"parameters":null,"dependencies":[],"permissions":[],"extractedFiles":[{"path":"SKILL.md","content":"---\nname: vmware-aiops\ndescription: >\n  Use this skill whenever the user needs to manage VMs in VMware/vSphere/ESXi — it's the entry point for all VM operations.\n  Directly handles: power on/off, clone, snapshot, migrate, deploy from OVA or templates, run commands inside VMs, batch operations, cluster management, vCenter alarm acknowledgment, a one-glance cluster-health triage (\"is anything on fire?\"), and VM/host/datastore investigation drill-downs.\n  Always use this skill for any \"power on\", \"clone\", \"deploy\", \"migrate\", \"batch\", \"guest exec\", \"alarm\", or VM lifecycle task, and for triage like \"is anything on fire\" / \"what needs attention now\" / \"investigate this VM\", when the context is explicitly VMware, vSphere, or ESXi.\n  Do NOT use for general read-only queries (inventory/events/VM details — use vmware-monitor), NSX networking (use vmware-nsx), storage/iSCSI/vSAN (use vmware-storage), or Kubernetes cluster lifecycle (use vmware-vks).\n  For multi-step workflows use vmware-pilot. For load balancing/AVI/AKO use vmware-avi.\ninstaller:\n  kind: uv\n  package: vmware-aiops\nargument-hint: \"[vm-name or describe your task]\"\nallowed-tools:\n  - Bash\nmetadata: {\"openclaw\":{\"requires\":{\"anyBins\":[\"vmware-aiops\",\"uvx\"]},\"optional\":{\"env\":[\"VMWARE_AIOPS_CONFIG\",\"VMWARE_TARGET_PASSWORD\",\"VMWARE_<TARGET>_USERNAME\",\"SLACK_WEBHOOK_URL\",\"DISCORD_WEBHOOK_URL\",\"VMWARE_AUDIT_APPROVED_BY\"],\"bins\":[\"vmware-policy\"]},\"homepage\":\"https://github.com/vmware-skills/VMware-AIops\",\"emoji\":\"🖥️\",\"os\":[\"macos\",\"linux\"]}}\ncompatibility: >\n  vmware-policy auto-installed as Python dependency (provides @vmware_tool decorator and audit logging). All write operations audited to ~/.vmware/audit.db.\n  Credentials: Each vCenter/ESXi target requires a per-target password env var in ~/.vmware-aiops/.env following the pattern VMWARE_<TARGET_NAME_UPPER>_PASSWORD. Passwords are never logged or echoed.\n  Destructive operations: All write tools require explicit parameters and pass through the @vmware_tool decorator (policy check + audit + sanitize). Every MCP write tool annotated destructive (22 of 43: power-off, delete, migrate, snapshot revert/delete, guest exec/upload/provision, cluster delete/remove-host, TTL, Clean Slate, plan apply/rollback, host-network and DRS) takes one confirm argument whose default returns a no-write blast-radius preview (HLD section 7); confirm=True is refused, and audited as a failure, on a blocker or an unreadable measurement, and vm_delete also requires the preview's acknowledgement echoed back and still matching; the other 21 write tools (create, clone, deploy, power-on, reconfigure) act on the first call. The enforcement boundary is the RBAC of the vCenter/ESXi account the server connects with, so run it under a dedicated least-privilege service account (a read-only role makes it read-only). Optional deny rules in ~/.vmware/rules.yaml are checked before every MCP call and remote CLI command; the shipped baseline denies nothing. CLI destructive commands addi"},{"path":"_meta.json","content":"{\n  \"ownerId\": \"kn7b067awq2s97bn3d7p5qfhw5827pxc\",\n  \"slug\": \"vmware-aiops\",\n  \"version\": \"1.12.0\",\n  \"publishedAt\": 1789915898030\n}"},{"path":"references/agent-guardrails.md","content":"# Operating vmware-aiops with a local / small model\n\nClaude-class models drive this skill without special instruction. Smaller and\nlocally-hosted models — Llama 3.3 70B, Qwen, Mistral, and similar, served\nthrough Goose, Ollama, or OpenShift AI — need explicit operating rules to call\ntools reliably.\n\nThis page exists because an operator wrote those rules by hand first. The\nguardrails below are adapted, with thanks, from the working configuration\n[@juanpf-ha](https://github.com/juanpf-ha) developed while running\nvmware-monitor and vmware-aria against a production vSphere estate with Llama\n3.3 70B FP8 on an on-prem H100\n([VMware-AIops#31](https://github.com/vmware-skills/VMware-AIops/issues/31)). The\ncross-skill rules are identical across this family; the parts below marked\nvmware-aiops are specific to this skill.\n\nvmware-aiops carries the family's largest write surface — 43 of its 60 MCP\ntools change state, including `vm_delete`, cluster deletion, host VMkernel\nremoval and guest command execution. Of every skill here, this is the one where a model's discipline\nshould not be the only thing standing between a prompt and a destroyed VM — and\nover MCP, apart from optional deny rules, the only enforcement is the RBAC of\nthe vCenter/ESXi account the server connects with. Run it under a dedicated,\nleast-privilege service account scoped to what the agent may change.\n\n> **Disclaimer**: This is a community-maintained open-source project and is\n> **not affiliated with, endorsed by, or sponsored by VMware, Inc. or Broadcom\n> Inc.** \"VMware\" and \"vSphere\" are trademarks of Broadcom.\n\n---\n\n## First: the rules you no longer need to write\n\nSeveral guardrails from the original configuration are now enforced by the\nskill itself. Prompt instructions are advisory — a model can ignore them.\nThese are structural, so it cannot.\n\n| Guardrail you would otherwise prompt for | Now enforced by |\n|---|---|\n| \"Use explicit limits for queries that may return large amounts of data\" | **The list envelope.** `browse_datastore`, `list_vcenter_alarms`, `vm_list_plans`, `vm_list_snapshots` and `vm_list_ttl` return `{items, returned, limit, total, truncated, hint}`, so the model reads truncation instead of guessing at it. |\n| \"If a listing came back empty, say so rather than claiming the call failed\" | Same envelope. Empty `items` with `truncated: false` means checked-and-none — a stated result, not a silence the model has to interpret. |\n| \"Log every state change you make\" | **The `@vmware_tool` decorator.** Every write is recorded to `~/.vmware/audit.db` before the model sees the result, and policy rules are evaluated ahead of execution. Neither depends on the model cooperating. |\n| \"Block state-changing writes against a production target\" | **Policy.** An opt-in environment-scoped `deny` rule in `~/.vmware/rules.yaml` matches a target's `environment:` label and refuses matching writes before execution. |\n\n---\n\n## The system prompt\n\nEverything below still benefits from being stated e"},{"path":"references/capabilities.md","content":"# Capabilities Reference\n\n## Automation Level Reference\n\nEach operation is classified by autonomy level per the Enterprise Harness Engineering framework. This tells AI agents how much human gating each tool needs:\n\n| Level | Meaning | Agent autonomy | Examples in this skill |\n|:-:|---|---|---|\n| **L1** | Read-only, raw data | Always auto-run | `cluster_info`, `browse_datastore`, `scan_datastore_images`, `list_vcenter_alarms`, `vm_list_snapshots`, `vm_list_ttl`, `vm_task_status` |\n| **L2** | Read + analysis / recommendation | Always auto-run | `cluster_health_summary`, `cross_vcenter_attention`, `vm_investigation_bundle`, `host_investigation_bundle`, `datastore_investigation_bundle`; scheduled scan reports, alarm/event correlation, log pattern analysis |\n| **L3** | Single write | The level is a statement about blast radius, not about an enforced gate. On the CLI the destructive ones double-confirm (`vm power-on` and `vm snapshot-create` do not); over MCP every destructive one (`vm_power_off`, `vm_delete`, `vm_migrate`, …) returns a no-write preview of its blast radius unless called with `confirm=True`, while `vm_power_on`, `vm_create_snapshot` and `vm_clone` act on the first call — see [What gates a write](#what-gates-a-write) | `vm_power_on`, `vm_power_off`, `vm_delete`, `vm_create_snapshot`, `vm_clone`, `vm_migrate` |\n| **L4** | Multi-step plan / apply workflow | Plan generation auto. Review with the user before applying — an agent convention, not enforced: `vm_apply_plan` takes only a plan id | `vm_create_plan` → `vm_apply_plan` → `vm_rollback_plan`, batch-clone, batch-deploy YAML |\n| **L5** | Auto-remediation from learned pattern | Pattern library only; requires `risk:low` + `reversible:true` + `repeatable:true` + signed approval | *(roadmap — not implemented; candidates: snapshot consolidation, orphaned VM cleanup)* |\n\n**Notes**:\n- L1/L2 tools are read-only and safe for agents to call unprompted.\n- **The levels describe risk, not enforcement.** The MCP previews stop an agent acting blind, and `confirm=True` is refused on a blocker, but nothing in this skill stops an agent that passes `confirm=True` from calling an L3 or L4 tool the account may call. What decides whether the write lands is the vCenter account — see [What gates a write](#what-gates-a-write).\n- **List envelope**: the read list tools (`browse_datastore`, `list_vcenter_alarms`, `vm_list_plans`, `vm_list_snapshots`, `vm_list_ttl`) return `{items, returned, limit, total, truncated, hint}` instead of a bare array, so an agent can tell a complete answer from a first page rather than inferring it (issue #31). All five enumerate their collection in full before any limit is applied, so `total` is always the real count; only `list_vcenter_alarms` takes a `limit` and can therefore report `truncated: true`. The write `batch_*` tools deliberately keep a bare list — each row is a per-item result of work already done, complete by construction. Errors from these read tools are `{error, hint}` ("},{"path":"references/cli-reference.md","content":"# CLI Reference\n\nDestructive and deploy commands ask for two confirmations and most write\ncommands take `--dry-run` (not `deploy iso`, `deploy mark-template`,\n`vm cancel-ttl`, `vm guest-download`). These are CLI-only: over MCP the 22 destructive write tools take `confirm` and preview by\ndefault, the other 21 write tools act immediately, and the enforcement boundary there is the RBAC of the vCenter/ESXi account — see\n`capabilities.md` → \"What gates a write\".\n\n```bash\n# Diagnostics\nvmware-aiops doctor [--skip-auth]   # --skip-auth only skips doctor's own vSphere login check; no other command has it\n\n# MCP Config Generator\nvmware-aiops mcp-config generate --agent <goose|cursor|claude-code|continue|vscode-copilot|localcowork|mcp-agent>\nvmware-aiops mcp-config list\n\n# VM Operations\nvmware-aiops vm power-on <vm-name>\nvmware-aiops vm power-off <vm-name> [--force]\nvmware-aiops vm create <name> [--cpu <n>] [--memory <mb>] [--disk <gb>]\nvmware-aiops vm delete <vm-name>\nvmware-aiops vm reconfigure <vm-name> [--cpu <n>] [--memory <mb>]\nvmware-aiops vm snapshot-create <vm-name> --name <snap-name> [--description <text>] [--memory]\nvmware-aiops vm snapshot-list <vm-name>\nvmware-aiops vm snapshot-revert <vm-name> --name <snap-name>\nvmware-aiops vm snapshot-delete <vm-name> --name <snap-name> [--remove-children]\nvmware-aiops vm clone <vm-name> --new-name <name> [--to-host <host>] [--to-datastore <ds>] [--power-on]\nvmware-aiops vm migrate <vm-name> --to-host <host> [--to-datastore <ds>]\nvmware-aiops vm set-ttl <vm-name> --minutes <n>\nvmware-aiops vm cancel-ttl <vm-name>\nvmware-aiops vm list-ttl\nvmware-aiops vm clean-slate <vm-name> [--snapshot baseline]\n\n# Guest Operations (requires VMware Tools)\nvmware-aiops vm guest-exec <vm-name> --cmd /bin/bash --args \"-c 'ls -la /tmp'\" --user root\nvmware-aiops vm guest-upload <vm-name> --local ./script.sh --guest /tmp/script.sh --user root\nvmware-aiops vm guest-download <vm-name> --guest /var/log/syslog --local ./syslog.txt --user root\n\n# Plan → Apply (multi-step operations)\nvmware-aiops plan list\n\n# Deploy\nvmware-aiops deploy ova <path> --name <vm-name> [--datastore <ds>] [--network <net>]\nvmware-aiops deploy template <template-name> --name <vm-name> [--datastore <ds>]\nvmware-aiops deploy linked-clone --source <vm> --snapshot <snap> --name <new-name>\nvmware-aiops deploy iso <vm-name> --iso \"[datastore] path/file.iso\"\nvmware-aiops deploy mark-template <vm-name>\nvmware-aiops deploy batch-clone --source <vm> --count <n> [--prefix <prefix>]\nvmware-aiops deploy batch <spec.yaml>\n\n# Cluster\nvmware-aiops cluster info <name>\nvmware-aiops cluster create <name> [--ha] [--drs] [--drs-behavior fullyAutomated|partiallyAutomated|manual] [--datacenter <dc>]\nvmware-aiops cluster delete <name>\nvmware-aiops cluster add-host <cluster> --host <hostname>\nvmware-aiops cluster remove-host <cluster> --host <hostname>   # host must be in maintenance mode; moved to datacenter host folder as standalone\nvmware-aiops cluster configure <name> [--ha/--no-"}],"languages":[],"docsSourceLabel":"CLAWHUB","editorialOverview":null,"editorialQuality":{"score":100,"threshold":65,"status":"thin","wordCount":2452,"uniquenessScore":37,"reasons":["uniqueness-below-45"]}},"media":{"evidence":{"source":"no-media","verified":false,"confidence":"low","updatedAt":"2026-10-09T04:26:49.944Z","emptyReason":"No screenshots, media assets, or demo links are available."},"primaryImageUrl":null,"mediaAssetCount":0,"assets":[],"demoUrl":null},"ownerResources":{"evidence":{"source":"unclaimed","verified":false,"confidence":"low","updatedAt":"2026-10-09T04:26:49.944Z","emptyReason":"This page has not been claimed by the agent owner."},"hasCustomPage":false,"customPageUpdatedAt":null,"customLinks":[],"structuredLinks":{"docsUrl":null,"demoUrl":null,"supportUrl":null,"pricingUrl":null,"statusUrl":null},"customPage":null},"relatedAgents":{"evidence":{"source":"protocol-neighbors","verified":false,"confidence":"medium","updatedAt":"2026-10-09T13:58:52.055Z","emptyReason":null},"items":[{"id":"b917f68a-ebff-438e-84f8-3f4b2494c0bc","entityType":"agent","canonicalPath":"/agent/activepieces-activepieces","slug":"activepieces-activepieces","name":"activepieces","description":"AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents","url":"https://github.com/activepieces/activepieces","homepage":"https://www.activepieces.com","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-15T02:22:12.426Z","createdAt":"2026-02-25T03:38:12.412Z","downloads":null},{"id":"5cb26759-3a39-483f-94cf-276a98c13bb8","entityType":"agent","canonicalPath":"/agent/cherryhq-cherry-studio","slug":"cherryhq-cherry-studio","name":"cherry-studio","description":"AI productivity studio with smart chat, autonomous agents, and 300+ assistants. Unified access to frontier LLMs","url":"https://github.com/CherryHQ/cherry-studio","homepage":"https://cherry-ai.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-11T14:38:40.986Z","createdAt":"2026-02-25T03:38:19.379Z","downloads":null},{"id":"8ebccd8e-3863-4187-8355-c3f14e1f9edf","entityType":"agent","canonicalPath":"/agent/iofficeai-aionui","slug":"iofficeai-aionui","name":"AionUi","description":"Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!","url":"https://github.com/iOfficeAI/AionUi","homepage":"https://www.aionui.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-10T18:48:31.762Z","createdAt":"2026-02-25T03:38:16.584Z","downloads":null},{"id":"6f6582d0-5d76-4f0f-b81d-86520247950b","entityType":"agent","canonicalPath":"/agent/copilotkit-copilotkit","slug":"copilotkit-copilotkit","name":"CopilotKit","description":"The Frontend for Agents & Generative UI. React + Angular","url":"https://github.com/CopilotKit/CopilotKit","homepage":"https://docs.copilotkit.ai","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-03-25T09:50:57.846Z","createdAt":"2026-02-25T03:39:14.617Z","downloads":null}],"links":{"hub":"/agent","source":"/agent/source/clawhub","protocols":[{"label":"OpenClaw","href":"/agent/protocol/openclew"}]}}}