{"id":"f9660d79-3b6d-42c5-b125-9971e80032e3","entityType":"agent","slug":"clawhub-zw008-k8s-aiops","name":"k8s-aiops","canonicalUrl":"https://www.xpersona.co/agent/clawhub-zw008-k8s-aiops","canonicalPath":"/agent/clawhub-zw008-k8s-aiops","generatedAt":"2026-10-10T10:53:18.745Z","source":"CLAWHUB","claimStatus":"UNCLAIMED","verificationTier":"NONE","summary":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T06:28:09.610Z","emptyReason":null},"description":"Use this skill whenever the user needs to operate a Kubernetes cluster — list/inspect pods, deployments, statefulsets, daemonsets, replicasets, jobs, cronjobs, services, ingresses, endpoints, configmaps, secrets (names/keys only), PVCs/PVs/storageclasses, nodes, namespaces, and events; read pod logs; describe pods/nodes; pod/node top (metrics); read-only diagnostics / RCA (pod-health, workload-readiness); scale deployments/statefulsets; rollout status/history/undo/pause/resume and set image; delete pods/deployments/jobs; create/delete namespaces; and cordon/uncordon/drain nodes. Works with any kubeconfig-reachable cluster (standard Kubernetes, k3s, EKS, GKE, AKS). Always use this skill for \"list k8s pods\", \"scale deployment\", \"kubernetes pod logs\", \"describe pod\", \"why is my pod crashing\", \"diagnose pods\", \"which deployments are unhealthy\", \"rollout undo\", \"set image\", \"top pods\", \"drain node\", \"cordon node\", \"restart deployment\", \"k3s\", or \"kubectl\"-style tasks when the context is explicitly Kubernetes / a cluster. Do NOT use when the target is not a Kubernetes cluster (hypervisor VM lifecycle, backup products, or cloud-provider consoles are out of scope). Common Kubernetes operations with a built-in governance harness (audit, token budget, undo, risk-tier labels).","descriptionLabel":"Source description","evidenceSummary":"Capability contract not published. No trust telemetry is available yet. 1.6K downloads reported by the source. Last updated 10/10/2026.","installCommand":"clawhub skill install s171xgnmqse0nqvgqvqnaq5f9183kyre:k8s-aiops","sourceUrl":"https://clawhub.ai/zw008/k8s-aiops","homepage":"https://clawhub.ai/zw008/skills/k8s-aiops","primaryLinks":[{"label":"View on ClawHub","url":"https://clawhub.ai/zw008/k8s-aiops","kind":"source"},{"label":"Homepage","url":"https://clawhub.ai/zw008/skills/k8s-aiops","kind":"homepage"}],"safetyScore":84,"overallRank":62,"popularityScore":64,"trustScore":null,"claimedByName":null,"isOwner":false,"seoDescription":"k8s-aiops technical dossier on Xpersona with agent coverage, OPENCLEW support, and live trust metadata."},"coverage":{"evidence":{"source":"public-profile","verified":false,"confidence":"medium","updatedAt":"2026-10-10T06:28:09.610Z","emptyReason":null},"protocols":[{"protocol":"OPENCLEW","label":"OpenClaw","status":"self-declared","notes":"Declared in the public agent profile."}],"capabilities":[],"verifiedCount":0,"selfDeclaredCount":1,"capabilityMatrix":{"rows":[{"key":"OPENCLEW","type":"protocol","support":"unknown","confidenceSource":"profile","notes":"Listed on profile"}],"flattenedTokens":"protocol:OPENCLEW|unknown|profile"}},"adoption":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T06:28:09.610Z","emptyReason":null},"stars":null,"forks":null,"downloads":1619,"packageName":null,"latestVersion":"0.13.3","tractionLabel":"1.6K downloads"},"release":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T06:28:09.610Z","emptyReason":null},"lastUpdatedAt":"2026-10-10T06:28:09.610Z","lastCrawledAt":"2026-10-10T06:28:09.610Z","lastIndexedAt":null,"nextCrawlAt":"2026-10-11T06:28:09.610Z","lastVerifiedAt":null,"highlights":[{"version":"0.13.3","createdAt":"2026-09-15T06:02:54.189Z","changelog":"- Removed the skill description file (skill-card.md). - No user-facing changes to functionality or features. - Documentation is unchanged except for the file removal.","fileCount":7,"zipByteSize":19602},{"version":"0.13.2","createdAt":"2026-09-12T14:26:40.832Z","changelog":"- skill-card.md file removed. - SKILL.md updated with new OpenClaw plugin installation instructions (changed plugin repo/user). - No code or functional changes to the skill logic itself. - Documentation clarifies that the skill-card.md is no longer referenced.","fileCount":7,"zipByteSize":19696},{"version":"0.13.1","createdAt":"2026-09-12T10:10:28.688Z","changelog":"- Documentation updated for plugin and OpenClaw usage; added installation info for OpenClaw plugin flow. - Explicit instructions for installing as an OpenClaw plugin and visibility settings. - Mentions requirement for `uvx` on `PATH` for MCP server install. - Removed obsolete skill-card.md file.","fileCount":7,"zipByteSize":19579},{"version":"0.13.0","createdAt":"2026-09-12T00:59:03.454Z","changelog":"## k8s-aiops 0.13.0 – Changelog - Updated environment and binary requirements for improved compatibility (now supports either `k8s-aiops` or `uvx` as valid binaries). - Expanded and clarified `metadata.openclaw` fields, adjusting environment variable recommendations. - Removed the `skill-card.md` file. - No new features or breaking changes to user-facing commands.","fileCount":7,"zipByteSize":19471},{"version":"0.12.0","createdAt":"2026-08-10T06:51:30.796Z","changelog":"- Removed the file: skill-card.md - No functional or behavior changes to the skill itself - Documentation and feature set remain unchanged","fileCount":7,"zipByteSize":19605},{"version":"0.11.0","createdAt":"2026-08-10T03:58:12.650Z","changelog":"- Removed the file skill-card.md. - No changes to features, functionality, or compatibility.","fileCount":7,"zipByteSize":19466},{"version":"0.10.0","createdAt":"2026-08-03T05:53:12.435Z","changelog":"- Removed the file skill-card.md. - No changes to functionality or user-facing features. - Documentation and usage remain unchanged.","fileCount":7,"zipByteSize":19556},{"version":"0.9.0","createdAt":"2026-08-02T09:39:56.859Z","changelog":"- Removed the skill-card.md file from the repository. - No functional or behavioral changes to the skill itself. - Documentation and core functionality remain unchanged.","fileCount":7,"zipByteSize":19515}]},"execution":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No published capability contract is available yet."},"installCommand":"clawhub skill install s171xgnmqse0nqvgqvqnaq5f9183kyre:k8s-aiops","setupComplexity":"low","setupSteps":["Install using `clawhub skill install s171xgnmqse0nqvgqvqnaq5f9183kyre:k8s-aiops` in an isolated environment before connecting it to live workloads.","No published capability contract is available yet, so validate auth and request/response behavior manually.","Review the upstream CLAWHUB listing at https://clawhub.ai/zw008/k8s-aiops before using production credentials."],"contract":{"contractStatus":"missing","authModes":[],"requires":[],"forbidden":[],"supportsMcp":false,"supportsA2a":false,"supportsStreaming":false,"inputSchemaRef":null,"outputSchemaRef":null,"dataRegion":null,"contractUpdatedAt":null,"sourceUpdatedAt":null,"freshnessSeconds":null},"invocationGuide":{"preferredApi":{"snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-k8s-aiops/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-k8s-aiops/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-k8s-aiops/trust"},"curlExamples":["curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-zw008-k8s-aiops/snapshot\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-zw008-k8s-aiops/contract\"","curl -s \"https://www.xpersona.co/api/v1/agents/clawhub-zw008-k8s-aiops/trust\""],"jsonRequestTemplate":{"query":"summarize this repo","constraints":{"maxLatencyMs":2000,"protocolPreference":["OPENCLEW"]}},"jsonResponseTemplate":{"ok":true,"result":{"summary":"...","confidence":0.9},"meta":{"source":"CLAWHUB","generatedAt":"2026-10-10T10:53:18.741Z"}},"retryPolicy":{"maxAttempts":3,"backoffMs":[500,1500,3500],"retryableConditions":["HTTP_429","HTTP_503","NETWORK_TIMEOUT"]}},"endpoints":{"dossierUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-k8s-aiops/dossier","snapshotUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-k8s-aiops/snapshot","contractUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-k8s-aiops/contract","trustUrl":"https://www.xpersona.co/api/v1/agents/clawhub-zw008-k8s-aiops/trust"}},"reliability":{"evidence":{"source":"runtime-metrics","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No trust, reliability, or runtime telemetry is available."},"trust":{"status":"unavailable","handshakeStatus":"UNKNOWN","verificationFreshnessHours":null,"reputationScore":null,"p95LatencyMs":null,"successRate30d":null,"fallbackRate":null,"attempts30d":null,"trustUpdatedAt":null,"trustConfidence":"unknown","sourceUpdatedAt":null,"freshnessSeconds":null},"decisionGuardrails":{"doNotUseIf":["Contract metadata is missing or unavailable for deterministic execution."],"safeUseWhen":[],"riskFlags":["missing_or_unavailable_contract","trust_data_unavailable","schema_references_missing"],"operationalConfidence":"low"},"executionMetrics":{"observedLatencyMsP50":null,"observedLatencyMsP95":null,"estimatedCostUsd":null,"uptime30d":null,"rateLimitRpm":null,"rateLimitBurst":null,"lastVerifiedAt":null,"verificationSource":null},"runtimeMetrics":{"successRate":null,"avgLatencyMs":null,"avgCostUsd":null,"hallucinationRate":null,"retryRate":null,"disputeRate":null,"p50Latency":null,"p95Latency":null,"lastUpdated":null}},"benchmarks":{"evidence":{"source":"no-benchmark-data","verified":false,"confidence":"low","updatedAt":null,"emptyReason":"No benchmark suites or observed failure patterns are available."},"suites":[],"failurePatterns":[]},"artifacts":{"evidence":{"source":"CLAWHUB","verified":false,"confidence":"medium","updatedAt":"2026-10-10T06:28:09.610Z","emptyReason":null},"readme":"Skill: k8s-aiops\n\nOwner: zw008\n\nSummary: Use this skill whenever the user needs to operate a Kubernetes cluster — list/inspect pods, deployments, statefulsets, daemonsets, replicasets, jobs, cronjobs, services, ingresses, endpoints, configmaps, secrets (names/keys only), PVCs/PVs/storageclasses, nodes, namespaces, and events; read pod logs; describe pods/nodes; pod/node top (metrics); read-only diagnostics / RCA (pod-health, workload-readiness); scale deployments/statefulsets; rollout status/history/undo/pause/resume and set image; delete pods/deployments/jobs; create/delete namespaces; and cordon/uncordon/drain nodes. Works with any kubeconfig-reachable cluster (standard Kubernetes, k3s, EKS, GKE, AKS). Always use this skill for \"list k8s pods\", \"scale deployment\", \"kubernetes pod logs\", \"describe pod\", \"why is my pod crashing\", \"diagnose pods\", \"which deployments are unhealthy\", \"rollout undo\", \"set image\", \"top pods\", \"drain node\", \"cordon node\", \"restart deployment\", \"k3s\", or \"kubectl\"-style tasks when the context is explicitly Kubernetes / a cluster. Do NOT use when the target is not a Kubernetes cluster (hypervisor VM lifecycle, backup products, or cloud-provider consoles are out of scope). Common Kubernetes operations with a built-in governance harness (audit, token budget, undo, risk-tier labels).\n\nTags: latest:0.13.3\n\nVersion history:\n\nv0.13.3 | 2026-09-15T06:02:54.189Z | auto\n\n- Removed the skill description file (skill-card.md).\n- No user-facing changes to functionality or features.\n- Documentation is unchanged except for the file removal.\n\nv0.13.2 | 2026-09-12T14:26:40.832Z | auto\n\n- skill-card.md file removed.\n- SKILL.md updated with new OpenClaw plugin installation instructions (changed plugin repo/user).\n- No code or functional changes to the skill logic itself.\n- Documentation clarifies that the skill-card.md is no longer referenced.\n\nv0.13.1 | 2026-09-12T10:10:28.688Z | auto\n\n- Documentation updated for plugin and OpenClaw usage; added installation info for OpenClaw plugin flow.\n- Explicit instructions for installing as an OpenClaw plugin and visibility settings.\n- Mentions requirement for `uvx` on `PATH` for MCP server install.\n- Removed obsolete skill-card.md file.\n\nv0.13.0 | 2026-09-12T00:59:03.454Z | auto\n\n## k8s-aiops 0.13.0 – Changelog\n\n- Updated environment and binary requirements for improved compatibility (now supports either `k8s-aiops` or `uvx` as valid binaries).\n- Expanded and clarified `metadata.openclaw` fields, adjusting environment variable recommendations.\n- Removed the `skill-card.md` file.\n- No new features or breaking changes to user-facing commands.\n\nv0.12.0 | 2026-08-10T06:51:30.796Z | auto\n\n- Removed the file: skill-card.md\n- No functional or behavior changes to the skill itself\n- Documentation and feature set remain unchanged\n\nv0.11.0 | 2026-08-10T03:58:12.650Z | auto\n\n- Removed the file skill-card.md.\n- No changes to features, functionality, or compatibility.\n\nv0.10.0 | 2026-08-03T05:53:12.435Z | auto\n\n- Removed the file skill-card.md.\n- No changes to functionality or user-facing features.\n- Documentation and usage remain unchanged.\n\nv0.9.0 | 2026-08-02T09:39:56.859Z | auto\n\n- Removed the skill-card.md file from the repository.\n- No functional or behavioral changes to the skill itself.\n- Documentation and core functionality remain unchanged.\n\nv0.8.0 | 2026-07-21T09:41:26.711Z | auto\n\nk8s-aiops v0.8.0\n\n- Updated governance harness: all write tools now use risk-tier labels in audit logs, replacing previous tiering scheme.\n- Revised documentation for clarity: audit and policy guardrails are now described as using token budgets and risk-tier labels.\n- Improved compatibility and skill description: clarified audit labeling and focus on token/runaway budget controls.\n- Removed obsolete documentation file (skill-card.md).\n\nv0.7.0 | 2026-07-20T11:15:39.715Z | auto\n\nk8s-aiops 0.7.0\n\n- Updated documentation in references/capabilities.md.\n- Removed the skill-card.md file.\n\nv0.6.0 | 2026-07-19T03:51:45.442Z | auto\n\nk8s-aiops 0.6.0 introduces diagnostic/RCA tools and expands governance:\n\n- Added read-only diagnostic/RCA tools: pod-health and workload-readiness checks.\n- Increased tool coverage to 55 MCP tools, including new diagnostics.\n- New documentation on agent guardrails (see references/agent-guardrails.md).\n- Improved CLI references and capability docs.\n- Trimmed and clarified skill documentation; removed legacy skill-card.md.\n\nv0.5.0 | 2026-07-17T05:54:35.305Z | auto\n\n- Removed the file: skill-card.md\n- No user-facing features were added or changed in this release.\n- Skill documentation and functionality remain unchanged.\n\nv0.4.0 | 2026-07-16T15:07:54.951Z | auto\n\n- Removed the skill-card.md file.\n- SKILL.md updated; no user-facing feature or behavioral changes described.\n- No changes to capabilities, compatibility, or installation process.\n\nv0.3.0 | 2026-07-13T13:07:40.220Z | auto\n\nk8s-aiops 0.3.0\n\n- Updated documentation in SKILL.md with more details about usage, capabilities, and governance.\n- Expanded descriptions of when and how to use the skill for common Kubernetes operations.\n- Clarified out-of-scope scenarios and related skill routing for user guidance.\n- No changes to code or features—documentation only.\n\nv0.2.0 | 2026-06-27T02:23:13.061Z | auto\n\nk8s-aiops 0.2.0\n\n- Major expansion: increases from 15 to 51 MCP tools, now covering advanced resources and operations.\n- Adds support for listing/inspecting statefulsets, daemonsets, replicasets, jobs, cronjobs, ingresses, endpoints, configmaps, secrets (names/keys only), PVCs/PVs/storageclasses, and additional rollout controls (status, history, undo, pause/resume, set-image).\n- New features: pod/node/namespace describe, top (metrics), and node drain are included.\n- Enhanced namespace, job, and node operations, with reversible write/undo logic improved and high-risk operations clearly tagged.\n- New onboarding wizard (`k8s-aiops init`) for easier context/target setup.\n- Documentation and CLI references updated to reflect new capabilities; deprecated skill-card.md removed.\n\nv0.1.0 | 2026-06-22T06:50:22.002Z | user\n\nv0.1.0 first release: standalone governed Kubernetes ops — 15 MCP tools with built-in audit/budget/undo/risk-tier harness\n\nArchive index:\n\nArchive v0.13.3: 7 files, 19602 bytes\n\nFiles: references/agent-guardrails.md (9366b), references/capabilities.md (6127b), references/cli-reference.md (3495b), references/setup-guide.md (3987b), skill-card.md (2400b), SKILL.md (18486b), _meta.json (129b)\n\nFile v0.13.3:SKILL.md\n\n---\nname: k8s-aiops\nslug: k8s-aiops\ndisplayName: \"k8s AIops\"\nsummary: \"Governed Kubernetes ops — 55 MCP tools with audit, budget, undo, risk-tier audit labels.\"\nlicense: MIT\nhomepage: https://github.com/AIops-tools/K8s-AIops\ntags: [aiops, mcp, governance, k8s]\ndescription: >\n  Use this skill whenever the user needs to operate a Kubernetes cluster — list/inspect pods, deployments, statefulsets, daemonsets, replicasets, jobs, cronjobs, services, ingresses, endpoints, configmaps, secrets (names/keys only), PVCs/PVs/storageclasses, nodes, namespaces, and events; read pod logs; describe pods/nodes; pod/node top (metrics); read-only diagnostics / RCA (pod-health, workload-readiness); scale deployments/statefulsets; rollout status/history/undo/pause/resume and set image; delete pods/deployments/jobs; create/delete namespaces; and cordon/uncordon/drain nodes. Works with any kubeconfig-reachable cluster (standard Kubernetes, k3s, EKS, GKE, AKS).\n  Always use this skill for \"list k8s pods\", \"scale deployment\", \"kubernetes pod logs\", \"describe pod\", \"why is my pod crashing\", \"diagnose pods\", \"which deployments are unhealthy\", \"rollout undo\", \"set image\", \"top pods\", \"drain node\", \"cordon node\", \"restart deployment\", \"k3s\", or \"kubectl\"-style tasks when the context is explicitly Kubernetes / a cluster.\n  Do NOT use when the target is not a Kubernetes cluster (hypervisor VM lifecycle, backup products, or cloud-provider consoles are out of scope).\n  Common Kubernetes operations with a built-in governance harness (audit, token budget, undo, risk-tier labels).\ninstaller:\n  kind: uv\n  package: k8s-aiops\nargument-hint: \"[resource name or describe your Kubernetes task]\"\nallowed-tools:\n  - Bash\nmetadata: {\"openclaw\":{\"requires\":{\"anyBins\":[\"k8s-aiops\",\"uvx\"]},\"optional\":{\"env\":[\"K8S_AIOPS_CONFIG\",\"KUBECONFIG\",\"K8S_AIOPS_HOME\"]},\"homepage\":\"https://github.com/AIops-tools/K8s-AIops\",\"emoji\":\"☸️\",\"os\":[\"macos\",\"linux\"]}}\ncompatibility: >\n  Standalone, self-governed Kubernetes operations. The governance harness (audit, token/runaway budget, undo, risk-tier labels) is bundled in the package — no external skill-family dependency.\n  All write operations are audited to a local SQLite DB under ~/.k8s-aiops/ (relocatable via K8S_AIOPS_HOME).\n  Credentials: k8s-aiops handles NO credentials directly — authentication is delegated to the kubeconfig (KUBECONFIG env or ~/.kube/config), which may hold client certs, bearer tokens, or exec plugins (EKS/GKE/AKS). Underlying credentials are never read, logged, or echoed. The state dir ~/.k8s-aiops should be chmod 700.\n  Destructive operations (deployment/job/namespace delete, node cordon/drain, rollout undo) require double confirmation at the CLI layer and support --dry-run. All write tools pass through the @governed_tool decorator (budget/runaway guard + audit + a descriptive risk-tier label). Reversible writes record an inverse undo descriptor (scale_deployment/scale_statefulset restore the previous replica count; set_deployment_image restores the previous image; cordon_node ↔ uncordon_node and rollout_pause ↔ rollout_resume; create_namespace ↔ delete_namespace); delete_* and rollout_undo record none. risk_level=high: delete_deployment, delete_job, delete_namespace, drain_node, rollout_undo_deployment. Secret VALUES are never read or returned by any tool.\n  Webhooks: none — no outbound network calls beyond the configured Kubernetes API server.\n  TLS: follows the kubeconfig (certificate-authority / insecure-skip-tls-verify); the skill does not weaken it.\n  Transitive dependencies: the official kubernetes Python client and the MCP SDK. No post-install scripts or background services.\n---\n\n# k8s AIops\n\n> **Disclaimer**: This is a community-maintained open-source project and is **not affiliated with, endorsed by, or sponsored by the Cloud Native Computing Foundation, the Kubernetes project, or k3s/Rancher.** \"Kubernetes\" and \"k3s\" are trademarks of their respective owners. Source code is publicly auditable at [github.com/AIops-tools/K8s-AIops](https://github.com/AIops-tools/K8s-AIops) under the MIT license.\n\nGoverned Kubernetes operations — **55 MCP tools**, every one wrapped with the bundled `@governed_tool` harness: a local unified audit log under `~/.k8s-aiops/`, a token/runaway budget guard, undo-token recording, and a descriptive risk-tier label on every audit row. Works with any kubeconfig-reachable cluster (standard Kubernetes, k3s, EKS, GKE, AKS). Run `k8s-aiops init` for a friendly onboarding wizard that registers your kube contexts as named targets.\n\n> **Standalone**: the governance harness is bundled in the package (`k8s_aiops.governance`) — k8s-aiops has no external skill-family dependency. Coverage focuses on common operations and is not yet exhaustive.\n\n## What This Skill Does\n\n| Category | Tools | Count | Read or Write |\n|----------|-------|:-----:|:-------------:|\n| **Pods** | list, get, logs, describe, delete | 5 | 4 read / 1 write |\n| **Deployments** | list, get, scale, rollout restart, delete | 5 | 2 read / 3 write |\n| **Rollout** | status, history, undo, pause, resume, set-image | 6 | 2 read / 4 write |\n| **StatefulSets** | list, get, scale | 3 | 2 read / 1 write |\n| **DaemonSets** | list, get | 2 | 2 read |\n| **ReplicaSets** | list | 1 | 1 read |\n| **Jobs / CronJobs** | job list/get/delete, cronjob list/get | 5 | 4 read / 1 write |\n| **Services / Ingress / Endpoints** | service list, ingress list/get, endpoints list | 4 | 4 read |\n| **Config / Secrets** | configmap list/get, secret list (names/keys only) | 3 | 3 read |\n| **Storage** | pvc list/get, pv list, storageclass list | 4 | 4 read |\n| **Nodes** | list, describe, cordon, uncordon, drain | 5 | 2 read / 3 write |\n| **Namespaces** | list, create, delete | 3 | 1 read / 2 write |\n| **Metrics (top)** | pod, node | 2 | 2 read |\n| **Cluster** | cluster_info, api_resources | 2 | 2 read |\n| **Events** | list | 1 | 1 read |\n| **Diagnostics / RCA** | pod-health, workload-readiness | 2 | 2 read |\n\n## Quick Install\n\n```bash\nuv tool install k8s-aiops\nk8s-aiops init            # friendly wizard: register your kube contexts as targets\nk8s-aiops doctor          # or skip init — works with your current kube-context too\n```\n\nOr as an OpenClaw plugin, which installs this skill and its MCP server together:\n\n```bash\nopenclaw plugins install clawhub:@zw008/k8s-aiops\nopenclaw skills info k8s-aiops          # expect: Visible to model: yes\n```\n\nNeeds `uvx` on `PATH`: the MCP server is fetched with uv, pinned to this release.\n\n## When to Use This Skill\n\n- List/inspect pods, deployments, services, nodes, namespaces and recent events\n- Read a pod's recent log lines to diagnose a crash loop\n- Run a read-only RCA sweep (`diagnose pod-health` / `diagnose workload-readiness`) to find the root cause worst-first\n- Scale a deployment up/down, or trigger a rolling restart\n- Delete a stuck pod (a controller recreates it) or a deployment\n- Cordon a node before maintenance, then uncordon it after\n\n**Do NOT use when** the target is not a Kubernetes cluster (hypervisor VM lifecycle, backup products, or cloud-provider consoles are out of scope for this skill).\n\n## Related Skills — Skill Routing\n\n| If the user wants… | Use |\n|--------------------|-----|\n| Kubernetes pods / deployments / nodes | **k8s-aiops** (this skill) |\n| Hypervisor VM lifecycle (power, snapshot, migrate) | a hypervisor ops skill |\n| Backup & restore | a backup ops skill |\n\n## Common Workflows\n\n### Diagnose a crash-looping pod and restart its deployment\n\n1. `k8s-aiops pod list -n prod` → find the pod with high `restarts` / non-Running `phase`\n2. `k8s-aiops pod logs <pod> -n prod --tail 200` → read the recent logs for the crash cause\n3. `k8s-aiops events -n prod` → check for `FailedScheduling` / image-pull events\n4. `k8s-aiops deployment restart <deploy> -n prod` → roll the deployment after fixing the cause\n5. **Failure branch**: if logs/events show an RBAC `403`, the kube context lacks the verb — run `kubectl auth can-i get pods -n prod` and switch to a context with adequate RBAC; the skill never retries a denied auth.\n\n### Triage an unhealthy namespace with RCA, then act on the worst finding\n\n1. `k8s-aiops diagnose pod-health -n prod` → worst-first findings; a `critical` `CrashLoopBackOff` on `prod/api` cites `restarts=9` and the exact `kubectl logs … --previous` action\n2. `k8s-aiops diagnose workload-readiness -n prod` → confirm the blast radius: e.g. `Deployment web ready 0/3` (`critical`, under-replicated)\n3. `k8s-aiops pod logs api-<hash> -n prod --tail 200 --previous`-equivalent via `k8s-aiops pod describe api-<hash> -n prod` → read the crash cause the RCA pointed you at\n4. `k8s-aiops deployment restart web -n prod` → roll the deployment once the root cause is fixed\n5. **Failure branch**: if `diagnose` returns an RBAC `403`, the kube context cannot list pods/deployments in that namespace — run `kubectl auth can-i list pods -n prod` and switch to a context with adequate RBAC; the RCA tools are read-only and never retry a denied auth.\n\n### Drain a node for maintenance, safely reversible\n\n1. `k8s-aiops node list` → identify the node and confirm it is `Ready`/schedulable\n2. `k8s-aiops node cordon <node> --dry-run` → preview, then `k8s-aiops node cordon <node>` (double confirm) — records an inverse `uncordon_node` undo descriptor\n3. After maintenance: `k8s-aiops node uncordon <node>` → re-enable scheduling\n4. **Failure branch**: if `doctor` shows the cluster unreachable, fix the kubeconfig context (`kubectl config get-contexts`) before retrying — cordon is never issued against an unauthenticated session.\n\n## Usage Mode\n\n| Scenario | Recommended | Why |\n|----------|:-----------:|-----|\n| Local/small models | **CLI** | fewer tokens than MCP |\n| Cloud models (Claude, GPT) | Either | MCP gives structured JSON I/O |\n| Automated pipelines | **MCP** | type-safe parameters, audited |\n\n> **MCP environment caveat**: MCP clients spawn the server with a CLEAN\n> environment — shell exports may not reach it. Set `K8S_AIOPS_HOME`,\n> `K8S_AUDIT_APPROVED_BY`, `K8S_AUDIT_RATIONALE` (and `KUBECONFIG` when the\n> kubeconfig is not at `~/.kube/config`) in the MCP server config's `env`\n> block, not just in your terminal.\n\n## MCP Tools (55 — 39 read, 16 write)\n\n| Category | Tools | R/W |\n|----------|-------|:---:|\n| Pods | `pod_list`, `pod_get`, `pod_logs`, `pod_describe` | Read |\n| | `delete_pod` | Write |\n| Deployments | `deployment_list`, `deployment_get` | Read |\n| | `scale_deployment`, `rollout_restart_deployment`, `delete_deployment` | Write |\n| Rollout | `rollout_status`, `rollout_history` | Read |\n| | `rollout_undo_deployment`, `rollout_pause`, `rollout_resume`, `set_deployment_image` | Write |\n| StatefulSets | `statefulset_list`, `statefulset_get` | Read |\n| | `scale_statefulset` | Write |\n| DaemonSets / ReplicaSets | `daemonset_list`, `daemonset_get`, `replicaset_list` | Read |\n| Jobs / CronJobs | `job_list`, `job_get`, `cronjob_list`, `cronjob_get` | Read |\n| | `delete_job` | Write |\n| Services / Ingress | `service_list`, `ingress_list`, `ingress_get`, `endpoints_list` | Read |\n| Config / Secrets | `configmap_list`, `configmap_get`, `secret_list` (names/keys only) | Read |\n| Storage | `pvc_list`, `pvc_get`, `pv_list`, `storageclass_list` | Read |\n| Nodes | `node_list`, `node_describe` | Read |\n| | `cordon_node`, `uncordon_node`, `drain_node` | Write |\n| Namespaces | `namespace_list` | Read |\n| | `create_namespace`, `delete_namespace` | Write |\n| Metrics (top) | `node_top`, `pod_top` | Read |\n| Cluster | `cluster_info`, `api_resources` | Read |\n| Events | `event_list` | Read |\n| Diagnostics / RCA | `pod_health_rca`, `workload_readiness_rca` | Read |\n| Undo | `undo_list` | Read |\n| | `undo_apply` | Write |\n\n**Security — secrets**: `secret_list` returns secret names, types, and key NAMES only. Secret VALUES are never read, returned, or logged, and there is deliberately no tool that returns secret values.\n\n**Dry-run previews**: every write tool takes `dry_run: bool = False`. A dry run returns a `{\"dryRun\": true, \"wouldX\": ...}` preview without touching the cluster, and no undo descriptor is recorded for a preview.\n\n**Harness features that light up**: write tools with a clean inverse pass an `undo=` lambda so the harness records an inverse descriptor (with `_undo_id`) to the undo store — `scale_deployment`/`scale_statefulset` record a scale-back to their returned `previous_replicas`, `set_deployment_image` records a restore to the captured `previous_image`, `cordon_node` ↔ `uncordon_node` and `rollout_pause` ↔ `rollout_resume` are mutual inverses, and `create_namespace` records a `delete_namespace`. `drain_node` records a partial `uncordon_node` inverse (cordon is reversible; evictions are not). `delete_*` and `rollout_undo_deployment` declare no undo. `risk_level=high`: `delete_deployment`, `delete_job`, `delete_namespace`, `drain_node`, `rollout_undo_deployment`. `undo_list` (read) lists recorded reversible writes whose undo tokens have not been applied yet, and `undo_apply` (write) executes a recorded inverse — itself governed, single-use, and supports `dry_run`. All 55 tools are audit-logged under `~/.k8s-aiops/` and pass through the budget/runaway guard, each recorded with a descriptive risk-tier label. `pod_top`/`node_top` return a clear \"metrics-server not installed\" message (not an error) when metrics-server is absent. Avoid tight poll loops (re-listing pods every second) — the runaway breaker backs this up.\n\n## CLI Quick Reference\n\n```bash\nk8s-aiops init                                            # interactive onboarding wizard\nk8s-aiops pod list [-n <ns>] [-t <target>]\nk8s-aiops pod get <name> [-n <ns>]\nk8s-aiops pod describe <name> [-n <ns>]                   # status, container states, events\nk8s-aiops pod logs <name> [-n <ns>] [--tail 200] [-c <container>]\nk8s-aiops pod delete <name> [-n <ns>] [--dry-run]        # double confirm\nk8s-aiops deployment list|get|scale|restart|delete ...    # scale/restart: single confirm + --dry-run; delete: double confirm\nk8s-aiops rollout status|history|pause|resume <name> [-n <ns>]\nk8s-aiops rollout set-image <name> <container> <image> [-n <ns>]\nk8s-aiops rollout undo <name> [--to-revision N] [--dry-run]   # double confirm\nk8s-aiops statefulset list|get|scale ...\nk8s-aiops daemonset list|get ...\nk8s-aiops job list|get|delete ...                         # delete: double confirm\nk8s-aiops cronjob list|get ...\nk8s-aiops service list [-n <ns>]\nk8s-aiops ingress list|get [-n <ns>]\nk8s-aiops configmap list|get [-n <ns>]\nk8s-aiops secret list [-n <ns>]                           # names/keys only — never values\nk8s-aiops storage pvc-list|pvc-get|pv-list|class-list\nk8s-aiops top pod|node                                    # requires metrics-server\nk8s-aiops node list|describe\nk8s-aiops node cordon|drain <name> [--dry-run]           # double confirm\nk8s-aiops node uncordon <name>\nk8s-aiops namespace list|create\nk8s-aiops namespace delete <name> [--dry-run]            # double confirm\nk8s-aiops cluster-info\nk8s-aiops api-resources\nk8s-aiops events [-n <ns>]\nk8s-aiops diagnose pod-health [-n <ns>] [-l <selector>]         # read-only RCA: crashloop/imagepull/OOM/unschedulable/restarts\nk8s-aiops diagnose workload-readiness [-n <ns>]                 # read-only RCA: ready<desired / stuck rollouts\nk8s-aiops doctor\nk8s-aiops mcp                                             # start MCP server (stdio)\n```\n\nSee `references/cli-reference.md` for the full command list.\n\n## Troubleshooting\n\n### \"Could not load kubeconfig … context not found\"\nThe named context does not exist in your kubeconfig. Run `kubectl config get-contexts` and set the target's `context:` to a listed name (or omit it to use current-context).\n\n### \"Authentication/authorization failed (401/403)\"\nThe kube context lacks the RBAC verb for the resource. Check with `kubectl auth can-i <verb> <resource> -n <ns>` and switch to a context/ServiceAccount with adequate roles. For EKS/GKE/AKS, confirm the exec-plugin (aws/gcloud/az CLI) is installed and logged in.\n\n### \"Resource not found (404)\"\nThe pod/deployment/node name or namespace is wrong, or the object was deleted. List the parent collection first (`pod list`, `deployment list`, `node list`) to get a current name. Remember most commands default to the `default` namespace unless `-n` is given.\n\n### \"Conflict (409)\"\nThe object changed concurrently (or already exists). Re-read it and retry the write.\n\n### Logs are empty or truncated\n`pod logs` returns the trailing `--tail` lines (default 100); raise `--tail`. For a multi-container pod, pass `-c <container>` or the API returns an error naming the available containers.\n\n## Audit & Safety\n\nAll operations are automatically audited via the bundled `@governed_tool` decorator (`k8s_aiops.governance`):\n- Every tool call logged to `~/.k8s-aiops/audit.db` (local SQLite audit DB; relocate with `K8S_AIOPS_HOME`)\n- Budget / runaway guard caps cumulative tool calls and wall-time, and trips on tight poll/retry loops — a safety backstop, not authorization\n- Undo store records inverse descriptors for reversible writes (scale → previous replicas; cordon ↔ uncordon)\n- Each write carries a descriptive risk-tier label into its audit row — a label, not a gate; `K8S_AUDIT_APPROVED_BY` / `K8S_AUDIT_RATIONALE` are optional annotations recorded when set, never required\n\n**Authorization is not this tool's job.** There is no read-only switch, policy file, or approval gate. Whether a write is permitted is the agent's judgement or the RBAC of the kubeconfig context you connect with — give it a read-only ServiceAccount and writes fail at the apiserver, the place that owns the permission.\n\nThe harness is bundled in the package — no external dependency, no manual setup. See `references/setup-guide.md` for security details.\n\nDriving these tools with a smaller / local model? See `references/agent-guardrails.md` — which guardrails the tool now enforces for you, plus a ready-to-paste system prompt.\n\n## Contributing & feature requests\n\nCoverage is intentionally focused. **Missing a device, action, or feature you need?** Open an issue or pull request at [github.com/AIops-tools/K8s-AIops](https://github.com/AIops-tools/K8s-AIops/issues) — feature requests, contributions, and comments are all welcome.\n\n## License\n\nMIT — [github.com/AIops-tools/K8s-AIops](https://github.com/AIops-tools/K8s-AIops)\n\nFile v0.13.3:_meta.json\n\n{\n  \"ownerId\": \"kn7b067awq2s97bn3d7p5qfhw5827pxc\",\n  \"slug\": \"k8s-aiops\",\n  \"version\": \"0.13.3\",\n  \"publishedAt\": 1789452174189\n}\n\nFile v0.13.3:references/agent-guardrails.md\n\n# Agent guardrails — running k8s-aiops with a smaller / local model\n\nIf you drive these tools with a local model (Llama, Qwen, Mistral … via Goose,\nOllama, LM Studio, or any OpenAI-compatible runtime), you will get noticeably\nbetter results with a short system prompt. This page gives you one, and — more\nimportantly — tells you which guardrails you **no longer need to write**, because\nthe tool now enforces them itself.\n\nThe distinction matters. A guardrail in a prompt is a request. A guardrail in the\nharness is a guarantee. Anything below that we could move into the harness, we did.\n\n## Authorization is not this tool's job — decide it where it belongs\n\nWhether a write should happen is your decision, or the account's. The tool does\nnot gate it — there is no read-only switch and no approval prompt to configure.\nThe two right places to control read vs write:\n\n- **The kubeconfig context you connect with.** Bind it to a ServiceAccount or\n  user whose RBAC grants only `get`/`list`/`watch`, and every write fails at the\n  apiserver — the only place the permission actually lives. No skill-side flag\n  can be argued around by a model, but a revoked RBAC verb cannot be. This is\n  strictly stronger than any in-process switch: it is enforced at the cluster.\n- **Your agent's system prompt.** If you want an observe-only session, tell the\n  model not to call the write tools (they are clearly tagged `[WRITE]`).\n\nWhat the tool *does* guarantee is that you can always see what happened:\n\n## What the tool enforces — do not waste prompt budget on these\n\n| You might be tempted to prompt | Why you don't need to |\n|---|---|\n| \"Log everything you do, over both MCP and the CLI\" | Every operation is audited to `~/.k8s-aiops/audit.db` regardless of what the model says it did — and the CLI writes the same row the MCP path does, so there is no unaudited entry point. Reversible writes also record an undo token capturing the *prior* state. |\n| \"Don't invent a value when a field is missing\" | A field the apiserver did not return comes back as `null`, never as `\"\"`. An unscheduled pod's `node` is `null`; a pod with no phase yet has `phase: null`; an object with no readable creation timestamp has `age: null`. Absent and empty are distinguishable in the payload. |\n| \"Tell me if the output was cut off\" | The limit-bearing reads return an envelope: `event_list` → `{\"events\": [...], \"returned\": N, \"limit\": L, \"truncated\": true/false}`, and `undo_list` → `{\"undos\": [...], \"returned\": N, \"limit\": L, \"truncated\": ...}`. Truncation is **measured** (one extra row is fetched), not guessed from `len(rows) == limit`. |\n| \"Preserve the ordering / tell me what's most urgent\" | `pod_health_rca` and `workload_readiness_rca` findings carry an explicit 1-based `rank`, worst-first, and each finding's `detail` cites the measured signal (the waiting reason, the restart count, the ready/desired ratio). Priority is in the payload, not implied by list position. |\n| \"Confirm before anything destructive\" | The destructive CLI commands (deployment/job/namespace delete, node cordon/drain, rollout undo) are `--dry-run`-able and require double confirmation. |\n| \"Don't get stuck retrying\" | The runaway guard trips a circuit breaker if the same call is hammered in a tight loop — a stuck agent is stopped rather than left to burn calls and time. It is a safety backstop, not an authorization gate. |\n\n## What still needs a prompt\n\nThese are model-behaviour problems the harness cannot fix from the outside.\nCopy this into your agent's system prompt:\n\n```text\nYou operate a Kubernetes cluster through the k8s-aiops MCP tools.\n\nTOOL USE\n- Before answering any question about the current cluster, you MUST call a\n  tool. Never answer from memory or assumption.\n- Actually invoke the tool. Do not describe the call you would make, and do not\n  emit an example JSON response in place of calling it.\n- If a tool call fails, report the real error verbatim. Never fill the gap with\n  a plausible-sounding answer.\n\nREADING RESULTS\n- Read the whole result before concluding. If a result contains a \"truncated\"\n  field that is true, say so and re-run with a higher limit instead of treating\n  the partial result as complete.\n- A null field means the apiserver did not return that value. Report it as \"not\n  available\" — never infer it. In particular, a null \"node\" means the pod is\n  not scheduled; it does not mean the node is unknown or missing.\n- Report values exactly as returned. Do not normalise, translate, or prettify\n  phases, conditions, container-state reasons, or resource names.\n- When an RCA result has findings, work in \"rank\" order and cite the measured\n  number in each finding's \"detail\".\n\nSCOPE AND IDENTIFIERS\n- Always state the namespace you are talking about. A bare pod or deployment\n  name is ambiguous — the same name exists in many namespaces.\n- Omitting the namespace means ALL namespaces, not the default one. Never\n  silently widen a namespaced question into a cluster-wide answer.\n- Do not confuse a namespace with a context (a cluster), a pod name with its\n  deployment name, or a container name with the pod that contains it. A pod\n  name generated by a ReplicaSet is not a stable identifier.\n- Separate observation from interpretation. State what the tools returned, then\n  any interpretation, clearly marked as such.\n- Do not assert a capacity, performance, or availability problem unless a tool\n  result supports it. Do not add generic advice that does not follow from the\n  tool output.\n```\n\n## Recommended setup for a local model\n\nStart with a kubeconfig context that *cannot* write — a ServiceAccount whose\nRBAC grants only `get`/`list`/`watch` — verify, and widen its permission only\nwhen you trust the setup. RBAC is enforced at the apiserver, so a write fails at\nthe cluster no matter what the model attempts:\n\n```bash\n# Point KUBECONFIG at a read-only ServiceAccount context, then:\nk8s-aiops doctor\n```\n\nOptionally annotate the audit trail with who is operating and why — recorded on\nevery row, never required:\n\n```bash\nexport K8S_AUDIT_APPROVED_BY=\"your.name@example.com\"\nexport K8S_AUDIT_RATIONALE=\"incident INC-1234 — restart the stuck payments rollout\"\n```\n\n## Kubernetes-specific notes\n\n- **Namespace confusion is the single most common failure mode.** Every\n  namespaced tool takes an optional `namespace`; omitting it lists across **all**\n  namespaces. A smaller model reliably drops the namespace from a follow-up call\n  and then reports a cluster-wide count as if it were the namespace's. Pin the\n  namespace in your prompt (or in the target's `namespace:` in\n  `~/.k8s-aiops/config.yaml`) and ask the model to restate it in every answer.\n- **Context confusion is the expensive one.** `target` selects a *cluster*, not a\n  namespace. If you have prod and staging contexts in the same kubeconfig, give\n  the agent a config with only the cluster it should touch — the model has no way\n  to notice it is on the wrong one, because both look plausible.\n- **Authentication is delegated to your kubeconfig.** Unlike the other tools in\n  this line, k8s-aiops has **no credential store**: there is nothing to encrypt\n  and no master password. (`k8s-aiops secret ...` lists Kubernetes *Secret\n  resources* by name — it is not a credential manager.) The agent inherits the RBAC of the\n  kubeconfig user. That makes RBAC your strongest guardrail — a read-only\n  ServiceAccount kubeconfig enforces read-only at the *cluster*, the only place\n  the permission truly lives.\n- **`drain_node` is the most dangerous tool here.** It evicts pods cluster-wide\n  in effect, and its blast radius is not visible in its arguments. Keep it out of\n  the RBAC role you hand the agent unless you specifically intend node drains.\n- **Pod names are not stable.** A ReplicaSet-generated pod name changes on every\n  rollout, so an id the model cached earlier in the conversation may already be\n  gone. Prefer `deployment_get` / `rollout_status` over re-using a pod name.\n- **Prefer the RCA tools over multi-step chains.** `pod_health_rca` and\n  `workload_readiness_rca` do the list-then-correlate work inside one call, so a\n  smaller model does not have to chain reads and keep names/namespaces straight.\n- **Secrets are names-only by design.** `secret_list` returns key *names*, never\n  values — there is no tool that can exfiltrate a secret payload.\n\n## If your model still struggles\n\nSome behaviours are model-capacity limits rather than prompt problems:\n\n- **Multi-tool workflows time out or drift.** Prefer the `*_rca` tools — they do\n  the multi-step correlation inside one call.\n- **The model ignores later tool results in a long context.** Ask narrower\n  questions, always scope to a namespace, and use `limit` deliberately rather\n  than pulling whole-cluster inventories (`pod_list` with no namespace on a busy\n  cluster will bury everything else in the context).\n- **The model describes calls instead of making them.** This is usually a\n  runtime/tool-calling-format mismatch, not a prompt problem — check that your\n  client advertises the tools in the format your model was trained on.\n\nFeedback on running this with a specific local model is genuinely useful —\nopen an issue at\n[github.com/AIops-tools/K8s-AIops](https://github.com/AIops-tools/K8s-AIops/issues)\nwith the model, runtime, and what went wrong.\n\nFile v0.13.3:references/capabilities.md\n\n# k8s-aiops Capabilities\n\n55 MCP tools (39 read / 16 write). Every tool is wrapped with `@governed_tool`\n(audit + policy + budget + risk-tier; undo where a clean inverse exists). Returns\nare high-signal summaries — `_get` / `_describe` tools add detail for a single object.\n\n## Read tools\n\n| Tool | Returns | Risk |\n|------|---------|:----:|\n| `pod_list` | name, namespace, phase, ready, restarts, node, age | low |\n| `pod_get` | + host_ip, pod_ip, containers | low |\n| `pod_describe` | status, conditions, container states + restart counts, recent events | low |\n| `pod_logs` | trailing log lines (default 100) | low |\n| `deployment_list` / `deployment_get` | replicas summary / + strategy, images | low |\n| `rollout_status` | desired/updated/available/unavailable + paused | low |\n| `rollout_history` | revisions (from replicasets) with images | low |\n| `statefulset_list` / `statefulset_get` | desired/ready/current / + service, images | low |\n| `daemonset_list` / `daemonset_get` | desired/ready/available / + images | low |\n| `replicaset_list` | name, namespace, desired/ready, age | low |\n| `job_list` / `job_get` | completions, succeeded/failed/active | low |\n| `cronjob_list` / `cronjob_get` | schedule, suspend, active, last schedule | low |\n| `service_list` | name, namespace, type, cluster IP, ports | low |\n| `ingress_list` / `ingress_get` | class, hosts / + path→backend rules | low |\n| `endpoints_list` | ready addresses + ports | low |\n| `configmap_list` / `configmap_get` | key count / keys + values | low |\n| `secret_list` | names, types, key NAMES only (values redacted) | low |\n| `pvc_list` / `pvc_get` | status, capacity, class / + access modes | low |\n| `pv_list` | capacity, status, claim, class | low |\n| `storageclass_list` | provisioner, reclaim policy, default | low |\n| `node_list` / `node_describe` | status, roles / capacity, allocatable, conditions, taints | low |\n| `namespace_list` | name, phase, age | low |\n| `pod_top` / `node_top` | CPU/mem via metrics-server (graceful if absent) | low |\n| `cluster_info` | server version, node/ready/namespace counts | low |\n| `api_resources` | available API groups + versions | low |\n| `event_list` | type, reason, object, namespace, message, age | low |\n| `pod_health_rca` | worst-first findings: CrashLoopBackOff, image-pull, OOMKilled, unschedulable, high restarts (each cites the reason/count) | low |\n| `workload_readiness_rca` | worst-first findings: ready<desired, zero-ready outages, stuck rollouts (Deployment/StatefulSet/DaemonSet) | low |\n| `undo_list` | recorded reversible writes / not-yet-applied undo tokens | low |\n\n## Write tools\n\n| Tool | Effect | Risk | Undo |\n|------|--------|:----:|------|\n| `scale_deployment` | set replica count | medium | scale back to `previous_replicas` |\n| `scale_statefulset` | set replica count | medium | scale back to `previous_replicas` |\n| `rollout_restart_deployment` | patch `restartedAt` annotation | medium | none (pods already rolling) |\n| `rollout_pause` / `rollout_resume` | toggle `spec.paused` | medium | each other |\n| `rollout_undo_deployment` | roll back to a prior revision | **high** | none |\n| `set_deployment_image` | update a container image | medium | restore `previous_image` |\n| `delete_pod` | delete a pod | medium | none (controller recreates) |\n| `delete_deployment` | delete deployment + pods | **high** | none |\n| `delete_job` | delete a job + pods | **high** | none |\n| `create_namespace` | create a namespace | medium | `delete_namespace` |\n| `delete_namespace` | delete namespace + everything in it | **high** | none |\n| `cordon_node` / `uncordon_node` | toggle schedulability | medium | each other |\n| `drain_node` | cordon + evict pods (skips DaemonSet/mirror) | **high** | partial: `uncordon_node` |\n| `undo_apply` | execute a recorded inverse (itself governed, single-use, supports `dry_run`) | medium | n/a (is the undo) |\n\n### `delete_namespace` refusals\n\nTwo targets are refused before anything is deleted, on the real call and on the\n`dry_run` preview alike (a preview never green-lights a delete that would be\nrejected):\n\n- **Control-plane namespaces** — `kube-system`, `kube-public`, `kube-node-lease`.\n  Deleting one takes CoreDNS, kube-proxy, the CNI and node heartbeats with it.\n  Pass `confirm=true` (CLI `--confirm`) to proceed deliberately, e.g. when\n  tearing a cluster down or clearing a namespace stuck `Terminating`.\n- **The namespace holding this target's own ServiceAccount credential** — that\n  delete revokes the credential it is running on, and `delete_namespace` has no\n  undo. **`confirm` does not override this one.** Re-run from a context whose\n  credential lives elsewhere. Only applies when the kubeconfig authenticates\n  with a ServiceAccount token; certificate and `exec` (EKS/GKE/AKS) credentials\n  are not namespace-bound, so the check stands down rather than guessing.\n\n## Token-budget notes\n\n- List tools accept a `namespace` filter to keep responses small; events and pod\n  listings also accept `limit` / `label_selector` where applicable.\n- Prefer `pod_get` / `deployment_get` over re-listing when you already have a name.\n- The runaway guard trips on tight poll loops — wait between repeated list calls.\n\n## Design notes / Kubernetes-client assumptions\n\n- Authentication is delegated to the kubeconfig; the skill never touches raw\n  credentials (works with client certs, tokens, and EKS/GKE/AKS exec plugins).\n- Typed Api clients (`CoreV1Api`, `AppsV1Api`, `BatchV1Api`, `NetworkingV1Api`,\n  `StorageV1Api`, `CustomObjectsApi`, `ApisApi`, `VersionApi`) are cached per kube\n  context in a module dict — third-party client objects are never monkey-patched.\n- `secret_list` reads only key NAMES from `secret.data` — secret values are never\n  read, returned, or logged, and no tool exposes them.\n- `pod_top` / `node_top` use the `metrics.k8s.io/v1beta1` API; when metrics-server\n  is absent the 404/503 is caught and returned as `{available: false, message}`.\n- `ApiException` is translated centrally at the connection layer into a teaching\n  `K8sApiError` (404/403/409/5xx), so agents see actionable messages, not tracebacks.\n\nFile v0.13.3:references/cli-reference.md\n\n# k8s-aiops CLI Reference\n\nAll commands accept `-t/--target <name>` to select a configured target (a kube\ncontext). Namespaced commands accept `-n/--namespace <ns>`; omit it to use the\ntarget's default namespace (read lists fall back to all-namespaces).\n\n## Onboarding\n\n```bash\nk8s-aiops init                    # interactive wizard: register kube contexts as targets\n```\n\n## Pods\n\n```bash\nk8s-aiops pod list [-n <ns>] [-t <target>]\nk8s-aiops pod get <name> [-n <ns>]\nk8s-aiops pod describe <name> [-n <ns>]              # status, container states, events\nk8s-aiops pod logs <name> [-n <ns>] [--tail N] [-c <container>]\nk8s-aiops pod delete <name> [-n <ns>] [--dry-run]    # destructive: double confirm\n```\n\n## Deployments & Rollouts\n\n```bash\nk8s-aiops deployment list [-n <ns>]\nk8s-aiops deployment get <name> [-n <ns>]\nk8s-aiops deployment scale <name> <replicas> [-n <ns>]\nk8s-aiops deployment restart <name> [-n <ns>]        # rolling restart\nk8s-aiops deployment delete <name> [-n <ns>] [--dry-run]   # HIGH RISK: double confirm\nk8s-aiops rollout status <name> [-n <ns>]\nk8s-aiops rollout history <name> [-n <ns>]\nk8s-aiops rollout pause|resume <name> [-n <ns>]\nk8s-aiops rollout set-image <name> <container> <image> [-n <ns>]\nk8s-aiops rollout undo <name> [--to-revision N] [--dry-run]   # HIGH RISK: double confirm\n```\n\n## StatefulSets / DaemonSets / Jobs / CronJobs\n\n```bash\nk8s-aiops statefulset list|get [-n <ns>]\nk8s-aiops statefulset scale <name> <replicas> [-n <ns>]\nk8s-aiops daemonset list|get [-n <ns>]\nk8s-aiops job list|get [-n <ns>]\nk8s-aiops job delete <name> [-n <ns>] [--dry-run]    # destructive: double confirm\nk8s-aiops cronjob list|get [-n <ns>]\n```\n\n## Services, Ingress, Config, Storage\n\n```bash\nk8s-aiops service list [-n <ns>]\nk8s-aiops ingress list|get [-n <ns>]\nk8s-aiops configmap list|get [-n <ns>]\nk8s-aiops secret list [-n <ns>]                      # names + key NAMES only, never values\nk8s-aiops storage pvc-list|pvc-get [-n <ns>]\nk8s-aiops storage pv-list|class-list\n```\n\n## Nodes & Metrics\n\n```bash\nk8s-aiops node list\nk8s-aiops node describe <name>\nk8s-aiops node cordon <name> [--dry-run]             # destructive: double confirm\nk8s-aiops node uncordon <name>\nk8s-aiops node drain <name> [--dry-run]              # HIGH RISK: double confirm\nk8s-aiops top pod|node                               # requires metrics-server\n```\n\n## Namespaces, Cluster & Events\n\n```bash\nk8s-aiops namespace list\nk8s-aiops namespace create <name>\nk8s-aiops namespace delete <name> [--dry-run]        # HIGH RISK: double confirm\nk8s-aiops cluster-info\nk8s-aiops api-resources\nk8s-aiops events [-n <ns>]\n```\n\n## Diagnostics & MCP\n\n```bash\nk8s-aiops diagnose pod-health [-n <ns>] [-l <selector>]  # RCA: crashloop/imagepull/OOM/unschedulable/restarts (read-only)\nk8s-aiops diagnose workload-readiness [-n <ns>]          # RCA: ready<desired / stuck rollouts (read-only)\nk8s-aiops doctor [--skip-auth]    # check config + cluster reachability\nk8s-aiops mcp                     # start the MCP server over stdio\n```\n\n## Flags summary\n\n| Flag | Meaning |\n|------|---------|\n| `-t, --target` | Target name from `~/.k8s-aiops/config.yaml` |\n| `-n, --namespace` | Namespace scope |\n| `--tail` | Trailing log lines (pod logs, default 100) |\n| `-c, --container` | Container name (pod logs) |\n| `--dry-run` | Preview a destructive op without executing |\n| `--to-revision` | Rollout revision (`rollout undo`, 0 = previous) |\n| `--skip-auth` | Skip the connectivity check in `doctor` |\n\nFile v0.13.3:references/setup-guide.md\n\n# k8s-aiops Setup Guide\n\n## Install\n\n```bash\nuv tool install k8s-aiops\nk8s-aiops doctor\n```\n\n`k8s-aiops` requires Python ≥ 3.11. If `uv` picked an older interpreter:\n\n```bash\nuv python install 3.12\nuv tool install --python 3.12 --force k8s-aiops\n```\n\n## Connecting to a cluster\n\nk8s-aiops reads your **kubeconfig** — the same one `kubectl` uses\n(`KUBECONFIG` env var, else `~/.kube/config`). Out of the box it talks to your\ncurrent kube-context, so a fresh install needs no config file:\n\n```bash\nk8s-aiops pod list\n```\n\n### Named targets (multiple clusters)\n\nThe fastest way is the interactive wizard, which discovers the contexts in your\nkubeconfig and registers the ones you pick (writing `config.yaml`, dir chmod 700):\n\n```bash\nk8s-aiops init\n```\n\nOr create `~/.k8s-aiops/config.yaml` by hand to give contexts friendly names:\n\n```yaml\ntargets:\n  - name: prod\n    context: prod-eks            # a context from `kubectl config get-contexts`\n    namespace: default           # optional default namespace\n    # kubeconfig: /path/to/kubeconfig   # optional, overrides KUBECONFIG/~/.kube/config\n  - name: lab\n    context: k3s-lab\n```\n\nThen select with `-t`:\n\n```bash\nk8s-aiops -t prod pod list -n payments\n```\n\nNo secrets are stored here — authentication lives entirely in the kubeconfig.\n\n### Works with\n\nStandard Kubernetes, k3s, EKS, GKE, AKS, kind, minikube. For managed clusters\n(EKS/GKE/AKS) the kubeconfig uses an exec plugin (`aws`/`gcloud`/`az`); make sure\nthat CLI is installed and logged in.\n\n## Security\n\n> **Disclaimer**: This is a community-maintained open-source project and is **not affiliated with, endorsed by, or sponsored by the Cloud Native Computing Foundation, the Kubernetes project, or k3s/Rancher.** \"Kubernetes\" and \"k3s\" are trademarks of their respective owners. Source is auditable at [github.com/AIops-tools/K8s-AIops](https://github.com/AIops-tools/K8s-AIops) under the MIT license.\n\n1. **Source code** — [github.com/AIops-tools/K8s-AIops](https://github.com/AIops-tools/K8s-AIops), MIT.\n2. **Config file contents** — `config.yaml` holds only target names, kube context\n   names, and optional namespaces/paths. No credentials.\n3. **Credentials** — delegated to the kubeconfig; never read, logged, or echoed by\n   k8s-aiops. Keep `~/.k8s-aiops` owner-only (`chmod 700`).\n4. **TLS verification** — follows the kubeconfig (`certificate-authority` /\n   `insecure-skip-tls-verify`); the skill does not weaken it.\n5. **Prompt-injection protection** — all API-returned text (names, log lines, event\n   messages) is run through `sanitize()` (truncation + control-character stripping).\n6. **Least privilege** — bind the kube context to a ServiceAccount/user with only the\n   RBAC verbs you need (read-only needs `get`/`list`/`watch`; writes add\n   `patch`/`delete`).\n\n## Governance harness\n\nBundled under `k8s_aiops.governance` — no external dependency. State lives under\n`~/.k8s-aiops/` (override with `K8S_AIOPS_HOME`):\n\n- `audit.db` — every tool call (skill, tool, params, status, duration, agent),\n  each carrying a descriptive risk-tier label derived from the tool's\n  `risk_level`. The tier is a label, not a gate.\n- Token/runaway budget guard (`K8S_MAX_TOOL_CALLS`, `K8S_MAX_TOOL_SECONDS`,\n  `K8S_RUNAWAY_MAX`, `K8S_RUNAWAY_WINDOW_SEC`) — a safety backstop, not\n  authorization.\n- Undo store — inverse descriptors for reversible writes.\n- Accountability: `K8S_AUDIT_APPROVED_BY` / `K8S_AUDIT_RATIONALE` are optional\n  annotations recorded on the audit row when set — never required, never\n  blocking. Authorization is the RBAC of the kubeconfig context, not this tool.\n\n## MCP client config\n\n```jsonc\n{\n  \"command\": \"k8s-aiops\",\n  \"args\": [\"mcp\"],\n  \"env\": { \"K8S_AIOPS_CONFIG\": \"~/.k8s-aiops/config.yaml\" }\n}\n```\n\nFallback (no `uv tool install`): `uvx --from k8s-aiops k8s-aiops-mcp`. Prefer the\ninstalled entry point — it does not re-resolve PyPI at launch.\n\n## Static analysis\n\n```bash\nuvx bandit -r k8s_aiops/ mcp_server/\n```\n\nFile v0.13.3:skill-card.md\n\n## Description:\n\nk8s-aiops helps agents inspect, diagnose, and operate kubeconfig-reachable Kubernetes clusters through governed CLI and MCP workflows.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[zw008](https://clawhub.ai/user/zw008)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers and platform engineers use this skill to inspect Kubernetes resources, read logs, run RCA, and perform governed operational actions such as scaling, rollouts, deletes, namespace changes, and node maintenance.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: The skill can make destructive Kubernetes changes using the access granted by the active kubeconfig.\n\nMitigation: Use a dedicated least-privilege Kubernetes context, preferably read-only by default, and avoid production admin kubeconfigs unless writes are intended.\n\nRisk: MCP write tools can perform real cluster changes even when audit, dry-run, and undo features are available.\n\nMitigation: Review planned write actions, use dry-run where available, confirm the target context and namespace, and rely on Kubernetes RBAC as the enforcement boundary.\n\nRisk: The package installed at runtime may differ from the reviewed release if it is not pinned or verified.\n\nMitigation: Pin and verify the k8s-aiops package version before installation or execution.\n\n## Reference(s):\n\n- [k8s-aiops source repository](https://github.com/AIops-tools/K8s-AIops)\n- [Capabilities](references/capabilities.md)\n- [CLI Reference](references/cli-reference.md)\n- [Setup Guide](references/setup-guide.md)\n- [Agent Guardrails](references/agent-guardrails.md)\n\n## Skill Output:\n\n**Output Type(s):** [text, markdown, shell commands, configuration, guidance]\n\n**Output Format:** [Markdown with inline shell commands and Kubernetes operation guidance]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [May guide Kubernetes read and write operations through k8s-aiops; effects depend on the provided kubeconfig, RBAC, and selected dry-run or confirmation path.]\n\n## Skill Version(s):\n\n0.13.3 (source: server release metadata)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nArchive v0.13.2: 7 files, 19696 bytes\n\nFiles: references/agent-guardrails.md (9366b), references/capabilities.md (6127b), references/cli-reference.md (3495b), references/setup-guide.md (3987b), skill-card.md (2700b), SKILL.md (18486b), _meta.json (129b)\n\nFile v0.13.2:SKILL.md\n\n---\nname: k8s-aiops\nslug: k8s-aiops\ndisplayName: \"k8s AIops\"\nsummary: \"Governed Kubernetes ops — 55 MCP tools with audit, budget, undo, risk-tier audit labels.\"\nlicense: MIT\nhomepage: https://github.com/AIops-tools/K8s-AIops\ntags: [aiops, mcp, governance, k8s]\ndescription: >\n  Use this skill whenever the user needs to operate a Kubernetes cluster — list/inspect pods, deployments, statefulsets, daemonsets, replicasets, jobs, cronjobs, services, ingresses, endpoints, configmaps, secrets (names/keys only), PVCs/PVs/storageclasses, nodes, namespaces, and events; read pod logs; describe pods/nodes; pod/node top (metrics); read-only diagnostics / RCA (pod-health, workload-readiness); scale deployments/statefulsets; rollout status/history/undo/pause/resume and set image; delete pods/deployments/jobs; create/delete namespaces; and cordon/uncordon/drain nodes. Works with any kubeconfig-reachable cluster (standard Kubernetes, k3s, EKS, GKE, AKS).\n  Always use this skill for \"list k8s pods\", \"scale deployment\", \"kubernetes pod logs\", \"describe pod\", \"why is my pod crashing\", \"diagnose pods\", \"which deployments are unhealthy\", \"rollout undo\", \"set image\", \"top pods\", \"drain node\", \"cordon node\", \"restart deployment\", \"k3s\", or \"kubectl\"-style tasks when the context is explicitly Kubernetes / a cluster.\n  Do NOT use when the target is not a Kubernetes cluster (hypervisor VM lifecycle, backup products, or cloud-provider consoles are out of scope).\n  Common Kubernetes operations with a built-in governance harness (audit, token budget, undo, risk-tier labels).\ninstaller:\n  kind: uv\n  package: k8s-aiops\nargument-hint: \"[resource name or describe your Kubernetes task]\"\nallowed-tools:\n  - Bash\nmetadata: {\"openclaw\":{\"requires\":{\"anyBins\":[\"k8s-aiops\",\"uvx\"]},\"optional\":{\"env\":[\"K8S_AIOPS_CONFIG\",\"KUBECONFIG\",\"K8S_AIOPS_HOME\"]},\"homepage\":\"https://github.com/AIops-tools/K8s-AIops\",\"emoji\":\"☸️\",\"os\":[\"macos\",\"linux\"]}}\ncompatibility: >\n  Standalone, self-governed Kubernetes operations. The governance harness (audit, token/runaway budget, undo, risk-tier labels) is bundled in the package — no external skill-family dependency.\n  All write operations are audited to a local SQLite DB under ~/.k8s-aiops/ (relocatable via K8S_AIOPS_HOME).\n  Credentials: k8s-aiops handles NO credentials directly — authentication is delegated to the kubeconfig (KUBECONFIG env or ~/.kube/config), which may hold client certs, bearer tokens, or exec plugins (EKS/GKE/AKS). Underlying credentials are never read, logged, or echoed. The state dir ~/.k8s-aiops should be chmod 700.\n  Destructive operations (deployment/job/namespace delete, node cordon/drain, rollout undo) require double confirmation at the CLI layer and support --dry-run. All write tools pass through the @governed_tool decorator (budget/runaway guard + audit + a descriptive risk-tier label). Reversible writes record an inverse undo descriptor (scale_deployment/scale_statefulset restore the previous replica count; set_deployment_image restores the previous image; cordon_node ↔ uncordon_node and rollout_pause ↔ rollout_resume; create_namespace ↔ delete_namespace); delete_* and rollout_undo record none. risk_level=high: delete_deployment, delete_job, delete_namespace, drain_node, rollout_undo_deployment. Secret VALUES are never read or returned by any tool.\n  Webhooks: none — no outbound network calls beyond the configured Kubernetes API server.\n  TLS: follows the kubeconfig (certificate-authority / insecure-skip-tls-verify); the skill does not weaken it.\n  Transitive dependencies: the official kubernetes Python client and the MCP SDK. No post-install scripts or background services.\n---\n\n# k8s AIops\n\n> **Disclaimer**: This is a community-maintained open-source project and is **not affiliated with, endorsed by, or sponsored by the Cloud Native Computing Foundation, the Kubernetes project, or k3s/Rancher.** \"Kubernetes\" and \"k3s\" are trademarks of their respective owners. Source code is publicly auditable at [github.com/AIops-tools/K8s-AIops](https://github.com/AIops-tools/K8s-AIops) under the MIT license.\n\nGoverned Kubernetes operations — **55 MCP tools**, every one wrapped with the bundled `@governed_tool` harness: a local unified audit log under `~/.k8s-aiops/`, a token/runaway budget guard, undo-token recording, and a descriptive risk-tier label on every audit row. Works with any kubeconfig-reachable cluster (standard Kubernetes, k3s, EKS, GKE, AKS). Run `k8s-aiops init` for a friendly onboarding wizard that registers your kube contexts as named targets.\n\n> **Standalone**: the governance harness is bundled in the package (`k8s_aiops.governance`) — k8s-aiops has no external skill-family dependency. Coverage focuses on common operations and is not yet exhaustive.\n\n## What This Skill Does\n\n| Category | Tools | Count | Read or Write |\n|----------|-------|:-----:|:-------------:|\n| **Pods** | list, get, logs, describe, delete | 5 | 4 read / 1 write |\n| **Deployments** | list, get, scale, rollout restart, delete | 5 | 2 read / 3 write |\n| **Rollout** | status, history, undo, pause, resume, set-image | 6 | 2 read / 4 write |\n| **StatefulSets** | list, get, scale | 3 | 2 read / 1 write |\n| **DaemonSets** | list, get | 2 | 2 read |\n| **ReplicaSets** | list | 1 | 1 read |\n| **Jobs / CronJobs** | job list/get/delete, cronjob list/get | 5 | 4 read / 1 write |\n| **Services / Ingress / Endpoints** | service list, ingress list/get, endpoints list | 4 | 4 read |\n| **Config / Secrets** | configmap list/get, secret list (names/keys only) | 3 | 3 read |\n| **Storage** | pvc list/get, pv list, storageclass list | 4 | 4 read |\n| **Nodes** | list, describe, cordon, uncordon, drain | 5 | 2 read / 3 write |\n| **Namespaces** | list, create, delete | 3 | 1 read / 2 write |\n| **Metrics (top)** | pod, node | 2 | 2 read |\n| **Cluster** | cluster_info, api_resources | 2 | 2 read |\n| **Events** | list | 1 | 1 read |\n| **Diagnostics / RCA** | pod-health, workload-readiness | 2 | 2 read |\n\n## Quick Install\n\n```bash\nuv tool install k8s-aiops\nk8s-aiops init            # friendly wizard: register your kube contexts as targets\nk8s-aiops doctor          # or skip init — works with your current kube-context too\n```\n\nOr as an OpenClaw plugin, which installs this skill and its MCP server together:\n\n```bash\nopenclaw plugins install clawhub:@zw008/k8s-aiops\nopenclaw skills info k8s-aiops          # expect: Visible to model: yes\n```\n\nNeeds `uvx` on `PATH`: the MCP server is fetched with uv, pinned to this release.\n\n## When to Use This Skill\n\n- List/inspect pods, deployments, services, nodes, namespaces and recent events\n- Read a pod's recent log lines to diagnose a crash loop\n- Run a read-only RCA sweep (`diagnose pod-health` / `diagnose workload-readiness`) to find the root cause worst-first\n- Scale a deployment up/down, or trigger a rolling restart\n- Delete a stuck pod (a controller recreates it) or a deployment\n- Cordon a node before maintenance, then uncordon it after\n\n**Do NOT use when** the target is not a Kubernetes cluster (hypervisor VM lifecycle, backup products, or cloud-provider consoles are out of scope for this skill).\n\n## Related Skills — Skill Routing\n\n| If the user wants… | Use |\n|--------------------|-----|\n| Kubernetes pods / deployments / nodes | **k8s-aiops** (this skill) |\n| Hypervisor VM lifecycle (power, snapshot, migrate) | a hypervisor ops skill |\n| Backup & restore | a backup ops skill |\n\n## Common Workflows\n\n### Diagnose a crash-looping pod and restart its deployment\n\n1. `k8s-aiops pod list -n prod` → find the pod with high `restarts` / non-Running `phase`\n2. `k8s-aiops pod logs <pod> -n prod --tail 200` → read the recent logs for the crash cause\n3. `k8s-aiops events -n prod` → check for `FailedScheduling` / image-pull events\n4. `k8s-aiops deployment restart <deploy> -n prod` → roll the deployment after fixing the cause\n5. **Failure branch**: if logs/events show an RBAC `403`, the kube context lacks the verb — run `kubectl auth can-i get pods -n prod` and switch to a context with adequate RBAC; the skill never retries a denied auth.\n\n### Triage an unhealthy namespace with RCA, then act on the worst finding\n\n1. `k8s-aiops diagnose pod-health -n prod` → worst-first findings; a `critical` `CrashLoopBackOff` on `prod/api` cites `restarts=9` and the exact `kubectl logs … --previous` action\n2. `k8s-aiops diagnose workload-readiness -n prod` → confirm the blast radius: e.g. `Deployment web ready 0/3` (`critical`, under-replicated)\n3. `k8s-aiops pod logs api-<hash> -n prod --tail 200 --previous`-equivalent via `k8s-aiops pod describe api-<hash> -n prod` → read the crash cause the RCA pointed you at\n4. `k8s-aiops deployment restart web -n prod` → roll the deployment once the root cause is fixed\n5. **Failure branch**: if `diagnose` returns an RBAC `403`, the kube context cannot list pods/deployments in that namespace — run `kubectl auth can-i list pods -n prod` and switch to a context with adequate RBAC; the RCA tools are read-only and never retry a denied auth.\n\n### Drain a node for maintenance, safely reversible\n\n1. `k8s-aiops node list` → identify the node and confirm it is `Ready`/schedulable\n2. `k8s-aiops node cordon <node> --dry-run` → preview, then `k8s-aiops node cordon <node>` (double confirm) — records an inverse `uncordon_node` undo descriptor\n3. After maintenance: `k8s-aiops node uncordon <node>` → re-enable scheduling\n4. **Failure branch**: if `doctor` shows the cluster unreachable, fix the kubeconfig context (`kubectl config get-contexts`) before retrying — cordon is never issued against an unauthenticated session.\n\n## Usage Mode\n\n| Scenario | Recommended | Why |\n|----------|:-----------:|-----|\n| Local/small models | **CLI** | fewer tokens than MCP |\n| Cloud models (Claude, GPT) | Either | MCP gives structured JSON I/O |\n| Automated pipelines | **MCP** | type-safe parameters, audited |\n\n> **MCP environment caveat**: MCP clients spawn the server with a CLEAN\n> environment — shell exports may not reach it. Set `K8S_AIOPS_HOME`,\n> `K8S_AUDIT_APPROVED_BY`, `K8S_AUDIT_RATIONALE` (and `KUBECONFIG` when the\n> kubeconfig is not at `~/.kube/config`) in the MCP server config's `env`\n> block, not just in your terminal.\n\n## MCP Tools (55 — 39 read, 16 write)\n\n| Category | Tools | R/W |\n|----------|-------|:---:|\n| Pods | `pod_list`, `pod_get`, `pod_logs`, `pod_describe` | Read |\n| | `delete_pod` | Write |\n| Deployments | `deployment_list`, `deployment_get` | Read |\n| | `scale_deployment`, `rollout_restart_deployment`, `delete_deployment` | Write |\n| Rollout | `rollout_status`, `rollout_history` | Read |\n| | `rollout_undo_deployment`, `rollout_pause`, `rollout_resume`, `set_deployment_image` | Write |\n| StatefulSets | `statefulset_list`, `statefulset_get` | Read |\n| | `scale_statefulset` | Write |\n| DaemonSets / ReplicaSets | `daemonset_list`, `daemonset_get`, `replicaset_list` | Read |\n| Jobs / CronJobs | `job_list`, `job_get`, `cronjob_list`, `cronjob_get` | Read |\n| | `delete_job` | Write |\n| Services / Ingress | `service_list`, `ingress_list`, `ingress_get`, `endpoints_list` | Read |\n| Config / Secrets | `configmap_list`, `configmap_get`, `secret_list` (names/keys only) | Read |\n| Storage | `pvc_list`, `pvc_get`, `pv_list`, `storageclass_list` | Read |\n| Nodes | `node_list`, `node_describe` | Read |\n| | `cordon_node`, `uncordon_node`, `drain_node` | Write |\n| Namespaces | `namespace_list` | Read |\n| | `create_namespace`, `delete_namespace` | Write |\n| Metrics (top) | `node_top`, `pod_top` | Read |\n| Cluster | `cluster_info`, `api_resources` | Read |\n| Events | `event_list` | Read |\n| Diagnostics / RCA | `pod_health_rca`, `workload_readiness_rca` | Read |\n| Undo | `undo_list` | Read |\n| | `undo_apply` | Write |\n\n**Security — secrets**: `secret_list` returns secret names, types, and key NAMES only. Secret VALUES are never read, returned, or logged, and there is deliberately no tool that returns secret values.\n\n**Dry-run previews**: every write tool takes `dry_run: bool = False`. A dry run returns a `{\"dryRun\": true, \"wouldX\": ...}` preview without touching the cluster, and no undo descriptor is recorded for a preview.\n\n**Harness features that light up**: write tools with a clean inverse pass an `undo=` lambda so the harness records an inverse descriptor (with `_undo_id`) to the undo store — `scale_deployment`/`scale_statefulset` record a scale-back to their returned `previous_replicas`, `set_deployment_image` records a restore to the captured `previous_image`, `cordon_node` ↔ `uncordon_node` and `rollout_pause` ↔ `rollout_resume` are mutual inverses, and `create_namespace` records a `delete_namespace`. `drain_node` records a partial `uncordon_node` inverse (cordon is reversible; evictions are not). `delete_*` and `rollout_undo_deployment` declare no undo. `risk_level=high`: `delete_deployment`, `delete_job`, `delete_namespace`, `drain_node`, `rollout_undo_deployment`. `undo_list` (read) lists recorded reversible writes whose undo tokens have not been applied yet, and `undo_apply` (write) executes a recorded inverse — itself governed, single-use, and supports `dry_run`. All 55 tools are audit-logged under `~/.k8s-aiops/` and pass through the budget/runaway guard, each recorded with a descriptive risk-tier label. `pod_top`/`node_top` return a clear \"metrics-server not installed\" message (not an error) when metrics-server is absent. Avoid tight poll loops (re-listing pods every second) — the runaway breaker backs this up.\n\n## CLI Quick Reference\n\n```bash\nk8s-aiops init                                            # interactive onboarding wizard\nk8s-aiops pod list [-n <ns>] [-t <target>]\nk8s-aiops pod get <name> [-n <ns>]\nk8s-aiops pod describe <name> [-n <ns>]                   # status, container states, events\nk8s-aiops pod logs <name> [-n <ns>] [--tail 200] [-c <container>]\nk8s-aiops pod delete <name> [-n <ns>] [--dry-run]        # double confirm\nk8s-aiops deployment list|get|scale|restart|delete ...    # scale/restart: single confirm + --dry-run; delete: double confirm\nk8s-aiops rollout status|history|pause|resume <name> [-n <ns>]\nk8s-aiops rollout set-image <name> <container> <image> [-n <ns>]\nk8s-aiops rollout undo <name> [--to-revision N] [--dry-run]   # double confirm\nk8s-aiops statefulset list|get|scale ...\nk8s-aiops daemonset list|get ...\nk8s-aiops job list|get|delete ...                         # delete: double confirm\nk8s-aiops cronjob list|get ...\nk8s-aiops service list [-n <ns>]\nk8s-aiops ingress list|get [-n <ns>]\nk8s-aiops configmap list|get [-n <ns>]\nk8s-aiops secret list [-n <ns>]                           # names/keys only — never values\nk8s-aiops storage pvc-list|pvc-get|pv-list|class-list\nk8s-aiops top pod|node                                    # requires metrics-server\nk8s-aiops node list|describe\nk8s-aiops node cordon|drain <name> [--dry-run]           # double confirm\nk8s-aiops node uncordon <name>\nk8s-aiops namespace list|create\nk8s-aiops namespace delete <name> [--dry-run]            # double confirm\nk8s-aiops cluster-info\nk8s-aiops api-resources\nk8s-aiops events [-n <ns>]\nk8s-aiops diagnose pod-health [-n <ns>] [-l <selector>]         # read-only RCA: crashloop/imagepull/OOM/unschedulable/restarts\nk8s-aiops diagnose workload-readiness [-n <ns>]                 # read-only RCA: ready<desired / stuck rollouts\nk8s-aiops doctor\nk8s-aiops mcp                                             # start MCP server (stdio)\n```\n\nSee `references/cli-reference.md` for the full command list.\n\n## Troubleshooting\n\n### \"Could not load kubeconfig … context not found\"\nThe named context does not exist in your kubeconfig. Run `kubectl config get-contexts` and set the target's `context:` to a listed name (or omit it to use current-context).\n\n### \"Authentication/authorization failed (401/403)\"\nThe kube context lacks the RBAC verb for the resource. Check with `kubectl auth can-i <verb> <resource> -n <ns>` and switch to a context/ServiceAccount with adequate roles. For EKS/GKE/AKS, confirm the exec-plugin (aws/gcloud/az CLI) is installed and logged in.\n\n### \"Resource not found (404)\"\nThe pod/deployment/node name or namespace is wrong, or the object was deleted. List the parent collection first (`pod list`, `deployment list`, `node list`) to get a current name. Remember most commands default to the `default` namespace unless `-n` is given.\n\n### \"Conflict (409)\"\nThe object changed concurrently (or already exists). Re-read it and retry the write.\n\n### Logs are empty or truncated\n`pod logs` returns the trailing `--tail` lines (default 100); raise `--tail`. For a multi-container pod, pass `-c <container>` or the API returns an error naming the available containers.\n\n## Audit & Safety\n\nAll operations are automatically audited via the bundled `@governed_tool` decorator (`k8s_aiops.governance`):\n- Every tool call logged to `~/.k8s-aiops/audit.db` (local SQLite audit DB; relocate with `K8S_AIOPS_HOME`)\n- Budget / runaway guard caps cumulative tool calls and wall-time, and trips on tight poll/retry loops — a safety backstop, not authorization\n- Undo store records inverse descriptors for reversible writes (scale → previous replicas; cordon ↔ uncordon)\n- Each write carries a descriptive risk-tier label into its audit row — a label, not a gate; `K8S_AUDIT_APPROVED_BY` / `K8S_AUDIT_RATIONALE` are optional annotations recorded when set, never required\n\n**Authorization is not this tool's job.** There is no read-only switch, policy file, or approval gate. Whether a write is permitted is the agent's judgement or the RBAC of the kubeconfig context you connect with — give it a read-only ServiceAccount and writes fail at the apiserver, the place that owns the permission.\n\nThe harness is bundled in the package — no external dependency, no manual setup. See `references/setup-guide.md` for security details.\n\nDriving these tools with a smaller / local model? See `references/agent-guardrails.md` — which guardrails the tool now enforces for you, plus a ready-to-paste system prompt.\n\n## Contributing & feature requests\n\nCoverage is intentionally focused. **Missing a device, action, or feature you need?** Open an issue or pull request at [github.com/AIops-tools/K8s-AIops](https://github.com/AIops-tools/K8s-AIops/issues) — feature requests, contributions, and comments are all welcome.\n\n## License\n\nMIT — [github.com/AIops-tools/K8s-AIops](https://github.com/AIops-tools/K8s-AIops)\n\nFile v0.13.2:_meta.json\n\n{\n  \"ownerId\": \"kn7b067awq2s97bn3d7p5qfhw5827pxc\",\n  \"slug\": \"k8s-aiops\",\n  \"version\": \"0.13.2\",\n  \"publishedAt\": 1789223200832\n}\n\nFile v0.13.2:references/agent-guardrails.md\n\n# Agent guardrails — running k8s-aiops with a smaller / local model\n\nIf you drive these tools with a local model (Llama, Qwen, Mistral … via Goose,\nOllama, LM Studio, or any OpenAI-compatible runtime), you will get noticeably\nbetter results with a short system prompt. This page gives you one, and — more\nimportantly — tells you which guardrails you **no longer need to write**, because\nthe tool now enforces them itself.\n\nThe distinction matters. A guardrail in a prompt is a request. A guardrail in the\nharness is a guarantee. Anything below that we could move into the harness, we did.\n\n## Authorization is not this tool's job — decide it where it belongs\n\nWhether a write should happen is your decision, or the account's. The tool does\nnot gate it — there is no read-only switch and no approval prompt to configure.\nThe two right places to control read vs write:\n\n- **The kubeconfig context you connect with.** Bind it to a ServiceAccount or\n  user whose RBAC grants only `get`/`list`/`watch`, and every write fails at the\n  apiserver — the only place the permission actually lives. No skill-side flag\n  can be argued around by a model, but a revoked RBAC verb cannot be. This is\n  strictly stronger than any in-process switch: it is enforced at the cluster.\n- **Your agent's system prompt.** If you want an observe-only session, tell the\n  model not to call the write tools (they are clearly tagged `[WRITE]`).\n\nWhat the tool *does* guarantee is that you can always see what happened:\n\n## What the tool enforces — do not waste prompt budget on these\n\n| You might be tempted to prompt | Why you don't need to |\n|---|---|\n| \"Log everything you do, over both MCP and the CLI\" | Every operation is audited to `~/.k8s-aiops/audit.db` regardless of what the model says it did — and the CLI writes the same row the MCP path does, so there is no unaudited entry point. Reversible writes also record an undo token capturing the *prior* state. |\n| \"Don't invent a value when a field is missing\" | A field the apiserver did not return comes back as `null`, never as `\"\"`. An unscheduled pod's `node` is `null`; a pod with no phase yet has `phase: null`; an object with no readable creation timestamp has `age: null`. Absent and empty are distinguishable in the payload. |\n| \"Tell me if the output was cut off\" | The limit-bearing reads return an envelope: `event_list` → `{\"events\": [...], \"returned\": N, \"limit\": L, \"truncated\": true/false}`, and `undo_list` → `{\"undos\": [...], \"returned\": N, \"limit\": L, \"truncated\": ...}`. Truncation is **measured** (one extra row is fetched), not guessed from `len(rows) == limit`. |\n| \"Preserve the ordering / tell me what's most urgent\" | `pod_health_rca` and `workload_readiness_rca` findings carry an explicit 1-based `rank`, worst-first, and each finding's `detail` cites the measured signal (the waiting reason, the restart count, the ready/desired ratio). Priority is in the payload, not implied by list position. |\n| \"Confirm before anything destructive\" | The destructive CLI commands (deployment/job/namespace delete, node cordon/drain, rollout undo) are `--dry-run`-able and require double confirmation. |\n| \"Don't get stuck retrying\" | The runaway guard trips a circuit breaker if the same call is hammered in a tight loop — a stuck agent is stopped rather than left to burn calls and time. It is a safety backstop, not an authorization gate. |\n\n## What still needs a prompt\n\nThese are model-behaviour problems the harness cannot fix from the outside.\nCopy this into your agent's system prompt:\n\n```text\nYou operate a Kubernetes cluster through the k8s-aiops MCP tools.\n\nTOOL USE\n- Before answering any question about the current cluster, you MUST call a\n  tool. Never answer from memory or assumption.\n- Actually invoke the tool. Do not describe the call you would make, and do not\n  emit an example JSON response in place of calling it.\n- If a tool call fails, report the real error verbatim. Never fill the gap with\n  a plausible-sounding answer.\n\nREADING RESULTS\n- Read the whole result before concluding. If a result contains a \"truncated\"\n  field that is true, say so and re-run with a higher limit instead of treating\n  the partial result as complete.\n- A null field means the apiserver did not return that value. Report it as \"not\n  available\" — never infer it. In particular, a null \"node\" means the pod is\n  not scheduled; it does not mean the node is unknown or missing.\n- Report values exactly as returned. Do not normalise, translate, or prettify\n  phases, conditions, container-state reasons, or resource names.\n- When an RCA result has findings, work in \"rank\" order and cite the measured\n  number in each finding's \"detail\".\n\nSCOPE AND IDENTIFIERS\n- Always state the namespace you are talking about. A bare pod or deployment\n  name is ambiguous — the same name exists in many namespaces.\n- Omitting the namespace means ALL namespaces, not the default one. Never\n  silently widen a namespaced question into a cluster-wide answer.\n- Do not confuse a namespace with a context (a cluster), a pod name with its\n  deployment name, or a container name with the pod that contains it. A pod\n  name generated by a ReplicaSet is not a stable identifier.\n- Separate observation from interpretation. State what the tools returned, then\n  any interpretation, clearly marked as such.\n- Do not assert a capacity, performance, or availability problem unless a tool\n  result supports it. Do not add generic advice that does not follow from the\n  tool output.\n```\n\n## Recommended setup for a local model\n\nStart with a kubeconfig context that *cannot* write — a ServiceAccount whose\nRBAC grants only `get`/`list`/`watch` — verify, and widen its permission only\nwhen you trust the setup. RBAC is enforced at the apiserver, so a write fails at\nthe cluster no matter what the model attempts:\n\n```bash\n# Point KUBECONFIG at a read-only ServiceAccount context, then:\nk8s-aiops doctor\n```\n\nOptionally annotate the audit trail with who is operating and why — recorded on\nevery row, never required:\n\n```bash\nexport K8S_AUDIT_APPROVED_BY=\"your.name@example.com\"\nexport K8S_AUDIT_RATIONALE=\"incident INC-1234 — restart the stuck payments rollout\"\n```\n\n## Kubernetes-specific notes\n\n- **Namespace confusion is the single most common failure mode.** Every\n  namespaced tool takes an optional `namespace`; omitting it lists across **all**\n  namespaces. A smaller model reliably drops the namespace from a follow-up call\n  and then reports a cluster-wide count as if it were the namespace's. Pin the\n  namespace in your prompt (or in the target's `namespace:` in\n  `~/.k8s-aiops/config.yaml`) and ask the model to restate it in every answer.\n- **Context confusion is the expensive one.** `target` selects a *cluster*, not a\n  namespace. If you have prod and staging contexts in the same kubeconfig, give\n  the agent a config with only the cluster it should touch — the model has no way\n  to notice it is on the wrong one, because both look plausible.\n- **Authentication is delegated to your kubeconfig.** Unlike the other tools in\n  this line, k8s-aiops has **no credential store**: there is nothing to encrypt\n  and no master password. (`k8s-aiops secret ...` lists Kubernetes *Secret\n  resources* by name — it is not a credential manager.) The agent inherits the RBAC of the\n  kubeconfig user. That makes RBAC your strongest guardrail — a read-only\n  ServiceAccount kubeconfig enforces read-only at the *cluster*, the only place\n  the permission truly lives.\n- **`drain_node` is the most dangerous tool here.** It evicts pods cluster-wide\n  in effect, and its blast radius is not visible in its arguments. Keep it out of\n  the RBAC role you hand the agent unless you specifically intend node drains.\n- **Pod names are not stable.** A ReplicaSet-generated pod name changes on every\n  rollout, so an id the model cached earlier in the conversation may already be\n  gone. Prefer `deployment_get` / `rollout_status` over re-using a pod name.\n- **Prefer the RCA tools over multi-step chains.** `pod_health_rca` and\n  `workload_readiness_rca` do the list-then-correlate work inside one call, so a\n  smaller model does not have to chain reads and keep names/namespaces straight.\n- **Secrets are names-only by design.** `secret_list` returns key *names*, never\n  values — there is no tool that can exfiltrate a secret payload.\n\n## If your model still struggles\n\nSome behaviours are model-capacity limits rather than prompt problems:\n\n- **Multi-tool workflows time out or drift.** Prefer the `*_rca` tools — they do\n  the multi-step correlation inside one call.\n- **The model ignores later tool results in a long context.** Ask narrower\n  questions, always scope to a namespace, and use `limit` deliberately rather\n  than pulling whole-cluster inventories (`pod_list` with no namespace on a busy\n  cluster will bury everything else in the context).\n- **The model describes calls instead of making them.** This is usually a\n  runtime/tool-calling-format mismatch, not a prompt problem — check that your\n  client advertises the tools in the format your model was trained on.\n\nFeedback on running this with a specific local model is genuinely useful —\nopen an issue at\n[github.com/AIops-tools/K8s-AIops](https://github.com/AIops-tools/K8s-AIops/issues)\nwith the model, runtime, and what went wrong.\n\nFile v0.13.2:references/capabilities.md\n\n# k8s-aiops Capabilities\n\n55 MCP tools (39 read / 16 write). Every tool is wrapped with `@governed_tool`\n(audit + policy + budget + risk-tier; undo where a clean inverse exists). Returns\nare high-signal summaries — `_get` / `_describe` tools add detail for a single object.\n\n## Read tools\n\n| Tool | Returns | Risk |\n|------|---------|:----:|\n| `pod_list` | name, namespace, phase, ready, restarts, node, age | low |\n| `pod_get` | + host_ip, pod_ip, containers | low |\n| `pod_describe` | status, conditions, container states + restart counts, recent events | low |\n| `pod_logs` | trailing log lines (default 100) | low |\n| `deployment_list` / `deployment_get` | replicas summary / + strategy, images | low |\n| `rollout_status` | desired/updated/available/unavailable + paused | low |\n| `rollout_history` | revisions (from replicasets) with images | low |\n| `statefulset_list` / `statefulset_get` | desired/ready/current / + service, images | low |\n| `daemonset_list` / `daemonset_get` | desired/ready/available / + images | low |\n| `replicaset_list` | name, namespace, desired/ready, age | low |\n| `job_list` / `job_get` | completions, succeeded/failed/active | low |\n| `cronjob_list` / `cronjob_get` | schedule, suspend, active, last schedule | low |\n| `service_list` | name, namespace, type, cluster IP, ports | low |\n| `ingress_list` / `ingress_get` | class, hosts / + path→backend rules | low |\n| `endpoints_list` | ready addresses + ports | low |\n| `configmap_list` / `configmap_get` | key count / keys + values | low |\n| `secret_list` | names, types, key NAMES only (values redacted) | low |\n| `pvc_list` / `pvc_get` | status, capacity, class / + access modes | low |\n| `pv_list` | capacity, status, claim, class | low |\n| `storageclass_list` | provisioner, reclaim policy, default | low |\n| `node_list` / `node_describe` | status, roles / capacity, allocatable, conditions, taints | low |\n| `namespace_list` | name, phase, age | low |\n| `pod_top` / `node_top` | CPU/mem via metrics-server (graceful if absent) | low |\n| `cluster_info` | server version, node/ready/namespace counts | low |\n| `api_resources` | available API groups + versions | low |\n| `event_list` | type, reason, object, namespace, message, age | low |\n| `pod_health_rca` | worst-first findings: CrashLoopBackOff, image-pull, OOMKilled, unschedulable, high restarts (each cites the reason/count) | low |\n| `workload_readiness_rca` | worst-first findings: ready<desired, zero-ready outages, stuck rollouts (Deployment/StatefulSet/DaemonSet) | low |\n| `undo_list` | recorded reversible writes / not-yet-applied undo tokens | low |\n\n## Write tools\n\n| Tool | Effect | Risk | Undo |\n|------|--------|:----:|------|\n| `scale_deployment` | set replica count | medium | scale back to `previous_replicas` |\n| `scale_statefulset` | set replica count | medium | scale back to `previous_replicas` |\n| `rollout_restart_deployment` | patch `restartedAt` annotation | medium | none (pods already rolling) |\n| `rollout_pause` / `rollout_resume` | toggle `spec.paused` | medium | each other |\n| `rollout_undo_deployment` | roll back to a prior revision | **high** | none |\n| `set_deployment_image` | update a container image | medium | restore `previous_image` |\n| `delete_pod` | delete a pod | medium | none (controller recreates) |\n| `delete_deployment` | delete deployment + pods | **high** | none |\n| `delete_job` | delete a job + pods | **high** | none |\n| `create_namespace` | create a namespace | medium | `delete_namespace` |\n| `delete_namespace` | delete namespace + everything in it | **high** | none |\n| `cordon_node` / `uncordon_node` | toggle schedulability | medium | each other |\n| `drain_node` | cordon + evict pods (skips DaemonSet/mirror) | **high** | partial: `uncordon_node` |\n| `undo_apply` | execute a recorded inverse (itself governed, single-use, supports `dry_run`) | medium | n/a (is the undo) |\n\n### `delete_namespace` refusals\n\nTwo targets are refused before anything is deleted, on the real call and on the\n`dry_run` preview alike (a preview never green-lights a delete that would be\nrejected):\n\n- **Control-plane namespaces** — `kube-system`, `kube-public`, `kube-node-lease`.\n  Deleting one takes CoreDNS, kube-proxy, the CNI and node heartbeats with it.\n  Pass `confirm=true` (CLI `--confirm`) to proceed deliberately, e.g. when\n  tearing a cluster down or clearing a namespace stuck `Terminating`.\n- **The namespace holding this target's own ServiceAccount credential** — that\n  delete revokes the credential it is running on, and `delete_namespace` has no\n  undo. **`confirm` does not override this one.** Re-run from a context whose\n  credential lives elsewhere. Only applies when the kubeconfig authenticates\n  with a ServiceAccount token; certificate and `exec` (EKS/GKE/AKS) credentials\n  are not namespace-bound, so the check stands down rather than guessing.\n\n## Token-budget notes\n\n- List tools accept a `namespace` filter to keep responses small; events and pod\n  listings also accept `limit` / `label_selector` where applicable.\n- Prefer `pod_get` / `deployment_get` over re-listing when you already have a name.\n- The runaway guard trips on tight poll loops — wait between repeated list calls.\n\n## Design notes / Kubernetes-client assumptions\n\n- Authentication is delegated to the kubeconfig; the skill never touches raw\n  credentials (works with client certs, tokens, and EKS/GKE/AKS exec plugins).\n- Typed Api clients (`CoreV1Api`, `AppsV1Api`, `BatchV1Api`, `NetworkingV1Api`,\n  `StorageV1Api`, `CustomObjectsApi`, `ApisApi`, `VersionApi`) are cached per kube\n  context in a module dict — third-party client objects are never monkey-patched.\n- `secret_list` reads only key NAMES from `secret.data` — secret values are never\n  read, returned, or logged, and no tool exposes them.\n- `pod_top` / `node_top` use the `metrics.k8s.io/v1beta1` API; when metrics-server\n  is absent the 404/503 is caught and returned as `{available: false, message}`.\n- `ApiException` is translated centrally at the connection layer into a teaching\n  `K8sApiError` (404/403/409/5xx), so agents see actionable messages, not tracebacks.\n\nFile v0.13.2:references/cli-reference.md\n\n# k8s-aiops CLI Reference\n\nAll commands accept `-t/--target <name>` to select a configured target (a kube\ncontext). Namespaced commands accept `-n/--namespace <ns>`; omit it to use the\ntarget's default namespace (read lists fall back to all-namespaces).\n\n## Onboarding\n\n```bash\nk8s-aiops init                    # interactive wizard: register kube contexts as targets\n```\n\n## Pods\n\n```bash\nk8s-aiops pod list [-n <ns>] [-t <target>]\nk8s-aiops pod get <name> [-n <ns>]\nk8s-aiops pod describe <name> [-n <ns>]              # status, container states, events\nk8s-aiops pod logs <name> [-n <ns>] [--tail N] [-c <container>]\nk8s-aiops pod delete <name> [-n <ns>] [--dry-run]    # destructive: double confirm\n```\n\n## Deployments & Rollouts\n\n```bash\nk8s-aiops deployment list [-n <ns>]\nk8s-aiops deployment get <name> [-n <ns>]\nk8s-aiops deployment scale <name> <replicas> [-n <ns>]\nk8s-aiops deployment restart <name> [-n <ns>]        # rolling restart\nk8s-aiops deployment delete <name> [-n <ns>] [--dry-run]   # HIGH RISK: double confirm\nk8s-aiops rollout status <name> [-n <ns>]\nk8s-aiops rollout history <name> [-n <ns>]\nk8s-aiops rollout pause|resume <name> [-n <ns>]\nk8s-aiops rollout set-image <name> <container> <image> [-n <ns>]\nk8s-aiops rollout undo <name> [--to-revision N] [--dry-run]   # HIGH RISK: double confirm\n```\n\n## StatefulSets / DaemonSets / Jobs / CronJobs\n\n```bash\nk8s-aiops statefulset list|get [-n <ns>]\nk8s-aiops statefulset scale <name> <replicas> [-n <ns>]\nk8s-aiops daemonset list|get [-n <ns>]\nk8s-aiops job list|get [-n <ns>]\nk8s-aiops job delete <name> [-n <ns>] [--dry-run]    # destructive: double confirm\nk8s-aiops cronjob list|get [-n <ns>]\n```\n\n## Services, Ingress, Config, Storage\n\n```bash\nk8s-aiops service list [-n <ns>]\nk8s-aiops ingress list|get [-n <ns>]\nk8s-aiops configmap list|get [-n <ns>]\nk8s-aiops secret list [-n <ns>]                      # names + key NAMES only, never values\nk8s-aiops storage pvc-list|pvc-get [-n <ns>]\nk8s-aiops storage pv-list|class-list\n```\n\n## Nodes & Metrics\n\n```bash\nk8s-aiops node list\nk8s-aiops node describe <name>\nk8s-aiops node cordon <name> [--dry-run]             # destructive: double confirm\nk8s-aiops node uncordon <name>\nk8s-aiops node drain <name> [--dry-run]              # HIGH RISK: double confirm\nk8s-aiops top pod|node                               # requires metrics-server\n```\n\n## Namespaces, Cluster & Events\n\n```bash\nk8s-aiops namespace list\nk8s-aiops namespace create <name>\nk8s-aiops namespace delete <name> [--dry-run]        # HIGH RISK: double confirm\nk8s-aiops cluster-info\nk8s-aiops api-resources\nk8s-aiops events [-n <ns>]\n```\n\n## Diagnostics & MCP\n\n```bash\nk8s-aiops diagnose pod-health [-n <ns>] [-l <selector>]  # RCA: crashloop/imagepull/OOM/unschedulable/restarts (read-only)\nk8s-aiops diagnose workload-readiness [-n <ns>]          # RCA: ready<desired / stuck rollouts (read-only)\nk8s-aiops doctor [--skip-auth]    # check config + cluster reachability\nk8s-aiops mcp                     # start the MCP server over stdio\n```\n\n## Flags summary\n\n| Flag | Meaning |\n|------|---------|\n| `-t, --target` | Target name from `~/.k8s-aiops/config.yaml` |\n| `-n, --namespace` | Namespace scope |\n| `--tail` | Trailing log lines (pod logs, default 100) |\n| `-c, --container` | Container name (pod logs) |\n| `--dry-run` | Preview a destructive op without executing |\n| `--to-revision` | Rollout revision (`rollout undo`, 0 = previous) |\n| `--skip-auth` | Skip the connectivity check in `doctor` |\n\nFile v0.13.2:references/setup-guide.md\n\n# k8s-aiops Setup Guide\n\n## Install\n\n```bash\nuv tool install k8s-aiops\nk8s-aiops doctor\n```\n\n`k8s-aiops` requires Python ≥ 3.11. If `uv` picked an older interpreter:\n\n```bash\nuv python install 3.12\nuv tool install --python 3.12 --force k8s-aiops\n```\n\n## Connecting to a cluster\n\nk8s-aiops reads your **kubeconfig** — the same one `kubectl` uses\n(`KUBECONFIG` env var, else `~/.kube/config`). Out of the box it talks to your\ncurrent kube-context, so a fresh install needs no config file:\n\n```bash\nk8s-aiops pod list\n```\n\n### Named targets (multiple clusters)\n\nThe fastest way is the interactive wizard, which discovers the contexts in your\nkubeconfig and registers the ones you pick (writing `config.yaml`, dir chmod 700):\n\n```bash\nk8s-aiops init\n```\n\nOr create `~/.k8s-aiops/config.yaml` by hand to give contexts friendly names:\n\n```yaml\ntargets:\n  - name: prod\n    context: prod-eks            # a context from `kubectl config get-contexts`\n    namespace: default           # optional default namespace\n    # kubeconfig: /path/to/kubeconfig   # optional, overrides KUBECONFIG/~/.kube/config\n  - name: lab\n    context: k3s-lab\n```\n\nThen select with `-t`:\n\n```bash\nk8s-aiops -t prod pod list -n payments\n```\n\nNo secrets are stored here — authentication lives entirely in the kubeconfig.\n\n### Works with\n\nStandard Kubernetes, k3s, EKS, GKE, AKS, kind, minikube. For managed clusters\n(EKS/GKE/AKS) the kubeconfig uses an exec plugin (`aws`/`gcloud`/`az`); make sure\nthat CLI is installed and logged in.\n\n## Security\n\n> **Disclaimer**: This is a community-maintained open-source project and is **not affiliated with, endorsed by, or sponsored by the Cloud Native Computing Foundation, the Kubernetes project, or k3s/Rancher.** \"Kubernetes\" and \"k3s\" are trademarks of their respective owners. Source is auditable at [github.com/AIops-tools/K8s-AIops](https://github.com/AIops-tools/K8s-AIops) under the MIT license.\n\n1. **Source code** — [github.com/AIops-tools/K8s-AIops](https://github.com/AIops-tools/K8s-AIops), MIT.\n2. **Config file contents** — `config.yaml` holds only target names, kube context\n   names, and optional namespaces/paths. No credentials.\n3. **Credentials** — delegated to the kubeconfig; never read, logged, or echoed by\n   k8s-aiops. Keep `~/.k8s-aiops` owner-only (`chmod 700`).\n4. **TLS verification** — follows the kubeconfig (`certificate-authority` /\n   `insecure-skip-tls-verify`); the skill does not weaken it.\n5. **Prompt-injection protection** — all API-returned text (names, log lines, event\n   messages) is run through `sanitize()` (truncation + control-character stripping).\n6. **Least privilege** — bind the kube context to a ServiceAccount/user with only the\n   RBAC verbs you need (read-only needs `get`/`list`/`watch`; writes add\n   `patch`/`delete`).\n\n## Governance harness\n\nBundled under `k8s_aiops.governance` — no external dependency. State lives under\n`~/.k8s-aiops/` (override with `K8S_AIOPS_HOME`):\n\n- `audit.db` — every tool call (skill, tool, params, status, duration, agent),\n  each carrying a descriptive risk-tier label derived from the tool's\n  `risk_level`. The tier is a label, not a gate.\n- Token/runaway budget guard (`K8S_MAX_TOOL_CALLS`, `K8S_MAX_TOOL_SECONDS`,\n  `K8S_RUNAWAY_MAX`, `K8S_RUNAWAY_WINDOW_SEC`) — a safety backstop, not\n  authorization.\n- Undo store — inverse descriptors for reversible writes.\n- Accountability: `K8S_AUDIT_APPROVED_BY` / `K8S_AUDIT_RATIONALE` are optional\n  annotations recorded on the audit row when set — never required, never\n  blocking. Authorization is the RBAC of the kubeconfig context, not this tool.\n\n## MCP client config\n\n```jsonc\n{\n  \"command\": \"k8s-aiops\",\n  \"args\": [\"mcp\"],\n  \"env\": { \"K8S_AIOPS_CONFIG\": \"~/.k8s-aiops/config.yaml\" }\n}\n```\n\nFallback (no `uv tool install`): `uvx --from k8s-aiops k8s-aiops-mcp`. Prefer the\ninstalled entry point — it does not re-resolve PyPI at launch.\n\n## Static analysis\n\n```bash\nuvx bandit -r k8s_aiops/ mcp_server/\n```\n\nFile v0.13.2:skill-card.md\n\n## Description:\n\nk8s-aiops helps agents inspect, diagnose, and operate kubeconfig-reachable Kubernetes clusters through CLI and MCP tools for resources, logs, metrics, rollouts, scaling, deletion, namespace management, and node maintenance.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[zw008](https://clawhub.ai/user/zw008)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers, platform engineers, and SREs use this skill to inspect Kubernetes resources, diagnose pod and workload health, read logs and events, and perform governed operational changes such as scaling, rollouts, deletion, namespace operations, and node maintenance.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: The skill can perform high-impact Kubernetes changes, including workload deletion, namespace deletion, node drain, rollout undo, and other write operations.\n\nMitigation: Install and run it with a dedicated least-privilege kubeconfig, preferably read-only by default; only grant write RBAC for clusters and namespaces where the operator intends those changes.\n\nRisk: The skill relies on kubeconfig RBAC and agent judgment rather than a built-in read-only mode or approval gate.\n\nMitigation: Use Kubernetes RBAC as the primary control boundary, review write requests before execution, and avoid admin or production contexts unless the operational intent is explicit.\n\nRisk: Using an unpinned or unverified package could execute a release different from the reviewed artifact.\n\nMitigation: Prefer a pinned and verified package release before using the skill against real clusters.\n\n## Reference(s):\n\n- [k8s-aiops project homepage](https://github.com/AIops-tools/K8s-AIops)\n- [k8s-aiops Capabilities](references/capabilities.md)\n- [k8s-aiops CLI Reference](references/cli-reference.md)\n- [k8s-aiops Setup Guide](references/setup-guide.md)\n- [Agent guardrails](references/agent-guardrails.md)\n\n## Skill Output:\n\n**Output Type(s):** [text, markdown, shell commands, configuration, guidance]\n\n**Output Format:** [Markdown with inline shell commands and structured CLI or MCP tool guidance]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [May include Kubernetes resource summaries, diagnostics, log excerpts, dry-run previews, and operational recommendations scoped to the selected kubeconfig context.]\n\n## Skill Version(s):\n\n0.13.2 (source: server release metadata)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nArchive v0.13.1: 7 files, 19579 bytes\n\nFiles: references/agent-guardrails.md (9366b), references/capabilities.md (6127b), references/cli-reference.md (3495b), references/setup-guide.md (3987b), skill-card.md (2385b), SKILL.md (18492b), _meta.json (129b)\n\nFile v0.13.1:SKILL.md\n\n---\nname: k8s-aiops\nslug: k8s-aiops\ndisplayName: \"k8s AIops\"\nsummary: \"Governed Kubernetes ops — 55 MCP tools with audit, budget, undo, risk-tier audit labels.\"\nlicense: MIT\nhomepage: https://github.com/AIops-tools/K8s-AIops\ntags: [aiops, mcp, governance, k8s]\ndescription: >\n  Use this skill whenever the user needs to operate a Kubernetes cluster — list/inspect pods, deployments, statefulsets, daemonsets, replicasets, jobs, cronjobs, services, ingresses, endpoints, configmaps, secrets (names/keys only), PVCs/PVs/storageclasses, nodes, namespaces, and events; read pod logs; describe pods/nodes; pod/node top (metrics); read-only diagnostics / RCA (pod-health, workload-readiness); scale deployments/statefulsets; rollout status/history/undo/pause/resume and set image; delete pods/deployments/jobs; create/delete namespaces; and cordon/uncordon/drain nodes. Works with any kubeconfig-reachable cluster (standard Kubernetes, k3s, EKS, GKE, AKS).\n  Always use this skill for \"list k8s pods\", \"scale deployment\", \"kubernetes pod logs\", \"describe pod\", \"why is my pod crashing\", \"diagnose pods\", \"which deployments are unhealthy\", \"rollout undo\", \"set image\", \"top pods\", \"drain node\", \"cordon node\", \"restart deployment\", \"k3s\", or \"kubectl\"-style tasks when the context is explicitly Kubernetes / a cluster.\n  Do NOT use when the target is not a Kubernetes cluster (hypervisor VM lifecycle, backup products, or cloud-provider consoles are out of scope).\n  Common Kubernetes operations with a built-in governance harness (audit, token budget, undo, risk-tier labels).\ninstaller:\n  kind: uv\n  package: k8s-aiops\nargument-hint: \"[resource name or describe your Kubernetes task]\"\nallowed-tools:\n  - Bash\nmetadata: {\"openclaw\":{\"requires\":{\"anyBins\":[\"k8s-aiops\",\"uvx\"]},\"optional\":{\"env\":[\"K8S_AIOPS_CONFIG\",\"KUBECONFIG\",\"K8S_AIOPS_HOME\"]},\"homepage\":\"https://github.com/AIops-tools/K8s-AIops\",\"emoji\":\"☸️\",\"os\":[\"macos\",\"linux\"]}}\ncompatibility: >\n  Standalone, self-governed Kubernetes operations. The governance harness (audit, token/runaway budget, undo, risk-tier labels) is bundled in the package — no external skill-family dependency.\n  All write operations are audited to a local SQLite DB under ~/.k8s-aiops/ (relocatable via K8S_AIOPS_HOME).\n  Credentials: k8s-aiops handles NO credentials directly — authentication is delegated to the kubeconfig (KUBECONFIG env or ~/.kube/config), which may hold client certs, bearer tokens, or exec plugins (EKS/GKE/AKS). Underlying credentials are never read, logged, or echoed. The state dir ~/.k8s-aiops should be chmod 700.\n  Destructive operations (deployment/job/namespace delete, node cordon/drain, rollout undo) require double confirmation at the CLI layer and support --dry-run. All write tools pass through the @governed_tool decorator (budget/runaway guard + audit + a descriptive risk-tier label). Reversible writes record an inverse undo descriptor (scale_deployment/scale_statefulset restore the previous replica count; set_deployment_image restores the previous image; cordon_node ↔ uncordon_node and rollout_pause ↔ rollout_resume; create_namespace ↔ delete_namespace); delete_* and rollout_undo record none. risk_level=high: delete_deployment, delete_job, delete_namespace, drain_node, rollout_undo_deployment. Secret VALUES are never read or returned by any tool.\n  Webhooks: none — no outbound network calls beyond the configured Kubernetes API server.\n  TLS: follows the kubeconfig (certificate-authority / insecure-skip-tls-verify); the skill does not weaken it.\n  Transitive dependencies: the official kubernetes Python client and the MCP SDK. No post-install scripts or background services.\n---\n\n# k8s AIops\n\n> **Disclaimer**: This is a community-maintained open-source project and is **not affiliated with, endorsed by, or sponsored by the Cloud Native Computing Foundation, the Kubernetes project, or k3s/Rancher.** \"Kubernetes\" and \"k3s\" are trademarks of their respective owners. Source code is publicly auditable at [github.com/AIops-tools/K8s-AIops](https://github.com/AIops-tools/K8s-AIops) under the MIT license.\n\nGoverned Kubernetes operations — **55 MCP tools**, every one wrapped with the bundled `@governed_tool` harness: a local unified audit log under `~/.k8s-aiops/`, a token/runaway budget guard, undo-token recording, and a descriptive risk-tier label on every audit row. Works with any kubeconfig-reachable cluster (standard Kubernetes, k3s, EKS, GKE, AKS). Run `k8s-aiops init` for a friendly onboarding wizard that registers your kube contexts as named targets.\n\n> **Standalone**: the governance harness is bundled in the package (`k8s_aiops.governance`) — k8s-aiops has no external skill-family dependency. Coverage focuses on common operations and is not yet exhaustive.\n\n## What This Skill Does\n\n| Category | Tools | Count | Read or Write |\n|----------|-------|:-----:|:-------------:|\n| **Pods** | list, get, logs, describe, delete | 5 | 4 read / 1 write |\n| **Deployments** | list, get, scale, rollout restart, delete | 5 | 2 read / 3 write |\n| **Rollout** | status, history, undo, pause, resume, set-image | 6 | 2 read / 4 write |\n| **StatefulSets** | list, get, scale | 3 | 2 read / 1 write |\n| **DaemonSets** | list, get | 2 | 2 read |\n| **ReplicaSets** | list | 1 | 1 read |\n| **Jobs / CronJobs** | job list/get/delete, cronjob list/get | 5 | 4 read / 1 write |\n| **Services / Ingress / Endpoints** | service list, ingress list/get, endpoints list | 4 | 4 read |\n| **Config / Secrets** | configmap list/get, secret list (names/keys only) | 3 | 3 read |\n| **Storage** | pvc list/get, pv list, storageclass list | 4 | 4 read |\n| **Nodes** | list, describe, cordon, uncordon, drain | 5 | 2 read / 3 write |\n| **Namespaces** | list, create, delete | 3 | 1 read / 2 write |\n| **Metrics (top)** | pod, node | 2 | 2 read |\n| **Cluster** | cluster_info, api_resources | 2 | 2 read |\n| **Events** | list | 1 | 1 read |\n| **Diagnostics / RCA** | pod-health, workload-readiness | 2 | 2 read |\n\n## Quick Install\n\n```bash\nuv tool install k8s-aiops\nk8s-aiops init            # friendly wizard: register your kube contexts as targets\nk8s-aiops doctor          # or skip init — works with your current kube-context too\n```\n\nOr as an OpenClaw plugin, which installs this skill and its MCP server together:\n\n```bash\nopenclaw plugins install clawhub:@aiops-tools/k8s-aiops\nopenclaw skills info k8s-aiops          # expect: Visible to model: yes\n```\n\nNeeds `uvx` on `PATH`: the MCP server is fetched with uv, pinned to this release.\n\n## When to Use This Skill\n\n- List/inspect pods, deployments, services, nodes, namespaces and recent events\n- Read a pod's recent log lines to diagnose a crash loop\n- Run a read-only RCA sweep (`diagnose pod-health` / `diagnose workload-readiness`) to find the root cause worst-first\n- Scale a deployment up/down, or trigger a rolling restart\n- Delete a stuck pod (a controller recreates it) or a deployment\n- Cordon a node before maintenance, then uncordon it after\n\n**Do NOT use when** the target is not a Kubernetes cluster (hypervisor VM lifecycle, backup products, or cloud-provider consoles are out of scope for this skill).\n\n## Related Skills — Skill Routing\n\n| If the user wants… | Use |\n|--------------------|-----|\n| Kubernetes pods / deployments / nodes | **k8s-aiops** (this skill) |\n| Hypervisor VM lifecycle (power, snapshot, migrate) | a hypervisor ops skill |\n| Backup & restore | a backup ops skill |\n\n## Common Workflows\n\n### Diagnose a crash-looping pod and restart its deployment\n\n1. `k8s-aiops pod list -n prod` → find the pod with high `restarts` / non-Running `phase`\n2. `k8s-aiops pod logs <pod> -n prod --tail 200` → read the recent logs for the crash cause\n3. `k8s-aiops events -n prod` → check for `FailedScheduling` / image-pull events\n4. `k8s-aiops deployment restart <deploy> -n prod` → roll the deployment after fixing the cause\n5. **Failure branch**: if logs/events show an RBAC `403`, the kube context lacks the verb — run `kubectl auth can-i get pods -n prod` and switch to a context with adequate RBAC; the skill never retries a denied auth.\n\n### Triage an unhealthy namespace with RCA, then act on the worst finding\n\n1. `k8s-aiops diagnose pod-health -n prod` → worst-first findings; a `critical` `CrashLoopBackOff` on `prod/api` cites `restarts=9` and the exact `kubectl logs … --previous` action\n2. `k8s-aiops diagnose workload-readiness -n prod` → confirm the blast radius: e.g. `Deployment web ready 0/3` (`critical`, under-replicated)\n3. `k8s-aiops pod logs api-<hash> -n prod --tail 200 --previous`-equivalent via `k8s-aiops pod describe api-<hash> -n prod` → read the crash cause the RCA pointed you at\n4. `k8s-aiops deployment restart web -n prod` → roll the deployment once the root cause is fixed\n5. **Failure branch**: if `diagnose` returns an RBAC `403`, the kube context cannot list pods/deployments in that namespace — run `kubectl auth can-i list pods -n prod` and switch to a context with adequate RBAC; the RCA tools are read-only and never retry a denied auth.\n\n### Drain a node for maintenance, safely reversible\n\n1. `k8s-aiops node list` → identify the node and confirm it is `Ready`/schedulable\n2. `k8s-aiops node cordon <node> --dry-run` → preview, then `k8s-aiops node cordon <node>` (double confirm) — records an inverse `uncordon_node` undo descriptor\n3. After maintenance: `k8s-aiops node uncordon <node>` → re-enable scheduling\n4. **Failure branch**: if `doctor` shows the cluster unreachable, fix the kubeconfig context (`kubectl config get-contexts`) before retrying — cordon is never issued against an unauthenticated session.\n\n## Usage Mode\n\n| Scenario | Recommended | Why |\n|----------|:-----------:|-----|\n| Local/small models | **CLI** | fewer tokens than MCP |\n| Cloud models (Claude, GPT) | Either | MCP gives structured JSON I/O |\n| Automated pipelines | **MCP** | type-safe parameters, audited |\n\n> **MCP environment caveat**: MCP clients spawn the server with a CLEAN\n> environment — shell exports may not reach it. Set `K8S_AIOPS_HOME`,\n> `K8S_AUDIT_APPROVED_BY`, `K8S_AUDIT_RATIONALE` (and `KUBECONFIG` when the\n> kubeconfig is not at `~/.kube/config`) in the MCP server config's `env`\n> block, not just in your terminal.\n\n## MCP Tools (55 — 39 read, 16 write)\n\n| Category | Tools | R/W |\n|----------|-------|:---:|\n| Pods | `pod_list`, `pod_get`, `pod_logs`, `pod_describe` | Read |\n| | `delete_pod` | Write |\n| Deployments | `deployment_list`, `deployment_get` | Read |\n| | `scale_deployment`, `rollout_restart_deployment`, `delete_deployment` | Write |\n| Rollout | `rollout_status`, `rollout_history` | Read |\n| | `rollout_undo_deployment`, `rollout_pause`, `rollout_resume`, `set_deployment_image` | Write |\n| StatefulSets | `statefulset_list`, `statefulset_get` | Read |\n| | `scale_statefulset` | Write |\n| DaemonSets / ReplicaSets | `daemonset_list`, `daemonset_get`, `replicaset_list` | Read |\n| Jobs / CronJobs | `job_list`, `job_get`, `cronjob_list`, `cronjob_get` | Read |\n| | `delete_job` | Write |\n| Services / Ingress | `service_list`, `ingress_list`, `ingress_get`, `endpoints_list` | Read |\n| Config / Secrets | `configmap_list`, `configmap_get`, `secret_list` (names/keys only) | Read |\n| Storage | `pvc_list`, `pvc_get`, `pv_list`, `storageclass_list` | Read |\n| Nodes | `node_list`, `node_describe` | Read |\n| | `cordon_node`, `uncordon_node`, `drain_node` | Write |\n| Namespaces | `namespace_list` | Read |\n| | `create_namespace`, `delete_namespace` | Write |\n| Metrics (top) | `node_top`, `pod_top` | Read |\n| Cluster | `cluster_info`, `api_resources` | Read |\n| Events | `event_list` | Read |\n| Diagnostics / RCA | `pod_health_rca`, `workload_readiness_rca` | Read |\n| Undo | `undo_list` | Read |\n| | `undo_apply` | Write |\n\n**Security — secrets**: `secret_list` returns secret names, types, and key NAMES only. Secret VALUES are never read, returned, or logged, and there is deliberately no tool that returns secret values.\n\n**Dry-run previews**: every write tool takes `dry_run: bool = False`. A dry run returns a `{\"dryRun\": true, \"wouldX\": ...}` preview without touching the cluster, and no undo descriptor is recorded for a preview.\n\n**Harness features that light up**: write tools with a clean inverse pass an `undo=` lambda so the harness records an inverse descriptor (with `_undo_id`) to the undo store — `scale_deployment`/`scale_statefulset` record a scale-back to their returned `previous_replicas`, `set_deployment_image` records a restore to the captured `previous_image`, `cordon_node` ↔ `uncordon_node` and `rollout_pause` ↔ `rollout_resume` are mutual inverses, and `create_namespace` records a `delete_namespace`. `drain_node` records a partial `uncordon_node` inverse (cordon is reversible; evictions are not). `delete_*` and `rollout_undo_deployment` declare no undo. `risk_level=high`: `delete_deployment`, `delete_job`, `delete_namespace`, `drain_node`, `rollout_undo_deployment`. `undo_list` (read) lists recorded reversible writes whose undo tokens have not been applied yet, and `undo_apply` (write) executes a recorded inverse — itself governed, single-use, and supports `dry_run`. All 55 tools are audit-logged under `~/.k8s-aiops/` and pass through the budget/runaway guard, each recorded with a descriptive risk-tier label. `pod_top`/`node_top` return a clear \"metrics-server not installed\" message (not an error) when metrics-server is absent. Avoid tight poll loops (re-listing pods every second) — the runaway breaker backs this up.\n\n## CLI Quick Reference\n\n```bash\nk8s-aiops init                                            # interactive onboarding wizard\nk8s-aiops pod list [-n <ns>] [-t <target>]\nk8s-aiops pod get <name> [-n <ns>]\nk8s-aiops pod describe <name> [-n <ns>]                   # status, container states, events\nk8s-aiops pod logs <name> [-n <ns>] [--tail 200] [-c <container>]\nk8s-aiops pod delete <name> [-n <ns>] [--dry-run]        # double confirm\nk8s-aiops deployment list|get|scale|restart|delete ...    # scale/restart: single confirm + --dry-run; delete: double confirm\nk8s-aiops rollout status|history|pause|resume <name> [-n <ns>]\nk8s-aiops rollout set-image <name> <container> <image> [-n <ns>]\nk8s-aiops rollout undo <name> [--to-revision N] [--dry-run]   # double confirm\nk8s-aiops statefulset list|get|scale ...\nk8s-aiops daemonset list|get ...\nk8s-aiops job list|get|delete ...                         # delete: double confirm\nk8s-aiops cronjob list|get ...\nk8s-aiops service list [-n <ns>]\nk8s-aiops ingress list|get [-n <ns>]\nk8s-aiops configmap list|get [-n <ns>]\nk8s-aiops secret list [-n <ns>]                           # names/keys only — never values\nk8s-aiops storage pvc-list|pvc-get|pv-list|class-list\nk8s-aiops top pod|node                                    # requires metrics-server\nk8s-aiops node list|describe\nk8s-aiops node cordon|drain <name> [--dry-run]           # double confirm\nk8s-aiops node uncordon <name>\nk8s-aiops namespace list|create\nk8s-aiops namespace delete <name> [--dry-run]            # double confirm\nk8s-aiops cluster-info\nk8s-aiops api-resources\nk8s-aiops events [-n <ns>]\nk8s-aiops diagnose pod-health [-n <ns>] [-l <selector>]         # read-only RCA: crashloop/imagepull/OOM/unschedulable/restarts\nk8s-aiops diagnose workload-readiness [-n <ns>]                 # read-only RCA: ready<desired / stuck rollouts\nk8s-aiops doctor\nk8s-aiops mcp                                             # start MCP server (stdio)\n```\n\nSee `references/cli-reference.md` for the full command list.\n\n## Troubleshooting\n\n### \"Could not load kubeconfig … context not found\"\nThe named context does not exist in your kubeconfig. Run `kubectl config get-contexts` and set the target's `context:` to a listed name (or omit it to use current-context).\n\n### \"Authentication/authorization failed (401/403)\"\nThe kube context lacks the RBAC verb for the resource. Check with `kubectl auth can-i <verb> <resource> -n <ns>` and switch to a context/ServiceAccount with adequate roles. For EKS/GKE/AKS, confirm the exec-plugin (aws/gcloud/az CLI) is installed and logged in.\n\n### \"Resource not found (404)\"\nThe pod/deployment/node name or namespace is wrong, or the object was deleted. List the parent collection first (`pod list`, `deployment list`, `node list`) to get a current name. Remember most commands default to the `default` namespace unless `-n` is given.\n\n### \"Conflict (409)\"\nThe object changed concurrently (or already exists). Re-read it and retry the write.\n\n### Logs are empty or truncated\n`pod logs` returns the trailing `--tail` lines (default 100); raise `--tail`. For a multi-container pod, pass `-c <container>` or the API returns an error naming the available containers.\n\n## Audit & Safety\n\nAll operations are automatically audited via the bundled `@governed_tool` decorator (`k8s_aiops.governance`):\n- Every tool call logged to `~/.k8s-aiops/audit.db` (local SQLite audit DB; relocate with `K8S_AIOPS_HOME`)\n- Budget / runaway guard caps cumulative tool calls and wall-time, and trips on tight poll/retry loops — a safety backstop, not authorization\n- Undo store records inverse descriptors for reversible writes (scale → previous replicas; cordon ↔ uncordon)\n- Each write carries a descriptive risk-tier label into its audit row — a label, not a gate; `K8S_AUDIT_APPROVED_BY` / `K8S_AUDIT_RATIONALE` are optional annotations recorded when set, never required\n\n**Authorization is not this tool's job.** There is no read-only switch, policy file, or approval gate. Whether a write is permitted is the agent's judgement or the RBAC of the kubeconfig context you connect with — give it a read-only ServiceAccount and writes fail at the apiserver, the place that owns the permission.\n\nThe harness is bundled in the package — no external dependency, no manual setup. See `references/setup-guide.md` for security details.\n\nDriving these tools with a smaller / local model? See `references/agent-guardrails.md` — which guardrails the tool now enforces for you, plus a ready-to-paste system prompt.\n\n## Contributing & feature requests\n\nCoverage is intentionally focused. **Missing a device, action, or feature you need?** Open an issue or pull request at [github.com/AIops-tools/K8s-AIops](https://github.com/AIops-tools/K8s-AIops/issues) — feature requests, contributions, and comments are all welcome.\n\n## License\n\nMIT — [github.com/AIops-tools/K8s-AIops](https://github.com/AIops-tools/K8s-AIops)\n\nFile v0.13.1:_meta.json\n\n{\n  \"ownerId\": \"kn7b067awq2s97bn3d7p5qfhw5827pxc\",\n  \"slug\": \"k8s-aiops\",\n  \"version\": \"0.13.1\",\n  \"publishedAt\": 1789207828688\n}\n\nFile v0.13.1:references/agent-guardrails.md\n\n# Agent guardrails — running k8s-aiops with a smaller / local model\n\nIf you drive these tools with a local model (Llama, Qwen, Mistral … via Goose,\nOllama, LM Studio, or any OpenAI-compatible runtime), you will get noticeably\nbetter results with a short system prompt. This page gives you one, and — more\nimportantly — tells you which guardrails you **no longer need to write**, because\nthe tool now enforces them itself.\n\nThe distinction matters. A guardrail in a prompt is a request. A guardrail in the\nharness is a guarantee. Anything below that we could move into the harness, we did.\n\n## Authorization is not this tool's job — decide it where it belongs\n\nWhether a write should happen is your decision, or the account's. The tool does\nnot gate it — there is no read-only switch and no approval prompt to configure.\nThe two right places to control read vs write:\n\n- **The kubeconfig context you connect with.** Bind it to a ServiceAccount or\n  user whose RBAC grants only `get`/`list`/`watch`, and every write fails at the\n  apiserver — the only place the permission actually lives. No skill-side flag\n  can be argued around by a model, but a revoked RBAC verb cannot be. This is\n  strictly stronger than any in-process switch: it is enforced at the cluster.\n- **Your agent's system prompt.** If you want an observe-only session, tell the\n  model not to call the write tools (they are clearly tagged `[WRITE]`).\n\nWhat the tool *does* guarantee is that you can always see what happened:\n\n## What the tool enforces — do not waste prompt budget on these\n\n| You might be tempted to prompt | Why you don't need to |\n|---|---|\n| \"Log everything you do, over both MCP and the CLI\" | Every operation is audited to `~/.k8s-aiops/audit.db` regardless of what the model says it did — and the CLI writes the same row the MCP path does, so there is no unaudited entry point. Reversible writes also record an undo token capturing the *prior* state. |\n| \"Don't invent a value when a field is missing\" | A field the apiserver did not return comes back as `null`, never as `\"\"`. An unscheduled pod's `node` is `null`; a pod with no phase yet has `phase: null`; an object with no readable creation timestamp has `age: null`. Absent and empty are distinguishable in the payload. |\n| \"Tell me if the output was cut off\" | The limit-bearing reads return an envelope: `event_list` → `{\"events\": [...], \"returned\": N, \"limit\": L, \"truncated\": true/false}`, and `undo_list` → `{\"undos\": [...], \"returned\": N, \"limit\": L, \"truncated\": ...}`. Truncation is **measured** (one extra row is fetched), not guessed from `len(rows) == limit`. |\n| \"Preserve the ordering / tell me what's most urgent\" | `pod_health_rca` and `workload_readiness_rca` findings carry an explicit 1-based `rank`, worst-first, and each finding's `detail` cites the measured signal (the waiting reason, the restart count, the ready/desired ratio). Priority is in the payload, not implied by list position. |\n| \"Confirm before anything destructive\" | The destructive CLI commands (deployment/job/namespace delete, node cordon/drain, rollout undo) are `--dry-run`-able and require double confirmation. |\n| \"Don't get stuck retrying\" | The runaway guard trips a circuit breaker if the same call is hammered in a tight loop — a stuck agent is stopped rather than left to burn calls and time. It is a safety backstop, not an authorization gate. |\n\n## What still needs a prompt\n\nThese are model-behaviour problems the harness cannot fix from the outside.\nCopy this into your agent's system prompt:\n\n```text\nYou operate a Kubernetes cluster through the k8s-aiops MCP tools.\n\nTOOL USE\n- Before answering any question about the current cluster, you MUST call a\n  tool. Never answer from memory or assumption.\n- Actually invoke the tool. Do not describe the call you would make, and do not\n  emit an example JSON response in place of calling it.\n- If a tool call fails, report the real error verbatim. Never fill the gap with\n  a plausible-sounding answer.\n\nREADING RESULTS\n- Read the whole result before concluding. If a result contains a \"truncated\"\n  field that is true, say so and re-run with a higher limit instead of treating\n  the partial result as complete.\n- A null field means the apiserver did not return that value. Report it as \"not\n  available\" — never infer it. In particular, a null \"node\" means the pod is\n  not scheduled; it does not mean the node is unknown or missing.\n- Report values exactly as returned. Do not normalise, translate, or prettify\n  phases, conditions, container-state reasons, or resource names.\n- When an RCA result has findings, work in \"rank\" order and cite the measured\n  number in each finding's \"detail\".\n\nSCOPE AND IDENTIFIERS\n- Always state the namespace you are talking about. A bare pod or deployment\n  name is ambiguous — the same name exists in many namespaces.\n- Omitting the namespace means ALL namespaces, not the default one. Never\n  silently widen a namespaced question into a cluster-wide answer.\n- Do not confuse a namespace with a context (a cluster), a pod name with its\n  deployment name, or a container name with the pod that contains it. A pod\n  name generated by a ReplicaSet is not a stable identifier.\n- Separate observation from interpretation. State what the tools returned, then\n  any interpretation, clearly marked as such.\n- Do not assert a capacity, performance, or availability problem unless a tool\n  result supports it. Do not add generic advice that does not follow from the\n  tool output.\n```\n\n## Recommended setup for a local model\n\nStart with a kubeconfig context that *cannot* write — a ServiceAccount whose\nRBAC grants only `get`/`list`/`watch` — verify, and widen its permission only\nwhen you trust the setup. RBAC is enforced at the apiserver, so a write fails at\nthe cluster no matter what the model attempts:\n\n```bash\n# Point KUBECONFIG at a read-only ServiceAccount context, then:\nk8s-aiops doctor\n```\n\nOptionally annotate the audit trail with who is operating and why — recorded on\nevery row, never required:\n\n```bash\nexport K8S_AUDIT_APPROVED_BY=\"your.name@example.com\"\nexport K8S_AUDIT_RATIONALE=\"incident INC-1234 — restart the stuck payments rollout\"\n```\n\n## Kubernetes-specific notes\n\n- **Namespace confusion is the single most common failure mode.** Every\n  namespaced tool takes an optional `namespace`; omitting it lists across **all**\n  namespaces. A smaller model reliably drops the namespace from a follow-up call\n  and then reports a cluster-wide count as if it were the namespace's. Pin the\n  namespace in your prompt (or in the target's `namespace:` in\n  `~/.k8s-aiops/config.yaml`) and ask the model to restate it in every answer.\n- **Context confusion is the expensive one.** `target` selects a *cluster*, not a\n  namespace. If you have prod and staging contexts in the same kubeconfig, give\n  the agent a config with only the cluster it should touch — the model has no way\n  to notice it is on the wrong one, because both look plausible.\n- **Authentication is delegated to your kubeconfig.** Unlike the other tools in\n  this line, k8s-aiops has **no credential store**: there is nothing to encrypt\n  and no master password. (`k8s-aiops secret ...` lists Kubernetes *Secret\n  resources* by name — it is not a credential manager.) The agent inherits the RBAC of the\n  kubeconfig user. That makes RBAC your strongest guardrail — a read-only\n  ServiceAccount kubeconfig enforces read-only at the *cluster*, the only place\n  the permission truly lives.\n- **`drain_node` is the most dangerous tool here.** It evicts pods cluster-wide\n  in effect, and its blast radius is not visible in its arguments. Keep it out of\n  the RBAC role you hand the agent unless you specifically intend node drains.\n- **Pod names are not stable.** A ReplicaSet-generated pod name changes on every\n  rollout, so an id the model cached earlier in the conversation may already be\n  gone. Prefer `deployment_get` / `rollout_status` over re-using a pod name.\n- **Prefer the RCA tools over multi-step chains.** `pod_health_rca` and\n  `workload_readiness_rca` do the list-then-correlate work inside one call, so a\n  smaller model does not have to chain reads and keep names/namespaces straight.\n- **Secrets are names-only by design.** `secret_list` returns key *names*, never\n  values — there is no tool that can exfiltrate a secret payload.\n\n## If your model still struggles\n\nSome behaviours are model-capacity limits rather than prompt problems:\n\n- **Multi-tool workflows time out or drift.** Prefer the `*_rca` tools — they do\n  the multi-step correlation inside one call.\n- **The model ignores later tool results in a long context.** Ask narrower\n  questions, always scope to a namespace, and use `limit` deliberately rather\n  than pulling whole-cluster inventories (`pod_list` with no namespace on a busy\n  cluster will bury everything else in the context).\n- **The model describes calls instead of making them.** This is usually a\n  runtime/tool-calling-format mismatch, not a prompt problem — check that your\n  client advertises the tools in the format your model was trained on.\n\nFeedback on running this with a specific local model is genuinely useful —\nopen an issue at\n[github.com/AIops-tools/K8s-AIops](https://github.com/AIops-tools/K8s-AIops/issues)\nwith the model, runtime, and what went wrong.\n\nFile v0.13.1:references/capabilities.md\n\n# k8s-aiops Capabilities\n\n55 MCP tools (39 read / 16 write). Every tool is wrapped with `@governed_tool`\n(audit + policy + budget + risk-tier; undo where a clean inverse exists). Returns\nare high-signal summaries — `_get` / `_describe` tools add detail for a single object.\n\n## Read tools\n\n| Tool | Returns | Risk |\n|------|---------|:----:|\n| `pod_list` | name, namespace, phase, ready, restarts, node, age | low |\n| `pod_get` | + host_ip, pod_ip, containers | low |\n| `pod_describe` | status, conditions, container states + restart counts, recent events | low |\n| `pod_logs` | trailing log lines (default 100) | low |\n| `deployment_list` / `deployment_get` | replicas summary / + strategy, images | low |\n| `rollout_status` | desired/updated/available/unavailable + paused | low |\n| `rollout_history` | revisions (from replicasets) with images | low |\n| `statefulset_list` / `statefulset_get` | desired/ready/current / + service, images | low |\n| `daemonset_list` / `daemonset_get` | desired/ready/available / + images | low |\n| `replicaset_list` | name, namespace, desired/ready, age | low |\n| `job_list` / `job_get` | completions, succeeded/failed/active | low |\n| `cronjob_list` / `cronjob_get` | schedule, suspend, active, last schedule | low |\n| `service_list` | name, namespace, type, cluster IP, ports | low |\n| `ingress_list` / `ingress_get` | class, hosts / + path→backend rules | low |\n| `endpoints_list` | ready addresses + ports | low |\n| `configmap_list` / `configmap_get` | key count / keys + values | low |\n| `secret_list` | names, types, key NAMES only (values redacted) | low |\n| `pvc_list` / `pvc_get` | status, capacity, class / + access modes | low |\n| `pv_list` | capacity, status, claim, class | low |\n| `storageclass_list` | provisioner, reclaim policy, default | low |\n| `node_list` / `node_describe` | status, roles / capacity, allocatable, conditions, taints | low |\n| `namespace_list` | name, phase, age | low |\n| `pod_top` / `node_top` | CPU/mem via metrics-server (graceful if absent) | low |\n| `cluster_info` | server version, node/ready/namespace counts | low |\n| `api_resources` | available API groups + versions | low |\n| `event_list` | type, reason, object, namespace, message, age | low |\n| `pod_health_rca` | worst-first findings: CrashLoopBackOff, image-pull, OOMKilled, unschedulable, high restarts (each cites the reason/count) | low |\n| `workload_readiness_rca` | worst-first findings: ready<desired, zero-ready outages, stuck rollouts (Deployment/StatefulSet/DaemonSet) | low |\n| `undo_list` | recorded reversible writes / not-yet-applied undo tokens | low |\n\n## Write tools\n\n| Tool | Effect | Risk | Undo |\n|------|--------|:----:|------|\n| `scale_deployment` | set replica count | medium | scale back to `previous_replicas` |\n| `scale_statefulset` | set replica count | medium | scale back to `previous_replicas` |\n| `rollout_restart_deployment` | patch `restartedAt` annotation | medium | none (pods already rolling) |\n| `rollout_pause` / `rollout_resume` | toggle `spec.paused` | medium | each other |\n| `rollout_undo_deployment` | roll back to a prior revision | **high** | none |\n| `set_deployment_image` | update a container image | medium | restore `previous_image` |\n| `delete_pod` | delete a pod | medium | none (controller recreates) |\n| `delete_deployment` | delete deployment + pods | **high** | none |\n| `delete_job` | delete a job + pods | **high** | none |\n| `create_namespace` | create a namespace | medium | `delete_namespace` |\n| `delete_namespace` | delete namespace + everything in it | **high** | none |\n| `cordon_node` / `uncordon_node` | toggle schedulability | medium | each other |\n| `drain_node` | cordon + evict pods (skips DaemonSet/mirror) | **high** | partial: `uncordon_node` |\n| `undo_apply` | execute a recorded inverse (itself governed, single-use, supports `dry_run`) | medium | n/a (is the undo) |\n\n### `delete_namespace` refusals\n\nTwo targets are refused before anything is deleted, on the real call and on the\n`dry_run` preview alike (a preview never green-lights a delete that would be\nrejected):\n\n- **Control-plane namespaces** — `kube-system`, `kube-public`, `kube-node-lease`.\n  Deleting one takes CoreDNS, kube-proxy, the CNI and node heartbeats with it.\n  Pass `confirm=true` (CLI `--confirm`) to proceed deliberately, e.g. when\n  tearing a cluster down or clearing a namespace stuck `Terminating`.\n- **The namespace holding this target's own ServiceAccount credential** — that\n  delete revokes the credential it is running on, and `delete_namespace` has no\n  undo. **`confirm` does not override this one.** Re-run from a context whose\n  credential lives elsewhere. Only applies when the kubeconfig authenticates\n  with a ServiceAccount token; certificate and `exec` (EKS/GKE/AKS) credentials\n  are not namespace-bound, so the check stands down rather than guessing.\n\n## Token-budget notes\n\n- List tools accept a `namespace` filter to keep responses small; events and pod\n  listings also accept `limit` / `label_selector` where applicable.\n- Prefer `pod_get` / `deployment_get` over re-listing when you already have a name.\n- The runaway guard trips on tight poll loops — wait between repeated list calls.\n\n## Design notes / Kubernetes-client assumptions\n\n- Authentication is delegated to the kubeconfig; the skill never touches raw\n  credentials (works with client certs, tokens, and EKS/GKE/AKS exec plugins).\n- Typed Api clients (`CoreV1Api`, `AppsV1Api`, `BatchV1Api`, `NetworkingV1Api`,\n  `StorageV1Api`, `CustomObjectsApi`, `ApisApi`, `VersionApi`) are cached per kube\n  context in a module dict — third-party client objects are never monkey-patched.\n- `secret_list` reads only key NAMES from `secret.data` — secret values are never\n  read, returned, or logged, and no tool exposes them.\n- `pod_top` / `node_top` use the `metrics.k8s.io/v1beta1` API; when metrics-server\n  is absent the 404/503 is caught and returned as `{available: false, message}`.\n- `ApiException` is translated centrally at the connection layer into a teaching\n  `K8sApiError` (404/403/409/5xx), so agents see actionable messages, not tracebacks.\n\nFile v0.13.1:references/cli-reference.md\n\n# k8s-aiops CLI Reference\n\nAll commands accept `-t/--target <name>` to select a configured target (a kube\ncontext). Namespaced commands accept `-n/--namespace <ns>`; omit it to use the\ntarget's default namespace (read lists fall back to all-namespaces).\n\n## Onboarding\n\n```bash\nk8s-aiops init                    # interactive wizard: register kube contexts as targets\n```\n\n## Pods\n\n```bash\nk8s-aiops pod list [-n <ns>] [-t <target>]\nk8s-aiops pod get <name> [-n <ns>]\nk8s-aiops pod describe <name> [-n <ns>]              # status, container states, events\nk8s-aiops pod logs <name> [-n <ns>] [--tail N] [-c <container>]\nk8s-aiops pod delete <name> [-n <ns>] [--dry-run]    # destructive: double confirm\n```\n\n## Deployments & Rollouts\n\n```bash\nk8s-aiops deployment list [-n <ns>]\nk8s-aiops deployment get <name> [-n <ns>]\nk8s-aiops deployment scale <name> <replicas> [-n <ns>]\nk8s-aiops deployment restart <name> [-n <ns>]        # rolling restart\nk8s-aiops deployment delete <name> [-n <ns>] [--dry-run]   # HIGH RISK: double confirm\nk8s-aiops rollout status <name> [-n <ns>]\nk8s-aiops rollout history <name> [-n <ns>]\nk8s-aiops rollout pause|resume <name> [-n <ns>]\nk8s-aiops rollout set-image <name> <container> <image> [-n <ns>]\nk8s-aiops rollout undo <name> [--to-revision N] [--dry-run]   # HIGH RISK: double confirm\n```\n\n## StatefulSets / DaemonSets / Jobs / CronJobs\n\n```bash\nk8s-aiops statefulset list|get [-n <ns>]\nk8s-aiops statefulset scale <name> <replicas> [-n <ns>]\nk8s-aiops daemonset list|get [-n <ns>]\nk8s-aiops job list|get [-n <ns>]\nk8s-aiops job delete <name> [-n <ns>] [--dry-run]    # destructive: double confirm\nk8s-aiops cronjob list|get [-n <ns>]\n```\n\n## Services, Ingress, Config, Storage\n\n```bash\nk8s-aiops service list [-n <ns>]\nk8s-aiops ingress list|get [-n <ns>]\nk8s-aiops configmap list|get [-n <ns>]\nk8s-aiops secret list [-n <ns>]                      # names + key NAMES only, never values\nk8s-aiops storage pvc-list|pvc-get [-n <ns>]\nk8s-aiops storage pv-list|class-list\n```\n\n## Nodes & Metrics\n\n```bash\nk8s-aiops node list\nk8s-aiops node describe <name>\nk8s-aiops node cordon <name> [--dry-run]             # destructive: double confirm\nk8s-aiops node uncordon <name>\nk8s-aiops node drain <name> [--dry-run]              # HIGH RISK: double confirm\nk8s-aiops top pod|node                               # requires metrics-server\n```\n\n## Namespaces, Cluster & Events\n\n```bash\nk8s-aiops namespace list\nk8s-aiops namespace create <name>\nk8s-aiops namespace delete <name> [--dry-run]        # HIGH RISK: double confirm\nk8s-aiops cluster-info\nk8s-aiops api-resources\nk8s-aiops events [-n <ns>]\n```\n\n## Diagnostics & MCP\n\n```bash\nk8s-aiops diagnose pod-health [-n <ns>] [-l <selector>]  # RCA: crashloop/imagepull/OOM/unschedulable/restarts (read-only)\nk8s-aiops diagnose workload-readiness [-n <ns>]          # RCA: ready<desired / stuck rollouts (read-only)\nk8s-aiops doctor [--skip-auth]    # check config + cluster reachability\nk8s-aiops mcp                     # start the MCP server over stdio\n```\n\n## Flags summary\n\n| Flag | Meaning |\n|------|---------|\n| `-t, --target` | Target name from `~/.k8s-aiops/config.yaml` |\n| `-n, --namespace` | Namespace scope |\n| `--tail` | Trailing log lines (pod logs, default 100) |\n| `-c, --container` | Container name (pod logs) |\n| `--dry-run` | Preview a destructive op without executing |\n| `--to-revision` | Rollout revision (`rollout undo`, 0 = previous) |\n| `--skip-auth` | Skip the connectivity check in `doctor` |\n\nFile v0.13.1:references/setup-guide.md\n\n# k8s-aiops Setup Guide\n\n## Install\n\n```bash\nuv tool install k8s-aiops\nk8s-aiops doctor\n```\n\n`k8s-aiops` requires Python ≥ 3.11. If `uv` picked an older interpreter:\n\n```bash\nuv python install 3.12\nuv tool install --python 3.12 --force k8s-aiops\n```\n\n## Connecting to a cluster\n\nk8s-aiops reads your **kubeconfig** — the same one `kubectl` uses\n(`KUBECONFIG` env var, else `~/.kube/config`). Out of the box it talks to your\ncurrent kube-context, so a fresh install needs no config file:\n\n```bash\nk8s-aiops pod list\n```\n\n### Named targets (multiple clusters)\n\nThe fastest way is the interactive wizard, which discovers the contexts in your\nkubeconfig and registers the ones you pick (writing `config.yaml`, dir chmod 700):\n\n```bash\nk8s-aiops init\n```\n\nOr create `~/.k8s-aiops/config.yaml` by hand to give contexts friendly names:\n\n```yaml\ntargets:\n  - name: prod\n    context: prod-eks            # a context from `kubectl config get-contexts`\n    namespace: default           # optional default namespace\n    # kubeconfig: /path/to/kubeconfig   # optional, overrides KUBECONFIG/~/.kube/config\n  - name: lab\n    context: k3s-lab\n```\n\nThen select with `-t`:\n\n```bash\nk8s-aiops -t prod pod list -n payments\n```\n\nNo secrets are stored here — authentication lives entirely in the kubeconfig.\n\n### Works with\n\nStandard Kubernetes, k3s, EKS, GKE, AKS, kind, minikube. For managed clusters\n(EKS/GKE/AKS) the kubeconfig uses an exec plugin (`aws`/`gcloud`/`az`); make sure\nthat CLI is installed and logged in.\n\n## Security\n\n> **Disclaimer**: This is a community-maintained open-source project and is **not affiliated with, endorsed by, or sponsored by the Cloud Native Computing Foundation, the Kubernetes project, or k3s/Rancher.** \"Kubernetes\" and \"k3s\" are trademarks of their respective owners. Source is auditable at [github.com/AIops-tools/K8s-AIops](https://github.com/AIops-tools/K8s-AIops) under the MIT license.\n\n1. **Source code** — [github.com/AIops-tools/K8s-AIops](https://github.com/AIops-tools/K8s-AIops), MIT.\n2. **Config file contents** — `config.yaml` holds only target names, kube context\n   names, and optional namespaces/paths. No credentials.\n3. **Credentials** — delegated to the kubeconfig; never read, logged, or echoed by\n   k8s-aiops. Keep `~/.k8s-aiops` owner-only (`chmod 700`).\n4. **TLS verification** — follows the kubeconfig (`certificate-authority` /\n   `insecure-skip-tls-verify`); the skill does not weaken it.\n5. **Prompt-injection protection** — all API-returned text (names, log lines, event\n   messages) is run through `sanitize()` (truncation + control-character stripping).\n6. **Least privilege** — bind the kube context to a ServiceAccount/user with only the\n   RBAC verbs you need (read-only needs `get`/`list`/`watch`; writes add\n   `patch`/`delete`).\n\n## Governance harness\n\nBundled under `k8s_aiops.governance` — no external dependency. State lives under\n`~/.k8s-aiops/` (override with `K8S_AIOPS_HOME`):\n\n- `audit.db` — every tool call (skill, tool, params, status, duration, agent),\n  each carrying a descriptive risk-tier label derived from the tool's\n  `risk_level`. The tier is a label, not a gate.\n- Token/runaway budget guard (`K8S_MAX_TOOL_CALLS`, `K8S_MAX_TOOL_SECONDS`,\n  `K8S_RUNAWAY_MAX`, `K8S_RUNAWAY_WINDOW_SEC`) — a safety backstop, not\n  authorization.\n- Undo store — inverse descriptors for reversible writes.\n- Accountability: `K8S_AUDIT_APPROVED_BY` / `K8S_AUDIT_RATIONALE` are optional\n  annotations recorded on the audit row when set — never required, never\n  blocking. Authorization is the RBAC of the kubeconfig context, not this tool.\n\n## MCP client config\n\n```jsonc\n{\n  \"command\": \"k8s-aiops\",\n  \"args\": [\"mcp\"],\n  \"env\": { \"K8S_AIOPS_CONFIG\": \"~/.k8s-aiops/config.yaml\" }\n}\n```\n\nFallback (no `uv tool install`): `uvx --from k8s-aiops k8s-aiops-mcp`. Prefer the\ninstalled entry point — it does not re-resolve PyPI at launch.\n\n## Static analysis\n\n```bash\nuvx bandit -r k8s_aiops/ mcp_server/\n```\n\nFile v0.13.1:skill-card.md\n\n## Description:\n\nk8s-aiops helps agents inspect, diagnose, and operate kubeconfig-reachable Kubernetes clusters with audited read and write workflows.\n\nThis skill is ready for commercial/non-commercial use.\n\n## Publisher:\n\n[zw008](https://clawhub.ai/user/zw008)\n\n### License/Terms of Use:\n\nMIT-0\n\n## Use Case:\n\nDevelopers, platform engineers, and operators use this skill to list Kubernetes resources, review logs and events, run read-only diagnostics, and perform governed operational actions such as scaling, rollouts, deletion, cordon, and drain. It is intended for clusters reachable through the user's kubeconfig, including standard Kubernetes, k3s, EKS, GKE, and AKS.\n\n### Deployment Geography for Use:\n\nGlobal\n\n## Known Risks and Mitigations:\n\nRisk: The skill can expose destructive Kubernetes actions without an enforced skill-side read-only or approval gate.\n\nMitigation: Use a deliberately scoped kubeconfig, prefer a read-only ServiceAccount by default, and withhold RBAC for node drain and namespace deletion unless those operations are explicitly needed.\n\nRisk: An agent using a broad kubeconfig could perform unintended cluster changes.\n\nMitigation: Review proposed write actions, use dry-run previews where available, and verify the exact k8s-aiops package version and source before use.\n\n## Reference(s):\n\n- [ClawHub skill page](https://clawhub.ai/zw008/skills/k8s-aiops)\n- [Project homepage](https://github.com/AIops-tools/K8s-AIops)\n- [agent-guardrails.md](references/agent-guardrails.md)\n- [capabilities.md](references/capabilities.md)\n- [cli-reference.md](references/cli-reference.md)\n- [setup-guide.md](references/setup-guide.md)\n\n## Skill Output:\n\n**Output Type(s):** [guidance, markdown, shell commands, configuration]\n\n**Output Format:** [Markdown with inline shell commands and structured Kubernetes operation guidance]\n\n**Output Parameters:** [1D]\n\n**Other Properties Related to Output:** [May include dry-run recommendations, RBAC checks, audit context, and Kubernetes resource names scoped to the selected kubeconfig target.]\n\n## Skill Version(s):\n\n0.13.1 (source: server release metadata)\n\n## Ethical Considerations:\n\nUsers should evaluate whether this skill is appropriate for their environment, review any generated or modified files before relying on them, and apply their organization's safety, security, and compliance requirements before deployment.\n\nArchive v0.13.0: 7 files, 19471 bytes\n\nFiles: references/agent-guardrails.md (9366b), references/capabilities.md (6127b), references/cli-reference.md (3495b), references/setup-guide.md (3987b), skill-card.md (2379b), SKILL.md (18186b), _meta.json (129b)\n\nFile v0.13.0:SKILL.md\n\n---\nname: k8s-aiops\nslug: k8s-aiops\ndisplayName: \"k8s AIops\"\nsummary: \"Governed Kubernetes ops — 55 MCP tools with audit, budget, undo, risk-tier audit labels.\"\nlicense: MIT\nhomepage: https://github.com/AIops-tools/K8s-AIops\ntags: [aiops, mcp, governance, k8s]\ndescription: >\n  Use this skill whenever the user needs to operate a Kubernetes cluster — list/inspect pods, deployments, statefulsets, daemonsets, replicasets, jobs, cronjobs, services, ingresses, endpoints, configmaps, secrets (names/keys only), PVCs/PVs/storageclasses, nodes, namespaces, and events; read pod logs; describe pods/nodes; pod/node top (metrics); read-only diagnostics / RCA (pod-health, workload-readiness); scale deployments/statefulsets; rollout status/history/undo/pause/resume and set image; delete pods/deployments/jobs; create/delete namespaces; and cordon/uncordon/drain nodes. Works with any kubeconfig-reachable cluster (standard Kubernetes, k3s, EKS, GKE, AKS).\n  Always use this skill for \"list k8s pods\", \"scale deployment\", \"kubernetes pod logs\", \"describe pod\", \"why is my pod crashing\", \"diagnose pods\", \"which deployments are unhealthy\", \"rollout undo\", \"set image\", \"top pods\", \"drain node\", \"cordon node\", \"restart deployment\", \"k3s\", or \"kubectl\"-style tasks when the context is explicitly Kubernetes / a cluster.\n  Do NOT use when the target is not a Kubernetes cluster (hypervisor VM lifecycle, backup products, or cloud-provider consoles are out of scope).\n  Common Kubernetes operations with a built-in governance harness (audit, token budget, undo, risk-tier labels).\ninstaller:\n  kind: uv\n  package: k8s-aiops\nargument-hint: \"[resource name or describe your Kubernetes task]\"\nallowed-tools:\n  - Bash\nmetadata: {\"openclaw\":{\"requires\":{\"anyBins\":[\"k8s-aiops\",\"uvx\"]},\"optional\":{\"env\":[\"K8S_AIOPS_CONFIG\",\"KUBECONFIG\",\"K8S_AIOPS_HOME\"]},\"homepage\":\"https://github.com/AIops-tools/K8s-AIops\",\"emoji\":\"☸️\",\"os\":[\"macos\",\"linux\"]}}\ncompatibility: >\n  Standalone, self-governed Kubernetes operations. The governance harness (audit, token/runaway budget, undo, risk-tier labels) is bundled in the package — no external skill-family dependency.\n  All write operations are audited to a local SQLite DB under ~/.k8s-aiops/ (relocatable via K8S_AIOPS_HOME).\n  Credentials: k8s-aiops handles NO credentials directly — authentication is delegated to the kubeconfig (KUBECONFIG env or ~/.kube/config), which may hold client certs, bearer tokens, or exec plugins (EKS/GKE/AKS). Underlying credentials are never read, logged, or echoed. The state dir ~/.k8s-aiops should be chmod 700.\n  Destructive operations (deployment/job/namespace delete, node cordon/drain, rollout undo) require double confirmation at the CLI layer and support --dry-run. All write tools pass through the @governed_tool decorator (budget/runaway guard + audit + a descriptive risk-tier label). Reversible writes record an inverse undo descriptor (scale_deployment/scale_statefulset restore the previous replica count; set_deployment_image restores the previous image; cordon_node ↔ uncordon_node and rollout_pause ↔ rollout_resume; create_namespace ↔ delete_namespace); delete_* and rollout_undo record none. risk_level=high: delete_deployment, delete_job, delete_namespace, drain_node, rollout_undo_deployment. Secret VALUES are never read or returned by any tool.\n  Webhooks: none — no outbound network calls beyond the configured Kubernetes API server.\n  TLS: follows the kubeconfig (certificate-authority / insecure-skip-tls-verify); the skill does not weaken it.\n  Transitive dependencies: the official kubernetes Python client and the MCP SDK. No post-install scripts or background services.\n---\n\n# k8s AIops\n\n> **Disclaimer**: This is a community-maintained open-source project and is **not affiliated with, endorsed by, or sponsored by the Cloud Native Computing Foundation, the Kubernetes project, or k3s/Rancher.** \"Kubernetes\" and \"k3s\" are trademarks of their respective owners. Source code is publicly auditable at [github.com/AIops-tools/K8s-AIops](https://github.com/AIops-tools/K8s-AIops) under the MIT license.\n\nGoverned Kubernetes operations — **55 MCP tools**, every one wrapped with the bundled `@governed_tool` harness: a local unified audit log under `~/.k8s-aiops/`, a token/runaway budget guard, undo-token recording, and a descriptive risk-tier label on every audit row. Works with any kubeconfig-reachable cluster (standard Kubernetes, k3s, EKS, GKE, AKS). Run `k8s-aiops init` for a friendly onboarding wizard that registers your kube contexts as named targets.\n\n> **Standalone**: the governance harness is bundled in the package (`k8s_aiops.governance`) — k8s-aiops has no external skill-family dependency. Coverage focuses on common operations and is not yet exhaustive.\n\n## What This Skill Does\n\n| Category | Tools | Count | Read or Write |\n|----------|-------|:-----:|:-------------:|\n| **Pods** | list, get, logs, describe, delete | 5 | 4 read / 1 write |\n| **Deployments** | list, get, scale, rollout restart, delete | 5 | 2 read / 3 write |\n| **Rollout** | status, history, undo, pause, resume, set-image | 6 | 2 read / 4 write |\n| **StatefulSets** | list, get, scale | 3 | 2 read / 1 write |\n| **DaemonSets** | list, get | 2 | 2 read |\n| **ReplicaSets** | list | 1 | 1 read |\n| **Jobs / CronJobs** | job list/get/delete, cronjob list/get | 5 | 4 read / 1 write |\n| **Services / Ingress / Endpoints** | service list, ingress list/get, endpoints list | 4 | 4 read |\n| **Config / Secrets** | configmap list/get, secret list (names/keys only) | 3 | 3 read |\n| **Storage** | pvc list/get, pv list, storageclass list | 4 | 4 read |\n| **Nodes** | list, describe, cordon, uncordon, drain | 5 | 2 read / 3 write |\n| **Namespaces** | list, create, delete | 3 | 1 read / 2 write |\n| **Metrics (top)** | pod, node | 2 | 2 read |\n| **Cluster** | cluster_info, api_resources | 2 | 2 read |\n| **Events** | list | 1 | 1 read |\n| **Diagnostics / RCA** | pod-health, workload-readiness | 2 | 2 read |\n\n## Quick Install\n\n```bash\nuv tool install k8s-aiops\nk8s-aiops init            # friendly wizard: register your kube contexts as targets\nk8s-aiops doctor          # or skip init — works with your current kube-context too\n```\n\n## When to Use This Skill\n\n- List/inspect pods, deployments, services, nodes, namespaces and recent events\n- Read a pod's recent log lines to diagnose a crash loop\n- Run a read-only RCA sweep (`diagnose pod-health` / `diagnose workload-readiness`) to find the root cause worst-first\n- Scale a deployment up/down, or trigger a rolling restart\n- Delete a stuck pod (a controller recreates it) or a deployment\n- Cordon a node before maintenance, then uncordon it after\n\n**Do NOT use when** the target is not a Kubernetes cluster (hypervisor VM lifecycle, backup products, or cloud-provider consoles are out of scope for this skill).\n\n## Related Skills — Skill Routing\n\n| If the user wants… | Use |\n|--------------------|-----|\n| Kubernetes pods / deployments / nodes | **k8s-aiops** (this skill) |\n| Hypervisor VM lifecycle (power, snapshot, migrate) | a hypervisor ops skill |\n| Backup & restore | a backup ops skill |\n\n## Common Workflows\n\n### Diagnose a crash-looping pod and restart its deployment\n\n1. `k8s-aiops pod list -n prod` → find the pod with high `restarts` / non-Running `phase`\n2. `k8s-aiops pod logs <pod> -n prod --tail 200` → read the recent logs for the crash cause\n3. `k8s-aiops events -n prod` → check for `FailedScheduling` / image-pull events\n4. `k8s-aiops deployment restart <deploy> -n prod` → roll the deployment after fixing the cause\n5. **Failure branch**: if logs/events show an RBAC `403`, the kube context lacks the verb — run `kubectl auth can-i get pods -n prod` and switch to a context with adequate RBAC; the skill never retries a denied auth.\n\n### Triage an unhealthy namespace with RCA, then act on the worst finding\n\n1. `k8s-aiops diagnose pod-health -n prod` → worst-first findings; a `critical` `CrashLoopBackOff` on `prod/api` cites `restarts=9` and the exact `kubectl logs … --previous` action\n2. `k8s-aiops diagnose workload-readiness -n prod` → confirm the blast radius: e.g. `Deployment web ready 0/3` (`critical`, under-replicated)\n3. `k8s-aiops pod logs api-<hash> -n prod --tail 200 --previous`-equivalent via `k8s-aiops pod describe api-<hash> -n prod` → read the crash cause the RCA poi\n\nArchive v0.12.0: 7 files, 19605 bytes\n\nFiles: references/agent-guardrails.md (9366b), references/capabilities.md (6127b), references/cli-reference.md (3495b), references/setup-guide.md (3987b), skill-card.md (2815b), SKILL.md (18255b), _meta.json (129b)\n\nArchive v0.11.0: 7 files, 19466 bytes\n\nFiles: references/agent-guardrails.md (9366b), references/capabilities.md (6127b), references/cli-reference.md (3495b), references/setup-guide.md (3987b), skill-card.md (2369b), SKILL.md (18255b), _meta.json (129b)\n\nArchive v0.10.0: 7 files, 19556 bytes\n\nFiles: references/agent-guardrails.md (9366b), references/capabilities.md (6127b), references/cli-reference.md (3495b), references/setup-guide.md (3987b), skill-card.md (2730b), SKILL.md (18255b), _meta.json (129b)\n\nArchive v0.9.0: 7 files, 19515 bytes\n\nFiles: references/agent-guardrails.md (9366b), references/capabilities.md (6127b), references/cli-reference.md (3495b), references/setup-guide.md (3987b), skill-card.md (2604b), SKILL.md (18255b), _meta.json (128b)\n\nArchive v0.8.0: 7 files, 19589 bytes\n\nFiles: references/agent-guardrails.md (9366b), references/capabilities.md (6127b), references/cli-reference.md (3495b), references/setup-guide.md (3987b), skill-card.md (2709b), SKILL.md (18255b), _meta.json (128b)\n\nArchive v0.7.0: 7 files, 19275 bytes\n\nFiles: references/agent-guardrails.md (8588b), references/capabilities.md (6127b), references/cli-reference.md (3495b), references/setup-guide.md (3790b), skill-card.md (2923b), SKILL.md (18314b), _meta.json (128b)","readmeExcerpt":"Skill: k8s-aiops Owner: zw008 Summary: Use this skill whenever the user needs to operate a Kubernetes cluster — list/inspect pods, deployments, statefulsets, daemonsets, replicasets, jobs, cronjobs, services, ingresses, endpoints, configmaps, secrets (names/keys only), PVCs/PVs/storageclasses, nodes, namespaces, and events; read pod logs; describe pods/nodes; pod/node top (metrics); read-only diagnostics / RCA (pod-h","codeSnippets":[],"executableExamples":[{"language":"bash","snippet":"uv tool install k8s-aiops\nk8s-aiops init            # friendly wizard: register your kube contexts as targets\nk8s-aiops doctor          # or skip init — works with your current kube-context too"},{"language":"bash","snippet":"openclaw plugins install clawhub:@zw008/k8s-aiops\nopenclaw skills info k8s-aiops          # expect: Visible to model: yes"},{"language":"bash","snippet":"k8s-aiops init                                            # interactive onboarding wizard\nk8s-aiops pod list [-n <ns>] [-t <target>]\nk8s-aiops pod get <name> [-n <ns>]\nk8s-aiops pod describe <name> [-n <ns>]                   # status, container states, events\nk8s-aiops pod logs <name> [-n <ns>] [--tail 200] [-c <container>]\nk8s-aiops pod delete <name> [-n <ns>] [--dry-run]        # double confirm\nk8s-aiops deployment list|get|scale|restart|delete ...    # scale/restart: single confirm + --dry-run; delete: double confirm\nk8s-aiops rollout status|history|pause|resume <name> [-n <ns>]\nk8s-aiops rollout set-image <name> <container> <image> [-n <ns>]\nk8s-aiops rollout undo <name> [--to-revision N] [--dry-run]   # double confirm\nk8s-aiops statefulset list|get|scale ...\nk8s-aiops daemonset list|get ...\nk8s-aiops job list|get|delete ...                         # delete: double confirm\nk8s-aiops cronjob list|get ...\nk8s-aiops service list [-n <ns>]\nk8s-aiops ingress list|get [-n <ns>]\nk8s-aiops configmap list|get [-n <ns>]\nk8s-aiops secret list [-n <ns>]                           # names/keys only — never values\nk8s-aiops storage pvc-list|pvc-get|pv-list|class-list\nk8s-aiops top pod|node                                    # requires metrics-server\nk8s-aiops node list|describe\nk8s-aiops node cordon|drain <name> [--dry-run]           # double confirm\nk8s-aiops node uncordon <name>\nk8s-aiops namespace list|create\nk8s-aiops namespace delete <name> [--dry-run]            # double confirm\nk8s-aiops cluster-info\nk8s-aiops api-resources\nk8s-aiops events [-n <ns>]\nk8s-aiops diagnose pod-health [-n <ns>] [-l <selector>]         # read-only RCA: crashloop/imagepull/OOM/unschedulable/restarts\nk8s-aiops diagnose workload-readiness [-n <ns>]                 # read-only RCA: ready<desired / stuck rollouts\nk8s-aiops doctor\nk8s-aiops mcp                                             # start MCP server (stdio)"},{"language":"text","snippet":"You operate a Kubernetes cluster through the k8s-aiops MCP tools.\n\nTOOL USE\n- Before answering any question about the current cluster, you MUST call a\n  tool. Never answer from memory or assumption.\n- Actually invoke the tool. Do not describe the call you would make, and do not\n  emit an example JSON response in place of calling it.\n- If a tool call fails, report the real error verbatim. Never fill the gap with\n  a plausible-sounding answer.\n\nREADING RESULTS\n- Read the whole result before concluding. If a result contains a \"truncated\"\n  field that is true, say so and re-run with a higher limit instead of treating\n  the partial result as complete.\n- A null field means the apiserver did not return that value. Report it as \"not\n  available\" — never infer it. In particular, a null \"node\" means the pod is\n  not scheduled; it does not mean the node is unknown or missing.\n- Report values exactly as returned. Do not normalise, translate, or prettify\n  phases, conditions, container-state reasons, or resource names.\n- When an RCA result has findings, work in \"rank\" order and cite the measured\n  number in each finding's \"detail\".\n\nSCOPE AND IDENTIFIERS\n- Always state the namespace you are talking about. A bare pod or deployment\n  name is ambiguous — the same name exists in many namespaces.\n- Omitting the namespace means ALL namespaces, not the default one. Never\n  silently widen a namespaced question into a cluster-wide answer.\n- Do not confuse a namespace with a context (a cluster), a pod name with its\n  deployment name, or a container name with the pod that contains it. A pod\n  name generated by a ReplicaSet is not a stable identifier.\n- Separate observation from interpretation. State what the tools returned, then\n  any interpretation, clearly marked as such.\n- Do not assert a capacity, performance, or availability problem unless a tool\n  result supports it. Do not add generic advice that does not follow from the\n  tool output."},{"language":"bash","snippet":"# Point KUBECONFIG at a read-only ServiceAccount context, then:\nk8s-aiops doctor"},{"language":"bash","snippet":"export K8S_AUDIT_APPROVED_BY=\"your.name@example.com\"\nexport K8S_AUDIT_RATIONALE=\"incident INC-1234 — restart the stuck payments rollout\""}],"parameters":null,"dependencies":[],"permissions":[],"extractedFiles":[{"path":"SKILL.md","content":"---\nname: k8s-aiops\nslug: k8s-aiops\ndisplayName: \"k8s AIops\"\nsummary: \"Governed Kubernetes ops — 55 MCP tools with audit, budget, undo, risk-tier audit labels.\"\nlicense: MIT\nhomepage: https://github.com/AIops-tools/K8s-AIops\ntags: [aiops, mcp, governance, k8s]\ndescription: >\n  Use this skill whenever the user needs to operate a Kubernetes cluster — list/inspect pods, deployments, statefulsets, daemonsets, replicasets, jobs, cronjobs, services, ingresses, endpoints, configmaps, secrets (names/keys only), PVCs/PVs/storageclasses, nodes, namespaces, and events; read pod logs; describe pods/nodes; pod/node top (metrics); read-only diagnostics / RCA (pod-health, workload-readiness); scale deployments/statefulsets; rollout status/history/undo/pause/resume and set image; delete pods/deployments/jobs; create/delete namespaces; and cordon/uncordon/drain nodes. Works with any kubeconfig-reachable cluster (standard Kubernetes, k3s, EKS, GKE, AKS).\n  Always use this skill for \"list k8s pods\", \"scale deployment\", \"kubernetes pod logs\", \"describe pod\", \"why is my pod crashing\", \"diagnose pods\", \"which deployments are unhealthy\", \"rollout undo\", \"set image\", \"top pods\", \"drain node\", \"cordon node\", \"restart deployment\", \"k3s\", or \"kubectl\"-style tasks when the context is explicitly Kubernetes / a cluster.\n  Do NOT use when the target is not a Kubernetes cluster (hypervisor VM lifecycle, backup products, or cloud-provider consoles are out of scope).\n  Common Kubernetes operations with a built-in governance harness (audit, token budget, undo, risk-tier labels).\ninstaller:\n  kind: uv\n  package: k8s-aiops\nargument-hint: \"[resource name or describe your Kubernetes task]\"\nallowed-tools:\n  - Bash\nmetadata: {\"openclaw\":{\"requires\":{\"anyBins\":[\"k8s-aiops\",\"uvx\"]},\"optional\":{\"env\":[\"K8S_AIOPS_CONFIG\",\"KUBECONFIG\",\"K8S_AIOPS_HOME\"]},\"homepage\":\"https://github.com/AIops-tools/K8s-AIops\",\"emoji\":\"☸️\",\"os\":[\"macos\",\"linux\"]}}\ncompatibility: >\n  Standalone, self-governed Kubernetes operations. The governance harness (audit, token/runaway budget, undo, risk-tier labels) is bundled in the package — no external skill-family dependency.\n  All write operations are audited to a local SQLite DB under ~/.k8s-aiops/ (relocatable via K8S_AIOPS_HOME).\n  Credentials: k8s-aiops handles NO credentials directly — authentication is delegated to the kubeconfig (KUBECONFIG env or ~/.kube/config), which may hold client certs, bearer tokens, or exec plugins (EKS/GKE/AKS). Underlying credentials are never read, logged, or echoed. The state dir ~/.k8s-aiops should be chmod 700.\n  Destructive operations (deployment/job/namespace delete, node cordon/drain, rollout undo) require double confirmation at the CLI layer and support --dry-run. All write tools pass through the @governed_tool decorator (budget/runaway guard + audit + a descriptive risk-tier label). Reversible writes record an inverse undo descriptor (scale_deployment/scale_statefulset restore the previous replica count; set_deployment_image"},{"path":"_meta.json","content":"{\n  \"ownerId\": \"kn7b067awq2s97bn3d7p5qfhw5827pxc\",\n  \"slug\": \"k8s-aiops\",\n  \"version\": \"0.13.3\",\n  \"publishedAt\": 1789452174189\n}"},{"path":"references/agent-guardrails.md","content":"# Agent guardrails — running k8s-aiops with a smaller / local model\n\nIf you drive these tools with a local model (Llama, Qwen, Mistral … via Goose,\nOllama, LM Studio, or any OpenAI-compatible runtime), you will get noticeably\nbetter results with a short system prompt. This page gives you one, and — more\nimportantly — tells you which guardrails you **no longer need to write**, because\nthe tool now enforces them itself.\n\nThe distinction matters. A guardrail in a prompt is a request. A guardrail in the\nharness is a guarantee. Anything below that we could move into the harness, we did.\n\n## Authorization is not this tool's job — decide it where it belongs\n\nWhether a write should happen is your decision, or the account's. The tool does\nnot gate it — there is no read-only switch and no approval prompt to configure.\nThe two right places to control read vs write:\n\n- **The kubeconfig context you connect with.** Bind it to a ServiceAccount or\n  user whose RBAC grants only `get`/`list`/`watch`, and every write fails at the\n  apiserver — the only place the permission actually lives. No skill-side flag\n  can be argued around by a model, but a revoked RBAC verb cannot be. This is\n  strictly stronger than any in-process switch: it is enforced at the cluster.\n- **Your agent's system prompt.** If you want an observe-only session, tell the\n  model not to call the write tools (they are clearly tagged `[WRITE]`).\n\nWhat the tool *does* guarantee is that you can always see what happened:\n\n## What the tool enforces — do not waste prompt budget on these\n\n| You might be tempted to prompt | Why you don't need to |\n|---|---|\n| \"Log everything you do, over both MCP and the CLI\" | Every operation is audited to `~/.k8s-aiops/audit.db` regardless of what the model says it did — and the CLI writes the same row the MCP path does, so there is no unaudited entry point. Reversible writes also record an undo token capturing the *prior* state. |\n| \"Don't invent a value when a field is missing\" | A field the apiserver did not return comes back as `null`, never as `\"\"`. An unscheduled pod's `node` is `null`; a pod with no phase yet has `phase: null`; an object with no readable creation timestamp has `age: null`. Absent and empty are distinguishable in the payload. |\n| \"Tell me if the output was cut off\" | The limit-bearing reads return an envelope: `event_list` → `{\"events\": [...], \"returned\": N, \"limit\": L, \"truncated\": true/false}`, and `undo_list` → `{\"undos\": [...], \"returned\": N, \"limit\": L, \"truncated\": ...}`. Truncation is **measured** (one extra row is fetched), not guessed from `len(rows) == limit`. |\n| \"Preserve the ordering / tell me what's most urgent\" | `pod_health_rca` and `workload_readiness_rca` findings carry an explicit 1-based `rank`, worst-first, and each finding's `detail` cites the measured signal (the waiting reason, the restart count, the ready/desired ratio). Priority is in the payload, not implied by list position. |\n| \"Confirm before anything destructive\" | Th"},{"path":"references/capabilities.md","content":"# k8s-aiops Capabilities\n\n55 MCP tools (39 read / 16 write). Every tool is wrapped with `@governed_tool`\n(audit + policy + budget + risk-tier; undo where a clean inverse exists). Returns\nare high-signal summaries — `_get` / `_describe` tools add detail for a single object.\n\n## Read tools\n\n| Tool | Returns | Risk |\n|------|---------|:----:|\n| `pod_list` | name, namespace, phase, ready, restarts, node, age | low |\n| `pod_get` | + host_ip, pod_ip, containers | low |\n| `pod_describe` | status, conditions, container states + restart counts, recent events | low |\n| `pod_logs` | trailing log lines (default 100) | low |\n| `deployment_list` / `deployment_get` | replicas summary / + strategy, images | low |\n| `rollout_status` | desired/updated/available/unavailable + paused | low |\n| `rollout_history` | revisions (from replicasets) with images | low |\n| `statefulset_list` / `statefulset_get` | desired/ready/current / + service, images | low |\n| `daemonset_list` / `daemonset_get` | desired/ready/available / + images | low |\n| `replicaset_list` | name, namespace, desired/ready, age | low |\n| `job_list` / `job_get` | completions, succeeded/failed/active | low |\n| `cronjob_list` / `cronjob_get` | schedule, suspend, active, last schedule | low |\n| `service_list` | name, namespace, type, cluster IP, ports | low |\n| `ingress_list` / `ingress_get` | class, hosts / + path→backend rules | low |\n| `endpoints_list` | ready addresses + ports | low |\n| `configmap_list` / `configmap_get` | key count / keys + values | low |\n| `secret_list` | names, types, key NAMES only (values redacted) | low |\n| `pvc_list` / `pvc_get` | status, capacity, class / + access modes | low |\n| `pv_list` | capacity, status, claim, class | low |\n| `storageclass_list` | provisioner, reclaim policy, default | low |\n| `node_list` / `node_describe` | status, roles / capacity, allocatable, conditions, taints | low |\n| `namespace_list` | name, phase, age | low |\n| `pod_top` / `node_top` | CPU/mem via metrics-server (graceful if absent) | low |\n| `cluster_info` | server version, node/ready/namespace counts | low |\n| `api_resources` | available API groups + versions | low |\n| `event_list` | type, reason, object, namespace, message, age | low |\n| `pod_health_rca` | worst-first findings: CrashLoopBackOff, image-pull, OOMKilled, unschedulable, high restarts (each cites the reason/count) | low |\n| `workload_readiness_rca` | worst-first findings: ready<desired, zero-ready outages, stuck rollouts (Deployment/StatefulSet/DaemonSet) | low |\n| `undo_list` | recorded reversible writes / not-yet-applied undo tokens | low |\n\n## Write tools\n\n| Tool | Effect | Risk | Undo |\n|------|--------|:----:|------|\n| `scale_deployment` | set replica count | medium | scale back to `previous_replicas` |\n| `scale_statefulset` | set replica count | medium | scale back to `previous_replicas` |\n| `rollout_restart_deployment` | patch `restartedAt` annotation | medium | none (pods already rolling) |\n| `rollout_pause` / `rollout_resume"},{"path":"references/cli-reference.md","content":"# k8s-aiops CLI Reference\n\nAll commands accept `-t/--target <name>` to select a configured target (a kube\ncontext). Namespaced commands accept `-n/--namespace <ns>`; omit it to use the\ntarget's default namespace (read lists fall back to all-namespaces).\n\n## Onboarding\n\n```bash\nk8s-aiops init                    # interactive wizard: register kube contexts as targets\n```\n\n## Pods\n\n```bash\nk8s-aiops pod list [-n <ns>] [-t <target>]\nk8s-aiops pod get <name> [-n <ns>]\nk8s-aiops pod describe <name> [-n <ns>]              # status, container states, events\nk8s-aiops pod logs <name> [-n <ns>] [--tail N] [-c <container>]\nk8s-aiops pod delete <name> [-n <ns>] [--dry-run]    # destructive: double confirm\n```\n\n## Deployments & Rollouts\n\n```bash\nk8s-aiops deployment list [-n <ns>]\nk8s-aiops deployment get <name> [-n <ns>]\nk8s-aiops deployment scale <name> <replicas> [-n <ns>]\nk8s-aiops deployment restart <name> [-n <ns>]        # rolling restart\nk8s-aiops deployment delete <name> [-n <ns>] [--dry-run]   # HIGH RISK: double confirm\nk8s-aiops rollout status <name> [-n <ns>]\nk8s-aiops rollout history <name> [-n <ns>]\nk8s-aiops rollout pause|resume <name> [-n <ns>]\nk8s-aiops rollout set-image <name> <container> <image> [-n <ns>]\nk8s-aiops rollout undo <name> [--to-revision N] [--dry-run]   # HIGH RISK: double confirm\n```\n\n## StatefulSets / DaemonSets / Jobs / CronJobs\n\n```bash\nk8s-aiops statefulset list|get [-n <ns>]\nk8s-aiops statefulset scale <name> <replicas> [-n <ns>]\nk8s-aiops daemonset list|get [-n <ns>]\nk8s-aiops job list|get [-n <ns>]\nk8s-aiops job delete <name> [-n <ns>] [--dry-run]    # destructive: double confirm\nk8s-aiops cronjob list|get [-n <ns>]\n```\n\n## Services, Ingress, Config, Storage\n\n```bash\nk8s-aiops service list [-n <ns>]\nk8s-aiops ingress list|get [-n <ns>]\nk8s-aiops configmap list|get [-n <ns>]\nk8s-aiops secret list [-n <ns>]                      # names + key NAMES only, never values\nk8s-aiops storage pvc-list|pvc-get [-n <ns>]\nk8s-aiops storage pv-list|class-list\n```\n\n## Nodes & Metrics\n\n```bash\nk8s-aiops node list\nk8s-aiops node describe <name>\nk8s-aiops node cordon <name> [--dry-run]             # destructive: double confirm\nk8s-aiops node uncordon <name>\nk8s-aiops node drain <name> [--dry-run]              # HIGH RISK: double confirm\nk8s-aiops top pod|node                               # requires metrics-server\n```\n\n## Namespaces, Cluster & Events\n\n```bash\nk8s-aiops namespace list\nk8s-aiops namespace create <name>\nk8s-aiops namespace delete <name> [--dry-run]        # HIGH RISK: double confirm\nk8s-aiops cluster-info\nk8s-aiops api-resources\nk8s-aiops events [-n <ns>]\n```\n\n## Diagnostics & MCP\n\n```bash\nk8s-aiops diagnose pod-health [-n <ns>] [-l <selector>]  # RCA: crashloop/imagepull/OOM/unschedulable/restarts (read-only)\nk8s-aiops diagnose workload-readiness [-n <ns>]          # RCA: ready<desired / stuck rollouts (read-only)\nk8s-aiops doctor [--skip-auth]    # check config + cluster reachability\nk8s-aiops mcp                     # st"}],"languages":[],"docsSourceLabel":"CLAWHUB","editorialOverview":null,"editorialQuality":{"score":100,"threshold":65,"status":"thin","wordCount":2134,"uniquenessScore":41,"reasons":["uniqueness-below-45"]}},"media":{"evidence":{"source":"no-media","verified":false,"confidence":"low","updatedAt":"2026-10-10T06:28:09.610Z","emptyReason":"No screenshots, media assets, or demo links are available."},"primaryImageUrl":null,"mediaAssetCount":0,"assets":[],"demoUrl":null},"ownerResources":{"evidence":{"source":"unclaimed","verified":false,"confidence":"low","updatedAt":"2026-10-10T06:28:09.610Z","emptyReason":"This page has not been claimed by the agent owner."},"hasCustomPage":false,"customPageUpdatedAt":null,"customLinks":[],"structuredLinks":{"docsUrl":null,"demoUrl":null,"supportUrl":null,"pricingUrl":null,"statusUrl":null},"customPage":null},"relatedAgents":{"evidence":{"source":"protocol-neighbors","verified":false,"confidence":"medium","updatedAt":"2026-10-10T10:53:18.745Z","emptyReason":null},"items":[{"id":"8ebccd8e-3863-4187-8355-c3f14e1f9edf","entityType":"agent","canonicalPath":"/agent/iofficeai-aionui","slug":"iofficeai-aionui","name":"AionUi","description":"Free, local, open-source 24/7 Cowork app and OpenClaw for Gemini CLI, Claude Code, Codex, OpenCode, Qwen Code, Goose CLI, Auggie, and more | 🌟 Star if you like it!","url":"https://github.com/iOfficeAI/AionUi","homepage":"https://www.aionui.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-10-09T19:11:12.944Z","createdAt":"2026-02-25T03:38:16.584Z","downloads":null},{"id":"b917f68a-ebff-438e-84f8-3f4b2494c0bc","entityType":"agent","canonicalPath":"/agent/activepieces-activepieces","slug":"activepieces-activepieces","name":"activepieces","description":"AI Agents & MCPs & AI Workflow Automation • (~400 MCP servers for AI agents) • AI Automation / AI Agent with MCPs • AI Workflows & AI Agents • MCPs for AI Agents","url":"https://github.com/activepieces/activepieces","homepage":"https://www.activepieces.com","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-15T02:22:12.426Z","createdAt":"2026-02-25T03:38:12.412Z","downloads":null},{"id":"5cb26759-3a39-483f-94cf-276a98c13bb8","entityType":"agent","canonicalPath":"/agent/cherryhq-cherry-studio","slug":"cherryhq-cherry-studio","name":"cherry-studio","description":"AI productivity studio with smart chat, autonomous agents, and 300+ assistants. Unified access to frontier LLMs","url":"https://github.com/CherryHQ/cherry-studio","homepage":"https://cherry-ai.com","source":"GITHUB_REPOS","protocols":["MCP","OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-04-11T14:38:40.986Z","createdAt":"2026-02-25T03:38:19.379Z","downloads":null},{"id":"6f6582d0-5d76-4f0f-b81d-86520247950b","entityType":"agent","canonicalPath":"/agent/copilotkit-copilotkit","slug":"copilotkit-copilotkit","name":"CopilotKit","description":"The Frontend for Agents & Generative UI. React + Angular","url":"https://github.com/CopilotKit/CopilotKit","homepage":"https://docs.copilotkit.ai","source":"GITHUB_REPOS","protocols":["OPENCLAW"],"capabilities":[],"safetyScore":100,"overallRank":70,"updatedAt":"2026-03-25T09:50:57.846Z","createdAt":"2026-02-25T03:39:14.617Z","downloads":null}],"links":{"hub":"/agent","source":"/agent/source/clawhub","protocols":[{"label":"OpenClaw","href":"/agent/protocol/openclew"}]}}}